Data Cleaning, Feature Engineering & Model Preparation
Step-by-Step Data Preprocessing Approach for Academic Research & Real-World Applications
Built for faculty members and research scholars from every discipline.
Registration is open 24×7. Register any time, from anywhere.
survey_data.csv 4 problems found
| ID | Age | Income | City |
|---|---|---|---|
| 101 | NaN | 42,000 | Pune |
| 102 | 29 | 9,90,000 | pune |
| 103 | 41 | 51,500 | Hyd. |
| 103 | 41 | 51,500 | Hyderabad |
| 104 | 37 | 47,200 | Kochi |
Missing value · outlier · messy label · duplicate row
Built for faculty members and research scholars from every discipline
Science, commerce, humanities, engineering, health, education: bring your own kind of data and leave with a clean, model-ready workflow.
Why this FDP matters for you
Most of a research project is preparing the data, not running the model. Weak preparation quietly weakens every result that follows, and reviewers notice.
Stronger publications
Clean, well-documented data is what lets your paper survive peer review and replication.
More reliable results
Missing values, outliers and duplicates can flip a conclusion. You will learn to find and fix them.
Works in every discipline
Commerce, science, humanities, health, engineering: the same workflow applies to your data.
Guide scholars better
Supervise students with confidence and teach data skills that employers and funders ask for.
Reproducible research
Learn documentation and validation habits so anyone can rerun and trust your work.
Ready for machine learning
Go from a raw spreadsheet to model-ready data with a clear, repeatable pipeline.
No technical skill required
Start from zero. Every concept is explained in plain language, then shown step by step in R and Python. If you can open a spreadsheet, you can follow along.
What you will learn, day by day
Each session builds on the last, ending with a full hands-on project.
Data understanding & cleaning
- Data collection and sources
- Exploratory Data Analysis (EDA)
- Missing value treatment
- Outlier detection and handling
- Duplicate identification and removal
Transformation & feature engineering
- Data transformation techniques
- Scaling and normalization
- Categorical data encoding
- Feature engineering
- Feature selection and dimensionality reduction
Model-ready data & best practices
- Training, validation and testing data
- Data splitting strategies
- Data validation and quality checks
- Preprocessing workflow for ML
- Documentation and reproducibility
- Hands-on end-to-end preprocessing
Live demonstrations in R Python
Check the session time in your country
Sessions run 7:00 PM to 8:30 PM IST on all three days. Pick your country or region to see your local time.
Programme details
Registration fee
- Certificate of participation
- Recordings of all sessions
Dr. A. Rajini
Assistant Professor, Department of Mathematics and Statistics
Bhavan’s Vivekananda College of Science, Humanities and Commerce, Sainikpuri, Secunderabad, Telangana, India
Reserve your seat for 27 October
Certificate and session recordings included. Choose the option that matches your country.
