Data Wrangling (SQL + Cleaning)
Real data is messy. This module builds the SQL fluency and data-cleaning discipline to extract, join and prepare raw data into a reliable, analysis-ready form — the unglamorous work every data science project depends on.
What You Will Learn
A detailed, industry-aligned breakdown of every topic covered in this module.
- Relational databases & SQL fundamentals
- SELECT, JOIN, GROUP BY & subqueries
- Handling missing values & duplicates
- Outlier detection & data standardization
- Merging, reshaping & pivoting datasets
- Feature engineering basics
- Working with APIs & scraped data
- Building clean, analysis-ready datasets
Tools You Will Use
Hands-on time with the same tools used by working data analysts and data scientists today.
SQL
Query language used to extract, filter and aggregate data from relational databases.
MySQL / PostgreSQL
Relational database systems used to store and query structured data.
Pandas
Data manipulation library used to clean, transform and analyse structured datasets.
OpenRefine
Data-cleaning tool used to explore, standardise and transform messy datasets.
Hands-On Labs
Production-style data science lab scenarios, built using real, messy datasets.
Write SQL queries using SELECT, JOIN, GROUP BY and subqueries against a real database.
Clean a messy dataset — handling missing values, duplicates and outliers.
Merge, reshape and pivot multiple datasets into a single analysis-ready table.
Pull data from a public API and merge it with an existing dataset.
Engineer basic features from raw columns for downstream analysis.
Assessment
Knowledge Assessment
Quiz covering SQL joins, subqueries and data-cleaning techniques.
Practical Evaluation
Students must extract, clean and merge data from multiple sources into one analysis-ready dataset.
Projects
Industry-style deliverables added directly to your project portfolio.
Data Cleaning Lab
Clean and prepare a messy real-world dataset using SQL and Pandas.
What This Module Builds
Students learn to extract, clean and prepare raw data into a reliable, analysis-ready form — the foundation of every downstream data science task.
