Festival Season Offer15% off on all our programmes — claim it before you enrol
MODULE 3 OF 8  ·  20 Hrs  ·  2 Weeks

Data Wrangling (SQL + Cleaning)

Real data is messy. This module builds the SQL fluency and data-cleaning discipline to extract, join and prepare raw data into a reliable, analysis-ready form — the unglamorous work every data science project depends on.

Who This Module Is For
Students who have completed the Python module and are ready to work with real, messy datasets.
Real-World Relevance
SQL is one of the most consistently tested skills in data analyst and data scientist interviews, and clean, analysis-ready data is the foundation every downstream model depends on.
Program OverviewView Hands-On Labs
Curriculum

What You Will Learn

A detailed, industry-aligned breakdown of every topic covered in this module.

  • Relational databases & SQL fundamentals
  • SELECT, JOIN, GROUP BY & subqueries
  • Handling missing values & duplicates
  • Outlier detection & data standardization
  • Merging, reshaping & pivoting datasets
  • Feature engineering basics
  • Working with APIs & scraped data
  • Building clean, analysis-ready datasets
Technology Stack

Tools You Will Use

Hands-on time with the same tools used by working data analysts and data scientists today.

SQL

Query language used to extract, filter and aggregate data from relational databases.

MySQL / PostgreSQL

Relational database systems used to store and query structured data.

Pandas

Data manipulation library used to clean, transform and analyse structured datasets.

OpenRefine

Data-cleaning tool used to explore, standardise and transform messy datasets.

Practical Work

Hands-On Labs

Production-style data science lab scenarios, built using real, messy datasets.

01

Write SQL queries using SELECT, JOIN, GROUP BY and subqueries against a real database.

02

Clean a messy dataset — handling missing values, duplicates and outliers.

03

Merge, reshape and pivot multiple datasets into a single analysis-ready table.

04

Pull data from a public API and merge it with an existing dataset.

05

Engineer basic features from raw columns for downstream analysis.

Evaluation

Assessment

Knowledge Assessment

Quiz covering SQL joins, subqueries and data-cleaning techniques.

Practical Evaluation

Students must extract, clean and merge data from multiple sources into one analysis-ready dataset.

Portfolio

Projects

Industry-style deliverables added directly to your project portfolio.

Portfolio Project 01

Data Cleaning Lab

Clean and prepare a messy real-world dataset using SQL and Pandas.

Module Outcome

What This Module Builds

Students learn to extract, clean and prepare raw data into a reliable, analysis-ready form — the foundation of every downstream data science task.

Maps to job roles
Data AnalystJunior Data EngineerData Wrangling SpecialistETL Analyst (Trainee)

Continue building your data science portfolio

Next up: Module 4 — Exploratory Data Analysis (EDA)

Go to Module 4Full Roadmap