Festival Season Offer15% off on all our programmes — claim it before you enrol
← All Career InsightsData Science

What is feature engineering and how do I do it?

Feature engineering is turning raw data into inputs a machine learning model can learn from. It covers filling gaps, converting categories into numbers, scaling values, and creating new columns such as a customer's days since last purchase. Good features often improve a model more than switching algorithms, as long as every transformation is learned from training data only.

Why features matter as much as the model

A model only sees the columns you give it. If a churn dataset holds raw order dates, the model cannot easily tell that a customer who has not ordered for ninety days is at risk. A column such as days since last order states that pattern directly, so even a simple model can use it.

Feature engineering is where your understanding of the business problem enters the model. It is also where many projects quietly go wrong, by using information that would not be available at prediction time.

Common feature engineering techniques

Most tabular projects use a mix of these:

  • Handling missing values: fill with a median or a most frequent value, and add a flag column when the fact that a value is missing is itself informative.
  • Encoding categories: one-hot encoding for a small number of categories, and ordinal encoding only when the order means something, such as low, medium and high.
  • Scaling numbers: standardising or normalising features for models that are sensitive to scale, such as logistic regression or k-nearest neighbours.
  • Date and time features: day of week, month, hour, or time since an event, extracted from a raw timestamp.
  • Aggregations: counts, totals and averages per customer or per product, such as number of orders in the last thirty days.
  • Ratios and interactions: values that combine two columns, such as loan amount divided by income.
  • Text features: simple word counts or TF-IDF to turn short text into numbers.

A method that keeps your features honest

Use a customer churn dataset as the example as you work through these steps.

  1. Split the data first

    Separate training and test data before any transformation. For time-based problems, split by date so the test period comes after the training period.

  2. Explore with the target in mind

    Look at missing values, outliers and how each column relates to churn, using the training data only.

  3. Create candidate features

    Add features from business logic, such as days since last order, orders in the last thirty days and support tickets raised.

  4. Put transformations in a pipeline

    Use a pipeline, for example scikit-learn's Pipeline and ColumnTransformer, so imputation, encoding and scaling are fitted on training data and simply applied to new data.

  5. Check each feature is available at prediction time

    Ask whether you would really know this value at the moment the model is used. If not, remove it; it is leaking the answer.

  6. Compare with a baseline

    Train a simple model without the new features, then with them, using cross-validation, and keep only the features that improve results on validation data.

Quick answers about feature engineering

Short answers to what learners ask most.

What is a feature in machine learning?

A feature is one input column the model learns from, such as age, city or number of past orders. The value the model predicts, such as whether a customer churns, is called the target.

Is feature engineering still needed with deep learning?

Less for images, audio and long text, where deep networks learn their own features. For tabular business data it remains one of the most effective ways to improve a model.

What is the difference between feature engineering and feature selection?

Feature engineering creates or transforms columns. Feature selection chooses which of the available columns to keep, to reduce noise, overfitting and training time.

How do I know if a new feature helps?

Compare model performance with and without it on validation data or with cross-validation. If the score does not improve, or only improves on training data, leave it out.

Which tools are used for feature engineering?

Python with pandas for creating columns and scikit-learn for pipelines, encoders and scalers covers most projects. SQL is often used to build aggregate features from databases.

Continue with these data science guides

See where feature work sits in the Data Science syllabus, or read the guides that connect to it.

See the Data Science programmeSee the Exploratory Data Analysis moduleRead: detect and prevent data leakageRead: learn statistics and Python step by stepRead: a data science project from start to finishBrowse all Career Insights

Let the business question shape the columns

Start from what would really predict the outcome, build those features inside a pipeline, and keep only the ones that improve validation results. Clear, honest features make a model easier to trust and to explain.

Train for a Data Science role

The same programme, duration and fees, with the learning path built around one job role.

Data ScientistML EngineerData AnalystBI AnalystData EngineerAnalytics Consultant

Build real data science projects

Ask our team how the Data Science programme teaches modelling and feature work.

Our admissions team will call you back within 90 minutes.
AddressLR Towers, No. 3-535, 3rd Floor A Section, 100 Feet Road, Ayappa Society, Madhapur, Hyderabad, Telangana, India