MLOps, LLMOps & AI Platform Engineering
A model that only works in a notebook isn't an AI system — it's a demo. This module builds the operational discipline to run AI in production: tracking, versioning, testing, deploying and monitoring at scale.
What You Will Learn
A detailed, industry-aligned breakdown of every topic covered in this module.
- Experiment tracking and model registry
- Model and prompt versioning
- AI and RAG evaluation
- Regression testing for AI systems
- CI/CD for AI systems
- Tracing and observability
- Cloud IAM and secrets management
- Scaling and rollbacks
- Cost and latency management
- AI security and reliability
Tools You Will Use
Hands-on time with the same tools used in professional AI engineering and production ML workflows.
MLflow
Experiment tracking and model registry platform used to manage the ML model lifecycle.
Git
Version control system used to track code changes throughout every project.
GitHub Actions
CI/CD automation platform used to test, build and deploy code on every change.
Docker
Containerization platform used to package and deploy applications and AI services consistently.
Cloud Platform
Cloud infrastructure used to deploy, scale and monitor AI systems in production.
pgvector
PostgreSQL extension used to store and query vector embeddings for semantic search.
Monitoring Tools
Observability tooling used to track performance, cost and reliability of AI systems in production.
Hands-On Labs
Production-style AI engineering lab scenarios, built using the same stack real AI teams ship with.
Track experiments and register models and prompts using MLflow-based versioning.
Build a regression test suite for an AI/RAG system to catch quality drops before release.
Set up a CI/CD pipeline that tests and deploys an AI system automatically.
Add tracing and observability to monitor an AI system's behaviour in production.
Manage cloud IAM, secrets, scaling and rollback strategy for a deployed AI platform.
Assessment
Knowledge Assessment
Quiz covering experiment tracking, CI/CD for AI, observability and AI security fundamentals.
Practical Evaluation
Students must operationalize an AI system end to end — versioned, tested, deployed via CI/CD and monitored for cost, latency and reliability.
Projects
Industry-style deliverables added directly to your project portfolio.
AI Platform Operationalization
Apply MLOps/LLMOps practices — versioning, CI/CD, observability — to an existing AI service from the program.
What This Module Builds
Students learn to operationalize AI systems end-to-end — tracking experiments, versioning models and prompts, automating testing and deployment, and monitoring production AI for cost, latency, security and reliability.
