All programsData Science & Machine Learning

Real data is messy. Learn to work with it anyway.

Every other program on this page is about language models. This one is about the work underneath them: finding data, cleaning data that arrives broken, understanding what it actually says, training a model on it, and being honest about how well that model works. It is the oldest skill in AI and still the most employable one, because the hard part was never the algorithm — it was the four days before the algorithm. Ten weeks, in the evenings, built around projects rather than lectures. You will finish with work you can show someone.

Duration
10 weeks · around 50 hours
Mode
Live online — real sessions, not recordings
Timing
Evenings and weekends, designed around classes or a job
Cohort size
25–30
You need
A laptop. No GPU required.
Recordings
Every session recorded, yours to keep

Who this is for

  • Students and final-year candidates who want something real on a CV, not coursework
  • Anyone starting out who keeps being told to "learn data science" and cannot find where to start
  • Working professionals — analysts, operations, finance, marketing — who already live in spreadsheets
  • Software engineers who can code but have never trained or evaluated a model
  • Career-switchers who need one provable, portfolio-backed base to move on

Anyone who only wants to use AI assistants rather than build with data — Claude at Work needs no code at all. If you already train models and want GPUs, fine-tuning and serving, take Applied AI Engineering instead.

Prerequisites

  • Basic Python — variables, loops, functions, and installing a package
  • School-level maths. Nothing beyond it, and nothing assumed from a degree
  • A laptop. No GPU, no paid cloud account, no server

If you have never written Python at all, week 1 is built for you — it starts from installing it. What we cannot compress is the habit of reading an error message and trying again, so bring that. Everything we use is free and open source.

Modules

10 modules · register to see what is inside each one

Module 1Week 1

Python for data, from zero

Covered in detail after registration

Module 2Week 2

Getting the data out of wherever it lives

Covered in detail after registration

Module 3Week 3

Cleaning — the part nobody puts in the tutorial

Covered in detail after registration

Module 4Week 4

Looking at data honestly

Covered in detail after registration

Module 5Week 5

The statistics you will actually use

Covered in detail after registration

Module 6Week 6

Your first models: regression and classification

Covered in detail after registration

Module 7Week 7

Evaluation, and the ways a model lies to you

Covered in detail after registration

Module 8Week 8

Models that win: ensembles and feature engineering

Covered in detail after registration

Module 9Week 9

When there are no labels: clustering and time series

Covered in detail after registration

Module 10Week 10

Ship it, then defend it

Covered in detail after registration

Every module runs as a working session — you write code in it. Modules 6 onward run against your own capstone dataset wherever possible, so the project grows through the course rather than being built in a panic at the end.

What you walk away with

  • A deployed, reviewed data project, assessed against a published rubric
  • Five smaller projects built along the way, each one yours to show
  • Working Python, pandas, SQL and scikit-learn — used, not merely watched
  • The ability to clean a dataset nobody has cleaned before and document what you did
  • Honest evaluation: leakage, the right metric, and what a score does not prove
  • Every session recording, all notebooks, and your starter repositories
  • Codnov Certificate of Completion, with a unique ID and a public verification link

Sample projects

  • A price predictor — scrape or pull a real listings dataset, clean it, and predict a price with a stated error range
  • A churn or dropout model — predict who leaves, on imbalanced data, with the metric that matches the decision
  • A customer segmentation — cluster real transaction data and defend the segments to a business reader
  • A document-to-data pipeline — unstructured records in, validated structured rows out
  • A demand forecast — a time series with trend and seasonality, forecast with an honest error band
  • A public-data investigation — take an open government or health dataset and answer a question nobody has answered with it
  • A recommendation engine — build one, then measure whether it beats recommending the most popular item
  • A fraud or anomaly detector — rare events, where accuracy is the wrong metric entirely

Bring your own dataset or take one of these. Each comes with a prepared dataset and a starter repository — a project that is too ambitious is the most common reason people do not finish.

The guarantee

You finish with a deployed, reviewed AI project, or we re-run you free.

Requires minimum 75% attendance and on-time milestone submission. Full conditions at codnov.ai/training/terms.

Taught by Codnov engineers who build and run AI systems in production every day.

Price

Cohort pricing is confirmed per intake — enquire for the current fee and dates. Student and college-batch rates are quoted separately. All prices exclusive of 18% GST.

Nothing is paid on this website. Register your interest and we will confirm your seat with you directly.

Register your interest

We will come back to you with the full syllabus, the schedule, and whether there is a seat in the next cohort.

No payment is taken on this website. Registering interest does not enrol you or commit you to anything.

Codnov Certificate of Completion. Not a degree or diploma recognised under the UGC Act, 1956 or the AICTE Act, 1987. All prices exclusive of 18% GST. Project completion guarantee subject to minimum 75% attendance and on-time milestone submission; full terms at codnov.ai/training/terms. Open to participants aged 16 and above.