All programsIntroduction to AI Supercomputing

What the hardware under all of this actually is.

Everyone talks about models. Almost nobody can explain the machine underneath, and it shows the first time a job runs out of memory at three in the morning. This program opens the box: what a GPU is doing that a CPU cannot, how an AI supercomputer is actually put together, how work gets scheduled onto one, and where the money goes. Enough to make real decisions rather than repeat headlines.

Duration
4 weeks · around 16 hours
Mode
Live online, with lab sessions
Timing
Evenings, designed around a full-time job
Cohort size
15–20
Lab access
Real accelerator hardware; capacity confirmed per cohort
Recordings
Every session recorded, yours to keep

Who this is for

  • Engineers who use GPUs through an API and have never seen one up close
  • Students heading into ML, HPC or systems work
  • Technical leads deciding between cloud GPUs and hardware of their own
  • Anyone who has been handed a cluster and told to make it useful

Anyone looking for model theory — this is the hardware and the systems around it. Applied AI Engineering is where models get trained.

Prerequisites

  • Comfort on a Linux command line, and any one programming language.

No CUDA and no hardware background assumed. You will read and run code rather than write kernels from scratch.

Modules

Module 1Week 1

Why a GPU, and not a faster CPU

Thousands of small cores versus a few fast ones, and the kind of problem that suits each. Memory bandwidth as the thing that actually limits you. What VRAM is, why a model “does not fit”, and how to read the specification sheet of an accelerator without being sold to.

IN THE SESSION

Read a real accelerator spec sheet and predict whether a given model will fit and run — then load it and find out whether you were right.

You leave with: The ability to look at a model and a card and say whether it will run.

Module 2Week 2

Anatomy of an AI supercomputer

A single box, a node, a rack, a cluster. The interconnect and why it decides everything once you pass one machine. Storage that can keep thousands of cores fed. Power and cooling as real constraints rather than footnotes — and an honest look at what a serious machine costs to buy and to run.

IN THE SESSION

Map a workload onto hardware on paper: how many cards, how much memory, where the bottleneck lands, what it costs per hour.

You leave with: A mental map from one chip to a full cluster, with the bottlenecks marked.

Module 3Week 3

Getting your job onto the machine

Shared hardware means queues: schedulers, jobs, allocations, and why your work waited four hours. Containers and environments that reproduce. Running the same job on one GPU, then several — data parallel versus model parallel, in plain language and then in practice.

IN THE SESSION

Submit a job to a scheduler, watch it queue, run it in a container, then run the same job across several GPUs and compare.

You leave with: A job you submitted, queued, ran and collected results from yourself.

Module 4Week 4

Fast enough, and cheap enough

Profiling: finding out whether you are actually compute-bound, memory-bound or just waiting on disk. Mixed precision and quantisation, and what each costs you in quality. Then the decision most teams get wrong — rent or own — worked through with real numbers at your volume.

IN THE SESSION

Profile a real job, find out whether you are compute-bound, memory-bound or waiting on disk, and make one measurable improvement.

You leave with: A measured speed-up on a real job, and a defensible rent-versus-own answer.

What you walk away with

  • A working understanding of the hardware everything else sits on
  • Jobs you have scheduled, run and profiled yourself
  • A rent-versus-own calculation you can defend to a finance team
  • Vocabulary that makes vendor conversations survivable
  • Codnov Certificate of Completion, with a unique ID and a public verification link

Sample projects

  • Benchmark one model across precisions and report where the time goes
  • Take a single-GPU training job and run it across several
  • Profile an inference service and cut its cost per request
  • Size and cost a cluster for a stated workload, and justify every line
The guarantee

You finish with a deployed, reviewed AI project, or we re-run you free.

Requires minimum 75% attendance and on-time milestone submission. Full conditions at codnov.ai/training/terms.

Taught by Codnov engineers who build and run AI systems in production every day.

Price

Cohort pricing is confirmed per intake — enquire for the current fee, dates and the lab capacity available to that cohort. All prices exclusive of 18% GST.

Nothing is paid on this website. Register your interest and we will confirm your seat with you directly.

Register your interest

We will come back to you with the full syllabus, the schedule, and whether there is a seat in the next cohort.

No payment is taken on this website. Registering interest does not enrol you or commit you to anything.

Codnov Certificate of Completion. Not a degree or diploma recognised under the UGC Act, 1956 or the AICTE Act, 1987. All prices exclusive of 18% GST. Project completion guarantee subject to minimum 75% attendance and on-time milestone submission; full terms at codnov.ai/training/terms. Open to participants aged 16 and above.