← Back to the course home

📐 The 26 lessons as diagrams

Every lesson's core idea as one numbered entity/sequence diagram — follow the circled numbers 1 → 2 → 3 in each box-and-arrow picture. For the full ELI5 story and hands-on commands, open the lesson itself.

1 🍱 Containers & images — recipe → lunchbox → lunch

One recipe (Dockerfile) bakes one frozen lunchbox (image); every opened box (container) is identical.

📝 Dockerfile the recipe card 🧊 Image school-api:v1 — frozen box 🗄️ Registry (ECR) the shelf where images wait 🏃 Running containers box 1 box 2 box 3 1 docker build 2 docker push 3 docker run ×3 — identical

Read full lesson 01 →

2 🪑 Pods — one desk, one address, replace don't repair

Kubernetes never runs a bare container — it always sits at a desk (Pod) with its own IP and shared shelf.

🌐 Cluster network callers 🪑 Pod — IP 10.0.4.7 📦 school-api main container 📦 sidecar optional helper 🗄️ shared volume 1 talks to the Pod's IP 2 🪑 NEW Pod new name, new IP 10.0.9.9 3 💥 desks are replaced, never repaired

Read full lesson 02 →

3 🧑‍🏫 Deployments — declare a wish, the monitor enforces it forever

You say "always 2"; the ReplicaSet counts non-stop and replaces anything that dies.

🧑‍🏫 You declare replicas: 2 · image: v5 Deployment school-api ReplicaSet the monitor · count = 2 🪑 pod …abc12 💥 crashes! 🪑 pod …def34 healthy 🪑 pod …xyz99 auto-created replacement 1 2 3 claims pods by label app: school-api 4 sees 1 missing → makes a new one

Read full lesson 03 →

4 ☎️ Services — one number that never changes

Pods get new IPs all the time; callers dial the Service name and it forwards to whoever is present.

🐍 analytics pod calls http://school-api ☎️ Service school-api stable IP + DNS · port 80 → 3000 🪑 pod 10.0.4.7 ready 🪑 pod 10.0.9.2 ready 💀 old pod gone — auto-removed 1 DNS finds the Service 2 picks a READY pod by label 3

Read full lesson 04 →

5 🚪 Namespaces — classrooms inside one building

One shared cluster, partitioned into rooms — names only need to be unique inside a room.

🏫 One Kubernetes cluster — the building 🚪 namespace: school Deployments: school-api, analytics Services ×2 · ConfigMap · Secret HPA · Ingress 1 our room — everything in this repo 🚪 kube-system the school office: CoreDNS · metrics-server 2 Kubernetes' own machinery 🚪 default experiments land here if you forget -n 😅 3 Same name in two rooms is fine — "Aarav from 3A" vs "Aarav from 3B" · delete a room = clean sweep of only that room

Read full lesson 05 →

6 🔑 ConfigMaps & Secrets — settings live outside the lunchbox

Same image everywhere; the notice board (config) and locker key (secret) are injected at start-up.

📌 ConfigMap PORT: 3000 · APP_VERSION 🔑 Secret DATABASE_URL (password!) 🪑 Pod at start-up env: PORT, DATABASE_URL… app reads process.env 🍱 Image identical in dev & prod 1 envFrom: configMapRef 2 secretKeyRef (never in git) 3 change settings → restart pods → same image, new behaviour

Read full lesson 06 →

7 🙋 Health probes — two questions, two very different consequences

Liveness failure restarts the container; readiness failure only pauses its traffic.

🧑‍🏫 kubelet asks every few seconds GET /healthz — alive? liveness probe GET /readyz — ready? readiness probe 🔄 fails ×3 → RESTART same pod, fresh container ⏸️ fails → NO traffic out of the Service list, no restart 1 2 3 4 mixing the two up is one of the most common Kubernetes mistakes!

Read full lesson 07 →

8 🍛 Requests & limits — promised plates and capped seconds

Requests are reservations the scheduler counts; limits are hard caps with different penalties for CPU vs memory.

📋 Scheduler the cook counting plates 🖥️ Node — 2 CPU · 4 Gi 🪑 school-api — req 100m / 128Mi limit 500m / 256Mi 🪑 analytics — req 100m / 128Mi limit 500m / 256Mi 🟩 unreserved — room for more pods 🍲 CPU over limit → throttled (slower) 🫃 RAM over limit → OOMKilled 💀 1 places pods only where requests fit 2 3

Read full lesson 08 →

9 🚌 Autoscaling — the transport manager watches the buses

Above 70% average CPU it adds pods (max 5); when quiet it parks them (never below 2).

📊 metrics-server measures pod CPU 🧑‍💼 HPA target 70% · min 2 · max 5 Deployment replicas: ← edited 🚌 1 🚌 2 🚌 3 🚌 4 1 2 85% → scale up! 3 scale-up is fast · scale-down is slow on purpose (don't park the bus the second the rain stops)

Read full lesson 09 →

10 🏫 Ingress — one gate, routed by the signboard

One load balancer for everything; the URL path decides which Service the visitor reaches.

🌍 Internet parent's browser 🏫 Main gate — ALB built by the controller 🪧 path? the signboard ☎️ analytics svc /analytics/* ☎️ school-api svc /* everything else 1 2 3 4 specific paths first (/analytics before /) · health-checks reuse lesson 07's /healthz · one gate = one bill + one TLS setup

Read full lesson 10 →

11 ⚽ Rollouts — substitute one player at a time, keep the bench

A full team is always on the field; the old version stays benched for instant rollback.

⏱️ during rollout 🪑 v1 pod — playing ✅ 🪑 v1 pod — playing ✅ 🏃 v2 pod — warming up 1 maxSurge: 1 extra allowed 🙋 readiness gate v2 gets traffic ONLY after /readyz says yes 2 ✅ done — all v2 old pods left one by one, never below full team 3 🪑 old ReplicaSet v1 — scaled to 0, kept on the bench kubectl rollout undo = "come back on!" — a 10-second rollback 4

Read full lesson 11 →

12 📚 Storage — backpacks vanish, library shelves survive

A PVC gives disposable pods a durable shelf; this repo goes one further and rents the library (RDS).

🪑 postgres pod v1 💥 dies (backpack gone) 🪑 postgres pod v2 new desk, same books 📝 PVC "I claim 10 Gi of shelf" 📚 PV = EBS disk outlives every pod 1 2 same claim → same data 3 🏛️ RDS — the rented library what THIS repo uses (terraform/rds.tf) 4 rule of thumb: stateless pods in the cluster, state in managed services outside

Read full lesson 12 →

13 🏢 Under the hood — what kubectl apply really does

Five office roles, one register, one endless loop: wish → record → reconcile → run → report.

🧑 kubectl "I wish: 2 pods" 🧑‍💼 API server the ONLY door 📖 etcd the sacred register 🔍 controllers wish ≠ reality? fix it 🗓️ scheduler picks the best node 🧑‍🏫 kubelet runs it on the node 1 2 wish written down 3 sees Deployment → creates Pods (unassigned) 4 assigns each pod a node 5 6 5: kubelet pulls the image 🍱, starts the container, runs the probes 🙋 · 6: "running & ready" reported back → kubectl get pods shows 2/2 🎉

Read full lesson 13 →

14 🤖 Bonus: CI/CD & GitOps — the homework robot and the caretaker robot

CI tests & builds on every push; ArgoCD lives inside the cluster and keeps it matching the plan book (git).

🧑‍💻 dev git push 📮 Robot 1 — CI pipeline ✅ test → 🍱 build → 🗄️ push to ECR → ✍️ manual approval 📖 git repo — the master plan k8s/ manifests = single source of truth 🏫 Kubernetes cluster 🤖 Robot 2 — ArgoCD lives INSIDE, no outside key 🪑 school pods 1 2 CI commits the new image tag 3 pulls & compares every ~3 min 4 4: sync + self-heal — hand-edits get reverted, deploys = git commits, rollback = git revert · the book always wins 📖

Read full lesson 14 →

15 🏫🏫 Bonus: Multi-AZ & the scaling ladder

Never seat the whole class in one building — and climb the cheapest rung first on exam day.

🌍 ALB — stands outside all buildings sends visitors wherever is healthy 🏫 building A — AZ a 🖥️ desk 🪑 api-1 🪑 analytics-1 🖥️ new desk added by the autoscaler 2 🏫 building B — AZ b 🖥️ desk 🪑 api-2 · 🪑 analytics-2 3 topology spread: copies across buildings — one fire ≠ down 1 🪜 the ladder 1 HPA: more pods — seconds 2 autoscaler: more desks — minutes 3 spread: across buildings 4 PDB: never all away at once Pending pods = kids standing — the autoscaler's signal 4

Read full lesson 15 →

16 🩺 Debugging — the nurse's triage chart

Whatever walks in: describe first, Events always, logs --previous for crash loops.

🤒 sick podsomething is wrong🩺 the nurse asks:1 describe → EVENTS2 logs --previous · 3 exec1Pending 🪑 — no desk fits: requests? taints?cluster full? (L08 · L15 · L20)ImagePullBackOff 🍱 — wrong label, missingtag, or no registry permission (AWS L05)CrashLoopBackOff 💥 — starts, dies, repeats:logs --previous first, then config (L06)OOMKilled 🫃 / Service silent ☎️ — limit hit,or selector ≠ labels: get endpointslices (L04)2match the symptomthe drill never changes:describe → logs → events → exec3

Read full lesson 16 →

17 🪪 RBAC — hall passes

Roles are passes, bindings hand them over, ServiceAccounts are the robots' cards.

👥 whoperson · group ·🤖 ServiceAccount (pod card)🔗 RoleBindinghands the pass over🪪 Role — one roomget/list pods in school🏫 ClusterRole — whole buildingview · edit · admin · cluster-admin 🗝️12the API server checks the pass on EVERY request — no pass, no entry · test: kubectl auth can-i --as=…3

Read full lesson 17 →

18 ⏰ Jobs & CronJobs — homework and the bell

The bell creates homework; homework creates a kid; the kid finishes and that's the point.

🔔 CronJobschedule: 0 2 * * *suspend: true = off-switch📝 Jobtonight's homeworkbackoffLimit: 2 retries🪑 podpg_dump → uploadthen EXITS 01on schedule2✅ Completed — kept for autopsy, tidied by history limits 🧹 · fire one now: kubectl create job --from=cronjob/school-db-backup3

Read full lesson 18 →

19 🚫📝 NetworkPolicies — passing-notes rules

Default-deny first, then exactly the conversations the app needs.

😳 defaultevery pod whispers toevery pod — even the DB1 default-denyno notes at all (selects every pod)2 allow exactly what's neededanalytics → api :3000 · gate → apps1🚫 anything elsenote intercepted —dropped silently2⚠️ needs a CNI that enforces policies (Calico / VPC CNI switch) — no enforcer = rules silently ignored: always TEST3

Read full lesson 19 →

20 🎫 Taints & affinity — assigned seating

Signs on desks push away; chits permit; wishes attract; anti-affinity separates twins.

💎 GPU desktaint: gpu=true:NoSchedulethe sign that repels🖥️ normal desksno signs — anyone sits🤖 ML podtoleration 🎫 + nodeAffinity:allowed AND attracted🪑 ordinary podno chit → repelled ❌12👯 anti-affinity:never both twinson one desk —HA at desk level3

Read full lesson 20 →

21 🧯 DaemonSets — one on every floor

No replica count: the node list IS the count. New desk, new extinguisher, automatically.

🧯 DaemonSetone per desk, always —no replica count: nodes ARE it🖥️ desk 1🪑 apps + 🧯🖥️ desk 2🪑 apps + 🧯🖥️ NEW desk(autoscaler, L15)🧯 appears by itself12you already run them: kube-proxy (L04's routing!) and the CNI live in kube-system as DaemonSets · log collectors next (L26)3

Read full lesson 21 →

22 🏷️ StatefulSets — desks with name plates

Sticky name, sticky drawer, ordered arrival — the whole difference from a Deployment.

🏷️ StatefulSetnamed, ordered kids:postgres-0, then -1, then -2🪑 postgres-0dies → replacement hasthe SAME name🗄️ PVC data-postgres-0ITS OWN drawer — survivesthe pod, always reattaches12☎️ headless Service: reach postgres-0 BY NAME, no load-balancing · HA needs an operator (L25) —production stance unchanged: rent the library (RDS, L12)3

Read full lesson 22 →

23 🍽️ QoS & evictions — who leaves first

BestEffort, then over-promise Burstable, then Guaranteed — your requests ARE your ranking.

🖥️ desk under pressurememory running out —kubelet must free space🥉 BestEffort — no requests at allEVICTED FIRST🥈 Burstable — requests < limitsnext (our pods live here)🥇 Guaranteed — requests == limitsevicted LAST12👮 PriorityClass:system staff basicallynever leaves3

Read full lesson 23 →

24 🏗️ Cluster upgrades — renovating while open

Office first, then desk by desk: cordon, drain (PDBs on guard), replace, uncordon.

🏢 1 office firstEKS control plane —one Terraform change🚧 cordonno new kids🚚 drainPDB guards 🍽️🖥️ replace deskfresh EC2 + kubelet1✅ uncordonnext classroom2📶 skew rule:desks may trailthe office,never lead ·one minor at a time3

Read full lesson 24 →

25 🤖 CRDs & operators — new words for the office

A word in the register plus a robot who makes it true — ArgoCD demystified.

📖 CRDteach the office a new word:kind: BackupPlan + grammar🏢 office / etcdstores your wishes —kubectl & RBAC just work🤖 controllerwatches the word, reconcilesforever — L03's loop, your noun12CRD + controller + expertise = an OPERATOR · you've used them all along: ArgoCD's Application,cert-manager's Certificate, CloudNativePG's Cluster — words plus robots 🤯3

Read full lesson 25 →

26 📈 Observability — eyes on everything

Report cards (metrics), diaries (logs), alarm bells (alerts) — and one screen for all of it.

🪑 pods expose /metricsnumbers on every door🧯 log collectorDaemonSet (L21) ships diaries📊 Prometheusscrapes + keeps history📜 log storediaries outlive desks📺 Grafanaone screen: metrics + logs🔔 Alertmanagerrules ring the on-call phone12start with: is it up · is it erroring · is it slow · is the cluster healthy — four questions, one dashboard3

Read full lesson 26 →