1 🍱 Containers & images — recipe → lunchbox → lunch
One recipe (Dockerfile) bakes one frozen lunchbox (image); every opened box (container) is identical.
📝 Dockerfile
the recipe card
🧊 Image
school-api:v1 — frozen box
🗄️ Registry (ECR)
the shelf where images wait
🏃 Running containers
box 1
box 2
box 3
1
docker build
2
docker push
3
docker run ×3 — identical
2 🪑 Pods — one desk, one address, replace don't repair
Kubernetes never runs a bare container — it always sits at a desk (Pod) with its own IP and shared shelf.
🌐 Cluster network
callers
🪑 Pod — IP 10.0.4.7
📦 school-api
main container
📦 sidecar
optional helper
🗄️ shared volume
1
talks to the Pod's IP
2
🪑 NEW Pod
new name, new IP 10.0.9.9
3
💥 desks are replaced, never repaired
3 🧑🏫 Deployments — declare a wish, the monitor enforces it forever
You say "always 2"; the ReplicaSet counts non-stop and replaces anything that dies.
🧑🏫 You declare
replicas: 2 · image: v5
Deployment
school-api
ReplicaSet
the monitor · count = 2
🪑 pod …abc12
💥 crashes!
🪑 pod …def34
healthy
🪑 pod …xyz99
auto-created replacement
1
2
3
claims pods by label app: school-api
4
sees 1 missing → makes a new one
4 ☎️ Services — one number that never changes
Pods get new IPs all the time; callers dial the Service name and it forwards to whoever is present.
🐍 analytics pod
calls http://school-api
☎️ Service school-api
stable IP + DNS · port 80 → 3000
🪑 pod 10.0.4.7
ready
🪑 pod 10.0.9.2
ready
💀 old pod
gone — auto-removed
1
DNS finds the Service
2
picks a READY pod by label
3
5 🚪 Namespaces — classrooms inside one building
One shared cluster, partitioned into rooms — names only need to be unique inside a room.
🏫 One Kubernetes cluster — the building
🚪 namespace: school
Deployments: school-api, analytics
Services ×2 · ConfigMap · Secret
HPA · Ingress
1
our room — everything in this repo
🚪 kube-system
the school office:
CoreDNS · metrics-server
2
Kubernetes' own machinery
🚪 default
experiments land here
if you forget -n 😅
3
Same name in two rooms is fine — "Aarav from 3A" vs "Aarav from 3B" · delete a room = clean sweep of only that room
6 🔑 ConfigMaps & Secrets — settings live outside the lunchbox
Same image everywhere; the notice board (config) and locker key (secret) are injected at start-up.
📌 ConfigMap
PORT: 3000 · APP_VERSION
🔑 Secret
DATABASE_URL (password!)
🪑 Pod at start-up
env: PORT, DATABASE_URL…
app reads process.env
🍱 Image
identical in dev & prod
1
envFrom: configMapRef
2
secretKeyRef (never in git)
3
change settings → restart pods → same image, new behaviour
7 🙋 Health probes — two questions, two very different consequences
Liveness failure restarts the container; readiness failure only pauses its traffic.
🧑🏫 kubelet
asks every few seconds
GET /healthz — alive?
liveness probe
GET /readyz — ready?
readiness probe
🔄 fails ×3 → RESTART
same pod, fresh container
⏸️ fails → NO traffic
out of the Service list, no restart
1
2
3
4
mixing the two up is one of the most common Kubernetes mistakes!
8 🍛 Requests & limits — promised plates and capped seconds
Requests are reservations the scheduler counts; limits are hard caps with different penalties for CPU vs memory.
📋 Scheduler
the cook counting plates
🖥️ Node — 2 CPU · 4 Gi
🪑 school-api — req 100m / 128Mi
limit 500m / 256Mi
🪑 analytics — req 100m / 128Mi
limit 500m / 256Mi
🟩 unreserved — room for more pods
🍲 CPU over limit
→ throttled (slower)
🫃 RAM over limit
→ OOMKilled 💀
1
places pods only where requests fit
2
3
9 🚌 Autoscaling — the transport manager watches the buses
Above 70% average CPU it adds pods (max 5); when quiet it parks them (never below 2).
📊 metrics-server
measures pod CPU
🧑💼 HPA
target 70% · min 2 · max 5
Deployment
replicas: ← edited
🚌 1
🚌 2
🚌 3
🚌 4
1
2
85% → scale up!
3
scale-up is fast · scale-down is slow on purpose (don't park the bus the second the rain stops)
10 🏫 Ingress — one gate, routed by the signboard
One load balancer for everything; the URL path decides which Service the visitor reaches.
🌍 Internet
parent's browser
🏫 Main gate — ALB
built by the controller
🪧 path?
the signboard
☎️ analytics svc
/analytics/*
☎️ school-api svc
/* everything else
1
2
3
4
specific paths first (/analytics before /) · health-checks reuse lesson 07's /healthz · one gate = one bill + one TLS setup
11 ⚽ Rollouts — substitute one player at a time, keep the bench
A full team is always on the field; the old version stays benched for instant rollback.
⏱️ during rollout
🪑 v1 pod — playing ✅
🪑 v1 pod — playing ✅
🏃 v2 pod — warming up
1
maxSurge: 1 extra allowed
🙋 readiness gate
v2 gets traffic ONLY
after /readyz says yes
2
✅ done — all v2
old pods left one by one,
never below full team
3
🪑 old ReplicaSet v1 — scaled to 0, kept on the bench
kubectl rollout undo = "come back on!" — a 10-second rollback
4
12 📚 Storage — backpacks vanish, library shelves survive
A PVC gives disposable pods a durable shelf; this repo goes one further and rents the library (RDS).
🪑 postgres pod v1
💥 dies (backpack gone)
🪑 postgres pod v2
new desk, same books
📝 PVC
"I claim 10 Gi of shelf"
📚 PV = EBS disk
outlives every pod
1
2
same claim → same data
3
🏛️ RDS — the rented library
what THIS repo uses (terraform/rds.tf)
4
rule of thumb: stateless pods in the cluster, state in managed services outside
13 🏢 Under the hood — what kubectl apply really does
Five office roles, one register, one endless loop: wish → record → reconcile → run → report.
🧑 kubectl
"I wish: 2 pods"
🧑💼 API server
the ONLY door
📖 etcd
the sacred register
🔍 controllers
wish ≠ reality? fix it
🗓️ scheduler
picks the best node
🧑🏫 kubelet
runs it on the node
1
2
wish written down
3
sees Deployment →
creates Pods (unassigned)
4
assigns each pod
a node
5
6
5: kubelet pulls the image 🍱, starts the container, runs the probes 🙋 · 6: "running & ready" reported back → kubectl get pods shows 2/2 🎉
14 🤖 Bonus: CI/CD & GitOps — the homework robot and the caretaker robot
CI tests & builds on every push; ArgoCD lives inside the cluster and keeps it matching the plan book (git).
🧑💻 dev
git push
📮 Robot 1 — CI pipeline
✅ test → 🍱 build → 🗄️ push to ECR
→ ✍️ manual approval
📖 git repo — the master plan
k8s/ manifests = single source of truth
🏫 Kubernetes cluster
🤖 Robot 2 — ArgoCD
lives INSIDE, no outside key
🪑 school pods
1
2
CI commits the new image tag
3
pulls & compares every ~3 min
4
4: sync + self-heal — hand-edits get reverted, deploys = git commits, rollback = git revert · the book always wins 📖
15 🏫🏫 Bonus: Multi-AZ & the scaling ladder
Never seat the whole class in one building — and climb the cheapest rung first on exam day.
🌍 ALB — stands outside all buildings
sends visitors wherever is healthy
🏫 building A — AZ a
🖥️ desk
🪑 api-1
🪑 analytics-1
🖥️ new desk
added by the
autoscaler
2
🏫 building B — AZ b
🖥️ desk
🪑 api-2 · 🪑 analytics-2
3
topology spread: copies across buildings — one fire ≠ down
1
🪜 the ladder
1 HPA: more pods — seconds
2 autoscaler: more desks — minutes
3 spread: across buildings
4 PDB: never all away at once
Pending pods = kids standing —
the autoscaler's signal
4
16 🩺 Debugging — the nurse's triage chart
Whatever walks in: describe first, Events always, logs --previous for crash loops.
🤒 sick pod something is wrong 🩺 the nurse asks: 1 describe → EVENTS 2 logs --previous · 3 exec 1 Pending 🪑 — no desk fits: requests? taints? cluster full? (L08 · L15 · L20) ImagePullBackOff 🍱 — wrong label, missing tag, or no registry permission (AWS L05) CrashLoopBackOff 💥 — starts, dies, repeats: logs --previous first, then config (L06) OOMKilled 🫃 / Service silent ☎️ — limit hit, or selector ≠ labels: get endpointslices (L04) 2 match the symptom the drill never changes: describe → logs → events → exec 3
17 🪪 RBAC — hall passes
Roles are passes, bindings hand them over, ServiceAccounts are the robots' cards.
👥 who person · group · 🤖 ServiceAccount (pod card) 🔗 RoleBinding hands the pass over 🪪 Role — one room get/list pods in school 🏫 ClusterRole — whole building view · edit · admin · cluster-admin 🗝️ 1 2 the API server checks the pass on EVERY request — no pass, no entry · test: kubectl auth can-i --as=… 3
18 ⏰ Jobs & CronJobs — homework and the bell
The bell creates homework; homework creates a kid; the kid finishes and that's the point.
🔔 CronJob schedule: 0 2 * * * suspend: true = off-switch 📝 Job tonight's homework backoffLimit: 2 retries 🪑 pod pg_dump → upload then EXITS 0 1 on schedule 2 ✅ Completed — kept for autopsy, tidied by history limits 🧹 · fire one now: kubectl create job --from=cronjob/school-db-backup 3
19 🚫📝 NetworkPolicies — passing-notes rules
Default-deny first, then exactly the conversations the app needs.
😳 default every pod whispers to every pod — even the DB 1 default-deny no notes at all (selects every pod) 2 allow exactly what's needed analytics → api :3000 · gate → apps 1 🚫 anything else note intercepted — dropped silently 2 ⚠️ needs a CNI that enforces policies (Calico / VPC CNI switch) — no enforcer = rules silently ignored: always TEST 3
20 🎫 Taints & affinity — assigned seating
Signs on desks push away; chits permit; wishes attract; anti-affinity separates twins.
💎 GPU desk taint: gpu=true:NoSchedule the sign that repels 🖥️ normal desks no signs — anyone sits 🤖 ML pod toleration 🎫 + nodeAffinity: allowed AND attracted 🪑 ordinary pod no chit → repelled ❌ 1 2 👯 anti-affinity: never both twins on one desk — HA at desk level 3
21 🧯 DaemonSets — one on every floor
No replica count: the node list IS the count. New desk, new extinguisher, automatically.
🧯 DaemonSet one per desk, always — no replica count: nodes ARE it 🖥️ desk 1 🪑 apps + 🧯 🖥️ desk 2 🪑 apps + 🧯 🖥️ NEW desk (autoscaler, L15) 🧯 appears by itself 1 2 you already run them: kube-proxy (L04's routing!) and the CNI live in kube-system as DaemonSets · log collectors next (L26) 3
22 🏷️ StatefulSets — desks with name plates
Sticky name, sticky drawer, ordered arrival — the whole difference from a Deployment.
🏷️ StatefulSet named, ordered kids: postgres-0, then -1, then -2 🪑 postgres-0 dies → replacement has the SAME name 🗄️ PVC data-postgres-0 ITS OWN drawer — survives the pod, always reattaches 1 2 ☎️ headless Service: reach postgres-0 BY NAME, no load-balancing · HA needs an operator (L25) — production stance unchanged: rent the library (RDS, L12) 3
23 🍽️ QoS & evictions — who leaves first
BestEffort, then over-promise Burstable, then Guaranteed — your requests ARE your ranking.
🖥️ desk under pressure memory running out — kubelet must free space 🥉 BestEffort — no requests at all EVICTED FIRST 🥈 Burstable — requests < limits next (our pods live here) 🥇 Guaranteed — requests == limits evicted LAST 1 2 👮 PriorityClass: system staff basically never leaves 3
24 🏗️ Cluster upgrades — renovating while open
Office first, then desk by desk: cordon, drain (PDBs on guard), replace, uncordon.
🏢 1 office first EKS control plane — one Terraform change 🚧 cordon no new kids 🚚 drain PDB guards 🍽️ 🖥️ replace desk fresh EC2 + kubelet 1 ✅ uncordon next classroom 2 📶 skew rule: desks may trail the office, never lead · one minor at a time 3
25 🤖 CRDs & operators — new words for the office
A word in the register plus a robot who makes it true — ArgoCD demystified.
📖 CRD teach the office a new word: kind: BackupPlan + grammar 🏢 office / etcd stores your wishes — kubectl & RBAC just work 🤖 controller watches the word, reconciles forever — L03's loop, your noun 1 2 CRD + controller + expertise = an OPERATOR · you've used them all along: ArgoCD's Application, cert-manager's Certificate, CloudNativePG's Cluster — words plus robots 🤯 3
26 📈 Observability — eyes on everything
Report cards (metrics), diaries (logs), alarm bells (alerts) — and one screen for all of it.
🪑 pods expose /metrics numbers on every door 🧯 log collector DaemonSet (L21) ships diaries 📊 Prometheus scrapes + keeps history 📜 log store diaries outlive desks 📺 Grafana one screen: metrics + logs 🔔 Alertmanager rules ring the on-call phone 1 2 start with: is it up · is it erroring · is it slow · is the cluster healthy — four questions, one dashboard 3