Challenge catalog

Kubernetes LLM challenges with evidence-based checks.

Start with free guided labs. Each challenge has objectives, commands, expected signals, paste-output validation, progressive hints, and a final readiness check.

Challenge of the Week

vLLM Inference Challenge

Deploy a GPU-backed OpenAI-compatible endpoint and prove scheduling, health, TTFT, queueing, and rollback readiness.

Hard75 minModel servingAI infrastructure engineer

Challenge catalog

Operator lab index

12/12 visible
Topic
Difficulty
Path
01
Model servingvllm-production-serving / production-readiness-ai

vLLM Inference Challenge

Deploy a GPU-backed OpenAI-compatible endpoint and prove scheduling, health, TTFT, queueing, and rollback readiness.

DifficultyHard
Time75 min
Progress statusnot started
02
RAGrag-platform-engineering

RAG Retrieval Challenge

Operate ingestion, metadata filters, vector retrieval, answer evaluation, and failure drills for production RAG.

DifficultyMedium
Time60 min
Progress statusnot started
03
Productionproduction-readiness-ai / kubernetes-llm-foundations

Production Readiness Challenge

Run a launch review across security, quota, rollout, observability, cost, and ownership before live traffic.

DifficultyHard
Time50 min
Progress statusnot started
04
Observabilityllm-observability-cost / production-readiness-ai

LLM Observability Challenge

Build the signal model needed to debug user latency, runtime saturation, GPU pressure, traces, logs, and alerts.

DifficultyMedium
Time45 min
Progress statusnot started
05
Model servingvllm-production-serving

vLLM Kubernetes Deployment Lab

Design the deployment contract for vLLM with model cache, readiness, runtime flags, and service exposure.

DifficultyMedium
Time55 min
Progress statusnot started
06
Architecturekubernetes-llm-foundations / vllm-production-serving

KServe vs Ray Serve Decision Lab

Choose the serving abstraction by ownership model, CRDs, graph complexity, autoscaling, and rollout needs.

DifficultyMedium
Time35 min
Progress statusnot started
07
GPU capacitykubernetes-llm-foundations / vllm-production-serving

GPU Node Pool Scheduling Lab

Prove accelerator placement with labels, taints, tolerations, quotas, and unschedulable-pod debugging.

DifficultyHard
Time65 min
Progress statusnot started
08
RAGrag-platform-engineering / llm-observability-cost

RAG Retrieval Quality Lab

Measure retrieval recall, citation accuracy, tenant filtering, and reranking latency before generation.

DifficultyHard
Time70 min
Progress statusnot started
09
Costllm-observability-cost / production-readiness-ai

Inference Cost Model Lab

Calculate cost per request from input tokens, output tokens, GPU profile, utilization, and cache behavior.

DifficultyMedium
Time45 min
Progress statusnot started
10
Productionproduction-readiness-ai / vllm-production-serving

LLM Rollout and Rollback Lab

Design traffic shifting, readiness gates, rollback triggers, and model-version ownership for inference services.

DifficultyHard
Time60 min
Progress statusnot started
11
Securityproduction-readiness-ai / rag-platform-engineering

Multi-Tenant LLM Security Lab

Review tenant routing, namespace boundaries, secrets, NetworkPolicy, prompt logging, and retrieval authorization.

DifficultyHard
Time70 min
Progress statusnot started
12
Observabilityllm-observability-cost

LLM Observability and Cost Dashboard Lab

Create a dashboard model that joins user latency, queue wait, GPU pressure, token throughput, and cost signals.

DifficultyMedium
Time50 min
Progress statusnot started