vLLM Inference Challenge
Deploy a GPU-backed OpenAI-compatible endpoint and prove scheduling, health, TTFT, queueing, and rollback readiness.
AI infrastructure engineer
kubectl + vLLM + Prometheus
Challenge catalog
Start with free guided labs. Each challenge has objectives, commands, expected signals, paste-output validation, progressive hints, and a final readiness check.
Challenge of the Week
Deploy a GPU-backed OpenAI-compatible endpoint and prove scheduling, health, TTFT, queueing, and rollback readiness.
Challenge catalog
Deploy a GPU-backed OpenAI-compatible endpoint and prove scheduling, health, TTFT, queueing, and rollback readiness.
AI infrastructure engineer
kubectl + vLLM + Prometheus
Operate ingestion, metadata filters, vector retrieval, answer evaluation, and failure drills for production RAG.
MLOps engineer
kubectl + curl + vector database
Run a launch review across security, quota, rollout, observability, cost, and ownership before live traffic.
Platform lead
kubectl + policy engine + dashboard
Build the signal model needed to debug user latency, runtime saturation, GPU pressure, traces, logs, and alerts.
SRE
Prometheus + Grafana + OpenTelemetry
Design the deployment contract for vLLM with model cache, readiness, runtime flags, and service exposure.
AI infrastructure engineer
kubectl + vLLM + container registry
Choose the serving abstraction by ownership model, CRDs, graph complexity, autoscaling, and rollout needs.
Platform architect
decision matrix + runtime inventory
Prove accelerator placement with labels, taints, tolerations, quotas, and unschedulable-pod debugging.
Platform engineer
kubectl + NVIDIA device plugin + cluster autoscaler
Measure retrieval recall, citation accuracy, tenant filtering, and reranking latency before generation.
MLOps engineer
evaluation set + vector database + reranker
Calculate cost per request from input tokens, output tokens, GPU profile, utilization, and cache behavior.
AI platform lead
spreadsheet + metrics export + benchmark report
Design traffic shifting, readiness gates, rollback triggers, and model-version ownership for inference services.
SRE
Argo CD + gateway policy + metrics dashboard
Review tenant routing, namespace boundaries, secrets, NetworkPolicy, prompt logging, and retrieval authorization.
Security-minded platform engineer
kubectl + NetworkPolicy + admission policy
Create a dashboard model that joins user latency, queue wait, GPU pressure, token throughput, and cost signals.
SRE
Prometheus + Grafana + OpenTelemetry