notes/collection
inference engineering
17 notes — Backend → Inference-Platform track for Java + K8s engineers targeting 75-95k+ offshore-remote (5L/mo net).
- 00 Roadmap and Study Plan
- 00b Prerequisites Product and Model Selection
- 01 Fundamentals Prefill Decode KV Cache
- 02 GPU CUDA Basics for Backend Devs
- 03 vLLM Deep Dive
- 04 SGLang TGI Triton When Why
- 05 Quantization AWQ GPTQ FP8
- 06 Batching Caching Spec Decoding
- 06b Modalities
- 07 Kubernetes GPU Production
- 08 Routing Router LiteLLM Dynamo
- 09 Observability Benchmarking Cost
- 10 Java Bridge FastAPI Gateway
- 11 Projects P1 P2 P3
- 12 Interviews Resume Positioning
- 13 Resources Glossary
- overview· index of inference engineering