Deployment & Inference

Serving stacks, inference engines, GPU orchestration and cost-control tooling.

Tools in this category

9router

A routing layer that keeps your coding agent alive when one provider rate-limits you — worth it for the fallback alone.

Individuals and small teams hitting provider rate limits during heavy agent sessions Read more →

vLLM

High-throughput LLM inference engine with PagedAttention — the default choice when you need to squeeze a GPU.

Self-hosted production inference on NVIDIA/AMD GPUs Read more →

Notes in this category