LLM capacity planner

Turn a model and workload into a hardware boundary.

Choose what the model must do, then keep loaded weights, KV Cache or training activations, runtime reserve, and per-GPU capacity in the same decision path.

Start with the real intent

Choose the work that will pressure the memory budget.

Open the capacity planner →

Reviewed model profile

Start from source-bound model facts.

Each profile is held to an exact repository revision and opens the shared planner with its architecture and request envelope intact.

Learn the moving parts

Read only what changes the capacity plan.

Open learning center →