VRAM failure diagnostic
Turn an OOM into one controlled correction.
Recreate the attempted model load, engine start, long prompt, or additional request. The diagnostic follows allocation order, compares the same shape with nearby corrections, and keeps the current hardware profile available across VRAM tools.
OOM diagnostic
Find the first memory boundary in the failed shape.
Rebuild the attempted model, context, request pressure, and hardware topology, then test corrections against the same canonical memory model.
The values below are editable planning baselines. Change the few inputs you know; open Advanced assumptions only when they change your decision.
Model configuration
Import a Hugging Face model profile.
Load a public model revision to set its verified parameter total and decoder attention shape. The current workload context stays unchanged.
Current hardware
No hardware profile saved.
Save the GPU capacity you already have, then apply it across VRAM planning paths.
Ordered checks
Follow the allocation sequence.
| Boundary | Measured shape | State |
|---|---|---|
| — | — | — |
| — | — | — |
| — | — | — |
| — | — | — |
| — | — | — |
Controlled corrections
Change one input and compare the new capacity result.
| Change | Target | Required / GPU | Result |
|---|
The current record already preserves its implementation margin; reconcile the actual runtime shape with the recorded inputs.
Memory analysis
See what consumes the per-GPU capacity.
Compare the same memory-capacity model across workload pressure and weight formats, with the current planning point kept visible.
- ——
- ——
- ——
- ——
- ——
—
Pressure matrix
Context × simultaneous requests
Each cell shows required VRAM per GPU. The outlined cell is the current plan.
| Requests ↓ / context → | — | — | — | — |
|---|---|---|---|---|
| — | —— | —— | —— | —— |
| — | —— | —— | —— | —— |
| — | —— | —— | —— | —— |
| — | —— | —— | —— | —— |
Weight-format comparison
Keep the workload constant.
Weights change between rows; architecture, KV Cache pressure, runtime reserve, and GPU topology stay fixed.
| Weight format | Weights / GPU | Required / GPU | After margin |
|---|---|---|---|
| — | — | — | — |
| — | — | — | — |
| — | — | — | — |
Diagnostic brief
Carry the failed shape and first correction together.
—
Next decision
Plan tools
Keep this useful version close, shareable, and easy to revisit.
Recent saved plans 0
Rebuild the attempt
Use the weight format, architecture, cache precision, context, active requests, tensor parallelism, and usable-memory target from the failed run.
Follow allocation order
The result checks available topology, loaded weights, runtime base, request-driven KV Cache, and implementation margin in sequence.
Retest one lever
Choose one proposed change, preserve the rest of the workload, and repeat the failure-triggering prompt or request pattern.