VRAM failure diagnostic

Turn an OOM into one controlled correction.

Recreate the attempted model load, engine start, long prompt, or additional request. The diagnostic follows allocation order, compares the same shape with nearby corrections, and keeps the current hardware profile available across VRAM tools.

OOM diagnostic

Find the first memory boundary in the failed shape.

Rebuild the attempted model, context, request pressure, and hardware topology, then test corrections against the same canonical memory model.

How the diagnostic works →

The values below are editable planning baselines. Change the few inputs you know; open Advanced assumptions only when they change your decision.

Model configuration

Import a Hugging Face model profile.

Load a public model revision to set its verified parameter total and decoder attention shape. The current workload context stays unchanged.

Current hardware

No hardware profile saved.

Save the GPU capacity you already have, then apply it across VRAM planning paths.

Advanced assumptionsMatch the attempted architecture, cache precision, topology, reserve, and usable-memory target.
First boundary
Required per GPU
After implementation margin
First correctionChange one lever

Ordered checks

Follow the allocation sequence.

BoundaryMeasured shapeState

Controlled corrections

Change one input and compare the new capacity result.

ChangeTargetRequired / GPUResult

Memory analysis

See what consumes the per-GPU capacity.

Compare the same memory-capacity model across workload pressure and weight formats, with the current planning point kept visible.

Planned with margin
Usable per GPU

Pressure matrix

Context × simultaneous requests

Each cell shows required VRAM per GPU. The outlined cell is the current plan.

Requests ↓ / context →

Weight-format comparison

Keep the workload constant.

Weights change between rows; architecture, KV Cache pressure, runtime reserve, and GPU topology stay fixed.

Weight formatWeights / GPURequired / GPUAfter margin

Diagnostic brief

Carry the failed shape and first correction together.

Keep this useful version close, shareable, and easy to revisit.

Rebuild the attempt

Use the weight format, architecture, cache precision, context, active requests, tensor parallelism, and usable-memory target from the failed run.

Follow allocation order

The result checks available topology, loaded weights, runtime base, request-driven KV Cache, and implementation margin in sequence.

Retest one lever

Choose one proposed change, preserve the rest of the workload, and repeat the failure-triggering prompt or request pattern.