- Model or quantization is too large
- Context length or batch size is too high
- Other applications consume memory
- Training needs more memory than inference
local companion
Local model ran out of memory
The model, context, or training process requested more VRAM, unified memory, or system RAM than the device could safely provide.

- CUDA out of memory
- out of memory
- failed to allocate memory
- metal out of memory
- Do not use page-file or swap success as proof of acceptable speed
- Do not start training merely because inference barely fits