qwen2.5-coder:32b

Qwen2.5 Coder 32B on your own machine

The closest local stand-in for a cloud coder. Wants a 24 GB class GPU.

Pull it

ollama pull qwen2.5-coder:32b

Then pick it in Osmium under Settings → Local models. Nothing else to configure.

Parameters
32B
VRAM (4-bit)
20 GB
Disk
~15 GB
Context
33K tokens
Tool calling
Yes
Reads images
No
Generates images
No
Coding
9/10
Explaining
8/10

Scores are our own product guidance, not benchmark results.

Which hardware runs it

Checked against GPU memory. Osmium runs the same check on your real machine and also looks at system RAM and free disk. Unified-memory Macs need more headroom because the GPU shares one pool with the operating system.

  • 4 GB — GTX 1650 / MX570Needs 20 GB; this tier has 4 GB.Too big
  • 6 GB — RTX 2060 / 3050Needs 20 GB; this tier has 6 GB.Too big
  • 8 GB — RTX 3060 Ti / 4060Needs 20 GB; this tier has 8 GB.Too big
  • 12 GB — RTX 3060 / 4070Needs 20 GB; this tier has 12 GB.Too big
  • 16 GB — RTX 4060 Ti / 4080Needs 20 GB; this tier has 16 GB.Too big
  • 24 GB — RTX 3090 / 4090Runs comfortably on 24 GB.Fits
  • Mac 16 GB unifiedNeeds 20 GB; this tier has 16 GB.Too big
  • Mac 24 GB unifiedRuns comfortably on 24 GB.Fits
  • Mac 36 GB unifiedRuns comfortably on 36 GB.Fits
  • Mac 64 GB unifiedRuns comfortably on 64 GB.Fits