llama3.3:70b
Llama 3.3 70B on your own machine
The best local explainer if you have 48 GB or more. Slow.
Pull it
ollama pull llama3.3:70bThen pick it in Osmium under Settings → Local models. Nothing else to configure.
- Parameters
- 70B
- VRAM (4-bit)
- 40 GB
- Disk
- ~30 GB
- Context
- 128K tokens
- Tool calling
- Yes
- Reads images
- No
- Generates images
- No
- Coding
- 7/10
- Explaining
- 9/10
Scores are our own product guidance, not benchmark results.
Which hardware runs it
Checked against GPU memory. Osmium runs the same check on your real machine and also looks at system RAM and free disk. Unified-memory Macs need more headroom because the GPU shares one pool with the operating system.
- 4 GB — GTX 1650 / MX570Needs 40 GB; this tier has 4 GB.Too big
- 6 GB — RTX 2060 / 3050Needs 40 GB; this tier has 6 GB.Too big
- 8 GB — RTX 3060 Ti / 4060Needs 40 GB; this tier has 8 GB.Too big
- 12 GB — RTX 3060 / 4070Needs 40 GB; this tier has 12 GB.Too big
- 16 GB — RTX 4060 Ti / 4080Needs 40 GB; this tier has 16 GB.Too big
- 24 GB — RTX 3090 / 4090Needs 40 GB; this tier has 24 GB.Too big
- Mac 16 GB unifiedNeeds 40 GB; this tier has 16 GB.Too big
- Mac 24 GB unifiedNeeds 40 GB; this tier has 24 GB.Too big
- Mac 36 GB unifiedNeeds 40 GB; this tier has 36 GB.Too big
- Mac 64 GB unifiedRuns comfortably on 64 GB.Fits