Magnitude is an open-source inference engine for running large language models locally. Instead of shipping pre-compiled kernels, it compiles and tunes them on your own machine before the first token. According to the team’s own benchmarks, that makes decode 92% faster than llama.cpp on an M4 Pro Mac and 19% faster on NVIDIA. Magnitude is Apache 2.0 licensed and backed by Y Combinator (S25).

The project was released on September 30, 2026 and passed 6,400 GitHub stars within days. When the fireup.pro team tested it, one of our developers summed it up: „So basically, it’s a guide to turning your Mac into a small heater for winter.” Run a 27B model behind a coding agent for an afternoon and you’ll see what he meant.

What is Magnitude?

Magnitude is a desktop app with a bundled CLI that runs open-weight LLMs on your own hardware and exposes them to AI agents. It runs on macOS, Linux and Windows, on Apple Silicon, NVIDIA and AMD GPUs, and on CPU-only machines. There are no per-token fees. After you download a model, prompts, files and weights stay on the device, and it works offline.

Before you download anything, Magnitude profiles your hardware and estimates tokens per second for every model in its catalog. You see roughly how fast a model will run before you spend 40 GB of disk on it.

How does Magnitude work?

Most local runtimes, llama.cpp included, ship kernels built for broad hardware classes such as „Apple M-series” or „recent NVIDIA GPUs”. Magnitude builds and tunes its kernels on the target machine, for that exact chip. You pay a one-time compilation step and in return get kernels matched to your hardware.

Magnitude also targets multi-agent setups. Concurrent sessions share a prefix cache, so five agents working from the same system prompt don’t each recompute it. The team says each agent uses 27% less memory than it would otherwise, and that memory is released as soon as the agent stops.

Is Magnitude faster than llama.cpp?

On the vendor’s numbers, yes, and most of the gain is on Mac. These figures come from magnitude.dev and haven’t been independently verified yet.

HardwareMetricllama.cppMagnitudeDifference
Mac M4 Pro (Metal)Prefill466 tok/s507 tok/s+9%
Mac M4 Pro (Metal)Decode30 tok/s57 tok/s+92%
NVIDIA DGX Spark (CUDA)Prefill2,033 tok/s2,507 tok/s+23%
NVIDIA DGX Spark (CUDA)Decode49 tok/s58 tok/s+19%

Decode is the token-by-token generation phase. When an agent writes code, decode speed is what you wait on. Going from 30 to 57 tokens per second means a 500-token answer takes about 9 seconds instead of 17.

Which models does Magnitude support?

The catalog lists 15 open-weight models. The ones developers are most likely to try:

  • Qwen3.8 27B, a dense multimodal model and the strongest in the catalog by Magnitude’s own intelligence score
  • Gemma 4 31B, good at coding
  • Qwen3.6 35B-A3B, an MoE model that activates only about 3B parameters per token
  • Gemma 4 12B, a mid-size model with tool use
  • Qwen3.5 4B, for laptops with limited RAM

Downloads range from 0.7 GB (MiniCPM5 1B) to 40.4 GB (Qwen3.6 35B-A3B at Q8).

Does Magnitude work with Claude Code and Codex?

Yes. Magnitude has one-click connections to Claude Code, Codex, OpenCode, Cline, Pi, Hermes, OpenClaw and Oh My Pi. Other tools connect through its OpenAI-compatible API, which usually means changing the base URL and nothing else.

Will Magnitude make my Mac hot?

Under sustained load, yes. Local inference keeps the GPU and memory bandwidth busy for minutes or hours at a time, and faster kernels push the chip harder. A MacBook Pro will spin up its fans. A MacBook Air has no fan, so it will throttle instead and get slower as it heats up.

If you plan to run local agents all day, plug the laptop in and check your unified memory: a Q8 model plus several parallel agents can use most of it. For routine coding tasks, try a 4B or 12B model first. It answers faster and keeps your Mac much cooler.

How to get started with Magnitude

  1. Download the desktop app from magnitude.dev/download. The CLI comes with it.
  2. In Discover, pick a model and check the estimated speed on your hardware.
  3. In Connections, link your agent, or point any tool at the local OpenAI-compatible endpoint.

Who should try it?

Developers working on private codebases can now run a coding agent without sending code to an external API or paying per token. DevOps and platform teams get an Apache 2.0 alternative to llama.cpp that they can benchmark on their own GPU servers before committing to anything. QA engineers running many agent sessions in parallel will benefit most from shared prefix caching.

For architects in regulated industries such as healthcare and finance, Magnitude offers another way to keep data on the device. fireup.pro sees local models as a complement to cloud LLMs, not a replacement, and Magnitude makes it cheaper to test which workloads can move on-device.

Sources: magnitude.dev, GitHub: magnitudedev/magnitude, Magnitude model catalog, AI Weekly launch report