Until recently, a model with more than 100 billion parameters needed a server rack or a hefty cloud bill. Strata, an open-source project on GitHub, now runs Qwen3.8-Flash-Next (125B) on a typical gaming PC with a 12 GB graphics card and 64 GB of RAM. On an RTX 5070, it generates about 80–94 tokens per second. It runs fully offline, and nothing leaves your machine.
What is Strata?
Strata is an inference engine released under the MIT license. It is partly built on llama.cpp/ggml and tuned for one model family, Qwen3.8-Flash-Next. After installation you get:
- OpenAI-compatible and Anthropic-compatible APIs on
localhost. Existing tools work by pointing them athttp://127.0.0.1:8080/v1, and the Codex CLI and Anthropic-style clients are supported too. - A browser chat app with optional image input.
- An MCP server and adjustable reasoning effort (off/low/medium/high).
The project has already passed 16,000 GitHub stars.
How does a 125B model fit on a 12 GB GPU?
The trick lies in the model’s architecture and in how Strata spreads it across your hardware:
- Mixture of Experts. Qwen3.8-Flash-Next has 24,576 small “experts,” but each token uses only 10 of them.
- Tiered memory. The GPU holds the experts used most often, system RAM holds all of them, the CPU computes the rest in parallel, and an SSD stores a large lookup table.
- Aggressive quantization. The model ships in 2- to 4-bit GGUF variants (Q2_0, IQ2_XS, IQ3_XXS, IQ3_S, Unsloth UD-IQ4_XS).
- Speculative decoding. A small helper model drafts tokens and the large one verifies them, which makes answers 1.6–1.8x faster with identical output.
Hardware requirements
| Component | Minimum | Recommended |
|---|---|---|
| GPU | 12 GB VRAM (NVIDIA RTX 20–50, selected AMD Radeon) | 16–24 GB VRAM |
| RAM | 32 GB (Coder variant only) | 64 GB (runs every size) |
| Disk | ~80 GB free | NVMe SSD |
| OS | Windows 10/11 or Linux | — |
Setups with two or three GPUs are supported. Intel Arc, AMD Strix Halo and older cards work only experimentally.
How fast is it?
These are the author’s measurements with 32K-token prompts and 4K-token answers:
| Variant | RTX 5070 12 GB + 64 GB RAM | RX 9070 XT 16 GB + 47 GB RAM |
|---|---|---|
| Q2_0 | 94 tok/s | 60 tok/s |
| IQ2_XS (recommended) | 79 tok/s | 52 tok/s |
| IQ3_S | 53 tok/s | — |
| Coder | 55 tok/s | 44 tok/s |
The author estimates an RTX 3090 at around 100–140 tokens per second. Prompt processing runs at 1,000–2,600 tokens per second.
Which variant should you pick?
- 32 GB RAM: Coder. Half of its experts were removed, yet its authors report it keeps 91% of the full model’s SWE-bench Verified score. It is noticeably weaker outside code.
- 48 GB RAM: IQ2_XS or Q2_0.
- 64 GB RAM: IQ2_XS (the sweet spot), IQ3_XXS or IQ3_S.
- 96 GB+ RAM: IQ3_S or UD-IQ4_XS, which comes closest to the full model.
Limitations worth knowing
- Loading the model takes 1–3 minutes and 35–55 GB of RAM, and the PC may freeze during that time.
- The download is about 70 GB.
- By default Strata serves one request at a time, so it’s a personal workstation, not a team server.
- The first message of a long chat is processed slowly, at about 1 minute per 30,000 tokens.
- On Windows, AMD cards can’t read images yet.
Why developers will care?
For developers, testers and DevOps engineers, Strata means a large, recent model that is private by default. Your code, logs and client data never reach an external API. There are no per-token costs and no rate limits, and it plugs into the tools you already use through standard APIs. It’s an ideal setup for working offline with sensitive repositories, prototyping AI agents, and testing integrations before you pay for a cloud model.
At fireup.pro, the engineering team closely follows ways to run LLMs locally and on-premise, because data control is often the first question clients in regulated industries ask.
Sources:








