125-billion-parameter model runs on a 12 GB GPU with Strata
The open-source project Strata makes the Qwen3.8-Flash-Next model, with 125 billion parameters, run on common graphics cards by splitting the work between GPU, RAM, and processor.
The open-source project Strata makes the Qwen3.8-Flash-Next model, with 125 billion parameters, run on common graphics cards by splitting the work between GPU, RAM, and processor.
What Strata Is
Strata is an open-source inference engine maintained on GitHub by a developer identified as Niko1221. The project now has 9,700 stars and 865 forks on the repository, is published under the MIT license, and includes a one-click installer for Windows and Linux (START-HERE.bat or setup.sh).

What Strata does is run Qwen3.8-Flash-Next, a model from the Qwen team with 125 billion parameters, on a single consumer graphics card. Models at this scale normally require servers with hundreds of gigabytes of video memory; Strata manages to run on cards with as little as 12 GB of VRAM, splitting the work between GPU, RAM, and processor.
How a 125B Model Fits in 12 GB of VRAM
Qwen3.8-Flash-Next is organized as a mixture of experts (MoE): the full model has 24,576 small "experts," but each generated word uses only 10 of them at a time. Strata keeps the most frequently activated experts on the GPU, stores the full set in RAM, and lets the processor handle the rest in parallel. The repository itself compares this to a kitchen: what gets used all the time stays on the counter, and the rest waits in the pantry.
There's also speculative decoding: a smaller auxiliary model "guesses" the next words and the large model checks them all at once, which delivers the same final answer 1.6 to 1.8 times faster, according to the maintainers. Long prompts are read in chunks of up to 8,192 tokens at a time, at more than 1,000 tokens per second.
The Numbers Measured by the Project
The maintainers measured performance on two common machines: one with an RTX 5070 (12 GB), a Ryzen 5 7600, and 64 GB of RAM, the other with an AMD RX 9070 XT (16 GB). On the RTX 5070, the response-generation numbers (tokens per second) and the reading of a 32,000-token prompt vary depending on the model's compression level:
| Quantization | Generates response | Reads prompt |
|---|---|---|
| Q2_0 | 94 tokens/s | 2,650 tokens/s |
| IQ2_XS | 79 tokens/s | 2,090 tokens/s |
| IQ3_XXS | 62 tokens/s | 1,750 tokens/s |
| IQ3_S | 53 tokens/s | 1,620 tokens/s |
| Coder | 55 tokens/s | 2,180 tokens/s |
For comparison, the repository notes that 60 tokens per second is already faster than the average human reading speed. AMD's RX 9070 XT was tested separately and produced numbers in the same range; the full table, including the two engines used (0.1.26 and 0.1.36), is in DETAILS.md in the repository. A 24 GB RTX 3090, according to the maintainers' estimate (not measured directly), should write between 100 and 140 tokens per second.
What the Hardware Needs
The requirements stated in the repository:
- GPU: NVIDIA GeForce RTX 20, 30, 40, or 50 series, or AMD Radeon RX 7900 XT/XTX, RX 7800 XT/7700 XT, RX 9060 XT, RX 9070/9070 XT, Radeon AI PRO R9700, or RX 6800/6900, with 12 GB of VRAM or more.
- RAM: 32 GB or more; the amount of RAM determines which model size fits, and 64 GB runs any size.
- Disk: about 80 GB free, preferably on SSD (the model download is around 70 GB).
- System: Windows 10/11 or Linux, with an updated GPU driver.
Older GPUs (Tesla P40/V100, GTX 10, Radeon VII/MI50, RX 6700 XT, RX 5500 XT), Intel Arc cards, and processors without AVX2 also work, but are treated as experimental support, tested by the community and not by the main maintainers.
Which Model Variant to Choose
The same architecture comes in different sizes, depending on available RAM:
- 32 GB of RAM: the Coder version, built for code, with half the experts removed. According to the authors' own measurements, it reaches 91% of the full model's score on SWE-bench Verified, but is weaker outside of a programming context (including in Chinese and other CJK languages).
- 48 GB:
IQ2_XSorQ2_0, the fastest. - 64 GB:
IQ2_XS(recommended by the maintainers),IQ3_XXS, orIQ3_S(the slowest and "smartest" of the three). - 96 GB or more:
IQ3_Sor Unsloth's ~4-bit version, UD-IQ4_XS.
There's also the Swift 1.5 version, a fine-tune that reduces reasoning time before answering, delivering the response sooner with similar quality, and Unsloth's UD-Q4_K_XL, the closest to the original model in quality, but which Strata needs to read from the SSD during the response: on a PC with 64 GB of RAM it writes only 7 to 8.5 tokens per second, well below the other options.
The Hook for Those Already Using Code Agents
The point that matters most to developers is API compatibility. Strata exposes the model at http://127.0.0.1:8080/v1, in the same format as the OpenAI API, and also at /v1/messages, in the Anthropic format. This means pointing existing tools at the local address without switching tools.
In practice, the repository cites direct integration with Claude Code (via the environment variable ANTHROPIC_BASE_URL=http://127.0.0.1:8080), with Codex CLI and other apps that use the OpenAI Responses API (/v1/responses), and with any OpenAI-compatible client, including GitHub Copilot and Cursor. Strata also exposes an MCP server, which lets a code agent install, start, and stop Strata itself.
What Still Limits Usage
By default, Strata handles one request at a time: while it processes one question, the others wait in queue. You can set "parallel": 2 to handle two at once, but on a 12 GB card this makes each individual response slower.
The first startup also locks up the PC for 1 to 3 minutes while Strata loads 35 to 55 GB into RAM and reserves part of it for the GPU, which is expected and does not indicate a failure. On AMD cards, image reading only works on Linux, running through the processor; on Windows this feature is not yet available.
It remains an open question, for now, how well these numbers hold up on hardware more modest than what was tested (RTX 5070 and RX 9070 XT) and on older cards, where support is still labeled experimental by the maintainers themselves.
Translated from the Brazilian Portuguese original · Read the original
Critical GitLab vulnerability allows unauthenticated data exfiltration and is already under attack
CVE-2026-85706 has the maximum severity score (CVSS 10.0) and allows anyone, without logging in, to read arbitrary files from self-managed GitLab instances. The fix has existed since September, but CISA confirms active exploitation.