OpenAI's gpt-oss-20b runs in 16GB of memory, but the documentation leaves gaps
Released in August 2025, gpt-oss-20b promises to fit in 16GB of memory thanks to MXFP4 quantization. The official documentation explains the happy path; what happens below that, or outside Ollama's script, is left to the developer.
An open model designed to fit outside the data center
OpenAI released gpt-oss-20b in August 2025 as the smaller sibling in an open-weight pair: alongside gpt-oss-120b, it is marketed as the option for "low latency and local or specialized use," according to the official documentation on GitHub. It has 21 billion parameters in total, but only 3.6 billion active per inference, because the architecture is mixture-of-experts: the model activates only a fraction of its weights at each step, which reduces computational cost without discarding the model's total capacity.
The license is Apache 2.0, with no copyleft restrictions, which clears the way for commercial use and fine-tuning without asking permission. The model comes with native support for function calling, Python code execution, and a web search tool, plus configurable reasoning effort (low, medium, high) to trade speed against response quality. This matters for anyone thinking about agents: you can trade reasoning depth for latency without switching models.
MXFP4: the quantization that makes the 16GB figure possible
The reason gpt-oss-20b fits on home hardware has a name: MXFP4. OpenAI applied this quantization to the MoE layer weights during post-training, and that is what lets the model "run within 16GB of memory," in the documentation's own words. The important detail that's easy to miss on a quick read: every quality evaluation OpenAI has published was already done with this same quantization applied, not with the model at full precision and reduced afterward.
This changes how to read the announcement. It's not "the original model runs well and can also be compressed"; it's "the model that was tested and evaluated is already the compressed one." In practice, this narrows the room for discussion about quality loss from quantization, because there is no BF16 version of gpt-oss-20b being compared side by side in the official materials. Anyone who wants to measure that loss on their own needs to run the reference implementation in PyTorch, which the repository itself describes as unoptimized and educational.
Three paths, three different hardware requirements
The documentation lists multiple ways to run gpt-oss-20b, but they are not interchangeable in terms of hardware requirements. It's worth separating what works on a laptop from what works on a server:
| Path | Entry command | Target hardware |
|---|---|---|
| Ollama | ollama pull gpt-oss:20b and ollama run gpt-oss:20b | Consumer, CPU or modest GPU |
| LM Studio | lms get openai/gpt-oss-20b | Consumer, graphical interface |
| vLLM | vllm serve openai/gpt-oss-20b | Dedicated GPU, production service |
| PyTorch (reference) | torchrun --nproc-per-node=4 -m gpt_oss.generate | Minimum 4x H100, not suited for local use |
| Metal | python gpt_oss/metal/examples/generate.py | Apple Silicon, after converting the weights |
The Ollama path is the only one OpenAI itself explicitly recommends for anyone "trying to run gpt-oss on consumer hardware," including a direct caveat: the reference implementations in PyTorch and Triton have not been tested on Windows, and anyone on that system should use solutions like Ollama instead of them.
What the documentation doesn't say: below 16GB
The 16GB figure appears as the declared floor for gpt-oss-20b with MXFP4, but OpenAI's material doesn't detail what happens on machines with less memory than that. There's no degradation table, no even more aggressive quantization mode published for this model, and no official guidance on offloading to disk or swap. Anyone with 8GB of RAM who tries to run it via Ollama anyway is testing a scenario the documentation simply doesn't cover.
In short: the official playbook guarantees operation within a specific window (16GB+ of unified memory or VRAM, using the Ollama or LM Studio path) and stays silent about anything outside it. That doesn't mean it's impossible to run with less; it means that, if it does run, the behavior around speed and any context cuts isn't documented anywhere in the official source. It's trial-and-error territory, not a fulfilled promise.
Harmony: the format that trips up anyone who skips a step
One detail that the documentation's highlights section makes a point of repeating, in bold, is that both models "were trained using our harmony response format and should only be used with that format; otherwise, they will not work correctly." This isn't a stylistic detail: it's a formatting contract that, if ignored, produces broken output even with the model loaded and running without errors.
On the vLLM path, this shows up explicitly in the documentation's own example code: before generating any response, you need to load HarmonyEncodingName.HARMONY_GPT_OSS, assemble the conversation with Conversation.from_messages, render the prefill with encoding.render_conversation_for_completion, and only then parse the output back with encoding.parse_messages_from_completion_tokens. Skipping this encoding and sending a raw prompt to the model via vLLM is the most common recipe for anyone who opens an issue asking why gpt-oss "doesn't work right."
Anyone using Ollama or LM Studio doesn't deal directly with this manual encoding: OpenAI itself recommends these tools as the standard path for consumer hardware, sparing the user from assembling the harmony pipeline by hand as in the vLLM example above. The documentation doesn't detail how each tool handles the format internally, but OpenAI's explicit recommendation for this audience suggests that, in practice, this complexity stays abstracted away from anyone using the high-level interface.
Codex as a local client: a concrete use case
The documentation includes a practical example that's worth it for anyone already using Codex day to day: you can point Codex to a local gpt-oss-20b instance running via Ollama. Just configure ~/.codex/config.toml with a provider pointing to http://localhost:11434/v1 (Ollama's default port) and an oss profile using model = "gpt-oss:20b". After that, the codex -p oss command runs against the local model instead of hitting a cloud endpoint.
This works because the Ollama server exposes a chat-completions-compatible API, and any client that speaks that protocol can connect, not just Codex. It's a reasonable starting point for anyone who wants to measure, on their own machine, whether gpt-oss-20b can handle coding tasks without relying on a paid API key.
Where I'd start
For anyone who just wants to experiment with no commitment, the path of least friction is still ollama pull gpt-oss:20b followed by ollama run gpt-oss:20b, with at least 16GB of available memory (unified on Apple Silicon, or RAM plus VRAM on PCs with a GPU). If the idea is to expose the model as a service, with real batching and fine-grained control over parameters, vLLM is the right path, but it requires a dedicated GPU and handling the harmony format manually if you're not using the high-level interface.
What remains open, because the source doesn't settle it, is exactly the range below 16GB and any quality comparison between the MXFP4-quantized model and a hypothetical full-precision version. Until OpenAI publishes that data, the honest answer to anyone asking "will it run on my 8GB laptop" is: the documentation doesn't promise that, and no one has officially tested it.
Translated from the Brazilian Portuguese original · Read the original
OpenAI Launches Decisions API, dots, and GPT-6.1 Sol at DevDay 2026
At DevDay 2026, held on September 28 and 29, OpenAI reorganized its platform stack with a cheaper model, a fast classification API, and agents that run on their own in the cloud. Here's what changes for people building products on top of the API.