NEWS

Step 5 Preview: StepFun's MoE model with 1M tokens arrives on OpenRouter

StepFun released Step 5 Preview on October 8, 2026, a Mixture-of-Experts model with a 1-million-token window already available on OpenRouter. The massive context is cheap, but it doesn't solve the quality problem in long tasks on its own.

What is Step 5 Preview

StepFun, a Chinese AI startup, released Step 5 Preview on October 8, 2026, its flagship model for agentic tasks. It is already available via OpenRouter, listed as stepfun/step-5-preview.

The architecture is a sparse Mixture-of-Experts (MoE): 600 billion total parameters, but only 27 billion active per token. It's the same design principle used by other recent large-scale models: a router chooses, for each token, which subset of "experts" processes the input, reducing the computational cost per inference without giving up the model's total capacity.

According to the model's page on OpenRouter, Step 5 Preview performs strongly in software engineering and professional knowledge work, with a specific highlight for finance. The model was designed for tasks that span large codebases and documents, using tools and refining results across multiple steps.

Step 5 Preview's page on OpenRouter shows a 1M-token context, input price of $1/M, output of $2.70/M, and release date of October 8, 2026
Step 5 Preview's page on OpenRouter shows a 1M-token context, input price of $1/M, output of $2.70/M, and release date of October 8, 2026. Reprodução: openrouter.ai.

What a 1M-token window allows in practice

The context window is 1,000,000 input tokens, with up to 64,000 output tokens. The model accepts text, image, and video as input and returns text; it also supports tool calling and structured output via JSON schema.

In practice, 1M tokens gives enough room to paste an entire medium-sized monorepo, an extensive technical documentation set, or the complete history of a debugging thread in a single call, without needing to split the context across multiple requests. It's this scenario, more than short conversations, that justifies the investment in long-context MoE architectures.

Pricing and performance as measured by OpenRouter

Step 5 Preview is hosted by a single provider on OpenRouter (StepFun itself), with no routing between multiple vendors.

MetricValue
Input$1.00 / million tokens
Output$2.70 / million tokens
Cache read$0.05 / million tokens
Latency (P50)3.48 s
Throughput56 tokens/s
Uptime (3 days)99.62%
Availability (3 days)99.37%

The input price competes directly with Western long-context models, and the $0.05/M cache read makes repeated calls over the same large document even cheaper, a common case for agents that re-examine the same codebase at every step.

The trade-off: a huge context isn't guaranteed quality

The OpenRouter page doesn't include independent quality benchmarks for Step 5 Preview, only the manufacturer's description. That matters: it's well known in the LLM literature that models with very large context windows tend to suffer from "lost in the middle," losing accuracy on information positioned in the middle of a long prompt, even when they can technically process that volume of tokens.

For anyone evaluating the model on a real project, the practical recommendation is to test it with your own material, not to rely solely on the 1M-token figure. Pasting an entire repository is possible; extracting the right answer about a specific function buried in the middle of the context is a different matter, one that only testing in the concrete use case can answer.

The throughput of 56 tokens/s and the latency of 3.48s (P50) also deserve attention: for agents that make multiple calls in sequence, this response time accumulates and can weigh more heavily on total execution time than the size of the available context.

Who is already running the model

OpenRouter lists the public applications that send the most traffic to Step 5 Preview, a signal of what kind of workload the model is being used for in practice:

  • Hermes Agent (Nous Research), an open-source agent with persistent memory across sessions: 990 billion tokens
  • OpenClaw, an agent that connects to messaging apps and executes real actions: 313 billion tokens
  • Kilo Code, an open-source coding agent for VS Code, JetBrains, and CLI: 268 billion tokens
  • Cline, an AI agent that lives inside the editor and explores the codebase autonomously: 251 billion tokens
  • StepCode, a new app from StepFun itself: 74.5 billion tokens

The volume led by coding and automation agents reinforces the model's positioning for agentic engineering tasks, rather than simple conversational chat.

StepFun's other options

StepFun also maintains two smaller, cheaper models in the same family. Step 3.7 Flash combines a 196-billion-parameter backbone with a vision encoder, activating about 11 billion parameters per token, with a 262,000-token window and selectable reasoning levels (high, medium, low). It costs $0.20/M input and $1.15/M output.

Step 3.5 Flash, in turn, is the company's most capable open-source model, also an MoE with 196 billion total parameters and 11 billion active, geared toward efficient reasoning in long contexts. It goes for $0.10/M input and $0.30/M output.

What changes for developers in Brazil

For teams that already use OpenRouter as an abstraction layer over multiple providers, switching models is a matter of changing the slug string in the API call, without rewriting the integration. That lowers the barrier to testing Step 5 Preview alongside other long-context options already in use, such as Gemini or Claude, and comparing cost per real task.

The practical point of attention is that Step 5 Preview has a single provider hosting the model on OpenRouter, with no redundancy across vendors. That means spikes in StepFun's downtime directly affect the application, with no automatic failover to another provider, unlike what happens with models that have multiple hosts on the same platform.

Translated from the Brazilian Portuguese original · Read the original