AIARTICLE

OpenAI announces GPT-6.1 Sol with performance close to Astra at a fifth of the price

GPT-6.1 Sol arrives as the successor to GPT-6 Sol and narrows the gap to GPT-6 Astra in coding, computer use, and professional tasks, charging about a fifth of Astra's price per token.

OpenAI has introduced GPT-6.1 Sol, an update to GPT-6 Sol that promises performance close to GPT-6 Astra (the family's flagship model) in agentic coding tasks, computer use, and professional work, charging about a fifth of the price per token charged by Astra in standard mode, according to the official OpenAI News post. For those building AI products, this changes the math on which model to use in production: Astra is no longer the only viable option when the budget per task matters.

The price math, token by token

The official GPT-6.1 Sol numbers on the API are straightforward: $2 per million standard input tokens, $0.10 per million cached input tokens, and $10 per million output tokens. OpenAI states these figures equal a fifth of Astra's standard input and output price. Working the math backward, this suggests Astra runs at around $10 per million input tokens and $50 per million output tokens in standard mode (OpenAI doesn't disclose this figure directly; it's inferred from the reported ratio).

Input caching is the detail that matters most to those building agents: $0.10 per million tokens is 95% cheaper than standard input and 50% cheaper than GPT-6 Sol's own cache, which by this math ran around $0.20. In practice, this lowers the cost of keeping context alive across calls: an agent that reprocesses the same conversation history, the same RAG documents, or the same system prompt at every step benefits directly, because that's exactly the kind of repeated token that falls into the cache.

Where the gain shows up: coding and complex PDFs

On DeepSWE v1.1, a benchmark that evaluates software engineering tasks on real codebases, GPT-6.1 Sol ties with Astra at about a fifth of the cost, and still beats GPT-6 Sol's best result by 6.4 percentage points, with lower reasoning effort (and cost). For those using the model inside Codex or internal code copilots, this is the metric that matters most: refactoring or bug-fixing tasks on a real repository, not a synthetic benchmark.

On GDP.pdf, which measures reading of complex professional documents (tables, charts, fine print) in finance, healthcare, legal, and other fields, GPT-6.1 Sol scores above Opus 5.5 with fallbacks, at less than half the cost per task, and comes close to Astra's top-end result at about a fifth of the price. This is the use case of extracting data from contracts, reports, and financial statements, where cost per document processed at scale is usually the deciding factor.

Workflow agents and computer use

On AutomationBench, which tests agents in complete workflows using 47 tools across sales, marketing, operations, support, finance, and HR, GPT-6.1 Sol comes in 2.2 percentage points above Opus 5.5 at medium reasoning effort, at about a third of the cost, and rises 4.8 points over GPT-6 Sol in the same configuration. OpenAI itself makes an important caveat about the competition in this benchmark: the Claude Fable 5.1 result understates the real cost, because it doesn't count the fallbacks that occurred in about 40% of the tasks.

On OSWorld 2.0, which evaluates agents in long-running workflows interacting with real computer applications, GPT-6.1 Sol beats GPT-6 Sol by seven percentage points at maximum reasoning effort, at less than half the cost, and comes within 2.1 percentage points of Astra at about a seventh of the price per task. This is the scenario closest to an agent that opens applications, fills out forms, and navigates real interfaces, not just answering text.

Where Astra still wins

Not everything tips toward Sol. On Terminal-Bench Science 0.1, which evaluates scientific workflows such as data analysis, simulation, and theorem proving, GPT-6.1 Sol more than doubles GPT-6 Sol's result at maximum effort, at less than half the cost per task: $5.47 on average, against $23.21 for Opus 5.5 and $23.80 for Astra. Even so, Astra still holds the highest score among the models tested (68.1%), and OpenAI itself recommends using it for the hardest scientific research tasks.

This is the distinction that matters when choosing a model: for heavy scientific research, Astra remains the right choice despite the price; for everyday coding, document reading, and workflow automation, Sol 6.1 covers most of the ground at a fraction of the cost.

Factuality and safety

OpenAI reports improvement in the factual error rate on difficult prompts, measured in ChatGPT conversations where users flagged an error from a previous model (a deliberately difficult set, not representative of typical use). At low reasoning effort, the rate of responses with at least one factual error dropped from 11.4% to 7.7%, a reduction of about 32%. In the tested configurations, the gap to Astra was as small as 1.9 percentage point, at less than a fifth of the cost.

In alignment tests, GPT-6.1 Sol made fewer mistakes than GPT-6 Sol in transparency about broken tools, respecting explicit restrictions, and avoiding unauthorized actions during agentic tasks. In the specific test of flagging when a search tool is broken instead of guessing an answer, the failure rate was 2.1% for Sol 6.1, against 4.9% for GPT-6 Sol, 1.5% for Astra, and 28.7% for GPT-6 Luna. None of the first three models attempted to bypass the automated safety reviewer in the tests.

How to use it and what's missing

GPT-6.1 Sol is already available for Plus, Pro, Business, Enterprise, and Edu accounts within ChatGPT Work and Codex, but it hasn't reached regular Chat yet. On the API, the model identifier is gpt-6.1-sol, and the call follows the pattern already familiar from any other model in the GPT-6 line:

curl https://api.openai.com/v1/chat/completions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-6.1-sol",
    "messages": [{"role": "user", "content": "revise este patch e aponte riscos"}]
  }'

OpenAI also announced it will release, in the coming days, an Ultrafast variant of GPT-6.1 Sol, with token generation up to 8 times faster than the standard speed within Codex, useful for those running agents that need near real-time responses during pairing sessions with the model.

When it's not worth switching

If the product already runs on GPT-6 Sol and doesn't depend on reprocessing a lot of repeated context, the gain from switching to 6.1 tends to come more from quality (the extra percentage points in coding and computer use) than from price, since the previous model was also cheaper than Astra.

For those already using Astra and needing the best possible result in heavy scientific research, OpenAI's own recommendation is to stick with it. Outside these two cases, the most direct reading of the announcement is that the cost floor for running an agent with near-flagship quality has dropped considerably, and it's worth testing in any pipeline that today pays Astra's full price just for safety's sake.

Translated from the Brazilian Portuguese original · Read the original

View profile →
Read also