Interactions API becomes the standard in Gemini, and generateContent enters legacy mode
The official Gemini API documentation, updated on October 6, 2026, confirms that the Interactions API has been the standard interface since June and that the old generateContent is now legacy. Anyone maintaining agents or RAG built for Gemini 1.5/2.0 has code adjustments ahead, not just a model upgrade.
An X-ray of the documentation, not a launch announcement
The Gemini API documentation page (Google AI for Developers), with an update logged on October 6, 2026, isn't announcing a recent launch. It shows an already mature model catalog: Gemini 3.8 Flash as the newest in the Flash line, Gemini 3.1 Pro as "the most intelligent," Gemini 3.5 Flash-Lite for high-volume, low-cost tasks, plus Nano Banana (image) and Omni Flash (video).

This means that if the story angle is "Gemini 3 arrives with thinking and 1M tokens," that specific announcement is already old news: what the current documentation reveals is how that feature has become established and, more importantly, a structural change in the API itself.
For those who wrote agents or RAG pipelines back in the Gemini 1.5 or 2.0 era, the shock isn't in the model's name. It's in how you talk to it.
generateContent is now legacy: what that means in practice
The documentation is direct: "the Interactions API has become our standard interface since June 2026 and is the best way to build with Gemini models and agents going forward. If you're starting a new project, you should use the Interactions API. While still supported, the generateContent API is now considered legacy." There's no shutdown date announced on this page, but the legacy label is usually the first warning before a deprecation cycle.
In practice, anyone who built agents with generateContent() in the google.genai SDK (or its equivalent in JS, Java, Go) needs to review that code in light of the official migration guide, listed in the doc itself as "Migration Guide." The calling pattern shifts from a function that takes a prompt and returns text to an interactions object that manages conversation state, messages, and output formats more explicitly.
# Old style (generateContent, now legacy)
response = model.generate_content("Explain how AI works in a few words")
print(response.text)
# Current style (Interactions API)
from google import genai
client = genai.Client()
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="Explain how AI works in a few words"
)
print(interaction.output_text)The difference isn't just cosmetic. The Interactions API also exposes event Streaming: not just incremental text tokens, but partial "thoughts" and tool-call events, something the single-response model of generateContent didn't cover in the same way.
The "thinking" that's already built in, with no spotlight
The documentation lists "Thinking" as its own capability category, alongside Long Context and Function Calling: "explore how reasoning capabilities improve performance on complex tasks and agents." It's no longer presented as a standalone preview feature; it's treated as part of the standard package of Gemini 3.x models.
This matches the positioning of Gemini 3.8 Flash, described as "our most intelligent Flash model, designed for long-horizon software engineering, autonomous agents, and complex enterprise workflows." In other words: multi-step reasoning stopped being a differentiator of an expensive model and became a baseline expectation even in the family's cheapest tier.
For those who built agents on Gemini 1.5/2.0 calling tools in a manual loop (parsing the response, deciding the next action, making a new call), the practical question is: how much of that handcrafted orchestrator can be replaced by the current model's native thinking? The documentation doesn't provide a benchmark for how much this built-in reasoning reduces API calls or total latency, so that's a calculation each team needs to run on its own use case before deciding.
Context window: less magic number, more fine print
Here's an honesty flag worth raising. The "Long Context" section of the current documentation says only: "feed millions of tokens into Gemini models and extract understanding from images, videos, and unstructured documents." There's no fixed figure of 1 million tokens nailed down on this page for the 3.x generation specifically.
That doesn't mean the window shrank. Gemini 1.5 Pro was already known for operating with up to 1 million tokens of context, and the plural phrase "millions of tokens" suggests the ceiling may have even gone up. But anyone building RAG with large context windows as a competitive edge should check the technical card for the specific model (3.8 Flash, 3.1 Pro) before promising an exact number to the product team, rather than assuming the previous generation's limit simply carried over.
The documentation also reinforces the "Document Understanding" capability: processing up to 1,000 PDF pages with full multimodal comprehension, or other text-based formats. For RAG pipelines that today split large documents into chunks because of context limitations, this raises the question of when manual chunking still pays off versus simply feeding the entire document for the model to process.
Managed agents: the Antigravity example
The documentation presents a feature called "Managed Agents," which gives Gemini a work environment (remote sandbox) to plan and complete tasks on its own, including writing and running code, searching the web, and creating files. The official example uses a pre-built agent called antigravity-preview-09-2026:
agent = "antigravity-preview-09-2026"
input = (
"Research the top 5 sustainable fashion brands, "
"compare their materials and pricing tiers, and "
"build an interactive dashboard in analysis.html."
)
environment = "remote"The agent's name (preview-09-2026) indicates it's still a preview feature, not GA. For those currently orchestrating multiple specialized agents by hand (one for search, another for synthesis, another for artifact generation), this is a sign that part of that orchestration can move inside the API itself, trading glue code for environment configuration.
For those who already had agents on 1.5/2.0: what to do now
Three concrete steps are worth taking before any major refactor:
- Map all code that still calls
generateContentorgenerate_contentand read the official Migration Guide point by point, since the legacy API remains functional but without a roadmap for new features. - Swap the model name in the calls (from
gemini-1.5-proorgemini-2.0-flashto something likegemini-3.1-proorgemini-3.8-flash) and run the same regression prompts to compare output quality before switching in production. - Test the Interactions API's event streaming in a real agent flow, watching whether partial "thinking" events help debug model decisions that used to be a black box.
This is the path that would make sense to follow in a typical project, not a measured outcome: the documentation doesn't provide latency or cost numbers comparing the two generations, so any gain needs to be validated with your own traffic before it becomes an argument for the team.
When it's not worth migrating right now
If the agent in production is small, stable, and doesn't depend on advanced thinking or a larger context window, the "legacy" label alone isn't enough reason to stop everything and migrate today. The documentation itself confirms that generateContent "remains supported." It is worth, however, starting to budget engineering time for the migration, since the industry's track record with legacy APIs usually ends in a shutdown deadline, it's just that one hasn't been announced on this page yet.
Translated from the Brazilian Portuguese original · Read the original
Mistral Large 4 arrives in preview with 1 trillion parameters, but open weights only by the end of October
The benchmark leap is real, but the open weights, which underpin the promise of running on-premise, are only expected by the end of the month.