NEWS

Spanner Omni reaches general availability and runs outside Google's infrastructure

Google made Spanner Omni, the version of its distributed Spanner database that swaps the Colossus file system and TrueTime's atomic clocks for software equivalents, generally available. The trade-off: no availability SLA comes with it.

Google announced the general availability (GA) of Spanner Omni, the "deploy anywhere" version of its distributed database Spanner, according to a report from InfoQ. The product runs in the customer's own data center, in another cloud, or even on a laptop, and the engineering behind that is what makes this GA different from any other managed database launch: the original Spanner depends on two components exclusive to Google's infrastructure, the distributed file system Colossus and the time service TrueTime, based on atomic clocks and GPS. Taking Spanner out of there meant replacing both with software.

For those evaluating global databases with cross-region replication, Spanner Omni changes the calculus: it's no longer "managed Spanner versus another database," it's "managed Spanner versus Spanner Omni," with the price paid in tail latency and operational effort, not in feature parity.

What was replaced with software

In place of Colossus, Spanner Omni introduces what Google calls a Colossus-like abstraction layer: it writes to attached local file systems and exposes them over the network to other nodes, with automated shard splitting and rebalancing. Google is upfront about the limitation: this layer is not Colossus, but it works as a sufficient substitute to deliver performance comparable to the managed service for most workloads.

TrueTime got the same treatment. The software alternative offers time synchronization with a bounded error margin across servers, the same way the original does with atomic clocks and GPS. Google's explanation for why this works lies in how Spanner already behaves: the database overlaps time-uncertainty waits with other work, so it tolerates weaker uncertainty bounds than TrueTime delivers in practice. It's that slack that allows measuring time on heterogeneous hardware without limiting availability or performance.

Paxos consensus, automatic sharding, and synchronous replication remain unchanged, and Google's internal benchmarks claim millions of queries per second across petabytes of data in a single regional deployment.

The reaction from those who will operate it

Practitioners' reaction has focused on what this means for operating the system day to day. Carlos Pérez Martín, CTO of Q2BSTUDIO, argued that the change is not primarily a matter of architecture:

The interesting shift is operational, not architectural: once the same engine runs in your racks, the failure domains become yours.

Carlos Pérez Martín, CTO of Q2BSTUDIO

He pointed to three practical consequences. Quorum and witness topology needs to be recalculated for the local latency budget, and when an entire data center becomes the unit of failure, it's p99, not the average, that the application feels. Patching, versioning upgrades, and rolling back stop being another company's ticket queue and become a change-management problem with the customer's name on the label. And moving data residency to a private data center brings along the burden of backup, key management, and auditing.

The test he recommends is narrower than a feature comparison:

A like-for-like pilot against the managed service, measured on tail latency and ops toil instead of feature parity, is the cheapest way to price that trade.

Carlos Pérez Martín, CTO of Q2BSTUDIO

No SLA, with topologies in its place

Google's own documentation supports this reading. Because Spanner Omni runs on infrastructure managed by the customer, Google doesn't offer an availability SLA, and instead points to reference architectures that would help achieve comparable high availability. In place of the SLA come topologies:

  • Single server: upgrades cause downtime.
  • Single zone: minimum of three servers.
  • Multi-zone: at least three zones with three servers each.
  • Multi-cluster: three zones across two or more clusters.

For those operating in Brazil, where Google Cloud maintains only the São Paulo region, this list matters in a specific way: Google itself cites multi-regional high availability as a use case in jurisdictions where it operates only one data center but data sovereignty is required. That's exactly the scenario for anyone who needs to keep data in the country while also replicating across distinct physical facilities without depending on a second Google region here.

The operational load also shows up in the tooling. Routine maintenance, version upgrades, and infrastructure monitoring become internal responsibilities, with monitoring via Prometheus alerts and Grafana dashboards instead of Cloud Monitoring, plus a diagnostic command that collects logs, traces, and thread stacks.

Limits that weigh on the evaluation

Two boundaries matter when deciding whether to migrate. Integrations that depend on Google Cloud are left out, including BigQuery, Knowledge Catalog, and Gemini Enterprise, and Google acknowledges remaining feature gaps relative to the managed service, without naming which ones or when they'll close. The documentation lists Google Cloud and Amazon as supported public clouds, in addition to on-premises installations and laptops; Azure isn't mentioned, which weighs on anyone who has already standardized critical infrastructure on Microsoft's cloud.

What the database keeps

What doesn't change is the database itself. Spanner Omni supports GoogleSQL, PostgreSQL, and Spanner Graph Language, and combines relational, graph, key-value, and vector models with full-text search and a columnar engine. Support for the Model Context Protocol via MCP Toolbox lets agents inspect schemas and use the database as an operational memory layer across deployments.

The GA release adds TLS encryption, authentication and authorization, audit logs, backup and restore, and worker nodes, stateless compute nodes that offload background operations from the primary servers.

Licensing in two tiers

Licensing has two tiers. The Developer Edition is free for non-production use, with a standard 90-day license that excludes backups and worker nodes, although a single-server deployment with 4 vCPUs or fewer doesn't expire and includes backup and restore; extending beyond 90 days requires filling out a Google form. The Commercial Edition is an annual subscription based on vCPU count, with no published price.

Who's already using it

Deployment patterns reported by early adopters confirm the operational framing. One company runs managed Spanner as its primary database and Spanner Omni elsewhere as a hot-cold failover. Another standardizes a single database layer across different environments. A third modernizes on-premises systems using existing hardware.

Mercado Livre (Latin America's largest e-commerce marketplace), which built an internal developer gateway offering a NewSQL service based on Spanner, made the resilience argument when the preview version launched in April 2026. Senior technical manager Diego Oscar Narducci stated that the company had maintained long-standing vigilance against insider threats, ransomware, and cloud outages, and that Spanner Omni enables cross-cloud resilience that would be significantly more complex to implement with other cloud providers.

One timeline detail marks the rush of the launch: the Spanner Omni overview in Google's documentation was last updated on September 28, 2026, two days before the GA announcement, and still carried the Preview label with limitations the final version removes, including the absence of TLS encryption, of backups and restores, and deployments that stopped accepting writes after 90 days.

Translated from the Brazilian Portuguese original · Read the original