Uber Eats rebuilds search pipeline and cuts latency by 50%
The company swapped the metric that measured speed, reduced retrieval and ranking work, and used agentic AI to find optimizations. The result: half the end-to-end time in Uber Eats search.
Uber reworked core parts of the Uber Eats search pipeline and reports a 50% reduction in end-to-end latency. The change, reported on October 2, 2026 by InfoQ, touched retrieval, feature hydration, ranking, advertising, presentation, and infrastructure. Part of the optimizations was identified and validated with an agentic coding workflow.
For those building search systems in production, the case is interesting less for the final number and more for the process: Uber changed the metric that measured speed before optimizing any line of code, and only then started cutting unnecessary work layer by layer.
The metric changed before the architecture
The first move wasn't technical, it was about measurement. Instead of tracking backend API response time, Uber started measuring Above-the-Fold completion: the time until the first screen of results is rendered with images. That's the experience the user actually feels, not what the server reports as latency.
With the right metric, two presentation changes already produced a direct gain:
- Pagination with server-side caching reduced the size of the initial response.
- Asynchronous rendering started processing result items in parallel instead of sequentially.
Together, these two changes improved Above-the-Fold by more than 200 milliseconds, according to Uber.
Where the milliseconds went in retrieval and ranking
When investigating the pipeline, the team found that tens of thousands of candidates were hydrated before ranking, and most were discarded unused. Removing low-value retrieval strategies cut about 120 milliseconds. Product-level embeddings reduced data queries by more than 100x, saving another 50 milliseconds.
In ranking, separating hydration used for scoring from hydration used for presentation reduced latency by more than 100 milliseconds. Removing unnecessary dependencies and request hedging (firing the same call in parallel and using the fastest response) contributed 35 and 40 milliseconds, respectively.
The advertising path, which runs in parallel with organic ranking, was redesigned with bid data organized by column, in-memory access, and less serialization, cutting about 130 milliseconds. On the infrastructure side, Uber applied parallel encoding, smaller embeddings, connection management improvements, and changes to Go data structures to reduce garbage collection overhead.
Agentic AI entered the optimization process
Beyond the architectural changes, Uber used an agentic coding workflow to identify, benchmark, and validate additional optimizations in the pipeline. InfoQ doesn't detail which tool was used or the volume of optimizations attributed specifically to this workflow, but the data confirms a pattern that already appears in other reports of large-scale performance engineering: code agents are used to sweep hotspots and propose fixes, with humans validating the result before going to production.
Three principles, not a rewrite
Engineers who publicly commented on the work reinforce that there wasn't a single architectural decision behind the gain. Pratik Dhanave described it this way:
No single big idea behind it, but a long list of careful decisions across the full stack.
Pratik Dhanave
Anubhooti Nagar summarized the performance challenge in a similar way:
It's less about doing things faster and more about doing less work and avoiding unnecessary waiting.
Anubhooti Nagar
Nagar also highlighted Uber's Measure, Identify, Fix, Validate loop as a model for continuous performance optimization, rather than a single tuning project.
Vidya Pandey condensed the approach into three principles:
Do less work. Start work earlier. Remove unnecessary dependencies.
Vidya Pandey
Pandey also connected the microbatching Uber is planning with techniques used in AI systems to reduce synchronization between processing stages, a parallel worth noting: the same pipeline ideas used to train and serve models are coming back to optimize traditional search.
What's still being tested
The changes described build on Uber's existing search platform, built on Apache Lucene, Spark-based indexing, streaming updates with Kafka, and a distributed serving layer. It's on this foundation that the company is now exploring four new fronts:

- End-to-end microbatching, allowing processing stages to overlap instead of waiting for the previous stage to finish completely.
- Product-based retrieval, instead of restaurant-based.
- Zero Pass Ranking, reducing ranking passes.
- Multipart HTTP streaming.
Uber reports that initial tests of product-based search have already produced a reduction of more than 50% in p99 latency, a figure separate from the 50% gain in end-to-end latency already in production. Neither metric was detailed with a general availability date.
What this changes for anyone scaling search
The Uber Eats case doesn't bring a new recipe, but it validates a roadmap any high-volume search team can apply: measure the end-user experience instead of backend response time, audit how many candidates reach ranking unnecessarily, and separate what is scoring data from what is presentation data. In lower-volume systems, the absolute gain in milliseconds may be irrelevant, but the method of attacking excessive hydration and unnecessary dependencies applies even at a much smaller scale than a global delivery marketplace.
For those deciding on a stack today, the detail that stands out most is the use of product-level embeddings to cut data queries by more than 100x, a technique that doesn't depend on Uber's scale to work: reducing a multi-attribute search to a single vector comparison is a latency gain any product catalog can try to replicate with the vector tools already available today.
Translated from the Brazilian Portuguese original · Read the original
Epic pauses new product development to fix MyChart security flaws
Epic, the company behind MyChart, suspended most of its product development after an Anthropic AI model found flaws that allowed access to patient records without leaving a trace in the logs.