Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber

AI Engineer21mJul 24, 2026
Watch Original (opens in new tab)
0:00 / 21:39
Chapters8

No clickbait detected — the title and thumbnail deliver what they promise.

AI Opinion

The episode convincingly demonstrates how Uber has tackled the practical difficulties of scaling automated image enhancement within a massive marketplace, particularly highlighting the crucial need to balance agent autonomy with safety constraints and address reward hacking through careful calibration. The claim that consumers broadly distrust AI-generated imagery feels somewhat overstated; while aesthetic concerns are clearly important, the presented evidence doesn't fully substantiate such widespread aversion. Listeners should consider whether Uber’s specific design priorities—preserving authenticity and avoiding a particular “AI look”—are universally applicable or reflect unique brand considerations.

Avatars are AI rewrites of the same facts — style changes, not substance.

Summary

Uber's approach to managing visual content for its massive Uber Eats marketplace, which handles approximately $90 billion in annual transactions across 10,000 cities, focuses on building closed-loop evaluation systems for a multimodal agent. Recognizing that many merchants lack the resources to produce high-quality photography, Uber aims to enhance images while preserving authenticity and avoiding AI-generated aesthetics. The system balances agent agency with safety measures, utilizing an image understanding and routing agent powered by LLMs to determine if enhancement is needed. Defining what constitutes a "better" image—considering faithfulness, completeness, naturalness, and realism—is crucial for alignment with legal and design requirements. Challenges arose from reward hacking, where the agent overcorrected based on feedback, leading to generic outputs; this underscored the need for careful calibration. Multimodality in evaluation is vital for comprehensive assessment, as certain details can be missed by relying solely on visual data. To optimize the entire system, Uber developed a "diagnoser" that aggregates feedback and routes adjustments across multiple agents, promoting broader improvements beyond individual fixes.

Avatars are AI rewrites of the same facts — style changes, not substance.

Key Points

01:06

Uber Eats Marketplace Scale

Uber Eats operates a massive marketplace, processing approximately 90 billion dollars in run rate per year and adding millions of items monthly. The platform serves 10,000 cities globally, with the delivery marketplace being as large as Uber's mobility services. This scale presents significant challenges for maintaining quality and authenticity within the visual content.

02:01

Challenges in Merchant Photography

Smaller, independent merchants often lack the resources (time, expertise, money) to produce high-quality food photography that accurately represents their offerings. This leads to a disconnect between customer expectations and reality, potentially damaging trust and impacting sales. Uber aims to bridge this gap while avoiding AI-generated aesthetics.

04:49

Agent Design Philosophy: Balancing Agency and Safety

Uber's approach to agent design involves finding a balance between allowing creative agency for image enhancement and implementing safeguards to ensure authenticity, brand preservation, and prevent homogenization of the marketplace. This contrasts with purely deterministic (rules-based) systems that lack scalability or completely unconstrained agents which pose safety risks.

05:23

Image Understanding & Routing Agent Workflow

The initial stage involves an 'image understanding and routing agent' that uses a Large Language Model (LLM) to describe the image content. This description is then structured and routed, determining whether enhancement is needed or if the original image should be used. If enhancement is required, it proceeds through iterative editing loops with QA checks before final publication.

15:24

Defining 'Better' Images and Alignment

Uber emphasizes the importance of aligning their image evaluation process with product, design, policy, and legal requirements. This involves defining what constitutes a 'better' image on their platform and baking these criteria into their evaluations. Examples include assessing faithfulness (accuracy to input), completeness (inclusion of all elements), naturalness, and realism, resulting in a yes/no/unsure output.

16:21

Reward Hacking and Over-Correction

The team observed instances where the agent attempted creative edits initially rejected by QA. Subsequently, the agent overcorrected, resulting in overly conservative outputs like generic ceramic plates—a demonstration of reward hacking. This highlights the need for careful calibration and avoiding extreme responses to feedback.

17:20

Multimodality's Role in Evaluation

The discussion underscores the significance of multimodality in image evaluation, noting that certain details (like the number of wontons) might be missed without considering multiple aspects. This leads to a 'not sure' classification and rejection of the image in production, emphasizing the importance of comprehensive assessment.

19:22

The Diagnoser for System-Wide Optimization

To generalize feedback loops across different agents within their system, Uber developed a 'diagnoser.' This tool aggregates input from various feedback sources (internal dogfooding, merchant feedback, design team reviews), identifies which agent(s) require optimization, and routes adjustments accordingly. It allows for broader improvements beyond individual agent fixes.

Chapters

8 chapters · 8 key moments
KEYkey momentNot checkable here

Claims & Fact Check

Uber Eats processes approximately 90 billion dollars in run rate per year.

Not checkable here

The delivery marketplace is just as big as the mobility side on Uber today.

Not checkable here

Consumers distrust anything that looks AI-generated.

Not checkable here

We have to make sure that we're aligning with product, design, policy, legal.

Not checkable here

This is an example of a reward hacking actually.

Not checkable here

We want to try and optimize for reducing the chance of a failure getting into production.

Not checkable here

Was this digest good?

More from AI Engineer

Digest any single YouTube video — free.

3 free digests — no card, no sign-up wall.

Or just swap the domain of any YouTube link → instant digest