From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AI

AI Engineer20mJul 25, 2026
Watch Original (opens in new tab)
0:00 / 20:24
Chapters8

No clickbait detected — the title and thumbnail deliver what they promise.

AI Opinion

Feyzkhanov makes a compelling case that organizations must develop production-like benchmarks to reliably evaluate and improve agent systems, particularly highlighting the value of offline simulation for repeatable experimentation. While the concept of creating private benchmarks tailored to specific workflows is likely valid, the assertion that public benchmarks are merely useful for initial orientation feels somewhat dismissive of their potential for collaborative learning and standardization. Listeners should be mindful that the feasibility of fully automating benchmark creation with LLMs, as suggested, remains an area requiring significant refinement and validation against expert judgment.

Avatars are AI rewrites of the same facts — style changes, not substance.

Summary

Rustam Feyzkhanov discusses the critical role of benchmarks in developing and deploying reliable agent systems. He emphasizes that every organization should create a production-like benchmark to evaluate agents, moving beyond simple pass rates to consider factors like cost, latency, and retry counts. Offline simulation allows for repeatable experiments by transforming real-world traces into controlled environments, enabling comparisons between different configurations. These benchmarks serve three key purposes: evaluating agent functionality, integration testing during releases, and providing training data for continuous improvement through a two-loop system that incorporates observability data to expand the benchmark’s scope. Feyzkhanov advocates for a train/test split approach in benchmark creation, ensuring robust evaluation by including both common use cases and edge cases, with subject matter expert review crucial for resolving discrepancies between automated assessments and human judgment.

Avatars are AI rewrites of the same facts — style changes, not substance.

Key Points

00:47

The Importance of Production-Like Benchmarks

Rustam emphasizes that every company needs a benchmark to reliably evaluate, release, and improve their agents. This benchmark must closely resemble the production environment, including real tools, API services, policies, and workflows. It's not merely about model performance but ensuring the entire agent system functions effectively within the organization’s specific context.

02:38

Offline Simulation for Repeatable Experiments

Rustam introduces offline simulation as a method to transform production traces into repeatable experiments. This allows for comparing different agent configurations and metrics (cost, latency, retries) in a controlled setting, avoiding the inconsistencies of real-time environments due to fluctuating database states or tool versions. This ensures more accurate comparisons between agent variants.

04:10

Beyond Pass Rate: Comprehensive Agent Evaluation

While public benchmarks often focus on pass rate, Rustam argues that production releases require a broader evaluation. Companies need to consider cost, latency, and the number of retries alongside success rates. Offline simulations enable this comprehensive assessment by allowing for controlled testing of the entire agent stack, including prompting strategies and available tools.

06:09

Trifecta of Use Cases: Evaluation, Integration Testing, and Training

Rustam outlines three core uses for benchmarks within an organization. Firstly, they serve as agent evaluation tools to ensure functionality and handle edge cases. Secondly, they act as integration tests for releases, preventing regressions. Finally, these traces can be used as training data to improve the agents themselves, demonstrating a closed-loop system.

15:12

Two-Loop System for Production Agents

The speaker describes a two-loop system for agents in production. The first loop focuses on benchmark expansion, utilizing observability traces (recorded via tools like Rise) to identify and incorporate failures into the benchmark. The second loop involves running experiments with new agent configurations against this extended benchmark to evaluate performance before release.

17:23

Train/Test Split for Agent Configuration Verification

The speaker recommends a traditional machine learning approach of an 80/20 train-test split when structuring benchmarks. This allows for the creation of a standalone dataset that the agent hasn’t seen, enabling verification of agent configurations and ensuring robust evaluation beyond training data.

18:26

Benchmark Should Cover Both Common Use Cases & Edge Cases

Effective benchmarks should encompass both 'bread and butter' use cases – representing typical, well-functioning scenarios – and edge cases. Edge cases are crucial for testing an agent’s ability to handle unexpected situations like tool failures or database issues, ensuring overall system resilience.

19:47

Importance of Subject Matter Expert Review for Discrepancies

When using both LLMs and human experts as verifiers, the most valuable use case is identifying disagreements between them. These discrepancies highlight areas where the agent's performance or assessment of a trace may be incorrect, requiring further tuning by subject matter experts to improve accuracy.

Chapters

8 chapters · 8 key moments
KEYkey momentNot checkable herePartially supportedUnverified

Claims & Fact Check

Every company needs a benchmark.

Not checkable here

Offline simulation turns traces into repeatable experiments.

±Partially supported

Public benchmarks are useful to orient and build your prior, but your private benchmark is useful to ship.

Not checkable here

Companies need a benchmark.

Not checkable here

Traces are useful for finding edge cases in production.

Not checkable here

A classic train/test split of 80/20 is generally applicable for agent benchmarks.

±Partially supported

Automating benchmark creation with LLMs and handcrafted elements is possible.

?Unverified

Was this digest good?

More from AI Engineer

Digest any single YouTube video — free.

3 free digests — no card, no sign-up wall.

Or just swap the domain of any YouTube link → instant digest