Learning on the Job: The Future of Post-Training — Raymond Feng, Applied Compute

AI Engineer18mJul 31, 2026
Watch Original (opens in new tab)
0:00 / 18:20
Chapters8

No clickbait detected — the title and thumbnail deliver what they promise.

AI Opinion

Raymond Feng persuasively demonstrates how enterprise needs are driving a shift toward more adaptable and deployable AI agents, particularly through post-training frameworks mimicking human learning processes. The discussion around synthetic environments for tool usage and complex interactions feels well-supported by the described technical setup. However, the assertion that experiential data will soon surpass human-labeled datasets in importance requires further scrutiny given current trends in large language model training; it would be prudent to investigate the specific metrics Applied Compute uses to measure "experience." Ultimately, while the framework’s potential is clear, listeners should consider whether the claimed timeline for self-improving agents aligns with broader industry developments.

Avatars are AI rewrites of the same facts — style changes, not substance.

Summary

Raymond Feng of Applied Compute discusses a new approach to post-training for AI models, driven by enterprise demand for easily deployable ("plug-and-play") agents adaptable to varied and often proprietary systems. Their framework mirrors human learning, beginning with foundational Q&A tasks and progressing to complex scenarios using synthetic environments that allow for tool usage and multi-turn interactions. This process involves a continuous loop of graded interaction data used to update model weights, ultimately aiming for models capable of self-evaluation and introspection—essentially treating themselves as entities to be improved through every interaction. Feng argues this shift moves beyond addressing individual task failures ("Whac-A-Mole") toward proactive, holistic self-optimization and emphasizes the growing importance of experiential data in AI training, suggesting it will soon surpass human-labeled datasets.

Avatars are AI rewrites of the same facts — style changes, not substance.

Key Points

00:53

The Need for Adaptable Post-Training

Raymond Feng explains that enterprises increasingly desire 'plug-and-play' agents, meaning they want to train custom models within their existing agent workflows. This necessitates post-training methods capable of adapting to diverse harnesses – the code and infrastructure surrounding an agent – even when source code access is limited. The ability to adapt to these external systems is a key challenge in modern AI development.

01:36

Post-Training Framework Mirrors Human Learning

Applied Compute frames their post-training approach as analogous to human learning, emphasizing the importance of building foundational skills before tackling more complex tasks. This involves starting with simple, single-turn Q&A and gradually progressing to longer horizon synthetic environment tasks that require higher-order reasoning and tool usage. The framework prioritizes compounding understanding through iterative skill development.

02:54

Q&A Training Stack Components

The basic Q&A training setup involves an orchestrator that holds a task (prompt and answer), sends it to a model, receives the response, and then submits it for grading. A training engine uses these graded interactions to compute weight updates which are synced to inference engines, creating a continuous loop of improvement. The key is having graded chats in a specific format within a controlled environment.

04:52

Synthetic Environments Extend Training Capabilities

To handle longer and more complex tasks, Applied Compute utilizes synthetic environments where the environment state resides outside of the training stack. This allows for multi-turn interactions involving tool calls and sandbox modifications. The resulting task traces are then graded and used to update model weights via GRPO (Generalized Reinforcement Proximal Policy Optimization), enabling learning in more dynamic scenarios.

15:29

Model as a Self-Improving Agent Across Environments

Raymond envisions a future where models aren't limited to improving on specific tasks but instead function as agents capable of interacting and learning across diverse environments. This involves the model continuously evaluating its performance in different settings, essentially treating itself as an entity to be improved upon through every interaction it has. This shift moves beyond task-specific improvements towards a holistic self-optimization strategy.

16:03

Self-Evaluation and Introspection for Continuous Improvement

A crucial element of this future is the model's ability to perform introspection – assessing its performance based on different interaction types. The system would automatically compute weight updates from these self-evaluations, leading to continuous improvement without explicit human intervention. This allows for a more dynamic and adaptive learning process compared to traditional methods.

17:00

Overcoming the 'Whac-A-Mole' Problem in Model Improvement

Raymond highlights the current challenge of improving models as a “Whac-A-Mole” game. Focusing on individual tasks or failure modes requires constant data creation and model adjustments whenever new issues arise. A self-improving system, understanding all interactions within its environment, would circumvent this problem by proactively addressing emerging challenges.

17:34

The Ascendancy of Experience in AI Training

Raymond concludes with a quote from a recent paper emphasizing the shift towards experience as the primary driver of AI improvement. He believes that experiential data will soon surpass the scale of human-labeled datasets currently used, marking a new era in AI development where continuous interaction and learning become paramount.

Chapters

8 chapters · 8 key moments
KEYkey momentNot checkable hereUnverifiedPartially supported

Claims & Fact Check

Enterprises want agents that can be deployed 'plug-and-play' into their existing workflows.

Not checkable here

The training setup requires graded chats in a very specific format.

?Unverified

Synthetic environments allow for more complex, multi-turn interactions and tool usage.

±Partially supported

Models will eventually be able to continuously improve themselves across diverse environments.

Not checkable here

Focusing on individual tasks leads to a 'Whac-A-Mole' problem in model improvement.

Not checkable here

Experience will soon dwarf the scale of human data used in AI systems.

Not checkable here

Was this digest good?

More from AI Engineer

Digest any single YouTube video — free.

3 free digests — no card, no sign-up wall.

Or just swap the domain of any YouTube link → instant digest