
No clickbait detected — the title and thumbnail deliver what they promise.
AI Opinion
The episode persuasively argues that the industry is indeed shifting away from large, web-trained "base models" toward architectures prioritizing reinforcement learning and specialized capabilities like function calling; the demonstrated decrease in reliance on web text lends significant weight to this claim. However, the assertion that RL might eventually *overtake* supervised learning feels speculative, as it doesn’t fully address the ongoing need for broad knowledge grounding often provided by massive datasets. Listeners should consider how these new techniques balance specialized reasoning with maintaining a degree of general understanding and factual accuracy, which remains crucial for many applications.
Avatars are AI rewrites of the same facts — style changes, not substance.
Summary
The prevailing approach to developing large language models is undergoing a significant transformation, suggesting the traditional concept of a "base model" focused on broad world knowledge acquisition is becoming obsolete. Early base models heavily relied on vast datasets of web text, but newer models are increasingly incorporating reinforcement learning (RL) as a core component rather than an afterthought, leading to improved performance in tasks like application building and function calling. This shift has resulted in a substantial decrease in the proportion of web text used for training, with some models now utilizing less than 20%. While debates continue regarding the role of synthetic data, labs are also exploring novel techniques such as incorporating "novel" reasoning traces into supervised learning and optimizing model performance during inference to improve efficiency. Ultimately, base models are evolving beyond simply encapsulating general knowledge towards prioritizing reasoning abilities and agentic behavior to better suit their increasingly interactive roles.
Avatars are AI rewrites of the same facts — style changes, not substance.
Key Points
The Traditional Base Model Approach
Historically, base models were trained on vast amounts of web text like Common Crawl and WebText-2, aiming to capture a broad representation of human knowledge. This approach was dominant in early large language model development, with datasets like WebText comprising up to 85% of the training data for models such as GPT-3. The focus was on pre-training to accumulate world knowledge and build useful representations through next token prediction.
The Rise of Reinforcement Learning (RL) in Model Development
Recent advancements, particularly with OpenAI's 01 release and DeepSeek R1, have shifted the focus towards reinforcement learning. RL is no longer a mere 'cherry on top' for shaping model behavior but has become crucial for dramatically improving performance across various tasks. This shift enables models to interact with software environments, build applications through function calling, and learn complex interactions.
The Declining Importance of Web Text in Training Data
A significant change is the reduced reliance on web text as a primary training data source. While previously constituting up to 85% of the training mix, web text now represents only around 15% in newer models like MEI Thinking 1. This indicates a move away from solely relying on readily available web scrapes and suggests new approaches to data curation are emerging.
Diverging Philosophies on Synthetic Data
There's an ongoing debate about the role of synthetic data in training modern language models. The MEI Thinking 1 paper advocates for avoiding synthetic data and relying primarily on filtered web scripts to maintain a connection with human knowledge. However, this contrasts with other approaches, highlighting differing perspectives on how best to leverage diverse data sources to optimize model performance.
RL's Potential to Diminish Supervised Learning
The speaker contemplates whether reinforcement learning (RL) might eventually reduce the reliance on supervised learning in language models. They note that human language presents a complex distribution, making it challenging to learn solely through RL. However, they suggest that we may observe a decrease in the importance of supervised learning as RL becomes more prevalent, highlighting the value of base models possessing fundamental skills for RL.
Introducing Novel Data During Supervised Learning
Some labs are experimenting with incorporating novel data into supervised learning processes. This 'novel' data refers to information the model hasn’t encountered before, pushing it beyond familiar patterns. An example given is reasoning traces, which often differ significantly from typical human-generated text output and can help models develop more advanced capabilities.
Training for Test Time Compute
A strategy being explored involves training models to optimize performance during inference (test time). This is achieved by 'warming up' the model to specific compute schemes during supervised fine-tuning (SFT) or even pre-training. The goal is to improve efficiency and reduce latency when deploying these large language models in real-world applications.
Base Models Evolving Beyond General Knowledge
The speaker argues that the role of base models is shifting from encapsulating general human knowledge and world priors to emphasizing reasoning and agentic behavior. While this simplification overlooks the complexity of reasoning agents, it reflects how chatbots are increasingly utilized – as interactive tools requiring problem-solving capabilities. This shift underscores the importance of designing base models with these specific skills in mind.
Chapters
Claims & Fact Check
The idea of the base model is dead.
Models like OpenAI's 01 pioneered reasoning models.
RL is no longer a cherry on top, but can dramatically improve performance.
RL might eventually overtake supervised learning for language models.
Reasoning traces differ significantly from typical human-generated text output.
Base models are moving away from general knowledge towards reasoning and agentic behavior.
Was this digest good?
More from AI Engineer

Teaching AI to Find Real Vulnerabilities — David Brumley, Bugcrowd
Aug 1, 2026

Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software
Aug 1, 2026

What's Next After RLHF? — Diogo Almeida, TypeSafe AI
Jul 31, 2026

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI
Jul 31, 2026
Digest any single YouTube video — free.
3 free digests — no card, no sign-up wall.
Or just swap the domain of any YouTube link → instant digest