Why Large? Tiny LMs & Agents on Edge/Robotics — Cormac Brick, Google

AI Engineer21mJul 25, 2026
Watch Original (opens in new tab)
0:00 / 21:45
Chapters13

No clickbait detected — the title and thumbnail deliver what they promise.

AI Opinion

The episode convincingly demonstrates the potential of small language models for enabling AI functionality on resource-constrained devices, particularly through strategic synthetic data generation techniques—a point supported by their "Mobile Actions" dataset and offline dictation demo. However, the claim that 10,000 to 10 million samples consistently achieve “really really high” reliability in fine-tuning warrants further scrutiny; results likely vary significantly depending on task complexity and data quality. Listeners should also consider that while DRAM costs are a confirmed obstacle, the episode doesn't fully explore alternative hardware optimization strategies beyond simply shrinking model size.

Avatars are AI rewrites of the same facts — style changes, not substance.

Summary

Google's AI Edge team is focused on enabling broader accessibility of artificial intelligence through the development and optimization of small language models (SLMs), typically ranging from one to four billion parameters. A key strategy involves synthetically generating data—often between 10,000 and 10 million samples—to fine-tune these smaller models for specific tasks like voice-activated function calling on mobile devices, demonstrated through Google’s open-sourced "mobile actions" dataset. The team showcased an offline voice dictation app powered by tiny Gemma models capable of cleaning filler words and personalizing output, highlighting the potential for resource-efficient AI deployment even on devices with limited DRAM (4-8 GB). However, rising DRAM costs pose a significant challenge to wider adoption, particularly in lower-tier web browsers and consumer robotics. Future efforts are aimed at automating synthetic data generation using agents and exploring faster visual models, while addressing hardware limitations to expand the reach of these smaller, more accessible AI solutions.

Avatars are AI rewrites of the same facts — style changes, not substance.

Key Points

15:08

Mobile Actions Model for Function Calling

Google demonstrated a 'mobile actions model' capable of text input and function calling with 10 different output functions. This model achieves an 86% reliability rate when converting arbitrary text into appropriate function calls, enabling common tasks on mobile devices like scheduling calendar events or controlling Wi-Fi. The demo integrated an ASR model to create a voice-to-function calling feature.

16:30

Synthetic Data is Key for Fine-tuning

The most effective approach, according to Google's playbook, involves synthetically generating data to fine-tune smaller models. They have open-sourced a dataset called 'mobile actions' on Hugging Face that can be used to recreate the demo and fine-tune Gemma from scratch. The team found that 10,000 to 10 million samples of synthetic data are typically sufficient for achieving high reliability.

17:46

Offline Voice Dictation with Tiny Models

Google showcased an app offering offline voice dictation without a subscription, leveraging two fine-tuned tiny Gemma models. This application not only performs dictation but also cleans up filler words ('ums' and 'a's) and personalizes the output by biasing towards relevant names and terms. The backbone of this app consists of low single-digit millions of parameters.

20:35

Future Ambitions: Agent-Generated Synthetic Data

Looking ahead, Google aims to simplify voice-to-function calling by exploring agent-generated synthetic data. This would automate the process of creating training datasets and make it accessible to a wider audience. They also see potential for faster visual models capable of segmentation and other tasks, further expanding the capabilities of tiny models.

NaN:NaN

Small Models & Hardware Requirements

The definition of a 'small' model is typically between one to four billion parameters. Deploying these models often requires 4 to 8 gigabytes of DRAM, limiting their use to devices like laptops and mobile phones. This restricts broader adoption in lower-tier web browsers or consumer robotics due to cost considerations.

NaN:NaN

Challenges of Edge Deployment: DRAM Cost

A significant challenge in deploying AI on the edge is Dynamic Random Access Memory (DRAM) cost, which has been steadily increasing. The speaker notes that some manufacturers are even reducing DRAM allocation in new devices and cites a 2.5x increase in Raspberry Pi 36 gigabytes costs, highlighting its substantial impact on design choices and model size limitations.

NaN:NaN

Google's AI Edge Team Focus

The speaker explains that Google’s AI edge team develops open-source projects like 'Lighter TLM,' 'Lighter T Mediapipe,' and collaborates with the Gemma team. These initiatives aim to simplify AI deployment on edge devices, deliver core technology to Google's products, and ensure compatibility of models across various hardware platforms.

NaN:NaN

The Need for Tiny Models

Cormac emphasizes the necessity of tiny models to enable AI deployment across a wide range of devices, not just expensive robots. This is crucial for accessibility and affordability, as larger models are impractical for many applications due to resource constraints. The talk will focus on current state-of-the-art techniques and practical implementation strategies for these smaller models.

Chapters

13 chapters · 8 key moments
KEYkey momentNot checkable hereWell-supportedUnverified

Claims & Fact Check

DAM costs are a significant constraint on edge AI deployment.

Not checkable here

Raspberry Pi 36 gigabytes have increased in cost by a factor of 2.5x.

Not checkable here

Techniques like zero-shot prompting and LoRA adapters can improve performance with smaller models.

Well-supported

This model knows about 10 different output functions and can call them at over 86% uh reliability from a given arbitrary text input.

Not checkable here

We’ve generally found that in the range of 10,000 to 10 million samples of synthetically generated um data will be sufficient to fine-tune a smaller model to a really really high degree of reliability.

?Unverified

tiny models will enable reach to a much much larger pool of devices and voiceto function calling um can now be built to to be robust using tiny models.

Not checkable here

Was this digest good?

More from AI Engineer

Digest any single YouTube video — free.

3 free digests — no card, no sign-up wall.

Or just swap the domain of any YouTube link → instant digest