Video Has No Memory. Here's How We Built One. — James Le, TwelveLabs

AI Engineer20mJul 23, 2026
Watch Original (opens in new tab)
0:00 / 20:27
Chapters10

No clickbait detected — the title and thumbnail deliver what they promise.

AI Opinion

Le’s presentation convincingly argues that current AI approaches treat video too superficially, overlooking crucial temporal relationships—a point supported by the demonstrated limitations in existing systems. The claims regarding Pegasus's ability to consistently and reliably generate complex insights like nuanced safety event detection or precisely identifying key advertising moments rest on demonstrations that, while compelling, lack detailed explanation of error rates or edge cases. Viewers should investigate how TwelveLabs addresses potential biases within its training data and consider the computational resources required for "video cognition infrastructure" at scale before assuming broad applicability.

Avatars are AI rewrites of the same facts — style changes, not substance.

Summary

James Le of TwelveLabs presented their approach to building what they term "video memory," arguing that current video AI systems lack the ability to understand video as more than just a series of frames. Their system treats video as a “spatial temporal volume,” recognizing the importance of continuity and relationships between visual, audio, and textual elements across time. To address challenges like temporal dependency, multimodality, and ambiguity inherent in video data, TwelveLabs developed an architecture featuring "semantic chunks" and "Morango," which creates vector embeddings representing spatial-temporal relationships. A video context-aware language model called Pegasus then leverages these embeddings to enable reasoning and generate insights such as summaries or comparisons. Demonstrations included real-world security applications like object classification and safety event detection, as well as an advertising use case identifying key moments in sports footage. TwelveLabs positions its technology not as a finished application but as a foundational "video cognition infrastructure" intended to empower developers to build specialized solutions, with a private beta program targeted toward media professionals and content creators.

Avatars are AI rewrites of the same facts — style changes, not substance.

Key Points

01:00

Video as a Spatial Temporal Volume

The speaker emphasizes that video should not be treated as simply a collection of frames, but rather as a 'spatial temporal volume.' This concept highlights the importance of continuity and relationships between different elements within the video – visual information, speech, sound, motion, OCR data, scene transitions, and time. Treating video as just frames ignores its unique characteristic: the meaning derived from the sequence of events across space and time.

03:51

Five Key Challenges in Video Intelligence

The speaker outlines five key challenges when dealing with video data. These include temporal dependency (meaning relies on preceding and following frames), multimodality (video incorporates various signals like audio, visuals, and text), density (a short video can contain numerous shots and elements), ambiguity (people or brands may appear differently across different moments), and cost (tracing back to the source moment is often necessary in enterprise workflows). Addressing these challenges requires a specialized memory layer for video.

05:04

TwelveLabs' Architecture: Semantic Chunks & Morango

TwelveLabs has developed an architecture to address the complexities of video understanding. This includes 'semantic chunks' which capture meaningful temporal units within a video, and 'Morango,' a multimodal embedding encoder that transforms these chunks into spatial-temporal relationships represented as vector embeddings. These embeddings allow for efficient storage and retrieval of video content based on its semantic meaning.

05:34

Pegasus: Video Context Aware Language Model

To enable reasoning over video content, TwelveLabs has created 'Pegasus,' a video context-aware language model. This VLM serves as the reasoning layer, capable of generating summaries, metadata, comparisons, and other insights by leveraging the spatial temporal relationships captured in the underlying embeddings. It allows for more sophisticated understanding than traditional text models.

15:34

Real-World Security Applications with Object Classification

The presentation demonstrates the use of TwelveLabs' technology for security surveillance by analyzing publicly available camera footage. The system can ingest video, classify vehicles and pedestrians within an intersection, and detect safety events like near-collisions. This showcases a practical application beyond simple scene understanding, highlighting its potential in real-time monitoring and incident detection.

17:11

Advertising Use Case: Identifying Key Moments

The Adidas World Cup footage example illustrates how the technology can be used for advertising purposes. The system identifies key moments within a five-minute clip suitable for ad placement, such as impactful actions, player reveals, and logo appearances. This allows advertisers to strategically insert their content during high-engagement scenes, maximizing brand visibility.

19:04

Video Cognition Infrastructure: A Foundation for Innovation

TwelveLabs positions its product as a 'video cognition infrastructure' rather than a finished application. This means it provides the foundational layer – a knowledge store acting as video memory, configurable injection points, and a search API – that enables developers to build diverse solutions. This approach fosters innovation by empowering others to create specialized tools on top of their core technology.

19:52

Target Audience & Beta Program

The presentation concludes with a call to action, encouraging those working with video content – including media archivists, content creators on platforms like YouTube and TikTok, sports analysts, and professionals in media workflows – to register for the private beta program. This signals a focus on enabling broader adoption and gathering feedback from key user groups.

Chapters

10 chapters · 8 key moments
KEYkey momentPartially supportedUnverifiedNot checkable here

Claims & Fact Check

Most video AI systems do not have memory in a system sense.

±Partially supported

Existing stacks are not equipped to deal with video in the way humans do.

?Unverified

Video is not naturally a sequence of text tokens.

?Unverified

The system can classify vehicles and pedestrians within an intersection.

?Unverified

It can detect safety events, such as near-collisions.

?Unverified

The system identified key moments in the Adidas World Cup footage for ad placement.

Not checkable here

TwelveLabs positions its product as a 'video cognition infrastructure'.

Not checkable here

Was this digest good?

More from AI Engineer

Digest any single YouTube video — free.

3 free digests — no card, no sign-up wall.

Or just swap the domain of any YouTube link → instant digest