Pegasus 1.6 brings video understanding to physical AI, says TwelveLabs
Pegasus provides temporal context, spatial reasoning, and judgment about completing a task. Source: TwelveLabs
As the physical world rapidly digitizes, artificial intelligence teams are collecting large quantities of complex video, but turning raw footage into usable models remains a critical challenge, according to TwelveLabs Inc. The company today released its Pegasus 1.6 model with additional features for understanding and navigating complex real-world environments.
“Our mission has always been to help machines understand how the world works through video,” said Jae Lee, co-founder and CEO of TwelveLabs. “Physical AI is the next expression of that mission. Most of what people know about performing physical work, such as a change of grip or a recovery after something slips, has never been captured in a form that a machine can learn from.”
“With our latest model, we can now turn that footage into structured, reviewable knowledge, so that robotics and physics AI teams can train on real human experience instead of starting from scratch,” he added.
Founded in 2021, TwelveLabs said it has created a “comprehensive video intelligence platform” using its Marengo and Pegasus models, designed to see and understand videos like humans do in a fraction of the time. Developers can use this single system that gets smarter over time to access and act on all their video content, the Seoul-based company said.
Pegasus 1.6 focuses on self-centered videos
TwelveLabs said its latest version of Pegasus marks its expansion into physical AI, generating insights rich in real-world perspectives. It solves specific workflow challenges so that machines such as robots, drones and autonomous vehicles can sense, reason and act safely in the physical world like never before, the company said.
“We focus on the level of understanding of the video,” Jae Lee said The robot report. “You can’t just take millions of hours of video and feed it into a robotics model and expect useful training data to come out. You first have to understand what’s happening in the video, what action is taking place, when it happens, what the person is interacting with, how the behavior changes over time, and when other people or spectators appear in the field of view.”
“Pegasus 1.6 provides more precise understanding so that robotics and data teams can then turn it into the training data they need,” he added.
TwelveLabs noted that Pegasus 1.6 is its first AI model built to understand egocentric videos, which are shot from the point of view of the person doing the work. This could be someone cooking a meal, assembling parts on a factory line, or operating a robot remotely. The model does not require specific cameras, Lee said.
“We don’t require robotics teams to use a particular camera or proprietary hardware to capture that footage,” he said. “What matters is being able to capture the actions and interactions that happen from the operator’s point of view.”
Additionally, Pegasus 1.6 can work with existing video data and allows analysis of still images in addition to video.
“One of the reasons we focus on egocentric videos is that they are much easier to collect and scale than teleoperation data,” Lee acknowledged. “The goal is to take that footage and make it more useful for robotics teams by identifying actions that take place, breaking them down into precise time segments, and capturing finer details about the behavior.”
Head-mounted shooting can break models trained on broadcast, instructional, and cinematic videos. Source: TwelveLabs
TwelveLabs supports five workflows
Pegasus 1.6 currently supports five workflows based on its native video model. They include:
- Action segmentation and labeling: This capability accelerates model training with standardized datasets by automatically generating accurate, timestamped action labels for activities, steps, objects, and hand-object interactions from raw videos, mapped to the customer’s domain-specific taxonomy.
- Dense caption labeling: This enables natural language understanding for robots by producing rich, descriptive language for spatial relationships, scene context, and hand-object interactions to train advanced language-conditioned robot policies.
- Quality Score: Users can save time by filtering out low-quality videos by automatically evaluating and scoring video clips for clarity, framing and action stability before submitting footage to human reviewers.
- Research and care: Customers can discover critical edge cases by surfacing rare events, long-tail scenarios and duplicate clips across an entire video archive using simple natural language search queries.
- Consent and compliance reporting: Privacy and compliance can be maintained by detecting and reporting faces, viewers and sensitive data on screen or on paper before video footage enters downstream development pipelines.
Editor’s Note: Physical AI is one of the topics covered during the RoboBusiness 2026 session, which will be held on October 20 and 21 in Santa Clara, California. Register now to participate.
Register now and help us celebrate 20 years of RoboBusiness!
Customers will be able to take advantage of existing features
Pegasus 1.6 builds on the video understanding capabilities that TwelveLabs developed for companies that maintain massive video libraries.
The new version extends the features of Pegasus 1.5 that attracted numerous new customers. TwelveLabs cited Time-Based Metadata (TBM), which allows users to define a custom schema and automatically receive timestamped structured metadata from video content. This is especially useful for aiding contextual understanding, as people often narrate what they are doing in self-centered clips.
Pegasus 1.6 also improved entity recognition for more consistent tracking of hands, objects and tools in clips. According to TwelveLabs, the model also offers faster and cheaper processing for high-volume video workloads.
Lee said TwelveLabs’ system adds video understanding to other sensor modalities.
“We are focused on understanding the behavior we can observe in the video,” he said. “Pegasus 1.6 can identify the action taking place, segment it precisely in time, understand what the left and right hand are doing and how the surrounding environment is changing.”
“We can also provide a rough understanding of how the limbs move based on the video, but we are not trying to replace the tactile or actuator-level sensing of the robot,” he added. “Our role is to provide robotics teams with a deeper understanding of the behavior in the video that they can combine with their sensor data to build more precise trajectories and learning policies.”
TwelveLabs said Pegasus 1.6 expands its existing collaborations with robotics developers.
“We are working with robotics labs and data teams using video to broaden the data available for training robots,” Lee said. “Much of the work focuses on dexterity and manipulation, tasks such as packaging, assembly and cleaning, as well as more specialized industrial applications such as semiconductor quality control. The broader goal is to help these teams make much larger quantities of human behavioral videos useful for training, which is much more scalable than teleoperation data.”



Post Comment