Product Introduction
- Definition: Jockey by TwelveLabs is a multimodal video intelligence agent and API platform. Technically, it is an AI-powered agent layer built on top of TwelveLabs' proprietary foundation models, Pegasus and Marengo, designed to understand, search, and structure content within video and photo libraries.
- Core Value Proposition: It exists to transform unstructured video and image libraries into instantly queryable knowledge bases. Its primary value is enabling users to search their media by semantic meaning—such as specific moments, actions, objects, people, or spoken words—using natural language, bypassing the need for manual tagging or folder-based organization.
Main Features
- Multimodal Semantic Search: This feature allows users to search their entire media library using natural language queries, images, or concepts. How it works: The underlying Marengo model generates a unified embedding vector that captures meaning across video frames, audio waveforms, and on-screen text. This enables retrieval based on semantic similarity rather than filename or manual metadata. For example, searching "find all the hikes with mountain views" returns relevant clips regardless of their file names.
- Structured Data & Metadata Extraction: Jockey can analyze video content and output timestamped, structured data according to a user-defined JSON schema. How it works: The Pegasus model describes visual and auditory scenes, which Jockey's agent layer then parses to populate predefined fields like "sentiment," "format," "identified objects," or "speaker quotes." This automates the generation of machine-readable metadata for every video file.
- AI Agent Integration via MCP (Model Context Protocol): This feature integrates Jockey directly into existing AI assistant workflows, starting with Claude. How it works: Users connect their media library via the TwelveLabs MCP server. Once connected, they can ask the AI assistant (e.g., "Find the scene where I explain the product hook") and Jockey will retrieve the precise moments from the connected library, enabling conversational interaction with video archives without a separate interface.
- Automatic Content Analysis & Tagging: Jockey automatically watches and labels incoming video content with descriptive tags. How it works: Using its combined model stack, it identifies elements like "hook" (the opening 3 seconds), on-screen talent, logos, visual tone (e.g., "cinematic"), format (e.g., "single continuous take"), and emotional sentiment across the timeline, creating a searchable index as soon as media is uploaded.
Problems Solved
- Pain Point: The "needle in a haystack" problem in large video libraries. Professionals waste hours manually scrubbing through footage to find specific moments, quotes, or scenes, leading to inefficient creative and marketing workflows.
- Target Audience: Marketing Teams managing large ad libraries; AI-Native Content Creators & Videographers with extensive raw footage archives; Developers building applications that require video understanding (e.g., sports analytics, media platforms, training systems).
- Use Cases: Ad Performance Analysis: Automatically tagging thousands of past ads by hook, talent, and format to identify winning patterns. Content Creation: A creator searching their entire archive for "every time I said 'machine learning' on camera" to quickly build a compilation. Application Development: A sports tech startup using the API to build a tool that finds all "corner kick" moments and analyzes player formation from broadcast footage.
Unique Advantages
- Differentiation: Unlike traditional video management tools that rely on manual tagging or basic keyword search, Jockey understands the context and content within the video. Unlike some AI tools that only generate a text description, Jockey provides frame-accurate timestamps and structured, queryable data outputs.
- Key Innovation: The integration of two specialized, proprietary models—Marengo for multimodal indexing/retrieval and Pegasus for dense captioning/description—into a single agent layer ("Jockey"). This stack allows it to both find relevant moments based on deep semantic understanding and describe them in structured detail, enabling complex queries and data extraction that single-model systems cannot perform.
Frequently Asked Questions (FAQ)
- How does Jockey by TwelveLabs handle privacy and facial recognition? Jockey uses multimodal AI analysis (visual context, motion, audio) to group and identify people within your private library for search purposes. It does not use traditional facial recognition databases, compare your data against external records, or share identity information between user accounts, prioritizing on-device/private-cloud analysis principles.
- What video formats and library sizes does Jockey support? Jockey supports common video formats like MP4, MOV, AVI, MKV, and WebM, and image formats like JPEG, PNG, and HEIC. It offers tiered plans based on knowledge store size, from a 5GB Free tier up to a 500GB Pro tier, designed to scale from individual creators to enterprise media libraries.
- Can I use Jockey's API to build my own video search application? Yes. Developers can use the TwelveLabs SDK and API to integrate Jockey's video understanding capabilities—such as natural-language search, multimodal embeddings, and structured data extraction—directly into their own applications, enabling custom video intelligence features.
- What is the difference between Jockey and TwelveLabs' other products like Embed or Search? Jockey is the agent layer that combines the capabilities of the underlying models (Marengo for search, Pegasus for analysis) into a conversational, task-oriented interface. It is built on top of and orchestrates the core APIs (Embed for vectorization, Search for retrieval) to provide a higher-level, user-friendly product for end-users and a powerful agentic API for developers.
