Product Introduction
- Definition: Chert is a software-as-a-service (SaaS) platform for building and deploying interactive AI video agents. Technically, it is a specialized API and control plane that enables AI-powered video calling via Apple's FaceTime protocol.
- Core Value Proposition: Chert exists to solve visual communication problems where voice-only AI agents fail. Its primary value is enabling businesses to deploy scalable, AI-driven video support for scenarios that require visual context, effectively acting as "Vapi for FaceTime" to automate complex, visual customer and employee interactions.
Main Features
- AI-Powered Video Context: Unlike voice-only bots, Chert agents can process real-time video feed from a caller's camera. This allows the underlying AI model to analyze visual information—such as identifying objects, reading labels, or assessing a physical environment—and provide context-aware guidance. It works by streaming video frames to a configured vision-capable language model (like GPT-4V) for analysis.
- Programmable Agent Configuration: Users can define an agent's behavior through a detailed prompt, specifying its role, instructions, tone, and decision-making logic. The platform allows configuration of the AI model, voice synthesis, interruption handling, and publishing state, enabling precise control over the agent's personality and response patterns before deployment.
- Managed FaceTime Line Infrastructure: Chert abstracts the complexity of Apple's FaceTime infrastructure. It provides provisioned and managed FaceTime phone numbers (e.g., +1 628-228-1717) to which AI agents can be assigned. This handles the backend signaling, media negotiation, and compliance, allowing developers to deploy a live agent with just a few lines of code focused on business logic.
Problems Solved
- Pain Point: The inefficiency and ambiguity of troubleshooting visual or physical problems over voice-only support channels. Describing complex hardware, cable setups, or environmental issues is error-prone and time-consuming.
- Target Audience: Product and engineering leaders at SaaS and hardware companies, operations managers in field service and telecommunications, telehealth platform developers, and customer success teams needing scalable, visual first-line support.
- Use Cases: The product is essential for remote IT support (guiding a user to plug in a specific cable), field service dispatch (verifying a technician's work visually), telehealth patient intake (documenting visual symptoms), guided product onboarding (showing a user where a button is), and visual quality inspections in manufacturing or logistics.
Unique Advantages
- Differentiation: Chert directly competes with voice AI platforms like Vapi by adding the critical dimension of live video. Compared to traditional methods like human video calls or static tutorial videos, it offers 24/7 scalability, consistency, and interactive, context-sensitive guidance.
- Key Innovation: Its core innovation is the operationalization of the FaceTime protocol for B2B AI agent use cases. By providing a managed service that handles FaceTime line provisioning, media routing, and integration with vision-language models, Chert significantly lowers the technical barrier to creating production-ready, interactive video AI experiences.
Frequently Asked Questions (FAQ)
- Is there a FaceTime API for developers? Apple does not offer a public, generally available FaceTime API. Chert operates through a controlled private preview, providing managed access to FaceTime infrastructure for businesses building AI video agents, handling the underlying complexity and compliance.
- How does Chert's AI understand what it sees on video calls? Chert integrates with advanced multimodal AI models (like OpenAI's GPT-4 with vision capabilities). The agent streams video frames to this model, which analyzes the visual context in real-time alongside the conversation transcript to generate informed, situation-aware responses.
- Can I test a Chert AI agent without making a FaceTime call? Yes, Chert includes a browser-based preview mode. This allows developers to test the agent's prompt logic, voice, interruption behavior, and avatar performance using their computer's microphone and camera before publishing to a live FaceTime line.
- What kind of AI avatars or personas does Chert use? The platform offers selectable digital avatars with different framings (e.g., mid-shot). These personas are designed to fit various professional contexts, such as a calm remote support agent, providing a consistent and appropriate visual representation for the AI during calls.
- Does Chert support outbound FaceTime calls from an AI agent? Yes, Chert's control plane supports bounded workflows for both inbound and outbound test calls. For live production use, outbound calling capabilities are provisioned and explicitly authorized on a per-use-case basis, as accepting inbound calls is the default mode.
