Product Introduction
- Definition: SelfJev is an open-source, self-hosted 4-billion-parameter decision model. It is a specialized AI model designed for structured decision-making, not text generation. It processes text and image inputs to output calibrated probabilities for typed questions.
- Core Value Proposition: SelfJev exists to provide developers and enterprises with a private, efficient, and cost-effective AI decision engine. It enables structured data extraction and classification from unstructured text and images while maintaining full data sovereignty and eliminating reliance on external API calls.
Main Features
- Shared-Prefix Inference Architecture: The model reads the input context (text or image) only once per request. This shared representation is then reused to answer multiple subsequent questions simultaneously, drastically reducing computational overhead and latency compared to sequential processing.
- Typed Question & Answer Framework: SelfJev supports four structured answer types: Yes/No (Noul), Pick One (Choice), Pick Many (MultiChoice), and Score (Scale). This enforces a deterministic output schema, returning machine-readable labels and calibrated probabilities instead of free-form generated text.
- Jev-Compatible API & Self-Hosting: The model provides an API specification compatible with TypeSafe's Jev, allowing seamless integration by changing environment variables. It is designed to be deployed on private infrastructure, from a local Mac with MPS to cloud GPUs like NVIDIA L40S or H100, giving users full control over data, cost, and performance.
- Vision-Language Capabilities: Built on Qwen2.5-4B-Instruct, SelfJev inherits its vision encoder. It can process images (via file path, bytes, or base64) as input
state, enabling visual question answering through the same structured decision API without generating descriptive text. - Configurable Fine-Tuning: Users can fine-tune the model on their proprietary datasets using supervised fine-tuning (SFT) or Low-Rank Adaptation (LoRA). This adapts the model to specific domains, vocabularies, and decision logic, improving accuracy for specialized use cases like internal policy checks or product categorization.
Problems Solved
- Pain Point: Eliminates the inefficiency and latency of sending multiple prompts to a large language model (LLM) for related questions about the same document, which incurs repeated context processing costs.
- Pain Point: Solves the problem of data privacy and vendor lock-in associated with using closed, hosted decision APIs, allowing sensitive data to remain on-premises or within a private cloud.
- Pain Point: Addresses the unreliability of parsing free-form text generated by LLMs for production systems by providing structured, typed outputs with probabilities that integrate directly into application logic.
- Target Audience: DevOps and MLOps engineers building internal AI agents; SaaS companies needing scalable, private content moderation or ticket routing; financial and legal firms processing sensitive documents for classification.
- Use Cases: Automated customer support ticket triage and routing; AI response safety and quality review (e.g., checking agent outputs for policy compliance); document classification and metadata extraction; visual content moderation and attribute tagging.
Unique Advantages
- Differentiation: Unlike general-purpose chat models (e.g., GPT-4, Claude), SelfJev is architecturally optimized for multi-question inference on a single context, offering significantly lower latency and cost per decision at scale. Compared to hosted decision APIs, it offers full data control and no per-call fees.
- Key Innovation: The "shared-prefix" tree-based inference engine is its core innovation. By caching the intermediate activations from the initial context processing, it bypasses redundant computation for subsequent questions, a technique that yields sub-200ms response times for multiple questions on a single GPU.
- Key Innovation: It focuses exclusively on producing calibrated probabilities for predefined options. This shifts the objective from language modeling to discriminative classification, which is more directly aligned with reliable decision-making and allows for targeted fine-tuning on decision accuracy.
Frequently Asked Questions (FAQ)
- What is the difference between SelfJev and Jev? SelfJev is the open-source, self-hostable version of the decision model technology that powers TypeSafe's hosted Jev API. They share the same API specification and core capabilities, but SelfJev gives you full control over deployment, data, and model fine-tuning.
- What hardware is required to run SelfJev? SelfJev-4B can run on a modern Mac with Apple Silicon (using MPS) or a cloud GPU with at least 24GB VRAM (e.g., NVIDIA A10G, L40S). For optimal performance with long contexts or high concurrency, a GPU like the L40S (48GB) or H100 (80GB) is recommended.
- Can SelfJev process images and text together? Yes, SelfJev's vision-language model can accept images as part of the input
state. You can provide a mix of text and images (as file paths, PIL images, or base64 strings) and ask structured questions about the visual content, such as identifying attributes or classifying scenes. - How accurate is SelfJev compared to larger models? In its published evaluations on a decision benchmark, SelfJev-4B achieved 95.7% accuracy matching expected answers, closely approaching the 97.2% accuracy of the larger, hosted Jev model. Accuracy for specific tasks can be further improved through fine-tuning on your data.
- Is fine-tuning SelfJev difficult? The project includes tools for supervised fine-tuning (SFT) and supports LoRA adapters. You can fine-tune it using your own dataset of
(context, question, correct_answer)examples via the provided CLI tools or the HTTP API endpoint (/v1/fine_tuning/jobs).