Get started

Structured extraction from video

Lirovo is the structured-extraction layer for video. Hand it a video plus a JSON Schema and get back typed JSON, where every value is linked to the source moment and modality that produced it.

What Lirovo is#

You define the shape you want with a JSON Schema and point a job at a video. Lirovo runs the pipeline (transcription, frame sampling, vision, reasoning) and returns JSON that conforms to your schema. Every extracted value carries evidence: the exact timestamp and modality (audio or visual) it came from.

The same surface is callable three ways: as a REST API, through the TypeScript SDK, or over MCP so an AI agent can call it directly. Start with the Quickstart, then dig into the concepts and the reference.