Structured extraction from video
Lirovo is the structured-extraction layer for video. Hand it a video plus a JSON Schema and get back typed JSON, where every value is linked to the source moment and modality that produced it.
What Lirovo is#
You define the shape you want with a JSON Schema and point a job at a video. Lirovo runs the pipeline (transcription, frame sampling, vision, reasoning) and returns JSON that conforms to your schema. Every extracted value carries evidence: the exact timestamp and modality (audio or visual) it came from.
The same surface is callable three ways: as a REST API, through the TypeScript SDK, or over MCP so an AI agent can call it directly. Start with the Quickstart, then dig into the concepts and the reference.
Three REST calls: register a schema, submit a job, fetch the typed result.
Schemas, jobs, results, evidence, the knowledge graph, artifacts, destinations, and tenancy.
Every REST endpoint, the auth model, the error envelope, and request and response shapes.
The typed client for the Developer API: schemas, destinations, and jobs.
Connect any MCP client and video extraction becomes a set of tools your agent can call.
How the pipeline runs on Cloudflare: Worker, Workflow, Container, D1, R2, and AI Gateway.