Built on Cloudflare
Every piece of the platform maps to a Cloudflare primitive, with no external SaaS in the critical path. One Worker hosts the surfaces and the orchestration; a Container handles the work the Workers runtime cannot.
Why one platform#
Running compute, storage, the database, and the model gateway on a single platform is a deliberate trade. It costs some best-of-breed flexibility, and it buys three things that matter more for this workload: data stays close to the compute that processes it (no cross-cloud egress for large media), the attack surface stays small (one identity and network boundary instead of a patchwork), and every request sits behind a global network with always-on network-layer protection before it reaches our code. Durable execution and the object store are native here, so the orchestration layer is a primitive rather than a service to run.
Primitives and their roles#
WorkersThe REST API and the MCP server run on one Hono Worker. ASR, vision, reasoning, and Pass A also execute in the Worker via the Workflow.
WorkflowsEach extraction is a durable, retryable, checkpointed Workflow. A crash resumes from the last successful step rather than re-running the whole job.
ContainersHeavy media work (yt-dlp fetch, FFmpeg normalize and remux, scene detection, perceptual-hash dedup) runs in a stateless Container that reads and writes R2 by key.
D1Relational metadata: tenants, API keys, jobs, schemas, destinations, the graph mirror (nodes and edges), and the evidence chain.
R2All artifacts: source and normalized media, frames, transcript, vision descriptions, the canonical knowledge graph, and extraction results.
AI GatewayThe single choke point for model calls. Routing plus per-call logging of provider, model, latency, tokens, and cost.
Queues + Durable ObjectsThe planned scaling layer from ADR-010: per-tenant Durable Object admission feeding a Cloudflare Queue ahead of the Container pool. Designed, not yet shipped.
Why a separate Container for media#
The Workers runtime cannot run arbitrary binaries or sustained CPU-bound work, and the media stages are exactly that: shelling out to yt-dlp, running FFmpeg, and hashing frames. Those stages run in a Cloudflare Container instead. The Container is stateless: it reads inputs from R2 by key and writes outputs back by key, so it can be scaled, retried, or replaced without any coordination, while the durable job state stays in the Workflow, R2, and D1.
Scaling model#
The binding constraint under load is model throughput, not the Container pool: the provider rate limit on the reasoning and graph calls is reached well before compute runs out. The scaling design (ADR-010) is built around that fact rather than around raw concurrency.
It is three layers: per-tenant Durable Object admission decides whether a tenant may start another job, a Cloudflare Queue absorbs the burst, and the Container pool drains it. Per-plan concurrency caps keep one tenant from starving the rest. The admission and queue layers are designed and tracked; the per-job pipeline ships today.
Observability#
Every request gets a trace id at API ingress that propagates through the Workflow, the Container, and the result write. It is returned on every response as x-lirovo-request-id. Logs are structured (pino), and errors are captured in Sentry with the trace id attached, so a single id locates a request across the whole pipeline. Every model call is metered at the gateway with provider, model, latency, tokens, and cost.