Architecture

Models

The model layer is provider-agnostic and routed through Cloudflare AI Gateway. Lirovo is not tied to a single provider, the default stack can be overridden per request, and a degraded provider falls back rather than failing the job.

Lirovo · architecture / model routinglirovo.aiWorkerpipeline stepsAI Gateway · BYOKcircuit breakerclosed · open · half-opengemini-3.1-flash-litereasoning · Pass A · visionglm-4.7-flashWorkers AI fallbackGoogle GeminiproviderVoxtral · ASRdirect to Mistralreasoning · visionon openroutedaudio · directdata flowindirect
Reasoning, Pass A, and vision route through AI Gateway behind a per-model circuit breaker, with a Workers AI fallback. Voxtral ASR is the one call that goes direct to Mistral, outside the gateway.

Why route through AI Gateway#

Putting one gateway in front of every model call is what makes the model layer swappable. It gives a single place to log provider, model, latency, tokens, and cost; a single place to attach a per-model circuit breaker and a fallback; and a single key story (bring-your-own-key) instead of credentials scattered through the code. Swapping a model or adding a provider is a configuration change behind a stable interface, not a rewrite.

The one exception is Voxtral ASR, which goes direct to the Mistral API because AI Gateway does not cover Mistral audio. That call is metered with app-level logs so it is still observable.

The default stack#

Reasoning
google/gemini-3.1-flash-lite

Default reasoning model, via AI Gateway with bring-your-own-key (BYOK).

Pass A (graph)
google/gemini-3.1-flash-lite

Default knowledge-graph model, via AI Gateway BYOK.

Vision
google/gemini-3.1-flash-lite

Default vision model, via AI Gateway BYOK.

ASR
voxtral-mini-2602

Transcription, called direct to the Mistral API. AI Gateway does not cover Mistral audio, so this is the one path outside the gateway.

Reasoning / Pass A fallback
@cf/zai-org/glm-4.7-flash

Fallback reasoning and Pass A model on Cloudflare Workers AI.

Per-request overrides#

Send the X-Lirovo-Model-Stack header to override the default models for a single request. This is how you A/B a candidate model or pin a job to the fallback.

bash
curl https://api.lirovo.ai/v1/jobs \
  -H "Authorization: Bearer $LIROVO_API_KEY" \
  -H "X-Lirovo-Model-Stack: reasoning=@cf/zai-org/glm-4.7-flash" \
  -H "Content-Type: application/json" \
  -d '{ "source_uri": "https://www.youtube.com/watch?v=...", "schema_id": "sch_a1b2c3" }'

Changing a default safely#

Default model ids are constants, not scattered literals. A fraction of traffic can be routed through a candidate stack and logged side by side for comparison, and a default is only promoted after a run of green nightly evals. The change to flip a default is then a one-line constant change, kept separate from any behavior change.

Fallback and breaker
When the primary model errors or the circuit breaker is open, reasoning and Pass A fall back to the Workers AI model automatically. See Reliability and safety for the breaker states.