VORQ Docs
Guides

Serve Responses and Batch endpoints

Run long jobs through a background Responses endpoint or a Files + Batches endpoint, and give each SLA window its own backend.

Hour- and day-long SLA windows suit backends that accept work in the background and deliver it later, often at a lower price. Two presets cover the common OpenAI-compatible shapes. Both record the backend's id for the job, so a restart resumes the job instead of submitting it again.

Use a background Responses endpoint

    backend:
      preset: openai-responses
      base_url: https://inference.example.com/v1
      model: example-model
      api_key: env:INFERENCE_API_KEY
      service_tier: flex          # optional: flex or priority
      max_polls: 60               # status checks per SLA window

The daemon submits POST {base_url}/responses with background: true, then polls GET {base_url}/responses/{id} every sla / max_polls seconds (never under one second) until the status is completed. failed, cancelled and incomplete fail the job.

The request body is built like an openai-chat body and then respelled into Responses field names (messages becomes input, max_tokens becomes max_output_tokens, and so on). The finished object is converted back into a chat completion, so clients read the same shape from every text backend.

The default forwarded set is smaller than the chat one. Declare params_supported from the endpoint's own schema: some gateways drop an unknown Responses field silently, and the client then believes it set a control that never reached the model.

Use a Files + Batches endpoint

    backend:
      preset: openai-batch
      base_url: https://inference.example.com/v1
      model: example-model
      api_key: env:INFERENCE_API_KEY
      completion_window: 24h
      max_polls: 288              # a status check every five minutes over 24h

Each job becomes a batch of one line. The daemon uploads the chat-completions body as a one-line JSONL file, creates the batch, polls it until completed, and reads the job's line from the output file. The input file is deleted afterwards. Output and error files are left in place for you to inspect.

Keep retries: 0 (the default) on a model served through openai-batch. The batch id is only known once the create call answers, so a retry after a create that timed out can start a second batch for the same job, and both are billed. Limit keys are per model, so this applies to every window of that model.

A batch job holds one of the model's slots from upload to collection, which can be the whole window. Size capacity and concurrency for that.

Give one SLA window its own backend

sla_backends serves one window through a different backend. Each entry is a patch merged over backend: its keys win, null removes a key, and naming a preset drops an inherited raw mapping (and the reverse).

    slas:
      "1h":  { rate_in: "220000", rate_out: "750000" }
      "24h": { rate_in: "160000", rate_out: "550000" }
    backend:
      preset: openai-responses
      base_url: https://inference.example.com/v1
      model: example-model
      api_key: env:INFERENCE_API_KEY
    sla_backends:
      "24h":
        preset: openai-batch
        max_polls: 288

Here 1h jobs run in the background tier and 24h jobs through the batch tier. Limit keys (retries, concurrency, rate_limit and the rest) stay on backend and are refused in a patch. The model is offered only while every one of its backends passes its health check.

On this page