Set prices
Quote one ask per SLA window, and optionally accept cheaper bids while your backend is idle.
Quote one ask per SLA window
Each entry under slas publishes one ask: a window and its rates.
slas:
"24h": { rate_in: "160000", rate_out: "550000" }
"1h": { rate_in: "220000", rate_out: "750000" }- A window is a count and a unit:
s,m,hord. - A rate is a whole number, written as a string, of atomic USDC units per 106 units of
work.
"550000"on a text model is 0.55 USDC per million output tokens. A decimal is refused at startup. rate_outis required on every window.rate_inis omitted on a model that meters no input side, such as a text-to-image model.
The units depend on the modality: tokens for text and embedding, output pixels for
image, output pixel-seconds for video. See
Pricing and asks.
These rates are also your floor. The daemon claims a job only when the rates the client signed are at least the configured rates for that job's window. Jobs priced below are left for other providers.
Price changes take effect on restart. Stop the daemon with SIGTERM: a graceful shutdown
withdraws every ask it published, and the new process publishes the new book. After a hard
kill, a window you removed from the config stays on the book at its old price, because the new
process does not know it was ever published.
Accept cheaper bids while idle
Dynamic pricing is off by default. When configured, the daemon may claim bids priced under its published ask while the backend is quiet. The published ask does not change, and nothing about the discount is published or signed.
provider:
pricing:
max_discount_pct: 20 # the discount at a fully idle backend, shrinking to 0 at full load
bid_tolerance_pct: 10 # an extra discount while load is under low_load_pct
low_load_pct: 30
models:
- model: deepseek-ai/deepseek-v4-pro:fp8
# ...
load:
url: http://localhost:8000/metrics
metric: "vllm:kv_cache_usage_perc"
scale: 1.0At load 0.5 this accepts bids 10% under the ask; below 30% load it adds another 10%. The total
is capped at 90%. A model without a load block always claims at exactly its configured rates.
Choose a load source
- A Prometheus endpoint. Name the URL and the series.
scaleis the divisor that turns the series into 0..1:1.0for a ratio,100for a percentage such as a GPU exporter's utilization. When the series has several samples (one per GPU), the busiest one counts. - The entry's own occupancy. For a backend that exports nothing, such as a hosted API
behind a quota, use
load: { source: occupancy }. The load is the fuller of jobs held over what the entry may hold, and requests started over the budget of its longestrate_limitwindow. It needs aconcurrencyor arate_limiton the backend.
A probe that fails, returns no usable sample, or is older than 30 seconds counts as busy: the daemon then takes no discount. A failing probe never stops claiming or settling.
Check it
vorqd_backend_load{model} shows the last reading and vorqd_bid_floor_discount_pct{model}
the discount it currently earns. A discount of 0 while load looks low usually means the probe
is failing: check vorqd_load_probe_failures_total{model}.
The discount gauge reveals how far under your ask you will go. Keep the metrics port off the public internet.