Spend and models
What counts as spend, the caps that hold it, which model a run uses, and the model connections that pay for it.
Your keys, your bill
Runs use model connections you add: your own Claude or OpenRouter credentials. The model provider bills you directly. Zero Human does not resell model usage, and a cheap model is not a budget: caps are.
What counts as spend
Spend is the cost of model calls: every step of a run, and every chat reply. Tool calls are not spend. Each call is priced once, and the same figure is used by every cap, by the run's Cost, by today's total on Spend, and by an execution's total.
| Connection | What a call costs |
|---|---|
| OpenRouter API key | What OpenRouter reports it billed, plus the provider's own charge when you bring that provider's key to OpenRouter. |
| Claude Console API key | The model's list price for the tokens used, cache reads and writes included. |
| Claude Pro/Max seat | A share of the seat's price: the seat's weekly price, times the share of its weekly limit the calls used, spread across them by what they would have cost on an API key. A seat never costs more than an API key would. Usage past the seat's limit is priced at API rates. |
The seat's plan comes from Anthropic, or you set it on the connection (Seat plan and Monthly price (USD)).
A run's Cost is at the top right of its page. Hover it for the split between API and seat usage and the tokens in, out and cached.
Caps
Settings → Spend holds your caps. A cap sits on one layer: the enterprise, a team, a member, a role, or a task. Amounts are in US dollars.
| Limit | What it holds |
|---|---|
| Day $ | Money spent today (UTC). |
| Runs / day | Runs started today. |
| Run $ | What one run should cost. See Per-run budgets. |
| Max model tier | The most capable tier a run may use. |
A run is held by every cap that applies to it: the enterprise's, and those on its team, its member, its task and its member's roles. The tightest one wins.
Today's spend and today's runs are counted for the whole enterprise. A day cap on a team or a task is compared with everything the enterprise spent today, not only that team's or task's share; the Spend page says so on each row.
A new enterprise has no caps until you add one.
When a cap is reached
The caps are checked when a run is about to start, before its first model call. If today's spend (with room for the
next call) is at a day cap, or today's runs are past a Runs / day cap, the run does not start. It is parked as
paused_spend:
- It keeps its place: its task and its member stay busy, so later starts of that task wait or are skipped.
- It appears on Blockers under Paused spend, with Raise cap.
- A run already under way is not stopped by a day cap. It carries on to the end.
Paused runs start again when you raise a Day $ cap to a higher amount (every paused run in the enterprise is queued again and checks the caps afresh) or turn spend limits Off. Changing Runs / day or Run $ does not queue them again, and nothing restarts them at midnight: raise a day cap, turn limits off, or cancel them. Once spend reaches 80% of the tightest day cap, Spend and Enterprise warn you and Spend in the sidebar shows a badge.
Chat replies count toward today's spend too. Before each reply, the day caps on the enterprise and on that member are checked; at a cap the member says so and does not reply. See Chat and plans.
Per-run budgets
A task's Spend USD / run, and Run $ on a cap, set what one run should cost at most; the tighter of the two
applies. It is only partly enforced today. On an OpenRouter connection, each model reply is kept short enough for the
run to stay within it, and a reply that the limit cuts off fails the run (llm_max_tokens). On a Claude connection
it is not enforced during a run.
Turning spend limits off
The switch at the top of Spend turns spend limits On or Off for your enterprise. Off ignores every money cap, every run-count cap, the caps' model tiers and the tasks' per-run budgets, for runs and chat alike, and queues any paused runs again. Spend is still recorded, and a task's own highest tier still applies. With limits off, the rest of the Spend page and the Paused spend section on Blockers are hidden.
Planned
Caps still to come
- A month cap. Month $ can be saved on a cap today, but nothing holds a run to it yet.
- Stopping a run mid-way. A run that reaches a cap will stop before its next paid model call, instead of only new runs being held.
- A per-run cap on every connection. A run that reaches its per-run budget will stop, on a Claude connection as on OpenRouter.
- An approval threshold. A run that would spend more than an amount you set will wait at a gate for you before its next paid call.
Model tiers
A tier names a class of model, so a task can ask for "a cheap one" without naming it:
| Tier | For |
|---|---|
economical |
The cheapest model; fine for API and data work. |
mid |
A good mid-priced model. |
frontier |
The most capable, and the most expensive. |
Each tier stands for one model on each provider. The cap form shows which model a tier means on your connection.
Which model a run uses
A task's Model field decides:
- An exact model, picked from your providers' directories, is the model the run uses. A tier cap does not block it; money caps still apply.
- Default uses a tier: the tightest of the tiers on the caps that apply and the task's own highest tier, or
economicalwhen neither sets one.
Every cap carries a Max model tier, mid unless you choose otherwise. So adding a cap also sets the tier your
Default runs use.
Chat has its own setting, Chat model on Enterprise, economical unless you raise it, and tightened by caps
like a run.
Model connections
Settings → Agents lists every model connection in the enterprise, whose it is, and whether it can be read. Keys are never displayed. Add Agent adds one:
| Provider | How you connect it |
|---|---|
| Claude | A Console API key, a setup token, or Connect Claude Pro/Max to sign in with your seat. |
| OpenRouter | An API key. |
A connection belongs to the enterprise, shared by everyone, or to one member, used for that member's runs and chats. A run tries them in this order:
- The member's own Claude connection.
- The enterprise's Claude connection.
- The member's own OpenRouter connection.
- The enterprise's OpenRouter connection.
When a task names an exact model and both providers are connected, the provider that offers that model is tried first. The OS's own tasks have no member, so they use the enterprise's connections.
When a connection cannot be used
A run never fails for want of a model. It is blocked, says why, and resumes when the cause clears:
| What happened | Reason | Resumes |
|---|---|---|
| No connection at all | llm_not_bound |
As soon as you add or change a credential, and it tries again hourly |
| The provider refused the credential, or it is out of credit | llm_http_401, …402, …403 |
As soon as you change a credential, and it tries again hourly |
| A Claude seat's usage limit is used up | llm_quota_exhausted |
When the limit resets. Blockers lists it under No LLM capacity. |
| The model does not exist | llm_http_404 |
When the model answers again; after an hour it waits for you |
The OS also checks every ten minutes that each tier's model still exists on your connection. A model the provider says is gone is listed on Blockers under Models unavailable, before a run meets it. See Runs.
Plans
Your plan sets how long run history is kept (see Isolation and retention), and what your enterprise includes: members, runs, runner minutes, concurrent runs, schedules and storage. See pricing.
Planned
Plan limits
The amounts your plan includes will be enforced. Where a plan says stop, work past it will not start until you move to a plan that includes more, rather than running on and being billed. Today the plan's included amounts are not enforced; your own caps and your enterprise's Concurrent runs setting are what limit work.
Over the API
| Call | Scope | What it does |
|---|---|---|
GET /v1/spend |
spend:read |
Today's spend and runs, the tightest caps, and every cap. |
PATCH /v1/spend/caps |
spend:write |
Add or replace the cap on one layer. Raising a day cap queues paused runs again. |
PATCH /v1/spend/enabled |
spend:write |
Turn spend limits on or off: { "enabled": false }. |
GET /v1/llm/models |
llm:read |
The models your connections offer, for a task's Model. |
GET /v1/llm/connections |
llm:read |
Your model connections, without their keys. |