Skip to main content
The account response shows identity, access state, exact wallet balances, referrals, onboarding and effective tier limits. The usage response gives totals, the preceding period, daily values, and breakdowns by served model and API key. Its dates are UTC.

Access and keys

An account can be waitlisted, active or blocked. A waitlisted account can see its place, browse models and request early access; most API work requires an active account. Early access to one catalogue entry is separate from general account admission. Create keys in the console or with POST /v1/account/keys. The full key is returned once. List keys to inspect their names, prefixes and last-used times; revoke a key with DELETE /v1/account/keys/{key_id}. The list never returns the secret again. Use different keys for separate applications so usage is easier to understand and access easier to revoke.

Metering

usage.input_tokens counts your supplied input using the shared billing tokenizer, identified in usage.tokenizer. It is distinct from the model’s rendered prompt tokens. The number of decisions is also reported. Native decisions return distributions, not generated output text. decision_price_nano on model catalogue entries is the list price in nano-USD per billed input token. A fine-tuned model uses its base size’s price. The model that actually answers, including a fallback, determines metered model usage. Consult the current catalogue and account rather than hardcoding a price from a guide. Wallet balances, estimates, charge receipts and usage costs are exact USD decimal strings, with nine places for accounting amounts. Preserve that precision instead of converting money to binary floating point. decision_price_nano is an exact integer string or null; zero means free serving, while null means unpriced. Before execution, the wallet reserves the highest possible price across the allowed fallback set. The response’s charge records the operation ID, reserved amount and actual answering model’s charge. Settlement and the replay response are committed atomically before success is returned. Reuse an Idempotency-Key to recover the original outcome after a lost response; see reliable retries.

Tier and limits

automatic_tier derives from verified purchases net of reversals. minimum_tier is the administrative floor; tier is their effective result. effective_limits reports rpm, tpm, concurrency, read_rpm and estimate_tokens_per_minute, shared across the account’s keys and browser sessions. The Account page explains these alongside the wallet.

Purchases and referrals

Manual and automatic top-ups require at least $5, in USD whole-cent amounts. Checkout credit is posted only after verified payment state. Read /v1/account/payments for purchase history and receipts; automatic top-ups require explicit consent, a saved card, a settings revision and a monthly cap. Disabling them stops new submissions while already submitted charges continue reconciliation. Verified signup issues 5startercreditonce.Aqualifyingreferredperson′spurchaseofatleast5 starter credit once. A qualifying referred person's purchase of at least 5 issues one $5 promotional reward, up to five people per referrer for life. A successful partial or full refund revokes unused reward credit permanently; spent reward credit is not recovered from the referrer’s paid balance. Revoked rewards still count toward the lifetime limit. /v1/account/referrals shows the reward history.

Diagnose a request

The required start and end dates are inclusive UTC dates. This example reads today. This cursor-paginated history contains request IDs, status codes, requested and served models, timing, token counts, fallback reasons, costs and key identity. It does not return the request context or decision answers. Fields can be null for failures that occurred before scoring. requests counts successful metered requests; total_requests includes recorded 4xx and 5xx failures. p50_ms and p95_ms summarize inference-stage time from histogram buckets. request_ms_sum / request_ms_count gives average full server-request time where measured; neither includes the client’s network overhead.

Training and tuning costs

Before submitting work, read fine_tune_price_cents and prompt_tuning.price_cents from /v1/decisionone. A job records the price it was created with. Fine-tuning failures and cancellations refund that job’s charge. Prompt tuning’s refund depends on the result and whether the same dataset has been tested before; see prompt tuning. Job refunds return to their original paid or promotional funding. Promotional expiry or revocation remains in force after a refund. Labeling is free.