id, status, limits and availability information. The La Décision catalogue includes Dot, Bit, Core, Chunk and Fat. A listed model may be unavailable on the current deployment. Early-access entries are discoverable without implying that their capabilities are enabled.
Base models, adapters and names
Base models have stable catalogue IDs such assqwish-d1-core. A fine-tuning job produces an adapter ID. A named model such as support-router points to a production version; support-router@3 selects an exact version. Named entries have kind: "named" and include production, candidate, previous and version summaries.
For reproducible comparisons, select an exact version and set fallback: "none". Inspect the response’s provenance as well: a reference is only one part of recording an evaluation.
Choose fallback deliberately
The default isfallback: "auto". The server builds a ladder from the requested model and eligible models. After the requested model, automatic fallback tries serving base sizes in catalogue size order, largest first. It does not automatically try earlier fine-tuned versions. Fallback handles selected availability and capacity failures; it does not turn invalid inputs or policy errors into valid requests.
/v1/decide with your key. If the requested model cannot answer, none returns an error rather than accepting another model’s result. A list of up to four model references supplies an explicit alternative ladder after the requested model.
When fallback occurs, the answer contains fallback.from, fallback.to and fallback.reason. Route the result according to your application’s requirements: you might accept it, send it for review, or mark it in an evaluation. Do not silently attribute it to the requested model.
Readiness and cold versions
GET /v1/models/{model_id}/ready performs a readiness probe. Read its ready flag; HTTP 200 alone does not mean the model is ready. Probing can load an adapter and perform real work.
POST /v1/models/{name}/wake starts loading a deployed sleeping version and returns 202 with its state. Poll the model for loaded, loading or asleep where residency is tracked. keep-warm requests that production stay loaded, subject to capacity. A warm version still has no fixed latency guarantee.
See versions and promotion for deployment and rollback, and accounts and usage for how the model that answered affects metering.