> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sqwish.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Models and fallback

> Discover serving models, pin versions and decide how to handle unavailability.

```bash theme={null}
curl --fail-with-body https://console.sqwish.ai/v1/models \
  -H "Authorization: Bearer $D1_API_KEY"
```

Inspect each entry's `id`, `status`, limits and availability information. The La Décision catalogue includes Dot, Bit, Core, Chunk and Fat. A listed model may be unavailable on the current deployment. Early-access entries are discoverable without implying that their capabilities are enabled.

## Base models, adapters and names

Base models have stable catalogue IDs such as `sqwish-d1-core`. A fine-tuning job produces an adapter ID. A named model such as `support-router` points to a production version; `support-router@3` selects an exact version. Named entries have `kind: "named"` and include production, candidate, previous and version summaries.

For reproducible comparisons, select an exact version and set `fallback: "none"`. Inspect the response's provenance as well: a reference is only one part of recording an evaluation.

## Choose fallback deliberately

The default is `fallback: "auto"`. The server builds a ladder from the requested model and eligible models. After the requested model, automatic fallback tries serving base sizes in catalogue size order, largest first. It does not automatically try earlier fine-tuned versions. Fallback handles selected availability and capacity failures; it does not turn invalid inputs or policy errors into valid requests.

```json theme={null}
{
  "model": "sqwish-d1-core",
  "fallback": "none",
  "context": "Please refund the duplicate charge.",
  "decisions": [{"id": "refund", "kind": "binary", "question": "Is a refund requested?"}]
}
```

Send that body to `/v1/decide` with your key. If the requested model cannot answer, `none` returns an error rather than accepting another model's result. A list of up to four model references supplies an explicit alternative ladder after the requested model.

When fallback occurs, the answer contains `fallback.from`, `fallback.to` and `fallback.reason`. Route the result according to your application's requirements: you might accept it, send it for review, or mark it in an evaluation. Do not silently attribute it to the requested model.

## Readiness and cold versions

`GET /v1/models/{model_id}/ready` performs a readiness probe. Read its `ready` flag; HTTP 200 alone does not mean the model is ready. Probing can load an adapter and perform real work.

`POST /v1/models/{name}/wake` starts loading a deployed sleeping version and returns 202 with its state. Poll the model for `loaded`, `loading` or `asleep` where residency is tracked. `keep-warm` requests that production stay loaded, subject to capacity. A warm version still has no fixed latency guarantee.

See [versions and promotion](/guides/model-versions) for deployment and rollback, and [accounts and usage](/guides/accounts-and-usage) for how the model that answered affects metering.
