> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sqwish.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Train a model version

> Train on reviewed labels and inspect held-out evidence before serving the result.

After creating a representative dataset, set `D1_DATASET_ID` to its ID and check `training_available` in `/v1/decisionone`. Then create a named model's first version:

```bash theme={null}
curl --fail-with-body https://console.sqwish.ai/v1/fine-tuning/jobs \
  -H "Authorization: Bearer $D1_API_KEY" -H 'Content-Type: application/json' \
  -H 'Idempotency-Key: support-router-first-version' \
  -d "{\"dataset_id\":\"$D1_DATASET_ID\",\"model\":\"sqwish-d1-core\",\"method\":\"sft\",\"model_name\":\"support-router\"}"
```

The response is a job, not a ready production endpoint. Save its `id`, inspect its resolved hyperparameters and split, and poll `/v1/fine-tuning/jobs/{job_id}`. The example assumes target-labelled data; reward datasets require `method: "reward"`.

## Follow the job

| Status | Meaning |
| - | - |
| `queued` | Validated and waiting to start |
| `running` | Training or evaluating |
| `cancelling` | Cancellation requested; work is still stopping |
| `succeeded` | Output model and evaluation available |
| `failed` | Inspect the reported error and events |
| `cancelled` | Stopped without a successful output |

`output_model`, `metrics` and `progress` may be null before they exist. Read `/events` for lifecycle messages and `/records` for paginated held-out decisions. A client timeout or stopped poll does not cancel the job. Use the job's `/cancel` operation when you intend to stop it.

Leave hyperparameters out initially to use the base model's defaults. Accepted ranks and input limits depend on that base; being available for inference does not imply that a model is trainable. The job records the limits and split used for training.

## Evaluate before promoting

Inspect held-out probabilities beside targets or rewards. For supervised jobs, compare loss, accuracy and mistakes by task and class. For reward jobs, compare expected reward and the scaled-reward gate. A majority-class baseline or a simple rule can be more useful than a single headline accuracy number.

The gate compares a version to its base or parent and looks for regressions. It is a safety check on the held-out data, not a guarantee about future traffic. Check your own operational metrics, such as harmful actions or review workload, before promotion.

## Continue an existing model

For a new dataset, submit a job with `parent: "support-router"` to start from its production version. The model name follows the parent. The workflow mixes sampled prior training examples with new data and preserves held-out contexts. The `replay` setting controls the mix; use the default until your evaluations justify changing it.

A manually started job does not promote itself. Follow [versions and promotion](/guides/model-versions) to load, probe and move production deliberately. [Feedback and continual learning](/guides/feedback) can automate later rounds when explicitly enabled.
