> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sqwish.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Build a dataset

> Turn reviewed examples into immutable inputs for evaluation and training.

Create a tiny format example with two distinct contexts:

```bash theme={null}
curl --fail-with-body https://console.sqwish.ai/v1/datasets \
  -H "Authorization: Bearer $D1_API_KEY" -H 'Content-Type: application/json' \
  -H 'Idempotency-Key: refund-format-example-v1' \
  --data-binary @- <<'JSON'
{
  "name": "Refund format example",
  "examples": [
    {"context":"Please refund the duplicate charge.",
     "decisions":[{"id":"refund","kind":"binary","question":"Is a refund requested?"}],
     "targets":{"refund":{"yes":1,"no":0}}},
    {"context":"Where is my parcel?",
     "decisions":[{"id":"refund","kind":"binary","question":"Is a refund requested?"}],
     "targets":{"refund":{"yes":0,"no":1}}}
  ]
}
JSON
```

Save the returned dataset `id`. Inspect `rows`, `contexts`, `duplicates_removed`, `method` and `sha256`. Two rows demonstrate the format; they do not make a useful evaluation or fine-tune. Choose a new idempotency key for your next intended dataset.

## Labels and rewards

Each row has a context and decisions, plus one of:

* `targets`: a probability distribution over every outcome of each labelled decision. One-hot labels encode a single reviewed outcome; soft labels can encode disagreement.
* `rewards`: a finite reward for every outcome, for `method: "reward"` fine-tuning.

A dataset uses one supervision method throughout. Rows must not carry action policy fields such as weights, costs or abstention. Those policies belong to inference. Include at least two distinct contexts so a held-out split is possible. Identical examples can be deduplicated; conflicting labels for the same example are rejected rather than silently averaged.

## Upload JSONL and inspect the saved rows

For larger sets, save one example per line in `examples.jsonl`:

```bash theme={null}
curl --fail-with-body 'https://console.sqwish.ai/v1/datasets/upload?name=Refunds' \
  -H "Authorization: Bearer $D1_API_KEY" \
  -H 'Content-Type: application/x-ndjson' \
  -H 'Idempotency-Key: refunds-upload-v1' \
  --data-binary @examples.jsonl
```

JSON uploads are limited to 1 MiB; JSONL uploads to 16 MiB and 20,000 rows. Store the original upload and operation key for retries. `input` records the input format, byte count and hash when applicable. The dataset hash identifies its validated content.

Read `/v1/datasets/{dataset_id}/rows` to inspect a page. Download `/v1/datasets/{dataset_id}/download` for validated JSONL; this is not necessarily a byte-for-byte copy of the upload. Downloads are rate limited separately.

## Curate before training

Collect real cases with permission to use them, remove unnecessary identifying data, and agree on the labelling rule before assigning labels. Include difficult negatives, rare outcomes and unclear cases. Resolve disagreements rather than treating a teacher's answer as ground truth.

The split keeps identical contexts together, but that cannot detect all semantic duplicates. Remove near-duplicate leakage yourself. Keep a separate evaluation set representative of future traffic and compare against a simple baseline, such as the majority class.

Datasets are immutable. Upload a new dataset when labels or wording change. `DELETE` removes a dataset from new work and sets `deleted_at`; historical data and job references remain available. This is a soft deletion, not an erasure guarantee.

Next, [review teacher labels](/guides/labelling), [tune wording](/guides/prompt-tuning), or [train a version](/guides/fine-tuning).
