> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sqwish.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Images

> Send photos, screenshots and scanned documents with a decision, for Fat and Chub.

Fat and Chub (`sqwish-decision-fat` and `sqwish-decision-chub`) read images as well as text. Put each picture in the request's `images` and mark where it belongs in the context with an anchor, `<image:1>`, `<image:2>` and so on. The decisions are then answered with the pictures in view, exactly as the models were trained to read them. Core and Dot read text only.

```bash theme={null}
IMAGE=$(base64 < front.jpg | tr -d '\n')
curl --fail-with-body https://console.sqwish.ai/v1/decide \
  -H "Authorization: Bearer $SQWISH_API_KEY" \
  -H "Content-Type: application/json" \
  -d @- <<EOF
{
  "model": "sqwish-decision-chub",
  "context": "Return request 4411. Photo of the item as received: <image:1>",
  "images": [{"id": 1, "media_type": "image/jpeg", "data": "$IMAGE"}],
  "decisions": [
    {"id": "damaged", "kind": "binary", "question": "Is the item visibly damaged?"},
    {"id": "condition", "kind": "single", "question": "What condition is the item in?",
     "outcomes": ["new", "used", "broken"]}
  ]
}
EOF
```

The answer has the same shape as any other: one distribution per decision, with `usage` saying how many images were read.

## Anchors and the context

* `context` must be a string. Each image's `id` is the number in its anchor, from 1 to 64, and appears once in `images`.
* Every image needs its anchor, and every anchor needs its image.
* An image is shown where its anchor first appears. A later mention of the same anchor is read as text, so you can refer back to a picture ("the damage in `<image:2>`") without sending it twice.
* Text around the anchors works as in any context: say what each picture is ("Photo of the item as received:", "Screenshot of the error:").

## Limits

| | |
| - | - |
| Models | `sqwish-decision-fat`, `sqwish-decision-chub` |
| Images per request | up to 8 |
| Formats | JPEG, PNG and WebP, still images only (no GIF or animation, SVG or PDF) |
| Size of one image | up to 10 MiB, 64 megapixels, and at most 200 times longer than it is wide |
| Size of the whole request | up to 48 MiB of JSON |
| Encoding | inline base64 (standard alphabet, padded, no line breaks); URLs are not fetched |

Each image is resized as the model's own image processor does, to between 65,536 and 2,097,152 pixels, and becomes between 64 and 2,048 image tokens (one per 32 × 32 pixels). A 640 × 480 photo is 300 tokens, a 1920 × 1080 screenshot 2,040, and a 4000 × 3000 phone photo 2,028: anything of about 2 megapixels or more comes to just under 2,048. The image tokens count toward the model's prompt limit, and nothing is ever cut to fit: a request over the limit is refused with `context_too_long`. Send fewer or smaller images if you see it.

A request with images can ask as many decisions as one without. The public playground takes no images.

## Fallback

With images, the fallback ladder keeps only the models that read images, so an answer always comes from a model that saw your pictures: Chub may fall back to Fat and Fat to Chub, never to a text-only size. A fine-tuned model reads text only.

## Billing

Images add tokens, never a different price. `usage.input_tokens` counts your text as usual, the anchors included, plus each image's tokens once, plus a fixed 300 tokens per image. A picture is billed once however many times the context mentions it. `usage.images` is the number of images and `usage.image_tokens` their tokens. For example, one 640 × 480 photo adds 300 + 300 = 600 tokens to the request, and one 4000 × 3000 phone photo 2,028 + 300 = 2,328. `POST /v1/decide/estimate` takes the same body and quotes it before you send it.

## Errors

| Code | Status | Meaning |
| - | - | - |
| `images_not_supported` | 422 | The model doesn't read images. The message names the models that do |
| `image_anchor_mismatch` | 422 | An image without its anchor, an anchor without its image, or vision tokens typed into the context |
| `image_context_must_be_text` | 422 | The context is a JSON object or array; send a string |
| `image_too_large` | 422 | An image is over 10 MiB |
| `image_pixels_exceeded` | 422 | An image is over 64 megapixels |
| `image_unreadable` | 422 | Not valid base64, not JPEG, PNG or WebP, animated, too long and thin, not the `media_type` it was sent as, or a file that can't be decoded |
| `context_too_long` | 422 | The text and image tokens together are over the model's prompt limit |
| `request_too_large` | 413 | The request is over 48 MiB |

None of these is retryable: change the request first.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.