Skip to main content
Fat and Chub (sqwish-decision-fat and sqwish-decision-chub) read images as well as text. Put each picture in the request’s images and mark where it belongs in the context with an anchor, <image:1>, <image:2> and so on. The decisions are then answered with the pictures in view, exactly as the models were trained to read them. Core and Dot read text only.
The answer has the same shape as any other: one distribution per decision, with usage saying how many images were read.

Anchors and the context

  • context must be a string. Each image’s id is the number in its anchor, from 1 to 64, and appears once in images.
  • Every image needs its anchor, and every anchor needs its image.
  • An image is shown where its anchor first appears. A later mention of the same anchor is read as text, so you can refer back to a picture (“the damage in <image:2>”) without sending it twice.
  • Text around the anchors works as in any context: say what each picture is (“Photo of the item as received:”, “Screenshot of the error:”).

Limits

Each image is resized as the model’s own image processor does, to between 65,536 and 2,097,152 pixels, and becomes between 64 and 2,048 image tokens (one per 32 × 32 pixels). A 640 × 480 photo is 300 tokens, a 1920 × 1080 screenshot 2,040, and a 4000 × 3000 phone photo 2,028: anything of about 2 megapixels or more comes to just under 2,048. The image tokens count toward the model’s prompt limit, and nothing is ever cut to fit: a request over the limit is refused with context_too_long. Send fewer or smaller images if you see it. A request with images can ask as many decisions as one without. The public playground takes no images.

Fallback

With images, the fallback ladder keeps only the models that read images, so an answer always comes from a model that saw your pictures: Chub may fall back to Fat and Fat to Chub, never to a text-only size. A fine-tuned model reads text only.

Billing

Images add tokens, never a different price. usage.input_tokens counts your text as usual, the anchors included, plus each image’s tokens once, plus a fixed 300 tokens per image. A picture is billed once however many times the context mentions it. usage.images is the number of images and usage.image_tokens their tokens. For example, one 640 × 480 photo adds 300 + 300 = 600 tokens to the request, and one 4000 × 3000 phone photo 2,028 + 300 = 2,328. POST /v1/decide/estimate takes the same body and quotes it before you send it.

Errors

None of these is retryable: change the request first.