LLM Vision Token Calculator
Drop an image or enter its size to see how many input tokens it costs on each vision model, how OpenAI slices it into 512-pixel tiles, and what a batch of them costs.
Everything runs locally: your image never leaves your browser.
Calculate vision tokens for an image
| Model | Tokens | Accuracy | $ / 1M input | Cost |
|---|---|---|---|---|
| GPT-4o / GPT-4.1 OpenAI | 1,105 | Exact | $0.00276 | |
| GPT-4o mini OpenAI | 36,835 | Exact | $0.00553 | |
| o1 / o3 OpenAI | 975 | Exact | — | |
| GPT-4.1 mini OpenAI | 2,443 | Exact | $0.00098 | |
| GPT-4.1 nano OpenAI | 3,710 | Exact | $0.00037 | |
| o4-mini OpenAI | 2,594 | Exact | $0.00285 | |
| Claude (Sonnet, Opus, Haiku) Anthropic | 1,609 | ≈ Approx. | $0.00483 | |
| Gemini 2.x Google | 1,548 | ≈ Approx. | $0.00015 |
Prices are sample list prices and change often — edit them to match your plan. Token counts do not depend on price.
How each provider counts image tokens
OpenAI documents its method exactly. A 1920 × 1080 screenshot in high detail is scaled to 1365 × 768, which spans 3 × 2 tiles of 512 pixels, so it costs 6 × 170 + 85 = 1,105 tokens on GPT-4o. Newer mini and nano models use 32-pixel patches, capped at 1,536, times a per-model multiplier. Both methods here reproduce OpenAI’s own worked examples exactly.
Claude costs about width × height ÷ 750 tokens, after scaling down anything over 1568 pixels on the long edge or about 1,600 tokens. Gemini charges 258 tokens for a small image and 258 per 768-pixel tile above 384 pixels. Neither publishes its exact resizing rule, so those rows are marked as approximations.
A GPT-4o image token cost calculator that shows the tiles
The number people usually need from a GPT-4o image token cost calculator is not just the total but why it is that total — which is why the grid is drawn. Because the cost moves in whole tiles, an image a few pixels past a 512-pixel boundary after resizing pays for a whole extra row or column. As an OpenAI vision image tile calculator, it shows exactly where those boundaries fall.
Where these numbers come from
OpenAI’s tile and patch algorithms and their worked examples are from OpenAI’s vision documentation. Claude’s formula, resize limits and maximum-size table are from Anthropic’s vision documentation. Gemini’s 258-token and 768-pixel tile figures are from Google’s Gemini API documentation.