Skip to content
Owais Khan Software Reviews

LLM Vision Token Calculator

Drop an image or enter its size to see how many input tokens it costs on each vision model, how OpenAI slices it into 512-pixel tiles, and what a batch of them costs.

Everything runs locally: your image never leaves your browser.

Calculate vision tokens for an image

OpenAI high detail: resized to 1365 × 768, 3 × 2 = 6 tiles.
Input tokens per image
Model Tokens Accuracy $ / 1M input Cost
GPT-4o / GPT-4.1 OpenAI 1,105 Exact $0.00276
GPT-4o mini OpenAI 36,835 Exact $0.00553
o1 / o3 OpenAI 975 Exact —
GPT-4.1 mini OpenAI 2,443 Exact $0.00098
GPT-4.1 nano OpenAI 3,710 Exact $0.00037
o4-mini OpenAI 2,594 Exact $0.00285
Claude (Sonnet, Opus, Haiku) Anthropic 1,609 ≈ Approx. $0.00483
Gemini 2.x Google 1,548 ≈ Approx. $0.00015

Prices are sample list prices and change often — edit them to match your plan. Token counts do not depend on price.

How each provider counts image tokens

OpenAI documents its method exactly. A 1920 × 1080 screenshot in high detail is scaled to 1365 × 768, which spans 3 × 2 tiles of 512 pixels, so it costs 6 × 170 + 85 = 1,105 tokens on GPT-4o. Newer mini and nano models use 32-pixel patches, capped at 1,536, times a per-model multiplier. Both methods here reproduce OpenAI’s own worked examples exactly.

Claude costs about width × height ÷ 750 tokens, after scaling down anything over 1568 pixels on the long edge or about 1,600 tokens. Gemini charges 258 tokens for a small image and 258 per 768-pixel tile above 384 pixels. Neither publishes its exact resizing rule, so those rows are marked as approximations.

A GPT-4o image token cost calculator that shows the tiles

The number people usually need from a GPT-4o image token cost calculator is not just the total but why it is that total — which is why the grid is drawn. Because the cost moves in whole tiles, an image a few pixels past a 512-pixel boundary after resizing pays for a whole extra row or column. As an OpenAI vision image tile calculator, it shows exactly where those boundaries fall.

Where these numbers come from

OpenAI’s tile and patch algorithms and their worked examples are from OpenAI’s vision documentation. Claude’s formula, resize limits and maximum-size table are from Anthropic’s vision documentation. Gemini’s 258-token and 768-pixel tile figures are from Google’s Gemini API documentation.

Frequently asked questions

How does GPT-4o slice an image into 512x512 tiles?
In high detail, the image is first scaled to fit inside a 2048 x 2048 square, then scaled so its shortest side is 768 pixels — both steps only ever shrink it. The result is divided into 512 x 512 tiles, and the cost is 170 tokens per tile plus a flat 85. A 1024 x 1024 image becomes 768 x 768, which is 2 x 2 tiles, so 4 × 170 + 85 = 765 tokens. The grid above draws exactly those tiles on your image.
What is the difference in token cost between detail low and detail high?
With detail set to low, OpenAI charges a flat 85 tokens (for GPT-4o) regardless of size, because the model sees a single 512 x 512 downscaled version. High detail charges per tile, so a 1080p screenshot costs 1,105 tokens instead of 85 — thirteen times more. Low detail is enough for classifying or captioning; high detail is needed for reading small text or fine detail.
How does Claude resize images and count tokens?
Anthropic documents that an image costs roughly width × height ÷ 750 tokens, and that images more than 1568 pixels on the long edge or more than about 1,600 tokens are scaled down first. So nearly every large photo ends up near 1,600 tokens. Anthropic publishes the maximum sizes that avoid resizing but not the exact rule, so this calculator’s Claude figure is an approximation, within about 2% of their published table.
Can shrinking an image slightly lower vision API costs?
Sometimes dramatically. OpenAI’s cost jumps in whole tiles, so an image whose resized width is 1,030 pixels pays for three tiles across where 1,020 pixels would pay for two. Dropping a dimension just below a multiple of 512 after resizing can cut a third off the price. Claude and Gemini also cap large images, so sending anything much bigger than about 1,600 pixels on the long edge usually buys nothing.
Why does GPT-4o mini cost more tokens per image than GPT-4o?
OpenAI prices GPT-4o mini images at 2,833 base tokens plus 5,667 per tile — about 33 times GPT-4o’s count — which offsets its much lower price per token. The dollar cost of an image ends up similar on both, so picking the mini model does not make vision cheaper the way it makes text cheaper.
Is my image uploaded anywhere?
No. The calculator runs entirely as a small script inside this page; a dropped image is opened locally only to read its width and height, and never leaves your browser. There is no upload, no API call and no analytics here. That is enforced rather than promised: this site’s test suite scans the shipped HTML for every browser API capable of sending data off the page and fails the build if it finds one.