OpenAI model

o3 Mini

OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model supports the `reasoning_effort` parameter, which can be set to...

OpenAIVerified 05/10/2026Released 31/01/2025

01 / decision

Decision snapshot

Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 370 models tracked here, not against an absolute standard.

Capability

42.4/100

#26 of 35 with published scores

Input / 1M tokens

$1.10

field median $0.40

Output / 1M tokens

$4.40

field median $1.68

Context

200K

field median 262K

Read the coverage before the score. That figure rests on 20% of the weighting: the rest of the categories have no published evidence for this model, and are not counted rather than counted as zero.

02 / overview

What o3 Mini is, and when to reach for it

Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.

This model has no written explanation yet.

The page is created the moment a model appears in a provider's catalogue, and the writing follows. Until then the measured sections below are the whole page.

What counts as evidence

03 / evidence

How much of this is verified

Split by category, so a strong number never hides a thin evidence base. Verified means it was read on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.

Published rows

1

of 19 benchmarks tracked

Independently verified

1

0 from the provider

Weighting covered

20%

the rest is not measured, not zero

Agentic

0/3

Not measured

Coding

1/3

Verified

Reasoning

0/3

Not measured

Multimodal

0/3

Not measured

Knowledge

0/2

Not measured

Multilingual

0/2

Not measured

Instruction following

0/1

Not measured

Maths

0/2

Not measured

04 / ledger

Benchmark ledger

Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.

Coding1 row
BenchmarkScoreBest publishedWeightEvidence
SWE-bench VerifiedReal GitHub issues, resolved and tested.42.4%Gemini 3 Flash Preview75.8% · 33.4 behind40%Verified

05 / capability

Capability shape

Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against every tracked model: ranking against models nobody tested would rank who published, not who is better.

Category recordWriteWorks
CategoryScoreRankWeightBenchmarksEvidence
AgenticNot measuredNot ranked22%0 of 3Not measured
Coding42.4#26 of 3526th percentile20%1 of 3Verified
ReasoningNot measuredNot ranked17%0 of 3Not measured
MultimodalNot measuredNot ranked12%0 of 3Not measured
KnowledgeNot measuredNot ranked12%0 of 2Not measured
MultilingualNot measuredNot ranked7%0 of 2Not measured
Instruction followingNot measuredNot ranked5%0 of 1Not measured
MathsNot measuredNot ranked5%0 of 2Not measured

Category weights are WriteWorks' own and they are a judgement: weighted for what a team buys a model to do. Agentic and coding lead because that is where the work is. Published here rather than hidden, and set out in full on the methodology page.

06 / cost

What it costs

List API rates as last read, each with the date it was confirmed, plus every change recorded since tracking began.

o3 Mini API pricingWriteWorks
ChargePriceUnitRead on
Input$1.10per 1M tokens2026-09-30
Output$4.40per 1M tokens2026-09-30
What a month costsWriteWorks
WorkloadInput / monthOutput / monthCost
A small product team20M tokens5M tokens$44.00
A busy support assistant200M tokens40M tokens$396.00
A document pipeline1000M tokens100M tokens$1,540.00

List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.

07 / specs

Specifications

As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.

SpecificationWriteWorks
Context window200,000 tokens
Maximum output100,000 tokens
Modalitiestext, file
Released31/01/2025
StatusCurrent

08 / lineage

Lineage

What this model replaced, what replaced it, and what else its provider has in the field.

Also from OpenAI

All 65 OpenAI models

09 / line

The line

Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.

Nothing else in this line yet.

A line is worked out from the naming and the release dates across every model page on the tracker. It fills in as the provider ships successors, or as the models that came before this one are added.

What counts as evidence

10 / notes

WriteWorks notes

What this model changes for a brand trying to be cited in AI answers, and every change logged since it launched.

Change log

Every release, price and feature change for o3 Mini, newest first. They also appear on the OpenAI page.

Nothing published here yet. Changes appear within hours of a provider announcing them.

11 / questions

Questions

The things people ask about this model, answered from what is on this page rather than from anywhere else.

What does o3 Mini cost?
$1.10 per million input tokens and $4.40 per million output tokens, as last read from the provider. The cost section works that into a monthly figure.
How current is this page?
The catalogue behind it is re-read every three hours, and the stamp at the top says when it last confirmed. A re-check that finds nothing changed updates that stamp and deliberately does not touch the page’s modified date.
Why are some sections empty?
Because nothing has been published that can be linked to. An empty section is better than a number you cannot check. The methodology sets out what counts.
WriteWorks

Is o3 Mini recommending you?

Models change what gets cited. WriteWorks tracks whether 10+ AI platforms, from ChatGPT and Claude to Gemini and Perplexity, name your brand or your competitors, and shows you the content gaps to close.