01 / decision
Decision snapshot
Each figure sits next to the middle of the field, so it can be read as dear or cheap, wide or narrow, rather than floating on its own. Compared against the 370 models tracked here, not against an absolute standard.
Capability
64.6/100
#16 of 35 with published scores
Input / 1M tokens
$15.00
field median $0.40
Output / 1M tokens
$60.00
field median $1.68
Context
200K
field median 262K
Read the coverage before the score. That figure rests on 20% of the weighting: the rest of the categories have no published evidence for this model, and are not counted rather than counted as zero.
02 / overview
What o1 is, and when to reach for it
Four questions, answered separately, because somebody arrives at one of them rather than at the top of the page.
This model has no written explanation yet.
The page is created the moment a model appears in a provider's catalogue, and the writing follows. Until then the measured sections below are the whole page.
What counts as evidence03 / evidence
How much of this is verified
Split by category, so a strong number never hides a thin evidence base. Verified means it was read on the benchmark's own published results; a provider's claim about its own model is shown and labelled rather than dropped.
Published rows
1
of 19 benchmarks tracked
Independently verified
0
1 from the provider
Weighting covered
20%
the rest is not measured, not zero
Agentic
0/3
Not measured
Coding
1/3
Provider claim
Reasoning
0/3
Not measured
Multimodal
0/3
Not measured
Knowledge
0/2
Not measured
Multilingual
0/2
Not measured
Instruction following
0/1
Not measured
Maths
0/2
Not measured
04 / ledger
Benchmark ledger
Every published row, grouped by category, each compared with the best published score on the same benchmark. 'Is 64% good' is a question nobody can answer; '26 points behind the leader' is one anybody can.
Coding1 row
| Benchmark | Score | Best published | Weight | Evidence |
|---|---|---|---|---|
| SWE-bench VerifiedReal GitHub issues, resolved and tested. | 64.6% | Gemini 3 Flash Preview75.8% · 11.2 behind | 40% | Provider claim |
05 / capability
Capability shape
Where this model is strong, and against how many peers. Ranks are against models with evidence in that category, not against every tracked model: ranking against models nobody tested would rank who published, not who is better.
| Category | Score | Rank | Weight | Benchmarks | Evidence |
|---|---|---|---|---|---|
| Agentic | Not measured | Not ranked | 22% | 0 of 3 | Not measured |
| Coding | 64.6 | #16 of 3556th percentile | 20% | 1 of 3 | Provider claim |
| Reasoning | Not measured | Not ranked | 17% | 0 of 3 | Not measured |
| Multimodal | Not measured | Not ranked | 12% | 0 of 3 | Not measured |
| Knowledge | Not measured | Not ranked | 12% | 0 of 2 | Not measured |
| Multilingual | Not measured | Not ranked | 7% | 0 of 2 | Not measured |
| Instruction following | Not measured | Not ranked | 5% | 0 of 1 | Not measured |
| Maths | Not measured | Not ranked | 5% | 0 of 2 | Not measured |
Category weights are WriteWorks' own and they are a judgement: weighted for what a team buys a model to do. Agentic and coding lead because that is where the work is. Published here rather than hidden, and set out in full on the methodology page.
06 / cost
What it costs
List API rates as last read, each with the date it was confirmed, plus every change recorded since tracking began.
| Charge | Price | Unit | Read on |
|---|---|---|---|
| Input | $15.00 | per 1M tokens | 2026-09-30 |
| Output | $60.00 | per 1M tokens | 2026-09-30 |
| Workload | Input / month | Output / month | Cost |
|---|---|---|---|
| A small product team | 20M tokens | 5M tokens | $600.00 |
| A busy support assistant | 200M tokens | 40M tokens | $5,400.00 |
| A document pipeline | 1000M tokens | 100M tokens | $21,000.00 |
List API rates, no caching and no batch discount, which both providers offer and which change the answer a great deal. Treat these as the ceiling, not the bill.
07 / specs
Specifications
As listed by the provider's own catalogue and re-read every few hours. Anything absent is absent there too.
| Context window | 200,000 tokens |
|---|---|
| Maximum output | 100,000 tokens |
| Modalities | text, image, file |
| Released | 17/12/2024 |
| Status | Current |
08 / lineage
Lineage
What this model replaced, what replaced it, and what else its provider has in the field.
Also from OpenAI
- GPT-6.1 Sol29/09/2026
- GPT-6.1 Sol Pro29/09/2026
- GPT-6 Luna22/09/2026
- GPT-6 Luna Pro22/09/2026
- GPT-6 Sol22/09/2026
- GPT-6 Sol Pro22/09/2026
- GPT Astra Latest11/09/2026
- GPT Luna Latest11/09/2026
- GPT Sol Latest11/09/2026
- GPT Terra Latest11/09/2026
- GPT-6 Astra04/09/2026
- GPT-6 Astra Pro04/09/2026
09 / line
The line
Every model in this family in release order, so a page from eight months ago says in one glance that two newer ones exist.
Nothing else in this line yet.
A line is worked out from the naming and the release dates across every model page on the tracker. It fills in as the provider ships successors, or as the models that came before this one are added.
What counts as evidence10 / notes
WriteWorks notes
What this model changes for a brand trying to be cited in AI answers, and every change logged since it launched.
Change log
Every release, price and feature change for o1, newest first. They also appear on the OpenAI page.
Nothing published here yet. Changes appear within hours of a provider announcing them.
11 / questions
Questions
The things people ask about this model, answered from what is on this page rather than from anywhere else.
- What does o1 cost?
- $15.00 per million input tokens and $60.00 per million output tokens, as last read from the provider. The cost section works that into a monthly figure.
- How current is this page?
- The catalogue behind it is re-read every three hours, and the stamp at the top says when it last confirmed. A re-check that finds nothing changed updates that stamp and deliberately does not touch the page’s modified date.
- Why are some sections empty?
- Because nothing has been published that can be linked to. An empty section is better than a number you cannot check. The methodology sets out what counts.
Is o1 recommending you?
Models change what gets cited. WriteWorks tracks whether 10+ AI platforms, from ChatGPT and Claude to Gemini and Perplexity, name your brand or your competitors, and shows you the content gaps to close.