Facts come from the provider, never from another tracker
Prices, context windows, release dates and deprecations are read from the provider's own catalogue, documentation or announcement, and every published entry links to it. Other trackers are read only to notice what has been missed, and that is all they are used for: a lead from one opens a task to go and check the provider, and the software refuses to publish it as anything else.
Sources are re-read on a schedule, not on demand
Model catalogues are re-read every three hours, provider pricing pages every six, and the cross-checks daily. Each source records when it last answered and whether it answered cleanly, and a source that has been failing shows as failing rather than as quiet.
Four dates, each with one job
When it happened, how precisely that is known, when it was detected, and when the source was last re-read and found unchanged. They are kept apart on purpose: a re-check that finds nothing new updates the verification stamp and does not touch the page's modified date, because telling a search engine a page changed when it did not is a claim that gets discounted, and deserves to be.
Where the month in a title comes from
A title reading "(September 2026)" means the page was verified in September 2026, not that it was generated then. If a page has not been successfully re-checked inside the current month, the month does not roll and the page says when it was last confirmed.
What publishes without a person
Only a price change read from a structured catalogue, and only when the move is within a plausible range. A change larger than fivefold, a move to or from free, a page whose content hash shifted, and anything written from a news feed all wait for a person. A hash says a page differs; it does not say what differs, and a footer year rolling over produces the same signal as a price cut.
A page can exist without asking to be ranked
A model page is created as soon as the model appears in a provider's catalogue. Whether it is offered to search engines is a separate decision, made on whether anybody searches the model's name. Pages that do not clear that bar stay useful, stay linked and carry noindex, because a folder full of pages nobody searches for makes the pages people do search for harder to find.
Benchmark scores carry their evidence state
A score is labelled verified when it was read on the benchmark's own published results, and provider claim when the only source is the company that made the model. Both are shown. Dropping everything that cannot be independently verified would tell a reader less than showing it and saying which is which, and it would quietly favour the labs that publish least. Where sources disagree the row is marked disputed and both figures stay.
The category weights are a judgement, and here it is
A model page reports eight categories, weighted: agentic 22%, coding 20%, reasoning 17%, multimodal 12%, knowledge 12%, multilingual 7%, instruction following 5%, maths 5%. They are weighted for what a team buys a model to do, which is why agentic and coding lead and maths does not: a model that wins a maths olympiad and cannot use a tool is no use on a Tuesday. Argue with the weights, they are WriteWorks' own rather than anybody's standard.
An unmeasured category is not a zero
Where a model has no published evidence in a category, the category is marked not measured and left out of the headline figure, which is then re-weighted across what remains. Scoring an absence as nought would rank a model that was never tested below one that was tested and did badly, and there is no basis for that. Every page shows how much of the weighting its score actually rests on, next to the score.
Where the benchmark scores come from
Two sources today. Artificial Analysis run their own evaluations and publish the primary metrics, which is where MMLU-Pro, GPQA, LiveCodeBench, AIME, MATH-500 and Humanity's Last Exam come from; their data is used under their free API terms, which require attribution, and every page showing one of their figures credits them. SWE-bench supplies its own leaderboard, including whether it has checked a submission's logs, which is what decides whether a row is labelled verified or a submitter's claim. Nobody's composite index is imported: that is their weighting of the same tests, and taking it alongside the tests would count one piece of evidence twice.
A model with no published score is not ranked
The leaderboard lists models with published benchmark results on record, and nothing else. An untested model does not appear at the bottom, because appearing at the bottom is a claim that it was measured and found wanting. This is also why the ranking is small next to the number of models tracked: a short ranking that is checkable is worth more than a long one that is mostly inference.
No ranking on price or context window
Both are published on every model page and neither feeds the ranking. A table built from prices and context lengths sorts cleanly, updates itself and answers a different question from the one a leaderboard is asked. Where a benchmark score does not exist, the honest output is an empty row, not a proxy.
Corrections happen in public
A wrong number is corrected on the page it is on, and the change is logged as its own entry with the date. Nothing is quietly edited: a tracker whose history can be rewritten is a tracker with no history.