How the numbers work
What each figure on this site means, how the plans page estimates a weekly limit and compares models' response times, and where an estimate goes wrong.
Where the numbers come from
Codex writes a log of every session to the machine it runs on, with the tokens it recorded for each response. The token-counter plugin reads those logs on that machine. When a sharer asks, its share skill sends a summary here: counts per day, each month's top sessions, and the weekly limit readings Codex logged beside them.
These are the counts Codex recorded, not a bill. Every figure is self-reported: the server turns away arithmetic no real log produces, but it cannot verify a report.
Terms
- Tokens
- Recorded input plus output. The leaderboard, profiles and the plans page all count this way.
- Recorded input
- Every input token Codex recorded for a response, cached or not.
- Cached input
- The part of recorded input Codex logged as cached. It is a subset of recorded input, never added on top.
- Uncached input
- Recorded input less cached input.
- Cache hit
- Cached input as a share of recorded input.
- Output
- Tokens the model wrote, reasoning included.
- Reasoning
- The part of output spent reasoning. A subset of output.
- Active time
- The time between consecutive responses in a session, leaving out any gap over 30 minutes. Span is the wall-clock time from its first response to its last.
- Days and months
- The sharer's local calendar days, as in the plugin's report. A session belongs to the day its first response landed on.
- Weekly window
- One run of a plan's weekly limit, from one reset to the next. Beside each response, Codex logs the plan and how much of the weekly limit the server says is used. A window starts where that percentage drops, which is not always seven days after the last one started.
- Plan
- The plan type Codex logs beside those readings, such as
plus. It comes from the logs, not from the sharer's account. - Response time
- From the moment a prompt was complete, the sharer's message or the last tool output, to the moment Codex recorded the response: network, queueing, reading the prompt, writing the answer and any retry, together. Codex records no send time, so the plugin reads it off the logs' own timestamps.
- Turn time
- From the sharer's message to the last response of the turn, tools and approvals included.
- Output tokens/s, above pace
- The plugin's estimates, marked with a star wherever they appear. For each model and effort it fits a line under the fastest tenth of responses: a fixed overhead plus a time per output token and per uncached input token. Output tokens/s is that line's rate, and above pace is the share of response time over it. That share is mostly queueing and retries, but slow stretches of writing land there too, so it is not a measured queue time.
How the plans page estimates a weekly limit
- For each sharer and plan: the tokens counted in their recent weekly windows, divided by the percentage points of the limit Codex reported used in them, times 100. Weeks that used only a few points are left out, since rounding in the reported percentage would swamp them. Only recent weeks count, because limits change.
- That is how many tokens a full weekly limit held, as one sharer's logs saw it. The plans page takes the median across sharers, so one unusual account cannot drag it far.
- A plan shows no number until enough sharers, with enough weeks between them, are on it. Then it shows a provisional median, with the full range of sharers' estimates from lowest to highest. With more sharers the median becomes the headline, and its range the middle half of sharers. The plans page says, for each plan, how many sharers its next step takes.
- Every figure it publishes is rounded to two significant figures, and there is no mean: one account could drag a mean anywhere, and the mean before and after a share would give away that sharer's exact estimate.
- The result is an observed range: what sharers' weeks held. OpenAI's pricing page gives no limit in tokens. It says model choice, context, reasoning, tool use, retrieval and caching all affect usage, and that Fast mode draws on included usage at 2.5 times the Standard rate. So the same limit can hold more tokens in a cache-heavy week, and fewer in a week of Fast mode.
How the plans page compares response times
- Each sharer's plugin times their responses over the last 30 days and sends the median and 90th percentile per model and reasoning effort. A sharer counts on a model and effort with 40 or more timed responses on it.
- The plans page takes the median of those sharers' medians, with the same bands as the plans page: a provisional figure from six sharers, with every sharer's median in its range, and a headline from ten, with the middle half. Below six, a model is not listed at all, so one account cannot add one.
- The same goes for each hour of the day and day of the week, in UTC: the median of the sharers' medians for responses that started in it, once six sharers have 20 or more responses there. UTC, because load on Codex follows the clock everywhere at once. Each hour has its own mix of sharers and models, which moves its median as well as load does; the time above the pace, an estimate, takes the model and the size of each response out.
- Response times depend on more than the model: the size of each prompt, how much of it was cached, the sharer's network, and the time of day. Read a figure as what sharers' responses took, not as a model's speed in isolation.
Why an estimate can be off
- Low, most often. Usage the sharer's logs do not see can draw on the same limit: Codex on another machine, cloud chats, ChatGPT Work. OpenAI's pricing page says cloud chats and ChatGPT Work share a plan's limits with Codex. All of these raise the percentage without adding tokens here.
- High, sometimes. Tokens in the logs that did not draw on the plan add tokens without raising the percentage: a session signed in with an API key, which is billed at API rates, or GPT-5.3-Codex-Spark, which launched with limits of its own, until its deprecation in September.
- Out of date. Limits and models change, and an estimate pools several recent weeks. Changes lists the ones that can move it.
Questions
Is this official?
No. tokenusage.dev is an independent site, not affiliated with or endorsed by OpenAI. About says who runs it.
How accurate is it?
The token counts are as good as Codex's logs: the plugin counts a share from the same ledger as its local report, so the two agree day for day. The weekly limit estimate is rougher. It rests on the percentages Codex logged, on how much of a sharer's usage their logs saw, and on a limit that does not weigh every token alike. Read it as the range sharers observed, not as a promise of what a plan holds.
Why do the estimates err low?
Because what the logs miss usually still draws on the limit. Codex on a second machine, a cloud chat or ChatGPT Work raises the percentage and adds no tokens here, so the tokens per point come out smaller than they were.
What is sent?
Counts per day, each month's top sessions by active time and by tokens under a hashed id, the weekly limit readings with the tokens counted beside them, response times per model, hour of the day and day of the week over the last 30 days, the handle you pick and, unless you leave it out, your report page. Never prompts, outputs, file contents, paths, session titles or your account. Privacy says what is kept and what is public.
How do I delete my data?
Ask Codex to remove your data from tokenusage.dev, or run python3 scripts/share.py --delete --yes from the plugin's token-share directory, ~/.codex/plugins/cache/jack-beanstalk-2022/token-counter/*/skills/token-share. It deletes everything you shared, your report page included, and releases your handle. It needs the token from your first share, which the plugin keeps on your machine; if you lost it, see Contact.