An Ordinary Upgrade Exposes a Default’s Time Lag
Simon Willison released ttok 1.0 on October 9, 2026. Ttok is a command-line tool that counts text using a selected tokenizer, allowing a user to check how many tokens an input would occupy under a model’s tokenization rules. It is not a complex model-calling system. It addresses an easy-to-overlook preliminary question in model workflows: before sending text to a model, whose rules are being used to estimate its length?
The release began with a small discovery. After publishing ttok 0.4, Willison ran `uv tool upgrade ttok` and piped a file into the new version. He found that the default was still the GPT-4 tokenizer, judged that it should clearly move to GPT-5/GPT-6, and used that change as a reason to ship 1.0. The change was not in the user’s input or in model inference. It was in the counting assumption used by the same command after an upgrade.
That is why defaults deserve a technical lead’s attention. A default hides a choice that users might otherwise have to make explicitly. For a quick check, avoiding an extra option is convenient. In an automated workflow, however, a default is a dependency that may not appear in configuration, and that dependency can shift when the tool version changes.
Thirty-One Fixtures Support the Choice, Not an Official Claim
Willison supports the new default with a comparison experiment by William Liu. The test covered seven models: GPT-5.5; the Sol, Terra, and Luna variants of GPT-5.6; and the Astra, Sol, and Luna variants of GPT-6. The report says that all seven matched on each of 31 fixtures, with a total of 44,794 tokens. On those inputs, GPT-6 made no difference to the input count.
Those results make the default change understandable. If a tool needs a convenient default for GPT-5 and GPT-6, matching counts across all 31 fixtures for seven models is more persuasive than choosing based only on a model name or a guess. The experiment supports an operational judgment: on the fixtures tested, the GPT-5-family tokenizer can serve as a practical approximation for GPT-6.
But OpenAI has not confirmed that GPT-6 uses the same tokenizer as the GPT-5 family, and the source notes an angry GitHub issue about the question. Matching fixtures do not establish that the underlying implementations are identical, nor do they guarantee matching results for every text, special token, or model-side processing path. The supplied material does not describe the composition or coverage of the 31 fixtures. So 44,794 is evidence about this experiment, not proof of universal compatibility.
A Token Count Is an Input to the Workflow
A tokenizer determines how text is divided into tokens, so a count is not a property of text independent of the model. It is the result of applying particular tokenization rules. If a tool selects the wrong tokenizer, its number may no longer represent the counting behavior of the intended model. Even when a discrepancy affects only some inputs, it can become a workflow decision if downstream steps use the count as a threshold.
For example, a team might use token counts to estimate whether an input fits a budget or to decide whether text should be truncated. In that case, the counting rule helps control what is actually sent to the model. If counts are used to filter data or compare models, the same number can also affect which examples remain and which results are considered comparable. These are possible roles for counts in automated workflows, not claims that the source proves all ttok users rely on them in this way.
A default update therefore brings both a benefit and a migration cost. For everyday checks aimed at GPT-5/GPT-6, moving away from GPT-4 can reduce the mismatch between the default and the user’s likely target. But when scripts, documentation, or team habits rely on “whatever happens if nothing is specified,” upgrading the tool can change counts without changing the call itself. The command still looks familiar, while the rule behind its output has changed.
Treat the New Default as Convenience, Not a Compatibility Verdict
For interactive users, ttok 1.0’s choice has a clear practical logic. Based on the comparison available to him, the author moved the default tokenizer closer to the likely target of GPT-5/GPT-6 users instead of retaining GPT-4. For a command-line tool used to check a file quickly, that reduces the friction of reconsidering the default model on every invocation. A default can be useful before every uncertainty has been resolved, provided users understand the limits of the evidence behind it.
Teams that rely on counts for budgets, truncation, or regression tests should be more explicit. Pinning the target model or tokenizer in configuration or at the call site makes the assumption visible and reduces the chance that a tool upgrade silently changes it on the team’s behalf. When upgrading, teams should recheck results on text representative of their own work, especially when counts trigger hard limits. The supplied material offers neither broader testing nor official confirmation, so it cannot establish that this approach covers every edge case.
The practical judgment has two parts. Using the GPT-5-family tokenizer as ttok 1.0’s convenient default is supported by an experiment. Stating that GPT-6 has been confirmed to use the same tokenizer goes beyond the evidence. A technical lead need not reject the new release, but should distinguish “convenient to use for now” from “proven fully compatible.” When a count determines system behavior, explicit configuration is safer than reliance on a default.