Evidence at a glance
The Problem Has Shifted from Nuisance to Independent Investigation
On September 27, 2026, Simon Willison published Bluesky reply bot checker. It is a web tool for examining a Bluesky profile for signs of automated replies. Rather than classifying a single reply as bot-written, it reads an account’s interaction history and checks for unusual combinations of timing, posting behavior, and interaction targets. Willison used Claude Opus 5.5 to generate the tool, and the code is available through his simonw/tools repository.
The case matters to technical leaders not because another small bot detector exists, but because ownership of anti-abuse investigation is beginning to shift. Willison describes receiving dozens of meaningless automated replies on Twitter, with similar bots later appearing on Bluesky. Whether a platform contains bots remains a platform-governance problem, but Bluesky still provides a free and useful API. That gives users and third-party developers a way to inspect abnormal interaction without waiting for the platform to expose an internal label.
The Evidence Comes from Behavioral Combinations
The checker uses a set of behavioral heuristics. It looks at whether replies appear within seconds of other posts, whether an account publishes little or no original text, images, or links, and whether it repeatedly replies to users with larger follower counts. It also treats question marks as a signal because a question is more likely to make a real person stop and respond, giving an automated account attention or engagement.
None of these clues proves an account’s identity on its own. Rapid replies can come from a highly active human participant. An account with little original material may simply be used that way, and replying to a larger account can have an entirely legitimate social motive. The tool is useful because it combines weak signals rather than turning one signal into a decisive rule. Its output is therefore a level of suspicion that merits inspection, not a confirmed bot identity.
The evidence can be understood as four separate views. Timing asks whether an account responds in dense bursts over very short intervals. Content structure asks whether it lacks a pattern of original posting. Interaction direction asks whether it mainly pursues larger accounts. Question marks add a clue about possible engagement bait. The material provides no signal weights, thresholds, test set, or false-positive rate, so the checker cannot be described as a validated classifier.
The LLM Accelerated the Prototype, Not the Responsibility
The way the tool was built is equally important. Willison did not begin by constructing a complete anti-abuse platform. He built a focused investigative utility around an open API and used Opus 5.5 to help generate the implementation. For this class of small tool, an LLM shortens the distance between an idea and a working prototype, allowing an individual developer to turn behavioral hypotheses into code and an interface more quickly.
Development speed and governance reliability are different things. A model can help generate code for reading data, organizing rules, and displaying results, but it cannot decide what level of anomaly should trigger intervention. It cannot take responsibility for wrongly targeting a legitimate account either. The material does not explain what tests were run on the generated code or what accuracy it achieved on real account samples. The public availability of the tool therefore does not establish that it is suitable for automated bans or large-scale operations.
For a technical leader, the more precise conclusion is that LLMs increase the supply of anti-abuse prototypes. As implementation becomes cheaper, scarce effort moves toward signal selection, evidence preservation, and review design. A public checker that only returns “bot-like” without showing the behavior that triggered the result would replace one opaque platform label with another opaque AI label.
An Open Interface Moves Governance into Third-Party Workflows
In operations, the checker is best positioned as a ranking layer before human review. A team can first use reply timing, account structure, and interaction direction to identify accounts that deserve attention, then return to the individual posts and their context. The benefit is lower investigation cost, not a direct conversion from a heuristic score to a ban.
The value of an open API is also more than convenient data access. It allows users, researchers, and independent developers to form different explanations of the same abnormal interaction and to build replaceable anti-abuse tooling. The platform still controls service and data boundaries, but it is no longer the only party able to observe behavior. Part of governance can therefore move from an internal feature into an externally auditable workflow, which is useful when an investigation needs to preserve evidence rather than wait for an opaque official conclusion.
Such a workflow can turn a one-time suspicion into a reviewable record. An operator can see why an account was ranked highly, verify whether the replies actually form an unusual pattern, and then decide whether to report, block, or keep watching it. The material does not explain how Bluesky itself handles these bots or how this checker relates to official enforcement. It should therefore not be presented as a replacement for platform governance.
Detection and Evasion Become Easier at the Same Time
The paradox of an open interface is clear here. It lowers the barrier for third-party investigation, but it also lowers the barrier for automated accounts to read the network, study the rules, and probe the defense. Once detection logic is visible, a bot may alter its reply intervals, add a small amount of original content, or avoid repeatedly targeting obviously larger accounts. Fixed rules can then become easier to evade.
Anti-abuse tooling must therefore be treated as a continuously updated investigation layer rather than a security boundary that remains effective after one deployment. The more understandable the behavioral clues are to users, the easier they are to use and the easier they are for adversaries to study. Teams have to balance explainability against resistance to evasion, but the material does not provide enough evidence to judge how this tool performs in an adversarial environment.
The actionable conclusion should remain restrained. The checker can narrow the scope of human investigation, serve as a replaceable component built on a public API, and help test whether a platform allows independent anti-abuse perspectives. But no single clue, including a question mark, a reply within seconds, or a lack of original posts, should become a banning rule. Nor should the fact that Opus 5.5 generated the prototype be mistaken for proof that the detector is reliable.
The case ultimately describes a new division of governance rather than a product that has solved bots. The platform exposes an observable interface. Developers organize behavioral clues into tools. Ope