Evidence at a glance
A Composition Test That Also Builds a Tool
On October 6, 2026, Simon Willison introduced Scrimshaw Jukebox as a test of whether Claude Opus 5.5 could compose computer game music. His prompt did not simply ask for a tune. It asked the model to design a simple text-based music format, build an interactive page that could play example tracks, and aim for the sound of the original The Secret of Monkey Island score. The result was not just a recording, but a browser tool that could hold and play compositions.
Willison described the result as “surprisingly good” and said it leaned more heavily into The Secret of Monkey Island than he had intended. Those observations explain both the appeal of the demonstration and the tension in evaluating it. A close stylistic resemblance can make an output feel immediately on target, but it does not by itself show that the model has a stable ability to compose. To understand the case, separate the question of whether the music sounds good from whether the model can assemble a workable, editable creative process.
Text Notation Makes the Music Operable
The tool does not turn a text prompt directly into audio. It first represents the music as structured text, with fields for tempo, time signature, voices, and sections. A parser checks and compiles that notation, after which the browser plays it. Users can edit the text, play the revised version, or mute an individual voice. The composition is therefore not a sealed finished product, but something that can be inspected, separated into parts, and changed repeatedly.
This representation also puts listening and composing in the same workflow. Because the prompt asked the model to design both the notation and the player, the task involved defining expressive rules, organizing music according to those rules, and providing an interface for hearing the result. For a technical lead, that adds a capability worth examining beyond “the model wrote a tune”: can it define a clear intermediate representation and deliver a usable tool around it? The available material does not show that this format can be imported into other music software, so for now it is best understood as an editable demonstration environment.
The Browser Synthesizes the Sound, Not the Model Alone
The playback mechanism changes how the result should be attributed. The project source says the page uses no sampled audio. Instead, it generates sound in the browser through the Web Audio API, using oscillators to produce tones and filters and volume envelopes to shape timbre and note dynamics. What the listener hears is therefore the combined result of the text score, its parsing and compilation logic, and the synthesizer. It is not equivalent to the model directly generating a recording.
That distinction does not diminish the demonstration. It clarifies what the system actually shows: the model contributes to structured musical expression and tool construction, while the synthesizer determines how those symbols become sound. A team considering a similar workflow for game-music drafts should evaluate more than the final playback. It should also ask whether the score is easy to edit, whether voices can be adjusted independently, and whether the work still meets expectations with a different playback implementation. The available material demonstrates the first two interaction points, but provides no comparison across players or evidence about timbral quality.
Six Tracks Show Range, Not an Evaluation Result
The page lists six example tracks, with durations of roughly 56 seconds to 2 minutes 11 seconds. Their displayed tempos range from about 66 to 152 BPM, and the examples use 3/4, 4/4, and 6/8 time signatures. The format guide allows tempos from 20 to 400 BPM. It organizes time in steps per bar: in 4/4, dividing each beat into four steps gives 16 steps per bar. These figures show that the tool can present playable examples with varied rhythms and lengths, and that it treats music as a time structure that can be manipulated explicitly.
But track count, tempo range, and meter variety are not substitutes for musical-quality measures. They show that the page offers several playable examples with different settings. They do not tell us whether the melodies are original, whether the arrangements are mature, or whether the model can maintain quality across repeated tasks. Willison's listening impression is a useful observation, but it remains one person's subjective judgment. The available material includes no independent listening panel, blind test, or comparison between newer and older models, so it cannot establish that composition ability has only recently emerged.
Treat It as a Creative Prototype, Not a Verdict
For a working team, the most useful lesson is to expand evaluation from a finished track to the whole creative path. Check whether the model can propose readable musical notation, generate tracks that parse correctly, and provide a tool that makes listening and revision convenient. Then assess separately whether the resulting music suits the intended use. Breaking the workflow into parts helps prevent the playback system's sound from being mistaken for the model's compositional quality, and makes it clearer where a person needs to intervene.
To determine whether capability has improved, an experiment would need to control for model version, prompt, notation format, and playback mechanism, and have reviewers listen under reasonably consistent conditions. In particular, stylistic fit should be evaluated separately from originality. This result leaned more heavily into The Secret of Monkey Island than expected, but the material provides no melody-similarity review and does not demonstrate that the style can be dialed up or down as requested. Scrimshaw Jukebox is therefore a persuasive end-to-end prototype, not complete evidence about a model's musical ability. It can be used to explore an editable draft-music workflow, while style control, originality, and performance across models remain questions for dedicated evaluation.