Evidence at a glance
What Sona Consolidates Is the User State, Not Every Decision
In its Sona technical report, Yandex describes a generative music recommender for My Vibe on Yandex Music smart speakers. The system can start playback without requiring a user to choose an artist, genre, or mood first. In an online experiment, Sona replaced more than 15 candidate generators in the previous pipeline, along with its pre-ranking and ranking stages. The system targets the conventional cascade, where candidates are found first and then filtered through successive stages.
This change merits close attention from technical leaders because it does not simply rename a pipeline as “one model.” Instead, it attempts to make candidate discovery and the choice of what to play first use the same representation of a listener's history. In a cascade, stages often optimize separate objectives. A final ranker can only work with songs admitted by upstream stages, and it cannot recover candidates that were omitted. Sona puts generation and ranking in a shared context to reduce objective mismatch, but it still generates candidates and still scores them.
Songs First Need a Discrete Representation the Model Can Generate
Sona does not generate natural-language descriptions for every track in a large catalog and treat those descriptions as recommendations. It first maps each track to three Semantic IDs, a discrete code that the model can generate autoregressively. The report describes a pipeline in which a frozen multimodal model reads the first 90 seconds of audio along with the track's title, artist, and tags. A refinement model then aligns the representation with collaborative listening data, and residual K-means quantizes it into three codebooks, each containing 32,000 codes.
These codes are not the final recommendation decision. They turn the catalog into a space in which candidates can be generated and searched. The report gives Semantic IDs a Recall@1000 of 0.8524, compared with 0.8111 for the CLMR audio baseline. That result supports the usefulness of the representation for candidate recall, but it does not by itself establish an improvement in the online experience. Encoding also creates a boundary: if a track has no suitable representation or is not effectively retrieved, the ranking module never gets the chance to promote it.
Even with Shared History, Candidate Generation and Ranking Remain Distinct
For an online request, the encoder reads a user's interaction history, up to 8,192 events. Computation is not distributed evenly across that history. The most recent 2,048 events receive deeper self-attention, while older records take a lighter cross-attention path, followed by a layer that processes the full history. The report says this compression preserves most of the quality of full-history attention at about half the inference cost. The trade-off is explicit: older behavior remains available to the model, but recent behavior receives more computation.
Next, a two-layer decoder generates Semantic IDs using beam search with a width of 1,024. A catalog trie blocks invalid code prefixes, and the generated codes are mapped back to tracks. A Ranking Module with four cross-attention layers then scores those tracks. “A single generative recommender” therefore does not mean that serving consists of one undifferentiated computation block. More precisely, the encoder, decoder, and ranking module form one serving pipeline and share the user state produced by the encoder.
The Teacher Stays in Training, but Online Updates Still Take Time
Sona does not gain ranking capability by eliminating a teacher model. During training, a 0.6-billion-parameter Teacher Ranker scores candidates generated by the decoder as well as tracks present in logged impressions. Sona's Ranking Module learns those scores, while the teacher itself is absent from online inference. The training objectives also include next-item prediction, rollout distillation on generated candidates, and learning from impression data. Together, these objectives update the shared encoder. In an ablation, removing teacher pre-training reduced weighted pair accuracy from 0.6215 to 0.6153, indicating that the teacher's supervision is not incidental to this training setup.
Separating training from serving reduces online dependence on the teacher, but it does not make the system an instantly adapting black box. Events are aggregated into sessions over 15-minute windows and sent to a GPU trainer. New weights are delivered to serving every 10 minutes. The report gives an end-to-end update latency of 45 minutes at the median and 60 minutes at p99. For rapidly changing user intent, new behavior will not immediately affect recommendations. For operations teams, the refresh cadence and feedback delay must therefore be treated as system design constraints, not confused with the latency of a single inference.
Short-Term Gains Are Evident, but Coverage and Generalization Remain Open
The final online experiment ran for seven days on My Vibe on smart speakers, not across every recommendation surface in Yandex Music. The paper says that the control and Sona groups each contained 15% of randomly selected users. Relative to the contemporaneous production control, Sona increased active users by 4.53%, total listening time by 6.30%, likes by 11.42%, Repeat commands by 17.99%, and deeply engaged users by 7.37%. The report describes all five results as statistically significant. These figures provide online evidence for the combined system design, but they do not isolate the contribution of each component.
The experiment's scope and its coverage limitation must be read alongside those numbers. My Vibe is a passive-consumption setting where users do not specify content before playback begins. Plays, skips, likes, and repeat listening have particular meanings in that context, and the results cannot directly stand in for search or other recommendation surfaces. Yandex also discloses that Sona's catalog coverage is lower than the previous production system, without giving the size or cause of the gap. The system has not been deployed to all traffic and needs further experiments in other settings. For teams with existing cascades, the practical judgment is not to merge every stage immediately. Test shared user state first on a small share of traffic in a comparable passive-consumption setting, make catalog coverage a guardrail, and check whether engagement gains come at the expense of long-tail content.