AI News Feed
Market watch
Products & Applications

Yandex Unveils Sona, a Single Generative Recommender That Replaces the Full Recommendation Cascade

Yandex introduced Sona, a generative recommender that merges candidate generation and ranking into one served transformer. A seven-day smart-speaker A/B test replaced over 15 candidate generators and two ranking stages, lifting active users 4.53% and listening time 6.30%.

Most production recommenders use cascades: candidate generators feed a pre-ranker, which feeds a heavy ranker built on hundreds of engineered features. Each stage is trained separately and optimizes its own objective, and the ranker sees only candidates that upstream stages pass through. On Yandex Music, the previous stack used hundreds of features, including signals from Argus, an earlier Yandex recommender transformer. Sona instead uses one shared user representation. An encoder reads the listener's history once per request, a decoder generates candidates, and a ranking module scores them against the same encoder states. No component uses hand-engineered features. Inputs are logged event fields such as track ID, artist ID, duration, likes, played time and surface flags, plus learned Semantic IDs.

On smart speakers, playback can begin without the user selecting an artist, genre or mood, a setting the research team calls pure recommendation. Sona's semantic tokenizer turns each track into a tuple of three discrete codes. A frozen multimodal LLM reads the mel-spectrogram of the first 90 seconds along with title, artists and tags in prefill-only mode. A four-layer refinement transformer aligns those features with listening behavior using InfoNCE on collaborative track pairs. Residual K-means quantizes the result into three codebooks of 32,000 entries each. Yandex says this beat a CLMR audio baseline, with Recall@1000 rising from 0.8111 to 0.8524.

The encoder attends to 8,192 past events, but full attention over that length is expensive. It spends depth unevenly: the recent 2,048 events pass through a seven-layer self-attention stack, while older events pass through cross-attention and one full-history layer. The paper reports that this keeps most of the quality of full attention at about half the inference cost. A two-layer decoder emits Semantic ID tuples through constrained beam search with width 1,024, and a catalog trie blocks invalid prefixes. Each tuple expands to every track that shares it. The ranking module, made of four cross-attention layers, then scores those tracks against the shared encoder memory.

The ranking module learns from a frozen teacher ranker, a 0.6-billion-parameter transformer that also avoids hand-engineered features. The teacher is trained on a year of engagement events in two stages: next-item-prediction pre-training and multi-head ranking fine-tuning. Removing pre-training dropped weighted pair accuracy from 0.6215 to 0.6153. Yandex calls its distillation method Rollout Distillation. During training, the current decoder generates beam candidates, and the teacher scores them together with logged impressions. The ranking module regresses onto those scores with mean absolute error. The joint loss is L = L_NTP + L_rollout + L_impression, and both losses update the shared encoder. The teacher is removed at serving time.

Training remains online. Events aggregate into sessions over a 15-minute window, feed a GPU trainer, and new weights reach serving every 10 minutes. End-to-end latency is 45 minutes at the median and 60 minutes at p99. Serving runs on NVIDIA Triton Inference Server with CUDA graphs and reaches 41% model FLOPs utilization.

In the final online A/B test, which ran for seven days on 15% of randomly selected users per split, Yandex reported statistically significant gains relative to the production control. Active users, the primary metric, rose 4.53%. Total listening time rose 6.30%, likes rose 11.42%, repeat commands rose 17.99%, and deeply engaged users rose 7.37%. Yandex says these gains stack on improvements retained from earlier deployments, and that Sona's uplift in active users is 2.35 times the 1.93% increment Argus previously delivered on the same surface.

Sona is not the first end-to-end generative recommender in production. Kuaishou's OneRec already serves a single encoder-decoder model, and Meta's HSTU Generative Recommenders reframed recommendation as sequential transduction over user actions in 2024. According to MarkTechPost, Sona combines full cascade replacement, no hand-engineered features and a distilled ranker, validated online. The report places Sona in music streaming, OneRec in short video and HSTU GR in a large internet platform across multiple domains.