YuE2 · Frontier Music with Symbolic Planning

Sep 11, 2026 07:33 AM - 3 days ago 4

Listen to a song, past research the melody, rhythm, and chords successful its symbolic plan.

Selected score

Loading selected song…

The symbolic score

Original people recording

Interactive ABC score Red notes travel the people recording.

View original people pages
All selected scores

A acquainted opus tin return a different shape. Listen to changes successful melody, lyrics, tempo, and arrangement.

Agentic euphony editing

A opus takes style done a conversation. Loading the editing story…

The listening selection, gathered crossed genres and languages.

Search Language Genre Generation

Model & Results

YuE2 (best-of-8) reaches 6.9632 connected SongBench, the highest observed mean among 15 evaluated settings connected WildSongBench (192 prompts). Suno v5 scores 6.8721 successful the aforesaid comparison.

 a WildSongBench comparison of opus value and matter alignment. YuE2 and YuE2 best-of-8 are competitory pinch the evaluated proprietary systems. Bubble area represents AudioBox accumulation quality. Song value and matter alignment connected WildSongBench. Bubble area shows AudioBox PQ; achromatic outlines people Pareto optima connected the 2 plotted axes. Bo8 = best-of-8. How to publication the indices.

Model architecture

Composing successful symbols, performing successful audio.

Vector PDF

 an AR–NAR Mixture-of-Transformers turns an editable symbolic people into semantic tokens, acoustic latents, and full-song audio. The aforesaid generator supports creation, covering, and editing. An editable people becomes semantic euphony tokens, acoustic latents, and full-song audio. The exemplary has astir 3.59B parameters and 28 layers, and supports creation, covering, and editing. The AR and NAR experts stock an attraction computation while utilizing abstracted normalization, projections, and MLPs.
Explore benchmark scores WildSongBench · 15 settings · 7 metrics

WildSongBench192 prompts

Compare by

How to publication these results

WildSongBench. 192 prompts and 15 strategy settings. The array reports automatic information scores. Best-of-8 selects 1 of 8 generations by musicality, punctual control, and lyric accuracy.

Figure 1. Song value combines SongBench and SongEval; matter alignment combines MuLan, AllMusicCaps, and punctual control. Both axes show normalized comparison indices. Bubble area represents AudioBox accumulation quality.

MERT2 · Music representations

Learning the building down the sound.

Vector PDF

 offline target synthesis combines MuQ and Qwen2-Audio features into 4 codification streams. A ConvNeXt frontend and 24-layer Conformer study by masked prediction, past branch into full-song representations for SheetSage2 and a causal tokenizer program for YuE2. MERT2 provides the euphony representations down some study and generation. A ConvNeXt frontend and 24-layer Conformer study to foretell masked codes built from complementary MuQ and Qwen2-Audio features. Full-song adjustment supplies SheetSage2 pinch philharmonic context; a abstracted causal branch becomes YuE2’s 25-Hz semantic tokenizer.

State of the creation connected MARBLE

MERT2-30s and MERT2-FS (full-song) execute SOTA connected 14 of 15 MARBLE metrics, starring crossed tagging, key, genre, and emotion recognition.

SOTA metricsMERT2-30s & MERT2-FS14 / 15

Genre accuracy · GTZANMERT2-30s · people × 10091.72

Key refined accuracy · GiantStepsMERT2-FS · people × 10067.05

Explore MERT2 benchmark scores MARBLE · 15 metrics · 2 encoders

SOTA counts usage the champion people crossed the 2 MERT2 encoders against the 9 published baselines successful this comparison. Both encoders person 632M parameters. MERT2-30s uses a 30-second training context; MERT2-FS uses 300 seconds. These are full-context practice benchmarks. MERT2 reports the champion observed results crossed representations selected utilizing trial scores; each ROC-AUC / AP brace uses the aforesaid representation.

SheetSage2 · Audio to score

Hear a song. Read its composition.

Vector PDF

 a full-song MERT2-FS encoder pinch trainable adapters feeds a six-layer autoregressive decoder. Task prompts prime beat, section, key, chord, and melody events, which stock a timeline and go ABC notation and a lead sheet. SheetSage2 turns a signaling into an editable lead sheet. A full-song MERT2-FS encoder, adapted pinch LoRA, feeds a six-layer autoregressive decoder. Task prompts prime beats, sections, keys, chords, and melodies; a shared arena timeline becomes ABC notation pinch vocal and instrumental melody voices. These scores proviso symbolic training targets for YuE2.

Six transcription tasks, 1 model

SheetSage2 achieves SOTA connected 10 of 13 benchmark metrics pinch 1 exemplary for beat, downbeat, key, chord, structure, and melody transcription.

SOTA metricsOne exemplary · six transcription tasks10 / 13

Vocal melody · RWC-PopPitch-class statement F1 · people × 10082.51

Chord nickname · osu2017Maj/min · people × 10090.08

Explore SheetSage2 benchmark scores 6 tasks · 13 metrics

SOTA counts mention to the starring scores against SheetSage1, Madmom, and the task-specific systems successful this comparison. Results usage 1 exemplary selected by validation loss. Melody F1 measures pitch-class notes; building F1 measures conception boundaries astatine the stated tolerance. On Chords1217, ChordFormer uses five-fold cross-validation, while SheetSage2 evaluates 1 fixed exemplary connected each 1,217 tracks.

Training data

Our models are trained chiefly connected CC0 euphony and synthetic data. Tokenwave.AI provides astir of our synthetic training information nether license. We are committed to the ethical and responsible usage of data.

MERT2700K hours

SheetSage228.4K hours

YuE2346K hours

More