Vevo2: separating content, timbre and prosody in zero-shot singing
Content-style and prosody tokenizers, flow-matching acoustics and vocoder reconstruction in a reference-driven singing path.

SINGING SYNTHESIS / TECHNICAL NOTE
SoulX-Singer accepts lyrics and MIDI as an explicit score. Vevo2 is reference-driven: content-style and prosody tokens are extracted from example audio before an acoustic model builds the target waveform. The two paths should not share one input contract.
Why two tokenizers
If one token stream carries phonemes, pitch, rhythm and timbre, replacing one factor perturbs the others. The Vevo route reduces that coupling with separate quantizers. For acoustic latent z, content-style tokens c and prosody tokens r, flow matching learns:
L_FM = E ||v_θ(z_t,t,c,r)-u_t||²
Inference integrates the conditional velocity field and a vocoder reconstructs audio. Separation improves control, but accompaniment, room tails and timing errors in the reference can still enter the tokens.
Not a score editor
Vevo2 is suited to carrying singing style and prosody from a reference. It does not expose every MIDI onset, offset, velocity and lyric-syllable alignment as an equally strict editing contract. SoulX-Singer remains the direct MIDI route; Vevo2 is the corresponding reference-driven profile.
Zero-shot similarity and permission
Vevo2's content, prosody and timbre representations can still leak into one another, while reverb, backing vocals and accompaniment contaminate the prompt. Voice similarity describes model behaviour; it does not grant permission to use a singer's identity. References and outputs need authorization and clear synthetic disclosure.
From fixed sources to conditioned separation
Vevo2 is Amphion's zero-shot singing route. It does not read MIDI as an exact score; it separates content, timbre, style and prosody representations and recombines prompt audio with content conditions. It is adjacent to score-conditioned SoulX/YingMusic tools but not interchangeable.
Earlier conversion and generation systems often packed speaker, emotion and rhythm into one latent. Vevo uses hierarchical quantization and generation to reduce coupling. Zero-shot ability comes from data and prompt encoding; an arbitrary short reference does not guarantee stable identity, lyrics and melody.

Layered conditions for content, prosody and timbre
ŷ=G(q_content(x_c),q_prosody(x_r),e_timbre(x_r))q_contentcontent representationq_prosodyprosody representatione_timbretimbre embedding
A content tokenizer preserves lyrics and articulation, a prosody tokenizer carries rhythm and contour, and a timbre embedding supplies identity cues. The generator recombines them; disentanglement is an objective, not a guarantee.
Vevo2's prompt-driven route is not MIDI score following. Comparison with SoulX/YingMusic must first fix whether the goal is prompt-based style transfer or editable notes and exact duration.
Paper, code and weights can be followed through Amphion Vevo2 source / Vevo2 weights / Vevo2 paper.

Comparable results under one stem definition
Each row retains its source architecture, metric or task definition; results without a shared protocol remain separate.
| Model | Architecture or evidence | Scope |
|---|---|---|
| Vevo2 | content plus zero-shot style/timbre prompt | disentanglement, identity and sampling variance |
| SoulX-Singer | reference-driven singing | timbre and lyrics under the same prompt |
| YingMusic-Singer-Plus | lyrics plus melody/score | notes, duration and editability |
Stem listening and mixture consistency
Sample repeatedly from the same prompt and inspect lyric clarity, pitch contour, rhythm and timbre separately. Cherry-picking one result hides variance; accompaniment, backing vocals and reverb make disentanglement harder.
Fix the target when comparing score-driven systems: Vevo2 suits prompt-based style/timbre transfer, while MIDI/lyric models suit editable notes and exact duration. They do not share one accuracy value.
Edges of the dataset stem definition
Representation disentanglement is not identity removal. Source and target traits can leak, and output is not guaranteed anonymous.
Singing prompts require permission, and generated audio needs clear disclosure. Similarity does not authorise impersonation.
Papers and public files for Amphion Vevo2 source
Model names, numbers and limitations trace to the papers, repositories or model cards below.