SoulX-Singer and YingMusic-Singer-Plus: score control versus melody reference
Zero-shot timbre prompts, MIDI metadata alignment, melody encoders and conditional flow matching with non-interchangeable inputs.

SINGING SYNTHESIS
SoulX score mode consumes MIDI-derived metadata. YingMusic-Singer-Plus consumes a singing clip that supplies melody plus target lyrics. One needs note–lyric alignment; the other encodes melody from audio.
Conditional generation
The generative path follows dx_t/dt = v_θ(x_t,t | l,m,r), with the velocity field conditioned by lyrics, melody and timbre reference. Misalignment bends that integrated path toward a poor compromise between pronunciation, rhythm and identity.
Alignment
SoulX converts MIDI into segmented metadata; syllables and note onsets must agree. YingMusic avoids explicit MIDI but needs a clean melody vocal, because accompaniment leakage contaminates its melody representation.
The published interfaces differ
SoulX-Singer centres on MIDI, lyrics and a timbre prompt. YingMusic-Singer-Plus additionally emphasizes flexible lyric manipulation and melody guidance without note-by-note annotation. Comparison must separate pitch, duration, articulation and reference timbre rather than reducing both systems to one singing switch.
From fixed sources to conditioned separation
The input contract determines what a singing system can do. SoulX-Singer emphasizes zero-shot timbre and singing generation; YingMusic-Singer-Plus combines lyrics with melody or score conditions. Both emit singing, but one depends more on reference audio and the other on structured music.
Singing synthesis progressed from concatenation and acoustic models through neural vocoders, codec tokens and large generators. The persistent problem is alignment among phonemes, note duration, vibrato, breaths and timbre. One fluent chorus can hide wrong lyrics and drift across bars.

Score conditions and melody references are not interchangeable
ŝ=G(lyrics, pitch, duration, e_timbre, c_style)pitchpitch or melody contourdurationnote duratione_timbretimbre from an authorized prompt
SoulX-Singer uses lyrics, MIDI/metadata and a timbre prompt; YingMusic-Singer-Plus accepts more flexible melody guidance. Pitch and duration are editable score conditions, while a melody reference carries additional performance information.
Pitch accuracy, duration, lyric clarity, timbre similarity and breathing are separate checks. One selected high-note sample cannot establish song-length stability.
Paper, code and weights can be followed through SoulX-Singer / SoulX-Singer paper / YingMusic-Singer-Plus.
Comparable results under one stem definition
Each row retains its source architecture, metric or task definition; results without a shared protocol remain separate.
| Model | Architecture or evidence | Scope |
|---|---|---|
| SoulX-Singer | reference-driven zero-shot singing | identity, lyrics and long-form stability |
| YingMusic-Singer-Plus | lyrics plus melody/score | alignment, range and tempo |
| traditional singing pipeline | explicit acoustic features plus vocoder | clear control but limited cross-domain naturalness |

Stem listening and mixture consistency
Compare every word for omissions, swallowed consonants, repetition and consonant timing, then inspect pitch entry, slides and vibrato. Naturalness does not replace content accuracy.
Hold key, BPM, meter and reference crop constant. Accompaniment and reverb in the reference can be copied into the output.
Edges of the dataset stem definition
Singer cloning requires explicit permission and must not support impersonation, deception or rights evasion.
The model does not replace arrangement or vocal direction. Range, diction, emotion and language habits still need human work.
Papers and public files for SoulX-Singer
Model names, numbers and limitations trace to the papers, repositories or model cards below.