HQ-SVC and YingMusic-SVC: disentangling content, F0, energy and timbre
Content encoding, pitch conditioning, speaker statistics and rectified flow in zero-shot singing voice conversion.

SINGING VOICE CONVERSION
SVC does not mix the target reference into the source. It must preserve source lyrics, timing and melody while replacing timbre statistics. Content leakage becomes pronunciation errors or background bleed.
Pitch conditioning
Pitch can be expressed in semitones as p=69+12log₂(f₀/440). Automatic range shifting helps large register differences, but F0 errors turn slides, breath and harmony into wrong melody cues.
What the two releases actually provide
The HQ-SVC model card bundles inference code, FACodec, RMVPE and vocoder assets. YingMusic-SVC publishes an F0-aware adaptor, energy-balanced flow matching and training for real-world harmonic contamination. Without a shared test set, the public materials do not support one cross-project winner.
Two forms of input contamination
Accompaniment leakage in the source is recoloured by conversion, while reverb or background sound in the target prompt can be mistaken for identity. Both the lead-vocal source and the target reference should be clean, and converted audio needs clear synthetic disclosure.
From fixed sources to conditioned separation
Singing voice conversion retains lyrics and melody from a performance while replacing timbre-related representation. HQ-SVC emphasizes a high-quality conversion chain; YingMusic-SVC provides another content, F0 and speaker-conditioning design. Unlike synthesis, the input is an existing sung performance.
SVC evolved from spectral mapping toward content encoders, explicit F0, generative decoders and stronger speaker representations. Disentanglement remains incomplete: source accent, dynamics and formants can leak, while room sound from the target reference can be copied.

Singing conversion separates content from target timbre
ŷ=G(c(x_src),F0(x_src),E(x_src),e(x_ref))c(x_src)linguistic contentF0,Emelody and energye(x_ref)target-timbre condition
Content, F0 and energy come from the source performance; target timbre comes from the reference. Imperfect disentanglement leaves source identity or transfers room coloration.
A full mix first needs a clean lead-vocal stem. Separation leakage is recolored by conversion and can become harder to remove, so both stages must be audited together.
Paper, code and weights can be followed through HQ-SVC / YingMusic-SVC / YingMusic-SVC paper.

Comparable results under one stem definition
Each row retains its source architecture, metric or task definition; results without a shared protocol remain separate.
| Model | Architecture or evidence | Scope |
|---|---|---|
| HQ-SVC | high-quality default chain | fidelity, clarity and high-register stability |
| YingMusic-SVC | content/F0/timbre comparison | cross-singer and cross-language retention |
| source separation + SVC | complete path for mixed input | separation leakage propagates into conversion |
Stem listening and mixture consistency
Verify lyrics and rhythm first, then pitch jumps, sibilants, breathiness and high-register breaks. Automatic transposition changes vocal stress and cannot be summarized by mean similarity.
Use a dry, solo, accompaniment-free target reference with useful range. A short reverberant sample can turn room coloration into apparent identity.
Edges of the dataset stem definition
Conversion does not clear copyrights or performer rights. Public release needs authorization for both source performance and target identity.
The system does not guarantee anonymization; source traits may remain, so SVC is not a privacy filter.
Papers and public files for HQ-SVC
Model names, numbers and limitations trace to the papers, repositories or model cards below.