Demucs v3 for saxophone: 14.22 dB and narrow-domain catalogue A/B testing
Time-domain encoding, sequence bottlenecks and waveform reconstruction, with the reported 14.22 dB result kept inside its narrow jazz test domain.

INSTRUMENT SEPARATION / TECHNICAL NOTE
Saxophone is not a stable frequency band. Fundamentals, reed noise, harmonics and room tails overlap vocals, brass and distorted guitars. The checkpoint's value comes from its training distribution, not a hand-written band-pass filter.
Waveform reconstruction
For x=s_sax+s_back, an encoder reduces the temporal rate, the sequence bottleneck exchanges long-range context, and the decoder upsamples two estimates. A compact training objective is:
L = ||ŝ_sax-s_sax||₁ + ||ŝ_back-s_back||₁ + λ||ŝ_sax+ŝ_back-x||₁
The consistency term discourages missing mixture energy. High registers, long reverberation and sustained notes that coincide with a lead vocal can still leak—limitations acknowledged by the model report.
Keeping 14.22 dB in scope
The reported result is about 14.22 dB SDR on a narrow jazz-saxophone set, versus roughly 14.03 dB for the cited commercial Wind baseline. The small gap and overlapping variation support a narrow same-table comparison, not a universal-SOTA label across genres and registers.
Why separation appears in a transcription study
The work supports reconstruction of the Charlie Parker Omnibook: historical recordings are separated before FiloSax transcription. High register, brass unisons, heavy reverb and distorted guitars remain difficult, so leakage must be read alongside downstream note omissions and onset errors.
From fixed sources to conditioned separation
The saxophone specialist keeps the Demucs v3 waveform encoder, sequence bottleneck and decoder, but its gains come from a narrower source: FiloSax data and a sax-specific objective concentrate capacity on reed noise, harmonics, slides and accompaniment overlap. The checkpoint filename records 14.22 dB SDR and the model card flags the high register as a weakness.
The paper places separation in a Charlie Parker and Omnibook research chain: isolate saxophone from historical jazz recordings, then transcribe with FiloSax. Separation is an upstream step. The 14.03 dB LALAL.AI Wind result appears in the same table; the small gap supports comparison, not a sweeping victory claim.

How a waveform encoder reconstructs two stems
[ŝ_sax,ŝ_back]=D(B(E(x))), L=Σ_k ℓ(ŝ_k,s_k)Ewaveform encoderBbidirectional sequence bottleneckDtarget-stem decoder
The encoder compresses the waveform, a sequence bottleneck exchanges long context, and two decoder paths reconstruct saxophone and backing. Specialist data concentrates capacity on reed noise, slides and jazz overlap.
The 14.22 and 14.03 dB rows share one narrow paper test set and differ by only 0.19 dB. High register, brass unisons and distorted guitar still need catalogue A/B.
Paper, code and weights can be followed through Demucs v3 saxophone weights / Saxophone separation report / Demucs source.

Comparable results under one stem definition
Each row retains its source architecture, metric or task definition; results without a shared protocol remain separate.
| Model | Architecture or evidence | Scope |
|---|---|---|
| Demucs v3 + FiloSax | 14.22 dB SDR | specialist row in the paper table |
| LALAL.AI Wind | 14.03 dB SDR | commercial baseline in the same table |
| Mega53 saxophone | no matching table value | same-track listening comparison only |
Stem listening and mixture consistency
High notes, heavy reverb, brass unisons and distorted guitar are common leakage cases. Beyond SDR, inspect softened attacks, sax ghosts in backing and mixture consistency when stems are summed.
Mega53 saxophone can provide a same-track listening comparison but has no row in the paper table. Model choice should be revisited for each catalogue and instrument register.
Edges of the dataset stem definition
The test material is jazz-saxophone-heavy and does not represent every genre, effect chain or saxophone register. Reported deviation also indicates substantial track variance.
Specialist separation changes the waveform and is not forensic recovery. Retain originals and processing records for remix or dataset work.
Papers and public files for Demucs v3 saxophone weights
Model names, numbers and limitations trace to the papers, repositories or model cards below.