BandIt and TIGER-DnR: where two three-stem separators spend their compute
A same-protocol comparison of band-split modelling and TIGER's efficient sequence design on DnR cinematic separation.

CINEMATIC AUDIO / DnR
Cinematic separation must preserve three broad source classes while they coexist for long passages. BandIt spends capacity on non-uniform spectral bands; TIGER pursues a more efficient sequence model.
Scale-invariant evaluation
Project the estimate onto the reference, s_target=<ŝ,s>s/||s||², then compute SI-SNR=10log10(||s_target||²/||ŝ-s_target||²). Gain is removed, but transient smearing and spatial damage still require listening tests.
Why BandIt remains available
BandIt models information-dense regions using non-uniform bands and communicates across them. The same-protocol comparison reports 10.9 dB, below the released SFC-Locoformer medium checkpoint at 11.8 dB. BandIt is therefore an architectural comparison rather than the default.
How TIGER changes the cost profile
TIGER-DnR is reported around 9.8 dB in the corresponding DnR context and emphasizes efficient time–frequency interaction. That number is useful only with its named protocol; speed, licence and separation quality remain separate axes.
From fixed sources to conditioned separation
DnR established dialogue, music and effects as a shared cinematic-separation target. BandIt models non-uniform frequency bands, while TIGER reduces long-sequence cost; they answer different engineering questions rather than replacing one another.
SFC-Locoformer later reports 11.8 dB in the same DnR v2 table where BandIt reports 10.9 dB. TIGER's public evidence emphasizes efficiency and separate experiments. Without a matching row it remains a selectable model, not a fabricated leaderboard position.

The compute ledger of band splitting and sequence modeling
h_b=Enc_b(X[F_b,:]), Ŝ=Dec({h_b}_{b=1}^B)F_bfrequency bins in band bEnc_bwithin-band encodingDeccross-band reconstruction
BandIt organizes the complex spectrum into non-uniform bands before exchanging information within and across bands. TIGER emphasizes a cheaper sequence representation. Both emit DnR stems while spending compute differently.
Numbers belong together only when split, sample rate, stem definitions and metric implementation match. Leaving TIGER blank without a matching protocol is more rigorous than borrowing another table.
Paper, code and weights can be followed through BandIt / TIGER / DnR comparison.

Comparable results under one stem definition
Each row retains its source architecture, metric or task definition; results without a shared protocol remain separate.
| Model | Architecture or evidence | Scope |
|---|---|---|
| SFC-Locoformer medium | 11.8 dB SI-SDR | current quality reference in the same DnR v2 table |
| BandIt | 10.9 dB SI-SDR | same-protocol band-split baseline |
| TIGER-DnR | efficiency evidence shown separately | no rank without a matching table row |

Stem listening and mixture consistency
On dialogue, audit sibilants, breaths and distant lines; on music, speech ghosts and damaged percussion; on effects, continuity of footsteps, doors and ambience. Mean SI-SDR does not perform those listening checks.
Delivery masters contain ducking, limiting, channel fold-down and loudness processing. A paper's mono protocol is a reproducible baseline, not every master. File and model metadata make track-level A/B possible.
Edges of the dataset stem definition
Effects remains a broad container rather than footsteps, rain, traffic or applause. Three stems are not object-level sound understanding.
Published scores depend on dataset version, sample rate and computation. Every leaderboard row retains metric direction, protocol and source.
Papers and public files for BandIt
Model names, numbers and limitations trace to the papers, repositories or model cards below.