MossFormer2 and TIGER-speech: separating speakers is not speech restoration
A task-level analysis of permutation-invariant separation, enhancement and super-resolution across the ClearVoice and TIGER profiles.

SPEECH FRONT END
SS 16K estimates who said what. SE 48K estimates a cleaner version of one speaker. SR 48K reconstructs bandwidth. Their supervision and output cardinality are different.
Permutation-invariant training
Outputs have no intrinsic speaker order: L_PIT=min_{π∈S_K}Σ_k L(ŝ_k,s_{π(k)}). MossFormer2 SS handles general overlap. TIGER-speech is retained for noisy, echoic and reverberant two-person recordings.
Enhancement and super-resolution
Enhancement produces one cleaned waveform; super-resolution predicts missing upper-band information and can hallucinate plausible detail. Evaluation must state input and output sample rates and listen separately for denoising damage and invented high-frequency texture.
How acoustic representations evolved
Speech front ends grew in three separate communities: meeting separation targets overlap, enhancement targets noise and reverberation, and super-resolution targets missing bandwidth. ClearerVoice-Studio places MossFormer2 SS 16K, SE 48K and bandwidth restoration behind one engineering entry point, but their objectives and output contracts remain different.
TIGER-speech applies band-aware sequence modeling to reduce long-context cost. It is an efficiency comparison, not a version-number replacement for MossFormer2. Deployment starts by deciding whether the recording contains overlapping speakers, masked speech, or missing bandwidth; the wrong task produces a polished answer to the wrong question.

A permutation-invariant objective for unknown speaker order
L_PIT=min_{π∈S_K} Σ_k ℓ(ŝ_k,s_{π(k)})S_Kpermutations of K speakersπone output/reference assignmentℓsingle-signal reconstruction loss
Two-speaker separation has no natural first speaker. PIT chooses the lowest-loss output/reference assignment so channel numbering does not create contradictory supervision.
SS 16K separates speakers; SE 48K and SR 48K repair one signal. A shared leaderboard would mix task, bandwidth and datasets.
Paper, code and weights can be followed through ClearerVoice-Studio / MossFormer2 SS 16K / MossFormer2 SE 48K.

What each speech metric answers
Each row retains its source architecture, metric or task definition; results without a shared protocol remain separate.
| Model | Architecture or evidence | Scope |
|---|---|---|
| MossFormer2 SS 16K | overlapping-speaker separation | cross-talk, identity swaps, silent regions |
| MossFormer2 SE 48K | wideband denoise and dereverb | consonant damage, music removal, room tails |
| TIGER-speech | efficiency-oriented band sequence model | measure latency and quality on the same files |

Content, timbre and long-form listening
For separation, check identity swaps, sibilants, breaths and reverb tails. Enhancement can remove consonants together with keyboard or HVAC noise. Super-resolution must be judged as plausible bandwidth extension, not recovery of every original high-frequency sample.
Permutation-invariant training removes output ordering from the loss but does not guarantee identity continuity over a long meeting. Record sample rate, resampling and channel fold-down before A/B runs, and inspect speaker swaps by time segment.
Risks left by voice similarity
A mono mixture may have no unique solution when speakers have similar timbre and pitch, and source count cannot be inferred without limit.
Bandwidth extension synthesizes plausible content. For forensic, biometric or medical use, generated high frequencies are not original evidence.
Papers and public files for ClearerVoice-Studio
Model names, numbers and limitations trace to the papers, repositories or model cards below.