Readable model briefings with source links, model cards, repositories, and visual context.
A technical reading of learned spectral compression, local/long-range sequence modelling and the same-protocol 11.8 dB DnR result.
Latent flow matching, few-step distillation and edit preservation across three released checkpoints, with runtime and evidence boundaries.
Typography, layout and semantics are evaluated separately, with MoT routing and the BizGenEval hard result kept in scope.
Conditional diffusion, APG/CFG, 32 function evaluations and the limits of Seed speaker-similarity measurements.
CLAP conditioning, time-frequency masking and mixture consistency in one query-driven model trained across 283 sound concepts.
A same-protocol comparison of band-split modelling and TIGER's efficient sequence design on DnR cinematic separation.
A task-level analysis of permutation-invariant separation, enhancement and super-resolution across the ClearVoice and TIGER profiles.
Event CRNNs, F0 segmentation and technique-aware multi-task transcription, with dataset boundaries kept explicit.
An evidence-led account of the flagship checkpoint, event modelling, private and cross-domain evaluation, Type-1 MIDI output, and gated non-commercial license.
A technical history of the 20B MMDiT, 6B single-stream DiT and 9B rectified-flow routes across typography, distillation and multi-reference editing.
Four different objectives—matting, masked inpainting, generative SR and colour regression—kept as separate product contracts.
A technical walkthrough of AudioVAE, semantic re-encoding, the Qwen backbone and AR flow matching in SOAR and MeanFlow profiles.
Zero-shot timbre prompts, MIDI metadata alignment, melody encoders and conditional flow matching with non-interchangeable inputs.
Content encoding, pitch conditioning, speaker statistics and rectified flow in zero-shot singing voice conversion.
Multi-resolution spectra, RoFormer sequence modelling and shift-tolerant objectives, with beat timing kept separate from note transcription.
Time-domain encoding, sequence bottlenecks and waveform reconstruction, with the reported 14.22 dB result kept inside its narrow jazz test domain.
Voice design, discrete audio tokens, long-context scaling and zero-shot references compared as distinct inference contracts rather than one score.
Content-style and prosody tokenizers, flow-matching acoustics and vocoder reconstruction in a reference-driven singing path.
TelkNet's current four-stem tool uses Huge-SCNet-4stems V1.2 to split a full mix into vocals, drums, bass, and other. The article explains the choice through MSS architecture history, Mel-Band RoFormer strengths and limits, SCNet subband modeling, and public four-stem scores.
Krea AI released Krea 2 Raw and Krea 2 Turbo as open weights. The technical report describes Krea 2 as a 12B text-to-image diffusion model trained from scratch and details data filtering, captioning, post-training, reinforcement learning, and safety-license boundaries.
ChordEdit is a CVPR 2026 Oral / Best Student Paper Honorable Mention paper. It treats one-step text-guided image editing as a dynamic optimal-transport problem between source and target prompts, then uses a Chord Control Field to reduce high-energy drift that can warp objects and damage backgrounds. The official demo exposes source_prompt, target_prompt, seed, sample count, time-window, and step-scale controls.
Zyphra released ZONOS2 under Apache-2.0. The model uses an 8B-total, roughly 900M-active MoE TTS architecture trained on more than 6M hours of speech. Its official materials document benchmark framing, language tiers, and voice-cloning limits.
A source-scoped technical overview of MRT2's 40ms frame-level autoregression, low-latency MIDI/text/audio control, 2.4B/230M model split, SpectroStream codec path, MusicCoCa style embeddings, and the practical boundary between live steering and one-shot song generation.