VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models
arXiv - 2026
Voie par défaut pour produire un MIDI vocal structuré, proche d’une partition, à partir d’un chant plus propre ; le checkpoint fixe n’expose aucun paramètre officiel d’inférence.