Supports inputs in formats such as mp3 / wav / flac / ogg / m4a / aac / aif / aiff / mp4 / mov / mkv
Supported formats
mp3wavflacoggm4aaacaifaiffmp4movmkv
Output
MIDI
Use case
Edit notes, instrumentation and performance details.
Choose by input and output needs, then comparable tests, published evaluations and capabilities. Missing scores do not imply lower quality; more parameters or a newer release do not prove superiority. This is a suggested usage order, not an overall leaderboard.
Cost, time, formats and parameters depend on the selected model.
Models in this field
18 entries
Published results include models on and off this site. The overall ranking opens first; switch the benchmark or category to see its ranking. Each benchmark keeps its own evaluation conditions.
Overall rankingHigher is better · best first
Uses the published primary metric for this test; scores are not pooled across benchmarks.
SwiftF0 0.2.0 uses the author's latest pitch-benchmark: 10 corpora, 8 scored conditions and 50 cents tolerance. This is pitch F1; the difference from RMVPE is statistically undetermined. Asymmetric 95% CIs are retained. REAPER is unranked due to incomplete results; the older 0.1.x paper does not score 0.2.0.
Markers identify models or components used here. Scores and settings come from the cited tests, not a new evaluation of the complete site pipeline.
Scores apply only to the cited dataset, metric and model version. Leading this table does not mean a current world record or a result from this service. Capabilities and tests can justify prioritizing models without comparable scores.
Benchmark chart: Model comparison
Scores come from the cited public sources; blue rows are models available in this tool.
Paper results retain their datasets. FiloBass Table 5 evaluates other algorithms on that dataset; the current FiloBass checkpoint has no published score.
Ranks are computed within each table and evaluation context, retaining ties. Change the metric to reorder; scroll horizontally for every column. Missing values are unranked, never zero.
Paper results retain their datasets. FiloBass Table 5 evaluates other algorithms on that dataset; the current FiloBass checkpoint has no published score.
Ranks are computed within each table and evaluation context, retaining ties. Change the metric to reorder; scroll horizontally for every column. Missing values are unranked, never zero.
Paper results retain their datasets. FiloBass Table 5 evaluates other algorithms on that dataset; the current FiloBass checkpoint has no published score.
Ranks are computed within each table and evaluation context, retaining ties. Change the metric to reorder; scroll horizontally for every column. Missing values are unranked, never zero.
Paper results retain their datasets. FiloBass Table 5 evaluates other algorithms on that dataset; the current FiloBass checkpoint has no published score.
Ranks are computed within each table and evaluation context, retaining ties. Change the metric to reorder; scroll horizontally for every column. Missing values are unranked, never zero.
Paper results retain their datasets. FiloBass Table 5 evaluates other algorithms on that dataset; the current FiloBass checkpoint has no published score.
Ranks are computed within each table and evaluation context, retaining ties. Change the metric to reorder; scroll horizontally for every column. Missing values are unranked, never zero.
Paper results retain their datasets. FiloBass Table 5 evaluates other algorithms on that dataset; the current FiloBass checkpoint has no published score.
Ranks are computed within each table and evaluation context, retaining ties. Change the metric to reorder; scroll horizontally for every column. Missing values are unranked, never zero.
SwiftF0 0.2.0 uses the author's latest pitch-benchmark: 10 corpora, 8 scored conditions and 50 cents tolerance. This is pitch F1; the difference from RMVPE is statistically undetermined. Asymmetric 95% CIs are retained. REAPER is unranked due to incomplete results; the older 0.1.x paper does not score 0.2.0.
Ranks are computed within each table and evaluation context, retaining ties. Change the metric to reorder; scroll horizontally for every column. Missing values are unranked, never zero.
Tsumugi Other and Vocal Harmony report COnP, not COnF1. Published results cover v1.5; default Vocal Harmony v1.6 has no published score. Instrument refinement remains experimental, with datasets shown separately.
Ranks are computed within each table and evaluation context, retaining ties. Change the metric to reorder; scroll horizontally for every column. Missing values are unranked, never zero.
Tsumugi Other and Vocal Harmony report COnP, not COnF1. Published results cover v1.5; default Vocal Harmony v1.6 has no published score. Instrument refinement remains experimental, with datasets shown separately.
Ranks are computed within each table and evaluation context, retaining ties. Change the metric to reorder; scroll horizontally for every column. Missing values are unranked, never zero.
Tsumugi Other and Vocal Harmony report COnP, not COnF1. Published results cover v1.5; default Vocal Harmony v1.6 has no published score. Instrument refinement remains experimental, with datasets shown separately.
Ranks are computed within each table and evaluation context, retaining ties. Change the metric to reorder; scroll horizontally for every column. Missing values are unranked, never zero.
Tsumugi Other and Vocal Harmony report COnP, not COnF1. Published results cover v1.5; default Vocal Harmony v1.6 has no published score. Instrument refinement remains experimental, with datasets shown separately.
Ranks are computed within each table and evaluation context, retaining ties. Change the metric to reorder; scroll horizontally for every column. Missing values are unranked, never zero.
Public sources checked: 2026-09-23
Showing 2/2 rows · 1 metrics
Tsumugi · instrument refinement
Rank
Test setting
Model
ACC (%)
#1
Held-out RWC-I
Instrument refinement model
74.50
#2
Held-out RWC-I
Base instrument classifier
71.30
Models, parameters, and sources
SwiftF0 0.2.0 uses its official pitch tracking and note segmentation for monophonic recordings. Outputs MIDI, frame-level pitch CSV and a plot.
SwiftF0 0.2.0 uses the author's latest pitch-benchmark: 10 corpora, 8 scored conditions and 50 cents tolerance. This is pitch F1; the difference from RMVPE is statistically undetermined. Asymmetric 95% CIs are retained. REAPER is unranked due to incomplete results; the older 0.1.x paper does not score 0.2.0.
Basic Pitch
Spotify Basic Pitch transcribes polyphonic performances from nearly any instrument, including voice, into MIDI and optional note-event CSV with pitch bends. It works best on one instrument at a time.
Shortest note retained by Basic Pitch, in milliseconds.
Discard Basic Pitch notes below this frequency; leave blank for no limit.
Discard Basic Pitch notes above this frequency; leave blank for no limit.
Keep separate pitch-bend curves for simultaneous Basic Pitch notes.
Enable Basic Pitch’s official Melodia-inspired post-processing.
After inference, Beat This final0 writes a valid fixed tempo or beat-by-beat tempo map into the MIDI while preserving every event's position in seconds. If valid beat evidence cannot be produced, the task fails instead of delivering MIDI with the model's default tempo.
GAPS
GAPS converts clean or separated solo guitar into note events and MIDI, without string positions, frets, continuous F0 or Pitch Bend. Adjust MIDI velocity by comparing with the original recording.
Paper results retain their datasets. FiloBass Table 5 evaluates other algorithms on that dataset; the current FiloBass checkpoint has no published score.
Guitar-FL
Guitar-FL converts guitar solos from the François Leduc data domain into discrete note events and MIDI; FL does not mean frame-level. It does not output continuous F0 or Pitch Bend.
Paper results retain their datasets. FiloBass Table 5 evaluates other algorithms on that dataset; the current FiloBass checkpoint has no published score.
FiloSax
FiloSax targets monophonic jazz saxophone solos and exports discrete pitch, onset, and offset events as MIDI. It does not output continuous pitch, slide, or vibrato tracks.
Paper results retain their datasets. FiloBass Table 5 evaluates other algorithms on that dataset; the current FiloBass checkpoint has no published score.
FiloBass
FiloBass targets monophonic jazz double bass (upright bass), not general or electric bass. Use a clean or separated stem; kick and low-frequency bleed can create false notes.
Paper results retain their datasets. FiloBass Table 5 evaluates other algorithms on that dataset; the current FiloBass checkpoint has no published score.
VioPTT
VioPTT outputs violin note events and one of four technique labels—détaché, flageolet, spiccato, pizzicato—or none. MIDI velocity is fixed, not measured performance dynamics.
Separate first: split vocals/accompaniment or stems before transcribing the target audio.
Select the released checkpoint that matches the source stem; this is routing, not automatic instrument recognition.
STrADi · Offline
STRAdi Offline is the matched non-causal, full-context reference route. It exports the same discrete MIDI and onset, offset and pitch CSV; it is not a published score-following system either.
Separate first: split vocals/accompaniment or stems before transcribing the target audio.
Select the released checkpoint that matches the source stem; this is routing, not automatic instrument recognition.
STrADi · Online
STRAdi Online targets clean solo violin with fully causal inference and no future context. It exports discrete MIDI plus onset, offset and pitch CSV; the public repository does not ship a score-following system.
GAPS: A Large and Diverse Classical Guitar Dataset and Benchmark Transcription Model
ISMIR - 2024
GAPS converts clean or separated solo guitar into note events and MIDI, without string positions, frets, continuous F0 or Pitch Bend. Adjust MIDI velocity by comparing with the original recording.
High Resolution Guitar Transcription via Domain Adaptation
ICASSP - 2024
Guitar-FL converts guitar solos from the François Leduc data domain into discrete note events and MIDI; FL does not mean frame-level. It does not output continuous F0 or Pitch Bend.
Reconstructing the Charlie Parker Omnibook using an audio-to-score automatic transcription pipeline
SMC - 2024
FiloSax targets monophonic jazz saxophone solos and exports discrete pitch, onset, and offset events as MIDI. It does not output continuous pitch, slide, or vibrato tracks.
FiloBass: A Dataset and Corpus Based Study of Jazz Basslines
ISMIR - 2023
FiloBass targets monophonic jazz double bass (upright bass), not general or electric bass. Use a clean or separated stem; kick and low-frequency bleed can create false notes.
VioPTT: Violin Playing Technique-Aware Transcription from Synthetic Data Augmentation
ICASSP - 2026
VioPTT outputs violin note events and one of four technique labels—détaché, flageolet, spiccato, pizzicato—or none. MIDI velocity is fixed, not measured performance dynamics.
A Lightweight Instrument-Agnostic Model for Polyphonic Note Transcription and Multipitch Estimation
ICASSP - 2022
Spotify Basic Pitch transcribes polyphonic performances from nearly any instrument, including voice, into MIDI and optional note-event CSV with pitch bends. It works best on one instrument at a time.
Instrument-Agnostic AMT: released checkpoints and inference guide
Official technical guide - 2026
Convert isolated instrument, vocal or drum tracks to MIDI using the matching official model. Vocal harmony also supports the v1.6 checkpoint. Instrument label refinement remains experimental.
SwiftF0 0.2.0 uses the author's latest pitch-benchmark: 10 corpora, 8 scored conditions and 50 cents tolerance. This is pitch F1; the difference from RMVPE is statistically undetermined. Asymmetric 95% CIs are retained. REAPER is unranked due to incomplete results; the older 0.1.x paper does not score 0.2.0.
Workflow
1Upload a single-instrument recording and select its specialist model.
2Import the MIDI into a DAW to review pitches, onsets and note lengths.
Before you submit
Select the released checkpoint that matches the source stem; this is routing, not automatic instrument recognition.
If other instruments are mixed in, separate them first before transcribing
Quick feedback
Include the tool name, steps and the problem in your feedback. Screenshots can help us locate it.