MusicListenBenchAMAAI Lab

Can audio LLMs hear the difference yet?

Each test item is two short clips and one question about melody, rhythm, timbre or harmony. Every item comes in two variants: in FLIP the music changes and the answer switches, in STAY only the sound changes (room echo, EQ, hiss, transposition) and the answer stays. A model that ignores the audio scores 50%.

Same or different melody?

Rankings

Test split: 1,000 clean items (250 per task) and 2,000 FLIP/STAY items. All numbers are percentages; the thin tick on each bar marks 50%. — models, last updated —.

#
Loading results…

Melody … Clean (all): accuracy on the clean items; “all” is the mean of the four tasks. AccFS: accuracy on the FLIP and STAY items, 100 − (MR + FFR) / 2; a model that ignores the audio scores 50% whatever letter it prefers. MR (miss rate): share of FLIP items answered wrongly. FFR (false-flip rate): share of STAY items answered wrongly. Lower is better for MR and FFR. A rate: share of ‘A’ answers on the clean items, so letter bias is visible. Models are ranked by AccFS within each group. Self-reported results have not yet been reproduced by the maintainers.