Can audio LLMs hear the difference yet?
Each test item is two short clips and one question about melody, rhythm, timbre or harmony. Every item comes in two variants: in FLIP the music changes and the answer switches, in STAY only the sound changes (room echo, EQ, hiss, transposition) and the answer stays. A model that ignores the audio scores 50%.
Same or different melody?
Rankings
Test split: 1,000 clean items (250 per task) and 2,000 FLIP/STAY items. All numbers are percentages; the thin tick on each bar marks 50%. — models, last updated —.
| # | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Loading results… | |||||||||||
Melody … Clean (all): accuracy on the clean items; “all” is the mean of the four tasks. AccFS: accuracy on the FLIP and STAY items, 100 − (MR + FFR) / 2; a model that ignores the audio scores 50% whatever letter it prefers. MR (miss rate): share of FLIP items answered wrongly. FFR (false-flip rate): share of STAY items answered wrongly. Lower is better for MR and FFR. A rate: share of ‘A’ answers on the clean items, so letter bias is visible. Models are ranked by AccFS within each group. Self-reported results have not yet been reproduced by the maintainers.