MusicListenBenchAMAAI Lab

Submit a result

The leaderboard is a single CSV file on GitHub. You score your model with the official script, add one row with a pull request, and this page and the Hugging Face Space update when it merges.

  1. Run the official evaluation

    Answer the 3,000 test items of the v1.0 test split (1,000 clean, 1,000 FLIP and 1,000 STAY items) and write one JSON Lines file: a first line with the model name and version, how the answer was read (letter probabilities or generated text), whether the model used the training split, the date and an optional link to code, then one record per item with the item ID and the letter ‘A’ or ‘B’. Missing or invalid answers count as wrong. Training on test items is not allowed. The format is described at the top of musiclistenbench/scoring/score_submission.py in the GitHub repository, and data/example_submission.jsonl is an example. Score the file with the official script, run from the repository root:

    python -m musiclistenbench.scoring.score_submission submission.jsonl
  2. Add one row per model and method

    Append to leaderboard.csv. Scores are percent accuracy with one decimal. Set verified to no; maintainers change it after reproducing your numbers.

    type,model,organization,model_url,access,params_b,method,melody,rhythm,timbre,harmony,overall,acc_fs,miss_rate,false_flip_rate,a_rate,readout,verified,submitted_by,date,source_url,benchmark_version,notes
    model,My-Audio-LLM,My Lab,https://huggingface.co/...,open,7,zero-shot,60.0,55.2,66.8,52.0,58.5,57.6,40.1,44.7,52.0,logprob,no,@your-github,2026-10-01,https://arxiv.org/...,1.0,

    The first four score columns are clean accuracy per task, overall is clean accuracy on all 1,000 clean items, and acc_fs, miss_rate, false_flip_rate and a_rate are FLIP/STAY accuracy, miss rate, false-flip rate and the share of ‘A’ answers on the clean items. method is zero-shot, or how you used the training split (for example GRPO (LoRA)); readout is logprob or generated.

  3. Open a pull request

    A check validates the file automatically and explains any problem. Fix it and push again; the check reruns.

  4. Maintainers review and merge

    TODO: what we check before merging, and how long review usually takes.