Dated benchmark · August 9, 2026
AI Music Detector Benchmark: 60-Track Blind Test
We tested the detector used by aimusicdetect.com on 60 independent, anonymously named clips: 20 human performances, 20 Suno generations, and 20 Udio generations. It classified all 60 as expected in this specific test set. That is a batch result, not a claim of universal 100% accuracy.
60 / 60
Expected classifications on independent tracks
0 / 20
Human tracks falsely classified as AI
0 / 15
Verdict flips after 128 kbps mono compression
Results by source
| Known source | Tracks | Correct | Observed result |
|---|---|---|---|
| Human performance | 20 | 20 | 20 Human, 0 AI |
| Suno | 20 | 20 | 20 AI, 0 Human |
| Udio | 20 | 20 | 20 AI, 0 Human |
Overall agreement was 60/60. The 95% Wilson lower confidence bound is 94.0%, which is one reason a perfect observed batch must not be read as proof that future accuracy is 100%. Median provider latency was 7.3 seconds and P95 latency was 16.5 seconds in the complete primary run.
How the blind test worked
- We selected 20 independently sourced tracks for each of the three known labels.
- Every track was clipped to 30 seconds, normalized to MP3 at 192 kbps stereo and stripped of metadata.
- Files were renamed to anonymous IDs such as B001.mp3 before being submitted to ACRCloud model h2zt2l57.
- Ground-truth labels stayed in a local manifest and were compared only after the provider returned its verdict.
- Five tracks per class were also encoded at 128 kbps mono. All 15 retained the same verdict as their parent clip.
Sample provenance
- Human performances came from the Open Music Archive. These are mostly older recordings distributed as public domain or under the archive's stated open license.
- Udio tracks came from the AIME dataset, using its published model labels and dataset revision.
- Suno tracks were sampled from a public research collection linking to Suno-hosted files. They were analyzed locally and are not redistributed by this site.
Download the underlying results
The public files contain anonymous sample IDs, SHA-256 hashes, known labels, returned verdicts, probabilities, model version, latency, and run identity. They do not contain audio, credentials, private paths, or source download URLs.
Focused analysis
See the paired MP3 compression robustness test for every original and 128 kbps mono probability comparison. The 20-track Suno analysispublishes the complete Suno subset, probability range, per-track latency, and practical interpretation. A separate Udio page will only be published when it adds material analysis beyond this benchmark.
Want to judge a track before relying on a score? Follow the five checks in how to tell if a song is AI generated.
Important limitations
- This was an internal engineering benchmark, not an independent audit.
- The older human recordings do not represent every modern genre, language, production chain, remix, or mastering style.
- The AI cohort covers sampled Suno and Udio outputs, not every generator or model version.
- The detector provider and generation models can change after the test date.
- Compression testing covered one 128 kbps mono condition; it does not cover every edit or codec.
Never use one detector score alone to accuse a creator, reject a submission, or make a legal or financial decision. Check provenance, project files, disclosures, and listening evidence as well.