Transparent testing evidence

Accuracy and Testing Methodology

A detector percentage without a dated sample set is easy to misinterpret. We publish our protocol, raw result files, and limitations so the evidence can be checked instead of taken on trust.

Current benchmark

On August 9, 2026, we ran a blind test containing 60 independent 30-second clips: 20 human performances, 20 Suno tracks, and 20 Udio tracks. The production provider classified all 60 as expected. It also retained the same verdict on 15 related 128 kbps mono variants.

This is a result for one curated batch and one provider model. It is not evidence that all future music will be classified correctly.

For the paired robustness subset, read our 15-track MP3 compression test.

What counts toward accuracy

Earlier pilot and correction

Our initial 18-file engineering pilot reported 16/18 agreement. It mixed three source files per label with three compressed copies, so it did not contain 18 independent tracks. We also later confirmed that the SoundHelix files used in its non-AI group were algorithmically composed, not a clean human-performance cohort. That pilot remains useful as an integration check, but it has been superseded by the documented 60-track benchmark and must not support a general accuracy claim.

Remaining limitations

Interpretation rule

Treat every detector result as one signal. Do not use it alone to accuse a creator, reject a submission, or make a legal or financial decision.