Sonicert/Research/Compression study
Bitrate is a false-negative minefield
2026-10-09 · Developer-authored study (we make Sonicert) · Samples tested 2026-09-30; per-file scores published on our methodology page.
The setup
During our engine pilot we ran the same labeled Suno track through a bitrate ladder — the original download plus progressively heavier compressions at 192, 128, 96 and 64 kbps — on two commercial engines: a paid third-party detection API, and our own production engine (open-weights logistic regression on spectral features, self-hosted). Both engines were given identical files, and every track's ground truth was known: all of them were AI-generated.
The result: opposite verdicts on the same files
| Sample | Paid API engine | Sonicert engine |
|---|---|---|
| Clean Suno originals (no re-encode) | Detected 6/6 ✓ | Detected 6/6 ✓ |
| zh1 re-encode @ 192 kbps | Detected ✓ | Detected ✓ |
| zh1 re-encode @ 128 kbps | Missed — verdict flipped to “human” ✗ | Detected ✓ |
| zh1 re-encode @ 96 kbps | Missed ✗ | Detected ✓ |
| 64 kbps re-encodes (zh1 + ja1) | Missed ✗ | Detected ✓ |
All five Sonicert detections held at P ≥ 0.9983 (lowest of the ladder), tested 2026-09-30. On human tracks the engines differ in what was measured: our engine ran 14 human tracks with zero false positives — including 128 kbps and 64 kbps transcodes of real recordings; the paid API's published run covered 6 human tracks (no compressed-human cases), also zero false positives. The failure mode that matters is the other direction.
Why a false negative is worse than a false positive
A false positive gets argued with. A false negative gets used: someone walks away with a confident “this is human” verdict about an AI track and cites it as evidence. Messaging apps and social platforms commonly re-encode shared audio into exactly the bitrate range where our tested engine flipped — the 128 kbps-and-below files in this study are the kind of copies that circulate.
What we do about it
- · Sonicert computes effective bitrate (file size ÷ duration) on upload; below ~160 kbps it returns a weak-signal “cannot judge” result instead of a “human” verdict.
- · Our engine held 5/5 on the full ladder in this study; we re-run the ladder against every engine update, and the published scoreboard includes the compressed variants.
- · We still label scores as probabilistic signals — compression robustness is not the same as calibrated accuracy.
Practical advice if you vet music
- · Test the original download, not a messaging-app copy, whenever provenance matters.
- · Treat any “human” verdict on a heavily compressed file as unsupported.
- · Ask your detector vendor for their compression-ladder data. If they don't have one, they don't know their own false-negative rate.
Scores are uncalibrated probabilistic signals, not legal proof of authorship. Per-file scores, source notes and SHA-256 prefixes for the pilot are published on our methodology page — anyone holding the same files can verify the scores. Try the engine free: sonicert.com — unlimited checks, no signup.