Music & Audio Forensics
AI music and voice cloning models operate by predicting audio waveforms based on vast datasets. While they are becoming incredibly convincing to the human ear, they often leave behind microscopic "digital fingerprints" in the frequency spectrum that SUS IT is designed to detect.
How We Detect It:
- High-Frequency Cutoffs: Many AI models struggle to generate consistent data above 16kHz-20kHz, creating an unnatural hard shelf or "void" in the spectrogram that rarely exists in natural recordings.
- Phase Incoherence: Natural sound has a physical relationship between frequencies. AI generation often treats frequencies mathematically but ignores their phase alignment, leading to a subtle "metallic" or "smeary" quality known as phase vocoder artifacts.
- Transient Smearing: The initial attack of a sound (like a snare drum or a consonant 't') is sharp and defined in real audio. In AI generation, these transients are often softer or "mushier" due to the diffusion process.