AI Forensics Knowledge Base

Deep dive into the science of digital verification. Learn how SUS IT detects manipulation in audio, video, images, and documents using state-of-the-art forensic technology.

Music & Audio Forensics

Music & Audio Forensics

AI music and voice cloning models operate by predicting audio waveforms based on vast datasets. While they are becoming incredibly convincing to the human ear, they often leave behind microscopic "digital fingerprints" in the frequency spectrum that SUS IT is designed to detect.

How We Detect It:

  • High-Frequency Cutoffs: Many AI models struggle to generate consistent data above 16kHz-20kHz, creating an unnatural hard shelf or "void" in the spectrogram that rarely exists in natural recordings.
  • Phase Incoherence: Natural sound has a physical relationship between frequencies. AI generation often treats frequencies mathematically but ignores their phase alignment, leading to a subtle "metallic" or "smeary" quality known as phase vocoder artifacts.
  • Transient Smearing: The initial attack of a sound (like a snare drum or a consonant 't') is sharp and defined in real audio. In AI generation, these transients are often softer or "mushier" due to the diffusion process.
Video & Deepfake Detection

Video & Deepfake Detection

Deepfakes use neural networks to map one face onto another. SUS IT looks beyond the surface image to the physics of the video itself. We analyze frame-by-frame consistency to spot the glitches that occur when the AI loses tracking.

Key Indicators:

  • Temporal Flickering: AI often regenerates the face slightly differently in every frame, causing skin texture to "shimmer" or flicker imperceptibly (temporal instability).
  • Biological Signals (rPPG): Real faces have subtle color changes due to heartbeat (remote photoplethysmography). Deepfakes currently struggle to replicate this invisible biological signal accurately across a timeline.
  • Lip-Sync Mismatches: We compare the phonemes (sounds) in the audio track with the visemes (mouth shapes) in the video to detect misalignment that occurs in AI-dubbed content.
Image Analysis

Image Analysis

Generative Adversarial Networks (GANs) and Diffusion models create images from noise. This process leaves behind a regular, grid-like noise pattern that is distinct from the random noise found in camera sensors (ISO grain).

Visual Tells:

  • Anatomical Logic: While models are improving, AI still struggles with complex connected geometry—hands merging into objects, nonsensical jewelry, or pupils that aren't perfectly circular (non-circular iris reflection).
  • Lighting Physics: AI often approximates lighting rather than simulating it. We look for shadows that don't align with light sources or reflections in eyes that don't match the environment.
  • Upscaling Artifacts: Many AI images are generated at low resolution and upscaled. This often leaves a distinct "checkerboard" pattern visible at high magnification levels.
Document & Text Verification

Document & Text Verification

AI Large Language Models (LLMs) are statistical prediction engines. They tend to choose the most "probable" next word, resulting in text that is grammatically perfect but statistically "flat" compared to human writing.

Our Metrics:

  • Perplexity: A measure of how "surprised" a model is by the text. Low perplexity suggests the text follows standard AI patterns; high perplexity suggests human idiosyncrasy.
  • Burstiness: Humans vary their sentence structure and length dynamically (high burstiness). We write short sentences. And then we write long, complex sentences that meander through thought processes. AI tends to be more monotonous and uniform in rhythm.

Optimizing Your Analysis

How to get the most accurate forensic results and a look at what happens under the hood.

Upload Best Practices

1

Avoid MP3s & Compression

MP3 compression cuts off high frequencies (>16kHz) in a way that looks identical to early AI models. For 100% accuracy, always upload WAV or FLAC files.

2

Original Files > Screen Recordings

Screen recording a video introduces "frame blending" and re-encoding artifacts that can mask the original AI glitches or create false positives. Always upload the source file when possible.

3

Raw vs. Processed

Heavily edited media (filters, auto-tune, color grading) is harder to analyze. If you are a creator proving authenticity, upload the raw stems or raw camera footage.

Technical Deep Dive: The "Thinking" Engine

SUS IT isn't just a simple classifier. We utilize next-generation reasoning models configured with an extended forensic thinking budget.

Why "Thinking" Matters

Standard AI detectors guess instantly based on patterns. Our system engages in a Chain-of-Thought forensic process. It explicitly debates evidence: "Is this artifact a compression error or a generative flaw?" or "Is this visual glitch a rendering error or just a rolling shutter effect?" before delivering a final verdict.

  • Multi-Modal Analysis: We analyze audio spectrograms (frequency), video frame vectors (motion), and pixel grids (texture) simultaneously.
  • False Positive Filtering: We explicitly prompt the model to identify and ignore common human editing techniques like Warp Stabilizer, Noise Reduction, or Auto-Tune.

Constantly Evolving Accuracy

AI models evolve daily, and so does SUS IT. Our detection algorithms are continuously updated to recognize the latest generation techniques, from Flux-based image generators to the newest voice cloning models like ElevenLabs and Suno.

We don't just check for known artifacts; we analyze the fundamental statistical properties of media to stay ahead of the curve. Our team is constantly refining our tools to give you more accurate, reliable, and granular results with every update.