Detecting AI-Generated Images and Audio: Practical Techniques for Researchers
A forensic guide to identifying diffusion model artifacts, corneal reflection geometry, acoustic vocal falloffs, and C2PA Content Credentials in the era of generative AI.
The democratization of generative artificial intelligence—from text-to-image latent diffusion models (Midjourney, Flux, Stable Diffusion) to zero-shot neural voice cloners (ElevenLabs, Tortoise)—has fundamentally transformed digital evidence analysis. Where investigators previously looked for crude digital clone stamping or edge blending, we now encounter photorealistic synthetic media capable of fabricating entire geopolitical incidents out of thin air.
Yet, despite widespread anxiety about “undetectable deepfakes,” synthetic generation algorithms operate under mathematical and architectural constraints. They do not simulate the real physical world; they assemble probabilistic arrangements of pixels and acoustic waveforms that consistently leave characteristic forensic signatures.
As part of the investigative methodology established in our OSINT Verification Handbook, this guide provides an actionable, technical diagnostic protocol for identifying AI-generated imagery and synthetic voice audio.
1. The Fallacy of Automated “AI Detector” Scores
Before examining forensic markers, researchers must understand an operational truth: Automated “AI Detection” web tools are inherently unreliable as sole evidentiary proof.
Commercial web detectors that return confidence percentages (e.g., “89.4% Probable AI Generated”) operate by evaluating high-frequency noise distributions and statistical pixel gradients. However: * When an authentic digital photo is uploaded to WhatsApp or X, social media compression introduces quantization artifacts that automated detectors routinely misclassify as AI generation (a false positive). * When a generative AI image is intentionally re-compressed, downsampled, or overlaid with subtle film grain, automated detectors regularly fail to identify it (a false negative).
Investigative Standard: An automated detection score is merely an indicator. Definite attribution requires identifying reproducible physical, anatomical, acoustic, and cryptographic anomalies.
2. Anatomical and Optical Forensics in Generative Images
Modern diffusion models excel at overall scene composition, but consistently stumble over micro-scale physical optics and human biological anatomy.
CORNEAL REFLECTION FORENSIC TEST
Authentic Photograph Synthetic AI Generation
┌─────────┐ ┌─────────┐ ┌─────────┐ ┌─────────┐
│ ☀️ │ │ ☀️ │ │ ☀️ │ │ 🪟 │
│ ( • ) │ │ ( • ) │ │ ( • ) │ │ ( • ) │
└─────────┘ └─────────┘ └─────────┘ └─────────┘
Left Eye Right Eye Left Eye Right Eye
Identical Light Source Vector Conflicting Illumination Physics
(Valid Physical Optics) (Probabilistic Neural Artifact)
1. Corneal Specular Reflections (The Gold Standard)
In a genuine photograph captured in the physical world, the subject’s eyes act as convex mirrors reflecting the ambient environment. Both pupils must reflect identical light sources, shapes, and angles. * The Anomaly: Latent diffusion models generate each eye as an independent probability cluster. Zoom in closely to the pupil: if the left cornea reflects a single circular outdoor sun while the right cornea reflects an indoor rectangular softbox window, the image is mathematically proven to be synthetic.
2. Pupil Boundary Circularity
- Human pupils dilate into near-perfect circles or smooth ovals depending on the angle of incidence.
- Diffusion models frequently produce irregular, jagged, or undulating pupil borders, with the black pupil bleeding into the colored iris striations.
3. Dental Micro-Anatomy
- Count and inspect the incisors, canines, and premolars.
- AI models frequently generate “supernumerary teeth”—individuals possessing six front incisors instead of four, or teeth that fuse together without distinct gum roots or interdental papillae.
4. Ear Cartilage and Jewelry Continuity
- The human ear possesses intricate, highly individual cartilage folds (helix, antihelix, tragus). AI generators routinely blur or simplify internal ear canals into amorphous flesh mounds.
- Examine earrings and necklaces: AI generation often renders an earring attached to one earlobe while failing to render the corresponding earring on the opposite ear, or shows a gold chain fading directly into skin tissue without a clasp.
3. Background Semantic Physics: The Architecture of Diffusion
While human evaluators focus instinctively on the foreground subject (e.g., a prominent politician or soldier), the background exposes generative synthesis:
- Text and Typography Hallucinations: Diffusion models do not understand written language; they model visual text as decorative textures. Background street signs, vehicle license plates, store banners, and book spines in AI images consistently display nonsensical pseudo-Latin or distorted Cyrillic glyphs.
- Structural Vanishing Points: Architectural lines (window frames, railroad tracks, sidewalk edges) in real physical cameras converge toward coherent perspective vanishing points. Generative models frequently produce parallel lines that bend, merge, or terminate abruptly in mid-air.
- Hand and Fingernail Anomalies: While state-of-the-art models (Flux, Midjourney v6) have improved hand rendering, inspect fingernail beds and skin knuckle wrinkles: generative systems frequently generate fingernails facing contradictory directions or fingers of unnatural proportional lengths.
4. Synthetic Audio & Voice Clone Forensics
Deepfake audio poses an acute risk in elections and breaking news, where cloned phone recordings or leaked voice memos can manipulate financial markets or civil stability.
When auditing a suspected voice recording:
1. Spectrogram Frequency Ceilings
Open the audio in Audacity and toggle the Spectrogram view: * Most commercial neural text-to-speech (TTS) engines are optimized for computational speed and bandwidth, operating at sample rates of 22.05 kHz or 24 kHz. * Consequently, synthetic audio tracks frequently exhibit a hard brick-wall cutoff at 11 kHz or 12 kHz, with zero acoustic energy existing in the upper frequencies. Authentic human voice recordings captured in physical acoustic spaces show natural harmonic overtones extending gracefully beyond 16 kHz.
2. Biological Respiration Inconsistencies
- A human speaker must breathe. Real conversational speech features organic inhalation pauses, subtle chest movements audible in the microphone, and tongue micro-clicks preceding plosive consonants (P, T, K, B).
- AI voice synthesis often delivers sentences spanning 25 to 30 words without a single inhalation pause, or inserts breathing audio samples at grammatically illogical moments where a human would never inhale.
3. Room Acoustic Impulse Response
- In authentic recordings, the acoustic reverberation corresponds directly to the space (e.g., the echo of a high-ceiling hall vs. the deadened acoustic of an automobile).
- Voice clones layered over fabricated background noise frequently exhibit “dry”, studio-close microphone proximity that acoustic filters fail to match with the simulated room acoustics.
5. Cryptographic Provenance: C2PA Content Credentials
The long-term defense against synthetic media is not reactive visual detection, but cryptographic provenance at the hardware capture point.
The C2PA (Coalition for Content Provenance and Authenticity) standard embeds tamper-evident cryptographic manifests into digital media files: * When a photograph is taken on a supported device (Leica M11-P, Sony Alpha, Canon C2PA firmware), a cryptographic hardware chip signs the file with the camera’s digital certificate. * Any subsequent edits in Adobe Photoshop or generative fill additions are recorded as cryptographically verified layers in the manifest.
# Verify C2PA provenance manifests using the open-source CLI
c2patool verify_evidence.jpg
If the cryptographic signature is intact, you can verify the camera serial number, creation timestamp, and edit history with mathematical certainty. If the manifest has been stripped, you must rely on the manual anatomical and reverse-search techniques detailed in our Reverse Image Search Guide.
Synthetic Media Diagnostic Checklist
| Verification Check | Target Asset | Forensic Anomaly Indicator |
|---|---|---|
| Corneal Reflection | Photographic Portrait | Conflicting light sources in left vs right eye. |
| Pupil Geometry | Photographic Portrait | Non-circular, jagged, or undulating pupil borders. |
| Background Typography | Scene / Architecture | Nonsensical glyphs or unreadable pseudo-text. |
| Acoustic Spectrogram | Audio Recording | Sharp artificial frequency cutoff below 16 kHz. |
| Respiration Patterns | Voice Memo | Continuous speech lacking organic inhalation pauses. |
| C2PA Manifest | File Container | Missing or broken cryptographic signatures. |
Conclusion: The Investigative Mindset
As generative AI continues to evolve toward higher visual fidelity, single visual markers will gradually diminish. The investigator’s ultimate defense is not looking for a single magic flaw, but applying the comprehensive multi-layered methodology of the Verification Navigator: testing the file container, cross-examining reverse database indices, calculating solar chronolocation, and corroborating with ground eyewitness reality.
Run private client-side Error Level Analysis (ELA) and inspect local metadata with the Digital Media Verification Navigator.
Put This Methodology Into Practice
Test these forensic workflows directly inside our client-side verification engine. Inspect EXIF headers in memory, calculate cryptographic file fingerprints, and run automated error level analysis with zero data leaving your device.
Launch Digital Verification Navigator →About the Contributor
The Dawat Forensic Research Desk analyzes synthetic media generation, algorithmic influence operations, and machine learning provenance standards.