Back to Blog

Deepfakes have moved out of research labs and political disinformation and into everyday fraud. A finance employee in Hong Kong was convinced by a realistic AI video conference call — featuring a convincing simulacrum of his company's CFO — to wire $25 million to criminal accounts. Voice clone technology now requires as little as three seconds of audio to replicate someone's voice convincingly enough to fool people who know them well. AI-generated profile photos are used in romance scams, fake job listings, and corporate impersonation. Knowing what to look for is not optional anymore. It is a basic literacy skill for anyone who interacts with digital media.

This article breaks down the specific tells — visual, audio, and behavioral — that AI-generated content leaves behind, and gives you the tools to verify what you're seeing before you act on it.

$25M Lost in a single deepfake video-call fraud, Hong Kong Reuters, Feb 2024 ↗
10× Increase in deepfakes detected globally, 2022 to 2023 Sumsub Identity Fraud Report ↗
3 sec Of audio needed to clone a human voice with AI tools McAfee Artificial Imposters, 2023 ↗
2022 Year the FBI issued its first formal deepfake employment fraud warning FBI IC3 PSA, Jun 2022 ↗

What Is a Deepfake — and Why It Matters Now

A deepfake is any media — image, video, audio, or text — that has been synthetically generated or manipulated using AI to make it appear authentic. The term originally referred specifically to face-swap video technology, but it now covers a spectrum: AI-generated photographs of people who don't exist, video footage with one person's face replaced by another, cloned voices, and hybrid content that combines real and fabricated elements.

The technology has become dramatically easier to use. Several years ago, producing a convincing deepfake video required significant technical skill and compute resources. Today, detection firms such as Sensity AI track a landscape that independent surveys now count in the thousands of tools, capable of generating photorealistic faces, lip-synced video and voice clones without any specialised knowledge. The barrier is gone. The output quality is not.

Why Detection Tools Are Not Enough on Their Own

Automated detectors are useful, and they are not definitive. The GenImage benchmark, published at NeurIPS in 2023, found detectors score above 98.5% on the generator they were trained on — and drop to averages of roughly 60 to 70% on generators they have never seen, collapsing to near coin-flip on the least familiar ones. Compression makes it worse: the same detector that scored well on a full-resolution file can fall to chance on a JPEG, which is what every image you actually see has already been through.

The arms race is real and documented. In 2020, Nicholas Carlini and Hany Farid showed that a detector achieving an AUC of 0.95 could be driven to near-zero accuracy by imperceptible changes — flipping the lowest bit of each pixel was enough. Even an attacker with no access to the detector at all got it down to 0.22, worse than guessing.

What follows is not a checklist that will make you a detector. It is a description of where the remaining signal is, and where it has already gone.

🖼️
AI-Generated Images — What Looking Still Tells You
Applies to: profile photos, news images, social media content

Start with the uncomfortable part. A pre-registered study of 1,276 people published in Communications of the ACM measured how well ordinary people separate real media from synthetic. Mean accuracy across everything was 51.2%. On images specifically it was 49.4% — very slightly worse than a coin. People were right 64.6% of the time about genuinely authentic content and only 38.8% of the time about synthetic content, because the reflex is to assume things are real. Familiarity with AI made no measurable difference: people who described themselves as highly familiar scored 51.9%, the unfamiliar 51.1%.

The authors' conclusion is worth quoting exactly, because it is stronger than most coverage of this subject will tell you: people's perceptual detection capabilities are “no longer a suitable defense against deceptive synthetic media.”

So treat what follows as a tiebreaker, not a test. It can raise your suspicion. It cannot settle the question, and the next two sections are the ones that actually can.

01
Things that could not work. The most useful category, and the least intuitive. A 2025 study at CHI annotated hundreds of AI images and found that functional implausibilities — guitar strings that are not taut, a tennis racket strung asymmetrically, a backpack strap merging into a jacket — were the hardest for people to catch. Not because they are subtle, but because nobody is looking. Ask whether the object in the picture could actually do its job.
02
Background semantics. Foregrounds get the model’s attention; backgrounds get filled in. Look for architecture that does not join up, signage that means nothing, crowds where individual people dissolve, and objects that begin as one thing and end as another.
03
Light, shadow and reflection. Check that shadows fall consistently with a single light source, and that reflections in glass, water and eyes agree with what is in front of them. This is geometry rather than perception — it can be reasoned about rather than felt — which is why forensic analysts use it and casual viewers do not.
04
Hands, but only in one direction. This is the tell everyone knows, and it is now half broken. Malformed hands still occur, so bad hands remain evidence of AI. But Midjourney was rendering five-fingered hands reliably by March 2023, and every major model since does it routinely, so good hands are evidence of nothing at all. If you count five fingers and conclude the photo is real, you have not performed a test. That asymmetry — useful when it fails, worthless when it passes — applies to nearly every item on this list.
05
Text, with the same caveat. Garbled lettering used to be a reliable giveaway. It is not any more: current models render short text accurately, and recent evaluation found leading systems holding up to around 800 characters of English before accuracy degrades. Non-Latin scripts remain notably weaker. So garbled text still points at AI; clean text does not point anywhere.
06
Time, which is the only lever that reliably helps. In the CHI study, accuracy on AI images rose from 72% when people had one second to look, to 82% at twenty seconds. The share of images that people judged worse than chance fell from 43% to 17%. “Look longer” is defensible advice supported by measurement. “Check the hands” is not.
What This Article Used to Say

An earlier version of this piece told you that malformed hands were the single most reliable tell in static images, and that legible text in an image meant it was probably real. Both were true when they were written and both are now wrong — the second one dangerously so, because it invited you to conclude that a clean image was authentic. If you learned those rules here, this is me retracting them.

🎬
Deepfake Video — Motion Artifacts and Temporal Tells
Applies to: news clips, video calls, political content, corporate impersonation

Video deepfakes must maintain consistency across every frame — a much harder problem than generating a single image. That computational challenge is where the tells concentrate.

01
Face boundary blur. The edge where the generated face meets the neck, hairline, or background is where compositing artifacts are most visible. Look for a subtle "halo" effect, inconsistent sharpness, or a slight color fringing around the face perimeter — particularly when the subject moves.
02
Unnatural blinking. Early deepfake models rarely blinked, because blinking was underrepresented in training data. Modern models do blink, but the frequency, duration, and timing of blinks may still be abnormal — either too infrequent, too fast, or synchronized oddly with speech pauses.
03
Lighting inconsistency between face and background. If the light source for the face appears to come from a different direction than the light source in the rest of the scene, the face has been composited. Pay attention to where highlights fall on the nose and forehead versus how light falls on nearby objects.
04
Accessories across frames. Earrings, glasses, and chains are rendered fresh on each frame by most deepfake systems. Watch for jewelry that appears to slightly shift position, change shape, or briefly disappear between frames — particularly during motion or fast cuts.
05
Lip sync inconsistency. Watch whether mouth movements precisely match the audio phonemes. Minor delays (even 50–100ms), or mouth shapes that don't quite match the sounds being produced, are common. This is particularly visible on words with distinctive mouth shapes — "P" and "B" sounds require lips to touch and separate.
06
Emotional misalignment. The emotional expression on the AI-generated face may not quite sync with the emotional content of the words or the vocal tone. A genuine speaker's face shows micro-expressions that naturally accompany what they're saying — AI-composited faces can feel slightly "flat" or emotionally incoherent on close inspection.
07
Temporal flickering on the face. Pause a video at several random points and compare. The face region in a deepfake may show subtle frame-to-frame variations in skin tone, texture, or sharpness that the rest of the video does not — particularly in lower-quality productions. This is sometimes called "temporal inconsistency."
08
Resolution mismatch. The synthetic face may render at a slightly different effective resolution than the background or the rest of the body. This is more common in lower-effort deepfakes but is worth checking by zooming in — the face may appear cleaner or softer than the surroundings in a way that doesn't make physical sense for the shooting conditions.
🎙️
Voice Clones — What to Listen For
Applies to: phone calls, voicemails, audio messages, video call audio

Voice cloning technology has improved dramatically. As of 2023, tools available to anyone can clone a recognizable voice from as little as three seconds of audio, according to McAfee's Artificial Imposters study. But every synthesis system leaves acoustic signatures — patterns in how it handles the things that make real speech sound human.

01
Missing disfluencies. Real human speech contains natural "disfluencies" — brief hesitations ("um," "uh"), false starts, self-corrections, and subtle pauses mid-sentence. AI-synthesized speech tends to flow too cleanly. Every word is pronounced correctly, every sentence completes without error. This perfection is unnatural.
02
Unnatural cadence or pacing. Listen for speech rhythm that doesn't sound quite right — emphasis that lands in slightly odd places, pauses between words or sentences that don't match how the person naturally speaks, or a slightly mechanical syllabic rhythm even within otherwise fluent speech.
03
Breathing patterns. Real speakers breathe audibly — brief inhales before sentences, slight exhales mid-phrase. AI voice synthesis commonly omits breathing sounds entirely, or inserts them in unnatural positions. Listen for where the speaker "takes breath" — or whether they seem not to breathe at all.
04
Emotional flatness or inconsistency. Even AI systems trained to inject emotion often produce emotional tones that don't quite match the content of the words. A voice that sounds slightly too neutral for a distressed message, or emotional peaks that don't land on the words they logically should, are flags.
05
"Hollow" or digital vocal quality. A subtle metallic, hollow, or slightly reverb-like quality in the voice — particularly on sustained vowels — can indicate synthesis. It may be most noticeable when the voice is speaking loudly or with emphasis, as amplitude variations expose synthesis artifacts more readily.
06
Speech pattern inconsistencies. If you know the supposed speaker, consider: do they actually use those words? That phrasing? That accent? Voice cloning replicates vocal timbre but is constrained by the text it's given to speak — the clone may "sound like" someone but say things that person would never say.
07
Audio environment mismatch. Real calls and voice messages carry room acoustics — slight reverb from walls, background noise that changes when the speaker moves. If the voice sounds like it was recorded in an acoustic vacuum while the caller claims to be calling from a busy airport, that's worth noting.

Provenance — The Question That Replaced “Does It Look Real?”

If perception no longer settles it, the alternative is to ask where a file came from rather than what it looks like. That is what C2PA does. It is a specification for attaching a cryptographically signed record to an image or video — what device made it, when, what software touched it since — maintained by a body whose steering committee includes Adobe, Google, Microsoft, OpenAI, Sony, the BBC, Meta and TikTok.

You check it by dragging the file onto contentcredentials.org/verify, which opens the Content Authenticity Initiative’s Verify tool. If a credential is present you see the signer, the timestamp, the capturing device and a list of the edits recorded since. OpenAI runs a separate checker that recognises only its own output.

Where it exists, this works. Leica shipped the first camera to sign at the point of capture in October 2023, the M11-P. Google’s Pixel 10 is the first phone to sign every JPEG its camera takes, and Pixel Camera is the only mobile app so far to reach Assurance Level 2, the highest tier C2PA currently defines. Canon began rolling out its own system in May 2026 for the EOS R1 and R5 Mark II — initially in Europe, the Middle East and Africa, and, per Canon’s own footnote, requiring paid activation.

Four Things to Know Before You Trust a Credential

The cryptography is not the weak point. The signer is. In August 2025 Nikon shipped Content Credentials on the Z6III. Within about a week the service was suspended, and by late September Nikon had invalidated every certificate it had issued — after a researcher demonstrated that the camera could be induced to sign an image it had never taken, and later, to sign a wholly AI-generated one. No cryptography was broken. The camera was persuaded to certify a lie, and the result was a mathematically perfect credential attesting to something false.

Revocation barely helped. When those certificates were withdrawn, most validators carried on accepting them, because checking for revocation is optional by design — the specification makes it so out of privacy concerns. A file signed by a revoked certificate can still come back clean.

Almost nothing survives to reach you. Social platforms strip image metadata as standard, and publishing pipelines strip it unless someone deliberately configures them not to. Credentials exist; you will rarely meet one.

A photograph of a screen defeats it entirely. This has a name — air-gapping — and YouTube names it in its own help pages, warning that someone can point a certified camera at a monitor displaying synthetic content.

Two more caveats worth carrying. C2PA is not an ISO standard, despite what a great deal of coverage says; the ISO project, 22144, is still a committee draft. And in April 2026 a security team including researchers from UMBC and Hacker Factor published the first independent analysis of the specifications, concluding they “fail to achieve their claimed security goals” and should not yet be relied on for journalism or legal evidence.

The most important rule is the one about absence. No credential tells you nothing whatsoever. Google says so about YouTube’s camera-capture label; OpenAI lists five innocent reasons a signal may be missing, starting with the file simply predating the technology and ending with a platform having stripped it in transit. Missing credentials are the normal case, not a red flag.

Context — What Actually Settles It

When professional fact-checkers evaluate something, the first thing they do is leave. Sam Wineburg and Sarah McGrew put historians, Stanford undergraduates and professional fact-checkers in front of unfamiliar websites and watched how they worked. The historians and students read vertically: they stayed on the page, studied its logo, its domain, its design — and were fooled by all three. The fact-checkers read laterally, leaving within seconds to open new tabs and find out what the rest of the internet said. They reached better-supported conclusions in a fraction of the time.

Mike Caulfield turned that behaviour into four moves, taught in university libraries under the name SIFT:

01
Stop. Before you share, before you react. The emotional pull of an image is the mechanism, not a side effect.
02
Investigate the source. Not the page in front of you — what others say about it. Open a tab.
03
Find better coverage. If a photograph shows something enormous, ask who else is reporting it. A genuine event has more than one witness.
04
Trace claims, quotes and media to the original context. This is the one that does the work on images — not is this fake but where did this first appear, and what was it of.

For that last move, reverse image search is the tool, and the engines do different jobs. TinEye returns only exact or modified copies, never merely similar ones, and its Oldest sort surfaces the earliest crawl it has — which is how you date an image’s first appearance. Google Lens has the broadest index and is best at identifying places and objects. Bing and Yandex sometimes surface what the others miss. Practitioners run the same image through several rather than trusting one, and the free Search by Image extension does it in parallel.

And the thing this whole article risks obscuring: outright fakery is rarer than recontextualisation. Bellingcat, whose entire practice is this work, has put it plainly — manipulated images are far rarer than old images posted out of context or deliberately mislabelled. The Aldi panic-buying video that went round during Covid was a real video, of a real shop, from Kiel in 2011. Nothing in it was synthetic. Reverse image search caught it in under a minute; no amount of squinting at pixels ever would have.

Behavioral Red Flags — What the Attacker Is Asking You to Do

Technical tells become less reliable as generation quality improves. Behavioral patterns are far more durable — because the goal of a deepfake attack is to get you to do something, and that goal creates predictable behaviors regardless of how convincing the media is.

"Hi, it's Dad. I'm in an emergency situation — I can't talk long. I've been in an accident, my phone is damaged, I'm using someone else's phone. I need you to send $800 via Zelle right now, I'll explain everything later. Please don't call the family, I don't want to worry everyone. Can you do it now?"

This script — or close variants — is used in what the FTC calls "family emergency scams," increasingly paired with AI voice cloning of the supposed family member's voice.

Urgency + money request in the same communication. Legitimate emergencies involving family do not require you to wire money to an unfamiliar account within minutes. The pressure to act before you can think is by design.
Request to not verify through other channels. "Don't call your brother, it'll worry him" or "I can't have you calling the hospital right now" are manufactured reasons to prevent you from calling a known number and confirming the situation independently.
Refusal to appear on video or switch platforms. A deepfake voice caller cannot show their face live and unscripted. Requests to keep communication on voice only, especially with technical excuses ("my camera is broken"), are a warning sign.
Inability to answer personal verification questions. Ask something only the real person would know — a family memory, a pet's name, a shared private experience. Voice clone systems cannot generate correct answers to questions the attacker doesn't know.
Payment method specificity. Requests for wire transfer, gift cards, cryptocurrency, or Zelle are chosen because they are difficult to reverse. Legitimate urgent situations don't require you to buy Amazon gift cards.
Contact through an unexpected or unfamiliar channel. Receiving a distress call from an unfamiliar number, a LinkedIn message from an unusual email address, or a WhatsApp message claiming to be from a known contact is a structural red flag — attackers can't use the real person's established channels.
Extreme story inconsistencies. AI attackers research targets but imperfectly. The "CFO" in a business email compromise may not know a project name. The "family member" may not remember what city they supposedly live in. Probe with specific details and watch for vagueness.

Detection Tools — What to Use and When

No single tool is definitive, but several are fast, free, and catch a significant proportion of what circulates. Use them as a first check — not as a final verdict.

Google Reverse Image Search Free
Image verification
Drag any image to images.google.com to find where it appears online. If a "journalist" profile photo returns results from a completely unrelated context, it's likely stolen or synthetic. The fastest first check for any suspicious image.
images.google.com
TinEye Free / Premium
Reverse image search
Indexes images differently than Google — often finds older instances or sources that Google misses. Especially useful for tracking the original source of a manipulated image and identifying where it first appeared.
tineye.com
FotoForensics Free
Error-level analysis
Performs error-level analysis (ELA) on images — areas that have been digitally altered appear at different compression levels than the original. Academic in nature but freely accessible. Highly useful for detecting spliced or manipulated images. Developed by Neal Krawetz.
fotoforensics.com
AI or Not Free / $10 mo
AI image detection
Upload any image and receive a probability score estimating whether it was AI-generated. Covers Midjourney, DALL-E, Stable Diffusion, and other generators. Free tier allows a limited number of checks; paid tier removes limits.
aiornot.com
InVID WeVerify Free (browser ext)
Video + image verification
A browser extension built and maintained by AFP’s Medialab, used by AFP, Bellingcat and fact-checking organisations to verify video content. Extracts keyframes from video for reverse image search, checks metadata, and connects to multiple verification services. Developed by the InVID/WeVerify consortium with EU research funding — built for real-world use.
invid-project.eu
Deepware Scanner Free
Video deepfake detection
Submit a video URL or upload a clip to receive a deepfake probability score. Scans frame-by-frame and highlights segments flagged as likely synthetic. One of the more accessible free options for analyzing video deepfakes directly.
deepware.ai
Illuminarty Free / $8 mo
AI image + localization
Unlike most AI detectors, Illuminarty also attempts to highlight which specific regions of an image are likely AI-generated — useful for detecting partially manipulated images where only a face or background was replaced.
illuminarty.ai
Reality Defender Enterprise
Multi-modal deepfake detection
Enterprise-grade platform covering video, audio, and image detection simultaneously. Used by media organizations, financial institutions, and governments. Not consumer-facing, but relevant context — it represents the state of the art in commercial detection capability as of 2026.
realitydefender.com

For academic research context on detection technology, DARPA’s Media Forensics (MediFor) programme funded much of the foundational work before concluding around 2020, and its successor SemaFor has since wound down as well, and Hany Farid at UC Berkeley is among the most cited researchers in deepfake detection. Both are worth following if you want to track the technical frontier of the field.

What to Do When You're Not Sure

Slow down deliberately. The most effective defense against deepfake fraud is time. Urgency is manufactured. If anyone — a caller claiming to be a family member, a CEO in a video message, a government official in a social post — is creating pressure to act immediately, that pressure is the attack. Take the time to verify.
Call back on a known number. Never call back on the number that contacted you. If a family member calls in distress, hang up and call their actual phone number from your contacts. If a company calls you, hang up and call the main number from the company's official website. This simple step defeats most voice clone attacks.
Establish a family code word. Choose a word or phrase that only your family knows, and agree to use it to verify identity in any suspicious emergency communication. A voice clone cannot produce a code word the attacker doesn't know. This is recommended by the AARP Fraud Watch Network.
Run the four-second video call test. In any video call where identity matters, ask the person to do something unrehearsed — turn their head slowly to the side, hold up a specific object, or perform a simple physical action. Current real-time deepfake systems degrade visibly during unexpected movements and object interaction.
Use a verification tool before sharing or acting. Before forwarding any image or video that seems significant or shocking, run it through Google Reverse Image Search and one AI detection tool. This takes 60 seconds and catches a substantial proportion of synthetic media in active circulation.
Look for C2PA watermarks on trustworthy content. The Coalition for Content Provenance and Authenticity (C2PA) is developing a technical standard for cryptographic content credentials that certify the origin and editing history of media. Major camera manufacturers, Adobe, and some news organizations are implementing this. An image or video with verified C2PA credentials is significantly harder to fake.
Ask the right questions about the content itself. Who benefits from you believing this? What would you need to do if it were true? Why is this reaching you this way, at this time? The answers to these questions reveal a lot about whether content is legitimate or manufactured.

If You've Been Targeted

If you've sent money, shared personal information, or been manipulated by content you now believe was AI-generated, act quickly — not to recover the money (which is often unrecoverable) but to document and report, which helps track and potentially stop the operation targeting others.

Report to the FTC at reportfraud.ftc.gov. Fraud reports feed into the Consumer Sentinel Network used by law enforcement.
File an IC3 complaint at ic3.gov — the FBI's Internet Crime Complaint Center. There is no minimum — IC3 accepts complaints at any value, and filing is what triggers the FBI’s recovery process for wire transfers.
Call your bank immediately if a wire transfer was involved. Some international transfers can be recalled within 72 hours if reported fast enough — call, do not email.
Contact the AARP Fraud Helpline at 877-908-3360 — available to all ages, not only AARP members. Staff are trained specifically on AI fraud schemes and can advise on next steps.
Conclusion
The tells are real. The defense is learnable.

Deepfake quality will continue to improve, and some of the specific visual artifacts described here will eventually be resolved by better generation systems. But the behavioral patterns of deepfake attacks — urgency, financial pressure, manufactured reasons to avoid verification — are structural features of the fraud, not features of the technology. Those won't change.

The most durable protection isn't a detection tool. It's the habit of slowing down before acting on anything that arrived unexpectedly, claimed to be urgent, and asked you to do something you'd normally verify. A phone call, a code word, an extra 90 seconds — these are not inconveniences. They are the difference between a close call and a $25 million wire transfer.

Sources and References
[1]
South China Morning Post — "Everyone looked real: multinational firm’s Hong Kong office loses HK$200 million after scammers stage deepfake video meeting," Feb 2024.
reuters.com ↗
[2]
FBI Internet Crime Complaint Center (IC3) — "Deepfakes and Stolen PII Utilized to Apply for Remote Work Positions" (I-062822-PSA), June 28, 2022.
ic3.gov ↗
[3]
McAfee — "Artificial Imposters: Cybercriminals Turn to AI Voice Cloning for a New Breed of Scam," 2023.
mcafee.com ↗
[4]
Sumsub — "The Sumsub Identity Fraud Report," 2023 — documents 10× increase in detected deepfakes year-over-year.
sumsub.com ↗
[5]
Europol — "Malicious Uses and Abuses of Artificial Intelligence" — co-authored with UNICRI and Trend Micro; covers deepfake threat landscape.
europol.europa.eu ↗
[6]
European Parliament — "Tackling deepfakes in European policy," STOA study, July 2021. EPRS_STU(2021)690039.
europarl.europa.eu ↗
[7]
Stanford Tech Impact & Policy Center — successor to the Stanford Internet Observatory, which was wound down in 2024.
tip.fsi.stanford.edu ↗
[8]
DARPA — Media Forensics (MediFor) program; automated detection of manipulated media from 2016 onward.
darpa.mil ↗
[9]
Hany Farid, UC Berkeley — prolific deepfake detection researcher; research page covers detection methodology and published findings.
hfarid.org ↗
[10]
Sensity AI (formerly Deeptrace) — tracks and categorizes deepfake content globally; publishes annual state-of-deepfakes reports.
sensity.ai ↗
[11]
InVID WeVerify Project — EU-funded video verification toolkit used by professional journalists and fact-checkers.
invid-project.eu ↗
[12]
FTC — "Scammers use AI to enhance their family emergency schemes" (March 2023) — recommends hanging up and calling back on a number you already have. Note the FTC does not recommend code words; that advice comes from AARP.
consumer.ftc.gov ↗
[13]
WITNESS — media preservation and video verification resources for human rights and general public use.
witness.org ↗
[14]
C2PA — Coalition for Content Provenance and Authenticity; open technical standard for verifiable content credentials.
c2pa.org ↗
S
Sheldon Valentine
Founder · Dear Tech

Sheldon writes about AI safety, tools, and the practical knowledge people need to navigate an AI-saturated world. Dear Tech's editorial commitment is honest, specific, and never alarmist — but also never dismissive of risks that are real and present.

Share
Back to Blog