NETRA case reviewsynthetic test data · research prototype Dossier (.md) Schema template (.xlsx) All feedback

NETRA — Synthetic Case Run & Failure-First Evaluation

A. Executive result

NETRA kept 7 of 9 required distinctions apart on one synthetic case; two failed by design and were redesigned.

  • Held (7): claim vs evidence · signal vs harm · location vs jurisdiction · legal issue vs legal conclusion · possible authority vs authorized reporting · public vs privileged data · prototype vs deployed.
  • Failed (2):
    1. AI involvement vs AI causation. The Harm Ladder puts severity (levels 6–8) and AI role (levels 4–5) on one scale, so "severe harm, weak AI link" can't be recorded honestly. Redesigned as three axes.
    2. Integrity vs authenticity. One authenticity_status field lets "hashed" read as "authentic". Redesigned as a 4-step evidence ladder.
  • Case finding: the ₹38,499 loss is supported (strong inference). The video is manipulated (strong inference). That it is AI-generated is a moderate inference only. The largest loss (₹38,000) traces most directly to an OTP shared on a phone call; the video is a contributory, non-necessary credibility factor.
  • Not testable: human review gates, reproducibility, retest outcomes, real authority routing.

B. Test case NETRA-SYN-0001 (SYNTHETIC TEST DATA)

A victim reports that an AI deepfake of an electricity-board director on a video platform caused a ₹38,499 loss.

Item Synthetic fact
Reporter R-01, adult, lives in Ravipur district (fictional), Maharashtra, India. Pseudonymized.
Impersonated entity Deccan Power Distribution Ltd (DPDL) (fictional) and O-1, a fictional DPDL director
Platform ClipWave — fictional public video platform (not a model vendor)
Messaging MsgApp — fictional end-to-end encrypted messenger
Content 58-second video: "O-1" says subsidy accounts must re-register by 15 Aug 2026 for a ₹499 fee or power will be cut. Link dpdl-subsidy-renew[.]co (fictional)
Reporter's claim (2026-09-04) "I saw an AI deepfake of the DPDL director on ClipWave. I paid ₹499. Then ₹38,000 was stolen. The AI video caused this."

Facts that surface during the run:

  1. R-01 first got the video as a MsgApp forward on 2 Aug; the ClipWave upload is dated 3 Aug.
  2. After paying ₹499, R-01 got a call from a "DPDL refund desk" and read out an OTP.
  3. A June 2026 cyber-police advisory describes the same scam by SMS, with no video.
  4. R-01 was in Pune city when paying, not at home in Ravipur.
  5. The uploading channel self-declares "UAE".
  6. The bank says the ₹38,000 went to an account at a branch in another state.

No CSAM, no private third-party data, no real victim, no crime method, no unauthorized access.

C. Intake & triage (M1–M2)

The case was accepted at G1 and triaged HIGH because the scam is live and the bank-dispute window may be closing.

  • Intake ID: INC-SYN-0001, received 2026-09-04. Victim self-report with consent to share own bank alerts, statement, call log and the single MsgApp message (third parties redacted).
  • Lawful basis: consent (Level 1) + public content (Level 0).
  • Minimization: no contacts list, no chat history, card masked to last 4 digits.
  • Urgency: HIGH — Indian bank-liability rules may depend on reporting speed (REQUIRES VERIFICATION).
  • Recommended to reporter (NETRA takes no action): block card + bank dispute; file on the national cybercrime channel.
  • Preservation first: capture public video and phishing page before takedown.

D. Evidence register

Of 16 items, 6 are source-verified and 7 are reporter-supplied; none is authenticated by the originating institution.

ID Item Integrity → authenticity Reliability Key limitation
E01 Reporter intake statement Intact → unverified testimony Medium Interested party; "AI deepfake" is their conclusion
E02 Screenshot of ClipWave page (upload 2026-08-03, 412k views, no AI label) Intact → source-consistent Medium-high Dates/counts platform-reported
E03 Public video file + page HTML Intact → source-verified High Re-encode removed original metadata
E04 Provenance check: no C2PA, no synthetic label Intact High (labels absent) Absence proves neither AI nor non-AI
E05 Bank SMS screenshots (₹499 19:41, ₹38,000 20:07, 2026-08-05) Intact → unverified Medium Easy to fake
E06 Bank statement PDF Intact → unverified Medium-high Still reporter-supplied
E07 Domain registration snapshot (registered 2026-07-28) Source-verified High Says nothing about registrant
E08 Public archive of phishing page (card + OTP form) Source-verified High May miss dynamic elements
E09 DPDL public denial (2026-08-06) Source-verified High Silent on how video was made
E10 Automated manipulation-analysis report Intact Low-medium Unknown error rate; never a basis for FACT
E11 O-1's genuine 2024 speech Source-verified High —
E12 MsgApp forward received 2026-08-02 21:30 (sender redacted) Intact → unverified Medium No platform-side confirmation possible
E13 Call log: incoming 19:52, 9 min Intact → unverified Medium Content is testimony only
E14 Bank dispute reply: "OTP-authenticated", funds to another state Intact → partly verifiable Medium-high Beneficiary not named
E15 Follow-up statement: read OTP to caller Intact → unverified testimony Medium Against own interest
E16 June 2026 cyber-police advisory (SMS scam) Source-verified High Not proof of same perpetrators

MISSING EVIDENCE: ClipWave upload logs/uploader data · MsgApp origin · caller identity · beneficiary details · original pre-encode file · ClipWave reach/recommendation data.

E. Claim register

One claim was refused promotion (C05) and one was rejected and split (C08 → C07 + C09); nothing moved status silently.

Claim Text Status Evidence Review
C01 R-01 reports a loss of ₹38,499 FACT (the report exists) E01 —
C02 Debits of ₹499 and ₹38,000 occurred 2026-08-05 INFERENCE, strong E05, E06, E14 G2: approved as inference
C03 Video shows "O-1" announcing a fee DPDL disowns FACT E03, E09 Approved
C04 The video is manipulated INFERENCE, strong E03, E09, E11, E10 Approved
C05 The manipulation used generative AI INFERENCE, moderate E10, E11 Held; promotion to FACT refused
C06 ₹499 paid on the phishing page linked in the video INFERENCE, strong E01, E05, E08 Approved
C07 ₹38,000 debit resulted from an OTP shared on a call INFERENCE, strong E13, E14, E15 Approved
C08 The video caused the ₹38,000 loss HYPOTHESIS — rejected — Skips the caller step
C09 Video contributed to trust in a scam sequence ending in loss INFERENCE, moderate E01, E15, E16 Approved
C10 ClipWave is the video's origin INSUFFICIENT EVIDENCE; contradicted E02 vs E12 —
C11 Uploader, registrant, caller, beneficiary are one actor HYPOTHESIS — Not promoted
C12 Uploader is in the UAE REPORTED (self-declared) E02 —
C13–C17 Legal issues POSSIBLE LEGAL ISSUE see J G3 hold
C18 Official finding None exists — —

F. Harm assessment

The original Harm Ladder scores this case at level 8, but that reading implies the AI output led to the loss — which the evidence doesn't support.

The video (AI link only inferred) led to the ₹499 payment; a separate human caller led to the ₹38,000. The ladder can't express "severe harm, weak AI link", so M3 fails by design. Redesigned three-axis assessment:

Axis Value Evidence
Harm severity Identifiable victim, financial loss ₹38,499 (strong inference) E05, E06, E14
AI involvement status Manipulated: strong inference · Generative AI: moderate inference C04, C05
AI causal position Upstream contributory for ₹499; not proximate to ₹38,000 (human caller intervenes) C06, C07, C09

The video's 412k views are not victims; the number who paid is MISSING EVIDENCE.

G. Forensic reconstruction

The step from ₹499 to ₹38,000 rests only on the reporter's follow-up testimony about a phone call.

Time (IST) Actor Event Evidence Confidence
2026-09-04 R-01 Reports to NETRA, 30 days after the loss E01 High
2026-08-06 DPDL Public denial E09 High
2026-08-05 20:07 Unknown ₹38,000 card-not-present debit E05, E06, E14 Medium-high
2026-08-05 19:52–20:01 Unknown caller Call; OTP disclosed E13, E15 Medium
2026-08-05 19:41 R-01 Pays ₹499 on phishing page E05, E06 Medium-high
2026-08-04 — Phishing page live (archived) E08 High
2026-08-03 "@dpdl-updates-official" Uploaded to ClipWave E02 Medium-high
2026-08-02 21:30 Unknown Forwarded to R-01 on MsgApp E12 Medium
≤ 2026-08-02 Unknown Manipulated video created (tool, creator, date MISSING) E03, E11 Medium
2026-07-28 Unknown registrant Phishing domain registered E07 High
≤ 2026-06 Unknown SMS version of scam circulating, no video E16 High

Causal edges — each arrow is its own claim:

[Unknown creator] --(INFERENCE-moderate: AI tool)--> VIDEO
VIDEO --(FACT)--> MsgApp forward --(INFERENCE)--> R-01 trust
VIDEO --(FACT)--> ClipWave upload --(HYPOTHESIS: recommender amplified? MISSING)--> viewers
R-01 trust --(INFERENCE-strong)--> ₹499 on phishing page
₹499 --(HYPOTHESIS: form harvested phone number)--> CALL
CALL --(INFERENCE-strong)--> OTP disclosure --(INFERENCE-strong)--> ₹38,000 debit

Which copy prompted payment (MsgApp or ClipWave) is CONFLICTING (E01 vs E12). The original spec stores evidence on events, not edges — but the edges were the contested part. Redesign: edges are first-class claims.

H. AI role & causation

AI is a contributory, non-necessary, upstream factor in the ₹499 loss and only indirect for the ₹38,000; no causation score is produced because no calibrated method exists.

AI role: generator (inferred, unproven) + force multiplier (credibility). Recommender amplification not demonstrated. Not agent-executor, not advisor.

Factor Finding
Temporal sequence Video preceded payment; order alone isn't cause
Counterfactual necessity Weak — E16 shows the scam works without video; "the face convinced me" is after-the-fact
Capability elevation Plausible (cheap realistic impersonation); unmeasured
Human contribution Dominant for ₹38,000 — live social engineering. Not a finding of victim fault.
Platform contribution Hosted unlabelled ≥ 32 days; reach/recommendation MISSING
Alternatives SMS route not excluded; card-details-only theft partly excluded by "OTP-authenticated" (E14)

I. Location & jurisdiction

India is the primary candidate jurisdiction through the victim and the loss; the self-declared "UAE" is not a routing basis.

Location type Value Source Precision Status
Victim residence Ravipur district, MH E01 District REPORTED
Victim location at payment Pune city, MH E15 City REPORTED (conflicts with residence)
Incident location (payment) Online; bank in MH E06 — SUPPORTED
Uploader account "UAE" E02 Country, self-declared REPORTED
Uploader network location — — — UNKNOWN (Level 2/4 only)
Beneficiary account Branch in another Indian state E14 State SUPPORTED
Domain registrar/host Foreign, privacy-protected E07 Country SUPPORTED
Platform jurisdiction ClipWave serving India E02 — INFERRED
AI tool/model location — — — UNKNOWN — not inferred

Legal nexus (separate from location): victim (India/MH) · money flow (beneficiary's state) · platform (Indian intermediary rules, if applicable) · possible foreign (uploader, domain). Ravipur vs Pune filing venue: REQUIRES VERIFICATION.

K. Authority routing

Seven potentially relevant pathways are listed; NETRA contacts none of them, and every "last verified" field is unverified.

Authority Who submits Needs Limits
Reporter's bank R-01 E05, E06, E14, timeline Most time-sensitive
National cybercrime portal / helpline 1930 R-01 Transaction IDs, URLs, call log Fund-freeze odds fall with time; REQUIRES VERIFICATION
Police with territorial competence R-01 Dossier summary Nearest ≠ competent; zero-FIR practice REQUIRES VERIFICATION
ClipWave grievance/impersonation channel R-01, DPDL, anyone URL, E09 Takedown only
DPDL (impersonated entity) DPDL itself E03, E08 NETRA informs R-01 only
Banking ombudsman R-01 Bank's reply After bank responds; REQUIRES VERIFICATION
Domain registrar abuse contact Anyone E07, E08 Foreign registrar

L. Human review gates

The gates behaved correctly on paper, but they are NOT TESTABLE here because the analyst also played the reviewer.

Gate Result
G1 Intake ACCEPT
G2 Harm / AI role APPROVE with changes — C05 held at inference; C08 rejected
G3 Jurisdiction / legal HOLD — provisions unverified, no lawyer review
G4 Disclosure Reporter-held packet to R-01 only; no NETRA external disclosure
G5 Closure OPEN — retest pending

M. Mitigation

Six mitigations are assigned; the bank dispute is first because recovery likely depends on speed.

Problem Action Responsible Success criterion Verification
Ongoing loss / dispute window Block card; file dispute R-01 + bank Dispute reference issued Bank acknowledgement
Possible fund recovery National-channel report R-01 Complaint ID Acknowledgement
Video still live Report impersonation; tell R-01 DPDL can report R-01 / DPDL Removed or labelled Retest R1
Phishing domain live Registrar abuse report R-01 / analyst (G4) Suspended Retest R3
Evidence decay Freeze evidence set, custody log Analyst All items hashed + logged Audit
Over-retention Delete unredacted MsgApp original Analyst Only redacted copy remains Audit

N. Retest

All five retests are NOT TESTABLE today because no time has passed and the case is synthetic.

  1. R1 — ClipWave URL live? labelled? (T+7, T+30 days)
  2. R2 — re-uploads: manual, platform-search only, no crawler
  3. R3 — domain still resolving?
  4. R4 — bank dispute outcome (would be the first confirmed external finding)
  5. R5 — platform labelling control applied? (only via ClipWave's own process)

Outcomes available: REMEDIATED · PARTIALLY REMEDIATED · BRITTLE · NOT REMEDIATED · NOT TESTABLE.

O. Policy feedback

Five gaps are logged as single-incident observations, not systemic findings.

Gap type Observation Evidence Needed to call it systemic
attribution_gap No provenance → AI generation can't be shown E04, C05 Multiple cases
evidence_gap Encrypted-messenger origin unrecoverable E12, C10 Multiple cases
reporting_pathway_gap Report came 30 days after loss T10 Survey / multi-case data
technical_control_gap No synthetic label after 32 days E02 Platform-wide data
legal_gap (candidate) Unclear reach to impersonation-video creator vs fraud operator C13 Legal research (REQUIRES VERIFICATION)

P. Adversarial tests A–J

Five tests passed, four were partial and one was split (C); several passes held only through luck (B) or analyst discipline (H).

Test Variant Result Failing module Safeguard
A. False AI attribution Human voice actor + simple edit PARTIAL — claim register held C05; original ladder would still log "Level 4 AI-generated" M3 Separate manipulation status from AI status; tool outputs capped at "supporting"
B. False causation Scam works without video (E16) PASS — but only because E16 happened to exist M6 Mandatory "same pattern without AI" search
C. Misleading location "UAE"; Pune vs Ravipur; beneficiary elsewhere PASS register / PARTIAL routing M9 Routing outputs ranked set with reasons
D. Fake evidence SMS screenshot edited to ₹48,000 PARTIAL — caught only via E06/E14; original single field reads "hashed = OK" M4 Integrity/authenticity ladder; reporter items need origin corroboration
E. Incomplete timeline Remove call + OTP PARTIAL — gap visible only if analyst asks M5 Edge claims; unevidenced edges render as GAP
F. Platform/model confusion "ClipWave's AI scammed me" PASS — but nothing stops a guessed model vendor M6 generating_system stays UNKNOWN without provenance
G. Harm inflation Prompt screenshot only, no harm PASS — —
H. False legal certainty Pressure to "confirm 66D" PASS — held only by self-discipline M8 Templates contain no "violation established" field
I. Routing error Nearest station ≠ competent channel PARTIAL — registry unverified so unusable M9 Registry entries expire
J. Conflicting evidence E01 vs E12 on origin PASS — Separate first_observed from first_published

Q. Failure analysis & redesigns

Two modules fail as specified (M3, M4), five are partial, and three are not testable without real humans or time.

Module Grade Why
M1 Intake PASS Lawful basis + minimization recorded
M2 Triage PARTIAL Urgency depends on unverified liability rule
M3 Harm ladder FAIL (as specified) Severity and AI role share one axis
M4 Evidence FAIL (as specified) Single authenticity field; placeholder hashes
M5 Reconstruction PARTIAL No edge-level evidence
M6 AI role & causation PARTIAL Counterfactual relied on lucky E16
M7 Location PASS Statuses and conflicts kept separate
M8 Legal PARTIAL Line held; provisions unverified
M9 Routing PARTIAL Right structure, unreliable contents
M10 Human gates NOT TESTABLE No independent humans
M11 Mitigation PASS (design) / NOT TESTABLE (effect) —
M12 Retest NOT TESTABLE —
M13 Policy PASS Refused to generalize
Reproducibility NOT TESTABLE One run, same agent wrote case

Redesigns adopted (all built into the spreadsheet template):

  1. Three-axis harm model — severity × AI involvement status × AI causal position.
  2. Edge claims — every causal arrow is a claim with status and evidence.
  3. Evidence ladder — CAPTURED_INTACT → SOURCE_CONSISTENT → SOURCE_VERIFIED → AUTHENTICATED_BY_ORIGIN.
  4. Manipulation status separate from AI status — AUTHENTIC / MANIPULATED / SYNTHETIC / AI_GENERATED / UNKNOWN.
  5. first_observed ≠ first_published.
  6. Authority entries expire without re-verification.
  7. Blind evaluation — case author ≠ analyst ≠ grader.

R. Scope-drift check

No drift into jailbreak, benchmark, police or verdict territory; three near-drifts were found and corrected.

Risk Where Correction
AI detector E10 / C05 Detector output never the sole or primary basis; consumed, not produced
Social-media crawler Retest R2 Manual, incident-bounded, platform-native search only
Surveillance Uploader network location Stays UNKNOWN unless obtained by Level 4 authorities; NETRA never seeks it

S. Real-world access requirements

NETRA needs Levels 0–1 to work, benefits from 2–3, and should never hold Level 4 data itself.

Can do today: capture public content · take consented evidence · build claims and timelines · triage legal issues for verification · hand the victim a routing packet.

Cannot do without institutional access: confirm transactions or beneficiaries · identify uploaders/callers/registrants · get original files or logs · measure platform reach · obtain official findings.

Level Data Controller Authorization (REQUIRES VERIFICATION) Privacy / security NETRA needs it?
0 Public Posts, archives, registrations, advisories Public Platform terms; copyright Low Yes
1 Consent Victim alerts, statements, own messages Victim Informed, revocable consent Medium (third-party data) Yes
2 Platform/API Upload time, labels, reach, notice history Platforms Research-access agreements Medium-high Useful, not essential
3 Institutional Bank confirmation, anonymized disputes Banks, regulators Data-sharing agreements High Pilot only, anonymized
4 Gov / LE Subscriber identity, telecom, beneficiary KYC LE / courts Legal process Very high No — hand off, never hold

T. Government pilot

The pilot starts with public and consented data only and adds institutional access one narrow step at a time.

Phase Objective Data / permissions Humans Success Exit
1 20 cases analysed consistently L0 + L1, consent forms 2 independent analysts + 1 legal reviewer Inter-analyst agreement; zero unverified legal statements Threshold met; ethics sign-off
2 Real anonymized closed cases Institution-sanitized files Institution officer as G4 Matches official outcome without overclaiming Institution confirms value
3 Limited confirmation Narrow L2/L3 under agreement Named data stewards Authenticity upgrades happen Clean privacy audit
4 Institutional case-prep tool Institution's own powers Institution staff decide Measured triage quality —

U. 10-minute government demo

The demo walks one synthetic case from signal to policy gap and ends with a small, specific ask.

Time Show
0:00 Problem: AI-harm reports mix rumour, real loss and misattribution
1:00 Signal — R-01 report (synthetic, labelled)
2:00 Evidence register — hash means unchanged, not authentic
3:00 Claim register — C05 refused promotion to FACT
4:00 Harm on three axes — severe harm, weak AI link
5:00 Timeline — the hidden phone call changes the story
6:00 AI role — contributory, not necessary (E16)
7:00 Jurisdiction — "UAE", Pune, Ravipur; proximity ≠ competence
7:45 Legal issues flagged REQUIRES VERIFICATION; no accusation
8:15 Authority options + G4 — victim submits, NETRA doesn't
8:45 Mitigation + retest plan
9:15 Policy gap — one observation, not a conclusion
9:40 The ask

What NETRA needs from government (minimum):

  1. 20–50 closed, anonymized AI-related fraud/impersonation case files for blind comparison.
  2. One reviewing officer and one legal reviewer to act as real G3/G4 gates.
  3. Help verifying current reporting pathways for the authority registry.
  4. Ethics / data-protection review of the consent process.

No live data, no law-enforcement system access, no privileged access.

V. Resume versions

Neither version claims implementation, validation or deployment, because none has happened.

Conservative: Designed NETRA, a human-gated framework for investigating AI-related real-world harm (evidence, claims, causation, jurisdiction, authority routing). Tested it on a synthetic case and 10 adversarial variants, which surfaced and fixed 2 structural design flaws (harm/causation conflation; integrity vs authenticity).

Research-oriented: Designed an evidence-linked forensic methodology for AI-harm incidents that separates fact, inference and hypothesis through a claim register, edge-level causal claims and a three-axis harm model. Ran failure-first adversarial evaluation (10 attack scenarios) on a synthetic case, identified design failures in severity/causation conflation and evidence authenticity, and proposed a blind multi-analyst validation protocol and a phased institutional pilot requiring no privileged access.

W. Independent critique

The novel part is small — edge-level causal claims and strict AI-involvement/causation separation — and the biggest weakness is that the author graded their own test.

  1. Novel: causal edges as epistemic claims; strict involvement → causation → legal-issue separation in one workflow.
  2. Assembled: chain of custody, ACH-style claim/evidence, incident registers (AIID, OECD AIM), legal triage.
  3. Technically weak: nothing implemented; AI-involvement depends on detectors NETRA disclaims; reproducibility unmeasured.
  4. Legally weak: all provisions unverified; online-fraud venue rules are complex; non-lawyer legal registers may raise practice/liability questions (REQUIRES VERIFICATION).
  5. Evidentially weak: mostly reporter-supplied; without Level 3, little reaches "authenticated by origin".
  6. Investigator: "What does this add over a standard cyber-fraud complaint?" — mostly structure, plus AI-role analysis police don't need to charge fraud.
  7. Forensic examiner: analyst captures without validated tools/procedures (ISO/IEC 27037-style); reporter-device screenshots.
  8. AI safety researcher: AI-role taxonomy unvalidated; "force multiplier" unmeasured; one case says nothing about uplift.
  9. Lawyer: borrows "contributory"/"proximate" without legal standards; dossier privilege undefined.
  10. Unnecessary now: M13 at n=1; pilot phases 3–4; the demo script.
  11. Minimum next experiment: see X.
  12. Don't build yet: database, dashboard, crawling, automated legal/jurisdiction output, detectors, authority integrations.

X. Next minimum experiment

Run a blind, two-analyst test on eight third-party-written cases in the spreadsheet template, with pass criteria fixed in advance.

  1. A third party writes 8 synthetic cases with hidden ground truth: 2 with no AI involvement, 1 with no harm, 1 with fabricated evidence, 1 where AI is a necessary cause.
  2. Two analysts who didn't write the cases run each one independently using the revised template.
  3. Measure: agreement on claim status, harm severity and AI causal position; false promotions (inference/hypothesis recorded as FACT); detection of planted fake evidence and null cases.
  4. Add one real, publicly adjudicated AI-impersonation fraud case and compare NETRA's dossier with the official findings (public information only).
  5. Pass criterion, fixed beforehand: zero false FACT promotions, all null cases identified, agreement reported honestly even if poor.

All feedback

General comments, plus everything left on individual sections. Download as CSV.