Skip to case study
Adhik AdhikariSelected work

Project 01

AI Content Attribution Pipeline

Provenance Guard

Provenance Guard explores how a content platform could make AI-attribution decisions without presenting a single opaque verdict as certainty.

Provenance Guard interface with text analysis, attribution analytics, and recent submissions
The working Provenance Guard interface, captured from the public project.
2
independent signals
3
attribution outcomes
50+
automated tests

Problem and approach

The system accepts submitted text and evaluates it through two independent signals. A Groq-powered LLM classifier looks at tone, phrasing, and structure, while pure-Python stylometric heuristics measure sentence-length uniformity, vocabulary diversity, and punctuation density.

A confidence scorer combines both signals and returns one of three outcomes: likely AI, uncertain, or likely human. The uncertain state is intentional; it prevents a borderline score from being presented as a definitive accusation.

Transparent product behavior

Each result includes a confidence score and a plain-language transparency label. Submissions are recorded in SQLite with both input signals, the combined decision, and its timestamp. An appeal workflow attaches a creator’s reasoning to the same record for later human review.

Evaluation changed the implementation

More than 50 automated tests cover scoring, labels, storage, classifier failures, and stylometric behavior. Evaluation exposed unreliable behavior on short inputs, so I recalibrated signal weighting instead of treating the first implementation as final.

Known limits

Stylometric statistics become unstable when there is too little text, and lightly edited AI writing remains difficult for both signals. The system is a decision-support prototype with explicit uncertainty and human review concepts—not proof of authorship.

Working stack

  • Python
  • Flask
  • Groq API
  • SQLite
  • Pytest