LLM evaluation pipeline

Source-Grounded Answer Verifier

When does verification actually make an LLM answer better?

I built a reproducible Python pipeline to compare direct answers with source-verified revisions — testing where a second pass improves grounding, and where it over-edits an answer that was already good.

Reader asks“What does this mean?”
Direct draft6.5/12

…After the assistant was removed, both groups completed three identical problems to measure their independent performance. The study aimed to assess how AI assistance affected motivation and persistence, as indicated by the choice to skip problems. The findings suggest that the presence of AI may influence how participants engage with the problems, but the study also acknowledges limitations in its design…

1 unsupported claim · 3 missing qualifications
After verification11.5/12

…The study aimed to assess how AI assistance affected motivation and persistence, as indicated by the choice to skip problems. The method supports causal comparison of AI access versus control within this experiment. However, the interpretation of skipping as persistence is an operationalization chosen by the authors, and the study acknowledges limitations in its design, particularly in Experiment 1…

Unsupported claim removed · qualifications restored
Real output from the experiment (case A1-G, gpt-4o-mini). Highlights mark what the verifier flagged and what it restored.
Time
Summer 2026
Role
Solo: experiment design, Python pipeline, evaluation
Stack
Python · OpenAI / Anthropic APIs · pydantic · pandas · pytest
Code
GitHub repo ↗

The problem

LLM explanations drift past their source

Graduate students increasingly ask an LLM to explain a dense passage before they read it closely. The answers sound right, but they can quietly add claims the source never made, drop the authors' qualifications, or turn a correlation into a cause — exactly what a time-pressed reader won't catch.

The approach

Draft first, then check against the source

Instead of asking for a better answer up front, the system lets the model answer directly, then runs a separate verifier that sees only the same source window and the exact draft. It either keeps the draft or makes the smallest possible repair. The question was whether that second pass earns its cost, and for which kinds of questions.

System

Five stages, one command

Every stage is a separate module with a structured output, so each step can be inspected, re-run, or swapped. The scorer never sees which workflow produced an answer.

Workflow A: direct draft goes straight to scoring Workflow B: same draft, verified Source + question Generate Verify Blind + judge Analyze fixed passage window grounded draft KEEP or minimal revise independent LLM, /12 tables + charts
  • Modelgpt-4o-mini, temperature 0
  • Structured outputspydantic schemas, JSON repair, retries
  • Mock moderuns end-to-end with no API key
  • Tests18 pytest checks on the pipeline

Experiment

Same passage, three ways of asking

How a reader phrases the question decides how much the model has to choose on its own. I held each passage fixed and varied only the question, then ran every version through both workflows: 10 passages × 3 question types × 2 workflows = 60 scored answers.

Generic

“What does this mean?”

The model decides what matters. Most room to drift.

Single concept

Names one idea

The question points at one claim in the passage.

Relational

Links two ideas

The answer has to connect two source claims correctly.

A design decision worth noting. My pilot added a planning step before answering. Planning didn't raise quality (9.8 vs 10.0 for direct), and it tangled planning and verification together, so I couldn't tell which one helped. The final design verifies the exact direct draft instead, so any change is attributable to verification alone.

Results

Verification helped most where the question was vaguest

Mean quality out of 12, scored blind. Descriptive results from a small exploratory study, not significance tests.

9.87 → 10.67mean quality, direct → verified
2 → 0boundary violations (unsupported, contradicted, overstated)
18 / 30drafts revised; 12 kept as-is
1regression: a strong draft made worse
View as table
Question typeDirectVerifiedΔ
Generic9.510.8+1.3
Single concept10.410.8+0.4
Relational9.710.4+0.7
All9.8710.67+0.80

Two cases

A repair, and an over-edit

A mean hides the interesting part. I tracked every change the verifier made, including the ones that hurt.

A1-G · Generic question

6.5 → 11.5 +5.0

The direct answer summarized the study correctly but softened a claim the source didn't support and dropped three of the authors' qualifications. The verifier listed each one against the passage and restored them with minimal edits (shown at the top of this page).

M4-R · Relational question

11.5 → 10.5 −1.0

The direct answer already connected both metrics cleanly. The verifier still revised it, adding true but unnecessary limitations, and the judge docked it for concision. Verification can add correct content and still make a good answer worse.

Recommendations

When to verify, and when not to

Each recommendation is tied to a specific result, not a general “always verify” rule.

Verify

Generic or under-specified questions

Biggest gain (+1.3); the model has the most room to drift.

Skip

Clear single-concept questions

Only +0.4; the direct answer is usually already right.

Verify with care

Relational questions

Removed both boundary violations, but produced the one regression.

Report it all

Not just the mean

Keeps, fixes, unnecessary edits, and regressions, side by side.

Limits. 10 passages from 3 papers, one model setting, one LLM judge without human agreement checks, and question variants I wrote myself. This measures answer quality, not student learning.

Run it yourself

No API key needed

Mock mode runs the whole pipeline end-to-end with canned model responses, so anyone can inspect every stage in about a minute.

git clone https://github.com/zichunxia78-creator/final-project-17630-xia
cd final-project-17630-xia && pip install -r requirements.txt
./run_pipeline.sh experiment --all --mock --fresh
./run_pipeline.sh test

The outcome

A reproducible evaluation that says not just whether a verifier helps, but for which questions, and at what cost.