…After the assistant was removed, both groups completed three identical problems to measure their independent performance. The study aimed to assess how AI assistance affected motivation and persistence, as indicated by the choice to skip problems. The findings suggest that the presence of AI may influence how participants engage with the problems, but the study also acknowledges limitations in its design…
LLM evaluation pipeline
Source-Grounded Answer Verifier
When does verification actually make an LLM answer better?
I built a reproducible Python pipeline to compare direct answers with source-verified revisions — testing where a second pass improves grounding, and where it over-edits an answer that was already good.
…The study aimed to assess how AI assistance affected motivation and persistence, as indicated by the choice to skip problems. The method supports causal comparison of AI access versus control within this experiment. However, the interpretation of skipping as persistence is an operationalization chosen by the authors, and the study acknowledges limitations in its design, particularly in Experiment 1…
gpt-4o-mini). Highlights mark what the verifier flagged and what it restored.The problem
LLM explanations drift past their source
Graduate students increasingly ask an LLM to explain a dense passage before they read it closely. The answers sound right, but they can quietly add claims the source never made, drop the authors' qualifications, or turn a correlation into a cause — exactly what a time-pressed reader won't catch.
The approach
Draft first, then check against the source
Instead of asking for a better answer up front, the system lets the model answer directly, then runs a separate verifier that sees only the same source window and the exact draft. It either keeps the draft or makes the smallest possible repair. The question was whether that second pass earns its cost, and for which kinds of questions.
System
Five stages, one command
Every stage is a separate module with a structured output, so each step can be inspected, re-run, or swapped. The scorer never sees which workflow produced an answer.
- Model
gpt-4o-mini, temperature 0 - Structured outputspydantic schemas, JSON repair, retries
- Mock moderuns end-to-end with no API key
- Tests18 pytest checks on the pipeline
Experiment
Same passage, three ways of asking
How a reader phrases the question decides how much the model has to choose on its own. I held each passage fixed and varied only the question, then ran every version through both workflows: 10 passages × 3 question types × 2 workflows = 60 scored answers.
Generic
“What does this mean?”
The model decides what matters. Most room to drift.
Single concept
Names one idea
The question points at one claim in the passage.
Relational
Links two ideas
The answer has to connect two source claims correctly.
A design decision worth noting. My pilot added a planning step before answering. Planning didn't raise quality (9.8 vs 10.0 for direct), and it tangled planning and verification together, so I couldn't tell which one helped. The final design verifies the exact direct draft instead, so any change is attributable to verification alone.
Results
Verification helped most where the question was vaguest
Mean quality out of 12, scored blind. Descriptive results from a small exploratory study, not significance tests.
View as table
| Question type | Direct | Verified | Δ |
|---|---|---|---|
| Generic | 9.5 | 10.8 | +1.3 |
| Single concept | 10.4 | 10.8 | +0.4 |
| Relational | 9.7 | 10.4 | +0.7 |
| All | 9.87 | 10.67 | +0.80 |
Two cases
A repair, and an over-edit
A mean hides the interesting part. I tracked every change the verifier made, including the ones that hurt.
A1-G · Generic question
6.5 → 11.5 +5.0
The direct answer summarized the study correctly but softened a claim the source didn't support and dropped three of the authors' qualifications. The verifier listed each one against the passage and restored them with minimal edits (shown at the top of this page).
M4-R · Relational question
11.5 → 10.5 −1.0
The direct answer already connected both metrics cleanly. The verifier still revised it, adding true but unnecessary limitations, and the judge docked it for concision. Verification can add correct content and still make a good answer worse.
Recommendations
When to verify, and when not to
Each recommendation is tied to a specific result, not a general “always verify” rule.
Generic or under-specified questions
Biggest gain (+1.3); the model has the most room to drift.
Clear single-concept questions
Only +0.4; the direct answer is usually already right.
Relational questions
Removed both boundary violations, but produced the one regression.
Not just the mean
Keeps, fixes, unnecessary edits, and regressions, side by side.
Limits. 10 passages from 3 papers, one model setting, one LLM judge without human agreement checks, and question variants I wrote myself. This measures answer quality, not student learning.
Run it yourself
No API key needed
Mock mode runs the whole pipeline end-to-end with canned model responses, so anyone can inspect every stage in about a minute.
git clone https://github.com/zichunxia78-creator/final-project-17630-xia
cd final-project-17630-xia && pip install -r requirements.txt
./run_pipeline.sh experiment --all --mock --fresh
./run_pipeline.sh test
The outcome