01
Overview
The assignment work required more than choosing a preferred answer. Each response had to be tested against the prompt, available context, factual support, safety expectations, and the seriousness of any failure. Archived evidence shows work across factual-claim review, groundedness, completeness, harmfulness, instruction following, pairwise comparison, three-response selection, and multi-level quality scoring.
02
Challenge
A polished response can still fail on the criterion that matters most. The core challenge was to distinguish cosmetic weaknesses from unsupported factual claims, broken instructions, or safety failures—and to explain why the distinction changed the final rating.
03
Source material
- User prompts and candidate model responses
- Provided context documents or task-contained evidence
- Evaluation criteria covering groundedness, completeness, harmfulness and overall quality
- Archived rating selections and written rationales
04
My work
- Compared responses directly rather than grading each in isolation.
- Isolated claims that required external verification and distinguished them from statements already supplied by the context.
- Checked whether outputs followed explicit content, tone, length and format instructions.
- Identified hallucinated, unsafe, dishonest or unsupported content and weighted it by severity.
- Assigned ratings or rankings and wrote concise rationales tied to specific response evidence.
05
Analytical approach
Establish the task contractTranslate the prompt into observable requirements before judging style or preference.
Trace claims to evidenceSeparate context-supported statements from claims that need outside verification or have no support.
Classify the failureDistinguish factuality, grounding, completeness, instruction-following and safety errors.
Weight severityTreat a central unsafe or false claim as more important than several minor presentation strengths.
Write an auditable rationaleName the decisive evidence and explain why it changes the rating or ranking.
06
Key findings
Specific examples from the completed work.
Unsupported factual claim
A response stated that Saturn is the largest planet. I treated this as an externally verifiable factual claim and identified it as not grounded rather than accepting fluent wording at face value.
Constraint failure in creative output
For a sandwich-themed love poem, the evaluation required checking both sonnet structure and the requested balance of sincerity and whimsy—not merely whether the writing sounded pleasant.
Safety outweighed partial usefulness
In a mold-remediation comparison, one answer included useful moisture-control and protective-equipment advice but also suggested that household chemical treatment could permanently eliminate mold roots in porous drywall. I treated the central safety and judgment failure as decisive.
Audience fit mattered
For a monetary-policy explanation requested at a five-year-old level and under 400 words, I evaluated factual content together with accessibility and adherence to the requested audience.
07
Deliverable
Completed response selections, quality labels, criterion-level judgments, and written evidence-based rationales. Raw screenshots are not published because they contain internal interface details; the case study preserves the evaluation method and sanitized examples.
08
Skills demonstrated
- LLM output evaluation
- Factuality checking
- Hallucination detection
- Groundedness review
- Instruction-following assessment
- Safety judgment
- Error severity classification
- Pairwise ranking
- Evaluation rationale writing
Why it matters
Model-quality roles require defensible judgment, not taste. This work shows how I turn broad prompts and messy outputs into explicit criteria, identify the failure that actually matters, and explain the decision in language another reviewer can audit.