← All work samples

AI Evaluation / Model Quality Work Sample

LLM Response Evaluation

Factuality, safety & instruction-following review

Compared candidate responses, separated grounded claims from unsupported assertions, weighted error severity, and wrote evidence-based ratings and rationales.

Scope

10archived evaluation screenshots reviewed for this work sample
2–3candidate responses compared in representative tasks
3 & 5level overall-quality scales applied
Multiplecriteria: grounding, completeness, safety, honesty and fit

01

Overview

The assignment work required more than choosing a preferred answer. Each response had to be tested against the prompt, available context, factual support, safety expectations, and the seriousness of any failure. Archived evidence shows work across factual-claim review, groundedness, completeness, harmfulness, instruction following, pairwise comparison, three-response selection, and multi-level quality scoring.

02

Challenge

A polished response can still fail on the criterion that matters most. The core challenge was to distinguish cosmetic weaknesses from unsupported factual claims, broken instructions, or safety failures—and to explain why the distinction changed the final rating.

03

Source material

  • User prompts and candidate model responses
  • Provided context documents or task-contained evidence
  • Evaluation criteria covering groundedness, completeness, harmfulness and overall quality
  • Archived rating selections and written rationales

04

My work

  • Compared responses directly rather than grading each in isolation.
  • Isolated claims that required external verification and distinguished them from statements already supplied by the context.
  • Checked whether outputs followed explicit content, tone, length and format instructions.
  • Identified hallucinated, unsafe, dishonest or unsupported content and weighted it by severity.
  • Assigned ratings or rankings and wrote concise rationales tied to specific response evidence.

05

Analytical approach

01

Establish the task contractTranslate the prompt into observable requirements before judging style or preference.

02

Trace claims to evidenceSeparate context-supported statements from claims that need outside verification or have no support.

03

Classify the failureDistinguish factuality, grounding, completeness, instruction-following and safety errors.

04

Weight severityTreat a central unsafe or false claim as more important than several minor presentation strengths.

05

Write an auditable rationaleName the decisive evidence and explain why it changes the rating or ranking.

06

Key findings

Specific examples from the completed work.

Unsupported factual claim

A response stated that Saturn is the largest planet. I treated this as an externally verifiable factual claim and identified it as not grounded rather than accepting fluent wording at face value.

Constraint failure in creative output

For a sandwich-themed love poem, the evaluation required checking both sonnet structure and the requested balance of sincerity and whimsy—not merely whether the writing sounded pleasant.

Safety outweighed partial usefulness

In a mold-remediation comparison, one answer included useful moisture-control and protective-equipment advice but also suggested that household chemical treatment could permanently eliminate mold roots in porous drywall. I treated the central safety and judgment failure as decisive.

Audience fit mattered

For a monetary-policy explanation requested at a five-year-old level and under 400 words, I evaluated factual content together with accessibility and adherence to the requested audience.

07

Deliverable

Completed response selections, quality labels, criterion-level judgments, and written evidence-based rationales. Raw screenshots are not published because they contain internal interface details; the case study preserves the evaluation method and sanitized examples.

08

Skills demonstrated

  • LLM output evaluation
  • Factuality checking
  • Hallucination detection
  • Groundedness review
  • Instruction-following assessment
  • Safety judgment
  • Error severity classification
  • Pairwise ranking
  • Evaluation rationale writing

Why it matters

Model-quality roles require defensible judgment, not taste. This work shows how I turn broad prompts and messy outputs into explicit criteria, identify the failure that actually matters, and explain the decision in language another reviewer can audit.