LLM Response Evaluation
Factuality, safety & instruction-following review
Compared candidate responses, separated grounded claims from unsupported assertions, weighted error severity, and wrote evidence-based ratings and rationales.
View case studyAI evaluation · data quality · analytical QA
Christopher LyversAI Quality & Data Analyst
I evaluate LLM outputs, design prompts and rubrics, reconcile messy multi-file records, verify calculations, and explain exactly why a conclusion is—or is not—supported.
Featured projects
Each case study names the files reviewed, the reasoning performed, the errors found, and the deliverable produced.
Factuality, safety & instruction-following review
Compared candidate responses, separated grounded claims from unsupported assertions, weighted error severity, and wrote evidence-based ratings and rationales.
View case study209-result reconciliation & decision-logic audit
Reconciled 209 laboratory results across 10 files, validated timing and QC rules, corrected percentile logic, and separated screening findings from formal decision points.
View case studyMulti-file prompts, rubrics & gold standards
Designed failure-revealing multi-file tasks, explicit scoring criteria, reference calculations, edge cases, and grading guidance across three complex domains.
View case studyAdditional work samples
Seven-file heart-failure record review
Matched records across care settings, excluded a wrong-patient lab, trended renal data, checked a derived clearance value, and classified 11 active orders.
View case studySeven-source midnight case
Rebuilt a crossing-midnight timeline, distinguished issued from administered products, verified ratios and thresholds, and recalculated weight- and rate-based values.
View case studyDenied-date analysis & record exclusion
Audited 14 indexed records across a 19-day denied range, excluded contaminated evidence, separated defensible dates from weak ones, and removed unsupported diagnoses.
View case studyENT referral timeline, findings & evidence gaps
Aligned a referral, medication list, imaging report and laboratory data; separated confirmed findings from suspected conditions; and preserved missing-evidence boundaries.
View case studyCapabilities
My strongest work sits where model evaluation and data quality meet: source files disagree, rules are easy to misapply, and a confident answer is not enough.
Factuality, hallucination detection, groundedness, safety, completeness, instruction following and severity-weighted scoring.
Prompts, rubrics, reference answers, edge cases, adversarial records, failure modes and evaluator QA.
Multi-file reconciliation, calculation verification, units, thresholds, percentiles, trends, ratios and timeline math.
Identity checks, duplicate and scope control, validity qualification, evidence lineage, exclusions and unsupported-claim detection.
Role alignment
About
I work through unfamiliar subject matter by making the evidence structure explicit: which source governs, what can be combined, what must be excluded, which calculation controls the decision, and where the available information stops.
Selected projects in this portfolio are adapted from completed AI evaluation and data-quality assignments. Some assignments used synthetic records, fictional organizations, and simulated real-world scenarios to test model reasoning, data reconciliation, and quality assurance. The analytical and evaluation work shown reflects work I performed. Internal platform instructions and identifiers have been removed.
Selected work