01
Overview
The strongest evaluation tasks do not reward confident prose; they require a model to reconcile evidence correctly. I developed and reviewed task structures in medication reconciliation, trauma resuscitation, and water-quality data QA, each built around realistic source conflicts and an auditable expected answer.
02
Challenge
A useful rubric must distinguish an answer that merely sounds plausible from one that reconciles the files, computes the correct values, excludes bad evidence, follows the output contract, and surfaces unprompted failure modes. It also must separate major reasoning failures from minor presentation issues.
03
Source material
- Task prompts and domain-bounded role instructions
- Synthetic multi-file source packages in PDF, DOCX, CSV, XLSX and TXT formats
- Reference calculations and expected-answer guidance
- Two model responses per task plus evaluator write-ups
- Output-file and visible-response consistency checks
04
My work
- Defined the objective, decision boundaries, source hierarchy and required output structure for complex tasks.
- Embedded adversarial but realistic discrepancies that forced identity, timing, scope and calculation checks.
- Built reference answers around exact expected values rather than vague quality descriptions.
- Specified critical failure conditions, acceptable partial credit and lower-severity formatting deductions.
- Reviewed multiple model runs, independently recomputed key values, and compared saved artifacts against visible responses.
- Refined task guidance where a model exposed ambiguity or an unintended shortcut.
05
Analytical approach
Choose a real decisionAnchor the task in a concrete output—verify orders, reconcile a resuscitation, or classify a results ledger.
Build source tensionIntroduce mismatched identity, conflicting drafts, duplicate-like records, stale references or competing clocks.
Define a gold standardRecord the exact calculations, inclusions, exclusions and unsupported conclusions expected.
Weight failure modesMark the errors that invalidate the answer as critical and keep style or wrapper issues proportionate.
Evaluate the evaluatorRecompute values independently and check both the response and any required output file.
06
Key findings
Specific examples from the completed work.
Identity as a gating criterion
The medication task included a same-surname lab record with a different middle initial, transposed MRN digits and different birth date. The rubric required exclusion before any downstream reasoning.
Two-clock timeline trap
The trauma task separated wall-clock timestamps from offsets anchored at activation and crossed midnight. The gold standard required 41 minutes since activation, 50 since arrival and a nine-unit red-cell tally that excluded an issued-but-not-administered unit.
Authority and method traps
The water task combined an obsolete reference, a superseded draft, analyte-specific holding times, an out-of-scope row and a percentile-method conflict. The rubric rewarded the traceable row ledger, not a generic compliance summary.
Severity calibration
Evaluator notes treated a missed platelet-imbalance implication as a substantive but bounded deduction while treating one sentence outside the required root wrapper as presentation-only. In another run, rounding 0.0104 upward when the task required three-decimal reporting was caught because it changed the classification.
07
Deliverable
Sanitized task designs, rubric logic, gold-standard calculations, model-run evaluations, and QA notes. Proprietary interface language, hidden instructions, platform IDs and workflow mechanics are not reproduced.
08
Skills demonstrated
- Prompt design
- Rubric development
- Gold-standard authoring
- Adversarial task design
- Edge-case design
- Failure-mode analysis
- Evaluator QA
- Reference calculation
- Scoring calibration
Why it matters
Prompt engineering and AI training depend on tasks that expose meaningful model failures. This work shows that I can design for auditability, create objective grading anchors, and distinguish a flawed answer from a merely imperfect one.