Eval Forge

Task and AI output in — verdict, scorecard, rubric, findings, refined output out.

Back to SkillSafe
Or pick files: read locally, nothing uploads until you run.
Your rubric — one criterion per line (optional; derived from the task if empty)
The rubric you last ran with is remembered on this device, so a standing rubric never needs re-pasting.
Reference — ground truth to check against (optional)
How it works

Nothing to hand? Load the — a weekly update with an unsupported metric, an empty risks section and chat boilerplate — the , where the function looks right but drops a requirement and mishandles an edge case, or the , which should come back ship.

1

Paste the task and the output

The prompt, ticket or brief the AI was answering, and the output it produced — a report, generated code, a drafted email, JSON. Add your own rubric (one criterion per line) if you have one, and a reference (source material, expected result, spec) if grounding matters. The instant prescan reads the output for free while you type and lists what it mechanically found: TODO and placeholder markers, unresolved template variables, assistant boilerplate, broken JSON, unclosed code fences, empty sections, repeated paragraphs, truncated endings, hedge clusters and unverifiable citations — plus the requirement-shaped lines it extracted from the task.

2

The AI evaluates it

A ship/revise/rework verdict with a 0–100 weighted score; a five-dimension scorecard across Instruction adherence, Correctness & grounding, Completeness & coverage, Clarity & structure and Safety & honesty, each scored 1–5 with a note pointing at the paste; your rubric scored criterion by criterion — or 5–8 criteria derived from the task when you brought none; and severity-ranked findings, each quoting the sentence, claim, code or field it concerns, with corrected text where a mechanical fix exists. Every prescan hit is confirmed or explicitly set aside — including the false positives, like a TODO the task itself asked to leave in.

3

Take the refined output and go

The output rewritten to actually satisfy the task — every high and medium finding fixed, failed rubric criteria met, fabricated specifics replaced with what the reference supports or an honest marked gap — as a download or a unified diff against the paste that git apply accepts, plus the ordered next steps, the findings as CSV, a review-thread comment, and evaluation history kept on your account with restore. The prescan re-reads the refinement for free and says which of the smells it matched are actually gone; paste your agent's next attempt (or the refinement) back into the box and it keeps scoring against the last evaluation, so you only pay for a second run once it is worth running. Markdown and JSON export throughout.

Derived from the @github/agentic-eval skill (MIT).