| 1. Claim Identification |
All claims correctly identified and separated |
Most claims identified correctly |
Some claims identified; 2 missed |
Few or no claims identified |
| 2. Source Quality |
Independent, credible; SOURCE test justified |
Most sources credible and independent |
Relies on a single source for multiple claims |
Sources absent, unreliable, or AI verifies AI |
| 3. Verification Accuracy |
All verdicts accurate and fully justified |
Most verdicts accurate |
Some verdicts accurate; skips hard claims |
Verdicts mostly absent or all marked ✅ |
| 4. Category Labels |
All claims correctly labeled F/O/I, incl. plausible inventions |
Most labels correct |
Struggles to distinguish Opinion from Fact |
Labels absent or mostly incorrect |
| 5. Reliability Verdict |
Clearly reasoned, directly supported by evidence |
Appropriate rating; relies on 1–2 pieces of evidence |
Vague or inconsistent with the Verification Log |
Absent or contradicts the evidence |
| 6. Verification Decision |
Clear, specific decision with a concrete change |
Reasonable decision; change mentioned but not explained |
Superficial ("I would check it first") |
Absent, or would share without verification |
| 7. Reflection Depth |
Genuine self-assessment; specific behavior change named |
Present and honest; not fully reasoned |
Surface-level ("I will check more") |
Absent or restates the task instructions |