Research topic

Evaluation beyond impressive demos

Reliable systems require realistic test cases, negative examples and metrics that expose both improvements and regressions.

What is measured

  • Groundedness and answer correctness
  • Context precision and recall
  • Hard-negative handling
  • Structured-output validity
  • Citation precision and recall

Evaluation principle

Base and fine-tuned models are compared on the same held-out examples. Aggregate scores are complemented by per-task analysis, malformed-output checks and manual inspection.