What is measured
- Groundedness and answer correctness
- Context precision and recall
- Hard-negative handling
- Structured-output validity
- Citation precision and recall
Reliable systems require realistic test cases, negative examples and metrics that expose both improvements and regressions.
Base and fine-tuned models are compared on the same held-out examples. Aggregate scores are complemented by per-task analysis, malformed-output checks and manual inspection.