Research questions
- Can hard-negative behavior be learned reliably?
- How does full fine-tuning affect grounding and structured output?
- Which gains remain stable across model families and sizes?
- How should base and tuned models be evaluated fairly?
A systematic research programme investigating whether small and medium language models can be trained for reliable, evidence-grounded RAG behavior.
Models from the Qwen, Gemma and Llama families are trained on a shared task format and compared with their base counterparts using identical test data and metrics.
The work emphasizes transparent datasets, reproducible settings and publication of model cards and evaluation results.
Released models are collected on Hugging Face; scripts and research material are maintained on GitHub.