[https://drive.google.com/file/d/1-7IOw2E2fNlaJOXZH8lunZ4YP3qZSSbI/view?usp=sharing](https://drive.google.com/file/d/1-7IOw2E2fNlaJOXZH8lunZ4YP3qZSSbI/view?usp=sharing)
Reliable legal AI remains underexplored in academia despite its importance for high-impact domains such as law and medicine. Current language models often hallucinate, over-rely on memorised knowledge, and struggle to determine when they have enough information to answer, particularly in retrieval-augmented generation (RAG) systems that depend on external evidence. Reward models and LLM judges provide scalable supervision signals for reinforcement learning and agentic systems, enabling models to become more grounded and context-aware without constant human feedback. This talk explores how lightweight legal judge models can be built using small, open-source language models, and presents a pipeline for converting legal question answering datasets into reward modelling benchmarks that evaluate answerability and faithfulness to context under realistic academic and startup compute constraints.
Dr Fabio J. Fehr is a postdoctoral researcher at the University of Oxford working on AI for legal reasoning, agentic workflows, and large language models. His background spans statistical modelling at the University of Cape Town, theoretical machine learning and NLP at EPFL and Idiap, and industry experience at Amazon working in agentic code generation. At Oxford, he is part of the OXAI committee and leads a research team working on reward modelling and reinforcement learning for high-stakes applications in law and clinical medicine. He is also co-organising the inaugural ICML AI4Law workshop in 2026. His broader research focuses on building socially grounded AI systems that are both theoretically principled and practically useful.
.png)