[https://docs.google.com/presentation/d/1me6Ir0WTAiskTtRKGZSbirJPpkWUN2-E/edit?usp=sharing&ouid=100090609942785485941&rtpof=true&sd=true](https://docs.google.com/presentation/d/1me6Ir0WTAiskTtRKGZSbirJPpkWUN2-E/preview?usp=sharing&ouid=100090609942785485941&rtpof=true&sd=true)
Researchers and benchmarking platforms are often confronted with the question of whether rubric- or preference-based evaluations are better suited to capture "quality" in legal work. In collaboration with various law firms, this research empirically addresses this question by comparing these evaluation methods across a set of real-world legal tasks. We generate legal work product of varying quality using LLMs and test each method’s ability to reliably recover that quality.
Pierce Kelaita is a Research Scientist and Head of Architecture at Stanford Law School’s Legal Innovation Through Frontier Technology Lab (liftlab). His work focuses on computational models of subjectivity and decision-making, and he leads both fundamental and applied research initiatives in this area. Prior to joining Stanford, he was a software investor at FPV Ventures and previously a software engineer at Meta. Outside of work, he co-founded the AI Collective, a globally recognized professional network, where he remains an occasional advisor. Pierce holds a B.S. in Computer Science from the University of Washington.
Russell is an AI Engineering Fellow at Stanford Law School, focusing on research and product development in collaboration with law firms and legal tech companies. He received a BS in Electrical Engineering and Computer Science and Molecular Biophysics and Biochemistry from Yale.

