PILOT-Bench: A Benchmark for Legal Reasoning in the Patent Domain with IRAC-Aligned Classification Tasks
This study addresses the absence of a systematic evaluation benchmark for large language models (LLMs) in the domain of patent law reasoning. The authors construct the first benchmark centered on decisions from the U.S. Patent Trial and Appeal Board (PTAB), aligning PTAB rulings with USPTO patent data to formulate three structured classification tasks grounded in the IRAC legal analysis framework: issue type, cited authority, and sub-decision. The benchmark enables multidimensional evaluation across input variations, model families, and error analyses, offering a comprehensive assessment of both open- and closed-source LLMs. Experimental results reveal a substantial performance gap: the best closed-source model achieves a Micro-F1 score of 0.75 on the issue-type task, whereas the strongest open-source model, Qwen-8B, attains only 0.56, highlighting significant limitations in current models’ capacity for patent-related legal reasoning.