🤖 AI Summary
Large language models (LLMs) exhibit significant potential for legal text analysis and generation, yet their reliability in law-specific tasks remains hindered by hallucinations, domain bias, and inconsistent performance. Method: This paper systematically reviews prompt engineering techniques for LLMs—including GPT-4, BERT, Llama 2, and Legal-Pegasus—evaluating few-shot, zero-shot, and chain-of-thought prompting across legal summarization, classification, and retrieval tasks. Contribution/Results: Empirical results demonstrate strong performance of multiple LLMs on structured legal tasks; however, model bias and factual hallucination persist as critical barriers to real-world deployment. To address these challenges, the paper proposes a legal-domain-oriented prompt engineering framework emphasizing domain-adaptive prompt design and rigorous trustworthiness verification. The framework provides both methodological guidance and empirical evidence to support robust, accountable LLM integration into judicial practice.
📝 Abstract
Large Language Models (LLMs) have been increasingly used to optimize the analysis and synthesis of legal documents, enabling the automation of tasks such as summarization, classification, and retrieval of legal information. This study aims to conduct a systematic literature review to identify the state of the art in prompt engineering applied to LLMs in the legal context. The results indicate that models such as GPT-4, BERT, Llama 2, and Legal-Pegasus are widely employed in the legal field, and techniques such as Few-shot Learning, Zero-shot Learning, and Chain-of-Thought prompting have proven effective in improving the interpretation of legal texts. However, challenges such as biases in models and hallucinations still hinder their large-scale implementation. It is concluded that, despite the great potential of LLMs for the legal field, there is a need to improve prompt engineering strategies to ensure greater accuracy and reliability in the generated results.