Institution profile

NeuralMind

Industry research
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Legal Nugget Extraction for Granular Retrieval over Long Jurisprudential Texts

Jul 24, 2026

Judicial case documents are typically lengthy, with critical legal arguments constituting only a small fraction of the text, which limits the effectiveness of conventional full-text retrieval. This work proposes a fine-grained retrieval framework that, for the first time, integrates legal nugget extraction with dense retrieval. The approach first automatically extracts self-contained legal nuggets from case documents, constructs an embedding index at the nugget level, and performs dense retrieval over these nuggets; results are then aggregated to produce document-level rankings. Evaluated on the JUA-Juris and JurisTCU datasets, the method achieves NDCG@10 scores of 0.20461 and 0.32696, respectively, significantly outperforming baseline approaches and enabling precise case retrieval tailored to legal argument queries.

0 citationsRead paper

Domain-Adaptive Dense Retrieval for Brazilian Legal Search

May 05, 2026

This study addresses the challenge of balancing domain-specific expertise and cross-task generalization in Brazilian legal retrieval, where heterogeneous tasks—such as case law search, legislative document retrieval, and question answering—demand both legal precision and broader adaptability. To tackle this, the authors propose a hybrid fine-tuning strategy based on Qwen3-Embedding-4B, integrating legal corpora with Portuguese SQuAD-pt question-answering data. This approach preserves strong performance on legal tasks while substantially enhancing generalization for question-based retrieval. Experimental results across six Brazilian legal datasets demonstrate an average NDCG@10 of 0.447, MRR@10 of 0.595, and MAP@10 of 0.308, with particularly significant gains on the Quati question-answering benchmark compared to models fine-tuned exclusively on legal data.

0 citationsRead paper

Aplicac{c}~ao de Large Language Models na An'alise e S'intese de Documentos Jur'idicos: Uma Revis~ao de Literatura

Apr 01, 2025

Large language models (LLMs) exhibit significant potential for legal text analysis and generation, yet their reliability in law-specific tasks remains hindered by hallucinations, domain bias, and inconsistent performance. Method: This paper systematically reviews prompt engineering techniques for LLMs—including GPT-4, BERT, Llama 2, and Legal-Pegasus—evaluating few-shot, zero-shot, and chain-of-thought prompting across legal summarization, classification, and retrieval tasks. Contribution/Results: Empirical results demonstrate strong performance of multiple LLMs on structured legal tasks; however, model bias and factual hallucination persist as critical barriers to real-world deployment. To address these challenges, the paper proposes a legal-domain-oriented prompt engineering framework emphasizing domain-adaptive prompt design and rigorous trustworthiness verification. The framework provides both methodological guidance and empirical evidence to support robust, accountable LLM integration into judicial practice.

0 citationsRead paper
Recent publications

Latest Papers

Legal Nugget Extraction for Granular Retrieval over Long Jurisprudential Texts

Jul 24, 2026

Judicial case documents are typically lengthy, with critical legal arguments constituting only a small fraction of the text, which limits the effectiveness of conventional full-text retrieval. This work proposes a fine-grained retrieval framework that, for the first time, integrates legal nugget extraction with dense retrieval. The approach first automatically extracts self-contained legal nuggets from case documents, constructs an embedding index at the nugget level, and performs dense retrieval over these nuggets; results are then aggregated to produce document-level rankings. Evaluated on the JUA-Juris and JurisTCU datasets, the method achieves NDCG@10 scores of 0.20461 and 0.32696, respectively, significantly outperforming baseline approaches and enabling precise case retrieval tailored to legal argument queries.

0 citationsRead paper

Domain-Adaptive Dense Retrieval for Brazilian Legal Search

May 05, 2026

This study addresses the challenge of balancing domain-specific expertise and cross-task generalization in Brazilian legal retrieval, where heterogeneous tasks—such as case law search, legislative document retrieval, and question answering—demand both legal precision and broader adaptability. To tackle this, the authors propose a hybrid fine-tuning strategy based on Qwen3-Embedding-4B, integrating legal corpora with Portuguese SQuAD-pt question-answering data. This approach preserves strong performance on legal tasks while substantially enhancing generalization for question-based retrieval. Experimental results across six Brazilian legal datasets demonstrate an average NDCG@10 of 0.447, MRR@10 of 0.595, and MAP@10 of 0.308, with particularly significant gains on the Quati question-answering benchmark compared to models fine-tuned exclusively on legal data.

0 citationsRead paper

Aplicac{c}~ao de Large Language Models na An'alise e S'intese de Documentos Jur'idicos: Uma Revis~ao de Literatura

Apr 01, 2025

Large language models (LLMs) exhibit significant potential for legal text analysis and generation, yet their reliability in law-specific tasks remains hindered by hallucinations, domain bias, and inconsistent performance. Method: This paper systematically reviews prompt engineering techniques for LLMs—including GPT-4, BERT, Llama 2, and Legal-Pegasus—evaluating few-shot, zero-shot, and chain-of-thought prompting across legal summarization, classification, and retrieval tasks. Contribution/Results: Empirical results demonstrate strong performance of multiple LLMs on structured legal tasks; however, model bias and factual hallucination persist as critical barriers to real-world deployment. To address these challenges, the paper proposes a legal-domain-oriented prompt engineering framework emphasizing domain-adaptive prompt design and rigorous trustworthiness verification. The framework provides both methodological guidance and empirical evidence to support robust, accountable LLM integration into judicial practice.

0 citationsRead paper