LaMSUM: Amplifying Voices Against Harassment through LLM Guided Extractive Summarization of User Incident Reports
To address the challenge of manually reviewing large-scale, code-mixed sexual harassment reports in India’s Safe City platform, this paper proposes the first LLM-driven extractive summarization framework tailored to this domain. Methodologically, it introduces a multi-model collaborative architecture integrating Llama, Mistral, and GPT-4o, enhanced by hierarchical text segmentation, prompt-engineered fine-grained extraction decisions, and an ensemble voting mechanism—effectively mitigating LLMs’ abstraction bias and context window limitations. Contributions include: (1) the first explainable and traceable extractive summarization system for code-mixed harassment reports; (2) state-of-the-art performance on the Safe City dataset, significantly outperforming existing baselines; and (3) generation of high-fidelity, structured event overviews that directly inform evidence-based policymaking and targeted anti-harassment interventions.