Institution profile

National Institute of Mental Health

Academic institutionnorthamerica · us
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Revealing the core dimensions underlying representations in brains, behavior and AI

May 26, 2026

Current approaches struggle to extract interpretable low-dimensional representations from sparse or incomplete similarity data, limiting our understanding of representational structures in neural, behavioral, and artificial intelligence systems. This work proposes Similarity Representation Factorization (SRF), a novel method that integrates non-negative matrix factorization with low-dimensional embedding to enable, for the first time, generalizable and interpretable extraction of representational dimensions. SRF effectively recovers task-specific model dimensions, accurately predicts independent behavioral attributes, and substantially enhances both exploratory analysis capabilities and statistical power in hypothesis testing. The method is broadly applicable to heterogeneous, multi-source similarity data, offering a robust framework for uncovering latent structure across diverse domains.

0 citationsRead paper

The role of neuromorphic principles in the future of biomedicine and healthcare

Mar 29, 2026

This study addresses the challenges of applying neuromorphic engineering in biomedical and healthcare domains by convening, for the first time, a diverse coalition of stakeholders—including academia, industry, clinicians, and funding agencies—to systematically assess the current state and critical bottlenecks of the technology through interdisciplinary expert workshops. By integrating insights from neuroscience, biomedical engineering, and brain-inspired computing, the project delineates key development pathways and collaborative strategies, establishing a community-wide consensus. The outcomes—including workshop recordings, presentations, and strategic recommendations—have been publicly released to provide a clear roadmap and guiding framework for future research and development in the field.

0 citationsRead paper

Lost without translation -- Can transformer (language models) understand mood states?

Nov 28, 2025

This study investigates the capacity of large language models (LLMs) to semantically represent affective states—depression, euthymia, euphoric mania, and irritable mania—in Indian languages, exposing critical bottlenecks in cross-lingual mental health modeling. We evaluate native multilingual embeddings (IndicBERT, Sarvam-M) on emotion classification and clustering tasks, comparing them against translation-augmented pipelines: machine translation (Gemini) and human translation followed by English-to-Chinese embedding. Native Indian-language embeddings fail to distinguish affective states (clustering score: 0.002), whereas Gemini-translated embeddings achieve markedly improved performance (0.60), and human-translated English embeddings yield the highest score (0.67). Results indicate that current multilingual LLMs lack native proficiency in interpreting non-English affective expressions; high-fidelity translation is thus essential for robust cross-lingual mental state representation. This work provides the first systematic quantification of language-specific distress expression effects on LLM affective representations, establishing a methodological benchmark and empirical foundation for equitable, generalizable multilingual mental health AI.

0 citationsRead paper

Local Obfuscation by GLINER for Impartial Context Aware Lineage: Development and evaluation of PII Removal system

Oct 22, 2025

This study addresses the privacy-preserving de-identification of electronic health record (EHR) clinical text in resource-constrained environments. We propose a lightweight, localized, and context-aware PII anonymization method. Our approach fine-tunes the GLiNER model (modern-gliner-bi-large-v1.0) to support fine-grained recognition and precise substitution of nine sensitive entity types—without requiring GPU acceleration and enabling efficient execution on commodity laptops. Under rigorous character-level evaluation, our model achieves a micro-averaged F1 score of 0.980, substantially outperforming Azure NER, Microsoft Presidio, and zero-shot prompting of Gemini-Pro-2.5 and Llama-3.3-70B-Instruct. Notably, 95% of documents are fully and correctly anonymized (vs. a maximum of 64% for baselines), with only a 2% false-negative rate and minimal human review overhead. These advances facilitate practical “source-level de-identification” deployment in primary-care settings and AI-driven clinical research.

0 citationsRead paper
Recent publications

Latest Papers

Revealing the core dimensions underlying representations in brains, behavior and AI

May 26, 2026

Current approaches struggle to extract interpretable low-dimensional representations from sparse or incomplete similarity data, limiting our understanding of representational structures in neural, behavioral, and artificial intelligence systems. This work proposes Similarity Representation Factorization (SRF), a novel method that integrates non-negative matrix factorization with low-dimensional embedding to enable, for the first time, generalizable and interpretable extraction of representational dimensions. SRF effectively recovers task-specific model dimensions, accurately predicts independent behavioral attributes, and substantially enhances both exploratory analysis capabilities and statistical power in hypothesis testing. The method is broadly applicable to heterogeneous, multi-source similarity data, offering a robust framework for uncovering latent structure across diverse domains.

0 citationsRead paper

The role of neuromorphic principles in the future of biomedicine and healthcare

Mar 29, 2026

This study addresses the challenges of applying neuromorphic engineering in biomedical and healthcare domains by convening, for the first time, a diverse coalition of stakeholders—including academia, industry, clinicians, and funding agencies—to systematically assess the current state and critical bottlenecks of the technology through interdisciplinary expert workshops. By integrating insights from neuroscience, biomedical engineering, and brain-inspired computing, the project delineates key development pathways and collaborative strategies, establishing a community-wide consensus. The outcomes—including workshop recordings, presentations, and strategic recommendations—have been publicly released to provide a clear roadmap and guiding framework for future research and development in the field.

0 citationsRead paper

Lost without translation -- Can transformer (language models) understand mood states?

Nov 28, 2025

This study investigates the capacity of large language models (LLMs) to semantically represent affective states—depression, euthymia, euphoric mania, and irritable mania—in Indian languages, exposing critical bottlenecks in cross-lingual mental health modeling. We evaluate native multilingual embeddings (IndicBERT, Sarvam-M) on emotion classification and clustering tasks, comparing them against translation-augmented pipelines: machine translation (Gemini) and human translation followed by English-to-Chinese embedding. Native Indian-language embeddings fail to distinguish affective states (clustering score: 0.002), whereas Gemini-translated embeddings achieve markedly improved performance (0.60), and human-translated English embeddings yield the highest score (0.67). Results indicate that current multilingual LLMs lack native proficiency in interpreting non-English affective expressions; high-fidelity translation is thus essential for robust cross-lingual mental state representation. This work provides the first systematic quantification of language-specific distress expression effects on LLM affective representations, establishing a methodological benchmark and empirical foundation for equitable, generalizable multilingual mental health AI.

0 citationsRead paper

Local Obfuscation by GLINER for Impartial Context Aware Lineage: Development and evaluation of PII Removal system

Oct 22, 2025

This study addresses the privacy-preserving de-identification of electronic health record (EHR) clinical text in resource-constrained environments. We propose a lightweight, localized, and context-aware PII anonymization method. Our approach fine-tunes the GLiNER model (modern-gliner-bi-large-v1.0) to support fine-grained recognition and precise substitution of nine sensitive entity types—without requiring GPU acceleration and enabling efficient execution on commodity laptops. Under rigorous character-level evaluation, our model achieves a micro-averaged F1 score of 0.980, substantially outperforming Azure NER, Microsoft Presidio, and zero-shot prompting of Gemini-Pro-2.5 and Llama-3.3-70B-Instruct. Notably, 95% of documents are fully and correctly anonymized (vs. a maximum of 64% for baselines), with only a 2% false-negative rate and minimal human review overhead. These advances facilitate practical “source-level de-identification” deployment in primary-care settings and AI-driven clinical research.

0 citationsRead paper