Institution profile

Invertible AI

Industry researchnorthamerica · us
Official website
Research library7linked papers
Opportunities0open roles
Selected work

Representative Papers

Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches

Aug 11, 2025

This study addresses the severe phoneme distortions and high inter-speaker variability in dysarthric speech, which drastically degrade automatic speech recognition (ASR) performance. We propose a novel collaborative decoding paradigm integrating self-supervised speech models with large language models (LLMs). Specifically, we jointly leverage front-end acoustic models—including Wav2Vec 2.0, HuBERT, and Whisper—with both CTC and sequence-to-sequence outputs; backend constrained decoding is performed using LLMs (BART, GPT-2, and Vicuna) to restore phonemes and enforce syntactic and semantic consistency. Extensive experiments across multiple severity levels and cross-dataset benchmarks (e.g., UA-Speech, TORGO) demonstrate that LLM-augmented decoding significantly improves word error rate (WER) and intelligibility—achieving up to a 23.6% relative reduction in WER for severely dysarthric speech over baseline ASR systems—while markedly enhancing generalization and robustness. To our knowledge, this is the first systematic investigation validating the critical role of LLMs in joint semantic–acoustic modeling of dysarthric speech.

0 citationsRead paper

Beyond Content: How Grammatical Gender Shapes Visual Representation in Text-to-Image Models

Aug 05, 2025

Prior bias research in multimodal AI has largely overlooked the influence of linguistic structure—particularly grammatical gender—on visual representations generated by text-to-image (T2I) models. Method: We construct a cross-lingual benchmark covering five grammatical-gender languages and two grammatical-gender-neutral languages, and systematically evaluate three state-of-the-art T2I models, generating 28,800 images for quantitative analysis. Contribution/Results: We provide the first empirical evidence that grammatical gender induces systematic visual biases: masculine grammatical gender increases male representation to 73%, while feminine gender elevates female representation to 38%—both significantly deviating from the English gender-neutral baseline. We formally introduce “grammatical gender” as a critical new dimension of multimodal fairness, offering both theoretical grounding and empirical validation for how syntactic properties of language shape AI-generated visual content. This work bridges a longstanding gap in bias assessment by integrating linguistic typology into multimodal fairness evaluation.

0 citationsRead paper

Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset

May 28, 2025

Existing large vision-language models (LVLMs) exhibit pervasive cultural biases, particularly lacking robust Arabic cultural grounding due to the absence of high-quality, culturally rich Arabic multimodal data. Method: We introduce Pearl—the first large-scale, culture-aware Arabic multimodal instruction dataset—covering all 22 Arab states and ten major cultural domains. We propose a novel fine-grained Arabic cultural annotation framework; design the Pearl-X subset to quantify regional cultural variability; and employ agent-driven data generation, cross-regional collaborative annotation by 45 annotators, and a three-tier benchmark suite (Pearl, Pearl-Lite, Pearl-X). Contribution/Results: Empirical analysis demonstrates that instruction alignment significantly improves cultural grounding more than model scaling alone. Comprehensive evaluation across leading open- and closed-source multimodal LLMs shows substantial gains in cultural understanding and reasoning. All data, annotations, and benchmarks are publicly released.

0 citationsRead paper

Voice of a Continent: Mapping Africa's Speech Technology Frontier

May 24, 2025

African languages are severely underrepresented in speech technology, hindering digital inclusion. To address this, we systematically map the African speech technology ecosystem and introduce SimbaBench—the first multilingual, multitask benchmark covering 30+ African languages and 10+ speech tasks—alongside the Simba model series. Methodologically, we establish a unified evaluation framework tailored to African languages, revealing for the first time how data quality, domain diversity, and genetic language relatedness jointly influence cross-lingual performance; we further propose a standardized data curation pipeline and an efficient multilingual pretraining–adaptation paradigm. Results show that Simba models achieve state-of-the-art performance on ASR and TTS tasks across multiple African languages, significantly enhancing modeling capabilities for low-resource languages. This work advances a fairness-centered research paradigm in speech technology grounded in linguistic diversity.

0 citationsRead paper

Jawaher: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking

Feb 28, 2025

Current large language models (LLMs) exhibit significant biases in cross-cultural figurative language understanding—particularly for non-Western, dialectally diverse cultural expressions such as Arabic proverbs—due to insufficient cultural sensitivity and contextual adaptability. To address this, we introduce Jawaher, the first benchmark dataset covering multiple Arabic dialects, featuring four-layer annotations: original proverb, dialect identification, idiomatic translation, and expert-curated cultural explanations. Distinct from prior work, Jawaher is the first to systematically integrate multi-dialectal proverbs, idiomatic translation, and deep cultural contextualization, thereby filling a critical gap in non-English figurative language evaluation. Through comprehensive automated and human evaluation across both open- and closed-source LLMs, our experiments reveal that while current models generate accurate literal translations, they consistently fail to produce culturally appropriate or contextually grounded interpretations—highlighting the core bottleneck in cross-cultural figurative language comprehension.

0 citationsRead paper
Recent publications

Latest Papers

Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches

Aug 11, 2025

This study addresses the severe phoneme distortions and high inter-speaker variability in dysarthric speech, which drastically degrade automatic speech recognition (ASR) performance. We propose a novel collaborative decoding paradigm integrating self-supervised speech models with large language models (LLMs). Specifically, we jointly leverage front-end acoustic models—including Wav2Vec 2.0, HuBERT, and Whisper—with both CTC and sequence-to-sequence outputs; backend constrained decoding is performed using LLMs (BART, GPT-2, and Vicuna) to restore phonemes and enforce syntactic and semantic consistency. Extensive experiments across multiple severity levels and cross-dataset benchmarks (e.g., UA-Speech, TORGO) demonstrate that LLM-augmented decoding significantly improves word error rate (WER) and intelligibility—achieving up to a 23.6% relative reduction in WER for severely dysarthric speech over baseline ASR systems—while markedly enhancing generalization and robustness. To our knowledge, this is the first systematic investigation validating the critical role of LLMs in joint semantic–acoustic modeling of dysarthric speech.

0 citationsRead paper

Beyond Content: How Grammatical Gender Shapes Visual Representation in Text-to-Image Models

Aug 05, 2025

Prior bias research in multimodal AI has largely overlooked the influence of linguistic structure—particularly grammatical gender—on visual representations generated by text-to-image (T2I) models. Method: We construct a cross-lingual benchmark covering five grammatical-gender languages and two grammatical-gender-neutral languages, and systematically evaluate three state-of-the-art T2I models, generating 28,800 images for quantitative analysis. Contribution/Results: We provide the first empirical evidence that grammatical gender induces systematic visual biases: masculine grammatical gender increases male representation to 73%, while feminine gender elevates female representation to 38%—both significantly deviating from the English gender-neutral baseline. We formally introduce “grammatical gender” as a critical new dimension of multimodal fairness, offering both theoretical grounding and empirical validation for how syntactic properties of language shape AI-generated visual content. This work bridges a longstanding gap in bias assessment by integrating linguistic typology into multimodal fairness evaluation.

0 citationsRead paper

Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset

May 28, 2025

Existing large vision-language models (LVLMs) exhibit pervasive cultural biases, particularly lacking robust Arabic cultural grounding due to the absence of high-quality, culturally rich Arabic multimodal data. Method: We introduce Pearl—the first large-scale, culture-aware Arabic multimodal instruction dataset—covering all 22 Arab states and ten major cultural domains. We propose a novel fine-grained Arabic cultural annotation framework; design the Pearl-X subset to quantify regional cultural variability; and employ agent-driven data generation, cross-regional collaborative annotation by 45 annotators, and a three-tier benchmark suite (Pearl, Pearl-Lite, Pearl-X). Contribution/Results: Empirical analysis demonstrates that instruction alignment significantly improves cultural grounding more than model scaling alone. Comprehensive evaluation across leading open- and closed-source multimodal LLMs shows substantial gains in cultural understanding and reasoning. All data, annotations, and benchmarks are publicly released.

0 citationsRead paper

Voice of a Continent: Mapping Africa's Speech Technology Frontier

May 24, 2025

African languages are severely underrepresented in speech technology, hindering digital inclusion. To address this, we systematically map the African speech technology ecosystem and introduce SimbaBench—the first multilingual, multitask benchmark covering 30+ African languages and 10+ speech tasks—alongside the Simba model series. Methodologically, we establish a unified evaluation framework tailored to African languages, revealing for the first time how data quality, domain diversity, and genetic language relatedness jointly influence cross-lingual performance; we further propose a standardized data curation pipeline and an efficient multilingual pretraining–adaptation paradigm. Results show that Simba models achieve state-of-the-art performance on ASR and TTS tasks across multiple African languages, significantly enhancing modeling capabilities for low-resource languages. This work advances a fairness-centered research paradigm in speech technology grounded in linguistic diversity.

0 citationsRead paper

Jawaher: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking

Feb 28, 2025

Current large language models (LLMs) exhibit significant biases in cross-cultural figurative language understanding—particularly for non-Western, dialectally diverse cultural expressions such as Arabic proverbs—due to insufficient cultural sensitivity and contextual adaptability. To address this, we introduce Jawaher, the first benchmark dataset covering multiple Arabic dialects, featuring four-layer annotations: original proverb, dialect identification, idiomatic translation, and expert-curated cultural explanations. Distinct from prior work, Jawaher is the first to systematically integrate multi-dialectal proverbs, idiomatic translation, and deep cultural contextualization, thereby filling a critical gap in non-English figurative language evaluation. Through comprehensive automated and human evaluation across both open- and closed-source LLMs, our experiments reveal that while current models generate accurate literal translations, they consistently fail to produce culturally appropriate or contextually grounded interpretations—highlighting the core bottleneck in cross-cultural figurative language comprehension.

0 citationsRead paper