Institution profile

Group 42

Industry researchasia · ae
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Intelligence Impact Quotient (IIQ): A Framework for Measuring Organizational AI Impact

May 14, 2026

Existing metrics—such as access counts or total token usage—struggle to accurately capture the depth of AI integration and its actual business value within organizations. This work proposes the Intelligence Impact Quotient (IIQ), a novel framework that integrates multiple dimensions—including semantic novelty, temporal decay, task complexity, organizational leverage, usage frequency, recency of activity, and autonomy—into a unified index. Through composite modeling, normalized mapping, and sub-daily update mechanisms, IIQ produces a standardized 0–1000 score that effectively distinguishes between high-frequency, low-value interactions and high-impact AI collaborations. The framework enables comparable evaluations across users and departments and demonstrates strong sensitivity and discriminative power across diverse usage patterns in synthetic validation scenarios.

0 citationsRead paper

Nomad: Autonomous Exploration and Discovery

Mar 31, 2026

Traditional query-driven research systems are constrained by predefined questions and struggle to proactively uncover novel insights. This work proposes the first closed-loop exploration–verification framework designed for autonomous discovery. By constructing a domain exploration map, the framework guides multi-tool agents—capable of retrieving information from documents, the web, and databases—to systematically traverse the knowledge space, balancing breadth and depth. An independent hypothesis validation mechanism coupled with automated report generation produces a primary report with citations alongside a meta-report. Experiments on corpora from the United Nations and the World Health Organization demonstrate that the proposed approach significantly outperforms baseline methods, yielding marked improvements in report credibility, quality, and diversity of insights.

0 citationsRead paper

Towards More Standardized AI Evaluation: From Models to Agents

Feb 20, 2026

This study addresses the limitations of traditional, static, model-centric evaluation methods in effectively assessing the behavioral reliability and trustworthiness of dynamic, tool-using agents. Moving beyond the prevailing paradigm centered on static benchmarks and aggregated scores, the work uncovers hidden failure modes in current evaluation practices and proposes a novel assessment framework tailored for non-deterministic agent systems. This framework emphasizes continuous, transparent monitoring of behavioral performance, reconceptualizing evaluation not as a one-time performance test but as an ongoing measurement discipline that supports trust formation, system iteration, and governance. The research demonstrates that high benchmark scores are often misleading and advocates for behavioral trustworthiness as a core metric, offering both theoretical foundations and practical pathways toward building trustworthy and governable agent systems.

0 citationsRead paper

BYOL: Bring Your Own Language Into LLMs

Jan 15, 2026

This work addresses the poor performance of large language models on low-resource and extremely low-resource languages, which stems from data scarcity and cultural misalignment. The authors propose the BYOL framework, which introduces a novel four-tier classification of languages based on their digital footprint. For low-resource languages, they design an end-to-end pipeline encompassing data cleaning, synthetic data generation, continued pretraining, and supervised fine-tuning. For extremely low-resource languages, they innovatively employ a translation-mediated adaptation pathway and integrate model weight fusion to balance local linguistic performance with multilingual generalization. Experiments demonstrate an average 12% performance gain on Chichewa and Māori and a 4-point BLEU improvement for Inuktitut translation. The study also releases the trilingual Global MMLU-Lite benchmark alongside open-sourced code and models.

0 citationsRead paper
Recent publications

Latest Papers

Intelligence Impact Quotient (IIQ): A Framework for Measuring Organizational AI Impact

May 14, 2026

Existing metrics—such as access counts or total token usage—struggle to accurately capture the depth of AI integration and its actual business value within organizations. This work proposes the Intelligence Impact Quotient (IIQ), a novel framework that integrates multiple dimensions—including semantic novelty, temporal decay, task complexity, organizational leverage, usage frequency, recency of activity, and autonomy—into a unified index. Through composite modeling, normalized mapping, and sub-daily update mechanisms, IIQ produces a standardized 0–1000 score that effectively distinguishes between high-frequency, low-value interactions and high-impact AI collaborations. The framework enables comparable evaluations across users and departments and demonstrates strong sensitivity and discriminative power across diverse usage patterns in synthetic validation scenarios.

0 citationsRead paper

Nomad: Autonomous Exploration and Discovery

Mar 31, 2026

Traditional query-driven research systems are constrained by predefined questions and struggle to proactively uncover novel insights. This work proposes the first closed-loop exploration–verification framework designed for autonomous discovery. By constructing a domain exploration map, the framework guides multi-tool agents—capable of retrieving information from documents, the web, and databases—to systematically traverse the knowledge space, balancing breadth and depth. An independent hypothesis validation mechanism coupled with automated report generation produces a primary report with citations alongside a meta-report. Experiments on corpora from the United Nations and the World Health Organization demonstrate that the proposed approach significantly outperforms baseline methods, yielding marked improvements in report credibility, quality, and diversity of insights.

0 citationsRead paper

Towards More Standardized AI Evaluation: From Models to Agents

Feb 20, 2026

This study addresses the limitations of traditional, static, model-centric evaluation methods in effectively assessing the behavioral reliability and trustworthiness of dynamic, tool-using agents. Moving beyond the prevailing paradigm centered on static benchmarks and aggregated scores, the work uncovers hidden failure modes in current evaluation practices and proposes a novel assessment framework tailored for non-deterministic agent systems. This framework emphasizes continuous, transparent monitoring of behavioral performance, reconceptualizing evaluation not as a one-time performance test but as an ongoing measurement discipline that supports trust formation, system iteration, and governance. The research demonstrates that high benchmark scores are often misleading and advocates for behavioral trustworthiness as a core metric, offering both theoretical foundations and practical pathways toward building trustworthy and governable agent systems.

0 citationsRead paper

BYOL: Bring Your Own Language Into LLMs

Jan 15, 2026

This work addresses the poor performance of large language models on low-resource and extremely low-resource languages, which stems from data scarcity and cultural misalignment. The authors propose the BYOL framework, which introduces a novel four-tier classification of languages based on their digital footprint. For low-resource languages, they design an end-to-end pipeline encompassing data cleaning, synthetic data generation, continued pretraining, and supervised fine-tuning. For extremely low-resource languages, they innovatively employ a translation-mediated adaptation pathway and integrate model weight fusion to balance local linguistic performance with multilingual generalization. Experiments demonstrate an average 12% performance gain on Chichewa and Māori and a 4-point BLEU improvement for Inuktitut translation. The study also releases the trilingual Global MMLU-Lite benchmark alongside open-sourced code and models.

0 citationsRead paper