Measuring AI Leadership: Development and Validation of a Multidimensional Measure for AI-Native Organizations
研究开发了AI领导力电池,通过多维度衡量AI原生组织中的领导行为,解决了现有方法无法准确评估AI环境下领导力的问题。
研究开发了AI领导力电池,通过多维度衡量AI原生组织中的领导行为,解决了现有方法无法准确评估AI环境下领导力的问题。
Existing disparity filters for backbone extraction in weighted directed networks neglect hierarchical dependencies among nodes. To address this, we propose the “Difference-of-Differences” (D₂) filtering method, which explicitly models relative dependencies and hierarchical structure by normalizing each edge’s weight against its expected disparity under a null model. D₂ integrates weighted directed network modeling, expectation-based disparity estimation, and statistical significance testing to yield a hierarchy-aware backbone extraction algorithm. Experiments on four real-world networks—academic journal citations, airport flight routes, corporate email communications, and international trade flows—demonstrate that D₂ more accurately identifies empirically grounded hierarchical organization and core–periphery structures. Quantitatively, it significantly outperforms conventional disparity filters in both structural fidelity and interpretability, while preserving directional and weighted characteristics of the original networks.
This study investigates whether multimodal large language models (MLLMs) can match or surpass human experts in emotion recognition, and evaluates the performance gains from human-AI collaboration and collective intelligence. Using the standardized Reading the Mind in the Eyes Test (RMET) and its multi-ethnic extension (MRMET), we benchmark individual MLLMs, individual humans, human groups (aggregating independent judgments), and human-AI collaborative frameworks. Results show that individual MLLMs significantly outperform individual humans in accuracy; however, human group decisions consistently exceed all single-model baselines. A novel human-AI co-reasoning framework—integrating model-generated explanations with human group consensus—achieves the highest overall accuracy. This work is the first to systematically demonstrate the superiority of collective intelligence in affective understanding, introducing the “augmented intelligence” paradigm. It provides both theoretical foundations and practical design principles for developing trustworthy, human-aligned affective AI systems.
In organizational research, built-in LLM moderation mechanisms often over-censor responses to microaggressions and hate speech—either refusing to answer or diluting harmful content—thereby compromising analytical validity. To address this, we propose an Elo-based dynamic evaluation framework, the first to adapt the Elo rating system for LLM-based harmful content detection. Our method comprises dual-stage annotation validation, task-specific fine-tuning, and real-time response quality calibration—overcoming key limitations of static classification models and conventional prompt engineering. Evaluated on benchmark datasets for microaggressions and hate speech, our approach significantly improves accuracy, precision, and F1-score while substantially reducing false positive rates. It enables high-fidelity, scalable toxicity analysis and real-time assessment of workplace textual data.
This study investigates the persistent impact of the COVID-19 pandemic on birth rates and cesarean delivery (CD) rates in New York State, along with underlying economic drivers. Using statewide inpatient obstetric delivery data from 2012–2022, we employ fixed-effects regression and statistical analysis. Results show a sharp, non-reverting decline in birth rates—7.61 percentage points below pre-pandemic trends—with vaginal deliveries falling disproportionately more than CDs, thereby increasing the CD share. Crucially, we document for the first time the irreversibility of this fertility decline. We further propose and substantiate a novel hypothesis: financial incentives may bias clinical decision-making toward CDs, as hospitals generate 61% higher per-case revenue from CDs than from vaginal deliveries. These findings provide critical empirical evidence and a theoretical framework for understanding how public health crises reshape reproductive behavior and distort clinical practice through economic mechanisms.
研究开发了AI领导力电池,通过多维度衡量AI原生组织中的领导行为,解决了现有方法无法准确评估AI环境下领导力的问题。
Existing disparity filters for backbone extraction in weighted directed networks neglect hierarchical dependencies among nodes. To address this, we propose the “Difference-of-Differences” (D₂) filtering method, which explicitly models relative dependencies and hierarchical structure by normalizing each edge’s weight against its expected disparity under a null model. D₂ integrates weighted directed network modeling, expectation-based disparity estimation, and statistical significance testing to yield a hierarchy-aware backbone extraction algorithm. Experiments on four real-world networks—academic journal citations, airport flight routes, corporate email communications, and international trade flows—demonstrate that D₂ more accurately identifies empirically grounded hierarchical organization and core–periphery structures. Quantitatively, it significantly outperforms conventional disparity filters in both structural fidelity and interpretability, while preserving directional and weighted characteristics of the original networks.
This study investigates whether multimodal large language models (MLLMs) can match or surpass human experts in emotion recognition, and evaluates the performance gains from human-AI collaboration and collective intelligence. Using the standardized Reading the Mind in the Eyes Test (RMET) and its multi-ethnic extension (MRMET), we benchmark individual MLLMs, individual humans, human groups (aggregating independent judgments), and human-AI collaborative frameworks. Results show that individual MLLMs significantly outperform individual humans in accuracy; however, human group decisions consistently exceed all single-model baselines. A novel human-AI co-reasoning framework—integrating model-generated explanations with human group consensus—achieves the highest overall accuracy. This work is the first to systematically demonstrate the superiority of collective intelligence in affective understanding, introducing the “augmented intelligence” paradigm. It provides both theoretical foundations and practical design principles for developing trustworthy, human-aligned affective AI systems.
In organizational research, built-in LLM moderation mechanisms often over-censor responses to microaggressions and hate speech—either refusing to answer or diluting harmful content—thereby compromising analytical validity. To address this, we propose an Elo-based dynamic evaluation framework, the first to adapt the Elo rating system for LLM-based harmful content detection. Our method comprises dual-stage annotation validation, task-specific fine-tuning, and real-time response quality calibration—overcoming key limitations of static classification models and conventional prompt engineering. Evaluated on benchmark datasets for microaggressions and hate speech, our approach significantly improves accuracy, precision, and F1-score while substantially reducing false positive rates. It enables high-fidelity, scalable toxicity analysis and real-time assessment of workplace textual data.
This study investigates the persistent impact of the COVID-19 pandemic on birth rates and cesarean delivery (CD) rates in New York State, along with underlying economic drivers. Using statewide inpatient obstetric delivery data from 2012–2022, we employ fixed-effects regression and statistical analysis. Results show a sharp, non-reverting decline in birth rates—7.61 percentage points below pre-pandemic trends—with vaginal deliveries falling disproportionately more than CDs, thereby increasing the CD share. Crucially, we document for the first time the irreversibility of this fertility decline. We further propose and substantiate a novel hypothesis: financial incentives may bias clinical decision-making toward CDs, as hospitals generate 61% higher per-case revenue from CDs than from vaginal deliveries. These findings provide critical empirical evidence and a theoretical framework for understanding how public health crises reshape reproductive behavior and distort clinical practice through economic mechanisms.