Institution profile

CENTAI

Industry researchnorthamerica · us
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Auditing LLM Editorial Bias in News Media Exposure

Oct 31, 2025

This study addresses the underexplored issue of *agentic editorial bias*—systematic, implicit information curation—when large language models (LLMs) function as news gatekeepers. Method: We conduct the first systematic audit of four state-of-the-art LLMs (GPT-4o-Mini, Claude-3.7-Sonnet, Gemini-2.0-Flash) against Google News, employing a multi-layered algorithmic framework integrating topic-based querying, media outlet classification, ideological positioning, and factual accuracy assessment—rigorously validated across diverse prompting strategies and reliability benchmarks. Results: All LLMs exhibit statistically significant, robust ideological skew and uneven attention allocation: they amplify ideologically aligned outlets while suppressing others, yielding lower media diversity and narrower exposure sets than conventional news aggregators. Crucially, models differ markedly in directional bias. We introduce the concept of *agentic editorial policy* to formalize LLMs’ latent, systemic filtering mechanisms—revealing their emergent role as high-stakes news intermediaries with substantial information manipulation potential. This work provides foundational empirical evidence and a theoretical framework for LLM content governance.

0 citationsRead paper

Online Minimization of Polarization and Disagreement via Low-Rank Matrix Bandits

Oct 01, 2025

This paper studies online intervention in Friedkin-Johnsen opinion dynamics to minimize group polarization and disagreement under realistic conditions where agents’ innate opinions are unknown and only sequentially observable. The problem is formalized as a low-rank matrix bandit regret minimization task with scalar feedback—marking the first formulation of social opinion intervention as a low-rank matrix bandit learning problem. We propose a two-stage algorithm: first, estimating the low-dimensional subspace of innate opinions via low-rank approximation; second, deploying a linear bandit policy within this compressed subspace. We theoretically establish a cumulative regret bound of Õ(√T). Experiments demonstrate that our method significantly outperforms baseline linear bandit approaches in both regret performance and computational efficiency.

0 citationsRead paper

Learning Individual Behavior in Agent-Based Models with Graph Diffusion Networks

May 27, 2025

To address the challenge of jointly optimizing non-differentiable Agent-Based Models (ABMs) with real-world data, this paper introduces the first end-to-end differentiable ABM framework. Unlike conventional surrogate modeling approaches—which approximate only system-level outputs—our method innovatively learns local, decentralized behavioral rules for each agent, thereby fully preserving the bottom-up dynamical essence of ABMs. Technically, we integrate graph diffusion models—to capture stochasticity in agent behavior—with graph neural networks—to encode multi-agent interaction structures—enabling joint modeling of individual agent trajectories and emergent population-level patterns. Evaluated on the Schelling segregation and predator–prey models, our framework reduces extrapolation prediction error by 42% compared to state-of-the-art surrogate methods, demonstrating substantial improvements in both fidelity and generalization.

0 citationsRead paper

On the Inference of Sociodemographics on Reddit

Feb 07, 2025

This study systematically evaluates methods for inferring sociodemographic attributes (age, gender, political affiliation) of Reddit users. We construct a large-scale, self-reported dataset comprising over 850,000 labeled posts. We benchmark embedding-based models against probabilistic models on two tasks: attribute classification (measured by ROC AUC) and population-level proportion estimation (measured by MAE). Results show that a Naïve Bayes classifier with bag-of-words features significantly outperforms state-of-the-art embedding approaches—achieving up to a 19% absolute improvement in classification AUC and consistently sub-15% MAE for proportion estimation under large-sample conditions. We propose the CSS (Coverage, Soundness, Scalability) framework—a best-practice guideline for computational social science research—emphasizing interpretable modeling and rigorous evaluation. All code, data, and model weights are publicly released, establishing a reproducible, interpretable, and robust methodological benchmark for bridging online behavior with offline demographic characteristics.

0 citationsRead paper

Minimizing Polarization and Disagreement in the Friedkin-Johnsen Model with Unknown Innate Opinions

Jan 27, 2025

This paper addresses the robust optimization of opinion polarization and disagreement in the Friedkin-Johnsen model under the realistic constraint that individuals’ innate opinions are unknown. We propose a three-stage framework: “limited querying → global reconstruction → objective optimization.” First, we actively query a small fraction (5–10%) of nodes to obtain local observations; second, we reconstruct the latent innate opinions via matrix completion and regression; third, we minimize polarization and disagreement objectives using convex optimization and heuristic algorithms. This work constitutes the first systematic study of opinion control under partial observability of innate opinions, establishing rigorous theoretical bounds on error propagation to quantify how reconstruction inaccuracies affect optimization performance. Experiments on diverse synthetic and real-world networks demonstrate that our method achieves over 92% of the optimization performance of the full-information baseline—significantly outperforming existing benchmark strategies.

0 citationsRead paper
Recent publications

Latest Papers

Auditing LLM Editorial Bias in News Media Exposure

Oct 31, 2025

This study addresses the underexplored issue of *agentic editorial bias*—systematic, implicit information curation—when large language models (LLMs) function as news gatekeepers. Method: We conduct the first systematic audit of four state-of-the-art LLMs (GPT-4o-Mini, Claude-3.7-Sonnet, Gemini-2.0-Flash) against Google News, employing a multi-layered algorithmic framework integrating topic-based querying, media outlet classification, ideological positioning, and factual accuracy assessment—rigorously validated across diverse prompting strategies and reliability benchmarks. Results: All LLMs exhibit statistically significant, robust ideological skew and uneven attention allocation: they amplify ideologically aligned outlets while suppressing others, yielding lower media diversity and narrower exposure sets than conventional news aggregators. Crucially, models differ markedly in directional bias. We introduce the concept of *agentic editorial policy* to formalize LLMs’ latent, systemic filtering mechanisms—revealing their emergent role as high-stakes news intermediaries with substantial information manipulation potential. This work provides foundational empirical evidence and a theoretical framework for LLM content governance.

0 citationsRead paper

Online Minimization of Polarization and Disagreement via Low-Rank Matrix Bandits

Oct 01, 2025

This paper studies online intervention in Friedkin-Johnsen opinion dynamics to minimize group polarization and disagreement under realistic conditions where agents’ innate opinions are unknown and only sequentially observable. The problem is formalized as a low-rank matrix bandit regret minimization task with scalar feedback—marking the first formulation of social opinion intervention as a low-rank matrix bandit learning problem. We propose a two-stage algorithm: first, estimating the low-dimensional subspace of innate opinions via low-rank approximation; second, deploying a linear bandit policy within this compressed subspace. We theoretically establish a cumulative regret bound of Õ(√T). Experiments demonstrate that our method significantly outperforms baseline linear bandit approaches in both regret performance and computational efficiency.

0 citationsRead paper

Learning Individual Behavior in Agent-Based Models with Graph Diffusion Networks

May 27, 2025

To address the challenge of jointly optimizing non-differentiable Agent-Based Models (ABMs) with real-world data, this paper introduces the first end-to-end differentiable ABM framework. Unlike conventional surrogate modeling approaches—which approximate only system-level outputs—our method innovatively learns local, decentralized behavioral rules for each agent, thereby fully preserving the bottom-up dynamical essence of ABMs. Technically, we integrate graph diffusion models—to capture stochasticity in agent behavior—with graph neural networks—to encode multi-agent interaction structures—enabling joint modeling of individual agent trajectories and emergent population-level patterns. Evaluated on the Schelling segregation and predator–prey models, our framework reduces extrapolation prediction error by 42% compared to state-of-the-art surrogate methods, demonstrating substantial improvements in both fidelity and generalization.

0 citationsRead paper

On the Inference of Sociodemographics on Reddit

Feb 07, 2025

This study systematically evaluates methods for inferring sociodemographic attributes (age, gender, political affiliation) of Reddit users. We construct a large-scale, self-reported dataset comprising over 850,000 labeled posts. We benchmark embedding-based models against probabilistic models on two tasks: attribute classification (measured by ROC AUC) and population-level proportion estimation (measured by MAE). Results show that a Naïve Bayes classifier with bag-of-words features significantly outperforms state-of-the-art embedding approaches—achieving up to a 19% absolute improvement in classification AUC and consistently sub-15% MAE for proportion estimation under large-sample conditions. We propose the CSS (Coverage, Soundness, Scalability) framework—a best-practice guideline for computational social science research—emphasizing interpretable modeling and rigorous evaluation. All code, data, and model weights are publicly released, establishing a reproducible, interpretable, and robust methodological benchmark for bridging online behavior with offline demographic characteristics.

0 citationsRead paper

Minimizing Polarization and Disagreement in the Friedkin-Johnsen Model with Unknown Innate Opinions

Jan 27, 2025

This paper addresses the robust optimization of opinion polarization and disagreement in the Friedkin-Johnsen model under the realistic constraint that individuals’ innate opinions are unknown. We propose a three-stage framework: “limited querying → global reconstruction → objective optimization.” First, we actively query a small fraction (5–10%) of nodes to obtain local observations; second, we reconstruct the latent innate opinions via matrix completion and regression; third, we minimize polarization and disagreement objectives using convex optimization and heuristic algorithms. This work constitutes the first systematic study of opinion control under partial observability of innate opinions, establishing rigorous theoretical bounds on error propagation to quantify how reconstruction inaccuracies affect optimization performance. Experiments on diverse synthetic and real-world networks demonstrate that our method achieves over 92% of the optimization performance of the full-information baseline—significantly outperforming existing benchmark strategies.

0 citationsRead paper