Institution profile

Mizuho–DL Financial Technology Co., Ltd.

Industry researchasia · jp
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Which Algorithm Specification Formats Help Language Models Implement Machine Learning Algorithms?

Jul 03, 2026

This study investigates how the format of algorithmic descriptions influences the accuracy of machine learning algorithm implementations generated by large language models (LLMs), with a focus on critical yet often implicit details such as interfaces, computational steps, numerical rules, and boundary behaviors. Through controlled experiments across multiple models (GPT-4o mini, Gemma 2 27B, Llama 3.2 3B), five ML algorithms, and seven description formats—including LaTeX pseudocode, YAML, and Python code stubs—the authors evaluate implementation correctness using fine-grained hidden tests. Results reveal that content clarity outweighs format per se: under core information conditions, LaTeX pseudocode yields the best performance, followed by YAML and plain text; with complete information, some models become format-agnostic, and code stubs offer no significant advantage. This work is the first to systematically demonstrate the pivotal role of algorithmic description readability in LLM-based implementation and offers writing guidelines optimized for AI interpretability.

0 citationsRead paper

Causality Elicitation from Large Language Models

Mar 04, 2026

This work proposes the first systematic framework for treating large language models (LLMs) as implicit sources of causal knowledge, aiming to extract their latent assumptions about causal relationships among events within a given topic. The approach involves generating topic-relevant text, extracting and normalizing events, constructing binary event indicator vectors, and applying causal discovery algorithms to infer underlying causal graph structures. By representing LLMs’ internal causal beliefs in terms of testable variables and graphical models, the framework offers a novel and interpretable means of externalizing these implicit assumptions. Empirical results demonstrate that the method successfully produces a set of plausible, interpretable candidate causal graphs, thereby validating the presence of coherent causal reasoning embedded within the model’s representations.

0 citationsRead paper

Bayesian Portfolio Optimization by Predictive Synthesis

Oct 08, 2025

Portfolio optimization typically relies on the assumed distribution of asset returns; however, this distribution is unknown in practice, and model forecasting performance degrades under market non-stationarity, undermining the robustness of conventional approaches. To address this, we propose a Bayesian Predictive Synthesis (BPS)-based portfolio optimization framework—the first application of BPS to asset allocation. It dynamically aggregates forecasts from multiple heterogeneous models via time-varying weights, yielding a posterior predictive distribution for asset returns with time-varying means. Integrated with dynamic linear models, the framework supports both mean–variance and quantile-based portfolio construction. Empirical results demonstrate that our method significantly improves the robustness of distributional forecasts, effectively adapts to model performance drift, and achieves superior risk-adjusted returns in uncertain, evolving market environments.

0 citationsRead paper

PUATE: Semiparametric Efficient Average Treatment Effect Estimation from Treated (Positive) and Unlabeled Units

Jan 31, 2025

This paper addresses the estimation of the average treatment effect (ATE) under partial treatment-status missingness—where only the treated units are labeled, and all others constitute unlabeled data—a problem at the intersection of causal inference and positive-unlabeled (PU) weakly supervised learning. We first derive the semiparametric efficient bound for PU-type ATE estimation and construct a doubly robust estimator achieving this bound. Our estimator integrates influence function theory, regularized machine learning, and the PU learning framework, ensuring $sqrt{n}$-consistency and asymptotic normality. Theoretically, it attains the minimal asymptotic variance among all regular estimators. Empirical evaluations—including simulations and real-data analysis—demonstrate its substantial improvement over existing approaches that either discard unlabeled samples or rely on misspecified models. To our knowledge, this is the first method provably achieving semiparametric efficiency for ATE estimation under PU sampling, thereby providing the first theoretically optimal solution for weakly supervised causal inference.

0 citationsRead paper
Recent publications

Latest Papers

Which Algorithm Specification Formats Help Language Models Implement Machine Learning Algorithms?

Jul 03, 2026

This study investigates how the format of algorithmic descriptions influences the accuracy of machine learning algorithm implementations generated by large language models (LLMs), with a focus on critical yet often implicit details such as interfaces, computational steps, numerical rules, and boundary behaviors. Through controlled experiments across multiple models (GPT-4o mini, Gemma 2 27B, Llama 3.2 3B), five ML algorithms, and seven description formats—including LaTeX pseudocode, YAML, and Python code stubs—the authors evaluate implementation correctness using fine-grained hidden tests. Results reveal that content clarity outweighs format per se: under core information conditions, LaTeX pseudocode yields the best performance, followed by YAML and plain text; with complete information, some models become format-agnostic, and code stubs offer no significant advantage. This work is the first to systematically demonstrate the pivotal role of algorithmic description readability in LLM-based implementation and offers writing guidelines optimized for AI interpretability.

0 citationsRead paper

Causality Elicitation from Large Language Models

Mar 04, 2026

This work proposes the first systematic framework for treating large language models (LLMs) as implicit sources of causal knowledge, aiming to extract their latent assumptions about causal relationships among events within a given topic. The approach involves generating topic-relevant text, extracting and normalizing events, constructing binary event indicator vectors, and applying causal discovery algorithms to infer underlying causal graph structures. By representing LLMs’ internal causal beliefs in terms of testable variables and graphical models, the framework offers a novel and interpretable means of externalizing these implicit assumptions. Empirical results demonstrate that the method successfully produces a set of plausible, interpretable candidate causal graphs, thereby validating the presence of coherent causal reasoning embedded within the model’s representations.

0 citationsRead paper

Bayesian Portfolio Optimization by Predictive Synthesis

Oct 08, 2025

Portfolio optimization typically relies on the assumed distribution of asset returns; however, this distribution is unknown in practice, and model forecasting performance degrades under market non-stationarity, undermining the robustness of conventional approaches. To address this, we propose a Bayesian Predictive Synthesis (BPS)-based portfolio optimization framework—the first application of BPS to asset allocation. It dynamically aggregates forecasts from multiple heterogeneous models via time-varying weights, yielding a posterior predictive distribution for asset returns with time-varying means. Integrated with dynamic linear models, the framework supports both mean–variance and quantile-based portfolio construction. Empirical results demonstrate that our method significantly improves the robustness of distributional forecasts, effectively adapts to model performance drift, and achieves superior risk-adjusted returns in uncertain, evolving market environments.

0 citationsRead paper

PUATE: Semiparametric Efficient Average Treatment Effect Estimation from Treated (Positive) and Unlabeled Units

Jan 31, 2025

This paper addresses the estimation of the average treatment effect (ATE) under partial treatment-status missingness—where only the treated units are labeled, and all others constitute unlabeled data—a problem at the intersection of causal inference and positive-unlabeled (PU) weakly supervised learning. We first derive the semiparametric efficient bound for PU-type ATE estimation and construct a doubly robust estimator achieving this bound. Our estimator integrates influence function theory, regularized machine learning, and the PU learning framework, ensuring $sqrt{n}$-consistency and asymptotic normality. Theoretically, it attains the minimal asymptotic variance among all regular estimators. Empirical evaluations—including simulations and real-data analysis—demonstrate its substantial improvement over existing approaches that either discard unlabeled samples or rely on misspecified models. To our knowledge, this is the first method provably achieving semiparametric efficiency for ATE estimation under PU sampling, thereby providing the first theoretically optimal solution for weakly supervised causal inference.

0 citationsRead paper