Institution profile

Tokyo City University

Academic institutionasia · jp
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Closed-Form Steepest Descent Direction toward Flat Minima: Reducing Upper Bounds on the Loss Hessian Eigenspectrum in Neural Networks

Jun 26, 2026

This work addresses the challenge of improving neural network generalization by analyzing optimization dynamics through the lens of data distribution and parameter structure, with a focus on steering training toward flat minima. Leveraging the Wolkowicz–Styan inequality, the authors derive, for the first time, a closed-form expression for the gradient of the largest eigenvalue of the Hessian of the cross-entropy loss—enabling explicit optimization of its spectral upper bound without numerical approximation. By updating parameters along the steepest descent direction defined by this gradient, the method effectively compresses the Hessian eigenvalue spectrum in three-layer networks, thereby avoiding sharp minima and saddle points. This approach guides convergence toward flatter minima that exhibit superior generalization, offering a novel theoretical and algorithmic pathway for understanding and promoting flatness in deep learning optimization.

0 citationsRead paper

Analytical Evaluation of DCA Convergence Properties for Minimizing Prediction Functions of Gaussian RBF Support Vector Regression

Jun 02, 2026

This study addresses the lack of a priori convergence assessment for the Difference-of-Convex Algorithm (DCA) in Gaussian radial basis function kernel support vector regression (RBF-SVR). By exploiting the analytical structure of the RBF kernel, the authors construct an explicit DC decomposition and, for the first time, derive closed-form expressions for the lower bound μ of the strong convexity parameter and the upper bound L of the gradient Lipschitz constant of the DC components. They further identify the scalar Cαρ—determined by the hyperparameters C and γ—as the key quantity governing DCA’s convergence rate and dependence on initialization. Numerical experiments on six benchmark functions confirm that Cαρ alone effectively predicts DCA’s convergence behavior both before and after training, establishing the first analytical link between RBF-SVR hyperparameters and DCA convergence.

0 citationsRead paper

Wolkowicz-Styan Upper Bound on the Hessian Eigenspectrum for Cross-Entropy Loss in Nonlinear Smooth Neural Networks

Apr 11, 2026

Existing theoretical frameworks struggle to characterize the largest eigenvalue of the Hessian of the loss function in smooth nonlinear multilayer neural networks, hindering a deeper understanding of the relationship between loss sharpness and generalization. This work addresses this gap by deriving, for the first time, a closed-form upper bound on the largest Hessian eigenvalue for deep networks employing cross-entropy loss and smooth activation functions—extending beyond prior results limited to linear or ReLU networks. The bound is established through a combination of the Wolkowicz–Styan inequality, second-order Taylor expansion, and spectral analysis, and it explicitly depends on network parameters, hidden layer dimensions, and the orthogonality of training samples. Notably, it requires no numerical computation, thereby offering a new analytical tool for theoretically studying loss sharpness.

0 citationsRead paper

Generating High-Level Test Cases from Requirements using LLM: An Industry Study

Oct 03, 2025

Existing advanced test case generation approaches heavily rely on manual effort or customized RAG systems, suffering from poor generalizability and high construction costs. Method: This paper proposes a pure prompt-engineering–driven large language model (LLM) approach that automatically generates functional-level natural language test cases directly from requirements documents—without retrieval-augmented generation (RAG). It first guides the LLM to identify appropriate test design techniques for the given requirements, then leverages this selection to generate concrete test cases, enabling end-to-end zero-shot inference. Contribution/Results: The method significantly improves cross-domain generalization and industrial deployability. On the Bluetooth and Mozilla datasets, it achieves macro-recall scores of 0.81 and 0.37, respectively, demonstrating both effectiveness and practical potential. To our knowledge, this is the first work achieving high-level test case automation solely through structured prompting—without any retrieval augmentation.

0 citationsRead paper
Recent publications

Latest Papers

Closed-Form Steepest Descent Direction toward Flat Minima: Reducing Upper Bounds on the Loss Hessian Eigenspectrum in Neural Networks

Jun 26, 2026

This work addresses the challenge of improving neural network generalization by analyzing optimization dynamics through the lens of data distribution and parameter structure, with a focus on steering training toward flat minima. Leveraging the Wolkowicz–Styan inequality, the authors derive, for the first time, a closed-form expression for the gradient of the largest eigenvalue of the Hessian of the cross-entropy loss—enabling explicit optimization of its spectral upper bound without numerical approximation. By updating parameters along the steepest descent direction defined by this gradient, the method effectively compresses the Hessian eigenvalue spectrum in three-layer networks, thereby avoiding sharp minima and saddle points. This approach guides convergence toward flatter minima that exhibit superior generalization, offering a novel theoretical and algorithmic pathway for understanding and promoting flatness in deep learning optimization.

0 citationsRead paper

Analytical Evaluation of DCA Convergence Properties for Minimizing Prediction Functions of Gaussian RBF Support Vector Regression

Jun 02, 2026

This study addresses the lack of a priori convergence assessment for the Difference-of-Convex Algorithm (DCA) in Gaussian radial basis function kernel support vector regression (RBF-SVR). By exploiting the analytical structure of the RBF kernel, the authors construct an explicit DC decomposition and, for the first time, derive closed-form expressions for the lower bound μ of the strong convexity parameter and the upper bound L of the gradient Lipschitz constant of the DC components. They further identify the scalar Cαρ—determined by the hyperparameters C and γ—as the key quantity governing DCA’s convergence rate and dependence on initialization. Numerical experiments on six benchmark functions confirm that Cαρ alone effectively predicts DCA’s convergence behavior both before and after training, establishing the first analytical link between RBF-SVR hyperparameters and DCA convergence.

0 citationsRead paper

Wolkowicz-Styan Upper Bound on the Hessian Eigenspectrum for Cross-Entropy Loss in Nonlinear Smooth Neural Networks

Apr 11, 2026

Existing theoretical frameworks struggle to characterize the largest eigenvalue of the Hessian of the loss function in smooth nonlinear multilayer neural networks, hindering a deeper understanding of the relationship between loss sharpness and generalization. This work addresses this gap by deriving, for the first time, a closed-form upper bound on the largest Hessian eigenvalue for deep networks employing cross-entropy loss and smooth activation functions—extending beyond prior results limited to linear or ReLU networks. The bound is established through a combination of the Wolkowicz–Styan inequality, second-order Taylor expansion, and spectral analysis, and it explicitly depends on network parameters, hidden layer dimensions, and the orthogonality of training samples. Notably, it requires no numerical computation, thereby offering a new analytical tool for theoretically studying loss sharpness.

0 citationsRead paper

Generating High-Level Test Cases from Requirements using LLM: An Industry Study

Oct 03, 2025

Existing advanced test case generation approaches heavily rely on manual effort or customized RAG systems, suffering from poor generalizability and high construction costs. Method: This paper proposes a pure prompt-engineering–driven large language model (LLM) approach that automatically generates functional-level natural language test cases directly from requirements documents—without retrieval-augmented generation (RAG). It first guides the LLM to identify appropriate test design techniques for the given requirements, then leverages this selection to generate concrete test cases, enabling end-to-end zero-shot inference. Contribution/Results: The method significantly improves cross-domain generalization and industrial deployability. On the Bluetooth and Mozilla datasets, it achieves macro-recall scores of 0.81 and 0.37, respectively, demonstrating both effectiveness and practical potential. To our knowledge, this is the first work achieving high-level test case automation solely through structured prompting—without any retrieval augmentation.

0 citationsRead paper