Score
Performs cost-aware performance evaluation by designing risk-adjusted metrics and evaluation frameworks that trade off performance against cost or risk.
This paper addresses longevity risk, market volatility, high operational costs, and default probability in pension plan management. We propose a simulation-based modular optimization framework that innovatively integrates adaptive asset allocation with dynamic benefit adjustment, enabling real-time, joint liability–asset management. A customizable, multi-dimensional performance metric system supports systematic calibration of strategic parameters. Numerical experiments demonstrate that introducing limited flexibility—such as annual portfolio rebalancing and modest benefit adjustments—reduces operational costs by 18%–32% and decreases default probability by up to 76%. The framework thus provides a scalable, empirically validated methodology for designing robust, cost-efficient, and low-default pension systems.
Existing architecture evaluation methods (e.g., ATAM) focus on concrete systems and struggle with the generality and variability challenges arising from multi-level abstractions—software architectures, reference architectures, and architecture frameworks. This paper proposes ATRAF, a scenario-driven unified assessment framework, introducing the first cross-level trade-off analysis paradigm. It comprises three complementary methods: ATRAM (for software architecture), RATRAM (for reference architecture), and AFTRAM (for architecture framework), collectively supporting holistic quality attribute trade-offs (e.g., modifiability, performance, security) and risk identification. Leveraging iterative, spiral-style scenario modeling, sensitivity analysis, and feedback-driven optimization, ATRAF refines and extends ATAM. Evaluated on an RTS case study, ATRAF demonstrates robust incremental assessment capability, significantly improving early architectural decision quality and cross-level consistency assurance.
This work addresses the limitations of conventional experimental evaluation methods, which treat multiple metrics in isolation, ignore their interdependencies, and rely on subjective judgments when objectives conflict—hindering scalability. To overcome these challenges, the authors propose the first framework that integrates Bayesian decision theory with hierarchical priors. By designing a custom loss function that incorporates business preferences alongside observed evidence, and leveraging historical experiment data to construct informative priors, the approach enables automated and systematic trade-offs among multiple objectives. Evaluated on both real-world and simulated supply chain experiments at Amazon, the method significantly improves estimation efficiency, streamlines complex decision-making processes, and transcends the constraints of traditional hypothesis testing.
This study addresses the limitations of existing approaches in simultaneously accounting for the time value of money and the integrated effects of multidimensional decision criteria on financial risk in manufacturing firms, while also overlooking the interactions among economic, operational, and managerial factors. To bridge this gap, we propose an evaluation framework that integrates a compound discounting model with multicriteria linear regression. For the first time, a time-discounting mechanism is incorporated into multicriteria decision analysis, enabling unified treatment of one-time expenditures, proportional costs, and complex cost structures. The method effectively quantifies the present value of costs and benefits across different time points and reveals how synergistic interactions among multiple factors influence discounted performance. This approach significantly enhances the systematicity and accuracy of financial risk assessment, offering manufacturing enterprises quantifiable decision support for optimizing the economic efficiency of control systems.
This paper addresses the limitations of conventional classification models in decision optimization—specifically, their neglect of cost sensitivity and causal effects. We propose the first unified evaluation framework integrating cost-sensitive learning and causal inference. Methodologically, we formalize standard classification as a special case of single-action causal classification and, grounded in decision theory and axiomatic performance measurement, construct an extensible family of causal performance metrics. Theoretical contributions include: (i) the first systematic unification of cost-sensitive and causal learning paradigms; (ii) rigorous proof of their intrinsic consistency; and (iii) reconstruction and generalization of established industry metrics (e.g., Qini, ROI). Empirical evaluations demonstrate that our framework significantly improves profit-maximizing decisions in real-world business applications—including customer retention and response modeling—outperforming both standard classification and isolated causal or cost-sensitive approaches.
This work addresses the challenges of low quality and poor transparency in build-or-buy decisions within enterprise software development, which often stem from reliance on unstructured experiential knowledge. To overcome these limitations—particularly in cold-start scenarios lacking historical data—the authors propose a structured approach that integrates a decision-factor ontology, rule-based reasoning, and reference-class matching. This method enables transparent, auditable evaluation of alternatives and represents the first application of combined ontology modeling and rule reasoning to build-or-buy decision-making. By revealing critical decision thresholds and supporting traceability, the approach enhances the rationality, transparency, and auditability of choices. Its practical efficacy is demonstrated through a lightweight tool validated in a financial industry case study, showing significant improvements in decision quality.
This study addresses the limitations of traditional risk-adjusted performance measures—such as the Sharpe ratio—in effectively ranking financial or insurance positions. The authors develop an axiomatic framework that introduces monotonicity and cash quasi-concavity to define a novel class of ranking measures, which map positions directly to performance levels rather than standardized returns. This approach establishes theoretical connections with acceptance sets and risk measures, encompassing classical ratio-based metrics while extending to new ranking methodologies grounded in expected shortfall, Lambda quantiles, and bibliometric-inspired constructions. Empirical analyses involving portfolio rankings and climate risk insurance demonstrate both the theoretical coherence and practical applicability of the proposed framework.
This study addresses the high cost of full-scale performance regression testing in large-scale continuous integration (CI), where existing approaches struggle to balance submission heterogeneity and resource efficiency. The work proposes the first framework integrating commit-level regression risk prediction with dynamic batching, introducing novel risk-aware scheduling strategies such as Risk-Aged Priority Batching (RAPB). Leveraging real-world Mozilla Firefox datasets, the authors fine-tune ModernBERT, CodeBERT, and LLaMA-3.1 models to predict regression risk and validate their approach through CI simulation. The optimal configuration, RAPB-la, reduces test execution volume by 32.4%, shortens average feedback time by 3.8%, decreases maximum localization latency by 26.2%, and yields an estimated annual infrastructure cost saving of approximately $491,000.
This study challenges the conventional paradigm of risk minimization that relies on explicit risk measures, proposing instead a novel approach to risk reduction that does not require pre-specified risk metrics. The method leverages the full spectrum of information contained in the portfolio return matrix, employing generalized numerical rank and full-spectrum condition numbers to rank risk scenarios and selectively attenuate exposures associated with high risk. In contrast to traditional strategies that focus solely on the smallest eigenvalue, this approach holistically accounts for the entire spectral structure. Empirical results on real-world data demonstrate that the proposed strategy significantly reduces out-of-sample return volatility while maintaining average returns and Sharpe ratios comparable to those of benchmark portfolios based on standard risk measures, all within realistic transaction cost constraints.
This study addresses the lack of systematic evaluation regarding whether historical marketing budget allocations approximate ex post optimal spending. The authors propose a hindsight-regret-based retrospective auditing framework that estimates spend–response functions and computes feasible ex post optimal allocations under budget and stability constraints. By integrating Monte Carlo methods to quantify estimation uncertainty, the framework effectively disentangles true allocation inefficiency from model estimation error. Introducing hindsight regret into marketing budget auditing for the first time, this approach reveals a practical trade-off between allocation flexibility and detectability, enabling reliable diagnostics when online experimentation is infeasible. Empirical evaluation on real-world marketing logs demonstrates that the method produces interpretable diagnostic reports and that moderate reallocations can capture the majority of measurable gains.