Mutual information and sensitivity analysis for feature selection in customer targeting: a comparative study
研究对比了互信息和数据敏感性分析在银行电话营销中的特征选择效果,通过构建逻辑回归模型评估两者优劣,为降低成本同时保持成功率提供依据。
研究对比了互信息和数据敏感性分析在银行电话营销中的特征选择效果,通过构建逻辑回归模型评估两者优劣,为降低成本同时保持成功率提供依据。
This study addresses the challenge of balancing predictive utility and privacy preservation in sensitive scenarios by leveraging synthetic time series as substitutes for real data. The authors propose a “train-on-synthetic, test-on-real” evaluation protocol and systematically compare multiple generative methods against noise-based anonymization baselines across seven datasets. They introduce Grasynda-P, an extension of graph-based generation that integrates matrix ensembling with kernel density estimation. Their experiments provide the first comprehensive characterization of the trade-off between prediction accuracy and distance-based privacy risk, revealing that no synthetic method fully replaces real data; noise-based anonymization offers the strongest privacy but poorest utility; simple transformations consistently outperform complex deep generative models; and Grasynda-P achieves a superior privacy-utility balance, residing on the Pareto frontier.
This work addresses the challenge in conceptual design where relying solely on safety invariants often fails to uniquely determine reaction rules aligned with user intent, leading to inconsistent synthesis outcomes. To overcome this, the authors propose a novel synthesis approach that integrates formal semantics with an LLM-driven Counterexample-Guided Inductive Synthesis (CEGIS) framework. The method leverages either positive/negative example scenarios or natural language prompts to guide the generation of rules satisfying given safety invariants, and introduces, for the first time, an LLM-assisted scenario-based elicitation mechanism to support early-stage design exploration. As the first effort to combine formal verification with LLM-based synthesis in conceptual design, experiments demonstrate that scenario-based guidance more reliably reproduces intended designs than natural language alone; with sufficient scenarios, LLM-augmented elicitation effectively recovers expected behaviors for most variants, though behavior omission and non-determinism remain key obstacles to achieving full coverage.
This study addresses inference bias in geostatistics arising from preferential sampling—where sampling locations are stochastically dependent on the underlying spatial process—by proposing a simple yet effective covariate construction method. The approach characterizes the dependence between sampling locations and the spatial variable through the average distance to nearest neighbors of observed points and incorporates this metric into standard geostatistical models. Without requiring complex model extensions, the method remains compatible with existing inference tools. Monte Carlo simulations and empirical analyses of two real-world datasets—Portuguese fishery landings and Galician lead contamination biomonitoring—demonstrate that the proposed approach substantially mitigates preferential sampling bias and enhances the reliability of spatial inference.
This work addresses the challenge of limited annotated data in medical image segmentation, exacerbated by privacy constraints, by proposing a scalable prefix-based data augmentation strategy that enhances few-shot 2D segmentation performance without altering the nnU-Net architecture. The approach leverages a two-stage image registration pipeline to generate deformed images along with their corresponding labels, and further incorporates synthetic mask generation and an efficient disk management mechanism to substantially increase training data diversity. Experimental results across five 2D medical image datasets demonstrate significant improvements over the original nnU-Net baseline, with Dice similarity coefficients increasing by up to approximately 22%.
研究对比了互信息和数据敏感性分析在银行电话营销中的特征选择效果,通过构建逻辑回归模型评估两者优劣,为降低成本同时保持成功率提供依据。
This study addresses the challenge of balancing predictive utility and privacy preservation in sensitive scenarios by leveraging synthetic time series as substitutes for real data. The authors propose a “train-on-synthetic, test-on-real” evaluation protocol and systematically compare multiple generative methods against noise-based anonymization baselines across seven datasets. They introduce Grasynda-P, an extension of graph-based generation that integrates matrix ensembling with kernel density estimation. Their experiments provide the first comprehensive characterization of the trade-off between prediction accuracy and distance-based privacy risk, revealing that no synthetic method fully replaces real data; noise-based anonymization offers the strongest privacy but poorest utility; simple transformations consistently outperform complex deep generative models; and Grasynda-P achieves a superior privacy-utility balance, residing on the Pareto frontier.
This work addresses the challenge in conceptual design where relying solely on safety invariants often fails to uniquely determine reaction rules aligned with user intent, leading to inconsistent synthesis outcomes. To overcome this, the authors propose a novel synthesis approach that integrates formal semantics with an LLM-driven Counterexample-Guided Inductive Synthesis (CEGIS) framework. The method leverages either positive/negative example scenarios or natural language prompts to guide the generation of rules satisfying given safety invariants, and introduces, for the first time, an LLM-assisted scenario-based elicitation mechanism to support early-stage design exploration. As the first effort to combine formal verification with LLM-based synthesis in conceptual design, experiments demonstrate that scenario-based guidance more reliably reproduces intended designs than natural language alone; with sufficient scenarios, LLM-augmented elicitation effectively recovers expected behaviors for most variants, though behavior omission and non-determinism remain key obstacles to achieving full coverage.
This study addresses inference bias in geostatistics arising from preferential sampling—where sampling locations are stochastically dependent on the underlying spatial process—by proposing a simple yet effective covariate construction method. The approach characterizes the dependence between sampling locations and the spatial variable through the average distance to nearest neighbors of observed points and incorporates this metric into standard geostatistical models. Without requiring complex model extensions, the method remains compatible with existing inference tools. Monte Carlo simulations and empirical analyses of two real-world datasets—Portuguese fishery landings and Galician lead contamination biomonitoring—demonstrate that the proposed approach substantially mitigates preferential sampling bias and enhances the reliability of spatial inference.
This work addresses the challenge of limited annotated data in medical image segmentation, exacerbated by privacy constraints, by proposing a scalable prefix-based data augmentation strategy that enhances few-shot 2D segmentation performance without altering the nnU-Net architecture. The approach leverages a two-stage image registration pipeline to generate deformed images along with their corresponding labels, and further incorporates synthetic mask generation and an efficient disk management mechanism to substantially increase training data diversity. Experimental results across five 2D medical image datasets demonstrate significant improvements over the original nnU-Net baseline, with Dice similarity coefficients increasing by up to approximately 22%.