Redakto - The Incognito Tab for LLMs
为解决LLM使用中的隐私问题,本文提出Redakto工具,通过匿名化和伪匿名化文本处理,确保个人信息安全,同时保持文本实用性。
为解决LLM使用中的隐私问题,本文提出Redakto工具,通过匿名化和伪匿名化文本处理,确保个人信息安全,同时保持文本实用性。
研究针对高风险公共部门应用中的信息提取问题,通过评估开源OCR引擎、大语言模型和视觉-语言模型在处理复杂文档任务上的表现,揭示了现有模型的局限性和影响因素。
This work addresses the gap between theory and practice in error detection and cleaning for tabular data by proposing and implementing an interactive web-based demonstration system that integrates error modeling, injection, and intelligent cleaning. For the first time, the system unifies machine learning–based data cleaning methods with dependency-aware error generation models within a visual platform, enabling users to upload tables, inject realistic errors, and interactively experience the cleaning process alongside interpretable explanations of the underlying mechanisms. By synergizing machine learning, database technologies, and web development, the project enhances intuitive understanding of error repair principles and provides a publicly accessible demo platform (https://cured.demo.calgo-lab.de/) that serves as an effective validation tool for both research and practical applications in data cleaning.
To address low labeling efficiency for soil profile data under limited expert annotation resources, this paper proposes an uncertainty-guided human-in-the-loop annotation framework. Methodologically, we first integrate model-agnostic conformal prediction into SoilNet—a multimodal, multitask deep regression model for soil depth estimation—to achieve reliable and well-calibrated uncertainty quantification. Based on these uncertainty estimates, we design a budget-constrained active annotation pipeline that dynamically triggers expert intervention only when model prediction confidence falls below a predefined threshold. Our contributions are: (1) the first conformalized uncertainty calibration scheme tailored to soil profile regression; and (2) statistically significant improvement in regression accuracy (p < 0.01) under identical annotation budgets, while maintaining classification performance comparable to baselines—demonstrating the effectiveness of uncertainty-driven optimization of expert labeling effort.
This study addresses the quantification of uncertainty in explanation outputs within explainable artificial intelligence (XAI), specifically modeling the joint impact of input perturbations and model parameter variations on the explanation function $e_ heta(x,f)$. We propose the first unified framework that formally characterizes uncertainty propagation in XAI. Our analysis reveals, for the first time, systematic failures of mainstream methods—including LIME and SHAP—in capturing explanation uncertainty. To enable rigorous evaluation, we establish a benchmark that enables direct comparison between analytical (first-order uncertainty propagation) and empirical (Monte Carlo variance estimation) approaches, validating their complementary strengths across diverse, heterogeneous datasets. We further introduce an explanation consistency metric to systematically assess robustness. All evaluation protocols and open-source code are publicly released, providing both theoretical foundations and practical tools for reliability assessment in XAI.
为解决LLM使用中的隐私问题,本文提出Redakto工具,通过匿名化和伪匿名化文本处理,确保个人信息安全,同时保持文本实用性。
研究针对高风险公共部门应用中的信息提取问题,通过评估开源OCR引擎、大语言模型和视觉-语言模型在处理复杂文档任务上的表现,揭示了现有模型的局限性和影响因素。
This work addresses the gap between theory and practice in error detection and cleaning for tabular data by proposing and implementing an interactive web-based demonstration system that integrates error modeling, injection, and intelligent cleaning. For the first time, the system unifies machine learning–based data cleaning methods with dependency-aware error generation models within a visual platform, enabling users to upload tables, inject realistic errors, and interactively experience the cleaning process alongside interpretable explanations of the underlying mechanisms. By synergizing machine learning, database technologies, and web development, the project enhances intuitive understanding of error repair principles and provides a publicly accessible demo platform (https://cured.demo.calgo-lab.de/) that serves as an effective validation tool for both research and practical applications in data cleaning.
To address low labeling efficiency for soil profile data under limited expert annotation resources, this paper proposes an uncertainty-guided human-in-the-loop annotation framework. Methodologically, we first integrate model-agnostic conformal prediction into SoilNet—a multimodal, multitask deep regression model for soil depth estimation—to achieve reliable and well-calibrated uncertainty quantification. Based on these uncertainty estimates, we design a budget-constrained active annotation pipeline that dynamically triggers expert intervention only when model prediction confidence falls below a predefined threshold. Our contributions are: (1) the first conformalized uncertainty calibration scheme tailored to soil profile regression; and (2) statistically significant improvement in regression accuracy (p < 0.01) under identical annotation budgets, while maintaining classification performance comparable to baselines—demonstrating the effectiveness of uncertainty-driven optimization of expert labeling effort.
This study addresses the quantification of uncertainty in explanation outputs within explainable artificial intelligence (XAI), specifically modeling the joint impact of input perturbations and model parameter variations on the explanation function $e_ heta(x,f)$. We propose the first unified framework that formally characterizes uncertainty propagation in XAI. Our analysis reveals, for the first time, systematic failures of mainstream methods—including LIME and SHAP—in capturing explanation uncertainty. To enable rigorous evaluation, we establish a benchmark that enables direct comparison between analytical (first-order uncertainty propagation) and empirical (Monte Carlo variance estimation) approaches, validating their complementary strengths across diverse, heterogeneous datasets. We further introduce an explanation consistency metric to systematically assess robustness. All evaluation protocols and open-source code are publicly released, providing both theoretical foundations and practical tools for reliability assessment in XAI.