Automated pipeline for herbarium label digitization
为解决标本标签元数据难以大规模访问的问题,本文提出HERBIOME,一种自动化处理流程,利用YOLOv8、CRAFT Hezar、TrOCR和GPT-4o Mini等技术实现标签信息的自动提取与结构化。
为解决标本标签元数据难以大规模访问的问题,本文提出HERBIOME,一种自动化处理流程,利用YOLOv8、CRAFT Hezar、TrOCR和GPT-4o Mini等技术实现标签信息的自动提取与结构化。
为解决标本图像中背景元素干扰植物特征识别的问题,提出AT-ViT模型,采用多尺度、多视角交叉注意力融合和掩码引导的补丁加权机制,提高对植物特征的学习准确性。
This study addresses the incomplete depth coverage of temperature–salinity profiles collected by deep-diving marine mammals in the Indian Ocean sector of the Southern Ocean, which arises from behavioral differences among individuals. To overcome this limitation, the authors propose a multivariate functional principal component analysis method incorporating geographic covariates. By modeling the mean and covariance structure of complete bivariate profiles, they construct eigenfunction bases and integrate a measurement error model to estimate conditional functional principal scores, enabling high-fidelity reconstruction of truncated profiles across the full depth range. In simulations, the approach improves reconstruction accuracy by 30% for temperature and 33% for salinity in the 20–500 m layer when applied to profiles truncated at 250 m. The method successfully reconstructs approximately 90,000 profiles across a 3-million-square-kilometer region surrounding the French subantarctic islands, marking the first large-scale, accurate recovery of incomplete oceanographic profiles.
This study addresses the high computational cost of full-dataset scans for data quality assessment—such as detecting missing values, duplicates, and outliers—which hinders near real-time monitoring. The authors systematically evaluate nine progressive sampling strategies across diverse real-world and synthetic datasets, comparing blind sampling methods (e.g., uniform random, clustering-based, Yamane) against proxy-guided approaches (e.g., MCMC, DAG-based, stratified weighting). Contrary to the prevailing assumption that incorporating prior knowledge improves accuracy, large-scale empirical results reveal that representative blind sampling consistently outperforms proxy-guided techniques, primarily due to mismatches between proxy metrics and actual data quality issues. Notably, under a 5% sampling budget, uniform random sampling achieves an average relative error of merely 0.49% with near-linear scalability, whereas DAG-guided methods incur 11–49× higher errors and run 28–47× slower, demonstrating that simple sampling strategies are better suited for production-grade data quality monitoring.
This study addresses the challenges of legal metric computation—namely textual complexity, large scale, high interpretability demands, and variable data quality—which often lead existing methods to produce hallucinations and lack explainability or evidentiary support. To overcome these limitations, this work proposes the first agent-driven Retrieval-Augmented Generation (RAG) framework tailored for legal metric calculation. The approach employs a modular pipeline that integrates adaptive retrieval, large language model agents, and validation mechanisms to enable transparent and traceable evidence selection, statutory linkage, and binary judgment. Experiments on a newly constructed corpus of French maritime environmental law demonstrate that the proposed method significantly outperforms baseline systems and exhibits strong generalization across two injunction-related tasks, offering a novel paradigm for building trustworthy and scalable legal monitoring systems.
为解决标本标签元数据难以大规模访问的问题,本文提出HERBIOME,一种自动化处理流程,利用YOLOv8、CRAFT Hezar、TrOCR和GPT-4o Mini等技术实现标签信息的自动提取与结构化。
为解决标本图像中背景元素干扰植物特征识别的问题,提出AT-ViT模型,采用多尺度、多视角交叉注意力融合和掩码引导的补丁加权机制,提高对植物特征的学习准确性。
This study addresses the incomplete depth coverage of temperature–salinity profiles collected by deep-diving marine mammals in the Indian Ocean sector of the Southern Ocean, which arises from behavioral differences among individuals. To overcome this limitation, the authors propose a multivariate functional principal component analysis method incorporating geographic covariates. By modeling the mean and covariance structure of complete bivariate profiles, they construct eigenfunction bases and integrate a measurement error model to estimate conditional functional principal scores, enabling high-fidelity reconstruction of truncated profiles across the full depth range. In simulations, the approach improves reconstruction accuracy by 30% for temperature and 33% for salinity in the 20–500 m layer when applied to profiles truncated at 250 m. The method successfully reconstructs approximately 90,000 profiles across a 3-million-square-kilometer region surrounding the French subantarctic islands, marking the first large-scale, accurate recovery of incomplete oceanographic profiles.
This study addresses the high computational cost of full-dataset scans for data quality assessment—such as detecting missing values, duplicates, and outliers—which hinders near real-time monitoring. The authors systematically evaluate nine progressive sampling strategies across diverse real-world and synthetic datasets, comparing blind sampling methods (e.g., uniform random, clustering-based, Yamane) against proxy-guided approaches (e.g., MCMC, DAG-based, stratified weighting). Contrary to the prevailing assumption that incorporating prior knowledge improves accuracy, large-scale empirical results reveal that representative blind sampling consistently outperforms proxy-guided techniques, primarily due to mismatches between proxy metrics and actual data quality issues. Notably, under a 5% sampling budget, uniform random sampling achieves an average relative error of merely 0.49% with near-linear scalability, whereas DAG-guided methods incur 11–49× higher errors and run 28–47× slower, demonstrating that simple sampling strategies are better suited for production-grade data quality monitoring.
This study addresses the challenges of legal metric computation—namely textual complexity, large scale, high interpretability demands, and variable data quality—which often lead existing methods to produce hallucinations and lack explainability or evidentiary support. To overcome these limitations, this work proposes the first agent-driven Retrieval-Augmented Generation (RAG) framework tailored for legal metric calculation. The approach employs a modular pipeline that integrates adaptive retrieval, large language model agents, and validation mechanisms to enable transparent and traceable evidence selection, statutory linkage, and binary judgment. Experiments on a newly constructed corpus of French maritime environmental law demonstrate that the proposed method significantly outperforms baseline systems and exhibits strong generalization across two injunction-related tasks, offering a novel paradigm for building trustworthy and scalable legal monitoring systems.