Zero-I/O Fault Recovery for Sharded Deep Learning via Dynamic Framework Dependency Rebinding
本文提出AccelPact,通过动态框架依赖重绑定解决大规模分布式模型训练中由网络故障导致的I/O密集型恢复问题,实现零I/O内存故障恢复。
本文提出AccelPact,通过动态框架依赖重绑定解决大规模分布式模型训练中由网络故障导致的I/O密集型恢复问题,实现零I/O内存故障恢复。
SemBridge通过编译图和运行时事实为类型化契约,解决了分布式张量系统与集体系统间通信问题,生成消费者可见的结果及交付义务,并验证通信计划。
This work addresses the challenge of information extraction from visually rich documents, where diverse layouts and real-world interferences—such as stamps and low contrast—hinder performance. Existing approaches typically rely on extensive annotated data and layout-specific training. To overcome these limitations, the authors propose a classification-guided large vision-language model framework that decouples document-type classification from content extraction and leverages in-context learning to dynamically construct prompts, enabling zero-shot information extraction across multiple document types. Built upon Qwen2.5-VL-7B, the method employs conditional computation to reduce task uncertainty. Evaluated on a real-world tendering dataset comprising 16 certificate categories, it achieves a zero-shot F1 score of 86.43%, outperforming a strong supervised baseline by 18.35 percentage points; with fine-tuning, F1 further improves to 93.65%, attaining a normalized edit distance of 0.93.
This work addresses the challenge of cross-category 3D point cloud anomaly detection using only a few normal samples by proposing the first training-free, general-purpose framework. The method projects 3D point clouds into multi-view realistic depth maps and leverages a frozen CLIP vision encoder to extract features, enabling anomaly identification through weighted feature similarity without any category-specific fine-tuning or adaptation. Experimental results demonstrate that the approach achieves state-of-the-art performance under few-shot settings on the ShapeNetPart dataset, significantly enhancing the generality, practicality, and robustness of cross-category 3D anomaly detection.
本文提出AccelPact,通过动态框架依赖重绑定解决大规模分布式模型训练中由网络故障导致的I/O密集型恢复问题,实现零I/O内存故障恢复。
SemBridge通过编译图和运行时事实为类型化契约,解决了分布式张量系统与集体系统间通信问题,生成消费者可见的结果及交付义务,并验证通信计划。
This work addresses the challenge of information extraction from visually rich documents, where diverse layouts and real-world interferences—such as stamps and low contrast—hinder performance. Existing approaches typically rely on extensive annotated data and layout-specific training. To overcome these limitations, the authors propose a classification-guided large vision-language model framework that decouples document-type classification from content extraction and leverages in-context learning to dynamically construct prompts, enabling zero-shot information extraction across multiple document types. Built upon Qwen2.5-VL-7B, the method employs conditional computation to reduce task uncertainty. Evaluated on a real-world tendering dataset comprising 16 certificate categories, it achieves a zero-shot F1 score of 86.43%, outperforming a strong supervised baseline by 18.35 percentage points; with fine-tuning, F1 further improves to 93.65%, attaining a normalized edit distance of 0.93.
This work addresses the challenge of cross-category 3D point cloud anomaly detection using only a few normal samples by proposing the first training-free, general-purpose framework. The method projects 3D point clouds into multi-view realistic depth maps and leverages a frozen CLIP vision encoder to extract features, enabling anomaly identification through weighted feature similarity without any category-specific fine-tuning or adaptation. Experimental results demonstrate that the approach achieves state-of-the-art performance under few-shot settings on the ShapeNetPart dataset, significantly enhancing the generality, practicality, and robustness of cross-category 3D anomaly detection.