GeoSPRINT: Geometric Redundancy-Aware Step Pruning for Inference in Diffusion Trajectories
本文提出GeoSPRINT方法,通过检测去噪轨迹中的几何冗余来优化扩散模型的推理步骤分配,从而在不增加训练成本的情况下提高采样效率。
本文提出GeoSPRINT方法,通过检测去噪轨迹中的几何冗余来优化扩散模型的推理步骤分配,从而在不增加训练成本的情况下提高采样效率。
Existing AI benchmarks predominantly assess programming proficiency, exhibiting a critical gap in systematically evaluating high-economic-value knowledge work—such as investment banking, management consulting, legal practice, and primary healthcare. Method: We introduce APEX-v1.0, the first comprehensive benchmark for this domain, comprising 200 realistic, expert-designed tasks with fine-grained, rubric-based scoring. Leveraging an innovative “expert-defined tasks + LLM-based automated adjudication” paradigm, we conduct large-scale, reproducible evaluations across 23 state-of-the-art models. Contribution/Results: GPT-5 (Thinking=High) achieves the highest average accuracy (64.2%), while Qwen3-235B emerges as the top-performing open-weight model. Nevertheless, all models fall substantially short of human expert performance. APEX-v1.0 establishes the first standardized, scalable, and reproducible evaluation infrastructure for non-coding professional reasoning, advancing both technical development and responsible governance of AI in high-stakes knowledge-intensive domains.
Infectious and immune-mediated disease (IID) data are fragmented across disparate sources, lack standardized metadata schemas, and suffer from poor discoverability and reusability. Method: We developed the first unified metadata search platform specifically for the IID domain, harmonizing over 4 million dataset-level metadata records from 400+ specialized and general-purpose databases. Through format normalization, semantic integration, and construction of a domain-specific ontology, the platform enables natural-language search, predefined queries, faceted browsing, and programmatic API access. Contribution/Results: This work represents the first systematic, cross-source metadata aggregation and interoperability framework for IID data, substantially enhancing Findability, Accessibility, Interoperability, and Reusability (FAIRness). The platform is actively supporting NIH/NIAID-funded projects and global researchers in hypothesis-driven analysis, cross-cohort comparison, and secondary analysis of public datasets—thereby increasing the scientific return on investment in biomedical data infrastructure.
本文提出GeoSPRINT方法,通过检测去噪轨迹中的几何冗余来优化扩散模型的推理步骤分配,从而在不增加训练成本的情况下提高采样效率。
Existing AI benchmarks predominantly assess programming proficiency, exhibiting a critical gap in systematically evaluating high-economic-value knowledge work—such as investment banking, management consulting, legal practice, and primary healthcare. Method: We introduce APEX-v1.0, the first comprehensive benchmark for this domain, comprising 200 realistic, expert-designed tasks with fine-grained, rubric-based scoring. Leveraging an innovative “expert-defined tasks + LLM-based automated adjudication” paradigm, we conduct large-scale, reproducible evaluations across 23 state-of-the-art models. Contribution/Results: GPT-5 (Thinking=High) achieves the highest average accuracy (64.2%), while Qwen3-235B emerges as the top-performing open-weight model. Nevertheless, all models fall substantially short of human expert performance. APEX-v1.0 establishes the first standardized, scalable, and reproducible evaluation infrastructure for non-coding professional reasoning, advancing both technical development and responsible governance of AI in high-stakes knowledge-intensive domains.
Infectious and immune-mediated disease (IID) data are fragmented across disparate sources, lack standardized metadata schemas, and suffer from poor discoverability and reusability. Method: We developed the first unified metadata search platform specifically for the IID domain, harmonizing over 4 million dataset-level metadata records from 400+ specialized and general-purpose databases. Through format normalization, semantic integration, and construction of a domain-specific ontology, the platform enables natural-language search, predefined queries, faceted browsing, and programmatic API access. Contribution/Results: This work represents the first systematic, cross-source metadata aggregation and interoperability framework for IID data, substantially enhancing Findability, Accessibility, Interoperability, and Reusability (FAIRness). The platform is actively supporting NIH/NIAID-funded projects and global researchers in hypothesis-driven analysis, cross-cohort comparison, and secondary analysis of public datasets—thereby increasing the scientific return on investment in biomedical data infrastructure.