Institution profile

DP Technology

Industry researchasia · cn
Official website
Research library69linked papers
Opportunities0open roles
Selected work

Representative Papers

Toward Ultra-Long-Horizon Agentic Science: Cognitive Accumulation for Machine Learning Engineering

Jan 15, 2026

This work addresses the challenge of maintaining strategic coherence and iterative refinement in artificial intelligence systems over ultra-long scientific research cycles. To this end, it introduces the ML-Master 2.0 agent, which reconfigures context management as a cognitive accumulation process through a Hierarchical Cognitive Cache (HCC) architecture. Inspired by multi-level memory systems, HCC dynamically distills execution trajectories into stable knowledge, decoupling immediate actions from long-term strategy and thereby transcending the limitations of static context windows. Integrated with dynamic knowledge distillation, cross-task experience consolidation, and large language model–driven autonomous experiment planning, the proposed approach achieves a state-of-the-art medal rate of 56.44% on MLE-Bench under a 24-hour budget, demonstrating for the first time the feasibility of fully autonomous, ultra-long-horizon scientific discovery.

4 citationsRead paper

Innovator-VL: A Multimodal Large Language Model for Scientific Discovery

Jan 27, 2026

This work addresses the challenge of developing models that simultaneously exhibit strong scientific reasoning and general multimodal capabilities without relying on massive domain-specific datasets. The authors propose a data-efficient, transparent, and fully reproducible end-to-end training paradigm comprising high-quality scientific data curation, supervised fine-tuning, and reinforcement learning. Using fewer than five million samples, the resulting model significantly reduces dependence on large-scale pretraining data while achieving performance on scientific reasoning tasks comparable to that of much larger models. Moreover, it remains competitive on standard vision and multimodal benchmarks, demonstrating that scientific intelligence and general-purpose capabilities can effectively coexist within a single architecture.

2 citationsRead paper

Deploy-Master: Automating the Deployment of 50,000+ Agent-Ready Scientific Tools in One Day

Jan 07, 2026arXiv.org

This work addresses the critical barrier to reproducibility and integration with AI-for-Science (AI4S) agent workflows posed by the complex compilation and configuration of scientific software. We propose an end-to-end agent workflow that leverages a domain-specific taxonomy to filter repositories, automatically infers build specifications, containerizes the resulting artifacts, and validates their executability. For the first time, this pipeline enables the automated deployment of over 50,000 scientific tools within a single day. We successfully constructed reproducible execution environments for 50,112 tools, each verified by a minimal executable command, and released them in SciencePedia—the first trustworthy capability repository grounded in actual execution rather than documentation. Additionally, we publicly share the large-scale deployment traces to illuminate operational bottlenecks in scientific software packaging.

1 citationsRead paper

MolParser-Mobile: Ultrafast OCSR System for Large-Scale Chemical Literature Mining

Sep 05, 2026

Optical Chemical Structure Recognition (OCSR) is a fundamental component of chemical literature mining, enabling molecular database construction, reaction extraction, and AI-driven scientific discovery. Despite substantial progress in recognition accuracy with recent deep learning-based methods, inference throughput remains a critical bottleneck that limits web-scale deployment. To address this challenge, we propose MolParser-Mobile, an AutoML-optimized lightweight end-to-end OCSR framework. MolParser-Mobile contains only 9.98M parameters, while reaching a throughput of 1,520 molecules per second on a single NVIDIA RTX 4090D GPU. Despite its compact design, it maintains competitive and, on several benchmarks, superior recognition accuracy.

0 citationsRead paper
Recent publications

Latest Papers

MolParser-Mobile: Ultrafast OCSR System for Large-Scale Chemical Literature Mining

Sep 05, 2026

Optical Chemical Structure Recognition (OCSR) is a fundamental component of chemical literature mining, enabling molecular database construction, reaction extraction, and AI-driven scientific discovery. Despite substantial progress in recognition accuracy with recent deep learning-based methods, inference throughput remains a critical bottleneck that limits web-scale deployment. To address this challenge, we propose MolParser-Mobile, an AutoML-optimized lightweight end-to-end OCSR framework. MolParser-Mobile contains only 9.98M parameters, while reaching a throughput of 1,520 molecules per second on a single NVIDIA RTX 4090D GPU. Despite its compact design, it maintains competitive and, on several benchmarks, superior recognition accuracy.

0 citationsRead paper

Fluid-DiT: Graph-Free Diffusion Transformers for Fluid Flow Simulations Learning

Aug 07, 2026

High-fidelity fluid simulation is computationally expensive, and existing graph-based diffusion models are limited by handcrafted graph structures, restricted receptive fields, and difficulties in multiscale modeling. This work proposes a graph-free diffusion Transformer framework that eliminates explicit graph construction and leverages global attention mechanisms in a latent space to decouple geometric fidelity from distribution learning, enabling efficient modeling and sampling of the equilibrium distribution of chaotic flow fields. The method substantially reduces high-frequency artifacts and improves sampling efficiency. On benchmarks including laminar wakes, elliptic flows, and three-dimensional turbulent airfoils, it outperforms graph-based models in both sample quality and distributional accuracy, achieving higher R² correlation and lower Wasserstein distance, while generalizing effectively to unseen Reynolds numbers and geometric configurations.

0 citationsRead paper

AMPBench-MT: A Homology-Controlled Benchmark for Antimicrobial Peptide Potency, Spectrum, and Safety Prediction

Jul 28, 2026

Current computational assessments of antimicrobial peptides (AMPs) are largely confined to binary classification tasks and lack a unified, homology-controlled benchmark that jointly evaluates critical experimental endpoints such as potency, antimicrobial spectrum, hemolytic activity, toxicity, and selectivity. To address this gap, this work proposes AMPBench-MT, a comprehensive multi-task benchmark that integrates AMP identification, species-specific pMIC regression, and multiple safety-related endpoints under strict sequence homology control. Leveraging protein language model embeddings, graph neural networks, and a multi-task learning framework, the study systematically evaluates 161 endpoints and finds that frozen language model embeddings achieve superior performance in pMIC prediction. Importantly, the results demonstrate that high identification accuracy does not reliably translate to strong experimental performance, advocating for a paradigm shift in AMP evaluation from recognition-oriented approaches toward endpoint-aware, evidence-based auditing.

0 citationsRead paper