Institution profile

Guangzhou Laboratory

Academic institutionasia · cn
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

A Specialized Large Language Model for Clinical Reasoning and Diagnosis in Rare Diseases

Nov 18, 2025

Rare diseases suffer from prolonged diagnostic timelines and fragmented clinical evidence; moreover, general-purpose large language models exhibit limited clinical reasoning capabilities due to scarce real-world electronic health records (EHRs), outdated medical knowledge, and hallucination. To address these challenges, we propose a domain-specific clinical reasoning paradigm centered on “narrative-first, knowledge-enhanced” inference. Our approach comprises: (1) constructing a physician-validated rare-disease reasoning dataset and domain-specific corpus; (2) designing a knowledge graph–anchored retrieval mechanism and phased chain-of-thought training to integrate non-phenotypic evidence (e.g., imaging, functional tests); and (3) enhancing robustness under noisy conditions and phenotypic overlap via instruction fine-tuning, knowledge graph fusion, and structured reasoning. Evaluated on multicenter real-world EHRs and public benchmarks, our method achieves state-of-the-art performance—matching the diagnostic accuracy of senior clinicians while significantly shortening diagnostic pathways and enabling transparent, auditable clinical decision-making.

0 citationsRead paper

Decoding Translation-Related Functional Sequences in 5'UTRs Using Interpretable Deep Learning Models

Jul 22, 2025

Existing 5′UTR translation efficiency prediction models suffer from fixed-length input constraints and limited interpretability. To address these limitations, we propose UTR-STCNet—a novel, interpretable deep learning architecture designed for variable-length sequence modeling. It integrates saliency-aware token clustering with a lightweight saliency-guided Transformer to enable multi-scale semantic aggregation and capture both local and long-range dependencies. By eliminating the need for sequence truncation, UTR-STCNet balances computational efficiency with biological interpretability. On three benchmark datasets, UTR-STCNet consistently outperforms state-of-the-art methods in predicting ribosomal load. Moreover, it successfully identifies key regulatory motifs—including upstream AUGs and Kozak sequences—demonstrating its capacity for biologically meaningful pattern discovery. This work establishes a new paradigm for functional interpretation and rational design of 5′UTRs.

0 citationsRead paper
Recent publications

Latest Papers

A Specialized Large Language Model for Clinical Reasoning and Diagnosis in Rare Diseases

Nov 18, 2025

Rare diseases suffer from prolonged diagnostic timelines and fragmented clinical evidence; moreover, general-purpose large language models exhibit limited clinical reasoning capabilities due to scarce real-world electronic health records (EHRs), outdated medical knowledge, and hallucination. To address these challenges, we propose a domain-specific clinical reasoning paradigm centered on “narrative-first, knowledge-enhanced” inference. Our approach comprises: (1) constructing a physician-validated rare-disease reasoning dataset and domain-specific corpus; (2) designing a knowledge graph–anchored retrieval mechanism and phased chain-of-thought training to integrate non-phenotypic evidence (e.g., imaging, functional tests); and (3) enhancing robustness under noisy conditions and phenotypic overlap via instruction fine-tuning, knowledge graph fusion, and structured reasoning. Evaluated on multicenter real-world EHRs and public benchmarks, our method achieves state-of-the-art performance—matching the diagnostic accuracy of senior clinicians while significantly shortening diagnostic pathways and enabling transparent, auditable clinical decision-making.

0 citationsRead paper

Decoding Translation-Related Functional Sequences in 5'UTRs Using Interpretable Deep Learning Models

Jul 22, 2025

Existing 5′UTR translation efficiency prediction models suffer from fixed-length input constraints and limited interpretability. To address these limitations, we propose UTR-STCNet—a novel, interpretable deep learning architecture designed for variable-length sequence modeling. It integrates saliency-aware token clustering with a lightweight saliency-guided Transformer to enable multi-scale semantic aggregation and capture both local and long-range dependencies. By eliminating the need for sequence truncation, UTR-STCNet balances computational efficiency with biological interpretability. On three benchmark datasets, UTR-STCNet consistently outperforms state-of-the-art methods in predicting ribosomal load. Moreover, it successfully identifies key regulatory motifs—including upstream AUGs and Kozak sequences—demonstrating its capacity for biologically meaningful pattern discovery. This work establishes a new paradigm for functional interpretation and rational design of 5′UTRs.

0 citationsRead paper