Institution profile

Autonomous University of Nuevo Leon

Academic institutionnorthamerica · mx
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Enhancing Factual Accuracy and Citation Generation in LLMs via Multi-Stage Self-Verification

Sep 06, 2025

Large language models (LLMs) frequently generate hallucinated content and lack verifiable citations when producing fact-intensive text. To address this, we propose a multi-stage self-verification framework that orchestrates a sequential pipeline of *fact verification → reflective revision → citation integration*. The method synergistically combines chain-of-thought (CoT) reasoning with dual knowledge validation—leveraging both internal consistency checks and external authoritative source alignment—to dynamically perform fine-grained factual scrutiny during generation. When inconsistencies are detected, the model triggers reflective revision and automatically annotates traceable, context-aligned citations. Compared to state-of-the-art approaches, our framework substantially reduces hallucination rates while improving factual accuracy and citation reliability. Empirical evaluation demonstrates its effectiveness in high-fidelity applications such as scientific writing and news generation, where trustworthiness and evidential grounding are critical.

0 citationsRead paper

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation

Jul 05, 2025

Traditional text-to-image generation relies on large-scale manually annotated image-text pairs and struggles to precisely model fine-grained attributes and spatial relationships in complex prompts. To address this, we propose Hi-SSLVLM—a two-stage self-supervised large vision-language model that eliminates the need for manual annotations. It achieves fine-grained semantic decomposition and controllable generation via multi-granularity vision-language alignment, hierarchical captioning, and an internal compositional planning mechanism. We further introduce a semantic consistency loss and a self-refinement strategy to ensure cross-modal semantic fidelity and structural coherence. Extensive experiments on multiple benchmarks demonstrate that Hi-SSLVLM significantly outperforms state-of-the-art methods. Both quantitative metrics and human evaluations confirm its superiority in prompt fidelity, compositional accuracy, and aesthetic quality of generated images.

0 citationsRead paper

Efficient Few-Shot Medical Image Analysis via Hierarchical Contrastive Vision-Language Learning

Jan 16, 2025

To address the challenge of few-shot classification in medical imaging—particularly for chest X-ray and breast ultrasound—this paper proposes HiCA, a hierarchical contrastive alignment framework. First, domain-adaptive pretraining enhances the medical representation capability of large vision-language models (LVLMs). Second, a two-stage adaptive fine-tuning strategy is introduced, incorporating a novel hierarchical contrastive alignment mechanism that jointly optimizes vision–language cross-modal alignment at both feature-level and semantic-level. Third, a high-quality medical image–text paired dataset is constructed to support contrastive learning. Evaluated on ChestX-ray and Breast Ultrasound benchmarks, HiCA achieves state-of-the-art performance under both few-shot and zero-shot settings, significantly outperforming existing baselines in accuracy. Moreover, it demonstrates strong robustness, cross-modal interpretability, and promising clinical applicability.

0 citationsRead paper
Recent publications

Latest Papers

Enhancing Factual Accuracy and Citation Generation in LLMs via Multi-Stage Self-Verification

Sep 06, 2025

Large language models (LLMs) frequently generate hallucinated content and lack verifiable citations when producing fact-intensive text. To address this, we propose a multi-stage self-verification framework that orchestrates a sequential pipeline of *fact verification → reflective revision → citation integration*. The method synergistically combines chain-of-thought (CoT) reasoning with dual knowledge validation—leveraging both internal consistency checks and external authoritative source alignment—to dynamically perform fine-grained factual scrutiny during generation. When inconsistencies are detected, the model triggers reflective revision and automatically annotates traceable, context-aligned citations. Compared to state-of-the-art approaches, our framework substantially reduces hallucination rates while improving factual accuracy and citation reliability. Empirical evaluation demonstrates its effectiveness in high-fidelity applications such as scientific writing and news generation, where trustworthiness and evidential grounding are critical.

0 citationsRead paper

Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation

Jul 05, 2025

Traditional text-to-image generation relies on large-scale manually annotated image-text pairs and struggles to precisely model fine-grained attributes and spatial relationships in complex prompts. To address this, we propose Hi-SSLVLM—a two-stage self-supervised large vision-language model that eliminates the need for manual annotations. It achieves fine-grained semantic decomposition and controllable generation via multi-granularity vision-language alignment, hierarchical captioning, and an internal compositional planning mechanism. We further introduce a semantic consistency loss and a self-refinement strategy to ensure cross-modal semantic fidelity and structural coherence. Extensive experiments on multiple benchmarks demonstrate that Hi-SSLVLM significantly outperforms state-of-the-art methods. Both quantitative metrics and human evaluations confirm its superiority in prompt fidelity, compositional accuracy, and aesthetic quality of generated images.

0 citationsRead paper

Efficient Few-Shot Medical Image Analysis via Hierarchical Contrastive Vision-Language Learning

Jan 16, 2025

To address the challenge of few-shot classification in medical imaging—particularly for chest X-ray and breast ultrasound—this paper proposes HiCA, a hierarchical contrastive alignment framework. First, domain-adaptive pretraining enhances the medical representation capability of large vision-language models (LVLMs). Second, a two-stage adaptive fine-tuning strategy is introduced, incorporating a novel hierarchical contrastive alignment mechanism that jointly optimizes vision–language cross-modal alignment at both feature-level and semantic-level. Third, a high-quality medical image–text paired dataset is constructed to support contrastive learning. Evaluated on ChestX-ray and Breast Ultrasound benchmarks, HiCA achieves state-of-the-art performance under both few-shot and zero-shot settings, significantly outperforming existing baselines in accuracy. Moreover, it demonstrates strong robustness, cross-modal interpretability, and promising clinical applicability.

0 citationsRead paper