Institution profile

National Center for AI

Academic institution
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

EditCrafter: Tuning-free High-Resolution Image Editing via Pretrained Diffusion Model

Apr 11, 2026

Existing diffusion models struggle to edit high-resolution images with arbitrary aspect ratios or resolutions significantly exceeding their training scale (e.g., 512×512), as naive tiling often introduces structural distortions and content duplication. This work proposes a fine-tuning-free editing framework that integrates tiled latent-space inversion with an enhanced noise-damped classifier-free guidance strategy (NDCFG++). By effectively leveraging the generative priors of pre-trained text-to-image diffusion models, the method achieves coherent and photorealistic high-resolution edits while preserving image identity. It supports inputs of arbitrary dimensions and consistently produces structurally consistent and detail-rich results across diverse resolutions, without requiring model fine-tuning or additional optimization.

0 citationsRead paper

Quantigence: A Multi-Agent AI Framework for Quantum Security Research

Dec 15, 2025

Cryptographically relevant quantum computers (CRQCs) pose a structural threat to public-key infrastructure (PKI), as Shor’s and Grover’s algorithms could break widely deployed cryptographic primitives; the “store-now-decrypt-later” (SNDL) threat model necessitates urgent migration to post-quantum cryptography (PQC). Method: This work introduces the first multi-agent AI framework explicitly designed for PQC migration, integrating cryptographic analysis, threat modeling, standards alignment, and risk assessment. Contributions/Results: (1) A “cognitive parallelism” architecture ensures clean, isolated reasoning contexts across agents; (2) Quantum-Adapted Risk Scoring (QARS), a formal extension of Mosca’s theorem, quantifies migration urgency under quantum threat timelines; (3) A Model Context Protocol (MCP) enables dynamic knowledge injection and lightweight inference. Experiments demonstrate efficient execution on consumer-grade hardware (e.g., RTX 2060), reducing research cycle time by 67% and significantly outperforming manual workflows in both breadth and depth of literature coverage.

0 citationsRead paper

Let Multimodal Embedders Learn When to Augment Query via Adaptive Query Augmentation

Nov 04, 2025

Existing LLM-based embedders apply uniform query augmentation to all inputs, increasing latency and degrading performance on certain queries; no existing method supports adaptive query augmentation in multimodal settings. Method: We propose M-Solomon—the first general-purpose multimodal embedding model enabling adaptive query augmentation. It leverages conditional prefixes (/augment or /embed) to guide a multimodal large language model (MLLM) in dynamically determining whether a given query requires augmentation, generating synthetic augmented content only when necessary. Fine-grained control is achieved via prefix-guided generation and a two-stage training strategy. Results: Experiments demonstrate that M-Solomon significantly outperforms both non-augmented and uniformly augmented baselines, improving embedding quality while maintaining low latency. To our knowledge, it is the first approach to achieve efficient and precise adaptive augmentation in multimodal embedding.

0 citationsRead paper

VARCO-VISION-2.0 Technical Report

Sep 12, 2025

Existing bilingual vision-language models (VLMs) exhibit limitations in multi-image understanding and layout-aware OCR—i.e., joint modeling of textual content and its spatial coordinates. To address this, we propose VARCO-VISION-2.0, the first open-source Korean–English bilingual VLM. Methodologically, it introduces a novel four-stage curriculum learning strategy coupled with a memory-efficient training framework, integrating multimodal alignment enhancement, language capability preservation, safety optimization, and spatially aware OCR. Key contributions include: (1) robust bilingual understanding of complex multi-image inputs—including documents, charts, and tables; (2) release of two model variants—14B and 1.7B parameters—balancing high performance with edge-device deployability; and (3) strong empirical performance: the 14B model ranks eighth among same-scale models on the OpenCompass VLM leaderboard, with significant gains in bilingual spatial grounding accuracy.

0 citationsRead paper

Arabic Multimodal Machine Learning: Datasets, Applications, Approaches, and Challenges

Aug 16, 2025

The Arabic multimodal machine learning (MML) field lacks a systematic survey and structured taxonomy, leaving research gaps, critical bottlenecks, and future directions unclear. Method: We conduct a comprehensive literature review and multimodal (text/audio/visual) technical analysis to propose the first four-dimensional taxonomy for Arabic MML—covering datasets, application scenarios, modeling approaches, and core challenges. Contribution/Results: Our taxonomy reveals underexplored directions, including cross-modal alignment, low-resource robustness, and culturally adaptive modeling, while identifying key bottlenecks: data scarcity, inconsistent annotation practices, and absent standardized evaluation protocols. The framework delivers a structured knowledge graph and a reproducible research roadmap for Arabic MML, significantly enhancing the field’s conceptual clarity, methodological rigor, and scalability. This work establishes foundational infrastructure to accelerate principled, culturally grounded advances in Arabic multimodal AI.

0 citationsRead paper
Recent publications

Latest Papers

EditCrafter: Tuning-free High-Resolution Image Editing via Pretrained Diffusion Model

Apr 11, 2026

Existing diffusion models struggle to edit high-resolution images with arbitrary aspect ratios or resolutions significantly exceeding their training scale (e.g., 512×512), as naive tiling often introduces structural distortions and content duplication. This work proposes a fine-tuning-free editing framework that integrates tiled latent-space inversion with an enhanced noise-damped classifier-free guidance strategy (NDCFG++). By effectively leveraging the generative priors of pre-trained text-to-image diffusion models, the method achieves coherent and photorealistic high-resolution edits while preserving image identity. It supports inputs of arbitrary dimensions and consistently produces structurally consistent and detail-rich results across diverse resolutions, without requiring model fine-tuning or additional optimization.

0 citationsRead paper

Quantigence: A Multi-Agent AI Framework for Quantum Security Research

Dec 15, 2025

Cryptographically relevant quantum computers (CRQCs) pose a structural threat to public-key infrastructure (PKI), as Shor’s and Grover’s algorithms could break widely deployed cryptographic primitives; the “store-now-decrypt-later” (SNDL) threat model necessitates urgent migration to post-quantum cryptography (PQC). Method: This work introduces the first multi-agent AI framework explicitly designed for PQC migration, integrating cryptographic analysis, threat modeling, standards alignment, and risk assessment. Contributions/Results: (1) A “cognitive parallelism” architecture ensures clean, isolated reasoning contexts across agents; (2) Quantum-Adapted Risk Scoring (QARS), a formal extension of Mosca’s theorem, quantifies migration urgency under quantum threat timelines; (3) A Model Context Protocol (MCP) enables dynamic knowledge injection and lightweight inference. Experiments demonstrate efficient execution on consumer-grade hardware (e.g., RTX 2060), reducing research cycle time by 67% and significantly outperforming manual workflows in both breadth and depth of literature coverage.

0 citationsRead paper

Let Multimodal Embedders Learn When to Augment Query via Adaptive Query Augmentation

Nov 04, 2025

Existing LLM-based embedders apply uniform query augmentation to all inputs, increasing latency and degrading performance on certain queries; no existing method supports adaptive query augmentation in multimodal settings. Method: We propose M-Solomon—the first general-purpose multimodal embedding model enabling adaptive query augmentation. It leverages conditional prefixes (/augment or /embed) to guide a multimodal large language model (MLLM) in dynamically determining whether a given query requires augmentation, generating synthetic augmented content only when necessary. Fine-grained control is achieved via prefix-guided generation and a two-stage training strategy. Results: Experiments demonstrate that M-Solomon significantly outperforms both non-augmented and uniformly augmented baselines, improving embedding quality while maintaining low latency. To our knowledge, it is the first approach to achieve efficient and precise adaptive augmentation in multimodal embedding.

0 citationsRead paper

VARCO-VISION-2.0 Technical Report

Sep 12, 2025

Existing bilingual vision-language models (VLMs) exhibit limitations in multi-image understanding and layout-aware OCR—i.e., joint modeling of textual content and its spatial coordinates. To address this, we propose VARCO-VISION-2.0, the first open-source Korean–English bilingual VLM. Methodologically, it introduces a novel four-stage curriculum learning strategy coupled with a memory-efficient training framework, integrating multimodal alignment enhancement, language capability preservation, safety optimization, and spatially aware OCR. Key contributions include: (1) robust bilingual understanding of complex multi-image inputs—including documents, charts, and tables; (2) release of two model variants—14B and 1.7B parameters—balancing high performance with edge-device deployability; and (3) strong empirical performance: the 14B model ranks eighth among same-scale models on the OpenCompass VLM leaderboard, with significant gains in bilingual spatial grounding accuracy.

0 citationsRead paper

Arabic Multimodal Machine Learning: Datasets, Applications, Approaches, and Challenges

Aug 16, 2025

The Arabic multimodal machine learning (MML) field lacks a systematic survey and structured taxonomy, leaving research gaps, critical bottlenecks, and future directions unclear. Method: We conduct a comprehensive literature review and multimodal (text/audio/visual) technical analysis to propose the first four-dimensional taxonomy for Arabic MML—covering datasets, application scenarios, modeling approaches, and core challenges. Contribution/Results: Our taxonomy reveals underexplored directions, including cross-modal alignment, low-resource robustness, and culturally adaptive modeling, while identifying key bottlenecks: data scarcity, inconsistent annotation practices, and absent standardized evaluation protocols. The framework delivers a structured knowledge graph and a reproducible research roadmap for Arabic MML, significantly enhancing the field’s conceptual clarity, methodological rigor, and scalability. This work establishes foundational infrastructure to accelerate principled, culturally grounded advances in Arabic multimodal AI.

0 citationsRead paper