Institution profile

Yellow.ai

Industry researchasia · in
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

PRISM: Prompt Reliability via Iterative Simulation and Monitoring for Enterprise Conversational AI

May 15, 2026

This work addresses the challenge of behavioral drift in large language models (LLMs) within enterprise conversational AI systems, which often renders deployed prompts ineffective post-deployment due to the absence of continuous monitoring and remediation mechanisms. The study reframes prompt engineering as a continuous reliability engineering problem and introduces the first closed-loop, automated framework tailored for enterprise settings. The proposed approach automatically generates test cases from natural language requirements, employs high-fidelity multi-turn dialogue simulation, and integrates LLM-as-a-judge evaluation with root-cause diagnosis to enable precise, “surgical” prompt repairs. Evaluated across 35 enterprise conversational agents, the method reduces prompt development time from two days to under 30 minutes, achieves 99% production reliability, and effectively detects and rectifies regression issues caused by model drift within 24 hours.

0 citationsRead paper

Vision-Guided Chunking Is All You Need: Enhancing RAG with Multimodal Document Understanding

Jun 19, 2025

Traditional text chunking methods struggle with multi-page tables, embedded figures, and cross-page semantic dependencies in PDFs, degrading RAG performance. This paper proposes a large multimodal model (LMM)-based, batch-aware PDF chunking method featuring a novel vision-guided cross-page batching mechanism. It employs multimodal PDF parsing for joint textual and visual understanding, enabling multi-page table alignment and structural context modeling, while incorporating configurable page batching and cross-batch context preservation. Evaluated on a manually curated PDF question-answering dataset, our method significantly improves chunk structural integrity and semantic coherence, yielding substantially higher RAG accuracy than baseline approaches. The core contribution lies in transcending the pure-text chunking paradigm by deeply integrating visual perception with cross-page batch processing—establishing a new framework for RAG over complex, layout-rich documents.

0 citationsRead paper

Historic Scripts to Modern Vision: A Novel Dataset and A VLM Framework for Transliteration of Modi Script to Devanagari

Mar 17, 2025

To address the urgent need for digitizing endangered medieval Indian Modi-script Marathi manuscripts, this paper introduces MoScNet, the first end-to-end vision-language model for Modi-to-Devanagari script transcription. We present MoDeTrans, the first publicly available, paired benchmark dataset comprising 2,043 handwritten Modi images with accurate Devanagari transcriptions. Methodologically, MoScNet employs a hybrid CNN-Transformer encoder and a sequence-to-sequence decoder, augmented by OCR-specific pretraining and a novel knowledge distillation–driven lightweight architecture: the student model retains only 1/163 of the teacher’s parameters while surpassing its performance. Experiments demonstrate state-of-the-art accuracy on Modi transcription and competitive results on general OCR benchmarks. This work establishes a scalable, high-precision paradigm for historical manuscript digitization.

0 citationsRead paper
Recent publications

Latest Papers

PRISM: Prompt Reliability via Iterative Simulation and Monitoring for Enterprise Conversational AI

May 15, 2026

This work addresses the challenge of behavioral drift in large language models (LLMs) within enterprise conversational AI systems, which often renders deployed prompts ineffective post-deployment due to the absence of continuous monitoring and remediation mechanisms. The study reframes prompt engineering as a continuous reliability engineering problem and introduces the first closed-loop, automated framework tailored for enterprise settings. The proposed approach automatically generates test cases from natural language requirements, employs high-fidelity multi-turn dialogue simulation, and integrates LLM-as-a-judge evaluation with root-cause diagnosis to enable precise, “surgical” prompt repairs. Evaluated across 35 enterprise conversational agents, the method reduces prompt development time from two days to under 30 minutes, achieves 99% production reliability, and effectively detects and rectifies regression issues caused by model drift within 24 hours.

0 citationsRead paper

Vision-Guided Chunking Is All You Need: Enhancing RAG with Multimodal Document Understanding

Jun 19, 2025

Traditional text chunking methods struggle with multi-page tables, embedded figures, and cross-page semantic dependencies in PDFs, degrading RAG performance. This paper proposes a large multimodal model (LMM)-based, batch-aware PDF chunking method featuring a novel vision-guided cross-page batching mechanism. It employs multimodal PDF parsing for joint textual and visual understanding, enabling multi-page table alignment and structural context modeling, while incorporating configurable page batching and cross-batch context preservation. Evaluated on a manually curated PDF question-answering dataset, our method significantly improves chunk structural integrity and semantic coherence, yielding substantially higher RAG accuracy than baseline approaches. The core contribution lies in transcending the pure-text chunking paradigm by deeply integrating visual perception with cross-page batch processing—establishing a new framework for RAG over complex, layout-rich documents.

0 citationsRead paper

Historic Scripts to Modern Vision: A Novel Dataset and A VLM Framework for Transliteration of Modi Script to Devanagari

Mar 17, 2025

To address the urgent need for digitizing endangered medieval Indian Modi-script Marathi manuscripts, this paper introduces MoScNet, the first end-to-end vision-language model for Modi-to-Devanagari script transcription. We present MoDeTrans, the first publicly available, paired benchmark dataset comprising 2,043 handwritten Modi images with accurate Devanagari transcriptions. Methodologically, MoScNet employs a hybrid CNN-Transformer encoder and a sequence-to-sequence decoder, augmented by OCR-specific pretraining and a novel knowledge distillation–driven lightweight architecture: the student model retains only 1/163 of the teacher’s parameters while surpassing its performance. Experiments demonstrate state-of-the-art accuracy on Modi transcription and competitive results on general OCR benchmarks. This work establishes a scalable, high-precision paradigm for historical manuscript digitization.

0 citationsRead paper