Institution profile

Indian Institute of Technology Patna

Academic institutionasia · in
Official website
Research library100linked papers
Opportunities0open roles
Selected work

Representative Papers

Improving Neural Question Generation using World Knowledge

Sep 09, 2019arXiv.org

Neural question generation suffers from weak entity semantic representation and produces questions with limited human readability and semantic plausibility. Method: This paper introduces structured world knowledge—specifically Wikidata-linked entities and their fine-grained types—into an encoder-decoder framework for the first time. It jointly models entity linking, type-aware embeddings, and hierarchical attention to enhance semantic understanding of key passage entities and improve controlled question generation. The model is trained end-to-end on SQuAD and MS MARCO. Results: Experiments show substantial improvements in generation quality, with absolute BLEU-4 gains of +1.37 and +1.59 over strong baselines. The core contribution is a novel, interpretable, and scalable paradigm for injecting external world knowledge, advancing semantic fidelity and linguistic naturalness in question generation.

7 citationsRead paper

Optimal Designs in Multicomponent Stress Strength Reliability for the Unit Generalized Rayleigh Distribution

Aug 06, 2026

This study addresses the challenge of stress-strength reliability assessment for multi-component systems under progressive Type-II censoring by developing a unified inferential framework that integrates maximum likelihood estimation, maximum product spacing, and Bayesian methods under the assumption of unit generalized Rayleigh distributions. The work innovatively proposes an optimal progressive censoring scheme based on three optimality criteria and enhances estimation accuracy through the synergistic use of the EM algorithm, Fisher information matrix, missing information principle, Markov chain Monte Carlo (MCMC) techniques, and optimal experimental design. Extensive simulation studies and real-data analysis demonstrate that the proposed methodology yields highly accurate reliability estimates and credible/confidence intervals while effectively identifying efficient censoring plans.

0 citationsRead paper

Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations

Aug 06, 2026

This work addresses the limitations of conventional top-k embedding-based retrieval methods when applied to structured long documents such as financial reports, where chunking often severs critical contextual relationships—such as those between numerical values, units, and fiscal year headings—leading to loss of essential information. To overcome this, the authors propose READ, a novel framework that abandons the embedding-retrieval paradigm entirely and instead introduces a deterministic, replayable agent-driven mechanism. READ performs normalized lexical search, structural navigation, and bounded-span reading directly on the original document to retrieve information and generate traceable audit trails. Evaluated on 51 verification questions, READ achieves an accuracy of 58.8%, significantly outperforming dense retrieval (15.7%, p = 2×10⁻⁵), thereby demonstrating the efficacy of embedding-free approaches for precise and interpretable information extraction from structured documents.

0 citationsRead paper

Optimum Multiple Sampling Plan Based on the Process Capability Index $C_{py}$ Under Type-II Hybrid Censoring

Jul 27, 2026

This study addresses the inefficiency of traditional multiple sampling plans, which rely on prior inspection outcomes and often require large sample sizes, making it difficult to balance both producer and consumer risks with economic feasibility. To overcome these limitations, this work proposes a Stage-Independent Multiple Sampling Plan (SIMSP) that, for the first time, integrates the process capability index $C_{py}$ with a Type-II hybrid censoring scheme. A constrained optimization model is formulated to minimize total inspection cost, leveraging the asymptotic distribution of $C_{py}$ to derive the operating characteristic function and employing the exact Fisher information matrix to determine the optimal design. Numerical experiments demonstrate that the proposed SIMSP significantly reduces required sample sizes and total costs while rigorously controlling both manufacturer and consumer risks, offering an efficient and practical approach for warranty acceptance testing of non-repairable products.

0 citationsRead paper

A Step Towards Robust Unsupervised Domain Adaptation via Fine-Tuning and Reinforcement Learning

Jul 03, 2026

This work addresses the insufficient adversarial robustness in unsupervised domain adaptation caused by noisy pseudo-labels and source–target distribution shifts. To tackle this, the authors propose a two-stage SFT+RL framework built upon the CLIP vision encoder. In the first stage, they perform PGD-based adversarial fine-tuning of the linear classifier while partially unfreezing the projection layer. The second stage introduces a confidence-guided progressive pseudo-labeling strategy that employs a dynamically decaying threshold to select high-confidence target samples, which are then used to construct a hybrid dataset for reinforcement learning–driven adversarial training. This approach preserves CLIP’s semantic priors while substantially enhancing cross-domain robustness and accuracy, achieving an average improvement of 10.2% in clean accuracy and 15.8% in adversarial robustness over state-of-the-art methods on OfficeHome, PACS, and VisDA benchmarks.

0 citationsRead paper
Recent publications

Latest Papers

Optimal Designs in Multicomponent Stress Strength Reliability for the Unit Generalized Rayleigh Distribution

Aug 06, 2026

This study addresses the challenge of stress-strength reliability assessment for multi-component systems under progressive Type-II censoring by developing a unified inferential framework that integrates maximum likelihood estimation, maximum product spacing, and Bayesian methods under the assumption of unit generalized Rayleigh distributions. The work innovatively proposes an optimal progressive censoring scheme based on three optimality criteria and enhances estimation accuracy through the synergistic use of the EM algorithm, Fisher information matrix, missing information principle, Markov chain Monte Carlo (MCMC) techniques, and optimal experimental design. Extensive simulation studies and real-data analysis demonstrate that the proposed methodology yields highly accurate reliability estimates and credible/confidence intervals while effectively identifying efficient censoring plans.

0 citationsRead paper

Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations

Aug 06, 2026

This work addresses the limitations of conventional top-k embedding-based retrieval methods when applied to structured long documents such as financial reports, where chunking often severs critical contextual relationships—such as those between numerical values, units, and fiscal year headings—leading to loss of essential information. To overcome this, the authors propose READ, a novel framework that abandons the embedding-retrieval paradigm entirely and instead introduces a deterministic, replayable agent-driven mechanism. READ performs normalized lexical search, structural navigation, and bounded-span reading directly on the original document to retrieve information and generate traceable audit trails. Evaluated on 51 verification questions, READ achieves an accuracy of 58.8%, significantly outperforming dense retrieval (15.7%, p = 2×10⁻⁵), thereby demonstrating the efficacy of embedding-free approaches for precise and interpretable information extraction from structured documents.

0 citationsRead paper

Optimum Multiple Sampling Plan Based on the Process Capability Index $C_{py}$ Under Type-II Hybrid Censoring

Jul 27, 2026

This study addresses the inefficiency of traditional multiple sampling plans, which rely on prior inspection outcomes and often require large sample sizes, making it difficult to balance both producer and consumer risks with economic feasibility. To overcome these limitations, this work proposes a Stage-Independent Multiple Sampling Plan (SIMSP) that, for the first time, integrates the process capability index $C_{py}$ with a Type-II hybrid censoring scheme. A constrained optimization model is formulated to minimize total inspection cost, leveraging the asymptotic distribution of $C_{py}$ to derive the operating characteristic function and employing the exact Fisher information matrix to determine the optimal design. Numerical experiments demonstrate that the proposed SIMSP significantly reduces required sample sizes and total costs while rigorously controlling both manufacturer and consumer risks, offering an efficient and practical approach for warranty acceptance testing of non-repairable products.

0 citationsRead paper

A Step Towards Robust Unsupervised Domain Adaptation via Fine-Tuning and Reinforcement Learning

Jul 03, 2026

This work addresses the insufficient adversarial robustness in unsupervised domain adaptation caused by noisy pseudo-labels and source–target distribution shifts. To tackle this, the authors propose a two-stage SFT+RL framework built upon the CLIP vision encoder. In the first stage, they perform PGD-based adversarial fine-tuning of the linear classifier while partially unfreezing the projection layer. The second stage introduces a confidence-guided progressive pseudo-labeling strategy that employs a dynamically decaying threshold to select high-confidence target samples, which are then used to construct a hybrid dataset for reinforcement learning–driven adversarial training. This approach preserves CLIP’s semantic priors while substantially enhancing cross-domain robustness and accuracy, achieving an average improvement of 10.2% in clean accuracy and 15.8% in adversarial robustness over state-of-the-art methods on OfficeHome, PACS, and VisDA benchmarks.

0 citationsRead paper

GPUAlert: A Zero-Instrumentation Process-Boundary Monitor for Diagnosing GPU Training-Job Failures

Jul 01, 2026

This work addresses the frequent failure of GPU training tasks due to hardware faults, exacerbated by delayed operational responses and the absence of automated diagnostics. The authors propose a non-intrusive command-line wrapper that monitors job execution at process boundaries without modifying training scripts, automatically dispatching structured email notifications containing categorized failure reasons, persisted logs, and output artifacts. The core contributions are three reliability primitives: pre-launch log assurance, notifier isolation, and non-silent artifact budgeting, which collectively guarantee log durability, accurate exit codes, and controlled attachment sizes. Evaluation across 12 reproducible hardware fault types demonstrates a macro F1 score of 0.997—significantly outperforming keyword matching (0.830) and exit-code inspection (0.133)—with only ~3 ms overhead per task and robust exit-code reporting even when SMTP is unreachable.

0 citationsRead paper