Institution profile

Narsee Monjee Institute of Management and Studies

Academic institutionnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

XEns-CKD: An Explainable Ensemble-Based Approach for Chronic Kidney Disease Stage Detection

Aug 02, 2026

This study addresses the challenge of early chronic kidney disease (CKD) detection, which is hindered by its asymptomatic onset and consequent delays in staging and intervention. To overcome this, the authors propose a novel approach that integrates an ensemble of Vision Transformers (ViTs) with multiple interpretable AI techniques to enable precise classification of ultrasound images across normal kidneys and all five CKD stages—a first in the field. The method leverages LIME, Layer-wise Relevance Propagation (LRP), and a newly introduced Attention-Min/Max fusion strategy to effectively localize and interpret diagnostically critical regions. Evaluated on a private renal ultrasound dataset, the model achieves an overall accuracy of 86.36%, outperforming existing methods by 4% and demonstrating superior performance across multiple macro-level metrics while maintaining both high diagnostic accuracy and strong interpretability.

0 citationsRead paper

Advancing SLM Tool-Use Capability using Reinforcement Learning

Sep 03, 2025

Small language models (SLMs) exhibit significantly weaker tool-use capabilities than large language models (LLMs), hindering their deployment in dynamic interactive applications such as virtual assistants and automated workflows. To address this, we propose the first framework applying Group Relative Policy Optimization (GRPO)—a reinforcement learning method—to train SLMs for tool use, eliminating reliance on large-scale annotated datasets and intensive computational resources typical of supervised fine-tuning. Our approach incorporates dynamic environment feedback and jointly optimizes multi-tool selection, parameter generation, and invocation sequence planning. Experiments across multiple tool-use benchmarks demonstrate an average 23.6% improvement in accuracy, substantially narrowing the functional gap with LLMs. Moreover, inference overhead is reduced by 68%, enabling practical SLM deployment in resource-constrained settings. This work establishes a scalable, RL-driven paradigm for enhancing SLM tool competence without sacrificing efficiency.

0 citationsRead paper

A Surveillance Based Interactive Robot

Aug 18, 2025

This study addresses the limited real-time speech interaction and environmental understanding capabilities of mobile surveillance robots by proposing a lightweight intelligent surveillance system based on a dual-Raspberry Pi architecture. Methodologically, it integrates FFmpeg-based video streaming, CPU-efficient YOLOv3 object detection, Kinect RGB-D environmental sensing, and multilingual end-to-end automatic speech recognition (ASR) and text-to-speech (TTS) modules; autonomous navigation and event response are enabled via voice command parsing and semantic mapping. Key contributions include: (i) the first implementation of a closed-loop speech–vision–motion control system on a low-cost, open-source platform; (ii) cross-lingual voice command comprehension and execution; and (iii) full reproducibility, low end-to-end latency (<300 ms), and zero human intervention. Indoor experiments demonstrate real-time performance, robustness, and a task completion rate exceeding 92%.

0 citationsRead paper
Recent publications

Latest Papers

XEns-CKD: An Explainable Ensemble-Based Approach for Chronic Kidney Disease Stage Detection

Aug 02, 2026

This study addresses the challenge of early chronic kidney disease (CKD) detection, which is hindered by its asymptomatic onset and consequent delays in staging and intervention. To overcome this, the authors propose a novel approach that integrates an ensemble of Vision Transformers (ViTs) with multiple interpretable AI techniques to enable precise classification of ultrasound images across normal kidneys and all five CKD stages—a first in the field. The method leverages LIME, Layer-wise Relevance Propagation (LRP), and a newly introduced Attention-Min/Max fusion strategy to effectively localize and interpret diagnostically critical regions. Evaluated on a private renal ultrasound dataset, the model achieves an overall accuracy of 86.36%, outperforming existing methods by 4% and demonstrating superior performance across multiple macro-level metrics while maintaining both high diagnostic accuracy and strong interpretability.

0 citationsRead paper

Advancing SLM Tool-Use Capability using Reinforcement Learning

Sep 03, 2025

Small language models (SLMs) exhibit significantly weaker tool-use capabilities than large language models (LLMs), hindering their deployment in dynamic interactive applications such as virtual assistants and automated workflows. To address this, we propose the first framework applying Group Relative Policy Optimization (GRPO)—a reinforcement learning method—to train SLMs for tool use, eliminating reliance on large-scale annotated datasets and intensive computational resources typical of supervised fine-tuning. Our approach incorporates dynamic environment feedback and jointly optimizes multi-tool selection, parameter generation, and invocation sequence planning. Experiments across multiple tool-use benchmarks demonstrate an average 23.6% improvement in accuracy, substantially narrowing the functional gap with LLMs. Moreover, inference overhead is reduced by 68%, enabling practical SLM deployment in resource-constrained settings. This work establishes a scalable, RL-driven paradigm for enhancing SLM tool competence without sacrificing efficiency.

0 citationsRead paper

A Surveillance Based Interactive Robot

Aug 18, 2025

This study addresses the limited real-time speech interaction and environmental understanding capabilities of mobile surveillance robots by proposing a lightweight intelligent surveillance system based on a dual-Raspberry Pi architecture. Methodologically, it integrates FFmpeg-based video streaming, CPU-efficient YOLOv3 object detection, Kinect RGB-D environmental sensing, and multilingual end-to-end automatic speech recognition (ASR) and text-to-speech (TTS) modules; autonomous navigation and event response are enabled via voice command parsing and semantic mapping. Key contributions include: (i) the first implementation of a closed-loop speech–vision–motion control system on a low-cost, open-source platform; (ii) cross-lingual voice command comprehension and execution; and (iii) full reproducibility, low end-to-end latency (<300 ms), and zero human intervention. Indoor experiments demonstrate real-time performance, robustness, and a task completion rate exceeding 92%.

0 citationsRead paper