Institution profile

MS Ramaiah Institute of Technology

Academic institutionnorthamerica · us
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Investigating Thematic Patterns and User Preferences in LLM Interactions using BERTopic

Oct 08, 2025

This study investigates topic-level patterns between user prompts and LLM responses in the LMSYS-Chat-1M dataset and their correlation with human model preferences. Method: We pioneer the application of BERTopic to multilingual LLM comparative evaluation data, integrating dialogue cleaning, multilingual preprocessing, and topic distribution visualization to construct a model–topic preference matrix. Contribution/Results: We identify 29 semantically coherent topics and discover consistent user preference advantages for specific LLMs across domains such as technology, programming, and ethics—revealing a topic-dependent distribution of model strengths. This work establishes an interpretable, topic-level analytical framework for LLM capability assessment and enables domain-aware model selection and targeted fine-tuning, thereby advancing personalized LLM deployment.

0 citationsRead paper

MultiFluxAI Enhancing Platform Engineering with Advanced Agent-Orchestrated Retrieval Systems

Aug 28, 2025

To address the challenges of integrating massive, heterogeneous, cross-domain data in product engineering and supporting highly dynamic, context-dependent user service queries in digital ecosystems, this paper proposes a generative AI platform architecture based on agent orchestration. The architecture integrates vector embeddings, retrieval-augmented generation (RAG), and context-aware modeling, enabling semantic alignment and real-time fusion of multi-source data through a schedulable, collaborative agent mechanism. Its key innovation lies in embedding agent orchestration directly into the retrieval layer to enable situation-adaptive responses to complex queries. Experimental results demonstrate that the platform significantly improves query accuracy (+28.6%) and system scalability, while enabling seamless integration of legacy and new services—thereby enhancing user interaction efficiency and engagement depth within digital ecosystems.

0 citationsRead paper

Cardiovascular Disease Prediction using Machine Learning: A Comparative Analysis

Jul 29, 2025

This study addresses cardiovascular disease (CVD) risk prediction using a large-scale clinical dataset of 68,119 individuals. It systematically evaluates the impact of multiple risk factors—including age, blood pressure, cholesterol levels, smoking, and alcohol consumption—via statistical tests (t-tests, chi-square tests, and ANOVA) to identify significant associations. Notably, an unexpected negative correlation between smoking and alcohol use was detected, suggesting potential data bias and underscoring the need for careful preprocessing. Among several benchmark models, CatBoost achieved superior probabilistic calibration and discrimination: accuracy of 73.4%, Brier score of 0.1824, and expected calibration error (ECE) of only 0.0064—substantially outperforming logistic regression and other baselines. The work demonstrates CatBoost’s efficacy and reliability for CVD risk prediction while advancing interpretability and clinical trustworthiness through a hybrid statistical–machine learning framework that integrates rigorous hypothesis testing with high-performance modeling.

0 citationsRead paper

Solving Scene Understanding for Autonomous Navigation in Unstructured Environments

Jul 27, 2025

To address the challenge of scene understanding in unstructured road environments—particularly those characteristic of Indian urban and rural areas—this paper introduces a high-difficulty Indian driving dataset and proposes a multi-level semantic segmentation framework. Methodologically, we systematically evaluate five mainstream architectures—U-Net, U-Net+ResNet50, DeepLabV3, PSPNet, and SegNet—using mean Intersection-over-Union (mIoU) as the unified metric for pixel-wise drivable area, obstacle, and roadside object segmentation. Our contributions are threefold: (1) we present the first benchmark dataset specifically designed for typical Indian unstructured scenarios; (2) through cross-architecture analysis, we reveal differential robustness under challenging conditions including variable illumination, absent road signage, and ambiguous road boundaries; and (3) our framework achieves the state-of-the-art mIoU of 0.6496 on this dataset, significantly enhancing fine-grained scene parsing capability in complex traffic environments.

0 citationsRead paper
Recent publications

Latest Papers

Investigating Thematic Patterns and User Preferences in LLM Interactions using BERTopic

Oct 08, 2025

This study investigates topic-level patterns between user prompts and LLM responses in the LMSYS-Chat-1M dataset and their correlation with human model preferences. Method: We pioneer the application of BERTopic to multilingual LLM comparative evaluation data, integrating dialogue cleaning, multilingual preprocessing, and topic distribution visualization to construct a model–topic preference matrix. Contribution/Results: We identify 29 semantically coherent topics and discover consistent user preference advantages for specific LLMs across domains such as technology, programming, and ethics—revealing a topic-dependent distribution of model strengths. This work establishes an interpretable, topic-level analytical framework for LLM capability assessment and enables domain-aware model selection and targeted fine-tuning, thereby advancing personalized LLM deployment.

0 citationsRead paper

MultiFluxAI Enhancing Platform Engineering with Advanced Agent-Orchestrated Retrieval Systems

Aug 28, 2025

To address the challenges of integrating massive, heterogeneous, cross-domain data in product engineering and supporting highly dynamic, context-dependent user service queries in digital ecosystems, this paper proposes a generative AI platform architecture based on agent orchestration. The architecture integrates vector embeddings, retrieval-augmented generation (RAG), and context-aware modeling, enabling semantic alignment and real-time fusion of multi-source data through a schedulable, collaborative agent mechanism. Its key innovation lies in embedding agent orchestration directly into the retrieval layer to enable situation-adaptive responses to complex queries. Experimental results demonstrate that the platform significantly improves query accuracy (+28.6%) and system scalability, while enabling seamless integration of legacy and new services—thereby enhancing user interaction efficiency and engagement depth within digital ecosystems.

0 citationsRead paper

Cardiovascular Disease Prediction using Machine Learning: A Comparative Analysis

Jul 29, 2025

This study addresses cardiovascular disease (CVD) risk prediction using a large-scale clinical dataset of 68,119 individuals. It systematically evaluates the impact of multiple risk factors—including age, blood pressure, cholesterol levels, smoking, and alcohol consumption—via statistical tests (t-tests, chi-square tests, and ANOVA) to identify significant associations. Notably, an unexpected negative correlation between smoking and alcohol use was detected, suggesting potential data bias and underscoring the need for careful preprocessing. Among several benchmark models, CatBoost achieved superior probabilistic calibration and discrimination: accuracy of 73.4%, Brier score of 0.1824, and expected calibration error (ECE) of only 0.0064—substantially outperforming logistic regression and other baselines. The work demonstrates CatBoost’s efficacy and reliability for CVD risk prediction while advancing interpretability and clinical trustworthiness through a hybrid statistical–machine learning framework that integrates rigorous hypothesis testing with high-performance modeling.

0 citationsRead paper

Solving Scene Understanding for Autonomous Navigation in Unstructured Environments

Jul 27, 2025

To address the challenge of scene understanding in unstructured road environments—particularly those characteristic of Indian urban and rural areas—this paper introduces a high-difficulty Indian driving dataset and proposes a multi-level semantic segmentation framework. Methodologically, we systematically evaluate five mainstream architectures—U-Net, U-Net+ResNet50, DeepLabV3, PSPNet, and SegNet—using mean Intersection-over-Union (mIoU) as the unified metric for pixel-wise drivable area, obstacle, and roadside object segmentation. Our contributions are threefold: (1) we present the first benchmark dataset specifically designed for typical Indian unstructured scenarios; (2) through cross-architecture analysis, we reveal differential robustness under challenging conditions including variable illumination, absent road signage, and ambiguous road boundaries; and (3) our framework achieves the state-of-the-art mIoU of 0.6496 on this dataset, significantly enhancing fine-grained scene parsing capability in complex traffic environments.

0 citationsRead paper