Institution profile

Nutanix

Industry researchnorthamerica · us
Official website
Research library17linked papers
Opportunities0open roles
Selected work

Representative Papers

DeltaServe: Host-Agnostic Co-Serving of Inference and Fine-Tuning for LLMs

Jul 30, 2026

This work addresses the underutilization of GPU resources in large language model inference systems, which are often over-provisioned to handle peak loads, leaving spare capacity idle during off-peak periods. To this end, we propose DeltaServe, a host-agnostic co-serving architecture that dynamically repurposes idle compute for LoRA fine-tuning tasks while strictly meeting inference service-level objectives (SLOs). DeltaServe’s key innovation lies in sharing the execution structure between inference prefilling and LoRA forward passes, coupled with an SLO-aware scheduler that enables efficient co-location without additional hardware overhead. The system integrates lightweight hooks supporting multi-LoRA batching and CUDA Graph–based latency modeling, offering compatibility with mainstream serving engines such as vLLM and SGLang. Experiments on real-world production workloads show that DeltaServe achieves 2.9× higher fine-tuning throughput than LLMStation and 39% improvement over a vLLM+torchtune baseline, all while maintaining 100% SLO compliance for inference requests.

0 citationsRead paper

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability

Jul 07, 2026

Current AI safety evaluation frameworks are predominantly grounded in a Western-centric perspective, often overlooking regional legal, linguistic, and cultural differences, thereby introducing security vulnerabilities when deploying vision-language models globally. This work proposes the first multimodal, multilingual evaluation framework centered on cultural appropriateness, encompassing six Asia-Pacific countries and eight languages. By natively collecting localized image-text risk samples, the framework jointly surfaces both universally prohibited content and culturally sensitive issues. We introduce multimodal prompting strategies, construct an empirically driven cultural taxonomy, and integrate a Judge-Pluralis ensemble adjudication mechanism. Our findings reveal systemic model deficiencies in cross-cultural contexts—including image misinterpretation, lack of regional contextual awareness, and failure of refusal mechanisms—highlighting critical evaluation blind spots obscured by global aggregate metrics.

0 citationsRead paper

Predictable GRPO: A Closed-Form Model of Training Dynamics

Jun 29, 2026

This work addresses the lack of theoretical understanding of GRPO training dynamics, which currently rely on empirical hyperparameter tuning. By leveraging first-principles reasoning, the authors formulate GRPO dynamics as a physical potential system influenced by an inertia term. Through dimensionality reduction, mean-field approximation, and softmax-bandit simplification, they derive a closed-form trajectory model that elucidates both the overdamped limit and oscillatory transition mechanisms. This model serves as a diagnostic tool capable of distinguishing among multiple failure modes. Empirical validation across three models and two group sizes demonstrates reward trajectory fits with R² ≥ 0.91, confirming group-size invariance and the predictability of stability thresholds.

0 citationsRead paper

Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets

May 14, 2026

This work addresses the limitations of existing Croissant metadata generation approaches, which rely on public platforms and struggle to accommodate governed or large-scale local datasets. The authors propose the first open-source, local-first command-line tool that directly generates Croissant-compliant JSON-LD metadata from local directories via a modular processor registration mechanism, supporting mainstream formats such as Parquet. By eliminating dependence on external platforms, this method significantly enhances the discoverability and reusability of private, high-value datasets. Experimental evaluation across more than 140 datasets—including MIMIC-IV with 886 million rows—demonstrates that the generated metadata achieves 97–100% accuracy, matching or exceeding that of manual curation or standard methods.

0 citationsRead paper
Recent publications

Latest Papers

DeltaServe: Host-Agnostic Co-Serving of Inference and Fine-Tuning for LLMs

Jul 30, 2026

This work addresses the underutilization of GPU resources in large language model inference systems, which are often over-provisioned to handle peak loads, leaving spare capacity idle during off-peak periods. To this end, we propose DeltaServe, a host-agnostic co-serving architecture that dynamically repurposes idle compute for LoRA fine-tuning tasks while strictly meeting inference service-level objectives (SLOs). DeltaServe’s key innovation lies in sharing the execution structure between inference prefilling and LoRA forward passes, coupled with an SLO-aware scheduler that enables efficient co-location without additional hardware overhead. The system integrates lightweight hooks supporting multi-LoRA batching and CUDA Graph–based latency modeling, offering compatibility with mainstream serving engines such as vLLM and SGLang. Experiments on real-world production workloads show that DeltaServe achieves 2.9× higher fine-tuning throughput than LLMStation and 39% improvement over a vLLM+torchtune baseline, all while maintaining 100% SLO compliance for inference requests.

0 citationsRead paper

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability

Jul 07, 2026

Current AI safety evaluation frameworks are predominantly grounded in a Western-centric perspective, often overlooking regional legal, linguistic, and cultural differences, thereby introducing security vulnerabilities when deploying vision-language models globally. This work proposes the first multimodal, multilingual evaluation framework centered on cultural appropriateness, encompassing six Asia-Pacific countries and eight languages. By natively collecting localized image-text risk samples, the framework jointly surfaces both universally prohibited content and culturally sensitive issues. We introduce multimodal prompting strategies, construct an empirically driven cultural taxonomy, and integrate a Judge-Pluralis ensemble adjudication mechanism. Our findings reveal systemic model deficiencies in cross-cultural contexts—including image misinterpretation, lack of regional contextual awareness, and failure of refusal mechanisms—highlighting critical evaluation blind spots obscured by global aggregate metrics.

0 citationsRead paper

Predictable GRPO: A Closed-Form Model of Training Dynamics

Jun 29, 2026

This work addresses the lack of theoretical understanding of GRPO training dynamics, which currently rely on empirical hyperparameter tuning. By leveraging first-principles reasoning, the authors formulate GRPO dynamics as a physical potential system influenced by an inertia term. Through dimensionality reduction, mean-field approximation, and softmax-bandit simplification, they derive a closed-form trajectory model that elucidates both the overdamped limit and oscillatory transition mechanisms. This model serves as a diagnostic tool capable of distinguishing among multiple failure modes. Empirical validation across three models and two group sizes demonstrates reward trajectory fits with R² ≥ 0.91, confirming group-size invariance and the predictability of stability thresholds.

0 citationsRead paper

Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets

May 14, 2026

This work addresses the limitations of existing Croissant metadata generation approaches, which rely on public platforms and struggle to accommodate governed or large-scale local datasets. The authors propose the first open-source, local-first command-line tool that directly generates Croissant-compliant JSON-LD metadata from local directories via a modular processor registration mechanism, supporting mainstream formats such as Parquet. By eliminating dependence on external platforms, this method significantly enhances the discoverability and reusability of private, high-value datasets. Experimental evaluation across more than 140 datasets—including MIMIC-IV with 886 million rows—demonstrates that the generated metadata achieves 97–100% accuracy, matching or exceeding that of manual curation or standard methods.

0 citationsRead paper