Institution profile

Indian Institute of Technology Gandhinagar

Academic institutionasia · in
Official website
Research library101linked papers
Opportunities0open roles
Selected work

Representative Papers

Hand Shadow Art: A Differentiable Rendering Perspective

May 27, 2025

This work addresses the problem of computational hand shadow art generation. We propose the first differentiable rendering-based method for 3D hand deformation inversion: given a target 2D shadow image and illumination conditions, it jointly optimizes the geometry and pose of both hands along with lighting parameters to minimize the discrepancy between rendered and target shadows. Our approach integrates neural implicit hand representations, physically grounded shadow modeling, and a gradient-guided co-optimization framework—enabling simultaneous bilateral hand solving and smooth pose interpolation across semantically distinct shadows. Experiments demonstrate stable reconstruction of high-fidelity hand shadows, precise matching of intricate shadow structures, and seamless temporal transitions. This work establishes a novel paradigm and practical toolkit for applying differentiable graphics to digital artistic creation.

3 citationsRead paper

Where To Look? : Causal Tracing of Vision Encoders in VLM

Aug 11, 2026

This study investigates whether visual language models (VLMs) genuinely rely on visual information localized to target regions when generating answers. Employing causal tracing, the authors systematically analyze which token positions in the visual encoder exert causal influence on model outputs and evaluate the models’ capacity to understand visual structure under conditions where appearance cues are absent. The findings reveal that tokens with high causal impact are often distributed outside target regions, challenging the assumption that strong multimodal performance implies spatially localized causal representations. Furthermore, the work demonstrates that prevailing VLMs predominantly depend on superficial appearance cues to infer structural relationships, exposing a significant gap between their abilities to perceive, utilize, and reason about visual structure. This research establishes a causal framework for analyzing how visual information is transformed, preserved, and leveraged in VLMs.

0 citationsRead paper

Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation

Aug 06, 2026

Existing multilingual text embedding models typically employ a single training objective across diverse tasks, overlooking the fundamental differences in their optimization requirements. This work proposes the Task-Conditional Flow Matching (TCFM) framework, which introduces a task-conditional mechanism to tailor optimization objectives according to task-specific learning dynamics: leveraging flow matching for translation while designing more suitable objectives for retrieval, classification, and other tasks. The approach further integrates teacher-guided representations with a three-stage curriculum learning strategy to enable stable and efficient multitask adaptation. Evaluated on the Indic Massive Text Embedding Benchmark, TCFM achieves a new state of the art, significantly enhancing embedding quality across a wide range of multilingual tasks and demonstrating strong generalization across different model families.

0 citationsRead paper
Recent publications

Latest Papers

Where To Look? : Causal Tracing of Vision Encoders in VLM

Aug 11, 2026

This study investigates whether visual language models (VLMs) genuinely rely on visual information localized to target regions when generating answers. Employing causal tracing, the authors systematically analyze which token positions in the visual encoder exert causal influence on model outputs and evaluate the models’ capacity to understand visual structure under conditions where appearance cues are absent. The findings reveal that tokens with high causal impact are often distributed outside target regions, challenging the assumption that strong multimodal performance implies spatially localized causal representations. Furthermore, the work demonstrates that prevailing VLMs predominantly depend on superficial appearance cues to infer structural relationships, exposing a significant gap between their abilities to perceive, utilize, and reason about visual structure. This research establishes a causal framework for analyzing how visual information is transformed, preserved, and leveraged in VLMs.

0 citationsRead paper

Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation

Aug 06, 2026

Existing multilingual text embedding models typically employ a single training objective across diverse tasks, overlooking the fundamental differences in their optimization requirements. This work proposes the Task-Conditional Flow Matching (TCFM) framework, which introduces a task-conditional mechanism to tailor optimization objectives according to task-specific learning dynamics: leveraging flow matching for translation while designing more suitable objectives for retrieval, classification, and other tasks. The approach further integrates teacher-guided representations with a three-stage curriculum learning strategy to enable stable and efficient multitask adaptation. Evaluated on the Indic Massive Text Embedding Benchmark, TCFM achieves a new state of the art, significantly enhancing embedding quality across a wide range of multilingual tasks and demonstrating strong generalization across different model families.

0 citationsRead paper

Crushing the Evidence: A Dual-Penalty Evasion Framework for Fooling White-Box Explainable AI Auditors

Aug 01, 2026

This work addresses the vulnerability of existing explainable AI methods—such as LIME and SHAP—to adversarial attacks, which adversaries can exploit to conceal model biases or backdoors. The paper proposes a white-box gradient regularization evasion framework that, for the first time, embeds evasion logic directly into model parameters. By integrating a dual-penalty mechanism during training, the approach continuously suppresses gradients associated with trigger features, enabling the model to maintain accurate predictions while deceiving explanation systems. Notably, this method generates smooth, anomaly-free predictions without relying on out-of-distribution detection bypasses, thereby effectively evading conditional anomaly detection defenses. Evaluated on four benchmark tabular datasets, the technique reduces target feature attributions to below 0.02 and achieves attack success rates exceeding 90%, substantially outperforming current state-of-the-art approaches.

0 citationsRead paper