Institution profile

Walmart Global Tech

Industry researchnorthamerica · us
Official website
Research library67linked papers
Opportunities0open roles
Selected work

Representative Papers

Attribute-Conditioned Multimodal Slot Factorization for Controllable Fashion Retrieval

Aug 12, 2026

This work addresses the limitation of conventional fashion retrieval methods, which conflate multi-attribute information into a single embedding, thereby hindering fine-grained control over specific attributes such as color or category. To overcome this, the authors propose the MM-slotgate model, which introduces— for the first time—a named and intervenable multimodal attribute-slot architecture. This framework decomposes Fashion-CLIP’s joint image-text embeddings into four semantically distinct attribute slots, each equipped with an independent image-text gating mechanism that automatically learns modality weights without explicit supervision. The approach achieves both semantic interpretability and controllable retrieval performance, attaining a ConstraintSatisfied@10 score of 0.7566 on the H&M dataset, improving color retrieval success by 0.568, and yielding a 15.3× enhancement in color intervention efficacy through quantized slot representations.

0 citationsRead paper

Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy

Aug 12, 2026

This study addresses the challenge global retailers face in manually maintaining price consistency and constructing coherent price hierarchies across million-scale product catalogs. To this end, the work proposes a scalable, context-aware multi-agent framework that introduces, for the first time, a multi-agent architecture to the construction of retail price classification systems—commonly known as “Lines and Ladders.” The framework orchestrates multiple specialized large language model agents to collaboratively perform product attribute recognition, multimodal value extraction, and hierarchical grouping, thereby mitigating cognitive overload inherent in single-agent approaches. Experimental results on real-world enterprise data demonstrate that a three-agent system achieves an F1 score of 0.83 on Lines tasks, with over 90% precision and 75% recall for grocery and household categories, and an overall catalog assignment accuracy of 80.2%.

0 citationsRead paper

Demand Transfer Estimation at Scale via Restricted Logit Modeling

Aug 12, 2026

This study addresses the computational bottleneck in estimating demand substitution effects within large-scale product assortments by proposing an efficient modeling approach based on a constrained logit framework. Building upon independent demand forecasts for individual products, the method incorporates inter-product substitution relationships to adjust predictions, thereby circumventing the need to model each assortment combination separately. The approach enables, for the first time, scalable and accurate estimation of demand transfer coefficients at the million-product scale, overcoming the efficiency limitations of traditional discrete choice models in large-scale settings. Empirical validation using multi-category, multi-store historical transaction data demonstrates significant improvements in demand forecasting accuracy under plausible substitution assumptions.

0 citationsRead paper

Adversarial Prompting Framework for AI Safety Assessment

Jul 15, 2026

This study addresses emerging security threats such as adversarial prompt attacks (APAs) that challenge the robustness of generative AI systems in enterprise deployments, necessitating systematic evaluation frameworks. To this end, the paper proposes the Adversarial Prompt Framework (APF), the first multi-level, structured methodology for generating and evaluating adversarial prompts across diverse attack vectors—from direct harmful requests to sophisticated encoded attacks. APF integrates automated prompt generation, encoding transformations, and quantitative safety metrics to enable end-to-end security assessment. Experimental results demonstrate significant variability in model vulnerability across attack types, with encoded prompts achieving the highest success rates in bypassing built-in safety mechanisms, thereby validating APF’s effectiveness and practical utility for comprehensive robustness evaluation.

0 citationsRead paper

Bridging the Catalog-to-Real Gap: Scalable Product Recognition via Multi-Stage Contrastive Learning

Jul 10, 2026

This work addresses the challenge of scalable product matching in large-scale retail settings, where significant domain discrepancies exist between high-resolution catalog images and real in-store photographs. To bridge this domain gap, the authors propose Cat2Real, a multi-stage contrastive learning framework that reformulates product identification as a cross-domain embedding retrieval task. The approach integrates hierarchical hard negative mining guided by both item-level and image-level similarity, cross-domain contrastive losses, and visual embedding learning, enabling zero-shot generalization to novel products without requiring real-world training data for new items. Experimental results demonstrate that the model substantially improves matching accuracy and system scalability on unseen products and categories, effectively narrowing the domain gap between catalog and real-world imagery.

0 citationsRead paper
Recent publications

Latest Papers

Attribute-Conditioned Multimodal Slot Factorization for Controllable Fashion Retrieval

Aug 12, 2026

This work addresses the limitation of conventional fashion retrieval methods, which conflate multi-attribute information into a single embedding, thereby hindering fine-grained control over specific attributes such as color or category. To overcome this, the authors propose the MM-slotgate model, which introduces— for the first time—a named and intervenable multimodal attribute-slot architecture. This framework decomposes Fashion-CLIP’s joint image-text embeddings into four semantically distinct attribute slots, each equipped with an independent image-text gating mechanism that automatically learns modality weights without explicit supervision. The approach achieves both semantic interpretability and controllable retrieval performance, attaining a ConstraintSatisfied@10 score of 0.7566 on the H&M dataset, improving color retrieval success by 0.568, and yielding a 15.3× enhancement in color intervention efficacy through quantized slot representations.

0 citationsRead paper

Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy

Aug 12, 2026

This study addresses the challenge global retailers face in manually maintaining price consistency and constructing coherent price hierarchies across million-scale product catalogs. To this end, the work proposes a scalable, context-aware multi-agent framework that introduces, for the first time, a multi-agent architecture to the construction of retail price classification systems—commonly known as “Lines and Ladders.” The framework orchestrates multiple specialized large language model agents to collaboratively perform product attribute recognition, multimodal value extraction, and hierarchical grouping, thereby mitigating cognitive overload inherent in single-agent approaches. Experimental results on real-world enterprise data demonstrate that a three-agent system achieves an F1 score of 0.83 on Lines tasks, with over 90% precision and 75% recall for grocery and household categories, and an overall catalog assignment accuracy of 80.2%.

0 citationsRead paper

Demand Transfer Estimation at Scale via Restricted Logit Modeling

Aug 12, 2026

This study addresses the computational bottleneck in estimating demand substitution effects within large-scale product assortments by proposing an efficient modeling approach based on a constrained logit framework. Building upon independent demand forecasts for individual products, the method incorporates inter-product substitution relationships to adjust predictions, thereby circumventing the need to model each assortment combination separately. The approach enables, for the first time, scalable and accurate estimation of demand transfer coefficients at the million-product scale, overcoming the efficiency limitations of traditional discrete choice models in large-scale settings. Empirical validation using multi-category, multi-store historical transaction data demonstrates significant improvements in demand forecasting accuracy under plausible substitution assumptions.

0 citationsRead paper

Adversarial Prompting Framework for AI Safety Assessment

Jul 15, 2026

This study addresses emerging security threats such as adversarial prompt attacks (APAs) that challenge the robustness of generative AI systems in enterprise deployments, necessitating systematic evaluation frameworks. To this end, the paper proposes the Adversarial Prompt Framework (APF), the first multi-level, structured methodology for generating and evaluating adversarial prompts across diverse attack vectors—from direct harmful requests to sophisticated encoded attacks. APF integrates automated prompt generation, encoding transformations, and quantitative safety metrics to enable end-to-end security assessment. Experimental results demonstrate significant variability in model vulnerability across attack types, with encoded prompts achieving the highest success rates in bypassing built-in safety mechanisms, thereby validating APF’s effectiveness and practical utility for comprehensive robustness evaluation.

0 citationsRead paper

Bridging the Catalog-to-Real Gap: Scalable Product Recognition via Multi-Stage Contrastive Learning

Jul 10, 2026

This work addresses the challenge of scalable product matching in large-scale retail settings, where significant domain discrepancies exist between high-resolution catalog images and real in-store photographs. To bridge this domain gap, the authors propose Cat2Real, a multi-stage contrastive learning framework that reformulates product identification as a cross-domain embedding retrieval task. The approach integrates hierarchical hard negative mining guided by both item-level and image-level similarity, cross-domain contrastive losses, and visual embedding learning, enabling zero-shot generalization to novel products without requiring real-world training data for new items. Experimental results demonstrate that the model substantially improves matching accuracy and system scalability on unseen products and categories, effectively narrowing the domain gap between catalog and real-world imagery.

0 citationsRead paper