Institution profile

Traversaal.ai

Industry researchnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop

Jan 28, 2026

This work addresses the absence of standardized reasoning benchmarks for low-resource languages like Urdu and the inability of existing machine translation approaches to preserve contextual and structural integrity in reasoning tasks. The authors propose a context-integrated translation framework that combines outputs from multiple translation systems with human validation to construct UrduBench—the first high-quality Urdu reasoning benchmark spanning multiple difficulty levels and task types, including MGSM and MATH-500. Using this benchmark, they systematically evaluate various large language models under diverse prompting strategies, revealing significant performance degradation in multi-step and symbolic reasoning. Their findings underscore the critical role of linguistic consistency in enabling robust cross-lingual reasoning and establish a scalable paradigm for evaluating reasoning capabilities in low-resource languages.

0 citationsRead paper

Query Attribute Modeling: Improving search relevance with Semantic Search and Meta Data Filtering

Aug 06, 2025

Traditional open-ended textual queries struggle to precisely extract structured metadata and semantic elements, leading to suboptimal relevance in e-commerce search. To address this, we propose Query Attribute Modeling (QAM), a framework that automatically decomposes free-text queries into executable metadata filtering conditions and dense semantic representations. QAM integrates semantic search, BM25 retrieval, cross-encoder re-ranking, and reciprocal rank fusion (RRF) into a hybrid retrieval pipeline. By jointly optimizing structured constraints and semantic understanding, it overcomes the limitations of both pure keyword matching and end-to-end semantic models. Experiments on the Amazon Toys Reviews dataset demonstrate that QAM achieves 52.99% mAP@5—significantly outperforming BM25, standalone semantic search, and existing hybrid baselines. These results validate QAM’s effectiveness and practicality for real-world e-commerce search applications.

0 citationsRead paper
Recent publications

Latest Papers

UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop

Jan 28, 2026

This work addresses the absence of standardized reasoning benchmarks for low-resource languages like Urdu and the inability of existing machine translation approaches to preserve contextual and structural integrity in reasoning tasks. The authors propose a context-integrated translation framework that combines outputs from multiple translation systems with human validation to construct UrduBench—the first high-quality Urdu reasoning benchmark spanning multiple difficulty levels and task types, including MGSM and MATH-500. Using this benchmark, they systematically evaluate various large language models under diverse prompting strategies, revealing significant performance degradation in multi-step and symbolic reasoning. Their findings underscore the critical role of linguistic consistency in enabling robust cross-lingual reasoning and establish a scalable paradigm for evaluating reasoning capabilities in low-resource languages.

0 citationsRead paper

Query Attribute Modeling: Improving search relevance with Semantic Search and Meta Data Filtering

Aug 06, 2025

Traditional open-ended textual queries struggle to precisely extract structured metadata and semantic elements, leading to suboptimal relevance in e-commerce search. To address this, we propose Query Attribute Modeling (QAM), a framework that automatically decomposes free-text queries into executable metadata filtering conditions and dense semantic representations. QAM integrates semantic search, BM25 retrieval, cross-encoder re-ranking, and reciprocal rank fusion (RRF) into a hybrid retrieval pipeline. By jointly optimizing structured constraints and semantic understanding, it overcomes the limitations of both pure keyword matching and end-to-end semantic models. Experiments on the Amazon Toys Reviews dataset demonstrate that QAM achieves 52.99% mAP@5—significantly outperforming BM25, standalone semantic search, and existing hybrid baselines. These results validate QAM’s effectiveness and practicality for real-world e-commerce search applications.

0 citationsRead paper