Institution profile

Intuit

Industry researchnorthamerica · us
Official website
Research library66linked papers
Opportunities112open roles
Selected work

Representative Papers

COMBOOD: A Semiparametric Approach for Detecting Out-of-distribution Data for Image Classification

Feb 04, 2026SDM

This work addresses the challenge of effectively detecting near-distribution out-of-distribution (near-OOD) samples in image classification inference, a task where existing methods often fall short. To this end, the authors propose COMBOOD, an unsupervised semi-parametric framework that uniquely integrates non-parametric nearest-neighbor distances with parametric Mahalanobis distances in the feature embedding space to produce a unified confidence score. This fusion enables robust performance across both near-OOD and far-OOD scenarios. COMBOOD is compatible with diverse feature extractors and exhibits computational complexity that scales linearly with the embedding dimensionality. Extensive evaluations on OpenOOD v1/v1.5 benchmarks and document datasets demonstrate that COMBOOD consistently outperforms current state-of-the-art methods, with most improvements achieving statistical significance.

8 citations1 influentialRead paper

REMem: Reasoning with Episodic Memory in Language Agent

Feb 13, 2026

This work addresses the limited capacity of current language agents to recall and reason over interactive histories with contextual richness, as prevailing memory systems predominantly emphasize semantic memory while neglecting the temporal and spatial context of events and lack explicit modeling. To bridge this gap, the authors propose REMem, a novel framework that formally articulates the challenge of episodic memory in language agents for the first time. REMem introduces a two-stage architecture: offline, it constructs a hybrid memory graph integrating time-aware summaries and factual details; online, it employs a tool-augmented, intelligent retriever to perform iterative reasoning over this graph. Evaluated across four benchmarks, REMem significantly outperforms Mem0 and HippoRAG 2, achieving accuracy gains of 3.4% and 13.4% on episodic recall and reasoning tasks, respectively, while also demonstrating enhanced robustness in rejecting unanswerable queries.

1 citationsRead paper
Recent publications

Latest Papers

GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation

Aug 12, 2026

This study addresses the challenge of automatically generating actionable financial recommendations from enterprise operational data that integrate numerical reasoning, domain expertise, safety guarantees, and tangible business value, without relying on costly human annotations. The task is formulated as a reinforcement learning problem, and we propose Group Relative Policy Optimization (GRPO), a fine-tuning framework tailored for financial scenarios, incorporating an LLM-as-a-judge reward mechanism based on multidimensional binary scoring and a safety filter. Innovatively, we introduce CATE (Conditional Average Treatment Effect) auditing from causal inference to uncover real-world business impact that LLM-based evaluations fail to capture. Experiments demonstrate that our approach achieves a gross margin improvement of 0.0228—approximately twice that of the strongest commercial baseline—while simultaneously exhibiting the lowest downside risk and negative tail loss, all without any human-labeled data.

0 citationsRead paper