Institution profile

Oracle Corporation

Industry researchnorthamerica · us
Official website
Research library98linked papers
Opportunities180open roles
Selected work

Representative Papers

FS-DAG: Few Shot Domain Adapting Graph Networks for Visually Rich Document Understanding

May 22, 2025

Addressing key challenges in few-shot visual-rich document understanding (VRDU)—including poor cross-domain transferability, low robustness to OCR noise, and constrained deployment resources—this paper proposes a modular few-shot domain-adaptive graph network architecture. The method supports dynamic switching between language- and vision-based backbones, and integrates graph neural networks, multimodal feature alignment, domain-adaptive knowledge distillation, and OCR-robust encoding. With fewer than 90M parameters, it achieves both lightweight design and strong generalization. On information extraction tasks, the model significantly accelerates convergence and improves accuracy, consistently outperforming existing state-of-the-art approaches. It provides a scalable, highly robust solution to real-world issues such as OCR errors, spelling variants, and domain shift—enabling effective adaptation under limited supervision and noisy inputs.

2 citationsRead paper

SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use

May 22, 2025North American Chapter of the Association for Computational Linguistics

Existing safety and ethics evaluations for large language models (LLMs) in enterprise multilingual, cross-cultural communication—e.g., email drafting and sales copy—lack systematic assessment of resistance to abusive instructions, cultural contextual adaptation, and ethical alignment. Method: We introduce the first enterprise-oriented, multilingual adversarial instruction safety benchmark, pioneering an explicit induction-based safety evaluation paradigm. It formalizes profanity-embedded adversarial prompts, integrates real-world communicative contexts and tonal variations, and jointly measures model refusal capability and cultural-linguistic comprehension. Our methodology includes constructing a multilingual profanity lexicon, human-in-the-loop + rule-augmented evaluation protocols, an open-source dataset, and an automated evaluation framework. Results: Evaluated on 20+ mainstream LLMs, we observe significantly degraded compliance under informal contexts. The benchmark delivers reproducible, quantitative safety metrics—enabling rigorous risk assessment and governance for enterprise AI deployment.

1 citationsRead paper

Hierarchical Preference Optimization: Learning to achieve goals via feasible subgoals prediction

Nov 01, 2024arXiv.org

Hierarchical reinforcement learning (HRL) suffers from two key challenges: non-stationarity in high-level policies due to evolving low-level policies, and infeasible sub-goals generated by high-level policies that low-level policies cannot execute. To address these, we propose Hierarchical Preference Optimization (HPO), the first framework to integrate token-level direct preference optimization (DPO) into HRL—without requiring a pretrained reference policy. HPO jointly optimizes high-level goal generation and low-level action selection via a bilevel optimization formulation. We introduce a primitive-regularized DPO loss that mathematically enforces sub-goal feasibility and prevents degenerate solutions. Additionally, maximum entropy regularization is incorporated to enhance exploration robustness. Evaluated on robotic navigation and manipulation tasks, HPO achieves an average 35% performance gain over strong baselines, significantly mitigating both non-stationarity and sub-goal infeasibility. Ablation studies and quantitative analysis comprehensively validate its effectiveness.

1 citationsRead paper

Knowing When Not to Answer: Abstention and Refusal Reasoning in Vision--Language Models

Sep 02, 2026

Many medical conditions require diagnosis through detailed, multi-context clinical assessment rather than from visual appearance alone. Despite this, vision-language models (VLMs) are increasingly queried to interpret images in ways that touch on medical or diagnostic judgments, raising safety concerns when such inferences are unsupported. ASD diagnosis requires behavioral and developmental evidence, not static facial photographs. We audit whether VLMs abstain from this unanswerable paired-image query, and whether expressions sway non-abstaining choices. We introduce PARITY (Paired Assessment with Reused Identity), a synthetic, demographically balanced set of identity-controlled neutral/expression portrait pairs with neutral-neutral controls. All identities are synthetic and have no ASD status; because the query is unanswerable from images, any non-abstaining selection is treated as a harmful attribution. Across contemporary VLMs, we find a clear split between refusal-first models and speculative models; in the latter, certain expressions disproportionately trigger harmful selections. Clinical guardrails and single-image framing substantially increase abstention, suggesting actionable mitigations in both prompting and interface design

0 citationsRead paper
Recent publications

Latest Papers

Knowing When Not to Answer: Abstention and Refusal Reasoning in Vision--Language Models

Sep 02, 2026

Many medical conditions require diagnosis through detailed, multi-context clinical assessment rather than from visual appearance alone. Despite this, vision-language models (VLMs) are increasingly queried to interpret images in ways that touch on medical or diagnostic judgments, raising safety concerns when such inferences are unsupported. ASD diagnosis requires behavioral and developmental evidence, not static facial photographs. We audit whether VLMs abstain from this unanswerable paired-image query, and whether expressions sway non-abstaining choices. We introduce PARITY (Paired Assessment with Reused Identity), a synthetic, demographically balanced set of identity-controlled neutral/expression portrait pairs with neutral-neutral controls. All identities are synthetic and have no ASD status; because the query is unanswerable from images, any non-abstaining selection is treated as a harmful attribution. Across contemporary VLMs, we find a clear split between refusal-first models and speculative models; in the latter, certain expressions disproportionately trigger harmful selections. Clinical guardrails and single-image framing substantially increase abstention, suggesting actionable mitigations in both prompting and interface design

0 citationsRead paper