Institution profile

Yooz

Industry research
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

ConRTF: Edge-Constrained Boundary Distribution Refinement for Realtime TransFormer Table Structure Recognition

Jul 01, 2026

This work addresses the challenge of inaccurate row/column boundary localization in table structure recognition, which often leads to erroneous cell assignments. Existing methods typically overlook the inherent geometric asymmetry between rows and columns. To remedy this, the authors propose an Edge-constrained Fine-grained Localization (EFL) loss that incorporates geometric priors during training—emphasizing horizontal boundaries for rows and vertical boundaries for columns—and integrates a Distribution-aware Boundary Refinement module (D-FINE) to enhance localization accuracy without increasing inference overhead. This approach is the first to explicitly model row-column asymmetry as part of the training objective, enabling efficient, structure-aware boundary optimization. Evaluated within a Transformer-based real-time detection framework, the method outperforms RT-DETRv2 and YOLOv10–11 on PubTables-1M and two private datasets, achieving up to a 1.6-point gain in GriTS scores while maintaining robust performance with only 2k–3k annotated samples.

0 citationsRead paper

Active Learning for Cascaded Object Detection: Balancing Coverage and Uncertainty in Table Extraction Pipelines

Jul 01, 2026

This work addresses the high annotation cost in the structure recognition phase of cascaded table extraction and the neglect of inter-stage dependencies between detection and recognition in existing active learning approaches. To this end, it introduces Uncertainty Herding into this pipeline for the first time and proposes two pipeline-aware active learning methods, RankFusion and CAPA. These methods jointly model representativeness and uncertainty by integrating dual-space coverage from both detection and structural representation, incorporating a stage-aware gating mechanism, and calibrating task-specific uncertainty, thereby explicitly leveraging cross-stage dependencies. Experimental results demonstrate that CAPA significantly outperforms baseline methods on three out of four benchmark datasets, effectively reducing annotation costs while enhancing model performance.

0 citationsRead paper

QUEST: Quality-aware Semi-supervised Table Extraction for Business Documents

Jun 17, 2025

To address the challenges of sparse human annotations, multi-stage error accumulation, and unreliable pseudo-labels in commercial document table extraction (TE), this paper proposes a quality-aware semi-supervised framework. Methodologically, it introduces: (1) a novel F1-predictability assessment model grounded in structural and contextual modeling for fine-grained quality quantification; (2) a row- and column-level diversity sampling mechanism integrating Determinantal Point Processes (DPP), Vendi Score, and IntDiv to mitigate confirmation bias; and (3) an interpretable, robust quality-guided pseudo-label filtering strategy. Evaluated on a proprietary business dataset, the framework achieves an F1 score of 74% (+10 percentage points) and reduces empty-table false detection by 45%. On the DocILE benchmark, it attains an F1 score of 50% (+8 percentage points) and lowers empty-table prediction errors by 19%.

0 citationsRead paper

RAPTOR: Refined Approach for Product Table Object Recognition

Feb 19, 2025

DETR-based models (e.g., TATR) suffer from industrial-grade errors—such as region mislocalization and column overlap—in product table detection (TD) and table structure recognition (TSR), primarily due to diverse, unstructured document layouts. To address this, we propose RAPTOR, a modular post-processing framework tailored for product tables (e.g., invoices, quotations). Its key contributions are: (1) the first configurable, domain-specific post-processing architecture for product tables; (2) integration of geometric correction, row/column alignment, and cell merging modules to refine coarse model outputs; and (3) adoption of a genetic algorithm for end-to-end automatic hyperparameter optimization on proprietary datasets. Evaluated on multiple private product table benchmarks, RAPTOR achieves significant F1-score improvements over baseline methods, while maintaining robustness on public benchmarks—including DOCILE and ICDAR 2013/2019. Ablation studies confirm the effectiveness and complementarity of each module.

0 citationsRead paper
Recent publications

Latest Papers

ConRTF: Edge-Constrained Boundary Distribution Refinement for Realtime TransFormer Table Structure Recognition

Jul 01, 2026

This work addresses the challenge of inaccurate row/column boundary localization in table structure recognition, which often leads to erroneous cell assignments. Existing methods typically overlook the inherent geometric asymmetry between rows and columns. To remedy this, the authors propose an Edge-constrained Fine-grained Localization (EFL) loss that incorporates geometric priors during training—emphasizing horizontal boundaries for rows and vertical boundaries for columns—and integrates a Distribution-aware Boundary Refinement module (D-FINE) to enhance localization accuracy without increasing inference overhead. This approach is the first to explicitly model row-column asymmetry as part of the training objective, enabling efficient, structure-aware boundary optimization. Evaluated within a Transformer-based real-time detection framework, the method outperforms RT-DETRv2 and YOLOv10–11 on PubTables-1M and two private datasets, achieving up to a 1.6-point gain in GriTS scores while maintaining robust performance with only 2k–3k annotated samples.

0 citationsRead paper

Active Learning for Cascaded Object Detection: Balancing Coverage and Uncertainty in Table Extraction Pipelines

Jul 01, 2026

This work addresses the high annotation cost in the structure recognition phase of cascaded table extraction and the neglect of inter-stage dependencies between detection and recognition in existing active learning approaches. To this end, it introduces Uncertainty Herding into this pipeline for the first time and proposes two pipeline-aware active learning methods, RankFusion and CAPA. These methods jointly model representativeness and uncertainty by integrating dual-space coverage from both detection and structural representation, incorporating a stage-aware gating mechanism, and calibrating task-specific uncertainty, thereby explicitly leveraging cross-stage dependencies. Experimental results demonstrate that CAPA significantly outperforms baseline methods on three out of four benchmark datasets, effectively reducing annotation costs while enhancing model performance.

0 citationsRead paper

QUEST: Quality-aware Semi-supervised Table Extraction for Business Documents

Jun 17, 2025

To address the challenges of sparse human annotations, multi-stage error accumulation, and unreliable pseudo-labels in commercial document table extraction (TE), this paper proposes a quality-aware semi-supervised framework. Methodologically, it introduces: (1) a novel F1-predictability assessment model grounded in structural and contextual modeling for fine-grained quality quantification; (2) a row- and column-level diversity sampling mechanism integrating Determinantal Point Processes (DPP), Vendi Score, and IntDiv to mitigate confirmation bias; and (3) an interpretable, robust quality-guided pseudo-label filtering strategy. Evaluated on a proprietary business dataset, the framework achieves an F1 score of 74% (+10 percentage points) and reduces empty-table false detection by 45%. On the DocILE benchmark, it attains an F1 score of 50% (+8 percentage points) and lowers empty-table prediction errors by 19%.

0 citationsRead paper

RAPTOR: Refined Approach for Product Table Object Recognition

Feb 19, 2025

DETR-based models (e.g., TATR) suffer from industrial-grade errors—such as region mislocalization and column overlap—in product table detection (TD) and table structure recognition (TSR), primarily due to diverse, unstructured document layouts. To address this, we propose RAPTOR, a modular post-processing framework tailored for product tables (e.g., invoices, quotations). Its key contributions are: (1) the first configurable, domain-specific post-processing architecture for product tables; (2) integration of geometric correction, row/column alignment, and cell merging modules to refine coarse model outputs; and (3) adoption of a genetic algorithm for end-to-end automatic hyperparameter optimization on proprietary datasets. Evaluated on multiple private product table benchmarks, RAPTOR achieves significant F1-score improvements over baseline methods, while maintaining robustness on public benchmarks—including DOCILE and ICDAR 2013/2019. Ablation studies confirm the effectiveness and complementarity of each module.

0 citationsRead paper