Institution profile

Tomocube Inc.

Industry researchasia · kr
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

PILOT-Bench: A Benchmark for Legal Reasoning in the Patent Domain with IRAC-Aligned Classification Tasks

Jan 08, 2026Proceedings of the Natural Legal Language Processing Workshop 2025

This study addresses the absence of a systematic evaluation benchmark for large language models (LLMs) in the domain of patent law reasoning. The authors construct the first benchmark centered on decisions from the U.S. Patent Trial and Appeal Board (PTAB), aligning PTAB rulings with USPTO patent data to formulate three structured classification tasks grounded in the IRAC legal analysis framework: issue type, cited authority, and sub-decision. The benchmark enables multidimensional evaluation across input variations, model families, and error analyses, offering a comprehensive assessment of both open- and closed-source LLMs. Experimental results reveal a substantial performance gap: the best closed-source model achieves a Micro-F1 score of 0.75 on the issue-type task, whereas the strongest open-source model, Qwen-8B, attains only 0.56, highlighting significant limitations in current models’ capacity for patent-related legal reasoning.

1 citationsRead paper

Soft Filtering: Guiding Zero-shot Composed Image Retrieval with Prescriptive and Proscriptive Constraints

Dec 23, 2025

Zero-shot compositional image retrieval (ZS-CIR) faces three key challenges: ambiguous user intent, entanglement of positive and negative semantics, and the mismatch between the single-target assumption and real-world query ambiguity. To address these, we propose SoFT—a training-free soft filtering module that introduces dual-track text constraint modeling for the first time, explicitly distinguishing prescriptive (“must include”) from proscriptive (“must avoid”) semantics. Leveraging multimodal large language models (MLLMs), SoFT dynamically parses both constraint types from reference images and modification texts to enable zero-shot re-ranking of retrieval results. Additionally, we design a pipeline for generating multi-objective CIR benchmarks supporting fine-grained, ambiguity-robust evaluation. On CIRR, CIRCO, and FashionIQ, SoFT achieves improvements of +12.94 in R@5, +6.13 in mAP@50, and +4.59 in R@50, significantly enhancing both robustness and accuracy of ZS-CIR.

0 citationsRead paper

STELLAR: Scene Text Editor for Low-Resource Languages and Real-World Data

Nov 13, 2025

To address three key challenges in scene text editing (STE)—limited low-resource language support, domain shift between synthetic and real data, and the lack of quantitative evaluation for text style preservation—this paper proposes: (1) a language-adaptive glyph encoder with a multi-stage training strategy; (2) STIPLAR, the first real-world STE benchmark for low-resource languages; (3) a diffusion-based framework combining synthetic-data pretraining with real-image fine-tuning; and (4) Text Appearance Similarity (TAS), a novel metric unifying font, color, and background consistency quantification. Experiments demonstrate a 2.2% average cross-lingual TAS improvement, significantly enhancing visual fidelity and OCR recognition accuracy. The proposed approach delivers a scalable, quantitatively evaluable solution for multilingual STE.

0 citationsRead paper
Recent publications

Latest Papers

PILOT-Bench: A Benchmark for Legal Reasoning in the Patent Domain with IRAC-Aligned Classification Tasks

Jan 08, 2026Proceedings of the Natural Legal Language Processing Workshop 2025

This study addresses the absence of a systematic evaluation benchmark for large language models (LLMs) in the domain of patent law reasoning. The authors construct the first benchmark centered on decisions from the U.S. Patent Trial and Appeal Board (PTAB), aligning PTAB rulings with USPTO patent data to formulate three structured classification tasks grounded in the IRAC legal analysis framework: issue type, cited authority, and sub-decision. The benchmark enables multidimensional evaluation across input variations, model families, and error analyses, offering a comprehensive assessment of both open- and closed-source LLMs. Experimental results reveal a substantial performance gap: the best closed-source model achieves a Micro-F1 score of 0.75 on the issue-type task, whereas the strongest open-source model, Qwen-8B, attains only 0.56, highlighting significant limitations in current models’ capacity for patent-related legal reasoning.

1 citationsRead paper

Soft Filtering: Guiding Zero-shot Composed Image Retrieval with Prescriptive and Proscriptive Constraints

Dec 23, 2025

Zero-shot compositional image retrieval (ZS-CIR) faces three key challenges: ambiguous user intent, entanglement of positive and negative semantics, and the mismatch between the single-target assumption and real-world query ambiguity. To address these, we propose SoFT—a training-free soft filtering module that introduces dual-track text constraint modeling for the first time, explicitly distinguishing prescriptive (“must include”) from proscriptive (“must avoid”) semantics. Leveraging multimodal large language models (MLLMs), SoFT dynamically parses both constraint types from reference images and modification texts to enable zero-shot re-ranking of retrieval results. Additionally, we design a pipeline for generating multi-objective CIR benchmarks supporting fine-grained, ambiguity-robust evaluation. On CIRR, CIRCO, and FashionIQ, SoFT achieves improvements of +12.94 in R@5, +6.13 in mAP@50, and +4.59 in R@50, significantly enhancing both robustness and accuracy of ZS-CIR.

0 citationsRead paper

STELLAR: Scene Text Editor for Low-Resource Languages and Real-World Data

Nov 13, 2025

To address three key challenges in scene text editing (STE)—limited low-resource language support, domain shift between synthetic and real data, and the lack of quantitative evaluation for text style preservation—this paper proposes: (1) a language-adaptive glyph encoder with a multi-stage training strategy; (2) STIPLAR, the first real-world STE benchmark for low-resource languages; (3) a diffusion-based framework combining synthetic-data pretraining with real-image fine-tuning; and (4) Text Appearance Similarity (TAS), a novel metric unifying font, color, and background consistency quantification. Experiments demonstrate a 2.2% average cross-lingual TAS improvement, significantly enhancing visual fidelity and OCR recognition accuracy. The proposed approach delivers a scalable, quantitatively evaluable solution for multilingual STE.

0 citationsRead paper