Institution profile

MARSAIL (Motor AI Recognition Solution Artificial Intelligence Laboratory)

Academic institutionasia · cn
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

HOMEY: Heuristic Object Masking with Enhanced YOLO for Property Insurance Risk Detection

Mar 19, 2026

This study addresses the need for automated property risk identification in insurance underwriting by proposing a YOLO-based object detection method capable of efficiently recognizing 17 categories of structural damage, maintenance deficiencies, and safety hazards. The approach introduces a heuristic object masking mechanism to enhance detection of weak-signal targets and incorporates a risk-aware weighted loss function to mitigate challenges arising from class imbalance and varying risk severities. Experimental results on real-world property images demonstrate that the proposed method significantly outperforms baseline models in detection accuracy and reliability while preserving YOLO’s computational efficiency, thereby offering a cost-effective and interpretable solution for property risk assessment in insurance applications.

0 citationsRead paper

SLICK: Selective Localization and Instance Calibration for Knowledge-Enhanced Car Damage Segmentation in Automotive Insurance

Jun 12, 2025

Addressing challenges in real-world road scenarios—including occlusion, deformation, paint loss, part overlap, and fine-scale scratches/dents that are difficult to segment—this paper proposes a structure- and knowledge-cooperative automotive damage segmentation method. The approach introduces a selective part segmentation module, a localization-aware attention mechanism, an instance-sensitive refinement head, a cross-channel calibration module, and a multi-source knowledge fusion module. It integrates a high-resolution semantic backbone, dynamic localization attention, pan-instance-guided shape prior modeling, and a joint training strategy leveraging synthetic collision data and geometric priors. Evaluated on a large-scale automotive dataset, the method achieves significant improvements in boundary alignment accuracy and overall segmentation performance: detection rates for fine-scale damages (scratches/dents) increase by 23.6%, and noise robustness reaches state-of-the-art (SOTA) levels.

0 citationsRead paper

ALBERT: Advanced Localization and Bidirectional Encoder Representations from Transformers for Automotive Damage Evaluation

Jun 12, 2025

This study addresses the joint modeling challenge of authentic/forged damage discrimination and fine-grained vehicle body part segmentation for intelligent vehicle damage assessment. We propose the first unified end-to-end framework integrating a bidirectional Transformer encoder with an adaptive localization attention mechanism, supporting 26 classes of authentic damage recognition, 7 classes of forged damage classification, and 61-class instance-level part segmentation. Leveraging multi-task joint training and a large-scale, meticulously annotated dataset—including forged damage samples—our method achieves 82.3% mAP for part segmentation and 94.1% accuracy for damage classification on a custom benchmark, significantly outperforming Mask R-CNN and Swin-InstanceSegmenter. The core innovation lies in the first integration of bidirectional semantic representation, precise localization capability, and hierarchical joint understanding of damage and parts within a single instance segmentation architecture.

0 citationsRead paper

DOTA: Deformable Optimized Transformer Architecture for End-to-End Text Recognition with Retrieval-Augmented Generation

May 07, 2025

Scene text recognition remains a fundamental challenge in computer vision and multimodal understanding. This paper proposes an end-to-end OCR framework that enhances geometric robustness by embedding deformable convolutions into the third and fourth stages of a ResNet-ViT hybrid backbone. To jointly model structural constraints and semantic context, it introduces, for the first time, a synergistic integration of retrieval-augmented generation (RAG) and conditional random fields (CRFs). Additionally, an adaptive dropout mechanism is designed to improve generalization. Evaluated on six standard benchmarks—including IC13 and IC15—the framework achieves a mean accuracy of 77.77%, with 97.32% on IC13—setting new state-of-the-art results across multiple datasets. The proposed method significantly advances recognition accuracy and robustness in complex, real-world scenes.

0 citationsRead paper
Recent publications

Latest Papers

HOMEY: Heuristic Object Masking with Enhanced YOLO for Property Insurance Risk Detection

Mar 19, 2026

This study addresses the need for automated property risk identification in insurance underwriting by proposing a YOLO-based object detection method capable of efficiently recognizing 17 categories of structural damage, maintenance deficiencies, and safety hazards. The approach introduces a heuristic object masking mechanism to enhance detection of weak-signal targets and incorporates a risk-aware weighted loss function to mitigate challenges arising from class imbalance and varying risk severities. Experimental results on real-world property images demonstrate that the proposed method significantly outperforms baseline models in detection accuracy and reliability while preserving YOLO’s computational efficiency, thereby offering a cost-effective and interpretable solution for property risk assessment in insurance applications.

0 citationsRead paper

SLICK: Selective Localization and Instance Calibration for Knowledge-Enhanced Car Damage Segmentation in Automotive Insurance

Jun 12, 2025

Addressing challenges in real-world road scenarios—including occlusion, deformation, paint loss, part overlap, and fine-scale scratches/dents that are difficult to segment—this paper proposes a structure- and knowledge-cooperative automotive damage segmentation method. The approach introduces a selective part segmentation module, a localization-aware attention mechanism, an instance-sensitive refinement head, a cross-channel calibration module, and a multi-source knowledge fusion module. It integrates a high-resolution semantic backbone, dynamic localization attention, pan-instance-guided shape prior modeling, and a joint training strategy leveraging synthetic collision data and geometric priors. Evaluated on a large-scale automotive dataset, the method achieves significant improvements in boundary alignment accuracy and overall segmentation performance: detection rates for fine-scale damages (scratches/dents) increase by 23.6%, and noise robustness reaches state-of-the-art (SOTA) levels.

0 citationsRead paper

ALBERT: Advanced Localization and Bidirectional Encoder Representations from Transformers for Automotive Damage Evaluation

Jun 12, 2025

This study addresses the joint modeling challenge of authentic/forged damage discrimination and fine-grained vehicle body part segmentation for intelligent vehicle damage assessment. We propose the first unified end-to-end framework integrating a bidirectional Transformer encoder with an adaptive localization attention mechanism, supporting 26 classes of authentic damage recognition, 7 classes of forged damage classification, and 61-class instance-level part segmentation. Leveraging multi-task joint training and a large-scale, meticulously annotated dataset—including forged damage samples—our method achieves 82.3% mAP for part segmentation and 94.1% accuracy for damage classification on a custom benchmark, significantly outperforming Mask R-CNN and Swin-InstanceSegmenter. The core innovation lies in the first integration of bidirectional semantic representation, precise localization capability, and hierarchical joint understanding of damage and parts within a single instance segmentation architecture.

0 citationsRead paper

DOTA: Deformable Optimized Transformer Architecture for End-to-End Text Recognition with Retrieval-Augmented Generation

May 07, 2025

Scene text recognition remains a fundamental challenge in computer vision and multimodal understanding. This paper proposes an end-to-end OCR framework that enhances geometric robustness by embedding deformable convolutions into the third and fourth stages of a ResNet-ViT hybrid backbone. To jointly model structural constraints and semantic context, it introduces, for the first time, a synergistic integration of retrieval-augmented generation (RAG) and conditional random fields (CRFs). Additionally, an adaptive dropout mechanism is designed to improve generalization. Evaluated on six standard benchmarks—including IC13 and IC15—the framework achieves a mean accuracy of 77.77%, with 97.32% on IC13—setting new state-of-the-art results across multiple datasets. The proposed method significantly advances recognition accuracy and robustness in complex, real-world scenes.

0 citationsRead paper