🤖 AI Summary
Jewelry exhibits high intra-class visual diversity and structural complexity, making accurate identification and descriptive translation challenging for non-experts—particularly in cross-lingual terminology alignment. To address this, we propose a translation-oriented, fine-grained visual description framework for jewelry: a three-level natural language generation architecture (category → structure → material/technique) that emulates expert cognitive reasoning. Our method integrates CNN-based visual feature extraction with an encoder-decoder architecture, augmented by image captioning and object detection to enable end-to-end jewelry recognition and multi-granularity textual description generation. Evaluated on a curated jewelry image dataset, our model achieves 90.3% description accuracy, substantially outperforming baseline approaches. This work represents the first systematic application of fine-grained vision-language generation to support jewelry translation, establishing a scalable paradigm for domain-specific visual–linguistic co-understanding.
📝 Abstract
Identifying jewelry pieces presents a significant challenge due to the wide range of styles and designs. Currently, precise descriptions are typically limited to industry experts. However, translators and interpreters often require a comprehensive understanding of these items. In this study, we introduce an innovative approach to automatically identify and describe jewelry using neural networks. This method enables translators and interpreters to quickly access accurate information, aiding in resolving queries and gaining essential knowledge about jewelry. Our model operates at three distinct levels of description, employing computer vision techniques and image captioning to emulate expert analysis of accessories. The key innovation involves generating natural language descriptions of jewelry across three hierarchical levels, capturing nuanced details of each piece. Different image captioning architectures are utilized to detect jewels in images and generate descriptions with varying levels of detail. To demonstrate the effectiveness of our approach in recognizing diverse types of jewelry, we assembled a comprehensive database of accessory images. The evaluation process involved comparing various image captioning architectures, focusing particularly on the encoder decoder model, crucial for generating descriptive captions. After thorough evaluation, our final model achieved a captioning accuracy exceeding 90 per cent.