TY - GEN
T1 - Automated Image Description using VisualGPT
T2 - 2025 International Symposium on Intelligent Signal Processing and Communication Systems, ISPACS 2025
AU - Isnaini, Aghisna Nur
AU - Sulistyaningrum, Dwi Ratna
AU - Viadinugroho, Raden A.A.
AU - Subchan, Subchan
AU - Widianto, Muhammad Y.H.
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Fashion products have attracted many enthusiasts in the last decade, especially with the growing demand in ecommerce. This has led to the generation of automatic image descriptions for fashion products becoming a hot topic. However, effective multimodal data processing, integrating visual and textual information, remains a significant challenge. In this study, we implement VisualGPT using Vision Transformer (ViT) as visual feature extractor and GPT-2 as a text decoder. The FACAD170K dataset is used and divided into training, validation, and testing sets in a 7:2:1 ratio. Two experimental scenarios were conducted: (1) training non data augmentation and (2) training with data augmentation. We find that the augmented training scenario consistently outperformed the non-augmented scenario, achieving BLEU-1, BLEU-2, BLEU-3, and BLEU-4 scores of 46.66%, 30.85%, 24.37%, and 13.20 %, respectively, along with METEOR score 25 %, ROUGE-L score 38 %, and a CIDErr score of 82.52 %. These results show that VisualGPT has high accuracy and contextually relevant descriptions for fashion products.
AB - Fashion products have attracted many enthusiasts in the last decade, especially with the growing demand in ecommerce. This has led to the generation of automatic image descriptions for fashion products becoming a hot topic. However, effective multimodal data processing, integrating visual and textual information, remains a significant challenge. In this study, we implement VisualGPT using Vision Transformer (ViT) as visual feature extractor and GPT-2 as a text decoder. The FACAD170K dataset is used and divided into training, validation, and testing sets in a 7:2:1 ratio. Two experimental scenarios were conducted: (1) training non data augmentation and (2) training with data augmentation. We find that the augmented training scenario consistently outperformed the non-augmented scenario, achieving BLEU-1, BLEU-2, BLEU-3, and BLEU-4 scores of 46.66%, 30.85%, 24.37%, and 13.20 %, respectively, along with METEOR score 25 %, ROUGE-L score 38 %, and a CIDErr score of 82.52 %. These results show that VisualGPT has high accuracy and contextually relevant descriptions for fashion products.
KW - Clothing
KW - Fashion Products
KW - GPT-2
KW - Vision Transformer
KW - VisualGPT
UR - https://www.scopus.com/pages/publications/105034718664
U2 - 10.1109/ISPACS68724.2025.11382666
DO - 10.1109/ISPACS68724.2025.11382666
M3 - Conference contribution
AN - SCOPUS:105034718664
T3 - 2025 International Symposium on Intelligent Signal Processing and Communication Systems, ISPACS 2025
BT - 2025 International Symposium on Intelligent Signal Processing and Communication Systems, ISPACS 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 4 November 2025 through 7 November 2025
ER -