Skip to main navigation Skip to search Skip to main content

Automated Image Description using VisualGPT: Case Studies on Fashion Products

  • Institut Teknologi Sepuluh Nopember

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Fashion products have attracted many enthusiasts in the last decade, especially with the growing demand in ecommerce. This has led to the generation of automatic image descriptions for fashion products becoming a hot topic. However, effective multimodal data processing, integrating visual and textual information, remains a significant challenge. In this study, we implement VisualGPT using Vision Transformer (ViT) as visual feature extractor and GPT-2 as a text decoder. The FACAD170K dataset is used and divided into training, validation, and testing sets in a 7:2:1 ratio. Two experimental scenarios were conducted: (1) training non data augmentation and (2) training with data augmentation. We find that the augmented training scenario consistently outperformed the non-augmented scenario, achieving BLEU-1, BLEU-2, BLEU-3, and BLEU-4 scores of 46.66%, 30.85%, 24.37%, and 13.20 %, respectively, along with METEOR score 25 %, ROUGE-L score 38 %, and a CIDErr score of 82.52 %. These results show that VisualGPT has high accuracy and contextually relevant descriptions for fashion products.

Original languageEnglish
Title of host publication2025 International Symposium on Intelligent Signal Processing and Communication Systems, ISPACS 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331580667
DOIs
Publication statusPublished - 2025
Event2025 International Symposium on Intelligent Signal Processing and Communication Systems, ISPACS 2025 - Bandung, Indonesia
Duration: 4 Nov 20257 Nov 2025

Publication series

Name2025 International Symposium on Intelligent Signal Processing and Communication Systems, ISPACS 2025

Conference

Conference2025 International Symposium on Intelligent Signal Processing and Communication Systems, ISPACS 2025
Country/TerritoryIndonesia
CityBandung
Period4/11/257/11/25

Keywords

  • Clothing
  • Fashion Products
  • GPT-2
  • Vision Transformer
  • VisualGPT

Fingerprint

Dive into the research topics of 'Automated Image Description using VisualGPT: Case Studies on Fashion Products'. Together they form a unique fingerprint.

Cite this