ENGLISH VISUAL QUESTION ANSWERING: BUILDING A CULTURALLY RELEVANT DATASET FROM IMAGE CAPTIONS

Authors

  • S SABIYA SULTANA1 , THATIKONDA KOTESH2 , PEDDA VEERANAGARI JOGENDHAR REDDY 3 , THALLOJU ABHISHEK 4 , EPPA RAMYA 5 Author

DOI:

https://doi.org/10.5281/zenodo.21102297

Abstract

Visual Question Answering (VQA) is a challenging multimodal task that requires joint understanding of visual content and natural language questions. Traditional VQA systems rely on complex attentionbased neural architectures that require significant computational resources and extensive GPU training. In this project, we propose an efficient and scalable Visual Question Answering system using a pretrained CLIP (Contrastive Language– Image Pretraining) model for open-ended English language queries. The proposed approach extracts semantically aligned image and question embeddings using a frozen CLIP backbone and combines them through a lightweight Multi-Layer Perceptron classifier for answer prediction. Experiments are conducted on the VizWiz dataset, a realworld VQA benchmark consisting of images captured by blind users. Despite being trained entirely on CPU without end-to-end finetuning, the proposed system achieves competitive performance, attaining a Top-1 accuracy of 40.6% and a Top-5 accuracy of 71.7% on the validation set. The results demonstrate that pretrained vision–language representations can provide strong VQA performance while significantly reducing computational cost, making the system suitable for deployment in low-resource environments.

Downloads

Published

29-06-2026

How to Cite

ENGLISH VISUAL QUESTION ANSWERING: BUILDING A CULTURALLY RELEVANT DATASET FROM IMAGE CAPTIONS. (2026). International Journal of Engineering Research and Science & Technology, 22(2(4), 579-586. https://doi.org/10.5281/zenodo.21102297