Image captioning bilingual berbasis BLIP-2 dan MarianMT dengan integrasi text-to-speech untuk literasi visual siswa
DOI:
https://doi.org/10.35760/tr.2026.v31i2.152Kata Kunci:
bilingual, BLIP-2, image captioning, MarianMT, visual literacyAbstrak
Visual literacy is the ability to understand and interpret information from visual representations, which involves linguistic processing. In the context of elementary education, this ability is crucial for helping students associate visual objects with bilingual vocabulary. However, most image captioning research remains monolingual, few studies accommodate Indonesian, and none have integrated audio representations into a single multimodal learning system. This study develops a bilingual image captioning system based on BLIP-2 and MarianMT, integrated with Text-to-Speech (TTS) within a FastAPI-based web application. English captions are generated using a pretrained BLIP-2 model, then translated into Indonesian using MarianMT, and converted to audio using gTTS. Additionally, a QLoRA fine-tuning experiment was conducted to compare model performance. The dataset consists of 6400 animal images relevant to the elementary school learning context, with caption quality evaluated using the METEOR metric. The results show that the pretrained BLIP-2 model delivers relatively stable performance with a METEOR score of 0.3765 for English and 0.3295 for Indonesian. These scores indicate that the generated captions are semantically relevant, although the overall performance is still relatively moderate compared to advanced image captioning systems. Functional testing of the prototype involving five elementary school students showed that the system is capable of generating bilingual captions and audio in real time and is easy to use. This multimodal integration supports the visual–verbal association process and has the potential to enrich students’ bilingual vocabulary, although no controlled experimental testing has yet been conducted to quantitatively measure improvements in literacy.
Unduhan
Referensi
[1] T. Xian, Z. Zhou, W. Zhou, and Z. Zhang, “Refining visual token sequence for efficient image captioning,” Neural Netw., Art. no. 107759, vol. 191, Nov. 2025, doi: 10.1016/j.neunet.2025.107759.
[2] J. Li, D. Li, C. Xiong, and S. Hoi, “BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in Proc. 39th Int. Conf. Mach. Learn. (ICML), vol. 162, Baltimore, MD, USA, Jul. 17–23, 2022, pp. 12888–12900.
[3] M. Cornia, L. Baraldi, A. Tal, and R. Cucchiara, “Fully-attentive iterative networks for region-based controllable image and video captioning,” Comput. Vis. Image Underst., vol. 237, Art. no. 103857, Dec. 2023, doi: 10.1016/j.cviu.2023.103857.
[4] H. Fadhilah and N. P. Utama, “Systematic literature review on medical image captioning using CNN-LSTM and transformer-based models,” Jurnal Masyarakat Informatika, vol. 16, no. 1, pp. 32–53, May 2025, doi: 10.14710/jmasif.16.1.73127.
[5] Z. Gao et al., “AnDR-BLIP2: Enhanced semantic understanding framework for industrial image anomaly detection and report generation,” J. Franklin Inst., vol. 362, no. 12, Art. no. 107816, Aug. 2025, doi: 10.1016/j.jfranklin.2025.107816.
[6] R. Castro, I. Pineda, W. Lim, and M. E. Morocho-Cayamcela, “Deep learning approaches based on transformer architectures for image captioning tasks,” IEEE Access, vol. 10, pp. 33679–33694, 2022, doi: 10.1109/ACCESS.2022.3161428.
[7] R. Yilmaz and F. G. K. Yilmaz, “The effect of generative artificial intelligence (AI)-based tool use on students’ computational thinking skills, programming self-efficacy and motivation,” Comput. Educ. Artif. Intell., vol. 4, Art. no. 100147, Jun. 2023, doi: 10.1016/j.caeai.2023.100147.
[8] A. K. Poddar and R. Rani, “Hybrid architecture using CNN and LSTM for image captioning in Hindi language,” in Procedia Comput. Sci., vol. 218, pp. 686–696, 2023, doi: 10.1016/j.procs.2023.01.049.
[9] J. A. Alzubi, R. Jain, P. Nagrath, S. Satapathy, S. Taneja, and P. Gupta, “Deep image captioning using an ensemble of CNN and LSTM based deep neural networks,” J. Intell. Fuzzy Syst., vol. 40, no. 4, pp. 5761–5769, 2021, doi: 10.3233/JIFS-189415.
[10] I. A. Albadarneh, B. H. Hammo, and O. S. Al-Kadi, “Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation,” Comput. Sci. Rev., vol. 58, Art. no. 100766, Nov. 2025, doi: 10.1016/j.cosrev.2025.100766.
[11] M. A. Al-Malla, O. Hamdoun, and N. Ghneim, “A comprehensive survey on deep learning approaches for image captioning: A systematic review,” J. Big Data, vol. 13, Art. no. 48, Feb. 2026, doi: 10.1186/s40537-026-01377-w.
[12] L. Xu, Q. Tang, J. Lv, B. Zheng, X. Zeng, and W. Li, “Deep image captioning: A review of methods, trends and future challenges,” Neurocomputing, vol. 546, Art. no. 126287, Aug. 2023, doi: 10.1016/j.neucom.2023.126287.
[13] M. Cho, S. Kim, D. Choi, and Y. Sung, “Enhanced BLIP-2 optimization using LoRA for generating dashcam captions,” Appl. Sci., vol. 15, no. 7, Art. no. 3712, Apr. 2025, doi: 10.3390/app15073712.
[14] J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in Proc. 40th Int. Conf. Mach. Learn. (ICML), vol. 202, Honolulu, HI, USA, Jul. 2023, Art. no. 814, pp. 19730–19742.
[15] S. Tyagi et al., “Novel advance image caption generation utilizing vision transformer and generative adversarial networks,” Computers, vol. 13, no. 12, Art. no. 305, Nov. 2024, doi: 10.3390/computers13120305.
[16] B. Brzȩk and G. Dziczkowski, “Efficient medical question answering through QLoRA fine-tuning and knowledge distillation,” Procedia Comput. Sci., vol. 270, pp. 1468–1477, 2025, doi: 10.1016/j.procs.2025.09.268.
[17] F. L. Chen et al., “VLP: A survey on vision-language pre-training,” Mach. Intell. Res., vol. 20, pp. 38–56, Jan. 2023, doi: 10.1007/s11633-022-1369-5.
[18] R. Patil, P. Khot, and V. Gudivada, “Analyzing LLAMA3 performance on classification task using LoRA and QLoRA techniques,” Appl. Sci., vol. 15, no. 6, Art. no. 3087, Mar. 2025, doi: 10.3390/app15063087.
[19] R. Mulyawan, A. Sunyoto, and A. H. Muhammad, “Automatic Indonesian image captioning using CNN and transformer-Based model approach,” in Proc. 5th Int. Conf. Inf. Commun. Technol. (ICOIACT), Yogyakarta, Indonesia, 2022, pp. 355–360, doi: 10.1109/ICOIACT55506.2022.9971855.
[20] M. Lupaşcu, A.-C. Rogoz, M. S. Stupariu, and R. T. Ionescu, “Large multimodal models for low-resource languages: A survey,” Inf. Fusion, vol. 131, Art. no. 104189, Jul. 2026, doi: 10.1016/j.inffus.2026.104189.
[21] P. Choudhury, S. Nair, P. Guha, and S. Nandi, “Image captioning in low resource Assamese language with semantic information prior and spatially encoded transformer model,” Expert Syst. Appl., vol. 297, pt. C, Art. no. 129479, Feb. 2026, doi: 10.1016/j.eswa.2025.129479.
[22] Z. Tan et al., “Neural machine translation: A review of methods, resources, and tools,” AI Open, vol. 1, pp. 5–21, 2020, doi: 10.1016/j.aiopen.2020.11.001.
[23] F. I. Maulana, Y. Heryadi, G. P. Kusuma, and W. Budiharto, “Data augmentation English-Indonesia-Madurese parallel corpus dataset using neural machine translation,” Data Brief, vol. 62, Art. no. 112046, 2025, doi: 10.1016/j.dib.2025.112046.
[24] S. A. Roslyn, A. Negha, and R. V. Sekhar, “Enhancing accessibility for visually impaired users: A BLIP2-powered image description system in Tamil,” in Proc. Int. Conf. Adv. Data Eng. Intell. Comput. Syst. (ADICS), Chennai, India, 2024, pp. 1–6, doi: 10.1109/ADICS58448.2024.10533565.
[25] J.-H. Kim, S.-W. Park, J.-H. Huh, S.-H. Jung, and C.-B. Sim, “Human scene understanding mechanism-based image captioning for blind assistance,” IEEE Access, vol. 13, pp. 81933–81947, 2025, doi: 10.1109/ACCESS.2025.3564991.
[26] A. Alsayed, M. Arif, T. M. Qadah, and S. Alotaibi, “A systematic literature review on using the encoder-decoder models for image captioning in English and Arabic languages,” Appl. Sci., vol. 13, no. 19, Art. no. 10894, Sep. 2023, doi: 10.3390/app131910894.
[27] Y. Tao et al., “CMDF-TTS: Text-to-speech method with limited target speaker corpus,” Neural Netw., vol. 188, Art. no. 107432, Aug. 2025, doi: 10.1016/j.neunet.2025.107432.
[28] O. Ondeng, H. Ouma, and P. Akuon, “A review of transformer-based approaches for image captioning,” Appl. Sci., vol. 13, no. 19, Art. no. 11103, Oct. 2023, doi: 10.3390/app131911103.
[29] K. R. Narejo et al., “Optimizing sentiment integration in image captioning using transformer-based fusion strategies,” Comput. Mater. Continua (CMC), vol. 84, no. 2, pp. 3407–3429, Jul. 2025, doi: 10.32604/cmc.2025.065872.
[30] A. Thobhani et al., “A survey on enhancing image captioning with advanced strategies and techniques,” CMES Comput. Model. Eng. Sci., vol. 142, no. 3, pp. 2247–2280, Mar. 2025, doi: 10.32604/cmes.2025.059192.
[31] J. Peng and T. Tang, “A unified visual and linguistic semantics method for enhanced image captioning,” Appl. Sci., vol. 14, no. 6, Art. no. 2657, Mar. 2024, doi: 10.3390/app14062657.
[32] A. S. Al-Shamayleh, O. Adwan, M. A. Alsharaiah, A. H. Hussein, Q. M. Kharma, and C. I. Eke, “A comprehensive literature review on image captioning methods and metrics based on deep learning technique,” Multimed. Tools Appl., vol. 83, pp. 34219–34268, Apr. 2024, doi: 10.1007/s11042-024-18307-8.
[33] M. Kaur and H. Kaur, “An efficient CNN-LSTM based framework for improved image captioning,” in Procedia Comput. Sci., vol. 258, pp. 3601–3607, 2025, doi: 10.1016/j.procs.2025.04.615.
[34] S. Banerjee, “Animal Image Dataset (90 Different Animals),” Kaggle, 2021. [Online]. Available: https://www.kaggle.com/datasets/iamsouravbanerjee/animal-image-dataset-90-different-animals. (Accessed: Oct. 24, 2025).
Unduhan
Diterbitkan
Terbitan
Bagian
Lisensi
Hak Cipta (c) 2026 Jurnal Ilmiah Teknologi dan Rekayasa

Artikel ini berlisensi Creative Commons Attribution 4.0 International License.
Universitas Gunadarma 