Multimodal RAG: Text-Informed Cross-Modal Enhancement for Efficient Local Deployment
Journal
Lecture Notes in Computer Science
Journal Volume
16617 LNAI
Start Page
348
End Page
359
ISSN
0302-9743
1611-3349
ISBN
9789819219254
9789819219261
Date Issued
2026-06-30
Author(s)
Lin, Chia-Yi
Abstract
Multimodal Retrieval-Augmented Generation (RAG) systems face significant challenges in maintaining cross-modal coherence when processing documents containing both text and visual elements. Traditional approaches suffer from modal isolation, where text and image retrieval operate independently, often resulting in responses that combine unrelated textual and visual content from disparate sources. This problem is exacerbated in resource-constrained environments where small language models (SLMs) lack the sophisticated instruction-following capabilities required for complex cross-modal enhancement tasks. We propose Text-Informed Multimodal RAG, a novel approach that leverages text retrieval results to enhance image retrieval through text-informed cross-modal enhancement. Our method addresses the modal isolation problem by first performing text retrieval, then extracting discriminative keywords using TF-IDF analysis to create enhanced queries for image search. We introduce a new re-ranking mechanism that prioritizes images from the same source documents as retrieved text, ensuring cross-modal coherence. This approach circumvents the limitations of small vision-language models by employing deterministic keyword extraction rather than complex instruction following. Experimental evaluation under a paired-comparison design shows consistent improvements across multiple metrics within the evaluated setting. Our approach yields measurable improvements in image retrieval quality (MRR: +20.1%, nDCG: +22.0%), while maintaining computational efficiency suitable for local deployment scenarios. Comparative analysis indicates that our text-informed method provides more stable cross-modal retrieval performance than SLM-aided enhancement under small-model deployment constraints. These results suggest that structure-aware cross-modal enhancement can improve cross-modal consistency under resource-constrained multimodal RAG deployment settings.
Event(s)
30th Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2026
Subjects
Cross-Modal Enhancement
Document Understanding
Local Deployment
Multimodal RAG
Query Enhancement
Small Language Model
Vision-Language Models
Publisher
Springer Nature Singapore
Type
conference paper
