VISUAL AND TEXTUAL INTELLIGENCE: A DEEP LEARNING APPROACH TO

Authors

  • Ms. Bibi Noreen Ayesha,Dr. Fahmina Taranum Author

DOI:

https://doi.org/10.62643/

Abstract

In some extreme conditions, like low light environments, limited exposure, etc., images are highly noisy and have a low signal-to-noise ratio, which poses significant challenges in artificial intelligence, particularly for image denoising and reconstruction. The traditional approaches rely on handcrafted image priors and noise models, but they often fail to capture the fine details and structural integrity in highly degraded images. To overcome these limitations, we introduce a text-guided architecture with semantic scene description as additional prior knowledge for better understanding of the context of objects, textures and spatial relationships. It employs a diffusion-based network architecture called DDPM in the raw picture domain that is more suitable to capture the sensor-level noise properties. For multimodal learning, CLIP is realized to ensure that the textual and visual representations are aligned for effective cross-modal guiding in reconstruction. It is first pre-trained on synthetically created noisy data, and then fine-tuned on real-world noisy images to improve the generalization capability with different camera settings by incorporating low-rank adaptation (LoRA). Experimental results indicate that the inclusion of textual guidance causes the perceptual quality to be improved, fine details to be preserved, and that there is semantic consistency, as confirmed by the CLIP score and BLEU.

Downloads

Published

06-07-2026

How to Cite

VISUAL AND TEXTUAL INTELLIGENCE: A DEEP LEARNING APPROACH TO . (2026). International Journal of Engineering Research and Science & Technology, 22(3(1), 27-32. https://doi.org/10.62643/