TEXT RECOGNITION FROM IMAGE AND VIDEO FRAME
DOI:
https://doi.org/10.62643/ijerst.2026.v22.n1.pp89-94Keywords:
CNN, Image processing, OCR, Video-to-frames conversionAbstract
Recognizing text from natural scene images and video frames is a challenging task due to uncontrolled imaging conditions such as background clutter, illumination variation, blur, low resolution, and diversity in text appearance. Traditional recognition pipelines commonly apply binarization prior to Optical Character Recognition (OCR), which often fails to preserve discriminative information in complex real-world environments. This paper introduces an adaptive text recognition framework that exploits automatic color channel selection to improve recognition accuracy in scene images and video data. Instead of relying on a single-color channel for an entire word image, the proposed method dynamically selects the most informative color channel for each sliding window based on local image characteristics. Text recognition is performed using a Hidden Markov Model (HMM) with Pyramidal Histogram of Oriented Gradient (PHOG) features extracted from the selected channel. A multi-label Support Vector Machine (SVM) classifier is employed to identify the optimal color channel, and multiple feature descriptors are investigated for this task. Experimental evaluation indicates that wavelet-based features provide the most reliable channel selection. The proposed framework is language-independent and has been validated on several standard English scene text datasets as well as a newly constructed Devanagari dataset. The results demonstrate consistent improvements over conventional recognition strategies.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.













