arXiv:cs.LG(机器学习,全量分类)· Faias Satter, Sk. Md. Masudul Ahsan·· 7 小时前AI 评分25
从转录孟加拉语文本进行开放词汇词识别
Open Vocabulary Word Recognition From Transcribed Bangla Texts
AI 导读
一项研究用 SSD+MobileNetV2、Faster R-CNN+InceptionResNetV2 及二者集成模型识别手写孟加拉语词图像,并引入改进的非极大值抑制提升效果。研究自建 9841 张手写孟加拉语词图像数据集,集成模型表现最佳,F1-score 达 92.61%,词级识别正确率为 96.12%。后续可通过加入后处理阶段纠正系统错误来进一步改进。
正文
Abstract:An optical character recognition (OCR) can scan a paper and extract text using technology, making people's jobs easier. While various OCR systems are available in the software industry, finding a reliable equivalent solution for Bangla takes much work. When it comes to handwritten texts, the situation is much more unusual. Recognizing words from word images is the most critical stage in any OCR process. It is the second stage after segmenting words from text pictures. If this stage fails, the overall performance of the OCR will be poor, regardless of how well the other phases perform. This study aims to recognize words using deep learning in a handwritten Bangla word image. Three object detection models, SSD with MobileNetV2, Faster R-CNN with InceptionResNetV2, and an ensemble model of these two, have been used to train and test handwritten word images. A modified Non-Maximum Suppression has been introduced to enhance the effectiveness of the models' results. A customized dataset of 9841 handwritten Bangla word images has been compiled, featuring diverse handwriting styles from various individuals. All three models' performances have been checked against the test dataset, and the ensemble model has been the most impressive, with an F1-score of 92.61%. Also, at the word level, the ensemble model correctly recognizes 96.12% of the words to some extent. The system can be further improved by introducing a post-processing phase to correct errors generated by the system.
| Comments: | 6 pages, 4 figures, 5 tables. Accepted version of the paper published in the 2023 26th International Conference on Computer and Information Technology (ICCIT). Code: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.01134 [cs.CV] |
| (or arXiv:2610.01134v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01134 arXiv-issued DOI via DataCite (pending registration) |
|
| Journal reference: | 2023 26th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 2023 |
| Related DOI: | https://doi.org/10.1109/ICCIT60459.2023.10441393
DOI(s) linking to related resources |
Submission history
From: Faias Satter [view email]
[v1]
Thu, 1 Oct 2026 06:14:41 UTC (396 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org