arXiv:cs.AI· Mirza Raquib, Munazer Montasir Akash, Tawhid Ahmed, Saydul Akbar Murad, Farida Siddiqi Prity, Mohammad Amzad Hossain, Asif Pervez Polok, Nick Rahimi·· 4 小时前AI 评分28
BERT-CNN-BiLSTM 统一框架实现孟加拉语新闻标题分类与情感分析
A Unified BERT-CNN-BiLSTM Framework for Simultaneous Headline Classification and Sentiment Analysis of Bangla News
AI 导读
研究提出 BERT-CNN-BiLSTM 混合迁移学习模型,在 BAN-ABSA 数据集 9014 条孟加拉语新闻标题上同时完成标题分类与情感分析,首次在该数据集上做此联合实验。两种采样策略中,拆分前过采样取得标题分类 78.57%、情感分析 73.43%;直接在原始不平衡数据上训练则达 81.37% 和 64.46%。该模型显著优于所有基线,为低资源孟加拉语文本分类树立新 SOTA。
正文
Abstract:In our daily lives, newspapers are an essential information source that impacts how the public talks about present-day issues. However, effectively navigating the vast amount of news content from different newspapers and online news portals can be challenging. Newspaper headlines with sentiment analysis tell us what the news is about (e.g., politics, sports) and how the news makes us feel (positive, negative, neutral). This helps us quickly understand the emotional tone of the news. This research presents a state-of-the-art approach to Bangla news headline classification combined with sentiment analysis applying Natural Language Processing (NLP) techniques, particularly the hybrid transfer learning model BERT-CNN-BiLSTM. We have explored a dataset called BAN-ABSA of 9014 news headlines, which is the first time that has been experimented with simultaneously in the headline and sentiment categorization in Bengali newspapers. Over this imbalanced dataset, we applied two experimental strategies: technique-1, where undersampling and oversampling are applied before splitting, and technique-2, where undersampling and oversampling are applied after splitting on the In technique-1 oversampling provided the strongest performance, both headline and sentiment, that is 78.57\% and 73.43\% respectively, while technique-2 delivered the highest result when trained directly on the original imbalanced dataset, both headline and sentiment, that is 81.37\% and 64.46\% respectively. The proposed model BERT-CNN-BiLSTM significantly outperforms all baseline models in classification tasks, and achieves new state-of-the-art results for Bangla news headline classification and sentiment analysis. These results demonstrate the importance of leveraging both the headline and sentiment datasets, and provide a strong baseline for Bangla text classification in low-resource.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2511.18618 [cs.CL] |
| (or arXiv:2511.18618v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2511.18618 arXiv-issued DOI via DataCite |
Submission history
From: Saydul Akbar Murad [view email]
[v1]
Sun, 23 Nov 2025 21:22:56 UTC (622 KB)
[v2]
Fri, 2 Oct 2026 04:20:33 UTC (1,105 KB)
来源:arXiv:cs.AI · arxiv.org