arXiv:cs.CL· Rohit Kumar Sen, Anik Chowdhury·· 3 小时前AI 评分27
BanglaRhet:孟加拉语政治演讲中修辞与说服检测的经典模型和 Transformer 基准测试
BanglaRhet: Benchmarking Classical and Transformer Models for Rhetorical and Persuasion Detection in Bangla Political Speech
AI 导读
研究基于人工标注的 30,289 条孟加拉语政治演讲语料 BanglaRhet,对 BanglaBERT、BanglaBERT-Base、SahajBERT 和 XLM-RoBERTa-Base 四个 Transformer 模型与 TF-IDF 经典基线进行对比评测。
正文
Abstract:Political discourse often uses rhetorical and persuasive language to frame narratives, influence public opinion, and mobilize audiences. While Bangla natural language processing has made progress in sentiment analysis and opinion mining, systematic benchmarking of transformer models for fine-grained rhetorical and persuasion technique detection in Bangla political speech remains largely underexplored. This paper presents a benchmark study of transformer-based models for detecting rhetorical form and persuasive intent in Bangla political discourse. Using BanglaRhet, a manually annotated corpus of 30,289 Bangla political speech segments collected from publicly available political news sources, we formulate two supervised single-label classification tasks: rhetorical technique detection (contrast, repetition, exaggeration, metaphor, rhetorical questions) and persuasion technique detection (blame assignment, call to action, unity call, moral, emotional, and logical appeals). We evaluate four transformer-based models, BanglaBERT, BanglaBERT-Base, SahajBERT, and XLM-RoBERTa-Base, against classical TF-IDF baselines. BanglaBERT achieves the highest performance, with 65.40% macro-F1 for rhetorical technique detection and 66.46% for persuasion technique detection, outperforming the best tuned classical baseline by 19.2 and 13.8 macro-F1 points, respectively. Class-level analysis indicates that errors are mainly associated with semantic overlap among labels, figurative language, and class imbalance. The results provide initial benchmark baselines for Bangla rhetorical and persuasion-aware political discourse analysis and highlight the need for context-aware and multi-label modeling.
| Comments: | 6 pages, 2 figures, 6 tables. Accepted at the 2026 2nd International Conference on Advances in Computing, Communication, Electrical, and Smart Systems (iCACCESS), Dhaka, Bangladesh. Dataset: this https URL. Code: this https URL. Hugging Face: this https URL |
| Subjects: | Computation and Language (cs.CL) |
| ACM classes: | I.2.7 |
| Cite as: | arXiv:2610.09464 [cs.CL] |
| (or arXiv:2610.09464v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09464 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Rohit Kumar Sen [view email]
[v1]
Wed, 7 Oct 2026 05:22:19 UTC (1,616 KB)
来源:arXiv:cs.CL · arxiv.org