arXiv:cs.CL· Vladimir Beskorovainyi·· 3 小时前
AI Appeals Processor:政府服务中公民申诉自动分类的深度学习方案
AI Appeals Processor: A Deep Learning Approach to Automated Classification of Citizen Appeals in Government Services
AI 导读
AI Appeals Processor 在纯 CPU 政府环境中部署,对 1 万条俄语跨领域公民申诉进行三分类对比,BERT 在 1500 条测试集上达 82% 准确率,Word2Vec+LSTM 为 78%,均高于人工操作员的 67%。
正文
Abstract:Government agencies must register, classify and route every citizen appeal within statutory time limits, and much of this work is still done by hand. We describe AI Appeals Processor, a classification and routing component deployed in a CPU-only government environment, and report what its evaluation and deployment taught us. On 10,000 real Russian-language appeals from a cross-domain dataset, we compare Bag-of-Words and TF-IDF with SVM, fastText, Word2Vec+LSTM and multilingual BERT on a three-way appeal-type task. On a held-out test set of 1,500 appeals, BERT reaches 82% accuracy and Word2Vec+LSTM 78%, against 67% for individual operators measured on an expert-adjudicated gold standard. We deployed the LSTM: in a workflow where an operator verifies every prediction, its lower training cost made frequent retraining on operator-verified labels practical, while the four-point accuracy gap did not change the operator's task. End-to-end handling time fell by 53-56% across four appeal-length bands (unweighted mean 22.5 to 10.25 minutes); model inference takes under two seconds of this. Most residual errors trace to the label taxonomy rather than the model: the statutory definitions of complaints and applications overlap, and many appeals carry two intents. A post-deployment audit of production classifications, made after several retraining cycles by operators who saw the assigned category, judged more than 95% correct; we explain why this figure is not comparable with the test-set result.
| Comments: | 11 pages, 6 tables. v2 corrects the reference list, adds confidence intervals, a description of the production audit, and sections on lessons from deployment, limitations and ethical considerations; all test-set results are unchanged |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| ACM classes: | I.2.7 |
| Cite as: | arXiv:2604.03672 [cs.CL] |
| (or arXiv:2604.03672v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2604.03672 arXiv-issued DOI via DataCite |
|
| Related DOI: | https://doi.org/10.5281/zenodo.19393274
DOI(s) linking to related resources |
Submission history
From: Vladimir Beskorovainyi [view email]
[v1]
Sat, 4 Apr 2026 10:03:53 UTC (12 KB)
[v2]
Thu, 8 Oct 2026 06:07:59 UTC (17 KB)
来源:arXiv:cs.CL · arxiv.org