跳到正文
arXiv:cs.CL· Pierre-Carl Langlais, Pavel Chizhov, Yannick Detrois, Carlos Rosas-Hinostroza, Ivan P. Yamshchikov, Bastien Perroy·· 6 小时前AI 评分35

法国合成社交媒体上的情绪分析:Model in Distress

Model in Distress: Sentiment Analysis on French Synthetic Social Media

AI 导读

一项研究用反向翻译与微调模型,从少量种子语料生成 170 万条合成推文及合成推理轨迹,用于法语公共交通客户困扰检测。基于此训练的 600M 参数推理模型在人工标注评测数据上达到 77-79% 准确率,追平或超过 SOTA 专有 LLM 与专用编码器。该流程无需暴露敏感用户数据,可推广到其他场景与语言。

正文

View PDF HTML (experimental)

Abstract:Automated analysis of customer feedback on social media is hindered by three challenges: the high cost of annotated training data, the scarcity of evaluation sets, especially in multilingual settings, and privacy concerns that prevent data sharing and reproducibility. We address these issues by developing a generalizable synthetic data generation pipeline applied to a case study on customer distress detection in French public transportation. Our approach utilizes backtranslation with fine-tuned models to generate 1.7 million synthetic tweets from a small seed corpus, complemented by synthetic reasoning traces. We train 600M-parameter reasoners with English and French reasoning that achieve 77-79% accuracy on human-annotated evaluation data, matching or exceeding SOTA proprietary LLMs and specialized encoders. Beyond reducing annotation costs, our pipeline preserves privacy by eliminating the exposure of sensitive user data. Our methodology can be adopted for other use cases and languages.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2604.18226 [cs.CL]
  (or arXiv:2604.18226v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2604.18226

arXiv-issued DOI via DataCite

Submission history

From: Pavel Chizhov [view email]
[v1] Mon, 20 Apr 2026 13:10:32 UTC (431 KB)
[v2] Mon, 5 Oct 2026 18:27:12 UTC (423 KB)

来源:arXiv:cs.CL · arxiv.org