跳到正文
arXiv:cs.AI· Mohammad Zare, Pirooz Shamsinejadbabaki·· 3 小时前

ProtoSemImage:以图像原型和可变形行对齐实现可解释文档分类

ProtoSemImage: Image-Valued Prototypes with Deformable Row Alignment for Interpretable Document Classification

AI 导读

ProtoSemImage 用四通道 HSV 空间的图像原型替代向量原型进行文档分类,每个 token 作为像素、通道对应命名的语言学因素,并通过 Skip-Gram 目标端到端学习颜色空间。在十类任务上,它比同等条件的向量原型模型在三个配对随机种子上高出 4.3 至 11.8 分,但基于距离的 2D 模板匹配反而拖累性能,诊断显示 20.6 分的差距来自匹配而非颜色压缩。

正文

View PDF

Abstract:Prototypes in classification models are almost always vectors, and a vector has no readable form. This paper asks what happens when a prototype is an image. Documents give the question a natural form, because a document can be rendered as a multi-channel image in which every token becomes a pixel, so a class representative can take the same shape and the same channel semantics as the inputs it stands for. ProtoSemImage represents each class by one or more visual archetypes: prototype images in a four-channel HSV space whose channels carry named linguistic factors. A Skip-Gram objective learns that color space end to end through a four-dimensional bottleneck, discourse boundary rows become differentiable typed difference rows, and classification reduces to 2D visual template matching: a deformable row alignment between a document image and the archetype bank, in the spirit of dynamic time warping. Because the match is a spatial pattern comparison rather than a linear readout, the model reports where an input departs from its archetype and along which channel, and a generative head decodes each archetype back into text. The image representation works: it beats an otherwise identical model with vector prototypes in all three paired seeds, by between 4.3 and 11.8 points on a ten-class task. The distance-based matching does not. A diagnostic that keeps the representation fixed and swaps only the classifier recovers the sequence baselines, which locates a 20.6-point shortfall in the matching rather than in the color compression, and a benchmark built so that a pair of documents shares a bag of words and differs only in arrangement confirms the layout-preservation it was designed for. We report both directions, because for a representation whose whole purpose is inspect ability, the failure modes are as informative as the gains.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11460 [cs.CV]
  (or arXiv:2610.11460v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.11460

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Mohammad Zare [view email]
[v1] Thu, 8 Oct 2026 08:12:22 UTC (465 KB)

来源:arXiv:cs.AI · arxiv.org