跳到正文
arXiv:cs.AI· Cristian Martella, Angelo Martella, Antonella Longo, Motaz Saad·· 4 小时前AI 评分28

边缘环境下用小型语言模型做智能数据模型分类:一种成本感知的混合方法

Small Language Models for Smart Data Model Classification at the Edge: A Cost-Aware Hybrid Approach

AI 导读

一项研究评估了轻量级开源语言模型在资源受限边缘环境下将输入数据实体匹配到最合适智能数据模型(SDM)的表现,并系统性地对通用、推理专用和代码专用架构在多个领域数据集上进行了基准测试。研究还补充对比了LLM与TF-IDF及轻量级句子编码器两种近乎零成本的相似度基线。

正文

View PDF

Abstract:The rapid proliferation of heterogeneous data sources within the Internet of Things (IoT) across domains such as smart cities, energy management, and environmental monitoring necessitates efficient and scalable data standardization methods. Effective classification of smart data models (SDMs) is essential for facilitating interoperability. However, existing approaches are often limited by high resource consumption and lack applicability in edge environments with constrained computational capabilities. Aiming to bridge this gap, the proposed study evaluates the performance of lightweight open-source language models (LMs) to resolve an input data entity against its corresponding best fitting SDM representation under resource-constrained conditions. It systematically benchmarks a diverse array of models, including general purpose (GP), reasoning-specialized (RS), and code-specialized (CS) architectures, across multiple domain-specific datasets. Addressing the current omission of lightweight, resource-efficient solutions in the literature, the investigation provides significant and valuable insights into model selection, task formulation, and deployment strategies that optimize accuracy and efficiency. A complementary experiment also compares the surveyed large language models (LLMs) against two near-zero-cost similarity baselines (Term Frequency-Inverse Document Frequency (TF-IDF) and a lightweight sentence encoder) on the same task, providing a strong reference point for interpreting the practical value of LLM-based classification on edge platforms.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07093 [cs.AI]
  (or arXiv:2610.07093v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07093

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Motaz Saad [view email]
[v1] Mon, 5 Oct 2026 13:25:09 UTC (67 KB)

来源:arXiv:cs.AI · arxiv.org