跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Zilong Zhao, Abdul Raheem, Jiayu Li, Sohei Arisaka, Darius Lim Hong Yi, Milad Abdollahzadeh, Uzair Javaid, Biplab Sikdar·· 5 小时前AI 评分43

GENSCRIPT:无需训练的纯推理合成数据流水线,统一支持表格、时序与关系数据

Synthesis Without Training: An Inference-Only Pipeline for Tabular, Temporal, and Relational Synthetic Data

AI 导读

GENSCRIPT 提出一种免训练的纯推理合成数据流水线,用语言模型从源数据的统计画像中推断字段语义与跨列完整性约束,再由编码智能体编译成可执行、可审计的采样器。

正文

View PDF HTML (experimental)

Abstract:Synthetic data generation is dominated by the fit-then-sample paradigm: a generative model is trained on a private dataset and then sampled from. Despite its widespread adoption, this paradigm faces three challenges: (1) a new training run is required for every dataset; (2) different data modalities, such as single tables, time series, and relational databases, require task-specific models and feature engineering; and (3) the resulting model is opaque, making its behavior under data constraints difficult to inspect. We propose GENSCRIPT, an inference-only pipeline that eliminates model training. GENSCRIPT computes a deterministic statistical profile of the source data (column types, ranges, missingness, categories, correlations, etc.) and passes it--rather than raw rows--to a language model to infer field semantics and cross-column integrity constraints. A coding agent then compiles the profile and constraints into an executable, auditable sampler. This unified approach supports single-table, temporal, and relational data without task-specific modeling. Across four single-table benchmarks, GENSCRIPT builds generators in 2 minutes and samples 50k rows within 6 seconds, while remaining within a few points of leading methods in marginal fidelity. Notably, it is the only method that perfectly preserves a 1-to-1 mapping between columns in the Adult dataset. On a smart-building dataset, it produces conditional time series that more closely match the real distribution than two baselines and perfectly preserves primary- and foreign-key relationships in the corresponding relational database.
Comments: Accepted at NeurIPS 2026 workshop: Beyond Private Training: The New Landscape of AI Privacy
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.38414 [cs.LG]
  (or arXiv:2609.38414v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.38414

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zilong Zhao [view email]
[v1] Tue, 29 Sep 2026 19:09:42 UTC (640 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org