跳到正文
arXiv:cs.LG· Samrendra Roy, Jason Yoo, Souvik Chakraborty, Syed Bahauddin Alam·· 4 小时前AI 评分36

目标搜索表明随机设备测试低估了模拟波动神经网络算子的最坏情况误差

Targeted search shows that random-device testing underestimates worst-case error in a simulated wave-based neural operator

AI 导读

一项针对混合傅里叶神经算子的数值案例研究显示,在模拟相干4f处理器上,通过搜索找到的设备在120个模型中的留出误差达到200个随机Monte Carlo设备最大值的1.08-3.10倍。

正文

View PDF HTML (experimental)

Abstract:Wave-based processors promise fast, energy-efficient Fourier layers for neural operators. They are usually validated on randomly sampled devices, but using them requires knowing how large their error can become under fabrication and alignment variation. In a stylised numerical case study, a hybrid Fourier neural operator runs its four spectral layers on simulated coherent 4f processors with 32 toleranced knobs, whose half-widths are representative rather than calibrated. For 120 models (four tasks, six training methods, five seeds), we compared the worst of N random in-spec devices with a searched one. On a deterministic simulator with one frozen draw of the random static errors, the searched device's held-out error was 1.08-3.10 times the maximum over 200 Monte Carlo devices and 1.06-2.71 times that over 1000. With 20 fresh static draws, it still exceeded the maximum over 200 random devices in 116 of 120 models. Under uniform sampling, the probability of drawing such a device is at most 0.37% per model (two-sided 95% Clopper-Pearson), which says nothing about how large its error is. The gap persisted with uniform or Sobol' sampling at the search's budget, shared knobs, a second crosstalk model, box scales of 0.25-2 and a pixel-level device model. Models trained only with random static errors reached 3.7-39.9 times their nominal error on searched devices, and fine-tuning on random and gradient-searched devices gave the lowest searched error of the six in all 20 task-seed pairs. For two heat-exchanger quantities, a search targeted at each exceeded the worst of 1000 random devices in all 39 models, and hence the Wilks 95/95 limit (worst of 59). For the mean pressure of 11 models, no random device exceeded a 1% error threshold, but the searched device did. Random testing estimates how often errors exceed a threshold; worst-device search gives a lower bound on how large they can be.
Comments: 50 pages (19 main text and references, 31 Supplementary Information), 5 figures, 1 table
Subjects: Machine Learning (cs.LG); Emerging Technologies (cs.ET); Optics (physics.optics)
Cite as: arXiv:2610.07529 [cs.LG]
  (or arXiv:2610.07529v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07529

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Samrendra Roy [view email]
[v1] Mon, 5 Oct 2026 23:53:31 UTC (2,030 KB)

来源:arXiv:cs.LG · arxiv.org