arXiv:cs.LG· Yihao Ding, Daniel Yitian Su, Yiran Zhang, Christopher M. Gonzalez, Wei Liu·· 2 天前AI 评分32
DrillBench:分布偏移下的自回归钻孔建模基准
Autoregressive Drillhole Modelling Under Distribution Shift
AI 导读
研究团队提出 DrillBench,一个包含 49,671 个西澳大利亚钻孔的基准,用于下一层岩性预测与自回归地层生成,覆盖从局部预测、空间偏移到跨地质省迁移的递进迁移谱系。
正文
Abstract:Autoregressive modelling has achieved remarkable success in language and sequence tasks by learning to predict future states from previous observation. Mineral-exploration drillholes provide a natural but largely unexplored setting for this paradigm: as drilling proceeds, lithology is revealed sequentially from shallow to deep, making prediction of deeper strata inherently autoregressive. Existing drillhole modelling, however, is dominated by spatial interpolation and reconstruction, or largely rely on masked modelling, leaving strictly autoregressive prediction largely underexplored. We introduce DrillBench, a benchmark of 49,671 Western Australian drillholes for next-layer prediction and autoregressive stratigraphic generation across a graded transfer spectrum, from local prediction through spatial shift to cross geological province transfer. Benchmarking classical, geostatistical, and neural models reveals a clear \emph{transfer boundary}: spatial and geochemical conditioning provides large local gains but deteriorates sharply under stronger shift, whereas lithology-sequence autoregressive models transfer more robustly. Guided by this finding, we develop a backbone-agnostic recipe combining large-scale pretraining on historical drillholes with spatial retrieval of neighbouring lithology. Retrieval is most effective in weathered cover, when local spatial continuity remains informative, whereas pretraining contributes more strongly in bedrock and under broader geological shift. Together, they retain strong local performance while improving generalisation under spatial and cross-province shift, most markedly on the most distant splits. The benchmark and code are available at this https URL.
| Comments: | work in progress |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.01204 [cs.LG] |
| (or arXiv:2610.01204v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01204 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Daniel Su [view email]
[v1]
Thu, 1 Oct 2026 07:15:36 UTC (1,733 KB)
来源:arXiv:cs.LG · arxiv.org