跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Gabriele Martino, Denis Skibinski, Ivo L. Hofacker, Sebastian Tschiatschek·· 18 小时前AI 评分26

RiboUnmix:从有偏且含噪的 Ribo-seq 测量中学习共享翻译动态

RiboUnmix: Learning Shared Translational Dynamics from Biased and Noisy Ribo-seq Measurements

AI 导读

RiboUnmix 是一个概率多数据集框架,将每个预期测量谱表示为受数据集特定乘性因子调制的共享序列依赖信号,并用负二项观测模型刻画重复样本间变异。在合成基准上,其推断的共享成分与数据集特定效应均与目标强相关;在四个生物体真实数据基准上,RiboUnmix 预测测量谱的表现优于序列到谱基线。

正文

View PDF HTML (experimental)

Abstract:Ribosome profiling (Ribo-seq) measures ribosome distributions along mRNAs, but observed occupancy profiles also contain experiment-specific distortions and stochastic variability. Consequently, models that accurately predict measured profiles may reproduce technical effects rather than recover the underlying biology. We ask whether jointly modeling datasets collected under different experimental conditions can reveal shared, sequence-dependent patterns of ribosome occupancy. We introduce RiboUnmix, a probabilistic multi-dataset framework in which each expected measured profile is represented as a shared sequence-dependent signal modulated by a dataset-specific multiplicative factor. A negative-binomial observation model captures variability across replicates. We evaluate RiboUnmix on a controlled synthetic benchmark combining programmed translation kinetics, ribosome traffic, stochastic count sampling, and sequence-dependent experimental distortions. Because the underlying kinetics and distortions are known, recovery of the shared profile and dataset-specific effects can be assessed separately. Both inferred components correlate strongly with their targets, demonstrating that RiboUnmix can disentangle shared kinetic patterns from experimental effects. Across four organism-specific real-data benchmarks, RiboUnmix outperforms sequence-to-profile baselines in predicting measured profiles. Models trained independently on subsets of 114 HEK-derived datasets recover concordant shared profiles for held-out transcripts, and experiments varying the number and composition of training datasets show that the learned representation remains stable. RiboUnmix thus converts variation across experiments into evidence for reproducible sequence-dependent patterns of ribosome occupancy, supporting biological hypothesis generation from diverse Ribo-seq datasets.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.39644 [cs.LG]
  (or arXiv:2609.39644v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.39644

arXiv-issued DOI via DataCite

Submission history

From: Gabriele Martino [view email]
[v1] Wed, 30 Sep 2026 12:48:44 UTC (6,081 KB)
[v2] Thu, 1 Oct 2026 09:34:04 UTC (6,081 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org