跳到正文
arXiv:cs.LG· Melika Shirian, Kianoosh Vadaei, Kian Majlessi, Audrina Ebrahimi, Peyman Adibi, Hossein Karshenas·· 4 小时前AI 评分37

PrismSSL:一个接口支持多模态,面向多模态自监督学习的单接口 Python 库

PrismSSL: One Interface, Many Modalities; A Single-Interface Library for Multimodal Self-Supervised Learning

AI 导读

PrismSSL 是一个统一音频、视觉、图与跨模态自监督学习方法的 Python 库,已发布在 PyPI 并采用 MIT 许可证。

正文

View PDF HTML (experimental)

Abstract:We present PrismSSL, a Python library that unifies state-of-the-art self-supervised learning (SSL) methods across audio, vision, graphs, and cross-modal settings in a single, modular codebase. The goal of the demo is to show how researchers and practitioners can: (i) install, configure, and run pretext training with a few lines of code; (ii) reproduce compact benchmarks; and (iii) extend the framework with new modalities or methods through clean trainer and dataset abstractions. PrismSSL is packaged on PyPI, released under the MIT license, integrates tightly with HuggingFace Transformers, and provides quality-of-life features such as distributed training in PyTorch, Optuna-based hyperparameter search, LoRA fine-tuning for Transformer backbones, animated embedding visualizations for sanity checks, Weights & Biases logging, and colorful, structured terminal logs for improved usability and clarity. In addition, PrismSSL offers a graphical dashboard - built with Flask and standard web technologies - that enables users to configure and launch training pipelines with minimal coding. The artifact (code and data recipes) will be publicly available and reproducible.
Subjects: Machine Learning (cs.LG); Multimedia (cs.MM)
Cite as: arXiv:2511.17776 [cs.LG]
  (or arXiv:2511.17776v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2511.17776

arXiv-issued DOI via DataCite

Submission history

From: Arshia Hemmat [view email]
[v1] Fri, 21 Nov 2025 20:47:50 UTC (2,347 KB)
[v2] Wed, 7 Oct 2026 11:04:05 UTC (2,348 KB)

来源:arXiv:cs.LG · arxiv.org