Hugging Face 发布 Transformers.js v3:支持 WebGPU、120 种架构与多运行时
Transformers.js v3: WebGPU Support, New Models & Tasks, and More…
Hugging Face 发布 Transformers.js v3,经一年多开发,核心更新包括 WebGPU 支持(官方称最高比 WASM 快 100 倍)、新的 dtype 量化格式、支持 120 种架构、Hub 上超过 1200 个预转换模型,以及 25 个新示例项目。
官方发布说明列出 WebGPU 加速、量化选项和运行时兼容等具体变化,读者可据此评估浏览器端模型部署的可行路径。
Transformers.js v3:WebGPU 支持、新模型与任务等更多内容…
发布于 2024 年 10 月 22 日
经过一年多的开发,我们很高兴地宣布 🤗 Transformers.js v3 发布! 亮点包括:
- WebGPU 支持(比 WASM 快达 100 倍!)
- 新的量化格式(dtypes)
- 共支持 120 种架构
- 25 个新的示例项目和模板
- Hugging Face Hub 上超过 1200 个预转换模型
- Node.js(ESM + CJS)、Deno 和 Bun 兼容性
- GitHub 和 NPM 上的新家
安装
你可以通过以下命令从 NPM 安装 Transformers.js v3 来开始使用:
npm i @huggingface/transformers
然后,使用以下方式导入该库
import { pipeline } from "@huggingface/transformers";
或者,通过 CDN
import { pipeline } from "https://cdn.jsdelivr.net/npm/@huggingface/transformers@3.0.0";
如需更多信息,请查看文档。
WebGPU 支持
WebGPU 是一项用于加速图形和计算的新 Web 标准。该 API 使 Web 开发者能够利用底层系统的 GPU 直接在浏览器中执行高性能计算。WebGPU 是 WebGL 的继任者,并提供显著更好的性能,因为它允许与现代 GPU 进行更直接的交互。最后,它支持通用 GPU 计算,这使其非常适合机器学习!
截至 2024 年 10 月,全球 WebGPU 支持率约为 70%(根据 caniuse.com),这意味着一些用户可能无法使用该 API。
如果以下演示在你的浏览器中无法运行,你可能需要使用功能标志来启用它:
在 Transformers.js v3 中的用法
得益于我们与 ONNX Runtime Web 的合作,启用 WebGPU 加速就像在加载模型时设置 device: 'webgpu' 一样简单。让我们看一些示例!
示例:在 WebGPU 上计算文本嵌入(演示)
import { pipeline } from "@huggingface/transformers";
// Create a feature-extraction pipeline
const extractor = await pipeline(
"feature-extraction",
"mixedbread-ai/mxbai-embed-xsmall-v1",
{ device: "webgpu" },
);
// Compute embeddings
const texts = ["Hello world!", "This is an example sentence."];
const embeddings = await extractor(texts, { pooling: "mean", normalize: true });
console.log(embeddings.tolist());
// [
// [-0.016986183822155, 0.03228696808218956, -0.0013630966423079371, ... ],
// [0.09050482511520386, 0.07207386940717697, 0.05762749910354614, ... ],
// ]
示例:在 WebGPU 上使用 OpenAI whisper 执行自动语音识别(演示)
import { pipeline } from "@huggingface/transformers";
// Create automatic speech recognition pipeline
const transcriber = await pipeline(
"automatic-speech-recognition",
"onnx-community/whisper-tiny.en",
{ device: "webgpu" },
);
// Transcribe audio from a URL
const url = "https://huggingface.co/datasets/Xenova/transformers.js-docs/resolve/main/jfk.wav";
const output = await transcriber(url);
console.log(output);
// { text: ' And so my fellow Americans ask not what your country can do for you, ask what you can do for your country.' }
示例:在 WebGPU 上使用 MobileNetV4 进行图像分类(演示)
import { pipeline } from "@huggingface/transformers";
// Create image classification pipeline
const classifier = await pipeline(
"image-classification",
"onnx-community/mobilenetv4_conv_small.e2400_r224_in1k",
{ device: "webgpu" },
);
// Classify an image from a URL
const url = "https://huggingface.co/datasets/Xenova/transformers.js-docs/resolve/main/tiger.jpg";
const output = await classifier(url);
console.log(output);
// [
// { label: 'tiger, Panthera tigris', score: 0.6149784922599792 },
// { label: 'tiger cat', score: 0.30281734466552734 },
// { label: 'tabby, tabby cat', score: 0.0019135422771796584 },
// { label: 'lynx, catamount', score: 0.0012161266058683395 },
// { label: 'Egyptian cat', score: 0.0011465961579233408 }
// ]
新的量化格式(dtypes)
在 Transformers.js v3 之前,我们使用 quantized 选项来指定是使用量化(q8)还是全精度(fp32)的模型变体,分别将 quantized 设置为 true 或 false。现在,我们通过 dtype 参数新增了从更大的列表中进行选择的能力。
可用的量化列表取决于模型,但一些常见的有:全精度("fp32")、半精度("fp16")、8 位("q8"、"int8"、"uint8")和 4 位("q4"、"bnb4"、"q4f16")。
(例如 mixedbread-ai/mxbai-embed-xsmall-v1)
基本用法
示例:以 4 位量化运行 Qwen2.5-0.5B-Instruct(演示)
import { pipeline } from "@huggingface/transformers";
// Create a text generation pipeline
const generator = await pipeline(
"text-generation",
"onnx-community/Qwen2.5-0.5B-Instruct",
{ dtype: "q4", device: "webgpu" },
);
// Define the list of messages
const messages = [
{ role: "system", content: "You are a helpful assistant." },
{ role: "user", content: "Tell me a funny joke." },
];
// Generate a response
const output = await generator(messages, { max_new_tokens: 128 });
console.log(output[0].generated_text.at(-1).content);
按模块设置 dtypes
一些编码器-解码器模型,如 Whisper 或 Florence-2,对量化设置极为敏感:尤其是编码器。因此,我们新增了按模块选择 dtypes 的能力,可以通过提供从模块名称到 dtype 的映射来实现。
示例:在 WebGPU 上运行 Florence-2(演示)
import { Florence2ForConditionalGeneration } from "@huggingface/transformers";
const model = await Florence2ForConditionalGeneration.from_pretrained(
"onnx-community/Florence-2-base-ft",
{
dtype: {
embed_tokens: "fp16",
vision_encoder: "fp16",
encoder_model: "q4",
decoder_model_merged: "q4",
},
device: "webgpu",
},
);

查看完整代码示例
import {
Florence2ForConditionalGeneration,
AutoProcessor,
AutoTokenizer,
RawImage,
} from "@huggingface/transformers";
// Load model, processor, and tokenizer
const model_id = "onnx-community/Florence-2-base-ft";
const model = await Florence2ForConditionalGeneration.from_pretrained(
model_id,
{
dtype: {
embed_tokens: "fp16",
vision_encoder: "fp16",
encoder_model: "q4",
decoder_model_merged: "q4",
},
device: "webgpu",
},
);
const processor = await AutoProcessor.from_pretrained(model_id);
const tokenizer = await AutoTokenizer.from_pretrained(model_id);
// Load image and prepare vision inputs
const url = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/tasks/car.jpg";
const image = await RawImage.fromURL(url);
const vision_inputs = await processor(image);
// Specify task and prepare text inputs
const task = "<MORE_DETAILED_CAPTION>";
const prompts = processor.construct_prompts(task);
const text_inputs = tokenizer(prompts);
// Generate text
const generated_ids = await model.generate({
...text_inputs,
...vision_inputs,
max_new_tokens: 100,
});
// Decode generated text
const generated_text = tokenizer.batch_decode(generated_ids, {
skip_special_tokens: false,
})[0];
// Post-process the generated text
const result = processor.post_process_generation(
generated_text,
task,
image.size,
);
console.log(result);
// { '<MORE_DETAILED_CAPTION>': 'A green car is parked in front of a tan building. The building has a brown door and two brown windows. The car is a two door and the door is closed. The green car has black tires.' }
支持 120 种架构
此版本将支持的架构总数增加到 120 种(见完整列表),涵盖广泛的输入模态和任务。值得注意的新名称包括:Phi-3、Gemma & Gemma 2、LLaVa、Moondream、Florence-2、MusicGen、Sapiens、Depth Pro、PyAnnote 和 RT-DETR。

新模型列表
- Cohere(来自 Cohere)随论文 Command-R: Retrieval Augmented Generation at Production Scale 发布,作者为 Cohere。
- Decision Transformer(来自 Berkeley/Facebook/Google)随论文 Decision Transformer: Reinforcement Learning via Sequence Modeling 发布,作者为 Lili Chen、Kevin Lu、Aravind Rajeswaran、Kimin Lee、Aditya Grover、Michael Laskin、Pieter Abbeel、Aravind Srinivas、Igor Mordatch。
- Depth Pro(来自 Apple)随论文 Depth Pro: Sharp Monocular Metric Depth in Less Than a Second 发布,作者为 Aleksei Bochkovskii、Amaël Delaunoy、Hugo Germain、Marcel Santos、Yichao Zhou、Stephan R. Richter、Vladlen Koltun。
- Florence2(来自 Microsoft)随论文 Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks 发布,作者为 Bin Xiao、Haiping Wu、Weijian Xu、Xiyang Dai、Houdong Hu、Yumao Lu、Michael Zeng、Ce Liu、Lu Yuan。
- Gemma(来自 Google)随论文 Gemma: Open Models Based on Gemini Technology and Research 发布,作者为 Gemma Google 团队。
- Gemma2(来自 Google)随论文 Gemma2: Open Models Based on Gemini Technology and Research 发布,作者为 Gemma Google 团队。
- Granite(来自 IBM)随论文 Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler 发布,作者为 Yikang Shen、Matthew Stallone、Mayank Mishra、Gaoyuan Zhang、Shawn Tan、Aditya Prasad、Adriana Meza Soria、David D. Cox、Rameswar Panda。
- GroupViT(来自 UCSD、NVIDIA)随论文 GroupViT: Semantic Segmentation Emerges from Text Supervision 发布,作者为 Jiarui Xu、Shalini De Mello、Sifei Liu、Wonmin Byeon、Thomas Breuel、Jan Kautz、Xiaolong Wang。
- Hiera(来自 Meta)随论文 Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles 发布,作者为 Chaitanya Ryali、Yuan-Ting Hu、Daniel Bolya、Chen Wei、Haoqi Fan、Po-Yao Huang、Vaibhav Aggarwal、Arkabandhu Chowdhury、Omid Poursaeed、Judy Hoffman、Jitendra Malik、Yanghao Li、Christoph Feichtenhofer。
- JAIS(来自 Core42)随论文 Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models 发布,作者为 Neha Sengupta、Sunil Kumar Sahu、Bokang Jia、Satheesh Katipomu、Haonan Li、Fajri Koto、William Marshall、Gurpreet Gosal、Cynthia Liu、Zhiming Chen、Osama Mohammed Afzal、Samta Kamboj、Onkar Pandit、Rahul Pal、Lalit Pradhan、Zain Muhammad Mujahid、Massa Baali、Xudong Han、Sondos Mahmoud Bsharat、Alham Fikri Aji、Zhiqiang Shen、Zhengzhong Liu、Natalia Vassilieva、Joel Hestness、Andy Hock、Andrew Feldman、Jonathan Lee、Andrew Jackson、Hector Xuguang Ren、Preslav Nakov、Timothy Baldwin、Eric Xing。
- LLaVa(来自 Microsoft Research 与 University of Wisconsin-Madison)随论文 Visual Instruction Tuning 发布,作者为 Haotian Liu、Chunyuan Li、Yuheng Li 和 Yong Jae Lee。
- MaskFormer(来自 Meta 和 UIUC)随论文 Per-Pixel Classification is Not All You Need for Semantic Segmentation 发布,作者为 Bowen Cheng、Alexander G. Schwing、Alexander Kirillov。
- MusicGen(来自 Meta)随论文 Simple and Controllable Music Generation 发布,作者为 Jade Copet、Felix Kreuk、Itai Gat、Tal Remez、David Kant、Gabriel Synnaeve、Yossi Adi 和 Alexandre Défossez。
- MobileCLIP(来自 Apple)随论文 MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training 发布,作者为 Pavan Kumar Anasosalu Vasu、Hadi Pouransari、Fartash Faghri、Raviteja Vemulapalli、Oncel Tuzel。
- MobileNetV1(来自 Google Inc.)随论文 MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications 发布,作者为 Andrew G. Howard、Menglong Zhu、Bo Chen、Dmitry Kalenichenko、Weijun Wang、Tobias Weyand、Marco Andreetto、Hartwig Adam。
- MobileNetV2(来自 Google Inc.)随论文 MobileNetV2: Inverted Residuals and Linear Bottlenecks 发布,作者为 Mark Sandler、Andrew Howard、Menglong Zhu、Andrey Zhmoginov、Liang-Chieh Chen。
- MobileNetV3(来自 Google Inc.)随论文 Searching for MobileNetV3 发布,作者为 Andrew Howard、Mark Sandler、Grace Chu、Liang-Chieh Chen、Bo Chen、Mingxing Tan、Weijun Wang、Yukun Zhu、Ruoming Pang、Vijay Vasudevan、Quoc V. Le、Hartwig Adam。
- MobileNetV4(来自 Google Inc.)随论文 MobileNetV4 - Universal Models for the Mobile Ecosystem 发布,作者为 Danfeng Qin、Chas Leichner、Manolis Delakis、Marco Fornoni、Shixin Luo、Fan Yang、Weijun Wang、Colby Banbury、Chengxi Ye、Berkin Akin、Vaibhav Aggarwal、Tenghui Zhu、Daniele Moro、Andrew Howard。
- Moondream1 由 vikhyat 在仓库 moondream 中发布。
- OpenELM(来自 Apple)随论文 OpenELM: An Efficient Language Model Family with Open-source Training and Inference Framework 发布,作者为 Sachin Mehta、Mohammad Hossein Sekhavat、Qingqing Cao、Maxwell Horton、Yanzi Jin、Chenfan Sun、Iman Mirzadeh、Mahyar Najibi、Dmitry Belenko、Peter Zatloukal、Mohammad Rastegari。
- Phi3(来自 Microsoft)随论文 Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone 发布,作者为 Marah Abdin、Sam Ade Jacobs、Ammar Ahmad Awan、Jyoti Aneja、Ahmed Awadallah、Hany Awadalla、Nguyen Bach、Amit Bahree、Arash Bakhtiari、Harkirat Behl、Alon Benhaim、Misha Bilenko、Johan Bjorck、Sébastien Bubeck、Martin Cai、Caio César Teodoro Mendes、Weizhu Chen、Vishrav Chaudhary、Parul Chopra、Allie Del Giorno、Gustavo de Rosa、Matthew Dixon、Ronen Eldan、Dan Iter、Amit Garg、Abhishek Goswami、Suriya Gunasekar、Emman Haider、Junheng Hao、Russell J. Hewett、Jamie Huynh、Mojan Javaheripi、Xin Jin、Piero Kauffmann、Nikos Karampatziakis、Dongwoo Kim、Mahoud Khademi、Lev Kurilenko、James R. Lee、Yin Tat Lee、Yuanzhi Li、Chen Liang、Weishung Liu、Eric Lin、Zeqi Lin、Piyush Madan、Arindam Mitra、Hardik Modi、Anh Nguyen、Brandon Norick、Barun Patra、Daniel Perez-Becker、Thomas Portet、Reid Pryzant、Heyang Qin、Marko Radmilac、Corby Rosset、Sambudha Roy、Olatunji Ruwase、Olli Saarikivi、Amin Saied、Adil Salim、Michael Santacroce、Shital Shah、Ning Shang、Hiteshi Sharma、Xia Song、Masahiro Tanaka、Xin Wang、Rachel Ward、Guanhua Wang、Philipp Witte、Michael Wyatt、Can Xu、Jiahang Xu、Sonali Yadav、Fan Yang、Ziyi Yang、Donghan Yu、Chengruidong Zhang、Cyril Zhang、Jianwen Zhang、Li Lyna Zhang、Yi Zhang、Yue Zhang、Yunan Zhang、Xiren Zhou。
- PVT(来自南京大学、香港大学等)随论文 Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions 发布,作者为 Wenhai Wang、Enze Xie、Xiang Li、Deng-Ping Fan、Kaitao Song、Ding Liang、Tong Lu、Ping Luo、Ling Shao。
- PyAnnote 在仓库 pyannote/pyannote-audio 中发布,作者为 Hervé Bredin。
- RT-DETR(来自百度),与论文 DETRs Beat YOLOs on Real-time Object Detection 一同发布,作者为 Yian Zhao、Wenyu Lv、Shangliang Xu、Jinman Wei、Guanzhong Wang、Qingqing Dang、Yi Liu、Jie Chen。
- Sapiens(来自 Meta AI)随论文 Sapiens: Foundation for Human Vision Models 发布,作者为 Rawal Khirodkar、Timur Bagautdinov、Julieta Martinez、Su Zhaoen、Austin James、Peter Selednik、Stuart Anderson、Shunsuke Saito。
- ViTMAE(来自 Meta AI)随论文 Masked Autoencoders Are Scalable Vision Learners 发布,作者为 Kaiming He、Xinlei Chen、Saining Xie、Yanghao Li、Piotr Dollár、Ross Girshick。
- ViTMSN(来自 Meta AI)随论文 Masked Siamese Networks for Label-Efficient Learning 发布,作者为 Mahmoud Assran、Mathilde Caron、Ishan Misra、Piotr Bojanowski、Florian Bordes、Pascal Vincent、Armand Joulin、Michael Rabbat、Nicolas Ballas。
示例项目和模板
作为本次发布的一部分,我们发布了 25 个新的示例项目和模板,主要聚焦于展示 WebGPU 支持!其中包括如下所示的 Phi-3.5 WebGPU 和 Whisper WebGPU 等演示。
我们正在将所有示例项目和演示迁移到 https://github.com/huggingface/transformers.js-examples,敬请关注相关更新!
![]() |
![]() |
|---|
超过 1200 个预转换模型
截至今天发布,社区已将超过 1200 个模型转换为与 Transformers.js 兼容!你可以在这里找到可用模型的完整列表。
如果你想转换自己的模型或微调模型,可以使用我们的转换脚本,方法如下:
python -m scripts.convert --quantize --model_id <model_name_or_path>
将生成的文件上传到 Hugging Face Hub 后,记得添加 transformers.js 标签,以便其他人能轻松找到并使用你的模型!

Node.js(ESM + CJS)、Deno 和 Bun 兼容性
Transformers.js v3 现在兼容三种最流行的服务器端 JavaScript 运行时:
| 运行时 | 描述 | 示例 |
|---|---|---|
| Node.js | 一个广泛使用的 JavaScript 运行时,基于 Chrome 的 V8 构建。它拥有庞大的生态系统,并支持多种库和框架。 | ESM 示例 / CJS 示例 |
| Deno | 一个面向 JavaScript 和 TypeScript 的现代运行时,默认安全。它使用 ES 模块,甚至支持实验性的 WebGPU。 | Deno 示例 |
| Bun | 一个为性能而优化的快速 JavaScript 运行时。它内置打包器、转译器和包管理器。 | Bun 示例 |
NPM 和 GitHub 上的新家
最后,我们很高兴地宣布,Transformers.js 现在将以 @huggingface/transformers 的名义在 NPM 上的 Hugging Face 官方组织下发布(而不是 v1 和 v2 使用的 @xenova/transformers)。
我们还将仓库迁移到了 GitHub 上的 Hugging Face 官方组织(https://github.com/huggingface/transformers.js),这将是我们的新家——欢迎来打个招呼!我们期待听到你的反馈、回复你的问题并审查你的 PR!
这是一个重要的里程碑,我们非常感谢社区帮助我们实现这一长期目标!没有你们所有人,这一切都不可能实现……谢谢!🤗
博客中的更多文章
社区
编辑预览
通过拖入文本输入框、粘贴或点击此处来上传图片、音频和视频。
点击或粘贴此处以上传图片
系统主题
公司
网站
来源:Hugging Face:Blog(RSS) · huggingface.co

