Hugging Face Optimum 1.2 发布:支持 Transformers pipelines 加速推理
Accelerated Inference with Optimum and Transformers Pipelines
Hugging Face 发布 Optimum 1.2,为 Optimum 增加推理支持,可通过 ORTModelForXxx 类配合 ONNX Runtime 使用 transformers pipelines。
官方教程给出从转换、优化到量化的完整流程和实测数据,延迟约减半而精度保留 99.61%,方法可直接迁移。
Inference 已加入 Optimum,支持 Hugging Face Transformers pipelines,包括使用 ONNX Runtime 的文本生成。
BERT 和 Transformers 的采用持续增长。基于 Transformer 的模型如今不仅在自然语言处理领域达到最先进水平,在计算机视觉、语音和时间序列领域也是如此。💬 🖼 🎤 ⏳
企业正从实验和研究阶段转向生产阶段,以便将 Transformer 模型用于大规模工作负载。但默认情况下,与传统机器学习算法相比,BERT 及其同类模型相对较慢、较大且复杂。
为解决这一挑战,我们创建了 Optimum —— Hugging Face Transformers 的扩展,用于加速 BERT 等 Transformer 模型的训练和推理。
在这篇博客文章中,你将了解到:
- 1. 什么是 Optimum?用简单的话解释
- 2. 新的 Optimum 推理和 pipeline 功能
- 3. 端到端教程:加速用于问答的 RoBERTa,包括量化和优化
- 4. 当前限制
- 5. Optimum 推理常见问题
- 6. 下一步是什么?
让我们开始吧!🚀
1. 什么是 Optimum?用简单的话解释
Hugging Face Optimum 是一个开源库,也是 Hugging Face Transformers 的扩展,它提供性能优化工具的统一 API,以在加速硬件上实现最高效的模型训练和运行,包括用于在 Graphcore IPU 和 Habana Gaudi 上优化性能的工具包。Optimum 可用于加速训练、量化、图优化,现在还支持 transformers pipelines 的推理。
2. 新的 Optimum 推理和 pipeline 功能
随着 Optimum 1.2 的 发布,我们添加了对 推理 和 transformers pipelines 的支持。这使 Optimum 用户能够利用他们从 transformers 中熟悉的相同 API,同时借助 ONNX Runtime 等加速运行时的能力。
从 Transformers 切换到 Optimum Inference
Optimum Inference 模型 与 Hugging Face Transformers 模型 API 兼容。这意味着你只需将 AutoModelForXxx 类替换为 Optimum 中对应的 ORTModelForXxx 类。例如,以下是如何在 Optimum 中使用问答模型:
from transformers import AutoTokenizer, pipeline
-from transformers import AutoModelForQuestionAnswering
+from optimum.onnxruntime import ORTModelForQuestionAnswering
-model = AutoModelForQuestionAnswering.from_pretrained("deepset/roberta-base-squad2") # pytorch checkpoint
+model = ORTModelForQuestionAnswering.from_pretrained("optimum/roberta-base-squad2") # onnx checkpoint
tokenizer = AutoTokenizer.from_pretrained("deepset/roberta-base-squad2")
optimum_qa = pipeline("question-answering", model=model, tokenizer=tokenizer)
question = "What's my name?"
context = "My name is Philipp and I live in Nuremberg."
pred = optimum_qa(question, context)
在第一个版本中,我们添加了 对 ONNX Runtime 的支持,但还有更多功能即将推出!
这些新的 ORTModelForXX 现在可以与 transformers pipelines 一起使用。它们也完全集成到 Hugging Face Hub 中,以便从社区推送和拉取优化后的检查点。除此之外,你可以使用 ORTQuantizer 和 ORTOptimizer 先对模型进行量化和优化,然后对其运行推理。
查看 端到端教程:加速用于问答的 RoBERTa,包括量化和优化 了解更多详情。
3. 端到端教程:加速用于问答的 RoBERTa,包括量化和优化
在这个关于加速用于问答的 RoBERTa 的端到端教程中,你将学习如何:
- 为 ONNX Runtime 安装
Optimum - 将 Hugging Face
Transformers模型转换为 ONNX 以进行推理 - 使用
ORTOptimizer优化模型 - 使用
ORTQuantizer来应用动态量化 - 使用Transformers pipelines运行加速推理
- 评估性能和速度
让我们开始吧 🚀
本教程是在m5.xlarge AWS EC2实例上创建并运行的。
3.1 为Onnxruntime安装Optimum
我们的第一步是安装Optimum以及onnxruntime实用工具。
pip install "optimum[onnxruntime]==1.2.0"
这将为我们安装所有必需的包,包括transformers、torch和onnxruntime。如果你打算使用GPU,可以使用pip install optimum[onnxruntime-gpu]来安装optimum。
3.2 将Hugging Face Transformers模型转换为ONNX以进行推理**
在开始优化之前,我们需要将原始的transformers模型转换为onnx格式。为此,我们将使用新的ORTModelForQuestionAnswering类,调用from_pretrained()方法并传入from_transformers属性。我们使用的模型是deepset/roberta-base-squad2,这是一个在SQUAD2数据集上微调的RoBERTa模型,F1分数达到82.91,特征(任务)为question-answering。
from pathlib import Path
from transformers import AutoTokenizer, pipeline
from optimum.onnxruntime import ORTModelForQuestionAnswering
model_id = "deepset/roberta-base-squad2"
onnx_path = Path("onnx")
task = "question-answering"
# load vanilla transformers and convert to onnx
model = ORTModelForQuestionAnswering.from_pretrained(model_id, from_transformers=True)
tokenizer = AutoTokenizer.from_pretrained(model_id)
# save onnx checkpoint and tokenizer
model.save_pretrained(onnx_path)
tokenizer.save_pretrained(onnx_path)
# test the model with using transformers pipeline, with handle_impossible_answer for squad_v2
optimum_qa = pipeline(task, model=model, tokenizer=tokenizer, handle_impossible_answer=True)
prediction = optimum_qa(question="What's my name?", context="My name is Philipp and I live in Nuremberg.")
print(prediction)
# {'score': 0.9041663408279419, 'start': 11, 'end': 18, 'answer': 'Philipp'}
我们成功地将原始transformers转换为onnx,并使用该模型配合transformers.pipelines运行了第一次预测。现在让我们来优化它。🏎
如果你想了解更多关于导出transformers模型的信息,请查看文档:导出 🤗 Transformers 模型
3.3 使用ORTOptimizer优化模型
在我们将onnx检查点保存到onnx/之后,现在可以使用ORTOptimizer来应用图优化,例如算子融合和常量折叠,以加速延迟和推理。
from optimum.onnxruntime import ORTOptimizer
from optimum.onnxruntime.configuration import OptimizationConfig
# create ORTOptimizer and define optimization configuration
optimizer = ORTOptimizer.from_pretrained(model_id, feature=task)
optimization_config = OptimizationConfig(optimization_level=99) # enable all optimizations
# apply the optimization configuration to the model
optimizer.export(
onnx_model_path=onnx_path / "model.onnx",
onnx_optimized_model_output_path=onnx_path / "model-optimized.onnx",
optimization_config=optimization_config,
)
为了测试性能,我们可以再次使用ORTModelForQuestionAnswering类,并提供一个额外的file_name参数来加载我们优化后的模型。(这也适用于hub上可用的模型)。
from optimum.onnxruntime import ORTModelForQuestionAnswering
# load quantized model
opt_model = ORTModelForQuestionAnswering.from_pretrained(onnx_path, file_name="model-optimized.onnx")
# test the quantized model with using transformers pipeline
opt_optimum_qa = pipeline(task, model=opt_model, tokenizer=tokenizer, handle_impossible_answer=True)
prediction = opt_optimum_qa(question="What's my name?", context="My name is Philipp and I live in Nuremberg.")
print(prediction)
# {'score': 0.9041663408279419, 'start': 11, 'end': 18, 'answer': 'Philipp'}
我们将在步骤3.6 评估性能和速度中详细评估性能变化。
3.4 使用ORTQuantizer应用动态量化
在我们优化模型之后,可以通过使用ORTQuantizer对其进行量化来进一步加速。ORTOptimizer可用于应用动态量化,以减小模型大小并加速延迟和推理。
我们使用avx512_vnni,因为该实例由支持avx512的Intel cascade-lake CPU驱动。
from optimum.onnxruntime import ORTQuantizer
from optimum.onnxruntime.configuration import AutoQuantizationConfig
# create ORTQuantizer and define quantization configuration
quantizer = ORTQuantizer.from_pretrained(model_id, feature=task)
qconfig = AutoQuantizationConfig.avx512_vnni(is_static=False, per_channel=True)
# apply the quantization configuration to the model
quantizer.export(
onnx_model_path=onnx_path / "model-optimized.onnx",
onnx_quantized_model_output_path=onnx_path / "model-quantized.onnx",
quantization_config=qconfig,
)
我们现在可以比较这个模型的大小以及一些延迟性能
import os
# get model file size
size = os.path.getsize(onnx_path / "model.onnx")/(1024*1024)
print(f"Vanilla Onnx Model file size: {size:.2f} MB")
size = os.path.getsize(onnx_path / "model-quantized.onnx")/(1024*1024)
print(f"Quantized Onnx Model file size: {size:.2f} MB")
# Vanilla Onnx Model file size: 473.31 MB
# Quantized Onnx Model file size: 291.77 MB
我们将模型大小减少了近50%,从473MB降至291MB。要运行推理,我们可以再次使用ORTModelForQuestionAnswering类,并提供一个额外的file_name参数来加载我们量化后的模型。(这也适用于hub上可用的模型)。
# load quantized model
quantized_model = ORTModelForQuestionAnswering.from_pretrained(onnx_path, file_name="model-quantized.onnx")
# test the quantized model with using transformers pipeline
quantized_optimum_qa = pipeline(task, model=quantized_model, tokenizer=tokenizer, handle_impossible_answer=True)
prediction = quantized_optimum_qa(question="What's my name?", context="My name is Philipp and I live in Nuremberg.")
print(prediction)
# {'score': 0.9246969819068909, 'start': 11, 'end': 18, 'answer': 'Philipp'}
很好!模型预测出了相同的答案。
3.5 使用Transformers pipelines运行加速推理
Optimum内置支持transformers pipelines。这使我们能够利用与使用PyTorch和TensorFlow模型时相同的API。我们已经在步骤3.2、3.3和3.4中使用了此功能来测试我们转换和优化后的模型。在撰写本文时,我们支持ONNX Runtime,未来还会支持更多。下面是一个如何使用transformers pipelines的示例。
from transformers import AutoTokenizer, pipeline
from optimum.onnxruntime import ORTModelForQuestionAnswering
tokenizer = AutoTokenizer.from_pretrained(onnx_path)
model = ORTModelForQuestionAnswering.from_pretrained(onnx_path)
optimum_qa = pipeline("question-answering", model=model, tokenizer=tokenizer)
prediction = optimum_qa(question="What's my name?", context="My name is Philipp and I live in Nuremberg.")
print(prediction)
# {'score': 0.9041663408279419, 'start': 11, 'end': 18, 'answer': 'Philipp'}
除此之外,我们还为Optimum添加了一个pipelines API,以确保你的加速模型更加安全。这意味着如果你尝试将optimum.pipelines用于不支持的模型或任务,将会看到错误。你可以使用optimum.pipelines作为transformers.pipelines的替代。
from transformers import AutoTokenizer
from optimum.onnxruntime import ORTModelForQuestionAnswering
from optimum.pipelines import pipeline
tokenizer = AutoTokenizer.from_pretrained(onnx_path)
model = ORTModelForQuestionAnswering.from_pretrained(onnx_path)
optimum_qa = pipeline("question-answering", model=model, tokenizer=tokenizer, handle_impossible_answer=True)
prediction = optimum_qa(question="What's my name?", context="My name is Philipp and I live in Nuremberg.")
print(prediction)
# {'score': 0.9041663408279419, 'start': 11, 'end': 18, 'answer': 'Philipp'}
3.6 评估性能和速度
在这个关于加速 RoBERTa 问答任务(包括量化和优化)的端到端教程中,我们创建了 3 个不同的模型。一个原始转换模型、一个优化模型和一个量化模型。
作为教程的最后一步,我们希望详细查看我们模型的性能和准确率。应用优化技术,如图优化或量化,不仅会影响性能(延迟),还可能影响模型的准确率。因此,加速模型是有取舍的。
让我们评估我们的模型。我们的 transformers 模型 deepset/roberta-base-squad2 是在 SQUAD2 数据集上微调的。这将是我们用来评估模型的数据集。
from datasets import load_metric,load_dataset
metric = load_metric("squad_v2")
dataset = load_dataset("squad_v2")["validation"]
print(f"length of dataset {len(dataset)}")
#length of dataset 11873
我们现在可以利用 datasets 的 map 函数来遍历 squad 2 的验证集,并为每个数据点运行预测。因此,我们编写一个 evaluate 辅助方法,它使用我们的管道并应用一些转换以配合 squad v2 指标。
这可能需要相当长的时间(1.5 小时)
def evaluate(example):
default = optimum_qa(question=example["question"], context=example["context"])
optimized = opt_optimum_qa(question=example["question"], context=example["context"])
quantized = quantized_optimum_qa(question=example["question"], context=example["context"])
return {
'reference': {'id': example['id'], 'answers': example['answers']},
'default': {'id': example['id'],'prediction_text': default['answer'], 'no_answer_probability': 0.},
'optimized': {'id': example['id'],'prediction_text': optimized['answer'], 'no_answer_probability': 0.},
'quantized': {'id': example['id'],'prediction_text': quantized['answer'], 'no_answer_probability': 0.},
}
result = dataset.map(evaluate)
# COMMENT IN to run evaluation on 2000 subset of the dataset
# result = dataset.shuffle().select(range(2000)).map(evaluate)
现在让我们比较结果
default_acc = metric.compute(predictions=result["default"], references=result["reference"])
optimized = metric.compute(predictions=result["optimized"], references=result["reference"])
quantized = metric.compute(predictions=result["quantized"], references=result["reference"])
print(f"vanilla model: exact={default_acc['exact']}% f1={default_acc['f1']}%")
print(f"optimized model: exact={optimized['exact']}% f1={optimized['f1']}%")
print(f"quantized model: exact={quantized['exact']}% f1={quantized['f1']}%")
# vanilla model: exact=79.07858165585783% f1=82.14970024570314%
# optimized model: exact=79.07858165585783% f1=82.14970024570314%
# quantized model: exact=78.75010528088941% f1=81.82526107204629%
我们的优化和量化模型达到了精确匹配 78.75% 和 f1 分数 81.83%,这是原始准确率的 99.61%。达到原始模型的 99% 非常好,特别是因为我们使用了动态量化。
好的,让我们测试优化和量化模型的性能(延迟)。
但首先,让我们将上下文和问题扩展到更合理的序列长度 128。
context="Hello, my name is Philipp and I live in Nuremberg, Germany. Currently I am working as a Technical Lead at Hugging Face to democratize artificial intelligence through open source and open science. In the past I designed and implemented cloud-native machine learning architectures for fin-tech and insurance companies. I found my passion for cloud concepts and machine learning 5 years ago. Since then I never stopped learning. Currently, I am focusing myself in the area NLP and how to leverage models like BERT, Roberta, T5, ViT, and GPT2 to generate business value."
question="As what is Philipp working?"
为了简单起见,我们将使用一个 Python 循环,并计算原始模型以及优化和量化模型的平均/均值延迟。
from time import perf_counter
import numpy as np
def measure_latency(pipe):
latencies = []
# warm up
for _ in range(10):
_ = pipe(question=question, context=context)
# Timed run
for _ in range(100):
start_time = perf_counter()
_ = pipe(question=question, context=context)
latency = perf_counter() - start_time
latencies.append(latency)
# Compute run statistics
time_avg_ms = 1000 * np.mean(latencies)
time_std_ms = 1000 * np.std(latencies)
return f"Average latency (ms) - {time_avg_ms:.2f} +\- {time_std_ms:.2f}"
print(f"Vanilla model {measure_latency(optimum_qa)}")
print(f"Optimized & Quantized model {measure_latency(quantized_optimum_qa)}")
# Vanilla model Average latency (ms) - 117.61 +\- 8.48
# Optimized & Quantized model Average latency (ms) - 64.94 +\- 3.65
我们成功地将模型延迟从 117.61ms 加速到 64.94ms,大约快了 2 倍,同时保持了 99.61% 的准确率。我们应该记住的是,我们使用了一个具有 2 个物理核心的中等性能 CPU 实例。通过切换到 GPU 或更高性能的 CPU 实例,例如 ice-lake 驱动的,您可以将延迟数字降低到几毫秒。
4. 当前限制
我们刚刚开始在 https://github.com/huggingface/optimum 中支持推理,因此我们想分享当前的限制。所有这些限制都在路线图中,并将在不久的将来解决。
- 远程模型 > 2GB:目前,只能从 Hugging Face Hub 加载小于 2GB 的模型。我们正在努力添加对 > 2GB / 多文件模型的支持。
- Seq2Seq 任务/模型:我们不支持 seq2seq 任务,如摘要和 T5 等模型,主要是由于单模型支持的限制。但我们正在积极解决这个问题,为您提供您在 transformers 中熟悉的相同体验。
- 过去键值:像 GPT-2 这样的生成模型使用所谓的过去键值,这些是注意力块的预计算键值对,可用于加速解码。目前 ORTModelForCausalLM 没有使用过去键值。
- 无缓存:目前,加载优化模型(*.onnx)时,它不会在本地缓存。
5. Optimum 推理常见问题
支持哪些任务?
您可以在文档中找到所有支持任务的列表。目前支持的管道任务有 feature-extraction、text-classification、token-classification、question-answering、zero-shot-classification、text-generation
支持哪些模型?
任何可以使用 transformers.onnx 导出且具有受支持任务的模型都可以使用,这包括 BERT、ALBERT、GPT2、RoBERTa、XLM-RoBERTa、DistilBERT 等。
支持哪些运行时?
目前支持 ONNX Runtime。我们正在努力在未来添加更多支持。如果您对特定运行时感兴趣,请告诉我们。
如何将 Optimum 与 Transformers 一起使用?
您可以在我们的文档中找到示例和说明。
如何使用 GPU?
要使用 GPU,您只需安装 optimum[onnxruntine-gpu],它将安装所需的 GPU 提供程序并默认使用它们。
如何在 pipelines 中使用量化和优化后的模型?
您可以使用新的 ORTModelForXXX 类,通过 from_pretrained 方法加载优化或量化后的模型。您可以在我们的文档中了解更多信息。
6. 下一步是什么?
您问 Optimum 的下一步是什么?有很多事情。我们致力于让 Optimum 成为与 transformers 一起用于加速和优化的参考开源工具包。为了实现这一目标,我们将解决当前的限制,改进文档,创建更多内容和示例,并突破加速和优化 transformers 的极限。
Optimum 路线图上的一些重要功能,包括当前限制,有:
- 支持语音模型(Wav2vec2)和语音任务(自动语音识别)
- 支持视觉模型(ViT)和视觉任务(图像分类)
- 通过添加对 OrtValue 和 IOBinding 的支持来提高性能
- 更简便的方法来评估加速模型
- 添加对其他运行时和提供程序的支持,如 TensorRT 和 AWS-Neuron
感谢阅读!如果您和我一样对加速 Transformers、使其高效并扩展到数十亿请求感到兴奋。您应该申请,我们正在招聘。🚀
如果您有任何问题,请随时通过 Github 或论坛与我联系。您也可以在 Twitter 或 LinkedIn 上与我联系。
来源:Hugging Face:Blog(RSS) · huggingface.co