跳到正文
原文
Google Developers Blog(RSS)·· 13 小时前AI 评分37

Google Cloud 在 Cloud TPU 上为 vLLM 实现长上下文多模态嵌入推理

Enterprise-Grade Precision for Long-Context Multimodal Embedding Inference on Cloud TPU

AI 导读

Google Cloud 将原生 TPU 支持集成进 vLLM,并以 Qwen3-Embedding-8B 为目标模型完成多项优化,支持 4K+ token 文本与 15K+ token 多模态输入的嵌入推理。

来源:Google Developers Blog(RSS) · developers.googleblog.com