跳到正文
原文
AK· @_akhaliq · X·· 1 天前AI 评分42
AI 导读

Audio8 ASR Infinite 流式语音识别模型发布,采用原生流式架构,每秒解码 12.5 次,通过滚动 KV Cache 在 7×24 小时运行下保持内存和延迟恒定。模型每时钟步输出一个文本 token(12.5/8.3/6.25 次决策每秒),平衡感知粒度与资源开销。HuggingChat 的 ML 实习生已搭建 Gradio 工作流供试用。

正文

Audio8 ASR Infinite

streaming speech recognition model

the native streaming architecture decodes 12.5 times per second

a rolling KV Cache keeps both memory and latency constant, even in 24/7 operation

one text token per clock step (12.5 / 8.3 / 6.25 decisions per second), balancing perception granularity and resource cost

ML intern in huggingchat setup a gradio workflow to try it out: https://huggingface.co/spaces/akhaliq/audio8-asr-workflow

huggingchat: https://huggingface.co/chat/

model: https://huggingface.co/Edge0/Audio8-ASR-Infinite

来源:AK · x.com