AK· @_akhaliq · X·· 1 天前AI 评分42
AI 导读
Audio8 ASR Infinite 流式语音识别模型发布,采用原生流式架构,每秒解码 12.5 次,通过滚动 KV Cache 在 7×24 小时运行下保持内存和延迟恒定。模型每时钟步输出一个文本 token(12.5/8.3/6.25 次决策每秒),平衡感知粒度与资源开销。HuggingChat 的 ML 实习生已搭建 Gradio 工作流供试用。
正文
Audio8 ASR Infinite
streaming speech recognition model
the native streaming architecture decodes 12.5 times per second
a rolling KV Cache keeps both memory and latency constant, even in 24/7 operation
one text token per clock step (12.5 / 8.3 / 6.25 decisions per second), balancing perception granularity and resource cost
ML intern in huggingchat setup a gradio workflow to try it out: https://huggingface.co/spaces/akhaliq/audio8-asr-workflow
huggingchat: https://huggingface.co/chat/
model: https://huggingface.co/Edge0/Audio8-ASR-Infinite
来源:AK · x.com