跳到正文
原文
SemiAnalysis· @SemiAnalysis_ · X·· 2 天前AI 评分35
AI 导读

SemiAnalysis 对 GPU 集群健康检查做基础测试,发现多个集群的健康检查在逻辑上根本不可能通过。在 Amazon HyperPod Slurm 上,一项健康检查要求节点先健康才能运行,但该检查本身又是节点重新入集群的必要条件,形成死循环。SemiAnalysis 称测试中约有五到十个此类案例,这些健康检查比没有更糟,因为会主动干扰任务。

正文

The most basic test SemiAnalysis applies to a GPU cluster health check is whether it could possibly work on paper. Several could not.

"We've been through a few of these clusters where there's no way it could possibly work. These are health checks so bad they're worse than no health checks, because they're actively interfering with jobs."

"On Amazon HyperPod Slurm, the first time we tested, there was a health check that needed the node to be healthy before it could run. It could only run once the node was back in the fleet, and it was needed to bring the node back into the fleet."

"There are probably five or ten examples of this throughout testing where the health check never had a chance. It makes you wonder whether these people are really putting it through the paces before they hand it off to customers."

来源:SemiAnalysis · x.com