NCCL_ERROR_TIMEOUT
pytorch
network_error
ai_generated
partial
RuntimeError: [torch.distributed] 屏障超时 600000 毫秒后。NCCL 通信器在 rank 2 上被中止。原始失败原因:看门狗回调超时。
RuntimeError: [torch.distributed] Barrier timeout after 600000 ms. NCCL communicator was aborted on rank 2. Original reason for failure was: watchdog callback timed out.
ID: pytorch/nccl-timeout-barrier
75%修复率
86%置信度
1证据数
2024-01-15首次发现
版本兼容性
| 版本 | 状态 | 引入 | 弃用 | 备注 |
|---|---|---|---|---|
| torch>=1.11 | active | — | — | — |
| torch>=2.0 | active | — | — | — |
| NCCL 2.12+ | active | — | — | — |
| CUDA 11.x+ | active | — | — | — |
| NVIDIA A100/H100 | active | — | — | — |
根因分析
分布式屏障超时,因为一个或多个 rank(例如 rank 2)因 NCCL 看门狗超时而停滞,通常由硬件故障、网络拥塞或 GPU 计算挂起引起。
English
A distributed barrier timed out because one or more ranks (e.g., rank 2) stalled due to NCCL watchdog timeout, often caused by hardware failures, network congestion, or GPU compute hung.
官方文档
https://pytorch.org/docs/stable/distributed.html#troubleshooting解决方案
-
Set NCCL timeout in environment: export NCCL_TIMEOUT=600 (seconds); and ensure all GPUs are healthy with nvidia-smi -pm 1
-
Reduce network load by using NCCL_IB_DISABLE=1 and NCCL_SOCKET_IFNAME=eth0 to force TCP instead of InfiniBand
-
Run with torch.distributed.run --nproc_per_node=N with NCCL_DEBUG=INFO to identify the hanging rank and check its GPU memory or process status
无效尝试
常见但无效的做法:
-
Increasing barrier timeout to 1200 seconds via torch.distributed.init_process_group(timeout=timedelta(seconds=1200))
70% 失败
A hung rank will eventually hit the watchdog timeout regardless of barrier timeout; the root cause is not duration but a hang.
-
Restarting only the failed rank 2 without resetting NCCL communicator
90% 失败
NCCL communicator state is global; restarting one rank leaves the group in an inconsistent state, causing further errors.