NCCL_ERROR_TIMEOUT pytorch network_error ai_generated partial

RuntimeError: [torch.distributed] 屏障超时 600000 毫秒后。NCCL 通信器在 rank 2 上被中止。原始失败原因:看门狗回调超时。

RuntimeError: [torch.distributed] Barrier timeout after 600000 ms. NCCL communicator was aborted on rank 2. Original reason for failure was: watchdog callback timed out.

ID: pytorch/nccl-timeout-barrier

其他格式: JSON · Markdown 中文 · English
75%修复率
86%置信度
1证据数
2024-01-15首次发现

版本兼容性

版本状态引入弃用备注
torch>=1.11 active
torch>=2.0 active
NCCL 2.12+ active
CUDA 11.x+ active
NVIDIA A100/H100 active

根因分析

分布式屏障超时,因为一个或多个 rank(例如 rank 2)因 NCCL 看门狗超时而停滞,通常由硬件故障、网络拥塞或 GPU 计算挂起引起。

English

A distributed barrier timed out because one or more ranks (e.g., rank 2) stalled due to NCCL watchdog timeout, often caused by hardware failures, network congestion, or GPU compute hung.

generic

官方文档

https://pytorch.org/docs/stable/distributed.html#troubleshooting

解决方案

  1. Set NCCL timeout in environment: export NCCL_TIMEOUT=600 (seconds); and ensure all GPUs are healthy with nvidia-smi -pm 1
  2. Reduce network load by using NCCL_IB_DISABLE=1 and NCCL_SOCKET_IFNAME=eth0 to force TCP instead of InfiniBand
  3. Run with torch.distributed.run --nproc_per_node=N with NCCL_DEBUG=INFO to identify the hanging rank and check its GPU memory or process status

无效尝试

常见但无效的做法:

  1. Increasing barrier timeout to 1200 seconds via torch.distributed.init_process_group(timeout=timedelta(seconds=1200)) 70% 失败

    A hung rank will eventually hit the watchdog timeout regardless of barrier timeout; the root cause is not duration but a hang.

  2. Restarting only the failed rank 2 without resetting NCCL communicator 90% 失败

    NCCL communicator state is global; restarting one rank leaves the group in an inconsistent state, causing further errors.