# RuntimeError: [torch.distributed] 屏障超时 600000 毫秒后。NCCL 通信器在 rank 2 上被中止。原始失败原因：看门狗回调超时。

- **ID:** `pytorch/nccl-timeout-barrier`
- **领域:** pytorch
- **类别:** network_error
- **错误码:** `NCCL_ERROR_TIMEOUT`
- **验证级别:** ai_generated
- **修复率:** 75%

## 根因

分布式屏障超时，因为一个或多个 rank（例如 rank 2）因 NCCL 看门狗超时而停滞，通常由硬件故障、网络拥塞或 GPU 计算挂起引起。

## 版本兼容性

| 版本 | 状态 | 引入 | 弃用 |
|------|------|------|------|
| torch>=1.11 | active | — | — |
| torch>=2.0 | active | — | — |
| NCCL 2.12+ | active | — | — |
| CUDA 11.x+ | active | — | — |
| NVIDIA A100/H100 | active | — | — |

## 解决方案

1. ```
   Set NCCL timeout in environment: export NCCL_TIMEOUT=600 (seconds); and ensure all GPUs are healthy with nvidia-smi -pm 1
   ```
2. ```
   Reduce network load by using NCCL_IB_DISABLE=1 and NCCL_SOCKET_IFNAME=eth0 to force TCP instead of InfiniBand
   ```
3. ```
   Run with torch.distributed.run --nproc_per_node=N with NCCL_DEBUG=INFO to identify the hanging rank and check its GPU memory or process status
   ```

## 无效尝试

- **Increasing barrier timeout to 1200 seconds via torch.distributed.init_process_group(timeout=timedelta(seconds=1200))** — A hung rank will eventually hit the watchdog timeout regardless of barrier timeout; the root cause is not duration but a hang. (70% 失败率)
- **Restarting only the failed rank 2 without resetting NCCL communicator** — NCCL communicator state is global; restarting one rank leaves the group in an inconsistent state, causing further errors. (90% 失败率)
