diff --git a/README.md b/README.md index a1dd28f..3214bf7 100644 --- a/README.md +++ b/README.md @@ -15,6 +15,35 @@ Kubeflow Trainer v2 (v1alpha1)에는 다음 런타임만 기본 제공된다: **TensorFlow runtime은 없다.** 따라서 Custom ClusterTrainingRuntime을 직접 생성해야 한다. +## 실행 결과 요약 + +2노드 × 8 GPU (Tesla V100-SXM2-32GB) 분산학습 성공: + +``` +환경: + - 노드: 2개 + - GPU: 노드당 8장 (총 16장) + - GPU 모델: Tesla V100-SXM2-32GB + - 통신: NCCL + InfiniBand + GPUDirect RDMA + - 이미지: nvcr.io/nvidia/tensorflow:24.03-tf2-py3 + +학습 결과: + - 데이터셋: CIFAR-10 (로컬 hostPath) + - 모델: CNN (Conv2D + BatchNorm + Dropout) + - Strategy: MultiWorkerMirroredStrategy (NCCL) + - num_replicas_in_sync: 16 + - Global Batch Size: 64 × 16 = 1024 + - Epochs: 10 + - Test Loss: 1.6169 + - Test Accuracy: 44.59% + +NCCL 초기화 로그: + NCCL INFO Connected all trees + NCCL INFO 16 coll channels, 0 nvls channels, ... + NCCL INFO comm 0x... rank 0 nranks 16 ... busId ... COMPLETE + → 16개 rank 전부 NCCL Init COMPLETE, GPUDirect RDMA 활성화 확인 +``` + ## PyTorch vs TensorFlow Runtime 핵심 차이 ``` @@ -78,6 +107,8 @@ mlPolicy에 framework(torch/mpi) 설정 여부에 따라 Pod 네이밍이 달라 └────────────────────────────────────────────────────────────────────┘ ``` +이 차이를 모르면 DNS 프로빙 패턴이 잘못되어 워커를 1개만 발견하는 문제가 발생한다. + ## TF_CONFIG 자동 구성 원리 ### 1. Pod hostname에서 정보 추출 @@ -89,6 +120,14 @@ hostname: tf-distributed-training-node-0-1 └─── replica_index (항상 0) ``` +hostname 파싱 정규식: +```python +match = re.match(r'^(.+)-node-(\d+)-(\d+)$', hostname) +job_name = match.group(1) # tf-distributed-training +replica_index = match.group(2) # 0 (항상) +worker_index = int(match.group(3)) # completion_index = worker_index +``` + ### 2. Headless Service DNS로 워커 주소 구성 ``` @@ -134,6 +173,108 @@ import tensorflow as tf # 2. 그 다음 tensorflow import # "different incarnation" 에러가 발생한다. ``` +## TFJob (Trainer v1) vs TrainJob (Trainer v2) 비교 + +### chief/worker 패턴 비교 + +Trainer v1의 TFJob은 `chief`와 `worker` role을 명시적으로 분리했다: + +```yaml +# TFJob (Trainer v1) - chief/worker 명시적 분리 +apiVersion: kubeflow.org/v1 +kind: TFJob +spec: + tfReplicaSpecs: + Chief: + replicas: 1 + template: ... # chief 전용 설정 + Worker: + replicas: 3 + template: ... # worker 설정 + PS: + replicas: 2 + template: ... # Parameter Server (선택) +``` + +TrainJob (Trainer v2)은 단일 role 타입(`node`) 기반이다: + +```yaml +# TrainJob (Trainer v2) - 단일 role +apiVersion: trainer.kubeflow.org/v1alpha1 +kind: TrainJob +spec: + trainer: + numNodes: 4 # 모든 노드가 동일한 role +``` + +### 구현 가능성 분석 + +| 기능 | TFJob (v1) | TrainJob (v2) | 비고 | +|------|-----------|---------------|------| +| chief/worker 분리 | `tfReplicaSpecs.Chief/Worker` | 미지원 (단일 role) | TrainJob은 replicatedJobs로 부분 구현 가능 | +| Parameter Server | `tfReplicaSpecs.PS` | 미지원 | PS 전략 자체가 레거시 | +| Evaluator | `tfReplicaSpecs.Evaluator` | 미지원 | - | +| worker index 기반 분기 | role별로 자동 | `worker_index == 0`으로 직접 구현 | 실질적으로 동일한 효과 | +| 분산 전략 | PS Strategy / MirroredStrategy | MultiWorkerMirroredStrategy | v2는 AllReduce 기반 | + +### worker_index == 0 패턴 (현재 구현) + +TFJob의 chief가 하던 역할(체크포인트 저장, 로그 출력, 평가 결과 출력 등)을 +`worker_index == 0`으로 분기하여 동일하게 구현: + +```python +# worker_index == 0 이 사실상 chief 역할 +if worker_index == 0: + print(f"Test Loss: {loss:.4f}") + print(f"Test Accuracy: {accuracy:.4f}") + # 체크포인트 저장, TensorBoard 로그 등도 여기서 +``` + +**주의**: `model.evaluate()`는 모든 워커가 함께 호출해야 한다 (collective operation). +Worker 0만 호출하면 다른 워커가 heartbeat timeout으로 실패한다. + +```python +# 올바른 패턴: 모든 워커가 evaluate 참여, 출력만 worker 0 +loss, accuracy = model.evaluate(test_dataset, verbose=eval_verbose) # 모든 워커 +if worker_index == 0: # 출력만 worker 0 + print(f"Test Accuracy: {accuracy:.4f}") +``` + +### replicatedJobs를 이용한 multi-role 구현 (참고) + +TrainJob의 Runtime에서 replicatedJobs를 여러 개 정의하면 role 분리가 이론적으로 가능하다: + +```yaml +# 이론적 구현 (권장하지 않음) +replicatedJobs: +- name: chief # chief role + template: + spec: + template: + spec: + containers: + - name: node + env: + - name: TF_ROLE + value: "chief" +- name: worker # worker role + template: + spec: + template: + spec: + containers: + - name: node + env: + - name: TF_ROLE + value: "worker" +``` + +하지만 TrainJob API의 `trainer` 필드는 단일 role을 가정하므로, 여러 replicatedJobs에 대한 +`numNodes`, `resourcesPerNode` 등의 매핑이 명확하지 않다. + +**결론**: 현재 TrainJob에서는 `worker_index == 0` 패턴으로 chief 역할을 대체하는 것이 가장 실용적이다. +TFJob의 PS(Parameter Server) 전략 자체가 레거시이므로, MultiWorkerMirroredStrategy 기반 AllReduce가 현대적 표준이다. + ## CIFAR-10 데이터 로딩 `tf.keras.datasets.cifar10.load_data()`는 외부 URL에서 다운로드를 시도한다. @@ -154,13 +295,14 @@ volumes: ## 통신 구조 (InfiniBand + NCCL) TensorFlow도 PyTorch와 동일하게 NCCL + InfiniBand를 사용할 수 있다. +NCCL은 NVIDIA에서 제공하는 프레임워크 독립적인 GPU 간 통신 라이브러리이므로, TF/PyTorch 구분 없이 동일하게 동작한다. ``` -PyTorch: torchrun → torch.distributed → NCCL → InfiniBand -TensorFlow: gRPC (조정) → MultiWorkerMirroredStrategy → NCCL → InfiniBand +PyTorch: torchrun → torch.distributed → NCCL → InfiniBand/GPUDirect RDMA +TensorFlow: gRPC (조정) → MultiWorkerMirroredStrategy → NCCL → InfiniBand/GPUDirect RDMA ``` -NCCL은 프레임워크 독립적인 라이브러리이므로, 동일한 환경변수가 양쪽 모두에 적용된다: +동일한 NCCL 환경변수가 양쪽 모두에 적용된다: ```yaml env: @@ -170,6 +312,16 @@ env: value: "net" # OOB 통신용 인터페이스 - name: NCCL_DEBUG value: "INFO" # NCCL 디버그 로그 + - name: NCCL_DEBUG_SUBSYS + value: "INIT,NET,IB" # NCCL 서브시스템별 디버그 +``` + +실제 학습 시 확인된 NCCL 초기화 로그: +``` +NCCL INFO NET/IB : Using [0]mlx5_0:1/RoCE ... ; OOB net0: +NCCL INFO Connected all trees +NCCL INFO 16 coll channels, 0 nvls channels, ... +NCCL INFO comm 0x... rank 0 nranks 16 ... COMPLETE ``` ## 파일 구조 @@ -310,6 +462,23 @@ spec: value: "WARN" ``` +### model.evaluate()는 모든 워커가 참여해야 함 + +MultiWorkerMirroredStrategy는 collective operation 기반이다. +`model.evaluate()`를 worker 0만 호출하면 다른 워커가 heartbeat timeout으로 실패한다. + +```python +# 잘못된 패턴 (worker 0만 evaluate) +if worker_index == 0: + loss, accuracy = model.evaluate(test_dataset) # ← 다른 워커 timeout 발생 + +# 올바른 패턴 (모든 워커가 evaluate, 출력만 worker 0) +eval_verbose = 2 if worker_index == 0 else 0 +loss, accuracy = model.evaluate(test_dataset, verbose=eval_verbose) # 모든 워커 참여 +if worker_index == 0: + print(f"Test Accuracy: {accuracy:.4f}") # 출력만 worker 0 +``` + ## 트러블슈팅 ### "different incarnation" 에러 @@ -322,15 +491,37 @@ unexpectedly tried to connect with a different incarnation. It has likely restar **원인**: `TF_CONFIG`가 `import tensorflow` 이후에 설정됨 **해결**: 학습 스크립트에서 `setup_tf_config()`를 `import tensorflow` 전에 호출 -### DNS 프로빙에서 워커를 1개만 발견 +### Pod 네이밍 패턴 불일치로 워커 감지 실패 ``` [Discovery] attempt 1/60: found 1 workers, retrying in 5s... +(영원히 1개만 발견) +``` + +**원인**: DNS 프로빙 패턴이 `node-{i}-0` (PyTorch 패턴)으로 되어있으나, 실제는 `node-0-{i}` (completion 패턴) +**해결**: mlPolicy에 framework가 없으면 `numNodes`가 단일 Job의 `completions`로 매핑됨을 이해하고, DNS 패턴을 `{trainjob}-node-0-{i}`로 수정 + +### DNS 프로빙에서 워커를 1개만 발견 (정상 대기) + +``` +[Discovery] attempt 1/60: found 1 workers, retrying in 5s... +[Discovery] attempt 2/60: found 2 workers, retrying in 5s... +[Discovery] Discovered 2 workers via DNS ``` **원인**: 다른 워커 Pod이 아직 시작되지 않았거나 DNS 전파 지연 **해결**: Volcano gang scheduling 사용 시 모든 Pod이 동시에 스케줄링되지만, DNS 전파에 시간이 걸릴 수 있다. `discover_workers()`가 연속 2회 동일 결과를 확인한 후 진행하므로 정상적으로 기다린다. +### model.evaluate() heartbeat timeout + +``` +W tensorflow/core/distributed_runtime/...coordination_service_agent.cc: +Heartbeat timeout from ... +``` + +**원인**: Worker 0만 `model.evaluate()` 호출, 다른 워커는 이미 종료 +**해결**: MultiWorkerMirroredStrategy는 collective operation이므로, 모든 워커가 `model.evaluate()`에 참여해야 함. 출력만 `worker_index == 0`에서 처리. + ### CIFAR-10 데이터 다운로드 시도 ```