Fix torchrun injection: add ancestor-step label, use torchrun command

- Add trainer.kubeflow.org/trainjob-ancestor-step label to replicatedJob
  template so Trainer controller injects PET_* env vars
- Change TrainJob command from python to torchrun so PET_* env vars
  are used for distributed training setup
- Remove command from Runtime (keep it in TrainJob only)

Verified: 2-node 16-GPU distributed training running successfully

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
selee 2026-02-10 16:36:44 +09:00
parent 2b1ad10c2c
commit 7cd7c1af7e
1 changed files with 4 additions and 1 deletions

View File

@ -310,6 +310,9 @@ spec:
replicatedJobs:
- name: node
template:
metadata:
labels:
trainer.kubeflow.org/trainjob-ancestor-step: trainer
spec:
template:
metadata:
@ -376,7 +379,7 @@ spec:
trainer:
image: nvcr.io/nvidia/pytorch:24.10-py3
command:
- python
- torchrun
- /workspace/scripts/train_nccl.py
numNodes: 2
resourcesPerNode: