- Add trainer.kubeflow.org/trainjob-ancestor-step label to replicatedJob template so Trainer controller injects PET_* env vars - Change TrainJob command from python to torchrun so PET_* env vars are used for distributed training setup - Remove command from Runtime (keep it in TrainJob only) Verified: 2-node 16-GPU distributed training running successfully Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| kubeflow-trainer-values.yaml | ||
| volcano-trainjob-integration.yaml | ||
| volcano-values.yaml | ||