kubeflow_trainer_with_volcano/configs
selee 7cd7c1af7e Fix torchrun injection: add ancestor-step label, use torchrun command
- Add trainer.kubeflow.org/trainjob-ancestor-step label to replicatedJob
  template so Trainer controller injects PET_* env vars
- Change TrainJob command from python to torchrun so PET_* env vars
  are used for distributed training setup
- Remove command from Runtime (keep it in TrainJob only)

Verified: 2-node 16-GPU distributed training running successfully

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-10 16:36:44 +09:00
..
kubeflow-trainer-values.yaml Add nodeSelector(nodegroup: nd) to all components, fix PriorityClass mismatch 2026-02-09 15:34:19 +09:00
volcano-trainjob-integration.yaml Fix torchrun injection: add ancestor-step label, use torchrun command 2026-02-10 16:36:44 +09:00
volcano-values.yaml Add nodeSelector(nodegroup: nd) to all components, fix PriorityClass mismatch 2026-02-09 15:34:19 +09:00