- README: add Queue/ClusterTrainingRuntime/TrainJob explanations with key options
- README: note single-node testing before multi-node
- Integration: set nvidia.com/gpu to 8 per node, Queue capacity to 24 (3 nodes)
- Integration: TrainJob example set to numNodes: 1 for install testing
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Volcano: set default_ns.nodegroup: nd for all components
- Kubeflow Trainer: set manager.nodeSelector.nodegroup: nd
- JobSet: set controller.nodeSelector.nodegroup: nd
- Fix PodGroup priorityClassName: normal-priority -> training-priority
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>