Build the cluster, feature store, and job scheduler that the modeling teams should never have to think about.
What you will do
Operate GPU capacity, CI for training jobs, and the APIs scientists use to launch and observe experiments.
What you bring
Strong systems engineering, Kubernetes, and enough ML literacy to debug a failed training run without a scientist in the room.