Senior Software Engineer, AI Infrastructure

Company Description

Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.  https://www.mirantis.com/

Job Description

What you'll do

  • Design and build LLM serving infrastructure on Kubernetes: deployment, GPU scheduling, scaling, and model lifecycle management.

  • Package the platform for enterprise environments: Helm-based installs, upgrades, and restricted/offline networks.

  • Integrate the serving layer with the platform's API gateway, identity, and metering services.

  • Build the observability for operating GPU inference in production (serving metrics, GPU telemetry).

  • Contribute across a multi-service codebase and help set engineering direction through design docs and reviews.

Qualifications

What we're looking for

  • 5+ years of software engineering experience in infrastructure, platform, or distributed systems.

  • Deep hands-on Kubernetes experience: building and operating production workloads and Helm charts, not just consuming managed clusters.

  • Experience with GPU workloads or LLM inference, or strong adjacent systems experience and a track record of learning fast.

  • Strong Go programming skills; solid CI/CD and infrastructure-as-code skills.

  • Fluency with AI-assisted development tools (Claude Code, OpenAI Codex) as part of your daily engineering workflow.

  • Comfortable with high autonomy on a small, remote-first, written-culture team.

Nice to have

  • Inference performance work (quantization, batching, caching) or distributed serving frameworks.

  • Enterprise deployment experience: air-gapped installs, SSO/OIDC, supply-chain security.

  • UI development experience (e.g. React/TypeScript), useful as the product's management surfaces grow.

  • Open-source contributions in the Kubernetes or ML-infrastructure ecosystems.

Additional Information

What does Mirantis offer you?

  • Work with an established Silicon Valley leader in the cloud infrastructure industry;

  • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;

  • Be a part of cutting-edge, open-source innovation;

  • Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;

  • Professional development and training;

  • Attend conferences and working groups;

  • Company outings, happy hours, hackathons, and tech talks;

  • Receive a competitive compensation package with a strong benefits plan.

We are a Leader for Container Management in G2 (#2 after AWS)!

Summary

Design and build LLM serving infrastructure on Kubernetes covering GPU scheduling, scaling, model lifecycle, Helm packaging, and production observability. Requires 5+ years in infrastructure, platform, or distributed systems with deep Kubernetes, Helm, Go, and CI/CD experience plus GPU or LLM inference background.

Responsibilities

Design and build LLM serving infrastructure on Kubernetes: deployment, GPU scheduling, scaling, and model lifecycle management; Package platform for enterprise environments: Helm-based installs, upgrades, and restricted/offline networks; Integrate serving layer with API gateway, identity, and metering services; Build observability for GPU inference; Contribute across multi-service codebase

Qualifications

Deep hands-on Kubernetes experience building and operating production workloads and Helm charts; Experience with GPU workloads or LLM inference; Strong Go programming skills; solid CI/CD and infrastructure-as-code skills; Fluency with AI-assisted development tools

Experience requirements

5+ years of software engineering experience in infrastructure, platform, or distributed systems