Senior Engineering Manager, AI Infrastructure

ApplyApply
Posted about 1 hour ago
Share
Full Time
Seattle, WA
$146,880 - $220,320 Annually

Senior Engineering Manager, AI Infrastructure

Seattle, WA - Hybrid (3 days/week onsite)
Salary: $146,880 - $220,320 + 15% performance bonus + long-term incentive

About the Role

We are seeking a Senior Engineering Manager, AI Infrastructure to lead the day-to-day operation of the systems that power large-scale AI research and development. Reporting to senior engineering leadership, you will own the execution, reliability, and performance of a high-performance computing environment consisting of on-premise GPU clusters and a software orchestration layer that schedules workloads across a hybrid cloud environment.

This is a hands-on operational leadership role focused on ensuring the platform remains reliable, performant, and highly utilized while delivering against infrastructure roadmaps and business priorities.

Who You Are

Systems Expert

You possess deep expertise in Linux, containerized environments, distributed systems, and large-scale compute infrastructure. You understand the performance characteristics of GPU clusters, high-performance networking, and large-scale AI workloads.

Execution-Focused Leader

You excel at driving operational excellence, delivering against priorities, and translating strategic goals into reliable systems and measurable outcomes.

Pragmatic Operator

You understand how to balance technical idealism with operational realities. You can quickly assess risks, prioritize effectively, and make sound decisions in fast-moving environments.

Why This Opportunity?

  • Work on infrastructure that supports cutting-edge AI and machine learning research.
  • Lead highly complex compute environments at significant scale.
  • Partner directly with researchers, engineers, and infrastructure teams to enable breakthrough technical work.
  • Join a mission-driven organization focused on advancing technology through collaboration, innovation, and long-term thinking.

Key Responsibilities

Cluster Operations

Manage the availability, performance, and health of dense on-prem GPU clusters. Coordinate with hardware vendors and internal stakeholders to ensure infrastructure meets the demands of large-scale AI training workloads.

Orchestration & Scheduling

Operate and improve workload orchestration systems by optimizing resource allocation and driving efficient utilization across on-prem infrastructure and public cloud environments (AWS/GCP).

Storage Operations

Manage and continuously improve distributed storage environments, balancing performance requirements for active training workloads with scalable, durable storage for large research datasets.

Resource Management

Track GPU utilization, capacity planning, and infrastructure spend. Provide recommendations regarding cloud expansion, on-prem investment, and resource allocation strategies.

Research Enablement

Serve as a technical partner to research and engineering teams, ensuring infrastructure accelerates productivity rather than creating bottlenecks.

Team Leadership

Lead and develop a team of systems engineers, SREs, and software engineers. Establish high standards for operational excellence, technical quality, collaboration, and execution.

Qualifications

Required

  • 12+ years of experience in infrastructure engineering, systems engineering, cloud platforms, or HPC environments (or an advanced degree with 8+ years of relevant experience).
  • 2+ years managing engineering teams of 5 or more individuals.
  • Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent practical experience).
  • Experience operating large-scale NVIDIA GPU infrastructure and high-performance networking technologies such as InfiniBand or RoCE.
  • Strong experience with Kubernetes, Slurm, or comparable orchestration and scheduling platforms.
  • Hands-on experience managing distributed storage systems such as WEKA, Ceph, Lustre, or similar technologies.
  • Experience leading software development processes including sprint planning, design reviews, and engineering execution.
  • Strong proficiency with Go and/or Python.

Preferred

  • Experience supporting AI/ML training environments at scale.
  • Hybrid cloud infrastructure experience across AWS, GCP, or similar platforms.
  • Deep understanding of workload scheduling and resource optimization.
  • Experience scaling engineering organizations in research-intensive environments.
  • Familiarity with modern observability, reliability, and infrastructure automation practices.

Culture & Environment

This organization values:

  • Continuous learning and professional development
  • Technical excellence and operational rigor
  • Diversity, equity, inclusion, and belonging
  • Healthy work-life balance
  • Transparency and collaboration
  • High ownership and accountability

You'll have the opportunity to work alongside world-class engineers and researchers while helping build infrastructure that supports some of the most demanding computational workloads in the industry.

Benefits

  • Medical, dental, and vision coverage for employees and dependents
  • Employee Assistance Program (EAP)
  • Health Savings Account (HSA), Healthcare Reimbursement Arrangement (HRA), and Flexible Spending Accounts (FSA)
  • 401(k) retirement plan
  • $125/month commuting or internet stipend
  • $200/month fitness and wellness stipend
  • Generous vacation, sick leave, personal days, and paid holidays
  • Annual performance bonus
  • Long-term incentive program

Interview Process

  • Hiring Manager Screen (30 minutes)
  • Technical Leadership Interview (45 minutes)
  • Take-Home Assignment
  • Final Onsite Interview

Onsite interviews include:

  • Assignment presentation and Q&A
  • Research stakeholder meeting
  • Technical infrastructure interview
  • Hiring manager discussion
  • Meetings with program and engineering leaders
  • Team lunch and informal meet-and-greet

Apply