Senior Engineering Manager, AI Infrastructure
Senior Engineering Manager, AI Infrastructure
Seattle, WA - Hybrid (3 days/week onsite)
Salary: $146,880 - $220,320 + 15% performance bonus + long-term incentive
About the Role
We are seeking a Senior Engineering Manager, AI Infrastructure to lead the day-to-day operation of the systems that power large-scale AI research and development. Reporting to senior engineering leadership, you will own the execution, reliability, and performance of a high-performance computing environment consisting of on-premise GPU clusters and a software orchestration layer that schedules workloads across a hybrid cloud environment.
This is a hands-on operational leadership role focused on ensuring the platform remains reliable, performant, and highly utilized while delivering against infrastructure roadmaps and business priorities.
Who You Are
Systems Expert
You possess deep expertise in Linux, containerized environments, distributed systems, and large-scale compute infrastructure. You understand the performance characteristics of GPU clusters, high-performance networking, and large-scale AI workloads.
Execution-Focused Leader
You excel at driving operational excellence, delivering against priorities, and translating strategic goals into reliable systems and measurable outcomes.
Pragmatic Operator
You understand how to balance technical idealism with operational realities. You can quickly assess risks, prioritize effectively, and make sound decisions in fast-moving environments.
Why This Opportunity?
- Work on infrastructure that supports cutting-edge AI and machine learning research.
- Lead highly complex compute environments at significant scale.
- Partner directly with researchers, engineers, and infrastructure teams to enable breakthrough technical work.
- Join a mission-driven organization focused on advancing technology through collaboration, innovation, and long-term thinking.
Key Responsibilities
Cluster Operations
Manage the availability, performance, and health of dense on-prem GPU clusters. Coordinate with hardware vendors and internal stakeholders to ensure infrastructure meets the demands of large-scale AI training workloads.
Orchestration & Scheduling
Operate and improve workload orchestration systems by optimizing resource allocation and driving efficient utilization across on-prem infrastructure and public cloud environments (AWS/GCP).
Storage Operations
Manage and continuously improve distributed storage environments, balancing performance requirements for active training workloads with scalable, durable storage for large research datasets.
Resource Management
Track GPU utilization, capacity planning, and infrastructure spend. Provide recommendations regarding cloud expansion, on-prem investment, and resource allocation strategies.
Research Enablement
Serve as a technical partner to research and engineering teams, ensuring infrastructure accelerates productivity rather than creating bottlenecks.
Team Leadership
Lead and develop a team of systems engineers, SREs, and software engineers. Establish high standards for operational excellence, technical quality, collaboration, and execution.
Qualifications
Required
- 12+ years of experience in infrastructure engineering, systems engineering, cloud platforms, or HPC environments (or an advanced degree with 8+ years of relevant experience).
- 2+ years managing engineering teams of 5 or more individuals.
- Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent practical experience).
- Experience operating large-scale NVIDIA GPU infrastructure and high-performance networking technologies such as InfiniBand or RoCE.
- Strong experience with Kubernetes, Slurm, or comparable orchestration and scheduling platforms.
- Hands-on experience managing distributed storage systems such as WEKA, Ceph, Lustre, or similar technologies.
- Experience leading software development processes including sprint planning, design reviews, and engineering execution.
- Strong proficiency with Go and/or Python.
Preferred
- Experience supporting AI/ML training environments at scale.
- Hybrid cloud infrastructure experience across AWS, GCP, or similar platforms.
- Deep understanding of workload scheduling and resource optimization.
- Experience scaling engineering organizations in research-intensive environments.
- Familiarity with modern observability, reliability, and infrastructure automation practices.
Culture & Environment
This organization values:
- Continuous learning and professional development
- Technical excellence and operational rigor
- Diversity, equity, inclusion, and belonging
- Healthy work-life balance
- Transparency and collaboration
- High ownership and accountability
You'll have the opportunity to work alongside world-class engineers and researchers while helping build infrastructure that supports some of the most demanding computational workloads in the industry.
Benefits
- Medical, dental, and vision coverage for employees and dependents
- Employee Assistance Program (EAP)
- Health Savings Account (HSA), Healthcare Reimbursement Arrangement (HRA), and Flexible Spending Accounts (FSA)
- 401(k) retirement plan
- $125/month commuting or internet stipend
- $200/month fitness and wellness stipend
- Generous vacation, sick leave, personal days, and paid holidays
- Annual performance bonus
- Long-term incentive program
Interview Process
- Hiring Manager Screen (30 minutes)
- Technical Leadership Interview (45 minutes)
- Take-Home Assignment
- Final Onsite Interview
Onsite interviews include:
- Assignment presentation and Q&A
- Research stakeholder meeting
- Technical infrastructure interview
- Hiring manager discussion
- Meetings with program and engineering leaders
- Team lunch and informal meet-and-greet

