Senior Software Engineer, AI Infrastructure

ApplyApply
Posted about 2 hours ago
Share
Full Time
Seattle, WA
$126,000 - $189,000 Annually

Senior Software Engineer, AI Infrastructure

Seattle, WA - Hybrid (3 days/week onsite)
Salary: $126,000-$189,000 + performance bonus + long-term incentive

Who You Are

You are an expert systems engineer who thrives at the intersection of large-scale software orchestration and low-level systems performance. You're passionate about building reliable infrastructure that enables cutting-edge AI research and high-performance computing at scale.

You're equally comfortable designing resource allocation systems in Go or Python as you are debugging distributed GPU training issues. You lead by example, combining deep technical expertise with a hands-on approach to operating and improving large-scale compute environments.

About the Opportunity

This organization is dedicated to advancing open AI research and building high-performance infrastructure that supports large-scale machine learning. Rather than focusing on proprietary platforms, the team is committed to developing scalable, transparent systems that empower researchers and engineers to push the boundaries of AI.

You'll join a collaborative engineering organization that combines the speed of a startup with the technical depth of a research environment, building infrastructure that supports some of the world's most demanding AI workloads.

What You'll Be Working On

You'll serve as a senior technical contributor responsible for building and operating the infrastructure that powers large-scale AI training.

Your work will include:

  • Designing for Scale: Build and improve orchestration systems that efficiently allocate GPU resources across large compute clusters.
  • Operational Excellence: Automate infrastructure management and reduce manual operational overhead.
  • Performance Engineering: Partner with researchers and engineers to optimize distributed training performance and maximize hardware utilization.

Responsibilities

  • Design and deliver critical infrastructure spanning workload scheduling, orchestration, and execution systems.
  • Build automation and software-defined infrastructure that improves researcher productivity and cluster reliability.
  • Investigate and resolve complex distributed systems issues while optimizing performance for large-scale workloads.
  • Contribute to the roadmap for compute, networking, and storage infrastructure alongside engineering leadership.
  • Review code and design documents, mentor teammates, and help elevate engineering best practices.
  • Collaborate closely with research and engineering teams to gather feedback, communicate technical designs, and support implementation efforts.

Qualifications

  • 8+ years of experience building business-critical software and operating large-scale compute infrastructure.
  • Strong experience with Go and/or Python.
  • Bachelor's degree in Computer Science or a related field (or equivalent practical experience).
  • Expert knowledge of Linux systems and container technologies such as Docker.
  • Proven experience designing, debugging, and optimizing distributed systems.
  • Excellent written communication skills with the ability to collaborate across engineering and research organizations.
  • A thoughtful approach to software engineering with an interest in solving complex infrastructure challenges.

Preferred Qualifications

  • Experience with workload schedulers such as Kubernetes or Slurm.
  • Experience with GPU infrastructure, distributed training, NCCL, or InfiniBand.
  • Experience training or supporting large-scale AI or machine learning workloads.
  • Site Reliability Engineering (SRE) or HPC infrastructure experience.
  • Contributions to open-source infrastructure or orchestration projects.
  • Familiarity with distributed storage technologies such as WEKA or Ceph.

Apply