Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure
About the Company
AI is transforming digital industries, but progress in the physical sciences — including drug discovery, materials science, and climate modeling — remains constrained by access to high-quality data and infrastructure.
This company is building an infrastructure platform that enables organizations to securely share proprietary datasets, train models on high-performance GPU infrastructure, and deploy those models for inference and application development. The platform gives data owners control over their datasets while providing researchers and engineers with the compute and tooling needed to train, deploy, and use domain-specific AI models.
The team is building critical infrastructure at the intersection of AI, cloud computing, GPUs, and scientific data, with the opportunity to shape the platform from an early stage.
About the Role
You'll be an early engineer on the infrastructure team, owning the compute layer that powers GPU workloads and model inference.
You'll make architectural decisions across the platform, work hands-on with Kubernetes and cloud infrastructure, and partner directly with customers to understand their technical challenges and build reliable systems that operate at scale.
This is a highly hands-on role for someone who enjoys solving complex infrastructure problems and wants significant ownership over technical direction.
What You Will Do
- Build the GPU compute layer — Develop and operate infrastructure for GPU workloads on Kubernetes, including resource allocation, scheduling, multi-tenancy, and cost management.
- Build the inference layer — Design systems for model loading, autoscaling, batching, and serving, with a focus on latency, throughput, and reliability.
- Own ML infrastructure end to end — Build and improve data ingestion, preprocessing, training and fine-tuning workflows, and recovery mechanisms for distributed workloads.
- Work directly with customers — Troubleshoot training and fine-tuning jobs, identify performance bottlenecks, and build observability into the platform to monitor model and infrastructure health.
- Own reliability — Participate in incident response and on-call while building systems that remain reliable as usage and scale increase.
- Shape technical direction — Drive build-vs-buy decisions across infrastructure and security, establish engineering standards, and help shape the team's technical roadmap.
Qualifications
- 5+ years of experience building and operating production infrastructure, ideally supporting ML workloads such as training, inference, or data pipelines
- Hands-on experience with Kubernetes and a major cloud platform such as AWS or GCP
- Experience working with GPU infrastructure or ML/AI workloads
- Strong computer science fundamentals and system design skills
- Experience operating production systems where reliability, performance, and scalability matter
- Comfortable working in ambiguous, early-stage environments where the playbook doesn't exist yet
- Strong ownership mentality and willingness to work across architecture, implementation, operations, and customer-facing technical problem solving
Why This Role
- Join an early-stage team building infrastructure at the intersection of AI, GPUs, and scientific computing
- Own meaningful technical systems rather than a narrow slice of a large platform
- Work directly with customers and see the impact of your infrastructure decisions
- Help shape architecture, engineering standards, and the growth of the infrastructure team
- Solve challenging problems around distributed systems, Kubernetes, GPU orchestration, and AI workloads

