Software Engineer (Backend-Focused)
Backend Engineer – AI Infrastructure
Location: Remote, U.S. | Seattle-area team presence
Compensation: $140-239k + equity
About the Company
Our client is a fast-growing AI technology company building enterprise AI solutions for complex, data-intensive industries.
The team works at the intersection of applied AI, machine learning, and enterprise software, helping organizations deploy AI in environments where security, reliability, governance, and measurable business impact matter.
This is an opportunity to join an early, highly technical team working on sophisticated AI infrastructure and production systems used to solve real-world problems.
About the Role
As a Backend Engineer, you’ll build the core services powering the company’s AI products and customer solutions.
A major focus of the role is LLM infrastructure: building the layer that provides a governed interface across multiple hosted and self-hosted models while giving enterprise users visibility into model usage, cost, permissions, and performance.
You’ll work closely with engineers actively using the infrastructure you build, creating a tight feedback loop between platform development and real-world implementation.
What You’ll Do
-
Own core components of the LLM gateway, including routing, metering, budgets, guardrail composition, and provider/backend adapters across hosted and self-hosted models.
-
Build and manage retrieval and knowledge infrastructure, including connectors, chunking, hybrid search, reranking, grounded responses, and MCP integrations.
-
Develop evaluation infrastructure that measures retrieval and model performance with confidence intervals and segment-level analysis.
-
Own and evolve asynchronous execution frameworks, including public APIs, rate-limiting algorithms, administrative tooling, and durable execution for agent workflows.
-
Design authorization systems including token hierarchies, capabilities, and tenant isolation.
-
Build production systems with strong observability, tracing, evaluation, and cost-accounting practices.
-
Partner closely with technical teams to translate real-world requirements into scalable infrastructure.
What We’re Looking For
-
5+ years of deep, production-level Python experience, particularly asynchronous Python, including cancellation scopes, streaming lifecycles, and connection pooling.
-
Strong experience with Postgres as infrastructure, including concepts such as MVCC, advisory locks, and vacuum discipline.
-
Meaningful distributed systems experience and a strong understanding of when technologies such as Redis are and are not appropriate.
-
Experience building API surfaces used by other engineers, including versioning, idempotency, error handling, migrations, and documentation.
-
Deep experience in at least one of the following:
-
LLM infrastructure: routing, metering, guardrails, provider failover, or model gateways.
-
Retrieval engineering: hybrid search, reranking, evaluation methodology, or production RAG.
-
-
Working knowledge of the other area and an ability to operate across the broader AI infrastructure stack.
-
Experience building evaluation or testing systems capable of identifying production regressions.
-
Strong engineering fundamentals and the ability to quickly ramp on unfamiliar technologies.
Technical Environment
The stack includes technologies such as:
-
Python 3.12+
-
Async Python
-
FastAPI / Starlette
-
Pydantic
-
PostgreSQL / pgvector
-
asyncpg / SQLAlchemy
-
Redis-compatible infrastructure
-
OpenTelemetry
-
SSE and streaming APIs
-
LLM provider APIs
-
MCP
-
Background job and durable execution frameworks
-
Modern rate-limiting approaches
Experience with every technology listed is not required, but candidates should have meaningful production depth in a modern backend stack and the ability to learn quickly.
Nice to Have
-
Experience working with AI or ML systems in production.
-
Experience in complex, highly regulated, or data-intensive industries.
-
Background working in both startup and larger enterprise environments.
-
Experience building infrastructure for agentic AI applications.

