Staff / Principal Backend Engineer — Unstructured Data

ApplyApply
Posted about 2 hours ago
Share
Full Time
Seattle, Washington
$160,000 - $200,000 Annually

Staff / Principal Backend Engineer — Unstructured Data & Document Pipelines

Location: Seattle, WA — Hybrid, 2 days/week onsite
Compensation: $160,000–$200,000 base + up to 1% equity
Level: Staff / Principal

What is the core job?

Own, operate, and improve a major vertical of a distributed infrastructure platform.

The platform runs large-scale services that execute millions of queries, process hundreds of millions of tokens every minute, and ingest, transform, store, and retrieve substantial volumes of structured and unstructured data.

The difficult part is not simply adding capacity. Workloads vary enormously: one customer’s matter may be 1,000 times larger or more demanding than another’s while both run on shared infrastructure. You will design systems that remain fair, isolated, observable, and predictable under contention.

This is an end-to-end ownership role. There is no separate platform, SRE, infrastructure, or database team responsible for finishing the work. When you design a system, you will also own its infrastructure definitions, deployment configuration, production promotion, observability, operational behavior, and incident response.

This is not primarily an architecture or advisory position. You will write production code, investigate performance and reliability problems, operate what you build, and establish technical patterns that other engineers can use.

Success means

  • One customer’s workload cannot degrade another’s. Large jobs are isolated, admission is fair, and tail latency remains predictable under contention.
  • Critical services have clear ownership, strong observability, understood failure modes, and reliable recovery paths.
  • The system handles extreme variance in matter size, query patterns, ingestion volume, and processing demand without requiring manual intervention.
  • Bottlenecks across ingestion, storage, retrieval, orchestration, and AI-processing pipelines are identified and removed.
  • Infrastructure, application code, deployment configuration, and production operation are treated as one engineering responsibility rather than separate functions.
  • The engineering team makes better architectural decisions because you contribute both technical leadership and working implementations.

Your background

You have built systems that ingest, transform, index, store, and retrieve large volumes of unstructured data.

Your experience may include:

  • Document-ingestion pipelines
  • Search, indexing, crawling, or retrieval systems
  • Data modeling for complex document collections
  • Deduplication, threading, family relationships, metadata extraction, and document enrichment
  • Large asynchronous processing pipelines at scale
  • Durable execution and workflow orchestration
  • Large relational and NoSQL data stores

You might come from an eDiscovery or legal-tech platform, a search or web-scale crawling team, or anywhere built around processing hostile unstructured data that cannot afford to lose a file.

Apply