SECTION I · THE BRIEF
Brief #73461Updated 08 OCT 2026SAN FRANCISCO, CAAshbyY COMBINATOR
Employbl Company Profile

MTS - Pre-Training Data & Acquisition Engineer

Specializing in the automation and integration industry for both residential and commercial settings.

Location
San Francisco, CA
Company size
10–50
Posted
Today
Via
Ashby
Section II · Full ProfileFree with an account
  • 01Comp band & equity packageLocked
  • 02Seniority & experience requirementsLocked
  • 03Interview process & rubricLocked
  • 04Hiring manager & team contextLocked
  • 05Growth trajectory in this roleLocked
  • 06Offer & decision timelineLocked

Free account · no card · 2 minutes

MTS - Pre-Training Data & Acquisition Engineer

Prometheus Technologies· San Francisco View company profile


Job title
MTS - Pre-Training Data & Acquisition Engineer
Job location
San Francisco
Job description

The Role

We’re hiring a Pre-Training Data & Acquisition Engineer to build the data systems powering Prometheus’s foundation models for the physical world. You’ll work closely with research and infrastructure teams to acquire, process, and deliver large-scale training datasets across engineering, scientific, and multimodal domains. This role spans distributed crawling, source integration, data processing, and production operation, with end-to-end ownership from raw content to training-ready datasets.

What You’ll Do

  • Identify and integrate valuable data sources across engineering and scientific domains.

  • Build distributed crawlers, API integrations, and bulk ingestion systems with effective scheduling, rate limiting, retries, and incremental updates.

  • Develop pipelines for parsing, extraction, normalization, deduplication, quality filtering, and tokenization across heterogeneous formats.

  • Optimize throughput and cost across networking, compute, storage, and databases as acquisition and processing workloads scale.

  • Build monitoring and tooling to track source coverage, ingestion failures, processing throughput, and usable data yield.

  • Work closely with pre-training researchers to translate data requirements into reliable pipelines and deliver datasets ready for large-scale training.

  • Own dataset reproducibility, versioning, provenance, and recovery from acquisition through delivery.

What We’re Looking For

  • Experience building and operating large-scale distributed systems, web crawlers, or data processing pipelines.

  • Strong programming ability in Python and Rust, Go, C++, or a comparable systems language.

  • A practical understanding of web infrastructure, including HTTP, DNS, concurrency, caching, and common failure modes.

  • Hands-on experience with databases, object storage, and distributed batch or streaming processing.

  • Experience designing fault-tolerant systems that handle partial failures, resume interrupted work, and prevent unintended duplication or data loss.

  • Ability to profile and debug performance across CPU, memory, disk, and network usage.

  • Strong technical judgment when integrating unfamiliar sources, APIs, and file formats.

  • Bias toward fast iteration and end-to-end ownership, from initial implementation through reliable production operation.

  • Experience with search indexing, document extraction, or foundation-model data pipelines is a plus.

Why Join Us

  • Work with world-class researchers on frontier AI systems for the physical world.

  • Build the acquisition and processing systems that supply engineering, scientific, and multimodal data to large-scale model training.

  • Competitive compensation and flexible work arrangements.

  • High-impact, mission-driven environment.

View job listing ↗
The Saturday Briefing

Get the Saturday tech briefing

New company profiles, funding moves, and who’s hiring across the market — every Saturday morning.

Prometheus Technologies headquarters

San Jose, CA

Company size

10–50 employees

Founded

2020

Total raised

$3,550,000

View company profile ↗

Work at Prometheus Technologies? Claim its Employbl profile

Funding rounds