[Proto] MTS - Software Engineer, Data Acqusition & Crawling
- Salary
- Not stated
- Level
- Mid
- Work type
- Hybrid · San Francisco
- Visa
- Not stated
Open. First seen 7 October 2026.
About the role
The Role
The best models are trained on data nobody else bothered to collect. This role owns how we acquire the internet: a crawler and ingestion system that pulls text, code, images, video, audio, 3D assets, documents, and stranger formats from across the web at petabyte scale, and a standing search for sources other labs overlook.
You will design and run the distributed systems that do the fetching, extraction, dedup, and delivery, and you will spend real time hunting for new corpora about the physical world and the economy: engineering and manufacturing data, CAD and mesh repositories, sensor and telemetry streams, scientific archives, industrial and logistics records, public filings, maps and geospatial, and whatever else turns out to teach a model how the world actually works. The job is half systems engineering at scale and half data research: every acquisition decision is measured by what it does to the model.
What you'll do
- Crawl & ingest: build and operate a large-scale, polite, fault-tolerant web crawler and the ingestion pipelines behind it, from scheduling and fetch to extraction, normalization, versioning, and delivery to the pretraining data team.
- Odd modalities: build specialized acquisition paths for video, 3D (meshes, point clouds, CAD, scenes), audio, scanned documents, and structured or semi-structured datasets, including the format parsing and storage layout those need.
- New sources: systematically find, assess, and acquire high-value corpora tied to the physical world and the economy, and work with legal on licensing, compliance, robots and rate limits, and privacy.
- Data research: run experiments on crawling strategy, extraction quality, and coverage; analyze what we have for gaps and redundancy; close the loop with the pretraining and data teams on what moves evals.
- Infrastructure: own petabyte-scale storage, indexing and search over the corpus, and deployment on Kubernetes with infrastructure as code.
What we're looking for
- Deep expertise in large, stateful, high-throughput distributed systems, with production experience in Rust or C++ [plus Python for glue and analysis].
- Hands-on experience building or running a large web crawler or a comparable internet-scale acquisition system; you know the practical problems (politeness, dedup, trap detection, extraction, freshness) firsthand.
- Comfort with multi-TB to PB datasets and with making systems observable, testable, and reproducible at that scale.
- Genuine curiosity about data: you enjoy finding unusual sources, reading their formats, and arguing from evidence about whether they will matter for a model.
- Enough familiarity with how LLMs are trained and evaluated to connect acquisition choices to model outcomes.
- Bonus: experience with video, 3D, or scientific data formats and pipelines; background in web archiving, search infrastructure, or data engineering frameworks (Spark, Ray, Beam). [Domain bonus: simulation, CAD, geospatial, or industrial data.]
See Software Engineer salaries in San Francisco →
Similar jobs
- AI Developer at Salvo Software
Shares 7 skills: Prometheus, Rust, Golang… · same seniority - Software Engineering Manager (Edge Systems) at Sift Stack
Shares 5 skills: Prometheus, Rust, Golang… · same country - Software Engineer, Core Platforms at Bot Auto
Shares 5 skills: Prometheus, Golang, Kubernetes… · same country - Senior Software Engineer - Ground & Mission Systems at Star Catcher
Shares 5 skills: Prometheus, Golang, C++… · same country - Senior/Middle Data Engineer at Raiffeisen Bank Ukraine
Shares 5 skills: Prometheus, Rust, Golang… - Senior Front-End Infrastructure Engineer (Founding) at prometheus
Also at this company
Is this your posting and you want it taken down? Email info@careerholo.com.