Research Engineer, Synthetic Data
clera · Singapore · $150,000 to $250,000
Open. First seen 19 September 2026.
Description
About the Role
This is a hands-on research engineering role focused on building synthetic data pipelines that turn real-world, domain-specific workflows into structured training tasks for AI agents. You will join a small, high-caliber engineering team of Olympiad medalists and published researchers, working at the core of a platform that powers reinforcement learning environments and post-training data for AI labs.
What You'll Do
Design and build end-to-end synthetic data pipelines that convert domain-specific workflows into realistic, challenging training tasks. Collaborate with subject-matter experts to produce synthetic tasks for AI agents across professional and technical domains. Develop task generation methods that maximize diversity, realism, and learnability.
Build tooling to mutate, validate, and iteratively improve synthetic tasks at scale. Analyze model and agent performance on synthetic tasks to identify what they teach and where they break down. Define and implement metrics to quantify synthetic task quality across diversity, realism, and learnability dimensions.
What We're Looking For
2 to 4 years of experience in software engineering, machine learning engineering, or AI research, with a focus on data pipelines, ML infrastructure, or synthetic data systems. Proficiency in Python and hands-on experience with Docker and Linux environments. Demonstrated experience applying synthetic data research methods to build generation pipelines end-to-end.
Strong understanding of synthetic data quality criteria, including diversity, realism, and learnability, as well as their inherent limitations. Experience designing, implementing, or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models. Track record of independently owning and delivering technical projects with minimal predefined requirements or roadmap.
Sharp eye for edge cases, subtle inconsistencies, and quality issues in synthetic or algorithmically generated datasets. Familiarity with reinforcement learning paradigms, agentic AI workflows, or LLM post-training pipelines is a plus. Strong communication skills for effective collaboration across time zones.
Compensation & Benefits Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available. Location On-site in Singapore.
ABOUT THE ROLE
This is a hands-on research engineering role focused on building synthetic data pipelines that turn real-world, domain-specific workflows into structured training tasks for AI agents. You will join a small, high-caliber engineering team of Olympiad medalists and published researchers, working at the core of a platform that powers reinforcement learning environments and post-training data for AI labs.
WHAT YOU'LL DO
- Design and build end-to-end synthetic data pipelines that convert domain-specific workflows into realistic, challenging training tasks.
- Collaborate with subject-matter experts to produce synthetic tasks for AI agents across professional and technical domains.
- Develop task generation methods that maximize diversity, realism, and learnability.
- Build tooling to mutate, validate, and iteratively improve synthetic tasks at scale.
- Analyze model and agent performance on synthetic tasks to identify what they teach and where they break down.
- Define and implement metrics to quantify synthetic task quality across diversity, realism, and learnability dimensions.
WHAT WE'RE LOOKING FOR
- 2 to 4 years of experience in software engineering, machine learning engineering, or AI research, with a focus on data pipelines, ML infrastructure, or synthetic data systems.
- Proficiency in Python and hands-on experience with Docker and Linux environments.
- Demonstrated experience applying synthetic data research methods to build generation pipelines end-to-end.
- Strong understanding of synthetic data quality criteria, including diversity, realism, and learnability, as well as their inherent limitations.
- Experience designing, implementing, or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.
- Track record of independently owning and delivering technical projects with minimal predefined requirements or roadmap.
- Sharp eye for edge cases, subtle inconsistencies, and quality issues in synthetic or algorithmically generated datasets.
- Familiarity with reinforcement learning paradigms, agentic AI workflows, or LLM post-training pipelines is a plus.
- Strong communication skills for effective collaboration across time zones. COMPENSATION & BENEFITS Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available. LOCATION On-site in Singapore.
Posting history
Reposted once since 9 September 2026: a posting for this role in Singapore was taken down and the role was posted again.
- 19 September 2026 – open (this posting)
- 12 September 2026 – closed 18 September 2026
- 9 September 2026 – closed 12 September 2026
Similar jobs
- Research Engineer, Synthetic Data at hud
- AI Infrastructure / Product Engineer at clera
- Member of Technical Staff at clera
- Senior Integration Engineer at clera
- Forward Deployed Engineer at clera
- Product Engineer at clera
Is this your posting and you want it taken down? Email info@careerholo.com.