Research Engineer, Benchmarks
clera · Singapore · $150,000 to $250,000
Open. First seen 18 September 2026.
Description
About the Role
This is a core technical role on a small, high-caliber team building rigorous benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. You will own the design and implementation of evaluations that frontier labs and enterprise customers rely on to understand real-world agent performance. The work is critical to the credibility and impact of the company's benchmark platform.
What You'll Do
Design, implement, and maintain the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks. Partner with subject-matter experts to define realistic workflows and translate them into well-scoped evaluation tasks. Build reliable infrastructure to run models and agents against benchmark tasks at scale.
Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes. Validate that benchmark performance correlates with real-world evaluations and customer expectations. Write clear technical documentation and benchmark reports for research and engineering audiences.
What We're Looking For
2 to 4 years of experience in research engineering or machine learning engineering, with a focus on AI benchmarks, evaluation infrastructure, or agent environments. Strong proficiency in Python, Docker, and Linux for building research or production infrastructure. Hands-on experience designing and running benchmarks or evaluation environments for AI agents or large language models.
Experience developing metrics and validation studies to assess benchmark difficulty, reliability, and real-world correlation. Experience collaborating with domain experts to turn workflows into concrete evaluation criteria. Strong technical writing skills; published papers or blog posts on AI benchmarking, model evaluation, or failure modes are a plus.
Experience with reinforcement learning training pipelines, data generation, or RL agent evaluation is a plus. Background at a frontier AI lab, research institution, or on a widely used public benchmark project is a plus. Comfort working independently in fast-paced, early-stage environments with unstructured problem spaces.
Sharp attention to detail and the ability to reason from first principles about task design, scoring, and edge cases. Compensation & Benefits Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available.
Location On-site in Singapore .
ABOUT THE ROLE
This is a core technical role on a small, high-caliber team building rigorous benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. You will own the design and implementation of evaluations that frontier labs and enterprise customers rely on to understand real-world agent performance. The work is critical to the credibility and impact of the company's benchmark platform.
WHAT YOU'LL DO
- Design, implement, and maintain the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.
- Partner with subject-matter experts to define realistic workflows and translate them into well-scoped evaluation tasks.
- Build reliable infrastructure to run models and agents against benchmark tasks at scale.
- Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes.
- Validate that benchmark performance correlates with real-world evaluations and customer expectations.
- Write clear technical documentation and benchmark reports for research and engineering audiences.
WHAT WE'RE LOOKING FOR
- 2 to 4 years of experience in research engineering or machine learning engineering, with a focus on AI benchmarks, evaluation infrastructure, or agent environments.
- Strong proficiency in Python, Docker, and Linux for building research or production infrastructure.
- Hands-on experience designing and running benchmarks or evaluation environments for AI agents or large language models.
- Experience developing metrics and validation studies to assess benchmark difficulty, reliability, and real-world correlation.
- Experience collaborating with domain experts to turn workflows into concrete evaluation criteria.
- Strong technical writing skills; published papers or blog posts on AI benchmarking, model evaluation, or failure modes are a plus.
- Experience with reinforcement learning training pipelines, data generation, or RL agent evaluation is a plus.
- Background at a frontier AI lab, research institution, or on a widely used public benchmark project is a plus.
- Comfort working independently in fast-paced, early-stage environments with unstructured problem spaces.
- Sharp attention to detail and the ability to reason from first principles about task design, scoring, and edge cases. COMPENSATION & BENEFITS Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available. LOCATION On-site in Singapore.
Posting history
Reposted once since 9 September 2026: a posting for this role in Singapore was taken down and the role was posted again.
- 18 September 2026 – open (this posting)
- 11 September 2026 – closed 17 September 2026
- 9 September 2026 – closed 12 September 2026
Similar jobs
- Research Engineer, Benchmarks at hud
- AI Infrastructure / Product Engineer at clera
- Member of Technical Staff at clera
- Senior Integration Engineer at clera
- Forward Deployed Engineer at clera
- Product Engineer at clera
Is this your posting and you want it taken down? Email info@careerholo.com.