Research Engineer, Privacy and Anonymization
clera · San Francisco
Open. First seen 17 September 2026.
Description
About the Role
This is a Research Engineer role focused on building privacy and anonymization systems that make sensitive, real-world data safe and useful for AI training. You will own the full pipeline for protecting privacy without destroying the structure and signal that make data valuable, sitting at the intersection of applied research and production engineering in a fast-moving AI infrastructure company.
What You'll Do
Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information, designing transformations based on data type and downstream use case. Develop and benchmark detection approaches that combine rules, statistical models, classifiers, and LLM-based methods. Build production pipelines that anonymize raw data before it enters downstream processing, training, evaluation, or synthetic data generation workflows.
Create evaluation frameworks that measure privacy risk and retained data utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts. Design systems that remain robust to new data sources, schema drift, unusual formats, and sensitive information embedded in unexpected fields. Collaborate with engineering, research, operations, and customers to translate privacy requirements into practical technical policies and safeguards.
What We're Looking For
2+ years building production data or ML systems in Python, with strong proficiency in the language. Hands-on experience with information extraction, named-entity recognition, classification, or related methods for detecting sensitive or rare content. Strong experimental instincts and the ability to compare approaches across recall, precision, latency, cost, and downstream data utility.
Solid understanding of redaction, masking, pseudonymization, anonymization, and synthetic data generation, and when each technique is appropriate. Experience designing systems that are robust to schema drift, unusual data formats, and edge cases. End-to-end experience building data processing pipelines without a fully prescribed roadmap.
Familiarity with privacy-enhancing technologies such as differential privacy, k-anonymity, secure aggregation, or format-preserving encryption is a plus. Experience with low-latency or high-throughput ML inference and data-processing systems is a plus. Prior work with sensitive data in healthcare, finance, security, or related domains is a plus.
Location On-site in San Francisco, California. Visa sponsorship is available.
ABOUT THE ROLE
This is a Research Engineer role focused on building privacy and anonymization systems that make sensitive, real-world data safe and useful for AI training. You will own the full pipeline for protecting privacy without destroying the structure and signal that make data valuable, sitting at the intersection of applied research and production engineering in a fast-moving AI infrastructure company.
WHAT YOU'LL DO
- Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information, designing transformations based on data type and downstream use case.
- Develop and benchmark detection approaches that combine rules, statistical models, classifiers, and LLM-based methods.
- Build production pipelines that anonymize raw data before it enters downstream processing, training, evaluation, or synthetic data generation workflows.
- Create evaluation frameworks that measure privacy risk and retained data utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts.
- Design systems that remain robust to new data sources, schema drift, unusual formats, and sensitive information embedded in unexpected fields.
- Collaborate with engineering, research, operations, and customers to translate privacy requirements into practical technical policies and safeguards.
WHAT WE'RE LOOKING FOR
- 2+ years building production data or ML systems in Python, with strong proficiency in the language.
- Hands-on experience with information extraction, named-entity recognition, classification, or related methods for detecting sensitive or rare content.
- Strong experimental instincts and the ability to compare approaches across recall, precision, latency, cost, and downstream data utility.
- Solid understanding of redaction, masking, pseudonymization, anonymization, and synthetic data generation, and when each technique is appropriate.
- Experience designing systems that are robust to schema drift, unusual data formats, and edge cases.
- End-to-end experience building data processing pipelines without a fully prescribed roadmap.
- Familiarity with privacy-enhancing technologies such as differential privacy, k-anonymity, secure aggregation, or format-preserving encryption is a plus.
- Experience with low-latency or high-throughput ML inference and data-processing systems is a plus.
- Prior work with sensitive data in healthcare, finance, security, or related domains is a plus. LOCATION On-site in San Francisco, California. Visa sponsorship is available.
Other openings for this role
2 open postings for this role in San Francisco, first posted 17 September 2026. The company is likely hiring for more than one position. None has been taken down and posted again.
- 24 September 2026 – open
- 17 September 2026 – open (this posting)
Similar jobs
- Research Engineer, Privacy and Anonymization at hud
- AI Infrastructure / Product Engineer at clera
- Member of Technical Staff at clera
- Senior Integration Engineer at clera
- Forward Deployed Engineer at clera
- Product Engineer at clera
Is this your posting and you want it taken down? Email info@careerholo.com.