This role has closed. Clera has taken the posting down.
hirly last saw it live on 30 September 2026. See similar open roles below, or browse all jobs in San Francisco.
Clera
Research Engineer, Privacy and Anonymization
San Francisco
Similar open jobs
- Research Engineer, Robotics EvalsHud · San FranciscoFirst seen today
- Research Engineer, Post-trainingMedraai · San FranciscoFirst seen yesterday
- Research Engineer, Index IntelligenceExa · San Francisco, CaliforniaFirst seen yesterday
- Research Engineer, Content UnderstandingExa · San Francisco, CaliforniaFirst seen yesterday
- Research Engineer / Research Scientist, RL FrontiersAnthropic · San Francisco, CA | New York City, NY | Seattle, WAFirst seen 3d agoremote
- Research Engineer / Performance Engineer, RL Distributed SystemsAnthropic · San Francisco, CA | New York City, NY | Seattle, WAFirst seen 3d agoremote
- Research Engineer , Quick ScienceAmazon · Seattle, Washington, USA; New York, New York, USA; San Francisco, California, USAFirst seen today
- Operations Research Engineer Level 3/4 (AHT)Ngc · United States-California-NorthridgeFirst seen today
- Research Engineer - 6G AI-Enabled Systems and TestbedsInterdigital · Conshohocken, PAFirst seen today
- RAN Research Engineer / Scientist (WirelessInterdigital · 2 LocationsFirst seen today
- Temporary Research EngineerUofl · Belknap CampusFirst seen yesterday
- AI Research Engineer (Kernel & Inference Optimization)Jobgether · USFirst seen yesterdayremote
- Frontier AI Research EngineerBah · 2 LocationsFirst seen 2d ago
- Associate Research Engineer | TitleistAcushnetgolf · Carlsbad, California, United States of AmericaFirst seen 2d ago
- Research Engineer, Field SimulationMonarch · Emeryville, CaliforniaFirst seen 2d ago
hirly's read of this role
- Seniority
- Mid level
- Country
- US
- Work mode
- On-site / unstated
- First seen by hirly
- 28 Sept 2026
Derived automatically from the posting.
the posting
About the Role
This Research Engineer role sits at the intersection of privacy engineering and AI data infrastructure. You will own the full pipeline for detecting and removing sensitive information from raw, real-world data before it flows into processing, training, evaluation, and synthetic data workflows. The work is foundational: protecting privacy without destroying the structure and signal that make data valuable for training frontier AI agents.
What You'll Do
Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information, and design appropriate transformations based on data type and downstream use case.
Develop and benchmark detection approaches that combine rules, statistical models, classifiers, and LLM-based methods.
Build production pipelines that anonymize raw data before it enters downstream processing, training, evaluation, or synthetic data generation workflows.
Create evaluation frameworks that measure privacy risk and retained data utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts.
Design systems that remain robust to new data sources, schema drift, unusual formats, and sensitive information embedded in unexpected fields.
Collaborate with engineering, research, operations, and customers to translate privacy requirements into practical technical policies and safeguards.
What We're Looking For
2+ years of experience building reliable production data or ML systems in Python.
Hands-on experience with information extraction, named-entity recognition, classification, or related methods for detecting sensitive or rare content.
Experience building end-to-end data processing pipelines without a fully prescribed roadmap.
Strong experimental instincts with the ability to compare approaches across recall, precision, latency, cost, and downstream data utility.
Solid understanding of redaction, masking, pseudonymization, anonymization, and synthetic data generation, and when each is appropriate.
Experience designing systems that are robust to schema drift, unusual formats, and edge cases.
Familiarity with privacy-enhancing technologies such as differential privacy, k-anonymity, secure aggregation, or format-preserving encryption is a strong plus.
Experience with low-latency or high-throughput ML inference and data processing systems is a plus.
Prior work with sensitive data in healthcare, finance, security, or related domains is a plus.
Location
This role is on-site in San Francisco, California, USA. Visa sponsorship is available.