What the work is
About the Role
We are looking for an ML/Data Engineer to improve the quality, reliability, and compliance of enterprise data pipelines. You will focus on data quality validation and testing PII/PHI de-identification systems to ensure sanitized data is accurate, consistent, and free from sensitive information leaks.
This role combines data engineering, ML evaluation, test automation, and compliance-focused quality assurance.
What You’ll Do
- Design and automate validation suites for data pipelines, including schema checks, completeness validation, drift detection, and reconciliation across pipeline stages.
- Perform deep data quality analysis across enterprise data sources and connectors, including topic coherence, domain coverage, consistency, and depth.
- Build adversarial test datasets for PII/PHI de-identification systems, covering edge cases, obfuscated identifiers, multilingual entities, OCR noise, and unusual document formats.
- Evaluate NER and ML-based de-identification systems using precision, recall, F1 score, leak rates, and false-negative analysis.
- Identify and investigate PII leakage risks across raw, processed, and sanitized data.
- Implement automated regression gates in CI/CD to prevent pipeline changes from being deployed without passing data quality and privacy checks.
- Conduct sampling-based human-in-the-loop audits and maintain detailed audit trails for compliance evidence.
- Partner with data and engineering teams to perform root-cause analysis and resolve data inconsistencies, quality issues, and privacy leaks.
- Develop monitoring and reporting mechanisms for pipeline quality, de-identification performance, and compliance risks.
Who gets hired
- 5+ years of experience in data engineering, ML engineering, data quality, or a related field.
- Strong Python skills for test automation, data validation, and ML evaluation.
- Experience with testing frameworks and data-quality tools such as pytest, Great Expectations, Pandera, or similar.
- Strong SQL skills and experience validating data across multiple pipeline stages.
- Understanding of PII and PHI categories and de-identification concepts.
- Familiarity with privacy and compliance frameworks such as HIPAA Safe Harbor, GDPR, or LGPD.
- Experience evaluating NER or other ML-based systems using labeled datasets and precision/recall metrics.
- Ability to design reliable evaluations for non-deterministic ML or LLM-based systems.
- Experience integrating automated tests and quality checks into CI/CD pipelines.
- Familiarity with GCP services such as BigQuery, Google Cloud Storage, and Cloud Run Jobs.
- Strong analytical, debugging, and communication skills.
Perks of Freelancing With Turing
- Work in a fully remote environment.
- Opportunity to work on cutting-edge AI projects with leading LLM companies.
Offer Details
- Commitments Required: 40 hours per week with overlap of 6 hours per day with PST.
- Duration of Contract: 1 month (adjustable based on engagement)
Pay
See listing, fully remote. How payouts and tax work.