What the work is
Role Overview
We are looking for a hands-on DevOps and Cloud Infrastructure Engineer to build and operate the infrastructure behind Lazarus, a large-scale platform for PII detection, redaction, and human review.
You will design secure, scalable, observable, and cost-efficient infrastructure for processing sensitive datasets across text, documents, images, and other file formats. The role involves supporting CPU- and GPU-intensive workloads, large-scale batch processing, ML inference pipelines, and production cloud operations across GCP and AWS.
What You’ll Do
- Design, provision, and operate GCP infrastructure for PII detection and redaction pipelines.
- Build and manage workloads across Cloud Run, Cloud Run Jobs, Compute Engine, GKE, and GPU-backed infrastructure.
- Design scalable batch-processing systems capable of processing hundreds of thousands of files.
- Implement worker parallelism, queues, retries, checkpointing, idempotency, timeouts, and failure recovery.
- Manage data movement across Google Cloud Storage, Amazon S3, VMs, containers, and external storage systems.
- Secure sensitive datasets using IAM, service accounts, Secret Manager, private networking, controlled egress, IAP, encryption, bucket-level access controls, and audit logging.
- Deploy and operate containerized Python and ML workloads using Docker.
- Support PII and ML services such as Google Sensitive Data Protection, Presidio, OCR and vision systems, NER models, and LLM-based validation pipelines.
- Provision and manage GPU infrastructure, including drivers, CUDA, quotas, autoscaling, and model-serving environments.
- Support model-serving stacks such as vLLM, Hugging Face Transformers, and Triton.
- Build and maintain CI/CD pipelines for Cloud Run, VMs, containers, and related services.
- Establish observability through centralized logging, metrics, alerting, job tracking, worker-health monitoring, and infrastructure utilization dashboards.
- Optimize throughput and cost through infrastructure selection, concurrency tuning, autoscaling, storage access patterns, and API rate-limit management.
- Develop backup, recovery, migration, and disaster-recovery processes for large datasets and cloud resources.
- Troubleshoot production issues involving networking, storage, IAM, containers, APIs, compute, and ML infrastructure.
Who gets hired
- 5+ years of experience in DevOps, cloud infrastructure, platform engineering, or a related field.
- Strong hands-on experience with GCP, including:
- Cloud Run and Cloud Run Jobs
- Compute Engine and GKE
- Google Cloud Storage
- IAM and service accounts
- Secret Manager
- VPC and private networking
- Identity-Aware Proxy
- Artifact Registry
- Cloud Logging and Monitoring
- Working experience with AWS services, particularly Amazon S3 and cross-cloud data movement.
- Strong Linux administration and shell-scripting skills.
- Strong Python skills for infrastructure automation, operational tooling, and data-processing workflows.
- Strong Docker and containerization experience.
- Experience designing and operating large-scale batch or data-processing pipelines.
- Solid understanding of queues, worker parallelism, retries, checkpoints, idempotency, timeouts, and failure recovery.
- Experience securely processing and transferring sensitive or regulated data.
- Experience with CI/CD tools and production deployment workflows.
- Experience debugging CPU, memory, disk I/O, network, storage, API, and GPU performance issues.
- Exposure to ML infrastructure, GPU workloads, OCR, NLP, LLM serving, or model inference systems.
- Strong problem-solving skills and the ability to independently investigate and resolve production issues.
Perks of Freelancing With Turing
- Work in a fully remote environment.
- Opportunity to work on cutting-edge AI projects with leading LLM companies.
Offer Details
- Commitments Required: 40 hours per week with overlap of 6 hours per day with PST.
- Duration of Contract: 1 month (adjustable based on engagement)
Pay
See listing, fully remote. How payouts and tax work.