Summary
✨ AI‑Generated
Join an engineering team responsible for building and operating large-scale ML infrastructure. You will design model-serving systems, automate training and production promotion, implement safe deployment strategies, optimize inference performance, monitor model health and drift, and manage accelerated compute resources. The role suits an experienced engineer who enjoys production-scale ML systems, reliability, and performance optimization.
Highlights
Own end-to-end ML infrastructure for production computer vision and fraud-detection models, with a strong focus on low latency, reliability, auditability, automated delivery, observability, and cloud resource optimization.
Description
01About the role
Our forensic engine runs multiple vision models per request — ELA analysis, compression artifact detection, font consistency checks, and transformer-based fraud classifiers.
As an ML Infrastructure Engineer, you'll own the systems that train, version, deploy, and serve these models at scale.
You'll ensure every API call returns in under 200ms while maintaining full auditability and zero-downtime deployments.
02What you'll do
>Build and maintain the model serving infrastructure on GCP (Cloud Run, GKE, or Vertex AI).>Design the CI/CD pipeline for model training, evaluation, and promotion to production.>Implement model versioning, canary deployments, and A/B testing frameworks.>Optimize inference latency through batching, quantization, and hardware-aware compilation.>Build monitoring and alerting for model performance, drift detection, and throughput.>Manage GPU/TPU resource allocation and cost optimization across training and serving.
03What we're looking for
>3+ years in ML infrastructure, MLOps, or backend systems engineering.>Strong experience with Kubernetes, Docker, and cloud-native deployments on GCP or AWS.>Hands-on experience with model serving frameworks (Triton, TorchServe, TF Serving, or similar).>Proficiency in Python and at least one systems language (Go, Rust, or C++).>Experience with CI/CD pipelines for ML (MLflow, Kubeflow, or custom).>Understanding of GPU inference optimization (TensorRT, ONNX Runtime).
04Nice to have
+Experience with Vertex AI or SageMaker in production.+Background in real-time serving systems with strict latency SLAs (+Familiarity with cost modeling for GPU workloads.+Experience with distributed training across multiple GPUs/nodes.+Terraform or Pulumi for infrastructure-as-code.