Description
Job Title: Senior Platform Engineer
Location:San Jose CA /NY/ FL
Duration- Long Term
Job Overview
We are seeking a highly skilled Senior Platform Engineer to design, build, and operate scalable cloud infrastructure and internal developer platforms across GCP and AWS.
The ideal candidate will have strong hands-on experience with Terraform, Kubernetes, CI/CD, GitOps, cloud security, observability, and developer environment automation.
This role will focus on building reliable, secure, and highly automated platforms that improve developer productivity, operational efficiency, and application reliability.
Key Responsibilities
Design, implement, and maintain scalable cloud infrastructure across GCP and AWS.Develop and manage Infrastructure as Code (IaC) using Terraform, Terramate, and Ansible.Build and operate Kubernetes/GKE platforms, including cluster configuration, Helm deployments, node pools, workload management, and upgrades.Develop and maintain automated CI/CD and GitOps pipelines using GitHub Actions, Jenkins, GitLab CI, and Flux.Build internal developer platforms and self-service development environments using technologies such as Cloud Workstations, Coder, and ephemeral environments.Establish standardized infrastructure patterns and reusable Terraform modules across multiple cloud projects and environments.Implement cloud IAM, RBAC, secrets management, and least-privilege security controls.Drive DevSecOps practices, including infrastructure security, compliance automation, and integration of SAST/DAST/SCA tooling.Design and implement high-availability, disaster recovery, and reliability solutions for production platforms.Build and maintain observability solutions using Prometheus, Grafana, ELK, and Datadog.Define and monitor SLOs, SLAs, reliability metrics, and operational health.Troubleshoot complex infrastructure, networking, Kubernetes, and cloud-platform issues and participate in incident response and on-call operations.Identify opportunities for cloud cost optimization, automation, and infrastructure standardization.Partner closely with software engineers and development teams to improve developer experience, deployment velocity, and platform reliability.Automate repetitive operational processes using Python, Bash, YAML, and configuration-management tooling.Required Qualifications
8+ years of experience in Platform Engineering, DevOps, Cloud Infrastructure, or Site Reliability Engineering.Strong hands-on experience with Terraform and Infrastructure as Code.Strong experience with Kubernetes, preferably GKE and/or EKS.Extensive experience with GCP and AWS cloud platforms.Strong knowledge of CI/CD and GitOps practices using tools such as Jenkins, GitHub Actions, GitLab CI, or Flux.Experience with Ansible or other configuration-management technologies.Strong understanding of IAM, RBAC, secrets management, networking, and cloud security.Experience designing highly available and resilient infrastructure.Strong understanding of observability, monitoring, logging, alerting, and SLO/SLA management.Proficiency in Python, Bash, and YAML scripting/automation.Experience with developer platforms, self-service infrastructure, or ephemeral development environments is highly preferred.Strong troubleshooting, incident response, and production operations experience.Preferred Qualifications
Experience managing infrastructure across 50+ cloud projects/accounts.Experience with Terramate or similar infrastructure orchestration tools.Experience with Google Cloud Workstations, Coder, or internal developer platforms.Experience supporting AI/ML workloads on Kubernetes or cloud infrastructure.Experience with BigQuery, Pub/Sub, GCS, Dataproc, or other GCP data services.Knowledge of SOC 2, compliance-as-code, and DevSecOps practices.Experience implementing cloud cost optimization initiatives.Experience with private LLM/AI-assisted observability or infrastructure operations.