Description
Overview
The DevOps Engineer plays a pivotal role in driving operational efficiency and scalability across our development and IT teams.
This role automates and streamlines the software development and deployment processes behind our next-generation tolling and intelligent transportation platform — a cloud-native, event-driven microservices system — ensuring faster and more reliable delivery of applications.
By bridging the gap between development and operations, this role helps reduce bottlenecks and increase collaboration, ultimately accelerating the release cycle.
This role directly impacts the company by enhancing system reliability, reducing downtime, and ensuring our infrastructure can scale with growing business demands.
This role is crucial for maintaining a robust and agile technology environment, allowing the company to deliver high-quality products to customers more efficiently and remain competitive in the market.
This position participates in a 24/7 on-call rotation supporting revenue-critical systems.
Responsibilities
Key Responsibilities :
Infrastructure Management
Design, deploy, and maintain scalable cloud infrastructure (AWS, Azure, Google Cloud).
Manage cloud resources to ensure high availability, reliability, and security.
Operate Kubernetes-based platforms end to end, including stateful workloads such as distributed SQL databases and event-streaming clusters (e.g., NATS, Kafka).
Optimize cloud infrastructure for cost efficiency and performance.
CI/CD Pipeline Development
Build and maintain CI/CD pipelines, including code-first, containerized pipelines (e.g., Dagger, GitHub Actions, GitLab CI).
Automate build, test, and deployment for a polyglot monorepo (e.g., Rust services and TypeScript applications) with reproducible developer toolchains.
Collaborate with development teams to integrate CI/CD best practices, including ephemeral local clusters (e.g., k3d, kind) that mirror production.
Automation & Infrastructure as Code (IAC)
Implement and manage infrastructure as code (IaC) with tools such as Terraform or Ansible, plus Helm charts and GitOps workflows (e.g., Argo CD) for Kubernetes workloads.
Automate provisioning, configuration, and scaling of infrastructure.
Ensure repeatability and consistency in infrastructure setup and deployment.
Monitoring & Incident Response
Set up and manage monitoring, logging, tracing, and alerting tools (e.g., Prometheus, Grafana, OpenTelemetry, ELK Stack).
Proactively monitor systems for performance issues and outages.
Participate in a 24/7 on-call rotation and respond to incidents to minimize downtime for revenue-critical services.
Containerization & Orchestration
Manage containerized applications using Docker and Kubernetes.
Deploy and manage event-driven microservices and federated GraphQL API routers in a containerized environment.
Ensure scalability and fault tolerance using orchestration tools.
Collaboration & Continuous Improvement
Work closely with software developers, system administrators, and QA engineers to troubleshoot and optimize system performance.
Participate in code reviews and provide feedback on best practices for deployment and operations.
Identify areas for process improvements and implement solutions to enhance productivity and system efficiency.
Security & Compliance
Integrate security best practices into infrastructure design and deployments, including policy-as-code (e.g., OPA) and centralized secrets management (e.g., Vault-compatible tooling).
Manage secrets, permissions, and access control within cloud environments.
Ensure compliance with relevant industry standards (e.g., SOC 2, ISO 27001, PCI DSS).
Qualifications
Education:
Bachelor’s Degree in Computer Science, Information Technology, Software Engineering, or a related field.
Equivalent practical experience may be considered in lieu of a formal degree.
3+ years of experience in a DevOps, Site Reliability Engineer (SRE), or similar role, with a strong focus on cloud infrastructure and automation.
Proven experience with cloud platforms such as AWS, Azure, or Google Cloud.
Hands-on experience with CI/CD tools (e.g., Dagger, GitHub Actions, GitLab CI, Jenkins) and automation frameworks (e.g., Terraform, Ansible, Helm).
Experience with containerization and orchestration tools like Docker and Kubernetes.
Strong knowledge of scripting languages (Python, Bash, etc.) for automation.
Familiarity with monitoring and logging solutions (e.g., Prometheus, Grafana, ELK Stack).
Required Skills/Abilities
Cloud Infrastructure Management: Strong experience with cloud platforms such as AWS, Azure, or Google Cloud, including deploying and managing scalable and reliable systems.
CI/CD Pipeline Management: Proficiency in designing, implementing, and maintaining Continuous Integration and Continuous Delivery (CI/CD) pipelines, including code-first pipeline-as-code approaches (e.g., Dagger, GitHub Actions, GitLab CI).
Infrastructure as Code (IaC): Expertise in using tools like Terraform, Ansible, and Helm to automate infrastructure provisioning and configuration, with GitOps-based deployment (e.g., Argo CD).
Containerization & Orchestration: Deep knowledge of Docker and Kubernetes for managing and deploying containerized applications at scale.
Scripting & Automation: Proficiency in scripting languages such as Python, Bash, or PowerShell for automating routine tasks and enhancing operational efficiency.
Monitoring & Logging: Experience with monitoring tools (e.g., Prometheus, Grafana, OpenTelemetry) and logging systems (e.g., ELK Stack) to track system performance and resolve issues promptly.
Version Control: Strong knowledge of version control systems, particularly Git (GitHub, GitLab), for managing code and configuration changes.
Problem-Solving: Strong analytical and troubleshooting skills to identify and resolve system bottlenecks, deployment issues, or infrastructure challenges.
Collaboration & Communication: Ability to work effectively across cross-functional teams (developers, system admins, QA engineers) and communicate complex technical information clearly to both technical and non-technical stakeholders.
Adaptability: Willingness to quickly learn and adopt new technologies and tools to stay current with industry trends and business needs.
Attention to Detail: Precision in executing automation and deployment processes, ensuring minimal errors and system downtime.
Time Management: Strong organizational skills to prioritize tasks in a fast-paced environment, managing both short-term troubleshooting and long-term infrastructure improvements.
Benefits
We offer a Total Rewards plan designed with you and your family’s health and wellness in mind that includes:
Paid days off ( i.
e.
vacatio n, sick days, bereavement leave) Health and Dental plans Retirement plans Employee and Family Assistance Program (EFAP) Employee referral program
We welcome applicants from all backgrounds, regardless of race, color, religion, sex, veteran sta tus, sexual orientation, gender identity, national origin, age, or disabilit y or any other protected characteristics in accordance with applicable federal, state/provincial, and local laws .
We 're committed to creating a workplace where everyone feels valued and respected.
We appreciate all responses and will acknowledge only those being considered for an interview .
We respectfully request no calls or unsolicited resumes from Agencies .