Production Operations Engineer

Hcl Global Systems Inc — United States · Posted ~4 hours ago

Mid Full-time

Skills

GitLab CI/CD Kubernetes PostgreSQL database deployments application troubleshooting monitoring incident response Dynatrace

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A production operations engineering role responsible for maintaining reliable platforms, improving automation, monitoring system health, and resolving complex technical incidents in enterprise environments.

Highlights

Operations role supporting critical platforms with focus on reliability, automation, monitoring, and incident resolution. Opportunity to improve deployment processes and system stability.

Description

Position Summary Bank is seeking a highly motivated Production Operations (ProdOps) Engineer to manage and support mission-critical production platforms. The ideal candidate will have strong expertise in GitLab CI/CD, Kubernetes, PostgreSQL, Database Deployments, Application Troubleshooting, Dynatrace Monitoring, and On-Call Production Support. This role requires a proactive individual who can ensure platform stability, automate operational processes, improve deployment efficiency, and rapidly resolve production incidents. Essential Job Functions: Production Operations & Support • Provide operational support for production and non-production environments. • Participate in a 24x7 on-call rotation and respond to critical production incidents. • Monitor system health, application performance, and infrastructure availability. • Perform root cause analysis (RCA) and implement preventive measures. • Troubleshoot complex applications, databases, infrastructure, and deployment issues. Must have a Development Background CI/CD & Release Management • Design, maintain, and optimize GitLab CI/CD pipelines. • Automate application deployment, testing, and release processes. • Implement deployment best practices to improve reliability and reduce downtime. • Collaborate with development teams to streamline DevOps processes. Kubernetes Platform Management • Deploy, configure, and manage containerized applications on Kubernetes. • Troubleshoot pod, service, ingress, and cluster-level issues. • Manage scaling, availability, and performance optimization of Kubernetes workloads. • Work with Helm charts and Kubernetes manifests for application deployments. Database Deployment & Administration • Plan and execute database deployment activities across environments. • Support database schema changes, migrations, and rollback procedures. • Monitor and optimize database performance. • Collaborate with development teams on database release strategies. PostgreSQL Administration • Support and maintain PostgreSQL databases. • Troubleshoot database performance bottlenecks and connectivity issues. • Manage backup, recovery, replication, and high-availability configurations. • Ensure database security and compliance standards are maintained. Monitoring & Observability • Configure and maintain monitoring dashboards using Dynatrace. • Analyze application and infrastructure performance metrics. • Create alerts and proactive monitoring strategies to identify issues before customer impact. • Use observability tools to drive system reliability and operational excellence. Continuous Improvement • Develop automation scripts and operational tooling. • Create and maintain operational runbooks and standard operating procedures. • Identify opportunities for process improvement and operational efficiency. • Partner with engineering teams to improve platform reliability and scalability. Position Requirements: Technical Skills • GitLab CI/CD • Pipeline creation and optimization • GitLab Runners • Deployment automation • Kubernetes • Cluster operations and administration • Helm Charts • Pod, Service, Ingress troubleshooting • Database Deployments • Schema migrations • Release management • Rollback strategies • Monitoring & Observability • Dynatrace • Application Performance Monitoring (APM) • Log analysis and alerting • Troubleshooting • Application debugging • Infrastructure issue diagnosis • Performance analysis • Production incident management Operational Skills • 24x7 On-Call Support • Incident Management • Root Cause Analysis (RCA) • Problem Management • Change Management • Release Coordination ________________________________________ Preferred Qualifications • Bachelor’s degree in computer science, Information Technology, or related field. • 5+ years of experience in Production Support, DevOps, SRE, or Platform Engineering roles. • Experience with Linux administration and shell scripting. • Familiarity with cloud platforms such as Azure, AWS, or GCP. • Experience with Infrastructure as Code (Terraform preferred). • Understanding of microservices architecture and containerization technologies. ________________________________________ Success Metrics • Production availability and up time. • Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR). • Deployment success rate. • Reduction in recurring incidents. • Platform reliability and performance improvements. • Operational automation and efficiency gains. Keywords GitLab CI/CD, Kubernetes, PostgreSQL, Database Deployments, Dynatrace, Production Support, Incident Management, Troubleshooting, On-Call Support, DevOps, SRE, Platform Engineering, Monitoring, Release Management, Automation.