Summary
✨ AI‑Generated
Join a production engineering team as a Senior Kubernetes Platform Support Engineer. You will support business-critical applications, resolve complex incidents, improve system availability, automate operational processes, and mentor colleagues while collaborating with engineering and infrastructure specialists.
Highlights
Senior technical role supporting business-critical production systems while driving automation, operational excellence, and incident recovery. Includes technical leadership, mentoring, and close collaboration with engineering and infrastructure teams.
Description
Role : Senior Kubernetes Platform Support Engineer
Skills: Production Support, SQL Queries, Batch job scheduling, ITIL concepts, Kubernetes and Docker, MQ or Kafka, observability tool Grafana
Location : Krakow, Poland (Hybrid)
Type: B2b/Permanent
We are at Coforge hiring for Senior Production Support Specialist with Production Support, SQL Queries, Batch job scheduling, ITIL concepts, Kubernetes and Docker, MQ or Kafka, observability tool Grafana
Role purpose
Within the Production Engineering team, you'll play a senior role supporting business-critical payment applications while leading operational improvements and incident recovery activities.You'll combine strong technical expertise with leadership, helping maintain highly available production systems while driving automation and operational excellence.
Acting as a senior technical contributor, you'll coordinate incident resolution, mentor team members, and work closely with engineering and infrastructure teams to improve service reliability.
Job Description:
6+ years of experience in Production Support, Application Support or Software Engineering, preferably within Corporate Banking or Financial Services.Strong Unix/Linux skills for production troubleshooting, including navigating the operating system, checking logs and investigating application behaviour.Ability to write complex SQL queries.Experience developing scripts in Bash, PowerShell, or Python.Practical experience supporting and maintaining Kubernetes environments.Experience supporting Java/J2EE applications.Experience with Oracle, MQ and WebSphere.Experience using monitoring and observability tools such as Splunk and AppDynamics.Strong understanding of IT infrastructure and distributed systems.Experience leading incident management and coordinating technical teams.Demonstrated track record of implementing automation and operational improvements.Practical experience using AI productivity tools (e.g.
Copilot, Claude) to support automation, documentation or operational analysis.Strong communication, stakeholder management, and collaboration skills.Knowledge of payment systems or payment processing concepts.Network troubleshooting knowledge.ITIL certification.Service support & incident managementLead service recovery during production incidents and coordinate technical resolver teams.Investigate application and infrastructure issues using logs, SQL queries, and Unix/Linux tools.Respond to user queries regarding application behaviour and service usage.Communicate incident impact, progress, and recovery plans to stakeholders.Escalate critical incidents appropriately and drive timely resolution.Problem management & continuous improvementLead or contribute to Root Cause Analysis (RCA) activities.Identify recurring operational issues and implement permanent improvements.Develop operational documentation, runbooks and knowledge sharing materials.Improve monitoring, alerting and operational processes to increase service resilience.Automation & Operational ExcellenceDesign and implement automation using Bash, PowerShell, or Python.Develop internal operational tools that reduce manual effort.Demonstrate practical use of AI tools (e.g.
Copilot, Claude) to automate operational activities and improve team productivity.Change & Release SupportParticipate in change reviews ensuring operational readiness and compliance with CLIENT standards.Support patching, platform maintenance, and disaster recovery activities.Identify operational risks and recommend appropriate mitigations.