Summary
✨ AI‑Generated
A DevOps/SRE Engineer is sought to own daily operations, maintenance, reliability improvements, and automation of production systems in a fast-moving technology environment.
Highlights
Hands-on role managing production infrastructure with opportunities to improve reliability and work on scalable systems.
Description
Location: Hybrid (We need someone who is actively involved with the engineering team and production environment rather than operating as an occasional infrastructure consultant.)
Type: Full-time
Industry: SaaS, Fintech, iGaming Loyalty Infrastructure
About PromofyWe’re a fast-growing B2B SaaS startup powering loyalty infrastructure for iGaming platforms.
Our backend stack is modular, real-time, and event-driven, serving thousands of players across multiple tenant environments.
We move fast, ship weekly, and work in tight feedback loops with our clients.
If you’re passionate about solving real-world problems at scale and want to work on production systems that matter - you’ll fit right in.
Job DescriptionWe are looking for a hands-on DevOps / SRE Engineer to take ownership of the daily operation, maintenance, reliability, and continuous improvement of our production infrastructure.
We are looking for someone who is comfortable working with production systems, thinks systematically, and actively looks for ways to reduce operational workload through automation rather than repeatedly solving the same problems manually.
ResponsibilitiesDaily maintenance and operation of our AWS production infrastructureManaging and maintaining Kubernetes / EKSMaintaining and improving our Terraform infrastructureManaging Karpenter, autoscaling, resource allocation, and cluster capacityMaintaining and improving CI/CD pipelinesMonitoring production infrastructure and proactively identifying problemsTroubleshooting and resolving production infrastructure incidentsMaintaining and improving monitoring, logging, metrics, and alertingBuilding internal infrastructure and operational automation using PythonAutomating repetitive operational and maintenance tasksImproving infrastructure reliability, scalability, and cost efficiencyMaintaining infrastructure security, IAM, networking, and access controlsPerforming root-cause analysis and implementing permanent solutions for recurring issuesMaintaining backup and recovery proceduresProduction SystemsYou should be comfortable operating production infrastructure involving:
PostgreSQL / AWS AuroraRedisKafka / AWS MSKWe are not looking for a dedicated DBA or Kafka engineer, but you should understand how these systems operate in production and be capable of managing their infrastructure, monitoring their health and capacity, understanding connectivity and resource requirements, and troubleshooting infrastructure-related issues.
Required ExperienceStrong hands-on experience with:
AWSKubernetes / EKSTerraformKarpenterPythonCI/CDDocker / containerized workloadsLinuxAWS networking and IAMProduction monitoring, logging, and alertingTroubleshooting real production environmentsProduction-grade AWS and Kubernetes experience is essential.
Engineering & Automation MindsetWe don’t want someone who only maintains what already exists.
We are looking for someone who asks: “Why are we doing this manually, and how can we automate or eliminate it?"
You should be able to look at infrastructure as a complete system, identify unnecessary operational work, and independently propose and implement improvements.
We expect active use of AI-assisted engineering tools where they improve productivity - for infrastructure automation, troubleshooting, scripting, analysis, and operational workflows - while maintaining full understanding and ownership of changes made to production.
Production OwnershipThis role includes real production responsibility.
You should be:
Comfortable taking ownership of production issuesAvailable during agreed working hoursAble to respond quickly to critical production incidentsCapable of independently investigating infrastructure problemsResponsible for following critical incidents through to resolutionProactive about preventing recurring incidents rather than repeatedly fixing symptomsMinimal response time for critical production issues is an important requirement for this role.
What We’re Looking ForMost importantly, we’re looking for someone who is:
Hands-on — comfortable working directly with production infrastructureProactive — identifies problems without waiting for ticketsSystem-oriented — understands how infrastructure components affect each otherAutomation-driven — actively reduces manual operational workReliable — takes production ownership seriouslyInnovative — explores better tooling, automation, and approaches instead of maintaining the status quoIndependent — capable of investigating and solving infrastructure problems without constant direction
The goal of this role is simple:
Keep production reliable while continuously reducing the amount of manual work required to operate it.
If you are interested fill in information on this link -> https://forms.clickup.com/25640435/f/reffk-64292/O3C3EH19FFL0FJZ536