DevOps/SRE
Soliduslabs — United States · Posted ~1 month ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
About Solidus Labs
At Solidus, we are shaping the financial markets of tomorrow by providing cutting-edge trade surveillance technology that protects investors, enhances transparency, and ensures regulatory compliance across traditional assets, prediction markets, and crypto.
With over 20 years of experience in developing Wall Street-grade FinTech, our team delivers innovative solutions that financial institutions and regulators worldwide rely on to detect, investigate, and report market manipulation, financial crime, and fraud.
Headquartered in NYC, with offices in Singapore, Tel Aviv, and London, we safeguard millions of retail and institutional entities globally, monitoring over a trillion events each day.
The Role
We are seeking an experienced New York-based DevOps / Site Reliability Engineer to join our DevOps team and own the reliability, stability, and operational support of our production systems.
This role focuses on production ownership, monitoring, incident response, and on-call support, providing critical coverage.
You will work with a modern cloud-native stack and play a key role in keeping systems highly available, secure, and performant.
Day-to-Day Responsibilities
Own the reliability, availability, and performance of our production environments.Operate production Kubernetes (EKS), including cluster upgrades and Helm deployments.Manage scaling and capacity using KEDA, Karpenter, and HPA for resource optimization.Manage AWS Cloud environments including EC2, Lambda, AWS Batch, Elasticache, RDS, and more.
Evolve infrastructure as code using Terraform and Helm with security best practices.Support GitLab CI/CD pipelines, resolving deployment issues and improving stability.Design observability systems using Prometheus, Grafana, and EFK to reduce alert fatigue.Solve networking issues involving TLS, Load Balancing, VPCs, NAT, and VPN.Support compliance initiatives and respond to security-related incidents.Leverage AI-powered tools as a standard part of your workflow for automation and productivity.Lead incident response end-to-end, including troubleshooting, mitigation, and resolution.Perform deep-dive RCA to drive long-term corrective and preventive actions.Participate in on-call rotations to provide consistent operational coverage.
Requirements:
Minimum Qualifications
3+ years of hands-on DevOps / SRE experienceStrong production experience with Docker and KubernetesSolid knowledge of AWS (EKS, EC2, Organizations, RDS, S3, CloudWatch, Lambda, DynamoDB)Experience with monitoring, logging, and alerting systemsProficiency with Terraform, Helm, and GitLab CI (or similar)Strong troubleshooting skills across infrastructure, CI/CD, and networkingScripting experience with Bash and PythonWillingness to participate in on-call rotationsFamiliarity with pub/sub systems (SQS, Kafka, or similar)
Nice to Have
Experience with Redis, Airflow, Databricks, Spark/EMRGitOps workflows and advanced Git usageExperience supporting databases such as Postgres, Snowflake, or ClickHouse
Why Join Us?
Join a team where you’ll own and improve the reliability of critical production systems end to end, with real autonomy and impact, directly supporting premier clients globally.
You’ll work on a modern, cloud-native stack operating at scale, tackling meaningful performance and resilience challenges.
And you’ll do it alongside a highly collaborative, global DevOps and R&D team—sharing standards, tooling, and operational expertise across regions.
We have 88,191 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 88,191 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume