Systems Engineer / Site Reliability Engineer

Sharpatoms — United States · Posted ~3 hours ago

Mid Contract Remote

Skills

Linux administration Ubuntu Debian Systems engineering Site Reliability Engineering Infrastructure automation Datacenter operations Monitoring Incident triage Hardware troubleshooting RMA coordination Linux Datacenter networks

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Work remotely on large-scale Linux infrastructure supporting datacenter networks and shared services. You will independently handle operational issues, troubleshoot hardware and dashboards, coordinate repairs, escalate incidents, and support infrastructure deployment and automation. The role values autonomy and sound technical judgment rather than extensive seniority.

Highlights

Fully remote US role with a 12-month engagement, autonomous operational responsibility, large-scale Linux infrastructure work, automation, and exposure to datacenter environments.

Description

Job Title: Systems Engineer / Site Reliability Engineer Number of Roles: 1 Location: Remote - US Rate: DOE Duration: 12 months Responsibilities We are looking for someone who can independently handle day-to-day operational tasks, such as triaging hardware errors, coordinating RMAs with DCOPS, monitoring dashboards, addressing anything that comes up as "issue," and escalating that to the appropriate stakeholders when necessary. To better understand the workflow, the candidate may also participate in the site build-out process, where tasks are already automated and repetitive. Significant seniority is not required, but the candidate should be able to work autonomously and make sound technical decisions. Responsible for deploying, automating, and supporting large-scale Linux infrastructure supporting datacenter facility network and cross-functional shared services environments. Qualifications Required: Minimum 5+ years of experience administering Ubuntu/Debian-based servers in a production environment. Proven experience troubleshooting, repairing, and upgrading diverse hardware infrastructure. Strong proficiency with hardware diagnostic utilities, BIOS/UEFI configurations, and reading system/POST error logs. Knowledge of PXE boot, DCIM, Monitoring system such as Prometheus/Librenms/Grafana, Linux are required. Holding an LFCS (Linux Foundation Certified System Administrator) or Ubuntu Certified Professional credential is a strong advantage. Containerization & Orchestration: Experience with Kubernetes (K8s) is preferred.