Summary
✨ AI‑Generated
Help build and operate a platform layer for large-scale AI and machine learning workloads. You will work across Kubernetes, workload scheduling, container orchestration, MLOps, and model-serving infrastructure, while automating services, troubleshooting production issues, and collaborating with infrastructure and customer-facing engineering teams.
Highlights
Hands-on platform engineering role supporting demanding AI, machine learning, and high-performance computing workloads. Offers exposure to GPU infrastructure, Kubernetes, workload scheduling, MLOps, automation, and customer-facing technical operations.
Description
AI Platform Engineer L2 (GPUaaS – AI Neocloud)
📍 Sydney | Hybrid
About Sharon AI
Sharon AI is building the infrastructure powering the next generation of artificial intelligence.
Operating across AI infrastructure, high-performance compute, cloud platforms and large-scale technology environments, Sharon AI delivers scalable, secure and reliable infrastructure for demanding AI, ML and HPC workloads.
The Role
As an AI Platform Engineer L2, you'll help build, operate and support the platform layer powering Sharon AI's GPU-as-a-Service (GPUaaS) offering.
You'll work across Kubernetes, Slurm, container orchestration, MLOps tooling and model serving infrastructure to enable customers to train and run AI/ML workloads reliably and efficiently on Sharon AI's neocloud platform.
Reporting to the Head of Operations, you'll work closely with Network Engineering, Infrastructure and customer-facing teams to implement, automate and troubleshoot the platform services sitting above Sharon AI's underlying GPU and network fabric.
This is a hands-on opportunity for a platform, DevOps, MLOps or SRE engineer looking to deepen their expertise in GPU infrastructure and AI-native platform operations.
Key Responsibilities
Build and operate the AI platform layer, including Kubernetes, Slurm and container orchestration, supporting Sharon AI's GPU infrastructureDevelop and maintain CI/CD pipelines for model training, fine-tuning and inference workloadsImplement and support MLOps tooling for experiment tracking, model registry and deploymentConfigure and manage multi-tenant GPU resource scheduling and quota management across customer workloadsSupport model serving infrastructure for training and inference, ensuring reliability and performanceBuild monitoring, logging and alerting to track platform health, GPU utilisation and workload performanceAutomate platform provisioning and configuration using Infrastructure-as-Code tools such as Terraform and AnsibleCollaborate with Network Engineering and Infrastructure teams to ensure the platform layer aligns with the underlying InfiniBand/RDMA fabricTroubleshoot platform-level issues affecting customer AI/ML workloadsParticipate in on-call rotations and incident response for platform-related issuesContribute to platform documentation, runbooks and the internal knowledge baseSupport customer onboarding onto the GPUaaS platform, including workload configuration and troubleshooting
Skills & Experience
2–4 years' experience in platform engineering, DevOps, MLOps or SRE, ideally supporting GPU or AI/ML workloadsBachelor's degree in Computer Science or a related fieldHands-on production experience with KubernetesExperience with CI/CD and Infrastructure-as-CodeSolid understanding of Kubernetes and GPU scheduling frameworks, including Slurm, Kubernetes device plugins and NVIDIA GPU OperatorExperience with MLOps tooling and ML pipeline orchestrationProficiency in scripting and automation using Python and BashWorking knowledge of GPU infrastructure and distributed training concepts, including NCCL and data/model parallelismExperience with observability tools such as Prometheus and GrafanaUnderstanding of Linux systems administration and networking fundamentalsStrong troubleshooting and problem-solving skillsStrong communication and collaboration skills, with the ability to work across Infrastructure, Network Engineering and customer-facing teamsExposure to GPU-based infrastructure or high-performance computing environments
Experience with Slurm, NVIDIA GPU Operator, NCCL and distributed training frameworks such as PyTorch or TensorFlow is advantageous, as is experience with MLOps platforms including MLflow, Kubeflow or Ray.
Kubernetes certifications such as CKA/CKAD, cloud certifications across AWS/GCP/Azure, and exposure to InfiniBand/RDMA networking concepts are also advantageous.
Why Join Sharon AI
Help build and operate the platform powering a growing GPU-as-a-Service and AI neocloud businessWork hands-on with GPU infrastructure, Kubernetes, Slurm and AI-native platform technologiesDevelop deeper expertise across AI/HPC infrastructure and distributed workloadsWork closely with Network Engineering and Infrastructure teams across the underlying GPU and network fabricHelp enable customers to reliably train, fine-tune and run AI/ML workloads at scaleContribute to automation, observability and platform capabilities in a fast-moving environmentJoin a highly technical and ambitious team operating at the forefront of AI infrastructure
Our Values
Integrity | Innovation | Collaboration | Wellbeing | Inclusion