Summary
✨ AI‑Generated
A hands-on AI/ML engineering role focused on designing, implementing, and operating high-throughput, real-time AI and computer vision systems in a manufacturing environment. You will build distributed Python services, containerized workloads, and Kubernetes deployments while optimizing concurrency, latency, resource usage, networking, observability, and automated remediation. The role suits someone who enjoys investigating problems, developing recommendations, and turning ML concepts into reliable production systems.
Highlights
Hands-on AI/ML engineering role centered on real-time computer vision and high-throughput production systems. Provides opportunities to solve open-ended problems, develop recommendations, and work across backend engineering, distributed systems, observability, scalability, and performance.
Description
Strong Technical acument:
must have strong Pyhton coding, and ML
Curosr->IDE->CODEX
HEAVY HANDS ON ROLE:
find a problem->come up w/ recommendations, not just problems
CURIOSITY DRIVEN
ML MODELING EXPERIENCE
This role focuses on the design, implementation, and operation of high?throughput, real?time AI and computer vision systems in a manufacturing environment.
Work centers on distributed Python services, real?time data/vision streaming, performance engineering, observability, and automated remediation.
The position collaborates closely with cross?functional engineering teams to help deliver reliable, secure, and scalable production systems.
Key Responsibilities
Backend & Platform Engineering
Design and implement backend services using Python (FastAPI/Uvicorn/asyncio) for high?concurrency, low?latency workloads.Configure and maintain containerized services (Docker) deployed on Kubernetes using Helm.Tune resource limits/requests, autoscaling, and networking for stateful and streaming workloads.Implement secure deployment patterns using enterprise registries and CI/CD pipelines (e.g., GitHub Actions).Real?Time Data & Vision Streaming
Build and maintain ingestion services for real?time data streams, including industrial video protocols (e.g., RTSP).Integrate multi?camera and multi?stream inputs into Python backends for monitoring, diagnostics, and analytics.Optimize streaming, buffering, and processing behavior to meet strict per?frame or per?request latency targets.Numerical & Statistical Computing
Implement and optimize matrix?heavy and numerical workloads using NumPy, SciPy, and OpenCV.Translate statistical methods (e.g., distance metrics, anomaly scores, temporal smoothing, thresholds) into robust, production?grade code.Design and execute calculations that distinguish normal process variation from true degradation or drift in data and signals.Automated Remediation & Diagnostics
Implement backend decision logic that separates software?correctable issues (e.g., configuration, firmware, calibration) from physical or environmental issues (e.g., optics, vibration, misalignment).Develop automation scripts and workflows that execute programmatic fixes (e.g., rolling back versions, refreshing baselines, updating configurations).Generate precise, localized diagnostic payloads for maintenance and operations teams when physical remediation is required.Performance, Reliability & Resilience
Design and run realistic load, stress, and soak tests (e.g., with k6 or Locust) that include high?FPS streams, bursty workloads, and multi?session interactions.Implement and tune rate limiting, request throttling, queuing, and back?pressure strategies to protect downstream services.Apply resiliency patterns such as retries, circuit breakers, fallbacks, and state?recovery for failed or partial executions.Observability & Monitoring
Configure and extend observability stacks (e.g., OpenTelemetry or similar) for metrics, logs, and distributed tracing across services.Enable visualization and inspection of complex execution graphs and data flows to support diagnostics and performance tuning.Contribute to definition and tracking of SLOs/SLIs around latency, throughput, error rates, and system health.APIs & Integration
Expose clean, stable APIs and schemas (e.g., JSON payloads) for downstream consumers, including UI applications and other backend services.Provide deterministic structures for diagnostic outputs (e.g., spatial heatmaps, ROI metrics, temporal trends) to support operator?facing tools.Collaborate with frontend and integration teams to ensure APIs are well?documented, versioned, and testable.
Required Skills & Experience
Core Experience
6+ years in one or more of the following: Infrastructure/SRE, Core Platform Engineering, MLOps, Vision/Edge Platform Engineering, or high?throughput data/backend systems.Experience delivering and supporting production systems with strict performance, reliability, and availability requirements.Python & Backend
Expert?level production Python, including:Asynchronous programming with asyncio.Building APIs with FastAPI (or similar ASGI frameworks).Tuning Uvicorn/gunicorn or equivalent servers for high?concurrency workloads.Strong understanding of memory management, concurrency, and profiling in Python.Containers, Kubernetes & CI/CD
Deep experience with Docker (image design, security, performance, multi?stage builds).Strong Kubernetes knowledge: workloads, services, ingress, stateful workloads, autoscaling, and resource tuning.Practical experience with Helm for templating and deploying applications.Hands?on experience with CI/CD (e.g., GitHub Actions or similar) for automated builds, tests, deployments, and rollbacks.Streaming & Vision
Hands?on experience ingesting and processing industrial camera/video streams (e.g., RTSP or similar protocols).Experience working with real?time streaming or event?driven architectures in production environments.Strong proficiency with OpenCV (or equivalent) and image/matrix operations.Numerical & Statistical Skills
Advanced proficiency with NumPy and SciPy for vectorized, high?performance numerical computing.Experience implementing statistical techniques (e.g., distance measures, anomaly detection, trend analysis) in production pipelines.Ability to profile and optimize algorithmic hot paths and matrix operations.Performance, Testing & Observability
Proven use of load and stress testing tools (k6, Locust, or similar) for data?heavy or streaming APIs.Experience with observability tools and concepts, including metrics, logging, and distributed tracing (OpenTelemetry or equivalent).Demonstrated ability to analyze performance bottlenecks and implement targeted improvements.Collaboration & Ways of Working
Experience working as part of a cross?functional engineering team within an established architecture and technical direction.Strong communication skills for documenting designs, APIs, diagnostics, and test results.Ability to work in an iterative environment with clear deliverables, reviews, and handoffs.
Preferred Qualifications
Experience with industrial IoT, manufacturing environments, or other edge/plant?floor systems (cameras, sensors, PLCs, etc.).Background in automated rollback, closed?loop control, or other remediation/alerting frameworks.Familiarity with modern AI/ML or LLM/agent systems from an infrastructure or performance perspective (e.g., handling tool?calling, long?running workflows, or token?streaming patterns).