Summary
Develop production-grade computer vision pipelines, optimize inference on embedded GPU hardware, build evaluation workflows, and deploy reliable containerized solutions for real-world environments.
Highlights
Work on production computer vision systems, optimize embedded AI deployments, collaborate in a small team, and gain direct customer impact.
Description
We are an early-stage, stealth-mode computer vision company.
Our software runs on GPU hardware installed inside customer facilities, watches live camera feeds, and turns what it sees into structured events that people act on.
It is deployed with real customers, on real sites, where being wrong is expensive and being slow is useless.
This role owns the part of the system that actually looks at pixels: detection, inference on constrained hardware, and the measurement that tells us whether any of it is trustworthy.
Tasks
Own the detection pipeline end to end: RTSP ingest, per-frame detection, tracking, and event emission
Get models running fast on embedded GPU hardware.
Quantisation, TensorRT, batching, stream decode budgets, thermal and power headroom.
Making a model accurate is half the job.
Making it accurate at 15 streams on a box that costs under 1,000 JOD is the other half
Design and operate a two-stage cascade: a cheap detector on every frame, a heavier vision-language model queried narrowly and rate-capped
Build the evaluation loop.
Per-rule precision and recall against real customer footage, tracked over time, with the numbers written down and defensible to a customer who is about to sign
Own the labelling pipeline: what gets sampled, what gets annotated, how it feeds retraining, and how we know a change made things better rather than different
Handle the ugly reality of production video: cameras that drop, streams that stall, lighting that changes at shift handover, sites that lose internet for a day
Package and ship your work as versioned containers that a non-specialist can deploy to a new site
Requirements
What we are looking for
1+ years working with computer vision in production, not only in notebooks
Strong Python.
Comfortable in PyTorch, and comfortable reading someone else's inference code and finding the bottleneck
Real experience with detection models (YOLO family, RT-DETR, or equivalent): training, fine-tuning, and evaluating them on data you collected yourself
Hands-on with NVIDIA edge hardware (Jetson Orin or similar) or a strong track record of optimising inference for constrained devices.
TensorRT, INT8 or FP16 quantisation, DeepStream or GStreamer
Video pipeline fluency: RTSP, decode, frame sampling, and what happens to all of it when a stream misbehaves
Docker and Linux as everyday tools
The instinct to measure before claiming.
We would rather hear "68 percent precision on 400 labelled frames" than "it works well"
Nice to have
Vision-language models in production (Qwen2.5-VL, InternVL, or similar), including serving and cost control
Multi-object tracking and re-identification
Experience with industrial or safety-critical deployments
Data-centric practice: knowing that a better dataset usually beats a better architecture
How we work
Hybrid from Irbid, mixing in-office and remote days
Small team, direct feedback, very short distance between a decision and its consequences
You will visit customer sites.
This is not a role you can do entirely from a desk
How to apply
Send your CV plus anything you have built.
Then answer this in a few sentences, because we read it first:
Describe a model you have run on an embedded or constrained device.
What was too slow, and what did you actually do about it?
Applications without an answer to that question are unlikely to get a response.
We reply either way.