Summary
✨ AI‑Generated
A data engineer role focused on designing data architectures, building pipelines, and managing scalable storage solutions for analytics and AI workloads.
Highlights
Data engineering role with opportunities to build modern data platforms, analytics infrastructure, and AI-ready solutions.
Description
COMPANY OVERVIEW
At ARRAY, we're not just another software services company—we're a team of dreamers, innovators, and trailblazers! From startup grit to big-tech aspirations, we're on a mission to redefine technology, put Bahrain on the global tech map, and grow into a powerhouse that inspires.
If you're ready to be part of an exciting journey, we want you on our team!
KEY RESPONSIBILITIES
Data Modelling: Design schemas and semantic layers (medallion bronze/silver/gold) that power AI-driven applications and analytics.ETL & Pipeline Architecture: Build and orchestrate ETL/ELT pipelines (batch and streaming) using Airflow and dbt, feeding cloud data warehouses and lakes.Lakehouse & Storage Architecture: Design and maintain data lake and lakehouse solutions using open table formats (Apache Iceberg, Delta Lake, or Hudi) over object storage, balancing cost, performance, and query flexibility.Database Engineering: Manage relational and NoSQL databases (PostgreSQL, MySQL, Oracle) supporting both transactional and analytical workloads, including performance tuning, replication, and migration between systems.Data Governance: Implement role-based access control, permission-aware retrieval, reconciliation, and data quality monitoring across data platforms.Cloud & Platform Integration: Deploy and operate data workloads on our Kubernetes-native platform (ArgoCD, Terraform) alongside core AWS services.
MUST-HAVE SKILLS
Bachelor's degree in Computer Science or a STEM-based subject.5+ years of experience in data or software engineering.Strong SQL and Python (or Java) skills.Hands-on experience building ETL/ELT pipelines using orchestration tools such as Airflow.Experience with cloud data warehouses and lakehouse architectures (e.g., Snowflake, Redshift, BigQuery, Apache Iceberg, or Delta Lake).Strong database fundamentals — schema design, indexing, query optimization — across relational (PostgreSQL, MySQL, Oracle) and NoSQL systems.Experience with unstructured data stores (Vector DBs, Document DBs) supporting AI/ML retrieval use cases.Working knowledge of containerized environments (Docker, Kubernetes) for deploying data workloads.CI/CD pipeline experience (GitHub Actions, GitLab CI, Jenkins, or similar).Working knowledge of Infrastructure as Code (Terraform or similar).Proficiency with version control (Git).Strong written and verbal English communication skills.
NICE-TO-HAVE SKILLS
Apache Spark for large-scale batch and distributed data processing.Exposure to streaming/CDC tools (Kafka, Debezium, or Kinesis).Experience with data catalog and metadata management tools (Glue Data Catalog, DataHub, or Amundsen).Familiarity with Power BI dashboarding and reporting.AWS, Azure, or GCP certification.Experience migrating legacy on-prem databases to cloud-native platforms.Observability stack — Prometheus, Grafana.