Description
Get to know us better
CodiLime is a software and network engineering industry expert and the first-choice service partner for top global networking hardware providers, software providers and telecoms.
We create proofs-of-concept, help our clients build new products, nurture existing ones and provide services in production environments.
Our clients include both tech startups and big players in various industries and geographic locations (US, Japan, Israel, Europe).
While no longer a startup - we have 250+ people on board and have been operating since 2011 we’ve kept our people-oriented culture.
Our values are simple:
Act to deliver.Disrupt to grow.Team up to win.
Renumeration: B2B Contract 16 500 - 28 000 PLN + VAT
The project and the team
You will join the team behind a large-scale, centralized data platform built for a global consulting organization.
The platform is the shared source of company data behind several of the firm's internal products, and is used daily by consultants for company research, including support for Mergers & Acquisitions (M&A) engagements.
This is fundamentally a data engineering role combined with software engineering: you'll be designing, coding, testing, and operating production Python systems - pipelines, libraries, and services - that move, transform, and serve data at scale.
This is not a role focused on configuring tools or writing one-off queries.
You'll be building reliable, maintainable software that powers our data platform.
The goal is a unified, enterprise-grade dataset of 300M+ company records, integrated from 10+ external and internal sources.
The platform delivers firm-level and site-level data - firmographics, technographics, and hierarchical relationships (parent company, subsidiary, site) - alongside key business metrics such as revenue, CAGR, EBITDA, headcount, M&A activity, competitors, industry classification, and web traffic.
Data needs to stay accurate, well-structured, and fast to query as both the dataset and the number of consumers keep growing.
Technology stack:
Languages: Python, SQLData platform: Snowflake, dbtWorkflow orchestration: Apache Airflow (complex DAGs), running on KubernetesData processing: Apache Spark on Azure DatabricksData tooling and DBs: pandas, Polars, PyArrow, DuckDB, PySpark, PostgreSQL, RedisCloud: Azure (AKS, Blob Storage, ACR, Databricks, OpenSearch, Azure AI Search)API & services: FastAPI (REST, async), API GatewayTesting & code quality: pytest, mypy/pyright, ruff/black, sqlfluff, SonarQubeSchema validation: PydanticDependency & environment management: uv, PoetryCI/CD & infrastructure: GitHub Actions, Docker, KubernetesAI-Assisted Development: Cursor, Claude Code, ChatGPT EnterpriseFuture direction: agentic AI systems, LangChain, Azure OpenAI integration
What else you should know:
Team: Data Architecture Lead, Data Engineers, DataOps Engineers, Backend Engineer, Product Owner, collaboration with Frontend Engineers and Data Science and AI EngineersDistributed team across Europe and IndiaAgile, collaborative environment; given the platform's organization-wide impact, we're looking for a mature, proactive, results-driven approachCode quality is enforced through testing, typing, and tooling - not just code reviewWe work on multiple interesting projects at a time, so it may happen that we’ll invite you to an interview for another project if we see that your competencies and profile are well suited for it.
Your role
This is a results-driven contributor role, split roughly between hands-on engineering delivery and continuous improvement of the platform.
As a part of the project team, you will be responsible for:
Engineering & delivery (~70%)
Design, build, and maintain batch and streaming data pipelines in Python, including their orchestration, scheduling, and monitoring (Airflow)Write reusable, well-typed Python libraries and internal packages used by other engineers and analystsBuild Python services and APIs (FastAPI) that expose data to downstream applications, and integrate with third-party and internal APIsDevelop data transformation logic in SQL and dbt on Snowflake, and build data models that support fast, reliable queryingWrite unit, integration, and data-contract tests, and keep pipelines covered by automated CIProfile and optimize Python code and data processing jobs for runtime, memory, and costDeploy and operate code in the cloud using containers, infrastructure as code, and CI/CDEnforce access controls, secrets handling, and sensitive-data protections throughout the data lifecycleUsing AI coding assistants effectively while validating all generated output before it reaches production and helping build automated quality gates
Continuous improvement (~30%)
Find and fix efficiency, reliability, cost, and correctness issues in existing pipelines, refactoring toward simpler designsReplace one-off scripts and notebooks with tested, packaged, scheduled codeImplement data quality checks, validation, and monitoring that catch issues before consumers doCreate matching logic to deduplicate and connect entities across multiple data sourcesDocument data processes and system architecture, and maintain project documentation
Do we have a match?
As a Data Engineer, you must meet the following criteria:
Strong Python experience: data structures, typing, error handling, generators/iterators, context managers, and the standard librarySoftware engineering fundamentals: modular design, dependency management, packaging, and API designTesting discipline: pytest, fixtures, mocking, and writing code that's testable by constructionHands-on experience building and operating ETL/ELT pipelines in production, not just scripts or notebooksStrong experience with Snowflake and dbtExperience with Apache Airflow or similar code-based orchestration toolsSolid working knowledge of SQL and data modeling, sufficient to design robust database schemas and query them effectivelyExperience with Docker, Kubernetes, and CI/CD practicesDebugging and profiling skills; able to reason about performance, concurrency, and memory in PythonExperience with at least one public cloud (AWS or Azure)Experience with version control systems (Git)Able to explain technical trade-offs to both technical and business audiencesExperience using AI coding assistants such as Claude Code, Cursor, or similar on a daily basisStrong communication skills and good knowledge of English (minimum C1 level)
Beyond the criteria above, we would appreciate the following nice-to-haves:
Experience with Apache Spark, ideally on DatabricksExperience with Pydantic or similar schema-validation librariesPython web/API frameworks (FastAPI, Flask)Async Python, multiprocessing, or other concurrency patternsExperience with Azure AI Search or AWS OpenSearchA second language: Go, Rust, Scala, or TypeScriptFamiliarity with LLMs, Azure OpenAI, or agentic AI systems
More reasons to join us
Flexible working hours and approach to work: fully remotely, in the office or hybridProfessional growth supported by internal training sessions and a training budgetSolid onboarding with a hands-on approach to give you an easy startA great atmosphere among professionals who are passionate about their workThe ability to change the project you work on