Senior Data & Python Software Engineer

Ceartasdmca1 โ€” Germany ยท Posted ~1 day ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

At Ceartas, we lead the way in AI-powered brand protection, copyright law, and digital security, safeguarding the integrity of content creators, brands, and enterprises worldwide. As we scale rapidly, we're looking for a Data & Python Software engineer to drive innovation in our data pipelines and crawling technologies. In this pivotal role, you'll collaborate with our CTO and Head of Engineering, steering our Data Engineering Team toward developing groundbreaking solutions for digital security challenges. Build Scalable Web Data Extraction Pipelines: Design and develop web scraping systems that support large-scale web data extraction and brand protection workflows. Ensure that data moves reliably from collection through processing to storage while maintaining performance, resilience, and operational stability at scale. Ensure Data Quality and Governance: Own data validation, consistency, and governance across ingestion, storage, and serving layers. Establish clear standards for schema design, transformation logic, and monitoring to guarantee trustworthy, production-grade datasets that can be reliably consumed across the organization. Optimize Performance and Reliability: Continuously improve scraping system efficiency through performance tuning, cost optimization, and architectural enhancements. Implement logging, metrics, and tracing to monitor production systems, diagnose issues quickly, and maintain high reliability under growing workloads. Responsibilities: Design, build, and maintain high-performance web scraping systems as well backend services and data pipelines supporting web data extraction and brand protection use casesImplement and maintain scraping focused APIs and other data services that power internal products and external integrationsBuild reliable ingestion, processing, and storage workflows for large-scale web dataHandle cleaning of web data and ensure data quality, validation, and governance across ingestion, storage, and serving layersOptimize scraping systems for performance, scalability, reliability, and cost efficiencyMonitor, debug, and improve scraping system reliability using observability tools (logging, metrics, tracing)Collaborate closely with product and engineering teams to deliver features from design through full end-to-end production deploymentTake independent ownership of systems in production, including maintenance,iteration and performance management Core Technical Requirements: Experience with web scrapingStrong SQL skillsStrong Python experienceExperience with PostgreSQL or similar relational databasesExperience designing and building scalable APIs and backend services (e.g. FastAPI, Django, or similar frameworks)Experience designing efficient, scalable data models and database schemasHands-on experience deploying and operating systems in the cloud (AWS, GCP, or Azure)Experience working with Docker and containerized environments Preferred Technical Requirements: Experience with workflow orchestration tools such as AirflowExperience with browser-based automation tools (Playwright, Selenium, or similar)Experience with DBT or analytics-focused data transformation workflowsExperience building or operating high-concurrency systems and task queuesExperience designing and deploying cloud-native workflows on AWSFamiliarity with CI/CD pipelines and production deployment practicesExperience working in a high-growth, early-stage startup environmentExperience - University education in a technical field such as Computer Science, Engineering or similar. Masters level preferred. 4+ years ( or 2 year+ in a early stage startup)