Senior Product Data Engineer- Data Insights (remote, Europe)

Modash โ€” Slovenia ยท Posted ~3 weeks ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

Hey, I'm Tadas. I'm hiring a Senior Product Data Engineer for Data Insights, the team that builds Modash's data products. Modash helps brands find, understand, and work with creators on Instagram, TikTok, and YouTube. More than 2,700 companies, including Stanley 1913, Sennheiser, and NordVPN, use us to run and grow their creator partnerships. Data Insights turns raw social media data into the datapoints customers buy: which brands a creator has worked with, how to contact them, where they're based, and which other creators are like them. Most of the work is taking messy public data across 200M+ profiles and making it accurate enough that a brand will act on it. In this role you'll own datapoints end to end, from the raw signal to what the customer sees, and you'll be judged on their quality. Why we're hiring At Modash, the data is the product customers pay for. Brands use our collaboration, contact, and location data to decide which creators to work with, and our largest customers take it in regular bulk deliveries through our API. We need a senior engineer who can take one of these datapoints from a rough idea to a production pipeline and keep improving it after launch. Right now that means building our own creator-location data, using LLMs to find and validate creator contact emails, and matching sponsored posts to the right brand and its parent company. You'll join a team of four data engineers who work closely with four backend engineers. You'll make important technical decisions yourself, with strong teammates to back you up, and customers will see what you build. For a feel of how we build software, read our Engineering Blog. What You'll Own Build the datapoints customers pay for You'll turn raw posts, bios, and profile data into signals: brand collaborations, contact details, creator location, related creators, and links between one creator's accounts on different platforms. Make data quality measurable Every datapoint we ship has automated quality checks. You'll decide what good looks like for your datapoint, measure it per platform against real coverage numbers, and catch regressions before customers do. Take projects from idea to production You'll shape the problem, scope the work, design the architecture, write the code, ship it, and learn from how it performs. We expect senior engineers to own the outcome, even when the problem arrives half-defined. Use LLMs where they pay off We run LLM enrichment at large scale. You'll decide where an LLM beats a rule or a simpler model, and work out what it costs across 200M+ profiles before it ships. What The Day-to-day Looks Like Most weeks you're on one project for days at a time, with some of your time going to pipeline issues and code review. A typical week might look like this: Monday. You spot-check your datapoint's output and find a quality problem. You write a short plan for fixing it and get feedback from another data engineer before you startTuesday. You build the change in PySpark and test it. It's too slow on the full dataset, so you profile the job and fix itWednesday. A pipeline failed overnight. You find the cause, ship a fix, and rerun it. In the afternoon you review a teammate's PR, then get back to your projectThursday. You compare results before and after the change on each platform. One number moves more than expected, and you dig in until you can explain itFriday. You merge, watch the first production run, and check the quality reports. You answer a question from a backend engineer about the data, then plan next week Every meeting needs a reason, and we protect time for focused work. You'll have a short standup, pair with people when it helps, and get long stretches to plan, build, and ship. Requirements What you've done before Built data systems at meaningful scale. You have solid experience as a data engineer and have worked with large, complex datasets in productionWorked deeply with Spark. PySpark, Scala, or Databricks all count. We use PySpark, but strong Spark fundamentals matter more than the exact flavour. Here is the blog post that helps explain the kind of work we doTurned messy data into signals people rely on. Entity matching, classification, deduplication, or enrichment: you've built reliable output from data that started out incomplete, inconsistent, or wrongTreated data quality as part of the job. You've defined quality checks, tracked coverage and accuracy, and handled a quality drop like a bugOwned full features. You've taken a feature through planning, scoping, architecture, implementation, release, and iterationWritten production Python and SQL. You care about code quality, system design, and maintainabilityWorked with orchestration and cloud infrastructure. Experience with Airflow or AWS Step Functions, and with services such as EMR, Glue, Athena, DynamoDB, S3, Kinesis, Lambda, or ECS, will help you get moving quicklyWorked using agentic development. You use coding agents like Cursor in your daily work, give them clear context, and review what they write as carefully as a teammate's PRWorked autonomously without working alone. You make progress with incomplete information, say what you think, ask for feedback, and help teammates do better work Curiosity about the creator economy helps too, but we'll get you up to speed. Our stack AWS, with Pulumi for infrastructure as codeSpark on EMR, mostly PySparkIceberg tables on S3, queried through Glue and AthenaPostgreSQL on AuroraAirflowGreat Expectations for data quality checksLLM models for data enrichmentDynamoDB, Kinesis, Lambda, and ECSGitHub, Notion, Linear, Cursor The interview process We move quickly and can finish the process in under a week: Intro chatTwo technical interviews: an agentic PySpark coding challenge and a system design sessionTeam fit and project presentationCulture and alignment conversation with our CEO, Avery Shrader Benefits What we offer Fully remote in Europe ๐Ÿ  Work from wherever you do your best workCompensation. Your compensation is made up of salary and stock options. As we're growing fast, the stock option package is especially significant. Annual salary range is โ‚ฌ100,000 to โ‚ฌ130,000. We hire across Europe, so the exact number depends on your location, employment type, skills, and experienceFlexible hours โฑ We care about outcomes, not when you log onUnlimited paid vacation ๐ŸŒด Take the time you need to stay rested and do great workPersonal development support ๐Ÿง  Courses, books, and conferences are on usReal ownership ๐Ÿ’ก Take meaningful customer problems from ambiguity to impactRegular offsites โœˆ๏ธ We're remote-first, but we make time to connect in person And a Little More About Us... Founded in 2018 by a high-school dropout and a Canadian (yes, we're also shocked it's going so well), Modash is building a suite of tools that help brands scale partnerships with online content creators. 2,700+ companies like Stanley 1913, Sennheiser, and NordVPN already use Modash to manage and scale their influencer marketing work. And we're just getting started. Over the coming decade, brand investment in creators will continue to boom, and Modash will be at the centre of it all. Modash is here to stay. We have 8-figures in ARR across two products, a $12M series A investment, and we are default alive. We are building a company that will still be here in 20 years; not rushing towards an exit. We're almost 100 people distributed across 20+ countries, operating with a fast, async-first culture. If you join Modash, you'll be surrounded by people who truly want to be the greatest at their craft. People who make you better. Interesting people too, who have done everything from building solar cars, to hanging out with Metallica and Bon Jovi. Come join us. Be great, do great things, create great memories, all while making a great impact. Do it.