Data Engineer

Big Wave Digital — United States · Posted ~3 hours ago

Mid Full-time Remote

Skills

Python SQL data pipelines LLM extraction machine learning data processing LLM

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An early-stage technology company is seeking a Data Engineer to build production data pipelines and transform complex documents into reliable structured data. The role involves Python, SQL, AI-driven extraction, and end-to-end ownership.

Highlights

Remote role with strong compensation potential and ownership of important data infrastructure supporting advanced AI applications.

Description

Mike Ross had a photographic memory. You'll need something rarer. On television, the paperwork always turns up in the right folder, perfectly legible, exactly when the hero needs it. In real life, the paperwork is a scanned PDF from a county courthouse, skewed by three degrees, stamped twice and annotated in someone's handwriting. Somebody has to turn that page into clean, reliable, queryable data. We think that person might be you. We're recruiting for an early-stage legal technology company building AI products for litigators. Everything those products do rests on the court data layer underneath: what goes in, how cleanly it's structured, and how far it can be trusted. It's one of the most important hires the company is making right now, and they're ready to move quickly for the right person. The job, in plain English You will own the caselaw and docket data layer, start to finish. No handing your work to someone else's roadmap; this one is yours. Build production pipelines pulling from federal and state court systems, e-filing platforms and published opinions, at the scale of millions of recordsTurn PDFs, scans and uneven XML or HTML into clean, structured, queryable dataDesign LLM-assisted extraction workflows and, just as importantly, prove they're accurateWork with full-stack engineers so your data powers the product's analytical and AI featuresRun statistical analysis, tune SQL, and use RAG and semantic search to get more from the dataLook after a cloud data stack within enterprise security and compliance guardrails Your three big technical skills Python and SQL for production pipelines. Your pipelines run unattended, recover gracefully, and don't wake anyone at 3am. Your SQL holds up under scrutiny.Document extraction. PDF parsing, OCR and semi-structured text. Court data rarely arrives as neat rows, and you know it.Applied LLM work. RAG, semantic search and LLM-assisted extraction, with a healthy scepticism about model output. Vector database experience is a bonus. The legal and data experience that gets you noticed Direct experience with legal or court data is strongly preferred, and it saves months of ramp. Dockets, court records, opinions and federal filing systems all count, as does any work with regulatory or government data. You don't need a law degree. You do need scar tissue. You know the same judge can appear under a dozen spellings, that case captions change mid-case, and that a field which looks reliable often isn't. On the data side, you've spent your career with large, dirty datasets and you're quick at cleaning them and finding the signal. What this role is Real ownership of the layer the whole product depends onHands-on, build-and-ship engineering in a small teamA place where domain depth is rewarded and messy data is the interesting partFully autonomous work: you set your own direction What this role is not Not BI or analytics engineering on a tidy, pre-built warehouseNot a prompt-writing or research role; LLMs are a tool inside your pipeline, not the jobNot for anyone who treats LLM extraction as plug-and-play, or who can't review the code an AI assistant writesNot a junior role: 4 to 8 years of production pipeline experienceNot a managed role with a backlog handed to you each morningNot a law degree role: legal qualifications aren't expectedNot open to candidates outside the US, or without Eastern Time overlap The practical bits Location: Remote within the US, East Coast preferred. At least four hours of daily overlap with Eastern Time, and quarterly in-person onsites.Compensation: Up to $220K base plus meaningful equityVisas: Open to visa transfersTeam: More than one hire plannedEducation: Bachelor's in computer science, data science or a related fieldProcess: Three steps, including a live SQL and pipeline design session and a technical deep dive on pipelines you've built Still reading? If a badly scanned filing strikes you as a puzzle rather than a problem, you're exactly who we're looking for. Send us your CV and tell us about the messiest dataset you've ever tamed. The title no longer says "East Coast preferred", because option 1 had dropped it. It's still in the practical bits, so candidates see it before they apply.