Principal Developer - Oncology Data and AI Systems

Jnjinnovativemedicine — United States · Posted ~18 hours ago

Lead Full-time

Skills

data science data analytics AI software development data systems Python AWS

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A principal-level developer is sought to build and advance data and AI systems supporting oncology-focused scientific and healthcare initiatives. The position sits at the intersection of data science, computational technologies, software development, and applied AI, contributing to complex systems intended to accelerate innovation in healthcare.

Highlights

Principal-level opportunity focused on advanced data and AI systems supporting oncology research and healthcare innovation, with an emphasis on meaningful scientific and technological impact.

Description

At Johnson & Johnson, we believe health is everything. Our strength in healthcare innovation empowers us to build a world where complex diseases are prevented, treated, and cured, where treatments are smarter and less invasive, and solutions are personal. Through our expertise in Innovative Medicine and MedTech, we are uniquely positioned to innovate across the full spectrum of healthcare solutions today to deliver the breakthroughs of tomorrow, and profoundly impact health for humanity. Learn more at jnj.com. As guided by Our Credo, Johnson & Johnson is responsible to our employees who work with us throughout the world. We provide an inclusive work environment where each person is considered as an individual. At Johnson & Johnson, we respect the diversity and dignity of our employees and recognize their merit. Job Function Data Analytics & Computational Sciences Job Sub Function Data Science Job Category Scientific/Technology All Job Posting Locations: Cambridge, Massachusetts, United States of America, Raritan, New Jersey, United States of America, San Diego, California, United States of America, Spring House, Pennsylvania, United States of America, Titusville, New Jersey, United States of America Job Description Our expertise in Innovative Medicine is informed and inspired by patients, whose insights fuel our science-based advancements. Visionaries like you work on teams that save lives by developing the medicines of tomorrow. Join us in developing treatments, finding cures, and pioneering the path from lab to life while championing patients every step of the way. Learn more at https://www.jnj.com/innovative-medicine Johnson and Johnson Innovative Medicine is recruiting a Principal Developer, Oncology Data & AI Systems (2 available positions) to strengthen the AI-ready data foundation that powers Oncology R&D. This role will lead efforts to modernize and standardize data capture, translate scientific and business needs into engineering requirements, and design, build, and optimize scalable data pipelines and workflows that ensure data is high-quality, well-governed, interoperable, and fit for advanced analytics and AI/ML. This role focuses on data science initiatives across Oncology R&D— including Clinical, Pre-Clinical, RWD and ‘omics platforms by enabling trusted, reusable datasets and foundational data products that accelerate downstream insights and model development. This role will be a leading data science contributor and creative problem solver with developing AI-ready data and other routinely used data applications that improve speed, reliability and impact of data-driven decision making for Oncology R&D. This position will be located in either Cambridge, MA; Spring House, PA; Titusville, NJ; Raritan, NJ; or San Diego, CA (no remote option). Key Responsibilities Partner with Oncology R&D and Data Science stakeholders to identify, prioritize, and deliver high-impact data, AI/ML, and GenAI use cases that accelerate scientific discovery, study execution, and evidence generation.Own the end-to-end design, build, and lifecycle management of Oncology R&D data products, including requirements, architecture, ETL/ELT development, documentation, and operational support.Integrate and harmonize multi-modal R&D data across biomarker labs, translational platforms, clinical trials, real-world data/evidence (), genomics/other ‘omics, and pre-clinical research systems to create trusted, reusable, AI-ready datasets.Co-develop an AI-ready data ecosystem in collaboration with Data Product Engineers, Data Scientists, Knowledge Graph Engineers, Enterprise Architecture, and IT—enabling advanced analytics, ML, and GenAI applications.Translate complex scientific, clinical, and operational needs into scalable engineering solutions, ensuring alignment with Oncology R&D priorities, enterprise data standards, and interoperability requirements.Design and optimize data pipelines for structured and unstructured data, leveraging Python, R, SQL, AWS services, and other relevant technologies to improve throughput, reliability, and maintainability.Implement robust data quality and reliability controls, including validation frameworks, automated monitoring, and KPI-driven measurement of data product performance, adoption, and business impact.Establish strong data governance foundations by maintaining data lineage, metadata, and versioning practices that support transparency, traceability, reproducibility, and compliance with internal and regulatory expectations. Embed FAIR data principles and compliance requirements into workflows Required Qualifications Bachelor’s Degree in Computer Science, Engineering, Life Sciences, or other relevant field. Advanced degree preferred5+ years of experience in data engineering, including data modeling, schema design and database architecture, preferably in the healthcare industryDemonstrated proficiency in data engineering tools such as Python, R and SQL for data processing, transformation, and automation across large-scale datasetsHands-on experience designing and operating cloud-based data platforms, preferably on AWS (e.g., Redshift, FSx, Glue, Lambda) or equivalent servicesExperience working with multiple data storage patterns, including relational, NoSQL/unstructured, and graph data technologies (e.g., knowledge graph–enabling platforms). Strong analytical and problem-solving skills, with the ability to troubleshoot complex data pipeline, quality, and performance issues in production environments.Proven capability to lead cross-functional delivery and continuous improvement initiatives with multidisciplinary, distributed teams; experience coordinating with external vendors/partners.Demonstrated strength in stakeholder management, including requirements discovery, business analysis, and planning; ability to translate conversations into clear user stories, engineering requirements, and executable delivery plans.Ability to manage multiple concurrent projects, prioritize work, exhibit organizational skills and flexibility to deliver maximum business value.Willingness to conduct periodic travel (