US Jobs US Jobs     UK Jobs UK Jobs     EU Jobs EU Jobs


Senior Scientist - GenAI Evaluation

At Johnson & Johnson,we believe health is everything.

Our strength in healthcare innovation empowers us to build aworld where complex diseases are prevented, treated, and cured,where treatments are smarter and less invasive, andsolutions are personal.Through our expertise in Innovative Medicine and MedTech, we are uniquely positioned to innovate across the full spectrum of healthcare solutions today to deliver the breakthroughs of tomorrow, and profoundly impact health for humanity.Learn more at jnj.com .

As guided by Our Credo, Johnson & Johnson is responsible to our employees who work with us throughout the world.

We provide an inclusive work environment where each person is considered as an individual.

At Johnson & Johnson, we respect the diversity and dignity of our employees and recognize their merit.

Job Function:
Data Analytics & Computational Sciences

Job Sub Function:
Data Science

Job Category:
Scientific/Technology

All Job Posting Locations:
Cornellà de Llobregat, Barcelona, Spain, Madrid, Spain

Job Description:

At J&J we are building Generative AI solutions to support pharmaceutical R&D - literature review, evidence synthesis, document Q&A, therapeutic area knowledge search, translational science workflows, and R&D decision support.

These systems need to be evaluated before teams rely on them in scientific workflows.

In pharma, a useful AI response depends on the question, user, source material, therapeutic area, and risk of error - so quality must be measurable, repeatable, traceable, and scientifically defensible.

As a Senior Scientist in our Generative AI Evaluation & Standards (EQS) function within Data Science and Digital Health, you will design, build, and run the evaluation methods we depend on to assess these systems across R&D.

You will author rubrics, curate benchmark and golden datasets, validate AI judges, and produce the quality readouts that inform release decisions.

Want to shape how a top pharmaceutical company determines whether its AI is ready for real scientific work? This is that role!

KEY RESPONSIBILITIES:

We need someone who can build the evaluation assets our teams count on, and continuously improve them based on what we learn.

Day to day, you will:


* Design, build, and maintain automated evaluation pipelines for LLM quality, RAG performance, agent reliability, safety, and scientific accuracy.


* Author evaluation rubrics and scoring criteria, curate golden and synthetic datasets with domain experts, and maintain our registry of reusable evaluation assets.


* Validate AI judges against human expert agreement and run model, prompt, retriever, and agent benchmarks that produce standardized quality readouts.


* Analyze failure patterns - hallucination, unsupported claims, weak traceability - and turn findings into actionable recommendations.


* Develop therapeutic-area-specific evaluation criteria with scientific, clinical, and regulatory partners, refining them based on real-world...




Share Job