Case study 2: Shaping AI-enabled scientific workflows at scale

Case study 2: Shaping AI-enabled scientific workflows at scale

01 Summary


A six-month engagement translating an early LLM product vision into a tangible and evidence-based direction for a scientific hypothesis-review platform.
page icon
Product + AI Strategy
 
page icon
Product, Science, Engineering + Leadership
page icon
Design sprint, vision prototyping, MVP definition and scientist validation
page icon
Human-in-the-loop scientific workflow and strategic product direction

Voyager was a six-month initiative exploring how an LLM-enabled workflow could support scientists reviewing potential biological hypotheses between Initiation and Hit Identification. The existing process required substantial specialist effort, relied on fragmented tools and captured critical decisions inconsistently. I helped translate the initial AI opportunity into a tangible product direction through workflow modelling, sprint facilitation, vision prototyping, MVP definition and triangulated user research.
warning_1 Challenge
💡
Thousands of hours of specialist review effort to arrive at Go / No go / Hold decisions were required to meet pipeline goals, with work fragmented across Slack, Hex dashboards, Google forms, Google sheets, Monday boards, 3rd party literature applications (DepMap, Pub Med), Google slides and manual decision records.
help_center Key questions
  • How could agent-generated content reduce manual review work?
  • How could decisions be captured in-app rather than in spreadsheets?
  • How could review context support future agentic workflows?
conversion Approach
Current state workflow mapping, product discovery, design sprint facilitation, prototype creation, MVP definition and validation.
 
inventory Output
  • HITL workflow
  • Trust framework
  • MVP scope
  • Capability themes

02 Context


The Nomination workflow sits between early hypothesis Initiation and progression towards Hit Identification. It requires scientists to review biological and chemical evidence, assess relevance and risk, validate information and decide whether a potential programme should progress to the interim board (Scientific review committee) before proceeding to Hit Identification. (Go / No Go / Hold decisions).
Scientific hypothesis review is a high-effort, high-judgement workflow involving evidence gathering, biological assessment, feasibility checks, confidence evaluation and programme-level decision-making.
The previous process relied on a mixture of Hex, Looker, Google Sheets, Monday, PubMed, DepMap, Enhanced Chat and other internal and external sources. Reviewers repeatedly cross-checked outputs and copied information between systems.
Voyager was conceived as an intelligent, LLM-based interactive application to support scientists completing manual hypothesis nomination reviews.
The longer-term vision was to streamline the Nomination of potential Targets to Hit ID at the Initiation process, reduce the time required for each review and capture scientific decisions in a structured form that the system could learn from.
A phased delivery approach was planned, beginning with an MVP focused on the core Biology review and decision-capture workflow. The Chemistry review would follow at a later stage, and the application would lay the foundations for Agentic workflows across the Stage Gates.
 
🌟 The platform was intended to do more than speed up the manual Human in the Loop [HITL] review process. It needed to establish the decision-capture foundation required for future agentic workflows.


03 Challenge


One of the most significant constraints on achieving the organisation’s 2025 pipeline goals (50 human-approved nominations ) was the number of manual hours Scientists (Biologists and Chemists) were required to spend approving or rejecting (Go / No-Go / Hold) the Nomination of a single Targets moving between the Initiation and Hit ID stage gates.
Scientists were spending a significant amount of time [9.7 hours at the latest review] searching, comparing, copying and validating information across Hex, Looker, Google Forms, Sheets, PubMed, DepMap, Enhanced Chat and other tools.
The challenge was not simply to generate scientific content via an agent. It was to reduce avoidable manual effort while preserving human judgement, improving decision traceability and creating trustworthy contextual data for future agent-enabled workflows, leading to Cost and Time savings for the business.
 
A complex web of external systems and tools……..

notion image
notion image
notion image
notion image
notion image
 

Fragmented workflows
Evidence was distributed across multiple applications creating repeated searching, copying and validation.
Inconsistent review outcomes
Prior knowledge and reviewer experience influenced how evidence was interpreted and whether a nomination progressed.
 
Weak decision traceability
Reasons for stopping, starting, pausing or advancing programmes were not captured consistently.
Late feasibility risks
Chemistry constraints could surface only after substantial Biology effort had already been invested.
 
 

04 Role and contribution

I joined after the original opportunity and early vision had been identified. I acted as the AI Product Strategy Lead for the initiative, translating an early AI opportunity into a tangible product direction through workflow discovery, cross-functional facilitation, vision prototyping, MVP definition and user validation.

strategy Discover

Understand the problem and scientific workflow

notion image
→ Mapped the as-is Nomination workflow, tools and handoffs
→ Identified pain points, workarounds and decision capture gaps
→ Led discovery interviews with Scientists across Biology and Chemistry
→ Explored Trust, Usability, Traceability and Adoption barriers
→ Gathered evidence from current processes and existing tools

strategy_2 Define

Frame the opportunity and strategic direction

→ Synthesised research into prioritised problem and opportunity themes
→ Defined core user and business needs
→ Reframed the opportunity beyond simple content generation
→ Identified immediate value and longer term agentic foundations
→ Established Design sprint goals and key HMW (How Might We) questions


design_services Develop
Explore, prototype and shape the solution

→ Facilitated the cross-functional Design sprint
→ Guided ideation, solution sketching and concept selection
→ Translated the product vision into a tangible Figma prototype
→ Designed the Human in the loop [HITL] review experience
→ Explored generated content, citations, risk assessment and decision capture
 
fact_check Deliver and validate
Define a bounded MVP and test the direction

→ Co-authored the MVP scope, capabilities and success metrics
→ Co-authored the PRD
→ Separated immediate delivery from longer term agentic ambition
→ Led moderated usability testing and post session survey research
→ Synthesised triangulated findings into strategic capability themes and product recommendations

05 Objectives

The MVP was designed to create measurable near-term value while establishing the foundations for future agent-enabled scientific workflows
 
notion image
 
01 Reduce review effort
Reduce time spent on manual Biology nomination reviews.
02 Support pipeline goals
Improve review efficiency so throughput between Initiation and Hit could increase without increasing staff headcount
03 Standardise decision capture
Capture Scientific decisions, risk ratings and rationale more consistently within the workflow
04 Test generated content value
Reduce time spent on manual Biology nomination reviews.
05 Establish an Agentic foundation
Create structured contextual data that could support future agent-enabled workflows
 

06 Product strategy

Strategic principle: create immediate workflow value while capturing the contextual decision data required for more capable agentic workflows later.

07 MVP hypothesis and success metrics

MVP hypothesis
If Voyager could generate a credible first draft of the Biology review and allow scientists to validate, edit and capture their decision within the same workflow, it could reduce manual effort while creating contextual decision data for future agentic capabilities.

100%

Target adoption of application among 1st cohort of Biologist Nomination reviewers.
 

75%

Target retention of Agent-generated module content.
 

50%

Time spent on modules with agent-generated content is reduced
 

08 Approach

09 MVP definition

In scope:

  • Biology review
  • Generated modules
  • Editing and citations
  • LLM judge
  • Decision capture
  • Embedded data access
August 2025

  • Agentic module proof of concept and prototype testing
Out of scope:

  • Chemistry review
  • Advanced assignment and prioritisation
  • Leadership approval routing
  • Automated cross-module updates
  • Full backlog management
October 2025

  • Generated module content and UX improvement
Enabled later:

  • Context-aware agents
  • Connected regeneration
  • Decision learning loops
  • Adaptive workflow automation

September 2025

  • Core decision capture in Voyager
 

10 Research and triangulation

7

Discovery interviews

5

Usability testing sessions

1

Post-session survey

9

Detailed research themes
Discovery interviews
Explored current tools, workflow stages, friction, confidence, workarounds, feedback, traceability and ideal future-state needs.
Usability testing
Tested starting a review, evaluating generated content, adding sources, assigning risk, saving, submitting and reopening a review.
Survey
Measured adoption experience, time and effort, trust, external validation, error handling, traceability and future-state expectations.
Synthesis
POV statements, pain points, opportunities and user needs were tagged and consolidated into shared themes.

11 Survey signal

How would you describe your level of Trust in Agent-generated content?

25%

High
 

50%

Medium
 

25%

Low
 
 
 

 

12 Triangulated themes

The combined evidence showed that the problem extended beyond content generation. Scientists needed more actionable synthesis, stronger provenance, clearer decision policy, better governance, integrated evidence, experimental continuity, usable interactions, reliable performance and sustained change support.
1. From organised information to actionable decision support

Outputs were neatly assembled but remained too high-level. Scientists still had to reconcile conflicting evidence, decide what mattered and translate the result into experiments.
Underlying research themes
Actionable synthesis gap · Biology-to-experiment bridge
Product implication
Create decision-ready briefs, reconciled recommendations and suggested assays, models and experimental plans.
2. From opaque outputs to inspectable evidence

Scientists routinely double-checked outputs. Trust depended on citations, explainable scoring, model fit, confidence, versions and provenance.
Underlying research themes
Trust and transparency · Performance and reliability
Product implication
Expose per-claim citations, confidence, model and data versions, snapshot dates, change history and correction routes.
3. From informal judgement to explicit decision policy

Leadership appetite, strategic fit, competition thresholds and stage-gate ownership were not sufficiently explicit or timely.
Underlying research themes
Decision policy and risk policy · Decision engine and governance
Product implication
Encode criteria, RACI, Go / Consider / No-go logic, early feasibility gates, override rules and reopening triggers.
4. From fragmented tools to a decision-centred workflow

Reviewers moved repeatedly between scientific tools and spreadsheets. They needed evidence embedded in the active hypothesis context.
Underlying research themes
Decision-centric data integration · UI and interaction clarity
Product implication
Bring relevant sources, explanations, citations and actions into one decision workspace without replacing every underlying tool.
 
5. From product launch to operating-model change

Processes and interpretation guidance changed over time, while teams relied heavily on a small number of experienced people.
Underlying research themes
Onboarding, communication and change management
Product implication
Provide role-based guidance, current policy, named owners, change history and in-product support.

13 What product discovery changed

The research expanded the opportunity beyond content generation. Scientists valued a credible first draft, but efficiency alone would not make the workflow trustworthy or adoptable.
From: Initial assumption
To: Evidence-led direction
→ Generate useful content
→ Generate decision-ready synthesis
→ Reduce manual drafting
→ Reduce searching, reconciliation and rework
→ Display confidence
→ Make evidence, provenance and model fit inspectable
→ Capture final answers
→ Capture rationale, edits and expert feedback
→ Create a new application
→ Create a decision-centred layer across existing sources
→ Launch usable software
→ Establish governance, policy and change support

14 Outcome and strategic value

The work converted an early LLM opportunity into a defined and testable product direction grounded in scientific workflow evidence. It established the MVP scope, interaction model, success criteria and evaluation approach; clarified where agent-generated content could create immediate value; and identified the broader trust, governance, integration and operating-model capabilities required for the workflow to scale.
notion image
Product direction
A bounded MVP, explicit scope and phased delivery plan.
notion image
User evidence
Validated needs, trust conditions and adoption risks.
notion image
Strategic direction
A shift towards accountable, decision-centred scientific AI.

15 Key artefacts

Artefact 1: Pipeline and application scope
Artefact 1: Pipeline and application scope
 
Artefact 2: User journey map
Artefact 2: User journey map
 
Artefact 3: Defining the MVP slice
Artefact 3: Defining the MVP slice
 
Artefact 4: PRD
Artefact 4: PRD
 
Artefact 5: MVP prototype
Artefact 5: MVP prototype
 
Artefact 6: User research plan
Artefact 6: User research plan
 
Artefact 7: Discovery playback
Artefact 7: Discovery playback
 
Strategic direction prototype
 

16 What this demonstrates

notion image
Translating AI ambition into product strategy
notion image
Structuring complex scientific workflows

notion image
Defining accountable HITL systems
notion image
Aligning Product, Science, AI and Engineering

notion image
Triangulating qualitative and quantitative evidence
notion image
Balancing near-term value with platform ambition