AI4HM

A protocol template for observational studies that develop, validate, or evaluate AI/ML-based models, algorithms, and tools in health and medicine.

Version 0.3 License: Freely available

What it is

Structure where you need it, flexibility where the science demands it.

Protocols are living documents that describe the intended conduct of a study. Having a pre-specified protocol is scientific best practice, and is commonly expected by data stewards, ethics boards, journals, and other research stakeholders.

Existing templates, such as the EU PAS Register template and HARPER, were built for real-world evidence studies with pre-specified analyses and well-defined components. They do not fully address the iterative development, validation, and deployment considerations unique to AI/ML research.

AI4HM is a fit-for-purpose template that provides structure while preserving methodological flexibility. It accommodates iterative model development, feature engineering as part of the research process, multiple validation stages, performance metrics beyond traditional epidemiology, and fairness, bias, and interpretability considerations.

ModularFlexible with rigorDownstream-compatibleAccessibleLiving document
Scope note. This template is not designed for studies that support regulatory submissions or that are considered Software as a Medical Device (SaMD). Support for regulatory submissions is planned for future versions.

Who it's for

Built for everyone at the table.

Researchers

Design and document AI/ML studies with a defensible, reproducible structure from the outset.

Sponsors

Standardize how AI/ML projects are scoped, planned, and governed across teams.

Regulators & reviewers

Assess study design, validity, and limitations against a consistent framework.

Ethics boards

Evaluate human-subjects protection, consent, and intended use with clarity.

Data custodians

Understand data provenance, transformation, and management before granting access.

The template, section by section

Seventeen core sections.

Expand any section to see what it covers. Headings marked Optional can be removed when not relevant to a study.

1Title Page
Protocol version and date, AI/ML tool name, investigators, sponsor, registration, and conflict-of-interest declaration.
2Abstract
A structured summary covering background, objectives, data sources, study design for training, outcome label, model training approach, study design for evaluation, and evaluation metrics.
3Amendments & Updates
Living-document tracking: version, date, sections updated, and a description and rationale for each change.
4List of AbbreviationsOptional
A reference list of abbreviations used throughout the protocol.
5MilestonesOptional
Key project milestones and timelines, where useful to the audience.
6Background & Rationale
The problem or unmet need the AI/ML tool addresses, existing tools and approaches, current gaps and limitations, and how this research addresses them.
7Study Scope & Intended UseOptional
Intended-use context, setting restrictions, and out-of-scope areas for the tool.
8Study Objectives
Primary, secondary, and exploratory objectives, presented in a structured table linking each Objective to its Population(s) and Metric(s).
9Research Methods
The methodological core of the protocol, organized into modular sub-sections included as relevant:
  • Existing Data Sources. Registries, EHR, claims, prior study data; validity, accuracy, provenance, and data governance.
  • Original Data Collection. New data from participants or human raters; ethics, consent, rater qualifications, and adjudication processes.
  • Data Cleaning & Transformation. Preparation, common data models (OMOP, FHIR), quality procedures, linkage.
  • Model Development: Study Design & Setting. Design classification, target vs. training population, selection criteria, time periods, design diagram.
  • Model Development: Training Task & Outcome Label. Task setup, label definition, rater credentials and annotation, label validation.
  • Model Development: Features. High-level feature types; an exhaustive list is not required, recognizing selection as iterative.
  • Model Development: Training & Fine-tuning. Iterative development, architectures, ensembling, cross-validation, software, and model version control.
  • Evaluation: Study Design & Setting. Design and rationale, comparators, determination of ground truth or a gold standard, evaluation populations and subgroups, external validity, and prompting and human-in-the-loop practices for large language models and agentic AI systems.
  • Evaluation: Metrics. Performance metrics matched to objectives, with rationale and references.
  • Evaluation: Analyses. Descriptive analyses, whether evaluation data was used in training, statistical tests, confidence intervals, fairness and bias analyses, software.
  • Sample Size Rationale & Calculations. Power and sample size with assumptions and justifications.
10Data Management
Data management and data governance procedures for transfer, storage, backup, and security, alongside model version control.
11Limitations
Data source, study design, model development, and analytic-method limitations, with mitigation steps.
12Protection of Human Subjects
Legal compliance, data privacy, and ethics review, including exemption status where applicable.
13Adverse Events
For patient-facing tools: adverse-event collection, management, and reporting, with regulatory requirements noted.
14Plans for Communicating ResultsOptional
How and where results will be disseminated.
15Other ConsiderationsOptional
Topics that are not required elements of a protocol, kept brief and high-level, for authors to address where relevant to the study and its intended use:
  • Patient & public involvement. Involvement of patients, carers, or the public in shaping the question, outcomes, conduct, or dissemination.
  • Multistakeholder engagement. Input from clinicians and other intended end users, data stewards, informaticians, health system leaders, developers, or regulators, and where in the lifecycle it occurred.
  • Post-deployment monitoring & model drift. Anticipated performance monitoring, drift detection, and what would prompt re-training.
  • Traceability. Audit logs, version control for models and datasets, and records of model updates.
  • Model cards. An existing or planned model card, referenced here and included as an appendix.
16References
Cited literature and guidance documents.
17Appendices
Code lists, algorithms, detailed methods, and a model card where one is available, with a structured index table.

Get the AI4HM template

Microsoft Word (.docx) · Version 0.3 · For feedback · Freely available

↓ Download template

About, instructions & glossary

Companion document · How to use the template, plus definitions for the terms it uses

↓ Download companion
NOTE: Guidance to authors appears in blue italic text and should be removed from the finalized protocol. The glossary lives in the companion document and may be adapted as an appendix.

How to cite

Citing AI4HM

Manuscript in preparation
Rough K, Diaz-Decaro J, Dickinson H, Divan HA, Naidoo N, Shiwa V, Smith M, Stettenbauer E, Teltsch D, Welby S. AI4HM: A Protocol Template for Observational Studies Involving Artificial Intelligence and Machine Learning in Health and Medicine. 2026. Manuscript in preparation.

Authors & acknowledgements

Created by

Kathryn Rough
John Diaz-Decaro
Harriet Dickinson
Hozefa A. Divan
Nivantha Naidoo
Veronique Shiwa
Meredith Smith
Elizabeth Stettenbauer
Dana Teltsch
Sarah Welby

This template was created and revised based on the collective experience of the authors in designing and evaluating AI/ML-based tools for healthcare and medicine.

Version history

v0.3Adds an optional Other considerations section, confirms the scope as observational designs, and expands the evaluation guidance to cover ground truth, large language models and agentic AI, data governance, and model version control.Revised 10 August 2026 to add authoring-team feedback missed at release. Please replace copies downloaded before this date.Current
v0.2Released for community feedback.↓ .docxArchived
v0.1Initial internal draft.Archived

New versions and the rationale for each change will be recorded here as the template evolves alongside AI/ML in healthcare and medicine.

Feedback & contact

Feedback and ideas for improvement are appreciated and shape future versions of the template.

info@kathrynrough.net