What it is
Structure where you need it, flexibility where the science demands it.
Protocols are living documents that describe the intended conduct of a study. Having a pre-specified protocol is scientific best practice, and is commonly expected by data stewards, ethics boards, journals, and other research stakeholders.
Existing templates, such as the EU PAS Register template and HARPER, were built for real-world evidence studies with pre-specified analyses and well-defined components. They do not fully address the iterative development, validation, and deployment considerations unique to AI/ML research.
AI4HM is a fit-for-purpose template that provides structure while preserving methodological flexibility. It accommodates iterative model development, feature engineering as part of the research process, multiple validation stages, performance metrics beyond traditional epidemiology, and fairness, bias, and interpretability considerations.
Who it's for
Built for everyone at the table.
Researchers
Design and document AI/ML studies with a defensible, reproducible structure from the outset.
Sponsors
Standardize how AI/ML projects are scoped, planned, and governed across teams.
Regulators & reviewers
Assess study design, validity, and limitations against a consistent framework.
Ethics boards
Evaluate human-subjects protection, consent, and intended use with clarity.
Data custodians
Understand data provenance, transformation, and management before granting access.
The template, section by section
Seventeen core sections.
Expand any section to see what it covers. Headings marked Optional can be removed when not relevant to a study.
1Title Page›
2Abstract›
3Amendments & Updates›
4List of AbbreviationsOptional›
5MilestonesOptional›
6Background & Rationale›
7Study Scope & Intended UseOptional›
8Study Objectives›
9Research Methods›
- Existing Data Sources. Registries, EHR, claims, prior study data; validity, accuracy, provenance, and data governance.
- Original Data Collection. New data from participants or human raters; ethics, consent, rater qualifications, and adjudication processes.
- Data Cleaning & Transformation. Preparation, common data models (OMOP, FHIR), quality procedures, linkage.
- Model Development: Study Design & Setting. Design classification, target vs. training population, selection criteria, time periods, design diagram.
- Model Development: Training Task & Outcome Label. Task setup, label definition, rater credentials and annotation, label validation.
- Model Development: Features. High-level feature types; an exhaustive list is not required, recognizing selection as iterative.
- Model Development: Training & Fine-tuning. Iterative development, architectures, ensembling, cross-validation, software, and model version control.
- Evaluation: Study Design & Setting. Design and rationale, comparators, determination of ground truth or a gold standard, evaluation populations and subgroups, external validity, and prompting and human-in-the-loop practices for large language models and agentic AI systems.
- Evaluation: Metrics. Performance metrics matched to objectives, with rationale and references.
- Evaluation: Analyses. Descriptive analyses, whether evaluation data was used in training, statistical tests, confidence intervals, fairness and bias analyses, software.
- Sample Size Rationale & Calculations. Power and sample size with assumptions and justifications.
10Data Management›
11Limitations›
12Protection of Human Subjects›
13Adverse Events›
14Plans for Communicating ResultsOptional›
15Other ConsiderationsOptional›
- Patient & public involvement. Involvement of patients, carers, or the public in shaping the question, outcomes, conduct, or dissemination.
- Multistakeholder engagement. Input from clinicians and other intended end users, data stewards, informaticians, health system leaders, developers, or regulators, and where in the lifecycle it occurred.
- Post-deployment monitoring & model drift. Anticipated performance monitoring, drift detection, and what would prompt re-training.
- Traceability. Audit logs, version control for models and datasets, and records of model updates.
- Model cards. An existing or planned model card, referenced here and included as an appendix.
16References›
17Appendices›
Get the AI4HM template
Microsoft Word (.docx) · Version 0.3 · For feedback · Freely available
About, instructions & glossary
Companion document · How to use the template, plus definitions for the terms it uses
How to cite
Citing AI4HM
Authors & acknowledgements
Created by
This template was created and revised based on the collective experience of the authors in designing and evaluating AI/ML-based tools for healthcare and medicine.
Version history
New versions and the rationale for each change will be recorded here as the template evolves alongside AI/ML in healthcare and medicine.
Feedback & contact
Feedback and ideas for improvement are appreciated and shape future versions of the template.
info@kathrynrough.net