What you will learn
By the end of this topic, you should be able to structure clinical evaluation of MDR software and performance evaluation of IVDR software, connect evidence to specific claims, distinguish analytical or technical performance from clinical performance, and maintain the evaluation as the software and state of the art evolve.
The guidance adapts evidence principles to medical device software
MDCG 2020-1 explains clinical evaluation under the MDR and performance evaluation under the IVDR for medical device software. It recognises that software may create medical information without direct physical contact with the patient and that evidence must still establish valid medical meaning, reliable technical operation and clinical benefit.
The guidance does not replace the MDR or IVDR. It helps manufacturers interpret the evidence chain for software, including stand-alone software and software operating with hardware or external data sources.
Does the software's output have a valid medical basis, is it produced accurately and reliably, and does its use achieve the claimed clinical purpose for the intended population and context?
Precise claims determine the evidence needed
Define the condition, population, user, input data, output, clinical role and intended decision. Distinguish screening, detection, diagnosis, prediction, prognosis, monitoring, treatment recommendation and workflow-support claims.
- What medical question does the software answer?
- Which input data and acquisition conditions are required?
- What output is produced and how should it be interpreted?
- Who acts on the output and with what other information?
- What clinical benefit or performance is claimed?
- What limitations, exclusions and uncertainty apply?
Broad claims require broad evidence. Align the evaluation with MTL-102 — Intended Purpose, Users and Use Environments and the qualification decision in MTL-314 — MDCG 2019-11 Software Qualification and Classification.
MDR and IVDR use related but distinct evidence frameworks
MDR: valid clinical association
Evidence that the software output is associated with the targeted clinical condition or physiological state.
MDR: technical performance
Evidence that the software reliably generates the intended output from the specified input.
MDR: clinical performance
Evidence that the output achieves the intended clinical purpose in the target population and context.
IVDR: scientific validity
Association of an analyte with a clinical condition or physiological state.
IVDR: analytical performance
Ability to detect or measure the analyte correctly under defined conditions.
IVDR: clinical performance
Ability to yield results correlated with the clinical condition or process for the intended purpose.
Use the terminology of the applicable Regulation. The evidence elements interact; a strong algorithm test cannot compensate for an invalid clinical association, and a plausible association cannot compensate for unreliable software output.
Establish the medical or scientific basis
Demonstrate that the relationship between the software's input, output and targeted medical condition is accepted or adequately supported. Evidence can include peer-reviewed literature, professional guidance, clinical practice, consensus standards, original studies and high-quality datasets.
Evaluate relevance to the exact population, input modality, output and intended clinical role. A general association in the literature may not support the product's threshold, prediction horizon, subgroup or recommended action.
Technical performance demonstrates correct and reliable computation
Show that software produces accurate, repeatable and robust outputs from specified inputs. Evidence may address data ingestion, preprocessing, algorithms, calculations, units, thresholds, interfaces, timing, error handling and platform compatibility.
Connect this evidence to MTL-106 — Verification and Validation and MTL-128 — Statistical Methods and Measurement Assurance.
Clinical performance must support use in the intended context
Evaluate whether the software output can achieve the intended clinical purpose when used by representative users, with representative patients and data, in the expected workflow. Relevant measures depend on the claim and may include sensitivity, specificity, predictive values, agreement, calibration, clinical outcome or decision impact.
Predefine clinically meaningful acceptance criteria and justify them against state of the art, benefit–risk and alternative methods. Report uncertainty and subgroup performance, not only an overall average. Consider the effect of prevalence, missing data, operator behaviour and downstream clinical interpretation.
Data suitability is part of the evidence
Describe data provenance, inclusion criteria, acquisition, labelling, reference methods, quality controls, representativeness and independence. Identify bias, confounding, missingness, leakage and limitations.
- Does the dataset represent the intended population and use environment?
- Are ground truth and reference standards clinically credible?
- Are training, tuning and evaluation data appropriately separated?
- Are important demographic, clinical and technical subgroups represented?
- Can records be traced to software and model versions?
- Are privacy, consent and data-governance obligations addressed?
MTL-129 — Privacy and Data Protection by Design explains related controls.
Algorithm evidence must address the released implementation
Explain the algorithm, features, thresholds, training process where relevant, assumptions and known failure modes. For adaptive or machine-learning systems, define what is locked at release, how changes are controlled and how performance drift will be detected.
Validate the complete software pipeline rather than only the model in isolation. Preprocessing, data mapping, user selection, interface behaviour and post-processing can materially affect clinical output.
Literature and equivalence need critical appraisal
Published evidence can support clinical association, state of the art and sometimes performance, but relevance must be demonstrated. Assess study quality, population, technology, inputs, outputs, endpoints, software version and clinical context.
Claims of equivalence require access to sufficient technical, biological where relevant, and clinical characteristics under the MDR. Similarity to a marketed app or algorithm is not automatically regulatory equivalence and does not remove the need for product-specific technical evidence.
Close evidence gaps deliberately
Map every claim and GSPR to available evidence and identify residual gaps. Options can include bench or software studies, retrospective data analysis, prospective clinical investigation, performance study, usability work or narrowed claims.
Choose the least burdensome method that still provides sufficient evidence. If the available data only support a narrower population or output, constrain intended purpose and labelling rather than overstating the conclusion.
Evaluation continues after market placement
Update the clinical or performance evaluation with post-market surveillance, PMCF or PMPF, complaints, incidents, literature, state-of-the-art changes and software performance monitoring. Compare real-world inputs and users with the original evidence base.
Software changes, model retraining, new data sources, platform updates or altered workflows can affect validity and performance. Use MTL-127 — Configuration and Change Management and MTL-111 — Post-market Support to maintain the evidence connection.
Common misconceptions
“Accurate code proves clinical performance.”
Technical correctness does not establish that the output is clinically valid or beneficial.
“Published research validates our algorithm.”
Literature must be relevant to the specific implementation, population, data and claim.
“One large dataset is sufficient.”
Representativeness, independence, ground truth, subgroups and bias matter as much as record count.
“A retrospective study is always enough for software.”
The required design follows the evidence gap and intended clinical role; prospective evidence may be necessary.
“Model performance equals product performance.”
The complete software pipeline, interface and workflow affect the result.
“Evaluation ends at CE marking.”
Software, data, practice and state of the art change, requiring lifecycle review and post-market evidence.
Practical evidence checklist
- Are medical claims specific and traceable to intended purpose?
- Is the correct MDR or IVDR evidence terminology used?
- Is the clinical association or scientific validity supported?
- Does technical or analytical performance cover the released implementation?
- Does clinical performance reflect representative users, patients and workflows?
- Are acceptance criteria clinically meaningful and predefined?
- Are datasets representative, independent and traceable?
- Are subgroup performance, uncertainty and limitations reported?
- Are literature and equivalence claims critically appraised?
- Does the post-market plan address drift, change and evidence gaps?
Authoritative references
- MDCG 2020-1 — Clinical evaluation / performance evaluation of medical device software
- Regulation (EU) 2017/745 — Medical Device Regulation
- Regulation (EU) 2017/746 — In Vitro Diagnostic Medical Device Regulation
- European Commission — Current MDCG-endorsed guidance
Confirm the current legislation, guidance, common specifications and device-specific evidence expectations when defining the evaluation strategy.
Software evidence must connect medical meaning, technical truth and clinical value
A defensible evaluation shows that the claimed association is valid, the released software produces reliable output and that output achieves its intended clinical purpose in the target population and workflow.