Inter-rater reliabilityResearch

How consistently different assessors get the same result when rating the same thing.

Inter-rater reliability is the degree to which two or more independent assessors agree when scoring, classifying, or measuring the same subject or data using the same tool or protocol. It is usually reported as a statistic such as Cohen's kappa (categorical ratings), an intraclass correlation coefficient (continuous or ordinal ratings), or percentage agreement, and there is no universal threshold for what counts as "acceptable" — cutoffs vary by field and by the stakes of the decision being made. Low inter-rater reliability suggests a measure depends heavily on rater judgement and may need clearer criteria, training, or a more objective tool.

Dose-Response critically appraises the sports-medicine, rehab and performance research that matters — in plain English, every week. Subscribe free →

This guide was auto-drafted and is pending editorial review.