A statistic that measures agreement between two raters beyond what chance alone would produce.
Cohen's Kappa quantifies inter-rater reliability for categorical ratings (e.g., two clinicians independently classifying the same movement fault or imaging finding), correcting for the agreement expected by chance. Values typically range from 0 (agreement no better than chance) to 1 (perfect agreement), with common but non-universal benchmarks labelling 0.60-0.80 as substantial and above 0.80 as near-perfect; these cutoffs are conventional rather than statistically derived, so interpret alongside prevalence and the number of categories, as kappa can be misleadingly low when one category is rare.
This guide was auto-drafted and is pending editorial review.