Back to Help Center
Calibration and inter-rater reliability
For challenges with several judges, consistent scoring matters — if two judges interpret a criterion very differently, the final ranking can end up reflecting who happened to judge a submission rather than its actual quality. Calibration rounds address this by having judges score a shared sample set before real judging opens, so obvious inconsistencies surface early.
Behind the scenes, the platform can measure inter-rater reliability (how consistently judges agree with each other) using Krippendorff's Alpha, and can flag an individual judge's scores as an outlier if they diverge significantly from the rest of the panel on the same submission — that flag is visible to challenge creators, not to other judges.
Was this helpful?