Work out inter annotator agreement instantly with clear inputs, formula shown and shareable results.
Raw agreement flatters annotators on skewed tasks: if 90 percent of items are negative, two people guessing would agree most of the time. Cohen's kappa subtracts the chance agreement and rescales, so it measures how much of the achievable agreement above chance was actually reached. Values below 0.4 mean the guidelines, not the annotators, need fixing.
Cohen's kappa
kappa = (Po - Pe) / (1 - Pe)
Aim for 0.7 or above on a pilot batch. Below that, disagreements will end up in the training labels as noise and cap the achievable model accuracy.
Fleiss' kappa for multiple raters and nominal labels, or Krippendorff's alpha when raters differ per item or the labels are ordinal.