Fast computation of Krippendorff's alpha agreement measure in Python.
-
Updated
Aug 2, 2026 - Python
Fast computation of Krippendorff's alpha agreement measure in Python.
CrowdTruth framework for crowdsourcing ground truth for training & evaluation of AI systems
Inter-annotator agreement for Brat annotation projects
Exploring the inter-annotator agreement between ISIC Archive segmentation masks
This repository holds code to annotate textual data using LLMs, and calculate different measures of Inter-Annotator Agreement (IAA).
command line stuff to calculate inter annotator agreement
Official Repository of 'Leveraging Epistemic Uncertainty to Improve Tumour Segmentation in Breast MRI: An exploratory analysis' In SPIE Medical Imaging 2024
Python package implementing measure from my working paper "Kappa-IoU: Inter-Rater Reliability for Spatial Annotation"
A Python Library to compute Disagreement Systematicity (σ)
Evaluation toolkit for multi-annotator human annotation research with agreement metrics, disagreement analysis, reports, and plots.
Measures whether your graders and your rubric are trustworthy: inter-annotator agreement, anchor diagnostics, grader calibration against gold, and an adjudication queue.
[MICCAI ISIC 2024] Code for "Segmentation Style Discovery: Application to Skin Lesion Images"
Digita Literacy for School of Foreign Languages (NRU HSE, 2018)
Measure inter-annotator agreement with modern metrics, uncertainty, ordinal support, and missing-data handling.
Browser-based IAA and annotation workflow tools for Yongle Palace mural patches
Concordia: reproducible inter-annotator agreement toolkit for human and machine cultural-heritage annotation
Production IAA engine — Krippendorff's Alpha, SBERT semantic agreement, Fleiss' Kappa, Shannon entropy, 3-layer collusion detection. Ran behind RawEval's 9-annotator workbench.
Reproducible audit of the MAD dataset (MAST failure taxonomy): three undocumented taxonomy versions, renumbered codes, and what that means for reported inter-annotator agreement.
Generate controlled synthetic annotator disagreement before real reviewer data is collected.
Production-grade pipeline for validating annotation consistency and evaluating LLM output quality using agreement metrics, schema validation, LLM-as-judge scoring, SQLite logging, and a Streamlit dashboard.
To associate your repository with the inter-annotator-agreement topic, visit your repo's landing page and select "manage topics."