OMK — Observe. Measure. Know. Make every knowledge change in your AI application evidence-backed.
-
Updated
Jul 29, 2026 - TypeScript
OMK — Observe. Measure. Know. Make every knowledge change in your AI application evidence-backed.
Measure whether your model judge agrees with human raters. Chance-corrected agreement statistics, bootstrap confidence intervals, and calibration gates for Swift Testing and CI. Zero dependencies.
Tool-agnostic inter-coder reliability (Krippendorff alpha, Cohen/Fleiss kappa) and disagreement adjudication for qualitative coding
Evaluation toolkit for multi-annotator human annotation research with agreement metrics, disagreement analysis, reports, and plots.
Classificazione dei commenti YouTube di Breaking Italy e La Repubblica tramite il modello User Needs (Shishkin & SmartOcto, 2021). Tesi magistrale in Comunicazione, ICT e Media — UniTO.
Measure how much your LLM judges actually agree. Inter-judge agreement metrics for LLM-as-a-judge evaluations.
Browser-based inter-rater reliability calculator for systematic literature reviews. Computes Krippendorff's Alpha using a pooled coincidence matrix across all (Paper, RQ, Field) units. No installation required — single HTML file, fully client-side. Built for the GenAI Evidence Hub at Learning Data Insights.
A reliability and DIF report card for LLM-judge and human-rater scoring instruments.
Add a description, image, and links to the krippendorff-alpha topic page so that developers can more easily learn about it.
To associate your repository with the krippendorff-alpha topic, visit your repo's landing page and select "manage topics."