Track progress & agreement

Track how a queue's work is coming along, including annotator agreement

The Analytics and Agreement tabs sit after Items on a queue’s detail view (managers also see Settings there). Analytics tells you how the work is going: how much is done, how fast, and who’s doing it. Agreement tells you something Analytics can’t: whether two people looking at the same item score it the same way. This walks both tabs using Support quality review, the queue from Create a queue, which carries the Response quality label and two submissions required per item.

Read the Analytics tab

Open Support quality review and switch to the Analytics tab.

Headline numbers

Four cards summarize the queue at a glance:

  • Total Items: a raw count of everything in the queue
  • Completed: a raw count of items that are done
  • Completion Rate: Completed turned into a percentage of Total Items
  • Avg / Day: completions averaged over the last 30 days

Status breakdown

Below the headline cards, a bar breaks total items into six buckets, each with its own count:

  • Completed
  • In Review
  • Needs Changes
  • Resubmitted
  • Pending Annotation
  • Skipped

Needs Changes and Resubmitted come out of the review workflow: an item only moves through them when the queue requires reviewer approval. In Review holds both: items part-way to the queue’s required submissions, and items that have already reached those submissions on a review-enabled queue but are still waiting on a reviewer’s verdict. On Support quality review, which requires two submissions per item, the first annotator’s submission alone puts the item in In Review, no reviewer needed.

Throughput over time

Daily Throughput (Last 30 Days) charts completions per day over that same window. Use it to spot a slowdown, or to confirm that adding annotators actually moved the queue faster.

Label distribution

Label Distribution shows one card per label, breaking down every value annotators have submitted for it. For Response quality, that’s a bar for each option, Good, Needs work, Wrong, with a count of how many times annotators picked it. A numeric or star label shows the same idea by rating instead of option, and a thumbs label shows up versus down (Label types & values covers every type).

Annotator performance

Annotator Performance lists everyone with activity in the queue. Completed counts items that have met the queue’s required submissions across every required label, so it credits every annotator whose submission contributed to that item, not just the one who finished it.

Read the Agreement tab

Switch to the Agreement tab.

Note

Agreement only has something to compare once two different annotators have actually scored the same item, not just been assigned to it. That means the queue’s submissions-per-item setting has to be above its default of 1 (Queue settings & limits covers where that lives), with submissions from two or more people on the same item. Support quality review is already set to 2, so its Agreement tab fills in as soon as a second annotator submits on the same item.

Agreement is also a gated feature that needs the Agreement Metrics entitlement: without it, the tab comes up empty instead of loading.

Overall agreement

At the top, Overall Agreement shows one percentage: the share of item/label pairs where every annotator who scored it landed on the same value. Until the precondition above is met, the number reads N/A, with “Need at least 2 annotators per item to calculate agreement” underneath it.

Per-label agreement

Per-Label Agreement breaks that number down one row per label. Agreement here is the same raw percentage, scoped to that one label; Disagreements counts how many items its annotators didn’t match on.

Cohen’s Kappa is the only agreement statistic the tab reports, whether two annotators scored an item or five. It isn’t swapped for a different multi-rater statistic once a third annotator joins. Kappa corrects the raw Agreement percentage for how often annotators would land on the same value purely by chance, so it usually reads lower, and more honestly, than Agreement alone, especially on a label with few options. As a rough guide: below 0.20 is poor agreement, 0.21–0.40 fair, 0.41–0.60 moderate, 0.61–0.80 substantial, and above 0.80 almost perfect. Kappa only computes for categorical, numeric, star, and thumbs labels; a free-text label shows a dash instead.

Annotator pair agreement

As soon as any single pair of annotators has overlapping work, Annotator Pair Agreement lists that pair with their agreement percentage and a Comparisons count: the number of item/label comparisons they share, not items, so one item with three labels counts as three. It’s the fastest way to tell an annotator who disagrees with everyone else apart from a label that’s just genuinely hard to agree on.

Dive deeper

Was this page helpful?

Questions & Discussion