Annotation & evaluation
Native-speaker judgment, at production scale
Transcription, translation, labeling, red-teaming and model evaluation — on data we collect or data you already own. Double-blind QA, reported agreement, and a written accuracy SLA.
Services
Human judgment where the model needs it
Audio & speech annotation
Verbatim and clean transcription, timestamping, forced alignment, speaker diarization, emotion and intent tagging, and audio event labeling.
Text annotation
Named entities, intent and slot filling, sentiment and stance, toxicity, relation extraction, summarisation quality and coreference — in-script and transliterated.
Visual annotation
Boxes, polygons, segmentation, keypoints, 3D cuboids, multi-object tracking, OCR with reading order, and native-language image captioning.
Model evaluation
Side-by-side preference, rubric-scored grading, factuality and citation checking, and localisation review by raters who live in the target market.
Safety & red-teaming
In-language adversarial probing, culturally specific harm taxonomies, jailbreak collection and severity grading, run by a welfare-protected specialist pool.
Localisation QA
Register, honorific and politeness review, cultural appropriateness, date/number/address formatting, and idiom checking for product surfaces.
Quality system
Three independent layers, all reported to you
Most vendors report one number and hope you do not ask how it was derived. We report the method, the sample size and the disagreements.
Gold seeding
Pre-adjudicated items are injected invisibly into live queues at a set rate. Every annotator carries a rolling accuracy score; falling below threshold pauses their queue automatically and triggers retraining, not silent degradation.
Double-blind passes
A defined proportion of items — 10% baseline, 100% for high-stakes schemas — are labeled independently by two annotators who cannot see each other's work. Agreement is computed per batch and per label class, so you can see which classes are genuinely ambiguous.
Senior adjudication
Disagreements route to a senior linguist or domain expert whose decision is final and is fed back into the written rubric. Rubric changes are versioned, so you can trace why a label convention shifted in week six.
What you receive with every batch: per-annotator accuracy distribution, inter-annotator agreement by label class, adjudication log, rubric version, throughput against plan, and the list of items we rejected internally before you ever saw them.
Engagement models
Buy it the way that fits your team
Per-unit, dedicated team or managed program. The quality system is identical; only the commercial shape changes.
| Model | Best for | Commitment | |
|---|---|---|---|
| Per-unit | Defined, stable schema and predictable volume | None — pay per item | Published unit rates |
| Dedicated pod | Evolving schema, ongoing iteration with your ML team | Monthly, 4 FTE minimum | Per-seat monthly |
| Managed program | Multi-language, multi-task, SLA-bound production | Quarterly or annual | Blended program rate |
| Secure facility | Regulated or highly sensitive data | Quarterly minimum | Per-seat + facility fee |
Answers
Annotation & evaluation questions
Yes — roughly a third of our work is annotation-only on client-owned data. We operate as a processor under your DPA, work inside your environment where required, and apply the same double-blind QA standard we use on our own collection. You keep every right to the underlying data; we assign the rights to the labels.
Ours or yours. We run our own annotation platform for speed and QA telemetry, but we staff on client platforms constantly — Label Studio, CVAT, Scale, Labelbox, SuperAnnotate, V7 and in-house tools. If you need a bespoke interface for an unusual schema we will build it; that is usually a one to two week engineering task.
Three layers. Gold questions seeded invisibly into live work measure per-annotator accuracy continuously. Double-blind passes on a defined proportion of items give inter-annotator agreement, reported per batch as Cohen's or Krippendorff's alpha depending on schema. A senior adjudicator resolves disagreements and their decisions feed back into the rubric. You see all three numbers in every delivery report.
Yes. Side-by-side model comparison with calibrated native raters, rubric-scored single-response grading, adversarial red-teaming in-language, and structured harm-taxonomy evaluation. For safety work we use an enhanced-screening pool with welfare protections: capped exposure hours, briefed and debriefed, paid rest, and unconditional opt-out.
From a standing start: 20 annotators in a week, 100 in three weeks, 500+ in six to eight weeks for a well-specified task in a high-resource language. Low-resource languages are the constraint — there may only be a few hundred qualified annotators for a given dialect worldwide, and we will tell you honestly what the ceiling is rather than quietly missing your deadline.
Send us 500 items and your schema.
We will label them free, return the accuracy and agreement numbers, and show you the disagreements. If the numbers are not better than what you have now, we have not earned the work.
Average first response: under 6 business hours. NDAs signed same day.