<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <title>The Dataa — AI training data research and field notes</title>
  <link>https://www.thedataa.com/blog.html</link>
  <atom:link href="https://www.thedataa.com/rss.xml" rel="self" type="application/rss+xml"/>
  <description>Practical writing on AI training data: how much you need to fine-tune, licensing and consent, low-resource languages, ASR and TTS data, annotation quality.</description>
  <language>en</language>
  <lastBuildDate>Tue, 01 Sep 2026 00:00:00 +0000</lastBuildDate>
  <item>
    <title>Measuring Data Annotation Quality</title>
    <link>https://www.thedataa.com/ai-data-annotation-quality.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/ai-data-annotation-quality.html</guid>
    <pubDate>Wed, 17 Jun 2026 00:00:00 +0000</pubDate>
    <category>Practical guides</category>
    <description>Gold seeding, double-blind passes and senior adjudication — the three layers that produce trustworthy labels, and what a delivery report must contain.</description>
  </item>
  <item>
    <title>Pashto NLP Resources &amp; Datasets</title>
    <link>https://www.thedataa.com/pashto-nlp-guide.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/pashto-nlp-guide.html</guid>
    <pubDate>Wed, 10 Jun 2026 00:00:00 +0000</pubDate>
    <category>Language guides</category>
    <description>Pashto spans Pakistan and Afghanistan with two major dialect groups that differ enough to degrade a model tuned on only one. What exists and what a spec must state.</description>
  </item>
  <item>
    <title>Punjabi NLP Resources &amp; Datasets</title>
    <link>https://www.thedataa.com/punjabi-nlp-resources.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/punjabi-nlp-resources.html</guid>
    <pubDate>Wed, 03 Jun 2026 00:00:00 +0000</pubDate>
    <category>Language guides</category>
    <description>Punjabi is among the world&#x27;s most spoken languages and among the least resourced in AI. What exists, what does not, and how Shahmukhi and Gurmukhi complicate everything.</description>
  </item>
  <item>
    <title>How to Run a Speech Data Collection Project</title>
    <link>https://www.thedataa.com/speech-data-collection-guide.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/speech-data-collection-guide.html</guid>
    <pubDate>Wed, 27 May 2026 00:00:00 +0000</pubDate>
    <category>Practical guides</category>
    <description>End to end: writing the spec, recruiting to quota, choosing capture conditions, transcription standards, QA gates and acceptance testing.</description>
  </item>
  <item>
    <title>Synthetic vs Real AI Training Data</title>
    <link>https://www.thedataa.com/synthetic-vs-real-training-data.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/synthetic-vs-real-training-data.html</guid>
    <pubDate>Wed, 20 May 2026 00:00:00 +0000</pubDate>
    <category>Research</category>
    <description>Synthetic data is cheap and increasingly good. Where it genuinely substitutes for real collection, where it quietly degrades models, and how to combine them.</description>
  </item>
  <item>
    <title>How to Choose an AI Training Data Vendor</title>
    <link>https://www.thedataa.com/how-to-choose-ai-data-vendor.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/how-to-choose-ai-data-vendor.html</guid>
    <pubDate>Wed, 13 May 2026 00:00:00 +0000</pubDate>
    <category>Practical guides</category>
    <description>Twelve questions that separate serious data vendors from resellers: provenance, consent, QA methodology, pay transparency and what happens when delivery misses spec.</description>
  </item>
  <item>
    <title>Urdu to Hindi Transfer Learning</title>
    <link>https://www.thedataa.com/urdu-hindi-transfer-learning.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/urdu-hindi-transfer-learning.html</guid>
    <pubDate>Wed, 06 May 2026 00:00:00 +0000</pubDate>
    <category>Research</category>
    <description>Urdu and Hindi are largely mutually intelligible in speech and written in different scripts. What that means for ASR, OCR and LLM transfer, with the honest limits.</description>
  </item>
  <item>
    <title>Inter-Annotator Agreement Explained</title>
    <link>https://www.thedataa.com/inter-annotator-agreement-guide.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/inter-annotator-agreement-guide.html</guid>
    <pubDate>Wed, 29 Apr 2026 00:00:00 +0000</pubDate>
    <category>Definitions</category>
    <description>Cohen&#x27;s kappa, Krippendorff&#x27;s alpha and raw agreement — what each measures, what counts as acceptable, and why low agreement usually means bad guidelines.</description>
  </item>
  <item>
    <title>RLHF in Non-English Languages</title>
    <link>https://www.thedataa.com/rlhf-outside-english.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/rlhf-outside-english.html</guid>
    <pubDate>Wed, 22 Apr 2026 00:00:00 +0000</pubDate>
    <category>Research</category>
    <description>Collecting preference data outside English needs calibrated native raters and a rubric written in their language. Why translated rubrics produce noise, and how to fix it.</description>
  </item>
  <item>
    <title>SFT Data Quality: Why Translation Fails</title>
    <link>https://www.thedataa.com/sft-data-quality-guide.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/sft-data-quality-guide.html</guid>
    <pubDate>Wed, 15 Apr 2026 00:00:00 +0000</pubDate>
    <category>Practical guides</category>
    <description>What separates useful instruction fine-tuning data from filler, how many pairs you need, and why translating an English seed set produces a model that sounds foreign.</description>
  </item>
  <item>
    <title>TTS Dataset Requirements: Hours Per Voice</title>
    <link>https://www.thedataa.com/tts-dataset-requirements.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/tts-dataset-requirements.html</guid>
    <pubDate>Wed, 08 Apr 2026 00:00:00 +0000</pubDate>
    <category>Practical guides</category>
    <description>How many studio hours a text-to-speech voice actually needs, why TTS data differs completely from ASR data, and what phonetic balance means in practice.</description>
  </item>
  <item>
    <title>Quota Sampling for AI Training Data</title>
    <link>https://www.thedataa.com/quota-sampling-training-data.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/quota-sampling-training-data.html</guid>
    <pubDate>Wed, 01 Apr 2026 00:00:00 +0000</pubDate>
    <category>Practical guides</category>
    <description>Why 500 well-distributed hours beat 5,000 convenience-sampled ones, which axes to quota on, and how to write a data specification that survives contact with reality.</description>
  </item>
  <item>
    <title>Code-Switching in Speech Recognition</title>
    <link>https://www.thedataa.com/code-switching-speech-recognition.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/code-switching-speech-recognition.html</guid>
    <pubDate>Wed, 25 Mar 2026 00:00:00 +0000</pubDate>
    <category>Research</category>
    <description>Bilingual speakers mix languages mid-sentence constantly. Why monolingual ASR fails on it, and how to specify and collect code-switched training data.</description>
  </item>
  <item>
    <title>Word Error Rate (WER) Explained</title>
    <link>https://www.thedataa.com/word-error-rate-explained.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/word-error-rate-explained.html</guid>
    <pubDate>Wed, 18 Mar 2026 00:00:00 +0000</pubDate>
    <category>Definitions</category>
    <description>What WER measures, how to compute it, what counts as good, and why a model at 8% WER in testing can exceed 40% on real production audio.</description>
  </item>
  <item>
    <title>Informed Consent for AI Training Data</title>
    <link>https://www.thedataa.com/informed-consent-ai-training-data.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/informed-consent-ai-training-data.html</guid>
    <pubDate>Wed, 11 Mar 2026 00:00:00 +0000</pubDate>
    <category>Legal &amp; ethics</category>
    <description>Valid consent must name machine-learning training, be understood by the person giving it, and be withdrawable. How that works at low literacy and across languages.</description>
  </item>
  <item>
    <title>Free vs Commissioned AI Datasets: When to Pay</title>
    <link>https://www.thedataa.com/open-vs-commissioned-datasets.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/open-vs-commissioned-datasets.html</guid>
    <pubDate>Wed, 04 Mar 2026 00:00:00 +0000</pubDate>
    <category>Practical guides</category>
    <description>Common Voice, FLEURS and OpenSLR are free and often sufficient. Here is exactly when open data runs out and commissioned collection becomes the cheaper option.</description>
  </item>
  <item>
    <title>Why Urdu OCR Is Hard: The Nastaliq Problem</title>
    <link>https://www.thedataa.com/urdu-ocr-nastaliq-hard.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/urdu-ocr-nastaliq-hard.html</guid>
    <pubDate>Wed, 25 Feb 2026 00:00:00 +0000</pubDate>
    <category>Research</category>
    <description>Nastaliq breaks OCR pipelines built for Arabic. Sloping baselines, extreme ligature context and overlapping glyphs explain why, and what training data fixes it.</description>
  </item>
  <item>
    <title>AI Training Data Licensing Explained</title>
    <link>https://www.thedataa.com/ai-training-data-licensing-explained.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/ai-training-data-licensing-explained.html</guid>
    <pubDate>Wed, 18 Feb 2026 00:00:00 +0000</pubDate>
    <category>Legal &amp; ethics</category>
    <description>What a training-rights licence must contain, why most content licences do not cover model training, and the clauses that stop deals in enterprise legal review.</description>
  </item>
  <item>
    <title>Why AI Fails on Low-Resource Languages</title>
    <link>https://www.thedataa.com/why-ai-fails-low-resource-languages.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/why-ai-fails-low-resource-languages.html</guid>
    <pubDate>Wed, 11 Feb 2026 00:00:00 +0000</pubDate>
    <category>Research</category>
    <description>Low-resource languages are a data problem, not a modelling problem. What actually breaks, why speaker count and corpus size are unrelated, and what fixes it.</description>
  </item>
  <item>
    <title>How Much Data to Fine-Tune an ASR Model?</title>
    <link>https://www.thedataa.com/how-much-data-to-fine-tune-asr.html</link>
    <guid isPermaLink="true">https://www.thedataa.com/how-much-data-to-fine-tune-asr.html</guid>
    <pubDate>Wed, 04 Feb 2026 00:00:00 +0000</pubDate>
    <category>Practical guides</category>
    <description>How many hours of speech data you actually need to fine-tune an ASR model, with realistic thresholds for 200, 500, 2,000 and 10,000 hours.</description>
  </item>
</channel>
</rss>
