SAID: Multi-Domain Sinhala Datasets for AI-Text Detection and Authorship-Boundary Detection
Abstract
Sinhala lacks a dedicated public benchmark for detecting text generated by artificial intelligence (AI) across domains and locating authorship changes within mixed documents. We present the first Sinhala benchmark dataset, SAID (Sinhala AI-text Identification Dataset), consisting of two core datasets. The first, Sinhala-HAT (Human and AI-generated Text), contains 9,000 documents: 4,500 human texts paired one-to-one with 4,500 AI-generated texts across three domains: news, Wikipedia, and social media. We use title-to-text and paraphrased generation strategies for news and Wikipedia, while social media uses paraphrased generation only. The second dataset, Sinhala-BOUND (Authorship Boundary Dataset), targets sentence-level authorship-boundary detection. It contains 4,244 mixed Wikipedia documents with 39,358 labelled sentences, constructed using continuation, single-span replacement, and multi-span replacement strategies. Sentence-level labels identify human-written and AI-generated contributions and their authorship boundaries. Human sources predate the release of ChatGPT, while AI-generated text is produced by three large language models (LLMs): GPT-4o, DeepSeek V3, and Gemini 2.5 Pro. Task instructions are written in Sinhala without an intermediate translation step. Length control and symmetric cleaning reduce formatting and length shortcuts, while source and generation metadata support source-grouped splits and cross-generator evaluation. The contribution is dataset construction and benchmark curation, complemented by baseline evaluation. Four pretrained encoders achieve 0.871–0.907 F1 in-distribution on Sinhala-HAT, while leave-one-domain-out evaluation shows substantially larger degradation than leave-one-generator-out evaluation. Together, these resources support document-level AI-text detection and sentence-level authorship-boundary research in an underserved language.