OncID: Sequence-Intrinsic Identification of Cancer-Emergent Small RNAs
Ria Garg ⋅ Nuo Liu ⋅ Hani Goodarzi
Abstract
Orphan non-coding RNAs (oncRNAs) are small non-coding RNAs that appear in cancer cells and are largely absent from healthy tissue, which makes them barcodes of cancer identity. They arise from cancer's transcriptional dysregulation, in which normally silent genomic regions become active and precursor transcripts are aberrantly processed into stable fragments, unmasking sequences a normal cell rarely expresses. Existing oncRNA catalogs are produced by cohort-based statistical enrichment, which assigns a label by counting samples and cannot interpret a sequence it has not already seen. We introduce OncID, a cancer-specialized RNA sequence model that reads an oncRNA's nucleotide sequence to predict whether the sequence is cancer-emergent. OncID is trained on a corpus we created of roughly $800$K loci discovered from roughly $11$K small-RNA libraries. We evaluate the representative read, flanking genomic context, domain-adaptive pretraining, and tissue-conditioned pooling. Without tissue metadata, domain-adapted OncID with genomic flanks achieves a per-tissue macro AUROC of $0.83$. In silico mutagenesis assigns $6$-fold greater per-base prediction sensitivity to the observed read than to its genomic flanks. Together, these results show that oncRNA identity is substantially intrinsic to sequence, enabling cohort-free scoring and detection of cancer-emergent small RNAs.
Chat is not available.
Successful Page Load