Coevolutionary mining of prokaryotic non-coding elements with a genome language model
A preprint describes Minerva, a genome-language-model framework for mapping sequence interactions; the reported viral repeat patterns still lack confirmed function.
TL;DR
- The unreviewed bioRxiv preprint introduces Minerva, a genome-language-model framework applied to 150 bacterial genomes.
- The authors report that 84.3% of predicted intergenic base-pairing fell outside existing annotations and describe repeat-associated RNA patterns in prophages.
- Nature reports that the viral repeats have not been shown to perform a CRISPR-like function, so the finding remains a computational lead.
The authors introduce Minerva, which uses genome language models to predict local sequence interactions, and report applying it to 150 bacterial genomes. They say 84.3% of predicted intergenic base-pairing fell outside known annotations. [1]
In prophages, the preprint reports arrays of structurally conserved but sequence-diverse non-coding RNAs. Nature says the newly reported viral repeats have not been shown to function like CRISPR systems and that no partnered DNA-slicing enzyme is known. [1] [2]
Why it matters
The work tests whether genome language models can surface candidate biological features at scale. The results are still unreviewed predictions and do not establish a new genome-editing mechanism.
Editor's note
bioRxiv labels the paper as a preprint that has not been peer reviewed. Nature reports that the viral repeat function remains uncharacterized; avoid implying a therapeutic or CRISPR capability.