RNA (ribonucleic acid) might be called the diva of biomolecules: It performs numerous key functions in cells and is also increasingly becoming the focus of attention in the development of new drugs. At the same time, it is fragile and difficult to study experimentally. This leads to a remarkable paradox: RNA sequences exist in vast quantities, yet we know comparatively little about the structures they form and the functions they perform. In particular, many RNAs that do not serve as templates for proteins remain poorly understood and, in a sense, constitute the “dark matter” of molecular biology.
The challenge is that AI methods that work remarkably well for proteins cannot simply be transferred to RNA. Protein models can draw on extensive structural data and many similar, evolutionarily related sequences. For RNA, such information is significantly scarcer. This is exactly where NucleicBERT comes in: Trained on approximately 25 million RNA sequences, NucleicBERT learns patterns and relationships directly from the data.
Researchers at KIT and Forschungszentrum Jülich were able to demonstrate that this acquired knowledge can be applied to a wide variety of tasks, ranging from predicting RNA structure to its biological function. This opens up new possibilities for basic research, biotechnology, and pharmaceuticals. The findings have been published in Nature Machine Intelligence.
A particularly noteworthy feature of NucleicBERT is that it requires only a single RNA sequence as input. It does not rely on hundreds or thousands of related sequences to make useful predictions. A closer look inside the model also reveals that it has identified correlations between RNA building blocks that provide insights into RNA structure, even though no structural information was supplied during training.
Contact at the SCC: Dr. Alexander Schug
To the open-access article “NucleicBERT interprets RNA sequence space through self-supervised language modeling”
