For over two decades, the human genome has been a bit like a book with 3 billion letters, where only 1-2% actually make sense. The rest? A vast, mysterious expanse of what scientists charmingly call "junk DNA" and regulatory bits that control how our genes express themselves. Turns out, this "junk" might hold the keys to everything from cancer to why you can (or can't) roll your tongue.
Cracking that code — understanding how tiny changes in these non-coding regions affect our traits and diseases — has been a monumental task. But now, UC Berkeley researchers have unveiled GPN-Star, a new AI model that's not just better at finding these crucial genetic changes; it's also a lot less of a power hog. Think of it as the ultra-efficient, super-smart genetic detective we've been waiting for.
"Our model is excellent at predicting if genetic changes are harmful and at finding functional parts of the genome," explains Yun Song, a Berkeley professor and senior author. And they've already released its genome-wide predictions, highlighting changes that are likely to have the biggest impact on inherited traits. Biologists, start your engines.
We're a new kind of news feed.
Regular news is designed to drain you. We're a non-profit built to restore you. Every story we publish is scored for impact, progress, and hope.
Start Your News DetoxChatbots for DNA
Genomic language models are essentially chatbots for DNA, but instead of learning from your questionable grammar on Twitter, they feast on vast amounts of DNA sequences. This allows them to spot repeating patterns and sequences at speeds no human could ever hope to match, revealing crucial parts of the genome that were previously invisible.
"Mathematically, a DNA sequence is just a string of letters — A, C, G, and T," Song clarifies. "We don't know beforehand which parts are functional, and only a tiny percentage is. By training a DNA language model on many different sequences, it can recognize patterns. People use this to learn what we call the 'grammar' of the genome."
Previous models, like the massive Evo 2, required an absurd amount of computing power — 2,000 powerful processors and months of training on over 100,000 species. Imagine the electricity bill. GPN-Star, however, takes a different, more enlightened approach.
Instead of individual genomes, Song's team trained GPN-Star on whole-genome alignments (WGAs). These are like pre-digested comparisons of hundreds of species to one (say, humans), showing where the code has remained stubbornly the same or decided to switch things up over evolutionary time. Because WGAs already highlight the conserved, important bits, GPN-Star can be trained in mere hours or days with just a few processors. It's also less likely to be distracted by all that "junk DNA" floating around.
"We tried to help the model learn by using data more likely to contain functional elements," Song notes. It's a bit like giving a student a textbook that's already highlighted the important parts, rather than making them read every single word and hope they figure it out.
They trained GPN-Star on different evolutionary timescales — primates, mammals, and other vertebrates — and found that models optimized for different timescales were better at interpreting different kinds of genetic variants. Which, if you think about it, makes perfect sense. Some genomic elements evolve at a snail's pace, while others are practically sprinting.
For instance, the model trained on longer timescales was a whiz at predicting the impact of rare genetic changes in proteins (the slow evolvers). But for complex traits like schizophrenia risk, which can be linked to as many as 10,000 different mutations (many in those non-coding regions), the model trained on shorter, primate-specific timescales was surprisingly accurate. Apparently, recent human evolution has its own kind of genetic drama.
Because GPN-Star is so resource-friendly, the hope is that other teams can jump in, tweak it, and make it even better. This could dramatically accelerate our understanding of genetics in humans and beyond. Because while we're making great progress, as co-first author Gonzalo Benegas puts it, "the more people who can work with these models, the better they will get." Let the genomic deciphering commence.











