When AI Starts Reading the Genome: The Biotech Signals to Watch

I keep a running list of biotech news that catches my eye. Three items from the past few weeks won't leave that list alone, and once you see the pattern, you can't unsee it: AI has quietly stopped being a bolt-on tool for biotech and started becoming its plumbing.
Start with the money. Eli Lilly and NVIDIA are building what Lilly calls the most powerful supercomputer ever owned by a drug company. Around the same time, Lilly signed a deal with Insilico Medicine worth up to $2.75 billion, with $115 million paid upfront, to bring Insilico's AI-discovered drug candidates to market. And Bristol Myers Squibb just put another $10 million behind its six-year partnership with insitro, adding two new ALS programs. None of these companies are "trying out" AI anymore. They're building the compute and data pipes to make AI-native discovery a permanent part of how they work.
Then there's structural biology, which has a new hard problem to chase. AlphaFold2 already solved the "what shape is this protein" question. Now labs are going after two harder ones: how proteins move between shapes over time, and how to design a brand-new protein binder from scratch instead of stumbling into one by luck. That second problem needs huge amounts of training data, and a new high-throughput sequencing method just made that possible — over 10 million data points from a single protein-engineering experiment. That's not a small bump. That's an unlock.
The one that actually made me sit up straight: a model trained on roughly 500,000 human DNA sequences went looking for the initiator element, the molecular switch that flips a gene from off to on. It found the switch in about 60% of human genes, including plenty that had resisted characterization by hand for years. The same kind of pattern recognition that folds proteins can apparently also read the part of the genome that decides when a gene turns on, not just what it builds.
Put these three together and a shape emerges. AI is moving from tool to substrate — in compute, in molecule design, and now in the regulatory grammar of DNA itself. People who track this space are already talking about the first Phase 2 trials for AI-nominated biologics landing around 2027–2028. That's a short runway from "the model found it" to "it's in a patient." If you're building anything in computational genomics right now, the gap between an interesting model and a production pipeline is closing faster than most roadmaps assume.
Sources: C2Bio Weekly AI/ML & Biotech Digest · Drug Discovery News · CNBC — Lilly/Insilico deal · pharmaphorum · phys.org — protein engineering data breakthrough · Communications Biology