AlphaFold in Drug Target Identification
AI-predicted protein structures accelerate drug discovery by revealing previously invisible targets.

Most drugs work by locking onto a specific site on a protein, the way a key fits a lock. Knowing that lock's shape has always been step one in drug design, and for decades, that step was the slow one. AlphaFold changed the speed and cost of getting a protein's 3D shape, and in doing so it changed which proteins researchers even bother going after. This piece follows that shift through the actual workflow of finding a drug target: from a raw amino acid sequence to a screened candidate sitting in a lab freezer somewhere.
Protein sequences got cheap once genome sequencing costs fell off a cliff in the 2000s. Figuring out what shape a given sequence folds into stayed expensive. X-ray crystallography and cryo-EM can pin down a structure atom by atom, but each run takes months of setup, a protein willing to crystallize (plenty aren't), and a budget that doesn't flinch at the cost. That mismatch, cheap sequences against expensive structures, left huge stretches of the human proteome structurally blank, which meant they sat off-limits as drug targets by default. The biology was rarely the obstacle; the geometry was the missing piece, and geometry is the thing you actually need to design a molecule against.
That gap sits upstream of the number everyone in the industry recites: many billions of dollars to get one new drug to market. A fair share of that time and money traces back to not knowing what a target looks like early enough to do anything with it.
What AlphaFold achieved and how quickly the field gained access to it
AlphaFold 2's showing at CASP14 in 2020 is the moment people point to. It predicted protein structures directly from amino acid sequences at an accuracy that competition organizers and outside scientists called, without much hedging, a solution to a problem structural biologists had chased for 50 years. CASP is a blind competition, and the result held up under scrutiny from people whose whole job is finding the flaws.
The bigger move came in 2021, when DeepMind and EMBL-EBI opened the AlphaFold Protein Structure Database to anyone with a browser. It now holds over 200 million predicted structures, covering the entire human proteome plus 47 other organisms picked for their relevance to research and global health. An analysis from the Innovation Growth Lab found that researchers using AlphaFold 2 increased their submissions of new experimental protein structures by more than 40%. Predicted structures pointed people toward what was worth confirming by hand, supplementing the lab work rather than replacing it.
The 2024 Nobel Prize in Chemistry made the institutional verdict official, split between Demis Hassabis and John Jumper for structure prediction and David Baker for computational protein design. But the number that matters more to drug hunters than any prize is this one: roughly half of the human proteins that had no experimental structure on record, the ones biologists call the "dark proteome," now have a predicted structure sitting in a public database. That's visibility into territory that was closed off entirely, not more data piled onto familiar targets.
How AlphaFold changes the starting point of target identification
Old-school target identification starts with known biology, sure, but it's quietly filtered by what structural data already exists. A researcher picks a disease pathway, looks at the proteins involved, and in practice works only with the ones crystallography or cryo-EM already solved. Everything else sits unaddressed: biologically interesting, structurally invisible, functionally untouchable.
AlphaFold breaks that filter because running a prediction is now cheap enough to do speculatively, before anyone commits a crystallography budget to it. A researcher can take a disease of interest, pull every protein in its pathway, and get a structural hypothesis for the whole list, not just the subset that happened to crystallize nicely.
CDK20 makes this concrete instead of theoretical. It had no crystal structure and almost no inhibitor literature to speak of, and that second fact follows directly from the first: nobody can design a molecule against a shape nobody has. AlphaFold's prediction gave CDK20 a shape for the first time, and that alone made it a target worth chasing. A similar story runs through neglected tropical disease research. The AlphaFold database now holds 19,036 predicted protein models from Trypanosoma cruzi, the parasite behind Chagas disease, whose protein space had been too structurally incomplete to support the kind of systematic, target-by-target drug discovery that better-funded diseases take for granted.
Using predicted structures to find binding sites and screen compounds at scale
Predicted structures don't just identify targets, they make large-scale compound screening faster and more accurate. Once a structure exists, predicted or otherwise, the next question is where on that protein a small molecule could grab hold, and whether anything in a compound library fits. Virtual screening answers this by docking huge numbers of candidate molecules against a predicted pocket and ranking how well each one sits.
TAAR1 shows what that looks like in practice. It's a G protein-coupled receptor with no experimentally solved structure, and researchers docked more than 16 million compounds against an AlphaFold model of it. The screen came back with a 60% hit rate, more than double the rate from the homology-model comparison run alongside it, and turned up 25 confirmed TAAR1 agonists. Hit rate isn't a vanity metric here. It decides how many compounds move into wet-lab testing afterward, and wet-lab testing is where the real money and time get spent. A screen that's twice as accurate cuts that downstream load roughly in half.
CDK20 shows up again here too. Pairing its AlphaFold-predicted structure with a generative AI design platform surfaced a small molecule with a measured binding affinity, a Kd, of 8.9 micromolar: the first confirmed CDK20 inhibitor on record, and reportedly the first case of an AlphaFold-predicted structure leading to a confirmed hit against a genuinely novel target in early discovery. Put the two cases together and a pattern emerges. AlphaFold opens the target, and it also makes hit-finding on that target doable inside a timeline that wasn't realistic five years ago.
AlphaFold 3's extension to complexes and what it adds to target validation
AlphaFold 2 was, at its core, a single-protein predictor. AlphaFold 3, released in May 2024, runs on a diffusion-based architecture built to model interactions between proteins, DNA, RNA, small molecules, and ions all at once, in a single prediction.
That matters because disease rarely involves a protein sitting alone. It involves a protein gripping a stretch of DNA to switch a gene on, or binding an RNA strand, or forming a complex with two or three other proteins that only makes sense together. AF3 handles that case. On the PoseBusters benchmark, a standard test of how well a model places a drug-like molecule into its binding site, AF3 hit 76% accuracy. Against static docking software, it does a better job on side-chain orientation, meaning it predicts how the pocket actually flexes to fit a molecule, not just whether the molecule landed in the right neighborhood.
The upgrade this gives researchers is subtle but real. Instead of asking "does this compound fit this protein," they can ask "does this compound fit this protein while it's bound to the DNA sequence it regulates," which is much closer to what's happening inside an actual cell. That said, AF3 is still better at freezing a single moment than showing the full sequence of events. Modeling how a protein flexes between conformational states over time remains a much harder problem, and plenty of targets are only druggable in one specific state out of several. Current predictions don't close that gap, and it's worth saying so plainly rather than glossing past it.
Where AlphaFold has opened targets that were previously out of reach
AlphaFold has made previously uncharacterizable proteins into genuine drug targets across disease areas, from cancer to cardiovascular disease to neglected tropical infections. Kinases rank second only to G protein-coupled receptors as a drug target class, and abnormal kinase activity drives a long list of cancers. AlphaFold has let researchers pursue kinase targets with structural hypotheses that support drug design across a wider range of the kinase family.
Cardiovascular disease gives a sharper example. Apolipoprotein B100, apoB100, is the central protein in LDL, the "bad cholesterol" every doctor warns about, and it had resisted structural characterization for decades despite being a long-standing target for atherosclerosis therapy. AlphaFold 2 helped reveal its complex architecture, handing structural biologists a blueprint against a condition responsible for a substantial share of global mortality.
Then there's the neglected tropical disease work. A 2025 bioinformatics study used AlphaFold-predicted structures for T. cruzi proteins to screen roughly 30,000 compounds computationally, tested 24 of them experimentally, and found two already-approved drugs, pimecrolimus and ledipasvir, with real antiparasitic activity against targets nobody had characterized before. None of that screen happens without a predicted structure to start from.
Line these three cases up and the common thread is access to targets that a structural blank spot had made functionally invisible, more than speed on familiar ground. Worth noting too: that access isn't spread evenly. Oncology and cardiovascular programs, backed by serious commercial funding, move fastest. Neglected disease work leans much more on the open public database than on any commercial platform.
How pharmaceutical companies are integrating AlphaFold into discovery programs
Major pharmaceutical companies have begun folding AlphaFold 3 into structure-based drug design, cutting how much they lean on experimental crystallography early on. AstraZeneca reports target drug design and validation running more than 50% faster in early programs where the tools got applied.
Isomorphic Labs, the DeepMind spinoff that now serves as the primary commercial steward of AF3, has struck partnerships worth $1.7 billion with Eli Lilly and $1.2 billion with Novartis, north of $3.6 billion combined, aimed at multiple drug programs, with first human trials of AI-designed oncology drugs expected sometime in 2025 or 2026. The investment pattern says something on its own: Isomorphic raised a $600 million Series A, and Xaira Therapeutics raised $1 billion, and neither company had a single clinical-stage asset at the time. That's money betting on target identification as a capability, placed on the platform well before any clinical proof exists.
Here's the part worth saying plainly: AF2 was freely available, and AF3's commercial use runs through Isomorphic Labs, and that has built something closer to a two-tier system. Academic labs work off the open database, while industrial programs increasingly need a licensing deal or a Lilly-sized partnership to get full access, and pretending otherwise doesn't help anyone plan around it.
What AlphaFold still cannot do in the target identification process
AlphaFold narrows a major bottleneck in drug discovery, but it leaves several critical gaps in the target identification workflow that wet-lab science still has to fill. A predicted structure is not a validated target, and that distinction is easy to lose in the excitement. Knowing a protein's shape says nothing about whether hitting it changes disease biology. That question belongs to cell assays, animal models, and eventually human trials.
Conformational dynamics remain a real limit. Plenty of proteins are only druggable in one specific shape, an "open" state versus a "closed" one, and AlphaFold generally hands back a static or narrowly limited picture rather than the full range of shapes a protein cycles through inside a living cell. Intrinsically disordered proteins are a related weak spot, and current AlphaFold models are noticeably less reliable for proteins that lack a stable fixed shape, which matters given how many disease-relevant proteins fall into this category.
Cellular context is another gap. AF3 improves complex prediction, but it doesn't capture the full biological context a protein operates in inside a living cell. And on the numbers: a 60% hit rate against a homology-model comparator's lower rate is a real leap, but flip it around and 40% of the top-ranked compounds in that TAAR1 screen still didn't pan out. Wet-lab validation isn't optional and isn't cheap, no matter how sharp the prediction gets. Access is its own limit too: researchers without a commercial license get a restricted version of AF3, especially around protein-ligand prediction. The academic release of model weights in November 2024 helped, but it didn't restore the fully open setup AF2 had.
The honest summary: AlphaFold removes one major bottleneck, structural ignorance, but it doesn't collapse the rest of the target identification workflow. It moves where the hard work sits, and the hard work itself stays substantial.
Where the AlphaFold-enabled target identification pipeline stands heading into late-stage trials
As of early 2026, the broader AI-assisted drug discovery pipeline has a growing number of drugs somewhere in clinical trials. Not all of them trace back to AlphaFold directly, but structural prediction sits underneath a large share of the target identification calls that got each program started. Isomorphic's planned 2025 to 2026 oncology trials will be the first direct test of whether AF3-based target identification and molecule design hold up in humans, which makes them a real checkpoint for the whole field, not just for one company.
The more interesting question isn't market size. It's the success rate coming out the other end: whether AI-assisted programs will clear Phase I and Phase II at better rates than the historical baseline for traditional drug programs remains an open question the coming trials will begin to answer. The group of AI-designed drugs that has actually reached Phase II is still small. There's a signal here, but it needs more data before anyone calls it settled.
What's changed, concretely, is where protein structure sits in the process. It used to be a bottleneck that quietly decided which targets were even worth considering. Now it's a resource researchers can query at the very start of a program, the moment they're still deciding what to chase. The ceiling on what counts as targetable has moved up, and the question left standing is which of those newly visible targets are actually worth the years and the billions it takes to push one through clinical development, a judgment no algorithm makes for anyone. Research programs built to prioritize well among a much longer list of candidates are in a better position to answer that, and answer it faster, than they were five years ago.

