HealthTechCrunch

Biotech Communicating Computational Drug Discovery to Investors

Computational drug discovery works best when you explain which bottleneck it actually solves.

Contributing Editor · · 11 min read
Cover illustration for “Biotech Communicating Computational Drug Discovery to Investors”
Computational Biology Drug Discovery · September 1, 2026 · 11 min read · 2,424 words

The vast majority of drug candidates that enter clinical trials fail, and each approved drug costs upward of $2.5 billion to bring to market. Computational drug discovery exists to fix the front end of that math, and communicating it well to investors means explaining exactly which part of the funnel it touches, not just gesturing at "AI" and hoping the valuation follows.

The full funnel runs through thousands of compounds in preclinical screening down to one approval. That attrition is not noise. It reflects specific, repeated failures in how candidates get chosen in the first place: wrong target, poor binding affinity, toxicity that shows up too late to matter cheaply. Computational methods enrich the pool that reaches the lab bench; they don't replace the bench itself. Any investor pitch that starts with the software before naming that problem has the order backward.

What computational drug discovery actually does, in terms investors can reason about

"Computational drug discovery" gets used as one label for at least five different jobs, and each one carries a different risk profile and a different timeline to prove itself.

Target identification is the search for which biological mechanism to hit at all, before any molecule exists. Virtual screening filters millions of existing compounds against a target computationally, cutting the number that ever touch a lab. Generative chemistry, or de novo design, goes further: instead of filtering what already exists, the model proposes molecular structures that have never been synthesized. ADMET prediction estimates how a compound will be absorbed, distributed, broken down, and cleared by the body, along with its toxicity, all before a single dose is given to an animal. Structure prediction, the AlphaFold lineage, lets researchers understand a target's shape before anyone solves it experimentally in a lab.

That last category picked up the closest thing biotech has to an Oscar: the 2024 Nobel Prize in Chemistry went to David Baker, Demis Hassabis, and John Jumper for computational protein design and structure prediction. That's not marketing copy; it's the Nobel committee, and it's worth citing plainly in investor materials because institutional validation of this kind is rare and hard to fake.

The performance numbers matter too, but only when tied to the right category. Research published in 2025 found that combining pharmacophoric features with protein-ligand interaction data lifted hit enrichment rates by more than 50-fold compared to traditional screening. That's a virtual screening result, and it says nothing about whether generative chemistry works. Pitch decks that blur the two invite the exact kind of question a founder doesn't want in a term sheet meeting: which one are you actually doing?

How the market's size and structure actually signal opportunity to investors — and where the numbers mislead

Market-size slides are where good pitches go to die, mostly because the numbers on offer don't agree with each other. The overall global drug discovery market, every method included, was valued at $71.96 billion in 2025 and is projected to reach $174.14 billion by 2035, a 9.24% compound annual growth rate. This is the baseline, the whole pie before anyone mentions AI.

The AI-specific slice gets murkier fast. One estimate puts the AI-in-drug-discovery market at $19.89 billion in 2025, climbing to $133.92 billion by 2034 at a 23.22% CAGR. Another puts it at $4.6 billion in 2025, reaching $49.5 billion by 2034 at a 30% CAGR. Same industry, same rough decade, and the two numbers disagree by a factor of four at the starting line. The gap usually comes down to scope: whether "platform value" and partnership deal flow get counted alongside actual software revenue, or just the software.

Here's the number that actually matters for underwriting a check: actual AI drug discovery software and platform revenue sits around $2 to 3 billion as of 2025 to 2026, against roughly $60 billion in cumulative investment tracked since 2019. That's a wide gap between money in and money out, and on its face it looks alarming. It isn't, necessarily, because most of these companies make money through milestone-laden partnerships and platform licensing, not through selling software licenses or shipping a product. An upfront payment, a per-program milestone, a royalty on downstream sales: that's the revenue model, and you need to name it outright rather than leave an investor to guess from a thin income statement.

Where investor capital is concentrating and what recent funding patterns reveal about conviction levels

Global VC funding for health-related AI peaked at $22 billion in 2021, then fell to $10.5 billion in 2024, according to PitchBook data reported by Axios. That's a more than 50% drop, and the obvious read is retreat. A better read is maturity: investors stopped funding the story and started asking for evidence the platforms actually worked.

Zoom out to biopharma VC broadly and the picture looks sturdier. JPMorgan characterized 2024 as a strong year at $26 billion across 416 rounds, up from $23.3 billion across 462 rounds in 2023. Fewer rounds, more dollars per round; bigger checks going to fewer bets, which is its own signal about where conviction concentrates.

AI-native biotechs have recently commanded a valuation premium approaching 100% over biopharma peers more broadly, per PitchBook. Investors are pricing in platform optionality, the idea that one platform might generate several drug candidates rather than just one, not merely the pipeline sitting in front of them today. Over a recent 12-month stretch, VCs put $3.2 billion into 135 AI drug development startups, and by October 2025 the ecosystem had grown past 530 companies worldwide focused on AI-powered drug discovery. This is no longer an experiment; it's an industry layer with its own conference circuit and its own competitive dynamics.

Some of the individual rounds are almost cartoonishly large. Xaira Therapeutics launched with a $1.0 billion debut round and no prior financing history at all, the biggest first check of its kind in the space. Chai Discovery raised a $30 million seed in September 2024, then a $130 million Series B by December 2025, three rounds totaling $230 million in fifteen months. Eikon Therapeutics and Generate:Biomedicines both went public in February 2026, raising $381 million and $400 million respectively. After a historic low of just 8 biotech IPOs in 2025, the market looks to be turning a corner in 2026, with AI widely cited as a factor improving how investors calculate risk on these deals.

Big checks are getting written, but it's worth asking what each of those investors demanded to see before signing. Evidence of efficacy is increasingly the gate, and the number on the term sheet only tells half the story.

What the clinical pipeline data actually shows about where AI's advantage holds and where it doesn't yet

Diagram: AI Drug Discovery: Phase 1 Win, Phase 2 Gap. Visualizes: Visualize the split-level performance story of AI-discovered drugs versus the industry historic average across two clinical phases.

As of April 2024, eight leading AI drug discovery companies had 31 drugs in human clinical trials: 17 in Phase I, 5 in Phase I/II, and additional candidates in later stages. The sample is small, but it's real and growing. The pipeline of AI-originated molecules entering human trials has continued to grow across multiple disease areas.

Now the number that should open every honest pitch on this topic: analyses circulating in the field have suggested Phase 1 success rates for AI-discovered molecules may run well above the industry's historic average. Those results, where they hold up, are genuinely strong and deserve to be cited plainly and without hedging.

Then comes the harder number. Phase 2 success rates for those same molecules have appeared roughly in line with the industry's long-standing historic average. Emerging research has found something similar: AI-discovered drugs appear to fail in Phase 2 at rates comparable to non-AI drugs. So AI gets cleaner molecules through the safety gate at Phase 1, but it has not yet shown, at scale, that it can predict whether a drug will actually work in humans, which is precisely the stage where most of the industry's money and time get destroyed. This is the honest shape of the data, and it's more interesting than a simple success story because it draws a clean line between the problem AI has solved and the one it hasn't.

There are specific proof points worth naming directly: (i) at least one company has reported positive Phase 2 topline results for a candidate it describes as the first drug discovered and designed through an end-to-end generative AI process to reach that stage, with the program advancing toward later-stage trials; (ii) physics-based computational design has produced candidates reaching late-stage trials, offering examples of the approach working beyond early development; and (iii) some platform companies have used their tools to identify therapeutic angles in diseases with no approved therapies, translating computational findings into early clinical signals.

How leading companies frame their platforms in investor materials — what works and what invites skepticism

Here's the practical problem you'll face underneath all of this: investors tend to be fluent in either deep tech or healthcare, rarely both at once. Explaining how your proprietary model differs from an off-the-shelf large language model, in terms that land with someone who knows biotech cold but treats "transformer" as a car part, is one of the harder tasks you will face in fundraising.

Overpromising has real costs here, and they're visible in the public record. Reporting in 2025 flagged that claims made by several AI drug discovery companies about their AI-generated drug candidates drew scrutiny from the scientific community. It's not a hypothetical risk sitting somewhere off in the future; it already happened, in public, to named companies, and it's a preview of the kind of question every platform will eventually face once results are in hand instead of promised.

Contrast that with companies anchoring their story in numbers a spreadsheet can hold. Recursion Pharmaceuticals frames progress around measurable partnership milestones: publicly disclosed program payments with large pharma partners, with stated cash runway extending well into the future. That's a scoreboard, not an abstraction. Some AI drug discovery companies have anchored public listings and financing narratives around disclosed revenue figures and upfront licensing payments from larger partners, using deal economics as a stand-in for platform credibility. Several computational drug discovery companies have secured FDA regulatory designations for specific programs, milestones that function as third-party validation nobody inside the company had to write themselves.

Then there's the consortium signal, which might be the quietest but most telling data point in this whole section. By October 2025, multiple large pharma companies have joined open consortia pooling proprietary structural data to advance shared computational infrastructure. When major incumbents agree to share data infrastructure in public, that's not charity. It's an implicit bet on the approach, and it reads to investors as validation none of the AI companies could have generated on their own.

The pattern across every example that works: concrete milestones, third-party validation, deal economics. Not capability claims sitting alone on a slide with no receipt attached.

The specific translation moves that convert technical claims into investor-ready language

Six moves, in rough order of how they should appear in a deck.

First, lead with the attrition problem, not the technology. Every capability gets introduced as a fix for a named failure mode: late-stage toxicity surprises, weak target validation, slow hit identification. Name the wound before naming the bandage.

Second, distinguish the specific tool category from the catch-all "AI in drug discovery" label. Generative de novo design, physics-based simulation, and ADMET prediction carry different risks and different timelines, and conflating them in your pitch is how you can end up fielding the wrong question in the room.

Third, map every capability to the metrics investors already track: probability of success and cost per candidate. The BCG data point, an 80 to 90% Phase 1 success rate, is a quantified probability-of-success lever. The 50-fold hit enrichment figure from the 2025 research translates directly into fewer dead-end compounds reaching the lab bench, which is a cost-per-candidate lever. Say it in those terms, not in terms of "our model is really good."

Fourth, let milestone economics carry the narrative. Upfront licensing fees, per-program milestones, royalty structures: these are the units investors already know how to value. Insilico's substantial upfront payment from Exelixis and Recursion's nine-figure cumulative Sanofi milestones are templates for exactly how to present this kind of progress.

Fifth, name the Phase 2 efficacy gap before anyone else brings it up. Sophisticated investors doing diligence will find the BCG and Nature Medicine data regardless. Teams that surface it first, and frame it as a known and bounded risk rather than a surprise buried in a footnote, come across as more credible than teams that dodge it. Something close to: strong evidence exists for Phase 1 enrichment; efficacy prediction at Phase 2 remains an open problem across the whole industry, and here's the specific work underway on it.

Sixth, lean on third-party anchors instead of self-reported claims. A Nobel Prize, a pharma consortium membership, an FDA designation, a peer-reviewed publication: each one is an outside party vouching for the work, which can matter more than any adjective you could choose.

What the next two years of clinical readouts will demand from computational drug discovery narratives

The pipeline is thickening at exactly the stage that will settle this argument. With multiple AI-originated molecules now moving into Phase 2 and Phase 3, 2026 and 2027 will produce the first substantial batch of late-stage clinical readouts specifically for AI-designed drugs. This is the test everyone in this space has been waiting to see run.

If Phase 2 success rates stay near their historic average, which is what the current data suggests they'll do, any investor narrative that promised transformative efficacy gains rather than transformative Phase 1 enrichment will face direct scrutiny. This isn't a hypothetical risk invented for the sake of a tidy conclusion. STAT News already documented one instance of scrutiny landing on named companies in 2025, and readouts over the next two years will generate plenty more chances for that pattern to repeat.

The valuation premium AI-native biotechs are commanding right now, close to 100% over non-AI peers, is a forward bet that Phase 2 performance eventually pulls ahead of the historic baseline. Nothing in the current data confirms that bet yet; it's priced in as a hope, not a result. Narrative discipline, saying exactly what's proven and exactly what isn't, is what protects that premium while the readouts roll in. Overclaim now, and a mediocre Phase 2 result reads as betrayal. Underclaim, price the risk honestly, and the same result reads as exactly what was promised. Same data, wildly different reception, and the only variable is what got said in the deck eighteen months earlier.

Sources

  1. intuitionlabs.ai
  2. biospace.com

More in Computational Biology Drug Discovery