The Most Valuable Data in Pharma Is the Data Nobody Keeps

Negative results are the missing fuel for AI drug discovery. The companies that win will be the ones with the highest failure-capture rate — and the machine they are feeding wants something the industry was never built to give

For a century, an experiment that did not work was a dead end. In the age of machine learning, it may be the most useful thing a laboratory produces — and the one asset the industry has spent a hundred years throwing away.

19-minute read

A library written by the survivors

Every model is a mirror of its training data, and the training data underneath artificial intelligence in drug discovery has a quiet defect. It is written almost entirely by the winners. A medicinal chemistry paper announces the compound that bound, the reaction that ran, the formulation that held. The far larger record — the compound that did nothing, the synthesis that collapsed, the candidate quietly abandoned in month nine — is rarely written down and almost never published. “Compound X showed no activity against target Y” is not a sentence that journals compete to print.

The bias is measurable, not merely cultural. The public bioactivity databases that supply much of computational drug discovery are themselves skewed toward what worked. In ChEMBL, one of the most widely used, an analysis found that at a common activity threshold almost 90% of the extracted compounds are labelled “active” — a proportion no bench scientist would recognise as real. The map of pharmacology is being drawn almost entirely in the colour of success.

The result is a corpus that looks like knowledge but behaves like a hall of fame. A model trained on it learns what success looks like in fine detail and has almost no idea what failure looks like at all. Reviews of the field now name this publication bias as the most structurally damaging limitation in AI for drug discovery, not a footnote to it. And the practitioners feel it: in surveys of technology leaders, 68% point to poor data quality and governance — not algorithms, not compute — as the main reason their AI initiatives fail.

The scientific literature is a hall of fame. A model needs the whole league table.

Why a model that never fails is a dangerous one

In most domains, success and failure are roughly balanced and the imbalance is forgivable. Drug discovery is not most domains. Of the candidates that reach a first human trial, around 90% never make it to approval. The reasons are well mapped: a lack of clinical efficacy accounts for 40–50% of failures, unmanageable toxicity for about 30%, poor drug-like properties for 10–15%, and weak commercial or strategic logic for the rest. Every approved medicine now carries an estimated development cost of roughly $2.3 billion — a figure that exists largely because it is shared out across everything that failed on the way.

That is the asymmetry that should worry anyone building a model. In this industry the negative result is not the exception; it is the overwhelming base rate. A system trained only on the winners is therefore optimised on the rarest 10% of outcomes and is, by construction, blind to the 90% that defines the real terrain. It can describe the summit beautifully and has never been shown the cliffs. Failure data is what draws the boundary of where things stop working — and a model that cannot see the boundary will report confidence exactly where caution is most needed.

The proof that failure teaches

This is not a thought experiment. In 2016, a team led by researchers at Haverford College and Purdue published a result in Nature that has aged into a manifesto. They took thousands of “dark” reactions — hydrothermal syntheses that had failed and were sitting unloved in laboratory notebooks — and trained a machine-learning model on them. Asked to predict the conditions under which entirely new crystals would form, the model succeeded 89% of the time. Trained chemists, relying on intuition built over careers, managed 78%.

The lesson is not that machines beat people. It is that the information chemists routinely throw away was worth more than the information they kept. A failed experiment is not noise to be discarded; it is a precisely labelled example of where the edge of the possible lies. Once that is taken seriously, the discard pile stops looking like waste and starts looking like inventory.

An experiment that failed is a labelled example. Failure, in other words, is an asset class.

The number almost nobody measures

If failure is inventory, then a laboratory can be judged by how much of it survives the day. Call it the failure-capture rate: the share of everything an organisation tries — including, and especially, everything that does not work — that is recorded cleanly enough for a model to learn from. Not filed in a notebook a colleague cannot find. Not summarised in a slide. Captured in structured, machine-readable form, with the conditions, the negative result and the context intact.

For most of the industry that number is close to zero, and almost no one measures it. It is about to become one of the most consequential figures in pharmaceutical research, because in a world where the model is the engine of discovery, a high failure-capture rate is simply a larger, less biased training set than a rival can assemble. Two companies can run the same experiments and own radically different assets, depending only on what they bothered to keep. The cleverness has quietly migrated from the chemistry to the bookkeeping.

Why the discard pile got discarded

If failure is so useful, why has it been thrown away for so long? Because every incentive in science has pointed the other way. Journals reward novelty and positive findings, and a paper reporting that nothing happened is a hard sell; there are strong rewards for publishing the active compound and almost none for the inactive one. Patents and competitive secrecy make a company’s list of dead compounds among its most sensitive documents, because the negatives reveal precisely where rivals have already searched. And a culture that prizes the clean result quietly discourages anyone from cataloguing the messy near-miss. None of these incentives is irrational. Together they produced a century-long habit of paying in full for failure and then deleting it. What has changed is not a sudden realisation that failure matters — chemists always suspected it did — but the arrival of a system, the model, that can finally make use of it, and an economics that finally rewards keeping it.

Three ways research is being rebuilt to raise the rate

Lifting a failure-capture rate from near-zero is not a filing exercise; it is a research-design problem, because the existing machinery of science was built to publish success, not to record failure. A small number of organisations have understood this and are re-engineering how experiments are run. Three patterns stand out.

Capture everything by default. The self-driving laboratory closes the loop between an algorithm that designs an experiment, a robot that runs it, and a model that learns from whatever comes back — success or failure, recorded identically. Novartis rebuilt a high-throughput chemistry system into a closed-loop platform called MicroCycle that autonomously synthesises compounds, purifies them, assays them, and chooses what to make next, with every intermediate result logged. Pfizer has installed autonomous laboratories of its own. The point is not only speed; it is that a machine has no incentive, and no ego, that makes it forget the runs that did not work.

Generate your own unbiased data at industrial scale. Rather than inherit the literature’s blind spots, Recursion runs its own experiments by the million — capturing on the order of millions of cell-based measurements per week through robotics and computer vision, and accumulating one of the largest proprietary biological datasets in the world, reported at more than 50 petabytes. Crucially, those experiments are designed to map biology completely, which means the non-effects are recorded with the same rigour as the hits. It is the difference between a record built to be published and a record built to train a model.

Pool the failures without surrendering the secrets. The richest negative data sits locked inside companies that would never share it, because the list of compounds that did not work is competitively sensitive. The MELLODDY consortium engineered a way around the deadlock: ten rival pharmaceutical companies — Amgen, Astellas, AstraZeneca, Bayer, Boehringer Ingelheim, GSK, Janssen, Merck KGaA, Novartis and Servier — trained a shared model across more than 2.6 billion confidential activity measurements on over 21 million molecules, using federated learning and a blockchain ledger so that no raw data ever left its owner. The shared model outperformed any single company’s own. They called it coopetition; it is also the first proof that an industry can pool its collective failures without exposing them.

A parallel movement is attempting the same thing in the open. The Open Reaction Database, launched in 2021, built a shared, public schema designed deliberately to capture the failed and low-yielding reactions that journals leave out — precisely because their absence, in its founders’ words, gives chemistry an overly optimistic view of itself. The common thread is an admission that the negative result has been the most under-collected asset in science, and that the tools to exploit it have at last caught up with the cost of throwing it away.

The failure moat

Here is the uncomfortable corollary the open-science movement would rather not dwell on. The most valuable negative data is also the most competitively sensitive. A list of targets a company has quietly abandoned is a map of where not to waste a decade — which is exactly why no firm will donate it. So the open repositories, for all their virtue, will tend to fill with low-stakes failures, while the high-value ones stay locked inside the organisations that can afford to generate them at industrial scale.

Failure-capture, in other words, rewards incumbency. The companies with the robots, the petabytes and the federated alliances will compound an advantage that has little to do with better chemistry and a great deal to do with better bookkeeping of their own mistakes. If the last era’s moat was the patent, the next era’s may be a proprietary archive of what did not work — harder to invent around, because a competitor cannot even see it. A field congratulating itself on democratising discovery could, through this one mechanism, concentrate it in fewer hands than before.

The next moat in pharma will not be a patent. It will be a private library of failure.

What the machine would ask for

For all the talk of feeding the machine, the industry rarely stops to ask the machine what it is hungry for. Pose the question seriously and the answer is unsettling, because a learning system and a pharmaceutical company optimise for almost opposite things. A company is built to maximise the probability that each experiment succeeds, publishes or patents. A model is built to maximise information — and information is richest precisely where the outcome is least certain. The experiment a seasoned chemist skips because it will obviously fail is, to the model, the most valuable hour in the building.

Asked to redesign pharmaceutical research for its own learning, a model would make demands no human research culture has ever rewarded. It would request the coin-flips — the experiments it rates at even odds — and refuse the safe bets, because a success it already predicted teaches it almost nothing. It would ask for the boundaries: the marginally soluble, the nearly toxic, the almost-druggable, where the cliffs of biology actually sit. It would want failures recorded with the same fidelity as triumphs — every concentration, the controls, the noise floor — not the single tidy number that survives into a figure. And it would be supremely indifferent to commercial value: a target with no market is, to a model, as informative as a blockbuster — so it would happily explore the unfanciable corners of biology that human portfolios never touch, and to which its map is therefore blind.

The experiment a chemist skips because it will obviously fail is, to the model, the most valuable hour in the building.

That indifference is the catch. A model set loose to maximise its own learning would rationally spend a research budget on questions of no near-term use to any patient, because that is where its uncertainty — and so its appetite — is greatest. Hand it the research agenda and you have quietly swapped one objective function, cure this disease this decade, for another, reduce my uncertainty about biology in general. The two often point the same way; when they diverge, someone has to decide which one wins, and that decision — not the algorithm — is the real frontier. The machines being built to learn from failure will, if allowed, begin to choose which failures are worth having. The governance question of the next decade is not how to make the model cleverer. It is who owns its curiosity.

What changes next

Publishing goes first. If the model’s appetite reshapes what is worth doing, it reshapes the institutions built around the old appetite. The unit of record shrinks from the paper — a narrative of success — to the result, positive or negative, deposited in machine-readable form with its own identifier. Scientists begin to be cited for the failures others learned from, and a “negative h-index” stops being a joke. Pre-registration becomes the default rather than the exception, because a registered failure, free of the file drawer and the post-hoc story, is the cleanest training label there is.

Incentives are rewritten. Capturing failure stops being a virtue and becomes a requirement. Funders turn the data-management plan into a failure-management plan and hold back the final tranche until the negatives are deposited. Companies pay scientists for well-logged dead ends the way they reward patents now, and the annual review counts data contributed alongside compounds advanced. The most radical version is a data dividend: provenance tracking that pays a researcher a royalty when the failure they recorded helps train a model that yields a drug.

The scoreboard inverts. Success rate has always been the number to maximise; the model turns it over. Because a decision boundary needs far more clean negatives than positives to locate — plausibly many times more — a laboratory that fails more, and captures it well, may train a better model than one that fails less. “Optimal failure rate” becomes a design parameter rather than a flaw to drive toward zero, and negative-by-design experiments — studies run because they are expected to fail informatively — become a respectable use of a budget for the first time in the history of the discipline.

Even the patent feels it. The claim drifts from the molecule to the model and the dataset behind it, and “data exclusivity” becomes the moat that matters — except that a patent demands disclosure while the most valuable failures derive their worth from being unseen, so the best negatives will stay trade secrets, un-patentable and un-inspectable by design. Which leads to the bet this article will stand or fall by: by 2030, a company’s negative-data corpus will be a line item in its deals — diligenced in licensing, valued in acquisitions, protected like any other asset. The first time a partnership is priced on the quality of a partner’s failures rather than the promise of its successes, the transition will be complete. Failure will have a market price.

The method itself is rewritten. For a century, the design of experiments — Fisher’s discipline — has meant arranging the fewest runs needed to estimate an effect a human can interpret and publish. When the consumer of the result is a model rather than a person, the objective inverts: each experiment is chosen not to confirm a hypothesis but to reduce the model’s uncertainty by as much as possible — which often means running precisely the conditions a researcher would judge least worth their time. Machine learning already has names for the mechanism — active learning, optimal experimental design — but as a research philosophy it deserves its own. Call it design for learning. Its deliverable is not a conclusion; it is a labelled example. The experiment stops being the end of the enquiry and becomes one data point in a training set that never stops growing.

Design of experiments was built to end an argument. Design for learning is built to feed a machine that never stops asking.

What it does to the university

Nowhere will the shift land harder than in the university, because the academy is built almost entirely on the opposite of what is becoming valuable. Start with the doctorate. A PhD is, by tradition, a monument to a single positive finding — and every year thousands of them quietly fail because the project “did not work,” three years of negative results binned and a career stalled with them. Reverse the value of failure and the cruellest waste in science is redeemed: a thesis that produces a complete, clean, machine-readable map of what does not work becomes a successful one. The viva begins to examine what a candidate captured, not only what they confirmed, and the failed PhD, as a category, starts to disappear.

Pushed to its conclusion, the doctorate itself is redefined. If the most valuable output of three years at the bench is a clean, complete, machine-readable record of what was tried and what happened, a growing share of PhDs become, in effect, dataset-creation projects — the candidate judged less by the single result they proved than by the coverage, quality and reusability of the data they leave behind for a model to learn from. The thesis as argument gives way to the thesis as corpus.

The training inverts with it. Students taught to design experiments likely to yield a publishable positive must learn the opposite craft — designing experiments that maximise information, the failures included — and the bench scientist becomes, in part, a data engineer. Funding feels it next, at the most sensitive point of all: accountability. Grants are awarded on promised breakthroughs and renewed on papers, but once captured failure has value a funder can begin to buy information yield — uncertainty reduced per pound — and a grant may count as a success even when its hypothesis was wrong, provided the failure was deposited cleanly. Dedicated null-result funding lines, today a curiosity, become a serious instrument, and research-assessment exercises that reward high-impact papers will have to learn to credit a reused negative dataset as impact in its own right.

And the university gains a new asset to sell. Alongside patents and spinouts, an institution’s archive of what did not work becomes licensable — and the first university to place a negative-results dataset under a commercial licence will quietly reprice the value of its own “failed” research. Reputation follows the asset: appointment and tenure committees begin to ask for a researcher’s failure-capture rate and their most-reused negative dataset, not only their citation count. There is even a cure hidden here for the replication crisis, whose root cause — the file drawer — is precisely what capturing negatives by default removes. The risks are real and worth naming: institutions without industrial capture infrastructure could be relegated to producing cheap data for richer models, and any metric that rewards failure will tempt people to manufacture cheap, uninformative failures to hit it. But the direction is set. The apparatus of science — the thesis, the grant, the journal, the prize — was built to reward being right. It is about to discover that the machines need it to be usefully wrong, and that almost none of those institutions yet know how to pay for that.

That asset will not stay a by-product for long. The logical end point is the university as a data vendor — institutions standing up dataset-and-model product lines the way they once built technology-transfer offices, packaging their captured failures into training corpora and selling not only the data but the models trained on it. A spin-out whose entire product is a proprietary dataset and the predictor built from it is no stranger than one built on a molecule; it is arguably more durable, because a dataset compounds while a patent expires. The richest research universities could become to biological data what foundries are to silicon — the places that manufacture the substrate everyone else builds on, and charge accordingly.

The cruellest waste in science — the unpublishable thesis — is about to become an asset.

What it does to the regulator

The regulator is already moving in the same direction, and for the same reason. In January 2025 the US Food and Drug Administration issued its first draft guidance on using AI to support regulatory decisions for drugs and biologics, built around a risk-based assessment of a model’s credibility for a defined “context of use.” What it proposes to examine is the tell: not only the model, but the data used to train it and the governance around it. The European Medicines Agency adopted a parallel reflection paper in September 2024, spanning the whole lifecycle from discovery to post-authorisation. The centre of regulatory gravity is shifting from the molecule to the model — and to the data beneath it.

Follow that logic and the questions a reviewer will start to ask come into focus. If a model’s credibility depends on how faithfully its training data represents reality, the regulator must eventually ask where the negatives are — whether the dataset is the survivor-biased hall of fame the literature has always been, or something closer to complete. Expect the regulatory cousin of the failure-capture rate: a requirement to characterise what a model was shown not to do, because its safety is defined at the boundary it was tested against. Expect, too, the qualification of datasets — agencies already qualify biomarkers and novel methodologies, and the natural next step is a sanctioned reference corpus, negatives included, that a model must be measured against. Approval itself narrows: a model is credible for a context of use, not in general, so “an approved model” comes to mean one qualified for a specific job rather than cleared for all of them.

Approval stops being a verdict on a molecule and becomes a verdict on a model — and only for the job it was tested to do.

Two further shifts follow. Because these models keep learning as new failures arrive, regulators need a way to permit updates without re-reviewing every version — the predetermined change control plan, the primitive the MHRA has been testing in its AI Airlock sandbox for medical devices and which will transfer directly to drug-development models. And because a model is only as trustworthy as its lineage, expect demands for provenance and audit — which experiments, which versions, no leakage between the data used to train and the data used to test — alongside continuous post-approval monitoring for drift: a pharmacovigilance not of the drug, but of the model that helped approve it.

The deepest consequence is one the agencies have not yet reckoned with in public. A regulator sits on every sponsor’s submitted data, and increasingly on the failures buried within it. An agency that aggregated the whole industry’s negatives could build the most capable predictive model in the field — and could require sponsors to benchmark against it. The referee would then hold the best model on the pitch. That may be the most powerful position this entire shift creates, and the one with the least settled answer to a simple question: who governs the body that governs the models? The institution that once filed the failures away may become the one that most depends on them — and the one with the most to decide about who else is allowed to learn from them.

The record of being wrong

Every laboratory will tell you it values learning from failure. Almost none can tell you its failure-capture rate, which is the same claim made honest. As models become the engine of discovery, the durable advantage will not belong to whoever runs the cleverest algorithm or owns the most successes; those are the easiest things in the system to copy. It will belong to whoever keeps the most complete, most trustworthy record of being wrong — which is, not by coincidence, exactly what the machine has been asking for all along.

For a hundred years the failed experiment was the one thing a laboratory was sure it could throw away. It is turning out to be the most valuable data in pharma — and the question for every research organisation is no longer whether that is true, but how much of it they have already lost.

Sources