Can England Map The Genome Of Every Living Thing With AI?

Lizards, whales, and bacteria have already solved problems medicine still fights. Their answers lie in DNA almost nobody has read. Britain has the labs, the sequencers, the AI, and a record of paying for costly moral causes. Reading all life on Earth should be the next one. Then giving it away free.

Share
Can England Map The Genome Of Every Living Thing With AI?

In 1992 John Eng, a biochemist at a veterans' hospital in the Bronx, was screening the venom of the Gila monster, a slow, beaded lizard of the American south-west. He found a molecule which mimicked a human gut hormone yet survived far longer in the blood. A synthetic copy called exenatide was approved in 2005 as Byetta, the first of the drug family now famous as Ozempic and Wegovy.

Gene editing exists because bacteria evolved a way to recognise and cut the DNA of invading viruses. The bowhead whale (which can live beyond 200 years) repairs broken DNA more accurately than human cells do, and researchers now want to know whether we can do it too.

Each of these finds depended on a scientist happening to study the right creature. Evolution has spent roughly four billion years testing chemical answers to infection, ageing, cold, hunger, and cancer, and it has written every result into DNA.

Humanity has read a sliver of God's brush strokes.

A national goal to sequence the genome of every species on Earth, and to teach artificial intelligence to learn from the result, would give medicine the largest library of tested answers and potential cures it has ever had.

Britain is better placed than almost any country to lead it.

It also has a history of paying for long, costly causes whose benefits flowed mostly to other people.

Our Great Cause can be understanding the natural world at its most intimate level yet, so we can program it to cure disease and save lives.

Four Billion Years of Nature's Experiments

A widely used estimate puts the number of animal, plant, fungal, and protist species at about 8.7 million, of which roughly 86% have never been formally described.

A 2016 study in PNAS put the number of microbial species as high as one trillion.

The Earth BioGenome Project, launched on 1st November 2018, aims to produce high-quality genomes for every known species of complex life. By September 2024 some 3,400 genomes from across all of science met its quality standard.

Its second phase, part of a plan to cover all 1.67 million named species by 2035, calls for 3,000 new genomes a month or more than ten times the current rate.

Measure Figure
Estimated species of animals, plants, fungi, and protists About 8.7 million (Mora et al.)
Share not yet formally described About 86%
Named species of complex life targeted by the Earth BioGenome Project 1.67 million (EMBL)
High-quality genomes of complex life publicly available, September 2024 About 3,400 (Earth BioGenome Project)
Estimated microbial species Up to 1 trillion (Locey and Lennon)
Species of complex life in Britain and Ireland About 70,000 (Sanger Institute)
Reference genomes released by the Sanger Tree of Life programme, August 2025 Over 3,000 (Wellcome Open Research)

The gaps matter because nature has already solved problems medicine still struggles with. The aforementioned bowhead whale carries vastly more cells than a human, lives far longer, and yet rarely develops cancer.

A team at the University of Rochester found a DNA repair protein in the whale at about 100 times the level seen in other mammals. Added to human cells, it improved their repair. Given to fruit flies, it lengthened their lives.

Elephants took another route, carrying extra copies of a tumour-suppressing gene.

Both discoveries came from years of work on one species. A complete library would let researchers work in the opposite direction, asking which creatures resist scarring of the organs, tolerate low oxygen, shrug off viruses, or regrow damaged tissue, and what their DNA has in common.

AI Learns Biology Only As Fast As It Is Fed DNA

In 2024 Demis Hassabis and John Jumper of Google DeepMind in London shared the Nobel Prize in Chemistry for AlphaFold, which predicts the three-dimensional structure of proteins. The Nobel committee noted it has been used to predict the structure of virtually all 200 million known proteins.

AlphaFold learned from the Protein Data Bank which contains half a century of structures shared freely by laboratories around the world. Jumper has credited public data as essential to its success.

DNA comes next.

In March 2026 Nature published Evo 2, a model from California's Arc Institute trained on about 9 trillion letters of DNA from every domain of life. Without being taught the task, it picks out harmful changes in human genes and classifies variants of the breast cancer gene BRCA1 with over 90% accuracy.

London's Basecamp Research has pushed further.

Its EDEN models, trained on about 9.7 trillion letters of DNA gathered by the company itself, designed working gene-insertion proteins for every disease-related site in the human genome its team tested, given nothing but the target location.

Of 33 antibacterial molecules EDEN designed, 32 worked against drug-resistant bacteria on the World Health Organization's most urgent list.

Glen Gowers, Basecamp's chief executive, says today's biological AI learns from a narrow slice of life on Earth. By one count, 80% of models built on gene sequences train on a single public database of fewer than 250 million sequences.

On 18th March 2026 the company launched the Trillion Gene Atlas, which aims to gather genetic material from more than 100 million species within two years, with sequencing from PacBio and Ultima Genomics, computing power from NVIDIA, and AI work with Anthropic (maker of Claude AI).

A complete library has become thinkable because the price of reading DNA has collapsed.

When Approximate cost
2003, first human genome About $3 billion (Sanger Institute)
Late 2010s $1,000 (Sanger Institute)
2022, at the Sanger Institute About $500 (Sanger Institute)
2024, launch of the Ultima UG 100 $100 (STAT)
2026, Ultima's advertised price $80 (Ultima Genomics)

Reading DNA has stopped being the hard part.

The difficult work now lies in finding specimens, naming them correctly, securing lawful access, and assembling complete genomes from the raw readings.

Hinxton, Oxford, and London Are Doing It

The Sanger Centre at Hinxton (Cambs.), founded in 1992 with Wellcome money, produced roughly a third of the final Human Genome Project sequence under John Sulston. Sulston and Wellcome also drove the 1996 Bermuda Principles, which led to the daily release of new human sequence into the public domain.

When the work was done, Sulston gloriously declared the genome open for any purpose, “without restraint or fee”. Few British exports have done more good.

The same institute now leads the Darwin Tree of Life project, which aims to read every one of the estimated 70,000 species of complex life in Britain and Ireland. The islands were chosen because their wildlife is probably the most thoroughly studied anywhere. The project passed 1,000 genomes in November 2023.

Its collaborators include the Natural History Museum, the Royal Botanic Gardens at Kew and Edinburgh, the Marine Biological Association, and EMBL's European Bioinformatics Institute, which stores and annotates the genomes on the same Hinxton campus.

Oxford Nanopore builds sequencers ranging from portable devices to benchtop machines, and its software uses neural networks, including transformer models of the kind behind modern chatbots, to turn electrical signals into DNA letters. In November 2024 the company and the government agreed to create an early-warning system for new pathogens across as many as 30 NHS hospitals.

On the human side, the Sanger Institute and Iceland's deCODE read 500,000 whole genomes for UK Biobank in a £200 million programme. Genomics England's Generation Study is sequencing 100,000 newborns, and the 10 Year Health Plan, published on 3rd July 2025, set an ambition to offer whole-genome sequencing to every baby in England within a decade.

Genomics plc, spun out of Oxford in 2014, combines roughly a million small DNA differences into risk scores for common diseases.

Nucleome Therapeutics, also in Oxford, maps which stretches of DNA outside genes switch particular genes on, and in which cells.

Few countries combine a national tree-of-life programme, population-scale human genomes linked to lifelong health records, a home-grown sequencing manufacturer, and leading AI laboratories.

It is extraordinary what our country can do when it is not wasting time on fraudulent sociology and toilet access for cross-dressing men.

On top, the money increasingly comes from elsewhere. On 8th June 2026 Google DeepMind and Google.org committed $5 million a year for five years to a Sanger consortium producing genomic data designed for training AI.

On the same day the government published its life sciences AI champion's adoption plan, which warned Britain's window to lead is closing “right now”. A generous gift from Google, and a pointed reminder of whose chequebook is open.

Brilliant Science, Hopeless Plumbing

The trouble starts, as one can imagine, where discoveries meet the British state. On 19th May 2026 Professor Sir Peter Donnelly, chief executive of Genomics plc, gave evidence to the House of Lords Science and Technology Committee.

UK Biobank charges a company like his £5,000 to £10,000 a year for data. Our Future Health, a cohort built with heavy public funding partly to strengthen British life sciences, charges £300,000.

Donnelly warned the gap risks pricing out young British firms while richer foreign rivals pay up, calling it potentially “a bit of an own goal”.

Genomics England turned down the company's renewed application for access because it also works with insurers. NICE, he said, had advised the firm to return only after randomised trials with ten years of follow-up, for a test designed to spot risk years before disease appears.

Genedrive, a Manchester company, has two NICE-recommended rapid genetic tests. One of them has already spared more than 30 newborns from deafness caused by a common antibiotic. Scotland rolled the newborn test out nationally.

In England, each trust or integrated care board must write its own business case, find its own money, and redesign its own services, while the next NICE decision on the newborn test is not due until July 2027.

The company's summary is hard to improve on: Britain has:

less of a discovery problem, more a deployment problem.

Mendelian, whose software searches health records for signs of undiagnosed rare disease, told the same inquiry it had scanned more than 10 million NHS records and flagged 1,500 patients for review.

The next step, moving from an anonymous finding to a call to the patient's GP, stalled inside NHS data environments built for research. Genetic results, it added, still reach doctors as PDF documents.

When NICE approved the gene therapy Luxturna in 2019, fewer than half of the estimated 86 eligible patients in England had even been identified.

In January 2025 the government's absurd Regulatory Horizons Council reported the Nagoya Protocol (an international treaty on access to genetic resources) was the one regulation causing trouble for most of the people it interviewed. It described a PhD student at the John Innes Centre who spent three years negotiating with a Vietnamese university for access to tropical trees; the talks collapsed and the project settled for a less suitable species.

An official review had found 100% compliance with the British regulations. The council suggested a likely reason: researchers were steering clear of any work the rules might cover.

Perfect compliance is easy when nobody tries.

In April 2026 de-identified UK Biobank data shared with three research institutions turned up for sale on an Alibaba marketplace. The charity closed its research platform and paused new applications until late 2026.

Any mission involving human DNA has to earn confidence before it can move fast.

Organisation What it has shown Where it stalled
Genomics plc GP trial in which 98.5% of patients found genetic risk results helpful £300,000 a year for Our Future Health data; renewed Genomics England access refused
Genedrive NICE-recommended newborn test; more than 30 babies spared hearing loss A separate business case in each part of England; next NICE decision due July 2027
Mendelian 1,500 possible rare-disease patients found in over 10 million records Research-only data rules blocked the step to contacting GPs
John Innes Centre student A lawful route to tropical tree samples Three years of talks collapsed
UK Biobank 500,000 whole genomes read Research platform shut after data surfaced for sale

$6,000 Against A Billion A Year

Most of the world's species live in tropical countries far poorer than Britain, and a mission to read them all must answer an old question: who gains? At Cali, Colombia, on 1st November 2024, governments under the Convention on Biological Diversity agreed to a global fund through which firms profiting from genetic data drawn from nature would share the proceeds.

Large companies should pay 1% of profits or 0.1% of revenue, and at least half the money must support indigenous peoples and local communities. The Cali Fund launched in Rome in February 2025.

By August 2026 it had received two contributions totalling $6,000, against the billion dollars a year its designers expected.

Both came from small British companies.

  1. TierraViva AI paid $1,000 on 19th November 2025, its chief executive calling the payment an “ice-breaker”.
  2. Prozomix, a British enzyme company, paid $5,000 in July 2026 without being required to.

Astrid Schomaker (executive secretary of the convention) offered a blunt explanation: “Nobody knows about the fund.”

Freedom of information requests show Defra contacted AstraZeneca in December 2024 hoping to recruit early contributors.

Two small British firms have now contributed more than every multinational drug company on Earth combined, which is to say, more than nothing.

Governments meet again in Yerevan in October 2026. Britain could arrive with something better than warm words: a pledge to build fair payment into any national sequencing effort, and a Treasury willing to match what small British firms contribute.

Basecamp already collects under commercial agreements with communities abroad.

A country which once paid to end a trade in human beings has standing to argue the people guarding the forests and reefs of poorer nations deserve a share of what their DNA makes possible.

Seven Commitments for the Next Ten Years

A serious national commitment would be measurable, and most of its starting points already exist.

Commitment Target date Where things stand
Reference genomes for all 70,000 species of complex life in Britain and Ireland 2032 Over 3,000 genomes across Sanger's Tree of Life projects by August 2025
British teams to read a fifth of the 1.67 million named species worldwide 2035 Sanger delivered a third of the human genome
Paying into the Cali Fund, with the Treasury matching small firms 2027 $6,000 received worldwide
One secure national route for approved AI tools to reach linked genomic and NHS records 2028 Research-only data environments; genetic results sent as PDFs
Our Future Health data for young British firms at UK Biobank prices 2027 £300,000 against £5,000 to £10,000 a year
Funded national rollout within a fixed period of any NICE recommendation 2027 A fresh business case in each part of England
Whole-genome sequencing offered to every newborn in England 2035 Generation Study sequencing 100,000 babies

Collection is the hardest step, and Britain has the institutions for it. Kew, the Natural History Museum, and the Royal Botanic Garden Edinburgh have spent generations gathering and naming specimens.

Their teams, equipped with portable sequencers and agreements which pay source countries fairly, could work beside scientists in the tropics as colleagues.

Every genome should come with notes on where the organism lives, what it eats, how long it survives, and what threatens it, because an AI can only find patterns in what it is given.

Predictions then need testing, and the government's own AI plan points to automated laboratories pioneered in Liverpool which run hundreds of experiments without human hands.

Non-human genomes should be released openly, in the Bermuda tradition. Human records need the opposite treatment: strict security, clear consent, and heavy penalties for leaks.

The same library would serve far more than medicine. The Earth BioGenome Project expects its genomes to help conservation, food security, and pandemic prevention, and farmers, park rangers, and public health teams would all draw on it.

Slavery abolition took years of parliamentary defeats and six decades of naval patrols. The Human Genome Project finished in 2003, more than two years early.

Reading the whole of life will take decades, and no country will manage it alone. A nation which once kept warships off Africa for sixty years, for people it would never meet, can spend a generation reading the living world for everyone.

Why? Because it's who and what we are.

Read more