Most of biology's data is public, which is a problem if you want to build a business on it. Basecamp Research, a London startup founded in 2019, has spent the last five years on expedition, building what it calls the world's largest proprietary protein sequence database from samples collected in 31 countries [PRNewswire]. Its bet is simple: the best way to design a novel protein for an industrial catalyst or a new drug is to train an AI on nature's own, vast, and previously uncatalogued blueprint. The company just raised $60 million in a Series B led by Singular to prove it [CB Insights].
The data moat is a field kit
The core of Basecamp's operation isn't just software; it's logistics. While others fine-tune models on public data, Basecamp's team collects novel biodiversity using off-grid DNA sequencing technologies in remote ecosystems [PRNewswire]. This generates BaseData™, a knowledge graph the company says now contains over one million new species [Astrobiology.com, June 2025]. The commercial wedge is that this data is proprietary and context-rich, tagged with environmental and evolutionary metadata absent from public repositories [EquityZen].
Training a GPT-4 for genes
On top of this growing dataset, Basecamp has built EDEN (Environmentally-Derived Evolutionary Network), a suite of biological foundation models. The flagship EDEN-28B model is a "GPT-4-scale biology model" trained on 9.7 trillion nucleotide tokens and 10 billion novel genes [basecamp-research.com]. The most audacious technical claim is for BaseFold, a model that reportedly outperforms DeepMind's AlphaFold 2 in predicting large, complex protein structures and their interactions with small molecules [TechCrunch, Oct 2024].
The founding team blends scientific ambition with operational pragmatism. Co-founders Oliver Vince and Will Pelton started the company in 2019 [Rory Cellan-Jones]. Vince, the CEO, studied physics at Oxford and worked in finance. Pelton brings the technical and biological depth. A third name, prominent computational biologist Oliver Stegle, was listed as a significant controller at incorporation but ceased that role in early 2020 [Companies House]. The team has since grown to between 11 and 50 employees [LinkedIn].
The partnership path to revenue
Basecamp's business model is partnership-driven, focusing on R&D collaborations. It has announced deals with industrial leaders like Johnson Matthey and a biodiscovery program with the government of Malawi [Redalpine] [GlobeNewswire, 2025]. The company generates revenue by providing partners with access to its knowledge graph and the AI-designed sequences it produces [Dealroom]. The recent $60 million Series B, following a $20 million Series A in late 2022, provides a long runway to convert these research partnerships into larger, recurring commercial agreements [Crunchbase] [UK Tech News, 2022].
Where the model could misfold
The ambition is clear, but the path is paved with non-technical challenges:
- The verification gap. The claim that BaseFold outperforms AlphaFold 2 remains a company-sourced benchmark [TechCrunch, Oct 2024].
- The benefit-sharing burden. Basecamp emphasizes equitable frameworks with data source countries, which adds operational complexity [PRNewswire].
- The commercialization clock. The company must demonstrate that its AI-designed proteins consistently outperform those found through traditional methods in real-world industrial settings.
The unit economics of biodiscovery hinge on how many of the 10 billion genes can be matched to a partner's multi-million-dollar problem. The incumbent Basecamp must beat isn't another AI startup; it's the entire, entrenched, and expensive practice of directed evolution in a corporate lab.