Drawing on Metix AI's 860 million+ global talent pool, a full-coverage profile of 35 AI-drug-discovery companies and big-pharma AI groups: who holds the ML talent that can fold proteins, where the scarce dry-lab × wet-lab crossover profiles sit, where the AlphaFold network is spilling over to, an engineer-level org map of big-pharma AI groups.
The figures below follow Metix AI database methodology (data through H1 2026). The population = talent currently employed at the 35 target companies and based in the United States, United Kingdom, Switzerland, Denmark, France, Germany, Canada, or Sweden. For big pharma, the scope is limited to the AI/computational subset.
The AI/computational groups at 15 big-pharma companies total 3,332 people, 3.2x the 1,042 across the 20 AI-native startups. The single largest AI group is Genentech (442 people, including the Prescient Design antibody-design team). For buyers: to source AI-pharma talent at scale, the main battleground is the computational groups inside incumbent pharma, not just the marquee startups.
55.6% of computational talent has both a wet-lab background (biology/biochemistry) and a dry-lab one (ML/programming), confirming this field is inherently interdisciplinary. But the truly scarce profile is the specialist in protein structure/design: across the full sample, only 125 people (2.9%) work explicitly on protein structure prediction or protein/antibody design. These 125 are the most fought-over core profile of the AlphaFold era.
Of the visible profiles, 56 have a DeepMind / Isomorphic history; 49 of them are now at Isomorphic (a wholesale DeepMind spinout), with a few spilling over to Xaira, Latent Labs, and insitro. Isomorphic's current technical talent has a median tenure of just 11 months, and 54.1% have been there under a year, a fast-scaling "new guard."
Unlike pure AI (where the Bay Area dominates), AI-pharma talent is spread across the US (571 Bay Area / 478 Boston), the UK (AstraZeneca and others), Switzerland (257 people, Roche/Novartis), and Denmark (169 people, Novo Nordisk). Cross-region, cross-time-zone sourcing is the norm in this space.
The points below were verified one by one against 2025-2026 public sources (full sourcing in the research memos), keeping only the facts that affect talent decisions. Dollar figures follow the reported amounts.
AI drug discovery drew roughly $11 billion in VC in 2025 (348 rounds, per DealForma). Isomorphic raised another $2.1 billion in May 2026 (the second-largest round in biotech history); Chai Discovery raised $130 million in December 2025 at a $1.3 billion valuation; Xaira launched with $1 billion and a ~$4 billion valuation; Generate Biomedicines and Eikon both IPO'd in February 2026 (raising roughly $780 million combined). The funding environment is clearly recovering.
Structure prediction (closed-source AlphaFold 3 → an open-source surge of Boltz-1/2, Chai-1, ESM3) → de novo protein/antibody design (RFdiffusion, Chai-2 with a hit rate near 20%, roughly 100x older methods) → molecular generation (Boltz-2 jointly modeling structure + affinity, approaching FEP but a thousand times faster) → virtual cell / phenotypic world models (Arc State, CZI, Xaira). Each stage maps to a new scarce talent profile.
Anthropic acquired Coefficient Bio (a team from Genentech's Prescient Design) for $400 million in April 2026; EvolutionaryScale was acquired by CZ Biohub in November 2025 (Alex Rives became Biohub's head of science); NVIDIA is deeply tied to Lilly/Roche/CZI, and OpenAI invested in Chai. Proprietary biological data has become the new moat for general-purpose AI companies, directly intensifying the fight for ML people who understand biology.
As of the end of 2025, not a single AI-discovered drug has been approved. The first generation of AI biotechs is contracting: BenevolentAI laid off staff and saw its valuation fall to about $145 million; Atomwise reorganized into Numerion in October 2025; Recursion cut its pipeline to control costs after merging with Exscientia. Massive funding for the new entrants coexisting with no approved drug to show for it yet is the real state of this space today.
Population = 4,374 computational/ML professionals (protein structure/folding, protein/antibody design, molecular generation/chemistry, target/systems biology, ML platform, and computational/ML science).
Reading: the single largest AI/computational group is Genentech (442 people, including Prescient Design antibody design), followed by big-pharma players such as AbbVie, AstraZeneca, Roche, Sanofi, Merck, and Novo Nordisk. The 15 big-pharma companies total 3,332 people, far exceeding the 1,042 across the 20 AI-native startups. Among the AI-natives, Recursion (211), Isomorphic Labs (187), AbCellera (144), and Schrödinger (138) lead on scale, while most protein-design startups (Cradle/Latent/Profluent/Nabla/Dyno/Chai) run lean teams (on the order of 10-30 people, with few visible profiles).
Reading: dual-background density correlates with company type but not absolutely. Density is highest at companies that build the wet-lab loop into their platform (Bristol Myers Squibb 74.6%, Novartis 74.3%, AbCellera 72.2%, AstraZeneca 71.1%); the more pure-ML, weak-wet-lab-signal companies are mostly pure foundation-model / protein-language-model startups (Genesis Therapeutics 30.4%, Isomorphic Labs 33.2%, Schrödinger 44.9%, insitro 53.6%). Sourcing takeaway: to find ML people who can read wet-lab data directly, recruit from the former; for pure-algorithm/foundation-model people, recruit from the latter.
Reading: starkly unlike pure AI (where the Bay Area dominates), AI-pharma talent is highly dispersed geographically. The US accounts for about 60% (the Bay Area, Boston/Cambridge, San Diego, New Jersey, and other pharma clusters), with the rest spread across the UK (London/Oxbridge), Switzerland (Basel/Zurich, Roche/Novartis), Denmark (Copenhagen, Novo Nordisk), France (Paris, Sanofi/Bioptimus), Germany (Mainz, BioNTech/InstaDeep), Canada (Vancouver, AbCellera), and Sweden (Gothenburg, AstraZeneca). Cross-region, cross-time-zone sourcing is the norm in this space, and localized outreach matters more than betting on a single cluster.
Reading: most profiles carry a title like "computational scientist / ML scientist" with no specified sub-discipline and fall into the general column. Specialists explicitly tagged to protein structure (19 people) or protein/antibody design (106 people) total just 125 across the entire industry, the truly scarce core; target/systems-biology ML (473 people) sits mostly in big pharma and at Recursion/insitro.
Reading: two main intake pipes. ① Universities/research institutes are the leading source of AI-pharma talent (Genentech, Novo, and Merck all hire PhDs straight from academia in volume), confirming that academic PIs and PhDs are the core supply; ② lateral movement between big pharma and among biotech peers is heavy. Isomorphic's intake pipe is dominated by "other companies," which is really a wholesale internal transfer from DeepMind (its spinout nature). The visible AlphaFold-lineage spillover beyond Isomorphic is very small, with most still inside the parent.
Reading: 1122 people were hired in 2025, 1.6x the 698 in 2023, reflecting the steady ramp in hiring driven by the 2024-2025 funding rebound (not the explosive surge seen in pure AI). The hiring peak for newcomers like Isomorphic is concentrated in 2024-2025.
Reading: the top-right "veteran zone" = AstraZeneca (median 37 months, 51.4% past 3 years) and Roche, where big pharma runs deep; the bottom-left "new-guard zone" = Isomorphic (median 11 months, 54.1% under a year), fast-scaling, where moving people in the honeymoon phase is hard but the first wobble window opens 12-24 months in. Genentech, Merck, and Novo sit in the middle (24-28 months).
This is the report's differentiating angle. "ML people who can fold proteins" = hybrid talent who understand both the wet lab (biology/biochemistry intuition) and the dry lab (ML/programming). Under the database methodology, 2,432 people (55.6% of the computational pool) carry a dual-background signal.
Reading: a dual background is the norm rather than the exception in AI pharma (55.6% average), which is precisely why being interdisciplinary is the ticket of admission to this space. Density is highest at companies that build the wet-lab loop into their platform (Bristol Myers Squibb, Novartis, AbCellera, AstraZeneca); pure foundation-model / protein-language-model companies (Genesis Therapeutics, Isomorphic Labs, Schrödinger, insitro) have a lower dual-background share, with more pure-ML teams.
Dual backgrounds may be common, but specialists explicitly working on protein structure prediction or protein/antibody design number just 125 across the full sample (2.9%). Roughly 45.0% of them hold a PhD. These 125 are the core everyone fights over in the AlphaFold era: people who can do de novo design with tools like RFdiffusion/ESM/Boltz and also tie into wet-lab validation. For buyers: ① this is a profile that genuinely requires active sourcing (not passively waiting for applications); ② supply comes mainly from the UW Institute for Protein Design (the Baker lineage), DeepMind/Isomorphic, the Meta FAIR protein team (now disbanded and spilling over), and a handful of academic PI labs; ③ knowing ML alone or biology alone is not enough, you must have both.
The market has no engineer-level version of big-pharma AI org charts. This section uses the full set of profiles to reconstruct the AI computational groups at Genentech / Novartis / AstraZeneca down to the team-lead layer and IC depth. Levels are classified from public job titles, are not official org structure, and serve only as an overview of team-tier structure. In the public version, names are masked by default.
Part of Roche, it includes the Prescient Design antibody/protein-design team (the birthplace of the lab-in-the-loop concept), led by Aviv Regev. It is one of the strongest AI protein-design groups inside big pharma and the source of the Coefficient Bio founding team.
Built a Generative Chemistry pipeline with Microsoft and partners with Isomorphic (6 projects), Generate, and Schrödinger. Among the highest dual-background densities in the full sample.
Deeply invested in AI drug discovery and partnered with Absci and others. One of the most tenured big-pharma AI groups (high veteran density).
Three groups totaling 32 people, screened by level, direction scarcity, and history strength. Profile facts come from the Metix AI database; those marked "publicly verified" have had their current role confirmed against 2025-2026 public sources (frontier-lab profiles lag in updates, so public sources take precedence). All figures in this section come from public professional profiles. Group A is founders, executives, and public technical leads; Group B is senior technical backbone; Group C is scarce-direction profiles. In the public version, names are masked by default.
The structural feature of AI-pharma compensation is "fighting frontier AI labs for the same ML people but unable to match the money." Figures follow 2025-2026 market data, not individual offers.
| Group | Total comp / pay range | Notes |
|---|---|---|
| AI-pharma ML scientist (most) | $98K-$176K base | Per ZipRecruiter / Takeda, AI-drug-discovery scientists average about $123K |
| AI-pharma specialist role (high end) | $200K-$240K | D.E. Shaw Research drug-discovery AI/ML data scientist |
| Frontier AI lab SWE median | $600K-$795K | Per levels.fyi, 3-5x the AI-pharma level |
| OpenAI / Anthropic equivalent level | $600K-$1.15M | Top-tier researchers go higher, with named special packages bypassing the leveling system |
| Return-to-China package at Chinese AI-pharma firms | Case by case | XtalPi / DP Technology / Helixon and others offer hybrid computational + wet-lab roles |
AI pharma and frontier AI labs fight for the same ML people, but the median pay gap is 3-5x. The result: AI pharma can't keep its pure-algorithm stars (they get hired by OpenAI/Anthropic) and must retain people differentially through "a sense of mission (curing disease / Nobel-grade science) + academic prestige (working alongside Baker/Koller/Hassabis) + equity upside + a dual-background bar (pure AI labs have no use for people who understand the wet lab)." This is also why this field runs on active sourcing + high fees rather than passively waiting for applications.
① Protein-structure/design specialists require active sourcing and must be attracted with non-cash advantages (quality of the scientific problem, the wet-lab loop, platform data); ② dual-background talent is a differentiated profile that AI labs struggle to attract and should be a priority; ③ European pay benchmarks (Switzerland/Denmark/UK) sit below the US, a value pocket for budget-constrained buyers building computational groups; ④ senior talent released by contracting companies is one of the few sources you can talk to right now.
Sources: ZipRecruiter, levels.fyi, the PwC 2025 AI skills premium, and the IntuitionLabs life-sciences employment report (retrieved 2025-2026). See the research memos for details.
Turning the map into action: biotech recruiters look at scarce profiles and windows, HealthTech HR at benchmarking and defense, healthcare VCs at team-diligence signals.
① Protein-structure/design specialists (just 125 across the full sample) are a high-fee area well suited to active sourcing; ② wet-lab × dry-lab talent (2,432 people) is a differentiated profile that AI labs struggle to attract and deserves priority; ③ big-pharma AI groups (3,332 people, 73%) are where talent is mainly concentrated, so don't focus only on marquee startups; ④ overlaying tenure windows with company type lets you pinpoint the more mobile groups.
① Position yourself against the 3.7 tenure benchmark: for newcomers (Isomorphic and others), the retention-defense window opens 12-24 months in; ② use non-cash advantages (the scientific problem, the wet-lab data loop, working with top PIs) to close the pay gap; ③ watch your own "dual-background + over-30-month-tenure" cohort (the most mobile group in the market); ④ The computational pool has a 45.0% PhD rate, and alumni chains + top-conference networks are the highest-hit outreach surface.
① The AlphaFold-lineage spillover is still early (most remain inside the Isomorphic parent), so the next wave of spinouts is worth tracking; ② academic PIs going commercial is the main axis of startup formation in this space (Baker → Xaira, Koller → insitro, Jian Peng → Earendil), so watch the moves of top protein-design / virtual-cell PIs closely; ③ for team diligence, use this report's dual-background density and org maps to judge whether a team is truly dual-capable or purely algorithmic.
This report's search, profiling, and flow analysis were all done by Metix AI. We can generate a custom map for any company under the same methodology: full long-list export, wet-lab × dry-lab filtering, org maps, email unlock, and multi-channel outreach, billed on a "pay only for qualified interviews" basis. No interview, no charge.
860 million+ global talent profiles4,374-person computational pool + 2,432-person dual-background long listProtein-design profile analysisPay only for qualified interviewsAI-native (all computational roles, 20 companies): Isomorphic Labs, Recursion, Genesis Therapeutics, Iambic, Chai Discovery, EvolutionaryScale, Xaira, Generate Biomedicines, insitro, Cradle, Latent Labs, Profluent, Schrödinger, AbCellera, Absci, Insilico Medicine, Cellarity, Dyno Therapeutics, Nabla Bio, Bioptimus. Big pharma (AI/computational subset only, 15 companies): Genentech, Roche, Novartis, AstraZeneca, Pfizer, Merck, Eli Lilly, Novo Nordisk, Sanofi, GSK, Amgen, BioNTech · InstaDeep, Bristol Myers Squibb, AbbVie, Johnson & Johnson. Geographic scope = profiles based in the United States, United Kingdom, Switzerland, Denmark, France, Germany, Canada, or Sweden.
Big pharma is enormous, so in this report "(AI/computational)" = those currently at the company whose title/headline is identifiable as AI/ML/computational/data-science-related, an identifiable subset rather than the full headcount, making the absolute numbers conservative. Among AI-native startups, protein-design teams are lean (10-20 people), so visible profiles are naturally few.
Computational/ML pool = protein structure/folding, protein/antibody design, molecular generation/chemistry, target/systems biology, ML platform, and computational/ML science (general). Wet-lab / pure-biology roles are excluded from the computational pool. Wet-lab × dry-lab = education/skills/history containing both wet-lab (biology/biochemistry/chemistry/experimental) and dry-lab (ML/programming/computational/statistics) signals, a probabilistic determination.
Data is through H1 2026; profile updates lag, representative figures have been re-checked against public information, and the latest 2025-2026 role changes are annotated based on public sources.
| Company | Current Profiles | Computational pool | Dual-background rate | PhD rate |
|---|---|---|---|---|
| Genentech (AI/computational) | 444 | 442 | 50.7% | 48.6% |
| AbbVie (AI/computational) | 353 | 353 | 71.1% | 58.1% |
| AstraZeneca (AI/computational) | 335 | 332 | 71.1% | 53.6% |
| Roche (AI/computational) | 281 | 281 | 51.6% | 50.9% |
| Sanofi (AI/computational) | 252 | 249 | 26.9% | 17.7% |
| Merck (AI/computational) | 246 | 246 | 54.5% | 45.9% |
| Novo Nordisk (AI/computational) | 243 | 243 | 66.7% | 46.1% |
| Johnson & Johnson (AI/computational) | 238 | 238 | 42.0% | 46.6% |
| Recursion | 612 | 211 | 54.0% | 44.5% |
| Novartis (AI/computational) | 206 | 206 | 74.3% | 33.0% |
| Bristol Myers Squibb (AI/computational) | 189 | 189 | 74.6% | 42.3% |
| Isomorphic Labs | 334 | 187 | 33.2% | 44.4% |
| GSK (AI/computational) | 156 | 156 | 62.2% | 44.2% |
| AbCellera | 458 | 144 | 72.2% | 40.3% |
| Schrödinger | 668 | 138 | 44.9% | 41.3% |
| Eli Lilly (AI/computational) | 131 | 131 | 51.9% | 45.8% |
| BioNTech · InstaDeep (AI/computational) | 126 | 126 | 27.8% | 29.4% |
| Amgen (AI/computational) | 108 | 108 | 50.0% | 44.4% |
| Xaira | 158 | 74 | 67.6% | 66.2% |
| insitro | 230 | 69 | 53.6% | 42.0% |
| Iambic | 115 | 37 | 67.6% | 48.6% |
| Pfizer (AI/computational) | 33 | 32 | 65.6% | 37.5% |
| Cellarity | 88 | 26 | 69.2% | 42.3% |
| Genesis Therapeutics | 71 | 23 | 30.4% | 52.2% |
| Absci | 64 | 23 | 56.5% | 43.5% |
| Generate Biomedicines | 68 | 19 | 73.7% | 73.7% |
| Cradle | 42 | 19 | 26.3% | 26.3% |
| Bioptimus | 30 | 17 | 5.9% | 29.4% |
| Profluent | 24 | 13 | 69.2% | 61.5% |
| Dyno Therapeutics | 44 | 11 | 90.9% | 54.5% |
| Nabla Bio | 17 | 9 | 55.6% | 22.2% |
| Insilico Medicine | 21 | 7 | 42.9% | 28.6% |
| EvolutionaryScale | 8 | 6 | 33.3% | 50.0% |
| Latent Labs | 21 | 6 | 33.3% | 66.7% |
| Chai Discovery | 6 | 3 | 33.3% | 66.7% |
① Snapshot currency: data is through H1 2026, with recent personnel changes lagging; representative figures have been re-checked against public information, and we recommend a second confirmation before using the long list.
② Coverage: this report is built from aggregated public professional profiles, with big pharma counted as its identifiable AI/computational subset; small protein-design startups (Chai/Latent/Cradle/EvolutionaryScale) run lean teams with lower coverage. All figures follow the database methodology and should be cross-referenced with companies' public headcounts.
③ Function and dual background are inferred: based on title/headline/skills/education keywords; many profiles carry a title like "computational/ML scientist" with no specified sub-discipline and fall into the general column, so the absolute number of protein-structure/design specialists is a lower bound.
④ Research memos: two research memos with all source URLs (industry landscape / talent ecosystem) are delivered in the same directory as this report.