MEMO · LAB · AUGUST 2026

Anthropic built an AI for scientists, and the hard part was trust

9 min read · every claim sourced below

The short version: Claude Science is a beta application for macOS and Linux that installs where your data already sits and runs a local daemon, with its interface in a browser. Anthropic’s own interviews found 91% of scientists want more AI and 79% name trust as the thing holding them back, so the product is built around evidence rather than output. Every figure, table and notebook carries four layers of history, and a second agent reads the session while the first one works and flags any claim it cannot trace. The parts a product guide puts last: it is token-heavy, it is not HIPAA-ready, it is not a validated system, and it does not run on Windows.

None of this is a leak or a secret. It is Anthropic’s own 26-page deployment guide, published alongside the product, plus the launch post and the three product screenshots it ships with, which are reproduced below. What the guide does not do is put the mechanism and the fine print on the same page, which is what this is.

hmm.
91% of scientists want more AI. 79% name trust as the reason they hold back.
Anthropic's title card for Claude Science: a hand-drawn monitor with a node graph on its screen, on a coral background.
Anthropic's title card for the launch, 30 June 2026. Image: Anthropic

The two numbers the product is built around

0%
of scientists want more AI in their research
0%
name trust and reliability as their number one barrier

Both figures come from Anthropic’s own internal research, drawn from interviews with researchers across chemistry, physics, biology and computational fields. That is worth saying plainly: this is not an independent survey, and no sample size is published, so treat them as directional rather than as a measurement.

The independent number in the guide is Deloitte’s, from a survey of 280 biopharma and medtech leaders. 78% expect AI to play a central role in driving major change this year, only 14% report full implementation of AI tools in daily workflows, and another 40% are still working toward it. The appetite is nearly universal. The deployment is not.

Where it runs, and where your data actually goes

Claude Science is a standalone application for macOS and Linux that runs a local daemon with its UI in the browser. Anthropic’s own comparison is a Jupyter notebook, except the agent is driving. You install it wherever the data lives: a scientist’s laptop, a lab Linux box, an HPC login node, or a cloud VM in your own tenancy. Data, compute environments and agents stay on that machine, and when the daemon runs remotely the scientist connects from a laptop browser over an SSH tunnel.

When a job needs bigger hardware, it dispatches from the same session to the lab’s own GPU box, an SSH host, a SLURM cluster (it auto-detects SLURM and writes the batch directives), or a serverless GPU account you supply. Every dispatch target is gated by the same approval broker, so Research IT can restrict which hosts a group’s install is allowed to reach.

The sentence most write-ups skip

Files remain on the host and are read in place. But content the agent reads as part of an analysis is sent to Anthropic’s API as context, subject to your plan’s data-use and retention commitments. An OS-level sandbox with deny-by-default network egress controls all other traffic leaving the host.

So “it runs next to your data” is true, and it is the reason this is deployable where a SaaS tool would not be. “Nothing leaves your machine” is not true, and nobody at Anthropic claims it.

Once installed, agents can be pointed at any local folder, including FASTQ files, AnnData objects and Seurat objects, and connect natively to S3, GCS, GitHub and institutional literature access. Conda and pip environments are managed per specialist, and sessions, kernels and artifacts persist across reboots. It installs as a signed binary running as a user-space daemon, with no kernel-level components.

Claude Science dispatching an eight-arm scVI hyperparameter sweep to a lab cluster, with a remote job list showing eight runs at about sixteen minutes each, and a live Python notebook kernel on the right.
Anthropic’s caption: “Claude Science builds environments and manages compute on your laptop, your cluster, or GPUs on demand.” What is actually on screen is an eight-arm hyperparameter sweep dispatched to the lab’s own A100s, the eight remote jobs listed with their run times, and a live kernel the scientist can type into while the agent works. Screenshot: Anthropic

The five design choices that make it defensible

This is the list the guide leads with, and it is the one worth memorising. Five things make Claude Science suitable for scientific work, and each one exists because a result you cannot check is a result you cannot publish.

1

Persistent kernels, and it sees its own plots

Agents load a dataset into a persistent Python or R kernel once, and after that they are exploring it rather than reloading it. Variables, dataframes and loaded models stay in memory across the whole analysis. The part that matters for trust: every figure generated is fed back into the agent’s own context, so it runs QC on its own output, spots the outlier cluster in its own UMAP, and filters it before moving on.

2

Full provenance on every artifact

Figures, tables, reports and notebooks are first-class objects rather than files to dig up afterwards. Each ships with four layers of history: a human-readable description of what was done, the exact reproducible code that produced it, the conversation and reasoning that led there, and a snapshot of every package and version used.

That bundle is what a methods section gets drafted from, and the guide recommends treating it as a controlled record for anything feeding a publication or a submission. Decide where the bundles live and for how long before you scale, not after.

Every chart ships with a receipt: what it did, the exact code, the conversation, and every package used. A second AI checks it before you see it.
A Claude Science figure of a single-cell atlas, with the artifact panel beside it showing tabs for Code, Execution Log, Messages, Environment and Review, and an inline comment on the figure reading 'these labels are hard to see'.
The four layers, as tabs. The figure on the left is the artifact; the panel on the right carries Code, Execution Log, Messages, Environment and a Review tick, with the input CSVs named above the script. The speech bubble over the plot is the other half of the idea: you annotate the figure itself, in plain language, and the agent edits the code that drew it. Screenshot: Anthropic
3

A background reviewer that runs every session

A separate reviewer agent reads each session’s transcript while the primary agent is still working, and flags any claim it cannot trace to evidence. Findings surface inline at the suspect sentence, and the agent fixes them before finishing. It runs every session by default and can also be triggered manually at any point.

A Claude Science literature review session. A Reviewer panel shows one finding, warning that PMID 31178118 was assigned to two different papers in the plan, and the agent's next message acknowledges the swap and gives the corrected pair. A compiled review PDF is open on the right.
The reviewer, caught in the act. Its one finding is that the same PubMed ID had been attached to two different methods papers, with the exec-log rows it checked listed underneath, and the agent’s next message acknowledges the swap and carries the corrected pair. Anthropic publishes this shot under “domain-ready on day one”, for the five parallel retrieval tracks running across PubMed, bioRxiv, OpenAlex and CELLxGENE on the left. Screenshot: Anthropic
4

Plans before actions, and permissions you can see

Plan mode is on by default: Claude Science drafts each task as a step-by-step plan and waits for approval before executing, and the plan stays visible as a checklist you can edit or scope down while work runs. Whenever an agent needs to reach a new website, open a folder or run code, an approval card appears with allow once, for this project, or always, and every decision can be reviewed or revoked from one permissions screen.

Underneath, a human-approval broker gates thirteen action kinds, including code execution, network grants, host file access, deletions, MCP tool calls, remote compute dispatch and skill persistence. The OS-level sandbox adds allowlist-proxy network egress, SSRF and DNS-rebind defences, and seccomp hardening.

5

Biosecurity safeguards specific to biology

Biology is a dual-use domain, so this layer sits on top of the sandbox rather than inside it. Biosecurity rules ship unconditionally in every agent’s system prompt, and a per-turn bio trajectory classifier runs inside the binary and cannot be disabled by the user or the admin. Authentication is OAuth-only, with no anonymous or API-key access against the product, and Anthropic states it completed external red-teaming and an internal Safeguards review against CBRNE risk before public release.

More than sixty databases, and about 150 skills

Claude Science ships with optional connections to more than sixty scientific databases and roughly 150 curated skills. The important detail is how they work: these are pre-built Python skills that run code rather than retrieve documents, so Claude writes and executes a query against the source of record and returns results with provenance. Because they are code, it can chain them: pull a gene’s variants from one database, cross-reference drug interactions in another, and check expression across cell types in a third, inside a single analysis.

Each skill is open source, so a computational team can inspect the query logic, pin versions, or extend a skill with the organisation’s own filters and output formats. The catalog groups by the kind of question it answers.

Genes, variants and annotation

  • gene-database (NCBI Gene and Datasets), for RefSeqs, GO terms, genomic location, associated phenotypes and batch retrieval.
  • biothings-database (MyGene, MyVariant, MyChem, MyDisease), for resolving identifiers across databases in one call.
  • ena-database (European Nucleotide Archive), for raw reads, assemblies and sample metadata by accession.
  • harmonizome-database, for what a gene is associated with across 170+ functional genomics resources at once.

Expression and single cell

  • cellxgene-census (CZ CELLxGENE Census), filtering 125M+ cells by type, tissue or disease straight into scanpy or PyTorch.
  • immgen-database, expression across mouse immune cell populations.
  • allen-brain-database, expression, connectivity and spatial transcriptomics across mouse and human brain regions.

Oncology

  • cosmic-database, somatic mutations, the Cancer Gene Census, mutational signatures and fusions. Requires institutional authentication.
  • tcia-database, DICOM imaging series from public cancer imaging collections.

Pharmacology, safety and metabolomics

  • clinpgx-database (ClinPGx, formerly PharmGKB), gene-drug interactions, CPIC dosing guidelines and allele function.
  • fda-database (openFDA), adverse event reports, recalls, drug labels, 510(k) and PMA records, and UNII identifiers.
  • hmdb-database, 220K+ human metabolites with properties, biomarker associations and NMR/MS spectra.
  • metabolomics-workbench-database, 4,200+ public studies, RefMet nomenclature and m/z lookups.

Neuroscience

  • neuromorpho-database, neuron morphology reconstructions and morphometrics.
  • openneuro-database, BIDS-formatted MRI, fMRI, EEG, MEG and PET datasets.

Multi-database toolkits

  • bioservices, 40+ services including UniProt, ChEMBL, PubChem, Reactome, QuickGO, KEGG, Ensembl and BioMart.
  • gget, 20+ databases for one-liner lookups: Ensembl, UniProt, NCBI, ARCHS4, Enrichr, OpenTargets, PDB, AlphaFold and BLAST.
  • biopython (Bio.Entrez), anything behind NCBI E-utilities.
  • hmmer for profile and sequence searches, and foldseek for searching by 3D structure to find homologs the sequence would miss.
The licensing line to read before you scale

Public databases and other third-party resources are governed by their own licences and use terms. Anthropic is explicit that organisations are responsible for ensuring their use, including commercial use, complies with those terms and their own entitlements. COSMIC, for one, needs institutional authentication.

The catalog is opinionated but open. A lab can save any working pipeline as a reusable skill, build a custom specialist for its own methods from inside a session, or wrap an internal API once so every future session inherits it. The rule of thumb the guide gives: use a connector when the answer lives in your own systems and entitlements matter, and a skill when the answer lives in the public record. Most real questions use both.

Which Claude is for which job

Claude Science is one surface of seven, and the guide expects most organisations to deploy more than one.

  • Claude Science, end-to-end research analysis with provenance. A local app on macOS and Linux that dispatches to SSH, SLURM or cloud compute.
  • Claude Chat, everyday questions and drafting in a browser, desktop or mobile.
  • Claude Cowork, a desktop app for cross-app document work spanning folders and systems such as Benchling, Veeva, Microsoft 365 and PubMed.
  • Claude Code, agentic software engineering inside a repository, for the production pipelines downstream of an analysis.
  • Claude for Microsoft 365, in-place drafting and redlining across Word, Outlook, Excel and PowerPoint.
  • Claude Platform and Claude Managed Agents, for embedding Claude into ELN, LIMS, CTMS, safety or RWE systems, and for running custom agents as hosted services.

The one-line test worth keeping: reach for Claude Code when the output is software that ships to other teams, and Claude Science when the output is an analysis, a figure, or a result. If you already live in Claude Code, our memo on keeping a Claude Code session cheap applies here almost word for word, because Anthropic says the usage profile is the same.

The rollout, in three phases

Foundation. Because it runs locally, this phase is about getting the daemon next to the right data and compute rather than standing up cloud tenancy. Decide the daemon host pattern, confirm scientists can reach it from a browser, and get IT to review the sandbox, the deny-by-default egress allowlist and the approval broker. Account setup is the usual SSO and SCIM work with one extra step: an admin must enable Claude Science before anyone in the org can download or sign in. On Team that is admin settings then capabilities. On Enterprise the admin creates a role carrying the Claude Science permission, assigns it to the pilot group, then enables the capability, so access is scoped from day one.

Pilot. Champions run real analyses on real lab data against criteria set upfront. Three metrics decide whether it is working, and the third is the one specific to this product:

  • Cycle time, how long the same class of analysis took before and after.
  • Keep rate, how often a scientist or PI trusts the result without re-running it by hand.
  • Cold reproduce, take an artifact produced in week one, hand its provenance bundle to a different scientist in week four, and confirm they can re-run it cold.

The signal Anthropic says to watch for is champions starting to save their own skills: a bioinformatician wrapping the lab’s normalisation pipeline, a group lead wrapping the LIMS API. Those become the lab’s catalog, and they can be shared across the organisation.

Scale. Research IT moves from per-lab installs to a managed pattern: a standard daemon host per group, a vetted network allowlist, a curated skill catalog seeded from the pilot, and a defined set of compute-dispatch targets. Skills compound across groups because so much computational biology shares structure. The guide is firm that governance is settled before scale, not after: who owns each skill, how it is QC’d before it is shared beyond its author’s group, and how provenance bundles are retained for anything feeding a regulatory or publication output.

The pro tip that decides adoption

A scientist who opens Claude Science, points it at a folder of FASTQ files, approves the plan and gets a clustered UMAP with the code and environment captured underneath will come back. A scientist who opens it without data in reach will close it. Make sure the install lands next to real data.

The fine print, in one place

All of this is in the guide, mostly in the CIO FAQ near the back. None of it is disqualifying, and all of it is the sort of thing you would rather know in week one than week four.

  • It is beta, available on first-party Claude paid plans and Claude for Enterprise, including customers who procure through AWS Marketplace.
  • macOS and Linux only. Windows is not currently supported.
  • Not on Amazon Bedrock, Google Vertex AI or Microsoft Foundry. If you need workloads inside your own cloud perimeter, that is a Claude Platform conversation instead.
  • No separate licence and no free tier. It draws down the usage limits of the user’s existing Pro, Max, Team or Enterprise plan. A subsidised Team plan exists for academic and nonprofit research labs.
  • It is token-intensive. Anthropic says heavy users consume at a rate comparable to heavy Claude Code use, and that heavy individuals should expect to need Max-tier limits. Consumption is visible in the app under Settings then Usage.
  • Not HIPAA-ready at launch. HIPAA readiness is on the post-launch roadmap, and until then it should not be used to process protected health information.
  • NIH controlled-access data is not cleared. dbGaP and similar typically require the analysis environment to meet NIST SP 800-171, and Claude Science has not yet been assessed against that standard.
  • Not a validated system. For GxP-regulated work it sits in research, analysis and draft-support roles, with a qualified reviewer approving every output before it enters a validated record. Pair it with your own CSV or CSA assessment.
  • Zero Data Retention does not apply. It is a stateful product, because sessions, artifacts and provenance bundles need storage to exist. Team and Enterprise plans support custom retention, and ZDR remains available on the Claude Platform and Claude Code for approved customers.
  • Research use only. Anthropic states it is not designed for clinical or diagnostic decision-making, and that roadmap items are subject to change and do not represent a commitment.

Three labs that used it before you could

These are the beta users Anthropic names in the launch post, and they are the closest thing to a benchmark that exists so far. Each one is the lab’s own account of its own work, not a measured comparison, so read them as what practitioners report rather than as a result anyone has replicated.

Stephen Francis, an associate professor and epidemiologist at the UCSF Brain Tumor Center, studies the molecular epidemiology of glioma and how thousands of small-effect germline variants combine into individual susceptibility. The work predates Claude Science, but he says the app enabled comprehensive germline workups across multiple approaches in roughly a tenth of the time it previously took, and his group independently validated the results.

Jérôme Lecoq, a neuroscientist at the Allen Institute, built a computational review template out of about twenty custom skills. Sub-agents read thousands of papers, pull each central claim and key quantitative finding into an evidence database, and then a narrative arc is constructed section by section, with an actor-critic pair on every section: one agent writes, a separate reviewer checks accuracy and citation fidelity. Reviews like that used to take his team up to two years. He now has around ten of them, many over a hundred pages.

Manifold Bio designs tissue-targeting medicines and used Claude Science to nominate targets for its latest experiments, assessing surface expression, trafficking and safety per tissue and ranking candidates against criteria learned from its own proprietary data. What separated it from a general coding assistant, the company said, was doing that end to end with the context of past programmes built in.

What the wider Claude family has already done in life sciences

Worth separating from Claude Science itself, because these are the surrounding surfaces rather than the new product. Novo Nordisk built NovoScribe, a Claude-powered platform that automates clinical study reports, device verification protocols and patient materials. Documentation that previously took more than ten weeks now reaches a reviewable first draft in roughly ten minutes, and Waheed Jowiya, Digitalization Strategy Director, is quoted saying Claude cut writing times on CSRs by 90% so documentation gets into human hands for review sooner.

The Garvan Institute has more than twenty software engineers and data scientists using agentic development tools for drug discovery, rare disease diagnosis through multi-agent systems that interpret genetic variants, and genomics data analysis. Sanofi has deployed Claude enterprise-wide inside its internal Concierge application, used by the majority of employees daily.

Questions people actually ask

What is Claude Science?

It is a beta Anthropic application for the digital steps of research: literature review, experiment design, data analysis, figure generation and writeups. Unlike Claude Chat or Claude Cowork it is not a hosted product. It is a standalone app for macOS and Linux that runs a local daemon on the machine where your data already lives, with its interface in a browser. The comparison Anthropic uses is a Jupyter notebook where the agent is driving.

Does my data leave my machine?

Partly, and the distinction matters. Files remain on the host and are read in place, so a folder of FASTQ files is never uploaded. But content the agent reads as part of an analysis is sent to Anthropic's API as context, subject to your plan's data-use and retention commitments. An OS-level sandbox with deny-by-default network egress controls everything else leaving the host. So it is accurate to say Claude Science runs next to your data. It is not accurate to say nothing leaves your machine.

Can I run Claude Science on Windows?

No. Windows is not currently supported. It installs as a signed user-space daemon on macOS or Linux with no kernel-level components, and on Linux it can run headless on a lab box or an HPC login node and be reached from your laptop browser over an SSH tunnel. Claude Cowork, a different product, is the one that ships a desktop app for macOS and Windows.

Is Claude Science HIPAA-ready or usable with dbGaP data?

Not at launch, on both counts. Anthropic states HIPAA readiness is on the post-launch roadmap and that until then Claude Science should not be used to process protected health information. NIH controlled-access datasets such as dbGaP typically require the analysis environment to meet NIST SP 800-171 controls, and Claude Science has not yet been assessed against that standard. Formal NIH controlled-access compliance is described as on the roadmap, and Anthropic notes roadmap items are subject to change and are not a commitment.

How much does Claude Science cost?

There is no separate licence and no free tier. It draws down the usage limits of the user's existing Pro, Max, Team or Enterprise plan, and it is available in beta on first-party Claude paid plans and Claude for Enterprise, including enterprise customers who procure through AWS Marketplace. Anthropic describes it as token-intensive, with heavy users consuming at a rate comparable to heavy Claude Code use, and says to expect heavy individual users to need Max-tier limits. A subsidised Team plan exists for academic and nonprofit research labs.

What is the difference between Claude Science and Claude Code?

The guide gives a one-line test: use Claude Code when the output is software that ships to other teams, and reach for Claude Science when the output is an analysis, a figure or a result. Claude Code is a terminal tool for repositories under version control. Claude Science is an analysis workbench with persistent kernels, scientific renderers and a provenance bundle attached to every artifact it produces.

Can Claude Science results go into a regulatory submission?

Not directly. Anthropic is explicit that Claude Science is not a validated system and that it is intended for research use rather than clinical or diagnostic decision-making. Organisations deploy it in research, analysis and draft-support roles with a qualified scientist or reviewer approving every output before it enters a validated record, a submission or a publication. The four-layer provenance bundle is what that reviewer reads.

Want this built for you?

We write these memos because we build this stuff every day. If you want it working in your business instead of sitting on your reading list, that is literally our job.