top of page

Introducing SynthID Bio: Google DeepMind Puts Watermarks Inside AI-Designed Proteins

1 day ago
12 min read

Google DeepMind is introducing SynthID Bio after laboratory tests produced the first watermarked protein binders that retained their intended biological function.

The September 30 announcement moves AI watermarking beyond images, text, audio, and video. Its signature can remain detectable in a synthesized physical protein, not merely in a digital design file. That makes the work more consequential than another content-labeling experiment.

The central conflict is provenance versus control. SynthID Bio can identify outputs from participating models, but it cannot establish that every unmarked protein is natural or safe. Researchers also found that determined users can remove its sequence watermark by redesigning the protein.

The result is a credible proof of concept with a demanding path to deployment. It pressures model developers, biological databases, and DNA synthesis providers to decide whether shared provenance standards belong in the infrastructure of AI-assisted biology.

Introducing SynthID Bio as a Physical Provenance Test

Google DeepMind has shown that an AI-generated protein can carry a detectable signature without losing the function researchers designed it to perform.

SynthID Bio is a family of methods for watermarking protein sequences and predicted biomolecular structures. A watermark is a hidden statistical signal that authorized software can detect later. It indicates an output came from a compatible AI system.

The project has two distinct components. SynthIDBio-sequence influences which amino acids a sequence-generation model selects. SynthIDBio-structure introduces small, detectable patterns into the atomic coordinates predicted by a folding model.

That distinction matters because a protein sequence and a predicted structure represent different things. The sequence specifies an ordered chain of amino acids. The structure describes how atoms in that chain are arranged in three dimensions.

For the sequence test, researchers combined AlphaProteo-designed backbones with a modified version of ProteinMPNN. ProteinMPNN generates amino acid sequences expected to fold into a supplied protein structure.

The team evaluated binders against three targets: VEGF-A, PD-L1, and the receptor-binding domain of the SARS-CoV-2 spike protein. Protein binders are molecules designed to attach selectively to particular biological targets.

Researchers started with 15 previously validated structural backbones for each target. For every backbone, they tested a parent design, five unwatermarked sequences, and two sets of six watermarked sequences.

The resulting laboratory evaluation covered 222 unwatermarked designs and 267 designs under each watermark setting, excluding controls. Surface plasmon resonance measurements tested whether those molecules still bound their targets.

According to the peer-reviewed protein watermarking study, researchers found no significant population-level difference in binding affinity between the watermarked and unwatermarked groups. Hit rates and sequence diversity also remained comparable.

Those results do not show that every protein class can accept a watermark safely. They do establish that watermarking survived contact with physical biology in a controlled binder-design experiment.

The SynthID Bio announcement therefore makes a narrower claim than its striking premise might suggest. Google DeepMind calls the system a proof of concept, not a universal provenance service ready for production.

That caution is essential. The experiment begins with known binder backbones and resequences them to collect useful affinity data. It does not test an unrestricted collection of enzymes, receptors, antibodies, or complete organisms.

Still, physical validation changes the status of the idea. Earlier biological watermarking proposals often focused on computational quality metrics. SynthID Bio tests whether a sequence-level signature can coexist with measured biological activity.

Why Protein Provenance Has Become an Infrastructure Problem

AI protein design is creating unfamiliar biological sequences faster than existing provenance systems were built to classify them.

Generative models can propose protein sequences that differ sharply from known natural examples. That ability supports drug discovery and biological research, but it complicates screening at the digital-to-physical boundary.

A researcher normally sends a DNA sequence to a synthesis provider before producing a designed protein. Providers screen orders against databases containing known biological risks. Traditional screening often depends partly on similarity between an ordered sequence and known hazardous material.

Novel designs weaken that assumption. An unfamiliar sequence might represent legitimate research, an undiscovered natural organism, or an engineered object requiring closer review. Low similarity does not automatically settle which explanation is correct.

Google DeepMind argues that a detectable watermark can provide additional context. A synthesis provider could determine that a sequence came from a participating model with defined safeguards and customer controls.

That signal would not certify safety. It could help screeners allocate their attention by separating recognized model outputs from unexplained sequences. James Diggans, Twist Bioscience’s vice president for policy and biosecurity, described watermarking as an additional tool for focusing screening resources.

The pressure is not purely hypothetical, although estimates of immediate danger remain contested. Previous computational work found that AI-assisted redesign could produce sequences intended to evade similarity-based screening.

A later NIST evaluation added important restraint. Its researchers found that contemporary systems could generate structurally similar synthetic homologs without reliably preserving biological activity.

NIST concluded that current tools were not yet consistently capable of rewriting a protein while retaining activity and evading screening. It also found that experimental validation requires substantial time, expertise, and resources.

SynthID Bio should therefore be understood as anticipatory infrastructure. It addresses a provenance problem that becomes more pressing as design models improve, not proof of an uncontrolled biological crisis today.

Scientific databases face another version of the same problem. Protein Data Bank, UniProt, and GenBank records feed research tools, bioinformatics pipelines, and later generations of machine-learning models.

Public submission supports broad scientific participation. It also creates opportunities for synthetic entries to be mislabeled as natural observations or supported by false metadata.

Incorrect records can contaminate search results and training corpora. They can also distort security analysis when screening tools rely on the same shared reference collections.

A watermark could prompt database curators to request supporting evidence or label an entry as AI-generated. That would preserve useful synthetic data without presenting it as a naturally observed sequence.

The stakes extend beyond one Google system. Provenance only becomes infrastructure when independent model developers, synthesis companies, repositories, and researchers can interpret compatible signals under agreed rules.

Without that coordination, SynthID Bio remains a technically interesting marker inside a fragmented biological information system.

How Google DeepMind Watermarks Sequences and Structures

SynthID Bio hides provenance in the choices a model already makes, but sequence and structure watermarking require different compromises.

Protein sequence generation is probabilistic. At each position, a model can often choose among several plausible amino acids while preserving the intended fold. That flexibility creates room for a statistical watermark.

SynthIDBio-sequence adapts tournament sampling from SynthID’s text technology. The sampling process gently favors selected amino-acid choices according to a secret watermarking key.

A detector later tests whether those favored choices appear more often than chance would predict. The sequence does not carry a visible label or an appended tag. Its pattern is distributed across the generated design.

Google DeepMind tested both distortionary and non-distortionary settings. Distortionary sampling changes the output distribution to strengthen the signal. Non-distortionary sampling aims to preserve the model’s original distribution across many generations.

The system also filters generated candidates using a watermark score. Designs below a chosen detectability threshold can be rejected before more expensive structural validation begins.

This creates a direct cost tradeoff. In one tested configuration, guaranteeing full detection at a 0.1 percent false-positive rate reduced the candidate pass rate by 41.7 percent.

That reduction requires about 71.5 percent more initial candidates to produce the same number passing the watermark filter. The researchers say practical overhead remained negligible because rejected candidates avoided the costlier AlphaFold 3 validation stage.

Structure watermarking follows another route. AlphaFold 3 predicts three-dimensional atomic coordinates through a diffusion process, which iteratively turns noise into a structured molecular model.

The researchers fine-tuned part of that diffusion network with an added watermarking objective. A separate detector learned to distinguish watermarked structures from ordinary predictions and experimental ground-truth structures.

Embedding the behavior in model weights offers an important advantage. A person running the modified model receives watermarked predictions without applying a separate post-processing step.

Google DeepMind reports detection rates above 99.8 percent at a 0.1 percent false-positive rate across three tested models. At the recommended setting, key structural accuracy measurements did not decline against the AlphaFold 3 baseline.

The detector uses geometric features such as atom-to-atom distances and torsion angles. Those features make the signature naturally resistant to rotation or translation of the complete structure.

It also remained detectable after minor coordinate noise. This is important because common file formats round atomic coordinates, and scientific pipelines frequently transform structural files.

The two methods therefore distribute control differently.

Sequence watermarking

  • Integrates into the ProteinMPNN sampling process.

  • Requires no change to the underlying model weights.

  • Can generalize to similar autoregressive sequence generators.

  • Can also be disabled or removed more easily.

Structure watermarking

  • Requires fine-tuning part of a diffusion model.

  • Embeds watermarking behavior within model weights.

  • Remains active when those modified weights are shared.

  • Depends more closely on the model architecture and training process.

The mechanism is related to earlier research, but it is not the first proposal for protein watermarking. The 2024 FoldMark project tested structure watermarks across several prediction and design models.

FoldMark reported high recovery rates for short embedded codes and resistance to several simulated modifications. However, it remained a computational study and reported weaker performance under major structural changes.

SynthID Bio’s clearest contribution is the connection between provenance and experimental function. The laboratory binder tests give the sequence approach evidence that structural scores alone cannot provide.

The Real Opponent Is Removable Provenance

The central test is not whether SynthID Bio can mark a protein, but whether the mark survives users who have reasons to remove it.

Researchers directly tested a resequencing attack using ProteinMPNN. An attacker starts with a watermarked protein structure and generates a new amino acid sequence expected to preserve that structure.

Across 38,396 binders, resequencing largely removed the sequence watermark. This is not a peripheral limitation. Browser-accessible versions of ProteinMPNN lower the expertise required to perform the transformation.

The attack also creates a biological cost, which gives the watermark some defensive value. According to estimates derived from the laboratory experiments, resequencing reduced expected binder success rates.

With structural filters, estimated hit rates fell to 97 percent for the SARS-CoV-2 target, 70 percent for PD-L1, and 66 percent for VEGF-A. Without those filters, rates fell to 33, 20, and 3 percent.

These figures describe selected binders and an evaluated attack pipeline. They should not be generalized to every protein, target, or adversarial strategy.

They nevertheless illustrate Google DeepMind’s security argument. A watermark does not need to be impossible to remove if removal imposes useful cost, uncertainty, or functional degradation.

That logic resembles the gene-synthesis industry’s broader security model. Screening raises barriers and increases the chance of detection. It does not guarantee that every harmful attempt will fail.

However, provenance works differently from access control. A malicious user can choose an unwatermarked model, disable sequence watermarking, or regenerate an existing design. The absence of a signal then says very little.

This asymmetry creates the adoption problem. Cooperative developers can label their outputs, while deliberately uncooperative actors remain outside the system.

The structure watermark also has failure modes. Mild coordinate noise left much of its signal intact, but constrained molecular relaxation destroyed the tested watermark.

Relaxation adjusts atomic coordinates toward a locally favorable physical arrangement. It is a normal scientific operation, not necessarily evidence of malicious tampering.

Google DeepMind says future training could incorporate relaxation to improve resilience. Until then, routine downstream processing can erase the structural signal under some conditions.

Both current methods are zero-bit watermarks. They communicate the presence of a signature, but they do not encode detailed information such as a user identity, generation time, or project record.

Detection also relies on secret keys shared with trusted parties. Public verification would require new public-key methods and standards for distinguishing watermarks from different providers.

False results carry operational consequences. A false negative can allow mislabeled data to enter a repository. A false positive can trigger extra review or wrongly suggest that researchers used AI.

In a biosecurity workflow, the consequences can reverse. Treating a detected mark as evidence of trust could reduce scrutiny for a sequence that deserves deeper inspection.

The research paper explicitly discusses these threshold choices. High detection guarantees can require rejecting more generated candidates, increasing compute and slowing legitimate work.

SynthID Bio is therefore not a safety certificate. It does not determine whether a protein is harmful, effective, ethically produced, or supported by valid experimental evidence.

It answers a narrower question: does this sequence or structure carry the statistical signature associated with a participating generation pipeline?

That answer can support a security decision, but it cannot replace one.

AI-Generated Protein Detection Needs Shared Rules

A watermark becomes useful only when institutions agree who can detect it, what detection means, and what action follows.

Google DeepMind presents SynthID Bio as one layer in a broader defense system. Other layers include model safeguards, customer verification, sequence screening, access controls, and experimental review.

That layered approach is appropriate because each intervention has blind spots. A watermark travels with the designed object, while account records remain with the service that generated it.

Metadata can provide richer context, including model versions and timestamps. Yet ordinary metadata can be stripped, altered, or separated from the underlying sequence.

Central registries offer another option. Developers could record hashes or complete sequences of AI-generated designs in a controlled repository. Screeners could query that record when reviewing an order.

Registries raise privacy and intellectual-property concerns. They also require a trusted operator and reliable matching at enormous scale. Slightly altered sequences can complicate exact lookup.

Watermarks trade detailed records for embedded evidence. They can survive copying because the signal resides in the sequence or predicted coordinates. Their usefulness still depends on access to a valid detector.

This makes governance as important as detection accuracy. Database maintainers need procedures for disputed findings. Synthesis providers need rules preventing a watermark from becoming an automatic approval.

Model developers need compatible schemes that do not interfere with one another. Researchers need disclosure policies that distinguish helpful provenance from punitive surveillance.

The field also lacks a settled division between sequence and structure provenance. A model can generate a structure, another tool can design its sequence, and a third system can optimize laboratory properties.

Each transformation can introduce a new creator and remove evidence from an earlier step. A single binary signature cannot represent that complete chain.

A future system may combine watermarks, signed metadata, audit logs, and repository records. Each layer could document a different part of the object’s history.

That approach resembles provenance work for digital media, but biology raises harder questions. A media file remains information, while a protein sequence can become a physical molecule and reproduce through encoded genetic material.

Google DeepMind is already testing that boundary. It says ongoing work with Stanford’s Hie lab and Arc Institute added SynthID Bio to Evo 2.

The team used it to watermark the genome of an AI-designed bacteriophage, a virus that infects bacteria. Early culture tests reportedly found that the watermarked phages remained functional.

That result has not yet established a general method for genomic provenance. A genome contains interacting genes, regulatory regions, and evolutionary pressures far beyond a single designed binder.

It does show where the research is heading. The target is not just a label for isolated protein files. It is persistent provenance for increasingly complex AI-designed biological systems.

Independent technical efforts will also shape any standard. FoldMark explores user-specific structural codes, while other researchers propose sequence databases, signed records, and model-level safeguards.

No single company should define biological provenance alone. A credible system requires participation from scientific repositories, synthesis providers, independent laboratories, standards organizations, and multiple model developers.

Introducing SynthID Bio starts that negotiation with physical evidence. It does not finish the institutional work required to make the signal dependable.

What Comes After Introducing SynthID Bio

Three signals will determine whether this proof of concept becomes useful infrastructure or remains a laboratory result.

The first signal is independent replication across broader protein classes. Researchers need to test enzymes, therapeutic candidates, antibodies, and larger complexes under realistic optimization workflows.

Success would strengthen the claim that watermarking can preserve function beyond selected binders. Failures concentrated in particular protein families would define where provenance must use other methods.

The second signal is improved resistance to routine transformation. Sequence marks must better survive partial resequencing, while structure marks must withstand relaxation and other standard scientific processing.

Resistance cannot come at any cost. Stronger marks that reduce binding, stability, diversity, or model accuracy would undermine the research they are intended to protect.

The third signal is institutional adoption. DNA synthesis providers and biological repositories must publish clear policies for detecting watermarks and responding to results.

That adoption must include safeguards against false confidence. A recognized signature should contribute evidence about origin, not excuse safety review. An absent signature should not be treated as evidence of malicious intent.

Public-key detection would also materially change the adoption picture. It could let more institutions verify marks without receiving a model developer’s secret key.

Yet broader verification introduces new risks. Attackers could gain more information about detector behavior, and incompatible providers could generate overlapping or confusing signals.

The research community will need benchmarks covering detectability, functional preservation, removal cost, and false-positive behavior. Wet-lab testing should remain central because computational quality scores cannot guarantee biological activity.

It should also separate current capability from projected risk. The AI protein risk study behind NIST’s evaluation found serious screening questions, but experimental results showed important limits in present systems.

That evidence supports measured preparation. Better provenance can arrive before design tools reliably overcome every biological constraint.

For developers and research organizations, the immediate lesson is practical. Record model versions, design steps, filters, and experimental validation alongside generated sequences. Do not wait for one watermark to carry the full history.

For database maintainers, the priority is defining how AI-generated entries should be labeled and reviewed. The provenance problem exists even when no deliberate attack occurs.

For synthesis providers, SynthID Bio offers a potential triage signal. Its value will depend on integration with sequence screening and customer checks, not replacement of either layer.

Introducing SynthID Bio makes a hidden assumption visible: AI-designed biology needs a trustworthy memory of how each object was created. The open question is whether the field will build that memory collectively.

Over the next several months, watch for independent wet-lab replications, stronger tamper resistance, and published adoption plans from repositories or synthesis companies. Those developments will show whether the watermark can move from a Google DeepMind model into shared scientific practice.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page