Google DeepMind creates watermarked AI-designed proteins that still work in lab tests

SynthID Bio embeds a detectable signature into AI-generated protein sequences and structures, with Nature research showing the watermark can survive without significantly disrupting biological function

Futuristic illustration of blue protein structures moving through a transparent digital scanning checkpoint ETIH coverage of Google DeepMind’s SynthID Bio research into watermarking AI-designed proteins.

Google DeepMind’s SynthID Bio embeds detectable watermarks into AI-generated protein sequences and predicted structures while aiming to preserve their biological function, Image credit: Google

Google DeepMind researchers have demonstrated a way to watermark AI-designed proteins while preserving their biological function, an attempt to bring the provenance systems already used for AI-generated media into synthetic biology.

The new method, SynthID Bio, puts a detectable signature directly into either a protein’s amino acid sequence or its predicted three-dimensional structure.

The central challenge is very different from watermarking an AI-generated image. Changing a few pixels may leave a picture effectively unchanged. Altering a protein can change how it folds, binds to other molecules or functions altogether.

DeepMind’s researchers had to show not simply that their watermark could be detected, but that the resulting biological designs still worked.

In laboratory testing reported in Nature, watermarked protein binders remained functional across three targets: VEGF-A, PD-L1 and the receptor-binding domain of the SARS-CoV-2 spike protein.

Pushmeet Kohli, Chief Scientist at Google Cloud and VP Science at Google DeepMind, described the result on LinkedIn as the “successful synthesis of AI-designed proteins that are both functional and watermarked.”

How do you watermark a protein?

SynthID Bio actually covers two different approaches. For protein sequences, the system subtly changes how amino acids are selected while an AI model generates the sequence. The resulting signature can then be detected using a secret watermarking key.

For three-dimensional protein structures, the researchers took a different route. They fine-tuned part of AlphaFold 3 so the watermark is built into the predicted atomic coordinates produced by the model.

The recommended version of that structural method achieved a detection rate above 99.8% at a 0.1% false-positive rate in the researchers’ tests, without reducing key structural accuracy measures compared with the AlphaFold 3 baseline.

The more important test for the sequence watermark happened outside the computer.

Researchers synthesized watermarked and non-watermarked protein binders and tested their ability to attach to the three target proteins. The study found no significant population-level differences in binding affinity between watermarked and non-watermarked binders.

The results were not identical across every measure. At one less stringent binding-affinity threshold, non-watermarked proteins recorded a significantly higher hit rate than one of the watermarked settings.

That makes SynthID Bio a proof of concept rather than a finished biological provenance system. The Nature paper explicitly describes it that way, arguing that operational use in biosecurity and scientific databases would require further technical work, coordination and standardization.

AI-generated biology creates a provenance problem

The reason DeepMind is pursuing the technology is straightforward: generative AI can now create biological sequences that did not previously exist.

Systems including AlphaProteo and other protein-design models can generate new proteins, while genomic models are beginning to move into more complex biological design.

That complicates existing biosecurity systems. DNA synthesis companies screen orders against databases of known threats. A completely novel AI-generated sequence may have little resemblance to something already cataloged, making it more difficult to determine whether the design is benign without additional review.

SynthID Bio is designed to add another signal. If a synthesis provider can verify that an unfamiliar sequence was generated by a known AI system with established safeguards, that information could become part of the screening process.

Sarah Carter, Principal at Science Policy Consulting, describes it as “an important piece of the puzzle for tracking the provenance of biological designs.”

The same issue applies to scientific databases. Resources such as the Protein Data Bank, UniProt and GenBank are used to train and operate biological research tools. The researchers argue that large volumes of AI-generated material entered without accurate labeling could undermine the integrity of those datasets.

A detectable watermark could make AI-generated sequences or structures easier to flag before they are treated as naturally occurring biological data.

Kohli put the limits of the approach clearly in his LinkedIn article: “Watermarking won’t solve biosecurity by itself, but it is an important step towards a safer and more resilient world.”

The watermark can still be removed

The research also identifies weaknesses that matter if SynthID Bio is eventually used as a security mechanism.

The sequence watermark can be removed through resequencing with ProteinMPNN. In tests covering 38,396 binders, that approach effectively stripped out the watermark.

Doing so can come at a cost to the resulting protein’s likelihood of functioning, particularly without access to the researchers’ filtering process, but the watermark itself is not resistant to the attack.

The structure watermark has a different weakness. It survived small amounts of digital noise and basic transformations, but a constrained structural relaxation process was able to destroy it.

Both versions currently use what researchers call a zero-bit watermark. In simple terms, the detector can establish whether the watermark is present, but it does not carry richer information that could distinguish between multiple users or encode additional provenance data.

The study was funded by Alphabet, and all of its authors are Alphabet employees who may own company stock as part of their compensation.

DeepMind is making its methods paper, code and in vitro data available, alongside access information for the recommended SynthID Bio structure model weights.

The work is already being extended beyond individual proteins. In an ongoing collaboration with the Hie lab at Stanford University and Arc Institute, researchers have integrated SynthID Bio into the Evo 2 genomic model to watermark the genome of an AI-designed bacteriophage.

Early laboratory testing in bacterial cultures found that those watermarked bacteriophages remained functional. DeepMind says further technical details will follow.

Previous
Previous

Former US Surgeon General Vivek Murthy to lead Common Sense Media’s youth AI safety push

Next
Next

Cambridge-backed AI startup Zenithon raises $10m to accelerate fusion, aerospace and chip design