Tricky Cases in Protein Preparation

by Ari Wagen · August 20, 2026

Preparing proteins is tedious, time-consuming, and absolutely essential to working with experimental protein structures. Our protein-preparation workflow aims to turn it into a routine, callable function so human and AI scientists can spend their time on higher-level science instead.

To get there, we've been stress-testing it against crystal structures from the RCSB PDB to expose weaknesses in our pipeline. Since launching the protein preparation workflow last month, we've already made substantial progress.

Improvements

Before getting into the results, we want to highlight two major changes we've made since announcing the workflow.

First, we've added PROPKA 3 as a protonation method, alongside our existing OpenMM and protonate_utils methods. PROPKA predicts the pKa values of ionizable sidechains, and we combine these predictions with the requested pH to assign protonation states to each residue. More information about PROPKA is available on GitHub.

Second, we've replaced our templated Boltz-2 method with an inpainting-based approach adapted from PATCHR. This approach better preserves the input crystal structure while generating more realistic structures for unresolved regions. More information about the original PATCHR program is available on GitHub.

We've also made performance improvements throughout the workflow, substantially reducing the cost of preparation.

Performance

We ran all of the systems below through Rowan's protein preparation workflow using PROPKA 3 protonation, ACE/NME caps, pH 7.4, and "Retain existing protonation?" enabled.

For 5FQD, we reduced the system from a dimer to a monomer; for 2E32, we kept the full system.

All ligands and ions were retained during preparation with one exception: the ligand with residue name AIJ1T in 9RR8. Our workflow currently cannot retain non-polymer residues with names that are five characters or longer, so this ligand was dropped and a warning was issued. (The PDB file format specification only allocates three characters for residue names, though some programs break this; OpenMM does not handle natively the five character case.)

We didn't run the old templated Boltz-2 method on 8HNC: we identified this test case after already retiring that pipeline in favor of Boltz-2 inpainting.

Complete Structures

PDB IDResiduesMethodBackbone breaksHeavy-atom RMSD (Å)Credits
TotalMissing
9RR81380PDBFixer00.150.94
Boltz-2 inpainting00.279.41
Templated Boltz-200.5348.85
1A422590PDBFixer00.010.85
Boltz-2 inpainting00.199.74
Templated Boltz-200.6150.91
1OTP4400PDBFixer00.141.02
Boltz-2 inpainting00.267.52
Templated Boltz-200.7536.53

Incomplete Structures

PDB IDResiduesMethodBackbone breaksHeavy-atom RMSD (Å)Credits
TotalMissing
8HNC711141PDBFixer01.373.10
Boltz-2 inpainting11.2018.14
2E32926192PDBFixer20.153.88
Boltz-2 inpainting00.4511.68
Templated Boltz-21917.3654.68
5FQD1623163PDBFixer150.155.22
Boltz-2 inpainting10.3017.55
Templated Boltz-2180.8770.40

For systems with no unresolved residues, PDBFixer and Boltz-2 inpainting perform similarly. For harder systems, PDBFixer and the old templated approach both start to fall apart. A few of these systems are particularly useful for showing where the workflow performs well and where there's still room to improve.

Here are three interesting and tricky cases:

8HNC: Protein Adducts & Large Unresolved Domains

The 8HNC crystal structure is a membrane protein with a bilirubin ligand bound in its transmembrane domain and two NAG protein adducts in its intracellular domain.

The 8HNC crystal structure

The 8HNC crystal structure

141 of this structure's 711 residues are unresolved. PDBFixer reconstructs this large extracellular region poorly, while Boltz-2 inpainting produces a much more plausible structure.

8HNC prepared with PDBFixer and Boltz-2 inpainting

Both methods produce minor geometric issues, but neither prevents us from successfully building a force field for the resulting system:

Neither PDBFixer nor Boltz-2 inpainting is able to retain the covalent protein adducts.

1A42: Metal Ions

Many proteins contain bound metal ions, and it's important to preserve local coordination geometry and protonate nearby sidechains in a metal-aware manner.

Our previous Boltz-2 preparation scheme (left) did not handle the mercury in 1A42 correctly. The combination of our new Boltz-2 inpainting scheme and updated protonation logic (right) handles this test case appropriately:

Detail of a cysteine-mercury bond in templated Boltz-2 and Boltz-2 inpainting prepared
structures

5FQD: Pushing Scale

Finally, we wanted to push the new Boltz-2 inpainting approach beyond the scale of a typical protein-preparation job by preparing a 1,623-residue system derived from 5FQD. Inpainting completed 4x faster than our old templated co-folding method (17.55 vs. 70.40 credits) and produced a structure with fewer potential issues.

However, Boltz-2 can become less reliable beyond 1,000 residues, and this stress test did expose two lingering issues. The first was minor: a single C–CA bond with a length of 2.1 Å.

The more interesting issue involved the LVY ligand. Part of the ligand lost its expected ring geometry, leaving several floating atoms:

To further explore this strange behavior, we ran the same preparation twice more:

Each of the 3 runs cost between 17 and 22 credits. This suggests that Boltz-2 inpainting can be stochastic for these very large system sizes. For systems of this scale, we recommend inspecting the resulting structure and falling back to PDBFixer if needed.

Conclusion

With these improvements, protein preparation in Rowan is fast, automated, and robust across a wide range of experimental structures. Because preparation sits upstream of docking, MD, and FEP, we see making this step reliable and routine as a core part of the platform.

Our recommended starting point is to use Boltz-2 inpainting to add missing atoms and PROPKA 3 to assign protonation states. This entire process can also be run through Rowan's Python API, making it easy to incorporate protein preparation into larger automated workflows.

There are still a few known limitations we hope to address:

While we're quite happy with our work here, we'd love to keep improving. If you spot something that doesn't work well, please tell us!

Banner background image

Start running calculations in minutes!

Our platform lets you submit, view, analyze, and share calculations using cutting-edge methods trusted by hundreds of leading scientists. We give every new user 500 free credits to start, plus more every week. Making an account and running your first calculation takes only seconds: start using Rowan today!

Start computing →

What to read next

Tricky Cases in Protein Preparation

Tricky Cases in Protein Preparation

How we're using Boltz-2 inpainting to improve protein preparation.
Aug 20, 2026 · Ari Wagen
Running a Full FEP Campaign in Python with Rowan

Running a Full FEP Campaign in Python with Rowan

Learn how to run an iterative FEP campaign programmatically with Rowan's Python SDK.
Aug 19, 2026 · Eli Mann
Phonons

Phonons

band structures and density of states; material waves; sound, light, and heat
Aug 18, 2026 · Raphael Stone and Jonathon Vandezande
Binder Optimization with LLMs and Specialized Models

Binder Optimization with LLMs and Specialized Models

Testing LLMs and specialized models on generating strongly binding ligands.
Aug 17, 2026 · Ishaan Ganti
Performance Optimization

Performance Optimization

or, how to get more Rowan for your dollar
Aug 10, 2026 · Corin Wagen and Eli Mann
Hits From a Hackathon

Hits From a Hackathon

How certain schemes to identify TBXT-binding compounds have succeeded.
Aug 7, 2026 · Kat Yenko, Corin Wagen, Ari Wagen, and Derek Alia
Testing Different Pose-Ranking Methods for RBFE Calculations

Testing Different Pose-Ranking Methods for RBFE Calculations

Benchmarking how well Rowan's analogue-docking pose scoring picks the best starting structure for RBFE, and how much a ranking miss actually matters downstream.
Aug 6, 2026 · Zachary Fried
LogP, API Key Budgets, and a User Survey

LogP, API Key Budgets, and a User Survey

the golden mean of logP; three approaches to predicting logP; API key budgets for low-trust delegation; a user survey and some blog posts
Aug 5, 2026 · Nick Casetti, Ari Wagen, Spencer Schneider, and Corin Wagen
How to Find Conformers

How to Find Conformers

A conceptual overview of conformer-generation and conformer-search methodology.
Aug 3, 2026 · Nicholas Casetti
Simulation Tools Improve Agent Problem-Solving

Simulation Tools Improve Agent Problem-Solving

External simulation tools do noticeably improve agent performance at 13C NMR structural elucidation.
Jul 31, 2026 · Corin Wagen