INTRODUCTION
Proteases are among biology’s most powerful regulators
1,2. By cutting specific peptide bonds they activate hormones
3, terminate signaling cascades
4, remodel extracellular matrices
5, and control protein homeostasis
1. Their catalytic activity also underpins a wide spectrum of industrial and medical applications — from laundry detergents and food processing to wound debridement and thrombolytic therapy
6,7. Yet the utility of natural proteases is circumscribed by their hard-wired substrate preferences
8,9. An enduring objective is therefore to create “molecular scalpels” whose cleavage site can be programmed at will, enabling precise dissection of cellular pathways and highly selective therapeutic intervention
10,11. Despite decades of effort, very few platforms deliver custom proteases that combine robust catalytic efficiency with single-bond selectivity
12-14.
Metalloproteases are prime candidates for such re-engineering, and their catalytic chemistry has been well studied
15-17. In a typical zinc metalloprotease, the zinc ion is coordinated by three amino-acid ligands and positioned adjacent to a general base that activates a water molecule for nucleophilic attack on the peptide bond, while a neighbouring residue stabilises the tetrahedral oxyanion intermediate
17 (Fig. 1a, b). This minimalist engine is embedded in diverse scaffolds — including matrix metalloproteinases
18, ADAMs
19, astacins
20, thermolysins
21, and bacterial virulence factors
22 — that collectively drive processes from extracellular-matrix turnover to cytokine maturation. Despite their natural versatility, repurposing these enzymes for novel targets remains challenging. While efforts to reprogram selectivity through directed evolution or rational mutagenesis have yielded promising advances
12,13,23, certain structural and functional limitations persist. As a result, repurposed natural metalloproteases have achieved only modest improvements in selectivity and therapeutic potential
24.
De novo protein design holds the promise of sidestepping the intrinsic limitations of natural enzymes by constructing entirely new scaffolds around a chosen chemical task
25-29. Recent breakthroughs in deep-learning-based structure generation have indeed delivered artificial metallohydrolases whose catalytic efficiencies approach those of native enzymes
26,28,30. However, these successes have been confined almost exclusively to model reactions such as the hydrolysis of activated ester bonds in small, hydrophobic substrates — reactions that present relatively low activation barriers and low specificity demands. Extending this progress to true proteolysis is vastly more challenging. Amide bonds are orders of magnitude more inert than esters
31,32, and a protease must not only accelerate their cleavage but also capture a highly flexible, polar peptide and orient one particular bond in a near-attack conformation. Moreover, practical applications require selective sequence recognition, such that the designed substrate-binding groove preferentially accommodates the target sequence while minimizing activity toward closely related peptides. Natural proteases achieve this balance with folds that have been evolutionarily honed for cognate substrates; even slight structural perturbations can disrupt the sub-angstrom precision of the catalytic residues
9,33. Consequently, conventional sequence redesign has limited success in retargeting these enzymes while maintaining native-like catalytic efficiency
33-36. The central challenge is to design new structures that can simultaneously preserve the strict metal-dependent catalytic geometry and tightly lock the scissile bond into position, a goal that has remained previously unattainable.
RESULTS
Computational strategy for designing metalloproteases with programmed substrate selectivity
We set out to develop a general strategy for the
de novo design of substrate-selective metalloproteases. To establish catalytic motifs for design, we focused on two representative zinc-dependent endopeptidases: PPEP-1 (PDB: 6R4W)
37 and Fungalysin (PDB: 4K90)
38. Both enzymes are secreted by microorganisms and play critical roles in survival or pathogenicity by hydrolyzing extracellular host or cell-surface proteins. Despite sharing a common catalytic mechanism, these two enzymes exhibit distinct active site architectures. Notably, their oxyanion hole compositions differ: PPEP-1 employs a tyrosine, whereas Fungalysin utilizes a histidine residue. Structurally, PPEP-1 features a broad, clamp-like groove formed by two flexible loops that close over the substrate, conferring high sequence selectivity for cleavage at Pro–Pro amide bonds. To reprogram this sequence selectivity, we developed a strategy that removes the peptide-binding loops from PPEP-1, while retaining the core protein structure (Fig. 1c). The excised regions were then rebuilt, enabling the design of new peptide-binding interfaces. Proteases designed through this approach are denoted “PP” (peptide-binding pocket designed protease) followed by a numerical identifier. In contrast, Fungalysin lacks an extended, pre-formed substrate-binding pocket and instead features an open substrate binding groove, resulting in a broader substrate selectivity for peptide bonds flanked by hydrophobic side chains. For this scaffold, we extracted only the minimalist catalytic motif — comprising the zinc-coordinating residues, the general base glutamate, and the oxyanion hole histidine (Fig. 1c). This minimalist approach necessitates
de novo generation of the entire surrounding protein scaffold to accurately position the catalytic motif and design a novel peptide-binding interface for substrate recognition. Proteases generated using this minimalist catalytic motif are designated “DP” (
de novo protease).
To enable precise co-design of enzyme functionality and substrate recognition, we employed ProteusFlow, our generative flow-based model significantly advanced over its predecessor
39. To explicitly endow the structural generative process with a deep understanding of substrate sequence constraints, we integrated a customized protein language model pre-trained on UniRef50
40 with an auxiliary structural decoding objective (see Materials and Methods). ProteusFlow was trained on experimental structures from the Protein Data Bank
41 and high-confidence AlphaFold2
42 predictions using a dual-modality motif scaffolding task. In this framework, sequence and structure conditions are treated as independent masking tracks, leveraging geometry-aware sequence semantics to control structure generation. This decoupled conditioning strategy empowers the model to process hybrid inputs comprising structurally defined catalytic motifs and coordinate-free substrate sequences, thereby generating metalloproteases with stabilized catalytic sites and selective substrate recognition.
We leveraged ProteusFlow to co-generate metalloprotease–substrate complex structures using the extracted active motif configurations as templates. To determine the catalytically competent conformation of the substrate relative to the catalytic residues, we used AlphaFold3
43 to model the substrate peptide configuration within the template protease. The predictions revealed that the substrate adopted a β-strand conformation in both enzymes (Supplementary Fig. S1). From these predicted structures, we extracted the P2–P1′ backbone atoms of the Fungalysin substrate peptide (where P1 is the residue immediately preceding, and P1′ immediately following, the cleavage site), and the P3–P1 backbone atoms of the PPEP-1 substrate. In both cases, these segments form strand pairing interactions with the enzyme's strand adjacent to the active site, thereby enabling the construction of a complete target active site. To maximize structural complementarity for the substrate peptide, we developed a “two-step encapsulation” strategy based on ProteusFlow (Fig. 1c, d). Unlike conventional one-step generative approaches that frequently yield suboptimal packing and limited control over global enzyme topology, our method operates in a sequential manner. The first step generates the protein backbone surrounding the catalytic glutamate, stabilizing the general base and achieving partial pre-wrapping of the substrate. In the second step, the scaffold is constructed around the oxyanion hole to complete substrate peptide encapsulation. This stepwise process minimizes steric clashes, substantially enhances the packing density at the enzyme–substrate interface, and consistently generates clamp-like structures with improved substrate recognition (Supplementary Fig. S2). Following backbone generation, protein sequences were optimized using LigandMPNN
44, and the resulting protease–substrate–zinc complex structures were validated with AlphaFold3 to ensure convergence to the intended catalytic geometry.
De novo proteases facilitate the direct hydrolysis of amyloid-β peptide
We selected the amyloid-β (Aβ) peptide as the substrate for design, given its central role in Alzheimer’s disease (AD) pathology, where the aggregation of neurotoxic Aβ plaques is a hallmark event
45. The pathological self-assembly of Aβ is primarily driven by intermolecular interactions within two hydrophobic regions: the central hydrophobic core (residues 17–21) and the C-terminal segment (residues 30–42) (Fig. 2a). These domains act as critical nucleation sites for β-sheet stacking, facilitating the formation of stable, neurotoxic fibrils
46,47. Consequently, these regions represent key therapeutic targets for catalytic interventions aimed at destabilizing amyloid architecture and mitigating AD progression
48. To demonstrate the generalizability of our computational methodology, we selected three distinct cleavage sites within Aβ: one within the central hydrophobic core and two within the C-terminal hydrophobic region. We employed both the PP and DP strategies to design metalloproteases targeting each site. On average, 26 designs were selected from each strategy–substrate pair for experimental characterization (Supplementary Table S2). These designed enzymes were expressed in
Escherichia coli, purified using Ni-NTA magnetic beads, with the His-tag subsequently removed via HRV3C protease. Final purification was performed by size exclusion chromatography. To assess catalytic activity, we engineered fusion substrates in which the targeted Aβ fragment was inserted between the mannose-binding protein (MBP) and a small ubiquitin-like modifier (SUMO) tag. Successful proteolytic cleavage results in two lower-molecular-weight fragments, readily detectable by SDS-PAGE (Fig. 2b). Upon overnight incubation of the designed proteases with their respective substrates in the presence of zinc, we observed that the DP strategy yielded active enzymes for all three targeted sites: DP720-S1, DP622-S2, and DP221-S3. In contrast, the PP strategy generated active designs only for the first site, resulting in PP507-S1 and PP532-S1 (Fig. 2c). We attribute the limited success of the PP designs at the remaining sites to steric constraints imposed by excessive preservation of the template backbone, which may be incompatible with the sequence and geometry of these cleavage sites. In contrast, the DP strategy allows broader sampling of the conformational space, facilitating the discovery of active variants. Mass spectrometry analysis of the cleavage products confirmed proteolytic cleavage-site fidelity, as the observed molecular weights perfectly matched theoretical predictions (Supplementary Fig. S3). Notably, for DP221-S3, in addition to the major mass spectrometric peak corresponding to the intended cleavage products, a very minor peak — representing approximately 10% of the main product’s abundance — was also observed, indicating an alternative cleavage site shifted by one residue toward the C-terminus (Supplementary Fig. S3e). We next characterized their catalytic efficiency by monitoring the time course of substrate cleavage (Supplementary Fig. S4). At a 1:2 enzyme-to-substrate molar ratio, DP622-S2 exhibited the highest activity, achieving complete substrate degradation within 4 h (Supplementary Fig. S4b). DP720-S1 and PP507-S1 achieved full digestion in 8 h (Supplementary Fig. S4a, c), while PP532-S1 required nearly 12 h for complete cleavage (Supplementary Fig. S4e). In contrast, DP221-S3 displayed slower kinetics, cleaving more than half of the substrate after 12 h (Supplementary Fig. S4d). Collectively, these results indicate that our design strategies enable the
de novo generation of metalloproteases with selective sequence recognition, capable of targeting disease-relevant sites within Aβ.
Validation of catalytic mechanism
To elucidate the catalytic mechanism and confirm the functional roles of key active site residues, we conducted a series of site-directed mutagenesis experiments (Fig. 3a). For designs generated by both the DP and PP strategies, the catalytic glutamate and oxyanion hole residues (histidine in the DP designs and tyrosine in the PP designs) were mutated to alanine (Fig. 3b). Additionally, in the DP designs, we mutated the newly generated second-sphere residues that form hydrogen bonds with the zinc-coordinating residues, to evaluate their contribution to catalytic activity. Functional assays revealed that substitution of the catalytic glutamate with alanine universally abolished proteolytic activity across both DP and PP designs, highlighting its indispensable role as the general base in catalysis (Fig. 3b). Interestingly, mutagenesis of the oxyanion hole residues revealed complex functional responses. In the DP designs, mutation of the oxyanion hole histidine to alanine had minimal or even positive effects on protease function; notably, the H191A variant of DP221-S3 exhibited enhanced substrate degradation. By contrast, in the PP designs, alanine substitution of the oxyanion hole tyrosine resulted in partial reduction of activity. These divergent responses suggest that the mechanisms of transition state stabilization are highly complex. Furthermore, mutating the second coordination sphere in the DP designs yielded distinct activity profiles: the D128A mutation in DP720-S1 and the N111A mutation in DP221-S3 markedly reduced activity, whereas in DP622-S2, single mutations (D126A or Y91F) had minimal impact but the double mutant (D126A/Y91F) dramatically decreased activity (Fig. 3b). Thus, peripheral interactions remain context-dependent; although mutational data align with the designed active sites, the inherent structural complexity highlights the challenge of achieving fully predictive catalytic control.
Characterization of the catalytic kinetics
To quantitatively assess the kinetics of our designed proteases, we developed a specialized fluorescence-based assay. The intrinsic hydrophobicity of the Aβ target sequences posed significant challenges for synthesizing peptides compatible with standard fluorophore-quencher systems, due to marked solubility and aggregation issues. As an alternative, we engineered a FRET-based biosensor by inserting the targeted Aβ sequences directly between the fluorescent protein pair mTurquoise2 (donor) and mVenus (acceptor) (Supplementary Table S3). In this system, enzymatic cleavage of the substrate separates the fluorescent proteins, resulting in a measurable decrease in FRET efficiency. Kinetic parameters were determined by fitting initial reaction velocities to the Michaelis-Menten model (Fig. 3c). Among the designed proteases, DP622-S2 demonstrated the highest catalytic efficiency, achieving a
kcat/
Km of 325.26 M
−1s
−1, driven by a favorable combination of a high
kcat (0.00191 s
−1) and a low
Km (5.87 μM). The PP507-S1 and PP532-S1 exhibited moderate efficiencies of 54.67 and 48.09 M
−1s
−1, respectively, while DP221-S3 and DP720-S1 displayed the lowest efficiencies at 9.50 and 6.06 M
−1s
−1. With the exception of DP720-S1, these kinetic findings were in close concordance with our time-course degradation assays. The discrepancy observed for DP720-S1, which displayed effective substrate degradation comparable to PP507-S1 in previous assays but low catalytic efficiency in the FRET system, may be attributable to steric hindrance wherein the bulky fluorescent modifications restrict substrate access to the active site. Notably, the
Km value of our lead design, DP622-S2, closely parallels those of naturally occurring Aβ-degrading proteases such as Insulin-Degrading Enzyme, which also exhibit
Km values in the low micromolar range
49. Although these designed enzymes exhibit lower turnover rates than natural biocatalysts, achieving micromolar substrate affinity confirms that our computational strategy can successfully design target-selective recognition.
To further elucidate the functional roles of the oxyanion hole and second coordination sphere residues, we characterized the catalytic kinetics of the corresponding mutants, including single mutants of DP622-S2 (Y91F, D126A, or H172A) and DP221-S3 (H191A), as well as the DP622-S2 double mutant (Y91F/D126A) (Supplementary Fig. S5). Intriguingly, the tested DP622-S2 single mutants (Y91F and D126A) exhibited enhanced catalytic efficiency (kcat/Km), driven predominantly by increased turnover (kcat) (Supplementary Fig. S5a, b). Conversely, the H172A variant and the Y91F/D126A double mutant showed decreased efficiency (Supplementary Fig. S5c, d), consistent with the SDS-PAGE results. Furthermore, the DP221-S3 H191A mutation achieved an approximately twofold increase in catalytic efficiency through a coupled improvement in both turnover rate (kcat) and substrate affinity (Km) (Supplementary Fig. S5e). These findings show that reshaping the local catalytic environment simultaneously refines both turnover rates and binding parameters, directly shaping overall enzyme efficiency.
Characterization the cleavage-site fidelity
We next assessed the substrate discrimination capability of the designed metalloproteases by evaluating their cleavage activity against three distinct MBP–Aβ_segment–SUMO fusion constructs (Substrates 1, 2, and 3) (Supplementary Fig. S6). SDS-PAGE analysis of the cleaved products revealed varying degrees of selectivity. PP507-S1 was strictly selective, exclusively cleaving Substrate 1 (Supplementary Fig. S6b), whereas PP532-S1 displayed broad tolerance, processing all three substrates (Supplementary Fig. S6c). The remaining designs exhibited intermediate discrimination: DP720-S1 showed reduced off-target activity against Substrate 2 (Supplementary Fig. S6a), while DP622-S2 and DP221-S3 displayed weak cross-reactivity toward each other’s substrates (Supplementary Fig. S6d, e). Subsequent mass spectrometric analysis of the cross-reactive products confirmed that the enzymes cleaved the alternative substrates at distinct, non-intended sites (Supplementary Fig. S7). These cross-reactivities are likely related to substrate sequence differences. The S1 cleavage site is flanked by a high frequency of bulky aromatic and large hydrophilic residues, contrasting sharply with the S2 and S3 flanking regions that are predominantly composed of small aliphatic amino acids and glycines. This divergence might explain the exceptional cleavage-site fidelity of PP507. To rationalize these profiles computationally, evaluation using ProteinMPNN was able to predict some selectivity trends of the experimentally validated designs to a certain degree, such as the strict target preference of PP507 and the high promiscuity of PP532 (Supplementary Fig. S8). However, the correlation between ProteinMPNN scores and substrate selectivity remains weak, as seen in the intermediate cross-reactivities of the DP series. Additional experimental results are necessary to clarify substrate selectivity and computational prediction methods.
We next evaluated the substrate discrimination capability of the active site mutants (Supplementary Fig. S9). The second-shell and oxyanion hole mutants in DP622 and DP720 exhibited cleavage profiles similar to those of their respective parent designs (Supplementary Fig. S9a, b). Notably, the DP221-S3 H191A variant showed a remarkable increase in selectivity, virtually abolishing cleavage activity toward non-intended substrates (Supplementary Fig. S9c). Mass spectrometry analysis of the cleavage products revealed that the DP622-S2 Y91F mutant retained the intended cleavage site, consistent with the parental enzyme (Supplementary Fig. S10a). In contrast, the scissile bond in DP221-S3 H191A variant shifted almost entirely to the original minor cleavage site — one residue closer to the C-terminus (Supplementary Fig. S10b).
Characterization the cleavage-site fidelity on Aβ42
To evaluate the cleavage activity of our designs on the native substrate, we incubated the designed metalloproteases (PP507-S1, PP532-S1, DP720-S1, DP622-S2, DP221-S3) and mutant variants (DP622-S2 Y91F and DP221-S3 H191A) with a full-length synthetic Aβ42 peptide (Supplementary Figs. S11–15). Mass spectrometry analysis demonstrated that all tested enzymes successfully cleaved the Aβ42 peptide, with the sole exception of DP720-S1 which exhibited no activity. PP507-S1, DP622-S2, and the two mutants demonstrated sequence selectivity, generating the intended cleavage products (Supplementary Figs. S11a, S12b). Furthermore, the peptide fragments generated by PP532-S1 and DP221-S3 aligned with the multi-site cleavage profiles previously observed for these designs (Supplementary Figs. S11b, S13). Next, we incubated the substrate with a cocktail of selected variants (PP507, DP622 Y91F, and DP221 H191A), demonstrating that these enzymes could simultaneously process the Aβ42 peptide without mutual interference (Supplementary Fig. S16). Together, these results demonstrate that our de novo metalloproteases can cleave the full-length Aβ42 peptide, with most designs successfully retaining recognition of their intended cleavage sites.
Structural characterization of the protease–substrate complex
To elucidate the structural details, we determined the cryo-EM structures of PP507-S1, DP622-S2, and DP221-S3 in complex with the Aβ42 peptide (Fig. 4a). To overcome the resolution limitations inherent to the small molecular weight of the designed proteases, each protease was rigidly fused with a large helical bundle protein generated by ProteusFlow (Supplementary Fig. S17). Minimal interfacial interactions were computationally introduced to rigidify the fusion construct, ensuring that neither the overall fold nor the active site architecture of the protease was perturbed. To capture the pre-catalytic Aβ42-bound state, the catalytic glutamate in each protease was mutated to glutamine (E to Q). The resulting fusion proteins, with inactivated active sites, were incubated with the Aβ42 peptide and subjected to cryo-EM structure determination (Supplementary Fig. S18). Focused refinement of the protease region within the fusion protein yielded structures at resolutions ranging from 2.98 to 3.44 Å (Supplementary Fig. S19). The reconstructions revealed clear, continuous density for the Aβ42 substrate, allowing confident assignment of the peptide backbone and local side chains within the binding cleft (Fig. 4b). The resolved peptide sequence confirms that all three proteases recognize the designed Aβ segment, consistent with their high substrate cleavage-site fidelity.
The cryo-EM structures closely matched their computational models, with Cα RMSD values ranging from 1.13 to 1.88 Å (Fig. 4c). Moreover, the designed catalytic geometries were largely preserved in the experimental active sites (Fig. 4d). Evaluation of active-site geometries revealed small local structural deviation in PP507-S1 and DP221-S3, where P1 oxygen–zinc and P1–oxyanion hole distances expanded to 5.1/4.4 Å and 3.2/4.4 Å, respectively, exceeding their designed ~2 Å values. In contrast, DP622-S2 maintained its intended spatial arrangement, exhibiting a P1 oxygen–zinc distance of 2.3 Å (designed 2.1 Å) and a 2.8 Å distance to the oxyanion hole. This structural arrangement aligns with our kinetic observations (Fig. 3c), where DP622-S2 exhibits higher substrate affinity and improved turnover compared to PP507-S1 and DP221-S3. While our structures strongly validate the successful design of global protein scaffolds, they also reveal that achieving reliable sub-angstrom catalytic preorganization remains a significant challenge.
Structure-guided optimization of catalytic activity
To further optimize the catalytic efficiency, we developed a two-stage structural refinement strategy driven by ProteusFlow, aiming to enhance both the overall architecture and substrate binding capabilities (Fig. 5a). DP622-S2 was selected as the optimization target. In the first stage, we strictly preserved the geometry of the catalytic motif while applying spatially weighted structural perturbations based on binding potential. Regions with close substrate contacts underwent minimal perturbation, whereas regions lacking direct interactions were subjected to more extensive structural resampling. In the second stage, we further optimized the most promising variant from the first stage. While keeping the core structural scaffold and the binding interface strictly fixed, we selectively applied generative refinement to the adjacent regions. Specifically, by extending the N terminus of the Aβ42 substrate, the flow model was guided to construct additional β-sheet motifs that pair with the extended sequence, thereby systematically enhancing overall binding affinity.
The initial round of structural optimization yielded the OP609-S2 variant, which exhibited superior substrate selectivity compared to DP622-S2 by exclusively cleaving the S2 substrate (Fig. 5d). Its catalytic efficiency (kcat/Km) reached 3045.14 M−1s−1 (Fig. 5c), a nearly ten-fold increase over DP622-S2 that stems largely from an elevated kcat. Mass spectrometry validated that OP609-S2 accurately cleaved both the MBP-SUMO fusion protein and Aβ42 into their intended products (Supplementary Fig. S20). A second round of generative refinement on OP609-S2 produced the OP669-S2 variant. Although its overall efficiency fell slightly below OP609-S2, it remained higher than the original DP622-S2 (Fig. 5c). OP669-S2 was characterized by a notably smaller Km, reflecting enhanced substrate affinity, which plausibly caused a slight trade-off in its kcat. OP669-S2 cleaved both S2 and S3 MBP-SUMO substrates (Fig. 5d). However, mass spectrometry analysis revealed a consistent cleavage site at S2 for both S2 and S3 MBP-SUMO substrates (Supplementary Fig. S21a, b). These results suggest that the generative structural optimization successfully granted the enzyme enhanced recognition of the substrate N-terminus. It also showed a more complete degradation of Aβ42 into intended fragments compared to OP609-S2 (Supplementary Fig. S21c). This enhanced clearance confirms that tighter substrate engagement directly promotes thorough proteolytic processing.
DISCUSSION
By explicitly integrating target sequence information into the flow-based structure generation process, ProteusFlow enables the generation of metalloproteases tailored for precise sequence recognition. Our two-step encapsulation strategy facilitates the creation of clamp-like proteases that envelop the substrate peptide, maximizing cleavage site fidelity. The successful design of metalloproteases capable of cleaving three distinct native Aβ peptide sequences extends de novo protein design beyond stable binding into functional proteolysis. Structurally, the global agreement between design models and experimental complexes highlights the accuracy of our backbone generation; however, localized deviations in the active site reveal that precise sub-angstrom catalytic geometry remains a critical area for future improvement.
While our computational framework successfully generates enzymes capable of processing custom targets with native-like substrate affinities (Km), optimizing the catalytic turnover rate (kcat) proves exceptionally challenging. The near-ideal β-strand substrate conformation observed in our designs likely reflects an inherent bias of current deep-learning generators toward regular secondary structures. These extensive main-chain hydrogen bonds may over-stabilize the ground state, restricting the dynamic flexibility needed for optimal catalytic kinetics. Moreover, our mutational insights demonstrate that the complex microenvironmental elements required for efficient turnover remain difficult to predict with static design frameworks. This indicates that the exact rules for rational enzyme design are not yet complete. To overcome this theoretical gap, we transitioned toward a more empirical exploration of the local conformational landscape. This structure-guided introduction of targeted structural perturbations successfully produced a remarkable increase in kcat for a selected variant. Although its inherent stochasticity makes this process difficult to replicate exactly, this approach offers a viable tool for functional optimization.
Achieving absolute substrate selectivity remains a nuanced challenge in de novo enzyme design. Our biochemical data indicate that while our clamp-like architectures successfully enforce strict cleavage site fidelity, macroscopic substrate selectivity varies among the designed variants. For instance, PP507-S1 exhibits exclusive selectivity for its intended target, measurable cross reactivity was observed in other variants. While sequence scoring with ProteinMPNN, to some extent, can predict broad cross-reactivity trends among alternative substrates for some designs, it struggles to capture subtle discrimination nuances. Therefore, to truly predict selectivity, future algorithms must move beyond static sequences and account for the dynamic catalytic processes. Additionally, future iterations must address our current reliance on biased extended β-strand conformations. This structural requirement restricts the approach primarily to unstructured targets like Aβ, as applying it to folded domains would impose a substantial energetic penalty for substrate unfolding.
Although monoclonal antibodies remain the mainstream clinical strategy for Aβ clearance, direct catalytic degradation represents a valuable exploratory direction. In this context, our
de novo design framework provides a highly programmable tool for targeted peptide hydrolysis. These engineered proteases selectively hydrolyze aggregation-prone regions within monomeric Aβ, providing a robust structural proof of concept. Demonstrating exquisite cleavage site fidelity alongside a relatively comprehensive degradation profile, our most selective variant, OP669-S2, effectively converts pathogenic Aβ42 into the benign Aβ38 isoform
50. While the present designs serve primarily as an early proof of concept, iteratively optimizing these enzymes to achieve superior catalytic efficiency, tighter substrate binding, and absolute selectivity could enable this approach to become a viable therapeutic strategy.
MATERIALS AND METHODS
Computational design of zinc metalloproteases
Design pipeline overview
We developed a computational framework for designing zinc metalloproteases targeting Aβ (Supplementary Fig. S22). The design workflow employed ProteusFlow
51, a structure-based generative model, to implement a two-step backbone generation strategy. This process separates the complex task of active site construction into two spatially distinct stages: first, prioritizing deep encapsulation of the substrate from the side of the catalytic glutamate residue; and second, precisely scaffolding the oxyanion hole to complete the active site and achieve robust substrate encapsulation. The foundation of this enzyme–substrate co-design strategy is the specialized architecture of ProteusFlow — a substantially new generative protein design model based on our original Proteus model. A key innovation in ProteusFlow is its capacity to process decoupled sequence and structure constraints, as detailed below.
ProteusFlow generative model
The ProteusFlow architecture integrates a structure-enhanced protein language model with a flow-matching structure generator
52,53. To fuse the modality between one-dimensional sequence semantics and three-dimensional geometry, the language model incorporates a structure-aligned encoder trained with an auxiliary structural decoding objective
54. This dual-task training strategy ensures that the resulting embeddings capture both evolutionary information and structural compatibility. These geometry-aware embeddings are then used as conditioning signals for the structural generative module. Importantly, the model is trained using a decoupled multi-modality masking strategy
55, treating sequence and structure as independently maskable variables. This framework enables flexible conditional generation, allowing for the simultaneous incorporation of fixed catalytic motifs and sequence-conditioned substrate peptide configurations — critical for the enzyme–substrate co-design strategy implemented in our metalloprotease design.
Enzyme structural templates and active motif extraction
To establish the structural foundation for enzyme generation, we selected two structurally distinct proteases as parental templates: (PDB: 6R4W)
37 and Fungalysin (PDB: 4K90)
38. Although both share a common catalytic mechanism, these enzymes exhibit different active site architectures. Notably, their oxyanion hole compositions differ — PPEP-1 relies on a tyrosine residue, while Fungalysin utilizes a histidine. To obtain the catalytic motif of the native enzyme–substrate complex required for template generation, we used AlphaFold3
43 to model the structures of each enzyme in complex with its endogenous substrate peptide.
To rigorously explore metalloprotease design methodology, we investigated two complementary approaches based on distinct motif extraction strategies: PP, which preserves the global fold of a native enzyme while regenerating the substrate-binding region; and DP, which assembles the entire enzyme around a minimalist catalytic core. Both strategies were applied to the non-native Aβ sequence, which bears no obvious similarity to the native substrates of the parental enzymes.
For the PPEP-1-based designs, we applied the PP strategy. We extracted the P3–P1 segment (substrate: VNPPVP
37, where P3–P1 corresponds to VNP; P1 is the residue immediately preceding, and P1′ is the residue immediately following, the cleavage site) backbone atoms from AF3-predicted PPEP-1–substrate complexes. This segment participates in β-strand pairing interactions with PPEP-1. Additionally, we retained most of the template enzyme structure, specifically residues 31–82, 115–117, 121–123, and 139–220, as the starting scaffold. Within this framework, the generative model was tasked primarily with redesigning the masked substrate-binding regions to accommodate the new Aβ target sequence.
For the Fungalysin–based designs, we implemented the DP strategy. We extracted the P2–P1′ substrate segment (native substrate: CGERGFFYTPKA
38; the specific extracted segment is RGF) and isolated a minimalist catalytic motif from the enzyme structure, strictly restricted to the essential catalytic core. This core included: the HEXXH helical segment (VIHEYTHGLS) housing two zinc-coordinating histidines; a second helix (GMGEGWSD) positioning the third zinc-coordinating glutamate; and a third helix (VHAIGTVWAS) providing the histidine-based oxyanion hole necessary for stabilization of the catalytic tetrahedral intermediate, together with a short β-strand pairing with the substrate. ProteusFlow was tasked with generating the entire supporting scaffold to incorporate this minimalist catalytic motif and bind with the substrate.
Two-step enzyme backbone generation
To achieve high substrate selectivity and catalytic efficiency, our design strategy must satisfy two key objectives: well-substrate encapsulation to maximize sequence-level recognition and accurate structural support for the catalytic motif. We observed that a direct, one-step motif scaffolding approach often resulted in insufficient substrate burial. While such one-step designs typically maintained essential catalytic geometry, they frequently exhibited suboptimal packing with the substrate peptide, leaving many substrate sidechains solvent-exposed. Such partial encapsulation might hinder the enzyme from sampling all substrate residues, consequently dampening substrate recognition. To address this limitation, we implemented a two-step generative strategy that partitions protease design into sequential structure generation stages, enabling precise, stepwise organization of the enzyme–substrate interface.
In the first stage, the primary objective was to anchor the Aβ segment of the substrate within the enzyme and generate new protein structures to support the catalytic glutamate (the general base) while partially encapsulating the substrate peptide. For both the PP and DP strategies, the oxyanion hole motif was treated as a spatially disconnected and positionally fixed structure that remained temporarily unscaffolded. This generation focused exclusively on the glutamate-proximal interface, enabling the model to prioritize high-quality scaffolding and packing on one side of the catalytic site and encapsulate the substrate peptide without geometric clashes from the opposing side. This setup allowed us to screen and select only those intermediate structures that provided precise support for the catalytic glutamate and optimal packing of the substrate from the glutamate-proximal interface.
The second generative stage built upon these intermediate structures, focusing on scaffolding the oxyanion hole motif while further encapsulating any remaining solvent-exposed regions of the substrate. ProteusFlow extended the intermediate designs to structurally integrate the previously "floating" oxyanion hole motif. This focused generation allowed for extensive sampling and selection of designs that achieved precise geometric accommodation of the oxyanion hole residues, alongside improved substrate packing from the other interface. Final protease candidates were selected based on their ability to provide dual-sided, clamp-like encapsulation of the substrate surface together with precise scaffolding of the catalytic residues.
The boundary between the two generation stages was established through an empirical, structure-based evaluation rather than utilizing rigid a priori definitions. During Stage 1, a variable-length sequence sampling strategy was employed. Preliminary models were computationally screened, and the precise boundary was determined by selecting intermediate structures whose variable flanking regions formed optimal geometric interactions with the substrate. Upon validation of the Stage 1 motif, the coordinates of its terminal residues were fixed. These endpoints subsequently served as the exact structural anchors for Stage 2, directing the continuous generation of the remaining protein scaffold.
Sequence-conditioned generation of substrate conformation
While enzyme reactivity requires precise positioning of the substrate’s scissile bond within the catalytic center, the overall substrate configuration must be broadly sampled during enzyme–substrate complex generation and conditioned on the target sequence. In the initial stage of generation, a minimal three-residue substrate motif was initialized based on AlphaFold3 predictions of the enzyme–substrate complex. The P1 residue, which coordinates the zinc ion via its backbone carbonyl oxygen, was rigidly fixed to ensure its correct spatial orientation relative to the catalytic motifs. To stabilize the scissile bond, the fixed motif was expanded from the P1 residue to include the two flanking residues connected to it; their Cα coordinates were fixed, while their distal backbone torsion angles are sampled along the generative process. While this constraint effectively rigidified the catalytic scissile bond, the motif’s termini torsion angles were left unconstrained to enhance substrate structural flexibility. The remainder of the substrate structure was generated from scratch, conditioned on the target sequence.
The substrate peptide sequence information was embedded using our pretrained protein language model, together with the enzyme catalytic motif residues. Sequence inputs for newly generated regions were represented as mask tokens. These semantically rich embeddings of substrate and catalytic residues were then provided to the ProteusFlow model to guide the generative process. By leveraging these embeddings, ProteusFlow can sample substrate conformations that are consistent with the target sequence identity, while simultaneously generating enzyme structures that precisely scaffold the active motifs and establish favorable interactions with the substrate peptide.
Sequence design and filtering for the first structure generation stage
Following the initial backbone generation, intermediate structures from both the PP and DP strategies were filtered based on the number of contact residues between the enzyme scaffold and the Aβ segment substrate to ensure adequate substrate packing. For the PP strategy, given the intrinsic structural preservation of the intact fold, this contact-based metric served as the primary selection criterion. In contrast, partial scaffolds generated by the DP strategy underwent additional filtering steps. These structures were subjected to sequence design using LigandMPNN
44, followed by structural prediction with AlphaFold2
42. Candidates were selected for the next design stage based on high prediction confidence (pLDDT) and strong structural self-consistency, as indicated by a low RMSD between the design model and the predicted structure.
Sequence design and filtering for the second structure generation stage
The final designs from the second backbone generation stage underwent three iterative cycles
28,56 of sequence optimization using LigandMPNN and structural refinement with Rosetta FastRelax
57,58. To strictly maintain the catalytic zinc coordination environment during sequence design, we implemented a 15 Å “histidine exclusion zone” around the zinc ion, preventing the introduction of non-catalytic histidines. Additionally, for solvent-exposed positions, we applied a bias toward hydrophilic residues, addressing the tendency of LigandMPNN to favor small hydrophobic residues (e.g., alanine) on suboptimal structures. During structural relaxation, we enforced strict distance and angular constraints to maintain the correct orientation of the catalytic residues and zinc, thereby preventing distortion of the catalytic center geometry.
After sequence design and refinement, the resulting complexes were filtered using a comprehensive set of metrics to identify the best candidates. First, Rosetta scoring metrics were used to assess enzyme–substrate interface quality and complementarity, including binding energy, interface shape complementarity, and contact molecular surface area. Next, we used AlphaFold2 and AlphaFold3 to predict monomeric and complex structures of the designed enzymes. Candidates were selected based on high structural self-consistency, defined by high pLDDT and pTM scores and low RMSD for monomer predictions (AlphaFold2), and high ipTM and low RMSD for complex predictions (AlphaFold3) relative to the design models. Finally, to ensure ideal catalytic geometry, we filtered designs based on precise distances between the scissile carbonyl, zinc ion, and oxyanion hole residues in the AlphaFold3-predicted models, as these parameters are critical for supporting the tetrahedral intermediate geometry required for catalytic activity
59. Designs passing all criteria were selected for experimental validation.
Structure guided optimization
A two-stage structural refinement strategy driven by flow-based models was implemented for the computational optimization of DP622-S2. In the first stage, spatially weighted structural perturbations based on binding potential were applied to the scaffold while strictly preserving the predefined catalytic geometry. Regions maintaining tight contact with the substrate were subjected to minimal perturbation, whereas regions lacking direct interactions underwent greater levels of structural resampling. Consistent with the previously described method, the resampled scaffolds underwent three iterative cycles of LigandMPNN sequence optimization and Rosetta FastRelax structural minimization. The resulting variants were evaluated using AlphaFold2 and AlphaFold3, allowing the selection of top candidates with high structural confidence and target selectivity for experimental validation.
In the second stage, the top performing variant from the initial step was selected for further refinement. With the core structural scaffold and the established binding interface strictly fixed, generative refinement was applied to adjacent regions. Specifically, the N-terminus of the Aβ42 substrate was extended to direct the flow model to construct supplementary beta sheet motifs capable of engaging this sequence. The generated structures were subsequently optimized through multiple iterative rounds of LigandMPNN sequence design and AlphaFold3 structural prediction, allowing the selection of top candidates for experimental validation.
Experimental characterization of the designed zinc metalloproteases
Protein expression and purification
For small-scale screening, genes encoding the designed proteins were synthesized by Universe Gene Technology (Tianjin) Co., Ltd. These genes were cloned into the pET28a vector, incorporating an N-terminal His-tag and an HRV 3C protease cleavage site. Plasmids were transformed into BL21 (DE3) E. coli competent cells, which were cultured in 50 mL of LB medium supplemented with 50 μg/mL kanamycin at 37 °C. Upon reaching an OD600 of 0.8–1.0, the cultures were equilibrated at 20 °C, and protein expression was induced by adding 0.6 mM IPTG. Cultures were incubated overnight at 20 °C. Cells were harvested by centrifugation and resuspended in lysis buffer (25 mM Tris, pH 7.4; 150 mM NaCl), then lysed by sonication. Lysates were clarified by centrifugation at 12,000 rpm, and the supernatant was purified using Ni2+-immobilized metal affinity chromatography (IMAC) with Ni-NTA Magrose Beads (TargetMol). The beads were washed twice with 10 mL of wash buffer (lysis buffer plus 30 mM imidazole) to remove contaminating proteins. For tag removal, beads were resuspended in 1 mL of lysis buffer containing His-tagged HRV 3C protease, and the mixture was incubated overnight to ensure complete cleavage of the His-tag. After settling the beads, the supernatant containing the cleaved target protein was carefully collected and further purified by size exclusion chromatography using a Superdex 75 Increase 10/300 GL column.
For large-scale production of designed proteases or fusion substrate proteins, plasmids were transformed into BL21 (DE3) E. coli competent cells and initially cultured overnight in 10 mL LB medium containing 50 μg/mL kanamycin at 37 °C. Cultures were subsequently scaled up to 0.3–1 L LB medium with the same kanamycin concentration and incubated at 37 °C until OD600 reached 0.6–1.0. Protein expression was induced with 0.3–0.6 mM IPTG and continued at 20 °C overnight. Cells were lysed by sonication, and lysates were clarified by centrifugation at 12,000 rpm. The supernatant was purified using Ni2+-IMAC with Ni-NTA Superflow resin (Qiagen), washed five times with 5 mL of wash buffer, and eluted with 6 mL of elution buffer (lysis buffer plus 300 mM imidazole). Eluates were analyzed on a 4–20% gradient SDS-PAGE gel to verify purity. N-terminal His-tags were removed by overnight incubation with His-tagged HRV 3C protease during dialysis against lysis buffer. A second IMAC purification was performed to remove uncleaved protein and residual HRV 3C protease. Final protein samples were further purified by SEC using either a Superdex 75 Increase 10/300 GL or Superdex 200 Increase 10/300 GL column (GE Healthcare) in lysis buffer.
Protease cleavage reaction assays
The proteolytic activity of the designed proteases was evaluated using MBP–Aβ_segment–SUMO fusion proteins as substrates. Reactions were carried out in a final volume of 20–50 μL in a buffer containing 25 mM Tris-HCl (pH 7.4), 150 mM NaCl, and 50 μM ZnSO4. Purified protease and fusion substrate were mixed to final concentrations of 5 μM and 10 μM, respectively. For time-course experiments, reactions were incubated at room temperature, and aliquots for SDS-PAGE analysis were collected at designated intervals: 0.25, 0.5, 1, 2, 4, 8, and 12 h. To verify catalytic activity, both the designed proteases and active-site mutants were incubated overnight. Reactions were quenched with 10 mM EDTA and 5× SDS loading buffer, and samples were analyzed on 4–20% gradient SDS-PAGE gels.
Mass spectrometry
Proteolytic reactions were initiated by incubating the purified designed protease with the fusion substrate in reaction buffer overnight at room temperature prior to mass spectrometry analysis. The resulting cleavage products were analyzed using a Waters Time-of-Flight (TOF) mass spectrometer (Waters Corp.) coupled with an integrated liquid chromatography system. Mass spectra were acquired over a broad m/z range, and the raw data were deconvoluted to determine the experimental monoisotopic masses of the cleavage fragments. Cleavage sites were precisely mapped by comparing the observed fragment masses to the theoretical molecular weights of the MBP–Aβ_segment–SUMO sequence segments.
FRET-based proteolytic kinetic assay
The catalytic kinetics of the designed proteases were characterized using FRET substrates, mTurquoise2–Aβ_segment–mVenus
60, in which the target Aβ segment serves as a cleavable linker. To convert raw fluorescence signals into substrate conversion rates, a standard curve was established by mixing fully hydrolyzed product (generated using DP622, which demonstrated near-complete digestion in SDS-PAGE assays) with intact substrate at various molar ratios. The emission ratio (R = I
527/I
474) was recorded for each sample to correlate R with the extent of substrate hydrolysis, ensuring normalization for instrument-specific variations.
For kinetic assays, purified proteases (2.5 μM or 5 μM) were incubated with the FRET substrate in reaction buffer containing 25 mM Tris-HCl (pH 7.4) and 150 mM NaCl. ZnSO4 was supplemented at a tenfold molar excess relative to protease concentration (25 μM or 50 μM) to ensure optimal catalytic activity. Substrate concentrations were prepared as a serial twofold dilution, ranging from 20 μM to 0.156 μM, in black, flat-bottom 96-well plates. Fluorescence was monitored at room temperature with excitation at 434 nm and dual-emission detection at 474 nm (donor) and 527 nm (acceptor). Initial reaction velocities (v0) were derived from the linear phase of product generation and fitted to the Michaelis–Menten equation using nonlinear regression to determine kcat, Km, and catalytic efficiency (kcat/Km). All assays were performed in three independent replicates.
Preparation of monomeric Aβ1-42 peptide
Lyophilized Aβ1-42 peptide was first dissolved in 1,1,1,3,3,3-hexafluoro-2-propanol (HFIP) at a concentration of 1 mM to disrupt any pre-existing aggregates. This solution was incubated at room temperature for 1 h to ensure complete monomerization, and then the HFIP was allowed to evaporate overnight in a fume hood. To remove any residual solvent, the resulting peptide film was further dried using a CT 02-50 vacuum centrifugal concentrator (Christ) for 1 h. The dried film was then dissolved in dimethyl sulfoxide (DMSO), yielding a monomeric stock solution for subsequent experiments.
Enzymatic cleavage of Aβ1-42 by designed metalloproteases
The proteolytic activity of the designed proteases was tested using monomeric Aβ1-42 as substrate. Reactions were performed by incubating 20 μM Aβ1-42 with the designed enzymes or their mutant variants, each at a final concentration of 10 μM. The reaction was incubated overnight in a buffer containing 25 mM Tris-HCl (pH 7.4), 150 mM NaCl, and ZnSO4 at an equimolar concentration to the enzyme. Cleavage products were analyzed by mass spectrometry.
Computational evaluation of substrate compatibility via ProteinMPNN
To computationally assess the substrate selectivity and compatibility of the designed proteases, substrate substitution analysis was performed on all selected designs from the four independent batches (DP S1, DP S2, DP S3, and PP S1). Structural models of the enzyme substrate complexes were generated by replacing the original target peptide with two other substrates while maintaining the fixed backbone coordinates of the designed protease. ProteinMPNN sequence scores were subsequently calculated for each redesigned complex configuration. These scores were utilized to evaluate the sequence fitness and structural compatibility of the designed enzymes across different substrate profiles, providing a computational metric to quantify and compare target selectivity.
Computational design of the rigid fusion protein for cryo-EM
To overcome the molecular weight limitations of small de novo proteases for high-resolution single-particle cryo-EM reconstruction, we employed a rigid-body fusion strategy utilizing a large scaffolding base protein. This base protein — a compact helical bundle consisting of 1,200 amino acids — was de novo designed using ProteusFlow. Initial rigid-body docking and spatial arrangement were performed manually in PyMOL, positioning the C-terminal helix of the designed protease adjacent to the N-terminus of the scaffolding base protein. Throughout this process, the orientation of the two domains was carefully constrained to prevent steric clashes and to ensure that the protease’s catalytic pocket remained fully exposed and accessible for substrate binding. Next, ProteusFlow was used to de novo generate a structured linker to bridge the two domains. Unlike conventional flexible linkers, which often introduce conformational heterogeneity and blur cryo-EM density, our generative model sampled rigid backbone configurations to create a seamless, rigid connection between the domains. Finally, to further stabilize the fusion and minimize inter-domain hinge motion, we computationally redesigned the amino acid sequence at the junction and surrounding inter-domain interface. This interface optimization maximized shape complementarity and favorable energetic interactions, yielding a highly stable, monolithic fusion architecture well suited for high-resolution cryo-EM imaging.
Cryo-EM data processing and model building
The cryo-EM samples were prepared by formulating a complex solution containing the enzyme-base protein fusion concentrated to 5 mg/mL, supplemented with a 1.5-fold molar excess of the Aβ42 peptide and an equimolar concentration of zinc ions. Subsequently, 2.5 μL of this mixture was applied to glow-discharged holey amorphous nickel titanium grids (ANTcryo™ Cu300-R1.2/1.3). The grids were blotted for 3.5 seconds and flash-frozen in liquid ethane cooled by liquid nitrogen using a Vitrobot (Mark IV, Thermo Fisher Scientific). The grids were then loaded onto either a 300 kV Titan Krios G4 equipped with a Falcon4i detector (Thermo Fisher Scientific). Images were automatically collected using EPU software (Thermo Fisher Scientific) at nominal magnifications of ×215,000, with a defocus series ranging from −1.0 μm to −2.0 μm. The total electron dose for each stack was approximately 50 e−/Å2. The pixel sizes of the resulting micrographs were 0.57 Å.
The flowchart for data processing is presented in Supplementary Fig. S18. All data processing steps were performed in cryoSPARC
61. Patch motion correction was applied to compensate for beam-induced motion, and contrast transfer function (CTF) parameters were estimated using patch CTF estimation. Micrographs were curated based on ice thickness, defocus range, and CTF resolution estimates to ensure data quality. Subsequently, particles were automatically picked using the template picker. The extracted particle images were subjected to 2D classification, using simulated data generated from the base protein as a reference. Selected particle sets were used for ab initio reconstruction, followed by heterogeneous refinement to remove poorly aligned particles. Particles from the best-resolved classes were subsequently subjected to non-uniform refinement. To optimize the alignment of the region of interest, the particle coordinates were re-centered onto the target protease and re-extracted. A mask focusing on the target protein region was generated, and 3D classification was performed to progressively enhance the clarity of the reconstructed density map. Finally, particles from the optimal classes underwent reference-based motion correction, 3D classification, particle subtraction, and local refinement. The final resolution was estimated using the gold-standard Fourier shell correlation (FSC) 0.143 criterion, with high-resolution noise substitution applied to prevent overfitting. Cryo-EM density visualization was performed in UCSF ChimeraX
62.
Design model was used as the initial structural templates. The design model was fitted into the cryo-EM map using UCSF ChimeraX. The models were then manually adjusted in Coot. Structure refinement was performed using phenix.real_space_refine in PHENIX, applying secondary structure and geometry restraints
63. Overfitting of the overall model was monitored by refining the model against one of the two independent maps from the gold-standard refinement approach and testing the refined model against the other map. Statistics for the map reconstruction and model refinement can be found in Supplementary Table S4.
DATA AVAILABILITY
The density maps and structure files have been deposited into the Electron Microscopy Data Bank and the RCSB Protein Data Bank with the following accession codes: EMD-69321 and 23WM (PP507 E110Q); EMD-69322 and 23WN (DP622 E96Q); EMD-69323 and 23WO (DP221 E85Q).
CODE AVAILABILITY
The Rosetta macromolecular modelling suite (
rosettacommons.org) is freely available to academic and commercial users. Design protocols and analysis scripts used in this paper are available in the Methods and will be made completely public at
github.com/LongxingLab/metalloprotease. The code for ProteusFlow is available at
doi.org/10.5281/zenodo.21272937.
The Author(s) 2026. Published by Higher Education Press. This is an Open Access article distributed under the terms of the CC BY license (https://creativecommons.org/licenses/by/4.0/).