Designing non-natural RNA-guided nucleases
Thank you for visiting nature.com. You are using a browser version with limited support for CSS. To obtain the best experience, we recommend you use a more up to date browser (or turn off compatibility mode in Internet Explorer). In the meantime, to ensure continued support, we a...
Thank you for visiting nature.com. You are using a browser version with limited support for CSS. To obtain the best experience, we recommend you use a more up to date browser (or turn off compatibility mode in Internet Explorer). In the meantime, to ensure continued support, we are displaying the site without styles and JavaScript.
Protein design is particularly problematic for multidomain enzymes such as RNA-guided nucleases, whose functions depend on coordinated interactions with guide RNAs, target DNAs and multiple conformational states. Protein sequence-based language models have opened new possibilities for enzyme engineering, but the active nucleases they produce often remain closely related to the reference sequences they are trained on. The authors addressed this challenge by combining the inverse protein-folding model ESM-IF1 with evolutionary information derived from natural TnpB homologs. TnpB is a small transposon-associated protein that functions as a genome-editing nuclease and is considered an evolutionary ancestor of Cas12 enzymes. This enzyme served as an excellent template because it combines programmable DNA targeting with diverse natural functions in a compact architecture. Initial computational experiments showed that ESM-IF1 could redesign the minimal RNA-guided nuclease ISDra2 TnpB. It preserved the DED catalytic triad, the catalytic residues, of the RuvC domain in TnpB.
From an evolutionary perspective, phylogenetic analysis showed that the ESM-IF1-generated sequences shared only 50–60% amino acid identity with ISDra2 TnpB. However, the sequences contained nonsynonymous changes at residues in a region analogous to the protospacer-adjacent motif (PAM) recognized by CRISPR–Cas9. To overcome this limitation, the researchers incorporated a masking strategy based on positional conservation and evolutionary coupling analyses. The coupling signals were derived from a Potts model (GREMLIN) trained on either paired TnpB–RNA or TnpB–DNA sequences obtained from genomic databases. Residues exceeding the positional conservation and evolutionary coupling thresholds served as sequence masks for conditional ESM-IF1 generation. This strategy effectively balanced structural input with functional conservation, allowing the production of highly divergent sequences while preserving key biological contacts.