top of page
Search

AlphaFold & Co. struggle predicting complexes for non-host-beneficial proteins?

  • Mar 11
  • 1 min read

Figure was adapted from ColabFold presentation


Why might AlphaFold & Co. struggle predicting complexes for non-host-beneficial proteins?


I work on transposons - mobile DNA elements that encode proteins capable of cutting and pasting their own sequence within a genome. Fascinating biology, but also quite “selfish” and potentially harmful to the host.


I have noticed that AlphaFold Multimer often struggles to predict protein–protein or protein-DNA complexes for these systems.


One of the major strengths of AlphaFold Multimer comes from Multiple Sequence Alignments (MSAs), ideally built from large numbers of related sequences. Proteins are linear amino acid chains that fold into 3D structures, where residues far apart in sequence can interact physically. During evolution, if one amino acid mutates, interacting residues often co-mutate to maintain function (coevolution). These correlated changes encode structural information that AlphaFold can learn from.


However, transposon-encoded proteins may experience very different evolutionary pressures. Instead of being optimized for host benefit, hosts may select mutations that weaken or deactivate these proteins. This could mean MSAs for these “selfish” proteins contain many non-functional or partially functional variants - reducing the strength of the co-evolutionary signal needed for accurate multimer prediction.


So maybe it’s not just about the number of sequences in an MSA, but also about what kind of evolutionary pressure shaped them.


Curious to hear thoughts from others working on:


• mobile genetic elements

• host–parasite molecular evolution

• structure prediction and MSA signal quality


Am I missing something in my reasoning?



 
 
bottom of page