[q-bio.BM] The last five-plus years have seen many protein engineering disciplines transformed by advances in machine learning (ML), but the same cannot be said for directed evolution.

Reflecting on a previously co-authored perspective, I discuss why I believe this to be the case, arguing that a disconnect between the goals of machine-learning-assisted directed evolution (MLDE) researchers–“identify an optimal protein”–and the goals of directed evolution more broadly–“identify a sufficient protein given time and resource constraints”–is a principal culprit.

As an example, I highlight how nearly all current MLDE methods neglect to account for the cost of DNA synthesis, resulting in strategies that have limited practical applicability regardless of the underlying models’ capabilities.

I close by discussing recent works that are exceptions to this overarching trend, and emphasize that the last five years of efforts in ML-assisted protein engineering and the prescribed reframe of MLDE objectives need not be mutually exclusive.

Bruce J. Wittmann

Subjects: Biomolecules (q-bio.BM); Machine Learning (cs.LG)
Cite as: arXiv:2609.03046 [q-bio.BM] (or arXiv:2609.03046v1 [q-bio.BM] for this version)
https://doi.org/10.48550/arXiv.2609.03046
Focus to learn more
Submission history
From: Bruce Wittmann
[v1] Wed, 2 Sep 2026 18:17:41 UTC (684 KB)
https://arxiv.org/abs/2609.03046

Astrobiology

Explorers Club Fellow, ex-NASA Space Station Payload manager/space biologist, Away Teams, Journalist, Lapsed climber, Synaesthete, Na’Vi-Jedi-Freman-Buddhist-mix, ASL, Devon Island and Everest Base Camp...

Leave a comment

Your email address will not be published. Required fields are marked *