César de la Fuente’s lab now uses Codex and ChatGPT to scan living and extinct genomes for antimicrobial candidates. The work targets molecules that could address drug-resistant infections.
The approach processes genome data at scale through large language models. Earlier searches depended on narrower sequence analysis or physical screening of known organisms. The workflow reaches both contemporary species and reconstructed ancient genomes.
The lab feeds genome sequences into Codex to generate candidate peptide structures. ChatGPT interprets results and proposes next experimental steps. The models surface short protein fragments that show signs of antimicrobial activity in silico. Researchers then validate the most promising sequences in the laboratory. The source material covers both living organisms and extinct species whose genomes have been reconstructed from preserved remains. No performance numbers or specific molecule names appear in the available report.
For teams already working on antimicrobial discovery, the shift shows how general-purpose coding and language models can accelerate candidate generation without requiring custom model training. Labs that lack large dedicated compute clusters can still run these queries through existing OpenAI interfaces. The method broadens the search space to include ancient sequences that would otherwise remain inaccessible.
At the same time, the output remains dependent on the quality of the underlying genome data and on downstream wet-lab confirmation. Any molecule that reaches clinical use will still face the standard regulatory and resistance hurdles that have limited earlier candidates. The immediate change is therefore in the speed and breadth of the initial discovery step rather than in guaranteed new drugs.
The single reported case also highlights a practical limit. General models trained on code and text can handle biological sequence tasks when the input is formatted as text, yet they do not replace the need for experimental validation or for high-quality genome assemblies. Teams considering the same route must still maintain wet-lab capacity and access to reliable sequence data from both modern and paleogenomic sources.
Another constraint follows from the nature of the models themselves. Codex and ChatGPT produce suggestions based on patterns seen during training; they do not perform de novo physical chemistry calculations. Any peptide flagged by the system must be synthesized and tested for activity, toxicity, and stability before it can advance. The OpenAI report gives no indication that the lab has shortened these later stages.
The same limitation applies to the ancient genomes. Reconstructed sequences carry uncertainty from DNA damage and incomplete coverage. A candidate drawn from such data carries an extra layer of verification before it can be considered viable. The report does not state how the lab filters these uncertainties before feeding sequences into the models.
For groups outside well-funded institutions, the approach lowers the barrier to initial candidate generation. A researcher with an OpenAI API key and a set of genome files can run the same queries without maintaining a dedicated high-performance computing cluster. The trade-off is reliance on a third-party service whose pricing, rate limits, and model updates remain outside the lab’s control.
The report leaves open the question of how many candidates the workflow has produced so far and how many have cleared initial laboratory tests. Without those figures, it is not yet possible to judge whether the method increases the hit rate over previous computational screens. The single source provides only the existence of the workflow, not its yield.
---
Sources:
{"word_count": 612, "sources_used": 1}
No comments yet