Aadi Trivedi
← Back to Research

Independent Research, 2026

Using Protein Language Models to Analyze Missense Mutations in the Human TP53 Gene

Aadi Trivedi

Overview

TP53 is one of the most important tumor suppressor genes in human biology, and many cancer-associated TP53 mutations are missense mutations that change a single amino acid in the p53 protein. This project asked whether ESM-C, a protein language model, can identify biologically important and clinically pathogenic TP53 missense mutations using only protein sequence context, without any structural or clinical information. All 7,467 possible single amino acid substitutions across the 393-amino-acid canonical human p53 sequence were generated and scored, then compared against known ClinVar variant classifications ranging from benign to pathogenic.

Abstract

Protein language models (PLMs) may help researchers study and interpret disease variants to guide biomedical work in areas including cancer biology and vaccine development. This study asked whether ESM-C, a protein language model by Biohub, can identify biologically important and clinically pathogenic missense mutations in the TP53 gene using only protein sequences. The canonical 393 amino acid human p53 sequence was used to generate all 7,467 possible single amino acid substitutions. Each mutation was scored based on how well the substituted amino acid fit within the normal p53 protein sequence, and scores were then compared with known data from ClinVar, ranging from benign to pathogenic. More negative ESM-C scores indicated lower predicted sequence compatibility and greater possible disruption, which could mean more pathogenic. ESM-C scores were unevenly distributed across p53, with some regions showing greater mutation sensitivity than others. ClinVar pathogenic and likely pathogenic variants had substantially more negative scores than benign and likely benign variants, suggesting that ESM-C captures scores that are relevant to p53 function. Overall, these results suggest that ESM-C can help highlight TP53 mutations that are more likely to disrupt p53 function and may require further biological study.

Workflow diagram for Using Protein Language Models to Analyze Missense Mutations in the Human TP53 Gene

Figures

Figure 1: Overview of the TP53 missense variant scoring workflow.
Figure 1: Overview of the TP53 missense variant scoring workflow.
Figure 2: Distribution of ESM-C scores for all 7,467 possible TP53 missense mutations. More negative scores indicate substitutions predicted to be more disruptive.
Figure 2: Distribution of ESM-C scores for all 7,467 possible TP53 missense mutations. More negative scores indicate substitutions predicted to be more disruptive.
Figure 3: Average ESM-C score by TP53 position. Lower average scores indicate positions where substitutions were predicted to be less tolerated.
Figure 3: Average ESM-C score by TP53 position. Lower average scores indicate positions where substitutions were predicted to be less tolerated.
Figure 4: Heatmap of ESM-C scores by TP53 position and substituted amino acid. Blue indicates more negative, disruptive scores; red indicates positive scores.
Figure 4: Heatmap of ESM-C scores by TP53 position and substituted amino acid. Blue indicates more negative, disruptive scores; red indicates positive scores.