
Predicting TCR-pMHC reactivity is useful for a variety of immunotherapy research tasks, including identifying relevant neoantigens and developing vaccines. In work recently published in Immunity, Messemaker and Kwee et al. developed a TCR rapid assembly platform (T-RAP) and used it to generate a dataset of functionally validated TCR-pMHC pairs, which enabled the assembly of diverse and standardized TCR libraries and allowed for structure-based prediction of TCR-pMHC reactivity.
To begin, Messemaker and Kwee et al. re-annotated TCRs from the public VDJdb database, which includes ~6300 TCRs with known specificities. They focused on entries that are reactive to a densely covered set of 10 epitopes (VDJdb-10) that encompassed approximately half of the full VDJdb database (3693 TCRs). Experimentally, they then created TCR rapid assembly platform (T-RAP) – a robotics-based high- throughput platform for the arrayed assembly of TCRs from only their sequence information, uncoupling analysis of TCR reactivity from clonotype abundance, T cell state, or T cell dysfunction. TCRs could be assembled quickly and at a low cost, enabling the creation of customized TCR libraries.
To confirm the efficacy of T-RAP, the researchers created a set of MHC-I and MHC-II TCRs, including TCRs from the VDJdb and TCR-s1–TCR-s4 libraries. Sanger sequencing of nearly 100 randomly selected individual TCR products validated that the correct sequence was the dominant product while Oxford Nanopore Technologies (ONT) sequencing validated the successful assembly of TCR libraries at scale.
Next, the researchers evaluated the reactivity of the T-RAP library by applying a pooled genetic screening platform and introducing the TCRs into TCR-null CD8+ Jurkat T cells. These cells were then exposed to target cells expressing relevant epitopes and corresponding HLA alleles, and activated (CD69+) T cells were isolated. When TCR abundance was determined by sequencing, the researchers found that only 56% of the TCRs on average showed their originally annotated reactivity.
To determine the accuracy of pooled TCR reactivity screens, the team selected a total of 77 TCRs for which the screen either did or did not validate the previously reported reactivity. When tested in individual arrayed co-cultures, reactivity was highly concordant with their pooled screening results. A low number of false negatives were observed, but no false positives. Some TCRs were found to be poly-reactive or reactive to an epitope that was distinct from what was reported in VDJdb.
Next, the team used their large-scale functional screening data to evaluate validation rates segregated by the original study in which the TCRs were identified, and found that there was wide variation, ranging from less than 10% to over 95% validation, with even VDJdb’s confidence scores showing only moderate precision and low sensitivity for TCRs with the claimed antigen reactivity.
Investigating whether the low TCR validation rates were due to differences between TCR-pMHC binding and actual pMHC-induced TCR signaling, the researchers performed pooled pMHC-multimer screens (rather than CD69 activation screening) for two HLA*02:01-restricted model epitopes. When VDJdb-10 Jurkat library was stained with MHC multimers for the corresponding epitopes, only a low fraction of TCRs showed detectable binding, showing concordance with the pooled functional screening data. When the researchers evaluated a set of 18 discordant TCRs, most were found to be false negatives in either the pooled functional screen or pooled pMHC-multimer screen, with no false positives. The remaining TCRs were rare LQ-MHC binders that failed to signal in both pooled screens and arrayed testing, likely representing a limited capacity to induce T cell activation. These results suggest that both non-specific pMHC-multimer binding and incorrect TCRα chain calling for dual TCRα T cells could be sources of error in the original VDJdb dataset.
Next, Messemaker and Kwee et al. used their at-scale TCR-pMHC reactivity model to evaluate the efficacy of two current prediction models – tcrdist3 and AlphaFold3. When they evaluated tcrdist3, which predicts epitope reactivity based on TCR sequence similarity to reference TCRs with known epitope reactivities, they found that tcrdist3 had variable performance for predicting reactivity to different model epitopes, ranging from high to modest. Similar variable prediction was observed for AlphaFold3, which predicts TCR reactivity based on minimum predicted aligned error (min-PAE) between residues at the TCR-peptide interface. This was observed both when using TCR signaling and pMHC binding as the readout.
One way that prediction tools are utilized is in identifying unknown pMHC epitopes that might be recognized by TCRs isolated from disease settings. To test whether AlphaFold3 models could accomplish this (prediction is not possible for this type of epitope with tcrdist3 as no reference TCRs are available), the researchers created AlphaFold3 models for each of their validating TCRs in complex with either their cognate pMHC or other pMHCs. For each complex, min-PAE scores were calculated and used to determine a relative TCR reactivity score (RTR). These scores consistently ranked TCRs paired to their cognate pMHC higher than those paired to non-cognate pMHCs, though performance levels varied from high (AUROC 0.83-0.90) to modest (AUROC 0.65-0.66). TCRs with a TCR-pMHC complex structure present in the Protein Data Bank (PDB) dataset that was used for AlphaFold3 training did not show better reactivity scores, suggesting that data leakage was not likely a driver of the modest performance. Further, no performance was observed for TCRs with reactivities that did not validate in functional screens, serving as a control that demonstrated how performance estimates were dependent on the quality of the data used.
In recent work, the researchers had identified TCRs from tumor-infiltrating CD8+ T cells from a patient with melanoma, including eight that were reactive to the HLA-B*44:01-restricted TNFAIP2P>A neoantigen and eleven that were reactive to the HLA-A*02:01-restricted CCSER2P>L neoantigen. Using AlphaFold3, the team found that these TCRs could effectively be ranked as reactive to these neoantigens compared to other non-reactive TCRs using min-PAE scores, without any task-specific training. The neoantigen-reactive TCRs could also be detected using precision-recall analysis or RTR scores.
These results show that T-RAP is an effective platform for creating large TCR libraries, and that pooled screening of these TCR libraries can be used to generate high-quality, validated datasets of TCR reactivity under standardized conditions. These datasets could then be used to evaluate the quality of existing reactivity prediction models. The researchers showed that both tcrdist3 and AlphaFold3 show strong capacities for distinguishing between validating and non-validating TCRs and could be used both to pair TCRs with cognate pMHC epitopes and to shortlist TCRs with experimentally confirmed reactivity towards patient-specific neoantigens, without any task-specific training. This technology for at-scale evaluation of TCR reactivity under standardized conditions could be used to validate existing prediction models and extend the ability to accurately predict TCR–pMHC pairs in a variety of immunotherapy research settings.
Write-up and image by Lauren Hitchings
Meet the researcher
This week, co-first authors Marius Messemaker and Bjørn Kwee answered our questions.

What was the most surprising finding of this study for you?
The observation that poor data aren’t a thing from the past: we – as a field – are mostly creating more data these days, really not so much creating better data, and this really is something that needs to be addressed.
What is the outlook?
Our work shows that prediction of recognition of previously unseen antigens by previously unseen TCRs from genetic sequence information alone is more feasible than previously estimated, and that working toward a broadly applicable prediction of T cell recognition is therefore a more realistic goal. In the longer term, it points to a future in which a routine tumor biopsy or a blood draw could be enough to predict what the TCRs in that sample recognize, without any wet-lab measurement.
If you could go back in time and give your early-career self one piece of advice for navigating a scientific career, what would it be?
With new screening methods and models appearing all the time, it can feel like you need to rush. Resist that pressure and take your time to aim for quality. For a new screening method, rigorously validate its precision by validating top-ranked pairs and testing on multiple samples. For a new model, build biological domain knowledge, pick the right task, take appropriate precautions to not overfit on your data, and make sure that improvements remain significant on truly held-out data. In the end, quality will stand the test of time and will, hopefully, make another study like this one unnecessary.
