Food safety is the assessment layer a cultured product must clear before it can enter the food chain, and the most computationally tractable part of that assessment is allergenicity. The recombinant growth factors, media proteins, and scaffold or matrix proteins introduced during cultivation are novel proteins, and FAO/WHO Codex Alimentarius guidance requires that any novel food protein be screened for allergenic potential by sequence homology against known allergens. The tools that run that screen live in Software.md / Food Safety & Allergenicity, and the reference allergen databases they query are in Databases.md / Food Safety & Allergen Databases. This page collects the labeled sequence data those predictors are trained and benchmarked on.
Allergen sequence & epitope datasets
Sequence-based allergenicity classifiers are trained on curated sets of allergen and non-allergen protein sequences, often annotated with the specific IgE-binding epitopes that drive the allergic response. The dataset below is the labeled corpus most widely used to train and benchmark that class of model.
The labeled corpus behind the AlgPred 2.0 allergen predictor: 10,075 allergen and 10,075 non-allergen protein sequences, together with 10,451 experimentally validated IgE epitopes and a 297-sequence independent validation set (plus a stricter 56-sequence non-redundant subset). It is a standard benchmark for training and evaluating sequence-based allergenicity classifiers, and directly relevant to screening the novel proteins introduced by cultivated meat and precision fermentation. Companion to Papers.md ref #290 (Sharma et al. 2021, Briefings in Bioinformatics); the predictor itself is catalogued in Software.md / Food Safety & Allergenicity.
cited by361
Further reading