Skip to content
CAAIL

CAAIL · Cellular Agriculture AI Library

Cell Ag × AI

The curated library at the intersection of cellular agriculture and AI, and a grounding database an AI assistant can query, so research and industry questions can be answered from a fixed, citable index rather than from model memory.

Indexing 346 papers, 142 software tools, 150 databases, 238 datasets.

Counted at build time

Point an AI agent at CAAIL

No download, no command line. In Claude Science:

  1. Settings → Skills → Add skill → Import from GitHub
  2. tucca-cellag/caail
  3. Preview → Import

Claude Science

Beta, macOS or Linux, Pro plan or above. Imports the query and contribute skills, untick either; pinned to that commit, and Check for updates moves them forward. Claude Science skills docs →

Dedicated ChatGPT and Gemini integrations are in progress. Until then the Any agent prompt works in both; it just does not persist.

Explore the library

Six collections, all cross-linked back to the canonical Markdown.

Start here

New to cell ag?

Watch the field foundations and pick up the reference reading.

New to AI?

Start with ML learning playlists built for bench scientists.

Looking for a method?

Filter the Papers matrix by AI method or research area.

Need training data?

Browse datasets by species, or benchmarks for model evaluation.

Recently added

2026-09-02 Paper The two matrix cells a reviewer confirmed and nothing landed
2026-08-25 Paper The Hybrid Mechanistic-ML and Comparative Studies rows
2026-08-20 Paper Metabolic Modeling and Food Safety Prediction columns
2026-08-20 Software Three allergenicity predictors
2026-08-12 Dataset The five bovine deposits the SuperSeries audit surfaced
View the full changelog →

Topics

Is your own area in here, and how thinly?

8 subject themes cut across papers, software, databases and dataset entries. The counts are uneven, and none of them is offered as completeness: a column with nothing in it is a result CAAIL returns rather than a gap it hides. Tagging is curator-assigned and has not been audited end to end, so if your own area looks wrong here, that is worth telling us.

16 finer tags sit beneath these 8 themes, each minted only where enough work clustered to earn one. Open the topic hub →

Why CAAIL

The same question, answered from memory and from an index

CAAIL is a grounding database: a static index an AI agent queries directly, so an answer can be assembled from the corpus instead of recalled from training. It holds 230 primary-research papers, each placed in at least one cell of a 27 × 8 matrix of AI method against research area, alongside 74 reviews and perspectives, 42 reference-work entries, 142 software tools, 150 databases and 238 dataset entries. Every figure in this section is read from that corpus when the page is built.

Answered from model memory

You

Which papers applied deep learning to cell-line engineering for cultivated meat, and which of them released code I can run?

AI

Convolutional and transformer models have both been used to guide cell-line engineering, for example Zhang et al. (2021) in Metab. Eng.[nothing here tells you whether this exists]

  • A citation is generated, not retrievedThe reference is produced token by token, so a well-formed DOI is not evidence that the work exists
  • Ask twice, get two answersDecoding samples from a distribution, so a pinned model on the same day need not agree with itself
  • There is no result for "found none"A gap in the training data and a gap in the field produce the same reply
  • Frequency stands in for relevanceWhat is heavily cited and plainly titled is what is recalled; a tool with an obscure name is not
  • Reuse terms do not travel with the answerA tool arrives named but unlicensed, so nothing separates one you can ship from one whose terms your product would inherit

Answered from CAAIL

You

Which papers applied deep learning to cell-line engineering for cultivated meat, and which of them released code I can run?

API

matrix.json → Deep Learning × Cellular Engineering

AI

6 papers: Li 2020 (#5), Magnusson 2024 (#122), Adduri 2025 (#57) and 3 more. 3 link runnable code. [ids resolve, corpus dated 2026-09-02]

  • All 216 method × area cells are enumerated138 hold no indexed paper, and each is returned as an empty result rather than skipped
  • An empty cell is returned with its own limit attachedA curated subset, not a census: CAAIL has not measured its own recall
  • Entries are placed by a curator, not ranked by a signalAn obscure tool and a well-known one are listed on the same basis, because inclusion is a recorded judgement
  • One subject axis across papers, software, databases and dataset entriesA design question returns them together, not one type at a time
  • A licence tier per tool, and a stated gap where there is noneOf 142 software tools, 76 are permissive and 22 copyleft, so a derivative inherits the terms; 8 are restricted and 36 record none
  • The same query returns the same rowsA fixed corpus dated 2026-09-02, so a result can be quoted and re-checked by someone else

The left exchange is constructed and its citation is invented: it illustrates the failure, it does not report one. The right exchange is generated when this page is built, from the same matrix.json an agent fetches.

Every reference id above resolves to an entry with its DOI. Browse the matrix →

What you can ask

Ask something specific

Four questions a working researcher would actually type, run against the endpoints an agent would query.

These are what the endpoints return, not finished answers; your agent reasons on top of them. Nothing here is a mock-up: every figure in the answers is computed from the corpus when this page is built.

Experiment design × existing deposits

I am planning an RNA-seq timecourse on primary bovine satellite cells through differentiation, budget for about 12 libraries. I would like this to be integrable with existing data rather than standing alone. Which public datasets could mine realistically be combined with, and what do I need to match in the design (timepoints, depth, annotation build) for that to be possible?

Answered from datasets.json

The grid you were going to spend 12 libraries on is already public. What is thin is the trigger you would be spending them under

  • Already public, at full gridGSE173198 is bulk RNA-seq at 0/24/48/72/96 h in biological quadruplicates, with differentiation induced by serum starvation. Repeating that grid under that trigger buys you replication, not reach
  • Where the public grid is thinUnder serum-free induction there is no comparable series. GSE173196 is bulk but samples only 0 h and 72 h, and GSE240556 covers all five timepoints in a single multiplexed single-nucleus library
  • What you actually have to matchThe induction trigger, not the timing. Match serum starvation and you can pool with the existing series directly. Choose serum-free and your 12 libraries become the first replicated timecourse under that trigger, which is the more valuable of the two things you could buy
  • Check before you rely on itThe bovine FAANG epigenome atlas is the obvious annotation reference and its new deposits are embargoed until journal acceptance

Both arms are deposited inside one parent accession that names neither, so a list of deposits will not tell you this and the design lines will. Read the thin side as thin in what is indexed here: CAAIL holds 47 bovine datasets as a curated subset and has not measured its own recall, so this is not evidence that nobody has run it.

Media optimization × metabolic modeling

We are trying to cut the cost of our serum-free medium and have reached the point where one-factor-at-a-time screening has stopped paying. I have a draft genome-scale model for our species but no experience closing the loop between a model and a screen. What work couples a metabolic model to an experimental design loop, and what did those groups actually have to measure to make the model useful?

Answered from matrix.jsonpapers.json

13 papers here are genome-scale or metabolic-network work. The matrix reaches 3 of them

  • The cell you would searchBayesian Optimization × Media Optimization holds 8 papers. 4 link released code, and 1 of the 8 links both code and a data deposit, which is what re-running a loop yourself requires
  • Where the rest of it sitsThe other 10 carry no matrix method at all, because they are filed as reference works and reviews rather than as method-to-area research. Searching the matrix will not surface them
  • The closest answer to your questionGomez Romero 2026 (#240), "iSsus3744: A genome-scale model-guided strategy for rational media design for cultivated pork", constrains its model with experimentally determined biomass composition and uptake and excretion fluxes from a muscle satellite-cell line, then validates it by amino-acid supplementation. That is the measurement burden you asked about. It is a preprint and has not been peer reviewed

What each group had to measure is in their methods sections rather than in any metadata field, which is why this ends in papers to read rather than in a filter. Read the scope precisely: this is where the coupling is indexed, not a claim about who has done it.

Allergenicity × novel-food dossier

We have engineered a bovine line expressing a couple of non-native proteins and we are scoping a novel-food submission. Before I commission wet-lab allergenicity work, what can I screen the sequences against computationally, and has anyone in cultivated meat already been through this whose submission I can read?

Answered from catalog.json

Only one of these implements the Codex criterion itself. The other 9 score allergen similarity by other means

  • The screen that is the criterionAllermatch implements the FAO/WHO Codex Alimentarius test directly: a sliding 80-mer window at or above 35% identity, plus an exact 6-mer match, run against curated allergen sets
  • The classifiers, which answer "allergen-like"AllerCatPro 2.0, AllerTOP v2, AllergenFP, AlgPred 2.0, ALLERDET, AllergenAI, AllerTrans, NetAllergen, ChAlPred. Alignment-free, homology-plus-structure and deep-learning methods, each trained on a different set, so agreement between them is not confirmation
  • What they are all standing onWHO/IUIS Allergen Nomenclature Database, AllergenOnline (FARRP), COMPARE Database, SDAP 2.0 (Structural Database of Allergenic Proteins), Allergome, AllFam (Database of Allergen Families). Which reference set you screen against moves the answer more than which classifier you run
  • Who has already been through itFDA Inventory of Completed Pre-market Consultations for Human Food Made with Cultured Animal Cells publishes, per product, the file number, the sponsor's own safety submission and the agency's response

The screens are cheap, and the choice that matters is the reference set rather than the classifier. What none of them settles is whether your specific protein needs serum-IgE work, which is the spend you were scoping. CAAIL links what each resource says about itself, so check scope and update dates at the source before relying on one.

Sensory prediction × small-n reality

We run a small cultured-pork line and have GC-MS volatile panels plus trained-panel scores for about 40 samples, nowhere near enough to train on. Before I commit to another year of panel sessions, which approaches have actually predicted sensory outcomes from instrumental measurements in meat, and what did those groups need in sample counts and reference standards to make it work?

Answered from matrix.jsonpapers.json

The methods are settled and small-n friendly. What is thin is anything you could build on

  • What the field actually uses32 papers across the ensemble, SVM, chemometrics and nearest-neighbour rows, against 17 on deep learning, CNN and GNN combined. The largest single cell in the matrix is Ensemble Learning × Sensory Prediction at 23 papers, so your problem is not the architecture
  • What those methods needed in samplesColantonio 2022 (#72) ran your question as a subsampling analysis: training sets from 50 to 170 with 39 held out, on 209 samples, accuracy still climbing across that range though some traits were already usable at 50. The smallest configuration it reports is about 90 samples
  • Why you are starting the baseline, not joining oneSensory papers release code at 11 of 52 (21%), against 81 of 178 (46%) across the rest of the primary literature. Every sensory reference is primary research, so that is like for like

What each group had to measure is in their methods sections, which is the reading this ends in. The useful read is not that 40 samples is fine, because the one paper that measured it says otherwise. It is that the methods you would use are the cheap ones, and the panel sessions are buying the thing the literature will not hand you.

Open the Papers Explorer → betaCorpus dated 2026-09-02

Connect your agent

CAAIL is static JSON on GitHub, so there is nothing to host and nothing to authenticate. The paths differ in how long the setup lasts, and in what each one asks of you.

No download, no command line. In Claude Science:

  1. Settings → Skills → Add skill → Import from GitHub
  2. tucca-cellag/caail
  3. Preview → Import

Claude Science

Beta, macOS or Linux, Pro plan or above. Imports the query and contribute skills, untick either; pinned to that commit, and Check for updates moves them forward. Claude Science skills docs →

Dedicated ChatGPT and Gemini integrations are in progress. Until then the Any agent prompt works in both; it just does not persist.

Linked external resources are independent of TUCCA and Tufts University and remain under their own licenses.

Last updated: