Experiment design × existing deposits
I am planning an RNA-seq timecourse on primary bovine satellite cells through differentiation, budget for about 12 libraries. I would like this to be integrable with existing data rather than standing alone. Which public datasets could mine realistically be combined with, and what do I need to match in the design (timepoints, depth, annotation build) for that to be possible?
datasets.json The grid you were going to spend 12 libraries on is already public. What is thin is the trigger you would be spending them under
- Already public, at full gridGSE173198 is bulk RNA-seq at 0/24/48/72/96 h in biological quadruplicates, with differentiation induced by serum starvation. Repeating that grid under that trigger buys you replication, not reach
- Where the public grid is thinUnder serum-free induction there is no comparable series. GSE173196 is bulk but samples only 0 h and 72 h, and GSE240556 covers all five timepoints in a single multiplexed single-nucleus library
- What you actually have to matchThe induction trigger, not the timing. Match serum starvation and you can pool with the existing series directly. Choose serum-free and your 12 libraries become the first replicated timecourse under that trigger, which is the more valuable of the two things you could buy
- Check before you rely on itThe bovine FAANG epigenome atlas is the obvious annotation reference and its new deposits are embargoed until journal acceptance
Both arms are deposited inside one parent accession that names neither, so a list of deposits will not tell you this and the design lines will. Read the thin side as thin in what is indexed here: CAAIL holds 47 bovine datasets as a curated subset and has not measured its own recall, so this is not evidence that nobody has run it.