end-to-end-study
end-to-end-study
Chain: best data for novelty -> low-hanging fruit -> most-likely paper (venue + story) -> author instructions -> method under rigor checklist -> reader/reviewer-oriented writing -> reviewer cycle -> private repo + tagged release.
Seven phases. Each phase links to a dedicated reference file.
Phase 0 - Orientation (30 s)
Confirm with the user: domain (oncology? immunology? neurology?), topic vagueness (is a topic given, or should the skill scan datasets first?), claim type preference (method / translational finding / benchmark). Decline phases outside scope (wet-lab validation, actual journal submission).
Phase 1 - Best data for novelty (data-first scan)
Novelty usually lives in under-mined assets of recently released large cohorts, not in re-analyses of classic datasets. Scan for (1) recent large-cohort releases with a secondary modality that has not been systematically exploited, (2) paired-modality data where one modality is under-used, (3) public drug-response or perturbation screens paired with rich clinical metadata.
See references/data-first-novelty.md for the scan heuristics, a catalog of currently-underutilised public datasets, and worked example (BeatAML ex-vivo drug-sensitivity table in Tyner 2018 supplementary).
Exit criterion: a primary dataset is named, its under-mined asset is identified, and a short list of 2-3 candidate claims is drafted.