Module 1.2.1: Design
Design is the process of turning a biological question into a study that can provide meaningful evidence.
Every omics study begins with a scientific question: a clearly defined knowledge gap that the experiment is designed to address. The question must be specific enough to determine which molecular layer is relevant, what comparison to make, and what a meaningful result looks like. From the question, four interconnected elements follow: a testable hypothesis or clearly defined objective, study variables, a cohort design, and platform selection. These elements influence one another and should be consider before samples are collected or selected and data is generated.
Omics studies also carry costs across three dimensions that make upfront design particularly important:
| Cost dimension | Examples |
|---|---|
| Financial | Sequencing runs, reagents, platform fees |
| Time | Sample processing, analysis pipelines, validation |
| Irreplaceability | Clinical biopsies, rare cohorts, and longitudinal samples cannot simply be recollected |
Foundations of study design
Omics studies rest on four interconnected elements mentioned above.
The question identifies the knowledge gap; the hypothesis specifies what is expected and at which molecular level; the variables define what is being compared and what needs to be considered; the cohort design determines which biological units and groups will be included; and the platform determines what the experiment can actually measure. Each element constrains the next.
Scientific question
The scientific question defines the scope and purpose of the study. A well formed question makes clear:
- What biological material or population is being studied?
- What condition, exposure or comparison is being investigated?
- Which molecular layer could provide relevant answer?
A useful test
Broad question such as "what differs between cases and controls?" can be useful starting point, but they do not yet define what will be compared, which molecular layer is relevant and make it difficult to evaluate whether your results are meaningful.
Before designing the study, the question should be specific enough to justify which molecular layer or combination of layers is most likely to provide relevant evidence.
Hypothesis
A testable hypothesis specifies what is expected to differ, at which molecular level, and in which biological context. A hypothesis framed at the wrong molecular layer will produce a study that cannot address the underlying question regardless of execution quality.
For example, genome sequencing alone would not directly measure changes in chromatin accessibility or post-translational modification. Although genetic variation can influence these processes, additional epigenomic or proteomic measurements may be needed to address the biological question.
Variables
Three types of variable should be considered when planing a study:
| Type | Other common names | Definition | Examples |
|---|---|---|---|
| Condition/exposure | Explanatory variable; independent variable; experimental factor (when assigned) | The biological variable whose relationship with the omics features is being investigated | Disease status, treatment, developmental stage, environmental exposure |
| Omics features | Response variables; dependent variables; outcomes | The molecular measurement generated by the platform | Gene expression, protein abundance, methylation state, metabolite concentration |
| Potential confounder | Confounding variable | Variables associated with condition/exposure that may also affects the omics features | Age, sex, batch, tissue composition, collection site |
Identifying and planning for potential confounders is part of study design. Important variables that are neither controlled nor recorded at the time of sample collection cannot be evaluated or accounted for during analysis.
Platform selection
Platform selection follows from the hypothesis. The question help identify which molecular layer or combination of layers is relevant; while the molecular layer, required resolution, sensitivity and scope guide the choice of platform. The platform also determines throughput and cost per sample. Under a fixed budget, a higher cost per sample may reduce cohort size and therefore statistical power. Platform selection is covered in detail in Key Consideration 2 below.
Consideration 1: Cohort design and confounding
Design principle
A potential confounder that is neither controlled nor recorded cannot be evaluated or accounted for during analysis. Record important variables that could be associated with the condition/exposure and affect the omics measurements.
Key terms
| Term | Definition |
|---|---|
| Matching | Selecting cases and controls so that relevant characteristics, such as age or sex, are similarly distributed between the groups |
| Stratification | Dividing the study population into subgroups defined by a variable, then making comparisons within each subgroup |
| Randomisation | Randomly distributing samples from each condition or exposure group across processing batches and run order to reduce systematic imbalance. |
| Metadata | Structured information describing the samples, conditions, collection and processing e.g. demographics, collection site, processing batch, storage conditions, and any other variable that might influence the measurement |
Because omics measurements are sensitive to both biological and technical variation, cohort composition and sample processing must be planned together.
Approaches for managing confounders: matching, stratification, randomisation, and balanced sampling, are covered in Module 2: Confounding. Not all confounders can be controlled in advance, particularly in retrospective studies where samples were collected before the study was designed. The minimum requirement is that important sources of variation are recorded in study metadata whenever possible, so they can be evaluated during analysis.
Example: A study recruiting cases from a specialist hospital and controls from a community health screen may differ systematically in age, medication use, comorbidity burden, and health seeking behaviour. These differences may affect the omics data, and be mistaken for differences associated with condition/exposure status (e.g. case/control).
Case study: Reference datasets carry their own sampling biases
The GTEx project is a widely used reference atlas of gene expression across human tissues. A 2020 analysis of GTEx data from 44 tissues found that 37% of genes showed sex biased expression in at least one tissue, although most effects were small and tissue-specific. Because the GTEx donor cohort contains more males than females and has particular age and ancestry distributions, studies using it as an external comparison should consider whether its population composition is appropriate for their study.
This issue is not specific to GTEx or gene expression data. Any population derived reference, such as a methylation atlas, plasma proteomic reference range, metabolite reference interval or population allele-frequency database reflects the cohort from which it was built. A 2011 review of studies in non-human mammals across ten biological fields found that male only studies outnumbered female only studies by 5.5:1 in neuroscience and approximately 5:1 in pharmacology, with the sex of the animals frequently unreported. Reflecting broader concerns about sex bias in research, an NIH policy effective from 2016 required sex to be considered in the design, analysis and reporting of relevant NIH-funded studies involving vertebrate animals and humans.
Beery & Zucker. Neuroscience & Biobehavioral Reviews 35, 565–572 (2011). doi:10.1016/j.neubiorev.2010.07.002
Oliva et al. Science 369, eaba3066 (2020). doi:10.1126/science.aba3066
National Institutes of Health. Consideration of Sex as a Biological Variable in NIH-funded Research (2015; effective 2016). NOT-OD-15-102
Consideration 2: Platform selection
Design principle
Platform selection is a biological design decision driven by the scientific question and the information required. Once data have been generated, information the selected platform did not capture cannot be recovered through downstream analysis.
The platform determines what biological information the experiment can capture. A platform that cannot measure the signal of interest at the required resolution or sensitivity will produce data that cannot answer the question, regardless of downstream analysis.
| Mismatch type | Description | Example |
|---|---|---|
| Scope | The platform measures a molecular layer that does not directly capture the biological process of interest | Genome sequencing identifies DNA variants but does not directly measure chromatin accessibility or current regulatory activity |
| Resolution | The platform measures at a level that obscures the relevant biology | Bulk approaches average across heterogeneous cell populations; cell-type specific responses cannot be recovered |
| Sensitivity | The platform may not reliably detect molecules at the abundance range relevant to the question | In proteomics, acquisition mode determines which proteins are measured at all; key targets may be absent rather than under quantified |
| Technical scope | The platform has limited ability to capture or resolve the molecular feature of interest | Short-read sequencing can resolve only few structural variants or full-length isoforms regardless of sequencing depth |
| Novelty over fit | A more sophisticated platform than the question requires is used, then analysed as though a simpler platform had been used | Single-cell omics applied to a bulk-level question, with no cell-type level analysis performed |
Consideration 3: Statistical power
Design principle
Statistical significance in an underpowered study does not indicate a robust finding. Planning sample size and assessing statistical power are therefore important parts of study design.
Key terms
| Term | Definition |
|---|---|
| Independent biological unit | The entity that provides independent biological information for the comparison, for example, a participant, animal or independently treated culture |
| Effect size | The magnitude of a difference or relationship being investigated e.g. fold change |
| p-value | The probability of observing a result at least as extreme as the one obtained, assuming no true effect exists. A small p-value indicates the result is unlikely under the null hypothesis; it does not indicate the effect is large or biologically meaningful |
| Statistical power | The probability that a study will detect a true effect of a given size, when the effect is present. Power depends on independent sample size, effect size, variability and the significance threshold |
| Multiple testing | When many hypotheses are tested simultaneously, the expected number of false positives increases proportionally. At p < 0.05 with 20,000 features tested, approximately 1,000 false positives are expected by chance alone |
| False discovery rate (FDR) | The expected proportion of statistically significant results that are false positives. Commonly controlled using method like Benjamini-Hochberg in omics |
| Pseudoreplication | Treating non-independent observations as independent. In single-cell and spatial omics, measurements from the same biological donor are not independent; the true sample size is the number of donors, not cells or spots |
In many omics studies, sample size is determined by budget or sample availability rather than by statistical need. This is particularly costly, where thousands of molecular features are tested simultaneously and multiple testing correction reduces the effective power per feature dramatically the sample size required to detect true signal is far higher than most researchers expect.
Genomics
In genome-wide association studies, variants with the strongest associations are selected for further investigation. In a study with limited power, their initial effect sizes are often overestimated simply because only the largest observed effects pass the significance threshold. This is called the winner’s curse. The association may also not be detected again if the replication cohort is too small or differs in its population or phenotype definition.
Zou et al. G3 (2022); Huffman. Nature Communications (2018)
Transcriptomics
A yeast RNA-seq study showed that 3 biological replicates per condition detected only 20–40% of the differentially expressed genes found with 42 replicates. Genes with > 4 fold changes were detected much more reliably. This shows that small studies are particularly likely to miss smaller effects.
In single-cell studies, cells from the same donor are not independent biological replicates. Counting them as independent makes the results appear more certain than they are, this is called pseudoreplication. When comparing donor groups, the independent sample size is the number of donors, not the number of cells.
Schurch et al. RNA (2016); Murphy et al. eLife (2023)
Proteomics
Missing values are common in proteomics and can reduce statistical power and introduce bias. Consequently, the usable information for a particular protein may come from fewer samples than were run, even though the number of independent biological units has not changed.
Metabolomics
A 2024 meta-analysis of 244 clinical cancer metabolomics studies of human serum identified 2,206 metabolites reported as statistically significant. Of these, 72% appeared in only one study, and metabolites reported by multiple studies frequently showed inconsistent directions of change. These findings demonstrate substantial instability across the published studies, although this cannot be attributed to sample size alone. They also illustrate why findings from small, single-cohort discovery studies require independent validation.
Cochran et al. Trends in Analytical Chemistry (2024)
In all cases, the result is the same: findings that appear statistically significant but do not replicate. Sample size requirements vary substantially by study type, discovery versus validation, rare versus common variants, large versus small effect sizes. Power calculations should be performed before data collection begins. check out module 2.2.1 for more details
Module 1.2.1 takeaways
- The scientific question determines which molecular layer to measure and what contrast to draw, platform and design follow from this, not the other way around.
- Confounders should be identified and controlled where possible and recorded in the metadata, so they can be evaluated during analysis.
- Platform choice involves trade-offs in scope, resolution, sensitivity, and cost; no single platform is optimal for all questions.
- Sample size requirements differ by platform and effect size; rules of thumb are unreliable substitutes for a power calculation grounded in pilot data or published estimates.
- Design decisions are interconnected, and limitations introduced early cannot always be corrected through downstream analysis.