Class set · a printable, ages 15–18
FairQuestion
A research question on one sheet. The order is the point: you write down which things you will count before you look at any numbers, so the answer cannot steer the sample. Print it and fill it in by hand.
The template
- My question. One sentence, about something I can count.
- My sampling rule, written first. Which proteins (or residues, or tools) count, and which do not. No changes after I start measuring.
- What I measure, and with which tool. Name the number, its unit and where it comes from.
- My results. A table or a chart. Say n, the number of things measured.
- What I am not sure about. At least two things that could make my answer wrong.
- My defence. The strongest objection a classmate could raise, and my reply.
A worked example
My question. Are longer proteins more likely to have a lot of residues that AlphaFold is very unsure about?
My sampling rule, written first. Every protein in Fold Commons' saved copy of AlphaFold DB confidence scores (snapshot of 2026-10-04) that has a model. No protein added or removed after looking. That gives n = 65 (1 protein in the set has no AlphaFold DB model and is left out by the rule).
What I measure. The chain length in residues, and the share of residues in AlphaFold's very-low band (pLDDT below 50), as AlphaFold DB reports it. "A lot" means 10% of the chain or more.
My results. Split at the middle length:
| Shorter half (up to 393 residues) | Longer half (443 residues and up) | |
|---|---|---|
| Proteins | 32 | 32 |
| With 10% or more very low | 9 | 15 |
The one protein exactly in the middle (399 residues) sits in neither half.
Rank correlation (Spearman's ρ, which compares the order of the two lists, not their sizes) between length and very-low share: ρ = 0.36. A value of 0 would mean no link in order; 1 would mean the longer protein always has the larger share.
the Sketcher's drawing Prediction Computed by a model, with how sure it is. Not an experiment. The scatter plots AlphaFold DB's own confidence scores, not measurements.
What I am not sure about.
- The sample is not random. These proteins are in Fold Commons because they are well known, so they may be more studied, more often human, or more often folded than proteins in general.
- n = 65 is small. A handful of proteins moving across the line would change the table.
- The confidence is AlphaFold's, not a measurement. Very low often marks flexible or disordered parts, but not always.
- Length may not be the cause. Long proteins often have several domains joined by flexible linkers; the linkers, not the length, might be what matters.
My defence. Objection: "You picked famous proteins." Reply: yes, and the rule was written before looking, so I did not pick them for this answer; the result is about this set, and to say more I would repeat it on a random sample of AlphaFold DB.
Numbers: AlphaFold DB (EMBL-EBI and Google DeepMind), per-residue pLDDT, from Fold Commons' snapshot of 2026-10-04, CC BY 4.0. Bands as AlphaFold DB defines them: very low is pLDDT below 50 (AlphaFold DB FAQ). Spearman's rank correlation: Spearman 1904, American Journal of Psychology 15:72 (doi:10.2307/1412159).