Fold Commons

Class set · a printable, ages 15–18

FairQuestion

A research question on one sheet. The order is the point: you write down which things you will count before you look at any numbers, so the answer cannot steer the sample. Print it and fill it in by hand.

The template

  1. My question. One sentence, about something I can count.
  2. My sampling rule, written first. Which proteins (or residues, or tools) count, and which do not. No changes after I start measuring.
  3. What I measure, and with which tool. Name the number, its unit and where it comes from.
  4. My results. A table or a chart. Say n, the number of things measured.
  5. What I am not sure about. At least two things that could make my answer wrong.
  6. My defence. The strongest objection a classmate could raise, and my reply.

A worked example

My question. Are longer proteins more likely to have a lot of residues that AlphaFold is very unsure about?

My sampling rule, written first. Every protein in Fold Commons' saved copy of AlphaFold DB confidence scores (snapshot of 2026-10-04) that has a model. No protein added or removed after looking. That gives n = 65 (1 protein in the set has no AlphaFold DB model and is left out by the rule).

What I measure. The chain length in residues, and the share of residues in AlphaFold's very-low band (pLDDT below 50), as AlphaFold DB reports it. "A lot" means 10% of the chain or more.

My results. Split at the middle length:

Shorter half (up to 393 residues)Longer half (443 residues and up)
Proteins3232
With 10% or more very low915

The one protein exactly in the middle (399 residues) sits in neither half.

Rank correlation (Spearman's ρ, which compares the order of the two lists, not their sizes) between length and very-low share: ρ = 0.36. A value of 0 would mean no link in order; 1 would mean the longer protein always has the larger share.

10020050010002000 0%20%40%60%80% P04002: 82 residues, 15% very lowP99999: 105 residues, 0% very lowP01308: 110 residues, 51% very lowQ13541: 118 residues, 18% very lowP68431: 136 residues, 1% very lowP37840: 140 residues, 13% very lowP69905: 142 residues, 0% very lowP00698: 147 residues, 1% very lowP68871: 147 residues, 0% very lowP61626: 148 residues, 1% very lowP0DP23: 149 residues, 1% very lowP00441: 154 residues, 0% very lowP02144: 154 residues, 0% very lowP02489: 173 residues, 15% very lowP02754: 178 residues, 5% very lowP02794: 183 residues, 3% very lowP01116: 189 residues, 2% very lowP02662: 214 residues, 71% very lowP01241: 217 residues, 11% very lowP0CG47: 229 residues, 1% very lowQ07817: 233 residues, 31% very lowP42212: 238 residues, 0% very lowP07477: 247 residues, 5% very lowP62753: 249 residues, 0% very lowP00918: 260 residues, 0% very lowP29972: 269 residues, 7% very lowP08100: 348 residues, 5% very lowP04075: 364 residues, 0% very lowP68133: 377 residues, 1% very lowP01012: 386 residues, 2% very lowP0DJD7: 388 residues, 2% very lowP04637: 393 residues, 30% very lowP01857: 399 residues, 4% very lowP14416: 443 residues, 26% very lowQ71U36: 451 residues, 2% very lowP01106: 454 residues, 41% very lowP00875: 475 residues, 2% very lowP14672: 509 residues, 7% very lowP0DUB6: 511 residues, 3% very lowP04040: 527 residues, 4% very lowP06576: 529 residues, 11% very lowP14679: 529 residues, 8% very lowP12931: 536 residues, 16% very lowP08659: 550 residues, 1% very lowP03452: 565 residues, 11% very lowP02768: 609 residues, 4% very lowP22303: 614 residues, 6% very lowP04264: 644 residues, 49% very lowP02788: 710 residues, 3% very lowP10636: 758 residues, 66% very lowQ9BYF1: 805 residues, 5% very lowP19367: 917 residues, 1% very lowP05023: 1023 residues, 3% very lowP00519: 1130 residues, 49% very lowP00533: 1210 residues, 23% very lowP0DTC2: 1273 residues, 24% very lowQ99ZW2: 1368 residues, 1% very lowP09884: 1462 residues, 24% very lowP02452: 1464 residues, 67% very lowP13569: 1480 residues, 17% very lowP38398: 1863 residues, 80% very lowP09848: 1927 residues, 6% very lowP12883: 1935 residues, 5% very lowP12882: 1939 residues, 4% very lowP24928: 1970 residues, 25% very low Chain length, residues (log scale) Share very low
Each dot is one protein. The dashed line is the 10% mark.

the Sketcher's drawing Prediction Computed by a model, with how sure it is. Not an experiment. The scatter plots AlphaFold DB's own confidence scores, not measurements.

What I am not sure about.

My defence. Objection: "You picked famous proteins." Reply: yes, and the rule was written before looking, so I did not pick them for this answer; the result is about this set, and to say more I would repeat it on a random sample of AlphaFold DB.

Numbers: AlphaFold DB (EMBL-EBI and Google DeepMind), per-residue pLDDT, from Fold Commons' snapshot of 2026-10-04, CC BY 4.0. Bands as AlphaFold DB defines them: very low is pLDDT below 50 (AlphaFold DB FAQ). Spearman's rank correlation: Spearman 1904, American Journal of Psychology 15:72 (doi:10.2307/1412159).