Six true stories of people who waited a long time for a shape. Some waited thirty-four years for one
protein; the last one changed how long anyone waits.
Every paragraph is marked. Documented: in a published source, numbered in the list at
the end. Reconstructed: how such work went, in our words, where there is no record of the day.
Disputed: sources or people disagree, and we say how.
Dorothy Hodgkin and insulin 1934–1969
Documented In 1934, after Dorothy Crowfoot (later Dorothy Hodgkin) came back to Oxford, the chemist Robert Robinson passed her a small sample of insulin crystals that the Boots company had given him. Insulin crystals had only recently become easy to grow, once it was found that zinc helps them form. She grew larger ones with zinc present. [1][2]
Documented In April 1935 she published the first X-ray photographs of single insulin crystals, in a short letter to Nature. Earlier X-ray work on insulin powder had shown no clear crystal pattern. [2]
Documented Insulin is small for a protein: two chains, of 21 and 30 amino acids, held together by disulfide bonds. Frederick Sanger finished reading its sequence of amino acids in 1955, the first protein whose full sequence was known, and received the 1958 Nobel Prize in Chemistry for it. Knowing the letters did not give the shape. [3]
Documented An X-ray photograph of a crystal is a pattern of spots, not a picture of the molecule. To turn spots into a shape you also need the phase of each scattered wave, and the photograph does not record it. For a molecule as large as a protein, nobody could get those phases in 1935. A general way through, attaching heavy atoms to the protein, came from Max Perutz's group in the 1950s. [4]
Documented While insulin waited, she solved other structures: penicillin during the Second World War, and vitamin B12, published in the 1950s. In 1964 she received the Nobel Prize in Chemistry for determining the structures of important biochemical substances by X-ray techniques. [5]
Reconstructed We do not have her day-by-day record of the insulin years, so this is how such work went rather than what she did on a given day. Grow crystals. Soak some in a heavy-metal salt and hope the metal sits in one fixed place without spoiling the crystal. Photograph them, measure the darkness of thousands of spots, and calculate, at first by hand and later on the early computers that crystallographers queued to use.
Documented In 1969 her group, ten authors in all, published the structure of insulin in its zinc-containing crystals in Nature. It was thirty-four years after the first photograph. [6]
Disputed When was the first photograph taken? Most accounts, and her own later talk, put it in 1935, the year the letter appeared. Some biographies say 1934, the year the crystals arrived. The paper itself only fixes the date it was published. [1][2]
Documented In 1958 John Kendrew and his colleagues in Cambridge published a three-dimensional model of myoglobin, the oxygen-storing protein of muscle, taken from sperm whale. It was the first protein seen in three dimensions by X-ray analysis. [7]
Documented Myoglobin was chosen because it is small, a single chain of about 150 amino acids holding one haem group, and because myoglobin from sperm whale, among the many animals surveyed, gave crystals good enough for the work. [8]
Documented The phase problem was crossed with heavy atoms. In 1954 Perutz's group showed that a heavy atom bound to a protein in its crystal, without changing the crystal, shifts the spot intensities enough to work out phases. Kendrew's team made several heavy-atom versions of myoglobin crystals in the same way. [9][7]
Documented The 1958 map was coarse, at a resolution of 6 Å, too coarse to see single atoms. It showed dense rods, which were stretches of helix, packed into a compact body, and one very dense spot for the iron of the haem group. The paper remarked on how complicated and irregular the arrangement was. [7]
Documented Two years later, a map at 2 Å showed the chain residue by residue. The rods were α-helices of the kind Linus Pauling had proposed in 1951. [10][11]
Documented Perutz's own target, haemoglobin, is four chains and about four times larger. He had started on it in 1937; his group's model at 5.5 Å came out in 1960, beside Kendrew's 2 Å myoglobin, and showed that each haemoglobin chain is folded much like myoglobin. [12]
Reconstructed We have no record of what was said in the room when the rods first appeared, so this is how such a model came about rather than a record of that day. The map was drawn as contour lines on a stack of clear sheets, one slice of the crystal each. People read the stack slice by slice and built the shape by hand on a frame of rods. Each new heavy-atom crystal meant weeks of photographs and measurements before anyone could say whether it helped.
Documented In 1962 Kendrew and Perutz shared the Nobel Prize in Chemistry for their studies of the structures of globular proteins. [13]
Documented Ribonuclease is an enzyme from cow pancreas that cuts RNA. Its chain of 124 amino acids is held in shape partly by four disulfide bonds, each linking two of its eight cysteines. [14]
Documented The experiments ran through the late 1950s and early 1960s in Anfinsen's laboratory at the US National Institutes of Health in Bethesda, Maryland. [15][14]
Documented Christian Anfinsen and his colleagues unfolded it with concentrated urea and broke the four disulfide bonds with a reducing agent. The enzyme stopped working. Then they removed the urea and the reducing agent and let air re-form the bonds. The activity came back, close to fully. [16][14]
Documented Eight cysteines can pair up in 105 different ways. If the bonds re-formed at random, only about one chain in a hundred would come back right. Far more did. [14]
Documented In a second test, the bonds were re-formed while the urea was still present, and the chains came back scrambled and inactive. Adding a trace of reducing agent let the wrong bonds break and re-form, and the chains drifted back to the active shape. The working shape is the most stable one under normal conditions. [14]
Documented Anfinsen drew the conclusion that the information for a protein's shape is in its sequence of amino acids. He received half of the 1972 Nobel Prize in Chemistry for this work; Stanford Moore and William Stein shared the other half. [14][15]
Reconstructed We do not have a record of the moment the activity came back, so this is how such a test went. A sample of the refolding enzyme is added to RNA, and the rate at which the RNA is cut shows how much of the enzyme is working again. The measurement is repeated as hours pass. Before that, the urea is taken out slowly, so the chain has time to settle, and the sample is left open to the air so the bonds can close.
Disputed How far does the rule reach? Inside cells, many proteins fold with help from chaperone proteins, which stop chains clumping but do not carry the shape. Some proteins have two stable folds. So the idea that one sequence makes one shape holds for many proteins, and is argued over for others. [17][18]
Documented In 1994 John Moult and colleagues ran the first Critical Assessment of protein Structure Prediction, CASP. Experimental groups who were about to solve a structure sent in its sequence first. Predictors sent in their models before the structure was made public. Only then were the models compared with the experiment. [19]
Documented The point was blindness. A method checked on structures its authors could already see is easy to over-rate without anyone meaning to. A blind test removes that. [19]
Documented The first round already sorted predictions into kinds that CASP still uses in some form: models built from a related protein whose structure was known, models made by threading a sequence onto known folds, and models made from scratch with no template at all. The from-scratch ones were the hardest, and the furthest from experiment. [19]
Documented The results were discussed at a meeting at Asilomar, California, in December 1994, and written up in a special issue of the journal Proteins in 1995. [19]
Documented How close is close? CASP now scores a model mostly with a measure called GDT_TS: roughly, the share of the model's residues that can be laid within a few ångströms of their places in the experiment. 100 is a perfect match. [20]
Documented CASP has run every two years since, with the same rule: predictions first, structures after. [21]
Documented The method that finally won was built on the work of the people in the earlier stories. AlphaFold2 was trained on structures in the Protein Data Bank, the open archive that structural biologists have filled since 1971, one experiment at a time. [22][23]
Reconstructed We have no record of how any one predictor spent those months, so this is how a round goes. A sequence is posted with a deadline. A group runs its methods, picks the models it trusts most and submits them. Weeks or months later the experimental structure is released, and the scores arrive for everyone at once. Most models in a round are some way off; a round is judged by the whole field, not by one good model.
Documented Twenty-six years passed between the first round and the round in which one method's models matched experiment for most targets. [19][23]
Documented The protease of Mason-Pfizer monkey virus, a retrovirus, is an enzyme the virus needs to mature. Crystallographers had crystals and X-ray data, but every attempt to solve the structure by molecular replacement, which starts from a model of a similar protein, failed. The problem had been open for more than a decade. [24][25]
Documented Foldit came from groups at the University of Washington. A 2010 paper in Nature showed that top players beat the Rosetta computer program on some puzzles, especially ones that needed large rearrangements of the chain. [26]
Documented Most retroviral proteases work as a pair of identical chains. This one crystallised as a single chain, and that is part of why models built from its relatives did not fit the data. [24][27]
Documented The team put the protein to the players of Foldit, an online game in which people fold protein chains by hand, scored by an energy function. In a three-week competition, players produced models good enough for molecular replacement to work. [24][25]
Documented The authors wrote that the refined structure gives new clues for designing drugs against retroviruses, the family that includes HIV. [24]
Documented With those models the structure was solved and refined, and published in 2011. Two groups of players are listed among the authors of the paper. [24]
Documented A higher-resolution structure, at 1.6 Å, followed the same year. [27]
Reconstructed We have no record of any one player's evening at the screen, so this is how a Foldit puzzle goes. The game starts you from a rough model. You pull loops, rotate side chains, wiggle the whole chain and keep an eye on the score; the score is the guide, and many people share their best moves. Nobody needs to be a crystallographer to play; the game hides the chemistry behind shapes, colours and points. Players can join teams and share their best shapes with team-mates, so one person's opening moves can become another's starting point for the last tidy-up.
Disputed Who solved it? Many news stories said the gamers did. The paper says the players' models made molecular replacement work, and the crystallographers then solved and refined the structure. Both are true; how to share the credit is argued. [24][25]
Documented The ground had been moving for a decade. From about 2011, groups used evolution as a clue: pairs of residues whose letters change together across related sequences are often in contact in the fold. At CASP13 in 2018 the first AlphaFold, from DeepMind, came top using distances predicted by a neural network from clues of this kind. [28][29]
Documented At CASP14 in 2020, AlphaFold2 from DeepMind placed protein backbones with a median error of 0.96 Å (Cα RMSD over 95% of residues) across the targets, against 2.8 Å for the next best method. Its models were often close to experimental accuracy. [23]
Documented Every AlphaFold model carries a confidence score for each residue, pLDDT, which says how sure the model is about where that part sits. A model is a prediction, with its own error bars. [23]
Documented The AlphaFold2 paper came out in July 2021 with its code, so other groups could run the method themselves and check it. [23]
Documented In 2021 EMBL-EBI and DeepMind opened the AlphaFold Protein Structure Database, free to all. It has since grown to more than 200 million predicted structures. [30][31]
Documented For the human proteome, AlphaFold gave a confident prediction (pLDDT 70 or more) for 58% of residues, and a very high one (90 or more) for about 36%. The rest is a reminder of how much of a protein is not a single fixed shape. [32]
Documented In 2024 Demis Hassabis and John Jumper shared half of the Nobel Prize in Chemistry for protein structure prediction; David Baker received the other half for computational protein design. [33]
Reconstructed We have no record of what any one lab said when the CASP14 results came out, so this is how such a day goes for a structural biologist: open the results, look up the targets you have worked on, and compare the models with the structures you spent years on. Then look at the confidence colours, because a structure you have studied for years shows you where the model is right to be unsure.
Disputed Was the protein-folding problem solved? Some CASP organisers said that, for single protein chains, it largely was. Others point to what remains: proteins with more than one shape, disordered parts, complexes, and how the chain actually folds over time. [34][18]
Adams M. J., Blundell T. L., Dodson E. J., et al. (1969). Structure of rhombohedral 2 zinc insulin crystals. Nature 224, 491–495. doi.org/10.1038/224491a0
Kendrew J. C., Bodo G., Dintzis H. M., et al. (1958). A three-dimensional model of the myoglobin molecule obtained by X-ray analysis. Nature 181, 662–666. doi.org/10.1038/181662a0
Green D. W., Ingram V. M., Perutz M. F. (1954). The structure of haemoglobin IV. Sign determination by the isomorphous replacement method. Proc. R. Soc. A 225, 287–307. doi.org/10.1098/rspa.1954.0203
Kendrew J. C., Dickerson R. E., Strandberg B. E., et al. (1960). Structure of myoglobin: a three-dimensional Fourier synthesis at 2 Å resolution. Nature 185, 422–427. doi.org/10.1038/185422a0
Pauling L., Corey R. B., Branson H. R. (1951). The structure of proteins: two hydrogen-bonded helical configurations of the polypeptide chain. PNAS 37, 205–211. doi.org/10.1073/pnas.37.4.205
Perutz M. F., Rossmann M. G., Cullis A. F., et al. (1960). Structure of haemoglobin: a three-dimensional Fourier synthesis at 5.5 Å resolution. Nature 185, 416–422. doi.org/10.1038/185416a0
Anfinsen C. B. (1973). Principles that govern the folding of protein chains. Science 181, 223–230 (Nobel lecture). doi.org/10.1126/science.181.4096.223
Anfinsen C. B., Haber E., Sela M., White F. H. (1961). The kinetics of formation of native ribonuclease during oxidation of the reduced polypeptide chain. PNAS 47, 1309–1314. doi.org/10.1073/pnas.47.9.1309
Hartl F. U., Bracher A., Hayer-Hartl M. (2011). Molecular chaperones in protein folding and proteostasis. Nature 475, 324–332. doi.org/10.1038/nature10317
Chakravarty D., Porter L. L. (2022). AlphaFold2 fails to predict protein fold switching. Protein Science 31, e4353. doi.org/10.1002/pro.4353
Moult J., Pedersen J. T., Judson R., Fidelis K. (1995). A large-scale experiment to assess protein structure prediction methods. Proteins 23, ii–v. doi.org/10.1002/prot.340230303
Zemla A. (2003). LGA: a method for finding 3D similarities in protein structures. Nucleic Acids Res. 31, 3370–3374. doi.org/10.1093/nar/gkg571
Protein Structure Prediction Center (CASP), list of experiments. predictioncenter.org/
Berman H. M., Westbrook J., Feng Z., et al. (2000). The Protein Data Bank. Nucleic Acids Res. 28, 235–242. doi.org/10.1093/nar/28.1.235
Jumper J., Evans R., Pritzel A., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589. doi.org/10.1038/s41586-021-03819-2
Khatib F., DiMaio F., Foldit Contenders Group, Foldit Void Crushers Group, et al. (2011). Crystal structure of a monomeric retroviral protease solved by protein folding game players. Nat. Struct. Mol. Biol. 18, 1175–1177. doi.org/10.1038/nsmb.2119
Cooper S., Khatib F., Treuille A., et al. (2010). Predicting protein structures with a multiplayer online game. Nature 466, 756–760. doi.org/10.1038/nature09304
Gilski M., Kazmierczyk M., Krzywda S., et al. (2011). High-resolution structure of a retroviral protease folded as a monomer. Acta Cryst. D 67, 907–914. doi.org/10.1107/S0907444911035943
Marks D. S., Colwell L. J., Sheridan R., et al. (2011). Protein 3D structure computed from evolutionary sequence variation. PLoS ONE 6, e28766. doi.org/10.1371/journal.pone.0028766
Senior A. W., Evans R., Jumper J., et al. (2020). Improved protein structure prediction using potentials from deep learning. Nature 577, 706–710. doi.org/10.1038/s41586-019-1923-7
Varadi M., Anyango S., Deshpande M., et al. (2022). AlphaFold Protein Structure Database. Nucleic Acids Res. 50, D439–D444. doi.org/10.1093/nar/gkab1061
Varadi M., Bertoni D., Magana P., et al. (2024). AlphaFold Protein Structure Database in 2024. Nucleic Acids Res. 52, D368–D375. doi.org/10.1093/nar/gkad1011
Tunyasuvunakool K., Adler J., Wu Z., et al. (2021). Highly accurate protein structure prediction for the human proteome. Nature 596, 590–596. doi.org/10.1038/s41586-021-03828-1
Callaway E. (2020). 'It will change everything': DeepMind's AI makes gigantic leap in solving protein structures. Nature 588, 203–204. doi.org/10.1038/d41586-020-03348-4
Fold Commons · reading · updated 2026-10-07 · CC BY 4.0