Fold Commons

CASPTracker

Thirty years of the Critical Assessment of Structure Prediction (CASP) — the biennial blind experiment that has measured protein-structure prediction since 1994. Each card below is one CASP experiment: its organiser, the widely-reported headline result, the group(s) publicly credited with topping the primary assessment, and a link to a public source. Nothing is installed; nothing about you is collected.

This is the tool — an honest headline-results record, not a mirror of the per-target GDT_TS / lDDT / TM-score / RMSD tables (those live on predictioncenter.org) — see the coverage gauge below. We report what the organiser and assessors reported; we show the metric, not a synthesised “winner”; and where a specific GDT_TS figure is not a clear matter of public record, we leave it blank rather than invent it. Editorial is CC BY 4.0; each experiment links a public source.

Headline median GDT_TS across CASP

Where a top-group median GDT_TS is a clear matter of public record it is plotted; most experiments are marked “not itemised” because this edition mirrors only headline results, not the full assessment tables.

Enable JavaScript to draw the GDT_TS-over-CASP chart — every figure is also listed in the cards below.

A headline-results record — not the full assessment tables

This edition tracks all 16 CASP experiments (CASP1–CASP16) at the level of headline outcomes — an estimated ≈40% of what the organiser publishes. The full per-target GDT_TS / lDDT / TM-score / RMSD tables are on predictioncenter.org; we show what we curate and name what we omit, so you can calibrate. As of 2026-09-17.

What’s covered

  • One record per CASP experiment, CASP1 (1994) through CASP16 (2024)
  • The organiser / host and the widely-reported headline result of each experiment
  • The group(s) publicly credited with topping the primary (tertiary-structure / regular-target) assessment where that is a matter of public record
  • A single headline numeric field (top-group median GDT_TS) only where a specific figure is a matter of clear public record — otherwise null

Known gaps

  • Per-target and per-group GDT_TS / lDDT / TM-score / RMSD tables are NOT mirrored here — only headline results (the full tables are published on predictioncenter.org)
  • CAPRI protein-protein docking rounds run alongside CASP are not itemised in this headline record
  • Sub-category winners (assembly / RNA / ligand / accuracy-estimation / contact) are summarised in prose, not enumerated per category with numbers
  • Exact top-group median GDT_TS is left null for every experiment except where a specific figure is a clear matter of public record (CASP14); we do not invent numbers
  • Early-experiment (CASP1–CASP5) group-level winners are reported at the level the organiser assessments state; fine-grained per-category leaders are omitted
Era
GDT_TS

16 of 16 experiments

  1. CASP1 1994 Founding (1994–1998) GDT_TS not itemised

    First CASP experiment, organised by John Moult and colleagues; results presented at Asilomar, California (December 1994).

    The first blind, community-wide assessment of protein structure prediction. Predictions were made for target sequences whose experimental structures were not yet public, establishing the double-blind format that still defines CASP. Comparative (homology) modelling showed early promise; ab initio prediction remained largely unsolved.

    Topped the primary assessmentNo single overall winner was crowned; CASP1 established the assessment framework rather than a leaderboard.

  2. CASP2 1996 Founding (1994–1998) GDT_TS not itemised

    Second CASP experiment; results meeting at Asilomar, California (December 1996). Organised by the Protein Structure Prediction Center.

    CASP2 consolidated the assessment categories (comparative modelling, fold recognition / threading, and ab initio). Threading methods drew particular attention for recognising folds with low sequence identity.

    Topped the primary assessmentAssessed by category by independent assessors; no single overall champion. Reported in the CASP2 special issue of Proteins.

  3. CASP3 1998 Founding (1994–1998) GDT_TS not itemised

    Third CASP experiment; results meeting at Asilomar, California (December 1998).

    CASP3 saw the introduction of fragment-assembly approaches to ab initio prediction; David Baker's Rosetta fragment-insertion method emerged as a leading approach for free-modelling targets.

    Topped the primary assessmentRosetta (Baker lab) was widely noted for its performance on ab initio / new-fold targets; category assessment reported in the Proteins CASP3 supplement.

  4. CASP4 2000 Founding (1994–1998) GDT_TS not itemised

    Fourth CASP experiment; results meeting at Asilomar, California (December 2000).

    CASP4 is widely regarded as the point where fragment-assembly ab initio methods, led by Rosetta, produced genuinely useful models for some new-fold targets. The GDT_TS score became a central assessment metric around this period.

    Topped the primary assessmentRosetta (Baker lab) was a standout on the ab initio / new-fold category; overall assessment reported in the Proteins CASP4 supplement.

  5. CASP5 2002 Template era (2000–2010) GDT_TS not itemised

    Fifth CASP experiment; results meeting at Asilomar, California (December 2002).

    CASP5 continued the rise of fragment-assembly methods and saw increasing use of meta-servers that combine multiple prediction servers. Comparative modelling remained the most reliable category.

    Topped the primary assessmentAssessed by category; Baker-lab Rosetta and leading server groups featured prominently. Reported in the Proteins CASP5 supplement.

  6. CASP6 2004 Template era (2000–2010) GDT_TS not itemised

    Sixth CASP experiment; results meeting at Gaeta, Italy (December 2004).

    CASP6 emphasised template-based modelling accuracy and model refinement. Automated servers narrowed the gap with human expert groups on many template-based targets.

    Topped the primary assessmentAssessed by category; no single overall champion. Reported in the Proteins CASP6 supplement.

  7. CASP7 2006 Template era (2000–2010) GDT_TS not itemised

    Seventh CASP experiment; results meeting at Asilomar, California (November 2006).

    CASP7 introduced a dedicated model-quality-assessment (QA) category and continued the trend of high-accuracy template-based modelling. Server + human hybrid pipelines dominated the reliable categories.

    Topped the primary assessmentAssessed by category; the Baker (Rosetta) and Zhang groups were among the strongest performers. Reported in the Proteins CASP7 supplement.

  8. CASP8 2008 Template era (2000–2010) GDT_TS not itemised

    Eighth CASP experiment; results meeting in Cagliari, Sardinia, Italy (December 2008).

    CASP8 confirmed the strength of consensus / meta-server approaches. The Zhang group's I-TASSER pipeline was widely recognised as a leading automated method for tertiary-structure prediction.

    Topped the primary assessmentZhang lab (I-TASSER) was among the top-performing groups for automated tertiary-structure prediction. Reported in the Proteins CASP8 supplement.

  9. CASP9 2010 Template era (2000–2010) GDT_TS not itemised

    Ninth CASP experiment; results meeting in Asilomar, California (December 2010).

    CASP9 continued incremental gains in template-based modelling and refinement. Contact-assisted and consensus methods matured, but ab initio prediction of large novel folds remained the central unsolved problem.

    Topped the primary assessmentAssessed by category; leading automated groups (including Zhang / I-TASSER and Baker / Rosetta lineages) featured. Reported in the Proteins CASP9 supplement.

  10. CASP10 2012 Co-evolution era (2010–2016) GDT_TS not itemised

    Tenth CASP experiment; results meeting in Gaeta, Italy (December 2012).

    CASP10 marked the emergence of co-evolution / residue–residue contact prediction from multiple sequence alignments as a route to fold prediction for targets lacking templates — the methodological seed of the later deep-learning surge.

    Topped the primary assessmentAssessed by category; the Baker and Zhang groups again featured among the strongest. Reported in the Proteins CASP10 supplement.

  11. CASP11 2014 Co-evolution era (2010–2016) GDT_TS not itemised

    Eleventh CASP experiment; results meeting in Riviera Maya, Mexico (December 2014).

    CASP11 saw contact-prediction-assisted modelling improve free-modelling results, with co-evolution methods increasingly integrated into fold prediction pipelines.

    Topped the primary assessmentAssessed by category; leading server and human groups performed strongly on template-based targets. Reported in the Proteins CASP11 supplement.

  12. CASP12 2016 Co-evolution era (2010–2016) GDT_TS not itemised

    Twelfth CASP experiment; results meeting in Gaeta, Italy (December 2016).

    CASP12 is widely cited as the point at which deep-learning-based contact prediction (e.g. RaptorX-Contact using residual neural networks) delivered a clear, measurable jump in free-modelling accuracy, foreshadowing the CASP13 deep-learning breakthrough.

    Topped the primary assessmentAssessed by category; deep-learning contact-prediction groups drew particular notice on hard targets. Reported in the Proteins CASP12 supplement.

  13. CASP13 2018 Deep-learning breakthrough (2018–2020) GDT_TS not itemised

    Thirteenth CASP experiment; results meeting in Riviera Maya, Cancún, Mexico (December 2018).

    The first AlphaFold (DeepMind), entered as group "A7D", topped the tertiary-structure (free-modelling) rankings — the first time a deep-learning system led CASP. AlphaFold 1 predicted inter-residue distance distributions with deep neural networks and used them to guide folding, producing a clear step-change on the hardest free-modelling targets.

    Topped the primary assessmentAlphaFold (DeepMind, group A7D) placed first overall on tertiary-structure prediction; the Zhang group and other established labs led among non-DeepMind entrants.

  14. CASP14 2020 Deep-learning breakthrough (2018–2020) GDT_TS ≈ 92.4

    Fourteenth CASP experiment; held virtually (December 2020) owing to the COVID-19 pandemic; organised by the Protein Structure Prediction Center.

    AlphaFold 2 (DeepMind, group "AF2") achieved a median backbone accuracy on the order of a GDT_TS of about 92.4 across targets — a level of accuracy competitive with experiment for many targets — and CASP organisers described protein structure prediction for single domains as, in large part, a solved problem. This is the result widely reported as the field's breakthrough moment.

    Topped the primary assessmentAlphaFold 2 (DeepMind) topped the tertiary-structure rankings by a wide margin; the Baker-lab and other groups led among non-DeepMind entrants.

  15. CASP15 2022 Post-AlphaFold (2022– ) GDT_TS not itemised

    Fifteenth CASP experiment; results meeting in Antalya, Türkiye (December 2022).

    With AlphaFold 2 open-source and widely available, CASP15 became a contest of AlphaFold-derived and enhanced pipelines rather than of a single system: top groups improved on plain AF2 by MSA engineering, sampling, and ensembling. CASP15 also ran a substantially expanded RNA / nucleic-acid category and emphasised protein assemblies (multimers).

    Topped the primary assessmentFor single-chain tertiary structure, AF2-based and AF2-enhanced groups (e.g. the Yang-Zhang / Zheng lineage and MSA-engineering groups) led; RoseTTAFold-family and specialised RNA-modelling groups featured in the newer categories. No single system swept every category.

  16. CASP16 2024 Post-AlphaFold (2022– ) GDT_TS not itemised

    Sixteenth CASP experiment; targets collected summer 2024; results conference in Punta Cana, Dominican Republic (December 1–4, 2024).

    CASP16 reaffirmed that single-domain protein folding is largely solved and shifted the frontier to complexes, ligands, and nucleic acids. AlphaFold 3, RoseTTAFold-All-Atom, and the Boltz / Chai families were newly available to all. Notably, in the protein-complex and antibody–antigen categories, a group layering traditional docking and extensive sampling on top of AF (the Kozakov–Vajda team) topped the field, showing hand-engineered pipelines still lead on the hardest complex cases.

    Topped the primary assessmentProtein complexes / multimers and antibody–antigen: Kozakov (Stony Brook) + Vajda (Boston University) team ("kozakovvajda"), the latter without using AF3/AFM directly on the antibody–antigen set. Oligomers: Cheng lab and Kihara lab. Stoichiometry: PreStoi (used by MULTICOM). Single-domain folding remained near-ceiling for AF-based methods.

Methods & limits

This page is the tool — free forever, on a path to open source, from a non-profit. A native version, when it ships, reads the full organiser exports and adds offline use; it never gates the web.