Metric reference
Every threshold in the table comes from official documentation or the original papers. Metrics marked "none" have no official threshold; do not borrow thresholds from other metrics.
Basis for the pLDDT bands: on a recent PDB test set, about 80% of χ1 side-chain angles are correct for residues with pLDDT > 90, and pLDDT > 70 corresponds to a generally correct backbone. Used as a disorder predictor (1 − 0.01×pLDDT), pLDDT reaches an AUC of 0.897 on the CAID benchmark (Tunyasuvunakool et al., 2021).
AF3 computes pLDDT per atom. For ligand atoms only the distance errors to polymers are counted; proteins use a 15 Å radius and DNA/RNA a 30 Å radius. Ligand pLDDT and protein-residue pLDDT are therefore not on the same scale.
The global AF3 iptm is an average over all interfaces. A developer confirmed in GitHub issue #565 that, when only one interface matters, you should read the corresponding element of chain_pair_iptm.
| Metric | What it measures | Range | Threshold | Source |
|---|---|---|---|---|
| pLDDT | Predicted accuracy of local distances around a residue (an atom in AF3), based on lDDT-Cα | 0–100, higher is better | >90 backbone and side chains accurate; 70–90 backbone generally correct; 50–70 low; <50 very low, long stretches read as disorder | AF2 human proteome paper; EBI course; AF3 output.md |
| PAE | Expected position error of residue j after aligning on the local frame of residue i | Å, lower is better; matrix is asymmetric | No universal threshold; contiguous low values in an inter-chain block mean the relative position is reliable | AF3 output.md; EBI course |
| pTM | Predicted TM-score of the whole structure | 0–1 | >0.5 the overall fold may be close to the true structure; forced below 0.05 when fewer than 20 tokens are involved | AF3 output.md |
| ipTM | Predicted TM-score of inter-chain relative positions, averaged over all interfaces | 0–1 | >0.8 high confidence; <0.6 usually failed; 0.6–0.8 grey zone | AF3 output.md; issue #565 |
| ranking_score | AF3 ranking score | −100 to 1.5 | None; only for ranking within one job | AF3 output.md |
| chain_pair_iptm | ipTM restricted to chains i and j; diagonal is single-chain pTM | 0–1 matrix | Same thresholds as ipTM; rank a known interface with it | AF3 output.md |
| chain_iptm | Average ipTM of the interfaces between one chain and all other chains | 0–1 | None; rank ligands or chains whose partner is unknown | AF3 output.md; issue #418 |
| chain_pair_pae_min | Minimum PAE over rows of chain i and columns of chain j | Å | None; officially correlated with whether two chains interact and in some cases separates binders from non-binders | AF3 output.md |
| fraction_disordered | Fraction of the structure that is disordered, estimated from accessible surface area | 0–1 | None; weighted +0.5 in ranking_score | AF3 output.md |
| has_clash | Clashing atoms exceed 50% of a chain, or a chain has more than 100 clashing atoms | 0 or 1 | When 1, ranking_score drops by 100; discard the model | AF3 output.md |
| contact_probs | Predicted probability that the representative atoms of two tokens are within 8 Å | 0–1 matrix | None | AF3 output.md |
Output files
The four common tools use different file names and different ranking rules for the "best model". First confirm that you are looking at the top-ranked model.
ColabFold's scores JSON keeps only two decimals for iptm and ptm. For more digits, log.txt gives three significant figures, or add --save-all to keep the raw pickle.
AF2 monomer ranked_0.pdb is ranked by mean pLDDT, not pTM, and ColabFold monomers are also ranked by pLDDT. Before comparing "best models", confirm that the tools used the same ranking metric.
| Tool | File | Contents | Notes |
|---|---|---|---|
| AF3 local | <job>_model.cif | Top-ranked structure; the B-factor column holds per-atom pLDDT | Highest ranking_score across all seeds × samples |
| AF3 local | <job>_summary_confidences.json | ptm, iptm, ranking_score, chain_pair_iptm, chain_pair_pae_min, chain_iptm, chain_ptm, fraction_disordered, has_clash, chain_ids | Read this file for judging and ranking. Tested on this page (AF3 repository as of 2026-10-09): chain_ids is a per-token list of chain IDs, as long as the number of tokens; the Server download's summary has no chain_ids |
| AF3 local | <job>_confidences.json | pae, atom_plddts, contact_probs, token_chain_ids, token_res_ids, atom_chain_ids | Used to plot PAE |
| AF3 local | <job>_ranking_scores.csv | ranking_score for every seed and sample | Shows how much seeds differ |
| AF3 local | seed-<s>_sample-<n>/ | One cif and two JSON files per sample | 5 samples per seed by default |
| AF3 local | <job>_data.json | Input with MSA and templates added | Reuse it with --run_data_pipeline=false to skip the MSA search when changing seeds |
| AlphaFold Server | fold_<job>_model_<N>.cif | N = 0–4, 0 is top-ranked | 1 seed and 5 models per job |
| AlphaFold Server | fold_<job>_summary_confidences_<N>.json / fold_<job>_full_data_<N>.json | Same fields as the local summary_confidences / confidences | full_data contains the PAE matrix |
| AlphaFold Server | fold_<job>_job_request.json | Submission parameters | Change the seed and resubmit |
| AF2 official | ranked_0.pdb | Best model | Monomers ranked by mean pLDDT; multimers by 0.8·ipTM + 0.2·pTM |
| AF2 official | ranking_debug.json | Ranking scores mapped to model names | Key is plddts (monomer) or iptm+ptm (multimer) |
| AF2 official | result_model_*.pkl | Raw outputs including plddt, ptm, iptm, predicted_aligned_error | Multimer runs 5 seeds per model by default, 25 predictions in total |
| ColabFold 1.6.3 | <job>_unrelaxed_rank_001_alphafold2_multimer_v3_model_<m>_seed_<s>.pdb | rank_001 is the best; B-factor column holds pLDDT | Complexes ranked by 80·ipTM + 20·pTM by default, monomers by pLDDT |
| ColabFold 1.6.3 | <job>_relaxed_rank_001_….pdb | Amber-relaxed structure | Only written with --amber |
| ColabFold 1.6.3 | <job>_scores_rank_00X_….json | plddt, pae, max_pae, ptm, iptm; complexes also have ipsae, pdockq and pdockq2, stored as dictionaries keyed by chain pair, e.g. "ipsae": {"A-B": 0.78, "B-A": 0.87}, where A-B means aligned on chain A and scored on chain B; log.txt shows the larger of the two directions | With --calc-extra-ptm also actifptm, pairwise_iptm, pairwise_actifptm and per_chain_ptm |
| ColabFold 1.6.3 | <job>_predicted_aligned_error_v1.json | PAE of the top-ranked model in AlphaFold DB format | Read directly by ChimeraX |
| ColabFold 1.6.3 | log.txt, <job>_pae.png, <job>_plddt.png, <job>_coverage.png, config.json, <job>.a3m, <job>_ext_metrics.png (with --calc-extra-ptm) | Per-model metrics, plots of the top 5 models, MSA coverage, run parameters, MSA | When ipTM is low, check MSA depth in coverage.png first |
PAE plot
- 01
Confirm the axes
AF3 defines element (i, j) as the position error of token j after alignment on the local frame of token i. Rows are the aligned residues and columns the scored residues. The matrix is asymmetric, so read the A→B and B→A inter-chain blocks separately.
- 02
Ignore the diagonal
The diagonal is each residue aligned on itself; it is always low and carries no information.
- 03
Read the intra-chain blocks
Two low-value squares within one chain separated by high values mean each domain is reliable but their relative orientation is not. In that case the relative position of the domains cannot be discussed even if every domain has pLDDT above 90.
- 04
Read the inter-chain blocks
Contiguous low values in the rows and columns of the interface mean the model is confident about the relative position of that part of the interface. If the whole inter-chain block is high, the relative position of the chains is essentially unconstrained, and two chains touching in the structure is not evidence of interaction.
- 05
Cross-check with pLDDT
Disordered segments with low pLDDT usually also have high PAE. Analyse individual interface residues only when residues on both sides have high pLDDT and inter-chain PAE is low.
Interaction
- 01
Discard clashing models
In AF3, check has_clash first and discard the model if it is 1. AF2 and ColabFold have no such field; check the structure for interpenetrating chains.
- 02
Read chain_pair_iptm for the target pair
For a dimer read ipTM; for multi-chain complexes read the corresponding element of chain_pair_iptm. Above 0.8, continue; below 0.6, treat the interaction as not predicted; 0.6–0.8 needs evidence from the following steps.
- 03
Read inter-chain PAE
chain_pair_pae_min reflects only the single best residue pair. Also count inter-chain residue pairs with PAE below 10 Å (the pairs<10A column of the script below). If only a few pairs are low, the interface evidence is weak.
- 04
Read interface pLDDT
Discuss which residues are in contact only when residues on both sides have pLDDT ≥ 70; discussing side-chain interactions such as salt bridges and hydrogen bonds needs pLDDT > 90.
- 05
Check agreement across seeds
Run at least 5 seeds: list 5 values in modelSeeds for local AF3; AlphaFold Server runs one seed per job, so resubmit with different seeds; use --num-seeds 5 in ColabFold. Compare whether the top-ranked models share the same binding site; DockQ can score one model against another used as reference.
- 06
Add non-polymer context
If the native complex contains metal ions, cofactors or ligands, add them to the AF3 input. AF3 developers explained in issues #142 and #474 that confidence depends on this context and may be lower overall without it.
- 07
Run a negative control
Predict the target protein with the same settings against a protein known not to interact, or against a scrambled peptide with the same amino-acid composition. The control scores are the background for your system. In the Dunbrack 2025 benchmark, the ipTM and ipSAE distributions of true and false dimers overlap between 0.3 and 0.7, so a single score cannot separate them.
ipTM limitations
Long disordered regions or accessory domains: ipTM averages over whole chains and d0 is set from the total chain length, so residues that do not bind pull the score down. In the KRAS–RAF1 RBD example of Dunbrack 2025, ipTM is 0.9 with domains only; after adding 120 disordered residues to each chain, AF2 ipTM drops to 0.59 although the interface is unchanged, while ipSAE is 0.8. Trimming to domains raises ipTM, but the higher score then describes only the trimmed construct.
Short chains and peptides: with fewer than 20 tokens pTM is forced below 0.05, and the documentation recommends reading PAE and pLDDT instead. Flexible flanks at peptide ends also lower ipTM; actifpTM was designed for this case.
Large complexes: the global iptm mixes all interfaces, so one chain pair can be diluted or masked by the others. Use chain_pair_iptm for a single interface and chain_iptm for ligands.
pTM and ipTM are computed from the PAE probability distribution; the PAE stored in JSON is only the expected value, so ipTM cannot be recomputed exactly from it (ColabFold issue #194). The alternatives below are approximations based on expected PAE or pLDDT.
| Metric | How it is computed | When to use | Threshold | Source |
|---|---|---|---|---|
| ipSAE | Counts only inter-chain residue pairs with PAE below a cutoff, sets d0 from those residues, takes the maximum over directions | Interactions with disordered regions or accessory domains; domain–peptide interactions | No universal threshold in the paper; PAE cutoff of 10 or 15 Å recommended | Dunbrack 2025; built into ColabFold 1.6.3 |
| pDockQ | Mean interface pLDDT × log10(number of interface contacts) mapped through a sigmoid to predicted DockQ; contacts are Cβ pairs ≤ 8 Å | AF2 heterodimers | AUC 0.95 for separating acceptable models (DockQ ≥ 0.23); SpeedPPI saves structures above 0.5 by default; not suited to overlapping chains such as homomers | Bryant 2022; FoldDock, SpeedPPI |
| pDockQ2 | Combines interface pLDDT and interface PAE, per chain pair | Single interfaces in multi-chain complexes | No universal threshold verified here | Zhu 2023; built into ColabFold 1.6.3 |
| actifpTM | Computes ipTM only over residues actually in the interface | Peptides and motif-mediated interactions with flexible flanks | None | Varga 2025; ColabFold --calc-extra-ptm |
| LIS / iLIS | Mean confidence over inter-chain regions with PAE ≤ 12 Å; iLIS adds a Cβ ≤ 8 Å contact filter | Large-scale PPI screens, flexible interactions | iLIS ≥ 0.223; or LIS ≥ 0.203 and LIA ≥ 3432 | Kim 2024, 2026; AFM-LIS |
Misreadings
Treating pLDDT or ipTM as binding strength
These metrics are trained to predict structural error (lDDT, TM-score) and contain no affinity information. Pak et al. 2023 found almost no correlation between AlphaFold output metrics and stability changes caused by point mutations. Comparing binding strength requires docking scores, free-energy calculations or experimental measurement.
Low-pLDDT regions are simply wrong
Long stretches with pLDDT < 50 appear as ribbons in AF2 and mean "probably disordered in isolation". The human proteome paper also found that low pLDDT concentrates in regions with a high fraction of contacts to other chains: adding the binding partner to the prediction can raise pLDDT there.
High pLDDT always means a stable structure
For disordered regions that fold only on binding, AF2 tends to predict the folded state with high pLDDT (EBI course). AF3 is a diffusion model and can generate ordered structure in disordered regions; these regions usually have very low pLDDT but lack the AF2 ribbon appearance, so they are easy to mistake for real structure.
High pTM means the interface is right
pTM is computed over all residues. When a large protein is predicted accurately it dominates the score, so pTM can stay high even if a small partner is misplaced. Judge interfaces with ipTM or chain_pair_iptm.
Comparing ranking_score across jobs
ranking_score contains +0.5·fraction_disordered and −100·has_clash, and the documentation states it is only for ranking. Compare different constructs or partners with chain_pair_iptm and similar metrics.
Reporting the best seed when seeds disagree
Different seeds giving different binding sites means the model is not confident about the interface. In the AF3 paper, antibody–antigen predictions kept improving with more seeds (up to 1,000), while other classes generally did not. For antibodies, run more seeds and rank by ipTM; for other systems, report disagreement across seeds as uncertainty. Agreement across seeds shows only that the result is stable: a developer stressed in issue #142 that confidence correlates with accuracy but does not guarantee it.
Code
One script handles local AF3, AlphaFold Server and ColabFold and depends only on numpy and matplotlib. It prints mean pLDDT per chain, minimum and mean PAE for each chain pair, the number of residue pairs below the PAE cutoff, and the summary metrics.
Black lines in the heatmap mark chain boundaries. Rows are aligned tokens and columns scored tokens, matching the AF3 definition of pae[i][j]. In AF3 every ligand atom is a token, so a ligand occupies as many rows as it has heavy atoms.
#!/usr/bin/env python3
"""Read confidence JSON from AlphaFold 3 / AlphaFold Server / ColabFold, print chain-pair metrics and plot the PAE heatmap.
AlphaFold 3 local: python af_pae.py job_confidences.json --summary job_summary_confidences.json
AlphaFold Server: python af_pae.py fold_job_full_data_0.json --summary fold_job_summary_confidences_0.json
ColabFold: python af_pae.py job_scores_rank_001_xxx.json --lengths 110,89
"""
import argparse
import json
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
import numpy as np
def segments(ids):
"""Split per-token chain IDs into [(chain, start, end)], end exclusive."""
segs, start = [], 0
for i in range(1, len(ids) + 1):
if i == len(ids) or ids[i] != ids[start]:
segs.append((ids[start], start, i))
start = i
return segs
def main():
ap = argparse.ArgumentParser()
ap.add_argument("conf", help="JSON containing the pae matrix")
ap.add_argument("--summary", help="AF3 summary_confidences JSON")
ap.add_argument("--lengths", help="ColabFold only: chain lengths, comma-separated")
ap.add_argument("--cutoff", type=float, default=10.0, help="PAE cutoff (A) for counting interface pairs")
ap.add_argument("--out", default="pae.png")
args = ap.parse_args()
with open(args.conf) as f:
d = json.load(f)
pae = np.asarray(d.get("pae", d.get("predicted_aligned_error")), dtype=float)
n = pae.shape[0]
if "token_chain_ids" in d: # AF3 local and AlphaFold Server
chain_of_token = d["token_chain_ids"]
elif args.lengths: # ColabFold: scores JSON has no chain IDs
lens = [int(x) for x in args.lengths.split(",")]
if sum(lens) != n:
raise SystemExit(f"sum of --lengths ({sum(lens)}) does not match PAE size {n}")
chain_of_token = [chr(65 + k) for k, m in enumerate(lens) for _ in range(m)]
else:
chain_of_token = ["A"] * n
segs = segments(chain_of_token)
# Mean pLDDT per chain
if "atom_plddts" in d and "atom_chain_ids" in d: # AF3: per atom
pl = np.asarray(d["atom_plddts"])
ids = np.asarray(d["atom_chain_ids"])
for c in dict.fromkeys(d["atom_chain_ids"]):
print(f"chain {c}: mean pLDDT (atoms) = {pl[ids == c].mean():.1f}")
elif "plddt" in d: # ColabFold: per residue
pl = np.asarray(d["plddt"])
for c, s, e in segs:
print(f"chain {c}: mean pLDDT (residues) = {pl[s:e].mean():.1f}")
# Chain-pair PAE: rows = aligned chain, columns = scored chain
print(f"\naligned scored min_PAE mean_PAE pairs<{args.cutoff:g}A")
for ci, si, ei in segs:
for cj, sj, ej in segs:
if ci == cj:
continue
b = pae[si:ei, sj:ej]
print(f"{ci:>7} {cj:>6} {b.min():8.2f} {b.mean():9.2f} {int((b < args.cutoff).sum()):10d}")
# Summary metrics
for key in ("ptm", "iptm", "ipsae", "pdockq", "pdockq2", "actifptm"):
if key in d:
print(f"{key}: {d[key]}")
if args.summary:
with open(args.summary) as f:
s = json.load(f)
for key in ("ptm", "iptm", "ranking_score", "fraction_disordered", "has_clash"):
print(f"{key}: {s.get(key)}")
# In local AF3 chain_ids is a per-token list and the Server omits it; deduplicate to get chain order
cids = list(dict.fromkeys(s.get("chain_ids") or [])) or [c for c, _, _ in segs]
for key in ("chain_pair_iptm", "chain_pair_pae_min"):
if key in s:
print(f"\n{key} (rows/cols = {cids})")
for c, row in zip(cids, s[key]):
print(c, " ".join(" nan" if v is None else f"{v:5.2f}" for v in row))
# PAE heatmap
fig, ax = plt.subplots(figsize=(6, 5), dpi=200)
im = ax.imshow(pae, cmap="Greens_r", vmin=0, vmax=30)
for _, s, _ in segs[1:]:
ax.axhline(s - 0.5, color="black", lw=0.8)
ax.axvline(s - 0.5, color="black", lw=0.8)
ax.set_xlabel("Scored token")
ax.set_ylabel("Aligned token")
fig.colorbar(im, ax=ax, label="Expected position error (A)")
fig.tight_layout()
fig.savefig(args.out)
print(f"\nsaved {args.out}")
if __name__ == "__main__":
main()# Requires only numpy and matplotlib
pip install numpy matplotlib
# AlphaFold 3 local run (top-ranked model)
python af_pae.py myjob/myjob_confidences.json --summary myjob/myjob_summary_confidences.json --out myjob_pae.png
# AlphaFold Server (after unzipping; N=0 is the top-ranked model)
python af_pae.py fold_myjob_full_data_0.json --summary fold_myjob_summary_confidences_0.json --out pae_0.png
# ColabFold: the scores JSON has no chain IDs, so pass chain lengths in order with --lengths
python af_pae.py out/barnase_barstar_scores_rank_001_*.json --lengths 110,89 --out pae_rank1.pngVisualisation
AF2 and ColabFold PDB files and AF3 mmCIF files store pLDDT in the B-factor column; AF3 values are per atom. Higher pLDDT means more reliable, the opposite of a crystallographic B-factor, so convert it before using a model for molecular replacement.
ChimeraX uses a continuous colour map by default, unlike the four-band AlphaFold DB scheme. The ChimeraX documentation discourages the four-band scheme because residues at pLDDT 71 and 69 appear in very different colours.
load fold_myjob_model_0.cif, m
hide everything, m
show cartoon, m
# The B-factor column holds pLDDT. AlphaFold DB four-band colours: <=50 orange, 50-70 yellow, 70-90 light blue, >90 dark blue
color 0xFF7D45, m
color 0xFFDB13, m and b > 50
color 0x65CBF3, m and b > 70
color 0x0053D6, m and b > 90
# Continuous colouring (optional): red below 50, blue above 90
# spectrum b, red_yellow_green_cyan_blue, m, minimum=50, maximum=90
# Interface residues: any atom within 5 A of the partner chain
select ifA, byres ((m and chain A) within 5 of (m and chain B))
select ifB, byres ((m and chain B) within 5 of (m and chain A))
select iface, ifA or ifB
show sticks, iface and not name N+C+O
# Print pLDDT of interface residues (CA atoms)
iterate iface and name CA, print(chain, resi, resn, round(b, 1))
# Save as PDB
save model_0.pdb, mopen fold_myjob_model_0.cif
# ChimeraX default continuous pLDDT colouring
color bfactor #1 palette alphafold
# Load PAE: AF3 full_data / confidences JSON and ColabFold scores JSON are read directly
alphafold pae #1 file fold_myjob_full_data_0.json plot true
# Draw pseudobonds between chain A and B residue pairs within 4 A and with PAE <= 5 A
alphafold contacts #1/A toAtoms #1/B distance 4 maxPae 5
# Cluster rigid domains by PAE and colour them (default connectMaxPae 5 A)
alphafold pae #1 colorDomains truepip install gemmi
# --shorten trims chain names to 1-2 characters; add --shorten-tlc if a ligand uses a 5-character monomer name
gemmi convert --shorten fold_myjob_model_0.cif model_0.pdbExample
With a system that has an experimental structure, DockQ lets you set confidence metrics next to the true error and calibrate your own reading of the thresholds. The experimental barnase–barstar structure is PDB 1BRS.
One asymmetric unit of 1BRS contains three complexes (A–D, B–E, C–F), and barstar is the C40A/C82A mutant; the sequences above match 1BRS. --mapping AB:AD maps model chains A and B to 1BRS chains A and D.
DockQ classes: < 0.23 incorrect, 0.23–0.49 acceptable, 0.49–0.80 medium, ≥ 0.80 high quality. Tabulating ipTM, ipSAE and DockQ for all 25 models shows directly whether high-ipTM models also have high DockQ. Then check whether barnase Arg59 and His102 appear at the model interface.
By default ColabFold sends sequences to the public MMseqs2 server to build the MSA. Complexes use --pair-mode unpaired_paired by default, combining paired and unpaired MSAs; for protein pairs from different species with little co-evolutionary signal the paired MSA is shallow and ipTM is often low, so check coverage.png first.
Tested on this page (2026-10-10): ColabFold 1.6.3 on an 8-core macOS arm64 CPU (no GPU), MSA from the public MMseqs2 server, with --num-models 5 --num-seeds 1 --calc-extra-ptm (5 models only, no --amber). Every model met the early-stop criterion after the first recycle (tol 0.285 < 0.5). Each model took about 10–12 minutes (other jobs were using the CPU at the same time; the first model included JAX compilation and took about 59 minutes), about 105 minutes for all 5. DockQ 2 was computed with the --mapping AB:AD command above.
All 5 models are high quality (DockQ ≥ 0.80) and their ipTM values differ by only 0.004, yet the top-ranked model has the lowest DockQ and the fifth-ranked model the highest. When all scores are high and close, the ranking order does not reflect the order of structural accuracy. The interface selected with the PyMOL commands in the previous section includes barnase Arg59 (pLDDT 97.1) and His102 (pLDDT 98.8). The scores JSON stores the top model's iptm as 0.92; log.txt shows 0.919.
For comparison, we predicted the same complex locally with AF3 (CPU backend, 1 seed × 5 samples) using the same ColabFold MSA: ipTM 0.93–0.94, chain_pair_pae_min 0.78–0.79 Å and DockQ 0.97–0.98; all 5 samples scored higher than the best ColabFold model.
cat > barnase_barstar.fasta <<'EOF'
>barnase_barstar
AQVINTFDGVADYLQTYHKLPDNYITKSEAQALGWVASKGNLADVAPGKSIGGDIFSNREGKLPGKSGRTWREADINYTSGFRNSDRILYSSDWLIYKTTDHYQTFTKIR:KKAVINGEQIRSISDLHQTLKKELALPEYYGENLDALWDALTGWVEYPLVLEWRQFEQSKQLTENGAESVLQVFREAKAEGADITIILS
EOF
# 5 models x 5 seeds = 25 predictions; complexes are ranked by 80*ipTM + 20*pTM by default
colabfold_batch --num-models 5 --num-seeds 5 --calc-extra-ptm \
--amber --num-relax 1 --use-gpu-relax \
barnase_barstar.fasta out/
# Tabulate ipTM, pTM, ipSAE and pDockQ2 for all 25 models (the last two are in the scores JSON from 1.6.3 on)
# ipsae and pdockq2 are dictionaries keyed by chain pair (A-B, B-A); take the larger value, as log.txt does
python - <<'PY'
import glob, json
for f in sorted(glob.glob("out/barnase_barstar_scores_rank_*.json")):
s = json.load(open(f))
rank = f.split("_scores_")[1].split("_alphafold2")[0]
mx = lambda v: max(v.values()) if isinstance(v, dict) and v else v
print(rank, s.get("iptm"), s.get("ptm"), mx(s.get("ipsae")), mx(s.get("pdockq2")))
PY
# Compare with the experimental structure 1BRS: model chains A, B map to 1BRS chains A (barnase) and D (barstar)
wget https://files.rcsb.org/download/1BRS.pdb
pip install DockQ
for f in out/barnase_barstar_unrelaxed_rank_*.pdb; do
DockQ "$f" 1BRS.pdb --mapping AB:AD --short | tail -1
done| Rank | Model | ipTM | pTM | ipSAE (larger direction) | pDockQ | pDockQ2 (larger direction) | DockQ (vs 1BRS A/D) |
|---|---|---|---|---|---|---|---|
| rank_001 | model_3 | 0.919 | 0.930 | 0.869 | 0.462 | 0.925 | 0.845 |
| rank_002 | model_4 | 0.917 | 0.930 | 0.866 | 0.481 | 0.915 | 0.875 |
| rank_003 | model_1 | 0.917 | 0.929 | 0.865 | 0.488 | 0.911 | 0.883 |
| rank_004 | model_2 | 0.915 | 0.929 | 0.863 | 0.490 | 0.915 | 0.877 |
| rank_005 | model_5 | 0.915 | 0.928 | 0.862 | 0.477 | 0.908 | 0.901 |
Chinese tutorials
These problems appear in several Chinese tutorials; the corrections follow the official documentation.
Colour direction opposite to AlphaFold DB
The common spectrum b, blue_white_red, and color_b.py with gradient=bgr (whose author notes that pLDDT runs blue, green, red from low to high), both show high pLDDT in red. This carries over the crystallographic B-factor convention where red means flexible, the opposite of AlphaFold DB where dark blue means high confidence, so readers see the most reliable regions as the least reliable. color_b.py with mode=ramp puts the same number of atoms in each colour bin, so the bin edges depend on the structure and the same colour means different pLDDT in different models. For publication figures, colour by fixed thresholds (the four-band scheme in the Visualization section, or spectrum with minimum=50, maximum=90) and include a colour bar.
Select by pLDDT with b
In PyMOL, pLDDT sits in the B-factor column; there is no separate pLDDT property. Select low-confidence regions with select low, m and b < 50; expressions such as plddt < 50 fail with a selection syntax error.
There is no ranked_0.cif in the Server download
Some tutorials say to look for ranked_0.cif after unzipping. ranked_0 is the AF2 local file name; in the AlphaFold Server download the top structure is fold_<job>_model_0.cif, and in local AF3 it is <job>_model.cif.
chain_ptm does not describe interfaces
Some tutorials describe chain_ptm as the confidence of a chain's interface with other chains, and treat chain_pair_iptm and chain_pair_pae_min as similar. AF3 output.md defines chain_ptm as the pTM restricted to chain i, which reflects only that chain's own fold; for interfaces use chain_iptm or chain_pair_iptm. chain_pair_pae_min is a minimum PAE in Å and is not on the same scale as ipTM.
Two interface PAE filters, computed differently
Protein binder design papers use two values. pae_interaction averages the means of the two inter-chain blocks (A→B and B→A); Bennett et al. 2023 used AF2 pae_interaction < 10 as a filter. min PAE interaction takes the smaller of the two inter-chain block minima; the off-diagonal elements of chain_pair_pae_min in AF3 summary_confidences.json are exactly these block minima and can be read directly (a Chinese community example [[0.76, 1.29], [1.4, 0.76]] gives 1.29). These thresholds were calibrated on de novo designed binders and have not been calibrated in the same way for natural protein interactions, so still follow the multi-metric workflow in the interaction section.
Hand off to an agent
Example instruction: "Predict the barnase–barstar complex with ColabFold, 5 models × 5 seeds, compare with PDB 1BRS, report ipTM, pTM, ipSAE, pDockQ2, inter-chain PAE and DockQ for every model, and check whether barnase Arg59 and His102 are at the interface."
- 01
Prepare the environment
Rent a GPU, select the bioinformatics environment and run the environment self-check.
- 02
Run the prediction
Write the FASTA and run script, execute colabfold_batch, and keep log.txt and config.json.
- 03
Tabulate and compare
Use the script in this guide to collect metrics and plot PAE for every model, download 1BRS, run DockQ and build a summary table.
- 04
Adversarial review
Check that all 25 models were evaluated, that the DockQ chain mapping is correct and that the numbers in the report match the JSON, then release the GPU.
- Deliverables: FASTA, run commands, all models, scores JSON, PAE and pLDDT plots, metric summary table, DockQ results, analysis report
- Still check yourself: the chain correspondence with the experimental structure
- Still check yourself: whether the interface residues agree with the literature
- Still check yourself: wording in the report that presents predictions as conclusions
Writing the paper
The minimum information a reviewer needs to reproduce the prediction and judge its reliability:
- Methods: software and version (e.g. ColabFold 1.6.3, AlphaFold 3 v3.0.1), model weights (e.g. alphafold2_multimer_v3), MSA source (MMseqs2 server or local databases and their versions), whether templates were used, seeds × models or samples, number of recycles, whether Amber relaxation was applied, ranking metric
- Which model was used and why: rank_001 or highest ranking_score, or re-ranked by chain_pair_iptm
- Table of values: pTM, ipTM (chain_pair_iptm of the target pair for multi-chain complexes), mean pLDDT per chain and mean pLDDT of interface residues for every reported model; with disordered regions, add ipSAE (state the PAE cutoff) or pDockQ2
- Distribution across seeds: median and range of ipTM over all models, or whether the top 5 models agree on the binding site
- Figures: structure coloured by pLDDT with a colour key; PAE heatmap with chain boundaries, a colour bar in Å and the same colour range across figures
- If an experimental structure exists: DockQ or interface RMSD
- Data: model coordinates and confidence JSON in the supplementary material, or deposited in ModelArchive, which assigns each model a DOI that can be cited in the paper
References
- AlphaFold 3 output documentation (output.md) — File layout, metric definitions, ipTM and pTM thresholds, ranking_score formula
- Abramson et al. 2024, Nature (AlphaFold 3) — Hallucination, chirality violation rate, antibody–antigen multi-seed results
- Tunyasuvunakool et al. 2021, Nature (human proteome) — Basis for pLDDT bands, disorder prediction AUC
- AlphaFold 2 README — AF2 output files, ranked_0 ranking rule, pLDDT in the B-factor column
- ColabFold 1.6.3 release notes — ipSAE and pDockQ2 output
- ColabFold batch.py (v1.6.3) — Output file naming, ranking rules, command-line options
- EBI AlphaFold training course — Interpreting pLDDT and PAE, AlphaFold Server output files
- Dunbrack 2025, ipSAE (bioRxiv) — Dilution of ipTM by disordered regions and the ipSAE fix
- DunbrackLab/IPSAE — ipSAE script and PAE cutoff examples
- Bryant et al. 2022, Nature Communications (pDockQ) — Definition of pDockQ
- FoldDock README — pDockQ parameters, AUC 0.95, homomer limitation
- Varga et al. 2025, Bioinformatics (actifpTM) — Interface confidence with flexible flanks
- AFM-LIS — LIS and iLIS definitions and thresholds
- Pak et al. 2023, PLoS One — AlphaFold metrics do not track mutational stability changes
- AlphaFold 3 issue #492 — Server vs local ipTM difference caused by the MSA (reproduced by a developer)
- AlphaFold 3 issue #385 — ipTM variation across genetic databases
- AlphaFold 3 issues #142, #474 — Confidence depends on ligand and ion context (developer replies)
- AlphaFold 3 issues #565, #418 — Relationship between iptm, chain_pair_iptm and chain_iptm
- AlphaFold 3 issue #606 — Community thread: scrambled peptides as negative controls
- ColabFold issue #194 — Community thread: ipTM cannot be recomputed exactly from expected PAE
- ChimeraX alphafold command documentation — alphafold pae, alphafold contacts and pLDDT colouring
- gemmi command-line documentation — --shorten option of gemmi convert
- DockQ — DockQ classes and chain mapping
- ModelArchive — Archiving computed models with DOIs
- CSDN: colouring proteins by AlphaFold pLDDT in PyMOL — Community post: color_b.py bgr gradient and ramp binning
- CSDN: complex prediction and interface visualization (AlphaFold + PyMOL) — Community post: spectrum b, blue_white_red and ranked_0.cif
- CSDN: interpreting AF3 structures — Community post: pae_interaction vs min PAE interaction, chain_pair_pae_min example
- Bennett et al. 2023, Nature Communications — AF2 pae_interaction < 10 as a binder design filter