Choice
The routes differ in which molecules you can enter, compute requirements and setup cost. Typical consequences of the wrong choice: spending a week downloading databases before finding out the GPU is unsupported, or reaching the Server before noticing the ligand has no CCD code.
| Route | When it fits | Limits |
|---|---|---|
| AlphaFold Server | A few jobs; proteins, DNA, RNA and ligands already in the CCD; no environment setup | 30 jobs per day; 5,000 tokens per job; no SMILES or custom covalent bonds; ions and modifications limited to a list |
| Local AF3 (v3.0.4) | SMILES ligands, covalent ligands, custom MSAs/templates, batches or more than 5,000 tokens | Linux; GPU with compute capability 8.0+; ~630 GB of databases; 64 GB RAM |
| ColabFold 1.6.3 (AF2-multimer) | Protein chains only; GPUs with about 16 GB; using the MMseqs2 server instead of local databases | No ligands, nucleic acids or modifications; ~2,000 residues on 16 GB |
| Boltz-2 | Protein-small molecule with affinity prediction | The affinity output has two fields: binder probability for separating binders from non-binders, and log10(IC50) for comparing binders |
| Chai-1 | AF3-style all-atom prediction installed with pip | Recommended A100 80GB, H100 80GB or L40S 48GB; users report success on RTX 4090 |
| Protenix v1 | Reimplementation with the same training cutoff as AF3, plus an applied model with a 2025-06 cutoff | Needs a GPU; input JSON is close to, but not identical with, the AF3 format |
| OpenFold3 | Open implementation aligned with the AF3 architecture; multi-GPU batches and a low-memory mode | Still labelled preview; default weights since 0.5.0 are OpenBind-0 |
Online
The figures below come from the AlphaFold Server FAQ. Tutorials from 2024 that cite 10 or 20 jobs per day and a fixed list of 19 ligands are out of date.
Signing in to the Server requires a Google account, and Google services are usually not directly reachable from networks in mainland China. If the Server is not an option, use local AF3, an AF3 installation on your university cluster (for example, Shanghai Jiao Tong University's Siyuan-1 provides a v3.0.2 container and shared parameters), or one of the open-source models above.
| Item | Current rule |
|---|---|
| Daily quota | 30 jobs per day; when the quota is used up you can save a draft and submit it after the quota refreshes |
| Job size | 5,000 tokens. 1 token per protein residue, 1 per DNA/RNA base, 1 per ligand atom, 1 per ion; modified residues count all their atoms; glycans count per atom |
| Minimum chain length | Each protein and nucleic acid chain needs at least 4 residues or bases |
| Ligands | 19 common ligands can be picked from a list; for others add a “CCD Code” entity and enter a 1–5 character CCD code (dictionary version 2024_10_28), up to 50 copies per ligand; SMILES cannot be entered |
| Ions | 10 ions: Ca²⁺, Co²⁺, Cu²⁺, Fe³⁺, K⁺, Mg²⁺, Mn²⁺, Na⁺, Zn²⁺, Cl⁻ |
| Glycosylation | Attach to N, S or T; residues limited to BGC, BMA, GLC, MAN, NAG (plus FUC on S/T); up to 8 glycan residues per glycan; glycosidic bond atoms cannot be specified |
| Not predicted | Water and hydrogen atoms; no awareness of membrane planes; non-standard amino acid codes B, J, O, U, X |
| Output | 5 samples per seed; the zip contains mmCIF and confidence JSON for every sample, job_request.json, paired and unpaired MSAs and up to 4 templates |
| Batches | JSON upload with up to 100 jobs per file; up to 500 saved drafts |
Online
[
{
"name": "kinase_atp_mg",
"modelSeeds": [],
"sequences": [
{
"proteinChain": {
"sequence": "MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPILSRVGDGTQDNLSGAEKAVQVKVKALPDAQ",
"count": 1,
"modifications": [{"ptmType": "CCD_SEP", "ptmPosition": 13}],
"glycans": [{"residues": "NAG(NAG)", "position": 47}]
}
},
{"ligand": {"ligand": "CCD_ATP", "count": 1}},
{"ion": {"ion": "MG", "count": 2}}
],
"dialect": "alphafoldserver",
"version": 1
}
]- 01
Add molecules per entity
Set copies for multiple copies of the same protein instead of pasting the sequence again. For double-stranded DNA tick “+ Reverse complement” to add the complementary strand; otherwise only a single strand is modelled. Replace unknown residues X with A, and unknown nucleotides N with T in DNA or U in RNA.
- 02
Trim disordered tails first
Long disordered termini produce low-confidence spurious helices, pull down pLDDT in nearby regions and use up tokens. The FAQ suggests removing disordered tails to get a clearer picture of confidence in ordered regions.
- 03
Check template settings
By default PDB templates released before 2021-09-30 are used. For a template-free control, or to steer the result towards a known structure, turn templates off, set a cutoff date or upload custom templates (up to 4 per chain) in the chain menu. With custom templates, the FAQ suggests also providing a shallow MSA of 10–100 sequences, otherwise the coevolution signal from the MSA overrides the template.
- 04
Use several jobs instead of several seeds
Each job submitted through the interface uses one seed and returns 5 samples. For difficult systems (antibody-antigen, protein-nucleic acid), use “Clone and reuse” with automatic seeds to get more candidates; each extra job adds 5 samples. For reproducibility, turn off automatic seeds on the preview page and enter an integer between 0 and 4,294,967,295.
- 05
Upload batches as JSON
Submit one job through the interface, download the zip, use its <job>_job_request.json as a template and generate the other jobs with a script, up to 100 per file; they appear as drafts after upload. Jobs beyond the daily quota stay in drafts until the next day. For CCD ligands outside the preset list, copy the JSON form from a job_request.json downloaded after adding a CCD Code entity in the interface.
- 06
Keep the zip for local reuse
The MSAs (a3m) and templates in the zip can be fed directly to local AF3; see the section on results that disagree with the Server. The Server notes that rerunning the same JSON later is not guaranteed to be bit-identical because of compiler optimisation changes. Tested on this page (AF3 repository as of 2026-10-09): a job_request.json downloaded from the Server now contains "version": 3, and local run_alphafold.py rejects it with AlphaFold Server input JSON has unsupported version: 3, expected 1; after changing it to 1 the file loads. A Server JSON with glycans fails locally with Specifying glycans in the `alphafoldserver` format is not supported and must be rewritten in the local dialect as shown below.
Local installation
Single-GPU inference time (excluding compilation): 1,024 tokens take 62 s on an A100 80GB and 34 s on an H100; 5,120 tokens take 2,547 s and 1,416 s.
Multiple GPUs do not speed up a single input. Each inference uses one GPU, selected with --gpu_device; on multi-GPU machines run several inputs at once. The developers report that a single A100 80GB is at least 2x more efficient in GPU time than 16 A100 40GB cards.
| Hardware | Capacity or problem | Required settings |
|---|---|---|
| A100 80GB / H100 80GB | 5,120 tokens on one GPU; numerical accuracy verified by the developers | Defaults |
| A100 40GB | 4,352 tokens at lower throughput | Unified memory plus a modified pair_transition_shard_spec |
| RTX 3090 / 4090 / A5000 / A6000 (compute capability 8.6/8.9) | Works; measurements in the next section | Large inputs may need the xla attention implementation or unified memory |
| RTX 50 series and other Blackwell GPUs | Supported out of the box since v3.0.2; v3.0.4 fixed unified memory on Blackwell | Use v3.0.4 |
| V100, RTX 20 series, Titan RTX (compute capability 7.x) | Without settings the structures look random with many clashes and ranking score ≤ -99; with them V100 handles 1,280 tokens | custom-kernel-fusion-rewriter in XLA_FLAGS and the xla attention implementation |
| P100 (compute capability 6.0) | 1,024 tokens | No changes |
| Compute capability < 6.0 or AMD GPUs | The program refuses to run | — |
| Disk | ~252 GB database download, ~630 GB unpacked, up to 1 TB for a full installation | Local SSD; the directory must not be inside the alphafold3 repository |
| RAM | The MSA search is memory-hungry, more so for long sequences | At least 64 GB |
| Operating system | Linux only; the maintainers state that WSL is not supported | Since v3.0.4, CPU (~100x slower) or Apple Silicon (mps, experimental) |
Field experience
These are results reported by users in GitHub issues, each from one or a few people. Where the conditions differ from your machine, treat them as order-of-magnitude references.
| GPU and settings | System | Reported result | Source |
|---|---|---|---|
| RTX 4090 24GB, defaults, 300 W power limit | Official 2PV7 example (homodimer, 2×298 residues) | Jackhmmer ~6 min, Hmmsearch ~3.5 min, GPU inference ~90 s | Issue #9 (user report) |
| RTX 3090 24GB, defaults | 2PV7 | 99.6 s inference per seed | Issue #9 (user report) |
| RTX 4090, defaults | 3 chains, ~800 residues in total; three copies of that trimer (~2,400 residues) | Both completed; the former essentially matched the Server result | Issue #9 (user report) |
| RTX 4090, --flash_attention_implementation=xla | Large complex (size not stated) | Completed, ~1 hour of inference, slowed by the xla implementation | Issue #9 (user report) |
| RTX 3090, 4090, A100 40GB/80GB | 2PV7 | ranking_score 0.67 on all of them, matching A100; compute capability 7.x cards gave -99 in the same test | Issue #59 (user report, ETH cluster) |
| Quadro P3000 6GB laptop (compute capability 6.1) | 167–334 tokens | 150–190 s per seed below 256 tokens; 618 s at 334 tokens | Issue #59 (user report) |
- nvidia-smi showing about 23 GB in use does not mean the card is nearly full. The Docker image sets XLA_CLIENT_MEM_FRACTION=0.95 and preallocates 95% of memory at start, so this number cannot be used to estimate how large a system still fits (maintainer explanation in issue #9).
- On workstations most of the time goes to the CPU MSA search. In the 4090 example above the MSA took about 9.5 minutes and inference only 1.5 minutes, so reusing MSAs comes before buying a faster GPU.
- If a 24 GB card runs out of memory on a large system, try --flash_attention_implementation=xla or unified memory first; both are noticeably slower. In 2024 users reported unified memory being unstable on the 4090 (issues #209 and #213); if it fails, first rule out WSL and beta drivers.
- On multi-GPU machines pin one GPU with --gpus device=0 (Docker) or --gpu_device and run different inputs in separate processes.
Field experience
AF3 supports Linux only, and the maintainers state that WSL is not supported. In issue #209 a 4090 user kept hitting memory errors under WSL, and the maintainers advised running on native Linux. Common alternatives are below.
Dual-boot Ubuntu or a separate Linux machine
The most reliable option. The official instructions are based on Ubuntu 22.04, the Docker image is based on Ubuntu 24.04 with CUDA 12.6.3, and the host needs a driver that supports CUDA 12.6. Put the databases on a separate 1 TB SSD.
Your university or institute cluster
Many clusters already provide an AF3 container, parameters and databases. For example, Shanghai Jiao Tong University's Siyuan-1 provides a v3.0.2 Singularity image with configurations for A100 40GB and A800 80GB and staged run scripts. Check your platform's user manual first.
Rent a Linux GPU cloud server
GPU rental platforms in China have community-built AF3 images. For example, one AutoDL image on GitHub requires driver 580 or newer and at least 700 GB free on the data disk (the databases take 627 GB). Download the databases once and keep them on the data disk for later jobs.
Test small systems on a Mac or CPU first
Since v3.0.4, --jax_backend=cpu or mps works for checking that an input JSON is correct; it is not suited to production runs.
Local installation
Docker is the route the developers verify. The official Dockerfile is based on CUDA 12.6.3 and Python 3.12 and sets XLA_FLAGS=--xla_gpu_enable_triton_gemm=false by default to avoid very long compile times. On clusters without Docker, build the Docker image first and convert it to a Singularity/Apptainer image.
# 1. Source code (keep databases and parameters outside the repository directory)
git clone https://github.com/google-deepmind/alphafold3.git
cd alphafold3
# 2. Databases: ~252 GB download, ~630 GB unpacked; an SSD is recommended
sudo apt install -y wget zstd
./fetch_databases.sh /data/af3_db
sudo chmod 755 --recursive /data/af3_db
# 3. Model parameters (direct download since July 2026; keep a single model file in the directory, no need to unpack .zst)
mkdir -p /data/af3_models
wget -P /data/af3_models https://storage.googleapis.com/alphafold3/af3.bin.zst
# 4. Build the image (host needs an NVIDIA driver, CUDA 12.6 and nvidia-container-toolkit)
docker build -t alphafold3 -f docker/Dockerfile .
# On RHEL / Rocky / AlmaLinux, if you hit "No file descriptors available (os error 24)":
# docker build --ulimit nofile=65535:65535 -t alphafold3 -f docker/Dockerfile .
# 5. Run
mkdir -p $HOME/af_input $HOME/af_output && chmod 755 $HOME/af_input $HOME/af_output
docker run -it \
--volume $HOME/af_input:/root/af_input \
--volume $HOME/af_output:/root/af_output \
--volume /data/af3_models:/root/models \
--volume /data/af3_db:/root/public_databases \
--gpus device=0 \
alphafold3 \
python run_alphafold.py \
--json_path=/root/af_input/fold_input.json \
--model_dir=/root/models \
--output_dir=/root/af_output# Non-Docker route (documented for CPU-only and macOS; usable on Linux GPU machines if CUDA/cuDNN/JAX work)
# Install HMMER (jackhmmer, nhmmer, etc. on PATH) and uv first
git clone https://github.com/google-deepmind/alphafold3.git
cd alphafold3
uv venv --python 3.12
source .venv/bin/activate
uv sync
uv run build_data # builds chemical component data; without it even --help fails with FileNotFoundError: ...chemical_component_sets.msgpack
uv run python run_alphafold_data_test.py # data pipeline self-test
# CPU only (~100x slower than GPU). The docs say Apple Silicon can use --jax_backend=mps; in our test mps produced wrong structures, see the checklist above
uv run run_alphafold.py \
--json_path=fold_input.json \
--model_dir=/data/af3_models \
--db_dir=/data/af3_db \
--output_dir=af_output \
--jax_backend=cpu \
--flash_attention_implementation=xla# On a machine where the build succeeds
docker build -t alphafold3 -f docker/Dockerfile .
docker save alphafold3 | gzip > alphafold3_image.tar.gz
# After copying it to the target machine
docker load < alphafold3_image.tar.gz- The parameter directory must contain a single model file. Putting af3.bin.zst and af3_synthid.bin.zst together raises Multiple models matched; the program reads .zst directly, so there is no need to unpack it.
- The parameter file is compatible with every 3.0.x release, so upgrading the code does not require a new download.
- The two most common build failures in China are timeouts while the Dockerfile installs Python dependencies, and CMake failing to download source archives such as abseil-cpp, pybind11 and libcifpp from GitHub/GitLab (a CSDN write-up got stuck at step 14/15 and at abseil-cpp). The reliable fix is to build once on a machine with a good connection, export with docker save and import with docker load.
- If docker run reports permission denied or Cannot connect to the Docker daemon, run sudo systemctl start docker, then add your user to the docker group (sudo usermod -aG docker $USER, effective after logging in again).
- The pip install -r dev-requirements.txt, Python 3.11 and conda instructions in older tutorials no longer apply: the repository no longer has dev-requirements.txt and pyproject.toml requires Python 3.12 or newer. The maintainers attribute errors such as DNN library initialization failed or ptxas version mismatches in conda environments to JAX/CUDA installation problems.
- Tested on this page (2026-10-10, AF3 repository as of 2026-10-09, jax 0.10.2, jax-mps 0.10.9, 8-core macOS arm64 with 16 GB): barnase–barstar (199 tokens) with MSAs generated by ColabFold, 1 seed × 5 samples. With --jax_backend=cpu, inference took 2,709 seconds, ipTM was 0.93–0.94 and DockQ against the experimental structure 1BRS was 0.97–0.98. The same input with --jax_backend=mps took 778 seconds and gave ipTM 0.17, a mean barstar pLDDT of 44.5 and DockQ 0.03–0.15: the structures were wrong, and the log showed no error. The mps path is marked experimental in the official docs; on a Mac, compare it with a CPU run on a system with an experimental structure before relying on it.
Field experience
The official script pipes wget into decompression with no resume and downloads 9 files in parallel; if any one breaks, it starts over. The 45 minutes quoted officially were measured on a GCP machine.
In the ColabFold route the MSAs come from MMseqs2, not from AF3's own Jackhmmer pipeline, so results differ; state this in the methods section of a paper. The public MSA server expects serial queries from a single IP, so do not submit from several machines at once.
Tested on this page: running --af3-json with ColabFold 1.6.3 on barnase–barstar (110 + 89 residues) produced the JSON in about 8 minutes, most of it spent queuing on the public server. The JSON is version 2; each chain has unpairedMsa and pairedMsa, and templates is an empty list, which means no templates are used during inference.
# The official fetch_databases.sh pipes wget into the decompressor; an interruption means re-downloading the whole file.
# On unstable networks, download the .zst files resumably first, then decompress them one by one.
SRC=https://storage.googleapis.com/alphafold-databases/v3.0
mkdir -p /data/af3_zst /data/af3_db && cd /data/af3_zst
for f in pdb_2022_09_28_mmcif_files.tar \
mgy_clusters_2022_05.fa \
bfd-first_non_consensus_sequences.fasta \
uniref90_2022_05.fa \
uniprot_all_2021_04.fa \
pdb_seqres_2022_09_28.fasta \
rnacentral_active_seq_id_90_cov_80_linclust.fasta \
nt_rna_2023_02_23_clust_seq_id_90_cov_80_rep_seq.fasta \
rfam_14_9_clust_seq_id_90_cov_80_rep_seq.fasta; do
aria2c -c -x 8 -s 8 "$SRC/$f.zst" # -c resumes; without aria2 use wget -c
done
# Decompress (delete each .zst once done, otherwise both copies need about 880 GB)
for f in *.fa.zst *.fasta.zst; do
zstd -d "$f" -o "/data/af3_db/${f%.zst}" && rm "$f"
done
tar --no-same-owner --no-same-permissions --use-compress-program=zstd \
-xf pdb_2022_09_28_mmcif_files.tar.zst -C /data/af3_db
chmod 755 --recursive /data/af3_db
# If someone already has a copy, rsync it; interrupted transfers resume
# rsync -avP user@host:/path/to/af3_db/ /data/af3_db/# Generate AF3-format JSON with MSAs via the ColabFold MMseqs2 server (JSON only, no prediction)
# Separate components in the FASTA with ":"; write non-protein components as "type|sequence" or "type|sequence|copies"; types: dna, rna, ccd, smiles
# colabfold_batch upper-cases the whole input, so lowercase aromatic atoms in SMILES become aliphatic:
# write SMILES in Kekulé form (e.g. CC(=O)OC1=CC=CC=C1C(=O)O) or add them to the JSON afterwards
colabfold_batch input_sequences.fasta msa_out --af3-json
# Then skip AF3's own data pipeline and run inference directly
python run_alphafold.py --json_path=msa_out/<name>.json --output_dir=out \
--model_dir=/data/af3_models --run_data_pipeline=false- Run the download inside tmux or screen so an SSH disconnect does not stop it.
- Tested on this page (HEAD requests for the 9 files on 2026-10-10): the .zst files total 238 GB (10^9 bytes); the largest are mgy_clusters_2022_05.fa.zst (69 GB) and pdb_2022_09_28_mmcif_files.tar.zst (57 GB). The bucket returns Accept-Ranges: bytes, so aria2c -c or wget -c can resume.
- When downloading first and decompressing later, keeping both the .zst files and the unpacked files needs about 880 GB; decompress one file at a time and delete its .zst to limit peak usage.
- If your group or university already has a copy, rsync -avP it; it is faster than downloading again and resumes after interruptions. The author of the AutoDL image also distributes the databases through cloud drives (user report); check that the file names match the list above.
- To get protein systems running before downloading the databases, generate MSAs with the ColabFold MMseqs2 server (code below) and run inference directly with --run_data_pipeline=false.
- Run chmod 755 --recursive after the download; with insufficient permissions the MSA tools fail with opaque errors that are hard to trace.
Input
The file below runs as is. The sequence is a format example; replace it with yours and update the modification and bond positions to match. Tested on this page: the file parses with folding_input from the AlphaFold 3 repository as of 2026-10-09 and featurises on macOS arm64; the modification position and bonded atom names pass the CCD checks.
{
"name": "kinase_atp_glyco_demo",
"modelSeeds": [1, 2, 3, 4, 5],
"sequences": [
{
"protein": {
"id": "A",
"sequence": "MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPILSRVGDGTQDNLSGAEKAVQVKVKALPDAQ",
"modifications": [{"ptmType": "SEP", "ptmPosition": 13}],
"description": "format demo: pSer13, N-glycan on Asn47"
}
},
{"ligand": {"id": "B", "ccdCodes": ["ATP"]}},
{"ligand": {"id": ["C", "D"], "ccdCodes": ["MG"]}},
{"ligand": {"id": "E", "smiles": "CC(=O)Oc1ccccc1C(=O)O"}},
{"ligand": {"id": "F", "ccdCodes": ["NAG", "NAG"]}}
],
"bondedAtomPairs": [
[["A", 47, "ND2"], ["F", 1, "C1"]],
[["F", 1, "O4"], ["F", 2, "C1"]]
],
"dialect": "alphafold3",
"version": 4
}- One job per file; a file whose top level is a list is read as the Server dialect.
- id accepts uppercase letters only (lowercase fails with IDs must be upper case letters); write multiple copies as a list such as ["C", "D"]. Multi-letter chain IDs skip the RASA calculation.
- In the local dialect, modification codes have no CCD_ prefix (SEP locally, CCD_SEP in the Server dialect); with the prefix the run fails (Protein ptms must not contain the "CCD_" prefix). This is the step most often missed when rewriting a Server JSON by hand.
- ptmPosition and bond residue numbers are 1-based; single-residue ligands use residue 1.
- Ions are ligands with ccdCodes, for example ["MG"].
- SMILES ligands cannot appear in bondedAtomPairs because they have no atom names; including one fails with Bond ... involves an unsupported SMILES ligand. Covalent ligands need a CCD code or a userCCD component.
- Backslashes in SMILES must be doubled in JSON; escape with jq -R . or Python json.dumps.
- Covalent bonds between or within polymers are not supported, for example disulfide-linked chains or head-to-tail cyclic peptides.
- For Failed to construct RDKit reference structure, raise --conformer_max_iterations or provide ideal coordinates in a userCCD.
- The name field must differ between jobs; use --input_dir to process a whole directory.
Sampling
Each seed produces 5 diffusion samples by default (--num_diffusion_samples=5). The output directory holds seeds × samples subdirectories, and the top-ranked structure is copied to the root. The AF3 paper benchmarks used 5 seeds × 5 samples, 25 candidates in total.
For single-chain proteins or high-confidence complexes, 1–5 seeds are enough. For antibody-antigen complexes the paper shows the top-ranked structure still improving at 1,000 seeds; for such systems start with at least 10–20 seeds and check whether ranking_score and ipTM level off.
Extra seeds do not require repeating the MSA search. Run the data pipeline alone with --run_inference=false, then run inference alone on the resulting _data.json.
# Step 1: CPU-only MSA/template search; writes <name>_data.json (can run on a machine without a GPU)
# To use --num_seeds in step 2, put exactly 1 value in modelSeeds of fold_input.json (e.g. [1]);
# with several values step 2 fails with Input must have one rng seed to set multiple seeds.
python run_alphafold.py --json_path=fold_input.json --output_dir=out \
--model_dir=/data/af3_models --db_dir=/data/af3_db --run_inference=false
# Step 2: GPU inference only from the JSON with MSAs; --num_seeds=10 generates 10 consecutive seeds starting from that seed
python run_alphafold.py --json_path=out/kinase_atp_glyco_demo/kinase_atp_glyco_demo_data.json \
--output_dir=out --model_dir=/data/af3_models \
--run_data_pipeline=false --force_output_dir=true \
--num_seeds=10 --num_diffusion_samples=5 \
--jax_compilation_cache_dir=$HOME/.cache/af3_jaxField experience
- Reuse MSAs: when screening all pairs between n and m proteins, run each chain once with --run_inference=false, then copy each chain's unpairedMsa, pairedMsa and templates fields into the dimer JSONs; the data pipeline drops from n×m runs to n+m (official performance.md).
- When swapping ligands or partner chains against the same target, compute the MSA of the fixed chain once and copy it into each new JSON, leaving the changing chain empty for the program to compute.
- CPU count: --jackhmmer_n_cpu and --nhmmer_n_cpu default to min(CPU cores, 8), and more than 8 gives almost no speed-up. On machines with many cores, run several inputs at once.
- Keep databases on a local SSD. On network storage such as NFS, parallel jobs tend to slow down or fail (maintainer assessment in issue #452).
- On many-core machines, shard the databases with src/alphafold3/scripts/shard_databases.sh and keep the shards on an SSD or a RAM-backed filesystem; the developers report 10–30x faster MSA search. Run ulimit -n 65535 before sharding.
- Turn on --jax_compilation_cache_dir: in the developers' test the second run of the same input dropped from 148 s to 32 s.
- For batches of similar-sized inputs, add --buckets entries close to the actual token counts to reduce padding; inputs larger than the largest bucket are recompiled at their exact size.
- Unified memory (TF_FORCE_UNIFIED_MEMORY=true) and the xla attention implementation both avoid out-of-memory errors at the cost of speed; turn them on only when you hit the error.
- Shanghai Jiao Tong University's cluster measured the data pipeline on A100-40GB dropping from 2,475 s in v3.0.0 to 677 s in v3.0.1, so upgrade old installations first.
Troubleshooting
# Not enough GPU memory (A100 40GB, 24 GB consumer cards) or more than 5,120 tokens: enable unified memory
export XLA_PYTHON_CLIENT_PREALLOCATE=false
export TF_FORCE_UNIFIED_MEMORY=true
export XLA_CLIENT_MEM_FRACTION=3.2
# On A100 40GB, also change pair_transition_shard_spec in model_config.py to:
# (2048, None), (3072, 1024), (None, 512)
# Compute capability 7.x (V100, RTX 20 series, Titan RTX, Quadro RTX): both settings are required, otherwise structures look random
export XLA_FLAGS="--xla_disable_hlo_passes=custom-kernel-fusion-rewriter"
python run_alphafold.py ... --flash_attention_implementation=xla
# In Docker pass them with -e, e.g. docker run -e XLA_FLAGS="..." -e TF_FORCE_UNIFIED_MEMORY=true ...| Symptom or error | Cause | Fix |
|---|---|---|
| Structure looks like a random coil, ranking score ≤ -99 | GPU compute capability 7.x | Set XLA_FLAGS=--xla_disable_hlo_passes=custom-kernel-fusion-rewriter and add --flash_attention_implementation=xla |
| implementation='triton' is unsupported on this GPU generation | GPUs below compute capability 8.0 do not support Triton flash attention | Add --flash_attention_implementation=xla |
| RESOURCE_EXHAUSTED / CUDA_ERROR_OUT_OF_MEMORY | Token count exceeds GPU memory | Enable unified memory (see the code below); on multi-GPU machines pin one GPU with --gpus device=0; under WSL, switch to native Linux |
| ptxas fatal: Unsupported .version 8.4; current version is '8.3' | The CUDA toolchain from conda or the system is older than JAX needs | Reinstall following the JAX instructions, or use Docker |
| unsupported version: 2, expected 1 | Outdated code or a third-party fork | Update to the current release; versions 1–4 are supported |
| Protein ptms must not contain the "CCD_" prefix | Server-dialect modification codes used in a local JSON | Remove the CCD_ prefix |
| Multiple models matched | More than one model file in the parameter directory | Keep only af3.bin.zst |
| Opaque errors from the MSA tools | Insufficient permissions on the database directory | chmod 755 --recursive the database directory |
| docker build is slow and the image is huge | Databases or parameters inside the repository directory | Move them out and rebuild |
| No file descriptors available (os error 24) | Low default file descriptor limit on RHEL/Rocky/AlmaLinux | Add --ulimit nofile=65535:65535 to docker build |
| NotImplementedError: Not supported on gpu. (Tokamax gated_linear_unit) | Some backends lack this kernel | Safe to ignore; another implementation is used automatically |
Field experience
The Server and local AF3 use the same model parameters; differences come mainly from the MSA. The Server runs Jackhmmer on sharded databases without --domZ, which effectively loosens --domE about 100x and gives deeper MSAs for some inputs. In issue #492 the same protein-DNA complex had ipTM 0.86 on the Server and 0.1 locally; with the Server's MSA the local result matched the Server.
{
"name": "use_server_msa",
"modelSeeds": [1, 2, 3, 4, 5],
"sequences": [
{
"protein": {
"id": "A",
"sequence": "<exactly the same sequence as the Server job>",
"unpairedMsaPath": "server_job/msas/fold_<job>_unpaired_msa_chains_a.a3m",
"pairedMsaPath": "server_job/msas/fold_<job>_paired_msa_chains_a.a3m"
}
}
],
"dialect": "alphafold3",
"version": 4
}- 01
Run several seeds on both sides first
In issue #385 one Server job gave ipTM 0.27 while 20 local samples ranged from 0.73 to 0.92; a maintainer reran on the Server with another seed and got 0.90, matching the local run. Each job submitted through the Server interface uses one seed, so run at least 3–5 seeds on both sides before comparing, to confirm the difference is systematic.
- 02
Compare MSA depth
Count the sequences in the unpairedMsa of the local <name>_data.json against the a3m files in the Server zip. In issue #492 the two differed by only 12 sequences (100 versus 112), and the results were completely different.
- 03
Use the Server's MSA directly
Put the paths of each chain's unpaired and paired a3m files from the zip into the local JSON (version 2 or later), omit the templates field so the local data pipeline searches templates again (templates: [] means no templates), and run as usual. The a3m files are in the msas/ folder of the zip. Tested on this page: a public Server download rewritten this way loads correctly in the local version.
- 04
Or loosen the local search
The developers' other option is to raise --domE for Jackhmmer/Nhmmer 100x. --domE is not a run_alphafold.py command-line flag and requires editing the tool configuration in the data pipeline, so in most cases using the Server's MSA is simpler.
- 05
When structures agree and only ipTM differs
The user in issue #385 built MSAs from one database at a time and got ipTM 0.75 (BFD), 0.91 (UniRef90), 0.70 (UniProt) and 0.90 (MGnify), with essentially the same structure. ipTM is inherently sensitive to the MSA source; when the structures superimpose, judge the interface from the distribution across seeds and from PAE rather than a single ipTM value.
Special systems
Protein-peptide
Enter the peptide as a second protein chain. Chains shorter than 16 residues get depressed pTM values (for example 0.03) because of how the TM-score formula scales for short chains; judge peptide binding from inter-chain PAE and the peptide's pLDDT.
Cyclic peptides
AF3 does not support covalent bonds within polymers, so a head-to-tail cyclic peptide cannot be written as a protein chain. The maintainers suggest writing the whole cyclic peptide as SMILES (as a small molecule) or defining it in a userCCD. The Server accepts no SMILES, so cyclic peptides need the local version.
Glycosylation and covalent ligands
Locally, define glycans as a multi-CCD ligand plus bondedAtomPairs; see the official examples/rnaseb_glycosylated.json. Except for glycans, leaving atoms of covalent ligands are not removed automatically; delete them after prediction and check bond lengths at the linkage.
RNA, DNA and protein complexes
Double-stranded DNA needs two chains. Clashes are most common in protein-nucleic acid complexes with more than 100 nucleotides and more than 2,000 residues in total. The documented Server-versus-local discrepancy (issue #492) is also a protein-DNA complex; see the previous section.
AF3 versus AF2 / AF2-multimer
AF2-multimer predicts protein chains only. AF3 adds nucleic acids, ligands, ions and modifications and is more accurate on protein complexes, but it can produce low-confidence spurious helices in disordered regions, where AF2 usually outputs ribbon-like extended chains. For protein-only complexes on limited GPU memory, AF2-multimer through ColabFold has the lowest setup cost.
Relation to molecular docking
AF3 outputs the ligand position directly, which is itself a binding-mode prediction, and the paper reports it significantly outperforms Vina on the PoseBusters benchmark. Its chirality violation rate is 4.4%, so check chiral centres, bond lengths and clashes before using ligand poses for docking scores or molecular dynamics. AF3 tends to predict the single conformation most common in the PDB (for cereblon it gives the closed state for both apo and holo), so it does not replace conformational sampling.
Results
<name>_model.cif in the root is the sample with the highest ranking_score, where ranking_score = 0.8×ipTM + 0.2×pTM + 0.5×disorder fraction − 100×has_clash and is meant only for ranking samples against each other. <name>_summary_confidences.json holds overall, per-chain and per-chain-pair pTM, ipTM and chain_pair_pae_min; <name>_confidences.json holds per-atom pLDDT and the full PAE matrix. Output is mmCIF only, with no PDB format; convert with tools such as gemmi if needed.
Official ipTM thresholds: above 0.8 is confident, below 0.6 is likely a failed prediction, and 0.6–0.8 needs PAE to decide. How each metric is computed, where the thresholds come from and common misreadings are covered in “How to read AlphaFold results”.
Practice in China
These practices come from Chinese community posts and university HPC manuals, and were checked against the official repositories or a second source.
Protenix web server: email sign-up, SMILES ligands
ByteDance's Protenix Server (protenix-server.com) uses email registration and accepts ligands as SMILES; AlphaFold Server supports neither. Ions still need CCD codes, for example FE2 for ferrous iron. Results are CIF files organised much like AF3 output (CSDN community post).
Protenix locally: weights and MSA do not go through Google
Model weights and CCD files are hosted on Volcengine object storage in Beijing (protenix.tos-cn-beijing.volces.com); protenix_base_default_v1.0.0.pt is 1.48 GB (checked with a HEAD request on 2026-10-10). Protein MSAs are sent to protenix-server.com/api/msa by default (--msa_server_mode protenix, or colabfold), and template search is off by default, so it runs without the 630 GB database. Official memory table: 1,000 tokens peak at 18.2 GB and 59 s; 2,000 tokens at 66.6 GB. A 24 GB card suits systems up to about 1,000 tokens. The current README lists two model generations, protenix-v2 (2026-04-08) and v1.0.0.
Installing Protenix: bypass a lagging pip mirror
With a pip mirror such as Tsinghua configured, the mirrored protenix package can lag behind GitHub, and the protenix pred commands in the README will not match. The README's fix is to point at the official index for this install: pip install --upgrade protenix --index-url https://pypi.org/simple.
University HPC: use the platform's container and shared databases
Besides SJTU's Siyuan-1, the South China University of Technology computing platform provides an AF3 Singularity image on its hpckapok1 and hpckapok2 clusters, with model parameters and databases in shared directories. The job script requests one card on the gpuA800 partition and uses singularity exec --nv to bind your own input and output directories to /root/af_input and /root/af_output. Before deploying anything, search your university's HPC manual for AlphaFold3.
Non-Docker builds: check the GCC version first
The non-Docker route most often fails at pip install ., where the C++ extensions are compiled. A record from a Japanese HPC user (Qiita) states that libcifpp needs GCC 11 or later and does not build below 9.4. A user in China, working in a conda environment on Ubuntu 24.04, installed gcc/gxx 12.4.0 and cmake 3.30.2 from conda-forge, added pybind11, and pip install --no-deps . then compiled (CSDN community post). If cloning from GitHub is slow, use the gitcode mirror https://gitcode.com/gh_mirrors/alp/alphafold3.git, and add -i https://pypi.tuna.tsinghua.edu.cn/simple to pip.
Hand it to an agent
Scientify's research environment comes with ColabFold 1.6.3 and AlphaFold2 weights preinstalled; it does not include AF3 parameters or databases. The preinstalled environment covers protein-only complex prediction; systems with ligands, nucleic acids or modifications need AF3 or an open-source model installed separately.
Example instruction: “Use ColabFold to predict the complex of the two protein chains in the attachment, run 5 models with 3 seeds each, rank by ipTM, and give PAE plots and inter-chain contact residue lists for the top three structures.”
- 01
Prepare the input
The agent formats the sequences as ColabFold FASTA or CSV and checks for non-standard residues and the number of chains.
- 02
Rent a GPU and run
Once you have authorised it, the agent rents a GPU as needed (billed per second) and runs colabfold_batch, keeping logs and intermediate files in the workspace. The task keeps running after you close your computer.
- 03
Summarise results
It produces PDB files for each model, a ranking table, PAE plots and interface residue lists, and reviews the results adversarially to check that the task is actually complete.
- 04
What you still check
Whether chain copy numbers match the real stoichiometry, whether interfaces with ipTM between 0.6 and 0.8 have experimental support, and whether disordered regions should be trimmed and the prediction rerun.
References
- google-deepmind/alphafold3 installation.md — System requirements, database size, Docker and uv installation, parameter download URL
- AlphaFold 3 performance.md — Token limits per GPU, unified memory, inference times, staged runs, MSA reuse, database sharding, compilation cache
- AlphaFold 3 known_issues.md — Wrong results on compute capability 7.x, Server versus local MSA differences
- AlphaFold 3 input.md — Input JSON fields, three ways to specify ligands, bonds and glycans, userCCD, external MSA paths
- AlphaFold 3 release notes (v3.0.1–v3.0.4) — Blackwell support, CPU and Apple Silicon support, new flags
- fetch_databases.sh and params.py — Database download URLs and method; matching rules for files in the parameter directory (src/alphafold3/model/params.py)
- AlphaFold Server FAQ — 30 jobs per day, token counting, CCD ligands, glycan rules, template and shallow-MSA advice, output contents
- AlphaFold Server JSON format — Server dialect fields and the allowed ligand, ion and modification lists
- Issue #9: running on 24 GB GPUs — User report: RTX 3090/4090 timings and system sizes
- Issue #59: wrong output on compute capability 7.x GPUs — User reports and maintainer replies: 2PV7 results across GPU models, P3000 laptop timings
- Issue #209: 4090 memory errors under WSL — Maintainers advise native Linux
- Issue #492: Server and local results disagree — Reproduction of ipTM 0.86 versus 0.1, MSAs differing by 12 sequences
- Issue #385: troubleshooting local versus Server differences — User report: low ipTM from a single Server seed, ipTM per database
- Issue #452: parallel runs and network storage — User report: databases on NFS
- Issue #629: entering head-to-tail cyclic peptides — Maintainers recommend SMILES
- Shanghai Jiao Tong University HPC manual: AlphaFold3 — User report: A100-40GB and A800 configurations, staged scripts, data pipeline time by version
- CSDN: AlphaFold3 installation guide (Docker build failures) — User report: dependency download timeouts and abseil-cpp download failures when building in China
- GitHub: AlphaFold3 AutoDL image notes — User report: driver and disk requirements of a cloud GPU image, database distribution via cloud drives
- Abramson et al., Nature 2024 — 5 seeds × 5 samples benchmark setup, 1,000 seeds for antibodies, 4.4% chirality violations, conformational bias
- ColabFold README — --af3-json for AF3 inputs, MSA server usage expectations
- Boltz, Chai-1, Protenix and OpenFold3 repositories — Features and hardware notes of open-source alternatives (see also jwohlwend/boltz, chaidiscovery/chai-lab, aqlaboratory/openfold-3)
- Protenix inference instructions and weight URLs — MSA server mode, memory and latency table; weight URLs in protenix/web_service/dependency_url.py
- CSDN: using the Protenix web server and deploying on Linux — Community post: email sign-up, SMILES ligands, CCD codes for ions
- SCUT computing platform: AlphaFold3 — University HPC manual: Singularity image, shared parameters and databases, A800 job script
- CSDN: local AlphaFold3 deployment notes (conda) — Community post: gitcode mirror, Tsinghua mirror, gcc 12.4.0, cmake 3.30.2, pybind11
- Qiita: AlphaFold3 non-Docker install notes — Community post: libcifpp needs GCC 11 or later