Route
Version information in the table was checked on 2026-10-10. If you need GPU acceleration, only the first two routes give you the current release and your own choice of CUDA version.
| Route | Available version | NVIDIA GPU | When it fits | Main limitations |
|---|---|---|---|---|
| WSL2 + source build | Any version, currently 2026.4 | Yes (CUDA) | Windows 10/11 PC with an NVIDIA GPU | nvidia-smi is limited under WSL2; devices cannot be filtered by index on multi-GPU systems; sources must live on the Linux file system |
| Linux source build | Any version, currently 2026.4 | Yes (CUDA/SYCL/HIP) | Workstations, servers, clusters | Needs CMake ≥ 3.28 and GCC ≥ 11; CentOS 7 (GCC 4.8.5) and Rocky/RHEL 8 (GCC 8.5) need a newer compiler first |
| conda-forge | 2026.3 (2026.4 is out upstream) | CUDA builds on linux-64 only | No root access, want something running quickly | No Windows builds; the CUDA build depends on cuda-toolkit 12.9, so the driver must support 12.9; x86 SIMD tops out at AVX2_256; regression tests are not run during packaging |
| Ubuntu apt | 2023.3 on 24.04, 2021.4 on 22.04 | No | Teaching or small CPU-only systems | Two to three years behind; no CUDA runtime among the dependencies |
| NVIDIA NGC container | Latest tag 2023.2 (updated 2023-08) | Yes | GPU servers that already run Docker/Singularity | No longer updated; GMX_ENABLE_DIRECT_GPU_COMM in the example command is a 2022/2023 setting |
| Native Windows (MSVC) | Build it yourself | Yes (CUDA, tested upstream) | Must run natively on Windows | No official Windows binaries; GMX_BUILD_OWN_FFTW does not work, so use MKL, fftpack or your own FFTW build; no PLUMED; CMake < 4.3 fails to detect OpenMP_CUDA for CUDA builds |
Version changes
Most tutorials online use versions 2018–2021. The table lists requirements and syntax that have changed in 2026.x.
| Item | GROMACS 2026.x | Old tutorial syntax | What happens if you copy it |
|---|---|---|---|
| CMake | ≥ 3.28 (since 2025) | cmake3, 3.13–3.17 | Configuration fails immediately; apt on Ubuntu 22.04 gives 3.22.1 (too old), 24.04 gives 3.28.3 (enough) |
| C++ compiler | GCC ≥ 11, Clang ≥ 14, MSVC 2019; the classic Intel compiler icc is no longer supported | GCC 5–7 | The system GCC on CentOS 7 and Rocky 8 cannot build it |
| Enable CUDA | -DGMX_GPU=CUDA | -DGMX_GPU=ON | Error: Invalid value for GMX_GPU: ON. Pick one of: OFF, CUDA, OpenCL, SYCL, HIP |
| CUDA path | -DCUDAToolkit_ROOT=/usr/local/cuda-12.9 | -DCUDA_TOOLKIT_ROOT_DIR=… | The GROMACS 2026 CMake scripts no longer read this variable; configuration only ends with Manually-specified variables were not used by the project: CUDA_TOOLKIT_ROOT_DIR (tested on this page with CMake 4.4). With several CUDA versions installed another nvcc may be found, so check the actual version in the cmake output |
| GPU architecture | -DCMAKE_CUDA_ARCHITECTURES=89 | -DGMX_CUDA_TARGET_SM=… | Still accepted with a deprecation warning; 2026 switched to the standard CMake variable |
| CUDA version | ≥ 12.1, compute capability ≥ 5.0 | CUDA 9–11 | CUDA 11.x fails at configuration |
| CUDA 13 | Supported since 2025.3; CUDA 13 cannot build for GPUs below compute capability 7.5 | Not mentioned | GTX 10 series, P100 and V100 cannot use the GPU after a CUDA 13 build; 2025.0–2025.2 with CUDA 13 fails at link time |
| AMD GPUs | HIP backend with full offload since 2026, or SYCL + AdaptiveCpp | OpenCL | OpenCL is deprecated and does not support RDNA AMD GPUs or NVIDIA GPUs from Volta onwards |
| GPU update and constraints | -update auto runs on the GPU by default (since 2023) | Add -update gpu by hand or set GMX_FORCE_UPDATE_DEFAULT_GPU | The old variables GMX_GPU_DD_COMMS and GMX_GPU_PME_PP_COMMS have been removed and have no effect |
| MPI | MPI 3.0 required (since 2026) | MPI 2.x | Single-node runs only need the built-in thread-MPI; no MPI build is required |
| PLUMED | Interface bundled since 2025, loaded at run time | Always plumed patch first | Routine use no longer needs a patch; see the PLUMED section |
Linux
GMX_BUILD_OWN_FFTW=ON downloads FFTW 3.3.10 and builds it in single precision with the SIMD options GROMACS recommends. The GROMACS team advises against distribution FFTW packages because they default to double precision and are not compiled with options tuned for GROMACS.
Check two lines in the cmake output: whether the SIMD instruction set matches the CPU (for example AVX2_256 or AVX_512), and, for CUDA builds, "Compiling GROMACS for CUDA architectures: …". When you compile on a login node and run on compute nodes, set -DGMX_SIMD for the compute nodes; a binary built for AVX-512 exits with an illegal instruction on a machine that only supports AVX2. For GPU runs the GROMACS team recommends AVX2_256.
In a CUDA build, nvcc uses CMAKE_CXX_COMPILER as its host compiler. If the system GCC is newer than your CUDA supports, specify an older compiler for both, for example -DCMAKE_C_COMPILER=gcc-12 -DCMAKE_CXX_COMPILER=g++-12, instead of editing CUDA's host_config.h.
# Ubuntu 24.04 (including Ubuntu 24.04 under WSL2); on 22.04 apt ships cmake 3.22, upgrade first with pip install cmake
sudo apt update
sudo apt install -y build-essential cmake wget perl
wget https://ftp.gromacs.org/gromacs/gromacs-2026.4.tar.gz
tar xfz gromacs-2026.4.tar.gz
cd gromacs-2026.4
mkdir build && cd build
cmake .. \
-DGMX_BUILD_OWN_FFTW=ON \
-DREGRESSIONTEST_DOWNLOAD=ON \
-DCMAKE_INSTALL_PREFIX=$HOME/opt/gromacs-2026.4
make -j $(nproc)
make check
make install
echo 'source $HOME/opt/gromacs-2026.4/bin/GMXRC' >> ~/.bashrc
source $HOME/opt/gromacs-2026.4/bin/GMXRC
gmx --version# Prerequisites: nvcc works (nvcc --version) and nvidia-smi lists the GPU
export PATH=/usr/local/cuda-12.9/bin:$PATH
cd gromacs-2026.4
mkdir build-cuda && cd build-cuda
# CMAKE_CUDA_ARCHITECTURES is the compute capability without the dot:
# RTX 20/T4=75, A100=80, RTX 30/A40=86, RTX 40/L40=89, H100=90, RTX 50=120
# If omitted, every architecture from 50 to 120 that nvcc accepts is built: slower build, larger binary
cmake .. \
-DGMX_GPU=CUDA \
-DCUDAToolkit_ROOT=/usr/local/cuda-12.9 \
-DCMAKE_CUDA_ARCHITECTURES=89 \
-DGMX_BUILD_OWN_FFTW=ON \
-DREGRESSIONTEST_DOWNLOAD=ON \
-DCMAKE_INSTALL_PREFIX=$HOME/opt/gromacs-2026.4-cuda
make -j $(nproc)
make check
make install
source $HOME/opt/gromacs-2026.4-cuda/bin/GMXRC
gmx --version | grep -E "GPU support|CUDA targets|CUDA driver|CUDA runtime|SIMD"
# If the server cannot reach fftw.org or ftp.gromacs.org, download these two files elsewhere and copy them over
# http://www.fftw.org/fftw-3.3.10.tar.gz
# https://ftp.gromacs.org/regressiontests/regressiontests-2026.4.tar.gz
tar xfz regressiontests-2026.4.tar.gz -C $HOME/src
cmake .. \
-DGMX_GPU=CUDA \
-DGMX_BUILD_OWN_FFTW=ON \
-DGMX_BUILD_OWN_FFTW_URL=$HOME/src/fftw-3.3.10.tar.gz \
-DREGRESSIONTEST_PATH=$HOME/src/regressiontests-2026.4 \
-DCMAKE_INSTALL_PREFIX=$HOME/opt/gromacs-2026.4-cudaWindows
After these steps, build with the CUDA commands from the previous section. Known WSL2 limitations: nvidia-smi offers a limited feature set; GPUs cannot be filtered by index number on multi-GPU machines; concurrent CPU/GPU access to the same memory is not supported. Single-GPU GROMACS runs are not affected by these limitations.
# Windows PowerShell (administrator)
wsl --install -d Ubuntu-24.04
wsl --update
# Optional: %UserProfile%\.wslconfig; run wsl --shutdown after editing, then reopen WSL
# [wsl2]
# memory=16GB
# processors=8# Run inside Ubuntu on WSL; first confirm the Windows driver is passed through
nvidia-smi
wget https://developer.download.nvidia.com/compute/cuda/repos/wsl-ubuntu/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt-get update
# Install only the toolkit metapackage; cuda, cuda-12-9 and cuda-drivers pull in a Linux driver
sudo apt-get install -y cuda-toolkit-12-9
echo 'export PATH=/usr/local/cuda-12.9/bin:$PATH' >> ~/.bashrc
source ~/.bashrc
nvcc --version
# Keep sources and build directories on the Linux file system, not under /mnt/c
mkdir -p ~/src && cd ~/src- 01
Install or update the NVIDIA driver on Windows
NVIDIA's WSL guide states this is the only driver you need, and that no Linux display driver should be installed inside WSL. The Windows driver is exposed inside WSL2 as libcuda.so, and nvidia-smi lives in /usr/lib/wsl/lib. The CUDA Version shown at the top right of nvidia-smi inside WSL should be at least the toolkit version you plan to install.
- 02
Install WSL2 and Ubuntu 24.04
CUDA is supported on WSL 2 only, not WSL 1. Ubuntu 24.04 ships CMake 3.28.3 and GCC 13.2 via apt, which meet GROMACS 2026's requirements, so you can skip a manual CMake upgrade.
- 03
Install only the cuda-toolkit-12-x metapackage
Use NVIDIA's wsl-ubuntu repository. The cuda, cuda-12-x and cuda-drivers metapackages install a Linux driver that overrides the WSL libcuda mapping and leaves the GPU unusable. CUDA 12.9 can build for compute capability 5.0 through 12.0, covering the RTX 50 series as well as the GTX 10 series.
- 04
Keep sources and build directories under ~
Microsoft recommends storing files under /home/<user> when working from a Linux command line, not under /mnt/c. Building under /mnt/c is slow, and a GROMACS forum user building there got "Clock skew detected" warnings.
- 05
Lower the parallelism if the build runs out of memory
By default WSL2 gets only 50% of Windows memory. If make -j stops with "c++: fatal error: Killed signal terminated program cc1plus", memory ran out; lower the -j value, or raise memory in .wslconfig and run wsl --shutdown.
conda
# linux-64 only (including WSL2). The CUDA build needs a driver that supports CUDA 12.9
conda create -n gmx -c conda-forge "gromacs=2026.3=nompi_cuda*"
conda activate gmx
gmx --version | grep -E "GPU support|SIMD"
# CPU/OpenCL build
# conda create -n gmx-cpu -c conda-forge "gromacs=2026.3=nompi_h*"
# ~/.condarc (conda-forge setup from the TUNA help page; run conda clean -i afterwards)
# channels:
# - conda-forge
# - nodefaults
# show_channel_urls: true
# custom_channels:
# conda-forge: https://mirrors.tuna.tsinghua.edu.cn/anaconda/cloud
conda clean -i
conda create -n gmx -c conda-forge "gromacs=2026.3=nompi_cuda*"
# With mamba, use mirrored_channels instead
# channels:
# - conda-forge
# mirrored_channels:
# conda-forge:
# - https://mirrors.tuna.tsinghua.edu.cn/anaconda/cloud/conda-forge
# For source builds, upgrade CMake with pip (3.28 or newer is required)
pip install -i https://pypi.tuna.tsinghua.edu.cn/simple "cmake>=3.28"
One patch release behind upstream
On the check date conda-forge had 2026.3 while upstream had 2026.4. The bioconda gromacs package stopped at 2021.3, so do not install from bioconda.
CUDA builds exist for linux-64 only
aarch64, ppc64le and macOS have no CUDA builds, and there are no Windows builds at all. The CUDA build installs cuda-toolkit 12.9 into the environment and checks the driver through the __cuda virtual package.
SIMD tops out at AVX2_256
On x86 the recipe builds only SSE2, AVX_256 and AVX2_256 variants and picks one at start-up based on the CPU. On CPUs with AVX-512, CPU-only runs may be slower than a self-compiled build.
Regression tests are not run during packaging
REGRESSIONTEST_DOWNLOAD is commented out in the recipe, and the package test only runs gmx -version. If you need trustworthy results, download regressiontests-2026.3 yourself and run gmxtest.pl.
Testing
The official installation guide's rule: an individual failed test may be a compiler bug, or a tolerance that is slightly too tight. Inspect the output files the script points to; if the failure is real, rebuild with a newer compiler. If the tests still fail, post on the GROMACS forum with a description of the hardware and the full output of gmx mdrun -version.
The most common case on the GROMACS forum is a GPU build where many tests report Timeout. A developer explained that tests use every core and GPU they detect by default, so they time out when the CPU or GPU is busy with other jobs. In one 2021.1 case, 23 tests timed out with GPU detection enabled and only 1 remained with GMX_DISABLE_GPU_DETECTION=1. If the failures are only timeouts and the CPU-only run passes, confirm the GPU works with the water box self-test in the next section; numerical comparison failures (output shows reference values that do not match the computed ones) should not be ignored.
# Show only the failed tests
make check 2>&1 | tee check.log
grep -A40 "tests FAILED" check.log
# 1. Rule out the GPU: if the CPU run passes and the GPU run times out, the GPU is usually busy
GMX_DISABLE_GPU_DETECTION=1 make check
# 2. On multi-GPU machines, expose a single GPU to the tests
CUDA_VISIBLE_DEVICES=0 make check
# 3. Timeouts on slow machines or under WSL: raise the timeout factor and reconfigure
cmake .. -DGMX_TEST_TIMEOUT_FACTOR=3 && make check
# 4. Rerun a single test with full output
ctest -R MdrunTests --output-on-failureErrors
| Error text (excerpt) | Cause | Fix |
|---|---|---|
| Invalid value for GMX_GPU: ON | Syntax from before 2021 | Use -DGMX_GPU=CUDA |
| Could not find `nvcc` executable in path specified by variable CUDAToolkit_ROOT=/usr/local/cuda-12.9 | There is no bin/nvcc under the directory given in CUDAToolkit_ROOT (CMake text tested on this page on a machine without CUDA) | Check the actual CUDA directory with ls /usr/local/ and correct CUDAToolkit_ROOT |
| CMake 3.28 or higher is required | System CMake is too old (3.22.1 on Ubuntu 22.04, older on CentOS 7) | pip install cmake, or use Kitware's official binaries |
| #error -- unsupported GNU version! gcc versions later than … are not supported! | GCC is newer than the host compilers your CUDA supports | Install a GCC your CUDA supports and select it with -DCMAKE_C_COMPILER/-DCMAKE_CXX_COMPILER, or upgrade CUDA |
| nvcc fatal : Unsupported gpu architecture 'compute_xx' | CMAKE_CUDA_ARCHITECTURES or the old GMX_CUDA_TARGET_SM lists an architecture the current CUDA does not support (for example 61 with CUDA 13) | Use the GPU's actual compute capability; for GPUs below compute capability 7.5 switch to CUDA 12.x |
| undefined reference to `void gmx::nbnxn_kernel_prune_cuda<true>(…)' | GROMACS 2025.0–2025.2 with CUDA 13 | Upgrade to 2025.3 or later, use CUDA 12.9, or add -DCMAKE_CUDA_FLAGS=-static-global-template-stub=false |
| Unknown CUDA architecture: native / ptxas error : Value of threads per SM … is out of range | Old CMake or outdated source | Update the source, delete the build directory, reconfigure, and set an explicit value such as -DCMAKE_CUDA_ARCHITECTURES=86-real |
| Cannot find FFTW 3 (with correct precision - libfftw3f for mixed-precision GROMACS …) | Only double-precision FFTW is installed, or CMAKE_PREFIX_PATH does not point to FFTW | Use -DGMX_BUILD_OWN_FFTW=ON; for offline builds add -DGMX_BUILD_OWN_FFTW_URL=<absolute path> |
| Cannot build FFTW3 automatically (GMX_BUILD_OWN_FFTW=ON) in Visual Studio / with ninja | Automatic FFTW builds are not supported for native Windows builds or the Ninja generator | Use -DGMX_FFT_LIBRARY=mkl, or build FFTW yourself first |
| WARNING: The gmx binary does not include support for the CUDA architecture of the GPU ID #0 | The CMAKE_CUDA_ARCHITECTURES used at build time does not include this GPU | Rebuild with the value suggested in the log, for example -DCMAKE_CUDA_ARCHITECTURES=89 |
| Cannot run short-ranged nonbonded interactions on a GPU because no GPU is detected. | The binary has GPU support, but no usable GPU is found at run time: the driver is unavailable, or a Linux driver was installed inside WSL | Check nvidia-smi; under WSL remove cuda-drivers-type packages |
| Nonbonded interactions on the GPU were requested with -nb gpu, but the GROMACS binary has been built without GPU support. | The binary itself has no GPU support; the GPU support line of gmx --version says disabled (tested on this page with the conda-forge 2026.3 CPU build) | Install a CUDA build or rebuild as described above |
| make stops with Killed signal terminated program cc1plus | Out of memory during compilation, common on WSL2 and small VMs | Lower the -j value or add memory |
| mdrun -plumed aborts at start-up with MPI_ERR_COMM in MPI-parallel runs | Bug in 2026.0–2026.3 | Upgrade to 2026.4 |
Verification
First run gmx --version and check that the GPU support line says CUDA, CUDA targets includes your GPU's architecture, and CUDA driver is not older than CUDA runtime. Builds with the PLUMED interface show its status on the Plumed support line.
Then run a water box. The -nb, -pme and -update options of mdrun all default to auto: if no usable GPU is detected, mdrun silently falls back to the CPU and only runs slower. Request gpu explicitly for the self-test so that an unusable GPU stops the run with an error.
mkdir -p ~/gmx-gpu-test && cd ~/gmx-gpu-test
# 5 nm cubic water box, about 4000 SPC/E waters
gmx solvate -cs spc216.gro -box 5 5 5 -o water.gro
cat > topol.top <<'EOF'
#include "oplsaa.ff/forcefield.itp"
#include "oplsaa.ff/spce.itp"
[ system ]
SPC/E water box
[ molecules ]
EOF
echo "SOL $(grep -c OW water.gro)" >> topol.top
cat > em.mdp <<'EOF'
integrator = steep
emtol = 1000
nsteps = 5000
cutoff-scheme = Verlet
coulombtype = PME
rcoulomb = 1.0
rvdw = 1.0
EOF
cat > md.mdp <<'EOF'
integrator = md
dt = 0.002
nsteps = 25000
cutoff-scheme = Verlet
coulombtype = PME
rcoulomb = 1.0
rvdw = 1.0
tcoupl = V-rescale
tc-grps = System
tau-t = 0.1
ref-t = 300
gen-vel = yes
gen-temp = 300
constraints = h-bonds
nstcalcenergy = 100
nstenergy = 1000
nstlog = 5000
EOF
gmx grompp -f em.mdp -c water.gro -p topol.top -o em.tpr
gmx mdrun -s em.tpr -c em.gro -g em.log
gmx grompp -f md.mdp -c em.gro -p topol.top -o md.tpr
# Request gpu explicitly: if no GPU is usable, mdrun stops with an error instead of silently falling back to the CPU
gmx mdrun -s md.tpr -g md.log -nb gpu -pme gpu -update gpu -ntmpi 1
grep -A2 "Mapping of GPU IDs" md.log
grep -E "PP task|PME tasks" md.log
grep "Performance:" md.log GPU info:
Number of GPUs detected: 1
#0: NVIDIA NVIDIA GeForce RTX 4090, compute cap.: 8.9, ECC: no, stat: compatible
1 GPU selected for this run.
Mapping of GPU IDs to the 2 GPU tasks in the 1 rank on this node:
PP:0,PME:0
PP tasks will do non-perturbed short-ranged interactions on the GPU
PP task will update and constrain coordinates on the GPU
PME tasks will do all aspects on the GPU- stat: compatible: the GPU was recognized and its architecture is compiled in. "not in set of targeted devices" means you need to rebuild for that architecture.
- The Mapping of GPU IDs line lists both PP and PME: both task types run on the GPU. With PP only, PME is computed on the CPU.
- PP task will update and constrain coordinates on the GPU: GPU-resident mode is active. CPU here means the system or settings do not qualify, for example virtual sites, mass or constraint free-energy perturbation, replica exchange, or constraints = all-bonds under domain decomposition.
- A NOTE containing "does not include kernels compiled natively": GPU kernels are JIT-compiled from PTX and may run slower; rebuild with the architecture value given in the log.
- ns/day on the Performance line: compare with a run using -nb cpu -pme cpu -update cpu; on the same system the GPU run should be much faster, and similar numbers mean the GPU did not take part.
- Tested on this page (CPU reference, not GPU speed): the script above with the conda-forge GROMACS 2026.3 CPU build on an Apple M2 (8 cores, other jobs running on the machine at the same time). gmx solvate produced 4,055 water molecules and 12,165 atoms; energy minimization reached Fmax < 1000 in 12 steps; mdrun with -nb gpu stopped with Nonbonded interactions on the GPU were requested with -nb gpu, but the GROMACS binary has been built without GPU support.; with -nb cpu -pme cpu -update cpu -ntmpi 1 the 50 ps run reached 14.4 ns/day under heavy load and 27.8 ns/day at lower load (1-minute load average about 13, -nsteps 10000 -resethway), and md.log showed Using SIMD4xM 4x4 nonbonded short-range kernels with no Mapping of GPU IDs or PP task lines. A GPU build should be much faster on the same system; similar numbers mean the GPU is not doing the work.
Run options
In GPU-resident mode (-update gpu), the GROMACS team recommends setting nstcalcenergy and the temperature and pressure coupling intervals to at least 50–100 steps, because the overhead of frequent virial and energy computation does not show up in the cycle counters of the log. Use constraints = h-bonds as well: it is one of the prerequisites for GPU-resident mode and matches how most force fields were parametrized.
When running several simulations on one GPU without -update gpu and without MPS or MIG, NVIDIA's tests showed no throughput gain. One A100 ran the 24,000-atom RNAse system at 1,083 ns/day and the 96,000-atom ADH system at 378 ns/day (GROMACS 2021.2), which can serve as a reference when estimating the speed of a comparable GPU.
| Scenario | Recommended command | Basis |
|---|---|---|
| One GPU, one simulation | gmx mdrun -s md.tpr -nb gpu -pme gpu -update gpu -ntmpi 1 -ntomp <physical cores> | With a GPU the default is already one rank per GPU; writing it out makes failures visible |
| Strong CPU, weaker GPU | gmx mdrun -s md.tpr -ntmpi 4 -nb gpu -pme cpu | Official example: long-range PME stays on the CPU and bonded work is assigned to the GPU automatically |
| Several GPUs in one node | gmx mdrun -s md.tpr -ntmpi 4 -nb gpu -pme gpu -npme 1 -update gpu | With PME on a GPU, -npme must be 1; use 1–3 ranks per GPU, and 1 rank per GPU is best with GPU-resident mode and direct GPU communication |
| Several small systems on one GPU | Each simulation with -update gpu, CUDA MPS enabled, and CPU cores separated with -pin on -pinoffset | NVIDIA measured on an A100 a 1.8× total throughput gain for a 24,000-atom system and 1.3× for 96,000 atoms; without -update gpu performance was about half |
| Node shared with other users | gmx mdrun -s md.tpr -gpu_id 1 -pin on -pinoffset 0 -nt 8 | -gpu_id restricts which GPUs are used; -pin keeps threads from competing for cores |
Experience
The following comes from Sobereva's blog (Computational Chemistry Commune) and the GROMACS forum; most points can be checked against first-hand documentation.
Do not build an MPI version for single-node runs
Common practice: on a single machine the default thread-MPI plus OpenMP needs one fewer installation step than an MPI build and is also more efficient; build with -DGMX_MPI=ON, which produces gmx_mpi, only for multi-node runs. The official installation guide likewise says multi-core runs on one workstation need no MPI setup.
Slow FFTW download in China: give a local path
Common practice: GMX_BUILD_OWN_FFTW often stalls when fetching from fftw.org inside mainland China. Download fftw-3.3.10.tar.gz with a browser first and pass it to CMake with -DGMX_BUILD_OWN_FFTW_URL=<absolute path>; the MD5 is checked automatically.
make -j hangs inside a virtual machine
Common practice: parallel builds inside VMware and similar VMs occasionally hang or fail; drop -j or use -j 4 and run make again. Under WSL2 the corresponding symptom is usually cc1plus being killed because memory ran out.
Rebuild when the log shows plain-C kernels or Compiled SIMD is None
If md.log shows "Using plain-C-4x4 4x4 nonbonded short-range kernels", or start-up prints "Compiled SIMD is None, but AVX2_256 might be faster (see log).", SIMD is not in use and the run will be several times slower; a normal build shows a SIMD kernel such as "Using SIMD4xM 4x4 nonbonded short-range kernels" (texts from the GROMACS 2026.4 source; the SIMD4xM line was tested on this page). "Using the slow plain C kernels" in older tutorials is the wording of old versions. Re-run cmake with an explicit -DGMX_SIMD=AVX2_256 (or the highest instruction set your CPU supports). If an old gcc fails on AVX-512, falling back to AVX2_256 also works.
Install a separate double-precision build
Common practice: energy minimization and Hessian diagonalization for normal mode analysis need double precision, so build a second copy with -DGMX_DOUBLE=ON. The program is named gmx_d and can sit in the same directory as the single-precision build. The double-precision build has no GPU support, runs at about half the speed, and writes trajectory and energy files twice as large.
Same RTX 4090, 2× speed difference between hosts
GROMACS forum measurements (version 2023, -nb gpu -bonded gpu -update gpu -ntomp 12): a ~30,000-atom system ran at about 900 and 1500 ns/day on two RTX 4090 hosts, and a ~100,000-atom system at 170 and 360 ns/day. A developer advised comparing with identical inputs and commands (for example -nsteps 100000 -ntmpi 1 -pin on -nb gpu -pme gpu -update gpu -bonded gpu) and then inspecting GPU kernel times with Nsight Systems. Treat these numbers as a rough range for that GPU generation when estimating your own machine's speed.
Sobereva's prebuilt native Windows binaries stop at 2020.6
Sobereva provides Windows 64-bit builds he compiled himself at sobereva.com/458: 2018.8 CPU, 2019.6 and 2020.3 CUDA (AVX), and 2020.6 CUDA (AVX2, NVIDIA driver ≥ 471.11). To use them, unzip and add the bin directory to Path; on machines without Visual Studio 2019, install VC_redist.x64 first; the CUDA builds need no CUDA toolkit, only a recent enough driver. They are command-line programs, and double-clicking gmx.exe just closes the window. These versions lack gmx dssp (2023+), the new gmx hbond (2024+) and the bundled amber19sb.ff (2026+), so they suit preparing input files and practice on Windows; for production runs with the current version, build under WSL2 as described above.
Widely cited Chinese install guides: what is outdated
Jerkwin's "GROMACS program compilation" page is marked "outdated and no longer updated" at the top; its example is 2016.4 and uses -DGMX_GPU=on, -DFFTWF_INCLUDE_DIR and similar options. His Chinese GROMACS manual still contains chapters on implicit solvent (removed in 2019) and the group cutoff scheme (removed in 2020). Sobereva's "How to install GROMACS" was last updated in May 2026 and its gcc and cmake requirements are current, so it is a usable reference; its CUDA part still uses -DCUDA_TOOLKIT_ROOT_DIR, which GROMACS 2026 no longer reads, so use -DCUDAToolkit_ROOT instead (see the version table above). The same article notes that a new gcc can also fail on an old GROMACS: gcc 11.2.1 on Rocky Linux 9 cannot build 2018.8.
PLUMED
Since 2025 the GROMACS source includes the PLUMED 2.10 interface, which the GROMACS documentation describes as compatible with any PLUMED version. It is enabled by default on non-Windows systems (GMX_USE_PLUMED=AUTO); building only requires dlopen from the system, and PLUMED need not be installed beforehand. At run time point the PLUMED_KERNEL environment variable to the kernel library and add -plumed plumed.dat.
The native interface does not support the ENERGY collective variable, replica exchange or lambda dynamics; PLUMED exits with an error when more than 1 thread-MPI rank is used, so add -ntmpi 1 on a single node and build an MPI version for multiple ranks. In 2026.0–2026.3, MPI-parallel mdrun -plumed aborts at start-up; 2026.4 fixes this.
ENERGY, multiple walkers or OPES multithermal still require the patch. Patches are provided for specific GROMACS versions: PLUMED 2.10.1 (released July 2026) ships patches for gromacs-2022.5, 2023.5, 2024.3 and 2025.0; a gromacs-2026.0 patch currently exists only in PLUMED's development branches. To patch, run plumed patch -p in the GROMACS source root before running cmake; the PLUMED documentation recommends also setting -DGMX_THREAD_MPI=OFF -DGMX_MPI=ON.
The GROMACS documentation names the kernel library libPlumedKernel.so, while PLUMED installs it as libplumedKernel.so by default. Linux file names are case-sensitive, so copying the documentation path verbatim fails to find the file.
# 1. Build PLUMED (no ordering requirement relative to GROMACS)
wget https://github.com/plumed/plumed2/releases/download/v2.10.1/plumed-2.10.1.tgz
tar xzf plumed-2.10.1.tgz && cd plumed-2.10.1
./configure --prefix=$HOME/opt/plumed-2.10.1
make -j $(nproc) && make install
# 2. Force the interface on when building GROMACS 2025/2026 (AUTO only prints a STATUS message and disables the interface when dlopen is missing; ON stops with an error)
cmake .. -DGMX_GPU=CUDA -DGMX_USE_PLUMED=ON -DGMX_BUILD_OWN_FFTW=ON
gmx --version | grep -i plumed
# 3. Point to the kernel library at run time. The real file name is libplumedKernel.so (lowercase p)
export PLUMED_KERNEL=$HOME/opt/plumed-2.10.1/lib/libplumedKernel.so
gmx mdrun -s md.tpr -plumed plumed.dat -ntmpi 1Agent
Example instruction: "Prepare a GROMACS 2026 GPU environment on one GPU, run a 5 nm water box for 50 ps each with -update gpu and -update cpu, report the ns/day of both, and include the GPU mapping lines from md.log and the gmx --version output."
In an isolated cloud computer, the agent uses the preinstalled GROMACS 2026.3 GPU build, or compiles another version with the commands on this page when needed; rents a GPU for the task; generates topol.top, em.mdp, md.mdp and the tpr files; runs both simulations; extracts ns/day and the GPU mapping lines from the logs; and finally reviews the results adversarially to check that the GPU really took part. The workspace keeps scripts, parameter files, logs and results, and the task keeps running after you shut down your own computer.
You still need to check: that the GROMACS version stated in your paper's methods matches the version actually used; that md.log shows update and PME running where you expect; and the force field, water model and mdp settings of your production system, which are your decisions. A water box self-test does not replace equilibration checks on the production system.
References
- GROMACS 2026.4 Installation guide — Build requirements, CMake options, GPU backends, FFTW, Windows builds, make check guidance
- GROMACS 2026 Release notes: Miscellaneous — Renamed CUDA-related CMake variables
- GROMACS 2025 Release notes: Portability — Minimum CMake 3.28, GCC 11, CUDA 12.1
- GROMACS 2026.4 release notes — Fix for MPI-parallel mdrun -plumed aborting at start-up
- GROMACS Known issues — OpenMP_CUDA detection problem in Windows CUDA builds
- Getting good performance from mdrun — Semantics of -nb/-pme/-update/-gputasks, GPU-resident conditions, ranks per GPU
- GROMACS Environment variables — GMX_DISABLE_GPU_DETECTION; removed GMX_GPU_DD_COMMS
- Using PLUMED (GROMACS reference manual) — PLUMED_KERNEL, -plumed and native interface limitations
- CUDA on WSL User Guide (NVIDIA) — Driver on Windows only, toolkit metapackage only, WSL2 limitations
- CUDA Features Archive (NVIDIA) — CUDA 13.0 removed offline compilation for Maxwell, Pascal and Volta
- Working across file systems (Microsoft) — Keep files on the Linux file system
- Advanced settings configuration in WSL (Microsoft) — .wslconfig default memory is 50% of Windows memory
- conda-forge gromacs-feedstock — Platforms, CUDA dependency and SIMD variants of the conda builds
- NGC GROMACS container — Latest container tag and example command
- GROMACS forum: Installation Error on WSL2 with CUDA Toolkit — Link error with 2025.2 and CUDA 13, and workarounds
- GROMACS forum: CUDA compilation error (WSL2, RTX 3050) — native architecture and ptxas errors
- GROMACS forum: Make check failing when GPU enabled — Cause of GPU test timeouts and comparison with GMX_DISABLE_GPU_DETECTION
- GROMACS forum: State of native PLUMED interface — Features missing from the native interface
- PLUMED patch for gromacs-2025.0 — Features added by the patch and thread-MPI recommendation
- NVIDIA: Maximizing GROMACS Throughput with Multiple Simulations per GPU — ns/day on A100, multi-simulation throughput and the effect of -update gpu
- Sobereva: Installing GROMACS (Computational Chemistry Commune blog) — Experience post: MPI build choice, FFTW download, VM builds, SIMD and double-precision tips
- GROMACS forum: Low GROMACS performance on RTX 4090 — Experience post: ns/day measured on two RTX 4090 hosts and the developer's comparison method
- Sobereva: compiling and installing native Windows GROMACS (Chinese) — Versions, driver requirements and usage of the prebuilt Windows binaries
- Jerkwin: GROMACS program compilation (Chinese) — Marked outdated; 2016.4 example with old options
- Jerkwin: Chinese GROMACS manual — Marked outdated; contains implicit-solvent and group-scheme chapters
- GROMACS 2019 Release notes: Removed functionality — Implicit solvent removed in 2019
- GROMACS 2020 Release notes: Removed functionality — Group cutoff scheme removed in 2020
- Tsinghua TUNA mirror: Anaconda mirror help — custom_channels for conda-forge and mirrored_channels for mamba