GROMACS / Installation and build

Installing GROMACS: Windows (WSL2), Linux and Common GPU Build Problems

This page is checked against the official installation guide and source code of GROMACS 2026.4 (released October 2, 2026). It covers how to choose an installation route, ready-to-run CPU and CUDA build commands, GPU prerequisites under WSL2, how to judge make check failures, a lookup table of common errors, and the log lines that confirm the GPU is actually doing the work.

Short answer

On Windows, build from source inside Ubuntu 24.04 on WSL2: install the NVIDIA driver on the Windows side only, install only cuda-toolkit-12-x inside WSL, then run cmake .. -DGMX_GPU=CUDA -DCMAKE_CUDA_ARCHITECTURES=<compute capability> -DGMX_BUILD_OWN_FFTW=ON -DREGRESSIONTEST_DOWNLOAD=ON, followed by make, make check and make install. The Linux steps are the same. GROMACS 2026 requires CMake ≥ 3.28, GCC ≥ 11, CUDA ≥ 12.1 and a GPU with compute capability ≥ 5.0; GPUs below compute capability 7.5 (GTX 10 series, V100) can only be built with CUDA 12.x. After installing, run a water box with gmx mdrun -nb gpu -pme gpu -update gpu and check that md.log contains "Mapping of GPU IDs" and "update and constrain coordinates on the GPU".

Route

Version information in the table was checked on 2026-10-10. If you need GPU acceleration, only the first two routes give you the current release and your own choice of CUDA version.

RouteAvailable versionNVIDIA GPUWhen it fitsMain limitations
WSL2 + source buildAny version, currently 2026.4Yes (CUDA)Windows 10/11 PC with an NVIDIA GPUnvidia-smi is limited under WSL2; devices cannot be filtered by index on multi-GPU systems; sources must live on the Linux file system
Linux source buildAny version, currently 2026.4Yes (CUDA/SYCL/HIP)Workstations, servers, clustersNeeds CMake ≥ 3.28 and GCC ≥ 11; CentOS 7 (GCC 4.8.5) and Rocky/RHEL 8 (GCC 8.5) need a newer compiler first
conda-forge2026.3 (2026.4 is out upstream)CUDA builds on linux-64 onlyNo root access, want something running quicklyNo Windows builds; the CUDA build depends on cuda-toolkit 12.9, so the driver must support 12.9; x86 SIMD tops out at AVX2_256; regression tests are not run during packaging
Ubuntu apt2023.3 on 24.04, 2021.4 on 22.04NoTeaching or small CPU-only systemsTwo to three years behind; no CUDA runtime among the dependencies
NVIDIA NGC containerLatest tag 2023.2 (updated 2023-08)YesGPU servers that already run Docker/SingularityNo longer updated; GMX_ENABLE_DIRECT_GPU_COMM in the example command is a 2022/2023 setting
Native Windows (MSVC)Build it yourselfYes (CUDA, tested upstream)Must run natively on WindowsNo official Windows binaries; GMX_BUILD_OWN_FFTW does not work, so use MKL, fftpack or your own FFTW build; no PLUMED; CMake < 4.3 fails to detect OpenMP_CUDA for CUDA builds
Sources: GROMACS 2026.4 installation guide and Known issues; conda-forge gromacs-feedstock recipe; packages.ubuntu.com; NGC GROMACS container page; NVIDIA CUDA on WSL User Guide.

Version changes

Most tutorials online use versions 2018–2021. The table lists requirements and syntax that have changed in 2026.x.

ItemGROMACS 2026.xOld tutorial syntaxWhat happens if you copy it
CMake≥ 3.28 (since 2025)cmake3, 3.13–3.17Configuration fails immediately; apt on Ubuntu 22.04 gives 3.22.1 (too old), 24.04 gives 3.28.3 (enough)
C++ compilerGCC ≥ 11, Clang ≥ 14, MSVC 2019; the classic Intel compiler icc is no longer supportedGCC 5–7The system GCC on CentOS 7 and Rocky 8 cannot build it
Enable CUDA-DGMX_GPU=CUDA-DGMX_GPU=ONError: Invalid value for GMX_GPU: ON. Pick one of: OFF, CUDA, OpenCL, SYCL, HIP
CUDA path-DCUDAToolkit_ROOT=/usr/local/cuda-12.9-DCUDA_TOOLKIT_ROOT_DIR=…The GROMACS 2026 CMake scripts no longer read this variable; configuration only ends with Manually-specified variables were not used by the project: CUDA_TOOLKIT_ROOT_DIR (tested on this page with CMake 4.4). With several CUDA versions installed another nvcc may be found, so check the actual version in the cmake output
GPU architecture-DCMAKE_CUDA_ARCHITECTURES=89-DGMX_CUDA_TARGET_SM=…Still accepted with a deprecation warning; 2026 switched to the standard CMake variable
CUDA version≥ 12.1, compute capability ≥ 5.0CUDA 9–11CUDA 11.x fails at configuration
CUDA 13Supported since 2025.3; CUDA 13 cannot build for GPUs below compute capability 7.5Not mentionedGTX 10 series, P100 and V100 cannot use the GPU after a CUDA 13 build; 2025.0–2025.2 with CUDA 13 fails at link time
AMD GPUsHIP backend with full offload since 2026, or SYCL + AdaptiveCppOpenCLOpenCL is deprecated and does not support RDNA AMD GPUs or NVIDIA GPUs from Volta onwards
GPU update and constraints-update auto runs on the GPU by default (since 2023)Add -update gpu by hand or set GMX_FORCE_UPDATE_DEFAULT_GPUThe old variables GMX_GPU_DD_COMMS and GMX_GPU_PME_PP_COMMS have been removed and have no effect
MPIMPI 3.0 required (since 2026)MPI 2.xSingle-node runs only need the built-in thread-MPI; no MPI build is required
PLUMEDInterface bundled since 2025, loaded at run timeAlways plumed patch firstRoutine use no longer needs a patch; see the PLUMED section
Sources: GROMACS 2026.4 installation guide; 2025 and 2026 release notes (Portability, Miscellaneous); 2026.4 source CMakeLists.txt and cmake/gmxManageCuda.cmake; NVIDIA CUDA Features Archive.

Linux

GMX_BUILD_OWN_FFTW=ON downloads FFTW 3.3.10 and builds it in single precision with the SIMD options GROMACS recommends. The GROMACS team advises against distribution FFTW packages because they default to double precision and are not compiled with options tuned for GROMACS.

Check two lines in the cmake output: whether the SIMD instruction set matches the CPU (for example AVX2_256 or AVX_512), and, for CUDA builds, "Compiling GROMACS for CUDA architectures: …". When you compile on a login node and run on compute nodes, set -DGMX_SIMD for the compute nodes; a binary built for AVX-512 exits with an illegal instruction on a machine that only supports AVX2. For GPU runs the GROMACS team recommends AVX2_256.

In a CUDA build, nvcc uses CMAKE_CXX_COMPILER as its host compiler. If the system GCC is newer than your CUDA supports, specify an older compiler for both, for example -DCMAKE_C_COMPILER=gcc-12 -DCMAKE_CXX_COMPILER=g++-12, instead of editing CUDA's host_config.h.

CPU build (single precision, thread-MPI)bash
# Ubuntu 24.04 (including Ubuntu 24.04 under WSL2); on 22.04 apt ships cmake 3.22, upgrade first with pip install cmake
sudo apt update
sudo apt install -y build-essential cmake wget perl

wget https://ftp.gromacs.org/gromacs/gromacs-2026.4.tar.gz
tar xfz gromacs-2026.4.tar.gz
cd gromacs-2026.4
mkdir build && cd build

cmake .. \
  -DGMX_BUILD_OWN_FFTW=ON \
  -DREGRESSIONTEST_DOWNLOAD=ON \
  -DCMAKE_INSTALL_PREFIX=$HOME/opt/gromacs-2026.4

make -j $(nproc)
make check
make install
echo 'source $HOME/opt/gromacs-2026.4/bin/GMXRC' >> ~/.bashrc
source $HOME/opt/gromacs-2026.4/bin/GMXRC
gmx --version
CUDA buildbash
# Prerequisites: nvcc works (nvcc --version) and nvidia-smi lists the GPU
export PATH=/usr/local/cuda-12.9/bin:$PATH

cd gromacs-2026.4
mkdir build-cuda && cd build-cuda

# CMAKE_CUDA_ARCHITECTURES is the compute capability without the dot:
# RTX 20/T4=75, A100=80, RTX 30/A40=86, RTX 40/L40=89, H100=90, RTX 50=120
# If omitted, every architecture from 50 to 120 that nvcc accepts is built: slower build, larger binary
cmake .. \
  -DGMX_GPU=CUDA \
  -DCUDAToolkit_ROOT=/usr/local/cuda-12.9 \
  -DCMAKE_CUDA_ARCHITECTURES=89 \
  -DGMX_BUILD_OWN_FFTW=ON \
  -DREGRESSIONTEST_DOWNLOAD=ON \
  -DCMAKE_INSTALL_PREFIX=$HOME/opt/gromacs-2026.4-cuda

make -j $(nproc)
make check
make install
source $HOME/opt/gromacs-2026.4-cuda/bin/GMXRC
gmx --version | grep -E "GPU support|CUDA targets|CUDA driver|CUDA runtime|SIMD"
When the server has no internet accessbash
# If the server cannot reach fftw.org or ftp.gromacs.org, download these two files elsewhere and copy them over
#   http://www.fftw.org/fftw-3.3.10.tar.gz
#   https://ftp.gromacs.org/regressiontests/regressiontests-2026.4.tar.gz
tar xfz regressiontests-2026.4.tar.gz -C $HOME/src

cmake .. \
  -DGMX_GPU=CUDA \
  -DGMX_BUILD_OWN_FFTW=ON \
  -DGMX_BUILD_OWN_FFTW_URL=$HOME/src/fftw-3.3.10.tar.gz \
  -DREGRESSIONTEST_PATH=$HOME/src/regressiontests-2026.4 \
  -DCMAKE_INSTALL_PREFIX=$HOME/opt/gromacs-2026.4-cuda

Windows

After these steps, build with the CUDA commands from the previous section. Known WSL2 limitations: nvidia-smi offers a limited feature set; GPUs cannot be filtered by index number on multi-GPU machines; concurrent CPU/GPU access to the same memory is not supported. Single-GPU GROMACS runs are not affected by these limitations.

Windows sidepowershell
# Windows PowerShell (administrator)
wsl --install -d Ubuntu-24.04
wsl --update

# Optional: %UserProfile%\.wslconfig; run wsl --shutdown after editing, then reopen WSL
# [wsl2]
# memory=16GB
# processors=8
Installing the CUDA toolkit inside WSLbash
# Run inside Ubuntu on WSL; first confirm the Windows driver is passed through
nvidia-smi

wget https://developer.download.nvidia.com/compute/cuda/repos/wsl-ubuntu/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt-get update
# Install only the toolkit metapackage; cuda, cuda-12-9 and cuda-drivers pull in a Linux driver
sudo apt-get install -y cuda-toolkit-12-9

echo 'export PATH=/usr/local/cuda-12.9/bin:$PATH' >> ~/.bashrc
source ~/.bashrc
nvcc --version

# Keep sources and build directories on the Linux file system, not under /mnt/c
mkdir -p ~/src && cd ~/src
  1. 01

    Install or update the NVIDIA driver on Windows

    NVIDIA's WSL guide states this is the only driver you need, and that no Linux display driver should be installed inside WSL. The Windows driver is exposed inside WSL2 as libcuda.so, and nvidia-smi lives in /usr/lib/wsl/lib. The CUDA Version shown at the top right of nvidia-smi inside WSL should be at least the toolkit version you plan to install.

  2. 02

    Install WSL2 and Ubuntu 24.04

    CUDA is supported on WSL 2 only, not WSL 1. Ubuntu 24.04 ships CMake 3.28.3 and GCC 13.2 via apt, which meet GROMACS 2026's requirements, so you can skip a manual CMake upgrade.

  3. 03

    Install only the cuda-toolkit-12-x metapackage

    Use NVIDIA's wsl-ubuntu repository. The cuda, cuda-12-x and cuda-drivers metapackages install a Linux driver that overrides the WSL libcuda mapping and leaves the GPU unusable. CUDA 12.9 can build for compute capability 5.0 through 12.0, covering the RTX 50 series as well as the GTX 10 series.

  4. 04

    Keep sources and build directories under ~

    Microsoft recommends storing files under /home/<user> when working from a Linux command line, not under /mnt/c. Building under /mnt/c is slow, and a GROMACS forum user building there got "Clock skew detected" warnings.

  5. 05

    Lower the parallelism if the build runs out of memory

    By default WSL2 gets only 50% of Windows memory. If make -j stops with "c++: fatal error: Killed signal terminated program cc1plus", memory ran out; lower the -j value, or raise memory in .wslconfig and run wsl --shutdown.

conda

conda-forge CUDA buildbash
# linux-64 only (including WSL2). The CUDA build needs a driver that supports CUDA 12.9
conda create -n gmx -c conda-forge "gromacs=2026.3=nompi_cuda*"
conda activate gmx
gmx --version | grep -E "GPU support|SIMD"

# CPU/OpenCL build
# conda create -n gmx-cpu -c conda-forge "gromacs=2026.3=nompi_h*"
Networks in China: TUNA mirrors for conda-forge and pipbash
# ~/.condarc (conda-forge setup from the TUNA help page; run conda clean -i afterwards)
# channels:
#   - conda-forge
#   - nodefaults
# show_channel_urls: true
# custom_channels:
#   conda-forge: https://mirrors.tuna.tsinghua.edu.cn/anaconda/cloud
conda clean -i
conda create -n gmx -c conda-forge "gromacs=2026.3=nompi_cuda*"

# With mamba, use mirrored_channels instead
# channels:
#   - conda-forge
# mirrored_channels:
#   conda-forge:
#     - https://mirrors.tuna.tsinghua.edu.cn/anaconda/cloud/conda-forge

# For source builds, upgrade CMake with pip (3.28 or newer is required)
pip install -i https://pypi.tuna.tsinghua.edu.cn/simple "cmake>=3.28"

One patch release behind upstream

On the check date conda-forge had 2026.3 while upstream had 2026.4. The bioconda gromacs package stopped at 2021.3, so do not install from bioconda.

CUDA builds exist for linux-64 only

aarch64, ppc64le and macOS have no CUDA builds, and there are no Windows builds at all. The CUDA build installs cuda-toolkit 12.9 into the environment and checks the driver through the __cuda virtual package.

SIMD tops out at AVX2_256

On x86 the recipe builds only SSE2, AVX_256 and AVX2_256 variants and picks one at start-up based on the CPU. On CPUs with AVX-512, CPU-only runs may be slower than a self-compiled build.

Regression tests are not run during packaging

REGRESSIONTEST_DOWNLOAD is commented out in the recipe, and the package test only runs gmx -version. If you need trustworthy results, download regressiontests-2026.3 yourself and run gmxtest.pl.

Testing

The official installation guide's rule: an individual failed test may be a compiler bug, or a tolerance that is slightly too tight. Inspect the output files the script points to; if the failure is real, rebuild with a newer compiler. If the tests still fail, post on the GROMACS forum with a description of the hardware and the full output of gmx mdrun -version.

The most common case on the GROMACS forum is a GPU build where many tests report Timeout. A developer explained that tests use every core and GPU they detect by default, so they time out when the CPU or GPU is busy with other jobs. In one 2021.1 case, 23 tests timed out with GPU detection enabled and only 1 remained with GMX_DISABLE_GPU_DETECTION=1. If the failures are only timeouts and the CPU-only run passes, confirm the GPU works with the water box self-test in the next section; numerical comparison failures (output shows reference values that do not match the computed ones) should not be ignored.

Locating failed testsbash
# Show only the failed tests
make check 2>&1 | tee check.log
grep -A40 "tests FAILED" check.log

# 1. Rule out the GPU: if the CPU run passes and the GPU run times out, the GPU is usually busy
GMX_DISABLE_GPU_DETECTION=1 make check

# 2. On multi-GPU machines, expose a single GPU to the tests
CUDA_VISIBLE_DEVICES=0 make check

# 3. Timeouts on slow machines or under WSL: raise the timeout factor and reconfigure
cmake .. -DGMX_TEST_TIMEOUT_FACTOR=3 && make check

# 4. Rerun a single test with full output
ctest -R MdrunTests --output-on-failure

Errors

Error text (excerpt)CauseFix
Invalid value for GMX_GPU: ONSyntax from before 2021Use -DGMX_GPU=CUDA
Could not find `nvcc` executable in path specified by variable CUDAToolkit_ROOT=/usr/local/cuda-12.9There is no bin/nvcc under the directory given in CUDAToolkit_ROOT (CMake text tested on this page on a machine without CUDA)Check the actual CUDA directory with ls /usr/local/ and correct CUDAToolkit_ROOT
CMake 3.28 or higher is requiredSystem CMake is too old (3.22.1 on Ubuntu 22.04, older on CentOS 7)pip install cmake, or use Kitware's official binaries
#error -- unsupported GNU version! gcc versions later than … are not supported!GCC is newer than the host compilers your CUDA supportsInstall a GCC your CUDA supports and select it with -DCMAKE_C_COMPILER/-DCMAKE_CXX_COMPILER, or upgrade CUDA
nvcc fatal : Unsupported gpu architecture 'compute_xx'CMAKE_CUDA_ARCHITECTURES or the old GMX_CUDA_TARGET_SM lists an architecture the current CUDA does not support (for example 61 with CUDA 13)Use the GPU's actual compute capability; for GPUs below compute capability 7.5 switch to CUDA 12.x
undefined reference to `void gmx::nbnxn_kernel_prune_cuda<true>(…)'GROMACS 2025.0–2025.2 with CUDA 13Upgrade to 2025.3 or later, use CUDA 12.9, or add -DCMAKE_CUDA_FLAGS=-static-global-template-stub=false
Unknown CUDA architecture: native / ptxas error : Value of threads per SM … is out of rangeOld CMake or outdated sourceUpdate the source, delete the build directory, reconfigure, and set an explicit value such as -DCMAKE_CUDA_ARCHITECTURES=86-real
Cannot find FFTW 3 (with correct precision - libfftw3f for mixed-precision GROMACS …)Only double-precision FFTW is installed, or CMAKE_PREFIX_PATH does not point to FFTWUse -DGMX_BUILD_OWN_FFTW=ON; for offline builds add -DGMX_BUILD_OWN_FFTW_URL=<absolute path>
Cannot build FFTW3 automatically (GMX_BUILD_OWN_FFTW=ON) in Visual Studio / with ninjaAutomatic FFTW builds are not supported for native Windows builds or the Ninja generatorUse -DGMX_FFT_LIBRARY=mkl, or build FFTW yourself first
WARNING: The gmx binary does not include support for the CUDA architecture of the GPU ID #0The CMAKE_CUDA_ARCHITECTURES used at build time does not include this GPURebuild with the value suggested in the log, for example -DCMAKE_CUDA_ARCHITECTURES=89
Cannot run short-ranged nonbonded interactions on a GPU because no GPU is detected.The binary has GPU support, but no usable GPU is found at run time: the driver is unavailable, or a Linux driver was installed inside WSLCheck nvidia-smi; under WSL remove cuda-drivers-type packages
Nonbonded interactions on the GPU were requested with -nb gpu, but the GROMACS binary has been built without GPU support.The binary itself has no GPU support; the GPU support line of gmx --version says disabled (tested on this page with the conda-forge 2026.3 CPU build)Install a CUDA build or rebuild as described above
make stops with Killed signal terminated program cc1plusOut of memory during compilation, common on WSL2 and small VMsLower the -j value or add memory
mdrun -plumed aborts at start-up with MPI_ERR_COMM in MPI-parallel runsBug in 2026.0–2026.3Upgrade to 2026.4
Sources: error strings in the GROMACS 2026.4 source; GROMACS forum threads 12528, 13012 and 1837; 2025.3 and 2026.4 release notes.

Verification

First run gmx --version and check that the GPU support line says CUDA, CUDA targets includes your GPU's architecture, and CUDA driver is not older than CUDA runtime. Builds with the PLUMED interface show its status on the Plumed support line.

Then run a water box. The -nb, -pme and -update options of mdrun all default to auto: if no usable GPU is detected, mdrun silently falls back to the CPU and only runs slower. Request gpu explicitly for the self-test so that an unusable GPU stops the run with an error.

Water box GPU self-test (about 12,000 atoms, 50 ps)bash
mkdir -p ~/gmx-gpu-test && cd ~/gmx-gpu-test

# 5 nm cubic water box, about 4000 SPC/E waters
gmx solvate -cs spc216.gro -box 5 5 5 -o water.gro
cat > topol.top <<'EOF'
#include "oplsaa.ff/forcefield.itp"
#include "oplsaa.ff/spce.itp"

[ system ]
SPC/E water box

[ molecules ]
EOF
echo "SOL $(grep -c OW water.gro)" >> topol.top

cat > em.mdp <<'EOF'
integrator    = steep
emtol         = 1000
nsteps        = 5000
cutoff-scheme = Verlet
coulombtype   = PME
rcoulomb      = 1.0
rvdw          = 1.0
EOF

cat > md.mdp <<'EOF'
integrator     = md
dt             = 0.002
nsteps         = 25000
cutoff-scheme  = Verlet
coulombtype    = PME
rcoulomb       = 1.0
rvdw           = 1.0
tcoupl         = V-rescale
tc-grps        = System
tau-t          = 0.1
ref-t          = 300
gen-vel        = yes
gen-temp       = 300
constraints    = h-bonds
nstcalcenergy  = 100
nstenergy      = 1000
nstlog         = 5000
EOF

gmx grompp -f em.mdp -c water.gro -p topol.top -o em.tpr
gmx mdrun -s em.tpr -c em.gro -g em.log
gmx grompp -f md.mdp -c em.gro -p topol.top -o md.tpr

# Request gpu explicitly: if no GPU is usable, mdrun stops with an error instead of silently falling back to the CPU
gmx mdrun -s md.tpr -g md.log -nb gpu -pme gpu -update gpu -ntmpi 1

grep -A2 "Mapping of GPU IDs" md.log
grep -E "PP task|PME tasks" md.log
grep "Performance:" md.log
Lines expected in md.log (single-GPU example)text
  GPU info:
    Number of GPUs detected: 1
    #0: NVIDIA NVIDIA GeForce RTX 4090, compute cap.: 8.9, ECC:  no, stat: compatible

1 GPU selected for this run.
Mapping of GPU IDs to the 2 GPU tasks in the 1 rank on this node:
  PP:0,PME:0
PP tasks will do non-perturbed short-ranged interactions on the GPU
PP task will update and constrain coordinates on the GPU
PME tasks will do all aspects on the GPU
  • stat: compatible: the GPU was recognized and its architecture is compiled in. "not in set of targeted devices" means you need to rebuild for that architecture.
  • The Mapping of GPU IDs line lists both PP and PME: both task types run on the GPU. With PP only, PME is computed on the CPU.
  • PP task will update and constrain coordinates on the GPU: GPU-resident mode is active. CPU here means the system or settings do not qualify, for example virtual sites, mass or constraint free-energy perturbation, replica exchange, or constraints = all-bonds under domain decomposition.
  • A NOTE containing "does not include kernels compiled natively": GPU kernels are JIT-compiled from PTX and may run slower; rebuild with the architecture value given in the log.
  • ns/day on the Performance line: compare with a run using -nb cpu -pme cpu -update cpu; on the same system the GPU run should be much faster, and similar numbers mean the GPU did not take part.
  • Tested on this page (CPU reference, not GPU speed): the script above with the conda-forge GROMACS 2026.3 CPU build on an Apple M2 (8 cores, other jobs running on the machine at the same time). gmx solvate produced 4,055 water molecules and 12,165 atoms; energy minimization reached Fmax < 1000 in 12 steps; mdrun with -nb gpu stopped with Nonbonded interactions on the GPU were requested with -nb gpu, but the GROMACS binary has been built without GPU support.; with -nb cpu -pme cpu -update cpu -ntmpi 1 the 50 ps run reached 14.4 ns/day under heavy load and 27.8 ns/day at lower load (1-minute load average about 13, -nsteps 10000 -resethway), and md.log showed Using SIMD4xM 4x4 nonbonded short-range kernels with no Mapping of GPU IDs or PP task lines. A GPU build should be much faster on the same system; similar numbers mean the GPU is not doing the work.

Run options

In GPU-resident mode (-update gpu), the GROMACS team recommends setting nstcalcenergy and the temperature and pressure coupling intervals to at least 50–100 steps, because the overhead of frequent virial and energy computation does not show up in the cycle counters of the log. Use constraints = h-bonds as well: it is one of the prerequisites for GPU-resident mode and matches how most force fields were parametrized.

When running several simulations on one GPU without -update gpu and without MPS or MIG, NVIDIA's tests showed no throughput gain. One A100 ran the 24,000-atom RNAse system at 1,083 ns/day and the 96,000-atom ADH system at 378 ns/day (GROMACS 2021.2), which can serve as a reference when estimating the speed of a comparable GPU.

ScenarioRecommended commandBasis
One GPU, one simulationgmx mdrun -s md.tpr -nb gpu -pme gpu -update gpu -ntmpi 1 -ntomp <physical cores>With a GPU the default is already one rank per GPU; writing it out makes failures visible
Strong CPU, weaker GPUgmx mdrun -s md.tpr -ntmpi 4 -nb gpu -pme cpuOfficial example: long-range PME stays on the CPU and bonded work is assigned to the GPU automatically
Several GPUs in one nodegmx mdrun -s md.tpr -ntmpi 4 -nb gpu -pme gpu -npme 1 -update gpuWith PME on a GPU, -npme must be 1; use 1–3 ranks per GPU, and 1 rank per GPU is best with GPU-resident mode and direct GPU communication
Several small systems on one GPUEach simulation with -update gpu, CUDA MPS enabled, and CPU cores separated with -pin on -pinoffsetNVIDIA measured on an A100 a 1.8× total throughput gain for a 24,000-atom system and 1.3× for 96,000 atoms; without -update gpu performance was about half
Node shared with other usersgmx mdrun -s md.tpr -gpu_id 1 -pin on -pinoffset 0 -nt 8-gpu_id restricts which GPUs are used; -pin keeps threads from competing for cores
Sources: GROMACS user guide, Getting good performance from mdrun; NVIDIA technical blog, Maximizing GROMACS Throughput with Multiple Simulations per GPU Using MPS and MIG (GROMACS 2021.2, DGX A100).

Experience

The following comes from Sobereva's blog (Computational Chemistry Commune) and the GROMACS forum; most points can be checked against first-hand documentation.

Do not build an MPI version for single-node runs

Common practice: on a single machine the default thread-MPI plus OpenMP needs one fewer installation step than an MPI build and is also more efficient; build with -DGMX_MPI=ON, which produces gmx_mpi, only for multi-node runs. The official installation guide likewise says multi-core runs on one workstation need no MPI setup.

Slow FFTW download in China: give a local path

Common practice: GMX_BUILD_OWN_FFTW often stalls when fetching from fftw.org inside mainland China. Download fftw-3.3.10.tar.gz with a browser first and pass it to CMake with -DGMX_BUILD_OWN_FFTW_URL=<absolute path>; the MD5 is checked automatically.

make -j hangs inside a virtual machine

Common practice: parallel builds inside VMware and similar VMs occasionally hang or fail; drop -j or use -j 4 and run make again. Under WSL2 the corresponding symptom is usually cc1plus being killed because memory ran out.

Rebuild when the log shows plain-C kernels or Compiled SIMD is None

If md.log shows "Using plain-C-4x4 4x4 nonbonded short-range kernels", or start-up prints "Compiled SIMD is None, but AVX2_256 might be faster (see log).", SIMD is not in use and the run will be several times slower; a normal build shows a SIMD kernel such as "Using SIMD4xM 4x4 nonbonded short-range kernels" (texts from the GROMACS 2026.4 source; the SIMD4xM line was tested on this page). "Using the slow plain C kernels" in older tutorials is the wording of old versions. Re-run cmake with an explicit -DGMX_SIMD=AVX2_256 (or the highest instruction set your CPU supports). If an old gcc fails on AVX-512, falling back to AVX2_256 also works.

Install a separate double-precision build

Common practice: energy minimization and Hessian diagonalization for normal mode analysis need double precision, so build a second copy with -DGMX_DOUBLE=ON. The program is named gmx_d and can sit in the same directory as the single-precision build. The double-precision build has no GPU support, runs at about half the speed, and writes trajectory and energy files twice as large.

Same RTX 4090, 2× speed difference between hosts

GROMACS forum measurements (version 2023, -nb gpu -bonded gpu -update gpu -ntomp 12): a ~30,000-atom system ran at about 900 and 1500 ns/day on two RTX 4090 hosts, and a ~100,000-atom system at 170 and 360 ns/day. A developer advised comparing with identical inputs and commands (for example -nsteps 100000 -ntmpi 1 -pin on -nb gpu -pme gpu -update gpu -bonded gpu) and then inspecting GPU kernel times with Nsight Systems. Treat these numbers as a rough range for that GPU generation when estimating your own machine's speed.

Sobereva's prebuilt native Windows binaries stop at 2020.6

Sobereva provides Windows 64-bit builds he compiled himself at sobereva.com/458: 2018.8 CPU, 2019.6 and 2020.3 CUDA (AVX), and 2020.6 CUDA (AVX2, NVIDIA driver ≥ 471.11). To use them, unzip and add the bin directory to Path; on machines without Visual Studio 2019, install VC_redist.x64 first; the CUDA builds need no CUDA toolkit, only a recent enough driver. They are command-line programs, and double-clicking gmx.exe just closes the window. These versions lack gmx dssp (2023+), the new gmx hbond (2024+) and the bundled amber19sb.ff (2026+), so they suit preparing input files and practice on Windows; for production runs with the current version, build under WSL2 as described above.

Widely cited Chinese install guides: what is outdated

Jerkwin's "GROMACS program compilation" page is marked "outdated and no longer updated" at the top; its example is 2016.4 and uses -DGMX_GPU=on, -DFFTWF_INCLUDE_DIR and similar options. His Chinese GROMACS manual still contains chapters on implicit solvent (removed in 2019) and the group cutoff scheme (removed in 2020). Sobereva's "How to install GROMACS" was last updated in May 2026 and its gcc and cmake requirements are current, so it is a usable reference; its CUDA part still uses -DCUDA_TOOLKIT_ROOT_DIR, which GROMACS 2026 no longer reads, so use -DCUDAToolkit_ROOT instead (see the version table above). The same article notes that a new gcc can also fail on an old GROMACS: gcc 11.2.1 on Rocky Linux 9 cannot build 2018.8.

PLUMED

Since 2025 the GROMACS source includes the PLUMED 2.10 interface, which the GROMACS documentation describes as compatible with any PLUMED version. It is enabled by default on non-Windows systems (GMX_USE_PLUMED=AUTO); building only requires dlopen from the system, and PLUMED need not be installed beforehand. At run time point the PLUMED_KERNEL environment variable to the kernel library and add -plumed plumed.dat.

The native interface does not support the ENERGY collective variable, replica exchange or lambda dynamics; PLUMED exits with an error when more than 1 thread-MPI rank is used, so add -ntmpi 1 on a single node and build an MPI version for multiple ranks. In 2026.0–2026.3, MPI-parallel mdrun -plumed aborts at start-up; 2026.4 fixes this.

ENERGY, multiple walkers or OPES multithermal still require the patch. Patches are provided for specific GROMACS versions: PLUMED 2.10.1 (released July 2026) ships patches for gromacs-2022.5, 2023.5, 2024.3 and 2025.0; a gromacs-2026.0 patch currently exists only in PLUMED's development branches. To patch, run plumed patch -p in the GROMACS source root before running cmake; the PLUMED documentation recommends also setting -DGMX_THREAD_MPI=OFF -DGMX_MPI=ON.

The GROMACS documentation names the kernel library libPlumedKernel.so, while PLUMED installs it as libplumedKernel.so by default. Linux file names are case-sensitive, so copying the documentation path verbatim fails to find the file.

Using the native interfacebash
# 1. Build PLUMED (no ordering requirement relative to GROMACS)
wget https://github.com/plumed/plumed2/releases/download/v2.10.1/plumed-2.10.1.tgz
tar xzf plumed-2.10.1.tgz && cd plumed-2.10.1
./configure --prefix=$HOME/opt/plumed-2.10.1
make -j $(nproc) && make install

# 2. Force the interface on when building GROMACS 2025/2026 (AUTO only prints a STATUS message and disables the interface when dlopen is missing; ON stops with an error)
cmake .. -DGMX_GPU=CUDA -DGMX_USE_PLUMED=ON -DGMX_BUILD_OWN_FFTW=ON
gmx --version | grep -i plumed

# 3. Point to the kernel library at run time. The real file name is libplumedKernel.so (lowercase p)
export PLUMED_KERNEL=$HOME/opt/plumed-2.10.1/lib/libplumedKernel.so
gmx mdrun -s md.tpr -plumed plumed.dat -ntmpi 1

Agent

Example instruction: "Prepare a GROMACS 2026 GPU environment on one GPU, run a 5 nm water box for 50 ps each with -update gpu and -update cpu, report the ns/day of both, and include the GPU mapping lines from md.log and the gmx --version output."

In an isolated cloud computer, the agent uses the preinstalled GROMACS 2026.3 GPU build, or compiles another version with the commands on this page when needed; rents a GPU for the task; generates topol.top, em.mdp, md.mdp and the tpr files; runs both simulations; extracts ns/day and the GPU mapping lines from the logs; and finally reviews the results adversarially to check that the GPU really took part. The workspace keeps scripts, parameter files, logs and results, and the task keeps running after you shut down your own computer.

You still need to check: that the GROMACS version stated in your paper's methods matches the version actually used; that md.log shows update and PME running where you expect; and the force field, water model and mdp settings of your production system, which are your decisions. A water box self-test does not replace equilibration checks on the production system.

References

FAQ

Can I install GROMACS on Windows without WSL?

You can build natively with MSVC, and CUDA has been tested on Windows upstream, but there are no official Windows binaries. Native builds cannot use GMX_BUILD_OWN_FFTW and do not support PLUMED; with CMake below 4.3, CUDA builds also need -G Ninja or manually set OpenMP_CUDA flags. If you need a GPU and later extensions, the WSL2 route causes fewer problems.

Can a conda-installed GROMACS use the GPU?

conda-forge provides CUDA builds on linux-64 (including WSL2); select one with "gromacs=2026.3=nompi_cuda*". The driver must support CUDA 12.9. There are no CUDA builds for Windows or macOS, and the bioconda version stopped at 2021.3.

Can older GPUs such as the GTX 1080 or 1660 run the GPU build?

The GTX 1080 has compute capability 6.1 and works, but must be built with CUDA 12.1–12.9; CUDA 13 can no longer generate code for compute capability below 7.5. The GTX 1660 has compute capability 7.5, so both CUDA 12 and 13 work. GROMACS 2026 requires compute capability 5.0 or higher.

What should I do about installation errors on CentOS 7?

CentOS 7 ships GCC 4.8.5 and CMake 2.8, while GROMACS 2025 and later need GCC 11 and CMake 3.28; the GCC 8.5 on Rocky/RHEL 8 is also too old. Install a newer GCC (for example gcc-toolset) and CMake from pip before building, or switch to Ubuntu 24.04 or a container.

How do I confirm GROMACS is using the GPU?

Run gmx mdrun -nb gpu -pme gpu -update gpu and check md.log for "Mapping of GPU IDs … PP:0,PME:0" and "PP task will update and constrain coordinates on the GPU". With the default auto settings, a missing GPU produces no error, only a slower run.

A few make check tests failed. Can I still use the build?

If they are only timeouts and GMX_DISABLE_GPU_DETECTION=1 make check passes, the GPU or CPU was most likely busy during testing; rerun when the machine is idle or with a single GPU. For numerical comparison failures follow the official advice: rebuild with a newer compiler, and if they persist, ask on the forum with the output of gmx mdrun -version.

Hand the GROMACS environment and GPU to Scientify

Scientify's science agent runs in an isolated cloud computer. Its molecular simulation environment comes with GROMACS 2026.3 GPU build preinstalled and self-tested with a water box run on a GPU before delivery. GPUs are rented per task, billed per second and released when done; simulations keep running after you shut down your own computer, and parameter files and logs stay in the workspace. New users get $5 of free credit.