LAMMPS / Installation and getting started

LAMMPS Installation and Getting Started: From Choosing an Install Route to Your First Tensile Simulation

This page reflects the state of LAMMPS in October 2026. It covers how to choose an install route, how to build with CMake and for GPUs, which command-line acceleration flags to use, and a line-by-line Cu uniaxial tension input file with a post-processing script. The commands work with both the 22Jul2025 and 30Sep2026 stable releases, and the differences between them are listed separately.

Short answer

If you only want to learn input files on Windows, the official LAMMPS-GUI installer is enough. It includes the command-line lmp and the GPU package in OpenCL mode, but no MPI. For multi-core parallel runs or CUDA acceleration, build from source with CMake on Linux or WSL2: -C ../cmake/presets/basic.cmake enables MANYBODY, MOLECULE, KSPACE and RIGID. For the GPU package set -D GPU_API=cuda and GPU_ARCH=sm_XX explicitly, and run with -sf gpu -pk gpu 1. For KOKKOS, run with -k on g 1 -sf kk. The new stable release of 2026-09-30 requires CMake 3.20 and C++17, and KOKKOS also needs GCC 10.4 and CUDA 12.2; older systems can stay on stable_22Jul2025_update6.

Step 1

Choose a route by three questions: do you need MPI parallel runs, do you need CUDA, and do you have root access. The typical cost of a wrong choice is finding out after installation that you cannot run in parallel, or that the GPU is not used at all.

RouteWhen to use itMPIGPUNotes
Windows LAMMPS-GUI installer (GitHub release, LAMMPS-Win10-x64-GUI-30Sep2026.exe)Learning and debugging input files, small systemsNoGPU package, OpenCL mixed precisionIncludes command-line lmp; no KIM and some other packages; SmartScreen shows a warning during installation
Windows MS-MPI installer (rpm.lammps.org/windows, LAMMPS-64bit-30Sep2026-MSMPI.exe)Multi-process parallel runs on WindowsMS-MPI only; install msmpisetup.exe separatelyGPU package, OpenCLRun with mpiexec -n 4 lmp -in ...; the installer sets PATH and LAMMPS_POTENTIALS; MS-MPI only since 2Aug2023
WSL2 + CMake source buildWindows machine that needs CUDA or MPIOpenMPIGPU package or KOKKOS (CUDA)Install no Linux display driver inside WSL; keep files on the Linux file system
conda-forge (conda install lammps)Linux/macOS without compilingopenmpi/mpich/nompi variantsCUDA variants use KOKKOS, no GPU packageVersion is 2025.07.22; no Windows build; includes KIM; ships no potentials/ folder and does not set LAMMPS_POTENTIALS (tested with the osx-arm64 package on this page)
Linux static binary (download.lammps.org/static)Cluster login nodes, no root, just need it to runNoNoneNo dependence on system libraries; no Python module
apt / dnf distribution packagesQuick trialYesNoneapt version is 20240207 on Ubuntu 24.04 and 20220106 on 22.04
CMake source build (Linux, clusters)Production runs, custom packages or GPUYesGPU package or KOKKOSMain route of this page; can install into $HOME without root
Docker/Apptainer (NVIDIA NGC nvcr.io/hpc/lammps)GPU machines with the NVIDIA Container ToolkitBuilt into the imageYes, CUDARun lmp -h inside the container first to see whether the accelerator is the GPU package or KOKKOS
Sources: LAMMPS manual Install pages, LAMMPS-GUI installation docs, packages.lammps.org/windows.html, conda-forge lammps-feedstock, packages.ubuntu.com.

Version

The new stable release of 30 September 2026 raised the toolchain requirements. On clusters with older GCC or CMake, building the new version fails at the configure step.

How to decide: run gcc --version and cmake --version first. If GCC is older than 10.4 and you need KOKKOS, use stable_22Jul2025_update6 or install a newer compiler first. If CMake is too old, python3 -m pip install --user cmake installs a recent version without root.

If your input must be compared with results from a supervisor or a published study, record the full version string printed by lmp -h. 30Sep2026 fixed platform-dependent results on ARM64 and on platforms with FMA enabled, so the two versions can differ in the last digits.

Itemstable_22Jul2025_update6stable_30Sep2026
Minimum CMake3.163.20
C++ standardC++11 by default; KOKKOS needs C++17C++17; KOKKOS needs C++20
Bundled Kokkos4.6.2 (GCC ≥ 8.2, CUDA ≥ 11.0)5.2 (GCC ≥ 10.4, NVCC ≥ 12.2)
GPU package default GPU_ARCHsm_50sm_75
Traditional make buildWorks for KOKKOSKOKKOS is CMake-only since 11Feb2026; make only builds packages that need no extra steps in lib/
dump image / dump movieCore featureMoved to the new GRAPHICS package, included in the basic.cmake preset
Removed packages—ATC, AWPMD, POEMS, ML-RANN, VTK; pair agni (dump vtk moved to EXTRA-DUMP)
Sources: stable_30Sep2026 release notes; cmake/CMakeLists.txt, lib/kokkos/CMakeLists.txt and Build_extras at both tags; Kokkos requirements page.

Windows

The LAMMPS-GUI installer suits learning: the editor runs the current input file directly and shows thermo plots and snapshots. It has three limits: no MPI parallel runs; GPU support only through the OpenCL GPU package, which needs a graphics driver with an OpenCL runtime; and KOKKOS only in serial or OpenMP mode. The installer is self-signed, so the browser and SmartScreen show a risk warning; choose "Run anyway" at the prompt.

There are two ways to use multiple cores on Windows. With the GUI installer, use OpenMP: lmp -sf omp -pk omp 8 -in in.file. With the MSMPI installer, use MPI: mpiexec -n 8 lmp -in in.file. The official notes say combining the two is usually slower, so pick one.

Switch to WSL2 when you need CUDA, your own package selection, or the same scripts as on a cluster. CUDA on WSL needs only the NVIDIA driver on the Windows side (R495 or later); inside WSL install only the cuda-toolkit-12-x metapackage. Do not install the cuda, cuda-12-x or cuda-drivers metapackages, because they overwrite the driver mapped into WSL. GPU support in WSL2 requires Pascal or newer GPUs.

Install WSL2 and prepare CUDApowershell
# Windows PowerShell (administrator)
wsl --install -d Ubuntu-24.04
wsl --set-default-version 2

# Inside Ubuntu: install only the CUDA Toolkit, no Linux display driver
# On the NVIDIA download page choose Linux > x86_64 > WSL-Ubuntu and install the cuda-toolkit-12-x metapackage
nvidia-smi                 # if not found: /usr/lib/wsl/lib/nvidia-smi

# keep the working directory on the Linux file system, not under /mnt/c
mkdir -p ~/md && cd ~/md

Step 2

The basic.cmake preset enables KSPACE, MANYBODY, MOLECULE and RIGID, and 30Sep2026 also enables GRAPHICS. EAM is in MANYBODY, which is the only package the tension example on this page needs.

CPU + MPI + OpenMP buildbash
# Dependencies (Ubuntu 22.04/24.04 or Ubuntu on WSL2)
sudo apt update
sudo apt install -y build-essential cmake git libopenmpi-dev openmpi-bin python3-pip
cmake --version            # 30Sep2026 needs >= 3.20; if older: python3 -m pip install --user cmake

# Get the source: the stable branch is the latest stable release (30Sep2026 as of 2026-10)
git clone -b stable --depth 1 https://github.com/lammps/lammps.git lammps
# To match 2025.07.22 use: git clone -b stable_22Jul2025_update6 --depth 1 ...

cd lammps && mkdir build && cd build
cmake -C ../cmake/presets/basic.cmake \
      -D BUILD_MPI=yes -D BUILD_OMP=yes -D PKG_OPENMP=yes \
      -D CMAKE_INSTALL_PREFIX=$HOME/.local \
      ../cmake
cmake --build . -j 8
cmake --install .            # installs into ~/.local, no root needed

# add to ~/.bashrc
export PATH=$HOME/.local/bin:$PATH
export LAMMPS_POTENTIALS=$HOME/.local/share/lammps/potentials

# Check: the MPI library name and packages such as MANYBODY must appear
lmp -h | sed -n '/MPI v/p;/Installed packages/,/List of individual/p'
  • The lmp -h output must show an MPI library name. If CMake cannot find the MPI development package (libopenmpi-dev), it silently builds a serial executable. mpirun -np 4 then runs 4 independent serial copies: the log shows 1 by 1 by 1 MPI processor grid and every output line appears 4 times.
  • mpirun must come from the same MPI used for the build. If the MPI loaded with module load on a cluster differs from the one used at compile time, you get the same N-copies symptom.
  • Enable extra packages with -D PKG_<NAME>=yes, for example -D PKG_REAXFF=yes, -D PKG_MEAM=yes, -D PKG_EXTRA-COMPUTE=yes. To enable most common packages, use -C ../cmake/presets/most.cmake.
  • After changing the package selection, rerun cmake in the same build directory and rebuild; existing settings are kept in CMakeCache.txt.

Step 3

Both run EAM on NVIDIA GPUs, but they split the work differently: the GPU package puts the pair computation on the GPU and leaves the rest on the CPU, while KOKKOS tries to keep the whole time step on the GPU.

The default for GPU_API is opencl. Setting only -D PKG_GPU=on gives an OpenCL build. It runs on NVIDIA GPUs too, but it is not the CUDA back end. This is the setting online tutorials most often leave out.

Set GPU_ARCH and Kokkos_ARCH from the GPU compute capability: V100 is sm_70 / VOLTA70; T4 and RTX 20 series are sm_75 / TURING75; A100 is sm_80 / AMPERE80; RTX 30 series and A40 are sm_86 / AMPERE86; RTX 40 series and L40S are sm_89 / ADA89 (not sm_90); H100 is sm_90 / HOPPER90; RTX 50 series is sm_120 / BLACKWELL120. CUDA 13 can no longer compile offline for Maxwell, Pascal or Volta. The 22Jul2025 default is sm_50, so with CUDA 13 you must set GPU_ARCH explicitly.

A KOKKOS executable moved to a GPU with a different major architecture version (for example from 7.x to 8.x) exits with an error; within the same major version it only adds a JIT compilation at startup. To support several GPU types, build for the oldest one or build one executable per type.

Two GPU builds (sm_86 for RTX 3090 / A40 as an example)bash
# Look up the compute capability first: 8.6 -> sm_86, 8.9 -> sm_89
nvidia-smi --query-gpu=name,compute_cap --format=csv

# Option A: GPU package (CUDA back end)
mkdir build-gpu && cd build-gpu
cmake -C ../cmake/presets/basic.cmake \
      -D PKG_GPU=on -D GPU_API=cuda -D GPU_ARCH=sm_86 -D GPU_PREC=mixed \
      -D BUILD_MPI=yes -D BUILD_OMP=yes -D PKG_OPENMP=yes \
      -D CMAKE_INSTALL_PREFIX=$HOME/.local \
      ../cmake
cmake --build . -j 16

# Option B: KOKKOS (CUDA back end); the two -C presets can be combined
mkdir build-kokkos && cd build-kokkos
cmake -C ../cmake/presets/basic.cmake \
      -C ../cmake/presets/kokkos-cuda.cmake \
      -D Kokkos_ARCH_AMPERE86=yes \
      -D BUILD_MPI=yes \
      ../cmake
cmake --build . -j 16
AspectGPU packageKOKKOS
What is acceleratedpair and parts of PPPM; integration and most fixes stay on the CPUevery pair, fix and compute with a /kk variant runs on the GPU; commands without a /kk variant copy data between CPU and GPU
MPI ranks per GPUseveral ranks sharing one GPU are usually faster; the manual suggests 2–101 by default; sharing a GPU among ranks requires CUDA MPS
Precisionsingle / mixed (default) / doubledouble by default; single and mixed also supported since 30Sep2026
Best formany CPU cores and few GPUs; the OpenCL build on Windowslarge systems on one GPU; the manual says KOKKOS is usually faster with many atoms in double precision
Run flags-sf gpu -pk gpu N-k on g N -sf kk
Key build options-D PKG_GPU=on -D GPU_API=cuda -D GPU_ARCH=sm_XX-C kokkos-cuda.cmake -D Kokkos_ARCH_<GPU>=yes
Sources: LAMMPS manual Speed_gpu, Speed_kokkos, Build_extras (30Sep2026).
  • Common practice (several MatSci forum threads): with CMake, do not also run make yes-gpu in src or build lib/gpu separately. CMake stops with an error when the two build systems are mixed; fix it by running make no-all purge in src or unpacking the source again, and use only -D PKG_GPU=on.
  • Common practice: a runtime error like CUDA driver version is insufficient usually means the CUDA Toolkit is newer than the driver. nvidia-smi shows the highest CUDA version the driver supports in its top-right corner, and nvcc --version shows the Toolkit used for the build; the former should not be lower than the latter.
  • If nvcc is not found during the build, add the CUDA bin directory to PATH (for example export PATH=/usr/local/cuda/bin:$PATH), delete the build directory and rerun cmake.

Step 4

Common ways to runbash
# CPU only, 8 MPI ranks
mpirun -np 8 lmp -in in.cu_tensile

# CPU only, OPENMP package, 1 process with 8 threads (also works with the Windows GUI installer)
lmp -sf omp -pk omp 8 -in in.cu_tensile

# GPU package: 4 MPI ranks share 1 GPU
mpirun -np 4 lmp -sf gpu -pk gpu 1 -in in.cu_tensile

# KOKKOS: 1 MPI rank per GPU
lmp -k on g 1 -sf kk -in in.cu_tensile
# On Pascal and newer GPUs, half neighbor lists are often faster
lmp -k on g 1 -sf kk -pk kokkos newton on neigh half -in in.cu_tensile
# Multiple GPUs with an MPI that is not GPU-aware (common with distribution OpenMPI)
mpirun -np 2 lmp -k on g 2 -sf kk -pk kokkos gpu/aware off -in in.cu_tensile
  • Use -in filename rather than < redirection. With some MPI implementations, < does not deliver the input in parallel runs.
  • The conda mpich variant ships its own mpirun (MPICH Hydra) in the environment's bin folder. After activating the environment, run mpirun -np 8 lmp -in in.cu_tensile directly and do not mix in a system OpenMPI. The MPI line of lmp -h shows MPICH Version. Tested on macOS for this page: with a VPN active, MPICH picks the utun virtual interface and fails with OFI poll failed (default nic=utun...); export FI_PROVIDER=tcp first and it runs normally.
  • The conda CPU package contains neither the GPU nor the KOKKOS package. Tested on this page: -k on fails with Cannot use -kokkos on without KOKKOS installed, and -sf gpu -pk gpu 1 fails with Package gpu command without GPU package installed; -sf omp -pk omp N works.
  • -sf gpu appends the suffix to every command that has a /gpu variant, so eam/alloy becomes eam/alloy/gpu; -sf kk works the same way. You can switch between CPU and GPU from the command line without editing the input file.
  • KOKKOS assumes a GPU-aware MPI by default. A single rank on a single GPU is unaffected. If multiple ranks segfault, or you see Turning off GPU-aware MPI since it is not detected, add -pk kokkos gpu/aware off.
  • Mixed precision is the GPU package default. Before publishing, run the same simulation for a while with a double-precision build or the CPU version and compare energy and stress.
  • To estimate run time, first set run to 2000 steps, read timesteps/s on the Performance line at the end of the log, and extrapolate to the full step count.

Step 5

The system is a periodic bulk Cu crystal of 32000 atoms with the Mishin 2001 EAM potential. The workflow is: energy minimization → equilibration at 300 K and zero pressure for 20 ps → tension along x at 1×10⁹ s⁻¹ for 300 ps with y and z held at zero pressure.

Tested on this page (LAMMPS 22Jul2025 update6, conda-forge CPU build, macOS arm64): with region changed to 0 10 0 10 0 10 (4000 atoms) and all other parameters unchanged, the full 20 ps equilibration and 300 ps tension were run. After equilibration at 300 K, lx is 36.31 Å, a lattice constant of 3.631 Å. plot_ss.py gives a 0–1% strain modulus of 67.3 GPa, which agrees with the theoretical [100] Young's modulus: the same potential at 0 K and 0.1% strain gives C11 = 169.7 GPa and C12 = 122.3 GPa (the Mishin 2001 paper reports 169.9 and 122.6 GPa), and E = (C11−C12)(C11+2C12)/(C11+C12) gives 67.2 GPa. The first yield stress is 8.80 GPa at a strain of 0.113, after which the stress drops abruptly to about 4.3 GPa; σyy and σzz stay within ±0.3 GPa, showing that the zero-pressure control in y and z works. With 864 atoms (6×6×6 cells) the results are 68.6 GPa and 8.85 GPa (strain 0.115). This yield stress is the homogeneous nucleation stress of a defect-free periodic crystal and is about two orders of magnitude above macroscopic experimental values, which is normal for an ideal-crystal model.

Single-core cost (same version and machine, 4000 atoms, 1000 steps each, 99.9% CPU use): the NPT equilibration runs at 204 steps/s (0.82 M atom-step/s); the tension stage with fix deform, fix print and dump runs at 156 steps/s, i.e. 6.4 ms per step (0.62 M atom-step/s). Extrapolated, 32000 atoms take about 51 ms per step on one core, and the 320,000 steps of equilibration plus tension take about 4.5 hours. Other jobs were running on the machine during testing, so multi-process speedups were distorted and are not reported here.

in.cu_tensilelammps
# in.cu_tensile -- uniaxial tension of single-crystal Cu along [100]
# Works with LAMMPS 22Jul2025 and 30Sep2026; requires the MANYBODY package
# ---------- 1. Initialization ----------
units           metal          # length Å, time ps, energy eV, pressure bar
dimension       3
boundary        p p p
atom_style      atomic

variable        T     equal 300.0               # temperature in K
variable        erate equal 1.0e9*1.0e-12       # 1e9 s^-1 converted to 1/ps

# ---------- 2. Build the model ----------
lattice         fcc 3.615                       # matches the Cu lattice constant in the potential file
region          box block 0 20 0 20 0 20        # 20x20x20 unit cells, 32000 atoms
create_box      1 box
create_atoms    1 box

# ---------- 3. Potential ----------
pair_style      eam/alloy
pair_coeff      * * Cu_mishin1.eam.alloy Cu     # file ships in the LAMMPS potentials/ folder
neighbor        2.0 bin
neigh_modify    delay 0 every 1 check yes

# ---------- 4. Energy minimization, relaxing box pressure to 0 ----------
thermo          100
thermo_style    custom step pe press lx ly lz
fix             relax all box/relax iso 0.0 vmax 0.001
min_style       cg
minimize        1.0e-12 1.0e-12 10000 100000
unfix           relax

# ---------- 5. Equilibrate at 300 K, 0 bar for 20 ps ----------
reset_timestep  0
timestep        0.001                           # 1 fs
velocity        all create ${T} 4928459 mom yes rot yes dist gaussian
fix             eq all npt temp ${T} ${T} 0.1 iso 0.0 0.0 1.0
thermo          1000
thermo_style    custom step temp pe press lx ly lz
run             20000
unfix           eq

# ---------- 6. Constant strain-rate tension along x, y and z held at 0 bar ----------
variable        tmp equal lx
variable        L0 equal ${tmp}                 # evaluated immediately: length before loading
reset_timestep  0
fix             nptyz all npt temp ${T} ${T} 0.1 y 0.0 0.0 1.0 z 0.0 0.0 1.0 drag 1.0
fix             pull all deform 1 x erate ${erate} units box remap x

variable        strain equal (lx-v_L0)/v_L0
variable        sxx equal -pxx/10000            # bar to GPa; sign flipped so tension is positive
variable        syy equal -pyy/10000
variable        szz equal -pzz/10000
fix             out all print 100 "${strain} ${sxx} ${syy} ${szz}" file stress_strain.txt screen no title "# strain sxx_GPa syy_GPa szz_GPa"

compute         peatom all pe/atom
dump            traj all custom 2000 tensile.lammpstrj id type x y z c_peatom
thermo_style    custom step v_strain temp v_sxx v_syy v_szz pe press
restart         50000 tensile.*.restart
run             300000                          # 300 ps, engineering strain reaches 0.30
write_data      final.data
  1. 01

    units metal

    Sets the units for every number that follows: distance in Å, time in ps, energy in eV, pressure in bar. Most EAM, Tersoff and SW potential files are given in metal units.

  2. 02

    lattice fcc 3.615 and region/create_box/create_atoms

    lattice sets both the lattice and the length unit, so 0 20 in region means 20 unit cells, about 72 Å. Use the equilibrium lattice constant of the potential, otherwise there is residual stress from the start. The header of Cu_mishin1.eam.alloy gives 3.615 Å.

  3. 03

    pair_style eam/alloy and pair_coeff * * file Cu

    eam/alloy reads setfl files (.eam.alloy). It takes exactly one pair_coeff * * line followed by element names in atom-type order, one name per atom type. The setfl file contains the mass, so no mass command is needed. .eam files (funcfl format) go with pair_style eam. Mixing the two reads the file incorrectly.

  4. 04

    neighbor and neigh_modify

    The default skin in metal units is 2.0 Å. Since 2Aug2023 the neigh_modify default is delay 0 every 1 check yes. The delay 10 in older tutorials can cause Dangerous builds warnings; remove it.

  5. 05

    fix box/relax + minimize

    Adjusts the box size during minimization so the 0 K pressure is zero. vmax 0.001 limits the volume change per iteration and keeps the minimization from diverging.

  6. 06

    velocity + fix npt iso

    Equilibrates at 300 K and 0 bar for 20 ps so the lattice reaches its thermal expansion. Skipping this step puts thermal stress at the start of the curve. The damping parameters 0.1 ps (Tdamp) and 1.0 ps (Pdamp) are 100 and 1000 steps, the order of magnitude the manual recommends.

  7. 07

    variable L0 equal ${tmp}

    ${} is evaluated when the line is read, so L0 stays fixed at the length before loading. With variable L0 equal lx, L0 would change with the box and the strain would always be 0.

  8. 08

    fix npt y/z + fix deform x erate

    deform stretches x at a constant engineering strain rate. erate is in 1/time units, which is 1/ps in metal units, so 1×10⁹ s⁻¹ is written 0.001. y and z are held at 0 bar with npt, allowing Poisson contraction and giving a uniaxial stress state. With deform alone and fixed y and z, you get uniaxial strain and higher stresses. Do not barostat the pulled direction x at the same time, because the barostat counteracts the deformation; this is one of the most common causes of odd tensile curves on the MatSci forum.

  9. 09

    variable sxx equal -pxx/10000

    pxx is a component of the pressure tensor in bar and is negative under tension. The minus sign makes tensile stress positive, and dividing by 10000 converts to GPa. When a variable uses pxx, pyy or pzz, thermo_style must contain a pressure keyword such as press so that LAMMPS computes the pressure tensor. Tested on this page: with press removed from the end of thermo_style, the tensile run stops immediately with ERROR: Thermo keyword pxx in variable requires thermo to use/init press.

  10. 10

    fix print, dump, restart

    The fix print string is in double quotes, so ${strain} inside it is evaluated at each output. dump writes a .lammpstrj file that OVITO and VMD open directly. restart writes a restart file every 50000 steps.

Variant

Nanowire tension is the most common modification of the bulk script. Three things change: the model, the barostat directions, and the volume used to normalize stress.

Keep periodic boundaries in y and z, but leave a vacuum layer much wider than the cutoff (5.5 Å for the Mishin potential) so the wire does not interact with its periodic images. In the example the box is about 72 Å wide, the wire about 40 Å in diameter, and the vacuum about 32 Å.

pxx is the virial stress of the whole box divided by the box volume. The box volume includes vacuum, so plain -pxx/10000 clearly underestimates the stress in the wire. The conversion above assumes the cross-section does not change during loading. Alternatively, sum compute stress/atom and divide by the number of atoms times the volume per atom. Whichever you use, state the definition of the normalization volume in the paper.

Tested on this page (same version; length shortened to 20 cells, 7540 atoms; 4 ps of equilibration at 300 K, then tension at 1×10⁹ s⁻¹ to a strain of 0.03): after equilibration lx is 71.52 Å, 1.1% shorter than the built length of 72.30 Å because surface stress contracts the wire axially, so take L0 after equilibration. The modulus fitted over 0–1% strain is 55.8 GPa with the volume conversion above and only 13.2 GPa with plain -pxx/10000; the ratio of about 4.2 equals the ratio of the box cross-section to the wire cross-section. Summing compute stress/atom and dividing by the number of atoms times the atomic volume a³/4 (corrected for the axial elongation) gives a stress within 0.5% of the conversion above. stress/atom has the opposite sign to pressure, so do not negate the sum again.

Changes to in.cu_tensilelammps
# Replace step 2 with these lines: [100] Cu nanowire, about 4 nm in diameter and 20 nm long
lattice         fcc 3.615
region          box block 0 55 -10 10 -10 10    # lattice units; integer number of cells along x
create_box      1 box
region          wire cylinder x 0 0 5.5 INF INF # radius of 5.5 lattice constants, about 19.9 Å
create_atoms    1 region wire

# Step 4: relax the box along x only
fix             relax all box/relax x 0.0 vmax 0.001
# Step 5: barostat along x only
fix             eq all npt temp ${T} ${T} 0.1 x 0.0 0.0 1.0
# Step 6: y and z are vacuum, so no barostat there; use NVT
fix             nvt1 all nvt temp ${T} ${T} 0.1
fix             pull all deform 1 x erate ${erate} units box remap x

# pxx is normalized by the whole box volume (including vacuum); rescale to the wire volume
variable        R    equal 5.5*3.615
variable        sxx  equal -pxx*vol/(PI*v_R^2*lx)/10000   # GPa, ignoring Poisson contraction

Potentials

You do not need to download a Cu EAM potential: Cu_mishin1.eam.alloy (Mishin et al., PRB 63, 224106, 2001), Cu_u3.eam (Foiles 1986) and Cu_zhou.eam.alloy are all in the LAMMPS potentials/ folder. On the NIST Interatomic Potentials Repository (ctcms.nist.gov/potentials/system/Cu), the same Mishin potential is listed as Cu01.eam.alloy. NIST notes that the Mendelev 2013 Cu2 potential improves stacking fault energies, and that the ipr2 version of Zhou 2004 corrects spurious fluctuations in the tabulated functions. For dislocation and twinning studies, choose a potential with these notes in mind and cite the source in your methods.

A conda installation has no such folder; download the file from the GitHub tag that matches your version: curl -LO https://raw.githubusercontent.com/lammps/lammps/stable_22Jul2025_update6/potentials/Cu_mishin1.eam.alloy, and put it next to the input file, or in a fixed folder that LAMMPS_POTENTIALS points to.

OpenKIM potentials need no file download; reference the model ID in the input: kim init EAM_Dynamo_MishinMehlPapaconstantopoulos_2001_Cu__MO_346334655118_006 metal, then kim interactions Cu after the model is built. This requires the KIM package. The conda and apt builds include KIM; the Windows prebuilt packages do not.

Recent versions read the UNITS tag on the first line of the potential file. If the units in the input file do not match, LAMMPS reports Potential file ... requires metal units but real units are in use. Fix it by changing units or by using a potential file in the matching units.

SystemCommon potentials (pair_style)PackagesUnitsSources
Metals and alloyseam, eam/alloy, eam/fs, meamMANYBODY, MEAMmetalNIST IPR, OpenKIM, LAMMPS potentials/ folder
Covalent crystals (Si, Ge, SiC)sw, tersoffMANYBODYmetalpotentials/ folder, NIST IPR, OpenKIM
Carbon materials (graphene, CNT, hydrocarbons)airebo, tersoffMANYBODYmetalpotentials/CH.airebo
Polymers and organic molecules (fixed topology)lj/cut/coul/long + bond, angle, dihedral terms (OPLS-AA, GAFF); lj/class2 (PCFF, COMPASS)MOLECULE, KSPACE, CLASS2realData files generated with moltemplate, LigParGen or msi2lmp
Reactive systems (combustion, oxidation, bond breaking and forming)reaxff + fix qeq/reaxff; or fix bond/reactREAXFF; REACTIONrealReaxFF parameters in the potentials/ folder and in paper supplements
Near-DFT accuracysnap, pace, mliapML-SNAP, ML-PACE, ML-IAPmetalFiles supplied with papers, OpenKIM
The Units column refers to the potential files shipped with LAMMPS. Source: Restrictions sections of each pair_style page.

Units

Quantitymetalreal
DistanceÅÅ
Timeps (default timestep 0.001 = 1 fs)fs (default timestep 1.0 = 1 fs)
EnergyeVkcal/mol
Pressurebar (1 GPa = 10000 bar)atm (1 GPa = 9869.2 atm)
fix deform erate1/ps1/fs
Default neighbor skin2.0 Å2.0 Å
Sources: LAMMPS manual units, neighbor, fix_deform.
  • When switching from metal to real, multiply the erate value by a further 10⁻³: 1×10⁹ s⁻¹ is 1.0e-6 (1/fs) in real units.
  • compute stress/atom is in pressure×volume units (bar·Å³) and has the opposite sign of pressure. It is not a per-atom stress; divide by a per-atom volume before reading it as stress.
  • press in thermo is (pxx+pyy+pzz)/3, which is about -σxx/3 in uniaxial tension. press cannot stand in for σxx.
  • Strain rates in MD tension are typically 10⁷–10¹⁰ s⁻¹, many orders of magnitude above experiment, so yield stresses come out high. Common practice: before production, run once at 10⁹ and once at 10⁸ s⁻¹ to see how much the yield stress changes, then decide whether a lower rate is worth the cost. When comparing sizes or temperatures, keep the strain rate the same and state it in the paper.

Step 6

In log.lammps, the thermo block of each run starts with a header line beginning with Step and ends with a line starting Loop time of. The header columns are the thermo_style keywords, and variables appear as v_strain and v_sxx. After Loop time comes a timing breakdown (shares of Pair, Neigh, Comm, Output and so on): a low Pair share with a high Comm or Output share means too many processes or too frequent output.

OVITO opens tensile.lammpstrj directly (and also reads .gz). Add the Common neighbor analysis modifier to the pipeline to color atoms as FCC, HCP or Other. In FCC copper, a single layer of HCP atoms marks a twin boundary and two adjacent HCP layers mark an intrinsic stacking fault. VMD opens the file with vmd -lammpstrj tensile.lammpstrj; it cannot read compressed files.

Tested on this page: running plot_ss.py on the 4000-atom case above prints E(0-1%) = 67.3 GPa and first yield = 8.80 GPa at strain 0.113. Small cells reload after the first yield; the 864-atom case reaches a higher peak of 9.9 GPa at a strain of 0.23, so the script takes the maximum before the first clear drop instead of the global maximum. The runs read by lammps.formats.LogFile correspond to minimization, equilibration and tension in that order, and the keys of runs[-1] are Step, v_strain, Temp, v_sxx, v_syy, v_szz, PotEng and Press.

Dislocation density versus strain is a common post-processing request in the Chinese community; several CSDN articles demonstrate it, and the scripts are mostly shared in QQ groups. The dxa_density.py script below uses only the public interface of the OVITO Python module (pip install ovito): DislocationAnalysisModifier outputs the total line length DislocationAnalysis.total_line_length (Å) and the cell volume DislocationAnalysis.cell_volume (ų); dividing them and multiplying by 10²⁰ gives m⁻². DislocationAnalysis.length.1/6<112> is the length of Shockley partials.

Tested on this page (OVITO 3.16.1, the 151-frame tension trajectory of the 4000-atom case above, about 11 s): DXA finds no dislocations before a strain of 0.116, while the FCC fraction is 0.936 at a strain of 0.100 and drops to 0.625 at 0.116. The first dislocations appear at a strain of 0.118 with a density of 4.1×10¹⁷ m⁻², and the HCP fraction rises to 0.149. After that the density jumps between 0 and 1.4×10¹⁸ m⁻² and is 9.6×10¹⁷ m⁻² at a strain of 0.300, while the HCP fraction stays between 0.15 and 0.33. The jumps come from the small box: with an edge of about 36 Å, one dislocation line spanning the box corresponds to about 8×10¹⁶ m⁻², so nucleation and annihilation across periodic boundaries change single-frame values several-fold. To compare dislocation densities in a paper, use a larger box or average over a strain interval.

plot_ss.pypython
# plot_ss.py: read the fix print output, plot the stress-strain curve, fit the initial modulus and find the first yield
import numpy as np
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt

d = np.loadtxt("stress_strain.txt", comments="#")
strain, sxx = d[:, 0], d[:, 1]

# 20-point moving average to suppress thermal noise (100 steps per point); valid mode avoids the false drop caused by zero padding at both ends
w = 20
sxx_s = np.convolve(sxx, np.ones(w) / w, mode="valid")
strain_s = strain[w // 2 : w // 2 + len(sxx_s)]

lin = strain <= 0.01                     # fit only 0–1% strain
E = np.polyfit(strain[lin], sxx[lin], 1)[0]

# First yield: the maximum before the stress first falls 20% below its running maximum (after exceeding 1 GPa)
# Small cells reload after yielding and the global maximum can come later, so plain argmax is not enough
run_max = np.maximum.accumulate(sxx_s)
drop = np.nonzero((run_max > 1.0) & (run_max - sxx_s > 0.2 * run_max))[0]
j = drop[0] if drop.size else len(sxx_s)
i = np.argmax(sxx_s[:j])
print(f"E(0-1%) = {E:.1f} GPa")
print(f"first yield = {sxx_s[i]:.2f} GPa at strain {strain_s[i]:.3f}")

plt.plot(strain, sxx, lw=0.5, alpha=0.4, label="raw")
plt.plot(strain_s, sxx_s, lw=1.5, label=f"moving avg ({w})")
plt.xlabel("Engineering strain")
plt.ylabel("Stress σxx (GPa)")
plt.legend()
plt.tight_layout()
plt.savefig("stress_strain.png", dpi=200)

# You can also parse log.lammps directly (needs the LAMMPS Python module; the conda build includes it)
# from lammps.formats import LogFile
# run = LogFile("log.lammps").runs[-1]   # keys are the thermo column names, e.g. "v_strain", "v_sxx"
dxa_density.py: dislocation density per framepython
# dxa_density.py: run DXA on every frame and print dislocation density and FCC/HCP fractions
import sys
from ovito.io import import_file
from ovito.modifiers import DislocationAnalysisModifier

pipeline = import_file(sys.argv[1] if len(sys.argv) > 1 else "tensile.lammpstrj")
dxa = DislocationAnalysisModifier(input_crystal_structure=DislocationAnalysisModifier.Lattice.FCC)
pipeline.modifiers.append(dxa)

print("# step  rho_total(1/m^2)  rho_shockley(1/m^2)  FCC_frac  HCP_frac")
for frame in range(pipeline.source.num_frames):
    data = pipeline.compute(frame)
    a = data.attributes
    vol = a["DislocationAnalysis.cell_volume"]                 # Å^3
    rho = a["DislocationAnalysis.total_line_length"] / vol * 1e20   # Å/Å^3 -> m^-2
    rho_s = a.get("DislocationAnalysis.length.1/6<112>", 0.0) / vol * 1e20
    n = data.particles.count
    fcc = a["DislocationAnalysis.counts.FCC"] / n
    hcp = a["DislocationAnalysis.counts.HCP"] / n
    print(f"{a['Timestep']:8d}  {rho:.3e}  {rho_s:.3e}  {fcc:.3f}  {hcp:.3f}")

Continuation

Restart files store atoms, the box, units, atom_style and similar settings. What they do not store must be re-specified in the continuation script: all fixes and computes, and pair_coeff for many-body potentials that read parameters from files (EAM, Tersoff and others).

When fix deform uses the same fix-ID as the original script (pull here), it restores the box from the start of loading from the restart file, but only uses it when the run command has start and stop. start is the step at which the original loading run began (0 after reset_timestep 0), and stop is the total step count. With only run 300000 upto, deform recomputes the elongation from the box length at the restart. Tested on this page (an 864-atom, 1×10¹⁰ s⁻¹ shortened case, continued from strain 0.10 to the planned final step): with upto alone the final strain is 0.32; with start 0 stop it is 0.30, and the final stress of 6.80 GPa agrees with 6.78 GPa from the uninterrupted run. Continuing the 4000-atom case from the “Your first input file” section from step 150000 for 1000 steps with this script (only L0 changed) gives the same strain as the original run and stresses that agree to 7 significant digits. Continuing the page's parameters from step 150000 without start/stop would end at a strain of 0.3225. L0 must be the length before the original loading, not lx at the restart, otherwise the strain starts again from 0. run N upto stops at total step N, so any number of interruptions still end at the same point.

in.cu_tensile.continuelammps
# in.cu_tensile.continue -- continue tension from the restart at step 150000
read_restart    tensile.150000.restart
pair_style      eam/alloy                        # many-body potentials are not stored in restart files; re-specify
pair_coeff      * * Cu_mishin1.eam.alloy Cu
neighbor        2.0 bin

variable        T     equal 300.0
variable        erate equal 1.0e9*1.0e-12
variable        L0    equal 72.5                 # set to the number after variable L0 equal in the original log
timestep        0.001
fix             nptyz all npt temp ${T} ${T} 0.1 y 0.0 0.0 1.0 z 0.0 0.0 1.0 drag 1.0
fix             pull all deform 1 x erate ${erate} units box remap x   # same fix-ID as the original script

variable        strain equal (lx-v_L0)/v_L0
variable        sxx equal -pxx/10000
fix             out all print 100 "${strain} ${sxx}" append stress_strain.txt screen no
thermo_style    custom step v_strain temp v_sxx pe press
thermo          1000
restart         50000 tensile.*.restart
run             300000 upto start 0 stop 300000  # run to total step 300000; start/stop keep deform on the original box length

Errors

Causes and fixes for Out of range atoms, Too many neighbor bins, NaN and other errors are covered in LAMMPS Common Errors and Fixes (/learn/lammps-errors).

Error message (excerpt)CauseFix
Unrecognized pair style 'eam/alloy' is part of the MANYBODY package which is not enabled in this LAMMPS binaryMANYBODY was not enabled at build timeRerun cmake with -D PKG_MANYBODY=yes, or use the basic.cmake preset
cannot open eam/alloy potential file Cu_mishin1.eam.alloy: No such file or directoryThe potential file is neither in the current directory nor in the directory LAMMPS_POTENTIALS points toCopy the file next to the input file, or export LAMMPS_POTENTIALS=<install prefix>/share/lammps/potentials; the conda build has no such folder, so download the potential from GitHub first (see the potentials section)
Potential file Cu_mishin1.eam.alloy requires metal units but real units are in useunits does not match the potential fileSwitch to units metal and rewrite timestep and erate in metal units
ERROR: Lost atoms: original 32000 current 31998Time step too large, overlapping atoms in the initial structure, or wrong potential or unitsCheck the lattice constant and timestep first; see the LAMMPS errors page for the full checklist

Working from China

This section collects practices that recur on the Computational Chemistry Commune forum (keinsci), CSDN, Zhihu and university HPC documentation, checked against the current documentation or source code of LAMMPS and the related tools. Numbers marked as tested on this page come from actual queries or runs on 2026-10-10.

~/.condarc (Tsinghua mirror)yaml
# ~/.condarc: route conda-forge through the Tsinghua TUNA mirror (community-only setup from the TUNA help page)
channels:
  - conda-forge
  - nodefaults
show_channel_urls: true
custom_channels:
  conda-forge: https://mirrors.tuna.tsinghua.edu.cn/anaconda/cloud
Install a specific variant from the mirrorbash
conda clean -i
# build string = accelerator + MPI: cpu, cuda129, cuda130, cuda134 combined with nompi, mpich, openmpi
conda create -n lammps -y "lammps=2025.07.22=cpu*openmpi*"          # about 410 MB total for a new env
conda create -n lammps-gpu -y "lammps=2025.07.22=cuda129*openmpi*"  # about 1 GB total for a new env
conda activate lammps && lmp -h | head -20
Cluster without internet: supply the third-party archive yourselfbash
# 1) On a machine with internet access, download the third-party archive (URL as in cmake/Modules/Packages/VORONOI.cmake)
curl -LO https://download.lammps.org/thirdparty/voro++-0.4.6.tar.gz
# 2) Copy it to the cluster and point <prefix>_URL at the local file; the default SHA256 is still checked, so the version must match
cd lammps/build
cmake ../cmake -C ../cmake/presets/basic.cmake -D PKG_VORONOI=yes -D DOWNLOAD_VORO=yes \
      -D VORO_URL=file://$HOME/src/voro++-0.4.6.tar.gz
cmake --build . -j 8

# Many downloads: run tools/offline/init_caches.sh on the online machine (default cache ~/.cache/lammps),
# copy the cache to the same path on the cluster, source tools/offline/use_caches.sh and add the -D LAMMPS_DOWNLOADS_URL=... options it prints
VMD TopoTools: pdb to data filetcl
# VMD Tk Console: convert a Packmol pdb to a LAMMPS data file (TopoTools)
package require topotools
mol new system.pdb
# for pdb files VMD takes the type from the atom name, so C1, C2, H1 become separate types; merge them by element first
[atomselect top {name "C.*"}] set type C
[atomselect top {name "H.*"}] set type H
pbc set {40.0 40.0 40.0}        ;# writelammpsdata refuses to write without a box
topo writelammpsdata system.data full
# type IDs follow the alphabetical order of type names: C=1, H=2; the comment at the end of each Masses line is the type name

Use the Tsinghua conda mirror and pick the variant with a build string

When conda-forge downloads are slow in China, follow the Tsinghua TUNA mirror help page and point conda-forge at mirrors.tuna.tsinghua.edu.cn/anaconda/cloud in ~/.condarc. Tested on this page (micromamba query of the mirror on 2026-10-10): linux-64 has lammps 2025.07.22 in four build families, cpu, cuda129, cuda130 and cuda134, each split into nompi, mpich and openmpi. With a bare conda install lammps the solver decides which MPI to use and whether to include CUDA; a build string such as "lammps=2025.07.22=cpu*openmpi*" pins the variant. Total download for a new environment: about 410 MB for cpu+openmpi and about 1 GB for cuda129+openmpi.

On a university cluster, check the modules before building your own

University clusters usually provide LAMMPS already, and the executable name differs between modules. The SCUT platform provides a source-built lammps/2Aug2023 (executable lmp_intel_cpu_intelmpi) and a conda environment lammps20230802-cpu (executable lmp). SJTU's Siyuan-1 provides CPU modules such as lammps/20230328-intel-2021.4.0-omp and lammps/20230802-intel-2021.4.0-gpu, with the example run command mpirun lmp -pk intel 0 omp 2 -sf intel -i in.lj; drop -pk intel and -sf intel when the force field is not supported by the INTEL package. Run module avail lammps first, then lmp -h after loading to see the installed packages. These modules are mostly 2023 versions, so new keywords written from the 30Sep2026 documentation may not be recognized; when that happens, check the documentation for that version. The SJTU documentation requires builds to run on compute nodes, since parallel compilation is not allowed on login nodes.

Third-party downloads fail during the build

With VORONOI, PLUMED, ML-PACE, KIM and similar packages enabled, CMake downloads source archives from GitHub, download.lammps.org or OpenKIM. On an unstable campus network or a cluster without internet access, CMake reports Each download failed! (the exact text in a CSDN write-up of a build on an offline cluster). That author placed the archive by hand and faked CMake's step stamp files. A more direct fix is to point the cache variable <prefix>_URL at a local file, for example VORO_URL, PLUMED_URL, PACELIB_URL or KIM_URL; the prefix and default URL are on the SetDownloadSettings line in cmake/Modules/Packages/<package>.cmake. The GPU package and KOKKOS libraries ship with the LAMMPS source (DOWNLOAD_KOKKOS is off by default), so a GPU build does not need to download either of them.

Pitfalls when converting from Packmol, VMD TopoTools or Materials Studio to a data file

Packmol only keeps atoms inside the box at least tolerance apart; it ignores periodic images. In his Packmol tutorial on the keinsci forum, Sobereva uses add_box_sides 1.2 to add 1.2 Å to the box in each direction; from Packmol 20.15.0 the pbc keyword packs directly into a periodic box. VMD TopoTools' topo writelammpsdata fails with need to have non-zero box sizes when no box is set, so run pbc set first; it assigns type IDs in alphabetical order of the VMD type names, and for pdb files the type name is taken from the atom name. Data files exported through OVITO after building in Materials Studio also often have a type order that differs from the force-field file, and a keinsci thread reorders the type IDs by hand in Word and Excel. For potentials written as pair_coeff * * file element…, such as eam/alloy, tersoff and reaxff, the data file does not need to change: list the element names in the order of data-file types 1, 2, 3 and so on. The example in the LAMMPS manual: with elements ordered C H O N in the ffield file and types 1 to 4 being C, C, N, H, write pair_coeff * * ffield.reax C C N H.

Hand it to an agent

Example one-sentence instruction: "Use LAMMPS to simulate uniaxial tension of a [100] single-crystal Cu nanowire about 4 nm in diameter and 20 nm long at 300 K and 1×10⁹ s⁻¹, with the Mishin EAM potential and GPU acceleration; output the stress-strain curve, Young's modulus, yield strength and common neighbor analysis snapshots, and compare with a nanowire about 2 nm in diameter to show the size effect."

The agent carries out these steps: rents a GPU and runs the environment self-test; writes the input files for model building, relaxation and tension; runs a short test to estimate run time before submitting the production run; extracts stress-strain data with Python and fits the modulus; marks dislocations and twins with common neighbor analysis; and reviews the results adversarially to confirm the task is actually done. The deliverables are the input files, post-processing scripts, logs, plots, structure snapshots, an environment record and a report. Compute is released when the run finishes.

You still need to check these yourself: whether the potential suits the deformation mechanism you study; the definition of the stress normalization volume; whether the effects of strain rate and size on yield strength are stated in the conclusions; and whether repeats with different random seeds are needed to estimate statistical error.

References

FAQ

Can LAMMPS use GPU acceleration on Windows?

The official installers include the GPU package in OpenCL mode. If the graphics driver provides an OpenCL runtime, run lmp -sf gpu -pk gpu 1 -in in.file. The Windows prebuilt packages do not support the KOKKOS GPU back end. For a CUDA build, compile from source in WSL2.

What is the minimum content of a LAMMPS data file?

The first line is a comment. Then come the atom count (32000 atoms), the number of atom types (1 atom types) and three lines of box bounds (0.0 72.3 xlo xhi and so on). After a blank line comes the Masses section, then after another blank line the Atoms # atomic section, one line per atom: atom-ID type x y z. The Atoms columns depend on atom_style; full, for example, needs a molecule ID and a charge. Before reading the file with read_data, check that atom_style in the input matches the data file.

How do I run LAMMPS on multiple cores?

With an MPI build, use mpirun -np 8 lmp -in in.file. With an OpenMP-only build (such as the Windows GUI installer), use lmp -sf omp -pk omp 8 -in in.file. To confirm parallel execution, look at MPI processor grid in the log: with 8 ranks it should show a decomposition whose product is 8, such as 2 by 2 by 2.

What etol and ftol should I use for minimize?

For pre-relaxing a bulk crystal before tension, etol 1e-12 and ftol 1e-12 with generous iteration limits are common. After minimization, check Stopping criterion in the Minimization stats block of the log: energy tolerance or force tolerance means it converged; max iterations means it did not, so raise the limits or check the structure.

Can I install LAMMPS with Docker?

Yes. NVIDIA NGC provides the nvcr.io/hpc/lammps image. The host needs the NVIDIA driver and the NVIDIA Container Toolkit; run with --gpus all and mount your working directory into the container. Inside the container, run lmp -h first to check the version and whether the built-in accelerator is the GPU package or KOKKOS, then choose -sf gpu or -k on g 1 -sf kk accordingly.

Why does -sf gpu do nothing with the conda build of LAMMPS?

The conda-forge CUDA builds disable the GPU package and enable KOKKOS, so use -k on g 1 -sf kk. Also, conda-forge currently provides version 2025.07.22 and has no Windows build. The conda-forge CPU build has neither the GPU package nor KOKKOS; tested on this page, -sf gpu -pk gpu 1 fails with Package gpu command without GPU package installed.

Hand your LAMMPS simulation to Scientify

Describe the system, the potential and the quantities you need. The scientific agent rents a GPU in a cloud computer, uses the preinstalled LAMMPS 2025.07.22 GPU build for model building, simulation and post-processing, and keeps every script and log. New users get USD 5 of free credit.