Step 1
Choose a route by three questions: do you need MPI parallel runs, do you need CUDA, and do you have root access. The typical cost of a wrong choice is finding out after installation that you cannot run in parallel, or that the GPU is not used at all.
| Route | When to use it | MPI | GPU | Notes |
|---|---|---|---|---|
| Windows LAMMPS-GUI installer (GitHub release, LAMMPS-Win10-x64-GUI-30Sep2026.exe) | Learning and debugging input files, small systems | No | GPU package, OpenCL mixed precision | Includes command-line lmp; no KIM and some other packages; SmartScreen shows a warning during installation |
| Windows MS-MPI installer (rpm.lammps.org/windows, LAMMPS-64bit-30Sep2026-MSMPI.exe) | Multi-process parallel runs on Windows | MS-MPI only; install msmpisetup.exe separately | GPU package, OpenCL | Run with mpiexec -n 4 lmp -in ...; the installer sets PATH and LAMMPS_POTENTIALS; MS-MPI only since 2Aug2023 |
| WSL2 + CMake source build | Windows machine that needs CUDA or MPI | OpenMPI | GPU package or KOKKOS (CUDA) | Install no Linux display driver inside WSL; keep files on the Linux file system |
| conda-forge (conda install lammps) | Linux/macOS without compiling | openmpi/mpich/nompi variants | CUDA variants use KOKKOS, no GPU package | Version is 2025.07.22; no Windows build; includes KIM; ships no potentials/ folder and does not set LAMMPS_POTENTIALS (tested with the osx-arm64 package on this page) |
| Linux static binary (download.lammps.org/static) | Cluster login nodes, no root, just need it to run | No | None | No dependence on system libraries; no Python module |
| apt / dnf distribution packages | Quick trial | Yes | None | apt version is 20240207 on Ubuntu 24.04 and 20220106 on 22.04 |
| CMake source build (Linux, clusters) | Production runs, custom packages or GPU | Yes | GPU package or KOKKOS | Main route of this page; can install into $HOME without root |
| Docker/Apptainer (NVIDIA NGC nvcr.io/hpc/lammps) | GPU machines with the NVIDIA Container Toolkit | Built into the image | Yes, CUDA | Run lmp -h inside the container first to see whether the accelerator is the GPU package or KOKKOS |
Version
The new stable release of 30 September 2026 raised the toolchain requirements. On clusters with older GCC or CMake, building the new version fails at the configure step.
How to decide: run gcc --version and cmake --version first. If GCC is older than 10.4 and you need KOKKOS, use stable_22Jul2025_update6 or install a newer compiler first. If CMake is too old, python3 -m pip install --user cmake installs a recent version without root.
If your input must be compared with results from a supervisor or a published study, record the full version string printed by lmp -h. 30Sep2026 fixed platform-dependent results on ARM64 and on platforms with FMA enabled, so the two versions can differ in the last digits.
| Item | stable_22Jul2025_update6 | stable_30Sep2026 |
|---|---|---|
| Minimum CMake | 3.16 | 3.20 |
| C++ standard | C++11 by default; KOKKOS needs C++17 | C++17; KOKKOS needs C++20 |
| Bundled Kokkos | 4.6.2 (GCC ≥ 8.2, CUDA ≥ 11.0) | 5.2 (GCC ≥ 10.4, NVCC ≥ 12.2) |
| GPU package default GPU_ARCH | sm_50 | sm_75 |
| Traditional make build | Works for KOKKOS | KOKKOS is CMake-only since 11Feb2026; make only builds packages that need no extra steps in lib/ |
| dump image / dump movie | Core feature | Moved to the new GRAPHICS package, included in the basic.cmake preset |
| Removed packages | — | ATC, AWPMD, POEMS, ML-RANN, VTK; pair agni (dump vtk moved to EXTRA-DUMP) |
Windows
The LAMMPS-GUI installer suits learning: the editor runs the current input file directly and shows thermo plots and snapshots. It has three limits: no MPI parallel runs; GPU support only through the OpenCL GPU package, which needs a graphics driver with an OpenCL runtime; and KOKKOS only in serial or OpenMP mode. The installer is self-signed, so the browser and SmartScreen show a risk warning; choose "Run anyway" at the prompt.
There are two ways to use multiple cores on Windows. With the GUI installer, use OpenMP: lmp -sf omp -pk omp 8 -in in.file. With the MSMPI installer, use MPI: mpiexec -n 8 lmp -in in.file. The official notes say combining the two is usually slower, so pick one.
Switch to WSL2 when you need CUDA, your own package selection, or the same scripts as on a cluster. CUDA on WSL needs only the NVIDIA driver on the Windows side (R495 or later); inside WSL install only the cuda-toolkit-12-x metapackage. Do not install the cuda, cuda-12-x or cuda-drivers metapackages, because they overwrite the driver mapped into WSL. GPU support in WSL2 requires Pascal or newer GPUs.
# Windows PowerShell (administrator)
wsl --install -d Ubuntu-24.04
wsl --set-default-version 2
# Inside Ubuntu: install only the CUDA Toolkit, no Linux display driver
# On the NVIDIA download page choose Linux > x86_64 > WSL-Ubuntu and install the cuda-toolkit-12-x metapackage
nvidia-smi # if not found: /usr/lib/wsl/lib/nvidia-smi
# keep the working directory on the Linux file system, not under /mnt/c
mkdir -p ~/md && cd ~/mdStep 2
The basic.cmake preset enables KSPACE, MANYBODY, MOLECULE and RIGID, and 30Sep2026 also enables GRAPHICS. EAM is in MANYBODY, which is the only package the tension example on this page needs.
# Dependencies (Ubuntu 22.04/24.04 or Ubuntu on WSL2)
sudo apt update
sudo apt install -y build-essential cmake git libopenmpi-dev openmpi-bin python3-pip
cmake --version # 30Sep2026 needs >= 3.20; if older: python3 -m pip install --user cmake
# Get the source: the stable branch is the latest stable release (30Sep2026 as of 2026-10)
git clone -b stable --depth 1 https://github.com/lammps/lammps.git lammps
# To match 2025.07.22 use: git clone -b stable_22Jul2025_update6 --depth 1 ...
cd lammps && mkdir build && cd build
cmake -C ../cmake/presets/basic.cmake \
-D BUILD_MPI=yes -D BUILD_OMP=yes -D PKG_OPENMP=yes \
-D CMAKE_INSTALL_PREFIX=$HOME/.local \
../cmake
cmake --build . -j 8
cmake --install . # installs into ~/.local, no root needed
# add to ~/.bashrc
export PATH=$HOME/.local/bin:$PATH
export LAMMPS_POTENTIALS=$HOME/.local/share/lammps/potentials
# Check: the MPI library name and packages such as MANYBODY must appear
lmp -h | sed -n '/MPI v/p;/Installed packages/,/List of individual/p'- The lmp -h output must show an MPI library name. If CMake cannot find the MPI development package (libopenmpi-dev), it silently builds a serial executable. mpirun -np 4 then runs 4 independent serial copies: the log shows 1 by 1 by 1 MPI processor grid and every output line appears 4 times.
- mpirun must come from the same MPI used for the build. If the MPI loaded with module load on a cluster differs from the one used at compile time, you get the same N-copies symptom.
- Enable extra packages with -D PKG_<NAME>=yes, for example -D PKG_REAXFF=yes, -D PKG_MEAM=yes, -D PKG_EXTRA-COMPUTE=yes. To enable most common packages, use -C ../cmake/presets/most.cmake.
- After changing the package selection, rerun cmake in the same build directory and rebuild; existing settings are kept in CMakeCache.txt.
Step 3
Both run EAM on NVIDIA GPUs, but they split the work differently: the GPU package puts the pair computation on the GPU and leaves the rest on the CPU, while KOKKOS tries to keep the whole time step on the GPU.
The default for GPU_API is opencl. Setting only -D PKG_GPU=on gives an OpenCL build. It runs on NVIDIA GPUs too, but it is not the CUDA back end. This is the setting online tutorials most often leave out.
Set GPU_ARCH and Kokkos_ARCH from the GPU compute capability: V100 is sm_70 / VOLTA70; T4 and RTX 20 series are sm_75 / TURING75; A100 is sm_80 / AMPERE80; RTX 30 series and A40 are sm_86 / AMPERE86; RTX 40 series and L40S are sm_89 / ADA89 (not sm_90); H100 is sm_90 / HOPPER90; RTX 50 series is sm_120 / BLACKWELL120. CUDA 13 can no longer compile offline for Maxwell, Pascal or Volta. The 22Jul2025 default is sm_50, so with CUDA 13 you must set GPU_ARCH explicitly.
A KOKKOS executable moved to a GPU with a different major architecture version (for example from 7.x to 8.x) exits with an error; within the same major version it only adds a JIT compilation at startup. To support several GPU types, build for the oldest one or build one executable per type.
# Look up the compute capability first: 8.6 -> sm_86, 8.9 -> sm_89
nvidia-smi --query-gpu=name,compute_cap --format=csv
# Option A: GPU package (CUDA back end)
mkdir build-gpu && cd build-gpu
cmake -C ../cmake/presets/basic.cmake \
-D PKG_GPU=on -D GPU_API=cuda -D GPU_ARCH=sm_86 -D GPU_PREC=mixed \
-D BUILD_MPI=yes -D BUILD_OMP=yes -D PKG_OPENMP=yes \
-D CMAKE_INSTALL_PREFIX=$HOME/.local \
../cmake
cmake --build . -j 16
# Option B: KOKKOS (CUDA back end); the two -C presets can be combined
mkdir build-kokkos && cd build-kokkos
cmake -C ../cmake/presets/basic.cmake \
-C ../cmake/presets/kokkos-cuda.cmake \
-D Kokkos_ARCH_AMPERE86=yes \
-D BUILD_MPI=yes \
../cmake
cmake --build . -j 16| Aspect | GPU package | KOKKOS |
|---|---|---|
| What is accelerated | pair and parts of PPPM; integration and most fixes stay on the CPU | every pair, fix and compute with a /kk variant runs on the GPU; commands without a /kk variant copy data between CPU and GPU |
| MPI ranks per GPU | several ranks sharing one GPU are usually faster; the manual suggests 2–10 | 1 by default; sharing a GPU among ranks requires CUDA MPS |
| Precision | single / mixed (default) / double | double by default; single and mixed also supported since 30Sep2026 |
| Best for | many CPU cores and few GPUs; the OpenCL build on Windows | large systems on one GPU; the manual says KOKKOS is usually faster with many atoms in double precision |
| Run flags | -sf gpu -pk gpu N | -k on g N -sf kk |
| Key build options | -D PKG_GPU=on -D GPU_API=cuda -D GPU_ARCH=sm_XX | -C kokkos-cuda.cmake -D Kokkos_ARCH_<GPU>=yes |
- Common practice (several MatSci forum threads): with CMake, do not also run make yes-gpu in src or build lib/gpu separately. CMake stops with an error when the two build systems are mixed; fix it by running make no-all purge in src or unpacking the source again, and use only -D PKG_GPU=on.
- Common practice: a runtime error like CUDA driver version is insufficient usually means the CUDA Toolkit is newer than the driver. nvidia-smi shows the highest CUDA version the driver supports in its top-right corner, and nvcc --version shows the Toolkit used for the build; the former should not be lower than the latter.
- If nvcc is not found during the build, add the CUDA bin directory to PATH (for example export PATH=/usr/local/cuda/bin:$PATH), delete the build directory and rerun cmake.
Step 4
# CPU only, 8 MPI ranks
mpirun -np 8 lmp -in in.cu_tensile
# CPU only, OPENMP package, 1 process with 8 threads (also works with the Windows GUI installer)
lmp -sf omp -pk omp 8 -in in.cu_tensile
# GPU package: 4 MPI ranks share 1 GPU
mpirun -np 4 lmp -sf gpu -pk gpu 1 -in in.cu_tensile
# KOKKOS: 1 MPI rank per GPU
lmp -k on g 1 -sf kk -in in.cu_tensile
# On Pascal and newer GPUs, half neighbor lists are often faster
lmp -k on g 1 -sf kk -pk kokkos newton on neigh half -in in.cu_tensile
# Multiple GPUs with an MPI that is not GPU-aware (common with distribution OpenMPI)
mpirun -np 2 lmp -k on g 2 -sf kk -pk kokkos gpu/aware off -in in.cu_tensile- Use -in filename rather than < redirection. With some MPI implementations, < does not deliver the input in parallel runs.
- The conda mpich variant ships its own mpirun (MPICH Hydra) in the environment's bin folder. After activating the environment, run mpirun -np 8 lmp -in in.cu_tensile directly and do not mix in a system OpenMPI. The MPI line of lmp -h shows MPICH Version. Tested on macOS for this page: with a VPN active, MPICH picks the utun virtual interface and fails with OFI poll failed (default nic=utun...); export FI_PROVIDER=tcp first and it runs normally.
- The conda CPU package contains neither the GPU nor the KOKKOS package. Tested on this page: -k on fails with Cannot use -kokkos on without KOKKOS installed, and -sf gpu -pk gpu 1 fails with Package gpu command without GPU package installed; -sf omp -pk omp N works.
- -sf gpu appends the suffix to every command that has a /gpu variant, so eam/alloy becomes eam/alloy/gpu; -sf kk works the same way. You can switch between CPU and GPU from the command line without editing the input file.
- KOKKOS assumes a GPU-aware MPI by default. A single rank on a single GPU is unaffected. If multiple ranks segfault, or you see Turning off GPU-aware MPI since it is not detected, add -pk kokkos gpu/aware off.
- Mixed precision is the GPU package default. Before publishing, run the same simulation for a while with a double-precision build or the CPU version and compare energy and stress.
- To estimate run time, first set run to 2000 steps, read timesteps/s on the Performance line at the end of the log, and extrapolate to the full step count.
Step 5
The system is a periodic bulk Cu crystal of 32000 atoms with the Mishin 2001 EAM potential. The workflow is: energy minimization → equilibration at 300 K and zero pressure for 20 ps → tension along x at 1×10⁹ s⁻¹ for 300 ps with y and z held at zero pressure.
Tested on this page (LAMMPS 22Jul2025 update6, conda-forge CPU build, macOS arm64): with region changed to 0 10 0 10 0 10 (4000 atoms) and all other parameters unchanged, the full 20 ps equilibration and 300 ps tension were run. After equilibration at 300 K, lx is 36.31 Å, a lattice constant of 3.631 Å. plot_ss.py gives a 0–1% strain modulus of 67.3 GPa, which agrees with the theoretical [100] Young's modulus: the same potential at 0 K and 0.1% strain gives C11 = 169.7 GPa and C12 = 122.3 GPa (the Mishin 2001 paper reports 169.9 and 122.6 GPa), and E = (C11−C12)(C11+2C12)/(C11+C12) gives 67.2 GPa. The first yield stress is 8.80 GPa at a strain of 0.113, after which the stress drops abruptly to about 4.3 GPa; σyy and σzz stay within ±0.3 GPa, showing that the zero-pressure control in y and z works. With 864 atoms (6×6×6 cells) the results are 68.6 GPa and 8.85 GPa (strain 0.115). This yield stress is the homogeneous nucleation stress of a defect-free periodic crystal and is about two orders of magnitude above macroscopic experimental values, which is normal for an ideal-crystal model.
Single-core cost (same version and machine, 4000 atoms, 1000 steps each, 99.9% CPU use): the NPT equilibration runs at 204 steps/s (0.82 M atom-step/s); the tension stage with fix deform, fix print and dump runs at 156 steps/s, i.e. 6.4 ms per step (0.62 M atom-step/s). Extrapolated, 32000 atoms take about 51 ms per step on one core, and the 320,000 steps of equilibration plus tension take about 4.5 hours. Other jobs were running on the machine during testing, so multi-process speedups were distorted and are not reported here.
# in.cu_tensile -- uniaxial tension of single-crystal Cu along [100]
# Works with LAMMPS 22Jul2025 and 30Sep2026; requires the MANYBODY package
# ---------- 1. Initialization ----------
units metal # length Å, time ps, energy eV, pressure bar
dimension 3
boundary p p p
atom_style atomic
variable T equal 300.0 # temperature in K
variable erate equal 1.0e9*1.0e-12 # 1e9 s^-1 converted to 1/ps
# ---------- 2. Build the model ----------
lattice fcc 3.615 # matches the Cu lattice constant in the potential file
region box block 0 20 0 20 0 20 # 20x20x20 unit cells, 32000 atoms
create_box 1 box
create_atoms 1 box
# ---------- 3. Potential ----------
pair_style eam/alloy
pair_coeff * * Cu_mishin1.eam.alloy Cu # file ships in the LAMMPS potentials/ folder
neighbor 2.0 bin
neigh_modify delay 0 every 1 check yes
# ---------- 4. Energy minimization, relaxing box pressure to 0 ----------
thermo 100
thermo_style custom step pe press lx ly lz
fix relax all box/relax iso 0.0 vmax 0.001
min_style cg
minimize 1.0e-12 1.0e-12 10000 100000
unfix relax
# ---------- 5. Equilibrate at 300 K, 0 bar for 20 ps ----------
reset_timestep 0
timestep 0.001 # 1 fs
velocity all create ${T} 4928459 mom yes rot yes dist gaussian
fix eq all npt temp ${T} ${T} 0.1 iso 0.0 0.0 1.0
thermo 1000
thermo_style custom step temp pe press lx ly lz
run 20000
unfix eq
# ---------- 6. Constant strain-rate tension along x, y and z held at 0 bar ----------
variable tmp equal lx
variable L0 equal ${tmp} # evaluated immediately: length before loading
reset_timestep 0
fix nptyz all npt temp ${T} ${T} 0.1 y 0.0 0.0 1.0 z 0.0 0.0 1.0 drag 1.0
fix pull all deform 1 x erate ${erate} units box remap x
variable strain equal (lx-v_L0)/v_L0
variable sxx equal -pxx/10000 # bar to GPa; sign flipped so tension is positive
variable syy equal -pyy/10000
variable szz equal -pzz/10000
fix out all print 100 "${strain} ${sxx} ${syy} ${szz}" file stress_strain.txt screen no title "# strain sxx_GPa syy_GPa szz_GPa"
compute peatom all pe/atom
dump traj all custom 2000 tensile.lammpstrj id type x y z c_peatom
thermo_style custom step v_strain temp v_sxx v_syy v_szz pe press
restart 50000 tensile.*.restart
run 300000 # 300 ps, engineering strain reaches 0.30
write_data final.data
- 01
units metal
Sets the units for every number that follows: distance in Å, time in ps, energy in eV, pressure in bar. Most EAM, Tersoff and SW potential files are given in metal units.
- 02
lattice fcc 3.615 and region/create_box/create_atoms
lattice sets both the lattice and the length unit, so 0 20 in region means 20 unit cells, about 72 Å. Use the equilibrium lattice constant of the potential, otherwise there is residual stress from the start. The header of Cu_mishin1.eam.alloy gives 3.615 Å.
- 03
pair_style eam/alloy and pair_coeff * * file Cu
eam/alloy reads setfl files (.eam.alloy). It takes exactly one pair_coeff * * line followed by element names in atom-type order, one name per atom type. The setfl file contains the mass, so no mass command is needed. .eam files (funcfl format) go with pair_style eam. Mixing the two reads the file incorrectly.
- 04
neighbor and neigh_modify
The default skin in metal units is 2.0 Å. Since 2Aug2023 the neigh_modify default is delay 0 every 1 check yes. The delay 10 in older tutorials can cause Dangerous builds warnings; remove it.
- 05
fix box/relax + minimize
Adjusts the box size during minimization so the 0 K pressure is zero. vmax 0.001 limits the volume change per iteration and keeps the minimization from diverging.
- 06
velocity + fix npt iso
Equilibrates at 300 K and 0 bar for 20 ps so the lattice reaches its thermal expansion. Skipping this step puts thermal stress at the start of the curve. The damping parameters 0.1 ps (Tdamp) and 1.0 ps (Pdamp) are 100 and 1000 steps, the order of magnitude the manual recommends.
- 07
variable L0 equal ${tmp}
${} is evaluated when the line is read, so L0 stays fixed at the length before loading. With variable L0 equal lx, L0 would change with the box and the strain would always be 0.
- 08
fix npt y/z + fix deform x erate
deform stretches x at a constant engineering strain rate. erate is in 1/time units, which is 1/ps in metal units, so 1×10⁹ s⁻¹ is written 0.001. y and z are held at 0 bar with npt, allowing Poisson contraction and giving a uniaxial stress state. With deform alone and fixed y and z, you get uniaxial strain and higher stresses. Do not barostat the pulled direction x at the same time, because the barostat counteracts the deformation; this is one of the most common causes of odd tensile curves on the MatSci forum.
- 09
variable sxx equal -pxx/10000
pxx is a component of the pressure tensor in bar and is negative under tension. The minus sign makes tensile stress positive, and dividing by 10000 converts to GPa. When a variable uses pxx, pyy or pzz, thermo_style must contain a pressure keyword such as press so that LAMMPS computes the pressure tensor. Tested on this page: with press removed from the end of thermo_style, the tensile run stops immediately with ERROR: Thermo keyword pxx in variable requires thermo to use/init press.
- 10
fix print, dump, restart
The fix print string is in double quotes, so ${strain} inside it is evaluated at each output. dump writes a .lammpstrj file that OVITO and VMD open directly. restart writes a restart file every 50000 steps.
Variant
Nanowire tension is the most common modification of the bulk script. Three things change: the model, the barostat directions, and the volume used to normalize stress.
Keep periodic boundaries in y and z, but leave a vacuum layer much wider than the cutoff (5.5 Å for the Mishin potential) so the wire does not interact with its periodic images. In the example the box is about 72 Å wide, the wire about 40 Å in diameter, and the vacuum about 32 Å.
pxx is the virial stress of the whole box divided by the box volume. The box volume includes vacuum, so plain -pxx/10000 clearly underestimates the stress in the wire. The conversion above assumes the cross-section does not change during loading. Alternatively, sum compute stress/atom and divide by the number of atoms times the volume per atom. Whichever you use, state the definition of the normalization volume in the paper.
Tested on this page (same version; length shortened to 20 cells, 7540 atoms; 4 ps of equilibration at 300 K, then tension at 1×10⁹ s⁻¹ to a strain of 0.03): after equilibration lx is 71.52 Å, 1.1% shorter than the built length of 72.30 Å because surface stress contracts the wire axially, so take L0 after equilibration. The modulus fitted over 0–1% strain is 55.8 GPa with the volume conversion above and only 13.2 GPa with plain -pxx/10000; the ratio of about 4.2 equals the ratio of the box cross-section to the wire cross-section. Summing compute stress/atom and dividing by the number of atoms times the atomic volume a³/4 (corrected for the axial elongation) gives a stress within 0.5% of the conversion above. stress/atom has the opposite sign to pressure, so do not negate the sum again.
# Replace step 2 with these lines: [100] Cu nanowire, about 4 nm in diameter and 20 nm long
lattice fcc 3.615
region box block 0 55 -10 10 -10 10 # lattice units; integer number of cells along x
create_box 1 box
region wire cylinder x 0 0 5.5 INF INF # radius of 5.5 lattice constants, about 19.9 Å
create_atoms 1 region wire
# Step 4: relax the box along x only
fix relax all box/relax x 0.0 vmax 0.001
# Step 5: barostat along x only
fix eq all npt temp ${T} ${T} 0.1 x 0.0 0.0 1.0
# Step 6: y and z are vacuum, so no barostat there; use NVT
fix nvt1 all nvt temp ${T} ${T} 0.1
fix pull all deform 1 x erate ${erate} units box remap x
# pxx is normalized by the whole box volume (including vacuum); rescale to the wire volume
variable R equal 5.5*3.615
variable sxx equal -pxx*vol/(PI*v_R^2*lx)/10000 # GPa, ignoring Poisson contractionPotentials
You do not need to download a Cu EAM potential: Cu_mishin1.eam.alloy (Mishin et al., PRB 63, 224106, 2001), Cu_u3.eam (Foiles 1986) and Cu_zhou.eam.alloy are all in the LAMMPS potentials/ folder. On the NIST Interatomic Potentials Repository (ctcms.nist.gov/potentials/system/Cu), the same Mishin potential is listed as Cu01.eam.alloy. NIST notes that the Mendelev 2013 Cu2 potential improves stacking fault energies, and that the ipr2 version of Zhou 2004 corrects spurious fluctuations in the tabulated functions. For dislocation and twinning studies, choose a potential with these notes in mind and cite the source in your methods.
A conda installation has no such folder; download the file from the GitHub tag that matches your version: curl -LO https://raw.githubusercontent.com/lammps/lammps/stable_22Jul2025_update6/potentials/Cu_mishin1.eam.alloy, and put it next to the input file, or in a fixed folder that LAMMPS_POTENTIALS points to.
OpenKIM potentials need no file download; reference the model ID in the input: kim init EAM_Dynamo_MishinMehlPapaconstantopoulos_2001_Cu__MO_346334655118_006 metal, then kim interactions Cu after the model is built. This requires the KIM package. The conda and apt builds include KIM; the Windows prebuilt packages do not.
Recent versions read the UNITS tag on the first line of the potential file. If the units in the input file do not match, LAMMPS reports Potential file ... requires metal units but real units are in use. Fix it by changing units or by using a potential file in the matching units.
| System | Common potentials (pair_style) | Packages | Units | Sources |
|---|---|---|---|---|
| Metals and alloys | eam, eam/alloy, eam/fs, meam | MANYBODY, MEAM | metal | NIST IPR, OpenKIM, LAMMPS potentials/ folder |
| Covalent crystals (Si, Ge, SiC) | sw, tersoff | MANYBODY | metal | potentials/ folder, NIST IPR, OpenKIM |
| Carbon materials (graphene, CNT, hydrocarbons) | airebo, tersoff | MANYBODY | metal | potentials/CH.airebo |
| Polymers and organic molecules (fixed topology) | lj/cut/coul/long + bond, angle, dihedral terms (OPLS-AA, GAFF); lj/class2 (PCFF, COMPASS) | MOLECULE, KSPACE, CLASS2 | real | Data files generated with moltemplate, LigParGen or msi2lmp |
| Reactive systems (combustion, oxidation, bond breaking and forming) | reaxff + fix qeq/reaxff; or fix bond/react | REAXFF; REACTION | real | ReaxFF parameters in the potentials/ folder and in paper supplements |
| Near-DFT accuracy | snap, pace, mliap | ML-SNAP, ML-PACE, ML-IAP | metal | Files supplied with papers, OpenKIM |
Units
| Quantity | metal | real |
|---|---|---|
| Distance | Å | Å |
| Time | ps (default timestep 0.001 = 1 fs) | fs (default timestep 1.0 = 1 fs) |
| Energy | eV | kcal/mol |
| Pressure | bar (1 GPa = 10000 bar) | atm (1 GPa = 9869.2 atm) |
| fix deform erate | 1/ps | 1/fs |
| Default neighbor skin | 2.0 Å | 2.0 Å |
- When switching from metal to real, multiply the erate value by a further 10⁻³: 1×10⁹ s⁻¹ is 1.0e-6 (1/fs) in real units.
- compute stress/atom is in pressure×volume units (bar·Å³) and has the opposite sign of pressure. It is not a per-atom stress; divide by a per-atom volume before reading it as stress.
- press in thermo is (pxx+pyy+pzz)/3, which is about -σxx/3 in uniaxial tension. press cannot stand in for σxx.
- Strain rates in MD tension are typically 10⁷–10¹⁰ s⁻¹, many orders of magnitude above experiment, so yield stresses come out high. Common practice: before production, run once at 10⁹ and once at 10⁸ s⁻¹ to see how much the yield stress changes, then decide whether a lower rate is worth the cost. When comparing sizes or temperatures, keep the strain rate the same and state it in the paper.
Step 6
In log.lammps, the thermo block of each run starts with a header line beginning with Step and ends with a line starting Loop time of. The header columns are the thermo_style keywords, and variables appear as v_strain and v_sxx. After Loop time comes a timing breakdown (shares of Pair, Neigh, Comm, Output and so on): a low Pair share with a high Comm or Output share means too many processes or too frequent output.
OVITO opens tensile.lammpstrj directly (and also reads .gz). Add the Common neighbor analysis modifier to the pipeline to color atoms as FCC, HCP or Other. In FCC copper, a single layer of HCP atoms marks a twin boundary and two adjacent HCP layers mark an intrinsic stacking fault. VMD opens the file with vmd -lammpstrj tensile.lammpstrj; it cannot read compressed files.
Tested on this page: running plot_ss.py on the 4000-atom case above prints E(0-1%) = 67.3 GPa and first yield = 8.80 GPa at strain 0.113. Small cells reload after the first yield; the 864-atom case reaches a higher peak of 9.9 GPa at a strain of 0.23, so the script takes the maximum before the first clear drop instead of the global maximum. The runs read by lammps.formats.LogFile correspond to minimization, equilibration and tension in that order, and the keys of runs[-1] are Step, v_strain, Temp, v_sxx, v_syy, v_szz, PotEng and Press.
Dislocation density versus strain is a common post-processing request in the Chinese community; several CSDN articles demonstrate it, and the scripts are mostly shared in QQ groups. The dxa_density.py script below uses only the public interface of the OVITO Python module (pip install ovito): DislocationAnalysisModifier outputs the total line length DislocationAnalysis.total_line_length (Å) and the cell volume DislocationAnalysis.cell_volume (ų); dividing them and multiplying by 10²⁰ gives m⁻². DislocationAnalysis.length.1/6<112> is the length of Shockley partials.
Tested on this page (OVITO 3.16.1, the 151-frame tension trajectory of the 4000-atom case above, about 11 s): DXA finds no dislocations before a strain of 0.116, while the FCC fraction is 0.936 at a strain of 0.100 and drops to 0.625 at 0.116. The first dislocations appear at a strain of 0.118 with a density of 4.1×10¹⁷ m⁻², and the HCP fraction rises to 0.149. After that the density jumps between 0 and 1.4×10¹⁸ m⁻² and is 9.6×10¹⁷ m⁻² at a strain of 0.300, while the HCP fraction stays between 0.15 and 0.33. The jumps come from the small box: with an edge of about 36 Å, one dislocation line spanning the box corresponds to about 8×10¹⁶ m⁻², so nucleation and annihilation across periodic boundaries change single-frame values several-fold. To compare dislocation densities in a paper, use a larger box or average over a strain interval.
# plot_ss.py: read the fix print output, plot the stress-strain curve, fit the initial modulus and find the first yield
import numpy as np
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
d = np.loadtxt("stress_strain.txt", comments="#")
strain, sxx = d[:, 0], d[:, 1]
# 20-point moving average to suppress thermal noise (100 steps per point); valid mode avoids the false drop caused by zero padding at both ends
w = 20
sxx_s = np.convolve(sxx, np.ones(w) / w, mode="valid")
strain_s = strain[w // 2 : w // 2 + len(sxx_s)]
lin = strain <= 0.01 # fit only 0–1% strain
E = np.polyfit(strain[lin], sxx[lin], 1)[0]
# First yield: the maximum before the stress first falls 20% below its running maximum (after exceeding 1 GPa)
# Small cells reload after yielding and the global maximum can come later, so plain argmax is not enough
run_max = np.maximum.accumulate(sxx_s)
drop = np.nonzero((run_max > 1.0) & (run_max - sxx_s > 0.2 * run_max))[0]
j = drop[0] if drop.size else len(sxx_s)
i = np.argmax(sxx_s[:j])
print(f"E(0-1%) = {E:.1f} GPa")
print(f"first yield = {sxx_s[i]:.2f} GPa at strain {strain_s[i]:.3f}")
plt.plot(strain, sxx, lw=0.5, alpha=0.4, label="raw")
plt.plot(strain_s, sxx_s, lw=1.5, label=f"moving avg ({w})")
plt.xlabel("Engineering strain")
plt.ylabel("Stress σxx (GPa)")
plt.legend()
plt.tight_layout()
plt.savefig("stress_strain.png", dpi=200)
# You can also parse log.lammps directly (needs the LAMMPS Python module; the conda build includes it)
# from lammps.formats import LogFile
# run = LogFile("log.lammps").runs[-1] # keys are the thermo column names, e.g. "v_strain", "v_sxx"
# dxa_density.py: run DXA on every frame and print dislocation density and FCC/HCP fractions
import sys
from ovito.io import import_file
from ovito.modifiers import DislocationAnalysisModifier
pipeline = import_file(sys.argv[1] if len(sys.argv) > 1 else "tensile.lammpstrj")
dxa = DislocationAnalysisModifier(input_crystal_structure=DislocationAnalysisModifier.Lattice.FCC)
pipeline.modifiers.append(dxa)
print("# step rho_total(1/m^2) rho_shockley(1/m^2) FCC_frac HCP_frac")
for frame in range(pipeline.source.num_frames):
data = pipeline.compute(frame)
a = data.attributes
vol = a["DislocationAnalysis.cell_volume"] # Å^3
rho = a["DislocationAnalysis.total_line_length"] / vol * 1e20 # Å/Å^3 -> m^-2
rho_s = a.get("DislocationAnalysis.length.1/6<112>", 0.0) / vol * 1e20
n = data.particles.count
fcc = a["DislocationAnalysis.counts.FCC"] / n
hcp = a["DislocationAnalysis.counts.HCP"] / n
print(f"{a['Timestep']:8d} {rho:.3e} {rho_s:.3e} {fcc:.3f} {hcp:.3f}")
Continuation
Restart files store atoms, the box, units, atom_style and similar settings. What they do not store must be re-specified in the continuation script: all fixes and computes, and pair_coeff for many-body potentials that read parameters from files (EAM, Tersoff and others).
When fix deform uses the same fix-ID as the original script (pull here), it restores the box from the start of loading from the restart file, but only uses it when the run command has start and stop. start is the step at which the original loading run began (0 after reset_timestep 0), and stop is the total step count. With only run 300000 upto, deform recomputes the elongation from the box length at the restart. Tested on this page (an 864-atom, 1×10¹⁰ s⁻¹ shortened case, continued from strain 0.10 to the planned final step): with upto alone the final strain is 0.32; with start 0 stop it is 0.30, and the final stress of 6.80 GPa agrees with 6.78 GPa from the uninterrupted run. Continuing the 4000-atom case from the “Your first input file” section from step 150000 for 1000 steps with this script (only L0 changed) gives the same strain as the original run and stresses that agree to 7 significant digits. Continuing the page's parameters from step 150000 without start/stop would end at a strain of 0.3225. L0 must be the length before the original loading, not lx at the restart, otherwise the strain starts again from 0. run N upto stops at total step N, so any number of interruptions still end at the same point.
# in.cu_tensile.continue -- continue tension from the restart at step 150000
read_restart tensile.150000.restart
pair_style eam/alloy # many-body potentials are not stored in restart files; re-specify
pair_coeff * * Cu_mishin1.eam.alloy Cu
neighbor 2.0 bin
variable T equal 300.0
variable erate equal 1.0e9*1.0e-12
variable L0 equal 72.5 # set to the number after variable L0 equal in the original log
timestep 0.001
fix nptyz all npt temp ${T} ${T} 0.1 y 0.0 0.0 1.0 z 0.0 0.0 1.0 drag 1.0
fix pull all deform 1 x erate ${erate} units box remap x # same fix-ID as the original script
variable strain equal (lx-v_L0)/v_L0
variable sxx equal -pxx/10000
fix out all print 100 "${strain} ${sxx}" append stress_strain.txt screen no
thermo_style custom step v_strain temp v_sxx pe press
thermo 1000
restart 50000 tensile.*.restart
run 300000 upto start 0 stop 300000 # run to total step 300000; start/stop keep deform on the original box lengthErrors
Causes and fixes for Out of range atoms, Too many neighbor bins, NaN and other errors are covered in LAMMPS Common Errors and Fixes (/learn/lammps-errors).
| Error message (excerpt) | Cause | Fix |
|---|---|---|
| Unrecognized pair style 'eam/alloy' is part of the MANYBODY package which is not enabled in this LAMMPS binary | MANYBODY was not enabled at build time | Rerun cmake with -D PKG_MANYBODY=yes, or use the basic.cmake preset |
| cannot open eam/alloy potential file Cu_mishin1.eam.alloy: No such file or directory | The potential file is neither in the current directory nor in the directory LAMMPS_POTENTIALS points to | Copy the file next to the input file, or export LAMMPS_POTENTIALS=<install prefix>/share/lammps/potentials; the conda build has no such folder, so download the potential from GitHub first (see the potentials section) |
| Potential file Cu_mishin1.eam.alloy requires metal units but real units are in use | units does not match the potential file | Switch to units metal and rewrite timestep and erate in metal units |
| ERROR: Lost atoms: original 32000 current 31998 | Time step too large, overlapping atoms in the initial structure, or wrong potential or units | Check the lattice constant and timestep first; see the LAMMPS errors page for the full checklist |
Working from China
This section collects practices that recur on the Computational Chemistry Commune forum (keinsci), CSDN, Zhihu and university HPC documentation, checked against the current documentation or source code of LAMMPS and the related tools. Numbers marked as tested on this page come from actual queries or runs on 2026-10-10.
# ~/.condarc: route conda-forge through the Tsinghua TUNA mirror (community-only setup from the TUNA help page)
channels:
- conda-forge
- nodefaults
show_channel_urls: true
custom_channels:
conda-forge: https://mirrors.tuna.tsinghua.edu.cn/anaconda/cloud
conda clean -i
# build string = accelerator + MPI: cpu, cuda129, cuda130, cuda134 combined with nompi, mpich, openmpi
conda create -n lammps -y "lammps=2025.07.22=cpu*openmpi*" # about 410 MB total for a new env
conda create -n lammps-gpu -y "lammps=2025.07.22=cuda129*openmpi*" # about 1 GB total for a new env
conda activate lammps && lmp -h | head -20
# 1) On a machine with internet access, download the third-party archive (URL as in cmake/Modules/Packages/VORONOI.cmake)
curl -LO https://download.lammps.org/thirdparty/voro++-0.4.6.tar.gz
# 2) Copy it to the cluster and point <prefix>_URL at the local file; the default SHA256 is still checked, so the version must match
cd lammps/build
cmake ../cmake -C ../cmake/presets/basic.cmake -D PKG_VORONOI=yes -D DOWNLOAD_VORO=yes \
-D VORO_URL=file://$HOME/src/voro++-0.4.6.tar.gz
cmake --build . -j 8
# Many downloads: run tools/offline/init_caches.sh on the online machine (default cache ~/.cache/lammps),
# copy the cache to the same path on the cluster, source tools/offline/use_caches.sh and add the -D LAMMPS_DOWNLOADS_URL=... options it prints
# VMD Tk Console: convert a Packmol pdb to a LAMMPS data file (TopoTools)
package require topotools
mol new system.pdb
# for pdb files VMD takes the type from the atom name, so C1, C2, H1 become separate types; merge them by element first
[atomselect top {name "C.*"}] set type C
[atomselect top {name "H.*"}] set type H
pbc set {40.0 40.0 40.0} ;# writelammpsdata refuses to write without a box
topo writelammpsdata system.data full
# type IDs follow the alphabetical order of type names: C=1, H=2; the comment at the end of each Masses line is the type name
Use the Tsinghua conda mirror and pick the variant with a build string
When conda-forge downloads are slow in China, follow the Tsinghua TUNA mirror help page and point conda-forge at mirrors.tuna.tsinghua.edu.cn/anaconda/cloud in ~/.condarc. Tested on this page (micromamba query of the mirror on 2026-10-10): linux-64 has lammps 2025.07.22 in four build families, cpu, cuda129, cuda130 and cuda134, each split into nompi, mpich and openmpi. With a bare conda install lammps the solver decides which MPI to use and whether to include CUDA; a build string such as "lammps=2025.07.22=cpu*openmpi*" pins the variant. Total download for a new environment: about 410 MB for cpu+openmpi and about 1 GB for cuda129+openmpi.
On a university cluster, check the modules before building your own
University clusters usually provide LAMMPS already, and the executable name differs between modules. The SCUT platform provides a source-built lammps/2Aug2023 (executable lmp_intel_cpu_intelmpi) and a conda environment lammps20230802-cpu (executable lmp). SJTU's Siyuan-1 provides CPU modules such as lammps/20230328-intel-2021.4.0-omp and lammps/20230802-intel-2021.4.0-gpu, with the example run command mpirun lmp -pk intel 0 omp 2 -sf intel -i in.lj; drop -pk intel and -sf intel when the force field is not supported by the INTEL package. Run module avail lammps first, then lmp -h after loading to see the installed packages. These modules are mostly 2023 versions, so new keywords written from the 30Sep2026 documentation may not be recognized; when that happens, check the documentation for that version. The SJTU documentation requires builds to run on compute nodes, since parallel compilation is not allowed on login nodes.
Third-party downloads fail during the build
With VORONOI, PLUMED, ML-PACE, KIM and similar packages enabled, CMake downloads source archives from GitHub, download.lammps.org or OpenKIM. On an unstable campus network or a cluster without internet access, CMake reports Each download failed! (the exact text in a CSDN write-up of a build on an offline cluster). That author placed the archive by hand and faked CMake's step stamp files. A more direct fix is to point the cache variable <prefix>_URL at a local file, for example VORO_URL, PLUMED_URL, PACELIB_URL or KIM_URL; the prefix and default URL are on the SetDownloadSettings line in cmake/Modules/Packages/<package>.cmake. The GPU package and KOKKOS libraries ship with the LAMMPS source (DOWNLOAD_KOKKOS is off by default), so a GPU build does not need to download either of them.
Pitfalls when converting from Packmol, VMD TopoTools or Materials Studio to a data file
Packmol only keeps atoms inside the box at least tolerance apart; it ignores periodic images. In his Packmol tutorial on the keinsci forum, Sobereva uses add_box_sides 1.2 to add 1.2 Å to the box in each direction; from Packmol 20.15.0 the pbc keyword packs directly into a periodic box. VMD TopoTools' topo writelammpsdata fails with need to have non-zero box sizes when no box is set, so run pbc set first; it assigns type IDs in alphabetical order of the VMD type names, and for pdb files the type name is taken from the atom name. Data files exported through OVITO after building in Materials Studio also often have a type order that differs from the force-field file, and a keinsci thread reorders the type IDs by hand in Word and Excel. For potentials written as pair_coeff * * file element…, such as eam/alloy, tersoff and reaxff, the data file does not need to change: list the element names in the order of data-file types 1, 2, 3 and so on. The example in the LAMMPS manual: with elements ordered C H O N in the ffield file and types 1 to 4 being C, C, N, H, write pair_coeff * * ffield.reax C C N H.
Hand it to an agent
Example one-sentence instruction: "Use LAMMPS to simulate uniaxial tension of a [100] single-crystal Cu nanowire about 4 nm in diameter and 20 nm long at 300 K and 1×10⁹ s⁻¹, with the Mishin EAM potential and GPU acceleration; output the stress-strain curve, Young's modulus, yield strength and common neighbor analysis snapshots, and compare with a nanowire about 2 nm in diameter to show the size effect."
The agent carries out these steps: rents a GPU and runs the environment self-test; writes the input files for model building, relaxation and tension; runs a short test to estimate run time before submitting the production run; extracts stress-strain data with Python and fits the modulus; marks dislocations and twins with common neighbor analysis; and reviews the results adversarially to confirm the task is actually done. The deliverables are the input files, post-processing scripts, logs, plots, structure snapshots, an environment record and a report. Compute is released when the run finishes.
You still need to check these yourself: whether the potential suits the deformation mechanism you study; the definition of the stress normalization volume; whether the effects of strain rate and size on yield strength are stated in the conclusions; and whether repeats with different random seeds are needed to estimate statistical error.
References
- LAMMPS stable_30Sep2026 release notes (GitHub) — C++17, CMake 3.20, KOKKOS C++20, GRAPHICS package and removed packages
- LAMMPS manual: Build with CMake — Minimum CMake version and preset files
- LAMMPS manual: Packages with extra build options (GPU, KOKKOS) — GPU_API and GPU_ARCH defaults, Kokkos architecture IDs
- LAMMPS manual: GPU package — -sf gpu -pk gpu usage and MPI ranks per GPU
- LAMMPS manual: KOKKOS package — -k on g usage, GPU-aware MPI, C++20 and CUDA 12.2 requirements
- LAMMPS-GUI installation docs — Prebuilt packages include the OpenCL GPU package, no MPI, Windows signing prompts
- LAMMPS Windows installer notes — MS-MPI builds and OpenMP usage
- conda-forge lammps-feedstock — Package list of the conda build; CUDA builds use KOKKOS
- NVIDIA CUDA on WSL User Guide — No Linux driver inside WSL; install only the cuda-toolkit metapackage
- LAMMPS manual: pair_style eam — funcfl vs setfl formats, pair_coeff syntax, potential repositories
- LAMMPS manual: fix deform — erate units, barostatting the two lateral directions during tension, restart behavior
- LAMMPS manual: compute stress/atom — Units of pressure×volume, sign opposite to pressure
- NIST Interatomic Potentials Repository: Cu — List of Cu potentials and OpenKIM IDs
- Mishin et al., Phys. Rev. B 63, 224106 (2001) — Original paper for Cu_mishin1.eam.alloy
- MatSci forum: mpirun launches multiple runs — Diagnosing N serial copies when MPI is not linked
- MatSci forum: fix npt + fix deform — Experience thread: which directions to barostat during tension, and MD strain-rate magnitudes
- MatSci forum: Regarding uniaxial tensile test — Experience thread: do not barostat the pulled direction
- MatSci forum: GPU package compilation — Experience thread: mixing CMake with make yes-gpu fails; GPU_API defaults to OpenCL
- MatSci forum: There is an error with the GPU implemented — Experience thread: CUDA Toolkit newer than the driver breaks GPU runs
- Gravelle et al., A Set of Tutorials for the LAMMPS Simulation Package, LiveCoMS 6, 3037 (2025) — LAMMPS-GUI based tutorial collection co-written by LAMMPS developers
- Tsinghua TUNA mirror: Anaconda mirror help — custom_channels setup for conda-forge and conda clean -i
- SCUT scientific computing platform: LAMMPS — Module and conda-environment deployments, executable names and Slurm scripts
- SJTU HPC documentation: LAMMPS — Module versions per cluster, -pk intel -sf intel usage, building on compute nodes
- CSDN: building LAMMPS with the quip package on an offline cluster without root — Experience post: CMake Each download failed! on an offline cluster and placing archives by hand
- LAMMPS source: tools/offline — init_caches.sh and use_caches.sh for offline build caches
- Sobereva: using Packmol to build initial MD structures (keinsci forum) — Experience post: tolerance 2.0 and add_box_sides 1.2 to avoid bad contacts across periodic boundaries
- Packmol User's Guide — pbc keyword available from 20.15.0
- TopoTools source topolammps.tcl — writelammpsdata requires a non-zero box; type IDs follow the alphabetical order of type names
- keinsci forum: a way to change the element order in a LAMMPS data file — Experience post: type order of data files exported from Materials Studio via OVITO differs from the force field
- LAMMPS manual: pair_style reaxff — Example of mapping elements to types in pair_coeff (C C N H)
- OVITO Python reference: DislocationAnalysisModifier — Global attributes total_line_length, cell_volume, length.1/n<ijk>
- CSDN: computing dislocation line length with Python ovito — Experience post: per-frame dislocation line length with OVITO Python
- CSDN: phase fractions and dislocation density with ovito+python — Experience post: workflow for per-frame dislocation density output in OVITO