Quick reference
The “Official explanation” column is the address printed after For more information see at the end of the error, and can be opened directly. Tested on this page: the first 18 rows were all triggered with small inputs on conda-forge LAMMPS 22Jul2025 update6 (macOS arm64, CPU build), and the error texts and err numbers match the table; the last 4 rows (MS-MPI, msmpi.dll, make and GPU) need environments not available on the test machine and were not reproduced.
| Error text | Most common actual cause | What to check first | Official explanation |
|---|---|---|---|
| ERROR: Lost atoms: original N current M | Atoms move too far in one step: initial overlaps, timestep too large, wrong units or parameters, boundary settings | Set thermo to 10 and find the step where PotEng, Press and Temp start to look abnormal | docs.lammps.org/err0008 |
| Out of range atoms - cannot compute PPPM | A charged atom moves out of its process's subdomain plus skin in one step | Whether Dangerous builds at the end of the log is non-zero; if it fails at step 0, check for overlaps | docs.lammps.org/err0004 |
| Bond atoms X Y missing on proc P at step S | Bonded atoms are farther apart than the communication cutoff; usually the system has already blown up | Temperature and pressure before the error; whether the pair cutoff is very short (e.g. repulsive-only LJ) | docs.lammps.org/err0005 |
| Non-numeric pressure - simulation unstable (and the atom coords / box dimensions variants) | Forces overflow to NaN or Inf | Whether NaN already appears at step 0; how many atoms delete_atoms overlap removes | docs.lammps.org/err0006 |
| Temp or Press shows nan or -nan in thermo output, but the run continues | Same as above; LAMMPS does not always stop immediately | Same as above; rerun with double-precision GPU settings | Section Pressure, forces, positions becoming NaN or Inf in Errors_details |
| Too many neighbor bins | The box has expanded too much, or the largest cutoff is too small | Add lx ly lz to thermo and see whether the box is growing quickly | docs.lammps.org/err0009 |
| Domain too large for neighbor bins | Some atoms are blown far away (s boundaries), or fix deform or a barostat has grown the box too much | Whether boundaries are s; the rate and units of fix deform | docs.lammps.org/err0017 |
| Cannot use neighbor bins - box size << cutoff | One box dimension is much smaller than the cutoff, common in thin-layer systems | Box length in the thinnest direction versus the pair cutoff | docs.lammps.org/err0015 |
| Neighbor list overflow, boost neigh_modify one | An atom has more neighbors than one (default 2000); a wrong box size can also inflate the density | Whether the box dimensions are correct; whether the data file mistakenly uses fractional coordinates | docs.lammps.org/err0036 |
| Did not assign all atoms correctly | With non-periodic boundaries, some atoms in the data file lie outside the box | Whether xlo xhi / ylo yhi / zlo zhi enclose all coordinates | docs.lammps.org/err0016 |
| Invalid atom ID in Bonds section of data file | A bond references an atom ID ≤ 0, larger than the maximum ID, or the same atom at both ends | The atoms count in the header versus lines in the Atoms section; whether IDs from the conversion tool start at 1 | Errors_messages |
| Incorrect format in Atoms section of data file | atom_style does not match the number of columns in the data file, or the header count exceeds the actual lines | Whether atom_style matches the comment on the Atoms section header (e.g. # full) | docs.lammps.org/err0002 |
| Unknown identifier in data file | A header keyword or section name is misspelled, or the first line contains a header keyword (the first line is always skipped as a comment) | The first line of the data file; capitalization of section names | docs.lammps.org/err0001 |
| Unrecognized pair style 'xxx' is part of the YYY package which is not enabled in this LAMMPS binary | The executable was built without package YYY | Installed packages in the lmp -h output | docs.lammps.org/err0010 |
| Incorrect args for pair coefficients | pair_coeff format does not match pair_style, e.g. mixing up eam and eam/alloy, or missing the sub-style name with hybrid | Compare with the pair_coeff examples on that pair_style's doc page | docs.lammps.org/err0021 |
| Numeric index X is out of bounds | A type number in pair_coeff or bond_coeff exceeds the number of types declared in the data file | atom types and bond types in the data file header | docs.lammps.org/err0019 |
| XXX command before / after simulation box is defined | Wrong command order: units and atom_style must come before read_data, pair_coeff after it | Position of the command relative to read_data / create_box | docs.lammps.org/err0033, err0034 |
| Substitution for illegal variable | A multi-letter variable without braces: $cutoff is parsed as $c | Change it to ${cutoff} | docs.lammps.org/err0013 |
| job aborted: [ranks] message ... abort code 1 (MS-MPI) | MPI-level summary; LAMMPS has already exited because of an input error | Remove -screen none and read the last ERROR in log.lammps | MatSci forum thread (see references) |
| Windows reports that msmpi.dll cannot be found | The MSMPI build of LAMMPS is being run, but the MS-MPI runtime is not installed or the machine was not rebooted after installing it | Whether msmpisetup.exe was installed and the machine rebooted | docs.lammps.org/Run_windows.html |
| mpicxx: command not found or mpi.h: No such file or directory during make mpi | The system has no MPI compiler wrapper or development headers | which mpicxx; without MPI, use make serial instead | docs.lammps.org/Build_basics.html |
| GPU library not compiled for this accelerator | The architecture or CUDA version the GPU library was built for does not match the current GPU and driver | The GPU_ARCH setting versus the GPU and driver version shown by nvidia-smi | docs.lammps.org/Build_extras.html |
Reading the error
Source location: the parentheses at the end of the ERROR line give the source file and line number, e.g. (src/thermo.cpp:526). Line numbers change between versions, so compare against the source of the same version.
ERROR versus ERROR on proc N: the former is an error detected consistently by all processes; the latter is detected and raised by a single process, and a serial run also prints it as ERROR on proc 0 (tested on this page: Substitution for illegal variable, Neighbor list overflow and Too many neighbor bins all use this format in serial). Lost atoms, Out of range and Bond atoms missing all depend on domain decomposition, so the error may appear or disappear with a different number of processes. That means the input itself is unstable; it does not mean “changing the core count fixed it”.
Last input line: most input syntax errors print the command being executed. Run-time errors such as Lost atoms are flagged NOLASTLINE and do not print this line, so read the thermo output before the error.
If a parallel run exits without any error message, the message may still be in the output buffer (4096 or 8192 bytes). Rerun with the command-line flag -nb to disable buffering; it slows the run noticeably, so use it only for debugging.
Lost atoms
By default LAMMPS checks the total atom count only on steps that write thermo output. With thermo 1000, the reported step can be up to 999 steps after the atoms were actually lost, so the first step is to shrink the thermo interval.
Tested on this page (LAMMPS 22Jul2025 update6, CPU, 1 core): in an 864-atom LJ liquid, adding one atom about 0.15σ from an existing atom gives a step-0 potential energy of 5.2×10⁷, and an NVE run reports Lost atoms: original 865 current 863 at the first thermo step. Applying the in.debug above to the same structure, minimization followed by 5000 steps of nve/limit and 2000 steps of NVE runs without errors, with Dangerous builds 0. With two atoms at exactly the same position, the step-0 potential energy is inf and the pressure is nan; fix nve keeps running and the temperature stays nan afterwards.
thermo_modify lost accepts error (default), warn and ignore. The official documentation says ignore is only for simulations where atoms are supposed to leave the system. With ignore the atom count changes, output formats that require a constant atom count such as dcd will fail, and statistics will be biased.
When the same input starts losing atoms after a LAMMPS version change, the developers' explanation is that the original input was only marginally stable and the old version happened not to trigger it; treat it as a new problem and troubleshoot it.
# Diagnostic settings: place after read_data / create_atoms and the force field, before run
thermo 10 # Lost atoms is only checked on thermo steps; shrink the interval first
thermo_style custom step temp pe ke etotal press vol lx ly lz
thermo_modify lost error norm no # keep the default error; do not switch to ignore
neigh_modify delay 0 every 1 check yes
# 1) Minimize first to remove close contacts
min_style cg
minimize 1.0e-4 1.0e-6 1000 10000
reset_timestep 0 # must come before the dump is defined, otherwise: Cannot reset timestep with active dump
# 2) Write coordinates, velocities and forces every 10 steps
dump d1 all custom 10 dump.debug id type x y z vx vy vz fx fy fz
dump_modify d1 delay 0 # if the error is known to occur near step N, change 0 to N-200
# 3) Limit per-step displacement during the initial stage (distance units; do not combine with fix shake)
fix f_lim all nve/limit 0.1
run 5000
unfix f_lim
# Check for overlapping atoms: run after pair_style / pair_coeff are defined
# Requirement: pair cutoff + neighbor skin >= 0.5 (distance units here); otherwise: Delete_atoms cutoff > max neighbor cutoff
# Bonded pairs with a special_bonds weight of 0 are not in the neighbor list and are not checked
# Use only on a diagnostic copy; it really deletes atoms
delete_atoms overlap 0.5 all all
# The log prints "Deleted N atoms, new total = ..."; N > 0 means the structure contains overlaps
| Cause | Typical signal | Diagnosis | Fix |
|---|---|---|---|
| Initial overlaps or close contacts | Potential energy and pressure are already huge at step 0; temperature jumps in the first steps | See how many atoms delete_atoms overlap 0.5 all all removes; search for close neighbors by distance in OVITO | Run minimize first, or run a few thousand steps with fix nve/limit 0.1, then switch back to fix nve/nvt |
| Timestep too large | Temperature keeps rising over the first few hundred steps; total energy drifts one way in NVE | Rerun with half the timestep; if the error is delayed or disappears, this is confirmed | Choose the timestep by system: 0.25 fs for flexible water, 1 fs with SHAKE on bonds to H, about 1 fs for organic systems with C/O/N |
| Wrong units or mismatched parameter units | Runs a few steps, but temperature and pressure are absurd; a Changing timestep warning appears in the log | Check units against the units of the potential file and literature parameters | See the section on mixed units below |
| Wrong element mapping in pair_coeff | Two atom types get abnormally close or pass through each other | Check that the element order after pair_coeff * * filename matches types 1, 2, ... in the data file | List element names in type order; use NULL for types handled by other styles (with hybrid) |
| Boundary conditions | Atoms fly off the surface with f boundaries; with a large box and s boundaries, the box collapses at the first step | Confirm whether the lost atoms are supposed to leave the system | For sputtering or evaporation, use thermo_modify lost warn; change s to m boundaries |
| Wall potential too close to the box edge | Atoms near fix wall move abnormally fast | Find where the lost atoms were in the dump | In a forum case, switching to fix wall/reflect solved it |
| fix deform rate too high or wrong unit conversion | lx changes a lot each step in thermo, with rising temperature | Convert erate to 1/s: in metal units erate 0.001 is 1e9 s⁻¹ | Lower the strain rate; confirm units box or lattice |
| Thermostat or barostat damping too small | Large oscillations in temperature or volume | Convert Tdamp and Pdamp into numbers of steps | Tdamp about 100 steps and Pdamp about 1000 steps; write them as $(100.0*dt) and $(1000.0*dt) |
PPPM and bonds
PPPM spreads each charge over several surrounding grid points. In parallel, each process holds only the grid for its subdomain plus the skin. If an atom moves out of that range before the next neighbor list rebuild, LAMMPS reports Out of range atoms - cannot compute PPPM. The “range” is the subdomain, which equals the whole box only in serial runs, so enlarging the box does not fix this error.
A widely shared practice in the Chinese community (a PPPM troubleshooting article by the same author on CSDN and Zhihu) splits the problem in two: if atoms are not moving very fast, rebuilding more often with neigh_modify delay 0 every 1 check yes is enough; if forces and velocities are very high, also reduce the timestep (the article goes down to 0.1 fs) and revisit the model and force field parameters; early in structure optimization, you can leave PPPM off. This agrees with the official explanation, provided you first confirm there are no overlapping atoms; otherwise a smaller timestep only delays the error.
Tested on this page (same version, 4 MPI processes): 800 atoms with ±1 charges placed randomly in a 30 Å cubic box (with severe overlaps), pair_style lj/cut/coul/long 10.0 plus pppm 1.0e-4. With neigh_modify delay 10 every 10 check no the run reports Out of range atoms - cannot compute PPPM; with the default rebuild settings the same structure reports Lost atoms instead; with the conservative settings below (skin 3.0 Å, 0.5 fs) it still reports Lost atoms; minimizing first makes it run normally. The conservative settings only fix late neighbor rebuilds; when the structure has overlaps you must minimize first.
When the same input runs on Windows but fails on Linux, the usual reason is a different FFT grid or domain decomposition; the cause is still a high-energy configuration in the system.
Bond atoms missing means the other atom of a bond is outside the communication cutoff. By default the communication cutoff is the pair cutoff plus the skin. Force fields with very short cutoffs, such as repulsive-only LJ, or pair_style none need a larger cutoff via comm_modify cutoff; if the system has already blown up, a larger cutoff does not help.
Tested on this page: with two bonded atoms 10 Å apart, pair cutoff 2.0 Å, skin 0.5 Å and 2 processes, LAMMPS reports Bond atom missing in image check (err0014) before the run starts; adding comm_modify cutoff 12.0 lets it run. When two bonded atoms are pulled apart during the run, the message is Bond atoms 1 2 missing on proc 0 at step 4 (err0005).
# Neighbor list statistics at the end of each run in log.lammps
Neighbor list builds = 1229
Dangerous builds = 0 # non-zero: atoms moved more than half the skin before a rebuild; rebuild more often or enlarge the skin
# Conservative settings for Out of range atoms / Bond atoms missing (real or metal units)
neighbor 3.0 bin # raise skin from the default 2.0 A to 3.0 A
neigh_modify delay 0 every 1 check yes
timestep 0.5 # 0.5 fs in real units; return to 1-2 fs after SHAKE-constraining bonds to H
# Early in equilibration, use cutoff Coulomb; switch back to long + pppm once the system is stable
# pair_style lj/cut/coul/cut 10.0
# (no kspace_style)
| Settings (real units, 1 fs baseline) | Neighbor list builds | Dangerous builds | Relative run time |
|---|---|---|---|
| delay 10, skin 2.0 Å, 1.0 fs | 1000 | 1000 | 1.00× |
| delay 0, skin 2.0 Å, 1.0 fs | 1229 | 0 | 1.04× |
| delay 10, skin 3.0 Å, 1.0 fs | 617 | 0 | 1.00× |
| delay 10, skin 2.0 Å, 0.5 fs | 1252 | 0 | 1.84× |
Neighbor lists
Too many neighbor bins
By default a bin is half the largest pair cutoff. The error occurs when the box grows so large that the number of bins exceeds the integer limit. Most often the system has blown up and the box grew with it; output lx ly lz in thermo to confirm. If the cutoff really is small, enlarge the bins with neigh_modify binsize. The neigh_modify bin/hash option in the current development docs does not exist in the 22Jul2025 stable release; tested on this page, it fails with Unknown neigh_modify keyword: bin/hash.
Domain too large for neighbor bins
Some atoms are blown far away and s boundaries stretch the box with them; fix deform or a barostat can also grow the box too much. Check atom velocities and box size before the error, then the fix deform rate.
Cannot use neighbor bins - box size << cutoff
One box dimension is much smaller than the cutoff, common in monolayer or thin-film systems. Enlarge the box in that direction (a vacuum layer), or use neighbor 2.0 nsq, which is slower.
Neighbor list overflow, boost neigh_modify one
Defaults are one 2000 and page 100000; page must be at least 10 times one, and 50 to 100 times is recommended. Before raising them, confirm the density is reasonable: a wrong box size or fractional coordinates in the data file both inflate the density.
NaN
- 01
Find the first step with NaN
Rerun with thermo 1. If NaN appears at step 0, the most common cause is overlapping atoms, which look like a single atom in a visualization. With periodic boundaries, if the box is set to the coordinate extremes without padding, atoms on opposite sides overlap through the boundary. Tested on this page: with two atoms at identical coordinates, fix nve alone raises no error and thermo keeps printing nan; with fix npt the run stops at step 0 with Non-numeric pressure - simulation unstable; a 4-process fix nve run stops with Non-numeric atom coords - simulation unstable.
- 02
Check for overlaps and minimize first
Use delete_atoms overlap on a diagnostic copy. If it deletes any atoms, go back to model building and leave spacing, or run minimize first.
- 03
Thermostat first, then barostat
The official docs note that NaN is more likely when a Nose-Hoover barostat is used directly. Run with only a thermostat until the potential energy stabilizes, then add fix npt; you can also minimize again after some equilibration MD.
- 04
Use double precision first on GPU
The GPU package is built in mixed precision by default (GPU_PREC=mixed), and single or mixed precision overflows more easily. Do the initial relaxation in double precision or on CPU, then switch back to mixed precision for production.
- 05
Start with a soft potential
pair_style soft or soft-core potentials are insensitive to close contacts and can push apart a badly overlapped initial structure before switching to the production force field.
Build and packages
Unknown pair style is the error text from versions before 2019. Since 2020 it reads Unrecognized pair style 'xxx' is part of the YYY package which is not enabled in this LAMMPS binary, naming the missing package directly. If the package is reported as installed but the style is still missing, a package it depends on is missing; check the Restrictions section of that style's documentation.
You must rebuild after installing a package. With traditional make, you cannot install a package and compile in the same command: make yes-colloid mpi does not take effect and must be split into two commands. Changing packages without rebuilding, or without replacing the old executable, is a common oversight named in the official documentation.
Building 22Jul2025 with CMake requires CMake 3.16 or later and a compiler with at least C++11; C++17 is used automatically when supported, and KOKKOS requires C++17. make mpi calls mpicxx by default; without MPI, use make serial, or install the OpenMPI or MPICH development packages first.
# Show which packages and styles the executable contains, and whether it has MPI
lmp -h | less # search for "Installed packages" and "pair styles"
# CMake build: add -D PKG_xxx=on for each missing package, then you must rebuild
cmake -S cmake -B build -D PKG_MANYBODY=on -D PKG_KSPACE=on -D PKG_MOLECULE=on -D BUILD_MPI=on
cmake --build build -j 8
# Traditional make: installing packages and compiling must be two separate commands
cd src
make yes-manybody yes-kspace
make mpi -j 8 # produces lmp_mpi; "make yes-manybody mpi" does not take effect
Parallel and Windows
REM Running the MSMPI build of LAMMPS in parallel on Windows
REM First use dir to confirm the full input file name (File Explorer hides extensions by default)
dir
mpiexec -localonly 4 lmp -in in.CHO.lmp
REM Do not add -screen none, or the actual ERROR line will not appear in the terminal
job aborted: [ranks] message ... abort code 1
This is MS-MPI's summary, meaning LAMMPS exited with its own error. The real cause is the ERROR line in the terminal or log.lammps. With -screen none, the terminal shows no ERROR. In one forum thread a user spent four days debugging MS-MPI; the actual cause was that the input file was named in.CHO.lmp while the command used -in in.CHO.
msmpi.dll not found
Among the official Windows installers, only the LAMMPS-MSMPI builds support MPI parallel runs, and they require Microsoft's msmpisetup.exe plus a reboot. msmpisdk.msi is only needed to compile from source with Visual Studio. For multi-threaded speed-up only, use the non-MSMPI build with -pk omp 4 -sf omp.
Output from the MSMPI build appears late
In MPI parallel mode, screen output is block-buffered and only shows after enough bytes accumulate. This is documented known behavior and does not mean the run is stuck.
Fine on one core on Linux, fails on several
A serial run has no domain decomposition, so every atom is accessible and the problem is hidden. Lost atoms or Out of range on multiple cores means the dynamics are wrong. If two decompositions give very different energies, neither result can be trusted.
MPI_ABORT raised in the OpenMPI layer
lmp -h shows which MPI library the executable is linked against. If the problem is in the MPI library itself and you only need serial runs on that machine, reconfigure with -D BUILD_MPI=off.
No error
Unit defaults: with units real, the default timestep is 1.0 fs and skin 2.0 Å; with units metal, the default timestep is 0.001 ps and skin 2.0 Å. Converting a real input to metal while keeping timestep 1.0 makes the timestep 1 ps, 1000 times larger. ReaxFF potentials use real; most EAM and Tersoff potential files use metal.
| Symptom | Common cause | Check and fix |
|---|---|---|
| Process uses full CPU, but no new output for a long time | thermo not set (output only at first and last step); output buffered under MPI; a large system with pppm or ewald, which do not scale linearly with atom count | Set thermo 100; rerun with -nb; measure time per step on a small system before extrapolating |
| Genuinely stuck at one step | A value became NaN and a loop cannot exit, or a code defect | Attach with gdb -p PID and type where to get a stack trace; include it in your forum post |
| Total energy drifts one way in NVE | Timestep too large, or neighbor lists rebuilt too late | Compare the drift at half the timestep; check Dangerous builds |
| Temperature keeps rising | Tdamp too small or wrong unit conversion; thermostat group differs from integration group; the same atoms integrated by two fixes | Look for the One or more atoms are time integrated more than once warning; use Tdamp of about 100 steps |
| Energy or pressure off by orders of magnitude | Mixing units metal and real: eV and kcal/mol differ by about 23 times, and bar differs from atm | Check the units of the source of your pair_coeff parameters; a potential file with a UNITS: tag errors out on a mismatch, one without the tag does not |
| Timestep silently changed | timestep placed before units; units resets it to the default | Search the log for a warning such as WARNING: Changing timestep from 2 to 1 due to changing units to real (exact text from a test on this page); put timestep after units |
| Same input, different results on different core counts | Round-off makes trajectories diverge after a few hundred to a few thousand steps; with create_atoms, atom IDs depend on the core count, so velocity assigns different velocities to the same atom | Compare statistical averages, not step-by-step values; add loop geom to the velocity command if initial velocities must be independent of the core count |
Workflow
# Minimal reproducer: start from the failing restart or data file and keep only commands needed to trigger the error
units real
atom_style full
read_data small.data # use the smallest cell before replicate, or remove most of the solvent
pair_style lj/cut/coul/long 10.0
kspace_style pppm 1.0e-4
include ff.params # keep force field parameters in a separate file for easy comparison
thermo 1
thermo_style custom step temp pe press vol
dump d1 all custom 1 dump.min id type x y z fx fy fz
fix f1 all nve
run 200 # reproducing within a few hundred steps is enough
- 01
Read the whole error
Note the error text, whether it is ERROR or ERROR on proc N, the source location and the err number, and open the official explanation. Read all earlier WARNINGs; warnings issued during a run are printed only to the screen, not to the log file.
- 02
Decide whether it is the input stage or the run stage
Errors before or after box definition and before run are usually syntax, command order or data file problems. Errors in the middle of a run are usually dynamics problems.
- 03
Shrink the interval and find where it starts
Set thermo to 1 to 10 and output pe, press, temp and lx. Write a dump every 1 to 10 steps, and use dump_modify delay to record only the stretch before the error.
- 04
Build a minimal reproducer
Reproduce the same error with a small system, one core and the fewest commands. If the small system fails within a few hundred steps, each iteration takes only minutes.
- 05
Change one variable at a time
Change one thing per run: minimize first, then check units and parameters, then the timestep, and only then the neighbor list. Changing several things at once hides the real cause.
- 06
Confirm the fix
A vanished error does not make the result trustworthy. Run at least a stretch of NVE to check energy conservation, and check that Dangerous builds is 0.
Frequent in Chinese forums
The entries in this section come from widely shared LAMMPS error round-ups and help threads on Zhihu and the Computational Chemistry Commune forum (keinsci). They were tested with LAMMPS 22Jul2025 update6 (conda-forge CPU build) or checked against the current source and documentation.
# Three ways to handle hbondchk failed / bondchk failed in sparse or many-rank ReaxFF runs
# 1) Enlarge the preallocation (defaults: safezone 1.2, mincap 50, minhbonds 25); memory use grows accordingly
pair_style reaxff NULL safezone 1.6 mincap 100 minhbonds 50
pair_coeff * * ffield.reax C H O N
# 2) Use the KOKKOS version of ReaxFF (a CPU-only KOKKOS build also works); it ignores safezone and related keywords
# mpirun -np 4 lmp -k on -sf kk -in in.reaxff
# 3) When mixing with non-ReaxFF potentials, apply charge equilibration to ReaxFF atoms only
# pair_style hybrid reaxff NULL lj/cut 10.0
# pair_coeff * * reaxff ffield.reax C H O NULL
# group reax type 1 2 3
# fix qeq reax qeq/reaxff 1 0.0 10.0 1.0e-6 reaxff maxiter 500
step N: hbondchk failed or bondchk failed (err0018)
ReaxFF preallocates its bond and hydrogen-bond lists from the current system plus a safety factor. With vigorous reactions, a sparse system or very few atoms per process, the lists overflow. Tested on this page: 8 RDX molecules (168 atoms) in a periodic box with 120 Å edges, NVT at 3000 K with a 0.25 fs timestep. With default settings, 1 process reported hbondchk failed at step 13617, 2 processes at step 13014 and 8 processes at step 5219. With safezone 1.6 mincap 100 minhbonds 50, both 1 and 8 processes completed 20,000 steps. The two other remedies in the official documentation are splitting a long run into segments and switching to the KOKKOS version of ReaxFF. Splitting the same system into 20 segments of 1000 steps on 8 processes still ended in a segmentation fault between steps 3000 and 4000, so splitting is not always enough.
Too few atoms per process with ReaxFF
The LAMMPS error round-up on Zhihu gives a rule of thumb of at least 100 atoms per core for ReaxFF, with 300 to 500 being better; in a pyrolysis help thread on keinsci, recurring lost atoms disappeared after the author changed the core count. More cores is not always faster: for the sparse 168-atom system above with the larger safezone, 20,000 steps took 8.0 s on 1 process and 111.5 s on 8 processes on the same machine. For small systems, measure the time per step on 1 to 2 processes before adding cores.
Fix qeq/reaxff CG convergence failed after N iterations
This is a warning: charge equilibration did not converge within maxiter iterations. The default maxiter is 200; append maxiter 500 to the fix to raise it. Chinese round-ups suggest silencing the warning with warn no, which in the current version stops with ERROR: Illegal fix qeq/reaxff command (tested on this page); the keyword that silences the warning is nowarn. The official documentation describes nowarn for serial-versus-parallel comparisons that need a fixed number of iterations; when the warning keeps appearing during a run, first check whether the system has already become unstable, as described earlier on this page.
Mixing ReaxFF with other potentials (pair_style hybrid)
In pair_coeff, write NULL as a placeholder for types that do not use ReaxFF. Apply fix qeq/reaxff only to a group made of ReaxFF atoms: with charge equilibration on all, types without QEq parameters trigger No QEq parameters for atom type N provided by pair reaxff, and a group with non-zero total charge gives the warning Fix qeq/reaxff group is not charge neutral. When charge equilibration is not needed, turn the check off with pair_style reaxff NULL checkqeq no; the static charges from the data file are then used.
Triclinic box skew is too large and box tilt large
Chinese round-ups fix this error by adding box tilt large to the script. Since version 22Dec2022 the box command has been removed, and a tilt larger than half the box length only gives WARNING: Triclinic box skew is large. LAMMPS will run inefficiently. The calculation continues. Tested on this page with 22Jul2025: an xy tilt of 2.5 times the box length also gives only this warning, and keeping box tilt large in the script just prints The 'box' command has been removed and will be ignored. If you still see skew too large, you are running a version older than 2022.
Hand off to an agent
Example instruction: “Here are my in.npt, system.data and error log. At step 3200 it reports Out of range atoms - cannot compute PPPM. Reproduce the error, check minimization, units and parameters, timestep and neighbor list in that order, changing one thing at a time. Once you find the cause, give me the corrected input file and an NVE energy conservation check.”
Steps the agent performs: reproduce the error in an environment with LAMMPS 2025.07.22 GPU build preinstalled; shrink the thermo and dump intervals to find where things go wrong; build a minimal reproducer; change and rerun one item at a time; run an energy conservation check on the corrected input and adversarially review the result.
Output files: the corrected in file, the log and dump from each troubleshooting round, Dangerous builds statistics and energy drift data, and a record explaining the cause and the basis for each change.
You still need to check: whether the force field and potential suit your system, whether units and parameter sources are consistent, and whether the corrected timestep and thermostat/barostat settings are adequate for the physical quantities you will analyze in your paper.
Asking for help
The official place to ask for help is the LAMMPS category of the MatSci community forum. For posts missing information, developers usually only reply “please attach a complete input”.
- LAMMPS version (first line of the log, e.g. LAMMPS (22 Jul 2025)) and platform, plus the installed packages and MPI information from lmp -h
- The complete error text, including source location and err number, as text rather than a screenshot
- A complete input that runs as is: in file, data file and potential files in a zip; avoid binary restart files where possible
- The smallest reproducing system: fewest atoms, fewest processes, shortest run
- The run command (mpirun -np 4 lmp -in ..., whether -sf gpu or -k on is used)
- Thermo output from the steps before the error
- Changes already tried and their results, e.g. “lowering the timestep from 1 fs to 0.5 fs delayed the error from step 3200 to step 9000”
- Whether serial and parallel runs agree
- Search the forum archive first: Out of range atoms is one of the most frequently asked questions on the forum
References
- LAMMPS docs: Errors and warnings details — Official explanations for err0001-err0039 and general troubleshooting advice
- LAMMPS docs: Common issues that are often regarded as bugs — Invalid style, -nb to disable buffering, cross-platform differences
- LAMMPS docs: Debugging crashes / appears to be stuck — Getting a stack trace with gdb; reasons a run appears stuck
- LAMMPS docs: Error messages — Short errors such as Invalid atom ID
- LAMMPS docs: thermo_modify — Values and defaults of the lost and lost/bond keywords
- LAMMPS docs: neigh_modify — Defaults for one, page and binsize
- LAMMPS docs: units — Default timestep and skin for each unit style
- LAMMPS docs: fix nvt/npt — Recommendation of about 100 steps for Tdamp and 1000 steps for Pdamp
- LAMMPS docs: Running LAMMPS on Windows — Installing MS-MPI and using mpiexec
- LAMMPS docs: Include packages in build — Rebuild after installing packages; make cannot install and build in one command
- LAMMPS source stable_22Jul2025 — Error texts, errorurl numbers, units resetting the timestep
- MatSci: Tracing atom that causes Out of range atoms — Experience post: Dangerous builds comparison and timestep guidance (developer replies)
- MatSci: Out of range atoms - cannot compute PPPM (pppm.cpp:1887) — Experience post: Out of range refers to parallel subdomains (developer reply)
- MatSci: ERROR: Lost atoms — Experience post: wall too close causing lost atoms; results differing between decompositions cannot be trusted
- MatSci: job aborted with the MSMPI build of LAMMPS — Experience post: job aborted caused by a hidden input file extension
- CSDN: Out of range atoms - cannot compute PPPM, causes and fixes (Chinese) — Experience post: handling the PPPM error in two cases by atom speed
- MatSci: Lost Atoms When I Changed the LAMMPS Version — Experience post: why lost atoms appear after a version change
- Zhihu: LAMMPS error round-up — Experience post: box tilt large, hbonds and atoms per core, qeq warning, ReaxFF hybrid entries
- keinsci forum: lost atoms in a ReaxFF simulation — Experience post: lost atoms in ReaxFF pyrolysis disappeared after changing the core count
- LAMMPS documentation: pair_style reaxff — Defaults of safezone, mincap, minhbonds; KOKKOS and run-splitting advice; checkqeq
- LAMMPS documentation: fix qeq/reaxff — maxiter default 200, nowarn keyword, zero total charge in the group
- LAMMPS documentation: Removed commands and packages — box command removed in 22Dec2022