Check the version first
The CellChat repository moved from sqjin/CellChat to jinworks/CellChat. The current GitHub main is 2.2.0.9001 (last commit 2026-03-04), and the latest release tag is v2.1.2. What is labeled "v3" online is the separate SpatialCellChat package for spatial transcriptomics; single-cell data still use CellChat v2.
Record packageVersion("CellChat") and the installation date in the methods section. GitHub main also changes between releases: PR #438, merged on 2026-02-19, changed the parallel behavior, and rankNet was modified on 2026-03-04.
| Item | v1-era tutorials (2021–2023) | Current v2 | Result of copying the old code |
|---|---|---|---|
| Database size | CellChatDB.human about 1,939 pairs | 3,233 pairs: 1,920 from v1, 1,313 added in v2 (994 Non-protein Signaling, 233 Cell-Cell Contact, 83 Secreted); mouse 3,379 pairs | Using the full database mixes metabolite and neurotransmitter signals into the results |
| Over-expression test | Seurat-style Wilcoxon | do.fast = TRUE by default calls presto; without presto the function stops | Same PBMC 3k data: presto gives 400 over-expressed signaling genes, do.fast = FALSE gives 744; v1 and v2 results are not directly comparable |
| PPI smoothing | projectData(cellchat, PPI.human) | Renamed smoothData; adj = PPI.human must be named | Error: is(adj, "matrix") || is(adj, "sparseMatrix") is not TRUE (GitHub #185) |
| Object slot | data.project | data.smooth | Old objects fail with no slot of name "data.smooth"; not fixed by the maintainer, users set obj@data.smooth <- matrix(nrow = 0, ncol = 0) (GitHub #202) |
| meta | Label column only | Also a samples column; if missing, it is set to sample1 with a warning | After merging samples, min.samples filtering is unavailable |
| Parallel | plan("multiprocess") | future has removed the multiprocess backend | Error: No such backend for futures: 'multiprocess' |
| Spatial transcriptomics | Visium support since v1.6 | v2 supports several spatial technologies; single-cell-resolution data use SpatialCellChat (v3) | — |
Step 1
CellChat is installed only from GitHub and contains C++ code, so it needs a compiler. The dependencies that fail most often are NMF (CRAN, requires Biobase), ComplexHeatmap and BiocNeighbors (Bioconductor), and presto (GitHub). With the other dependencies installed from conda, compiling NMF, presto and CellChat took 2.5 minutes in total on this page, without errors. Nature Protocols gives a dual-core CPU and 16 GB of RAM as the minimum.
# Option A: dependencies from conda (tested on this page, macOS arm64, about 4 min + 2.5 min compiling)
micromamba create -y -n cellchat -c conda-forge -c bioconda \
r-base=4.5 r-seurat r-remotes r-rcppeigen c-compiler cxx-compiler fortran-compiler make \
r-future r-future.apply r-pbapply r-irlba r-ggalluvial r-svglite r-ggrepel r-circlize \
bioconductor-complexheatmap bioconductor-biocneighbors bioconductor-biobase \
r-rspectra r-reticulate r-sna r-fnn r-shape r-ggpubr r-ggnetwork r-plotly r-shiny r-bslib \
r-collapse r-pkgmaker r-registry r-rngtools r-gridbase r-doparallel r-cowplot r-uwot
micromamba run -n cellchat Rscript -e '
options(repos = c(CRAN = "https://cloud.r-project.org"))
install.packages("NMF") # conda-forge has no r-nmf
remotes::install_github("immunogenomics/presto", upgrade = "never")
remotes::install_github("jinworks/CellChat", upgrade = "never")
packageVersion("CellChat")'
# Option B: an existing R 4.4 or later
# install.packages(c("BiocManager", "remotes", "NMF"))
# BiocManager::install(c("ComplexHeatmap", "BiocNeighbors", "Biobase"))
# remotes::install_github(c("immunogenomics/presto", "jinworks/CellChat"), upgrade = "never")| Error or symptom | Cause | Fix |
|---|---|---|
| Skipping 3 packages not available: Biobase, BiocNeighbors, BiocGenerics | remotes does not install Bioconductor packages | First run BiocManager::install(c("Biobase", "BiocNeighbors", "BiocGenerics", "ComplexHeatmap")) |
| Failed to connect to github.com / download timeout | Unstable GitHub access from China | Install dependencies from conda or a Bioconductor mirror so that only presto and CellChat come from GitHub; if that still fails, download the source zip and run R CMD INSTALL |
| installation of package ... had non-zero exit status (macOS) | Missing or broken Xcode command line tools | Reinstall the Xcode command line tools (confirmed by two users in GitHub #226), or use conda's c-compiler and cxx-compiler |
| no member named 'Rlog1p' in namespace 'std' | Old R with the compiler on Apple silicon | export PKG_CPPFLAGS="-DHAVE_WORKING_LOG1P" and reinstall (Chinese community post, 2023) |
| CMake was not found on the PATH | A dependency needs CMake | conda install cmake |
| dependencies 'RSpectra', 'RcppEigen', 'ggpubr', 'BiocNeighbors' are not available | CRAN and Bioconductor dependencies incomplete | Install them with BiocManager::install(), or use the r-* packages from conda-forge |
| For a faster implementation of the Wilcoxon Test, please install the presto package | v2 uses presto by default | remotes::install_github("immunogenomics/presto"); if it cannot be installed, identifyOverExpressedGenes(do.fast = FALSE), which gives different results from presto |
Step 2
CellChat needs two inputs: a log-normalized expression matrix (genes in rows, cells in columns) and a group label for each cell. The code below was checked on Seurat 5.5.1; the comments name the error each line prevents.
library(Seurat)
library(CellChat)
seu <- readRDS("pbmc3k_annot.rds") # annotated Seurat v5 object
seu <- JoinLayers(seu) # join layers first, otherwise only the first layer is read
seu$celltype <- droplevels(factor(seu$celltype))
stopifnot(!anyNA(seu$celltype)) # labels must not contain NA
data.input <- LayerData(seu, assay = "RNA", layer = "data") # log-normalized values, not counts
stopifnot(min(data.input) >= 0) # scale.data or integrated cannot be used
meta <- data.frame(
labels = seu$celltype,
samples = factor(seu$orig.ident), # v2 expects a samples column
row.names = colnames(seu)
)
# prefix numeric cluster IDs; CellChat rejects the label "0"
# meta$labels <- factor(paste0("C", seu$seurat_clusters))
cellchat <- createCellChat(object = data.input, meta = meta, group.by = "labels")
table(cellchat@idents)Why normalized data
CellChat divides the matrix by its global maximum and then applies a Hill function (Kh = 0.5) to compute probabilities. With counts the maximum is in the hundreds, so almost all scaled values are close to 0. Tested on this page: PBMC 3k with counts runs without error and gives 170 interactions (152 with correct input) and the same pathways, but the total strength is about 1/800 of the correct value. Strength comparisons between groups and datasets become meaningless.
Neither scale.data nor integrated
When a Seurat object is passed directly, createCellChat checks for negative values and stops; when a matrix is passed, it does not check. On this page, passing scale.data (2,000 variable genes, negative values) as a matrix ran to the end and gave 201 interactions without any message. The integrated assay holds only a subset of genes after correction, and CellChat warns about it. The data layer of the SCT assay contains log1p-transformed corrected counts and can be used.
Objects with split layers
A multi-sample Seurat v5 object can be split (data.A, data.B) before or after IntegrateLayers. Tested on this page: passing the object directly gives invalid 'row.names' length; LayerData(layer = "data") only warns and returns the first layer with 1,319 of 2,638 cells. Run JoinLayers first.
Four requirements for labels
No NA; no unused factor levels; no label 0 (Seurat cluster IDs used directly give Cell labels cannot contain `0`!); preferably more than 10 cells per group. The first two give the same error: Please check `unique(object@idents)` and ensure that the factor levels are correct!
Label granularity: major types or subclusters
CellChat computes between groups and averages away heterogeneity within a group. Nature Protocols lists this as a limitation and suggests subclustering before running CellChat when needed. More groups increase run time and plot complexity: in GitHub #167 a user with 133 groups and 11,000 cells reported more than 11 hours. A common approach is to obtain the global network with major types first, then split the types of interest into subclusters and run them separately. No specific group-number recommendation from the maintainers was found.
Organizing multiple samples
In GitHub #62 the maintainer sqjin recommends merging all samples of one condition into one object, as in integrated clustering. To reduce interactions found in only one sample, add a samples column to meta and use filterCommunication(min.samples = 2) to require an interaction in at least two samples (since v2.1.2). A user in the same issue reported that the merged run had 4–5% unique interactions not found in any per-sample run.
Extra input for spatial transcriptomics
createCellChat(datatype = "spatial", coordinates = ..., spatial.factors = data.frame(ratio = ..., tol = ...)): coordinates are pixel positions in the full-resolution image, ratio converts pixels to microns (65 / spot_diameter_fullres for Visium), and tol is half the spot diameter (32.5). In computeCommunProb, interaction.range = 250 and contact.range = 100 for Visium; about 10 for single-cell-resolution technologies. Single-cell-resolution spatial data use SpatialCellChat.
Step 3
CellChatDB sorts ligand-receptor pairs into four categories. The official tutorial uses only Secreted Signaling as an example. The choice of subset changes the number of results more than any algorithm parameter.
Pros and cons of Secreted Signaling only
It gives fewer results, cleaner plots and comparability with earlier papers. The cost is that the main contact signals between immune cells (MHC-I/II, CD45, ICAM, SELPLG and others) are missing: on PBMC 3k, adding Cell-Cell Contact raised the pathways from 13 to 34. For immune synapses, T cell–APC interactions or tumor immune checkpoints, use subsetDB(CellChatDB).
Non-protein Signaling is excluded by default
The official tutorial advises against using the whole CellChatDB. Ligand levels in this category are inferred from synthesizing enzymes and transporters because metabolites themselves are not measured. When neurotransmitter signals are needed in nervous system studies, add them separately and interpret them apart from protein signals.
Species and gene names
Use CellChatDB.human for human, CellChatDB.mouse for mouse; CellChatDB.zebrafish also exists. Gene names must be official symbols: with Ensembl IDs or mixed human and mouse names, almost no signaling genes remain after subsetData, and downstream steps fail with errors such as data.use[RsubunitsV, ]. Convert orthologs first for other species.
Adding custom ligand-receptor pairs
Use updateCellChatDB() to merge your own interaction, complex, cofactor and geneInfo tables (official Update-CellChatDB tutorial). Pairs without a pathway assignment do not appear in pathway-level plots.
| Category | Human pairs | Description | PBMC 3k on this page (triMean / truncatedMean 0.1) |
|---|---|---|---|
| Secreted Signaling | 1,280 | Secreted paracrine and autocrine signals | 152 / 241 interactions, 6 / 13 pathways |
| ECM-Receptor | 424 | Extracellular matrix-receptor; important in tissues, rare in blood | — |
| Cell-Cell Contact | 535 | Requires direct contact, such as MHC, ICAM, CD86 | First three categories together (subsetDB(CellChatDB)): 437 / 967 interactions, 14 / 34 pathways |
| Non-protein Signaling | 994 | Metabolites and neurotransmitters; ligand level inferred from synthesizing enzymes and transporters | Full database: 445 / 1,055 interactions, of which 8 / 88 Non-protein |
Step 4
The core workflow below runs on PBMC 3k. The table after the code explains each parameter and how to choose its value; the next section shows the measured effects.
CellChatDB <- CellChatDB.human # use CellChatDB.mouse for mouse
cellchat@DB <- subsetDB(CellChatDB) # all protein ligand-receptor pairs, without Non-protein Signaling
# cellchat@DB <- subsetDB(CellChatDB, search = "Secreted Signaling", key = "annotation")
future::plan("sequential") # sequential is faster on small datasets
cellchat <- subsetData(cellchat)
cellchat <- identifyOverExpressedGenes(cellchat) # uses presto for the Wilcoxon test by default
cellchat <- identifyOverExpressedInteractions(cellchat)
cellchat <- computeCommunProb(cellchat,
type = "truncatedMean", trim = 0.1, # trim only takes effect with truncatedMean
population.size = FALSE) # TRUE is an option for unsorted samples
cellchat <- filterCommunication(cellchat, min.cells = 10)
cellchat <- computeCommunProbPathway(cellchat)
cellchat <- aggregateNet(cellchat)
df.net <- subsetCommunication(cellchat) # ligand-receptor level
df.netP <- subsetCommunication(cellchat, slot.name = "netP") # pathway level
nrow(df.net); cellchat@netP$pathways
saveRDS(cellchat, "cellchat_pbmc3k.rds")| Parameter | Default | Role and how to choose |
|---|---|---|
| thresh.p in identifyOverExpressedGenes | 0.05 | Selects over-expressed signaling genes per group by unadjusted p value; thresh.fc = 0 and thresh.pc = 0. A pair enters the calculation if either the ligand or the receptor is over-expressed |
| type in computeCommunProb | "triMean" | triMean is the mean of the quantiles Q1, Q2, Q2 and Q3: if fewer than 25% of cells express a gene, the group average is 0; at 25%–50% only Q3/4 remains. Fewer, stronger interactions |
| trim with type = "truncatedMean" | 0.1 | The mean after removing the fraction trim from each end: the group average is 0 if the expressing fraction is below trim. 0.1 means a 10% threshold, 0.05 means 5%. trim only takes effect with truncatedMean (and the undocumented thresholdedMean). Nature Protocols recommends trim = 0.1 in general and lowering trim when well-known signals of the system are missing |
| computeAveExpr(features = , type = , trim = ) | — | Shows group averages of ligands and receptors of interest before fixing trim; a gene that is 0 under triMean but non-zero at trim = 0.1 is expressed in 10%–25% of the cells (the approach recommended in Nature Protocols) |
| population.size | FALSE | TRUE multiplies probabilities by the proportions of the two groups, intended for unsorted samples; use FALSE for sorted or enriched samples |
| nboot | 100 | Number of permutations; p value resolution is 1/nboot. Measured: 20 runs 4.3 s and 938 interactions, 100 runs 22.6 s and 967, 1,000 runs 493 s and 972 |
| min.cells in filterCommunication | 10 | Removes groups with at most min.cells cells (the source uses <=), so a group of exactly 10 cells is removed |
| raw.use | TRUE | FALSE uses expression smoothed by smoothData, meant to compensate for missing subunits in shallow data, at the risk of false positives |
Parameter tests
Tested on this page: CellChat 2.2.0.9001, 2,638 cells in 9 groups from the official Seurat PBMC 3k workflow (13 platelets, 32 DCs), nboot = 100. "Interactions" are the rows returned by subsetCommunication (source group–target group–ligand-receptor pair). Nature Protocols compared triMean, trim 0.1 and trim 0.05 on atopic dermatitis skin data (5,011 cells, 12 groups) with the same direction: smaller trim gives more pairs and pathways, mostly weak signals.
triMean misses known signals in shallow data
On PBMC 3k (Cell Ranger 1.1, median 817 genes per cell) triMean finds no CCL, CXCL or IFN-II pathway. With trim = 0.1, IFNG–IFNGR1/IFNGR2 appears from NK cells to CD14 monocytes, FCGR3A monocytes and DCs, consistent with known IFN-γ secretion by NK cells. The gap between the two methods is smaller in deeply sequenced data.
Smaller trim, more false positives
At trim = 0.1 the sources of CCL5–CCR1 include B cells and naive CD4 T cells, in which only a few cells express CCL5. Significance is relative to random label permutation, so a ligand just above the threshold can be significant. Results at trim = 0.05 suit exploration; conclusions for a paper should hold at both trim = 0.1 and triMean.
population.size changes rankings, not the list
With truncatedMean 0.1, the significant interactions for TRUE and FALSE match row by row (241), while the total strength drops from 5.71 to 0.058. Sender ranking: DC first with FALSE; Memory CD4 T first and DC seventh with TRUE. DCs have only 32 cells, so this parameter decides whether a small group with strong signals ranks high.
min.cells mainly affects small groups
Platelets have 13 cells: min.cells = 10 keeps them, and at trim = 0.05 there are 45 interactions involving platelets; min.cells = 20 removes them all. identifyOverExpressedGenes also has min.cells = 10, which affects only the differential test.
| Database | type / trim | Interactions | L-R pairs | Pathways | Involving platelets (min.cells = 10) | Interactions at min.cells = 20 |
|---|---|---|---|---|---|---|
| Secreted | triMean | 152 | 10 | 6 | 1 | 151 |
| Secreted | truncatedMean 0.25 | 151 | 10 | 6 | 3 | 148 |
| Secreted | truncatedMean 0.1 | 241 | 19 | 13 | 13 | 228 |
| Secreted | truncatedMean 0.05 | 402 | 37 | 23 | 45 | 357 |
| Without Non-protein | triMean | 437 | 41 | 14 | 36 | 401 |
| Without Non-protein | truncatedMean 0.1 | 967 | 79 | 34 | 98 | 869 |
| Without Non-protein | truncatedMean 0.05 | 1,375 | 116 | 51 | 186 | 1,189 |
| Full database | truncatedMean 0.05 | 1,527 | 130 | 57 | 203 | 1,324 |
Step 5
aggregateNet produces two matrices: net$count is the number of significant ligand-receptor pairs between two groups, and net$weight is the sum of their communication probabilities. Count reflects how many kinds of signals there are, weight reflects the total signal, and their rankings often differ.
groupSize <- as.numeric(table(cellchat@idents))
par(mfrow = c(1, 2), xpd = TRUE)
netVisual_circle(cellchat@net$count, vertex.weight = groupSize, weight.scale = TRUE,
label.edge = FALSE, title.name = "Number of interactions")
netVisual_circle(cellchat@net$weight, vertex.weight = groupSize, weight.scale = TRUE,
label.edge = FALSE, title.name = "Interaction strength")
pathways.show <- "MIF" # check that it is in cellchat@netP$pathways first
netVisual_aggregate(cellchat, signaling = pathways.show, layout = "chord")
netVisual_heatmap(cellchat, signaling = pathways.show, color.heatmap = "Reds")
netVisual_bubble(cellchat, sources.use = c("CD14 Mono", "FCGR3A Mono"),
targets.use = c("CD8 T", "NK", "B"), remove.isolate = FALSE)
cellchat <- netAnalysis_computeCentrality(cellchat, slot.name = "netP")
netAnalysis_signalingRole_scatter(cellchat)
netAnalysis_signalingRole_heatmap(cellchat, pattern = "outgoing") +
netAnalysis_signalingRole_heatmap(cellchat, pattern = "incoming")
library(NMF); library(ggalluvial)
selectK(cellchat, pattern = "outgoing", k.range = 2:6, nrun = 10) # the default 2:10 with nrun = 30 is slow
cellchat <- identifyCommunicationPatterns(cellchat, pattern = "outgoing", k = 3)
netAnalysis_river(cellchat, pattern = "outgoing")Total strength comes from a few pathways
On PBMC 3k (trim = 0.1), MIF, GALECTIN and CypA account for 88% of the total strength. The "strongest sender" in the global circle plot mostly reflects these three pathways. Before interpreting the global network, look at the strength of each pathway in netP, and use signaling.exclude or plot only the pathways of interest when needed.
Signaling role analysis
netAnalysis_computeCentrality computes out-degree (sender), in-degree (receiver), flow betweenness (mediator) and information centrality (influencer) on each pathway network. The 2D scatter plot has outgoing strength on the x axis and incoming strength on the y axis, with dot size as the number of links; it summarizes which groups mainly send signals.
Pattern recognition and selectK
identifyCommunicationPatterns decomposes groups and pathways into k patterns with NMF. selectK defaults to k = 2:10 with 30 runs each; on this page 13 pathways and 9 groups took 78 seconds, and k = 2:6 with nrun = 5 took 34 seconds. With more than 50 pathways and 20 groups, selectK takes tens of minutes to hours. The official guidance is to choose the k before Cophenetic and Silhouette start to drop sharply; with very few pathways (fewer than 10) the pattern analysis adds little and can be skipped.
Probabilities have no absolute meaning
Communication probabilities depend on the global maximum expression, the Hill function parameters and the averaging method. Changing type or trim on the same data shifts all probabilities; probabilities are comparable between datasets only under the same parameters and normalization.
| Plot | Function | Question answered | Notes |
|---|---|---|---|
| Circle plot (global) | netVisual_circle(net$count or net$weight) | Which groups communicate most and which group is the main sender | Use edge.weight.max to share the scale when comparing two datasets |
| Pathway circle, chord or hierarchy plot | netVisual_aggregate(signaling = , layout = ) | Who sends and who receives one pathway | If the pathway is not significant: There is no significant communication of XXX; check cellchat@netP$pathways first |
| Heatmap | netVisual_heatmap(signaling = ) | Strength matrix of one pathway across all group pairs | Colors are scaled within the pathway and cannot be compared across pathways |
| Bubble plot | netVisual_bubble(sources.use, targets.use) | Which ligand-receptor pairs connect given groups, with probabilities and p values | Best suited for the main text of a paper; groups can be given by name or index |
| L-R chord diagram | netVisual_chord_gene | All pairs sent or received by one group | With many pairs the labels overlap; adjust small.gap or restrict signaling |
| Contribution plot | netAnalysis_contribution | Which pair contributes most to one pathway | — |
| Signaling roles | netAnalysis_signalingRole_scatter / _heatmap / _network | Whether each group is a sender, receiver, mediator or influencer | Run netAnalysis_computeCentrality first; the heatmap is row-scaled |
Step 6
CellChat compares groups by running each group separately and then merging. The prerequisites are: both groups use the same annotation (same column, same names), the same database subset and the same parameters; the cell type sets must match, otherwise run liftCellChat first.
run_cellchat <- function(obj) {
meta <- data.frame(labels = droplevels(obj$celltype), # drop types absent from this group
samples = factor(obj$sample), row.names = colnames(obj))
cc <- createCellChat(LayerData(obj, assay = "RNA", layer = "data"), meta = meta, group.by = "labels")
cc@DB <- subsetDB(CellChatDB.human, search = "Secreted Signaling", key = "annotation")
cc <- subsetData(cc)
cc <- identifyOverExpressedGenes(cc)
cc <- identifyOverExpressedInteractions(cc)
cc <- computeCommunProb(cc, type = "truncatedMean", trim = 0.1) # use identical parameters for both groups
cc <- filterCommunication(cc, min.cells = 10)
cc <- computeCommunProbPathway(cc)
cc <- aggregateNet(cc)
netAnalysis_computeCentrality(cc, slot.name = "netP")
}
# both groups use the same labels (same annotation column, same spelling)
seu$celltype <- factor(seu$celltype)
cc.ctrl <- run_cellchat(subset(seu, stim == "CTRL"))
cc.stim <- run_cellchat(subset(seu, stim == "STIM"))
# if one group lacks a cell type, lift both to the same set of labels
grp <- union(levels(cc.ctrl@idents), levels(cc.stim@idents))
if (!identical(levels(cc.ctrl@idents), levels(cc.stim@idents))) {
cc.ctrl <- liftCellChat(cc.ctrl, group.new = grp)
cc.stim <- liftCellChat(cc.stim, group.new = grp)
}
object.list <- list(CTRL = cc.ctrl, STIM = cc.stim)
cellchat <- mergeCellChat(object.list, add.names = names(object.list))
compareInteractions(cellchat, show.legend = FALSE, group = c(1, 2)) +
compareInteractions(cellchat, show.legend = FALSE, group = c(1, 2), measure = "weight")
netVisual_diffInteraction(cellchat, weight.scale = TRUE, measure = "weight") # red = stronger in STIM
rankNet(cellchat, mode = "comparison", measure = "weight", stacked = TRUE, do.stat = TRUE)
netVisual_bubble(cellchat, sources.use = "CD14 Mono", comparison = c(1, 2),
max.dataset = 2, title.name = "Increased in STIM", angle.x = 45, remove.isolate = TRUE)Factor levels: drop per group, then lift
A common approach is to give both groups the same factor levels to keep the order. When one group lacks a cell type, this makes computeCommunProb fail with the factor-levels error. The correct order is droplevels in each group, then liftCellChat(group.new = union(...)) after computing. Without lifting, mergeCellChat itself does not fail, but netVisual_diffInteraction later gives non-conformable arrays.
Two-group test on this page
Kang 2018 data (6,548 CTRL cells, 7,451 STIM cells, 13 types), Secreted, trim = 0.1: CTRL 421 interactions and 13 pathways, STIM 542 and 14; total strength from 4.30 to 10.79. Pairs increased from CD14 monocytes include CXCL10/CXCL11–CXCR3, CCL8–CCR1 and CCL7–CCR1, as expected for interferon-induced chemokines. Each group took 26–43 seconds, and the two-group workflow peaked at 3.3–3.6 GB of memory.
Two kinds of "differential"
netVisual_bubble(comparison, max.dataset) and rankNet compare communication probabilities; netMappingDEG selects pairs whose ligand or receptor is up-regulated between conditions. The logFC computed by presto in v2 is smaller than in v1, so the official tutorial lowered thresh.fc from 0.1 to 0.05. The same pair can appear among both up- and down-regulated results because the differential test runs within each group; set group.DE.combined = TRUE to ignore groups.
The unit of the statistical test
The paired Wilcoxon test in rankNet(do.stat = TRUE) uses group pairs as units and ignores biological replicates. With one merged sample per group, a "significantly increased pathway" only describes this sequencing run. With several samples, filter with min.samples, or add MultiNicheNet or LIANA+, which model at the sample level.
Datasets with very different composition
For example, embryonic skin versus adult wound: netVisual_diffInteraction and functional similarity (computeNetSimilarityPairwise(type = "functional")) no longer apply; structural similarity (type = "structural") and separate circle plots still work.
Use a shared scale in plots
When drawing circle plots for each group, take the maxima across groups with getMaxWeight(object.list, attribute = c("idents", "count")) and pass them to edge.weight.max and vertex.weight.max; otherwise each plot is scaled separately and line widths cannot be compared.
Troubleshooting
The error messages below were reproduced on this page or come from GitHub issues and the Nature Protocols troubleshooting table, in workflow order.
Parallel runs: on 2,638 cells, future::plan("multisession", workers = 4) made identifyOverExpressedGenes go from 0.1 to 25.1 seconds and computeCommunProb from 5.8 to 26.3 seconds on this page. Several users in GitHub #357 report the same (0.9 to 368 seconds on Ubuntu). PR #438 (merged 2026-02-19) parallelized the permutation loop, and its author reports about 46 hours down to 5.2 hours for 40,000 cells on 64 cores. Use plan("sequential") below about 10,000 cells; for large data, first confirm that main after 2026-02-19 is installed, then use multisession with options(future.globals.maxSize = 4e9). Sources disagree on run time: Nature Protocols reports about 15 minutes of inference on a skin atlas of about 300,000 cells, while several GitHub users report hours for tens of thousands of cells; the differences come from the number of groups, database subset, nboot, version and parallel settings. Another user reports that steps up to computeCommunProb took about 4 hours for 90,000 cells and 6 groups on one core with 500 GB of memory (GitHub #167).
Large data can also be downsampled first: subset(obj, downsample = 500) in Seurat or sketchData in CellChat, sampling large groups and keeping small ones. The list of significant interactions is usually stable after downsampling, while strengths change.
| Error message | Cause | Fix |
|---|---|---|
| Cell labels cannot contain `0`! | Seurat cluster IDs used directly as labels | factor(paste0("C", seurat_clusters)) |
| Please check `unique(object@idents)` and ensure that the factor levels are correct! | NA labels or unused factor levels (most often after subset) | Remove cells with NA; meta$labels <- droplevels(meta$labels) |
| invalid 'row.names' length (Seurat object input) | Seurat v5 layers are split | Run JoinLayers(obj) before creating the object |
| Error in if (sum(P1) == 0) : missing value where TRUE/FALSE needed | Negative values or NA in the input, such as scale.data, integrated or a matrix scaled in Python (GitHub #79) | Use the log-normalized data layer |
| Error in integer(n) : invalid 'length' argument | In v2 builds before November 2023, the LTE4-DPEP1 complex with the full database (GitHub #1); Seurat also has a function named subsetData | Update CellChat; call CellChat::subsetData(cellchat) |
| There is no significant communication of XXX | The pathway is not significant in this data | Choose pathways from cellchat@netP$pathways |
| netVisual_circle: need finite 'xlim' values | net$count is all zero: no significant interactions | Check the rows of subsetCommunication(cellchat); check species database, gene names and labels; relax type or trim |
| A group has no interactions after aggregateNet | The group has at most min.cells cells and was filtered, or it truly has no significant signals | Check table(cellchat@idents); keep groups aligned with liftCellChat when comparing |
| The input `color.use` should be a named vector! | Custom colors unnamed or not matching the groups | color.use = setNames(colors, levels(cellchat@idents)) |
| Length of new attribute value... (netVisual_circle) | Compatibility with igraph 1.4.0 (Nature Protocols Table 1) | Downgrade igraph to 1.3.5, run updateCellChat(), or reinstall CellChat |
| netAnalysis_computeCentrality fails on a merged object | The function must run on each single object (Nature Protocols Table 1) | Run netAnalysis_computeCentrality on every object in object.list, then mergeCellChat |
| Several times slower after enabling multiple cores | Inter-process transfer overhead of multisession; versions before February 2026 parallelized only a small part | See the next paragraph |
Tool choice
Dimitrov et al. (Nature Communications 2022) compared 16 resources and 7 methods and concluded that both the method and the resource strongly change the predictions. Liu et al. (Genome Biology 2022) evaluated 16 methods against spatial transcriptomics; CellChat, CellPhoneDB, NicheNet and ICELLNET performed better overall, and the authors recommend cross-checking with at least two methods.
| Tool | Language | Question answered | Suited for | Not suited for |
|---|---|---|---|---|
| CellChat v2 | R | Which ligand-receptor pairs and pathways are significant between groups; network roles and patterns | Seurat users; many ready-made plots and two-group comparisons | Statistical modeling of many samples and conditions |
| CellPhoneDB v5 | Python | Which pairs are specific to given group pairs, accounting for multi-subunit complexes | Human data; linking to Vento-Tormo lab atlases and spatial microenvironment methods | Mouse data (requires ortholog conversion) |
| NicheNet / MultiNicheNet | R | Which ligands explain downstream gene changes in receiver cells | Existing treatment-versus-control DE genes, looking for upstream ligands; MultiNicheNet supports multiple samples and conditions | Exploratory description without a control group |
| LIANA+ | Python (R version liana also available) | Consensus ranking across methods and resources; multi-sample analysis via Tensor-cell2cell and MOFA | Scanpy users; scores from CellChat, CellPhoneDB and other methods at once | When only CellChat-style plots are needed |
- CellChat results come from mRNA expression and a prior database. They show that expression permits communication; they do not prove an interaction at the protein level.
- Single-cell data lose spatial position. Two groups expressing a ligand and its receptor are not necessarily adjacent in the tissue; spatial data or in situ experiments fill this gap.
- Pairs outside the database never appear, and CellChatDB favors well-studied signals.
- Strength values depend on parameters and normalization and should be compared only under the same parameters. A paper should report the database subset, type, trim, population.size, min.cells and version.
- Cross-condition analysis in CellChat is mainly pairwise; changes across several conditions or time series have to be organized by the user (a limitation listed in Nature Protocols).
- Common follow-up validation includes RNAscope of ligands or receptors, immunofluorescence co-localization, and receptor blocking or neutralizing antibody experiments.
Tested on this page
Tested on this page: macOS arm64 (Apple M2, 8 cores, 16 GB), R 4.5.3, Seurat 5.5.1, CellChat 2.2.0.9001, presto 1.1.0, NMF 0.28, ComplexHeatmap 2.26.1, future 1.76.0. Other jobs were running on the same machine, so times indicate orders of magnitude only.
| Item | PBMC 3k | Kang 2018, two groups |
|---|---|---|
| Cells and groups | 2,638 cells, 9 groups | 6,548 CTRL / 7,451 STIM cells, 13 types |
| Default workflow (Secreted, triMean) | 152 interactions, 6 pathways; computeCommunProb 9.6 s; peak memory 1.28 GB | CTRL 158 / STIM 254 |
| truncatedMean 0.1 | 241 interactions, 13 pathways | CTRL 421 / STIM 542; 26–43 s per group |
| Full workflow of the code blocks on this page | Without Non-protein + trim 0.1: 967 interactions, 34 pathways; 56 s including plots, 1.49 GB | Two-group comparison 73 s, 3.30 GB |
| selectK | Default 78 s; k = 2:6, nrun = 5 took 34 s | Not run |
- Errors reproduced on this page: label 0, NA labels, unused factor levels, split layers, identical factor levels with a missing type in one group, plotting after merging without lifting.
- Silent errors: counts input, scale.data passed as a matrix, and LayerData on split layers returning only the first layer all run without error.
- Spatial transcriptomics (datatype = "spatial", SpatialCellChat), smoothData, CellPhoneDB, NicheNet and LIANA+ were not run here; only function and parameter names were checked against the documentation.
Chinese community
The following come from Chinese blog posts republished on the Tencent Cloud developer community. Only items with error messages or the author's own runs, and consistent with the official documentation or the tests on this page, are included.
Cluster ID 0 cannot be a label
生信补给站 (November 2023) renamed clusters 0–9 to C0–C9 on Visium data before the run succeeded. The error reproduced on this page is Cell labels cannot contain `0`!, and the maintainer also noted in GitHub #16 that labels cannot be "0".
The "unused factor" error was really NA
天意生信云 (March 2025) hit the factor-levels error, found that every level had cells, and traced it to NA values in the annotation column; removing those cells fixed it. On this page NA and unused levels give the same error, so check both: sum(is.na(meta$labels)) and setdiff(levels(meta$labels), unique(meta$labels)).
Compiler error on Apple silicon Macs
R小白 (May 2023) hit no member named 'Rlog1p' in namespace 'std' on an M-series Mac, and export PKG_CPPFLAGS="-DHAVE_WORKING_LOG1P" fixed it; CMake was not found on the PATH was fixed with conda install cmake. The installation on this page with conda compilers and R 4.5.3 did not hit these, but they still apply to older R setups.
nboot = 20 for quick checks
The spatial tutorial by 生信技能树 (June 2025) uses nboot = 20 while tuning. On this page nboot = 20 and 100 shared 96% of significant interactions (930 of 967) at about one fifth of the time, which is enough for choosing the database and trim; use 100 for final results.
"v3" is a separate spatial package
KS科研 (May 2026) points out that SpatialCellChat is a spatial transcriptomics package split off from CellChat, with dependencies MERINGUE, ALRA, BiocNeighbors and RcppML, consistent with the official README; single-cell data still use CellChat v2.
Two outdated habits
Many Chinese tutorials still write computeCommunProb(trim = 0.1) without changing type, and use plan("multiprocess") on Windows. In the first case trim has no effect; the second fails in the current future with No such backend for futures: 'multiprocess'. Use plan("multisession") or run sequentially.
Hand it to an agent
You can describe the analysis in one sentence and let the Scientify scientific agent complete it in a cloud computer.
Example instruction: "Run a CellChat v2 analysis on data/annotated.rds (Seurat v5; the celltype column holds annotations, the condition column holds control and treatment, the sample column holds 6 samples): run the two conditions separately, exclude Non-protein Signaling, use both triMean and truncatedMean 0.1 and compare the results, set min.samples = 2 in filterCommunication; after merging, output the count and strength comparison, increased and decreased pathways, and a bubble plot of up-regulated pairs sent by macrophages, and summarize them in tables."
The agent installs CellChat and its dependencies in the cloud computer and records versions; checks whether layers are split, whether labels contain NA or unused levels, and whether the cell types match between groups, running liftCellChat when needed; runs both parameter sets and lists interactions found under only one of them; and outputs circle plots, bubble plots, signaling role plots and differential tables. The agent reviews the results adversarially, for example checking whether the strength is dominated by a few pathways such as MIF and whether significant interactions come from a single sample. The workspace keeps scripts, parameters, logs and rds files for reproduction, and the task continues after you shut down your computer.
You still need to check yourself whether the annotation is reliable, whether the database subset matches the research question, whether candidate pairs are plausible given the literature and the tissue, and which conclusions need experimental validation.
References
- CellChat GitHub (jinworks): README and installation — Current version, dependencies, relation between v2 and v3 (SpatialCellChat)
- CellChat NEWS — Changes in v2.0–v2.1.2: samples column, min.samples, CellChatDB v2
- CellChat single-dataset tutorial (CellChat-vignette) — Four input options, database composition, triMean and population.size, plotting functions
- CellChat tutorial: comparison of multiple datasets — mergeCellChat, rankNet, DEG mapping and thresh.fc = 0.05
- CellChat tutorial: datasets with different cellular compositions — liftCellChat; analyses unavailable for very different compositions
- CellChat source R/modeling.R — Definitions of triMean, truncatedMean and thresholdedMean; the <= test in filterCommunication
- Jin et al. 2024, Nature Protocols: CellChat for systematic analysis of cell–cell communication — CellChat v2 methods and procedures; trim advice, limitations, troubleshooting table, timing
- Jin et al. 2021, Nature Communications: Inference and analysis of cell-cell communication using CellChat — Communication probability model, permutation test, network analysis
- SpatialCellChat(CellChat v3) — Single-cell-resolution spatial transcriptomics
- GitHub issue #357: multisession is slower — Community report: several users find parallel runs slower; downgrading future
- GitHub PR #438: parallel permutation loop in computeCommunProb — Merged February 2026; 40k cells from 46 h to 5.2 h (PR author's data)
- GitHub issue #62: best practice for multiple samples — Maintainer recommends merged runs; origin of min.samples
- GitHub issue #185: projectData renamed smoothData — Community report: adj = PPI.human must be named
- GitHub issue #202: no slot of name data.smooth — Community report: adding an empty matrix to old objects
- GitHub issues #79, #16, #1: computeCommunProb errors — Negative input, NA and 0 labels, early v2 full-database error
- GitHub issue #167: speeding up CellChat — Community report: run time and memory for 90k cells and for 133 groups
- Dimitrov et al. 2022, Nature Communications: Comparison of methods and resources for cell-cell communication inference — Method and resource choices strongly affect results; LIANA
- Liu et al. 2022, Genome Biology: Evaluation of cell-cell interaction methods by integrating single-cell RNA sequencing data with spatial information — Benchmark of 16 methods; recommends cross-checking with at least two
- Kang et al. 2018, Nature Biotechnology (ifnb dataset) — Public data used for the two-group comparison on this page
- CellPhoneDB — Tool choice table
- MultiNicheNet — Ligand activity analysis for multiple samples and conditions
- LIANA+ documentation — Multi-method consensus and multi-sample analysis
- 生信补给站: CellChat v2 spatial transcriptomics communication analysis (Tencent Cloud, 2023-11, Chinese) — Community post: cluster IDs 0 renamed to C0–C9
- 天意生信云: CellChat factor-level error (Tencent Cloud, 2025-03, Chinese) — Community post: the real cause of the error was NA in the annotation column
- R小白: installing CellChat (Tencent Cloud, 2023-05, Chinese) — Community post: Rlog1p compile error on Apple silicon, missing CMake; applies to older R setups
- 生信技能树: CellChat v2 spatial transcriptomics analysis (Tencent Cloud, 2025-06, Chinese) — Community post: nboot = 20 for quick checks, Visium parameters
- KS科研: CellChat v3 spatial communication analysis (Tencent Cloud, 2026-05, Chinese) — Community post: SpatialCellChat dependencies and input