Single-cell transcriptomics / Cell-cell communication

CellChat Tutorial: From a Seurat Object to Communication Networks, Pathways and Group Comparisons

This page is written for CellChat 2.2.0.9001 (GitHub main, March 2026) and Seurat 5.5.1, and its code runs on 10x PBMC 3k and on the two-condition interferon-stimulated PBMC data from Kang 2018. It focuses on what the official tutorials leave out: the differences between v1 and v2, how to choose parameters and how much they change the number of interactions, how to read the plots, and how to handle group comparisons and common errors.

Short answer

The standard CellChat v2 workflow is: extract the log-normalized matrix from a Seurat v5 object with LayerData(layer = "data") (run JoinLayers first; do not use counts or scale.data), and remove NA and unused factor levels from the cell type labels; use CellChatDB.human for human and CellChatDB.mouse for mouse, usually with subsetDB(CellChatDB) to exclude Non-protein Signaling; then run subsetData, identifyOverExpressedGenes, identifyOverExpressedInteractions, computeCommunProb, filterCommunication(min.cells = 10), computeCommunProbPathway and aggregateNet. The default type = "triMean" requires a gene to be expressed in at least 25% of the cells of a group, so shallow data yield few interactions; type = "truncatedMean", trim = 0.1 lowers the threshold to 10%. For group comparisons, run each group separately with identical parameters; if the cell types differ, run liftCellChat before mergeCellChat.

Check the version first

The CellChat repository moved from sqjin/CellChat to jinworks/CellChat. The current GitHub main is 2.2.0.9001 (last commit 2026-03-04), and the latest release tag is v2.1.2. What is labeled "v3" online is the separate SpatialCellChat package for spatial transcriptomics; single-cell data still use CellChat v2.

Record packageVersion("CellChat") and the installation date in the methods section. GitHub main also changes between releases: PR #438, merged on 2026-02-19, changed the parallel behavior, and rankNet was modified on 2026-03-04.

Itemv1-era tutorials (2021–2023)Current v2Result of copying the old code
Database sizeCellChatDB.human about 1,939 pairs3,233 pairs: 1,920 from v1, 1,313 added in v2 (994 Non-protein Signaling, 233 Cell-Cell Contact, 83 Secreted); mouse 3,379 pairsUsing the full database mixes metabolite and neurotransmitter signals into the results
Over-expression testSeurat-style Wilcoxondo.fast = TRUE by default calls presto; without presto the function stopsSame PBMC 3k data: presto gives 400 over-expressed signaling genes, do.fast = FALSE gives 744; v1 and v2 results are not directly comparable
PPI smoothingprojectData(cellchat, PPI.human)Renamed smoothData; adj = PPI.human must be namedError: is(adj, "matrix") || is(adj, "sparseMatrix") is not TRUE (GitHub #185)
Object slotdata.projectdata.smoothOld objects fail with no slot of name "data.smooth"; not fixed by the maintainer, users set obj@data.smooth <- matrix(nrow = 0, ncol = 0) (GitHub #202)
metaLabel column onlyAlso a samples column; if missing, it is set to sample1 with a warningAfter merging samples, min.samples filtering is unavailable
Parallelplan("multiprocess")future has removed the multiprocess backendError: No such backend for futures: 'multiprocess'
Spatial transcriptomicsVisium support since v1.6v2 supports several spatial technologies; single-cell-resolution data use SpatialCellChat (v3)—

Step 1

CellChat is installed only from GitHub and contains C++ code, so it needs a compiler. The dependencies that fail most often are NMF (CRAN, requires Biobase), ComplexHeatmap and BiocNeighbors (Bioconductor), and presto (GitHub). With the other dependencies installed from conda, compiling NMF, presto and CellChat took 2.5 minutes in total on this page, without errors. Nature Protocols gives a dual-core CPU and 16 GB of RAM as the minimum.

Installation tested on this pagebash
# Option A: dependencies from conda (tested on this page, macOS arm64, about 4 min + 2.5 min compiling)
micromamba create -y -n cellchat -c conda-forge -c bioconda \
  r-base=4.5 r-seurat r-remotes r-rcppeigen c-compiler cxx-compiler fortran-compiler make \
  r-future r-future.apply r-pbapply r-irlba r-ggalluvial r-svglite r-ggrepel r-circlize \
  bioconductor-complexheatmap bioconductor-biocneighbors bioconductor-biobase \
  r-rspectra r-reticulate r-sna r-fnn r-shape r-ggpubr r-ggnetwork r-plotly r-shiny r-bslib \
  r-collapse r-pkgmaker r-registry r-rngtools r-gridbase r-doparallel r-cowplot r-uwot
micromamba run -n cellchat Rscript -e '
options(repos = c(CRAN = "https://cloud.r-project.org"))
install.packages("NMF")                                        # conda-forge has no r-nmf
remotes::install_github("immunogenomics/presto", upgrade = "never")
remotes::install_github("jinworks/CellChat", upgrade = "never")
packageVersion("CellChat")'

# Option B: an existing R 4.4 or later
# install.packages(c("BiocManager", "remotes", "NMF"))
# BiocManager::install(c("ComplexHeatmap", "BiocNeighbors", "Biobase"))
# remotes::install_github(c("immunogenomics/presto", "jinworks/CellChat"), upgrade = "never")
Error or symptomCauseFix
Skipping 3 packages not available: Biobase, BiocNeighbors, BiocGenericsremotes does not install Bioconductor packagesFirst run BiocManager::install(c("Biobase", "BiocNeighbors", "BiocGenerics", "ComplexHeatmap"))
Failed to connect to github.com / download timeoutUnstable GitHub access from ChinaInstall dependencies from conda or a Bioconductor mirror so that only presto and CellChat come from GitHub; if that still fails, download the source zip and run R CMD INSTALL
installation of package ... had non-zero exit status (macOS)Missing or broken Xcode command line toolsReinstall the Xcode command line tools (confirmed by two users in GitHub #226), or use conda's c-compiler and cxx-compiler
no member named 'Rlog1p' in namespace 'std'Old R with the compiler on Apple siliconexport PKG_CPPFLAGS="-DHAVE_WORKING_LOG1P" and reinstall (Chinese community post, 2023)
CMake was not found on the PATHA dependency needs CMakeconda install cmake
dependencies 'RSpectra', 'RcppEigen', 'ggpubr', 'BiocNeighbors' are not availableCRAN and Bioconductor dependencies incompleteInstall them with BiocManager::install(), or use the r-* packages from conda-forge
For a faster implementation of the Wilcoxon Test, please install the presto packagev2 uses presto by defaultremotes::install_github("immunogenomics/presto"); if it cannot be installed, identifyOverExpressedGenes(do.fast = FALSE), which gives different results from presto

Step 2

CellChat needs two inputs: a log-normalized expression matrix (genes in rows, cells in columns) and a group label for each cell. The code below was checked on Seurat 5.5.1; the comments name the error each line prevents.

Create a CellChat object from a Seurat v5 objectr
library(Seurat)
library(CellChat)

seu <- readRDS("pbmc3k_annot.rds")          # annotated Seurat v5 object
seu <- JoinLayers(seu)                       # join layers first, otherwise only the first layer is read
seu$celltype <- droplevels(factor(seu$celltype))
stopifnot(!anyNA(seu$celltype))              # labels must not contain NA

data.input <- LayerData(seu, assay = "RNA", layer = "data")   # log-normalized values, not counts
stopifnot(min(data.input) >= 0)              # scale.data or integrated cannot be used
meta <- data.frame(
  labels  = seu$celltype,
  samples = factor(seu$orig.ident),          # v2 expects a samples column
  row.names = colnames(seu)
)
# prefix numeric cluster IDs; CellChat rejects the label "0"
# meta$labels <- factor(paste0("C", seu$seurat_clusters))

cellchat <- createCellChat(object = data.input, meta = meta, group.by = "labels")
table(cellchat@idents)

Why normalized data

CellChat divides the matrix by its global maximum and then applies a Hill function (Kh = 0.5) to compute probabilities. With counts the maximum is in the hundreds, so almost all scaled values are close to 0. Tested on this page: PBMC 3k with counts runs without error and gives 170 interactions (152 with correct input) and the same pathways, but the total strength is about 1/800 of the correct value. Strength comparisons between groups and datasets become meaningless.

Neither scale.data nor integrated

When a Seurat object is passed directly, createCellChat checks for negative values and stops; when a matrix is passed, it does not check. On this page, passing scale.data (2,000 variable genes, negative values) as a matrix ran to the end and gave 201 interactions without any message. The integrated assay holds only a subset of genes after correction, and CellChat warns about it. The data layer of the SCT assay contains log1p-transformed corrected counts and can be used.

Objects with split layers

A multi-sample Seurat v5 object can be split (data.A, data.B) before or after IntegrateLayers. Tested on this page: passing the object directly gives invalid 'row.names' length; LayerData(layer = "data") only warns and returns the first layer with 1,319 of 2,638 cells. Run JoinLayers first.

Four requirements for labels

No NA; no unused factor levels; no label 0 (Seurat cluster IDs used directly give Cell labels cannot contain `0`!); preferably more than 10 cells per group. The first two give the same error: Please check `unique(object@idents)` and ensure that the factor levels are correct!

Label granularity: major types or subclusters

CellChat computes between groups and averages away heterogeneity within a group. Nature Protocols lists this as a limitation and suggests subclustering before running CellChat when needed. More groups increase run time and plot complexity: in GitHub #167 a user with 133 groups and 11,000 cells reported more than 11 hours. A common approach is to obtain the global network with major types first, then split the types of interest into subclusters and run them separately. No specific group-number recommendation from the maintainers was found.

Organizing multiple samples

In GitHub #62 the maintainer sqjin recommends merging all samples of one condition into one object, as in integrated clustering. To reduce interactions found in only one sample, add a samples column to meta and use filterCommunication(min.samples = 2) to require an interaction in at least two samples (since v2.1.2). A user in the same issue reported that the merged run had 4–5% unique interactions not found in any per-sample run.

Extra input for spatial transcriptomics

createCellChat(datatype = "spatial", coordinates = ..., spatial.factors = data.frame(ratio = ..., tol = ...)): coordinates are pixel positions in the full-resolution image, ratio converts pixels to microns (65 / spot_diameter_fullres for Visium), and tol is half the spot diameter (32.5). In computeCommunProb, interaction.range = 250 and contact.range = 100 for Visium; about 10 for single-cell-resolution technologies. Single-cell-resolution spatial data use SpatialCellChat.

Step 3

CellChatDB sorts ligand-receptor pairs into four categories. The official tutorial uses only Secreted Signaling as an example. The choice of subset changes the number of results more than any algorithm parameter.

Pros and cons of Secreted Signaling only

It gives fewer results, cleaner plots and comparability with earlier papers. The cost is that the main contact signals between immune cells (MHC-I/II, CD45, ICAM, SELPLG and others) are missing: on PBMC 3k, adding Cell-Cell Contact raised the pathways from 13 to 34. For immune synapses, T cell–APC interactions or tumor immune checkpoints, use subsetDB(CellChatDB).

Non-protein Signaling is excluded by default

The official tutorial advises against using the whole CellChatDB. Ligand levels in this category are inferred from synthesizing enzymes and transporters because metabolites themselves are not measured. When neurotransmitter signals are needed in nervous system studies, add them separately and interpret them apart from protein signals.

Species and gene names

Use CellChatDB.human for human, CellChatDB.mouse for mouse; CellChatDB.zebrafish also exists. Gene names must be official symbols: with Ensembl IDs or mixed human and mouse names, almost no signaling genes remain after subsetData, and downstream steps fail with errors such as data.use[RsubunitsV, ]. Convert orthologs first for other species.

Adding custom ligand-receptor pairs

Use updateCellChatDB() to merge your own interaction, complex, cofactor and geneInfo tables (official Update-CellChatDB tutorial). Pairs without a pathway assignment do not appear in pathway-level plots.

CategoryHuman pairsDescriptionPBMC 3k on this page (triMean / truncatedMean 0.1)
Secreted Signaling1,280Secreted paracrine and autocrine signals152 / 241 interactions, 6 / 13 pathways
ECM-Receptor424Extracellular matrix-receptor; important in tissues, rare in blood—
Cell-Cell Contact535Requires direct contact, such as MHC, ICAM, CD86First three categories together (subsetDB(CellChatDB)): 437 / 967 interactions, 14 / 34 pathways
Non-protein Signaling994Metabolites and neurotransmitters; ligand level inferred from synthesizing enzymes and transportersFull database: 445 / 1,055 interactions, of which 8 / 88 Non-protein

Step 4

The core workflow below runs on PBMC 3k. The table after the code explains each parameter and how to choose its value; the next section shows the measured effects.

Infer the communication networkr
CellChatDB <- CellChatDB.human               # use CellChatDB.mouse for mouse
cellchat@DB <- subsetDB(CellChatDB)          # all protein ligand-receptor pairs, without Non-protein Signaling
# cellchat@DB <- subsetDB(CellChatDB, search = "Secreted Signaling", key = "annotation")

future::plan("sequential")                   # sequential is faster on small datasets
cellchat <- subsetData(cellchat)
cellchat <- identifyOverExpressedGenes(cellchat)          # uses presto for the Wilcoxon test by default
cellchat <- identifyOverExpressedInteractions(cellchat)

cellchat <- computeCommunProb(cellchat,
                              type = "truncatedMean", trim = 0.1,  # trim only takes effect with truncatedMean
                              population.size = FALSE)              # TRUE is an option for unsorted samples
cellchat <- filterCommunication(cellchat, min.cells = 10)
cellchat <- computeCommunProbPathway(cellchat)
cellchat <- aggregateNet(cellchat)

df.net  <- subsetCommunication(cellchat)                      # ligand-receptor level
df.netP <- subsetCommunication(cellchat, slot.name = "netP")  # pathway level
nrow(df.net); cellchat@netP$pathways
saveRDS(cellchat, "cellchat_pbmc3k.rds")
ParameterDefaultRole and how to choose
thresh.p in identifyOverExpressedGenes0.05Selects over-expressed signaling genes per group by unadjusted p value; thresh.fc = 0 and thresh.pc = 0. A pair enters the calculation if either the ligand or the receptor is over-expressed
type in computeCommunProb"triMean"triMean is the mean of the quantiles Q1, Q2, Q2 and Q3: if fewer than 25% of cells express a gene, the group average is 0; at 25%–50% only Q3/4 remains. Fewer, stronger interactions
trim with type = "truncatedMean"0.1The mean after removing the fraction trim from each end: the group average is 0 if the expressing fraction is below trim. 0.1 means a 10% threshold, 0.05 means 5%. trim only takes effect with truncatedMean (and the undocumented thresholdedMean). Nature Protocols recommends trim = 0.1 in general and lowering trim when well-known signals of the system are missing
computeAveExpr(features = , type = , trim = )—Shows group averages of ligands and receptors of interest before fixing trim; a gene that is 0 under triMean but non-zero at trim = 0.1 is expressed in 10%–25% of the cells (the approach recommended in Nature Protocols)
population.sizeFALSETRUE multiplies probabilities by the proportions of the two groups, intended for unsorted samples; use FALSE for sorted or enriched samples
nboot100Number of permutations; p value resolution is 1/nboot. Measured: 20 runs 4.3 s and 938 interactions, 100 runs 22.6 s and 967, 1,000 runs 493 s and 972
min.cells in filterCommunication10Removes groups with at most min.cells cells (the source uses <=), so a group of exactly 10 cells is removed
raw.useTRUEFALSE uses expression smoothed by smoothData, meant to compensate for missing subunits in shallow data, at the risk of false positives

Parameter tests

Tested on this page: CellChat 2.2.0.9001, 2,638 cells in 9 groups from the official Seurat PBMC 3k workflow (13 platelets, 32 DCs), nboot = 100. "Interactions" are the rows returned by subsetCommunication (source group–target group–ligand-receptor pair). Nature Protocols compared triMean, trim 0.1 and trim 0.05 on atopic dermatitis skin data (5,011 cells, 12 groups) with the same direction: smaller trim gives more pairs and pathways, mostly weak signals.

triMean misses known signals in shallow data

On PBMC 3k (Cell Ranger 1.1, median 817 genes per cell) triMean finds no CCL, CXCL or IFN-II pathway. With trim = 0.1, IFNG–IFNGR1/IFNGR2 appears from NK cells to CD14 monocytes, FCGR3A monocytes and DCs, consistent with known IFN-γ secretion by NK cells. The gap between the two methods is smaller in deeply sequenced data.

Smaller trim, more false positives

At trim = 0.1 the sources of CCL5–CCR1 include B cells and naive CD4 T cells, in which only a few cells express CCL5. Significance is relative to random label permutation, so a ligand just above the threshold can be significant. Results at trim = 0.05 suit exploration; conclusions for a paper should hold at both trim = 0.1 and triMean.

population.size changes rankings, not the list

With truncatedMean 0.1, the significant interactions for TRUE and FALSE match row by row (241), while the total strength drops from 5.71 to 0.058. Sender ranking: DC first with FALSE; Memory CD4 T first and DC seventh with TRUE. DCs have only 32 cells, so this parameter decides whether a small group with strong signals ranks high.

min.cells mainly affects small groups

Platelets have 13 cells: min.cells = 10 keeps them, and at trim = 0.05 there are 45 interactions involving platelets; min.cells = 20 removes them all. identifyOverExpressedGenes also has min.cells = 10, which affects only the differential test.

Databasetype / trimInteractionsL-R pairsPathwaysInvolving platelets (min.cells = 10)Interactions at min.cells = 20
SecretedtriMean1521061151
SecretedtruncatedMean 0.251511063148
SecretedtruncatedMean 0.1241191313228
SecretedtruncatedMean 0.05402372345357
Without Non-proteintriMean437411436401
Without Non-proteintruncatedMean 0.1967793498869
Without Non-proteintruncatedMean 0.051,375116511861,189
Full databasetruncatedMean 0.051,527130572031,324
With population.size TRUE or FALSE the interaction rows were identical for every combination, so they are not listed separately.

Step 5

aggregateNet produces two matrices: net$count is the number of significant ligand-receptor pairs between two groups, and net$weight is the sum of their communication probabilities. Count reflects how many kinds of signals there are, weight reflects the total signal, and their rankings often differ.

Plots and systems analysisr
groupSize <- as.numeric(table(cellchat@idents))
par(mfrow = c(1, 2), xpd = TRUE)
netVisual_circle(cellchat@net$count,  vertex.weight = groupSize, weight.scale = TRUE,
                 label.edge = FALSE, title.name = "Number of interactions")
netVisual_circle(cellchat@net$weight, vertex.weight = groupSize, weight.scale = TRUE,
                 label.edge = FALSE, title.name = "Interaction strength")

pathways.show <- "MIF"                       # check that it is in cellchat@netP$pathways first
netVisual_aggregate(cellchat, signaling = pathways.show, layout = "chord")
netVisual_heatmap(cellchat, signaling = pathways.show, color.heatmap = "Reds")
netVisual_bubble(cellchat, sources.use = c("CD14 Mono", "FCGR3A Mono"),
                 targets.use = c("CD8 T", "NK", "B"), remove.isolate = FALSE)

cellchat <- netAnalysis_computeCentrality(cellchat, slot.name = "netP")
netAnalysis_signalingRole_scatter(cellchat)
netAnalysis_signalingRole_heatmap(cellchat, pattern = "outgoing") +
  netAnalysis_signalingRole_heatmap(cellchat, pattern = "incoming")

library(NMF); library(ggalluvial)
selectK(cellchat, pattern = "outgoing", k.range = 2:6, nrun = 10)   # the default 2:10 with nrun = 30 is slow
cellchat <- identifyCommunicationPatterns(cellchat, pattern = "outgoing", k = 3)
netAnalysis_river(cellchat, pattern = "outgoing")

Total strength comes from a few pathways

On PBMC 3k (trim = 0.1), MIF, GALECTIN and CypA account for 88% of the total strength. The "strongest sender" in the global circle plot mostly reflects these three pathways. Before interpreting the global network, look at the strength of each pathway in netP, and use signaling.exclude or plot only the pathways of interest when needed.

Signaling role analysis

netAnalysis_computeCentrality computes out-degree (sender), in-degree (receiver), flow betweenness (mediator) and information centrality (influencer) on each pathway network. The 2D scatter plot has outgoing strength on the x axis and incoming strength on the y axis, with dot size as the number of links; it summarizes which groups mainly send signals.

Pattern recognition and selectK

identifyCommunicationPatterns decomposes groups and pathways into k patterns with NMF. selectK defaults to k = 2:10 with 30 runs each; on this page 13 pathways and 9 groups took 78 seconds, and k = 2:6 with nrun = 5 took 34 seconds. With more than 50 pathways and 20 groups, selectK takes tens of minutes to hours. The official guidance is to choose the k before Cophenetic and Silhouette start to drop sharply; with very few pathways (fewer than 10) the pattern analysis adds little and can be skipped.

Probabilities have no absolute meaning

Communication probabilities depend on the global maximum expression, the Hill function parameters and the averaging method. Changing type or trim on the same data shifts all probabilities; probabilities are comparable between datasets only under the same parameters and normalization.

PlotFunctionQuestion answeredNotes
Circle plot (global)netVisual_circle(net$count or net$weight)Which groups communicate most and which group is the main senderUse edge.weight.max to share the scale when comparing two datasets
Pathway circle, chord or hierarchy plotnetVisual_aggregate(signaling = , layout = )Who sends and who receives one pathwayIf the pathway is not significant: There is no significant communication of XXX; check cellchat@netP$pathways first
HeatmapnetVisual_heatmap(signaling = )Strength matrix of one pathway across all group pairsColors are scaled within the pathway and cannot be compared across pathways
Bubble plotnetVisual_bubble(sources.use, targets.use)Which ligand-receptor pairs connect given groups, with probabilities and p valuesBest suited for the main text of a paper; groups can be given by name or index
L-R chord diagramnetVisual_chord_geneAll pairs sent or received by one groupWith many pairs the labels overlap; adjust small.gap or restrict signaling
Contribution plotnetAnalysis_contributionWhich pair contributes most to one pathway—
Signaling rolesnetAnalysis_signalingRole_scatter / _heatmap / _networkWhether each group is a sender, receiver, mediator or influencerRun netAnalysis_computeCentrality first; the heatmap is row-scaled

Step 6

CellChat compares groups by running each group separately and then merging. The prerequisites are: both groups use the same annotation (same column, same names), the same database subset and the same parameters; the cell type sets must match, otherwise run liftCellChat first.

Two-group comparison (run on the Kang 2018 interferon-stimulated PBMC data)r
run_cellchat <- function(obj) {
  meta <- data.frame(labels = droplevels(obj$celltype),  # drop types absent from this group
                      samples = factor(obj$sample), row.names = colnames(obj))
  cc <- createCellChat(LayerData(obj, assay = "RNA", layer = "data"), meta = meta, group.by = "labels")
  cc@DB <- subsetDB(CellChatDB.human, search = "Secreted Signaling", key = "annotation")
  cc <- subsetData(cc)
  cc <- identifyOverExpressedGenes(cc)
  cc <- identifyOverExpressedInteractions(cc)
  cc <- computeCommunProb(cc, type = "truncatedMean", trim = 0.1)  # use identical parameters for both groups
  cc <- filterCommunication(cc, min.cells = 10)
  cc <- computeCommunProbPathway(cc)
  cc <- aggregateNet(cc)
  netAnalysis_computeCentrality(cc, slot.name = "netP")
}
# both groups use the same labels (same annotation column, same spelling)
seu$celltype <- factor(seu$celltype)
cc.ctrl <- run_cellchat(subset(seu, stim == "CTRL"))
cc.stim <- run_cellchat(subset(seu, stim == "STIM"))

# if one group lacks a cell type, lift both to the same set of labels
grp <- union(levels(cc.ctrl@idents), levels(cc.stim@idents))
if (!identical(levels(cc.ctrl@idents), levels(cc.stim@idents))) {
  cc.ctrl <- liftCellChat(cc.ctrl, group.new = grp)
  cc.stim <- liftCellChat(cc.stim, group.new = grp)
}
object.list <- list(CTRL = cc.ctrl, STIM = cc.stim)
cellchat <- mergeCellChat(object.list, add.names = names(object.list))

compareInteractions(cellchat, show.legend = FALSE, group = c(1, 2)) +
  compareInteractions(cellchat, show.legend = FALSE, group = c(1, 2), measure = "weight")
netVisual_diffInteraction(cellchat, weight.scale = TRUE, measure = "weight")  # red = stronger in STIM
rankNet(cellchat, mode = "comparison", measure = "weight", stacked = TRUE, do.stat = TRUE)
netVisual_bubble(cellchat, sources.use = "CD14 Mono", comparison = c(1, 2),
                 max.dataset = 2, title.name = "Increased in STIM", angle.x = 45, remove.isolate = TRUE)

Factor levels: drop per group, then lift

A common approach is to give both groups the same factor levels to keep the order. When one group lacks a cell type, this makes computeCommunProb fail with the factor-levels error. The correct order is droplevels in each group, then liftCellChat(group.new = union(...)) after computing. Without lifting, mergeCellChat itself does not fail, but netVisual_diffInteraction later gives non-conformable arrays.

Two-group test on this page

Kang 2018 data (6,548 CTRL cells, 7,451 STIM cells, 13 types), Secreted, trim = 0.1: CTRL 421 interactions and 13 pathways, STIM 542 and 14; total strength from 4.30 to 10.79. Pairs increased from CD14 monocytes include CXCL10/CXCL11–CXCR3, CCL8–CCR1 and CCL7–CCR1, as expected for interferon-induced chemokines. Each group took 26–43 seconds, and the two-group workflow peaked at 3.3–3.6 GB of memory.

Two kinds of "differential"

netVisual_bubble(comparison, max.dataset) and rankNet compare communication probabilities; netMappingDEG selects pairs whose ligand or receptor is up-regulated between conditions. The logFC computed by presto in v2 is smaller than in v1, so the official tutorial lowered thresh.fc from 0.1 to 0.05. The same pair can appear among both up- and down-regulated results because the differential test runs within each group; set group.DE.combined = TRUE to ignore groups.

The unit of the statistical test

The paired Wilcoxon test in rankNet(do.stat = TRUE) uses group pairs as units and ignores biological replicates. With one merged sample per group, a "significantly increased pathway" only describes this sequencing run. With several samples, filter with min.samples, or add MultiNicheNet or LIANA+, which model at the sample level.

Datasets with very different composition

For example, embryonic skin versus adult wound: netVisual_diffInteraction and functional similarity (computeNetSimilarityPairwise(type = "functional")) no longer apply; structural similarity (type = "structural") and separate circle plots still work.

Use a shared scale in plots

When drawing circle plots for each group, take the maxima across groups with getMaxWeight(object.list, attribute = c("idents", "count")) and pass them to edge.weight.max and vertex.weight.max; otherwise each plot is scaled separately and line widths cannot be compared.

Troubleshooting

The error messages below were reproduced on this page or come from GitHub issues and the Nature Protocols troubleshooting table, in workflow order.

Parallel runs: on 2,638 cells, future::plan("multisession", workers = 4) made identifyOverExpressedGenes go from 0.1 to 25.1 seconds and computeCommunProb from 5.8 to 26.3 seconds on this page. Several users in GitHub #357 report the same (0.9 to 368 seconds on Ubuntu). PR #438 (merged 2026-02-19) parallelized the permutation loop, and its author reports about 46 hours down to 5.2 hours for 40,000 cells on 64 cores. Use plan("sequential") below about 10,000 cells; for large data, first confirm that main after 2026-02-19 is installed, then use multisession with options(future.globals.maxSize = 4e9). Sources disagree on run time: Nature Protocols reports about 15 minutes of inference on a skin atlas of about 300,000 cells, while several GitHub users report hours for tens of thousands of cells; the differences come from the number of groups, database subset, nboot, version and parallel settings. Another user reports that steps up to computeCommunProb took about 4 hours for 90,000 cells and 6 groups on one core with 500 GB of memory (GitHub #167).

Large data can also be downsampled first: subset(obj, downsample = 500) in Seurat or sketchData in CellChat, sampling large groups and keeping small ones. The list of significant interactions is usually stable after downsampling, while strengths change.

Error messageCauseFix
Cell labels cannot contain `0`!Seurat cluster IDs used directly as labelsfactor(paste0("C", seurat_clusters))
Please check `unique(object@idents)` and ensure that the factor levels are correct!NA labels or unused factor levels (most often after subset)Remove cells with NA; meta$labels <- droplevels(meta$labels)
invalid 'row.names' length (Seurat object input)Seurat v5 layers are splitRun JoinLayers(obj) before creating the object
Error in if (sum(P1) == 0) : missing value where TRUE/FALSE neededNegative values or NA in the input, such as scale.data, integrated or a matrix scaled in Python (GitHub #79)Use the log-normalized data layer
Error in integer(n) : invalid 'length' argumentIn v2 builds before November 2023, the LTE4-DPEP1 complex with the full database (GitHub #1); Seurat also has a function named subsetDataUpdate CellChat; call CellChat::subsetData(cellchat)
There is no significant communication of XXXThe pathway is not significant in this dataChoose pathways from cellchat@netP$pathways
netVisual_circle: need finite 'xlim' valuesnet$count is all zero: no significant interactionsCheck the rows of subsetCommunication(cellchat); check species database, gene names and labels; relax type or trim
A group has no interactions after aggregateNetThe group has at most min.cells cells and was filtered, or it truly has no significant signalsCheck table(cellchat@idents); keep groups aligned with liftCellChat when comparing
The input `color.use` should be a named vector!Custom colors unnamed or not matching the groupscolor.use = setNames(colors, levels(cellchat@idents))
Length of new attribute value... (netVisual_circle)Compatibility with igraph 1.4.0 (Nature Protocols Table 1)Downgrade igraph to 1.3.5, run updateCellChat(), or reinstall CellChat
netAnalysis_computeCentrality fails on a merged objectThe function must run on each single object (Nature Protocols Table 1)Run netAnalysis_computeCentrality on every object in object.list, then mergeCellChat
Several times slower after enabling multiple coresInter-process transfer overhead of multisession; versions before February 2026 parallelized only a small partSee the next paragraph

Tool choice

Dimitrov et al. (Nature Communications 2022) compared 16 resources and 7 methods and concluded that both the method and the resource strongly change the predictions. Liu et al. (Genome Biology 2022) evaluated 16 methods against spatial transcriptomics; CellChat, CellPhoneDB, NicheNet and ICELLNET performed better overall, and the authors recommend cross-checking with at least two methods.

ToolLanguageQuestion answeredSuited forNot suited for
CellChat v2RWhich ligand-receptor pairs and pathways are significant between groups; network roles and patternsSeurat users; many ready-made plots and two-group comparisonsStatistical modeling of many samples and conditions
CellPhoneDB v5PythonWhich pairs are specific to given group pairs, accounting for multi-subunit complexesHuman data; linking to Vento-Tormo lab atlases and spatial microenvironment methodsMouse data (requires ortholog conversion)
NicheNet / MultiNicheNetRWhich ligands explain downstream gene changes in receiver cellsExisting treatment-versus-control DE genes, looking for upstream ligands; MultiNicheNet supports multiple samples and conditionsExploratory description without a control group
LIANA+Python (R version liana also available)Consensus ranking across methods and resources; multi-sample analysis via Tensor-cell2cell and MOFAScanpy users; scores from CellChat, CellPhoneDB and other methods at onceWhen only CellChat-style plots are needed
  • CellChat results come from mRNA expression and a prior database. They show that expression permits communication; they do not prove an interaction at the protein level.
  • Single-cell data lose spatial position. Two groups expressing a ligand and its receptor are not necessarily adjacent in the tissue; spatial data or in situ experiments fill this gap.
  • Pairs outside the database never appear, and CellChatDB favors well-studied signals.
  • Strength values depend on parameters and normalization and should be compared only under the same parameters. A paper should report the database subset, type, trim, population.size, min.cells and version.
  • Cross-condition analysis in CellChat is mainly pairwise; changes across several conditions or time series have to be organized by the user (a limitation listed in Nature Protocols).
  • Common follow-up validation includes RNAscope of ligands or receptors, immunofluorescence co-localization, and receptor blocking or neutralizing antibody experiments.

Tested on this page

Tested on this page: macOS arm64 (Apple M2, 8 cores, 16 GB), R 4.5.3, Seurat 5.5.1, CellChat 2.2.0.9001, presto 1.1.0, NMF 0.28, ComplexHeatmap 2.26.1, future 1.76.0. Other jobs were running on the same machine, so times indicate orders of magnitude only.

ItemPBMC 3kKang 2018, two groups
Cells and groups2,638 cells, 9 groups6,548 CTRL / 7,451 STIM cells, 13 types
Default workflow (Secreted, triMean)152 interactions, 6 pathways; computeCommunProb 9.6 s; peak memory 1.28 GBCTRL 158 / STIM 254
truncatedMean 0.1241 interactions, 13 pathwaysCTRL 421 / STIM 542; 26–43 s per group
Full workflow of the code blocks on this pageWithout Non-protein + trim 0.1: 967 interactions, 34 pathways; 56 s including plots, 1.49 GBTwo-group comparison 73 s, 3.30 GB
selectKDefault 78 s; k = 2:6, nrun = 5 took 34 sNot run
  • Errors reproduced on this page: label 0, NA labels, unused factor levels, split layers, identical factor levels with a missing type in one group, plotting after merging without lifting.
  • Silent errors: counts input, scale.data passed as a matrix, and LayerData on split layers returning only the first layer all run without error.
  • Spatial transcriptomics (datatype = "spatial", SpatialCellChat), smoothData, CellPhoneDB, NicheNet and LIANA+ were not run here; only function and parameter names were checked against the documentation.

Chinese community

The following come from Chinese blog posts republished on the Tencent Cloud developer community. Only items with error messages or the author's own runs, and consistent with the official documentation or the tests on this page, are included.

Cluster ID 0 cannot be a label

生信补给站 (November 2023) renamed clusters 0–9 to C0–C9 on Visium data before the run succeeded. The error reproduced on this page is Cell labels cannot contain `0`!, and the maintainer also noted in GitHub #16 that labels cannot be "0".

The "unused factor" error was really NA

天意生信云 (March 2025) hit the factor-levels error, found that every level had cells, and traced it to NA values in the annotation column; removing those cells fixed it. On this page NA and unused levels give the same error, so check both: sum(is.na(meta$labels)) and setdiff(levels(meta$labels), unique(meta$labels)).

Compiler error on Apple silicon Macs

R小白 (May 2023) hit no member named 'Rlog1p' in namespace 'std' on an M-series Mac, and export PKG_CPPFLAGS="-DHAVE_WORKING_LOG1P" fixed it; CMake was not found on the PATH was fixed with conda install cmake. The installation on this page with conda compilers and R 4.5.3 did not hit these, but they still apply to older R setups.

nboot = 20 for quick checks

The spatial tutorial by 生信技能树 (June 2025) uses nboot = 20 while tuning. On this page nboot = 20 and 100 shared 96% of significant interactions (930 of 967) at about one fifth of the time, which is enough for choosing the database and trim; use 100 for final results.

"v3" is a separate spatial package

KS科研 (May 2026) points out that SpatialCellChat is a spatial transcriptomics package split off from CellChat, with dependencies MERINGUE, ALRA, BiocNeighbors and RcppML, consistent with the official README; single-cell data still use CellChat v2.

Two outdated habits

Many Chinese tutorials still write computeCommunProb(trim = 0.1) without changing type, and use plan("multiprocess") on Windows. In the first case trim has no effect; the second fails in the current future with No such backend for futures: 'multiprocess'. Use plan("multisession") or run sequentially.

Hand it to an agent

You can describe the analysis in one sentence and let the Scientify scientific agent complete it in a cloud computer.

Example instruction: "Run a CellChat v2 analysis on data/annotated.rds (Seurat v5; the celltype column holds annotations, the condition column holds control and treatment, the sample column holds 6 samples): run the two conditions separately, exclude Non-protein Signaling, use both triMean and truncatedMean 0.1 and compare the results, set min.samples = 2 in filterCommunication; after merging, output the count and strength comparison, increased and decreased pathways, and a bubble plot of up-regulated pairs sent by macrophages, and summarize them in tables."

The agent installs CellChat and its dependencies in the cloud computer and records versions; checks whether layers are split, whether labels contain NA or unused levels, and whether the cell types match between groups, running liftCellChat when needed; runs both parameter sets and lists interactions found under only one of them; and outputs circle plots, bubble plots, signaling role plots and differential tables. The agent reviews the results adversarially, for example checking whether the strength is dominated by a few pathways such as MIF and whether significant interactions come from a single sample. The workspace keeps scripts, parameters, logs and rds files for reproduction, and the task continues after you shut down your computer.

You still need to check yourself whether the annotation is reliable, whether the database subset matches the research question, whether candidate pairs are plausible given the literature and the tissue, and which conclusions need experimental validation.

References

FAQ

CellChat finds very few interactions, only pathways such as MIF and GALECTIN. Is that normal?

It is common with the default triMean on shallow data. triMean requires a gene to be expressed in at least 25% of the cells of a group. On PBMC 3k on this page triMean gave 6 pathways, type = "truncatedMean", trim = 0.1 gave 13, and adding Cell-Cell Contact to the database gave 34. Also check the species database, gene names and labels.

Should CellChat input be counts or data?

Use the log-normalized data layer. Counts do not cause an error, but all probabilities become tiny (on this page the total strength was about 1/800 of the correct value), so strength comparisons fail. If only counts are available, use CellChat::normalizeData(counts), which gives exactly the same result as Seurat's LogNormalize.

Should population.size be TRUE or FALSE?

TRUE is an option for unsorted tissue samples; use FALSE for sorted or enriched samples. On this page it did not change the list of significant interactions, only strengths and rankings: with TRUE large groups rank higher as senders and small groups (such as DCs) rank lower.

After mergeCellChat, plotting fails with non-conformable arrays. What should I do?

The cell type sets of the two groups differ. Run droplevels when computing each group, then run liftCellChat(obj, group.new = union(levels(a@idents), levels(b@idents))) on the object lacking types, and then mergeCellChat.

CellChat is slow, and multiple cores make it slower. Why?

Below about 10,000 cells use future::plan("sequential"); on this page multisession made two steps more than 4 times slower. For large data, update to GitHub main after 2026-02-19 (PR #438 parallelized the permutation loop) before using multisession; you can also tune with nboot = 20 and downsample large groups.

CellChat or CellPhoneDB?

Use CellChat if you work in Seurat and need group comparisons and many plots; use CellPhoneDB for human data when complex specificity or links to Vento-Tormo lab atlases matter. Method and resource choices change results substantially, so cross-check key conclusions with two methods or the LIANA+ consensus ranking.

Hand cell-cell communication analysis to Scientify

The scientific agent installs CellChat in an isolated cloud computer, checks the input object, runs and compares several parameter sets, completes group comparisons and plots, and keeps all scripts, parameters and logs. New users receive 5 USD of free credit.