Finding disease clusters with SaTScan

SaTScan is free software for the spatial, temporal and space-time scan statistics. Public health teams use it to test whether cases are unusually concentrated in an area or a period, to evaluate reported cluster alarms, and to run ongoing outbreak surveillance. This guide covers how it works, which model to pick, and how to run it from its interface or from R.

How the scan statistic works

Imagine a circle centred on one location on the map. SaTScan grows it step by step, from zero up to a maximum size, and repeats this for every centre. Each circle is a candidate cluster. For each one it compares the cases observed inside with the number expected if risk were the same everywhere.

The candidate with the highest likelihood ratio is the most likely cluster. Under the Poisson model, for a window with c observed cases, E[c] expected cases and C cases in total:

LLR = c log(c / E[c]) + (C − c) log((C − c) / (C − E[c])) when c > E[c]

Because thousands of windows are tested, an ordinary p-value would be far too optimistic. SaTScan instead simulates many random datasets under the null hypothesis, runs the same search on each, and asks how often a random dataset produces a cluster as strong as the real one. That Monte Carlo p-value already accounts for the many windows searched.

Scan windows over a map of locations Dots mark locations. Three nested circles grow around one location; the middle circle, shaded, contains several high-rate locations and is the most likely cluster.
Windows grow around each location. The shaded circle holds more cases than expected and has the highest likelihood ratio.

Choosing a probability model

The model follows from the data you have, not from the question alone.

If your data are…UseHealth data example
Case counts with a known population at riskPoissonAsthma ED visits per local geographic area, with population from the census
Cases and controls, coded 0/1BernoulliAdmitted versus not admitted ED visits, or birth defects against unaffected births
Cases only, with dates, and no reliable populationSpace-time permutationDaily syndromic counts for outbreak surveillance
Ordered categoriesOrdinalCancer stage at diagnosis, or triage level
Unordered categoriesMultinomialPopulation age structure, or type of injury
Survival times, with or without censoringExponentialTime to death after diagnosis, or length of stay
Other continuous measurementsNormalAverage birth weight or a lab value by area

Poisson and Bernoulli analyses can adjust for categorical covariates such as age group and sex. SaTScan can also adjust for temporal trends, known clusters and missing data, and can scan several datasets together.

Analysis types

  • Purely spatial

    Where are cases concentrated, ignoring time? The usual first look at incidence or prevalence over a fixed period.

  • Purely temporal

    When did cases rise, ignoring place? Windows slide along the time axis instead of across the map.

  • Retrospective space-time

    Cylinders instead of circles: the base is an area and the height a time window. Answers "was there a cluster somewhere, at some time, in this historical period?"

  • Prospective space-time

    For repeated surveillance. Only clusters still active at the end of the data count, and the significance can be adjusted for the analyses already run in earlier weeks.

  • Spatial variation in temporal trends

    Finds areas where the trend over time differs from the rest, such as a place where rates are falling more slowly.

Running an analysis

A purely spatial Poisson analysis of ED asthma visits by Alberta local geographic area makes a good first project.

  1. Install SaTScan

    Download it from satscan.org. The software is free, but you register first and receive a download password by email. The standard Windows installer includes the graphical interface, a bundled Java runtime, the command-line program, the user guide and sample data. There are separate downloads for macOS and Linux.

  2. Prepare the input files

    Plain-text files with one record per line. Location IDs must match exactly across files.

    Case file (.cas): location, cases, date

    LGA_101  42  2025
    LGA_102  17  2025
    LGA_103  88  2025

    Population file (.pop): location, year, population

    LGA_101  2025  31250
    LGA_102  2025  9870
    LGA_103  2025  56400

    Coordinates file (.geo): location, latitude, longitude

    LGA_101  53.54  -113.49
    LGA_102  52.27  -113.81
    LGA_103  51.05  -114.07

    Bernoulli analyses add a control file (.ctl). A grid file (.grd) lets circles centre somewhere other than the data locations. Covariates go in extra columns of the case and population files. For areas, coordinates are usually population-weighted centroids. The values above are illustrative.

  3. Set up the analysis

    The interface has three tabs. On Input, point to the files, choose the coordinate system and set the study period. On Analysis, choose the analysis type, the probability model and whether to scan for high rates, low rates or both. On Output, name the results file and choose extra outputs such as maps.

  4. Check the advanced settings

    Two matter most. The maximum spatial cluster size defaults to 50% of the population at risk; a smaller cap, such as 10–25%, stops one huge circle from swallowing half the province. The number of Monte Carlo replications defaults to 999; use more when you need a precise small p-value.

  5. Run and read the report

    The main results file lists the most likely cluster and any secondary clusters: the locations included, centre and radius, observed and expected cases, observed/expected ratio, relative risk, log likelihood ratio and p-value. Companion files give the same information per cluster and per location for joining back to your data.

  6. Map the clusters

    SaTScan can write KML for Google Earth and shapefiles for GIS software such as ArcGIS or QGIS, so clusters can be drawn over the original areas.

Running it from R

The rsatscan package on CRAN does not reimplement the method. It writes R data frames into SaTScan's file formats, writes the parameter file, calls your installed copy of SaTScan, and reads the results back into R. That makes analyses scriptable and reproducible in R Markdown or Quarto.

Parameters are numeric codes from SaTScan's parameter file, as in the comments in the example. Check them against your version: set the analysis up once in the interface, save the parameter file, and read the codes and their explanations in a text editor.

If you only need a quick purely spatial Poisson or Bernoulli scan inside R, the kulldorff() function in the SpatialEpi package is a lighter alternative.

Purely spatial Poisson scan with rsatscan

library(rsatscan)
td <- tempdir()

# cas: data frame of LGA, cases, year
# pop: data frame of LGA, year, population
# geo: data frame of LGA, latitude, longitude
write.cas(cas, td, "asthma")
write.pop(pop, td, "asthma")
write.geo(geo, td, "asthma")

invisible(ss.options(reset = TRUE))
ss.options(list(
  CaseFile        = "asthma.cas",
  PopulationFile  = "asthma.pop",
  CoordinatesFile = "asthma.geo",
  CoordinatesType = 1,  # latitude/longitude
  StartDate       = "2025/1/1",
  EndDate         = "2025/12/31",
  AnalysisType    = 1,  # purely spatial
  ModelType       = 0,  # discrete Poisson
  ScanAreas       = 1   # high rates
))
write.ss.prm(td, "asthma")

res <- satscan(td, "asthma", sslocation = "path/to/SaTScan")
summary(res)
res$col   # one row per cluster

Reading the results responsibly

  • A cluster is not a cause

    A significant cluster says cases are more concentrated than chance would suggest under the model. It does not say why. Differences in testing, coding, access to care or population counts can all produce one.

  • The expected count is only as good as the denominator

    Poisson results depend on the population file. Out-of-date or misallocated populations create spurious clusters, so use the boundary vintage that matches your case data.

  • Circles are an approximation

    A long, thin cluster along a river or highway will be split or padded by circular windows. Elliptic windows or a non-Euclidean neighbours file can help.

  • Aggregation changes the answer

    The same cases grouped by LGA, by health status area or by postal code can give different clusters. Report the level you used and, where you can, check whether the finding holds at another.

  • Adjust before you scan

    Without age and sex adjustment, an older or younger area can look like a cluster purely because of its population structure.

  • Secondary clusters need care

    Their p-values are conservative and can overlap the most likely cluster. Decide in advance how you will report them.

Learn more

The method paper to cite is Kulldorff M. A spatial scan statistic. Communications in Statistics: Theory and Methods 1997;26:1481–1496. Cite the software version you used as well. SaTScan is a trademark of Martin Kulldorff; this page is an independent guide.