Skip to contents

treescanr turns emergency department (ED) visits into a TreeScan analysis in three steps: standardize the visits, build the count file, and run TreeScan. Each step takes the output of the previous one, so the whole analysis is a single |> pipeline.

library(treescanr)
ex <- function(f) system.file("extdata", f, package = "treescanr")

Where to store things

treescanr is designed to run routinely (e.g., every day), so it helps to keep inputs and outputs in a known, persistent directory. We recommend the user data directory, but any path works:

dir <- tools::R_user_dir("treescanr", "data")
dir.create(dir, recursive = TRUE, showWarnings = FALSE)

A suggested layout:

<dir>/
  reference/                 # tree files and not-evaluated nodes
    Tree_File_2027.csv
    Tree_File_2026_wide_format.txt
    Do_not_evaluate_nodes.csv
  2026-06-30/                # one folder per analysis date...
    lag1/                    # ...and per lag
      counts.txt  parameters.prm  results.csv  results.html  treescan.log

The code in this vignette that writes to disk is not run when the vignette is built. Chunks that are run use the toy data bundled with the package.

1. Visit-level data

ts_visits() accepts any data with one row per visit and tells it which columns hold the patient identifier, the visit date, the diagnosis codes, and whether the patient was admitted:

visits <- read.csv(ex("toy_visits.csv")) |>
  ts_visits(key = "key", date = "date", codes = "diagnosis_codes",
            admitted = "severity")
head(visits)
#>       key       date     diagnosis_codes severity
#>    <char>     <Date>              <char>   <char>
#> 1: P00016 2025-03-01        R10822 A0221        A
#> 2: P00027 2025-03-01              R62.59        V
#> 3: P01136 2025-03-01 R62.51 A08.11 R19.0        V
#> 4: P01366 2025-03-01   R1931 A0811 R1905        V
#> 5: P01420 2025-03-01               R19.2        V
#> 6: P01585 2025-03-01   R6883 A0811 R1902        V

For NSSP ESSENCE DataDetails extracts (e.g., the daily files downloaded by treescan_project/code/1_download_data.R), use ts_visits_nssp():

files <- list.files(file.path(dir, "raw_data"), pattern = "\\.csv$", full.names = TRUE)
visits <- data.table::rbindlist(lapply(files, data.table::fread)) |>
  ts_visits_nssp()

The data should cover the study period (90 days by default) plus one year of lookback, which is used to find incident diagnoses.

2. Count file

ts_counts() keeps incident diagnoses and aggregates them into daily counts per node. 0- nodes are ED visits and 1- nodes are admissions. The end_date is usually the analysis date minus a reporting lag:

counts <- visits |>
  ts_counts(end_date = "2026-06-30", tree_wide = ex("toy_tree_wide.txt"), seed = 1)
counts[order(-n)] |> head()
#>       code       date     n
#>     <char>     <char> <int>
#> 1: 0-A08.4 2026/06/24     6
#> 2: 0-R11.2 2026/06/24     6
#> 3: 0-A08.4 2026/06/29     5
#> 4: 0-R11.2 2026/06/29     5
#> 5: 0-A08.4 2026/06/28     5
#> 6: 0-R11.2 2026/06/28     5

3. Parameter file

The package bundles a parameter template for a tree-temporal scan (TreeScan 2.4.1). Change any setting by name:

prm <- ts_prm_template() |>
  ts_prm_set("monte-carlo-replications" = 999)
ts_prm_get(prm, "monte-carlo-replications")
#> [1] "999"

ts_run() fills in the file paths, the date ranges, and the number of processes (2 by default).

4. Run TreeScan

options(treescanr.binary = "~/TreeScan/treescan64")
ref <- file.path(dir, "reference")

res <- counts |>
  ts_run(
    tree = file.path(ref, "Tree_File_2027.csv"),
    not_evaluated = file.path(ref, "Do_not_evaluate_nodes.csv"),
    prm = prm,
    dir = file.path(dir, "2026-06-30", "lag1")
  )
res

A routine run with several lags

run_date <- Sys.Date()
ref <- file.path(dir, "reference")

results <- lapply(c(1, 4), function(lag) {
  visits |>
    ts_counts(
      end_date = run_date - lag,
      tree_wide = file.path(ref, "Tree_File_2026_wide_format.txt")
    ) |>
    ts_run(
      tree = file.path(ref, "Tree_File_2027.csv"),
      not_evaluated = file.path(ref, "Do_not_evaluate_nodes.csv"),
      dir = file.path(dir, run_date, paste0("lag", lag))
    )
})

Results from earlier runs can be reloaded at any time:

ts_results(file.path(dir, "2026-06-30", "lag1"))