Walks a directory recursively, reads every file matching pattern
with cr_read_export() and row-binds the result. Design facts that
live in the directory layout and in the file names — rather than
inside the files — are recovered by a parser and joined back onto the
cells.
Usage
cr_read_exports(
root,
pattern = "\\.(csv|tsv|xls|xlsx)$",
column_map = NULL,
spec = NULL,
parser = NULL,
recursive = TRUE,
drop_empty_rows = TRUE,
col_types = NULL,
progress = TRUE,
call = rlang::caller_env()
)Arguments
- root
Directory to walk.
- pattern
Regular expression selecting export files. Default matches
.csv,.tsv,.xlsand.xlsx.- column_map
Optional
cr_column_map()applied to every export.- spec
Optional
cr_path_spec()describing the directory and file-name grammar. Ignored whenparseris supplied.- parser
Optional function of the file paths returning a tibble with one row per path.
- recursive
Logical. Walk sub-directories. Default
TRUE.- drop_empty_rows
Logical. Passed to
cr_read_export().- col_types
Optional column-type specification passed to
cr_read_export().- progress
Logical. Show a progress bar while reading. Default
TRUE.- call
The execution environment of the calling function. Used for error reporting; experts only.
Value
A tibble of all cells, with source_file and source_path
provenance columns and any columns produced by the parser. The paths
that were read are attached as the "files" attribute.
Details
No naming convention is hardcoded. Either supply a cr_path_spec()
via spec, which is handed to cr_parse_paths(), or supply an
arbitrary parser function. A parser receives the character vector of
file paths and must return one row per path; if it returns a
source_path column the join is made on that column, otherwise row
order is assumed to match.
Row-binding is tolerant of exports that carry different subsets of the
available columns: absent columns are filled with NA.
See also
cr_read_export(), cr_path_spec(), cr_parse_paths(),
cr_dataset().
Other import:
cr_assign_units(),
cr_centroid_overlap(),
cr_column_map(),
cr_extract_markers(),
cr_filename_grammar(),
cr_marker_rules(),
cr_merge_rules(),
cr_parse_paths(),
cr_path_spec(),
cr_read_cellprofiler(),
cr_read_cells(),
cr_read_design(),
cr_read_export(),
cr_read_qupath(),
cr_read_segmantr(),
cr_unit_map()
Examples
# Build a small two-file export tree.
root <- file.path(tempdir(), "cr_exports_demo")
unlink(root, recursive = TRUE)
leaf <- file.path(root, "Run1", "CompoundA", "Plate_1")
dir.create(leaf, recursive = TRUE, showWarnings = FALSE)
one <- function(n) {
data.frame("Event Label" = seq_len(n),
"Signal - Mean Intensity" = seq_len(n) * 10,
check.names = FALSE)
}
utils::write.csv(one(3), file.path(leaf, "CompoundA_vehicle_1.csv"),
row.names = FALSE)
utils::write.csv(one(4), file.path(leaf, "CompoundA_5min_10uM_treated_1.csv"),
row.names = FALSE)
spec <- cr_path_spec(
levels = c("run", "compound", "plate"),
grammar = cr_filename_grammar(
tokens = list(interval = "[0-9]+min", dose = "[0-9]+uM"),
defaults = list(interval = "none", dose = "vehicle"),
prefix_strip = "CompoundA"
)
)
cells <- cr_read_exports(root, spec = spec, progress = FALSE)
cells
#> # A tibble: 7 × 19
#> source_file source_path `Event Label` Signal - Mean Intens…¹ run compound
#> <chr> <chr> <dbl> <dbl> <chr> <chr>
#> 1 CompoundA_5mi… /tmp/Rtmpb… 1 10 Run1 Compoun…
#> 2 CompoundA_5mi… /tmp/Rtmpb… 2 20 Run1 Compoun…
#> 3 CompoundA_5mi… /tmp/Rtmpb… 3 30 Run1 Compoun…
#> 4 CompoundA_5mi… /tmp/Rtmpb… 4 40 Run1 Compoun…
#> 5 CompoundA_veh… /tmp/Rtmpb… 1 10 Run1 Compoun…
#> 6 CompoundA_veh… /tmp/Rtmpb… 2 20 Run1 Compoun…
#> 7 CompoundA_veh… /tmp/Rtmpb… 3 30 Run1 Compoun…
#> # ℹ abbreviated name: ¹`Signal - Mean Intensity`
#> # ℹ 13 more variables: plate <chr>, merge_unit <lgl>, partial_plate <lgl>,
#> # omitted_reagent <lgl>, reacquisition <lgl>, lot <lgl>, variant <chr>,
#> # core <chr>, replicate <chr>, interval <chr>, dose <chr>, parse_ok <lgl>,
#> # parse_error <chr>