A cr_filename_grammar states how one export file name decomposes
into design facts: which trailing token is the replicate index, which
leading labels are stripped, which substrings are tokens, and which
whole-name shapes are legal at all.
Arguments
- tokens
Named list (or named character vector) of regular expressions. Each is searched for in the name core; the matched text becomes the value of a column named after the list element.
- defaults
Named list of values to use when a token is absent. Elements with no entry default to
NA.- core_patterns
Character vector of regular expressions. The core (the name with extension, markers, replicate index and prefix removed) must match at least one of them. Empty means no whitelist.
- typo_fixes
Named character vector of
pattern = replacementrepairs applied before anything else is parsed.- prefix_strip
Character vector of leading labels to strip from the core, longest first.
- replicate
Regular expression for the trailing replicate token. Default
"[0-9]+(?:\\.[0-9]+)?", which accepts both1and1.1.- sep
Token separator. Default
"_".- normalise_space
Logical. Convert runs of whitespace to
sepand collapse repeated separators. DefaultTRUE.- x
A
cr_filename_grammar.- ...
Ignored.
Details
The core_patterns whitelist exists because the absence of a token is
itself meaningful — a name with no interval token is the untreated
reference arm, a name with no concentration token is a vehicle
control. A mistyped token therefore matches nothing, falls through to
the defaults, and silently reclassifies a treated unit as a vehicle
control, pulling it into the control denominator of its own batch.
With a whitelist in place such a name is a hard error instead.
prefix_strip entries are removed longest-first, so a short label
cannot consume a longer one that starts with the same characters.
See also
cr_parse_paths(), cr_path_spec().
Other import:
cr_assign_units(),
cr_centroid_overlap(),
cr_column_map(),
cr_extract_markers(),
cr_marker_rules(),
cr_merge_rules(),
cr_parse_paths(),
cr_path_spec(),
cr_read_cellprofiler(),
cr_read_cells(),
cr_read_design(),
cr_read_export(),
cr_read_exports(),
cr_read_qupath(),
cr_read_segmantr(),
cr_unit_map()
Examples
g <- cr_filename_grammar(
tokens = list(interval = "[0-9]+min", dose = "[0-9]+uM",
mode = "treated|vehicle"),
defaults = list(interval = "none", dose = "vehicle"),
core_patterns = c("^vehicle$",
"^[0-9]+min_vehicle$",
"^[0-9]+uM_treated$",
"^[0-9]+min_[0-9]+uM_treated$"),
prefix_strip = c("CompoundA", "CompoundB")
)
g
#> <cr_filename_grammar>
#> • tokens: interval, dose, and mode
#> • core patterns: 4
#> • prefixes stripped: 2
#> • replicate: "[0-9]+(?:\\.[0-9]+)?"