library(OpenSpecy)This example combines a few files bundled with OpenSpecy into a small reference library. Real library builds usually need larger lookup tables and more curation, but the same helper functions apply.
c_spec() and build_lib() default to the
widest range represented by the source spectra at resolution 6. Values
outside each source’s original range are kept as NA, so
useful spectral regions are not discarded. This compact vignette example
uses the shared overlapping range to keep the rendered output small.
mini_files <- c(
read_extdata("raman_hdpe.csv"),
read_extdata("ftir_ldpe_soil.asp"),
read_extdata("raman_atacamit.spc")
)
mini_sources <- lapply(mini_files, read_any)
mini_sources <- lapply(mini_sources, function(x) {
x$metadata$intensity_units <- "absorbance"
attr(x, "intensity_unit") <- "absorbance"
x
})
mini_raw <- c_spec(mini_sources, range = "common", res = 6)
check_OpenSpecy(mini_raw)
#> [1] TRUE
dim(mini_raw$spectra)
#> [1] 67 3
mini_raw$metadata[, "file_name", with = FALSE]
#> file_name
#> <char>
#> 1: raman_hdpe.csv
#> 2: ftir_ldpe_soil.asp
#> 3: raman_atacamit.spcLookup templates help users see which metadata values need curation.
When path is not supplied, the template is returned as a
data.table; users can write it to CSV by supplying
path. Lookup joins use exact values, so edit the template
until the key column matches the object metadata.
template <- make_lib_lookup_template(
mini_raw,
columns = "file_name",
add = c("material", "material_type")
)
template
#> file_name material material_type
#> <char> <char> <char>
#> 1: raman_hdpe.csv <NA> <NA>
#> 2: ftir_ldpe_soil.asp <NA> <NA>
#> 3: raman_atacamit.spc <NA> <NA>For this small example, create the lookup table directly in R. The same shape could come from a CSV edited outside R.
lookup <- data.table::data.table(
file_name = basename(mini_files),
library_type = c("example", "example", "example"),
spectrum_type = c("raman", "ftir", "raman"),
material = c("hdpe", "ldpe in soil", "atacamite")
)
hierarchy <- data.table::data.table(
material = c("hdpe", "ldpe in soil", "atacamite"),
material_class = c("polyethylene", "polyethylene", "copper mineral"),
material_type = c("plastic", "plastic", "mineral")
)
join_lib_metadata(mini_raw, lookup, by = "file_name",
require_complete = TRUE)$metadata[
, c("file_name", "library_type", "spectrum_type", "material"),
with = FALSE
]
#> file_name library_type spectrum_type material
#> <char> <char> <char> <char>
#> 1: raman_hdpe.csv example raman hdpe
#> 2: ftir_ldpe_soil.asp example ftir ldpe in soil
#> 3: raman_atacamit.spc example raman atacamitebuild_lib() can run the ordinary lookup and material
hierarchy joins before applying named recipes. Recipe names become names
in the returned list. Empty recipes keep the merged spectra unchanged;
other recipe lists are passed to process_spec(). Missing
values and processing attributes are handled automatically. Metadata
column names are also cleaned to lowercase underscore names. Known
aliases are coalesced using an editable lookup table. Variants that
differ only by underscores or one terminal plural s match
automatically.
Before sources are merged, build_lib() also converts
declared reflectance and transmittance spectra to absorbance by default.
A nonempty attr(x, "intensity_unit") is the primary truth
for the whole object; otherwise, metadata$intensity_units
is evaluated spectrum by spectrum. Unknown or missing units are left
unchanged with a warning. Use convert_intensity = FALSE
when unit handling has already been completed outside the builder.
The normal input is one or more file paths; one
OpenSpecy or a list of OpenSpecy objects
supports sources already loaded in memory. An RDS path may store either
one object or a list of objects. Other formats are read with
read_any(). Large same-axis source lists are prepared in
bulk, which is much faster for legacy RDS files that store many
one-spectrum objects. Named stages and elapsed time are reported while
the library is built; use progress = FALSE when quiet
output is preferred. Supplying restrict_range_args triggers
the existing restrict_range() operation before
deduplication and recipes; multiple retained ranges can exclude a known
silent region without custom workflow code.
name_lookup <- lib_metadata_name_lookup(
project_code = c("campaign id", "study code"),
regex = list(instrument_mode = "^method_[0-9]+$")
)
name_lookup[
canonical_name %in% c("material_color", "number_of_accumulations")
]
#> canonical_name source_name regex
#> <char> <char> <char>
#> 1: material_color material_color <NA>
#> 2: material_color color <NA>
#> 3: material_color colour <NA>
#> 4: number_of_accumulations number_of_accumulations <NA>
#> 5: number_of_accumulations number_of_sample_scans <NA>
#> 6: number_of_accumulations coadded_scans <NA>
lib_clean_name(c("User Name", "Laser (%)", "Method...3"))
#> [1] "user_name" "laser_perc" "method_3"Named arguments add exact aliases to the defaults, while
regex adds patterns evaluated against cleaned names.
Overlapping regex patterns produce an error that identifies the source
column and matching rules. Pass the result as
metadata_name_lookup. Set
clean_metadata_values = TRUE in build_lib() or
clean_values = TRUE in lib_clean_metadata() to
lowercase, trim, and ASCII-normalize character metadata values before
joins. Ordinary and hierarchical joins run whenever their corresponding
lookup input is non-NULL. spectrum_identity
receives one additional deterministic cleanup: recognizable paths are
reduced to their basename and trailing file extensions supported by
read_any() are removed. Numeric OPUS extensions include
.10 and any other terminal period followed only by digits.
Exact lookup keys receive the same cleanup. Each built library records
changed values and counts in its
spectrum_identity_cleanup_report attribute. Keep flexible
class patterns in a separate table and apply them afterward with
predict_class_reference(). Automatic ordinary lookups use
the single shared column with overlapping values and unique lookup keys.
Lookups with no usable shared key are skipped with a message; lookups
with multiple usable shared keys are treated as ambiguous, so wrap a
lookup as list(lookup = table, by = "key") when an explicit
key is needed. Named by vectors map metadata names to
different lookup names. Canonical source metadata is standardized before
these external joins. Every spectrum receives library_name
from a populated organization, otherwise from
user_name; this is the canonical source-library grouping
used by build assessments. A blank organization is still
filled from the reviewed internal user_name alias for the
existing type lookup. The older lookup-level fallback_by
field is deprecated. Set fill_only = TRUE to fill blank
values such as library_type or spectrum_type
without replacing populated source metadata.
mini_libs <- build_lib(
mini_files,
recipes = list(
raw = list(),
derivative = list(
conform_spec = FALSE,
smooth_intens = TRUE,
smooth_intens_args = list(window = 15, derivative = 1),
make_rel = TRUE
),
nobaseline = list(
conform_spec = FALSE,
smooth_intens = FALSE,
subtr_baseline = TRUE,
make_rel = TRUE
)
),
metadata_lookups = lookup,
material_hierarchy = hierarchy,
clean_metadata_values = TRUE,
convert_intensity = FALSE,
assess = TRUE,
dedupe = FALSE
)
#> build_lib [0.0s]: starting
#> build_lib [0.0s]: reading path 1/3 (raman_hdpe.csv)
#> build_lib [0.3s]: reading path 2/3 (ftir_ldpe_soil.asp)
#> build_lib [0.6s]: reading path 3/3 (raman_atacamit.spc)
#> build_lib [1.0s]: streamed 3 source object(s) from 3 path(s)
#> build_lib [1.0s]: merging 3 source object(s) with c_spec()
#> build_lib [1.3s]: joining metadata lookup 1/1
#> build_lib [1.3s]: metadata lookup 1/1 complete (matched=3; unmatched=0)
#> build_lib [1.3s]: joining the material hierarchy
#> build_lib [1.3s]: material hierarchy complete (unmatched=0)
#> build_lib [1.3s]: processing recipe 1/3 (raw)
#> build_lib [1.3s]: calculating signal-to-noise (raw)
#> build_lib [1.3s]: assessing spectra (raw)
#> build_lib [1.4s]: processing recipe 2/3 (derivative)
#> build_lib [1.4s]: calculating signal-to-noise (derivative)
#> build_lib [1.4s]: assessing spectra (derivative)
#> build_lib [1.4s]: processing recipe 3/3 (nobaseline)
#> build_lib [1.4s]: calculating signal-to-noise (nobaseline)
#> build_lib [1.4s]: assessing spectra (nobaseline)
#> build_lib [1.5s]: complete
names(mini_libs)
#> [1] "raw" "derivative" "nobaseline"
check_OpenSpecy(mini_libs$raw)
#> [1] TRUE
check_OpenSpecy(mini_libs$derivative)
#> [1] TRUE
attr(mini_libs$derivative, "derivative_order")
#> [1] "1"
attr(mini_libs$nobaseline, "baseline")
#> [1] "nobaseline"
mini_libs$raw$metadata[
, .(file_name, material, material_class, material_type, sn,
assessment_flag, assessment_checks)
]
#> file_name material material_class material_type sn
#> <char> <char> <char> <char> <num>
#> 1: raman_hdpe.csv hdpe polyethylene plastic 5.542373
#> 2: ftir_ldpe_soil.asp ldpe in soil polyethylene plastic 43.439904
#> 3: raman_atacamit.spc atacamite copper mineral mineral 5.220850
#> assessment_flag assessment_checks
#> <lgcl> <char>
#> 1: TRUE missing_values
#> 2: TRUE missing_values
#> 3: TRUE missing_valuesprune_lib() performs the spectrum-supported cleanup used
before medoid or model creation. With cross_class = TRUE,
it first flags correlations strictly above
cross_class_threshold between different reviewed classes
within each source library. The spectrum with the most active
wrong-class neighbors is removed first and scores are recalculated;
adjacent maximum-score ties remove both endpoints. A second pass then
uses the number of independent opposing libraries, removing the
higher-evidence endpoint or both on a tie. Generic and unclassified
labels are excluded. The audit records phase, view, component,
identities/classes/libraries, active degree or independent evidence,
decision round, threshold, and reason. Within the same
spectral-technique pool, other may then be reassigned to
the nearest established class, other plastic only to a
candidate whose material_type is plastic, and
other material only to organic matter or
mineral. The matched material_type is copied
with the class. It then evaluates material classes from largest to
smallest using bounded correlation blocks. FTIR and NIR share a
candidate pool; Raman is separate; 2200–2420 cm-1 is excluded
by default. After generic reassignment, each spectrum_type
and material-class group must contain at least min_n
spectra across the complete input database; min_n is not
applied separately to each source library. Smaller resolved groups are
reassigned as a whole when an eligible destination exists and removed
only when no destination exists; groups exactly at the threshold are
retained whole and larger groups are never reduced below it. Constant or
unmatched spectra in eligible classes are retained, and deterministic
IDs resolve correlation ties. Use return = "report" for the
retained IDs, frozen class schedule, reassignment correlations,
spectrum-level removals, and excluded_classes, which
reports observed support and the number of additional or reassigned
spectra needed. End-to-end builds combine these class rows in
assessments$cleanup$summary with the recipe name.
build_lib(prune = ...) maps recipe names to
prune_lib() argument lists, so each selected spectral
representation is pruned independently. Resolved class/type groups below
min_n are reassigned as a whole to the most-correlated
established class in the same technique pool and material type; they are
dropped only when no eligible correlated destination exists. The pruning
assessment records the destination, mean class correlation, action, and
reason. In the official workflow, an otherwise unresolved identity is
temporarily labeled other. The default
remove_other = TRUE removes blank identities and the
unresolved literal other class before quality control.
Reviewed broad other plastic and
other material categories stay in the reference libraries
and remain eligible for constrained reassignment during
derivative/no-baseline pruning. The builder keeps reviewed identifiers,
source metadata, prior labels, reasons, actions, and typed before/after
counts in assessments$cleanup$summary. The separate
assessments$cleanup$dropped_spectrum_identities table
contains only the sorted distinct source identities removed anywhere in
cleanup. Set remove_other = FALSE to retain generic rows
for the constrained semisupervised prune_lib() reassignment
described above.
pruned <- prune_lib(
mini_libs$derivative,
min_n = 1,
return = "report",
progress = FALSE
)
pruned$summary
#> before after reassigned classes_excluded cross_class_removed
#> <int> <int> <int> <int> <int>
#> 1: 3 3 0 0 0
#> threshold_removed removed
#> <int> <int>
#> 1: 0 0
pruned$schedule
#> pool material_class initial_n is_protected schedule_order
#> <char> <char> <int> <lgcl> <int>
#> 1: ftir_nir polyethylene 1 FALSE 1
#> 2: raman copper mineral 1 FALSE 1
#> 3: raman polyethylene 1 FALSE 2
pruned$excluded_classes
#> Empty data.table (0 rows and 11 cols): spectrum_type,pool,material_class,observed_n,minimum_spectra,shortfall...The composable default leaves the new first pass off. Enable it
explicitly for a library carrying canonical library_name
metadata; the official derivative and no-baseline recipes do this by
default:
prune_lib(
reference_library,
cross_class = TRUE,
cross_class_threshold = 0.9
)The version-controlled workflows/OpenSpecy_reference_library.R
script retains the source and output path finders and passes those paths
to one build_lib() call. The complete visual flow,
including checkpoint reuse and old/new assessments, is maintained in .specify/memory/build-lib-diagram.html.
Canonically named exact-class, regex-class, library-type,
material-hierarchy, known-bad-ID, material-form-regex, and common-use
tables live under workflows/data/. Raw-data corrections are
completed externally before this workflow runs. The source-library paths
and output directory are explicit inputs. When
workflow_data is not supplied, build_lib()
looks for the helper tables under data/ beside the calling
script, then under data/ or workflows/data/ in
the current working directory. Calling build_lib() with no
x now stops with an actionable error instead of guessing
external paths. The official build derives library_name
from organization first and user name second, fills a missing
organization from user_name before one source-type lookup,
and asserts complete library/spectrum-type coverage. Per-recipe
retention rows in assessments$cleanup$summary report counts
at preparation, exclusion, deduplication, broad-category review, quality
control, pruning, post-transform, and final partition stages. A
completely dropped source names the first empty stage and its reason.
The exact class lookup runs first;
predict_class_reference() then evaluates the separate regex
table only for blank materials. Exact/regex overlaps are audited without
overwriting exact values, and conflicting regex predictions stop for
review. material_class describes chemistry rather than
physical form. Every reviewed standard plastic class begins with
poly for discoverability (other plastic is the
explicit catch-all exception); polyethylene and polypropylene remain
separate, and polyhydroxy(meth)acrylates retains its
chemically meaningful optional group notation. Poly(styrene-butadiene)
rubber and poly(ethylene-propylene-diene) rubber have distinct classes.
Paint binder labels such as acrylic, alkyd, urethane, nitrocellulose,
and vinyl combinations remain in spectrum_identity,
paint remains in material_form, and the
hierarchy maps each reviewed binder to its relevant chemical polymer
family. Remaining uncertain identities follow the selected generic-row
review policy.
The official workflow also standardizes two enrichment fields. It
normalizes and concatenates all atomic metadata values once per
spectrum, then evaluates the reviewable
material_form_regex.csv. A uniquely matching existing
material_form takes precedence; otherwise the full metadata
row is searched. Cross-category matches remain NA and are
retained in the clash audit. The controlled vocabulary currently covers
paint, rubber, hard plastic, fiber, pellet, film plastic, fragment,
foam, and sphere/bead. common_use_reference.csv is a
material-class lookup for the proposed predominant global ultimate
end-market. Quantitative rows retain consumer, industrial, and
unresolved mass shares: consumer and
industrial require more than 50%, while mixed
requires at least 25% on each side, at least 75% classified, and neither
side above 50%. Where comparable mass shares do not exist, a qualitative
proposal is allowed only with a cited application source, retrieval
date, and explicit review note; share cells remain blank. Consumer
endpoints include products used directly by individuals, such as
passenger tires, household goods, clothing, and patient-facing
healthcare products. Classes with substantial consumer and industrial
endpoints use mixed. Ambiguous catch-alls and poorly
supported specialty groups remain NA. Commodity shares use
the OECD
2019 polymer-by-application mass dataset; PTFE uses a separate
material-specific end-use source recorded in the CSV. Coverage,
provenance, matched evidence, and clashes are retained with the upstream
build assessments. Derivative and nobaseline spectra pass three quality
gates before pruning: FTIR CO2 is flattened when the 2200–2420 maximum
divided by the 2420–2550 silent-region maximum is greater than two;
high-tail detection ignores NA padding, trims each spectrum on its
finite support, and drops failed corrections; and finite running SNR
below two (or unavailable SNR) is removed. Same-library majority closure
then precedes independent-library evidence, generic reassignment, and
conventional top-match pruning. After rounding and type partitioning,
the full and model-range views are closed again. Raw remains unpruned.
Every spectrum excluded by cross-class closure is preserved in the
review quarantine. Each completed library, medoid, model, and assessment
component is written with a SHA-256 manifest under
output_dir/checkpoints. reuse = TRUE loads it
only when the source/lookup signatures, arguments, package version, and
builder contract still match. Validated output is promoted into a
versioned release directory, which contains the seven legacy RDS names,
explicit algorithm-named model RDS files, one aggregate build RDS, and a
release manifest. Existing versioned payloads are immutable: a different
hash stops promotion. The release also contains
quarantined_spectra.rds, a versioned bundle of valid
OpenSpecy objects by recipe/type, row-level removal
metadata, the long conflict audit, and build provenance. Its review copy
and CSV mirrors are checkpointed under output_dir/review
before downstream model and comparison stages.
The returned object has four top-level elements:
libraries, medoids, models, and
assessments. Libraries are recipe-first and then keyed by
ftir, raman, or nir: full Raman
uses 200–4000, FTIR uses 400–4000, and NIR uses 4000–12000
cm-1. FTIR/Raman medoids and models use 800–3200; the NIR
identification interval is derived from the longest region with at least
90% finite coverage (falling back to available coverage) and recorded on
the object. All-blank metadata columns are dropped independently by
type. The enrichment columns are retained even when every value for one
type is missing; other all-blank metadata columns are dropped
independently by type. Each spectrum must observe at least 10% of its
type-specific identification axis to enter medoid selection or model
fitting. PAM operates on a temporary relative matrix where each
spectrum’s missing values are filled by that spectrum’s finite mean.
Selected medoid IDs are then pulled from the original range-restricted
object, so published medoids retain their original NA
positions. Groups with more than 3,000 spectra use five deterministic
1,000-spectrum PAM samples; each candidate set is scored against the
complete group with the same correlation distance, avoiding a full
oversized dissimilarity matrix. models is algorithm-first
(logistic_regression or random_forest), then
recipe and type. Logistic regression retains derivative/nobaseline
medoid training and fills restored gaps with the finite mean at each
wavenumber across training medoids. Random forests use the complete
eligible raw, derivative, and nobaseline type libraries, relative
normalization, wavenumber-mean filling, inverse-frequency balanced case
sampling, 500 probability trees, permutation importance, and all
available ranger threads. Balanced sampling improves minority
representation in each bootstrap without also applying a class-vote
correction. Because this resampling can make out-of-bag accuracy
optimistic, treat OOB results as fit diagnostics. The release assessment
separately applies each deployed model to its complete corresponding
source dataset. train_spec_model() exposes both engines;
build_model_lib() remains a compatible wrapper. Each model
stores its filler, support audit, training diagnostics, and one tidy
tests table. Logistic lambda selection remains maximum
out-of-fold macro class accuracy with its calibrated alpha,
no-intercept, grouped multinomial, and weighting settings unchanged.
Full assessment uses every candidate and downloaded legacy artifact.
Candidate and legacy artifacts independently use a seeded approximately
ten-percent class/type-stratified sample. Rows sharing a physical
identifier or exact spectral content are one group and cannot cross
training and test. This source-local policy tolerates taxonomy evolution
without fuzzy cross-version class matching; its results describe each
artifact on its own source population and are not paired causal deltas.
Held-out groups are removed from full reference libraries before their
identification assessment. Medoid artifacts are instead tested as
deployed: each medoid library identifies every spectrum in its complete
corresponding processed library (for example,
medoid_derivative.rds searches
derivative.rds). Existing logistic and random-forest
artifacts likewise identify their complete corresponding source
libraries once; assessment does not select new medoids or call
train_spec_model(). match_spec() uses stored
fillers for partial spectra. Macro class accuracy is the primary metric,
followed by evaluated coverage, overall accuracy, and clear old/new
assess_spec(report = "all") shifts. The review surface
contains at most ten tables nested under cleanup,
ref_lib, medoid, model, and
functionality. Old and new values occupy adjacent columns;
accuracy, misidentification counts, absolute diagnostic correlations,
and warning/error rate shifts are sorted from largest to smallest. Pass
rows are omitted from quality shifts. Machine-level split and row-test
evidence is available through
attr(reference_library_build$assessments, "evidence").
Published accuracy tables contain aggregate overall and macro metrics,
not per-class accuracy rows. A versioned release writes all global
cleanup, quality, pruning, model, and comparison evidence to
assessments.rds; its library, medoid, and model files
contain only runtime data and prediction state. The accompanying
reference_library_build.rds is a lightweight index that can
be supplied to rebuild_lib_artifacts(). The smaller
1,000-spectrum run in the manual benchmark is development-only and is
not part of build_lib() or final acceptance evidence.
The standard call therefore has this shape:
reference_library_build <- build_lib(
files,
output_dir = output_dir,
previous_library_dir = "system",
remove_other = TRUE
)After upstream libraries have already completed, rebuild only the downstream artifacts into a new explicit output location:
updated_reference_library_build <- rebuild_lib_artifacts(
completed_build_dir,
output_dir = downstream_output_dir,
previous_library_dir = "system"
)completed_build_dir may instead be a
reference_library_build.rds, a libraries.rds
checkpoint, or the corresponding in-memory build object.
The script is excluded from package builds but remains available in the GitHub repository so library releases can be reviewed and reproduced.