| Type: | Package |
| Title: | Reproducible Analysis of Freshwater Microplastic Data |
| Version: | 0.1.0 |
| Author: | Chanikya Naidu [aut, cre] |
| Maintainer: | Chanikya Naidu <thefisherieschanikyaneeti@gmail.com> |
| Description: | Provides validation, harmonization, descriptive analysis, compositional analysis, transparent risk components, grouped cross-validation, visualization, and predictive modelling tools for freshwater microplastic datasets. The package includes synthetic demonstration data conforming to the LIMPID-India data model and emphasizes explicit units, provenance, percentage closure, non-imputation of missing environmental covariates, and leakage-aware model evaluation. |
| License: | MIT + file LICENSE |
| Encoding: | UTF-8 |
| Depends: | R (≥ 4.1.0) |
| Imports: | ggplot2 |
| Suggests: | knitr, rmarkdown, readxl, testthat (≥ 3.0.0) |
| Config/testthat/edition: | 3 |
| VignetteBuilder: | knitr |
| RoxygenNote: | 7.3.2 |
| NeedsCompilation: | no |
| Packaged: | 2026-09-22 08:50:13 UTC; chani |
| Repository: | CRAN |
| Date/Publication: | 2026-09-30 11:50:02 UTC |
Reproducible Analysis of Freshwater Microplastic Data
Description
Tools for loading, validating, harmonizing, analysing, modelling and visualizing freshwater microplastic datasets, with bundled synthetic demonstration tables that follow the LIMPID-India schema.
Details
The package emphasizes explicit units, provenance, composition closure, grouped validation, and transparent assumptions. Missing environmental values are not silently imputed.
Author(s)
Chanikya Naidu <thefisherieschanikyaneeti@gmail.com>
Transparent risk components and reproducibility metadata
Description
Calculate assumption-explicit risk components and report reproducibility/citation metadata.
Usage
calculate_risk(data, abundance_reference, hazard_scores = NULL,
weights = NULL, component_max = NULL, thresholds = NULL)
limpid_session_info()
limpid_citation(doi = NULL, creators = "Chanikya Naidu")
Arguments
data |
A |
abundance_reference |
Positive reference abundance in the same unit as MP_Mean. |
hazard_scores |
Optional named polymer hazard scores supplied by the researcher. |
weights |
Optional named composite-score weights. |
component_max |
Explicit positive scaling maxima required when weights are used. |
thresholds |
Optional ascending numeric composite-score cut points. |
doi |
Assigned LIMPID-India dataset DOI. |
creators |
Creator string for citation output. |
Details
No universal polymer hazard scores, composite weights or risk classes are hard-coded. Assumptions must be supplied explicitly.
Value
A risk-component data frame, sessionInfo object, or citation string.
Load, clean and validate freshwater microplastic data
Description
Functions for reading LIMPID-India tables, unit harmonization and structural validation.
Usage
limpid_tables()
load_limpid(path = NULL, tables = limpid_tables(), validate = TRUE, quiet = FALSE)
check_database(data, strict = FALSE, tolerance = 0.2)
validate_mp_data(data, type = c("abundance", "morphology", "size", "polymer",
"events", "lake"), tolerance = 0.2)
clean_mp_data(data, abundance_col = "MP_Mean", unit_col = "Unit",
target_unit = "particles L^-1", invalid_action = c("error", "flag", "na"))
convert_mp_units(x, from, to = "particles L^-1")
Arguments
path |
NULL, a compatible CSV directory, or an Excel workbook. Excel input requires readxl. |
tables |
Character vector of table names. |
validate |
Logical; validate after loading. |
quiet |
Logical; suppress load message. |
data |
A data frame or |
strict |
Logical; stop when a failing database check is found. |
tolerance |
Allowed percentage deviation from 100 for composition closure. |
type |
Expected table schema. |
abundance_col |
Name of abundance column. |
unit_col |
Name of unit column. |
target_unit |
Target abundance unit. |
invalid_action |
Action for invalid abundance values. |
x |
Numeric abundance vector. |
from |
Source abundance unit. |
to |
Target abundance unit. |
Value
A loaded database, validation table, cleaned object, or converted numeric vector.
Examples
db <- load_limpid()
check_database(db)
convert_mp_units(1000, "particles m^-3", "particles L^-1")
Model and validate freshwater microplastic abundance
Description
Lognormal or Gamma regression with original-scale prediction and grouped cross-validation.
Usage
model_mp_abundance(data, formula = NULL,
method = c("lognormal_lm", "gamma_glm"), na_action = c("omit", "fail"))
predict_mp(model, newdata, interval = c("none", "confidence", "prediction"), level = 0.95)
cross_validate_mp(data, formula = NULL,
method = c("lognormal_lm", "gamma_glm"), group = "Lake_ID")
Arguments
data |
A |
formula |
Model formula with untransformed abundance response. |
method |
Model family. |
na_action |
Missing-data handling for fitting. |
model |
A |
newdata |
Data frame for prediction. |
interval |
Prediction interval type. |
level |
Interval confidence level. |
group |
Grouping column for held-out folds. |
Details
The default lognormal model uses a log1p response and Duan smearing for back-transformation. Grouped cross-validation defaults to Lake_ID to reduce spatial leakage.
Value
A limpid_model, prediction data frame, or cross-validation result list.
Examples
db <- load_limpid()
m <- model_mp_abundance(db, MP_Mean ~ Season_Global + Lake_Type)
cross_validate_mp(db, MP_Mean ~ Season_Global + Lake_Type)
Visualize freshwater microplastic patterns
Description
Seasonal, composition, depth and coordinate-based point visualizations.
Usage
plot_lake_map(data, value_col = "MP_Mean")
classify_hotspots(data, value_col = "MP_Mean", method = c("quantile", "robust_z"))
map_mp_hotspots(data, value_col = "MP_Mean", method = c("quantile", "robust_z"))
plot_seasonality(data, lake = NULL, value_col = "MP_Mean")
plot_morphology_profile(data, id_col = "Lake_Name")
plot_polymer_profile(data, id_col = "Lake_Name")
plot_depth_profile(data, depth_col = "Depth_m", value_col = "MP_Mean", group_col = NULL)
Arguments
data |
A |
value_col |
Numeric indicator column. |
method |
Relative hotspot classification method. |
lake |
Optional lake name or identifier filter. |
id_col |
Identity/group column for profile bars. |
depth_col |
Depth column in metres. |
group_col |
Optional depth-profile grouping column. |
Details
Hotspot functions provide relative point classification and do not imply spatial interpolation. Depth plotting requires genuinely depth-resolved data and will reject the bundled event-level database.
Value
A ggplot2 object, except classify_hotspots() which returns a data frame.
Descriptive and compositional microplastic analysis
Description
Summaries, event-level joins, composition closure, CLR transformation and Aitchison distance.
Usage
summarise_abundance(data, by = c("Lake_Name", "Season_Global"),
value_col = "MP_Mean", conf_level = 0.95)
analyse_morphology(data, group_by = c("Lake_Name", "Season_Global"))
analyse_size_distribution(data, group_by = c("Lake_Name", "Season_Global"))
analyse_polymers(data, group_by = c("Lake_Name", "Season_Global"))
composition_closure(data, cols, tolerance = 0.2,
action = c("flag", "renormalize", "error"))
clr_transform(data, cols, pseudocount = 1e-06)
aitchison_distance(data, cols, pseudocount = 1e-06)
build_model_data(data, include_environment = TRUE, primary_eligible_only = FALSE)
provenance_summary(data)
Arguments
data |
A |
by |
Grouping columns for abundance summaries. |
value_col |
Numeric abundance or indicator column. |
conf_level |
Confidence level for mean confidence intervals. |
group_by |
Optional composition grouping columns. |
cols |
Composition component columns. |
tolerance |
Allowed deviation from 100 percent. |
action |
Closure action. |
pseudocount |
Positive zero replacement on the proportion scale. |
include_environment |
Join environmental context into model data. |
primary_eligible_only |
Restrict to explicitly primary-model-eligible rows. |
Value
A data frame, distance object, or named list, depending on function.
Examples
db <- load_limpid()
summarise_abundance(db)
analyse_polymers(db)