| Type: | Package |
| Title: | Machine Learning Imputation, Clustering and Survival Analysis for Longitudinal Proteomic Data |
| Version: | 0.1.0 |
| Description: | Imputes missing biomarker measurements in a wide longitudinal serum panel with gradient-boosted decision trees, groups the completed panel by Bayesian consensus clustering, and compares the resulting patient subgroups by Kaplan-Meier, log-rank and Cox analysis. The imputation learner is described in Ke et al. (2017) https://papers.nips.cc/paper/6907-lightgbm-a-highly-efficient-gradient-boosting-decision-tree and the clustering method in Lock and Dunson (2013) <doi:10.1093/bioinformatics/btt425>. Imputed values are conditional-mean predictions, so the procedure is a machine-learning single imputation; the completions carry no between-imputation variance and must not be pooled by Rubin's rules. Two panels from Gene Expression Omnibus accession 'GSE65622' are included, one for each survival endpoint. |
| License: | GPL-3 |
| Encoding: | UTF-8 |
| Depends: | R (≥ 4.1.0) |
| Imports: | lightgbm (≥ 3.3.0), data.table, survival, stats, utils |
| Suggests: | BCClong, testthat (≥ 3.0.0) |
| Config/testthat/edition: | 3 |
| LazyData: | true |
| LazyDataCompression: | xz |
| Config/roxygen2/version: | 8.1.0 |
| NeedsCompilation: | no |
| Packaged: | 2026-09-21 12:50:13 UTC; DELL |
| Author: | Neelesh Kumar [aut, cre], Atanu Bhattacharjee [aut], Gajendra K. Vishwakarma [aut], Tanmoy Majumdar [aut] |
| Maintainer: | Neelesh Kumar <neelesh2302@gmail.com> |
| Repository: | CRAN |
| Date/Publication: | 2026-09-30 10:10:02 UTC |
MIML: machine-learning imputation, clustering and survival for GSE65622
Description
Two functions run the published MIML workflow end to end: miml_rfs() on
RFS_data against recurrence-free survival, and miml_mfs() on MFS_data
against metastasis-free survival. Each performs LightGBM imputation of the
incomplete biomarker panel, Bayesian consensus clustering of the completed
panel, and a survival comparison of the resulting clusters.
Details
Imputed values are conditional-mean predictions, so the completions carry no between-imputation variance and must not be combined by Rubin's rules. This is machine-learning single imputation.
Author(s)
Maintainer: Neelesh Kumar neelesh2302@gmail.com
Authors:
Neelesh Kumar neelesh2302@gmail.com
Atanu Bhattacharjee
Gajendra K. Vishwakarma
Tanmoy Majumdar
References
Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q. and Liu, T.-Y. (2017). LightGBM: a highly efficient gradient boosting decision tree. Advances in Neural Information Processing Systems 30, 3146-3154. https://papers.nips.cc/paper/6907-lightgbm-a-highly-efficient-gradient-boosting-decision-tree
Lock, E. F. and Dunson, D. B. (2013). Bayesian consensus clustering. Bioinformatics 29(20), 2610-2616. doi:10.1093/bioinformatics/btt425
Metastasis-free survival panel (GSE65622)
Description
The same patients, visits and 507 biomarkers as RFS_data, paired instead
with metastasis-free survival. The default input to miml_mfs().
Usage
MFS_data
Format
A data frame of 288 rows and 516 columns, identical in structure to
RFS_data except that the survival column is MFS and event
is the metastasis indicator. Twenty-six of the 80 patients record a
metastasis.
Source
Gene Expression Omnibus accession GSE65622, released April 2016.
See Also
Examples
data(MFS_data)
dim(MFS_data)
MFS_data[1:5, 1:6]
Recurrence-free survival panel (GSE65622)
Description
Longitudinal serum proteomic measurements from patients with locally
advanced rectal cancer undergoing multimodal treatment, paired with
recurrence-free survival. The default input to miml_rfs().
Usage
RFS_data
Format
A data frame of 288 rows and 516 columns: 80 patients observed at
three or four sampling points each. Patient_ID is an integer
identifier repeated across a patient's visits; RFS is
recurrence-free survival in days measured from that visit, so it
decreases across a patient's rows; event is the recurrence
indicator, constant within a patient. Then 507 protein biomarkers named
as in the original submission, then six phenotype columns:
diarrhea_grade, sampling_point, slide_print_batch,
trg_score, ypn_stage, ypt_stage.
Biomarker names are not syntactic R names, so quote them or use
[[ ]]. About 10.3 percent of the biomarker measurements are
missing. Five of the 80 patients record a recurrence.
Source
Gene Expression Omnibus accession GSE65622, released April 2016.
See Also
Examples
data(RFS_data)
dim(RFS_data)
RFS_data[1:5, 1:6]
Metastasis-free survival analysis of the GSE65622 panel
Description
Identical to miml_rfs() but run on MFS_data against metastasis-free
survival. Twenty-six of the eighty patients record a metastasis, against
five recurrences for the other endpoint, so this is the better powered of
the two comparisons.
Usage
miml_mfs(
data = NULL,
m = 5L,
maxit = 10L,
nrounds = 200L,
n_clusters = 2L,
n_bcc_markers = 3L,
predictors = c("manuscript", "all"),
seed = 123L,
verbose = TRUE
)
Arguments
data |
Panel to analyse. Defaults to MFS_data. |
m |
Number of completed datasets. The published analysis used 5. |
maxit |
Refit cycles per completion. The published analysis used 10. |
nrounds |
Maximum boosting rounds. The published analysis used 200. |
n_clusters |
Number of consensus clusters. The published analysis used 2. |
n_bcc_markers |
Biomarkers carried into the clustering model, chosen by descending variance. The published analysis used 3. |
predictors |
Predictor pool for the imputation models.
|
seed |
Random seed. The caller's RNG state is restored on exit. |
verbose |
Print progress. |
Value
An object of class miml_analysis, as for miml_rfs().
Predictor pool and outcome leakage
Under predictors = "manuscript" the survival time and the event
indicator are inside the predictor pool, and no biomarker predicts another.
Any association later estimated between the imputed values and survival is
therefore inflated by construction. The default reproduces the published
numbers; use predictors = "all" for new work.
Run time
At the defaults the imputation fits 507 biomarkers over 10 cycles for each of 5 completions and takes roughly an hour.
See Also
Examples
data(MFS_data)
small <- MFS_data[1:60, c("Patient_ID", "MFS", "event",
names(MFS_data)[4:6])]
miml_mfs(small, m = 1, maxit = 1, nrounds = 3,
n_clusters = 0, verbose = FALSE)
sub <- MFS_data[, c("Patient_ID", "MFS", "event", names(MFS_data)[4:23])]
res <- miml_mfs(sub, m = 1, maxit = 2, nrounds = 20)
res
## the published analysis is miml_mfs() at its defaults.
Recurrence-free survival analysis of the GSE65622 panel
Description
Runs the published MIML workflow on RFS_data: LightGBM imputation of every incomplete biomarker, Bayesian consensus clustering of the completed panel, and a Kaplan-Meier, log-rank and Cox comparison of the resulting clusters against recurrence-free survival.
Usage
miml_rfs(
data = NULL,
m = 5L,
maxit = 10L,
nrounds = 200L,
n_clusters = 2L,
n_bcc_markers = 3L,
predictors = c("manuscript", "all"),
seed = 123L,
verbose = TRUE
)
Arguments
data |
Panel to analyse. Defaults to RFS_data. |
m |
Number of completed datasets. The published analysis used 5. |
maxit |
Refit cycles per completion. The published analysis used 10. |
nrounds |
Maximum boosting rounds. The published analysis used 200. |
n_clusters |
Number of consensus clusters. The published analysis used 2. |
n_bcc_markers |
Biomarkers carried into the clustering model, chosen by descending variance. The published analysis used 3. |
predictors |
Predictor pool for the imputation models.
|
seed |
Random seed. The caller's RNG state is restored on exit. |
verbose |
Print progress. |
Details
Imputed values are the conditional-mean predictions of the fitted LightGBM models. No residual is drawn and no donor is sampled, so the completions carry no between-imputation variance and must not be combined by Rubin's rules. This is machine-learning single imputation.
Value
An object of class miml_analysis: a list with
endpoint, missingness, completed, clusters,
survival and settings.
Predictor pool and outcome leakage
Under predictors = "manuscript" the survival time and the event
indicator are inside the predictor pool, and no biomarker predicts another.
Any association later estimated between the imputed values and survival is
therefore inflated by construction. The default reproduces the published
numbers; use predictors = "all" for new work.
Run time
At the defaults the imputation fits 507 biomarkers over 10 cycles for each of 5 completions and takes roughly an hour.
See Also
Examples
data(RFS_data)
## small and fast: three biomarkers, one cycle, no clustering
small <- RFS_data[1:60, c("Patient_ID", "RFS", "event",
names(RFS_data)[4:6])]
miml_rfs(small, m = 1, maxit = 1, nrounds = 3,
n_clusters = 0, verbose = FALSE)
## twenty biomarkers, two cycles, clustered and compared by survival
sub <- RFS_data[, c("Patient_ID", "RFS", "event", names(RFS_data)[4:23])]
res <- miml_rfs(sub, m = 1, maxit = 2, nrounds = 20)
res
## the published analysis is miml_rfs() at its defaults: 507 biomarkers,
## 10 cycles, 5 completions, roughly an hour.