| Type: | Package |
| Title: | Methods for Handling Missing Values in Linear Modeling |
| Version: | 0.1.0 |
| Maintainer: | Billy Ouattara <ouattarabilly33@hotmail.com> |
| Description: | Provides user-friendly methods for handling missing data in regression modeling, including available-case linear regression, multiple imputation, random-forest imputation, and the Tower method. Implemented approaches include chained-equation imputation described by van Buuren and Groothuis-Oudshoorn (2011) <doi:10.18637/jss.v045.i03>, multiple imputation described by Honaker, King and Blackwell (2011) <doi:10.18637/jss.v045.i07>, random-forest imputation described by Stekhoven and Buehlmann (2012) <doi:10.1093/bioinformatics/btr597>, and the Tower method described by Matloff and Mohanty (2023) https://CRAN.R-project.org/package=toweranNA. |
| Depends: | R (≥ 4.0.0) |
| Imports: | regtools, mice, Amelia, missForest, toweranNA, stats, utils, qeML |
| License: | MIT + file LICENSE |
| URL: | https://github.com/matloff/mvPred |
| BugReports: | https://github.com/matloff/mvPred/issues |
| Encoding: | UTF-8 |
| LazyData: | true |
| RoxygenNote: | 7.3.3 |
| NeedsCompilation: | no |
| Packaged: | 2026-09-20 20:17:20 UTC; ouatt |
| Author: | Norman Matloff [aut], Cynthia Mascarenhas [aut], Billy Ouattara [aut, cre], Alekya Veluri [aut] |
| Repository: | CRAN |
| Date/Publication: | 2026-09-30 09:10:09 UTC |
Datasets Included in mvPred
Description
Datasets distributed with mvPred for examples, experiments, and package demonstrations.
Usage
data("auto-mpg")
data(english)
data(NHkids)
Format
auto_mpg: a data frame based on the Auto MPG data, including fuel-efficiency and vehicle-characteristic variables.
english: a data frame used for predictive modeling examples involving vocabulary outcomes and demographic predictors.
NHkids: a data frame used for predictive modeling examples involving child weight, height, and age in months.
Details
These datasets are included to support examples and methodological comparisons within the package.
Examples
data("auto-mpg", package = "mvPred")
data("english", package = "mvPred")
data("NHkids", package = "mvPred")
Cross-Validation for Predictive Modeling with Missing Data
Description
Run k-fold cross-validation for regression while applying one of the package's missing-data handling strategies. Depending on the selected method, the function can fit complete-case, available-case, PREFILL, or TOWER-based models and return fold-level summaries.
Usage
bootstrap(
data,
yName,
k = 5,
task = c("regression", "classification"),
method = c("AC", "CC", "PREFILL", "TOWER"),
threshold = 0.5,
seed = 42,
test_split = 0.2,
impute_method = c("mice", "amelia", "missforest", "complete"),
mice_method = NULL,
m = 5,
use_dummies = FALSE,
tower_regFtnName = "lm",
tower_opts = list(),
tower_scaling = NULL,
tower_yesYVal = NULL,
...
)
Arguments
data |
A data frame containing predictors and the response variable. |
yName |
The name of the response column in |
k |
Number of folds used for cross-validation. |
task |
Whether to run |
method |
Missing-data handling strategy: |
threshold |
Classification threshold used when |
seed |
Random seed used for split reproducibility. |
test_split |
Proportion of the full dataset reserved for the final holdout test set. |
impute_method |
Imputation method used when |
mice_method |
Optional mice method specification for PREFILL runs. |
m |
Number of multiple imputations used by applicable PREFILL methods. |
use_dummies |
Logical indicating whether categorical variables should be converted to dummies for Amelia-based PREFILL runs. |
tower_regFtnName |
Regression routine passed to the TOWER method. |
tower_opts |
Optional list of control settings passed to TOWER. |
tower_scaling |
Optional scaling object passed to TOWER. |
tower_yesYVal |
Optional flag passed to TOWER. |
... |
Additional arguments forwarded to the selected fitting routine. |
Details
The function first optionally creates a holdout split controlled by test_split. Cross-validation is then run on the training portion using the requested missing-data handling method. For PREFILL runs, the supported imputation methods are mice, Amelia, missForest, and complete-case preprocessing.
Value
A list containing fold-level performance summaries and fitted model objects. Depending on the selected task, the returned metrics correspond to regression or classification performance.
Common returned components include:
-
fold_models: list of fitted models, one per cross-validation fold. -
final_model: fitted model from the final training split whentest_split > 0. -
test_results: performance on the final holdout test set whentest_split > 0. -
training_missingness_by_fold: missingness tables for each training fold. -
training_missingness_average: average missingness across training folds.
Examples
data("auto-mpg", package = "mvPred")
df <- auto_mpg
df$car_name <- NULL
for (nm in names(df)) {
suppressWarnings(df[[nm]] <- as.numeric(df[[nm]]))
}
res <- bootstrap(
data = df,
yName = "mpg",
k = 3,
task = "regression",
method = "CC"
)
res$RMSE_mean
Extract Complete Cases from a Data Frame
Description
Remove rows containing missing values and report which rows were retained or discarded.
Usage
compCases(data)
Arguments
data |
A data frame or matrix that may contain missing values. |
Value
A list with components:
-
intactData: rows ofdatawith no missing values. -
intactRows: logical vector indicating complete rows. -
nonintactRows: integer indices of rows containing at least one missing value.
Examples
df <- data.frame(x = c(1, NA, 3), y = c(2, 4, 6))
compCases(df)
Linear Regression Using Available Cases Methods
Description
This suite of functions implements linear regression using the Available Cases (pairwise deletion) approach.
It makes use of the avaliable cases in the dataset to fit the model and make predictions rather than deleting
any rows with missing values. The main function lm_ac fits the model even when the predictors or response contain missing values.
The helper functions compute matrix products using only complete pairs of observations. S3 methods are included for summary and prediction.
Usage
lm_ac(data, yName, holdout = NULL, ...)
ac_mul(A)
ac_vec(X, y)
## S3 method for class 'lm_ac'
summary(object, ...)
## S3 method for class 'lm_ac'
predict(object, newdata, ...)
Arguments
data |
A data.frame containing predictors and the response variable. |
yName |
The name of the response column in |
holdout |
Optional holdout specification used to define testing rows. This may be a logical vector of length |
... |
Additional arguments passed to |
A |
A numeric matrix or data frame to be used in |
X |
A numeric matrix of predictors for |
y |
A numeric vector representing the response variable for |
object |
An object of class |
newdata |
A data frame with predictor variables for prediction. |
Details
The function lm_ac fits a linear regression model using pairwise deletion for missing values.
Instead of removing entire rows with missing values (as in complete-case analysis), this approach computes each component
of the normal equations (X'X and X'y) using the average of all available pairs.
ac_mul returns a symmetric matrix of average products between all pairs of columns.
ac_vec returns a vector of average products between each column of X and y.
summary.lm_ac prints the estimated coefficients.
predict.lm_ac generates predictions using the fitted model on new data.
Value
lm_ac returns an object of class lm_ac, a list containing:
data |
Cleaned input data used in modeling. |
yName |
Response variable name. |
formula |
The model formula. |
fit_obj |
A list with components |
ac_mul returns a numeric matrix.
ac_vec returns a numeric vector.
summary.lm_ac returns the coefficient vector (invisible).
predict.lm_ac returns a numeric vector of predictions.
Examples
data(airquality)
# Fit model
mod <- lm_ac(airquality, "Ozone")
# Print coefficients
summary(mod)
# Predict on new data
predict(mod, newdata = airquality)
Compare Missing-Data Methods with Bootstrap Summaries
Description
Run bootstrap() across one or more missing-data handling methods and collect the resulting cross-validation summary metrics in a single data frame.
Usage
lm_compare(
data,
predictors = NULL,
yName,
k = 5,
task = c("regression", "classification"),
methods = c("CC", "AC", "TOWER", "PREFILL"),
seed = 42,
impute_method = c("mice", "amelia", "missforest", "complete"),
mice_method = NULL,
m = 5,
use_dummies = FALSE,
tower_regFtnName = "lm",
tower_opts = list(),
tower_scaling = NULL,
tower_yesYVal = NULL,
...
)
Arguments
data |
A data frame containing predictors and the response variable. |
predictors |
Optional character vector of predictor names to keep. If |
yName |
The name of the response column in |
k |
Number of folds used for cross-validation. |
task |
Whether to summarize regression or classification performance. |
methods |
Character vector of methods to evaluate. Supported values are |
seed |
Random seed used for reproducibility. |
impute_method |
Imputation method used when |
mice_method |
Optional mice method specification passed to PREFILL runs. |
m |
Number of multiple imputations used by applicable PREFILL methods. |
use_dummies |
Logical indicating whether categorical variables should be converted to dummies for Amelia-based PREFILL runs. |
tower_regFtnName |
Regression routine passed to TOWER. |
tower_opts |
Optional list of TOWER options. |
tower_scaling |
Optional scaling object passed to TOWER. |
tower_yesYVal |
Optional flag passed to TOWER. |
... |
Additional arguments forwarded to |
Details
lm_compare() is intended for method comparison under cross-validation. Internally it calls bootstrap() with test_split = 0, so the reported summaries correspond to cross-validation results rather than final holdout-test performance.
Value
A data frame containing one row per requested method and the corresponding average performance metrics.
Examples
data("auto-mpg", package = "mvPred")
df <- auto_mpg
df$car_name <- NULL
for (nm in names(df)) {
suppressWarnings(df[[nm]] <- as.numeric(df[[nm]]))
}
lm_compare(df, yName = "mpg", k = 3, task = "regression", methods = c("CC", "AC"))
Linear Model Fitting with Pre-Imputation of Missing Data
Description
Fit linear models to data with missing values by first imputing missing data using multiple imputation methods,
then fitting linear models to the imputed datasets. Supports multiple imputation methods including mice,
Amelia, and missForest, as well as complete-case analysis.
Usage
lm_prefill(
data,
yName,
impute_method = "mice",
m = 5,
use_dummies = FALSE,
holdout = NULL,
...
)
## S3 method for class 'lm_prefill'
summary(object, ...)
## S3 method for class 'lm_prefill'
predict(object, newdata, type = "response", ...)
Arguments
data |
A data frame containing the variables for modeling, possibly with missing values. |
yName |
A character string specifying the name of the response variable in |
impute_method |
Character string specifying the imputation method to use. One of
|
m |
Integer specifying the number of multiple imputations to generate. This is
only used for |
use_dummies |
Logical indicating whether to convert categorical predictors to
dummy variables before imputation. This is only relevant for |
holdout |
Optional holdout specification used to define testing rows. This may be
a logical vector of length |
... |
Additional arguments passed to the imputation function
( |
object |
An object of class |
newdata |
A data frame containing new data for prediction. Rows with missing values in |
type |
Type of prediction; passed to |
Details
lm_prefill is a wrapper function that first imputes missing values in the
dataset using one of several supported methods, then fits linear models to the
imputed datasets.
Supported imputation methods:
-
"mice": Multiple Imputation by Chained Equations via the mice package. -
"amelia": Multiple imputation using the Amelia package. -
"missforest": Nonparametric imputation using random forests via the missForest package. -
"complete": Complete-case analysis by omitting rows with missing values.
The use_dummies argument controls whether categorical variables are converted
to dummy variables before imputation. If use_dummies = TRUE, the package
regtools is required for dummy variable creation.
Additional arguments passed via ... are forwarded to the underlying imputation
or modeling functions as appropriate.
The returned object is of class "lm_prefill" and contains the imputed data, fitted models, and metadata.
summary.lm_prefill returns model summaries for each imputed dataset or a
single summary for complete-case or missForest methods.
predict.lm_prefill returns predictions averaged across imputations for
multiple imputation methods, or direct predictions for single imputation or
complete-case models.
If newdata contains missing values, those rows are omitted with a warning before prediction.
Value
An object of class "lm_prefill" containing at least the following components:
data |
Original input data. |
yName |
Response variable name. |
formula |
Model formula used for fitting. |
imputed_data |
Imputed dataset(s) or imputation object. |
impute_method |
Imputation method used. |
fit_obj |
Fitted linear model(s). |
missforest_OOBerror |
If |
Author(s)
Cynthia Mascarenhas cynmascarenhas@ucdavis.edu
See Also
mice, amelia,
missForest, lm,
predict.lm, na.omit,
factorsToDummies
Examples
library(mice)
library(Amelia)
library(missForest)
# Example dataset with missing values
data(airquality)
airquality$Ozone[1:10] <- NA
# Fit linear model with mice imputation
lm_obj <- lm_prefill(
airquality,
yName = "Ozone",
impute_method = "mice",
m = 5
)
# Summarize fitted models
summary(lm_obj)
# Predict on new data (complete cases only)
newdata <- airquality[11:20, ]
predict(lm_obj, newdata)
# Fit with Amelia and use dummy variables for categorical predictors
lm_obj2 <- lm_prefill(
airquality,
yName = "Ozone",
impute_method = "amelia",
m = 5,
use_dummies = TRUE
)
summary(lm_obj2)
Fit Linear Models with Missing Data Using the Tower Method
Description
lm_tower fits a linear regression model using the Tower Method implemented in the toweranNA package.
This method handles missing data in both training and new datasets without explicit imputation by leveraging regression averaging.
Usage
lm_tower(data, yName, regFtnName = "lm", opts = list(), scaling = NULL, yesYVal = NULL)
## S3 method for class 'lm_tower'
summary(object, ...)
## S3 method for class 'lm_tower'
predict(object, newdata, ...)
Arguments
data |
A data frame containing the training data with possible missing values. |
yName |
A character string specifying the name of the response variable in |
regFtnName |
A character string specifying the regression function to use (default is |
opts |
A list of options passed to |
scaling |
Optional scaling parameters passed to |
yesYVal |
Optional parameter passed to |
object |
An object of class |
newdata |
A data frame containing new observations for prediction. May contain missing values. |
... |
Additional arguments (currently unused). |
Details
The Tower Method implemented in the toweranNA package fits regression models that can handle missing data without explicit imputation. It uses regression averaging over subsets of complete cases to provide predictions even when new data contain missing values.
This function provides a convenient wrapper around toweranNA::makeTower() and associated prediction methods.
lm_tower is particularly useful when the primary goal is prediction rather than inference.
Value
lm_tower returns an object of class "lm_tower" containing the fitted Tower model and related information.
summary.lm_tower prints a brief summary of the fitted model, including the response variable, regression function, and number of complete cases used.
predict.lm_tower returns a numeric vector of predicted values for newdata. The Tower method handles missing values internally, so predictions are returned for all rows of newdata.
Author(s)
Cynthia Mascarenhas cynmascarenhas@ucdavis.edu
See Also
Examples
library(toweranNA)
data(airquality)
airquality$Ozone[1:10] <- NA
airquality$Solar.R[5:15] <- NA
# Fit the tower model
lm_tower_obj <- lm_tower(airquality, yName = "Ozone")
# Summarize the model
summary(lm_tower_obj)
# Prepare new data with missing values
newdata <- airquality[20:30, -which(names(airquality) == "Ozone")]
newdata[1, "Solar.R"] <- NA
# Predict on new data
preds <- predict(lm_tower_obj, newdata = newdata)
print(preds)
Interface qeML Models with Missing-Value Preprocessing
Description
Apply a missing-value handling strategy before fitting a predictive model from the qeML package.
The helper addListElement() appends or replaces an element in a list.
Usage
qeMLna(
data,
yName,
qeMLftn,
mvFtn,
qeMLopts = NULL,
mvPredOpts = NULL,
retainMVFtnOut = TRUE,
seed = 9999,
holdout = 1000
)
addListElement(l, newEltName, newElt)
Arguments
data |
A data frame containing predictors and the response variable. |
yName |
The name of the response column in |
qeMLftn |
Character string naming the predictive function from qeML to call. |
mvFtn |
Character string naming the missing-value handling routine to apply first. |
qeMLopts |
Optional list of arguments passed to the selected qeML function. |
mvPredOpts |
Optional list of arguments passed to the missing-value handling routine. |
retainMVFtnOut |
Logical indicating whether to store the missing-value routine output in the returned object. |
seed |
Random seed passed through to applicable routines. |
holdout |
Holdout size used by the predictive routine if applicable. |
l |
A list to modify. If |
newEltName |
Name of the element to add or replace. |
newElt |
Value of the element to add or replace. |
Value
qeMLna() returns an object of class "MVqeMLout" containing the qeML fit and, optionally, the missing-value handling output.
addListElement() returns the modified list.
Examples
data("english", package = "mvPred")
eng <- english[, c("age", "birth_order", "ethnicity", "sex", "mom_ed", "vocab")]
qeMLna(eng, "vocab", "qeLin", "compCases", retainMVFtnOut = FALSE)
addListElement(list(a = 1), "b", 2)