pls() estimates Partial Least Squares Structural Equation Models (PLS-SEM)
and their consistent (PLSc) variants. The function accepts lavaan-style
syntax, handles ordered indicators through polychoric correlations and probit
factor scores, and estimates two-level (multilevel) models, with separate
models for the within (level: 1) and between (level: 2) levels.
pls(
syntax,
data,
standardize = TRUE,
consistent = NULL,
bootstrap = FALSE,
ordered = NULL,
missing = c("listwise", "mean", "kNN"),
knn.k = 5,
mcpls = NULL,
probit = NULL,
tolerance = 1e-05,
max.iter.0_5 = 500L,
boot.ncores = 1L,
boot.ncpus = NULL,
boot.parallel = c("no", "multicore", "multisession", "snow"),
boot.R = 500L,
boot.iseed = NULL,
sample = NULL,
mc.min.iter = 50L,
mc.max.iter = 1000L,
mc.reps = 20000L,
mc.fixed.seed = FALSE,
mc.polyak.juditsky = TRUE,
mc.pj.extrapolate = TRUE,
mc.tol = if (mc.polyak.juditsky) 1e-04 else 0.001,
mc.delta.se = TRUE,
mc.delta.jacobian.k = max(floor(boot.R/100L), 1),
mc.fn.args = list(),
mc.rescov = c("auto", "reduced", "full"),
mc.diag.secant = FALSE,
mc.small.sample = FALSE,
mc.small.sample.max.k = 100L,
mc.small.sample.point.estimate = c("median", "mean"),
verbose = interactive(),
boot.optimize = TRUE,
boot.drop.inadmissible = FALSE,
mc.boot.control = list(min.iter = mc.min.iter, max.iter = mc.max.iter, mc.reps =
floor(0.5 * mc.reps), tol = mc.tol, polyak.juditsky = mc.polyak.juditsky,
pj.extrapolate = FALSE, verbose = FALSE, fixed.seed = TRUE, reuse.p.start = TRUE),
reliabilities = NULL,
default.path.estimator = c("ols", "gls", "lmer"),
cluster = NULL,
inner.weights = c("path", "centroid", "factorial"),
approach.weights = c("pls", "pca"),
mlm = NULL,
level2.cov = c("muml", "means"),
level2.approach.weights = c("pca", "pls"),
...
)Character string with lavaan-style model syntax describing
both measurement (=~) and structural (~) relations. Two-level
models are specified with level: 1 and level: 2 blocks (see
mlm), where random slopes are specified with rv() modifiers at
level 1 (e.g., fw ~ rv(s1)*x1). The random slopes (e.g., s1)
are variables at level 2.
A data.frame or coercible object containing the manifest
indicators referenced in syntax. Ordered factors are automatically
detected, but can also be supplied explicitly through ordered.
Logical; if TRUE, indicators are standardized before
estimation so that factor scores have comparable scales.
Logical; TRUE requests PLSc corrections, whereas FALSE
fits the traditional PLS model. If NULL (default), FALSE is
used for MC-PLS (including two-level models) and for models using lmer
as the path estimator, and TRUE otherwise.
Logical; if TRUE, nonparametric bootstrap standard errors
are computed with boot.R resamples.
Optional character vector naming manifest indicators that should be treated as ordered when computing polychoric correlations.
Character string specifying how to handle missing indicator data.
"listwise" removes rows with missing values (listwise deletion).
"mean" imputes missing indicator values using simple univariate
imputation: the mean for continuous variables, the median for ordered variables
with more than two categories, and the mode for binary ordered variables (two
categories) or nominal variables.
"kNN" (or "knn") imputes missing indicator values using
k-nearest neighbors imputation (kNN). When missing = "kNN", rows with
all indicators missing are removed prior to imputation. Rows with missing
cluster values are always removed.
Integer specifying the number of neighbors (k) used when
missing = "kNN".
Should the model be estimated using the Monte-Carlo Consistent Partial Least Squares (MC-PLSc) algorithm?
Logical; overrides the automatic choice of probit factor scores that is based on whether ordered indicators are present.
Numeric; Convergence criteria/tolerance.
Maximum number of PLS iterations performed when estimating the measurement and structural models.
Integer: number of workers to be used for parallel bootstrapping.
Parallel bootstrapping is enabled when boot.ncores > 1.
Deprecated alias for boot.ncores.
The type of parallel operation to be used (if any). The
default is "no". "multisession" runs the bootstrap using multiple
background R sessions (works on all platforms), while "multicore"
uses forked processes (not available on Windows). "snow" is kept for
backwards compatibility and is treated as an alias for "multisession".
Internally this is implemented using the future package
Integer giving the number of bootstrap resamples drawn when
bootstrap = TRUE.
An integer to set the bootstrap seed. Or NULL if no
reproducible results are needed. This works for both serial and parallel
settings. When boot.iseed is not NULL, .Random.seed (if it
exists) in the global environment is left untouched.
DEPRECATED. Integer giving the number of bootstrap resamples drawn when
bootstrap = TRUE.
Minimum number of iterations in MC-PLS algorithm.
Maximum number of iterations in MC-PLS algorithm.
Monte-Carlo sample size in MC-PLS algorithm.
Should a fixed seed be used in the MC-PLS algorithm? Setting a fixed seed will likely yield less accurate estimates, but can substantially improve the stability and computational efficiency of the algorithm.
Should the polyak.juditsky running average method be applied in the MC-PLS algorithm?
Logical; if TRUE (the default), the Polyak-Juditsky
convergence point is estimated via NLS exponential extrapolation (with
Aitken \(\delta^2\) as a fallback). If FALSE a warm start is performed
instead, and the plain Polyak-Juditsky average is used.
Only relevant when mc.polyak.juditsky = TRUE.
Tolerance in MC-PLS algorithm.
Should delta-method standard errors be computed for MC-PLS estimates?
Integer number of Monte-Carlo Jacobians to average when computing delta-method standard errors. Defaults to one per 100 bootstrap resamples, with a minimum of 1.
Additional arguments to MC-PLS algorithm, mainly for controlling the step size.
How residual covariances are treated in MC-PLS. One of
"auto" (the default), "reduced", or "full". In
"reduced" mode residual covariances are not treated as free
parameters; they are identified from the rest of the structural model and
the disturbances are simulated independently. In "full" mode the
endogenous residual covariances (eta ~~ eta and xi ~~ eta)
are free parameters, explicitly simulated. "auto" uses "full"
when the structural model is estimated by GLS (i.e. the model contains
residual covariances) and "reduced" otherwise.
Logical; if TRUE, the MC-PLS root-finding algorithm
uses a per-coordinate diagonal-secant step (estimating each coordinate's local
slope from consecutive iterates), instead of the standard Robbins-Monro step.
This might be usefull if the binding funtion is non-monotone.
Logical; if TRUE, average MC-PLS
estimating equations over simulated samples of the observed sample size.
If FALSE (default), fit one sample using all simulated observations.
This assumes that the auxiliary estimator (i.e., traditional PLS) has
negligible finite sample bias. I.e., that the bias of the estimator is
not affected by the sample size.
Maximum number of simulated samples to average
when mc.small.sample = TRUE. Defaults to 100. The number of samples
is also limited by mc.reps, rounded down to a multiple of the
observed sample size, with at least one sample.
Which point estimate of the simulated
auxiliary parameters the root equation matches to the observed ones, when
mc.small.sample = TRUE? "mean" solves
\(E[\theta^{*}|\theta] = \hat{\theta}^{*}\). "median" (the default)
solves \(median[\theta^{*}|\theta] = \hat{\theta}^{*}\).
The median commutes with the (monotone) binding function where the mean
does not, so median-matching targets a median-unbiased estimator. This
removes the finite-sample bias. which can be introduced by the curvature
size models, and when the indicators are uninformative (small loadings,
few categories, strongly assymetric thresholds). It makes the estimating
function somewhat noisier for a given number of simulated samples.
Ignored when mc.small.sample = FALSE.
Should verbose output be printed?
Logical; if TRUE and bootstrap = TRUE, applies
the settings in mc.boot.control inside each bootstrap replicate (MC-PLS only).
In general it will lead to slightly larger and less accurate standard errors.
Logical; if TRUE and bootstrap = TRUE,
bootstrap replicates that yield an inadmissible solution (e.g. a Heywood case)
are discarded before computing standard errors, just like replicates where the
estimation procedure fails outright. Defaults to FALSE, which keeps
inadmissible replicates. In either case the number of inadmissible replicates
is reported. Note that dropping inadmissible solutions conditions the bootstrap
distribution on well-behaved resamples and may bias the standard errors downward.
List of control parameters passed to the MC-PLS algorithm
inside each bootstrap replicate when boot.optimize = TRUE.
Optional named numeric vector of user-supplied reliabilities used for the PLSc consistency correction.
Character string selecting the estimator used for
the structural (path) model when the model does not require Generalized Least
Squares (GLS), or Linear Mixed-Effects Regression (LMER).
The default "ols" uses Ordinary Least Squares whenever
possible, falling back to GLS automatically when the model contains residual
covariances. Setting default.path.estimator = "gls"
forces GLS estimation of the structural model even when OLS would otherwise be
used.
Optional character vector naming the cluster variable(s) in
data. The cluster variables are stored alongside the (standardized)
indicators, and bootstrapping resamples whole clusters instead of rows.
Required for two-level models (a single cluster variable).
Character string selecting the inner weighting scheme
used with approach.weights = "pls". One of: "centroid",
"factorial", or "path". Defaults to "path".
Character string selecting the approach used to
estimate the outer weights. "pls" (default) uses the PLS algorithm,
with the inner weighting scheme given by inner.weights. "pca"
forms the weights from each construct's own indicators only, ignoring the
structural.
Should the model be estimated as a two-level (multilevel) model?
If NULL (default), this is detected from syntax, i.e., whether
it has level: 1 and level: 2 blocks. Two-level models are
estimated using an extension of the MC-PLSc estimator.
The mc.* arguments (e.g., mc.max.iter, mc.reps) are used when they are
specified, whereas arguments like standardize
are (currently) not supported. Standard errors (bootstrap = TRUE)
are computed using the delta method.
Two-level models only. How the (co-)variances of the
variables at level 2 are computed. "means" uses the
covariances of the cluster means. "muml" (default) uses Muthen's (1994)
estimator of the between-cluster covariance matrix, which corrects for the
within-cluster variation in the cluster means.
Two-level models only. The approach used to
estimate the outer weights of the level 2 model (approach.weights
applies to the level 1 model). Defaults to "pca".
Internal arguments. For advanced users only.
A PlsModel object containing the estimated parameters, fit measures,
factor scores, and any bootstrap results (a PlsMultilevelModel object
for two-level models). Methods such as summary(), coef(), and
parameter_estimates() can be applied to inspect the fit.
# \donttest{
library(plssem)
library(modsem)
tpb <- '
ATT =~ att1 + att2 + att3 + att4 + att5
SN =~ sn1 + sn2
PBC =~ pbc1 + pbc2 + pbc3
INT =~ int1 + int2 + int3
BEH =~ b1 + b2
INT ~ ATT + SN + PBC
BEH ~ INT + PBC
'
fit <- pls(tpb, TPB, bootstrap = TRUE)
summary(fit)
#> plssem (0.2.0) ended normally after 3 iterations
#> Estimator PLSc
#> Link LINEAR
#>
#> Number of observations 2000
#> Number of iterations 3
#> Number of latent variables 5
#> Number of observed variables 15
#>
#> Fit Measures:
#> Chi-Square 106.316
#> Degrees of Freedom 82
#> SRMR 0.008
#> RMSEA 0.012
#>
#> R-squared [indicators]:
#> att1 0.847
#> att2 0.825
#> att3 0.805
#> att4 0.745
#> att5 0.845
#> sn1 0.817
#> sn2 0.863
#> pbc1 0.856
#> pbc2 0.859
#> pbc3 0.787
#> int1 0.816
#> int2 0.827
#> int3 0.742
#> b1 0.762
#> b2 0.821
#>
#> R-squared [latents]:
#> INT 0.367
#> BEH 0.210
#>
#> Latent Variables:
#> Estimate Std.Error z.value P(>|z|)
#> ATT =~
#> att1 0.921 0.013 69.630 0.000
#> att2 0.908 0.015 62.202 0.000
#> att3 0.897 0.016 55.832 0.000
#> att4 0.863 0.018 47.108 0.000
#> att5 0.919 0.015 62.303 0.000
#> SN =~
#> sn1 0.904 0.012 75.378 0.000
#> sn2 0.929 0.012 76.951 0.000
#> PBC =~
#> pbc1 0.925 0.011 87.046 0.000
#> pbc2 0.927 0.012 80.415 0.000
#> pbc3 0.887 0.012 73.299 0.000
#> INT =~
#> int1 0.903 0.011 82.355 0.000
#> int2 0.909 0.011 82.309 0.000
#> int3 0.861 0.013 64.016 0.000
#> BEH =~
#> b1 0.873 0.016 56.246 0.000
#> b2 0.906 0.016 57.702 0.000
#>
#> Regressions:
#> Estimate Std.Error z.value P(>|z|)
#> INT ~
#> ATT 0.243 0.028 8.778 0.000
#> SN 0.201 0.030 6.791 0.000
#> PBC 0.240 0.031 7.628 0.000
#> BEH ~
#> PBC 0.308 0.026 12.033 0.000
#> INT 0.210 0.026 7.951 0.000
#>
#> Covariances:
#> Estimate Std.Error z.value P(>|z|)
#> ATT ~~
#> SN 0.633 0.015 41.846 0.000
#> PBC 0.692 0.012 56.920 0.000
#> SN ~~
#> PBC 0.696 0.013 52.154 0.000
#>
#> Variances:
#> Estimate Std.Error z.value P(>|z|)
#> ATT 1.000
#> SN 1.000
#> PBC 1.000
#> .INT 0.633 0.019 33.063 0.000
#> .BEH 0.790 0.019 40.696 0.000
#> .att1 0.153 0.024 6.284 0.000
#> .att2 0.175 0.027 6.599 0.000
#> .att3 0.195 0.029 6.752 0.000
#> .att4 0.255 0.032 8.073 0.000
#> .att5 0.155 0.027 5.724 0.000
#> .sn1 0.183 0.022 8.423 0.000
#> .sn2 0.137 0.022 6.090 0.000
#> .pbc1 0.144 0.020 7.320 0.000
#> .pbc2 0.141 0.021 6.608 0.000
#> .pbc3 0.213 0.021 9.896 0.000
#> .int1 0.184 0.020 9.287 0.000
#> .int2 0.173 0.020 8.628 0.000
#> .int3 0.258 0.023 11.151 0.000
#> .b1 0.238 0.027 8.763 0.000
#> .b2 0.179 0.029 6.262 0.000
#>
# }