
dataprep: Fast, Efficient, and Versatile Data Preprocessing and Reshaping with C++, OpenMP & SIMD
Source:R/dataprep-package.R
dataprep-package.RdFast, efficient, and versatile preprocessing and reshaping of tabular and time-series data. Most heavy routines are implemented in C++ via 'Rcpp', with optional OpenMP parallelization and SIMD acceleration (AVX2 / AVX-512) on supported hardware. The 0.1.7 release rewrites the cleaning routines in C++ and delivers a 1.1–1146× speedup over 0.1.5. The 'melt()' and 'dcast()' reshaping functions achieve a 0.5–1187× speedup for 'melt()' and a 2.0–639× speedup for 'dcast()' relative to every one of the seven major alternatives in the R and Python ecosystems, at every tested scale (from 1,000 to 100,000,000 rows), and produce output identical to 'reshape2', 'data.table', 'tidyr', 'pandas', 'polars', 'dask', and 'duckdb'. Core preprocessing steps include variable deletion by missing fraction, observation deletion by consecutive missing runs, point-by-point weighted outlier removal via conditional extremum, traditional percentile-based outlier removal, and linear interpolation within short time periods. The package also provides fast reshaping, descriptive statistics, missing-value diagnosis, multiple imputation strategies, winsorization, several outlier detection methods (IQR, MAD, percentile), data transformation and standardization, categorical encoding, duplicate removal, data validation, data quality reporting, and stratified sampling. Feature-engineering helpers cover binning, high-correlation and low-variance filtering, and string cleaning. Time-series tools cover detrending, diurnal-cycle removal, rolling statistics, lag creation, resampling, simple decomposition, day/night and season flags, log returns, drift detection, and panel balancing. Fit/transform-style machine-learning interfaces prevent data leakage during preprocessing. Methods are based on, and improved from: Liang, C.-S., Wu, H., Li, H.-Y., Zhang, Q., Li, Z. & He, K.-B. (2020) doi:10.1016/j.scitotenv.2020.140923 . This work was supported by the National Natural Science Foundation of China (No. 12301674).
Details
The package is organised around a four-step cleaning pipeline:
Variable deletion ([varidele()]) – drop columns whose missing fraction exceeds a threshold.
Observation deletion ([obsedele()]) – drop rows with a consecutive missing run longer than `half` minutes on both sides.
Outlier removal ([condextr()]) – point-by-point weighted conditional extremum detection.
Short-period interpolation ([shorvalu()]) – fill remaining short gaps from nearby valid anchors.
Steps 1-4 are wrapped by [dataprep()] for one-call use. The design reasoning behind the ordering is documented in `vignette("dataprep-philosophy")`.
The package also provides fast wide-to-long and long-to-wide reshaping ([melt()], [dcast()]), a full set of descriptive and diagnostic helpers ([descdata()], [descplot()], [percdata()], [percplot()], [na_diagnose()], [data_report()]), and fit/transform-style preprocessing plans that prevent data leakage ([prep_fit()], [prep_transform()]).
This work was supported by the National Natural Science Foundation of China (No. 12301674).
See also
Core pipeline: [dataprep()], [varidele()], [obsedele()], [condextr()], [shorvalu()].
Reshaping: [melt()], [dcast()].
Leakage-free workflow: [prep_fit()], [prep_transform()].
Diagnostics and reporting: [descdata()], [descplot()], [percdata()], [percplot()], [na_diagnose()], [data_report()].
Author
Maintainer: Chun-Sheng Liang chun-shengliang@qq.com (ORCID)
Authors:
Chun-Sheng Liang chun-shengliang@qq.com (ORCID)
Hao Wu
Hai-Yan Li
Qiang Zhang
Zhanqing Li
Ke-Bin He