
Package index
Complete preprocessing workflow
One-call pipelines and fit/transform interfaces that combine several steps while preventing data leakage.
-
dataprep-package - dataprep: Fast, Efficient, and Versatile Data Preprocessing and Reshaping with C++, OpenMP & SIMD
-
dataprep() - Data preprocessing with multiple steps in one function
-
prep_fit() - Build a preprocessing plan on training data to prevent data leakage
-
prep_transform() - Apply a preprocessing plan to new data
-
dry_run() - Simulate preprocessing and report changes without modifying data
-
data_report() - Generate a simple data quality report
Variable and observation deletion
Remove variables by missing fraction, remove observations by consecutive missing runs, and diagnose the effect beforehand.
-
varidele() - Delete variables containing too many missing values
-
obsedele() - Delete observations with excessive consecutive missing values
-
na_diagnose() - Diagnose missing value patterns in data
-
balance_panel() - Balance panel data
Outlier removal and detection
Point-by-point weighted conditional extremum, traditional percentile removal, mask-based detection, and winsorization.
-
condextr() - Remove outliers using point-by-point weighed outlier removal by conditional extremum
-
percoutl() - Traditional percentile-based outlier removal
-
optisolu() - Find optimal combination of interval and times for condextr
-
detect_outliers() - Detect outliers using multiple methods
-
winsorize() - Winsorize outliers by capping extreme values
-
phys_filter() - Physical limit filtering
-
shorvalu() - Interpolation with values to refer to within short periods
-
impute_missing() - Impute missing values
Variable selection and encoding
Drop redundant variables, encode categorical columns, and discretize continuous variables.
-
filter_high_cor() - Remove highly correlated variables
-
filter_low_var() - Remove low-variance (near-constant) variables
-
encode_categorical() - Encode categorical variables
-
bin_data() - Discretize continuous variables into bins
Transformation and standardization
Log / Box-Cox / Yeo-Johnson transformations and z-score / min-max / robust scaling.
-
transform_data() - Transform and standardize numeric variables
-
log_returns() - Logarithmic returns for financial time series
-
zerona() - Turn zeros to missing values
Time series tools
Detrending, diurnal-cycle removal, rolling statistics, lags, resampling, decomposition, drift detection, and time flags.
-
detrend_ts() - Remove linear trend from time series
-
remove_diurnal_cycle() - Remove diurnal cycle
-
roll_apply() - Apply rolling window statistics
-
create_lags() - Create lagged variables
-
resample_time() - Resample time series to a coarser period
-
decompose_ts() - Simple time series decomposition
-
drift_detect() - Sensor drift detection
-
day_night_flag() - Day/night flag
-
season_flag() - Season flag
-
clean_strings() - Clean and standardize character columns
-
deduplicate() - Remove duplicate observations
-
validate_data() - Validate data against a set of rules
Sampling and summary
Stratified sampling, descriptive statistics, and percentile summaries with matching plots.
-
sample_data() - Random sampling with optional stratification
-
descdata() - Fast descriptive statistics
-
descplot() - View descriptive statistics via plot
-
percdata() - Calculate top and bottom percentiles of selected variables
-
percplot() - Plot top and bottom percentiles of selected variables