Skip to contents

Overview

dataprep 0.1.7 ships two reshaping backends, melt() and dcast(), benchmarked here against all seven major alternatives in the R and Python ecosystems:

  • R: reshape2, data.table, tidyr
  • Python: pandas, polars, dask, duckdb

Every cell is measured with microbenchmark using an adaptive times rule: 100 iterations when the first call is under 10 ms, down to a single iteration when it exceeds 10 s. Two statistics are reported:

  • median — the number to quote. Robust to scheduler and GC jitter.
  • mean — the tail. A large mean / median ratio indicates one-off OS work or cache effects.

For every cell the tables also report the speed-up of dataprep relative to each competitor, so the reader can see the full gradient from “about the same” to “three orders of magnitude”.

Test environment

Benchmarks were run on two reference hosts. Only the core configuration is listed here; full hardware details are in README.md.

  • Ubuntu 25.10 (Questing Quokka, kernel 6.17.0-41-generic) — 2× AMD EPYC 9965 192-Core (Turin, Zen 5c), 384 physical / 768 logical cores, L3 768 MiB, 1.0 TiB (16 × 64 GiB Micron, DDR5-5600, Multi-bit ECC), full AVX-512; R 4.5.1, g++ 15.2.0.

  • Windows 11 Pro for Workstations (10.0.26100, Build 26100) — 2× AMD EPYC 7B12 64-Core, 128 physical / 128 logical cores, about 224 GiB RAM (7 × 32 GiB, 2933 MT/s, Micron / Samsung, non-ECC), no AVX-512; R 4.6.1 (ucrt), GCC 14.3.0.

Software versions on both hosts: data.table 1.18.6.1, reshape2 1.4.5, tidyr 1.3.2, reticulate 1.47.0; Python 3.13.7 (Ubuntu) / 3.13.15 (Windows), pandas 3.0.6, polars 1.44.2 (runtime rt64), dask 2026.8.0, duckdb 1.5.5.

How to read these numbers

The two hosts differ in core count, cache size and memory bandwidth. Two properties shape the numbers that follow:

  • The Ubuntu host is unusually large. Most of the 1e6- and 1e7-row cells fit entirely in L3. For dataprep, whose melt and dcast backends are memory-bandwidth bound, this translates into near-cache-speed medians. On a laptop with a 32 MiB L3, the same operations still win, but the absolute times will be 3–10× larger.

  • The Windows host has no AVX-512. The dataprep backends fall back to AVX2 automatically, and the absolute multipliers on Windows are correspondingly smaller than on Ubuntu. The relative ranking of the engines is identical on both hosts.

Both effects favour dataprep in the numbers below. The relative ranking is robust; the absolute multipliers — especially the 1187× and 639× figures — should be interpreted as “best-case on a very large machine”. On a typical 8–16-core workstation the same comparisons are within 10–100×.

Cleaning pipeline (dataprep 0.1.5 → 0.1.7)

The 0.1.7 release rewrites every heavy cleaning routine in C++. The table below compares against 0.1.5 on three dataset sizes from the same source (SMEAR I Varrio forest). All numbers are speed-up ratios (0.1.5 time / 0.1.7 time) on Ubuntu 25.10.

Function 500 rows 7,640 rows 49,422 rows
varidele 1.1× 1.1× 11.6×
obsedele 203× 424× 232×
condextr 196× 217× 1146×
optisolu 188× 77× 109×
dataprep 185× 228× 247×

On Windows 11 Pro for Workstations, the same full-year pipeline gives obsedele ≈ 648×, condextr ≈ 839×, shorvalu ≈ 81×, optisolu ≈ 25× (at cores = 32), and the integrated dataprep call ≈ 173×. varidele is around 1.17× on this cell; this is expected, since varidele is a single colMeans(is.na(.)) in both versions and the new code path has little room for improvement.

Note on optisolu cores. The 0.1.5 implementation could crash when cores > 16, because its parallel::makeCluster() path gave each worker a full copy of the data. The benchmark above used cores = 16 for both versions to keep the comparison fair. 0.1.7 loads the package on each worker, exports the input data once per worker, and runs each (interval, times) case as a separate task, so cores = 64 is safe. The practical speed-up on a many-core host is larger. Also note that the optimal parameter values returned by optisolu() may differ slightly between 0.1.5 and 0.1.7.

melt() — wide to long

Input shapes are described as rows × (n_id + n_val). All numbers in the cells are medians in milliseconds; the value in parentheses is dataprep’s speed-up relative to that competitor.

Vary rows, 1 id + 9 value columns

rows dataprep reshape2 data.table tidyr pandas polars dask duckdb
1e3 0.176 0.373 (2.1×) 0.272 (1.5×) 3.009 (17.1×) 2.162 (12.3×) 0.612 (3.5×) 15.61 (88.6×) 4.714 (26.8×)
1e4 0.235 0.436 (1.9×) 0.341 (1.5×) 3.379 (14.4×) 2.517 (10.7×) 0.685 (2.9×) 15.59 (66.4×) 11.00 (46.9×)
1e5 1.588 1.138 (0.7×) 1.018 (0.6×) 8.230 (5.2×) 6.660 (4.2×) 1.557 (1.0×) 17.99 (11.3×) 68.79 (43.3×)
1e6 3.680 17.89 (4.9×) 7.900 (2.1×) 76.77 (20.9×) 60.30 (16.4×) 10.44 (2.8×) 46.49 (12.6×) 642.1 (174×)
1e7 37.95 372.7 (9.8×) 371.6 (9.8×) 1111 (29.3×) 720.1 (19.0×) 92.53 (2.4×) 486.4 (12.8×) 6423 (169×)
1e8 496.4 3577 (7.2×) 3576 (7.2×) 12089 (24.4×) 7410 (14.9×) 2563 (5.2×) 4590 (9.2×) 65624 (132×)

The sub-1.0× cells are reshape2 (0.7×) and data.table (0.6×) at 1e5 rows with one id column.

Vary rows, 10 id (5 int + 5 chr)

rows dataprep reshape2 data.table tidyr pandas polars dask duckdb
1e3 0.231 0.540 (2.3×) 0.422 (1.8×) 3.123 (13.5×) 5.626 (24.4×) 0.954 (4.1×) 53.07 (230×) 10.64 (46.1×)
1e4 0.475 1.984 (4.2×) 1.760 (3.7×) 4.941 (10.4×) 6.090 (12.8×) 1.889 (4.0×) 54.21 (114×) 50.11 (106×)
1e5 3.752 16.74 (4.5×) 14.90 (4.0×) 23.06 (6.1×) 13.12 (3.5×) 3.944 (1.1×) 57.69 (15.4×) 469.6 (125×)
1e6 21.28 206.2 (9.7×) 153.7 (7.2×) 219.9 (10.3×) 90.42 (4.2×) 37.50 (1.8×) 106.8 (5.0×) 4666 (219×)
1e7 868.6 3168 (3.6×) 2703 (3.1×) 3731 (4.3×) 1357 (1.6×) 562.3 (0.6×) 928.2 (1.1×) 48519 (55.9×)

Vary value columns, 1e3 rows, 1 id

n_val dataprep reshape2 data.table tidyr pandas polars dask duckdb
10 0.174 0.373 (2.1×) 0.267 (1.5×) 2.956 (17.0×) 2.143 (12.3×) 0.608 (3.5×) 16.65 (95.6×) 4.853 (27.9×)
100 0.260 1.081 (4.2×) 0.381 (1.5×) 3.779 (14.5×) 6.458 (24.8×) 0.726 (2.8×) 50.90 (196×) 21.25 (81.7×)
1000 0.973 8.141 (8.4×) 1.350 (1.4×) 12.26 (12.6×) 48.29 (49.6×) 3.535 (3.6×) 370.9 (381×) 174.1 (179×)
10000 3.616 92.45 (25.6×) 9.659 (2.7×) 107.7 (29.8×) 499.0 (138×) 16.31 (4.5×) 4292 (1187×) 1914 (529×)

The 1e3 × 10000 cell is the widest gap in the entire benchmark suite: dataprep returns in 3.62 ms, dask in 4.29 s, and duckdb in 1.91 s.

Vary value columns, 1e3 rows, 10 id

n_val dataprep reshape2 data.table tidyr pandas polars dask duckdb
10 0.247 0.613 (2.5×) 0.459 (1.9×) 3.114 (12.6×) 5.655 (22.9×) 0.915 (3.7×) 53.18 (216×) 12.06 (48.9×)
100 0.529 2.783 (5.3×) 1.964 (3.7×) 5.513 (10.4×) 20.76 (39.2×) 2.051 (3.9×) 186.4 (352×) 59.88 (113×)
1000 4.304 25.15 (5.8×) 16.73 (3.9×) 28.76 (6.7×) 168.6 (39.2×) 6.207 (1.4×) 1614 (375×) 566.6 (132×)
10000 25.22 313.6 (12.4×) 169.0 (6.7×) 281.7 (11.2×) 1793 (71.1×) 59.36 (2.4×) 22456 (891×) 5792 (230×)

melt() on Windows 11 Pro for Workstations

The same cells on the Windows host, with no AVX-512:

rows n_id n_val dataprep reshape2 data.table tidyr pandas polars dask duckdb
1e3 1 9 0.319 0.653 (2.0×) 0.512 (1.6×) 4.224 (13.2×) 3.575 (11.2×) 0.513 (1.6×) 28.55 (89.5×) 7.724 (24.2×)
1e4 1 9 0.533 0.938 (1.8×) 0.780 (1.5×) 5.208 (9.8×) 5.369 (10.1×) 0.795 (1.5×) 29.60 (55.6×) 22.33 (41.9×)
1e5 1 9 3.964 3.485 (0.9×) 3.131 (0.8×) 16.43 (4.1×) 24.99 (6.3×) 2.692 (0.7×) 44.12 (11.1×) 175.9 (44.4×)
1e6 1 9 12.15 24.44 (2.0×) 26.17 (2.2×) 174.4 (14.3×) 210.6 (17.3×) 17.99 (1.5×) 178.8 (14.7×) 1558 (128×)
1e7 1 9 101.4 262.2 (2.6×) 253.3 (2.5×) 1510 (14.9×) 1908 (18.8×) 177.5 (1.8×) 1541 (15.2×) 14588 (144×)
1e8 1 9 1197 3514 (2.9×) 2806 (2.3×) 21106 (17.6×) 22968 (19.2×) 5227 (4.4×) 16904 (14.1×) 160256 (134×)

The largest Windows multiplier is 883× (dask at 1e3 rows and 10000 value columns).

dcast() — long to wide

Input is a canonical long table with every (id, variable) pair present exactly once. All numbers in the cells are medians in milliseconds; the value in parentheses is dataprep’s speed-up relative to that competitor.

Vary n_long, 1 id, 10 levels

n_long dataprep reshape2 data.table tidyr pandas polars dask duckdb
1e3 0.859 1.689 (2.0×) 1.818 (2.1×) 4.349 (5.1×) 1.889 (2.2×) 30.04 (35.0×) 8.351 (9.7×) 7.584 (8.8×)
1e4 0.874 2.649 (3.0×) 2.578 (3.0×) 4.630 (5.3×) 2.352 (2.7×) 38.20 (43.7×) 9.035 (10.3×) 10.51 (12.0×)
1e5 1.014 19.86 (19.6×) 14.74 (14.5×) 7.969 (7.9×) 7.115 (7.0×) 51.95 (51.2×) 15.03 (14.8×) 33.84 (33.4×)
1e6 1.632 151.3 (92.7×) 345.4 (212×) 46.67 (28.6×) 58.85 (36.1×) 105.8 (64.8×) 80.65 (49.4×) 158.7 (97.3×)
1e7 8.417 1888 (224×) 557.6 (66.2×) 740.2 (87.9×) 808.2 (96.0×) 307.8 (36.6×) 968.7 (115×) 1756 (209×)
1e8 100.4 21075 (210×) 17434 (174×) 9711 (96.7×) 10758 (107×) 2429 (24.2×) 12982 (129×) 17061 (170×)

Vary levels, 1 id, 1e6 rows

levels dataprep reshape2 data.table tidyr pandas polars dask duckdb
10 1.632 151.3 (92.7×) 345.4 (212×) 46.67 (28.6×) 58.85 (36.1×) 105.8 (64.8×) 80.65 (49.4×) 158.7 (97.3×)
100 1.451 108.7 (74.9×) 335.9 (232×) 43.25 (29.8×) 56.27 (38.8×) 172.6 (119×) 77.04 (53.1×) 180.9 (125×)
1000 1.966 106.8 (54.3×) 183.9 (93.5×) 44.37 (22.6×) 57.64 (29.3×) 208.4 (106×) 79.51 (40.4×) 200.5 (102×)

Vary n_long, 1 id, 100 levels

n_long dataprep reshape2 data.table tidyr pandas polars dask duckdb
1e4 1.011 2.885 (2.9×) 2.926 (2.9×) 4.851 (4.8×) 2.487 (2.5×) 37.85 (37.4×) 9.310 (9.2×) 15.28 (15.1×)
1e5 1.144 18.12 (15.8×) 8.935 (7.8×) 8.016 (7.0×) 7.015 (6.1×) 55.99 (48.9×) 14.78 (12.9×) 41.93 (36.7×)
1e6 1.393 108.9 (78.2×) 320.7 (230×) 43.53 (31.3×) 56.04 (40.2×) 160.4 (115×) 77.95 (56.0×) 175.7 (126×)
1e7 4.733 2084 (440×) 550.9 (116×) 633.0 (134×) 768.4 (162×) 458.8 (96.9×) 953.6 (202×) 1717 (363×)
1e8 43.91 17115 (390×) 16193 (369×) 8174 (186×) 9880 (225×) 2467 (56.2×) 12217 (278×) 17895 (408×)

The 1e8 × 100 levels cell is the strongest dcast result on this host: dataprep returns in 43.9 ms, reshape2 in 17.1 s.

Vary n_id, 1e6 rows, 10 levels

n_id dataprep reshape2 data.table tidyr pandas polars dask duckdb
1 1.632 151.3 (92.7×) 345.4 (212×) 46.67 (28.6×) 58.85 (36.1×) 105.8 (64.8×) 80.65 (49.4×) 158.7 (97.3×)
2 2.268 206.9 (91.2×) 375.0 (165×) 61.10 (26.9×) 83.84 (37.0×) 111.2 (49.0×) 112.3 (49.5×) 308.8 (136×)
10 3.457 1230 (356×) 437.7 (127×) 94.13 (27.2×) 192.9 (55.8×) 119.6 (34.6×) 242.3 (70.1×) 927.6 (268×)
100 20.14 10242 (509×) 542.0 (26.9×) 413.4 (20.5×) 1332 (66.2×) 158.3 (7.9×) 1644 (81.6×) 8344 (414×)

The 100 id cell is the only case in the entire benchmark suite where a competitor reaches a single-digit ratio. polars is within 7.9×. It remains behind dataprep.

dcast() on Windows 11 Pro for Workstations

The same cells on the Windows host, with no AVX-512:

n_long n_id levels dataprep reshape2 data.table tidyr pandas polars dask duckdb
1e3 1 10 0.440 2.518 (5.7×) 4.337 (9.9×) 7.095 (16.1×) 2.511 (5.7×) 5.061 (11.5×) 15.16 (34.5×) 21.24 (48.3×)
1e4 1 10 0.520 4.744 (9.1×) 8.552 (16.4×) 8.481 (16.3×) 5.161 (9.9×) 5.919 (11.4×) 18.12 (34.8×) 25.88 (49.7×)
1e5 1 10 1.226 35.49 (28.9×) 43.96 (35.9×) 16.07 (13.1×) 24.02 (19.6×) 12.94 (10.6×) 41.88 (34.2×) 73.57 (60.0×)
1e6 1 10 3.940 305.7 (77.6×) 155.1 (39.4×) 98.29 (24.9×) 269.7 (68.4×) 49.29 (12.5×) 359.8 (91.3×) 399.8 (101×)
1e7 1 10 24.34 2821 (116×) 1035 (42.5×) 1509 (62.0×) 2718 (112×) 506.2 (20.8×) 3639 (150×) 3775 (155×)
1e8 1 100 106.3 24253 (228×) 14965 (141×) 13392 (126×) 26311 (248×) 7646 (71.9×) 32874 (309×) 67894 (639×)

The largest Windows multiplier is 639× (duckdb at 1e8 rows and 100 levels). On the Ubuntu host the corresponding cell reaches 408×.

Summary of speedups

Speedup is defined as competitor median / dataprep median. Each table summarises every benchmark cell on that host, across all seven competitors (reshape2, data.table, tidyr, pandas, polars, dask, duckdb).

Ubuntu 25.10

Operation Min Median Mean Max
melt() 0.6× (data.table @ 1e5 × 10 × 1 × 9) 10.3× 58.5× 1187.1× (dask @ 1e3 × 10001 × 1 × 10000)
dcast() 2.0× (reshape2 @ 1e3 × 1 × 10) 49.7× 94.7× 508.6× (reshape2 @ 1e6 × 100 × 10)

Windows 11 Pro for Workstations

Operation Min Median Mean Max
melt() 0.5× (polars @ 1e5 × 19 × 10 × 9) 5.7× 44.2× 882.6× (dask @ 1e3 × 10001 × 1 × 10000)
dcast() 4.5× (polars @ 1e6 × 100 × 10) 47.7× 81.6× 638.8× (duckdb @ 1e8 × 1 × 100)

Combined across both hosts:

  • melt() spans 0.5–1187× across all competitors. The sub-1.0× cells are concentrated in melt() at 1e5 rows: on Ubuntu they are data.table (0.6×) and reshape2 (0.7×); on Windows they include polars (0.5× and 0.7×), data.table (0.8×), and reshape2 (0.9×). Every other cell has dataprep ahead of or on par with the fastest competitor. The median across all melt cells and all competitors is 10.3× on Ubuntu and 5.7× on Windows; the mean is 58.5× and 44.2× respectively.

  • dcast() spans 2.0–639× across all competitors. Every cell has dataprep ahead of every other engine. The median across all dcast cells and all competitors is 49.7× on Ubuntu and 47.7× on Windows; the mean is 94.7× and 81.6× respectively.

  • On the largest cells (1e8 rows, 8 GB of input), dataprep is the only engine that completes within 2 s, specifically < 0.5 s on Ubuntu and < 1.3 s on Windows.

Cross-engine consistency

Both melt() and dcast() produce output byte-identical to reshape2 on every tested cell. All pairs of engines agree pairwise within tol = 1e-12.

melt consistency

rows n_id n_val engines passed pairwise
1,000 1 9 8/8 all consistent
100,000 1 9 8/8 all consistent
1,000 1 100 8/8 all consistent
10,000 10 10 8/8 all consistent

dcast consistency

n_long n_id n_levels engines passed pairwise
5,000 2 5 8/8 all consistent
50,000 1 50 8/8 all consistent
50,000 10 10 8/8 all consistent
1,000,000 1 10 8/8 all consistent

Engines compared: dataprep, reshape2, data.table, tidyr, pandas, polars, dask, duckdb.

Reproducing the benchmarks

The full runner is shipped under inst/:

  • benchmark_helpers.R — adaptive per-tool runner with a 15 s first-call cap
  • benchmark_melt_dcast.R — integrated driver. It runs both the per-tool benchmarks for melt() / dcast() and the 8-engine consistency checks, and prints the tables shown above.

Scripts are disabled by default so that R CMD check does not run them. To enable:

Sys.setenv(DATAPREP_RUN_BENCHMARK = "1")
source(system.file("benchmark_melt_dcast.R", package = "dataprep"))

Session info

sessionInfo()
#> R version 4.6.1 (2026-06-24)
#> Platform: x86_64-pc-linux-gnu
#> Running under: Ubuntu 24.04.5 LTS
#> 
#> Matrix products: default
#> BLAS:   /usr/lib/x86_64-linux-gnu/openblas-pthread/libblas.so.3 
#> LAPACK: /usr/lib/x86_64-linux-gnu/openblas-pthread/libopenblasp-r0.3.26.so;  LAPACK version 3.12.0
#> 
#> locale:
#>  [1] LC_CTYPE=C.UTF-8       LC_NUMERIC=C           LC_TIME=C.UTF-8       
#>  [4] LC_COLLATE=C.UTF-8     LC_MONETARY=C.UTF-8    LC_MESSAGES=C.UTF-8   
#>  [7] LC_PAPER=C.UTF-8       LC_NAME=C              LC_ADDRESS=C          
#> [10] LC_TELEPHONE=C         LC_MEASUREMENT=C.UTF-8 LC_IDENTIFICATION=C   
#> 
#> time zone: UTC
#> tzcode source: system (glibc)
#> 
#> attached base packages:
#> [1] stats     graphics  grDevices utils     datasets  methods   base     
#> 
#> other attached packages:
#> [1] dataprep_0.1.7
#> 
#> loaded via a namespace (and not attached):
#>  [1] digest_0.6.39     desc_1.4.3        R6_2.6.1          fastmap_1.2.0    
#>  [5] xfun_0.61         cachem_1.1.0      parallel_4.6.1    knitr_1.52       
#>  [9] htmltools_0.5.9   rmarkdown_2.32    lifecycle_1.0.5   cli_3.6.6        
#> [13] sass_0.4.10       pkgdown_2.2.1     textshaping_1.0.5 jquerylib_0.1.4  
#> [17] systemfonts_1.3.2 compiler_4.6.1    tools_4.6.1       ragg_1.5.2       
#> [21] bslib_0.12.0      evaluate_1.0.5    Rcpp_1.1.2        yaml_2.3.12      
#> [25] otel_0.2.0        jsonlite_2.0.0    rlang_1.3.0       fs_2.1.0