
dataprep: performance and cross-engine consistency
Source:vignettes/dataprep-performance.Rmd
dataprep-performance.RmdOverview
dataprep 0.1.7 ships two reshaping backends,
melt() and dcast(), benchmarked here against
all seven major alternatives in the R and Python ecosystems:
-
R:
reshape2,data.table,tidyr -
Python:
pandas,polars,dask,duckdb
Every cell is measured with microbenchmark using an
adaptive times rule: 100 iterations when the first call is
under 10 ms, down to a single iteration when it exceeds 10 s. Two
statistics are reported:
- median — the number to quote. Robust to scheduler and GC jitter.
- mean — the tail. A large mean / median ratio indicates one-off OS work or cache effects.
For every cell the tables also report the speed-up of
dataprep relative to each competitor, so the reader can see
the full gradient from “about the same” to “three orders of
magnitude”.
Test environment
Benchmarks were run on two reference hosts. Only the core
configuration is listed here; full hardware details are in
README.md.
Ubuntu 25.10 (Questing Quokka, kernel 6.17.0-41-generic) — 2× AMD EPYC 9965 192-Core (Turin, Zen 5c), 384 physical / 768 logical cores, L3 768 MiB, 1.0 TiB (16 × 64 GiB Micron, DDR5-5600, Multi-bit ECC), full AVX-512; R 4.5.1, g++ 15.2.0.
Windows 11 Pro for Workstations (10.0.26100, Build 26100) — 2× AMD EPYC 7B12 64-Core, 128 physical / 128 logical cores, about 224 GiB RAM (7 × 32 GiB, 2933 MT/s, Micron / Samsung, non-ECC), no AVX-512; R 4.6.1 (ucrt), GCC 14.3.0.
Software versions on both hosts: data.table 1.18.6.1,
reshape2 1.4.5, tidyr 1.3.2,
reticulate 1.47.0; Python 3.13.7 (Ubuntu) / 3.13.15
(Windows), pandas 3.0.6, polars 1.44.2
(runtime rt64), dask 2026.8.0, duckdb
1.5.5.
How to read these numbers
The two hosts differ in core count, cache size and memory bandwidth. Two properties shape the numbers that follow:
The Ubuntu host is unusually large. Most of the 1e6- and 1e7-row cells fit entirely in L3. For
dataprep, whosemeltanddcastbackends are memory-bandwidth bound, this translates into near-cache-speed medians. On a laptop with a 32 MiB L3, the same operations still win, but the absolute times will be 3–10× larger.The Windows host has no AVX-512. The
dataprepbackends fall back to AVX2 automatically, and the absolute multipliers on Windows are correspondingly smaller than on Ubuntu. The relative ranking of the engines is identical on both hosts.
Both effects favour dataprep in the numbers below. The
relative ranking is robust; the absolute multipliers — especially the
1187× and 639× figures — should be interpreted as “best-case on a very
large machine”. On a typical 8–16-core workstation the same comparisons
are within 10–100×.
Cleaning pipeline (dataprep 0.1.5 → 0.1.7)
The 0.1.7 release rewrites every heavy cleaning routine in C++. The table below compares against 0.1.5 on three dataset sizes from the same source (SMEAR I Varrio forest). All numbers are speed-up ratios (0.1.5 time / 0.1.7 time) on Ubuntu 25.10.
| Function | 500 rows | 7,640 rows | 49,422 rows |
|---|---|---|---|
varidele |
1.1× | 1.1× | 11.6× |
obsedele |
203× | 424× | 232× |
condextr |
196× | 217× | 1146× |
optisolu |
188× | 77× | 109× |
dataprep |
185× | 228× | 247× |
On Windows 11 Pro for Workstations, the same full-year pipeline gives
obsedele ≈ 648×, condextr ≈ 839×,
shorvalu ≈ 81×, optisolu ≈ 25× (at
cores = 32), and the integrated dataprep call
≈ 173×. varidele is around 1.17× on this cell; this is
expected, since varidele is a single
colMeans(is.na(.)) in both versions and the new code path
has little room for improvement.
Note on
optisolucores. The 0.1.5 implementation could crash whencores > 16, because itsparallel::makeCluster()path gave each worker a full copy of the data. The benchmark above usedcores = 16for both versions to keep the comparison fair. 0.1.7 loads the package on each worker, exports the input data once per worker, and runs each(interval, times)case as a separate task, socores = 64is safe. The practical speed-up on a many-core host is larger. Also note that the optimal parameter values returned byoptisolu()may differ slightly between 0.1.5 and 0.1.7.
melt() — wide to long
Input shapes are described as rows × (n_id + n_val). All
numbers in the cells are medians in milliseconds; the value in
parentheses is dataprep’s speed-up relative to that
competitor.
Vary rows, 1 id + 9 value columns
| rows | dataprep | reshape2 | data.table | tidyr | pandas | polars | dask | duckdb |
|---|---|---|---|---|---|---|---|---|
| 1e3 | 0.176 | 0.373 (2.1×) | 0.272 (1.5×) | 3.009 (17.1×) | 2.162 (12.3×) | 0.612 (3.5×) | 15.61 (88.6×) | 4.714 (26.8×) |
| 1e4 | 0.235 | 0.436 (1.9×) | 0.341 (1.5×) | 3.379 (14.4×) | 2.517 (10.7×) | 0.685 (2.9×) | 15.59 (66.4×) | 11.00 (46.9×) |
| 1e5 | 1.588 | 1.138 (0.7×) | 1.018 (0.6×) | 8.230 (5.2×) | 6.660 (4.2×) | 1.557 (1.0×) | 17.99 (11.3×) | 68.79 (43.3×) |
| 1e6 | 3.680 | 17.89 (4.9×) | 7.900 (2.1×) | 76.77 (20.9×) | 60.30 (16.4×) | 10.44 (2.8×) | 46.49 (12.6×) | 642.1 (174×) |
| 1e7 | 37.95 | 372.7 (9.8×) | 371.6 (9.8×) | 1111 (29.3×) | 720.1 (19.0×) | 92.53 (2.4×) | 486.4 (12.8×) | 6423 (169×) |
| 1e8 | 496.4 | 3577 (7.2×) | 3576 (7.2×) | 12089 (24.4×) | 7410 (14.9×) | 2563 (5.2×) | 4590 (9.2×) | 65624 (132×) |
The sub-1.0× cells are reshape2 (0.7×) and
data.table (0.6×) at 1e5 rows with one id
column.
Vary rows, 10 id (5 int + 5 chr)
| rows | dataprep | reshape2 | data.table | tidyr | pandas | polars | dask | duckdb |
|---|---|---|---|---|---|---|---|---|
| 1e3 | 0.231 | 0.540 (2.3×) | 0.422 (1.8×) | 3.123 (13.5×) | 5.626 (24.4×) | 0.954 (4.1×) | 53.07 (230×) | 10.64 (46.1×) |
| 1e4 | 0.475 | 1.984 (4.2×) | 1.760 (3.7×) | 4.941 (10.4×) | 6.090 (12.8×) | 1.889 (4.0×) | 54.21 (114×) | 50.11 (106×) |
| 1e5 | 3.752 | 16.74 (4.5×) | 14.90 (4.0×) | 23.06 (6.1×) | 13.12 (3.5×) | 3.944 (1.1×) | 57.69 (15.4×) | 469.6 (125×) |
| 1e6 | 21.28 | 206.2 (9.7×) | 153.7 (7.2×) | 219.9 (10.3×) | 90.42 (4.2×) | 37.50 (1.8×) | 106.8 (5.0×) | 4666 (219×) |
| 1e7 | 868.6 | 3168 (3.6×) | 2703 (3.1×) | 3731 (4.3×) | 1357 (1.6×) | 562.3 (0.6×) | 928.2 (1.1×) | 48519 (55.9×) |
Vary value columns, 1e3 rows, 1 id
| n_val | dataprep | reshape2 | data.table | tidyr | pandas | polars | dask | duckdb |
|---|---|---|---|---|---|---|---|---|
| 10 | 0.174 | 0.373 (2.1×) | 0.267 (1.5×) | 2.956 (17.0×) | 2.143 (12.3×) | 0.608 (3.5×) | 16.65 (95.6×) | 4.853 (27.9×) |
| 100 | 0.260 | 1.081 (4.2×) | 0.381 (1.5×) | 3.779 (14.5×) | 6.458 (24.8×) | 0.726 (2.8×) | 50.90 (196×) | 21.25 (81.7×) |
| 1000 | 0.973 | 8.141 (8.4×) | 1.350 (1.4×) | 12.26 (12.6×) | 48.29 (49.6×) | 3.535 (3.6×) | 370.9 (381×) | 174.1 (179×) |
| 10000 | 3.616 | 92.45 (25.6×) | 9.659 (2.7×) | 107.7 (29.8×) | 499.0 (138×) | 16.31 (4.5×) | 4292 (1187×) | 1914 (529×) |
The 1e3 × 10000 cell is the widest gap in the entire benchmark suite:
dataprep returns in 3.62 ms, dask in 4.29 s,
and duckdb in 1.91 s.
Vary value columns, 1e3 rows, 10 id
| n_val | dataprep | reshape2 | data.table | tidyr | pandas | polars | dask | duckdb |
|---|---|---|---|---|---|---|---|---|
| 10 | 0.247 | 0.613 (2.5×) | 0.459 (1.9×) | 3.114 (12.6×) | 5.655 (22.9×) | 0.915 (3.7×) | 53.18 (216×) | 12.06 (48.9×) |
| 100 | 0.529 | 2.783 (5.3×) | 1.964 (3.7×) | 5.513 (10.4×) | 20.76 (39.2×) | 2.051 (3.9×) | 186.4 (352×) | 59.88 (113×) |
| 1000 | 4.304 | 25.15 (5.8×) | 16.73 (3.9×) | 28.76 (6.7×) | 168.6 (39.2×) | 6.207 (1.4×) | 1614 (375×) | 566.6 (132×) |
| 10000 | 25.22 | 313.6 (12.4×) | 169.0 (6.7×) | 281.7 (11.2×) | 1793 (71.1×) | 59.36 (2.4×) | 22456 (891×) | 5792 (230×) |
melt() on Windows 11 Pro for Workstations
The same cells on the Windows host, with no AVX-512:
| rows | n_id | n_val | dataprep | reshape2 | data.table | tidyr | pandas | polars | dask | duckdb |
|---|---|---|---|---|---|---|---|---|---|---|
| 1e3 | 1 | 9 | 0.319 | 0.653 (2.0×) | 0.512 (1.6×) | 4.224 (13.2×) | 3.575 (11.2×) | 0.513 (1.6×) | 28.55 (89.5×) | 7.724 (24.2×) |
| 1e4 | 1 | 9 | 0.533 | 0.938 (1.8×) | 0.780 (1.5×) | 5.208 (9.8×) | 5.369 (10.1×) | 0.795 (1.5×) | 29.60 (55.6×) | 22.33 (41.9×) |
| 1e5 | 1 | 9 | 3.964 | 3.485 (0.9×) | 3.131 (0.8×) | 16.43 (4.1×) | 24.99 (6.3×) | 2.692 (0.7×) | 44.12 (11.1×) | 175.9 (44.4×) |
| 1e6 | 1 | 9 | 12.15 | 24.44 (2.0×) | 26.17 (2.2×) | 174.4 (14.3×) | 210.6 (17.3×) | 17.99 (1.5×) | 178.8 (14.7×) | 1558 (128×) |
| 1e7 | 1 | 9 | 101.4 | 262.2 (2.6×) | 253.3 (2.5×) | 1510 (14.9×) | 1908 (18.8×) | 177.5 (1.8×) | 1541 (15.2×) | 14588 (144×) |
| 1e8 | 1 | 9 | 1197 | 3514 (2.9×) | 2806 (2.3×) | 21106 (17.6×) | 22968 (19.2×) | 5227 (4.4×) | 16904 (14.1×) | 160256 (134×) |
The largest Windows multiplier is 883× (dask at 1e3 rows
and 10000 value columns).
dcast() — long to wide
Input is a canonical long table with every
(id, variable) pair present exactly once. All numbers in
the cells are medians in milliseconds; the value in parentheses is
dataprep’s speed-up relative to that competitor.
Vary n_long, 1 id, 10 levels
| n_long | dataprep | reshape2 | data.table | tidyr | pandas | polars | dask | duckdb |
|---|---|---|---|---|---|---|---|---|
| 1e3 | 0.859 | 1.689 (2.0×) | 1.818 (2.1×) | 4.349 (5.1×) | 1.889 (2.2×) | 30.04 (35.0×) | 8.351 (9.7×) | 7.584 (8.8×) |
| 1e4 | 0.874 | 2.649 (3.0×) | 2.578 (3.0×) | 4.630 (5.3×) | 2.352 (2.7×) | 38.20 (43.7×) | 9.035 (10.3×) | 10.51 (12.0×) |
| 1e5 | 1.014 | 19.86 (19.6×) | 14.74 (14.5×) | 7.969 (7.9×) | 7.115 (7.0×) | 51.95 (51.2×) | 15.03 (14.8×) | 33.84 (33.4×) |
| 1e6 | 1.632 | 151.3 (92.7×) | 345.4 (212×) | 46.67 (28.6×) | 58.85 (36.1×) | 105.8 (64.8×) | 80.65 (49.4×) | 158.7 (97.3×) |
| 1e7 | 8.417 | 1888 (224×) | 557.6 (66.2×) | 740.2 (87.9×) | 808.2 (96.0×) | 307.8 (36.6×) | 968.7 (115×) | 1756 (209×) |
| 1e8 | 100.4 | 21075 (210×) | 17434 (174×) | 9711 (96.7×) | 10758 (107×) | 2429 (24.2×) | 12982 (129×) | 17061 (170×) |
Vary levels, 1 id, 1e6 rows
| levels | dataprep | reshape2 | data.table | tidyr | pandas | polars | dask | duckdb |
|---|---|---|---|---|---|---|---|---|
| 10 | 1.632 | 151.3 (92.7×) | 345.4 (212×) | 46.67 (28.6×) | 58.85 (36.1×) | 105.8 (64.8×) | 80.65 (49.4×) | 158.7 (97.3×) |
| 100 | 1.451 | 108.7 (74.9×) | 335.9 (232×) | 43.25 (29.8×) | 56.27 (38.8×) | 172.6 (119×) | 77.04 (53.1×) | 180.9 (125×) |
| 1000 | 1.966 | 106.8 (54.3×) | 183.9 (93.5×) | 44.37 (22.6×) | 57.64 (29.3×) | 208.4 (106×) | 79.51 (40.4×) | 200.5 (102×) |
Vary n_long, 1 id, 100 levels
| n_long | dataprep | reshape2 | data.table | tidyr | pandas | polars | dask | duckdb |
|---|---|---|---|---|---|---|---|---|
| 1e4 | 1.011 | 2.885 (2.9×) | 2.926 (2.9×) | 4.851 (4.8×) | 2.487 (2.5×) | 37.85 (37.4×) | 9.310 (9.2×) | 15.28 (15.1×) |
| 1e5 | 1.144 | 18.12 (15.8×) | 8.935 (7.8×) | 8.016 (7.0×) | 7.015 (6.1×) | 55.99 (48.9×) | 14.78 (12.9×) | 41.93 (36.7×) |
| 1e6 | 1.393 | 108.9 (78.2×) | 320.7 (230×) | 43.53 (31.3×) | 56.04 (40.2×) | 160.4 (115×) | 77.95 (56.0×) | 175.7 (126×) |
| 1e7 | 4.733 | 2084 (440×) | 550.9 (116×) | 633.0 (134×) | 768.4 (162×) | 458.8 (96.9×) | 953.6 (202×) | 1717 (363×) |
| 1e8 | 43.91 | 17115 (390×) | 16193 (369×) | 8174 (186×) | 9880 (225×) | 2467 (56.2×) | 12217 (278×) | 17895 (408×) |
The 1e8 × 100 levels cell is the strongest dcast result
on this host: dataprep returns in 43.9 ms,
reshape2 in 17.1 s.
Vary n_id, 1e6 rows, 10 levels
| n_id | dataprep | reshape2 | data.table | tidyr | pandas | polars | dask | duckdb |
|---|---|---|---|---|---|---|---|---|
| 1 | 1.632 | 151.3 (92.7×) | 345.4 (212×) | 46.67 (28.6×) | 58.85 (36.1×) | 105.8 (64.8×) | 80.65 (49.4×) | 158.7 (97.3×) |
| 2 | 2.268 | 206.9 (91.2×) | 375.0 (165×) | 61.10 (26.9×) | 83.84 (37.0×) | 111.2 (49.0×) | 112.3 (49.5×) | 308.8 (136×) |
| 10 | 3.457 | 1230 (356×) | 437.7 (127×) | 94.13 (27.2×) | 192.9 (55.8×) | 119.6 (34.6×) | 242.3 (70.1×) | 927.6 (268×) |
| 100 | 20.14 | 10242 (509×) | 542.0 (26.9×) | 413.4 (20.5×) | 1332 (66.2×) | 158.3 (7.9×) | 1644 (81.6×) | 8344 (414×) |
The 100 id cell is the only case in the entire benchmark
suite where a competitor reaches a single-digit ratio.
polars is within 7.9×. It remains behind
dataprep.
dcast() on Windows 11 Pro for Workstations
The same cells on the Windows host, with no AVX-512:
| n_long | n_id | levels | dataprep | reshape2 | data.table | tidyr | pandas | polars | dask | duckdb |
|---|---|---|---|---|---|---|---|---|---|---|
| 1e3 | 1 | 10 | 0.440 | 2.518 (5.7×) | 4.337 (9.9×) | 7.095 (16.1×) | 2.511 (5.7×) | 5.061 (11.5×) | 15.16 (34.5×) | 21.24 (48.3×) |
| 1e4 | 1 | 10 | 0.520 | 4.744 (9.1×) | 8.552 (16.4×) | 8.481 (16.3×) | 5.161 (9.9×) | 5.919 (11.4×) | 18.12 (34.8×) | 25.88 (49.7×) |
| 1e5 | 1 | 10 | 1.226 | 35.49 (28.9×) | 43.96 (35.9×) | 16.07 (13.1×) | 24.02 (19.6×) | 12.94 (10.6×) | 41.88 (34.2×) | 73.57 (60.0×) |
| 1e6 | 1 | 10 | 3.940 | 305.7 (77.6×) | 155.1 (39.4×) | 98.29 (24.9×) | 269.7 (68.4×) | 49.29 (12.5×) | 359.8 (91.3×) | 399.8 (101×) |
| 1e7 | 1 | 10 | 24.34 | 2821 (116×) | 1035 (42.5×) | 1509 (62.0×) | 2718 (112×) | 506.2 (20.8×) | 3639 (150×) | 3775 (155×) |
| 1e8 | 1 | 100 | 106.3 | 24253 (228×) | 14965 (141×) | 13392 (126×) | 26311 (248×) | 7646 (71.9×) | 32874 (309×) | 67894 (639×) |
The largest Windows multiplier is 639× (duckdb at 1e8
rows and 100 levels). On the Ubuntu host the corresponding cell reaches
408×.
Summary of speedups
Speedup is defined as
competitor median / dataprep median. Each table summarises
every benchmark cell on that host, across all seven competitors
(reshape2, data.table, tidyr,
pandas, polars, dask,
duckdb).
Ubuntu 25.10
| Operation | Min | Median | Mean | Max |
|---|---|---|---|---|
melt() |
0.6× (data.table @ 1e5 × 10 × 1 × 9) | 10.3× | 58.5× | 1187.1× (dask @ 1e3 × 10001 × 1 × 10000) |
dcast() |
2.0× (reshape2 @ 1e3 × 1 × 10) | 49.7× | 94.7× | 508.6× (reshape2 @ 1e6 × 100 × 10) |
Windows 11 Pro for Workstations
| Operation | Min | Median | Mean | Max |
|---|---|---|---|---|
melt() |
0.5× (polars @ 1e5 × 19 × 10 × 9) | 5.7× | 44.2× | 882.6× (dask @ 1e3 × 10001 × 1 × 10000) |
dcast() |
4.5× (polars @ 1e6 × 100 × 10) | 47.7× | 81.6× | 638.8× (duckdb @ 1e8 × 1 × 100) |
Combined across both hosts:
melt()spans 0.5–1187× across all competitors. The sub-1.0× cells are concentrated inmelt()at 1e5 rows: on Ubuntu they aredata.table(0.6×) andreshape2(0.7×); on Windows they includepolars(0.5× and 0.7×),data.table(0.8×), andreshape2(0.9×). Every other cell hasdataprepahead of or on par with the fastest competitor. The median across allmeltcells and all competitors is 10.3× on Ubuntu and 5.7× on Windows; the mean is 58.5× and 44.2× respectively.dcast()spans 2.0–639× across all competitors. Every cell hasdataprepahead of every other engine. The median across alldcastcells and all competitors is 49.7× on Ubuntu and 47.7× on Windows; the mean is 94.7× and 81.6× respectively.On the largest cells (1e8 rows, 8 GB of input),
dataprepis the only engine that completes within 2 s, specifically < 0.5 s on Ubuntu and < 1.3 s on Windows.
Cross-engine consistency
Both melt() and dcast() produce output
byte-identical to reshape2 on every tested cell. All pairs
of engines agree pairwise within tol = 1e-12.
Reproducing the benchmarks
The full runner is shipped under inst/:
-
benchmark_helpers.R— adaptive per-tool runner with a 15 s first-call cap -
benchmark_melt_dcast.R— integrated driver. It runs both the per-tool benchmarks formelt()/dcast()and the 8-engine consistency checks, and prints the tables shown above.
Scripts are disabled by default so that R CMD check does
not run them. To enable:
Sys.setenv(DATAPREP_RUN_BENCHMARK = "1")
source(system.file("benchmark_melt_dcast.R", package = "dataprep"))Session info
sessionInfo()
#> R version 4.6.1 (2026-06-24)
#> Platform: x86_64-pc-linux-gnu
#> Running under: Ubuntu 24.04.5 LTS
#>
#> Matrix products: default
#> BLAS: /usr/lib/x86_64-linux-gnu/openblas-pthread/libblas.so.3
#> LAPACK: /usr/lib/x86_64-linux-gnu/openblas-pthread/libopenblasp-r0.3.26.so; LAPACK version 3.12.0
#>
#> locale:
#> [1] LC_CTYPE=C.UTF-8 LC_NUMERIC=C LC_TIME=C.UTF-8
#> [4] LC_COLLATE=C.UTF-8 LC_MONETARY=C.UTF-8 LC_MESSAGES=C.UTF-8
#> [7] LC_PAPER=C.UTF-8 LC_NAME=C LC_ADDRESS=C
#> [10] LC_TELEPHONE=C LC_MEASUREMENT=C.UTF-8 LC_IDENTIFICATION=C
#>
#> time zone: UTC
#> tzcode source: system (glibc)
#>
#> attached base packages:
#> [1] stats graphics grDevices utils datasets methods base
#>
#> other attached packages:
#> [1] dataprep_0.1.7
#>
#> loaded via a namespace (and not attached):
#> [1] digest_0.6.39 desc_1.4.3 R6_2.6.1 fastmap_1.2.0
#> [5] xfun_0.61 cachem_1.1.0 parallel_4.6.1 knitr_1.52
#> [9] htmltools_0.5.9 rmarkdown_2.32 lifecycle_1.0.5 cli_3.6.6
#> [13] sass_0.4.10 pkgdown_2.2.1 textshaping_1.0.5 jquerylib_0.1.4
#> [17] systemfonts_1.3.2 compiler_4.6.1 tools_4.6.1 ragg_1.5.2
#> [21] bslib_0.12.0 evaluate_1.0.5 Rcpp_1.1.2 yaml_2.3.12
#> [25] otel_0.2.0 jsonlite_2.0.0 rlang_1.3.0 fs_2.1.0