
Cast a long-format data.frame into a wide format
dcast.RdStructural inverse of melt: a long-format data frame
with one row per (id, variable) pair is reshaped into a
wide-format data frame with one row per id combination and
one column per variable level. The C++ backend builds
compact integer lookup tables for both the row keys (id columns)
and the column keys (variable column), then writes values by
output column in strictly sequential order.
Usage
dcast(data, id = NULL, formula = NULL,
variable = NULL, value = NULL, value.var = NULL,
fill = NA_real_, fun.aggregate = NULL,
na.rm = FALSE, cores = 0L, verbose = FALSE)Arguments
- data
A data frame in long format. It must contain one column identifying the variable (e.g.
"variable"), one column holding the value (e.g."value"), and one or more id columns.- id
Identifier columns. Character names, integer indices, a logical mask, or
NULL. IfNULL, all columns other than the variable and value columns are used as ids.- formula
Optional formula of the form
id1 + id2 ~ variable. When supplied, the left-hand side overridesidand the right-hand side overridesvariable.- variable
Name (or index) of the column identifying the variable. If
NULL, a column named"variable","variables","Variable","Variables","VARIABLE", or"VARIABLES"is used when exactly one such column exists.- value
Name (or index) of the column holding the values. If
NULL, a column named"value","values","Value","Values","VALUE", or"VALUES"is used when exactly one such column exists.- value.var
Alias for
value. Used only whenvalueisNULL; if both are supplied,valuetakes precedence.- fill
Value used to fill cells for
(id, variable)pairs that do not appear in the input. Cells that appear in the input with valueNAkeepNAunlessna.rm = TRUE. DefaultNA_real_.- fun.aggregate
Optional function used to reduce duplicate
(id, variable)pairs. Called with a single argument (the vector of values for that pair) plusna.rm = TRUE. Without it, the last occurrence wins.- na.rm
Logical; if
TRUE, rows whose value isNAorNaNare skipped during scatter, so the corresponding output cell keeps thefillvalue. DefaultFALSE.- cores
Number of OpenMP threads.
0(default) lets the C++ backend choose based on data size.- verbose
Logical; if
TRUE, prints timing information.
Details
dcast() is the structural inverse of melt. The
C++ backend detects canonical melt() output automatically
(the variable column is periodic and every id column is constant
within one period) and switches to a block-path tile
transpose: each tile of TILE x period doubles is read
contiguously into an L1 buffer, transposed in place, and written
contiguously to the output columns. This keeps both reads and
writes sequential and enables OpenMP parallelisation. The cost per
tile is O(TILE * period), independent of the number of
levels, which is why wide-level tables scale well.
When the input is not block-aligned, dcast() builds an
open-addressing hash of 64-bit packed row keys. If the combined bit
budget of the id columns exceeds 64, it falls back to a 96-bit
fingerprint (uint64_t + uint32_t) computed by a 4-way
parallel FNV-1a and stored in a 16-byte slot table. For
shuffle-friendly input, a 4-pass LSD radix sort over the packed
keys replaces the hash table entirely; and when the block path is
taken and the id column is a permutation of 1..n_blocks, a
direct-index shortcut bypasses both.
The output is identical to reshape2::dcast,
data.table::dcast, tidyr::pivot_wider,
pandas.pivot, polars.pivot, duckdb PIVOT, and
dask on every tested shape, within tol = 1e-12. See
vignette("dataprep-melt-dcast") for the implementation notes
and vignette("dataprep-performance") for the benchmark
tables.
On the canonical long-to-wide shape, dcast() is faster than
every tested alternative. The speed-up relative to reshape2,
data.table, tidyr, pandas, polars,
dask, and duckdb spans 2.0x (against
reshape2 on 1e3 rows and 10 levels, Ubuntu 25.10) to
639x (against duckdb on 1e8 rows and 100 levels,
Windows 11 Pro for Workstations). The median across all tested
cells and all competitors is 49.7x on Ubuntu 25.10 and
47.7x on Windows 11 Pro for Workstations; the mean is
94.7x and 81.6x, respectively.
Value
A wide-format data frame with one row per unique id
combination and one column per unique value of the variable column,
plus the id columns. Column names are the string representation of
the variable values.
References
Liang, C.-S., Wu, H., Li, H.-Y., Zhang, Q., Li, Z. & He, K.-B. (2020). Efficient data preprocessing, episode classification, and source apportionment of particle number concentrations. Science of the Total Environment, 741, 140923. doi:10.1016/j.scitotenv.2020.140923
Author
Chun-Sheng Liang chun-shengliang@qq.com
See also
melt for the inverse operation.
vignette("dataprep-melt-dcast") for usage notes and
implementation.
Examples
## --- Basic usage --------------------------------------------------
long <- data.frame(
id = rep(1:3, each = 2),
variable = rep(c("x", "y"), 3),
value = c(1, 2, 3, 4, 5, 6)
)
dcast(long, id = "id", variable = "variable", value = "value")
#> id x y
#> 1 1 1 2
#> 2 2 3 4
#> 3 3 5 6
## --- Formula interface --------------------------------------------
dcast(long, formula = id ~ variable)
#> id x y
#> 1 1 1 2
#> 2 2 3 4
#> 3 3 5 6
## --- Multiple id columns ------------------------------------------
long2 <- data.frame(
year = rep(2020:2021, each = 4),
city = rep(c("A", "B"), each = 2, times = 2),
variable = rep(c("temp", "rain"), 4),
value = c(15.2, 210, 14.8, 180, 16.1, 230, 15.5, 195)
)
# Explicit id columns
dcast(long2, id = c("year", "city"),
variable = "variable", value = "value")
#> year city temp rain
#> 1 2020 A 15.2 210
#> 2 2020 B 14.8 180
#> 3 2021 A 16.1 230
#> 4 2021 B 15.5 195
# Formula form (equivalent)
dcast(long2, formula = year + city ~ variable)
#> year city temp rain
#> 1 2020 A 15.2 210
#> 2 2020 B 14.8 180
#> 3 2021 A 16.1 230
#> 4 2021 B 15.5 195
## --- Fill missing cells -------------------------------------------
# (2, "y") is missing from the input; fill = 0 gives it 0.
long3 <- data.frame(
id = c(1, 1, 2),
variable = c("x", "y", "x"),
value = c(1, 2, 3)
)
dcast(long3, id = "id",
variable = "variable", value = "value",
fill = 0)
#> id x y
#> 1 1 1 2
#> 2 2 3 0
## --- Aggregate duplicate (id, variable) pairs ---------------------
# id = 1 appears twice with variable = "x"; the mean is 1.5.
long4 <- data.frame(
id = c(1, 1, 2),
variable = c("x", "x", "x"),
value = c(1, 2, 3)
)
dcast(long4, id = "id",
variable = "variable", value = "value",
fun.aggregate = mean)
#> id x
#> 1 1 1.5
#> 2 2 3.0
# sum and length work the same way
dcast(long4, id = "id",
variable = "variable", value = "value",
fun.aggregate = sum)
#> id x
#> 1 1 3
#> 2 2 3
## --- Skip NA values during scatter --------------------------------
long5 <- data.frame(
id = c(1, 1, 2, 2),
variable = c("x", "y", "x", "y"),
value = c(1, NA, 3, 4)
)
# Default: NA stays as NA
dcast(long5, id = "id",
variable = "variable", value = "value")
#> id x y
#> 1 1 1 NA
#> 2 2 3 4
# na.rm = TRUE: (1, "y") is skipped, cell keeps the fill value
dcast(long5, id = "id",
variable = "variable", value = "value",
na.rm = TRUE, fill = -1)
#> id x y
#> 1 1 1 NA
#> 2 2 3 4
## --- Round-trip with melt() ---------------------------------------
wide <- data.frame(id = 1:3, a = c(1.5, 2.5, 3.5), b = c(4.5, 5.5, 6.5))
back <- dcast(melt(wide, id.vars = "id"),
id = "id", variable = "variable", value = "value")
back[order(back$id), ]
#> id a b
#> 1 1 1.5 4.5
#> 2 2 2.5 5.5
#> 3 3 3.5 6.5
## --- Larger example: 1000 rows, 50 levels -------------------------
set.seed(1)
n_rows <- 1000L
n_levels <- 50L
long_big <- data.frame(
id = rep(seq_len(n_rows / n_levels), each = n_levels),
variable = rep(sprintf("v%02d", seq_len(n_levels)),
times = n_rows / n_levels),
value = rnorm(n_rows)
)
wide_big <- dcast(long_big, id = "id",
variable = "variable", value = "value")
dim(wide_big)
#> [1] 20 51