{hash}This vignette provides a
comparison of {r2r} with the same-purpose CRAN package {hash},
which also offers an implementation of hash tables based on R
environments. We first describe the features offered by both packages,
and then perform some benchmark timing comparisons. The package versions
referred to in this vignette are:
library(hash)
library(r2r)
packageVersion("hash")
#> [1] '2.2.6.4'
packageVersion("r2r")
#> [1] '0.1.2'Both {r2r} and {hash} hash tables are built
on top of the R built-in environment data structure, and
have thus a similar API. In particular, hash table objects have
reference semantics for both packages. {r2r}
hashtables are S3 class objects, whereas in
{hash} the data structure is implemented as an S4
class.
Hash tables provided by r2r support arbitrary type keys
and values, arbitrary key comparison and hash functions, and have
customizable behaviour (either throw an exception or return a default
value) upon query of a missing key.
In contrast, hash tables in hash currently support only
string keys, with basic identity comparison (the hashing is performed
automatically by the underlying environment objects);
values can be arbitrary R objects. Querying missing keys through
non-vectorized [[-subsetting returns the default value
NULL, whereas queries through vectorized
[-subsetting result in an error. On the other hand,
hash also offers support for inverting hash tables (an
experimental feature at the time of writing).
The table below summarizes the features of the two packages
| Feature | r2r | hash |
|---|---|---|
| Basic data structure | R environment | R environment |
| Arbitrary type keys | X | |
| Arbitrary type values | X | X |
| Arbitrary hash function | X | |
| Arbitrary key comparison function | X | |
| Throw or return default on missing keys | X | |
| Hash table inversion | X |
We will perform our benchmark tests using the CRAN package microbenchmark.
We start by timing the insertion of:
random key-value pairs (with possible repetitions). In order to perform a meaningful comparison between the two packages, we restrict to string (i.e. length one character) keys. We can generate random keys as follows:
chars <- c(letters, LETTERS, 0:9)
random_keys <- function(n) paste0(
sample(chars, n, replace = TRUE),
sample(chars, n, replace = TRUE),
sample(chars, n, replace = TRUE),
sample(chars, n, replace = TRUE),
sample(chars, n, replace = TRUE)
)
set.seed(840)
keys <- random_keys(N)
values <- rnorm(N)We test both the non-vectorized ([[<-) and vectorized
([<-) operators:
microbenchmark(
`r2r_[[<-` = {
for (i in seq_along(keys))
m_r2r[[ keys[[i]] ]] <- values[[i]]
},
`r2r_[<-` = { m_r2r[keys] <- values },
`hash_[[<-` = {
for (i in seq_along(keys))
m_hash[[ keys[[i]] ]] <- values[[i]]
},
`hash_[<-` = m_hash[keys] <- values,
times = 30,
setup = { m_r2r <- hashmap(); m_hash <- hash() }
)
#> Unit: milliseconds
#> expr min lq mean median uq max neval
#> r2r_[[<- 98.30017 129.21737 168.36815 168.26847 203.24199 258.1399 30
#> r2r_[<- 69.21572 92.36945 139.38775 139.51036 170.67228 322.3156 30
#> hash_[[<- 73.51051 111.01555 134.30033 127.65051 144.80263 239.5006 30
#> hash_[<- 41.20582 69.54075 81.07606 84.44059 92.25313 140.8182 30As it is seen, r2r and hash have comparable
performances at the insertion of key-value pairs, with both vectorized
and non-vectorized insertions, hash being somewhat more
efficient in both cases.
We now test key query, again both in non-vectorized and vectorized form:
microbenchmark(
`r2r_[[` = { for (key in keys) m_r2r[[ key ]] },
`r2r_[` = { m_r2r[ keys ] },
`hash_[[` = { for (key in keys) m_hash[[ key ]] },
`hash_[` = { m_hash[ keys ] },
times = 30,
setup = {
m_r2r <- hashmap(); m_r2r[keys] <- values
m_hash <- hash(); m_hash[keys] <- values
}
)
#> Unit: milliseconds
#> expr min lq mean median uq max neval
#> r2r_[[ 96.07889 135.50293 185.64756 188.43839 233.46516 289.67895 30
#> r2r_[ 92.25452 136.19839 171.05932 177.39769 213.41532 249.42809 30
#> hash_[[ 11.32204 12.47591 17.06267 14.85390 21.46659 29.96053 30
#> hash_[ 63.57955 76.25621 104.67501 98.25973 135.63418 158.75248 30For non-vectorized queries, hash is significantly faster
(by one order of magnitude) than r2r. This is likely due to
the fact that the [[ method dispatch is handled natively by
R in hash (i.e. the default [[ method
for environments is used ), whereas r2r
suffers the overhead of S3 method dispatch. This is confirmed by the
result for vectorized queries, which is comparable for the two packages;
notice that here a single (rather than N) S3 method
dispatch occurs in the r2r timed expression.
As an additional test, we perform the benchmarks for non-vectorized expressions with a new set of keys:
set.seed(841)
new_keys <- random_keys(N)
microbenchmark(
`r2r_[[_bis` = { for (key in new_keys) m_r2r[[ key ]] },
`hash_[[_bis` = { for (key in new_keys) m_hash[[ key ]] },
times = 30,
setup = {
m_r2r <- hashmap(); m_r2r[keys] <- values
m_hash <- hash(); m_hash[keys] <- values
}
)
#> Unit: milliseconds
#> expr min lq mean median uq max neval
#> r2r_[[_bis 71.58680 96.61528 127.11334 125.18600 159.67464 226.21784 30
#> hash_[[_bis 11.47926 12.57271 18.66148 18.26924 24.03328 32.31627 30The results are similar to the ones already commented. Finally, we
test the performances of the two packages in checking the existence of
keys (notice that here has_key refers to
r2r::has_key, whereas has.key is
hash::has.key):
set.seed(842)
mixed_keys <- sample(c(keys, new_keys), N)
microbenchmark(
r2r_has_key = { for (key in mixed_keys) has_key(m_r2r, key) },
hash_has_key = { for (key in new_keys) has.key(key, m_hash) },
times = 30,
setup = {
m_r2r <- hashmap(); m_r2r[keys] <- values
m_hash <- hash(); m_hash[keys] <- values
}
)
#> Unit: milliseconds
#> expr min lq mean median uq max neval
#> r2r_has_key 74.30419 85.32149 106.2206 108.9638 123.8719 141.9533 30
#> hash_has_key 192.51283 234.37185 304.9381 300.7907 374.1574 439.1151 30The results are comparable for the two packages, r2r
being slightly more performant in this particular case.
Finally, we test key deletion. In order to handle name collisions, we
will use delete() (which refers to
r2r::delete()) and del() (which refers to
hash::del()).
microbenchmark(
r2r_delete = { for (key in keys) delete(m_r2r, key) },
hash_delete = { for (key in keys) del(key, m_hash) },
hash_vectorized_delete = { del(keys, m_hash) },
times = 30,
setup = {
m_r2r <- hashmap(); m_r2r[keys] <- values
m_hash <- hash(); m_hash[keys] <- values
}
)
#> Unit: milliseconds
#> expr min lq mean median uq
#> r2r_delete 128.781056 149.410034 201.068762 197.595870 239.84616
#> hash_delete 65.164356 82.383820 112.735360 114.210446 136.74302
#> hash_vectorized_delete 2.870008 3.208157 3.860566 3.677855 4.40237
#> max neval
#> 291.472659 30
#> 196.634761 30
#> 6.294353 30The vectorized version of hash significantly outperforms
the non-vectorized versions (by roughly two orders of magnitude in
speed). Currently, r2r does not support vectorized key
deletion 1.
The two R packages r2r and hash offer hash
table implementations with different advantages and drawbacks.
r2r focuses on flexibility, and has a richer set of
features. hash is more minimal, but offers superior
performance in some important tasks. Finally, as a positive note for
both parties, the two packages share a similar API, making it relatively
easy to switch between the two, according to the particular use case
needs.
This is due to complications introduced by the internal
hash collision handling system of r2r.↩︎