| Title: | DuckDB-Backed DataFrame and Table Structures |
|---|---|
| Description: | Provides DuckDB-backed implementations of core tabular data structures including DuckDBTable, DuckDBDataFrame, and DuckDBColumn. These classes extend Bioconductor's S4Vectors framework to enable efficient, out-of-memory operations on large tabular datasets using DuckDB's analytical database engine. The package includes SQL function discovery, connection management, and support for list columns and embeddings stored in ARRAY[] columns. |
| Authors: | Patrick Aboyoun [aut, cre], Genentech, Inc. [cph] |
| Maintainer: | Patrick Aboyoun <[email protected]> |
| License: | MIT + file LICENSE |
| Version: | 0.99.6 |
| Built: | 2026-07-23 10:58:12 UTC |
| Source: | https://github.com/BiocStaging/DuckDBDataFrame |
Utilities for narrowing integer types and converting between
DataType objects and Arrow type names. Used by
BiocDuckDB and DuckDBArray for Parquet schema selection
and parallel-safe type passing.
arrowIntType(range) arrowType(x) arrowTypeToName(type) arrowTypeFromName(name)arrowIntType(range) arrowType(x) arrowTypeToName(type) arrowTypeFromName(name)
range |
Length-two numeric vector |
x |
An R vector for |
type |
An |
name |
Length-one character Arrow type name (e.g. |
arrowIntType selects the narrowest unsigned or signed Arrow integer
type that can represent range. arrowType infers an Arrow type
from an R vector, narrowing integers when possible.
arrowTypeFromName and arrowTypeToName implement a small
registry of scalar Arrow types used when passing type information to
BiocParallel workers (where DataType external pointers do not
serialize reliably).
arrowIntType, arrowType
An DataType.
arrowTypeToNameA length-one character Arrow type name.
arrowTypeFromNameAn DataType.
Patrick Aboyoun
parquet-io,
writeParquet,
writeCoordArray
arrowIntType(c(0L, 10L)) arrowType(1:10L) arrowTypeToName(arrow::uint16()) arrowTypeFromName("uint16")arrowIntType(c(0L, 10L)) arrowType(1:10L) arrowTypeToName(arrow::uint16()) arrowTypeFromName("uint16")
The DuckDBAtomicList class extends AtomicList and DuckDBColumn to represent list columns extracted from a DuckDBDataFrame object.
Objects of class DuckDBAtomicList extend AtomicList.
In the code snippets below, x is a DuckDBAtomicList object:
length(x):Get the number of list elements in x.
names(x):Get the names of the list elements of x.
elementNROWS(x):Get the length of each list element in x.
elementType(x):Get the type of the list elements; one of "logical",
"integer", "integer64", "numeric",
"character", or "factor".
dimtbls(x, drop = TRUE), dimtbls(x) <- value:Get or set the list of dimension tables used to define partitions for
efficient queries. If drop = TRUE, then it returns a named
DataFrameList object, else it returns an environment containing
a dimtbls named DataFrameList element.
as.list(x):Coerces x to a list.
realize(x, BACKEND = getAutoRealizationBackend()):Realize an object into memory or on disk using the equivalent of
realize(as.list(x), BACKEND).
In the code snippets below, x is a DuckDBAtomicList object:
x[i]:Returns a DuckDBAtomicList object containing the selected list elements.
x[[i]]:Extracts the list element at position i as an atomic vector.
head(x, n = 6L):If n is non-negative, returns the first n list elements of x.
If n is negative, returns all but the last abs(n) list
elements of x.
tail(x, n = 6L):If n is non-negative, returns the last n list elements of x.
If n is negative, returns all but the first abs(n) list
elements of x.
Patrick Aboyoun
DuckDBColumn-class for atomic columns
AtomicList for the base class
List for the List class hierarchy
tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) df <- data.frame( id = sprintf("gene%02d", 1:5), int_list = I(lapply(1:5, function(i) seq_len(i))) ) arrow::write_parquet(df, tf) ddf <- DuckDBDataFrame(tf, keycol = "id") lst <- ddf[["int_list"]] lst elementNROWS(lst)tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) df <- data.frame( id = sprintf("gene%02d", 1:5), int_list = I(lapply(1:5, function(i) seq_len(i))) ) arrow::write_parquet(df, tf) ddf <- DuckDBDataFrame(tf, keycol = "id") lst <- ddf[["int_list"]] lst elementNROWS(lst)
The DuckDBColumn class extends Vector to represent a column extracted from a DuckDBDataFrame object.
Objects of class DuckDBColumn extend Vector.
In the code snippets below, x is a DuckDBColumn object:
length(x):Get the number of elements in x.
names(x):Get the names of the elements of x.
dimtbls(x, drop = TRUE), dimtbls(x) <- value:Get or set the list of dimension tables used to define partitions for
efficient queries. If drop = TRUE, then it returns a named
DataFrameList object, else it returns an environment containing
a dimtbls named DataFrameList element.
type(x), type(x) <- value:Get or set the data type of the elements; one of "logical",
"integer", "integer64", "double", or
"character".
as.vector(x):Coerces x to a vector.
realize(x, BACKEND = getAutoRealizationBackend()):Realize an object into memory or on disk using the equivalent of
realize(as.vector(x), BACKEND).
In the code snippets below, x is a DuckDBColumn object:
x[i]:Returns a DuckDBColumn object containing the selected elements.
head(x, n = 6L):If n is non-negative, returns the first n elements of x.
If n is negative, returns all but the last abs(n) elements
of x.
tail(x, n = 6L):If n is non-negative, returns the last n elements of x.
If n is negative, returns all but the first abs(n) elements
of x.
Patrick Aboyoun
DuckDBColumn-utils for the utilities
Vector for the base class
tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) arrow::write_parquet(cbind(model = rownames(mtcars), mtcars), tf) df <- DuckDBDataFrame(tf, datacols = colnames(mtcars), keycol = "model") cyl <- df[["cyl"]] cyl length(cyl) mean(cyl) head(cyl, 3)tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) arrow::write_parquet(cbind(model = rownames(mtcars), mtcars), tf) df <- DuckDBDataFrame(tf, datacols = colnames(mtcars), keycol = "model") cyl <- df[["cyl"]] cyl length(cyl) mean(cyl) head(cyl, 3)
Common operations on DuckDBColumn objects.
Method return types are documented in the sections above.
In the code snippets below, x is a DuckDBColumn object:
sql_call(x, fun, ...):Applies the specified SQL function to all data columns of x.
sql_fun(x, function_type = NULL, return_type = NULL, description = FALSE):Finds DuckDB SQL functions compatible with the data type of x.
Returns a data.frame of matching functions with optional filtering by
function type and return type.
DuckDBColumn objects have support for S4 group generic functionality:
Arith"+", "-", "*", "^",
"%%", "%/%", "/"
Compare"==", ">", "<", "!=",
"<=", ">="
Logic"&", "|", "!"
Ops"Arith", "Compare", "Logic"
Math"abs", "sign", "sqrt",
"ceiling", "floor", "trunc", "log",
"log10", "log2", "acos", "acosh",
"asin", "asinh", "atan", "atanh",
"exp", "expm1", "cos", "cosh",
"sin", "sinh", "tan", "tanh",
"gamma", "lgamma"
Summary"max", "min", "range",
"prod", "sum", "any", "all"
See S4groupGeneric for more details.
In the code snippets below, x is a DuckDBColumn object:
is.finite(x):Returns a DuckDBColumn containing logicals that indicate which values are finite.
is.infinite(x):Returns a DuckDBColumn containing logicals that indicate which values are infinite.
is.nan(x):Returns a DuckDBColumn containing logicals that indicate which values are Not a Number.
mean(x):Calculates the mean of x.
var(x):Calculates the variance of x.
sd(x):Calculates the standard deviation of x.
median(x):Calculates the median of x.
quantile(x, probs = seq(0, 1, 0.25), names = TRUE, type = 7):Calculates the specified quantiles of x.
probsA numeric vector of probabilities with values in [0,1].
namesIf TRUE, the result has names describing the
quantiles.
typeEither 1 or 7 that specifies the quantile algorithm
detailed in quantile.
mad(x, constant = 1.4826):Calculates the median absolute deviation of x.
constantThe scale factor.
IQR(x, type = 7):Calculates the interquartile range of x.
typeEither 1 or 7 that specifies the quantile algorithm
detailed in quantile.
pmax(..., na.rm = FALSE):Returns the parallel maxima of multiple DuckDBColumn objects. All arguments must be DuckDBColumn objects with compatible dimensions.
pmin(..., na.rm = FALSE):Returns the parallel minima of multiple DuckDBColumn objects. All arguments must be DuckDBColumn objects with compatible dimensions.
In the code snippets below, x is a DuckDBColumn object:
nchar(x):Returns a DuckDBColumn containing the number of characters in each string.
tolower(x):Returns a DuckDBColumn with all strings converted to lowercase.
toupper(x):Returns a DuckDBColumn with all strings converted to uppercase.
chartr(old, new, x):Returns a DuckDBColumn with characters translated.
oldCharacters to be translated.
newCharacters to translate to.
substr(x, start, stop):Returns a DuckDBColumn containing substrings extracted by position.
startInteger starting position (1-indexed).
stopInteger ending position (inclusive).
substring(x, first, last = 1000000L):Returns a DuckDBColumn containing substrings extracted by position.
firstInteger starting position (1-indexed).
lastInteger ending position (inclusive).
grepl(pattern, x, ignore.case = FALSE, fixed = FALSE):Returns a DuckDBColumn containing logicals indicating pattern matches.
patternCharacter string containing a regular expression.
ignore.caseIf TRUE, case-insensitive matching.
fixedIf TRUE, pattern is a fixed string not regex.
sub(pattern, replacement, x, ignore.case = FALSE, fixed = FALSE):Returns a DuckDBColumn with first match of pattern replaced.
patternCharacter string containing a regular expression.
replacementReplacement string.
ignore.caseIf TRUE, case-insensitive matching.
fixedIf TRUE, pattern is a fixed string not regex.
gsub(pattern, replacement, x, ignore.case = FALSE, fixed = FALSE):Returns a DuckDBColumn with all matches of pattern replaced.
patternCharacter string containing a regular expression.
replacementReplacement string.
ignore.caseIf TRUE, case-insensitive matching.
fixedIf TRUE, pattern is a fixed string not regex.
startsWith(x, prefix):Returns a DuckDBColumn containing logicals indicating if strings start with the specified prefix.
endsWith(x, suffix):Returns a DuckDBColumn containing logicals indicating if strings end with the specified suffix.
paste(..., sep = " ", collapse = NULL):Concatenates multiple DuckDBColumn objects using the specified separator.
The collapse argument is not supported.
All arguments must be DuckDBColumn objects with compatible dimensions.
In the code snippets below, x is a DuckDBColumn object:
unique(x):Returns a DuckDBColumn containing the distinct rows.
x %in% table:Returns a DuckDBColumn containing logicals that indicate the elements of
x in table.
table(...):Returns a table containing the counts across the distinct values.
In the code snippets below, x is a DuckDBColumn object:
is_nonzero(x):Returns a DuckDBColumn containing logicals that indicate if the elements
of x are non-zero.
nzcount(x):Returns the total number of non-zero values.
Patrick Aboyoun
DuckDBColumn-class for the main class
Vector for the base class
tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) arrow::write_parquet(cbind(model = rownames(mtcars), mtcars), tf) df <- DuckDBDataFrame(tf, datacols = colnames(mtcars), keycol = "model") mpg <- df[["mpg"]] mean(mpg) sd(mpg) unique(mpg)[1:5]tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) arrow::write_parquet(cbind(model = rownames(mtcars), mtcars), tf) df <- DuckDBDataFrame(tf, datacols = colnames(mtcars), keycol = "model") mpg <- df[["mpg"]] mean(mpg) sd(mpg) unique(mpg)[1:5]
Acquire and Release the DuckDB connection used by the BiocDuckDB package.
loadExtension(conn, extension, optional = FALSE) configureOutOfCore(conn) configureCloud(conn) acquireDuckDBConn( conn = dbConnect(duckdb(), bigint = "integer64", array = "matrix") ) releaseDuckDBConn()loadExtension(conn, extension, optional = FALSE) configureOutOfCore(conn) configureCloud(conn) acquireDuckDBConn( conn = dbConnect(duckdb(), bigint = "integer64", array = "matrix") ) releaseDuckDBConn()
conn |
A |
extension |
A character string naming a DuckDB extension (e.g.
|
optional |
If |
acquireDuckDBConn will add a duckdb_connection object to
the BiocDuckDB package cache.
releaseDuckDBConn will remove the duckdb_connection object
from that cache.
loadExtension installs and loads a DuckDB extension on the shared
connection; it is used by companion packages (e.g. DuckDBSpatial) to
ensure their required extension ("spatial") is available.
configureOutOfCore is called automatically by acquireDuckDBConn
when the shared connection is first created; call it again (on
acquireDuckDBConn()) after changing an option to re-apply. Each
setting is read from an R option first, then an environment variable, and
left at the DuckDB default when neither is set:
| Setting | Option | Environment variable |
| memory limit | DuckDBDataFrame.memory_limit |
BIOCDUCKDB_MEMORY_LIMIT |
| threads |
DuckDBDataFrame.threads |
BIOCDUCKDB_THREADS |
| spill directory | DuckDBDataFrame.temp_directory |
BIOCDUCKDB_TEMP_DIRECTORY |
| preserve insertion order | \
DuckDBDataFrame.preserve_insertion_order |
BIOCDUCKDB_PRESERVE_INSERTION_ORDER
|
Setting memory_limit (e.g. "16GB" or "80%") and
temp_directory guards against the OS killing the process before DuckDB
can spill a large aggregation/sort/join; preserve_insertion_order =
FALSE avoids buffering a whole result to preserve row order on a Tahoe-scale
export where order is not significant.
acquireDuckDBConn returns a cached duckdb_connection object.
releaseDuckDBConn returns NULL, invisibly.
loadExtension installs (if needed) and loads extension on
conn, returning the load status invisibly.
configureOutOfCore applies out-of-core engine settings (buffer-pool
memory limit, worker threads, spill directory, insertion-order preservation)
to conn from R options or environment variables, returning NULL
invisibly.
Patrick Aboyoun
releaseDuckDBConn() conn <- acquireDuckDBConn() identical(conn, acquireDuckDBConn()) releaseDuckDBConn() releaseDuckDBConn()releaseDuckDBConn() conn <- acquireDuckDBConn() identical(conn, acquireDuckDBConn()) releaseDuckDBConn() releaseDuckDBConn()
The DuckDBDataFrame class extends both DuckDBTable and DataFrame to represent a DuckDB table as a DataFrame object.
DuckDBDataFrame adds DataFrame semantics to a DuckDB table. It achieves a balance between the flexibility of a tbl_duckdb_connection object and the familiarity of a DataFrame object.
Objects of class DuckDBDataFrame extend DataFrame.
DuckDBDataFrame(conn, datacols = colnames(conn), keycol = NULL, dimtbl = NULL, type = NULL, collevels = NULL):Creates a DuckDBDataFrame object.
connEither a character vector containing the paths to parquet, csv, or
gzipped csv data files; a string that defines a duckdb read_*
data source; a DuckDBDataFrame object; or a tbl_duckdb_connection
object.
datacolsEither a character vector of column names from conn or a
named expression that will be evaluated in the context of
conn that defines the data.
keycolAn optional string specifying the column name from conn that
will define the foreign key in the underlying table, or a named list
containing a character vector where the name of the list element
defines the foreign key and the character vector set the distinct
values for that key. If missing, a row_number column is
created as an identifier.
dimtblA optional named DataFrameList that specifies the dimension
table associated with the keycol. The name of the list
element must match the name of the keycol list. Additionally,
the DataFrame object must have row names that match the
distinct values of the keycol list element and columns
that define partitions in the data table for efficient querying.
typeAn optional named character vector where the names specify the
column names and the values specify the column type; one of
"logical", "integer", "integer64",
"double", or "character".
collevelsAn optional named list, keyed by data column name, that restores
factor columns on materialization (each element a list with a
character levels vector and a TRUE/FALSE
ordered flag). Primarily used by readParquet().
In the code snippets below, x is a DuckDBDataFrame object:
dim(x):Length two integer vector defined as c(nrow(x), ncol(x)).
nrow(x), ncol(x):Get the number of rows and columns, respectively.
NROW(x), NCOL(x):Same as nrow(x) and ncol(x), respectively.
dimnames(x):Length two list of character vectors defined as
list(rownames(x), colnames(x)).
rownames(x), colnames(x):Get the names of the rows and columns, respectively.
coltypes(x), coltypes(x) <- value:Get or set the data type of the columns; one of "logical",
"integer", "integer64", "double", or
"character".
dimtbls(x, drop = TRUE), dimtbls(x) <- value:Get or set the list of dimension tables used to define partitions for
efficient queries. If drop = TRUE, then it returns a named
DataFrameList object, else it returns an environment containing
a dimtbls named DataFrameList element.
as.data.frame(x):Coerces x to a data.frame.
as.matrix(x):Coerces x to a matrix.
as(from, "DFrame"):Converts a DuckDBDataFrame object to a DFrame object. This process
involves first loading the data into memory using as.data.frame.
The resulting data.frame is then coerced into a DFrame. Additionally,
any associated metadata and mcols (metadata columns) are preserved and
added to the DFrame, if they exist.
realize(x, BACKEND = getAutoRealizationBackend()):Realize an object into memory or on disk using the equivalent of
realize(as(x, "DFrame"), BACKEND).
In the code snippets below, x is a DuckDBDataFrame object:
x[i, j, drop = TRUE]:Returns either a new DuckDBDataFrame object or a DuckDBColumn if
selecting a single column and drop = TRUE.
x[[i]]:Extracts a DuckDBColumn object from x.
x[[i]] <- value:Replaces column i in x with value.
head(x, n = 6L):If n is non-negative, returns the first n rows of x.
If n is negative, returns all but the last abs(n) rows of
x.
tail(x, n = 6L):If n is non-negative, returns the last n rows of x.
If n is negative, returns all but the first abs(n) rows of
x.
subset(x, subset, select, drop = FALSE):Return a new DuckDBDataFrame using:
logical expression indicating rows to keep, where missing values are taken as FALSE.
expression indicating columns to keep.
passed on to [ indexing operator.
cbind(...):Creates a new DuckDBDataFrame object by concatenating the columns of the
input objects.
The returned DuckDBDataFrame object concatenates the metadata across the
input objects.
The metadata columns of the returned DuckDBDataFrame object are obtained
by combining the metadata columns of the input object with
combineRows().
The show() method for DuckDBDataFrame objects obeys global options
showHeadLines and showTailLines for controlling the number of
head and tail rows and columns to display.
Patrick Aboyoun, Aaron Lun
DuckDBTable-class for the table base class
DuckDBTable-utils for the table base class utilities
DuckDBColumn-class for the extracted column class
DuckDBColumn-utils for the extracted column utilities
DuckDBTransposedDataFrame-class for the transposed
DataFrame class
DuckDBDataFrameList-class for the split DataFrame class
DataFrame for the base class
# Mocking up a file: tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) arrow::write_parquet(cbind(model = rownames(mtcars), mtcars), tf) # Creating our DuckDB-backed data frame: df <- DuckDBDataFrame(tf, datacols = colnames(mtcars), keycol = "model") df # Extraction yields a DuckDBColumn: df$carb # Slicing DuckDBDataFrame objects: df[,1:5] df[1:5,] # Combining by DuckDBDataFrame and DuckDBColumn objects: combined <- cbind(df, df) class(combined) combined2 <- cbind(df, some_new_name=df[,1]) class(combined2)# Mocking up a file: tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) arrow::write_parquet(cbind(model = rownames(mtcars), mtcars), tf) # Creating our DuckDB-backed data frame: df <- DuckDBDataFrame(tf, datacols = colnames(mtcars), keycol = "model") df # Extraction yields a DuckDBColumn: df$carb # Slicing DuckDBDataFrame objects: df[,1:5] df[1:5,] # Combining by DuckDBDataFrame and DuckDBColumn objects: combined <- cbind(df, df) class(combined) combined2 <- cbind(df, some_new_name=df[,1]) class(combined2)
The DuckDBDataFrame class extends SplitDataFrameList.
Objects of class DuckDBDataFrameList extend List.
split(x, f):Creates a DuckDBDataFrameList object.
xA DuckDBDataFrame object to split.
fA DuckDBColumn object to split x by.
In the code snippets below, x is a DuckDBDataFrameList object:
length(x):Get the number of elements in x.
names(x), names(x) <- value:Get or set the names of the elements of x.
mcols(x), mcols(x) <- value:Get or set the metadata columns.
elementNROWS(x):Get the length (or nb of row for a matrix-like object) of each of the elements.
dimtbls(x, drop = TRUE), dimtbls(x) <- value:Get or set the list of dimension tables used to define partitions for
efficient queries. If drop = TRUE, then it returns a named
DataFrameList object, else it returns an environment containing
a dimtbls named DataFrameList element.
dims(x):Get the two-column matrix indicating the number of rows and columns for each of the list elements.
nrows(x), ncols(x):Get the number of rows and columns, respectively, for each of the list elements.
dimnames(x):Get the list of two CharacterLists, the first holding the rownames and the second the column names.
rownames(x), colnames(x):Get the CharacterList of row and colum names, respectively.
commonColnames(x), commonColnames(x) <- value:Get or set the character vector of column names present in the individual
DataFrames in x.
columnMetadata(x), columnMetadata(x) <- value:Get the DataFrame of metadata along the columns, i.e.,
where each column in x is represented by a row in the metadata.
The metadata is common across all elements of x. Note that calling
mcols(x) returns the metadata on the DataFrame
elements of x.
In the code snippets below, x is a DuckDBDataFrameList object:
unlist(x):Returns the underlying DuckDBDataFrame object.
as(from, "DFrameList"):Converts a DuckDBDataFrameList object to a DFrameList object. This
process involves first loading the data into memory using
as(from, "DFrame"). The resulting DFrame is then split into a
DFrameList. Additionally, any associated metadata and mcols (metadata
columns) are preserved and added to the DFrameList, if they exist.
realize(x, BACKEND = getAutoRealizationBackend()):Realize an object into memory or on disk using the equivalent of
realize(as(x, "DFrameList"), BACKEND).
In the code snippets below, x is a DuckDBDataFrameList object:
x[i]:Returns a DuckDBDataFrameList object containing the selected elements.
x[[i]]:Return the selected DuckDBDataFrame by i, where i is an
numeric or character vector of length 1.
x$name:Similar to x[[name]], but name is taken literally as an
element name.
head(x, n = 6L):If n is non-negative, returns the first n elements of x.
If n is negative, returns all but the last abs(n) elements
of x.
tail(x, n = 6L):If n is non-negative, returns the last n elements of x.
If n is negative, returns all but the first abs(n) elements
of x.
Patrick Aboyoun
DuckDBDataFrame-class for the unsplit class
SplitDataFrameList for the base class
# Mocking up a file: tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) arrow::write_parquet(cbind(model = rownames(mtcars), mtcars), tf) # Creating our DuckDB-backed data frame: df <- DuckDBDataFrame(tf, datacols = colnames(mtcars), keycol = "model") dflist <- split(df, df$cyl) dflist# Mocking up a file: tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) arrow::write_parquet(cbind(model = rownames(mtcars), mtcars), tf) # Creating our DuckDB-backed data frame: df <- DuckDBDataFrame(tf, datacols = colnames(mtcars), keycol = "model") dflist <- split(df, df$cyl) dflist
The DuckDBEmbeddings class extends RectangularData and DuckDBColumn to represent a single dimensionality reduction embedding (e.g., PCA, UMAP, tSNE) stored in DuckDB using an ARRAY[] column.
DuckDBEmbeddings provides a DuckDB-backed container for storing a single embedding as a 2D rectangular object where rows represent cells and columns represent dimensions within that embedding. The embedding is stored as a single ARRAY[] column in the underlying DuckDB table, enabling efficient storage and lazy evaluation.
Objects of class DuckDBEmbeddings extend RectangularData.
In the code snippets below, x is a DuckDBEmbeddings object:
dim(x):Returns integer vector of length 2: c(nrow(x), ncol(x)) where
nrow is the number of cells and ncol is the ARRAY size (number of
dimensions in the embedding).
nrow(x), ncol(x):Get the number of rows (cells) and columns (dimensions), respectively.
NROW(x), NCOL(x):Same as nrow(x) and ncol(x), respectively.
dimnames(x):Returns list of length 2: list(rownames(x), colnames(x)).
rownames(x), colnames(x):Get the names of the rows (cell IDs) and columns (dimension names), respectively.
length(x):Get the number of cells (same as nrow(x)).
names(x):Get the cell IDs (same as rownames(x)).
type(x):Get the ARRAY data type. Returns enhanced format "array<double,50>"
or raw DuckDB format "DOUBLE[50]" depending on schema detection.
coltypes(x):Get the R element type (e.g., "numeric" for DOUBLE arrays,
"integer" for INTEGER arrays). Handles both type formats.
dimtbls(x, drop = TRUE), dimtbls(x) <- value:Get or set the list of dimension tables used to define partitions for
efficient queries. If drop = TRUE, then it returns a named
DataFrameList object, else it returns an environment containing
a dimtbls named DataFrameList element.
In the code snippets below, x is a DuckDBEmbeddings object:
x[i, j, drop=TRUE]:Subset cells (rows) and/or dimensions (columns). When both i
and j are missing, returns x. When drop=TRUE
and the result is a single cell or single dimension, returns a
vector. Otherwise returns a matrix or DuckDBEmbeddings object.
x[i]:Subset cells (rows), returns a DuckDBEmbeddings object with the selected cells.
head(x, n = 6L):Returns the first n cells of the embedding.
tail(x, n = 6L):Returns the last n cells of the embedding.
as.matrix(x):Coerces the embedding to a matrix with cells as rows and dimensions as columns.
realize(x, BACKEND = getAutoRealizationBackend()):Realize an object into memory or on disk using the equivalent of
realize(as.matrix(x), BACKEND).
Patrick Aboyoun
DuckDBNumericList-class for the base class
DuckDBTable-class for the underlying table representation
tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) n <- 10L mat <- matrix(rnorm(n * 5), nrow = n) df <- data.frame( cell_id = sprintf("cell_%02d", seq_len(n)), pca = I(asplit(mat, 1L)), stringsAsFactors = FALSE ) arrow::write_parquet(df, tf) emb <- DuckDBDataFrame(tf, keycol = "cell_id")[["pca"]] emb dim(emb) head(emb, 3)tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) n <- 10L mat <- matrix(rnorm(n * 5), nrow = n) df <- data.frame( cell_id = sprintf("cell_%02d", seq_len(n)), pca = I(asplit(mat, 1L)), stringsAsFactors = FALSE ) arrow::write_parquet(df, tf) emb <- DuckDBDataFrame(tf, keycol = "cell_id")[["pca"]] emb dim(emb) head(emb, 3)
The DuckDBSelfHits class extends the SelfHits-class class from S4Vectors for DuckDB tables.
DuckDBSelfHits provides lazy evaluation for large graph structures (e.g., = KNN graphs, similarity matrices) by storing edge lists in DuckDB-backed tables. This enables efficient storage and querying of pairwise relationships without materializing all edges in memory.
The class stores the parallel slots ('from', 'to', and any metadata columns) in a DuckDBDataFrame, while keeping the scalar 'nLnode' and 'nRnode' slots materialized for fast validation.
Objects of class DuckDBSelfHits extend SelfHits-class.
DuckDBSelfHits(conn, from, to, nnode, mcols = NULL,
keycol = NULL, dimtbl = NULL, nodes = NULL):Creates a DuckDBSelfHits object.
connEither a character vector containing the paths to parquet, csv, or
gzipped csv data files; a string that defines a duckdb read_*
data source; a DuckDBDataFrame object; or a tbl_duckdb_connection
object.
fromEither NULL or a string specifying the column from
conn that defines the query/left nodes (1-indexed).
toEither NULL or a string specifying the column from
conn that defines the subject/right nodes (1-indexed).
nnodeSingle integer specifying the number of nodes in the graph. Must be positive. This is the only slot that remains materialized.
mcolsOptional character vector specifying additional columns that define edge metadata (e.g., "weight", "distance", "x").
keycolAn optional string specifying the column name from conn that
will define the foreign key in the underlying table, or a named list
containing a character vector where the name of the list element
defines the foreign key and the character vector sets the distinct
values for that key. If missing, a row_number column is
created as an identifier for each edge.
dimtblAn optional named DataFrameList that specifies the dimension
table associated with the keycol. The name of the list
element must match the name of the keycol list. Additionally,
the DataFrame object must have row names that match the
distinct values of the keycol list element.
nodesAn optional integer vector specifying the node ids. If NULL
(default), uses implicit encoding c(NA_integer_, -nnode)
representing 1:nnode. Can be an explicit integer vector for
non-contiguous node subsets, optionally named for node aliases.
In the code snippets below, x is a DuckDBSelfHits object:
length(x):Get the number of edges (hits).
from(x), queryHits(x):Get the query/left node indices as a DuckDBColumn.
to(x), subjectHits(x):Get the subject/right node indices as a DuckDBColumn.
nnode(x), nLnode(x), nRnode(x):Get the number of nodes. All three return the same value for SelfHits.
queryLength(x), subjectLength(x):Get the number of left/right nodes. For SelfHits, both return
nnode(x).
countLnodeHits(x), countQueryHits(x):Count the number of hits for each left/query node. Returns an integer
vector of length nLnode(x). When nodes are implicit (1:nnode),
returns an unnamed vector with positional indexing. When nodes are
explicit, returns a named vector where names are the node IDs as
characters. This operation materializes the from column.
countRnodeHits(x), countSubjectHits(x):Count the number of hits for each right/subject node. Returns an integer
vector of length nRnode(x). When nodes are implicit (1:nnode),
returns an unnamed vector with positional indexing. When nodes are
explicit, returns a named vector where names are the node IDs as
characters. This operation materializes the to column.
mcols(x), mcols(x) <- value:Get or set the edge metadata columns.
as(from, "DuckDBDataFrame"):Creates a DuckDBDataFrame object with from, to, and mcols.
as.data.frame(x):Coerces x to a data.frame.
as(from, "SelfHits"):Converts a DuckDBSelfHits object to a materialized SelfHits object. This will load all edges into memory.
as(from, "dgCMatrix"):Converts to a sparse matrix. Equivalent to
as(as(from, "SelfHits"), "dgCMatrix").
realize(x, BACKEND = getAutoRealizationBackend()):Realize an object into memory or on disk using the equivalent of
realize(as(x, "SelfHits"), BACKEND).
In the code snippets below, x is a DuckDBSelfHits object:
x[i] or extractROWS(x, i):Returns a DuckDBSelfHits object with edge-based subsetting (subsets edges directly by index, does NOT filter by nodes).
extractNODES(x, i):Returns a DuckDBSelfHits object with node-based subsetting. Updates
the nodes slot to track the subset, and filters edges lazily
via SQL to include only edges where BOTH from and to are in the node
subset. Filtering is applied when data is materialized via
tblconn(), as.data.frame(), or coercion methods.
head(x, n = 6L):If n is non-negative, returns the first n edges.
If n is negative, returns all but the last abs(n) edges.
tail(x, n = 6L):If n is non-negative, returns the last n edges.
If n is negative, returns all but the first abs(n) edges.
sort(x, decreasing = FALSE):Returns a DuckDBSelfHits with edges sorted by from node, then by to node. The sorting is performed lazily using SQL ORDER BY.
t(x):Transpose the graph by swapping from and to columns.
c(x, ...):Concatenate multiple DuckDBSelfHits objects.
The show() method for DuckDBSelfHits objects obeys global options
showHeadLines and showTailLines for controlling the number of
edges to display.
Patrick Aboyoun
SelfHits-class for the base class
colPairs for usage in SingleCellExperiment
DuckDBDataFrame for the underlying storage
# Create an example edge list: edges <- data.frame( from = c(1, 1, 2, 3, 4, 4, 5), to = c(2, 5, 3, 4, 2, 5, 1), weight = c(0.8, 0.6, 0.9, 0.7, 0.85, 0.75, 0.65), distance = c(1.2, 2.3, 1.5, 1.8, 1.4, 2.1, 2.5) ) tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) arrow::write_parquet(edges, tf) # Create the DuckDBSelfHits object hits <- DuckDBSelfHits(tf, from = "from", to = "to", nnode = 5, mcols = c("weight", "distance")) hits # Access edge data from(hits) to(hits) mcols(hits) # Convert to sparse matrix (materializes) mat <- as(hits, "dgCMatrix")# Create an example edge list: edges <- data.frame( from = c(1, 1, 2, 3, 4, 4, 5), to = c(2, 5, 3, 4, 2, 5, 1), weight = c(0.8, 0.6, 0.9, 0.7, 0.85, 0.75, 0.65), distance = c(1.2, 2.3, 1.5, 1.8, 1.4, 2.1, 2.5) ) tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) arrow::write_parquet(edges, tf) # Create the DuckDBSelfHits object hits <- DuckDBSelfHits(tf, from = "from", to = "to", nnode = 5, mcols = c("weight", "distance")) hits # Access edge data from(hits) to(hits) mcols(hits) # Convert to sparse matrix (materializes) mat <- as(hits, "dgCMatrix")
The DuckDBTable class extends the RectangularData virtual class for DuckDB tables by wrapping a tbl_duckdb_connection object.
The DuckDBTable class provides a way to define a DuckDB table as a
RectangularData object. It supports standard 2D API
such as dim(), nrow(), ncol(), dimnames(),
x[i, j] and cbind(), but does not support rbind().
Objects of class DuckDBTable extend RectangularData.
DuckDBTable(conn, datacols = colnames(conn), keycols = NULL, dimtbls = NULL, type = NULL, collevels = NULL):Creates a DuckDBTable object.
connEither a character vector containing the paths to parquet, csv, or
gzipped csv data files; a string that defines a duckdb read_*
data source; a DuckDBDataFrame object; or a tbl_duckdb_connection
object.
datacolsEither a character vector of column names from conn or a
named expression that will be evaluated in the context of
conn that defines the data.
keycolsAn optional character vector of column names from conn that
will define the set of foreign keys in the underlying table, or a
named list of character vectors where the names of the list define
the foreign keys and the character vectors set the distinct values
for those keys. If missing, a row_number column is created
as an identifier.
dimtblsEither NULL, a named DataFrameList object, or an environment
containing a named DataFrameList "dimtbls" element that
specifies the dimension tables associated with the keycols.
The name of the list elements match the names of the keycols
list. Additionally, the DataFrame objects have row names that
match the distinct values of the corresponding keycols list
element and columns that define partitions in the data table for
efficient querying.
typeAn optional named character vector where the names specify the
column names and the values specify the column type; one of
"logical", "integer", "integer64",
"double", or "character".
collevelsAn optional named list, keyed by data column name, that restores
factor columns on materialization. Each element is a list with
a character levels vector and a TRUE/FALSE
ordered flag; columns not named are left unchanged. Primarily
used by readParquet() to recover factors recorded in a
product's schema.
In the code snippets below, x is a DuckDBTable object:
dim(x):Length two integer vector defined as c(nrow(x), ncol(x)).
nrow(x), ncol(x):Get the number of rows and columns, respectively.
NROW(x), NCOL(x):Same as nrow(x) and ncol(x), respectively.
dimnames(x):Length two list of character vectors defined as
list(rownames(x), colnames(x)).
rownames(x), colnames(x):Get the names of the rows and columns, respectively.
coltypes(x), coltypes(x) <- value:Get or set the data type of the columns. Getter returns one of
"logical", "integer", "integer64", "double",
"character", or complex type strings like "list<integer>",
"array<numeric,128>", "struct", "map", "blob",
or "geometry". Setter accepts atomic types only.
dimtbls(x, drop = TRUE), dimtbls(x) <- value:Get or set the list of dimension tables used to define partitions for
efficient queries. If drop = TRUE, then it returns a named
DataFrameList object, else it returns an environment containing
a dimtbls named DataFrameList element.
as.data.frame(x):Coerces x to a data.frame.
In the code snippets below, x is a DuckDBTable object:
x[i, j]:Return a new DuckDBTable of the same class as x made of the
selected rows and columns.
head(x, n = 6L):If n is non-negative, returns the first n rows of x.
If n is negative, returns all but the last abs(n) rows of
x.
tail(x, n = 6L):If n is non-negative, returns the last n rows of x.
If n is negative, returns all but the first abs(n) rows of
x.
Patrick Aboyoun
DuckDBTable-utils for the utilities
RectangularData for the base class
# Create a data.frame from the Titanic data df <- do.call(expand.grid, c(dimnames(Titanic), stringsAsFactors = FALSE)) df$fate <- as.integer(Titanic[as.matrix(df)]) # Write data to a parquet file tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) arrow::write_parquet(df, tf) tbl <- DuckDBTable(tf, datacols = "fate", keycols = c("Class", "Sex", "Age", "Survived"))# Create a data.frame from the Titanic data df <- do.call(expand.grid, c(dimnames(Titanic), stringsAsFactors = FALSE)) df$fate <- as.integer(Titanic[as.matrix(df)]) # Write data to a parquet file tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) arrow::write_parquet(df, tf) tbl <- DuckDBTable(tf, datacols = "fate", keycols = c("Class", "Sex", "Age", "Survived"))
Common operations on DuckDBTable objects.
Method return types are documented in the sections above.
In the code snippets below, x is a DuckDBTable object:
sql_call(x, fun, ...):Applies the specified SQL function to all data columns of x.
sql_fun(x, function_type = NULL, return_type = NULL, description = FALSE):Finds DuckDB SQL functions compatible with the data type of x.
Returns a data.frame of matching functions with optional filtering by
function type and return type.
DuckDBTable objects have support for S4 group generic functionality:
Arith"+", "-", "*", "^",
"%%", "%/%", "/"
Compare"==", ">", "<", "!=",
"<=", ">="
Logic"&", "|", "!"
Ops"Arith", "Compare", "Logic"
Math"abs", "sign", "sqrt",
"ceiling", "floor", "trunc", "log",
"log10", "log2", "acos", "acosh",
"asin", "asinh", "atan", "atanh",
"exp", "expm1", "cos", "cosh",
"sin", "sinh", "tan", "tanh",
"gamma", "lgamma"
Summary"max", "min", "range",
"prod", "sum", "any", "all"
See S4groupGeneric for more details.
In the code snippets below, x is a DuckDBTable object:
is.finite(x):Returns a DuckDBTable containing logicals that indicate which values are finite.
is.infinite(x):Returns a DuckDBTable containing logicals that indicate which values are infinite.
is.nan(x):Returns a DuckDBTable containing logicals that indicate which values are Not a Number.
mean(x):Calculates the mean of x.
var(x):Calculates the variance of x.
sd(x):Calculates the standard deviation of x.
median(x):Calculates the median of x.
quantile(x, probs = seq(0, 1, 0.25), names = TRUE, type = 7):Calculates the specified quantiles of x.
probsA numeric vector of probabilities with values in [0,1].
namesIf TRUE, the result has names describing the
quantiles.
typeEither 1 or 7 that specifies the quantile algorithm
detailed in quantile.
mad(x, constant = 1.4826):Calculates the median absolute deviation of x.
constantThe scale factor.
IQR(x, type = 7):Calculates the interquartile range of x.
typeEither 1 or 7 that specifies the quantile algorithm
detailed in quantile.
pmax(..., na.rm = FALSE):Returns the parallel maxima of multiple DuckDBTable objects. All arguments must be DuckDBTable objects with compatible dimensions.
pmin(..., na.rm = FALSE):Returns the parallel minima of multiple DuckDBTable objects. All arguments must be DuckDBTable objects with compatible dimensions.
In the code snippets below, x is a DuckDBTable object:
nchar(x):Returns a DuckDBTable containing the number of characters in each string.
tolower(x):Returns a DuckDBTable with all strings converted to lowercase.
toupper(x):Returns a DuckDBTable with all strings converted to uppercase.
chartr(old, new, x):Returns a DuckDBTable with characters translated.
oldCharacters to be translated.
newCharacters to translate to.
substr(x, start, stop):Returns a DuckDBTable containing substrings extracted by position.
startInteger starting position (1-indexed).
stopInteger ending position (inclusive).
substring(x, first, last = 1000000L):Returns a DuckDBTable containing substrings extracted by position.
firstInteger starting position (1-indexed).
lastInteger ending position (inclusive).
grepl(pattern, x, ignore.case = FALSE, fixed = FALSE):Returns a DuckDBTable containing logicals indicating pattern matches.
patternCharacter string containing a regular expression.
ignore.caseIf TRUE, case-insensitive matching.
fixedIf TRUE, pattern is a fixed string not regex.
sub(pattern, replacement, x, ignore.case = FALSE, fixed = FALSE):Returns a DuckDBTable with first match of pattern replaced.
patternCharacter string containing a regular expression.
replacementReplacement string.
ignore.caseIf TRUE, case-insensitive matching.
fixedIf TRUE, pattern is a fixed string not regex.
gsub(pattern, replacement, x, ignore.case = FALSE, fixed = FALSE):Returns a DuckDBTable with all matches of pattern replaced.
patternCharacter string containing a regular expression.
replacementReplacement string.
ignore.caseIf TRUE, case-insensitive matching.
fixedIf TRUE, pattern is a fixed string not regex.
startsWith(x, prefix):Returns a DuckDBTable containing logicals indicating if strings start with the specified prefix.
endsWith(x, suffix):Returns a DuckDBTable containing logicals indicating if strings end with the specified suffix.
paste(..., sep = " ", collapse = NULL):Concatenates multiple DuckDBTable objects column-wise using the
specified separator. The collapse argument is not supported.
All arguments must be DuckDBTable objects with compatible dimensions.
In the code snippets below, x is a DuckDBTable object:
unique(x):Returns a DuckDBTable containing the distinct rows.
x %in% table:Returns a DuckDBTable containing logicals that indicate if the
values in each of the columns of x are in table.
table(...):Returns a table containing the counts across the distinct values.
In the code snippets below, x is a DuckDBTable object:
elementNROWS(x):Returns a DuckDBTable containing the number of elements in each list
column of x.
In the code snippets below, x is a DuckDBTable object:
is_nonzero(x):Returns a DuckDBTable containing logicals that indicate if the
values in each of the columns of x are non-zero.
nzcount(x):Returns the total number of non-zero values.
is_sparse(x):Returns TRUE since data are stored in a sparse array representation.
Patrick Aboyoun
DuckDBTable-class for the main class
RectangularData for the base class
tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) arrow::write_parquet(cbind(model = rownames(mtcars), mtcars), tf) tbl <- DuckDBTable(tf, datacols = colnames(mtcars), keycol = "model") mean(tbl[, "mpg"]) head(tbl, 3)tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) arrow::write_parquet(cbind(model = rownames(mtcars), mtcars), tf) tbl <- DuckDBTable(tf, datacols = colnames(mtcars), keycol = "model") mean(tbl[, "mpg"]) head(tbl, 3)
The DuckDBTransposedDataFrame class extends TransposedDataFrame to represent a DuckDB table as a TransposedDataFrame object.
DuckDBTransposedDataFrame objects are constructed by calling t() on a
DuckDBDataFrame object.
Objects of class DuckDBTransposedDataFrame extend
RectangularData.
t(x):Creates a DuckDBTransposedDataFrame object from a DuckDBDataFrame object.
In the code snippets below, x is a DuckDBTransposedDataFrame object:
dim(x):Length two integer vector defined as c(nrow(x), ncol(x)).
nrow(x), ncol(x):Get the number of rows and columns, respectively.
NROW(x), NCOL(x):Same as nrow(x) and ncol(x), respectively.
dimnames(x):Length two list of character vectors defined as
list(rownames(x), colnames(x)).
rownames(x), colnames(x):Get the names of the rows and columns, respectively.
In the code snippets below, x is a DuckDBTransposedDataFrame object:
x[i, j, drop = TRUE]:Returns either a new DuckDBTransposedDataFrame object or a DuckDBColumn
if selecting a single row and drop = TRUE.
The show() method for DuckDBTransposedDataFrame objects obeys global
options showHeadLines and showTailLines for controlling the
number of head and tail rows to display.
Patrick Aboyoun
DuckDBDataFrame-class for the DataFrame class
TransposedDataFrame for the base class
# Mocking up a file: tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) arrow::write_parquet(cbind(model = rownames(mtcars), mtcars), tf) # Creating our DuckDB-backed data frame: tdf <- t(DuckDBDataFrame(tf, datacols = colnames(mtcars), keycol = "model")) tdf # Slicing DuckDBTransposedDataFrame objects: tdf[,1:5] tdf[1:5,]# Mocking up a file: tf <- tempfile(fileext = ".parquet") on.exit(unlink(tf)) arrow::write_parquet(cbind(model = rownames(mtcars), mtcars), tf) # Creating our DuckDB-backed data frame: tdf <- t(DuckDBDataFrame(tf, datacols = colnames(mtcars), keycol = "model")) tdf # Slicing DuckDBTransposedDataFrame objects: tdf[,1:5] tdf[1:5,]
Get the keycols names, keycols dimension names, or keycols count of an object.
keynames(x) keydimnames(x, value) keydimnames(x) <- value nkey(x) nkeydim(x)keynames(x) keydimnames(x, value) keydimnames(x) <- value nkey(x) nkeydim(x)
x |
An object to get the keycols related information. |
value |
A character vector of keycols dimension names. |
keynames and keydimnames return character vectors.
nkey and nkeydim return a single integer.
Patrick Aboyoun
keynames showMethods("keynames") keydimnames showMethods("keydimnames") nkey showMethods("nkey")keynames showMethods("keynames") keydimnames showMethods("keydimnames") nkey showMethods("nkey")
Clustering keys for writeParquet's
cluster_by argument, plus the in-memory reorder they imply. Writing
with a clustering key physically orders rows so each Parquet row group covers
a tight region, letting DuckDB's row-group zonemaps prune groups outside a
range query.
zorder(cols, bits = 16L, by = NULL) hilbert(cols, bits = 16L, by = NULL) clusterSort(df, cluster_by)zorder(cols, bits = 16L, by = NULL) hilbert(cols, bits = 16L, by = NULL) clusterSort(df, cluster_by)
cols |
Character vector of numeric column names to cluster by.
|
bits |
Grid resolution per axis (default 16; a |
by |
Optional character vector of columns ordered lexicographically
before the curve (a composite categorical-prefix key); |
df |
A |
cluster_by |
A |
zorder() orders by a Morton (Z-order) code over cols (any
number of numeric columns), lowered SQL-side to a generated bit-interleave
ORDER BY (no extension needed). hilbert() orders by the native
DuckDB spatial ST_Hilbert curve (exactly two numeric columns,
better locality, requires the spatial extension). A plain character vector is
a lexicographic ordering (best for the leading column). clusterSort()
is the in-memory counterpart used by the materializing data.frame /
DataFrame write path (Hilbert falls back to Morton there). Its
double-precision Morton code is exact only for
length(cols) * bits <= 52; the lazy SQL path (unsigned 64-bit) is
exact to 62, so for a total above 52 bits the two write paths may order rows
differently. Keep length(cols) * bits at or below 52 when both paths
must agree byte-for-byte.
by adds a composite key: the named (typically low-cardinality
categorical) columns are ordered lexicographically before the
space-filling curve, i.e. ORDER BY by..., curve(cols). This keeps each
by-group's rows contiguous so the group's own zonemaps / run-length
encoding stay tight (e.g. zorder(c("x","y"), by = "gene") for a
viewport+gene workload, or zorder(c("x","y"), by = "__sample__") to
keep a COO assay's sample-slice pruning). The prefix columns are not
interleaved into the Morton code, so they cost none of the bits
budget.
zorder, hilbert
A DuckDBClusterSpec for
writeParquet(cluster_by=).
clusterSortdf reordered by the clustering key (same
class, same columns); a no-op when cluster_by is NULL,
df is empty, or a key column is absent.
zorder(c("x", "y")) hilbert(c("x", "y"), bits = 12L) clusterSort(data.frame(x = c(9, 1, 5), y = c(9, 1, 5)), zorder(c("x", "y")))zorder(c("x", "y")) hilbert(c("x", "y"), bits = 12L) clusterSort(data.frame(x = c(9, 1, 5), y = c(9, 1, 5)), zorder(c("x", "y")))
Apply any DuckDB SQL function to a BiocDuckDB object. This is a low-level interface that provides direct access to DuckDB's extensive SQL function library. The function is applied lazily and only executed when the result is materialized.
sql_call(x, fun, ...)sql_call(x, fun, ...)
x |
A BiocDuckDB object (DuckDBTable, DuckDBColumn, DuckDBDataFrame, or DuckDBAtomicList). |
fun |
A character string naming the SQL function to call. |
... |
Additional arguments passed to the SQL function. Arguments are coerced to SQL literals or column references as appropriate. |
This function applies the specified SQL function to all data columns in the
object. For common operations, consider using dedicated methods or
convenience wrappers (e.g., sort, unique,
%in%).
Available SQL functions are documented in DuckDB's function reference: https://duckdb.org/docs/sql/functions/overview
An object of the same class as x with the SQL function applied to
all data columns. The operation is lazy and only executed when the result
is materialized via as.vector(), as.list(), or similar.
Arguments in ... are converted to SQL expressions as follows:
Character: SQL string literal (automatically quoted)
Numeric: SQL numeric literal
Logical: SQL boolean literal (TRUE/FALSE)
NULL: SQL NULL
name/symbol: SQL column reference (unquoted, from as.name())
sql(): Raw SQL expression (from dplyr::sql(), not quoted)
Patrick Aboyoun
sql_call showMethods("sql_call")sql_call showMethods("sql_call")
Query DuckDB's function catalog to find SQL functions that are compatible with a BiocDuckDB object's data type. This is useful for discovering what operations can be performed on a particular column type.
sql_fun(x, function_type = NULL, return_type = NULL, description = FALSE)sql_fun(x, function_type = NULL, return_type = NULL, description = FALSE)
x |
A BiocDuckDB object (DuckDBTable, DuckDBColumn, DuckDBDataFrame, or DuckDBAtomicList). |
function_type |
Optional character vector of function types to filter for. If NULL (default), returns all function types. Options: "scalar", "aggregate", "macro", "table", "pragma", "table_macro" |
return_type |
Optional character vector of return types to filter for. If NULL (default), returns all functions. Common values: "BOOLEAN", "INTEGER", "DOUBLE", "VARCHAR", "ANY", "INTEGER[]", "DOUBLE[]", "VARCHAR[]", "ANY[]", etc. |
description |
Logical; if TRUE, include function descriptions in output. Default is FALSE. |
This function queries DuckDB directly via DESCRIBE to get the column's exact type (e.g., INTEGER[], DOUBLE[], VARCHAR[]), then searches duckdb_functions() for functions that accept that type as the first parameter. It matches both the exact type and generic types (T[], ANY[], ANY) that work with any type.
The results are grouped by function name with distinct return types collected into a list for functions that have multiple overloads.
A data.frame with the following columns:
function_name: Name of the function
alias_of: If this is an alias, the canonical function name
function_type: Type of function (scalar, aggregate, etc.)
return_type: List of distinct return types for this function
description: (if description = TRUE) Function description
Patrick Aboyoun
sql_fun showMethods("sql_fun")sql_fun showMethods("sql_fun")
Get a database table connection object
tblconn(x, select = TRUE, filter = TRUE)tblconn(x, select = TRUE, filter = TRUE)
x |
An object with a database table connection. |
select |
A logical value indicating whether to apply column selections to the database table connection. |
filter |
A logical value indicating whether to apply key column filtering to the database table connection. |
returns a database table connection object.
Patrick Aboyoun
tblconn showMethods("tableconn")tblconn showMethods("tableconn")