A GPU is most useful when data are not repeatedly copied between R and the GPU. This guide shows the simple pattern: upload once, run several CUDA tasks, and bring back only the result you need.
Use a strict selection before a long analysis:
If CUDA is not ready, this call stops and explains what is missing. It never turns a requested GPU workflow into an unnoticed non-GPU run.
Construct a cudatensor once, then apply arithmetic,
reshape, transpose, subsetting, reduction, and matrix multiplication
without reconstructing the input for every call.
set.seed(1)
x <- matrix(rnorm(10000 * 100), nrow = 10000)
x_gpu <- cuda_tensor(x, device = "cuda", dtype = "float32")
means_gpu <- tensor_mean(x_gpu, dim = 1)
centered_gpu <- x_gpu - means_gpu
selected_gpu <- centered_gpu[, 1:50, drop = FALSE]
gram_gpu <- tensor_matmul(t(selected_gpu), selected_gpu)
tensor_device(gram_gpu)
cuda_provenance(gram_gpu)Supported reshape and selection operations stay on the GPU, so they can feed the next calculation without downloading the full tensor.
Only cross the host boundary when an ordinary R object is required:
Calling to_cpu() inside each stage would add transfers
and erase much of the reason to use the GPU.
The native PCA object retains device storage for its scores, so exact
kNN can consume pca$x directly.
pca <- cuda_pca(
x_gpu,
n_components = 20,
device = "cuda"
)
neighbors <- cuda_knn(
pca$x,
k = 15,
batch_size = 256,
device = "cuda"
)
cuda_provenance(pca)
cuda_provenance(neighbors)
head(neighbors$index)The public kNN result contains ordinary R index and distance matrices because those are normally the final analysis output. The expensive PCA, distance blocks, and stable top-k stages remain CUDA stages.
For sparse data, upload a Matrix object once and
continue with normalization, PCA, and exact kNN:
counts <- Matrix::rsparsematrix(50000, 128, density = 0.01)
counts@x <- abs(counts@x)
counts_gpu <- cuda_sparse(counts, device = "cuda")
normalized_gpu <- sparse_normalize(
counts_gpu,
margin = "rows",
scale_factor = 10000,
log1p = TRUE
)
sparse_pca <- cuda_pca(normalized_gpu, n_components = 20, device = "cuda")
sparse_neighbors <- cuda_knn(
sparse_pca$x,
k = 15,
device = "cuda"
)
sparse_info(normalized_gpu)
cuda_provenance(sparse_neighbors)Use to_dgCMatrix() only when a downstream R package
specifically needs a host dgCMatrix.
cuda_provenance() shows where each part of a calculation
ran. The most useful fields are:
| Field | Question answered |
|---|---|
operation |
What did this stage do? |
device |
Where did the computation run? |
backend |
Which cudaverse implementation performed it? |
output_device |
Where was the result left for the next step? |
Multi-stage algorithms retain a stage list. Inspect it directly when checking residency:
A strict device = "cuda" request cannot silently fall
back. A hybrid high-level workflow can still contain a deliberately
documented host stage; that stage has its own record rather than being
described as GPU-resident.
float32 when its numerical precision is
sufficient; it halves the value storage compared with
float64.cudatensor and cudasparse objects
across operations.cuda_knn() instead of cuda_distance()
when a full dense distance matrix is not needed.batch_size when exact distance or kNN blocks
approach the available device memory.The lightweight path does not require R torch. A typical
native result should show CUDA and native in its stage record:
stages <- cuda_provenance(neighbors)$stages
vapply(stages, `[[`, character(1), "device")
vapply(stages, `[[`, character(1), "backend")Some records describe preparation or returning the final R object rather than a numerical calculation. The operation coverage guide explains what to expect for each task.