Use this page to decide whether cudaverse can accelerate the part of your analysis that matters. “Yes” means the numerical work runs with CUDA. “Partly” means cudaverse uses CUDA for some steps and an established R package for the rest.
| Task | Runs with CUDA? | What you receive in R |
|---|---|---|
| Dense tensor operations | Yes | A GPU tensor until you call to_cpu() |
| Matrix multiplication and summaries | Yes | A GPU tensor |
| SVD and PCA | Yes | An R model that can reuse GPU results |
| Pairwise distances | Yes | A regular R distance matrix |
| Exact kNN | Yes | Neighbour indices and distances |
| k-means | Yes | An R clustering result |
| Sparse storage, normalization, and multiplication | Yes | A GPU sparse object |
| Sparse PCA followed by exact kNN | Yes | Final neighbour results |
| kNN graph construction | Partly | A sparse Matrix graph |
| Louvain and Leiden | Partly | Community labels |
| UMAP and t-SNE | Partly | Embedding coordinates |
| Diffusion map | Partly | Embedding coordinates |
Use device = "cuda" for tasks that must run on an
eligible CUDA backend and inspect cuda_provenance() for the
actual stage boundary.
library(cudaverse)
cuda_select_device("cuda")
set.seed(1)
x <- matrix(rnorm(10000 * 100), nrow = 10000)
x_gpu <- cuda_tensor(x, device = "cuda", dtype = "float32")
product_gpu <- tensor_matmul(t(x_gpu), x_gpu)
reduced_gpu <- tensor_sum(product_gpu, dim = 1)
pca <- cuda_pca(x_gpu, n_components = 20, device = "cuda")
neighbors <- cuda_knn(pca$x, k = 15, device = "cuda")
tensor_device(reduced_gpu)
cuda_provenance(neighbors)The tensor operations and PCA scores stay on the GPU. Only the final neighbour matrices are returned as ordinary R objects.
counts <- Matrix::rsparsematrix(10000, 100, density = 0.03)
counts@x <- abs(counts@x)
counts_gpu <- cuda_sparse(counts, device = "cuda")
normalized_gpu <- sparse_normalize(
counts_gpu,
margin = "rows",
scale_factor = 10000,
log1p = TRUE
)
feature_totals_gpu <- sparse_row_sums(t(normalized_gpu))
sparse_pca <- cuda_pca(normalized_gpu, n_components = 20, device = "cuda")
sparse_neighbors <- cuda_knn(
sparse_pca$x,
k = 15,
device = "cuda"
)
sparse_info(normalized_gpu)
cuda_provenance(sparse_neighbors)GPU acceleration of PCA and exact kNN can still reduce the cost before graph or embedding stages, but the current graph clustering and UMAP/t-SNE implementations are not described as native CUDA.
pca <- cuda_pca(x, n_components = 20, device = "cuda")
neighbors <- cuda_knn(pca$x, k = 15, device = "cuda")
graph <- cuda_knn_graph(neighbors)
communities <- cuda_leiden(graph)
embedding <- cuda_umap(pca$x)
cuda_provenance(communities)
cuda_provenance(embedding)The provenance record shows the CUDA preprocessing stages and the intentional host stages separately. This makes it possible to use the accelerated portion without presenting the complete workflow as GPU-resident.
The native backend loads the NVIDIA libraries already installed on your computer and does not require LibTorch. The selected backend is always visible in diagnostics and provenance.
An explicit CUDA request is strict: if CUDA is not ready, cudaverse stops and shows diagnostic guidance instead of silently changing where the work runs.