![]() |
Eigen-Contrib
5.0.1
|
#include <contrib/Eigen/src/GPU/GpuContext.h>
Unified GPU execution context owning a CUDA stream and library handles.
Each Context creates a dedicated CUDA stream and an eager cuBLAS handle bound to it; multiple contexts run concurrently on independent streams. The cuSOLVER, cuBLASLt, and cuSPARSE handles are created on first use, so a translation unit that never touches cuSOLVER (the cuFFT test, say) does not need it at link time. threadLocal() supplies a lazily-created default for simple single-stream usage.
A Context is not thread-safe: the library handles are not thread-safe per handle and the lazy initialization above is racy, so use one per thread or synchronize externally.
Public Member Functions | |
| Context () | |
| Context (cudaStream_t stream) | |
| cublasLtHandle_t | cublasLtHandle () |
| std::size_t | cublasLtMaxWorkspaceBytes () const |
| cusolverDnHandle_t | cusolverHandle () |
| cusparseHandle_t | cusparseHandle () |
| internal::CublasLtPlanCache & | gemmPlanCache () |
| internal::DeviceBuffer & | gemmWorkspace () |
| const NppStreamContext & | nppStreamContext () const |
| internal::OneShotSolverScratch & | oneshotSolverScratch () |
| void | setCublasLtMaxWorkspaceBytes (std::size_t bytes) |
Static Public Member Functions | |
| static void | setThreadLocal (Context *ctx) |
| static Context & | threadLocal () |
|
inline |
Create a new context with a dedicated CUDA stream.
|
inlineexplicit |
Create a context on an existing stream (e.g., stream 0 = nullptr). The caller retains ownership of the stream — this context will not destroy it.
|
inline |
cuBLASLt handle (lazy-initialized on first GEMM call).
|
inline |
Workspace ceiling passed to the cublasLtMatmul heuristic at plan-creation time. Defaults to internal::kCublasLtMaxWorkspaceBytes (compile-time configurable via EIGEN_CUDA_CUBLASLT_MAX_WORKSPACE_BYTES).
|
inline |
Returns the cuSOLVER handle, creating it on first call.
|
inline |
cuSPARSE handle, created on first use.
|
inline |
Plan cache for cublasLtMatmul (caches descriptors and selected algorithm by shape to avoid per-call overhead). Same thread-safety as workspace.
|
inline |
Workspace buffer for cublasLtMatmul (grown lazily by cublaslt_gemm). Not thread-safe — all GEMM calls must be on this context's stream.
|
inline |
NPP stream context for stream(), filled in at construction. Filling it queries the stream, which fails while the stream is being captured on some drivers, so NPP calls inside a capture must use this one.
|
inline |
Grow-only scratch for the one-shot solver expressions (d_A.llt().solve(d_B), d_A.lu().solve(d_B)). Same thread-safety rules as the GEMM workspace: all uses must be on this context's stream.
|
inline |
Override the workspace ceiling for future plan-cache misses on this context. The cap is consulted at plan-creation time only; pre-existing cached plans keep the cap they were built with. Call gemmPlanCache().clear() to force re-selection under the new cap.
|
inlinestatic |
Override the thread-local default context for this thread. The caller retains ownership of ctx — it must outlive all uses. Pass nullptr to restore the lazily-created default.
|
inlinestatic |
Get the thread-local default context.
If setThreadLocal() has been called, returns that context. Otherwise lazily creates a new context with a dedicated stream.