Eigen-Contrib  5.0.1
 
Loading...
Searching...
No Matches
Eigen::gpu::Context Class Reference

#include <contrib/Eigen/src/GPU/GpuContext.h>

Detailed Description

Unified GPU execution context owning a CUDA stream and library handles.

Each Context creates a dedicated CUDA stream and an eager cuBLAS handle bound to it; multiple contexts run concurrently on independent streams. The cuSOLVER, cuBLASLt, and cuSPARSE handles are created on first use, so a translation unit that never touches cuSOLVER (the cuFFT test, say) does not need it at link time. threadLocal() supplies a lazily-created default for simple single-stream usage.

A Context is not thread-safe: the library handles are not thread-safe per handle and the lazy initialization above is racy, so use one per thread or synchronize externally.

Public Member Functions

 Context ()
 
 Context (cudaStream_t stream)
 
cublasLtHandle_t cublasLtHandle ()
 
std::size_t cublasLtMaxWorkspaceBytes () const
 
cusolverDnHandle_t cusolverHandle ()
 
cusparseHandle_t cusparseHandle ()
 
internal::CublasLtPlanCache & gemmPlanCache ()
 
internal::DeviceBuffer & gemmWorkspace ()
 
const NppStreamContext & nppStreamContext () const
 
internal::OneShotSolverScratch & oneshotSolverScratch ()
 
void setCublasLtMaxWorkspaceBytes (std::size_t bytes)
 

Static Public Member Functions

static void setThreadLocal (Context *ctx)
 
static Context & threadLocal ()
 

Constructor & Destructor Documentation

◆ Context() [1/2]

Eigen::gpu::Context::Context ( )
inline

Create a new context with a dedicated CUDA stream.

◆ Context() [2/2]

Eigen::gpu::Context::Context ( cudaStream_t stream)
inlineexplicit

Create a context on an existing stream (e.g., stream 0 = nullptr). The caller retains ownership of the stream — this context will not destroy it.

Member Function Documentation

◆ cublasLtHandle()

cublasLtHandle_t Eigen::gpu::Context::cublasLtHandle ( )
inline

cuBLASLt handle (lazy-initialized on first GEMM call).

◆ cublasLtMaxWorkspaceBytes()

std::size_t Eigen::gpu::Context::cublasLtMaxWorkspaceBytes ( ) const
inline

Workspace ceiling passed to the cublasLtMatmul heuristic at plan-creation time. Defaults to internal::kCublasLtMaxWorkspaceBytes (compile-time configurable via EIGEN_CUDA_CUBLASLT_MAX_WORKSPACE_BYTES).

◆ cusolverHandle()

cusolverDnHandle_t Eigen::gpu::Context::cusolverHandle ( )
inline

Returns the cuSOLVER handle, creating it on first call.

◆ cusparseHandle()

cusparseHandle_t Eigen::gpu::Context::cusparseHandle ( )
inline

cuSPARSE handle, created on first use.

◆ gemmPlanCache()

internal::CublasLtPlanCache & Eigen::gpu::Context::gemmPlanCache ( )
inline

Plan cache for cublasLtMatmul (caches descriptors and selected algorithm by shape to avoid per-call overhead). Same thread-safety as workspace.

◆ gemmWorkspace()

internal::DeviceBuffer & Eigen::gpu::Context::gemmWorkspace ( )
inline

Workspace buffer for cublasLtMatmul (grown lazily by cublaslt_gemm). Not thread-safe — all GEMM calls must be on this context's stream.

◆ nppStreamContext()

const NppStreamContext & Eigen::gpu::Context::nppStreamContext ( ) const
inline

NPP stream context for stream(), filled in at construction. Filling it queries the stream, which fails while the stream is being captured on some drivers, so NPP calls inside a capture must use this one.

◆ oneshotSolverScratch()

internal::OneShotSolverScratch & Eigen::gpu::Context::oneshotSolverScratch ( )
inline

Grow-only scratch for the one-shot solver expressions (d_A.llt().solve(d_B), d_A.lu().solve(d_B)). Same thread-safety rules as the GEMM workspace: all uses must be on this context's stream.

◆ setCublasLtMaxWorkspaceBytes()

void Eigen::gpu::Context::setCublasLtMaxWorkspaceBytes ( std::size_t bytes)
inline

Override the workspace ceiling for future plan-cache misses on this context. The cap is consulted at plan-creation time only; pre-existing cached plans keep the cap they were built with. Call gemmPlanCache().clear() to force re-selection under the new cap.

◆ setThreadLocal()

static void Eigen::gpu::Context::setThreadLocal ( Context * ctx)
inlinestatic

Override the thread-local default context for this thread. The caller retains ownership of ctx — it must outlive all uses. Pass nullptr to restore the lazily-created default.

◆ threadLocal()

static Context & Eigen::gpu::Context::threadLocal ( )
inlinestatic

Get the thread-local default context.

If setThreadLocal() has been called, returns that context. Otherwise lazily creates a new context with a dedicated stream.

Note
The thread-local instance is destroyed when the thread exits (or at static destruction time for the main thread). On some CUDA driver configurations this may print "CUDA_ERROR_DEINITIALIZED" to stderr if the CUDA context has already been torn down. These errors are harmless and are suppressed in the destructor, but they can produce noise in test output. To avoid this, call cudaDeviceReset() only after all Context instances (including thread-local ones) have been destroyed — or create and own a Context and install it with setThreadLocal(): the lazily-created default is then never constructed, and teardown order is fully under application control.

The documentation for this class was generated from the following file: