![]() |
Eigen-Contrib
5.0.1
|
#include <contrib/Eigen/src/GPU/GpuLLT.h>
GPU Cholesky (LL^T) decomposition via cuSOLVER.
| Scalar_ | Element type: float, double, complex<float>, complex<double> |
| UpLo_ | Triangle used: Lower (default) or Upper |
Factorizes a symmetric positive-definite matrix A = LL^H on the GPU and caches the factor L in device memory. Each subsequent solve(B) uploads only B, calls cusolverDnXpotrs, and downloads the result — the factor is not re-transferred.
Each LLT object owns a dedicated CUDA stream and cuSOLVER handle, enabling concurrent factorizations from multiple objects on the same host thread.
Public Member Functions | |
| template<typename InputType> | |
| LLT & | compute (const DenseBase< InputType > &A) |
| LLT & | compute (const DeviceMatrix< Scalar > &d_A) |
| LLT & | compute (DeviceMatrix< Scalar > &&d_A) |
| LLT ()=default | |
| template<typename InputType> | |
| LLT (const DenseBase< InputType > &A) | |
| LLT (const DeviceMatrix< Scalar > &d_A) | |
| LLT (Context &ctx) | |
| template<typename InputType> | |
| LLT (Context &ctx, const DenseBase< InputType > &A) | |
| LLT (Context &ctx, const DeviceMatrix< Scalar > &d_A) | |
| LLT (Context &ctx, DeviceMatrix< Scalar > &&d_A) | |
| LLT (DeviceMatrix< Scalar > &&d_A) | |
| DeviceMatrix< Scalar > | solve (const DeviceMatrix< Scalar > &d_B) const |
| template<typename Rhs> | |
| PlainMatrix | solve (const MatrixBase< Rhs > &B) const |
| DeviceMatrix< Scalar > | solve (DeviceMatrix< Scalar > &&d_B) const |
|
default |
|
inlineexplicit |
Bind to ctx: run on its stream with its cuSOLVER/cuBLAS handles, so solver work chains with other work on the same Context without cross-stream event waits. ctx must outlive this object.
|
inlineexplicit |
Factor A immediately. Equivalent to LLT llt; llt.compute(A).
|
inlineexplicit |
Factor a device-resident A immediately (D2D copy).
|
inlineexplicit |
Factor a device-resident A immediately (adopt, no copy; a view is copied).
|
inline |
Bind to ctx and factor A immediately.
|
inline |
Bind to ctx and factor a device-resident A (D2D copy).
|
inline |
Bind to ctx and factor a device-resident A (adopt, no copy; a view is copied).
|
inline |
Compute the Cholesky factorization of A (host matrix). The upload is complete on return; factorization remains asynchronous.
|
inline |
Compute the Cholesky factorization from a device-resident matrix (D2D copy).
|
inline |
Compute the Cholesky factorization from a device matrix (move, no copy). A view is copied instead: its storage belongs to another object.
|
inline |
|
inline |
Solve A * X = B using the cached Cholesky factor (host → host).
|
inline |
Solve in place: consumes d_B and returns it holding the solution — no RHS copy and no allocation (potrs overwrites its RHS). A view is copied instead: its storage belongs to another object.