![]() |
Eigen-Contrib
5.0.1
|
#include <contrib/Eigen/src/GPU/GpuLU.h>
GPU LU decomposition with partial pivoting via cuSOLVER.
| Scalar_ | Element type: float, double, complex<float>, complex<double> |
Decomposes a square matrix A = P L U on the GPU and retains the factored matrix and pivot array in device memory. Solves A*X=B, A^T*X=B, or A^H*X=B by passing the appropriate gpu::GpuOp.
Each LU object owns a dedicated CUDA stream and cuSOLVER handle.
Public Member Functions | |
| template<typename InputType> | |
| LU & | compute (const DenseBase< InputType > &A) |
| LU & | compute (const DeviceMatrix< Scalar > &d_A) |
| LU & | compute (DeviceMatrix< Scalar > &&d_A) |
| LU (const DeviceMatrix< Scalar > &d_A) | |
| LU (Context &ctx) | |
| template<typename InputType> | |
| LU (Context &ctx, const DenseBase< InputType > &A) | |
| LU (Context &ctx, const DeviceMatrix< Scalar > &d_A) | |
| LU (Context &ctx, DeviceMatrix< Scalar > &&d_A) | |
| LU (DeviceMatrix< Scalar > &&d_A) | |
| DeviceMatrix< Scalar > | solve (const DeviceMatrix< Scalar > &d_B, GpuOp op=GpuOp::NoTrans) const |
| template<typename Rhs> | |
| PlainMatrix | solve (const MatrixBase< Rhs > &B, GpuOp op=GpuOp::NoTrans) const |
| DeviceMatrix< Scalar > | solve (DeviceMatrix< Scalar > &&d_B, GpuOp op=GpuOp::NoTrans) const |
|
inlineexplicit |
Bind to ctx: run on its stream with its cuSOLVER/cuBLAS handles, so solver work chains with other work on the same Context without cross-stream event waits. ctx must outlive this object.
|
inlineexplicit |
Factor a device-resident A immediately (D2D copy).
|
inlineexplicit |
Factor a device-resident A immediately (adopt, no copy; a view is copied).
|
inline |
Bind to ctx and factor A immediately.
|
inline |
Bind to ctx and factor a device-resident A (D2D copy).
|
inline |
Bind to ctx and factor a device-resident A (adopt, no copy; a view is copied).
|
inline |
Compute the LU factorization of A (host matrix, must be square). The upload is complete on return; factorization remains asynchronous.
|
inline |
Compute the LU factorization from a device-resident matrix (D2D copy).
|
inline |
Compute the LU factorization from a device matrix (move, no copy). A view is copied instead: its storage belongs to another object.
|
inline |
|
inline |
|
inline |
Solve in place: consumes d_B and returns it holding the solution — no RHS copy and no allocation (getrs overwrites its RHS). A view is copied instead: its storage belongs to another object.