Eigen-Contrib  5.0.1
 
Loading...
Searching...
No Matches
Eigen::gpu::LU< Scalar_ > Class Template Reference

#include <contrib/Eigen/src/GPU/GpuLU.h>

Detailed Description

template<typename Scalar_>
class Eigen::gpu::LU< Scalar_ >

GPU LU decomposition with partial pivoting via cuSOLVER.

Template Parameters
Scalar_Element type: float, double, complex<float>, complex<double>

Decomposes a square matrix A = P L U on the GPU and retains the factored matrix and pivot array in device memory. Solves A*X=B, A^T*X=B, or A^H*X=B by passing the appropriate gpu::GpuOp.

Each LU object owns a dedicated CUDA stream and cuSOLVER handle.

Public Member Functions

template<typename InputType>
LU & compute (const DenseBase< InputType > &A)
 
LU & compute (const DeviceMatrix< Scalar > &d_A)
 
LU & compute (DeviceMatrix< Scalar > &&d_A)
 
 LU (const DeviceMatrix< Scalar > &d_A)
 
 LU (Context &ctx)
 
template<typename InputType>
 LU (Context &ctx, const DenseBase< InputType > &A)
 
 LU (Context &ctx, const DeviceMatrix< Scalar > &d_A)
 
 LU (Context &ctx, DeviceMatrix< Scalar > &&d_A)
 
 LU (DeviceMatrix< Scalar > &&d_A)
 
DeviceMatrix< Scalar > solve (const DeviceMatrix< Scalar > &d_B, GpuOp op=GpuOp::NoTrans) const
 
template<typename Rhs>
PlainMatrix solve (const MatrixBase< Rhs > &B, GpuOp op=GpuOp::NoTrans) const
 
DeviceMatrix< Scalar > solve (DeviceMatrix< Scalar > &&d_B, GpuOp op=GpuOp::NoTrans) const
 

Constructor & Destructor Documentation

◆ LU() [1/6]

template<typename Scalar_>
Eigen::gpu::LU< Scalar_ >::LU ( Context & ctx)
inlineexplicit

Bind to ctx: run on its stream with its cuSOLVER/cuBLAS handles, so solver work chains with other work on the same Context without cross-stream event waits. ctx must outlive this object.

◆ LU() [2/6]

template<typename Scalar_>
Eigen::gpu::LU< Scalar_ >::LU ( const DeviceMatrix< Scalar > & d_A)
inlineexplicit

Factor a device-resident A immediately (D2D copy).

◆ LU() [3/6]

template<typename Scalar_>
Eigen::gpu::LU< Scalar_ >::LU ( DeviceMatrix< Scalar > && d_A)
inlineexplicit

Factor a device-resident A immediately (adopt, no copy; a view is copied).

◆ LU() [4/6]

template<typename Scalar_>
template<typename InputType>
Eigen::gpu::LU< Scalar_ >::LU ( Context & ctx,
const DenseBase< InputType > & A )
inline

Bind to ctx and factor A immediately.

◆ LU() [5/6]

template<typename Scalar_>
Eigen::gpu::LU< Scalar_ >::LU ( Context & ctx,
const DeviceMatrix< Scalar > & d_A )
inline

Bind to ctx and factor a device-resident A (D2D copy).

◆ LU() [6/6]

template<typename Scalar_>
Eigen::gpu::LU< Scalar_ >::LU ( Context & ctx,
DeviceMatrix< Scalar > && d_A )
inline

Bind to ctx and factor a device-resident A (adopt, no copy; a view is copied).

Member Function Documentation

◆ compute() [1/3]

template<typename Scalar_>
template<typename InputType>
LU & Eigen::gpu::LU< Scalar_ >::compute ( const DenseBase< InputType > & A)
inline

Compute the LU factorization of A (host matrix, must be square). The upload is complete on return; factorization remains asynchronous.

◆ compute() [2/3]

template<typename Scalar_>
LU & Eigen::gpu::LU< Scalar_ >::compute ( const DeviceMatrix< Scalar > & d_A)
inline

Compute the LU factorization from a device-resident matrix (D2D copy).

◆ compute() [3/3]

template<typename Scalar_>
LU & Eigen::gpu::LU< Scalar_ >::compute ( DeviceMatrix< Scalar > && d_A)
inline

Compute the LU factorization from a device matrix (move, no copy). A view is copied instead: its storage belongs to another object.

◆ solve() [1/3]

template<typename Scalar_>
DeviceMatrix< Scalar > Eigen::gpu::LU< Scalar_ >::solve ( const DeviceMatrix< Scalar > & d_B,
GpuOp op = GpuOp::NoTrans ) const
inline

Solve op(A) * X = B with device-resident RHS. Fully asynchronous: returns immediately after enqueuing the solve. Debug builds verify the factorization status first (one host sync on the first solve after compute()); release builds do not — use info() when failure must be detected.

◆ solve() [2/3]

template<typename Scalar_>
template<typename Rhs>
PlainMatrix Eigen::gpu::LU< Scalar_ >::solve ( const MatrixBase< Rhs > & B,
GpuOp op = GpuOp::NoTrans ) const
inline

Solve op(A) * X = B using the cached LU factorization (host → host).

Parameters
BRight-hand side (n x nrhs host matrix).
opgpu::GpuOp::NoTrans (default), Trans, or ConjTrans.

◆ solve() [3/3]

template<typename Scalar_>
DeviceMatrix< Scalar > Eigen::gpu::LU< Scalar_ >::solve ( DeviceMatrix< Scalar > && d_B,
GpuOp op = GpuOp::NoTrans ) const
inline

Solve in place: consumes d_B and returns it holding the solution — no RHS copy and no allocation (getrs overwrites its RHS). A view is copied instead: its storage belongs to another object.


The documentation for this class was generated from the following files: