Eigen-Contrib  5.0.1
 
Loading...
Searching...
No Matches
Eigen::gpu::LLT< Scalar_, UpLo_ > Class Template Reference

#include <contrib/Eigen/src/GPU/GpuLLT.h>

Detailed Description

template<typename Scalar_, int UpLo_ = Lower>
class Eigen::gpu::LLT< Scalar_, UpLo_ >

GPU Cholesky (LL^T) decomposition via cuSOLVER.

Template Parameters
Scalar_Element type: float, double, complex<float>, complex<double>
UpLo_Triangle used: Lower (default) or Upper

Factorizes a symmetric positive-definite matrix A = LL^H on the GPU and caches the factor L in device memory. Each subsequent solve(B) uploads only B, calls cusolverDnXpotrs, and downloads the result — the factor is not re-transferred.

Each LLT object owns a dedicated CUDA stream and cuSOLVER handle, enabling concurrent factorizations from multiple objects on the same host thread.

Public Member Functions

template<typename InputType>
LLT & compute (const DenseBase< InputType > &A)
 
LLT & compute (const DeviceMatrix< Scalar > &d_A)
 
LLT & compute (DeviceMatrix< Scalar > &&d_A)
 
 LLT ()=default
 
template<typename InputType>
 LLT (const DenseBase< InputType > &A)
 
 LLT (const DeviceMatrix< Scalar > &d_A)
 
 LLT (Context &ctx)
 
template<typename InputType>
 LLT (Context &ctx, const DenseBase< InputType > &A)
 
 LLT (Context &ctx, const DeviceMatrix< Scalar > &d_A)
 
 LLT (Context &ctx, DeviceMatrix< Scalar > &&d_A)
 
 LLT (DeviceMatrix< Scalar > &&d_A)
 
DeviceMatrix< Scalar > solve (const DeviceMatrix< Scalar > &d_B) const
 
template<typename Rhs>
PlainMatrix solve (const MatrixBase< Rhs > &B) const
 
DeviceMatrix< Scalar > solve (DeviceMatrix< Scalar > &&d_B) const
 

Constructor & Destructor Documentation

◆ LLT() [1/8]

template<typename Scalar_, int UpLo_ = Lower>
Eigen::gpu::LLT< Scalar_, UpLo_ >::LLT ( )
default

Default constructor. Does not factorize; call compute() before solve().

◆ LLT() [2/8]

template<typename Scalar_, int UpLo_ = Lower>
Eigen::gpu::LLT< Scalar_, UpLo_ >::LLT ( Context & ctx)
inlineexplicit

Bind to ctx: run on its stream with its cuSOLVER/cuBLAS handles, so solver work chains with other work on the same Context without cross-stream event waits. ctx must outlive this object.

◆ LLT() [3/8]

template<typename Scalar_, int UpLo_ = Lower>
template<typename InputType>
Eigen::gpu::LLT< Scalar_, UpLo_ >::LLT ( const DenseBase< InputType > & A)
inlineexplicit

Factor A immediately. Equivalent to LLT llt; llt.compute(A).

◆ LLT() [4/8]

template<typename Scalar_, int UpLo_ = Lower>
Eigen::gpu::LLT< Scalar_, UpLo_ >::LLT ( const DeviceMatrix< Scalar > & d_A)
inlineexplicit

Factor a device-resident A immediately (D2D copy).

◆ LLT() [5/8]

template<typename Scalar_, int UpLo_ = Lower>
Eigen::gpu::LLT< Scalar_, UpLo_ >::LLT ( DeviceMatrix< Scalar > && d_A)
inlineexplicit

Factor a device-resident A immediately (adopt, no copy; a view is copied).

◆ LLT() [6/8]

template<typename Scalar_, int UpLo_ = Lower>
template<typename InputType>
Eigen::gpu::LLT< Scalar_, UpLo_ >::LLT ( Context & ctx,
const DenseBase< InputType > & A )
inline

Bind to ctx and factor A immediately.

◆ LLT() [7/8]

template<typename Scalar_, int UpLo_ = Lower>
Eigen::gpu::LLT< Scalar_, UpLo_ >::LLT ( Context & ctx,
const DeviceMatrix< Scalar > & d_A )
inline

Bind to ctx and factor a device-resident A (D2D copy).

◆ LLT() [8/8]

template<typename Scalar_, int UpLo_ = Lower>
Eigen::gpu::LLT< Scalar_, UpLo_ >::LLT ( Context & ctx,
DeviceMatrix< Scalar > && d_A )
inline

Bind to ctx and factor a device-resident A (adopt, no copy; a view is copied).

Member Function Documentation

◆ compute() [1/3]

template<typename Scalar_, int UpLo_ = Lower>
template<typename InputType>
LLT & Eigen::gpu::LLT< Scalar_, UpLo_ >::compute ( const DenseBase< InputType > & A)
inline

Compute the Cholesky factorization of A (host matrix). The upload is complete on return; factorization remains asynchronous.

◆ compute() [2/3]

template<typename Scalar_, int UpLo_ = Lower>
LLT & Eigen::gpu::LLT< Scalar_, UpLo_ >::compute ( const DeviceMatrix< Scalar > & d_A)
inline

Compute the Cholesky factorization from a device-resident matrix (D2D copy).

◆ compute() [3/3]

template<typename Scalar_, int UpLo_ = Lower>
LLT & Eigen::gpu::LLT< Scalar_, UpLo_ >::compute ( DeviceMatrix< Scalar > && d_A)
inline

Compute the Cholesky factorization from a device matrix (move, no copy). A view is copied instead: its storage belongs to another object.

◆ solve() [1/3]

template<typename Scalar_, int UpLo_ = Lower>
DeviceMatrix< Scalar > Eigen::gpu::LLT< Scalar_, UpLo_ >::solve ( const DeviceMatrix< Scalar > & d_B) const
inline

Solve A * X = B with device-resident RHS. Fully asynchronous: returns immediately after enqueuing the solve. Debug builds verify the factorization status first (one host sync on the first solve after compute()); release builds do not — use info() when failure must be detected.

◆ solve() [2/3]

template<typename Scalar_, int UpLo_ = Lower>
template<typename Rhs>
PlainMatrix Eigen::gpu::LLT< Scalar_, UpLo_ >::solve ( const MatrixBase< Rhs > & B) const
inline

Solve A * X = B using the cached Cholesky factor (host → host).

◆ solve() [3/3]

template<typename Scalar_, int UpLo_ = Lower>
DeviceMatrix< Scalar > Eigen::gpu::LLT< Scalar_, UpLo_ >::solve ( DeviceMatrix< Scalar > && d_B) const
inline

Solve in place: consumes d_B and returns it holding the solution — no RHS copy and no allocation (potrs overwrites its RHS). A view is copied instead: its storage belongs to another object.


The documentation for this class was generated from the following files: