Eigen-Contrib  5.0.1
 
Loading...
Searching...
No Matches
Eigen::gpu::internal::SparseSolverBase< Scalar_, Derived > Class Template Reference

#include <contrib/Eigen/src/GPU/GpuSparseSolverBase.h>

Detailed Description

template<typename Scalar_, typename Derived>
class Eigen::gpu::internal::SparseSolverBase< Scalar_, Derived >

CRTP base for GPU sparse direct solvers.

Template Parameters
Scalar_Element type (passed explicitly to avoid incomplete-type issues with CRTP).
DerivedThe concrete solver class (SparseLLT, SparseLDLT, SparseLU). Must provide:
  • static constexpr cudssMatrixType_t cudss_matrix_type()
  • static constexpr cudssMatrixViewType_t cudss_matrix_view()
  • static constexpr bool needs_csr_conversion()

Public Member Functions

template<typename InputType>
Derived & analyzePattern (const SparseMatrixBase< InputType > &A)
 
template<typename InputType>
Derived & compute (const SparseMatrixBase< InputType > &A)
 
const SparseSolverConfig & config () const
 
template<typename InputType>
Derived & factorize (const SparseMatrixBase< InputType > &A)
 
Derived & setConfig (const SparseSolverConfig &cfg)
 
DeviceMatrix< Scalar > solve (const DeviceMatrix< Scalar > &d_B) const
 
template<typename Rhs>
DenseMatrix solve (const MatrixBase< Rhs > &B) const
 
 SparseSolverBase (Context &ctx)
 

Constructor & Destructor Documentation

◆ SparseSolverBase()

template<typename Scalar_, typename Derived>
Eigen::gpu::internal::SparseSolverBase< Scalar_, Derived >::SparseSolverBase ( Context & ctx)
inlineexplicit

Borrow ctx's stream: solver work runs on the same stream as the caller's other GPU operations (device-resident solves chain with SpMV / cuBLAS work without cross-stream event waits). The cuDSS handle itself is always owned by this solver. ctx must outlive this object.

Member Function Documentation

◆ analyzePattern()

template<typename Scalar_, typename Derived>
template<typename InputType>
Derived & Eigen::gpu::internal::SparseSolverBase< Scalar_, Derived >::analyzePattern ( const SparseMatrixBase< InputType > & A)
inline

Symbolic analysis only. Uploads sparsity structure to device. This phase is synchronous (blocks until complete).

◆ compute()

template<typename Scalar_, typename Derived>
template<typename InputType>
Derived & Eigen::gpu::internal::SparseSolverBase< Scalar_, Derived >::compute ( const SparseMatrixBase< InputType > & A)
inline

Symbolic analysis + numeric factorization.

◆ config()

template<typename Scalar_, typename Derived>
const SparseSolverConfig & Eigen::gpu::internal::SparseSolverBase< Scalar_, Derived >::config ( ) const
inline

The configuration most recently passed to setConfig().

◆ factorize()

template<typename Scalar_, typename Derived>
template<typename InputType>
Derived & Eigen::gpu::internal::SparseSolverBase< Scalar_, Derived >::factorize ( const SparseMatrixBase< InputType > & A)
inline

Numeric factorization using the symbolic analysis from analyzePattern.

Warning
The sparsity pattern (outerIndexPtr, innerIndexPtr) must be identical to the one passed to analyzePattern(). Only the numerical values may change. Passing a different pattern is undefined behavior. This matches the contract of CHOLMOD, UMFPACK, and cuDSS's own API.

This phase is asynchronous — info() lazily synchronizes.

◆ setConfig()

template<typename Scalar_, typename Derived>
Derived & Eigen::gpu::internal::SparseSolverBase< Scalar_, Derived >::setConfig ( const SparseSolverConfig & cfg)
inline

Apply cfg to this solver. Most knobs are consumed by the phase they affect — reordering and matching by analyzePattern(), pivoting by factorize(), refinement by solve() — so call setConfig() before the first phase whose behavior it changes; phases already executed are unaffected.

The two hybrid modes are the exception. cuDSS builds different analysis state for them, and documents that analysis has to be redone when hybrid memory is enabled afterwards, so hybridMemory and hybridExecute must be set before analyzePattern(). Changing either one after analysis invalidates it: analyzePattern() has to be called again before factorize(), which would otherwise consume analysis state built for a different execution and memory mode. hybridMemoryDeviceLimit is only a budget within hybridMemory and may be changed between factorizations.

Non-default fields require cuDSS >= 0.8, which EIGEN_HAS_CUDSS_SOLVER_CONFIG reports. Below that, the algorithm enumerators are not declared, so selecting one does not compile; a non-default value of the remaining fields is rejected rather than applied, asserting and leaving info() reporting InvalidInput until the config is reset to default.

◆ solve() [1/2]

template<typename Scalar_, typename Derived>
DeviceMatrix< Scalar > Eigen::gpu::internal::SparseSolverBase< Scalar_, Derived >::solve ( const DeviceMatrix< Scalar > & d_B) const
inline

Solve A * X = B with device-resident RHS. Returns an n × nrhs DeviceMatrix that stays on device — no H2D/D2H transfer and no host synchronization (debug builds verify the factorization status first, which syncs once after each factorize()). Chain the result directly into SpMV / cuBLAS work, e.g. for iterative refinement.

◆ solve() [2/2]

template<typename Scalar_, typename Derived>
template<typename Rhs>
DenseMatrix Eigen::gpu::internal::SparseSolverBase< Scalar_, Derived >::solve ( const MatrixBase< Rhs > & B) const
inline

Solve A * X = B (host → host). Returns X as a dense matrix. Supports single or multiple right-hand sides.


The documentation for this class was generated from the following file: