Eigen  5.0.1
 
Loading...
Searching...
No Matches
Using AMD® AOCL from Eigen

Since Eigen version 3.4 and later, users can benefit from built-in AMD® Optimizing CPU Libraries (AOCL) optimizations with an installed copy of AOCL 5.0 (or later).

AMD AOCL provides highly optimized, multi-threaded mathematical routines for x86-64 processors with a focus on AMD "Zen"-based architectures. AOCL is available on Linux and Windows for x86-64 architectures.

Note
AMD® AOCL is freely available software, but it is the responsibility of users to download, install, and ensure their product's license allows linking to the AOCL libraries. AOCL is distributed under a permissive license that allows commercial use.

Using AMD AOCL through Eigen is straightforward:

  1. export AOCL_ROOT into your environment
  2. define one of the AOCL macros before including any Eigen headers (see table below)
  3. link your program to AOCL libraries (BLIS, FLAME, LibM)
  4. ensure your system supports the target architecture optimizations

When doing so, a number of Eigen's algorithms are silently substituted with calls to AMD AOCL routines. These substitutions apply only for Dynamic or large enough objects with one of the following standard scalar types: float, double, complex<float>, and complex<double>. Operations on other scalar types or mixing reals and complexes will continue to use the built-in algorithms.

The AOCL integration targets three core components:

  • BLIS: High-performance BLAS implementation optimized for modern cache hierarchies
  • FLAME: Dense linear algebra algorithms providing LAPACK functionality
  • LibM: Optimized standard math routines with vectorized implementations

Configuration Macros

You can choose which parts will be substituted by defining one or multiple of the following macros:

EIGEN_USE_BLAS Enables the use of external BLAS level 2 and 3 routines (AOCL-BLIS)
EIGEN_USE_LAPACKE Enables the use of external LAPACK routines via the LAPACKE C interface (AOCL-FLAME)
EIGEN_USE_LAPACKE_STRICT Same as EIGEN_USE_LAPACKE but algorithms of lower robustness are disabled.
This currently concerns only JacobiSVD which would be replaced by gesvd.
EIGEN_USE_AOCL_VML Enables the use of AOCL LibM vector math operations for coefficient-wise functions
EIGEN_USE_AOCL_ALL Defines EIGEN_USE_BLAS, EIGEN_USE_LAPACKE, and EIGEN_USE_AOCL_VML
EIGEN_USE_AOCL_MT Equivalent to EIGEN_USE_AOCL_ALL, but ensures multi-threaded BLIS (libblis-mt) is used.
Recommended for most applications.
Note
The AOCL integration automatically enables optimizations when the matrix/vector size exceeds EIGEN_AOCL_VML_THRESHOLD (default: 128 elements). For smaller operations, Eigen's built-in vectorization may be faster due to function call overhead.

Performance Considerations

The EIGEN_USE_BLAS and EIGEN_USE_LAPACKE macros can be combined with AOCL-specific optimizations:

  • Multi-threading: Use EIGEN_USE_AOCL_MT to automatically select the multi-threaded BLIS library
  • Architecture targeting: AOCL libraries are optimized for AMD Zen architectures (Zen, Zen2, Zen3, Zen4, Zen5)
  • Vector Math Library: AOCL LibM provides vectorized implementations that can operate on entire arrays simultaneously
  • Memory layout: Eigen's column-major storage directly matches AOCL's expected data layout for zero-copy operation

Supported Data Types and Sizes

AOCL acceleration is applied to:

  • Scalar types: float, double, complex<float>, complex<double>
  • Matrix/Vector sizes: Dynamic size or compile-time size ≥ EIGEN_AOCL_VML_THRESHOLD
  • Storage order: Both column-major (default) and row-major layouts
  • Memory alignment: Eigen's data pointers are directly compatible with AOCL function signatures

The current AOCL Vector Math Library integration is specialized for double precision, with automatic fallback to scalar implementations for float.

Vector Math Functions

The following table summarizes coefficient-wise operations accelerated by EIGEN_USE_AOCL_VML:

Code exampleAOCL routines
v2 = v1.array().exp();
v2 = v1.array().sin();
v2 = v1.array().cos();
v2 = v1.array().tan();
v2 = v1.array().log();
v2 = v1.array().log10();
v2 = v1.array().log2();
v2 = v1.array().sqrt();
v2 = v1.array().pow(1.5);
v2 = v1.array() + v2.array();
amd_vrda_exp
amd_vrda_sin
amd_vrda_cos
amd_vrda_tan
amd_vrda_log
amd_vrda_log10
amd_vrda_log2
amd_vrda_sqrt
amd_vrda_pow
amd_vrda_add

In the examples, v1 and v2 are dense vectors of type VectorXd with size ≥ EIGEN_AOCL_VML_THRESHOLD.

Complete Example

#define EIGEN_USE_AOCL_MT
#include <iostream>
#include <Eigen/Dense>
int main() {
const int n = 2048;
// Large matrices automatically use AOCL-BLIS for multiplication
Eigen::MatrixXd A = Eigen::MatrixXd::Random(n, n);
Eigen::MatrixXd B = Eigen::MatrixXd::Random(n, n);
Eigen::MatrixXd C = A * B; // Dispatched to dgemm
// Large vectors automatically use AOCL LibM for math functions
Eigen::VectorXd v = Eigen::VectorXd::LinSpaced(10000, 0, 10);
Eigen::VectorXd result = v.array().sin(); // Dispatched to amd_vrda_sin
// LAPACK decompositions use AOCL-FLAME
Eigen::LLT<Eigen::MatrixXd> llt(A); // Dispatched to dpotrf
std::cout << "Matrix norm: " << C.norm() << std::endl;
std::cout << "Vector result norm: " << result.norm() << std::endl;
return 0;
}
Standard Cholesky decomposition (LL^T) of a matrix and associated features.
Definition LLT.h:85
Matrix< double, Dynamic, Dynamic > MatrixXd
Dynamic×Dynamic matrix of type double.
Definition Matrix.h:489
Matrix< double, Dynamic, 1 > VectorXd
Dynamic×1 vector of type double.
Definition Matrix.h:489

Building and Linking

To compile with AOCL support, set the AOCL_ROOT environment variable and link against the required libraries:

export AOCL_ROOT=/path/to/aocl
clang++ -O3 -g -DEIGEN_USE_AOCL_ALL \
-I./install/include -I${AOCL_ROOT}/include \
-Wno-parentheses my_app.cpp \
-L${AOCL_ROOT} -lamdlibm -lflame -lblis \
-lpthread -lrt -lm -lomp \
-o eigen_aocl_example

For multi-threaded performance, use the multi-threaded BLIS library:

clang++ -O3 -g -DEIGEN_USE_AOCL_MT \
-I./install/include -I${AOCL_ROOT}/include \
-Wno-parentheses my_app.cpp \
-L${AOCL_ROOT} -lamdlibm -lflame -lblis-mt \
-lpthread -lrt -lm -lomp \
-o eigen_aocl_example

Key compiler and linker flags:

  • -DEIGEN_USE_AOCL_ALL: Enable all AOCL accelerations (BLAS, LAPACK, VML)
  • -DEIGEN_USE_AOCL_MT: Enable multi-threaded version (uses -lblis-mt)
  • -lblis: Single-threaded BLIS library
  • -lblis-mt: Multi-threaded BLIS library (recommended for performance)
  • -lflame: FLAME LAPACK implementation
  • -lamdlibm: AMD LibM vector math library
  • -lomp: OpenMP runtime for multi-threading support
  • -lpthread -lrt: System threading and real-time libraries
  • -Wno-parentheses: Suppress common warnings when using AOCL headers

Building Eigen with AOCL Support

To build Eigen with AOCL Support, use the following CMake configuration:

cmake .. -DCMAKE_BUILD_TYPE=Release \
-DCMAKE_C_COMPILER=clang \
-DCMAKE_CXX_COMPILER=clang++ \
-DCMAKE_INSTALL_PREFIX=$PWD/install \
-DINCLUDE_INSTALL_DIR=$PWD/install/include \
&& make install -j$(nproc)

To build Eigen with AOCL integration, use the following CMake configuration:

cmake .. -DCMAKE_BUILD_TYPE=Release \
-DCMAKE_C_COMPILER=clang \
-DCMAKE_CXX_COMPILER=clang++ \
-DCMAKE_INSTALL_PREFIX=$PWD/install \
-DINCLUDE_INSTALL_DIR=$PWD/install/include \
&& make install -j$(nproc)

CMake Configuration Parameters:**

ParameterExpected ValuesDescription
CMAKE_BUILD_TYPE Release, Debug, RelWithDebInfo Build configuration (Release recommended for benchmarks)
CMAKE_C_COMPILER clang, gcc C compiler (clang recommended for AOCL)
CMAKE_CXX_COMPILER clang++, g++ C++ compiler (clang++ recommended for AOCL)
CMAKE_INSTALL_PREFIX Installation pathWhere to install Eigen headers
INCLUDE_INSTALL_DIR Header pathSpecific path for Eigen headers

Architecture Selection Guide:**

  • znver3: AMD Zen 3 (EPYC 7003, Ryzen 5000 series)
  • znver4: AMD Zen 4 (EPYC 9004, Ryzen 7000 series)
  • znver5: AMD Zen 5 (EPYC 9005, Ryzen 9000 series)
  • native: Auto-detect current CPU architecture
  • generic: Generic x86-64 without specific optimizations

    Custom Compiler Flags Explanation:**

  • -O3: Maximum optimization level
  • -mavx512f: Enable AVX-512 instruction set (if supported)
  • -fveclib=AMDLIBM: Use AMD LibM for vectorized math functions

Building the AOCL Benchmark

After configuring Eigen, build the AOCL benchmark executable:

cmake --build . --target benchmark_aocl -j$(nproc)

This creates the benchmark_aocl executable that demonstrates AOCL acceleration with various matrix sizes and operations.

Running the Benchmark:**

./benchmark_aocl

The benchmark will automatically compare:

  • Eigen's native performance vs AOCL-accelerated operations
  • Matrix multiplication performance (BLIS vs Eigen)
  • Vector math functions performance (LibM vs Eigen)
  • Memory bandwidth utilization and cache efficiency

CMake Integration

When using CMake, you can use a FindAOCL module:

find_package(AOCL REQUIRED)
target_compile_definitions(my_target PRIVATE EIGEN_USE_AOCL_MT)
target_link_libraries(my_target PRIVATE AOCL::BLIS_MT AOCL::FLAME AOCL::LIBM)

Troubleshooting

Common issues and solutions:

  • Link errors: Ensure AOCL_ROOT is set and libraries are in LD_LIBRARY_PATH
  • Performance not improved: Verify you're using matrices/vectors larger than the threshold
  • Thread contention: Set OMP_NUM_THREADS to match your CPU core count
  • Architecture mismatch: Use appropriate -march flag for your AMD processor

Links

  • AMD AOCL can be downloaded for free here
  • AOCL User Guide and documentation available on the AMD Developer Portal
  • AOCL is also available through package managers and containerized environments