![]() |
Eigen
5.0.1
|
Since Eigen version 3.4 and later, users can benefit from built-in AMD® Optimizing CPU Libraries (AOCL) optimizations with an installed copy of AOCL 5.0 (or later).
AMD AOCL provides highly optimized, multi-threaded mathematical routines for x86-64 processors with a focus on AMD "Zen"-based architectures. AOCL is available on Linux and Windows for x86-64 architectures.
Using AMD AOCL through Eigen is straightforward:
AOCL_ROOT into your environmentWhen doing so, a number of Eigen's algorithms are silently substituted with calls to AMD AOCL routines. These substitutions apply only for Dynamic or large enough objects with one of the following standard scalar types: float, double, complex<float>, and complex<double>. Operations on other scalar types or mixing reals and complexes will continue to use the built-in algorithms.
The AOCL integration targets three core components:
You can choose which parts will be substituted by defining one or multiple of the following macros:
EIGEN_USE_BLAS | Enables the use of external BLAS level 2 and 3 routines (AOCL-BLIS) |
EIGEN_USE_LAPACKE | Enables the use of external LAPACK routines via the LAPACKE C interface (AOCL-FLAME) |
EIGEN_USE_LAPACKE_STRICT | Same as EIGEN_USE_LAPACKE but algorithms of lower robustness are disabled. This currently concerns only JacobiSVD which would be replaced by gesvd. |
EIGEN_USE_AOCL_VML | Enables the use of AOCL LibM vector math operations for coefficient-wise functions |
EIGEN_USE_AOCL_ALL | Defines EIGEN_USE_BLAS, EIGEN_USE_LAPACKE, and EIGEN_USE_AOCL_VML |
EIGEN_USE_AOCL_MT | Equivalent to EIGEN_USE_AOCL_ALL, but ensures multi-threaded BLIS (libblis-mt) is used. Recommended for most applications. |
EIGEN_AOCL_VML_THRESHOLD (default: 128 elements). For smaller operations, Eigen's built-in vectorization may be faster due to function call overhead.The EIGEN_USE_BLAS and EIGEN_USE_LAPACKE macros can be combined with AOCL-specific optimizations:
EIGEN_USE_AOCL_MT to automatically select the multi-threaded BLIS libraryAOCL acceleration is applied to:
float, double, complex<float>, complex<double> EIGEN_AOCL_VML_THRESHOLD The current AOCL Vector Math Library integration is specialized for double precision, with automatic fallback to scalar implementations for float.
The following table summarizes coefficient-wise operations accelerated by EIGEN_USE_AOCL_VML:
| Code example | AOCL routines |
|---|---|
v2 = v1.array().exp();
v2 = v1.array().sin();
v2 = v1.array().cos();
v2 = v1.array().tan();
v2 = v1.array().log();
v2 = v1.array().log10();
v2 = v1.array().log2();
v2 = v1.array().sqrt();
v2 = v1.array().pow(1.5);
v2 = v1.array() + v2.array();
| amd_vrda_exp
amd_vrda_sin
amd_vrda_cos
amd_vrda_tan
amd_vrda_log
amd_vrda_log10
amd_vrda_log2
amd_vrda_sqrt
amd_vrda_pow
amd_vrda_add
|
In the examples, v1 and v2 are dense vectors of type VectorXd with size ≥ EIGEN_AOCL_VML_THRESHOLD.
To compile with AOCL support, set the AOCL_ROOT environment variable and link against the required libraries:
For multi-threaded performance, use the multi-threaded BLIS library:
Key compiler and linker flags:
-DEIGEN_USE_AOCL_ALL: Enable all AOCL accelerations (BLAS, LAPACK, VML)-DEIGEN_USE_AOCL_MT: Enable multi-threaded version (uses -lblis-mt)-lblis: Single-threaded BLIS library-lblis-mt: Multi-threaded BLIS library (recommended for performance)-lflame: FLAME LAPACK implementation -lamdlibm: AMD LibM vector math library-lomp: OpenMP runtime for multi-threading support-lpthread -lrt: System threading and real-time libraries-Wno-parentheses: Suppress common warnings when using AOCL headersTo build Eigen with AOCL Support, use the following CMake configuration:
To build Eigen with AOCL integration, use the following CMake configuration:
CMake Configuration Parameters:**
| Parameter | Expected Values | Description |
|---|---|---|
CMAKE_BUILD_TYPE | Release, Debug, RelWithDebInfo | Build configuration (Release recommended for benchmarks) |
CMAKE_C_COMPILER | clang, gcc | C compiler (clang recommended for AOCL) |
CMAKE_CXX_COMPILER | clang++, g++ | C++ compiler (clang++ recommended for AOCL) |
CMAKE_INSTALL_PREFIX | Installation path | Where to install Eigen headers |
INCLUDE_INSTALL_DIR | Header path | Specific path for Eigen headers |
Architecture Selection Guide:**
znver3: AMD Zen 3 (EPYC 7003, Ryzen 5000 series)znver4: AMD Zen 4 (EPYC 9004, Ryzen 7000 series) znver5: AMD Zen 5 (EPYC 9005, Ryzen 9000 series)native: Auto-detect current CPU architecturegeneric: Generic x86-64 without specific optimizations
Custom Compiler Flags Explanation:**
-O3: Maximum optimization level-mavx512f: Enable AVX-512 instruction set (if supported)-fveclib=AMDLIBM: Use AMD LibM for vectorized math functionsAfter configuring Eigen, build the AOCL benchmark executable:
This creates the benchmark_aocl executable that demonstrates AOCL acceleration with various matrix sizes and operations.
Running the Benchmark:**
The benchmark will automatically compare:
When using CMake, you can use a FindAOCL module:
Common issues and solutions:
AOCL_ROOT is set and libraries are in LD_LIBRARY_PATH OMP_NUM_THREADS to match your CPU core count-march flag for your AMD processor