![]() |
Eigen
5.0.1
|
You can control some aspects of Eigen by defining the preprocessor tokens using #define. These macros should be defined before any Eigen headers are included. Often they are best set in the project options.
This page lists the preprocessor tokens recognized by Eigen.
These macros have a major effect and typically break the API (Application Programming Interface) and/or the ABI (Application Binary Interface). This can be rather dangerous: if parts of your program are compiled with one option, and other parts (or libraries that you use) are compiled with another option, your program may fail to link or exhibit subtle bugs. Nevertheless, these options can be useful for people who know what they are doing.
std::ptrdiff_t by default.1x1 (resp. 2x1 or 1x2) fixed size matrices is always interpreted as an initialization constructor where the argument(s) are the coefficient values and not the sizes. For instance, 2x1 vector initialized with zeros (i.e., [0,0]). If such cases might occur, then it is recommended to use the default constructor with a explicit call to resize: EIGEN_INITIALIZE_MATRICES_BY_ZERO for a discussion on a limitations of these macros when applied to 1x1, 1x2, and 2x1 fixed-size matrices.a = b have to be of the same size; otherwise, Eigen automatically resizes a so that it is of the correct size. Not defined by default.By default, Eigen strives to automatically detect and enable language features at compile-time based on the information provided by the compiler.
Individual features can be explicitly enabled or disabled by defining the following token to 0 or 1 respectively. For instance, one might disable use of std::invoke_result by defining EIGEN_HAS_STD_INVOKE_RESULT=0.
std::invoke_result (C++17). When disabled, std::result_of is used as a fallback.<iostreams>.The Eigen library contains many assertions to guard against programming errors, both at compile time and at run time. However, these assertions do cost time and can thus be turned off.
NDEBUG macro is defined (this is a standard C++ macro which disables all asserts).assert, which aborts the program if the assertion is violated. Redefine this macro if you want to do something else, like throwing an exception.EIGEN_MALLOC_ALREADY_ALIGNED - Can be set to 0 or 1 to tell whether default system malloc already returns aligned buffers. If not defined, then this information is automatically deduced from the compiler and system preprocessor tokens.EIGEN_MAX_ALIGN_BYTES - Must be a power of two, or 0. Defines an upper bound on the memory boundary in bytes on which dynamically and statically allocated data may be aligned by Eigen. If not defined, a default value is automatically computed based on architecture, compiler, and OS. This option is typically used to enforce binary compatibility between code/libraries compiled with different SIMD options. For instance, one may compile AVX code and enforce ABI compatibility with existing SSE code by defining EIGEN_MAX_ALIGN_BYTES=16. In the other way round, since by default AVX implies 32 bytes alignment for best performance, one can compile SSE code to be ABI compatible with AVX code by defining EIGEN_MAX_ALIGN_BYTES=32.EIGEN_MAX_STATIC_ALIGN_BYTES - Same as EIGEN_MAX_ALIGN_BYTES but for statically allocated data only. By default, if only EIGEN_MAX_ALIGN_BYTES is defined, then EIGEN_MAX_STATIC_ALIGN_BYTES == EIGEN_MAX_ALIGN_BYTES, otherwise a default value is automatically computed based on architecture, compiler, and OS (can be smaller than the default value of EIGEN_MAX_ALIGN_BYTES on architectures that do not support stack alignment). Let us emphasize that EIGEN_MAX_*_ALIGN_BYTES define only a desirable upper bound. In practice data is aligned to largest power-of-two common divisor of EIGEN_MAX_STATIC_ALIGN_BYTES and the size of the data, such that memory is not wasted.EIGEN_DONT_PARALLELIZE - if defined, this disables multi-threading. This is only relevant if you enabled OpenMP. See Eigen and multi-threading for details.EIGEN_GEMM_THREADPOOL - if defined, the general matrix-matrix product is parallelized using a user-provided Eigen::ThreadPool instead of OpenMP; register the pool with Eigen::setGemmThreadPool(). Mutually exclusive with OpenMP parallelization. See Eigen and multi-threading for details.EIGEN_DONT_VECTORIZE - disables explicit vectorization when defined. Not defined by default, unless alignment is disabled by Eigen's platform test or the user defining EIGEN_DONT_ALIGN. The implication also runs the other way outside GPU compilation: defining it sets the ideal alignment to zero, so EIGEN_MAX_ALIGN_BYTES and EIGEN_MAX_STATIC_ALIGN_BYTES default to 0 and fixed-size vectorizable types lose their over-alignment. It therefore affects the ABI and must be defined consistently across translation units. See Vectorization.EIGEN_UNALIGNED_VECTORIZE - disables/enables vectorization with unaligned stores. Default is 1 (enabled). If set to 0 (disabled), then expression for which the destination cannot be aligned are not vectorized (e.g., unaligned small fixed size vectors or matrices)EIGEN_FAST_MATH - enables optimizations which might affect the accuracy of the result. On the SIMD backends where such faster but slightly less accurate implementations exist, this currently controls the vectorization of sin(), cos(), tan(), tanh(), erf(), erfc(), and reciprocal, and selects faster sqrt()/rsqrt() implementations (e.g., Newton-Raphson iterations on AVX512). Defined to 1 by default. Define it to 0 to disable.EIGEN_UNROLLING_LIMIT - defines the size of a loop to enable meta unrolling. Set it to zero to disable unrolling. The size of a loop here is expressed in Eigen's own notion of "number of FLOPS", it does not correspond to the number of iterations or the number of instructions. The default is value 110.EIGEN_GEMM_TO_COEFFBASED_THRESHOLD - a matrix-matrix product that reaches the general (GEMM) product path falls back at run time to the coefficient-based product when rows + cols + depth is below this value. Default is 20. In an ARM SME build, the scalar types the SME GEMM kernel handles use EIGEN_SME_GEMM_TO_COEFFBASED_THRESHOLD instead when it is defined, and also fall back when the result is not a vector and has at most EIGEN_SME_GEMM_TO_COEFFBASED_OUTPUT_AREA_THRESHOLD(Scalar) coefficients (default 96 / sizeof(Scalar)).EIGEN_FIXED_SIZE_GEMM_TO_COEFFBASED_THRESHOLD - a product that would take the GEMM path and whose three dimensions are all fixed at compile time uses the coefficient-based product instead, selected at compile time, when rows + cols + depth is below this value. Default is 40. It does not follow EIGEN_GEMM_TO_COEFFBASED_THRESHOLD, which still applies at run time to the fixed-size products this one sends to the GEMM path, so their crossover is the larger of the two values.EIGEN_SME_FIXED_SIZE_GEMM_TO_COEFFBASED_THRESHOLD - the same compile-time bound for the scalar types the ARM SME GEMM kernel handles (float, std::complex<float>, and with FEAT_SME_F64F64 double and std::complex<double>). Default is 61, measured on Apple M4 Pro; other scalar types use EIGEN_FIXED_SIZE_GEMM_TO_COEFFBASED_THRESHOLD in an SME build too.EIGEN_STACK_ALLOCATION_LIMIT - defines the maximum bytes for a buffer to be allocated on the stack. For internal temporary buffers, dynamic memory allocation is employed as a fall back; setting the limit to 0 disables stack allocation for them, so every such temporary is heap-allocated. For fixed-size matrices or arrays, exceeding this threshold raises a compile time assertion; there, 0 disables the check. Default is 128 KB.EIGEN_STACK_ALIGN_BYTES - alignment of the internal temporary buffers: the ones on the stack, their heap fallback above EIGEN_STACK_ALLOCATION_LIMIT and the GEMM blocking buffers. Defaults to EIGEN_DEFAULT_ALIGN_BYTES, except in ARM SME builds where it is 64: the SME GEMM kernel reads its packed panels a streaming vector at a time and slows by 35-50% when they straddle 64-byte lines. It does not affect the ABI.EIGEN_NO_CUDA - disables CUDA support when defined. Might be useful in .cu files for which Eigen is used on the host only, and never called from device code.EIGEN_NO_HIP - disables HIP support when defined. Analogous to EIGEN_NO_CUDA for AMD's HIP compiler.EIGEN_STRONG_INLINE - This macro is used to qualify critical functions and methods that we expect the compiler to inline. By default it is defined to __forceinline for MSVC and ICC, and to inline for other compilers. A typical usage is to define it to inline for MSVC users wanting faster compilation times, at the risk of performance degradations in some rare cases for which MSVC inliner fails to do a good job.EIGEN_DEFAULT_L1_CACHE_SIZE - Sets the default L1 cache size that is used in Eigen's GEBP kernel when the correct cache size cannot be determined at runtime.EIGEN_DEFAULT_L2_CACHE_SIZE - Sets the default L2 cache size that is used in Eigen's GEBP kernel when the correct cache size cannot be determined at runtime.EIGEN_DEFAULT_L3_CACHE_SIZE - Sets the default L3 cache size that is used in Eigen's GEBP kernel when the correct cache size cannot be determined at runtime.EIGEN_SME_MAX_KC - Maximum depth (k) blocking size used by the ARM SME GEMM kernel, expressed in float elements and scaled by the scalar width for wider types. The default (2048) is empirically tuned for Apple M4; override to retune for other SME implementations.EIGEN_SME_NEON_MAX_DEPTH, EIGEN_SME_NEON_MAX_PANEL, EIGEN_SME_NEON_MAX_WIDTH, EIGEN_SME_NEON_SHALLOW_DEPTH, EIGEN_SME_NEON_SHALLOW_WIDTH, EIGEN_SME_NEON_THIN_DIM - Bounds below which the ARM SME GEMM backend packs and multiplies a block with NEON instead of the ZA tiles (panels no deeper than the depth bound and no wider than the panel bound; a block of such panels no wider than the width bound, or the shallow width up to the shallow depth, or narrower than the thin bound on one side). One value overrides the per-scalar defaults, which are fitted for Apple M4. EIGEN_SME_NO_NEON_SMALL_BLOCKS and EIGEN_SME_FORCE_NEON_SMALL_BLOCKS pin the path for A/B runs; both also turn off the NEON kernel for results of at most 8 x 8 (float) or 4 x 8 (double).EIGEN_SME_UNITS - Number of ARM SME units a product on the SME GEMM kernel spreads over; more threads than units only contend for them. Detected from the core topology by default (see Eigen::nbSmeUnits() and Eigen::setNbSmeUnits()): on macOS from the performance clusters, and 0 (no cap) elsewhere. Define it on other systems whose SME units are shared by a cluster of cores, such as Linux on Apple silicon.EIGEN_SME_MIN_TASK_SIZE - Minimum work per thread, in FP32 multiply-adds, before a product on the ARM SME GEMM kernel is split across threads. Other types count their kernel time in FP32 multiply-adds: double and complex<float> four each, complex<double> sixteen. The default (2^20) is tuned for Apple M4 and applies where the SME unit count is known (see EIGEN_SME_UNITS). There, when the maximum sizes of the result and of the inner dimension are not all fixed at compile time, each thread computes a disjoint part of the result.EIGEN_SME_DIRECT_LHS_MAX_STRIDE_BYTES and EIGEN_SME_DIRECT_LHS_MAX_BLOCK_BYTES - Largest column stride and largest block (rows times depth, both in bytes) at which the ARM SME GEMM kernel reads a column-major real LHS block from its source instead of a packed panel, whatever its alignment. The defaults are 16 KB and 4 MB; strides that are a multiple of 4 KB always pack, since they alias in L1. 0 for either disables the in-place read. These limits do not apply when the RHS is a single panel (see EIGEN_SME_DIRECT_LHS_MAX_SPAN_BYTES).EIGEN_SME_DIRECT_LHS_MAX_SPAN_BYTES - When the RHS is a single panel, each LHS element is read once, so the ARM SME GEMM kernel reads a column-major real LHS in place: always up to half a panel wide, and up to a full panel while depth times column stride (in bytes) stays within this limit, past which the panel spans more pages than the TLB covers. The default is 16 MB; 0 disables the in-place read for single-panel products.EIGEN_ARM64_NO_SME_F64F64 - Disables the double-precision ARM SME kernels, which need the optional FEAT_SME_F64F64 extension. Eigen enables it whenever the compiler reports that extension; define this macro when the build target is wider than the run target, so that double keeps the generic kernels for matrix products, matrix-vector products, dot products, and scaled vector additions.EIGEN_SME_PACKED_RHS_BUDGET_BYTES - Byte budget for the packed RHS panel in the ARM SME GEMM kernel. The default (32 MB) is empirically tuned for Apple M4.EIGEN_SME_LHS_WORKING_SET_BUDGET_BYTES - Byte budget for the LHS working set in the ARM SME GEMM kernel. The default (7 MB) is empirically tuned for Apple M4.EIGEN_SME_SINGLE_PASS_RHS_BUDGET_BYTES - Byte budget for the packed RHS panel in the ARM SME GEMM kernel when all rows of the product fit one LHS block, so that each packed panel is read once and should stay in L2 between packing and the kernel. The default (4 MB) is empirically tuned for Apple M4 (16 MB L2 per cluster).EIGEN_DONT_ALIGN - Deprecated, it is a synonym for EIGEN_MAX_ALIGN_BYTES=0. It disables alignment completely. Eigen will not try to align its objects and does not expect that any objects passed to it are aligned. This will turn off vectorization if EIGEN_UNALIGNED_VECTORIZE=1. Not defined by default.EIGEN_DONT_ALIGN_STATICALLY - Deprecated, it is a synonym for EIGEN_MAX_STATIC_ALIGN_BYTES=0. It disables alignment of arrays on the stack. Not defined by default, unless EIGEN_DONT_ALIGN is defined.EIGEN_ALTIVEC_ENABLE_MMA_DYNAMIC_DISPATCH - Controls whether to use Eigen's dynamic dispatching for Altivec MMA or not.EIGEN_ALTIVEC_DISABLE_MMA - Overrides the usage of Altivec MMA instructions.EIGEN_ALTIVEC_USE_CUSTOM_PACK - Controls whether to use Eigen's custom packing for Altivec or not.It is possible to add new methods to many fundamental classes in Eigen by writing a plugin. As explained in the section Extending MatrixBase (and other classes), the plugin is specified by defining a EIGEN_xxx_PLUGIN macro. The following macros are supported; none of them are defined by default.
These macros are mainly meant for people developing Eigen and for testing purposes. Even though, they might be useful for power users and the curious for debugging and testing purpose, they should not be used by real-word code.
set_is_malloc_allowed(bool). If malloc is not allowed and Eigen tries to allocate memory dynamically anyway, an assertion failure results. Not defined by default.