Eigen  5.0.1
 
Loading...
Searching...
No Matches
Vectorization

Eigen performs explicit vectorization: instead of relying on the compiler to auto-vectorize scalar loops, Eigen's evaluators generate SIMD instructions directly through a portable wrapper layer (the "packet math" layer). This page gives an overview of which instruction sets are supported, how vectorization is enabled, what gets vectorized, and how to control it.

Supported instruction sets

On CPUs, Eigen provides vectorized kernels for the following instruction sets:

ArchitectureInstruction setsNotes
x86 / x86-64SSE2, SSE3, SSSE3, SSE4.1, SSE4.2, AVX, AVX2+FMA, AVX512 (incl. the DQ, VL, FP16, and BF16 extensions) Selected from the compiler's target flags (e.g. -mavx2 -mfma, /arch:AVX2)
ARM / AArch64NEON, SME; SVE as an opt-in backend SVE requires EIGEN_ARM64_USE_SVE and a fixed vector length (-msve-vector-bits=N). The SME backend accelerates real and complex matrix products as well as large real dot products, scaled vector additions and column-major matrix-vector products. It is selected automatically when the compiler targets SME2 (-march=armv9.2-a+sme2, or a -mcpu that implies it) without -msve-vector-bits, since a fixed length would pin the kernels to one runtime streaming vector length; everything else keeps the NEON path. Define EIGEN_ARM64_NO_SME to opt out, or EIGEN_ARM64_USE_SME to turn a toolchain that cannot provide it into a build error instead of a NEON fallback. Double precision additionally needs the optional FEAT_SME_F64F64 extension (+sme-f64f64, or a -mcpu that implies it); without it double and std::complex<double> keep the generic kernels. Floating-point exceptions raised in the streaming SME kernels (matrix products, and the vector and matrix-vector kernels) are not reported: FMOPA sets no flags, and the caller's flags are restored on exit, discarding those of the kernels' other arithmetic. Products on the NEON paths report exceptions as usual
PowerPCAltiVec, VSX, MMA
IBM Z (s390x)ZVector
MIPSMSA
LoongArchLSX
RISC-VRVV 1.0 Requires EIGEN_RISCV64_USE_RVV10 and a fixed vector length (-mrvv-vector-bits=zvl)
Qualcomm HexagonHVX

In addition, the same packet abstraction is used to generate efficient device code for CUDA and HIP GPUs (see Using Eigen in CUDA kernels) and for SYCL devices.

Vectorization is enabled automatically whenever the compiler's target flags advertise one of the supported instruction sets. You can check what a given translation unit ended up using by printing Eigen::SimdInstructionSetsInUse():

std::cout << Eigen::SimdInstructionSetsInUse() << std::endl;

What gets vectorized

Vectorization applies across Eigen, most notably to:

  • assignment of coefficient-wise expressions (e.g. a = 2*b + c.cwiseProduct(d)), including the fused evaluation loops described in Expression templates in Eigen;
  • reductions such as sum(), minCoeff(), squaredNorm(), and their partial (colwise/rowwise) variants;
  • matrix-matrix and matrix-vector products, and through them most decompositions;
  • many coefficient-wise math functions; the math function catalogue lists, for every function, the instruction sets and scalar types with a vectorized implementation.

The scalar types with SIMD support depend on the instruction set; float, double, 32- and 64-bit integers, and std::complex<float>/std::complex<double> are covered on the main backends, with Eigen::half and Eigen::bfloat16 vectorized where hardware support exists (e.g. AVX512FP16, NEON). Types without SIMD support, such as long double or custom scalar types, simply use scalar code.

Fixed-size objects are vectorized when their size is a multiple of the packet size; this is where the alignment rules for fixed-size vectorizable types come from. Dynamic-size objects use aligned heap allocation plus vectorized loops with scalar prologues and epilogues where needed. Unaligned loads and stores are used where alignment cannot be established; this is enabled by default and controlled by EIGEN_UNALIGNED_VECTORIZE.

Controlling vectorization

Vectorization is meant to be transparent, but the following macros, all documented in Preprocessor directives, let you control it:

  • EIGEN_DONT_VECTORIZE disables explicit vectorization entirely.
  • EIGEN_MAX_ALIGN_BYTES and EIGEN_MAX_STATIC_ALIGN_BYTES bound the alignment Eigen requests for heap and static data respectively; note that changing them affects the ABI.
  • EIGEN_UNALIGNED_VECTORIZE=0 restricts vectorization to expressions whose destination is aligned.
  • EIGEN_FAST_MATH selects between faster and more accurate vectorized math functions.

Disabling vectorization also drops Eigen's alignment: outside GPU compilation, EIGEN_DONT_VECTORIZE sets the ideal alignment to zero, and the defaults for EIGEN_MAX_STATIC_ALIGN_BYTES and EIGEN_MAX_ALIGN_BYTES follow it. Fixed-size vectorizable types then lose their over-alignment — alignof(Vector4f) drops from 16 to 4 on a typical x86-64 build — so translation units that disagree about EIGEN_DONT_VECTORIZE disagree about the ABI. Define the alignment macros explicitly and identically everywhere if you need the boundary to stay fixed. See Explanation of the assertion on unaligned arrays for what the alignment is there for.

When compiling CUDA sources, host-side SIMD must be disabled; see Using Eigen in CUDA kernels.

Under the hood

Vectorized kernels are written in terms of generic packet primitives — pload(), pstore(), padd(), pmul(), predux(), and friends — declared in Eigen/src/Core/GenericPacketMath.h and specialized for each instruction set under Eigen/src/Core/arch/. Expression evaluators query each expression's packet support and alignment at compile time to choose between scalar, linear, sliced, and unrolled vectorized evaluation loops. Custom functors can participate in vectorization by providing a packetOp() and advertising PacketAccess in their traits; see Writing custom functors.