Beaver. MLIR. Dialect. X86
(beaver v0.4.8)
Copy Markdown
Summary
Functions
Return op name x86.amx.tile_load as a bitstring.
x86.amx.tile_load - tile load operation
Return op name x86.amx.tile_mulf as a bitstring.
x86.amx.tile_mulf - tile multiplication operation (floating-point)
Return op name x86.amx.tile_muli as a bitstring.
x86.amx.tile_muli - tile multiplication operation (integer)
Return op name x86.amx.tile_store as a bitstring.
x86.amx.tile_store - tile store operation
Return op name x86.amx.tile_zero as a bitstring.
x86.amx.tile_zero - tile zero operation
Return op name x86.avx10.dot.i8 as a bitstring.
x86.avx10.dot.i8 - AVX10 Dot Int8 op
Return op name x86.avx512.cvt.packed.f32_to_bf16 as a bitstring.
x86.avx512.cvt.packed.f32_to_bf16 - Convert packed F32 to packed BF16 Data.
Return op name x86.avx512.dot as a bitstring.
x86.avx512.dot - Dot BF16 op
Return op name x86.avx512.mask.compress as a bitstring.
x86.avx512.mask.compress - Masked compress op
Return op name x86.avx512.mask.rndscale as a bitstring.
x86.avx512.mask.rndscale - Masked roundscale op
Return op name x86.avx512.mask.scalef as a bitstring.
x86.avx512.mask.scalef - ScaleF op
Return op name x86.avx512.vp2intersect as a bitstring.
x86.avx512.vp2intersect - Vp2Intersect op
Return op name x86.avx.bcst_to_f32.packed as a bitstring.
x86.avx.bcst_to_f32.packed - AVX: Broadcasts BF16/F16 into packed F32 Data.
Return op name x86.avx.cvt.packed.even.indexed_to_f32 as a bitstring.
x86.avx.cvt.packed.even.indexed_to_f32 - AVX: Convert packed BF16/F16 even-indexed elements into packed F32 Data.
Return op name x86.avx.cvt.packed.odd.indexed_to_f32 as a bitstring.
x86.avx.cvt.packed.odd.indexed_to_f32 - AVX: Convert packed BF16/F16 odd-indexed elements into packed F32 Data.
Return op name x86.avx.dot.i8 as a bitstring.
x86.avx.dot.i8 - Dot Int8 op
Return op name x86.avx.intr.dot as a bitstring.
x86.avx.intr.dot - Dot
Return op name x86.avx.rsqrt as a bitstring.
x86.avx.rsqrt - Rsqrt
Functions
Return op name x86.amx.tile_load as a bitstring.
x86.amx.tile_load - tile load operation
Operands
base- Single,AnyMemRef, memref of any non-token type valuesindices- Variadic,Index, variadic of indexstride- Optional,Index, index
Results
res- Single,AnyAMXTile, tile of 32-bit float or 16-bit float or bfloat16 type or 32-bit signless integer or 8-bit signless integer or f8E4M3FN type or f8E5M2 type values
Description
Loads a tile from memory defined by a base and indices, with the
shape defined by the 2-dim vector type of the result.
The tile's rows are populated by reading contiguous elements starting
at the base. For each tile row, the base is incremented by stride
number of elements.
The tile is loaded using the following indexing scheme:
for row in enumerate(tile_rows):
mem_row = base[i0, i1, ..., iN + row * stride]
for col in enumerate(tile_cols):
tile[row, col] = mem_row[col]If the stride is not provided, then the base buffer must be at least
2-dimensional, and the stride is automatically inferred and corresponds
to the stride of the buffer's second innermost dimension.
The operation is eventually lowered into the "tileloadd" instruction with the corresponding tile configuration.
With the write memory effect, each x86.amx.tile_load operation serves as
a compilation hint to use a separate tile register.
Example:
// Tile load from a 2-D memref with implicit stride.
%0 = x86.amx.tile_load %arg0[%c0, %c0] : memref<?x?xi8> into !x86.amx.tile<16x64xi8>
// Tile load from a 1-D memref with explicit stride.
%0 = x86.amx.tile_load %arg0[%c0], %stride : memref<?xi8> into !x86.amx.tile<16x64xi8>
Return op name x86.amx.tile_mulf as a bitstring.
x86.amx.tile_mulf - tile multiplication operation (floating-point)
This op has support for result type inference.
Operands
lhs- Single,AMXTileF16OrBF16OrF8, tile of 16-bit float or bfloat16 type or f8E4M3FN type or f8E5M2 type valuesrhs- Single,AMXTileF16OrBF16OrF8, tile of 16-bit float or bfloat16 type or f8E4M3FN type or f8E5M2 type valuesacc- Single,AMXTileF32, tile of 32-bit float values
Results
res- Single,AMXTileF32, tile of 32-bit float values
Description
Multiplies a "m x k" tile with a "k x n" tile and accumulates the results into a "m x n" destination tile. Supports "f32 <- bf16 x bf16" (with pairs of "bf16") and "f32 <- f8E5M2/f8E4M3FN x f8E5M2/f8E4M3FN".
The operation is eventually lowered into the "tdpbf16ps/tdpbf8ps/tdpbhf8ps/ tdphbf8ps/tdphf8ps" instruction with the corresponding tile configuration.
Example:
%0 = x86.amx.tile_mulf %a, %b, %c
: !x86.amx.tile<16x32xbf16>, !x86.amx.tile<16x32xbf16>, !x86.amx.tile<16x16xf32>
Return op name x86.amx.tile_muli as a bitstring.
x86.amx.tile_muli - tile multiplication operation (integer)
This op has support for result type inference.
Attributes
isZextLhs- Optional,UnitAttr, unit attributeisZextRhs- Optional,UnitAttr, unit attribute
Operands
lhs- Single,AMXTileI8, tile of 8-bit signless integer valuesrhs- Single,AMXTileI8, tile of 8-bit signless integer valuesacc- Single,AMXTileI32, tile of 32-bit signless integer values
Results
res- Single,AMXTileI32, tile of 32-bit signless integer values
Description
Multiplies a "m x k" tile with a "k x n" tile and accumulates the results into a "m x n" destination tile. Supports all "si32 <- s/ui8 x s/ui8" combinations (4 bytes packed into dwords in the columns of both the source operand tiles; the zero or sign extension is specified with the attributes and default to sign extended).
The operation is eventually lowered into one of the "tdpbssd", "tdpbsud", "tdpbusd", or "tdpbuud" instructions with the corresponding tile configuration.
Example:
%0 = x86.amx.tile_muli %a zext, %b zext, %c
: !x86.amx.tile<16x64xi8>, !x86.amx.tile<16x64xi8>, !x86.amx.tile<16x16xi32>
Return op name x86.amx.tile_store as a bitstring.
x86.amx.tile_store - tile store operation
Operands
base- Single,AnyMemRef, memref of any non-token type valuesindices- Variadic,Index, variadic of indexval- Single,AnyAMXTile, tile of 32-bit float or 16-bit float or bfloat16 type or 32-bit signless integer or 8-bit signless integer or f8E4M3FN type or f8E5M2 type valuesstride- Optional,Index, index
Description
Stores a tile to memory defined by a base and indices, with the
shape defined by the 2-dim vector type of the value.
The tile's rows are written contiguously to the buffer starting at
the base. For each tile row, the base is incremented by stride
number of elements.
The tile is stored using the following indexing scheme:
for row in enumerate(tile_rows):
mem_row = base[i0, i1, ..., iN + row * stride]
for col in enumerate(tile_cols):
mem_row[col] = tile[row, col]If the stride is not provided, then the base buffer must be at least
2-dimensional, and the stride is automatically inferred and corresponds
to the stride of the buffer's second innermost dimension.
The operation is eventually lowered into the "tilestored" instruction with the corresponding tile configuration.
Example:
// Tile store to a 2-D memref with implicit stride.
x86.amx.tile_store %arg1[%c0, %c0], %0 : memref<?x?xi8>, !x86.amx.tile<16x64xi8>
// Tile store to a 1-D memref with explicit stride.
x86.amx.tile_store %arg1[%c0], %0, %stride : memref<?xi8>, !x86.amx.tile<16x64xi8>
Return op name x86.amx.tile_zero as a bitstring.
x86.amx.tile_zero - tile zero operation
Results
res- Single,AnyAMXTile, tile of 32-bit float or 16-bit float or bfloat16 type or 32-bit signless integer or 8-bit signless integer or f8E4M3FN type or f8E5M2 type values
Description
Zeroes the destination tile, with the shape defined by the 2-dim vector type of the result.
The operation is eventually lowered into the "tilezero" instruction with the corresponding tile configuration.
With the write memory effect, each x86.amx.tile_zero operation serves as
a compilation hint to use a separate tile register.
Example:
%0 = x86.amx.tile_zero : !x86.amx.tile<16x16xbf16>
Return op name x86.avx10.dot.i8 as a bitstring.
x86.avx10.dot.i8 - AVX10 Dot Int8 op
This op has support for result type inference.
Operands
w- Single, anonymous/composite constraint, vector of 32-bit signless integer values of length 16a- Single, anonymous/composite constraint, vector of 8-bit signless integer values of length 64b- Single, anonymous/composite constraint, vector of 8-bit signless integer values of length 64
Results
dst- Single, anonymous/composite constraint, vector of 32-bit signless integer values of length 16
Description
The dot op is an AVX10-Int8 specific op that can lower to the proper
LLVMAVX10-INT8 operation llvm.vpdpbssd.512.
Multiply groups of 4 adjacent pairs of signed 8-bit integers in a with
corresponding signed 8-bit integers in b, producing 4 intermediate signed 16-bit
results. Sum these 4 results with the corresponding 32-bit integer in w, and
store the packed 32-bit results in dst.
Example:
%dst = x86.avx10.dot.i8 %w, %a, %b : vector<64xi8> -> vector<16xi32>
Return op name x86.avx512.cvt.packed.f32_to_bf16 as a bitstring.
x86.avx512.cvt.packed.f32_to_bf16 - Convert packed F32 to packed BF16 Data.
Operands
a- Single, anonymous/composite constraint, vector of 32-bit float values of length 8/16
Results
dst- Single, anonymous/composite constraint, vector of bfloat16 type values of length 8/16
Description
The convert_f32_to_bf16 op is an AVX512-BF16 specific op that can lower
to the proper LLVMAVX512BF16 operation llvm.cvtneps2bf16 depending on
the width of MLIR vectors it is applied to.
From the Intel Intrinsics Guide:
Convert packed single-precision (32-bit) floating-point elements in a to
packed BF16 (16-bit) floating-point elements, and store the results in dst.
Example:
%dst = x86.avx512.cvt.packed.f32_to_bf16 %a : vector<8xf32> -> vector<8xbf16>
Return op name x86.avx512.dot as a bitstring.
x86.avx512.dot - Dot BF16 op
This op has support for result type inference.
Operands
src- Single, anonymous/composite constraint, vector of 32-bit float values of length 4/8/16a- Single, anonymous/composite constraint, vector of bfloat16 type values of length 8/16/32b- Single, anonymous/composite constraint, vector of bfloat16 type values of length 8/16/32
Results
dst- Single, anonymous/composite constraint, vector of 32-bit float values of length 4/8/16
Description
The dot op is an AVX512-BF16 specific op that can lower to the proper
LLVMAVX512BF16 operation llvm.dpbf16ps depending on the width of MLIR
vectors it is applied to.
From the Intel Intrinsics Guide:
Compute dot-product of BF16 (16-bit) floating-point pairs in a and b,
accumulating the intermediate single-precision (32-bit) floating-point
elements with elements in src, and store the results in dst.
Example:
%dst = x86.avx512.dot %src, %a, %b : vector<32xbf16> -> vector<16xf32>
Return op name x86.avx512.mask.compress as a bitstring.
x86.avx512.mask.compress - Masked compress op
This op has support for result type inference.
Attributes
constant_src- Optional,ElementsAttr, constant vector/tensor attribute
Operands
k- Single, anonymous/composite constraint, vector of 1-bit signless integer values of length 16/8a- Single, anonymous/composite constraint, vector of 32-bit float or 32-bit signless integer or 64-bit float or 64-bit signless integer values of length 16/8src- Optional, anonymous/composite constraint, vector of 32-bit float or 32-bit signless integer or 64-bit float or 64-bit signless integer values of length 16/8
Results
dst- Single, anonymous/composite constraint, vector of 32-bit float or 32-bit signless integer or 64-bit float or 64-bit signless integer values of length 16/8
Description
The mask.compress op is an AVX512 specific op that can lower to the
llvm.mask.compress instruction. Instead of src, a constant vector
vector attribute constant_src may be specified. If neither src nor
constant_src is specified, the remaining elements in the result vector are
set to zero.
From the Intel Intrinsics Guide:
Contiguously store the active integer/floating-point elements in a (those
with their respective bit set in writemask k) to dst, and pass through the
remaining elements from src.
Return op name x86.avx512.mask.rndscale as a bitstring.
x86.avx512.mask.rndscale - Masked roundscale op
This op has support for result type inference.
Operands
src- Single, anonymous/composite constraint, vector of 32-bit float or 64-bit float values of length 16/8k- Single,I32, 32-bit signless integera- Single, anonymous/composite constraint, vector of 32-bit float or 64-bit float values of length 16/8imm- Single, anonymous/composite constraint, 16-bit signless integer or 8-bit signless integerrounding- Single,I32, 32-bit signless integer
Results
dst- Single, anonymous/composite constraint, vector of 32-bit float or 64-bit float values of length 16/8
Description
The mask.rndscale op is an AVX512 specific op that can lower to the proper
LLVMAVX512 operation: llvm.mask.rndscale.ps.512 or
llvm.mask.rndscale.pd.512 instruction depending on the type of vectors it
is applied to.
From the Intel Intrinsics Guide:
Round packed floating-point elements in a to the number of fraction bits
specified by imm, and store the results in dst using writemask k
(elements are copied from src when the corresponding mask bit is not set).
Return op name x86.avx512.mask.scalef as a bitstring.
x86.avx512.mask.scalef - ScaleF op
This op has support for result type inference.
Operands
src- Single, anonymous/composite constraint, vector of 32-bit float or 64-bit float values of length 16/8a- Single, anonymous/composite constraint, vector of 32-bit float or 64-bit float values of length 16/8b- Single, anonymous/composite constraint, vector of 32-bit float or 64-bit float values of length 16/8k- Single, anonymous/composite constraint, 16-bit signless integer or 8-bit signless integerrounding- Single,I32, 32-bit signless integer
Results
dst- Single, anonymous/composite constraint, vector of 32-bit float or 64-bit float values of length 16/8
Description
The mask.scalef op is an AVX512 specific op that can lower to the proper
LLVMAVX512 operation: llvm.mask.scalef.ps.512 or
llvm.mask.scalef.pd.512 depending on the type of MLIR vectors it is
applied to.
From the Intel Intrinsics Guide:
Scale the packed floating-point elements in a using values from b, and
store the results in dst using writemask k (elements are copied from src
when the corresponding mask bit is not set).
Return op name x86.avx512.vp2intersect as a bitstring.
x86.avx512.vp2intersect - Vp2Intersect op
This op has support for result type inference.
Operands
a- Single, anonymous/composite constraint, vector of 32-bit signless integer or 64-bit signless integer values of length 16/8b- Single, anonymous/composite constraint, vector of 32-bit signless integer or 64-bit signless integer values of length 16/8
Results
k1- Single, anonymous/composite constraint, vector of 1-bit signless integer values of length 16/8k2- Single, anonymous/composite constraint, vector of 1-bit signless integer values of length 16/8
Description
The vp2intersect op is an AVX512 specific op that can lower to the proper
LLVMAVX512 operation: llvm.vp2intersect.d.512 or
llvm.vp2intersect.q.512 depending on the type of MLIR vectors it is
applied to.
From the Intel Intrinsics Guide:
Compute intersection of packed integer vectors a and b, and store
indication of match in the corresponding bit of two mask registers
specified by k1 and k2. A match in corresponding elements of a and
b is indicated by a set bit in the corresponding bit of the mask
registers.
Return op name x86.avx.bcst_to_f32.packed as a bitstring.
x86.avx.bcst_to_f32.packed - AVX: Broadcasts BF16/F16 into packed F32 Data.
Operands
a- Single, anonymous/composite constraint, memref of bfloat16 type or 16-bit float values
Results
dst- Single, anonymous/composite constraint, vector of 32-bit float values of length 4/8
Description
From the Intel Intrinsics Guide:
Convert scalar BF16 or F16 (16-bit) floating-point element stored at memory locations
starting at location __A to a single-precision (32-bit) floating-point,
broadcast it to packed single-precision (32-bit) floating-point elements,
and store the results in dst.
Example:
%dst = x86.avx.bcst_to_f32.packed %a : memref<1xbf16> -> vector<8xf32>
%dst = x86.avx.bcst_to_f32.packed %a : memref<1xf16> -> vector<8xf32>
Return op name x86.avx.cvt.packed.even.indexed_to_f32 as a bitstring.
x86.avx.cvt.packed.even.indexed_to_f32 - AVX: Convert packed BF16/F16 even-indexed elements into packed F32 Data.
Operands
a- Single, anonymous/composite constraint, memref of bfloat16 type or 16-bit float values
Results
dst- Single, anonymous/composite constraint, vector of 32-bit float values of length 4/8
Description
From the Intel Intrinsics Guide:
Convert packed BF16 or F16 (16-bit) floating-point even-indexed elements stored at
memory locations starting at location __A to packed single-precision
(32-bit) floating-point elements, and store the results in dst.
Example:
%dst = x86.avx.cvt.packed.even.indexed_to_f32 %a : memref<16xbf16> -> vector<8xf32>
%dst = x86.avx.cvt.packed.even.indexed_to_f32 %a : memref<16xf16> -> vector<8xf32>
Return op name x86.avx.cvt.packed.odd.indexed_to_f32 as a bitstring.
x86.avx.cvt.packed.odd.indexed_to_f32 - AVX: Convert packed BF16/F16 odd-indexed elements into packed F32 Data.
Operands
a- Single, anonymous/composite constraint, memref of bfloat16 type or 16-bit float values
Results
dst- Single, anonymous/composite constraint, vector of 32-bit float values of length 4/8
Description
From the Intel Intrinsics Guide:
Convert packed BF16 or F16 (16-bit) floating-point odd-indexed elements stored at
memory locations starting at location __A to packed single-precision
(32-bit) floating-point elements, and store the results in dst.
Example:
%dst = x86.avx.cvt.packed.odd.indexed_to_f32 %a : memref<16xbf16> -> vector<8xf32>
%dst = x86.avx.cvt.packed.odd.indexed_to_f32 %a : memref<16xf16> -> vector<8xf32>
Return op name x86.avx.dot.i8 as a bitstring.
x86.avx.dot.i8 - Dot Int8 op
This op has support for result type inference.
Operands
w- Single, anonymous/composite constraint, vector of 32-bit signless integer values of length 4/8a- Single, anonymous/composite constraint, vector of 8-bit signless integer values of length 16/32b- Single, anonymous/composite constraint, vector of 8-bit signless integer values of length 16/32
Results
dst- Single, anonymous/composite constraint, vector of 32-bit signless integer values of length 4/8
Description
The dot op is an AVX2-Int8 specific op that can lower to the proper
LLVMAVX2-INT8 operation llvm.vpdpbssd depending on the width of MLIR
vectors it is applied to.
From the Intel Intrinsics Guide:
Multiply groups of 4 adjacent pairs of signed 8-bit integers in a with
corresponding signed 8-bit integers in b, producing 4 intermediate signed 16-bit
results. Sum these 4 results with the corresponding 32-bit integer in w, and
store the packed 32-bit results in dst.
Example:
%dst = x86.avx.dot.i8 %w, %a, %b : vector<32xi8> -> vector<8xi32>
Return op name x86.avx.intr.dot as a bitstring.
x86.avx.intr.dot - Dot
This op has support for result type inference.
Operands
a- Single, anonymous/composite constraint, vector of 32-bit float values of length 8b- Single, anonymous/composite constraint, vector of 32-bit float values of length 8
Results
res- Single, anonymous/composite constraint, vector of 32-bit float values of length 8
Description
Computes the 4-way dot products of the lower and higher parts of the source vectors and broadcasts the two results to the lower and higher elements of the destination vector, respectively. Adding one element of the lower part to one element of the higher part in the destination vector yields the full dot product of the two source vectors.
Example:
%0 = x86.avx.intr.dot %a, %b : vector<8xf32>
%1 = vector.extract %0[%i0] : f32 from vector<8xf32>
%2 = vector.extract %0[%i4] : f32 from vector<8xf32>
%d = arith.addf %1, %2 : f32
Return op name x86.avx.rsqrt as a bitstring.
x86.avx.rsqrt - Rsqrt
This op has support for result type inference.
Operands
a- Single, anonymous/composite constraint, vector of 32-bit float values of length 8
Results
b- Single, anonymous/composite constraint, vector of 32-bit float values of length 8