Beaver. MLIR. Dialect. ArmSME
(beaver v0.4.8)
Copy Markdown
Summary
Functions
Return op name arm_sme.copy_tile as a bitstring.
arm_sme.copy_tile - Copies an SME tile value
Return op name arm_sme.extract_tile_slice as a bitstring.
arm_sme.extract_tile_slice - Extract 1-D scalable vector from slice of 2-D tile
Return op name arm_sme.fmopa_2way as a bitstring.
arm_sme.fmopa_2way - Floating-point sum of 2 outer products and accumulate
Return op name arm_sme.fmops_2way as a bitstring.
arm_sme.fmops_2way - Floating-point sum of 2 outer products and subtract
Return op name arm_sme.get_tile as a bitstring.
arm_sme.get_tile - Creates an undefined value of SME virtual tile type
Return op name arm_sme.insert_tile_slice as a bitstring.
arm_sme.insert_tile_slice - Insert 1-D scalable vector into slice of 2-D tile
Return op name arm_sme.intr.cntsd as a bitstring.
arm_sme.intr.cntsd
Return op name arm_sme.intr.ld1b.horiz as a bitstring.
arm_sme.intr.ld1b.horiz
Return op name arm_sme.intr.ld1b.vert as a bitstring.
arm_sme.intr.ld1b.vert
Return op name arm_sme.intr.ld1d.horiz as a bitstring.
arm_sme.intr.ld1d.horiz
Return op name arm_sme.intr.ld1d.vert as a bitstring.
arm_sme.intr.ld1d.vert
Return op name arm_sme.intr.ld1h.horiz as a bitstring.
arm_sme.intr.ld1h.horiz
Return op name arm_sme.intr.ld1h.vert as a bitstring.
arm_sme.intr.ld1h.vert
Return op name arm_sme.intr.ld1q.horiz as a bitstring.
arm_sme.intr.ld1q.horiz
Return op name arm_sme.intr.ld1q.vert as a bitstring.
arm_sme.intr.ld1q.vert
Return op name arm_sme.intr.ld1w.horiz as a bitstring.
arm_sme.intr.ld1w.horiz
Return op name arm_sme.intr.ld1w.vert as a bitstring.
arm_sme.intr.ld1w.vert
Return op name arm_sme.intr.mopa as a bitstring.
arm_sme.intr.mopa
Return op name arm_sme.intr.mopa.wide as a bitstring.
arm_sme.intr.mopa.wide
Return op name arm_sme.intr.mops as a bitstring.
arm_sme.intr.mops
Return op name arm_sme.intr.mops.wide as a bitstring.
arm_sme.intr.mops.wide
Return op name arm_sme.intr.read.horiz as a bitstring.
arm_sme.intr.read.horiz
Return op name arm_sme.intr.read.vert as a bitstring.
arm_sme.intr.read.vert
Return op name arm_sme.intr.smopa.wide as a bitstring.
arm_sme.intr.smopa.wide
Return op name arm_sme.intr.smopa.za32 as a bitstring.
arm_sme.intr.smopa.za32
Return op name arm_sme.intr.smops.wide as a bitstring.
arm_sme.intr.smops.wide
Return op name arm_sme.intr.smops.za32 as a bitstring.
arm_sme.intr.smops.za32
Return op name arm_sme.intr.st1b.horiz as a bitstring.
arm_sme.intr.st1b.horiz
Return op name arm_sme.intr.st1b.vert as a bitstring.
arm_sme.intr.st1b.vert
Return op name arm_sme.intr.st1d.horiz as a bitstring.
arm_sme.intr.st1d.horiz
Return op name arm_sme.intr.st1d.vert as a bitstring.
arm_sme.intr.st1d.vert
Return op name arm_sme.intr.st1h.horiz as a bitstring.
arm_sme.intr.st1h.horiz
Return op name arm_sme.intr.st1h.vert as a bitstring.
arm_sme.intr.st1h.vert
Return op name arm_sme.intr.st1q.horiz as a bitstring.
arm_sme.intr.st1q.horiz
Return op name arm_sme.intr.st1q.vert as a bitstring.
arm_sme.intr.st1q.vert
Return op name arm_sme.intr.st1w.horiz as a bitstring.
arm_sme.intr.st1w.horiz
Return op name arm_sme.intr.st1w.vert as a bitstring.
arm_sme.intr.st1w.vert
Return op name arm_sme.intr.str as a bitstring.
arm_sme.intr.str
Return op name arm_sme.intr.sumopa.wide as a bitstring.
arm_sme.intr.sumopa.wide
Return op name arm_sme.intr.sumops.wide as a bitstring.
arm_sme.intr.sumops.wide
Return op name arm_sme.intr.umopa.wide as a bitstring.
arm_sme.intr.umopa.wide
Return op name arm_sme.intr.umopa.za32 as a bitstring.
arm_sme.intr.umopa.za32
Return op name arm_sme.intr.umops.wide as a bitstring.
arm_sme.intr.umops.wide
Return op name arm_sme.intr.umops.za32 as a bitstring.
arm_sme.intr.umops.za32
Return op name arm_sme.intr.usmopa.wide as a bitstring.
arm_sme.intr.usmopa.wide
Return op name arm_sme.intr.usmops.wide as a bitstring.
arm_sme.intr.usmops.wide
Return op name arm_sme.intr.write.horiz as a bitstring.
arm_sme.intr.write.horiz
Return op name arm_sme.intr.write.vert as a bitstring.
arm_sme.intr.write.vert
Return op name arm_sme.intr.zero as a bitstring.
arm_sme.intr.zero
Return op name arm_sme.load_tile_slice as a bitstring.
arm_sme.load_tile_slice - Tile slice load and update operation
Return op name arm_sme.outerproduct as a bitstring.
arm_sme.outerproduct - Outer product with optional fused add/sub
Return op name arm_sme.smopa_2way as a bitstring.
arm_sme.smopa_2way - Signed integer sum of 2 outer products and accumulate
Return op name arm_sme.smopa_4way as a bitstring.
arm_sme.smopa_4way - Signed integer sum of 4 outer products and accumulate
Return op name arm_sme.smops_2way as a bitstring.
arm_sme.smops_2way - Signed integer sum of 2 outer products and subtract
Return op name arm_sme.smops_4way as a bitstring.
arm_sme.smops_4way - Signed integer sum of 4 outer products and subtract
Return op name arm_sme.store_tile_slice as a bitstring.
arm_sme.store_tile_slice - Tile slice store operation
Return op name arm_sme.streaming_vl as a bitstring.
arm_sme.streaming_vl - Query the streaming vector length
Return op name arm_sme.sumopa_4way as a bitstring.
arm_sme.sumopa_4way - Signed by unsigned integer sum of 4 outer products and accumulate
Return op name arm_sme.sumops_4way as a bitstring.
arm_sme.sumops_4way - Signed by unsigned integer sum of 4 outer products and subtract
Return op name arm_sme.tile_load as a bitstring.
arm_sme.tile_load - Tile load operation
Return op name arm_sme.tile_store as a bitstring.
arm_sme.tile_store - Tile store operation
Return op name arm_sme.umopa_2way as a bitstring.
arm_sme.umopa_2way - Unsiged integer sum of 2 outer products and accumulate
Return op name arm_sme.umopa_4way as a bitstring.
arm_sme.umopa_4way - Unsigned integer sum of 4 outer products and accumulate
Return op name arm_sme.umops_2way as a bitstring.
arm_sme.umops_2way - Unsiged integer sum of 2 outer products and subtract
Return op name arm_sme.umops_4way as a bitstring.
arm_sme.umops_4way - Unsigned integer sum of 4 outer products and subtract
Return op name arm_sme.usmopa_4way as a bitstring.
arm_sme.usmopa_4way - Unsigned by signed integer sum of 4 outer products and accumulate
Return op name arm_sme.usmops_4way as a bitstring.
arm_sme.usmops_4way - Unsigned by signed integer sum of 4 outer products and subtract
Return op name arm_sme.zero as a bitstring.
arm_sme.zero - Creates a zero-initialized value of SME virtual tile type
Functions
Return op name arm_sme.copy_tile as a bitstring.
arm_sme.copy_tile - Copies an SME tile value
This op has support for result type inference.
Operands
tile- Single,SMETile, a vector type that fits into a SME tile
Results
result- Single,SMETile, a vector type that fits into a SME tile
Description
Copies an SME "virtual tile" value to a new SSA value. This operation is primarily intended to be used to normalize the IR prior to tile allocation.
Example:
%copy = arm_sme.copy_tile %tile : vector<[4]x[4]xf32>
Return op name arm_sme.extract_tile_slice as a bitstring.
arm_sme.extract_tile_slice - Extract 1-D scalable vector from slice of 2-D tile
This op has support for result type inference.
Attributes
layout- Single,ArmSME_TileSliceLayoutAttr, Layout of a tile slice
Operands
tile- Single,SMETile, a vector type that fits into a SME tiletile_slice_index- Single,Index, index
Results
result- Single,SVEVector, a vector type that matches the size of a SVE vector
Description
Extracts a 1-D scalable slice from a 2-D scalable tile at the given index. A tile slice is a 1-D vector of horizontally or vertically contiguous elements within a ZA tile.
An optional tile slice layout attribute specifies whether the tile slice is horizontal (default) or vertical.
Example 1: Extract vector<[16]xi8> from tile horizontally at the given index.
%slice = arm_sme.extract_tile_slice %tile[%tile_slice_index] : vector<[16]xi8> from vector<[16]x[16]xi8>Example 2: Extract vector<[2]xf64> from tile vertically at the given index.
%slice = arm_sme.extract_tile_slice %tile[%tile_slice_index] layout<vertical> : vector<[2]xf64> from vector<[2]x[2]xf64>
Return op name arm_sme.fmopa_2way as a bitstring.
arm_sme.fmopa_2way - Floating-point sum of 2 outer products and accumulate
Operands
lhs- Single, anonymous/composite constraint, of ranks 1scalable vector of 16-bit float or bfloat16 type values of length 8rhs- Single,AnyVectorOfNonZeroRank, vector of any non-token type valueslhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesrhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesacc- Optional,AnyVectorOfNonZeroRank, vector of any non-token type values
Results
result- Single, anonymous/composite constraint, vector<[4]x[4]xf32> of 32-bit float values
Description
This operation represents a sum of 2 widened outer products. It takes 2 1-D scalable vectors as input and a 2-D scalable vector (ZA tile) as output.
For example (fp16 to fp32):
%result = arm_sme.fmopa_2way %lhs, %rhs :
vector<[8]xf16>, vector<[8]xf16> into vector<[4]x[4]xf32>The lhs encodes a matrix of shape SVLSx2 and the rhs a matrix of
2xSVLS, where SVLS (spec [1], section B2.1) is the number of 32-bit
elements in a vector of SVL bits. To illustrate, below is a breakdown of
this operation for fp16 to fp32, SVL=128 (i.e., vscale=1):
LHS RHS
[A0 A1 A2 A3 A4 A5 A6 A7] [B0 B1 B2 B3 B4 B5 B6 B7]
----------------------------------------------------------------------------
implicit layout
[A0 A1] |
[A2 A3] | [B0 B2 B4 B6]
[A4 A5] | [B1 B3 B5 B7]
[A6 A7] |
----------------------------------------------------------------------------
2 outer products
Acol0 ⊗ Brow0 | Acol1 ⊗ Brow1
------------- | -------------
|
[B0 B2 B4 B6] | [B1 B3 B5 B7]
|
[A0 [A0B0 A0B2 A0B4 A0B6] | [A1 [A1B1 A1B3 A1B5 A1B7]
A2 [A2B0 A2B2 A2B4 A2B6] | A3 [A3B1 A3B3 A3B5 A3B7]
A4 [A4B0 A4B2 A4B4 A4B6] | A5 [A5B1 A5B3 A5B5 A5B7]
A6] [A6B0 A6B2 A6B4 A6B6] | A7] [A7B1 A7B3 A7B5 A7B7]
|
----------------------------------------------------------------------------
sum of 2 outer products
Acol0 ⊗ Brow0 + Acol1 ⊗ Brow1
[A0B0 + A1B1 A0B2 + A1B3 A0B4 + A1B5 A0B6 + A1B7]
[A2B0 + A3B1 A2B2 + A3B3 A2B4 + A3B5 A2B6 + A3B7]
[A4B0 + A5B1 A4B2 + A5B3 A4B4 + A5B5 A4B6 + A5B7]
[A6B0 + A7B1 A6B2 + A7B3 A6B4 + A7B5 A6B6 + A7B7]
----------------------------------------------------------------------------This operation enables the folding of 2 outer products chained via the accumulator into a single outer product.
For example:
%a0_ext = arith.extf %a0 : vector<[4]xf16> to vector<[4]xf32>
%b0_ext = arith.extf %b0 : vector<[4]xf16> to vector<[4]xf32>
%a1_ext = arith.extf %a1 : vector<[4]xf16> to vector<[4]xf32>
%b1_ext = arith.extf %b1 : vector<[4]xf16> to vector<[4]xf32>
%0 = arm_sme.outerproduct %a0_ext, %b0_ext : vector<[4]xf32>, vector<[4]xf32>
%1 = arm_sme.outerproduct %a1_ext, %b1_ext acc(%0) : vector<[4]xf32>, vector<[4]xf32>The 2 outer products in the example above can be fused into a single outer product as follows:
%a_packed = vector.interleave %a0, %a1 : vector<[4]xf16> -> vector<[8]xf16>
%b_packed = vector.interleave %b0, %b1 : vector<[4]xf16> -> vector<[8]xf16>
%0 = arm_sme.fmopa_2way %a_packed, %b_packed : vector<[8]xf16>, vector<[8]xf16> into vector<[4]x[4]xf32>This is implemented in the -arm-sme-outer-product-fusion pass.
Example: FP16 to FP32
%result = arm_sme.fmopa_2way $lhs, $rhs : vector<[8]xf16>, vector<[8]xf16> into vector<[4]x[4]xf32>Example: BF16 to FP32
%result = arm_sme.fmopa_2way $lhs, $rhs : vector<[8]xbf16>, vector<[8]xbf16> into vector<[4]x[4]xf32>| Spec | Features |
|---|---|
| FMOPA (widening, 2-way, FP16 to FP32) | +sme |
| BFMOPA (widening, 2-way, BF16 to FP32) | +sme |
Return op name arm_sme.fmops_2way as a bitstring.
arm_sme.fmops_2way - Floating-point sum of 2 outer products and subtract
Operands
lhs- Single, anonymous/composite constraint, of ranks 1scalable vector of 16-bit float or bfloat16 type values of length 8rhs- Single,AnyVectorOfNonZeroRank, vector of any non-token type valueslhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesrhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesacc- Optional,AnyVectorOfNonZeroRank, vector of any non-token type values
Results
result- Single, anonymous/composite constraint, vector<[4]x[4]xf32> of 32-bit float values
Description
Equivalent to fmopa_2way but outer products are subtracted from
destination result.
Example: FP16 to FP32
%result = arm_sme.fmops_2way $lhs, $rhs : vector<[8]xf16>, vector<[8]xf16> into vector<[4]x[4]xf32>Example: BF16 to FP32
%result = arm_sme.fmops_2way $lhs, $rhs : vector<[8]xbf16>, vector<[8]xbf16> into vector<[4]x[4]xf32>Refer to fmopa_2way for a detailed description of 2-way outer products.
| Spec | Features |
|---|---|
| FMOPS (widening, 2-way, FP16 to FP32) | +sme |
| BFMOPS (widening, 2-way, BF16 to FP32) | +sme |
Return op name arm_sme.get_tile as a bitstring.
arm_sme.get_tile - Creates an undefined value of SME virtual tile type
Results
tile- Single,SMETile, a vector type that fits into a SME tile
Description
Creates a new SME "virtual tile" value within a function. The contents of the tile returned from this operation are undefined.
Example 1:
// Create an 8-bit element "virtual tile" value:
%za0_b = arm_sme.get_tile: vector<[16]x[16]xi8>Example 2:
// Create two 16-bit element "virtual tiles" values:
%za0_h = arm_sme.get_tile : vector<[8]x[8]xi16>
%za1_h = arm_sme.get_tile : vector<[8]x[8]xi16>Example 3:
// Create an 128-bit element "virtual tile" value:
%za0_q = arm_sme.get_tile : vector<[1]x[1]xi128>
Return op name arm_sme.insert_tile_slice as a bitstring.
arm_sme.insert_tile_slice - Insert 1-D scalable vector into slice of 2-D tile
This op has support for result type inference.
Attributes
layout- Single,ArmSME_TileSliceLayoutAttr, Layout of a tile slice
Operands
vector- Single,SVEVector, a vector type that matches the size of a SVE vectortile- Single,SMETile, a vector type that fits into a SME tiletile_slice_index- Single,Index, index
Results
result- Single,SMETile, a vector type that fits into a SME tile
Description
Inserts a 1-D scalable vector into a slice of a 2-D scalable vector tile at the given index. The type of the 1-D scalable vector to be inserted must match the type of the tile slice. A tile slice is a 1-D vector of horizontally or vertically contiguous elements within a ZA tile. The updated tile is returned as the result.
An optional tile slice layout attribute specifies whether the tile slice is horizontal (default) or vertical.
Example 1: Insert vector<[16]xi8> into tile horizontally at the given index.
%tile_update = arm_sme.insert_tile_slice %vector, %tile[%tile_slice_index] : vector<[16]xi8> into vector<[16]x[16]xi8>Example 2: Insert vector<[2]xf64> into tile vertically at the given index.
%tile_update = arm_sme.insert_tile_slice %vector, %tile[%tile_slice_index] layout<vertical> : vector<[2]xf64> into vector<[2]x[2]xf64>
Return op name arm_sme.intr.cntsd as a bitstring.
arm_sme.intr.cntsd
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name arm_sme.intr.ld1b.horiz as a bitstring.
arm_sme.intr.ld1b.horiz
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicateload_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.ld1b.vert as a bitstring.
arm_sme.intr.ld1b.vert
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicateload_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.ld1d.horiz as a bitstring.
arm_sme.intr.ld1d.horiz
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicateload_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.ld1d.vert as a bitstring.
arm_sme.intr.ld1d.vert
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicateload_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.ld1h.horiz as a bitstring.
arm_sme.intr.ld1h.horiz
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicateload_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.ld1h.vert as a bitstring.
arm_sme.intr.ld1h.vert
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicateload_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.ld1q.horiz as a bitstring.
arm_sme.intr.ld1q.horiz
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicateload_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.ld1q.vert as a bitstring.
arm_sme.intr.ld1q.vert
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicateload_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.ld1w.horiz as a bitstring.
arm_sme.intr.ld1w.horiz
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicateload_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.ld1w.vert as a bitstring.
arm_sme.intr.ld1w.vert
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicateload_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.mopa as a bitstring.
arm_sme.intr.mopa
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
lhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionsrhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionslhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructionsrhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructions
Return op name arm_sme.intr.mopa.wide as a bitstring.
arm_sme.intr.mopa.wide
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
lhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionsrhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionslhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructionsrhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructions
Return op name arm_sme.intr.mops as a bitstring.
arm_sme.intr.mops
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
lhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionsrhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionslhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructionsrhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructions
Return op name arm_sme.intr.mops.wide as a bitstring.
arm_sme.intr.mops.wide
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
lhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionsrhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionslhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructionsrhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructions
Return op name arm_sme.intr.read.horiz as a bitstring.
arm_sme.intr.read.horiz
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
vector- Single,SVEVector, a vector type that matches the size of a SVE vectorpredicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicatetile_slice_index- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name arm_sme.intr.read.vert as a bitstring.
arm_sme.intr.read.vert
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
vector- Single,SVEVector, a vector type that matches the size of a SVE vectorpredicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicatetile_slice_index- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name arm_sme.intr.smopa.wide as a bitstring.
arm_sme.intr.smopa.wide
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
lhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionsrhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionslhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructionsrhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructions
Return op name arm_sme.intr.smopa.za32 as a bitstring.
arm_sme.intr.smopa.za32
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
lhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionsrhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionslhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructionsrhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructions
Return op name arm_sme.intr.smops.wide as a bitstring.
arm_sme.intr.smops.wide
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
lhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionsrhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionslhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructionsrhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructions
Return op name arm_sme.intr.smops.za32 as a bitstring.
arm_sme.intr.smops.za32
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
lhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionsrhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionslhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructionsrhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructions
Return op name arm_sme.intr.st1b.horiz as a bitstring.
arm_sme.intr.st1b.horiz
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicatestore_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.st1b.vert as a bitstring.
arm_sme.intr.st1b.vert
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicatestore_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.st1d.horiz as a bitstring.
arm_sme.intr.st1d.horiz
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicatestore_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.st1d.vert as a bitstring.
arm_sme.intr.st1d.vert
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicatestore_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.st1h.horiz as a bitstring.
arm_sme.intr.st1h.horiz
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicatestore_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.st1h.vert as a bitstring.
arm_sme.intr.st1h.vert
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicatestore_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.st1q.horiz as a bitstring.
arm_sme.intr.st1q.horiz
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicatestore_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.st1q.vert as a bitstring.
arm_sme.intr.st1q.vert
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicatestore_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.st1w.horiz as a bitstring.
arm_sme.intr.st1w.horiz
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicatestore_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.st1w.vert as a bitstring.
arm_sme.intr.st1w.vert
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
predicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicatestore_address- Single,LLVM_AnyPointer, LLVM pointer typetile_slice_index- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.str as a bitstring.
arm_sme.intr.str
Operands
index- Single,I32, 32-bit signless integerstore_address- Single,LLVM_AnyPointer, LLVM pointer typeoffset- Single,I32, 32-bit signless integer
Return op name arm_sme.intr.sumopa.wide as a bitstring.
arm_sme.intr.sumopa.wide
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
lhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionsrhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionslhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructionsrhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructions
Return op name arm_sme.intr.sumops.wide as a bitstring.
arm_sme.intr.sumops.wide
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
lhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionsrhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionslhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructionsrhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructions
Return op name arm_sme.intr.umopa.wide as a bitstring.
arm_sme.intr.umopa.wide
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
lhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionsrhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionslhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructionsrhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructions
Return op name arm_sme.intr.umopa.za32 as a bitstring.
arm_sme.intr.umopa.za32
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
lhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionsrhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionslhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructionsrhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructions
Return op name arm_sme.intr.umops.wide as a bitstring.
arm_sme.intr.umops.wide
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
lhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionsrhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionslhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructionsrhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructions
Return op name arm_sme.intr.umops.za32 as a bitstring.
arm_sme.intr.umops.za32
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
lhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionsrhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionslhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructionsrhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructions
Return op name arm_sme.intr.usmopa.wide as a bitstring.
arm_sme.intr.usmopa.wide
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
lhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionsrhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionslhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructionsrhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructions
Return op name arm_sme.intr.usmops.wide as a bitstring.
arm_sme.intr.usmops.wide
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
lhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionsrhs_predicate- Single,MOPPredicate, a vector type that is a supported predicate for the SME MOP instructionslhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructionsrhs_vector- Single,MOPVector, a vector type that is a supported input for the SME MOP instructions
Return op name arm_sme.intr.write.horiz as a bitstring.
arm_sme.intr.write.horiz
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
tile_slice_index- Single,I32, 32-bit signless integerpredicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicatevector- Single,SVEVector, a vector type that matches the size of a SVE vector
Return op name arm_sme.intr.write.vert as a bitstring.
arm_sme.intr.write.vert
Attributes
tile_id- Single,I32Attr, 32-bit signless integer attribute
Operands
tile_slice_index- Single,I32, 32-bit signless integerpredicate- Single,SVEPredicate, a vector type that matches the size of a SVE predicatevector- Single,SVEVector, a vector type that matches the size of a SVE vector
Return op name arm_sme.intr.zero as a bitstring.
arm_sme.intr.zero
Attributes
tile_mask- Single,I32Attr, 32-bit signless integer attribute
Return op name arm_sme.load_tile_slice as a bitstring.
arm_sme.load_tile_slice - Tile slice load and update operation
This op has support for result type inference.
Attributes
layout- Single,ArmSME_TileSliceLayoutAttr, Layout of a tile slice
Operands
base- Single,AnyMemRef, memref of any non-token type valuesmask- Single,SVEPredicate, a vector type that matches the size of a SVE predicatetile- Single,SMETile, a vector type that fits into a SME tileindices- Variadic,Index, variadic of indextile_slice_index- Single,Index, index
Results
result- Single,SMETile, a vector type that fits into a SME tile
Description
Loads a 1D tile slice from memory into a 2D SME "virtual tile". The tile slice is defined by the dimension of the 2D scalable vector type pointed by the index. A tile slice index describes where in the input tile the tile slice is loaded to. An optional tile slice layout attribute specifies whether the tile slice being loaded at the given index is horizontal (default) or vertical. The updated tile is returned as the result.
The slice of memory read is defined by a base and indices and must be contiguous. The memref must be either rank 1 or rank 2, have dynamic dimensions since the operation is scalable, and the element type must be a scalar that matches the element type of the result.
The provided mask is used to specify which elements of the tile slice
will be loaded.
Example 1: Load a vector<[16]xi8> tile slice from memory into tile horizontally (default) at given index.
%tile_update = arm_sme.load_tile_slice %base[%c0], %mask, %tile, %tile_slice_index : memref<?x?xi8>, vector<[16]xi1>, vector<[16]x[16]xi8>Example 2: Load a vector<[4]xf32> tile slice from memory into tile vertically at given index.
%tile_update = arm_sme.load_tile_slice %base[%c0], %mask, %tile, %tile_slice_index layout<vertical> : memref<?x?xf32>, vector<[4]xi1>, vector<[4]x[4]xf32>Example 3: Load a vector<[1]xi128> tile slice from memory into tile vertically at given index.
%tile_update = arm_sme.load_tile_slice %base[%c0], %mask, %tile, %tile_slice_index layout<vertical> : memref<?x?xi128>, vector<[1]xi1>, vector<[1]x[1]xi128>
Return op name arm_sme.outerproduct as a bitstring.
arm_sme.outerproduct - Outer product with optional fused add/sub
This op has support for result type inference.
Attributes
kind- Single,ArmSME_CombiningKindAttr, Kind of combining function
Operands
lhs- Single,SVEVector, a vector type that matches the size of a SVE vectorrhs- Single,SVEVector, a vector type that matches the size of a SVE vectorlhsMask- Optional,SVEPredicate, a vector type that matches the size of a SVE predicaterhsMask- Optional,SVEPredicate, a vector type that matches the size of a SVE predicateacc- Optional,SMETile, a vector type that fits into a SME tile
Results
result- Single,SMETile, a vector type that fits into a SME tile
Description
This operation represents an outer product that fits within an SME tile.
All operands must be SVE vectors and the result a SME tile. Unlike
vector.outerproduct masking is on the operands (rather than the result),
which mirrors the SME instructions.
Example 1: Unmasked outerproduct (without accumulator)
// Not specifying an accumulator implicitly zeros the destination tile.
%result = arm_sme.outerproduct $lhs, $rhs : vector<[4]xf32>, vector<[4]xf32>Example 2: Unmasked outerproduct (with accumulator)
%result = arm_sme.outerproduct $lhs, $rhs acc($accumulator)
: vector<[4]xf32>, vector<[4]xf32>Example 3: Masked outerproduct
%result = arm_sme.outerproduct $lhs, $rhs masks($lhsMask, $rhsMask)
: vector<[4]xf32>, vector<[4]xf32>Example 4: Masked outerproduct (with accumulator)
%result = arm_sme.outerproduct $lhs, $rhs acc($accumulator) masks($lhsMask, $rhsMask)
: vector<[4]xf32>, vector<[4]xf32>
Return op name arm_sme.smopa_2way as a bitstring.
arm_sme.smopa_2way - Signed integer sum of 2 outer products and accumulate
Operands
lhs- Single, anonymous/composite constraint, of ranks 1scalable vector of 16-bit signless integer values of length 8rhs- Single,AnyVectorOfNonZeroRank, vector of any non-token type valueslhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesrhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesacc- Optional,AnyVectorOfNonZeroRank, vector of any non-token type values
Results
result- Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values
Description
Example:
%result = arm_sme.smopa_2way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[4]x[4]xi32>Refer to fmopa_2way for a detailed description of 2-way outer products.
| Spec | Features |
|---|---|
| SMOPA (2-way) | +sme2 |
Return op name arm_sme.smopa_4way as a bitstring.
arm_sme.smopa_4way - Signed integer sum of 4 outer products and accumulate
Operands
lhs- Single, anonymous/composite constraint, of ranks 1scalable vector of 8-bit signless integer values of length 16 or of ranks 1scalable vector of 16-bit signless integer values of length 8rhs- Single,AnyVectorOfNonZeroRank, vector of any non-token type valueslhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesrhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesacc- Optional,AnyVectorOfNonZeroRank, vector of any non-token type values
Results
result- Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values or vector<[2]x[2]xi64> of 64-bit signless integer values
Description
This operation represents a sum of 4 widened outer products. It takes 2 1-D scalable vectors as input and a 2-D scalable vector (ZA tile) as output.
For example (i8 to i32):
%result = arm_sme.smopa_4way $lhs, $rhs :
vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>The lhs encodes a matrix of shape SVLSx4 and the rhs a matrix of
4xSVLS, where SVLS (spec [1], section B2.1) is the number of 32-bit
elements in a vector of SVL bits. To illustrate, below is a breakdown of
this operation for i8 to i32, SVL=128 (i.e., vscale=1):
LHS
[A0 A1 A2 A3 A4 A5 A6 A7 A8 A9 A10 A11 A12 A15 A14 A15]
RHS
[B0 B1 B2 B3 B4 B5 B6 B7 B8 B9 B10 B11 B12 B13 B14 B15]
----------------------------------------------------------------------------
implicit layout
[A0 A1 A2 A3] | [B0 B4 B8 B12]
[A4 A5 A6 A7] | [B1 B5 B9 B13]
[A8 A9 A10 A11] | [B2 B6 B10 B14]
[A12 A13 A14 A15] | [B3 B7 B11 B15]
----------------------------------------------------------------------------
4 outer products
Acol0 ⊗ Brow0 | Acol1 ⊗ Brow1
------------- | -------------
|
[B0 B4 B8 B12] | [B1 B5 B9 B13]
|
[A0 [ A0B0 A0B4 A0B8 A0B12] | [A1 [ A1B1 A1B5 A1B9 A1B13]
A4 [ A4B0 A4B4 A4B8 A4B12] | A5 [ A5B1 A5B5 A5B9 A5B13]
A8 [ A8B0 A8B4 A8B8 A8B12] | A9 [ A9B1 A9B5 A9B9 A9B13]
A12] [A12B0 A12B4 A12B8 A12B12] | A13] [A13B1 A13B5 A13B9 A13B13]
|
Acol2 ⊗ Brow2 | Acol3 ⊗ Brow3
------------- | -------------
|
[B2, B6, B10, B14] | [B3 B7 B11 B15]
|
[A2 [ A2B2 A2B6 A2B10 A2B14] | [A3 [ A3B3 A3B7 A3B11 A3B15]
A6 [ A6B2 A6B6 A6B10 A6B14] | A7 [ A7B3 A7B7 A7B11 A7B15]
A10 [A10B2 A10B6 A10B10 A10B14] | A11 [A11B3 A11B7 A11B11 A11B15]
A14] [A14B2 A14B6 A14B10 A14B14] | A15] [A15B3 A15B7 A15B11 A15B15]
|
----------------------------------------------------------------------------
sum of 4 outer products
Acol0 ⊗ Brow0 + Acol1 ⊗ Brow1 + Acol2 ⊗ Brow2 + Acol3 ⊗ Brow3
[ A0B0 + A1B1 + A2B2 + A3B3 ... ... A0B12 + A1B13 + A2B14 + A3B15]
[ A4B0 + A5B1 + A6B2 + A7B3 ... ... A4B12 + A5B13 + A6B14 + A7B15]
[ A8B0 + A9B1 + A10B2 + A11B3 ... ... A8B12 + A9B13 + A10B14 + A11B15]
[A12B0 + A13B1 + A14B2 + A15B3 ... ... A12B12 + A13B13 + A14B14 + A15B15]
----------------------------------------------------------------------------This operation enables the folding of 4 outer products chained via the accumulator into a single outer product.
For example:
%a0_ext = arith.extsi %a0 : vector<[4]xi8> to vector<[4]xi32>
%b0_ext = arith.extsi %b0 : vector<[4]xi8> to vector<[4]xi32>
%a1_ext = arith.extsi %a1 : vector<[4]xi8> to vector<[4]xi32>
%b1_ext = arith.extsi %b1 : vector<[4]xi8> to vector<[4]xi32>
%a2_ext = arith.extsi %a2 : vector<[4]xi8> to vector<[4]xi32>
%b2_ext = arith.extsi %b2 : vector<[4]xi8> to vector<[4]xi32>
%a3_ext = arith.extsi %a3 : vector<[4]xi8> to vector<[4]xi32>
%b3_ext = arith.extsi %b3 : vector<[4]xi8> to vector<[4]xi32>
%0 = arm_sme.outerproduct %a0_ext, %b0_ext : vector<[4]xi32>, vector<[4]xi32>
%1 = arm_sme.outerproduct %a1_ext, %b1_ext acc(%0) : vector<[4]xi32>, vector<[4]xi32>
%2 = arm_sme.outerproduct %a2_ext, %b2_ext acc(%1) : vector<[4]xi32>, vector<[4]xi32>
%3 = arm_sme.outerproduct %a3_ext, %b3_ext acc(%2) : vector<[4]xi32>, vector<[4]xi32>The 4 outer products in the example above can be fused into a single outer product as follows:
%lhs0 = vector.interleave %a0, %a2 : vector<[4]xi8> -> vector<[8]xi8>
%lhs1 = vector.interleave %a1, %a3 : vector<[4]xi8> -> vector<[8]xi8>
%lhs = vector.interleave %lhs0, %lhs1 : vector<[8]xi8> -> vector<[16]xi8>
%rhs0 = vector.interleave %b0, %b2 : vector<[4]xi8> -> vector<[8]xi8>
%rhs1 = vector.interleave %b1, %b3 : vector<[4]xi8> -> vector<[8]xi8>
%rhs = vector.interleave %rhs0, %rhs1 : vector<[8]xi8> -> vector<[16]xi8>
%0 = arm_sme.smopa_4way %lhs, %rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>This is implemented in the -arm-sme-outer-product-fusion pass.
Example: I8 to I32
%result = arm_sme.smopa_4way $lhs, $rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>Example: I16 to I64
%result = arm_sme.smopa_4way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[2]x[2]xi64>| Spec | Features |
|---|---|
| SMOPA (4-way) | +sme (32-bit), +sme-i16i64 (64-bit) |
Return op name arm_sme.smops_2way as a bitstring.
arm_sme.smops_2way - Signed integer sum of 2 outer products and subtract
Operands
lhs- Single, anonymous/composite constraint, of ranks 1scalable vector of 16-bit signless integer values of length 8rhs- Single,AnyVectorOfNonZeroRank, vector of any non-token type valueslhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesrhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesacc- Optional,AnyVectorOfNonZeroRank, vector of any non-token type values
Results
result- Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values
Description
Example:
%result = arm_sme.smops_2way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[4]x[4]xi32>Refer to fmopa_2way for a detailed description of 2-way outer products.
| Spec | Features |
|---|---|
| SMOPS (2-way) | +sme2 |
Return op name arm_sme.smops_4way as a bitstring.
arm_sme.smops_4way - Signed integer sum of 4 outer products and subtract
Operands
lhs- Single, anonymous/composite constraint, of ranks 1scalable vector of 8-bit signless integer values of length 16 or of ranks 1scalable vector of 16-bit signless integer values of length 8rhs- Single,AnyVectorOfNonZeroRank, vector of any non-token type valueslhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesrhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesacc- Optional,AnyVectorOfNonZeroRank, vector of any non-token type values
Results
result- Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values or vector<[2]x[2]xi64> of 64-bit signless integer values
Description
Equivalent to smopa_4way but outer products are subtracted from
destination result.
Example: I8 to I32
%result = arm_sme.smops_4way $lhs, $rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>Example: I16 to I64
%result = arm_sme.smops_4way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[2]x[2]xi64>Refer to smopa_4way for a detailed description of 4-way outer products.
| Spec | Features |
|---|---|
| SMOPS (4-way) | +sme (32-bit), +sme-i16i64 (64-bit) |
Return op name arm_sme.store_tile_slice as a bitstring.
arm_sme.store_tile_slice - Tile slice store operation
Attributes
layout- Single,ArmSME_TileSliceLayoutAttr, Layout of a tile slice
Operands
tile- Single,SMETile, a vector type that fits into a SME tiletile_slice_index- Single,Index, indexmask- Single,SVEPredicate, a vector type that matches the size of a SVE predicatebase- Single,AnyMemRef, memref of any non-token type valuesindices- Variadic,Index, variadic of index
Description
Stores a 1D tile slice from a 2D SME "virtual tile" into memory. The tile slice is defined by the dimension of the 2D scalable vector type pointed by the index. A tile slice index describes where in the input tile the tile slice is stored from. An optional tile slice layout attribute specifies whether the tile slice being stored from the given index is horizontal (default) or vertical.
The slice of memory written is defined by a base and indices and must be contiguous. The memref must be either rank 1 or rank 2, have dynamic dimensions since the operation is scalable, and the element type must be a scalar that matches the element type of the input tile.
The provided mask is used to specify which elements of the tile slice
will be stored.
Example 1: Store vector<[16]xi8> horizontal (default) tile slice from tile at given index to memory.
arm_sme.store_tile_slice %tile, %tile_slice_index, %mask, %base[%c0] : vector<[16]x[16]xi8>, vector<[16]xi1>, memref<?x?xi8>Example 2: Store vector<[4]xf32> vertical tile slice from tile at given index to memory.
arm_sme.store_tile_slice %tile, %tile_slice_index, %mask, %base[%c0] layout<vertical> : vector<[4]x[4]xf32>, vector<[4]xi1>, memref<?x?xf32>Example 3: Store a vector<[1]xi128> vertical tile slice from tile at given index to memory.
arm_sme.store_tile_slice %tile, %tile_slice_index, %mask, %base[%c0] layout<vertical> : vector<[1]x[1]xi128>, vector<[1]xi1>, memref<?x?xi128>
Return op name arm_sme.streaming_vl as a bitstring.
arm_sme.streaming_vl - Query the streaming vector length
This op has support for result type inference.
Attributes
type_size- Single,ArmSME_TypeSizeAttr, Size of a vector element type
Results
- anonymous - Single,
Index, index
Description
This operation returns the streaming vector length (SVL) for a given type
size. Unlike vector.vscale the value returned is invariant to the
streaming mode.
Example:
// Streaming vector length in:
// - bytes (8-bit, SVL.B)
%svl_b = arm_sme.streaming_vl <byte>
// - half words (16-bit, SVL.H)
%svl_h = arm_sme.streaming_vl <half>
// - words (32-bit, SVL.W)
%svl_w = arm_sme.streaming_vl <word>
// - double words (64-bit, SVL.D)
%svl_d = arm_sme.streaming_vl <double>
Return op name arm_sme.sumopa_4way as a bitstring.
arm_sme.sumopa_4way - Signed by unsigned integer sum of 4 outer products and accumulate
Operands
lhs- Single, anonymous/composite constraint, of ranks 1scalable vector of 8-bit signless integer values of length 16 or of ranks 1scalable vector of 16-bit signless integer values of length 8rhs- Single,AnyVectorOfNonZeroRank, vector of any non-token type valueslhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesrhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesacc- Optional,AnyVectorOfNonZeroRank, vector of any non-token type values
Results
result- Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values or vector<[2]x[2]xi64> of 64-bit signless integer values
Description
Example: I8 to I32
%result = arm_sme.sumopa_4way $lhs, $rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>Example: I16 to I64
%result = arm_sme.sumopa_4way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[2]x[2]xi64>Refer to smopa_4way for a detailed description of 4-way outer products.
| Spec | Features |
|---|---|
| SUMOPA (4-way) | +sme (32-bit), +sme-i16i64 (64-bit) |
Return op name arm_sme.sumops_4way as a bitstring.
arm_sme.sumops_4way - Signed by unsigned integer sum of 4 outer products and subtract
Operands
lhs- Single, anonymous/composite constraint, of ranks 1scalable vector of 8-bit signless integer values of length 16 or of ranks 1scalable vector of 16-bit signless integer values of length 8rhs- Single,AnyVectorOfNonZeroRank, vector of any non-token type valueslhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesrhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesacc- Optional,AnyVectorOfNonZeroRank, vector of any non-token type values
Results
result- Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values or vector<[2]x[2]xi64> of 64-bit signless integer values
Description
Example: I8 to I32
%result = arm_sme.sumops_4way $lhs, $rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>Example: I16 to I64
%result = arm_sme.sumops_4way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[2]x[2]xi64>Refer to smopa_4way for a detailed description of 4-way outer products.
| Spec | Features |
|---|---|
| SUMOPS (4-way) | +sme (32-bit), +sme-i16i64 (64-bit) |
Return op name arm_sme.tile_load as a bitstring.
arm_sme.tile_load - Tile load operation
Attributes
layout- Single,ArmSME_TileSliceLayoutAttr, Layout of a tile slice
Operands
base- Single, anonymous/composite constraint, 2D memref of any non-token type valuesindices- Variadic,Index, variadic of indexpadding- Optional,AnyType, any non-token typemask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type values
Results
result- Single,SMETile, a vector type that fits into a SME tile
Description
Loads a 2D SME "virtual tile" from memory defined by a base and indices, with the shape defined by the 2D scalable vector type of the result tile. An optional tile slice layout attribute specifies whether the slices of the tile being loaded are horizontal (default) or vertical. The slice of memory must be contiguous. The memref must be either rank 1 or rank 2 with dynamic dimensions, since the operation is scalable, and the element type must be a scalar that matches the element type of the result.
An optional SSA value padding of the same elemental type as the MemRef is
provided to specify a fallback value in the case of masking.
An optional SSA value mask may be specified to mask out elements read
from the MemRef. The mask type is an i1 vector with a shape that
matches how elements are read from the MemRef. Elements whose corresponding
mask element is 0 are masked out and replaced with padding.
If either padding or mask are specified, both must be specified.
Example 1: Load an 8-bit element ZA tile with horizontal layout (default) from memory (ZA0.B).
%tile = arm_sme.tile_load %base[%c0, %c0] : memref<?x?xi8>, vector<[16]x[16]xi8>Example 2: Load a FP 32-bit element ZA tile with vertical layout from memory.
%tile = arm_sme.tile_load %base[%c0, %c0] layout<vertical> : memref<?x?xf32>, vector<[4]x[4]xf32>Example 3: Load a 128-bit element ZA tile with horizontal layout (default) from memory.
%tile = arm_sme.tile_load %base[%c0, %c0] layout<horizontal> : memref<?x?xi128>, vector<[1]x[1]xi128>Example 4: Masked load of int 32-bit element ZA tile with horizontal layout (default) from memory.
%tile = arm_sme.tile_load %base[%c0, %c0], %pad, %mask : memref<?x?xf32>, vector<[4]x[4]xf32>
Return op name arm_sme.tile_store as a bitstring.
arm_sme.tile_store - Tile store operation
Attributes
layout- Single,ArmSME_TileSliceLayoutAttr, Layout of a tile slice
Operands
valueToStore- Single,SMETile, a vector type that fits into a SME tilebase- Single, anonymous/composite constraint, 2D memref of any non-token type valuesindices- Variadic,Index, variadic of indexmask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type values
Description
Stores a 2D SME "virtual tile" to memory defined by a base and indices, with the shape defined by the 2D scalable vector type of the tile being stored. An optional tile slice layout attribute specifies whether the slices of the tile being stored are horizontal (default) or vertical. The slice of memory must be contiguous. The memref must be either rank 1 or rank 2 with dynamic dimensions, since the operation is scalable, and the element type must be a scalar that matches the element type of the result.
An optional mask may be provided, the shape of which corresponds to the
tile, and selects which elements of the tile will be stored.
Example 1: Store an 8-bit element ZA tile with horizontal (default) layout to memory (ZA0.B).
arm_sme.tile_store %tile, %base[%c0, %c0] : vector<[16]x[16]xi8>, memref<?x?xi8>Example 2: Store a FP 32-bit element ZA tile with vertical layout to memory.
arm_sme.tile_store %tile, %base[%c0, %c0] layout<vertical> : vector<[4]x[4]xf32>, memref<?x?xf32>Example 3: Store a 128-bit element ZA tile with horizontal (default) layout to memory.
arm_sme.tile_store %tile, %base[%c0, %c0] layout<horizontal> : vector<[1]x[1]xi128>, memref<?x?xi128>Example 4: Masked store a int 32-bit element ZA tile with vertical layout to memory.
arm_sme.tile_store %tile, %base[%c0, %c0], %mask layout<vertical> : vector<[4]x[4]xf32>, memref<?x?xf32>
Return op name arm_sme.umopa_2way as a bitstring.
arm_sme.umopa_2way - Unsiged integer sum of 2 outer products and accumulate
Operands
lhs- Single, anonymous/composite constraint, of ranks 1scalable vector of 16-bit signless integer values of length 8rhs- Single,AnyVectorOfNonZeroRank, vector of any non-token type valueslhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesrhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesacc- Optional,AnyVectorOfNonZeroRank, vector of any non-token type values
Results
result- Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values
Description
Example:
%result = arm_sme.umopa_2way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[4]x[4]xi32>Refer to fmopa_2way for a detailed description of 2-way outer products.
| Spec | Features |
|---|---|
| UMOPA (2-way) | +sme2 |
Return op name arm_sme.umopa_4way as a bitstring.
arm_sme.umopa_4way - Unsigned integer sum of 4 outer products and accumulate
Operands
lhs- Single, anonymous/composite constraint, of ranks 1scalable vector of 8-bit signless integer values of length 16 or of ranks 1scalable vector of 16-bit signless integer values of length 8rhs- Single,AnyVectorOfNonZeroRank, vector of any non-token type valueslhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesrhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesacc- Optional,AnyVectorOfNonZeroRank, vector of any non-token type values
Results
result- Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values or vector<[2]x[2]xi64> of 64-bit signless integer values
Description
Example: I8 to I32
%result = arm_sme.umopa_4way $lhs, $rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>Example: I16 to I64
%result = arm_sme.umopa_4way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[2]x[2]xi64>Refer to smopa_4way for a detailed description of 4-way outer products.
| Spec | Features |
|---|---|
| UMOPA (4-way) | +sme (32-bit), +sme-i16i64 (64-bit) |
Return op name arm_sme.umops_2way as a bitstring.
arm_sme.umops_2way - Unsiged integer sum of 2 outer products and subtract
Operands
lhs- Single, anonymous/composite constraint, of ranks 1scalable vector of 16-bit signless integer values of length 8rhs- Single,AnyVectorOfNonZeroRank, vector of any non-token type valueslhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesrhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesacc- Optional,AnyVectorOfNonZeroRank, vector of any non-token type values
Results
result- Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values
Description
Example:
%result = arm_sme.umops_2way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[4]x[4]xi32>Refer to fmopa_2way for a detailed description of 2-way outer products.
| Spec | Features |
|---|---|
| UMOPS (2-way) | +sme2 |
Return op name arm_sme.umops_4way as a bitstring.
arm_sme.umops_4way - Unsigned integer sum of 4 outer products and subtract
Operands
lhs- Single, anonymous/composite constraint, of ranks 1scalable vector of 8-bit signless integer values of length 16 or of ranks 1scalable vector of 16-bit signless integer values of length 8rhs- Single,AnyVectorOfNonZeroRank, vector of any non-token type valueslhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesrhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesacc- Optional,AnyVectorOfNonZeroRank, vector of any non-token type values
Results
result- Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values or vector<[2]x[2]xi64> of 64-bit signless integer values
Description
Example: I8 to I32
%result = arm_sme.umops_4way $lhs, $rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>Example: I16 to I64
%result = arm_sme.umops_4way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[2]x[2]xi64>Refer to smopa_4way for a detailed description of 4-way outer products.
| Spec | Features |
|---|---|
| UMOPS (4-way) | +sme (32-bit), +sme-i16i64 (64-bit) |
Return op name arm_sme.usmopa_4way as a bitstring.
arm_sme.usmopa_4way - Unsigned by signed integer sum of 4 outer products and accumulate
Operands
lhs- Single, anonymous/composite constraint, of ranks 1scalable vector of 8-bit signless integer values of length 16 or of ranks 1scalable vector of 16-bit signless integer values of length 8rhs- Single,AnyVectorOfNonZeroRank, vector of any non-token type valueslhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesrhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesacc- Optional,AnyVectorOfNonZeroRank, vector of any non-token type values
Results
result- Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values or vector<[2]x[2]xi64> of 64-bit signless integer values
Description
Example: I8 to I32
%result = arm_sme.usmopa_4way $lhs, $rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>Example: I16 to I64
%result = arm_sme.usmopa_4way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[2]x[2]xi64>Refer to smopa_4way for a detailed description of 4-way outer products.
| Spec | Features |
|---|---|
| USMOPA (4-way) | +sme (32-bit), +sme-i16i64 (64-bit) |
Return op name arm_sme.usmops_4way as a bitstring.
arm_sme.usmops_4way - Unsigned by signed integer sum of 4 outer products and subtract
Operands
lhs- Single, anonymous/composite constraint, of ranks 1scalable vector of 8-bit signless integer values of length 16 or of ranks 1scalable vector of 16-bit signless integer values of length 8rhs- Single,AnyVectorOfNonZeroRank, vector of any non-token type valueslhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesrhsMask- Optional,AnyVectorOfNonZeroRank, vector of any non-token type valuesacc- Optional,AnyVectorOfNonZeroRank, vector of any non-token type values
Results
result- Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values or vector<[2]x[2]xi64> of 64-bit signless integer values
Description
Example: I8 to I32
%result = arm_sme.usmops_4way $lhs, $rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>Example: I16 to I64
%result = arm_sme.usmops_4way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[2]x[2]xi64>Refer to smopa_4way for a detailed description of 4-way outer products.
| Spec | Features |
|---|---|
| USMOPS (4-way) | +sme (32-bit), +sme-i16i64 (64-bit) |
Return op name arm_sme.zero as a bitstring.
arm_sme.zero - Creates a zero-initialized value of SME virtual tile type
Results
res- Single,SMETile, a vector type that fits into a SME tile
Description
Creates a new SME "virtual tile" value within a function. The contents of the tile returned from this operation are zero-initialized.
Example 1: Zero an 8-bit element ZA tile.
%0 = arm_sme.zero : vector<[16]x[16]xi8>Example 2: Zero a 64-bit element ZA tile.
%0 = arm_sme.zero : vector<[2]x[2]xi64>