Beaver.MLIR.Dialect.ArmSME (beaver v0.4.8)

Copy Markdown

Summary

Functions

Return op name arm_sme.copy_tile as a bitstring.

arm_sme.copy_tile - Copies an SME tile value

Return op name arm_sme.extract_tile_slice as a bitstring.

arm_sme.extract_tile_slice - Extract 1-D scalable vector from slice of 2-D tile

Return op name arm_sme.fmopa_2way as a bitstring.

arm_sme.fmopa_2way - Floating-point sum of 2 outer products and accumulate

Return op name arm_sme.fmops_2way as a bitstring.

arm_sme.fmops_2way - Floating-point sum of 2 outer products and subtract

Return op name arm_sme.get_tile as a bitstring.

arm_sme.get_tile - Creates an undefined value of SME virtual tile type

Return op name arm_sme.insert_tile_slice as a bitstring.

arm_sme.insert_tile_slice - Insert 1-D scalable vector into slice of 2-D tile

Return op name arm_sme.intr.cntsd as a bitstring.

arm_sme.intr.cntsd

Return op name arm_sme.intr.ld1b.horiz as a bitstring.

arm_sme.intr.ld1b.horiz

Return op name arm_sme.intr.ld1b.vert as a bitstring.

arm_sme.intr.ld1b.vert

Return op name arm_sme.intr.ld1d.horiz as a bitstring.

arm_sme.intr.ld1d.horiz

Return op name arm_sme.intr.ld1d.vert as a bitstring.

arm_sme.intr.ld1d.vert

Return op name arm_sme.intr.ld1h.horiz as a bitstring.

arm_sme.intr.ld1h.horiz

Return op name arm_sme.intr.ld1h.vert as a bitstring.

arm_sme.intr.ld1h.vert

Return op name arm_sme.intr.ld1q.horiz as a bitstring.

arm_sme.intr.ld1q.horiz

Return op name arm_sme.intr.ld1q.vert as a bitstring.

arm_sme.intr.ld1q.vert

Return op name arm_sme.intr.ld1w.horiz as a bitstring.

arm_sme.intr.ld1w.horiz

Return op name arm_sme.intr.ld1w.vert as a bitstring.

arm_sme.intr.ld1w.vert

Return op name arm_sme.intr.mopa as a bitstring.

arm_sme.intr.mopa

Return op name arm_sme.intr.mopa.wide as a bitstring.

arm_sme.intr.mopa.wide

Return op name arm_sme.intr.mops as a bitstring.

arm_sme.intr.mops

Return op name arm_sme.intr.mops.wide as a bitstring.

arm_sme.intr.mops.wide

Return op name arm_sme.intr.read.horiz as a bitstring.

arm_sme.intr.read.horiz

Return op name arm_sme.intr.read.vert as a bitstring.

arm_sme.intr.read.vert

Return op name arm_sme.intr.smopa.wide as a bitstring.

arm_sme.intr.smopa.wide

Return op name arm_sme.intr.smopa.za32 as a bitstring.

arm_sme.intr.smopa.za32

Return op name arm_sme.intr.smops.wide as a bitstring.

arm_sme.intr.smops.wide

Return op name arm_sme.intr.smops.za32 as a bitstring.

arm_sme.intr.smops.za32

Return op name arm_sme.intr.st1b.horiz as a bitstring.

arm_sme.intr.st1b.horiz

Return op name arm_sme.intr.st1b.vert as a bitstring.

arm_sme.intr.st1b.vert

Return op name arm_sme.intr.st1d.horiz as a bitstring.

arm_sme.intr.st1d.horiz

Return op name arm_sme.intr.st1d.vert as a bitstring.

arm_sme.intr.st1d.vert

Return op name arm_sme.intr.st1h.horiz as a bitstring.

arm_sme.intr.st1h.horiz

Return op name arm_sme.intr.st1h.vert as a bitstring.

arm_sme.intr.st1h.vert

Return op name arm_sme.intr.st1q.horiz as a bitstring.

arm_sme.intr.st1q.horiz

Return op name arm_sme.intr.st1q.vert as a bitstring.

arm_sme.intr.st1q.vert

Return op name arm_sme.intr.st1w.horiz as a bitstring.

arm_sme.intr.st1w.horiz

Return op name arm_sme.intr.st1w.vert as a bitstring.

arm_sme.intr.st1w.vert

Return op name arm_sme.intr.str as a bitstring.

arm_sme.intr.str

Return op name arm_sme.intr.sumopa.wide as a bitstring.

arm_sme.intr.sumopa.wide

Return op name arm_sme.intr.sumops.wide as a bitstring.

arm_sme.intr.sumops.wide

Return op name arm_sme.intr.umopa.wide as a bitstring.

arm_sme.intr.umopa.wide

Return op name arm_sme.intr.umopa.za32 as a bitstring.

arm_sme.intr.umopa.za32

Return op name arm_sme.intr.umops.wide as a bitstring.

arm_sme.intr.umops.wide

Return op name arm_sme.intr.umops.za32 as a bitstring.

arm_sme.intr.umops.za32

Return op name arm_sme.intr.usmopa.wide as a bitstring.

arm_sme.intr.usmopa.wide

Return op name arm_sme.intr.usmops.wide as a bitstring.

arm_sme.intr.usmops.wide

Return op name arm_sme.intr.write.horiz as a bitstring.

arm_sme.intr.write.horiz

Return op name arm_sme.intr.write.vert as a bitstring.

arm_sme.intr.write.vert

Return op name arm_sme.intr.zero as a bitstring.

arm_sme.intr.zero

Return op name arm_sme.load_tile_slice as a bitstring.

arm_sme.load_tile_slice - Tile slice load and update operation

Return op name arm_sme.outerproduct as a bitstring.

arm_sme.outerproduct - Outer product with optional fused add/sub

Return op name arm_sme.smopa_2way as a bitstring.

arm_sme.smopa_2way - Signed integer sum of 2 outer products and accumulate

Return op name arm_sme.smopa_4way as a bitstring.

arm_sme.smopa_4way - Signed integer sum of 4 outer products and accumulate

Return op name arm_sme.smops_2way as a bitstring.

arm_sme.smops_2way - Signed integer sum of 2 outer products and subtract

Return op name arm_sme.smops_4way as a bitstring.

arm_sme.smops_4way - Signed integer sum of 4 outer products and subtract

Return op name arm_sme.store_tile_slice as a bitstring.

arm_sme.store_tile_slice - Tile slice store operation

Return op name arm_sme.streaming_vl as a bitstring.

arm_sme.streaming_vl - Query the streaming vector length

Return op name arm_sme.sumopa_4way as a bitstring.

arm_sme.sumopa_4way - Signed by unsigned integer sum of 4 outer products and accumulate

Return op name arm_sme.sumops_4way as a bitstring.

arm_sme.sumops_4way - Signed by unsigned integer sum of 4 outer products and subtract

Return op name arm_sme.tile_load as a bitstring.

arm_sme.tile_load - Tile load operation

Return op name arm_sme.tile_store as a bitstring.

arm_sme.tile_store - Tile store operation

Return op name arm_sme.umopa_2way as a bitstring.

arm_sme.umopa_2way - Unsiged integer sum of 2 outer products and accumulate

Return op name arm_sme.umopa_4way as a bitstring.

arm_sme.umopa_4way - Unsigned integer sum of 4 outer products and accumulate

Return op name arm_sme.umops_2way as a bitstring.

arm_sme.umops_2way - Unsiged integer sum of 2 outer products and subtract

Return op name arm_sme.umops_4way as a bitstring.

arm_sme.umops_4way - Unsigned integer sum of 4 outer products and subtract

Return op name arm_sme.usmopa_4way as a bitstring.

arm_sme.usmopa_4way - Unsigned by signed integer sum of 4 outer products and accumulate

Return op name arm_sme.usmops_4way as a bitstring.

arm_sme.usmops_4way - Unsigned by signed integer sum of 4 outer products and subtract

Return op name arm_sme.zero as a bitstring.

arm_sme.zero - Creates a zero-initialized value of SME virtual tile type

Functions

copy_tile()

Return op name arm_sme.copy_tile as a bitstring.

copy_tile(ssa)

arm_sme.copy_tile - Copies an SME tile value

This op has support for result type inference.

Operands

  • tile - Single, SMETile, a vector type that fits into a SME tile

Results

  • result - Single, SMETile, a vector type that fits into a SME tile

Description

Copies an SME "virtual tile" value to a new SSA value. This operation is primarily intended to be used to normalize the IR prior to tile allocation.

Example:

%copy = arm_sme.copy_tile %tile : vector<[4]x[4]xf32>

extract_tile_slice()

Return op name arm_sme.extract_tile_slice as a bitstring.

extract_tile_slice(ssa)

arm_sme.extract_tile_slice - Extract 1-D scalable vector from slice of 2-D tile

This op has support for result type inference.

Attributes

  • layout - Single, ArmSME_TileSliceLayoutAttr, Layout of a tile slice

Operands

  • tile - Single, SMETile, a vector type that fits into a SME tile
  • tile_slice_index - Single, Index, index

Results

  • result - Single, SVEVector, a vector type that matches the size of a SVE vector

Description

Extracts a 1-D scalable slice from a 2-D scalable tile at the given index. A tile slice is a 1-D vector of horizontally or vertically contiguous elements within a ZA tile.

An optional tile slice layout attribute specifies whether the tile slice is horizontal (default) or vertical.

Example 1: Extract vector<[16]xi8> from tile horizontally at the given index.

%slice = arm_sme.extract_tile_slice %tile[%tile_slice_index] : vector<[16]xi8> from vector<[16]x[16]xi8>

Example 2: Extract vector<[2]xf64> from tile vertically at the given index.

%slice = arm_sme.extract_tile_slice %tile[%tile_slice_index] layout<vertical> : vector<[2]xf64> from vector<[2]x[2]xf64>

fmopa_2way()

Return op name arm_sme.fmopa_2way as a bitstring.

fmopa_2way(ssa)

arm_sme.fmopa_2way - Floating-point sum of 2 outer products and accumulate

Operands

  • lhs - Single, anonymous/composite constraint, of ranks 1scalable vector of 16-bit float or bfloat16 type values of length 8
  • rhs - Single, AnyVectorOfNonZeroRank, vector of any non-token type values
  • lhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • rhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • acc - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values

Results

  • result - Single, anonymous/composite constraint, vector<[4]x[4]xf32> of 32-bit float values

Description

This operation represents a sum of 2 widened outer products. It takes 2 1-D scalable vectors as input and a 2-D scalable vector (ZA tile) as output.

For example (fp16 to fp32):

%result = arm_sme.fmopa_2way %lhs, %rhs :
  vector<[8]xf16>, vector<[8]xf16> into vector<[4]x[4]xf32>

The lhs encodes a matrix of shape SVLSx2 and the rhs a matrix of 2xSVLS, where SVLS (spec [1], section B2.1) is the number of 32-bit elements in a vector of SVL bits. To illustrate, below is a breakdown of this operation for fp16 to fp32, SVL=128 (i.e., vscale=1):

                      LHS                          RHS
           [A0 A1 A2 A3 A4 A5 A6 A7]    [B0 B1 B2 B3 B4 B5 B6 B7]

----------------------------------------------------------------------------

                              implicit layout

                          [A0 A1]    |
                          [A2 A3]    |    [B0 B2 B4 B6]
                          [A4 A5]    |    [B1 B3 B5 B7]
                          [A6 A7]    |

----------------------------------------------------------------------------

                              2 outer products

                  Acol0  Brow0      |           Acol1  Brow1
                  -------------      |           -------------
                                     |
              [B0 B2 B4 B6]          |       [B1 B3 B5 B7]
                                     |
         [A0  [A0B0 A0B2 A0B4 A0B6]  |  [A1  [A1B1 A1B3 A1B5 A1B7]
          A2  [A2B0 A2B2 A2B4 A2B6]  |   A3  [A3B1 A3B3 A3B5 A3B7]
          A4  [A4B0 A4B2 A4B4 A4B6]  |   A5  [A5B1 A5B3 A5B5 A5B7]
          A6] [A6B0 A6B2 A6B4 A6B6]  |   A7] [A7B1 A7B3 A7B5 A7B7]
                                     |

----------------------------------------------------------------------------

                          sum of 2 outer products

                       Acol0  Brow0 + Acol1  Brow1

             [A0B0 + A1B1 A0B2 + A1B3 A0B4 + A1B5 A0B6 + A1B7]
             [A2B0 + A3B1 A2B2 + A3B3 A2B4 + A3B5 A2B6 + A3B7]
             [A4B0 + A5B1 A4B2 + A5B3 A4B4 + A5B5 A4B6 + A5B7]
             [A6B0 + A7B1 A6B2 + A7B3 A6B4 + A7B5 A6B6 + A7B7]

----------------------------------------------------------------------------

This operation enables the folding of 2 outer products chained via the accumulator into a single outer product.

For example:

%a0_ext = arith.extf %a0 : vector<[4]xf16> to vector<[4]xf32>
%b0_ext = arith.extf %b0 : vector<[4]xf16> to vector<[4]xf32>
%a1_ext = arith.extf %a1 : vector<[4]xf16> to vector<[4]xf32>
%b1_ext = arith.extf %b1 : vector<[4]xf16> to vector<[4]xf32>

%0 = arm_sme.outerproduct %a0_ext, %b0_ext : vector<[4]xf32>, vector<[4]xf32>
%1 = arm_sme.outerproduct %a1_ext, %b1_ext acc(%0) : vector<[4]xf32>, vector<[4]xf32>

The 2 outer products in the example above can be fused into a single outer product as follows:

%a_packed = vector.interleave %a0, %a1 : vector<[4]xf16> -> vector<[8]xf16>
%b_packed = vector.interleave %b0, %b1 : vector<[4]xf16> -> vector<[8]xf16>
%0 = arm_sme.fmopa_2way %a_packed, %b_packed : vector<[8]xf16>, vector<[8]xf16> into vector<[4]x[4]xf32>

This is implemented in the -arm-sme-outer-product-fusion pass.

Example: FP16 to FP32

%result = arm_sme.fmopa_2way $lhs, $rhs : vector<[8]xf16>, vector<[8]xf16> into vector<[4]x[4]xf32>

Example: BF16 to FP32

%result = arm_sme.fmopa_2way $lhs, $rhs : vector<[8]xbf16>, vector<[8]xbf16> into vector<[4]x[4]xf32>
SpecFeatures
FMOPA (widening, 2-way, FP16 to FP32)+sme
BFMOPA (widening, 2-way, BF16 to FP32)+sme

[1] https://developer.arm.com/documentation/ddi0616

fmops_2way()

Return op name arm_sme.fmops_2way as a bitstring.

fmops_2way(ssa)

arm_sme.fmops_2way - Floating-point sum of 2 outer products and subtract

Operands

  • lhs - Single, anonymous/composite constraint, of ranks 1scalable vector of 16-bit float or bfloat16 type values of length 8
  • rhs - Single, AnyVectorOfNonZeroRank, vector of any non-token type values
  • lhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • rhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • acc - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values

Results

  • result - Single, anonymous/composite constraint, vector<[4]x[4]xf32> of 32-bit float values

Description

Equivalent to fmopa_2way but outer products are subtracted from destination result.

Example: FP16 to FP32

%result = arm_sme.fmops_2way $lhs, $rhs : vector<[8]xf16>, vector<[8]xf16> into vector<[4]x[4]xf32>

Example: BF16 to FP32

%result = arm_sme.fmops_2way $lhs, $rhs : vector<[8]xbf16>, vector<[8]xbf16> into vector<[4]x[4]xf32>

Refer to fmopa_2way for a detailed description of 2-way outer products.

SpecFeatures
FMOPS (widening, 2-way, FP16 to FP32)+sme
BFMOPS (widening, 2-way, BF16 to FP32)+sme

get_tile()

Return op name arm_sme.get_tile as a bitstring.

get_tile(ssa)

arm_sme.get_tile - Creates an undefined value of SME virtual tile type

Results

  • tile - Single, SMETile, a vector type that fits into a SME tile

Description

Creates a new SME "virtual tile" value within a function. The contents of the tile returned from this operation are undefined.

Example 1:

// Create an 8-bit element "virtual tile" value:
%za0_b = arm_sme.get_tile: vector<[16]x[16]xi8>

Example 2:

// Create two 16-bit element "virtual tiles" values:
%za0_h = arm_sme.get_tile : vector<[8]x[8]xi16>
%za1_h = arm_sme.get_tile : vector<[8]x[8]xi16>

Example 3:

// Create an 128-bit element "virtual tile" value:
%za0_q = arm_sme.get_tile : vector<[1]x[1]xi128>

insert_tile_slice()

Return op name arm_sme.insert_tile_slice as a bitstring.

insert_tile_slice(ssa)

arm_sme.insert_tile_slice - Insert 1-D scalable vector into slice of 2-D tile

This op has support for result type inference.

Attributes

  • layout - Single, ArmSME_TileSliceLayoutAttr, Layout of a tile slice

Operands

  • vector - Single, SVEVector, a vector type that matches the size of a SVE vector
  • tile - Single, SMETile, a vector type that fits into a SME tile
  • tile_slice_index - Single, Index, index

Results

  • result - Single, SMETile, a vector type that fits into a SME tile

Description

Inserts a 1-D scalable vector into a slice of a 2-D scalable vector tile at the given index. The type of the 1-D scalable vector to be inserted must match the type of the tile slice. A tile slice is a 1-D vector of horizontally or vertically contiguous elements within a ZA tile. The updated tile is returned as the result.

An optional tile slice layout attribute specifies whether the tile slice is horizontal (default) or vertical.

Example 1: Insert vector<[16]xi8> into tile horizontally at the given index.

%tile_update = arm_sme.insert_tile_slice %vector, %tile[%tile_slice_index] : vector<[16]xi8> into vector<[16]x[16]xi8>

Example 2: Insert vector<[2]xf64> into tile vertically at the given index.

%tile_update = arm_sme.insert_tile_slice %vector, %tile[%tile_slice_index] layout<vertical> : vector<[2]xf64> into vector<[2]x[2]xf64>

intr_cntsd()

Return op name arm_sme.intr.cntsd as a bitstring.

intr_cntsd(ssa)

arm_sme.intr.cntsd

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

intr_ld1b_horiz()

Return op name arm_sme.intr.ld1b.horiz as a bitstring.

intr_ld1b_horiz(ssa)

arm_sme.intr.ld1b.horiz

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • load_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_ld1b_vert()

Return op name arm_sme.intr.ld1b.vert as a bitstring.

intr_ld1b_vert(ssa)

arm_sme.intr.ld1b.vert

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • load_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_ld1d_horiz()

Return op name arm_sme.intr.ld1d.horiz as a bitstring.

intr_ld1d_horiz(ssa)

arm_sme.intr.ld1d.horiz

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • load_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_ld1d_vert()

Return op name arm_sme.intr.ld1d.vert as a bitstring.

intr_ld1d_vert(ssa)

arm_sme.intr.ld1d.vert

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • load_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_ld1h_horiz()

Return op name arm_sme.intr.ld1h.horiz as a bitstring.

intr_ld1h_horiz(ssa)

arm_sme.intr.ld1h.horiz

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • load_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_ld1h_vert()

Return op name arm_sme.intr.ld1h.vert as a bitstring.

intr_ld1h_vert(ssa)

arm_sme.intr.ld1h.vert

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • load_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_ld1q_horiz()

Return op name arm_sme.intr.ld1q.horiz as a bitstring.

intr_ld1q_horiz(ssa)

arm_sme.intr.ld1q.horiz

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • load_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_ld1q_vert()

Return op name arm_sme.intr.ld1q.vert as a bitstring.

intr_ld1q_vert(ssa)

arm_sme.intr.ld1q.vert

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • load_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_ld1w_horiz()

Return op name arm_sme.intr.ld1w.horiz as a bitstring.

intr_ld1w_horiz(ssa)

arm_sme.intr.ld1w.horiz

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • load_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_ld1w_vert()

Return op name arm_sme.intr.ld1w.vert as a bitstring.

intr_ld1w_vert(ssa)

arm_sme.intr.ld1w.vert

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • load_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_mopa()

Return op name arm_sme.intr.mopa as a bitstring.

intr_mopa(ssa)

arm_sme.intr.mopa

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • lhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • rhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • lhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions
  • rhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions

intr_mopa_wide()

Return op name arm_sme.intr.mopa.wide as a bitstring.

intr_mopa_wide(ssa)

arm_sme.intr.mopa.wide

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • lhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • rhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • lhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions
  • rhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions

intr_mops()

Return op name arm_sme.intr.mops as a bitstring.

intr_mops(ssa)

arm_sme.intr.mops

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • lhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • rhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • lhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions
  • rhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions

intr_mops_wide()

Return op name arm_sme.intr.mops.wide as a bitstring.

intr_mops_wide(ssa)

arm_sme.intr.mops.wide

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • lhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • rhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • lhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions
  • rhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions

intr_read_horiz()

Return op name arm_sme.intr.read.horiz as a bitstring.

intr_read_horiz(ssa)

arm_sme.intr.read.horiz

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • vector - Single, SVEVector, a vector type that matches the size of a SVE vector
  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • tile_slice_index - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

intr_read_vert()

Return op name arm_sme.intr.read.vert as a bitstring.

intr_read_vert(ssa)

arm_sme.intr.read.vert

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • vector - Single, SVEVector, a vector type that matches the size of a SVE vector
  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • tile_slice_index - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

intr_smopa_wide()

Return op name arm_sme.intr.smopa.wide as a bitstring.

intr_smopa_wide(ssa)

arm_sme.intr.smopa.wide

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • lhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • rhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • lhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions
  • rhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions

intr_smopa_za32()

Return op name arm_sme.intr.smopa.za32 as a bitstring.

intr_smopa_za32(ssa)

arm_sme.intr.smopa.za32

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • lhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • rhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • lhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions
  • rhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions

intr_smops_wide()

Return op name arm_sme.intr.smops.wide as a bitstring.

intr_smops_wide(ssa)

arm_sme.intr.smops.wide

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • lhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • rhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • lhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions
  • rhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions

intr_smops_za32()

Return op name arm_sme.intr.smops.za32 as a bitstring.

intr_smops_za32(ssa)

arm_sme.intr.smops.za32

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • lhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • rhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • lhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions
  • rhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions

intr_st1b_horiz()

Return op name arm_sme.intr.st1b.horiz as a bitstring.

intr_st1b_horiz(ssa)

arm_sme.intr.st1b.horiz

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • store_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_st1b_vert()

Return op name arm_sme.intr.st1b.vert as a bitstring.

intr_st1b_vert(ssa)

arm_sme.intr.st1b.vert

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • store_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_st1d_horiz()

Return op name arm_sme.intr.st1d.horiz as a bitstring.

intr_st1d_horiz(ssa)

arm_sme.intr.st1d.horiz

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • store_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_st1d_vert()

Return op name arm_sme.intr.st1d.vert as a bitstring.

intr_st1d_vert(ssa)

arm_sme.intr.st1d.vert

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • store_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_st1h_horiz()

Return op name arm_sme.intr.st1h.horiz as a bitstring.

intr_st1h_horiz(ssa)

arm_sme.intr.st1h.horiz

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • store_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_st1h_vert()

Return op name arm_sme.intr.st1h.vert as a bitstring.

intr_st1h_vert(ssa)

arm_sme.intr.st1h.vert

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • store_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_st1q_horiz()

Return op name arm_sme.intr.st1q.horiz as a bitstring.

intr_st1q_horiz(ssa)

arm_sme.intr.st1q.horiz

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • store_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_st1q_vert()

Return op name arm_sme.intr.st1q.vert as a bitstring.

intr_st1q_vert(ssa)

arm_sme.intr.st1q.vert

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • store_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_st1w_horiz()

Return op name arm_sme.intr.st1w.horiz as a bitstring.

intr_st1w_horiz(ssa)

arm_sme.intr.st1w.horiz

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • store_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_st1w_vert()

Return op name arm_sme.intr.st1w.vert as a bitstring.

intr_st1w_vert(ssa)

arm_sme.intr.st1w.vert

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • store_address - Single, LLVM_AnyPointer, LLVM pointer type
  • tile_slice_index - Single, I32, 32-bit signless integer

intr_str()

Return op name arm_sme.intr.str as a bitstring.

intr_str(ssa)

arm_sme.intr.str

Operands

  • index - Single, I32, 32-bit signless integer
  • store_address - Single, LLVM_AnyPointer, LLVM pointer type
  • offset - Single, I32, 32-bit signless integer

intr_sumopa_wide()

Return op name arm_sme.intr.sumopa.wide as a bitstring.

intr_sumopa_wide(ssa)

arm_sme.intr.sumopa.wide

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • lhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • rhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • lhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions
  • rhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions

intr_sumops_wide()

Return op name arm_sme.intr.sumops.wide as a bitstring.

intr_sumops_wide(ssa)

arm_sme.intr.sumops.wide

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • lhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • rhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • lhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions
  • rhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions

intr_umopa_wide()

Return op name arm_sme.intr.umopa.wide as a bitstring.

intr_umopa_wide(ssa)

arm_sme.intr.umopa.wide

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • lhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • rhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • lhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions
  • rhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions

intr_umopa_za32()

Return op name arm_sme.intr.umopa.za32 as a bitstring.

intr_umopa_za32(ssa)

arm_sme.intr.umopa.za32

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • lhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • rhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • lhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions
  • rhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions

intr_umops_wide()

Return op name arm_sme.intr.umops.wide as a bitstring.

intr_umops_wide(ssa)

arm_sme.intr.umops.wide

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • lhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • rhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • lhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions
  • rhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions

intr_umops_za32()

Return op name arm_sme.intr.umops.za32 as a bitstring.

intr_umops_za32(ssa)

arm_sme.intr.umops.za32

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • lhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • rhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • lhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions
  • rhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions

intr_usmopa_wide()

Return op name arm_sme.intr.usmopa.wide as a bitstring.

intr_usmopa_wide(ssa)

arm_sme.intr.usmopa.wide

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • lhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • rhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • lhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions
  • rhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions

intr_usmops_wide()

Return op name arm_sme.intr.usmops.wide as a bitstring.

intr_usmops_wide(ssa)

arm_sme.intr.usmops.wide

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • lhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • rhs_predicate - Single, MOPPredicate, a vector type that is a supported predicate for the SME MOP instructions
  • lhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions
  • rhs_vector - Single, MOPVector, a vector type that is a supported input for the SME MOP instructions

intr_write_horiz()

Return op name arm_sme.intr.write.horiz as a bitstring.

intr_write_horiz(ssa)

arm_sme.intr.write.horiz

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • tile_slice_index - Single, I32, 32-bit signless integer
  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • vector - Single, SVEVector, a vector type that matches the size of a SVE vector

intr_write_vert()

Return op name arm_sme.intr.write.vert as a bitstring.

intr_write_vert(ssa)

arm_sme.intr.write.vert

Attributes

  • tile_id - Single, I32Attr, 32-bit signless integer attribute

Operands

  • tile_slice_index - Single, I32, 32-bit signless integer
  • predicate - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • vector - Single, SVEVector, a vector type that matches the size of a SVE vector

intr_zero()

Return op name arm_sme.intr.zero as a bitstring.

intr_zero(ssa)

arm_sme.intr.zero

Attributes

  • tile_mask - Single, I32Attr, 32-bit signless integer attribute

load_tile_slice()

Return op name arm_sme.load_tile_slice as a bitstring.

load_tile_slice(ssa)

arm_sme.load_tile_slice - Tile slice load and update operation

This op has support for result type inference.

Attributes

  • layout - Single, ArmSME_TileSliceLayoutAttr, Layout of a tile slice

Operands

  • base - Single, AnyMemRef, memref of any non-token type values
  • mask - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • tile - Single, SMETile, a vector type that fits into a SME tile
  • indices - Variadic, Index, variadic of index
  • tile_slice_index - Single, Index, index

Results

  • result - Single, SMETile, a vector type that fits into a SME tile

Description

Loads a 1D tile slice from memory into a 2D SME "virtual tile". The tile slice is defined by the dimension of the 2D scalable vector type pointed by the index. A tile slice index describes where in the input tile the tile slice is loaded to. An optional tile slice layout attribute specifies whether the tile slice being loaded at the given index is horizontal (default) or vertical. The updated tile is returned as the result.

The slice of memory read is defined by a base and indices and must be contiguous. The memref must be either rank 1 or rank 2, have dynamic dimensions since the operation is scalable, and the element type must be a scalar that matches the element type of the result.

The provided mask is used to specify which elements of the tile slice will be loaded.

Example 1: Load a vector<[16]xi8> tile slice from memory into tile horizontally (default) at given index.

%tile_update = arm_sme.load_tile_slice %base[%c0], %mask, %tile, %tile_slice_index : memref<?x?xi8>, vector<[16]xi1>, vector<[16]x[16]xi8>

Example 2: Load a vector<[4]xf32> tile slice from memory into tile vertically at given index.

%tile_update = arm_sme.load_tile_slice %base[%c0], %mask, %tile, %tile_slice_index layout<vertical> : memref<?x?xf32>, vector<[4]xi1>, vector<[4]x[4]xf32>

Example 3: Load a vector<[1]xi128> tile slice from memory into tile vertically at given index.

%tile_update = arm_sme.load_tile_slice %base[%c0], %mask, %tile, %tile_slice_index layout<vertical> : memref<?x?xi128>, vector<[1]xi1>, vector<[1]x[1]xi128>

outerproduct()

Return op name arm_sme.outerproduct as a bitstring.

outerproduct(ssa)

arm_sme.outerproduct - Outer product with optional fused add/sub

This op has support for result type inference.

Attributes

  • kind - Single, ArmSME_CombiningKindAttr, Kind of combining function

Operands

  • lhs - Single, SVEVector, a vector type that matches the size of a SVE vector
  • rhs - Single, SVEVector, a vector type that matches the size of a SVE vector
  • lhsMask - Optional, SVEPredicate, a vector type that matches the size of a SVE predicate
  • rhsMask - Optional, SVEPredicate, a vector type that matches the size of a SVE predicate
  • acc - Optional, SMETile, a vector type that fits into a SME tile

Results

  • result - Single, SMETile, a vector type that fits into a SME tile

Description

This operation represents an outer product that fits within an SME tile. All operands must be SVE vectors and the result a SME tile. Unlike vector.outerproduct masking is on the operands (rather than the result), which mirrors the SME instructions.

Example 1: Unmasked outerproduct (without accumulator)

// Not specifying an accumulator implicitly zeros the destination tile.
%result = arm_sme.outerproduct $lhs, $rhs : vector<[4]xf32>, vector<[4]xf32>

Example 2: Unmasked outerproduct (with accumulator)

%result = arm_sme.outerproduct $lhs, $rhs acc($accumulator)
            : vector<[4]xf32>, vector<[4]xf32>

Example 3: Masked outerproduct

%result = arm_sme.outerproduct $lhs, $rhs masks($lhsMask, $rhsMask)
            : vector<[4]xf32>, vector<[4]xf32>

Example 4: Masked outerproduct (with accumulator)

%result = arm_sme.outerproduct $lhs, $rhs acc($accumulator) masks($lhsMask, $rhsMask)
            : vector<[4]xf32>, vector<[4]xf32>

smopa_2way()

Return op name arm_sme.smopa_2way as a bitstring.

smopa_2way(ssa)

arm_sme.smopa_2way - Signed integer sum of 2 outer products and accumulate

Operands

  • lhs - Single, anonymous/composite constraint, of ranks 1scalable vector of 16-bit signless integer values of length 8
  • rhs - Single, AnyVectorOfNonZeroRank, vector of any non-token type values
  • lhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • rhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • acc - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values

Results

  • result - Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values

Description

Example:

%result = arm_sme.smopa_2way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[4]x[4]xi32>

Refer to fmopa_2way for a detailed description of 2-way outer products.

SpecFeatures
SMOPA (2-way)+sme2

smopa_4way()

Return op name arm_sme.smopa_4way as a bitstring.

smopa_4way(ssa)

arm_sme.smopa_4way - Signed integer sum of 4 outer products and accumulate

Operands

  • lhs - Single, anonymous/composite constraint, of ranks 1scalable vector of 8-bit signless integer values of length 16 or of ranks 1scalable vector of 16-bit signless integer values of length 8
  • rhs - Single, AnyVectorOfNonZeroRank, vector of any non-token type values
  • lhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • rhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • acc - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values

Results

  • result - Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values or vector<[2]x[2]xi64> of 64-bit signless integer values

Description

This operation represents a sum of 4 widened outer products. It takes 2 1-D scalable vectors as input and a 2-D scalable vector (ZA tile) as output.

For example (i8 to i32):

%result = arm_sme.smopa_4way $lhs, $rhs :
  vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>

The lhs encodes a matrix of shape SVLSx4 and the rhs a matrix of 4xSVLS, where SVLS (spec [1], section B2.1) is the number of 32-bit elements in a vector of SVL bits. To illustrate, below is a breakdown of this operation for i8 to i32, SVL=128 (i.e., vscale=1):

                                    LHS
          [A0 A1 A2 A3 A4 A5 A6 A7 A8 A9 A10 A11 A12 A15 A14 A15]

                                    RHS
          [B0 B1 B2 B3 B4 B5 B6 B7 B8 B9 B10 B11 B12 B13 B14 B15]

----------------------------------------------------------------------------

                              implicit layout

                [A0   A1  A2  A3]    |    [B0 B4  B8 B12]
                [A4   A5  A6  A7]    |    [B1 B5  B9 B13]
                [A8   A9 A10 A11]    |    [B2 B6 B10 B14]
                [A12 A13 A14 A15]    |    [B3 B7 B11 B15]

----------------------------------------------------------------------------

                              4 outer products

             Acol0  Brow0           |            Acol1  Brow1
             -------------           |            -------------
                                     |
         [B0 B4 B8 B12]              |        [B1 B5 B9 B13]
                                     |
   [A0   [ A0B0  A0B4  A0B8  A0B12]  |  [A1   [ A1B1  A1B5  A1B9  A1B13]
    A4   [ A4B0  A4B4  A4B8  A4B12]  |   A5   [ A5B1  A5B5  A5B9  A5B13]
    A8   [ A8B0  A8B4  A8B8  A8B12]  |   A9   [ A9B1  A9B5  A9B9  A9B13]
    A12] [A12B0 A12B4 A12B8 A12B12]  |   A13] [A13B1 A13B5 A13B9 A13B13]
                                     |
             Acol2  Brow2           |            Acol3  Brow3
             -------------           |            -------------
                                     |
         [B2, B6, B10, B14]          |        [B3 B7 B11 B15]
                                     |
   [A2   [ A2B2  A2B6  A2B10  A2B14] |  [A3   [ A3B3  A3B7  A3B11  A3B15]
    A6   [ A6B2  A6B6  A6B10  A6B14] |   A7   [ A7B3  A7B7  A7B11  A7B15]
    A10  [A10B2 A10B6 A10B10 A10B14] |   A11  [A11B3 A11B7 A11B11 A11B15]
    A14] [A14B2 A14B6 A14B10 A14B14] |   A15] [A15B3 A15B7 A15B11 A15B15]
                                     |

----------------------------------------------------------------------------

                          sum of 4 outer products

       Acol0  Brow0 + Acol1  Brow1 + Acol2  Brow2 + Acol3  Brow3

 [ A0B0 +  A1B1 +  A2B2 +  A3B3 ... ...  A0B12 +  A1B13 +  A2B14 +  A3B15]
 [ A4B0 +  A5B1 +  A6B2 +  A7B3 ... ...  A4B12 +  A5B13 +  A6B14 +  A7B15]
 [ A8B0 +  A9B1 + A10B2 + A11B3 ... ...  A8B12 +  A9B13 + A10B14 + A11B15]
 [A12B0 + A13B1 + A14B2 + A15B3 ... ... A12B12 + A13B13 + A14B14 + A15B15]

----------------------------------------------------------------------------

This operation enables the folding of 4 outer products chained via the accumulator into a single outer product.

For example:

%a0_ext = arith.extsi %a0 : vector<[4]xi8> to vector<[4]xi32>
%b0_ext = arith.extsi %b0 : vector<[4]xi8> to vector<[4]xi32>

%a1_ext = arith.extsi %a1 : vector<[4]xi8> to vector<[4]xi32>
%b1_ext = arith.extsi %b1 : vector<[4]xi8> to vector<[4]xi32>

%a2_ext = arith.extsi %a2 : vector<[4]xi8> to vector<[4]xi32>
%b2_ext = arith.extsi %b2 : vector<[4]xi8> to vector<[4]xi32>

%a3_ext = arith.extsi %a3 : vector<[4]xi8> to vector<[4]xi32>
%b3_ext = arith.extsi %b3 : vector<[4]xi8> to vector<[4]xi32>

%0 = arm_sme.outerproduct %a0_ext, %b0_ext : vector<[4]xi32>, vector<[4]xi32>
%1 = arm_sme.outerproduct %a1_ext, %b1_ext acc(%0) : vector<[4]xi32>, vector<[4]xi32>
%2 = arm_sme.outerproduct %a2_ext, %b2_ext acc(%1) : vector<[4]xi32>, vector<[4]xi32>
%3 = arm_sme.outerproduct %a3_ext, %b3_ext acc(%2) : vector<[4]xi32>, vector<[4]xi32>

The 4 outer products in the example above can be fused into a single outer product as follows:

%lhs0 = vector.interleave %a0, %a2 : vector<[4]xi8> -> vector<[8]xi8>
%lhs1 = vector.interleave %a1, %a3 : vector<[4]xi8> -> vector<[8]xi8>
%lhs = vector.interleave %lhs0, %lhs1 : vector<[8]xi8> -> vector<[16]xi8>

%rhs0 = vector.interleave %b0, %b2 : vector<[4]xi8> -> vector<[8]xi8>
%rhs1 = vector.interleave %b1, %b3 : vector<[4]xi8> -> vector<[8]xi8>
%rhs = vector.interleave %rhs0, %rhs1 : vector<[8]xi8> -> vector<[16]xi8>

%0 = arm_sme.smopa_4way %lhs, %rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>

This is implemented in the -arm-sme-outer-product-fusion pass.

Example: I8 to I32

%result = arm_sme.smopa_4way $lhs, $rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>

Example: I16 to I64

%result = arm_sme.smopa_4way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[2]x[2]xi64>
SpecFeatures
SMOPA (4-way)+sme (32-bit), +sme-i16i64 (64-bit)

smops_2way()

Return op name arm_sme.smops_2way as a bitstring.

smops_2way(ssa)

arm_sme.smops_2way - Signed integer sum of 2 outer products and subtract

Operands

  • lhs - Single, anonymous/composite constraint, of ranks 1scalable vector of 16-bit signless integer values of length 8
  • rhs - Single, AnyVectorOfNonZeroRank, vector of any non-token type values
  • lhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • rhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • acc - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values

Results

  • result - Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values

Description

Example:

%result = arm_sme.smops_2way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[4]x[4]xi32>

Refer to fmopa_2way for a detailed description of 2-way outer products.

SpecFeatures
SMOPS (2-way)+sme2

smops_4way()

Return op name arm_sme.smops_4way as a bitstring.

smops_4way(ssa)

arm_sme.smops_4way - Signed integer sum of 4 outer products and subtract

Operands

  • lhs - Single, anonymous/composite constraint, of ranks 1scalable vector of 8-bit signless integer values of length 16 or of ranks 1scalable vector of 16-bit signless integer values of length 8
  • rhs - Single, AnyVectorOfNonZeroRank, vector of any non-token type values
  • lhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • rhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • acc - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values

Results

  • result - Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values or vector<[2]x[2]xi64> of 64-bit signless integer values

Description

Equivalent to smopa_4way but outer products are subtracted from destination result.

Example: I8 to I32

%result = arm_sme.smops_4way $lhs, $rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>

Example: I16 to I64

%result = arm_sme.smops_4way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[2]x[2]xi64>

Refer to smopa_4way for a detailed description of 4-way outer products.

SpecFeatures
SMOPS (4-way)+sme (32-bit), +sme-i16i64 (64-bit)

store_tile_slice()

Return op name arm_sme.store_tile_slice as a bitstring.

store_tile_slice(ssa)

arm_sme.store_tile_slice - Tile slice store operation

Attributes

  • layout - Single, ArmSME_TileSliceLayoutAttr, Layout of a tile slice

Operands

  • tile - Single, SMETile, a vector type that fits into a SME tile
  • tile_slice_index - Single, Index, index
  • mask - Single, SVEPredicate, a vector type that matches the size of a SVE predicate
  • base - Single, AnyMemRef, memref of any non-token type values
  • indices - Variadic, Index, variadic of index

Description

Stores a 1D tile slice from a 2D SME "virtual tile" into memory. The tile slice is defined by the dimension of the 2D scalable vector type pointed by the index. A tile slice index describes where in the input tile the tile slice is stored from. An optional tile slice layout attribute specifies whether the tile slice being stored from the given index is horizontal (default) or vertical.

The slice of memory written is defined by a base and indices and must be contiguous. The memref must be either rank 1 or rank 2, have dynamic dimensions since the operation is scalable, and the element type must be a scalar that matches the element type of the input tile.

The provided mask is used to specify which elements of the tile slice will be stored.

Example 1: Store vector<[16]xi8> horizontal (default) tile slice from tile at given index to memory.

arm_sme.store_tile_slice %tile, %tile_slice_index, %mask, %base[%c0] : vector<[16]x[16]xi8>, vector<[16]xi1>, memref<?x?xi8>

Example 2: Store vector<[4]xf32> vertical tile slice from tile at given index to memory.

arm_sme.store_tile_slice %tile, %tile_slice_index, %mask, %base[%c0] layout<vertical> : vector<[4]x[4]xf32>, vector<[4]xi1>, memref<?x?xf32>

Example 3: Store a vector<[1]xi128> vertical tile slice from tile at given index to memory.

arm_sme.store_tile_slice %tile, %tile_slice_index, %mask, %base[%c0] layout<vertical> : vector<[1]x[1]xi128>, vector<[1]xi1>, memref<?x?xi128>

streaming_vl()

Return op name arm_sme.streaming_vl as a bitstring.

streaming_vl(ssa)

arm_sme.streaming_vl - Query the streaming vector length

This op has support for result type inference.

Attributes

  • type_size - Single, ArmSME_TypeSizeAttr, Size of a vector element type

Results

  • anonymous - Single, Index, index

Description

This operation returns the streaming vector length (SVL) for a given type size. Unlike vector.vscale the value returned is invariant to the streaming mode.

Example:

// Streaming vector length in:
// - bytes (8-bit, SVL.B)
%svl_b = arm_sme.streaming_vl <byte>
// - half words (16-bit, SVL.H)
%svl_h = arm_sme.streaming_vl <half>
// - words (32-bit, SVL.W)
%svl_w = arm_sme.streaming_vl <word>
// - double words (64-bit, SVL.D)
%svl_d = arm_sme.streaming_vl <double>

sumopa_4way()

Return op name arm_sme.sumopa_4way as a bitstring.

sumopa_4way(ssa)

arm_sme.sumopa_4way - Signed by unsigned integer sum of 4 outer products and accumulate

Operands

  • lhs - Single, anonymous/composite constraint, of ranks 1scalable vector of 8-bit signless integer values of length 16 or of ranks 1scalable vector of 16-bit signless integer values of length 8
  • rhs - Single, AnyVectorOfNonZeroRank, vector of any non-token type values
  • lhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • rhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • acc - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values

Results

  • result - Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values or vector<[2]x[2]xi64> of 64-bit signless integer values

Description

Example: I8 to I32

%result = arm_sme.sumopa_4way $lhs, $rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>

Example: I16 to I64

%result = arm_sme.sumopa_4way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[2]x[2]xi64>

Refer to smopa_4way for a detailed description of 4-way outer products.

SpecFeatures
SUMOPA (4-way)+sme (32-bit), +sme-i16i64 (64-bit)

sumops_4way()

Return op name arm_sme.sumops_4way as a bitstring.

sumops_4way(ssa)

arm_sme.sumops_4way - Signed by unsigned integer sum of 4 outer products and subtract

Operands

  • lhs - Single, anonymous/composite constraint, of ranks 1scalable vector of 8-bit signless integer values of length 16 or of ranks 1scalable vector of 16-bit signless integer values of length 8
  • rhs - Single, AnyVectorOfNonZeroRank, vector of any non-token type values
  • lhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • rhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • acc - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values

Results

  • result - Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values or vector<[2]x[2]xi64> of 64-bit signless integer values

Description

Example: I8 to I32

%result = arm_sme.sumops_4way $lhs, $rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>

Example: I16 to I64

%result = arm_sme.sumops_4way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[2]x[2]xi64>

Refer to smopa_4way for a detailed description of 4-way outer products.

SpecFeatures
SUMOPS (4-way)+sme (32-bit), +sme-i16i64 (64-bit)

tile_load()

Return op name arm_sme.tile_load as a bitstring.

tile_load(ssa)

arm_sme.tile_load - Tile load operation

Attributes

  • layout - Single, ArmSME_TileSliceLayoutAttr, Layout of a tile slice

Operands

  • base - Single, anonymous/composite constraint, 2D memref of any non-token type values
  • indices - Variadic, Index, variadic of index
  • padding - Optional, AnyType, any non-token type
  • mask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values

Results

  • result - Single, SMETile, a vector type that fits into a SME tile

Description

Loads a 2D SME "virtual tile" from memory defined by a base and indices, with the shape defined by the 2D scalable vector type of the result tile. An optional tile slice layout attribute specifies whether the slices of the tile being loaded are horizontal (default) or vertical. The slice of memory must be contiguous. The memref must be either rank 1 or rank 2 with dynamic dimensions, since the operation is scalable, and the element type must be a scalar that matches the element type of the result.

An optional SSA value padding of the same elemental type as the MemRef is provided to specify a fallback value in the case of masking.

An optional SSA value mask may be specified to mask out elements read from the MemRef. The mask type is an i1 vector with a shape that matches how elements are read from the MemRef. Elements whose corresponding mask element is 0 are masked out and replaced with padding.

If either padding or mask are specified, both must be specified.

Example 1: Load an 8-bit element ZA tile with horizontal layout (default) from memory (ZA0.B).

%tile = arm_sme.tile_load %base[%c0, %c0] : memref<?x?xi8>, vector<[16]x[16]xi8>

Example 2: Load a FP 32-bit element ZA tile with vertical layout from memory.

%tile = arm_sme.tile_load %base[%c0, %c0] layout<vertical> : memref<?x?xf32>, vector<[4]x[4]xf32>

Example 3: Load a 128-bit element ZA tile with horizontal layout (default) from memory.

%tile = arm_sme.tile_load %base[%c0, %c0] layout<horizontal> : memref<?x?xi128>, vector<[1]x[1]xi128>

Example 4: Masked load of int 32-bit element ZA tile with horizontal layout (default) from memory.

%tile = arm_sme.tile_load %base[%c0, %c0], %pad, %mask : memref<?x?xf32>, vector<[4]x[4]xf32>

tile_store()

Return op name arm_sme.tile_store as a bitstring.

tile_store(ssa)

arm_sme.tile_store - Tile store operation

Attributes

  • layout - Single, ArmSME_TileSliceLayoutAttr, Layout of a tile slice

Operands

  • valueToStore - Single, SMETile, a vector type that fits into a SME tile
  • base - Single, anonymous/composite constraint, 2D memref of any non-token type values
  • indices - Variadic, Index, variadic of index
  • mask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values

Description

Stores a 2D SME "virtual tile" to memory defined by a base and indices, with the shape defined by the 2D scalable vector type of the tile being stored. An optional tile slice layout attribute specifies whether the slices of the tile being stored are horizontal (default) or vertical. The slice of memory must be contiguous. The memref must be either rank 1 or rank 2 with dynamic dimensions, since the operation is scalable, and the element type must be a scalar that matches the element type of the result.

An optional mask may be provided, the shape of which corresponds to the tile, and selects which elements of the tile will be stored.

Example 1: Store an 8-bit element ZA tile with horizontal (default) layout to memory (ZA0.B).

arm_sme.tile_store %tile, %base[%c0, %c0] : vector<[16]x[16]xi8>, memref<?x?xi8>

Example 2: Store a FP 32-bit element ZA tile with vertical layout to memory.

arm_sme.tile_store %tile, %base[%c0, %c0] layout<vertical> : vector<[4]x[4]xf32>, memref<?x?xf32>

Example 3: Store a 128-bit element ZA tile with horizontal (default) layout to memory.

arm_sme.tile_store %tile, %base[%c0, %c0] layout<horizontal> : vector<[1]x[1]xi128>, memref<?x?xi128>

Example 4: Masked store a int 32-bit element ZA tile with vertical layout to memory.

arm_sme.tile_store %tile, %base[%c0, %c0], %mask layout<vertical> : vector<[4]x[4]xf32>, memref<?x?xf32>

umopa_2way()

Return op name arm_sme.umopa_2way as a bitstring.

umopa_2way(ssa)

arm_sme.umopa_2way - Unsiged integer sum of 2 outer products and accumulate

Operands

  • lhs - Single, anonymous/composite constraint, of ranks 1scalable vector of 16-bit signless integer values of length 8
  • rhs - Single, AnyVectorOfNonZeroRank, vector of any non-token type values
  • lhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • rhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • acc - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values

Results

  • result - Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values

Description

Example:

%result = arm_sme.umopa_2way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[4]x[4]xi32>

Refer to fmopa_2way for a detailed description of 2-way outer products.

SpecFeatures
UMOPA (2-way)+sme2

umopa_4way()

Return op name arm_sme.umopa_4way as a bitstring.

umopa_4way(ssa)

arm_sme.umopa_4way - Unsigned integer sum of 4 outer products and accumulate

Operands

  • lhs - Single, anonymous/composite constraint, of ranks 1scalable vector of 8-bit signless integer values of length 16 or of ranks 1scalable vector of 16-bit signless integer values of length 8
  • rhs - Single, AnyVectorOfNonZeroRank, vector of any non-token type values
  • lhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • rhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • acc - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values

Results

  • result - Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values or vector<[2]x[2]xi64> of 64-bit signless integer values

Description

Example: I8 to I32

%result = arm_sme.umopa_4way $lhs, $rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>

Example: I16 to I64

%result = arm_sme.umopa_4way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[2]x[2]xi64>

Refer to smopa_4way for a detailed description of 4-way outer products.

SpecFeatures
UMOPA (4-way)+sme (32-bit), +sme-i16i64 (64-bit)

umops_2way()

Return op name arm_sme.umops_2way as a bitstring.

umops_2way(ssa)

arm_sme.umops_2way - Unsiged integer sum of 2 outer products and subtract

Operands

  • lhs - Single, anonymous/composite constraint, of ranks 1scalable vector of 16-bit signless integer values of length 8
  • rhs - Single, AnyVectorOfNonZeroRank, vector of any non-token type values
  • lhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • rhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • acc - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values

Results

  • result - Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values

Description

Example:

%result = arm_sme.umops_2way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[4]x[4]xi32>

Refer to fmopa_2way for a detailed description of 2-way outer products.

SpecFeatures
UMOPS (2-way)+sme2

umops_4way()

Return op name arm_sme.umops_4way as a bitstring.

umops_4way(ssa)

arm_sme.umops_4way - Unsigned integer sum of 4 outer products and subtract

Operands

  • lhs - Single, anonymous/composite constraint, of ranks 1scalable vector of 8-bit signless integer values of length 16 or of ranks 1scalable vector of 16-bit signless integer values of length 8
  • rhs - Single, AnyVectorOfNonZeroRank, vector of any non-token type values
  • lhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • rhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • acc - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values

Results

  • result - Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values or vector<[2]x[2]xi64> of 64-bit signless integer values

Description

Example: I8 to I32

%result = arm_sme.umops_4way $lhs, $rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>

Example: I16 to I64

%result = arm_sme.umops_4way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[2]x[2]xi64>

Refer to smopa_4way for a detailed description of 4-way outer products.

SpecFeatures
UMOPS (4-way)+sme (32-bit), +sme-i16i64 (64-bit)

usmopa_4way()

Return op name arm_sme.usmopa_4way as a bitstring.

usmopa_4way(ssa)

arm_sme.usmopa_4way - Unsigned by signed integer sum of 4 outer products and accumulate

Operands

  • lhs - Single, anonymous/composite constraint, of ranks 1scalable vector of 8-bit signless integer values of length 16 or of ranks 1scalable vector of 16-bit signless integer values of length 8
  • rhs - Single, AnyVectorOfNonZeroRank, vector of any non-token type values
  • lhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • rhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • acc - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values

Results

  • result - Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values or vector<[2]x[2]xi64> of 64-bit signless integer values

Description

Example: I8 to I32

%result = arm_sme.usmopa_4way $lhs, $rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>

Example: I16 to I64

%result = arm_sme.usmopa_4way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[2]x[2]xi64>

Refer to smopa_4way for a detailed description of 4-way outer products.

SpecFeatures
USMOPA (4-way)+sme (32-bit), +sme-i16i64 (64-bit)

usmops_4way()

Return op name arm_sme.usmops_4way as a bitstring.

usmops_4way(ssa)

arm_sme.usmops_4way - Unsigned by signed integer sum of 4 outer products and subtract

Operands

  • lhs - Single, anonymous/composite constraint, of ranks 1scalable vector of 8-bit signless integer values of length 16 or of ranks 1scalable vector of 16-bit signless integer values of length 8
  • rhs - Single, AnyVectorOfNonZeroRank, vector of any non-token type values
  • lhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • rhsMask - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values
  • acc - Optional, AnyVectorOfNonZeroRank, vector of any non-token type values

Results

  • result - Single, anonymous/composite constraint, vector<[4]x[4]xi32> of 32-bit signless integer values or vector<[2]x[2]xi64> of 64-bit signless integer values

Description

Example: I8 to I32

%result = arm_sme.usmops_4way $lhs, $rhs : vector<[16]xi8>, vector<[16]xi8> into vector<[4]x[4]xi32>

Example: I16 to I64

%result = arm_sme.usmops_4way $lhs, $rhs : vector<[8]xi16>, vector<[8]xi16> into vector<[2]x[2]xi64>

Refer to smopa_4way for a detailed description of 4-way outer products.

SpecFeatures
USMOPS (4-way)+sme (32-bit), +sme-i16i64 (64-bit)

zero()

Return op name arm_sme.zero as a bitstring.

zero(ssa)

arm_sme.zero - Creates a zero-initialized value of SME virtual tile type

Results

  • res - Single, SMETile, a vector type that fits into a SME tile

Description

Creates a new SME "virtual tile" value within a function. The contents of the tile returned from this operation are zero-initialized.

Example 1: Zero an 8-bit element ZA tile.

%0 = arm_sme.zero : vector<[16]x[16]xi8>

Example 2: Zero a 64-bit element ZA tile.

%0 = arm_sme.zero : vector<[2]x[2]xi64>