Beaver.MLIR.Dialect.ROCDL (beaver v0.4.8)

Copy Markdown

Summary

Functions

Return op name rocdl.asyncmark as a bitstring.

rocdl.asyncmark - Mark the end of a group of asynchronous operations

Return op name rocdl.ballot as a bitstring.

rocdl.ballot - Vote across thread group

Return op name rocdl.barrier as a bitstring.

rocdl.barrier

Return op name rocdl.cluster.id.x as a bitstring.

rocdl.cluster.id.x

Return op name rocdl.cluster.id.y as a bitstring.

rocdl.cluster.id.y

Return op name rocdl.cluster.id.z as a bitstring.

rocdl.cluster.id.z

Return op name rocdl.cluster.load.async.to.lds.b8 as a bitstring.

rocdl.cluster.load.async.to.lds.b8

Return op name rocdl.cluster.load.async.to.lds.b32 as a bitstring.

rocdl.cluster.load.async.to.lds.b32

Return op name rocdl.cluster.load.async.to.lds.b64 as a bitstring.

rocdl.cluster.load.async.to.lds.b64

Return op name rocdl.cluster.load.async.to.lds.b128 as a bitstring.

rocdl.cluster.load.async.to.lds.b128

Return op name rocdl.cluster.workgroup.id.x as a bitstring.

rocdl.cluster.workgroup.id.x

Return op name rocdl.cluster.workgroup.id.y as a bitstring.

rocdl.cluster.workgroup.id.y

Return op name rocdl.cluster.workgroup.id.z as a bitstring.

rocdl.cluster.workgroup.id.z

Return op name rocdl.cos as a bitstring.

rocdl.cos

Return op name rocdl.cvt.f32.bf8 as a bitstring.

rocdl.cvt.f32.bf8 - Convert bf8 to f32

Return op name rocdl.cvt.f32.fp8 as a bitstring.

rocdl.cvt.f32.fp8 - Convert fp8 to f32

Return op name rocdl.cvt.pk.bf8.f32 as a bitstring.

rocdl.cvt.pk.bf8.f32 - Convert two f32's to bf8

Return op name rocdl.cvt.pk.f32.bf8 as a bitstring.

rocdl.cvt.pk.f32.bf8 - Convert packed bf8 to packed f32

Return op name rocdl.cvt.pk.f32.fp8 as a bitstring.

rocdl.cvt.pk.f32.fp8 - Convert packed fp8 to packed f32

Return op name rocdl.cvt.pk.fp8.f32 as a bitstring.

rocdl.cvt.pk.fp8.f32 - Convert two f32's to fp8

Return op name rocdl.cvt.pkrtz as a bitstring.

rocdl.cvt.pkrtz - Convert two f32 input into a vector<2xf16>

Return op name rocdl.cvt.scale.pk8.bf16.bf8 as a bitstring.

rocdl.cvt.scale.pk8.bf16.bf8 - Scales 8 bf8 and converts them to 8 bf16.

Return op name rocdl.cvt.scale.pk8.bf16.fp4 as a bitstring.

rocdl.cvt.scale.pk8.bf16.fp4 - Scales 8 fp4 and converts them to 8 bf16.

Return op name rocdl.cvt.scale.pk8.bf16.fp8 as a bitstring.

rocdl.cvt.scale.pk8.bf16.fp8 - Scales 8 fp8 and converts them to 8 bf16.

Return op name rocdl.cvt.scale.pk8.f16.bf8 as a bitstring.

rocdl.cvt.scale.pk8.f16.bf8 - Scales 8 bf8 and converts them to 8 f16.

Return op name rocdl.cvt.scale.pk8.f16.fp4 as a bitstring.

rocdl.cvt.scale.pk8.f16.fp4 - Scales 8 fp4 and converts them to 8 f16.

Return op name rocdl.cvt.scale.pk8.f16.fp8 as a bitstring.

rocdl.cvt.scale.pk8.f16.fp8 - Scales 8 fp8 and converts them to 8 f16.

Return op name rocdl.cvt.scale.pk8.f32.bf8 as a bitstring.

rocdl.cvt.scale.pk8.f32.bf8 - Scales 8 bf8 and converts them to 8 f32.

Return op name rocdl.cvt.scale.pk8.f32.fp4 as a bitstring.

rocdl.cvt.scale.pk8.f32.fp4 - Scales 8 fp4 and converts them to 8 f32.

Return op name rocdl.cvt.scale.pk8.f32.fp8 as a bitstring.

rocdl.cvt.scale.pk8.f32.fp8 - Scales 8 fp8 and converts them to 8 f32.

Return op name rocdl.cvt.scale.pk16.bf16.bf6 as a bitstring.

rocdl.cvt.scale.pk16.bf16.bf6 - Scales 16 bf6 and converts them to 16 bf16.

Return op name rocdl.cvt.scale.pk16.bf16.fp6 as a bitstring.

rocdl.cvt.scale.pk16.bf16.fp6 - Scales 16 fp6 and converts them to 16 bf16.

Return op name rocdl.cvt.scale.pk16.f16.bf6 as a bitstring.

rocdl.cvt.scale.pk16.f16.bf6 - Scales 16 bf6 and converts them to 16 f16.

Return op name rocdl.cvt.scale.pk16.f16.fp6 as a bitstring.

rocdl.cvt.scale.pk16.f16.fp6 - Scales 16 fp6 and converts them to 16 f16.

Return op name rocdl.cvt.scale.pk16.f32.bf6 as a bitstring.

rocdl.cvt.scale.pk16.f32.bf6 - Scales 16 bf6 and converts them to 16 f32.

Return op name rocdl.cvt.scale.pk16.f32.fp6 as a bitstring.

rocdl.cvt.scale.pk16.f32.fp6 - Scales 16 fp6 and converts them to 16 f32.

Return op name rocdl.cvt.scalef32.2xpk16.bf6.f32 as a bitstring.

rocdl.cvt.scalef32.2xpk16.bf6.f32 - Scale and convert two vector<16xf32> to 32 packed bf6

Return op name rocdl.cvt.scalef32.2xpk16.fp6.f32 as a bitstring.

rocdl.cvt.scalef32.2xpk16.fp6.f32 - Scale and convert two vector<16xf32> to 32 packed fp6

Return op name rocdl.cvt.scalef32.f16.bf8 as a bitstring.

rocdl.cvt.scalef32.f16.bf8 - Scaled convert bf8 from packed vector to f16, updating tied result

Return op name rocdl.cvt.scalef32.f16.fp8 as a bitstring.

rocdl.cvt.scalef32.f16.fp8 - Scaled convert fp8 from packed vector to f16, updating tied result

Return op name rocdl.cvt.scalef32.f32.bf8 as a bitstring.

rocdl.cvt.scalef32.f32.bf8 - Scaled convert bf8 from packed vector to f32

Return op name rocdl.cvt.scalef32.f32.fp8 as a bitstring.

rocdl.cvt.scalef32.f32.fp8 - Scaled convert fp8 from packed vector to f32

Return op name rocdl.cvt.scalef32.pk8.bf8.bf16 as a bitstring.

rocdl.cvt.scalef32.pk8.bf8.bf16 - Scale and convert packed bf16 to packed bf8

Return op name rocdl.cvt.scalef32.pk8.bf8.f16 as a bitstring.

rocdl.cvt.scalef32.pk8.bf8.f16 - Scale and convert packed f16 to packed bf8

Return op name rocdl.cvt.scalef32.pk8.bf8.f32 as a bitstring.

rocdl.cvt.scalef32.pk8.bf8.f32 - Scale and convert packed f32 to packed bf8

Return op name rocdl.cvt.scalef32.pk8.fp4.bf16 as a bitstring.

rocdl.cvt.scalef32.pk8.fp4.bf16 - Scale and convert packed bf16 to packed fp4

Return op name rocdl.cvt.scalef32.pk8.fp4.f16 as a bitstring.

rocdl.cvt.scalef32.pk8.fp4.f16 - Scale and convert packed f16 to packed fp4

Return op name rocdl.cvt.scalef32.pk8.fp4.f32 as a bitstring.

rocdl.cvt.scalef32.pk8.fp4.f32 - Scale and convert packed f32 to packed fp4

Return op name rocdl.cvt.scalef32.pk8.fp8.bf16 as a bitstring.

rocdl.cvt.scalef32.pk8.fp8.bf16 - Scale and convert packed bf16 to packed fp8

Return op name rocdl.cvt.scalef32.pk8.fp8.f16 as a bitstring.

rocdl.cvt.scalef32.pk8.fp8.f16 - Scale and convert packed f16 to packed fp8

Return op name rocdl.cvt.scalef32.pk8.fp8.f32 as a bitstring.

rocdl.cvt.scalef32.pk8.fp8.f32 - Scale and convert packed f32 to packed fp8

Return op name rocdl.cvt.scalef32.pk16.bf6.bf16 as a bitstring.

rocdl.cvt.scalef32.pk16.bf6.bf16 - Scale and convert packed bf16 to packed bf6

Return op name rocdl.cvt.scalef32.pk16.bf6.f16 as a bitstring.

rocdl.cvt.scalef32.pk16.bf6.f16 - Scale and convert packed f16 to packed bf6

Return op name rocdl.cvt.scalef32.pk16.bf6.f32 as a bitstring.

rocdl.cvt.scalef32.pk16.bf6.f32 - Scale and convert packed f32 to packed bf6

Return op name rocdl.cvt.scalef32.pk16.fp6.bf16 as a bitstring.

rocdl.cvt.scalef32.pk16.fp6.bf16 - Scale and convert packed bf16 to packed fp6

Return op name rocdl.cvt.scalef32.pk16.fp6.f16 as a bitstring.

rocdl.cvt.scalef32.pk16.fp6.f16 - Scale and convert packed f16 to packed fp6

Return op name rocdl.cvt.scalef32.pk16.fp6.f32 as a bitstring.

rocdl.cvt.scalef32.pk16.fp6.f32 - Scale and convert packed f32 to packed fp6

Return op name rocdl.cvt.scalef32.pk32.bf6.bf16 as a bitstring.

rocdl.cvt.scalef32.pk32.bf6.bf16 - Scale and convert packed bf16 to packed bf6

Return op name rocdl.cvt.scalef32.pk32.bf6.f16 as a bitstring.

rocdl.cvt.scalef32.pk32.bf6.f16 - Scale and convert packed f16 to packed bf6

Return op name rocdl.cvt.scalef32.pk32.bf16.bf6 as a bitstring.

rocdl.cvt.scalef32.pk32.bf16.bf6 - Scale and convert packed bf6 to packed bf16

Return op name rocdl.cvt.scalef32.pk32.bf16.fp6 as a bitstring.

rocdl.cvt.scalef32.pk32.bf16.fp6 - Scale and convert packed fp6 to packed bf16

Return op name rocdl.cvt.scalef32.pk32.f16.bf6 as a bitstring.

rocdl.cvt.scalef32.pk32.f16.bf6 - Scale and convert packed bf6 to packed f16

Return op name rocdl.cvt.scalef32.pk32.f16.fp6 as a bitstring.

rocdl.cvt.scalef32.pk32.f16.fp6 - Scale and convert packed fp6 to packed f16

Return op name rocdl.cvt.scalef32.pk32.f32.bf6 as a bitstring.

rocdl.cvt.scalef32.pk32.f32.bf6 - Scale and convert packed bf6 to packed f32

Return op name rocdl.cvt.scalef32.pk32.f32.fp6 as a bitstring.

rocdl.cvt.scalef32.pk32.f32.fp6 - Scale and convert packed fp6 to packed f32

Return op name rocdl.cvt.scalef32.pk32.fp6.bf16 as a bitstring.

rocdl.cvt.scalef32.pk32.fp6.bf16 - Scale and convert packed bf16 to packed fp6

Return op name rocdl.cvt.scalef32.pk32.fp6.f16 as a bitstring.

rocdl.cvt.scalef32.pk32.fp6.f16 - Scale and convert packed f16 to packed fp6

Return op name rocdl.cvt.scalef32.pk.bf8.bf16 as a bitstring.

rocdl.cvt.scalef32.pk.bf8.bf16 - Scaled convert two bf16to two bf8, updating packed vector

Return op name rocdl.cvt.scalef32.pk.bf8.f16 as a bitstring.

rocdl.cvt.scalef32.pk.bf8.f16 - Scaled convert two f16to two bf8, updating packed vector

Return op name rocdl.cvt.scalef32.pk.bf8.f32 as a bitstring.

rocdl.cvt.scalef32.pk.bf8.f32 - Scaled convert two f32 to two bf8, updating packed vector

Return op name rocdl.cvt.scalef32.pk.bf16.bf8 as a bitstring.

rocdl.cvt.scalef32.pk.bf16.bf8 - Scaled convert two bf8to two bf16

Return op name rocdl.cvt.scalef32.pk.bf16.fp4 as a bitstring.

rocdl.cvt.scalef32.pk.bf16.fp4 - Scale and convert two packed fp4 to packed bf16

Return op name rocdl.cvt.scalef32.pk.bf16.fp8 as a bitstring.

rocdl.cvt.scalef32.pk.bf16.fp8 - Scaled convert two fp8to two bf16

Return op name rocdl.cvt.scalef32.pk.f16.bf8 as a bitstring.

rocdl.cvt.scalef32.pk.f16.bf8 - Scaled convert two bf8to two f16

Return op name rocdl.cvt.scalef32.pk.f16.fp4 as a bitstring.

rocdl.cvt.scalef32.pk.f16.fp4 - Scale and convert two packed fp4 to packed f16

Return op name rocdl.cvt.scalef32.pk.f16.fp8 as a bitstring.

rocdl.cvt.scalef32.pk.f16.fp8 - Scaled convert two fp8to two f16

Return op name rocdl.cvt.scalef32.pk.f32.bf8 as a bitstring.

rocdl.cvt.scalef32.pk.f32.bf8 - Scaled convert two bf8to two f32

Return op name rocdl.cvt.scalef32.pk.f32.fp4 as a bitstring.

rocdl.cvt.scalef32.pk.f32.fp4 - Scale and convert two packed fp4 to packed f32

Return op name rocdl.cvt.scalef32.pk.f32.fp8 as a bitstring.

rocdl.cvt.scalef32.pk.f32.fp8 - Scaled convert two fp8to two f32

Return op name rocdl.cvt.scalef32.pk.fp4.bf16 as a bitstring.

rocdl.cvt.scalef32.pk.fp4.bf16 - Scale and convert two bf16 to packed fp4, updating tied vector

Return op name rocdl.cvt.scalef32.pk.fp4.f16 as a bitstring.

rocdl.cvt.scalef32.pk.fp4.f16 - Scale and convert two f16 to packed fp4, updating tied vector

Return op name rocdl.cvt.scalef32.pk.fp4.f32 as a bitstring.

rocdl.cvt.scalef32.pk.fp4.f32 - Scale and convert two f32 values to two packed fp4, updating tied vector

Return op name rocdl.cvt.scalef32.pk.fp8.bf16 as a bitstring.

rocdl.cvt.scalef32.pk.fp8.bf16 - Scaled convert two bf16to two fp8, updating packed vector

Return op name rocdl.cvt.scalef32.pk.fp8.f16 as a bitstring.

rocdl.cvt.scalef32.pk.fp8.f16 - Scaled convert two f16to two fp8, updating packed vector

Return op name rocdl.cvt.scalef32.pk.fp8.f32 as a bitstring.

rocdl.cvt.scalef32.pk.fp8.f32 - Scaled convert two f32 to two fp8, updating packed vector

Return op name rocdl.cvt.scalef32.sr.bf8.bf16 as a bitstring.

rocdl.cvt.scalef32.sr.bf8.bf16 - Scaled convert bf16to bf8 with stochiastic rounding, updating packed vector

Return op name rocdl.cvt.scalef32.sr.bf8.f16 as a bitstring.

rocdl.cvt.scalef32.sr.bf8.f16 - Scaled convert f16to bf8 with stochiastic rounding, updating packed vector

Return op name rocdl.cvt.scalef32.sr.bf8.f32 as a bitstring.

rocdl.cvt.scalef32.sr.bf8.f32 - Scaled convert f32to bf8 with stochiastic rounding, updating packed vector

Return op name rocdl.cvt.scalef32.sr.fp8.bf16 as a bitstring.

rocdl.cvt.scalef32.sr.fp8.bf16 - Scaled convert bf16to fp8 with stochiastic rounding, updating packed vector

Return op name rocdl.cvt.scalef32.sr.fp8.f16 as a bitstring.

rocdl.cvt.scalef32.sr.fp8.f16 - Scaled convert f16to fp8 with stochiastic rounding, updating packed vector

Return op name rocdl.cvt.scalef32.sr.fp8.f32 as a bitstring.

rocdl.cvt.scalef32.sr.fp8.f32 - Scaled convert f32to fp8 with stochiastic rounding, updating packed vector

Return op name rocdl.cvt.scalef32.sr.pk8.bf8.bf16 as a bitstring.

rocdl.cvt.scalef32.sr.pk8.bf8.bf16 - Scale and convert packed bf16 to packed bf8 with stochastic rounding

Return op name rocdl.cvt.scalef32.sr.pk8.bf8.f16 as a bitstring.

rocdl.cvt.scalef32.sr.pk8.bf8.f16 - Scale and convert packed f16 to packed bf8 with stochastic rounding

Return op name rocdl.cvt.scalef32.sr.pk8.bf8.f32 as a bitstring.

rocdl.cvt.scalef32.sr.pk8.bf8.f32 - Scale and convert packed f32 to packed bf8 with stochastic rounding

Return op name rocdl.cvt.scalef32.sr.pk8.fp4.bf16 as a bitstring.

rocdl.cvt.scalef32.sr.pk8.fp4.bf16 - Scale and convert packed bf16 to packed fp4 with stochastic rounding

Return op name rocdl.cvt.scalef32.sr.pk8.fp4.f16 as a bitstring.

rocdl.cvt.scalef32.sr.pk8.fp4.f16 - Scale and convert packed f16 to packed fp4 with stochastic rounding

Return op name rocdl.cvt.scalef32.sr.pk8.fp4.f32 as a bitstring.

rocdl.cvt.scalef32.sr.pk8.fp4.f32 - Scale and convert packed f32 to packed fp4 with stochastic rounding

Return op name rocdl.cvt.scalef32.sr.pk8.fp8.bf16 as a bitstring.

rocdl.cvt.scalef32.sr.pk8.fp8.bf16 - Scale and convert packed bf16 to packed fp8 with stochastic rounding

Return op name rocdl.cvt.scalef32.sr.pk8.fp8.f16 as a bitstring.

rocdl.cvt.scalef32.sr.pk8.fp8.f16 - Scale and convert packed f16 to packed fp8 with stochastic rounding

Return op name rocdl.cvt.scalef32.sr.pk8.fp8.f32 as a bitstring.

rocdl.cvt.scalef32.sr.pk8.fp8.f32 - Scale and convert packed f32 to packed fp8 with stochastic rounding

Return op name rocdl.cvt.scalef32.sr.pk16.bf6.bf16 as a bitstring.

rocdl.cvt.scalef32.sr.pk16.bf6.bf16 - Scale and convert packed bf16 to packed bf6 with stochastic rounding

Return op name rocdl.cvt.scalef32.sr.pk16.bf6.f16 as a bitstring.

rocdl.cvt.scalef32.sr.pk16.bf6.f16 - Scale and convert packed f16 to packed bf6 with stochastic rounding

Return op name rocdl.cvt.scalef32.sr.pk16.bf6.f32 as a bitstring.

rocdl.cvt.scalef32.sr.pk16.bf6.f32 - Scale and convert packed f32 to packed bf6 with stochastic rounding

Return op name rocdl.cvt.scalef32.sr.pk16.fp6.bf16 as a bitstring.

rocdl.cvt.scalef32.sr.pk16.fp6.bf16 - Scale and convert packed bf16 to packed fp6 with stochastic rounding

Return op name rocdl.cvt.scalef32.sr.pk16.fp6.f16 as a bitstring.

rocdl.cvt.scalef32.sr.pk16.fp6.f16 - Scale and convert packed f16 to packed fp6 with stochastic rounding

Return op name rocdl.cvt.scalef32.sr.pk16.fp6.f32 as a bitstring.

rocdl.cvt.scalef32.sr.pk16.fp6.f32 - Scale and convert packed f32 to packed fp6 with stochastic rounding

Return op name rocdl.cvt.scalef32.sr.pk32.bf6.bf16 as a bitstring.

rocdl.cvt.scalef32.sr.pk32.bf6.bf16 - Scale and convert packed bf16 to packed bf6 with stochiastic rounding

Return op name rocdl.cvt.scalef32.sr.pk32.bf6.f16 as a bitstring.

rocdl.cvt.scalef32.sr.pk32.bf6.f16 - Scale and convert packed f16 to packed bf6 with stochiastic rounding

Return op name rocdl.cvt.scalef32.sr.pk32.bf6.f32 as a bitstring.

rocdl.cvt.scalef32.sr.pk32.bf6.f32 - Scale and convert packed f32 to packed bf6 with stochiastic rounding

Return op name rocdl.cvt.scalef32.sr.pk32.fp6.bf16 as a bitstring.

rocdl.cvt.scalef32.sr.pk32.fp6.bf16 - Scale and convert packed bf16 to packed fp6 with stochiastic rounding

Return op name rocdl.cvt.scalef32.sr.pk32.fp6.f16 as a bitstring.

rocdl.cvt.scalef32.sr.pk32.fp6.f16 - Scale and convert packed f16 to packed fp6 with stochiastic rounding

Return op name rocdl.cvt.scalef32.sr.pk32.fp6.f32 as a bitstring.

rocdl.cvt.scalef32.sr.pk32.fp6.f32 - Scale and convert packed f32 to packed fp6 with stochiastic rounding

Return op name rocdl.cvt.scalef32.sr.pk.fp4.bf16 as a bitstring.

rocdl.cvt.scalef32.sr.pk.fp4.bf16 - Scale and convert two bf16 to packed fp4 with stochiastic rounding, updating tied vector

Return op name rocdl.cvt.scalef32.sr.pk.fp4.f16 as a bitstring.

rocdl.cvt.scalef32.sr.pk.fp4.f16 - Scale and convert two f16 to packed fp4 with stochiastic rounding, updating tied vector

Return op name rocdl.cvt.scalef32.sr.pk.fp4.f32 as a bitstring.

rocdl.cvt.scalef32.sr.pk.fp4.f32 - Scale and convert two f32 to packed fp4 with stochiastic rounding, updating tied vector

Return op name rocdl.cvt.sr.bf8.f32 as a bitstring.

rocdl.cvt.sr.bf8.f32 - Convert f32 to bf8, stochiastic rounding

Return op name rocdl.cvt.sr.fp8.f32 as a bitstring.

rocdl.cvt.sr.fp8.f32 - Convert f32 to fp8, stochiastic rounding

Return op name rocdl.dot4.f32.bf8.bf8 as a bitstring.

rocdl.dot4.f32.bf8.bf8

Return op name rocdl.dot4.f32.bf8.fp8 as a bitstring.

rocdl.dot4.f32.bf8.fp8

Return op name rocdl.dot4.f32.fp8.bf8 as a bitstring.

rocdl.dot4.f32.fp8.bf8

Return op name rocdl.dot4.f32.fp8.fp8 as a bitstring.

rocdl.dot4.f32.fp8.fp8

Return op name rocdl.ds.atomic.async.barrier.arrive.b64 as a bitstring.

rocdl.ds.atomic.async.barrier.arrive.b64

Return op name rocdl.ds.atomic.barrier.arrive.rtn.b64 as a bitstring.

rocdl.ds.atomic.barrier.arrive.rtn.b64

Return op name rocdl.ds_bpermute as a bitstring.

rocdl.ds_bpermute

Return op name rocdl.ds.load.tr4.b64 as a bitstring.

rocdl.ds.load.tr4.b64 - Loads and transposes a matrix from ds memory to registers (available in gfx1250+).

Return op name rocdl.ds.load.tr6.b96 as a bitstring.

rocdl.ds.load.tr6.b96 - Loads and transposes a matrix from ds memory to registers (available in gfx1250+).

Return op name rocdl.ds.load.tr8.b64 as a bitstring.

rocdl.ds.load.tr8.b64 - Loads and transposes a matrix from ds memory to registers (available in gfx1250+).

Return op name rocdl.ds.load.tr16.b128 as a bitstring.

rocdl.ds.load.tr16.b128 - Loads and transposes a matrix from ds memory to registers (available in gfx1250+).

Return op name rocdl.ds.read.tr4.b64 as a bitstring.

rocdl.ds.read.tr4.b64

Return op name rocdl.ds.read.tr6.b96 as a bitstring.

rocdl.ds.read.tr6.b96

Return op name rocdl.ds.read.tr8.b64 as a bitstring.

rocdl.ds.read.tr8.b64

Return op name rocdl.ds.read.tr16.b64 as a bitstring.

rocdl.ds.read.tr16.b64

Return op name rocdl.ds_swizzle as a bitstring.

rocdl.ds_swizzle

Return op name rocdl.exp2 as a bitstring.

rocdl.exp2

Return op name rocdl.exp as a bitstring.

rocdl.exp

Return op name rocdl.fdot2 as a bitstring.

rocdl.fdot2

Return op name rocdl.fdot2.bf16.bf16 as a bitstring.

rocdl.fdot2.bf16.bf16

Return op name rocdl.fdot2.f16.f16 as a bitstring.

rocdl.fdot2.f16.f16

Return op name rocdl.fdot2.f32.bf16 as a bitstring.

rocdl.fdot2.f32.bf16

Return op name rocdl.flat.prefetch as a bitstring.

rocdl.flat.prefetch

Return op name rocdl.fmed3 as a bitstring.

rocdl.fmed3 - Median of three float/half values

Return op name rocdl.global.load.async.lds as a bitstring.

rocdl.global.load.async.lds - Version of rocdl.load.async.to.lds specialized to global pointers

Return op name rocdl.global.load.async.to.lds.b8 as a bitstring.

rocdl.global.load.async.to.lds.b8

Return op name rocdl.global.load.async.to.lds.b32 as a bitstring.

rocdl.global.load.async.to.lds.b32

Return op name rocdl.global.load.async.to.lds.b64 as a bitstring.

rocdl.global.load.async.to.lds.b64

Return op name rocdl.global.load.async.to.lds.b128 as a bitstring.

rocdl.global.load.async.to.lds.b128

Return op name rocdl.global.load.lds as a bitstring.

rocdl.global.load.lds

Return op name rocdl.global.load.tr4.b64 as a bitstring.

rocdl.global.load.tr4.b64 - Loads and transposes a matrix from global memory to registers (available in gfx1250+).

Return op name rocdl.global.load.tr6.b96 as a bitstring.

rocdl.global.load.tr6.b96 - Loads and transposes a matrix from global memory to registers (available in gfx1250+).

Return op name rocdl.global.load.tr.b64 as a bitstring.

rocdl.global.load.tr.b64 - Loads and transposes a matrix from global memory to registers (available in gfx1250+).

Return op name rocdl.global.load.tr.b128 as a bitstring.

rocdl.global.load.tr.b128 - Loads and transposes a matrix from global memory to registers (available in gfx1250+).

Return op name rocdl.global.prefetch as a bitstring.

rocdl.global.prefetch

Return op name rocdl.global.store.async.from.lds.b8 as a bitstring.

rocdl.global.store.async.from.lds.b8

Return op name rocdl.global.store.async.from.lds.b32 as a bitstring.

rocdl.global.store.async.from.lds.b32

Return op name rocdl.global.store.async.from.lds.b64 as a bitstring.

rocdl.global.store.async.from.lds.b64

Return op name rocdl.global.store.async.from.lds.b128 as a bitstring.

rocdl.global.store.async.from.lds.b128

Return op name rocdl.iglp.opt as a bitstring.

rocdl.iglp.opt

Return op name rocdl.load.async.to.lds as a bitstring.

rocdl.load.async.to.lds - Gathering load to LDS that requires explicit async memory tracking

Return op name rocdl.load.to.lds as a bitstring.

rocdl.load.to.lds

Return op name rocdl.log as a bitstring.

rocdl.log

Return op name rocdl.make.buffer.rsrc as a bitstring.

rocdl.make.buffer.rsrc

Return op name rocdl.mbcnt.hi as a bitstring.

rocdl.mbcnt.hi

Return op name rocdl.mbcnt.lo as a bitstring.

rocdl.mbcnt.lo

Return op name rocdl.mfma.f32.4x4x1f32 as a bitstring.

rocdl.mfma.f32.4x4x1f32

Return op name rocdl.mfma.f32.4x4x2bf16 as a bitstring.

rocdl.mfma.f32.4x4x2bf16

Return op name rocdl.mfma.f32.4x4x4bf16.1k as a bitstring.

rocdl.mfma.f32.4x4x4bf16.1k

Return op name rocdl.mfma.f32.4x4x4f16 as a bitstring.

rocdl.mfma.f32.4x4x4f16

Return op name rocdl.mfma.f32.16x16x1f32 as a bitstring.

rocdl.mfma.f32.16x16x1f32

Return op name rocdl.mfma.f32.16x16x2bf16 as a bitstring.

rocdl.mfma.f32.16x16x2bf16

Return op name rocdl.mfma.f32.16x16x4bf16.1k as a bitstring.

rocdl.mfma.f32.16x16x4bf16.1k

Return op name rocdl.mfma.f32.16x16x4f16 as a bitstring.

rocdl.mfma.f32.16x16x4f16

Return op name rocdl.mfma.f32.16x16x4f32 as a bitstring.

rocdl.mfma.f32.16x16x4f32

Return op name rocdl.mfma.f32.16x16x8.xf32 as a bitstring.

rocdl.mfma.f32.16x16x8.xf32

Return op name rocdl.mfma.f32.16x16x8bf16 as a bitstring.

rocdl.mfma.f32.16x16x8bf16

Return op name rocdl.mfma.f32.16x16x16bf16.1k as a bitstring.

rocdl.mfma.f32.16x16x16bf16.1k

Return op name rocdl.mfma.f32.16x16x16f16 as a bitstring.

rocdl.mfma.f32.16x16x16f16

Return op name rocdl.mfma.f32.16x16x32.bf8.bf8 as a bitstring.

rocdl.mfma.f32.16x16x32.bf8.bf8

Return op name rocdl.mfma.f32.16x16x32.bf8.fp8 as a bitstring.

rocdl.mfma.f32.16x16x32.bf8.fp8

Return op name rocdl.mfma.f32.16x16x32.bf16 as a bitstring.

rocdl.mfma.f32.16x16x32.bf16

Return op name rocdl.mfma.f32.16x16x32.f16 as a bitstring.

rocdl.mfma.f32.16x16x32.f16

Return op name rocdl.mfma.f32.16x16x32.fp8.bf8 as a bitstring.

rocdl.mfma.f32.16x16x32.fp8.bf8

Return op name rocdl.mfma.f32.16x16x32.fp8.fp8 as a bitstring.

rocdl.mfma.f32.16x16x32.fp8.fp8

Return op name rocdl.mfma.f32.32x32x1f32 as a bitstring.

rocdl.mfma.f32.32x32x1f32

Return op name rocdl.mfma.f32.32x32x2bf16 as a bitstring.

rocdl.mfma.f32.32x32x2bf16

Return op name rocdl.mfma.f32.32x32x2f32 as a bitstring.

rocdl.mfma.f32.32x32x2f32

Return op name rocdl.mfma.f32.32x32x4.xf32 as a bitstring.

rocdl.mfma.f32.32x32x4.xf32

Return op name rocdl.mfma.f32.32x32x4bf16 as a bitstring.

rocdl.mfma.f32.32x32x4bf16

Return op name rocdl.mfma.f32.32x32x4bf16.1k as a bitstring.

rocdl.mfma.f32.32x32x4bf16.1k

Return op name rocdl.mfma.f32.32x32x4f16 as a bitstring.

rocdl.mfma.f32.32x32x4f16

Return op name rocdl.mfma.f32.32x32x8bf16.1k as a bitstring.

rocdl.mfma.f32.32x32x8bf16.1k

Return op name rocdl.mfma.f32.32x32x8f16 as a bitstring.

rocdl.mfma.f32.32x32x8f16

Return op name rocdl.mfma.f32.32x32x16.bf8.bf8 as a bitstring.

rocdl.mfma.f32.32x32x16.bf8.bf8

Return op name rocdl.mfma.f32.32x32x16.bf8.fp8 as a bitstring.

rocdl.mfma.f32.32x32x16.bf8.fp8

Return op name rocdl.mfma.f32.32x32x16.bf16 as a bitstring.

rocdl.mfma.f32.32x32x16.bf16

Return op name rocdl.mfma.f32.32x32x16.f16 as a bitstring.

rocdl.mfma.f32.32x32x16.f16

Return op name rocdl.mfma.f32.32x32x16.fp8.bf8 as a bitstring.

rocdl.mfma.f32.32x32x16.fp8.bf8

Return op name rocdl.mfma.f32.32x32x16.fp8.fp8 as a bitstring.

rocdl.mfma.f32.32x32x16.fp8.fp8

Return op name rocdl.mfma.f64.4x4x4f64 as a bitstring.

rocdl.mfma.f64.4x4x4f64

Return op name rocdl.mfma.f64.16x16x4f64 as a bitstring.

rocdl.mfma.f64.16x16x4f64

Return op name rocdl.mfma.i32.4x4x4i8 as a bitstring.

rocdl.mfma.i32.4x4x4i8

Return op name rocdl.mfma.i32.16x16x4i8 as a bitstring.

rocdl.mfma.i32.16x16x4i8

Return op name rocdl.mfma.i32.16x16x16i8 as a bitstring.

rocdl.mfma.i32.16x16x16i8

Return op name rocdl.mfma.i32.16x16x32.i8 as a bitstring.

rocdl.mfma.i32.16x16x32.i8

Return op name rocdl.mfma.i32.16x16x64.i8 as a bitstring.

rocdl.mfma.i32.16x16x64.i8

Return op name rocdl.mfma.i32.32x32x4i8 as a bitstring.

rocdl.mfma.i32.32x32x4i8

Return op name rocdl.mfma.i32.32x32x8i8 as a bitstring.

rocdl.mfma.i32.32x32x8i8

Return op name rocdl.mfma.i32.32x32x16.i8 as a bitstring.

rocdl.mfma.i32.32x32x16.i8

Return op name rocdl.mfma.i32.32x32x32.i8 as a bitstring.

rocdl.mfma.i32.32x32x32.i8

Return op name rocdl.mfma.scale.f32.16x16x128.f8f6f4 as a bitstring.

rocdl.mfma.scale.f32.16x16x128.f8f6f4

Return op name rocdl.mfma.scale.f32.32x32x64.f8f6f4 as a bitstring.

rocdl.mfma.scale.f32.32x32x64.f8f6f4

Return op name rocdl.permlane16.swap as a bitstring.

rocdl.permlane16.swap

Return op name rocdl.permlane16.var as a bitstring.

rocdl.permlane16.var

Return op name rocdl.permlane32.swap as a bitstring.

rocdl.permlane32.swap

Return op name rocdl.permlanex16 as a bitstring.

rocdl.permlanex16

Return op name rocdl.permlanex16.var as a bitstring.

rocdl.permlanex16.var

Return op name rocdl.ptr.s.buffer.load as a bitstring.

rocdl.ptr.s.buffer.load

Return op name rocdl.raw.buffer.atomic.cmpswap as a bitstring.

rocdl.raw.buffer.atomic.cmpswap

Return op name rocdl.raw.buffer.atomic.fadd as a bitstring.

rocdl.raw.buffer.atomic.fadd

Return op name rocdl.raw.buffer.atomic.fmax as a bitstring.

rocdl.raw.buffer.atomic.fmax

Return op name rocdl.raw.buffer.atomic.smax as a bitstring.

rocdl.raw.buffer.atomic.smax

Return op name rocdl.raw.buffer.atomic.umin as a bitstring.

rocdl.raw.buffer.atomic.umin

Return op name rocdl.raw.buffer.load as a bitstring.

rocdl.raw.buffer.load

Return op name rocdl.raw.buffer.store as a bitstring.

rocdl.raw.buffer.store

Return op name rocdl.raw.ptr.buffer.atomic.cmpswap as a bitstring.

rocdl.raw.ptr.buffer.atomic.cmpswap

Return op name rocdl.raw.ptr.buffer.atomic.fadd as a bitstring.

rocdl.raw.ptr.buffer.atomic.fadd

Return op name rocdl.raw.ptr.buffer.atomic.fmax as a bitstring.

rocdl.raw.ptr.buffer.atomic.fmax

Return op name rocdl.raw.ptr.buffer.atomic.smax as a bitstring.

rocdl.raw.ptr.buffer.atomic.smax

Return op name rocdl.raw.ptr.buffer.atomic.umin as a bitstring.

rocdl.raw.ptr.buffer.atomic.umin

Return op name rocdl.raw.ptr.buffer.load as a bitstring.

rocdl.raw.ptr.buffer.load

Return op name rocdl.raw.ptr.buffer.load.async.lds as a bitstring.

rocdl.raw.ptr.buffer.load.async.lds - Async variant of raw.ptr.buffer.load.lds

Return op name rocdl.raw.ptr.buffer.load.lds as a bitstring.

rocdl.raw.ptr.buffer.load.lds

Return op name rocdl.raw.ptr.buffer.store as a bitstring.

rocdl.raw.ptr.buffer.store

Return op name rocdl.rcp as a bitstring.

rocdl.rcp

Return op name rocdl.readfirstlane as a bitstring.

rocdl.readfirstlane - Get the value in first active lane.

Return op name rocdl.readlane as a bitstring.

rocdl.readlane - Get the value in the specific lane.

Return op name rocdl.rsq as a bitstring.

rocdl.rsq

Return op name rocdl.s.barrier as a bitstring.

rocdl.s.barrier

Return op name rocdl.s.barrier.init as a bitstring.

rocdl.s.barrier.init

Return op name rocdl.s.barrier.join as a bitstring.

rocdl.s.barrier.join

Return op name rocdl.s.barrier.leave as a bitstring.

rocdl.s.barrier.leave

Return op name rocdl.s.barrier.signal as a bitstring.

rocdl.s.barrier.signal

Return op name rocdl.s.barrier.signal.isfirst as a bitstring.

rocdl.s.barrier.signal.isfirst

Return op name rocdl.s.barrier.signal.var as a bitstring.

rocdl.s.barrier.signal.var

Return op name rocdl.s.barrier.wait as a bitstring.

rocdl.s.barrier.wait

Return op name rocdl.s.get.barrier.state as a bitstring.

rocdl.s.get.barrier.state

Return op name rocdl.s.get.named.barrier.state as a bitstring.

rocdl.s.get.named.barrier.state

Return op name rocdl.s.nop as a bitstring.

rocdl.s.nop

Return op name rocdl.s.setprio as a bitstring.

rocdl.s.setprio

Return op name rocdl.s.sleep as a bitstring.

rocdl.s.sleep

Return op name rocdl.s.wait.asynccnt as a bitstring.

rocdl.s.wait.asynccnt - Wait until ASYNCCNT is less than or equal to count

Return op name rocdl.s.wait.dscnt as a bitstring.

rocdl.s.wait.dscnt - Wait until DSCNT is less than or equal to count

Return op name rocdl.s.wait.expcnt as a bitstring.

rocdl.s.wait.expcnt - Wait until EXPCNT is less than or equal to count

Return op name rocdl.s.wait.loadcnt as a bitstring.

rocdl.s.wait.loadcnt - Wait until LOADCNT is less than or equal to count

Return op name rocdl.s.wait.storecnt as a bitstring.

rocdl.s.wait.storecnt - Wait until STORECNT is less than or equal to count

Return op name rocdl.s.wait.tensorcnt as a bitstring.

rocdl.s.wait.tensorcnt - Wait until TENSORCNT is less than or equal to count

Return op name rocdl.s.waitcnt as a bitstring.

rocdl.s.waitcnt

Return op name rocdl.s.wakeup.barrier as a bitstring.

rocdl.s.wakeup.barrier

Return op name rocdl.sched.barrier as a bitstring.

rocdl.sched.barrier

Return op name rocdl.sched.group.barrier as a bitstring.

rocdl.sched.group.barrier

Return op name rocdl.sdot2 as a bitstring.

rocdl.sdot2

Return op name rocdl.sdot4 as a bitstring.

rocdl.sdot4

Return op name rocdl.sdot8 as a bitstring.

rocdl.sdot8

Return op name rocdl.sin as a bitstring.

rocdl.sin

Return op name rocdl.smfmac.f32.16x16x32.bf16 as a bitstring.

rocdl.smfmac.f32.16x16x32.bf16

Return op name rocdl.smfmac.f32.16x16x32.f16 as a bitstring.

rocdl.smfmac.f32.16x16x32.f16

Return op name rocdl.smfmac.f32.16x16x64.bf8.bf8 as a bitstring.

rocdl.smfmac.f32.16x16x64.bf8.bf8

Return op name rocdl.smfmac.f32.16x16x64.bf8.fp8 as a bitstring.

rocdl.smfmac.f32.16x16x64.bf8.fp8

Return op name rocdl.smfmac.f32.16x16x64.bf16 as a bitstring.

rocdl.smfmac.f32.16x16x64.bf16

Return op name rocdl.smfmac.f32.16x16x64.f16 as a bitstring.

rocdl.smfmac.f32.16x16x64.f16

Return op name rocdl.smfmac.f32.16x16x64.fp8.bf8 as a bitstring.

rocdl.smfmac.f32.16x16x64.fp8.bf8

Return op name rocdl.smfmac.f32.16x16x64.fp8.fp8 as a bitstring.

rocdl.smfmac.f32.16x16x64.fp8.fp8

Return op name rocdl.smfmac.f32.16x16x128.bf8.bf8 as a bitstring.

rocdl.smfmac.f32.16x16x128.bf8.bf8

Return op name rocdl.smfmac.f32.16x16x128.bf8.fp8 as a bitstring.

rocdl.smfmac.f32.16x16x128.bf8.fp8

Return op name rocdl.smfmac.f32.16x16x128.fp8.bf8 as a bitstring.

rocdl.smfmac.f32.16x16x128.fp8.bf8

Return op name rocdl.smfmac.f32.16x16x128.fp8.fp8 as a bitstring.

rocdl.smfmac.f32.16x16x128.fp8.fp8

Return op name rocdl.smfmac.f32.32x32x16.bf16 as a bitstring.

rocdl.smfmac.f32.32x32x16.bf16

Return op name rocdl.smfmac.f32.32x32x16.f16 as a bitstring.

rocdl.smfmac.f32.32x32x16.f16

Return op name rocdl.smfmac.f32.32x32x32.bf8.bf8 as a bitstring.

rocdl.smfmac.f32.32x32x32.bf8.bf8

Return op name rocdl.smfmac.f32.32x32x32.bf8.fp8 as a bitstring.

rocdl.smfmac.f32.32x32x32.bf8.fp8

Return op name rocdl.smfmac.f32.32x32x32.bf16 as a bitstring.

rocdl.smfmac.f32.32x32x32.bf16

Return op name rocdl.smfmac.f32.32x32x32.f16 as a bitstring.

rocdl.smfmac.f32.32x32x32.f16

Return op name rocdl.smfmac.f32.32x32x32.fp8.bf8 as a bitstring.

rocdl.smfmac.f32.32x32x32.fp8.bf8

Return op name rocdl.smfmac.f32.32x32x32.fp8.fp8 as a bitstring.

rocdl.smfmac.f32.32x32x32.fp8.fp8

Return op name rocdl.smfmac.f32.32x32x64.bf8.bf8 as a bitstring.

rocdl.smfmac.f32.32x32x64.bf8.bf8

Return op name rocdl.smfmac.f32.32x32x64.bf8.fp8 as a bitstring.

rocdl.smfmac.f32.32x32x64.bf8.fp8

Return op name rocdl.smfmac.f32.32x32x64.fp8.bf8 as a bitstring.

rocdl.smfmac.f32.32x32x64.fp8.bf8

Return op name rocdl.smfmac.f32.32x32x64.fp8.fp8 as a bitstring.

rocdl.smfmac.f32.32x32x64.fp8.fp8

Return op name rocdl.smfmac.i32.16x16x64.i8 as a bitstring.

rocdl.smfmac.i32.16x16x64.i8

Return op name rocdl.smfmac.i32.16x16x128.i8 as a bitstring.

rocdl.smfmac.i32.16x16x128.i8

Return op name rocdl.smfmac.i32.32x32x32.i8 as a bitstring.

rocdl.smfmac.i32.32x32x32.i8

Return op name rocdl.smfmac.i32.32x32x64.i8 as a bitstring.

rocdl.smfmac.i32.32x32x64.i8

Return op name rocdl.sqrt as a bitstring.

rocdl.sqrt

Return op name rocdl.sudot4 as a bitstring.

rocdl.sudot4

Return op name rocdl.sudot8 as a bitstring.

rocdl.sudot8

Return op name rocdl.swmmac.bf16.16x16x32.bf16 as a bitstring.

rocdl.swmmac.bf16.16x16x32.bf16

Return op name rocdl.swmmac.bf16.16x16x64.bf16 as a bitstring.

rocdl.swmmac.bf16.16x16x64.bf16

Return op name rocdl.swmmac.bf16f32.16x16x64.bf16 as a bitstring.

rocdl.swmmac.bf16f32.16x16x64.bf16

Return op name rocdl.swmmac.f16.16x16x32.f16 as a bitstring.

rocdl.swmmac.f16.16x16x32.f16

Return op name rocdl.swmmac.f16.16x16x64.f16 as a bitstring.

rocdl.swmmac.f16.16x16x64.f16

Return op name rocdl.swmmac.f16.16x16x128.bf8.bf8 as a bitstring.

rocdl.swmmac.f16.16x16x128.bf8.bf8

Return op name rocdl.swmmac.f16.16x16x128.bf8.fp8 as a bitstring.

rocdl.swmmac.f16.16x16x128.bf8.fp8

Return op name rocdl.swmmac.f16.16x16x128.fp8.bf8 as a bitstring.

rocdl.swmmac.f16.16x16x128.fp8.bf8

Return op name rocdl.swmmac.f16.16x16x128.fp8.fp8 as a bitstring.

rocdl.swmmac.f16.16x16x128.fp8.fp8

Return op name rocdl.swmmac.f32.16x16x32.bf8.bf8 as a bitstring.

rocdl.swmmac.f32.16x16x32.bf8.bf8

Return op name rocdl.swmmac.f32.16x16x32.bf8.fp8 as a bitstring.

rocdl.swmmac.f32.16x16x32.bf8.fp8

Return op name rocdl.swmmac.f32.16x16x32.bf16 as a bitstring.

rocdl.swmmac.f32.16x16x32.bf16

Return op name rocdl.swmmac.f32.16x16x32.f16 as a bitstring.

rocdl.swmmac.f32.16x16x32.f16

Return op name rocdl.swmmac.f32.16x16x32.fp8.bf8 as a bitstring.

rocdl.swmmac.f32.16x16x32.fp8.bf8

Return op name rocdl.swmmac.f32.16x16x32.fp8.fp8 as a bitstring.

rocdl.swmmac.f32.16x16x32.fp8.fp8

Return op name rocdl.swmmac.f32.16x16x64.bf16 as a bitstring.

rocdl.swmmac.f32.16x16x64.bf16

Return op name rocdl.swmmac.f32.16x16x64.f16 as a bitstring.

rocdl.swmmac.f32.16x16x64.f16

Return op name rocdl.swmmac.f32.16x16x128.bf8.bf8 as a bitstring.

rocdl.swmmac.f32.16x16x128.bf8.bf8

Return op name rocdl.swmmac.f32.16x16x128.bf8.fp8 as a bitstring.

rocdl.swmmac.f32.16x16x128.bf8.fp8

Return op name rocdl.swmmac.f32.16x16x128.fp8.bf8 as a bitstring.

rocdl.swmmac.f32.16x16x128.fp8.bf8

Return op name rocdl.swmmac.f32.16x16x128.fp8.fp8 as a bitstring.

rocdl.swmmac.f32.16x16x128.fp8.fp8

Return op name rocdl.swmmac.i32.16x16x32.iu4 as a bitstring.

rocdl.swmmac.i32.16x16x32.iu4

Return op name rocdl.swmmac.i32.16x16x32.iu8 as a bitstring.

rocdl.swmmac.i32.16x16x32.iu8

Return op name rocdl.swmmac.i32.16x16x64.iu4 as a bitstring.

rocdl.swmmac.i32.16x16x64.iu4

Return op name rocdl.swmmac.i32.16x16x128.iu8 as a bitstring.

rocdl.swmmac.i32.16x16x128.iu8

Return op name rocdl.tanh as a bitstring.

rocdl.tanh

Return op name rocdl.tensor.load.to.lds as a bitstring.

rocdl.tensor.load.to.lds - Base class for ROCDL tensor load/store to/from LDS.

Return op name rocdl.tensor.store.from.lds as a bitstring.

rocdl.tensor.store.from.lds - Base class for ROCDL tensor load/store to/from LDS.

Return op name rocdl.udot2 as a bitstring.

rocdl.udot2

Return op name rocdl.udot4 as a bitstring.

rocdl.udot4

Return op name rocdl.udot8 as a bitstring.

rocdl.udot8

Return op name rocdl.update.dpp as a bitstring.

rocdl.update.dpp

Return op name rocdl.wait.asyncmark as a bitstring.

rocdl.wait.asyncmark - Wait until N or fewer async operation groups are unexecuted

Return op name rocdl.wave.barrier as a bitstring.

rocdl.wave.barrier

Return op name rocdl.wave.id as a bitstring.

rocdl.wave.id

Return op name rocdl.wavefrontsize as a bitstring.

rocdl.wavefrontsize

Return op name rocdl.wmma.bf16.16x16x16.bf16 as a bitstring.

rocdl.wmma.bf16.16x16x16.bf16

Return op name rocdl.wmma.bf16.16x16x32.bf16 as a bitstring.

rocdl.wmma.bf16.16x16x32.bf16

Return op name rocdl.wmma.bf16f32.16x16x32.bf16 as a bitstring.

rocdl.wmma.bf16f32.16x16x32.bf16

Return op name rocdl.wmma.f16.16x16x16.f16 as a bitstring.

rocdl.wmma.f16.16x16x16.f16

Return op name rocdl.wmma.f16.16x16x32.f16 as a bitstring.

rocdl.wmma.f16.16x16x32.f16

Return op name rocdl.wmma.f16.16x16x64.bf8_bf8 as a bitstring.

rocdl.wmma.f16.16x16x64.bf8_bf8

Return op name rocdl.wmma.f16.16x16x64.bf8_fp8 as a bitstring.

rocdl.wmma.f16.16x16x64.bf8_fp8

Return op name rocdl.wmma.f16.16x16x64.fp8_bf8 as a bitstring.

rocdl.wmma.f16.16x16x64.fp8_bf8

Return op name rocdl.wmma.f16.16x16x64.fp8_fp8 as a bitstring.

rocdl.wmma.f16.16x16x64.fp8_fp8

Return op name rocdl.wmma.f16.16x16x128.bf8_bf8 as a bitstring.

rocdl.wmma.f16.16x16x128.bf8_bf8

Return op name rocdl.wmma.f16.16x16x128.bf8_fp8 as a bitstring.

rocdl.wmma.f16.16x16x128.bf8_fp8

Return op name rocdl.wmma.f16.16x16x128.fp8_bf8 as a bitstring.

rocdl.wmma.f16.16x16x128.fp8_bf8

Return op name rocdl.wmma.f16.16x16x128.fp8_fp8 as a bitstring.

rocdl.wmma.f16.16x16x128.fp8_fp8

Return op name rocdl.wmma.f32.16x16x4.f32 as a bitstring.

rocdl.wmma.f32.16x16x4.f32

Return op name rocdl.wmma.f32.16x16x16.bf8_bf8 as a bitstring.

rocdl.wmma.f32.16x16x16.bf8_bf8

Return op name rocdl.wmma.f32.16x16x16.bf8_fp8 as a bitstring.

rocdl.wmma.f32.16x16x16.bf8_fp8

Return op name rocdl.wmma.f32.16x16x16.bf16 as a bitstring.

rocdl.wmma.f32.16x16x16.bf16

Return op name rocdl.wmma.f32.16x16x16.f16 as a bitstring.

rocdl.wmma.f32.16x16x16.f16

Return op name rocdl.wmma.f32.16x16x16.fp8_bf8 as a bitstring.

rocdl.wmma.f32.16x16x16.fp8_bf8

Return op name rocdl.wmma.f32.16x16x16.fp8_fp8 as a bitstring.

rocdl.wmma.f32.16x16x16.fp8_fp8

Return op name rocdl.wmma.f32.16x16x32.bf16 as a bitstring.

rocdl.wmma.f32.16x16x32.bf16

Return op name rocdl.wmma.f32.16x16x32.f16 as a bitstring.

rocdl.wmma.f32.16x16x32.f16

Return op name rocdl.wmma.f32.16x16x64.bf8_bf8 as a bitstring.

rocdl.wmma.f32.16x16x64.bf8_bf8

Return op name rocdl.wmma.f32.16x16x64.bf8_fp8 as a bitstring.

rocdl.wmma.f32.16x16x64.bf8_fp8

Return op name rocdl.wmma.f32.16x16x64.fp8_bf8 as a bitstring.

rocdl.wmma.f32.16x16x64.fp8_bf8

Return op name rocdl.wmma.f32.16x16x64.fp8_fp8 as a bitstring.

rocdl.wmma.f32.16x16x64.fp8_fp8

Return op name rocdl.wmma.f32.16x16x128.bf8_bf8 as a bitstring.

rocdl.wmma.f32.16x16x128.bf8_bf8

Return op name rocdl.wmma.f32.16x16x128.bf8_fp8 as a bitstring.

rocdl.wmma.f32.16x16x128.bf8_fp8

Return op name rocdl.wmma.f32.16x16x128.fp8_bf8 as a bitstring.

rocdl.wmma.f32.16x16x128.fp8_bf8

Return op name rocdl.wmma.f32.16x16x128.fp8_fp8 as a bitstring.

rocdl.wmma.f32.16x16x128.fp8_fp8

Return op name rocdl.wmma.i32.16x16x16.iu4 as a bitstring.

rocdl.wmma.i32.16x16x16.iu4

Return op name rocdl.wmma.i32.16x16x16.iu8 as a bitstring.

rocdl.wmma.i32.16x16x16.iu8

Return op name rocdl.wmma.i32.16x16x32.iu4 as a bitstring.

rocdl.wmma.i32.16x16x32.iu4

Return op name rocdl.wmma.i32.16x16x64.iu8 as a bitstring.

rocdl.wmma.i32.16x16x64.iu8

Return op name rocdl.wmma.scale16.f32.16x16x128.f8f6f4 as a bitstring.

rocdl.wmma.scale16.f32.16x16x128.f8f6f4

Return op name rocdl.wmma.scale16.f32.32x16x128.f4 as a bitstring.

rocdl.wmma.scale16.f32.32x16x128.f4

Return op name rocdl.wmma.scale.f32.16x16x128.f8f6f4 as a bitstring.

rocdl.wmma.scale.f32.16x16x128.f8f6f4

Return op name rocdl.wmma.scale.f32.32x16x128.f4 as a bitstring.

rocdl.wmma.scale.f32.32x16x128.f4

Return op name rocdl.workgroup.id.x as a bitstring.

rocdl.workgroup.id.x

Return op name rocdl.workgroup.id.y as a bitstring.

rocdl.workgroup.id.y

Return op name rocdl.workgroup.id.z as a bitstring.

rocdl.workgroup.id.z

Return op name rocdl.workitem.id.x as a bitstring.

rocdl.workitem.id.x

Return op name rocdl.workitem.id.y as a bitstring.

rocdl.workitem.id.y

Return op name rocdl.workitem.id.z as a bitstring.

rocdl.workitem.id.z

Functions

asyncmark()

Return op name rocdl.asyncmark as a bitstring.

asyncmark(ssa)

rocdl.asyncmark - Mark the end of a group of asynchronous operations

Description

This operation, in conjunction with rocdl.wait.asyncmark, forms the compiler-provided framework for tracking explicitly asynchronous memory operations, such as copies to LDS that use async intrinsics and gfx1250's tensor loads.

Details of its behavior can be found in the LLVM documentation on async tracking.

See rocdl.wait.asyncmark's documentation for a usage example.

Example:

// Mark the end of an async operation group.
rocdl.asyncmark

Available on gfx9 and later.

ballot()

Return op name rocdl.ballot as a bitstring.

ballot(ssa)

rocdl.ballot - Vote across thread group

Operands

  • pred - Single, I1, 1-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Ballot provides a bit mask containing the 1-bit predicate value from each lane. The nth bit of the result contains the 1 bit contributed by the nth warp lane.

Example:

// Ballot across thread group.
%0 = rocdl.ballot %pred : i64

barrier()

Return op name rocdl.barrier as a bitstring.

barrier(ssa)

rocdl.barrier

Description

An operation with the same expansion as HIP's __synchthreads();

DEPRECATION NOTICE: Use gpu.barrier, which will expand to these operations, instead.

Example:

// Workgroup barrier with acquire/release fences.
rocdl.barrier

cluster_id_x()

Return op name rocdl.cluster.id.x as a bitstring.

cluster_id_x(ssa)

rocdl.cluster.id.x

Attributes

  • range - Optional, LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Read a hardware register for thread/workgroup/cluster identification. An optional range attribute can constrain the returned value.

Example:

// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32

// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32

cluster_id_y()

Return op name rocdl.cluster.id.y as a bitstring.

cluster_id_y(ssa)

rocdl.cluster.id.y

Attributes

  • range - Optional, LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Read a hardware register for thread/workgroup/cluster identification. An optional range attribute can constrain the returned value.

Example:

// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32

// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32

cluster_id_z()

Return op name rocdl.cluster.id.z as a bitstring.

cluster_id_z(ssa)

rocdl.cluster.id.z

Attributes

  • range - Optional, LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Read a hardware register for thread/workgroup/cluster identification. An optional range attribute can constrain the returned value.

Example:

// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32

// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32

cluster_load_async_to_lds_b8()

Return op name rocdl.cluster.load.async.to.lds.b8 as a bitstring.

cluster_load_async_to_lds_b8(ssa)

rocdl.cluster.load.async.to.lds.b8

Attributes

  • offset - Single, I32Attr, 32-bit signless integer attribute
  • cpol - Single, ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • globalPtr - Single, ROCDLGlobalBuffer, LLVM pointer in address space 1
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3
  • mask - Single, I32, 32-bit signless integer

Description

Broadcasts memory load of 8 bits of data for a cluster of workgroups.

Available on gfx1250+.

Example:

// Cluster broadcast 8-bit load to LDS.
rocdl.cluster.load.async.to.lds.b8 %src, %dst, 0, 0, %mask : !llvm.ptr<1>, !llvm.ptr<3>

cluster_load_async_to_lds_b32()

Return op name rocdl.cluster.load.async.to.lds.b32 as a bitstring.

cluster_load_async_to_lds_b32(ssa)

rocdl.cluster.load.async.to.lds.b32

Attributes

  • offset - Single, I32Attr, 32-bit signless integer attribute
  • cpol - Single, ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • globalPtr - Single, ROCDLGlobalBuffer, LLVM pointer in address space 1
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3
  • mask - Single, I32, 32-bit signless integer

Description

Broadcasts memory load of 32 bits of data for a cluster of workgroups.

Available on gfx1250+.

Example:

// Cluster broadcast 32-bit load to LDS.
rocdl.cluster.load.async.to.lds.b32 %src, %dst, 0, 0, %mask : !llvm.ptr<1>, !llvm.ptr<3>

cluster_load_async_to_lds_b64()

Return op name rocdl.cluster.load.async.to.lds.b64 as a bitstring.

cluster_load_async_to_lds_b64(ssa)

rocdl.cluster.load.async.to.lds.b64

Attributes

  • offset - Single, I32Attr, 32-bit signless integer attribute
  • cpol - Single, ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • globalPtr - Single, ROCDLGlobalBuffer, LLVM pointer in address space 1
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3
  • mask - Single, I32, 32-bit signless integer

Description

Broadcasts memory load of 64 bits of data for a cluster of workgroups.

Available on gfx1250+.

Example:

// Cluster broadcast 64-bit load to LDS.
rocdl.cluster.load.async.to.lds.b64 %src, %dst, 0, 0, %mask : !llvm.ptr<1>, !llvm.ptr<3>

cluster_load_async_to_lds_b128()

Return op name rocdl.cluster.load.async.to.lds.b128 as a bitstring.

cluster_load_async_to_lds_b128(ssa)

rocdl.cluster.load.async.to.lds.b128

Attributes

  • offset - Single, I32Attr, 32-bit signless integer attribute
  • cpol - Single, ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • globalPtr - Single, ROCDLGlobalBuffer, LLVM pointer in address space 1
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3
  • mask - Single, I32, 32-bit signless integer

Description

Broadcasts memory load of 128 bits of data for a cluster of workgroups.

Available on gfx1250+.

Example:

// Cluster broadcast 128-bit load to LDS.
rocdl.cluster.load.async.to.lds.b128 %src, %dst, 0, 0, %mask : !llvm.ptr<1>, !llvm.ptr<3>

cluster_workgroup_id_x()

Return op name rocdl.cluster.workgroup.id.x as a bitstring.

cluster_workgroup_id_x(ssa)

rocdl.cluster.workgroup.id.x

Attributes

  • range - Optional, LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Read a hardware register for thread/workgroup/cluster identification. An optional range attribute can constrain the returned value.

Example:

// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32

// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32

cluster_workgroup_id_y()

Return op name rocdl.cluster.workgroup.id.y as a bitstring.

cluster_workgroup_id_y(ssa)

rocdl.cluster.workgroup.id.y

Attributes

  • range - Optional, LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Read a hardware register for thread/workgroup/cluster identification. An optional range attribute can constrain the returned value.

Example:

// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32

// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32

cluster_workgroup_id_z()

Return op name rocdl.cluster.workgroup.id.z as a bitstring.

cluster_workgroup_id_z(ssa)

rocdl.cluster.workgroup.id.z

Attributes

  • range - Optional, LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Read a hardware register for thread/workgroup/cluster identification. An optional range attribute can constrain the returned value.

Example:

// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32

// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32

cos()

Return op name rocdl.cos as a bitstring.

cos(ssa)

rocdl.cos

Operands

  • arg - Single, LLVM_AnyFloat, floating point LLVM type

Results

  • res - Single, LLVM_AnyFloat, floating point LLVM type

Description

Note: In the general case, prefer the conventional arith, math, or llvm ops over this. Use this ROCDL-specific operation only when you fully understand its implication and when it is strictly necessary. This op is usually chosen when a small loss in precision is acceptable in exchange for higher execution speed.

Example:

%0 = rocdl.cos %a f32 -> f32

cvt_f32_bf8()

Return op name rocdl.cvt.f32.bf8 as a bitstring.

cvt_f32_bf8(ssa)

rocdl.cvt.f32.bf8 - Convert bf8 to f32

Attributes

  • byteSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • srcA - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Convert 8-bit bf8 value from the byteSelth bit of srcA to fp32.

Example:

// Convert bf8 byte 0 to f32.
%0 = rocdl.cvt.f32.bf8 %src[0] : f32

cvt_f32_fp8()

Return op name rocdl.cvt.f32.fp8 as a bitstring.

cvt_f32_fp8(ssa)

rocdl.cvt.f32.fp8 - Convert fp8 to f32

Attributes

  • byteSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • srcA - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Convert 8-bit fp8 value from the byteSelth bit of srcA to fp32.

Example:

// Convert fp8 byte 0 to f32.
%0 = rocdl.cvt.f32.fp8 %src[0] : f32

cvt_pk_bf8_f32()

Return op name rocdl.cvt.pk.bf8.f32 as a bitstring.

cvt_pk_bf8_f32(ssa)

rocdl.cvt.pk.bf8.f32 - Convert two f32's to bf8

Attributes

  • wordSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • srcA - Single, F32, 32-bit float
  • srcB - Single, F32, 32-bit float
  • old - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Convert srcA and srcB to bf8 and store into the low/high word of old, preserving the other word.

Example:

// Pack two f32 values into bf8 in the low word of old.
%0 = rocdl.cvt.pk.bf8.f32 %a, %b -> %old[false] : i32

cvt_pk_f32_bf8()

Return op name rocdl.cvt.pk.f32.bf8 as a bitstring.

cvt_pk_f32_bf8(ssa)

rocdl.cvt.pk.f32.bf8 - Convert packed bf8 to packed f32

Attributes

  • wordSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • src - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Convert src based on $wordSel to packed fp32.

Example:

// Unpack bf8 word to packed f32.
%0 = rocdl.cvt.pk.f32.bf8 %src[false] : vector<2xf32>

cvt_pk_f32_fp8()

Return op name rocdl.cvt.pk.f32.fp8 as a bitstring.

cvt_pk_f32_fp8(ssa)

rocdl.cvt.pk.f32.fp8 - Convert packed fp8 to packed f32

Attributes

  • wordSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • src - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Convert src based on $wordSel to packed fp32.

Example:

// Unpack fp8 word to packed f32.
%0 = rocdl.cvt.pk.f32.fp8 %src[false] : vector<2xf32>

cvt_pk_fp8_f32()

Return op name rocdl.cvt.pk.fp8.f32 as a bitstring.

cvt_pk_fp8_f32(ssa)

rocdl.cvt.pk.fp8.f32 - Convert two f32's to fp8

Attributes

  • wordSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • srcA - Single, F32, 32-bit float
  • srcB - Single, F32, 32-bit float
  • old - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Convert srcA and srcB to fp8 and store into the low/high word of old, preserving the other word.

Example:

// Pack two f32 values into fp8 in the low word of old.
%0 = rocdl.cvt.pk.fp8.f32 %a, %b -> %old[false] : i32

cvt_pkrtz()

Return op name rocdl.cvt.pkrtz as a bitstring.

cvt_pkrtz(ssa)

rocdl.cvt.pkrtz - Convert two f32 input into a vector<2xf16>

Operands

  • srcA - Single, F32, 32-bit float
  • srcB - Single, F32, 32-bit float

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Convert two f32 values into a packed vector<2xf16>.

Example:

// Pack two f32 values into a vector<2xf16> with round-to-zero.
%0 = rocdl.cvt.pkrtz %a, %b : vector<2xf16>

cvt_scale_pk8_bf16_bf8()

Return op name rocdl.cvt.scale.pk8.bf16.bf8 as a bitstring.

cvt_scale_pk8_bf16_bf8(ssa)

rocdl.cvt.scale.pk8.bf16.bf8 - Scales 8 bf8 and converts them to 8 bf16.

Attributes

  • scaleSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
  • scale - Single, I32, 32-bit signless integer

Results

  • res - Single, ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8

Description

Available on gfx1250+.

cvt_scale_pk8_bf16_fp4()

Return op name rocdl.cvt.scale.pk8.bf16.fp4 as a bitstring.

cvt_scale_pk8_bf16_fp4(ssa)

rocdl.cvt.scale.pk8.bf16.fp4 - Scales 8 fp4 and converts them to 8 bf16.

Attributes

  • scaleSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, I32, 32-bit signless integer
  • scale - Single, I32, 32-bit signless integer

Results

  • res - Single, ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8

Description

Available on gfx1250+.

cvt_scale_pk8_bf16_fp8()

Return op name rocdl.cvt.scale.pk8.bf16.fp8 as a bitstring.

cvt_scale_pk8_bf16_fp8(ssa)

rocdl.cvt.scale.pk8.bf16.fp8 - Scales 8 fp8 and converts them to 8 bf16.

Attributes

  • scaleSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
  • scale - Single, I32, 32-bit signless integer

Results

  • res - Single, ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8

Description

Available on gfx1250+.

cvt_scale_pk8_f16_bf8()

Return op name rocdl.cvt.scale.pk8.f16.bf8 as a bitstring.

cvt_scale_pk8_f16_bf8(ssa)

rocdl.cvt.scale.pk8.f16.bf8 - Scales 8 bf8 and converts them to 8 f16.

Attributes

  • scaleSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
  • scale - Single, I32, 32-bit signless integer

Results

  • res - Single, ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8

Description

Available on gfx1250+.

cvt_scale_pk8_f16_fp4()

Return op name rocdl.cvt.scale.pk8.f16.fp4 as a bitstring.

cvt_scale_pk8_f16_fp4(ssa)

rocdl.cvt.scale.pk8.f16.fp4 - Scales 8 fp4 and converts them to 8 f16.

Attributes

  • scaleSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, I32, 32-bit signless integer
  • scale - Single, I32, 32-bit signless integer

Results

  • res - Single, ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8

Description

Available on gfx1250+.

cvt_scale_pk8_f16_fp8()

Return op name rocdl.cvt.scale.pk8.f16.fp8 as a bitstring.

cvt_scale_pk8_f16_fp8(ssa)

rocdl.cvt.scale.pk8.f16.fp8 - Scales 8 fp8 and converts them to 8 f16.

Attributes

  • scaleSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
  • scale - Single, I32, 32-bit signless integer

Results

  • res - Single, ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8

Description

Available on gfx1250+.

cvt_scale_pk8_f32_bf8()

Return op name rocdl.cvt.scale.pk8.f32.bf8 as a bitstring.

cvt_scale_pk8_f32_bf8(ssa)

rocdl.cvt.scale.pk8.f32.bf8 - Scales 8 bf8 and converts them to 8 f32.

Attributes

  • scaleSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
  • scale - Single, I32, 32-bit signless integer

Results

  • res - Single, ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8

Description

Available on gfx1250+.

cvt_scale_pk8_f32_fp4()

Return op name rocdl.cvt.scale.pk8.f32.fp4 as a bitstring.

cvt_scale_pk8_f32_fp4(ssa)

rocdl.cvt.scale.pk8.f32.fp4 - Scales 8 fp4 and converts them to 8 f32.

Attributes

  • scaleSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, I32, 32-bit signless integer
  • scale - Single, I32, 32-bit signless integer

Results

  • res - Single, ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8

Description

Available on gfx1250+.

cvt_scale_pk8_f32_fp8()

Return op name rocdl.cvt.scale.pk8.f32.fp8 as a bitstring.

cvt_scale_pk8_f32_fp8(ssa)

rocdl.cvt.scale.pk8.f32.fp8 - Scales 8 fp8 and converts them to 8 f32.

Attributes

  • scaleSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
  • scale - Single, I32, 32-bit signless integer

Results

  • res - Single, ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8

Description

Available on gfx1250+.

cvt_scale_pk16_bf16_bf6()

Return op name rocdl.cvt.scale.pk16.bf16.bf6 as a bitstring.

cvt_scale_pk16_bf16_bf6(ssa)

rocdl.cvt.scale.pk16.bf16.bf6 - Scales 16 bf6 and converts them to 16 bf16.

Attributes

  • scaleSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
  • scale - Single, I32, 32-bit signless integer

Results

  • res - Single, ROCDL_V16BF16Type, fixed-length vector of bfloat16 type values of length 16

Description

Available on gfx1250+.

cvt_scale_pk16_bf16_fp6()

Return op name rocdl.cvt.scale.pk16.bf16.fp6 as a bitstring.

cvt_scale_pk16_bf16_fp6(ssa)

rocdl.cvt.scale.pk16.bf16.fp6 - Scales 16 fp6 and converts them to 16 bf16.

Attributes

  • scaleSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
  • scale - Single, I32, 32-bit signless integer

Results

  • res - Single, ROCDL_V16BF16Type, fixed-length vector of bfloat16 type values of length 16

Description

Available on gfx1250+.

cvt_scale_pk16_f16_bf6()

Return op name rocdl.cvt.scale.pk16.f16.bf6 as a bitstring.

cvt_scale_pk16_f16_bf6(ssa)

rocdl.cvt.scale.pk16.f16.bf6 - Scales 16 bf6 and converts them to 16 f16.

Attributes

  • scaleSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
  • scale - Single, I32, 32-bit signless integer

Results

  • res - Single, ROCDL_V16F16Type, fixed-length vector of 16-bit float values of length 16

Description

Available on gfx1250+.

cvt_scale_pk16_f16_fp6()

Return op name rocdl.cvt.scale.pk16.f16.fp6 as a bitstring.

cvt_scale_pk16_f16_fp6(ssa)

rocdl.cvt.scale.pk16.f16.fp6 - Scales 16 fp6 and converts them to 16 f16.

Attributes

  • scaleSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
  • scale - Single, I32, 32-bit signless integer

Results

  • res - Single, ROCDL_V16F16Type, fixed-length vector of 16-bit float values of length 16

Description

Available on gfx1250+.

cvt_scale_pk16_f32_bf6()

Return op name rocdl.cvt.scale.pk16.f32.bf6 as a bitstring.

cvt_scale_pk16_f32_bf6(ssa)

rocdl.cvt.scale.pk16.f32.bf6 - Scales 16 bf6 and converts them to 16 f32.

Attributes

  • scaleSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
  • scale - Single, I32, 32-bit signless integer

Results

  • res - Single, ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16

Description

Available on gfx1250+.

cvt_scale_pk16_f32_fp6()

Return op name rocdl.cvt.scale.pk16.f32.fp6 as a bitstring.

cvt_scale_pk16_f32_fp6(ssa)

rocdl.cvt.scale.pk16.f32.fp6 - Scales 16 fp6 and converts them to 16 f32.

Attributes

  • scaleSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
  • scale - Single, I32, 32-bit signless integer

Results

  • res - Single, ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16

Description

Available on gfx1250+.

cvt_scalef32_2xpk16_bf6_f32()

Return op name rocdl.cvt.scalef32.2xpk16.bf6.f32 as a bitstring.

cvt_scalef32_2xpk16_bf6_f32(ssa)

rocdl.cvt.scalef32.2xpk16.bf6.f32 - Scale and convert two vector<16xf32> to 32 packed bf6

Operands

  • src0 - Single, ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16
  • src1 - Single, ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6

Description

Convert 32 single-precision float values, packed into two length-16 vectors that will be logically concanenated, to packed bf6, dividing by the exponent part of scale before doing so.

cvt_scalef32_2xpk16_fp6_f32()

Return op name rocdl.cvt.scalef32.2xpk16.fp6.f32 as a bitstring.

cvt_scalef32_2xpk16_fp6_f32(ssa)

rocdl.cvt.scalef32.2xpk16.fp6.f32 - Scale and convert two vector<16xf32> to 32 packed fp6

Operands

  • src0 - Single, ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16
  • src1 - Single, ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6

Description

Convert 32 single-precision float values, packed into two length-16 vectors that will be logically concanenated, to packed fp6, dividing by the exponent part of scale before doing so.

cvt_scalef32_f16_bf8()

Return op name rocdl.cvt.scalef32.f16.bf8 as a bitstring.

cvt_scalef32_f16_bf8(ssa)

rocdl.cvt.scalef32.f16.bf8 - Scaled convert bf8 from packed vector to f16, updating tied result

Attributes

  • srcSelIndex - Single, I32Attr, 32-bit signless integer attribute
  • dstLoHiSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • oldVdst - Single, ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2
  • src - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2

Description

Convert a bf8 byte from src, selected by srcSelIndex, to f16 while multiplying it by the expontent of scale, and place it into the dstLoHiSelth bit of oldVdst preserving the other element of that vector in the return value.

The bytes are stored as an i32 and not a <4 x i8>.

cvt_scalef32_f16_fp8()

Return op name rocdl.cvt.scalef32.f16.fp8 as a bitstring.

cvt_scalef32_f16_fp8(ssa)

rocdl.cvt.scalef32.f16.fp8 - Scaled convert fp8 from packed vector to f16, updating tied result

Attributes

  • srcSelIndex - Single, I32Attr, 32-bit signless integer attribute
  • dstLoHiSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • oldVdst - Single, ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2
  • src - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2

Description

Convert a fp8 byte from src, selected by srcSelIndex, to f16 while multiplying it by the expontent of scale, and place it into the dstLoHiSelth bit of oldVdst preserving the other element of that vector in the return value.

The bytes are stored as an i32 and not a <4 x i8>.

cvt_scalef32_f32_bf8()

Return op name rocdl.cvt.scalef32.f32.bf8 as a bitstring.

cvt_scalef32_f32_bf8(ssa)

rocdl.cvt.scalef32.f32.bf8 - Scaled convert bf8 from packed vector to f32

Attributes

  • srcSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, F32, 32-bit float

Description

Convert a bf8 byte from src, selected by srcSelIndex, to f32, multiplying it by the exponent of scale.

The bytes are stored in an i32, not a <4 x i8>.

cvt_scalef32_f32_fp8()

Return op name rocdl.cvt.scalef32.f32.fp8 as a bitstring.

cvt_scalef32_f32_fp8(ssa)

rocdl.cvt.scalef32.f32.fp8 - Scaled convert fp8 from packed vector to f32

Attributes

  • srcSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, F32, 32-bit float

Description

Convert a fp8 byte from src, selected by srcSelIndex, to f32, multiplying it by the exponent of scale.

The bytes are stored in an i32, not a <4 x i8>.

cvt_scalef32_pk8_bf8_bf16()

Return op name rocdl.cvt.scalef32.pk8.bf8.bf16 as a bitstring.

cvt_scalef32_pk8_bf8_bf16(ssa)

rocdl.cvt.scalef32.pk8.bf8.bf16 - Scale and convert packed bf16 to packed bf8

Operands

  • src - Single, ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2

Description

Convert 8 packed bf16 values to packed bf8, multiplying by the exponent part of scale before doing so. This op is for gfx1250+ arch.

cvt_scalef32_pk8_bf8_f16()

Return op name rocdl.cvt.scalef32.pk8.bf8.f16 as a bitstring.

cvt_scalef32_pk8_bf8_f16(ssa)

rocdl.cvt.scalef32.pk8.bf8.f16 - Scale and convert packed f16 to packed bf8

Operands

  • src - Single, ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2

Description

Convert 8 packed f16 values to packed bf8, multiplying by the exponent part of scale before doing so. This op is for gfx1250+ arch.

cvt_scalef32_pk8_bf8_f32()

Return op name rocdl.cvt.scalef32.pk8.bf8.f32 as a bitstring.

cvt_scalef32_pk8_bf8_f32(ssa)

rocdl.cvt.scalef32.pk8.bf8.f32 - Scale and convert packed f32 to packed bf8

Operands

  • src - Single, ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2

Description

Convert 8 packed f32 values to packed bf8, multiplying by the exponent part of scale before doing so. This op is for gfx1250+ arch.

cvt_scalef32_pk8_fp4_bf16()

Return op name rocdl.cvt.scalef32.pk8.fp4.bf16 as a bitstring.

cvt_scalef32_pk8_fp4_bf16(ssa)

rocdl.cvt.scalef32.pk8.fp4.bf16 - Scale and convert packed bf16 to packed fp4

Operands

  • src - Single, ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert 8 packed bf16 values to packed fp4, multiplying by the exponent part of scale before doing so. This op is for gfx1250+ arch.

cvt_scalef32_pk8_fp4_f16()

Return op name rocdl.cvt.scalef32.pk8.fp4.f16 as a bitstring.

cvt_scalef32_pk8_fp4_f16(ssa)

rocdl.cvt.scalef32.pk8.fp4.f16 - Scale and convert packed f16 to packed fp4

Operands

  • src - Single, ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert 8 packed f16 values to packed fp4, multiplying by the exponent part of scale before doing so. This op is for gfx1250+ arch.

cvt_scalef32_pk8_fp4_f32()

Return op name rocdl.cvt.scalef32.pk8.fp4.f32 as a bitstring.

cvt_scalef32_pk8_fp4_f32(ssa)

rocdl.cvt.scalef32.pk8.fp4.f32 - Scale and convert packed f32 to packed fp4

Operands

  • src - Single, ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert 8 packed f32 values to packed fp4, multiplying by the exponent part of scale before doing so. This op is for gfx1250+ arch.

cvt_scalef32_pk8_fp8_bf16()

Return op name rocdl.cvt.scalef32.pk8.fp8.bf16 as a bitstring.

cvt_scalef32_pk8_fp8_bf16(ssa)

rocdl.cvt.scalef32.pk8.fp8.bf16 - Scale and convert packed bf16 to packed fp8

Operands

  • src - Single, ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2

Description

Convert 8 packed bf16 values to packed fp8, multiplying by the exponent part of scale before doing so. This op is for gfx1250+ arch.

cvt_scalef32_pk8_fp8_f16()

Return op name rocdl.cvt.scalef32.pk8.fp8.f16 as a bitstring.

cvt_scalef32_pk8_fp8_f16(ssa)

rocdl.cvt.scalef32.pk8.fp8.f16 - Scale and convert packed f16 to packed fp8

Operands

  • src - Single, ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2

Description

Convert 8 packed f16 values to packed fp8, multiplying by the exponent part of scale before doing so. This op is for gfx1250+ arch.

cvt_scalef32_pk8_fp8_f32()

Return op name rocdl.cvt.scalef32.pk8.fp8.f32 as a bitstring.

cvt_scalef32_pk8_fp8_f32(ssa)

rocdl.cvt.scalef32.pk8.fp8.f32 - Scale and convert packed f32 to packed fp8

Operands

  • src - Single, ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2

Description

Convert 8 packed f32 values to packed fp8, multiplying by the exponent part of scale before doing so. This op is for gfx1250+ arch.

cvt_scalef32_pk16_bf6_bf16()

Return op name rocdl.cvt.scalef32.pk16.bf6.bf16 as a bitstring.

cvt_scalef32_pk16_bf6_bf16(ssa)

rocdl.cvt.scalef32.pk16.bf6.bf16 - Scale and convert packed bf16 to packed bf6

Operands

  • src - Single, ROCDL_V16BF16Type, fixed-length vector of bfloat16 type values of length 16
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3

Description

Convert 8 packed bf16 values to packed bf6, multiplying by the exponent part of scale before doing so. This op is for gfx1250+ arch.

cvt_scalef32_pk16_bf6_f16()

Return op name rocdl.cvt.scalef32.pk16.bf6.f16 as a bitstring.

cvt_scalef32_pk16_bf6_f16(ssa)

rocdl.cvt.scalef32.pk16.bf6.f16 - Scale and convert packed f16 to packed bf6

Operands

  • src - Single, ROCDL_V16F16Type, fixed-length vector of 16-bit float values of length 16
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3

Description

Convert 8 packed f16 values to packed bf6, multiplying by the exponent part of scale before doing so. This op is for gfx1250+ arch.

cvt_scalef32_pk16_bf6_f32()

Return op name rocdl.cvt.scalef32.pk16.bf6.f32 as a bitstring.

cvt_scalef32_pk16_bf6_f32(ssa)

rocdl.cvt.scalef32.pk16.bf6.f32 - Scale and convert packed f32 to packed bf6

Operands

  • src - Single, ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3

Description

Convert 8 packed f32 values to packed bf6, multiplying by the exponent part of scale before doing so. This op is for gfx1250+ arch.

cvt_scalef32_pk16_fp6_bf16()

Return op name rocdl.cvt.scalef32.pk16.fp6.bf16 as a bitstring.

cvt_scalef32_pk16_fp6_bf16(ssa)

rocdl.cvt.scalef32.pk16.fp6.bf16 - Scale and convert packed bf16 to packed fp6

Operands

  • src - Single, ROCDL_V16BF16Type, fixed-length vector of bfloat16 type values of length 16
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3

Description

Convert 8 packed bf16 values to packed fp6, multiplying by the exponent part of scale before doing so. This op is for gfx1250+ arch.

cvt_scalef32_pk16_fp6_f16()

Return op name rocdl.cvt.scalef32.pk16.fp6.f16 as a bitstring.

cvt_scalef32_pk16_fp6_f16(ssa)

rocdl.cvt.scalef32.pk16.fp6.f16 - Scale and convert packed f16 to packed fp6

Operands

  • src - Single, ROCDL_V16F16Type, fixed-length vector of 16-bit float values of length 16
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3

Description

Convert 8 packed f16 values to packed fp6, multiplying by the exponent part of scale before doing so. This op is for gfx1250+ arch.

cvt_scalef32_pk16_fp6_f32()

Return op name rocdl.cvt.scalef32.pk16.fp6.f32 as a bitstring.

cvt_scalef32_pk16_fp6_f32(ssa)

rocdl.cvt.scalef32.pk16.fp6.f32 - Scale and convert packed f32 to packed fp6

Operands

  • src - Single, ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3

Description

Convert 8 packed f32 values to packed fp6, multiplying by the exponent part of scale before doing so. This op is for gfx1250+ arch.

cvt_scalef32_pk32_bf6_bf16()

Return op name rocdl.cvt.scalef32.pk32.bf6.bf16 as a bitstring.

cvt_scalef32_pk32_bf6_bf16(ssa)

rocdl.cvt.scalef32.pk32.bf6.bf16 - Scale and convert packed bf16 to packed bf6

Operands

  • src - Single, ROCDL_V32BF16Type, fixed-length vector of bfloat16 type values of length 32
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6

Description

Convert 32 packed bf16 values to packed bf6, dividing by the exponent part of scale before doing so.

cvt_scalef32_pk32_bf6_f16()

Return op name rocdl.cvt.scalef32.pk32.bf6.f16 as a bitstring.

cvt_scalef32_pk32_bf6_f16(ssa)

rocdl.cvt.scalef32.pk32.bf6.f16 - Scale and convert packed f16 to packed bf6

Operands

  • src - Single, ROCDL_V32F16Type, fixed-length vector of 16-bit float values of length 32
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6

Description

Convert 32 packed f16 values to packed bf6, dividing by the exponent part of scale before doing so.

cvt_scalef32_pk32_bf16_bf6()

Return op name rocdl.cvt.scalef32.pk32.bf16.bf6 as a bitstring.

cvt_scalef32_pk32_bf16_bf6(ssa)

rocdl.cvt.scalef32.pk32.bf16.bf6 - Scale and convert packed bf6 to packed bf16

Operands

  • src - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V32BF16Type, fixed-length vector of bfloat16 type values of length 32

Description

Convert 32 packed bf6 values to packed bf16, multiplying by the exponent part of scale before doing so.

cvt_scalef32_pk32_bf16_fp6()

Return op name rocdl.cvt.scalef32.pk32.bf16.fp6 as a bitstring.

cvt_scalef32_pk32_bf16_fp6(ssa)

rocdl.cvt.scalef32.pk32.bf16.fp6 - Scale and convert packed fp6 to packed bf16

Operands

  • src - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V32BF16Type, fixed-length vector of bfloat16 type values of length 32

Description

Convert 32 packed fp6 values to packed bf16, multiplying by the exponent part of scale before doing so.

cvt_scalef32_pk32_f16_bf6()

Return op name rocdl.cvt.scalef32.pk32.f16.bf6 as a bitstring.

cvt_scalef32_pk32_f16_bf6(ssa)

rocdl.cvt.scalef32.pk32.f16.bf6 - Scale and convert packed bf6 to packed f16

Operands

  • src - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V32F16Type, fixed-length vector of 16-bit float values of length 32

Description

Convert 32 packed bf6 values to packed f16, multiplying by the exponent part of scale before doing so.

cvt_scalef32_pk32_f16_fp6()

Return op name rocdl.cvt.scalef32.pk32.f16.fp6 as a bitstring.

cvt_scalef32_pk32_f16_fp6(ssa)

rocdl.cvt.scalef32.pk32.f16.fp6 - Scale and convert packed fp6 to packed f16

Operands

  • src - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V32F16Type, fixed-length vector of 16-bit float values of length 32

Description

Convert 32 packed fp6 values to packed f16, multiplying by the exponent part of scale before doing so.

cvt_scalef32_pk32_f32_bf6()

Return op name rocdl.cvt.scalef32.pk32.f32.bf6 as a bitstring.

cvt_scalef32_pk32_f32_bf6(ssa)

rocdl.cvt.scalef32.pk32.f32.bf6 - Scale and convert packed bf6 to packed f32

Operands

  • src - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V32F32Type, fixed-length vector of 32-bit float values of length 32

Description

Convert 32 packed bf6 values to packed f32, multiplying by the exponent part of scale before doing so.

cvt_scalef32_pk32_f32_fp6()

Return op name rocdl.cvt.scalef32.pk32.f32.fp6 as a bitstring.

cvt_scalef32_pk32_f32_fp6(ssa)

rocdl.cvt.scalef32.pk32.f32.fp6 - Scale and convert packed fp6 to packed f32

Operands

  • src - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V32F32Type, fixed-length vector of 32-bit float values of length 32

Description

Convert 32 packed fp6 values to packed f32, multiplying by the exponent part of scale before doing so.

cvt_scalef32_pk32_fp6_bf16()

Return op name rocdl.cvt.scalef32.pk32.fp6.bf16 as a bitstring.

cvt_scalef32_pk32_fp6_bf16(ssa)

rocdl.cvt.scalef32.pk32.fp6.bf16 - Scale and convert packed bf16 to packed fp6

Operands

  • src - Single, ROCDL_V32BF16Type, fixed-length vector of bfloat16 type values of length 32
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6

Description

Convert 32 packed bf16 values to packed fp6, dividing by the exponent part of scale before doing so.

cvt_scalef32_pk32_fp6_f16()

Return op name rocdl.cvt.scalef32.pk32.fp6.f16 as a bitstring.

cvt_scalef32_pk32_fp6_f16(ssa)

rocdl.cvt.scalef32.pk32.fp6.f16 - Scale and convert packed f16 to packed fp6

Operands

  • src - Single, ROCDL_V32F16Type, fixed-length vector of 16-bit float values of length 32
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6

Description

Convert 32 packed f16 values to packed fp6, dividing by the exponent part of scale before doing so.

cvt_scalef32_pk_bf8_bf16()

Return op name rocdl.cvt.scalef32.pk.bf8.bf16 as a bitstring.

cvt_scalef32_pk_bf8_bf16(ssa)

rocdl.cvt.scalef32.pk.bf8.bf16 - Scaled convert two bf16to two bf8, updating packed vector

Attributes

  • dstLoHiSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • oldVdst - Single, ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2
  • src0 - Single, ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2

Description

Convert two bf16 values in src0 to two bf8 bytes, dividing by the exponent in scale. The bytes are packed into a 16-bit value which is inserted into oldVdst at the dstLoHiSel position, with the entire updated vector being returned.

cvt_scalef32_pk_bf8_f16()

Return op name rocdl.cvt.scalef32.pk.bf8.f16 as a bitstring.

cvt_scalef32_pk_bf8_f16(ssa)

rocdl.cvt.scalef32.pk.bf8.f16 - Scaled convert two f16to two bf8, updating packed vector

Attributes

  • dstLoHiSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • oldVdst - Single, ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2
  • src0 - Single, ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2

Description

Convert two f16 values in src0 to two bf8 bytes, dividing by the exponent in scale. The bytes are packed into a 16-bit value which is inserted into oldVdst at the dstLoHiSel position, with the entire updated vector being returned.

cvt_scalef32_pk_bf8_f32()

Return op name rocdl.cvt.scalef32.pk.bf8.f32 as a bitstring.

cvt_scalef32_pk_bf8_f32(ssa)

rocdl.cvt.scalef32.pk.bf8.f32 - Scaled convert two f32 to two bf8, updating packed vector

Attributes

  • dstLoHiSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • oldVdst - Single, ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2
  • src0 - Single, F32, 32-bit float
  • src1 - Single, F32, 32-bit float
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2

Description

Convert two f32 values in src0 and src1 to two bf8 bytes, dividing by the exponent in scale. The bytes are packed into a 16-bit value which is inserted into oldVdst at the dstLoHiSel position, with the entire updated vector being returned.

cvt_scalef32_pk_bf16_bf8()

Return op name rocdl.cvt.scalef32.pk.bf16.bf8 as a bitstring.

cvt_scalef32_pk_bf16_bf8(ssa)

rocdl.cvt.scalef32.pk.bf16.bf8 - Scaled convert two bf8to two bf16

Attributes

  • srcLoHiSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • src - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2

Description

Convert two packed bf8 values in src0 to two bf16 values, multiplying by the exponent in scale. The two values to be converted are selected from the low or high half of src (a packed vector represented as an i32) on the basis of srcLoHiSel.

cvt_scalef32_pk_bf16_fp4()

Return op name rocdl.cvt.scalef32.pk.bf16.fp4 as a bitstring.

cvt_scalef32_pk_bf16_fp4(ssa)

rocdl.cvt.scalef32.pk.bf16.fp4 - Scale and convert two packed fp4 to packed bf16

Attributes

  • srcSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2

Description

Convert two packed fp4 (f4E2M1) values stored as one byte of a 32-bit integer to packed bf16, multiplying by the exponent part of scale before doing so.

The byte to convert is chosen by srcSelIndex.

cvt_scalef32_pk_bf16_fp8()

Return op name rocdl.cvt.scalef32.pk.bf16.fp8 as a bitstring.

cvt_scalef32_pk_bf16_fp8(ssa)

rocdl.cvt.scalef32.pk.bf16.fp8 - Scaled convert two fp8to two bf16

Attributes

  • srcLoHiSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • src - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2

Description

Convert two packed fp8 values in src0 to two bf16 values, multiplying by the exponent in scale. The two values to be converted are selected from the low or high half of src (a packed vector represented as an i32) on the basis of srcLoHiSel.

cvt_scalef32_pk_f16_bf8()

Return op name rocdl.cvt.scalef32.pk.f16.bf8 as a bitstring.

cvt_scalef32_pk_f16_bf8(ssa)

rocdl.cvt.scalef32.pk.f16.bf8 - Scaled convert two bf8to two f16

Attributes

  • srcLoHiSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • src - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2

Description

Convert two packed bf8 values in src0 to two f16 values, multiplying by the exponent in scale. The two values to be converted are selected from the low or high half of src (a packed vector represented as an i32) on the basis of srcLoHiSel.

cvt_scalef32_pk_f16_fp4()

Return op name rocdl.cvt.scalef32.pk.f16.fp4 as a bitstring.

cvt_scalef32_pk_f16_fp4(ssa)

rocdl.cvt.scalef32.pk.f16.fp4 - Scale and convert two packed fp4 to packed f16

Attributes

  • srcSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2

Description

Convert two packed fp4 (f4E2M1) values stored as one byte of a 32-bit integer to packed f16, multiplying by the exponent part of scale before doing so.

The byte to convert is chosen by srcSelIndex.

cvt_scalef32_pk_f16_fp8()

Return op name rocdl.cvt.scalef32.pk.f16.fp8 as a bitstring.

cvt_scalef32_pk_f16_fp8(ssa)

rocdl.cvt.scalef32.pk.f16.fp8 - Scaled convert two fp8to two f16

Attributes

  • srcLoHiSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • src - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2

Description

Convert two packed fp8 values in src0 to two f16 values, multiplying by the exponent in scale. The two values to be converted are selected from the low or high half of src (a packed vector represented as an i32) on the basis of srcLoHiSel.

cvt_scalef32_pk_f32_bf8()

Return op name rocdl.cvt.scalef32.pk.f32.bf8 as a bitstring.

cvt_scalef32_pk_f32_bf8(ssa)

rocdl.cvt.scalef32.pk.f32.bf8 - Scaled convert two bf8to two f32

Attributes

  • srcLoHiSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • src - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2F32Type, fixed-length vector of 32-bit float values of length 2

Description

Convert two packed bf8 values in src0 to two f32 values, multiplying by the exponent in scale. The two values to be converted are selected from the low or high half of src (a packed vector represented as an i32) on the basis of srcLoHiSel.

cvt_scalef32_pk_f32_fp4()

Return op name rocdl.cvt.scalef32.pk.f32.fp4 as a bitstring.

cvt_scalef32_pk_f32_fp4(ssa)

rocdl.cvt.scalef32.pk.f32.fp4 - Scale and convert two packed fp4 to packed f32

Attributes

  • srcSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • src - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2F32Type, fixed-length vector of 32-bit float values of length 2

Description

Convert two packed fp4 (f4E2M1) values stored as one byte of a 32-bit integer to packed f32, multiplying by the exponent part of scale before doing so.

The byte to convert is chosen by srcSelIndex.

cvt_scalef32_pk_f32_fp8()

Return op name rocdl.cvt.scalef32.pk.f32.fp8 as a bitstring.

cvt_scalef32_pk_f32_fp8(ssa)

rocdl.cvt.scalef32.pk.f32.fp8 - Scaled convert two fp8to two f32

Attributes

  • srcLoHiSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • src - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2F32Type, fixed-length vector of 32-bit float values of length 2

Description

Convert two packed fp8 values in src0 to two f32 values, multiplying by the exponent in scale. The two values to be converted are selected from the low or high half of src (a packed vector represented as an i32) on the basis of srcLoHiSel.

cvt_scalef32_pk_fp4_bf16()

Return op name rocdl.cvt.scalef32.pk.fp4.bf16 as a bitstring.

cvt_scalef32_pk_fp4_bf16(ssa)

rocdl.cvt.scalef32.pk.fp4.bf16 - Scale and convert two bf16 to packed fp4, updating tied vector

Attributes

  • dstSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • oldVdst - Single, I32, 32-bit signless integer
  • src - Single, ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert two packed bf16 values to packed fp4, dividing by the exponent part of scale before doing so.

The two scaled values are packed into a byte. That byte is used to update the dstSelIndexth byte of oldVdst, which is returned in its entirity.

cvt_scalef32_pk_fp4_f16()

Return op name rocdl.cvt.scalef32.pk.fp4.f16 as a bitstring.

cvt_scalef32_pk_fp4_f16(ssa)

rocdl.cvt.scalef32.pk.fp4.f16 - Scale and convert two f16 to packed fp4, updating tied vector

Attributes

  • dstSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • oldVdst - Single, I32, 32-bit signless integer
  • src - Single, ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert two packed f16 values to packed fp4, dividing by the exponent part of scale before doing so.

The two scaled values are packed into a byte. That byte is used to update the dstSelIndexth byte of oldVdst, which is returned in its entirity.

cvt_scalef32_pk_fp4_f32()

Return op name rocdl.cvt.scalef32.pk.fp4.f32 as a bitstring.

cvt_scalef32_pk_fp4_f32(ssa)

rocdl.cvt.scalef32.pk.fp4.f32 - Scale and convert two f32 values to two packed fp4, updating tied vector

Attributes

  • dstSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • oldVdst - Single, I32, 32-bit signless integer
  • src0 - Single, F32, 32-bit float
  • src1 - Single, F32, 32-bit float
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert two single-precision float values, passed in src0 and src1 into two fp4 values, dividing them by the expontent part of scale before doing so.

The two scaled values are packed into a byte. That byte is used to update the dstSelIndexth byte of oldVdst, which is returned in its entirity.

Example:

// Scaled convert two f32 values to packed fp4 in byte 0 of old.
%0 = rocdl.cvt.scalef32.pk.fp4.f32 %a, %b, %scale -> %old[0] : i32

cvt_scalef32_pk_fp8_bf16()

Return op name rocdl.cvt.scalef32.pk.fp8.bf16 as a bitstring.

cvt_scalef32_pk_fp8_bf16(ssa)

rocdl.cvt.scalef32.pk.fp8.bf16 - Scaled convert two bf16to two fp8, updating packed vector

Attributes

  • dstLoHiSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • oldVdst - Single, ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2
  • src0 - Single, ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2

Description

Convert two bf16 values in src0 to two fp8 bytes, dividing by the exponent in scale. The bytes are packed into a 16-bit value which is inserted into oldVdst at the dstLoHiSel position, with the entire updated vector being returned.

cvt_scalef32_pk_fp8_f16()

Return op name rocdl.cvt.scalef32.pk.fp8.f16 as a bitstring.

cvt_scalef32_pk_fp8_f16(ssa)

rocdl.cvt.scalef32.pk.fp8.f16 - Scaled convert two f16to two fp8, updating packed vector

Attributes

  • dstLoHiSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • oldVdst - Single, ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2
  • src0 - Single, ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2

Description

Convert two f16 values in src0 to two fp8 bytes, dividing by the exponent in scale. The bytes are packed into a 16-bit value which is inserted into oldVdst at the dstLoHiSel position, with the entire updated vector being returned.

cvt_scalef32_pk_fp8_f32()

Return op name rocdl.cvt.scalef32.pk.fp8.f32 as a bitstring.

cvt_scalef32_pk_fp8_f32(ssa)

rocdl.cvt.scalef32.pk.fp8.f32 - Scaled convert two f32 to two fp8, updating packed vector

Attributes

  • dstLoHiSel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • oldVdst - Single, ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2
  • src0 - Single, F32, 32-bit float
  • src1 - Single, F32, 32-bit float
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2

Description

Convert two f32 values in src0 and src1 to two fp8 bytes, dividing by the exponent in scale. The bytes are packed into a 16-bit value which is inserted into oldVdst at the dstLoHiSel position, with the entire updated vector being returned.

cvt_scalef32_sr_bf8_bf16()

Return op name rocdl.cvt.scalef32.sr.bf8.bf16 as a bitstring.

cvt_scalef32_sr_bf8_bf16(ssa)

rocdl.cvt.scalef32.sr.bf8.bf16 - Scaled convert bf16to bf8 with stochiastic rounding, updating packed vector

Attributes

  • dstSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • oldVdst - Single, I32, 32-bit signless integer
  • src0 - Single, BF16, bfloat16 type
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert a bf16 value in src0 to a bf8 bytes, dividing by the exponent in scale and using seed for stochiastic rounding. Place the resulting byte in the dstSelIndexth bit of oldVdst and return the entire packed vector, which is stored as an i32.

cvt_scalef32_sr_bf8_f16()

Return op name rocdl.cvt.scalef32.sr.bf8.f16 as a bitstring.

cvt_scalef32_sr_bf8_f16(ssa)

rocdl.cvt.scalef32.sr.bf8.f16 - Scaled convert f16to bf8 with stochiastic rounding, updating packed vector

Attributes

  • dstSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • oldVdst - Single, I32, 32-bit signless integer
  • src0 - Single, F16, 16-bit float
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert a f16 value in src0 to a bf8 bytes, dividing by the exponent in scale and using seed for stochiastic rounding. Place the resulting byte in the dstSelIndexth bit of oldVdst and return the entire packed vector, which is stored as an i32.

cvt_scalef32_sr_bf8_f32()

Return op name rocdl.cvt.scalef32.sr.bf8.f32 as a bitstring.

cvt_scalef32_sr_bf8_f32(ssa)

rocdl.cvt.scalef32.sr.bf8.f32 - Scaled convert f32to bf8 with stochiastic rounding, updating packed vector

Attributes

  • dstSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • oldVdst - Single, I32, 32-bit signless integer
  • src0 - Single, F32, 32-bit float
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert a f32 value in src0 to a bf8 bytes, dividing by the exponent in scale and using seed for stochiastic rounding. Place the resulting byte in the dstSelIndexth bit of oldVdst and return the entire packed vector, which is stored as an i32.

cvt_scalef32_sr_fp8_bf16()

Return op name rocdl.cvt.scalef32.sr.fp8.bf16 as a bitstring.

cvt_scalef32_sr_fp8_bf16(ssa)

rocdl.cvt.scalef32.sr.fp8.bf16 - Scaled convert bf16to fp8 with stochiastic rounding, updating packed vector

Attributes

  • dstSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • oldVdst - Single, I32, 32-bit signless integer
  • src0 - Single, BF16, bfloat16 type
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert a bf16 value in src0 to a fp8 bytes, dividing by the exponent in scale and using seed for stochiastic rounding. Place the resulting byte in the dstSelIndexth bit of oldVdst and return the entire packed vector, which is stored as an i32.

cvt_scalef32_sr_fp8_f16()

Return op name rocdl.cvt.scalef32.sr.fp8.f16 as a bitstring.

cvt_scalef32_sr_fp8_f16(ssa)

rocdl.cvt.scalef32.sr.fp8.f16 - Scaled convert f16to fp8 with stochiastic rounding, updating packed vector

Attributes

  • dstSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • oldVdst - Single, I32, 32-bit signless integer
  • src0 - Single, F16, 16-bit float
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert a f16 value in src0 to a fp8 bytes, dividing by the exponent in scale and using seed for stochiastic rounding. Place the resulting byte in the dstSelIndexth bit of oldVdst and return the entire packed vector, which is stored as an i32.

cvt_scalef32_sr_fp8_f32()

Return op name rocdl.cvt.scalef32.sr.fp8.f32 as a bitstring.

cvt_scalef32_sr_fp8_f32(ssa)

rocdl.cvt.scalef32.sr.fp8.f32 - Scaled convert f32to fp8 with stochiastic rounding, updating packed vector

Attributes

  • dstSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • oldVdst - Single, I32, 32-bit signless integer
  • src0 - Single, F32, 32-bit float
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert a f32 value in src0 to a fp8 bytes, dividing by the exponent in scale and using seed for stochiastic rounding. Place the resulting byte in the dstSelIndexth bit of oldVdst and return the entire packed vector, which is stored as an i32.

cvt_scalef32_sr_pk8_bf8_bf16()

Return op name rocdl.cvt.scalef32.sr.pk8.bf8.bf16 as a bitstring.

cvt_scalef32_sr_pk8_bf8_bf16(ssa)

rocdl.cvt.scalef32.sr.pk8.bf8.bf16 - Scale and convert packed bf16 to packed bf8 with stochastic rounding

Operands

  • src - Single, ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2

Description

Convert 8 packed bf16 values to packed bf8, multiplying by the exponent part of scale before doing so and apply stochastic rounding. This op is for gfx1250+ arch.

cvt_scalef32_sr_pk8_bf8_f16()

Return op name rocdl.cvt.scalef32.sr.pk8.bf8.f16 as a bitstring.

cvt_scalef32_sr_pk8_bf8_f16(ssa)

rocdl.cvt.scalef32.sr.pk8.bf8.f16 - Scale and convert packed f16 to packed bf8 with stochastic rounding

Operands

  • src - Single, ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2

Description

Convert 8 packed f16 values to packed bf8, multiplying by the exponent part of scale before doing so and apply stochastic rounding. This op is for gfx1250+ arch.

cvt_scalef32_sr_pk8_bf8_f32()

Return op name rocdl.cvt.scalef32.sr.pk8.bf8.f32 as a bitstring.

cvt_scalef32_sr_pk8_bf8_f32(ssa)

rocdl.cvt.scalef32.sr.pk8.bf8.f32 - Scale and convert packed f32 to packed bf8 with stochastic rounding

Operands

  • src - Single, ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2

Description

Convert 8 packed f32 values to packed bf8, multiplying by the exponent part of scale before doing so and apply stochastic rounding. This op is for gfx1250+ arch.

cvt_scalef32_sr_pk8_fp4_bf16()

Return op name rocdl.cvt.scalef32.sr.pk8.fp4.bf16 as a bitstring.

cvt_scalef32_sr_pk8_fp4_bf16(ssa)

rocdl.cvt.scalef32.sr.pk8.fp4.bf16 - Scale and convert packed bf16 to packed fp4 with stochastic rounding

Operands

  • src - Single, ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert 8 packed bf16 values to packed fp4, multiplying by the exponent part of scale before doing so and apply stochastic rounding. This op is for gfx1250+ arch.

cvt_scalef32_sr_pk8_fp4_f16()

Return op name rocdl.cvt.scalef32.sr.pk8.fp4.f16 as a bitstring.

cvt_scalef32_sr_pk8_fp4_f16(ssa)

rocdl.cvt.scalef32.sr.pk8.fp4.f16 - Scale and convert packed f16 to packed fp4 with stochastic rounding

Operands

  • src - Single, ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert 8 packed f16 values to packed fp4, multiplying by the exponent part of scale before doing so and apply stochastic rounding. This op is for gfx1250+ arch.

cvt_scalef32_sr_pk8_fp4_f32()

Return op name rocdl.cvt.scalef32.sr.pk8.fp4.f32 as a bitstring.

cvt_scalef32_sr_pk8_fp4_f32(ssa)

rocdl.cvt.scalef32.sr.pk8.fp4.f32 - Scale and convert packed f32 to packed fp4 with stochastic rounding

Operands

  • src - Single, ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert 8 packed f32 values to packed fp4, multiplying by the exponent part of scale before doing so and apply stochastic rounding. This op is for gfx1250+ arch.

cvt_scalef32_sr_pk8_fp8_bf16()

Return op name rocdl.cvt.scalef32.sr.pk8.fp8.bf16 as a bitstring.

cvt_scalef32_sr_pk8_fp8_bf16(ssa)

rocdl.cvt.scalef32.sr.pk8.fp8.bf16 - Scale and convert packed bf16 to packed fp8 with stochastic rounding

Operands

  • src - Single, ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2

Description

Convert 8 packed bf16 values to packed fp8, multiplying by the exponent part of scale before doing so and apply stochastic rounding. This op is for gfx1250+ arch.

cvt_scalef32_sr_pk8_fp8_f16()

Return op name rocdl.cvt.scalef32.sr.pk8.fp8.f16 as a bitstring.

cvt_scalef32_sr_pk8_fp8_f16(ssa)

rocdl.cvt.scalef32.sr.pk8.fp8.f16 - Scale and convert packed f16 to packed fp8 with stochastic rounding

Operands

  • src - Single, ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2

Description

Convert 8 packed f16 values to packed fp8, multiplying by the exponent part of scale before doing so and apply stochastic rounding. This op is for gfx1250+ arch.

cvt_scalef32_sr_pk8_fp8_f32()

Return op name rocdl.cvt.scalef32.sr.pk8.fp8.f32 as a bitstring.

cvt_scalef32_sr_pk8_fp8_f32(ssa)

rocdl.cvt.scalef32.sr.pk8.fp8.f32 - Scale and convert packed f32 to packed fp8 with stochastic rounding

Operands

  • src - Single, ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2

Description

Convert 8 packed f32 values to packed fp8, multiplying by the exponent part of scale before doing so and apply stochastic rounding. This op is for gfx1250+ arch.

cvt_scalef32_sr_pk16_bf6_bf16()

Return op name rocdl.cvt.scalef32.sr.pk16.bf6.bf16 as a bitstring.

cvt_scalef32_sr_pk16_bf6_bf16(ssa)

rocdl.cvt.scalef32.sr.pk16.bf6.bf16 - Scale and convert packed bf16 to packed bf6 with stochastic rounding

Operands

  • src - Single, ROCDL_V16BF16Type, fixed-length vector of bfloat16 type values of length 16
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3

Description

Convert 8 packed bf16 values to packed bf6, multiplying by the exponent part of scale before doing so and apply stochastic rounding. This op is for gfx1250+ arch.

cvt_scalef32_sr_pk16_bf6_f16()

Return op name rocdl.cvt.scalef32.sr.pk16.bf6.f16 as a bitstring.

cvt_scalef32_sr_pk16_bf6_f16(ssa)

rocdl.cvt.scalef32.sr.pk16.bf6.f16 - Scale and convert packed f16 to packed bf6 with stochastic rounding

Operands

  • src - Single, ROCDL_V16F16Type, fixed-length vector of 16-bit float values of length 16
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3

Description

Convert 8 packed f16 values to packed bf6, multiplying by the exponent part of scale before doing so and apply stochastic rounding. This op is for gfx1250+ arch.

cvt_scalef32_sr_pk16_bf6_f32()

Return op name rocdl.cvt.scalef32.sr.pk16.bf6.f32 as a bitstring.

cvt_scalef32_sr_pk16_bf6_f32(ssa)

rocdl.cvt.scalef32.sr.pk16.bf6.f32 - Scale and convert packed f32 to packed bf6 with stochastic rounding

Operands

  • src - Single, ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3

Description

Convert 8 packed f32 values to packed bf6, multiplying by the exponent part of scale before doing so and apply stochastic rounding. This op is for gfx1250+ arch.

cvt_scalef32_sr_pk16_fp6_bf16()

Return op name rocdl.cvt.scalef32.sr.pk16.fp6.bf16 as a bitstring.

cvt_scalef32_sr_pk16_fp6_bf16(ssa)

rocdl.cvt.scalef32.sr.pk16.fp6.bf16 - Scale and convert packed bf16 to packed fp6 with stochastic rounding

Operands

  • src - Single, ROCDL_V16BF16Type, fixed-length vector of bfloat16 type values of length 16
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3

Description

Convert 8 packed bf16 values to packed fp6, multiplying by the exponent part of scale before doing so and apply stochastic rounding. This op is for gfx1250+ arch.

cvt_scalef32_sr_pk16_fp6_f16()

Return op name rocdl.cvt.scalef32.sr.pk16.fp6.f16 as a bitstring.

cvt_scalef32_sr_pk16_fp6_f16(ssa)

rocdl.cvt.scalef32.sr.pk16.fp6.f16 - Scale and convert packed f16 to packed fp6 with stochastic rounding

Operands

  • src - Single, ROCDL_V16F16Type, fixed-length vector of 16-bit float values of length 16
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3

Description

Convert 8 packed f16 values to packed fp6, multiplying by the exponent part of scale before doing so and apply stochastic rounding. This op is for gfx1250+ arch.

cvt_scalef32_sr_pk16_fp6_f32()

Return op name rocdl.cvt.scalef32.sr.pk16.fp6.f32 as a bitstring.

cvt_scalef32_sr_pk16_fp6_f32(ssa)

rocdl.cvt.scalef32.sr.pk16.fp6.f32 - Scale and convert packed f32 to packed fp6 with stochastic rounding

Operands

  • src - Single, ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3

Description

Convert 8 packed f32 values to packed fp6, multiplying by the exponent part of scale before doing so and apply stochastic rounding. This op is for gfx1250+ arch.

cvt_scalef32_sr_pk32_bf6_bf16()

Return op name rocdl.cvt.scalef32.sr.pk32.bf6.bf16 as a bitstring.

cvt_scalef32_sr_pk32_bf6_bf16(ssa)

rocdl.cvt.scalef32.sr.pk32.bf6.bf16 - Scale and convert packed bf16 to packed bf6 with stochiastic rounding

Operands

  • src - Single, ROCDL_V32BF16Type, fixed-length vector of bfloat16 type values of length 32
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6

Description

Convert 32 packed bf16 values to packed bf6, dividing by the exponent part of scale before doing so and applying random rounding derived from seed.

cvt_scalef32_sr_pk32_bf6_f16()

Return op name rocdl.cvt.scalef32.sr.pk32.bf6.f16 as a bitstring.

cvt_scalef32_sr_pk32_bf6_f16(ssa)

rocdl.cvt.scalef32.sr.pk32.bf6.f16 - Scale and convert packed f16 to packed bf6 with stochiastic rounding

Operands

  • src - Single, ROCDL_V32F16Type, fixed-length vector of 16-bit float values of length 32
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6

Description

Convert 32 packed f16 values to packed bf6, dividing by the exponent part of scale before doing so and applying random rounding derived from seed.

cvt_scalef32_sr_pk32_bf6_f32()

Return op name rocdl.cvt.scalef32.sr.pk32.bf6.f32 as a bitstring.

cvt_scalef32_sr_pk32_bf6_f32(ssa)

rocdl.cvt.scalef32.sr.pk32.bf6.f32 - Scale and convert packed f32 to packed bf6 with stochiastic rounding

Operands

  • src - Single, ROCDL_V32F32Type, fixed-length vector of 32-bit float values of length 32
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6

Description

Convert 32 packed f32 values to packed bf6, dividing by the exponent part of scale before doing so and applying random rounding derived from seed.

cvt_scalef32_sr_pk32_fp6_bf16()

Return op name rocdl.cvt.scalef32.sr.pk32.fp6.bf16 as a bitstring.

cvt_scalef32_sr_pk32_fp6_bf16(ssa)

rocdl.cvt.scalef32.sr.pk32.fp6.bf16 - Scale and convert packed bf16 to packed fp6 with stochiastic rounding

Operands

  • src - Single, ROCDL_V32BF16Type, fixed-length vector of bfloat16 type values of length 32
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6

Description

Convert 32 packed bf16 values to packed fp6, dividing by the exponent part of scale before doing so and applying random rounding derived from seed.

cvt_scalef32_sr_pk32_fp6_f16()

Return op name rocdl.cvt.scalef32.sr.pk32.fp6.f16 as a bitstring.

cvt_scalef32_sr_pk32_fp6_f16(ssa)

rocdl.cvt.scalef32.sr.pk32.fp6.f16 - Scale and convert packed f16 to packed fp6 with stochiastic rounding

Operands

  • src - Single, ROCDL_V32F16Type, fixed-length vector of 16-bit float values of length 32
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6

Description

Convert 32 packed f16 values to packed fp6, dividing by the exponent part of scale before doing so and applying random rounding derived from seed.

cvt_scalef32_sr_pk32_fp6_f32()

Return op name rocdl.cvt.scalef32.sr.pk32.fp6.f32 as a bitstring.

cvt_scalef32_sr_pk32_fp6_f32(ssa)

rocdl.cvt.scalef32.sr.pk32.fp6.f32 - Scale and convert packed f32 to packed fp6 with stochiastic rounding

Operands

  • src - Single, ROCDL_V32F32Type, fixed-length vector of 32-bit float values of length 32
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6

Description

Convert 32 packed f32 values to packed fp6, dividing by the exponent part of scale before doing so and applying random rounding derived from seed.

cvt_scalef32_sr_pk_fp4_bf16()

Return op name rocdl.cvt.scalef32.sr.pk.fp4.bf16 as a bitstring.

cvt_scalef32_sr_pk_fp4_bf16(ssa)

rocdl.cvt.scalef32.sr.pk.fp4.bf16 - Scale and convert two bf16 to packed fp4 with stochiastic rounding, updating tied vector

Attributes

  • dstSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • oldVdst - Single, I32, 32-bit signless integer
  • src - Single, ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert two packed bf16 values to packed fp4, dividing by the exponent part of scale before doing so and using seed as the random seed for stochiastic rounding.

The two scaled values are packed (little-endian) into a byte. That byte is used to update the dstSelIndexth byte of oldVdst, which is returned in its entirity.

cvt_scalef32_sr_pk_fp4_f16()

Return op name rocdl.cvt.scalef32.sr.pk.fp4.f16 as a bitstring.

cvt_scalef32_sr_pk_fp4_f16(ssa)

rocdl.cvt.scalef32.sr.pk.fp4.f16 - Scale and convert two f16 to packed fp4 with stochiastic rounding, updating tied vector

Attributes

  • dstSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • oldVdst - Single, I32, 32-bit signless integer
  • src - Single, ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert two packed f16 values to packed fp4, dividing by the exponent part of scale before doing so and using seed as the random seed for stochiastic rounding.

The two scaled values are packed (little-endian) into a byte. That byte is used to update the dstSelIndexth byte of oldVdst, which is returned in its entirity.

cvt_scalef32_sr_pk_fp4_f32()

Return op name rocdl.cvt.scalef32.sr.pk.fp4.f32 as a bitstring.

cvt_scalef32_sr_pk_fp4_f32(ssa)

rocdl.cvt.scalef32.sr.pk.fp4.f32 - Scale and convert two f32 to packed fp4 with stochiastic rounding, updating tied vector

Attributes

  • dstSelIndex - Single, I32Attr, 32-bit signless integer attribute

Operands

  • oldVdst - Single, I32, 32-bit signless integer
  • src - Single, ROCDL_V2F32Type, fixed-length vector of 32-bit float values of length 2
  • seed - Single, I32, 32-bit signless integer
  • scale - Single, F32, 32-bit float

Results

  • res - Single, I32, 32-bit signless integer

Description

Convert two packed f32 values to packed fp4, dividing by the exponent part of scale before doing so and using seed as the random seed for stochiastic rounding.

The two scaled values are packed (little-endian) into a byte. That byte is used to update the dstSelIndexth byte of oldVdst, which is returned in its entirity.

cvt_sr_bf8_f32()

Return op name rocdl.cvt.sr.bf8.f32 as a bitstring.

cvt_sr_bf8_f32(ssa)

rocdl.cvt.sr.bf8.f32 - Convert f32 to bf8, stochiastic rounding

Attributes

  • byteSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • srcA - Single, F32, 32-bit float
  • srcB - Single, I32, 32-bit signless integer
  • old - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Convert srcA to bf8, adding the rounding factor from srcB, and store into the byteSelth byte of old, preserving the others.

Example:

// Stochastic rounding convert f32 to bf8 in byte 2 of old.
%0 = rocdl.cvt.sr.bf8.f32 %val, %stoch -> %old[2] : i32

cvt_sr_fp8_f32()

Return op name rocdl.cvt.sr.fp8.f32 as a bitstring.

cvt_sr_fp8_f32(ssa)

rocdl.cvt.sr.fp8.f32 - Convert f32 to fp8, stochiastic rounding

Attributes

  • byteSel - Single, I32Attr, 32-bit signless integer attribute

Operands

  • srcA - Single, F32, 32-bit float
  • srcB - Single, I32, 32-bit signless integer
  • old - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Convert srcA to fp8, adding the rounding factor from srcB, and store into the byteSelth byte of old, preserving the others.

Example:

// Stochastic rounding convert f32 to fp8 in byte 3 of old.
%0 = rocdl.cvt.sr.fp8.f32 %val, %stoch -> %old[3] : i32

dot4_f32_bf8_bf8()

Return op name rocdl.dot4.f32.bf8.bf8 as a bitstring.

dot4_f32_bf8_bf8(ssa)

rocdl.dot4.f32.bf8.bf8

Operands

  • a - Single, anonymous/composite constraint, 32-bit signless integer
  • b - Single, anonymous/composite constraint, 32-bit signless integer
  • c - Single, anonymous/composite constraint, 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float

Description

Packed intra-lane dot-product with no clamp control. Computes res = sum_i a[i]*b[i] + c. Covers the full-f16/bf16 accumulator forms (fdot2.f16.f16, fdot2.bf16.bf16) and the FP8/BF8 dot4.f32.* variants, whose hardware instructions have no CLAMP bit in their modifier word.

Example:

%r = rocdl.dot4.f32.bf8.bf8 %a, %b, %c : (i32, i32, f32) -> f32

dot4_f32_bf8_fp8()

Return op name rocdl.dot4.f32.bf8.fp8 as a bitstring.

dot4_f32_bf8_fp8(ssa)

rocdl.dot4.f32.bf8.fp8

Operands

  • a - Single, anonymous/composite constraint, 32-bit signless integer
  • b - Single, anonymous/composite constraint, 32-bit signless integer
  • c - Single, anonymous/composite constraint, 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float

Description

Packed intra-lane dot-product with no clamp control. Computes res = sum_i a[i]*b[i] + c. Covers the full-f16/bf16 accumulator forms (fdot2.f16.f16, fdot2.bf16.bf16) and the FP8/BF8 dot4.f32.* variants, whose hardware instructions have no CLAMP bit in their modifier word.

Example:

%r = rocdl.dot4.f32.bf8.fp8 %a, %b, %c : (i32, i32, f32) -> f32

dot4_f32_fp8_bf8()

Return op name rocdl.dot4.f32.fp8.bf8 as a bitstring.

dot4_f32_fp8_bf8(ssa)

rocdl.dot4.f32.fp8.bf8

Operands

  • a - Single, anonymous/composite constraint, 32-bit signless integer
  • b - Single, anonymous/composite constraint, 32-bit signless integer
  • c - Single, anonymous/composite constraint, 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float

Description

Packed intra-lane dot-product with no clamp control. Computes res = sum_i a[i]*b[i] + c. Covers the full-f16/bf16 accumulator forms (fdot2.f16.f16, fdot2.bf16.bf16) and the FP8/BF8 dot4.f32.* variants, whose hardware instructions have no CLAMP bit in their modifier word.

Example:

%r = rocdl.dot4.f32.fp8.bf8 %a, %b, %c : (i32, i32, f32) -> f32

dot4_f32_fp8_fp8()

Return op name rocdl.dot4.f32.fp8.fp8 as a bitstring.

dot4_f32_fp8_fp8(ssa)

rocdl.dot4.f32.fp8.fp8

Operands

  • a - Single, anonymous/composite constraint, 32-bit signless integer
  • b - Single, anonymous/composite constraint, 32-bit signless integer
  • c - Single, anonymous/composite constraint, 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float

Description

Packed intra-lane dot-product with no clamp control. Computes res = sum_i a[i]*b[i] + c. Covers the full-f16/bf16 accumulator forms (fdot2.f16.f16, fdot2.bf16.bf16) and the FP8/BF8 dot4.f32.* variants, whose hardware instructions have no CLAMP bit in their modifier word.

Example:

%r = rocdl.dot4.f32.fp8.fp8 %a, %b, %c : (i32, i32, f32) -> f32

ds_atomic_async_barrier_arrive_b64()

Return op name rocdl.ds.atomic.async.barrier.arrive.b64 as a bitstring.

ds_atomic_async_barrier_arrive_b64(ssa)

rocdl.ds.atomic.async.barrier.arrive.b64

Attributes

  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • barrierPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Description

Waits on a given DS barrier and decrements pending count by -1. Stays in order with ASYNC loads to LDS, and uses ASYNCcnt to track its completion. Available on gfx1250+.

Example:

// Async atomic barrier arrive (fire-and-forget).
rocdl.ds.atomic.async.barrier.arrive.b64 %ptr : !llvm.ptr<3>

ds_atomic_barrier_arrive_rtn_b64()

Return op name rocdl.ds.atomic.barrier.arrive.rtn.b64 as a bitstring.

ds_atomic_barrier_arrive_rtn_b64(ssa)

rocdl.ds.atomic.barrier.arrive.rtn.b64

Attributes

  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • barrierPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3
  • val - Single, I64, 64-bit signless integer

Results

  • res - Single, I64, 64-bit signless integer

Description

Waits on a given DS barrier and decrements its pending count by a given value. Note, the barrier state is given as a 64-bit structure containing pending count, phase and init count. The op returns the old barrier state. The op is executed as an ordinary LDS operations and it is ordered with other LDS operations. Thus, check DSCNT to determine when this instruction has executed. Available on gfx1250+.

Example:

// Atomic barrier arrive with return of old barrier state.
%res = rocdl.ds.atomic.barrier.arrive.rtn.b64 %ptr, %val : !llvm.ptr<3>, i64 -> i64

ds_bpermute()

Return op name rocdl.ds_bpermute as a bitstring.

ds_bpermute(ssa)

rocdl.ds_bpermute

Operands

  • index - Single, I32, 32-bit signless integer
  • src - Single, I32, 32-bit signless integer

Results

  • res - Single, I32, 32-bit signless integer

Description

Perform a backward permute (pull) operation across lanes using DS/LDS permute hardware.

Each lane reads the value of src from the lane whose byte address is given by index (i.e. lane id = index / 4).

This is “backward” (pull) in contrast to ds_permute_b32, which is “forward” (push/scatter).

Example:

// Backward permute across lanes (pull from selected lane).
%0 = rocdl.ds_bpermute %index, %src : (i32, i32) -> i32

ds_load_tr4_b64()

Return op name rocdl.ds.load.tr4.b64 as a bitstring.

ds_load_tr4_b64(ssa)

rocdl.ds.load.tr4.b64 - Loads and transposes a matrix from ds memory to registers (available in gfx1250+).

Attributes

  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • ptr - Single, anonymous/composite constraint, LLVM pointer in address space 3

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Load a matrix of 4-bit data from the ds memory, transpose data between row-major and column-major order, and store the result into a 64-bit vector register.

Available in gfx1250+.

Example (concrete mnemonics depend on address space and element size):

// 64-bit transpose load from global memory.
%0 = rocdl.global.load.tr4.b64 %ptr : !llvm.ptr<1> -> vector<2xi32>

// 128-bit transpose load from global memory with f16 result.
%1 = rocdl.global.load.tr.b128 %ptr : !llvm.ptr<1> -> vector<8xf16>

// 64-bit transpose load from LDS.
%2 = rocdl.ds.load.tr4.b64 %ptr : !llvm.ptr<3> -> vector<2xi32>

// 128-bit transpose load from LDS with bf16 result.
%3 = rocdl.ds.load.tr16.b128 %ptr : !llvm.ptr<3> -> vector<8xbf16>

ds_load_tr6_b96()

Return op name rocdl.ds.load.tr6.b96 as a bitstring.

ds_load_tr6_b96(ssa)

rocdl.ds.load.tr6.b96 - Loads and transposes a matrix from ds memory to registers (available in gfx1250+).

Attributes

  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • ptr - Single, anonymous/composite constraint, LLVM pointer in address space 3

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Load a matrix of 6-bit data from the ds memory, transpose data between row-major and column-major order, and store the result into a 96-bit vector register.

Available in gfx1250+.

Example (concrete mnemonics depend on address space and element size):

// 64-bit transpose load from global memory.
%0 = rocdl.global.load.tr4.b64 %ptr : !llvm.ptr<1> -> vector<2xi32>

// 128-bit transpose load from global memory with f16 result.
%1 = rocdl.global.load.tr.b128 %ptr : !llvm.ptr<1> -> vector<8xf16>

// 64-bit transpose load from LDS.
%2 = rocdl.ds.load.tr4.b64 %ptr : !llvm.ptr<3> -> vector<2xi32>

// 128-bit transpose load from LDS with bf16 result.
%3 = rocdl.ds.load.tr16.b128 %ptr : !llvm.ptr<3> -> vector<8xbf16>

ds_load_tr8_b64()

Return op name rocdl.ds.load.tr8.b64 as a bitstring.

ds_load_tr8_b64(ssa)

rocdl.ds.load.tr8.b64 - Loads and transposes a matrix from ds memory to registers (available in gfx1250+).

Attributes

  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • ptr - Single, anonymous/composite constraint, LLVM pointer in address space 3

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Load a matrix of 8-bit data from the ds memory, transpose data between row-major and column-major order, and store the result into a 64-bit vector register.

Available in gfx1250+.

Example (concrete mnemonics depend on address space and element size):

// 64-bit transpose load from global memory.
%0 = rocdl.global.load.tr4.b64 %ptr : !llvm.ptr<1> -> vector<2xi32>

// 128-bit transpose load from global memory with f16 result.
%1 = rocdl.global.load.tr.b128 %ptr : !llvm.ptr<1> -> vector<8xf16>

// 64-bit transpose load from LDS.
%2 = rocdl.ds.load.tr4.b64 %ptr : !llvm.ptr<3> -> vector<2xi32>

// 128-bit transpose load from LDS with bf16 result.
%3 = rocdl.ds.load.tr16.b128 %ptr : !llvm.ptr<3> -> vector<8xbf16>

ds_load_tr16_b128()

Return op name rocdl.ds.load.tr16.b128 as a bitstring.

ds_load_tr16_b128(ssa)

rocdl.ds.load.tr16.b128 - Loads and transposes a matrix from ds memory to registers (available in gfx1250+).

Attributes

  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • ptr - Single, anonymous/composite constraint, LLVM pointer in address space 3

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Load a matrix of 16-bit data from the ds memory, transpose data between row-major and column-major order, and store the result into a 128-bit vector register.

Available in gfx1250+.

Example (concrete mnemonics depend on address space and element size):

// 64-bit transpose load from global memory.
%0 = rocdl.global.load.tr4.b64 %ptr : !llvm.ptr<1> -> vector<2xi32>

// 128-bit transpose load from global memory with f16 result.
%1 = rocdl.global.load.tr.b128 %ptr : !llvm.ptr<1> -> vector<8xf16>

// 64-bit transpose load from LDS.
%2 = rocdl.ds.load.tr4.b64 %ptr : !llvm.ptr<3> -> vector<2xi32>

// 128-bit transpose load from LDS with bf16 result.
%3 = rocdl.ds.load.tr16.b128 %ptr : !llvm.ptr<3> -> vector<8xbf16>

ds_read_tr4_b64()

Return op name rocdl.ds.read.tr4.b64 as a bitstring.

ds_read_tr4_b64(ssa)

rocdl.ds.read.tr4.b64

Attributes

  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • ptr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

ds_read_tr6_b96()

Return op name rocdl.ds.read.tr6.b96 as a bitstring.

ds_read_tr6_b96(ssa)

rocdl.ds.read.tr6.b96

Attributes

  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • ptr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

ds_read_tr8_b64()

Return op name rocdl.ds.read.tr8.b64 as a bitstring.

ds_read_tr8_b64(ssa)

rocdl.ds.read.tr8.b64

Attributes

  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • ptr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

ds_read_tr16_b64()

Return op name rocdl.ds.read.tr16.b64 as a bitstring.

ds_read_tr16_b64(ssa)

rocdl.ds.read.tr16.b64

Attributes

  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • ptr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

ds_swizzle()

Return op name rocdl.ds_swizzle as a bitstring.

ds_swizzle(ssa)

rocdl.ds_swizzle

Operands

  • src - Single, I32, 32-bit signless integer
  • offset - Single, I32, 32-bit signless integer

Results

  • res - Single, I32, 32-bit signless integer

Description

Perform a data-sharing swizzle operation within a wavefront.

The offset operand encodes the swizzle pattern that will be placed in the instruction's offset field (i.e., the pattern used by ds_swizzle_b32). See https://llvm.org/docs/AMDGPUModifierSyntax.html#swizzle-pattern for how this 16-bit pattern is constructed.

Example:

// Swizzle data within a wavefront.
%0 = rocdl.ds_swizzle %src, %offset : (i32, i32) -> i32

exp2()

Return op name rocdl.exp2 as a bitstring.

exp2(ssa)

rocdl.exp2

Operands

  • arg - Single, LLVM_AnyFloat, floating point LLVM type

Results

  • res - Single, LLVM_AnyFloat, floating point LLVM type

Description

Note: In the general case, prefer the conventional arith, math, or llvm ops over this. Use this ROCDL-specific operation only when you fully understand its implication and when it is strictly necessary. This op is usually chosen when a small loss in precision is acceptable in exchange for higher execution speed.

Example:

%0 = rocdl.exp2 %a f32 -> f32

exp()

Return op name rocdl.exp as a bitstring.

exp(ssa)

rocdl.exp

Operands

  • arg - Single, LLVM_AnyFloat, floating point LLVM type

Results

  • res - Single, LLVM_AnyFloat, floating point LLVM type

Description

Note: In the general case, prefer the conventional arith, math, or llvm ops over this. Use this ROCDL-specific operation only when you fully understand its implication and when it is strictly necessary. This op is usually chosen when a small loss in precision is acceptable in exchange for higher execution speed.

Example:

%0 = rocdl.exp %a f32 -> f32

fdot2()

Return op name rocdl.fdot2 as a bitstring.

fdot2(ssa)

rocdl.fdot2

Attributes

  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2
  • b - Single, ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2
  • c - Single, anonymous/composite constraint, 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float

Description

Packed intra-lane dot-product with optional result clamping (clamp). Computes res = sum_i a[i]*b[i] + c, where a and b hold packed 4/8/16-bit data (for dot2,dot4,dot8).

Example:

%r = rocdl.fdot2 %a, %b, %c {clamp = true} :
     (vector<2xf16>, vector<2xf16>, f32) -> f32

fdot2_bf16_bf16()

Return op name rocdl.fdot2.bf16.bf16 as a bitstring.

fdot2_bf16_bf16(ssa)

rocdl.fdot2.bf16.bf16

Operands

  • a - Single, ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2
  • b - Single, ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2
  • c - Single, anonymous/composite constraint, bfloat16 type

Results

  • res - Single, anonymous/composite constraint, bfloat16 type

Description

Packed intra-lane dot-product with no clamp control. Computes res = sum_i a[i]*b[i] + c. Covers the full-f16/bf16 accumulator forms (fdot2.f16.f16, fdot2.bf16.bf16) and the FP8/BF8 dot4.f32.* variants, whose hardware instructions have no CLAMP bit in their modifier word.

Example:

%r = rocdl.fdot2.bf16.bf16 %a, %b, %c : (vector<2xbf16>, vector<2xbf16>, bf16) -> bf16

fdot2_f16_f16()

Return op name rocdl.fdot2.f16.f16 as a bitstring.

fdot2_f16_f16(ssa)

rocdl.fdot2.f16.f16

Operands

  • a - Single, ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2
  • b - Single, ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2
  • c - Single, anonymous/composite constraint, 16-bit float

Results

  • res - Single, anonymous/composite constraint, 16-bit float

Description

Packed intra-lane dot-product with no clamp control. Computes res = sum_i a[i]*b[i] + c. Covers the full-f16/bf16 accumulator forms (fdot2.f16.f16, fdot2.bf16.bf16) and the FP8/BF8 dot4.f32.* variants, whose hardware instructions have no CLAMP bit in their modifier word.

Example:

%r = rocdl.fdot2.f16.f16 %a, %b, %c : (vector<2xf16>, vector<2xf16>, f16) -> f16

fdot2_f32_bf16()

Return op name rocdl.fdot2.f32.bf16 as a bitstring.

fdot2_f32_bf16(ssa)

rocdl.fdot2.f32.bf16

Attributes

  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2
  • b - Single, ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2
  • c - Single, anonymous/composite constraint, 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float

Description

Packed intra-lane dot-product with optional result clamping (clamp). Computes res = sum_i a[i]*b[i] + c, where a and b hold packed 4/8/16-bit data (for dot2,dot4,dot8).

Example:

%r = rocdl.fdot2.f32.bf16 %a, %b, %c {clamp = true} :
     (vector<2xbf16>, vector<2xbf16>, f32) -> f32

flat_prefetch()

Return op name rocdl.flat.prefetch as a bitstring.

flat_prefetch(ssa)

rocdl.flat.prefetch

Attributes

  • cachePolicy - Single, ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • ptr - Single, anonymous/composite constraint, LLVM pointer in address space 0

Description

Prefetches 1 byte of data per lane using flat-memory addresses into the WGP-cache or L2-cache. Available on gfx1250+.

Example:

// Prefetch from flat memory into cache.
rocdl.flat.prefetch %ptr, 0 : !llvm.ptr

fmed3()

Return op name rocdl.fmed3 as a bitstring.

fmed3(ssa)

rocdl.fmed3 - Median of three float/half values

Operands

  • src0 - Single, anonymous/composite constraint, floating point LLVM type or LLVM dialect-compatible vector of floating point LLVM type
  • src1 - Single, anonymous/composite constraint, floating point LLVM type or LLVM dialect-compatible vector of floating point LLVM type
  • src2 - Single, anonymous/composite constraint, floating point LLVM type or LLVM dialect-compatible vector of floating point LLVM type

Results

  • res - Single, anonymous/composite constraint, floating point LLVM type or LLVM dialect-compatible vector of floating point LLVM type

Description

Computes the median of three floating-point values using the AMDGPU fmed3 intrinsic. This operation is equivalent to max(min(a, b), min(max(a, b), c)) but uses the hardware-accelerated V_MED3_F16/V_MED3_F32 instruction for better performance.

The operation supports both scalar and vector floating-point types (f16, f32).

Example:

// Scalar f32 median
%result = rocdl.fmed3 %a, %b, %c : f32

// Vector f16 median
%result = rocdl.fmed3 %va, %vb, %vc : vector<4xf16>

global_load_async_lds()

Return op name rocdl.global.load.async.lds as a bitstring.

global_load_async_lds(ssa)

rocdl.global.load.async.lds - Version of rocdl.load.async.to.lds specialized to global pointers

Attributes

  • size - Single, I32Attr, 32-bit signless integer attribute
  • offset - Single, I32Attr, 32-bit signless integer attribute
  • aux - Single, ROCDL_PreGfx12CachePolicyCompatAttr, pre-gfx12 or gfx942 AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • globalPtr - Single, ROCDLGlobalBuffer, LLVM pointer in address space 1
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Description

This operation works identically to rocdl.load.async.to.lds except that the global pointer argument is limited to pointers in address space 1 (pure global pointers) instead of also allowing fat buffer pointers.

Available on gfx9 and gfx10.

For the operation introduced in gfx1250, see rocdl.global.load.async.to.lds.bN.

Example:

// Async load from global pointer to LDS (address space 1 only).
rocdl.load.async.to.lds %global, %shared, 4, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>

global_load_async_to_lds_b8()

Return op name rocdl.global.load.async.to.lds.b8 as a bitstring.

global_load_async_to_lds_b8(ssa)

rocdl.global.load.async.to.lds.b8

Attributes

  • offset - Single, I32Attr, 32-bit signless integer attribute
  • aux - Single, ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • globalPtr - Single, ROCDLGlobalBuffer, LLVM pointer in address space 1
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Description

Asynchronously loads 8 bits of data from a global memory pointer to a Local Data Share (LDS) pointer.

Available on gfx1250+.

Example:

// Async 8-bit load from global to LDS.
rocdl.global.load.async.to.lds.b8 %src, %dst, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>

global_load_async_to_lds_b32()

Return op name rocdl.global.load.async.to.lds.b32 as a bitstring.

global_load_async_to_lds_b32(ssa)

rocdl.global.load.async.to.lds.b32

Attributes

  • offset - Single, I32Attr, 32-bit signless integer attribute
  • aux - Single, ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • globalPtr - Single, ROCDLGlobalBuffer, LLVM pointer in address space 1
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Description

Asynchronously loads 32 bits of data from a global memory pointer to a Local Data Share (LDS) pointer.

Available on gfx1250+.

Example:

// Async 32-bit load from global to LDS.
rocdl.global.load.async.to.lds.b32 %src, %dst, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>

global_load_async_to_lds_b64()

Return op name rocdl.global.load.async.to.lds.b64 as a bitstring.

global_load_async_to_lds_b64(ssa)

rocdl.global.load.async.to.lds.b64

Attributes

  • offset - Single, I32Attr, 32-bit signless integer attribute
  • aux - Single, ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • globalPtr - Single, ROCDLGlobalBuffer, LLVM pointer in address space 1
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Description

Asynchronously loads 64 bits of data from a global memory pointer to a Local Data Share (LDS) pointer.

Available on gfx1250+.

Example:

// Async 64-bit load from global to LDS.
rocdl.global.load.async.to.lds.b64 %src, %dst, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>

global_load_async_to_lds_b128()

Return op name rocdl.global.load.async.to.lds.b128 as a bitstring.

global_load_async_to_lds_b128(ssa)

rocdl.global.load.async.to.lds.b128

Attributes

  • offset - Single, I32Attr, 32-bit signless integer attribute
  • aux - Single, ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • globalPtr - Single, ROCDLGlobalBuffer, LLVM pointer in address space 1
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Description

Asynchronously loads 128 bits of data from a global memory pointer to a Local Data Share (LDS) pointer.

Available on gfx1250+.

Example:

// Async 128-bit load from global to LDS.
rocdl.global.load.async.to.lds.b128 %src, %dst, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>

global_load_lds()

Return op name rocdl.global.load.lds as a bitstring.

global_load_lds(ssa)

rocdl.global.load.lds

Attributes

  • size - Single, I32Attr, 32-bit signless integer attribute
  • offset - Single, I32Attr, 32-bit signless integer attribute
  • aux - Single, ROCDL_PreGfx12CachePolicyCompatAttr, pre-gfx12 or gfx942 AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • globalPtr - Single, ROCDLGlobalBuffer, LLVM pointer in address space 1
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

global_load_tr4_b64()

Return op name rocdl.global.load.tr4.b64 as a bitstring.

global_load_tr4_b64(ssa)

rocdl.global.load.tr4.b64 - Loads and transposes a matrix from global memory to registers (available in gfx1250+).

Attributes

  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • ptr - Single, anonymous/composite constraint, LLVM pointer in address space 1

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Load a matrix of 4-bit data from the global memory, transpose data between row-major and column-major order, and store the result into a 64-bit vector register.

Available in gfx1250+.

Example (concrete mnemonics depend on address space and element size):

// 64-bit transpose load from global memory.
%0 = rocdl.global.load.tr4.b64 %ptr : !llvm.ptr<1> -> vector<2xi32>

// 128-bit transpose load from global memory with f16 result.
%1 = rocdl.global.load.tr.b128 %ptr : !llvm.ptr<1> -> vector<8xf16>

// 64-bit transpose load from LDS.
%2 = rocdl.ds.load.tr4.b64 %ptr : !llvm.ptr<3> -> vector<2xi32>

// 128-bit transpose load from LDS with bf16 result.
%3 = rocdl.ds.load.tr16.b128 %ptr : !llvm.ptr<3> -> vector<8xbf16>

global_load_tr6_b96()

Return op name rocdl.global.load.tr6.b96 as a bitstring.

global_load_tr6_b96(ssa)

rocdl.global.load.tr6.b96 - Loads and transposes a matrix from global memory to registers (available in gfx1250+).

Attributes

  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • ptr - Single, anonymous/composite constraint, LLVM pointer in address space 1

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Load a matrix of 6-bit data from the global memory, transpose data between row-major and column-major order, and store the result into a 96-bit vector register.

Available in gfx1250+.

Example (concrete mnemonics depend on address space and element size):

// 64-bit transpose load from global memory.
%0 = rocdl.global.load.tr4.b64 %ptr : !llvm.ptr<1> -> vector<2xi32>

// 128-bit transpose load from global memory with f16 result.
%1 = rocdl.global.load.tr.b128 %ptr : !llvm.ptr<1> -> vector<8xf16>

// 64-bit transpose load from LDS.
%2 = rocdl.ds.load.tr4.b64 %ptr : !llvm.ptr<3> -> vector<2xi32>

// 128-bit transpose load from LDS with bf16 result.
%3 = rocdl.ds.load.tr16.b128 %ptr : !llvm.ptr<3> -> vector<8xbf16>

global_load_tr_b64()

Return op name rocdl.global.load.tr.b64 as a bitstring.

global_load_tr_b64(ssa)

rocdl.global.load.tr.b64 - Loads and transposes a matrix from global memory to registers (available in gfx1250+).

Attributes

  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • ptr - Single, anonymous/composite constraint, LLVM pointer in address space 1

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Load a matrix of 8-bit data from the global memory, transpose data between row-major and column-major order, and store the result into a 64-bit vector register.

Available in gfx1250+.

Example (concrete mnemonics depend on address space and element size):

// 64-bit transpose load from global memory.
%0 = rocdl.global.load.tr4.b64 %ptr : !llvm.ptr<1> -> vector<2xi32>

// 128-bit transpose load from global memory with f16 result.
%1 = rocdl.global.load.tr.b128 %ptr : !llvm.ptr<1> -> vector<8xf16>

// 64-bit transpose load from LDS.
%2 = rocdl.ds.load.tr4.b64 %ptr : !llvm.ptr<3> -> vector<2xi32>

// 128-bit transpose load from LDS with bf16 result.
%3 = rocdl.ds.load.tr16.b128 %ptr : !llvm.ptr<3> -> vector<8xbf16>

global_load_tr_b128()

Return op name rocdl.global.load.tr.b128 as a bitstring.

global_load_tr_b128(ssa)

rocdl.global.load.tr.b128 - Loads and transposes a matrix from global memory to registers (available in gfx1250+).

Attributes

  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • ptr - Single, anonymous/composite constraint, LLVM pointer in address space 1

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Load a matrix of 16-bit data from the global memory, transpose data between row-major and column-major order, and store the result into a 128-bit vector register.

Available in gfx1250+.

Example (concrete mnemonics depend on address space and element size):

// 64-bit transpose load from global memory.
%0 = rocdl.global.load.tr4.b64 %ptr : !llvm.ptr<1> -> vector<2xi32>

// 128-bit transpose load from global memory with f16 result.
%1 = rocdl.global.load.tr.b128 %ptr : !llvm.ptr<1> -> vector<8xf16>

// 64-bit transpose load from LDS.
%2 = rocdl.ds.load.tr4.b64 %ptr : !llvm.ptr<3> -> vector<2xi32>

// 128-bit transpose load from LDS with bf16 result.
%3 = rocdl.ds.load.tr16.b128 %ptr : !llvm.ptr<3> -> vector<8xbf16>

global_prefetch()

Return op name rocdl.global.prefetch as a bitstring.

global_prefetch(ssa)

rocdl.global.prefetch

Attributes

  • cachePolicy - Single, ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • ptr - Single, anonymous/composite constraint, LLVM pointer in address space 1

Description

Prefetches 1 byte of data per lane from global memory into the WGP-cache or L2-cache. Available on gfx1250+.

Example:

// Prefetch from global memory into cache.
rocdl.global.prefetch %ptr, 0 : !llvm.ptr<1>

global_store_async_from_lds_b8()

Return op name rocdl.global.store.async.from.lds.b8 as a bitstring.

global_store_async_from_lds_b8(ssa)

rocdl.global.store.async.from.lds.b8

Attributes

  • offset - Single, I32Attr, 32-bit signless integer attribute
  • aux - Single, ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • globalPtr - Single, ROCDLGlobalBuffer, LLVM pointer in address space 1
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Description

Asynchronously stores 8 bits of data from a Local Data Share (LDS) pointer to a global memory pointer.

Available on gfx1250+.

Example:

// Async 8-bit store from LDS to global.
rocdl.global.store.async.from.lds.b8 %dst, %src, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>

global_store_async_from_lds_b32()

Return op name rocdl.global.store.async.from.lds.b32 as a bitstring.

global_store_async_from_lds_b32(ssa)

rocdl.global.store.async.from.lds.b32

Attributes

  • offset - Single, I32Attr, 32-bit signless integer attribute
  • aux - Single, ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • globalPtr - Single, ROCDLGlobalBuffer, LLVM pointer in address space 1
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Description

Asynchronously stores 32 bits of data from a Local Data Share (LDS) pointer to a global memory pointer.

Available on gfx1250+.

Example:

// Async 32-bit store from LDS to global.
rocdl.global.store.async.from.lds.b32 %dst, %src, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>

global_store_async_from_lds_b64()

Return op name rocdl.global.store.async.from.lds.b64 as a bitstring.

global_store_async_from_lds_b64(ssa)

rocdl.global.store.async.from.lds.b64

Attributes

  • offset - Single, I32Attr, 32-bit signless integer attribute
  • aux - Single, ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • globalPtr - Single, ROCDLGlobalBuffer, LLVM pointer in address space 1
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Description

Asynchronously stores 64 bits of data from a Local Data Share (LDS) pointer to a global memory pointer.

Available on gfx1250+.

Example:

// Async 64-bit store from LDS to global.
rocdl.global.store.async.from.lds.b64 %dst, %src, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>

global_store_async_from_lds_b128()

Return op name rocdl.global.store.async.from.lds.b128 as a bitstring.

global_store_async_from_lds_b128(ssa)

rocdl.global.store.async.from.lds.b128

Attributes

  • offset - Single, I32Attr, 32-bit signless integer attribute
  • aux - Single, ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • globalPtr - Single, ROCDLGlobalBuffer, LLVM pointer in address space 1
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Description

Asynchronously stores 128 bits of data from a Local Data Share (LDS) pointer to a global memory pointer.

Available on gfx1250+.

Example:

// Async 128-bit store from LDS to global.
rocdl.global.store.async.from.lds.b128 %dst, %src, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>

iglp_opt()

Return op name rocdl.iglp.opt as a bitstring.

iglp_opt(ssa)

rocdl.iglp.opt

Attributes

  • variant - Single, I32Attr, 32-bit signless integer attribute

Description

Instruction-group-level parallelism optimization hint.

Example:

// IGLP optimization hint variant 0.
rocdl.iglp.opt 0

load_async_to_lds()

Return op name rocdl.load.async.to.lds as a bitstring.

load_async_to_lds(ssa)

rocdl.load.async.to.lds - Gathering load to LDS that requires explicit async memory tracking

Attributes

  • size - Single, I32Attr, 32-bit signless integer attribute
  • offset - Single, I32Attr, 32-bit signless integer attribute
  • aux - Single, ROCDL_PreGfx12CachePolicyCompatAttr, pre-gfx12 or gfx942 AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • globalPtr - Single, LLVM_AnyPointer, LLVM pointer type
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Description

Load size bytes (the valid sizes vary by architecture) from the global memory pointed to by globalPtr and put them at ldsPtr, concantenating (and applying padding for sizes less than 4 bytes, along with padding out 12-byte reads to 16-byte writes). The value of globalPtr can vary between lanes, while sharedPtr must be subgroup-uniform (the values from each lane are concatentated before being written to LDS with appropriate padding applied.)

offset is a constant offset applied to both pointers, and aux sets the cache policy. Unlike rocdl.load.to.lds, the compiler will not automatically inserts waits for this load to complete at the point it thinks you're using a region of LDS you've stored values to - you need to use the rocdl.asyncmark and rocdl.wait.asyncmark operations to explicitly group these operations and wait for their completion.

Available on gfx10 and earlier with varying suppported values of size.

Example:

// Async load 4 bytes from global pointer to LDS.
rocdl.load.async.to.lds %global, %shared, 4, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>

// Async load 4 bytes from fat buffer pointer to LDS.
rocdl.load.async.to.lds %fatBuffer, %shared, 4, 0, 0 : !llvm.ptr<7>, !llvm.ptr<3>

load_to_lds()

Return op name rocdl.load.to.lds as a bitstring.

load_to_lds(ssa)

rocdl.load.to.lds

Attributes

  • size - Single, I32Attr, 32-bit signless integer attribute
  • offset - Single, I32Attr, 32-bit signless integer attribute
  • aux - Single, ROCDL_PreGfx12CachePolicyCompatAttr, pre-gfx12 or gfx942 AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • globalPtr - Single, LLVM_AnyPointer, LLVM pointer type
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

log()

Return op name rocdl.log as a bitstring.

log(ssa)

rocdl.log

Operands

  • arg - Single, LLVM_AnyFloat, floating point LLVM type

Results

  • res - Single, LLVM_AnyFloat, floating point LLVM type

Description

Note: In the general case, prefer the conventional arith, math, or llvm ops over this. Use this ROCDL-specific operation only when you fully understand its implication and when it is strictly necessary. This op is usually chosen when a small loss in precision is acceptable in exchange for higher execution speed.

Example:

%0 = rocdl.log %a f32 -> f32

make_buffer_rsrc()

Return op name rocdl.make.buffer.rsrc as a bitstring.

make_buffer_rsrc(ssa)

rocdl.make.buffer.rsrc

Operands

  • base - Single, LLVM_AnyPointer, LLVM pointer type
  • stride - Single, I16, 16-bit signless integer
  • numRecords - Single, I64, 64-bit signless integer
  • flags - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_AnyPointer, LLVM pointer type

mbcnt_hi()

Return op name rocdl.mbcnt.hi as a bitstring.

mbcnt_hi(ssa)

rocdl.mbcnt.hi

Attributes

  • arg_attrs - Optional, DictArrayAttr, Array of dictionary attributes
  • res_attrs - Optional, DictArrayAttr, Array of dictionary attributes

Operands

  • in0 - Single, I32, 32-bit signless integer
  • in1 - Single, I32, 32-bit signless integer

Results

  • res - Single, I32, 32-bit signless integer

Description

Masked bit count of threads below the current lane in a wavefront.

in0 is a 32-bit mask that is AND-ed with the relevant half of the execution mask and the bits below the current lane; in1 is added to the resulting popcount:

  • lo: in1 + popcount(in0 & exec_lo & ((1 << min(lane_id, 32)) - 1))
  • hi: in1 + popcount(in0 & exec_hi & ((1 << saturating_usub(lane_id, 32)) - 1))

To obtain a unique thread index within a wave64, chain the two ops with in0 = -1 (all bits set):

Example:

%all_ones = arith.constant -1 : i32
%zero = arith.constant 0 : i32

// Count active threads below this lane in the low 32 lanes.
%lo = rocdl.mbcnt.lo %all_ones, %zero : (i32, i32) -> i32

// Add the count from the high 32 lanes to get the full lane index.
%hi = rocdl.mbcnt.hi %all_ones, %lo : (i32, i32) -> i32

mbcnt_lo()

Return op name rocdl.mbcnt.lo as a bitstring.

mbcnt_lo(ssa)

rocdl.mbcnt.lo

Attributes

  • arg_attrs - Optional, DictArrayAttr, Array of dictionary attributes
  • res_attrs - Optional, DictArrayAttr, Array of dictionary attributes

Operands

  • in0 - Single, I32, 32-bit signless integer
  • in1 - Single, I32, 32-bit signless integer

Results

  • res - Single, I32, 32-bit signless integer

Description

Masked bit count of threads below the current lane in a wavefront.

in0 is a 32-bit mask that is AND-ed with the relevant half of the execution mask and the bits below the current lane; in1 is added to the resulting popcount:

  • lo: in1 + popcount(in0 & exec_lo & ((1 << min(lane_id, 32)) - 1))
  • hi: in1 + popcount(in0 & exec_hi & ((1 << saturating_usub(lane_id, 32)) - 1))

To obtain a unique thread index within a wave64, chain the two ops with in0 = -1 (all bits set):

Example:

%all_ones = arith.constant -1 : i32
%zero = arith.constant 0 : i32

// Count active threads below this lane in the low 32 lanes.
%lo = rocdl.mbcnt.lo %all_ones, %zero : (i32, i32) -> i32

// Add the count from the high 32 lanes to get the full lane index.
%hi = rocdl.mbcnt.hi %all_ones, %lo : (i32, i32) -> i32

mfma_f32_4x4x1f32()

Return op name rocdl.mfma.f32.4x4x1f32 as a bitstring.

mfma_f32_4x4x1f32(ssa)

rocdl.mfma.f32.4x4x1f32

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 32-bit float
  • b - Single, anonymous/composite constraint, 32-bit float
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.4x4x1f32 %a0, %b0, %c0, 0, 0, none : (f32, f32, vector<4xf32>) -> vector<4xf32>

mfma_f32_4x4x2bf16()

Return op name rocdl.mfma.f32.4x4x2bf16 as a bitstring.

mfma_f32_4x4x2bf16(ssa)

rocdl.mfma.f32.4x4x2bf16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.4x4x2bf16 %a0, %b0, %c0, 0, 0, none : (vector<2xi16>, vector<2xi16>, vector<4xf32>) -> vector<4xf32>

mfma_f32_4x4x4bf16_1k()

Return op name rocdl.mfma.f32.4x4x4bf16.1k as a bitstring.

mfma_f32_4x4x4bf16_1k(ssa)

rocdl.mfma.f32.4x4x4bf16.1k

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.4x4x4bf16.1k %a0, %b0, %c0, 0, 0, none : (vector<4xi16>, vector<4xi16>, vector<4xf32>) -> vector<4xf32>

mfma_f32_4x4x4f16()

Return op name rocdl.mfma.f32.4x4x4f16 as a bitstring.

mfma_f32_4x4x4f16(ssa)

rocdl.mfma.f32.4x4x4f16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.4x4x4f16 %a0, %b0, %c0, 0, 0, none : (vector<4xf16>, vector<4xf16>, vector<4xf32>) -> vector<4xf32>

mfma_f32_16x16x1f32()

Return op name rocdl.mfma.f32.16x16x1f32 as a bitstring.

mfma_f32_16x16x1f32(ssa)

rocdl.mfma.f32.16x16x1f32

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 32-bit float
  • b - Single, anonymous/composite constraint, 32-bit float
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.16x16x1f32 %a0, %b0, %c0, 0, 0, none : (f32, f32, vector<16xf32>) -> vector<16xf32>

mfma_f32_16x16x2bf16()

Return op name rocdl.mfma.f32.16x16x2bf16 as a bitstring.

mfma_f32_16x16x2bf16(ssa)

rocdl.mfma.f32.16x16x2bf16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.16x16x2bf16 %a0, %b0, %c0, 0, 0, none : (vector<2xi16>, vector<2xi16>, vector<16xf32>) -> vector<16xf32>

mfma_f32_16x16x4bf16_1k()

Return op name rocdl.mfma.f32.16x16x4bf16.1k as a bitstring.

mfma_f32_16x16x4bf16_1k(ssa)

rocdl.mfma.f32.16x16x4bf16.1k

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.16x16x4bf16.1k %a0, %b0, %c0, 0, 0, none : (vector<4xi16>, vector<4xi16>, vector<16xf32>) -> vector<16xf32>

mfma_f32_16x16x4f16()

Return op name rocdl.mfma.f32.16x16x4f16 as a bitstring.

mfma_f32_16x16x4f16(ssa)

rocdl.mfma.f32.16x16x4f16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.16x16x4f16 %a0, %b0, %c0, 0, 0, none : (vector<4xf16>, vector<4xf16>, vector<16xf32>) -> vector<16xf32>

mfma_f32_16x16x4f32()

Return op name rocdl.mfma.f32.16x16x4f32 as a bitstring.

mfma_f32_16x16x4f32(ssa)

rocdl.mfma.f32.16x16x4f32

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 32-bit float
  • b - Single, anonymous/composite constraint, 32-bit float
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.16x16x4f32 %a0, %b0, %c0, 0, 0, none : (f32, f32, vector<4xf32>) -> vector<4xf32>

mfma_f32_16x16x8_xf32()

Return op name rocdl.mfma.f32.16x16x8.xf32 as a bitstring.

mfma_f32_16x16x8_xf32(ssa)

rocdl.mfma.f32.16x16x8.xf32

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 2
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.16x16x8.xf32 %a0, %b0, %c0, 0, 0, none : (vector<2xf32>, vector<2xf32>, vector<4xf32>) -> vector<4xf32>

mfma_f32_16x16x8bf16()

Return op name rocdl.mfma.f32.16x16x8bf16 as a bitstring.

mfma_f32_16x16x8bf16(ssa)

rocdl.mfma.f32.16x16x8bf16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.16x16x8bf16 %a0, %b0, %c0, 0, 0, none : (vector<2xi16>, vector<2xi16>, vector<4xf32>) -> vector<4xf32>

mfma_f32_16x16x16bf16_1k()

Return op name rocdl.mfma.f32.16x16x16bf16.1k as a bitstring.

mfma_f32_16x16x16bf16_1k(ssa)

rocdl.mfma.f32.16x16x16bf16.1k

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.16x16x16bf16.1k %a0, %b0, %c0, 0, 0, none : (vector<4xi16>, vector<4xi16>, vector<4xf32>) -> vector<4xf32>

mfma_f32_16x16x16f16()

Return op name rocdl.mfma.f32.16x16x16f16 as a bitstring.

mfma_f32_16x16x16f16(ssa)

rocdl.mfma.f32.16x16x16f16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.16x16x16f16 %a0, %b0, %c0, 0, 0, none : (vector<4xf16>, vector<4xf16>, vector<4xf32>) -> vector<4xf32>

mfma_f32_16x16x32_bf8_bf8()

Return op name rocdl.mfma.f32.16x16x32.bf8.bf8 as a bitstring.

mfma_f32_16x16x32_bf8_bf8(ssa)

rocdl.mfma.f32.16x16x32.bf8.bf8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 64-bit signless integer
  • b - Single, anonymous/composite constraint, 64-bit signless integer
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.16x16x32.bf8.bf8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<4xf32>) -> vector<4xf32>

mfma_f32_16x16x32_bf8_fp8()

Return op name rocdl.mfma.f32.16x16x32.bf8.fp8 as a bitstring.

mfma_f32_16x16x32_bf8_fp8(ssa)

rocdl.mfma.f32.16x16x32.bf8.fp8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 64-bit signless integer
  • b - Single, anonymous/composite constraint, 64-bit signless integer
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.16x16x32.bf8.fp8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<4xf32>) -> vector<4xf32>

mfma_f32_16x16x32_bf16()

Return op name rocdl.mfma.f32.16x16x32.bf16 as a bitstring.

mfma_f32_16x16x32_bf16(ssa)

rocdl.mfma.f32.16x16x32.bf16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of bfloat16 type values of length 8
  • b - Single, anonymous/composite constraint, fixed-length vector of bfloat16 type values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.16x16x32.bf16 %a0, %b0, %c0, 0, 0, none : (vector<8xbf16>, vector<8xbf16>, vector<4xf32>) -> vector<4xf32>

mfma_f32_16x16x32_f16()

Return op name rocdl.mfma.f32.16x16x32.f16 as a bitstring.

mfma_f32_16x16x32_f16(ssa)

rocdl.mfma.f32.16x16x32.f16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 8
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.16x16x32.f16 %a0, %b0, %c0, 0, 0, none : (vector<8xf16>, vector<8xf16>, vector<4xf32>) -> vector<4xf32>

mfma_f32_16x16x32_fp8_bf8()

Return op name rocdl.mfma.f32.16x16x32.fp8.bf8 as a bitstring.

mfma_f32_16x16x32_fp8_bf8(ssa)

rocdl.mfma.f32.16x16x32.fp8.bf8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 64-bit signless integer
  • b - Single, anonymous/composite constraint, 64-bit signless integer
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.16x16x32.fp8.bf8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<4xf32>) -> vector<4xf32>

mfma_f32_16x16x32_fp8_fp8()

Return op name rocdl.mfma.f32.16x16x32.fp8.fp8 as a bitstring.

mfma_f32_16x16x32_fp8_fp8(ssa)

rocdl.mfma.f32.16x16x32.fp8.fp8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 64-bit signless integer
  • b - Single, anonymous/composite constraint, 64-bit signless integer
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.16x16x32.fp8.fp8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<4xf32>) -> vector<4xf32>

mfma_f32_32x32x1f32()

Return op name rocdl.mfma.f32.32x32x1f32 as a bitstring.

mfma_f32_32x32x1f32(ssa)

rocdl.mfma.f32.32x32x1f32

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 32-bit float
  • b - Single, anonymous/composite constraint, 32-bit float
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 32

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 32

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.32x32x1f32 %a0, %b0, %c0, 0, 0, none : (f32, f32, vector<32xf32>) -> vector<32xf32>

mfma_f32_32x32x2bf16()

Return op name rocdl.mfma.f32.32x32x2bf16 as a bitstring.

mfma_f32_32x32x2bf16(ssa)

rocdl.mfma.f32.32x32x2bf16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 32

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 32

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.32x32x2bf16 %a0, %b0, %c0, 0, 0, none : (vector<2xi16>, vector<2xi16>, vector<32xf32>) -> vector<32xf32>

mfma_f32_32x32x2f32()

Return op name rocdl.mfma.f32.32x32x2f32 as a bitstring.

mfma_f32_32x32x2f32(ssa)

rocdl.mfma.f32.32x32x2f32

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 32-bit float
  • b - Single, anonymous/composite constraint, 32-bit float
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.32x32x2f32 %a0, %b0, %c0, 0, 0, none : (f32, f32, vector<16xf32>) -> vector<16xf32>

mfma_f32_32x32x4_xf32()

Return op name rocdl.mfma.f32.32x32x4.xf32 as a bitstring.

mfma_f32_32x32x4_xf32(ssa)

rocdl.mfma.f32.32x32x4.xf32

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 2
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.32x32x4.xf32 %a0, %b0, %c0, 0, 0, none : (vector<2xf32>, vector<2xf32>, vector<16xf32>) -> vector<16xf32>

mfma_f32_32x32x4bf16()

Return op name rocdl.mfma.f32.32x32x4bf16 as a bitstring.

mfma_f32_32x32x4bf16(ssa)

rocdl.mfma.f32.32x32x4bf16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.32x32x4bf16 %a0, %b0, %c0, 0, 0, none : (vector<2xi16>, vector<2xi16>, vector<16xf32>) -> vector<16xf32>

mfma_f32_32x32x4bf16_1k()

Return op name rocdl.mfma.f32.32x32x4bf16.1k as a bitstring.

mfma_f32_32x32x4bf16_1k(ssa)

rocdl.mfma.f32.32x32x4bf16.1k

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 32

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 32

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.32x32x4bf16.1k %a0, %b0, %c0, 0, 0, none : (vector<4xi16>, vector<4xi16>, vector<32xf32>) -> vector<32xf32>

mfma_f32_32x32x4f16()

Return op name rocdl.mfma.f32.32x32x4f16 as a bitstring.

mfma_f32_32x32x4f16(ssa)

rocdl.mfma.f32.32x32x4f16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 32

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 32

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.32x32x4f16 %a0, %b0, %c0, 0, 0, none : (vector<4xf16>, vector<4xf16>, vector<32xf32>) -> vector<32xf32>

mfma_f32_32x32x8bf16_1k()

Return op name rocdl.mfma.f32.32x32x8bf16.1k as a bitstring.

mfma_f32_32x32x8bf16_1k(ssa)

rocdl.mfma.f32.32x32x8bf16.1k

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.32x32x8bf16.1k %a0, %b0, %c0, 0, 0, none : (vector<4xi16>, vector<4xi16>, vector<16xf32>) -> vector<16xf32>

mfma_f32_32x32x8f16()

Return op name rocdl.mfma.f32.32x32x8f16 as a bitstring.

mfma_f32_32x32x8f16(ssa)

rocdl.mfma.f32.32x32x8f16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.32x32x8f16 %a0, %b0, %c0, 0, 0, none : (vector<4xf16>, vector<4xf16>, vector<16xf32>) -> vector<16xf32>

mfma_f32_32x32x16_bf8_bf8()

Return op name rocdl.mfma.f32.32x32x16.bf8.bf8 as a bitstring.

mfma_f32_32x32x16_bf8_bf8(ssa)

rocdl.mfma.f32.32x32x16.bf8.bf8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 64-bit signless integer
  • b - Single, anonymous/composite constraint, 64-bit signless integer
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.32x32x16.bf8.bf8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<16xf32>) -> vector<16xf32>

mfma_f32_32x32x16_bf8_fp8()

Return op name rocdl.mfma.f32.32x32x16.bf8.fp8 as a bitstring.

mfma_f32_32x32x16_bf8_fp8(ssa)

rocdl.mfma.f32.32x32x16.bf8.fp8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 64-bit signless integer
  • b - Single, anonymous/composite constraint, 64-bit signless integer
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.32x32x16.bf8.fp8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<16xf32>) -> vector<16xf32>

mfma_f32_32x32x16_bf16()

Return op name rocdl.mfma.f32.32x32x16.bf16 as a bitstring.

mfma_f32_32x32x16_bf16(ssa)

rocdl.mfma.f32.32x32x16.bf16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of bfloat16 type values of length 8
  • b - Single, anonymous/composite constraint, fixed-length vector of bfloat16 type values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.32x32x16.bf16 %a0, %b0, %c0, 0, 0, none : (vector<8xbf16>, vector<8xbf16>, vector<16xf32>) -> vector<16xf32>

mfma_f32_32x32x16_f16()

Return op name rocdl.mfma.f32.32x32x16.f16 as a bitstring.

mfma_f32_32x32x16_f16(ssa)

rocdl.mfma.f32.32x32x16.f16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 8
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.32x32x16.f16 %a0, %b0, %c0, 0, 0, none : (vector<8xf16>, vector<8xf16>, vector<16xf32>) -> vector<16xf32>

mfma_f32_32x32x16_fp8_bf8()

Return op name rocdl.mfma.f32.32x32x16.fp8.bf8 as a bitstring.

mfma_f32_32x32x16_fp8_bf8(ssa)

rocdl.mfma.f32.32x32x16.fp8.bf8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 64-bit signless integer
  • b - Single, anonymous/composite constraint, 64-bit signless integer
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.32x32x16.fp8.bf8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<16xf32>) -> vector<16xf32>

mfma_f32_32x32x16_fp8_fp8()

Return op name rocdl.mfma.f32.32x32x16.fp8.fp8 as a bitstring.

mfma_f32_32x32x16_fp8_fp8(ssa)

rocdl.mfma.f32.32x32x16.fp8.fp8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 64-bit signless integer
  • b - Single, anonymous/composite constraint, 64-bit signless integer
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.f32.32x32x16.fp8.fp8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<16xf32>) -> vector<16xf32>

mfma_f64_4x4x4f64()

Return op name rocdl.mfma.f64.4x4x4f64 as a bitstring.

mfma_f64_4x4x4f64(ssa)

rocdl.mfma.f64.4x4x4f64

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMANegModifierAttr, negation modifier bitfield for gfx94x double-precision MFMA

Operands

  • a - Single, anonymous/composite constraint, 64-bit float
  • b - Single, anonymous/composite constraint, 64-bit float
  • c - Single, anonymous/composite constraint, 64-bit float

Results

  • res - Single, anonymous/composite constraint, 64-bit float

Description

Double-precision matrix fused multiply-add (MFMA) intrinsic. On gfx94x, the blgp immarg is a NEG bitfield rather than a B-lane permutation.

Example:

%r0 = mfma.f64.4x4x4f64 %a0, %b0, %c0, 0, 0, neg_a|neg_b : (f64, f64, f64) -> f64

mfma_f64_16x16x4f64()

Return op name rocdl.mfma.f64.16x16x4f64 as a bitstring.

mfma_f64_16x16x4f64(ssa)

rocdl.mfma.f64.16x16x4f64

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMANegModifierAttr, negation modifier bitfield for gfx94x double-precision MFMA

Operands

  • a - Single, anonymous/composite constraint, 64-bit float
  • b - Single, anonymous/composite constraint, 64-bit float
  • c - Single, anonymous/composite constraint, fixed-length vector of 64-bit float values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 64-bit float values of length 4

Description

Double-precision matrix fused multiply-add (MFMA) intrinsic. On gfx94x, the blgp immarg is a NEG bitfield rather than a B-lane permutation.

Example:

%r0 = mfma.f64.16x16x4f64 %a0, %b0, %c0, 0, 0, neg_a|neg_b : (f64, f64, vector<4xf64>) -> vector<4xf64>

mfma_i32_4x4x4i8()

Return op name rocdl.mfma.i32.4x4x4i8 as a bitstring.

mfma_i32_4x4x4i8(ssa)

rocdl.mfma.i32.4x4x4i8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 32-bit signless integer
  • b - Single, anonymous/composite constraint, 32-bit signless integer
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.i32.4x4x4i8 %a0, %b0, %c0, 0, 0, none : (i32, i32, vector<4xi32>) -> vector<4xi32>

mfma_i32_16x16x4i8()

Return op name rocdl.mfma.i32.16x16x4i8 as a bitstring.

mfma_i32_16x16x4i8(ssa)

rocdl.mfma.i32.16x16x4i8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 32-bit signless integer
  • b - Single, anonymous/composite constraint, 32-bit signless integer
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.i32.16x16x4i8 %a0, %b0, %c0, 0, 0, none : (i32, i32, vector<16xi32>) -> vector<16xi32>

mfma_i32_16x16x16i8()

Return op name rocdl.mfma.i32.16x16x16i8 as a bitstring.

mfma_i32_16x16x16i8(ssa)

rocdl.mfma.i32.16x16x16i8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 32-bit signless integer
  • b - Single, anonymous/composite constraint, 32-bit signless integer
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.i32.16x16x16i8 %a0, %b0, %c0, 0, 0, none : (i32, i32, vector<4xi32>) -> vector<4xi32>

mfma_i32_16x16x32_i8()

Return op name rocdl.mfma.i32.16x16x32.i8 as a bitstring.

mfma_i32_16x16x32_i8(ssa)

rocdl.mfma.i32.16x16x32.i8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 64-bit signless integer
  • b - Single, anonymous/composite constraint, 64-bit signless integer
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.i32.16x16x32.i8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<4xi32>) -> vector<4xi32>

mfma_i32_16x16x64_i8()

Return op name rocdl.mfma.i32.16x16x64.i8 as a bitstring.

mfma_i32_16x16x64_i8(ssa)

rocdl.mfma.i32.16x16x64.i8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.i32.16x16x64.i8 %a0, %b0, %c0, 0, 0, none : (vector<4xi32>, vector<4xi32>, vector<4xi32>) -> vector<4xi32>

mfma_i32_32x32x4i8()

Return op name rocdl.mfma.i32.32x32x4i8 as a bitstring.

mfma_i32_32x32x4i8(ssa)

rocdl.mfma.i32.32x32x4i8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 32-bit signless integer
  • b - Single, anonymous/composite constraint, 32-bit signless integer
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 32

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 32

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.i32.32x32x4i8 %a0, %b0, %c0, 0, 0, none : (i32, i32, vector<32xi32>) -> vector<32xi32>

mfma_i32_32x32x8i8()

Return op name rocdl.mfma.i32.32x32x8i8 as a bitstring.

mfma_i32_32x32x8i8(ssa)

rocdl.mfma.i32.32x32x8i8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 32-bit signless integer
  • b - Single, anonymous/composite constraint, 32-bit signless integer
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.i32.32x32x8i8 %a0, %b0, %c0, 0, 0, none : (i32, i32, vector<16xi32>) -> vector<16xi32>

mfma_i32_32x32x16_i8()

Return op name rocdl.mfma.i32.32x32x16.i8 as a bitstring.

mfma_i32_32x32x16_i8(ssa)

rocdl.mfma.i32.32x32x16.i8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, 64-bit signless integer
  • b - Single, anonymous/composite constraint, 64-bit signless integer
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.i32.32x32x16.i8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<16xi32>) -> vector<16xi32>

mfma_i32_32x32x32_i8()

Return op name rocdl.mfma.i32.32x32x32.i8 as a bitstring.

mfma_i32_32x32x32_i8(ssa)

rocdl.mfma.i32.32x32x32.i8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute
  • blgp - Single, ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16

Description

Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C with matrix operands. The cbsz, abid, and blgp attributes control broadcast and block layout modes.

Example:

%r0 = mfma.i32.32x32x32.i8 %a0, %b0, %c0, 0, 0, none : (vector<4xi32>, vector<4xi32>, vector<16xi32>) -> vector<16xi32>

mfma_scale_f32_16x16x128_f8f6f4()

Return op name rocdl.mfma.scale.f32.16x16x128.f8f6f4 as a bitstring.

mfma_scale_f32_16x16x128_f8f6f4(ssa)

rocdl.mfma.scale.f32.16x16x128.f8f6f4

Attributes

  • cbsz - Single, ROCDL_MatrixFormatAttr, matrix operand formats selected by scaled MFMA/WMMA format fields
  • blgp - Single, ROCDL_MatrixFormatAttr, matrix operand formats selected by scaled MFMA/WMMA format fields
  • opselA - Single, I32Attr, 32-bit signless integer attribute
  • opselB - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit signless integer
  • b - Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit signless integer
  • c - Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit float
  • scaleA - Single, I32, 32-bit signless integer
  • scaleB - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Scaled matrix fused multiply-add (MFMA) intrinsic with per-operand scaling. The opselA/opselB and scaleA/scaleB arguments control the scaling of input operands.

Example:

// Scaled MFMA with fp8 * fp8 inputs.
%r0 = rocdl.mfma.scale.f32.32x32x64.f8f6f4 %a, %a, %c, fp8_e4m3, fp8_e4m3, 0, %scaleA, 0, %scaleB :
  (vector<8xi32>, vector<8xi32>, vector<16xf32>, i32, i32) -> vector<16xf32>

// Scaled MFMA with fp8 * bf8 inputs.
%r1 = rocdl.mfma.scale.f32.32x32x64.f8f6f4 %a, %a, %c, fp8_e4m3, fp8_e5m2, 0, %scaleA, 0, %scaleB :
  (vector<8xi32>, vector<8xi32>, vector<16xf32>, i32, i32) -> vector<16xf32>

// Scaled MFMA with fp8 * fp6 inputs (6xi32 operand B).
%r2 = rocdl.mfma.scale.f32.32x32x64.f8f6f4 %a, %b6, %c, fp8_e4m3, fp6_e2m3, 0, %scaleA, 0, %scaleB :
  (vector<8xi32>, vector<6xi32>, vector<16xf32>, i32, i32) -> vector<16xf32>

mfma_scale_f32_32x32x64_f8f6f4()

Return op name rocdl.mfma.scale.f32.32x32x64.f8f6f4 as a bitstring.

mfma_scale_f32_32x32x64_f8f6f4(ssa)

rocdl.mfma.scale.f32.32x32x64.f8f6f4

Attributes

  • cbsz - Single, ROCDL_MatrixFormatAttr, matrix operand formats selected by scaled MFMA/WMMA format fields
  • blgp - Single, ROCDL_MatrixFormatAttr, matrix operand formats selected by scaled MFMA/WMMA format fields
  • opselA - Single, I32Attr, 32-bit signless integer attribute
  • opselB - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit signless integer
  • b - Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit signless integer
  • c - Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit float
  • scaleA - Single, I32, 32-bit signless integer
  • scaleB - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Scaled matrix fused multiply-add (MFMA) intrinsic with per-operand scaling. The opselA/opselB and scaleA/scaleB arguments control the scaling of input operands.

Example:

// Scaled MFMA with fp8 * fp8 inputs.
%r0 = rocdl.mfma.scale.f32.32x32x64.f8f6f4 %a, %a, %c, fp8_e4m3, fp8_e4m3, 0, %scaleA, 0, %scaleB :
  (vector<8xi32>, vector<8xi32>, vector<16xf32>, i32, i32) -> vector<16xf32>

// Scaled MFMA with fp8 * bf8 inputs.
%r1 = rocdl.mfma.scale.f32.32x32x64.f8f6f4 %a, %a, %c, fp8_e4m3, fp8_e5m2, 0, %scaleA, 0, %scaleB :
  (vector<8xi32>, vector<8xi32>, vector<16xf32>, i32, i32) -> vector<16xf32>

// Scaled MFMA with fp8 * fp6 inputs (6xi32 operand B).
%r2 = rocdl.mfma.scale.f32.32x32x64.f8f6f4 %a, %b6, %c, fp8_e4m3, fp6_e2m3, 0, %scaleA, 0, %scaleB :
  (vector<8xi32>, vector<6xi32>, vector<16xf32>, i32, i32) -> vector<16xf32>

permlane16_swap()

Return op name rocdl.permlane16.swap as a bitstring.

permlane16_swap(ssa)

rocdl.permlane16.swap

Attributes

  • fi - Single, I1Attr, 1-bit signless integer attribute
  • boundControl - Single, I1Attr, 1-bit signless integer attribute

Operands

  • old - Single, I32, 32-bit signless integer
  • src - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, LLVM dialect-compatible struct of 32-bit signless integerand32-bit signless integer

Description

Performs a permlane16.swap operation with the given operands, applying the permutation specified by $fi to the provided inputs.

Example:

// Swap lanes between groups of 16 threads.
%res = rocdl.permlane16.swap %src, %src, 0, -1 : (i32, i32) -> !llvm.struct<(i32, i32)>

permlane16_var()

Return op name rocdl.permlane16.var as a bitstring.

permlane16_var(ssa)

rocdl.permlane16.var

Attributes

  • fi - Single, I1Attr, 1-bit signless integer attribute
  • boundControl - Single, I1Attr, 1-bit signless integer attribute

Operands

  • old - Single, I32, 32-bit signless integer
  • src0 - Single, I32, 32-bit signless integer
  • src1 - Single, I32, 32-bit signless integer

Results

  • res - Single, I32, 32-bit signless integer

Description

Performs a permlane16.var operation: a per-lane variable-selector intra-row permutation (within each 16-lane row). Maps to llvm.amdgcn.permlane16.var.

Each destination lane within a 16-lane row reads its value from the source lane whose index is given by the corresponding per-lane entry in $src1 (a VGPR). $fi and $boundControl are immediate i1 attrs matching the underlying intrinsic's modifiers.

Example:

%res = rocdl.permlane16.var %old, %src, %selector, false, true : (i32, i32, i32) -> i32

permlane32_swap()

Return op name rocdl.permlane32.swap as a bitstring.

permlane32_swap(ssa)

rocdl.permlane32.swap

Attributes

  • fi - Single, I1Attr, 1-bit signless integer attribute
  • boundControl - Single, I1Attr, 1-bit signless integer attribute

Operands

  • old - Single, I32, 32-bit signless integer
  • src - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, LLVM dialect-compatible struct of 32-bit signless integerand32-bit signless integer

Description

Performs a permlane32.swap operation with the given operands, applying the permutation specified by $fi to the provided inputs.

Example:

// Swap lanes between groups of 32 threads.
%res = rocdl.permlane32.swap %src, %src, 0, -1 : (i32, i32) -> !llvm.struct<(i32, i32)>

permlanex16()

Return op name rocdl.permlanex16 as a bitstring.

permlanex16(ssa)

rocdl.permlanex16

Attributes

  • fi - Single, I1Attr, 1-bit signless integer attribute
  • boundControl - Single, I1Attr, 1-bit signless integer attribute

Operands

  • old - Single, LLVM_Type, LLVM dialect-compatible type
  • src0 - Single, LLVM_Type, LLVM dialect-compatible type
  • src1 - Single, LLVM_Type, LLVM dialect-compatible type
  • src2 - Single, LLVM_Type, LLVM dialect-compatible type

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Performs a permlanex16 operation with the given operands, applying the permutation specified by $fi to the provided inputs.

Example:

// Scalar permlanex16.
%ret0 = rocdl.permlanex16 %src0, %src0, %sel, %sel, 0, -1 : f32, i32

// Vector permlanex16.
%ret1 = rocdl.permlanex16 %src1, %src1, %sel, %sel, 0, -1 : vector<2xf32>, i32

permlanex16_var()

Return op name rocdl.permlanex16.var as a bitstring.

permlanex16_var(ssa)

rocdl.permlanex16.var

Attributes

  • fi - Single, I1Attr, 1-bit signless integer attribute
  • boundControl - Single, I1Attr, 1-bit signless integer attribute

Operands

  • old - Single, I32, 32-bit signless integer
  • src0 - Single, I32, 32-bit signless integer
  • src1 - Single, I32, 32-bit signless integer

Results

  • res - Single, I32, 32-bit signless integer

Description

Performs a permlanex16.var operation: a per-lane variable-selector cross-row permutation (each lane in one 16-lane row reads from a per-lane-chosen source lane in the opposite row). Maps to llvm.amdgcn.permlanex16.var.

With per-lane "identity" selectors this realises an XOR-16 swap pattern in pure VALU; with arbitrary selectors it enables general per-lane cross-half-wave permutations. $fi and $boundControl are immediate i1 attrs matching the underlying intrinsic's modifiers.

Example:

%res = rocdl.permlanex16.var %old, %src, %selector, false, true : (i32, i32, i32) -> i32

ptr_s_buffer_load()

Return op name rocdl.ptr.s.buffer.load as a bitstring.

ptr_s_buffer_load(ssa)

rocdl.ptr.s.buffer.load

Attributes

  • aux - Single, ROCDL_NonAtomicBufferCachePolicyCompatAttr, non-atomic AMDGPU buffer cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • rsrc - Single, ROCDLBufferRsrc, LLVM pointer in address space 8
  • offset - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

raw_buffer_atomic_cmpswap()

Return op name rocdl.raw.buffer.atomic.cmpswap as a bitstring.

raw_buffer_atomic_cmpswap(ssa)

rocdl.raw.buffer.atomic.cmpswap

Attributes

  • aux - Single, ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attribute

Operands

  • src - Single, LLVM_Type, LLVM dialect-compatible type
  • cmp - Single, LLVM_Type, LLVM dialect-compatible type
  • rsrc - Single, ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4
  • offset - Single, I32, 32-bit signless integer
  • soffset - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

raw_buffer_atomic_fadd()

Return op name rocdl.raw.buffer.atomic.fadd as a bitstring.

raw_buffer_atomic_fadd(ssa)

rocdl.raw.buffer.atomic.fadd

Attributes

  • aux - Single, ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attribute

Operands

  • vdata - Single, LLVM_Type, LLVM dialect-compatible type
  • rsrc - Single, ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4
  • offset - Single, I32, 32-bit signless integer
  • soffset - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

raw_buffer_atomic_fmax()

Return op name rocdl.raw.buffer.atomic.fmax as a bitstring.

raw_buffer_atomic_fmax(ssa)

rocdl.raw.buffer.atomic.fmax

Attributes

  • aux - Single, ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attribute

Operands

  • vdata - Single, LLVM_Type, LLVM dialect-compatible type
  • rsrc - Single, ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4
  • offset - Single, I32, 32-bit signless integer
  • soffset - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

raw_buffer_atomic_smax()

Return op name rocdl.raw.buffer.atomic.smax as a bitstring.

raw_buffer_atomic_smax(ssa)

rocdl.raw.buffer.atomic.smax

Attributes

  • aux - Single, ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attribute

Operands

  • vdata - Single, LLVM_Type, LLVM dialect-compatible type
  • rsrc - Single, ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4
  • offset - Single, I32, 32-bit signless integer
  • soffset - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

raw_buffer_atomic_umin()

Return op name rocdl.raw.buffer.atomic.umin as a bitstring.

raw_buffer_atomic_umin(ssa)

rocdl.raw.buffer.atomic.umin

Attributes

  • aux - Single, ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attribute

Operands

  • vdata - Single, LLVM_Type, LLVM dialect-compatible type
  • rsrc - Single, ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4
  • offset - Single, I32, 32-bit signless integer
  • soffset - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

raw_buffer_load()

Return op name rocdl.raw.buffer.load as a bitstring.

raw_buffer_load(ssa)

rocdl.raw.buffer.load

Attributes

  • aux - Single, ROCDL_NonAtomicBufferCachePolicyCompatAttr, non-atomic AMDGPU buffer cache policy attribute

Operands

  • rsrc - Single, ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4
  • offset - Single, I32, 32-bit signless integer
  • soffset - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

raw_buffer_store()

Return op name rocdl.raw.buffer.store as a bitstring.

raw_buffer_store(ssa)

rocdl.raw.buffer.store

Attributes

  • aux - Single, ROCDL_NonAtomicBufferCachePolicyCompatAttr, non-atomic AMDGPU buffer cache policy attribute

Operands

  • vdata - Single, LLVM_Type, LLVM dialect-compatible type
  • rsrc - Single, ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4
  • offset - Single, I32, 32-bit signless integer
  • soffset - Single, I32, 32-bit signless integer

raw_ptr_buffer_atomic_cmpswap()

Return op name rocdl.raw.ptr.buffer.atomic.cmpswap as a bitstring.

raw_ptr_buffer_atomic_cmpswap(ssa)

rocdl.raw.ptr.buffer.atomic.cmpswap

Attributes

  • aux - Single, ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • src - Single, LLVM_Type, LLVM dialect-compatible type
  • cmp - Single, LLVM_Type, LLVM dialect-compatible type
  • rsrc - Single, ROCDLBufferRsrc, LLVM pointer in address space 8
  • offset - Single, I32, 32-bit signless integer
  • soffset - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

raw_ptr_buffer_atomic_fadd()

Return op name rocdl.raw.ptr.buffer.atomic.fadd as a bitstring.

raw_ptr_buffer_atomic_fadd(ssa)

rocdl.raw.ptr.buffer.atomic.fadd

Attributes

  • aux - Single, ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • vdata - Single, LLVM_Type, LLVM dialect-compatible type
  • rsrc - Single, ROCDLBufferRsrc, LLVM pointer in address space 8
  • offset - Single, I32, 32-bit signless integer
  • soffset - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

raw_ptr_buffer_atomic_fmax()

Return op name rocdl.raw.ptr.buffer.atomic.fmax as a bitstring.

raw_ptr_buffer_atomic_fmax(ssa)

rocdl.raw.ptr.buffer.atomic.fmax

Attributes

  • aux - Single, ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • vdata - Single, LLVM_Type, LLVM dialect-compatible type
  • rsrc - Single, ROCDLBufferRsrc, LLVM pointer in address space 8
  • offset - Single, I32, 32-bit signless integer
  • soffset - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

raw_ptr_buffer_atomic_smax()

Return op name rocdl.raw.ptr.buffer.atomic.smax as a bitstring.

raw_ptr_buffer_atomic_smax(ssa)

rocdl.raw.ptr.buffer.atomic.smax

Attributes

  • aux - Single, ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • vdata - Single, LLVM_Type, LLVM dialect-compatible type
  • rsrc - Single, ROCDLBufferRsrc, LLVM pointer in address space 8
  • offset - Single, I32, 32-bit signless integer
  • soffset - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

raw_ptr_buffer_atomic_umin()

Return op name rocdl.raw.ptr.buffer.atomic.umin as a bitstring.

raw_ptr_buffer_atomic_umin(ssa)

rocdl.raw.ptr.buffer.atomic.umin

Attributes

  • aux - Single, ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • vdata - Single, LLVM_Type, LLVM dialect-compatible type
  • rsrc - Single, ROCDLBufferRsrc, LLVM pointer in address space 8
  • offset - Single, I32, 32-bit signless integer
  • soffset - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

raw_ptr_buffer_load()

Return op name rocdl.raw.ptr.buffer.load as a bitstring.

raw_ptr_buffer_load(ssa)

rocdl.raw.ptr.buffer.load

Attributes

  • aux - Single, ROCDL_NonAtomicBufferCachePolicyCompatAttr, non-atomic AMDGPU buffer cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • rsrc - Single, ROCDLBufferRsrc, LLVM pointer in address space 8
  • offset - Single, I32, 32-bit signless integer
  • soffset - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

raw_ptr_buffer_load_async_lds()

Return op name rocdl.raw.ptr.buffer.load.async.lds as a bitstring.

raw_ptr_buffer_load_async_lds(ssa)

rocdl.raw.ptr.buffer.load.async.lds - Async variant of raw.ptr.buffer.load.lds

Attributes

  • aux - Single, ROCDL_PreGfx12CachePolicyCompatAttr, pre-gfx12 or gfx942 AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • rsrc - Single, ROCDLBufferRsrc, LLVM pointer in address space 8
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3
  • size - Single, I32, 32-bit signless integer
  • voffset - Single, I32, 32-bit signless integer
  • soffset - Single, I32, 32-bit signless integer
  • offset - Single, I32, 32-bit signless integer

Description

Load from a buffer resource rsrc to ldsPtr, which must be uniform.

See rocdl.load.async.to.lds for overall semantics of such loads, noting that here voffset can be lane-varying and that rsrc (which holds the base addres) must, as always, be uniform.

Available on gfx9 and gfx10.

Example:

// Async buffer load to LDS via buffer resource pointer.
rocdl.raw.ptr.buffer.load.async.lds %rsrc, %ldsPtr, %size, %voffset, %soffset, %offset, 0

raw_ptr_buffer_load_lds()

Return op name rocdl.raw.ptr.buffer.load.lds as a bitstring.

raw_ptr_buffer_load_lds(ssa)

rocdl.raw.ptr.buffer.load.lds

Attributes

  • aux - Single, ROCDL_NonAtomicBufferCachePolicyCompatAttr, non-atomic AMDGPU buffer cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • rsrc - Single, ROCDLBufferRsrc, LLVM pointer in address space 8
  • ldsPtr - Single, ROCDLBufferLDS, LLVM pointer in address space 3
  • size - Single, I32, 32-bit signless integer
  • voffset - Single, I32, 32-bit signless integer
  • soffset - Single, I32, 32-bit signless integer
  • offset - Single, I32, 32-bit signless integer

raw_ptr_buffer_store()

Return op name rocdl.raw.ptr.buffer.store as a bitstring.

raw_ptr_buffer_store(ssa)

rocdl.raw.ptr.buffer.store

Attributes

  • aux - Single, ROCDL_NonAtomicBufferCachePolicyCompatAttr, non-atomic AMDGPU buffer cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • vdata - Single, LLVM_Type, LLVM dialect-compatible type
  • rsrc - Single, ROCDLBufferRsrc, LLVM pointer in address space 8
  • offset - Single, I32, 32-bit signless integer
  • soffset - Single, I32, 32-bit signless integer

rcp()

Return op name rocdl.rcp as a bitstring.

rcp(ssa)

rocdl.rcp

Operands

  • arg - Single, LLVM_AnyFloat, floating point LLVM type

Results

  • res - Single, LLVM_AnyFloat, floating point LLVM type

Description

Note: In the general case, prefer the conventional arith, math, or llvm ops over this. Use this ROCDL-specific operation only when you fully understand its implication and when it is strictly necessary. This op is usually chosen when a small loss in precision is acceptable in exchange for higher execution speed.

Example:

%0 = rocdl.rcp %a f32 -> f32

readfirstlane()

Return op name rocdl.readfirstlane as a bitstring.

readfirstlane(ssa)

rocdl.readfirstlane - Get the value in first active lane.

Operands

  • src - Single, LLVM_Type, LLVM dialect-compatible type

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Returns the value in the lowest active lane of the input operand.

Example:

// Scalar readfirstlane.
%0 = rocdl.readfirstlane %src0 : f32

// Vector readfirstlane.
%1 = rocdl.readfirstlane %src1 : vector<2xf32>

readlane()

Return op name rocdl.readlane as a bitstring.

readlane(ssa)

rocdl.readlane - Get the value in the specific lane.

Operands

  • src0 - Single, LLVM_Type, LLVM dialect-compatible type
  • src1 - Single, I32, 32-bit signless integer

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Get the value in lane src1 from input src0.

Example:

// Scalar readlane.
%0 = rocdl.readlane %src0, %idx : (f32, i32) -> f32

// Vector readlane.
%1 = rocdl.readlane %src1, %idx : (vector<2xf32>, i32) -> vector<2xf32>

rsq()

Return op name rocdl.rsq as a bitstring.

rsq(ssa)

rocdl.rsq

Operands

  • arg - Single, LLVM_AnyFloat, floating point LLVM type

Results

  • res - Single, LLVM_AnyFloat, floating point LLVM type

Description

Note: In the general case, prefer the conventional arith, math, or llvm ops over this. Use this ROCDL-specific operation only when you fully understand its implication and when it is strictly necessary. This op is usually chosen when a small loss in precision is acceptable in exchange for higher execution speed.

Example:

%0 = rocdl.rsq %a f32 -> f32

s_barrier()

Return op name rocdl.s.barrier as a bitstring.

s_barrier(ssa)

rocdl.s.barrier

Description

Insert a workgroup barrier without memory fences.

Available on gfx9 and later but deprecated on gfx12+; see rocdl.s.barrier.signal and rocdl.s.barrier.wait instead.

Example:

// Synchronize threads within a workgroup.
rocdl.s.barrier

s_barrier_init()

Return op name rocdl.s.barrier.init as a bitstring.

s_barrier_init(ssa)

rocdl.s.barrier.init

Attributes

  • memberCnt - Single, I32Attr, 32-bit signless integer attribute

Operands

  • ptr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Description

Available on gfx1250+.

Example:

// Initialize a named barrier with member count.
rocdl.s.barrier.init %ptr member_cnt = 1 : !llvm.ptr<3>

s_barrier_join()

Return op name rocdl.s.barrier.join as a bitstring.

s_barrier_join(ssa)

rocdl.s.barrier.join

Operands

  • ptr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Description

Available on gfx1250+.

Example:

// Join a named barrier.
rocdl.s.barrier.join %ptr : !llvm.ptr<3>

s_barrier_leave()

Return op name rocdl.s.barrier.leave as a bitstring.

s_barrier_leave(ssa)

rocdl.s.barrier.leave

Attributes

  • id - Single, I16Attr, 16-bit signless integer attribute

Description

Available on gfx1250+.

Example:

// Leave a named barrier by id.
rocdl.s.barrier.leave id = 1

s_barrier_signal()

Return op name rocdl.s.barrier.signal as a bitstring.

s_barrier_signal(ssa)

rocdl.s.barrier.signal

Attributes

  • id - Single, I32Attr, 32-bit signless integer attribute

Description

Signal a barrier by id. Available on gfx1250+.

Example:

// Signal barrier with id -1 (all barriers).
rocdl.s.barrier.signal id = -1

s_barrier_signal_isfirst()

Return op name rocdl.s.barrier.signal.isfirst as a bitstring.

s_barrier_signal_isfirst(ssa)

rocdl.s.barrier.signal.isfirst

Attributes

  • id - Single, I32Attr, 32-bit signless integer attribute

Results

  • res - Single, I1, 1-bit signless integer

Description

Available on gfx1200+.

Example:

// Signal barrier and check if this wave is first to arrive.
%0 = rocdl.s.barrier.signal.isfirst id = 1 -> i1

s_barrier_signal_var()

Return op name rocdl.s.barrier.signal.var as a bitstring.

s_barrier_signal_var(ssa)

rocdl.s.barrier.signal.var

Attributes

  • memberCnt - Single, I32Attr, 32-bit signless integer attribute

Operands

  • ptr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Description

Available on gfx1250+.

If memberCnt is 0, the member count is retained from a previous initialization.

Example:

// Signal a named barrier with variable ID.
rocdl.s.barrier.signal.var %ptr member_cnt = 1 : !llvm.ptr<3>

s_barrier_wait()

Return op name rocdl.s.barrier.wait as a bitstring.

s_barrier_wait(ssa)

rocdl.s.barrier.wait

Attributes

  • id - Single, I16Attr, 16-bit signless integer attribute

Description

Wait on a barrier by id. Available on gfx1200+.

Example:

// Wait on barrier with id -1 (all barriers).
rocdl.s.barrier.wait id = -1

s_get_barrier_state()

Return op name rocdl.s.get.barrier.state as a bitstring.

s_get_barrier_state(ssa)

rocdl.s.get.barrier.state

Attributes

  • id - Single, I32Attr, 32-bit signless integer attribute

Results

  • res - Single, I32, 32-bit signless integer

Description

Available on gfx1200+.

Example:

// Query barrier state by id.
%0 = rocdl.s.get.barrier.state id = 1 -> i32

s_get_named_barrier_state()

Return op name rocdl.s.get.named.barrier.state as a bitstring.

s_get_named_barrier_state(ssa)

rocdl.s.get.named.barrier.state

Operands

  • ptr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Results

  • res - Single, I32, 32-bit signless integer

Description

Available on gfx1250+.

Example:

// Query named barrier state by pointer.
%0 = rocdl.s.get.named.barrier.state %ptr : !llvm.ptr<3> -> i32

s_nop()

Return op name rocdl.s.nop as a bitstring.

s_nop(ssa)

rocdl.s.nop

Attributes

  • count - Single, I16Attr, 16-bit signless integer attribute

Description

Insert a number of NOP cycles.

Example:

// Insert a no-op.
rocdl.s.nop 0

s_setprio()

Return op name rocdl.s.setprio as a bitstring.

s_setprio(ssa)

rocdl.s.setprio

Attributes

  • priority - Single, I16Attr, 16-bit signless integer attribute

Description

Set the wavefront scheduling priority.

Example:

// Set priority to 0.
rocdl.s.setprio 0

s_sleep()

Return op name rocdl.s.sleep as a bitstring.

s_sleep(ssa)

rocdl.s.sleep

Attributes

  • count - Single, I32Attr, 32-bit signless integer attribute

Description

Sleep for a number of clock cycles.

Example:

// Sleep for a minimum duration.
rocdl.s.sleep 0

s_wait_asynccnt()

Return op name rocdl.s.wait.asynccnt as a bitstring.

s_wait_asynccnt(ssa)

rocdl.s.wait.asynccnt - Wait until ASYNCCNT is less than or equal to count

Attributes

  • count - Single, I16Attr, 16-bit signless integer attribute

Description

Wait for the counter specified to be less-than or equal-to the count before continuing.

Available on gfx1250+.

Example:

// Wait for async counter to drain.
rocdl.s.wait.asynccnt 0

s_wait_dscnt()

Return op name rocdl.s.wait.dscnt as a bitstring.

s_wait_dscnt(ssa)

rocdl.s.wait.dscnt - Wait until DSCNT is less than or equal to count

Attributes

  • count - Single, I16Attr, 16-bit signless integer attribute

Description

Wait for the counter specified to be less-than or equal-to the count before continuing.

Available on gfx12+.

Example:

// Wait for data-sharing counter to drain.
rocdl.s.wait.dscnt 0

s_wait_expcnt()

Return op name rocdl.s.wait.expcnt as a bitstring.

s_wait_expcnt(ssa)

rocdl.s.wait.expcnt - Wait until EXPCNT is less than or equal to count

Attributes

  • count - Single, I16Attr, 16-bit signless integer attribute

Description

Wait for the counter specified to be less-than or equal-to the count before continuing.

Available on gfx12+.

Example:

// Wait for export counter to drain.
rocdl.s.wait.expcnt 0

s_wait_loadcnt()

Return op name rocdl.s.wait.loadcnt as a bitstring.

s_wait_loadcnt(ssa)

rocdl.s.wait.loadcnt - Wait until LOADCNT is less than or equal to count

Attributes

  • count - Single, I16Attr, 16-bit signless integer attribute

Description

Wait for the counter specified to be less-than or equal-to the count before continuing.

Available on gfx12+.

Example:

// Wait for load counter to drain.
rocdl.s.wait.loadcnt 0

s_wait_storecnt()

Return op name rocdl.s.wait.storecnt as a bitstring.

s_wait_storecnt(ssa)

rocdl.s.wait.storecnt - Wait until STORECNT is less than or equal to count

Attributes

  • count - Single, I16Attr, 16-bit signless integer attribute

Description

Wait for the counter specified to be less-than or equal-to the count before continuing.

Available on gfx12+.

Example:

// Wait for store counter to drain.
rocdl.s.wait.storecnt 0

s_wait_tensorcnt()

Return op name rocdl.s.wait.tensorcnt as a bitstring.

s_wait_tensorcnt(ssa)

rocdl.s.wait.tensorcnt - Wait until TENSORCNT is less than or equal to count

Attributes

  • count - Single, I16Attr, 16-bit signless integer attribute

Description

Wait for the counter specified to be less-than or equal-to the count before continuing.

Available on gfx1250+.

Example:

// Wait for tensor counter to drain.
rocdl.s.wait.tensorcnt 0

s_waitcnt()

Return op name rocdl.s.waitcnt as a bitstring.

s_waitcnt(ssa)

rocdl.s.waitcnt

Attributes

  • bitfield - Single, I32Attr, 32-bit signless integer attribute

Description

Wait for outstanding memory operations to complete, as specified by a bitfield whose semantics depend on the target chipset.

Example:

// Wait for all counters to reach zero.
rocdl.s.waitcnt 0

s_wakeup_barrier()

Return op name rocdl.s.wakeup.barrier as a bitstring.

s_wakeup_barrier(ssa)

rocdl.s.wakeup.barrier

Operands

  • ptr - Single, ROCDLBufferLDS, LLVM pointer in address space 3

Description

Wakes up waves associated with a given named barrier. Note, This op does not release waves waiting at the barrier. It just signal other waves in the same work-group waiting on the indicated named barrier to wake up. Available on gfx1250+.

Example:

// Wake up waves waiting on a named barrier.
rocdl.s.wakeup.barrier %ptr : !llvm.ptr<3>

sched_barrier()

Return op name rocdl.sched.barrier as a bitstring.

sched_barrier(ssa)

rocdl.sched.barrier

Attributes

  • mask - Single, ROCDL_SchedGroupMaskAttr, instruction type mask for scheduling barriers

Description

Insert a scheduling barrier with the given mask. The mask is a bitfield that controls which instruction types may be scheduled across the barrier. The mask values mirror the llvm.amdgcn.sched.barrier intrinsic's documented mask values and the AMDGPU backend's SchedGroupMask enum.

Example:

// Scheduling barrier with no instructions allowed to cross.
rocdl.sched.barrier none

// Allow VALU and all VMEM instructions to cross.
rocdl.sched.barrier valu|all_vmem

sched_group_barrier()

Return op name rocdl.sched.group.barrier as a bitstring.

sched_group_barrier(ssa)

rocdl.sched.group.barrier

Attributes

  • mask - Single, ROCDL_SchedGroupMaskAttr, instruction type mask for scheduling barriers
  • size - Single, I32Attr, 32-bit signless integer attribute
  • groupId - Single, I32Attr, 32-bit signless integer attribute

Description

Insert a scheduling group barrier. The first parameter uses the same scheduling group mask values as rocdl.sched.barrier.

Example:

// Schedule group barrier with mask, size, and group id.
rocdl.sched.group.barrier mfma_wmma, 1, 0

sdot2()

Return op name rocdl.sdot2 as a bitstring.

sdot2(ssa)

rocdl.sdot2

Attributes

  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2
  • b - Single, ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2
  • c - Single, anonymous/composite constraint, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit signless integer

Description

Packed intra-lane dot-product with optional result clamping (clamp). Computes res = sum_i a[i]*b[i] + c, where a and b hold packed 4/8/16-bit data (for dot2,dot4,dot8).

Example:

%r = rocdl.sdot2 %a, %b, %c {clamp = true} :
     (vector<2xi16>, vector<2xi16>, i32) -> i32

sdot4()

Return op name rocdl.sdot4 as a bitstring.

sdot4(ssa)

rocdl.sdot4

Attributes

  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, 32-bit signless integer
  • b - Single, anonymous/composite constraint, 32-bit signless integer
  • c - Single, anonymous/composite constraint, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit signless integer

Description

Packed intra-lane dot-product with optional result clamping (clamp). Computes res = sum_i a[i]*b[i] + c, where a and b hold packed 4/8/16-bit data (for dot2,dot4,dot8).

Example:

%r = rocdl.sdot4 %a, %b, %c {clamp = true} :
     (i32, i32, i32) -> i32

sdot8()

Return op name rocdl.sdot8 as a bitstring.

sdot8(ssa)

rocdl.sdot8

Attributes

  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, 32-bit signless integer
  • b - Single, anonymous/composite constraint, 32-bit signless integer
  • c - Single, anonymous/composite constraint, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit signless integer

Description

Packed intra-lane dot-product with optional result clamping (clamp). Computes res = sum_i a[i]*b[i] + c, where a and b hold packed 4/8/16-bit data (for dot2,dot4,dot8).

Example:

%r = rocdl.sdot8 %a, %b, %c {clamp = true} :
     (i32, i32, i32) -> i32

sin()

Return op name rocdl.sin as a bitstring.

sin(ssa)

rocdl.sin

Operands

  • arg - Single, LLVM_AnyFloat, floating point LLVM type

Results

  • res - Single, LLVM_AnyFloat, floating point LLVM type

Description

Note: In the general case, prefer the conventional arith, math, or llvm ops over this. Use this ROCDL-specific operation only when you fully understand its implication and when it is strictly necessary. This op is usually chosen when a small loss in precision is acceptable in exchange for higher execution speed.

Example:

%0 = rocdl.sin %a f32 -> f32

smfmac_f32_16x16x32_bf16()

Return op name rocdl.smfmac.f32.16x16x32.bf16 as a bitstring.

smfmac_f32_16x16x32_bf16(ssa)

rocdl.smfmac.f32.16x16x32.bf16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_16x16x32_f16()

Return op name rocdl.smfmac.f32.16x16x32.f16 as a bitstring.

smfmac_f32_16x16x32_f16(ssa)

rocdl.smfmac.f32.16x16x32.f16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_16x16x64_bf8_bf8()

Return op name rocdl.smfmac.f32.16x16x64.bf8.bf8 as a bitstring.

smfmac_f32_16x16x64_bf8_bf8(ssa)

rocdl.smfmac.f32.16x16x64.bf8.bf8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_16x16x64_bf8_fp8()

Return op name rocdl.smfmac.f32.16x16x64.bf8.fp8 as a bitstring.

smfmac_f32_16x16x64_bf8_fp8(ssa)

rocdl.smfmac.f32.16x16x64.bf8.fp8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_16x16x64_bf16()

Return op name rocdl.smfmac.f32.16x16x64.bf16 as a bitstring.

smfmac_f32_16x16x64_bf16(ssa)

rocdl.smfmac.f32.16x16x64.bf16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of bfloat16 type values of length 8
  • b - Single, anonymous/composite constraint, fixed-length vector of bfloat16 type values of length 16
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_16x16x64_f16()

Return op name rocdl.smfmac.f32.16x16x64.f16 as a bitstring.

smfmac_f32_16x16x64_f16(ssa)

rocdl.smfmac.f32.16x16x64.f16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 8
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 16
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_16x16x64_fp8_bf8()

Return op name rocdl.smfmac.f32.16x16x64.fp8.bf8 as a bitstring.

smfmac_f32_16x16x64_fp8_bf8(ssa)

rocdl.smfmac.f32.16x16x64.fp8.bf8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_16x16x64_fp8_fp8()

Return op name rocdl.smfmac.f32.16x16x64.fp8.fp8 as a bitstring.

smfmac_f32_16x16x64_fp8_fp8(ssa)

rocdl.smfmac.f32.16x16x64.fp8.fp8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_16x16x128_bf8_bf8()

Return op name rocdl.smfmac.f32.16x16x128.bf8.bf8 as a bitstring.

smfmac_f32_16x16x128_bf8_bf8(ssa)

rocdl.smfmac.f32.16x16x128.bf8.bf8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_16x16x128_bf8_fp8()

Return op name rocdl.smfmac.f32.16x16x128.bf8.fp8 as a bitstring.

smfmac_f32_16x16x128_bf8_fp8(ssa)

rocdl.smfmac.f32.16x16x128.bf8.fp8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_16x16x128_fp8_bf8()

Return op name rocdl.smfmac.f32.16x16x128.fp8.bf8 as a bitstring.

smfmac_f32_16x16x128_fp8_bf8(ssa)

rocdl.smfmac.f32.16x16x128.fp8.bf8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_16x16x128_fp8_fp8()

Return op name rocdl.smfmac.f32.16x16x128.fp8.fp8 as a bitstring.

smfmac_f32_16x16x128_fp8_fp8(ssa)

rocdl.smfmac.f32.16x16x128.fp8.fp8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_32x32x16_bf16()

Return op name rocdl.smfmac.f32.32x32x16.bf16 as a bitstring.

smfmac_f32_32x32x16_bf16(ssa)

rocdl.smfmac.f32.32x32x16.bf16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_32x32x16_f16()

Return op name rocdl.smfmac.f32.32x32x16.f16 as a bitstring.

smfmac_f32_32x32x16_f16(ssa)

rocdl.smfmac.f32.32x32x16.f16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_32x32x32_bf8_bf8()

Return op name rocdl.smfmac.f32.32x32x32.bf8.bf8 as a bitstring.

smfmac_f32_32x32x32_bf8_bf8(ssa)

rocdl.smfmac.f32.32x32x32.bf8.bf8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_32x32x32_bf8_fp8()

Return op name rocdl.smfmac.f32.32x32x32.bf8.fp8 as a bitstring.

smfmac_f32_32x32x32_bf8_fp8(ssa)

rocdl.smfmac.f32.32x32x32.bf8.fp8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_32x32x32_bf16()

Return op name rocdl.smfmac.f32.32x32x32.bf16 as a bitstring.

smfmac_f32_32x32x32_bf16(ssa)

rocdl.smfmac.f32.32x32x32.bf16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of bfloat16 type values of length 8
  • b - Single, anonymous/composite constraint, fixed-length vector of bfloat16 type values of length 16
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_32x32x32_f16()

Return op name rocdl.smfmac.f32.32x32x32.f16 as a bitstring.

smfmac_f32_32x32x32_f16(ssa)

rocdl.smfmac.f32.32x32x32.f16

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 8
  • b - Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 16
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_32x32x32_fp8_bf8()

Return op name rocdl.smfmac.f32.32x32x32.fp8.bf8 as a bitstring.

smfmac_f32_32x32x32_fp8_bf8(ssa)

rocdl.smfmac.f32.32x32x32.fp8.bf8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_32x32x32_fp8_fp8()

Return op name rocdl.smfmac.f32.32x32x32.fp8.fp8 as a bitstring.

smfmac_f32_32x32x32_fp8_fp8(ssa)

rocdl.smfmac.f32.32x32x32.fp8.fp8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_32x32x64_bf8_bf8()

Return op name rocdl.smfmac.f32.32x32x64.bf8.bf8 as a bitstring.

smfmac_f32_32x32x64_bf8_bf8(ssa)

rocdl.smfmac.f32.32x32x64.bf8.bf8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_32x32x64_bf8_fp8()

Return op name rocdl.smfmac.f32.32x32x64.bf8.fp8 as a bitstring.

smfmac_f32_32x32x64_bf8_fp8(ssa)

rocdl.smfmac.f32.32x32x64.bf8.fp8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_32x32x64_fp8_bf8()

Return op name rocdl.smfmac.f32.32x32x64.fp8.bf8 as a bitstring.

smfmac_f32_32x32x64_fp8_bf8(ssa)

rocdl.smfmac.f32.32x32x64.fp8.bf8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_f32_32x32x64_fp8_fp8()

Return op name rocdl.smfmac.f32.32x32x64.fp8.fp8 as a bitstring.

smfmac_f32_32x32x64_fp8_fp8(ssa)

rocdl.smfmac.f32.32x32x64.fp8.fp8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_i32_16x16x64_i8()

Return op name rocdl.smfmac.i32.16x16x64.i8 as a bitstring.

smfmac_i32_16x16x64_i8(ssa)

rocdl.smfmac.i32.16x16x64.i8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_i32_16x16x128_i8()

Return op name rocdl.smfmac.i32.16x16x128.i8 as a bitstring.

smfmac_i32_16x16x128_i8(ssa)

rocdl.smfmac.i32.16x16x128.i8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_i32_32x32x32_i8()

Return op name rocdl.smfmac.i32.32x32x32.i8 as a bitstring.

smfmac_i32_32x32x32_i8(ssa)

rocdl.smfmac.i32.32x32x32.i8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

smfmac_i32_32x32x64_i8()

Return op name rocdl.smfmac.i32.32x32x64.i8 as a bitstring.

smfmac_i32_32x32x64_i8(ssa)

rocdl.smfmac.i32.32x32x64.i8

Attributes

  • cbsz - Single, I32Attr, 32-bit signless integer attribute
  • abid - Single, I32Attr, 32-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
  • b - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8
  • c - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16

Description

Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4 structured sparsity. The index operand provides the sparsity metadata, and cbsz/abid control broadcast modes.

Example:

// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
  (vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
  (vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>

// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>

// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
  (vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>

sqrt()

Return op name rocdl.sqrt as a bitstring.

sqrt(ssa)

rocdl.sqrt

Operands

  • arg - Single, LLVM_AnyFloat, floating point LLVM type

Results

  • res - Single, LLVM_AnyFloat, floating point LLVM type

Description

Note: In the general case, prefer the conventional arith, math, or llvm ops over this. Use this ROCDL-specific operation only when you fully understand its implication and when it is strictly necessary. This op is usually chosen when a small loss in precision is acceptable in exchange for higher execution speed.

Example:

%0 = rocdl.sqrt %a f32 -> f32

sudot4()

Return op name rocdl.sudot4 as a bitstring.

sudot4(ssa)

rocdl.sudot4

Attributes

  • signA - Single, I1Attr, 1-bit signless integer attribute
  • signB - Single, I1Attr, 1-bit signless integer attribute
  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, I32, 32-bit signless integer
  • b - Single, I32, 32-bit signless integer
  • c - Single, I32, 32-bit signless integer

Results

  • res - Single, I32, 32-bit signless integer

Description

Mixed-signedness packed dot-product with per-operand sign controls. Computes res = sum_i a[i]*b[i] + c. Each lane of a is treated as signed when signA = true; when signA = false, the unsigned lane value is zero-extended into a wider signed integer. signB controls the same for b. clamp controls result clamping.

These ops correspond to RDNA's unified mixed-sign v_dot4_i32_iu8 and v_dot8_i32_iu4 instructions (gfx11+).

Example:

%r = rocdl.sudot4 %a, %b, %c
       {signA = true, signB = false, clamp = true} :
     (i32, i32, i32) -> i32

sudot8()

Return op name rocdl.sudot8 as a bitstring.

sudot8(ssa)

rocdl.sudot8

Attributes

  • signA - Single, I1Attr, 1-bit signless integer attribute
  • signB - Single, I1Attr, 1-bit signless integer attribute
  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, I32, 32-bit signless integer
  • b - Single, I32, 32-bit signless integer
  • c - Single, I32, 32-bit signless integer

Results

  • res - Single, I32, 32-bit signless integer

Description

Mixed-signedness packed dot-product with per-operand sign controls. Computes res = sum_i a[i]*b[i] + c. Each lane of a is treated as signed when signA = true; when signA = false, the unsigned lane value is zero-extended into a wider signed integer. signB controls the same for b. clamp controls result clamping.

These ops correspond to RDNA's unified mixed-sign v_dot4_i32_iu8 and v_dot8_i32_iu4 instructions (gfx11+).

Example:

%r = rocdl.sudot8 %a, %b, %c
       {signA = true, signB = false, clamp = true} :
     (i32, i32, i32) -> i32

swmmac_bf16_16x16x32_bf16()

Return op name rocdl.swmmac.bf16.16x16x32.bf16 as a bitstring.

swmmac_bf16_16x16x32_bf16(ssa)

rocdl.swmmac.bf16.16x16x32.bf16

Operands

  • a - Single, anonymous/composite constraint, LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, LLVM dialect-compatible vector of integer
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, LLVM dialect-compatible vector of integer

swmmac_bf16_16x16x64_bf16()

Return op name rocdl.swmmac.bf16.16x16x64.bf16 as a bitstring.

swmmac_bf16_16x16x64_bf16(ssa)

rocdl.swmmac.bf16.16x16x64.bf16

Attributes

  • signA - Single, I1Attr, 1-bit signless integer attribute
  • signB - Single, I1Attr, 1-bit signless integer attribute
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
  • b - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
  • c - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type

swmmac_bf16f32_16x16x64_bf16()

Return op name rocdl.swmmac.bf16f32.16x16x64.bf16 as a bitstring.

swmmac_bf16f32_16x16x64_bf16(ssa)

rocdl.swmmac.bf16f32.16x16x64.bf16

Attributes

  • signA - Single, I1Attr, 1-bit signless integer attribute
  • signB - Single, I1Attr, 1-bit signless integer attribute
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
  • b - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
  • c - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type

swmmac_f16_16x16x32_f16()

Return op name rocdl.swmmac.f16.16x16x32.f16 as a bitstring.

swmmac_f16_16x16x32_f16(ssa)

rocdl.swmmac.f16.16x16x32.f16

Operands

  • a - Single, anonymous/composite constraint, LLVM dialect-compatible vector of 16-bit float
  • b - Single, anonymous/composite constraint, LLVM dialect-compatible vector of 16-bit float
  • c - Single, anonymous/composite constraint, LLVM dialect-compatible vector of 16-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, LLVM dialect-compatible vector of 16-bit float

swmmac_f16_16x16x64_f16()

Return op name rocdl.swmmac.f16.16x16x64.f16 as a bitstring.

swmmac_f16_16x16x64_f16(ssa)

rocdl.swmmac.f16.16x16x64.f16

Attributes

  • signA - Single, I1Attr, 1-bit signless integer attribute
  • signB - Single, I1Attr, 1-bit signless integer attribute
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • b - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • c - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

swmmac_f16_16x16x128_bf8_bf8()

Return op name rocdl.swmmac.f16.16x16x128.bf8.bf8 as a bitstring.

swmmac_f16_16x16x128_bf8_bf8(ssa)

rocdl.swmmac.f16.16x16x128.bf8.bf8

Attributes

  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

swmmac_f16_16x16x128_bf8_fp8()

Return op name rocdl.swmmac.f16.16x16x128.bf8.fp8 as a bitstring.

swmmac_f16_16x16x128_bf8_fp8(ssa)

rocdl.swmmac.f16.16x16x128.bf8.fp8

Attributes

  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

swmmac_f16_16x16x128_fp8_bf8()

Return op name rocdl.swmmac.f16.16x16x128.fp8.bf8 as a bitstring.

swmmac_f16_16x16x128_fp8_bf8(ssa)

rocdl.swmmac.f16.16x16x128.fp8.bf8

Attributes

  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

swmmac_f16_16x16x128_fp8_fp8()

Return op name rocdl.swmmac.f16.16x16x128.fp8.fp8 as a bitstring.

swmmac_f16_16x16x128_fp8_fp8(ssa)

rocdl.swmmac.f16.16x16x128.fp8.fp8

Attributes

  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

swmmac_f32_16x16x32_bf8_bf8()

Return op name rocdl.swmmac.f32.16x16x32.bf8.bf8 as a bitstring.

swmmac_f32_16x16x32_bf8_bf8(ssa)

rocdl.swmmac.f32.16x16x32.bf8.bf8

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

swmmac_f32_16x16x32_bf8_fp8()

Return op name rocdl.swmmac.f32.16x16x32.bf8.fp8 as a bitstring.

swmmac_f32_16x16x32_bf8_fp8(ssa)

rocdl.swmmac.f32.16x16x32.bf8.fp8

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

swmmac_f32_16x16x32_bf16()

Return op name rocdl.swmmac.f32.16x16x32.bf16 as a bitstring.

swmmac_f32_16x16x32_bf16(ssa)

rocdl.swmmac.f32.16x16x32.bf16

Operands

  • a - Single, anonymous/composite constraint, LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit float

swmmac_f32_16x16x32_f16()

Return op name rocdl.swmmac.f32.16x16x32.f16 as a bitstring.

swmmac_f32_16x16x32_f16(ssa)

rocdl.swmmac.f32.16x16x32.f16

Operands

  • a - Single, anonymous/composite constraint, LLVM dialect-compatible vector of 16-bit float
  • b - Single, anonymous/composite constraint, LLVM dialect-compatible vector of 16-bit float
  • c - Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit float

swmmac_f32_16x16x32_fp8_bf8()

Return op name rocdl.swmmac.f32.16x16x32.fp8.bf8 as a bitstring.

swmmac_f32_16x16x32_fp8_bf8(ssa)

rocdl.swmmac.f32.16x16x32.fp8.bf8

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

swmmac_f32_16x16x32_fp8_fp8()

Return op name rocdl.swmmac.f32.16x16x32.fp8.fp8 as a bitstring.

swmmac_f32_16x16x32_fp8_fp8(ssa)

rocdl.swmmac.f32.16x16x32.fp8.fp8

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

swmmac_f32_16x16x64_bf16()

Return op name rocdl.swmmac.f32.16x16x64.bf16 as a bitstring.

swmmac_f32_16x16x64_bf16(ssa)

rocdl.swmmac.f32.16x16x64.bf16

Attributes

  • signA - Single, I1Attr, 1-bit signless integer attribute
  • signB - Single, I1Attr, 1-bit signless integer attribute
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
  • b - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

swmmac_f32_16x16x64_f16()

Return op name rocdl.swmmac.f32.16x16x64.f16 as a bitstring.

swmmac_f32_16x16x64_f16(ssa)

rocdl.swmmac.f32.16x16x64.f16

Attributes

  • signA - Single, I1Attr, 1-bit signless integer attribute
  • signB - Single, I1Attr, 1-bit signless integer attribute
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • b - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

swmmac_f32_16x16x128_bf8_bf8()

Return op name rocdl.swmmac.f32.16x16x128.bf8.bf8 as a bitstring.

swmmac_f32_16x16x128_bf8_bf8(ssa)

rocdl.swmmac.f32.16x16x128.bf8.bf8

Attributes

  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

swmmac_f32_16x16x128_bf8_fp8()

Return op name rocdl.swmmac.f32.16x16x128.bf8.fp8 as a bitstring.

swmmac_f32_16x16x128_bf8_fp8(ssa)

rocdl.swmmac.f32.16x16x128.bf8.fp8

Attributes

  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

swmmac_f32_16x16x128_fp8_bf8()

Return op name rocdl.swmmac.f32.16x16x128.fp8.bf8 as a bitstring.

swmmac_f32_16x16x128_fp8_bf8(ssa)

rocdl.swmmac.f32.16x16x128.fp8.bf8

Attributes

  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

swmmac_f32_16x16x128_fp8_fp8()

Return op name rocdl.swmmac.f32.16x16x128.fp8.fp8 as a bitstring.

swmmac_f32_16x16x128_fp8_fp8(ssa)

rocdl.swmmac.f32.16x16x128.fp8.fp8

Attributes

  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

swmmac_i32_16x16x32_iu4()

Return op name rocdl.swmmac.i32.16x16x32.iu4 as a bitstring.

swmmac_i32_16x16x32_iu4(ssa)

rocdl.swmmac.i32.16x16x32.iu4

Attributes

  • signA - Single, I1Attr, 1-bit signless integer attribute
  • signB - Single, I1Attr, 1-bit signless integer attribute
  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer

swmmac_i32_16x16x32_iu8()

Return op name rocdl.swmmac.i32.16x16x32.iu8 as a bitstring.

swmmac_i32_16x16x32_iu8(ssa)

rocdl.swmmac.i32.16x16x32.iu8

Attributes

  • signA - Single, I1Attr, 1-bit signless integer attribute
  • signB - Single, I1Attr, 1-bit signless integer attribute
  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer

swmmac_i32_16x16x64_iu4()

Return op name rocdl.swmmac.i32.16x16x64.iu4 as a bitstring.

swmmac_i32_16x16x64_iu4(ssa)

rocdl.swmmac.i32.16x16x64.iu4

Attributes

  • signA - Single, I1Attr, 1-bit signless integer attribute
  • signB - Single, I1Attr, 1-bit signless integer attribute
  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer

swmmac_i32_16x16x128_iu8()

Return op name rocdl.swmmac.i32.16x16x128.iu8 as a bitstring.

swmmac_i32_16x16x128_iu8(ssa)

rocdl.swmmac.i32.16x16x128.iu8

Attributes

  • signA - Single, I1Attr, 1-bit signless integer attribute
  • signB - Single, I1Attr, 1-bit signless integer attribute
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute
  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • index - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer

tanh()

Return op name rocdl.tanh as a bitstring.

tanh(ssa)

rocdl.tanh

Operands

  • arg - Single, LLVM_AnyFloat, floating point LLVM type

Results

  • res - Single, LLVM_AnyFloat, floating point LLVM type

Description

Note: In the general case, prefer the conventional arith, math, or llvm ops over this. Use this ROCDL-specific operation only when you fully understand its implication and when it is strictly necessary. This op is usually chosen when a small loss in precision is acceptable in exchange for higher execution speed.

Example:

%0 = rocdl.tanh %a f32 -> f32

tensor_load_to_lds()

Return op name rocdl.tensor.load.to.lds as a bitstring.

tensor_load_to_lds(ssa)

rocdl.tensor.load.to.lds - Base class for ROCDL tensor load/store to/from LDS.

Attributes

  • cachePolicy - Single, ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • dgroup0 - Single, ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4
  • dgroup1 - Single, ROCDL_V8I32Type, fixed-length vector of 32-bit signless integer values of length 8
  • dgroup2 - Single, ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4
  • dgroup3 - Single, ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4
  • dgroup4 - Single, ROCDL_V8I32Type, fixed-length vector of 32-bit signless integer values of length 8

Description

Moves tiles of tensor data between global memory and LDS. The tile is described by the $dgroup descriptors. 5 $dgroup descriptors allows for movement of up to 5D tensors. $cachePolicy describes the memory scope and an indicator of expected data re-use.

This op is for gfx1250+ architectures.

Example:

// Tensor load from global memory to LDS using 4 descriptor groups.
rocdl.tensor.load.to.lds %dg0, %dg1, %dg2, %dg3, %dg4, 0 : vector<4xi32>, vector<8xi32>

// Tensor store from LDS to global memory using 4 descriptor groups.
rocdl.tensor.store.from.lds %dg0, %dg1, %dg2, %dg3, %dg4, 0 : vector<4xi32>, vector<8xi32>

tensor_store_from_lds()

Return op name rocdl.tensor.store.from.lds as a bitstring.

tensor_store_from_lds(ssa)

rocdl.tensor.store.from.lds - Base class for ROCDL tensor load/store to/from LDS.

Attributes

  • cachePolicy - Single, ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attribute
  • alias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • noalias_scopes - Optional, LLVM_AliasScopeArrayAttr, LLVM dialect alias scope array
  • tbaa - Optional, LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array

Operands

  • dgroup0 - Single, ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4
  • dgroup1 - Single, ROCDL_V8I32Type, fixed-length vector of 32-bit signless integer values of length 8
  • dgroup2 - Single, ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4
  • dgroup3 - Single, ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4
  • dgroup4 - Single, ROCDL_V8I32Type, fixed-length vector of 32-bit signless integer values of length 8

Description

Moves tiles of tensor data between global memory and LDS. The tile is described by the $dgroup descriptors. 5 $dgroup descriptors allows for movement of up to 5D tensors. $cachePolicy describes the memory scope and an indicator of expected data re-use.

This op is for gfx1250+ architectures.

Example:

// Tensor load from global memory to LDS using 4 descriptor groups.
rocdl.tensor.load.to.lds %dg0, %dg1, %dg2, %dg3, %dg4, 0 : vector<4xi32>, vector<8xi32>

// Tensor store from LDS to global memory using 4 descriptor groups.
rocdl.tensor.store.from.lds %dg0, %dg1, %dg2, %dg3, %dg4, 0 : vector<4xi32>, vector<8xi32>

udot2()

Return op name rocdl.udot2 as a bitstring.

udot2(ssa)

rocdl.udot2

Attributes

  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2
  • b - Single, ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2
  • c - Single, anonymous/composite constraint, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit signless integer

Description

Packed intra-lane dot-product with optional result clamping (clamp). Computes res = sum_i a[i]*b[i] + c, where a and b hold packed 4/8/16-bit data (for dot2,dot4,dot8).

Example:

%r = rocdl.udot2 %a, %b, %c {clamp = true} :
     (vector<2xi16>, vector<2xi16>, i32) -> i32

udot4()

Return op name rocdl.udot4 as a bitstring.

udot4(ssa)

rocdl.udot4

Attributes

  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, 32-bit signless integer
  • b - Single, anonymous/composite constraint, 32-bit signless integer
  • c - Single, anonymous/composite constraint, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit signless integer

Description

Packed intra-lane dot-product with optional result clamping (clamp). Computes res = sum_i a[i]*b[i] + c, where a and b hold packed 4/8/16-bit data (for dot2,dot4,dot8).

Example:

%r = rocdl.udot4 %a, %b, %c {clamp = true} :
     (i32, i32, i32) -> i32

udot8()

Return op name rocdl.udot8 as a bitstring.

udot8(ssa)

rocdl.udot8

Attributes

  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, 32-bit signless integer
  • b - Single, anonymous/composite constraint, 32-bit signless integer
  • c - Single, anonymous/composite constraint, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit signless integer

Description

Packed intra-lane dot-product with optional result clamping (clamp). Computes res = sum_i a[i]*b[i] + c, where a and b hold packed 4/8/16-bit data (for dot2,dot4,dot8).

Example:

%r = rocdl.udot8 %a, %b, %c {clamp = true} :
     (i32, i32, i32) -> i32

update_dpp()

Return op name rocdl.update.dpp as a bitstring.

update_dpp(ssa)

rocdl.update.dpp

Attributes

  • dppCtrl - Single, I32Attr, 32-bit signless integer attribute
  • rowMask - Single, I32Attr, 32-bit signless integer attribute
  • bankMask - Single, I32Attr, 32-bit signless integer attribute
  • boundCtrl - Single, I1Attr, 1-bit signless integer attribute

Operands

  • old - Single, LLVM_Type, LLVM dialect-compatible type
  • src - Single, LLVM_Type, LLVM dialect-compatible type

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

wait_asyncmark()

Return op name rocdl.wait.asyncmark as a bitstring.

wait_asyncmark(ssa)

rocdl.wait.asyncmark - Wait until N or fewer async operation groups are unexecuted

Attributes

  • count - Single, I16Attr, 16-bit signless integer attribute

Description

This operation, along with rocdl.asyncmark, forms the compiler-provided framework for explicitly tracking asynchronous operations.

At the point where a wait.asyncmark operation is executed, all async operations that were parts of any async group (established by asyncmark in program order) other than the count previously-added ones will have finished executing.

For more detail, including on how this mechanism composes with function calls, see the LLVM documentation on async tracking.

Available on gfx9 and later.

Example:

// Wait until at most N async groups remain outstanding.
rocdl.wait.asyncmark 1

Usage example:

rocdl.tensor.load.to.lds ...
rocdl.global.async.load.to.lds ...

rocdl.asyncmark

rocdl.tensor.load.to.lds ...
rocdl.global.async.load.to.lds ...

rocdl.asyncmark

rocdl.wait.asyncmark 1 // First group of loads completes after this

wave_barrier()

Return op name rocdl.wave.barrier as a bitstring.

wave_barrier(ssa)

rocdl.wave.barrier

Description

Insert a wave-level (subgroup) barrier. Synchronizes lanes within a single wave/wavefront without any memory ordering guarantees.

Example:

rocdl.wave.barrier

wave_id()

Return op name rocdl.wave.id as a bitstring.

wave_id(ssa)

rocdl.wave.id

Attributes

  • range - Optional, LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Read a hardware register for thread/workgroup/cluster identification. An optional range attribute can constrain the returned value.

Example:

// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32

// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32

wavefrontsize()

Return op name rocdl.wavefrontsize as a bitstring.

wavefrontsize(ssa)

rocdl.wavefrontsize

Attributes

  • range - Optional, LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Read a hardware register for thread/workgroup/cluster identification. An optional range attribute can constrain the returned value.

Example:

// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32

// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32

wmma_bf16_16x16x16_bf16()

Return op name rocdl.wmma.bf16.16x16x16.bf16 as a bitstring.

wmma_bf16_16x16x16_bf16(ssa)

rocdl.wmma.bf16.16x16x16.bf16

Attributes

  • opsel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer

Results

  • res - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer

Description

Wave Matrix Multiply-Accumulate (WMMA) with output operand selection.

Example:

// WMMA f16 with opsel control.
%r = rocdl.wmma.f16.16x16x16.f16 %a, %b, %c {opsel = false} :
  (vector<16xf16>, vector<16xf16>, vector<16xf16>) -> vector<16xf16>

wmma_bf16_16x16x32_bf16()

Return op name rocdl.wmma.bf16.16x16x32.bf16 as a bitstring.

wmma_bf16_16x16x32_bf16(ssa)

rocdl.wmma.bf16.16x16x32.bf16

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
  • b - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
  • c - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type

Results

  • res - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_bf16f32_16x16x32_bf16()

Return op name rocdl.wmma.bf16f32.16x16x32.bf16 as a bitstring.

wmma_bf16f32_16x16x32_bf16(ssa)

rocdl.wmma.bf16f32.16x16x32.bf16

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
  • b - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type

Description

Wave Matrix Multiply-Accumulate (WMMA) with different C and D types.

Example:

// WMMA bf16 output from f32 accumulator with bf16 inputs.
%r = rocdl.wmma.bf16f32.16x16x32.bf16 %a, %b, %c, modC = none :
  (vector<16xbf16>, vector<16xbf16>, vector<8xf32>) -> vector<16xbf16>

wmma_f16_16x16x16_f16()

Return op name rocdl.wmma.f16.16x16x16.f16 as a bitstring.

wmma_f16_16x16x16_f16(ssa)

rocdl.wmma.f16.16x16x16.f16

Attributes

  • opsel - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • b - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • c - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Results

  • res - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with output operand selection.

Example:

// WMMA f16 with opsel control.
%r = rocdl.wmma.f16.16x16x16.f16 %a, %b, %c {opsel = false} :
  (vector<16xf16>, vector<16xf16>, vector<16xf16>) -> vector<16xf16>

wmma_f16_16x16x32_f16()

Return op name rocdl.wmma.f16.16x16x32.f16 as a bitstring.

wmma_f16_16x16x32_f16(ssa)

rocdl.wmma.f16.16x16x32.f16

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • b - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • c - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Results

  • res - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f16_16x16x64_bf8_bf8()

Return op name rocdl.wmma.f16.16x16x64.bf8_bf8 as a bitstring.

wmma_f16_16x16x64_bf8_bf8(ssa)

rocdl.wmma.f16.16x16x64.bf8_bf8

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Results

  • res - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f16_16x16x64_bf8_fp8()

Return op name rocdl.wmma.f16.16x16x64.bf8_fp8 as a bitstring.

wmma_f16_16x16x64_bf8_fp8(ssa)

rocdl.wmma.f16.16x16x64.bf8_fp8

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Results

  • res - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f16_16x16x64_fp8_bf8()

Return op name rocdl.wmma.f16.16x16x64.fp8_bf8 as a bitstring.

wmma_f16_16x16x64_fp8_bf8(ssa)

rocdl.wmma.f16.16x16x64.fp8_bf8

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Results

  • res - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f16_16x16x64_fp8_fp8()

Return op name rocdl.wmma.f16.16x16x64.fp8_fp8 as a bitstring.

wmma_f16_16x16x64_fp8_fp8(ssa)

rocdl.wmma.f16.16x16x64.fp8_fp8

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Results

  • res - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f16_16x16x128_bf8_bf8()

Return op name rocdl.wmma.f16.16x16x128.bf8_bf8 as a bitstring.

wmma_f16_16x16x128_bf8_bf8(ssa)

rocdl.wmma.f16.16x16x128.bf8_bf8

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Results

  • res - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f16_16x16x128_bf8_fp8()

Return op name rocdl.wmma.f16.16x16x128.bf8_fp8 as a bitstring.

wmma_f16_16x16x128_bf8_fp8(ssa)

rocdl.wmma.f16.16x16x128.bf8_fp8

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Results

  • res - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f16_16x16x128_fp8_bf8()

Return op name rocdl.wmma.f16.16x16x128.fp8_bf8 as a bitstring.

wmma_f16_16x16x128_fp8_bf8(ssa)

rocdl.wmma.f16.16x16x128.fp8_bf8

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Results

  • res - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f16_16x16x128_fp8_fp8()

Return op name rocdl.wmma.f16.16x16x128.fp8_fp8 as a bitstring.

wmma_f16_16x16x128_fp8_fp8(ssa)

rocdl.wmma.f16.16x16x128.fp8_fp8

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Results

  • res - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x4_f32()

Return op name rocdl.wmma.f32.16x16x4.f32 as a bitstring.

wmma_f32_16x16x4_f32(ssa)

rocdl.wmma.f32.16x16x4.f32

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
  • b - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x16_bf8_bf8()

Return op name rocdl.wmma.f32.16x16x16.bf8_bf8 as a bitstring.

wmma_f32_16x16x16_bf8_bf8(ssa)

rocdl.wmma.f32.16x16x16.bf8_bf8

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) intrinsic.

Example:

// WMMA with f16 inputs and f32 accumulator.
%r = rocdl.wmma.f32.16x16x16.f16 %a, %b, %c :
  (vector<16xf16>, vector<16xf16>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x16_bf8_fp8()

Return op name rocdl.wmma.f32.16x16x16.bf8_fp8 as a bitstring.

wmma_f32_16x16x16_bf8_fp8(ssa)

rocdl.wmma.f32.16x16x16.bf8_fp8

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) intrinsic.

Example:

// WMMA with f16 inputs and f32 accumulator.
%r = rocdl.wmma.f32.16x16x16.f16 %a, %b, %c :
  (vector<16xf16>, vector<16xf16>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x16_bf16()

Return op name rocdl.wmma.f32.16x16x16.bf16 as a bitstring.

wmma_f32_16x16x16_bf16(ssa)

rocdl.wmma.f32.16x16x16.bf16

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) intrinsic.

Example:

// WMMA with f16 inputs and f32 accumulator.
%r = rocdl.wmma.f32.16x16x16.f16 %a, %b, %c :
  (vector<16xf16>, vector<16xf16>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x16_f16()

Return op name rocdl.wmma.f32.16x16x16.f16 as a bitstring.

wmma_f32_16x16x16_f16(ssa)

rocdl.wmma.f32.16x16x16.f16

Operands

  • a - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • b - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) intrinsic.

Example:

// WMMA with f16 inputs and f32 accumulator.
%r = rocdl.wmma.f32.16x16x16.f16 %a, %b, %c :
  (vector<16xf16>, vector<16xf16>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x16_fp8_bf8()

Return op name rocdl.wmma.f32.16x16x16.fp8_bf8 as a bitstring.

wmma_f32_16x16x16_fp8_bf8(ssa)

rocdl.wmma.f32.16x16x16.fp8_bf8

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) intrinsic.

Example:

// WMMA with f16 inputs and f32 accumulator.
%r = rocdl.wmma.f32.16x16x16.f16 %a, %b, %c :
  (vector<16xf16>, vector<16xf16>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x16_fp8_fp8()

Return op name rocdl.wmma.f32.16x16x16.fp8_fp8 as a bitstring.

wmma_f32_16x16x16_fp8_fp8(ssa)

rocdl.wmma.f32.16x16x16.fp8_fp8

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) intrinsic.

Example:

// WMMA with f16 inputs and f32 accumulator.
%r = rocdl.wmma.f32.16x16x16.f16 %a, %b, %c :
  (vector<16xf16>, vector<16xf16>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x32_bf16()

Return op name rocdl.wmma.f32.16x16x32.bf16 as a bitstring.

wmma_f32_16x16x32_bf16(ssa)

rocdl.wmma.f32.16x16x32.bf16

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
  • b - Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x32_f16()

Return op name rocdl.wmma.f32.16x16x32.f16 as a bitstring.

wmma_f32_16x16x32_f16(ssa)

rocdl.wmma.f32.16x16x32.f16

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • b - Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x64_bf8_bf8()

Return op name rocdl.wmma.f32.16x16x64.bf8_bf8 as a bitstring.

wmma_f32_16x16x64_bf8_bf8(ssa)

rocdl.wmma.f32.16x16x64.bf8_bf8

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x64_bf8_fp8()

Return op name rocdl.wmma.f32.16x16x64.bf8_fp8 as a bitstring.

wmma_f32_16x16x64_bf8_fp8(ssa)

rocdl.wmma.f32.16x16x64.bf8_fp8

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x64_fp8_bf8()

Return op name rocdl.wmma.f32.16x16x64.fp8_bf8 as a bitstring.

wmma_f32_16x16x64_fp8_bf8(ssa)

rocdl.wmma.f32.16x16x64.fp8_bf8

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x64_fp8_fp8()

Return op name rocdl.wmma.f32.16x16x64.fp8_fp8 as a bitstring.

wmma_f32_16x16x64_fp8_fp8(ssa)

rocdl.wmma.f32.16x16x64.fp8_fp8

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x128_bf8_bf8()

Return op name rocdl.wmma.f32.16x16x128.bf8_bf8 as a bitstring.

wmma_f32_16x16x128_bf8_bf8(ssa)

rocdl.wmma.f32.16x16x128.bf8_bf8

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x128_bf8_fp8()

Return op name rocdl.wmma.f32.16x16x128.bf8_fp8 as a bitstring.

wmma_f32_16x16x128_bf8_fp8(ssa)

rocdl.wmma.f32.16x16x128.bf8_fp8

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x128_fp8_bf8()

Return op name rocdl.wmma.f32.16x16x128.fp8_bf8 as a bitstring.

wmma_f32_16x16x128_fp8_bf8(ssa)

rocdl.wmma.f32.16x16x128.fp8_bf8

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_f32_16x16x128_fp8_fp8()

Return op name rocdl.wmma.f32.16x16x128.fp8_fp8 as a bitstring.

wmma_f32_16x16x128_fp8_fp8(ssa)

rocdl.wmma.f32.16x16x128.fp8_fp8

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.

Example:

// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>

wmma_i32_16x16x16_iu4()

Return op name rocdl.wmma.i32.16x16x16.iu4 as a bitstring.

wmma_i32_16x16x16_iu4(ssa)

rocdl.wmma.i32.16x16x16.iu4

Attributes

  • signA - Single, I1Attr, 1-bit signless integer attribute
  • signB - Single, I1Attr, 1-bit signless integer attribute
  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer

Results

  • res - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer

Description

Wave Matrix Multiply-Accumulate (WMMA) for integer types with sign and clamp control.

Example:

// WMMA i32 with unsigned i8 inputs.
%r = rocdl.wmma.i32.16x16x16.iu8 %a, %b, %c
  {signA = false, signB = false, clamp = false} :
  (vector<4xi32>, vector<4xi32>, vector<8xi32>) -> vector<8xi32>

wmma_i32_16x16x16_iu8()

Return op name rocdl.wmma.i32.16x16x16.iu8 as a bitstring.

wmma_i32_16x16x16_iu8(ssa)

rocdl.wmma.i32.16x16x16.iu8

Attributes

  • signA - Single, I1Attr, 1-bit signless integer attribute
  • signB - Single, I1Attr, 1-bit signless integer attribute
  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer

Results

  • res - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer

Description

Wave Matrix Multiply-Accumulate (WMMA) for integer types with sign and clamp control.

Example:

// WMMA i32 with unsigned i8 inputs.
%r = rocdl.wmma.i32.16x16x16.iu8 %a, %b, %c
  {signA = false, signB = false, clamp = false} :
  (vector<4xi32>, vector<4xi32>, vector<8xi32>) -> vector<8xi32>

wmma_i32_16x16x32_iu4()

Return op name rocdl.wmma.i32.16x16x32.iu4 as a bitstring.

wmma_i32_16x16x32_iu4(ssa)

rocdl.wmma.i32.16x16x32.iu4

Attributes

  • signA - Single, I1Attr, 1-bit signless integer attribute
  • signB - Single, I1Attr, 1-bit signless integer attribute
  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer

Results

  • res - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer

Description

Wave Matrix Multiply-Accumulate (WMMA) for integer types with sign and clamp control.

Example:

// WMMA i32 with unsigned i8 inputs.
%r = rocdl.wmma.i32.16x16x16.iu8 %a, %b, %c
  {signA = false, signB = false, clamp = false} :
  (vector<4xi32>, vector<4xi32>, vector<8xi32>) -> vector<8xi32>

wmma_i32_16x16x64_iu8()

Return op name rocdl.wmma.i32.16x16x64.iu8 as a bitstring.

wmma_i32_16x16x64_iu8(ssa)

rocdl.wmma.i32.16x16x64.iu8

Attributes

  • signA - Single, I1Attr, 1-bit signless integer attribute
  • signB - Single, I1Attr, 1-bit signless integer attribute
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute
  • clamp - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer

Results

  • res - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer

Description

Wave Matrix Multiply-Accumulate (WMMA) for integer types with sign, reuse, and clamp controls.

Example:

// WMMA i32 with unsigned i8 inputs and reuse controls.
%r = rocdl.wmma.i32.16x16x64.iu8 %a, %b, %c
  {signA = false, signB = false, reuseA = false, reuseB = false, clamp = false} :
  (vector<8xi32>, vector<8xi32>, vector<8xi32>) -> vector<8xi32>

wmma_scale16_f32_16x16x128_f8f6f4()

Return op name rocdl.wmma.scale16.f32.16x16x128.f8f6f4 as a bitstring.

wmma_scale16_f32_16x16x128_f8f6f4(ssa)

rocdl.wmma.scale16.f32.16x16x128.f8f6f4

Attributes

  • fmtA - Single, ROCDL_MatrixFormatAttr, matrix operand formats selected by scaled MFMA/WMMA format fields
  • fmtB - Single, ROCDL_MatrixFormatAttr, matrix operand formats selected by scaled MFMA/WMMA format fields
  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • scaleAType - Single, ROCDL_WMMAMatrixScaleAttr, matrix scale row selector
  • fmtScaleA - Single, ROCDL_WMMAMatrixScaleFormatAttr, matrix scale exponent formats
  • scaleBType - Single, ROCDL_WMMAMatrixScaleAttr, matrix scale row selector
  • fmtScaleB - Single, ROCDL_WMMAMatrixScaleFormatAttr, matrix scale exponent formats
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
  • scaleA - Single, I64, 64-bit signless integer
  • scaleB - Single, I64, 64-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Scaled Wave Matrix Multiply-Accumulate (WMMA) with per-operand scaling.

Example:

// Scaled WMMA with f8f6f4 format inputs.
%r = rocdl.wmma.scale.f32.16x16x128.f8f6f4 %a, %b, %c, %scaleA, %scaleB
  fmtA = fp8_e4m3, fmtB = fp8_e4m3, modC = none,
  scaleAType = row0, fmtScaleA = e8, scaleBType = row0, fmtScaleB = e8 :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>, i32, i32) -> vector<8xf32>

wmma_scale16_f32_32x16x128_f4()

Return op name rocdl.wmma.scale16.f32.32x16x128.f4 as a bitstring.

wmma_scale16_f32_32x16x128_f4(ssa)

rocdl.wmma.scale16.f32.32x16x128.f4

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • scaleAType - Single, ROCDL_WMMAMatrixScaleAttr, matrix scale row selector
  • fmtScaleA - Single, ROCDL_WMMAMatrixScaleFormatAttr, matrix scale exponent formats
  • scaleBType - Single, ROCDL_WMMAMatrixScaleAttr, matrix scale row selector
  • fmtScaleB - Single, ROCDL_WMMAMatrixScaleFormatAttr, matrix scale exponent formats
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
  • scaleA - Single, I64, 64-bit signless integer
  • scaleB - Single, I64, 64-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Scaled Wave Matrix Multiply-Accumulate (WMMA) for F4 format inputs.

Example:

// Scaled WMMA with f4 format inputs.
%r = rocdl.wmma.scale.f32.16x16x128.f4 %a, %b, %c, %scaleA, %scaleB
  modC = none, scaleAType = row0, fmtScaleA = e8,
  scaleBType = row0, fmtScaleB = e8 :
  (vector<8xi32>, vector<8xi32>, vector<8xf32>, i32, i32) -> vector<8xf32>

wmma_scale_f32_16x16x128_f8f6f4()

Return op name rocdl.wmma.scale.f32.16x16x128.f8f6f4 as a bitstring.

wmma_scale_f32_16x16x128_f8f6f4(ssa)

rocdl.wmma.scale.f32.16x16x128.f8f6f4

Attributes

  • fmtA - Single, ROCDL_MatrixFormatAttr, matrix operand formats selected by scaled MFMA/WMMA format fields
  • fmtB - Single, ROCDL_MatrixFormatAttr, matrix operand formats selected by scaled MFMA/WMMA format fields
  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • scaleAType - Single, ROCDL_WMMAMatrixScaleAttr, matrix scale row selector
  • fmtScaleA - Single, ROCDL_WMMAMatrixScaleFormatAttr, matrix scale exponent formats
  • scaleBType - Single, ROCDL_WMMAMatrixScaleAttr, matrix scale row selector
  • fmtScaleB - Single, ROCDL_WMMAMatrixScaleFormatAttr, matrix scale exponent formats
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
  • scaleA - Single, I32, 32-bit signless integer
  • scaleB - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Scaled Wave Matrix Multiply-Accumulate (WMMA) with per-operand scaling.

Example:

// Scaled WMMA with f8f6f4 format inputs.
%r = rocdl.wmma.scale.f32.16x16x128.f8f6f4 %a, %b, %c, %scaleA, %scaleB
  fmtA = fp8_e4m3, fmtB = fp8_e4m3, modC = none,
  scaleAType = row0, fmtScaleA = e8, scaleBType = row0, fmtScaleB = e8 :
  (vector<16xi32>, vector<16xi32>, vector<8xf32>, i32, i32) -> vector<8xf32>

wmma_scale_f32_32x16x128_f4()

Return op name rocdl.wmma.scale.f32.32x16x128.f4 as a bitstring.

wmma_scale_f32_32x16x128_f4(ssa)

rocdl.wmma.scale.f32.32x16x128.f4

Attributes

  • modC - Single, ROCDL_WMMACModifierAttr, WMMA C operand modifiers
  • scaleAType - Single, ROCDL_WMMAMatrixScaleAttr, matrix scale row selector
  • fmtScaleA - Single, ROCDL_WMMAMatrixScaleFormatAttr, matrix scale exponent formats
  • scaleBType - Single, ROCDL_WMMAMatrixScaleAttr, matrix scale row selector
  • fmtScaleB - Single, ROCDL_WMMAMatrixScaleFormatAttr, matrix scale exponent formats
  • reuseA - Single, I1Attr, 1-bit signless integer attribute
  • reuseB - Single, I1Attr, 1-bit signless integer attribute

Operands

  • a - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • b - Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
  • c - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
  • scaleA - Single, I32, 32-bit signless integer
  • scaleB - Single, I32, 32-bit signless integer

Results

  • res - Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float

Description

Scaled Wave Matrix Multiply-Accumulate (WMMA) for F4 format inputs.

Example:

// Scaled WMMA with f4 format inputs.
%r = rocdl.wmma.scale.f32.16x16x128.f4 %a, %b, %c, %scaleA, %scaleB
  modC = none, scaleAType = row0, fmtScaleA = e8,
  scaleBType = row0, fmtScaleB = e8 :
  (vector<8xi32>, vector<8xi32>, vector<8xf32>, i32, i32) -> vector<8xf32>

workgroup_id_x()

Return op name rocdl.workgroup.id.x as a bitstring.

workgroup_id_x(ssa)

rocdl.workgroup.id.x

Attributes

  • range - Optional, LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Read a hardware register for thread/workgroup/cluster identification. An optional range attribute can constrain the returned value.

Example:

// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32

// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32

workgroup_id_y()

Return op name rocdl.workgroup.id.y as a bitstring.

workgroup_id_y(ssa)

rocdl.workgroup.id.y

Attributes

  • range - Optional, LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Read a hardware register for thread/workgroup/cluster identification. An optional range attribute can constrain the returned value.

Example:

// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32

// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32

workgroup_id_z()

Return op name rocdl.workgroup.id.z as a bitstring.

workgroup_id_z(ssa)

rocdl.workgroup.id.z

Attributes

  • range - Optional, LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Read a hardware register for thread/workgroup/cluster identification. An optional range attribute can constrain the returned value.

Example:

// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32

// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32

workitem_id_x()

Return op name rocdl.workitem.id.x as a bitstring.

workitem_id_x(ssa)

rocdl.workitem.id.x

Attributes

  • range - Optional, LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Read a hardware register for thread/workgroup/cluster identification. An optional range attribute can constrain the returned value.

Example:

// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32

// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32

workitem_id_y()

Return op name rocdl.workitem.id.y as a bitstring.

workitem_id_y(ssa)

rocdl.workitem.id.y

Attributes

  • range - Optional, LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Read a hardware register for thread/workgroup/cluster identification. An optional range attribute can constrain the returned value.

Example:

// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32

// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32

workitem_id_z()

Return op name rocdl.workitem.id.z as a bitstring.

workitem_id_z(ssa)

rocdl.workitem.id.z

Attributes

  • range - Optional, LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange

Results

  • res - Single, LLVM_Type, LLVM dialect-compatible type

Description

Read a hardware register for thread/workgroup/cluster identification. An optional range attribute can constrain the returned value.

Example:

// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32

// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32