Beaver. MLIR. Dialect. ROCDL
(beaver v0.4.8)
Copy Markdown
Summary
Functions
Return op name rocdl.asyncmark as a bitstring.
rocdl.asyncmark - Mark the end of a group of asynchronous operations
Return op name rocdl.ballot as a bitstring.
rocdl.ballot - Vote across thread group
Return op name rocdl.barrier as a bitstring.
rocdl.barrier
Return op name rocdl.cluster.id.x as a bitstring.
rocdl.cluster.id.x
Return op name rocdl.cluster.id.y as a bitstring.
rocdl.cluster.id.y
Return op name rocdl.cluster.id.z as a bitstring.
rocdl.cluster.id.z
Return op name rocdl.cluster.load.async.to.lds.b8 as a bitstring.
rocdl.cluster.load.async.to.lds.b8
Return op name rocdl.cluster.load.async.to.lds.b32 as a bitstring.
rocdl.cluster.load.async.to.lds.b32
Return op name rocdl.cluster.load.async.to.lds.b64 as a bitstring.
rocdl.cluster.load.async.to.lds.b64
Return op name rocdl.cluster.load.async.to.lds.b128 as a bitstring.
rocdl.cluster.load.async.to.lds.b128
Return op name rocdl.cluster.workgroup.id.x as a bitstring.
rocdl.cluster.workgroup.id.x
Return op name rocdl.cluster.workgroup.id.y as a bitstring.
rocdl.cluster.workgroup.id.y
Return op name rocdl.cluster.workgroup.id.z as a bitstring.
rocdl.cluster.workgroup.id.z
Return op name rocdl.cos as a bitstring.
rocdl.cos
Return op name rocdl.cvt.f32.bf8 as a bitstring.
rocdl.cvt.f32.bf8 - Convert bf8 to f32
Return op name rocdl.cvt.f32.fp8 as a bitstring.
rocdl.cvt.f32.fp8 - Convert fp8 to f32
Return op name rocdl.cvt.pk.bf8.f32 as a bitstring.
rocdl.cvt.pk.bf8.f32 - Convert two f32's to bf8
Return op name rocdl.cvt.pk.f32.bf8 as a bitstring.
rocdl.cvt.pk.f32.bf8 - Convert packed bf8 to packed f32
Return op name rocdl.cvt.pk.f32.fp8 as a bitstring.
rocdl.cvt.pk.f32.fp8 - Convert packed fp8 to packed f32
Return op name rocdl.cvt.pk.fp8.f32 as a bitstring.
rocdl.cvt.pk.fp8.f32 - Convert two f32's to fp8
Return op name rocdl.cvt.pkrtz as a bitstring.
rocdl.cvt.pkrtz - Convert two f32 input into a vector<2xf16>
Return op name rocdl.cvt.scale.pk8.bf16.bf8 as a bitstring.
rocdl.cvt.scale.pk8.bf16.bf8 - Scales 8 bf8 and converts them to 8 bf16.
Return op name rocdl.cvt.scale.pk8.bf16.fp4 as a bitstring.
rocdl.cvt.scale.pk8.bf16.fp4 - Scales 8 fp4 and converts them to 8 bf16.
Return op name rocdl.cvt.scale.pk8.bf16.fp8 as a bitstring.
rocdl.cvt.scale.pk8.bf16.fp8 - Scales 8 fp8 and converts them to 8 bf16.
Return op name rocdl.cvt.scale.pk8.f16.bf8 as a bitstring.
rocdl.cvt.scale.pk8.f16.bf8 - Scales 8 bf8 and converts them to 8 f16.
Return op name rocdl.cvt.scale.pk8.f16.fp4 as a bitstring.
rocdl.cvt.scale.pk8.f16.fp4 - Scales 8 fp4 and converts them to 8 f16.
Return op name rocdl.cvt.scale.pk8.f16.fp8 as a bitstring.
rocdl.cvt.scale.pk8.f16.fp8 - Scales 8 fp8 and converts them to 8 f16.
Return op name rocdl.cvt.scale.pk8.f32.bf8 as a bitstring.
rocdl.cvt.scale.pk8.f32.bf8 - Scales 8 bf8 and converts them to 8 f32.
Return op name rocdl.cvt.scale.pk8.f32.fp4 as a bitstring.
rocdl.cvt.scale.pk8.f32.fp4 - Scales 8 fp4 and converts them to 8 f32.
Return op name rocdl.cvt.scale.pk8.f32.fp8 as a bitstring.
rocdl.cvt.scale.pk8.f32.fp8 - Scales 8 fp8 and converts them to 8 f32.
Return op name rocdl.cvt.scale.pk16.bf16.bf6 as a bitstring.
rocdl.cvt.scale.pk16.bf16.bf6 - Scales 16 bf6 and converts them to 16 bf16.
Return op name rocdl.cvt.scale.pk16.bf16.fp6 as a bitstring.
rocdl.cvt.scale.pk16.bf16.fp6 - Scales 16 fp6 and converts them to 16 bf16.
Return op name rocdl.cvt.scale.pk16.f16.bf6 as a bitstring.
rocdl.cvt.scale.pk16.f16.bf6 - Scales 16 bf6 and converts them to 16 f16.
Return op name rocdl.cvt.scale.pk16.f16.fp6 as a bitstring.
rocdl.cvt.scale.pk16.f16.fp6 - Scales 16 fp6 and converts them to 16 f16.
Return op name rocdl.cvt.scale.pk16.f32.bf6 as a bitstring.
rocdl.cvt.scale.pk16.f32.bf6 - Scales 16 bf6 and converts them to 16 f32.
Return op name rocdl.cvt.scale.pk16.f32.fp6 as a bitstring.
rocdl.cvt.scale.pk16.f32.fp6 - Scales 16 fp6 and converts them to 16 f32.
Return op name rocdl.cvt.scalef32.2xpk16.bf6.f32 as a bitstring.
rocdl.cvt.scalef32.2xpk16.bf6.f32 - Scale and convert two vector<16xf32> to 32 packed bf6
Return op name rocdl.cvt.scalef32.2xpk16.fp6.f32 as a bitstring.
rocdl.cvt.scalef32.2xpk16.fp6.f32 - Scale and convert two vector<16xf32> to 32 packed fp6
Return op name rocdl.cvt.scalef32.f16.bf8 as a bitstring.
rocdl.cvt.scalef32.f16.bf8 - Scaled convert bf8 from packed vector to f16, updating tied result
Return op name rocdl.cvt.scalef32.f16.fp8 as a bitstring.
rocdl.cvt.scalef32.f16.fp8 - Scaled convert fp8 from packed vector to f16, updating tied result
Return op name rocdl.cvt.scalef32.f32.bf8 as a bitstring.
rocdl.cvt.scalef32.f32.bf8 - Scaled convert bf8 from packed vector to f32
Return op name rocdl.cvt.scalef32.f32.fp8 as a bitstring.
rocdl.cvt.scalef32.f32.fp8 - Scaled convert fp8 from packed vector to f32
Return op name rocdl.cvt.scalef32.pk8.bf8.bf16 as a bitstring.
rocdl.cvt.scalef32.pk8.bf8.bf16 - Scale and convert packed bf16 to packed bf8
Return op name rocdl.cvt.scalef32.pk8.bf8.f16 as a bitstring.
rocdl.cvt.scalef32.pk8.bf8.f16 - Scale and convert packed f16 to packed bf8
Return op name rocdl.cvt.scalef32.pk8.bf8.f32 as a bitstring.
rocdl.cvt.scalef32.pk8.bf8.f32 - Scale and convert packed f32 to packed bf8
Return op name rocdl.cvt.scalef32.pk8.fp4.bf16 as a bitstring.
rocdl.cvt.scalef32.pk8.fp4.bf16 - Scale and convert packed bf16 to packed fp4
Return op name rocdl.cvt.scalef32.pk8.fp4.f16 as a bitstring.
rocdl.cvt.scalef32.pk8.fp4.f16 - Scale and convert packed f16 to packed fp4
Return op name rocdl.cvt.scalef32.pk8.fp4.f32 as a bitstring.
rocdl.cvt.scalef32.pk8.fp4.f32 - Scale and convert packed f32 to packed fp4
Return op name rocdl.cvt.scalef32.pk8.fp8.bf16 as a bitstring.
rocdl.cvt.scalef32.pk8.fp8.bf16 - Scale and convert packed bf16 to packed fp8
Return op name rocdl.cvt.scalef32.pk8.fp8.f16 as a bitstring.
rocdl.cvt.scalef32.pk8.fp8.f16 - Scale and convert packed f16 to packed fp8
Return op name rocdl.cvt.scalef32.pk8.fp8.f32 as a bitstring.
rocdl.cvt.scalef32.pk8.fp8.f32 - Scale and convert packed f32 to packed fp8
Return op name rocdl.cvt.scalef32.pk16.bf6.bf16 as a bitstring.
rocdl.cvt.scalef32.pk16.bf6.bf16 - Scale and convert packed bf16 to packed bf6
Return op name rocdl.cvt.scalef32.pk16.bf6.f16 as a bitstring.
rocdl.cvt.scalef32.pk16.bf6.f16 - Scale and convert packed f16 to packed bf6
Return op name rocdl.cvt.scalef32.pk16.bf6.f32 as a bitstring.
rocdl.cvt.scalef32.pk16.bf6.f32 - Scale and convert packed f32 to packed bf6
Return op name rocdl.cvt.scalef32.pk16.fp6.bf16 as a bitstring.
rocdl.cvt.scalef32.pk16.fp6.bf16 - Scale and convert packed bf16 to packed fp6
Return op name rocdl.cvt.scalef32.pk16.fp6.f16 as a bitstring.
rocdl.cvt.scalef32.pk16.fp6.f16 - Scale and convert packed f16 to packed fp6
Return op name rocdl.cvt.scalef32.pk16.fp6.f32 as a bitstring.
rocdl.cvt.scalef32.pk16.fp6.f32 - Scale and convert packed f32 to packed fp6
Return op name rocdl.cvt.scalef32.pk32.bf6.bf16 as a bitstring.
rocdl.cvt.scalef32.pk32.bf6.bf16 - Scale and convert packed bf16 to packed bf6
Return op name rocdl.cvt.scalef32.pk32.bf6.f16 as a bitstring.
rocdl.cvt.scalef32.pk32.bf6.f16 - Scale and convert packed f16 to packed bf6
Return op name rocdl.cvt.scalef32.pk32.bf16.bf6 as a bitstring.
rocdl.cvt.scalef32.pk32.bf16.bf6 - Scale and convert packed bf6 to packed bf16
Return op name rocdl.cvt.scalef32.pk32.bf16.fp6 as a bitstring.
rocdl.cvt.scalef32.pk32.bf16.fp6 - Scale and convert packed fp6 to packed bf16
Return op name rocdl.cvt.scalef32.pk32.f16.bf6 as a bitstring.
rocdl.cvt.scalef32.pk32.f16.bf6 - Scale and convert packed bf6 to packed f16
Return op name rocdl.cvt.scalef32.pk32.f16.fp6 as a bitstring.
rocdl.cvt.scalef32.pk32.f16.fp6 - Scale and convert packed fp6 to packed f16
Return op name rocdl.cvt.scalef32.pk32.f32.bf6 as a bitstring.
rocdl.cvt.scalef32.pk32.f32.bf6 - Scale and convert packed bf6 to packed f32
Return op name rocdl.cvt.scalef32.pk32.f32.fp6 as a bitstring.
rocdl.cvt.scalef32.pk32.f32.fp6 - Scale and convert packed fp6 to packed f32
Return op name rocdl.cvt.scalef32.pk32.fp6.bf16 as a bitstring.
rocdl.cvt.scalef32.pk32.fp6.bf16 - Scale and convert packed bf16 to packed fp6
Return op name rocdl.cvt.scalef32.pk32.fp6.f16 as a bitstring.
rocdl.cvt.scalef32.pk32.fp6.f16 - Scale and convert packed f16 to packed fp6
Return op name rocdl.cvt.scalef32.pk.bf8.bf16 as a bitstring.
rocdl.cvt.scalef32.pk.bf8.bf16 - Scaled convert two bf16to two bf8, updating packed vector
Return op name rocdl.cvt.scalef32.pk.bf8.f16 as a bitstring.
rocdl.cvt.scalef32.pk.bf8.f16 - Scaled convert two f16to two bf8, updating packed vector
Return op name rocdl.cvt.scalef32.pk.bf8.f32 as a bitstring.
rocdl.cvt.scalef32.pk.bf8.f32 - Scaled convert two f32 to two bf8, updating packed vector
Return op name rocdl.cvt.scalef32.pk.bf16.bf8 as a bitstring.
rocdl.cvt.scalef32.pk.bf16.bf8 - Scaled convert two bf8to two bf16
Return op name rocdl.cvt.scalef32.pk.bf16.fp4 as a bitstring.
rocdl.cvt.scalef32.pk.bf16.fp4 - Scale and convert two packed fp4 to packed bf16
Return op name rocdl.cvt.scalef32.pk.bf16.fp8 as a bitstring.
rocdl.cvt.scalef32.pk.bf16.fp8 - Scaled convert two fp8to two bf16
Return op name rocdl.cvt.scalef32.pk.f16.bf8 as a bitstring.
rocdl.cvt.scalef32.pk.f16.bf8 - Scaled convert two bf8to two f16
Return op name rocdl.cvt.scalef32.pk.f16.fp4 as a bitstring.
rocdl.cvt.scalef32.pk.f16.fp4 - Scale and convert two packed fp4 to packed f16
Return op name rocdl.cvt.scalef32.pk.f16.fp8 as a bitstring.
rocdl.cvt.scalef32.pk.f16.fp8 - Scaled convert two fp8to two f16
Return op name rocdl.cvt.scalef32.pk.f32.bf8 as a bitstring.
rocdl.cvt.scalef32.pk.f32.bf8 - Scaled convert two bf8to two f32
Return op name rocdl.cvt.scalef32.pk.f32.fp4 as a bitstring.
rocdl.cvt.scalef32.pk.f32.fp4 - Scale and convert two packed fp4 to packed f32
Return op name rocdl.cvt.scalef32.pk.f32.fp8 as a bitstring.
rocdl.cvt.scalef32.pk.f32.fp8 - Scaled convert two fp8to two f32
Return op name rocdl.cvt.scalef32.pk.fp4.bf16 as a bitstring.
rocdl.cvt.scalef32.pk.fp4.bf16 - Scale and convert two bf16 to packed fp4, updating tied vector
Return op name rocdl.cvt.scalef32.pk.fp4.f16 as a bitstring.
rocdl.cvt.scalef32.pk.fp4.f16 - Scale and convert two f16 to packed fp4, updating tied vector
Return op name rocdl.cvt.scalef32.pk.fp4.f32 as a bitstring.
rocdl.cvt.scalef32.pk.fp4.f32 - Scale and convert two f32 values to two packed fp4, updating tied vector
Return op name rocdl.cvt.scalef32.pk.fp8.bf16 as a bitstring.
rocdl.cvt.scalef32.pk.fp8.bf16 - Scaled convert two bf16to two fp8, updating packed vector
Return op name rocdl.cvt.scalef32.pk.fp8.f16 as a bitstring.
rocdl.cvt.scalef32.pk.fp8.f16 - Scaled convert two f16to two fp8, updating packed vector
Return op name rocdl.cvt.scalef32.pk.fp8.f32 as a bitstring.
rocdl.cvt.scalef32.pk.fp8.f32 - Scaled convert two f32 to two fp8, updating packed vector
Return op name rocdl.cvt.scalef32.sr.bf8.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.bf8.bf16 - Scaled convert bf16to bf8 with stochiastic rounding, updating packed vector
Return op name rocdl.cvt.scalef32.sr.bf8.f16 as a bitstring.
rocdl.cvt.scalef32.sr.bf8.f16 - Scaled convert f16to bf8 with stochiastic rounding, updating packed vector
Return op name rocdl.cvt.scalef32.sr.bf8.f32 as a bitstring.
rocdl.cvt.scalef32.sr.bf8.f32 - Scaled convert f32to bf8 with stochiastic rounding, updating packed vector
Return op name rocdl.cvt.scalef32.sr.fp8.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.fp8.bf16 - Scaled convert bf16to fp8 with stochiastic rounding, updating packed vector
Return op name rocdl.cvt.scalef32.sr.fp8.f16 as a bitstring.
rocdl.cvt.scalef32.sr.fp8.f16 - Scaled convert f16to fp8 with stochiastic rounding, updating packed vector
Return op name rocdl.cvt.scalef32.sr.fp8.f32 as a bitstring.
rocdl.cvt.scalef32.sr.fp8.f32 - Scaled convert f32to fp8 with stochiastic rounding, updating packed vector
Return op name rocdl.cvt.scalef32.sr.pk8.bf8.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.bf8.bf16 - Scale and convert packed bf16 to packed bf8 with stochastic rounding
Return op name rocdl.cvt.scalef32.sr.pk8.bf8.f16 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.bf8.f16 - Scale and convert packed f16 to packed bf8 with stochastic rounding
Return op name rocdl.cvt.scalef32.sr.pk8.bf8.f32 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.bf8.f32 - Scale and convert packed f32 to packed bf8 with stochastic rounding
Return op name rocdl.cvt.scalef32.sr.pk8.fp4.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.fp4.bf16 - Scale and convert packed bf16 to packed fp4 with stochastic rounding
Return op name rocdl.cvt.scalef32.sr.pk8.fp4.f16 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.fp4.f16 - Scale and convert packed f16 to packed fp4 with stochastic rounding
Return op name rocdl.cvt.scalef32.sr.pk8.fp4.f32 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.fp4.f32 - Scale and convert packed f32 to packed fp4 with stochastic rounding
Return op name rocdl.cvt.scalef32.sr.pk8.fp8.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.fp8.bf16 - Scale and convert packed bf16 to packed fp8 with stochastic rounding
Return op name rocdl.cvt.scalef32.sr.pk8.fp8.f16 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.fp8.f16 - Scale and convert packed f16 to packed fp8 with stochastic rounding
Return op name rocdl.cvt.scalef32.sr.pk8.fp8.f32 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.fp8.f32 - Scale and convert packed f32 to packed fp8 with stochastic rounding
Return op name rocdl.cvt.scalef32.sr.pk16.bf6.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.pk16.bf6.bf16 - Scale and convert packed bf16 to packed bf6 with stochastic rounding
Return op name rocdl.cvt.scalef32.sr.pk16.bf6.f16 as a bitstring.
rocdl.cvt.scalef32.sr.pk16.bf6.f16 - Scale and convert packed f16 to packed bf6 with stochastic rounding
Return op name rocdl.cvt.scalef32.sr.pk16.bf6.f32 as a bitstring.
rocdl.cvt.scalef32.sr.pk16.bf6.f32 - Scale and convert packed f32 to packed bf6 with stochastic rounding
Return op name rocdl.cvt.scalef32.sr.pk16.fp6.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.pk16.fp6.bf16 - Scale and convert packed bf16 to packed fp6 with stochastic rounding
Return op name rocdl.cvt.scalef32.sr.pk16.fp6.f16 as a bitstring.
rocdl.cvt.scalef32.sr.pk16.fp6.f16 - Scale and convert packed f16 to packed fp6 with stochastic rounding
Return op name rocdl.cvt.scalef32.sr.pk16.fp6.f32 as a bitstring.
rocdl.cvt.scalef32.sr.pk16.fp6.f32 - Scale and convert packed f32 to packed fp6 with stochastic rounding
Return op name rocdl.cvt.scalef32.sr.pk32.bf6.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.pk32.bf6.bf16 - Scale and convert packed bf16 to packed bf6 with stochiastic rounding
Return op name rocdl.cvt.scalef32.sr.pk32.bf6.f16 as a bitstring.
rocdl.cvt.scalef32.sr.pk32.bf6.f16 - Scale and convert packed f16 to packed bf6 with stochiastic rounding
Return op name rocdl.cvt.scalef32.sr.pk32.bf6.f32 as a bitstring.
rocdl.cvt.scalef32.sr.pk32.bf6.f32 - Scale and convert packed f32 to packed bf6 with stochiastic rounding
Return op name rocdl.cvt.scalef32.sr.pk32.fp6.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.pk32.fp6.bf16 - Scale and convert packed bf16 to packed fp6 with stochiastic rounding
Return op name rocdl.cvt.scalef32.sr.pk32.fp6.f16 as a bitstring.
rocdl.cvt.scalef32.sr.pk32.fp6.f16 - Scale and convert packed f16 to packed fp6 with stochiastic rounding
Return op name rocdl.cvt.scalef32.sr.pk32.fp6.f32 as a bitstring.
rocdl.cvt.scalef32.sr.pk32.fp6.f32 - Scale and convert packed f32 to packed fp6 with stochiastic rounding
Return op name rocdl.cvt.scalef32.sr.pk.fp4.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.pk.fp4.bf16 - Scale and convert two bf16 to packed fp4 with stochiastic rounding, updating tied vector
Return op name rocdl.cvt.scalef32.sr.pk.fp4.f16 as a bitstring.
rocdl.cvt.scalef32.sr.pk.fp4.f16 - Scale and convert two f16 to packed fp4 with stochiastic rounding, updating tied vector
Return op name rocdl.cvt.scalef32.sr.pk.fp4.f32 as a bitstring.
rocdl.cvt.scalef32.sr.pk.fp4.f32 - Scale and convert two f32 to packed fp4 with stochiastic rounding, updating tied vector
Return op name rocdl.cvt.sr.bf8.f32 as a bitstring.
rocdl.cvt.sr.bf8.f32 - Convert f32 to bf8, stochiastic rounding
Return op name rocdl.cvt.sr.fp8.f32 as a bitstring.
rocdl.cvt.sr.fp8.f32 - Convert f32 to fp8, stochiastic rounding
Return op name rocdl.dot4.f32.bf8.bf8 as a bitstring.
rocdl.dot4.f32.bf8.bf8
Return op name rocdl.dot4.f32.bf8.fp8 as a bitstring.
rocdl.dot4.f32.bf8.fp8
Return op name rocdl.dot4.f32.fp8.bf8 as a bitstring.
rocdl.dot4.f32.fp8.bf8
Return op name rocdl.dot4.f32.fp8.fp8 as a bitstring.
rocdl.dot4.f32.fp8.fp8
Return op name rocdl.ds.atomic.async.barrier.arrive.b64 as a bitstring.
rocdl.ds.atomic.async.barrier.arrive.b64
Return op name rocdl.ds.atomic.barrier.arrive.rtn.b64 as a bitstring.
rocdl.ds.atomic.barrier.arrive.rtn.b64
Return op name rocdl.ds_bpermute as a bitstring.
rocdl.ds_bpermute
Return op name rocdl.ds.load.tr4.b64 as a bitstring.
rocdl.ds.load.tr4.b64 - Loads and transposes a matrix from ds memory to registers (available in gfx1250+).
Return op name rocdl.ds.load.tr6.b96 as a bitstring.
rocdl.ds.load.tr6.b96 - Loads and transposes a matrix from ds memory to registers (available in gfx1250+).
Return op name rocdl.ds.load.tr8.b64 as a bitstring.
rocdl.ds.load.tr8.b64 - Loads and transposes a matrix from ds memory to registers (available in gfx1250+).
Return op name rocdl.ds.load.tr16.b128 as a bitstring.
rocdl.ds.load.tr16.b128 - Loads and transposes a matrix from ds memory to registers (available in gfx1250+).
Return op name rocdl.ds.read.tr4.b64 as a bitstring.
rocdl.ds.read.tr4.b64
Return op name rocdl.ds.read.tr6.b96 as a bitstring.
rocdl.ds.read.tr6.b96
Return op name rocdl.ds.read.tr8.b64 as a bitstring.
rocdl.ds.read.tr8.b64
Return op name rocdl.ds.read.tr16.b64 as a bitstring.
rocdl.ds.read.tr16.b64
Return op name rocdl.ds_swizzle as a bitstring.
rocdl.ds_swizzle
Return op name rocdl.exp2 as a bitstring.
rocdl.exp2
Return op name rocdl.exp as a bitstring.
rocdl.exp
Return op name rocdl.fdot2 as a bitstring.
rocdl.fdot2
Return op name rocdl.fdot2.bf16.bf16 as a bitstring.
rocdl.fdot2.bf16.bf16
Return op name rocdl.fdot2.f16.f16 as a bitstring.
rocdl.fdot2.f16.f16
Return op name rocdl.fdot2.f32.bf16 as a bitstring.
rocdl.fdot2.f32.bf16
Return op name rocdl.flat.prefetch as a bitstring.
rocdl.flat.prefetch
Return op name rocdl.fmed3 as a bitstring.
rocdl.fmed3 - Median of three float/half values
Return op name rocdl.global.load.async.lds as a bitstring.
rocdl.global.load.async.lds - Version of rocdl.load.async.to.lds specialized to global pointers
Return op name rocdl.global.load.async.to.lds.b8 as a bitstring.
rocdl.global.load.async.to.lds.b8
Return op name rocdl.global.load.async.to.lds.b32 as a bitstring.
rocdl.global.load.async.to.lds.b32
Return op name rocdl.global.load.async.to.lds.b64 as a bitstring.
rocdl.global.load.async.to.lds.b64
Return op name rocdl.global.load.async.to.lds.b128 as a bitstring.
rocdl.global.load.async.to.lds.b128
Return op name rocdl.global.load.lds as a bitstring.
rocdl.global.load.lds
Return op name rocdl.global.load.tr4.b64 as a bitstring.
rocdl.global.load.tr4.b64 - Loads and transposes a matrix from global memory to registers (available in gfx1250+).
Return op name rocdl.global.load.tr6.b96 as a bitstring.
rocdl.global.load.tr6.b96 - Loads and transposes a matrix from global memory to registers (available in gfx1250+).
Return op name rocdl.global.load.tr.b64 as a bitstring.
rocdl.global.load.tr.b64 - Loads and transposes a matrix from global memory to registers (available in gfx1250+).
Return op name rocdl.global.load.tr.b128 as a bitstring.
rocdl.global.load.tr.b128 - Loads and transposes a matrix from global memory to registers (available in gfx1250+).
Return op name rocdl.global.prefetch as a bitstring.
rocdl.global.prefetch
Return op name rocdl.global.store.async.from.lds.b8 as a bitstring.
rocdl.global.store.async.from.lds.b8
Return op name rocdl.global.store.async.from.lds.b32 as a bitstring.
rocdl.global.store.async.from.lds.b32
Return op name rocdl.global.store.async.from.lds.b64 as a bitstring.
rocdl.global.store.async.from.lds.b64
Return op name rocdl.global.store.async.from.lds.b128 as a bitstring.
rocdl.global.store.async.from.lds.b128
Return op name rocdl.iglp.opt as a bitstring.
rocdl.iglp.opt
Return op name rocdl.load.async.to.lds as a bitstring.
rocdl.load.async.to.lds - Gathering load to LDS that requires explicit async memory tracking
Return op name rocdl.load.to.lds as a bitstring.
rocdl.load.to.lds
Return op name rocdl.log as a bitstring.
rocdl.log
Return op name rocdl.make.buffer.rsrc as a bitstring.
rocdl.make.buffer.rsrc
Return op name rocdl.mbcnt.hi as a bitstring.
rocdl.mbcnt.hi
Return op name rocdl.mbcnt.lo as a bitstring.
rocdl.mbcnt.lo
Return op name rocdl.mfma.f32.4x4x1f32 as a bitstring.
rocdl.mfma.f32.4x4x1f32
Return op name rocdl.mfma.f32.4x4x2bf16 as a bitstring.
rocdl.mfma.f32.4x4x2bf16
Return op name rocdl.mfma.f32.4x4x4bf16.1k as a bitstring.
rocdl.mfma.f32.4x4x4bf16.1k
Return op name rocdl.mfma.f32.4x4x4f16 as a bitstring.
rocdl.mfma.f32.4x4x4f16
Return op name rocdl.mfma.f32.16x16x1f32 as a bitstring.
rocdl.mfma.f32.16x16x1f32
Return op name rocdl.mfma.f32.16x16x2bf16 as a bitstring.
rocdl.mfma.f32.16x16x2bf16
Return op name rocdl.mfma.f32.16x16x4bf16.1k as a bitstring.
rocdl.mfma.f32.16x16x4bf16.1k
Return op name rocdl.mfma.f32.16x16x4f16 as a bitstring.
rocdl.mfma.f32.16x16x4f16
Return op name rocdl.mfma.f32.16x16x4f32 as a bitstring.
rocdl.mfma.f32.16x16x4f32
Return op name rocdl.mfma.f32.16x16x8.xf32 as a bitstring.
rocdl.mfma.f32.16x16x8.xf32
Return op name rocdl.mfma.f32.16x16x8bf16 as a bitstring.
rocdl.mfma.f32.16x16x8bf16
Return op name rocdl.mfma.f32.16x16x16bf16.1k as a bitstring.
rocdl.mfma.f32.16x16x16bf16.1k
Return op name rocdl.mfma.f32.16x16x16f16 as a bitstring.
rocdl.mfma.f32.16x16x16f16
Return op name rocdl.mfma.f32.16x16x32.bf8.bf8 as a bitstring.
rocdl.mfma.f32.16x16x32.bf8.bf8
Return op name rocdl.mfma.f32.16x16x32.bf8.fp8 as a bitstring.
rocdl.mfma.f32.16x16x32.bf8.fp8
Return op name rocdl.mfma.f32.16x16x32.bf16 as a bitstring.
rocdl.mfma.f32.16x16x32.bf16
Return op name rocdl.mfma.f32.16x16x32.f16 as a bitstring.
rocdl.mfma.f32.16x16x32.f16
Return op name rocdl.mfma.f32.16x16x32.fp8.bf8 as a bitstring.
rocdl.mfma.f32.16x16x32.fp8.bf8
Return op name rocdl.mfma.f32.16x16x32.fp8.fp8 as a bitstring.
rocdl.mfma.f32.16x16x32.fp8.fp8
Return op name rocdl.mfma.f32.32x32x1f32 as a bitstring.
rocdl.mfma.f32.32x32x1f32
Return op name rocdl.mfma.f32.32x32x2bf16 as a bitstring.
rocdl.mfma.f32.32x32x2bf16
Return op name rocdl.mfma.f32.32x32x2f32 as a bitstring.
rocdl.mfma.f32.32x32x2f32
Return op name rocdl.mfma.f32.32x32x4.xf32 as a bitstring.
rocdl.mfma.f32.32x32x4.xf32
Return op name rocdl.mfma.f32.32x32x4bf16 as a bitstring.
rocdl.mfma.f32.32x32x4bf16
Return op name rocdl.mfma.f32.32x32x4bf16.1k as a bitstring.
rocdl.mfma.f32.32x32x4bf16.1k
Return op name rocdl.mfma.f32.32x32x4f16 as a bitstring.
rocdl.mfma.f32.32x32x4f16
Return op name rocdl.mfma.f32.32x32x8bf16.1k as a bitstring.
rocdl.mfma.f32.32x32x8bf16.1k
Return op name rocdl.mfma.f32.32x32x8f16 as a bitstring.
rocdl.mfma.f32.32x32x8f16
Return op name rocdl.mfma.f32.32x32x16.bf8.bf8 as a bitstring.
rocdl.mfma.f32.32x32x16.bf8.bf8
Return op name rocdl.mfma.f32.32x32x16.bf8.fp8 as a bitstring.
rocdl.mfma.f32.32x32x16.bf8.fp8
Return op name rocdl.mfma.f32.32x32x16.bf16 as a bitstring.
rocdl.mfma.f32.32x32x16.bf16
Return op name rocdl.mfma.f32.32x32x16.f16 as a bitstring.
rocdl.mfma.f32.32x32x16.f16
Return op name rocdl.mfma.f32.32x32x16.fp8.bf8 as a bitstring.
rocdl.mfma.f32.32x32x16.fp8.bf8
Return op name rocdl.mfma.f32.32x32x16.fp8.fp8 as a bitstring.
rocdl.mfma.f32.32x32x16.fp8.fp8
Return op name rocdl.mfma.f64.4x4x4f64 as a bitstring.
rocdl.mfma.f64.4x4x4f64
Return op name rocdl.mfma.f64.16x16x4f64 as a bitstring.
rocdl.mfma.f64.16x16x4f64
Return op name rocdl.mfma.i32.4x4x4i8 as a bitstring.
rocdl.mfma.i32.4x4x4i8
Return op name rocdl.mfma.i32.16x16x4i8 as a bitstring.
rocdl.mfma.i32.16x16x4i8
Return op name rocdl.mfma.i32.16x16x16i8 as a bitstring.
rocdl.mfma.i32.16x16x16i8
Return op name rocdl.mfma.i32.16x16x32.i8 as a bitstring.
rocdl.mfma.i32.16x16x32.i8
Return op name rocdl.mfma.i32.16x16x64.i8 as a bitstring.
rocdl.mfma.i32.16x16x64.i8
Return op name rocdl.mfma.i32.32x32x4i8 as a bitstring.
rocdl.mfma.i32.32x32x4i8
Return op name rocdl.mfma.i32.32x32x8i8 as a bitstring.
rocdl.mfma.i32.32x32x8i8
Return op name rocdl.mfma.i32.32x32x16.i8 as a bitstring.
rocdl.mfma.i32.32x32x16.i8
Return op name rocdl.mfma.i32.32x32x32.i8 as a bitstring.
rocdl.mfma.i32.32x32x32.i8
Return op name rocdl.mfma.scale.f32.16x16x128.f8f6f4 as a bitstring.
rocdl.mfma.scale.f32.16x16x128.f8f6f4
Return op name rocdl.mfma.scale.f32.32x32x64.f8f6f4 as a bitstring.
rocdl.mfma.scale.f32.32x32x64.f8f6f4
Return op name rocdl.permlane16.swap as a bitstring.
rocdl.permlane16.swap
Return op name rocdl.permlane16.var as a bitstring.
rocdl.permlane16.var
Return op name rocdl.permlane32.swap as a bitstring.
rocdl.permlane32.swap
Return op name rocdl.permlanex16 as a bitstring.
rocdl.permlanex16
Return op name rocdl.permlanex16.var as a bitstring.
rocdl.permlanex16.var
Return op name rocdl.ptr.s.buffer.load as a bitstring.
rocdl.ptr.s.buffer.load
Return op name rocdl.raw.buffer.atomic.cmpswap as a bitstring.
rocdl.raw.buffer.atomic.cmpswap
Return op name rocdl.raw.buffer.atomic.fadd as a bitstring.
rocdl.raw.buffer.atomic.fadd
Return op name rocdl.raw.buffer.atomic.fmax as a bitstring.
rocdl.raw.buffer.atomic.fmax
Return op name rocdl.raw.buffer.atomic.smax as a bitstring.
rocdl.raw.buffer.atomic.smax
Return op name rocdl.raw.buffer.atomic.umin as a bitstring.
rocdl.raw.buffer.atomic.umin
Return op name rocdl.raw.buffer.load as a bitstring.
rocdl.raw.buffer.load
Return op name rocdl.raw.buffer.store as a bitstring.
rocdl.raw.buffer.store
Return op name rocdl.raw.ptr.buffer.atomic.cmpswap as a bitstring.
rocdl.raw.ptr.buffer.atomic.cmpswap
Return op name rocdl.raw.ptr.buffer.atomic.fadd as a bitstring.
rocdl.raw.ptr.buffer.atomic.fadd
Return op name rocdl.raw.ptr.buffer.atomic.fmax as a bitstring.
rocdl.raw.ptr.buffer.atomic.fmax
Return op name rocdl.raw.ptr.buffer.atomic.smax as a bitstring.
rocdl.raw.ptr.buffer.atomic.smax
Return op name rocdl.raw.ptr.buffer.atomic.umin as a bitstring.
rocdl.raw.ptr.buffer.atomic.umin
Return op name rocdl.raw.ptr.buffer.load as a bitstring.
rocdl.raw.ptr.buffer.load
Return op name rocdl.raw.ptr.buffer.load.async.lds as a bitstring.
rocdl.raw.ptr.buffer.load.async.lds - Async variant of raw.ptr.buffer.load.lds
Return op name rocdl.raw.ptr.buffer.load.lds as a bitstring.
rocdl.raw.ptr.buffer.load.lds
Return op name rocdl.raw.ptr.buffer.store as a bitstring.
rocdl.raw.ptr.buffer.store
Return op name rocdl.rcp as a bitstring.
rocdl.rcp
Return op name rocdl.readfirstlane as a bitstring.
rocdl.readfirstlane - Get the value in first active lane.
Return op name rocdl.readlane as a bitstring.
rocdl.readlane - Get the value in the specific lane.
Return op name rocdl.rsq as a bitstring.
rocdl.rsq
Return op name rocdl.s.barrier as a bitstring.
rocdl.s.barrier
Return op name rocdl.s.barrier.init as a bitstring.
rocdl.s.barrier.init
Return op name rocdl.s.barrier.join as a bitstring.
rocdl.s.barrier.join
Return op name rocdl.s.barrier.leave as a bitstring.
rocdl.s.barrier.leave
Return op name rocdl.s.barrier.signal as a bitstring.
rocdl.s.barrier.signal
Return op name rocdl.s.barrier.signal.isfirst as a bitstring.
rocdl.s.barrier.signal.isfirst
Return op name rocdl.s.barrier.signal.var as a bitstring.
rocdl.s.barrier.signal.var
Return op name rocdl.s.barrier.wait as a bitstring.
rocdl.s.barrier.wait
Return op name rocdl.s.get.barrier.state as a bitstring.
rocdl.s.get.barrier.state
Return op name rocdl.s.get.named.barrier.state as a bitstring.
rocdl.s.get.named.barrier.state
Return op name rocdl.s.nop as a bitstring.
rocdl.s.nop
Return op name rocdl.s.setprio as a bitstring.
rocdl.s.setprio
Return op name rocdl.s.sleep as a bitstring.
rocdl.s.sleep
Return op name rocdl.s.wait.asynccnt as a bitstring.
rocdl.s.wait.asynccnt - Wait until ASYNCCNT is less than or equal to count
Return op name rocdl.s.wait.dscnt as a bitstring.
rocdl.s.wait.dscnt - Wait until DSCNT is less than or equal to count
Return op name rocdl.s.wait.expcnt as a bitstring.
rocdl.s.wait.expcnt - Wait until EXPCNT is less than or equal to count
Return op name rocdl.s.wait.loadcnt as a bitstring.
rocdl.s.wait.loadcnt - Wait until LOADCNT is less than or equal to count
Return op name rocdl.s.wait.storecnt as a bitstring.
rocdl.s.wait.storecnt - Wait until STORECNT is less than or equal to count
Return op name rocdl.s.wait.tensorcnt as a bitstring.
rocdl.s.wait.tensorcnt - Wait until TENSORCNT is less than or equal to count
Return op name rocdl.s.waitcnt as a bitstring.
rocdl.s.waitcnt
Return op name rocdl.s.wakeup.barrier as a bitstring.
rocdl.s.wakeup.barrier
Return op name rocdl.sched.barrier as a bitstring.
rocdl.sched.barrier
Return op name rocdl.sched.group.barrier as a bitstring.
rocdl.sched.group.barrier
Return op name rocdl.sdot2 as a bitstring.
rocdl.sdot2
Return op name rocdl.sdot4 as a bitstring.
rocdl.sdot4
Return op name rocdl.sdot8 as a bitstring.
rocdl.sdot8
Return op name rocdl.sin as a bitstring.
rocdl.sin
Return op name rocdl.smfmac.f32.16x16x32.bf16 as a bitstring.
rocdl.smfmac.f32.16x16x32.bf16
Return op name rocdl.smfmac.f32.16x16x32.f16 as a bitstring.
rocdl.smfmac.f32.16x16x32.f16
Return op name rocdl.smfmac.f32.16x16x64.bf8.bf8 as a bitstring.
rocdl.smfmac.f32.16x16x64.bf8.bf8
Return op name rocdl.smfmac.f32.16x16x64.bf8.fp8 as a bitstring.
rocdl.smfmac.f32.16x16x64.bf8.fp8
Return op name rocdl.smfmac.f32.16x16x64.bf16 as a bitstring.
rocdl.smfmac.f32.16x16x64.bf16
Return op name rocdl.smfmac.f32.16x16x64.f16 as a bitstring.
rocdl.smfmac.f32.16x16x64.f16
Return op name rocdl.smfmac.f32.16x16x64.fp8.bf8 as a bitstring.
rocdl.smfmac.f32.16x16x64.fp8.bf8
Return op name rocdl.smfmac.f32.16x16x64.fp8.fp8 as a bitstring.
rocdl.smfmac.f32.16x16x64.fp8.fp8
Return op name rocdl.smfmac.f32.16x16x128.bf8.bf8 as a bitstring.
rocdl.smfmac.f32.16x16x128.bf8.bf8
Return op name rocdl.smfmac.f32.16x16x128.bf8.fp8 as a bitstring.
rocdl.smfmac.f32.16x16x128.bf8.fp8
Return op name rocdl.smfmac.f32.16x16x128.fp8.bf8 as a bitstring.
rocdl.smfmac.f32.16x16x128.fp8.bf8
Return op name rocdl.smfmac.f32.16x16x128.fp8.fp8 as a bitstring.
rocdl.smfmac.f32.16x16x128.fp8.fp8
Return op name rocdl.smfmac.f32.32x32x16.bf16 as a bitstring.
rocdl.smfmac.f32.32x32x16.bf16
Return op name rocdl.smfmac.f32.32x32x16.f16 as a bitstring.
rocdl.smfmac.f32.32x32x16.f16
Return op name rocdl.smfmac.f32.32x32x32.bf8.bf8 as a bitstring.
rocdl.smfmac.f32.32x32x32.bf8.bf8
Return op name rocdl.smfmac.f32.32x32x32.bf8.fp8 as a bitstring.
rocdl.smfmac.f32.32x32x32.bf8.fp8
Return op name rocdl.smfmac.f32.32x32x32.bf16 as a bitstring.
rocdl.smfmac.f32.32x32x32.bf16
Return op name rocdl.smfmac.f32.32x32x32.f16 as a bitstring.
rocdl.smfmac.f32.32x32x32.f16
Return op name rocdl.smfmac.f32.32x32x32.fp8.bf8 as a bitstring.
rocdl.smfmac.f32.32x32x32.fp8.bf8
Return op name rocdl.smfmac.f32.32x32x32.fp8.fp8 as a bitstring.
rocdl.smfmac.f32.32x32x32.fp8.fp8
Return op name rocdl.smfmac.f32.32x32x64.bf8.bf8 as a bitstring.
rocdl.smfmac.f32.32x32x64.bf8.bf8
Return op name rocdl.smfmac.f32.32x32x64.bf8.fp8 as a bitstring.
rocdl.smfmac.f32.32x32x64.bf8.fp8
Return op name rocdl.smfmac.f32.32x32x64.fp8.bf8 as a bitstring.
rocdl.smfmac.f32.32x32x64.fp8.bf8
Return op name rocdl.smfmac.f32.32x32x64.fp8.fp8 as a bitstring.
rocdl.smfmac.f32.32x32x64.fp8.fp8
Return op name rocdl.smfmac.i32.16x16x64.i8 as a bitstring.
rocdl.smfmac.i32.16x16x64.i8
Return op name rocdl.smfmac.i32.16x16x128.i8 as a bitstring.
rocdl.smfmac.i32.16x16x128.i8
Return op name rocdl.smfmac.i32.32x32x32.i8 as a bitstring.
rocdl.smfmac.i32.32x32x32.i8
Return op name rocdl.smfmac.i32.32x32x64.i8 as a bitstring.
rocdl.smfmac.i32.32x32x64.i8
Return op name rocdl.sqrt as a bitstring.
rocdl.sqrt
Return op name rocdl.sudot4 as a bitstring.
rocdl.sudot4
Return op name rocdl.sudot8 as a bitstring.
rocdl.sudot8
Return op name rocdl.swmmac.bf16.16x16x32.bf16 as a bitstring.
rocdl.swmmac.bf16.16x16x32.bf16
Return op name rocdl.swmmac.bf16.16x16x64.bf16 as a bitstring.
rocdl.swmmac.bf16.16x16x64.bf16
Return op name rocdl.swmmac.bf16f32.16x16x64.bf16 as a bitstring.
rocdl.swmmac.bf16f32.16x16x64.bf16
Return op name rocdl.swmmac.f16.16x16x32.f16 as a bitstring.
rocdl.swmmac.f16.16x16x32.f16
Return op name rocdl.swmmac.f16.16x16x64.f16 as a bitstring.
rocdl.swmmac.f16.16x16x64.f16
Return op name rocdl.swmmac.f16.16x16x128.bf8.bf8 as a bitstring.
rocdl.swmmac.f16.16x16x128.bf8.bf8
Return op name rocdl.swmmac.f16.16x16x128.bf8.fp8 as a bitstring.
rocdl.swmmac.f16.16x16x128.bf8.fp8
Return op name rocdl.swmmac.f16.16x16x128.fp8.bf8 as a bitstring.
rocdl.swmmac.f16.16x16x128.fp8.bf8
Return op name rocdl.swmmac.f16.16x16x128.fp8.fp8 as a bitstring.
rocdl.swmmac.f16.16x16x128.fp8.fp8
Return op name rocdl.swmmac.f32.16x16x32.bf8.bf8 as a bitstring.
rocdl.swmmac.f32.16x16x32.bf8.bf8
Return op name rocdl.swmmac.f32.16x16x32.bf8.fp8 as a bitstring.
rocdl.swmmac.f32.16x16x32.bf8.fp8
Return op name rocdl.swmmac.f32.16x16x32.bf16 as a bitstring.
rocdl.swmmac.f32.16x16x32.bf16
Return op name rocdl.swmmac.f32.16x16x32.f16 as a bitstring.
rocdl.swmmac.f32.16x16x32.f16
Return op name rocdl.swmmac.f32.16x16x32.fp8.bf8 as a bitstring.
rocdl.swmmac.f32.16x16x32.fp8.bf8
Return op name rocdl.swmmac.f32.16x16x32.fp8.fp8 as a bitstring.
rocdl.swmmac.f32.16x16x32.fp8.fp8
Return op name rocdl.swmmac.f32.16x16x64.bf16 as a bitstring.
rocdl.swmmac.f32.16x16x64.bf16
Return op name rocdl.swmmac.f32.16x16x64.f16 as a bitstring.
rocdl.swmmac.f32.16x16x64.f16
Return op name rocdl.swmmac.f32.16x16x128.bf8.bf8 as a bitstring.
rocdl.swmmac.f32.16x16x128.bf8.bf8
Return op name rocdl.swmmac.f32.16x16x128.bf8.fp8 as a bitstring.
rocdl.swmmac.f32.16x16x128.bf8.fp8
Return op name rocdl.swmmac.f32.16x16x128.fp8.bf8 as a bitstring.
rocdl.swmmac.f32.16x16x128.fp8.bf8
Return op name rocdl.swmmac.f32.16x16x128.fp8.fp8 as a bitstring.
rocdl.swmmac.f32.16x16x128.fp8.fp8
Return op name rocdl.swmmac.i32.16x16x32.iu4 as a bitstring.
rocdl.swmmac.i32.16x16x32.iu4
Return op name rocdl.swmmac.i32.16x16x32.iu8 as a bitstring.
rocdl.swmmac.i32.16x16x32.iu8
Return op name rocdl.swmmac.i32.16x16x64.iu4 as a bitstring.
rocdl.swmmac.i32.16x16x64.iu4
Return op name rocdl.swmmac.i32.16x16x128.iu8 as a bitstring.
rocdl.swmmac.i32.16x16x128.iu8
Return op name rocdl.tanh as a bitstring.
rocdl.tanh
Return op name rocdl.tensor.load.to.lds as a bitstring.
rocdl.tensor.load.to.lds - Base class for ROCDL tensor load/store to/from LDS.
Return op name rocdl.tensor.store.from.lds as a bitstring.
rocdl.tensor.store.from.lds - Base class for ROCDL tensor load/store to/from LDS.
Return op name rocdl.udot2 as a bitstring.
rocdl.udot2
Return op name rocdl.udot4 as a bitstring.
rocdl.udot4
Return op name rocdl.udot8 as a bitstring.
rocdl.udot8
Return op name rocdl.update.dpp as a bitstring.
rocdl.update.dpp
Return op name rocdl.wait.asyncmark as a bitstring.
rocdl.wait.asyncmark - Wait until N or fewer async operation groups are unexecuted
Return op name rocdl.wave.barrier as a bitstring.
rocdl.wave.barrier
Return op name rocdl.wave.id as a bitstring.
rocdl.wave.id
Return op name rocdl.wavefrontsize as a bitstring.
rocdl.wavefrontsize
Return op name rocdl.wmma.bf16.16x16x16.bf16 as a bitstring.
rocdl.wmma.bf16.16x16x16.bf16
Return op name rocdl.wmma.bf16.16x16x32.bf16 as a bitstring.
rocdl.wmma.bf16.16x16x32.bf16
Return op name rocdl.wmma.bf16f32.16x16x32.bf16 as a bitstring.
rocdl.wmma.bf16f32.16x16x32.bf16
Return op name rocdl.wmma.f16.16x16x16.f16 as a bitstring.
rocdl.wmma.f16.16x16x16.f16
Return op name rocdl.wmma.f16.16x16x32.f16 as a bitstring.
rocdl.wmma.f16.16x16x32.f16
Return op name rocdl.wmma.f16.16x16x64.bf8_bf8 as a bitstring.
rocdl.wmma.f16.16x16x64.bf8_bf8
Return op name rocdl.wmma.f16.16x16x64.bf8_fp8 as a bitstring.
rocdl.wmma.f16.16x16x64.bf8_fp8
Return op name rocdl.wmma.f16.16x16x64.fp8_bf8 as a bitstring.
rocdl.wmma.f16.16x16x64.fp8_bf8
Return op name rocdl.wmma.f16.16x16x64.fp8_fp8 as a bitstring.
rocdl.wmma.f16.16x16x64.fp8_fp8
Return op name rocdl.wmma.f16.16x16x128.bf8_bf8 as a bitstring.
rocdl.wmma.f16.16x16x128.bf8_bf8
Return op name rocdl.wmma.f16.16x16x128.bf8_fp8 as a bitstring.
rocdl.wmma.f16.16x16x128.bf8_fp8
Return op name rocdl.wmma.f16.16x16x128.fp8_bf8 as a bitstring.
rocdl.wmma.f16.16x16x128.fp8_bf8
Return op name rocdl.wmma.f16.16x16x128.fp8_fp8 as a bitstring.
rocdl.wmma.f16.16x16x128.fp8_fp8
Return op name rocdl.wmma.f32.16x16x4.f32 as a bitstring.
rocdl.wmma.f32.16x16x4.f32
Return op name rocdl.wmma.f32.16x16x16.bf8_bf8 as a bitstring.
rocdl.wmma.f32.16x16x16.bf8_bf8
Return op name rocdl.wmma.f32.16x16x16.bf8_fp8 as a bitstring.
rocdl.wmma.f32.16x16x16.bf8_fp8
Return op name rocdl.wmma.f32.16x16x16.bf16 as a bitstring.
rocdl.wmma.f32.16x16x16.bf16
Return op name rocdl.wmma.f32.16x16x16.f16 as a bitstring.
rocdl.wmma.f32.16x16x16.f16
Return op name rocdl.wmma.f32.16x16x16.fp8_bf8 as a bitstring.
rocdl.wmma.f32.16x16x16.fp8_bf8
Return op name rocdl.wmma.f32.16x16x16.fp8_fp8 as a bitstring.
rocdl.wmma.f32.16x16x16.fp8_fp8
Return op name rocdl.wmma.f32.16x16x32.bf16 as a bitstring.
rocdl.wmma.f32.16x16x32.bf16
Return op name rocdl.wmma.f32.16x16x32.f16 as a bitstring.
rocdl.wmma.f32.16x16x32.f16
Return op name rocdl.wmma.f32.16x16x64.bf8_bf8 as a bitstring.
rocdl.wmma.f32.16x16x64.bf8_bf8
Return op name rocdl.wmma.f32.16x16x64.bf8_fp8 as a bitstring.
rocdl.wmma.f32.16x16x64.bf8_fp8
Return op name rocdl.wmma.f32.16x16x64.fp8_bf8 as a bitstring.
rocdl.wmma.f32.16x16x64.fp8_bf8
Return op name rocdl.wmma.f32.16x16x64.fp8_fp8 as a bitstring.
rocdl.wmma.f32.16x16x64.fp8_fp8
Return op name rocdl.wmma.f32.16x16x128.bf8_bf8 as a bitstring.
rocdl.wmma.f32.16x16x128.bf8_bf8
Return op name rocdl.wmma.f32.16x16x128.bf8_fp8 as a bitstring.
rocdl.wmma.f32.16x16x128.bf8_fp8
Return op name rocdl.wmma.f32.16x16x128.fp8_bf8 as a bitstring.
rocdl.wmma.f32.16x16x128.fp8_bf8
Return op name rocdl.wmma.f32.16x16x128.fp8_fp8 as a bitstring.
rocdl.wmma.f32.16x16x128.fp8_fp8
Return op name rocdl.wmma.i32.16x16x16.iu4 as a bitstring.
rocdl.wmma.i32.16x16x16.iu4
Return op name rocdl.wmma.i32.16x16x16.iu8 as a bitstring.
rocdl.wmma.i32.16x16x16.iu8
Return op name rocdl.wmma.i32.16x16x32.iu4 as a bitstring.
rocdl.wmma.i32.16x16x32.iu4
Return op name rocdl.wmma.i32.16x16x64.iu8 as a bitstring.
rocdl.wmma.i32.16x16x64.iu8
Return op name rocdl.wmma.scale16.f32.16x16x128.f8f6f4 as a bitstring.
rocdl.wmma.scale16.f32.16x16x128.f8f6f4
Return op name rocdl.wmma.scale16.f32.32x16x128.f4 as a bitstring.
rocdl.wmma.scale16.f32.32x16x128.f4
Return op name rocdl.wmma.scale.f32.16x16x128.f8f6f4 as a bitstring.
rocdl.wmma.scale.f32.16x16x128.f8f6f4
Return op name rocdl.wmma.scale.f32.32x16x128.f4 as a bitstring.
rocdl.wmma.scale.f32.32x16x128.f4
Return op name rocdl.workgroup.id.x as a bitstring.
rocdl.workgroup.id.x
Return op name rocdl.workgroup.id.y as a bitstring.
rocdl.workgroup.id.y
Return op name rocdl.workgroup.id.z as a bitstring.
rocdl.workgroup.id.z
Return op name rocdl.workitem.id.x as a bitstring.
rocdl.workitem.id.x
Return op name rocdl.workitem.id.y as a bitstring.
rocdl.workitem.id.y
Return op name rocdl.workitem.id.z as a bitstring.
rocdl.workitem.id.z
Functions
Return op name rocdl.asyncmark as a bitstring.
rocdl.asyncmark - Mark the end of a group of asynchronous operations
Description
This operation, in conjunction with rocdl.wait.asyncmark, forms the
compiler-provided framework for tracking explicitly asynchronous
memory operations, such as copies to LDS that use async intrinsics
and gfx1250's tensor loads.
Details of its behavior can be found in the LLVM documentation on async tracking.
See rocdl.wait.asyncmark's documentation for a usage example.
Example:
// Mark the end of an async operation group.
rocdl.asyncmarkAvailable on gfx9 and later.
Return op name rocdl.ballot as a bitstring.
rocdl.ballot - Vote across thread group
Operands
pred- Single,I1, 1-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Ballot provides a bit mask containing the 1-bit predicate value from each lane. The nth bit of the result contains the 1 bit contributed by the nth warp lane.
Example:
// Ballot across thread group.
%0 = rocdl.ballot %pred : i64
Return op name rocdl.barrier as a bitstring.
rocdl.barrier
Description
An operation with the same expansion as HIP's __synchthreads();
DEPRECATION NOTICE: Use gpu.barrier, which will expand to these
operations, instead.
Example:
// Workgroup barrier with acquire/release fences.
rocdl.barrier
Return op name rocdl.cluster.id.x as a bitstring.
rocdl.cluster.id.x
Attributes
range- Optional,LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Read a hardware register for thread/workgroup/cluster identification.
An optional range attribute can constrain the returned value.
Example:
// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32
// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32
Return op name rocdl.cluster.id.y as a bitstring.
rocdl.cluster.id.y
Attributes
range- Optional,LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Read a hardware register for thread/workgroup/cluster identification.
An optional range attribute can constrain the returned value.
Example:
// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32
// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32
Return op name rocdl.cluster.id.z as a bitstring.
rocdl.cluster.id.z
Attributes
range- Optional,LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Read a hardware register for thread/workgroup/cluster identification.
An optional range attribute can constrain the returned value.
Example:
// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32
// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32
Return op name rocdl.cluster.load.async.to.lds.b8 as a bitstring.
rocdl.cluster.load.async.to.lds.b8
Attributes
offset- Single,I32Attr, 32-bit signless integer attributecpol- Single,ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
globalPtr- Single,ROCDLGlobalBuffer, LLVM pointer in address space 1ldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3mask- Single,I32, 32-bit signless integer
Description
Broadcasts memory load of 8 bits of data for a cluster of workgroups.
Available on gfx1250+.
Example:
// Cluster broadcast 8-bit load to LDS.
rocdl.cluster.load.async.to.lds.b8 %src, %dst, 0, 0, %mask : !llvm.ptr<1>, !llvm.ptr<3>
Return op name rocdl.cluster.load.async.to.lds.b32 as a bitstring.
rocdl.cluster.load.async.to.lds.b32
Attributes
offset- Single,I32Attr, 32-bit signless integer attributecpol- Single,ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
globalPtr- Single,ROCDLGlobalBuffer, LLVM pointer in address space 1ldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3mask- Single,I32, 32-bit signless integer
Description
Broadcasts memory load of 32 bits of data for a cluster of workgroups.
Available on gfx1250+.
Example:
// Cluster broadcast 32-bit load to LDS.
rocdl.cluster.load.async.to.lds.b32 %src, %dst, 0, 0, %mask : !llvm.ptr<1>, !llvm.ptr<3>
Return op name rocdl.cluster.load.async.to.lds.b64 as a bitstring.
rocdl.cluster.load.async.to.lds.b64
Attributes
offset- Single,I32Attr, 32-bit signless integer attributecpol- Single,ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
globalPtr- Single,ROCDLGlobalBuffer, LLVM pointer in address space 1ldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3mask- Single,I32, 32-bit signless integer
Description
Broadcasts memory load of 64 bits of data for a cluster of workgroups.
Available on gfx1250+.
Example:
// Cluster broadcast 64-bit load to LDS.
rocdl.cluster.load.async.to.lds.b64 %src, %dst, 0, 0, %mask : !llvm.ptr<1>, !llvm.ptr<3>
Return op name rocdl.cluster.load.async.to.lds.b128 as a bitstring.
rocdl.cluster.load.async.to.lds.b128
Attributes
offset- Single,I32Attr, 32-bit signless integer attributecpol- Single,ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
globalPtr- Single,ROCDLGlobalBuffer, LLVM pointer in address space 1ldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3mask- Single,I32, 32-bit signless integer
Description
Broadcasts memory load of 128 bits of data for a cluster of workgroups.
Available on gfx1250+.
Example:
// Cluster broadcast 128-bit load to LDS.
rocdl.cluster.load.async.to.lds.b128 %src, %dst, 0, 0, %mask : !llvm.ptr<1>, !llvm.ptr<3>
Return op name rocdl.cluster.workgroup.id.x as a bitstring.
rocdl.cluster.workgroup.id.x
Attributes
range- Optional,LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Read a hardware register for thread/workgroup/cluster identification.
An optional range attribute can constrain the returned value.
Example:
// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32
// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32
Return op name rocdl.cluster.workgroup.id.y as a bitstring.
rocdl.cluster.workgroup.id.y
Attributes
range- Optional,LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Read a hardware register for thread/workgroup/cluster identification.
An optional range attribute can constrain the returned value.
Example:
// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32
// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32
Return op name rocdl.cluster.workgroup.id.z as a bitstring.
rocdl.cluster.workgroup.id.z
Attributes
range- Optional,LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Read a hardware register for thread/workgroup/cluster identification.
An optional range attribute can constrain the returned value.
Example:
// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32
// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32
Return op name rocdl.cos as a bitstring.
rocdl.cos
Operands
arg- Single,LLVM_AnyFloat, floating point LLVM type
Results
res- Single,LLVM_AnyFloat, floating point LLVM type
Description
Note: In the general case, prefer the conventional arith, math, or llvm ops over this.
Use this ROCDL-specific operation only when you fully understand its implication and
when it is strictly necessary. This op is usually chosen when a small loss in precision is
acceptable in exchange for higher execution speed.
Example:
%0 = rocdl.cos %a f32 -> f32
Return op name rocdl.cvt.f32.bf8 as a bitstring.
rocdl.cvt.f32.bf8 - Convert bf8 to f32
Attributes
byteSel- Single,I32Attr, 32-bit signless integer attribute
Operands
srcA- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Convert 8-bit bf8 value from the byteSelth bit of srcA to fp32.
Example:
// Convert bf8 byte 0 to f32.
%0 = rocdl.cvt.f32.bf8 %src[0] : f32
Return op name rocdl.cvt.f32.fp8 as a bitstring.
rocdl.cvt.f32.fp8 - Convert fp8 to f32
Attributes
byteSel- Single,I32Attr, 32-bit signless integer attribute
Operands
srcA- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Convert 8-bit fp8 value from the byteSelth bit of srcA to fp32.
Example:
// Convert fp8 byte 0 to f32.
%0 = rocdl.cvt.f32.fp8 %src[0] : f32
Return op name rocdl.cvt.pk.bf8.f32 as a bitstring.
rocdl.cvt.pk.bf8.f32 - Convert two f32's to bf8
Attributes
wordSel- Single,I1Attr, 1-bit signless integer attribute
Operands
srcA- Single,F32, 32-bit floatsrcB- Single,F32, 32-bit floatold- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Convert srcA and srcB to bf8 and store into the low/high word of
old, preserving the other word.
Example:
// Pack two f32 values into bf8 in the low word of old.
%0 = rocdl.cvt.pk.bf8.f32 %a, %b -> %old[false] : i32
Return op name rocdl.cvt.pk.f32.bf8 as a bitstring.
rocdl.cvt.pk.f32.bf8 - Convert packed bf8 to packed f32
Attributes
wordSel- Single,I1Attr, 1-bit signless integer attribute
Operands
src- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Convert src based on $wordSel to packed fp32.
Example:
// Unpack bf8 word to packed f32.
%0 = rocdl.cvt.pk.f32.bf8 %src[false] : vector<2xf32>
Return op name rocdl.cvt.pk.f32.fp8 as a bitstring.
rocdl.cvt.pk.f32.fp8 - Convert packed fp8 to packed f32
Attributes
wordSel- Single,I1Attr, 1-bit signless integer attribute
Operands
src- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Convert src based on $wordSel to packed fp32.
Example:
// Unpack fp8 word to packed f32.
%0 = rocdl.cvt.pk.f32.fp8 %src[false] : vector<2xf32>
Return op name rocdl.cvt.pk.fp8.f32 as a bitstring.
rocdl.cvt.pk.fp8.f32 - Convert two f32's to fp8
Attributes
wordSel- Single,I1Attr, 1-bit signless integer attribute
Operands
srcA- Single,F32, 32-bit floatsrcB- Single,F32, 32-bit floatold- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Convert srcA and srcB to fp8 and store into the low/high word of
old, preserving the other word.
Example:
// Pack two f32 values into fp8 in the low word of old.
%0 = rocdl.cvt.pk.fp8.f32 %a, %b -> %old[false] : i32
Return op name rocdl.cvt.pkrtz as a bitstring.
rocdl.cvt.pkrtz - Convert two f32 input into a vector<2xf16>
Operands
srcA- Single,F32, 32-bit floatsrcB- Single,F32, 32-bit float
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Convert two f32 values into a packed vector<2xf16>.
Example:
// Pack two f32 values into a vector<2xf16> with round-to-zero.
%0 = rocdl.cvt.pkrtz %a, %b : vector<2xf16>
Return op name rocdl.cvt.scale.pk8.bf16.bf8 as a bitstring.
rocdl.cvt.scale.pk8.bf16.bf8 - Scales 8 bf8 and converts them to 8 bf16.
Attributes
scaleSel- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2scale- Single,I32, 32-bit signless integer
Results
res- Single,ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8
Description
Available on gfx1250+.
Return op name rocdl.cvt.scale.pk8.bf16.fp4 as a bitstring.
rocdl.cvt.scale.pk8.bf16.fp4 - Scales 8 fp4 and converts them to 8 bf16.
Attributes
scaleSel- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,I32, 32-bit signless integerscale- Single,I32, 32-bit signless integer
Results
res- Single,ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8
Description
Available on gfx1250+.
Return op name rocdl.cvt.scale.pk8.bf16.fp8 as a bitstring.
rocdl.cvt.scale.pk8.bf16.fp8 - Scales 8 fp8 and converts them to 8 bf16.
Attributes
scaleSel- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2scale- Single,I32, 32-bit signless integer
Results
res- Single,ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8
Description
Available on gfx1250+.
Return op name rocdl.cvt.scale.pk8.f16.bf8 as a bitstring.
rocdl.cvt.scale.pk8.f16.bf8 - Scales 8 bf8 and converts them to 8 f16.
Attributes
scaleSel- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2scale- Single,I32, 32-bit signless integer
Results
res- Single,ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8
Description
Available on gfx1250+.
Return op name rocdl.cvt.scale.pk8.f16.fp4 as a bitstring.
rocdl.cvt.scale.pk8.f16.fp4 - Scales 8 fp4 and converts them to 8 f16.
Attributes
scaleSel- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,I32, 32-bit signless integerscale- Single,I32, 32-bit signless integer
Results
res- Single,ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8
Description
Available on gfx1250+.
Return op name rocdl.cvt.scale.pk8.f16.fp8 as a bitstring.
rocdl.cvt.scale.pk8.f16.fp8 - Scales 8 fp8 and converts them to 8 f16.
Attributes
scaleSel- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2scale- Single,I32, 32-bit signless integer
Results
res- Single,ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8
Description
Available on gfx1250+.
Return op name rocdl.cvt.scale.pk8.f32.bf8 as a bitstring.
rocdl.cvt.scale.pk8.f32.bf8 - Scales 8 bf8 and converts them to 8 f32.
Attributes
scaleSel- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2scale- Single,I32, 32-bit signless integer
Results
res- Single,ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8
Description
Available on gfx1250+.
Return op name rocdl.cvt.scale.pk8.f32.fp4 as a bitstring.
rocdl.cvt.scale.pk8.f32.fp4 - Scales 8 fp4 and converts them to 8 f32.
Attributes
scaleSel- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,I32, 32-bit signless integerscale- Single,I32, 32-bit signless integer
Results
res- Single,ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8
Description
Available on gfx1250+.
Return op name rocdl.cvt.scale.pk8.f32.fp8 as a bitstring.
rocdl.cvt.scale.pk8.f32.fp8 - Scales 8 fp8 and converts them to 8 f32.
Attributes
scaleSel- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2scale- Single,I32, 32-bit signless integer
Results
res- Single,ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8
Description
Available on gfx1250+.
Return op name rocdl.cvt.scale.pk16.bf16.bf6 as a bitstring.
rocdl.cvt.scale.pk16.bf16.bf6 - Scales 16 bf6 and converts them to 16 bf16.
Attributes
scaleSel- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3scale- Single,I32, 32-bit signless integer
Results
res- Single,ROCDL_V16BF16Type, fixed-length vector of bfloat16 type values of length 16
Description
Available on gfx1250+.
Return op name rocdl.cvt.scale.pk16.bf16.fp6 as a bitstring.
rocdl.cvt.scale.pk16.bf16.fp6 - Scales 16 fp6 and converts them to 16 bf16.
Attributes
scaleSel- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3scale- Single,I32, 32-bit signless integer
Results
res- Single,ROCDL_V16BF16Type, fixed-length vector of bfloat16 type values of length 16
Description
Available on gfx1250+.
Return op name rocdl.cvt.scale.pk16.f16.bf6 as a bitstring.
rocdl.cvt.scale.pk16.f16.bf6 - Scales 16 bf6 and converts them to 16 f16.
Attributes
scaleSel- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3scale- Single,I32, 32-bit signless integer
Results
res- Single,ROCDL_V16F16Type, fixed-length vector of 16-bit float values of length 16
Description
Available on gfx1250+.
Return op name rocdl.cvt.scale.pk16.f16.fp6 as a bitstring.
rocdl.cvt.scale.pk16.f16.fp6 - Scales 16 fp6 and converts them to 16 f16.
Attributes
scaleSel- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3scale- Single,I32, 32-bit signless integer
Results
res- Single,ROCDL_V16F16Type, fixed-length vector of 16-bit float values of length 16
Description
Available on gfx1250+.
Return op name rocdl.cvt.scale.pk16.f32.bf6 as a bitstring.
rocdl.cvt.scale.pk16.f32.bf6 - Scales 16 bf6 and converts them to 16 f32.
Attributes
scaleSel- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3scale- Single,I32, 32-bit signless integer
Results
res- Single,ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16
Description
Available on gfx1250+.
Return op name rocdl.cvt.scale.pk16.f32.fp6 as a bitstring.
rocdl.cvt.scale.pk16.f32.fp6 - Scales 16 fp6 and converts them to 16 f32.
Attributes
scaleSel- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3scale- Single,I32, 32-bit signless integer
Results
res- Single,ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16
Description
Available on gfx1250+.
Return op name rocdl.cvt.scalef32.2xpk16.bf6.f32 as a bitstring.
rocdl.cvt.scalef32.2xpk16.bf6.f32 - Scale and convert two vector<16xf32> to 32 packed bf6
Operands
src0- Single,ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16src1- Single,ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
Description
Convert 32 single-precision float values, packed into two length-16
vectors that will be logically concanenated, to packed bf6, dividing by the exponent part of scale
before doing so.
Return op name rocdl.cvt.scalef32.2xpk16.fp6.f32 as a bitstring.
rocdl.cvt.scalef32.2xpk16.fp6.f32 - Scale and convert two vector<16xf32> to 32 packed fp6
Operands
src0- Single,ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16src1- Single,ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
Description
Convert 32 single-precision float values, packed into two length-16
vectors that will be logically concanenated, to packed fp6, dividing by the exponent part of scale
before doing so.
Return op name rocdl.cvt.scalef32.f16.bf8 as a bitstring.
rocdl.cvt.scalef32.f16.bf8 - Scaled convert bf8 from packed vector to f16, updating tied result
Attributes
srcSelIndex- Single,I32Attr, 32-bit signless integer attributedstLoHiSel- Single,I1Attr, 1-bit signless integer attribute
Operands
oldVdst- Single,ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2src- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2
Description
Convert a bf8 byte from src, selected by
srcSelIndex, to f16 while multiplying it by the expontent of scale,
and place it into the dstLoHiSelth bit
of oldVdst preserving the other element of that vector in
the return value.
The bytes are stored as an i32 and not a <4 x i8>.
Return op name rocdl.cvt.scalef32.f16.fp8 as a bitstring.
rocdl.cvt.scalef32.f16.fp8 - Scaled convert fp8 from packed vector to f16, updating tied result
Attributes
srcSelIndex- Single,I32Attr, 32-bit signless integer attributedstLoHiSel- Single,I1Attr, 1-bit signless integer attribute
Operands
oldVdst- Single,ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2src- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2
Description
Convert a fp8 byte from src, selected by
srcSelIndex, to f16 while multiplying it by the expontent of scale,
and place it into the dstLoHiSelth bit
of oldVdst preserving the other element of that vector in
the return value.
The bytes are stored as an i32 and not a <4 x i8>.
Return op name rocdl.cvt.scalef32.f32.bf8 as a bitstring.
rocdl.cvt.scalef32.f32.bf8 - Scaled convert bf8 from packed vector to f32
Attributes
srcSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,F32, 32-bit float
Description
Convert a bf8 byte from src, selected by
srcSelIndex, to f32, multiplying it by the exponent of scale.
The bytes are stored in an i32, not a <4 x i8>.
Return op name rocdl.cvt.scalef32.f32.fp8 as a bitstring.
rocdl.cvt.scalef32.f32.fp8 - Scaled convert fp8 from packed vector to f32
Attributes
srcSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,F32, 32-bit float
Description
Convert a fp8 byte from src, selected by
srcSelIndex, to f32, multiplying it by the exponent of scale.
The bytes are stored in an i32, not a <4 x i8>.
Return op name rocdl.cvt.scalef32.pk8.bf8.bf16 as a bitstring.
rocdl.cvt.scalef32.pk8.bf8.bf16 - Scale and convert packed bf16 to packed bf8
Operands
src- Single,ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
Description
Convert 8 packed bf16 values to packed bf8, multiplying by the exponent part of scale
before doing so. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.pk8.bf8.f16 as a bitstring.
rocdl.cvt.scalef32.pk8.bf8.f16 - Scale and convert packed f16 to packed bf8
Operands
src- Single,ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
Description
Convert 8 packed f16 values to packed bf8, multiplying by the exponent part of scale
before doing so. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.pk8.bf8.f32 as a bitstring.
rocdl.cvt.scalef32.pk8.bf8.f32 - Scale and convert packed f32 to packed bf8
Operands
src- Single,ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
Description
Convert 8 packed f32 values to packed bf8, multiplying by the exponent part of scale
before doing so. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.pk8.fp4.bf16 as a bitstring.
rocdl.cvt.scalef32.pk8.fp4.bf16 - Scale and convert packed bf16 to packed fp4
Operands
src- Single,ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8scale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert 8 packed bf16 values to packed fp4, multiplying by the exponent part of scale
before doing so. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.pk8.fp4.f16 as a bitstring.
rocdl.cvt.scalef32.pk8.fp4.f16 - Scale and convert packed f16 to packed fp4
Operands
src- Single,ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8scale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert 8 packed f16 values to packed fp4, multiplying by the exponent part of scale
before doing so. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.pk8.fp4.f32 as a bitstring.
rocdl.cvt.scalef32.pk8.fp4.f32 - Scale and convert packed f32 to packed fp4
Operands
src- Single,ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8scale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert 8 packed f32 values to packed fp4, multiplying by the exponent part of scale
before doing so. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.pk8.fp8.bf16 as a bitstring.
rocdl.cvt.scalef32.pk8.fp8.bf16 - Scale and convert packed bf16 to packed fp8
Operands
src- Single,ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
Description
Convert 8 packed bf16 values to packed fp8, multiplying by the exponent part of scale
before doing so. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.pk8.fp8.f16 as a bitstring.
rocdl.cvt.scalef32.pk8.fp8.f16 - Scale and convert packed f16 to packed fp8
Operands
src- Single,ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
Description
Convert 8 packed f16 values to packed fp8, multiplying by the exponent part of scale
before doing so. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.pk8.fp8.f32 as a bitstring.
rocdl.cvt.scalef32.pk8.fp8.f32 - Scale and convert packed f32 to packed fp8
Operands
src- Single,ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
Description
Convert 8 packed f32 values to packed fp8, multiplying by the exponent part of scale
before doing so. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.pk16.bf6.bf16 as a bitstring.
rocdl.cvt.scalef32.pk16.bf6.bf16 - Scale and convert packed bf16 to packed bf6
Operands
src- Single,ROCDL_V16BF16Type, fixed-length vector of bfloat16 type values of length 16scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
Description
Convert 8 packed bf16 values to packed bf6, multiplying by the exponent part of scale
before doing so. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.pk16.bf6.f16 as a bitstring.
rocdl.cvt.scalef32.pk16.bf6.f16 - Scale and convert packed f16 to packed bf6
Operands
src- Single,ROCDL_V16F16Type, fixed-length vector of 16-bit float values of length 16scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
Description
Convert 8 packed f16 values to packed bf6, multiplying by the exponent part of scale
before doing so. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.pk16.bf6.f32 as a bitstring.
rocdl.cvt.scalef32.pk16.bf6.f32 - Scale and convert packed f32 to packed bf6
Operands
src- Single,ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
Description
Convert 8 packed f32 values to packed bf6, multiplying by the exponent part of scale
before doing so. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.pk16.fp6.bf16 as a bitstring.
rocdl.cvt.scalef32.pk16.fp6.bf16 - Scale and convert packed bf16 to packed fp6
Operands
src- Single,ROCDL_V16BF16Type, fixed-length vector of bfloat16 type values of length 16scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
Description
Convert 8 packed bf16 values to packed fp6, multiplying by the exponent part of scale
before doing so. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.pk16.fp6.f16 as a bitstring.
rocdl.cvt.scalef32.pk16.fp6.f16 - Scale and convert packed f16 to packed fp6
Operands
src- Single,ROCDL_V16F16Type, fixed-length vector of 16-bit float values of length 16scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
Description
Convert 8 packed f16 values to packed fp6, multiplying by the exponent part of scale
before doing so. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.pk16.fp6.f32 as a bitstring.
rocdl.cvt.scalef32.pk16.fp6.f32 - Scale and convert packed f32 to packed fp6
Operands
src- Single,ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
Description
Convert 8 packed f32 values to packed fp6, multiplying by the exponent part of scale
before doing so. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.pk32.bf6.bf16 as a bitstring.
rocdl.cvt.scalef32.pk32.bf6.bf16 - Scale and convert packed bf16 to packed bf6
Operands
src- Single,ROCDL_V32BF16Type, fixed-length vector of bfloat16 type values of length 32scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
Description
Convert 32 packed bf16 values to packed bf6, dividing by the exponent part of scale
before doing so.
Return op name rocdl.cvt.scalef32.pk32.bf6.f16 as a bitstring.
rocdl.cvt.scalef32.pk32.bf6.f16 - Scale and convert packed f16 to packed bf6
Operands
src- Single,ROCDL_V32F16Type, fixed-length vector of 16-bit float values of length 32scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
Description
Convert 32 packed f16 values to packed bf6, dividing by the exponent part of scale
before doing so.
Return op name rocdl.cvt.scalef32.pk32.bf16.bf6 as a bitstring.
rocdl.cvt.scalef32.pk32.bf16.bf6 - Scale and convert packed bf6 to packed bf16
Operands
src- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V32BF16Type, fixed-length vector of bfloat16 type values of length 32
Description
Convert 32 packed bf6 values to packed bf16, multiplying by the exponent part of scale
before doing so.
Return op name rocdl.cvt.scalef32.pk32.bf16.fp6 as a bitstring.
rocdl.cvt.scalef32.pk32.bf16.fp6 - Scale and convert packed fp6 to packed bf16
Operands
src- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V32BF16Type, fixed-length vector of bfloat16 type values of length 32
Description
Convert 32 packed fp6 values to packed bf16, multiplying by the exponent part of scale
before doing so.
Return op name rocdl.cvt.scalef32.pk32.f16.bf6 as a bitstring.
rocdl.cvt.scalef32.pk32.f16.bf6 - Scale and convert packed bf6 to packed f16
Operands
src- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V32F16Type, fixed-length vector of 16-bit float values of length 32
Description
Convert 32 packed bf6 values to packed f16, multiplying by the exponent part of scale
before doing so.
Return op name rocdl.cvt.scalef32.pk32.f16.fp6 as a bitstring.
rocdl.cvt.scalef32.pk32.f16.fp6 - Scale and convert packed fp6 to packed f16
Operands
src- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V32F16Type, fixed-length vector of 16-bit float values of length 32
Description
Convert 32 packed fp6 values to packed f16, multiplying by the exponent part of scale
before doing so.
Return op name rocdl.cvt.scalef32.pk32.f32.bf6 as a bitstring.
rocdl.cvt.scalef32.pk32.f32.bf6 - Scale and convert packed bf6 to packed f32
Operands
src- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V32F32Type, fixed-length vector of 32-bit float values of length 32
Description
Convert 32 packed bf6 values to packed f32, multiplying by the exponent part of scale
before doing so.
Return op name rocdl.cvt.scalef32.pk32.f32.fp6 as a bitstring.
rocdl.cvt.scalef32.pk32.f32.fp6 - Scale and convert packed fp6 to packed f32
Operands
src- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V32F32Type, fixed-length vector of 32-bit float values of length 32
Description
Convert 32 packed fp6 values to packed f32, multiplying by the exponent part of scale
before doing so.
Return op name rocdl.cvt.scalef32.pk32.fp6.bf16 as a bitstring.
rocdl.cvt.scalef32.pk32.fp6.bf16 - Scale and convert packed bf16 to packed fp6
Operands
src- Single,ROCDL_V32BF16Type, fixed-length vector of bfloat16 type values of length 32scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
Description
Convert 32 packed bf16 values to packed fp6, dividing by the exponent part of scale
before doing so.
Return op name rocdl.cvt.scalef32.pk32.fp6.f16 as a bitstring.
rocdl.cvt.scalef32.pk32.fp6.f16 - Scale and convert packed f16 to packed fp6
Operands
src- Single,ROCDL_V32F16Type, fixed-length vector of 16-bit float values of length 32scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
Description
Convert 32 packed f16 values to packed fp6, dividing by the exponent part of scale
before doing so.
Return op name rocdl.cvt.scalef32.pk.bf8.bf16 as a bitstring.
rocdl.cvt.scalef32.pk.bf8.bf16 - Scaled convert two bf16to two bf8, updating packed vector
Attributes
dstLoHiSel- Single,I1Attr, 1-bit signless integer attribute
Operands
oldVdst- Single,ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2src0- Single,ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2
Description
Convert two bf16 values in src0 to two bf8 bytes, dividing by the exponent in scale. The bytes are
packed into a 16-bit value which is inserted into oldVdst at the
dstLoHiSel position, with the entire updated vector being returned.
Return op name rocdl.cvt.scalef32.pk.bf8.f16 as a bitstring.
rocdl.cvt.scalef32.pk.bf8.f16 - Scaled convert two f16to two bf8, updating packed vector
Attributes
dstLoHiSel- Single,I1Attr, 1-bit signless integer attribute
Operands
oldVdst- Single,ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2src0- Single,ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2
Description
Convert two f16 values in src0 to two bf8 bytes, dividing by the exponent in scale. The bytes are
packed into a 16-bit value which is inserted into oldVdst at the
dstLoHiSel position, with the entire updated vector being returned.
Return op name rocdl.cvt.scalef32.pk.bf8.f32 as a bitstring.
rocdl.cvt.scalef32.pk.bf8.f32 - Scaled convert two f32 to two bf8, updating packed vector
Attributes
dstLoHiSel- Single,I1Attr, 1-bit signless integer attribute
Operands
oldVdst- Single,ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2src0- Single,F32, 32-bit floatsrc1- Single,F32, 32-bit floatscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2
Description
Convert two f32 values in src0 and src1 to two bf8 bytes,
dividing by the exponent in scale. The bytes are packed into
a 16-bit value which is inserted into oldVdst at the dstLoHiSel
position, with the entire updated vector being returned.
Return op name rocdl.cvt.scalef32.pk.bf16.bf8 as a bitstring.
rocdl.cvt.scalef32.pk.bf16.bf8 - Scaled convert two bf8to two bf16
Attributes
srcLoHiSel- Single,I1Attr, 1-bit signless integer attribute
Operands
src- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2
Description
Convert two packed bf8 values in src0 to two bf16 values, multiplying by the exponent in scale.
The two values to be converted are selected from the low or high half
of src (a packed vector represented as an i32)
on the basis of srcLoHiSel.
Return op name rocdl.cvt.scalef32.pk.bf16.fp4 as a bitstring.
rocdl.cvt.scalef32.pk.bf16.fp4 - Scale and convert two packed fp4 to packed bf16
Attributes
srcSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2
Description
Convert two packed fp4 (f4E2M1) values stored as one byte of a 32-bit integer
to packed bf16, multiplying by the exponent part of scale
before doing so.
The byte to convert is chosen by srcSelIndex.
Return op name rocdl.cvt.scalef32.pk.bf16.fp8 as a bitstring.
rocdl.cvt.scalef32.pk.bf16.fp8 - Scaled convert two fp8to two bf16
Attributes
srcLoHiSel- Single,I1Attr, 1-bit signless integer attribute
Operands
src- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2
Description
Convert two packed fp8 values in src0 to two bf16 values, multiplying by the exponent in scale.
The two values to be converted are selected from the low or high half
of src (a packed vector represented as an i32)
on the basis of srcLoHiSel.
Return op name rocdl.cvt.scalef32.pk.f16.bf8 as a bitstring.
rocdl.cvt.scalef32.pk.f16.bf8 - Scaled convert two bf8to two f16
Attributes
srcLoHiSel- Single,I1Attr, 1-bit signless integer attribute
Operands
src- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2
Description
Convert two packed bf8 values in src0 to two f16 values, multiplying by the exponent in scale.
The two values to be converted are selected from the low or high half
of src (a packed vector represented as an i32)
on the basis of srcLoHiSel.
Return op name rocdl.cvt.scalef32.pk.f16.fp4 as a bitstring.
rocdl.cvt.scalef32.pk.f16.fp4 - Scale and convert two packed fp4 to packed f16
Attributes
srcSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2
Description
Convert two packed fp4 (f4E2M1) values stored as one byte of a 32-bit integer
to packed f16, multiplying by the exponent part of scale
before doing so.
The byte to convert is chosen by srcSelIndex.
Return op name rocdl.cvt.scalef32.pk.f16.fp8 as a bitstring.
rocdl.cvt.scalef32.pk.f16.fp8 - Scaled convert two fp8to two f16
Attributes
srcLoHiSel- Single,I1Attr, 1-bit signless integer attribute
Operands
src- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2
Description
Convert two packed fp8 values in src0 to two f16 values, multiplying by the exponent in scale.
The two values to be converted are selected from the low or high half
of src (a packed vector represented as an i32)
on the basis of srcLoHiSel.
Return op name rocdl.cvt.scalef32.pk.f32.bf8 as a bitstring.
rocdl.cvt.scalef32.pk.f32.bf8 - Scaled convert two bf8to two f32
Attributes
srcLoHiSel- Single,I1Attr, 1-bit signless integer attribute
Operands
src- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2F32Type, fixed-length vector of 32-bit float values of length 2
Description
Convert two packed bf8 values in src0 to two f32 values, multiplying by the exponent in scale.
The two values to be converted are selected from the low or high half
of src (a packed vector represented as an i32)
on the basis of srcLoHiSel.
Return op name rocdl.cvt.scalef32.pk.f32.fp4 as a bitstring.
rocdl.cvt.scalef32.pk.f32.fp4 - Scale and convert two packed fp4 to packed f32
Attributes
srcSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
src- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2F32Type, fixed-length vector of 32-bit float values of length 2
Description
Convert two packed fp4 (f4E2M1) values stored as one byte of a 32-bit integer
to packed f32, multiplying by the exponent part of scale
before doing so.
The byte to convert is chosen by srcSelIndex.
Return op name rocdl.cvt.scalef32.pk.f32.fp8 as a bitstring.
rocdl.cvt.scalef32.pk.f32.fp8 - Scaled convert two fp8to two f32
Attributes
srcLoHiSel- Single,I1Attr, 1-bit signless integer attribute
Operands
src- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2F32Type, fixed-length vector of 32-bit float values of length 2
Description
Convert two packed fp8 values in src0 to two f32 values, multiplying by the exponent in scale.
The two values to be converted are selected from the low or high half
of src (a packed vector represented as an i32)
on the basis of srcLoHiSel.
Return op name rocdl.cvt.scalef32.pk.fp4.bf16 as a bitstring.
rocdl.cvt.scalef32.pk.fp4.bf16 - Scale and convert two bf16 to packed fp4, updating tied vector
Attributes
dstSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
oldVdst- Single,I32, 32-bit signless integersrc- Single,ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2scale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert two packed bf16 values to packed
fp4, dividing by the exponent part of scale
before doing so.
The two scaled values are packed into a byte.
That byte is used to update the dstSelIndexth
byte of oldVdst, which is returned in its entirity.
Return op name rocdl.cvt.scalef32.pk.fp4.f16 as a bitstring.
rocdl.cvt.scalef32.pk.fp4.f16 - Scale and convert two f16 to packed fp4, updating tied vector
Attributes
dstSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
oldVdst- Single,I32, 32-bit signless integersrc- Single,ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2scale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert two packed f16 values to packed
fp4, dividing by the exponent part of scale
before doing so.
The two scaled values are packed into a byte.
That byte is used to update the dstSelIndexth
byte of oldVdst, which is returned in its entirity.
Return op name rocdl.cvt.scalef32.pk.fp4.f32 as a bitstring.
rocdl.cvt.scalef32.pk.fp4.f32 - Scale and convert two f32 values to two packed fp4, updating tied vector
Attributes
dstSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
oldVdst- Single,I32, 32-bit signless integersrc0- Single,F32, 32-bit floatsrc1- Single,F32, 32-bit floatscale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert two single-precision float values, passed in src0 and src1
into two fp4 values, dividing them by the expontent part of scale
before doing so.
The two scaled values are packed into a byte.
That byte is used to update the dstSelIndexth
byte of oldVdst, which is returned in its entirity.
Example:
// Scaled convert two f32 values to packed fp4 in byte 0 of old.
%0 = rocdl.cvt.scalef32.pk.fp4.f32 %a, %b, %scale -> %old[0] : i32
Return op name rocdl.cvt.scalef32.pk.fp8.bf16 as a bitstring.
rocdl.cvt.scalef32.pk.fp8.bf16 - Scaled convert two bf16to two fp8, updating packed vector
Attributes
dstLoHiSel- Single,I1Attr, 1-bit signless integer attribute
Operands
oldVdst- Single,ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2src0- Single,ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2
Description
Convert two bf16 values in src0 to two fp8 bytes, dividing by the exponent in scale. The bytes are
packed into a 16-bit value which is inserted into oldVdst at the
dstLoHiSel position, with the entire updated vector being returned.
Return op name rocdl.cvt.scalef32.pk.fp8.f16 as a bitstring.
rocdl.cvt.scalef32.pk.fp8.f16 - Scaled convert two f16to two fp8, updating packed vector
Attributes
dstLoHiSel- Single,I1Attr, 1-bit signless integer attribute
Operands
oldVdst- Single,ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2src0- Single,ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2scale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2
Description
Convert two f16 values in src0 to two fp8 bytes, dividing by the exponent in scale. The bytes are
packed into a 16-bit value which is inserted into oldVdst at the
dstLoHiSel position, with the entire updated vector being returned.
Return op name rocdl.cvt.scalef32.pk.fp8.f32 as a bitstring.
rocdl.cvt.scalef32.pk.fp8.f32 - Scaled convert two f32 to two fp8, updating packed vector
Attributes
dstLoHiSel- Single,I1Attr, 1-bit signless integer attribute
Operands
oldVdst- Single,ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2src0- Single,F32, 32-bit floatsrc1- Single,F32, 32-bit floatscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2
Description
Convert two f32 values in src0 and src1 to two fp8 bytes,
dividing by the exponent in scale. The bytes are packed into
a 16-bit value which is inserted into oldVdst at the dstLoHiSel
position, with the entire updated vector being returned.
Return op name rocdl.cvt.scalef32.sr.bf8.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.bf8.bf16 - Scaled convert bf16to bf8 with stochiastic rounding, updating packed vector
Attributes
dstSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
oldVdst- Single,I32, 32-bit signless integersrc0- Single,BF16, bfloat16 typeseed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert a bf16 value in src0 to a bf8 bytes, dividing by the exponent in scale and using seed
for stochiastic rounding. Place the resulting byte in the
dstSelIndexth bit of oldVdst and return the entire packed vector,
which is stored as an i32.
Return op name rocdl.cvt.scalef32.sr.bf8.f16 as a bitstring.
rocdl.cvt.scalef32.sr.bf8.f16 - Scaled convert f16to bf8 with stochiastic rounding, updating packed vector
Attributes
dstSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
oldVdst- Single,I32, 32-bit signless integersrc0- Single,F16, 16-bit floatseed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert a f16 value in src0 to a bf8 bytes, dividing by the exponent in scale and using seed
for stochiastic rounding. Place the resulting byte in the
dstSelIndexth bit of oldVdst and return the entire packed vector,
which is stored as an i32.
Return op name rocdl.cvt.scalef32.sr.bf8.f32 as a bitstring.
rocdl.cvt.scalef32.sr.bf8.f32 - Scaled convert f32to bf8 with stochiastic rounding, updating packed vector
Attributes
dstSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
oldVdst- Single,I32, 32-bit signless integersrc0- Single,F32, 32-bit floatseed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert a f32 value in src0 to a bf8 bytes, dividing by the exponent in scale and using seed
for stochiastic rounding. Place the resulting byte in the
dstSelIndexth bit of oldVdst and return the entire packed vector,
which is stored as an i32.
Return op name rocdl.cvt.scalef32.sr.fp8.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.fp8.bf16 - Scaled convert bf16to fp8 with stochiastic rounding, updating packed vector
Attributes
dstSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
oldVdst- Single,I32, 32-bit signless integersrc0- Single,BF16, bfloat16 typeseed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert a bf16 value in src0 to a fp8 bytes, dividing by the exponent in scale and using seed
for stochiastic rounding. Place the resulting byte in the
dstSelIndexth bit of oldVdst and return the entire packed vector,
which is stored as an i32.
Return op name rocdl.cvt.scalef32.sr.fp8.f16 as a bitstring.
rocdl.cvt.scalef32.sr.fp8.f16 - Scaled convert f16to fp8 with stochiastic rounding, updating packed vector
Attributes
dstSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
oldVdst- Single,I32, 32-bit signless integersrc0- Single,F16, 16-bit floatseed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert a f16 value in src0 to a fp8 bytes, dividing by the exponent in scale and using seed
for stochiastic rounding. Place the resulting byte in the
dstSelIndexth bit of oldVdst and return the entire packed vector,
which is stored as an i32.
Return op name rocdl.cvt.scalef32.sr.fp8.f32 as a bitstring.
rocdl.cvt.scalef32.sr.fp8.f32 - Scaled convert f32to fp8 with stochiastic rounding, updating packed vector
Attributes
dstSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
oldVdst- Single,I32, 32-bit signless integersrc0- Single,F32, 32-bit floatseed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert a f32 value in src0 to a fp8 bytes, dividing by the exponent in scale and using seed
for stochiastic rounding. Place the resulting byte in the
dstSelIndexth bit of oldVdst and return the entire packed vector,
which is stored as an i32.
Return op name rocdl.cvt.scalef32.sr.pk8.bf8.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.bf8.bf16 - Scale and convert packed bf16 to packed bf8 with stochastic rounding
Operands
src- Single,ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
Description
Convert 8 packed bf16 values to packed bf8, multiplying by the exponent part of scale
before doing so and apply stochastic rounding. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.sr.pk8.bf8.f16 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.bf8.f16 - Scale and convert packed f16 to packed bf8 with stochastic rounding
Operands
src- Single,ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
Description
Convert 8 packed f16 values to packed bf8, multiplying by the exponent part of scale
before doing so and apply stochastic rounding. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.sr.pk8.bf8.f32 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.bf8.f32 - Scale and convert packed f32 to packed bf8 with stochastic rounding
Operands
src- Single,ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
Description
Convert 8 packed f32 values to packed bf8, multiplying by the exponent part of scale
before doing so and apply stochastic rounding. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.sr.pk8.fp4.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.fp4.bf16 - Scale and convert packed bf16 to packed fp4 with stochastic rounding
Operands
src- Single,ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert 8 packed bf16 values to packed fp4, multiplying by the exponent part of scale
before doing so and apply stochastic rounding. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.sr.pk8.fp4.f16 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.fp4.f16 - Scale and convert packed f16 to packed fp4 with stochastic rounding
Operands
src- Single,ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert 8 packed f16 values to packed fp4, multiplying by the exponent part of scale
before doing so and apply stochastic rounding. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.sr.pk8.fp4.f32 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.fp4.f32 - Scale and convert packed f32 to packed fp4 with stochastic rounding
Operands
src- Single,ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert 8 packed f32 values to packed fp4, multiplying by the exponent part of scale
before doing so and apply stochastic rounding. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.sr.pk8.fp8.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.fp8.bf16 - Scale and convert packed bf16 to packed fp8 with stochastic rounding
Operands
src- Single,ROCDL_V8BF16Type, fixed-length vector of bfloat16 type values of length 8seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
Description
Convert 8 packed bf16 values to packed fp8, multiplying by the exponent part of scale
before doing so and apply stochastic rounding. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.sr.pk8.fp8.f16 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.fp8.f16 - Scale and convert packed f16 to packed fp8 with stochastic rounding
Operands
src- Single,ROCDL_V8F16Type, fixed-length vector of 16-bit float values of length 8seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
Description
Convert 8 packed f16 values to packed fp8, multiplying by the exponent part of scale
before doing so and apply stochastic rounding. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.sr.pk8.fp8.f32 as a bitstring.
rocdl.cvt.scalef32.sr.pk8.fp8.f32 - Scale and convert packed f32 to packed fp8 with stochastic rounding
Operands
src- Single,ROCDL_V8F32Type, fixed-length vector of 32-bit float values of length 8seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V2I32Type, fixed-length vector of 32-bit signless integer values of length 2
Description
Convert 8 packed f32 values to packed fp8, multiplying by the exponent part of scale
before doing so and apply stochastic rounding. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.sr.pk16.bf6.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.pk16.bf6.bf16 - Scale and convert packed bf16 to packed bf6 with stochastic rounding
Operands
src- Single,ROCDL_V16BF16Type, fixed-length vector of bfloat16 type values of length 16seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
Description
Convert 8 packed bf16 values to packed bf6, multiplying by the exponent part of scale
before doing so and apply stochastic rounding. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.sr.pk16.bf6.f16 as a bitstring.
rocdl.cvt.scalef32.sr.pk16.bf6.f16 - Scale and convert packed f16 to packed bf6 with stochastic rounding
Operands
src- Single,ROCDL_V16F16Type, fixed-length vector of 16-bit float values of length 16seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
Description
Convert 8 packed f16 values to packed bf6, multiplying by the exponent part of scale
before doing so and apply stochastic rounding. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.sr.pk16.bf6.f32 as a bitstring.
rocdl.cvt.scalef32.sr.pk16.bf6.f32 - Scale and convert packed f32 to packed bf6 with stochastic rounding
Operands
src- Single,ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
Description
Convert 8 packed f32 values to packed bf6, multiplying by the exponent part of scale
before doing so and apply stochastic rounding. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.sr.pk16.fp6.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.pk16.fp6.bf16 - Scale and convert packed bf16 to packed fp6 with stochastic rounding
Operands
src- Single,ROCDL_V16BF16Type, fixed-length vector of bfloat16 type values of length 16seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
Description
Convert 8 packed bf16 values to packed fp6, multiplying by the exponent part of scale
before doing so and apply stochastic rounding. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.sr.pk16.fp6.f16 as a bitstring.
rocdl.cvt.scalef32.sr.pk16.fp6.f16 - Scale and convert packed f16 to packed fp6 with stochastic rounding
Operands
src- Single,ROCDL_V16F16Type, fixed-length vector of 16-bit float values of length 16seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
Description
Convert 8 packed f16 values to packed fp6, multiplying by the exponent part of scale
before doing so and apply stochastic rounding. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.sr.pk16.fp6.f32 as a bitstring.
rocdl.cvt.scalef32.sr.pk16.fp6.f32 - Scale and convert packed f32 to packed fp6 with stochastic rounding
Operands
src- Single,ROCDL_V16F32Type, fixed-length vector of 32-bit float values of length 16seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V3I32Type, fixed-length vector of 32-bit signless integer values of length 3
Description
Convert 8 packed f32 values to packed fp6, multiplying by the exponent part of scale
before doing so and apply stochastic rounding. This op is for gfx1250+ arch.
Return op name rocdl.cvt.scalef32.sr.pk32.bf6.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.pk32.bf6.bf16 - Scale and convert packed bf16 to packed bf6 with stochiastic rounding
Operands
src- Single,ROCDL_V32BF16Type, fixed-length vector of bfloat16 type values of length 32seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
Description
Convert 32 packed bf16 values to packed bf6, dividing by the exponent part of scale
before doing so and applying random rounding derived from
seed.
Return op name rocdl.cvt.scalef32.sr.pk32.bf6.f16 as a bitstring.
rocdl.cvt.scalef32.sr.pk32.bf6.f16 - Scale and convert packed f16 to packed bf6 with stochiastic rounding
Operands
src- Single,ROCDL_V32F16Type, fixed-length vector of 16-bit float values of length 32seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
Description
Convert 32 packed f16 values to packed bf6, dividing by the exponent part of scale
before doing so and applying random rounding derived from
seed.
Return op name rocdl.cvt.scalef32.sr.pk32.bf6.f32 as a bitstring.
rocdl.cvt.scalef32.sr.pk32.bf6.f32 - Scale and convert packed f32 to packed bf6 with stochiastic rounding
Operands
src- Single,ROCDL_V32F32Type, fixed-length vector of 32-bit float values of length 32seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
Description
Convert 32 packed f32 values to packed bf6, dividing by the exponent part of scale
before doing so and applying random rounding derived from
seed.
Return op name rocdl.cvt.scalef32.sr.pk32.fp6.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.pk32.fp6.bf16 - Scale and convert packed bf16 to packed fp6 with stochiastic rounding
Operands
src- Single,ROCDL_V32BF16Type, fixed-length vector of bfloat16 type values of length 32seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
Description
Convert 32 packed bf16 values to packed fp6, dividing by the exponent part of scale
before doing so and applying random rounding derived from
seed.
Return op name rocdl.cvt.scalef32.sr.pk32.fp6.f16 as a bitstring.
rocdl.cvt.scalef32.sr.pk32.fp6.f16 - Scale and convert packed f16 to packed fp6 with stochiastic rounding
Operands
src- Single,ROCDL_V32F16Type, fixed-length vector of 16-bit float values of length 32seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
Description
Convert 32 packed f16 values to packed fp6, dividing by the exponent part of scale
before doing so and applying random rounding derived from
seed.
Return op name rocdl.cvt.scalef32.sr.pk32.fp6.f32 as a bitstring.
rocdl.cvt.scalef32.sr.pk32.fp6.f32 - Scale and convert packed f32 to packed fp6 with stochiastic rounding
Operands
src- Single,ROCDL_V32F32Type, fixed-length vector of 32-bit float values of length 32seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,ROCDL_V6I32Type, fixed-length vector of 32-bit signless integer values of length 6
Description
Convert 32 packed f32 values to packed fp6, dividing by the exponent part of scale
before doing so and applying random rounding derived from
seed.
Return op name rocdl.cvt.scalef32.sr.pk.fp4.bf16 as a bitstring.
rocdl.cvt.scalef32.sr.pk.fp4.bf16 - Scale and convert two bf16 to packed fp4 with stochiastic rounding, updating tied vector
Attributes
dstSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
oldVdst- Single,I32, 32-bit signless integersrc- Single,ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert two packed bf16 values to packed
fp4, dividing by the exponent part of scale
before doing so and using seed as the random seed for
stochiastic rounding.
The two scaled values are packed (little-endian)
into a byte. That byte is used to update the dstSelIndexth
byte of oldVdst, which is returned in its entirity.
Return op name rocdl.cvt.scalef32.sr.pk.fp4.f16 as a bitstring.
rocdl.cvt.scalef32.sr.pk.fp4.f16 - Scale and convert two f16 to packed fp4 with stochiastic rounding, updating tied vector
Attributes
dstSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
oldVdst- Single,I32, 32-bit signless integersrc- Single,ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert two packed f16 values to packed
fp4, dividing by the exponent part of scale
before doing so and using seed as the random seed for
stochiastic rounding.
The two scaled values are packed (little-endian)
into a byte. That byte is used to update the dstSelIndexth
byte of oldVdst, which is returned in its entirity.
Return op name rocdl.cvt.scalef32.sr.pk.fp4.f32 as a bitstring.
rocdl.cvt.scalef32.sr.pk.fp4.f32 - Scale and convert two f32 to packed fp4 with stochiastic rounding, updating tied vector
Attributes
dstSelIndex- Single,I32Attr, 32-bit signless integer attribute
Operands
oldVdst- Single,I32, 32-bit signless integersrc- Single,ROCDL_V2F32Type, fixed-length vector of 32-bit float values of length 2seed- Single,I32, 32-bit signless integerscale- Single,F32, 32-bit float
Results
res- Single,I32, 32-bit signless integer
Description
Convert two packed f32 values to packed
fp4, dividing by the exponent part of scale
before doing so and using seed as the random seed for
stochiastic rounding.
The two scaled values are packed (little-endian)
into a byte. That byte is used to update the dstSelIndexth
byte of oldVdst, which is returned in its entirity.
Return op name rocdl.cvt.sr.bf8.f32 as a bitstring.
rocdl.cvt.sr.bf8.f32 - Convert f32 to bf8, stochiastic rounding
Attributes
byteSel- Single,I32Attr, 32-bit signless integer attribute
Operands
srcA- Single,F32, 32-bit floatsrcB- Single,I32, 32-bit signless integerold- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Convert srcA to bf8, adding the rounding factor from srcB,
and store into the byteSelth byte of old, preserving the others.
Example:
// Stochastic rounding convert f32 to bf8 in byte 2 of old.
%0 = rocdl.cvt.sr.bf8.f32 %val, %stoch -> %old[2] : i32
Return op name rocdl.cvt.sr.fp8.f32 as a bitstring.
rocdl.cvt.sr.fp8.f32 - Convert f32 to fp8, stochiastic rounding
Attributes
byteSel- Single,I32Attr, 32-bit signless integer attribute
Operands
srcA- Single,F32, 32-bit floatsrcB- Single,I32, 32-bit signless integerold- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Convert srcA to fp8, adding the rounding factor from srcB,
and store into the byteSelth byte of old, preserving the others.
Example:
// Stochastic rounding convert f32 to fp8 in byte 3 of old.
%0 = rocdl.cvt.sr.fp8.f32 %val, %stoch -> %old[3] : i32
Return op name rocdl.dot4.f32.bf8.bf8 as a bitstring.
rocdl.dot4.f32.bf8.bf8
Operands
a- Single, anonymous/composite constraint, 32-bit signless integerb- Single, anonymous/composite constraint, 32-bit signless integerc- Single, anonymous/composite constraint, 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float
Description
Packed intra-lane dot-product with no clamp control.
Computes res = sum_i a[i]*b[i] + c. Covers the full-f16/bf16
accumulator forms (fdot2.f16.f16, fdot2.bf16.bf16) and the
FP8/BF8 dot4.f32.* variants, whose hardware instructions have no
CLAMP bit in their modifier word.
Example:
%r = rocdl.dot4.f32.bf8.bf8 %a, %b, %c : (i32, i32, f32) -> f32
Return op name rocdl.dot4.f32.bf8.fp8 as a bitstring.
rocdl.dot4.f32.bf8.fp8
Operands
a- Single, anonymous/composite constraint, 32-bit signless integerb- Single, anonymous/composite constraint, 32-bit signless integerc- Single, anonymous/composite constraint, 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float
Description
Packed intra-lane dot-product with no clamp control.
Computes res = sum_i a[i]*b[i] + c. Covers the full-f16/bf16
accumulator forms (fdot2.f16.f16, fdot2.bf16.bf16) and the
FP8/BF8 dot4.f32.* variants, whose hardware instructions have no
CLAMP bit in their modifier word.
Example:
%r = rocdl.dot4.f32.bf8.fp8 %a, %b, %c : (i32, i32, f32) -> f32
Return op name rocdl.dot4.f32.fp8.bf8 as a bitstring.
rocdl.dot4.f32.fp8.bf8
Operands
a- Single, anonymous/composite constraint, 32-bit signless integerb- Single, anonymous/composite constraint, 32-bit signless integerc- Single, anonymous/composite constraint, 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float
Description
Packed intra-lane dot-product with no clamp control.
Computes res = sum_i a[i]*b[i] + c. Covers the full-f16/bf16
accumulator forms (fdot2.f16.f16, fdot2.bf16.bf16) and the
FP8/BF8 dot4.f32.* variants, whose hardware instructions have no
CLAMP bit in their modifier word.
Example:
%r = rocdl.dot4.f32.fp8.bf8 %a, %b, %c : (i32, i32, f32) -> f32
Return op name rocdl.dot4.f32.fp8.fp8 as a bitstring.
rocdl.dot4.f32.fp8.fp8
Operands
a- Single, anonymous/composite constraint, 32-bit signless integerb- Single, anonymous/composite constraint, 32-bit signless integerc- Single, anonymous/composite constraint, 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float
Description
Packed intra-lane dot-product with no clamp control.
Computes res = sum_i a[i]*b[i] + c. Covers the full-f16/bf16
accumulator forms (fdot2.f16.f16, fdot2.bf16.bf16) and the
FP8/BF8 dot4.f32.* variants, whose hardware instructions have no
CLAMP bit in their modifier word.
Example:
%r = rocdl.dot4.f32.fp8.fp8 %a, %b, %c : (i32, i32, f32) -> f32
Return op name rocdl.ds.atomic.async.barrier.arrive.b64 as a bitstring.
rocdl.ds.atomic.async.barrier.arrive.b64
Attributes
alias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
barrierPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Description
Waits on a given DS barrier and decrements pending count by -1. Stays in order with ASYNC loads to LDS, and uses ASYNCcnt to track its completion. Available on gfx1250+.
Example:
// Async atomic barrier arrive (fire-and-forget).
rocdl.ds.atomic.async.barrier.arrive.b64 %ptr : !llvm.ptr<3>
Return op name rocdl.ds.atomic.barrier.arrive.rtn.b64 as a bitstring.
rocdl.ds.atomic.barrier.arrive.rtn.b64
Attributes
alias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
barrierPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3val- Single,I64, 64-bit signless integer
Results
res- Single,I64, 64-bit signless integer
Description
Waits on a given DS barrier and decrements its pending count by a given value. Note, the barrier state is given as a 64-bit structure containing pending count, phase and init count. The op returns the old barrier state. The op is executed as an ordinary LDS operations and it is ordered with other LDS operations. Thus, check DSCNT to determine when this instruction has executed. Available on gfx1250+.
Example:
// Atomic barrier arrive with return of old barrier state.
%res = rocdl.ds.atomic.barrier.arrive.rtn.b64 %ptr, %val : !llvm.ptr<3>, i64 -> i64
Return op name rocdl.ds_bpermute as a bitstring.
rocdl.ds_bpermute
Operands
index- Single,I32, 32-bit signless integersrc- Single,I32, 32-bit signless integer
Results
res- Single,I32, 32-bit signless integer
Description
Perform a backward permute (pull) operation across lanes using DS/LDS permute hardware.
Each lane reads the value of src from the lane whose byte address is
given by index (i.e. lane id = index / 4).
This is “backward” (pull) in contrast to ds_permute_b32, which is
“forward” (push/scatter).
Example:
// Backward permute across lanes (pull from selected lane).
%0 = rocdl.ds_bpermute %index, %src : (i32, i32) -> i32
Return op name rocdl.ds.load.tr4.b64 as a bitstring.
rocdl.ds.load.tr4.b64 - Loads and transposes a matrix from ds memory to registers (available in gfx1250+).
Attributes
alias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
ptr- Single, anonymous/composite constraint, LLVM pointer in address space 3
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Load a matrix of 4-bit data from the ds memory, transpose data between row-major and column-major order, and store the result into a 64-bit vector register.
Available in gfx1250+.
Example (concrete mnemonics depend on address space and element size):
// 64-bit transpose load from global memory.
%0 = rocdl.global.load.tr4.b64 %ptr : !llvm.ptr<1> -> vector<2xi32>
// 128-bit transpose load from global memory with f16 result.
%1 = rocdl.global.load.tr.b128 %ptr : !llvm.ptr<1> -> vector<8xf16>
// 64-bit transpose load from LDS.
%2 = rocdl.ds.load.tr4.b64 %ptr : !llvm.ptr<3> -> vector<2xi32>
// 128-bit transpose load from LDS with bf16 result.
%3 = rocdl.ds.load.tr16.b128 %ptr : !llvm.ptr<3> -> vector<8xbf16>
Return op name rocdl.ds.load.tr6.b96 as a bitstring.
rocdl.ds.load.tr6.b96 - Loads and transposes a matrix from ds memory to registers (available in gfx1250+).
Attributes
alias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
ptr- Single, anonymous/composite constraint, LLVM pointer in address space 3
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Load a matrix of 6-bit data from the ds memory, transpose data between row-major and column-major order, and store the result into a 96-bit vector register.
Available in gfx1250+.
Example (concrete mnemonics depend on address space and element size):
// 64-bit transpose load from global memory.
%0 = rocdl.global.load.tr4.b64 %ptr : !llvm.ptr<1> -> vector<2xi32>
// 128-bit transpose load from global memory with f16 result.
%1 = rocdl.global.load.tr.b128 %ptr : !llvm.ptr<1> -> vector<8xf16>
// 64-bit transpose load from LDS.
%2 = rocdl.ds.load.tr4.b64 %ptr : !llvm.ptr<3> -> vector<2xi32>
// 128-bit transpose load from LDS with bf16 result.
%3 = rocdl.ds.load.tr16.b128 %ptr : !llvm.ptr<3> -> vector<8xbf16>
Return op name rocdl.ds.load.tr8.b64 as a bitstring.
rocdl.ds.load.tr8.b64 - Loads and transposes a matrix from ds memory to registers (available in gfx1250+).
Attributes
alias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
ptr- Single, anonymous/composite constraint, LLVM pointer in address space 3
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Load a matrix of 8-bit data from the ds memory, transpose data between row-major and column-major order, and store the result into a 64-bit vector register.
Available in gfx1250+.
Example (concrete mnemonics depend on address space and element size):
// 64-bit transpose load from global memory.
%0 = rocdl.global.load.tr4.b64 %ptr : !llvm.ptr<1> -> vector<2xi32>
// 128-bit transpose load from global memory with f16 result.
%1 = rocdl.global.load.tr.b128 %ptr : !llvm.ptr<1> -> vector<8xf16>
// 64-bit transpose load from LDS.
%2 = rocdl.ds.load.tr4.b64 %ptr : !llvm.ptr<3> -> vector<2xi32>
// 128-bit transpose load from LDS with bf16 result.
%3 = rocdl.ds.load.tr16.b128 %ptr : !llvm.ptr<3> -> vector<8xbf16>
Return op name rocdl.ds.load.tr16.b128 as a bitstring.
rocdl.ds.load.tr16.b128 - Loads and transposes a matrix from ds memory to registers (available in gfx1250+).
Attributes
alias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
ptr- Single, anonymous/composite constraint, LLVM pointer in address space 3
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Load a matrix of 16-bit data from the ds memory, transpose data between row-major and column-major order, and store the result into a 128-bit vector register.
Available in gfx1250+.
Example (concrete mnemonics depend on address space and element size):
// 64-bit transpose load from global memory.
%0 = rocdl.global.load.tr4.b64 %ptr : !llvm.ptr<1> -> vector<2xi32>
// 128-bit transpose load from global memory with f16 result.
%1 = rocdl.global.load.tr.b128 %ptr : !llvm.ptr<1> -> vector<8xf16>
// 64-bit transpose load from LDS.
%2 = rocdl.ds.load.tr4.b64 %ptr : !llvm.ptr<3> -> vector<2xi32>
// 128-bit transpose load from LDS with bf16 result.
%3 = rocdl.ds.load.tr16.b128 %ptr : !llvm.ptr<3> -> vector<8xbf16>
Return op name rocdl.ds.read.tr4.b64 as a bitstring.
rocdl.ds.read.tr4.b64
Attributes
alias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
ptr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.ds.read.tr6.b96 as a bitstring.
rocdl.ds.read.tr6.b96
Attributes
alias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
ptr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.ds.read.tr8.b64 as a bitstring.
rocdl.ds.read.tr8.b64
Attributes
alias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
ptr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.ds.read.tr16.b64 as a bitstring.
rocdl.ds.read.tr16.b64
Attributes
alias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
ptr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.ds_swizzle as a bitstring.
rocdl.ds_swizzle
Operands
src- Single,I32, 32-bit signless integeroffset- Single,I32, 32-bit signless integer
Results
res- Single,I32, 32-bit signless integer
Description
Perform a data-sharing swizzle operation within a wavefront.
The offset operand encodes the swizzle pattern that will be placed in the
instruction's offset field (i.e., the pattern used by ds_swizzle_b32).
See https://llvm.org/docs/AMDGPUModifierSyntax.html#swizzle-pattern for
how this 16-bit pattern is constructed.
Example:
// Swizzle data within a wavefront.
%0 = rocdl.ds_swizzle %src, %offset : (i32, i32) -> i32
Return op name rocdl.exp2 as a bitstring.
rocdl.exp2
Operands
arg- Single,LLVM_AnyFloat, floating point LLVM type
Results
res- Single,LLVM_AnyFloat, floating point LLVM type
Description
Note: In the general case, prefer the conventional arith, math, or llvm ops over this.
Use this ROCDL-specific operation only when you fully understand its implication and
when it is strictly necessary. This op is usually chosen when a small loss in precision is
acceptable in exchange for higher execution speed.
Example:
%0 = rocdl.exp2 %a f32 -> f32
Return op name rocdl.exp as a bitstring.
rocdl.exp
Operands
arg- Single,LLVM_AnyFloat, floating point LLVM type
Results
res- Single,LLVM_AnyFloat, floating point LLVM type
Description
Note: In the general case, prefer the conventional arith, math, or llvm ops over this.
Use this ROCDL-specific operation only when you fully understand its implication and
when it is strictly necessary. This op is usually chosen when a small loss in precision is
acceptable in exchange for higher execution speed.
Example:
%0 = rocdl.exp %a f32 -> f32
Return op name rocdl.fdot2 as a bitstring.
rocdl.fdot2
Attributes
clamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single,ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2b- Single,ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2c- Single, anonymous/composite constraint, 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float
Description
Packed intra-lane dot-product with optional result clamping (clamp).
Computes res = sum_i a[i]*b[i] + c, where a and b hold packed
4/8/16-bit data (for dot2,dot4,dot8).
Example:
%r = rocdl.fdot2 %a, %b, %c {clamp = true} :
(vector<2xf16>, vector<2xf16>, f32) -> f32
Return op name rocdl.fdot2.bf16.bf16 as a bitstring.
rocdl.fdot2.bf16.bf16
Operands
a- Single,ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2b- Single,ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2c- Single, anonymous/composite constraint, bfloat16 type
Results
res- Single, anonymous/composite constraint, bfloat16 type
Description
Packed intra-lane dot-product with no clamp control.
Computes res = sum_i a[i]*b[i] + c. Covers the full-f16/bf16
accumulator forms (fdot2.f16.f16, fdot2.bf16.bf16) and the
FP8/BF8 dot4.f32.* variants, whose hardware instructions have no
CLAMP bit in their modifier word.
Example:
%r = rocdl.fdot2.bf16.bf16 %a, %b, %c : (vector<2xbf16>, vector<2xbf16>, bf16) -> bf16
Return op name rocdl.fdot2.f16.f16 as a bitstring.
rocdl.fdot2.f16.f16
Operands
a- Single,ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2b- Single,ROCDL_V2F16Type, fixed-length vector of 16-bit float values of length 2c- Single, anonymous/composite constraint, 16-bit float
Results
res- Single, anonymous/composite constraint, 16-bit float
Description
Packed intra-lane dot-product with no clamp control.
Computes res = sum_i a[i]*b[i] + c. Covers the full-f16/bf16
accumulator forms (fdot2.f16.f16, fdot2.bf16.bf16) and the
FP8/BF8 dot4.f32.* variants, whose hardware instructions have no
CLAMP bit in their modifier word.
Example:
%r = rocdl.fdot2.f16.f16 %a, %b, %c : (vector<2xf16>, vector<2xf16>, f16) -> f16
Return op name rocdl.fdot2.f32.bf16 as a bitstring.
rocdl.fdot2.f32.bf16
Attributes
clamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single,ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2b- Single,ROCDL_V2BF16Type, fixed-length vector of bfloat16 type values of length 2c- Single, anonymous/composite constraint, 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float
Description
Packed intra-lane dot-product with optional result clamping (clamp).
Computes res = sum_i a[i]*b[i] + c, where a and b hold packed
4/8/16-bit data (for dot2,dot4,dot8).
Example:
%r = rocdl.fdot2.f32.bf16 %a, %b, %c {clamp = true} :
(vector<2xbf16>, vector<2xbf16>, f32) -> f32
Return op name rocdl.flat.prefetch as a bitstring.
rocdl.flat.prefetch
Attributes
cachePolicy- Single,ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
ptr- Single, anonymous/composite constraint, LLVM pointer in address space 0
Description
Prefetches 1 byte of data per lane using flat-memory addresses into the WGP-cache or L2-cache. Available on gfx1250+.
Example:
// Prefetch from flat memory into cache.
rocdl.flat.prefetch %ptr, 0 : !llvm.ptr
Return op name rocdl.fmed3 as a bitstring.
rocdl.fmed3 - Median of three float/half values
Operands
src0- Single, anonymous/composite constraint, floating point LLVM type or LLVM dialect-compatible vector of floating point LLVM typesrc1- Single, anonymous/composite constraint, floating point LLVM type or LLVM dialect-compatible vector of floating point LLVM typesrc2- Single, anonymous/composite constraint, floating point LLVM type or LLVM dialect-compatible vector of floating point LLVM type
Results
res- Single, anonymous/composite constraint, floating point LLVM type or LLVM dialect-compatible vector of floating point LLVM type
Description
Computes the median of three floating-point values using the AMDGPU fmed3 intrinsic.
This operation is equivalent to max(min(a, b), min(max(a, b), c)) but uses the
hardware-accelerated V_MED3_F16/V_MED3_F32 instruction for better performance.
The operation supports both scalar and vector floating-point types (f16, f32).
Example:
// Scalar f32 median
%result = rocdl.fmed3 %a, %b, %c : f32
// Vector f16 median
%result = rocdl.fmed3 %va, %vb, %vc : vector<4xf16>
Return op name rocdl.global.load.async.lds as a bitstring.
rocdl.global.load.async.lds - Version of rocdl.load.async.to.lds specialized to global pointers
Attributes
size- Single,I32Attr, 32-bit signless integer attributeoffset- Single,I32Attr, 32-bit signless integer attributeaux- Single,ROCDL_PreGfx12CachePolicyCompatAttr, pre-gfx12 or gfx942 AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
globalPtr- Single,ROCDLGlobalBuffer, LLVM pointer in address space 1ldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Description
This operation works identically to rocdl.load.async.to.lds except that the
global pointer argument is limited to pointers in address space 1 (pure global
pointers) instead of also allowing fat buffer pointers.
Available on gfx9 and gfx10.
For the operation introduced in gfx1250, see rocdl.global.load.async.to.lds.bN.
Example:
// Async load from global pointer to LDS (address space 1 only).
rocdl.load.async.to.lds %global, %shared, 4, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>
Return op name rocdl.global.load.async.to.lds.b8 as a bitstring.
rocdl.global.load.async.to.lds.b8
Attributes
offset- Single,I32Attr, 32-bit signless integer attributeaux- Single,ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
globalPtr- Single,ROCDLGlobalBuffer, LLVM pointer in address space 1ldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Description
Asynchronously loads 8 bits of data from a global memory pointer to a Local Data Share (LDS) pointer.
Available on gfx1250+.
Example:
// Async 8-bit load from global to LDS.
rocdl.global.load.async.to.lds.b8 %src, %dst, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>
Return op name rocdl.global.load.async.to.lds.b32 as a bitstring.
rocdl.global.load.async.to.lds.b32
Attributes
offset- Single,I32Attr, 32-bit signless integer attributeaux- Single,ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
globalPtr- Single,ROCDLGlobalBuffer, LLVM pointer in address space 1ldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Description
Asynchronously loads 32 bits of data from a global memory pointer to a Local Data Share (LDS) pointer.
Available on gfx1250+.
Example:
// Async 32-bit load from global to LDS.
rocdl.global.load.async.to.lds.b32 %src, %dst, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>
Return op name rocdl.global.load.async.to.lds.b64 as a bitstring.
rocdl.global.load.async.to.lds.b64
Attributes
offset- Single,I32Attr, 32-bit signless integer attributeaux- Single,ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
globalPtr- Single,ROCDLGlobalBuffer, LLVM pointer in address space 1ldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Description
Asynchronously loads 64 bits of data from a global memory pointer to a Local Data Share (LDS) pointer.
Available on gfx1250+.
Example:
// Async 64-bit load from global to LDS.
rocdl.global.load.async.to.lds.b64 %src, %dst, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>
Return op name rocdl.global.load.async.to.lds.b128 as a bitstring.
rocdl.global.load.async.to.lds.b128
Attributes
offset- Single,I32Attr, 32-bit signless integer attributeaux- Single,ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
globalPtr- Single,ROCDLGlobalBuffer, LLVM pointer in address space 1ldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Description
Asynchronously loads 128 bits of data from a global memory pointer to a Local Data Share (LDS) pointer.
Available on gfx1250+.
Example:
// Async 128-bit load from global to LDS.
rocdl.global.load.async.to.lds.b128 %src, %dst, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>
Return op name rocdl.global.load.lds as a bitstring.
rocdl.global.load.lds
Attributes
size- Single,I32Attr, 32-bit signless integer attributeoffset- Single,I32Attr, 32-bit signless integer attributeaux- Single,ROCDL_PreGfx12CachePolicyCompatAttr, pre-gfx12 or gfx942 AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
globalPtr- Single,ROCDLGlobalBuffer, LLVM pointer in address space 1ldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Return op name rocdl.global.load.tr4.b64 as a bitstring.
rocdl.global.load.tr4.b64 - Loads and transposes a matrix from global memory to registers (available in gfx1250+).
Attributes
alias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
ptr- Single, anonymous/composite constraint, LLVM pointer in address space 1
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Load a matrix of 4-bit data from the global memory, transpose data between row-major and column-major order, and store the result into a 64-bit vector register.
Available in gfx1250+.
Example (concrete mnemonics depend on address space and element size):
// 64-bit transpose load from global memory.
%0 = rocdl.global.load.tr4.b64 %ptr : !llvm.ptr<1> -> vector<2xi32>
// 128-bit transpose load from global memory with f16 result.
%1 = rocdl.global.load.tr.b128 %ptr : !llvm.ptr<1> -> vector<8xf16>
// 64-bit transpose load from LDS.
%2 = rocdl.ds.load.tr4.b64 %ptr : !llvm.ptr<3> -> vector<2xi32>
// 128-bit transpose load from LDS with bf16 result.
%3 = rocdl.ds.load.tr16.b128 %ptr : !llvm.ptr<3> -> vector<8xbf16>
Return op name rocdl.global.load.tr6.b96 as a bitstring.
rocdl.global.load.tr6.b96 - Loads and transposes a matrix from global memory to registers (available in gfx1250+).
Attributes
alias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
ptr- Single, anonymous/composite constraint, LLVM pointer in address space 1
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Load a matrix of 6-bit data from the global memory, transpose data between row-major and column-major order, and store the result into a 96-bit vector register.
Available in gfx1250+.
Example (concrete mnemonics depend on address space and element size):
// 64-bit transpose load from global memory.
%0 = rocdl.global.load.tr4.b64 %ptr : !llvm.ptr<1> -> vector<2xi32>
// 128-bit transpose load from global memory with f16 result.
%1 = rocdl.global.load.tr.b128 %ptr : !llvm.ptr<1> -> vector<8xf16>
// 64-bit transpose load from LDS.
%2 = rocdl.ds.load.tr4.b64 %ptr : !llvm.ptr<3> -> vector<2xi32>
// 128-bit transpose load from LDS with bf16 result.
%3 = rocdl.ds.load.tr16.b128 %ptr : !llvm.ptr<3> -> vector<8xbf16>
Return op name rocdl.global.load.tr.b64 as a bitstring.
rocdl.global.load.tr.b64 - Loads and transposes a matrix from global memory to registers (available in gfx1250+).
Attributes
alias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
ptr- Single, anonymous/composite constraint, LLVM pointer in address space 1
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Load a matrix of 8-bit data from the global memory, transpose data between row-major and column-major order, and store the result into a 64-bit vector register.
Available in gfx1250+.
Example (concrete mnemonics depend on address space and element size):
// 64-bit transpose load from global memory.
%0 = rocdl.global.load.tr4.b64 %ptr : !llvm.ptr<1> -> vector<2xi32>
// 128-bit transpose load from global memory with f16 result.
%1 = rocdl.global.load.tr.b128 %ptr : !llvm.ptr<1> -> vector<8xf16>
// 64-bit transpose load from LDS.
%2 = rocdl.ds.load.tr4.b64 %ptr : !llvm.ptr<3> -> vector<2xi32>
// 128-bit transpose load from LDS with bf16 result.
%3 = rocdl.ds.load.tr16.b128 %ptr : !llvm.ptr<3> -> vector<8xbf16>
Return op name rocdl.global.load.tr.b128 as a bitstring.
rocdl.global.load.tr.b128 - Loads and transposes a matrix from global memory to registers (available in gfx1250+).
Attributes
alias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
ptr- Single, anonymous/composite constraint, LLVM pointer in address space 1
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Load a matrix of 16-bit data from the global memory, transpose data between row-major and column-major order, and store the result into a 128-bit vector register.
Available in gfx1250+.
Example (concrete mnemonics depend on address space and element size):
// 64-bit transpose load from global memory.
%0 = rocdl.global.load.tr4.b64 %ptr : !llvm.ptr<1> -> vector<2xi32>
// 128-bit transpose load from global memory with f16 result.
%1 = rocdl.global.load.tr.b128 %ptr : !llvm.ptr<1> -> vector<8xf16>
// 64-bit transpose load from LDS.
%2 = rocdl.ds.load.tr4.b64 %ptr : !llvm.ptr<3> -> vector<2xi32>
// 128-bit transpose load from LDS with bf16 result.
%3 = rocdl.ds.load.tr16.b128 %ptr : !llvm.ptr<3> -> vector<8xbf16>
Return op name rocdl.global.prefetch as a bitstring.
rocdl.global.prefetch
Attributes
cachePolicy- Single,ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
ptr- Single, anonymous/composite constraint, LLVM pointer in address space 1
Description
Prefetches 1 byte of data per lane from global memory into the WGP-cache or L2-cache. Available on gfx1250+.
Example:
// Prefetch from global memory into cache.
rocdl.global.prefetch %ptr, 0 : !llvm.ptr<1>
Return op name rocdl.global.store.async.from.lds.b8 as a bitstring.
rocdl.global.store.async.from.lds.b8
Attributes
offset- Single,I32Attr, 32-bit signless integer attributeaux- Single,ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
globalPtr- Single,ROCDLGlobalBuffer, LLVM pointer in address space 1ldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Description
Asynchronously stores 8 bits of data from a Local Data Share (LDS) pointer to a global memory pointer.
Available on gfx1250+.
Example:
// Async 8-bit store from LDS to global.
rocdl.global.store.async.from.lds.b8 %dst, %src, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>
Return op name rocdl.global.store.async.from.lds.b32 as a bitstring.
rocdl.global.store.async.from.lds.b32
Attributes
offset- Single,I32Attr, 32-bit signless integer attributeaux- Single,ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
globalPtr- Single,ROCDLGlobalBuffer, LLVM pointer in address space 1ldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Description
Asynchronously stores 32 bits of data from a Local Data Share (LDS) pointer to a global memory pointer.
Available on gfx1250+.
Example:
// Async 32-bit store from LDS to global.
rocdl.global.store.async.from.lds.b32 %dst, %src, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>
Return op name rocdl.global.store.async.from.lds.b64 as a bitstring.
rocdl.global.store.async.from.lds.b64
Attributes
offset- Single,I32Attr, 32-bit signless integer attributeaux- Single,ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
globalPtr- Single,ROCDLGlobalBuffer, LLVM pointer in address space 1ldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Description
Asynchronously stores 64 bits of data from a Local Data Share (LDS) pointer to a global memory pointer.
Available on gfx1250+.
Example:
// Async 64-bit store from LDS to global.
rocdl.global.store.async.from.lds.b64 %dst, %src, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>
Return op name rocdl.global.store.async.from.lds.b128 as a bitstring.
rocdl.global.store.async.from.lds.b128
Attributes
offset- Single,I32Attr, 32-bit signless integer attributeaux- Single,ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
globalPtr- Single,ROCDLGlobalBuffer, LLVM pointer in address space 1ldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Description
Asynchronously stores 128 bits of data from a Local Data Share (LDS) pointer to a global memory pointer.
Available on gfx1250+.
Example:
// Async 128-bit store from LDS to global.
rocdl.global.store.async.from.lds.b128 %dst, %src, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>
Return op name rocdl.iglp.opt as a bitstring.
rocdl.iglp.opt
Attributes
variant- Single,I32Attr, 32-bit signless integer attribute
Description
Instruction-group-level parallelism optimization hint.
Example:
// IGLP optimization hint variant 0.
rocdl.iglp.opt 0
Return op name rocdl.load.async.to.lds as a bitstring.
rocdl.load.async.to.lds - Gathering load to LDS that requires explicit async memory tracking
Attributes
size- Single,I32Attr, 32-bit signless integer attributeoffset- Single,I32Attr, 32-bit signless integer attributeaux- Single,ROCDL_PreGfx12CachePolicyCompatAttr, pre-gfx12 or gfx942 AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
globalPtr- Single,LLVM_AnyPointer, LLVM pointer typeldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Description
Load size bytes (the valid sizes vary by architecture) from the global memory
pointed to by globalPtr and put them at ldsPtr, concantenating (and applying
padding for sizes less than 4 bytes, along with padding out 12-byte reads
to 16-byte writes). The value of globalPtr can vary between lanes, while
sharedPtr must be subgroup-uniform (the values from each lane are concatentated
before being written to LDS with appropriate padding applied.)
offset is a constant offset applied to both pointers, and aux sets the cache
policy. Unlike rocdl.load.to.lds, the compiler will not automatically inserts waits
for this load to complete at the point it thinks you're using a region of LDS you've
stored values to - you need to use the rocdl.asyncmark and rocdl.wait.asyncmark
operations to explicitly group these operations and wait for their completion.
Available on gfx10 and earlier with varying suppported values of size.
Example:
// Async load 4 bytes from global pointer to LDS.
rocdl.load.async.to.lds %global, %shared, 4, 0, 0 : !llvm.ptr<1>, !llvm.ptr<3>
// Async load 4 bytes from fat buffer pointer to LDS.
rocdl.load.async.to.lds %fatBuffer, %shared, 4, 0, 0 : !llvm.ptr<7>, !llvm.ptr<3>
Return op name rocdl.load.to.lds as a bitstring.
rocdl.load.to.lds
Attributes
size- Single,I32Attr, 32-bit signless integer attributeoffset- Single,I32Attr, 32-bit signless integer attributeaux- Single,ROCDL_PreGfx12CachePolicyCompatAttr, pre-gfx12 or gfx942 AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
globalPtr- Single,LLVM_AnyPointer, LLVM pointer typeldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Return op name rocdl.log as a bitstring.
rocdl.log
Operands
arg- Single,LLVM_AnyFloat, floating point LLVM type
Results
res- Single,LLVM_AnyFloat, floating point LLVM type
Description
Note: In the general case, prefer the conventional arith, math, or llvm ops over this.
Use this ROCDL-specific operation only when you fully understand its implication and
when it is strictly necessary. This op is usually chosen when a small loss in precision is
acceptable in exchange for higher execution speed.
Example:
%0 = rocdl.log %a f32 -> f32
Return op name rocdl.make.buffer.rsrc as a bitstring.
rocdl.make.buffer.rsrc
Operands
base- Single,LLVM_AnyPointer, LLVM pointer typestride- Single,I16, 16-bit signless integernumRecords- Single,I64, 64-bit signless integerflags- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_AnyPointer, LLVM pointer type
Return op name rocdl.mbcnt.hi as a bitstring.
rocdl.mbcnt.hi
Attributes
arg_attrs- Optional,DictArrayAttr, Array of dictionary attributesres_attrs- Optional,DictArrayAttr, Array of dictionary attributes
Operands
in0- Single,I32, 32-bit signless integerin1- Single,I32, 32-bit signless integer
Results
res- Single,I32, 32-bit signless integer
Description
Masked bit count of threads below the current lane in a wavefront.
in0 is a 32-bit mask that is AND-ed with the relevant half of the
execution mask and the bits below the current lane; in1 is added
to the resulting popcount:
- lo:
in1 + popcount(in0 & exec_lo & ((1 << min(lane_id, 32)) - 1)) - hi:
in1 + popcount(in0 & exec_hi & ((1 << saturating_usub(lane_id, 32)) - 1))
To obtain a unique thread index within a wave64, chain the two ops
with in0 = -1 (all bits set):
Example:
%all_ones = arith.constant -1 : i32
%zero = arith.constant 0 : i32
// Count active threads below this lane in the low 32 lanes.
%lo = rocdl.mbcnt.lo %all_ones, %zero : (i32, i32) -> i32
// Add the count from the high 32 lanes to get the full lane index.
%hi = rocdl.mbcnt.hi %all_ones, %lo : (i32, i32) -> i32
Return op name rocdl.mbcnt.lo as a bitstring.
rocdl.mbcnt.lo
Attributes
arg_attrs- Optional,DictArrayAttr, Array of dictionary attributesres_attrs- Optional,DictArrayAttr, Array of dictionary attributes
Operands
in0- Single,I32, 32-bit signless integerin1- Single,I32, 32-bit signless integer
Results
res- Single,I32, 32-bit signless integer
Description
Masked bit count of threads below the current lane in a wavefront.
in0 is a 32-bit mask that is AND-ed with the relevant half of the
execution mask and the bits below the current lane; in1 is added
to the resulting popcount:
- lo:
in1 + popcount(in0 & exec_lo & ((1 << min(lane_id, 32)) - 1)) - hi:
in1 + popcount(in0 & exec_hi & ((1 << saturating_usub(lane_id, 32)) - 1))
To obtain a unique thread index within a wave64, chain the two ops
with in0 = -1 (all bits set):
Example:
%all_ones = arith.constant -1 : i32
%zero = arith.constant 0 : i32
// Count active threads below this lane in the low 32 lanes.
%lo = rocdl.mbcnt.lo %all_ones, %zero : (i32, i32) -> i32
// Add the count from the high 32 lanes to get the full lane index.
%hi = rocdl.mbcnt.hi %all_ones, %lo : (i32, i32) -> i32
Return op name rocdl.mfma.f32.4x4x1f32 as a bitstring.
rocdl.mfma.f32.4x4x1f32
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 32-bit floatb- Single, anonymous/composite constraint, 32-bit floatc- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.4x4x1f32 %a0, %b0, %c0, 0, 0, none : (f32, f32, vector<4xf32>) -> vector<4xf32>
Return op name rocdl.mfma.f32.4x4x2bf16 as a bitstring.
rocdl.mfma.f32.4x4x2bf16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.4x4x2bf16 %a0, %b0, %c0, 0, 0, none : (vector<2xi16>, vector<2xi16>, vector<4xf32>) -> vector<4xf32>
Return op name rocdl.mfma.f32.4x4x4bf16.1k as a bitstring.
rocdl.mfma.f32.4x4x4bf16.1k
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.4x4x4bf16.1k %a0, %b0, %c0, 0, 0, none : (vector<4xi16>, vector<4xi16>, vector<4xf32>) -> vector<4xf32>
Return op name rocdl.mfma.f32.4x4x4f16 as a bitstring.
rocdl.mfma.f32.4x4x4f16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.4x4x4f16 %a0, %b0, %c0, 0, 0, none : (vector<4xf16>, vector<4xf16>, vector<4xf32>) -> vector<4xf32>
Return op name rocdl.mfma.f32.16x16x1f32 as a bitstring.
rocdl.mfma.f32.16x16x1f32
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 32-bit floatb- Single, anonymous/composite constraint, 32-bit floatc- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.16x16x1f32 %a0, %b0, %c0, 0, 0, none : (f32, f32, vector<16xf32>) -> vector<16xf32>
Return op name rocdl.mfma.f32.16x16x2bf16 as a bitstring.
rocdl.mfma.f32.16x16x2bf16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.16x16x2bf16 %a0, %b0, %c0, 0, 0, none : (vector<2xi16>, vector<2xi16>, vector<16xf32>) -> vector<16xf32>
Return op name rocdl.mfma.f32.16x16x4bf16.1k as a bitstring.
rocdl.mfma.f32.16x16x4bf16.1k
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.16x16x4bf16.1k %a0, %b0, %c0, 0, 0, none : (vector<4xi16>, vector<4xi16>, vector<16xf32>) -> vector<16xf32>
Return op name rocdl.mfma.f32.16x16x4f16 as a bitstring.
rocdl.mfma.f32.16x16x4f16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.16x16x4f16 %a0, %b0, %c0, 0, 0, none : (vector<4xf16>, vector<4xf16>, vector<16xf32>) -> vector<16xf32>
Return op name rocdl.mfma.f32.16x16x4f32 as a bitstring.
rocdl.mfma.f32.16x16x4f32
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 32-bit floatb- Single, anonymous/composite constraint, 32-bit floatc- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.16x16x4f32 %a0, %b0, %c0, 0, 0, none : (f32, f32, vector<4xf32>) -> vector<4xf32>
Return op name rocdl.mfma.f32.16x16x8.xf32 as a bitstring.
rocdl.mfma.f32.16x16x8.xf32
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 2c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.16x16x8.xf32 %a0, %b0, %c0, 0, 0, none : (vector<2xf32>, vector<2xf32>, vector<4xf32>) -> vector<4xf32>
Return op name rocdl.mfma.f32.16x16x8bf16 as a bitstring.
rocdl.mfma.f32.16x16x8bf16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.16x16x8bf16 %a0, %b0, %c0, 0, 0, none : (vector<2xi16>, vector<2xi16>, vector<4xf32>) -> vector<4xf32>
Return op name rocdl.mfma.f32.16x16x16bf16.1k as a bitstring.
rocdl.mfma.f32.16x16x16bf16.1k
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.16x16x16bf16.1k %a0, %b0, %c0, 0, 0, none : (vector<4xi16>, vector<4xi16>, vector<4xf32>) -> vector<4xf32>
Return op name rocdl.mfma.f32.16x16x16f16 as a bitstring.
rocdl.mfma.f32.16x16x16f16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.16x16x16f16 %a0, %b0, %c0, 0, 0, none : (vector<4xf16>, vector<4xf16>, vector<4xf32>) -> vector<4xf32>
Return op name rocdl.mfma.f32.16x16x32.bf8.bf8 as a bitstring.
rocdl.mfma.f32.16x16x32.bf8.bf8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 64-bit signless integerb- Single, anonymous/composite constraint, 64-bit signless integerc- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.16x16x32.bf8.bf8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<4xf32>) -> vector<4xf32>
Return op name rocdl.mfma.f32.16x16x32.bf8.fp8 as a bitstring.
rocdl.mfma.f32.16x16x32.bf8.fp8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 64-bit signless integerb- Single, anonymous/composite constraint, 64-bit signless integerc- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.16x16x32.bf8.fp8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<4xf32>) -> vector<4xf32>
Return op name rocdl.mfma.f32.16x16x32.bf16 as a bitstring.
rocdl.mfma.f32.16x16x32.bf16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of bfloat16 type values of length 8b- Single, anonymous/composite constraint, fixed-length vector of bfloat16 type values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.16x16x32.bf16 %a0, %b0, %c0, 0, 0, none : (vector<8xbf16>, vector<8xbf16>, vector<4xf32>) -> vector<4xf32>
Return op name rocdl.mfma.f32.16x16x32.f16 as a bitstring.
rocdl.mfma.f32.16x16x32.f16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 8b- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.16x16x32.f16 %a0, %b0, %c0, 0, 0, none : (vector<8xf16>, vector<8xf16>, vector<4xf32>) -> vector<4xf32>
Return op name rocdl.mfma.f32.16x16x32.fp8.bf8 as a bitstring.
rocdl.mfma.f32.16x16x32.fp8.bf8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 64-bit signless integerb- Single, anonymous/composite constraint, 64-bit signless integerc- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.16x16x32.fp8.bf8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<4xf32>) -> vector<4xf32>
Return op name rocdl.mfma.f32.16x16x32.fp8.fp8 as a bitstring.
rocdl.mfma.f32.16x16x32.fp8.fp8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 64-bit signless integerb- Single, anonymous/composite constraint, 64-bit signless integerc- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.16x16x32.fp8.fp8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<4xf32>) -> vector<4xf32>
Return op name rocdl.mfma.f32.32x32x1f32 as a bitstring.
rocdl.mfma.f32.32x32x1f32
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 32-bit floatb- Single, anonymous/composite constraint, 32-bit floatc- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 32
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 32
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.32x32x1f32 %a0, %b0, %c0, 0, 0, none : (f32, f32, vector<32xf32>) -> vector<32xf32>
Return op name rocdl.mfma.f32.32x32x2bf16 as a bitstring.
rocdl.mfma.f32.32x32x2bf16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 32
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 32
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.32x32x2bf16 %a0, %b0, %c0, 0, 0, none : (vector<2xi16>, vector<2xi16>, vector<32xf32>) -> vector<32xf32>
Return op name rocdl.mfma.f32.32x32x2f32 as a bitstring.
rocdl.mfma.f32.32x32x2f32
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 32-bit floatb- Single, anonymous/composite constraint, 32-bit floatc- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.32x32x2f32 %a0, %b0, %c0, 0, 0, none : (f32, f32, vector<16xf32>) -> vector<16xf32>
Return op name rocdl.mfma.f32.32x32x4.xf32 as a bitstring.
rocdl.mfma.f32.32x32x4.xf32
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 2c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.32x32x4.xf32 %a0, %b0, %c0, 0, 0, none : (vector<2xf32>, vector<2xf32>, vector<16xf32>) -> vector<16xf32>
Return op name rocdl.mfma.f32.32x32x4bf16 as a bitstring.
rocdl.mfma.f32.32x32x4bf16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 2c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.32x32x4bf16 %a0, %b0, %c0, 0, 0, none : (vector<2xi16>, vector<2xi16>, vector<16xf32>) -> vector<16xf32>
Return op name rocdl.mfma.f32.32x32x4bf16.1k as a bitstring.
rocdl.mfma.f32.32x32x4bf16.1k
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 32
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 32
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.32x32x4bf16.1k %a0, %b0, %c0, 0, 0, none : (vector<4xi16>, vector<4xi16>, vector<32xf32>) -> vector<32xf32>
Return op name rocdl.mfma.f32.32x32x4f16 as a bitstring.
rocdl.mfma.f32.32x32x4f16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 32
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 32
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.32x32x4f16 %a0, %b0, %c0, 0, 0, none : (vector<4xf16>, vector<4xf16>, vector<32xf32>) -> vector<32xf32>
Return op name rocdl.mfma.f32.32x32x8bf16.1k as a bitstring.
rocdl.mfma.f32.32x32x8bf16.1k
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.32x32x8bf16.1k %a0, %b0, %c0, 0, 0, none : (vector<4xi16>, vector<4xi16>, vector<16xf32>) -> vector<16xf32>
Return op name rocdl.mfma.f32.32x32x8f16 as a bitstring.
rocdl.mfma.f32.32x32x8f16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.32x32x8f16 %a0, %b0, %c0, 0, 0, none : (vector<4xf16>, vector<4xf16>, vector<16xf32>) -> vector<16xf32>
Return op name rocdl.mfma.f32.32x32x16.bf8.bf8 as a bitstring.
rocdl.mfma.f32.32x32x16.bf8.bf8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 64-bit signless integerb- Single, anonymous/composite constraint, 64-bit signless integerc- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.32x32x16.bf8.bf8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<16xf32>) -> vector<16xf32>
Return op name rocdl.mfma.f32.32x32x16.bf8.fp8 as a bitstring.
rocdl.mfma.f32.32x32x16.bf8.fp8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 64-bit signless integerb- Single, anonymous/composite constraint, 64-bit signless integerc- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.32x32x16.bf8.fp8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<16xf32>) -> vector<16xf32>
Return op name rocdl.mfma.f32.32x32x16.bf16 as a bitstring.
rocdl.mfma.f32.32x32x16.bf16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of bfloat16 type values of length 8b- Single, anonymous/composite constraint, fixed-length vector of bfloat16 type values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.32x32x16.bf16 %a0, %b0, %c0, 0, 0, none : (vector<8xbf16>, vector<8xbf16>, vector<16xf32>) -> vector<16xf32>
Return op name rocdl.mfma.f32.32x32x16.f16 as a bitstring.
rocdl.mfma.f32.32x32x16.f16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 8b- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.32x32x16.f16 %a0, %b0, %c0, 0, 0, none : (vector<8xf16>, vector<8xf16>, vector<16xf32>) -> vector<16xf32>
Return op name rocdl.mfma.f32.32x32x16.fp8.bf8 as a bitstring.
rocdl.mfma.f32.32x32x16.fp8.bf8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 64-bit signless integerb- Single, anonymous/composite constraint, 64-bit signless integerc- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.32x32x16.fp8.bf8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<16xf32>) -> vector<16xf32>
Return op name rocdl.mfma.f32.32x32x16.fp8.fp8 as a bitstring.
rocdl.mfma.f32.32x32x16.fp8.fp8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 64-bit signless integerb- Single, anonymous/composite constraint, 64-bit signless integerc- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.f32.32x32x16.fp8.fp8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<16xf32>) -> vector<16xf32>
Return op name rocdl.mfma.f64.4x4x4f64 as a bitstring.
rocdl.mfma.f64.4x4x4f64
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMANegModifierAttr, negation modifier bitfield for gfx94x double-precision MFMA
Operands
a- Single, anonymous/composite constraint, 64-bit floatb- Single, anonymous/composite constraint, 64-bit floatc- Single, anonymous/composite constraint, 64-bit float
Results
res- Single, anonymous/composite constraint, 64-bit float
Description
Double-precision matrix fused multiply-add (MFMA) intrinsic. On gfx94x,
the blgp immarg is a NEG bitfield rather than a B-lane permutation.
Example:
%r0 = mfma.f64.4x4x4f64 %a0, %b0, %c0, 0, 0, neg_a|neg_b : (f64, f64, f64) -> f64
Return op name rocdl.mfma.f64.16x16x4f64 as a bitstring.
rocdl.mfma.f64.16x16x4f64
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMANegModifierAttr, negation modifier bitfield for gfx94x double-precision MFMA
Operands
a- Single, anonymous/composite constraint, 64-bit floatb- Single, anonymous/composite constraint, 64-bit floatc- Single, anonymous/composite constraint, fixed-length vector of 64-bit float values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 64-bit float values of length 4
Description
Double-precision matrix fused multiply-add (MFMA) intrinsic. On gfx94x,
the blgp immarg is a NEG bitfield rather than a B-lane permutation.
Example:
%r0 = mfma.f64.16x16x4f64 %a0, %b0, %c0, 0, 0, neg_a|neg_b : (f64, f64, vector<4xf64>) -> vector<4xf64>
Return op name rocdl.mfma.i32.4x4x4i8 as a bitstring.
rocdl.mfma.i32.4x4x4i8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 32-bit signless integerb- Single, anonymous/composite constraint, 32-bit signless integerc- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.i32.4x4x4i8 %a0, %b0, %c0, 0, 0, none : (i32, i32, vector<4xi32>) -> vector<4xi32>
Return op name rocdl.mfma.i32.16x16x4i8 as a bitstring.
rocdl.mfma.i32.16x16x4i8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 32-bit signless integerb- Single, anonymous/composite constraint, 32-bit signless integerc- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.i32.16x16x4i8 %a0, %b0, %c0, 0, 0, none : (i32, i32, vector<16xi32>) -> vector<16xi32>
Return op name rocdl.mfma.i32.16x16x16i8 as a bitstring.
rocdl.mfma.i32.16x16x16i8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 32-bit signless integerb- Single, anonymous/composite constraint, 32-bit signless integerc- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.i32.16x16x16i8 %a0, %b0, %c0, 0, 0, none : (i32, i32, vector<4xi32>) -> vector<4xi32>
Return op name rocdl.mfma.i32.16x16x32.i8 as a bitstring.
rocdl.mfma.i32.16x16x32.i8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 64-bit signless integerb- Single, anonymous/composite constraint, 64-bit signless integerc- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.i32.16x16x32.i8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<4xi32>) -> vector<4xi32>
Return op name rocdl.mfma.i32.16x16x64.i8 as a bitstring.
rocdl.mfma.i32.16x16x64.i8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.i32.16x16x64.i8 %a0, %b0, %c0, 0, 0, none : (vector<4xi32>, vector<4xi32>, vector<4xi32>) -> vector<4xi32>
Return op name rocdl.mfma.i32.32x32x4i8 as a bitstring.
rocdl.mfma.i32.32x32x4i8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 32-bit signless integerb- Single, anonymous/composite constraint, 32-bit signless integerc- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 32
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 32
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.i32.32x32x4i8 %a0, %b0, %c0, 0, 0, none : (i32, i32, vector<32xi32>) -> vector<32xi32>
Return op name rocdl.mfma.i32.32x32x8i8 as a bitstring.
rocdl.mfma.i32.32x32x8i8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 32-bit signless integerb- Single, anonymous/composite constraint, 32-bit signless integerc- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.i32.32x32x8i8 %a0, %b0, %c0, 0, 0, none : (i32, i32, vector<16xi32>) -> vector<16xi32>
Return op name rocdl.mfma.i32.32x32x16.i8 as a bitstring.
rocdl.mfma.i32.32x32x16.i8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, 64-bit signless integerb- Single, anonymous/composite constraint, 64-bit signless integerc- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.i32.32x32x16.i8 %a0, %b0, %c0, 0, 0, none : (i64, i64, vector<16xi32>) -> vector<16xi32>
Return op name rocdl.mfma.i32.32x32x32.i8 as a bitstring.
rocdl.mfma.i32.32x32x32.i8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attributeblgp- Single,ROCDL_MFMAPermBAttr, permutations of the lanes storing B in an MFMA
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16
Description
Matrix fused multiply-add (MFMA) intrinsic. Computes D = A * B + C
with matrix operands. The cbsz, abid, and blgp attributes control
broadcast and block layout modes.
Example:
%r0 = mfma.i32.32x32x32.i8 %a0, %b0, %c0, 0, 0, none : (vector<4xi32>, vector<4xi32>, vector<16xi32>) -> vector<16xi32>
Return op name rocdl.mfma.scale.f32.16x16x128.f8f6f4 as a bitstring.
rocdl.mfma.scale.f32.16x16x128.f8f6f4
Attributes
cbsz- Single,ROCDL_MatrixFormatAttr, matrix operand formats selected by scaled MFMA/WMMA format fieldsblgp- Single,ROCDL_MatrixFormatAttr, matrix operand formats selected by scaled MFMA/WMMA format fieldsopselA- Single,I32Attr, 32-bit signless integer attributeopselB- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit signless integerb- Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit signless integerc- Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit floatscaleA- Single,I32, 32-bit signless integerscaleB- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Scaled matrix fused multiply-add (MFMA) intrinsic with per-operand scaling.
The opselA/opselB and scaleA/scaleB arguments control the scaling
of input operands.
Example:
// Scaled MFMA with fp8 * fp8 inputs.
%r0 = rocdl.mfma.scale.f32.32x32x64.f8f6f4 %a, %a, %c, fp8_e4m3, fp8_e4m3, 0, %scaleA, 0, %scaleB :
(vector<8xi32>, vector<8xi32>, vector<16xf32>, i32, i32) -> vector<16xf32>
// Scaled MFMA with fp8 * bf8 inputs.
%r1 = rocdl.mfma.scale.f32.32x32x64.f8f6f4 %a, %a, %c, fp8_e4m3, fp8_e5m2, 0, %scaleA, 0, %scaleB :
(vector<8xi32>, vector<8xi32>, vector<16xf32>, i32, i32) -> vector<16xf32>
// Scaled MFMA with fp8 * fp6 inputs (6xi32 operand B).
%r2 = rocdl.mfma.scale.f32.32x32x64.f8f6f4 %a, %b6, %c, fp8_e4m3, fp6_e2m3, 0, %scaleA, 0, %scaleB :
(vector<8xi32>, vector<6xi32>, vector<16xf32>, i32, i32) -> vector<16xf32>
Return op name rocdl.mfma.scale.f32.32x32x64.f8f6f4 as a bitstring.
rocdl.mfma.scale.f32.32x32x64.f8f6f4
Attributes
cbsz- Single,ROCDL_MatrixFormatAttr, matrix operand formats selected by scaled MFMA/WMMA format fieldsblgp- Single,ROCDL_MatrixFormatAttr, matrix operand formats selected by scaled MFMA/WMMA format fieldsopselA- Single,I32Attr, 32-bit signless integer attributeopselB- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit signless integerb- Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit signless integerc- Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit floatscaleA- Single,I32, 32-bit signless integerscaleB- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Scaled matrix fused multiply-add (MFMA) intrinsic with per-operand scaling.
The opselA/opselB and scaleA/scaleB arguments control the scaling
of input operands.
Example:
// Scaled MFMA with fp8 * fp8 inputs.
%r0 = rocdl.mfma.scale.f32.32x32x64.f8f6f4 %a, %a, %c, fp8_e4m3, fp8_e4m3, 0, %scaleA, 0, %scaleB :
(vector<8xi32>, vector<8xi32>, vector<16xf32>, i32, i32) -> vector<16xf32>
// Scaled MFMA with fp8 * bf8 inputs.
%r1 = rocdl.mfma.scale.f32.32x32x64.f8f6f4 %a, %a, %c, fp8_e4m3, fp8_e5m2, 0, %scaleA, 0, %scaleB :
(vector<8xi32>, vector<8xi32>, vector<16xf32>, i32, i32) -> vector<16xf32>
// Scaled MFMA with fp8 * fp6 inputs (6xi32 operand B).
%r2 = rocdl.mfma.scale.f32.32x32x64.f8f6f4 %a, %b6, %c, fp8_e4m3, fp6_e2m3, 0, %scaleA, 0, %scaleB :
(vector<8xi32>, vector<6xi32>, vector<16xf32>, i32, i32) -> vector<16xf32>
Return op name rocdl.permlane16.swap as a bitstring.
rocdl.permlane16.swap
Attributes
fi- Single,I1Attr, 1-bit signless integer attributeboundControl- Single,I1Attr, 1-bit signless integer attribute
Operands
old- Single,I32, 32-bit signless integersrc- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, LLVM dialect-compatible struct of 32-bit signless integerand32-bit signless integer
Description
Performs a permlane16.swap operation with the given operands, applying the
permutation specified by $fi to the provided inputs.
Example:
// Swap lanes between groups of 16 threads.
%res = rocdl.permlane16.swap %src, %src, 0, -1 : (i32, i32) -> !llvm.struct<(i32, i32)>
Return op name rocdl.permlane16.var as a bitstring.
rocdl.permlane16.var
Attributes
fi- Single,I1Attr, 1-bit signless integer attributeboundControl- Single,I1Attr, 1-bit signless integer attribute
Operands
old- Single,I32, 32-bit signless integersrc0- Single,I32, 32-bit signless integersrc1- Single,I32, 32-bit signless integer
Results
res- Single,I32, 32-bit signless integer
Description
Performs a permlane16.var operation: a per-lane variable-selector
intra-row permutation (within each 16-lane row). Maps to
llvm.amdgcn.permlane16.var.
Each destination lane within a 16-lane row reads its value from the
source lane whose index is given by the corresponding per-lane entry
in $src1 (a VGPR). $fi and $boundControl are immediate i1 attrs
matching the underlying intrinsic's modifiers.
Example:
%res = rocdl.permlane16.var %old, %src, %selector, false, true : (i32, i32, i32) -> i32
Return op name rocdl.permlane32.swap as a bitstring.
rocdl.permlane32.swap
Attributes
fi- Single,I1Attr, 1-bit signless integer attributeboundControl- Single,I1Attr, 1-bit signless integer attribute
Operands
old- Single,I32, 32-bit signless integersrc- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, LLVM dialect-compatible struct of 32-bit signless integerand32-bit signless integer
Description
Performs a permlane32.swap operation with the given operands, applying the
permutation specified by $fi to the provided inputs.
Example:
// Swap lanes between groups of 32 threads.
%res = rocdl.permlane32.swap %src, %src, 0, -1 : (i32, i32) -> !llvm.struct<(i32, i32)>
Return op name rocdl.permlanex16 as a bitstring.
rocdl.permlanex16
Attributes
fi- Single,I1Attr, 1-bit signless integer attributeboundControl- Single,I1Attr, 1-bit signless integer attribute
Operands
old- Single,LLVM_Type, LLVM dialect-compatible typesrc0- Single,LLVM_Type, LLVM dialect-compatible typesrc1- Single,LLVM_Type, LLVM dialect-compatible typesrc2- Single,LLVM_Type, LLVM dialect-compatible type
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Performs a permlanex16 operation with the given operands, applying the
permutation specified by $fi to the provided inputs.
Example:
// Scalar permlanex16.
%ret0 = rocdl.permlanex16 %src0, %src0, %sel, %sel, 0, -1 : f32, i32
// Vector permlanex16.
%ret1 = rocdl.permlanex16 %src1, %src1, %sel, %sel, 0, -1 : vector<2xf32>, i32
Return op name rocdl.permlanex16.var as a bitstring.
rocdl.permlanex16.var
Attributes
fi- Single,I1Attr, 1-bit signless integer attributeboundControl- Single,I1Attr, 1-bit signless integer attribute
Operands
old- Single,I32, 32-bit signless integersrc0- Single,I32, 32-bit signless integersrc1- Single,I32, 32-bit signless integer
Results
res- Single,I32, 32-bit signless integer
Description
Performs a permlanex16.var operation: a per-lane variable-selector
cross-row permutation (each lane in one 16-lane row reads from a
per-lane-chosen source lane in the opposite row). Maps to
llvm.amdgcn.permlanex16.var.
With per-lane "identity" selectors this realises an XOR-16 swap pattern
in pure VALU; with arbitrary selectors it enables general per-lane
cross-half-wave permutations. $fi and $boundControl are immediate
i1 attrs matching the underlying intrinsic's modifiers.
Example:
%res = rocdl.permlanex16.var %old, %src, %selector, false, true : (i32, i32, i32) -> i32
Return op name rocdl.ptr.s.buffer.load as a bitstring.
rocdl.ptr.s.buffer.load
Attributes
aux- Single,ROCDL_NonAtomicBufferCachePolicyCompatAttr, non-atomic AMDGPU buffer cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
rsrc- Single,ROCDLBufferRsrc, LLVM pointer in address space 8offset- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.raw.buffer.atomic.cmpswap as a bitstring.
rocdl.raw.buffer.atomic.cmpswap
Attributes
aux- Single,ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attribute
Operands
src- Single,LLVM_Type, LLVM dialect-compatible typecmp- Single,LLVM_Type, LLVM dialect-compatible typersrc- Single,ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4offset- Single,I32, 32-bit signless integersoffset- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.raw.buffer.atomic.fadd as a bitstring.
rocdl.raw.buffer.atomic.fadd
Attributes
aux- Single,ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attribute
Operands
vdata- Single,LLVM_Type, LLVM dialect-compatible typersrc- Single,ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4offset- Single,I32, 32-bit signless integersoffset- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.raw.buffer.atomic.fmax as a bitstring.
rocdl.raw.buffer.atomic.fmax
Attributes
aux- Single,ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attribute
Operands
vdata- Single,LLVM_Type, LLVM dialect-compatible typersrc- Single,ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4offset- Single,I32, 32-bit signless integersoffset- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.raw.buffer.atomic.smax as a bitstring.
rocdl.raw.buffer.atomic.smax
Attributes
aux- Single,ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attribute
Operands
vdata- Single,LLVM_Type, LLVM dialect-compatible typersrc- Single,ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4offset- Single,I32, 32-bit signless integersoffset- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.raw.buffer.atomic.umin as a bitstring.
rocdl.raw.buffer.atomic.umin
Attributes
aux- Single,ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attribute
Operands
vdata- Single,LLVM_Type, LLVM dialect-compatible typersrc- Single,ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4offset- Single,I32, 32-bit signless integersoffset- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.raw.buffer.load as a bitstring.
rocdl.raw.buffer.load
Attributes
aux- Single,ROCDL_NonAtomicBufferCachePolicyCompatAttr, non-atomic AMDGPU buffer cache policy attribute
Operands
rsrc- Single,ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4offset- Single,I32, 32-bit signless integersoffset- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.raw.buffer.store as a bitstring.
rocdl.raw.buffer.store
Attributes
aux- Single,ROCDL_NonAtomicBufferCachePolicyCompatAttr, non-atomic AMDGPU buffer cache policy attribute
Operands
vdata- Single,LLVM_Type, LLVM dialect-compatible typersrc- Single,ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4offset- Single,I32, 32-bit signless integersoffset- Single,I32, 32-bit signless integer
Return op name rocdl.raw.ptr.buffer.atomic.cmpswap as a bitstring.
rocdl.raw.ptr.buffer.atomic.cmpswap
Attributes
aux- Single,ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
src- Single,LLVM_Type, LLVM dialect-compatible typecmp- Single,LLVM_Type, LLVM dialect-compatible typersrc- Single,ROCDLBufferRsrc, LLVM pointer in address space 8offset- Single,I32, 32-bit signless integersoffset- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.raw.ptr.buffer.atomic.fadd as a bitstring.
rocdl.raw.ptr.buffer.atomic.fadd
Attributes
aux- Single,ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
vdata- Single,LLVM_Type, LLVM dialect-compatible typersrc- Single,ROCDLBufferRsrc, LLVM pointer in address space 8offset- Single,I32, 32-bit signless integersoffset- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.raw.ptr.buffer.atomic.fmax as a bitstring.
rocdl.raw.ptr.buffer.atomic.fmax
Attributes
aux- Single,ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
vdata- Single,LLVM_Type, LLVM dialect-compatible typersrc- Single,ROCDLBufferRsrc, LLVM pointer in address space 8offset- Single,I32, 32-bit signless integersoffset- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.raw.ptr.buffer.atomic.smax as a bitstring.
rocdl.raw.ptr.buffer.atomic.smax
Attributes
aux- Single,ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
vdata- Single,LLVM_Type, LLVM dialect-compatible typersrc- Single,ROCDLBufferRsrc, LLVM pointer in address space 8offset- Single,I32, 32-bit signless integersoffset- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.raw.ptr.buffer.atomic.umin as a bitstring.
rocdl.raw.ptr.buffer.atomic.umin
Attributes
aux- Single,ROCDL_AtomicBufferCachePolicyCompatAttr, atomic AMDGPU buffer cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
vdata- Single,LLVM_Type, LLVM dialect-compatible typersrc- Single,ROCDLBufferRsrc, LLVM pointer in address space 8offset- Single,I32, 32-bit signless integersoffset- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.raw.ptr.buffer.load as a bitstring.
rocdl.raw.ptr.buffer.load
Attributes
aux- Single,ROCDL_NonAtomicBufferCachePolicyCompatAttr, non-atomic AMDGPU buffer cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
rsrc- Single,ROCDLBufferRsrc, LLVM pointer in address space 8offset- Single,I32, 32-bit signless integersoffset- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.raw.ptr.buffer.load.async.lds as a bitstring.
rocdl.raw.ptr.buffer.load.async.lds - Async variant of raw.ptr.buffer.load.lds
Attributes
aux- Single,ROCDL_PreGfx12CachePolicyCompatAttr, pre-gfx12 or gfx942 AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
rsrc- Single,ROCDLBufferRsrc, LLVM pointer in address space 8ldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3size- Single,I32, 32-bit signless integervoffset- Single,I32, 32-bit signless integersoffset- Single,I32, 32-bit signless integeroffset- Single,I32, 32-bit signless integer
Description
Load from a buffer resource rsrc to ldsPtr, which must be uniform.
See rocdl.load.async.to.lds for overall semantics of such loads, noting that
here voffset can be lane-varying and that rsrc (which holds the base addres)
must, as always, be uniform.
Available on gfx9 and gfx10.
Example:
// Async buffer load to LDS via buffer resource pointer.
rocdl.raw.ptr.buffer.load.async.lds %rsrc, %ldsPtr, %size, %voffset, %soffset, %offset, 0
Return op name rocdl.raw.ptr.buffer.load.lds as a bitstring.
rocdl.raw.ptr.buffer.load.lds
Attributes
aux- Single,ROCDL_NonAtomicBufferCachePolicyCompatAttr, non-atomic AMDGPU buffer cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
rsrc- Single,ROCDLBufferRsrc, LLVM pointer in address space 8ldsPtr- Single,ROCDLBufferLDS, LLVM pointer in address space 3size- Single,I32, 32-bit signless integervoffset- Single,I32, 32-bit signless integersoffset- Single,I32, 32-bit signless integeroffset- Single,I32, 32-bit signless integer
Return op name rocdl.raw.ptr.buffer.store as a bitstring.
rocdl.raw.ptr.buffer.store
Attributes
aux- Single,ROCDL_NonAtomicBufferCachePolicyCompatAttr, non-atomic AMDGPU buffer cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
vdata- Single,LLVM_Type, LLVM dialect-compatible typersrc- Single,ROCDLBufferRsrc, LLVM pointer in address space 8offset- Single,I32, 32-bit signless integersoffset- Single,I32, 32-bit signless integer
Return op name rocdl.rcp as a bitstring.
rocdl.rcp
Operands
arg- Single,LLVM_AnyFloat, floating point LLVM type
Results
res- Single,LLVM_AnyFloat, floating point LLVM type
Description
Note: In the general case, prefer the conventional arith, math, or llvm ops over this.
Use this ROCDL-specific operation only when you fully understand its implication and
when it is strictly necessary. This op is usually chosen when a small loss in precision is
acceptable in exchange for higher execution speed.
Example:
%0 = rocdl.rcp %a f32 -> f32
Return op name rocdl.readfirstlane as a bitstring.
rocdl.readfirstlane - Get the value in first active lane.
Operands
src- Single,LLVM_Type, LLVM dialect-compatible type
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Returns the value in the lowest active lane of the input operand.
Example:
// Scalar readfirstlane.
%0 = rocdl.readfirstlane %src0 : f32
// Vector readfirstlane.
%1 = rocdl.readfirstlane %src1 : vector<2xf32>
Return op name rocdl.readlane as a bitstring.
rocdl.readlane - Get the value in the specific lane.
Operands
src0- Single,LLVM_Type, LLVM dialect-compatible typesrc1- Single,I32, 32-bit signless integer
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Get the value in lane src1 from input src0.
Example:
// Scalar readlane.
%0 = rocdl.readlane %src0, %idx : (f32, i32) -> f32
// Vector readlane.
%1 = rocdl.readlane %src1, %idx : (vector<2xf32>, i32) -> vector<2xf32>
Return op name rocdl.rsq as a bitstring.
rocdl.rsq
Operands
arg- Single,LLVM_AnyFloat, floating point LLVM type
Results
res- Single,LLVM_AnyFloat, floating point LLVM type
Description
Note: In the general case, prefer the conventional arith, math, or llvm ops over this.
Use this ROCDL-specific operation only when you fully understand its implication and
when it is strictly necessary. This op is usually chosen when a small loss in precision is
acceptable in exchange for higher execution speed.
Example:
%0 = rocdl.rsq %a f32 -> f32
Return op name rocdl.s.barrier as a bitstring.
rocdl.s.barrier
Description
Insert a workgroup barrier without memory fences.
Available on gfx9 and later but deprecated on gfx12+; see
rocdl.s.barrier.signal and rocdl.s.barrier.wait instead.
Example:
// Synchronize threads within a workgroup.
rocdl.s.barrier
Return op name rocdl.s.barrier.init as a bitstring.
rocdl.s.barrier.init
Attributes
memberCnt- Single,I32Attr, 32-bit signless integer attribute
Operands
ptr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Description
Available on gfx1250+.
Example:
// Initialize a named barrier with member count.
rocdl.s.barrier.init %ptr member_cnt = 1 : !llvm.ptr<3>
Return op name rocdl.s.barrier.join as a bitstring.
rocdl.s.barrier.join
Operands
ptr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Description
Available on gfx1250+.
Example:
// Join a named barrier.
rocdl.s.barrier.join %ptr : !llvm.ptr<3>
Return op name rocdl.s.barrier.leave as a bitstring.
rocdl.s.barrier.leave
Attributes
id- Single,I16Attr, 16-bit signless integer attribute
Description
Available on gfx1250+.
Example:
// Leave a named barrier by id.
rocdl.s.barrier.leave id = 1
Return op name rocdl.s.barrier.signal as a bitstring.
rocdl.s.barrier.signal
Attributes
id- Single,I32Attr, 32-bit signless integer attribute
Description
Signal a barrier by id. Available on gfx1250+.
Example:
// Signal barrier with id -1 (all barriers).
rocdl.s.barrier.signal id = -1
Return op name rocdl.s.barrier.signal.isfirst as a bitstring.
rocdl.s.barrier.signal.isfirst
Attributes
id- Single,I32Attr, 32-bit signless integer attribute
Results
res- Single,I1, 1-bit signless integer
Description
Available on gfx1200+.
Example:
// Signal barrier and check if this wave is first to arrive.
%0 = rocdl.s.barrier.signal.isfirst id = 1 -> i1
Return op name rocdl.s.barrier.signal.var as a bitstring.
rocdl.s.barrier.signal.var
Attributes
memberCnt- Single,I32Attr, 32-bit signless integer attribute
Operands
ptr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Description
Available on gfx1250+.
If memberCnt is 0, the member count is retained from a previous initialization.
Example:
// Signal a named barrier with variable ID.
rocdl.s.barrier.signal.var %ptr member_cnt = 1 : !llvm.ptr<3>
Return op name rocdl.s.barrier.wait as a bitstring.
rocdl.s.barrier.wait
Attributes
id- Single,I16Attr, 16-bit signless integer attribute
Description
Wait on a barrier by id. Available on gfx1200+.
Example:
// Wait on barrier with id -1 (all barriers).
rocdl.s.barrier.wait id = -1
Return op name rocdl.s.get.barrier.state as a bitstring.
rocdl.s.get.barrier.state
Attributes
id- Single,I32Attr, 32-bit signless integer attribute
Results
res- Single,I32, 32-bit signless integer
Description
Available on gfx1200+.
Example:
// Query barrier state by id.
%0 = rocdl.s.get.barrier.state id = 1 -> i32
Return op name rocdl.s.get.named.barrier.state as a bitstring.
rocdl.s.get.named.barrier.state
Operands
ptr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Results
res- Single,I32, 32-bit signless integer
Description
Available on gfx1250+.
Example:
// Query named barrier state by pointer.
%0 = rocdl.s.get.named.barrier.state %ptr : !llvm.ptr<3> -> i32
Return op name rocdl.s.nop as a bitstring.
rocdl.s.nop
Attributes
count- Single,I16Attr, 16-bit signless integer attribute
Description
Insert a number of NOP cycles.
Example:
// Insert a no-op.
rocdl.s.nop 0
Return op name rocdl.s.setprio as a bitstring.
rocdl.s.setprio
Attributes
priority- Single,I16Attr, 16-bit signless integer attribute
Description
Set the wavefront scheduling priority.
Example:
// Set priority to 0.
rocdl.s.setprio 0
Return op name rocdl.s.sleep as a bitstring.
rocdl.s.sleep
Attributes
count- Single,I32Attr, 32-bit signless integer attribute
Description
Sleep for a number of clock cycles.
Example:
// Sleep for a minimum duration.
rocdl.s.sleep 0
Return op name rocdl.s.wait.asynccnt as a bitstring.
rocdl.s.wait.asynccnt - Wait until ASYNCCNT is less than or equal to count
Attributes
count- Single,I16Attr, 16-bit signless integer attribute
Description
Wait for the counter specified to be less-than or equal-to the count
before continuing.
Available on gfx1250+.
Example:
// Wait for async counter to drain.
rocdl.s.wait.asynccnt 0
Return op name rocdl.s.wait.dscnt as a bitstring.
rocdl.s.wait.dscnt - Wait until DSCNT is less than or equal to count
Attributes
count- Single,I16Attr, 16-bit signless integer attribute
Description
Wait for the counter specified to be less-than or equal-to the count
before continuing.
Available on gfx12+.
Example:
// Wait for data-sharing counter to drain.
rocdl.s.wait.dscnt 0
Return op name rocdl.s.wait.expcnt as a bitstring.
rocdl.s.wait.expcnt - Wait until EXPCNT is less than or equal to count
Attributes
count- Single,I16Attr, 16-bit signless integer attribute
Description
Wait for the counter specified to be less-than or equal-to the count
before continuing.
Available on gfx12+.
Example:
// Wait for export counter to drain.
rocdl.s.wait.expcnt 0
Return op name rocdl.s.wait.loadcnt as a bitstring.
rocdl.s.wait.loadcnt - Wait until LOADCNT is less than or equal to count
Attributes
count- Single,I16Attr, 16-bit signless integer attribute
Description
Wait for the counter specified to be less-than or equal-to the count
before continuing.
Available on gfx12+.
Example:
// Wait for load counter to drain.
rocdl.s.wait.loadcnt 0
Return op name rocdl.s.wait.storecnt as a bitstring.
rocdl.s.wait.storecnt - Wait until STORECNT is less than or equal to count
Attributes
count- Single,I16Attr, 16-bit signless integer attribute
Description
Wait for the counter specified to be less-than or equal-to the count
before continuing.
Available on gfx12+.
Example:
// Wait for store counter to drain.
rocdl.s.wait.storecnt 0
Return op name rocdl.s.wait.tensorcnt as a bitstring.
rocdl.s.wait.tensorcnt - Wait until TENSORCNT is less than or equal to count
Attributes
count- Single,I16Attr, 16-bit signless integer attribute
Description
Wait for the counter specified to be less-than or equal-to the count
before continuing.
Available on gfx1250+.
Example:
// Wait for tensor counter to drain.
rocdl.s.wait.tensorcnt 0
Return op name rocdl.s.waitcnt as a bitstring.
rocdl.s.waitcnt
Attributes
bitfield- Single,I32Attr, 32-bit signless integer attribute
Description
Wait for outstanding memory operations to complete, as specified by a bitfield whose semantics depend on the target chipset.
Example:
// Wait for all counters to reach zero.
rocdl.s.waitcnt 0
Return op name rocdl.s.wakeup.barrier as a bitstring.
rocdl.s.wakeup.barrier
Operands
ptr- Single,ROCDLBufferLDS, LLVM pointer in address space 3
Description
Wakes up waves associated with a given named barrier. Note, This op does not release waves waiting at the barrier. It just signal other waves in the same work-group waiting on the indicated named barrier to wake up. Available on gfx1250+.
Example:
// Wake up waves waiting on a named barrier.
rocdl.s.wakeup.barrier %ptr : !llvm.ptr<3>
Return op name rocdl.sched.barrier as a bitstring.
rocdl.sched.barrier
Attributes
mask- Single,ROCDL_SchedGroupMaskAttr, instruction type mask for scheduling barriers
Description
Insert a scheduling barrier with the given mask. The mask is a
bitfield that controls which instruction types may be scheduled
across the barrier. The mask values mirror the llvm.amdgcn.sched.barrier
intrinsic's documented mask values and the AMDGPU backend's
SchedGroupMask enum.
Example:
// Scheduling barrier with no instructions allowed to cross.
rocdl.sched.barrier none
// Allow VALU and all VMEM instructions to cross.
rocdl.sched.barrier valu|all_vmem
Return op name rocdl.sched.group.barrier as a bitstring.
rocdl.sched.group.barrier
Attributes
mask- Single,ROCDL_SchedGroupMaskAttr, instruction type mask for scheduling barrierssize- Single,I32Attr, 32-bit signless integer attributegroupId- Single,I32Attr, 32-bit signless integer attribute
Description
Insert a scheduling group barrier. The first parameter uses the same
scheduling group mask values as rocdl.sched.barrier.
Example:
// Schedule group barrier with mask, size, and group id.
rocdl.sched.group.barrier mfma_wmma, 1, 0
Return op name rocdl.sdot2 as a bitstring.
rocdl.sdot2
Attributes
clamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single,ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2b- Single,ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2c- Single, anonymous/composite constraint, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit signless integer
Description
Packed intra-lane dot-product with optional result clamping (clamp).
Computes res = sum_i a[i]*b[i] + c, where a and b hold packed
4/8/16-bit data (for dot2,dot4,dot8).
Example:
%r = rocdl.sdot2 %a, %b, %c {clamp = true} :
(vector<2xi16>, vector<2xi16>, i32) -> i32
Return op name rocdl.sdot4 as a bitstring.
rocdl.sdot4
Attributes
clamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, 32-bit signless integerb- Single, anonymous/composite constraint, 32-bit signless integerc- Single, anonymous/composite constraint, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit signless integer
Description
Packed intra-lane dot-product with optional result clamping (clamp).
Computes res = sum_i a[i]*b[i] + c, where a and b hold packed
4/8/16-bit data (for dot2,dot4,dot8).
Example:
%r = rocdl.sdot4 %a, %b, %c {clamp = true} :
(i32, i32, i32) -> i32
Return op name rocdl.sdot8 as a bitstring.
rocdl.sdot8
Attributes
clamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, 32-bit signless integerb- Single, anonymous/composite constraint, 32-bit signless integerc- Single, anonymous/composite constraint, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit signless integer
Description
Packed intra-lane dot-product with optional result clamping (clamp).
Computes res = sum_i a[i]*b[i] + c, where a and b hold packed
4/8/16-bit data (for dot2,dot4,dot8).
Example:
%r = rocdl.sdot8 %a, %b, %c {clamp = true} :
(i32, i32, i32) -> i32
Return op name rocdl.sin as a bitstring.
rocdl.sin
Operands
arg- Single,LLVM_AnyFloat, floating point LLVM type
Results
res- Single,LLVM_AnyFloat, floating point LLVM type
Description
Note: In the general case, prefer the conventional arith, math, or llvm ops over this.
Use this ROCDL-specific operation only when you fully understand its implication and
when it is strictly necessary. This op is usually chosen when a small loss in precision is
acceptable in exchange for higher execution speed.
Example:
%0 = rocdl.sin %a f32 -> f32
Return op name rocdl.smfmac.f32.16x16x32.bf16 as a bitstring.
rocdl.smfmac.f32.16x16x32.bf16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.16x16x32.f16 as a bitstring.
rocdl.smfmac.f32.16x16x32.f16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.16x16x64.bf8.bf8 as a bitstring.
rocdl.smfmac.f32.16x16x64.bf8.bf8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.16x16x64.bf8.fp8 as a bitstring.
rocdl.smfmac.f32.16x16x64.bf8.fp8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.16x16x64.bf16 as a bitstring.
rocdl.smfmac.f32.16x16x64.bf16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of bfloat16 type values of length 8b- Single, anonymous/composite constraint, fixed-length vector of bfloat16 type values of length 16c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.16x16x64.f16 as a bitstring.
rocdl.smfmac.f32.16x16x64.f16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 8b- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 16c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.16x16x64.fp8.bf8 as a bitstring.
rocdl.smfmac.f32.16x16x64.fp8.bf8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.16x16x64.fp8.fp8 as a bitstring.
rocdl.smfmac.f32.16x16x64.fp8.fp8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.16x16x128.bf8.bf8 as a bitstring.
rocdl.smfmac.f32.16x16x128.bf8.bf8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.16x16x128.bf8.fp8 as a bitstring.
rocdl.smfmac.f32.16x16x128.bf8.fp8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.16x16x128.fp8.bf8 as a bitstring.
rocdl.smfmac.f32.16x16x128.fp8.bf8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.16x16x128.fp8.fp8 as a bitstring.
rocdl.smfmac.f32.16x16x128.fp8.fp8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 4
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.32x32x16.bf16 as a bitstring.
rocdl.smfmac.f32.32x32x16.bf16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 16-bit signless integer values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.32x32x16.f16 as a bitstring.
rocdl.smfmac.f32.32x32x16.f16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.32x32x32.bf8.bf8 as a bitstring.
rocdl.smfmac.f32.32x32x32.bf8.bf8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.32x32x32.bf8.fp8 as a bitstring.
rocdl.smfmac.f32.32x32x32.bf8.fp8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.32x32x32.bf16 as a bitstring.
rocdl.smfmac.f32.32x32x32.bf16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of bfloat16 type values of length 8b- Single, anonymous/composite constraint, fixed-length vector of bfloat16 type values of length 16c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.32x32x32.f16 as a bitstring.
rocdl.smfmac.f32.32x32x32.f16
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 8b- Single, anonymous/composite constraint, fixed-length vector of 16-bit float values of length 16c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.32x32x32.fp8.bf8 as a bitstring.
rocdl.smfmac.f32.32x32x32.fp8.bf8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.32x32x32.fp8.fp8 as a bitstring.
rocdl.smfmac.f32.32x32x32.fp8.fp8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.32x32x64.bf8.bf8 as a bitstring.
rocdl.smfmac.f32.32x32x64.bf8.bf8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.32x32x64.bf8.fp8 as a bitstring.
rocdl.smfmac.f32.32x32x64.bf8.fp8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.32x32x64.fp8.bf8 as a bitstring.
rocdl.smfmac.f32.32x32x64.fp8.bf8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.f32.32x32x64.fp8.fp8 as a bitstring.
rocdl.smfmac.f32.32x32x64.fp8.fp8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit float values of length 16
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.i32.16x16x64.i8 as a bitstring.
rocdl.smfmac.i32.16x16x64.i8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.i32.16x16x128.i8 as a bitstring.
rocdl.smfmac.i32.16x16x128.i8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.i32.32x32x32.i8 as a bitstring.
rocdl.smfmac.i32.32x32x32.i8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 2b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4c- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.smfmac.i32.32x32x64.i8 as a bitstring.
rocdl.smfmac.i32.32x32x64.i8
Attributes
cbsz- Single,I32Attr, 32-bit signless integer attributeabid- Single,I32Attr, 32-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 4b- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 8c- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16index- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, fixed-length vector of 32-bit signless integer values of length 16
Description
Sparse matrix fused multiply-accumulate (SMFMAC) intrinsic with 2:4
structured sparsity. The index operand provides the sparsity metadata,
and cbsz/abid control broadcast modes.
Example:
// SMFMAC with f16 inputs.
%r0 = rocdl.smfmac.f32.16x16x32.f16 %a0, %b0, %c0, %idx, 0, 0 :
(vector<4xf16>, vector<8xf16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with bf16 inputs.
%r1 = rocdl.smfmac.f32.16x16x32.bf16 %a1, %b1, %c0, %idx, 0, 0 :
(vector<4xi16>, vector<8xi16>, vector<4xf32>, i32) -> vector<4xf32>
// SMFMAC with i8 inputs and i32 accumulator.
%r2 = rocdl.smfmac.i32.16x16x64.i8 %a2, %b2, %c2, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xi32>, i32) -> vector<4xi32>
// SMFMAC with fp8 inputs.
%r3 = rocdl.smfmac.f32.16x16x64.fp8.fp8 %a2, %b2, %c0, %idx, 0, 0 :
(vector<2xi32>, vector<4xi32>, vector<4xf32>, i32) -> vector<4xf32>
Return op name rocdl.sqrt as a bitstring.
rocdl.sqrt
Operands
arg- Single,LLVM_AnyFloat, floating point LLVM type
Results
res- Single,LLVM_AnyFloat, floating point LLVM type
Description
Note: In the general case, prefer the conventional arith, math, or llvm ops over this.
Use this ROCDL-specific operation only when you fully understand its implication and
when it is strictly necessary. This op is usually chosen when a small loss in precision is
acceptable in exchange for higher execution speed.
Example:
%0 = rocdl.sqrt %a f32 -> f32
Return op name rocdl.sudot4 as a bitstring.
rocdl.sudot4
Attributes
signA- Single,I1Attr, 1-bit signless integer attributesignB- Single,I1Attr, 1-bit signless integer attributeclamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single,I32, 32-bit signless integerb- Single,I32, 32-bit signless integerc- Single,I32, 32-bit signless integer
Results
res- Single,I32, 32-bit signless integer
Description
Mixed-signedness packed dot-product with per-operand sign controls.
Computes res = sum_i a[i]*b[i] + c. Each lane of a is treated as
signed when signA = true; when signA = false, the unsigned lane
value is zero-extended into a wider signed integer. signB controls
the same for b. clamp controls result clamping.
These ops correspond to RDNA's unified mixed-sign v_dot4_i32_iu8
and v_dot8_i32_iu4 instructions (gfx11+).
Example:
%r = rocdl.sudot4 %a, %b, %c
{signA = true, signB = false, clamp = true} :
(i32, i32, i32) -> i32
Return op name rocdl.sudot8 as a bitstring.
rocdl.sudot8
Attributes
signA- Single,I1Attr, 1-bit signless integer attributesignB- Single,I1Attr, 1-bit signless integer attributeclamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single,I32, 32-bit signless integerb- Single,I32, 32-bit signless integerc- Single,I32, 32-bit signless integer
Results
res- Single,I32, 32-bit signless integer
Description
Mixed-signedness packed dot-product with per-operand sign controls.
Computes res = sum_i a[i]*b[i] + c. Each lane of a is treated as
signed when signA = true; when signA = false, the unsigned lane
value is zero-extended into a wider signed integer. signB controls
the same for b. clamp controls result clamping.
These ops correspond to RDNA's unified mixed-sign v_dot4_i32_iu8
and v_dot8_i32_iu4 instructions (gfx11+).
Example:
%r = rocdl.sudot8 %a, %b, %c
{signA = true, signB = false, clamp = true} :
(i32, i32, i32) -> i32
Return op name rocdl.swmmac.bf16.16x16x32.bf16 as a bitstring.
rocdl.swmmac.bf16.16x16x32.bf16
Operands
a- Single, anonymous/composite constraint, LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, LLVM dialect-compatible vector of integerindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, LLVM dialect-compatible vector of integer
Return op name rocdl.swmmac.bf16.16x16x64.bf16 as a bitstring.
rocdl.swmmac.bf16.16x16x64.bf16
Attributes
signA- Single,I1Attr, 1-bit signless integer attributesignB- Single,I1Attr, 1-bit signless integer attributereuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 typeb- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 typec- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 typeindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
Return op name rocdl.swmmac.bf16f32.16x16x64.bf16 as a bitstring.
rocdl.swmmac.bf16f32.16x16x64.bf16
Attributes
signA- Single,I1Attr, 1-bit signless integer attributesignB- Single,I1Attr, 1-bit signless integer attributereuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 typeb- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 typec- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 typeindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
Return op name rocdl.swmmac.f16.16x16x32.f16 as a bitstring.
rocdl.swmmac.f16.16x16x32.f16
Operands
a- Single, anonymous/composite constraint, LLVM dialect-compatible vector of 16-bit floatb- Single, anonymous/composite constraint, LLVM dialect-compatible vector of 16-bit floatc- Single, anonymous/composite constraint, LLVM dialect-compatible vector of 16-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, LLVM dialect-compatible vector of 16-bit float
Return op name rocdl.swmmac.f16.16x16x64.f16 as a bitstring.
rocdl.swmmac.f16.16x16x64.f16
Attributes
signA- Single,I1Attr, 1-bit signless integer attributesignB- Single,I1Attr, 1-bit signless integer attributereuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatb- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatc- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Return op name rocdl.swmmac.f16.16x16x128.bf8.bf8 as a bitstring.
rocdl.swmmac.f16.16x16x128.bf8.bf8
Attributes
reuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Return op name rocdl.swmmac.f16.16x16x128.bf8.fp8 as a bitstring.
rocdl.swmmac.f16.16x16x128.bf8.fp8
Attributes
reuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Return op name rocdl.swmmac.f16.16x16x128.fp8.bf8 as a bitstring.
rocdl.swmmac.f16.16x16x128.fp8.bf8
Attributes
reuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Return op name rocdl.swmmac.f16.16x16x128.fp8.fp8 as a bitstring.
rocdl.swmmac.f16.16x16x128.fp8.fp8
Attributes
reuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Return op name rocdl.swmmac.f32.16x16x32.bf8.bf8 as a bitstring.
rocdl.swmmac.f32.16x16x32.bf8.bf8
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Return op name rocdl.swmmac.f32.16x16x32.bf8.fp8 as a bitstring.
rocdl.swmmac.f32.16x16x32.bf8.fp8
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Return op name rocdl.swmmac.f32.16x16x32.bf16 as a bitstring.
rocdl.swmmac.f32.16x16x32.bf16
Operands
a- Single, anonymous/composite constraint, LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit float
Return op name rocdl.swmmac.f32.16x16x32.f16 as a bitstring.
rocdl.swmmac.f32.16x16x32.f16
Operands
a- Single, anonymous/composite constraint, LLVM dialect-compatible vector of 16-bit floatb- Single, anonymous/composite constraint, LLVM dialect-compatible vector of 16-bit floatc- Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, LLVM dialect-compatible vector of 32-bit float
Return op name rocdl.swmmac.f32.16x16x32.fp8.bf8 as a bitstring.
rocdl.swmmac.f32.16x16x32.fp8.bf8
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Return op name rocdl.swmmac.f32.16x16x32.fp8.fp8 as a bitstring.
rocdl.swmmac.f32.16x16x32.fp8.fp8
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Return op name rocdl.swmmac.f32.16x16x64.bf16 as a bitstring.
rocdl.swmmac.f32.16x16x64.bf16
Attributes
signA- Single,I1Attr, 1-bit signless integer attributesignB- Single,I1Attr, 1-bit signless integer attributereuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 typeb- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 typec- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Return op name rocdl.swmmac.f32.16x16x64.f16 as a bitstring.
rocdl.swmmac.f32.16x16x64.f16
Attributes
signA- Single,I1Attr, 1-bit signless integer attributesignB- Single,I1Attr, 1-bit signless integer attributereuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatb- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Return op name rocdl.swmmac.f32.16x16x128.bf8.bf8 as a bitstring.
rocdl.swmmac.f32.16x16x128.bf8.bf8
Attributes
reuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Return op name rocdl.swmmac.f32.16x16x128.bf8.fp8 as a bitstring.
rocdl.swmmac.f32.16x16x128.bf8.fp8
Attributes
reuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Return op name rocdl.swmmac.f32.16x16x128.fp8.bf8 as a bitstring.
rocdl.swmmac.f32.16x16x128.fp8.bf8
Attributes
reuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Return op name rocdl.swmmac.f32.16x16x128.fp8.fp8 as a bitstring.
rocdl.swmmac.f32.16x16x128.fp8.fp8
Attributes
reuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit floatindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Return op name rocdl.swmmac.i32.16x16x32.iu4 as a bitstring.
rocdl.swmmac.i32.16x16x32.iu4
Attributes
signA- Single,I1Attr, 1-bit signless integer attributesignB- Single,I1Attr, 1-bit signless integer attributeclamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
Return op name rocdl.swmmac.i32.16x16x32.iu8 as a bitstring.
rocdl.swmmac.i32.16x16x32.iu8
Attributes
signA- Single,I1Attr, 1-bit signless integer attributesignB- Single,I1Attr, 1-bit signless integer attributeclamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
Return op name rocdl.swmmac.i32.16x16x64.iu4 as a bitstring.
rocdl.swmmac.i32.16x16x64.iu4
Attributes
signA- Single,I1Attr, 1-bit signless integer attributesignB- Single,I1Attr, 1-bit signless integer attributeclamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
Return op name rocdl.swmmac.i32.16x16x128.iu8 as a bitstring.
rocdl.swmmac.i32.16x16x128.iu8
Attributes
signA- Single,I1Attr, 1-bit signless integer attributesignB- Single,I1Attr, 1-bit signless integer attributereuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attributeclamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerindex- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
Return op name rocdl.tanh as a bitstring.
rocdl.tanh
Operands
arg- Single,LLVM_AnyFloat, floating point LLVM type
Results
res- Single,LLVM_AnyFloat, floating point LLVM type
Description
Note: In the general case, prefer the conventional arith, math, or llvm ops over this.
Use this ROCDL-specific operation only when you fully understand its implication and
when it is strictly necessary. This op is usually chosen when a small loss in precision is
acceptable in exchange for higher execution speed.
Example:
%0 = rocdl.tanh %a f32 -> f32
Return op name rocdl.tensor.load.to.lds as a bitstring.
rocdl.tensor.load.to.lds - Base class for ROCDL tensor load/store to/from LDS.
Attributes
cachePolicy- Single,ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
dgroup0- Single,ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4dgroup1- Single,ROCDL_V8I32Type, fixed-length vector of 32-bit signless integer values of length 8dgroup2- Single,ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4dgroup3- Single,ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4dgroup4- Single,ROCDL_V8I32Type, fixed-length vector of 32-bit signless integer values of length 8
Description
Moves tiles of tensor data between global memory and LDS. The tile is described by the $dgroup descriptors. 5 $dgroup descriptors allows for movement of up to 5D tensors. $cachePolicy describes the memory scope and an indicator of expected data re-use.
This op is for gfx1250+ architectures.
Example:
// Tensor load from global memory to LDS using 4 descriptor groups.
rocdl.tensor.load.to.lds %dg0, %dg1, %dg2, %dg3, %dg4, 0 : vector<4xi32>, vector<8xi32>
// Tensor store from LDS to global memory using 4 descriptor groups.
rocdl.tensor.store.from.lds %dg0, %dg1, %dg2, %dg3, %dg4, 0 : vector<4xi32>, vector<8xi32>
Return op name rocdl.tensor.store.from.lds as a bitstring.
rocdl.tensor.store.from.lds - Base class for ROCDL tensor load/store to/from LDS.
Attributes
cachePolicy- Single,ROCDL_Gfx12NonAtomicCachePolicyCompatAttr, gfx12 non-atomic AMDGPU cache policy attributealias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraynoalias_scopes- Optional,LLVM_AliasScopeArrayAttr, LLVM dialect alias scope arraytbaa- Optional,LLVM_TBAATagArrayAttr, LLVM dialect TBAA tag metadata array
Operands
dgroup0- Single,ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4dgroup1- Single,ROCDL_V8I32Type, fixed-length vector of 32-bit signless integer values of length 8dgroup2- Single,ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4dgroup3- Single,ROCDL_V4I32Type, fixed-length vector of 32-bit signless integer values of length 4dgroup4- Single,ROCDL_V8I32Type, fixed-length vector of 32-bit signless integer values of length 8
Description
Moves tiles of tensor data between global memory and LDS. The tile is described by the $dgroup descriptors. 5 $dgroup descriptors allows for movement of up to 5D tensors. $cachePolicy describes the memory scope and an indicator of expected data re-use.
This op is for gfx1250+ architectures.
Example:
// Tensor load from global memory to LDS using 4 descriptor groups.
rocdl.tensor.load.to.lds %dg0, %dg1, %dg2, %dg3, %dg4, 0 : vector<4xi32>, vector<8xi32>
// Tensor store from LDS to global memory using 4 descriptor groups.
rocdl.tensor.store.from.lds %dg0, %dg1, %dg2, %dg3, %dg4, 0 : vector<4xi32>, vector<8xi32>
Return op name rocdl.udot2 as a bitstring.
rocdl.udot2
Attributes
clamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single,ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2b- Single,ROCDL_V2I16Type, fixed-length vector of 16-bit signless integer values of length 2c- Single, anonymous/composite constraint, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit signless integer
Description
Packed intra-lane dot-product with optional result clamping (clamp).
Computes res = sum_i a[i]*b[i] + c, where a and b hold packed
4/8/16-bit data (for dot2,dot4,dot8).
Example:
%r = rocdl.udot2 %a, %b, %c {clamp = true} :
(vector<2xi16>, vector<2xi16>, i32) -> i32
Return op name rocdl.udot4 as a bitstring.
rocdl.udot4
Attributes
clamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, 32-bit signless integerb- Single, anonymous/composite constraint, 32-bit signless integerc- Single, anonymous/composite constraint, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit signless integer
Description
Packed intra-lane dot-product with optional result clamping (clamp).
Computes res = sum_i a[i]*b[i] + c, where a and b hold packed
4/8/16-bit data (for dot2,dot4,dot8).
Example:
%r = rocdl.udot4 %a, %b, %c {clamp = true} :
(i32, i32, i32) -> i32
Return op name rocdl.udot8 as a bitstring.
rocdl.udot8
Attributes
clamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, 32-bit signless integerb- Single, anonymous/composite constraint, 32-bit signless integerc- Single, anonymous/composite constraint, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit signless integer
Description
Packed intra-lane dot-product with optional result clamping (clamp).
Computes res = sum_i a[i]*b[i] + c, where a and b hold packed
4/8/16-bit data (for dot2,dot4,dot8).
Example:
%r = rocdl.udot8 %a, %b, %c {clamp = true} :
(i32, i32, i32) -> i32
Return op name rocdl.update.dpp as a bitstring.
rocdl.update.dpp
Attributes
dppCtrl- Single,I32Attr, 32-bit signless integer attributerowMask- Single,I32Attr, 32-bit signless integer attributebankMask- Single,I32Attr, 32-bit signless integer attributeboundCtrl- Single,I1Attr, 1-bit signless integer attribute
Operands
old- Single,LLVM_Type, LLVM dialect-compatible typesrc- Single,LLVM_Type, LLVM dialect-compatible type
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Return op name rocdl.wait.asyncmark as a bitstring.
rocdl.wait.asyncmark - Wait until N or fewer async operation groups are unexecuted
Attributes
count- Single,I16Attr, 16-bit signless integer attribute
Description
This operation, along with rocdl.asyncmark, forms the compiler-provided
framework for explicitly tracking asynchronous operations.
At the point where a wait.asyncmark operation is executed, all async operations
that were parts of any async group (established by asyncmark in program order)
other than the count previously-added ones will have finished executing.
For more detail, including on how this mechanism composes with function calls, see the LLVM documentation on async tracking.
Available on gfx9 and later.
Example:
// Wait until at most N async groups remain outstanding.
rocdl.wait.asyncmark 1Usage example:
rocdl.tensor.load.to.lds ...
rocdl.global.async.load.to.lds ...
rocdl.asyncmark
rocdl.tensor.load.to.lds ...
rocdl.global.async.load.to.lds ...
rocdl.asyncmark
rocdl.wait.asyncmark 1 // First group of loads completes after this
Return op name rocdl.wave.barrier as a bitstring.
rocdl.wave.barrier
Description
Insert a wave-level (subgroup) barrier. Synchronizes lanes within a single wave/wavefront without any memory ordering guarantees.
Example:
rocdl.wave.barrier
Return op name rocdl.wave.id as a bitstring.
rocdl.wave.id
Attributes
range- Optional,LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Read a hardware register for thread/workgroup/cluster identification.
An optional range attribute can constrain the returned value.
Example:
// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32
// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32
Return op name rocdl.wavefrontsize as a bitstring.
rocdl.wavefrontsize
Attributes
range- Optional,LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Read a hardware register for thread/workgroup/cluster identification.
An optional range attribute can constrain the returned value.
Example:
// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32
// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32
Return op name rocdl.wmma.bf16.16x16x16.bf16 as a bitstring.
rocdl.wmma.bf16.16x16x16.bf16
Attributes
opsel- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
Results
res- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
Description
Wave Matrix Multiply-Accumulate (WMMA) with output operand selection.
Example:
// WMMA f16 with opsel control.
%r = rocdl.wmma.f16.16x16x16.f16 %a, %b, %c {opsel = false} :
(vector<16xf16>, vector<16xf16>, vector<16xf16>) -> vector<16xf16>
Return op name rocdl.wmma.bf16.16x16x32.bf16 as a bitstring.
rocdl.wmma.bf16.16x16x32.bf16
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 typeb- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 typec- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
Results
res- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.bf16f32.16x16x32.bf16 as a bitstring.
rocdl.wmma.bf16f32.16x16x32.bf16
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 typeb- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 typec- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 type
Description
Wave Matrix Multiply-Accumulate (WMMA) with different C and D types.
Example:
// WMMA bf16 output from f32 accumulator with bf16 inputs.
%r = rocdl.wmma.bf16f32.16x16x32.bf16 %a, %b, %c, modC = none :
(vector<16xbf16>, vector<16xbf16>, vector<8xf32>) -> vector<16xbf16>
Return op name rocdl.wmma.f16.16x16x16.f16 as a bitstring.
rocdl.wmma.f16.16x16x16.f16
Attributes
opsel- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatb- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatc- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Results
res- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with output operand selection.
Example:
// WMMA f16 with opsel control.
%r = rocdl.wmma.f16.16x16x16.f16 %a, %b, %c {opsel = false} :
(vector<16xf16>, vector<16xf16>, vector<16xf16>) -> vector<16xf16>
Return op name rocdl.wmma.f16.16x16x32.f16 as a bitstring.
rocdl.wmma.f16.16x16x32.f16
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatb- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatc- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Results
res- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f16.16x16x64.bf8_bf8 as a bitstring.
rocdl.wmma.f16.16x16x64.bf8_bf8
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Results
res- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f16.16x16x64.bf8_fp8 as a bitstring.
rocdl.wmma.f16.16x16x64.bf8_fp8
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Results
res- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f16.16x16x64.fp8_bf8 as a bitstring.
rocdl.wmma.f16.16x16x64.fp8_bf8
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Results
res- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f16.16x16x64.fp8_fp8 as a bitstring.
rocdl.wmma.f16.16x16x64.fp8_fp8
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Results
res- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f16.16x16x128.bf8_bf8 as a bitstring.
rocdl.wmma.f16.16x16x128.bf8_bf8
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Results
res- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f16.16x16x128.bf8_fp8 as a bitstring.
rocdl.wmma.f16.16x16x128.bf8_fp8
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Results
res- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f16.16x16x128.fp8_bf8 as a bitstring.
rocdl.wmma.f16.16x16x128.fp8_bf8
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Results
res- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f16.16x16x128.fp8_fp8 as a bitstring.
rocdl.wmma.f16.16x16x128.fp8_fp8
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Results
res- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x4.f32 as a bitstring.
rocdl.wmma.f32.16x16x4.f32
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit floatb- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit floatc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x16.bf8_bf8 as a bitstring.
rocdl.wmma.f32.16x16x16.bf8_bf8
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) intrinsic.
Example:
// WMMA with f16 inputs and f32 accumulator.
%r = rocdl.wmma.f32.16x16x16.f16 %a, %b, %c :
(vector<16xf16>, vector<16xf16>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x16.bf8_fp8 as a bitstring.
rocdl.wmma.f32.16x16x16.bf8_fp8
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) intrinsic.
Example:
// WMMA with f16 inputs and f32 accumulator.
%r = rocdl.wmma.f32.16x16x16.f16 %a, %b, %c :
(vector<16xf16>, vector<16xf16>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x16.bf16 as a bitstring.
rocdl.wmma.f32.16x16x16.bf16
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) intrinsic.
Example:
// WMMA with f16 inputs and f32 accumulator.
%r = rocdl.wmma.f32.16x16x16.f16 %a, %b, %c :
(vector<16xf16>, vector<16xf16>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x16.f16 as a bitstring.
rocdl.wmma.f32.16x16x16.f16
Operands
a- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatb- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) intrinsic.
Example:
// WMMA with f16 inputs and f32 accumulator.
%r = rocdl.wmma.f32.16x16x16.f16 %a, %b, %c :
(vector<16xf16>, vector<16xf16>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x16.fp8_bf8 as a bitstring.
rocdl.wmma.f32.16x16x16.fp8_bf8
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) intrinsic.
Example:
// WMMA with f16 inputs and f32 accumulator.
%r = rocdl.wmma.f32.16x16x16.f16 %a, %b, %c :
(vector<16xf16>, vector<16xf16>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x16.fp8_fp8 as a bitstring.
rocdl.wmma.f32.16x16x16.fp8_fp8
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) intrinsic.
Example:
// WMMA with f16 inputs and f32 accumulator.
%r = rocdl.wmma.f32.16x16x16.f16 %a, %b, %c :
(vector<16xf16>, vector<16xf16>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x32.bf16 as a bitstring.
rocdl.wmma.f32.16x16x32.bf16
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 typeb- Single, anonymous/composite constraint, bfloat16 type or LLVM dialect-compatible vector of bfloat16 typec- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x32.f16 as a bitstring.
rocdl.wmma.f32.16x16x32.f16
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatb- Single, anonymous/composite constraint, 16-bit float or LLVM dialect-compatible vector of 16-bit floatc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x64.bf8_bf8 as a bitstring.
rocdl.wmma.f32.16x16x64.bf8_bf8
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x64.bf8_fp8 as a bitstring.
rocdl.wmma.f32.16x16x64.bf8_fp8
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x64.fp8_bf8 as a bitstring.
rocdl.wmma.f32.16x16x64.fp8_bf8
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x64.fp8_fp8 as a bitstring.
rocdl.wmma.f32.16x16x64.fp8_fp8
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x128.bf8_bf8 as a bitstring.
rocdl.wmma.f32.16x16x128.bf8_bf8
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x128.bf8_fp8 as a bitstring.
rocdl.wmma.f32.16x16x128.bf8_fp8
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x128.fp8_bf8 as a bitstring.
rocdl.wmma.f32.16x16x128.fp8_bf8
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.f32.16x16x128.fp8_fp8 as a bitstring.
rocdl.wmma.f32.16x16x128.fp8_fp8
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Wave Matrix Multiply-Accumulate (WMMA) with modC and reuse controls.
Example:
// WMMA f32 with fp8 inputs and modC/reuse controls.
%r = rocdl.wmma.f32.16x16x64.fp8_fp8 %a, %b, %c, modC = none :
(vector<16xi32>, vector<16xi32>, vector<8xf32>) -> vector<8xf32>
Return op name rocdl.wmma.i32.16x16x16.iu4 as a bitstring.
rocdl.wmma.i32.16x16x16.iu4
Attributes
signA- Single,I1Attr, 1-bit signless integer attributesignB- Single,I1Attr, 1-bit signless integer attributeclamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
Results
res- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
Description
Wave Matrix Multiply-Accumulate (WMMA) for integer types with sign and clamp control.
Example:
// WMMA i32 with unsigned i8 inputs.
%r = rocdl.wmma.i32.16x16x16.iu8 %a, %b, %c
{signA = false, signB = false, clamp = false} :
(vector<4xi32>, vector<4xi32>, vector<8xi32>) -> vector<8xi32>
Return op name rocdl.wmma.i32.16x16x16.iu8 as a bitstring.
rocdl.wmma.i32.16x16x16.iu8
Attributes
signA- Single,I1Attr, 1-bit signless integer attributesignB- Single,I1Attr, 1-bit signless integer attributeclamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
Results
res- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
Description
Wave Matrix Multiply-Accumulate (WMMA) for integer types with sign and clamp control.
Example:
// WMMA i32 with unsigned i8 inputs.
%r = rocdl.wmma.i32.16x16x16.iu8 %a, %b, %c
{signA = false, signB = false, clamp = false} :
(vector<4xi32>, vector<4xi32>, vector<8xi32>) -> vector<8xi32>
Return op name rocdl.wmma.i32.16x16x32.iu4 as a bitstring.
rocdl.wmma.i32.16x16x32.iu4
Attributes
signA- Single,I1Attr, 1-bit signless integer attributesignB- Single,I1Attr, 1-bit signless integer attributeclamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
Results
res- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
Description
Wave Matrix Multiply-Accumulate (WMMA) for integer types with sign and clamp control.
Example:
// WMMA i32 with unsigned i8 inputs.
%r = rocdl.wmma.i32.16x16x16.iu8 %a, %b, %c
{signA = false, signB = false, clamp = false} :
(vector<4xi32>, vector<4xi32>, vector<8xi32>) -> vector<8xi32>
Return op name rocdl.wmma.i32.16x16x64.iu8 as a bitstring.
rocdl.wmma.i32.16x16x64.iu8
Attributes
signA- Single,I1Attr, 1-bit signless integer attributesignB- Single,I1Attr, 1-bit signless integer attributereuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attributeclamp- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
Results
res- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integer
Description
Wave Matrix Multiply-Accumulate (WMMA) for integer types with sign, reuse, and clamp controls.
Example:
// WMMA i32 with unsigned i8 inputs and reuse controls.
%r = rocdl.wmma.i32.16x16x64.iu8 %a, %b, %c
{signA = false, signB = false, reuseA = false, reuseB = false, clamp = false} :
(vector<8xi32>, vector<8xi32>, vector<8xi32>) -> vector<8xi32>
Return op name rocdl.wmma.scale16.f32.16x16x128.f8f6f4 as a bitstring.
rocdl.wmma.scale16.f32.16x16x128.f8f6f4
Attributes
fmtA- Single,ROCDL_MatrixFormatAttr, matrix operand formats selected by scaled MFMA/WMMA format fieldsfmtB- Single,ROCDL_MatrixFormatAttr, matrix operand formats selected by scaled MFMA/WMMA format fieldsmodC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersscaleAType- Single,ROCDL_WMMAMatrixScaleAttr, matrix scale row selectorfmtScaleA- Single,ROCDL_WMMAMatrixScaleFormatAttr, matrix scale exponent formatsscaleBType- Single,ROCDL_WMMAMatrixScaleAttr, matrix scale row selectorfmtScaleB- Single,ROCDL_WMMAMatrixScaleFormatAttr, matrix scale exponent formatsreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit floatscaleA- Single,I64, 64-bit signless integerscaleB- Single,I64, 64-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Scaled Wave Matrix Multiply-Accumulate (WMMA) with per-operand scaling.
Example:
// Scaled WMMA with f8f6f4 format inputs.
%r = rocdl.wmma.scale.f32.16x16x128.f8f6f4 %a, %b, %c, %scaleA, %scaleB
fmtA = fp8_e4m3, fmtB = fp8_e4m3, modC = none,
scaleAType = row0, fmtScaleA = e8, scaleBType = row0, fmtScaleB = e8 :
(vector<16xi32>, vector<16xi32>, vector<8xf32>, i32, i32) -> vector<8xf32>
Return op name rocdl.wmma.scale16.f32.32x16x128.f4 as a bitstring.
rocdl.wmma.scale16.f32.32x16x128.f4
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersscaleAType- Single,ROCDL_WMMAMatrixScaleAttr, matrix scale row selectorfmtScaleA- Single,ROCDL_WMMAMatrixScaleFormatAttr, matrix scale exponent formatsscaleBType- Single,ROCDL_WMMAMatrixScaleAttr, matrix scale row selectorfmtScaleB- Single,ROCDL_WMMAMatrixScaleFormatAttr, matrix scale exponent formatsreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit floatscaleA- Single,I64, 64-bit signless integerscaleB- Single,I64, 64-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Scaled Wave Matrix Multiply-Accumulate (WMMA) for F4 format inputs.
Example:
// Scaled WMMA with f4 format inputs.
%r = rocdl.wmma.scale.f32.16x16x128.f4 %a, %b, %c, %scaleA, %scaleB
modC = none, scaleAType = row0, fmtScaleA = e8,
scaleBType = row0, fmtScaleB = e8 :
(vector<8xi32>, vector<8xi32>, vector<8xf32>, i32, i32) -> vector<8xf32>
Return op name rocdl.wmma.scale.f32.16x16x128.f8f6f4 as a bitstring.
rocdl.wmma.scale.f32.16x16x128.f8f6f4
Attributes
fmtA- Single,ROCDL_MatrixFormatAttr, matrix operand formats selected by scaled MFMA/WMMA format fieldsfmtB- Single,ROCDL_MatrixFormatAttr, matrix operand formats selected by scaled MFMA/WMMA format fieldsmodC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersscaleAType- Single,ROCDL_WMMAMatrixScaleAttr, matrix scale row selectorfmtScaleA- Single,ROCDL_WMMAMatrixScaleFormatAttr, matrix scale exponent formatsscaleBType- Single,ROCDL_WMMAMatrixScaleAttr, matrix scale row selectorfmtScaleB- Single,ROCDL_WMMAMatrixScaleFormatAttr, matrix scale exponent formatsreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit floatscaleA- Single,I32, 32-bit signless integerscaleB- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Scaled Wave Matrix Multiply-Accumulate (WMMA) with per-operand scaling.
Example:
// Scaled WMMA with f8f6f4 format inputs.
%r = rocdl.wmma.scale.f32.16x16x128.f8f6f4 %a, %b, %c, %scaleA, %scaleB
fmtA = fp8_e4m3, fmtB = fp8_e4m3, modC = none,
scaleAType = row0, fmtScaleA = e8, scaleBType = row0, fmtScaleB = e8 :
(vector<16xi32>, vector<16xi32>, vector<8xf32>, i32, i32) -> vector<8xf32>
Return op name rocdl.wmma.scale.f32.32x16x128.f4 as a bitstring.
rocdl.wmma.scale.f32.32x16x128.f4
Attributes
modC- Single,ROCDL_WMMACModifierAttr, WMMA C operand modifiersscaleAType- Single,ROCDL_WMMAMatrixScaleAttr, matrix scale row selectorfmtScaleA- Single,ROCDL_WMMAMatrixScaleFormatAttr, matrix scale exponent formatsscaleBType- Single,ROCDL_WMMAMatrixScaleAttr, matrix scale row selectorfmtScaleB- Single,ROCDL_WMMAMatrixScaleFormatAttr, matrix scale exponent formatsreuseA- Single,I1Attr, 1-bit signless integer attributereuseB- Single,I1Attr, 1-bit signless integer attribute
Operands
a- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerb- Single, anonymous/composite constraint, integer or LLVM dialect-compatible vector of integerc- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit floatscaleA- Single,I32, 32-bit signless integerscaleB- Single,I32, 32-bit signless integer
Results
res- Single, anonymous/composite constraint, 32-bit float or LLVM dialect-compatible vector of 32-bit float
Description
Scaled Wave Matrix Multiply-Accumulate (WMMA) for F4 format inputs.
Example:
// Scaled WMMA with f4 format inputs.
%r = rocdl.wmma.scale.f32.16x16x128.f4 %a, %b, %c, %scaleA, %scaleB
modC = none, scaleAType = row0, fmtScaleA = e8,
scaleBType = row0, fmtScaleB = e8 :
(vector<8xi32>, vector<8xi32>, vector<8xf32>, i32, i32) -> vector<8xf32>
Return op name rocdl.workgroup.id.x as a bitstring.
rocdl.workgroup.id.x
Attributes
range- Optional,LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Read a hardware register for thread/workgroup/cluster identification.
An optional range attribute can constrain the returned value.
Example:
// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32
// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32
Return op name rocdl.workgroup.id.y as a bitstring.
rocdl.workgroup.id.y
Attributes
range- Optional,LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Read a hardware register for thread/workgroup/cluster identification.
An optional range attribute can constrain the returned value.
Example:
// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32
// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32
Return op name rocdl.workgroup.id.z as a bitstring.
rocdl.workgroup.id.z
Attributes
range- Optional,LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Read a hardware register for thread/workgroup/cluster identification.
An optional range attribute can constrain the returned value.
Example:
// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32
// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32
Return op name rocdl.workitem.id.x as a bitstring.
rocdl.workitem.id.x
Attributes
range- Optional,LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Read a hardware register for thread/workgroup/cluster identification.
An optional range attribute can constrain the returned value.
Example:
// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32
// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32
Return op name rocdl.workitem.id.y as a bitstring.
rocdl.workitem.id.y
Attributes
range- Optional,LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Read a hardware register for thread/workgroup/cluster identification.
An optional range attribute can constrain the returned value.
Example:
// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32
// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32
Return op name rocdl.workitem.id.z as a bitstring.
rocdl.workitem.id.z
Attributes
range- Optional,LLVM_ConstantRangeAttr, A range of two integers, corresponding to LLVM's ConstantRange
Results
res- Single,LLVM_Type, LLVM dialect-compatible type
Description
Read a hardware register for thread/workgroup/cluster identification.
An optional range attribute can constrain the returned value.
Example:
// Read the workitem id in the x dimension.
%0 = rocdl.workitem.id.x : i32
// Read with a known range constraint.
%1 = rocdl.workitem.id.x range <i32, 0, 64> : i32