Inside a linalg.matmul: A Deep Dive into IREE's Compilation Pipeline
As part of my research, I recently joined both the LLVM and IREE communities on Github and started contributing to llvm-project and iree, respectively. I'm really honored to be part of these welcoming and friendly communities, that not only are at the forefront of compiler research, but also have a strong focus on collaboration and knowledge sharing. To know more about, just visit my GitHub profile.
Lately, I have been reading a lot of MLIR code generated by IREE, and I find it difficult to read at first glance—not because of IREE.
So, with the hope to help other folks understand the compilation pipeline of IREE, and to help myself understand it better, I decided to write this post.
It is a deep dive into the compilation pipeline of IREE, focusing on the linalg.matmul operation.
Well, the IREE blog has incredible posts, and I strongly recommend reading all of them.
The following post is partially inspired by IREE / MLIR / Linalg tutorial by B. Jacob, and Data Tiling Walkthrough by H. Wang.
The following assumes that the reader has a basic understanding of MLIR and IREE and the motivation behind them. As always, any feedback is welcome, and I will be happy to discuss any topic related to this post — Fede 🫶.
How We Compile the MLIR linalg.matmul with IREE
The linalg.matmul operation is a named, structured operation that represents the matrix multiplication of two matrices.
It takes two input matrices (%A and %B) and produces an output matrix (%C), following the mathematical definition of matrix multiplication.
Nowadays, matrix multiplication is a fundamental operation in many applications. For instance, it's at the heart of the transformer architecture, scientific computing, and many other domains.
// FILE: /tmp/matmul_500.mlir
func.func @matmul(%A: tensor<500x500xf32>,
%B: tensor<500x500xf32>,
%C: tensor<500x500xf32>) -> tensor<500x500xf32> {
%0 = linalg.matmul
ins(%A, %B : tensor<500x500xf32>, tensor<500x500xf32>)
outs(%C : tensor<500x500xf32>) -> tensor<500x500xf32>
func.return %0 : tensor<500x500xf32>
}
For sake of simplicity, we will focus on square matrices of size 500x500 and of type f32.
IREE accepts this representation expressed in the tensor dialect, and it generates the corresponding code for the target architecture.
We analyze the compilation pipeline behind the following command:
iree-compile /tmp/matmul_500.mlir \
--iree-hal-target-backends=llvm-cpu \
--iree-llvmcpu-target-cpu=apple-m4 \
--iree-opt-level=O3 \
--iree-llvmcpu-disable-distribution=true \
--iree-opt-data-tiling=true \
--iree-llvmcpu-enable-ukernels=none \
--mlir-disable-threading \
--mlir-print-ir-after-all \
--mlir-print-ir-after-change \
--mlir-print-ir-tree-dir=/tmp/matmul_500 \
-o /tmp/matmul_500/matmul_500.vmfb
Firstly, we're targeting the Apple M4 CPU (line 3) by using LLVM as the backend (line 2) with the optimization level set to O3 (line 4).
Importantly, we are disabling distribution (line 5) because I'm personally studying the performance in single-core scenarios, we are enabling data tiling (line 6) to adjust the data layout of our operands (formal parameters), and the microkernels, AKA ukernels, (line 7) are disabled because we want to analyze the generated code from scratch, without any pre-defined kernel (for more information on ukernels, see the Exploring CPU Microkernels on a Matmul Example post).
Finally, the remaining flags (lines 8-11) are used to dump the IR of the compilation pipeline—only when changes are made to the IR—in the /tmp/matmul_500 directory, which will be used for our analysis.
If you'd rather not reproduce the whole compilation, you can download the exact IR dump used in this post—the full pass-by-pass tree of /tmp/matmul_500, the standalone per-dispatch configuration files, and the input matmul_500.mlir—as a single archive: iree-matmul_500-compile-dumps.zip.
An additional useful flag is --iree-hal-dump-executable-configurations-to=/tmp/matmul_500/configs, which writes one standalone, self-contained .mlir file per hal.executable—one per dispatch—capturing the IR right after the codegen strategy has been selected (the translation_info and lowering_config attributes are already set) but before the actual tiling/vectorization/bufferization pipeline runs. Unlike the tree dump above, each of these files includes the full hal.executable.variant/hal.executable.export wrapping around the function, so it can be fed back into iree-opt on its own.
My version information is:
$ iree-compile --version
IREE (https://iree.dev):
IREE compiler version dev @ 442c605b3f85b0728389abb83ad842db38d5108e
LLVM version 24.0.0git
Optimized build with assertions
For readers familiar with Armv9-A and the AArch64 architecture, this post focuses on a baseline AArch64 target using Advanced SIMD (NEON), without considering SVE- or SME-specific features, which IREE nonetheless supports through dedicated flags.
From a Named Op to a Generic Op
Our tidy little linalg.matmul is already gone after the first few passes of the compilation pipeline.
If you look a couple of passes in, it has been rewritten into a linalg.generic:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/util_func_matmul/4_0_iree-global-opt-generalize-linalg-named-ops.mlir
%3 = linalg.generic {
indexing_maps = [affine_map<(d0, d1, d2) -> (d0, d2)>,
affine_map<(d0, d1, d2) -> (d2, d1)>,
affine_map<(d0, d1, d2) -> (d0, d1)>],
iterator_types = ["parallel", "parallel", "reduction"]
} ins(%0, %1 : tensor<500x500xf32>, tensor<500x500xf32>) outs(%2 : tensor<500x500xf32>) {
^bb0(%in: f32, %in_0: f32, %out: f32):
%5 = arith.mulf %in, %in_0 : f32
%6 = arith.addf %out, %5 : f32
linalg.yield %6 : f32
} -> tensor<500x500xf32>
linalg.matmul is what's called a named operation: its semantics are implicit in its name, and MLIR knows what it means without spelling anything out. linalg.generic, on the other hand, spells everything out explicitly: indexing_maps say how each operand's indices relate to the iteration space (d0, d1, d2) via affine_maps, iterator_types say which of those dimensions are parallel and which are a reduction, and the small region (^bb0) spells out, scalar by scalar, what happens at each point: multiply, then accumulate.
If you're not familiar with this notation, I recommend reading the Affine Dialect and OpenMP post by S. Diehl, the Linalg Dialect and Affine Dialect documentations, and last but not least, one of my favorite papers, Compiler transformations for high-performance computing by Bacon, Graham, and Sharp.
Anyway, why bother throwing away a perfectly good, compact name? Because almost everything the compiler does from this point on is implemented once, against this generic representation, instead of being reimplemented separately for every named operation.
This particular rewrite is done by a pass called GeneralizeLinalgNamedOpsPass.
Setting the Encoding
Before diving into the mechanics, it's worth asking a simple question: why does a matrix need to be touched at all before multiplying it? Why can't the CPU just multiply %A and %B as they are, row by row?
A modern CPU can process more than one number at a time—it has SIMD registers that crunch several numbers together in a single instruction, and a cache that rewards reading memory in small, contiguous chunks. A plain row-major matrix doesn't give the CPU either of these for free: to compute one output element you need one row of %A and one column of %B, and a column of a row-major matrix is scattered across memory, one element every 500 floats apart in our case.
IREE's answer to this is data tiling (sometimes called packing): before the actual multiplication happens, it physically rewrites each operand into small, regular blocks so that everything needed for one step of the computation is already sitting contiguous in memory. This is what --iree-opt-data-tiling=true turns on, and the mechanism that drives it, end to end, is exactly what this and the next few sections walk through: an encoding. See Data Tiling Walkthrough for a more detailed explanation.
Before any of that, though, a few smaller passes prepare the ground. Right after our matmul becomes a linalg.generic, InsertTensorBarriersPass wraps it with a pair of iree_tensor_ext.compute_barrier.start/.end markers—purely internal bookkeeping that fences off the boundaries of the computation region so that later passes (like reshape propagation) don't accidentally reach across it. They have no effect at runtime, and they're removed again once the dispatch region has settled. Then FormDispatchRegionsPass creates the very first flow.dispatch.region, wrapping just the matmul for now:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/util_func_matmul/5_2_iree-dispatch-creation-form-dispatch-regions.mlir
%1 = iree_tensor_ext.compute_barrier.start %0 : tensor<500x500xf32> -> tensor<500x500xf32>
%3 = iree_tensor_ext.compute_barrier.start %2 : tensor<500x500xf32> -> tensor<500x500xf32>
%5 = iree_tensor_ext.compute_barrier.start %4 : tensor<500x500xf32> -> tensor<500x500xf32>
%6 = flow.dispatch.region -> (tensor<500x500xf32>) {
%9 = linalg.generic {...} ins(%1, %3 : ...) outs(%5 : ...) { ... } -> tensor<500x500xf32>
flow.return %9 : tensor<500x500xf32>
}
%7 = iree_tensor_ext.compute_barrier.end %6 : tensor<500x500xf32> -> tensor<500x500xf32>
Next, AnnotateDataTilingHintsPass marks the matmul itself as a candidate for data tiling. The --iree-opt-data-tiling=true flag is what adds this pass (and the encoding steps that follow it) to the pipeline in the first place; the pass then tags the specific operations eligible for it:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/util_func_matmul/5_4_iree-dispatch-creation-annotate-data-tiling-hints.mlir
%9 = linalg.generic {...} ins(%1, %3 : ...) outs(%5 : ...) attrs = {iree.opt.data_tiling} { ... } -> tensor<500x500xf32>
Only now does SetEncodingPass attach an encoding to each operand—a label that says nothing about hardware yet, only about the operand's role in the computation. Its immediate follow-up, HoistEncodingOpsPass, then pulls those three new operations out of the way, so they're free to become their own independent dispatches later:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/7_iree-dispatch-creation-hoist-encoding-ops.mlir
#map = affine_map<(d0, d1, d2) -> (d0, d2)>
#map1 = affine_map<(d0, d1, d2) -> (d2, d1)>
#map2 = affine_map<(d0, d1, d2) -> (d0, d1)>
#encoding = #iree_encoding.encoding<operand_index = 0 : index, op_type = matmul,
element_types = [f32, f32, f32], user_indexing_maps = [#map, #map1, #map2],
iteration_sizes = [500, 500, 500]>
#encoding1 = #iree_encoding.encoding<operand_index = 1 : index, ...>
#encoding2 = #iree_encoding.encoding<operand_index = 2 : index, ...>
... (the same compute_barrier.start ops from before) ...
%6 = iree_encoding.set_encoding %1 : tensor<500x500xf32> -> tensor<500x500xf32, #encoding>
%7 = iree_encoding.set_encoding %3 : tensor<500x500xf32> -> tensor<500x500xf32, #encoding1>
%8 = iree_encoding.set_encoding %5 : tensor<500x500xf32> -> tensor<500x500xf32, #encoding2>
%9 = flow.dispatch.region -> (tensor<500x500xf32>) {
%12 = linalg.generic {...} ins(%6, %7 : ...) outs(%8 : ...) { ... } -> tensor<500x500xf32, #encoding2>
%13 = iree_encoding.unset_encoding %12 : tensor<500x500xf32, #encoding2> -> tensor<500x500xf32>
flow.return %13 : tensor<500x500xf32>
}
operand_index = 0 marks the LHS of a matmul (1 = RHS, 2 = the accumulator/output), op_type = matmul says which named operation this used to be, and user_indexing_maps/iteration_sizes capture the same information a linalg.generic already had.
Notice that the three iree_encoding.set_encoding ops already sit outside the flow.dispatch.region, while the matmul and the final iree_encoding.unset_encoding stay inside it.
By hoisting them out, each op can instead become its own independent dispatch, allowing the three encoding operations to run concurrently with one another—as we'll see later.
Formation of the First Dispatch Region, and Workgroups
A dispatch is IREE's basic unit of compiled, independently-invocable work—the closest analogy is a single kernel launch on a GPU, generalized to also run on CPU. One more small conversion happens first: ConvertEncodingToFlowPass converts the encoding ops that sit outside a flow.dispatch.region—exactly the iree_encoding.set_encoding ops we just hoisted—into the flow.tensor.encode form we're about to see below (which is why the hoisting had to come first). Only then does OutlineDispatchRegionsPass run, and only one dispatch exists so far—the matmul itself, pulled out into its own compiled unit, @matmul_dispatch_0:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/10_iree-flow-outline-dispatch-regions.mlir
flow.executable private @matmul_dispatch_0 {
flow.executable.export public @matmul_dispatch_0 workgroups() -> (index, index, index) {
%x, %y, %z = iree_tensor_ext.dispatch.workgroup_count_from_slice()
flow.return %x, %y, %z : index, index, index
}
builtin.module {
func.func @matmul_dispatch_0(%arg0: ..., %arg1: ..., %arg2: ..., %arg3: ...) {
...
%3 = linalg.generic {...} ins(%0, %1 : tensor<500x500xf32, #encoding>, tensor<500x500xf32, #encoding1>)
outs(%2 : tensor<500x500xf32, #encoding2>) { ... } -> tensor<500x500xf32, #encoding2>
%4 = iree_encoding.unset_encoding %3 : tensor<500x500xf32, #encoding2> -> tensor<500x500xf32>
...
}
}
}
util.func public @matmul(...) {
%0 = hal.tensor.import %arg0 "input0" : !hal.buffer_view -> tensor<500x500xf32>
...
%3 = flow.tensor.encode %0 : tensor<500x500xf32> -> tensor<500x500xf32, #encoding>
%4 = flow.tensor.encode %1 : tensor<500x500xf32> -> tensor<500x500xf32, #encoding1>
%5 = flow.tensor.encode %2 : tensor<500x500xf32> -> tensor<500x500xf32, #encoding2>
%6 = flow.dispatch @matmul_dispatch_0::@matmul_dispatch_0(%3, %4, %5) : ... -> tensor<500x500xf32>
...
}
Notice the three flow.tensor.encode ops sitting right in the caller, util.func @matmul—they are still plain, un-dispatched operations at this point (the same three ops introduced and hoisted in the previous section). Only the matmul's own computation was important enough to earn a dispatch of its own, straight away. We'll see why the other three catch up a bit later.
You'll also notice every dispatch declares a workgroups() -> (index, index, index) region, created a couple of passes earlier by MaterializeDefaultWorkgroupCountRegionPass. A workgroup is one independent, parallel instance of the dispatch's body—the CPU equivalent of a thread block in CUDA. This region computes how many such instances to launch before the dispatch actually runs. For now, don't worry about the details of how that count is chosen—we compiled with --iree-llvmcpu-disable-distribution=true, so it will always resolve to a trivial 1x1x1: one workgroup doing all the work.
However, see TileDispatchUsingForall.cpp for more information.
Resolving the Encoding for the Target CPU
Only after the matmul already has its own dispatch does the target-specific resolver enter the picture. A pass called SpecializeEncodingsPass takes the abstract encoding we saw earlier and, for the Apple M4 target we're compiling for, resolves it into a concrete physical layout:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/15_iree-stream-specialize-encodings.mlir
#encoding = #iree_encoding.layout<[#iree_cpu.cpu_encoding_resolver<
configuration = {encoding_info = {innerDimsPos = [0, 1], innerTileSizes = [8, 1], outerDimsPerm = [0, 1]}}>]>
#encoding1 = #iree_encoding.layout<[#iree_cpu.cpu_encoding_resolver<
configuration = {encoding_info = {innerDimsPos = [1, 0], innerTileSizes = [8, 1], outerDimsPerm = [1, 0]}}>]>
#encoding2 = #iree_encoding.layout<[#iree_cpu.cpu_encoding_resolver<
configuration = {encoding_info = {innerDimsPos = [0, 1], innerTileSizes = [8, 8], outerDimsPerm = [0, 1]}}>]>
#encoding3 = #iree_encoding.encoding<operand_index = 0 : index, op_type = matmul, ...>
#encoding4 = #iree_encoding.encoding<operand_index = 1 : index, op_type = matmul, ...>
#encoding5 = #iree_encoding.encoding<operand_index = 2 : index, op_type = matmul, ...>
Notice both layers now coexist side by side: #encoding3/#encoding4/#encoding5 are still the same abstract, target-agnostic labels from before. #encoding/#encoding1/#encoding2 are new: a #iree_cpu.cpu_encoding_resolver, specific to apple-m4, has decided the actual physical tile shape for each operand—innerTileSizes = [8, 1] for the LHS and RHS (8 rows/columns per panel), [8, 8] for the accumulator (8x8 output blocks), plus innerDimsPos/outerDimsPerm saying which axis gets tiled and in what order.
We'll come back to exactly how to read innerDimsPos, innerTileSizes, and outerDimsPerm a bit later, once we see the actual linalg.pack operation they produce.
These numbers come purely from the target's vector width and ISA—not from the 500x500 size of our problem. Change the target CPU and these tile shapes can change; change the matrix size and they won't.
The Encode Operations Become Dispatches of Their Own
Now that the tile shapes are known, a pass called MaterializeEncodingsPass gives the same treatment to the three flow.tensor.encode ops we've been carrying around since the very first section—until now, plain inline operations in the caller. Each one is promoted into its own independent dispatch, exactly like the matmul already was:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/17_iree-stream-materialize-encodings.mlir
stream.executable private @matmul_dispatch_0 {
stream.executable.export public @matmul_dispatch_0_matmul_like_500x500x500_f32 workgroups() -> (index, index, index) { ... }
builtin.module {
func.func @matmul_dispatch_0_matmul_like_500x500x500_f32( ... ) { ... }
}
}
stream.executable private @_encoding_0 {
stream.executable.export public @_encoding_0_encode_500x500xf32_to_500x500xf32 workgroups() -> (index, index, index) { ... }
builtin.module {
func.func @_encoding_0_encode_500x500xf32_to_500x500xf32(%arg0: !stream.binding, %arg1: !stream.binding) {
%c0 = arith.constant 0 : index
%0 = stream.binding.subspan %arg0[%c0] : !stream.binding -> !iree_tensor_ext.dispatch.tensor>
%1 = stream.binding.subspan %arg1[%c0] : !stream.binding -> !iree_tensor_ext.dispatch.tensor>
%2 = iree_tensor_ext.dispatch.tensor.load %0, offsets = [0, 0], sizes = [500, 500], strides = [1, 1] : !iree_tensor_ext.dispatch.tensor> -> tensor<500x500xf32>
%3 = iree_encoding.set_encoding %2 : tensor<500x500xf32> -> tensor<500x500xf32, #encoding>
iree_tensor_ext.dispatch.tensor.store %3, %1, offsets = [0, 0], sizes = [500, 500], strides = [1, 1] : tensor<500x500xf32, #encoding> -> !iree_tensor_ext.dispatch.tensor>
return
}
}
}
stream.executable private @_encoding_1 { ... }
stream.executable private @_encoding_2 { ... }
...
%6 = stream.async.dispatch on(...) @_encoding_0::@_encoding_0_encode_500x500xf32_to_500x500xf32(...) : (!stream.resource<*>{%c1000000}) -> !stream.resource<*>{%c1008000}
%7 = stream.async.dispatch on(...) @_encoding_1::@_encoding_1_encode_500x500xf32_to_500x500xf32(...) : (!stream.resource<*>{%c1000000}) -> !stream.resource<*>{%c1008000}
%8 = stream.async.dispatch on(...) @_encoding_2::@_encoding_2_encode_500x500xf32_to_500x500xf32(...) : (!stream.resource<*>{%c1000000}) -> !stream.resource<*>{%c1016064}
%9 = stream.async.dispatch
@matmul_dispatch_0::@matmul_dispatch_0_matmul_like_500x500x500_f32(%6, %7, %8) : (...) -> !stream.resource<*>{%c1000000}
From here on, @_encoding_0, @_encoding_1, and @_encoding_2 are first-class dispatches, exactly like @matmul_dispatch_0: each will get its own lowering strategy, its own tiling, its own compiled machine code—and, since HoistEncodingOpsPass already pulled them out of the matmul's region, its own opportunity to run concurrently with the others, as we'll see later.
Our single linalg.matmul is now, quite literally, four separate compiled kernels.
Sizes Already Known: Stream Resources and Slicing
Look again at the dispatch calls above (stream.async.dispatch). Each one takes a !stream.resource as input and produces a !stream.resource as output, and the type of each resource carries an exact byte size in its type:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/17_iree-stream-materialize-encodings.mlir
%c1016064 = arith.constant 1016064 : index
%c1008000 = arith.constant 1008000 : index
%c1000000 = arith.constant 1000000 : index
%6 = stream.async.dispatch @_encoding_0::@_encoding_0_encode_500x500xf32_to_500x500xf32(
%1[%c0_0 to %c1000000 for %c1000000]) : (!stream.resource<*>{%c1000000}) -> !stream.resource<*>{%c1008000}
%7 = stream.async.dispatch @_encoding_1::@_encoding_1_encode_500x500xf32_to_500x500xf32(
%3[%c0_1 to %c1000000 for %c1000000]) : (!stream.resource<*>{%c1000000}) -> !stream.resource<*>{%c1008000}
%8 = stream.async.dispatch @_encoding_2::@_encoding_2_encode_500x500xf32_to_500x500xf32(
%5[%c0_2 to %c1000000 for %c1000000]) : (!stream.resource<*>{%c1000000}) -> !stream.resource<*>{%c1016064}
%9 = stream.async.dispatch @matmul_dispatch_0::@matmul_dispatch_0_matmul_like_500x500x500_f32(
%6[%c0 to %c1008000 for %c1008000],
%7[%c0 to %c1008000 for %c1008000],
%8[%c0 to %c1016064 for %c1016064]) : (...) -> !stream.resource<*>{%c1000000}
The bracket syntax %resource[%offset to %end for %length] is a slice of a !stream.resource—IREE's underlying, untyped byte buffer abstraction. Note that the %resource is different for each dispatch (%1, %3, and %5).
Dispatches don't necessarily each get a fresh allocation; they read and write slices of resources that the stream layer manages and can reuse, and every value carries its exact byte size right in its type: !stream.resource<*>{%c1008000} means a resource of exactly 1.008.000 bytes.
Those numbers aren't arbitrary, and tracing exactly where they first appear reveals one more step we haven't covered yet. SpecializeEncodingsPass, from the previous section, only resolves the encoding on the types—no byte counts yet. It's the very next pass, EncodeHostTensorsPass, that computes these byte counts for the first time and produces a single, abstract stream.tensor.encode operation per operand:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/util_func_matmul/15_12_iree-stream-encode-host-tensors.mlir
%6 = stream.tensor.encode on(...) %1 : tensor<500x500xf32> in !stream.resource<*>{%c1000000}
-> tensor<500x500xf32, #iree_encoding.layout<...innerTileSizes = [8, 1]...>> in !stream.resource<*>{%c1008000}
Only afterwards does MaterializeEncodingsPass turn that single operation into a full dispatch, and the byte counts are carried forward into the dispatch's resource types.
Worth checking these numbers by hand, since they already tell us most of the story we're about to see in detail:
1.000.000bytes =500 × 500 × 4—a plain, unpackedf32buffer.1.008.000bytes =63 × 500 × 8 × 1 × 4—exactly the packed LHS/RHS shape, padding included.1.016.064bytes =63 × 63 × 8 × 8 × 4—exactly the packed accumulator shape.
In other words: by the time any dispatch exists at all, the compiler already knows exactly how big the padded, tiled buffers will be.
It goes further, actually: a pass called LayoutSlicesPass decides that the three packed buffers don't each need their own allocation—they can share one physical buffer of 3.032.064 bytes (1.008.000 + 1.008.000 + 1.016.064), with each dispatch simply writing to its own offset within it:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/util_func_matmul/23_17_iree-stream-layout-slices.mlir
stream.cmd.dispatch @_encoding_0::@..._0 { wo %arg7[%c0 for %c1008000] : ... } // LHS: [0, 1.008.000)
stream.cmd.dispatch @_encoding_1::@..._1 { wo %arg7[%c1008000 for %c1008000] : ... } // RHS: [1.008.000, 2.016.000)
stream.cmd.dispatch @_encoding_2::@..._2 { wo %arg7[%c2016000 for %c1016064] : ... } // Acc: [2.016.000, 3.032.064)
One buffer, three tenants, laid out back to back—one fewer allocation for the runtime to manage. What's still missing is the actual operation that turns 1.000.000 bytes of plain data into 1.008.000 bytes of tiled data. That's exactly what we look at next.
Concurrency Between Dispatches
Right after the four dispatches exist, but well before any of their internal codegen even starts, the compiler decides which of them can run at the same time. ScheduleExecutionPass first wraps everything—all four dispatches—inside a single stream.async.execute region, one after another, with no notion yet of which ones are independent:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/util_func_matmul/19_14_iree-stream-schedule-execution.mlir
%results, %result_timepoint = stream.async.execute on(...) with(...) -> !stream.resource<external>{%c1000000} {
%5 = stream.async.dispatch @_encoding_0::@_encoding_0_encode_500x500xf32_to_500x500xf32(...) : ... -> !stream.resource<transient>{%c1008000}
%6 = stream.async.dispatch @_encoding_1::@_encoding_1_encode_500x500xf32_to_500x500xf32(...) : ... -> !stream.resource<transient>{%c1008000}
%7 = stream.async.dispatch @_encoding_2::@_encoding_2_encode_500x500xf32_to_500x500xf32(...) : ... -> !stream.resource<transient>{%c1016064}
%8 = stream.async.dispatch @matmul_dispatch_0::@matmul_dispatch_0_matmul_like_500x500x500_f32(%5, %6, %7) : ... -> !stream.resource<external>{%c1000000}
}
The very next pass, ScheduleConcurrencyPass, looks at the data dependencies between those four dispatches and notices something we already knew from way back: the three _encoding_N dispatches don't read each other's output—each only reads one of the three original operands—while matmul_dispatch_0 needs all three of their results. So it wraps the three independent ones in a stream.async.concurrent block, and leaves the matmul dispatch outside it, after:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/util_func_matmul/19_15_iree-stream-schedule-concurrency.mlir
%results, %result_timepoint = stream.async.execute on(...) with(...) -> !stream.resource<external>{%c1000000} {
%5:3 = stream.async.concurrent with(...) -> (!stream.resource<transient>{%c1008000}, !stream.resource<transient>{%c1008000}, !stream.resource<transient>{%c1016064}) {
%7 = stream.async.dispatch @_encoding_0::@_encoding_0_encode_500x500xf32_to_500x500xf32(...) : ... -> !stream.resource<transient>{%c1008000}
%8 = stream.async.dispatch @_encoding_1::@_encoding_1_encode_500x500xf32_to_500x500xf32(...) : ... -> !stream.resource<transient>{%c1008000}
%9 = stream.async.dispatch @_encoding_2::@_encoding_2_encode_500x500xf32_to_500x500xf32(...) : ... -> !stream.resource<transient>{%c1016064}
stream.yield %7, %8, %9 : ...
}
%6 = stream.async.dispatch @matmul_dispatch_0::@matmul_dispatch_0_matmul_like_500x500x500_f32(%5#0, %5#1, %5#2) : ... -> !stream.resource<external>{%c1000000}
stream.yield %6 : !stream.resource<external>{%c1000000}
}
This is decided at the whole-module level, entirely independently of anything we've seen so far about tiling or vectorization inside a single dispatch—in fact it happens well before any per-dispatch codegen even begins. It's also worth noting what this concurrency is not: it says nothing about which CPU core runs what, or how many threads are involved. That's a separate, later decision (the workgroup distribution we touched on earlier), about splitting the work inside one dispatch across a hardware grid. This is a different, coarser axis: whether the runtime can even start issuing multiple dispatches before earlier ones finish.
A little further down the pipeline, ScheduleAllocationPass (iree-stream-schedule-allocation) inherits this same grouping and converts it into the buffer-allocated form, stream.cmd.concurrent—same three dispatches, same structure, just with real memory behind it. And since all three buffers are alive at the same time—written inside that concurrent block, and only released once the matmul has read them—this pass can do something clever: instead of three separate allocations, it issues a single stream.resource.pack—a genuine bin-packing request: "here are three slices, these sizes, all alive during the same window, figure out where each one goes":
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/23_iree-stream-schedule-allocation.mlir
%3:4 = stream.resource.pack on(...) slices({
[0, 1] = %c1008000,
[0, 1] = %c1008000,
[0, 1] = %c1016064
}) : index
// %3#0 = total size (3.032.064); %3#1, %3#2, %3#3 = one offset per slice, still symbolic here
LayoutSlicesPass then resolves those symbolic offsets into concrete constants:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/util_func_matmul/23_17_iree-stream-layout-slices.mlir
stream.cmd.dispatch @_encoding_0::@..._0 { wo %arg7[%c0 for %c1008000] : ... } // LHS: [0, 1.008.000)
stream.cmd.dispatch @_encoding_1::@..._1 { wo %arg7[%c1008000 for %c1008000] : ... } // RHS: [1.008.000, 2.016.000)
stream.cmd.dispatch @_encoding_2::@..._2 { wo %arg7[%c2016000 for %c1016064] : ... } // Acc: [2.016.000, 3.032.064)
One buffer, three tenants, laid out back to back—one fewer allocation for the runtime to manage. It works because the three slices are all alive over the same window, so none of them can reuse another's space, and they can simply sit side by side.
Annotating Dispatch Arguments
Remember the offset arguments FuseDispatchBindingsPass introduced back when the three encode operations became dispatches of their own? A pass called AnnotateDispatchArgumentsPass (iree-stream-annotate-dispatch-arguments) now records two facts about them that the compiler has just become certain of, thanks to LayoutSlicesPass resolving their exact offsets:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/29_iree-stream-annotate-dispatch-arguments.mlir
func.func @matmul_dispatch_0_matmul_like_500x500x500_f32(
%arg0: !stream.binding {stream.alignment = 64 : index}, // the shared 3.032.064-byte input
%arg1: !stream.binding {stream.alignment = 64 : index}, // the output
%arg2: index {stream.values = [0 : index]}, // LHS offset into %arg0
%arg3: index {stream.alignment = 128 : index, stream.values = [1008000 : index]}, // RHS offset into %arg0
%arg4: index {stream.alignment = 256 : index, stream.values = [2016000 : index]}, // accumulator offset into %arg0
%arg5: index {stream.values = [0 : index]}) { // output offset into %arg1
stream.values records that a parameter isn't just some index passed at runtime—it will always be exactly this one, every single time this dispatch is called: the RHS offset will always be 1.008.000, the accumulator's always 2.016.000—the same numbers LayoutSlicesPass computed one section ago. stream.alignment records the strongest alignment the compiler can prove for that value: 1.008.000 happens to be a multiple of 128, 2.016.000 a multiple of 256. These are facts the compiler now carries along for later passes to use; the immediate payoff is the one we're about to see.
This is a small, quiet pass—it doesn't change what the program computes, or even the shape of its arguments. It just writes down, in the type system, things the compiler now happens to know for certain. Those facts don't go to waste, as we're about to see.
Packing Dispatch Operands for the ABI
Right after those annotations, PackDispatchOperandsPass (iree-stream-pack-dispatch-operands) does something that has nothing to do with our matmul specifically, and everything to do with how IREE talks to any compute target: it packs every dispatch operand into i32-sized push constants. The pass's own comments give the reason: push constants are a very scarce resource (think of at most about 32 i32 values in total), so every operand is expressed in i32 words, and a wider value eats proportionally more of that budget. It isn't a GPU-only concern, either: the pass runs here too, on a CPU target. On our target index is 64 bits wide (a per-target setting), so every one of our offset arguments gets split into two i32 halves:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/31_iree-stream-pack-dispatch-operands.mlir
func.func @matmul_dispatch_0_matmul_like_500x500x500_f32(
%arg0: !stream.binding, %arg1: !stream.binding,
%arg2: i32, %arg3: i32, %arg4: i32, %arg5: i32, %arg6: i32, %arg7: i32, %arg8: i32, %arg9: i32) {
%0 = arith.extui %arg2 : i32 to i64 // low 32 bits
%1 = arith.extui %arg3 : i32 to i64 // high 32 bits
%2 = arith.shli %1, %c32_i64 : i64 // shift the high half into place
%3 = arith.ori %0, %2 : i64 // OR them back together
%4 = arith.index_castui %3 {stream.values = [0 : index]} : i64 to index
...
Every 64-bit offset becomes two 32-bit halves at the call site, and gets stitched back together—extui, shli, ori—right at the top of the callee. Four index arguments become eight i32 ones.
This sets up a satisfying payoff. Remember stream.values from the previous section—the compiler's proof that some of these arguments are always the same constant? FoldUniformOperandsPass (iree-stream-fold-uniform-operands) cashes it in: if an argument is provably uniform across every call site, there's no reason to pass it as an argument at all. It becomes a literal constant baked directly into the function body, and disappears from both the signature and the call site:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/34_iree-stream-fold-uniform-operands.mlir
func.func @matmul_dispatch_0_matmul_like_500x500x500_f32(%arg0: !stream.binding {stream.alignment = 64 : index}, %arg1: !stream.binding {stream.alignment = 64 : index}) {
%c0_i32 = arith.constant 0 : i32
%c1008000_i32 = arith.constant 1008000 : i32
%c2016000_i32 = arith.constant 2016000 : i32
...
}
...
stream.cmd.dispatch @matmul_dispatch_0::@matmul_dispatch_0_matmul_like_500x500x500_f32 {
...
}
Eight i32 arguments become zero. This is the full arc, start to finish: FuseDispatchBindingsPass turned per-operand buffers into a shared buffer plus offsets; AnnotateDispatchArgumentsPass proved those offsets were always the same value; PackDispatchOperandsPass reshaped them for the ABI; and now FoldUniformOperandsPass removes them entirely, because a value the compiler can prove ahead of time doesn't need to travel through an argument list at all.
The Gateway to Codegen
Between what we just saw and the actual per-executable code generation, there's a little more whole-module bookkeeping (folding globals, mostly—generic cleanup, nothing specific to our matmul) that doesn't earn its own section. The pass that actually matters is MaterializeInterfacesPass (iree-hal-materialize-interfaces). This is where each stream.executable we've been looking at finally becomes a proper hal.executable, with a concrete hal.executable.variant for our target:
// FILE: /tmp/matmul_500/builtin_module_no-symbol-name/37_iree-hal-materialize-interfaces.mlir
#pipeline_layout = #hal.pipeline.layout<
bindings = [
#hal.pipeline.binding<storage_buffer, "ReadOnly|Indirect">,
#hal.pipeline.binding<storage_buffer, Indirect>],
flags = Indirect>
hal.executable private @_encoding_0 {
hal.executable.variant public @embedded_elf_arm_64 target(#executable_target_embedded_elf_arm_64) {
...
}
}
This is the gateway: from here on, most of the passes we look at operate on one executable at a time—the dump tree reflects that, with everything nested inside its own hal_executable_.../hal_executable_variant_embedded_elf_arm_64/... subdirectory. This is where data tiling (the linalg.pack/unpack/mmt4d we've been promising since the very first section), loop tiling, workgroup distribution, and vectorization all happen. A few whole-module passes are still interleaved right alongside it, though—iree-hal-conversion is one, and it's where the hal.command_buffer.dispatch calls we'll see much later actually get created.
It doesn't stay per-executable forever, either: once each executable's codegen is done, the pipeline returns to the whole-module level to link all four into one binary, memoize the command buffer, and set up device/resource caches—we'll get there later.