Modern computers are no longer built around single processors. A capable system today is a heterogeneous ensemble: CPUs, GPUs, and an expanding zoo of specialized accelerators for machine learning, signal processing, and networking. Getting good performance out of that ensemble is hard, and getting it quickly on hardware the compiler has never seen before is harder still. DARPA MOCHA—DARPA’s Machine Learning and Optimization-guided Compilers for Heterogeneous Architectures program—exists to close that gap. The Advanced Computing Lab in the SEI’s AI Division has spent the last several months considering a question that will shape the next two years of the effort: which hardware should the program’s compilers be tested against?
This post walks through how we are approaching that question. We describe what MOCHA is trying to do, the role the SEI plays, how we selected candidate hardware, the list of that hardware and how we pressure-tested each candidate, and where the final choices landed.
Automating Computer Optimization
MOCHA is a program in DARPA’s Information Processing Techniques Office, managed by Dr. Howard Shrobe. A well-known frustration drives its work: traditional compilers were not designed to generate efficient machine code for heterogeneous mixes of CPUs, GPUs, and application accelerators. To exploit a new accelerator, developers typically hand-write specialized code and rely on vendor-tuned libraries. Although that approach works, it is slow and expensive, and it quietly encourages vendor lock-in: once an application is written against a proprietary library, moving it to different hardware means rewriting it.
Extending a compiler to support a genuinely new computational element is a manual job that can only be done by compiler experts. It is time-consuming and error-prone, and it does not scale to the pace at which novel silicon is appearing. MOCHA’s hypothesis is that data-driven methods, machine learning, and advanced optimization can accelerate that process, allowing compilers to be adapted to new hardware rapidly and with minimal human intervention. A key insight is that performance models of the target hardware drive every step of compilation, and that building those models by hand is the central bottleneck. If those models can instead be generated by measuring generated code on real hardware and by mining architectural documentation, the cost of supporting a new device drops dramatically.
A guiding principle for DARPA MOCHA is ALARA—keeping human involvement As Low As Reasonably Achievable. ALARA captures MOCHA’s emphasis on both rapidly enabling compilation for novel hardware and enabling compilation across heterogeneous hardware. Speed on one new chip is not sufficient; the program cares about how little human effort it takes to span a diverse collection of computational elements at once.
The SEI’s Role
The SEI’s Advanced Computing Lab, part of the AI Division, supports the government team, comprised of the DARPA program manager and several systems engineering and technical assistants (SETAs), with a focus on test and evaluation. In practice, that means we gather and assess the options and provide the program manager with the information he needs to decide which computational elements enter the program. We then stand up and maintain the evaluation machine where those elements are integrated, and we build the measurement methodology for MOCHA to compare performer results fairly. Performer teams develop the compiler technology; our job is to give them a well-characterized, representative, and appropriately challenging set of targets to aim at, and to maximize validity of the evaluation itself.
DARPA plans for MOCHA to include six distinct computing types by the end of the effort. A computing type is defined not just by a hardware architecture but by a distinct instruction set and programming model. Under that definition, a data-center GPU, a device that fuses a field programmable gate array (FPGA) fabric with a spatial AI-engine array, a long-vector processor, and a RISC-V-plus-dataflow AI accelerator are four different types, even though a casual observer might lump the last three together as “accelerators.” The point of the program is to demonstrate rapid, low-effort retargeting across architectural and instruction set architecture (ISA) boundaries, so architectural diversity in the target set is essential.
There is also a concrete constraint: any hardware chosen must physically fit inside the evaluation machine. That machine is a workstation-class tower built around an Intel Core Ultra 9 285K, which brings its own compute types: AVX2 SIMD on the CPU, an integrated Xe GPU, and a neural processing unit plus PCIe 5.0 connectivity and an NVIDIA RTX 4500 Ada card already installed as a baseline reference. Candidate accelerators therefore need to be available as PCIe cards that fit the chassis, power envelope, and cooling of a single tower.
5 Factors for Selecting a Candidate Accelerator
Five factors shaped our candidate list: availability, maturity, affordability, programmability, and the ability to host the device in the evaluation machine. Availability and hosting knocked out otherwise fascinating options, including wafer-scale engines and reconfigurable-dataflow systems that only ship as complete servers and cloud-only accelerators you cannot buy and install. Affordability kept us honest about parts that cost more than the rest of the machine combined.
The most influential factor was programmability and, specifically, the state of compiler and multi-level intermediate representation (MLIR) support for each target. MOCHA’s performers overwhelmingly build on the LLVM and MLIR ecosystem. The key innovation of LLVM was the extensible IR and tooling for a developer to interact with it. MLIR is a more recent innovation that has extended that capability by defining IRs at different abstraction levels. It has become the connective tissue of modern compiler infrastructure, and it lets a compiler express computation at several levels of abstraction and progressively lower it toward a specific device. If a target already has an MLIR or LLVM path, a performer can plausibly reach it and then focus on the interesting research: retargetable code generation, learned cost models, and optimization selection, and eventually partitioning work across heterogeneous elements. If a target is a sealed black box reachable only through a vendor’s high-level, pre-tuned inference stack, there may be very little surface area for a MOCHA compiler to work against, no matter how much machine learning is applied.
The tension between open, low-level access versus closed vendor libraries runs straight through the program. Programming to a proprietary library is convenient, but it is fundamentally at odds with the goal of rapidly supporting new hardware because the library only exists for hardware the vendor already chose to support. As we assessed each candidate, we looked closely at how open its programming model is and whether an MLIR-based path to the metal exists or is realistically within reach.
The Candidate List
With those criteria applied, our working short list of targets spanned GPUs, spatial FPGA-plus-AI-engine devices, a vector processor, several distinct AI accelerators, and networking silicon—the raw material for six or so genuinely different computing types:
- AMD Instinct MI350P — a data-center GPU (CDNA 4) programmed through ROCm/HIP, with a mature MLIR story via rocMLIR and the Triton path. Notably, it is a newly announced PCIe form factor that brings OAM-class Instinct compute into a standard slot.
- AMD Versal ACAP — a heterogeneous device combining an FPGA fabric, a spatial AI-engine array, and ARM cores. The open mlir-aie/IRON toolchain and its Peano LLVM back end make the AI-engine array a genuinely interesting, close-to-metal MLIR target.
- Intel Data Center GPU Max 1100 — an Xe-HPC GPU programmed through oneAPI/SYCL, reachable via MLIR through the Intel Triton XPU backend and SPIR-V. It is the only Max-series part offered as a PCIe card.
- Intel Gaudi 3 — an AI accelerator with matrix and VLIW-SIMD tensor engines. Custom kernels are written in TPC-C through an open TPC-LLVM compiler, and the graph compiler builds MLIR-based fused kernels, though the graph compiler itself remains proprietary.
- Intel Xe iGPU — the integrated GPU already present in the evaluation machine’s CPU. It shares the oneAPI/SYCL surface and MLIR path with the Max 1100, making it a zero-cost portability target.
- NEC SX-Aurora TSUBASA — a classic long-vector processor on a PCIe card, with an upstream LLVM back end. It offers architectural diversity for vectorization, autotuning, and bandwidth-bound HPC kernels.
- Qualcomm Cloud AI 100 — an inference accelerator whose main path is ONNX/PyTorch, but which—contrary to its "closed" reputation—also ships an open compiler based on upstream LLVM and supports registering low-level custom kernels.
- Tenstorrent Blackhole (p100a/p150a) — an affordable, unusually open AI accelerator pairing Tensix cores with RISC-V, with a native, fully open-source MLIR compiler (tt-mlir/tt-forge) that ingests models from PyTorch, JAX, and ONNX.
- MangoBoost BoostX DPU and GPUBoost RNIC — networking-focused parts (a SmartNIC/DPU and an RDMA NIC) included for completeness, but with no general-compute MLIR or LLVM kernel path.
Asking the Performers, and the Challenge We Set for Them
Before finalizing the candidate list above, we circulated a draft to the program performer teams and asked two questions: what did we miss that should be here, and which of these are impossible for your toolchain to support? The answers were candid and useful. Teams flagged which parts had real LLVM/MLIR back ends and which did not, pushed back on targets whose value depended entirely on closed inference stacks, and told us plainly when the networking parts were not compute targets they would pursue.
Underlying the candidate identification process was the principle that targets should be hard but not impossible. A target that is too easy—one with a mature, polished, vendor-optimized stack—does not really test MOCHA’s central claim about rapid, low-effort adaptation, because the hard work has already been done by the vendor. A target that is too hard—an undocumented black box with no low-level programming surface and no way to model its microarchitecture—simply blocks progress, and performers waste time and effort fighting the tooling rather than advancing the science. The goal is a device open enough to reach and reason about, but different enough from what performers already know that retargeting genuinely exercises their compilers, cost models, and kernel generators.
Hardware Selection: AMD RDNA 4 GPUs and Tenstorrent Tensix Cores
Between drafting that candidate list and this writing, identifying hardware elements that are feasibly obtainable further narrowed the field, and the program’s first tranche came into focus around two cards: the AMD Radeon AI PRO R9700 and the Tenstorrent Blackhole p150a.
A GPU on the AMD ROCm/HIP stack was always going to anchor the set. Across the performer teams, it was the consensus first choice: a serious, non-NVIDIA GPU with a real MLIR path, through rocMLIR and Triton, and direct relevance to the tensor, sparse, graph, and cost-modeling work at the heart of several performer proposals. Our initial pick was the newly announced MI350P, the PCIe form factor of AMD’s flagship Instinct part. But engineers at AMD Research informed us that they were themselves waiting on MI350P silicon and did not expect units until the spring of 2027, well past the window we need to begin year 2 evaluation. We therefore substituted the Radeon AI PRO R9700, which is a workstation card available now. It fits in a standard PCIe slot, speaks the same ROCm/HIP programming model, and reaches the same MLIR and Triton paths. It preserves the AMD-GPU computing type we wanted while being something a performer can actually put in a machine this year.
The R9700 turns out to make an unexpectedly good MOCHA target for a reason that goes to the heart of the program. Because it is built on a new architecture (RDNA 4), AMD’s own hand-tuned assembly libraries do not yet fully cover it. Several of them carry hardcoded lists of supported architectures that simply exclude the card and silently fall back to slow paths when they encounter it. The compiler route is what works: the Triton and MLIR path just-in-time generates native kernels for the new architecture at runtime, precisely where the pre-built vendor libraries fail. That is the MOCHA thesis in miniature, namely compiler-generated code retargeting to new silicon where hand-tuned libraries cannot, and it means there is genuine, measurable performance headroom for a MOCHA compiler to capture, rather than a vendor-polished baseline that is already close to optimal.
For architectural contrast we chose the Tenstorrent Blackhole p150a. Where the R9700 is a GPU on a mature LLVM backend, the Blackhole is something genuinely different: an array of Tensix cores paired with general-purpose RISC-V cores, programmed through Tenstorrent’s fully open, MLIR-native compiler stack (tt-mlir and tt-forge), which ingests models from PyTorch, JAX, and ONNX by way of StableHLO. Its MLIR story is arguably the strongest of anything we evaluated. Blackhole is a named target of an open-source compiler whose development happens entirely in public, with a documented StableHLO entry point where a performer’s compiler can plug in. The difficulty here is not getting in the door but learning a bespoke tower of dialects rather than the familiar LLVM-target model. That is exactly the kind of retargeting challenge MOCHA wants to time and, eventually, automate.
We had hoped to include a third architectural type in this first tranche: the AMD Versal, whose AI-engine array is a spatial dataflow fabric quite unlike either a GPU or the Tensix array, and which has an attractive open MLIR toolchain in mlir-aie. In the end we could not find a Versal part that both exposes the AI engines and ships as a PCIe card that fits the evaluation machine, so the AI-engine type falls to after the program.
Together the two cards give the program complementary retargeting problems rather than redundant ones. The R9700 tests retargeting within a mature ecosystem whose libraries happen to be immature for this specific, brand-new architecture, while the Blackhole tests retargeting into an entirely new architecture class with a young but fully open stack. Combined with the compute already inside the evaluation machine, namely AVX2 SIMD, the Xe iGPU, and the NPU, alongside the NVIDIA baseline, they move the program a meaningful step toward its six-computing-type goal.
Not all three PCIe cards fit into the machine at the same time. We will decide later in the process which additional cards to include concurrently to get to the six-computing-type metric.
Next-Gen Hardware: From Months of Tuning to Days of Measurement
Selecting the first tranche of hardware is only the opening move. Several questions will follow us into the rest of the program.
A key question relates to abstraction level. For a sufficiently opaque device, we may never be able to program at the ISA level more efficiently than the vendor’s own high-level tools. Yet leaning on those proprietary tools cuts against the goal of rapidly supporting new hardware. Finding the right level to target and improving our ability to model black-box microarchitectures, will shape which future devices are worth adding.
With the first two targets chosen, the work now shifts from selection to execution: integrating the Radeon AI PRO R9700 and the Tenstorrent Blackhole p150a into the evaluation machine, characterizing them, and standing up the measurement pipeline that will let us compare what the performers’ compilers can do with them. In addition, we will be specifying workloads that can benefit from the use of six compute types, doing manual implementations for the workloads across six compute types, and then challenging the performers to automatically perform the decomposition. Where MOCHA succeeds, the payoff is a world in which adopting the next novel accelerator is a matter of days of measurement rather than months of expert hand-tuning. Getting the hardware right is the first step toward finding out.