TL;DR — Key Takeaways

  • NVIDIA has released cuda-oxide and cutile-rs, enabling developers to write GPU kernels directly in native Rust.
  • cuda-oxide uses the familiar CUDA-style SIMT model, while cutile-rs takes a higher-level tile-based approach.
  • Rust’s ownership and borrow checker can catch memory aliasing, unsafe buffer use and other GPU programming errors at compile time.
  • cutile-rs already appears in projects including Hugging Face’s Grout inference engine and mistral.rs, while cuda-oxide remains early alpha.
  • The releases close one of the biggest remaining gaps in Rust’s growing role across the AI infrastructure stack.

Rust has been quietly taking over the systems layer under AI infrastructure for a couple of years now. NVIDIA’s own Nova Linux driver is written in it. So is the core of NVIDIA Dynamo. Inference engines, agent runtimes, and serving infrastructure keep landing on Rust for the same reason: its compiler catches memory bugs before code ever reaches production, which matters a lot when that code is managing GPU memory across thousands of concurrent threads.

But there’s always been a gap in that story. You could write your driver, your scheduler, and your serving layer in Rust, then launch a GPU kernel — and that kernel itself had to be written in CUDA C++. Rust could call the kernel. It couldn’t write it.

NVIDIA just closed that gap. The company has released two open-source projects, cuda-oxide and cutile-rs, that let developers write actual GPU kernels in native Rust and compile them straight to PTX, the same instruction set CUDA C++ targets. Neither is a wrapper. Neither is a DSL bolted onto Rust syntax. Both compile ordinary Rust code into code that runs on the GPU.

They take different routes to get there, and NVIDIA built both on purpose rather than picking a winner.

cuda-oxide targets the SIMT model most CUDA developers already know: you write what a single thread does, and the GPU launches thousands of copies of it. A custom rustc codegen backend pushes that code through Rust’s MIR, an open-source compiler framework called Pliron, and LLVM, down to PTX. It’s early alpha, needs a pinned nightly Rust toolchain, and still asks developers to think about thread indexing and shared memory the way CUDA C++ always has.

cutile-rs takes the newer tile-based approach instead. Rather than describing one thread, you describe an operation on a tile of data, and the compiler figures out how to map that tile across the GPU’s architecture. It JIT-compiles through CUDA Tile IR at launch time, runs on stable Rust with no nightly toolchain required, and is already far enough along that Hugging Face’s Grout inference engine and the mistral.rs project use it in production code today.

The part worth paying attention to isn’t which model is faster or more mature. It’s what both get from Rust’s ownership and borrow checker: they catch a whole category of GPU bugs at compile time instead of at runtime, or worse, in a customer’s production cluster. In cutile-rs, a tensor’s ownership follows it across kernel launches. Try to pass the same buffer in as both a readable input and a mutable output, and the code won’t compile — you get a borrow-checker error before you ever run it. cuda-oxide enforces something similar with a type called DisjointSlice, which grants each thread exclusive access to its own data and blocks the classic aliasing mistake where two threads quietly stomp on the same memory.

That’s a meaningfully different failure mode than what GPU programming has lived with for two decades. A race condition in CUDA C++ usually shows up as a flaky result, a wrong number three runs out of ten, or a silent corruption that nobody notices until it’s expensive. Pushing that class of bug into a compiler error, before the code ships, is the same trade Rust has been making in every other systems domain it’s colonized.

Mitch Ashley, vice president and practice lead for CIO & Technology Buyers, and Software Lifecycle Engineering at The Futurum Group, sees the timing as the real story. “Rust already won the argument for infrastructure code — drivers, schedulers, control planes,” he said. “The kernel was the last holdout because performance-critical GPU code always demanded a language with total control and zero abstraction cost. What NVIDIA is testing here is whether you can get that control without giving up the safety guarantees everywhere else in the stack. If cutile-rs holds up in production the way early adopters are reporting, that’s the more interesting signal than cuda-oxide, because it means the safety and the abstraction can coexist without a performance tax.”

Neither project is finished, and NVIDIA isn’t pretending otherwise. cuda-oxide is explicitly early alpha. cutile-rs is further along and already in production use, but its toolchain is still young. NVIDIA has also said it plans interoperability between CUDA Rust, CUDA C++, and CUDA Python, so picking Rust for a kernel today doesn’t mean cutting yourself off from the rest of the CUDA ecosystem tomorrow.

For platform and infrastructure teams already standardizing on Rust for AI systems work, this is worth a look now rather than a wait-and-see. Start with cutile-rs — it needs nothing but stable Rust and a GPU with compute capability 8.0 or newer, and the ownership model does most of the safety work without asking developers to think differently about how tiles map to hardware. cuda-oxide is the one to watch if your team needs the explicit thread-level control CUDA C++ has always offered, once it clears alpha. Either way, the direction is set: the last piece of the GPU stack that Rust couldn’t touch is now one it can.

Frequently Asked Questions

What are NVIDIA cuda-oxide and cutile-rs?
They are open-source NVIDIA projects that allow developers to write GPU kernels directly in Rust and compile them for execution on NVIDIA GPUs.
Which NVIDIA Rust GPU project is more mature?
cutile-rs is further along and already being used in production-oriented projects, while cuda-oxide is currently described as early alpha and requires a pinned nightly Rust toolchain.
Will Rust replace CUDA C++?
Not immediately. NVIDIA is positioning Rust as another option within the CUDA ecosystem and plans interoperability between CUDA Rust, CUDA C++ and CUDA Python.