Intro
I’ve recently taken up an interest in old console modding, so I decided to take my favorite console, the Xbox 360, and make the ultimate modchip for it. The idea of the modchip is two-fold, one of the main things it should do is have the ability to switch NANDs on the console, to be able to boot both the development OS and the retail version. This is simple enough to do, just add a transistor between the CE pins of the original chip and the second one, and you’re done. But there’s also another thing I would like it to do, that is flashing the actual NAND chips without having an external flasher, which is also the most interesting part of this project.
Flashing protocol
Naturally, the first step to flash anything is to identify which protocol it speaks, and how to properly use it. Thankfully this part has been figured out for quite a long time on the Xbox 360 as the modding scene is very mature. The console must have some way to apply firmware updates right? That’s the job of the SMC controller which communicates with the flash chip. The communication commands are also well known from reverse engineering the kernel. But how would we be able to communicate over those lines? Fortunately the console also needs to be flashed at the factory and conveniently the SMC exposes the SPI protocol on an unpopulated debug header on the board that we can solder to.
Now that we know where to send the data, let’s see what do we actually need to send in order for the SMC to understand us. The commands themselves effectively just read/write registers on the SMC so their format is quite simple:
- The read register command is
[u8(regnum), u8(0xff), u32_be(0x00)]where the last0x00u32is used to drive the SPI bus so that we can read the response on the MISO SPI pin - The write register command is
[u8(regnum), u32_be(value)]
Knowing the SMC communication protocol, we can derive an algorithm for flashing the NAND chip:
- Switch the console into debug mode by toggling the
RST_XDKpin to stop the SMC - Get the NAND flash config to know the block size for erasing
- Send erase commands every N blocks
- Send write+commit commands for each block
Who needs the SPI peripheral?
There are a couple ways of actually implementing the algorithm. The RP2040 integrates an SPI peripheral that we could use for communication; this however takes away some flexibility of the GPIO pins we could choose for outputs. So, there’s a second solution which is unique to the RP2xxx series of MCUs: the PIO peripheral. This peripheral is effectively a very constrained programmable I/O state machine that we can execute arbitrary code on. Sounds very interesting to mess around with and could potentially result in great performance and efficiency gains, let’s mess around with that!
The PIO peripheral on the RP2040 consists of 2 PIO blocks which themselves contain 4 state machines inside of them. Each PIO block has 32 instruction slots with fixed-length instruction encoding, this should be enough for everything we want to do. The PIOs also have a hidden trick up their sleeve, where you can stream instructions to them in the same stream as you do data. This could really help with extending the execution capacity if done right, we will not be using this, but it’s a very neat feature that’s worth mentioning.
Implementing basic SPI communication is trivial using PIOs, as the protocol itself is very simple: write a bit on the rising edge of the clock and read the data back on the falling edge. Optimizing this for fast block reading and writing is a more interesting challenge. This could be solved by creating a large number of buffer commands with the data we want to write, but that would be more expensive in terms of RAM usage, and in this scenario I wanted to go all in on speed and efficiency.
In this project I will be taking a very unexpected inspiration from modern graphics APIs, specifically command buffers and recording them, as we face the same issue here: figuring out what to do (what to write to which registers) is slower than actually doing it. As an additional benefit of doing command buffer-style communication quite a bit of RAM is saved due to padding bytes not being required. Padding would instead be synthesized on the fly, by the PIO state machine.
PIO commands
NAND R/W operations through the SMC require writing two registers; the first one for specifying the operation done, and the second one for committing it and optionally reading the response back. As a result of this, the algorithm that the PIOs would need to implement is as follows:
- Write the first register index
- Write the first register value
- Write the second register index
- If reading: output read marker and drive the SPI for the next 4 bytes streaming data in
- If writing: write the second register’s value
We could then construct a struct, which will be provided from the CPU:
struct LbaCommand {
first_reg: u8,
first_reg_value: Unalign<u32>,
second_reg: u8,
output_read_marker: u8,
second_reg_value: Option<u32>,
}
When implementing this algorithm and looking at the data actually being written to the SMC I noticed that the upper 24-bits of the second register are always 0. This is because it stores a boolean value, we could make an optimization that turns it into a u8, then pads the remaining space with zeroes on the state machine, saving us 1.5KiB per block.
Why not synthesize everything except for the actual block data on the PIOs? Ideally we would do this, but we quickly hit the 32 instruction slots limit per PIO block. We briefly explored doing ping-pong between two PIO blocks with IRQs effectively extending the instruction slots to 64. Unfortunately this feature is only available on the RP2350 which requires a more complicated power circuitry setup while not giving us any speed improvement deeming it unnecessary.
How rust helps!
As a result of shifting block writing to the PIOs, the writes will go through two different peripherals that cannot be active at the same time due to GPIO conflicts where the same pins are used. One obvious solution which could be used in a lot of other programming languages is to have a function that switches the internal state. The state change then affects how certain functions work:
impl SpiexState {
fn switch_mode(&mut self, to: SpiexMode) { ... }
fn read_reg(&mut self, reg: u8) -> Result<u32> {
if !self.is_spi() {
return Err(...);
}
...
}
fn read_lba(&mut self, data: &mut [u8; LBA_SIZE]) -> Result<()> {
if !self.is_lba() {
return Err(...);
}
...
}
}
This technique prevents the flasher from communicating when it shouldn’t, it does not prevent the programmer from introducing a bug where the flasher doesn’t work. Solving this can be done with type-state programming. After all, the second best time to catch errors is at compile time. To do this, the communication struct will have a generic that’s used to expose only the functions that are allowed to be called in the current mode:
impl SpiexState<SpiMode> {
fn read_reg(&mut self, reg: u8) -> u32 { ... }
fn write_reg(&mut self, reg: u8, value: u32) { ... }
fn into_lba(self) -> SpiexState<LbaMode> {
self.spi.disable();
self.pio.enable();
SpiexState { ... }
}
}
impl SpiexState<LbaMode> {
fn read_lba(&mut self, data: &mut [u8; LBA_SIZE]) { ... }
fn write_lba(&mut self, data: &mut LbaWriteData) { ... }
fn into_spi(self) -> SpiexState<SpiMode> {
self.pio.disable();
self.spi.enable();
SpiexState { ... }
}
}
This will ensure that users of the API will never call a function that’s invalid for the currently enabled mode, and will not introduce any unnecessary runtime checks that would increase code size.
USB peripheral
Having figured out Xbox communication, we shift focus onto communication with the computer that the flasher is connected to. This would of course be done using the USB protocol as its the most versatile and available one for external devices. And here’s where we hit our first problem, the NAND chip is typically 16MiB, but the RP2040 can only reach USB 2.0 FS speeds (12 Mbit/s) theoretically, in practice it’s usually more like ~9 Mbit/s due to USB overhead. This is our primary bottleneck and ideally we would reduce its effect as much as possible.
Besides just being fast, the communication protocol has to be resilient to host failures or intermittent USB cable disconnection, as well as being flexible enough to transfer arbitrary structs as commands. Something like this already exists if you are in an std environment: tokio_util::codec, this is quite a nice solution for command transfers as it does packet delimiting as well as serialization all in one nice to use package. Yet another issue that this solution solves for us is USB packet size limits. When doing USB FS communication you can only send 64-byte packets, and if your data is larger than that you would have to read it in until you have received the full data frame.
Verbatim implementations of this are of course not going to be performant on embedded for at least one reason: dynamic memory allocation. This introduces much unpredictability into an embedded system and is usually not advisable on low-power MCUs. Thankfully in our case we already know all packet sizes ahead of time, so we can reserve a fixed-size stack buffer for received data and use that without having to worry about resizing, our packet sizes are also small enough to always fully fit in memory allowing the buffers to be constant size.
Why is it slow?
Running this however, we see quite disappointing results:
Writing took 53002221us (53s)
With the expectation of write times being somewhere between 35-40 seconds, this seems very strange. What could be going wrong? Profiling embedded applications is usually quite a bad experience, state of the art tools are unwieldy and getting them to work on hardware that they don’t explicitly support feels like a sisyphean task. Can we do something better?
Of course! Because I do quite a lot of performance profiling for game engines, I am quite familiar with the Tracy frame profiler, which is a nanosecond precision profiler. Our goal is conceptually simple, monitor async task execution times and display them to the developer for further analysis. We will be querying the embassy executor, conveniently it already provides hooks for tracing task execution:
#[unsafe(no_mangle)]
unsafe extern "Rust" fn _embassy_trace_task_exec_begin(executor_id: u32, task_id: u32) {
let time_ticks = Instant::now().as_ticks()
write_event(Event {
ty: EventType::TaskExecBegin {
executor_id,
id: task_id,
},
time_ticks,
});
}
#[unsafe(no_mangle)]
unsafe extern "Rust" fn _embassy_trace_task_exec_end(executor_id: u32, task_id: u32) {
let time_ticks = Instant::now().as_ticks();
write_event(Event {
ty: EventType::TaskExecEnd {
executor_id,
id: task_id,
},
time_ticks,
});
}
Getting this data off device and into tracy will take some care as we can’t just send it over the same USB bus. It’s critical we keep overhead minimal while doing so. Serial was chosen as a simple initial method of transfer until there is data showing that it became a bottleneck. The profiler itself should not take up more hardware resources than it needs. Running it on a different core to isolate from the executor is not an option. Another downside, is that we lose the ability to, somewhat tautologically, profile the profiler. Our best bet is to run the profiler as an async task. Traces can also come in before the task is started, as well as coming in faster than the task has a chance to run.
Buffering will solve both of those issues as long as we can guarantee that the send task is executed at some regular rate to not lose profiling data. A ringbuffer is a perfect candidate for the backing structure allowing us to read and write data without having to reorder it in memory.
Ensuring that the task runs at a certain time interval is challenging in a cooperative environment, but we only need to guarantee that the task runs at least once between other tasks, regardless of when this actually happens in time. This is because trace events, and therefore, writes to the buffer, are only produced when scheduling happens, which in turn only happens when tasks are switching states. Fortunately the embassy executor is fair, guaranteeing that the data sending task will have a chance to run.
After transferring the data to the host and importing it into tracy, it’s quite easy to spot the fact that the USB processing task is taking up quite a bit of time:

What could be going wrong? After adding some more trace logging into the USB processing function, it was quickly apparent that we are leaving quite a bit of performance on the table. Why? Trying to send data without serialization resulted in much better performance. The cost of serialization, while not too high for individual writes, accumulates into quite a big performance penalty over 32k blocks and therefore needs to be eliminated. This is quite an easy task as we can just get the underlying codec data stream, that doesn’t apply serialization to structs, and use that as “downgrade” infrastructure to transfer blocks much faster.
Why not downgrade to before the codec is even applied? We still have to be resilient to host application failures and the codec infrastructure lets us achieve this goal. After applying this optimization, we can see a very big decrease in write times down to 41(!) seconds:
Writing took 41001208us (41s)
And the profiling timings are much nicer!

Run from ram!
But that’s not all the optimizations that we can apply here, the RP2040 chip doesn’t have an internal ROM for code storage and instead uses an external QSPI connected flash chip for storing code. This, of course, comes with a performance penalty especially for larger code that’s frequently executing. In the C++ SDK there is an option to put individual functions into RAM, but the Rust SDK doesn’t offer us such flexibility currently. Not all is lost however, the full size of our binary fits into the Pico’s RAM allowing the code to be linked in a way that can be copied to RAM and then executed. This is achieved with a linker script which links our code to expect execution in the RAM address space rather than external flash. It’s quite easy to derive it from the original linker script. We need to tell the linker to expect the section to be positioned at an address in RAM, while still being put into flash memory as we still need to copy it from there:
- } > FLASH
+ } > RAM AT > FLASH
We still need to copy the data from flash, but how would we do that if all code that we write expects to execute from RAM already? The boot process of the RP2040 goes as follows:
first stage bootloader (burned into the MCU)
|
| jumps to
\|/
second stage bootloader (BOOT2, flashed to the external QSPI flash)
|
| jumps to
\|/
flash/RAM depending on the implementation
We can use the second-stage bootloader to do exactly this! While the space for the bootloader is small (only 256 bytes), this is plenty of space however as the only things that the bootloader really needs to do is to just start at the beginning of the flash section and copy over the sizeof(RAM) amount of bytes to RAM, then jump to the first instruction in RAM. Bootloader pseudocode:
/// Memory layout of the RP2040
///
/// -------------------- 0x10000000 (FLASH)
/// | BOOT2 |
/// -------------------- 0x10000100
/// | vector table |
/// | |
/// | CODE |
/// -------------------- 0x101F4000
///
/// ----------- 0x20000000
/// | RAM |
/// ----------- 0x20040740
let size = 0x20040740 - 0x20000000;
for i in 0..size {
let dst = 0x20000000 + i;
let src = 0x10000100 + i;
*dst = *src;
}
ram_vector_table[0]();
There’s quite a nice implementation of the stage2 bootloader that already does what we need in the rp2040-boot2 repo. Using that and the derived linker script, we see even better performance:
Writing took 34005689us (34s)

Conclusion
I’m quite happy with the results, as the theoretical maximum for reading is ~29 seconds when doing it through the SMC.
Why not interface with the chip directly for flashing? This is a valid point as we already solder to the I/O pins on the console for the dual-NAND functionality, however we did not want to make the modchip larger/more expensive and felt like 30-40 seconds is a good enough target for flashing the chip.