When ldaxr Doesn't Work: Exclusive Accesses and Cacheability on AArch64
Introduction
I’ve now been working on floss, my own operating system (OS) from scratch for AArch64, which is really just a kernel so far, for a couple of months since my last post about Interrupts and the Generic Interrupt Controller on AArch64. Right now, floss targets QEMU’s virt and Raspberry Pi 4 (rpi4) boards, as well as real Raspberry Pi 5 (rpi5) hardware. Testing on real hardware has been rewarding and makes the project exciting. However, as we’ll cover in much detail about in this post, real hardware reveals a new dimension of possible failures/behavior that QEMU doesn’t necessarily reveal.
This post will cover my journey of implementing a basic spin lock and figuring out why executing a load exclusive raises an exception on real hardware but not in QEMU. We will cover an exciting combination of areas, which include: spin locks, exception handling, virtual memory, memory types, memory attributes, cacheability, the differences between QEMU and real hardware, and most notably, how exclusive memory operations work.
Implementing a Spin Lock
Once I got to the point of enabling and bringing up secondary cores on the CPU, it was a good time to get some form of mutual exclusion going. As a trivial example, each core prints its own id in the setup path, e.g., “Running core 0”. Without mutual exclusion, the cores will race with each other and interleave their output, producing a garbled mess as shown in the snippet below. To get legible output, we want to guard the print with some form of mutual exclusion, so that only one core will print at a time.
Running core 0
RRRununing cournnien ng 2cni
onrge 1co
re 3
To warm up on locks and synchronization, I’ve read and recapped Arm’s Implementation Software Synchronization Primitives in A64, as well as my favourite OS book Operating Systems: Three Easy Pieces. As a good starting point, I decided to implement a spin lock, which is extremely straightforward in its simplest form. I also decided it would be a good exercise to write the spin lock by hand in assembly, which is much more fun than writing it in C/C++.
My basic version uses the load exclusive and store exclusive instructions, which work in concert to achieve atomicity. The idea is that the stxr instruction will only succeed if no other writes to the same memory location were observed since the ldx[a]r instruction took place.
void SpinLock::lock() {
uint32_t lock_read;
uint32_t store_result;
asm volatile(
"1: \t\n\
ldaxr %w0, [%2] \t\n\
cbnz %w0, 1b \t\n\
mov %w0, #1 \t\n\
stxr %w1, %w0, [%2] \t\n\
cbnz %w1, 1b \t\n\
"
: "=&r"(lock_read), "=&r"(store_result) // outputs
: "r"(&_lock) // input
: "memory" // clobber
);
}
void SpinLock::unlock() {
asm volatile(
"stlr wzr, [%0]"
: // no output
: "r"(&_lock) // input
: "memory" // clobber
);
}
The lock uses a 32-bit locking word (hence the use of w registers), where a 0 indicates unlocked and 1 indicates locked. To satisfy correctness, the load exclusive is marked with acquire semantics (the “a” in ldaxr), which guarantees that all memory accesses after the load are not reordered before it. Similarly, the store in the unlock path is marked with release semantics (the “l” in stlr), which guarantees that all memory accesses before the store are not reordered after it.
Locks are usually evaluated in three categories: correctness, fairness, and performance. We’ll empirically validate correctness later, and we’ve already considered memory ordering. Fairness is a no-go, which is alright in this basic version, but a core might get starved and never get to take the lock if other cores manage to always get it first. Performance is not that good, since cores will spend cycles busy-waiting to get the lock.
One approach to address the performance issue is to use Arm’s wfe (wait for event) and sev (send event) pair, which puts the core in a low-power state waiting for an event, and wakes all cores up that were waiting. That is something for future work.
Validation
I initially tested the spin lock in the path with the garbled output shown in #Implementing a Spin Lock, to get legible output during the bringup of multiple cores. In QEMU’s virt and rpi4 boards, everything seemed to be working fine. I reran the program a couple of thousand times for good measure and never observed garbled output.
// SpinLockGuard is a RAII object that locks during
// construction and unlocks during destruction
{
SpinLockGuard guard(&_print_lock);
kprintf("Running core %d\n", cpu_id);
}
However, when I ran the same code on my rpi5 I got a mysterious exception. The output below is from my custom exception handler, where we see a Data Abort exception with the Exception Class (EC) value of 37 (0x25). According to Arm’s “Types of exception”: “A data abort exception happens as a result of a load or store instruction”.
Got an exception (from 0x84474)
- Syndrome is 0x96000410
- EC is 37 (Data Abort)
Register dump:
x0: 0x2bcf88
...
x30: 0x85554
Taking a closer look at the instruction the CPU was executing when the exception occurred (0x84474), we see that it is right on the ldaxr instruction. This aligns with the Data Abort value we also got, so something must have gone awry with the ldaxr memory access. As a sanity check, replacing the load exclusive and store exclusive with a “normal” load acquire (ldar) and a plain store (str) makes the problem go away and no exception is thrown. This tells me there is something blocking the exclusive operations from working.
0000000000084470 <_ZN13SpinLockGuardC1EP8SpinLock>:
84470: f9000001 str x1, [x0]
84474: 885ffc20 ldaxr w0, [x1] <-- The CPU was here
84478: 35ffffe0 cbnz w0, 84474
8447c: 52800020 mov w0, #0x1
84480: 88027c20 stxr w2, w0, [x1]
84484: 35ffff82 cbnz w2, 84474
84488: d65f03c0 ret
Figuring Out Why
Fortunately, since I had read Implementation Software Synchronization Primitives in A64, I remember there was a specific section on memory attributes. This section highlights the memory attributes for which it is architecturally guaranteed that atomic instructions are atomic. These are:
- Inner Shareable, Inner Write-Back, Outer Write-Back Normal memory with Read allocation hints and Write allocation hints and not transient
- Outer Shareable, Inner Write-Back, Outer Write-Back Normal memory with Read allocation hints and Write allocation hints and not transient
The page also lists memory types which might not support atomic instructions. The most notable bullet point of those is:
- Device, Non-cacheable memory, or memory that is treated as Non-cacheable, in an implementation that does support hardware cache coherency
Memory Types
The default memory type for all memory is Device memory. However, I set up and map virtual memory particularly early in floss, associating all regions of RAM with the “Normal” memory type, so that unaligned accesses can be performed (e.g., an 8-byte load from an address that is not 8-byte aligned). This is because Device memory requires alignment checking. I became aware of this when upgrading from QEMU v8.2.2 to v11.1.50, which included a (not so recent) patch to add alignment checking to device memory, which raised a rather infuriating exception that took a while to debug as well.
Normal memory is a pretty fuzzy term, but is generally all types of memory that are not Device. Normal memory can generally be cached, speculative, and reordered depending on attributes. Device memory that is accessed via memory mapped IO (MMIO) can generally not be speculative or cached like RAM.
All types of memory are defined by a set of attributes, which can be configured for individual virtual memory mappings.
Memory Attributes
Memory attributes are configured via the Memory Attribute Indirection Register (MAIR_EL1 for EL1). This is a 64-bit register that is divided into 8 different sets of attributes of 8 bits each. An entry in the translation table (page table in Linux terminology) that points to a block of physical memory (i.e., a Block or Page descriptor) has three bits that reference an index in the MAIR register, indicating the attributes for that specific memory mapping. Additionally, such an entry has bits that configure execute permissions, access permission, shareability, and more.
Relevant to this post, and specifically to exclusive memory operations, is the shareability configuration, which can be set to one of: Non-Shareable, Inner-Shareable, or Outer-Shareable. Loosely, shareability determines in which domain, and for which observers, memory operations should be observed to complete in a finite amount of time and without using explicit cache maintenance. This topic is pretty complex, and I refer you to the Arm ARM, section B 2.7.1 for more information. In floss, the Normal memory type is configured to be Inner-Shareable, which is what Arm expects operating systems to do.
Currently, floss only configures two sets of memory attributes, one for Device memory and one for Normal memory. Let’s ignore the specifics for Device and focus on Normal memory.
-
Device:
0b00000000, Device-nGnRnE memory -
Normal:
0b11111111, Normal memory, Outer Write-Back, Inner Write-Back, with Read allocation hints and Write allocation hints and Non-transient
The Write-Back property means that writes should be written to cache, and flushed to memory only when the cache-line is evicted. In summary, this means that the memory attributes I’ve configured for Normal memory has full support for caching. It also fully matches the architectural requirements for atomics to be atomic.
An alternative to Write-Back is Write-Through, which writes to both cache and memory immediately, without waiting for the cache-line to be evicted.
Global and Local Cacheability Scopes
As we’ve concluded, the memory attributes for the Normal memory type do have the cacheable Inner and Outer Write-Back property. Even so, with all the details available to us so far, the only logical explanation for why an exception is raised when executing the load exclusive is cacheability.
With the help of my LLM, I was able to figure out that there is a global setting in the System Control Register (SCTLR): the “C” bit, which controls cacheability for data accesses at EL1 and EL0 (floss runs at EL1). Querying this bit, I see that it’s set to its default value of 0 on all cores, which disables the data cache for all of them. Setting this bit to 1 makes the exception go away on the rpi5. Hoorah!
Reading up on the documentation of L1 memory system and cache behavior on the Cortex-A76 CPU (the CPU of the rpi5), there is some revealing information:
- When the data cache is disabled, instructions and operations are affected as follows:
- All load and store instructions to cacheable memory are treated as if they were Non-cacheable and are incoherent with the caches in both this core and other cores in the cluster. Software must take this into account.
Cacheability
We’ve figured out why the exception was raised on the rpi5, because SCTLR.C with a value of 0 effectively treated the memory as non-cacheable. Interestingly, if I enable the global cacheability setting for each core by setting SCTLR.C to 1, but change the MAIR attribute to something that does not have Write-Back enabled, the same exception is raised. I did a small test of different MAIR configurations for the memory being referenced by ldaxr, and this is the result when executing on the rpi5.
| Bit pattern | Description | Exception raised |
|---|---|---|
0b00110011 |
Normal memory, Outer/Inner Write-Through Transient | Yes |
0b01000100 |
Normal memory, Outer/Inner Non-cacheable | Yes |
0b01110111 |
Normal memory, Outer/Inner Write-Back Transient | No |
0b10111011 |
Normal memory, Outer/Inner Write-Through Non-transient | Yes |
0b11111111 |
Normal memory, Outer/Inner Write-Back Non-transient | No |
From this experiment, it appears as if the load exclusive instruction doesn’t work without Write-Back. In fact, the documentation of Memory attributes on the Cortex-A76 CPU says that: "… a page is cacheable only if the Inner and Outer memory attributes are Write-Back. In all other cases, all pages are downgraded to Non-cacheable Normal memory". This means that even though Write-Through technically means “write to cache”, the Cortex-A76 CPU does not in fact cache that type of memory.
From this, it appears as if the load exclusive operation only completes without an exception if cacheability is enabled, both globally with SCTLR.C set to 1, and locally with a Write-Back attribute for the translation table entry.
It’s a bit frustrating that I wasn’t able to catch this in QEMU. QEMU doesn’t have support for the rpi5 board yet, and there are likely a few differences between the actual rpi5 hardware and the available rpi4 QEMU board. Nonetheless, when digging through the QEMU source code, it turns out that cacheability is simply not being modelled at all (see references in commits fa2ef212df, b0fe242751), reinforced by the constant for SCTLR_C not being referenced anywhere at all. So, even though it runs fine in QEMU, don’t assume everything’s perfect until you test stuff out in the real world.
Exclusive Monitors
To achieve atomicity via the load exclusive and store exclusive instructions, one or more exclusive monitors are used. Each core has its own local monitor, and all cores have access to a global monitor. Each monitor has two possible states: open and exclusive. When executing a load exclusive instruction, the monitor(s) are moved to the exclusive state, and the address of the load is recorded. When executing a store exclusive instruction, the monitor(s) are checked to see if they are all in the exclusive state and if the address matches, in which case it succeeds.
More precisely, the address stored in the exclusive monitor(s) is only the upper address bits, referred to as the Exclusive Reservation Granule (ERG). Two address locations within the same ERG will appear as the same entry in the monitor(s). On a typical system, the ERG matches the size of a cache line.
Exclusive accesses to memory locations marked as Non-shareable will only access the local monitor, and memory locations that are marked shareable (either Inner or Outer) will access both the local and global monitors.
Again, Arm’s Implementation Software Synchronization Primitives in A64 highlights a very interesting aspect in its documentation of exclusive monitors:
- Although we describe the local and global monitors as being different things, they might actually share logic. For example, one implementation approach is to use a combination of the per-core local monitors and the cache coherency logic to provide global monitor functionality for cacheable locations.
This feels like the cherry on the top of this mystery. Apparently (or empirically?), the rpi5’s Cortex-A76 CPU seems to require cacheability to be enabled for a load exclusive instruction to not generate a Data Abort exception. This could very well be because it implements the global monitor via cache coherency logic, and therefore, any exclusive memory operation to a memory location marked as shareable (either Inner or Outer) that access the global monitor, will fail when cacheability is disabled.
Conclusion
Let’s recap: I wanted mutual exclusion and opted for a “hand-made” spin lock using load exclusive and store exclusive. It worked fine in QEMU, but I got an exception on rpi5 with real hardware. I then ended up scouring documentation (Arm ARM, Cortex-A76, Implementation Software Synchronization Primitives in A64) to understand what’s going on.
From this adventure, I can conclude that on this implementation and configuration of the Raspberry Pi 5 with an Arm Cortex-A76 CPU, I observe that a load exclusive operation on shareable memory (which presumably accesses the global exclusive monitor), does not work unless cacheability is enabled, both globally (SCTLR.C set to 1) and locally (memory attribute for translation table entry configured with Write-Back).
During development I iterate quickly with QEMU, but from time to time also verify on real hardware with the Raspberry Pi 5. This creates several layers, which make this a bit confusing, but very interesting at the same time. QEMU says one thing, the Arm ARM another, and then finally the implementation of the Cortex-A76 a third thing.
This is by far the most interesting bug I have encountered so far while developing floss, but likely not the last. Thank you for reading!