SHA-256/SHA-384 Type Confusion OOB Write Explained: When a struct Lies About Its Age, the Kernel Believes It
From 15 out-of-bounds bytes to full root and a walk out of a chroot jail: a technical teardown of a Linux kernel exploit chain that starts with a size mix-up between SHA-256 and SHA-384 inside a custom crypto module, travels through MSG_COPY, pipe_buffer, and struct page, and ends with a data-only write into cred followed by swapping task->fs for init_fs.
Picture a clerk in the records office of some venerable institution, the kind of person who respects procedure the way other people respect gravity. Every form in this office is the same size, every form carries a little box at the bottom labeled LINE COUNT, and our clerk has exactly one job, a job simple to the point of insult: copy one form onto another, and the number of lines you copy is whatever number sits in that little box. He does not check the size of the form. He does not ask where it came from. He reads the number and he starts copying.
One day a form arrives from a different department, a department that uses bigger forms, 48 lines instead of 32. The box he thinks of as LINE COUNT sits at line 32 of his small form, but on the big forms that spot isn’t a box at all, it’s just payload. So the “line count” he reads is actually digits from the big form’s data, digits nobody put there to serve as a counter, digits that become a counter for no better reason than that he read them as one.
He starts copying, and everything is pleasantly boring until line 32. That’s where the moment worth reading this entire article for happens: his pen lands on the LINE COUNT box itself. The counter stops being a counter and becomes a value the clerk wrote with his own hand. And from the next line onward he isn’t writing on his own form at all. He hops over to the form in the next box and writes at whatever “address” the original form’s data points him to.
And who chose the original form’s data? You did. Not directly, of course. You just played the content lottery a few thousand times until the digits you wanted came up in the positions you wanted them. The lottery here is cheap, and as we’ll see shortly, it pays out instantly.
Now the part that turns a tired clerk into a security incident. What if the form in the next box is your employee ID, the one that decides who you are and what you’re allowed to touch? And what if the institution has locked you into one wing of the building, and the map in your pocket is drawn so that it only leads to that wing? A handful of bytes written in the wrong place, plus one obedient clerk, is all it takes to turn “temporary contractor” into “Director General”, and then to redraw your map into a map of the whole building.
None of this is a strained metaphor. It is, with the names changed, exactly what happened. The clerk is called copy_hash. The forms are structs living in kernel space. The ID card is called cred. The map is called fs_struct. And the whole thing ran inside an isolated virtual test environment under our full control, against a kernel module somebody wrote carefully, and that’s the worst part.
The route map
Before we dive in, here’s the chain from the first byte to the last qword, so you can come back to it whenever you feel lost:
- A tuned OOB write: 15 bytes written past the end of a buffer at an offset we choose (we call it B), inside a neighboring object we also choose.
- Heap grooming with msg_msg: tagged System V IPC messages occupy the neighboring slab, and an E2BIG oracle tells us exactly where our shot landed.
- m_ts := 255, then MSG_COPY: a 255-byte read walks past HARDENED_USERCOPY like a customs officer who’s on the phone.
- pipe_buffer spray: tiny pipes move into the neighborhood, and the ops pointer the kernel fills in itself breaks KASLR.
- kread24: writing msg->next byte by byte grants an arbitrary read, 24 bytes from any kernel address.
- A chain of reads:
__per_cpu_offset, thencurrent_task, thencred, until we’re standing on the address of our own ID. - Forging an entire pipe_buffer: page, ops, offset/len, and flags=CAN_MERGE, which turns a single write() into an arbitrary physical write into any page of RAM.
- The write into cred: 72 bytes that make uid zero and fill in every capability. No ROP, no stack pivot, not one gadget. This is called data-only, and it is the aristocracy of exploitation.
- The last eight bytes:
task->fs := init_fs, and the doors of the real filesystem swing open outside the jail.
The setup: a kernel module that manages passwords
A small kernel module, written by somebody to manage user authentication data inside kernel space instead of userspace. You’ve heard the idea before, in plenty of places: keep the verification close to “the truth”, meaning the kernel itself, so nobody can tamper with it. And what could possibly go wrong when we place variable-sized structures at the top of the system’s privilege stack? A rhetorical question, obviously, and answering it will occupy the rest of this article.
The module exposes a misc device we talk to through ioctl, with a reasonable-looking interface of four operations: register a new user, delete one, update a password, and restore the old one. Each user keeps two copies of the hash: a curr copy that does the actual work, and an old backup copy that’s supposed to save the day if an update goes wrong. And the module supports two algorithms: SHA-256 and SHA-384.
And right there is the spot where any honest reviewer would love to catch it before some hungry attacker does: the two algorithms are stored in the same type of buffer, the distinction between them lives in a type field inside a different, distant struct, and copying between the copies relies on a function that reads the “length” of the source from inside the source’s own data. If your hand didn’t just twitch, go check your calendar, because you haven’t written enough C.
The data structures, with generic names as everywhere else in this article:
typedef struct {
u8 digest[32];
size_t len; /* at offset 32 */
} sha256_hash_t;
typedef struct {
u8 digest[48];
size_t len; /* at offset 48 */
} sha384_hash_t;
/* the heart of the problem: */
static void copy_sha256_hash(sha256_hash_t *dst, sha256_hash_t *src)
{
size_t n = min_t(size_t, src->len, HASH_MAX);
int i;
memset(dst, 0, sizeof(*dst));
for (i = 0; i < n; i++)
dst->digest[dst->len++] = src->digest[i]; /* <== the bug lives here */
}
The bug: when SHA-256 meets SHA-384 in one struct
The function above looks innocent to the point of pity: copy n bytes from a source to a destination, in a byte-by-byte style from the era before memcpy. But it carries two stacked problems, and either one alone is enough for a genuinely bad day.
The first problem is that src->len is read from offset 32 of the source buffer. If the source holds a SHA-384 hash (48 bytes), then bytes 32 through 39 are the beating heart of the digest itself, so arbitrary data gets read and interpreted as a length. The resulting number is usually enormous, but min_t(..., HASH_MAX) clips it to 48, always. Savor the irony here: the “validation” mechanism clips any value to the maximum allowed, which means it guarantees the attacker that the loop will complete all forty-eight of its iterations no matter what the original number was. The guard himself swings the door wide open and signs the guest book on your behalf.
The second problem, the lethal one, is that the index and the counter in dst->digest[dst->len++] are the same field. As long as len stays below 32, life is good. But at exactly i=32 the write lands at dst+32, which is the address of the len field itself. The counter that was measuring the copy becomes an output of the copy, like a clock that sets itself by reading a message that arrived too late.
(A side note for compiler-precision enthusiasts: the ordering between the post-increment and the store when both write to the same address is a matter of code generation. In the compiled binary in front of us the named store lands last, so the counter becomes the value of the byte just written. This is not undefined behavior in its folk sense; it’s stable behavior in a fixed binary, and everything we build on top of it stands on that stability.)
Now follow the loop’s path after the counter ate itself, where D is the digest actually present in the source:
| iteration i | what actually happens |
|---|---|
| 0 to 31 | a normal copy into digest[0..31], the counter crawls up to 32 |
| 32 | the write lands on the len field itself, so len = D[32], and we call that B |
| 33 to 47 | the writes land at digest[B] through digest[B+14], that is, the bytes D[33..47] |
And the practical conclusion deserves bold letters on the whiteboard:
15 bytes (they are D[33..47]), at an offset we control (it is B = D[32]), inside a buffer requested as 48 bytes, in a slot from the kmalloc-64 class.
B is a single byte, so it can take any value from 0 to 255. For the write to leave the hash’s slot and reach the neighbor in the slab, we want B close to 64 or above. And here’s the live layout we aim at:
kmalloc-64 slot (the hash) kmalloc-64 slot (the neighbor)
+-----------------------------+ +------------------------------+
| digest[0..31] | | neighbor+0 .. +15 |
| digest[32..47] (size view!) | | neighbor+16 .. +23 |
| (rest of the 64B slot) | | neighbor+24 .. +63 |
+-----------------------------+ +------------------------------+
hash+0 hash+64
shot at B=74 writes D[33..47] at hash+74 .. hash+88
(that is neighbor+10 .. neighbor+24)
Notice what a B=74 shot does to a neighbor that happens to be an IPC message (we’ll see shortly why messages specifically): the hash+74 to hash+88 window grazes parts of the m_list.prev and m_type fields, then ends exactly on the zeroth byte of m_ts. The last byte in the window is the only byte we fully control through the brute force, and landing it precisely on m_ts is what will finance our first leak. The choice here isn’t an aesthetic preference, it’s engineering.
Who wakes the bug: the story of a lying label
So far this has all been talk of “if a buffer holding SHA-384 ever got interpreted as SHA-256”. The practical question: how do we make the module do exactly that, on demand? The answer lives in the failed update cycle, and you should read it slowly, because it’s the most beautiful part of the bug:
- We register a new user of type SHA-256 with password P1. Now
currholds a clean SHA-256,oldholds an identical copy, and both labels are telling the truth. - We request a successful update to SHA-384 with password P2. The module authenticates us, copies
currintoold(and this copy is sound, because the types match at that moment), computes the new hash intocurr, then updates the labels because everything succeeded. - We request a second update with the correct current password but an empty new password. Authentication succeeds, so the content of
curr(a full SHA-384) gets copied intoold. Then computing the new hash fails, because an empty password is rejected, so the module restorescurrfromoldand returns the error. And the labels? Never touched. Updating them was conditional on the computation succeeding.
The result after step three: old holds 48 bytes of SHA-384 content, while its label still says “I am SHA-256” from step two. Content from one world, label from another.
- We call restore to bring back the old password. The module copies
oldintocurr“based on old’s type”, meaning it feeds a 48-byte buffer into copy_sha256_hash, a function that sees the world in 32-byte units. The shot fires.
The most beautiful part of the story is that it’s the failing operation that opens the door. The successful update fixes the labels carefully; the failed one leaves the two fields living in different worlds and then apologizes on its way out. Error paths are always where the programmers relax and the attackers loiter.
Cooking passwords: brute force on a budget
All we need from password P2 is for sha384(P2) to satisfy two tiny constraints on specific bytes: d[32] = B (the jump address) and d[47] = the byte we want dropped at the end of the window. Two constraints of eight bits each, 16 bits total, which means about 65,536 expected attempts before the shot we want comes up. With an implementation running two to three million hashes per second on a single core, that’s less time than you spend restarting your browser.
And here a small confession deserves its place in the article, because it’s a standalone lesson: the first SHA-384 implementation we carried with us shipped five broken round constants, copied wrong in some ancient era. The function ran without a single error, in total silence, producing completely wrong results. And that’s the worst thing a hash can do in life: lie to you confidently. We re-derived all eighty constants from first principles (floor(frac(cbrt(prime)) * 2^64)), re-tested against reference vectors, and the lottery started paying again. The lesson: when brute force fails you even though the math is “correct”, audit your constants before you audit your logic.
The patch
The fix is smaller than the bloodshed deserves, which is precisely what makes it insulting:
static void copy_hash(void *dst, const void *src, u32 type)
{
size_t sz;
switch (type) {
case HASH_SHA256: sz = SHA256_DIGEST_SIZE; break; /* 32 */
case HASH_SHA384: sz = SHA384_DIGEST_SIZE; break; /* 48 */
default: return;
}
memcpy(dst, src, sz); /* length from the type, never from the data */
}
Three rules scream from inside these few lines. The length must come from the declared type, not from a field living inside the data itself, because data that carries its own counter isn’t data, it’s a booby trap. The label and the content must transition atomically on every path, and specifically on the failure paths, because “I am SHA-256” is a sentence that must always be true or never said at all. And the third rule, the most bitter and beautiful one: the correct alternative was memcpy from day one, a function that existed before the module’s author was born, doing the job without a counter that reads itself. The simple thing was the correct thing all along.
Taming the heap: a neighborhood of 64-byte residents
An OOB write with no context is worth nothing: 15 bytes will burn themselves out on the first object they meet, and they’ll either fall asleep without a trace or wake the kernel up against you. We want them to hit a neighbor we know by name. And the usual solution in kernel exploitation: make the neighborhood the module lives in live by your rules.
The key is that everyone lives in the same building. The module’s hash is allocated with kzalloc(48), so it lives in kmalloc-64. A System V IPC message carrying 16 bytes of data has a header called msg_msg of 48 bytes (m_list, then m_type, then m_ts, then next, then security), for a total of exactly 64 bytes, which makes it a neighbor in the same slab cache, kmalloc-64. And even a single pipe_buffer (we’ll need it shortly) is 40 bytes, which is also kmalloc-64 class. The neighborhood is homogeneous, and you’re the one arranging its residents.
The plan: we spray hundreds of tagged messages across IPC queues, each message carrying an identification mark in m_type and in its data, so that later we know exactly whom we hit. Then we run consecutive register and unregister cycles: each cycle frees the hash and carves it out of the freelist again, and SLUB hands out slots in strict LIFO style (the last freed slot is the first one granted), so with every cycle the hash moves down one seat. We repeat until it sits next to one of our messages, in a slab packed with our messages. The odds here aren’t luck, they’re boring statistics working in our favor.
But how do we know the shot landed next to a message at all, and which message, and in which queue? Here enters the smallest and most elegant tool in this article, the E2BIG oracle:
/* non-destructive scan of all messages: with MSG_COPY, msgtyp becomes
* a sequence number inside the queue, so we walk the messages one by one
* without receiving them. */
errno = 0;
ssize_t r = msgrcv(q, &small, 16, idx, MSG_COPY | IPC_NOWAIT);
/* r == 16: healthy message, a normal copy.
* errno == E2BIG: m_ts bigger than bufsz, we found our victim. */
The logic is first-layer simple: msgrcv with a buffer smaller than the message size and without MSG_NOERROR answers E2BIG, and the check runs before the message gets unlinked from the list, so the message stays in place, untouched. All our messages are 16 bytes, so any message “apologizing for its size” has certainly tasted the shot. You’re not attacking anything here. You’re knocking on doors one by one, waiting for the one who opens with teary eyes.
The clean leak: MSG_COPY magic
Now the victim is known, and its m_ts reads 255 after having been 16. The obvious move: receive it with a big buffer and enjoy 255 bytes (239 of them outside the object’s bounds). And that exact obvious move will kill the VM on the spot.
The reason is that a normal receive copies from the slab directly into userspace, and HARDENED_USERCOPY stands at exactly this line: every copy to userspace must stay inside the bounds of a documented slab object. 255 bytes out of a 64-byte object? Panic, oops, a red splash in dmesg, and a reboot. (Yes, we tried it once, and dmesg was needlessly harsh about it.)
MSG_COPY rebuilds the whole process from the roots. This flag, which exists to serve checkpoint/restore, makes msgrcv not receive the message but copy it: a fresh message gets allocated with size bufsz (255 here), then copy_msg copies from the original into the fresh one with a kernel-to-kernel memcpy (no usercopy check at all, because there is no usercopy to begin with), and then store_msg copies the correctly-sized fresh one into userspace. The customs officer sees a box of perfectly legal dimensions, while the contraband had already been placed in the box inside the warehouse itself a moment earlier. In less metaphorical terms: HARDENED_USERCOPY inspects the final source of the copy, and MSG_COPY made the final source a correctly-sized object by design, while the out-of-bounds read happened in an internal kernel stage that nobody watches.
The single condition is that the new m_ts (255) doesn’t exceed the copy’s capacity, because copy_msg refuses (EINVAL) a source bigger than its destination. And here the wisdom of B=74 shows itself again: a shot at, say, B=80 would write the whole of m_ts with a huge random value that blows past the limits, turning the message into a bomb that can never be read. B=74 touches only the zeroth byte, so m_ts becomes 0xFF, that is, 255: big enough to read 239 bytes of the neighborhood, small enough to pass the check. One tuned byte, in the tuned position, for tuned reasons.
And what’s in the neighborhood? Here the picture completes: once we know our victim, we carefully drain the other queues (surgically, without touching our victim or the message preceding it in its list, for a reason we’ll explain in a moment), which frees slots, and then we spray tiny pipes: each pipe() with F_SETPIPE_SZ(4096) carves a ring of a single slot, that is, one 40-byte pipe_buffer, that is, another kmalloc-64 that settles into the gaps. Then we re-read the victim with the same MSG_COPY.
In the stretched 255-byte window we now see: our messages’ data, m_list pointers of live messages (direct heap addresses), and between the fields, the most precious thing in the whole picture: pipe_buffer fields that the kernel filled in with its own hand, containing page (a pointer to a struct page in the vmemmap region) and ops (a pointer to the pipe operations table, a static object in .data whose exact offset from the image in front of us we know). ops minus the known offset equals kbase. KASLR fell officially, and we never fired a single usercopy shot in its war.
A small promise, now kept: why don’t we touch the message preceding the victim in the list? Because the B=74 shot scrambled the victim’s m_list.prev (that’s its side effect), and CONFIG_DEBUG_LIST validates list links on unlink, so if you unlink the neighbor with any normal receive, the validation refuses the deletion, and you’re left holding a node that is simultaneously linked and freed (free-while-linked), which some poor soul receives later and then follows an encoded freelist pointer (SLAB_FREELIST_HARDENED) as if it were a healthy next, and the oops arrives a few steps down the road. We diagnosed this crash completely inside GDB before understanding it, and it’s the kind that teaches you to respect linked lists forever. MSG_COPY once again is the answer: it never unlinks the message from the list, not on success and not on failure. Our victim stays in its queue like a piece of furniture, read but never touched.
kread24: turning msg->next into a ticket to any address
The leak gave us kbase and heap addresses, but it’s a passive leak: we read whatever happens to sit nearby, not what we actually want. The required upgrade is an arbitrary read: 24 bytes from any kernel address we choose. And this time the key isn’t m_ts but the field right after it, msg->next.
The idea rests on how big IPC messages (bigger than 4048 bytes, which is DATALEN_MSG, a page minus the header) are stored as a chain of segments: the first message carries 4048 bytes, and the rest goes into a fresh msg_msgseg that the next field points to, and each segment’s data begins right after its pointer. When copy_msg copies a message shaped like this, it walks the chain and copies from each segment according to its share, all of it with kernel-to-kernel memcpy.
Now arrange the pieces: we write into the victim m_ts := 4072 (that is, 4048 plus 24) and next := X, where X is an address we choose. On an MSG_COPY with a big buffer, copy_msg will copy the first 4048-plus-24 bytes of the victim’s data (out of bounds, but internally and safely), then walk to the “segment” at X and copy 24 bytes from X+8. The segment is fake, nobody lives there, but the function doesn’t know and doesn’t care: it copies whatever m_ts and next tell it to. The single condition is that the first qword at X itself must be zero, because the function will read it as the fake segment’s ->next and should stop there rather than wander off to a random address. Hence our permanent rule: X always equals “the wanted address minus 8”, and we know in advance (from the image or from a previous read) that the qword before the target is zero.
24 bytes isn’t an arbitrary number: it’s exactly three pointers, and that’s enough to read three lines of the kernel’s map in one shot. And the first shot lands on an address we know from the image before any extra leak, collecting in a single strike: vmemmap_base, vmalloc_base, and page_offset_base. The first and third regions are the heart of the matter: a modern kernel with KASLR doesn’t hide just the base of the text, it also hides the base of the direct map (where physical memory lives at its virtual addresses) and the base of vmemmap (where the struct page descriptors of every page of RAM in the system live). Without those two values we can’t translate any virtual address into a physical page later on, and the entire final-writing project stalls at exactly this sentence.
Then the hunting chain, three consecutive reads, each one eating its 24-byte loaf:
kread24(kbase + OFF_PER_CPU_OFFSET - 8) --> __per_cpu_offset[0]
kread24(per_cpu_off + 0x1ed00 - 8) --> current_task (us, personally)
kread24(current_task + 0x698 - 8) --> cred, real_cred
The first read descends from a constant in .data into the per-CPU array. The second goes to the current_task slot inside the region of the CPU we’re pinned to (yes, we pinned the process to a single CPU from the start with sched_setaffinity, because this whole construction, from freelist to per-CPU, unravels if the process dances between processors). The third opens our own task_struct and pulls out the cred and real_cred pointers, and we verify that they’re identical (an internal sanity check for reliability: if the two differ, you’re reading something else, not a healthy task). The offsets 0x698 and 0x6e0 and their neighbors are specific to this build; we derived them from the disassembly. And the number that should matter to you: each of those reads cost sixteen consecutive OOB shots. Yes, one kernel read cost sixteen out-of-bounds writes, and here is its accounting:
Writing byte by byte: the arithmetic of accumulation
The OOB write gives us exactly one tuned byte per shot (the last byte of the window, d[47]) and the 14 bytes before it are random collateral that lands directly in front of it. So to write a full qword (8 bytes) at a given position we need 8 shots, writing the bytes from the highest to the lowest, so that each shot’s stray collateral lands on bytes we’ve already tuned or will tune later, repairing itself:
/* write a full qword at hash+QOFF: eight shots, from b7 down to b0 */
for (int b = 7; b >= 0; b--) {
uint32_t B = QOFF - 14 + b; /* d[47] lands on hash+QOFF+b */
uint8_t v = (val >> (8 * b)) & 0xff;
fire_oob_shot(B, v); /* brute: sha384 of a password where d[32]==B and d[47]==v */
}
But there’s a problem deeper than bytes: the chair problem. Every OOB shot goes through a full lifecycle: unregister frees the current hash, and a fresh register carves a new hash out of the freelist. Who guarantees that the new hash returns to the same slot, the same chair next to our victim? If the chair moved by even one shot, the next shot would write onto an object we don’t know, and we’d pay in panic or in eternal silence.
The solution is an ugly elegance called consumer rotation: we arrange the frees so that the victim-neighbor’s slot lands at the head of the freelist at the tuned moment, and part of it gets carved out by an intermediary allocation (a “consumer” message we send and receive in a special cycle) that swallows the slot we don’t want, while the hash returns to its wanted chair every cycle. And alongside it, continuous maintenance: a pool of messages from which we free one every cycle so the freelist never runs dry (an empty freelist means SLUB moves on to a fresh slab, and then things get freed in places our cycle can’t reach, and the chair slides away silently). And all of it verified under the microscope: we tracked the hash’s address across a hundred-plus consecutive restore calls and found it fixed in place like a government employee five minutes before opening time. The theory is boring and the result is decisive: this discipline is the difference between 18 consecutive writes onto the same object and 18 writes scattered at random, that is, the difference between a working exploit and a kernel panic you get to watch live.
The three deaths of the ROP chain
The seasoned reader now expects the familiar chapter: hijack an ops pointer, forge an operations table, stack pivot, then a ROP chain from commit_creds(&init_cred) through a KPTI trampoline to an iretq that drops us back into userspace wearing a crown. Correct expectation, and it’s exactly what we planned for, a full week of honest work. Then the plan died three times in a row, each time for a different reason, and this is the best part of the whole article, so read it twice:
The first death came from STATIC_USERMODEHELPER. Remember the classic hand-me-down weapons: write your program’s path into modprobe_path or core_pattern, and the kernel itself will call your program with full privileges at the first strange event. Those doors were all nailed shut in this build: the static_usermodehelper_path constant is set to an empty string, and call_usermodehelper ignores the requested path in the first place, reads the constant, and refuses to execute with EINVAL. The trick that fed generations of exploits is here just a footnote in a retirement memoir.
The second death came from the forged operations table itself. The idea of planting a fake ops table in writable .data dies simply, because STRICT_KERNEL_RWX makes .data non-executable. Fine, no big problem: ROP doesn’t execute .data in the first place, it executes gadgets from .text through pointers. But to reach the first gadget you need a stack pivot, and here we re-scanned .text in a way that resembled sifting sand: epilogues in every flavor, xchg, push/pop, indirect jumps on every base register, and all the famous gadget families. There is no viable pivot. The second grave.
The third death was the sharpest and most scientifically beautiful, and it came from a check that doesn’t exceed a single line in pipe_write. When the kernel tries to merge new data into an old buffer, it verifies that offset + len + chars doesn’t exceed PAGE_SIZE. The offset and len fields are unsigned int in the struct, but the compiled code in this build does a sign-extension and then an unsigned comparison. Any kernel address (they all start with 0xffff…) placed in one of those two fields turns, after sign-extension, into a cosmically large number, so the comparison fails immediately, and execution never even reaches confirm. In other words: merely placing a gadget into those two fields prevents you from reaching the gadget. Self-strangulation of rare skill.
Three roads closed, and each of them had been a whole door in another world. And here, instead of writing “and then we discovered the clever solution”, a literary lie, we’ll tell the truth: we didn’t discover anything ingenious. We just looked at the tool we’d been holding since the very beginning and read it again. The write that was going to plant the ROP payload is the same write that lands anywhere in physical memory. Who said we needed code execution in the first place?
The endgame: a direct physical write into cred
This is the moment the whole investigation turns from a battle into an audit. We don’t want to hijack any pointer, we don’t want to execute any instruction, and we don’t need a single gadget. We only want to write 72 bytes in one correct place: inside the cred of our own process. This is called data-only privilege escalation, and it’s the highest class of exploitation, because the kernel never executes a single thing of our choosing.
The machine used is the same pipe_buffer we mugged KASLR with, but this time from the other side: we’re the ones writing its fields. Recall the struct:
struct pipe_buffer { /* 40 bytes, kmalloc-64 */
struct page *page; /* +0 : the RAM page the buffer points at */
unsigned int offset, len; /* +8 : data position in the page and length */
const struct pipe_buf_operations *ops; /* +16 */
unsigned int flags; /* +24 : and CAN_MERGE lives here */
unsigned long private; /* +32 */
};
And on a write() to a pipe, if the last buffer in the ring carries a flag called CAN_MERGE, pipe_write merges the new data into the end of the buffer’s existing data, that is, at page + offset + len, directly inside the page. This is the offspring of the famous Dirty Pipe (CVE-2022-0847): there, the flag got set by mistake on a page-cache page, and read-only files got written into. Here it’s the same idea, but after we choose the page: we forge page to point at the struct page of any RAM page we want, we set offset and len so that the write lands at the wanted position inside it, we point ops at the standard pipe table (which contains no confirm, so there’s no objection raised in the middle), and flags at 0x58, where CAN_MERGE resides. Then we write to the pipe: one ordinary, modest, innocent-looking write(). The data lands in the physical page we chose, and all of it happens without a single instruction of our making being executed.
The full staging, abridged without betraying the spirit:
First, we prepare the weapon: a pipe with a single-slot ring, and we splice one byte into it from a dummy file of our own making (a file in /tmp, harming nobody). The splice hands the pipe_buffer live fields that the kernel fills in itself, and that reassures every code path that runs later. Then we perform the ballet on the freelist: we free the neighbor of the future hash’s seat, free a filler, unregister empties the two hash seats, the consumer message picks up one of them, register brings the hash back to its chair, and the single pipe_buffer array (40 bytes, that is, kmalloc-64) carves out the last available slot: the one sitting at exactly hash+64. From this moment on, every OOB shot fired from this hash chair is a sentence in the neighboring pipe_buffer’s constitution. (Two words summarize the SLUB philosophy: LIFO and timing. If the reader missed them, go back to the consumer rotation paragraph.)
Second, the first shot is flags: a single shot at B=88 drops the byte 0x58 into the flags field. And before we go on, we pre-verify: we write a single byte to the pipe. If the recruitment worked, CAN_MERGE is active, and the byte merges into the dummy file’s page itself (harmlessly, writing into our own file, nobody gets hurt), and write() returns 1. If the pre-verify fails, we withdraw immediately, and our loss is one shot of collateral instead of the dozens of shots that earlier versions of this exploit spent digging a cemetery into the heap before collapsing on top of it. This rule specifically (test with a byte before you pour by the ton) deserves to hang above every exploitation library in the world.
Third, the final accumulation in deliberate order: ops pointing to the standard pipe table (8 shots), then offset and len so that the merge position is exactly the target’s position in the page (8 shots), then page itself, to the struct page of the target page (8 shots). And the order is descending by address on purpose, because each qword’s random collateral lands in the fourteen bytes before it, so we arrange for it to always land on fields we’ll write later, or on the tail of the hash itself, where we don’t care. And through all of it, one strict rule: no read on the pipe (it consumes the buffer and destroys the disguise), and no write before the payload (it shifts the merge position). The payload write must be the first real operation on the weapon since the splice.
And the address arithmetic, shorter than you’d expect, because we already paid for it in full during the kread24 stage:
phys = cred_addr - page_offset_base /* the address in physical memory */
pfn = phys >> 12 /* the physical page number */
inpage = phys & 0xfff /* the target position inside the page */
page = vmemmap_base + pfn * 64 /* sizeof(struct page) == 64 here */
/* the values we plant into the pipe_buffer (with the same OOB shots, not from userspace): */
page = vmemmap_base + pfn * 64
offset = inpage - 1 ; len = 1 /* the merge lands at offset+len = inpage */
ops = &anon_pipe_buf_ops
flags = 0x58 /* CAN_MERGE */
Then the payload itself, 72 bytes starting at cred+0x14:
| fields at cred+ | new value |
|---|---|
| 0x14 to 0x30: uid and gid and the other eight relatives | zeros, all of them, no exceptions and no farewells |
| 0x34: securebits | zero |
| 0x38: cap_inheritable | zero |
| 0x40: cap_permitted | 0x000001ffffffffff, the full set |
| 0x48: cap_effective | 0x000001ffffffffff |
| 0x50: cap_bset | 0x000001ffffffffff |
| 0x58: cap_ambient (the part covered by the payload) | zero |
Why start at 0x14 instead of the top of the struct? Because this build enables CONFIG_DEBUG_CREDENTIALS, and at 0x10 lives a magic field that the kernel checks to make sure a cred is healthy. Writing that field would break the very check that was put there to catch people like us. So we started writing four bytes after it, and the magic stayed intact, testifying that everything was normal, exactly like a burglar who leaves the door locked behind him. And a side lesson from the same neighborhood: in an early version of the payload, the fields were written at a default qword alignment, so the capability sets landed four bytes late, and CapEff came out short, specifically without CAP_MKNOD, and mknod was politely refused (EPERM). Four bytes. In kernel space, four bytes separate owning the world from standing at a service window that never opens.
Then we write: write(weapon_fd, payload, 72). And the kernel, with its famously legendary honesty, merges the 72 bytes into the physical page we described to it. No ceremony, no crash, nothing. Just a getuid() afterward that returns zero, and that’s what root looks like in this language. And the only acceptable verdict is getuid() itself, because write() can “succeed” even if it merged into an ordinary pipe that slipped into the weapon’s place unnoticed, while uid flatters no one.
And before we move on, a word for the elephant in the room: the environment this whole story runs in raises no_new_privs (NNP) and runs seccomp. NNP closes the entire promotion-via-execve door: no su, no SUID, no file capabilities, and that’s deliberate, so that one logic bug in one module is the only road in. And seccomp blocks syscalls like mknod at the root. The design is logical from the environment owner’s point of view: he’s guarding all the known doors. But a write into kernel memory isn’t a door at all, it’s the wall the doors were made from. On our long road to uid 0, we didn’t use execve even once, and we never asked the system to promote us. We wrote the promotion ourselves, in the system’s own ledger.
Breaking out: eight bytes that redraw the map
Now you’re root. The blood whispers that you’re a king and the story is over. Then you open your eyes to the second uncomfortable truth: you’re root inside a cell. The whole process lives in a tight chroot jail: the filesystem root, from its point of view, is a subfolder somewhere, every path it tries to resolve lands inside that folder’s borders, and the “official” exit from the session powers down the entire machine on purpose. uid 0 answered “who are you?” with distinction, but one question remains unanswered: “where are you?”, and that question belongs to an entirely different field.
Every task_struct carries a pointer called fs that points to an fs_struct, and that struct carries root and pwd, the system’s permanent answer to “where does the world begin?” on every path the process walks. chroot doesn’t build walls around your process; it simply edits your copy of the answer. And the kernel itself keeps an authentic copy hanging on the constants: init_fs, the fs_struct of the first task since the system was born, whose root is the true root of the entire filesystem, the vantage point from which every layer of the prison looks like ordinary folders.
And the conclusion we paid the whole article for: this pointer is data. Eight bytes at task+0x6e0. And we own a machine that writes any bytes at any physical address, and we just proved it writes with four-byte precision (we didn’t forget the capabilities lesson). So we restaged everything from the top: the same grooming, the same chair, the same weapon, the same shot discipline, but the payload this time is a single value, the address of init_fs (which is kbase plus a constant offset we’ve known since KASLR fell, so not one extra read, because every extra read is extra risk, and the final write didn’t need a read at all). One write() of eight bytes, at task+0x6e0:
task->fs = init_fs
Then the quietest coup in history: from the first path resolution afterward, open and stat and readdir start resolving from the true root. The disk’s block device, which was impossible to see, appears as a node that already exists in the outer /dev, and opening it is allowed (seccomp blocked mknod, but opening an existing node apparently never occurred to anyone to block, and that’s exactly why we filled in the full capabilities: we need to read from the device, not create it). We read what we wanted from the outer world’s files directly, and then, out of narcissistic documentation, we wiped the raw disk itself block by block (a megabyte at a time, with a 256-byte overlap between blocks so that no text fingerprint would break at a read boundary) and confirmed that everything we saw through the filesystem we also found in the bare metal, and every result was cross-checked from two independent paths before being believed.
And notice what we didn’t do: we didn’t dance the classic chroot escape (make a directory, chroot into it, then climb ../../../.. past the root) or any of its cousins. Those tricks assume you’re asking the system for permission to leave, that the prison has locked your body. We didn’t ask permission and we never touched the walls: we changed the answer to “where am I”, so the definition of the prison stopped applying to us at all. The difference between the two is the difference between escaping a prison and buying the real estate the prison was built on.
If you sat on the other side of the table
Reread the journey now with a defender’s eye and ask: where should it have stopped? The answer is embarrassing, and that’s the most important thing about it:
| hardening | what it actually did to us |
|---|---|
| HARDENED_USERCOPY | blocked the normal receive of the inflated message, forcing us to launder the read through MSG_COPY |
| SLAB_FREELIST_HARDENED | encoded the freelist pointers, killing any classic freelist hijack and pushing us toward data-only |
| KASLR (with direct-map randomization) | forced a whole leak stage: an ops pointer into the text, and a read of all three bases |
| CONFIG_DEBUG_LIST | exposed the free-while-linked early and turned a mysterious crash into a diagnosis (but stopped nothing) |
| STATIC_USERMODEHELPER | killed the modprobe_path and core_pattern weapons in one sentence |
| STRICT_KERNEL_RWX | made .data non-executable, draining the forged table of its meaning |
| the sign-mangled merge check | strangled the idea of putting gadgets into offset/len in the crib |
| NNP with seccomp | closed su and SUID and mknod, and closed none of what we actually did |
Notice the disturbing pattern across the whole table: nothing in the list stopped the attack. Everything these hardenings did, collectively and with love, was force a more expensive, slower, smarter shape: from “copy the answer” to “launder the leak through messages that never get received”, from “hijack a pointer and dance the ROP” to “write the data directly and forget the dancing”. The hardenings were never a wall at any moment; they were switchbacks on a road that stayed open. And this is the central lesson of kernel exploitation for years now, repeating itself in every serious paper: writable memory is privilege. cred is data, task->fs is data, and data needs no gadgets to be interpreted, only whoever writes it.
And the real responsibility for the start of the whole story sits at the first door, at the module itself, and it gets three commandments from this journey: copy with memcpy and a length taken from the type, never from the data; make the label and the content transition together even on the failure paths; and never park variable-sized structures in one buffer, separated by a label living in a distant struct. Then a fourth commandment, deeper and less popular: think long and hard before putting password logic in kernel space at all, because you’re not buying security with it, you’re buying full privileges for the smallest typo you write.
The takeaway
We started with a copy function that reads its length from inside the very text it’s copying, and we ended up owning the kernel completely and holding a map of the world outside the prison, without executing a single instruction of our choosing in kernel space. No ROP, no stack pivot, not one gadget, not one hijacked pointer: just reads and writes. Between the start and the finish stood a hash lottery that pays out instantly, IPC messages that get read but never received, pipes that hand us a map of memory and then become a writing weapon, and one physical page that we write onto with a plain, modest write().
And the closing note worth carrying with you if you forget everything above: the whole story started because the kernel believed a struct that said “I’m 32” while carrying 48 inside. And the kernel, poor thing, isn’t stupid when it believes it; it’s quite the opposite: it believes it because nothing has ever lied to it before. Every piece of data in its memory was always honest, always exactly what it claimed to be, until some people started cooking passwords in userspace to make the hash jump wherever they pleased. So if you ever write code that distinguishes types by a detachable label, know that somebody, somewhere, is buying their ticket to the neighboring slot right now. And the neighbor, in kernel space, is always something worth reading.
References
- Dirty Pipe (CVE-2022-0847), the original disclosure by Max Kellermann
- CVE-2022-0847 in the NVD database
- The mainline fix commit for Dirty Pipe
- msgrcv(2), where MSG_COPY is documented
- The SLUB allocator, kernel documentation
- HARDENED_USERCOPY, kernel documentation
- Linux credentials (cred and init_cred), kernel documentation
- The physical memory model (struct page and vmemmap), kernel documentation
- seccomp(2) man page