| 0/0 |
2026/05/26 05:53 |
flow |
assessment-security |
47m
Results: map[DenialOfService:true Explanation:Based on the analysis of the bug report and the f2fs source code, here is the
assessment of the security impact:
### Bug Mechanism
The bug is a hung task (infinite loop) in `f2fs_sync_dirty_inodes` that occurs
during the mounting of a crafted f2fs image.
When mounting an f2fs filesystem, if there is fsync data to recover,
`f2fs_fill_super` calls `f2fs_recover_fsync_data`, which in turn calls
`f2fs_write_checkpoint` -> `block_operations` -> `f2fs_sync_dirty_inodes`.
If writing a dirty page fails during this process (e.g., due to a memory
allocation failure simulated by fault injection, or an `-EIO` / `-EFSCORRUPTED`
error caused by a corrupted node page in a malicious image),
`f2fs_write_single_data_page` redirties the page and returns the error. However,
`f2fs_sync_dirty_inodes` ignores the error returned by `filemap_fdatawrite` and
continues to loop as long as the inode's dirty list is not empty. Because the
page was redirtied and the error did not set the `CP_ERROR_FLAG` to abort the
checkpoint, the inode remains in the dirty list, resulting in an infinite loop.
### System-Wide Impact
During this infinite loop, the `mount` syscall holds the f2fs superblock's
`s_umount` semaphore for **write**.
Because `s_umount` is held for write indefinitely, any other system calls that
attempt to acquire `s_umount` for read on this superblock will block forever.
This includes system-wide operations like `quotactl` (which iterates over all
superblocks via `__iterate_supers` and takes `s_umount` for read) and `sync`.
This perfectly matches the syzkaller bug report where multiple tasks are blocked
for more than 143 seconds in `quotactl` -> `quota_sync_all` ->
`__iterate_supers` -> `super_lock`.
### Analysis Dimensions
* **Exploitable:** **No.** The bug is an infinite loop leading to a CPU spin and
a deadlock of the `s_umount` semaphore. There is no memory corruption,
use-after-free, or out-of-bounds access involved, meaning it cannot be exploited
for privilege escalation or arbitrary code execution.
* **Denial Of Service:** **Yes.** The infinite loop consumes a CPU core and
permanently blocks system-wide operations like `sync()` and `quotactl()`,
leading to a Denial of Service.
* **Accessible From Unprivileged Processes:** **No.** Mounting an f2fs
filesystem requires `CAP_SYS_ADMIN` in the initial user namespace. Unprivileged
users cannot trigger this directly.
* **Accessible From User Namespaces:** **No.** The f2fs filesystem does not have
the `FS_USERNS_MOUNT` flag, meaning it cannot be mounted from within an
unprivileged user namespace.
* **Peripheral Trigger:** **Yes.** An attacker can trigger this bug by plugging
in a malicious USB drive containing a crafted f2fs image, provided the desktop
environment (e.g., udisks2) is configured to auto-mount external media.
* **Malicious Filesystem Trigger:** **Yes.** The bug is triggered by parsing and
recovering a corrupted/malicious f2fs filesystem image.
* **VM Guest Trigger:** **No.**
* **VM Host Trigger in The Confidential Computing Context:** **No.**
* **Ethernet Network Trigger:** **No.**
* **Other Remote Trigger:** **No.** Exploitable:false FilesystemTrigger:true NetworkTrigger:false PeripheralTrigger:true RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
|
| 1/1 |
2026/05/26 05:53 |
action |
syz-repro-to-c-repro |
0m
Results: map[SimplifiedCRepro:// autogenerated by syzkaller (https://github.com/google/syzkaller)
#define _GNU_SOURCE
#include <endian.h>
#include <errno.h>
#include <fcntl.h>
#include <setjmp.h>
#include <stdbool.h>
#include <stddef.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/ioctl.h>
#include <sys/mman.h>
#include <sys/mount.h>
#include <sys/stat.h>
#include <sys/syscall.h>
#include <sys/types.h>
#include <unistd.h>
#include <linux/loop.h>
#ifndef __NR_memfd_create
#define __NR_memfd_create 319
#endif
static unsigned long long procid;
//% This code is derived from puff.{c,h}, found in the zlib development. The
//% original files come with the following copyright notice:
//% Copyright (C) 2002-2013 Mark Adler, all rights reserved
//% version 2.3, 21 Jan 2013
//% This software is provided 'as-is', without any express or implied
//% warranty. In no event will the author be held liable for any damages
//% arising from the use of this software.
//% Permission is granted to anyone to use this software for any purpose,
//% including commercial applications, and to alter it and redistribute it
//% freely, subject to the following restrictions:
//% 1. The origin of this software must not be misrepresented; you must not
//% claim that you wrote the original software. If you use this software
//% in a product, an acknowledgment in the product documentation would be
//% appreciated but is not required.
//% 2. Altered source versions must be plainly marked as such, and must not be
//% misrepresented as being the original software.
//% 3. This notice may not be removed or altered from any source distribution.
//% Mark Adler madler@alumni.caltech.edu
//% BEGIN CODE DERIVED FROM puff.{c,h}
#define MAXBITS 15
#define MAXLCODES 286
#define MAXDCODES 30
#define MAXCODES (MAXLCODES + MAXDCODES)
#define FIXLCODES 288
struct puff_state {
unsigned char* out;
unsigned long outlen;
unsigned long outcnt;
const unsigned char* in;
unsigned long inlen;
unsigned long incnt;
int bitbuf;
int bitcnt;
jmp_buf env;
};
static int puff_bits(struct puff_state* s, int need)
{
long val = s->bitbuf;
while (s->bitcnt < need) {
if (s->incnt == s->inlen)
longjmp(s->env, 1);
val |= (long)(s->in[s->incnt++]) << s->bitcnt;
s->bitcnt += 8;
}
s->bitbuf = (int)(val >> need);
s->bitcnt -= need;
return (int)(val & ((1L << need) - 1));
}
static int puff_stored(struct puff_state* s)
{
s->bitbuf = 0;
s->bitcnt = 0;
if (s->incnt + 4 > s->inlen)
return 2;
unsigned len = s->in[s->incnt++];
len |= s->in[s->incnt++] << 8;
if (s->in[s->incnt++] != (~len & 0xff) ||
s->in[s->incnt++] != ((~len >> 8) & 0xff))
return -2;
if (s->incnt + len > s->inlen)
return 2;
if (s->outcnt + len > s->outlen)
return 1;
for (; len--; s->outcnt++, s->incnt++) {
if (s->in[s->incnt])
s->out[s->outcnt] = s->in[s->incnt];
}
return 0;
}
struct puff_huffman {
short* count;
short* symbol;
};
static int puff_decode(struct puff_state* s, const struct puff_huffman* h)
{
int first = 0;
int index = 0;
int bitbuf = s->bitbuf;
int left = s->bitcnt;
int code = first = index = 0;
int len = 1;
short* next = h->count + 1;
while (1) {
while (left--) {
code |= bitbuf & 1;
bitbuf >>= 1;
int count = *next++;
if (code - count < first) {
s->bitbuf = bitbuf;
s->bitcnt = (s->bitcnt - len) & 7;
return h->symbol[index + (code - first)];
}
index += count;
first += count;
first <<= 1;
code <<= 1;
len++;
}
left = (MAXBITS + 1) - len;
if (left == 0)
break;
if (s->incnt == s->inlen)
longjmp(s->env, 1);
bitbuf = s->in[s->incnt++];
if (left > 8)
left = 8;
}
return -10;
}
static int puff_construct(struct puff_huffman* h, const short* length, int n)
{
int len;
for (len = 0; len <= MAXBITS; len++)
h->count[len] = 0;
int symbol;
for (symbol = 0; symbol < n; symbol++)
(h->count[length[symbol]])++;
if (h->count[0] == n)
return 0;
int left = 1;
for (len = 1; len <= MAXBITS; len++) {
left <<= 1;
left -= h->count[len];
if (left < 0)
return left;
}
short offs[MAXBITS + 1];
offs[1] = 0;
for (len = 1; len < MAXBITS; len++)
offs[len + 1] = offs[len] + h->count[len];
for (symbol = 0; symbol < n; symbol++)
if (length[symbol] != 0)
h->symbol[offs[length[symbol]]++] = symbol;
return left;
}
static int puff_codes(struct puff_state* s,
const struct puff_huffman* lencode,
const struct puff_huffman* distcode)
{
static const short lens[29] = {
3, 4, 5, 6, 7, 8, 9, 10, 11, 13, 15, 17, 19, 23, 27, 31,
35, 43, 51, 59, 67, 83, 99, 115, 131, 163, 195, 227, 258};
static const short lext[29] = {
0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 2, 2, 2, 2,
3, 3, 3, 3, 4, 4, 4, 4, 5, 5, 5, 5, 0};
static const short dists[30] = {
1, 2, 3, 4, 5, 7, 9, 13, 17, 25, 33, 49, 65, 97, 129, 193,
257, 385, 513, 769, 1025, 1537, 2049, 3073, 4097, 6145,
8193, 12289, 16385, 24577};
static const short dext[30] = {
0, 0, 0, 0, 1, 1, 2, 2, 3, 3, 4, 4, 5, 5, 6, 6,
7, 7, 8, 8, 9, 9, 10, 10, 11, 11,
12, 12, 13, 13};
int symbol;
do {
symbol = puff_decode(s, lencode);
if (symbol < 0)
return symbol;
if (symbol < 256) {
if (s->outcnt == s->outlen)
return 1;
if (symbol)
s->out[s->outcnt] = symbol;
s->outcnt++;
} else if (symbol > 256) {
symbol -= 257;
if (symbol >= 29)
return -10;
int len = lens[symbol] + puff_bits(s, lext[symbol]);
symbol = puff_decode(s, distcode);
if (symbol < 0)
return symbol;
unsigned dist = dists[symbol] + puff_bits(s, dext[symbol]);
if (dist > s->outcnt)
return -11;
if (s->outcnt + len > s->outlen)
return 1;
while (len--) {
if (dist <= s->outcnt && s->out[s->outcnt - dist])
s->out[s->outcnt] = s->out[s->outcnt - dist];
s->outcnt++;
}
}
} while (symbol != 256);
return 0;
}
static int puff_fixed(struct puff_state* s)
{
static int virgin = 1;
static short lencnt[MAXBITS + 1], lensym[FIXLCODES];
static short distcnt[MAXBITS + 1], distsym[MAXDCODES];
static struct puff_huffman lencode, distcode;
if (virgin) {
lencode.count = lencnt;
lencode.symbol = lensym;
distcode.count = distcnt;
distcode.symbol = distsym;
short lengths[FIXLCODES];
int symbol;
for (symbol = 0; symbol < 144; symbol++)
lengths[symbol] = 8;
for (; symbol < 256; symbol++)
lengths[symbol] = 9;
for (; symbol < 280; symbol++)
lengths[symbol] = 7;
for (; symbol < FIXLCODES; symbol++)
lengths[symbol] = 8;
puff_construct(&lencode, lengths, FIXLCODES);
for (symbol = 0; symbol < MAXDCODES; symbol++)
lengths[symbol] = 5;
puff_construct(&distcode, lengths, MAXDCODES);
virgin = 0;
}
return puff_codes(s, &lencode, &distcode);
}
static int puff_dynamic(struct puff_state* s)
{
static const short order[19] =
{16, 17, 18, 0, 8, 7, 9, 6, 10, 5, 11, 4, 12, 3, 13, 2, 14, 1, 15};
int nlen = puff_bits(s, 5) + 257;
int ndist = puff_bits(s, 5) + 1;
int ncode = puff_bits(s, 4) + 4;
if (nlen > MAXLCODES || ndist > MAXDCODES)
return -3;
short lengths[MAXCODES];
int index;
for (index = 0; index < ncode; index++)
lengths[order[index]] = puff_bits(s, 3);
for (; index < 19; index++)
lengths[order[index]] = 0;
short lencnt[MAXBITS + 1], lensym[MAXLCODES];
struct puff_huffman lencode = {lencnt, lensym};
int err = puff_construct(&lencode, lengths, 19);
if (err != 0)
return -4;
index = 0;
while (index < nlen + ndist) {
int symbol;
int len;
symbol = puff_decode(s, &lencode);
if (symbol < 0)
return symbol;
if (symbol < 16)
lengths[index++] = symbol;
else {
len = 0;
if (symbol == 16) {
if (index == 0)
return -5;
len = lengths[index - 1];
symbol = 3 + puff_bits(s, 2);
} else if (symbol == 17)
symbol = 3 + puff_bits(s, 3);
else
symbol = 11 + puff_bits(s, 7);
if (index + symbol > nlen + ndist)
return -6;
while (symbol--)
lengths[index++] = len;
}
}
if (lengths[256] == 0)
return -9;
err = puff_construct(&lencode, lengths, nlen);
if (err && (err < 0 || nlen != lencode.count[0] + lencode.count[1]))
return -7;
short distcnt[MAXBITS + 1], distsym[MAXDCODES];
struct puff_huffman distcode = {distcnt, distsym};
err = puff_construct(&distcode, lengths + nlen, ndist);
if (err && (err < 0 || ndist != distcode.count[0] + distcode.count[1]))
return -8;
return puff_codes(s, &lencode, &distcode);
}
static int puff(
unsigned char* dest,
unsigned long* destlen,
const unsigned char* source,
unsigned long sourcelen)
{
struct puff_state s = {
.out = dest,
.outlen = *destlen,
.outcnt = 0,
.in = source,
.inlen = sourcelen,
.incnt = 0,
.bitbuf = 0,
.bitcnt = 0,
};
int err;
if (setjmp(s.env) != 0)
err = 2;
else {
int last;
do {
last = puff_bits(&s, 1);
int type = puff_bits(&s, 2);
err = type == 0 ? puff_stored(&s) : (type == 1 ? puff_fixed(&s) : (type == 2 ? puff_dynamic(&s) : -1));
if (err != 0)
break;
} while (!last);
}
*destlen = s.outcnt;
return err;
}
//% END CODE DERIVED FROM puff.{c,h}
#define ZLIB_HEADER_WIDTH 2
static int puff_zlib_to_file(const unsigned char* source, unsigned long sourcelen, int dest_fd)
{
if (sourcelen < ZLIB_HEADER_WIDTH)
return 0;
source += ZLIB_HEADER_WIDTH;
sourcelen -= ZLIB_HEADER_WIDTH;
const unsigned long max_destlen = 132 << 20;
void* ret = mmap(0, max_destlen, PROT_WRITE | PROT_READ, MAP_PRIVATE | MAP_ANON, -1, 0);
if (ret == MAP_FAILED)
return -1;
unsigned char* dest = (unsigned char*)ret;
unsigned long destlen = max_destlen;
int err = puff(dest, &destlen, source, sourcelen);
if (err) {
munmap(dest, max_destlen);
errno = -err;
return -1;
}
if (write(dest_fd, dest, destlen) != (ssize_t)destlen) {
munmap(dest, max_destlen);
return -1;
}
return munmap(dest, max_destlen);
}
static int setup_loop_device(unsigned char* data, unsigned long size, const char* loopname, int* loopfd_p)
{
int err = 0, loopfd = -1;
int memfd = syscall(__NR_memfd_create, "syzkaller", 0);
if (memfd == -1) {
err = errno;
goto error;
}
if (puff_zlib_to_file(data, size, memfd)) {
err = errno;
goto error_close_memfd;
}
loopfd = open(loopname, O_RDWR);
if (loopfd == -1) {
err = errno;
goto error_close_memfd;
}
if (ioctl(loopfd, LOOP_SET_FD, memfd)) {
if (errno != EBUSY) {
err = errno;
goto error_close_loop;
}
ioctl(loopfd, LOOP_CLR_FD, 0);
usleep(1000);
if (ioctl(loopfd, LOOP_SET_FD, memfd)) {
err = errno;
goto error_close_loop;
}
}
close(memfd);
*loopfd_p = loopfd;
return 0;
error_close_loop:
close(loopfd);
error_close_memfd:
close(memfd);
error:
errno = err;
return -1;
}
static void reset_loop_device(const char* loopname)
{
int loopfd = open(loopname, O_RDWR);
if (loopfd == -1) {
return;
}
if (ioctl(loopfd, LOOP_CLR_FD, 0)) {
}
close(loopfd);
}
static long syz_mount_image(
volatile long fsarg,
volatile long dir,
volatile long flags,
volatile long optsarg,
volatile long change_dir,
volatile unsigned long size,
volatile long image)
{
unsigned char* data = (unsigned char*)image;
int res = -1, err = 0, need_loop_device = !!size;
char* mount_opts = (char*)optsarg;
char* target = (char*)dir;
char* fs = (char*)fsarg;
char* source = NULL;
char loopname[64];
if (need_loop_device) {
int loopfd;
memset(loopname, 0, sizeof(loopname));
snprintf(loopname, sizeof(loopname), "/dev/loop%llu", procid);
if (setup_loop_device(data, size, loopname, &loopfd) == -1)
return -1;
close(loopfd);
source = loopname;
}
mkdir(target, 0777);
char opts[256];
memset(opts, 0, sizeof(opts));
if (strlen(mount_opts) > (sizeof(opts) - 32)) {
}
strncpy(opts, mount_opts, sizeof(opts) - 32);
if (strcmp(fs, "iso9660") == 0) {
flags |= MS_RDONLY;
} else if (strncmp(fs, "ext", 3) == 0) {
bool has_remount_ro = false;
char* remount_ro_start = strstr(opts, "errors=remount-ro");
if (remount_ro_start != NULL) {
char after = *(remount_ro_start + strlen("errors=remount-ro"));
char before = remount_ro_start == opts ? '\0' : *(remount_ro_start - 1);
has_remount_ro = ((before == '\0' || before == ',') && (after == '\0' || after == ','));
}
if (strstr(opts, "errors=panic") || !has_remount_ro)
strcat(opts, ",errors=continue");
} else if (strcmp(fs, "xfs") == 0) {
strcat(opts, ",nouuid");
} else if (strncmp(fs, "gfs2", 4) == 0 && (strstr(opts, "errors=panic") || strstr(opts, "debug"))) {
strcat(opts, ",errors=withdraw");
}
res = mount(source, target, fs, flags, opts);
if (res == -1) {
err = errno;
goto error_clear_loop;
}
res = open(target, O_RDONLY | O_DIRECTORY);
if (res == -1) {
err = errno;
goto error_clear_loop;
}
if (change_dir) {
res = chdir(target);
if (res == -1) {
err = errno;
}
}
error_clear_loop:
if (need_loop_device)
reset_loop_device(loopname);
errno = err;
return res;
}
uint64_t r[1] = {0xffffffffffffffff};
int main(void)
{
syscall(__NR_mmap, /*addr=*/0x1ffffffff000ul, /*len=*/0x1000ul, /*prot=*/0ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
syscall(__NR_mmap, /*addr=*/0x200000000000ul, /*len=*/0x1000000ul, /*prot=PROT_WRITE|PROT_READ|PROT_EXEC*/7ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
syscall(__NR_mmap, /*addr=*/0x200001000000ul, /*len=*/0x1000ul, /*prot=*/0ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
const char* reason;
(void)reason;
intptr_t res = 0;
if (write(1, "executing program\n", sizeof("executing program\n") - 1)) {}
// syz_mount_image$f2fs arguments: [
// fs: ptr[in, buffer] {
// buffer: {66 32 66 73 00} (length 0x5)
// }
// dir: ptr[in, buffer] {
// buffer: {2e 2f 66 69 6c 65 30 00} (length 0x8)
// }
// flags: mount_flags = 0x8 (8 bytes)
// opts: ptr[in, fs_options[f2fs_options]] {
// fs_options[f2fs_options] {
// elems: array[fs_opt_elem[f2fs_options]] {
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// nodiscard: buffer: {6e 6f 64 69 73 63 61 72 64} (length 0x9)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// noinline_dentry: buffer: {6e 6f 69 6e 6c 69 6e 65 5f 64 65 6e 74 72 79} (length 0xf)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// alloc_mode_def: buffer: {61 6c 6c 6f 63 5f 6d 6f 64 65 3d 64 65 66 61 75 6c 74} (length 0x12)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// six_active_logs: buffer: {61 63 74 69 76 65 5f 6c 6f 67 73 3d 36} (length 0xd)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// quota: buffer: {71 75 6f 74 61} (length 0x5)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// fault_injection: fs_opt["fault_injection", fmt[dec, int32]] {
// name: buffer: {66 61 75 6c 74 5f 69 6e 6a 65 63 74 69 6f 6e} (length 0xf)
// eq: const = 0x3d (1 bytes)
// val: int32 = 0x10 (20 bytes)
// }
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// discard: buffer: {64 69 73 63 61 72 64} (length 0x7)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// noacl: buffer: {6e 6f 61 63 6c} (length 0x5)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// fsync_mode_posix: buffer: {66 73 79 6e 63 5f 6d 6f 64 65 3d 70 6f 73 69 78} (length 0x10)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// alloc_mode_def: buffer: {61 6c 6c 6f 63 5f 6d 6f 64 65 3d 64 65 66 61 75 6c 74} (length 0x12)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// noextent_cache: buffer: {6e 6f 65 78 74 65 6e 74 5f 63 61 63 68 65} (length 0xe)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// two_active_logs: buffer: {61 63 74 69 76 65 5f 6c 6f 67 73 3d 32} (length 0xd)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// fault_type: fs_opt["fault_type", fmt[dec, int32]] {
// name: buffer: {66 61 75 6c 74 5f 74 79 70 65} (length 0xa)
// eq: const = 0x3d (1 bytes)
// val: int32 = 0xfffffe (20 bytes)
// }
// }
// comma: const = 0x2c (1 bytes)
// }
// }
// common: array[fs_opt_elem[fs_options_common]] {
// }
// null: const = 0x0 (1 bytes)
// }
// }
// chdir: int8 = 0x1 (1 bytes)
// size: len = 0x5505 (8 bytes)
// img: ptr[in, buffer] {
// buffer: (compressed buffer with length 0x5505)
// }
// ]
// returns fd_dir
memcpy((void*)0x200000000000, "f2fs\000", 5);
memcpy((void*)0x200000000040, "./file0\000", 8);
memcpy((void*)0x200000000240, "nodiscard", 9);
*(uint8_t*)0x200000000249 = 0x2c;
memcpy((void*)0x20000000024a, "noinline_dentry", 15);
*(uint8_t*)0x200000000259 = 0x2c;
memcpy((void*)0x20000000025a, "alloc_mode=default", 18);
*(uint8_t*)0x20000000026c = 0x2c;
memcpy((void*)0x20000000026d, "active_logs=6", 13);
*(uint8_t*)0x20000000027a = 0x2c;
memcpy((void*)0x20000000027b, "quota", 5);
*(uint8_t*)0x200000000280 = 0x2c;
memcpy((void*)0x200000000281, "fault_injection", 15);
*(uint8_t*)0x200000000290 = 0x3d;
sprintf((char*)0x200000000291, "%020llu", (long long)0x10);
*(uint8_t*)0x2000000002a5 = 0x2c;
memcpy((void*)0x2000000002a6, "discard", 7);
*(uint8_t*)0x2000000002ad = 0x2c;
memcpy((void*)0x2000000002ae, "noacl", 5);
*(uint8_t*)0x2000000002b3 = 0x2c;
memcpy((void*)0x2000000002b4, "fsync_mode=posix", 16);
*(uint8_t*)0x2000000002c4 = 0x2c;
memcpy((void*)0x2000000002c5, "alloc_mode=default", 18);
*(uint8_t*)0x2000000002d7 = 0x2c;
memcpy((void*)0x2000000002d8, "noextent_cache", 14);
*(uint8_t*)0x2000000002e6 = 0x2c;
memcpy((void*)0x2000000002e7, "active_logs=2", 13);
*(uint8_t*)0x2000000002f4 = 0x2c;
memcpy((void*)0x2000000002f5, "fault_type", 10);
*(uint8_t*)0x2000000002ff = 0x3d;
sprintf((char*)0x200000000300, "%020llu", (long long)0xfffffe);
*(uint8_t*)0x200000000314 = 0x2c;
*(uint8_t*)0x200000000315 = 0;
memcpy((void*)0x200000000340, "... [truncated large byte array] ...", 21765);
syz_mount_image(/*fs=*/0x200000000000, /*dir=*/0x200000000040, /*flags=MS_NOEXEC*/8, /*opts=*/0x200000000240, /*chdir=*/1, /*size=*/0x5505, /*img=*/0x200000000340);
// openat$cgroup_ro arguments: [
// fd: fd_cgroup (resource)
// file: ptr[in, buffer] {
// buffer: {63 70 75 2e 73 74 61 74 00} (length 0x9)
// }
// flags: const = 0x275a (4 bytes)
// mode: const = 0x0 (2 bytes)
// ]
// returns fd
memcpy((void*)0x200000000040, "cpu.stat\000", 9);
syscall(__NR_openat, /*fd=*/0xffffff9c, /*file=*/0x200000000040ul, /*flags=*/0x275a, /*mode=*/0);
// openat arguments: [
// fd: fd_dir (resource)
// file: ptr[in, buffer] {
// buffer: {2e 2f 66 69 6c 65 31 00} (length 0x8)
// }
// flags: open_flags = 0x105042 (4 bytes)
// mode: open_mode = 0x0 (2 bytes)
// ]
// returns fd
memcpy((void*)0x200000000080, "./file1\000", 8);
syscall(__NR_openat, /*fd=*/0xffffff9c, /*file=*/0x200000000080ul, /*flags=O_SYNC|O_DIRECT|O_CREAT|O_RDWR*/0x105042, /*mode=*/0);
// openat$cgroup_ro arguments: [
// fd: fd_cgroup (resource)
// file: ptr[in, buffer] {
// buffer: {63 70 75 2e 73 74 61 74 00} (length 0x9)
// }
// flags: const = 0x275a (4 bytes)
// mode: const = 0x0 (2 bytes)
// ]
// returns fd
memcpy((void*)0x200000000040, "cpu.stat\000", 9);
res = syscall(__NR_openat, /*fd=*/0xffffff9c, /*file=*/0x200000000040ul, /*flags=*/0x275a, /*mode=*/0);
if (res != -1)
r[0] = res;
// write$UHID_INPUT arguments: [
// fd: fd_uhid (resource)
// data: ptr[in, uhid_req[UHID_INPUT, uhid_input_req]] {
// uhid_req[UHID_INPUT, uhid_input_req] {
// type: const = 0x8 (4 bytes)
// data: uhid_input_req {
// data: buffer: {85 f0 80 a4 93 3d 55 26 6e 07 e7 99 aa 0c c4 21 38 82 42 df 2a 3c 6b 63 1b 65 b1 c0 61 ed d2 aa 10 8c 35 28 fe 9b 0b b3 a5 3a b1 20 0f 5d 01 a6 8a 4a cd ec 8f ee 09 64 82 22 f9 08 c1 fe dc 30 00 34 2e 61 39 de 28 36 6c 13 50 93 06 d0 0e bc c6 74 97 18 1a c9 16 db 98 af 9d 36 6b 76 e4 27 d9 ab 5b b6 80 95 f0 fb 24 6d f3 2b 8a f0 78 36 53 13 6f 8a 04 c0 36 90 31 21 25 c7 de d6 a2 4f da 86 85 34 0c 57 5e ad 69 51 9e 35 83 f8 9d 46 7e c2 32 d6 a1 ff d0 46 3b a4 ea 3c ba e5 da e6 65 4b 55 47 b5 45 8f 02 ac 30 77 29 e5 7b 09 e1 34 f6 8b e4 4f 88 d7 25 17 b2 30 b0 66 f6 31 5b 5f b8 02 06 39 7b bf f8 cb c2 a3 6e 01 c2 e7 b3 aa db 32 bd 3d d5 28 8a 69 a9 91 d9 c6 74 71 7e 3a bb a7 16 72 80 b2 db 3b 1b 85 02 af a4 f3 f2 96 c5 32 51 0c 9d 2d d7 9b b5 ee b2 5a db 5e dd db dd 06 9c 09 d1 4d 15 c2 e7 e1 e2 bd 6c 10 8f ab 35 91 bb 22 e9 7d 69 92 23 6d 22 73 c8 bb 95 53 6f 71 18 d0 07 96 50 08 b1 25 c7 da ac 28 14 e6 bb e1 ad bf a3 57 2a d0 b7 ad 5c 26 c8 01 41 18 d8 37 4c a9 f2 85 77 9d fe e7 71 5a 40 39 08 14 6a 74 de 61 b3 85 39 14 c8 9f 44 4c 12 e7 a3 8b dd 46 c4 ed 36 eb 80 6e a5 98 f4 4d 1d ec 9e ff 9e 24 76 f4 38 02 21 1f 07 62 b6 66 73 b4 5d 23 6b 23 91 ce 32 2e 30 fb 9c 69 fe 0d 51 4d c1 f8 b6 e3 97 9c 12 05 fd 52 24 b0 7d 18 a4 4f ec 4f 6f 1a 6f 65 15 8b b6 ad cc 29 5b f2 dd 7d ea 10 7f 59 d7 e0 3c 61 fe 58 22 29 2e 45 96 89 56 b9 31 bd c4 d6 44 5f f1 63 1e 0b 98 e4 b4 44 87 74 dd 4b 9c d5 3a 45 89 6f db 3f 03 70 27 78 74 1a e2 b4 5a 25 bf 9a 23 fc 02 fb 97 a6 30 f1 32 bf 9d ef 6c 6d 4a 7b ae b6 29 72 f1 a8 14 f6 f2 37 7b cf c7 8e 2e 86 36 8c 13 85 10 a0 4c ed f7 17 5a f8 c2 03 4f ae 74 13 e3 ac e8 c7 1a b9 a0 af 1c a7 04 20 11 a6 ed 02 8e 20 56 48 53 5d ab f3 b2 f8 51 96 ae 18 d3 6b 83 9e 3c d5 4a e4 93 3a d5 29 88 8f da c7 bb 8a 70 c7 2b c0 fc 81 ba 06 50 6f 2d 5b c7 68 6e 21 9b be 52 83 95 9c be f9 95 0e 07 1c b6 d9 f3 41 fc 62 4a 51 10 34 1f 26 ce bd 71 00 59 9a 06 e6 1f 66 fa e1 20 c7 fc 2b 34 c6 22 12 00 eb a7 5b d1 27 71 14 67 1a 3f a8 f0 58 b2 7f d8 97 b0 52 f4 a5 2a fc ea 81 4d f5 26 18 1c 75 c4 49 72 10 a2 b8 b7 4e 26 60 15 61 e7 87 35 38 7c f1 23 65 4b 02 95 d1 d6 05 56 95 6b 36 d9 6d d0 38 86 6c 4b 4d b3 1e bd cd dd 69 29 bb c2 85 0c d4 90 13 89 e6 ea 6e 86 04 1e 0e fa 11 58 f3 34 e7 af da 0e 11 c2 fb 0e 6d f6 36 4c b9 56 59 f5 06 d5 c7 e6 3f b6 7c 81 16 57 7d 15 e4 a4 b1 fc 4c 27 de 2e 52 58 6c b1 f5 2b e9 c3 60 1f 50 66 54 9d e8 bd c3 ec 07 d1 a8 4c af 19 61 32 3e c2 48 7a 37 b7 51 ae ab af cd 64 7c e2 da e5 d9 49 9c 0f 96 94 67 e6 ca ba d1 98 66 9a c9 6b d1 48 89 54 ef f0 85 4e e0 c8 3d 7b 59 6d 27 36 25 bd b1 62 70 78 23 21 07 1f da 5d 98 0d ed 78 ff a9 dc 2b 56 03 7d 7c bf 94 25 47 f4 8a 51 31 f1 99 1f 6c 17 ae 1e d5 12 0c a6 87 8f 98 e6 8e 79 97 a9 a2 b7 0b e6 40 a7 0a 34 ad b8 0d e2 86 c6 69 2a bb 5f 09 2e 4e 3a 15 a8 32 17 e0 3d 02 a4 05 4f 34 af 3a 65 ff 6b 36 f3 95 b7 6a 05 79 cf fa fd 5d 3b b0 e7 04 c9 35 ca ec f3 a7 ab 75 6c 23 fd 60 c9 fe 3f 4f b2 be 75 04 f5 ba e2 2b 11 6f f1 58 8d cf 02 b3 27 d3 1b f0 48 8d ba 8a f5 b3 3c cf 2d 7d 87 f4 3b bc 48 fc d4 f1 91 ad 6a f9 31 3a d3 8b 8b 29 67 4b fa bd 66 51 bc 1f 6c e5 ab b4 a2 f1 41 31 94 f9 6b 26 d7 d6 ed c4 e0 13 fb a5 49 07 5c 97 ee f5 08 af 5c a7 87 36 64 b0 58 b7 bc f4 55 a8 a0 4b 59 1d 29 fa b6 36 6c 84 4b b7 55 76 ba c2 d5 23 23 e7 47 30 3d 00 a5 73 6c 98 12 92 2b 0e 17 bc ec 91 35 55 07 36 b5 4c f6 40 7d 61 e2 2e 62 d7 bb 75 f6 29 35 b6 65 ac f3 3e 75 f6 88 c3 6e f4 16 f1 b8 90 d0 f0 c8 ad 1d f0 0e 02 ec 45 96 78 34 d5 64 9c 8e 71 43 97 86 22 fa 37 04 67 29 70 b7 99 3a 87 e9 7d 3d 92 6a 14 26 56 47 bc 8b 8c 9e 6f 83 e2 95 72 60 8d 24 b4 2c 26 35 ef 4a bb d0 af 83 86 0e 99 c9 0d 74 71 cf 6e 8c e9 95 07 f5 ec 2b c5 72 21 2f a9 ee 3f 5a 9d fa 38 15 fe 55 f0 bb b1 19 ac ce 06 2a e3 7f 2f f9 21 70 7a bb a1 39 bc dd f4 2b fd 17 4d 29 b5 40 16 1b 41 13 c4 e1 a1 3f 3a 62 8c 63 8e c4 d3 a8 84 df bc 09 3e 23 ec 0d 06 71 b4 6b 41 dc 8b 42 d9 50 c8 61 5b a5 ee 87 f4 9b 5d 09 10 ff a4 87 12 07 99 50 01 92 0d b0 5a 95 19 99 67 f0 97 ba 7b 55 bb d2 71 d8 18 69 0c 42 38 40 6b 40 a3 df c4 2f a5 6a 67 17 3b 53 a9 6b 54 33 26 c5 67 38 b6 d0 43 19 59 34 01 86 96 f5 ab 49 34 7e 51 48 a7 8f 2d 13 69 a7 1a fa b8 33 02 73 d4 6e cf ba 4e e0 58 02 a5 38 56 49 85 1d b9 49 db fb 39 e2 90 94 16 41 c5 0b 1a c2 0f b3 10 27 54 a7 60 b0 97 f4 64 dd b0 b8 3f 81 68 ba df a7 1d b6 62 1d cf 22 fb 08 1e 34 03 f3 ba c5 c7 e6 59 05 ac a5 28 85 c8 07 f8 dd ab 18 bb 2f 12 ef 95 2c 50 48 3c 0e 25 19 68 bc 70 ff 0d 42 a6 38 ca 74 4d ea 4c 7e bb 4f ea 77 7c f6 63 bb 4f 15 05 ed 79 73 0c 45 bc 86 e4 88 a1 3f 92 43 77 a8 e2 ee 66 70 a0 2c a5 28 74 ae 1c 42 a3 5d 55 b9 76 57 57 04 7b 2c c3 74 2a a5 1f a3 e4 3f b2 c1 13 c9 2a d2 13 ba d2 52 c1 a8 29 66 dd 01 6f 12 a7 f1 c3 90 0c 0f 1a b4 55 03 51 63 f3 18 99 bd d3 0f 3f f4 3a d1 7d 9e 45 bb 74 38 c1 c9 86 71 27 36 f2 4b e1 4f 71 ab 1b fe 92 a2 5e c0 7f 08 6e e8 c7 97 1b 80 77 a1 3e 58 a8 e8 be a3 9c 8e 06 b2 51 90 9f 02 cb 00 80 ab f0 20 f2 7c a1 60 eb 26 c0 82 dd a1 fa 54 ea 40 94 df db cb 2f a7 bd dc cb 67 a8 44 e8 07 5f 4c c0 8d ad 35 75 70 06 d0 51 e1 83 dc ed 33 6b c0 c2 50 2f 93 ff c8 7d ca 62 22 86 ba 17 4c 24 e1 f5 3f 27 dc 27 77 ba af e1 70 34 8b 0e 8d 3e 74 3b 3a a9 06 bc 07 64 bb e7 da 08 ff 40 3e fe 22 12 62 7d 67 22 50 65 8b b5 13 b7 31 25 17 d1 f8 8c 61 c7 ba 5f 96 47 cd 61 92 81 c5 b3 90 b4 86 06 ee 39 fb 41 71 10 3d f2 e0 9d 7c fd 56 c0 6c 72 1f 7c 24 ad 8c ce 38 36 23 fc 2d cb 15 ac 56 43 8e a3 31 82 0a e5 9c 8c 47 4e 36 fc 73 f7 b1 b3 b8 6d f1 b4 24 90 81 55 13 68 1a af bf 7e 87 1b 4b 96 86 ef ae 6c 45 ec fc a6 0a 64 0a 6f 07 1d fd 31 f9 43 7c 3d 03 08 61 64 b4 8c 1e d8 02 98 68 64 bf e0 d4 9b dd 77 09 66 22 62 36 8d bc 3e cc 05 eb 24 0e cc 41 90 4c 76 d7 8a b5 c5 2b 66 af 5a 72 0f dd 6a 92 f5 2b e0 67 64 27 a5 6e 32 e5 bc 50 85 b2 5f 90 ad d2 8a 76 f2 fc e6 f8 f0 ef 74 f4 65 96 98 54 96 46 bd 63 17 5a df 77 b5 cd cf e6 76 e1 b1 a9 af 15 10 29 46 55 4b a6 13 6c bc 83 c6 26 8e e4 03 18 f3 c9 d4 71 80 25 68 8b 35 d2 26 5b f6 0b f8 89 ff 62 9f 78 34 58 6e f4 6e ab 7a 91 76 33 75 36 bb 60 01 e6 76 54 6b 98 7f 36 b1 fe 4b 9f 6e 46 a8 ce 73 eb 22 eb bb 9c 14 d8 e2 b4 3e a7 7e f8 87 e5 a2 64 48 f4 08 6f a8 19 a2 5e 27 72 5a c1 02 98 85 1c 8b c4 5f 2c e4 43 0b 07 91 7a de 5e a8 c4 34 c3 f2 57 6e ff be b5 21 17 37 36 e5 c9 55 74 50 64 30 68 b0 c0 fb 13 2a 7e 99 de 6c a2 92 24 6a 99 37 fa 7d 7e 06 e5 9c f5 9c e5 b9 f8 42 62 90 49 93 11 46 af 40 a8 a1 25 6b a3 73 a8 8d 09 dc 00 cd f4 45 3c c6 ba 78 57 2b f3 e1 f2 35 2a 97 8c db ad 60 22 0c b8 ac 37 d7 f6 14 a3 06 49 2a 4b 5e ee 92 44 b0 ca 84 b6 cf 2e 23 01 3b fb 1c b9 2b f6 d1 26 fe 55 0e 58 c1 9f 84 e7 a4 08 14 37 b7 5b 31 b2 b9 fb 65 8d cd 8b a0 77 96 2e 0f 33 59 72 1a 14 8d 4f ef e5 c9 79 41 ca 96 88 cb 85 ad f3 8f d1 0f 58 11 cd d8 e0 74 a2 1b bf c9 54 1c 71 46 5b 08 d7 32 12 81 b6 8e d5 2b fa b7 89 b9 c8 38 49 c0 9d 52 37 6d 41 9b 1e 7b a3 67 60 32 36 e1 19 cd f4 a7 b7 cf 9d 81 f2 22 96 01 de ac e5 3c ea 2f 14 a0 5f 7f a0 ca 04 b3 9e 31 c6 45 3e 33 2f 4b d0 91 5c 0e 09 e2 8f 4d 11 25 c3 90 c6 ff 08 33 a0 4b 6f c3 78 55 e6 5d e9 03 33 e5 05 b9 eb 66 e0 06 86 a3 ed 49 9c fb 7b 8b 21 5d bd c9 78 7b 5b aa 72 4c fa 71 ee 67 45 b4 1e 20 3d e8 b7 79 47 57 ac 32 8e c5 56 75 40 b9 51 b5 05 30 c3 d4 ee 34 70 5e a1 c6 6f d6 59 1e 88 56 10 83 e8 6d 48 c4 5e f3 b8 3a 30 29 31 9d 8f 3d 8e 65 ce 14 c1 dc 3c b9 2d 0a 7d be b6 09 a8 d2 79 39 28 ca a0 79 f0 fb bb 2b c9 0b 9f 05 8c c0 48 f4 03 20 41 d1 4c 5b ca 00 e9 9b 30 27 ec 3a 50 c4 95 71 99 cf 01 6a 45 94 06 9a f8 65 9d f0 97 3f 20 ff b1 5d bc 26 5a c5 b8 a2 20 3e 90 b1 14 a3 e9 44 1e 35 7c 60 ce 0b 55 0a 7f e6 6f c3 4f 57 02 ac 8e 89 92 a2 2e 89 19 4c 1d f6 9e 81 a9 b7 ad 3d 26 34 ea 8c 03 88 58 81 92 fd 47 d8 e8 03 b1 00 44 d5 58 61 7f b2 92 1b 69 eb 4d 85 c0 51 f8 6e f6 3a 2f 43 82 b9 be cd 87 0f b2 ec ad ca 69 02 71 2b 88 68 07 92 e2 f2 ec 89 59 1c fe bb 6d b3 ad 31 c2 a3 39 af 10 46 5f cf 79 88 51 9d 38 22 18 df 52 26 12 34 f2 6a 6f 66 ad 0d 18 59 de 50 5d 0f e8 19 ca f2 f8 d3 0a a9 fd 12 28 ac 91 d1 1c a6 7f 1f 8d 50 c8 ee fa 5c 44 15 14 32 15 07 df f6 c6 ea 3c ff 6f 34 0a 1c 11 e0 c4 0f 41 9e 8e 60 fc 94 d8 82 8f a4 7a 96 cd f7 ee 4f 61 e2 3f 40 75 1b 25 cf 9c a1 29 50 41 a3 50 f8 3f 0e 67 95 15 d6 b4 b4 6e 2c 9c e8 99 9e 07 f8 35 ab c1 66 3c ef cf 72 8d f3 78 31 f4 e1 7f 8c 8a 4f ea f1 fb f4 4c 38 c9 31 32 84 40 4a 50 ba 4c d8 ab e8 35 b3 3b fe cb 02 cd 6c 9d 7f 43 58 53 b4 c8 d5 05 ab 83 cf 46 51 27 39 11 66 94 76 56 58 ba e6 4b 31 27 15 2d 21 60 55 ae f9 b2 5c 70 a8 a3 b3 02 75 2d 7b 1e 87 91 c6 57 b9 f3 fc 90 01 ef 29 9f d1 a3 49 49 1a e6 ee 99 40 14 91 60 50 7f c4 13 0f b8 25 d4 7d 97 dc 2c 24 32 09 d2 40 35 83 ac 3e f6 dd be dc aa 76 43 22 55 48 7c 0a 06 e5 9e 04 3e 57 2a b3 ae c0 02 af 6a 6b 6a 2d c9 cf ae ef a7 05 57 88 6c 4d 12 92 4a 03 88 f2 f1 bc 8e 89 e4 cf a6 97 05 d1 ef 3c 46 58 f8 61 62 78 b5 88 01 1d 9d d9 14 be ec 0b 15 1d 65 b6 52 4f ba 3e 3f 23 5d 58 37 3e 02 16 99 b0 76 22 a5 15 04 ea de 74 7e 0b 2f 9c f3 8b c1 67 ca bc 8c b1 8c 70 8d 13 37 e2 56 48 70 7e 8c 08 72 87 65 14 c7 a4 9c 0b 2a ae e5 ed 9e 9e cf cb cc 23 e0 32 c4 de b6 3e 48 e7 12 01 88 05 64 68 fa d3 14 48 e4 b4 2e 7d 62 fb cf c1 c2 cf b0 1f cf 0d b5 e8 a1 62 bd b9 bd 82 0c 76 3f 17 b9 6c 23 f3 2d b9 d1 c1 d7 4d db e6 57 b4 f9 59 5a 97 96 98 2a 07 42 15 31 11 b1 5e 48 4d 8f fe be 47 ce 0a 78 a4 1e 47 0a 34 16 16 ec 2e b6 ee f8 13 fb 41 5f ab 50 fa 96 5b ec 6a 59 77 a0 ad e4 fe d6 7f 86 fc 24 e1 1b 0e 2f 53 64 07 9f 7c 7c 35 b5 60 ca c7 26 db c8 0f 29 db d2 48 f7 ec da 02 86 a2 3e f1 72 a2 8c 96 d9 bd b1 f5 98 ea b3 1c 6b af 03 21 31 2d a2 39 20 c0 74 b9 d2 d2 e4 42 71 7f 2c 21 00 11 42 a3 9c 2d a6 b8 bd 9d 06 b0 5a 6a 8d ea e1 bd 1b e4 10 8a 63 6d db eb c6 82 d1 13 e7 15 f2 f3 ee 50 6a bc 28 b1 c6 54 b3 d2 d2 8e 02 f7 3b 17 1e d0 de ed 71 dd a9 0c e4 a0 b7 28 cc 75 cb 57 63 85 e7 41 8b 54 5b 99 2b 1d d9 8e 2e e5 33 55 f4 64 f9 25 0a 2a 03 b3 d1 e4 d2 ac 1a a7 1d 01 de 25 73 53 03 24 e1 4d 15 a5 07 88 34 11 dd bd 37 be 21 dc 92 9d b9 b1 1f 11 01 0e 4d 2a 04 c7 32 5a 5a e6 d2 4d 19 90 0f f9 7a 8a 89 43 8f 86 76 45 7a 78 cf f0 52 01 52 8f 43 58 fe 67 f6 12 88 bf 04 2a 8f 3c 2e 0f 8b 49 97 b8 fb 74 99 6b 80 d4 65 48 9b 2d 78 07 a9 45 eb a7 2a 94 5e 8c f8 62 5d ba d6 f6 f6 e3 05 37 b2 9e 55 8c a8 a8 b0 62 5f 57 8c 76 6d 34 f2 d2 8d 70 47 15 f9 4d f1 f6 31 8a 30 80 42 aa 49 4b aa 29 56 40 67 9f 1e ab 1e 6a 83 08 af 8c e6 44 1d 5a d8 a2 f3 d4 77 eb 53 07 af 0d fa e6 64 44 93 f1 43 40 30 b8 36 16 21 34 0a d3 dc ef a6 d8 f0 90 ed 39 29 d8 9a d9 c0 ba 01 a6 90 30 33 42 8d d8 f9 86 19 30 4b 18 03 18 7d 2a 61 30 bf 1e 00 9b 5e b0 e7 e2 1c 75 59 4b 3b 84 70 f3 10 6c 92 a9 e5 5b fe b0 26 09 1c ed 12 7a 90 a1 f1 24 7f 3c 07 e3 6d 35 72 92 3e 0d e3 f7 35 18 d0 36 9a 25 fc d4 e6 5d 24 3b 7e ab 91 06 3a 7b fd 8b c8 ac 9e 39 fb bb 32 b5 c9 51 78 86 28 7a 18 ea cd 8d bf da 9b 91 db 69 3c b1 2e 42 cb a5 98 82 80 e0 9e 51 e7 2b 91 a0 f3 60 65 6d 8f 21 cc 1e b3 24 8a c3 45 ad 51 d1 a6 17 2b 18 41 92 77 85 1d ab 01 e0 28 c7 e8 e2 cf 34 09 5e fd 28 26 78 52 a0 8f cc f6 1c 45 b5 e4 69 30 16 0d af 50 28 2b e0 58 27 4e 7f f5 8c 48 b6 0b 86 d0 c8 cc 88 6a b8 77 8a 2a 2b 5f a2 55 7a d7 bf 7f 01 f2 17 55 23 ce 75 88 71 d3 34 c2 0a 98 39 f7 a8 9f e2 86 7c 06 28 9c 8a 2f 64 56 ac 7e 4f ba de e5 3a c4 1a ee 68 21 4b c7 66 69 85 3b aa 4f 58 25 5b 39 ba 58 14 14 f2 e8 c4 b4 93 03 f8 93 6a 33 ba e5 cb c9 6b 29 a5 c7 24 d5 b5 0e 16 14 14 4c 2a cd 03 bd 90 de 89 1c 3d 36 ce 04 0d 57 54 36 82 ec bf ca e6 6c 41 53 c3 37 c3 a5 d0 18 96 52 4c 8e 0c 27 e9 a0 89 19 82 1c a2 7b bf e3 fc 9b a2 4a 82 3a ee c8 d4 ad 8e f8 a6 5c 58 5b ee 4d cf 1a ca a3 da 50 1a 8c 11 a2 3b 2e 22 92 0c 1a d0 12 9a 03 8b 31 fe 16 c2 ab c8 05 89 db f7 c3 72 11 af c5 d1 a6 db 90 8b c5 fe 8a 69 20 60 06 9f cb cd cb a7 c5 23 d3 c4 24 aa 3b 0c 65 56 38 7e 0e 51 bc de 9e 5f 85 0a bf 62 d2 c2 10 1c 3a 27 86 a0 cb 94 93 28 77 a0 9c b6 b6 4a a6 1d a8 ca b3 b4 23 e7 ad c4 c4 70 04 18 a6 5e 87 22 57 10 e1 69 1f 6d 9c b2 eb 63 cc e5 b6 05 ce 0a 4a 89 cf 51 97 67 e0 08 45 39 7c 5e 38 11 41 a0 ed 8a 89 b0 10 64 b4 95 ec 8d 1e 2d a3 74 33 bf 15 97 d9 19 a6 96 10 d2 ad 26 bd f6 fc a8 de 42 2b b2 cb 80 d0 51 62 06 e8 19 4e e5 14 45 a3 dc b5 de e3 3c 0c 31 0b 47 51 e6 8e 58 be bd a2 fb 58 69 85 a5 a5 b0 64 56 75 6f 44 e6 db fe f4 bb 99 ca 73 2f 00 fb 9e a3 57 75 f7 41 96 81 bf be 6f 43 dc 7c 46 50 c1 3b 63 d9 3c 1d 49 0b f0 17 3f 28 7a 43 09 53 1f 13 ec b1 a7 75 d0 bd 88 1a} (length 0x1000)
// size: len = 0x1000 (2 bytes)
// }
// }
// }
// len: len = 0x1006 (8 bytes)
// ]
*(uint32_t*)0x200000000940 = 8;
memcpy((void*)0x200000000944, "... [truncated large byte array] ...", 4096);
*(uint16_t*)0x200000001944 = 0x1000;
syscall(__NR_write, /*fd=*/r[0], /*data=*/0x200000000940ul, /*len=*/0x1006ul);
// quotactl$Q_SYNC arguments: [
// cmd: quota_cmd_sync = 0xffffffff80000101 (8 bytes)
// special: const = 0x0 (8 bytes)
// id: const = 0x0 (8 bytes)
// addr: const = 0x0 (8 bytes)
// ]
syscall(__NR_quotactl, /*cmd=Q_SYNC_GRP*/0xffffffff80000101ul, /*special=*/0ul, /*id=*/0ul, /*addr=*/0ul);
return 0;
}
]
|
| 2/1 |
2026/05/26 05:53 |
action |
kernel-checkouter |
1m
Results: map[KernelSrc:/app/workdir/cache/src/d9debef4fa23b82f95b5d9bd809824a08680859f]
|
| 3/1 |
2026/05/26 05:54 |
action |
kernel-builder |
20m
Results: map[KernelObj:/app/workdir/cache/build/3ae46821403b49eebff5f82adc6e7d80aaf1284c]
|
| 4/1 |
2026/05/26 06:15 |
action |
codesearch-prepare |
17m
Results: map[Index:codesearch-index]
|
| 5/1 |
2026/05/26 06:33 |
agent |
expert |
8m
Model:
Results: map[DenialOfService:true Exploitable:false FilesystemTrigger:true NetworkTrigger:false PeripheralTrigger:true RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
Instruction:
You are an experienced Linux kernel security engineer. Your task is to analyze given kernel bug report
and determine its security impact based on the following dimensions.
Use the provided tools to examine the source code, check for capability checks (e.g., capable(), ns_capable()),
and understand the nature of the bug. Analyze the given kernel build and configuration.
You can check the kernel config by grepping ".config" file; you can check kernel cmdline by greeping
".config" file for "CONFIG_CMDLINE=". Assume sysctl parameters have default values.
But analyze for the corresponding production build w/o debugging tools enabled (like KASAN, KMSAN, UBSAN).
Try different strategies when analyzing the bug:
- think of ways in which the vulnerable code is unreachable
- or the other way around: try to come up with different ideas of how an unprivileged user can reach the bug
If still unsure err on the side of the bug being non-exploitable/not-accessible.
In the final reply, provide a reasoning for your assessment.
Analysis dimensions:
* Exploitable:
Determine if the bug can result in memory corruption or elevated privileges.
Memory safety issues are almost always exploitable (KASAN or UBSAN reports for use-after-free, out-of-bounds;
refcounting issues, corrupted lists, etc). When kernel is crashing on a completly wild pointer access
(e.g. user-space address, or non-canonical address, but not on NULL or address corresponding to KASAN shadow
for NULL address), including both data accesses and control tranfers, that's also usually implies possibility
of exploitation. Such reports usually say "unable to handle kernel paging request".
Uses of uninitialized values detected by KMSAN may be exploitable b/c attacker frequently can affect uninit
values with spraying techniques. However, for these exploitabability depends on how exactly the uninit value
is used in the code, and what it affects.
Think of what happens after the bug is triggered. Some bugs cause kernel panic and halt execution,
they are harder to exploit. For example, BUG reports halts the kernel. However, WARNING reports don't halt
execution in production builds. Debug bug detection tools (like KASAN, KMSAN, KCSAN, UBSAN) are also not enabled
in production builds, so attacker can freely exploit these bugs w/o being detected by these tools.
If you see an integer overflow, think how the overflowed value used later (if it's used as allocation size,
or an array index). If you see an out-of-bounds read, think if it's followed by an out-of-bounds write as well.
Some KCSAN data-races may be exploitable by skilled attackers as well. Think what data structures got corrupted
as the result of data races and how. However, note that kernel has lots of "benign" data races that don't lead
to any runtime misbehavior at all.
* Denial Of Service:
Determine if the bug can result in denial-of-service. Most bugs can, since they cause system crash,
hangs, deadlocks, or resource leaks. This is mostly applicable to WARNING bugs that won't cause system crash
in production. For these think what will be consequences of the violation of the kernel assumptions flagged
by the WARNING. In some cases the unexpected condition is also properly handled by the normal control flow
(e.g. with "if (WARN_ON(...))"), these won't cause denial-of-service. If the condition is not handled,
then it may or may not cause denial-of-service.
* Accessible From Unprivileged Processes:
Determine if the bug can be reached from a typical (non-root) user process that does NOT have any special capabilities
(like CAP_SYS_ADMIN, CAP_NET_ADMIN, CAP_NET_RAW, CAP_PERFMON) or access to device nodes restricted to root.
Assume that unprivileged_bpf_disabled=1, that is eBPF loading is not accessible. However, cBPF (classical BPF)
is still accessible to non-root processes.
Assume that user namespaces are not accessible, that is, the process cannot get the mentioned capabilities even
within a new user namespace (checked by ns_capable() function in the kernel sources).
* Accessible From User Namespaces:
Determine if the bug can be reached within a user-namespace where the process has all capabilities
(including CAP_SYS_ADMIN, CAP_NET_ADMIN, CAP_NET_RAW, CAP_PERFMON). Such capabilities are checked with ns_capable()
function in the kernel sources.
* VM Guest Trigger:
Determine if the bug can be triggered from the context of a typical KVM guest (e.g., set up by a QEMU VMM).
Consider accesses to standard Linux host paravirtualized features (virtio-blk, virtio-net, etc.),
and handling of VM exits in the KVM code.
* VM Host Trigger in The Confidetial Computing Context:
Determine if the bug can be triggered in a confidential computing guest kernel from the context of a KVM host.
Consider access to standard Linux guest paravirtualized features (virtio-blk, virtio-net, etc.).
* Ethernet Network Trigger:
Determine if the bug can be triggered by processing ingress network Ethernet traffic, either directly (network stack)
or via drivers exposed to network data.
* Other Remote Trigger:
Determine if the bug can be triggered by processing remote traffic other than Ethernet (Wifi, Bluetooth, NFC, etc).
* Peripheral Trigger:
Determine if the bug can be triggered via an untrusted peripheral device that can be physically plugged
into a system, such as a USB device or a niche hardware driver handling external hardware inputs.
This is particularly important for mobile and desktop environments where users can plug in unknown devices.
* Malicious Filesystem Trigger:
Determine if the bug can be triggered by the kernel mounting and parsing a malicious filesystem image.
This is highly critical for Desktop and Mobile environments where external media or downloaded images
might be auto-mounted.
Don't make assumptions about the kernel source code (it may be different from what you assume it is).
Extensively use the provided code access tools (codesearch-*, git-*, grepper, etc)
to examine the actual source code, and confirm any assumptions.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt:
The kernel bug report is:
INFO: task syz.2.19:6138 blocked for more than 143 seconds.
Not tainted syzkaller #0
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
task:syz.2.19 state:D stack:24520 pid:6138 tgid:6138 ppid:5969 task_flags:0x400140 flags:0x00004006
Call Trace:
<TASK>
context_switch kernel/sched/core.c:5357 [inline]
__schedule+0x16f3/0x4c20 kernel/sched/core.c:6961
__schedule_loop kernel/sched/core.c:7043 [inline]
rt_mutex_schedule+0x77/0xf0 kernel/sched/core.c:7339
rt_mutex_slowlock_block kernel/locking/rtmutex.c:1647 [inline]
__rt_mutex_slowlock kernel/locking/rtmutex.c:1721 [inline]
__rt_mutex_slowlock_locked+0x1e04/0x25e0 kernel/locking/rtmutex.c:1760
__rwbase_read_lock+0xbc/0x180 kernel/locking/rwbase_rt.c:114
rwbase_read_lock kernel/locking/rwbase_rt.c:147 [inline]
__down_read kernel/locking/rwsem.c:1466 [inline]
down_read+0x127/0x1f0 kernel/locking/rwsem.c:1539
__super_lock fs/super.c:59 [inline]
super_lock+0x2a9/0x3b0 fs/super.c:121
__iterate_supers+0x126/0x290 fs/super.c:923
quota_sync_all fs/quota/quota.c:69 [inline]
__do_sys_quotactl fs/quota/quota.c:938 [inline]
__se_sys_quotactl+0x353/0x950 fs/quota/quota.c:917
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0xfa/0x3b0 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7fa770f3eba9
RSP: 002b:00007fffd79c8678 EFLAGS: 00000246 ORIG_RAX: 00000000000000b3
RAX: ffffffffffffffda RBX: 00007fa771185fa0 RCX: 00007fa770f3eba9
RDX: 0000000000000000 RSI: 0000000000000000 RDI: ffffffff80000101
RBP: 00007fa770fc1e19 R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000
R13: 00007fa771185fa0 R14: 00007fa771185fa0 R15: 0000000000000004
</TASK>
INFO: task syz.1.18:6139 blocked for more than 143 seconds.
Not tainted syzkaller #0
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
task:syz.1.18 state:D stack:25352 pid:6139 tgid:6139 ppid:5967 task_flags:0x400140 flags:0x00004004
Call Trace:
<TASK>
context_switch kernel/sched/core.c:5357 [inline]
__schedule+0x16f3/0x4c20 kernel/sched/core.c:6961
__schedule_loop kernel/sched/core.c:7043 [inline]
rt_mutex_schedule+0x77/0xf0 kernel/sched/core.c:7339
rt_mutex_slowlock_block kernel/locking/rtmutex.c:1647 [inline]
__rt_mutex_slowlock kernel/locking/rtmutex.c:1721 [inline]
__rt_mutex_slowlock_locked+0x1e04/0x25e0 kernel/locking/rtmutex.c:1760
__rwbase_read_lock+0xbc/0x180 kernel/locking/rwbase_rt.c:114
rwbase_read_lock kernel/locking/rwbase_rt.c:147 [inline]
__down_read kernel/locking/rwsem.c:1466 [inline]
down_read+0x127/0x1f0 kernel/locking/rwsem.c:1539
__super_lock fs/super.c:59 [inline]
super_lock+0x2a9/0x3b0 fs/super.c:121
__iterate_supers+0x126/0x290 fs/super.c:923
quota_sync_all fs/quota/quota.c:69 [inline]
__do_sys_quotactl fs/quota/quota.c:938 [inline]
__se_sys_quotactl+0x353/0x950 fs/quota/quota.c:917
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0xfa/0x3b0 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7fca16aceba9
RSP: 002b:00007ffdcd5ef6b8 EFLAGS: 00000246 ORIG_RAX: 00000000000000b3
RAX: ffffffffffffffda RBX: 00007fca16d15fa0 RCX: 00007fca16aceba9
RDX: 0000000000000000 RSI: 0000000000000000 RDI: ffffffff80000101
RBP: 00007fca16b51e19 R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000
R13: 00007fca16d15fa0 R14: 00007fca16d15fa0 R15: 0000000000000004
</TASK>
INFO: task syz.3.20:6143 blocked for more than 143 seconds.
Not tainted syzkaller #0
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
task:syz.3.20 state:D stack:24520 pid:6143 tgid:6143 ppid:5973 task_flags:0x400140 flags:0x00004004
Call Trace:
<TASK>
context_switch kernel/sched/core.c:5357 [inline]
__schedule+0x16f3/0x4c20 kernel/sched/core.c:6961
__schedule_loop kernel/sched/core.c:7043 [inline]
rt_mutex_schedule+0x77/0xf0 kernel/sched/core.c:7339
rt_mutex_slowlock_block kernel/locking/rtmutex.c:1647 [inline]
__rt_mutex_slowlock kernel/locking/rtmutex.c:1721 [inline]
__rt_mutex_slowlock_locked+0x1e04/0x25e0 kernel/locking/rtmutex.c:1760
__rwbase_read_lock+0xbc/0x180 kernel/locking/rwbase_rt.c:114
rwbase_read_lock kernel/locking/rwbase_rt.c:147 [inline]
__down_read kernel/locking/rwsem.c:1466 [inline]
down_read+0x127/0x1f0 kernel/locking/rwsem.c:1539
__super_lock fs/super.c:59 [inline]
super_lock+0x2a9/0x3b0 fs/super.c:121
__iterate_supers+0x126/0x290 fs/super.c:923
quota_sync_all fs/quota/quota.c:69 [inline]
__do_sys_quotactl fs/quota/quota.c:938 [inline]
__se_sys_quotactl+0x353/0x950 fs/quota/quota.c:917
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0xfa/0x3b0 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7f98a2d0eba9
RSP: 002b:00007ffeed4a8968 EFLAGS: 00000246 ORIG_RAX: 00000000000000b3
RAX: ffffffffffffffda RBX: 00007f98a2f55fa0 RCX: 00007f98a2d0eba9
RDX: 0000000000000000 RSI: 0000000000000000 RDI: ffffffff80000101
RBP: 00007f98a2d91e19 R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000
R13: 00007f98a2f55fa0 R14: 00007f98a2f55fa0 R15: 0000000000000004
</TASK>
INFO: task syz.4.21:6145 blocked for more than 143 seconds.
Not tainted syzkaller #0
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
task:syz.4.21 state:D stack:23880 pid:6145 tgid:6145 ppid:5976 task_flags:0x400140 flags:0x00004004
Call Trace:
<TASK>
context_switch kernel/sched/core.c:5357 [inline]
__schedule+0x16f3/0x4c20 kernel/sched/core.c:6961
__schedule_loop kernel/sched/core.c:7043 [inline]
rt_mutex_schedule+0x77/0xf0 kernel/sched/core.c:7339
rt_mutex_slowlock_block kernel/locking/rtmutex.c:1647 [inline]
__rt_mutex_slowlock kernel/locking/rtmutex.c:1721 [inline]
__rt_mutex_slowlock_locked+0x1e04/0x25e0 kernel/locking/rtmutex.c:1760
__rwbase_read_lock+0xbc/0x180 kernel/locking/rwbase_rt.c:114
rwbase_read_lock kernel/locking/rwbase_rt.c:147 [inline]
__down_read kernel/locking/rwsem.c:1466 [inline]
down_read+0x127/0x1f0 kernel/locking/rwsem.c:1539
__super_lock fs/super.c:59 [inline]
super_lock+0x2a9/0x3b0 fs/super.c:121
__iterate_supers+0x126/0x290 fs/super.c:923
quota_sync_all fs/quota/quota.c:69 [inline]
__do_sys_quotactl fs/quota/quota.c:938 [inline]
__se_sys_quotactl+0x353/0x950 fs/quota/quota.c:917
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0xfa/0x3b0 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7f12cbceeba9
RSP: 002b:00007ffef7fa66c8 EFLAGS: 00000246 ORIG_RAX: 00000000000000b3
RAX: ffffffffffffffda RBX: 00007f12cbf35fa0 RCX: 00007f12cbceeba9
RDX: 0000000000000000 RSI: 0000000000000000 RDI: ffffffff80000101
RBP: 00007f12cbd71e19 R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000
R13: 00007f12cbf35fa0 R14: 00007f12cbf35fa0 R15: 0000000000000004
</TASK>
Showing all locks held in the system:
3 locks held by kworker/u8:0/12:
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3211 [inline]
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_scheduled_works+0x9b4/0x17b0 kernel/workqueue.c:3319
#1: ffffc90000117bc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3212 [inline]
#1: ffffc90000117bc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_scheduled_works+0x9ef/0x17b0 kernel/workqueue.c:3319
#2: ffff888026c160d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_trylock_shared+0x20/0xf0 fs/super.c:563
7 locks held by kworker/u8:1/13:
4 locks held by kworker/u8:2/37:
1 lock held by khungtaskd/39:
#0: ffffffff8d9a8d80 (rcu_read_lock){....}-{1:3}, at: rcu_lock_acquire include/linux/rcupdate.h:331 [inline]
#0: ffffffff8d9a8d80 (rcu_read_lock){....}-{1:3}, at: rcu_read_lock include/linux/rcupdate.h:841 [inline]
#0: ffffffff8d9a8d80 (rcu_read_lock){....}-{1:3}, at: debug_show_all_locks+0x2e/0x180 kernel/locking/lockdep.c:6775
3 locks held by kworker/u8:3/57:
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3211 [inline]
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_scheduled_works+0x9b4/0x17b0 kernel/workqueue.c:3319
#1: ffffc9000123fbc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3212 [inline]
#1: ffffc9000123fbc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_scheduled_works+0x9ef/0x17b0 kernel/workqueue.c:3319
#2: ffff8880489ba0d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_trylock_shared+0x20/0xf0 fs/super.c:563
3 locks held by kworker/u8:4/67:
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3211 [inline]
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_scheduled_works+0x9b4/0x17b0 kernel/workqueue.c:3319
#1: ffffc9000152fbc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3212 [inline]
#1: ffffc9000152fbc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_scheduled_works+0x9ef/0x17b0 kernel/workqueue.c:3319
#2: ffff8880352b20d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_trylock_shared+0x20/0xf0 fs/super.c:563
3 locks held by kworker/u8:6/170:
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3211 [inline]
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_scheduled_works+0x9b4/0x17b0 kernel/workqueue.c:3319
#1: ffffc90003bcfbc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3212 [inline]
#1: ffffc90003bcfbc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_scheduled_works+0x9ef/0x17b0 kernel/workqueue.c:3319
#2: ffff888030b000d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_trylock_shared+0x20/0xf0 fs/super.c:563
3 locks held by kworker/u8:7/1163:
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3211 [inline]
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_scheduled_works+0x9b4/0x17b0 kernel/workqueue.c:3319
#1: ffffc90004e17bc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3212 [inline]
#1: ffffc90004e17bc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_scheduled_works+0x9ef/0x17b0 kernel/workqueue.c:3319
#2: ffff88805a4880d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_trylock_shared+0x20/0xf0 fs/super.c:563
3 locks held by kworker/u8:10/1499:
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3211 [inline]
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_scheduled_works+0x9b4/0x17b0 kernel/workqueue.c:3319
#1: ffffc90005917bc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3212 [inline]
#1: ffffc90005917bc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_scheduled_works+0x9ef/0x17b0 kernel/workqueue.c:3319
#2: ffff88803d7240d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_trylock_shared+0x20/0xf0 fs/super.c:563
3 locks held by kworker/u8:11/1514:
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3211 [inline]
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_scheduled_works+0x9b4/0x17b0 kernel/workqueue.c:3319
#1: ffffc90005927bc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3212 [inline]
#1: ffffc90005927bc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_scheduled_works+0x9ef/0x17b0 kernel/workqueue.c:3319
#2: ffff88803c0320d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_trylock_shared+0x20/0xf0 fs/super.c:563
3 locks held by kworker/u8:14/3605:
3 locks held by kworker/u8:16/3632:
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3211 [inline]
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_scheduled_works+0x9b4/0x17b0 kernel/workqueue.c:3319
#1: ffffc9000d9efbc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3212 [inline]
#1: ffffc9000d9efbc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_scheduled_works+0x9ef/0x17b0 kernel/workqueue.c:3319
#2: ffff88805a60e0d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_trylock_shared+0x20/0xf0 fs/super.c:563
2 locks held by getty/5600:
#0: ffff88823bf748a0 (&tty->ldisc_sem){++++}-{0:0}, at: tty_ldisc_ref_wait+0x25/0x70 drivers/tty/tty_ldisc.c:243
#1: ffffc90003e832e0 (&ldata->atomic_read_lock){+.+.}-{4:4}, at: n_tty_read+0x444/0x1410 drivers/tty/n_tty.c:2222
2 locks held by syz-executor/5963:
1 lock held by syz.2.19/6138:
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: __super_lock fs/super.c:59 [inline]
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_lock+0x2a9/0x3b0 fs/super.c:121
1 lock held by syz.1.18/6139:
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: __super_lock fs/super.c:59 [inline]
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_lock+0x2a9/0x3b0 fs/super.c:121
3 locks held by f2fs_ckpt-7:2/6140:
1 lock held by syz.3.20/6143:
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: __super_lock fs/super.c:59 [inline]
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_lock+0x2a9/0x3b0 fs/super.c:121
1 lock held by syz.4.21/6145:
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: __super_lock fs/super.c:59 [inline]
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_lock+0x2a9/0x3b0 fs/super.c:121
6 locks held by f2fs_ckpt-7:1/6146:
3 locks held by f2fs_ckpt-7:3/6150:
3 locks held by f2fs_ckpt-7:4/6156:
1 lock held by syz.6.24/6292:
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: __super_lock fs/super.c:59 [inline]
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_lock+0x2a9/0x3b0 fs/super.c:121
1 lock held by syz.5.23/6294:
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: __super_lock fs/super.c:59 [inline]
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_lock+0x2a9/0x3b0 fs/super.c:121
2 locks held by f2fs_ckpt-7:6/6295:
2 locks held by f2fs_ckpt-7:5/6299:
1 lock held by syz.8.26/6303:
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: __super_lock fs/super.c:59 [inline]
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_lock+0x2a9/0x3b0 fs/super.c:121
1 lock held by syz.7.25/6306:
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: __super_lock fs/super.c:59 [inline]
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_lock+0x2a9/0x3b0 fs/super.c:121
4 locks held by f2fs_ckpt-7:8/6307:
3 locks held by f2fs_ckpt-7:7/6314:
1 lock held by syz.9.28/6319:
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: __super_lock fs/super.c:59 [inline]
#0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_lock+0x2a9/0x3b0 fs/super.c:121
2 locks held by f2fs_ckpt-7:9/6320:
3 locks held by kworker/u8:17/6449:
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3211 [inline]
#0: ffff8881432da138 ((wq_completion)writeback){+.+.}-{0:0}, at: process_scheduled_works+0x9b4/0x17b0 kernel/workqueue.c:3319
#1: ffffc90005a67bc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3212 [inline]
#1: ffffc90005a67bc0 ((work_completion)(&(&wb->dwork)->work)){+.+.}-{0:0}, at: process_scheduled_works+0x9ef/0x17b0 kernel/workqueue.c:3319
#2: ffff8880482820d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_trylock_shared+0x20/0xf0 fs/super.c:563
5 locks held by kworker/u8:18/6464:
1 lock held by syz.2.94/6519:
3 locks held by syz.3.96/6520:
2 locks held by syz.1.97/6522:
=============================================
NMI backtrace for cpu 0
CPU: 0 UID: 0 PID: 39 Comm: khungtaskd Not tainted syzkaller #0 PREEMPT_{RT,(full)}
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 08/18/2025
Call Trace:
<TASK>
dump_stack_lvl+0x189/0x250 lib/dump_stack.c:120
nmi_cpu_backtrace+0x39e/0x3d0 lib/nmi_backtrace.c:113
nmi_trigger_cpumask_backtrace+0x17a/0x300 lib/nmi_backtrace.c:62
trigger_all_cpu_backtrace include/linux/nmi.h:160 [inline]
check_hung_uninterruptible_tasks kernel/hung_task.c:328 [inline]
watchdog+0xf93/0xfe0 kernel/hung_task.c:491
kthread+0x70e/0x8a0 kernel/kthread.c:463
ret_from_fork+0x439/0x7d0 arch/x86/kernel/process.c:148
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
</TASK>
Sending NMI from CPU 0 to CPUs 1:
NMI backtrace for cpu 1
CPU: 1 UID: 0 PID: 13 Comm: kworker/u8:1 Not tainted syzkaller #0 PREEMPT_{RT,(full)}
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 08/18/2025
Workqueue: events_unbound cfg80211_wiphy_work
RIP: 0010:lock_acquire+0x15c/0x360 kernel/locking/lockdep.c:5871
Code: ff ff ff 65 0f c1 05 e3 7c 5a 10 83 f8 01 0f 85 c9 00 00 00 48 c7 44 24 30 00 00 00 00 9c 8f 44 24 30 f7 44 24 30 00 02 00 00 <0f> 85 cd 00 00 00 f7 44 24 08 00 02 00 00 74 01 fb 65 48 8b 05 1b
RSP: 0018:ffffc90000126b78 EFLAGS: 00000046
RAX: 0000000000000001 RBX: 0000000000000000 RCX: 6f95821df66f6e00
RDX: 0000000000000000 RSI: ffffffff8d21ad8b RDI: ffffffff8b621600
RBP: ffffffff8172c165 R08: 0000000000000000 R09: ffffffff8172c165
R10: ffffc90000126d38 R11: ffffffff81aaf310 R12: 0000000000000002
R13: ffffffff8d9a8d80 R14: 0000000000000000 R15: 0000000000000246
FS: 0000000000000000(0000) GS:ffff8881269bc000(0000) knlGS:0000000000000000
CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 00007f1ecf9a3000 CR3: 00000000325f6000 CR4: 00000000003526f0
Call Trace:
<TASK>
rcu_lock_acquire include/linux/rcupdate.h:331 [inline]
rcu_read_lock include/linux/rcupdate.h:841 [inline]
class_rcu_constructor include/linux/rcupdate.h:1155 [inline]
unwind_next_frame+0xc2/0x2390 arch/x86/kernel/unwind_orc.c:479
arch_stack_walk+0x11c/0x150 arch/x86/kernel/stacktrace.c:25
stack_trace_save+0x9c/0xe0 kernel/stacktrace.c:122
save_stack+0xf7/0x1f0 mm/page_owner.c:156
__set_page_owner+0x8d/0x490 mm/page_owner.c:329
set_page_owner include/linux/page_owner.h:32 [inline]
post_alloc_hook+0x240/0x2a0 mm/page_alloc.c:1851
prep_new_page mm/page_alloc.c:1859 [inline]
get_page_from_freelist+0x2119/0x21b0 mm/page_alloc.c:3858
__alloc_frozen_pages_noprof+0x181/0x370 mm/page_alloc.c:5148
alloc_pages_mpol+0xd1/0x380 mm/mempolicy.c:2416
alloc_slab_page mm/slub.c:2492 [inline]
allocate_slab+0x8a/0x370 mm/slub.c:2660
new_slab mm/slub.c:2714 [inline]
___slab_alloc+0x8d1/0xdc0 mm/slub.c:3901
__slab_alloc mm/slub.c:3992 [inline]
__slab_alloc_node mm/slub.c:4067 [inline]
slab_alloc_node mm/slub.c:4228 [inline]
__do_kmalloc_node mm/slub.c:4375 [inline]
__kmalloc_noprof+0x145/0x430 mm/slub.c:4388
kmalloc_noprof include/linux/slab.h:909 [inline]
kzalloc_noprof include/linux/slab.h:1039 [inline]
ieee802_11_parse_elems_full+0x152/0x2ab0 net/mac80211/parse.c:1011
ieee802_11_parse_elems_crc net/mac80211/ieee80211_i.h:2462 [inline]
ieee802_11_parse_elems net/mac80211/ieee80211_i.h:2469 [inline]
ieee80211_rx_mgmt_probe_beacon net/mac80211/ibss.c:1569 [inline]
ieee80211_ibss_rx_queued_mgmt+0x462/0x2ae0 net/mac80211/ibss.c:1600
ieee80211_iface_process_skb net/mac80211/iface.c:1699 [inline]
ieee80211_iface_work+0x85f/0x12d0 net/mac80211/iface.c:1753
cfg80211_wiphy_work+0x2b8/0x470 net/wireless/core.c:435
process_one_work kernel/workqueue.c:3236 [inline]
process_scheduled_works+0xade/0x17b0 kernel/workqueue.c:3319
worker_thread+0x8a0/0xda0 kernel/workqueue.c:3400
kthread+0x70e/0x8a0 kernel/kthread.c:463
ret_from_fork+0x439/0x7d0 arch/x86/kernel/process.c:148
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
</TASK>
It is reproducible with the followint program.
Keep in mind that it may lack the precise threading, sandboxing, and some arguments of a working reproducer.
But it should give an idea of the involved syscalls.
// autogenerated by syzkaller (https://github.com/google/syzkaller)
#define _GNU_SOURCE
#include <endian.h>
#include <errno.h>
#include <fcntl.h>
#include <setjmp.h>
#include <stdbool.h>
#include <stddef.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/ioctl.h>
#include <sys/mman.h>
#include <sys/mount.h>
#include <sys/stat.h>
#include <sys/syscall.h>
#include <sys/types.h>
#include <unistd.h>
#include <linux/loop.h>
#ifndef __NR_memfd_create
#define __NR_memfd_create 319
#endif
static unsigned long long procid;
//% This code is derived from puff.{c,h}, found in the zlib development. The
//% original files come with the following copyright notice:
//% Copyright (C) 2002-2013 Mark Adler, all rights reserved
//% version 2.3, 21 Jan 2013
//% This software is provided 'as-is', without any express or implied
//% warranty. In no event will the author be held liable for any damages
//% arising from the use of this software.
//% Permission is granted to anyone to use this software for any purpose,
//% including commercial applications, and to alter it and redistribute it
//% freely, subject to the following restrictions:
//% 1. The origin of this software must not be misrepresented; you must not
//% claim that you wrote the original software. If you use this software
//% in a product, an acknowledgment in the product documentation would be
//% appreciated but is not required.
//% 2. Altered source versions must be plainly marked as such, and must not be
//% misrepresented as being the original software.
//% 3. This notice may not be removed or altered from any source distribution.
//% Mark Adler madler@alumni.caltech.edu
//% BEGIN CODE DERIVED FROM puff.{c,h}
#define MAXBITS 15
#define MAXLCODES 286
#define MAXDCODES 30
#define MAXCODES (MAXLCODES + MAXDCODES)
#define FIXLCODES 288
struct puff_state {
unsigned char* out;
unsigned long outlen;
unsigned long outcnt;
const unsigned char* in;
unsigned long inlen;
unsigned long incnt;
int bitbuf;
int bitcnt;
jmp_buf env;
};
static int puff_bits(struct puff_state* s, int need)
{
long val = s->bitbuf;
while (s->bitcnt < need) {
if (s->incnt == s->inlen)
longjmp(s->env, 1);
val |= (long)(s->in[s->incnt++]) << s->bitcnt;
s->bitcnt += 8;
}
s->bitbuf = (int)(val >> need);
s->bitcnt -= need;
return (int)(val & ((1L << need) - 1));
}
static int puff_stored(struct puff_state* s)
{
s->bitbuf = 0;
s->bitcnt = 0;
if (s->incnt + 4 > s->inlen)
return 2;
unsigned len = s->in[s->incnt++];
len |= s->in[s->incnt++] << 8;
if (s->in[s->incnt++] != (~len & 0xff) ||
s->in[s->incnt++] != ((~len >> 8) & 0xff))
return -2;
if (s->incnt + len > s->inlen)
return 2;
if (s->outcnt + len > s->outlen)
return 1;
for (; len--; s->outcnt++, s->incnt++) {
if (s->in[s->incnt])
s->out[s->outcnt] = s->in[s->incnt];
}
return 0;
}
struct puff_huffman {
short* count;
short* symbol;
};
static int puff_decode(struct puff_state* s, const struct puff_huffman* h)
{
int first = 0;
int index = 0;
int bitbuf = s->bitbuf;
int left = s->bitcnt;
int code = first = index = 0;
int len = 1;
short* next = h->count + 1;
while (1) {
while (left--) {
code |= bitbuf & 1;
bitbuf >>= 1;
int count = *next++;
if (code - count < first) {
s->bitbuf = bitbuf;
s->bitcnt = (s->bitcnt - len) & 7;
return h->symbol[index + (code - first)];
}
index += count;
first += count;
first <<= 1;
code <<= 1;
len++;
}
left = (MAXBITS + 1) - len;
if (left == 0)
break;
if (s->incnt == s->inlen)
longjmp(s->env, 1);
bitbuf = s->in[s->incnt++];
if (left > 8)
left = 8;
}
return -10;
}
static int puff_construct(struct puff_huffman* h, const short* length, int n)
{
int len;
for (len = 0; len <= MAXBITS; len++)
h->count[len] = 0;
int symbol;
for (symbol = 0; symbol < n; symbol++)
(h->count[length[symbol]])++;
if (h->count[0] == n)
return 0;
int left = 1;
for (len = 1; len <= MAXBITS; len++) {
left <<= 1;
left -= h->count[len];
if (left < 0)
return left;
}
short offs[MAXBITS + 1];
offs[1] = 0;
for (len = 1; len < MAXBITS; len++)
offs[len + 1] = offs[len] + h->count[len];
for (symbol = 0; symbol < n; symbol++)
if (length[symbol] != 0)
h->symbol[offs[length[symbol]]++] = symbol;
return left;
}
static int puff_codes(struct puff_state* s,
const struct puff_huffman* lencode,
const struct puff_huffman* distcode)
{
static const short lens[29] = {
3, 4, 5, 6, 7, 8, 9, 10, 11, 13, 15, 17, 19, 23, 27, 31,
35, 43, 51, 59, 67, 83, 99, 115, 131, 163, 195, 227, 258};
static const short lext[29] = {
0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 2, 2, 2, 2,
3, 3, 3, 3, 4, 4, 4, 4, 5, 5, 5, 5, 0};
static const short dists[30] = {
1, 2, 3, 4, 5, 7, 9, 13, 17, 25, 33, 49, 65, 97, 129, 193,
257, 385, 513, 769, 1025, 1537, 2049, 3073, 4097, 6145,
8193, 12289, 16385, 24577};
static const short dext[30] = {
0, 0, 0, 0, 1, 1, 2, 2, 3, 3, 4, 4, 5, 5, 6, 6,
7, 7, 8, 8, 9, 9, 10, 10, 11, 11,
12, 12, 13, 13};
int symbol;
do {
symbol = puff_decode(s, lencode);
if (symbol < 0)
return symbol;
if (symbol < 256) {
if (s->outcnt == s->outlen)
return 1;
if (symbol)
s->out[s->outcnt] = symbol;
s->outcnt++;
} else if (symbol > 256) {
symbol -= 257;
if (symbol >= 29)
return -10;
int len = lens[symbol] + puff_bits(s, lext[symbol]);
symbol = puff_decode(s, distcode);
if (symbol < 0)
return symbol;
unsigned dist = dists[symbol] + puff_bits(s, dext[symbol]);
if (dist > s->outcnt)
return -11;
if (s->outcnt + len > s->outlen)
return 1;
while (len--) {
if (dist <= s->outcnt && s->out[s->outcnt - dist])
s->out[s->outcnt] = s->out[s->outcnt - dist];
s->outcnt++;
}
}
} while (symbol != 256);
return 0;
}
static int puff_fixed(struct puff_state* s)
{
static int virgin = 1;
static short lencnt[MAXBITS + 1], lensym[FIXLCODES];
static short distcnt[MAXBITS + 1], distsym[MAXDCODES];
static struct puff_huffman lencode, distcode;
if (virgin) {
lencode.count = lencnt;
lencode.symbol = lensym;
distcode.count = distcnt;
distcode.symbol = distsym;
short lengths[FIXLCODES];
int symbol;
for (symbol = 0; symbol < 144; symbol++)
lengths[symbol] = 8;
for (; symbol < 256; symbol++)
lengths[symbol] = 9;
for (; symbol < 280; symbol++)
lengths[symbol] = 7;
for (; symbol < FIXLCODES; symbol++)
lengths[symbol] = 8;
puff_construct(&lencode, lengths, FIXLCODES);
for (symbol = 0; symbol < MAXDCODES; symbol++)
lengths[symbol] = 5;
puff_construct(&distcode, lengths, MAXDCODES);
virgin = 0;
}
return puff_codes(s, &lencode, &distcode);
}
static int puff_dynamic(struct puff_state* s)
{
static const short order[19] =
{16, 17, 18, 0, 8, 7, 9, 6, 10, 5, 11, 4, 12, 3, 13, 2, 14, 1, 15};
int nlen = puff_bits(s, 5) + 257;
int ndist = puff_bits(s, 5) + 1;
int ncode = puff_bits(s, 4) + 4;
if (nlen > MAXLCODES || ndist > MAXDCODES)
return -3;
short lengths[MAXCODES];
int index;
for (index = 0; index < ncode; index++)
lengths[order[index]] = puff_bits(s, 3);
for (; index < 19; index++)
lengths[order[index]] = 0;
short lencnt[MAXBITS + 1], lensym[MAXLCODES];
struct puff_huffman lencode = {lencnt, lensym};
int err = puff_construct(&lencode, lengths, 19);
if (err != 0)
return -4;
index = 0;
while (index < nlen + ndist) {
int symbol;
int len;
symbol = puff_decode(s, &lencode);
if (symbol < 0)
return symbol;
if (symbol < 16)
lengths[index++] = symbol;
else {
len = 0;
if (symbol == 16) {
if (index == 0)
return -5;
len = lengths[index - 1];
symbol = 3 + puff_bits(s, 2);
} else if (symbol == 17)
symbol = 3 + puff_bits(s, 3);
else
symbol = 11 + puff_bits(s, 7);
if (index + symbol > nlen + ndist)
return -6;
while (symbol--)
lengths[index++] = len;
}
}
if (lengths[256] == 0)
return -9;
err = puff_construct(&lencode, lengths, nlen);
if (err && (err < 0 || nlen != lencode.count[0] + lencode.count[1]))
return -7;
short distcnt[MAXBITS + 1], distsym[MAXDCODES];
struct puff_huffman distcode = {distcnt, distsym};
err = puff_construct(&distcode, lengths + nlen, ndist);
if (err && (err < 0 || ndist != distcode.count[0] + distcode.count[1]))
return -8;
return puff_codes(s, &lencode, &distcode);
}
static int puff(
unsigned char* dest,
unsigned long* destlen,
const unsigned char* source,
unsigned long sourcelen)
{
struct puff_state s = {
.out = dest,
.outlen = *destlen,
.outcnt = 0,
.in = source,
.inlen = sourcelen,
.incnt = 0,
.bitbuf = 0,
.bitcnt = 0,
};
int err;
if (setjmp(s.env) != 0)
err = 2;
else {
int last;
do {
last = puff_bits(&s, 1);
int type = puff_bits(&s, 2);
err = type == 0 ? puff_stored(&s) : (type == 1 ? puff_fixed(&s) : (type == 2 ? puff_dynamic(&s) : -1));
if (err != 0)
break;
} while (!last);
}
*destlen = s.outcnt;
return err;
}
//% END CODE DERIVED FROM puff.{c,h}
#define ZLIB_HEADER_WIDTH 2
static int puff_zlib_to_file(const unsigned char* source, unsigned long sourcelen, int dest_fd)
{
if (sourcelen < ZLIB_HEADER_WIDTH)
return 0;
source += ZLIB_HEADER_WIDTH;
sourcelen -= ZLIB_HEADER_WIDTH;
const unsigned long max_destlen = 132 << 20;
void* ret = mmap(0, max_destlen, PROT_WRITE | PROT_READ, MAP_PRIVATE | MAP_ANON, -1, 0);
if (ret == MAP_FAILED)
return -1;
unsigned char* dest = (unsigned char*)ret;
unsigned long destlen = max_destlen;
int err = puff(dest, &destlen, source, sourcelen);
if (err) {
munmap(dest, max_destlen);
errno = -err;
return -1;
}
if (write(dest_fd, dest, destlen) != (ssize_t)destlen) {
munmap(dest, max_destlen);
return -1;
}
return munmap(dest, max_destlen);
}
static int setup_loop_device(unsigned char* data, unsigned long size, const char* loopname, int* loopfd_p)
{
int err = 0, loopfd = -1;
int memfd = syscall(__NR_memfd_create, "syzkaller", 0);
if (memfd == -1) {
err = errno;
goto error;
}
if (puff_zlib_to_file(data, size, memfd)) {
err = errno;
goto error_close_memfd;
}
loopfd = open(loopname, O_RDWR);
if (loopfd == -1) {
err = errno;
goto error_close_memfd;
}
if (ioctl(loopfd, LOOP_SET_FD, memfd)) {
if (errno != EBUSY) {
err = errno;
goto error_close_loop;
}
ioctl(loopfd, LOOP_CLR_FD, 0);
usleep(1000);
if (ioctl(loopfd, LOOP_SET_FD, memfd)) {
err = errno;
goto error_close_loop;
}
}
close(memfd);
*loopfd_p = loopfd;
return 0;
error_close_loop:
close(loopfd);
error_close_memfd:
close(memfd);
error:
errno = err;
return -1;
}
static void reset_loop_device(const char* loopname)
{
int loopfd = open(loopname, O_RDWR);
if (loopfd == -1) {
return;
}
if (ioctl(loopfd, LOOP_CLR_FD, 0)) {
}
close(loopfd);
}
static long syz_mount_image(
volatile long fsarg,
volatile long dir,
volatile long flags,
volatile long optsarg,
volatile long change_dir,
volatile unsigned long size,
volatile long image)
{
unsigned char* data = (unsigned char*)image;
int res = -1, err = 0, need_loop_device = !!size;
char* mount_opts = (char*)optsarg;
char* target = (char*)dir;
char* fs = (char*)fsarg;
char* source = NULL;
char loopname[64];
if (need_loop_device) {
int loopfd;
memset(loopname, 0, sizeof(loopname));
snprintf(loopname, sizeof(loopname), "/dev/loop%llu", procid);
if (setup_loop_device(data, size, loopname, &loopfd) == -1)
return -1;
close(loopfd);
source = loopname;
}
mkdir(target, 0777);
char opts[256];
memset(opts, 0, sizeof(opts));
if (strlen(mount_opts) > (sizeof(opts) - 32)) {
}
strncpy(opts, mount_opts, sizeof(opts) - 32);
if (strcmp(fs, "iso9660") == 0) {
flags |= MS_RDONLY;
} else if (strncmp(fs, "ext", 3) == 0) {
bool has_remount_ro = false;
char* remount_ro_start = strstr(opts, "errors=remount-ro");
if (remount_ro_start != NULL) {
char after = *(remount_ro_start + strlen("errors=remount-ro"));
char before = remount_ro_start == opts ? '\0' : *(remount_ro_start - 1);
has_remount_ro = ((before == '\0' || before == ',') && (after == '\0' || after == ','));
}
if (strstr(opts, "errors=panic") || !has_remount_ro)
strcat(opts, ",errors=continue");
} else if (strcmp(fs, "xfs") == 0) {
strcat(opts, ",nouuid");
} else if (strncmp(fs, "gfs2", 4) == 0 && (strstr(opts, "errors=panic") || strstr(opts, "debug"))) {
strcat(opts, ",errors=withdraw");
}
res = mount(source, target, fs, flags, opts);
if (res == -1) {
err = errno;
goto error_clear_loop;
}
res = open(target, O_RDONLY | O_DIRECTORY);
if (res == -1) {
err = errno;
goto error_clear_loop;
}
if (change_dir) {
res = chdir(target);
if (res == -1) {
err = errno;
}
}
error_clear_loop:
if (need_loop_device)
reset_loop_device(loopname);
errno = err;
return res;
}
uint64_t r[1] = {0xffffffffffffffff};
int main(void)
{
syscall(__NR_mmap, /*addr=*/0x1ffffffff000ul, /*len=*/0x1000ul, /*prot=*/0ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
syscall(__NR_mmap, /*addr=*/0x200000000000ul, /*len=*/0x1000000ul, /*prot=PROT_WRITE|PROT_READ|PROT_EXEC*/7ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
syscall(__NR_mmap, /*addr=*/0x200001000000ul, /*len=*/0x1000ul, /*prot=*/0ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
const char* reason;
(void)reason;
intptr_t res = 0;
if (write(1, "executing program\n", sizeof("executing program\n") - 1)) {}
// syz_mount_image$f2fs arguments: [
// fs: ptr[in, buffer] {
// buffer: {66 32 66 73 00} (length 0x5)
// }
// dir: ptr[in, buffer] {
// buffer: {2e 2f 66 69 6c 65 30 00} (length 0x8)
// }
// flags: mount_flags = 0x8 (8 bytes)
// opts: ptr[in, fs_options[f2fs_options]] {
// fs_options[f2fs_options] {
// elems: array[fs_opt_elem[f2fs_options]] {
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// nodiscard: buffer: {6e 6f 64 69 73 63 61 72 64} (length 0x9)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// noinline_dentry: buffer: {6e 6f 69 6e 6c 69 6e 65 5f 64 65 6e 74 72 79} (length 0xf)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// alloc_mode_def: buffer: {61 6c 6c 6f 63 5f 6d 6f 64 65 3d 64 65 66 61 75 6c 74} (length 0x12)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// six_active_logs: buffer: {61 63 74 69 76 65 5f 6c 6f 67 73 3d 36} (length 0xd)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// quota: buffer: {71 75 6f 74 61} (length 0x5)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// fault_injection: fs_opt["fault_injection", fmt[dec, int32]] {
// name: buffer: {66 61 75 6c 74 5f 69 6e 6a 65 63 74 69 6f 6e} (length 0xf)
// eq: const = 0x3d (1 bytes)
// val: int32 = 0x10 (20 bytes)
// }
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// discard: buffer: {64 69 73 63 61 72 64} (length 0x7)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// noacl: buffer: {6e 6f 61 63 6c} (length 0x5)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// fsync_mode_posix: buffer: {66 73 79 6e 63 5f 6d 6f 64 65 3d 70 6f 73 69 78} (length 0x10)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// alloc_mode_def: buffer: {61 6c 6c 6f 63 5f 6d 6f 64 65 3d 64 65 66 61 75 6c 74} (length 0x12)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// noextent_cache: buffer: {6e 6f 65 78 74 65 6e 74 5f 63 61 63 68 65} (length 0xe)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// two_active_logs: buffer: {61 63 74 69 76 65 5f 6c 6f 67 73 3d 32} (length 0xd)
// }
// comma: const = 0x2c (1 bytes)
// }
// fs_opt_elem[f2fs_options] {
// elem: union f2fs_options {
// fault_type: fs_opt["fault_type", fmt[dec, int32]] {
// name: buffer: {66 61 75 6c 74 5f 74 79 70 65} (length 0xa)
// eq: const = 0x3d (1 bytes)
// val: int32 = 0xfffffe (20 bytes)
// }
// }
// comma: const = 0x2c (1 bytes)
// }
// }
// common: array[fs_opt_elem[fs_options_common]] {
// }
// null: const = 0x0 (1 bytes)
// }
// }
// chdir: int8 = 0x1 (1 bytes)
// size: len = 0x5505 (8 bytes)
// img: ptr[in, buffer] {
// buffer: (compressed buffer with length 0x5505)
// }
// ]
// returns fd_dir
memcpy((void*)0x200000000000, "f2fs\000", 5);
memcpy((void*)0x200000000040, "./file0\000", 8);
memcpy((void*)0x200000000240, "nodiscard", 9);
*(uint8_t*)0x200000000249 = 0x2c;
memcpy((void*)0x20000000024a, "noinline_dentry", 15);
*(uint8_t*)0x200000000259 = 0x2c;
memcpy((void*)0x20000000025a, "alloc_mode=default", 18);
*(uint8_t*)0x20000000026c = 0x2c;
memcpy((void*)0x20000000026d, "active_logs=6", 13);
*(uint8_t*)0x20000000027a = 0x2c;
memcpy((void*)0x20000000027b, "quota", 5);
*(uint8_t*)0x200000000280 = 0x2c;
memcpy((void*)0x200000000281, "fault_injection", 15);
*(uint8_t*)0x200000000290 = 0x3d;
sprintf((char*)0x200000000291, "%020llu", (long long)0x10);
*(uint8_t*)0x2000000002a5 = 0x2c;
memcpy((void*)0x2000000002a6, "discard", 7);
*(uint8_t*)0x2000000002ad = 0x2c;
memcpy((void*)0x2000000002ae, "noacl", 5);
*(uint8_t*)0x2000000002b3 = 0x2c;
memcpy((void*)0x2000000002b4, "fsync_mode=posix", 16);
*(uint8_t*)0x2000000002c4 = 0x2c;
memcpy((void*)0x2000000002c5, "alloc_mode=default", 18);
*(uint8_t*)0x2000000002d7 = 0x2c;
memcpy((void*)0x2000000002d8, "noextent_cache", 14);
*(uint8_t*)0x2000000002e6 = 0x2c;
memcpy((void*)0x2000000002e7, "active_logs=2", 13);
*(uint8_t*)0x2000000002f4 = 0x2c;
memcpy((void*)0x2000000002f5, "fault_type", 10);
*(uint8_t*)0x2000000002ff = 0x3d;
sprintf((char*)0x200000000300, "%020llu", (long long)0xfffffe);
*(uint8_t*)0x200000000314 = 0x2c;
*(uint8_t*)0x200000000315 = 0;
memcpy((void*)0x200000000340, "... [truncated large byte array] ...", 21765);
syz_mount_image(/*fs=*/0x200000000000, /*dir=*/0x200000000040, /*flags=MS_NOEXEC*/8, /*opts=*/0x200000000240, /*chdir=*/1, /*size=*/0x5505, /*img=*/0x200000000340);
// openat$cgroup_ro arguments: [
// fd: fd_cgroup (resource)
// file: ptr[in, buffer] {
// buffer: {63 70 75 2e 73 74 61 74 00} (length 0x9)
// }
// flags: const = 0x275a (4 bytes)
// mode: const = 0x0 (2 bytes)
// ]
// returns fd
memcpy((void*)0x200000000040, "cpu.stat\000", 9);
syscall(__NR_openat, /*fd=*/0xffffff9c, /*file=*/0x200000000040ul, /*flags=*/0x275a, /*mode=*/0);
// openat arguments: [
// fd: fd_dir (resource)
// file: ptr[in, buffer] {
// buffer: {2e 2f 66 69 6c 65 31 00} (length 0x8)
// }
// flags: open_flags = 0x105042 (4 bytes)
// mode: open_mode = 0x0 (2 bytes)
// ]
// returns fd
memcpy((void*)0x200000000080, "./file1\000", 8);
syscall(__NR_openat, /*fd=*/0xffffff9c, /*file=*/0x200000000080ul, /*flags=O_SYNC|O_DIRECT|O_CREAT|O_RDWR*/0x105042, /*mode=*/0);
// openat$cgroup_ro arguments: [
// fd: fd_cgroup (resource)
// file: ptr[in, buffer] {
// buffer: {63 70 75 2e 73 74 61 74 00} (length 0x9)
// }
// flags: const = 0x275a (4 bytes)
// mode: const = 0x0 (2 bytes)
// ]
// returns fd
memcpy((void*)0x200000000040, "cpu.stat\000", 9);
res = syscall(__NR_openat, /*fd=*/0xffffff9c, /*file=*/0x200000000040ul, /*flags=*/0x275a, /*mode=*/0);
if (res != -1)
r[0] = res;
// write$UHID_INPUT arguments: [
// fd: fd_uhid (resource)
// data: ptr[in, uhid_req[UHID_INPUT, uhid_input_req]] {
// uhid_req[UHID_INPUT, uhid_input_req] {
// type: const = 0x8 (4 bytes)
// data: uhid_input_req {
// data: buffer: {85 f0 80 a4 93 3d 55 26 6e 07 e7 99 aa 0c c4 21 38 82 42 df 2a 3c 6b 63 1b 65 b1 c0 61 ed d2 aa 10 8c 35 28 fe 9b 0b b3 a5 3a b1 20 0f 5d 01 a6 8a 4a cd ec 8f ee 09 64 82 22 f9 08 c1 fe dc 30 00 34 2e 61 39 de 28 36 6c 13 50 93 06 d0 0e bc c6 74 97 18 1a c9 16 db 98 af 9d 36 6b 76 e4 27 d9 ab 5b b6 80 95 f0 fb 24 6d f3 2b 8a f0 78 36 53 13 6f 8a 04 c0 36 90 31 21 25 c7 de d6 a2 4f da 86 85 34 0c 57 5e ad 69 51 9e 35 83 f8 9d 46 7e c2 32 d6 a1 ff d0 46 3b a4 ea 3c ba e5 da e6 65 4b 55 47 b5 45 8f 02 ac 30 77 29 e5 7b 09 e1 34 f6 8b e4 4f 88 d7 25 17 b2 30 b0 66 f6 31 5b 5f b8 02 06 39 7b bf f8 cb c2 a3 6e 01 c2 e7 b3 aa db 32 bd 3d d5 28 8a 69 a9 91 d9 c6 74 71 7e 3a bb a7 16 72 80 b2 db 3b 1b 85 02 af a4 f3 f2 96 c5 32 51 0c 9d 2d d7 9b b5 ee b2 5a db 5e dd db dd 06 9c 09 d1 4d 15 c2 e7 e1 e2 bd 6c 10 8f ab 35 91 bb 22 e9 7d 69 92 23 6d 22 73 c8 bb 95 53 6f 71 18 d0 07 96 50 08 b1 25 c7 da ac 28 14 e6 bb e1 ad bf a3 57 2a d0 b7 ad 5c 26 c8 01 41 18 d8 37 4c a9 f2 85 77 9d fe e7 71 5a 40 39 08 14 6a 74 de 61 b3 85 39 14 c8 9f 44 4c 12 e7 a3 8b dd 46 c4 ed 36 eb 80 6e a5 98 f4 4d 1d ec 9e ff 9e 24 76 f4 38 02 21 1f 07 62 b6 66 73 b4 5d 23 6b 23 91 ce 32 2e 30 fb 9c 69 fe 0d 51 4d c1 f8 b6 e3 97 9c 12 05 fd 52 24 b0 7d 18 a4 4f ec 4f 6f 1a 6f 65 15 8b b6 ad cc 29 5b f2 dd 7d ea 10 7f 59 d7 e0 3c 61 fe 58 22 29 2e 45 96 89 56 b9 31 bd c4 d6 44 5f f1 63 1e 0b 98 e4 b4 44 87 74 dd 4b 9c d5 3a 45 89 6f db 3f 03 70 27 78 74 1a e2 b4 5a 25 bf 9a 23 fc 02 fb 97 a6 30 f1 32 bf 9d ef 6c 6d 4a 7b ae b6 29 72 f1 a8 14 f6 f2 37 7b cf c7 8e 2e 86 36 8c 13 85 10 a0 4c ed f7 17 5a f8 c2 03 4f ae 74 13 e3 ac e8 c7 1a b9 a0 af 1c a7 04 20 11 a6 ed 02 8e 20 56 48 53 5d ab f3 b2 f8 51 96 ae 18 d3 6b 83 9e 3c d5 4a e4 93 3a d5 29 88 8f da c7 bb 8a 70 c7 2b c0 fc 81 ba 06 50 6f 2d 5b c7 68 6e 21 9b be 52 83 95 9c be f9 95 0e 07 1c b6 d9 f3 41 fc 62 4a 51 10 34 1f 26 ce bd 71 00 59 9a 06 e6 1f 66 fa e1 20 c7 fc 2b 34 c6 22 12 00 eb a7 5b d1 27 71 14 67 1a 3f a8 f0 58 b2 7f d8 97 b0 52 f4 a5 2a fc ea 81 4d f5 26 18 1c 75 c4 49 72 10 a2 b8 b7 4e 26 60 15 61 e7 87 35 38 7c f1 23 65 4b 02 95 d1 d6 05 56 95 6b 36 d9 6d d0 38 86 6c 4b 4d b3 1e bd cd dd 69 29 bb c2 85 0c d4 90 13 89 e6 ea 6e 86 04 1e 0e fa 11 58 f3 34 e7 af da 0e 11 c2 fb 0e 6d f6 36 4c b9 56 59 f5 06 d5 c7 e6 3f b6 7c 81 16 57 7d 15 e4 a4 b1 fc 4c 27 de 2e 52 58 6c b1 f5 2b e9 c3 60 1f 50 66 54 9d e8 bd c3 ec 07 d1 a8 4c af 19 61 32 3e c2 48 7a 37 b7 51 ae ab af cd 64 7c e2 da e5 d9 49 9c 0f 96 94 67 e6 ca ba d1 98 66 9a c9 6b d1 48 89 54 ef f0 85 4e e0 c8 3d 7b 59 6d 27 36 25 bd b1 62 70 78 23 21 07 1f da 5d 98 0d ed 78 ff a9 dc 2b 56 03 7d 7c bf 94 25 47 f4 8a 51 31 f1 99 1f 6c 17 ae 1e d5 12 0c a6 87 8f 98 e6 8e 79 97 a9 a2 b7 0b e6 40 a7 0a 34 ad b8 0d e2 86 c6 69 2a bb 5f 09 2e 4e 3a 15 a8 32 17 e0 3d 02 a4 05 4f 34 af 3a 65 ff 6b 36 f3 95 b7 6a 05 79 cf fa fd 5d 3b b0 e7 04 c9 35 ca ec f3 a7 ab 75 6c 23 fd 60 c9 fe 3f 4f b2 be 75 04 f5 ba e2 2b 11 6f f1 58 8d cf 02 b3 27 d3 1b f0 48 8d ba 8a f5 b3 3c cf 2d 7d 87 f4 3b bc 48 fc d4 f1 91 ad 6a f9 31 3a d3 8b 8b 29 67 4b fa bd 66 51 bc 1f 6c e5 ab b4 a2 f1 41 31 94 f9 6b 26 d7 d6 ed c4 e0 13 fb a5 49 07 5c 97 ee f5 08 af 5c a7 87 36 64 b0 58 b7 bc f4 55 a8 a0 4b 59 1d 29 fa b6 36 6c 84 4b b7 55 76 ba c2 d5 23 23 e7 47 30 3d 00 a5 73 6c 98 12 92 2b 0e 17 bc ec 91 35 55 07 36 b5 4c f6 40 7d 61 e2 2e 62 d7 bb 75 f6 29 35 b6 65 ac f3 3e 75 f6 88 c3 6e f4 16 f1 b8 90 d0 f0 c8 ad 1d f0 0e 02 ec 45 96 78 34 d5 64 9c 8e 71 43 97 86 22 fa 37 04 67 29 70 b7 99 3a 87 e9 7d 3d 92 6a 14 26 56 47 bc 8b 8c 9e 6f 83 e2 95 72 60 8d 24 b4 2c 26 35 ef 4a bb d0 af 83 86 0e 99 c9 0d 74 71 cf 6e 8c e9 95 07 f5 ec 2b c5 72 21 2f a9 ee 3f 5a 9d fa 38 15 fe 55 f0 bb b1 19 ac ce 06 2a e3 7f 2f f9 21 70 7a bb a1 39 bc dd f4 2b fd 17 4d 29 b5 40 16 1b 41 13 c4 e1 a1 3f 3a 62 8c 63 8e c4 d3 a8 84 df bc 09 3e 23 ec 0d 06 71 b4 6b 41 dc 8b 42 d9 50 c8 61 5b a5 ee 87 f4 9b 5d 09 10 ff a4 87 12 07 99 50 01 92 0d b0 5a 95 19 99 67 f0 97 ba 7b 55 bb d2 71 d8 18 69 0c 42 38 40 6b 40 a3 df c4 2f a5 6a 67 17 3b 53 a9 6b 54 33 26 c5 67 38 b6 d0 43 19 59 34 01 86 96 f5 ab 49 34 7e 51 48 a7 8f 2d 13 69 a7 1a fa b8 33 02 73 d4 6e cf ba 4e e0 58 02 a5 38 56 49 85 1d b9 49 db fb 39 e2 90 94 16 41 c5 0b 1a c2 0f b3 10 27 54 a7 60 b0 97 f4 64 dd b0 b8 3f 81 68 ba df a7 1d b6 62 1d cf 22 fb 08 1e 34 03 f3 ba c5 c7 e6 59 05 ac a5 28 85 c8 07 f8 dd ab 18 bb 2f 12 ef 95 2c 50 48 3c 0e 25 19 68 bc 70 ff 0d 42 a6 38 ca 74 4d ea 4c 7e bb 4f ea 77 7c f6 63 bb 4f 15 05 ed 79 73 0c 45 bc 86 e4 88 a1 3f 92 43 77 a8 e2 ee 66 70 a0 2c a5 28 74 ae 1c 42 a3 5d 55 b9 76 57 57 04 7b 2c c3 74 2a a5 1f a3 e4 3f b2 c1 13 c9 2a d2 13 ba d2 52 c1 a8 29 66 dd 01 6f 12 a7 f1 c3 90 0c 0f 1a b4 55 03 51 63 f3 18 99 bd d3 0f 3f f4 3a d1 7d 9e 45 bb 74 38 c1 c9 86 71 27 36 f2 4b e1 4f 71 ab 1b fe 92 a2 5e c0 7f 08 6e e8 c7 97 1b 80 77 a1 3e 58 a8 e8 be a3 9c 8e 06 b2 51 90 9f 02 cb 00 80 ab f0 20 f2 7c a1 60 eb 26 c0 82 dd a1 fa 54 ea 40 94 df db cb 2f a7 bd dc cb 67 a8 44 e8 07 5f 4c c0 8d ad 35 75 70 06 d0 51 e1 83 dc ed 33 6b c0 c2 50 2f 93 ff c8 7d ca 62 22 86 ba 17 4c 24 e1 f5 3f 27 dc 27 77 ba af e1 70 34 8b 0e 8d 3e 74 3b 3a a9 06 bc 07 64 bb e7 da 08 ff 40 3e fe 22 12 62 7d 67 22 50 65 8b b5 13 b7 31 25 17 d1 f8 8c 61 c7 ba 5f 96 47 cd 61 92 81 c5 b3 90 b4 86 06 ee 39 fb 41 71 10 3d f2 e0 9d 7c fd 56 c0 6c 72 1f 7c 24 ad 8c ce 38 36 23 fc 2d cb 15 ac 56 43 8e a3 31 82 0a e5 9c 8c 47 4e 36 fc 73 f7 b1 b3 b8 6d f1 b4 24 90 81 55 13 68 1a af bf 7e 87 1b 4b 96 86 ef ae 6c 45 ec fc a6 0a 64 0a 6f 07 1d fd 31 f9 43 7c 3d 03 08 61 64 b4 8c 1e d8 02 98 68 64 bf e0 d4 9b dd 77 09 66 22 62 36 8d bc 3e cc 05 eb 24 0e cc 41 90 4c 76 d7 8a b5 c5 2b 66 af 5a 72 0f dd 6a 92 f5 2b e0 67 64 27 a5 6e 32 e5 bc 50 85 b2 5f 90 ad d2 8a 76 f2 fc e6 f8 f0 ef 74 f4 65 96 98 54 96 46 bd 63 17 5a df 77 b5 cd cf e6 76 e1 b1 a9 af 15 10 29 46 55 4b a6 13 6c bc 83 c6 26 8e e4 03 18 f3 c9 d4 71 80 25 68 8b 35 d2 26 5b f6 0b f8 89 ff 62 9f 78 34 58 6e f4 6e ab 7a 91 76 33 75 36 bb 60 01 e6 76 54 6b 98 7f 36 b1 fe 4b 9f 6e 46 a8 ce 73 eb 22 eb bb 9c 14 d8 e2 b4 3e a7 7e f8 87 e5 a2 64 48 f4 08 6f a8 19 a2 5e 27 72 5a c1 02 98 85 1c 8b c4 5f 2c e4 43 0b 07 91 7a de 5e a8 c4 34 c3 f2 57 6e ff be b5 21 17 37 36 e5 c9 55 74 50 64 30 68 b0 c0 fb 13 2a 7e 99 de 6c a2 92 24 6a 99 37 fa 7d 7e 06 e5 9c f5 9c e5 b9 f8 42 62 90 49 93 11 46 af 40 a8 a1 25 6b a3 73 a8 8d 09 dc 00 cd f4 45 3c c6 ba 78 57 2b f3 e1 f2 35 2a 97 8c db ad 60 22 0c b8 ac 37 d7 f6 14 a3 06 49 2a 4b 5e ee 92 44 b0 ca 84 b6 cf 2e 23 01 3b fb 1c b9 2b f6 d1 26 fe 55 0e 58 c1 9f 84 e7 a4 08 14 37 b7 5b 31 b2 b9 fb 65 8d cd 8b a0 77 96 2e 0f 33 59 72 1a 14 8d 4f ef e5 c9 79 41 ca 96 88 cb 85 ad f3 8f d1 0f 58 11 cd d8 e0 74 a2 1b bf c9 54 1c 71 46 5b 08 d7 32 12 81 b6 8e d5 2b fa b7 89 b9 c8 38 49 c0 9d 52 37 6d 41 9b 1e 7b a3 67 60 32 36 e1 19 cd f4 a7 b7 cf 9d 81 f2 22 96 01 de ac e5 3c ea 2f 14 a0 5f 7f a0 ca 04 b3 9e 31 c6 45 3e 33 2f 4b d0 91 5c 0e 09 e2 8f 4d 11 25 c3 90 c6 ff 08 33 a0 4b 6f c3 78 55 e6 5d e9 03 33 e5 05 b9 eb 66 e0 06 86 a3 ed 49 9c fb 7b 8b 21 5d bd c9 78 7b 5b aa 72 4c fa 71 ee 67 45 b4 1e 20 3d e8 b7 79 47 57 ac 32 8e c5 56 75 40 b9 51 b5 05 30 c3 d4 ee 34 70 5e a1 c6 6f d6 59 1e 88 56 10 83 e8 6d 48 c4 5e f3 b8 3a 30 29 31 9d 8f 3d 8e 65 ce 14 c1 dc 3c b9 2d 0a 7d be b6 09 a8 d2 79 39 28 ca a0 79 f0 fb bb 2b c9 0b 9f 05 8c c0 48 f4 03 20 41 d1 4c 5b ca 00 e9 9b 30 27 ec 3a 50 c4 95 71 99 cf 01 6a 45 94 06 9a f8 65 9d f0 97 3f 20 ff b1 5d bc 26 5a c5 b8 a2 20 3e 90 b1 14 a3 e9 44 1e 35 7c 60 ce 0b 55 0a 7f e6 6f c3 4f 57 02 ac 8e 89 92 a2 2e 89 19 4c 1d f6 9e 81 a9 b7 ad 3d 26 34 ea 8c 03 88 58 81 92 fd 47 d8 e8 03 b1 00 44 d5 58 61 7f b2 92 1b 69 eb 4d 85 c0 51 f8 6e f6 3a 2f 43 82 b9 be cd 87 0f b2 ec ad ca 69 02 71 2b 88 68 07 92 e2 f2 ec 89 59 1c fe bb 6d b3 ad 31 c2 a3 39 af 10 46 5f cf 79 88 51 9d 38 22 18 df 52 26 12 34 f2 6a 6f 66 ad 0d 18 59 de 50 5d 0f e8 19 ca f2 f8 d3 0a a9 fd 12 28 ac 91 d1 1c a6 7f 1f 8d 50 c8 ee fa 5c 44 15 14 32 15 07 df f6 c6 ea 3c ff 6f 34 0a 1c 11 e0 c4 0f 41 9e 8e 60 fc 94 d8 82 8f a4 7a 96 cd f7 ee 4f 61 e2 3f 40 75 1b 25 cf 9c a1 29 50 41 a3 50 f8 3f 0e 67 95 15 d6 b4 b4 6e 2c 9c e8 99 9e 07 f8 35 ab c1 66 3c ef cf 72 8d f3 78 31 f4 e1 7f 8c 8a 4f ea f1 fb f4 4c 38 c9 31 32 84 40 4a 50 ba 4c d8 ab e8 35 b3 3b fe cb 02 cd 6c 9d 7f 43 58 53 b4 c8 d5 05 ab 83 cf 46 51 27 39 11 66 94 76 56 58 ba e6 4b 31 27 15 2d 21 60 55 ae f9 b2 5c 70 a8 a3 b3 02 75 2d 7b 1e 87 91 c6 57 b9 f3 fc 90 01 ef 29 9f d1 a3 49 49 1a e6 ee 99 40 14 91 60 50 7f c4 13 0f b8 25 d4 7d 97 dc 2c 24 32 09 d2 40 35 83 ac 3e f6 dd be dc aa 76 43 22 55 48 7c 0a 06 e5 9e 04 3e 57 2a b3 ae c0 02 af 6a 6b 6a 2d c9 cf ae ef a7 05 57 88 6c 4d 12 92 4a 03 88 f2 f1 bc 8e 89 e4 cf a6 97 05 d1 ef 3c 46 58 f8 61 62 78 b5 88 01 1d 9d d9 14 be ec 0b 15 1d 65 b6 52 4f ba 3e 3f 23 5d 58 37 3e 02 16 99 b0 76 22 a5 15 04 ea de 74 7e 0b 2f 9c f3 8b c1 67 ca bc 8c b1 8c 70 8d 13 37 e2 56 48 70 7e 8c 08 72 87 65 14 c7 a4 9c 0b 2a ae e5 ed 9e 9e cf cb cc 23 e0 32 c4 de b6 3e 48 e7 12 01 88 05 64 68 fa d3 14 48 e4 b4 2e 7d 62 fb cf c1 c2 cf b0 1f cf 0d b5 e8 a1 62 bd b9 bd 82 0c 76 3f 17 b9 6c 23 f3 2d b9 d1 c1 d7 4d db e6 57 b4 f9 59 5a 97 96 98 2a 07 42 15 31 11 b1 5e 48 4d 8f fe be 47 ce 0a 78 a4 1e 47 0a 34 16 16 ec 2e b6 ee f8 13 fb 41 5f ab 50 fa 96 5b ec 6a 59 77 a0 ad e4 fe d6 7f 86 fc 24 e1 1b 0e 2f 53 64 07 9f 7c 7c 35 b5 60 ca c7 26 db c8 0f 29 db d2 48 f7 ec da 02 86 a2 3e f1 72 a2 8c 96 d9 bd b1 f5 98 ea b3 1c 6b af 03 21 31 2d a2 39 20 c0 74 b9 d2 d2 e4 42 71 7f 2c 21 00 11 42 a3 9c 2d a6 b8 bd 9d 06 b0 5a 6a 8d ea e1 bd 1b e4 10 8a 63 6d db eb c6 82 d1 13 e7 15 f2 f3 ee 50 6a bc 28 b1 c6 54 b3 d2 d2 8e 02 f7 3b 17 1e d0 de ed 71 dd a9 0c e4 a0 b7 28 cc 75 cb 57 63 85 e7 41 8b 54 5b 99 2b 1d d9 8e 2e e5 33 55 f4 64 f9 25 0a 2a 03 b3 d1 e4 d2 ac 1a a7 1d 01 de 25 73 53 03 24 e1 4d 15 a5 07 88 34 11 dd bd 37 be 21 dc 92 9d b9 b1 1f 11 01 0e 4d 2a 04 c7 32 5a 5a e6 d2 4d 19 90 0f f9 7a 8a 89 43 8f 86 76 45 7a 78 cf f0 52 01 52 8f 43 58 fe 67 f6 12 88 bf 04 2a 8f 3c 2e 0f 8b 49 97 b8 fb 74 99 6b 80 d4 65 48 9b 2d 78 07 a9 45 eb a7 2a 94 5e 8c f8 62 5d ba d6 f6 f6 e3 05 37 b2 9e 55 8c a8 a8 b0 62 5f 57 8c 76 6d 34 f2 d2 8d 70 47 15 f9 4d f1 f6 31 8a 30 80 42 aa 49 4b aa 29 56 40 67 9f 1e ab 1e 6a 83 08 af 8c e6 44 1d 5a d8 a2 f3 d4 77 eb 53 07 af 0d fa e6 64 44 93 f1 43 40 30 b8 36 16 21 34 0a d3 dc ef a6 d8 f0 90 ed 39 29 d8 9a d9 c0 ba 01 a6 90 30 33 42 8d d8 f9 86 19 30 4b 18 03 18 7d 2a 61 30 bf 1e 00 9b 5e b0 e7 e2 1c 75 59 4b 3b 84 70 f3 10 6c 92 a9 e5 5b fe b0 26 09 1c ed 12 7a 90 a1 f1 24 7f 3c 07 e3 6d 35 72 92 3e 0d e3 f7 35 18 d0 36 9a 25 fc d4 e6 5d 24 3b 7e ab 91 06 3a 7b fd 8b c8 ac 9e 39 fb bb 32 b5 c9 51 78 86 28 7a 18 ea cd 8d bf da 9b 91 db 69 3c b1 2e 42 cb a5 98 82 80 e0 9e 51 e7 2b 91 a0 f3 60 65 6d 8f 21 cc 1e b3 24 8a c3 45 ad 51 d1 a6 17 2b 18 41 92 77 85 1d ab 01 e0 28 c7 e8 e2 cf 34 09 5e fd 28 26 78 52 a0 8f cc f6 1c 45 b5 e4 69 30 16 0d af 50 28 2b e0 58 27 4e 7f f5 8c 48 b6 0b 86 d0 c8 cc 88 6a b8 77 8a 2a 2b 5f a2 55 7a d7 bf 7f 01 f2 17 55 23 ce 75 88 71 d3 34 c2 0a 98 39 f7 a8 9f e2 86 7c 06 28 9c 8a 2f 64 56 ac 7e 4f ba de e5 3a c4 1a ee 68 21 4b c7 66 69 85 3b aa 4f 58 25 5b 39 ba 58 14 14 f2 e8 c4 b4 93 03 f8 93 6a 33 ba e5 cb c9 6b 29 a5 c7 24 d5 b5 0e 16 14 14 4c 2a cd 03 bd 90 de 89 1c 3d 36 ce 04 0d 57 54 36 82 ec bf ca e6 6c 41 53 c3 37 c3 a5 d0 18 96 52 4c 8e 0c 27 e9 a0 89 19 82 1c a2 7b bf e3 fc 9b a2 4a 82 3a ee c8 d4 ad 8e f8 a6 5c 58 5b ee 4d cf 1a ca a3 da 50 1a 8c 11 a2 3b 2e 22 92 0c 1a d0 12 9a 03 8b 31 fe 16 c2 ab c8 05 89 db f7 c3 72 11 af c5 d1 a6 db 90 8b c5 fe 8a 69 20 60 06 9f cb cd cb a7 c5 23 d3 c4 24 aa 3b 0c 65 56 38 7e 0e 51 bc de 9e 5f 85 0a bf 62 d2 c2 10 1c 3a 27 86 a0 cb 94 93 28 77 a0 9c b6 b6 4a a6 1d a8 ca b3 b4 23 e7 ad c4 c4 70 04 18 a6 5e 87 22 57 10 e1 69 1f 6d 9c b2 eb 63 cc e5 b6 05 ce 0a 4a 89 cf 51 97 67 e0 08 45 39 7c 5e 38 11 41 a0 ed 8a 89 b0 10 64 b4 95 ec 8d 1e 2d a3 74 33 bf 15 97 d9 19 a6 96 10 d2 ad 26 bd f6 fc a8 de 42 2b b2 cb 80 d0 51 62 06 e8 19 4e e5 14 45 a3 dc b5 de e3 3c 0c 31 0b 47 51 e6 8e 58 be bd a2 fb 58 69 85 a5 a5 b0 64 56 75 6f 44 e6 db fe f4 bb 99 ca 73 2f 00 fb 9e a3 57 75 f7 41 96 81 bf be 6f 43 dc 7c 46 50 c1 3b 63 d9 3c 1d 49 0b f0 17 3f 28 7a 43 09 53 1f 13 ec b1 a7 75 d0 bd 88 1a} (length 0x1000)
// size: len = 0x1000 (2 bytes)
// }
// }
// }
// len: len = 0x1006 (8 bytes)
// ]
*(uint32_t*)0x200000000940 = 8;
memcpy((void*)0x200000000944, "... [truncated large byte array] ...", 4096);
*(uint16_t*)0x200000001944 = 0x1000;
syscall(__NR_write, /*fd=*/r[0], /*data=*/0x200000000940ul, /*len=*/0x1006ul);
// quotactl$Q_SYNC arguments: [
// cmd: quota_cmd_sync = 0xffffffff80000101 (8 bytes)
// special: const = 0x0 (8 bytes)
// id: const = 0x0 (8 bytes)
// addr: const = 0x0 (8 bytes)
// ]
syscall(__NR_quotactl, /*cmd=Q_SYNC_GRP*/0xffffffff80000101ul, /*special=*/0ul, /*id=*/0ul, /*addr=*/0ul);
return 0;
}
Reply: Based on the analysis of the bug report and the f2fs source code, here is the assessment of the security impact:
### Bug Mechanism
The bug is a hung task (infinite loop) in `f2fs_sync_dirty_inodes` that occurs during the mounting of a crafted f2fs image.
When mounting an f2fs filesystem, if there is fsync data to recover, `f2fs_fill_super` calls `f2fs_recover_fsync_data`, which in turn calls `f2fs_write_checkpoint` -> `block_operations` -> `f2fs_sync_dirty_inodes`.
If writing a dirty page fails during this process (e.g., due to a memory allocation failure simulated by fault injection, or an `-EIO` / `-EFSCORRUPTED` error caused by a corrupted node page in a malicious image), `f2fs_write_single_data_page` redirties the page and returns the error. However, `f2fs_sync_dirty_inodes` ignores the error returned by `filemap_fdatawrite` and continues to loop as long as the inode's dirty list is not empty. Because the page was redirtied and the error did not set the `CP_ERROR_FLAG` to abort the checkpoint, the inode remains in the dirty list, resulting in an infinite loop.
### System-Wide Impact
During this infinite loop, the `mount` syscall holds the f2fs superblock's `s_umount` semaphore for **write**.
Because `s_umount` is held for write indefinitely, any other system calls that attempt to acquire `s_umount` for read on this superblock will block forever. This includes system-wide operations like `quotactl` (which iterates over all superblocks via `__iterate_supers` and takes `s_umount` for read) and `sync`. This perfectly matches the syzkaller bug report where multiple tasks are blocked for more than 143 seconds in `quotactl` -> `quota_sync_all` -> `__iterate_supers` -> `super_lock`.
### Analysis Dimensions
* **Exploitable:** **No.** The bug is an infinite loop leading to a CPU spin and a deadlock of the `s_umount` semaphore. There is no memory corruption, use-after-free, or out-of-bounds access involved, meaning it cannot be exploited for privilege escalation or arbitrary code execution.
* **Denial Of Service:** **Yes.** The infinite loop consumes a CPU core and permanently blocks system-wide operations like `sync()` and `quotactl()`, leading to a Denial of Service.
* **Accessible From Unprivileged Processes:** **No.** Mounting an f2fs filesystem requires `CAP_SYS_ADMIN` in the initial user namespace. Unprivileged users cannot trigger this directly.
* **Accessible From User Namespaces:** **No.** The f2fs filesystem does not have the `FS_USERNS_MOUNT` flag, meaning it cannot be mounted from within an unprivileged user namespace.
* **Peripheral Trigger:** **Yes.** An attacker can trigger this bug by plugging in a malicious USB drive containing a crafted f2fs image, provided the desktop environment (e.g., udisks2) is configured to auto-mount external media.
* **Malicious Filesystem Trigger:** **Yes.** The bug is triggered by parsing and recovering a corrupted/malicious f2fs filesystem image.
* **VM Guest Trigger:** **No.**
* **VM Host Trigger in The Confidential Computing Context:** **No.**
* **Ethernet Network Trigger:** **No.**
* **Other Remote Trigger:** **No.**
|
| 6/2 |
2026/05/26 06:33 |
llm |
expert |
0m
Model:
Tokens:
input: 36576
output: 21
thoughts: 3560
Thoughts: **Analyzing a Syzkaller Report: A Deep Dive into a Blocked Task**
Alright, let's break this down. We've got a `syzkaller` bug report indicating a task, `syz.2.19`, has been blocked for quite some time, specifically more than 143 seconds. The error message suggests a potential deadlock or livelock situation. My initial thought is to disable the `hung_task_timeout_secs` message, but that's just a temporary measure, a patch for now. Let's dig in.
The stack trace points to `context_switch`, then `__schedule`, and eventually, a lock contention within the kernel's scheduler. The culprit seems to be `rt_mutex_schedule` and `__rwbase_read_lock`, which uses a `down_read` call. From the context, this read-lock is on `s_umount`, acquired during a call to `super_lock` inside `__iterate_supers`, which is itself called from `quota_sync_all`.
My instinct is to check who's holding the lock. I see several kworkers holding `s_umount_key`. Interesting. These kworkers seem to be calling `super_trylock_shared` from `bdi_split_work_to_wbs` or `writeback_inodes_wb`. I'll keep that in mind.
There's an NMI backtrace from CPU 1, which appears to be a kworker process caught in memory allocation. It doesn't seem directly related to the blocked task, just a coincidentally-triggered NMI.
Right. The reproducer calls, which seems to be the critical part, involves mounting f2fs, opening `cpu.stat`, opening a file and writing to a UHID device, and finally a `quotactl` syscall. Wait, there is something that stands out in the mount configuration: `fault_injection=16` and `fault_type=16777214`.
Aha! Fault injection! This is a known source of pain. The `0xfffffe` value for `fault_type` means almost all faults are injected except for `FAULT_KMALLOC`. This could explain the infinite loops or deadlocks. The `super_lock` in `__iterate_supers` is now starting to make more sense. My first instinct here is to attribute this to the fault injection.
Looking at the locks held, I see the task itself holds `s_umount_key` for read, which means something else has it for write. Since `super_lock` calls `down_read(&sb->s_umount)`, `s_umount` is held for write, likely during the mount, umount, remount, or freeze/unfreeze operations.
The reproducer performs a mount using `syz_mount_image`, and then `quotactl`. This suggests the fault injection during f2fs mount could be the cause of the lock contention. During a mount operation, particularly during `fill_super`, `s_umount` will be held in write mode. It seems `fill_super` is being delayed due to the fault injection.
A crafted image and fault injection. This has serious implications: a hang here could be exploited for Denial of Service. Now I'm checking if an unprivileged user can create a user namespace and mount a filesystem with `FS_USERNS_MOUNT` enabled. If f2fs has `FS_USERNS_MOUNT` and it does not handle all faults when mounting a crafted image, we may have a vulnerability! This could give an unprivileged user the ability to DoS the system, making it more serious than just a kernel bug. I need to know: does f2fs support user namespace mounts?
|
| 7/2 |
2026/05/26 06:33 |
tool |
grepper |
0m
Args: map[Expression:FS_USERNS_MOUNT]
Results: map[Output:drivers/android/binderfs.c=782=static struct file_system_type binder_fs_type = {
--
drivers/android/binderfs.c-786- .kill_sb = binderfs_kill_super,
drivers/android/binderfs.c:787: .fs_flags = FS_USERNS_MOUNT,
drivers/android/binderfs.c-788-};
--
fs/binfmt_misc.c=1026=static struct file_system_type bm_fs_type = {
--
fs/binfmt_misc.c-1029- .init_fs_context = bm_init_fs_context,
fs/binfmt_misc.c:1030: .fs_flags = FS_USERNS_MOUNT,
fs/binfmt_misc.c-1031- .kill_sb = kill_litter_super,
--
fs/devpts/inode.c=467=static struct file_system_type devpts_fs_type = {
--
fs/devpts/inode.c-471- .kill_sb = devpts_kill_sb,
fs/devpts/inode.c:472: .fs_flags = FS_USERNS_MOUNT,
fs/devpts/inode.c-473-};
--
fs/fuse/inode.c=2100=static struct file_system_type fuse_fs_type = {
--
fs/fuse/inode.c-2102- .name = "fuse",
fs/fuse/inode.c:2103: .fs_flags = FS_HAS_SUBTYPE | FS_USERNS_MOUNT | FS_ALLOW_IDMAP,
fs/fuse/inode.c-2104- .init_fs_context = fuse_init_fs_context,
--
fs/overlayfs/params.c=751=static const struct fs_context_operations ovl_context_ops = {
--
fs/overlayfs/params.c-760- * This is called during fsopen() and will record the user namespace of
fs/overlayfs/params.c:761: * the caller in fc->user_ns since we've raised FS_USERNS_MOUNT. We'll
fs/overlayfs/params.c-762- * need it when we actually create the superblock to verify that the
--
fs/overlayfs/super.c=1514=struct file_system_type ovl_fs_type = {
--
fs/overlayfs/super.c-1518- .parameters = ovl_parameter_spec,
fs/overlayfs/super.c:1519: .fs_flags = FS_USERNS_MOUNT,
fs/overlayfs/super.c-1520- .kill_sb = kill_anon_super,
--
fs/proc/root.c=277=static struct file_system_type proc_fs_type = {
--
fs/proc/root.c-281- .kill_sb = proc_kill_sb,
fs/proc/root.c:282: .fs_flags = FS_USERNS_MOUNT | FS_DISALLOW_NOTIFY_PERM,
fs/proc/root.c-283-};
--
fs/ramfs/inode.c=319=static struct file_system_type ramfs_fs_type = {
--
fs/ramfs/inode.c-323- .kill_sb = ramfs_kill_sb,
fs/ramfs/inode.c:324: .fs_flags = FS_USERNS_MOUNT,
fs/ramfs/inode.c-325-};
--
fs/super.c=691=bool mount_capable(struct fs_context *fc)
fs/super.c-692-{
fs/super.c:693: if (!(fc->fs_type->fs_flags & FS_USERNS_MOUNT))
fs/super.c-694- return capable(CAP_SYS_ADMIN);
--
fs/super.c=731=struct super_block *sget_fc(struct fs_context *fc,
--
fs/super.c-740- /*
fs/super.c:741: * Never allow s_user_ns != &init_user_ns when FS_USERNS_MOUNT is
fs/super.c-742- * not set, as the filesystem is likely unprepared to handle it.
--
fs/super.c-745- */
fs/super.c:746: if (user_ns != &init_user_ns && !(fc->fs_type->fs_flags & FS_USERNS_MOUNT)) {
fs/super.c-747- errorfc(fc, "VFS: Mounting from non-initial user namespace is not allowed");
--
fs/sysfs/mount.c=90=static struct file_system_type sysfs_fs_type = {
--
fs/sysfs/mount.c-93- .kill_sb = sysfs_kill_sb,
fs/sysfs/mount.c:94: .fs_flags = FS_USERNS_MOUNT,
fs/sysfs/mount.c-95-};
--
include/linux/fs.h=2672=struct file_system_type {
--
include/linux/fs.h-2677-#define FS_HAS_SUBTYPE 4
include/linux/fs.h:2678:#define FS_USERNS_MOUNT 8 /* Can be mounted by userns root */
include/linux/fs.h-2679-#define FS_DISALLOW_NOTIFY_PERM 16 /* Disable fanotify permission events */
--
ipc/mqueue.c=1638=static struct file_system_type mqueue_fs_type = {
--
ipc/mqueue.c-1641- .kill_sb = kill_litter_super,
ipc/mqueue.c:1642: .fs_flags = FS_USERNS_MOUNT,
ipc/mqueue.c-1643-};
--
kernel/bpf/inode.c=1087=static struct file_system_type bpf_fs_type = {
--
kernel/bpf/inode.c-1092- .kill_sb = bpf_kill_super,
kernel/bpf/inode.c:1093: .fs_flags = FS_USERNS_MOUNT,
kernel/bpf/inode.c-1094-};
--
kernel/cgroup/cgroup.c=2346=struct file_system_type cgroup_fs_type = {
--
kernel/cgroup/cgroup.c-2350- .kill_sb = cgroup_kill_sb,
kernel/cgroup/cgroup.c:2351: .fs_flags = FS_USERNS_MOUNT,
kernel/cgroup/cgroup.c-2352-};
--
kernel/cgroup/cgroup.c=2354=static struct file_system_type cgroup2_fs_type = {
--
kernel/cgroup/cgroup.c-2358- .kill_sb = cgroup_kill_sb,
kernel/cgroup/cgroup.c:2359: .fs_flags = FS_USERNS_MOUNT,
kernel/cgroup/cgroup.c-2360-};
--
kernel/cgroup/cgroup.c=2427=static struct file_system_type cpuset_fs_type = {
--
kernel/cgroup/cgroup.c-2430- .parameters = cpuset_fs_parameters,
kernel/cgroup/cgroup.c:2431: .fs_flags = FS_USERNS_MOUNT,
kernel/cgroup/cgroup.c-2432-};
--
mm/shmem.c=5391=static struct file_system_type shmem_fs_type = {
--
mm/shmem.c-5398- .kill_sb = kill_litter_super,
mm/shmem.c:5399: .fs_flags = FS_USERNS_MOUNT | FS_ALLOW_IDMAP | FS_MGTIME,
mm/shmem.c-5400-};
--
mm/shmem.c=5801=static struct file_system_type shmem_fs_type = {
--
mm/shmem.c-5805- .kill_sb = ramfs_kill_sb,
mm/shmem.c:5806: .fs_flags = FS_USERNS_MOUNT,
mm/shmem.c-5807-};
]
|
| 8/2 |
2026/05/26 06:33 |
llm |
expert |
0m
Model:
Tokens:
input: 43169
output: 17
thoughts: 411
Thoughts: **Analyzing the f2fs Mount Vulnerability**
Okay, so f2fs doesn't support `FS_USERNS_MOUNT`, meaning it can't be mounted directly from a user namespace. Therefore, an unprivileged user can't mount an f2fs image directly, that's clear. But wait, what about USB drives? If someone plugs in an f2fs-formatted drive, does that change things? Ah, yes, in some desktop environments, daemons like udisks2 run as root and auto-mount those devices. So, it's a "Peripheral Trigger" and "Malicious Filesystem Trigger" type of vulnerability.
But hold on, is this actually an f2fs bug, or just a side effect of the fault injection? The reproducer uses `fault_injection=16` and `fault_type=16777214`. If it *is* fault injection, itβs likely not a practical vulnerability in a production system because fault injection is a debugging feature. But, wait, I recall Syzkaller found this. That's a red flag! Syzkaller often leverages fault injection to uncover error handling bugs. If itβs just the βtask blocked because of an infinite loop caused by fault injectionβ kind of thing, it still might be a real bug, if that same loop can be triggered without fault injection, perhaps by a corrupted filesystem.
Let me revisit the bug report. Okay, the task is blocked in `quotactl`, specifically in `quota_sync_all`, which calls `__iterate_supers`, which eventually calls `super_lock`. So, `quotactl` is trying to iterate over all superblocks and lock them. If a superblock is locked for a long time, then `quotactl` will block. So who's holding that superblock lock? It would be the f2fs mount process. Therefore, if the f2fs mount process is slow or stuck, `quotactl` will block.
Is the f2fs mount process stuck *because* of the fault injection? I need to delve into the f2fs mount code to find out.
|
| 9/2 |
2026/05/26 06:33 |
tool |
grepper |
0m
Args: map[Expression:fault_injection]
Results: map[Output:Documentation/ABI/testing/sysfs-fs-f2fs=709=Description: Support configuring fault injection type, should be
Documentation/ABI/testing/sysfs-fs-f2fs:710: enabled with fault_injection option, fault type value
Documentation/ABI/testing/sysfs-fs-f2fs-711- is shown below, it supports single or combined type.
--
Documentation/filesystems/f2fs.rst=178=resgid=%d The group ID which may use the reserved blocks.
Documentation/filesystems/f2fs.rst:179:fault_injection=%d Enable fault injection in all supported types with
Documentation/filesystems/f2fs.rst-180- specified injection rate.
Documentation/filesystems/f2fs.rst=181=fault_type=%d Support configuring fault injection type, should be
Documentation/filesystems/f2fs.rst:182: enabled with fault_injection option, fault type value
Documentation/filesystems/f2fs.rst-183- is shown below, it supports single or combined type.
--
drivers/edac/zynqmp_edac.c=91=struct ecc_status {
--
drivers/edac/zynqmp_edac.c-107- * @ue_bitpos: Array to store UnCorrectable Error bit positions
drivers/edac/zynqmp_edac.c:108: * @fault_injection_cnt: Fault Injection Counter value
drivers/edac/zynqmp_edac.c-109- */
drivers/edac/zynqmp_edac.c=110=struct edac_priv {
--
drivers/edac/zynqmp_edac.c-119- u8 ue_bitpos[OCM_NUM_UE_BITPOS];
drivers/edac/zynqmp_edac.c:120: u32 fault_injection_cnt;
drivers/edac/zynqmp_edac.c-121-#endif
--
drivers/edac/zynqmp_edac.c=230=static void write_fault_count(struct edac_priv *priv)
drivers/edac/zynqmp_edac.c-231-{
drivers/edac/zynqmp_edac.c:232: u32 ficount = priv->fault_injection_cnt;
drivers/edac/zynqmp_edac.c-233-
--
drivers/edac/zynqmp_edac.c=349=static void setup_debugfs(struct edac_device_ctl_info *edac_dev)
--
drivers/edac/zynqmp_edac.c-357- edac_debugfs_create_x32("inject_fault_count", 0644, priv->debugfs_dir,
drivers/edac/zynqmp_edac.c:358: &priv->fault_injection_cnt);
drivers/edac/zynqmp_edac.c-359- edac_debugfs_create_file("inject_ue_bitpos", 0644, priv->debugfs_dir,
--
fs/f2fs/super.c=125=enum {
--
fs/f2fs/super.c-148- Opt_mode,
fs/f2fs/super.c:149: Opt_fault_injection,
fs/f2fs/super.c-150- Opt_fault_type,
--
fs/f2fs/super.c=247=static const struct fs_parameter_spec f2fs_param_specs[] = {
--
fs/f2fs/super.c-270- fsparam_enum("mode", Opt_mode, f2fs_param_mode),
fs/f2fs/super.c:271: fsparam_s32("fault_injection", Opt_fault_injection),
fs/f2fs/super.c-272- fsparam_u32("fault_type", Opt_fault_type),
--
fs/f2fs/super.c=307=static match_table_t f2fs_checkpoint_tokens = {
--
fs/f2fs/super.c-321-#define F2FS_SPEC_mode (1 << 6)
fs/f2fs/super.c:322:#define F2FS_SPEC_fault_injection (1 << 7)
fs/f2fs/super.c-323-#define F2FS_SPEC_fault_type (1 << 8)
--
fs/f2fs/super.c=711=static int f2fs_parse_param(struct fs_context *fc, struct fs_parameter *param)
--
fs/f2fs/super.c-856-#ifdef CONFIG_F2FS_FAULT_INJECTION
fs/f2fs/super.c:857: case Opt_fault_injection:
fs/f2fs/super.c-858- F2FS_CTX_INFO(ctx).fault_info.inject_rate = result.int_32;
fs/f2fs/super.c:859: ctx->spec_mask |= F2FS_SPEC_fault_injection;
fs/f2fs/super.c-860- ctx_set_opt(ctx, F2FS_MOUNT_FAULT_INJECTION);
--
fs/f2fs/super.c-870-#else
fs/f2fs/super.c:871: case Opt_fault_injection:
fs/f2fs/super.c-872- case Opt_fault_type:
--
fs/f2fs/super.c=1608=static void f2fs_apply_options(struct fs_context *fc, struct super_block *sb)
--
fs/f2fs/super.c-1632-#ifdef CONFIG_F2FS_FAULT_INJECTION
fs/f2fs/super.c:1633: if (ctx->spec_mask & F2FS_SPEC_fault_injection)
fs/f2fs/super.c-1634- (void)f2fs_build_fault_attr(sbi,
--
fs/f2fs/super.c=2257=static int f2fs_show_options(struct seq_file *seq, struct dentry *root)
--
fs/f2fs/super.c-2354- if (test_opt(sbi, FAULT_INJECTION)) {
fs/f2fs/super.c:2355: seq_printf(seq, ",fault_injection=%u",
fs/f2fs/super.c-2356- F2FS_OPTION(sbi).fault_info.inject_rate);
--
net/core/Makefile=48=obj-$(CONFIG_DEBUG_NET) += lock_debug.o
net/core/Makefile:49:obj-$(CONFIG_FAIL_SKB_REALLOC) += skb_fault_injection.o
--
tools/testing/selftests/iommu/iommufd_fail_nth.c-24-
tools/testing/selftests/iommu/iommufd_fail_nth.c:25:static bool have_fault_injection;
tools/testing/selftests/iommu/iommufd_fail_nth.c-26-
--
tools/testing/selftests/iommu/iommufd_fail_nth.c=42=static __attribute__((constructor)) void setup_buffer(void)
--
tools/testing/selftests/iommu/iommufd_fail_nth.c-58- */
tools/testing/selftests/iommu/iommufd_fail_nth.c:59:static __attribute__((constructor)) void setup_fault_injection(void)
tools/testing/selftests/iommu/iommufd_fail_nth.c-60-{
--
tools/testing/selftests/iommu/iommufd_fail_nth.c-83- closedir(debugfs);
tools/testing/selftests/iommu/iommufd_fail_nth.c:84: have_fault_injection = true;
tools/testing/selftests/iommu/iommufd_fail_nth.c-85-}
--
tools/testing/selftests/iommu/iommufd_fail_nth.c=162=void __fail_nth_enable(struct __test_metadata *_metadata,
--
tools/testing/selftests/iommu/iommufd_fail_nth.c-186- \
tools/testing/selftests/iommu/iommufd_fail_nth.c:187: if (!have_fault_injection) \
tools/testing/selftests/iommu/iommufd_fail_nth.c-188- SKIP(return, \
]
|
| 10/2 |
2026/05/26 06:33 |
llm |
expert |
0m
Model:
Tokens:
input: 46282
output: 18
thoughts: 1608
Thoughts: **Analyzing the F2FS Mount Hang**
Okay, so the reproducer uses `fault_injection=16` and `fault_type=16777214`. This is F2FS-specific, and the `fault_injection` option immediately sets off alarm bells. It injects faults into memory allocations, I/O, you name it, and that's a recipe for infinite loops and deadlocks if the error handling is even slightly off. But, the crucial question is whether `CONFIG_F2FS_FAULT_INJECTION` is even enabled in a production environment? Almost certainly not; it's a debug tool. If this bug *requires* fault injection to trigger, then it's not a real-world exploitable problem.
But can it be triggered without fault injection? If the bug is an infinite loop in f2fs mount due to a corrupted image, then fault injection might just be simulating a corrupted image or an I/O error.
Let's see where the mount is stuck. The task `syz.2.19` is blocked on `super_lock` inside `__iterate_supers`. The lock `s_umount` is possibly being held by `f2fs_ckpt-7:2/6140`, `f2fs_ckpt-7:1/6146`, or `syz-executor`. Since the problem is a mount, `syz-executor` seems like the more likely candidate.
Okay, let's look at the lock information of other tasks: `f2fs_ckpt` sounds like a kernel thread F2FS creates, which means the mount has actually succeeded. It is created at the end of `f2fs_fill_super`. Thus, it means that `s_umount` isn't held by `mount` anymore. So *who* is holding `s_umount` for *write*? Since the current task `syz.2.19` is blocked on `super_lock`, and `s_umount` is a `rw_semaphore`, someone is trying to acquire it in write mode. It could be `mount`, `umount`, `remount`, `freeze`, `unfreeze`, or `sync_fs`. The `umount` function does `down_write(&sb->s_umount)`, so maybe the `syz-executor` is doing the `umount`.
Checking the locks held by `syz-executor/5963` doesn't help: it doesn't show any. Maybe other `syz` tasks are doing something.
Wait, what about `s_umount` being held by a `kworker`? If multiple `kworker`s are holding it for read, and someone tries to take it for read, it should succeed. Except⦠there's a *pending writer*! `rw_semaphore` is fair. If there's a pending writer, new readers will block.
So, someone is trying to acquire `s_umount` for write and is blocked by the `kworker`s holding it for read. Who is the pending writer? There are several `syz.*` tasks, and those tasks are blocked on `super_lock`, so it seems the task is trying to call `super_lock`. Lockdep says task `syz.2.19/6138` holds the lock, but the task is actually blocked. But, `__super_lock` calls `down_read(&sb->s_umount)`, so it means `syz.2.19/6138` is trying to get a read lock. But `down_read` is called. It means it is waiting for the lock.
It looks like the `super_lock` is a custom lock. Let me go check `fs/super.c`.
|
| 11/2 |
2026/05/26 06:34 |
tool |
grepper |
0m
Args: map[Expression:__super_lock]
Results: map[Output:fs/super.c=48=static char *sb_writers_name[SB_FREEZE_LEVELS] = {
--
fs/super.c-53-
fs/super.c:54:static inline void __super_lock(struct super_block *sb, bool excl)
fs/super.c-55-{
--
fs/super.c=62=static inline void super_unlock(struct super_block *sb, bool excl)
--
fs/super.c-69-
fs/super.c:70:static inline void __super_lock_excl(struct super_block *sb)
fs/super.c-71-{
fs/super.c:72: __super_lock(sb, true);
fs/super.c-73-}
--
fs/super.c=110=static __must_check bool super_lock(struct super_block *sb, bool excl)
--
fs/super.c-120-
fs/super.c:121: __super_lock(sb, excl);
fs/super.c-122-
--
fs/super.c=503=void deactivate_super(struct super_block *s)
--
fs/super.c-505- if (!atomic_add_unless(&s->s_active, -1, 1)) {
fs/super.c:506: __super_lock_excl(s);
fs/super.c-507- deactivate_locked_super(s);
--
fs/super.c=589=void retire_super(struct super_block *sb)
--
fs/super.c-591- WARN_ON(!sb->s_bdev);
fs/super.c:592: __super_lock_excl(sb);
fs/super.c-593- if (sb->s_iflags & SB_I_PERSB_BDI) {
--
fs/super.c=1016=int reconfigure_super(struct fs_context *fc)
--
fs/super.c-1046- group_pin_kill(&sb->s_pins);
fs/super.c:1047: __super_lock_excl(sb);
fs/super.c-1048- if (!sb->s_root)
--
fs/super.c=2131=int freeze_super(struct super_block *sb, enum freeze_holder who, const void *freeze_owner)
--
fs/super.c-2175- sb_wait_write(sb, SB_FREEZE_WRITE);
fs/super.c:2176: __super_lock_excl(sb);
fs/super.c-2177-
]
|
| 12/2 |
2026/05/26 06:34 |
llm |
expert |
0m
Model:
Tokens:
input: 49421
output: 38
thoughts: 12
Thoughts: Here's my attempt at summarizing the thought process, tailored for an expert audience:
**Diving into the `__super_lock` Implementation**
Okay, time to dig into the `__super_lock` implementation. I need to refresh my understanding of its intricacies. I've been meaning to revisit this; there are likely some subtle edge cases and performance considerations that are worth another look. I'm particularly interested in how it handles contention and what underlying mechanisms it leverages for atomicity. Are they relying on atomic instructions directly, or are there more complex synchronization primitives at play? I'll need to trace the code path carefully, paying close attention to any potential lock hierarchies or deadlock vulnerabilities. I'm also curious about its interaction with other locking mechanisms in the system. Does it play nicely with spinlocks, mutexes, or reader-writer locks, or is it isolated? Specifically, I'll be looking for potential optimization opportunities, maybe there are things I can refactor for improved throughput under heavy load. I'll need to think about how to quantify any changes after making any. And what's with the naming convention? I should be looking at where it's called from and how frequently it is utilized, which can give insights into the design.
|
| 13/2 |
2026/05/26 06:34 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/super.c IncludeLines:true Name:__super_lock]
Results: map[SourceCode: 54: static inline void __super_lock(struct super_block *sb, bool excl)
55: {
56: if (excl)
57: down_write(&sb->s_umount);
58: else
59: down_read(&sb->s_umount);
60: }
SourceFile:fs/super.c]
|
| 14/2 |
2026/05/26 06:34 |
llm |
expert |
0m
Model:
Tokens:
input: 49849
output: 37
thoughts: 265
Thoughts: **Debugging a Lockdep Mystery**
Okay, so `__super_lock` just calls either `down_read` or `down_write`. That's straightforward enough. But hold on, why is `lockdep` showing it as held? That doesn't make sense if the task is blocked, right? Oh, right, I'm such a fool! `lockdep` tracks locks *held by the task*, not locks the task is *trying* to acquire.
So if a task is blocked inside of `down_read`, it's not actually holding the lock yet. But wait a minute... `syz.2.19` *is* blocked in `down_read`. And yet, `lockdep` is telling me: `"1 lock held by syz.2.19/6138: #0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: __super_lock fs/super.c:59 [inline] #0: ffff8880281d00d0 (&type->s_umount_key#72){++++}-{4:4}, at: super_lock+0x2a9/0x3b0 fs/super.c:121"`. That's a direct contradiction!
Okay, so the `super_lock` function is at `fs/super.c:121`. I need to go and check the source code for the `super_lock` function to see what's going on. Something doesn't add up here.
|
| 15/2 |
2026/05/26 06:34 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/super.c IncludeLines:true Name:super_lock]
Results: map[SourceCode: 94: /**
95: * super_lock - wait for superblock to become ready and lock it
96: * @sb: superblock to wait for
97: * @excl: whether exclusive access is required
98: *
99: * If the superblock has neither passed through vfs_get_tree() or
100: * generic_shutdown_super() yet wait for it to happen. Either superblock
101: * creation will succeed and SB_BORN is set by vfs_get_tree() or we're
102: * woken and we'll see SB_DYING.
103: *
104: * The caller must have acquired a temporary reference on @sb->s_count.
105: *
106: * Return: The function returns true if SB_BORN was set and with
107: * s_umount held. The function returns false if SB_DYING was
108: * set and without s_umount held.
109: */
110: static __must_check bool super_lock(struct super_block *sb, bool excl)
111: {
112: lockdep_assert_not_held(&sb->s_umount);
113:
114: /* wait until the superblock is ready or dying */
115: wait_var_event(&sb->s_flags, super_flags(sb, SB_BORN | SB_DYING));
116:
117: /* Don't pointlessly acquire s_umount. */
118: if (super_flags(sb, SB_DYING))
119: return false;
120:
121: __super_lock(sb, excl);
122:
123: /*
124: * Has gone through generic_shutdown_super() in the meantime.
125: * @sb->s_root is NULL and @sb->s_active is 0. No one needs to
126: * grab a reference to this. Tell them so.
127: */
128: if (sb->s_flags & SB_DYING) {
129: super_unlock(sb, excl);
130: return false;
131: }
132:
133: WARN_ON_ONCE(!(sb->s_flags & SB_BORN));
134: return true;
135: }
SourceFile:fs/super.c]
|
| 16/2 |
2026/05/26 06:34 |
llm |
expert |
0m
Model:
Tokens:
input: 51205
output: 23
thoughts: 1657
Thoughts: **Analyzing a Potential F2FS Deadlock with Lockdep**
Okay, so `lockdep` is flagging a lock held because `down_read` essentially reports the lock acquisition *before* it actually blocks, which makes sense for detecting potential deadlocks. `syz.2.19` is trying to grab `s_umount` for a read, and it's blocked. My immediate thought is, who's holding it for write?
I checked the lock list, and `kworker` has one instance of `&type->s_umount_key#72` but `syz.2.19` is trying to acquire *another* instance, `ffff8880281d00d0`, which is also `&type->s_umount_key#72`. It's a different lock instance.
Now I look for who's holding that specific lock, `ffff8880281d00d0`. And, surprise, it seems all the `syz` tasks are trying to acquire the same lock. It's almost as though they are waiting for each other, forming a read lock cycle. But no one is actually holding it.
That's when it hits me: the lock list is incomplete! The output is truncated, and it's hiding the task that holds the write lock. `syz-executor/5963` *should* be the one doing that.
I remember that `syz-executor` is the parent process performing the `mount` operation. `mount` calls `vfs_get_tree` and that calls `f2fs_fill_super`. And, during `f2fs_fill_super` the `s_umount` is indeed held for write.
So, if `f2fs_fill_super` is stuck, the write lock will be held forever. But why would it be stuck? I remember the reproducer is mounting f2fs with `fault_injection=16` and `fault_type=16777214`. Could this be due to a known f2fs fault injection issue that causes an infinite loop in `f2fs_fill_super`?
We don't have the stack trace of `syz-executor/5963` in this bug report. But there is a `f2fs_ckpt` process running. I recall that `f2fs_ckpt` is created during `f2fs_fill_super`. It seems like `f2fs_fill_super` is stuck in this instance. Now I have a better idea of what to do next.
|
| 17/2 |
2026/05/26 06:34 |
tool |
codesearch-find-references |
0m
Error:
Args:
Results: map[References:<nil> TruncatedOutput:false]
|
| 18/2 |
2026/05/26 06:34 |
llm |
expert |
0m
Model:
Tokens:
input: 53476
output: 19
thoughts: 39
Thoughts: Okay, here's my interpretation of those thoughts, crafted with the expectation that I'm talking to a fellow expert:
**Deep Dive into F2FS Checkpointing**
Alright, that thread name, `f2fs_ckpt-%u:%u`, is the key. Seems like we're zeroing in on a checkpointing process within the F2FS filesystem. First things first, I need to get a handle on what's going on here. The obvious starting point is a quick search for "f2fs_ckpt" within the `fs/f2fs/` directory. That's where all the core F2FS code resides, so it's the likeliest location for the functions or structures that are referenced in this thread name. Time to see what we can dig up. Hopefully, we'll quickly discover the function calls or data structures that are related to this checkpointing thread. From there we can move to understanding what the purpose of the checkpointing thread is.
|
| 19/2 |
2026/05/26 06:34 |
tool |
grepper |
0m
Args: map[Expression:f2fs_ckpt]
Results: map[Output:fs/f2fs/checkpoint.c=1879=int f2fs_start_ckpt_thread(struct f2fs_sb_info *sbi)
--
fs/f2fs/checkpoint.c-1887- cprc->f2fs_issue_ckpt = kthread_run(issue_checkpoint_thread, sbi,
fs/f2fs/checkpoint.c:1888: "f2fs_ckpt-%u:%u", MAJOR(dev), MINOR(dev));
fs/f2fs/checkpoint.c-1889- if (IS_ERR(cprc->f2fs_issue_ckpt)) {
]
|
| 20/2 |
2026/05/26 06:34 |
llm |
expert |
0m
Model:
Tokens:
input: 53910
output: 27
thoughts: 19
Thoughts: Okay, let's dive into this. Here's my thought process on tracking down where `f2fs_start_ckpt_thread` is called, presented as if it were my own internal monologue:
**Tracing the Checkpoint Thread's Origins**
Right, so I need to figure out the calling context for `f2fs_start_ckpt_thread`. This is a classic debugging scenario. Let's see... the most direct approach would be a code search. I'll need to use a tool that can navigate the kernel source and accurately identify cross-references. Probably `grep` with the right flags, or maybe `cscope` if I want something more interactive, would be the most efficient. I need to make sure I'm searching within the relevant `f2fs` directory and its subdirectories, or at least setting a good filter to limit false positives. I'll be looking for instances where the function name appears directly, and possibly any macro definitions that expand to it. I also need to be mindful of any potential indirect calls through function pointers or similar mechanisms. Need to ensure the tool is set to track those! Then I'll meticulously go through each call site and try to piece together the larger execution path that leads to starting the checkpoint thread. This will involve understanding what the context looks like, what the overall system is doing. Is it during mount, during a flush, or something else entirely? The comments should help clarify the broader picture. I'll focus on understanding the logic around each call, paying attention to any conditional statements or potential race conditions that might be in play. Ultimately, I'll build a mental map of how the checkpointing thread is initiated and its relation to other core `f2fs` operations.
|
| 21/2 |
2026/05/26 06:34 |
tool |
codesearch-find-references |
0m
Args: map[Name:f2fs_start_ckpt_thread]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:__f2fs_remount SourceFile:fs/f2fs/super.c SourceLine:2820] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:f2fs_fill_super SourceFile:fs/f2fs/super.c SourceLine:4985]] TruncatedOutput:false]
|
| 22/2 |
2026/05/26 06:34 |
llm |
expert |
0m
Model:
Tokens:
input: 54447
output: 45
thoughts: 50
Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience:
**Diving into F2FS Superblock Initialization**
Alright, so we're talking about F2FS superblock initialization, specifically within the `f2fs_fill_super` function. I see the reference β line 4985, got it. My immediate focus is on understanding the checkpoint thread initiation within that context. I need to trace the flow after `f2fs_start_ckpt_thread` is called within `f2fs_fill_super`. This is a crucial area because this function is the linchpin for setting up the filesystem metadata and recovery processes. I'm keen to dissect how the checkpoint thread is integrated with superblock initialization. I want to see exactly how the asynchronous checkpointing, which is the cornerstone of F2FS's crash consistency design, plays out during the mount operation.
|
| 23/2 |
2026/05/26 06:34 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/super.c IncludeLines:true Name:f2fs_fill_super]
Results: map[SourceCode:4770: static int f2fs_fill_super(struct super_block *sb, struct fs_context *fc)
4771: {
4772: struct f2fs_fs_context *ctx = fc->fs_private;
4773: struct f2fs_sb_info *sbi;
4774: struct f2fs_super_block *raw_super;
4775: struct inode *root;
4776: int err;
4777: bool skip_recovery = false, need_fsck = false;
4778: int recovery, i, valid_super_block;
4779: struct curseg_info *seg_i;
4780: int retry_cnt = 1;
4781: #ifdef CONFIG_QUOTA
4782: bool quota_enabled = false;
4783: #endif
4784:
4785: try_onemore:
4786: err = -EINVAL;
4787: raw_super = NULL;
4788: valid_super_block = -1;
4789: recovery = 0;
4790:
4791: /* allocate memory for f2fs-specific super block info */
4792: sbi = kzalloc(sizeof(struct f2fs_sb_info), GFP_KERNEL);
4793: if (!sbi)
4794: return -ENOMEM;
4795:
4796: sbi->sb = sb;
4797:
4798: /* initialize locks within allocated memory */
4799: init_f2fs_rwsem(&sbi->gc_lock);
4800: mutex_init(&sbi->writepages);
4801: init_f2fs_rwsem(&sbi->cp_global_sem);
4802: init_f2fs_rwsem(&sbi->node_write);
4803: init_f2fs_rwsem(&sbi->node_change);
4804: spin_lock_init(&sbi->stat_lock);
4805: init_f2fs_rwsem(&sbi->cp_rwsem);
4806: init_f2fs_rwsem(&sbi->quota_sem);
4807: init_waitqueue_head(&sbi->cp_wait);
4808: spin_lock_init(&sbi->error_lock);
4809:
4810: for (i = 0; i < NR_INODE_TYPE; i++) {
4811: INIT_LIST_HEAD(&sbi->inode_list[i]);
4812: spin_lock_init(&sbi->inode_lock[i]);
4813: }
4814: mutex_init(&sbi->flush_lock);
4815:
4816: /* set a block size */
4817: if (unlikely(!sb_set_blocksize(sb, F2FS_BLKSIZE))) {
4818: f2fs_err(sbi, "unable to set blocksize");
4819: goto free_sbi;
4820: }
4821:
4822: err = read_raw_super_block(sbi, &raw_super, &valid_super_block,
4823: &recovery);
4824: if (err)
4825: goto free_sbi;
4826:
4827: sb->s_fs_info = sbi;
4828: sbi->raw_super = raw_super;
4829:
4830: INIT_WORK(&sbi->s_error_work, f2fs_record_error_work);
4831: memcpy(sbi->errors, raw_super->s_errors, MAX_F2FS_ERRORS);
4832: memcpy(sbi->stop_reason, raw_super->s_stop_reason, MAX_STOP_REASON);
4833:
4834: /* precompute checksum seed for metadata */
4835: if (f2fs_sb_has_inode_chksum(sbi))
4836: sbi->s_chksum_seed = f2fs_chksum(~0, raw_super->uuid,
4837: sizeof(raw_super->uuid));
4838:
4839: default_options(sbi, false);
4840:
4841: err = f2fs_check_opt_consistency(fc, sb);
4842: if (err)
4843: goto free_sb_buf;
4844:
4845: f2fs_apply_options(fc, sb);
4846:
4847: err = f2fs_sanity_check_options(sbi, false);
4848: if (err)
4849: goto free_options;
4850:
4851: sb->s_maxbytes = max_file_blocks(NULL) <<
4852: le32_to_cpu(raw_super->log_blocksize);
4853: sb->s_max_links = F2FS_LINK_MAX;
4854:
4855: err = f2fs_setup_casefold(sbi);
4856: if (err)
4857: goto free_options;
4858:
4859: #ifdef CONFIG_QUOTA
4860: sb->dq_op = &f2fs_quota_operations;
4861: sb->s_qcop = &f2fs_quotactl_ops;
4862: sb->s_quota_types = QTYPE_MASK_USR | QTYPE_MASK_GRP | QTYPE_MASK_PRJ;
4863:
4864: if (f2fs_sb_has_quota_ino(sbi)) {
4865: for (i = 0; i < MAXQUOTAS; i++) {
4866: if (f2fs_qf_ino(sbi->sb, i))
4867: sbi->nquota_files++;
4868: }
4869: }
4870: #endif
4871:
4872: sb->s_op = &f2fs_sops;
4873: #ifdef CONFIG_FS_ENCRYPTION
4874: sb->s_cop = &f2fs_cryptops;
4875: #endif
4876: #ifdef CONFIG_FS_VERITY
4877: sb->s_vop = &f2fs_verityops;
4878: #endif
4879: sb->s_xattr = f2fs_xattr_handlers;
4880: sb->s_export_op = &f2fs_export_ops;
4881: sb->s_magic = F2FS_SUPER_MAGIC;
4882: sb->s_time_gran = 1;
4883: sb->s_flags = (sb->s_flags & ~SB_POSIXACL) |
4884: (test_opt(sbi, POSIX_ACL) ? SB_POSIXACL : 0);
4885: if (test_opt(sbi, INLINECRYPT))
4886: sb->s_flags |= SB_INLINECRYPT;
4887:
4888: if (test_opt(sbi, LAZYTIME))
4889: sb->s_flags |= SB_LAZYTIME;
4890: else
4891: sb->s_flags &= ~SB_LAZYTIME;
4892:
4893: super_set_uuid(sb, (void *) raw_super->uuid, sizeof(raw_super->uuid));
4894: super_set_sysfs_name_bdev(sb);
4895: sb->s_iflags |= SB_I_CGROUPWB;
4896:
4897: /* init f2fs-specific super block info */
4898: sbi->valid_super_block = valid_super_block;
4899:
4900: /* disallow all the data/node/meta page writes */
4901: set_sbi_flag(sbi, SBI_POR_DOING);
4902:
4903: err = f2fs_init_write_merge_io(sbi);
4904: if (err)
4905: goto free_bio_info;
4906:
4907: init_sb_info(sbi);
4908:
4909: err = f2fs_init_iostat(sbi);
4910: if (err)
4911: goto free_bio_info;
4912:
4913: err = init_percpu_info(sbi);
4914: if (err)
4915: goto free_iostat;
4916:
4917: /* init per sbi slab cache */
4918: err = f2fs_init_xattr_caches(sbi);
4919: if (err)
4920: goto free_percpu;
4921: err = f2fs_init_page_array_cache(sbi);
4922: if (err)
4923: goto free_xattr_cache;
4924:
4925: /* get an inode for meta space */
4926: sbi->meta_inode = f2fs_iget(sb, F2FS_META_INO(sbi));
4927: if (IS_ERR(sbi->meta_inode)) {
4928: f2fs_err(sbi, "Failed to read F2FS meta data inode");
4929: err = PTR_ERR(sbi->meta_inode);
4930: goto free_page_array_cache;
4931: }
4932:
4933: err = f2fs_get_valid_checkpoint(sbi);
4934: if (err) {
4935: f2fs_err(sbi, "Failed to get valid F2FS checkpoint");
4936: goto free_meta_inode;
4937: }
4938:
4939: if (__is_set_ckpt_flags(F2FS_CKPT(sbi), CP_QUOTA_NEED_FSCK_FLAG))
4940: set_sbi_flag(sbi, SBI_QUOTA_NEED_REPAIR);
4941: if (__is_set_ckpt_flags(F2FS_CKPT(sbi), CP_DISABLED_QUICK_FLAG)) {
4942: set_sbi_flag(sbi, SBI_CP_DISABLED_QUICK);
4943: sbi->interval_time[DISABLE_TIME] = DEF_DISABLE_QUICK_INTERVAL;
4944: }
4945:
4946: if (__is_set_ckpt_flags(F2FS_CKPT(sbi), CP_FSCK_FLAG))
4947: set_sbi_flag(sbi, SBI_NEED_FSCK);
4948:
4949: /* Initialize device list */
4950: err = f2fs_scan_devices(sbi);
4951: if (err) {
4952: f2fs_err(sbi, "Failed to find devices");
4953: goto free_devices;
4954: }
4955:
4956: err = f2fs_init_post_read_wq(sbi);
4957: if (err) {
4958: f2fs_err(sbi, "Failed to initialize post read workqueue");
4959: goto free_devices;
4960: }
4961:
4962: sbi->total_valid_node_count =
4963: le32_to_cpu(sbi->ckpt->valid_node_count);
4964: percpu_counter_set(&sbi->total_valid_inode_count,
4965: le32_to_cpu(sbi->ckpt->valid_inode_count));
4966: sbi->user_block_count = le64_to_cpu(sbi->ckpt->user_block_count);
4967: sbi->total_valid_block_count =
4968: le64_to_cpu(sbi->ckpt->valid_block_count);
4969: sbi->last_valid_block_count = sbi->total_valid_block_count;
4970: sbi->reserved_blocks = 0;
4971: sbi->current_reserved_blocks = 0;
4972: limit_reserve_root(sbi);
4973: adjust_unusable_cap_perc(sbi);
4974:
4975: f2fs_init_extent_cache_info(sbi);
4976:
4977: f2fs_init_ino_entry_info(sbi);
4978:
4979: f2fs_init_fsync_node_info(sbi);
4980:
4981: /* setup checkpoint request control and start checkpoint issue thread */
4982: f2fs_init_ckpt_req_control(sbi);
4983: if (!f2fs_readonly(sb) && !test_opt(sbi, DISABLE_CHECKPOINT) &&
4984: test_opt(sbi, MERGE_CHECKPOINT)) {
4985: err = f2fs_start_ckpt_thread(sbi);
4986: if (err) {
4987: f2fs_err(sbi,
4988: "Failed to start F2FS issue_checkpoint_thread (%d)",
4989: err);
4990: goto stop_ckpt_thread;
4991: }
4992: }
4993:
4994: /* setup f2fs internal modules */
4995: err = f2fs_build_segment_manager(sbi);
4996: if (err) {
4997: f2fs_err(sbi, "Failed to initialize F2FS segment manager (%d)",
4998: err);
4999: goto free_sm;
5000: }
5001: err = f2fs_build_node_manager(sbi);
5002: if (err) {
5003: f2fs_err(sbi, "Failed to initialize F2FS node manager (%d)",
5004: err);
5005: goto free_nm;
5006: }
5007:
5008: /* For write statistics */
5009: sbi->sectors_written_start = f2fs_get_sectors_written(sbi);
5010:
5011: /* get segno of first zoned block device */
5012: sbi->first_seq_zone_segno = get_first_seq_zone_segno(sbi);
5013:
5014: sbi->reserved_pin_section = f2fs_sb_has_blkzoned(sbi) ?
5015: ZONED_PIN_SEC_REQUIRED_COUNT :
5016: GET_SEC_FROM_SEG(sbi, overprovision_segments(sbi));
5017:
5018: /* Read accumulated write IO statistics if exists */
5019: seg_i = CURSEG_I(sbi, CURSEG_HOT_NODE);
5020: if (__exist_node_summaries(sbi))
5021: sbi->kbytes_written =
5022: le64_to_cpu(seg_i->journal->info.kbytes_written);
5023:
5024: f2fs_build_gc_manager(sbi);
5025:
5026: err = f2fs_build_stats(sbi);
5027: if (err)
5028: goto free_nm;
5029:
5030: /* get an inode for node space */
5031: sbi->node_inode = f2fs_iget(sb, F2FS_NODE_INO(sbi));
5032: if (IS_ERR(sbi->node_inode)) {
5033: f2fs_err(sbi, "Failed to read node inode");
5034: err = PTR_ERR(sbi->node_inode);
5035: goto free_stats;
5036: }
5037:
5038: /* read root inode and dentry */
5039: root = f2fs_iget(sb, F2FS_ROOT_INO(sbi));
5040: if (IS_ERR(root)) {
5041: f2fs_err(sbi, "Failed to read root inode");
5042: err = PTR_ERR(root);
5043: goto free_node_inode;
5044: }
5045: if (!S_ISDIR(root->i_mode) || !root->i_blocks ||
5046: !root->i_size || !root->i_nlink) {
5047: iput(root);
5048: err = -EINVAL;
5049: goto free_node_inode;
5050: }
5051:
5052: generic_set_sb_d_ops(sb);
5053: sb->s_root = d_make_root(root); /* allocate root dentry */
5054: if (!sb->s_root) {
5055: err = -ENOMEM;
5056: goto free_node_inode;
5057: }
5058:
5059: err = f2fs_init_compress_inode(sbi);
5060: if (err)
5061: goto free_root_inode;
5062:
5063: err = f2fs_register_sysfs(sbi);
5064: if (err)
5065: goto free_compress_inode;
5066:
5067: sbi->umount_lock_holder = current;
5068: #ifdef CONFIG_QUOTA
5069: /* Enable quota usage during mount */
5070: if (f2fs_sb_has_quota_ino(sbi) && !f2fs_readonly(sb)) {
5071: err = f2fs_enable_quotas(sb);
5072: if (err)
5073: f2fs_err(sbi, "Cannot turn on quotas: error %d", err);
5074: }
5075:
5076: quota_enabled = f2fs_recover_quota_begin(sbi);
5077: #endif
5078: /* if there are any orphan inodes, free them */
5079: err = f2fs_recover_orphan_inodes(sbi);
5080: if (err)
5081: goto free_meta;
5082:
5083: if (unlikely(is_set_ckpt_flags(sbi, CP_DISABLED_FLAG))) {
5084: skip_recovery = true;
5085: goto reset_checkpoint;
5086: }
5087:
5088: /* recover fsynced data */
5089: if (!test_opt(sbi, DISABLE_ROLL_FORWARD) &&
5090: !test_opt(sbi, NORECOVERY)) {
5091: /*
5092: * mount should be failed, when device has readonly mode, and
5093: * previous checkpoint was not done by clean system shutdown.
5094: */
5095: if (f2fs_hw_is_readonly(sbi)) {
5096: if (!is_set_ckpt_flags(sbi, CP_UMOUNT_FLAG)) {
5097: err = f2fs_recover_fsync_data(sbi, true);
5098: if (err > 0) {
5099: err = -EROFS;
5100: f2fs_err(sbi, "Need to recover fsync data, but "
5101: "write access unavailable, please try "
5102: "mount w/ disable_roll_forward or norecovery");
5103: }
5104: if (err < 0)
5105: goto free_meta;
5106: }
5107: f2fs_info(sbi, "write access unavailable, skipping recovery");
5108: goto reset_checkpoint;
5109: }
5110:
5111: if (need_fsck)
5112: set_sbi_flag(sbi, SBI_NEED_FSCK);
5113:
5114: if (skip_recovery)
5115: goto reset_checkpoint;
5116:
5117: err = f2fs_recover_fsync_data(sbi, false);
5118: if (err < 0) {
5119: if (err != -ENOMEM)
5120: skip_recovery = true;
5121: need_fsck = true;
5122: f2fs_err(sbi, "Cannot recover all fsync data errno=%d",
5123: err);
5124: goto free_meta;
5125: }
5126: } else {
5127: err = f2fs_recover_fsync_data(sbi, true);
5128:
5129: if (!f2fs_readonly(sb) && err > 0) {
5130: err = -EINVAL;
5131: f2fs_err(sbi, "Need to recover fsync data");
5132: goto free_meta;
5133: }
5134: }
5135:
5136: reset_checkpoint:
5137: #ifdef CONFIG_QUOTA
5138: f2fs_recover_quota_end(sbi, quota_enabled);
5139: #endif
5140: /*
5141: * If the f2fs is not readonly and fsync data recovery succeeds,
5142: * write pointer consistency of cursegs and other zones are already
5143: * checked and fixed during recovery. However, if recovery fails,
5144: * write pointers are left untouched, and retry-mount should check
5145: * them here.
5146: */
5147: if (skip_recovery)
5148: err = f2fs_check_and_fix_write_pointer(sbi);
5149: if (err)
5150: goto free_meta;
5151:
5152: /* f2fs_recover_fsync_data() cleared this already */
5153: clear_sbi_flag(sbi, SBI_POR_DOING);
5154:
5155: err = f2fs_init_inmem_curseg(sbi);
5156: if (err)
5157: goto sync_free_meta;
5158:
5159: if (test_opt(sbi, DISABLE_CHECKPOINT)) {
5160: err = f2fs_disable_checkpoint(sbi);
5161: if (err)
5162: goto sync_free_meta;
5163: } else if (is_set_ckpt_flags(sbi, CP_DISABLED_FLAG)) {
5164: f2fs_enable_checkpoint(sbi);
5165: }
5166:
5167: /*
5168: * If filesystem is not mounted as read-only then
5169: * do start the gc_thread.
5170: */
5171: if ((F2FS_OPTION(sbi).bggc_mode != BGGC_MODE_OFF ||
5172: test_opt(sbi, GC_MERGE)) && !f2fs_readonly(sb)) {
5173: /* After POR, we can run background GC thread.*/
5174: err = f2fs_start_gc_thread(sbi);
5175: if (err)
5176: goto sync_free_meta;
5177: }
5178:
5179: /* recover broken superblock */
5180: if (recovery) {
5181: err = f2fs_commit_super(sbi, true);
5182: f2fs_info(sbi, "Try to recover %dth superblock, ret: %d",
5183: sbi->valid_super_block ? 1 : 2, err);
5184: }
5185:
5186: f2fs_join_shrinker(sbi);
5187:
5188: f2fs_tuning_parameters(sbi);
5189:
5190: f2fs_notice(sbi, "Mounted with checkpoint version = %llx",
5191: cur_cp_version(F2FS_CKPT(sbi)));
5192: f2fs_update_time(sbi, CP_TIME);
5193: f2fs_update_time(sbi, REQ_TIME);
5194: clear_sbi_flag(sbi, SBI_CP_DISABLED_QUICK);
5195:
5196: sbi->umount_lock_holder = NULL;
5197: return 0;
5198:
5199: sync_free_meta:
5200: /* safe to flush all the data */
5201: sync_filesystem(sbi->sb);
5202: retry_cnt = 0;
5203:
5204: free_meta:
5205: #ifdef CONFIG_QUOTA
5206: f2fs_truncate_quota_inode_pages(sb);
5207: if (f2fs_sb_has_quota_ino(sbi) && !f2fs_readonly(sb))
5208: f2fs_quota_off_umount(sbi->sb);
5209: #endif
5210: /*
5211: * Some dirty meta pages can be produced by f2fs_recover_orphan_inodes()
5212: * failed by EIO. Then, iput(node_inode) can trigger balance_fs_bg()
5213: * followed by f2fs_write_checkpoint() through f2fs_write_node_pages(), which
5214: * falls into an infinite loop in f2fs_sync_meta_pages().
5215: */
5216: truncate_inode_pages_final(META_MAPPING(sbi));
5217: /* evict some inodes being cached by GC */
5218: evict_inodes(sb);
5219: f2fs_unregister_sysfs(sbi);
5220: free_compress_inode:
5221: f2fs_destroy_compress_inode(sbi);
5222: free_root_inode:
5223: dput(sb->s_root);
5224: sb->s_root = NULL;
5225: free_node_inode:
5226: f2fs_release_ino_entry(sbi, true);
5227: truncate_inode_pages_final(NODE_MAPPING(sbi));
5228: iput(sbi->node_inode);
5229: sbi->node_inode = NULL;
5230: free_stats:
5231: f2fs_destroy_stats(sbi);
5232: free_nm:
5233: /* stop discard thread before destroying node manager */
5234: f2fs_stop_discard_thread(sbi);
5235: f2fs_destroy_node_manager(sbi);
5236: free_sm:
5237: f2fs_destroy_segment_manager(sbi);
5238: stop_ckpt_thread:
5239: f2fs_stop_ckpt_thread(sbi);
5240: /* flush s_error_work before sbi destroy */
5241: flush_work(&sbi->s_error_work);
5242: f2fs_destroy_post_read_wq(sbi);
5243: free_devices:
5244: destroy_device_list(sbi);
5245: kvfree(sbi->ckpt);
5246: free_meta_inode:
5247: make_bad_inode(sbi->meta_inode);
5248: iput(sbi->meta_inode);
5249: sbi->meta_inode = NULL;
5250: free_page_array_cache:
5251: f2fs_destroy_page_array_cache(sbi);
5252: free_xattr_cache:
5253: f2fs_destroy_xattr_caches(sbi);
5254: free_percpu:
5255: destroy_percpu_info(sbi);
5256: free_iostat:
5257: f2fs_destroy_iostat(sbi);
5258: free_bio_info:
5259: for (i = 0; i < NR_PAGE_TYPE; i++)
5260: kfree(sbi->write_io[i]);
5261:
5262: #if IS_ENABLED(CONFIG_UNICODE)
5263: utf8_unload(sb->s_encoding);
5264: sb->s_encoding = NULL;
5265: #endif
5266: free_options:
5267: #ifdef CONFIG_QUOTA
5268: for (i = 0; i < MAXQUOTAS; i++)
5269: kfree(F2FS_OPTION(sbi).s_qf_names[i]);
5270: #endif
5271: /* no need to free dummy_enc_policy, we just keep it in ctx when failed */
5272: swap(F2FS_CTX_INFO(ctx).dummy_enc_policy, F2FS_OPTION(sbi).dummy_enc_policy);
5273: free_sb_buf:
5274: kfree(raw_super);
5275: free_sbi:
5276: kfree(sbi);
5277: sb->s_fs_info = NULL;
5278:
5279: /* give only one another chance */
5280: if (retry_cnt > 0 && skip_recovery) {
5281: retry_cnt--;
5282: shrink_dcache_sb(sb);
5283: goto try_onemore;
5284: }
5285: return err;
5286: }
SourceFile:fs/f2fs/super.c]
|
| 24/2 |
2026/05/26 06:34 |
llm |
expert |
0m
Model:
Tokens:
input: 63192
output: 37
thoughts: 208
Thoughts: **Analyzing the F2FS Filesystem Initialization**
Okay, so I'm tracing the boot process here and I've zeroed in on the `f2fs_start_ckpt_thread` function, called right at line 4985. Seems like that's the trigger for a significant chunk of the initialization. After that, we've got a sequence of builders: `f2fs_build_segment_manager`, `f2fs_build_node_manager`, `f2fs_build_gc_manager`, `f2fs_build_stats`. Makes sense, we need to get all the key structures set up. Then the `f2fs_iget` calls for the node and root inodes, followed by some recovery routines: `f2fs_recover_orphan_inodes`, and `f2fs_recover_fsync_data`. Standard stuff.
But here's a potential point of failure. If *any* of those initializations or recoveries hit a snag, it looks like we're heading straight to `free_meta` or `sync_free_meta`. And wait a minute... if it's `sync_free_meta` then we end up calling `sync_filesystem(sbi->sb)`. Hmmm, I was sure `sync_filesystem` had a different purpose. Ah, yes, I remember now, I got it mixed up. *`sync_filesystem` takes `s_umount` for read!* Right, right. I'm going to double-check that `sync_filesystem` call. Gotta be sure about this.
|
| 25/2 |
2026/05/26 06:34 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/sync.c IncludeLines:true Name:sync_filesystem]
Results: map[SourceCode: 25: /*
26: * Write out and wait upon all dirty data associated with this
27: * superblock. Filesystem data as well as the underlying block
28: * device. Takes the superblock lock.
29: */
30: int sync_filesystem(struct super_block *sb)
31: {
32: int ret = 0;
33:
34: /*
35: * We need to be protected against the filesystem going from
36: * r/o to r/w or vice versa.
37: */
38: WARN_ON(!rwsem_is_locked(&sb->s_umount));
39:
40: /*
41: * No point in syncing out anything if the filesystem is read-only.
42: */
43: if (sb_rdonly(sb))
44: return 0;
45:
46: /*
47: * Do the filesystem syncing work. For simple filesystems
48: * writeback_inodes_sb(sb) just dirties buffers with inodes so we have
49: * to submit I/O for these buffers via sync_blockdev(). This also
50: * speeds up the wait == 1 case since in that case write_inode()
51: * methods call sync_dirty_buffer() and thus effectively write one block
52: * at a time.
53: */
54: writeback_inodes_sb(sb, WB_REASON_SYNC);
55: if (sb->s_op->sync_fs) {
56: ret = sb->s_op->sync_fs(sb, 0);
57: if (ret)
58: return ret;
59: }
60: ret = sync_blockdev_nowait(sb->s_bdev);
61: if (ret)
62: return ret;
63:
64: sync_inodes_sb(sb);
65: if (sb->s_op->sync_fs) {
66: ret = sb->s_op->sync_fs(sb, 1);
67: if (ret)
68: return ret;
69: }
70: return sync_blockdev(sb->s_bdev);
71: }
SourceFile:fs/sync.c]
|
| 26/2 |
2026/05/26 06:34 |
llm |
expert |
0m
Model:
Tokens:
input: 64441
output: 43
thoughts: 125
Thoughts: **Analyzing the F2FS Filesystem Synchronization Process**
Okay, so I'm looking at this `sync_filesystem` function, and it seems like it *requires* the `s_umount` lock to be held. That's the starting point. But, hang on a minute... I recall that `f2fs_fill_super`βthe function responsible for actually filling in the superblockβis called with `s_umount` held for *write* access! So, that means `sync_filesystem` is indeed called *with* `s_umount` already held for write. Makes sense, in principle.
Now I'm diving deeper into the implications. It looks like `sync_filesystem` calls `writeback_inodes_sb`. And *that* calls `bdi_writeback`. Now, does it queue work to `bdi_writeback` and then *wait* for it? Or does it just queue the work and wait? Let me confirm. *No*, it is `writeback_inodes_sb` that queues the work and *then* waits for it. That's the critical distinction.
Okay, let's now check the implementation of `writeback_inodes_sb` to confirm my understanding. Let's make sure I didn't miss something.
|
| 27/2 |
2026/05/26 06:35 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/fs-writeback.c IncludeLines:true Name:writeback_inodes_sb]
Results: map[SourceCode:2769: /**
2770: * writeback_inodes_sb - writeback dirty inodes from given super_block
2771: * @sb: the superblock
2772: * @reason: reason why some writeback work was initiated
2773: *
2774: * Start writeback on some inodes on this super_block. No guarantees are made
2775: * on how many (if any) will be written, and this function does not wait
2776: * for IO completion of submitted IO.
2777: */
2778: void writeback_inodes_sb(struct super_block *sb, enum wb_reason reason)
2779: {
2780: writeback_inodes_sb_nr(sb, get_nr_dirty_pages(), reason);
2781: }
SourceFile:fs/fs-writeback.c]
|
| 28/2 |
2026/05/26 06:35 |
llm |
expert |
0m
Model:
Tokens:
input: 65136
output: 42
thoughts: 28
Thoughts: Here's my take on this, considering I'm trying to think like an expert in this field:
**Internal IO Synchronization Conundrum**
Okay, hold on a second. I need to make sure I'm following this correctly. `writeback_inodes_sb`... that's *not* a synchronous operation, is it? It schedules the IO, but it doesn't block waiting for it to actually finish. We're talking about background writeback, right? It's optimized for that asynchronous behavior.
But then, if thatβs the case, *`sync_inodes_sb` IS* supposed to wait, isn't it? It has to ensure data is actually on persistent storage, before proceeding and acknowledging that the sync is complete. That has always been the fundamental difference, hasn't it? One schedules, the other synchronizes. It's critical to understanding the guarantees we're offering (or not offering) to the application layer. I need to keep that distinction absolutely clear in my head. They're related, but definitely not the same. Let's make sure that's represented correctly downstream in the code.
|
| 29/2 |
2026/05/26 06:35 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/fs-writeback.c IncludeLines:true Name:sync_inodes_sb]
Results: map[SourceCode:2801: /**
2802: * sync_inodes_sb - sync sb inode pages
2803: * @sb: the superblock
2804: *
2805: * This function writes and waits on any dirty inode belonging to this
2806: * super_block.
2807: */
2808: void sync_inodes_sb(struct super_block *sb)
2809: {
2810: struct backing_dev_info *bdi = sb->s_bdi;
2811: DEFINE_WB_COMPLETION(done, bdi);
2812: struct wb_writeback_work work = {
2813: .sb = sb,
2814: .sync_mode = WB_SYNC_ALL,
2815: .nr_pages = LONG_MAX,
2816: .range_cyclic = 0,
2817: .done = &done,
2818: .reason = WB_REASON_SYNC,
2819: .for_sync = 1,
2820: };
2821:
2822: /*
2823: * Can't skip on !bdi_has_dirty() because we should wait for !dirty
2824: * inodes under writeback and I_DIRTY_TIME inodes ignored by
2825: * bdi_has_dirty() need to be written out too.
2826: */
2827: if (bdi == &noop_backing_dev_info)
2828: return;
2829: WARN_ON(!rwsem_is_locked(&sb->s_umount));
2830:
2831: /* protect against inode wb switch, see inode_switch_wbs_work_fn() */
2832: bdi_down_write_wb_switch_rwsem(bdi);
2833: bdi_split_work_to_wbs(bdi, &work, false);
2834: wb_wait_for_completion(&done);
2835: bdi_up_write_wb_switch_rwsem(bdi);
2836:
2837: wait_sb_inodes(sb);
2838: }
SourceFile:fs/fs-writeback.c]
|
| 30/2 |
2026/05/26 06:35 |
llm |
expert |
0m
Model:
Tokens:
input: 66087
output: 40
thoughts: 296
Thoughts: **Analyzing the Filesystem Sync Process**
Okay, so `sync_inodes_sb` is clearly submitting work to the `bdi` and then waiting. That makes sense. And, as expected, `kworker` threads are the ones actually executing that queued work. Now, let's see what locks these `kworker` threads are holding. Aha, I see: `kworker/u8:0/12` is reporting holding three locks. The first two seem tied to the workqueue completion itself, no surprises there. They are internal to the workqueue. But the third one, that's interesting: `&type->s_umount_key#72` is held read! Wait a minute... that suggests a shared lock, and it's being held by a `kworker`. Okay, let's trace this back. The `super_trylock_shared` function is the culprit. This function is the one taking the `s_umount` lock, in shared mode. Let's dig deeper into that function to understand the implications of this locking behavior.
|
| 31/2 |
2026/05/26 06:35 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/super.c IncludeLines:true Name:super_trylock_shared]
Results: map[SourceCode: 544: /*
545: * super_trylock_shared - try to grab ->s_umount shared
546: * @sb: reference we are trying to grab
547: *
548: * Try to prevent fs shutdown. This is used in places where we
549: * cannot take an active reference but we need to ensure that the
550: * filesystem is not shut down while we are working on it. It returns
551: * false if we cannot acquire s_umount or if we lose the race and
552: * filesystem already got into shutdown, and returns true with the s_umount
553: * lock held in read mode in case of success. On successful return,
554: * the caller must drop the s_umount lock when done.
555: *
556: * Note that unlike get_super() et.al. this one does *not* bump ->s_count.
557: * The reason why it's safe is that we are OK with doing trylock instead
558: * of down_read(). There's a couple of places that are OK with that, but
559: * it's very much not a general-purpose interface.
560: */
561: bool super_trylock_shared(struct super_block *sb)
562: {
563: if (down_read_trylock(&sb->s_umount)) {
564: if (!(sb->s_flags & SB_DYING) && sb->s_root &&
565: (sb->s_flags & SB_BORN))
566: return true;
567: super_unlock_shared(sb);
568: }
569:
570: return false;
571: }
SourceFile:fs/super.c]
|
| 32/2 |
2026/05/26 06:35 |
llm |
expert |
0m
Model:
Tokens:
input: 67185
output: 24
thoughts: 2172
Thoughts: **Analysis of a Potential F2FS Hang Vulnerability**
Okay, let's break this down. It looks like `super_trylock_shared` is being used, specifically `down_read_trylock` on `s_umount`. If this succeeds, a read lock is acquired. The problem arises if `f2fs_fill_super` holds `s_umount` for write, as `down_read_trylock` will fail. It seems there's a contradiction, because I see that a `kworker` *is* holding `s_umount` for read. The key here is there are multiple superblocks on this system. One is probably the F2FS filesystem, another may be ext4.
The system task `syz.2.19` is blocked on a different superblock; it's doing `quotactl` which translates into a call to `__iterate_supers` which then leads to a call to `super_lock`. This function iterates over *all* superblocks, so it's trying to lock the F2FS superblock, but it's blocked. The F2FS superblock is likely held for write by whatever task is performing the `mount` operation. This strongly indicates that `f2fs_fill_super` is stuck or caught in a deadlock during the `mount` process, holding the write lock indefinitely.
Let's dig into that `f2fs_fill_super` to check for problematic behavior. The trace shows that `f2fs_ckpt-7:2/6140` is running which means `f2fs_start_ckpt_thread` was called. That thread is called at line 4985 of `f2fs_fill_super`. After that the function then performs these operations: `f2fs_build_segment_manager`, `f2fs_build_node_manager`, `f2fs_build_gc_manager`, `f2fs_build_stats`, plus retrieval of the node and root inodes, plus calls to `f2fs_recover_orphan_inodes`, and `f2fs_recover_fsync_data`. Either of the last two functions could take a significant amount of time, or they could get hung up with a crafted image.
Based on the bug report's details which indicated a task blocked for over 143 seconds, the problem seems to be an indefinite hang. This means that `f2fs_fill_super` is holding `s_umount` for write forever. Consequently, operations that iterate over superblocks, like `quotactl`, and `sync` will also block. This creates a system-wide hang.
Is this exploitable? Yes, as a Denial of Service (DoS). An unprivileged user shouldn't be able to trigger this because F2FS doesn't use the `FS_USERNS_MOUNT` flag. Thus, unprivileged users can't mount it. However, if a user plugs in a USB drive and it auto-mounts, this could be exploited because auto-mounting is done by a privileged process (e.g. udisks2). So, it's a "peripheral trigger" and a "malicious filesystem trigger".
Since the reproducer uses `syz_mount_image`, which calls `mount()`, and `syz-executor` runs as root in the test environment, this indicates that the bug is *not* accessible to unprivileged users, or even user namespaces. I need to make sure I am not missing any other way to call `f2fs_fill_super`. Nope, only `mount` can trigger this code.
So:
- Exploitable: No (hang, not memory corruption).
- Denial Of Service: Yes (system-wide `sync` and `quotactl` hang).
- Unprivileged Access: No.
- User Namespace Access: No.
- VM Guest Trigger: No.
- VM Host Trigger: No.
- Network Trigger: No.
- Other Remote Trigger: No.
- Peripheral Trigger: Yes (USB auto-mount).
- Malicious Filesystem Trigger: Yes.
Now I need to re-examine my assumptions about how `f2fs_fill_super` uses `s_umount`. The `mount` operation calls `vfs_get_tree`, which in turn calls `fc->ops->get_tree` -> `f2fs_get_tree` -> `get_tree_bdev` -> `f2fs_fill_super`. `sget_dev` *does* return a superblock with `s_umount` held for write. `deactivate_locked_super` or the caller of `vfs_get_tree` will drop it. Specifically, `vfs_get_
|
| 33/2 |
2026/05/26 06:35 |
tool |
grepper |
0m
Args: map[Expression:f2fs_recover_fsync_data]
Results: map[Output:fs/f2fs/f2fs.h=4046=int f2fs_get_victim(struct f2fs_sb_info *sbi, unsigned int *result,
--
fs/f2fs/f2fs.h-4052- */
fs/f2fs/f2fs.h:4053:int f2fs_recover_fsync_data(struct f2fs_sb_info *sbi, bool check_only);
fs/f2fs/f2fs.h-4054-bool f2fs_space_for_roll_forward(struct f2fs_sb_info *sbi);
--
fs/f2fs/recovery.c=776=static int recover_data(struct f2fs_sb_info *sbi, struct list_head *inode_list,
--
fs/f2fs/recovery.c-869-
fs/f2fs/recovery.c:870:int f2fs_recover_fsync_data(struct f2fs_sb_info *sbi, bool check_only)
fs/f2fs/recovery.c-871-{
--
fs/f2fs/recovery.c-878-
fs/f2fs/recovery.c:879: f2fs_notice(sbi, "f2fs_recover_fsync_data: recovery fsync data, "
fs/f2fs/recovery.c-880- "check_only: %d", check_only);
--
fs/f2fs/super.c=4770=static int f2fs_fill_super(struct super_block *sb, struct fs_context *fc)
--
fs/f2fs/super.c-5096- if (!is_set_ckpt_flags(sbi, CP_UMOUNT_FLAG)) {
fs/f2fs/super.c:5097: err = f2fs_recover_fsync_data(sbi, true);
fs/f2fs/super.c-5098- if (err > 0) {
--
fs/f2fs/super.c-5116-
fs/f2fs/super.c:5117: err = f2fs_recover_fsync_data(sbi, false);
fs/f2fs/super.c-5118- if (err < 0) {
--
fs/f2fs/super.c-5126- } else {
fs/f2fs/super.c:5127: err = f2fs_recover_fsync_data(sbi, true);
fs/f2fs/super.c-5128-
--
fs/f2fs/super.c-5151-
fs/f2fs/super.c:5152: /* f2fs_recover_fsync_data() cleared this already */
fs/f2fs/super.c-5153- clear_sbi_flag(sbi, SBI_POR_DOING);
]
|
| 34/2 |
2026/05/26 06:35 |
llm |
expert |
0m
Model:
Tokens:
input: 71159
output: 48
thoughts: 28
Thoughts: Here's my take on the thought process, as if it were my own:
**Diving into F2FS fsync Recovery**
Okay, time to get my hands dirty. I need to understand the ins and outs of how F2FS recovers `fsync` data, particularly after a crash or power loss. That means heading straight into the `fs/f2fs/recovery.c` file. I'll need to drill down into the `f2fs_recover_fsync_data` function specifically. This is where the magic (or potentially the troubleshooting) happens. I'm expecting this function will be the core of handling any partially-written transactions or inconsistent metadata related to `fsync` operations during recovery. Let's see what gems are waiting to be uncovered.
|
| 35/2 |
2026/05/26 06:35 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/recovery.c IncludeLines:true Name:f2fs_recover_fsync_data]
Results: map[SourceCode: 870: int f2fs_recover_fsync_data(struct f2fs_sb_info *sbi, bool check_only)
871: {
872: struct list_head inode_list, tmp_inode_list;
873: struct list_head dir_list;
874: int err;
875: int ret = 0;
876: unsigned long s_flags = sbi->sb->s_flags;
877: bool need_writecp = false;
878:
879: f2fs_notice(sbi, "f2fs_recover_fsync_data: recovery fsync data, "
880: "check_only: %d", check_only);
881:
882: if (is_sbi_flag_set(sbi, SBI_IS_WRITABLE))
883: f2fs_info(sbi, "recover fsync data on readonly fs");
884:
885: INIT_LIST_HEAD(&inode_list);
886: INIT_LIST_HEAD(&tmp_inode_list);
887: INIT_LIST_HEAD(&dir_list);
888:
889: /* prevent checkpoint */
890: f2fs_down_write(&sbi->cp_global_sem);
891:
892: /* step #1: find fsynced inode numbers */
893: err = find_fsync_dnodes(sbi, &inode_list, check_only);
894: if (err || list_empty(&inode_list))
895: goto skip;
896:
897: if (check_only) {
898: ret = 1;
899: goto skip;
900: }
901:
902: need_writecp = true;
903:
904: /* step #2: recover data */
905: err = recover_data(sbi, &inode_list, &tmp_inode_list, &dir_list);
906: if (!err)
907: f2fs_bug_on(sbi, !list_empty(&inode_list));
908: else
909: f2fs_bug_on(sbi, sbi->sb->s_flags & SB_ACTIVE);
910: skip:
911: destroy_fsync_dnodes(&inode_list, err);
912: destroy_fsync_dnodes(&tmp_inode_list, err);
913:
914: /* truncate meta pages to be used by the recovery */
915: truncate_inode_pages_range(META_MAPPING(sbi),
916: (loff_t)MAIN_BLKADDR(sbi) << PAGE_SHIFT, -1);
917:
918: if (err) {
919: truncate_inode_pages_final(NODE_MAPPING(sbi));
920: truncate_inode_pages_final(META_MAPPING(sbi));
921: }
922:
923: /*
924: * If fsync data succeeds or there is no fsync data to recover,
925: * and the f2fs is not read only, check and fix zoned block devices'
926: * write pointer consistency.
927: */
928: if (!err)
929: err = f2fs_check_and_fix_write_pointer(sbi);
930:
931: if (!err)
932: clear_sbi_flag(sbi, SBI_POR_DOING);
933:
934: f2fs_up_write(&sbi->cp_global_sem);
935:
936: /* let's drop all the directory inodes for clean checkpoint */
937: destroy_fsync_dnodes(&dir_list, err);
938:
939: if (need_writecp) {
940: set_sbi_flag(sbi, SBI_IS_RECOVERED);
941:
942: if (!err) {
943: struct cp_control cpc = {
944: .reason = CP_RECOVERY,
945: };
946: stat_inc_cp_call_count(sbi, TOTAL_CALL);
947: err = f2fs_write_checkpoint(sbi, &cpc);
948: }
949: }
950:
951: sbi->sb->s_flags = s_flags; /* Restore SB_RDONLY status */
952:
953: return ret ? ret : err;
954: }
SourceFile:fs/f2fs/recovery.c]
|
| 36/2 |
2026/05/26 06:35 |
llm |
expert |
0m
Model:
Tokens:
input: 72759
output: 45
thoughts: 91
Thoughts: **Analyzing a Potential F2FS Hang During Filesystem Recovery**
Okay, so I'm digging into this F2FS filesystem recovery issue. My initial thought is that `f2fs_recover_fsync_data` seems to be calling `f2fs_write_checkpoint`. Right, and `f2fs_write_checkpoint` is responsible for, well, writing the checkpoint. Now, if `f2fs_write_checkpoint` hangs, then from what I'm seeing, it appears `f2fs_fill_super` is going to hang, which is obviously a problem. The million dollar question then becomes: why would `f2fs_write_checkpoint` hang in the first place? To get a handle on this, I need to take a closer look at the `f2fs_write_checkpoint` function itself. Let's see what's going on there.
|
| 37/2 |
2026/05/26 06:35 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/checkpoint.c IncludeLines:true Name:f2fs_write_checkpoint]
Results: map[SourceCode:1610: int f2fs_write_checkpoint(struct f2fs_sb_info *sbi, struct cp_control *cpc)
1611: {
1612: struct f2fs_checkpoint *ckpt = F2FS_CKPT(sbi);
1613: unsigned long long ckpt_ver;
1614: int err = 0;
1615:
1616: if (f2fs_readonly(sbi->sb) || f2fs_hw_is_readonly(sbi))
1617: return -EROFS;
1618:
1619: if (unlikely(is_sbi_flag_set(sbi, SBI_CP_DISABLED))) {
1620: if (cpc->reason != CP_PAUSE)
1621: return 0;
1622: f2fs_warn(sbi, "Start checkpoint disabled!");
1623: }
1624: if (cpc->reason != CP_RESIZE)
1625: f2fs_down_write(&sbi->cp_global_sem);
1626:
1627: if (!is_sbi_flag_set(sbi, SBI_IS_DIRTY) &&
1628: ((cpc->reason & CP_FASTBOOT) || (cpc->reason & CP_SYNC) ||
1629: ((cpc->reason & CP_DISCARD) && !sbi->discard_blks)))
1630: goto out;
1631: if (unlikely(f2fs_cp_error(sbi))) {
1632: err = -EIO;
1633: goto out;
1634: }
1635:
1636: trace_f2fs_write_checkpoint(sbi->sb, cpc->reason, "start block_ops");
1637:
1638: err = block_operations(sbi);
1639: if (err)
1640: goto out;
1641:
1642: trace_f2fs_write_checkpoint(sbi->sb, cpc->reason, "finish block_ops");
1643:
1644: f2fs_flush_merged_writes(sbi);
1645:
1646: /* this is the case of multiple fstrims without any changes */
1647: if (cpc->reason & CP_DISCARD) {
1648: if (!f2fs_exist_trim_candidates(sbi, cpc)) {
1649: unblock_operations(sbi);
1650: goto out;
1651: }
1652:
1653: if (NM_I(sbi)->nat_cnt[DIRTY_NAT] == 0 &&
1654: SIT_I(sbi)->dirty_sentries == 0 &&
1655: prefree_segments(sbi) == 0) {
1656: f2fs_flush_sit_entries(sbi, cpc);
1657: f2fs_clear_prefree_segments(sbi, cpc);
1658: unblock_operations(sbi);
1659: goto out;
1660: }
1661: }
1662:
1663: /*
1664: * update checkpoint pack index
1665: * Increase the version number so that
1666: * SIT entries and seg summaries are written at correct place
1667: */
1668: ckpt_ver = cur_cp_version(ckpt);
1669: ckpt->checkpoint_ver = cpu_to_le64(++ckpt_ver);
1670:
1671: /* write cached NAT/SIT entries to NAT/SIT area */
1672: err = f2fs_flush_nat_entries(sbi, cpc);
1673: if (err) {
1674: f2fs_err(sbi, "f2fs_flush_nat_entries failed err:%d, stop checkpoint", err);
1675: f2fs_bug_on(sbi, !f2fs_cp_error(sbi));
1676: goto stop;
1677: }
1678:
1679: f2fs_flush_sit_entries(sbi, cpc);
1680:
1681: /* save inmem log status */
1682: f2fs_save_inmem_curseg(sbi);
1683:
1684: err = do_checkpoint(sbi, cpc);
1685: if (err) {
1686: f2fs_err(sbi, "do_checkpoint failed err:%d, stop checkpoint", err);
1687: f2fs_bug_on(sbi, !f2fs_cp_error(sbi));
1688: f2fs_release_discard_addrs(sbi);
1689: } else {
1690: f2fs_clear_prefree_segments(sbi, cpc);
1691: }
1692:
1693: f2fs_restore_inmem_curseg(sbi);
1694: f2fs_reinit_atgc_curseg(sbi);
1695: stat_inc_cp_count(sbi);
1696: stop:
1697: unblock_operations(sbi);
1698:
1699: if (cpc->reason & CP_RECOVERY)
1700: f2fs_notice(sbi, "checkpoint: version = %llx", ckpt_ver);
1701:
1702: /* update CP_TIME to trigger checkpoint periodically */
1703: f2fs_update_time(sbi, CP_TIME);
1704: trace_f2fs_write_checkpoint(sbi->sb, cpc->reason, "finish checkpoint");
1705: out:
1706: if (cpc->reason != CP_RESIZE)
1707: f2fs_up_write(&sbi->cp_global_sem);
1708: return err;
1709: }
SourceFile:fs/f2fs/checkpoint.c]
|
| 38/2 |
2026/05/26 06:35 |
llm |
expert |
0m
Model:
Tokens:
input: 74785
output: 41
thoughts: 31
Thoughts: Okay, here's my take on those thoughts, framed as an expert's internal monologue:
**Investigating the F2FS Checkpoint Write**
Alright, hang on a second. I'm focusing on this `f2fs_write_checkpoint` function, and it seems like it's calling `block_operations(sbi)`. Interesting. That means it's likely interacting directly with the lower-level block device. Before I go further down that rabbit hole, I need to understand what `block_operations` *is* actually doing here. It's the key to understanding how they're handling this at the block level. I'll need to dig into the `block_operations` code to see how they're handling the raw I/O.
|
| 39/2 |
2026/05/26 06:35 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/checkpoint.c IncludeLines:true Name:block_operations]
Results: map[SourceCode:1204: /*
1205: * Freeze all the FS-operations for checkpoint.
1206: */
1207: static int block_operations(struct f2fs_sb_info *sbi)
1208: {
1209: struct writeback_control wbc = {
1210: .sync_mode = WB_SYNC_ALL,
1211: .nr_to_write = LONG_MAX,
1212: };
1213: int err = 0, cnt = 0;
1214:
1215: /*
1216: * Let's flush inline_data in dirty node pages.
1217: */
1218: f2fs_flush_inline_data(sbi);
1219:
1220: retry_flush_quotas:
1221: f2fs_lock_all(sbi);
1222: if (__need_flush_quota(sbi)) {
1223: bool need_lock = sbi->umount_lock_holder != current;
1224:
1225: if (++cnt > DEFAULT_RETRY_QUOTA_FLUSH_COUNT) {
1226: set_sbi_flag(sbi, SBI_QUOTA_SKIP_FLUSH);
1227: set_sbi_flag(sbi, SBI_QUOTA_NEED_FLUSH);
1228: goto retry_flush_dents;
1229: }
1230: f2fs_unlock_all(sbi);
1231:
1232: /* don't grab s_umount lock during mount/umount/remount/freeze/quotactl */
1233: if (!need_lock) {
1234: f2fs_do_quota_sync(sbi->sb, -1);
1235: } else if (down_read_trylock(&sbi->sb->s_umount)) {
1236: f2fs_do_quota_sync(sbi->sb, -1);
1237: up_read(&sbi->sb->s_umount);
1238: }
1239: cond_resched();
1240: goto retry_flush_quotas;
1241: }
1242:
1243: retry_flush_dents:
1244: /* write all the dirty dentry pages */
1245: if (get_pages(sbi, F2FS_DIRTY_DENTS)) {
1246: f2fs_unlock_all(sbi);
1247: err = f2fs_sync_dirty_inodes(sbi, DIR_INODE, true);
1248: if (err)
1249: return err;
1250: cond_resched();
1251: goto retry_flush_quotas;
1252: }
1253:
1254: /*
1255: * POR: we should ensure that there are no dirty node pages
1256: * until finishing nat/sit flush. inode->i_blocks can be updated.
1257: */
1258: f2fs_down_write(&sbi->node_change);
1259:
1260: if (get_pages(sbi, F2FS_DIRTY_IMETA)) {
1261: f2fs_up_write(&sbi->node_change);
1262: f2fs_unlock_all(sbi);
1263: err = f2fs_sync_inode_meta(sbi);
1264: if (err)
1265: return err;
1266: cond_resched();
1267: goto retry_flush_quotas;
1268: }
1269:
1270: retry_flush_nodes:
1271: f2fs_down_write(&sbi->node_write);
1272:
1273: if (get_pages(sbi, F2FS_DIRTY_NODES)) {
1274: f2fs_up_write(&sbi->node_write);
1275: atomic_inc(&sbi->wb_sync_req[NODE]);
1276: err = f2fs_sync_node_pages(sbi, &wbc, false, FS_CP_NODE_IO);
1277: atomic_dec(&sbi->wb_sync_req[NODE]);
1278: if (err) {
1279: f2fs_up_write(&sbi->node_change);
1280: f2fs_unlock_all(sbi);
1281: return err;
1282: }
1283: cond_resched();
1284: goto retry_flush_nodes;
1285: }
1286:
1287: /*
1288: * sbi->node_change is used only for AIO write_begin path which produces
1289: * dirty node blocks and some checkpoint values by block allocation.
1290: */
1291: __prepare_cp_block(sbi);
1292: f2fs_up_write(&sbi->node_change);
1293: return err;
1294: }
SourceFile:fs/f2fs/checkpoint.c]
|
| 40/2 |
2026/05/26 06:35 |
llm |
expert |
0m
Model:
Tokens:
input: 76480
output: 47
thoughts: 253
Thoughts: **Infinite Loop Concern in Block Operations**
Okay, so I'm looking at this `block_operations` code and immediately I see a loop, specifically around the quota and dent flushing:
```c
retry_flush_quotas:
f2fs_lock_all(sbi);
if (__need_flush_quota(sbi)) {
...
goto retry_flush_quotas;
}
retry_flush_dents:
if (get_pages(sbi, F2FS_DIRTY_DENTS)) {
f2fs_unlock_all(sbi);
err = f2fs_sync_dirty_inodes(sbi, DIR_INODE, true);
if (err)
return err;
cond_resched();
goto retry_flush_quotas;
}
```
My initial concern is that if `f2fs_sync_dirty_inodes` returns 0, indicating no apparent error, but *doesn't* actually clear `F2FS_DIRTY_DENTS`, this loop could go on forever. That's a critical issue, and I need to verify that scenario.
Then I am wondering, could there be fault injection, for example, that would result in a zero return from `f2fs_sync_dirty_inodes` that wouldn't necessarily result in it clearing all dirty dents? I need to check `f2fs_sync_dirty_inodes` itself and see if this is possible. I'll check that routine next.
|
| 41/2 |
2026/05/26 06:36 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/inode.c IncludeLines:true Name:f2fs_sync_dirty_inodes]
Results: map[SourceCode:1068: int f2fs_sync_dirty_inodes(struct f2fs_sb_info *sbi, enum inode_type type,
1069: bool from_cp)
1070: {
1071: struct list_head *head;
1072: struct inode *inode;
1073: struct f2fs_inode_info *fi;
1074: bool is_dir = (type == DIR_INODE);
1075: unsigned long ino = 0;
1076:
1077: trace_f2fs_sync_dirty_inodes_enter(sbi->sb, is_dir,
1078: get_pages(sbi, is_dir ?
1079: F2FS_DIRTY_DENTS : F2FS_DIRTY_DATA));
1080: retry:
1081: if (unlikely(f2fs_cp_error(sbi))) {
1082: trace_f2fs_sync_dirty_inodes_exit(sbi->sb, is_dir,
1083: get_pages(sbi, is_dir ?
1084: F2FS_DIRTY_DENTS : F2FS_DIRTY_DATA));
1085: return -EIO;
1086: }
1087:
1088: spin_lock(&sbi->inode_lock[type]);
1089:
1090: head = &sbi->inode_list[type];
1091: if (list_empty(head)) {
1092: spin_unlock(&sbi->inode_lock[type]);
1093: trace_f2fs_sync_dirty_inodes_exit(sbi->sb, is_dir,
1094: get_pages(sbi, is_dir ?
1095: F2FS_DIRTY_DENTS : F2FS_DIRTY_DATA));
1096: return 0;
1097: }
1098: fi = list_first_entry(head, struct f2fs_inode_info, dirty_list);
1099: inode = igrab(&fi->vfs_inode);
1100: spin_unlock(&sbi->inode_lock[type]);
1101: if (inode) {
1102: unsigned long cur_ino = inode->i_ino;
1103:
1104: if (from_cp)
1105: F2FS_I(inode)->cp_task = current;
1106: F2FS_I(inode)->wb_task = current;
1107:
1108: filemap_fdatawrite(inode->i_mapping);
1109:
1110: F2FS_I(inode)->wb_task = NULL;
1111: if (from_cp)
1112: F2FS_I(inode)->cp_task = NULL;
1113:
1114: iput(inode);
1115: /* We need to give cpu to another writers. */
1116: if (ino == cur_ino)
1117: cond_resched();
1118: else
1119: ino = cur_ino;
1120: } else {
1121: /*
1122: * We should submit bio, since it exists several
1123: * writebacking dentry pages in the freeing inode.
1124: */
1125: f2fs_submit_merged_write(sbi, DATA);
1126: cond_resched();
1127: }
1128: goto retry;
1129: }
SourceFile:fs/f2fs/checkpoint.c]
|
| 42/2 |
2026/05/26 06:36 |
llm |
expert |
0m
Model:
Tokens:
input: 78132
output: 47
thoughts: 723
Thoughts: **Infinite Loop Bug in F2FS Error Handling**
Okay, so I'm looking at this `f2fs_sync_dirty_inodes` function, and it's looping over this `sbi->inode_list[type]` structure, right? It calls `filemap_fdatawrite` on each inode's mapping. My initial thought was: if `filemap_fdatawrite` fails due to, say, fault injection, does it remove the inode from the list? No, it just *starts* writeback. The inode is removed later, when the pages are *actually* written and cleaned.
This is where it gets interesting. If fault injection causes the write to fail, the pages might stay dirty. This means the inode could stay on the dirty list! Then `f2fs_sync_dirty_inodes` will loop forever because the `list_empty(head)` check will *never* be true. Hold on, `filemap_fdatawrite` *does* return an error if it fails! But `f2fs_sync_dirty_inodes` *ignores* the return value! It just blindly calls `filemap_fdatawrite` and moves on. So if `filemap_fdatawrite` can't clean the pages, the inode stays dirty, and we've got an infinite loop. This is a classic infinite loop bug, a real problem in F2FS error handling!
The issue arises during `f2fs_fill_super`, then `f2fs_recover_fsync_data`, then `f2fs_write_checkpoint`, then `block_operations` and finally `f2fs_sync_dirty_inodes`.
I'm realizing that `f2fs_fill_super` holds `s_umount` for write. This is a significant issue because that means `s_umount` is held for write, potentially forever. This blocks `quotactl`, `sync`, and any operation that iterates over superblocks. Is this exploitable? No, it's just a hang.
Is this a Denial of Service (DoS)? Yes, absolutely. Can an unprivileged user trigger it? No, because `mount` needs `CAP_SYS_ADMIN` in the initial user namespace (since f2fs doesn't have `FS_USERNS_MOUNT`). However, could a malicious filesystem trigger it? Yes! If the filesystem can be crafted to make `filemap_fdatawrite` fail, such as by returning an error from `f2fs_writepages` due to corrupted metadata, then it could trigger the hang, even without fault injection!
So, if the filesystem is corrupt, `f2fs_writepages` could fail, and the pages might stay dirty, leading to this infinite loop. This is a Malicious Filesystem Trigger. And, since it could be triggered by plugging in a USB drive with a corrupted F2FS filesystem, it's a Peripheral Trigger, and so a real vulnerability!
I've got to verify that if `filemap_fdatawrite` fails, the inode remains on the dirty list. Okay, if `f2fs_writepages` fails, I'm thinking it calls `mapping_set_error` and possibly unlocks the pages without clearing the dirty bit. But actually, if `writepages` fails, it *usually* clears the dirty bit and *then* sets the error on the mapping. So, if the dirty bit is cleared, the inode *will* be removed. Now, let's verify in `f2fs_write_data_pages`.
|
| 43/2 |
2026/05/26 06:36 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/data.c IncludeLines:true Name:f2fs_write_data_pages]
Results: map[SourceCode:3312: static int f2fs_write_data_pages(struct address_space *mapping,
3313: struct writeback_control *wbc)
3314: {
3315: struct inode *inode = mapping->host;
3316:
3317: return __f2fs_write_data_pages(mapping, wbc,
3318: F2FS_I(inode)->cp_task == current ?
3319: FS_CP_DATA_IO : FS_DATA_IO);
3320: }
SourceFile:fs/f2fs/data.c]
|
| 44/2 |
2026/05/26 06:36 |
llm |
expert |
0m
Model:
Tokens:
input: 79839
output: 48
thoughts: 16
Thoughts: Okay, let's dive into this. Here's my thought process on this:
**Analyzing `__f2fs_write_data_pages`**
Alright, the next step is to examine `__f2fs_write_data_pages`. I need to understand exactly what's happening at this critical level. This function is, fundamentally, where the rubber meets the road for data persistence within the F2FS filesystem. I want to scrutinize the data flow, the error handling, and the overall efficiency of its operations. I'm looking for potential bottlenecks, areas for optimization, or any subtle issues that might impact performance or reliability. Given my understanding of block device interactions and filesystem internals, I need to evaluate how effectively the function handles write requests, considers data alignment, and manages the potential for write amplification. I'll need to pay close attention to the allocation strategies used for the data blocks and the interaction with the F2FS's journaling or logging mechanisms if it applies. Furthermore, I need to check how it handles any I/O errors and, most importantly, if it is working efficiently. This is the core of data writing, so thorough analysis is key.
|
| 45/2 |
2026/05/26 06:36 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/data.c IncludeLines:true Name:__f2fs_write_data_pages]
Results: map[SourceCode:3244: static int __f2fs_write_data_pages(struct address_space *mapping,
3245: struct writeback_control *wbc,
3246: enum iostat_type io_type)
3247: {
3248: struct inode *inode = mapping->host;
3249: struct f2fs_sb_info *sbi = F2FS_I_SB(inode);
3250: struct blk_plug plug;
3251: int ret;
3252: bool locked = false;
3253:
3254: /* skip writing if there is no dirty page in this inode */
3255: if (!get_dirty_pages(inode) && wbc->sync_mode == WB_SYNC_NONE)
3256: return 0;
3257:
3258: /* during POR, we don't need to trigger writepage at all. */
3259: if (unlikely(is_sbi_flag_set(sbi, SBI_POR_DOING)))
3260: goto skip_write;
3261:
3262: if ((S_ISDIR(inode->i_mode) || IS_NOQUOTA(inode)) &&
3263: wbc->sync_mode == WB_SYNC_NONE &&
3264: get_dirty_pages(inode) < nr_pages_to_skip(sbi, DATA) &&
3265: f2fs_available_free_memory(sbi, DIRTY_DENTS))
3266: goto skip_write;
3267:
3268: /* skip writing in file defragment preparing stage */
3269: if (is_inode_flag_set(inode, FI_SKIP_WRITES))
3270: goto skip_write;
3271:
3272: trace_f2fs_writepages(mapping->host, wbc, DATA);
3273:
3274: /* to avoid spliting IOs due to mixed WB_SYNC_ALL and WB_SYNC_NONE */
3275: if (wbc->sync_mode == WB_SYNC_ALL)
3276: atomic_inc(&sbi->wb_sync_req[DATA]);
3277: else if (atomic_read(&sbi->wb_sync_req[DATA])) {
3278: /* to avoid potential deadlock */
3279: if (current->plug)
3280: blk_finish_plug(current->plug);
3281: goto skip_write;
3282: }
3283:
3284: if (__should_serialize_io(inode, wbc)) {
3285: mutex_lock(&sbi->writepages);
3286: locked = true;
3287: }
3288:
3289: blk_start_plug(&plug);
3290: ret = f2fs_write_cache_pages(mapping, wbc, io_type);
3291: blk_finish_plug(&plug);
3292:
3293: if (locked)
3294: mutex_unlock(&sbi->writepages);
3295:
3296: if (wbc->sync_mode == WB_SYNC_ALL)
3297: atomic_dec(&sbi->wb_sync_req[DATA]);
3298: /*
3299: * if some pages were truncated, we cannot guarantee its mapping->host
3300: * to detect pending bios.
3301: */
3302:
3303: f2fs_remove_dirty_inode(inode);
3304: return ret;
3305:
3306: skip_write:
3307: wbc->pages_skipped += get_dirty_pages(inode);
3308: trace_f2fs_writepages(mapping->host, wbc, DATA);
3309: return 0;
3310: }
SourceFile:fs/f2fs/data.c]
|
| 46/2 |
2026/05/26 06:36 |
llm |
expert |
0m
Model:
Tokens:
input: 81245
output: 47
thoughts: 1321
Thoughts: **F2FS Recovery and Dirty Inodes: A Deep Dive**
Okay, let's break this down. My initial thought was: `__f2fs_write_data_pages` has a check for `SBI_POR_DOING`. If it's set, the write is skipped. That makes sense, but what happens to the dirty inode list? Ah, `f2fs_remove_dirty_inode(inode)` isn't called in the skip path. That's...interesting.
Now, when is `SBI_POR_DOING` set? In `f2fs_fill_super`. It's set at the beginning to disallow writes during the Power-On Recovery (POR) phase. And it's cleared *after* `f2fs_recover_fsync_data`. So `f2fs_recover_fsync_data` should handle the cleaning up, right?
But wait. Looking at the code for `f2fs_recover_fsync_data` again, `clear_sbi_flag(sbi, SBI_POR_DOING)` happens *before* `f2fs_write_checkpoint`. So, `SBI_POR_DOING` is definitely *not* set during `f2fs_write_checkpoint`. Okay, that's good.
However, the real gotcha is what happens on error. If `f2fs_recover_fsync_data` *fails*, `SBI_POR_DOING` *isn't* cleared and `f2fs_write_checkpoint` *isn't* called. The whole thing bails out with an error! Then, it jumps to `free_meta`. And in `free_meta`, there's no call to `sync_filesystem`. It does `truncate_inode_pages_final` and `evict_inodes`, but no explicit syncing.
But if we go to `sync_free_meta`, it *does* call `sync_filesystem`. When does that happen? It depends on `f2fs_init_inmem_curseg`, `DISABLE_CHECKPOINT` and GC. If any of those fail, or `DISABLE_CHECKPOINT` is set, or if we have a background GC thread, we jump to `sync_free_meta`. So, *before* we get there `SBI_POR_DOING` is already cleared, which is good.
So, if `f2fs_recover_fsync_data` succeeds, all is well. But if the recovery *fails* and we jump to `free_meta`, what then? It looks like we've got dirty inodes potentially hanging around, and the writes are not being attempted in the error case. However, in the case of a failed recovery, since `SBI_POR_DOING` will not be cleared, it looks like writes will still be skipped as expected.
Ok, now, if `f2fs_write_checkpoint` *is* called (meaning recovery succeeded), `SBI_POR_DOING` is cleared. So subsequent calls to `__f2fs_write_data_pages` won't skip writes due to that flag. However, there's always the possibility of failure in the subsequent `f2fs_write_cache_pages` call due to a fault, but that's a separate issue.
|
| 47/2 |
2026/05/26 06:36 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/data.c IncludeLines:true Name:f2fs_write_cache_pages]
Results: map[SourceCode:2934: /*
2935: * This function was copied from write_cache_pages from mm/page-writeback.c.
2936: * The major change is making write step of cold data page separately from
2937: * warm/hot data page.
2938: */
2939: static int f2fs_write_cache_pages(struct address_space *mapping,
2940: struct writeback_control *wbc,
2941: enum iostat_type io_type)
2942: {
2943: int ret = 0;
2944: int done = 0, retry = 0;
2945: struct page *pages_local[F2FS_ONSTACK_PAGES];
2946: struct page **pages = pages_local;
2947: struct folio_batch fbatch;
2948: struct f2fs_sb_info *sbi = F2FS_M_SB(mapping);
2949: struct bio *bio = NULL;
2950: sector_t last_block;
2951: #ifdef CONFIG_F2FS_FS_COMPRESSION
2952: struct inode *inode = mapping->host;
2953: struct compress_ctx cc = {
2954: .inode = inode,
2955: .log_cluster_size = F2FS_I(inode)->i_log_cluster_size,
2956: .cluster_size = F2FS_I(inode)->i_cluster_size,
2957: .cluster_idx = NULL_CLUSTER,
2958: .rpages = NULL,
2959: .nr_rpages = 0,
2960: .cpages = NULL,
2961: .valid_nr_cpages = 0,
2962: .rbuf = NULL,
2963: .cbuf = NULL,
2964: .rlen = PAGE_SIZE * F2FS_I(inode)->i_cluster_size,
2965: .private = NULL,
2966: };
2967: #endif
2968: int nr_folios, p, idx;
2969: int nr_pages;
2970: unsigned int max_pages = F2FS_ONSTACK_PAGES;
2971: pgoff_t index;
2972: pgoff_t end; /* Inclusive */
2973: pgoff_t done_index;
2974: int range_whole = 0;
2975: xa_mark_t tag;
2976: int nwritten = 0;
2977: int submitted = 0;
2978: int i;
2979:
2980: #ifdef CONFIG_F2FS_FS_COMPRESSION
2981: if (f2fs_compressed_file(inode) &&
2982: 1 << cc.log_cluster_size > F2FS_ONSTACK_PAGES) {
2983: pages = f2fs_kzalloc(sbi, sizeof(struct page *) <<
2984: cc.log_cluster_size, GFP_NOFS | __GFP_NOFAIL);
2985: max_pages = 1 << cc.log_cluster_size;
2986: }
2987: #endif
2988:
2989: folio_batch_init(&fbatch);
2990:
2991: if (get_dirty_pages(mapping->host) <=
2992: SM_I(F2FS_M_SB(mapping))->min_hot_blocks)
2993: set_inode_flag(mapping->host, FI_HOT_DATA);
2994: else
2995: clear_inode_flag(mapping->host, FI_HOT_DATA);
2996:
2997: if (wbc->range_cyclic) {
2998: index = mapping->writeback_index; /* prev offset */
2999: end = -1;
3000: } else {
3001: index = wbc->range_start >> PAGE_SHIFT;
3002: end = wbc->range_end >> PAGE_SHIFT;
3003: if (wbc->range_start == 0 && wbc->range_end == LLONG_MAX)
3004: range_whole = 1;
3005: }
3006: if (wbc->sync_mode == WB_SYNC_ALL || wbc->tagged_writepages)
3007: tag = PAGECACHE_TAG_TOWRITE;
3008: else
3009: tag = PAGECACHE_TAG_DIRTY;
3010: retry:
3011: retry = 0;
3012: if (wbc->sync_mode == WB_SYNC_ALL || wbc->tagged_writepages)
3013: tag_pages_for_writeback(mapping, index, end);
3014: done_index = index;
3015: while (!done && !retry && (index <= end)) {
3016: nr_pages = 0;
3017: again:
3018: nr_folios = filemap_get_folios_tag(mapping, &index, end,
3019: tag, &fbatch);
3020: if (nr_folios == 0) {
3021: if (nr_pages)
3022: goto write;
3023: break;
3024: }
3025:
3026: for (i = 0; i < nr_folios; i++) {
3027: struct folio *folio = fbatch.folios[i];
3028:
3029: idx = 0;
3030: p = folio_nr_pages(folio);
3031: add_more:
3032: pages[nr_pages] = folio_page(folio, idx);
3033: folio_get(folio);
3034: if (++nr_pages == max_pages) {
3035: index = folio->index + idx + 1;
3036: folio_batch_release(&fbatch);
3037: goto write;
3038: }
3039: if (++idx < p)
3040: goto add_more;
3041: }
3042: folio_batch_release(&fbatch);
3043: goto again;
3044: write:
3045: for (i = 0; i < nr_pages; i++) {
3046: struct page *page = pages[i];
3047: struct folio *folio = page_folio(page);
3048: bool need_readd;
3049: readd:
3050: need_readd = false;
3051: #ifdef CONFIG_F2FS_FS_COMPRESSION
3052: if (f2fs_compressed_file(inode)) {
3053: void *fsdata = NULL;
3054: struct page *pagep;
3055: int ret2;
3056:
3057: ret = f2fs_init_compress_ctx(&cc);
3058: if (ret) {
3059: done = 1;
3060: break;
3061: }
3062:
3063: if (!f2fs_cluster_can_merge_page(&cc,
3064: folio->index)) {
3065: ret = f2fs_write_multi_pages(&cc,
3066: &submitted, wbc, io_type);
3067: if (!ret)
3068: need_readd = true;
3069: goto result;
3070: }
3071:
3072: if (unlikely(f2fs_cp_error(sbi)))
3073: goto lock_folio;
3074:
3075: if (!f2fs_cluster_is_empty(&cc))
3076: goto lock_folio;
3077:
3078: if (f2fs_all_cluster_page_ready(&cc,
3079: pages, i, nr_pages, true))
3080: goto lock_folio;
3081:
3082: ret2 = f2fs_prepare_compress_overwrite(
3083: inode, &pagep,
3084: folio->index, &fsdata);
3085: if (ret2 < 0) {
3086: ret = ret2;
3087: done = 1;
3088: break;
3089: } else if (ret2 &&
3090: (!f2fs_compress_write_end(inode,
3091: fsdata, folio->index, 1) ||
3092: !f2fs_all_cluster_page_ready(&cc,
3093: pages, i, nr_pages,
3094: false))) {
3095: retry = 1;
3096: break;
3097: }
3098: }
3099: #endif
3100: /* give a priority to WB_SYNC threads */
3101: if (atomic_read(&sbi->wb_sync_req[DATA]) &&
3102: wbc->sync_mode == WB_SYNC_NONE) {
3103: done = 1;
3104: break;
3105: }
3106: #ifdef CONFIG_F2FS_FS_COMPRESSION
3107: lock_folio:
3108: #endif
3109: done_index = folio->index;
3110: retry_write:
3111: folio_lock(folio);
3112:
3113: if (unlikely(folio->mapping != mapping)) {
3114: continue_unlock:
3115: folio_unlock(folio);
3116: continue;
3117: }
3118:
3119: if (!folio_test_dirty(folio)) {
3120: /* someone wrote it for us */
3121: goto continue_unlock;
3122: }
3123:
3124: if (folio_test_writeback(folio)) {
3125: if (wbc->sync_mode == WB_SYNC_NONE)
3126: goto continue_unlock;
3127: f2fs_folio_wait_writeback(folio, DATA, true, true);
3128: }
3129:
3130: if (!folio_clear_dirty_for_io(folio))
3131: goto continue_unlock;
3132:
3133: #ifdef CONFIG_F2FS_FS_COMPRESSION
3134: if (f2fs_compressed_file(inode)) {
3135: folio_get(folio);
3136: f2fs_compress_ctx_add_page(&cc, folio);
3137: continue;
3138: }
3139: #endif
3140: submitted = 0;
3141: ret = f2fs_write_single_data_page(folio,
3142: &submitted, &bio, &last_block,
3143: wbc, io_type, 0, true);
3144: #ifdef CONFIG_F2FS_FS_COMPRESSION
3145: result:
3146: #endif
3147: nwritten += submitted;
3148: wbc->nr_to_write -= submitted;
3149:
3150: if (unlikely(ret)) {
3151: /*
3152: * keep nr_to_write, since vfs uses this to
3153: * get # of written pages.
3154: */
3155: if (ret == 1) {
3156: ret = 0;
3157: goto next;
3158: } else if (ret == -EAGAIN) {
3159: ret = 0;
3160: if (wbc->sync_mode == WB_SYNC_ALL) {
3161: f2fs_io_schedule_timeout(
3162: DEFAULT_IO_TIMEOUT);
3163: goto retry_write;
3164: }
3165: goto next;
3166: }
3167: done_index = folio_next_index(folio);
3168: done = 1;
3169: break;
3170: }
3171:
3172: if (wbc->nr_to_write <= 0 &&
3173: wbc->sync_mode == WB_SYNC_NONE) {
3174: done = 1;
3175: break;
3176: }
3177: next:
3178: if (need_readd)
3179: goto readd;
3180: }
3181: release_pages(pages, nr_pages);
3182: cond_resched();
3183: }
3184: #ifdef CONFIG_F2FS_FS_COMPRESSION
3185: /* flush remained pages in compress cluster */
3186: if (f2fs_compressed_file(inode) && !f2fs_cluster_is_empty(&cc)) {
3187: ret = f2fs_write_multi_pages(&cc, &submitted, wbc, io_type);
3188: nwritten += submitted;
3189: wbc->nr_to_write -= submitted;
3190: if (ret) {
3191: done = 1;
3192: retry = 0;
3193: }
3194: }
3195: if (f2fs_compressed_file(inode))
3196: f2fs_destroy_compress_ctx(&cc, false);
3197: #endif
3198: if (retry) {
3199: index = 0;
3200: end = -1;
3201: goto retry;
3202: }
3203: if (wbc->range_cyclic && !done)
3204: done_index = 0;
3205: if (wbc->range_cyclic || (range_whole && wbc->nr_to_write > 0))
3206: mapping->writeback_index = done_index;
3207:
3208: if (nwritten)
3209: f2fs_submit_merged_write_cond(F2FS_M_SB(mapping), mapping->host,
3210: NULL, 0, DATA);
3211: /* submit cached bio of IPU write */
3212: if (bio)
3213: f2fs_submit_merged_ipu_write(sbi, &bio, NULL);
3214:
3215: #ifdef CONFIG_F2FS_FS_COMPRESSION
3216: if (pages != pages_local)
3217: kfree(pages);
3218: #endif
3219:
3220: return ret;
3221: }
SourceFile:fs/f2fs/data.c]
|
| 48/2 |
2026/05/26 06:36 |
llm |
expert |
0m
Model:
Tokens:
input: 87558
output: 47
thoughts: 313
Thoughts: **F2FS Write Cache Analysis**
Okay, so I'm tracing the data flow here in the F2FS file system, specifically around the write caching mechanism. It looks like `f2fs_write_cache_pages` calls `f2fs_write_single_data_page`, which makes sense. I can see that if `f2fs_write_single_data_page` throws an error, it sets `done = 1` and exits. But, here's the kicker: `f2fs_write_single_data_page` is invoked *after* `folio_clear_dirty_for_io(folio)` is called. That means the page is no longer flagged as dirty!
If the page is no longer dirty, then `get_dirty_pages(inode)` should decrement its count. That's how it's designed to work, right? So, if all those pages are cleared, `get_dirty_pages(inode)` would eventually become zero, assuming no other writes. Okay, so would this affect the behavior of `f2fs_sync_dirty_inodes`?
No, wait, `f2fs_sync_dirty_inodes` looks at `list_empty(head)` to determine if it has anything to do. The dirty inodes are managed using a list, not by just checking that dirty page counter. Hold on, wasn't there something about how the inode is removed from a dirty list? Ah, right! It is removed from the dirty list by `f2fs_remove_dirty_inode(inode)`, which is called at the end of `__f2fs_write_data_pages`:
```c
f2fs_remove_dirty_inode(inode);
```
So `f2fs_remove_dirty_inode` handles the list removal. Okay, I need to check out what's happening within `f2fs_remove_dirty_inode` to completely understand the sequence and the implications for the dirty inode list.
|
| 49/2 |
2026/05/26 06:36 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/inode.c IncludeLines:true Name:f2fs_remove_dirty_inode]
Results: map[SourceCode:1051: void f2fs_remove_dirty_inode(struct inode *inode)
1052: {
1053: struct f2fs_sb_info *sbi = F2FS_I_SB(inode);
1054: enum inode_type type = S_ISDIR(inode->i_mode) ? DIR_INODE : FILE_INODE;
1055:
1056: if (!S_ISDIR(inode->i_mode) && !S_ISREG(inode->i_mode) &&
1057: !S_ISLNK(inode->i_mode))
1058: return;
1059:
1060: if (type == FILE_INODE && !test_opt(sbi, DATA_FLUSH))
1061: return;
1062:
1063: spin_lock(&sbi->inode_lock[type]);
1064: __remove_dirty_inode(inode, type);
1065: spin_unlock(&sbi->inode_lock[type]);
1066: }
SourceFile:fs/f2fs/checkpoint.c]
|
| 50/2 |
2026/05/26 06:36 |
llm |
expert |
0m
Model:
Tokens:
input: 88662
output: 44
thoughts: 45
Thoughts: Okay, let's break this down.
**Understanding Inode Dirty List Management**
Right, so I'm focusing on `__remove_dirty_inode`. The key is, I see this function is responsible for removing an inode from the dirty list. But there's a crucial conditional element here: this removal *only* happens if `get_dirty_pages(inode)` returns zero. In other words, the inode is removed from the dirty list only if it has no dirty pages associated with it. That's the core of the check I need to confirm. I have to verify that behavior. I need to dive into the definition of `__remove_dirty_inode` and see how it is working.
|
| 51/2 |
2026/05/26 06:36 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/inode.c IncludeLines:true Name:__remove_dirty_inode]
Results: map[SourceCode:1021: static void __remove_dirty_inode(struct inode *inode, enum inode_type type)
1022: {
1023: int flag = (type == DIR_INODE) ? FI_DIRTY_DIR : FI_DIRTY_FILE;
1024:
1025: if (get_dirty_pages(inode) || !is_inode_flag_set(inode, flag))
1026: return;
1027:
1028: list_del_init(&F2FS_I(inode)->dirty_list);
1029: clear_inode_flag(inode, flag);
1030: stat_dec_dirty_inode(F2FS_I_SB(inode), type);
1031: }
SourceFile:fs/f2fs/checkpoint.c]
|
| 52/2 |
2026/05/26 06:36 |
llm |
expert |
0m
Model:
Tokens:
input: 89132
output: 49
thoughts: 680
Thoughts: **The Infinite Loop of Dirty Inodes**
Okay, here's what's going through my head: I've traced this issue, and it looks like a nasty infinite loop in the F2FS filesystem code, and I'm pretty sure it's a denial of service vulnerability, and I am fairly certain I have found the root cause.
It all hinges on how `__remove_dirty_inode` interacts with `get_dirty_pages(inode)`. If `get_dirty_pages(inode)` isn't zero, the inode *isn't* removed from the dirty list. And that's where the problem arises. If `f2fs_write_cache_pages` fails *and* leaves pages dirty (say, because of an early abort due to an error), `get_dirty_pages(inode)` is going to be greater than zero. Then, `f2fs_remove_dirty_inode` *won't* remove that inode.
Then, `f2fs_sync_dirty_inodes` comes along...and loops. It loops over that dirty list. And because that inode is still dirty, `f2fs_sync_dirty_inodes` grabs it, and calls `filemap_fdatawrite`, which calls `__f2fs_write_data_pages`, which calls `f2fs_write_cache_pages`... which, if it encounters the same failure again, will fail, leave pages dirty, and leave the inode in the dirty list.
See? An INFINITE LOOP! And this happens inside `block_operations`, which is called by `f2fs_write_checkpoint`, which is called by `f2fs_recover_fsync_data`, which is called by `f2fs_fill_super`. And `f2fs_fill_super` holds the `s_umount` lock for writing. So this means `s_umount` is held for write forever, blocking `quotactl` and system-wide `sync` calls. It's a textbook Denial of Service.
Wait, the next step is to figure out *why* `f2fs_write_cache_pages` is failing. Oh! The reproducer's using fault injection (`fault_injection=16`, `fault_type=16777214`). So, if `f2fs_write_single_data_page` fails due to, say, a memory allocation error triggered by the fault injection, it returns an error. Then `f2fs_write_cache_pages` breaks out of its loop, it looks like this:
```c
if (unlikely(ret)) {
if (ret == 1) {
ret = 0;
goto next;
} else if (ret == -EAGAIN) {
...
}
done_index = folio_next_index(folio);
done = 1;
break;
}
```
...and it returns that error. Crucially, it doesn't write the remaining pages in the mapping, so they *stay* dirty. So `get_dirty_pages(inode)` remains greater than zero. And the inode *stays* in the dirty list. So, `f2fs_sync_dirty_inodes` keeps looping forever.
But, is this only a problem with fault injection? What if we encounter a *real* I/O error? I have to look into this more. Let's check `f2fs_write_single_data_page`...
|
| 53/2 |
2026/05/26 06:36 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/data.c IncludeLines:true Name:f2fs_write_single_data_page]
Results: map[SourceCode:2774: int f2fs_write_single_data_page(struct folio *folio, int *submitted,
2775: struct bio **bio,
2776: sector_t *last_block,
2777: struct writeback_control *wbc,
2778: enum iostat_type io_type,
2779: int compr_blocks,
2780: bool allow_balance)
2781: {
2782: struct inode *inode = folio->mapping->host;
2783: struct f2fs_sb_info *sbi = F2FS_I_SB(inode);
2784: loff_t i_size = i_size_read(inode);
2785: const pgoff_t end_index = ((unsigned long long)i_size)
2786: >> PAGE_SHIFT;
2787: loff_t psize = (loff_t)(folio->index + 1) << PAGE_SHIFT;
2788: unsigned offset = 0;
2789: bool need_balance_fs = false;
2790: bool quota_inode = IS_NOQUOTA(inode);
2791: int err = 0;
2792: struct f2fs_io_info fio = {
2793: .sbi = sbi,
2794: .ino = inode->i_ino,
2795: .type = DATA,
2796: .op = REQ_OP_WRITE,
2797: .op_flags = wbc_to_write_flags(wbc),
2798: .old_blkaddr = NULL_ADDR,
2799: .folio = folio,
2800: .encrypted_page = NULL,
2801: .submitted = 0,
2802: .compr_blocks = compr_blocks,
2803: .need_lock = compr_blocks ? LOCK_DONE : LOCK_RETRY,
2804: .meta_gc = f2fs_meta_inode_gc_required(inode) ? 1 : 0,
2805: .io_type = io_type,
2806: .io_wbc = wbc,
2807: .bio = bio,
2808: .last_block = last_block,
2809: };
2810:
2811: trace_f2fs_writepage(folio, DATA);
2812:
2813: /* we should bypass data pages to proceed the kworker jobs */
2814: if (unlikely(f2fs_cp_error(sbi))) {
2815: mapping_set_error(folio->mapping, -EIO);
2816: /*
2817: * don't drop any dirty dentry pages for keeping lastest
2818: * directory structure.
2819: */
2820: if (S_ISDIR(inode->i_mode) &&
2821: !is_sbi_flag_set(sbi, SBI_IS_CLOSE))
2822: goto redirty_out;
2823:
2824: /* keep data pages in remount-ro mode */
2825: if (F2FS_OPTION(sbi).errors == MOUNT_ERRORS_READONLY)
2826: goto redirty_out;
2827: goto out;
2828: }
2829:
2830: if (unlikely(is_sbi_flag_set(sbi, SBI_POR_DOING)))
2831: goto redirty_out;
2832:
2833: if (folio->index < end_index ||
2834: f2fs_verity_in_progress(inode) ||
2835: compr_blocks)
2836: goto write;
2837:
2838: /*
2839: * If the offset is out-of-range of file size,
2840: * this page does not have to be written to disk.
2841: */
2842: offset = i_size & (PAGE_SIZE - 1);
2843: if ((folio->index >= end_index + 1) || !offset)
2844: goto out;
2845:
2846: folio_zero_segment(folio, offset, folio_size(folio));
2847: write:
2848: /* Dentry/quota blocks are controlled by checkpoint */
2849: if (S_ISDIR(inode->i_mode) || quota_inode) {
2850: /*
2851: * We need to wait for node_write to avoid block allocation during
2852: * checkpoint. This can only happen to quota writes which can cause
2853: * the below discard race condition.
2854: */
2855: if (quota_inode)
2856: f2fs_down_read(&sbi->node_write);
2857:
2858: fio.need_lock = LOCK_DONE;
2859: err = f2fs_do_write_data_page(&fio);
2860:
2861: if (quota_inode)
2862: f2fs_up_read(&sbi->node_write);
2863:
2864: goto done;
2865: }
2866:
2867: need_balance_fs = true;
2868: err = -EAGAIN;
2869: if (f2fs_has_inline_data(inode)) {
2870: err = f2fs_write_inline_data(inode, folio);
2871: if (!err)
2872: goto out;
2873: }
2874:
2875: if (err == -EAGAIN) {
2876: err = f2fs_do_write_data_page(&fio);
2877: if (err == -EAGAIN) {
2878: f2fs_bug_on(sbi, compr_blocks);
2879: fio.need_lock = LOCK_REQ;
2880: err = f2fs_do_write_data_page(&fio);
2881: }
2882: }
2883:
2884: if (err) {
2885: file_set_keep_isize(inode);
2886: } else {
2887: spin_lock(&F2FS_I(inode)->i_size_lock);
2888: if (F2FS_I(inode)->last_disk_size < psize)
2889: F2FS_I(inode)->last_disk_size = psize;
2890: spin_unlock(&F2FS_I(inode)->i_size_lock);
2891: }
2892:
2893: done:
2894: if (err && err != -ENOENT)
2895: goto redirty_out;
2896:
2897: out:
2898: inode_dec_dirty_pages(inode);
2899: if (err) {
2900: folio_clear_uptodate(folio);
2901: folio_clear_f2fs_gcing(folio);
2902: }
2903: folio_unlock(folio);
2904: if (!S_ISDIR(inode->i_mode) && !IS_NOQUOTA(inode) &&
2905: !F2FS_I(inode)->wb_task && allow_balance)
2906: f2fs_balance_fs(sbi, need_balance_fs);
2907:
2908: if (unlikely(f2fs_cp_error(sbi))) {
2909: f2fs_submit_merged_write(sbi, DATA);
2910: if (bio && *bio)
2911: f2fs_submit_merged_ipu_write(sbi, bio, NULL);
2912: submitted = NULL;
2913: }
2914:
2915: if (submitted)
2916: *submitted = fio.submitted;
2917:
2918: return 0;
2919:
2920: redirty_out:
2921: folio_redirty_for_writepage(wbc, folio);
2922: /*
2923: * pageout() in MM translates EAGAIN, so calls handle_write_error()
2924: * -> mapping_set_error() -> set_bit(AS_EIO, ...).
2925: * file_write_and_wait_range() will see EIO error, which is critical
2926: * to return value of fsync() followed by atomic_write failure to user.
2927: */
2928: folio_unlock(folio);
2929: if (!err)
2930: return 1;
2931: return err;
2932: }
SourceFile:fs/f2fs/data.c]
|
| 54/2 |
2026/05/26 06:36 |
llm |
expert |
0m
Model:
Tokens:
input: 93133
output: 19
thoughts: 1620
Thoughts: **F2FS Writeback Infinite Loop Analysis**
Alright, let's break this down. It seems we have a nasty infinite loop condition within F2FS's writeback mechanism. I'm focusing on `f2fs_write_single_data_page`, and it's interactions with page redirtying.
Here's the scenario: if `f2fs_do_write_data_page` returns an error, such as `-ENOMEM` (memory allocation failure due to fault injection) or `-EIO` (I/O error), the page gets redirtied using `folio_redirty_for_writepage`. The key here is that when redirtied, the dirty page count for the inode isn't decremented right away ( `inode_dec_dirty_pages` gets skipped in the `redirty_out` path).
Therefore `get_dirty_pages(inode)` remains > 0.
So, `f2fs_write_single_data_page` returns the error, which propagates back up through `f2fs_write_cache_pages`, and eventually to `__f2fs_write_data_pages`. `__f2fs_write_data_pages` then calls `f2fs_remove_dirty_inode(inode)`. Because the inode still has `get_dirty_pages(inode) > 0`, `f2fs_remove_dirty_inode` *won't* remove the inode from the dirty list.
Then, `f2fs_sync_dirty_inodes` comes along, sees the inode still dirty, and calls `filemap_fdatawrite` again which in turn calls `__f2fs_write_data_pages`, and `f2fs_write_cache_pages`, and then `f2fs_write_single_data_page` all again, and the *same error happens again*.
This error can happen because of either a persistent I/O error (e.g. bad disk sector) or because of a persistent filesystem corruption that causes `f2fs_do_write_data_page` to fail (e.g. invalid block address or running out of space).
The critical thing is that this loop *continues* because the inode is never removed from the dirty list.
Now, if `f2fs_cp_error(sbi)` *is* set, then `f2fs_write_single_data_page` checks this. If `f2fs_cp_error(sbi)` is set it can either set the file to readonly or call the `out` path, in which `inode_dec_dirty_pages` *is* called. This would terminate the loop. However, by default, f2fs uses `MOUNT_ERRORS_CONTINUE`, so it should go to `out`.
However, the reproducer uses fault injection (e.g. `FAULT_KMALLOC`). And here's the kicker: fault injection *doesn't* set `f2fs_cp_error(sbi)`! Fault injection just returns `-ENOMEM`. This means the code goes to `redirty_out` and the loop is *never* terminated.
The reproducer sets `fault_type=0xfffffe`. This includes various fault injection types, including `FAULT_WRITE_IO`. If `FAULT_WRITE_IO` is set, `f2fs_submit_page_write` would fail with `-EIO`. However, because `f2fs_cp_error(sbi)` isn't set, the loop continues.
So, the infinite loop is *specifically* triggered by fault injection (or a persistent, but *undetected* error) because, while it generates errors, it doesn't trigger the error handling path that *would* clear the dirty page count and prevent the loop.
|
| 55/2 |
2026/05/26 06:37 |
tool |
grepper |
0m
Args: map[Expression:FAULT_WRITE_IO]
Results: map[Output:Documentation/ABI/testing/sysfs-fs-f2fs=709=Description: Support configuring fault injection type, should be
--
Documentation/ABI/testing/sysfs-fs-f2fs-729- FAULT_DISCARD 0x00002000
Documentation/ABI/testing/sysfs-fs-f2fs:730: FAULT_WRITE_IO 0x00004000
Documentation/ABI/testing/sysfs-fs-f2fs-731- FAULT_SLAB_ALLOC 0x00008000
--
Documentation/filesystems/f2fs.rst=181=fault_type=%d Support configuring fault injection type, should be
--
Documentation/filesystems/f2fs.rst-201- FAULT_DISCARD 0x00002000
Documentation/filesystems/f2fs.rst:202: FAULT_WRITE_IO 0x00004000
Documentation/filesystems/f2fs.rst-203- FAULT_SLAB_ALLOC 0x00008000
--
fs/f2fs/data.c=317=static void f2fs_write_end_io(struct bio *bio)
--
fs/f2fs/data.c-324-
fs/f2fs/data.c:325: if (time_to_inject(sbi, FAULT_WRITE_IO))
fs/f2fs/data.c-326- bio->bi_status = BLK_STS_IOERR;
--
fs/f2fs/f2fs.h=43=enum {
--
fs/f2fs/f2fs.h-57- FAULT_DISCARD,
fs/f2fs/f2fs.h:58: FAULT_WRITE_IO,
fs/f2fs/f2fs.h-59- FAULT_SLAB_ALLOC,
--
fs/f2fs/super.c=47=const char *f2fs_fault_name[FAULT_MAX] = {
--
fs/f2fs/super.c-61- [FAULT_DISCARD] = "discard error",
fs/f2fs/super.c:62: [FAULT_WRITE_IO] = "write IO error",
fs/f2fs/super.c-63- [FAULT_SLAB_ALLOC] = "slab alloc",
]
|
| 56/2 |
2026/05/26 06:37 |
llm |
expert |
0m
Model:
Tokens:
input: 96217
output: 18
thoughts: 215
Thoughts: **Analyzing Potential Infinite Loops with Fault Injection in F2FS**
Okay, let's break this down. My initial concern was a potential infinite loop scenario arising from injected faults, specifically `FAULT_WRITE_IO`, within the F2FS file system. My train of thought started with the observation that if this fault is injected, `bio->bi_status` gets set to `BLK_STS_IOERR` in `f2fs_write_end_io`. That makes sense. However, I noticed a discrepancy. While `f2fs_write_end_io` does flag the page with an error, it *doesn't* directly call `f2fs_cp_error(sbi)`. That felt odd.
Then, I remembered, `f2fs_write_end_io` *does* call `mapping_set_error` upon encountering an error. But this also raised a question; the page was already submitted, which indicates that `f2fs_write_single_data_page` returned a zero value, right? So the loop in `f2fs_write_cache_pages` proceeds as if nothing happened, and completes! That's concerning. That, in turn, means that `f2fs_sync_dirty_inodes` then sees that these pages were submitted and the inode is then removed from the dirty list.
Therefore, after going through these steps, `FAULT_WRITE_IO` *doesn't* appear to trigger an infinite loop here. That's a relief.
Now, a new question arises, what about the potential impact of `FAULT_CHECKPOINT`? I need to analyze this further.
|
| 57/2 |
2026/05/26 06:37 |
tool |
grepper |
0m
Args: map[Expression:FAULT_CHECKPOINT]
Results: map[Output:Documentation/ABI/testing/sysfs-fs-f2fs=709=Description: Support configuring fault injection type, should be
--
Documentation/ABI/testing/sysfs-fs-f2fs-727- FAULT_READ_IO 0x00000800
Documentation/ABI/testing/sysfs-fs-f2fs:728: FAULT_CHECKPOINT 0x00001000
Documentation/ABI/testing/sysfs-fs-f2fs-729- FAULT_DISCARD 0x00002000
--
Documentation/filesystems/f2fs.rst=181=fault_type=%d Support configuring fault injection type, should be
--
Documentation/filesystems/f2fs.rst-199- FAULT_READ_IO 0x00000800
Documentation/filesystems/f2fs.rst:200: FAULT_CHECKPOINT 0x00001000
Documentation/filesystems/f2fs.rst-201- FAULT_DISCARD 0x00002000
--
fs/f2fs/checkpoint.c-23-
fs/f2fs/checkpoint.c:24:#define DEFAULT_CHECKPOINT_IOPRIO (IOPRIO_PRIO_VALUE(IOPRIO_CLASS_RT, 3))
fs/f2fs/checkpoint.c-25-
--
fs/f2fs/checkpoint.c=1927=void f2fs_init_ckpt_req_control(struct f2fs_sb_info *sbi)
--
fs/f2fs/checkpoint.c-1933- atomic_set(&cprc->queued_ckpt, 0);
fs/f2fs/checkpoint.c:1934: cprc->ckpt_thread_ioprio = DEFAULT_CHECKPOINT_IOPRIO;
fs/f2fs/checkpoint.c-1935- init_waitqueue_head(&cprc->ckpt_wait_queue);
--
fs/f2fs/f2fs.h=43=enum {
--
fs/f2fs/f2fs.h-55- FAULT_READ_IO,
fs/f2fs/f2fs.h:56: FAULT_CHECKPOINT,
fs/f2fs/f2fs.h-57- FAULT_DISCARD,
--
fs/f2fs/gc.c=31=static int gc_thread_func(void *data)
--
fs/f2fs/gc.c-74-
fs/f2fs/gc.c:75: if (time_to_inject(sbi, FAULT_CHECKPOINT))
fs/f2fs/gc.c-76- f2fs_stop_checkpoint(sbi, false,
--
fs/f2fs/segment.c=424=void f2fs_balance_fs(struct f2fs_sb_info *sbi, bool need)
--
fs/f2fs/segment.c-428-
fs/f2fs/segment.c:429: if (time_to_inject(sbi, FAULT_CHECKPOINT))
fs/f2fs/segment.c-430- f2fs_stop_checkpoint(sbi, false, STOP_CP_REASON_FAULT_INJECT);
--
fs/f2fs/super.c=47=const char *f2fs_fault_name[FAULT_MAX] = {
--
fs/f2fs/super.c-59- [FAULT_READ_IO] = "read IO error",
fs/f2fs/super.c:60: [FAULT_CHECKPOINT] = "checkpoint error",
fs/f2fs/super.c-61- [FAULT_DISCARD] = "discard error",
]
|
| 58/2 |
2026/05/26 06:37 |
llm |
expert |
0m
Model:
Tokens:
input: 97755
output: 17
thoughts: 190
Thoughts: **Analyzing F2FS Error Handling**
Okay, so I'm tracing through this F2FS code, and it looks like a `FAULT_CHECKPOINT` triggers a call to `f2fs_stop_checkpoint`. That function, in turn, sets `f2fs_cp_error(sbi)`, which is interesting. Wait a minute... if `f2fs_cp_error(sbi)` is set, `f2fs_write_single_data_page` takes the `out` path, and critically, it clears the dirty bit. So, `FAULT_CHECKPOINT` actually *doesn't* lead to some endless loop, as I initially suspected. That's good to know.
Now, I'm shifting gears. What about the other fault conditions, specifically `FAULT_KMALLOC` or `FAULT_PAGE_ALLOC`? What happens if those failures occur within `f2fs_do_write_data_page`, specifically when it's calling `f2fs_allocate_data_block`? If that allocation fails due to `FAULT_KMALLOC`, `FAULT_PAGE_ALLOC`, or even `FAULT_BLOCK`, how does the error propagate and what's the ultimate behavior? I need to dig into the error handling paths from there.
|
| 59/2 |
2026/05/26 06:37 |
tool |
grepper |
0m
Args: map[Expression:FAULT_BLOCK]
Results: map[Output:Documentation/ABI/testing/sysfs-fs-f2fs=709=Description: Support configuring fault injection type, should be
--
Documentation/ABI/testing/sysfs-fs-f2fs-722- FAULT_ORPHAN 0x00000040
Documentation/ABI/testing/sysfs-fs-f2fs:723: FAULT_BLOCK 0x00000080
Documentation/ABI/testing/sysfs-fs-f2fs-724- FAULT_DIR_DEPTH 0x00000100
--
Documentation/filesystems/f2fs.rst=181=fault_type=%d Support configuring fault injection type, should be
--
Documentation/filesystems/f2fs.rst-194- FAULT_ORPHAN 0x00000040
Documentation/filesystems/f2fs.rst:195: FAULT_BLOCK 0x00000080
Documentation/filesystems/f2fs.rst-196- FAULT_DIR_DEPTH 0x00000100
--
drivers/net/ethernet/amazon/ena/ena_com.c-45-#define ENA_PHC_DEFAULT_EXPIRE_TIMEOUT_USEC 10
drivers/net/ethernet/amazon/ena/ena_com.c:46:#define ENA_PHC_DEFAULT_BLOCK_TIMEOUT_USEC 1000
drivers/net/ethernet/amazon/ena/ena_com.c-47-#define ENA_PHC_REQ_ID_OFFSET 0xDEAD
--
drivers/net/ethernet/amazon/ena/ena_com.c=1677=int ena_com_phc_config(struct ena_com_dev *ena_dev)
--
drivers/net/ethernet/amazon/ena/ena_com.c-1721- get_feat_resp.u.phc.block_timeout_usec :
drivers/net/ethernet/amazon/ena/ena_com.c:1722: ENA_PHC_DEFAULT_BLOCK_TIMEOUT_USEC;
drivers/net/ethernet/amazon/ena/ena_com.c-1723-
--
drivers/net/wireless/marvell/mwifiex/11ac.c=275=void mwifiex_set_11ac_ba_params(struct mwifiex_private *priv)
drivers/net/wireless/marvell/mwifiex/11ac.c-276-{
drivers/net/wireless/marvell/mwifiex/11ac.c:277: priv->add_ba_param.timeout = MWIFIEX_DEFAULT_BLOCK_ACK_TIMEOUT;
drivers/net/wireless/marvell/mwifiex/11ac.c-278-
--
drivers/net/wireless/marvell/mwifiex/11n.c=786=void mwifiex_set_ba_params(struct mwifiex_private *priv)
drivers/net/wireless/marvell/mwifiex/11n.c-787-{
drivers/net/wireless/marvell/mwifiex/11n.c:788: priv->add_ba_param.timeout = MWIFIEX_DEFAULT_BLOCK_ACK_TIMEOUT;
drivers/net/wireless/marvell/mwifiex/11n.c-789-
--
drivers/net/wireless/marvell/mwifiex/decl.h-74-
drivers/net/wireless/marvell/mwifiex/decl.h:75:#define MWIFIEX_DEFAULT_BLOCK_ACK_TIMEOUT 0xffff
drivers/net/wireless/marvell/mwifiex/decl.h-76-
--
drivers/scsi/st.c=1068=static int check_tape(struct scsi_tape *STp, struct file *filp)
--
drivers/scsi/st.c-1176- DEBC_printk(STp, "No Mode Sense.\n");
drivers/scsi/st.c:1177: STp->block_size = ST_DEFAULT_BLOCK; /* Educated guess (?) */
drivers/scsi/st.c-1178- (STp->buffer)->syscall_result = 0; /* Prevent error propagation */
--
drivers/scsi/st_options.h-36- SENSE. */
drivers/scsi/st_options.h:37:#define ST_DEFAULT_BLOCK 0
drivers/scsi/st_options.h-38-
--
fs/f2fs/f2fs.h=43=enum {
--
fs/f2fs/f2fs.h-50- FAULT_ORPHAN,
fs/f2fs/f2fs.h:51: FAULT_BLOCK,
fs/f2fs/f2fs.h-52- FAULT_DIR_DEPTH,
--
fs/f2fs/f2fs.h=2400=static inline int inc_valid_block_count(struct f2fs_sb_info *sbi,
--
fs/f2fs/f2fs.h-2410-
fs/f2fs/f2fs.h:2411: if (time_to_inject(sbi, FAULT_BLOCK)) {
fs/f2fs/f2fs.h-2412- release = *count;
--
fs/f2fs/f2fs.h=2740=static inline int inc_valid_node_count(struct f2fs_sb_info *sbi,
--
fs/f2fs/f2fs.h-2759-
fs/f2fs/f2fs.h:2760: if (time_to_inject(sbi, FAULT_BLOCK))
fs/f2fs/f2fs.h-2761- goto enospc;
--
fs/f2fs/super.c=47=const char *f2fs_fault_name[FAULT_MAX] = {
--
fs/f2fs/super.c-54- [FAULT_ORPHAN] = "orphan",
fs/f2fs/super.c:55: [FAULT_BLOCK] = "no more block",
fs/f2fs/super.c-56- [FAULT_DIR_DEPTH] = "too big dir depth",
--
fs/ntfs3/lib/lzx_decompress.c-78-/* Assumed block size when the encoded block size begins with a 0 bit. */
fs/ntfs3/lib/lzx_decompress.c:79:#define LZX_DEFAULT_BLOCK_SIZE 32768
fs/ntfs3/lib/lzx_decompress.c-80-
--
fs/ntfs3/lib/lzx_decompress.c=345=static int lzx_read_block_header(struct lzx_decompressor *d,
--
fs/ntfs3/lib/lzx_decompress.c-363- if (bitstream_pop_bits(is, 1)) {
fs/ntfs3/lib/lzx_decompress.c:364: block_size = LZX_DEFAULT_BLOCK_SIZE;
fs/ntfs3/lib/lzx_decompress.c-365- } else {
]
|
| 60/2 |
2026/05/26 06:37 |
llm |
expert |
0m
Model:
Tokens:
input: 100056
output: 31
thoughts: 2803
Thoughts: **F2FS Infinite Loop Vulnerability Analysis**
Okay, so here's what I've figured out. It looks like if an `ENOSPC` error occurs during data page writes in F2FS, specifically during a checkpoint (triggered by a `sync`), we can get into an infinite loop. This happens because:
1. If `FAULT_BLOCK` is injected, `inc_valid_block_count` or `inc_valid_node_count` returns `-ENOSPC`.
2. If `f2fs_do_write_data_page` gets `-ENOSPC`, it propagates that error up to `f2fs_write_single_data_page`.
3. `f2fs_write_single_data_page` redirties the page and returns `-ENOSPC`.
4. Critically, `f2fs_cp_error(sbi)` is *not* set.
5. This means the `f2fs_sync_dirty_inodes` loop just keeps retrying, because `filemap_fdatawrite`'s `-ENOSPC` return is ignored, the inode stays dirty, and the same inode gets picked up again in the next loop iteration. This also means if there's a *real* disk full situation and `inc_valid_block_count` returns `-ENOSPC`, the same loop occurs.
Tracing a full disk scenario, the write goes from `f2fs_do_write_data_page` which returns `-ENOSPC`, to `f2fs_write_single_data_page`, which redirties the page and returns `-ENOSPC`, to `f2fs_write_cache_pages`, which breaks and returns `-ENOSPC` to `__f2fs_write_data_pages` and then back to `f2fs_sync_dirty_inodes`, which loops again.
Looking at `f2fs_sync_dirty_inodes` code, I see `filemap_fdatawrite` is called, but its return value is ignored, and the inode remains in the dirty list. `filemap_fdatawrite` eventually calls `do_writepages`, which *can* set an error on the mapping if it fails. However, the pages themselves remain dirty, so the loop continues.
So yes, if the disk is full, the sync process will loop infinitely, and here's the kicker: it doesn't matter if we're talking about new allocations or overwrites (though it's more likely on new allocations). If a page is dirty, and the disk becomes full before the checkpoint can write it, this infinite loop will happen.
I checked `f2fs_sync_dirty_inodes`, but there's no protection against this infinite loop. It just keeps retrying until the dirty list is empty, which will never happen in the face of persistent `ENOSPC`. This is a big problem. This code is used during checkpoints, including those triggered by `sync()`.
Okay, let's explore the implications. If we can trigger this by filling up a disk and calling `sync()`, we can cause a hang. If a user does this, can they crash the system? Let's check `f2fs_write_single_data_page`, specifically if `-ENOSPC` causes the dirty bit to be dropped: the `redirty_out` path clearly redirties the page! This is the core problem! The page stays dirty.
So, if `f2fs_sync_dirty_inodes` loops infinitely, `sync()` will hang. Because it's called from `block_operations`. But `sync()` calls `f2fs_sync_fs`, which calls `f2fs_write_checkpoint`, which calls `block_operations`, so `sync` will hang. This means the `sync()` thread will hold `s_umount` for read. This prevents a `umount` or `remount` from happening which requires a write lock on `s_umount`.
The bug report says the process syz.2.19 blocked on `super_lock` which takes `s_umount` for write. This means that a `mount` is blocking `quotactl` or another `sync()`.
But hold on... the user process may not cause the hang on `umount` or `quotactl`, as `sync` only holds `s_umount` for read. But `quotactl` blocks on *write* to `s_umount`, whereas the bug report indicates that a thread is blocked on *read* to `super_lock` (which takes `s_umount` for read). This would seem to indicate that *another* thread is holding the write lock to `s_umount`.
|
| 61/2 |
2026/05/26 06:38 |
tool |
grepper |
0m
Error:
Args: map[Expression:f2fs_down_read(&sbi->cp_global_sem)]
Results:
|
| 62/2 |
2026/05/26 06:38 |
llm |
expert |
0m
Model:
Tokens:
input: 103929
output: 19
thoughts: 12
Thoughts: Here's my thought process, summarized as if I were thinking it:
**Digging into the Codebase**
Okay, time to dive in. My task is to understand the context of `cp_global_sem`. This suggests we're dealing with a global semaphore, which is a common synchronization primitive. The name suggests this one is particularly important, perhaps critical to the system. Where is it used? Let's start with a simple search for the exact string `cp_global_sem` within the code. That should give me the initial points of interest β where it's declared, initialized, and used. From there, I can trace its usage patterns. I'll need to look for calls to semaphore acquisition and release functions (e.g., `sem_wait`, `sem_post`, or whatever the specific implementation calls are). Let's see what we can find. I anticipate a reasonable number of hits, so I need to be ready to sift through them efficiently, focusing first on declarations and the highest-level interactions with the semaphore. This should allow me to understand the nature of `cp_global_sem` at a high level.
|
| 63/2 |
2026/05/26 06:38 |
tool |
grepper |
0m
Args: map[Expression:cp_global_sem]
Results: map[Output:fs/f2fs/checkpoint.c=378=static int f2fs_write_meta_pages(struct address_space *mapping,
--
fs/f2fs/checkpoint.c-393- /* if locked failed, cp will flush dirty pages instead */
fs/f2fs/checkpoint.c:394: if (!f2fs_down_write_trylock(&sbi->cp_global_sem))
fs/f2fs/checkpoint.c-395- goto skip_write;
--
fs/f2fs/checkpoint.c-399- written = f2fs_sync_meta_pages(sbi, META, wbc->nr_to_write, FS_META_IO);
fs/f2fs/checkpoint.c:400: f2fs_up_write(&sbi->cp_global_sem);
fs/f2fs/checkpoint.c-401- wbc->nr_to_write = max((long)0, wbc->nr_to_write - written - diff);
--
fs/f2fs/checkpoint.c=1610=int f2fs_write_checkpoint(struct f2fs_sb_info *sbi, struct cp_control *cpc)
--
fs/f2fs/checkpoint.c-1624- if (cpc->reason != CP_RESIZE)
fs/f2fs/checkpoint.c:1625: f2fs_down_write(&sbi->cp_global_sem);
fs/f2fs/checkpoint.c-1626-
--
fs/f2fs/checkpoint.c-1706- if (cpc->reason != CP_RESIZE)
fs/f2fs/checkpoint.c:1707: f2fs_up_write(&sbi->cp_global_sem);
fs/f2fs/checkpoint.c-1708- return err;
--
fs/f2fs/f2fs.h=1597=struct f2fs_sb_info {
--
fs/f2fs/f2fs.h-1631- struct inode *meta_inode; /* cache meta blocks */
fs/f2fs/f2fs.h:1632: struct f2fs_rwsem cp_global_sem; /* checkpoint procedure lock */
fs/f2fs/f2fs.h-1633- struct f2fs_rwsem cp_rwsem; /* blocking FS operations */
--
fs/f2fs/gc.c=2203=int f2fs_resize_fs(struct file *filp, __u64 block_count)
--
fs/f2fs/gc.c-2290- f2fs_down_write(&sbi->gc_lock);
fs/f2fs/gc.c:2291: f2fs_down_write(&sbi->cp_global_sem);
fs/f2fs/gc.c-2292-
--
fs/f2fs/gc.c-2338-out_err:
fs/f2fs/gc.c:2339: f2fs_up_write(&sbi->cp_global_sem);
fs/f2fs/gc.c-2340- f2fs_up_write(&sbi->gc_lock);
--
fs/f2fs/node.c=437=static void cache_nat_entry(struct f2fs_sb_info *sbi, nid_t nid,
--
fs/f2fs/node.c-443- /* Let's mitigate lock contention of nat_tree_lock during checkpoint */
fs/f2fs/node.c:444: if (f2fs_rwsem_is_locked(&sbi->cp_global_sem))
fs/f2fs/node.c-445- return;
--
fs/f2fs/node.c=557=int f2fs_get_node_info(struct f2fs_sb_info *sbi, nid_t nid,
--
fs/f2fs/node.c-595- */
fs/f2fs/node.c:596: if (!f2fs_rwsem_is_locked(&sbi->cp_global_sem) || checkpoint_context) {
fs/f2fs/node.c-597- down_read(&curseg->journal_rwsem);
--
fs/f2fs/recovery.c=870=int f2fs_recover_fsync_data(struct f2fs_sb_info *sbi, bool check_only)
--
fs/f2fs/recovery.c-889- /* prevent checkpoint */
fs/f2fs/recovery.c:890: f2fs_down_write(&sbi->cp_global_sem);
fs/f2fs/recovery.c-891-
--
fs/f2fs/recovery.c-933-
fs/f2fs/recovery.c:934: f2fs_up_write(&sbi->cp_global_sem);
fs/f2fs/recovery.c-935-
--
fs/f2fs/super.c=4770=static int f2fs_fill_super(struct super_block *sb, struct fs_context *fc)
--
fs/f2fs/super.c-4800- mutex_init(&sbi->writepages);
fs/f2fs/super.c:4801: init_f2fs_rwsem(&sbi->cp_global_sem);
fs/f2fs/super.c-4802- init_f2fs_rwsem(&sbi->node_write);
]
|
| 64/2 |
2026/05/26 06:38 |
llm |
expert |
0m
Model:
Tokens:
input: 105615
output: 30
thoughts: 64
Thoughts: **Analyzing F2FS Synchronization Mechanisms**
Okay, let me think this through. I'm focusing on the synchronization mechanisms within F2FS. Specifically, I was initially keyed in on `cp_global_sem`. Now, I'm recalling where it's actually *used* within the code, and it looks like it's primarily confined to a few critical areas: `f2fs_write_meta_pages`, `f2fs_write_checkpoint`, `f2fs_resize_fs`, and `f2fs_recover_fsync_data`. Hmm, that's a relatively limited scope for a global semaphore. It suggests a focused protection strategy for those particular operations.
Now, my mind is drifting to another potential synchronization primitive: `cp_rwsem`. I'm starting to wonder about the interplay between these two β what are the specific use cases of the read-write semaphore? Is there overlap? Or do they fulfill distinct roles in ensuring the consistency of the file system's metadata and checkpointing activities? This needs further investigation. I'll need to drill down and understand how they interact, and, critically, the data or regions they're guarding.
|
| 65/2 |
2026/05/26 06:38 |
tool |
grepper |
0m
Error:
Args: map[Expression:f2fs_down_read(&sbi->cp_rwsem)]
Results:
|
| 66/2 |
2026/05/26 06:38 |
llm |
expert |
0m
Model:
Tokens:
input: 105971
output: 18
thoughts: 11
Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience, and presented as if it were my own internal monologue:
**Investigating `cp_rwsem`**
Right, let's dive in. The phrase `cp_rwsem` has popped up in my mental radar, and my immediate instinct is to figure out what context it's relevant to. I need to get a clear understanding of its purpose. It's likely a kernel structure or function, probably related to synchronization, given the `rwsem` part. "cp_" is likely a prefix. I'm assuming that itβs probably a part of some implementation in the kernel. Let's start by doing a quick search for its definition and usage. Gotta understand the scope here.
|
| 67/2 |
2026/05/26 06:38 |
tool |
grepper |
0m
Args:
Results: map[Output:fs/f2fs/f2fs.h=1236=struct f2fs_io_info {
--
fs/f2fs/f2fs.h-1252- unsigned int compr_blocks; /* # of compressed block addresses */
fs/f2fs/f2fs.h:1253: unsigned int need_lock:8; /* indicate we need to lock cp_rwsem */
fs/f2fs/f2fs.h-1254- unsigned int version:8; /* version of the node */
--
fs/f2fs/f2fs.h=1597=struct f2fs_sb_info {
--
fs/f2fs/f2fs.h-1632- struct f2fs_rwsem cp_global_sem; /* checkpoint procedure lock */
fs/f2fs/f2fs.h:1633: struct f2fs_rwsem cp_rwsem; /* blocking FS operations */
fs/f2fs/f2fs.h-1634- struct f2fs_rwsem node_write; /* locking node writes */
--
fs/f2fs/f2fs.h=2295=static inline void f2fs_lock_op(struct f2fs_sb_info *sbi)
fs/f2fs/f2fs.h-2296-{
fs/f2fs/f2fs.h:2297: f2fs_down_read(&sbi->cp_rwsem);
fs/f2fs/f2fs.h-2298-}
--
fs/f2fs/f2fs.h=2300=static inline int f2fs_trylock_op(struct f2fs_sb_info *sbi)
--
fs/f2fs/f2fs.h-2303- return 0;
fs/f2fs/f2fs.h:2304: return f2fs_down_read_trylock(&sbi->cp_rwsem);
fs/f2fs/f2fs.h-2305-}
--
fs/f2fs/f2fs.h=2307=static inline void f2fs_unlock_op(struct f2fs_sb_info *sbi)
fs/f2fs/f2fs.h-2308-{
fs/f2fs/f2fs.h:2309: f2fs_up_read(&sbi->cp_rwsem);
fs/f2fs/f2fs.h-2310-}
--
fs/f2fs/f2fs.h=2312=static inline void f2fs_lock_all(struct f2fs_sb_info *sbi)
fs/f2fs/f2fs.h-2313-{
fs/f2fs/f2fs.h:2314: f2fs_down_write(&sbi->cp_rwsem);
fs/f2fs/f2fs.h-2315-}
--
fs/f2fs/f2fs.h=2317=static inline void f2fs_unlock_all(struct f2fs_sb_info *sbi)
fs/f2fs/f2fs.h-2318-{
fs/f2fs/f2fs.h:2319: f2fs_up_write(&sbi->cp_rwsem);
fs/f2fs/f2fs.h-2320-}
--
fs/f2fs/segment.c=470=static inline bool excess_dirty_threshold(struct f2fs_sb_info *sbi)
fs/f2fs/segment.c-471-{
fs/f2fs/segment.c:472: int factor = f2fs_rwsem_is_locked(&sbi->cp_rwsem) ? 3 : 2;
fs/f2fs/segment.c-473- unsigned int dents = get_pages(sbi, F2FS_DIRTY_DENTS);
--
fs/f2fs/segment.c=489=void f2fs_balance_fs_bg(struct f2fs_sb_info *sbi, bool from_bg)
--
fs/f2fs/segment.c-518- if (is_inflight_io(sbi, REQ_TIME) ||
fs/f2fs/segment.c:519: (!f2fs_time_over(sbi, REQ_TIME) && f2fs_rwsem_is_locked(&sbi->cp_rwsem)))
fs/f2fs/segment.c-520- return;
--
fs/f2fs/super.c=4770=static int f2fs_fill_super(struct super_block *sb, struct fs_context *fc)
--
fs/f2fs/super.c-4804- spin_lock_init(&sbi->stat_lock);
fs/f2fs/super.c:4805: init_f2fs_rwsem(&sbi->cp_rwsem);
fs/f2fs/super.c-4806- init_f2fs_rwsem(&sbi->quota_sem);
]
|
| 68/2 |
2026/05/26 06:38 |
llm |
expert |
0m
Model:
Tokens:
input: 107460
output: 49
thoughts: 1359
Thoughts: **Analyzing the F2FS Synchronization Issue**
Okay, let's break this down. My initial thought was that `f2fs_lock_all` uses `cp_rwsem` for write, and looking at the `block_operations` code, the retry logic for flushing quotas seemed relevant. The code calls `f2fs_lock_all`, checks for quota flushing, and then unlocks. This suggested that `cp_rwsem` might be held during inode syncing. However, upon closer inspection, `f2fs_sync_dirty_inodes` is called *after* the unlock. This means `cp_rwsem` isn't held during the crucial `f2fs_sync_dirty_inodes` operation.
But wait, what about `cp_global_sem`? That's acquired for write during `f2fs_write_checkpoint`. So, that *is* held during `f2fs_sync_dirty_inodes`. The question then becomes: does `cp_global_sem` block other filesystem operations? No, `cp_rwsem` is responsible for that (via `f2fs_lock_op`). `cp_global_sem` is just for checkpointing to prevent concurrency. This means that if `sync()` hangs inside `f2fs_sync_dirty_inodes`, it *only* blocks other checkpoints, not regular operations like writes or directory creation.
So, if `sync()` hangs, the issue is localized to that single `sync()` call, not a system-wide lockup. But a hang *does* involve holding `s_umount` for read. This is where things get interesting. `quotactl` also requires `s_umount` for read. Consequently, `quotactl` wouldn't directly hang. The only way `quotactl` would hang is if the `mount` operation is holding `s_umount` for write. This explains the bug report. That means that the root cause must be related to the mounting process. `f2fs_fill_super` calls `f2fs_recover_fsync_data`, which in turn calls `f2fs_write_checkpoint`, potentially hanging the mount process.
This system-wide hang is only triggered during `mount`, `umount`, `remount`, or `freeze` actions, all of which require `CAP_SYS_ADMIN` privileges. An unprivileged user, therefore, can't directly trigger it, unless it's via a peripheral trigger, like an auto-mount.
But wait, can an unprivileged user still cause a system-wide problem with a full filesystem? If an unprivileged user can make `sync()` hang, holding `s_umount` for read, and root later tries to unmount the filesystem, then `umount` will block waiting for `s_umount` for write. Because the semaphore is *fair*, a pending writer blocks subsequent readers. Therefore, after that point, any process trying to call `sync()` or `quotactl` on *any* filesystem will also hang, waiting on `s_umount` for read!
So, in essence, an unprivileged user *could* cause a system-wide denial of service if the following sequence occurs: 1) the user fills the disk to trigger `-ENOSPC` or uses a corrupted filesystem. 2) The user calls `sync()` which can hang in `f2fs_sync_dirty_inodes`, holding `s_umount` for read. 3) Root, or some automated process, tries to `umount`. 4) Now, any subsequent attempt to `sync()` or `quotactl` will hang.
The problem could escalate because `sync()` iterates over *all* superblocks. If one superblock is blocked, `sync()` blocks on that superblock and so on.
The next question is, can we trigger the infinite loop *without* explicit fault injection? If the disk is full, `f2fs_do_write_data_page` returns `-ENOSPC`. This doesn't usually happen with overwrites, but, because F2FS is a log-structured filesystem, *every* write allocates a new block. This makes `-ENOSPC` more likely! So if the disk is full, any dirty page will fail to write, resulting in `-ENOSPC`. Then, `f2fs_write_single_data_page` redirties the page and returns `-ENOSPC`. And then, `f2fs_sync_dirty_inodes` will get stuck in an infinite loop!
This means that *any* time an F2FS filesystem fills, `sync()` could potentially hang indefinitely! That seems like a massive and easily reproducible bug. But wait, is there any protection against this? In `f
|
| 69/2 |
2026/05/26 06:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/data.c IncludeLines:true Name:f2fs_do_write_data_page]
Results: map[SourceCode:2652: int f2fs_do_write_data_page(struct f2fs_io_info *fio)
2653: {
2654: struct folio *folio = fio->folio;
2655: struct inode *inode = folio->mapping->host;
2656: struct dnode_of_data dn;
2657: struct node_info ni;
2658: bool ipu_force = false;
2659: bool atomic_commit;
2660: int err = 0;
2661:
2662: /* Use COW inode to make dnode_of_data for atomic write */
2663: atomic_commit = f2fs_is_atomic_file(inode) &&
2664: folio_test_f2fs_atomic(folio);
2665: if (atomic_commit)
2666: set_new_dnode(&dn, F2FS_I(inode)->cow_inode, NULL, NULL, 0);
2667: else
2668: set_new_dnode(&dn, inode, NULL, NULL, 0);
2669:
2670: if (need_inplace_update(fio) &&
2671: f2fs_lookup_read_extent_cache_block(inode, folio->index,
2672: &fio->old_blkaddr)) {
2673: if (!f2fs_is_valid_blkaddr(fio->sbi, fio->old_blkaddr,
2674: DATA_GENERIC_ENHANCE))
2675: return -EFSCORRUPTED;
2676:
2677: ipu_force = true;
2678: fio->need_lock = LOCK_DONE;
2679: goto got_it;
2680: }
2681:
2682: /* Deadlock due to between page->lock and f2fs_lock_op */
2683: if (fio->need_lock == LOCK_REQ && !f2fs_trylock_op(fio->sbi))
2684: return -EAGAIN;
2685:
2686: err = f2fs_get_dnode_of_data(&dn, folio->index, LOOKUP_NODE);
2687: if (err)
2688: goto out;
2689:
2690: fio->old_blkaddr = dn.data_blkaddr;
2691:
2692: /* This page is already truncated */
2693: if (fio->old_blkaddr == NULL_ADDR) {
2694: folio_clear_uptodate(folio);
2695: folio_clear_f2fs_gcing(folio);
2696: goto out_writepage;
2697: }
2698: got_it:
2699: if (__is_valid_data_blkaddr(fio->old_blkaddr) &&
2700: !f2fs_is_valid_blkaddr(fio->sbi, fio->old_blkaddr,
2701: DATA_GENERIC_ENHANCE)) {
2702: err = -EFSCORRUPTED;
2703: goto out_writepage;
2704: }
2705:
2706: /* wait for GCed page writeback via META_MAPPING */
2707: if (fio->meta_gc)
2708: f2fs_wait_on_block_writeback(inode, fio->old_blkaddr);
2709:
2710: /*
2711: * If current allocation needs SSR,
2712: * it had better in-place writes for updated data.
2713: */
2714: if (ipu_force ||
2715: (__is_valid_data_blkaddr(fio->old_blkaddr) &&
2716: need_inplace_update(fio))) {
2717: err = f2fs_encrypt_one_page(fio);
2718: if (err)
2719: goto out_writepage;
2720:
2721: folio_start_writeback(folio);
2722: f2fs_put_dnode(&dn);
2723: if (fio->need_lock == LOCK_REQ)
2724: f2fs_unlock_op(fio->sbi);
2725: err = f2fs_inplace_write_data(fio);
2726: if (err) {
2727: if (fscrypt_inode_uses_fs_layer_crypto(inode))
2728: fscrypt_finalize_bounce_page(&fio->encrypted_page);
2729: folio_end_writeback(folio);
2730: } else {
2731: set_inode_flag(inode, FI_UPDATE_WRITE);
2732: }
2733: trace_f2fs_do_write_data_page(folio, IPU);
2734: return err;
2735: }
2736:
2737: if (fio->need_lock == LOCK_RETRY) {
2738: if (!f2fs_trylock_op(fio->sbi)) {
2739: err = -EAGAIN;
2740: goto out_writepage;
2741: }
2742: fio->need_lock = LOCK_REQ;
2743: }
2744:
2745: err = f2fs_get_node_info(fio->sbi, dn.nid, &ni, false);
2746: if (err)
2747: goto out_writepage;
2748:
2749: fio->version = ni.version;
2750:
2751: err = f2fs_encrypt_one_page(fio);
2752: if (err)
2753: goto out_writepage;
2754:
2755: folio_start_writeback(folio);
2756:
2757: if (fio->compr_blocks && fio->old_blkaddr == COMPRESS_ADDR)
2758: f2fs_i_compr_blocks_update(inode, fio->compr_blocks - 1, false);
2759:
2760: /* LFS mode write path */
2761: f2fs_outplace_write_data(&dn, fio);
2762: trace_f2fs_do_write_data_page(folio, OPU);
2763: set_inode_flag(inode, FI_APPEND_WRITE);
2764: if (atomic_commit)
2765: folio_clear_f2fs_atomic(folio);
2766: out_writepage:
2767: f2fs_put_dnode(&dn);
2768: out:
2769: if (fio->need_lock == LOCK_REQ)
2770: f2fs_unlock_op(fio->sbi);
2771: return err;
2772: }
SourceFile:fs/f2fs/data.c]
|
| 70/2 |
2026/05/26 06:38 |
llm |
expert |
0m
Model:
Tokens:
input: 111875
output: 48
thoughts: 81
Thoughts: Okay, let me think this through.
**Analyzing the Allocation and Error Handling in `f2fs_outplace_write_data`**
Right, so `f2fs_outplace_write_data`, as I recall, it allocates a *new* block for writing data, which is standard for out-of-place modifications. Now, a critical question arises: what happens if the allocation fails? For instance, what if we run out of space β `-ENOSPC`? Does the function signal that failure somehow?
The immediate problem is, **`f2fs_outplace_write_data` is declared as returning `void`!** No direct return value to indicate success or failure. That feels⦠problematic, or at least, requiring further investigation. I need to understand how the error handling is implemented. Are there side effects, global variables, or error codes set elsewhere that I need to be aware of? Let's dive into the code itself and see how it works internally. I'm going to check the implementation of `f2fs_outplace_write_data` to see where error checks are performed and how they are handled. I need to confirm whether it is possible for an error to propagate up the call chain.
|
| 71/2 |
2026/05/26 06:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/segment.c IncludeLines:true Name:f2fs_outplace_write_data]
Results: map[SourceCode:4006: void f2fs_outplace_write_data(struct dnode_of_data *dn,
4007: struct f2fs_io_info *fio)
4008: {
4009: struct f2fs_sb_info *sbi = fio->sbi;
4010: struct f2fs_summary sum;
4011:
4012: f2fs_bug_on(sbi, dn->data_blkaddr == NULL_ADDR);
4013: if (fio->io_type == FS_DATA_IO || fio->io_type == FS_CP_DATA_IO)
4014: f2fs_update_age_extent_cache(dn);
4015: set_summary(&sum, dn->nid, dn->ofs_in_node, fio->version);
4016: do_write_page(&sum, fio);
4017: f2fs_update_data_blkaddr(dn, fio->new_blkaddr);
4018:
4019: f2fs_update_iostat(sbi, dn->inode, fio->io_type, F2FS_BLKSIZE);
4020: }
SourceFile:fs/f2fs/segment.c]
|
| 72/2 |
2026/05/26 06:38 |
llm |
expert |
1m
Model:
Tokens:
input: 112610
output: 43
thoughts: 531
Thoughts: **Analyzing the F2FS Write Path for Out-of-Space Conditions**
Okay, let's break this down. My initial thought was that `f2fs_outplace_write_data` was the key failure point, but it turns out that's *not* the case. It appears it doesn't handle the allocation itself; it assumes the block is already there. So, the block allocation is happening upstream, in `f2fs_get_dnode_of_data`, which seems to be the critical function to examine.
Aha! `f2fs_get_dnode_of_data` allocates the block. Now, if *that* fails with `-ENOSPC`, then `f2fs_do_write_data_page` returns `-ENOSPC`. That makes sense. Then the error propagates: `-ENOSPC` from `f2fs_write_single_data_page` to `redirty_out`, then from `f2fs_write_cache_pages`, and eventually the process breaks. This is where I started thinking about the implications for `f2fs_sync_dirty_inodes`, which appears to loop infinitely if there's an out-of-space error.
But wait a minute... I need to consider `LOOKUP_NODE`. If the disk is *actually* full, `f2fs_get_dnode_of_data` with `LOOKUP_NODE` will *not* allocate anything. `LOOKUP_NODE` just looks up the node. So, if the block isn't allocated, `dn.data_blkaddr` becomes `NULL_ADDR`. And that leads us to the crucial part in `f2fs_do_write_data_page`:
```c
/* This page is already truncated */
if (fio->old_blkaddr == NULL_ADDR) {
folio_clear_uptodate(folio);
folio_clear_f2fs_gcing(folio);
goto out_writepage;
}
```
This clears the 'uptodate' flag and returns! And crucially, it returns `err = 0`, *not* `-ENOSPC`. That's the key. Why? Because the `writepage` function is only called for pages that *already* have blocks allocated, so there isn't an error. In f2fs, blocks are allocated during write_begin or fallocate, not in writepage. This makes sense from an implementation standpoint.
So, for out-of-place updates, `f2fs_outplace_write_data` will write to a new block, but the allocation isn't happening there. *That* allocation happens during `do_write_page` instead. Okay, I need to go back and check `do_write_page`. I must have missed something.
|
| 73/2 |
2026/05/26 06:40 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/segment.c IncludeLines:true Name:do_write_page]
Results: map[SourceCode:3932: static void do_write_page(struct f2fs_summary *sum, struct f2fs_io_info *fio)
3933: {
3934: struct folio *folio = fio->folio;
3935: enum log_type type = __get_segment_type(fio);
3936: int seg_type = log_type_to_seg_type(type);
3937: bool keep_order = (f2fs_lfs_mode(fio->sbi) &&
3938: seg_type == CURSEG_COLD_DATA);
3939:
3940: if (keep_order)
3941: f2fs_down_read(&fio->sbi->io_order_lock);
3942:
3943: if (f2fs_allocate_data_block(fio->sbi, folio, fio->old_blkaddr,
3944: &fio->new_blkaddr, sum, type, fio)) {
3945: if (fscrypt_inode_uses_fs_layer_crypto(folio->mapping->host))
3946: fscrypt_finalize_bounce_page(&fio->encrypted_page);
3947: folio_end_writeback(folio);
3948: if (f2fs_in_warm_node_list(fio->sbi, folio))
3949: f2fs_del_fsync_node_entry(fio->sbi, folio);
3950: f2fs_bug_on(fio->sbi, !is_set_ckpt_flags(fio->sbi,
3951: CP_ERROR_FLAG));
3952: goto out;
3953: }
3954:
3955: f2fs_bug_on(fio->sbi, !f2fs_is_valid_blkaddr_raw(fio->sbi,
3956: fio->new_blkaddr, DATA_GENERIC_ENHANCE));
3957:
3958: if (GET_SEGNO(fio->sbi, fio->old_blkaddr) != NULL_SEGNO)
3959: f2fs_invalidate_internal_cache(fio->sbi, fio->old_blkaddr, 1);
3960:
3961: /* writeout dirty page into bdev */
3962: f2fs_submit_page_write(fio);
3963:
3964: f2fs_update_device_state(fio->sbi, fio->ino, fio->new_blkaddr, 1);
3965: out:
3966: if (keep_order)
3967: f2fs_up_read(&fio->sbi->io_order_lock);
3968: }
SourceFile:fs/f2fs/segment.c]
|
| 74/2 |
2026/05/26 06:40 |
llm |
expert |
0m
Model:
Tokens:
input: 114580
output: 47
thoughts: 386
Thoughts: **F2FS Data Block Allocation - A Potential Disaster?**
Okay, hold on a second! I'm seeing something here that's making my expert senses tingle. The code in `f2fs_allocate_data_block`... it *can* fail, right? And when it does, look at this:
```c
if (f2fs_allocate_data_block(...)) {
...
folio_end_writeback(folio);
...
f2fs_bug_on(fio->sbi, !is_set_ckpt_flags(fio->sbi, CP_ERROR_FLAG));
goto out;
}
```
My brain is flashing red now because it *assumes* the `CP_ERROR_FLAG` is set after a failed allocation! But it doesn't return an error! `do_write_page` returns `void`, so `f2fs_outplace_write_data` is also `void`, and `f2fs_do_write_data_page` returns 0. If it's zero, `f2fs_write_single_data_page` thinks it was a success! And then it *clears* the dirty bit!
So, the loop in `f2fs_sync_dirty_inodes`... it WILL terminate prematurely! That's bad. Wait, it is even worse. Because if allocation fails, it calls `folio_end_writeback(folio)`, and `f2fs_write_single_data_page` calls `inode_dec_dirty_pages(inode)` and `folio_unlock(folio)`. This means the page is no longer dirty, and the loop is going to terminate thinking that the block was written.
This is a real problem. Wait, why does `f2fs_allocate_data_block` fail in the first place? And if it fails because of `-ENOSPC`, does it even set `CP_ERROR_FLAG`? I need to go and analyze the `f2fs_allocate_data_block` function. I need to be sure that it handles space allocation correctly.
|
| 75/2 |
2026/05/26 06:40 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/segment.c IncludeLines:true Name:f2fs_allocate_data_block]
Results: map[SourceCode:3749: int f2fs_allocate_data_block(struct f2fs_sb_info *sbi, struct folio *folio,
3750: block_t old_blkaddr, block_t *new_blkaddr,
3751: struct f2fs_summary *sum, int type,
3752: struct f2fs_io_info *fio)
3753: {
3754: struct sit_info *sit_i = SIT_I(sbi);
3755: struct curseg_info *curseg = CURSEG_I(sbi, type);
3756: unsigned long long old_mtime;
3757: bool from_gc = (type == CURSEG_ALL_DATA_ATGC);
3758: struct seg_entry *se = NULL;
3759: bool segment_full = false;
3760: int ret = 0;
3761:
3762: f2fs_down_read(&SM_I(sbi)->curseg_lock);
3763:
3764: mutex_lock(&curseg->curseg_mutex);
3765: down_write(&sit_i->sentry_lock);
3766:
3767: if (curseg->segno == NULL_SEGNO) {
3768: ret = -ENOSPC;
3769: goto out_err;
3770: }
3771:
3772: if (from_gc) {
3773: f2fs_bug_on(sbi, GET_SEGNO(sbi, old_blkaddr) == NULL_SEGNO);
3774: se = get_seg_entry(sbi, GET_SEGNO(sbi, old_blkaddr));
3775: sanity_check_seg_type(sbi, se->type);
3776: f2fs_bug_on(sbi, IS_NODESEG(se->type));
3777: }
3778: *new_blkaddr = NEXT_FREE_BLKADDR(sbi, curseg);
3779:
3780: f2fs_bug_on(sbi, curseg->next_blkoff >= BLKS_PER_SEG(sbi));
3781:
3782: f2fs_wait_discard_bio(sbi, *new_blkaddr);
3783:
3784: curseg->sum_blk->entries[curseg->next_blkoff] = *sum;
3785: if (curseg->alloc_type == SSR) {
3786: curseg->next_blkoff = f2fs_find_next_ssr_block(sbi, curseg);
3787: } else {
3788: curseg->next_blkoff++;
3789: if (F2FS_OPTION(sbi).fs_mode == FS_MODE_FRAGMENT_BLK)
3790: f2fs_randomize_chunk(sbi, curseg);
3791: }
3792: if (curseg->next_blkoff >= f2fs_usable_blks_in_seg(sbi, curseg->segno))
3793: segment_full = true;
3794: stat_inc_block_count(sbi, curseg);
3795:
3796: if (from_gc) {
3797: old_mtime = get_segment_mtime(sbi, old_blkaddr);
3798: } else {
3799: update_segment_mtime(sbi, old_blkaddr, 0);
3800: old_mtime = 0;
3801: }
3802: update_segment_mtime(sbi, *new_blkaddr, old_mtime);
3803:
3804: /*
3805: * SIT information should be updated before segment allocation,
3806: * since SSR needs latest valid block information.
3807: */
3808: update_sit_entry(sbi, *new_blkaddr, 1);
3809: update_sit_entry(sbi, old_blkaddr, -1);
3810:
3811: /*
3812: * If the current segment is full, flush it out and replace it with a
3813: * new segment.
3814: */
3815: if (segment_full) {
3816: if (type == CURSEG_COLD_DATA_PINNED &&
3817: !((curseg->segno + 1) % sbi->segs_per_sec)) {
3818: write_sum_page(sbi, curseg->sum_blk,
3819: GET_SUM_BLOCK(sbi, curseg->segno));
3820: reset_curseg_fields(curseg);
3821: goto skip_new_segment;
3822: }
3823:
3824: if (from_gc) {
3825: ret = get_atssr_segment(sbi, type, se->type,
3826: AT_SSR, se->mtime);
3827: } else {
3828: if (need_new_seg(sbi, type))
3829: ret = new_curseg(sbi, type, false);
3830: else
3831: ret = change_curseg(sbi, type);
3832: stat_inc_seg_type(sbi, curseg);
3833: }
3834:
3835: if (ret)
3836: goto out_err;
3837: }
3838:
3839: skip_new_segment:
3840: /*
3841: * segment dirty status should be updated after segment allocation,
3842: * so we just need to update status only one time after previous
3843: * segment being closed.
3844: */
3845: locate_dirty_segment(sbi, GET_SEGNO(sbi, old_blkaddr));
3846: locate_dirty_segment(sbi, GET_SEGNO(sbi, *new_blkaddr));
3847:
3848: if (IS_DATASEG(curseg->seg_type))
3849: atomic64_inc(&sbi->allocated_data_blocks);
3850:
3851: up_write(&sit_i->sentry_lock);
3852:
3853: if (folio && IS_NODESEG(curseg->seg_type)) {
3854: fill_node_footer_blkaddr(folio, NEXT_FREE_BLKADDR(sbi, curseg));
3855:
3856: f2fs_inode_chksum_set(sbi, folio);
3857: }
3858:
3859: if (fio) {
3860: struct f2fs_bio_info *io;
3861:
3862: INIT_LIST_HEAD(&fio->list);
3863: fio->in_list = 1;
3864: io = sbi->write_io[fio->type] + fio->temp;
3865: spin_lock(&io->io_lock);
3866: list_add_tail(&fio->list, &io->io_list);
3867: spin_unlock(&io->io_lock);
3868: }
3869:
3870: mutex_unlock(&curseg->curseg_mutex);
3871: f2fs_up_read(&SM_I(sbi)->curseg_lock);
3872: return 0;
3873:
3874: out_err:
3875: *new_blkaddr = NULL_ADDR;
3876: up_write(&sit_i->sentry_lock);
3877: mutex_unlock(&curseg->curseg_mutex);
3878: f2fs_up_read(&SM_I(sbi)->curseg_lock);
3879: return ret;
3880: }
SourceFile:fs/f2fs/segment.c]
|
| 76/2 |
2026/05/26 06:40 |
llm |
expert |
0m
Model:
Tokens:
input: 117717
output: 23
thoughts: 421
Thoughts: **F2FS Data Block Allocation - A Deep Dive**
Okay, here's what's running through my head about this `f2fs_allocate_data_block` function. It seems like a critical piece of the puzzle, and I'm seeing a potential for a really nasty bug. Right, so `f2fs_allocate_data_block` *can* return `-ENOSPC`, which, okay, I understand. If the `curseg->segno` is `NULL_SEGNO`, it's going to return `-ENOSPC`. Makes sense. And also if there's an error from either `new_curseg` or `change_curseg`, that also makes sense; it'll propagate the error.
Okay, and I see that when there's an error returned, the `do_write_page` function handles it like this:
```c
if (f2fs_allocate_data_block(...)) {
...
folio_end_writeback(folio);
...
f2fs_bug_on(fio->sbi, !is_set_ckpt_flags(fio->sbi, CP_ERROR_FLAG));
goto out;
}
```
Hold on... *that* `f2fs_bug_on` call is screaming at me! It checks `!is_set_ckpt_flags(fio->sbi, CP_ERROR_FLAG)`. So, if the allocation fails, and we *don't* have `CP_ERROR_FLAG` set, that's a bug! It's going to trigger a `BUG_ON`. My mind is racing. Does `f2fs_allocate_data_block` *set* `CP_ERROR_FLAG`? I don't see anything. Wait, it doesn't? If `f2fs_allocate_data_block` returns `-ENOSPC` *and* it doesn't set that flag, then the `f2fs_bug_on` *will* trigger.
If that `f2fs_bug_on` actually triggers, it means the system will either panic (if `panic_on_bug` is set) or print a warning and keep going. I'm going to quickly check how `f2fs_bug_on` is defined:
```c
#define f2fs_bug_on(sbi, condition) BUG_ON(condition)
```
Right, it's just a simple macro that calls `BUG_ON`. And that makes it a hard error. This is a potential for a really serious issue; a failing allocation that doesn't trigger a checkpoint error *will* result in a panic or a really bad situation. This needs immediate attention.
|
| 77/2 |
2026/05/26 06:40 |
tool |
grepper |
0m
Args: map[Expression:#define f2fs_bug_on]
Results: map[Output:fs/f2fs/f2fs.h=31=struct pagevec;
--
fs/f2fs/f2fs.h-33-#ifdef CONFIG_F2FS_CHECK_FS
fs/f2fs/f2fs.h:34:#define f2fs_bug_on(sbi, condition) BUG_ON(condition)
fs/f2fs/f2fs.h-35-#else
fs/f2fs/f2fs.h:36:#define f2fs_bug_on(sbi, condition) \
fs/f2fs/f2fs.h-37- do { \
]
|
| 78/2 |
2026/05/26 06:40 |
llm |
expert |
0m
Model:
Tokens:
input: 118929
output: 45
thoughts: 172
Thoughts: **Analyzing the F2FS Error Handling**
Okay, so I'm looking at this `f2fs_bug_on` macro. If `CONFIG_F2FS_CHECK_FS` isn't enabled, then this thing kicks in, right? It seems pretty straightforward. If a specified `condition` is `unlikely` (which is good practice for performance, I'll give them that), it triggers a `WARN_ON(1)`. Standard warning, nothing fancy there. Then it sets the `SBI_NEED_FSCK` flag, indicating that a file system check is required. Makes sense. But here's where things get interesting, it calls `f2fs_handle_error`. Ah, *that's* the bit I really need to focus on.
I know from experience with F2FS that `f2fs_handle_error` is a crucial function, as it is the error handling function. Wait... I need to be sure... I need to check `f2fs_handle_error` itself. I need to know *specifically* if `f2fs_handle_error` is going to set `CP_ERROR_FLAG`, because that can have big implications for how the recovery process proceeds. This entire bug-on mechanism relies on the consistent behaviour of this function. I need to be 100% certain how this function sets this flag. Gotta dig into that next.
|
| 79/2 |
2026/05/26 06:40 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/super.c IncludeLines:true Name:f2fs_handle_error]
Results: map[SourceCode:4485: void f2fs_handle_error(struct f2fs_sb_info *sbi, unsigned char error)
4486: {
4487: f2fs_save_errors(sbi, error);
4488: f2fs_record_errors(sbi, error);
4489: }
SourceFile:fs/f2fs/super.c]
|
| 80/2 |
2026/05/26 06:40 |
llm |
expert |
0m
Model:
Tokens:
input: 119568
output: 50
thoughts: 750
Thoughts: **F2FS Infinite Loop Analysis**
Okay, so here's what's going through my head as I dissect this F2FS issue. I see that `f2fs_handle_error` *doesn't* set `CP_ERROR_FLAG`, which immediately throws up a red flag. If `f2fs_allocate_data_block` fails, and `CP_ERROR_FLAG` isn't set, the system will just log a warning and continue, unless `CONFIG_F2FS_CHECK_FS` is enabled. Then the write functions, `do_write_page` and `f2fs_outplace_write_data`, return `void`, and `f2fs_do_write_data_page` returns `0`.
Hold on, if `f2fs_do_write_data_page` returns `0`, `f2fs_write_single_data_page` thinks it succeeded and clears the dirty bit. Which means the page is no longer considered dirty, and the loop in `f2fs_sync_dirty_inodes` *should* terminate. But wait, we have a problem. If `f2fs_do_write_data_page` returns `0`, why does `f2fs_sync_dirty_inodes` loop infinitely? The answer must be that there's an error before `f2fs_outplace_write_data` is called.
For example, `f2fs_get_dnode_of_data` could return an error, like `-ENOMEM`, due to fault injection. If that happens, `f2fs_do_write_data_page` returns `-ENOMEM`. Then `f2fs_write_single_data_page` goes to `redirty_out` and returns `-ENOMEM`, then `f2fs_write_cache_pages` breaks and returns `-ENOMEM`. This cascades up to `__f2fs_write_data_pages`, which returns `-ENOMEM`, and finally, the dreaded infinite loop in `f2fs_sync_dirty_inodes` appears!
The root cause is that `f2fs_get_dnode_of_data` returning `-ENOMEM` *doesn't* set `CP_ERROR_FLAG` and *doesn't* clear the dirty bit. So the page remains dirty, and the inode remains in the dirty list, leading to the infinite loop.
So the problem is, essentially, `f2fs_get_dnode_of_data` returning an error (like `-ENOMEM`) without properly handling the dirty bit or setting the error flag. But is it only fault injection that can trigger this condition? Not necessarily. Could be a straight-up `kmalloc` failure, that would also return `-ENOMEM`, but these kinds of failures are normally transient. What if `f2fs_get_dnode_of_data` returns `-ENOSPC`? Wait, though! It can't, since `LOOKUP_NODE` doesn't allocate blocks, only reads the node page. If reading the node page fails, for example returning `-EIO` or `-EFSCORRUPTED`, the error is returned. If it returns `-EFSCORRUPTED`, does that set the `CP_ERROR_FLAG`? I need to check `f2fs_get_dnode_of_data` again.
|
| 81/2 |
2026/05/26 06:40 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/f2fs/data.c IncludeLines:true Name:f2fs_get_dnode_of_data]
Results: map[SourceCode: 783: /*
784: * Caller should call f2fs_put_dnode(dn).
785: * Also, it should grab and release a rwsem by calling f2fs_lock_op() and
786: * f2fs_unlock_op() only if mode is set with ALLOC_NODE.
787: */
788: int f2fs_get_dnode_of_data(struct dnode_of_data *dn, pgoff_t index, int mode)
789: {
790: struct f2fs_sb_info *sbi = F2FS_I_SB(dn->inode);
791: struct folio *nfolio[4];
792: struct folio *parent = NULL;
793: int offset[4];
794: unsigned int noffset[4];
795: nid_t nids[4];
796: int level, i = 0;
797: int err = 0;
798:
799: level = get_node_path(dn->inode, index, offset, noffset);
800: if (level < 0)
801: return level;
802:
803: nids[0] = dn->inode->i_ino;
804:
805: if (!dn->inode_folio) {
806: nfolio[0] = f2fs_get_inode_folio(sbi, nids[0]);
807: if (IS_ERR(nfolio[0]))
808: return PTR_ERR(nfolio[0]);
809: } else {
810: nfolio[0] = dn->inode_folio;
811: }
812:
813: /* if inline_data is set, should not report any block indices */
814: if (f2fs_has_inline_data(dn->inode) && index) {
815: err = -ENOENT;
816: f2fs_folio_put(nfolio[0], true);
817: goto release_out;
818: }
819:
820: parent = nfolio[0];
821: if (level != 0)
822: nids[1] = get_nid(parent, offset[0], true);
823: dn->inode_folio = nfolio[0];
824: dn->inode_folio_locked = true;
825:
826: /* get indirect or direct nodes */
827: for (i = 1; i <= level; i++) {
828: bool done = false;
829:
830: if (nids[i] && nids[i] == dn->inode->i_ino) {
831: err = -EFSCORRUPTED;
832: f2fs_err_ratelimited(sbi,
833: "inode mapping table is corrupted, run fsck to fix it, "
834: "ino:%lu, nid:%u, level:%d, offset:%d",
835: dn->inode->i_ino, nids[i], level, offset[level]);
836: set_sbi_flag(sbi, SBI_NEED_FSCK);
837: goto release_pages;
838: }
839:
840: if (!nids[i] && mode == ALLOC_NODE) {
841: /* alloc new node */
842: if (!f2fs_alloc_nid(sbi, &(nids[i]))) {
843: err = -ENOSPC;
844: goto release_pages;
845: }
846:
847: dn->nid = nids[i];
848: nfolio[i] = f2fs_new_node_folio(dn, noffset[i]);
849: if (IS_ERR(nfolio[i])) {
850: f2fs_alloc_nid_failed(sbi, nids[i]);
851: err = PTR_ERR(nfolio[i]);
852: goto release_pages;
853: }
854:
855: set_nid(parent, offset[i - 1], nids[i], i == 1);
856: f2fs_alloc_nid_done(sbi, nids[i]);
857: done = true;
858: } else if (mode == LOOKUP_NODE_RA && i == level && level > 1) {
859: nfolio[i] = f2fs_get_node_folio_ra(parent, offset[i - 1]);
860: if (IS_ERR(nfolio[i])) {
861: err = PTR_ERR(nfolio[i]);
862: goto release_pages;
863: }
864: done = true;
865: }
866: if (i == 1) {
867: dn->inode_folio_locked = false;
868: folio_unlock(parent);
869: } else {
870: f2fs_folio_put(parent, true);
871: }
872:
873: if (!done) {
874: nfolio[i] = f2fs_get_node_folio(sbi, nids[i]);
875: if (IS_ERR(nfolio[i])) {
876: err = PTR_ERR(nfolio[i]);
877: f2fs_folio_put(nfolio[0], false);
878: goto release_out;
879: }
880: }
881: if (i < level) {
882: parent = nfolio[i];
883: nids[i + 1] = get_nid(parent, offset[i], false);
884: }
885: }
886: dn->nid = nids[level];
887: dn->ofs_in_node = offset[level];
888: dn->node_folio = nfolio[level];
889: dn->data_blkaddr = f2fs_data_blkaddr(dn);
890:
891: if (is_inode_flag_set(dn->inode, FI_COMPRESSED_FILE) &&
892: f2fs_sb_has_readonly(sbi)) {
893: unsigned int cluster_size = F2FS_I(dn->inode)->i_cluster_size;
894: unsigned int ofs_in_node = dn->ofs_in_node;
895: pgoff_t fofs = index;
896: unsigned int c_len;
897: block_t blkaddr;
898:
899: /* should align fofs and ofs_in_node to cluster_size */
900: if (fofs % cluster_size) {
901: fofs = round_down(fofs, cluster_size);
902: ofs_in_node = round_down(ofs_in_node, cluster_size);
903: }
904:
905: c_len = f2fs_cluster_blocks_are_contiguous(dn, ofs_in_node);
906: if (!c_len)
907: goto out;
908:
909: blkaddr = data_blkaddr(dn->inode, dn->node_folio, ofs_in_node);
910: if (blkaddr == COMPRESS_ADDR)
911: blkaddr = data_blkaddr(dn->inode, dn->node_folio,
912: ofs_in_node + 1);
913:
914: f2fs_update_read_extent_tree_range_compressed(dn->inode,
915: fofs, blkaddr, cluster_size, c_len);
916: }
917: out:
918: return 0;
919:
920: release_pages:
921: f2fs_folio_put(parent, true);
922: if (i > 1)
923: f2fs_folio_put(nfolio[0], false);
924: release_out:
925: dn->inode_folio = NULL;
926: dn->node_folio = NULL;
927: if (err == -ENOENT) {
928: dn->cur_level = i;
929: dn->max_level = level;
930: dn->ofs_in_node = offset[level];
931: }
932: return err;
933: }
SourceFile:fs/f2fs/node.c]
|
| 82/2 |
2026/05/26 06:40 |
llm |
expert |
0m
Model:
Tokens:
input: 123599
output: 66
thoughts: 2433
Thoughts: **Analysis of Potential Infinite Loop in f2fs Filesystem Synchronization**
Okay, let's break this down. My thought process is as follows: If `f2fs_get_node_folio` fails, it returns the error, *crucially* without setting `CP_ERROR_FLAG`. This absence of the flag is key. Now, if we get an I/O error or corrupted data during node page reads, `f2fs_get_dnode_of_data` returns an error, and this cascades through `f2fs_do_write_data_page`, `f2fs_write_single_data_page`, and `f2fs_write_cache_pages`. Ultimately, `__f2fs_write_data_pages` returns an error, and this is where things go south. `f2fs_sync_dirty_inodes` enters an infinite loop! Why? Because the page gets redirtied, but the crucial `CP_ERROR_FLAG` is *not* set. The loop continues indefinitely since the file's dirty bit stays set.
I need to circle back and confirm: if `f2fs_get_node_folio` fails, does it ever set `CP_ERROR_FLAG`? No, it just returns the error code, and if the failure is due to `-EFSCORRUPTED`, the `SBI_NEED_FSCK` is set, but not the critical flag. Now, looking at `f2fs_get_dnode_of_data`, if it detects corruption, it returns `-EFSCORRUPTED`, sets `SBI_NEED_FSCK`, and... *still* no `CP_ERROR_FLAG`. This confirms my suspicion: infinite loop potential. This is a legitimate vulnerability!
This infinite loop during `f2fs_sync_dirty_inodes` can occur during `sync()` or, significantly, during `mount` if recovering `fsync` data. The trigger is filesystem corruption, making it a "Malicious Filesystem Trigger." It also seems that inserting a USB drive could cause this to happen, making it a "Peripheral Trigger."
Is it exploitable? No, it just results in a DoS via an infinite loop. Could an unprivileged user exploit this? Mounting the filesystem requires `CAP_SYS_ADMIN`, so normally, no. However, what if a user *already* has a corrupted filesystem? Let's say, they write to a file, and a read fails because the filesystem is corrupted. Then, if they call `sync()`, it can cause the loop, too. However, this only impacts the calling process. It doesn't hang the whole system unless `umount` is called, which requires root privileges. So, to get a system-wide hang, you need a privileged operation, but the infinite loop consumes 100% of a CPU core, which is exploitable. So, it appears that an unprivileged user *can* trigger the loop, *if* the filesystem is already corrupted. The corrupted USB drive is already a Peripheral Trigger, and the unprivileged user merely serves as a way to "activate" it.
So, here's the assessment:
- Exploitable: No
- Denial of Service: Yes
- Accessible From Unprivileged Processes: No (requires a privileged mount to establish)
- Accessible From User Namespaces: No (lacks `FS_USERNS_MOUNT`)
- VM Guest/Host Trigger: No
- Network Trigger: No
- Peripheral Trigger: Yes
- Malicious Filesystem Trigger: Yes
Okay, let's do a sanity check. Could running out of disk space trigger this loop? If `f2fs_get_dnode_of_data` with `LOOKUP_NODE` is called, it won't allocate blocks, and there will be no `-ENOSPC`. However, what if `f2fs_do_write_data_page` needs to allocate a block for an out-of-place update, that calls `f2fs_allocate_data_block`. If *that* fails with `-ENOSPC`, it hits `f2fs_bug_on`, which, if `CONFIG_F2FS_CHECK_FS` is NOT set (which it usually isn't in production), calls `f2fs_handle_error`.
Crucially, `f2fs_handle_error` *doesn't* set `CP_ERROR_FLAG`. After which `f2fs_do_write_data_page` returns 0 which clears the dirty bit, and the loop *terminates*! That means the `-ENOSPC` scenario drops the data, but it *doesn't* cause the infinite loop. This known issue can cause data loss, but
|
| 83/2 |
2026/05/26 06:41 |
tool |
set-results |
0m
Args: map[DenialOfService:true Exploitable:false FilesystemTrigger:true NetworkTrigger:false PeripheralTrigger:true RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
Results: map[DenialOfService:true Exploitable:false FilesystemTrigger:true NetworkTrigger:false PeripheralTrigger:true RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
|
| 84/2 |
2026/05/26 06:41 |
llm |
expert |
0m
Model:
Tokens:
input: 127187
output: 788
thoughts: 137
Thoughts: Here's my take on the situation:
**Finalizing the Analysis: A Clear and Concise Response**
Okay, the results are in, and I've confirmed everything. Time to write up the final response. My reasoning is rock solid; I've meticulously considered every angle. Let me recap: This isn't just a hypothetical vulnerability; it's a real, demonstrable issue.
* **Exploitable?** No, but for a terrible reason - it's an infinite loop, resulting in a system hang.
* **Denial of Service?** Absolutely, unequivocally yes. This hang would cripple critical operations like `sync` and `quotactl`. This is a serious issue.
* **Unprivileged or User Namespace?** No, thankfully, exploiting the vulnerability requires `CAP_SYS_ADMIN` during the initial userns mount, which limits the scope.
* **Peripheral/Filesystem Trigger?** Yes, the most concerning aspect. A maliciously crafted USB drive, auto-mounted, could trigger this. This has serious implications for system security.
* **Network/Remote/VM?** No, that's not the attack vector here.
Now, I'll organize my response in a clear and easy-to-understand format. This deserves a concise and direct explanation of the implications.
|
| 85/1 |
2026/05/26 06:41 |
action |
format-explanation |
0m
Results: map[Explanation:Based on the analysis of the bug report and the f2fs source code, here is the
assessment of the security impact:
### Bug Mechanism
The bug is a hung task (infinite loop) in `f2fs_sync_dirty_inodes` that occurs
during the mounting of a crafted f2fs image.
When mounting an f2fs filesystem, if there is fsync data to recover,
`f2fs_fill_super` calls `f2fs_recover_fsync_data`, which in turn calls
`f2fs_write_checkpoint` -> `block_operations` -> `f2fs_sync_dirty_inodes`.
If writing a dirty page fails during this process (e.g., due to a memory
allocation failure simulated by fault injection, or an `-EIO` / `-EFSCORRUPTED`
error caused by a corrupted node page in a malicious image),
`f2fs_write_single_data_page` redirties the page and returns the error. However,
`f2fs_sync_dirty_inodes` ignores the error returned by `filemap_fdatawrite` and
continues to loop as long as the inode's dirty list is not empty. Because the
page was redirtied and the error did not set the `CP_ERROR_FLAG` to abort the
checkpoint, the inode remains in the dirty list, resulting in an infinite loop.
### System-Wide Impact
During this infinite loop, the `mount` syscall holds the f2fs superblock's
`s_umount` semaphore for **write**.
Because `s_umount` is held for write indefinitely, any other system calls that
attempt to acquire `s_umount` for read on this superblock will block forever.
This includes system-wide operations like `quotactl` (which iterates over all
superblocks via `__iterate_supers` and takes `s_umount` for read) and `sync`.
This perfectly matches the syzkaller bug report where multiple tasks are blocked
for more than 143 seconds in `quotactl` -> `quota_sync_all` ->
`__iterate_supers` -> `super_lock`.
### Analysis Dimensions
* **Exploitable:** **No.** The bug is an infinite loop leading to a CPU spin and
a deadlock of the `s_umount` semaphore. There is no memory corruption,
use-after-free, or out-of-bounds access involved, meaning it cannot be exploited
for privilege escalation or arbitrary code execution.
* **Denial Of Service:** **Yes.** The infinite loop consumes a CPU core and
permanently blocks system-wide operations like `sync()` and `quotactl()`,
leading to a Denial of Service.
* **Accessible From Unprivileged Processes:** **No.** Mounting an f2fs
filesystem requires `CAP_SYS_ADMIN` in the initial user namespace. Unprivileged
users cannot trigger this directly.
* **Accessible From User Namespaces:** **No.** The f2fs filesystem does not have
the `FS_USERNS_MOUNT` flag, meaning it cannot be mounted from within an
unprivileged user namespace.
* **Peripheral Trigger:** **Yes.** An attacker can trigger this bug by plugging
in a malicious USB drive containing a crafted f2fs image, provided the desktop
environment (e.g., udisks2) is configured to auto-mount external media.
* **Malicious Filesystem Trigger:** **Yes.** The bug is triggered by parsing and
recovering a corrupted/malicious f2fs filesystem image.
* **VM Guest Trigger:** **No.**
* **VM Host Trigger in The Confidential Computing Context:** **No.**
* **Ethernet Network Trigger:** **No.**
* **Other Remote Trigger:** **No.**]
|