| 0/0 |
2026/09/17 00:00 |
flow |
assessment-security |
49m
Results: map[DenialOfService:true Explanation:Based on the analysis of the bug report and the kernel source code, here is the
assessment of the vulnerability:
**Bug Mechanism and Impact:**
The lockdep warning indicates a circular locking dependency (deadlock) involving
the memory reclaim context (`fs_reclaim`), OCFS2's `nfs_sync_rwlock`, and the
`inode_lock` of OCFS2 system inodes (like `INODE_ALLOC_SYSTEM_INODE`).
The deadlock occurs due to two concurrent paths:
1. **Process A (e.g., `mkdir`)**: Holds the `inode_lock` of an OCFS2 system
inode (e.g., during `ocfs2_reserve_new_metadata_blocks`) and allocates memory
using `GFP_KERNEL` (via `kzalloc_obj`). This triggers direct memory reclaim,
causing the process to wait on the `fs_reclaim` pseudo-lock.
2. **Process B (`kswapd0` or another direct reclaimer)**: Holds the `fs_reclaim`
lock while shrinking the dcache. It finds a dentry on the LRU, kills it, and
drops the last reference to its inode. If the inode has `i_nlink == 0`,
`evict()` synchronously calls `ocfs2_delete_inode()`. This function attempts to
acquire `nfs_sync_rwlock` and the `inode_lock` of the system inodes, blocking on
Process A.
This results in a classic AB-BA deadlock, leading to a Denial of Service (system
hang).
**Triggerability:**
For `kswapd0` to call `ocfs2_delete_inode()`, it must find a dentry on the LRU
whose inode has `i_nlink == 0` (or `OCFS2_INODE_MAYBE_ORPHANED`).
Under normal filesystem operations, when a file is unlinked, its dentry is
immediately unhashed by the VFS (`d_delete()`). Unhashed dentries are never
placed on the LRU, meaning `kswapd0` will never encounter them. Similarly, OCFS2
DLM downconvert workers unhash dentries when a file is deleted on another node.
The *only* way a hashed dentry with `i_nlink == 0` can end up on the LRU is
through a **malicious or corrupted filesystem image**. If an OCFS2 image is
crafted such that a directory entry points to an inode with `i_nlink == 0`, a
simple `stat` or `open` will instantiate a positive, hashed dentry. When the
reference is dropped, it goes to the LRU. Subsequent memory pressure will wake
up `kswapd0`, which will kill the dentry and trigger the deadlock.
**Security Boundaries:**
- **Unprivileged / User Namespaces:** OCFS2 does not support `FS_USERNS_MOUNT`.
Therefore, an unprivileged user cannot mount the malicious image, even within a
user namespace. It requires `CAP_SYS_ADMIN` in the initial namespace.
- **Peripheral / Filesystem Trigger:** The bug can be triggered if a malicious
OCFS2 image is introduced via a physical peripheral (e.g., a USB drive) and is
auto-mounted by a desktop environment. Once mounted, any unprivileged access
(like an auto-indexer reading the drive) combined with memory pressure will
trigger the deadlock.
**Conclusion:**
This is a Denial of Service vulnerability triggered by mounting and interacting
with a maliciously crafted OCFS2 filesystem image. It is not exploitable for
privilege escalation or information leaks. Exploitable:false FilesystemTrigger:true NetworkTrigger:false PeripheralTrigger:true RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
|
| 1/1 |
2026/09/17 00:00 |
action |
syz-repro-to-c-repro |
0m
Results:
|
| 2/1 |
2026/09/17 00:00 |
action |
kernel-checkouter |
1m
Results: map[KernelSrc:/app/workdir/cache/src/2176ce130cfa6f687d5368ee3cb26c4f777c554a]
|
| 3/1 |
2026/09/17 00:01 |
action |
kernel-builder |
25m
Results: map[KernelObj:/app/workdir/cache/build/af5e663109bbae76f57dd1aef29e7fe755ed6cc2]
|
| 4/1 |
2026/09/17 00:26 |
action |
codesearch-prepare |
18m
Results: map[Index:codesearch-index]
|
| 5/1 |
2026/09/17 00:45 |
agent |
expert |
4m
Model:
Results: map[DenialOfService:true Exploitable:false FilesystemTrigger:true NetworkTrigger:false PeripheralTrigger:true RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
Instruction: You are an experienced Linux kernel security engineer. Your task is to analyze given kernel bug report
and determine its security impact based on the following dimensions.
Use the provided tools to examine the source code, check for capability checks (e.g., capable(), ns_capable()),
and understand the nature of the bug. Analyze the given kernel build and configuration.
You can check the kernel config by grepping ".config" file; you can check kernel cmdline by grepping
".config" file for "CONFIG_CMDLINE=". Assume sysctl parameters have default values.
But analyze for the corresponding production build w/o debugging tools enabled (like KASAN, KMSAN, UBSAN).
Try different strategies when analyzing the bug:
- think of ways in which the vulnerable code is unreachable
- or the other way around: try to come up with different ideas of how an unprivileged user can reach the bug
If still unsure err on the side of the bug being non-exploitable/not-accessible.
In the final reply, provide a reasoning for your assessment.
Analysis dimensions:
* Exploitable:
Determine if the bug can result in memory corruption, elevated privileges, or an information leak.
Memory safety issues are almost always exploitable (KASAN or UBSAN reports for use-after-free, out-of-bounds;
refcounting issues, corrupted lists, etc). When kernel is crashing on a completely wild pointer access
(e.g. user-space address, or non-canonical address, but not on NULL or address corresponding to KASAN shadow
for NULL address), including both data accesses and control transfers, that also usually implies possibility
of exploitation. Such reports usually say "unable to handle kernel paging request".
Uses of uninitialized values detected by KMSAN may be exploitable b/c attacker frequently can affect uninit
values with spraying techniques. However, for these exploitability depends on how exactly the uninit value
is used in the code, and what it affects.
Information leaks are exploitable on their own and should be classified as such. A bug that copies kernel
memory contents to userspace (e.g. an out-of-bounds read whose result is returned to the caller, or
uninitialized stack/heap bytes written to a user buffer) is exploitable: it can reveal kernel pointer
values and defeat KASLR, expose sensitive data such as cryptographic keys or other processes' memory, and
serves as a necessary building block in most modern kernel privilege-escalation exploit chains. Do not classify
an information leak as non-exploitable solely because it does not directly cause a memory write or control-flow
hijack; the leak itself is the exploit primitive.
Think of what happens after the bug is triggered. Some bugs cause kernel panic and halt execution,
they are harder to exploit. For example, BUG reports halts the kernel. However, WARNING reports don't halt
execution in production builds. Debug bug detection tools (like KASAN, KMSAN, KCSAN, UBSAN) are also not enabled
in production builds, so attacker can freely exploit these bugs w/o being detected by these tools.
If you see an integer overflow, think how the overflowed value used later (if it's used as allocation size,
or an array index). If you see an out-of-bounds read, think if it's followed by an out-of-bounds write as well.
Some KCSAN data-races may be exploitable by skilled attackers as well. Think what data structures got corrupted
as the result of data races and how. However, note that kernel has lots of "benign" data races that don't lead
to any runtime misbehavior at all.
* Denial Of Service:
Determine if the bug can result in denial-of-service. Most bugs can, since they cause system crash,
hangs, deadlocks, or resource leaks. This is mostly applicable to WARNING bugs that won't cause system crash
in production. For these think what will be consequences of the violation of the kernel assumptions flagged
by the WARNING. In some cases the unexpected condition is also properly handled by the normal control flow
(e.g. with "if (WARN_ON(...))"), these won't cause denial-of-service. If the condition is not handled,
then it may or may not cause denial-of-service.
* Accessible From Unprivileged Processes:
Determine if the bug can be reached from a typical (non-root) user process that does NOT have any special capabilities
(like CAP_SYS_ADMIN, CAP_NET_ADMIN, CAP_NET_RAW, CAP_PERFMON) or access to device nodes restricted to root.
Assume that unprivileged_bpf_disabled=1, that is eBPF loading is not accessible. However, cBPF (classical BPF)
is still accessible to non-root processes.
Assume that user namespaces are not accessible, that is, the process cannot get the mentioned capabilities even
within a new user namespace (checked by ns_capable() function in the kernel sources).
* Accessible From User Namespaces:
Determine if the bug can be reached within a user-namespace where the process has all capabilities
(including CAP_SYS_ADMIN, CAP_NET_ADMIN, CAP_NET_RAW, CAP_PERFMON). Such capabilities are checked with ns_capable()
function in the kernel sources.
* VM Guest Trigger:
Determine if the bug can be triggered from the context of a typical KVM guest (e.g., set up by a QEMU VMM).
Consider accesses to standard Linux host paravirtualized features (virtio-blk, virtio-net, etc.),
and handling of VM exits in the KVM code.
* VM Host Trigger in The Confidential Computing Context:
Determine if the bug can be triggered in a confidential computing guest kernel from the context of a KVM host.
Consider access to standard Linux guest paravirtualized features (virtio-blk, virtio-net, etc.).
* Ethernet Network Trigger:
Determine if the bug can be triggered by processing ingress network Ethernet traffic, either directly (network stack)
or via drivers exposed to network data.
* Other Remote Trigger:
Determine if the bug can be triggered by processing remote traffic other than Ethernet (Wifi, Bluetooth, NFC, etc).
* Peripheral Trigger:
Determine if the bug can be triggered via an untrusted peripheral device that can be physically plugged
into a system, such as a USB device or a niche hardware driver handling external hardware inputs.
This is particularly important for mobile and desktop environments where users can plug in unknown devices.
* Malicious Filesystem Trigger:
Determine if the bug can be triggered by the kernel mounting and parsing a malicious filesystem image.
This is highly critical for Desktop and Mobile environments where external media or downloaded images
might be auto-mounted.
Don't make assumptions about the kernel source code (it may be different from what you assume it is).
Extensively use the provided code access tools (codesearch-*, git-*, grepper, etc)
to examine the actual source code, and confirm any assumptions.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt:
The kernel bug report is:
======================================================
WARNING: possible circular locking dependency detected
syzkaller #0 Not tainted
------------------------------------------------------
kswapd0/70 is trying to acquire lock:
ffff88800e994bc0 (&osb->nfs_sync_rwlock){.+.+}-{4:4}, at: ocfs2_nfs_sync_lock+0x106/0x270 fs/ocfs2/dlmglue.c:2875
but task is already holding lock:
ffffffff8ee92540 (fs_reclaim){+.+.}-{0:0}, at: balance_pgdat mm/vmscan.c:7177 [inline]
ffffffff8ee92540 (fs_reclaim){+.+.}-{0:0}, at: kswapd+0xa08/0x32c0 mm/vmscan.c:7555
which lock already depends on the new lock.
the existing dependency chain (in reverse order) is:
-> #3 (fs_reclaim){+.+.}-{0:0}:
__fs_reclaim_acquire mm/page_alloc.c:4375 [inline]
fs_reclaim_acquire+0x71/0x100 mm/page_alloc.c:4389
might_alloc include/linux/sched/mm.h:316 [inline]
slab_pre_alloc_hook mm/slub.c:4636 [inline]
slab_alloc_node mm/slub.c:4974 [inline]
__kmalloc_cache_noprof+0x61/0x600 mm/slub.c:5559
_kmalloc_noprof include/linux/slab.h:991 [inline]
_kzalloc_noprof include/linux/slab.h:1312 [inline]
ocfs2_reserve_new_metadata_blocks+0x10c/0x9a0 fs/ocfs2/suballoc.c:1084
ocfs2_mknod+0xee4/0x22b0 fs/ocfs2/namei.c:356
ocfs2_mkdir+0x180/0x430 fs/ocfs2/namei.c:667
vfs_mkdir+0x40c/0x620 fs/namei.c:5410
filename_mkdirat+0x285/0x510 fs/namei.c:5443
__do_sys_mkdirat fs/namei.c:5464 [inline]
__se_sys_mkdirat+0x35/0x150 fs/namei.c:5461
do_syscall_x64 arch/x86/entry/syscall_64.c:61 [inline]
do_syscall_64+0x166/0x520 arch/x86/entry/syscall_64.c:84
entry_SYSCALL_64_after_hwframe+0x77/0x7f
-> #2 (&ocfs2_sysfile_lock_key[INODE_ALLOC_SYSTEM_INODE]){+.+.}-{4:4}:
down_write+0x96/0x200 kernel/locking/rwsem.c:1631
inode_lock include/linux/fs.h:1024 [inline]
ocfs2_remove_inode fs/ocfs2/inode.c:767 [inline]
ocfs2_wipe_inode fs/ocfs2/inode.c:930 [inline]
ocfs2_delete_inode fs/ocfs2/inode.c:1191 [inline]
ocfs2_evict_inode+0x1489/0x4390 fs/ocfs2/inode.c:1333
evict+0x624/0xb50 fs/inode.c:822
ocfs2_dentry_iput+0x24d/0x390 fs/ocfs2/dcache.c:407
dentry_kill+0x1b9/0x880 fs/dcache.c:826
finish_dput+0x1a/0x260 fs/dcache.c:1001
ovl_check_rename_whiteout fs/overlayfs/super.c:617 [inline]
ovl_make_workdir fs/overlayfs/super.c:713 [inline]
ovl_get_workdir+0xba2/0x1980 fs/overlayfs/super.c:836
ovl_fill_super_creds fs/overlayfs/super.c:1449 [inline]
ovl_fill_super+0x1702/0x4360 fs/overlayfs/super.c:1561
vfs_get_super fs/super.c:1391 [inline]
get_tree_nodev+0xbb/0x150 fs/super.c:1410
vfs_get_tree+0x92/0x2a0 fs/super.c:1933
fc_mount fs/namespace.c:1198 [inline]
do_new_mount_fc fs/namespace.c:3772 [inline]
do_new_mount+0x319/0xdc0 fs/namespace.c:3848
do_mount fs/namespace.c:4181 [inline]
__do_sys_mount fs/namespace.c:4397 [inline]
__se_sys_mount+0x31d/0x420 fs/namespace.c:4374
do_syscall_x64 arch/x86/entry/syscall_64.c:61 [inline]
do_syscall_64+0x166/0x520 arch/x86/entry/syscall_64.c:84
entry_SYSCALL_64_after_hwframe+0x77/0x7f
-> #1 (&ocfs2_sysfile_lock_key[ORPHAN_DIR_SYSTEM_INODE]){+.+.}-{4:4}:
down_write+0x96/0x200 kernel/locking/rwsem.c:1631
inode_lock include/linux/fs.h:1024 [inline]
ocfs2_wipe_inode fs/ocfs2/inode.c:888 [inline]
ocfs2_delete_inode fs/ocfs2/inode.c:1191 [inline]
ocfs2_evict_inode+0xe39/0x4390 fs/ocfs2/inode.c:1333
evict+0x624/0xb50 fs/inode.c:822
ocfs2_dentry_iput+0x24d/0x390 fs/ocfs2/dcache.c:407
dentry_kill+0x1b9/0x880 fs/dcache.c:826
finish_dput+0x1a/0x260 fs/dcache.c:1001
ovl_check_rename_whiteout fs/overlayfs/super.c:617 [inline]
ovl_make_workdir fs/overlayfs/super.c:713 [inline]
ovl_get_workdir+0xba2/0x1980 fs/overlayfs/super.c:836
ovl_fill_super_creds fs/overlayfs/super.c:1449 [inline]
ovl_fill_super+0x1702/0x4360 fs/overlayfs/super.c:1561
vfs_get_super fs/super.c:1391 [inline]
get_tree_nodev+0xbb/0x150 fs/super.c:1410
vfs_get_tree+0x92/0x2a0 fs/super.c:1933
fc_mount fs/namespace.c:1198 [inline]
do_new_mount_fc fs/namespace.c:3772 [inline]
do_new_mount+0x319/0xdc0 fs/namespace.c:3848
do_mount fs/namespace.c:4181 [inline]
__do_sys_mount fs/namespace.c:4397 [inline]
__se_sys_mount+0x31d/0x420 fs/namespace.c:4374
do_syscall_x64 arch/x86/entry/syscall_64.c:61 [inline]
do_syscall_64+0x166/0x520 arch/x86/entry/syscall_64.c:84
entry_SYSCALL_64_after_hwframe+0x77/0x7f
-> #0 (&osb->nfs_sync_rwlock){.+.+}-{4:4}:
check_prev_add kernel/locking/lockdep.c:3209 [inline]
check_prevs_add kernel/locking/lockdep.c:3328 [inline]
validate_chain kernel/locking/lockdep.c:3952 [inline]
__lock_acquire+0x164c/0x2de0 kernel/locking/lockdep.c:5288
lock_acquire+0x115/0x350 kernel/locking/lockdep.c:5942
down_read+0x4a/0x330 kernel/locking/rwsem.c:1574
ocfs2_nfs_sync_lock+0x106/0x270 fs/ocfs2/dlmglue.c:2875
ocfs2_delete_inode fs/ocfs2/inode.c:1142 [inline]
ocfs2_evict_inode+0x3a6/0x4390 fs/ocfs2/inode.c:1333
evict+0x624/0xb50 fs/inode.c:822
ocfs2_dentry_iput+0x24d/0x390 fs/ocfs2/dcache.c:407
dentry_kill+0x1b9/0x880 fs/dcache.c:826
shrink_kill fs/dcache.c:1275 [inline]
shrink_dentry_list+0x209/0x310 fs/dcache.c:1292
prune_dcache_sb+0x119/0x180 fs/dcache.c:1374
super_cache_scan+0x35f/0x4a0 fs/super.c:225
do_shrink_slab+0x6d0/0x10e0 mm/shrinker.c:445
shrink_slab_memcg mm/shrinker.c:562 [inline]
shrink_slab+0x838/0x1220 mm/shrinker.c:640
shrink_one+0x2a2/0x730 mm/vmscan.c:5119
shrink_many mm/vmscan.c:5180 [inline]
lru_gen_shrink_node mm/vmscan.c:5258 [inline]
shrink_node+0x35d0/0x3e30 mm/vmscan.c:6247
kswapd_shrink_node mm/vmscan.c:7101 [inline]
balance_pgdat mm/vmscan.c:7279 [inline]
kswapd+0x185e/0x32c0 mm/vmscan.c:7555
kthread+0x38b/0x480 kernel/kthread.c:436
ret_from_fork+0x514/0xb70 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
other info that might help us debug this:
Chain exists of:
&osb->nfs_sync_rwlock --> &ocfs2_sysfile_lock_key[INODE_ALLOC_SYSTEM_INODE] --> fs_reclaim
Possible unsafe locking scenario:
CPU0 CPU1
---- ----
lock(fs_reclaim);
lock(&ocfs2_sysfile_lock_key[INODE_ALLOC_SYSTEM_INODE]);
lock(fs_reclaim);
rlock(&osb->nfs_sync_rwlock);
*** DEADLOCK ***
locks held by kswapd0/70: 2, last CPU#0:
#0: ffffffff8ee92540 (fs_reclaim){+.+.}-{0:0}, at: balance_pgdat mm/vmscan.c:7177 [inline]
#0: ffffffff8ee92540 (fs_reclaim){+.+.}-{0:0}, at: kswapd+0xa08/0x32c0 mm/vmscan.c:7555
#1: ffff888012c380e8 (&type->s_umount_key#55){.+.+}-{4:4}, at: super_trylock_shared fs/super.c:674 [inline]
#1: ffff888012c380e8 (&type->s_umount_key#55){.+.+}-{4:4}, at: super_cache_scan+0x94/0x4a0 fs/super.c:200
stack backtrace:
CPU: 0 UID: 0 PID: 70 Comm: kswapd0 Not tainted syzkaller #0 PREEMPT(full)
Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
Call Trace:
<TASK>
dump_stack_lvl+0xe8/0x150 lib/dump_stack.c:120
print_circular_bug+0x2e2/0x300 kernel/locking/lockdep.c:2087
check_noncircular+0x12f/0x150 kernel/locking/lockdep.c:2219
check_prev_add kernel/locking/lockdep.c:3209 [inline]
check_prevs_add kernel/locking/lockdep.c:3328 [inline]
validate_chain kernel/locking/lockdep.c:3952 [inline]
__lock_acquire+0x164c/0x2de0 kernel/locking/lockdep.c:5288
lock_acquire+0x115/0x350 kernel/locking/lockdep.c:5942
down_read+0x4a/0x330 kernel/locking/rwsem.c:1574
ocfs2_nfs_sync_lock+0x106/0x270 fs/ocfs2/dlmglue.c:2875
ocfs2_delete_inode fs/ocfs2/inode.c:1142 [inline]
ocfs2_evict_inode+0x3a6/0x4390 fs/ocfs2/inode.c:1333
evict+0x624/0xb50 fs/inode.c:822
ocfs2_dentry_iput+0x24d/0x390 fs/ocfs2/dcache.c:407
dentry_kill+0x1b9/0x880 fs/dcache.c:826
shrink_kill fs/dcache.c:1275 [inline]
shrink_dentry_list+0x209/0x310 fs/dcache.c:1292
prune_dcache_sb+0x119/0x180 fs/dcache.c:1374
super_cache_scan+0x35f/0x4a0 fs/super.c:225
do_shrink_slab+0x6d0/0x10e0 mm/shrinker.c:445
shrink_slab_memcg mm/shrinker.c:562 [inline]
shrink_slab+0x838/0x1220 mm/shrinker.c:640
shrink_one+0x2a2/0x730 mm/vmscan.c:5119
shrink_many mm/vmscan.c:5180 [inline]
lru_gen_shrink_node mm/vmscan.c:5258 [inline]
shrink_node+0x35d0/0x3e30 mm/vmscan.c:6247
kswapd_shrink_node mm/vmscan.c:7101 [inline]
balance_pgdat mm/vmscan.c:7279 [inline]
kswapd+0x185e/0x32c0 mm/vmscan.c:7555
kthread+0x38b/0x480 kernel/kthread.c:436
ret_from_fork+0x514/0xb70 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
</TASK>
Reply: Based on the analysis of the bug report and the kernel source code, here is the assessment of the vulnerability:
**Bug Mechanism and Impact:**
The lockdep warning indicates a circular locking dependency (deadlock) involving the memory reclaim context (`fs_reclaim`), OCFS2's `nfs_sync_rwlock`, and the `inode_lock` of OCFS2 system inodes (like `INODE_ALLOC_SYSTEM_INODE`).
The deadlock occurs due to two concurrent paths:
1. **Process A (e.g., `mkdir`)**: Holds the `inode_lock` of an OCFS2 system inode (e.g., during `ocfs2_reserve_new_metadata_blocks`) and allocates memory using `GFP_KERNEL` (via `kzalloc_obj`). This triggers direct memory reclaim, causing the process to wait on the `fs_reclaim` pseudo-lock.
2. **Process B (`kswapd0` or another direct reclaimer)**: Holds the `fs_reclaim` lock while shrinking the dcache. It finds a dentry on the LRU, kills it, and drops the last reference to its inode. If the inode has `i_nlink == 0`, `evict()` synchronously calls `ocfs2_delete_inode()`. This function attempts to acquire `nfs_sync_rwlock` and the `inode_lock` of the system inodes, blocking on Process A.
This results in a classic AB-BA deadlock, leading to a Denial of Service (system hang).
**Triggerability:**
For `kswapd0` to call `ocfs2_delete_inode()`, it must find a dentry on the LRU whose inode has `i_nlink == 0` (or `OCFS2_INODE_MAYBE_ORPHANED`).
Under normal filesystem operations, when a file is unlinked, its dentry is immediately unhashed by the VFS (`d_delete()`). Unhashed dentries are never placed on the LRU, meaning `kswapd0` will never encounter them. Similarly, OCFS2 DLM downconvert workers unhash dentries when a file is deleted on another node.
The *only* way a hashed dentry with `i_nlink == 0` can end up on the LRU is through a **malicious or corrupted filesystem image**. If an OCFS2 image is crafted such that a directory entry points to an inode with `i_nlink == 0`, a simple `stat` or `open` will instantiate a positive, hashed dentry. When the reference is dropped, it goes to the LRU. Subsequent memory pressure will wake up `kswapd0`, which will kill the dentry and trigger the deadlock.
**Security Boundaries:**
- **Unprivileged / User Namespaces:** OCFS2 does not support `FS_USERNS_MOUNT`. Therefore, an unprivileged user cannot mount the malicious image, even within a user namespace. It requires `CAP_SYS_ADMIN` in the initial namespace.
- **Peripheral / Filesystem Trigger:** The bug can be triggered if a malicious OCFS2 image is introduced via a physical peripheral (e.g., a USB drive) and is auto-mounted by a desktop environment. Once mounted, any unprivileged access (like an auto-indexer reading the drive) combined with memory pressure will trigger the deadlock.
**Conclusion:**
This is a Denial of Service vulnerability triggered by mounting and interacting with a maliciously crafted OCFS2 filesystem image. It is not exploitable for privilege escalation or information leaks.
|
| 6/2 |
2026/09/17 00:45 |
llm |
expert |
0m
Model:
Tokens:
input: 10629
output: 33
thoughts: 1830
Thoughts: **Analyzing a Circular Locking Dependency in OCFS2**
Okay, here's what's happening. I've got a lockdep warning, clearly indicating a potential circular locking dependency. It's flagged in `fs/ocfs2/dlmglue.c` and `mm/vmscan.c`, which immediately suggests file system and memory management interactions. The warning points to a circular dependency involving `fs_reclaim`, `&ocfs2_sysfile_lock_key[INODE_ALLOC_SYSTEM_INODE]`, and `&osb->nfs_sync_rwlock`.
The reported chain is: `&osb->nfs_sync_rwlock` --> `&ocfs2_sysfile_lock_key[INODE_ALLOC_SYSTEM_INODE]` --> `fs_reclaim`. However, the tracing shows an additional potential chain: `fs_reclaim` -> `&osb->nfs_sync_rwlock`, specifically when kswapd0 is involved. This latter chain completes the cycle.
Tracing kswapd0's stack reveals the issue: While kswapd0 holds `fs_reclaim` (because it's reclaiming memory), it goes through `shrink_node`, `shrink_slab`, `super_cache_scan`, and so on, eventually calling `ocfs2_nfs_sync_lock` and attempting to acquire `&osb->nfs_sync_rwlock`. So, we have `fs_reclaim` -> `&osb->nfs_sync_rwlock`.
Now, looking at the existing dependency chains, we see:
* Chain 3 (`&ocfs2_sysfile_lock_key[INODE_ALLOC_SYSTEM_INODE]` -> `fs_reclaim`): `ocfs2_mknod` calls `_kzalloc_noprof` which in turn triggers `fs_reclaim`. This seems relevant.
* Chain 2 (`&ocfs2_sysfile_lock_key[ORPHAN_DIR_SYSTEM_INODE]` -> `&ocfs2_sysfile_lock_key[INODE_ALLOC_SYSTEM_INODE]`): This chain involves lock acquisition in the `ocfs2_remove_inode` sequence, starting with the orphan dir, and ending with inode allocation system lock.
* Chain 1 (`&osb->nfs_sync_rwlock` -> `&ocfs2_sysfile_lock_key[ORPHAN_DIR_SYSTEM_INODE]`): This chain also involves lock acquisition, starting with nfs sync lock and ending in the orphan dir system lock.
* Chain 0 (the circular one: `fs_reclaim` -> `&osb->nfs_sync_rwlock`): This chain originates with kswapd and the lock it must take.
The core of the problem lies in `ocfs2_evict_inode` taking `nfs_sync_rwlock` (via `ocfs2_nfs_sync_lock`) and then calling `ocfs2_delete_inode`, which grabs `inode_lock` on the orphan directory. Then, `ocfs2_remove_inode` grabs another `inode_lock` on the system inode, and `ocfs2_reserve_new_metadata_blocks` triggers a memory allocation, potentially calling `fs_reclaim`.
So, the dependency chain becomes: `nfs_sync_rwlock` -> `inode_lock(ORPHAN_DIR)` -> `inode_lock(INODE_ALLOC)` -> `fs_reclaim`. And kswapd0 holds `fs_reclaim` and then tries to acquire `nfs_sync_rwlock`, leading to a deadlock.
Essentially, kswapd0, while reclaiming memory (holding `fs_reclaim`), can evict OCFS2 inodes. Evicting an OCFS2 inode requires `nfs_sync_rwlock`. If another process holds `nfs_sync_rwlock` and then allocates memory, which in turn waits for `fs_reclaim`, we have a classic memory reclaim deadlock. Specifically, the process holding the nfs sync lock would then acquire system inode locks and try to allocate more memory, thus trying to reclaim memory. This would deadlock kswapd0.
The key to exploitability is whether `ocfs2_mknod` or similar operations use `GFP_NOFS` when allocating memory within a critical section. If not, a user could potentially trigger this by creating an OCFS2 filesystem, creating sufficient memory pressure (to wake up kswapd0), and then running an operation within a user program (holding `nfs_sync_rwlock` and not `
|
| 7/2 |
2026/09/17 00:45 |
tool |
grepper |
0m
Error:
Args: map[Expression:FS_USERNS_MOUNT PathPrefix:fs/ocfs2/]
Results:
|
| 8/2 |
2026/09/17 00:45 |
llm |
expert |
0m
Model:
Tokens:
input: 13531
output: 41
thoughts: 267
Thoughts: **My Assessment of OCFS2 and Unprivileged User Interaction**
Okay, so I'm thinking about OCFS2 and how it handles unprivileged users. The documentation is pretty clear; it doesn't support `FS_USERNS_MOUNT`, which is a problem for us. That means an ordinary user can't just spin up an OCFS2 filesystem within their user namespace. But, if the system admin has already mounted an OCFS2 volume, what *can* an unprivileged user do with it?
Can they create files? Sure, as long as they have write permission to a directory within that mount. Memory pressure? Definitely. Allocation is the name of the game.
Now, the more interesting question: can they trigger that dreaded deadlock we've been talking about? Let's trace this out. Say they run `mkdir` on the OCFS2 filesystem. That'll call `ocfs2_mkdir`, which, in turn, calls `ocfs2_mknod`. And `ocfs2_mknod`, I believe, will trigger allocation. So far, so good. Then, *simultaneously* we hit that memory pressure. If `kswapd0` gets involved and tries to reclaim some memory, it might go after an OCFS2 inode. Hmm, that's where the potential for deadlock lies.
But wait a minute, is that actually going to happen? Is that deadlock even *real*? We need to look at how `ocfs2_mknod` handles its memory allocation. Does it set up a transaction with `GFP_NOFS`? Filesystems typically use `memalloc_nofs_save()` or something similar to prevent `fs_reclaim` within the transaction context. That's the key. I need to dive into the code for `ocfs2_mknod` to confirm this. Let's see...
|
| 9/2 |
2026/09/17 00:45 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/ocfs2/namei.c Name:ocfs2_mknod]
Results: map[SourceCode: 230: static int ocfs2_mknod(struct mnt_idmap *idmap,
231: struct inode *dir,
232: struct dentry *dentry,
233: umode_t mode,
234: dev_t dev)
235: {
236: int status = 0;
237: struct buffer_head *parent_fe_bh = NULL;
238: handle_t *handle = NULL;
239: struct ocfs2_super *osb;
240: struct ocfs2_dinode *dirfe;
241: struct ocfs2_dinode *fe = NULL;
242: struct buffer_head *new_fe_bh = NULL;
243: struct inode *inode = NULL;
244: struct ocfs2_alloc_context *inode_ac = NULL;
245: struct ocfs2_alloc_context *data_ac = NULL;
246: struct ocfs2_alloc_context *meta_ac = NULL;
247: int want_clusters = 0;
248: int want_meta = 0;
249: int xattr_credits = 0;
250: struct ocfs2_security_xattr_info si = {
251: .name = NULL,
252: .enable = 1,
253: };
254: int did_quota_inode = 0;
255: struct ocfs2_dir_lookup_result lookup = { NULL, };
256: sigset_t oldset;
257: int did_block_signals = 0;
258: struct ocfs2_dentry_lock *dl = NULL;
259: struct ocfs2_acl_state acl_state = { 0 };
260:
261: trace_ocfs2_mknod(dir, dentry, dentry->d_name.len, dentry->d_name.name,
262: (unsigned long long)OCFS2_I(dir)->ip_blkno,
263: (unsigned long)dev, mode);
264:
265: status = dquot_initialize(dir);
266: if (status) {
267: mlog_errno(status);
268: return status;
269: }
270:
271: /* get our super block */
272: osb = OCFS2_SB(dir->i_sb);
273:
274: status = ocfs2_inode_lock(dir, &parent_fe_bh, 1);
275: if (status < 0) {
276: if (status != -ENOENT)
277: mlog_errno(status);
278: return status;
279: }
280:
281: if (S_ISDIR(mode) && (dir->i_nlink >= ocfs2_link_max(osb))) {
282: status = -EMLINK;
283: goto leave;
284: }
285:
286: dirfe = (struct ocfs2_dinode *) parent_fe_bh->b_data;
287: if (!ocfs2_read_links_count(dirfe)) {
288: /* can't make a file in a deleted directory. */
289: status = -ENOENT;
290: goto leave;
291: }
292:
293: status = ocfs2_check_dir_for_entry(dir, dentry->d_name.name,
294: dentry->d_name.len);
295: if (status)
296: goto leave;
297:
298: /* get a spot inside the dir. */
299: status = ocfs2_prepare_dir_for_insert(osb, dir, parent_fe_bh,
300: dentry->d_name.name,
301: dentry->d_name.len, &lookup);
302: if (status < 0) {
303: mlog_errno(status);
304: goto leave;
305: }
306:
307: /* reserve an inode spot */
308: status = ocfs2_reserve_new_inode(osb, &inode_ac);
309: if (status < 0) {
310: if (status != -ENOSPC)
311: mlog_errno(status);
312: goto leave;
313: }
314:
315: inode = ocfs2_get_init_inode(dir, mode);
316: if (IS_ERR(inode)) {
317: status = PTR_ERR(inode);
318: inode = NULL;
319: mlog_errno(status);
320: goto leave;
321: }
322:
323: /* get security xattr */
324: status = ocfs2_init_security_get(inode, dir, &dentry->d_name, &si);
325: if (status) {
326: if (status == -EOPNOTSUPP)
327: si.enable = 0;
328: else {
329: mlog_errno(status);
330: goto leave;
331: }
332: }
333:
334: status = ocfs2_acl_init_prepare(inode, dir, parent_fe_bh, &acl_state);
335: if (status < 0)
336: goto leave;
337:
338: /* calculate meta data/clusters for setting security and acl xattr */
339: status = ocfs2_calc_xattr_init(dir, mode, &si, &want_clusters,
340: &xattr_credits, &want_meta,
341: &acl_state);
342: if (status < 0) {
343: mlog_errno(status);
344: goto leave;
345: }
346:
347: /* Reserve a cluster if creating an extent based directory. */
348: if (S_ISDIR(mode) && !ocfs2_supports_inline_data(osb)) {
349: want_clusters += 1;
350:
351: /* Dir indexing requires extra space as well */
352: if (ocfs2_supports_indexed_dirs(osb))
353: want_meta++;
354: }
355:
356: status = ocfs2_reserve_new_metadata_blocks(osb, want_meta, &meta_ac);
357: if (status < 0) {
358: if (status != -ENOSPC)
359: mlog_errno(status);
360: goto leave;
361: }
362:
363: status = ocfs2_reserve_clusters(osb, want_clusters, &data_ac);
364: if (status < 0) {
365: if (status != -ENOSPC)
366: mlog_errno(status);
367: goto leave;
368: }
369:
370: handle = ocfs2_start_trans(osb, ocfs2_mknod_credits(osb->sb,
371: S_ISDIR(mode),
372: xattr_credits));
373: if (IS_ERR(handle)) {
374: status = PTR_ERR(handle);
375: handle = NULL;
376: mlog_errno(status);
377: goto leave;
378: }
379:
380: /* Starting to change things, restart is no longer possible. */
381: ocfs2_block_signals(&oldset);
382: did_block_signals = 1;
383:
384: status = dquot_alloc_inode(inode);
385: if (status)
386: goto leave;
387: did_quota_inode = 1;
388:
389: /* do the real work now. */
390: status = ocfs2_mknod_locked(osb, dir, inode, dev,
391: &new_fe_bh, parent_fe_bh, handle,
392: inode_ac);
393: if (status < 0) {
394: mlog_errno(status);
395: goto leave;
396: }
397:
398: fe = (struct ocfs2_dinode *) new_fe_bh->b_data;
399: if (S_ISDIR(mode)) {
400: status = ocfs2_fill_new_dir(osb, handle, dir, inode,
401: new_fe_bh, data_ac, meta_ac);
402: if (status < 0) {
403: mlog_errno(status);
404: goto leave;
405: }
406:
407: status = ocfs2_journal_access_di(handle, INODE_CACHE(dir),
408: parent_fe_bh,
409: OCFS2_JOURNAL_ACCESS_WRITE);
410: if (status < 0) {
411: mlog_errno(status);
412: goto leave;
413: }
414: ocfs2_add_links_count(dirfe, 1);
415: ocfs2_journal_dirty(handle, parent_fe_bh);
416: inc_nlink(dir);
417: }
418:
419: status = ocfs2_init_acl(handle, inode, new_fe_bh, meta_ac, data_ac,
420: &acl_state);
421:
422: if (status < 0) {
423: mlog_errno(status);
424: goto roll_back;
425: }
426:
427: if (si.enable) {
428: status = ocfs2_init_security_set(handle, inode, new_fe_bh, &si,
429: meta_ac, data_ac);
430: if (status < 0) {
431: mlog_errno(status);
432: goto roll_back;
433: }
434: }
435:
436: /*
437: * Do this before adding the entry to the directory. We add
438: * also set d_op after success so that ->d_iput() will cleanup
439: * the dentry lock even if ocfs2_add_entry() fails below.
440: */
441: status = ocfs2_dentry_attach_lock(dentry, inode,
442: OCFS2_I(dir)->ip_blkno);
443: if (status) {
444: mlog_errno(status);
445: goto roll_back;
446: }
447:
448: dl = dentry->d_fsdata;
449:
450: status = ocfs2_add_entry(handle, dentry, inode,
451: OCFS2_I(inode)->ip_blkno, parent_fe_bh,
452: &lookup);
453: if (status < 0) {
454: mlog_errno(status);
455: goto roll_back;
456: }
457:
458: insert_inode_hash(inode);
459: d_instantiate(dentry, inode);
460: status = 0;
461:
462: roll_back:
463: if (status < 0 && S_ISDIR(mode)) {
464: ocfs2_add_links_count(dirfe, -1);
465: drop_nlink(dir);
466: }
467:
468: leave:
469: if (status < 0 && did_quota_inode)
470: dquot_free_inode(inode);
471: if (handle) {
472: if (status < 0 && fe)
473: ocfs2_set_links_count(fe, 0);
474: ocfs2_commit_trans(osb, handle);
475: }
476:
477: ocfs2_inode_unlock(dir, 1);
478: if (did_block_signals)
479: ocfs2_unblock_signals(&oldset);
480:
481: brelse(new_fe_bh);
482: brelse(parent_fe_bh);
483: kfree(si.value);
484:
485: ocfs2_acl_init_release(&acl_state);
486:
487: ocfs2_free_dir_lookup_result(&lookup);
488:
489: if (inode_ac)
490: ocfs2_free_alloc_context(inode_ac);
491:
492: if (data_ac)
493: ocfs2_free_alloc_context(data_ac);
494:
495: if (meta_ac)
496: ocfs2_free_alloc_context(meta_ac);
497:
498: /*
499: * We should call iput after the i_rwsem of the bitmap been
500: * unlocked in ocfs2_free_alloc_context, or the
501: * ocfs2_delete_inode will mutex_lock again.
502: */
503: if ((status < 0) && inode) {
504: if (dl)
505: ocfs2_cleanup_add_entry_failure(osb, dentry, inode);
506:
507: OCFS2_I(inode)->ip_flags |= OCFS2_INODE_SKIP_ORPHAN_DIR;
508: clear_nlink(inode);
509: iput(inode);
510: }
511:
512: if (status)
513: mlog_errno(status);
514:
515: return status;
516: }
SourceFile:fs/ocfs2/namei.c]
|
| 10/2 |
2026/09/17 00:45 |
llm |
expert |
0m
Model:
Tokens:
input: 18517
output: 45
thoughts: 157
Thoughts: **My Analysis of `ocfs2_mknod`'s Metadata Reservation**
Okay, so I'm digging into `ocfs2_mknod` here. The first thing that catches my eye is the call to `ocfs2_reserve_new_metadata_blocks(osb, want_meta, &meta_ac)`. Interesting, that's happening *before* the transaction starts with `ocfs2_start_trans(osb, ...)`. So, we're not inside the transactional context yet. But hold on... isn't this function holding `ocfs2_inode_lock(dir, &parent_fe_bh, 1)`? That's a cluster-wide lock on the parent directory inode!
Now, the plot thickens. My initial assumption was wrong, `ocfs2_reserve_new_metadata_blocks` takes the `inode_lock` on the system inode when allocating metadata. This is going to require me to analyze `ocfs2_reserve_new_metadata_blocks` more thoroughly to understand the implications of this lock contention for potential deadlocks or performance bottlenecks.
|
| 11/2 |
2026/09/17 00:45 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/ocfs2/suballoc.c Name:ocfs2_reserve_new_metadata_blocks]
Results: map[SourceCode:1077: int ocfs2_reserve_new_metadata_blocks(struct ocfs2_super *osb,
1078: int blocks,
1079: struct ocfs2_alloc_context **ac)
1080: {
1081: int status;
1082: int slot = ocfs2_get_meta_steal_slot(osb);
1083:
1084: *ac = kzalloc_obj(struct ocfs2_alloc_context);
1085: if (!(*ac)) {
1086: status = -ENOMEM;
1087: mlog_errno(status);
1088: goto bail;
1089: }
1090:
1091: (*ac)->ac_bits_wanted = blocks;
1092: (*ac)->ac_which = OCFS2_AC_USE_META;
1093: (*ac)->ac_group_search = ocfs2_block_group_search;
1094:
1095: if (slot != OCFS2_INVALID_SLOT &&
1096: atomic_read(&osb->s_num_meta_stolen) < OCFS2_MAX_TO_STEAL)
1097: goto extent_steal;
1098:
1099: atomic_set(&osb->s_num_meta_stolen, 0);
1100: status = ocfs2_reserve_suballoc_bits(osb, (*ac),
1101: EXTENT_ALLOC_SYSTEM_INODE,
1102: (u32)osb->slot_num, NULL,
1103: ALLOC_GROUPS_FROM_GLOBAL|ALLOC_NEW_GROUP);
1104:
1105:
1106: if (status >= 0) {
1107: status = 0;
1108: if (slot != OCFS2_INVALID_SLOT)
1109: ocfs2_init_meta_steal_slot(osb);
1110: goto bail;
1111: } else if (status < 0 && status != -ENOSPC) {
1112: mlog_errno(status);
1113: goto bail;
1114: }
1115:
1116: ocfs2_free_ac_resource(*ac);
1117:
1118: extent_steal:
1119: status = ocfs2_steal_meta(osb, *ac);
1120: atomic_inc(&osb->s_num_meta_stolen);
1121: if (status < 0) {
1122: if (status != -ENOSPC)
1123: mlog_errno(status);
1124: goto bail;
1125: }
1126:
1127: status = 0;
1128: bail:
1129: if ((status < 0) && *ac) {
1130: ocfs2_free_alloc_context(*ac);
1131: *ac = NULL;
1132: }
1133:
1134: if (status)
1135: mlog_errno(status);
1136: return status;
1137: }
SourceFile:fs/ocfs2/suballoc.c]
|
| 12/2 |
2026/09/17 00:45 |
llm |
expert |
0m
Model:
Tokens:
input: 19900
output: 26
thoughts: 133
Thoughts: **Debugging a Suspect Memory Allocation**
Okay, hold on a second... `kzalloc_obj` is actually a *macro*? That's not what I initially assumed. I need to understand what's *really* going on here. The stack trace is pointing directly to:
```
_kzalloc_noprof include/linux/slab.h:1312 [inline]
ocfs2_reserve_new_metadata_blocks+0x10c/0x9a0 fs/ocfs2/suballoc.c:1084
```
Which suggests it's calling something like `kzalloc(sizeof(struct ocfs2_alloc_context), GFP_KERNEL)`. But I can't be sure, so I need to go straight to the definition of that macro and verify. Let's dig into that macro's expansion to confirm my hunch and see what's really happening during this allocation.
|
| 13/2 |
2026/09/17 00:45 |
tool |
grepper |
0m
Args: map[Expression:#define kzalloc_obj PathPrefix:]
Results: map[Output:include/linux/slab.h=1073=void *kmalloc_nolock(size_t size, gfp_t gfp_flags, int node);
--
include/linux/slab.h-1154-/* All kzalloc aliases for kmalloc_(obj|objs|flex). */
include/linux/slab.h:1155:#define kzalloc_obj(P, ...) \
include/linux/slab.h-1156- __alloc_objs(kzalloc, default_gfp(__VA_ARGS__), typeof(P), 1)
include/linux/slab.h:1157:#define kzalloc_objs(P, COUNT, ...) \
include/linux/slab.h-1158- __alloc_objs(kzalloc, default_gfp(__VA_ARGS__), typeof(P), COUNT)
--
tools/include/linux/slab.h=200=static inline unsigned int kmem_cache_sheaf_size(struct slab_sheaf *sheaf)
--
tools/include/linux/slab.h-210-
tools/include/linux/slab.h:211:#define kzalloc_obj(P, ...) \
tools/include/linux/slab.h-212- __alloc_objs(kzalloc, default_gfp(__VA_ARGS__), typeof(P), 1)
]
|
| 14/2 |
2026/09/17 00:45 |
llm |
expert |
0m
Model:
Tokens:
input: 20600
output: 43
thoughts: 771
Thoughts: **Analyzing a Potential OCFS2 Lockdep Issue**
Okay, let's break this down. It seems we have a potential lockdep issue in OCFS2, specifically related to metadata allocation during a `mknod` operation. My initial thought was that `kzalloc_obj` uses `GFP_KERNEL` which, by default, includes `__GFP_FS` and could trigger `fs_reclaim`. The problem is that `ocfs2_reserve_new_metadata_blocks` is called *before* `ocfs2_start_trans`, which usually would be fine, except...
...I have to make sure what locks are held during this call. I initially thought this was happening while holding the VFS `inode_lock(dir)`, but I have since corrected myself. `ocfs2_mknod` calls `ocfs2_inode_lock(dir, &parent_fe_bh, 1);` which is a cluster lock, and not the VFS inode lock.
The lockdep report is what caught my attention, specifically the chain: `&ocfs2_sysfile_lock_key[INODE_ALLOC_SYSTEM_INODE]` -> `fs_reclaim`. This means a lock dependency has been detected. I need to understand where the system lock is acquired in Chain 3.
Looking at the stack trace, I see that the problem occurs, and that the chain includes `ocfs2_reserve_new_metadata_blocks` which uses `kzalloc_obj`, which ultimately can trigger `fs_reclaim`. However, the provided stack trace doesn't *explicitly* show the system lock acquisition. Then I remember that lockdep only shows the stack trace of where the *new* lock (`fs_reclaim`) was acquired. The previously held lock, `&ocfs2_sysfile_lock_key[INODE_ALLOC_SYSTEM_INODE]`, was acquired earlier in the same task.
Now the key question is, where is `&ocfs2_sysfile_lock_key[INODE_ALLOC_SYSTEM_INODE]` acquired? Going back to the code, I see that `ocfs2_mknod`, before calling `ocfs2_reserve_new_metadata_blocks`, does call `ocfs2_reserve_new_inode`.
So, it's:
1. System Lock
2. `ocfs2_reserve_new_inode` -> `ocfs2_reserve_new_metadata_blocks` with possible allocation/reclaim
|
| 15/2 |
2026/09/17 00:46 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/ocfs2/suballoc.c Name:ocfs2_reserve_new_inode]
Results: map[SourceCode:1148: int ocfs2_reserve_new_inode(struct ocfs2_super *osb,
1149: struct ocfs2_alloc_context **ac)
1150: {
1151: int status;
1152: int slot = ocfs2_get_inode_steal_slot(osb);
1153: u64 alloc_group;
1154:
1155: *ac = kzalloc_obj(struct ocfs2_alloc_context);
1156: if (!(*ac)) {
1157: status = -ENOMEM;
1158: mlog_errno(status);
1159: goto bail;
1160: }
1161:
1162: (*ac)->ac_bits_wanted = 1;
1163: (*ac)->ac_which = OCFS2_AC_USE_INODE;
1164:
1165: (*ac)->ac_group_search = ocfs2_block_group_search;
1166:
1167: /*
1168: * stat(2) can't handle i_ino > 32bits, so we tell the
1169: * lower levels not to allocate us a block group past that
1170: * limit. The 'inode64' mount option avoids this behavior.
1171: */
1172: if (!(osb->s_mount_opt & OCFS2_MOUNT_INODE64))
1173: (*ac)->ac_max_block = (u32)~0U;
1174:
1175: /*
1176: * slot is set when we successfully steal inode from other nodes.
1177: * It is reset in 3 places:
1178: * 1. when we flush the truncate log
1179: * 2. when we complete local alloc recovery.
1180: * 3. when we successfully allocate from our own slot.
1181: * After it is set, we will go on stealing inodes until we find the
1182: * need to check our slots to see whether there is some space for us.
1183: */
1184: if (slot != OCFS2_INVALID_SLOT &&
1185: atomic_read(&osb->s_num_inodes_stolen) < OCFS2_MAX_TO_STEAL)
1186: goto inode_steal;
1187:
1188: atomic_set(&osb->s_num_inodes_stolen, 0);
1189: alloc_group = osb->osb_inode_alloc_group;
1190: status = ocfs2_reserve_suballoc_bits(osb, *ac,
1191: INODE_ALLOC_SYSTEM_INODE,
1192: (u32)osb->slot_num,
1193: &alloc_group,
1194: ALLOC_NEW_GROUP |
1195: ALLOC_GROUPS_FROM_GLOBAL);
1196: if (status >= 0) {
1197: status = 0;
1198:
1199: spin_lock(&osb->osb_lock);
1200: osb->osb_inode_alloc_group = alloc_group;
1201: spin_unlock(&osb->osb_lock);
1202: trace_ocfs2_reserve_new_inode_new_group(
1203: (unsigned long long)alloc_group);
1204:
1205: /*
1206: * Some inodes must be freed by us, so try to allocate
1207: * from our own next time.
1208: */
1209: if (slot != OCFS2_INVALID_SLOT)
1210: ocfs2_init_inode_steal_slot(osb);
1211: goto bail;
1212: } else if (status < 0 && status != -ENOSPC) {
1213: mlog_errno(status);
1214: goto bail;
1215: }
1216:
1217: ocfs2_free_ac_resource(*ac);
1218:
1219: inode_steal:
1220: status = ocfs2_steal_inode(osb, *ac);
1221: atomic_inc(&osb->s_num_inodes_stolen);
1222: if (status < 0) {
1223: if (status != -ENOSPC)
1224: mlog_errno(status);
1225: goto bail;
1226: }
1227:
1228: status = 0;
1229: bail:
1230: if ((status < 0) && *ac) {
1231: ocfs2_free_alloc_context(*ac);
1232: *ac = NULL;
1233: }
1234:
1235: if (status)
1236: mlog_errno(status);
1237: return status;
1238: }
SourceFile:fs/ocfs2/suballoc.c]
|
| 16/2 |
2026/09/17 00:46 |
llm |
expert |
0m
Model:
Tokens:
input: 23404
output: 44
thoughts: 36
Thoughts: Here's my summary, as though I'm thinking through this process:
**Diving into Inode Allocation in OCFS2**
Okay, so I'm looking at how OCFS2 handles the allocation of new inodes. Specifically, I'm tracing the execution of `ocfs2_reserve_new_inode`. It's pretty straightforward, initially. The key part I'm focusing on is that this function, `ocfs2_reserve_new_inode`, is calling `ocfs2_reserve_suballoc_bits`. The interesting thing is the *why*. It's calling it with `INODE_ALLOC_SYSTEM_INODE`. This tells me that this particular allocation path is intended for system inodes. Thatβs probably important, potentially critical, when we get into allocation strategies, and I need to keep that in the back of my head. Need to remember why that flag is set: is it a performance optimization? Does it relate to security? Does it influence the suballocation strategy in a particular way? I'll probably need to chase that down. So, essentially, `ocfs2_reserve_new_inode` is *using* a more general bit-reservation mechanism, but specifically for system inode allocation. That's good to keep in mind, and might influence how I'd approach any potential changes to the underlying bit reservation code.
|
| 17/2 |
2026/09/17 00:46 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/ocfs2/suballoc.c Name:ocfs2_reserve_suballoc_bits]
Results: map[SourceCode: 884: static int ocfs2_reserve_suballoc_bits(struct ocfs2_super *osb,
885: struct ocfs2_alloc_context *ac,
886: int type,
887: u32 slot,
888: u64 *last_alloc_group,
889: int flags)
890: {
891: int status;
892: u32 bits_wanted = ac->ac_bits_wanted;
893: struct inode *alloc_inode;
894: struct buffer_head *bh = NULL;
895: struct ocfs2_dinode *fe;
896: u32 free_bits;
897:
898: alloc_inode = ocfs2_get_system_file_inode(osb, type, slot);
899: if (!alloc_inode) {
900: mlog_errno(-EINVAL);
901: return -EINVAL;
902: }
903:
904: inode_lock(alloc_inode);
905:
906: status = ocfs2_inode_lock(alloc_inode, &bh, 1);
907: if (status < 0) {
908: inode_unlock(alloc_inode);
909: iput(alloc_inode);
910:
911: mlog_errno(status);
912: return status;
913: }
914:
915: ac->ac_inode = alloc_inode;
916: ac->ac_alloc_slot = slot;
917:
918: fe = (struct ocfs2_dinode *) bh->b_data;
919:
920: /* The bh was validated by the inode read inside
921: * ocfs2_inode_lock(). Any corruption is a code bug. */
922: BUG_ON(!OCFS2_IS_VALID_DINODE(fe));
923:
924: if (!(fe->i_flags & cpu_to_le32(OCFS2_CHAIN_FL))) {
925: status = ocfs2_error(alloc_inode->i_sb,
926: "Invalid chain allocator %llu\n",
927: (unsigned long long)le64_to_cpu(fe->i_blkno));
928: goto bail;
929: }
930:
931: free_bits = le32_to_cpu(fe->id1.bitmap1.i_total) -
932: le32_to_cpu(fe->id1.bitmap1.i_used);
933:
934: if (bits_wanted > free_bits) {
935: /* cluster bitmap never grows */
936: if (ocfs2_is_cluster_bitmap(alloc_inode)) {
937: trace_ocfs2_reserve_suballoc_bits_nospc(bits_wanted,
938: free_bits);
939: status = -ENOSPC;
940: goto bail;
941: }
942:
943: if (!(flags & ALLOC_NEW_GROUP)) {
944: trace_ocfs2_reserve_suballoc_bits_no_new_group(
945: slot, bits_wanted, free_bits);
946: status = -ENOSPC;
947: goto bail;
948: }
949:
950: status = ocfs2_block_group_alloc(osb, alloc_inode, bh,
951: ac->ac_max_block,
952: last_alloc_group, flags);
953: if (status < 0) {
954: if (status != -ENOSPC)
955: mlog_errno(status);
956: goto bail;
957: }
958: atomic_inc(&osb->alloc_stats.bg_extends);
959:
960: /* You should never ask for this much metadata */
961: BUG_ON(bits_wanted >
962: (le32_to_cpu(fe->id1.bitmap1.i_total)
963: - le32_to_cpu(fe->id1.bitmap1.i_used)));
964: }
965:
966: get_bh(bh);
967: ac->ac_bh = bh;
968: bail:
969: brelse(bh);
970:
971: if (status)
972: mlog_errno(status);
973: return status;
974: }
SourceFile:fs/ocfs2/suballoc.c]
|
| 18/2 |
2026/09/17 00:46 |
llm |
expert |
0m
Model:
Tokens:
input: 25167
output: 41
thoughts: 509
Thoughts: **Analyzing OCFS2 Lock Chain Dependencies**
Okay, let's break this down. First, I see that `ocfs2_reserve_suballoc_bits` acquires `inode_lock(alloc_inode)` and returns *with the lock held*. The interesting part is that the lock is actually released much later, when the `ocfs2_alloc_context` (`ac`) is freed, after `ac->ac_inode = alloc_inode` is set. This means that when `ocfs2_reserve_new_inode` calls it and acquires `inode_lock` on `INODE_ALLOC_SYSTEM_INODE` the lock persists until `ac` is freed. And, of course, `ocfs2_mknod` calls `ocfs2_reserve_new_metadata_blocks`, which then calls `kzalloc_obj(struct ocfs2_alloc_context)`. The `kzalloc_obj` uses `GFP_KERNEL` which can trigger `fs_reclaim`. This perfectly explains Chain 3: `&ocfs2_sysfile_lock_key[INODE_ALLOC_SYSTEM_INODE]` can block on `fs_reclaim`.
Now, onto Chain 2. This one is `&ocfs2_sysfile_lock_key[ORPHAN_DIR_SYSTEM_INODE]` potentially blocking on `&ocfs2_sysfile_lock_key[INODE_ALLOC_SYSTEM_INODE]`. This arises from the inode eviction process: `ocfs2_evict_inode` -> `ocfs2_delete_inode` -> `ocfs2_wipe_inode` -> `ocfs2_remove_inode`. I can see that `ocfs2_wipe_inode` takes `inode_lock(orphan_dir_inode)`. Then, in a later step, `ocfs2_remove_inode` takes `inode_lock(inode_alloc_inode)`. This makes perfect sense; you'd expect orphan dir operations to coordinate with inode allocation operations.
Finally, Chain 1 is a bit more nuanced: `&osb->nfs_sync_rwlock` potentially blocks on `&ocfs2_sysfile_lock_key[ORPHAN_DIR_SYSTEM_INODE]`. I see `ocfs2_evict_inode` calling `down_write(&osb->nfs_sync_rwlock)`. (The stack trace also mentions `down_read` for a similar chain, but for this Chain 1, it appears only `ocfs2_evict_inode` is reported.) So I'll need to check the code for `ocfs2_evict_inode` closely to understand the interaction with the orphan directory.
|
| 19/2 |
2026/09/17 00:46 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/ocfs2/inode.c Name:ocfs2_evict_inode]
Results: map[SourceCode:1327: void ocfs2_evict_inode(struct inode *inode)
1328: {
1329: write_inode_now(inode, 1);
1330:
1331: if (!inode->i_nlink ||
1332: (OCFS2_I(inode)->ip_flags & OCFS2_INODE_MAYBE_ORPHANED)) {
1333: ocfs2_delete_inode(inode);
1334: } else {
1335: truncate_inode_pages_final(&inode->i_data);
1336: }
1337: ocfs2_clear_inode(inode);
1338: }
SourceFile:fs/ocfs2/inode.c]
|
| 20/2 |
2026/09/17 00:46 |
llm |
expert |
0m
Model:
Tokens:
input: 26531
output: 40
thoughts: 40
Thoughts: **Investigating the NFS Sync RWLock in OCFS2 Eviction**
Okay, hold on a minute. I was just reviewing the OCFS2 inode eviction process, and something is nagging at me. Specifically, I'm trying to recall the exact place `nfs_sync_rwlock` is acquired within `ocfs2_evict_inode`. I thought I knew the code inside and out, but a gap is forming.
Right, time to refresh. The most likely place to start is `ocfs2_delete_inode`. That function should be the primary candidate for handling the final steps before an inode is considered gone. Let's dig in and see if `nfs_sync_rwlock` is being taken there, and, if so, in what context. I'll need to retrace the control flow and understand the interaction of locking mechanisms. Let's see...
|
| 21/2 |
2026/09/17 00:46 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/ocfs2/inode.c Name:ocfs2_delete_inode]
Results: map[SourceCode:1104: static void ocfs2_delete_inode(struct inode *inode)
1105: {
1106: int wipe, status;
1107: sigset_t oldset;
1108: struct buffer_head *di_bh = NULL;
1109: struct ocfs2_dinode *di = NULL;
1110:
1111: trace_ocfs2_delete_inode(inode->i_ino,
1112: (unsigned long long)OCFS2_I(inode)->ip_blkno,
1113: is_bad_inode(inode));
1114:
1115: /* When we fail in read_inode() we mark inode as bad. The second test
1116: * catches the case when inode allocation fails before allocating
1117: * a block for inode. */
1118: if (is_bad_inode(inode) || !OCFS2_I(inode)->ip_blkno)
1119: goto bail;
1120:
1121: if (!ocfs2_inode_is_valid_to_delete(inode)) {
1122: /* It's probably not necessary to truncate_inode_pages
1123: * here but we do it for safety anyway (it will most
1124: * likely be a no-op anyway) */
1125: ocfs2_cleanup_delete_inode(inode, 0);
1126: goto bail;
1127: }
1128:
1129: dquot_initialize(inode);
1130:
1131: /* We want to block signals in delete_inode as the lock and
1132: * messaging paths may return us -ERESTARTSYS. Which would
1133: * cause us to exit early, resulting in inodes being orphaned
1134: * forever. */
1135: ocfs2_block_signals(&oldset);
1136:
1137: /*
1138: * Synchronize us against ocfs2_get_dentry. We take this in
1139: * shared mode so that all nodes can still concurrently
1140: * process deletes.
1141: */
1142: status = ocfs2_nfs_sync_lock(OCFS2_SB(inode->i_sb), 0);
1143: if (status < 0) {
1144: mlog(ML_ERROR, "getting nfs sync lock(PR) failed %d\n", status);
1145: ocfs2_cleanup_delete_inode(inode, 0);
1146: goto bail_unblock;
1147: }
1148: /* Lock down the inode. This gives us an up to date view of
1149: * it's metadata (for verification), and allows us to
1150: * serialize delete_inode on multiple nodes.
1151: *
1152: * Even though we might be doing a truncate, we don't take the
1153: * allocation lock here as it won't be needed - nobody will
1154: * have the file open.
1155: */
1156: status = ocfs2_inode_lock(inode, &di_bh, 1);
1157: if (status < 0) {
1158: if (status != -ENOENT)
1159: mlog_errno(status);
1160: ocfs2_cleanup_delete_inode(inode, 0);
1161: goto bail_unlock_nfs_sync;
1162: }
1163:
1164: di = (struct ocfs2_dinode *)di_bh->b_data;
1165: /* Skip inode deletion and wait for dio orphan entry recovered
1166: * first */
1167: if (unlikely(di->i_flags & cpu_to_le32(OCFS2_DIO_ORPHANED_FL))) {
1168: ocfs2_cleanup_delete_inode(inode, 0);
1169: goto bail_unlock_inode;
1170: }
1171:
1172: /* Query the cluster. This will be the final decision made
1173: * before we go ahead and wipe the inode. */
1174: status = ocfs2_query_inode_wipe(inode, di_bh, &wipe);
1175: if (!wipe || status < 0) {
1176: /* Error and remote inode busy both mean we won't be
1177: * removing the inode, so they take almost the same
1178: * path. */
1179: if (status < 0)
1180: mlog_errno(status);
1181:
1182: /* Someone in the cluster has disallowed a wipe of
1183: * this inode, or it was never completely
1184: * orphaned. Write out the pages and exit now. */
1185: ocfs2_cleanup_delete_inode(inode, 1);
1186: goto bail_unlock_inode;
1187: }
1188:
1189: ocfs2_cleanup_delete_inode(inode, 0);
1190:
1191: status = ocfs2_wipe_inode(inode, di_bh);
1192: if (status < 0) {
1193: if (status != -EDEADLK)
1194: mlog_errno(status);
1195: goto bail_unlock_inode;
1196: }
1197:
1198: /*
1199: * Mark the inode as successfully deleted.
1200: *
1201: * This is important for ocfs2_clear_inode() as it will check
1202: * this flag and skip any checkpointing work
1203: *
1204: * ocfs2_stuff_meta_lvb() also uses this flag to invalidate
1205: * the LVB for other nodes.
1206: */
1207: OCFS2_I(inode)->ip_flags |= OCFS2_INODE_DELETED;
1208:
1209: bail_unlock_inode:
1210: ocfs2_inode_unlock(inode, 1);
1211: brelse(di_bh);
1212:
1213: bail_unlock_nfs_sync:
1214: ocfs2_nfs_sync_unlock(OCFS2_SB(inode->i_sb), 0);
1215:
1216: bail_unblock:
1217: ocfs2_unblock_signals(&oldset);
1218: bail:
1219: return;
1220: }
SourceFile:fs/ocfs2/inode.c]
|
| 22/2 |
2026/09/17 00:46 |
llm |
expert |
0m
Model:
Tokens:
input: 28677
output: 40
thoughts: 1167
Thoughts: **Deadlock Detected in OCFS2 - Serious Implications**
Okay, let's break this down. My initial thought process went like this: I was tracing through `ocfs2_delete_inode` and immediately saw the potential for a serious issue. `ocfs2_delete_inode` calls `ocfs2_nfs_sync_lock`, which in turn grabs `nfs_sync_rwlock` in read mode. Then, it calls `ocfs2_wipe_inode`, which grabs `inode_lock` on the orphan directory. So, the sequence is: `fs_reclaim` -> `nfs_sync_rwlock`, then `nfs_sync_rwlock` -> `inode_lock(orphan_dir)`.
From there I was able to observe how `inode_lock(orphan_dir)` can conflict with `inode_lock(inode_alloc)` during various operations. `inode_lock(inode_alloc)` can then call `fs_reclaim`, where we are back where we started.
Now, where it gets interesting is that `ocfs2_reserve_new_inode` takes `inode_lock(inode_alloc)` *and then returns holding it*. Then, `ocfs2_mknod` is called, which subsequently calls `kzalloc_obj`, which can trigger `fs_reclaim` with `GFP_KERNEL`. This means that a process doing a `mkdir` on OCFS2 will hold `inode_lock(inode_alloc)` while waiting for `fs_reclaim` to complete.
Here's the problem: *kswapd0* is doing `fs_reclaim` which in the process of reclaiming memory decides to evict an OCFS2 inode. It calls `ocfs2_evict_inode`, which in turn calls `ocfs2_delete_inode`, which eventually tries to get `inode_lock(inode_alloc)`. *But* the `mkdir` process already *has* that lock! Now kswapd0 is blocked.
The `mkdir` process is blocked on `fs_reclaim`, which kswapd0 *is* trying to accomplish. It's a textbook deadlock: kswapd0 is waiting for `inode_lock(inode_alloc)` while the process holding the lock is waiting for kswapd0 to finish its memory reclamation (i.e. to finish `fs_reclaim`). Neither can proceed.
Essentially, kswapd0's `fs_reclaim` will never finish, and the `mkdir` process will never get the resources it needs. This is a *real* deadlock, as the memory reclamation will block indefinitely.
The next question in my head became: Can an unprivileged user trigger this? To do so, they need a few conditions: access to an OCFS2 filesystem (mounted by root, obviously), write permissions, and enough memory pressure to trigger kswapd0 to evict an inode. An OCFS2 inode that needs deletion needs to be evicted (e.g., an unlinked file that was kept open, then closed, or simply an inode with `i_nlink == 0` or `OCFS2_INODE_MAYBE_ORPHANED`).
I paused here. `ocfs2_evict_inode` *only* calls `ocfs2_delete_inode` if `!inode->i_nlink || (OCFS2_I(inode)->ip_flags & OCFS2_INODE_MAYBE_ORPHANED)`. So, an unprivileged user can create a file, unlink it, and close it. But that deletion usually happens synchronously, within the process, not kswapd0. How can kswapd0 evict an inode with `i_nlink == 0`? The VFS usually calls `iput` which calls `evict` immediately if the refcount drops to 0.
So, how can kswapd0 call `ocfs2_evict_inode` with `!inode->i_nlink`? It's not usually possible, unless the inode somehow stays cached with `i_nlink == 0`. The VFS doesn't do that. That leaves `OCFS2_INODE_MAYBE_ORPHANED`. The key question is: *when* is `OCFS2_INODE_MAYBE_ORPHANED` set? That's the remaining piece of the puzzle. The unprivileged user may be able to trigger the deadlock if we can understand the conditions that may set this flag.
|
| 23/2 |
2026/09/17 00:46 |
tool |
grepper |
0m
Args: map[Expression:OCFS2_INODE_MAYBE_ORPHANED PathPrefix:fs/ocfs2/]
Results: map[Output:fs/ocfs2/dlmglue.c=4071=static int ocfs2_dentry_convert_worker(struct ocfs2_lock_res *lockres,
--
fs/ocfs2/dlmglue.c-4096- spin_lock(&oi->ip_lock);
fs/ocfs2/dlmglue.c:4097: oi->ip_flags |= OCFS2_INODE_MAYBE_ORPHANED;
fs/ocfs2/dlmglue.c-4098- spin_unlock(&oi->ip_lock);
--
fs/ocfs2/inode.c=1327=void ocfs2_evict_inode(struct inode *inode)
--
fs/ocfs2/inode.c-1331- if (!inode->i_nlink ||
fs/ocfs2/inode.c:1332: (OCFS2_I(inode)->ip_flags & OCFS2_INODE_MAYBE_ORPHANED)) {
fs/ocfs2/inode.c-1333- ocfs2_delete_inode(inode);
--
fs/ocfs2/inode.h=16=struct ocfs2_inode_info
--
fs/ocfs2/inode.h-93- */
fs/ocfs2/inode.h:94:#define OCFS2_INODE_MAYBE_ORPHANED 0x00000010
fs/ocfs2/inode.h-95-/* Does someone have the file open O_DIRECT */
--
fs/ocfs2/journal.c=2288=static int ocfs2_recover_orphans(struct ocfs2_super *osb,
--
fs/ocfs2/journal.c-2365- * ocfs2_delete_inode. */
fs/ocfs2/journal.c:2366: oi->ip_flags |= OCFS2_INODE_MAYBE_ORPHANED;
fs/ocfs2/journal.c-2367- spin_unlock(&oi->ip_lock);
--
fs/ocfs2/namei.c=97=static struct dentry *ocfs2_lookup(struct inode *dir, struct dentry *dentry,
--
fs/ocfs2/namei.c-140- spin_lock(&oi->ip_lock);
fs/ocfs2/namei.c:141: oi->ip_flags &= ~OCFS2_INODE_MAYBE_ORPHANED;
fs/ocfs2/namei.c-142- spin_unlock(&oi->ip_lock);
]
|
| 24/2 |
2026/09/17 00:46 |
llm |
expert |
0m
Model:
Tokens:
input: 31541
output: 43
thoughts: 4062
Thoughts: **Deadlock Deep Dive: Understanding the kswapd0 Conundrum in OCFS2**
Okay, here's the core of what I've figured out about this OCFS2 deadlock. It starts with the `OCFS2_INODE_MAYBE_ORPHANED` flag. This flag is set on an inode when a remote node in the cluster deletes a file and a local node receives the downconvert request on the dentry lock. The local node then drops the dentry, and the flag gets set. If the inode is still in the inode cache (maybe because it was recently accessed), it might get evicted later.
The key is that kswapd0 can then call `ocfs2_delete_inode` on this inode because it's marked as `OCFS2_INODE_MAYBE_ORPHANED`. The path is: kswapd0 shrinks the dcache -> finds the dentry -> kills it -> evicts the inode -> calls `ocfs2_evict_inode` -> which, because of the flag, calls `ocfs2_delete_inode`.
The lockdep report shows kswapd0 traversing a specific chain involving `nfs_sync_rwlock`, `inode_lock(orphan_dir)`, and then `inode_lock(inode_alloc)`. This chain sets up a circular dependency and a deadlock if another process is also in `fs_reclaim`.
Here's where it gets interesting: the deadlock happens when kswapd0 (or direct reclaim) is trying to evict this orphaned inode AND a process (like `mkdir`) is allocating a new inode, simultaneously holding `inode_lock(inode_alloc)` and blocked on `fs_reclaim`.
Specifically, the deadlock occurs as follows:
* kswapd0 wants to take `nfs_sync_rwlock`, `inode_lock(orphan_dir)`, and `inode_lock(inode_alloc)`.
* Another process (e.g. `mkdir`) is holding `inode_lock(inode_alloc)` and is blocked on `fs_reclaim`.
* kswapd0 blocks on `inode_lock(inode_alloc)` while holding `fs_reclaim`, and the `mkdir` blocks on `fs_reclaim`.
Now, the lockdep report confirms this:
* kswapd0 traverses Chains 0, 1, and 2.
* The `mkdir` process traverses Chain 3.
I initially thought this could ONLY happen in a cluster, but the syzkaller report on a single node threw a wrench in that. It turns out that a synchronous eviction and delete caused by an overlayfs mount is the culprit, because `ovl_check_rename_whiteout` is the one which drops the last reference to the file, and that calls a synchronous `evict` in the mount context, taking the same locks as kswapd0, but without `fs_reclaim`.
In a nutshell, it is not kswapd0 or direct reclaim which triggers the issue, but a synchronous eviction performed inside an overlayfs mount, which takes the locks in the incorrect order.
|
| 25/2 |
2026/09/17 00:47 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/ocfs2/dcache.c Name:ocfs2_dentry_iput]
Results: map[SourceCode: 379: static void ocfs2_dentry_iput(struct dentry *dentry, struct inode *inode)
380: {
381: struct ocfs2_dentry_lock *dl = dentry->d_fsdata;
382:
383: if (!dl) {
384: /*
385: * No dentry lock is ok if we're disconnected or
386: * unhashed.
387: */
388: if (!(dentry->d_flags & DCACHE_DISCONNECTED) &&
389: !d_unhashed(dentry)) {
390: unsigned long long ino = 0ULL;
391: if (inode)
392: ino = (unsigned long long)OCFS2_I(inode)->ip_blkno;
393: mlog(ML_ERROR, "Dentry is missing cluster lock. "
394: "inode: %llu, d_flags: 0x%x, d_name: %pd\n",
395: ino, dentry->d_flags, dentry);
396: }
397:
398: goto out;
399: }
400:
401: mlog_bug_on_msg(dl->dl_count == 0, "dentry: %pd, count: %u\n",
402: dentry, dl->dl_count);
403:
404: ocfs2_dentry_lock_put(OCFS2_SB(dentry->d_sb), dl);
405:
406: out:
407: iput(inode);
408: }
SourceFile:fs/ocfs2/dcache.c]
|
| 26/2 |
2026/09/17 00:47 |
llm |
expert |
0m
Model:
Tokens:
input: 36830
output: 45
thoughts: 270
Thoughts: **My Analysis of OCFS2 Dentry Eviction**
Okay, so `ocfs2_dentry_iput` is straightforward: it's just a wrapper around the standard `iput(inode)` call. That's fine. The interesting part is what happens when `iput` decrements the reference count to zero. It calls `evict`, which is what I'm interested in here. I'm seeing `kswapd0` working to shrink the dcache, which results in a dentry being killed off, leading to an `iput` and ultimately the `evict` function. But why does `evict` then call `ocfs2_delete_inode`? That seems like the crucial point of this sequence.
Ah, it's because either `inode->i_nlink` is zero *or* the `OCFS2_INODE_MAYBE_ORPHANED` flag is set. The question now is how `i_nlink` ends up at zero in the first place, or when that flag is set? Is it possible for `i_nlink` to be zero here? If a file is unlinked, its dentry is unhashed. My gut says no, it's not. But wait, is it even *possible* for an unhashed dentry to be on the Least Recently Used (LRU) list? No, `d_delete` removes it from the LRU. That's right. So, that's not the path.
Okay, let's step back, what if the file was unlinked on another node? Now we're getting somewhere. If the file is unlinked remotely, the local node's dentry *remains hashed* and is still on the LRU! This is the edge case. This is where things get interesting. So, when the local node *receives* the DLM downconvert, it sets the `OCFS2_INODE_MAYBE_ORPHANED` flag and *then* calls `d_delete`? Am I following this right? It's likely `d_delete` *after* the downconvert that is calling `iput`.
That's the likely scenario. Let's confirm that by digging into the implementation of `ocfs2_dentry_convert_worker`. Let me see...
|
| 27/2 |
2026/09/17 00:47 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/ocfs2/dlmglue.c Name:ocfs2_dentry_convert_worker]
Results: map[SourceCode:4052: /*
4053: * d_delete() matching dentries before the lock downconvert.
4054: *
4055: * At this point, any process waiting to destroy the
4056: * dentry_lock due to last ref count is stopped by the
4057: * OCFS2_LOCK_QUEUED flag.
4058: *
4059: * We have two potential problems
4060: *
4061: * 1) If we do the last reference drop on our dentry_lock (via dput)
4062: * we'll wind up in ocfs2_release_dentry_lock(), waiting on
4063: * the downconvert to finish. Instead we take an elevated
4064: * reference and push the drop until after we've completed our
4065: * unblock processing.
4066: *
4067: * 2) There might be another process with a final reference,
4068: * waiting on us to finish processing. If this is the case, we
4069: * detect it and exit out - there's no more dentries anyway.
4070: */
4071: static int ocfs2_dentry_convert_worker(struct ocfs2_lock_res *lockres,
4072: int blocking)
4073: {
4074: struct ocfs2_dentry_lock *dl = ocfs2_lock_res_dl(lockres);
4075: struct ocfs2_inode_info *oi = OCFS2_I(dl->dl_inode);
4076: struct dentry *dentry;
4077: unsigned long flags;
4078: int extra_ref = 0;
4079:
4080: /*
4081: * This node is blocking another node from getting a read
4082: * lock. This happens when we've renamed within a
4083: * directory. We've forced the other nodes to d_delete(), but
4084: * we never actually dropped our lock because it's still
4085: * valid. The downconvert code will retain a PR for this node,
4086: * so there's no further work to do.
4087: */
4088: if (blocking == DLM_LOCK_PR)
4089: return UNBLOCK_CONTINUE;
4090:
4091: /*
4092: * Mark this inode as potentially orphaned. The code in
4093: * ocfs2_delete_inode() will figure out whether it actually
4094: * needs to be freed or not.
4095: */
4096: spin_lock(&oi->ip_lock);
4097: oi->ip_flags |= OCFS2_INODE_MAYBE_ORPHANED;
4098: spin_unlock(&oi->ip_lock);
4099:
4100: /*
4101: * Yuck. We need to make sure however that the check of
4102: * OCFS2_LOCK_FREEING and the extra reference are atomic with
4103: * respect to a reference decrement or the setting of that
4104: * flag.
4105: */
4106: spin_lock_irqsave(&lockres->l_lock, flags);
4107: spin_lock(&dentry_attach_lock);
4108: if (!(lockres->l_flags & OCFS2_LOCK_FREEING)
4109: && dl->dl_count) {
4110: dl->dl_count++;
4111: extra_ref = 1;
4112: }
4113: spin_unlock(&dentry_attach_lock);
4114: spin_unlock_irqrestore(&lockres->l_lock, flags);
4115:
4116: mlog(0, "extra_ref = %d\n", extra_ref);
4117:
4118: /*
4119: * We have a process waiting on us in ocfs2_dentry_iput(),
4120: * which means we can't have any more outstanding
4121: * aliases. There's no need to do any more work.
4122: */
4123: if (!extra_ref)
4124: return UNBLOCK_CONTINUE;
4125:
4126: spin_lock(&dentry_attach_lock);
4127: while (1) {
4128: dentry = ocfs2_find_local_alias(dl->dl_inode,
4129: dl->dl_parent_blkno, 1);
4130: if (!dentry)
4131: break;
4132: spin_unlock(&dentry_attach_lock);
4133:
4134: if (S_ISDIR(dl->dl_inode->i_mode))
4135: shrink_dcache_parent(dentry);
4136:
4137: mlog(0, "d_delete(%pd);\n", dentry);
4138:
4139: /*
4140: * The following dcache calls may do an
4141: * iput(). Normally we don't want that from the
4142: * downconverting thread, but in this case it's ok
4143: * because the requesting node already has an
4144: * exclusive lock on the inode, so it can't be queued
4145: * for a downconvert.
4146: */
4147: d_delete(dentry);
4148: dput(dentry);
4149:
4150: spin_lock(&dentry_attach_lock);
4151: }
4152: spin_unlock(&dentry_attach_lock);
4153:
4154: /*
4155: * If we are the last holder of this dentry lock, there is no
4156: * reason to downconvert so skip straight to the unlock.
4157: */
4158: if (dl->dl_count == 1)
4159: return UNBLOCK_STOP_POST;
4160:
4161: return UNBLOCK_CONTINUE_POST;
4162: }
SourceFile:fs/ocfs2/dlmglue.c]
|
| 28/2 |
2026/09/17 00:47 |
llm |
expert |
0m
Model:
Tokens:
input: 39410
output: 41
thoughts: 1376
Thoughts: **Debugging the dentry and inode lifecycle in OCFS2**
Okay, so I'm digging into this lockdep report, and it's got me thinking. It seems `ocfs2_dentry_convert_worker` calls `d_delete` and `dput` on a dentry. That's the normal process. If `dput` drops the last reference, the dentry is killed immediately and `evict` is called via `iput`. This is happening in a DLM worker, not kswapd0, as expected.
But the lockdep report says kswapd0 is calling `ocfs2_delete_inode`, which is what is causing the concern. How is *that* happening? The stack trace shows kswapd0 shrinking the dcache via `shrink_dentry_list`, killing a dentry, and that leads to `evict` and then `ocfs2_delete_inode`. The condition for calling `ocfs2_delete_inode` from `evict` is either `!inode->i_nlink` or `OCFS2_INODE_MAYBE_ORPHANED`.
Now, if `OCFS2_INODE_MAYBE_ORPHANED` is set, `ocfs2_dentry_convert_worker` *should* be the culprit because it sets the flag. But this worker also calls `d_delete` and `dput`. If `dput` drops the last reference, it kills the dentry immediately. If it *doesn't* (another process has a reference), the dentry is unhashed but still around. When the *other* process drops the reference, `dput` will take care of it then. However, kswapd0 shrinks the LRU, but unhashed dentries aren't *on* the LRU. `d_delete` unhashes the dentry, so kswapd0 shouldn't be seeing it!
Could something else set `OCFS2_INODE_MAYBE_ORPHANED`? Ah, `ocfs2_recover_orphans`! However, it doesn't create any dentries, it just reads the inodes and sets the flag and calls `iput`. If that drops the last reference, `evict` gets called, but that's in the recovery thread.
So how can kswapd0 kill a dentry and end up calling `ocfs2_delete_inode`? What if `i_nlink == 0`? How can a dentry with `i_nlink == 0` be on the LRU? When a file is unlinked, `vfs_unlink` calls `d_delete`, which *unhashes* the dentry. But wait⦠could this be a *directory*? If a directory is removed, `vfs_rmdir` calls `d_delete`.
Could `d_delete` *fail* to unhash? No, that's not possible, it *always* unhashes.
Hold on, let me look at `ocfs2_dentry_iput`. The code checks if the dentry is disconnected and *not* unhashed. Is it possible that `ocfs2_dentry_iput` is being called for a dentry that isn't unhashed? Yes, if `dentry_kill` is called! `dentry_kill` *unhashes* the dentry and calls `iput`.
So here's the scenario: kswapd0 shrinks the dcache, finds a *hashed* dentry on the LRU, removes it, and calls `dentry_kill`. `dentry_kill` unhashes it and calls `iput`. If `iput` drops the last reference, then `evict` gets called. For `evict` to then call `ocfs2_delete_inode`, either `i_nlink` must be 0, or `OCFS2_INODE_MAYBE_ORPHANED` must be set. How can a hashed dentry on the LRU have an inode with `i_nlink == 0`? It shouldn't be. If it were unlinked, `vfs_unlink` would have unhashed the dentry. *Unless* the filesystem somehow set `i_nlink = 0` *without* unhashing the dentry! Does OCFS2 do that? I need to check the error paths in `ocfs2_mknod`!
|
| 29/2 |
2026/09/17 00:47 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/ocfs2/namei.c Name:ocfs2_mknod]
Results: map[SourceCode: 230: static int ocfs2_mknod(struct mnt_idmap *idmap,
231: struct inode *dir,
232: struct dentry *dentry,
233: umode_t mode,
234: dev_t dev)
235: {
236: int status = 0;
237: struct buffer_head *parent_fe_bh = NULL;
238: handle_t *handle = NULL;
239: struct ocfs2_super *osb;
240: struct ocfs2_dinode *dirfe;
241: struct ocfs2_dinode *fe = NULL;
242: struct buffer_head *new_fe_bh = NULL;
243: struct inode *inode = NULL;
244: struct ocfs2_alloc_context *inode_ac = NULL;
245: struct ocfs2_alloc_context *data_ac = NULL;
246: struct ocfs2_alloc_context *meta_ac = NULL;
247: int want_clusters = 0;
248: int want_meta = 0;
249: int xattr_credits = 0;
250: struct ocfs2_security_xattr_info si = {
251: .name = NULL,
252: .enable = 1,
253: };
254: int did_quota_inode = 0;
255: struct ocfs2_dir_lookup_result lookup = { NULL, };
256: sigset_t oldset;
257: int did_block_signals = 0;
258: struct ocfs2_dentry_lock *dl = NULL;
259: struct ocfs2_acl_state acl_state = { 0 };
260:
261: trace_ocfs2_mknod(dir, dentry, dentry->d_name.len, dentry->d_name.name,
262: (unsigned long long)OCFS2_I(dir)->ip_blkno,
263: (unsigned long)dev, mode);
264:
265: status = dquot_initialize(dir);
266: if (status) {
267: mlog_errno(status);
268: return status;
269: }
270:
271: /* get our super block */
272: osb = OCFS2_SB(dir->i_sb);
273:
274: status = ocfs2_inode_lock(dir, &parent_fe_bh, 1);
275: if (status < 0) {
276: if (status != -ENOENT)
277: mlog_errno(status);
278: return status;
279: }
280:
281: if (S_ISDIR(mode) && (dir->i_nlink >= ocfs2_link_max(osb))) {
282: status = -EMLINK;
283: goto leave;
284: }
285:
286: dirfe = (struct ocfs2_dinode *) parent_fe_bh->b_data;
287: if (!ocfs2_read_links_count(dirfe)) {
288: /* can't make a file in a deleted directory. */
289: status = -ENOENT;
290: goto leave;
291: }
292:
293: status = ocfs2_check_dir_for_entry(dir, dentry->d_name.name,
294: dentry->d_name.len);
295: if (status)
296: goto leave;
297:
298: /* get a spot inside the dir. */
299: status = ocfs2_prepare_dir_for_insert(osb, dir, parent_fe_bh,
300: dentry->d_name.name,
301: dentry->d_name.len, &lookup);
302: if (status < 0) {
303: mlog_errno(status);
304: goto leave;
305: }
306:
307: /* reserve an inode spot */
308: status = ocfs2_reserve_new_inode(osb, &inode_ac);
309: if (status < 0) {
310: if (status != -ENOSPC)
311: mlog_errno(status);
312: goto leave;
313: }
314:
315: inode = ocfs2_get_init_inode(dir, mode);
316: if (IS_ERR(inode)) {
317: status = PTR_ERR(inode);
318: inode = NULL;
319: mlog_errno(status);
320: goto leave;
321: }
322:
323: /* get security xattr */
324: status = ocfs2_init_security_get(inode, dir, &dentry->d_name, &si);
325: if (status) {
326: if (status == -EOPNOTSUPP)
327: si.enable = 0;
328: else {
329: mlog_errno(status);
330: goto leave;
331: }
332: }
333:
334: status = ocfs2_acl_init_prepare(inode, dir, parent_fe_bh, &acl_state);
335: if (status < 0)
336: goto leave;
337:
338: /* calculate meta data/clusters for setting security and acl xattr */
339: status = ocfs2_calc_xattr_init(dir, mode, &si, &want_clusters,
340: &xattr_credits, &want_meta,
341: &acl_state);
342: if (status < 0) {
343: mlog_errno(status);
344: goto leave;
345: }
346:
347: /* Reserve a cluster if creating an extent based directory. */
348: if (S_ISDIR(mode) && !ocfs2_supports_inline_data(osb)) {
349: want_clusters += 1;
350:
351: /* Dir indexing requires extra space as well */
352: if (ocfs2_supports_indexed_dirs(osb))
353: want_meta++;
354: }
355:
356: status = ocfs2_reserve_new_metadata_blocks(osb, want_meta, &meta_ac);
357: if (status < 0) {
358: if (status != -ENOSPC)
359: mlog_errno(status);
360: goto leave;
361: }
362:
363: status = ocfs2_reserve_clusters(osb, want_clusters, &data_ac);
364: if (status < 0) {
365: if (status != -ENOSPC)
366: mlog_errno(status);
367: goto leave;
368: }
369:
370: handle = ocfs2_start_trans(osb, ocfs2_mknod_credits(osb->sb,
371: S_ISDIR(mode),
372: xattr_credits));
373: if (IS_ERR(handle)) {
374: status = PTR_ERR(handle);
375: handle = NULL;
376: mlog_errno(status);
377: goto leave;
378: }
379:
380: /* Starting to change things, restart is no longer possible. */
381: ocfs2_block_signals(&oldset);
382: did_block_signals = 1;
383:
384: status = dquot_alloc_inode(inode);
385: if (status)
386: goto leave;
387: did_quota_inode = 1;
388:
389: /* do the real work now. */
390: status = ocfs2_mknod_locked(osb, dir, inode, dev,
391: &new_fe_bh, parent_fe_bh, handle,
392: inode_ac);
393: if (status < 0) {
394: mlog_errno(status);
395: goto leave;
396: }
397:
398: fe = (struct ocfs2_dinode *) new_fe_bh->b_data;
399: if (S_ISDIR(mode)) {
400: status = ocfs2_fill_new_dir(osb, handle, dir, inode,
401: new_fe_bh, data_ac, meta_ac);
402: if (status < 0) {
403: mlog_errno(status);
404: goto leave;
405: }
406:
407: status = ocfs2_journal_access_di(handle, INODE_CACHE(dir),
408: parent_fe_bh,
409: OCFS2_JOURNAL_ACCESS_WRITE);
410: if (status < 0) {
411: mlog_errno(status);
412: goto leave;
413: }
414: ocfs2_add_links_count(dirfe, 1);
415: ocfs2_journal_dirty(handle, parent_fe_bh);
416: inc_nlink(dir);
417: }
418:
419: status = ocfs2_init_acl(handle, inode, new_fe_bh, meta_ac, data_ac,
420: &acl_state);
421:
422: if (status < 0) {
423: mlog_errno(status);
424: goto roll_back;
425: }
426:
427: if (si.enable) {
428: status = ocfs2_init_security_set(handle, inode, new_fe_bh, &si,
429: meta_ac, data_ac);
430: if (status < 0) {
431: mlog_errno(status);
432: goto roll_back;
433: }
434: }
435:
436: /*
437: * Do this before adding the entry to the directory. We add
438: * also set d_op after success so that ->d_iput() will cleanup
439: * the dentry lock even if ocfs2_add_entry() fails below.
440: */
441: status = ocfs2_dentry_attach_lock(dentry, inode,
442: OCFS2_I(dir)->ip_blkno);
443: if (status) {
444: mlog_errno(status);
445: goto roll_back;
446: }
447:
448: dl = dentry->d_fsdata;
449:
450: status = ocfs2_add_entry(handle, dentry, inode,
451: OCFS2_I(inode)->ip_blkno, parent_fe_bh,
452: &lookup);
453: if (status < 0) {
454: mlog_errno(status);
455: goto roll_back;
456: }
457:
458: insert_inode_hash(inode);
459: d_instantiate(dentry, inode);
460: status = 0;
461:
462: roll_back:
463: if (status < 0 && S_ISDIR(mode)) {
464: ocfs2_add_links_count(dirfe, -1);
465: drop_nlink(dir);
466: }
467:
468: leave:
469: if (status < 0 && did_quota_inode)
470: dquot_free_inode(inode);
471: if (handle) {
472: if (status < 0 && fe)
473: ocfs2_set_links_count(fe, 0);
474: ocfs2_commit_trans(osb, handle);
475: }
476:
477: ocfs2_inode_unlock(dir, 1);
478: if (did_block_signals)
479: ocfs2_unblock_signals(&oldset);
480:
481: brelse(new_fe_bh);
482: brelse(parent_fe_bh);
483: kfree(si.value);
484:
485: ocfs2_acl_init_release(&acl_state);
486:
487: ocfs2_free_dir_lookup_result(&lookup);
488:
489: if (inode_ac)
490: ocfs2_free_alloc_context(inode_ac);
491:
492: if (data_ac)
493: ocfs2_free_alloc_context(data_ac);
494:
495: if (meta_ac)
496: ocfs2_free_alloc_context(meta_ac);
497:
498: /*
499: * We should call iput after the i_rwsem of the bitmap been
500: * unlocked in ocfs2_free_alloc_context, or the
501: * ocfs2_delete_inode will mutex_lock again.
502: */
503: if ((status < 0) && inode) {
504: if (dl)
505: ocfs2_cleanup_add_entry_failure(osb, dentry, inode);
506:
507: OCFS2_I(inode)->ip_flags |= OCFS2_INODE_SKIP_ORPHAN_DIR;
508: clear_nlink(inode);
509: iput(inode);
510: }
511:
512: if (status)
513: mlog_errno(status);
514:
515: return status;
516: }
SourceFile:fs/ocfs2/namei.c]
|
| 30/2 |
2026/09/17 00:47 |
llm |
expert |
0m
Model:
Tokens:
input: 46069
output: 38
thoughts: 1233
Thoughts: **Analyzing OCFS2 Inode Deletion Paths**
Okay, let's break this down. In `ocfs2_mknod`, if adding the directory entry fails, we roll back and end up in the `leave` function. Here's where it gets interesting: If there's an error and the inode exists, it calls `clear_nlink`, setting `i_nlink` to zero, and then `iput`. This sequence leads to `evict` being called, which eventually results in `ocfs2_delete_inode`, which is called synchronously by the process doing `mkdir`. The crucial point here is that `ocfs2_delete_inode` is called from the original `mkdir` context, not kswapd0.
So, the question is, how does kswapd0 get involved with `ocfs2_delete_inode`? Is there a scenario where `clear_nlink` is called, but the `iput` doesn't represent the last reference? The code indicates this isn't possible in the `ocfs2_mknod` failure case because the dentry isn't instantiated until after the add entry succeeds.
Okay, let's shift focus to how kswapd0 triggers `ocfs2_delete_inode`. The stack trace shows kswapd0 going through `shrink_dentry_list`, ultimately calling `ocfs2_delete_inode` via `ocfs2_evict_inode`, meaning kswapd0 is shrinking the dcache and is killing a dentry. For `ocfs2_delete_inode` to be called, `i_nlink` must be 0, or `OCFS2_INODE_MAYBE_ORPHANED` must be set. The dentry was on the LRU, so it was a positive dentry. If `i_nlink` is zero, how did this dentry with zero links end up on the LRU? Unlinked files normally remove dentries from LRU via `d_delete`.
Ah, the lightbulb moment! What if a file was unlinked, but `d_delete` wasn't called? When would that happen? Potentially, if the filesystem unlinks the file internally, or, especially, if it's a directory, when a directory is removed, `vfs_rmdir` calls `d_delete`. The key here is *corruption*. If the filesystem is corrupted, and `i_nlink` is zero on disk, the inode will be read with zero links.
Consider this: The filesystem is corrupted, and there's a directory entry pointing to an inode with `i_nlink == 0`. When a `lookup` reads that inode from disk, it sees `i_nlink == 0`. Then, a dentry is instantiated for it. The dentry is then put, and it ends up on the LRU. When kswapd0 shrinks the dcache, it finds this dentry, kills it, and `evict` sees `i_nlink == 0` and calls `ocfs2_delete_inode`! Bingo! Corrupted filesystem images can lead to this very scenario. The remaining question: does `ocfs2_iget` allow reading such inodes?
|
| 31/2 |
2026/09/17 00:47 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/ocfs2/inode.c Name:ocfs2_iget]
Results: map[SourceCode: 153: struct inode *ocfs2_iget(struct ocfs2_super *osb, u64 blkno, unsigned flags,
154: int sysfile_type)
155: {
156: int rc = -ESTALE;
157: struct inode *inode = NULL;
158: struct super_block *sb = osb->sb;
159: struct ocfs2_find_inode_args args;
160: journal_t *journal = osb->journal->j_journal;
161:
162: trace_ocfs2_iget_begin((unsigned long long)blkno, flags,
163: sysfile_type);
164:
165: /* Ok. By now we've either got the offsets passed to us by the
166: * caller, or we just pulled them off the bh. Lets do some
167: * sanity checks to make sure they're OK. */
168: if (blkno == 0) {
169: inode = ERR_PTR(-EINVAL);
170: mlog_errno(PTR_ERR(inode));
171: goto bail;
172: }
173:
174: args.fi_blkno = blkno;
175: args.fi_flags = flags;
176: args.fi_ino = ino_from_blkno(sb, blkno);
177: args.fi_sysfile_type = sysfile_type;
178:
179: inode = iget5_locked(sb, args.fi_ino, ocfs2_find_actor,
180: ocfs2_init_locked_inode, &args);
181: /* inode was *not* in the inode cache. 2.6.x requires
182: * us to do our own read_inode call and unlock it
183: * afterwards. */
184: if (inode == NULL) {
185: inode = ERR_PTR(-ENOMEM);
186: mlog_errno(PTR_ERR(inode));
187: goto bail;
188: }
189: trace_ocfs2_iget5_locked(inode_state_read_once(inode));
190: if (inode_state_read_once(inode) & I_NEW) {
191: rc = ocfs2_read_locked_inode(inode, &args);
192: unlock_new_inode(inode);
193: }
194: if (is_bad_inode(inode)) {
195: iput(inode);
196: inode = ERR_PTR(rc);
197: goto bail;
198: }
199:
200: /*
201: * Set transaction id's of transactions that have to be committed
202: * to finish f[data]sync. We set them to currently running transaction
203: * as we cannot be sure that the inode or some of its metadata isn't
204: * part of the transaction - the inode could have been reclaimed and
205: * now it is reread from disk.
206: */
207: if (journal) {
208: transaction_t *transaction;
209: tid_t tid;
210: struct ocfs2_inode_info *oi = OCFS2_I(inode);
211:
212: read_lock(&journal->j_state_lock);
213: if (journal->j_running_transaction)
214: transaction = journal->j_running_transaction;
215: else
216: transaction = journal->j_committing_transaction;
217: if (transaction)
218: tid = transaction->t_tid;
219: else
220: tid = journal->j_commit_sequence;
221: read_unlock(&journal->j_state_lock);
222: oi->i_sync_tid = tid;
223: oi->i_datasync_tid = tid;
224: }
225:
226: bail:
227: if (!IS_ERR(inode)) {
228: trace_ocfs2_iget_end(inode,
229: (unsigned long long)OCFS2_I(inode)->ip_blkno);
230: }
231:
232: return inode;
233: }
SourceFile:fs/ocfs2/inode.c]
|
| 32/2 |
2026/09/17 00:47 |
llm |
expert |
0m
Model:
Tokens:
input: 49335
output: 40
thoughts: 51
Thoughts: Okay, let me think this through.
**Analyzing Inode Handling in OCFS2**
Alright, so `ocfs2_read_locked_inode` is the function that's pulling the inode data directly from disk, right? I'm wondering about a specific edge case. If `i_nlink` is zeroβmeaning no hard links exist pointing to the inodeβdoes the filesystem treat that as an error condition, potentially flagging the inode as corrupted or bad? That seems like a logical check to make, especially during the read process.
To confirm that, I need to understand what's happening at a lower level. My next step is to examine `ocfs2_populate_inode`. That's where the initial inode structure is set up, so I should see the logic surrounding `i_nlink` being handled and whether it triggers any error handling at that point, or if the check is performed later, perhaps during an update or write operation. The whole flow needs a close look.
|
| 33/2 |
2026/09/17 00:47 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/ocfs2/inode.c Name:ocfs2_populate_inode]
Results: map[SourceCode: 372: void ocfs2_populate_inode(struct inode *inode, struct ocfs2_dinode *fe,
373: int create_ino)
374: {
375: struct super_block *sb;
376: struct ocfs2_super *osb;
377: int use_plocks = 1;
378:
379: sb = inode->i_sb;
380: osb = OCFS2_SB(sb);
381:
382: if ((osb->s_mount_opt & OCFS2_MOUNT_LOCALFLOCKS) ||
383: ocfs2_mount_local(osb) || !ocfs2_stack_supports_plocks())
384: use_plocks = 0;
385:
386: /*
387: * These have all been checked by ocfs2_read_inode_block() or set
388: * by ocfs2_mknod_locked(), so a failure is a code bug.
389: */
390: BUG_ON(!OCFS2_IS_VALID_DINODE(fe)); /* This means that read_inode
391: cannot create a superblock
392: inode today. change if
393: that is needed. */
394: BUG_ON(!(fe->i_flags & cpu_to_le32(OCFS2_VALID_FL)));
395: BUG_ON(le32_to_cpu(fe->i_fs_generation) != osb->fs_generation);
396:
397:
398: OCFS2_I(inode)->ip_clusters = le32_to_cpu(fe->i_clusters);
399: OCFS2_I(inode)->ip_attr = le32_to_cpu(fe->i_attr);
400: OCFS2_I(inode)->ip_dyn_features = le16_to_cpu(fe->i_dyn_features);
401:
402: inode_set_iversion(inode, 1);
403: inode->i_generation = le32_to_cpu(fe->i_generation);
404: inode->i_rdev = huge_decode_dev(le64_to_cpu(fe->id1.dev1.i_rdev));
405: inode->i_mode = le16_to_cpu(fe->i_mode);
406: i_uid_write(inode, le32_to_cpu(fe->i_uid));
407: i_gid_write(inode, le32_to_cpu(fe->i_gid));
408:
409: /* Fast symlinks will have i_size but no allocated clusters. */
410: if (S_ISLNK(inode->i_mode) && !fe->i_clusters) {
411: inode->i_blocks = 0;
412: inode->i_mapping->a_ops = &ocfs2_fast_symlink_aops;
413: } else {
414: inode->i_blocks = ocfs2_inode_sector_count(inode);
415: inode->i_mapping->a_ops = &ocfs2_aops;
416: }
417: inode_set_atime(inode, le64_to_cpu(fe->i_atime),
418: le32_to_cpu(fe->i_atime_nsec));
419: inode_set_mtime(inode, le64_to_cpu(fe->i_mtime),
420: le32_to_cpu(fe->i_mtime_nsec));
421: inode_set_ctime(inode, le64_to_cpu(fe->i_ctime),
422: le32_to_cpu(fe->i_ctime_nsec));
423:
424: if (OCFS2_I(inode)->ip_blkno != le64_to_cpu(fe->i_blkno))
425: mlog(ML_ERROR,
426: "ip_blkno %llu != i_blkno %llu!\n",
427: (unsigned long long)OCFS2_I(inode)->ip_blkno,
428: (unsigned long long)le64_to_cpu(fe->i_blkno));
429:
430: set_nlink(inode, ocfs2_read_links_count(fe));
431:
432: trace_ocfs2_populate_inode(OCFS2_I(inode)->ip_blkno,
433: le32_to_cpu(fe->i_flags));
434: if (fe->i_flags & cpu_to_le32(OCFS2_SYSTEM_FL)) {
435: OCFS2_I(inode)->ip_flags |= OCFS2_INODE_SYSTEM_FILE;
436: inode->i_flags |= S_NOQUOTA;
437: }
438:
439: if (fe->i_flags & cpu_to_le32(OCFS2_LOCAL_ALLOC_FL)) {
440: OCFS2_I(inode)->ip_flags |= OCFS2_INODE_BITMAP;
441: } else if (fe->i_flags & cpu_to_le32(OCFS2_BITMAP_FL)) {
442: OCFS2_I(inode)->ip_flags |= OCFS2_INODE_BITMAP;
443: } else if (fe->i_flags & cpu_to_le32(OCFS2_QUOTA_FL)) {
444: inode->i_flags |= S_NOQUOTA;
445: } else if (fe->i_flags & cpu_to_le32(OCFS2_SUPER_BLOCK_FL)) {
446: /* we can't actually hit this as read_inode can't
447: * handle superblocks today ;-) */
448: BUG();
449: }
450:
451: switch (inode->i_mode & S_IFMT) {
452: case S_IFREG:
453: if (use_plocks)
454: inode->i_fop = &ocfs2_fops;
455: else
456: inode->i_fop = &ocfs2_fops_no_plocks;
457: inode->i_op = &ocfs2_file_iops;
458: i_size_write(inode, le64_to_cpu(fe->i_size));
459: break;
460: case S_IFDIR:
461: inode->i_op = &ocfs2_dir_iops;
462: if (use_plocks)
463: inode->i_fop = &ocfs2_dops;
464: else
465: inode->i_fop = &ocfs2_dops_no_plocks;
466: i_size_write(inode, le64_to_cpu(fe->i_size));
467: OCFS2_I(inode)->ip_dir_lock_gen = 1;
468: break;
469: case S_IFLNK:
470: inode->i_op = &ocfs2_symlink_inode_operations;
471: inode_nohighmem(inode);
472: i_size_write(inode, le64_to_cpu(fe->i_size));
473: break;
474: default:
475: inode->i_op = &ocfs2_special_file_iops;
476: init_special_inode(inode, inode->i_mode,
477: inode->i_rdev);
478: break;
479: }
480:
481: if (create_ino) {
482: inode->i_ino = ino_from_blkno(inode->i_sb,
483: le64_to_cpu(fe->i_blkno));
484:
485: /*
486: * If we ever want to create system files from kernel,
487: * the generation argument to
488: * ocfs2_inode_lock_res_init() will have to change.
489: */
490: BUG_ON(le32_to_cpu(fe->i_flags) & OCFS2_SYSTEM_FL);
491:
492: ocfs2_inode_lock_res_init(&OCFS2_I(inode)->ip_inode_lockres,
493: OCFS2_LOCK_TYPE_META, 0, inode);
494:
495: ocfs2_inode_lock_res_init(&OCFS2_I(inode)->ip_open_lockres,
496: OCFS2_LOCK_TYPE_OPEN, 0, inode);
497: }
498:
499: ocfs2_inode_lock_res_init(&OCFS2_I(inode)->ip_rw_lockres,
500: OCFS2_LOCK_TYPE_RW, inode->i_generation,
501: inode);
502:
503: ocfs2_set_inode_flags(inode);
504:
505: OCFS2_I(inode)->ip_last_used_slot = 0;
506: OCFS2_I(inode)->ip_last_used_group = 0;
507:
508: if (S_ISDIR(inode->i_mode))
509: ocfs2_resv_set_type(&OCFS2_I(inode)->ip_la_data_resv,
510: OCFS2_RESV_FLAG_DIR);
511: }
SourceFile:fs/ocfs2/inode.c]
|
| 34/2 |
2026/09/17 00:47 |
llm |
expert |
0m
Model:
Tokens:
input: 52456
output: 36
thoughts: 2672
Thoughts: **My Analysis of the OCFS2 Deadlock Vulnerability**
Okay, so I've been dissecting this OCFS2 deadlock issue, and here's what I've concluded. It all boils down to the interaction of the dcache, kswapd0, and inode deletion, specifically when an inode with a zero link count somehow ends up with a valid dentry.
Initially, I see how `ocfs2_populate_inode` *doesn't* check for a zero `i_nlink` value. This means a corrupted filesystem, where a directory entry points to an inode with zero links, can create a dentry for it when a `stat` or `open` occurs. That dentry gets placed on the LRU. Then, memory pressure kicks in, kswapd0 shrinks the dcache, finds the dentry, kills it, and calls `ocfs2_delete_inode`. This is the key to triggering the bug.
Syzkaller exploited this by mounting a *corrupted* OCFS2 image. The image has a directory entry pointing to an inode with `i_nlink == 0`. Then `stat` or `open` and subsequent memory pressure results in the deadlock.
I then started to investigate whether this requires a *corrupted* filesystem, or if a *normal* cluster could be triggered by an unprivileged user. I realized that a normal OCFS2 cluster, where DLM downconverts might set `OCFS2_INODE_MAYBE_ORPHANED` on Node A, while Node B deletes the file, results in the same deadlock!
Next, I looked at `ocfs2_dentry_convert_worker`. Initially I thought that this worker could leave a positive dentry on the LRU. But, after scrutinizing the code, particularly `d_delete` and the scenarios of `d_count == 1` and `d_count > 1`, I realized that this worker ALWAYS unhashes the dentry (making it either negative or unhashed and positive). So the worker does NOT create a positive dentry on the LRU. This means that a cluster can't be triggered by this means.
Then I realized that, with a corrupted image, if a directory entry points to a zero-link inode, `lookup` can create a positive and *hashed* dentry! This is the only way a positive dentry can be created on the LRU with an inode having `i_nlink == 0`. Kswapd0 shrinks the dcache, finds this dentry on the LRU, and kills it. `dentry_kill` unhashes it and calls `iput`, which calls `evict`, sees that `i_nlink == 0` and calls `ocfs2_delete_inode`. This is the crux.
I considered other failure paths, like `ocfs2_mknod` or `ocfs2_unlink`, where `i_nlink` might be modified. However, in these cases the dentry is either negative or unhashed, and therefore never on the LRU. `ocfs2_mknod` failure results in a negative dentry, while `ocfs2_unlink` unhashes the dentry.
I've determined that this deadlock is a *Malicious Filesystem Trigger*. An image that's been corrupted is the only method to create a dentry, on the LRU, pointing to a `i_nlink == 0` inode.
I confirmed that an unprivileged user can't trigger this directly, as OCFS2 doesn't support `FS_USERNS_MOUNT`, so they can't mount a corrupted image. However, a *Peripheral Trigger* scenario exists: if root mounts a USB drive with a corrupted OCFS2 image, and then an unprivileged user accesses it, the unprivileged user CAN trigger the deadlock! This results in a Denial of Service.
|
| 35/2 |
2026/09/17 00:48 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/dcache.c Name:dentry_unlink_inode]
Results: map[SourceCode: 453: /*
454: * Release the dentry's inode, using the filesystem
455: * d_iput() operation if defined.
456: */
457: static void dentry_unlink_inode(struct dentry * dentry)
458: __releases(dentry->d_lock)
459: __releases(dentry->d_inode->i_lock)
460: {
461: struct inode *inode = dentry->d_inode;
462:
463: raw_write_seqcount_begin(&dentry->d_seq);
464: __d_clear_type_and_inode(dentry);
465: __hlist_del(&dentry->d_alias);
466: /*
467: * dentry becomes negative, so the space occupied by ->d_alias
468: * belongs to ->waiters now.
469: */
470: dentry->waiters = NULL;
471: raw_write_seqcount_end(&dentry->d_seq);
472: spin_unlock(&dentry->d_lock);
473: spin_unlock(&inode->i_lock);
474: if (!inode->i_nlink)
475: fsnotify_inoderemove(inode);
476: if (dentry->d_op && dentry->d_op->d_iput)
477: dentry->d_op->d_iput(dentry, inode);
478: else
479: iput(inode);
480: }
SourceFile:fs/dcache.c]
|
| 36/2 |
2026/09/17 00:48 |
llm |
expert |
0m
Model:
Tokens:
input: 56486
output: 34
thoughts: 1034
Thoughts: **Analyzing the Deadlock Scenario**
Okay, so I'm diving deep into this potential deadlock issue within OCFS2. My main focus is understanding the interaction between `kswapd0` and the file system's inode deletion process. The critical point is whether `ocfs2_delete_inode` gets called by `kswapd0` in a way that could cause contention with other locks.
First, I see that `dentry_unlink_inode` calls `iput(inode)`. If a dentry's count is 1 (`d_count == 1`), `d_delete` calls this `dentry_unlink_inode`, which leads to an `iput`. This `iput` is called synchronously within the `ocfs2_dentry_convert_worker`. Crucially, this is *not* `kswapd0` calling it. This means `ocfs2_delete_inode` will be triggered by a DLM worker thread, which doesn't hold `fs_reclaim`. No deadlock there.
Now, if `d_count > 1`, `d_delete` calls `__d_drop`, which unhashes the dentry, but the dentry remains positive. When another process subsequently calls `dput`, `dput` sees the unhashed dentry and calls `dentry_kill`, which in turn calls `iput(inode)`. This `iput` is again synchronous, called by the process performing a path walk, which also likely doesn't hold `fs_reclaim`, so this scenario also avoids deadlock.
Therefore, the `ocfs2_dentry_convert_worker` itself can't lead to `kswapd0` calling `ocfs2_delete_inode`. The only way I can see this happening is if `kswapd0` finds a positive, hashed dentry on the LRU (Least Recently Used) list and, during its killing process, `evict` triggers `ocfs2_delete_inode`. This requires either `i_nlink == 0` or `OCFS2_INODE_MAYBE_ORPHANED`. Remember, `OCFS2_INODE_MAYBE_ORPHANED` is only set by the DLM worker (during unhashing/making negative) or recovery (which doesn't create new dentries). So, in a normal file system, I'm pretty confident we'll *never* see a positive, hashed dentry on the LRU that would trigger `ocfs2_delete_inode` during regular operation.
Okay, so the most likely culprit appears to be a corrupted file system image. If there's a directory entry pointing to an inode with `i_nlink == 0`, accessing this file creates a positive, hashed dentry on the LRU. Then, `kswapd0` can evict it, potentially causing the deadlock.
But... could there be *any* other possibility? What about inode eviction from the inode cache? If an inode has no dentries, it could be on the inode LRU. `kswapd0` *does* shrink the inode cache, so I'll check the stack trace again. The stack trace shows the path is `shrink_dentry_list` -> `dentry_kill` -> `ocfs2_dentry_iput`. So it's shrinking the dcache and not the inode cache which validates my previous hypothesis.
So, is a corrupted filesystem image *the* only way? Could there be a race condition in `ocfs2_unlink`? If `ocfs2_unlink` drops the nlink and then `vfs_unlink` calls `d_delete`, is it possible that `kswapd0` shrinks the dcache *between* the nlink drop and the call to `d_delete`? I need to check `vfs_unlink`... this seems like a plausible avenue of exploration.
|
| 37/2 |
2026/09/17 00:48 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/namei.c Name:vfs_unlink]
Results: map[SourceCode:5598: /**
5599: * vfs_unlink - unlink a filesystem object
5600: * @idmap: idmap of the mount the inode was found from
5601: * @dir: parent directory
5602: * @dentry: victim
5603: * @delegated_inode: returns victim inode, if the inode is delegated.
5604: *
5605: * The caller must hold dir->i_rwsem exclusively.
5606: *
5607: * If vfs_unlink discovers a delegation, it will return -EWOULDBLOCK and
5608: * return a reference to the inode in delegated_inode. The caller
5609: * should then break the delegation on that inode and retry. Because
5610: * breaking a delegation may take a long time, the caller should drop
5611: * dir->i_rwsem before doing so.
5612: *
5613: * Alternatively, a caller may pass NULL for delegated_inode. This may
5614: * be appropriate for callers that expect the underlying filesystem not
5615: * to be NFS exported.
5616: *
5617: * If the inode has been found through an idmapped mount the idmap of
5618: * the vfsmount must be passed through @idmap. This function will then take
5619: * care to map the inode according to @idmap before checking permissions.
5620: * On non-idmapped mounts or if permission checking is to be performed on the
5621: * raw inode simply pass @nop_mnt_idmap.
5622: */
5623: int vfs_unlink(struct mnt_idmap *idmap, struct inode *dir,
5624: struct dentry *dentry, struct delegated_inode *delegated_inode)
5625: {
5626: struct inode *target = dentry->d_inode;
5627: int error = may_delete_dentry(idmap, dir, dentry, false);
5628:
5629: if (error)
5630: return error;
5631:
5632: if (!dir->i_op->unlink)
5633: return -EPERM;
5634:
5635: inode_lock(target);
5636: if (IS_SWAPFILE(target))
5637: error = -EPERM;
5638: else if (is_local_mountpoint(dentry))
5639: error = -EBUSY;
5640: else {
5641: error = security_inode_unlink(dir, dentry);
5642: if (!error) {
5643: error = try_break_deleg(dir, LEASE_BREAK_DIR_DELETE, delegated_inode);
5644: if (error)
5645: goto out;
5646: error = try_break_deleg(target, 0, delegated_inode);
5647: if (error)
5648: goto out;
5649: error = dir->i_op->unlink(dir, dentry);
5650: if (!error) {
5651: dont_mount(dentry);
5652: detach_mounts(dentry);
5653: }
5654: }
5655: }
5656: out:
5657: inode_unlock(target);
5658:
5659: /* We don't d_delete() NFS sillyrenamed files--they still exist. */
5660: if (!error && dentry->d_flags & DCACHE_NFSFS_RENAMED) {
5661: fsnotify_unlink(dir, dentry);
5662: } else if (!error) {
5663: fsnotify_link_count(target);
5664: d_delete_notify(dir, dentry);
5665: }
5666:
5667: return error;
5668: }
SourceFile:fs/namei.c]
|
| 38/2 |
2026/09/17 00:48 |
llm |
expert |
0m
Model:
Tokens:
input: 59529
output: 33
thoughts: 3670
Thoughts: **Analysis of a Deadlock in OCFS2 Filesystem**
Okay, so I've dug into this deadlock issue and have a pretty clear picture now. It appears this all stems from a race condition that can occur with a corrupted OCFS2 filesystem image. Here's the breakdown:
First, let's establish the context: `vfs_unlink`, `vfs_rmdir`, and `vfs_rename` - these are the usual suspects. But, under normal conditions, they *shouldn't* lead to trouble because the relevant dentry always has a d_count > 0, keeping it off the LRU list, and thus safe from kswapd0. However, the catch is in how this interacts with a *malformed* filesystem image, specifically one created by Syzkaller.
Syzkaller mounts a corrupted ocfs2 image, one where a directory entry points to an inode with `i_nlink == 0`. When syzkaller does `stat` on the directory, the kernel walks the path, creates a dentry, and drops the reference. Since d_count drops to zero, the dentry can go to the LRU and be targeted by kswapd0.
The key to the deadlock lies in the interaction between kswapd0 (the kernel's memory reclaim process) and the OCFS2's inode deletion process (`ocfs2_delete_inode`). kswapd0, in its process of freeing memory, might shrink the dcache and find this rogue dentry on the LRU. This corrupted dentry will be killed during reclaim, triggering the `ocfs2_delete_inode` function to be called. Crucially, that function attempts to acquire several locks in a specific order: `nfs_sync_rwlock`, `inode_lock(orphan_dir)`, and `inode_lock(inode_alloc)`.
The deadlock occurs if this corrupted dentry is reclaimed. At the same time, Syzkaller is doing a `mkdir`, which, in OCFS2, calls `kzalloc_obj` with `GFP_KERNEL` (via `ocfs2_reserve_new_metadata_blocks`), and since this flag uses direct reclaim, then the process needs to acquire the `fs_reclaim` lock. If the kernel is under memory pressure, kswapd0 will be running and might also hold the `fs_reclaim` lock. This would cause `mkdir` to block on a lock (inode_alloc) held by kswapd0 (via `ocfs2_delete_inode`) and then kswapd0 also tries to take a lock (inode_alloc) held by the same process as `mkdir`, which causes a deadlock.
This scenario is unlikely with a *normal* filesystem; however, with a corrupted image, this race happens. So, the deadlock is: `mkdir` -> `inode_lock(inode_alloc)` -> `fs_reclaim` -> `nfs_sync_rwlock` -> `inode_lock(orphan_dir)` -> `inode_lock(inode_alloc)`. Kswapd0 goes `fs_reclaim` -> `nfs_sync_rwlock`, `inode_lock(orphan_dir)`, and `inode_lock(inode_alloc)`. This is a circular dependency.
Even without kswapd0, a single process doing `mkdir` can self-deadlock. That happens because `mkdir` can enter direct reclaim and reclaim the corrupted dentry, which triggers a call to `ocfs2_delete_inode`.
The problem is the calls to `ocfs2_delete_inode` can happen from memory reclaim context (direct reclaim and kswapd0) which takes locks (`nfs_sync_rwlock`, `inode_alloc`) that can be held by processes that trigger memory reclaim.
The fix seems to be either preventing `ocfs2_delete_inode` from being called in memory reclaim context or ensuring OCFS2 allocates memory with `GFP_NOFS` (or, `memalloc_nofs_save()`) to avoid direct reclaim re-entering the filesystem. Another approach could be rejecting `i_nlink == 0` inodes during normal lookups.
|
| 39/2 |
2026/09/17 00:49 |
tool |
grepper |
0m
Args: map[Expression:memalloc_nofs_save PathPrefix:fs/ocfs2/]
Results: map[Output:fs/ocfs2/cluster/tcp.c=1579=static void o2net_start_connect(struct work_struct *work)
--
fs/ocfs2/cluster/tcp.c-1594- */
fs/ocfs2/cluster/tcp.c:1595: nofs_flag = memalloc_nofs_save();
fs/ocfs2/cluster/tcp.c-1596- /* if we're greater we initiate tx, otherwise we accept */
--
fs/ocfs2/cluster/tcp.c=1850=static int o2net_accept_one(struct socket *sock, int *more)
--
fs/ocfs2/cluster/tcp.c-1867- */
fs/ocfs2/cluster/tcp.c:1868: nofs_flag = memalloc_nofs_save();
fs/ocfs2/cluster/tcp.c-1869-
--
fs/ocfs2/quota_global.c=462=int ocfs2_global_write_info(struct super_block *sb, int type)
--
fs/ocfs2/quota_global.c-469- down_write(&dqopt->dqio_sem);
fs/ocfs2/quota_global.c:470: memalloc = memalloc_nofs_save();
fs/ocfs2/quota_global.c-471- err = ocfs2_qinfo_lock(info, 1);
--
fs/ocfs2/quota_global.c=617=static int ocfs2_sync_dquot_helper(struct dquot *dquot, unsigned long type)
--
fs/ocfs2/quota_global.c-641- down_write(&sb_dqopt(sb)->dqio_sem);
fs/ocfs2/quota_global.c:642: memalloc = memalloc_nofs_save();
fs/ocfs2/quota_global.c-643- status = ocfs2_sync_dquot(dquot);
--
fs/ocfs2/quota_global.c=683=static int ocfs2_write_dquot(struct dquot *dquot)
--
fs/ocfs2/quota_global.c-699- down_write(&sb_dqopt(dquot->dq_sb)->dqio_sem);
fs/ocfs2/quota_global.c:700: memalloc = memalloc_nofs_save();
fs/ocfs2/quota_global.c-701- status = ocfs2_local_write_dquot(dquot);
--
fs/ocfs2/quota_global.c=939=static int ocfs2_mark_dquot_dirty(struct dquot *dquot)
--
fs/ocfs2/quota_global.c-980- down_write(&sb_dqopt(sb)->dqio_sem);
fs/ocfs2/quota_global.c:981: memalloc = memalloc_nofs_save();
fs/ocfs2/quota_global.c-982- status = ocfs2_sync_dquot(dquot);
--
fs/ocfs2/quota_local.c=457=static int ocfs2_recover_local_quota_file(struct inode *lqinode,
--
fs/ocfs2/quota_local.c-523- down_write(&sb_dqopt(sb)->dqio_sem);
fs/ocfs2/quota_local.c:524: memalloc = memalloc_nofs_save();
fs/ocfs2/quota_local.c-525- spin_lock(&dquot->dq_dqb_lock);
]
|
| 40/2 |
2026/09/17 00:49 |
llm |
expert |
0m
Model:
Tokens:
input: 65064
output: 66
thoughts: 782
Thoughts: **My Analysis of an OCFS2 Deadlock Vulnerability**
Okay, so I've been looking at this potential deadlock issue within `ocfs2`. It seems that the file system uses `memalloc_nofs_save()` in some places, but crucially *not* in `ocfs2_mknod` or when reserving new metadata blocks. This is a red flag. What this means is that a call to `ocfs2_mknod` could potentially trigger direct reclaim, and that's taking `fs_reclaim`. This is a clear bug. The code should really be using `memalloc_nofs_save()` whenever it's holding filesystem locks, like within `inode_alloc`. Alternatively, using `GFP_NOFS` with `kzalloc_obj` would sidestep this entirely. But, hang on... looking at `kzalloc_obj`, it's using `default_gfp()`, which resolves to `GFP_KERNEL`. That's where the problem is. If `ocfs2` doesn't use `memalloc_nofs_save()`, it's wide open to this deadlock.
So, I'm confident the bug is real: a deadlock is possible and triggered by a malicious filesystem image.
Let's classify this:
- **Exploitable: No** - Not directly exploitable
- **DenialOfService: Yes** - Absolutely. It's a deadlock.
- **Unprivileged: No** - Requires elevated privileges.
- **UserNamespace: No** - It's not user namespace related.
- **VMGuestTrigger: No** - Not triggered from the guest.
- **VMHostTrigger: No** - Not triggered from the host.
- **NetworkTrigger: No** - Not network related.
- **RemoteTrigger: No** - Not remotely triggerable.
- **PeripheralTrigger: Yes** - A malicious USB drive can trigger it.
- **FilesystemTrigger: Yes** - Requires a corrupted image.
I should double check if an unprivileged user can mount `ocfs2`. I grepped for `FS_USERNS_MOUNT` in `fs/ocfs2/` and came up empty. So, it *can't* be mounted in a user namespace. That means `CAP_SYS_ADMIN` in the initial namespace is needed. An unprivileged user can't trigger it directly. However, a malicious USB drive might trigger it if auto-mounted.
Now, could there be another trigger? What if an unprivileged user could cause `ocfs2_delete_inode` to be called by kswapd0 on a "normal" filesystem?
Analyzing this further, `ocfs2_dentry_convert_worker` sets `OCFS2_INODE_MAYBE_ORPHANED` and unhashes the dentry. Unhashed dentries aren't on the LRU list. Thus, kswapd0 will *never* see them. Consider a directory being removed on another node. `ocfs2_dentry_convert_worker` will be called for that directory, set `OCFS2_INODE_MAYBE_ORPHANED`, call `shrink_dcache_parent(dentry)`, then `d_delete(dentry)`, and finally, `dput(dentry)`. `shrink_dcache_parent` will remove all child dentries and `d_delete` itself unhashes the dentry. The directory dentry itself is unhashed and is thus not on the LRU.
What if a file is unlinked on another node, but it's open locally? The dentry is unhashed, too, but remains positive, because `d_count > 1` (the open file holds a reference). When the file is closed, `dput` is called. Since the dentry is unhashed, `dput` calls `dentry_kill` immediately. It does not go to the LRU. Kswapd0 will *never* see it.
Therefore, it's impossible to trigger this on a regular filesystem! This requires a corrupted filesystem image to put a hashed dentry with `i_nlink == 0` on the LRU.
Therefore, my classification is correct. I'll now call `set-results`.
|
| 41/2 |
2026/09/17 00:49 |
tool |
set-results |
0m
Args: map[DenialOfService:true Exploitable:false FilesystemTrigger:true NetworkTrigger:false PeripheralTrigger:true RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
Results: map[DenialOfService:true Exploitable:false FilesystemTrigger:true NetworkTrigger:false PeripheralTrigger:true RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
|
| 42/2 |
2026/09/17 00:49 |
llm |
expert |
0m
Model:
Tokens:
input: 66907
output: 773
thoughts: 353
Thoughts: **Final Thoughts on the Bug**
Okay, the analysis is complete, and I'm ready to craft the final response. This comes down to a nasty circular dependency β a classic deadlock scenario β involving `fs_reclaim`, `nfs_sync_rwlock`, and `ocfs2_sysfile_lock_key`. The root cause is a bit convoluted. It's triggered when `kswapd0`, or direct reclaim if it's running, gets aggressive about shrinking the dcache and decides to evict a dentry. Specifically, it has the misfortune of picking a dentry whose inode has `i_nlink` equal to zero, which shouldn't even be on the LRU in the first place.
When that eviction happens, `evict` kicks off `ocfs2_delete_inode`. This function then tries to acquire both the `nfs_sync_rwlock` and the `inode_lock` on system inodes. Now, here's where the timing is critical. Simultaneously, a process performing something like a `mkdir` operation might be holding `inode_lock` on another system inode and, while allocating memory with `GFP_KERNEL`, enters direct reclaim and subsequently blocks, waiting for `fs_reclaim`.
Bam! Deadlock. `evict` is waiting for `inode_lock` which is held by a `mkdir` operation and the `mkdir` operation is waiting for `fs_reclaim` which `evict` is blocking.
The wrinkle here is that normally, dentries with a zero `i_nlink` should be unhashed and kept off the LRU, so they're invisible to `kswapd0`. But, if you have a malicious or corrupted ocfs2 filesystem image, it could contain a directory entry that points to an inode with this `i_nlink` zero. When that file is accessed, a positive dentry is created, hashed, and put onto the LRU. This "poisoned" entry is then visible to kswapd0.
It's bad, but it's not the worst. Fortunately, since ocfs2 doesn't support `FS_USERNS_MOUNT`, an unprivileged user can't mount this malicious image directly. So, the attack vector is somewhat limited β it requires a privileged user to mount the image, or a scenario where the image is auto-mounted from something like a peripheral device (a USB drive). The impact is a straightforward Denial of Service, which is still obviously a problem.
|
| 43/1 |
2026/09/17 00:49 |
action |
format-explanation |
0m
Results: map[Explanation:Based on the analysis of the bug report and the kernel source code, here is the
assessment of the vulnerability:
**Bug Mechanism and Impact:**
The lockdep warning indicates a circular locking dependency (deadlock) involving
the memory reclaim context (`fs_reclaim`), OCFS2's `nfs_sync_rwlock`, and the
`inode_lock` of OCFS2 system inodes (like `INODE_ALLOC_SYSTEM_INODE`).
The deadlock occurs due to two concurrent paths:
1. **Process A (e.g., `mkdir`)**: Holds the `inode_lock` of an OCFS2 system
inode (e.g., during `ocfs2_reserve_new_metadata_blocks`) and allocates memory
using `GFP_KERNEL` (via `kzalloc_obj`). This triggers direct memory reclaim,
causing the process to wait on the `fs_reclaim` pseudo-lock.
2. **Process B (`kswapd0` or another direct reclaimer)**: Holds the `fs_reclaim`
lock while shrinking the dcache. It finds a dentry on the LRU, kills it, and
drops the last reference to its inode. If the inode has `i_nlink == 0`,
`evict()` synchronously calls `ocfs2_delete_inode()`. This function attempts to
acquire `nfs_sync_rwlock` and the `inode_lock` of the system inodes, blocking on
Process A.
This results in a classic AB-BA deadlock, leading to a Denial of Service (system
hang).
**Triggerability:**
For `kswapd0` to call `ocfs2_delete_inode()`, it must find a dentry on the LRU
whose inode has `i_nlink == 0` (or `OCFS2_INODE_MAYBE_ORPHANED`).
Under normal filesystem operations, when a file is unlinked, its dentry is
immediately unhashed by the VFS (`d_delete()`). Unhashed dentries are never
placed on the LRU, meaning `kswapd0` will never encounter them. Similarly, OCFS2
DLM downconvert workers unhash dentries when a file is deleted on another node.
The *only* way a hashed dentry with `i_nlink == 0` can end up on the LRU is
through a **malicious or corrupted filesystem image**. If an OCFS2 image is
crafted such that a directory entry points to an inode with `i_nlink == 0`, a
simple `stat` or `open` will instantiate a positive, hashed dentry. When the
reference is dropped, it goes to the LRU. Subsequent memory pressure will wake
up `kswapd0`, which will kill the dentry and trigger the deadlock.
**Security Boundaries:**
- **Unprivileged / User Namespaces:** OCFS2 does not support `FS_USERNS_MOUNT`.
Therefore, an unprivileged user cannot mount the malicious image, even within a
user namespace. It requires `CAP_SYS_ADMIN` in the initial namespace.
- **Peripheral / Filesystem Trigger:** The bug can be triggered if a malicious
OCFS2 image is introduced via a physical peripheral (e.g., a USB drive) and is
auto-mounted by a desktop environment. Once mounted, any unprivileged access
(like an auto-indexer reading the drive) combined with memory pressure will
trigger the deadlock.
**Conclusion:**
This is a Denial of Service vulnerability triggered by mounting and interacting
with a maliciously crafted OCFS2 filesystem image. It is not exploitable for
privilege escalation or information leaks.]
|