| 0/0 |
2026/02/24 14:34 |
flow |
assessment-kcsan |
28m
Results: map[Benign:true Confident:true Explanation:The data race occurs between a read from a block device (via `blkdev_read_iter` and `filemap_read`) and a write to the same block device's page cache by the FAT filesystem (via `fat_mirror_bhs`).
### Analysis
1. **Nature of the Accesses**:
* **Reader**: A user process is reading the raw block device (e.g., a loop device). This path uses `filemap_read`, which copies data from the block device's page cache to userspace.
* **Writer**: The FAT filesystem is updating its File Allocation Table (FAT) mirrors. When the primary FAT is updated, `fat_mirror_bhs` copies the changes to the backup FAT(s) using `memcpy` on the buffer head's data.
2. **Synchronization**:
* The FAT filesystem uses `fat_lock` (a mutex) to synchronize its own metadata updates, ensuring that the filesystem remains consistent from its own perspective.
* In `fat_mirror_bhs`, the code uses `lock_buffer(c_bh)` before calling `memcpy`. This synchronizes with other kernel components that use buffer heads and respect the `BH_Lock` bit (e.g., `sb_bread` or other metadata operations).
* However, `filemap_read` (used for block device reads) does not check the `BH_Lock` bit on buffer heads if the underlying page (folio) is already marked as uptodate. It simply copies the data from the page cache.
* The comment in `fs/fat/fatent.c` at line 393 (`/* Avoid race with userspace read via bdev */`) indicates that the developers were aware of this potential race and attempted to mitigate it with `lock_buffer`. While `lock_buffer` is insufficient to block `filemap_read`, it is the standard way to protect buffer head data in the kernel.
3. **Impact**:
* **Data Integrity**: The only "victim" of this race is the user process performing a raw read of the block device. It might see a "torn" FAT sector (a mix of old and new FAT entries). This is expected behavior when reading a raw block device that is concurrently being modified by a mounted filesystem.
* **Kernel Stability**: The race involves a simple `memcpy` to a buffer and a `copy_to_user` from the same buffer. Neither operation can cause a kernel crash, memory corruption, or invalid memory access, as the buffer boundaries are fixed and managed by the block device layer.
* **Filesystem Consistency**: The filesystem's own consistency is not affected because the primary FAT update and the mirror updates are coordinated by the filesystem's internal locks. The data on disk will eventually be consistent when the page cache is flushed.
### Conclusion
This data race is **benign**. It is a known characteristic of the Linux block device layer that raw reads are not fully synchronized with filesystem-level metadata updates in the page cache. Such races do not lead to kernel instability or filesystem corruption; they only affect the consistency of data returned to a raw block device reader, which is already an inherently "racy" operation when the filesystem is mounted.]
|
| 1/1 |
2026/02/24 14:34 |
action |
kernel-checkouter |
0m
Results: map[KernelSrc:/usr/local/google/home/dvyukov/syzkaller/agent/workdir/cache/src/5f2d4fec15c6db42e498645e435c2de8e147d8c9]
|
| 2/1 |
2026/02/24 14:35 |
action |
kernel-builder |
4m
Results: map[KernelObj:/usr/local/google/home/dvyukov/syzkaller/agent/workdir/cache/build/d6cdd6e8e7c74d319baf487ac137c044830934b4]
|
| 3/1 |
2026/02/24 14:39 |
action |
codesearch-prepare |
20m
Results: map[Index:codesearch-index]
|
| 4/1 |
2026/02/24 15:00 |
agent |
expert |
2m
Model:
Results: map[Benign:true Confident:true]
Instruction:
You are an experienced Linux kernel developer tasked with determining if the given kernel
data race is benign or not. The data race report is from KCSAN tool.
It contains 2 stack traces of the memory accesses that constitute a data race.
A "benign" data races are on a simple int/bool variable or similar field,
and the accesses are not supposed to be protected by any mutual exclusion primitives.
Common examples of such "benign" data races are accesses to various flags fields,
statistics counters, and similar. A "benign" data race does not lead to memory corruption/crash
with a conservative compiler that compiles memory accesses to primitive types
effectively as atomic.
A non-benign (or "harmful" data race) can lead to corruption/crash even with
a conservative compiler that compiles memory accesses to primitive types
effectively as atomic. A common example of a "harmful" data races is race on
a complex container (list/hashmap/etc), where accesses are supposed to be protected
by a mutual exclusion primitive.
Also consider races that happen at the same time with the given one.
If there is no synchronization in between, other memory accesses in the involved threads
race with each other if they access the same memory. For example, if both threads execute:
some_struct->have_elements = true;
list_add(new_node, &some_struct->list_head);
the race on some_struct->have_elements may appear benign, however it also implies there
is a race on some_struct->list_head which is not benign, since the list is not thread-safe.
Take into account that on 32-bit systems 64-bit memory accesses may be split into two accesses,
and thus even with a conservative compiler may not be fully atomic. However, such races may
still be benign depending on how writes are done, and how read data is used.
In the final reply explain why you think the given data race is benign or is harmful.
Use the provided tools to confirm any assumptions, variables/fields being accessed, etc.
In particular, don't make assumptions about the kernel source code,
use codesearch tools to read the actual source code.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt:
The data race report is:
loop3: detected capacity change from 0 to 128
==================================================================
BUG: KCSAN: data-race in _copy_to_iter / fat_mirror_bhs
read to 0xffff88814b8c7000 of 1536 bytes by task 10167 on cpu 0:
instrument_copy_to_user include/linux/instrumented.h:113 [inline]
copy_to_user_iter lib/iov_iter.c:24 [inline]
iterate_ubuf include/linux/iov_iter.h:30 [inline]
iterate_and_advance2 include/linux/iov_iter.h:302 [inline]
iterate_and_advance include/linux/iov_iter.h:330 [inline]
_copy_to_iter+0x130/0xea0 lib/iov_iter.c:197
copy_page_to_iter+0x18f/0x2d0 lib/iov_iter.c:374
copy_folio_to_iter include/linux/uio.h:204 [inline]
filemap_read+0x407/0xa10 mm/filemap.c:2851
blkdev_read_iter+0x22d/0x2e0 block/fops.c:856
new_sync_read fs/read_write.c:491 [inline]
vfs_read+0x6c0/0x7f0 fs/read_write.c:572
ksys_read+0xdc/0x1a0 fs/read_write.c:715
__do_sys_read fs/read_write.c:724 [inline]
__se_sys_read fs/read_write.c:722 [inline]
__x64_sys_read+0x40/0x50 fs/read_write.c:722
x64_sys_call+0x2889/0x3000 arch/x86/include/generated/asm/syscalls_64.h:1
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0xc0/0x2a0 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
write to 0xffff88814b8c7000 of 2048 bytes by task 16943 on cpu 1:
fat_mirror_bhs+0x1df/0x320 fs/fat/fatent.c:395
fat_alloc_clusters+0xb48/0xc50 fs/fat/fatent.c:543
fat_add_cluster fs/fat/inode.c:108 [inline]
__fat_get_block fs/fat/inode.c:155 [inline]
fat_get_block+0x258/0x5e0 fs/fat/inode.c:190
__block_write_begin_int+0x400/0xf90 fs/buffer.c:2145
block_write_begin fs/buffer.c:2256 [inline]
cont_write_begin+0x5fe/0x970 fs/buffer.c:2594
fat_write_begin+0x4f/0xe0 fs/fat/inode.c:230
generic_perform_write+0x183/0x490 mm/filemap.c:4314
__generic_file_write_iter+0x9e/0x120 mm/filemap.c:4431
generic_file_write_iter+0x8d/0x310 mm/filemap.c:4457
new_sync_write fs/read_write.c:593 [inline]
vfs_write+0x5a6/0x9f0 fs/read_write.c:686
ksys_write+0xdc/0x1a0 fs/read_write.c:738
__do_sys_write fs/read_write.c:749 [inline]
__se_sys_write fs/read_write.c:746 [inline]
__x64_sys_write+0x40/0x50 fs/read_write.c:746
x64_sys_call+0x2847/0x3000 arch/x86/include/generated/asm/syscalls_64.h:2
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0xc0/0x2a0 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
Reported by Kernel Concurrency Sanitizer on:
CPU: 1 UID: 0 PID: 16943 Comm: syz.3.4099 Tainted: G W syzkaller #0 PREEMPT(voluntary)
Tainted: [W]=WARN
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 10/25/2025
==================================================================
loop2: detected capacity change from 0 to 7
Reply: The data race occurs between a read from a block device (via `blkdev_read_iter` and `filemap_read`) and a write to the same block device's page cache by the FAT filesystem (via `fat_mirror_bhs`).
### Analysis
1. **Nature of the Accesses**:
* **Reader**: A user process is reading the raw block device (e.g., a loop device). This path uses `filemap_read`, which copies data from the block device's page cache to userspace.
* **Writer**: The FAT filesystem is updating its File Allocation Table (FAT) mirrors. When the primary FAT is updated, `fat_mirror_bhs` copies the changes to the backup FAT(s) using `memcpy` on the buffer head's data.
2. **Synchronization**:
* The FAT filesystem uses `fat_lock` (a mutex) to synchronize its own metadata updates, ensuring that the filesystem remains consistent from its own perspective.
* In `fat_mirror_bhs`, the code uses `lock_buffer(c_bh)` before calling `memcpy`. This synchronizes with other kernel components that use buffer heads and respect the `BH_Lock` bit (e.g., `sb_bread` or other metadata operations).
* However, `filemap_read` (used for block device reads) does not check the `BH_Lock` bit on buffer heads if the underlying page (folio) is already marked as uptodate. It simply copies the data from the page cache.
* The comment in `fs/fat/fatent.c` at line 393 (`/* Avoid race with userspace read via bdev */`) indicates that the developers were aware of this potential race and attempted to mitigate it with `lock_buffer`. While `lock_buffer` is insufficient to block `filemap_read`, it is the standard way to protect buffer head data in the kernel.
3. **Impact**:
* **Data Integrity**: The only "victim" of this race is the user process performing a raw read of the block device. It might see a "torn" FAT sector (a mix of old and new FAT entries). This is expected behavior when reading a raw block device that is concurrently being modified by a mounted filesystem.
* **Kernel Stability**: The race involves a simple `memcpy` to a buffer and a `copy_to_user` from the same buffer. Neither operation can cause a kernel crash, memory corruption, or invalid memory access, as the buffer boundaries are fixed and managed by the block device layer.
* **Filesystem Consistency**: The filesystem's own consistency is not affected because the primary FAT update and the mirror updates are coordinated by the filesystem's internal locks. The data on disk will eventually be consistent when the page cache is flushed.
### Conclusion
This data race is **benign**. It is a known characteristic of the Linux block device layer that raw reads are not fully synchronized with filesystem-level metadata updates in the page cache. Such races do not lead to kernel instability or filesystem corruption; they only affect the consistency of data returned to a raw block device reader, which is already an inherently "racy" operation when the filesystem is mounted.
|
| 5/2 |
2026/02/24 15:00 |
llm |
expert |
0m
Model:
Tokens:
input: 4414
output: 43
thoughts: 355
Thoughts: **Data Race Deep Dive**
Okay, so the report's flashing red, pointing to a data race. My initial read of it suggests a standard read operation from a block device clashing with a write on a FAT filesystem. Specifically, the read, triggered through `_copy_to_iter`, trickles down from `filemap_read`, through `blkdev_read_iter` and finally `vfs_read`. This makes sense, a standard read path for block devices.
But then there's the write. It's happening within `fat_mirror_bhs`, which I immediately recognize as FAT metadata manipulation – the FAT's mirrors, to be precise. The call stack, leading through `fat_write_begin`, confirms it is metadata being updated. The write goes through `fat_alloc_clusters`, `fat_add_cluster` and finally to `fat_mirror_bhs`.
The real kicker is the address and byte count. We have a read of 1536 bytes at `0xffff88814b8c7000` going on at the same time as a 2048-byte write. This means something is reading the block device while FAT is updating its metadata, which is incredibly dangerous. `blkdev_read_iter` indicates the read is coming from the block device level. That is the issue.
My intuition's kicking in; the problem is almost certainly within `fat_mirror_bhs` or its call chain. It's touching the block device directly, while a concurrent read attempts to pull data. I'm going to jump into `fs/fat/fatent.c` and take a closer look at `fat_mirror_bhs` and trace the context of the underlying block writes.
|
| 6/2 |
2026/02/24 15:00 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/fat/fatent.c IncludeLines:true Name:fat_mirror_bhs]
Results: map[SourceCode: 375: /* FIXME: We can write the blocks as more big chunk. */
376: static int fat_mirror_bhs(struct super_block *sb, struct buffer_head **bhs,
377: int nr_bhs)
378: {
379: struct msdos_sb_info *sbi = MSDOS_SB(sb);
380: struct buffer_head *c_bh;
381: int err, n, copy;
382:
383: err = 0;
384: for (copy = 1; copy < sbi->fats; copy++) {
385: sector_t backup_fat = sbi->fat_length * copy;
386:
387: for (n = 0; n < nr_bhs; n++) {
388: c_bh = sb_getblk(sb, backup_fat + bhs[n]->b_blocknr);
389: if (!c_bh) {
390: err = -ENOMEM;
391: goto error;
392: }
393: /* Avoid race with userspace read via bdev */
394: lock_buffer(c_bh);
395: memcpy(c_bh->b_data, bhs[n]->b_data, sb->s_blocksize);
396: set_buffer_uptodate(c_bh);
397: unlock_buffer(c_bh);
398: mark_buffer_dirty_inode(c_bh, sbi->fat_inode);
399: if (sb->s_flags & SB_SYNCHRONOUS)
400: err = sync_dirty_buffer(c_bh);
401: brelse(c_bh);
402: if (err)
403: goto error;
404: }
405: }
406: error:
407: return err;
408: }
SourceFile:fs/fat/fatent.c]
|
| 7/2 |
2026/02/24 15:00 |
llm |
expert |
0m
Model:
Tokens:
input: 5773
output: 41
thoughts: 465
Thoughts: **Analyzing a Potential Race Condition in the FAT Filesystem's Mirroring**
Okay, so I'm looking at this FAT filesystem code and the comment specifically highlights a race condition with userspace reads via the block device (`bdev`). The code snippet, using `lock_buffer(c_bh)` is supposed to protect against concurrent access. However, my immediate thought is that this isn't sufficient. The locking mechanism, as far as I can see, only protects accesses to the *same* `buffer_head`. If a userspace program is reading directly from the block device using, for example, `blkdev_read_iter`, it could bypass this `lock_buffer` entirely. `blkdev_read_iter` will likely be using the page cache of the block device directly, which is a different access path.
The crucial link here is the page cache. The FAT filesystem utilizes the page cache for its metadata, including the FAT table itself. `sb_getblk` returns a `buffer_head`, which I'm sure is associated with a specific page within the block device's address space, right? If `blkdev_read_iter` is reading from the block device (and therefore, the page cache), it's reading the *same* pages that `fat_mirror_bhs` is writing to during mirroring.
So, `blkdev_read_iter`'s read operation (ultimately calling `filemap_read` then `copy_folio_to_iter`) is hitting the page, and `fat_mirror_bhs` is writing to that same page using `memcpy` with the `buffer_head`'s `b_data` pointer. The concern boils down to this: does `filemap_read` (or the subsequent `copy_folio_to_iter`) check the `BH_Lock` bit that is set by `lock_buffer(c_bh)`? No, from what I know, it doesn't directly check that. It might check `PageLocked(page)` but `lock_buffer` doesn't necessarily lock the *page* itself, only the `buffer_head`.
I think I need to dig deeper into the implementation of `lock_buffer` to confirm that. This is starting to look like a genuine race condition that isn't fully addressed by the current locking strategy. I need to be sure that the page cache coherence is properly maintained here, especially in light of the potential for concurrent reads and writes from both the filesystem and direct block device access.
|
| 8/2 |
2026/02/24 15:00 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:include/linux/buffer_head.h IncludeLines:true Name:lock_buffer]
Results: map[SourceCode: 428: static inline void lock_buffer(struct buffer_head *bh)
429: {
430: might_sleep();
431: if (!trylock_buffer(bh))
432: __lock_buffer(bh);
433: }
SourceFile:include/linux/buffer_head.h]
|
| 9/2 |
2026/02/24 15:00 |
llm |
expert |
0m
Model:
Tokens:
input: 6926
output: 39
thoughts: 910
Thoughts: **Analyzing a Potential Data Race in FAT Filesystem Metadata Updates**
Okay, so I'm looking at how the FAT filesystem mirrors its FAT tables, specifically within `fat_mirror_bhs`. It appears the code uses `lock_buffer`, but it's important to understand *why*. The original intent, based on the comment, seems to be to avoid data races with userspace reads directly from the block device. However, `lock_buffer` itself, via the `trylock_buffer`, doesn't guarantee protection from concurrent reads, especially if those reads use `blkdev_read_iter`.
My initial thought was, "Well, so what? If a user is reading the *raw* block device, they're responsible for handling potential inconsistencies during writes. That's inherent to block device operation." If the FAT filesystem is updating its metadata on the block device, a direct read might see an intermediate, potentially inconsistent, state. But KCSAN is reporting something, so there's more to consider here.
Let's dissect the read path more carefully. We're looking at `blkdev_read_iter` which calls `filemap_read`, which then goes down the path to `copy_folio_to_iter` and finally `_copy_to_iter`. So it's not a *direct* read; it's going through the page cache. This means we're dealing with buffer heads associated with the block device's address space that have been read into memory, and then written back.
Within `fat_mirror_bhs`, the key section is the `memcpy` which is copying data from one buffer head (`bhs[n]`) to another (`c_bh`). `sb_getblk` is being used to get the buffer heads. This means the buffer heads are associated with the block device. This code is doing this:
```c
388: c_bh = sb_getblk(sb, backup_fat + bhs[n]->b_blocknr);
...
394: lock_buffer(c_bh);
395: memcpy(c_bh->b_data, bhs[n]->b_data, sb->s_blocksize);
396: set_buffer_uptodate(c_bh);
397: unlock_buffer(c_bh);
```
My concern is whether this `memcpy` operation poses a problem. In a block device context, concurrent reads and writes are generally allowed, and that userspace should handle synchronization if they care about consistency. However, in the *kernel*, any potential for corruption needs a closer look.
The crucial question is, if this `memcpy` isn't atomic, could the reader see torn writes? With direct block device reads, torn writes are usually acceptable as a consequence of concurrency.
The context of this `memcpy` is mirroring the FAT tables. FAT filesystems have multiple copies to guard against corruption. `fat_mirror_bhs` mirrors the changes to all the copies.
Now, that comment, `/* Avoid race with userspace read via bdev */`, is nagging me. Why did they include `lock_buffer`? I still don't see how it prevents a race in this scenario. If the intention was race protection, they should have been locking the page, not just the buffer! And with `blkdev_read_iter`, it is indeed calling into the page cache, which *does* lock the page.
So, let's go back and look at that `filemap_read` again. I need to be sure about the page locking involved.
|
| 10/2 |
2026/02/24 15:01 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:mm/filemap.c IncludeLines:true Name:filemap_read]
Results: map[SourceCode:2750: /**
2751: * filemap_read - Read data from the page cache.
2752: * @iocb: The iocb to read.
2753: * @iter: Destination for the data.
2754: * @already_read: Number of bytes already read by the caller.
2755: *
2756: * Copies data from the page cache. If the data is not currently present,
2757: * uses the readahead and read_folio address_space operations to fetch it.
2758: *
2759: * Return: Total number of bytes copied, including those already read by
2760: * the caller. If an error happens before any bytes are copied, returns
2761: * a negative error number.
2762: */
2763: ssize_t filemap_read(struct kiocb *iocb, struct iov_iter *iter,
2764: ssize_t already_read)
2765: {
2766: struct file *filp = iocb->ki_filp;
2767: struct file_ra_state *ra = &filp->f_ra;
2768: struct address_space *mapping = filp->f_mapping;
2769: struct inode *inode = mapping->host;
2770: struct folio_batch fbatch;
2771: int i, error = 0;
2772: bool writably_mapped;
2773: loff_t isize, end_offset;
2774: loff_t last_pos = ra->prev_pos;
2775:
2776: if (unlikely(iocb->ki_pos < 0))
2777: return -EINVAL;
2778: if (unlikely(iocb->ki_pos >= inode->i_sb->s_maxbytes))
2779: return 0;
2780: if (unlikely(!iov_iter_count(iter)))
2781: return 0;
2782:
2783: iov_iter_truncate(iter, inode->i_sb->s_maxbytes - iocb->ki_pos);
2784: folio_batch_init(&fbatch);
2785:
2786: do {
2787: cond_resched();
2788:
2789: /*
2790: * If we've already successfully copied some data, then we
2791: * can no longer safely return -EIOCBQUEUED. Hence mark
2792: * an async read NOWAIT at that point.
2793: */
2794: if ((iocb->ki_flags & IOCB_WAITQ) && already_read)
2795: iocb->ki_flags |= IOCB_NOWAIT;
2796:
2797: if (unlikely(iocb->ki_pos >= i_size_read(inode)))
2798: break;
2799:
2800: error = filemap_get_pages(iocb, iter->count, &fbatch, false);
2801: if (error < 0)
2802: break;
2803:
2804: /*
2805: * i_size must be checked after we know the pages are Uptodate.
2806: *
2807: * Checking i_size after the check allows us to calculate
2808: * the correct value for "nr", which means the zero-filled
2809: * part of the page is not copied back to userspace (unless
2810: * another truncate extends the file - this is desired though).
2811: */
2812: isize = i_size_read(inode);
2813: if (unlikely(iocb->ki_pos >= isize))
2814: goto put_folios;
2815: end_offset = min_t(loff_t, isize, iocb->ki_pos + iter->count);
2816:
2817: /*
2818: * Once we start copying data, we don't want to be touching any
2819: * cachelines that might be contended:
2820: */
2821: writably_mapped = mapping_writably_mapped(mapping);
2822:
2823: /*
2824: * When a read accesses the same folio several times, only
2825: * mark it as accessed the first time.
2826: */
2827: if (!pos_same_folio(iocb->ki_pos, last_pos - 1,
2828: fbatch.folios[0]))
2829: folio_mark_accessed(fbatch.folios[0]);
2830:
2831: for (i = 0; i < folio_batch_count(&fbatch); i++) {
2832: struct folio *folio = fbatch.folios[i];
2833: size_t fsize = folio_size(folio);
2834: size_t offset = iocb->ki_pos & (fsize - 1);
2835: size_t bytes = min_t(loff_t, end_offset - iocb->ki_pos,
2836: fsize - offset);
2837: size_t copied;
2838:
2839: if (end_offset < folio_pos(folio))
2840: break;
2841: if (i > 0)
2842: folio_mark_accessed(folio);
2843: /*
2844: * If users can be writing to this folio using arbitrary
2845: * virtual addresses, take care of potential aliasing
2846: * before reading the folio on the kernel side.
2847: */
2848: if (writably_mapped)
2849: flush_dcache_folio(folio);
2850:
2851: copied = copy_folio_to_iter(folio, offset, bytes, iter);
2852:
2853: already_read += copied;
2854: iocb->ki_pos += copied;
2855: last_pos = iocb->ki_pos;
2856:
2857: if (copied < bytes) {
2858: error = -EFAULT;
2859: break;
2860: }
2861: }
2862: put_folios:
2863: for (i = 0; i < folio_batch_count(&fbatch); i++) {
2864: struct folio *folio = fbatch.folios[i];
2865:
2866: filemap_end_dropbehind_read(folio);
2867: folio_put(folio);
2868: }
2869: folio_batch_init(&fbatch);
2870: } while (iov_iter_count(iter) && iocb->ki_pos < isize && !error);
2871:
2872: file_accessed(filp);
2873: ra->prev_pos = last_pos;
2874: return already_read ? already_read : error;
2875: }
SourceFile:mm/filemap.c]
|
| 11/2 |
2026/02/24 15:01 |
llm |
expert |
0m
Model:
Tokens:
input: 10657
output: 41
thoughts: 568
Thoughts: **Analyzing a Potential Race Condition in the FAT Filesystem's Mirror Update**
Okay, so I'm looking at how the FAT filesystem handles mirror updates, specifically focusing on a potential race condition. Here's what's going through my mind:
I've noticed that `filemap_read` doesn't lock the folio before calling `copy_folio_to_iter`. It retrieves folios through `filemap_get_pages`, which conveniently returns folios that are up-to-date. In the context of `fat_mirror_bhs`, where we're writing to the buffer head data associated with a page (folio) on the block device. It locks the buffer head, but not the page itself.
Now, if a user's reading directly from the block device, they're essentially reading the same page that's undergoing these mirror updates. Since neither side is locking the page itself, there's a potential race. The question is, does this matter?
Generally, concurrent reads and writes are fine at the block device level. Users might get inconsistent data if they read while a write is in progress, but that's a normal expectation for block devices. But, we're dealing with a filesystem, and the FAT filesystem *does* care about metadata consistency. However, in this case, `fat_mirror_bhs` is *just* updating the mirrors of the FAT table; the primary FAT table has already been updated.
So, if a user reads the mirror *while* it's being updated, they might see a partially written mirror. Is this a kernel stability issue? No, I don't think so. Data integrity? Well, if they're bypassing the filesystem and reading the block device directly, they are already accepting the lack of consistency guarantees. If they want consistent data, they probably shouldn't be reading the device directly while it's mounted.
But wait, what about that comment in `fat_mirror_bhs` that says, "Avoid race with userspace read via bdev" with a `lock_buffer(c_bh)` call. It seems like someone *thought* `lock_buffer` would prevent a race with a user reading via the block device. But as I saw, `blkdev_read_iter` seems to be ignoring `lock_buffer`.
I need to confirm if `blkdev_read_iter` or `filemap_read` checks the `buffer_locked(bh)` status. `filemap_read` operates on folios, and if folios have buffer heads, does it actually care about whether those buffer heads are locked? In most cases, `filemap_read` doesn't usually concern itself with buffer heads.
I should probably dive into `blkdev_read_iter` to see if it does anything special with buffer locks. Let's see...
|
| 12/2 |
2026/02/24 15:01 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:block/fops.c IncludeLines:true Name:blkdev_read_iter]
Results: map[SourceCode: 812: static ssize_t blkdev_read_iter(struct kiocb *iocb, struct iov_iter *to)
813: {
814: struct inode *bd_inode = bdev_file_inode(iocb->ki_filp);
815: struct block_device *bdev = I_BDEV(iocb->ki_filp->f_mapping->host);
816: loff_t size = bdev_nr_bytes(bdev);
817: loff_t pos = iocb->ki_pos;
818: size_t shorted = 0;
819: ssize_t ret = 0;
820: size_t count;
821:
822: if (unlikely(pos + iov_iter_count(to) > size)) {
823: if (pos >= size)
824: return 0;
825: size -= pos;
826: shorted = iov_iter_count(to) - size;
827: iov_iter_truncate(to, size);
828: }
829:
830: count = iov_iter_count(to);
831: if (!count)
832: goto reexpand; /* skip atime */
833:
834: if (iocb->ki_flags & IOCB_DIRECT) {
835: ret = kiocb_write_and_wait(iocb, count);
836: if (ret < 0)
837: goto reexpand;
838: file_accessed(iocb->ki_filp);
839:
840: ret = blkdev_direct_IO(iocb, to);
841: if (ret > 0) {
842: iocb->ki_pos += ret;
843: count -= ret;
844: }
845: if (ret != -EIOCBQUEUED)
846: iov_iter_revert(to, count - iov_iter_count(to));
847: if (ret < 0 || !count)
848: goto reexpand;
849: }
850:
851: /*
852: * Take i_rwsem and invalidate_lock to avoid racing with set_blocksize
853: * changing i_blkbits/folio order and punching out the pagecache.
854: */
855: inode_lock_shared(bd_inode);
856: ret = filemap_read(iocb, to, ret);
857: inode_unlock_shared(bd_inode);
858:
859: reexpand:
860: if (unlikely(shorted))
861: iov_iter_reexpand(to, iov_iter_count(to) + shorted);
862: return ret;
863: }
SourceFile:block/fops.c]
|
| 13/2 |
2026/02/24 15:01 |
llm |
expert |
0m
Model:
Tokens:
input: 12721
output: 23
thoughts: 1852
Thoughts: **Analyzing a Potential Data Race in FAT Filesystem Block Device Reads**
Okay, let's break this down. First, it seems like we've identified a potential data race involving `inode_lock_shared(bd_inode)` in the FAT filesystem code. My initial thought is to trace the execution path. We're looking at `fat_mirror_bhs` which is called from several functions related to FAT metadata updates: `fat_alloc_clusters`, `fat_free_clusters`, and `fat_ent_write`. This hints at writes to the FAT table itself. The interesting part is, where is the lock taken? The call stack shows `inode_lock(inode)` is acquired in `generic_file_write_iter`, but that's the *file* inode, not the block device inode. It's the block device inode (`bd_inode`) that is the target of the read operation.
The core of the issue appears to be a read operation on the block device page cache concurrent with `fat_mirror_bhs` writing to that same page cache. `fat_mirror_bhs` uses `sb_getblk` which interacts with the block device's address space. A user process is reading via `ksys_read`.
Generally, raw block device reads concurrent with a live filesystem can lead to inconsistent data, and it is usually a "don't do that" situation. The kernel itself, however, has to be watertight. The question is, is this benign, or is there a way this could lead to a kernel crash? The concern is `memcpy` or `copy_to_user_iter` with potentially modified source data. `copy_to_user_iter` eventually calls `copy_to_user`. Getting inconsistent data to userspace is bad, but it shouldn't crash the kernel.
So, let's go back to `fat_mirror_bhs`. It's mirroring the FAT, which is critical. If the FAT gets corrupted, the filesystem is toast. But the reader is not writing to the FAT, and the writer (`fat_mirror_bhs`) is writing correct data. This suggests that the reader is the "victim" here, and inconsistent data is expected for a raw block device read.
However, KCSAN flagged this. Hmm, is `fat_mirror_bhs` racing with *another* kernel thread also using the block device? The reader is `ksys_read` from a user process. I need to think about locks. Is `fat_mirror_bhs` called with a lock that the reader *should* have taken? Nope. The reader is reading the block device raw, and the writer is updating FAT metadata. These are independent.
I'm starting to consider if `fat_mirror_bhs` is called anywhere else. It is used to update the FAT table. If this is truly benign, why the comment and `lock_buffer` in `fat_mirror_bhs`? Was that developer trying to avoid a race they *thought* was important?
I need to understand what `lock_buffer` really does in relation to how `blkdev_read_iter` and `filemap_read` work. `lock_buffer` will block any concurrent access using the same `buffer_head` and also calling `lock_buffer` or `wait_on_buffer`, but it looks like `blkdev_read_iter` uses `filemap_read`, which copies from the page. If the page is uptodate, it just copies the data. Otherwise, it calls `read_folio`, which might use `buffer_head`, but only if the page is not uptodate.
In `fat_mirror_bhs`, the `memcpy` is followed by `set_buffer_uptodate`. If the buffer *was* uptodate, then `set_buffer_uptodate` is a no-op. If it *wasn't* uptodate, then another thread might see it as not uptodate and wait, but `filemap_read` checks `FolioUptodate` and that's the page, and the page becomes uptodate when all buffers are. If the page is already uptodate, `filemap_read` will just copy it.
I also need to see if there's any other synchronization mechanism. FAT uses `fat_lock` (a mutex in `msdos_sb_info`) to protect FAT updates. But the raw block device read doesn't know about `fat_lock`.
If this is just the user getting inconsistent data from raw block device reads, this usually is benign.
Let's look at `fat_mirror_bhs` again. I see a `lock_buffer` and `unlock_buffer`. If the reader was
|
| 14/2 |
2026/02/24 15:01 |
tool |
grepper |
0m
Args: map[Expression:Avoid race with userspace read via bdev]
Results: map[Output:fs/fat/dir.c=1093=static int fat_zeroed_cluster(struct inode *dir, sector_t blknr, int nr_used,
--
fs/fat/dir.c-1108- }
fs/fat/dir.c:1109: /* Avoid race with userspace read via bdev */
fs/fat/dir.c-1110- lock_buffer(bhs[n]);
--
fs/fat/dir.c=1145=int fat_alloc_new_dir(struct inode *dir, struct timespec64 *ts)
--
fs/fat/dir.c-1169- de = (struct msdos_dir_entry *)bhs[0]->b_data;
fs/fat/dir.c:1170: /* Avoid race with userspace read via bdev */
fs/fat/dir.c-1171- lock_buffer(bhs[0]);
--
fs/fat/dir.c=1210=static int fat_add_new_entries(struct inode *dir, void *slots, int nr_slots,
--
fs/fat/dir.c-1251- copy = min(size, sb->s_blocksize);
fs/fat/dir.c:1252: /* Avoid race with userspace read via bdev */
fs/fat/dir.c-1253- lock_buffer(bhs[n]);
--
fs/fat/fatent.c=376=static int fat_mirror_bhs(struct super_block *sb, struct buffer_head **bhs,
--
fs/fat/fatent.c-392- }
fs/fat/fatent.c:393: /* Avoid race with userspace read via bdev */
fs/fat/fatent.c-394- lock_buffer(c_bh);
]
|
| 15/2 |
2026/02/24 15:01 |
llm |
expert |
0m
Model:
Tokens:
input: 16094
output: 42
thoughts: 931
Thoughts: **Analyzing Potential Race Conditions in FAT Metadata Updates**
Okay, so I'm seeing a pattern in the FAT filesystem where we have `lock_buffer(bh)`, a `memcpy`, and then `unlock_buffer(bh)`. It's used in a few places. The apparent intent is to prevent races with userspace reads via the block device. But KCSAN is still flagging a race. So either `lock_buffer` isn't the right tool, or `blkdev_read_iter` needs to check if the buffer is locked.
Wait a second. If `blkdev_read_iter` is using the page cache, it's reading the page itself, and buffer heads are just metadata attached to it. `lock_buffer` just modifies a bit in the `buffer_head`; it doesn't affect the actual page. To truly avoid races with page cache reads, someone would need to lock the page, not the buffer head. But `fat_mirror_bhs` is working with these buffer heads.
Can I lock the page from within `fat_mirror_bhs`? The buffer head, `c_bh`, points to the page with `c_bh->b_page` (or in newer kernels, `c_bh->b_folio->page`). If `fat_mirror_bhs` were to use `lock_page(c_bh->b_page)`, would it block `filemap_read`?
Let me look at `filemap_read`. I don't see `folio_lock` being called for every folio read. It looks like it only locks the folio if it needs to wait for it to become uptodate or something. I need to re-read that part.
Okay, I'm reading through it again. Line 2831 has a loop, and then it's just copying.
```c
2831: for (i = 0; i < folio_batch_count(&fbatch); i++) {
...
2851: copied = copy_folio_to_iter(folio, offset, bytes, iter);
```
So even if `fat_mirror_bhs` locked the folio, `filemap_read` could *still* race with it.
Wait, if `filemap_read` *doesn't* lock the folio, then *any* write to the page cache will race with `filemap_read`. Is that standard across all filesystems?
Yes, usually `filemap_read` races with `filemap_write` (or rather `generic_perform_write`). But even `generic_perform_write` doesn't lock the folio for the entire copy operation. Wait, it *does* call `a_ops->write_begin`, which usually locks the folio. And then it copies from userspace. Then `a_ops->write_end` unlocks it.
So if `filemap_read` doesn't lock the folio, it can race with `generic_perform_write`. Is this a known, "benign" race in the kernel?
Yes, it's a known data race. Userspace can see partial writes if it reads concurrently. KCSAN usually ignores these, especially if they're marked or considered harmless.
But here, the write is in `fat_mirror_bhs`, a *metadata* update, while the read is on the block device.
If the developers added `lock_buffer` and added a comment thinking they avoided a race, and it still exists, is this a bug?
I should check if `blkdev_read_iter` is supposed to be synchronized with something. It takes `inode_lock_shared(bd_inode)`. Does FAT take `inode_lock(bd_inode)`? Probably not. It uses `fat_lock` and the file's lock.
Wait, looking back at `fat_mirror_bhs`, it's mirroring the FAT, and the primary FAT was already updated using `lock_buffer` too. Let's look at `fat_ent_write`.
|
| 16/2 |
2026/02/24 15:01 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/fat/fatent.c IncludeLines:true Name:fat_ent_write]
Results: map[SourceCode: 410: int fat_ent_write(struct inode *inode, struct fat_entry *fatent,
411: int new, int wait)
412: {
413: struct super_block *sb = inode->i_sb;
414: const struct fatent_operations *ops = MSDOS_SB(sb)->fatent_ops;
415: int err;
416:
417: ops->ent_put(fatent, new);
418: if (wait) {
419: err = fat_sync_bhs(fatent->bhs, fatent->nr_bhs);
420: if (err)
421: return err;
422: }
423: return fat_mirror_bhs(sb, fatent->bhs, fatent->nr_bhs);
424: }
SourceFile:fs/fat/fatent.c]
|
| 17/2 |
2026/02/24 15:01 |
llm |
expert |
0m
Model:
Tokens:
input: 18220
output: 26
thoughts: 39
Thoughts: Okay, let's see...
**Looking into `ent_put` for FAT32**
Alright, so I'm digging into this code and I see `ops->ent_put`. Immediately, my mind jumps to the specifics of the filesystem. For FAT32, I suspect that the implementation might be `fat32_ent_put`, but I need to confirm this. Time to trace back the definitions! I need to pinpoint exactly where this `ent_put` function is defined so I can understand its specific behavior within the FAT32 context and how it interacts with the broader filesystem structure. Let's start the search...
|
| 18/2 |
2026/02/24 15:01 |
tool |
codesearch-find-references |
0m
Args: map[Name:fatent_operations::ent_put]
Results: map[References:[map[ReferenceKind:reads ReferencingEntityKind:function ReferencingEntityName:fat_alloc_clusters SourceFile:fs/fat/fatent.c SourceLine:503] map[ReferenceKind:reads ReferencingEntityKind:function ReferencingEntityName:fat_alloc_clusters SourceFile:fs/fat/fatent.c SourceLine:505] map[ReferenceKind:reads ReferencingEntityKind:function ReferencingEntityName:fat_ent_write SourceFile:fs/fat/fatent.c SourceLine:417] map[ReferenceKind:reads ReferencingEntityKind:function ReferencingEntityName:fat_free_clusters SourceFile:fs/fat/fatent.c SourceLine:597]] TruncatedOutput:false]
|
| 19/2 |
2026/02/24 15:01 |
llm |
expert |
0m
Model:
Tokens:
input: 18640
output: 44
thoughts: 14
Thoughts: Okay, let's dive into this `fat32_ent_put` function.
**Analyzing `fat32_ent_put`**
Right, `fat32_ent_put`. I need to understand its purpose within the FAT32 file system context. The name strongly suggests this is about writing an entry, likely a directory entry or a file entry, onto the FAT32 volume. This probably means it's responsible for the creation of new files and directories, or potentially updating existing ones. It could also involve resizing existing entries, depending on the implementation details. I immediately need to consider potential issues:
* **Concurrency:** How does it handle concurrent access to the FAT table and the directory structure? Locking mechanisms are a likely candidate. Are these granular or at the entire volume level?
* **Error Handling:** What kind of error checking is implemented? Invalid filenames, insufficient disk space, corruption, and permissions are all potential failure points. Detailed error codes are a must in such systems.
* **Atomicity:** Is this operation atomic? If the system crashes mid-way, will the file system be left in a consistent state? Transactions or journaling might be involved to ensure atomicity.
* **Performance:** How efficient is the allocation of clusters and updating the FAT table? This function is fundamental to any file system operation and performance is paramount. Cluster allocation strategies and write buffering will be crucial here.
* **Security:** How does it ensure that the new entry complies with any security policies, such as file permissions?
I'll be looking for how it interacts with the FAT (File Allocation Table) and directory structures, specifically how it updates the FAT, writes directory data, and manages cluster allocation. Any optimizations related to caching or buffering would be of immediate interest as well. I'm also curious to see how it handles long filenames and any special file attributes.
|
| 20/2 |
2026/02/24 15:02 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/fat/fatent.c IncludeLines:true Name:fat32_ent_put]
Results: map[SourceCode: 187: static void fat32_ent_put(struct fat_entry *fatent, int new)
188: {
189: WARN_ON(new & 0xf0000000);
190: new |= le32_to_cpu(*fatent->u.ent32_p) & ~0x0fffffff;
191: *fatent->u.ent32_p = cpu_to_le32(new);
192: mark_buffer_dirty_inode(fatent->bhs[0], fatent->fat_inode);
193: }
SourceFile:fs/fat/fatent.c]
|
| 21/2 |
2026/02/24 15:02 |
llm |
expert |
0m
Model:
Tokens:
input: 19285
output: 19
thoughts: 4916
Thoughts: **Benign Data Race in FAT Mirror Update**
Okay, so I've been looking at this KCSAN report. It's flagging a data race between `fat_mirror_bhs` and `blkdev_read_iter` on a FAT filesystem, specifically on the `memcpy` that updates the FAT mirror. My initial thought was, "Wait, `fat32_ent_put` doesn't use `lock_buffer`, so why is the mirror update even bothering?". Then it dawns on me: the mirror uses `lock_buffer` to, in theory, prevent userspace from seeing an inconsistent FAT, but *it's failing* because `filemap_read` doesn't respect it! This doesn't seem like a kernel bug and it's certainly not a security issue, it's just a byproduct of reading a block device that is being written to.
The core of the problem is that `lock_buffer` in `fat_mirror_bhs` *doesn't* actually synchronize with reads using `filemap_read`, which is called by the block device read path (`blkdev_read_iter`). It's just a raw block device read versus filesystem metadata update and this means the race is essentially benign. KCSAN is reporting this because it's a data race, but from a kernel perspective, it's pretty harmless. The user who reads the block device gets the current state, even if it's partially updated. It's the user's responsibility to understand the consequences of raw device access; the kernel provides no guarantees of atomicity in this scenario. The filesystem itself, however, *is* consistent in its operations using `fat_lock`.
The more I dig, the more this seems to be the intended behavior. The comment in `fat_mirror_bhs` from 2005 confirms that the goal *was* to avoid this race, but the mechanism doesn't work. The fix doesn't prevent `filemap_read` from reading the data. If the race is between `memcpy` and `copy_to_user` on a block device, it's generally fine. `data_race(memcpy(...))` could silence KCSAN, but I don't think it's necessary since the read path isn't marked either. This also isn't a race with another block device reader that *does* respect buffer locks like `sb_bread`. This race specifically involves `filemap_read` which bypasses that lock.
So, I'm leaning heavily towards "benign" because the race only affects the raw block device read; kernel-side there's no corruption or instability. And it's not harmful in the classic sense, there isn't corruption in the kernel side because the kernel doesn't guarantee a consistent snapshot on raw device access. The stack traces in the report, confirm that the read is happening from userspace. It isn't a `fat_read_static_metadata` call or something like that. Furthermore, the write is occurring during `fat_write_begin`, which happens when files are written to the FAT partition. This is a case of reading the raw block device at the same time a mounted filesystem is writing to it.
I'm confident that the intended target, to synchronize with `blkdev_read_iter` is failing. Checking the `fat_mirror_bhs` code, I see: `lock_buffer(c_bh); memcpy(c_bh->b_data, bhs[n]->b_data, sb->s_blocksize); set_buffer_uptodate(c_bh); unlock_buffer(c_bh);` If the goal of `lock_buffer` was to synchronize with `blkdev_read_iter`, this approach is not correct. However, it *does* synchronize with other kernel users of the buffer head, which is why it isn't useless, just insufficient.
The data on disk, is safe too because the `memcpy` happens before the page cache is flushed, so the written data is consistent. This is a "don't do that" operation from a userspace perspective. There's nothing to do to really fix the problem. The read isn't synchronized with the write, because they're on separate code paths. I can verify the race is happening on a raw block device read from the report. So it's safe to assume this is coming from a loop device, which is a block device. Reading a `/dev/loopN` device while a FAT filesystem is mounted on it can cause this race. So the conclusion is, the data race is happening on a block device because of how a FAT partition mirrors its data and that's expected.
I have come to the conclusion it's benign because this is happening
|
| 22/2 |
2026/02/24 15:02 |
tool |
set-results |
0m
Args: map[Benign:true Confident:true]
Results: map[Benign:true Confident:true]
|
| 23/2 |
2026/02/24 15:02 |
llm |
expert |
0m
Model:
Tokens:
input: 25262
output: 700
thoughts: 265
Thoughts: **My Assessment of the Benign Data Race**
Okay, so I've looked at this data race situation, and my conclusion is that it's benign. Here's how I've come to that decision, breaking it down step-by-step.
Firstly, we're dealing with a race condition between two specific functions. On one side we have a raw block device read happening via `blkdev_read_iter` which leads to `filemap_read`, and on the other, a filesystem metadata update, specifically `fat_mirror_bhs`. This is where my initial interest was piqued.
Essentially, the filesystem is in the process of updating its FAT table mirrors, and doing so within the block device's page cache. Now, the reader here is a direct user process accessing the block device, let's say a `/dev/loopN` device.
A core point is that the kernel *doesn't* synchronize concurrent access to a block device between a filesystem and a raw reader. We don't get a consistent view here. The filesystem is responsible for maintaining its own integrity. It's using its internal synchronization mechanism, the `fat_lock`, to ensure its metadata remains consistent.
Further analysis reveals the use of `lock_buffer` in `fat_mirror_bhs`. This synchronizes with other kernel users of `buffer_head`, but *not* with `filemap_read`, because `filemap_read` bypasses `buffer_head` entirely when dealing with up-to-date pages.
Now, this is where it gets interesting. The race essentially means that the raw reader might see a partially updated FAT sector. My judgement is that this isn't a kernel bug nor a security risk.
Ultimately, and this is the most crucial part, I can confidently say that this race *will not* lead to any kernel-side memory corruption or system instability. That's the key factor that makes me label this as benign. This race is unfortunate but is well within design.
|