| 0/0 |
2026/07/22 18:40 |
flow |
assessment-kcsan |
36m
Results: map[Benign:true Explanation:### Race Summary
* **Function A:** `delete_from_page_cache_batch` (via `page_cache_delete_batch`)
* **Function B:** `filemap_write_and_wait_range` (via `mapping_needs_writeback`)
* **Variable/Field:** `mapping->nrpages` (of `struct address_space`)
---
### Synchronization Context
* **Writer (`delete_from_page_cache_batch`):** Modifies `mapping->nrpages` under
the protection of the `xa_lock_irq(&mapping->i_pages)` spinlock.
* **Reader (`filemap_write_and_wait_range`):** Reads `mapping->nrpages`
locklessly via `mapping_needs_writeback` to check if there are any pages in the
page cache that might require writeback.
---
### Final Classification: **BENIGN**
---
### Reasoning
The data race is a classic **heuristic hint** race and is completely benign for
the following reasons:
1. **No Torn Reads:** `nrpages` is an `unsigned long` aligned to a word
boundary. On all supported architectures, word-sized reads and writes are
atomic, meaning there is no risk of reading a torn or corrupted value.
2. **Heuristic Hint/Fast-path Optimization:** The check in
`mapping_needs_writeback` is a fast-path optimization to avoid the overhead of
calling `filemap_fdatawrite_range` (which scans the page cache xarray) when the
page cache is empty.
3. **Tolerance to Stale Values:**
* **Case 1: `nrpages` decrements from 1 to 0.**
* If the reader reads `1` (stale): It calls `filemap_fdatawrite_range`,
which scans the empty xarray, finds nothing, and returns. This is a safe,
slightly sub-optimal fast-path.
* If the reader reads `0`: It skips writeback, which is correct because the
page cache is empty.
* **Case 2: `nrpages` increments from 0 to 1.**
* If the reader reads `0` (stale): It skips writeback. This is safe because
the page was added concurrently with or after the writeback request
started, so the caller has no expectation of it being written back.
* If the reader reads `1`: It calls writeback, which is also safe.
* **Case 3: `nrpages` changes between non-zero values (e.g., 6 to 5).**
* Both values are non-zero, so `mapping_needs_writeback` returns `true` in
both cases, leading to the correct behavior.
4. **No Compiler Optimization/Livelock Issues:** The read is not part of a
polling loop, so there is no risk of the compiler hoisting the load out of a
loop or causing a livelock.
---
### Recommended Annotations
To document the lockless access and silence KCSAN, the read of
`mapping->nrpages` in `mapping_needs_writeback` should be annotated with
`READ_ONCE()` or wrapped in `data_race()`:
```c
static bool mapping_needs_writeback(struct address_space *mapping)
{
if (dax_mapping(mapping))
return mapping->nrexceptional;
return READ_ONCE(mapping->nrpages);
}
```]
|
| 1/1 |
2026/07/22 18:40 |
action |
kernel-checkouter |
1m
Results: map[KernelSrc:/app/workdir/cache/src/c62f318aa22f666cd96518e3888547dd326dc082]
|
| 2/1 |
2026/07/22 18:41 |
action |
kernel-builder |
12m
Results: map[KernelObj:/app/workdir/cache/build/888c520c20d3c1b16e32d2fac4415f1dd58ea0be]
|
| 3/1 |
2026/07/22 18:53 |
action |
codesearch-prepare |
8m
Results: map[Index:codesearch-index]
|
| 4/1 |
2026/07/22 19:02 |
agent |
expert |
14m
Model:
Results:
Instruction: # KCSAN Data Race Severity Analysis Protocol
You are an expert Linux kernel concurrency engineer analyzing a Kernel
Concurrency Sanitizer (KCSAN) data race report to determine if it is
**BENIGN** or **HARMFUL**.
## 1. CLASSIFICATIONS
### **BENIGN (Truly Benign)**
The logic is sound and structurally tolerant to compiler optimizations or
stale/torn reads.
- **Diagnostics/Stats:** Reads used only for `/proc`, `/sys`, counters, or
`pr_info`.
- **Heuristic Hints:** A "hint" flag where an old value only causes a
slightly delayed update or a sub-optimal but safe fast-path.
- **Single-Writer Flag Updates:** A single writer updating flags where the
concurrent read is a simple bitwise check (e.g., `flags & MASK`). These are
historically tolerated, assuming neither "Fused Accesses" nor "Ordering
Violations" are relevant in this context.
- **Marked Reloads:** A load feeding into a `cmpxchg()` loop or checked
against a later `READ_ONCE()` reload.
- **Safe Overwrites:** Writing the same value already present.
### **HARMFUL (Logic Bug or Marking Required)**
The race causes incorrect behavior due to a synchronization failure or
because missing annotations allow the compiler to break the algorithm.
**Marking Required for Correctness:**
The algorithm is logically sound but requires annotations (`READ_ONCE()`,
`WRITE_ONCE()`, `smp_load_acquire()`, `smp_store_release()`, etc.) to be safe.
- **Fused Accesses:** The compiler might merge accesses or hoist a load out
of a loop, breaking polling/wait loops (livelocks).
- **Torn Accesses:** A large access (e.g., 64-bit on 32-bit arch) might be
split into multiple non-atomic accesses. Note that `READ_ONCE()` does **not**
guarantee atomicity for 64-bit variables on 32-bit architectures.
- **Ordering Violations:** The race breaks a "happens-before" relationship
(requires primitives with implied or explicit memory barriers).
**Logic Bugs:**
A fundamental synchronization failure. Marking accesses will **not** fix it;
the logic itself must change.
- **Pointers/Lifecycle:** The racing variable is a pointer being dereferenced
or a refcount governing object lifecycle (Use-After-Free risk).
- **Control Flow:** The variable guards a critical section, memory allocation,
or hardware command.
- **Bitfields:** Concurrent writes to different bits in the same word.
Compilers often use non-atomic read-modify-write sequences, meaning a
write to `bit_A` can "clobber" a concurrent write to `bit_B`. However,
do not blindly assume all bitfield accesses are harmful; you must prove
that a concurrent write actually clobbers another in a way that breaks
logic.
- **Complex Structures:** Races on shared lists, trees, or hashmaps.
- **Lossy Updates:** Concurrent plain RMW operations (e.g., `var++`) on
non-diagnostic variables where every increment must be preserved.
- **State Machines:** Races allowing a state machine to bypass transitions
or enter an invalid state.
- **Adjacent Unsynchronized Operations:** Consider races happening at the
same time. For example, if both threads execute `struct->has_elements = true;
list_add(node, &struct->list);`, the race on `has_elements` implies an
adjacent race on `list_head`, which is HARMFUL.
## 2. RESEARCH & ANALYSIS WORKFLOW
1. **Locate the Race:** Find the exact variables and functions in the stack
traces. Do not speculate about hypothetical compiler
behaviors or theoretical dangers (e.g., dismissing something as
"fundamentally unsafe") without tracing the actual data flow to a crash.
2. **Contextualize:** Identify held locks, RCU sections, or interrupt
contexts. Explain why the race is possible (e.g., "Thread A holds `lock`,
but Thread B is a lockless reader").
3. **Data Flow:** Follow the racing variable into its subsequent uses. If
the reader sees a stale or torn value, what branch is taken? Does it lead
to a pointer dereference or `BUG_ON()`? **You must prove harm via this
data flow.**
4. **Failure Sequencing (if HARMFUL):** Construct a two-column execution trace
demonstrating the failure.
*Example format:*
```text
CPU0 CPU1
function_a()
// lockless read
ptr = global_ptr
<PREEMPT>
function_b()
lock(&my_lock)
global_ptr = NULL
unlock(&my_lock)
kfree(ptr)
if (ptr)
*ptr = 1; // -> Use-After-Free!
```
## 3. OUTPUT FORMAT
- **Race Summary:** `[Function A]` vs `[Function B]` on `[Variable/Field]`.
- **Synchronization context:** List held locks and explain the lack of mutual
exclusion.
- **Final Classification:** **BENIGN** or **HARMFUL**.
**If BENIGN:**
- **Reasoning:** Briefly explain why the race is structurally tolerant to stale
or torn reads.
- **Recommended Annotations:** Suggest appropriate annotations (e.g.,
`READ_ONCE()`, `data_race()`).
**If HARMFUL:**
- **Failure Reasoning:** Explain and provide a two-column interleaving showing
exactly how the race leads to a failure such as corruption or kernel crash.
- **Recommended Fix (ONLY if fix is trivial):** Suggest a structural fix (e.g.,
"Hold `mapping->i_pages` lock", "Convert to `atomic_t`") or required memory
ordering annotations (e.g., "Wrap in `READ_ONCE()`", "Use
`smp_load_acquire()`").
Don't make assumptions about the kernel source code (it may be different from what you assume it is).
Extensively use the provided code access tools (codesearch-*, git-*, grepper, etc)
to examine the actual source code, and confirm any assumptions.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt:
The data race report is:
==================================================================
BUG: KCSAN: data-race in delete_from_page_cache_batch / filemap_write_and_wait_range
read-write to 0xffff88810610ce08 of 8 bytes by task 14787 on cpu 1:
page_cache_delete_batch mm/filemap.c:317 [inline]
delete_from_page_cache_batch+0x534/0x720 mm/filemap.c:336
truncate_inode_pages_range+0x1f3/0x7d0 mm/truncate.c:405
truncate_inode_pages mm/truncate.c:485 [inline]
truncate_pagecache+0x53/0x70 mm/truncate.c:798
fat_write_failed fs/fat/inode.c:218 [inline]
fat_write_end+0xa5/0x160 fs/fat/inode.c:247
generic_perform_write+0x314/0x490 mm/filemap.c:4379
__generic_file_write_iter+0x9e/0x120 mm/filemap.c:4475
generic_file_write_iter+0x77/0x1f0 mm/filemap.c:4501
new_sync_write fs/read_write.c:595 [inline]
vfs_write+0x57d/0x990 fs/read_write.c:687
ksys_pwrite64 fs/read_write.c:794 [inline]
__do_sys_pwrite64 fs/read_write.c:802 [inline]
__se_sys_pwrite64 fs/read_write.c:799 [inline]
__x64_sys_pwrite64+0xfd/0x150 fs/read_write.c:799
x64_sys_call+0xa2a/0x3020 arch/x86/include/generated/asm/syscalls_64.h:19
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x136/0x3c0 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
read to 0xffff88810610ce08 of 8 bytes by task 14788 on cpu 0:
mapping_needs_writeback mm/filemap.c:634 [inline]
filemap_write_and_wait_range+0x59/0x370 mm/filemap.c:683
kiocb_write_and_wait mm/filemap.c:2904 [inline]
generic_file_read_iter+0xf4/0x320 mm/filemap.c:2978
copy_splice_read+0x47a/0x6b0 fs/splice.c:362
do_splice_read fs/splice.c:979 [inline]
splice_direct_to_actor+0x27c/0x680 fs/splice.c:1084
do_splice_direct_actor fs/splice.c:1202 [inline]
do_splice_direct+0x119/0x1a0 fs/splice.c:1228
do_sendfile+0x37d/0x640 fs/read_write.c:1371
__do_sys_sendfile64 fs/read_write.c:1432 [inline]
__se_sys_sendfile64 fs/read_write.c:1418 [inline]
__x64_sys_sendfile64+0x105/0x150 fs/read_write.c:1418
x64_sys_call+0x2dc4/0x3020 arch/x86/include/generated/asm/syscalls_64.h:41
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x136/0x3c0 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
value changed: 0x0000000000000006 -> 0x0000000000000005
Reported by Kernel Concurrency Sanitizer on:
CPU: 0 UID: 0 PID: 14788 Comm: syz.7.2910 Not tainted syzkaller #0 PREEMPT(lazy)
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 06/25/2026
==================================================================
Reply: ### Race Summary
* **Function A:** `delete_from_page_cache_batch` (via `page_cache_delete_batch`)
* **Function B:** `filemap_write_and_wait_range` (via `mapping_needs_writeback`)
* **Variable/Field:** `mapping->nrpages` (of `struct address_space`)
---
### Synchronization Context
* **Writer (`delete_from_page_cache_batch`):** Modifies `mapping->nrpages` under the protection of the `xa_lock_irq(&mapping->i_pages)` spinlock.
* **Reader (`filemap_write_and_wait_range`):** Reads `mapping->nrpages` locklessly via `mapping_needs_writeback` to check if there are any pages in the page cache that might require writeback.
---
### Final Classification: **BENIGN**
---
### Reasoning
The data race is a classic **heuristic hint** race and is completely benign for the following reasons:
1. **No Torn Reads:** `nrpages` is an `unsigned long` aligned to a word boundary. On all supported architectures, word-sized reads and writes are atomic, meaning there is no risk of reading a torn or corrupted value.
2. **Heuristic Hint/Fast-path Optimization:** The check in `mapping_needs_writeback` is a fast-path optimization to avoid the overhead of calling `filemap_fdatawrite_range` (which scans the page cache xarray) when the page cache is empty.
3. **Tolerance to Stale Values:**
* **Case 1: `nrpages` decrements from 1 to 0.**
* If the reader reads `1` (stale): It calls `filemap_fdatawrite_range`, which scans the empty xarray, finds nothing, and returns. This is a safe, slightly sub-optimal fast-path.
* If the reader reads `0`: It skips writeback, which is correct because the page cache is empty.
* **Case 2: `nrpages` increments from 0 to 1.**
* If the reader reads `0` (stale): It skips writeback. This is safe because the page was added concurrently with or after the writeback request started, so the caller has no expectation of it being written back.
* If the reader reads `1`: It calls writeback, which is also safe.
* **Case 3: `nrpages` changes between non-zero values (e.g., 6 to 5).**
* Both values are non-zero, so `mapping_needs_writeback` returns `true` in both cases, leading to the correct behavior.
4. **No Compiler Optimization/Livelock Issues:** The read is not part of a polling loop, so there is no risk of the compiler hoisting the load out of a loop or causing a livelock.
---
### Recommended Annotations
To document the lockless access and silence KCSAN, the read of `mapping->nrpages` in `mapping_needs_writeback` should be annotated with `READ_ONCE()` or wrapped in `data_race()`:
```c
static bool mapping_needs_writeback(struct address_space *mapping)
{
if (dax_mapping(mapping))
return mapping->nrexceptional;
return READ_ONCE(mapping->nrpages);
}
```
|
| 5/2 |
2026/07/22 19:02 |
llm |
expert |
0m
Model:
Tokens:
input: 6391
output: 68
thoughts: 157
Thoughts: **Analyzing the Filemap Report**
Okay, so I have a report here mentioning some specific functions and line numbers within `mm/filemap.c`. First things first, I need to locate those lines and understand their immediate context. The report points to `page_cache_delete_batch` at line 317, and then `delete_from_page_cache_batch` at line 336. That seems to indicate a batch deletion process, likely related to cleaning the page cache. Then, the report mentions `mapping_needs_writeback` at line 634 and `filemap_write_and_wait_range` at line 683. Since the report focuses on these lines of code, it appears there might be issues related to writeback operations. I'll need to read the code in and around those lines, specifically in `mm/filemap.c`, to get a sense of what's happening. I'll focus on the interactions between page cache deletion and writeback, see if I can find a link between the batch deletion and the need for writeback. Time to dive into the code.
|
| 6/2 |
2026/07/22 19:02 |
tool |
read-file |
0m
Args: map[File:mm/filemap.c FirstLine:300 LineCount:50]
Results: map[Contents: 300: * possible because we're holding the PageLock.
301: */
302: if (folio != fbatch->folios[i]) {
303: VM_BUG_ON_FOLIO(folio->index >
304: fbatch->folios[i]->index, folio);
305: continue;
306: }
307:
308: WARN_ON_ONCE(!folio_test_locked(folio));
309:
310: folio->mapping = NULL;
311: /* Leave folio->index set: truncation lookup relies on it */
312:
313: i++;
314: xas_store(&xas, NULL);
315: total_pages += folio_nr_pages(folio);
316: }
317: mapping->nrpages -= total_pages;
318: }
319:
320: void delete_from_page_cache_batch(struct address_space *mapping,
321: struct folio_batch *fbatch)
322: {
323: int i;
324:
325: if (!folio_batch_count(fbatch))
326: return;
327:
328: spin_lock(&mapping->host->i_lock);
329: xa_lock_irq(&mapping->i_pages);
330: for (i = 0; i < folio_batch_count(fbatch); i++) {
331: struct folio *folio = fbatch->folios[i];
332:
333: trace_mm_filemap_delete_from_page_cache(folio);
334: filemap_unaccount_folio(mapping, folio);
335: }
336: page_cache_delete_batch(mapping, fbatch);
337: xa_unlock_irq(&mapping->i_pages);
338: if (mapping_shrinkable(mapping))
339: inode_lru_list_add(mapping->host);
340: spin_unlock(&mapping->host->i_lock);
341:
342: for (i = 0; i < folio_batch_count(fbatch); i++)
343: filemap_free_folio(mapping, fbatch->folios[i]);
344: }
345:
346: int filemap_check_errors(struct address_space *mapping)
347: {
348: int ret = 0;
349: /* Check for outstanding write errors */
]
|
| 7/2 |
2026/07/22 19:02 |
tool |
read-file |
0m
Args: map[File:mm/filemap.c FirstLine:620 LineCount:70]
Results: map[Contents: 620: * fsfreeze(8)
621: *
622: * Return: error status of the address space.
623: */
624: int filemap_fdatawait_keep_errors(struct address_space *mapping)
625: {
626: __filemap_fdatawait_range(mapping, 0, LLONG_MAX);
627: return filemap_check_and_keep_errors(mapping);
628: }
629: EXPORT_SYMBOL(filemap_fdatawait_keep_errors);
630:
631: /* Returns true if writeback might be needed or already in progress. */
632: static bool mapping_needs_writeback(struct address_space *mapping)
633: {
634: return mapping->nrpages;
635: }
636:
637: bool filemap_range_has_writeback(struct address_space *mapping,
638: loff_t start_byte, loff_t end_byte)
639: {
640: XA_STATE(xas, &mapping->i_pages, start_byte >> PAGE_SHIFT);
641: pgoff_t max = end_byte >> PAGE_SHIFT;
642: struct folio *folio;
643:
644: if (end_byte < start_byte)
645: return false;
646:
647: rcu_read_lock();
648: xas_for_each(&xas, folio, max) {
649: if (xas_retry(&xas, folio))
650: continue;
651: if (xa_is_value(folio))
652: continue;
653: if (folio_test_dirty(folio) || folio_test_locked(folio) ||
654: folio_test_writeback(folio))
655: break;
656: }
657: rcu_read_unlock();
658: return folio != NULL;
659: }
660: EXPORT_SYMBOL_GPL(filemap_range_has_writeback);
661:
662: /**
663: * filemap_write_and_wait_range - write out & wait on a file range
664: * @mapping: the address_space for the pages
665: * @lstart: offset in bytes where the range starts
666: * @lend: offset in bytes where the range ends (inclusive)
667: *
668: * Write out and wait upon file offsets lstart->lend, inclusive.
669: *
670: * Note that @lend is inclusive (describes the last byte to be written) so
671: * that this function can be used to write to the very end-of-file (end = -1).
672: *
673: * Return: error status of the address space.
674: */
675: int filemap_write_and_wait_range(struct address_space *mapping,
676: loff_t lstart, loff_t lend)
677: {
678: int err = 0, err2;
679:
680: if (lend < lstart)
681: return 0;
682:
683: if (mapping_needs_writeback(mapping)) {
684: err = filemap_fdatawrite_range(mapping, lstart, lend);
685: /*
686: * Even if the above returned error, the pages may be
687: * written partially (e.g. -ENOSPC), so we wait for it.
688: * But the -EIO is special case, it may indicate the worst
689: * thing (e.g. bug) happened, so we avoid waiting for it.
]
|
| 8/2 |
2026/07/22 19:02 |
llm |
expert |
0m
Model:
Tokens:
input: 8694
output: 32
thoughts: 396
Thoughts: **Analyzing a Potential Race Condition in the Page Cache**
Okay, let's break this down. I'm seeing a possible race condition, and my internal alarm bells are going off. Task 14787, running on CPU 1, is busy deleting pages from the page cache via `delete_from_page_cache_batch`. This, in turn, calls `page_cache_delete_batch`, which *directly decrements* `mapping->nrpages`. That's a write operation β specifically, a reduction of the page count associated with a particular mapping.
At the same time, Task 14788 on CPU 0 is calling `filemap_write_and_wait_range`. Within that function, `mapping_needs_writeback(mapping)` is called. And what does *that* do? It reads `mapping->nrpages` to decide if the mapping needs to be written back to disk.
So we have a classic read-modify-write scenario, and that, my friend, smells like a race. Task 14787 is decrementing, while Task 14788 is reading. But hold on, is this a *benign* race? I need to analyze this further.
The critical thing is what `mapping_needs_writeback` does with the value. If `mapping->nrpages` is non-zero, it calls `filemap_fdatawrite_range`. If itβs zero, it skips it. The outcome hinges on what `filemap_fdatawrite_range` *would* do in either case. And critically, I need to know if the read of `mapping->nrpages` is done with any sort of locking. If not, then we could have a nasty problem.
Okay, let's examine the code. Ah, the `filemap_write_and_wait_range` function does *not* hold a lock when calling `mapping_needs_writeback`. The read of `nrpages` is, therefore, potentially lockless. So, letβs see what else might touch this. I need to comprehensively search `mm/filemap.c` for instances of either `mapping_needs_writeback` or `nrpages` to see if there are other potential contenders.
|
| 9/2 |
2026/07/22 19:02 |
tool |
grepper |
0m
Args: map[Expression:mapping_needs_writeback PathPrefix:mm/filemap.c]
Results: map[Output:mm/filemap.c=629=EXPORT_SYMBOL(filemap_fdatawait_keep_errors);
--
mm/filemap.c-631-/* Returns true if writeback might be needed or already in progress. */
mm/filemap.c:632:static bool mapping_needs_writeback(struct address_space *mapping)
mm/filemap.c-633-{
--
mm/filemap.c=675=int filemap_write_and_wait_range(struct address_space *mapping,
--
mm/filemap.c-682-
mm/filemap.c:683: if (mapping_needs_writeback(mapping)) {
mm/filemap.c-684- err = filemap_fdatawrite_range(mapping, lstart, lend);
--
mm/filemap.c=777=int file_write_and_wait_range(struct file *file, loff_t lstart, loff_t lend)
--
mm/filemap.c-784-
mm/filemap.c:785: if (mapping_needs_writeback(mapping)) {
mm/filemap.c-786- err = filemap_fdatawrite_range(mapping, lstart, lend);
]
|
| 10/2 |
2026/07/22 19:02 |
llm |
expert |
0m
Model:
Tokens:
input: 9927
output: 24
thoughts: 23
Thoughts: Okay, here's my summary, presented as a stream of thought:
**Diving into the `nrpages` Hunt**
Alright, I need to understand how `nrpages` is used, specifically within the `mm/` directory. Let's see... my first instinct is to do a quick search within that directory. I'll need to use a tool to locate instances of `nrpages` and see where it's referenced and how it's being handled. This should help me quickly grasp the context surrounding it β are we talking about a simple variable, a macro, a function, a struct member, what? Understanding the access patterns is key. Hopefully, the search results will give me a clear picture of its role and significance. I'm expecting some pretty clear hits, given that it's a core concept.
|
| 11/2 |
2026/07/22 19:02 |
tool |
grepper |
0m
Args: map[Expression:nrpages PathPrefix:mm/]
Results: map[Output:mm/early_ioremap.c=108=__early_ioremap(resource_size_t phys_addr, unsigned long size, pgprot_t prot)
--
mm/early_ioremap.c-111- resource_size_t last_addr;
mm/early_ioremap.c:112: unsigned int nrpages;
mm/early_ioremap.c-113- enum fixed_addresses idx;
--
mm/early_ioremap.c-145- */
mm/early_ioremap.c:146: nrpages = size >> PAGE_SHIFT;
mm/early_ioremap.c:147: if (WARN_ON(nrpages > NR_FIX_BTMAPS))
mm/early_ioremap.c-148- return NULL;
--
mm/early_ioremap.c-156- idx = FIX_BTMAP_BEGIN - NR_FIX_BTMAPS*slot;
mm/early_ioremap.c:157: while (nrpages > 0) {
mm/early_ioremap.c-158- if (after_paging_init)
--
mm/early_ioremap.c-163- --idx;
mm/early_ioremap.c:164: --nrpages;
mm/early_ioremap.c-165- }
--
mm/early_ioremap.c=171=void __init early_iounmap(void __iomem *addr, unsigned long size)
--
mm/early_ioremap.c-174- unsigned long offset;
mm/early_ioremap.c:175: unsigned int nrpages;
mm/early_ioremap.c-176- enum fixed_addresses idx;
--
mm/early_ioremap.c-202- offset = offset_in_page(virt_addr);
mm/early_ioremap.c:203: nrpages = PAGE_ALIGN(offset + size) >> PAGE_SHIFT;
mm/early_ioremap.c-204-
mm/early_ioremap.c-205- idx = FIX_BTMAP_BEGIN - NR_FIX_BTMAPS*slot;
mm/early_ioremap.c:206: while (nrpages > 0) {
mm/early_ioremap.c-207- if (after_paging_init)
--
mm/early_ioremap.c-211- --idx;
mm/early_ioremap.c:212: --nrpages;
mm/early_ioremap.c-213- }
--
mm/fadvise.c=31=int generic_fadvise(struct file *file, loff_t offset, loff_t len, int advice)
--
mm/fadvise.c-38- pgoff_t end_index;
mm/fadvise.c:39: unsigned long nrpages;
mm/fadvise.c-40-
--
mm/fadvise.c-101- /* Careful about overflow on the "+1" */
mm/fadvise.c:102: nrpages = end_index - start_index + 1;
mm/fadvise.c:103: if (!nrpages)
mm/fadvise.c:104: nrpages = ~0UL;
mm/fadvise.c-105-
mm/fadvise.c:106: force_page_cache_readahead(mapping, file, start_index, nrpages);
mm/fadvise.c-107- break;
--
mm/filemap.c=129=static void page_cache_delete(struct address_space *mapping,
--
mm/filemap.c-146- /* Leave folio->index set: truncation lookup relies upon it */
mm/filemap.c:147: mapping->nrpages -= nr;
mm/filemap.c-148-}
--
mm/filemap.c=279=static void page_cache_delete_batch(struct address_space *mapping,
--
mm/filemap.c-316- }
mm/filemap.c:317: mapping->nrpages -= total_pages;
mm/filemap.c-318-}
--
mm/filemap.c=632=static bool mapping_needs_writeback(struct address_space *mapping)
mm/filemap.c-633-{
mm/filemap.c:634: return mapping->nrpages;
mm/filemap.c-635-}
--
mm/filemap.c=848=noinline int __filemap_add_folio(struct address_space *mapping,
--
mm/filemap.c-918-
mm/filemap.c:919: mapping->nrpages += nr;
mm/filemap.c-920-
--
mm/filemap.c=4263=void kiocb_invalidate_post_direct_write(struct kiocb *iocb, size_t count)
--
mm/filemap.c-4266-
mm/filemap.c:4267: if (mapping->nrpages &&
mm/filemap.c-4268- invalidate_inode_pages2_range(mapping,
--
mm/filemap.c=4556=int filemap_invalidate_inode(struct inode *inode, bool flush,
--
mm/filemap.c-4563-
mm/filemap.c:4564: if (!mapping || !mapping->nrpages || end < start)
mm/filemap.c-4565- goto out;
--
mm/filemap.c-4569-
mm/filemap.c:4570: if (!mapping->nrpages)
mm/filemap.c-4571- goto unlock;
--
mm/khugepaged.c=2242=static enum scan_result collapse_file(struct mm_struct *mm, unsigned long addr,
--
mm/khugepaged.c-2651- xas_lock_irq(&xas);
mm/khugepaged.c:2652: mapping->nrpages -= nr_none;
mm/khugepaged.c-2653- xas_unlock_irq(&xas);
--
mm/mprotect.c=725=mprotect_fixup(struct vma_iterator *vmi, struct mmu_gather *tlb,
--
mm/mprotect.c-731- vma_flags_t new_vma_flags = legacy_to_vma_flags(newflags);
mm/mprotect.c:732: long nrpages = (end - start) >> PAGE_SHIFT;
mm/mprotect.c-733- unsigned int mm_cp_flags = 0;
--
mm/mprotect.c-772- /* Check space limits when area turns into data. */
mm/mprotect.c:773: if (!may_expand_vm(mm, &new_vma_flags, nrpages) &&
mm/mprotect.c:774: may_expand_vm(mm, &old_vma_flags, nrpages))
mm/mprotect.c-775- return -ENOMEM;
--
mm/mprotect.c-778- VMA_SHARED_BIT, VMA_NORESERVE_BIT)) {
mm/mprotect.c:779: charged = nrpages;
mm/mprotect.c-780- if (security_vm_enough_memory_mm(mm, charged))
--
mm/mprotect.c-810- !vma_flags_test(&new_vma_flags, VMA_ACCOUNT_BIT))
mm/mprotect.c:811: vm_unacct_memory(nrpages);
mm/mprotect.c-812-
--
mm/mprotect.c-821-
mm/mprotect.c:822: vm_stat_account(mm, vma_flags_to_legacy(old_vma_flags), -nrpages);
mm/mprotect.c-823- newflags = vma_flags_to_legacy(new_vma_flags);
mm/mprotect.c:824: vm_stat_account(mm, newflags, nrpages);
mm/mprotect.c-825- perf_event_mmap(vma);
--
mm/readahead.c=222=void page_cache_ra_unbounded(struct readahead_control *ractl,
--
mm/readahead.c-228- unsigned long mark = ULONG_MAX, i = 0;
mm/readahead.c:229: unsigned int min_nrpages = mapping_min_folio_nrpages(mapping);
mm/readahead.c-230-
--
mm/readahead.c-249- /*
mm/readahead.c:250: * As iterator `i` is aligned to min_nrpages, round_up the
mm/readahead.c-251- * difference between nr_to_read and lookahead_size to mark the
--
mm/readahead.c-259- nr_to_read - lookahead_size,
mm/readahead.c:260: min_nrpages);
mm/readahead.c-261- mark = ra_folio_index - index;
--
mm/readahead.c-282- read_pages(ractl);
mm/readahead.c:283: ractl->_index += min_nrpages;
mm/readahead.c-284- i = ractl->_index - index;
--
mm/readahead.c-298- read_pages(ractl);
mm/readahead.c:299: ractl->_index += min_nrpages;
mm/readahead.c-300- i = ractl->_index - index;
--
mm/readahead.c-305- ractl->_workingset |= folio_test_workingset(folio);
mm/readahead.c:306: ractl->_nr_pages += min_nrpages;
mm/readahead.c:307: i += min_nrpages;
mm/readahead.c-308- }
--
mm/readahead.c=786=void readahead_expand(struct readahead_control *ractl,
--
mm/readahead.c-792- gfp_t gfp_mask = readahead_gfp_mask(mapping);
mm/readahead.c:793: unsigned long min_nrpages = mapping_min_folio_nrpages(mapping);
mm/readahead.c-794- unsigned int min_order = mapping_min_folio_order(mapping);
--
mm/readahead.c-798- * Readahead code should have aligned the ractl->_index to
mm/readahead.c:799: * min_nrpages before calling readahead aops.
mm/readahead.c-800- */
mm/readahead.c:801: VM_BUG_ON(!IS_ALIGNED(ractl->_index, min_nrpages));
mm/readahead.c-802-
--
mm/readahead.c-824- }
mm/readahead.c:825: ractl->_nr_pages += min_nrpages;
mm/readahead.c-826- ractl->_index = folio->index;
--
mm/readahead.c-853- }
mm/readahead.c:854: ractl->_nr_pages += min_nrpages;
mm/readahead.c-855- if (ra) {
mm/readahead.c:856: ra->size += min_nrpages;
mm/readahead.c:857: ra->async_size += min_nrpages;
mm/readahead.c-858- }
--
mm/shmem.c=413=static void shmem_free_inode(struct super_block *sb, size_t freed_ispace)
--
mm/shmem.c-431- *
mm/shmem.c:432: * But normally info->alloced == inode->i_mapping->nrpages + info->swapped
mm/shmem.c:433: * So mm freed is info->alloced - (inode->i_mapping->nrpages + info->swapped)
mm/shmem.c-434- *
--
mm/shmem.c=437=bool shmem_recalc_inode(struct inode *inode, long alloced, long swapped)
--
mm/shmem.c-446- freed = info->alloced - info->swapped -
mm/shmem.c:447: READ_ONCE(inode->i_mapping->nrpages);
mm/shmem.c-448- /*
mm/shmem.c-449- * Special case: whereas normally shmem_recalc_inode() is called
mm/shmem.c:450: * after i_mapping->nrpages has already been adjusted (up or down),
mm/shmem.c:451: * shmem_writeout() has to raise swapped before nrpages is lowered -
mm/shmem.c-452- * to stop a racing shmem_recalc_inode() from thinking that a page has
--
mm/shmem.c=470=bool shmem_charge(struct inode *inode, long pages)
--
mm/shmem.c-476-
mm/shmem.c:477: /* nrpages adjustment first, then shmem_recalc_inode() when balanced */
mm/shmem.c-478- xa_lock_irq(&mapping->i_pages);
mm/shmem.c:479: mapping->nrpages += pages;
mm/shmem.c-480- xa_unlock_irq(&mapping->i_pages);
--
mm/shmem.c=486=void shmem_uncharge(struct inode *inode, long pages)
--
mm/shmem.c-488- /* pages argument is currently unused: keep it to help debugging */
mm/shmem.c:489: /* nrpages adjustment done by __filemap_remove_folio() or caller */
mm/shmem.c-490-
--
mm/shmem.c=881=int shmem_add_to_page_cache(struct folio *folio,
--
mm/shmem.c-925- shmem_update_stats(folio, nr);
mm/shmem.c:926: mapping->nrpages += nr;
mm/shmem.c-927-unlock:
--
mm/shmem.c=943=static void shmem_delete_from_page_cache(struct folio *folio, void *radswap)
--
mm/shmem.c-951- folio->mapping = NULL;
mm/shmem.c:952: mapping->nrpages -= nr;
mm/shmem.c-953- shmem_update_stats(folio, -nr);
--
mm/shmem.c=1283=static int shmem_getattr(struct mnt_idmap *idmap,
--
mm/shmem.c-1289-
mm/shmem.c:1290: if (info->alloced - info->swapped != inode->i_mapping->nrpages)
mm/shmem.c-1291- shmem_recalc_inode(inode, 0, 0);
--
mm/shmem.c=1913=static struct folio *shmem_alloc_and_add_folio(struct vm_fault *vmf,
--
mm/shmem.c-1993- freed = pages + info->alloced - info->swapped -
mm/shmem.c:1994: READ_ONCE(mapping->nrpages);
mm/shmem.c-1995- if (freed > 0)
--
mm/truncate.c=468=EXPORT_SYMBOL(truncate_inode_pages_range);
--
mm/truncate.c-479- * deletion (inside __filemap_remove_folio()) in the specified range. Thus
mm/truncate.c:480: * mapping->nrpages can be non-zero when this function returns even after
mm/truncate.c-481- * truncation of the whole mapping.
--
mm/vma.c=1398=static int vms_gather_munmap_vmas(struct vma_munmap_struct *vms,
--
mm/vma.c-1439- for_each_vma_range(*(vms->vmi), next, vms->end) {
mm/vma.c:1440: long nrpages;
mm/vma.c-1441-
--
mm/vma.c-1458- vma_mark_detached(next);
mm/vma.c:1459: nrpages = vma_pages(next);
mm/vma.c-1460-
mm/vma.c:1461: vms->nr_pages += nrpages;
mm/vma.c-1462- if (vma_test(next, VMA_LOCKED_BIT))
mm/vma.c:1463: vms->locked_vm += nrpages;
mm/vma.c-1464-
mm/vma.c-1465- if (vma_test(next, VMA_ACCOUNT_BIT))
mm/vma.c:1466: vms->nr_accounted += nrpages;
mm/vma.c-1467-
mm/vma.c-1468- if (is_exec_mapping(next->vm_flags))
mm/vma.c:1469: vms->exec_vm += nrpages;
mm/vma.c-1470- else if (is_stack_mapping(next->vm_flags))
mm/vma.c:1471: vms->stack_vm += nrpages;
mm/vma.c-1472- else if (is_data_mapping_vma_flags(&next->flags))
mm/vma.c:1473: vms->data_vm += nrpages;
mm/vma.c-1474-
]
|
| 12/2 |
2026/07/22 19:02 |
llm |
expert |
0m
Model:
Tokens:
input: 14810
output: 21
thoughts: 733
Thoughts: **Analyzing a Potential Race Condition in the Linux Kernel's Filemap Code**
Okay, so I'm digging into this `mm/shmem.c` code, specifically lines 446-447 and 1993-1994, and it looks like `mapping->nrpages` is being read locklessly, using `READ_ONCE()` in some places. I've seen that before, which is fine, but then I stumbled upon `mapping_needs_writeback` in `mm/filemap.c`, and it's reading `mapping->nrpages` *without* `READ_ONCE()`. That's what has caught my attention and I need to figure out if it's a problem.
Let me refresh my memory on `mapping_needs_writeback`: it's a simple function that checks the value of `mapping->nrpages`. It's called from `filemap_write_and_wait_range`, which, in turn, is part of the `kiocb_write_and_wait` path and ultimately involved in `generic_file_read_iter`. So if `mapping_needs_writeback` returns false, it skips the `filemap_fdatawrite_range` call.
Now, here's the concern: If `nrpages` is being updated concurrently, and we read it without proper synchronization, we might get a stale or even a "torn" value. If `nrpages` is non-zero, and it gets read as zero, we might skip a crucial writeback operation. That could be bad.
But wait, let's think this through. If `nrpages` is being *decremented* to zero, reading it as zero or a non-zero value, is fine because there's nothing to write back. However, If `nrpages` is being *incremented* from zero to one, then we might read it as zero, skip the writeback, and that could be problematic because a page is being added to the mapping and it might need writeback. But hold on, the writeback would have nothing to write out at this point, so this should be fine.
The real worry here is if we get a torn read, meaning we read a value in the middle of an update. However, `nrpages` is an `unsigned long`. On 64-bit architectures, an aligned read/write of an `unsigned long` is atomic. *Is* it aligned, though? Let's check the definition of `struct address_space`.
Okay, I need to check the definition of `struct address_space` to confirm. If it's aligned, then, even in the worst-case scenario, the read *should* be atomic, eliminating the possibility of a truly torn value. If that is the case, then this whole thing might be benign after all. Still, this is something I need to look into further.
|
| 13/2 |
2026/07/22 19:02 |
tool |
codesearch-struct-layout |
0m
Args:
Results: map[Fields:[map[Name:host OffsetBits:0 SizeBits:64] map[Name:i_pages OffsetBits:64 SizeBits:128] map[Name:invalidate_lock OffsetBits:192 SizeBits:256] map[Name:gfp_mask OffsetBits:448 SizeBits:32] map[Name:i_mmap_writable OffsetBits:480 SizeBits:32] map[Name:i_mmap OffsetBits:512 SizeBits:128] map[Name:nrpages OffsetBits:640 SizeBits:64] map[Name:writeback_index OffsetBits:704 SizeBits:64] map[Name:a_ops OffsetBits:768 SizeBits:64] map[Name:flags OffsetBits:832 SizeBits:64] map[Name:wb_err OffsetBits:896 SizeBits:32] map[Name:i_private_lock OffsetBits:928 SizeBits:32] map[Name:i_mmap_rwsem OffsetBits:960 SizeBits:256]]]
|
| 14/2 |
2026/07/22 19:02 |
llm |
expert |
0m
Model:
Tokens:
input: 16483
output: 28
thoughts: 470
Thoughts: **Analyzing Potential Data Races with `nrpages`**
Okay, so I'm digging into this `nrpages` field and how it's being accessed. First thought: it's at offset 640 bits, which is 80 bytes. That 8-byte alignment on 64-bit architectures means reads and writes should be atomic, which is good. But wait, what about the 32-bit world? Ah, right, on 32-bit, `unsigned long` is also 32-bit, so atomic reads and writes still hold. No tearing issue there. That's a relief.
Now, I need to consider compiler optimizations. Are we potentially getting into trouble with the compiler rearranging accesses or caching values? Specifically, is `mapping_needs_writeback` called inside a loop, where such optimizations could be a problem? Let's check. Ah, no, in `filemap_write_and_wait_range`, and `file_write_and_wait_range` it's only called *once*. Okay, so less of a concern.
Alright, where else is `mapping_needs_writeback` being used? Let's be thorough and check all calls. Scanning the code, I find those two functions. That's it. And `kiocb_invalidate_post_direct_write` and `filemap_invalidate_inode`, they also perform simple checks against `mapping->nrpages`.
Alright, a deeper dive is required. This gets me thinking, are there *any* other places where `mapping->nrpages` is read locklessly? Need to make sure I'm not missing anything. I need to make sure I haven't missed anything. I'll need to check the entire codebase.
Let's use `git-log` to see if there's any history on this, any commits related to `mapping_needs_writeback` or any mentions of `nrpages` data races. That's the next step. Let's see what the commit history reveals.
|
| 15/2 |
2026/07/22 19:02 |
tool |
git-log |
10m
Error:
Args: map[CodeRegexp:mapping_needs_writeback Count:10]
Results:
|
| 16/2 |
2026/07/22 19:12 |
llm |
expert |
0m
Model:
Tokens:
input: 17455
output: 40
thoughts: 24
Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience:
**Refining the Git-Log Query**
Right, let's zero in on the relevant changes. The standard `git log` command is a powerful tool, but without refinement, it can return a deluge of information. Thinking through the current debugging process, I need to isolate the changes affecting the file map functionality within the `mm` subsystem, specifically, the `filemap.c` implementation. Therefore, the logical next step is to use a `PathPrefix` argument. I'll need to use `PathPrefix: "mm/filemap.c"` to effectively filter the `git log` output. This will significantly reduce the noise and allow me to focus directly on the commits touching that specific file. This targeted approach is much more efficient than sifting through potentially unrelated changes. I'll get the information I need quicker, which is the ultimate goal.
|
| 17/2 |
2026/07/22 19:12 |
tool |
git-log |
0m
Args: map[CodeRegexp:mapping_needs_writeback Count:10 PathPrefix:mm/filemap.c]
Results: map[Output:4bdcd1dd4d2f mm: move filemap_range_needs_writeback() into header
63135aa3866d mm: provide filemap_range_needs_writeback() helper
ddf8f376d137 mm/filemap.c: clean up filemap_write_and_wait()
9326c9b20dd8 mm: consolidate dax / non-dax checks for writeback
]
|
| 18/2 |
2026/07/22 19:12 |
llm |
expert |
0m
Model:
Tokens:
input: 17835
output: 26
thoughts: 36
Thoughts: Okay, here's my interpretation of those thoughts, presented as a first-person expert summary:
**Digging into the Revision History**
Alright, I need to understand some specific changes in this codebase. I'm going to start by examining two commits: `63135aa3866d` and `4bdcd1dd4d2f`. Those commit hashes should give me a pretty direct line into the exact diffs and contextual code modifications I'm interested in. I'm essentially bypassing general overviews and heading straight for the relevant changes in the revision history. I want to see what precisely happened during those specific commits. This is the fastest way to get to the root of the issue, whatever the issue may be!
|
| 19/2 |
2026/07/22 19:12 |
tool |
git-show |
0m
Args:
Results: map[Output:commit 63135aa3866db99fd923b716c5ff2e468879624a
Author: Jens Axboe <axboe@kernel.dk>
Date: Thu Apr 29 22:55:18 2021 -0700
mm: provide filemap_range_needs_writeback() helper
Patch series "Improve IOCB_NOWAIT O_DIRECT reads", v3.
An internal workload complained because it was using too much CPU, and
when I took a look, we had a lot of io_uring workers going to town.
For an async buffered read like workload, I am normally expecting _zero_
offloads to a worker thread, but this one had tons of them. I'd drop
caches and things would look good again, but then a minute later we'd
regress back to using workers. Turns out that every minute something
was reading parts of the device, which would add page cache for that
inode. I put patches like these in for our kernel, and the problem was
solved.
Don't -EAGAIN IOCB_NOWAIT dio reads just because we have page cache
entries for the given range. This causes unnecessary work from the
callers side, when the IO could have been issued totally fine without
blocking on writeback when there is none.
This patch (of 3):
For O_DIRECT reads/writes, we check if we need to issue a call to
filemap_write_and_wait_range() to issue and/or wait for writeback for any
page in the given range. The existing mechanism just checks for a page in
the range, which is suboptimal for IOCB_NOWAIT as we'll fallback to the
slow path (and needing retry) if there's just a clean page cache page in
the range.
Provide filemap_range_needs_writeback() which tries a little harder to
check if we actually need to issue and/or wait for writeback in the range.
Link: https://lkml.kernel.org/r/20210224164455.1096727-1-axboe@kernel.dk
Link: https://lkml.kernel.org/r/20210224164455.1096727-2-axboe@kernel.dk
Signed-off-by: Jens Axboe <axboe@kernel.dk>
Reviewed-by: Matthew Wilcox (Oracle) <willy@infradead.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
diff --git a/include/linux/fs.h b/include/linux/fs.h
index bf4e90d3ab18..12766edee81f 100644
--- a/include/linux/fs.h
+++ b/include/linux/fs.h
@@ -2878,6 +2878,8 @@ static inline int filemap_fdatawait(struct address_space *mapping)
extern bool filemap_range_has_page(struct address_space *, loff_t lstart,
loff_t lend);
+extern bool filemap_range_needs_writeback(struct address_space *,
+ loff_t lstart, loff_t lend);
extern int filemap_write_and_wait_range(struct address_space *mapping,
loff_t lstart, loff_t lend);
extern int __filemap_fdatawrite_range(struct address_space *mapping,
diff --git a/mm/filemap.c b/mm/filemap.c
index 151090fdcf29..1750742b5689 100644
--- a/mm/filemap.c
+++ b/mm/filemap.c
@@ -635,6 +635,49 @@ static bool mapping_needs_writeback(struct address_space *mapping)
return mapping->nrpages;
}
+/**
+ * filemap_range_needs_writeback - check if range potentially needs writeback
+ * @mapping: address space within which to check
+ * @start_byte: offset in bytes where the range starts
+ * @end_byte: offset in bytes where the range ends (inclusive)
+ *
+ * Find at least one page in the range supplied, usually used to check if
+ * direct writing in this range will trigger a writeback. Used by O_DIRECT
+ * read/write with IOCB_NOWAIT, to see if the caller needs to do
+ * filemap_write_and_wait_range() before proceeding.
+ *
+ * Return: %true if the caller should do filemap_write_and_wait_range() before
+ * doing O_DIRECT to a page in this range, %false otherwise.
+ */
+bool filemap_range_needs_writeback(struct address_space *mapping,
+ loff_t start_byte, loff_t end_byte)
+{
+ XA_STATE(xas, &mapping->i_pages, start_byte >> PAGE_SHIFT);
+ pgoff_t max = end_byte >> PAGE_SHIFT;
+ struct page *page;
+
+ if (!mapping_needs_writeback(mapping))
+ return false;
+ if (!mapping_tagged(mapping, PAGECACHE_TAG_DIRTY) &&
+ !mapping_tagged(mapping, PAGECACHE_TAG_WRITEBACK))
+ return false;
+ if (end_byte < start_byte)
+ return false;
+
+ rcu_read_lock();
+ xas_for_each(&xas, page, max) {
+ if (xas_retry(&xas, page))
+ continue;
+ if (xa_is_value(page))
+ continue;
+ if (PageDirty(page) || PageLocked(page) || PageWriteback(page))
+ break;
+ }
+ rcu_read_unlock();
+ return page != NULL;
+}
+EXPORT_SYMBOL_GPL(filemap_range_needs_writeback);
+
/**
* filemap_write_and_wait_range - write out & wait on a file range
* @mapping: the address_space for the pages
]
|
| 20/2 |
2026/07/22 19:12 |
llm |
expert |
0m
Model:
Tokens:
input: 19664
output: 36
thoughts: 33
Thoughts: Okay, here's my interpretation of those thoughts, framed as my own, and tailored for someone with expertise in this area:
**Digging into Filemap.c for Related Changes**
Alright, I'm diving into this. The current problem seems tied to `mm/filemap.c`, and I need to understand the potential ripple effects. My immediate thought is, let's cast a wider net. The core issue probably lies with interactions surrounding `mapping_needs_writeback` or `nrpages` within that module. So, I need to hunt down other commits that have, even tangentially, modified these variables or the code surrounding them. This means a commit history search is in order β specifically looking for changes that touch these variables within `mm/filemap.c`. A targeted search will save a lot of time and effort in the long run. Let's see what we can turn up!
|
| 21/2 |
2026/07/22 19:12 |
tool |
git-log |
0m
Args: map[CodeRegexp:nrpages Count:10 PathPrefix:mm/filemap.c]
Results: map[Output:74e797d79cf1 mm: Provide a means of invalidation without using launder_folio
c402a9a9430b filemap: add a kiocb_invalidate_post_direct_write helper
d68eccad3706 mm/filemap: Allow large folios to be added to the page cache
46be67b424ef mm: stop accounting shadow entries
9266a14033a8 mm/filemap.c: warn if stale pagecache is left after direct write
80c1fe902691 mm/filemap.c: remove redundant cache invalidation after async direct-io write
875d91b11a20 mm/filemap.c: rewrite mapping_needs_writeback in less fancy manner
8fa8e538e4be page cache: Convert filemap_range_has_page to XArray
74d609585d8b page cache: Add and replace pages using the XArray
aa65c29ce1b6 mm: batch radix tree operations when truncating pages
]
|
| 22/2 |
2026/07/22 19:12 |
llm |
expert |
0m
Model:
Tokens:
input: 20197
output: 27
thoughts: 34
Thoughts: Here's my take on those commit thoughts:
**Analyzing Commit 875d91b11a20: "rewrite mapping_needs_writeback in less fancy manner"**
Okay, let's dive into commit `875d91b11a20`. The commit message itself is intriguing, hinting at a simplification of the `mapping_needs_writeback` function. My initial thought is to assess the specific modifications. I'll need to dissect the code changes to understand what "less fancy" actually entails.
From my experience, "less fancy" often implies removing unnecessary abstractions, reducing complexity, or perhaps optimizing for performance by trading elegance for directness. I immediately want to compare the before-and-after versions, focusing on things like:
* **Readability:** Is the new code easier to grasp at a glance? Are there fewer nested conditionals or complex logic paths?
* **Performance:** Did they manage to squeeze out any performance improvements? Perhaps by reducing function call overhead or avoiding redundant computations?
* **Maintainability:** Will the simplified code be easier to debug and modify in the future? Avoiding overly clever tricks is often a key consideration here.
* **Correctness:** Did they introduce any regressions? That's the most critical thing to watch out for, as a simpler implementation can sometimes unintentionally break things.
Ultimately, I expect this to be a straightforward rewrite. I'll be looking to see if my intuitions about the simplification are correct, and whether any new techniques have been learned.
|
| 23/2 |
2026/07/22 19:12 |
tool |
git-show |
0m
Args:
Results: map[Output:commit 875d91b11a201276ac3a9ab79f8b0fa3dc4ee8fd
Author: Konstantin Khlebnikov <koct9i@gmail.com>
Date: Mon Sep 23 15:34:48 2019 -0700
mm/filemap.c: rewrite mapping_needs_writeback in less fancy manner
This actually checks that writeback is needed or in progress.
Link: http://lkml.kernel.org/r/156378817069.1087.1302816672037672488.stgit@buzz
Signed-off-by: Konstantin Khlebnikov <khlebnikov@yandex-team.ru>
Reviewed-by: Andrew Morton <akpm@linux-foundation.org>
Cc: Tejun Heo <tj@kernel.org>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: Jan Kara <jack@suse.cz>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
diff --git a/mm/filemap.c b/mm/filemap.c
index 82f3b9a3a940..d5462d706f76 100644
--- a/mm/filemap.c
+++ b/mm/filemap.c
@@ -618,10 +618,13 @@ int filemap_fdatawait_keep_errors(struct address_space *mapping)
}
EXPORT_SYMBOL(filemap_fdatawait_keep_errors);
+/* Returns true if writeback might be needed or already in progress. */
static bool mapping_needs_writeback(struct address_space *mapping)
{
- return (!dax_mapping(mapping) && mapping->nrpages) ||
- (dax_mapping(mapping) && mapping->nrexceptional);
+ if (dax_mapping(mapping))
+ return mapping->nrexceptional;
+
+ return mapping->nrpages;
}
int filemap_write_and_wait(struct address_space *mapping)
]
|
| 24/2 |
2026/07/22 19:12 |
llm |
expert |
0m
Model:
Tokens:
input: 21183
output: 35
thoughts: 236
Thoughts: **Analyzing a Potential Race Condition in the File System**
Okay, so `mapping_needs_writeback` is intended as a quick optimization, a shortcut to prevent unnecessary calls to `filemap_fdatawrite_range` when the page cache is empty. Makes sense. But let's dig a bit deeper. Essentially, it's a conditional: if true, we write; if false, we don't. Hmm, is this where a race condition could hide?
Let's seeβ¦ The KCSAN report might shed some light on the situation. Ah, there it is. One task, 14787 on CPU 1, is actively deleting pages from the page cacheβlikely due to truncation or deletion. Simultaneously, task 14788 on CPU 0 is calling `filemap_write_and_wait_range` during a read operation, specifically within `generic_file_read_iter`. This strikes me as odd. Why would a read function need to *write* and then *wait*?
I need to investigate `generic_file_read_iter` in `mm/filemap.c` to understand why this write-and-wait is happening during a read. This seems like a potential point of contention where `mapping_needs_writeback`'s conditional logic, if not handled carefully, might introduce an issue. Time to dive into the code.
|
| 25/2 |
2026/07/22 19:13 |
tool |
read-file |
0m
Args: map[File:mm/filemap.c FirstLine:2960 LineCount:40]
Results: map[Contents:2960: * Return:
2961: * * number of bytes copied, even for partial reads
2962: * * negative error code (or 0 if IOCB_NOIO) if nothing was read
2963: */
2964: ssize_t
2965: generic_file_read_iter(struct kiocb *iocb, struct iov_iter *iter)
2966: {
2967: size_t count = iov_iter_count(iter);
2968: ssize_t retval = 0;
2969:
2970: if (!count)
2971: return 0; /* skip atime */
2972:
2973: if (iocb->ki_flags & IOCB_DIRECT) {
2974: struct file *file = iocb->ki_filp;
2975: struct address_space *mapping = file->f_mapping;
2976: struct inode *inode = mapping->host;
2977:
2978: retval = kiocb_write_and_wait(iocb, count);
2979: if (retval < 0)
2980: return retval;
2981: file_accessed(file);
2982:
2983: retval = mapping->a_ops->direct_IO(iocb, iter);
2984: if (retval >= 0) {
2985: iocb->ki_pos += retval;
2986: count -= retval;
2987: }
2988: if (retval != -EIOCBQUEUED)
2989: iov_iter_revert(iter, count - iov_iter_count(iter));
2990:
2991: /*
2992: * Btrfs can have a short DIO read if we encounter
2993: * compressed extents, so if there was an error, or if
2994: * we've already read everything we wanted to, or if
2995: * there was a short read because we hit EOF, go ahead
2996: * and return. Otherwise fallthrough to buffered io for
2997: * the rest of the read. Buffered reads will not work for
2998: * DAX files, so don't bother trying.
2999: */
]
|
| 26/2 |
2026/07/22 19:13 |
llm |
expert |
0m
Model:
Tokens:
input: 22390
output: 34
thoughts: 997
Thoughts: **Analyzing a Potential Race Condition in Direct IO Read with Page Cache Writeback**
Okay, here's what's running through my head. `generic_file_read_iter` is doing a direct IO read. The code calls `kiocb_write_and_wait` *before* the direct IO to ensure data consistency. The reason is straightforward: if there are dirty pages in the page cache overlapping the read range, we *must* write them back to disk first so the direct IO reads the up-to-date data. Makes sense.
But then, I start thinking about edge cases, and that's where things get interesting. What if `mapping->nrpages` (the count of pages in the address space's mapping) is being decremented because pages are *being deleted* from the page cache, like during a truncation? These pages are no longer present, so writing them back is wasted effort and potentially harmful.
Now the deeper analysis kicks in. The code *shouldn't* be writing back pages that have been truncated! And I realize that, if `mapping->nrpages` reads as zero, but pages are *still* in the page cache, then we'd skip the writeback unnecessarily. This means we'd bypass the critical writeback operation when it's still needed, and we'd be reading stale data. How could this happen? Only if we get a zero value when the count is incrementing, or if we get a zero value when the count is decrementing.
Let's break it down. If the page count is incrementing from zero, then there are *no* dirty pages that need to be flushed. Direct IO doesn't need to write back a page that doesn't yet exist. If the page count is *decreasing* from non-zero, then this means that pages may or may not need to be written back.
Okay, let's look at the KCSAN report. The report gives us a change in page count, say `nrpages` changed from 6 to 5. Since both values are non-zero, the check to write back will work fine. But what if `nrpages` changes from 1 to 0?
Now we are in a potentially risky situation. The function `mapping_needs_writeback` might read a 1 or a 0. If it reads a 1, it calls the `filemap_fdatawrite_range` function to write back. If it reads a 0, the writeback will be skipped. But, if the page is gone, then the writeback is unnecessary! Skipping it is safe and correct.
Are there other scenarios? What about compiler optimization? Could the compiler load `mapping->nrpages` multiple times? In `mapping_needs_writeback`, it is just a simple function. I don't see any possibility of the compiler loading the variable outside of the function. No risk there.
Could we get a "fused access" or a "torn access"? Since `nrpages` is an aligned `unsigned long`, there's no risk of tearing. And it's not a loop, so the "fused access" concept doesn't apply. Ordering violations are what's important.
But is there a memory ordering requirement? If a page is added to the page cache, then the `mapping->nrpages` count will be incremented *after* the page is inserted into the xarray. Looking at the `__filemap_add_folio` function in `mm/filemap.c` will confirm this. So, in general, it looks like there's no race.
|
| 27/2 |
2026/07/22 19:13 |
tool |
read-file |
0m
Args: map[File:mm/filemap.c FirstLine:890 LineCount:40]
Results: map[Contents: 890: if (old) {
891: if (order > 0 && order > forder) {
892: unsigned int split_order = max(forder,
893: xas_try_split_min_order(order));
894:
895: /* How to handle large swap entries? */
896: BUG_ON(shmem_mapping(mapping));
897:
898: while (order > forder) {
899: xas_set_order(&xas, index, split_order);
900: xas_try_split(&xas, old, order);
901: if (xas_error(&xas))
902: goto unlock;
903: order = split_order;
904: split_order =
905: max(xas_try_split_min_order(
906: split_order),
907: forder);
908: }
909: xas_reset(&xas);
910: }
911: if (shadowp)
912: *shadowp = old;
913: }
914:
915: xas_store(&xas, folio);
916: if (xas_error(&xas))
917: goto unlock;
918:
919: mapping->nrpages += nr;
920:
921: /* hugetlb pages do not participate in page cache accounting */
922: if (!huge) {
923: lruvec_stat_mod_folio(folio, NR_FILE_PAGES, nr);
924: if (folio_test_pmd_mappable(folio))
925: lruvec_stat_mod_folio(folio,
926: NR_FILE_THPS, nr);
927: }
928:
929: unlock:
]
|
| 28/2 |
2026/07/22 19:13 |
llm |
expert |
0m
Model:
Tokens:
input: 24739
output: 27
thoughts: 744
Thoughts: **Analyzing the `mapping->nrpages` Concurrency**
Okay, so I'm thinking about the order of operations when dealing with `mapping->nrpages` and the xarray (XArray) in the page cache. Specifically, it seems pretty clear that the update to `mapping->nrpages` always happens *after* the `xas_store(&xas, folio)` or `xas_store(&xas, NULL)` calls, depending on whether we're adding or removing a page. I just confirmed this by quickly checking the `page_cache_delete_batch` function.
Therefore, when we decrement `mapping->nrpages`, the page has *already* been removed from the xarray. And when we increment it, the page has *already* been added.
Now, the critical question is what `mapping_needs_writeback` does with `mapping->nrpages`. It looks like if it reads a non-zero value for `mapping->nrpages`, then it initiates a `filemap_fdatawrite_range` scan of the xarray to find dirty pages. If it reads zero, it skips the scan, assuming there are no pages to write back.
Aha! This is where it gets interesting. I see a potential race condition, at least in theory. If a page is being *added*, there's a tiny window where the page is in the xarray *before* the `mapping->nrpages` is incremented. If `mapping_needs_writeback` is called during this window, it would read zero and skip the writeback. But I don't think this is a problem. The page has just been added, it can't be dirty (or if it is, it's newly added and there isn't an expectation for it to be written back yet).
What about the other way around: a page being *deleted*? `xas_store(&xas, NULL)` happens first, then `mapping->nrpages` is decremented. So there's a window where the page is removed from the xarray *before* `mapping->nrpages` is decremented. If `mapping_needs_writeback` is called in this window, it *will* read a non-zero value and call `filemap_fdatawrite_range`. But that's okay! `filemap_fdatawrite_range` will scan the now-empty xarray and find nothing. It'll just return. It's harmless, just a bit of wasted effort, and extremely rare in the first place.
I think we are good.
But wait, is there *any* other place in the kernel code where `mapping->nrpages` is read locklessly? I need to be absolutely sure. Let me see if there are any KCSAN reports or discussions about this sort of potential data race. I should start with a web/mailing list search, specifically for `mapping_needs_writeback` and `KCSAN` or mentions of `data-race`. Then, maybe a quick check of the kernel source to see if there are existing annotations like `data_race(mapping->nrpages)` or similar to see if anyone else has stumbled on this. Let's see...I'll grep for `data_race.*nrpages` and `READ_ONCE.*nrpages`.
|
| 29/2 |
2026/07/22 19:13 |
tool |
grepper |
0m
Args: map[Expression:(data_race|READ_ONCE)\(.*nrpages]
Results: map[Output:mm/shmem.c=437=bool shmem_recalc_inode(struct inode *inode, long alloced, long swapped)
--
mm/shmem.c-446- freed = info->alloced - info->swapped -
mm/shmem.c:447: READ_ONCE(inode->i_mapping->nrpages);
mm/shmem.c-448- /*
--
mm/shmem.c=1913=static struct folio *shmem_alloc_and_add_folio(struct vm_fault *vmf,
--
mm/shmem.c-1993- freed = pages + info->alloced - info->swapped -
mm/shmem.c:1994: READ_ONCE(mapping->nrpages);
mm/shmem.c-1995- if (freed > 0)
]
|
| 30/2 |
2026/07/22 19:13 |
llm |
expert |
0m
Model:
Tokens:
input: 26448
output: 31
thoughts: 594
Thoughts: **Analyzing Lockless Access to `nrpages` in the Linux Kernel**
Okay, so I've noticed `shmem.c` is using `READ_ONCE()` when accessing `mapping->nrpages`. That's good practice. But, I'm also seeing it's not being done in `mm/filemap.c`. Hmm, I wonder if there are other instances of lockless access to this variable. Let me check.
Right, `mapping_needs_writeback` reads `mapping->nrpages` directly.
```c
static bool mapping_needs_writeback(struct address_space *mapping)
{
if (dax_mapping(mapping))
return mapping->nrexceptional;
return mapping->nrpages;
}
```
Okay, since it's read locklessly, KCSAN is going to flag a data race if there's a concurrent write, which there definitely is in places like `page_cache_delete_batch` or `__filemap_add_folio`. The question is, is this race benign? Let's break it down.
First, this read is just a heuristic, a hint to skip calling `filemap_fdatawrite_range` if the page cache is empty. So, if it returns a stale value:
* If it returns 1 instead of 0, `filemap_fdatawrite_range` is called, sees an empty xarray, and returns immediately. Safe, albeit slightly suboptimal.
* If it returns 0 instead of 1, `filemap_fdatawrite_range` is skipped. This is only possible if a page was *just* added. Since it was added concurrently, it wouldn't have been expected to be written back by this request anyway, and it's not dirty yet.
Third, the read of `nrpages` is atomic because it's an aligned `unsigned long`, so there's no tearing. Also, there's no way a compiler could optimize in such a way that it would become a bug.
So, this looks pretty safe because the function is called once per writeback request.
Wait, let me double-check for other instances where `mapping->nrpages` is read locklessly and ensure this data race is addressed consistently. Even though it's benign, we should annotate those reads with `READ_ONCE()` or `data_race()` to prevent KCSAN from complaining and future compilers from causing problems. It's good defensive programming.
Alright, let's see if there are any existing reports or commits addressing this. Let's search for "mapping_needs_writeback" or "nrpages" in commit messages.
|
| 31/2 |
2026/07/22 19:13 |
tool |
git-log |
0m
Args: map[MessageRegexps:[nrpages] PathPrefix:mm/filemap.c]
Results: map[Output:7f0e07fb0289 dax: account DAX entries as nrpages
55635ba76ef9 fs: fix data invalidation in the cleancache during direct IO
3ddf40e8c319 mm: filemap: fix mapping->nrpages double accounting in fuse
865ffef3797d fs: fix fsync() error reporting
05fe478dd04e mm: write_cache_pages integrity fix
]
|
| 32/2 |
2026/07/22 19:13 |
llm |
expert |
1m
Model:
Tokens:
input: 27774
output: 23
thoughts: 28
Thoughts: Okay, here's my summary, taking the perspective of someone deeply familiar with kernel development:
**Investigating Potential Commit Overlap**
Alright, before I dive too deep into this specific issue, I need to get a broader perspective. Let's see... I'm going to run a quick search through the entire kernel history. My intuition tells me that the current behavior might have been addressed or touched upon in other commits β it's always possible a similar situation has been tackled before. I'm specifically looking for any previous work related to "mapping_needs_writeback" in commit messages. This will help me understand if there are any existing solutions, relevant discussions, or related design decisions that I should be aware of. I need to make sure I'm not reinventing the wheel or missing something crucial from earlier contexts. This comprehensive sweep is the smart way to begin.
|
| 33/2 |
2026/07/22 19:14 |
tool |
git-log |
0m
Args: map[MessageRegexps:[mapping_needs_writeback]]
Results: map[Output:7f0e07fb0289 dax: account DAX entries as nrpages
875d91b11a20 mm/filemap.c: rewrite mapping_needs_writeback in less fancy manner
]
|
| 34/2 |
2026/07/22 19:14 |
llm |
expert |
0m
Model:
Tokens:
input: 28073
output: 34
thoughts: 692
Thoughts: **KCSAN Warning Investigation: A Deep Dive into `mapping_needs_writeback`**
Alright, no immediate commit fixes this specific KCSAN warning, or perhaps it's a known benign race. First things first, let's consult the KCSAN documentation. We need to understand what constitutes "BENIGN" versus "HARMFUL" behavior in this context.
According to the documentation, "BENIGN" means the logic is sound, even with potential compiler optimizations or stale reads. Okay, let's see... it mentions things like reads for `/proc`, `/sys`, counters, or `pr_info`. It also highlights "Heuristic Hints." Now, *that* catches my eye.
"Heuristic Hints: A 'hint' flag where an old value only causes a slightly delayed update or a sub-optimal but safe fast-path." Aha! That sounds like a perfect fit for `mapping_needs_writeback`! This function serves as a hint to avoid the relatively expensive `filemap_fdatawrite_range` operation. It prevents that xarray scanning when there are no pages in the page cache.
Let's verify this. If `mapping_needs_writeback` returns a stale value, here's how it plays out:
* If it returns `true` (stale non-zero) when it should be `false` (zero), we just end up taking a sub-optimal but safe fast-path. We call `filemap_fdatawrite_range`, which does nothing and returns. No harm, no foul.
* If it returns `false` (stale zero) when it should be `true` (non-zero), this can only happen if `nrpages` was 0 and is now incrementing to 1. Because the page is being added concurrently, it couldn't have been dirtied yet, so we don't need to write anything back.
Hold on, let's think a bit more deeply. What if `nrpages` is decremented? Let's consider all the scenarios:
* `nrpages` decrements from 1 to 0, and we read 1? We call `filemap_fdatawrite_range` which will find nothing, which is safe.
* `nrpages` decrements from 1 to 0, and we read 0? We skip `filemap_fdatawrite_range`, which is also safe because no pages are left to write back.
Now, a quick thought about compiler optimizations. Could there be a compiler bug where the read of `mapping->nrpages` gets split into multiple instructions? No. `nrpages` is an `unsigned long`. That's word-sized, and on all supported architectures, word-sized reads and writes are atomic with respect to other CPUs. No tearing. However, to be fully compliant with the Linux kernel memory model and appease KCSAN, using `READ_ONCE()` or `data_race()` would be the recommended route.
Finally, to be absolutely sure, let's quickly check the definition of `mapping->nrpages` within the `struct address_space`. I'll go find the definition in `include/linux/fs.h` just to confirm my assumptions.
|
| 35/2 |
2026/07/22 19:15 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:include/linux/fs.h Name:address_space]
Results: map[SourceCode: 453: /**
454: * struct address_space - Contents of a cacheable, mappable object.
455: * @host: Owner, either the inode or the block_device.
456: * @i_pages: Cached pages.
457: * @invalidate_lock: Guards coherency between page cache contents and
458: * file offset->disk block mappings in the filesystem during invalidates.
459: * It is also used to block modification of page cache contents through
460: * memory mappings.
461: * @gfp_mask: Memory allocation flags to use for allocating pages.
462: * @i_mmap_writable: Number of VM_SHARED, VM_MAYWRITE mappings.
463: * @i_mmap: Tree of private and shared mappings.
464: * @i_mmap_rwsem: Protects @i_mmap and @i_mmap_writable.
465: * @nrpages: Number of page entries, protected by the i_pages lock.
466: * @writeback_index: Writeback starts here.
467: * @a_ops: Methods.
468: * @flags: Error bits and flags (AS_*).
469: * @wb_err: The most recent error which has occurred.
470: * @i_private_lock: For use by the owner of the address_space.
471: */
472: struct address_space {
473: struct inode *host;
474: struct xarray i_pages;
475: struct rw_semaphore invalidate_lock;
476: gfp_t gfp_mask;
477: atomic_t i_mmap_writable;
478: struct rb_root_cached i_mmap;
479: unsigned long nrpages;
480: pgoff_t writeback_index;
481: const struct address_space_operations *a_ops;
482: unsigned long flags;
483: errseq_t wb_err;
484: spinlock_t i_private_lock;
485: struct rw_semaphore i_mmap_rwsem;
486: } __attribute__((aligned(sizeof(long)))) __randomize_layout;
SourceFile:include/linux/fs.h]
|
| 36/2 |
2026/07/22 19:15 |
llm |
expert |
0m
Model:
Tokens:
input: 30150
output: 15
thoughts: 1257
Thoughts: **Analyzing the Data Race in `mapping_needs_writeback`**
Okay, so I'm looking at this data race reported by KCSAN concerning `mapping->nrpages`. I see that `nrpages` is an `unsigned long` and, according to the comment, it represents the number of page entries, protected by the `i_pages` lock. Hmm, is it *always* protected? Let me quickly verify that.
I'm checking `page_cache_delete_batch` and I see it holds `xa_lock_irq(&mapping->i_pages)` before decrementing `nrpages` indirectly. Also, `__filemap_add_folio` uses the same lock. But `mapping_needs_writeback` *reads* `mapping->nrpages` without any locking. That's the problem. KCSAN is right; that's a classic data race scenario.
Now, let's see where `mapping_needs_writeback` is called. Ah, `filemap_write_and_wait_range` seems important. It calls `mapping_needs_writeback` and *if* that returns `false`, it skips both the write and wait operations. That's a potential issue. Could we miss dirty pages because of a stale read of `nrpages`?
If there are dirty pages, `nrpages` must be at least 1, right? How could it read as zero if there *are* dirty pages? Maybe if pages are *just* added. If the dirty pages were *just* added to the page cache, and `mapping->nrpages` was zero before, but is now incrementing to one.
Let's think. If those pages were *just* added concurrently, the caller probably didn't expect them to be written back immediately. Also, if the pages were added *before* the call to `filemap_write_and_wait_range`, `nrpages` would already be non-zero. So there's no issue there.
But what if `nrpages` is decremented from 1 to 0? The page cache is now empty. Reading 1 (stale) makes `filemap_fdatawrite_range` do nothing. Reading 0 skips it which is also safe because the cache is empty.
What about decrementing from 2 to 1? Both stale and non-stale reads will return true, which is correct.
So, it's safe! This is a "heuristic hint" data race: the race is benign!
Let's quickly double-check that this is correct. Yes, it's benign. Time to call `set-results` with `Benign: true`.
|
| 37/2 |
2026/07/22 19:15 |
tool |
set-results |
0m
Args:
Results:
|
| 38/2 |
2026/07/22 19:15 |
llm |
expert |
1m
Model:
Tokens:
input: 32014
output: 756
thoughts: 682
Thoughts: **My Analysis of the Potential Race Condition**
Okay, here's how I see this, let's craft a really clear and well-structured analysis. My goal is to produce a report that's both comprehensive and concise.
First, the core of the issue: We're looking at a potential race condition involving `delete_from_page_cache_batch` (specifically `page_cache_delete_batch`), `filemap_write_and_wait_range` (specifically `mapping_needs_writeback`), and the `struct address_space::nrpages` variable.
Here's the breakdown:
1. **The Race Scenario:**
* `delete_from_page_cache_batch` modifies `mapping->nrpages` while holding the `xa_lock_irq(&mapping->i_pages)`.
* `filemap_write_and_wait_range` reads `mapping->nrpages` locklessly through a fast path via `mapping_needs_writeback` to decide if a writeback operation is needed.
2. **Synchronization Context:** `delete_from_page_cache_batch` is protected via locking, whereas `filemap_write_and_wait_range` uses a lockless read, which raises our concern.
3. **My Final Classification:** **BENIGN** (Thankfully!)
4. **Reasoning for Benign Classification:**
* This race is structurally tolerant of stale reads. We're safe from any corruption.
* **No Tearing:** `nrpages` is an `unsigned long` and, therefore, is an aligned, atomic read/write on all our supported architectures.
* **Heuristic Nature:** `mapping_needs_writeback` is a heuristic. It's an optimization to skip writeback calls if the page cache is empty.
* **Stale Read Analysis:** Let's trace all of the possible state transitions:
* If `nrpages` decrements from 1 to 0:
* Reading 1 (stale): We still call `filemap_fdatawrite_range`. Even though it's technically stale, it scans the empty xarray and returns without doing anything. It's safe. It's just a slightly slower fast path.
* Reading 0: We skip writeback and that's the correct behavior since the page cache is empty.
* If `nrpages` increments from 0 to 1:
* Reading 0 (stale): We skip writeback. This is safe, the page was added concurrently after the writeback request started, so we don't expect it to be written back.
* Reading 1: We call writeback. That's also perfectly safe.
* If `nrpages` decrements from $N$ to $N-1$ (where $N > 1$):
* Reading either $N$ or $N-1$ will return a non-zero value, so `mapping_needs_writeback` returns `true` in both cases, which is correct.
* **No Compiler Issues:** The read isn't within a polling loop, so we are not concerned with the compiler hoisting a value or producing any kind of livelock. We're safe here.
5. **Recommended Annotations:**
* I'd suggest we wrap the read of `mapping->nrpages` in `mapping_needs_writeback` with either `READ_ONCE()`, or use `data_race()`, to explicitly document the lockless access and quiet KCSAN. A small, explicit annotation helps make the code more obvious.
This should be a very complete and professional analysis! I believe this nails everything the prompt asked for.
|
| 39/1 |
2026/07/22 19:16 |
action |
format-explanation |
0m
Results: map[Explanation:### Race Summary
* **Function A:** `delete_from_page_cache_batch` (via `page_cache_delete_batch`)
* **Function B:** `filemap_write_and_wait_range` (via `mapping_needs_writeback`)
* **Variable/Field:** `mapping->nrpages` (of `struct address_space`)
---
### Synchronization Context
* **Writer (`delete_from_page_cache_batch`):** Modifies `mapping->nrpages` under
the protection of the `xa_lock_irq(&mapping->i_pages)` spinlock.
* **Reader (`filemap_write_and_wait_range`):** Reads `mapping->nrpages`
locklessly via `mapping_needs_writeback` to check if there are any pages in the
page cache that might require writeback.
---
### Final Classification: **BENIGN**
---
### Reasoning
The data race is a classic **heuristic hint** race and is completely benign for
the following reasons:
1. **No Torn Reads:** `nrpages` is an `unsigned long` aligned to a word
boundary. On all supported architectures, word-sized reads and writes are
atomic, meaning there is no risk of reading a torn or corrupted value.
2. **Heuristic Hint/Fast-path Optimization:** The check in
`mapping_needs_writeback` is a fast-path optimization to avoid the overhead of
calling `filemap_fdatawrite_range` (which scans the page cache xarray) when the
page cache is empty.
3. **Tolerance to Stale Values:**
* **Case 1: `nrpages` decrements from 1 to 0.**
* If the reader reads `1` (stale): It calls `filemap_fdatawrite_range`,
which scans the empty xarray, finds nothing, and returns. This is a safe,
slightly sub-optimal fast-path.
* If the reader reads `0`: It skips writeback, which is correct because the
page cache is empty.
* **Case 2: `nrpages` increments from 0 to 1.**
* If the reader reads `0` (stale): It skips writeback. This is safe because
the page was added concurrently with or after the writeback request
started, so the caller has no expectation of it being written back.
* If the reader reads `1`: It calls writeback, which is also safe.
* **Case 3: `nrpages` changes between non-zero values (e.g., 6 to 5).**
* Both values are non-zero, so `mapping_needs_writeback` returns `true` in
both cases, leading to the correct behavior.
4. **No Compiler Optimization/Livelock Issues:** The read is not part of a
polling loop, so there is no risk of the compiler hoisting the load out of a
loop or causing a livelock.
---
### Recommended Annotations
To document the lockless access and silence KCSAN, the read of
`mapping->nrpages` in `mapping_needs_writeback` should be annotated with
`READ_ONCE()` or wrapped in `data_race()`:
```c
static bool mapping_needs_writeback(struct address_space *mapping)
{
if (dax_mapping(mapping))
return mapping->nrexceptional;
return READ_ONCE(mapping->nrpages);
}
```]
|