| 0/0 |
2026/04/06 20:46 |
flow |
assessment-kcsan |
20m
Results: map[Benign:true Confident:true Explanation:- **Race Summary:** `pipe_release` vs `pipe_wait_writable` on `pipe->files`.
- **Synchronization context:** `pipe_release` (via `put_pipe_info`) decrements
`pipe->files` under `inode->i_lock`. `pipe_wait_writable` (via `pipe_lock`)
reads `pipe->files` without holding `inode->i_lock` to determine if it should
acquire `pipe->mutex`.
- **Final Classification:** **BENIGN**
### Reasoning
The race is on the `pipe->files` field, which tracks the number of `struct file`
objects referring to a pipe.
1. **Contextual Guarantee:** In the reported race, `pipe_wait_writable` is
called from `vmsplice_to_pipe`. This syscall holds a reference to the pipe's
file descriptor (via `fdget`). This ensures that `pipe->files` is at least 1 as
long as the syscall is active.
2. **Safe Transition:** The KCSAN report shows `pipe->files` changing from `2`
to `1`. This occurs because another thread closed a different file descriptor
for the same pipe.
3. **Logical Tolerance:** In `pipe_lock(pipe)`, the code checks `if
(pipe->files)`. Since the current thread holds a file reference, `pipe->files`
will be at least 1. Whether the read sees `2` or `1`, the result of the `if`
condition is `true`, and `mutex_lock(&pipe->mutex)` is correctly called.
4. **No Livelock/UAF:** The check in `pipe_lock` is a "belt and suspenders"
safety measure. If `pipe->files` were 0, it would mean the pipe is being
destroyed, but this is impossible while a thread holds a valid file reference.
Even if the compiler were to re-read the value, it would remain non-zero.
The race is structurally tolerant to stale or concurrent updates because the
outcome of the branch (`true`) is invariant in the context of an active file
operation.
### Recommended Annotations
The accesses to `pipe->files` in `pipe_lock()` and `pipe_unlock()` should be
marked with `data_race()` or `READ_ONCE()` to acknowledge the lockless read and
silence KCSAN.
```c
void pipe_lock(struct pipe_inode_info *pipe)
{
if (data_race(pipe->files))
mutex_lock(&pipe->mutex);
}
```]
|
| 1/1 |
2026/04/06 20:46 |
action |
kernel-checkouter |
1m
Results: map[KernelSrc:/app/workdir/cache/src/de2829276ac4973ff7bb50d225da160402a31231]
|
| 2/1 |
2026/04/06 20:47 |
action |
kernel-builder |
9m
Results: map[KernelObj:/app/workdir/cache/build/06092afd896eff1a64d01b734e354fc8b1ca5cac]
|
| 3/1 |
2026/04/06 20:57 |
action |
codesearch-prepare |
7m
Results: map[Index:codesearch-index]
|
| 4/1 |
2026/04/06 21:04 |
agent |
expert |
2m
Model:
Results: map[Benign:true Confident:true]
Instruction: # KCSAN Data Race Severity Analysis Protocol
You are an expert Linux kernel concurrency engineer analyzing a Kernel
Concurrency Sanitizer (KCSAN) data race report to determine if it is
**BENIGN** or **HARMFUL**.
## 1. CLASSIFICATIONS
### **BENIGN (Truly Benign)**
The logic is sound and structurally tolerant to compiler optimizations or
stale/torn reads.
- **Diagnostics/Stats:** Reads used only for '/proc', '/sys', counters, or
'pr_info'.
- **Heuristic Hints:** A "hint" flag where an old value only causes a
slightly delayed update or a sub-optimal but safe fast-path.
- **Single-Writer Flag Updates:** A single writer updating flags where the
concurrent read is a simple bitwise check (e.g., 'flags & MASK'). These are
historically tolerated, assuming neither "Fused Accesses" nor "Ordering
Violations" are relevant in this context.
- **Marked Reloads:** A load feeding into a 'cmpxchg()' loop or checked
against a later 'READ_ONCE()' reload.
- **Safe Overwrites:** Writing the same value already present.
### **HARMFUL (Logic Bug or Marking Required)**
The race causes incorrect behavior due to a synchronization failure or
because missing annotations allow the compiler to break the algorithm.
**Marking Required for Correctness:**
The algorithm is logically sound but requires annotations ('READ_ONCE()',
'WRITE_ONCE()', 'smp_load_acquire()', 'smp_store_release()', etc.) to be safe.
- **Fused Accesses:** The compiler might merge accesses or hoist a load out
of a loop, breaking polling/wait loops (livelocks).
- **Torn Accesses:** A large access (e.g., 64-bit on 32-bit arch) might be
split into multiple non-atomic accesses. Note that 'READ_ONCE()' does **not**
guarantee atomicity for 64-bit variables on 32-bit architectures.
- **Ordering Violations:** The race breaks a "happens-before" relationship
(requires primitives with implied or explicit memory barriers).
**Logic Bugs:**
A fundamental synchronization failure. Marking accesses will **not** fix it;
the logic itself must change.
- **Pointers/Lifecycle:** The racing variable is a pointer being dereferenced
or a refcount governing object lifecycle (Use-After-Free risk).
- **Control Flow:** The variable guards a critical section, memory allocation,
or hardware command.
- **Bitfields:** Concurrent writes to different bits in the same word.
Compilers often use non-atomic read-modify-write sequences, meaning a
write to 'bit_A' can "clobber" a concurrent write to 'bit_B'. However,
do not blindly assume all bitfield accesses are harmful; you must prove
that a concurrent write actually clobbers another in a way that breaks
logic.
- **Complex Structures:** Races on shared lists, trees, or hashmaps.
- **Lossy Updates:** Concurrent plain RMW operations (e.g., 'var++') on
non-diagnostic variables where every increment must be preserved.
- **State Machines:** Races allowing a state machine to bypass transitions
or enter an invalid state.
- **Adjacent Unsynchronized Operations:** Consider races happening at the
same time. For example, if both threads execute 'struct->has_elements = true;
list_add(node, &struct->list);', the race on 'has_elements' implies an
adjacent race on 'list_head', which is HARMFUL.
## 2. RESEARCH & ANALYSIS WORKFLOW
1. **Locate the Race:** Find the exact variables and functions in the stack
traces. **Use codesearch tools to read the actual source code and
confirm all assumptions.** Do not speculate about hypothetical compiler
behaviors or theoretical dangers (e.g., dismissing something as
"fundamentally unsafe") without tracing the actual data flow to a crash.
2. **Contextualize:** Identify held locks, RCU sections, or interrupt
contexts. Explain why the race is possible (e.g., "Thread A holds 'lock',
but Thread B is a lockless reader").
3. **Data Flow:** Follow the racing variable into its subsequent uses. If
the reader sees a stale or torn value, what branch is taken? Does it lead
to a pointer dereference or 'BUG_ON()'? **You must prove harm via this
data flow.**
4. **Failure Sequencing (if HARMFUL):** Construct a two-column execution trace
demonstrating the failure.
*Example format:*
CPU0 CPU1
function_a()
// lockless read
ptr = global_ptr
<PREEMPT>
function_b()
lock(&my_lock)
global_ptr = NULL
unlock(&my_lock)
kfree(ptr)
if (ptr)
*ptr = 1; // -> Use-After-Free!
## 3. OUTPUT FORMAT
- **Race Summary:** '[Function A]' vs '[Function B]' on '[Variable/Field]'.
- **Synchronization context:** List held locks and explain the lack of mutual
exclusion.
- **Final Classification:** **BENIGN** or **HARMFUL**.
**If BENIGN:**
- **Reasoning:** Briefly explain why the race is structurally tolerant to stale
or torn reads.
- **Recommended Annotations:** Suggest appropriate annotations (e.g.,
'READ_ONCE()', 'data_race()').
**If HARMFUL:**
- **Failure Reasoning:** Explain and provide a two-column interleaving showing
exactly how the race leads to a failure such as corruption or kernel crash.
- **Recommended Fix (ONLY if fix is trivial):** Suggest a structural fix (e.g.,
"Hold 'mapping->i_pages' lock", "Convert to 'atomic_t'") or required memory
ordering annotations (e.g., "Wrap in 'READ_ONCE()'", "Use
'smp_load_acquire()'").
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt:
The data race report is:
==================================================================
BUG: KCSAN: data-race in pipe_release / pipe_wait_writable
read-write to 0xffff88811bdfb5ac of 4 bytes by task 9866 on cpu 1:
put_pipe_info fs/pipe.c:714 [inline]
pipe_release+0x16a/0x1d0 fs/pipe.c:744
__fput+0x29b/0x650 fs/file_table.c:469
fput_close_sync+0xad/0x190 fs/file_table.c:574
__do_sys_close fs/open.c:1509 [inline]
__se_sys_close fs/open.c:1494 [inline]
__x64_sys_close+0x55/0xf0 fs/open.c:1494
x64_sys_call+0x275d/0x3020 arch/x86/include/generated/asm/syscalls_64.h:4
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x12c/0x370 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
read to 0xffff88811bdfb5ac of 4 bytes by task 9863 on cpu 0:
pipe_lock fs/pipe.c:90 [inline]
pipe_wait_writable+0x202/0x250 fs/pipe.c:1084
wait_for_space fs/splice.c:1272 [inline]
vmsplice_to_pipe fs/splice.c:1551 [inline]
__do_sys_vmsplice fs/splice.c:1608 [inline]
__se_sys_vmsplice+0x42b/0xfd0 fs/splice.c:1578
__x64_sys_vmsplice+0x55/0x70 fs/splice.c:1578
x64_sys_call+0x28e9/0x3020 arch/x86/include/generated/asm/syscalls_64.h:279
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x12c/0x370 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
value changed: 0x00000002 -> 0x00000001
Reported by Kernel Concurrency Sanitizer on:
CPU: 0 UID: 0 PID: 9863 Comm: syz.1.2745 Not tainted syzkaller #0 PREEMPT(full)
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 03/18/2026
==================================================================
Reply: - **Race Summary:** `pipe_release` vs `pipe_wait_writable` on `pipe->files`.
- **Synchronization context:** `pipe_release` (via `put_pipe_info`) decrements `pipe->files` under `inode->i_lock`. `pipe_wait_writable` (via `pipe_lock`) reads `pipe->files` without holding `inode->i_lock` to determine if it should acquire `pipe->mutex`.
- **Final Classification:** **BENIGN**
### Reasoning
The race is on the `pipe->files` field, which tracks the number of `struct file` objects referring to a pipe.
1. **Contextual Guarantee:** In the reported race, `pipe_wait_writable` is called from `vmsplice_to_pipe`. This syscall holds a reference to the pipe's file descriptor (via `fdget`). This ensures that `pipe->files` is at least 1 as long as the syscall is active.
2. **Safe Transition:** The KCSAN report shows `pipe->files` changing from `2` to `1`. This occurs because another thread closed a different file descriptor for the same pipe.
3. **Logical Tolerance:** In `pipe_lock(pipe)`, the code checks `if (pipe->files)`. Since the current thread holds a file reference, `pipe->files` will be at least 1. Whether the read sees `2` or `1`, the result of the `if` condition is `true`, and `mutex_lock(&pipe->mutex)` is correctly called.
4. **No Livelock/UAF:** The check in `pipe_lock` is a "belt and suspenders" safety measure. If `pipe->files` were 0, it would mean the pipe is being destroyed, but this is impossible while a thread holds a valid file reference. Even if the compiler were to re-read the value, it would remain non-zero.
The race is structurally tolerant to stale or concurrent updates because the outcome of the branch (`true`) is invariant in the context of an active file operation.
### Recommended Annotations
The accesses to `pipe->files` in `pipe_lock()` and `pipe_unlock()` should be marked with `data_race()` or `READ_ONCE()` to acknowledge the lockless read and silence KCSAN.
```c
void pipe_lock(struct pipe_inode_info *pipe)
{
if (data_race(pipe->files))
mutex_lock(&pipe->mutex);
}
```
|
| 5/2 |
2026/04/06 21:04 |
llm |
expert |
0m
Model:
Tokens:
input: 4794
output: 115
thoughts: 357
Thoughts: **Data Race Investigation: `pipe_release` and `pipe_wait_writable`**
Okay, so the KCSAN report just flagged a data race. Seems there's a problem between `pipe_release`, which is writing, and `pipe_wait_writable`, which is reading, on a 4-byte variable at that memory address, `0xffff88811bdfb5ac`. The value went from `0x2` down to `0x1`. Intriguing.
Let's dive into those stack traces. On the write side (Task 9866, CPU 1), we've got `put_pipe_info` called inlined within `pipe_release`, which is then called by `__fput` β seems like a standard file descriptor release path, and ultimately `sys_close`. This strongly suggests we're dealing with a reference count decrement.
On the read side (Task 9863, CPU 0), the trace goes through `pipe_wait_writable`, after a call to `pipe_lock`. This is followed by `wait_for_space` and finally `vmsplice_to_pipe`, which is called as part of `sys_vmsplice`. This means we're dealing with a pipe waiting for more data to be written.
Given the context, I'm betting this variable is something like `pipe->files` or related to the reference count on the pipe information itself. Hang on, `put_pipe_info` is *definitely* involved in decrementing a reference count. Thatβs probably it. I need to get those lines of code up and check.
|
| 6/2 |
2026/04/06 21:04 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/pipe.c IncludeLines:true Name:put_pipe_info]
Results: map[SourceCode: 709: static void put_pipe_info(struct inode *inode, struct pipe_inode_info *pipe)
710: {
711: int kill = 0;
712:
713: spin_lock(&inode->i_lock);
714: if (!--pipe->files) {
715: inode->i_pipe = NULL;
716: kill = 1;
717: }
718: spin_unlock(&inode->i_lock);
719:
720: if (kill)
721: free_pipe_info(pipe);
722: }
SourceFile:fs/pipe.c]
|
| 7/2 |
2026/04/06 21:04 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/pipe.c IncludeLines:true Name:pipe_wait_writable]
Results: map[SourceCode:1080: void pipe_wait_writable(struct pipe_inode_info *pipe)
1081: {
1082: pipe_unlock(pipe);
1083: wait_event_interruptible(pipe->wr_wait, pipe_writable(pipe));
1084: pipe_lock(pipe);
1085: }
SourceFile:fs/pipe.c]
|
| 8/2 |
2026/04/06 21:04 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/pipe.c IncludeLines:true Name:pipe_lock]
Results: map[SourceCode: 88: void pipe_lock(struct pipe_inode_info *pipe)
89: {
90: if (pipe->files)
91: mutex_lock(&pipe->mutex);
92: }
SourceFile:fs/pipe.c]
|
| 9/2 |
2026/04/06 21:04 |
llm |
expert |
0m
Model:
Tokens:
input: 6036
output: 37
thoughts: 483
Thoughts: **Race Condition in Pipe File Handling**
Okay, so I'm looking at this `pipe->files` variable and I'm seeing a potential race condition. The scenario involves the `pipe->files` counter, used to track the number of open file descriptors associated with a pipe.
Specifically, I see a write to `pipe->files` in `put_pipe_info`, which is called by `pipe_release`. This is a decrement operation using `--pipe->files`, protected by a spinlock `inode->i_lock`. This suggests that the decrement *should* be atomic, at least in the context of file release.
However, I'm concerned because `pipe_lock`, which is called by `pipe_wait_writable`, checks `pipe->files` directly, as in a read, without holding any apparent lock. The code in `pipe_lock` goes something like this: "If `pipe->files` is not zero, grab a mutex." The `pipe_wait_writable` function then calls `pipe_lock`. This is where the red flag goes up. If `pipe->files` is read in `pipe_lock` without synchronization, and then decremented by `pipe_release` without holding the mutex which `pipe_lock` later grabs if the value is not zero, it's possible for a race condition to occur. This could lead to a situation where the mutex might be acquired even after the file descriptors have been released, potentially causing issues.
To confirm my suspicions, I need to take a closer look at the source code of `pipe_release`. I need to make sure I am seeing the whole picture to be sure that the `put_pipe_info` calls are properly synchronized. There's also the question of whether other functions might access `pipe->files` without the necessary locks, and I will need to look further.
|
| 10/2 |
2026/04/06 21:04 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/pipe.c IncludeLines:true Name:pipe_release]
Results: map[SourceCode: 724: static int
725: pipe_release(struct inode *inode, struct file *file)
726: {
727: struct pipe_inode_info *pipe = file->private_data;
728:
729: mutex_lock(&pipe->mutex);
730: if (file->f_mode & FMODE_READ)
731: pipe->readers--;
732: if (file->f_mode & FMODE_WRITE)
733: pipe->writers--;
734:
735: /* Was that the last reader or writer, but not the other side? */
736: if (!pipe->readers != !pipe->writers) {
737: wake_up_interruptible_all(&pipe->rd_wait);
738: wake_up_interruptible_all(&pipe->wr_wait);
739: kill_fasync(&pipe->fasync_readers, SIGIO, POLL_IN);
740: kill_fasync(&pipe->fasync_writers, SIGIO, POLL_OUT);
741: }
742: mutex_unlock(&pipe->mutex);
743:
744: put_pipe_info(inode, pipe);
745: return 0;
746: }
SourceFile:fs/pipe.c]
|
| 11/2 |
2026/04/06 21:04 |
llm |
expert |
0m
Model:
Tokens:
input: 7331
output: 39
thoughts: 362
Thoughts: **Analyzing the Pipe Locking and Unlocking Sequence**
Alright, let's break this down. `pipe_release` seems straightforward: it releases the `pipe->mutex` and then calls `put_pipe_info`. `put_pipe_info` is the interesting part; it decrements `pipe->files` while holding `inode->i_lock`. That's a classic reference counting and resource management pattern, so far so good.
Now, `pipe_wait_writable`, called by `vmsplice_to_pipe`, is where it gets interesting. I see a `pipe_unlock(pipe)` call, which, as expected, unlocks the `pipe->mutex`. Then, there's a `wait_event_interruptible` call, waiting for `pipe_writable(pipe)` to become true. The code then re-acquires the lock with `pipe_lock(pipe)`. Standard blocking mechanism.
Here's the rub: `pipe_lock` *conditionally* locks the mutex. It only does so if `pipe->files` is non-zero. That's the key. If `pipe->files` is zero, there's no lock re-acquisition! But `pipe_wait_writable` is *supposed* to re-acquire the lock after the wait. That looks like a race condition waiting to happen.
If `pipe->files` drops to zero while `pipe_wait_writable` is waiting, that essentially signals that the pipe is being destroyed. The waiter might be unblocked, but without re-acquiring the lock, it could lead to all sorts of nasty consequences, like accessing freed memory, or concurrent access to inconsistent state.
I need to look at the `pipe_inode_info` struct definition to fully grasp the state transitions here. This looks like a classic synchronization problem, and potentially a critical one.
|
| 12/2 |
2026/04/06 21:04 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/pipe.c IncludeLines:true Name:pipe_inode_info]
Results: map[SourceCode: 61: /**
62: * struct pipe_inode_info - a linux kernel pipe
63: * @mutex: mutex protecting the whole thing
64: * @rd_wait: reader wait point in case of empty pipe
65: * @wr_wait: writer wait point in case of full pipe
66: * @pipe_index: the pipe indeces
67: * @note_loss: The next read() should insert a data-lost message
68: * @max_usage: The maximum number of slots that may be used in the ring
69: * @ring_size: total number of buffers (should be a power of 2)
70: * @nr_accounted: The amount this pipe accounts for in user->pipe_bufs
71: * @tmp_page: cached released page
72: * @readers: number of current readers of this pipe
73: * @writers: number of current writers of this pipe
74: * @files: number of struct file referring this pipe (protected by ->i_lock)
75: * @r_counter: reader counter
76: * @w_counter: writer counter
77: * @poll_usage: is this pipe used for epoll, which has crazy wakeups?
78: * @fasync_readers: reader side fasync
79: * @fasync_writers: writer side fasync
80: * @bufs: the circular array of pipe buffers
81: * @user: the user who created this pipe
82: * @watch_queue: If this pipe is a watch_queue, this is the stuff for that
83: **/
84: struct pipe_inode_info {
85: struct mutex mutex;
86: wait_queue_head_t rd_wait, wr_wait;
87:
88: union pipe_index;
89:
90: unsigned int max_usage;
91: unsigned int ring_size;
92: unsigned int nr_accounted;
93: unsigned int readers;
94: unsigned int writers;
95: unsigned int files;
96: unsigned int r_counter;
97: unsigned int w_counter;
98: bool poll_usage;
99: #ifdef CONFIG_WATCH_QUEUE
100: bool note_loss;
101: #endif
102: struct page *tmp_page[2];
103: struct fasync_struct *fasync_readers;
104: struct fasync_struct *fasync_writers;
105: struct pipe_buffer *bufs;
106: struct user_struct *user;
107: #ifdef CONFIG_WATCH_QUEUE
108: struct watch_queue *watch_queue;
109: #endif
110: };
SourceFile:include/linux/pipe_fs_i.h]
|
| 13/2 |
2026/04/06 21:04 |
llm |
expert |
0m
Model:
Tokens:
input: 8934
output: 37
thoughts: 4134
Thoughts: **Analyzing the Pipe Locking Mechanism**
Okay, let's break this down. The core of the issue seems to be a race condition involving the `pipe->files` counter within the pipe locking mechanisms. The comment indicates that this counter, which represents the number of file descriptors referencing a pipe, isn't protected by a lock when read in `pipe_lock`. This immediately strikes me as a potential problem.
Here's my thought process:
1. **The Code:** `pipe_lock` checks `pipe->files` before acquiring a mutex. `put_pipe_info`, which is called during file closure, *decrements* `pipe->files` under `i_lock`. The race happens when `pipe_lock` reads `pipe->files` while `put_pipe_info` is modifying it.
2. **The Potential Outcomes:** If `pipe_lock` reads a non-zero value, the mutex is taken. If it reads zero, the mutex isn't taken.
3. **The Context:** `pipe_lock` is called by functions operating on the pipe (e.g., `vmsplice`, `pipe_read`, `pipe_write`). In all of these cases, the existence of the call *should* imply that at least one file descriptor is referencing the pipe, so `pipe->files` *should* be at least 1. The key is to see where the call is coming from.
4. **`pipe_wait_writable`:** This function, called during `vmsplice_to_pipe`, is relevant. The task running the operation *must* have an open file descriptor. Hence, `pipe->files` should be at least 1.
5. **`pipe_release` and the Decrement:** When a file is closed, `pipe_release` is called, which decrements `pipe->files` through `put_pipe_info`. If multiple files point to the same pipe, `pipe->files` would be > 1.
6. **The KCSAN Report:** The KCSAN report showing the change from 2 to 1 for `pipe->files`. That shows that while the race is happening, it seems to not cause any issue because `pipe->files` will always be a non-zero value.
7. **The "Harmless" Scenario (Maybe):** If `pipe_lock` sees a non-zero value, it calls `mutex_lock`. Whether it sees 2 or 1 doesn't seem to matter. It's the *zero* case that concerns me. The call to `pipe_lock` means we have a valid `pipe` pointer which is the issue. If it sees zero, it *doesn't* take the mutex, yet we're supposed to have a reference to the pipe, which shouldn't happen.
8. **The "Destruction" Case:** If `pipe->files` becomes zero, the pipe is being freed by `free_pipe_info`. If `pipe_lock` reads zero, it skips taking the mutex. This is where I can see there could be a Use-After-Free scenario. However, the existing calls to the pipe functions mean that someone is still holding a valid pointer to the pipe.
9. **The Design Question:** Why does `pipe_lock` check `pipe->files` *at all*? If the pipe pointer is valid, shouldn't we always be taking the mutex? What's the historical reasoning behind this check? This seems like a potential design flaw.
10. **File Operations and References:** Functions like `pipe_read`, `pipe_write`, `pipe_poll`, etc., are all file operations, meaning they operate on a `struct file *`. These operations mean `pipe->files` should always be >= 1.
11. **FIFO and the "Open" Case:** I need to consider FIFO's. It looks like `fifo_open` increments the `pipe->files` counter when a FIFO file is opened. This means the counter should always be greater than 0 if a file is open.
12. **The Race Realized and Possible Harm:** The race is real: `pipe_lock` reads without a lock; `put_pipe_info` modifies with a lock. If `pipe->files` is zero, we might have a Use-After-Free. But the calling function still has a valid pointer, meaning this may not happen.
13. **Why no race when `vmsplice` is active?** `vmsplice` uses `fdget`. `fdget` should increment the file reference. So `fput` won't be called. So `pipe_release` won't be called and `pipe->files` shouldn
|
| 14/2 |
2026/04/06 21:05 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/fifo.c IncludeLines:true Name:fifo_open]
Results: map[SourceCode:1121: static int fifo_open(struct inode *inode, struct file *filp)
1122: {
1123: bool is_pipe = inode->i_fop == &pipeanon_fops;
1124: struct pipe_inode_info *pipe;
1125: int ret;
1126:
1127: filp->f_pipe = 0;
1128:
1129: spin_lock(&inode->i_lock);
1130: if (inode->i_pipe) {
1131: pipe = inode->i_pipe;
1132: pipe->files++;
1133: spin_unlock(&inode->i_lock);
1134: } else {
1135: spin_unlock(&inode->i_lock);
1136: pipe = alloc_pipe_info();
1137: if (!pipe)
1138: return -ENOMEM;
1139: pipe->files = 1;
1140: spin_lock(&inode->i_lock);
1141: if (unlikely(inode->i_pipe)) {
1142: inode->i_pipe->files++;
1143: spin_unlock(&inode->i_lock);
1144: free_pipe_info(pipe);
1145: pipe = inode->i_pipe;
1146: } else {
1147: inode->i_pipe = pipe;
1148: spin_unlock(&inode->i_lock);
1149: }
1150: }
1151: filp->private_data = pipe;
1152: /* OK, we have a pipe and it's pinned down */
1153:
1154: mutex_lock(&pipe->mutex);
1155:
1156: /* We can only do regular read/write on fifos */
1157: stream_open(inode, filp);
1158:
1159: switch (filp->f_mode & (FMODE_READ | FMODE_WRITE)) {
1160: case FMODE_READ:
1161: /*
1162: * O_RDONLY
1163: * POSIX.1 says that O_NONBLOCK means return with the FIFO
1164: * opened, even when there is no process writing the FIFO.
1165: */
1166: pipe->r_counter++;
1167: if (pipe->readers++ == 0)
1168: wake_up_partner(pipe);
1169:
1170: if (!is_pipe && !pipe->writers) {
1171: if ((filp->f_flags & O_NONBLOCK)) {
1172: /* suppress EPOLLHUP until we have
1173: * seen a writer */
1174: filp->f_pipe = pipe->w_counter;
1175: } else {
1176: if (wait_for_partner(pipe, &pipe->w_counter))
1177: goto err_rd;
1178: }
1179: }
1180: break;
1181:
1182: case FMODE_WRITE:
1183: /*
1184: * O_WRONLY
1185: * POSIX.1 says that O_NONBLOCK means return -1 with
1186: * errno=ENXIO when there is no process reading the FIFO.
1187: */
1188: ret = -ENXIO;
1189: if (!is_pipe && (filp->f_flags & O_NONBLOCK) && !pipe->readers)
1190: goto err;
1191:
1192: pipe->w_counter++;
1193: if (!pipe->writers++)
1194: wake_up_partner(pipe);
1195:
1196: if (!is_pipe && !pipe->readers) {
1197: if (wait_for_partner(pipe, &pipe->r_counter))
1198: goto err_wr;
1199: }
1200: break;
1201:
1202: case FMODE_READ | FMODE_WRITE:
1203: /*
1204: * O_RDWR
1205: * POSIX.1 leaves this case "undefined" when O_NONBLOCK is set.
1206: * This implementation will NEVER block on a O_RDWR open, since
1207: * the process can at least talk to itself.
1208: */
1209:
1210: pipe->readers++;
1211: pipe->writers++;
1212: pipe->r_counter++;
1213: pipe->w_counter++;
1214: if (pipe->readers == 1 || pipe->writers == 1)
1215: wake_up_partner(pipe);
1216: break;
1217:
1218: default:
1219: ret = -EINVAL;
1220: goto err;
1221: }
1222:
1223: /* Ok! */
1224: mutex_unlock(&pipe->mutex);
1225: return 0;
1226:
1227: err_rd:
1228: if (!--pipe->readers)
1229: wake_up_interruptible(&pipe->wr_wait);
1230: ret = -ERESTARTSYS;
1231: goto err;
1232:
1233: err_wr:
1234: if (!--pipe->writers)
1235: wake_up_interruptible_all(&pipe->rd_wait);
1236: ret = -ERESTARTSYS;
1237: goto err;
1238:
1239: err:
1240: mutex_unlock(&pipe->mutex);
1241:
1242: put_pipe_info(inode, pipe);
1243: return ret;
1244: }
SourceFile:fs/pipe.c]
|
| 15/2 |
2026/04/06 21:05 |
llm |
expert |
0m
Model:
Tokens:
input: 15841
output: 21
thoughts: 2263
Thoughts: **Analyzing the Pipe Locking Mechanism**
Okay, so I'm diving into this `pipe_lock` function, trying to understand why it includes that `pipe->files` check. My initial understanding is that `fifo_open` increments `pipe->files` while `put_pipe_info` decrements it, all under the protection of `inode->i_lock`. The `pipe_lock` function itself is pretty straightforward: it locks a mutex if `pipe->files` is non-zero; otherwise, it does nothing.
My first thought was, if there's a `struct file` referencing the pipe, `pipe->files` should be at least 1, so the check seems redundant. Then, I considered the possibility of a `pipe` existing without an associated file. But, the `pipe_inode_info` structure is freed when `pipe->files` drops to zero. That check looked suspicious.
I went digging and found the actual definition in `fs/pipe.c`, and confirmed it. If `pipe->files` is 0, the pipe is being freed. If we're calling `pipe_lock` on a freed pipe, that's a UAF (Use-After-Free) bug. So, the intent of the check must be to avoid this. But, it got me wondering if there was a case where that reference could be 0.
Looking at how `pipe_lock` is used, specifically within `pipe_wait_writable`, I realized a scenario: if `pipe->files` becomes zero *during* a `wait_event_interruptible`, the `pipe_lock` call might see 0 and skip the lock. But, if all files are closed, the thread in `pipe_wait_writable` must have closed its file too. However, the function `pipe_wait_writable` is called through `vmsplice_to_pipe`, which is a syscall on an open file. That should mean `pipe->files` can't be zero here!
Since it's always at least 1 in this context, the race is between the value being 2 and 1. In both cases, `mutex_lock` would be called, so the behavior is the same. I then considered if there were Torn Accesses, given that `pipe->files` is an `unsigned int` (4 bytes). On x86_64, 4-byte aligned reads/writes are atomic. Even if torn, any non-zero value would still trigger the lock. And KCSAN is reporting the address as aligned.
So, the read will see either 2 or 1. Both are non-zero. Thus, the lock will always be called. So why the warning? It must be a plain data race (read and write). Is this harmful? If the logic is "if it's non-zero, lock," and it's *always* non-zero in this specific context, it should be benign. So why the check? Why not just `mutex_lock(&pipe->mutex)`?
I went back and checked `pipe_lock` and `pipe_unlock` usage in `fs/pipe.c`. I looked at `free_pipe_info`, which is used to free the pipe info. It doesn't seem to explicitly destroy the mutex. So the check is probably to protect against UAF. But, reading `pipe->files` is already an access. If the `pipe` is freed, even reading `pipe->files` is UAF.
So `pipe_lock` must be called only when the `pipe` is still valid, and if we have a reference, `pipe->files` should be > 0. So what if the reference is from the `inode`? When `pipe->files` is 0, `inode->i_pipe` is set to NULL. So someone might look at `inode->i_pipe` *before* it becomes NULL. But they should hold `inode->i_lock`!
I looked for callers of `pipe_lock` outside of `fs/pipe.c`. I checked `pipe_read`, `pipe_write` which also require the file to be open, and `pipe_poll`. I see `pipe_lock` is used in `pipe_wait_writable` and `pipe_wait_readable`. Both are called from splice, vmsplice, read/write. All of these calls require a `struct file`.
So why the check? Maybe the disconnection from the inode? But `pipe_inode_info` isn't the inode! So it looks like it's guarding against a condition that isn't possible, but might have been added to guard a very specific, and probably flawed, understanding of how references to pipes could be lost.
|
| 16/2 |
2026/04/06 21:05 |
tool |
codesearch-find-references |
0m
Args:
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:ipipe_prep SourceFile:fs/splice.c SourceLine:1654] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:iter_file_splice_write SourceFile:fs/splice.c SourceLine:682] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:opipe_prep SourceFile:fs/splice.c SourceLine:1690] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:pipe_double_lock SourceFile:fs/pipe.c SourceLine:110] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:pipe_double_lock SourceFile:fs/pipe.c SourceLine:111] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:pipe_wait_readable SourceFile:fs/pipe.c SourceLine:1077] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:pipe_wait_writable SourceFile:fs/pipe.c SourceLine:1084] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:port_fops_splice_write SourceFile:drivers/char/virtio_console.c SourceLine:925] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:splice_file_to_pipe SourceFile:fs/splice.c SourceLine:1287] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:splice_from_pipe SourceFile:fs/splice.c SourceLine:640] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:splice_to_socket SourceFile:fs/splice.c SourceLine:805] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:vmsplice_to_pipe SourceFile:fs/splice.c SourceLine:1550] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:vmsplice_to_user SourceFile:fs/splice.c SourceLine:1518] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:wait_for_dump_helpers SourceFile:fs/coredump.c SourceLine:591] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:wait_for_dump_helpers SourceFile:fs/coredump.c SourceLine:604] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:wait_for_partner SourceFile:fs/pipe.c SourceLine:1109]] TruncatedOutput:false]
|
| 17/2 |
2026/04/06 21:05 |
llm |
expert |
0m
Model:
Tokens:
input: 19948
output: 38
thoughts: 4275
Thoughts: **Analyzing a Data Race in the Pipe Locking Mechanism**
Okay, so I'm digging into this KCSAN report related to a data race in the pipe code, specifically around `pipe_lock` and `pipe->files`. It seems like all the callers are getting a `struct file` or `struct pipe_inode_info` from a `struct file`, so that's where we need to start.
I see that `wait_for_partner` in `fs/pipe.c` calls `pipe_unlock` then waits. This function is used in `fifo_open`. In `fifo_open`, we have a `struct file` that's being opened and the counter `pipe->files` is incremented before `wait_for_partner` is called, so `pipe->files` should be at least 1.
The `pipe_lock` function checks `pipe->files`, and I understand that's because of interactions with `pipe_release` and `put_pipe_info`. If `pipe_release` is called, the file is being closed, and if it's the *last* file, `pipe->files` becomes 0. However, if we're *inside* a file operation, like read, write, or splice, the file *can't* be closed. Wait, what if another thread closes the file? The syscall holds a reference. `fput` won't be called until the last reference is gone.
Okay, in `pipe_release`, `put_pipe_info` decrements `pipe->files`. Then the code I found showed that `pipe_wait_writable` unlocks the mutex, waits, and re-locks. Another thread *could* call `pipe_release` on a *different* file descriptor for the same pipe while the other thread is waiting. It *can* decrement `pipe->files`. If it's at 2, it drops to 1, and the `pipe_lock` would see 1 and be fine. If it's at 1, it becomes 0, but it can only be 1 if the waiting thread is the *only* one with a reference.
Wait, what if the `inode` holds a reference? `inode->i_pipe` points to the pipe, but that's only set if `pipe->files > 0`. `pipe->files` is the number of `struct file` objects. Therefore, if `pipe_wait_writable` is running, there's at least one `struct file` open and so `pipe->files` is at least 1. `pipe_release` on a different file can only decrement it to 1 or more, not 0.
Wait, what if the thread calling `vmsplice` is also the one closing the file? No, that's impossible. What about a race between `fifo_open` and `pipe_release`? `fifo_open` increments `pipe->files` under `inode->i_lock`, while `put_pipe_info` decrements it, also under `inode->i_lock`. This keeps `inode->i_pipe` and `pipe->files` consistent.
So, the reported race is between `pipe_release` (write) and `pipe_wait_writable` (read). `pipe->files` changed from 2 to 1. One file closed, the `vmsplice` file is still open, meaning it's 1. `pipe_lock` reads `pipe->files`. If it's 2 or 1, it locks, which is correct. Can it read 0? Only if `pipe->files` was 1 and became 0. But I've argued that if `vmsplice` is running, `pipe->files` *must* be at least 1.
Wait, what if `vmsplice` is running on a file that was already closed? No, `fdget` would fail. What if the `pipe` pointer in `vmsplice_to_pipe` is stale? Nope, it's obtained from `get_pipe_info(file, true)`. So, it's a race between two valid operations via *different* files on the same pipe: `vmsplice` (waiting for space) and `close`.
The race on `pipe->files` is technically a data race because `pipe_lock` *doesn't* hold `inode->i_lock`. However, it's benign because, in the context of a file operation, `pipe->files` is guaranteed to be at least 1. The transition from 2 to 1 will always leave `pipe->files` > 0. Even if it goes from 1 to 0, it means
|
| 18/2 |
2026/04/06 21:06 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:fs/pipe.c IncludeLines:true Name:__pipe_lock]
Results: map[SourceCode: SourceFile:]
|
| 19/2 |
2026/04/06 21:06 |
llm |
expert |
0m
Model:
Tokens:
input: 25307
output: 18
thoughts: 31
Thoughts: Okay, here's my interpretation of those thoughts, tailored for an expert audience, and delivered as a first-person summary:
**Digging into Pipe Locking**
Right, I'm pretty sure I came across this specific issue within the `pipe_release` function, or something similar, earlier in my work. It's been bugging me in the back of my mind. The thing I want to understand is how the pipe is being locked. So, let me quickly jump over to `fs/pipe.c`. My hunch is that I'll find the details, the lock itself, defined within that file, specifically around something called `__pipe_lock`. Yep, that's where I need to start digging. Let's see what's happening there with those locks... this is likely where the relevant synchronization primitives and contention management reside.
|
| 20/2 |
2026/04/06 21:06 |
tool |
grepper |
0m
Error:
Args: map[Expression:__pipe_lock]
Results:
|
| 21/2 |
2026/04/06 21:06 |
llm |
expert |
0m
Model:
Tokens:
input: 25548
output: 20
thoughts: 3912
Thoughts: **Analysis of the Pipe Locking Mechanism**
Okay, so I've been digging into this pipe locking issue, specifically the interaction between `pipe_release`, `pipe_lock`, and `pipe_unlock`. I started by noticing that `pipe_release` directly uses `mutex_lock`, while `pipe_lock` has a check on `pipe->files` before locking the mutex. This immediately raised a question: why the difference?
Initially, I thought that `pipe_lock`'s check might be handling the destruction of a pipe, as it checks if `pipe->files` exists, suggesting that if there are no open file descriptors referring to this pipe, then `pipe->files` would be 0, and the lock is unnecessary. However, I quickly realized that if we have a file reference, the pipe shouldn't be in the process of being destroyed. Specifically, `pipe_wait_writable` uses the `pipe_lock` function, and a file reference must exist to call it.
I then revisited the code and considered the race condition between `pipe_release` and the other pipe functions. `pipe_release` locks, unlocks, and then decrements `pipe->files` via `put_pipe_info`. It is possible that `pipe->files` could become 0. I considered if the race could be between `pipe_release` (another file) and `pipe_lock`. My first impression was that this was a benign race because `pipe->files` should always remain greater than or equal to 1.
I found a possible scenario if something other than a file reference was using the pipe. I then turned to examine a related commit, `f367b096f90c`, which indicates that the check on `pipe->files` is there to handle the pipe being destroyed.
The more I thought about it, the more I was in agreement with the explanation in the commit. The issue is if the pipe can be destroyed while we are waiting for something. But I was still trying to determine how this was possible when a file is referencing it.
I then considered what it means to close the file or for it to be released. `sys_close` does not remove the file reference because of `fdget` in the syscall. I considered a kernel thread closing the reference, but rejected that.
I found `splice`, where `pipe_lock` is called from `splice_from_pipe`. That has a file reference, though. I then found that `wait_for_dump_helpers` also calls `pipe_lock` and has a `struct file`, too.
Then I realized that `pipe->files` must be a "belt and suspenders" approach. The check is there as a safeguard. That's why it is causing the race because it is redundant.
I tried to determine if the race was benign. I considered whether we could end up with `pipe->files` as 0. I then thought about what happens if `pipe->files` is 1 and goes to 0 during the lock. My conclusion was that this should not happen if we have a file reference.
I reviewed the `splice` code again, and saw the pipe reference comes from `fdget`. I reviewed `get_pipe_info` and it is from a file. Then, I reviewed all the callers of `pipe_lock`, and they all seem to have file references.
I then determined that I was correct in the assessment. `pipe->files` is always greater than or equal to 1. The race is between `pipe_release` (from 2 -> 1) and `pipe_lock` (reading `pipe->files`). The race is benign.
Then, I looked at whether `pipe->files` could be torn (changed in the middle of a read), but realized it is aligned and 4-bytes. So, it will not be torn.
I also considered whether a compiler might optimize the read away, but I realized the `mutex_lock` would act as a barrier. The only case I could see was inlining, which does occur. But the `mutex_lock` acts as a barrier there as well.
I looked at other uses of `pipe->files`, and realized it is mainly used with `pipe_lock`. There is a benign use with `put_pipe_info` too. `fifo_open` increments the counter, and `alloc_pipe_info` sets it to 1 initially.
I then tracked what happens to `pipe->files` when the pipe is created. It is set to 1 in `alloc_pipe_info`, and assigned to `inode->i_pipe`, but no file yet. Then, `fifo_open` happens, and the `struct file` is created. `pipe->files` is 1 then. If another `open` happens
|
| 22/2 |
2026/04/06 21:06 |
tool |
read-file |
0m
Args:
Results: map[Contents:// SPDX-License-Identifier: GPL-2.0
/*
* linux/fs/pipe.c
*
* Copyright (C) 1991, 1992, 1999 Linus Torvalds
*/
#include <linux/mm.h>
#include <linux/file.h>
#include <linux/poll.h>
#include <linux/slab.h>
#include <linux/module.h>
#include <linux/init.h>
#include <linux/fs.h>
#include <linux/log2.h>
#include <linux/mount.h>
#include <linux/pseudo_fs.h>
#include <linux/magic.h>
#include <linux/pipe_fs_i.h>
#include <linux/uio.h>
#include <linux/highmem.h>
#include <linux/pagemap.h>
#include <linux/audit.h>
#include <linux/syscalls.h>
#include <linux/fcntl.h>
#include <linux/memcontrol.h>
#include <linux/watch_queue.h>
#include <linux/sysctl.h>
#include <linux/sort.h>
#include <linux/uaccess.h>
#include <asm/ioctls.h>
#include "internal.h"
/*
* New pipe buffers will be restricted to this size while the user is exceeding
* their pipe buffer quota. The general pipe use case needs at least two
* buffers: one for data yet to be read, and one for new data. If this is less
* than two, then a write to a non-empty pipe may block even if the pipe is not
* full. This can occur with GNU make jobserver or similar uses of pipes as
* semaphores: multiple processes may be waiting to write tokens back to the
* pipe before reading tokens: https://lore.kernel.org/lkml/1628086770.5rn8p04n6j.none@localhost/.
*
* Users can reduce their pipe buffers with F_SETPIPE_SZ below this at their
* own risk, namely: pipe writes to non-full pipes may block until the pipe is
* emptied.
*/
#define PIPE_MIN_DEF_BUFFERS 2
/*
* The max size that a non-root user is allowed to grow the pipe. Can
* be set by root in /proc/sys/fs/pipe-max-size
*/
static unsigned int pipe_max_size = 1048576;
/* Maximum allocatable pages per user. Hard limit is unset by default, soft
* matches default values.
*/
static unsigned long pipe_user_pages_hard;
static unsigned long pipe_user_pages_soft = PIPE_DEF_BUFFERS * INR_OPEN_CUR;
/*
* We use head and tail indices that aren't masked off, except at the point of
* dereference, but rather they're allowed to wrap naturally. This means there
* isn't a dead spot in the buffer, but the ring has to be a power of two and
* <= 2^31.
* -- David Howells 2019-09-23.
*
* Reads with count = 0 should always return 0.
* -- Julian Bradfield 1999-06-07.
*
* FIFOs and Pipes now generate SIGIO for both readers and writers.
* -- Jeremy Elson <jelson@circlemud.org> 2001-08-16
*
* pipe_read & write cleanup
* -- Manfred Spraul <manfred@colorfullife.com> 2002-05-09
*/
#ifdef CONFIG_PROVE_LOCKING
static int pipe_lock_cmp_fn(const struct lockdep_map *a,
const struct lockdep_map *b)
{
return cmp_int((unsigned long) a, (unsigned long) b);
}
#endif
void pipe_lock(struct pipe_inode_info *pipe)
{
if (pipe->files)
mutex_lock(&pipe->mutex);
}
EXPORT_SYMBOL(pipe_lock);
void pipe_unlock(struct pipe_inode_info *pipe)
{
if (pipe->files)
mutex_unlock(&pipe->mutex);
}
EXPORT_SYMBOL(pipe_unlock);
void pipe_double_lock(struct pipe_inode_info *pipe1,
struct pipe_inode_info *pipe2)
{
BUG_ON(pipe1 == pipe2);
if (pipe1 > pipe2)
swap(pipe1, pipe2);
pipe_lock(pipe1);
pipe_lock(pipe2);
}
static struct page *anon_pipe_get_page(struct pipe_inode_info *pipe)
{
for (int i = 0; i < ARRAY_SIZE(pipe->tmp_page); i++) {
if (pipe->tmp_page[i]) {
struct page *page = pipe->tmp_page[i];
pipe->tmp_page[i] = NULL;
return page;
}
}
return alloc_page(GFP_HIGHUSER | __GFP_ACCOUNT);
}
static void anon_pipe_put_page(struct pipe_inode_info *pipe,
struct page *page)
{
if (page_count(page) == 1) {
for (int i = 0; i < ARRAY_SIZE(pipe->tmp_page); i++) {
if (!pipe->tmp_page[i]) {
pipe->tmp_page[i] = page;
return;
}
}
}
put_page(page);
}
static void anon_pipe_buf_release(struct pipe_inode_info *pipe,
struct pipe_buffer *buf)
{
struct page *page = buf->page;
anon_pipe_put_page(pipe, page);
}
static bool anon_pipe_buf_try_steal(struct pipe_inode_info *pipe,
struct pipe_buffer *buf)
{
struct page *page = buf->page;
if (page_count(page) != 1)
return false;
memcg_kmem_uncharge_page(page, 0);
__SetPageLocked(page);
return true;
}
/**
* generic_pipe_buf_try_steal - attempt to take ownership of a &pipe_buffer
* @pipe: the pipe that the buffer belongs to
* @buf: the buffer to attempt to steal
*
* Description:
* This function attempts to steal the &struct page attached to
* @buf. If successful, this function returns 0 and returns with
* the page locked. The caller may then reuse the page for whatever
* he wishes; the typical use is insertion into a different file
* page cache.
*/
bool generic_pipe_buf_try_steal(struct pipe_inode_info *pipe,
struct pipe_buffer *buf)
{
struct page *page = buf->page;
/*
* A reference of one is golden, that means that the owner of this
* page is the only one holding a reference to it. lock the page
* and return OK.
*/
if (page_count(page) == 1) {
lock_page(page);
return true;
}
return false;
}
EXPORT_SYMBOL(generic_pipe_buf_try_steal);
/**
* generic_pipe_buf_get - get a reference to a &struct pipe_buffer
* @pipe: the pipe that the buffer belongs to
* @buf: the buffer to get a reference to
*
* Description:
* This function grabs an extra reference to @buf. It's used in
* the tee() system call, when we duplicate the buffers in one
* pipe into another.
*/
bool generic_pipe_buf_get(struct pipe_inode_info *pipe, struct pipe_buffer *buf)
{
return try_get_page(buf->page);
}
EXPORT_SYMBOL(generic_pipe_buf_get);
/**
* generic_pipe_buf_release - put a reference to a &struct pipe_buffer
* @pipe: the pipe that the buffer belongs to
* @buf: the buffer to put a reference to
*
* Description:
* This function releases a reference to @buf.
*/
void generic_pipe_buf_release(struct pipe_inode_info *pipe,
struct pipe_buffer *buf)
{
put_page(buf->page);
}
EXPORT_SYMBOL(generic_pipe_buf_release);
static const struct pipe_buf_operations anon_pipe_buf_ops = {
.release = anon_pipe_buf_release,
.try_steal = anon_pipe_buf_try_steal,
.get = generic_pipe_buf_get,
};
/* Done while waiting without holding the pipe lock - thus the READ_ONCE() */
static inline bool pipe_readable(const struct pipe_inode_info *pipe)
{
union pipe_index idx = { .head_tail = READ_ONCE(pipe->head_tail) };
unsigned int writers = READ_ONCE(pipe->writers);
return !pipe_empty(idx.head, idx.tail) || !writers;
}
static inline unsigned int pipe_update_tail(struct pipe_inode_info *pipe,
struct pipe_buffer *buf,
unsigned int tail)
{
pipe_buf_release(pipe, buf);
/*
* If the pipe has a watch_queue, we need additional protection
* by the spinlock because notifications get posted with only
* this spinlock, no mutex
*/
if (pipe_has_watch_queue(pipe)) {
spin_lock_irq(&pipe->rd_wait.lock);
#ifdef CONFIG_WATCH_QUEUE
if (buf->flags & PIPE_BUF_FLAG_LOSS)
pipe->note_loss = true;
#endif
pipe->tail = ++tail;
spin_unlock_irq(&pipe->rd_wait.lock);
return tail;
}
/*
* Without a watch_queue, we can simply increment the tail
* without the spinlock - the mutex is enough.
*/
pipe->tail = ++tail;
return tail;
}
static ssize_t
anon_pipe_read(struct kiocb *iocb, struct iov_iter *to)
{
size_t total_len = iov_iter_count(to);
struct file *filp = iocb->ki_filp;
struct pipe_inode_info *pipe = filp->private_data;
bool wake_writer = false, wake_next_reader = false;
ssize_t ret;
/* Null read succeeds. */
if (unlikely(total_len == 0))
return 0;
ret = 0;
mutex_lock(&pipe->mutex);
/*
* We only wake up writers if the pipe was full when we started reading
* and it is no longer full after reading to avoid unnecessary wakeups.
*
* But when we do wake up writers, we do so using a sync wakeup
* (WF_SYNC), because we want them to get going and generate more
* data for us.
*/
for (;;) {
/* Read ->head with a barrier vs post_one_notification() */
unsigned int head = smp_load_acquire(&pipe->head);
unsigned int tail = pipe->tail;
#ifdef CONFIG_WATCH_QUEUE
if (pipe->note_loss) {
struct watch_notification n;
if (total_len < 8) {
if (ret == 0)
ret = -ENOBUFS;
break;
}
n.type = WATCH_TYPE_META;
n.subtype = WATCH_META_LOSS_NOTIFICATION;
n.info = watch_sizeof(n);
if (copy_to_iter(&n, sizeof(n), to) != sizeof(n)) {
if (ret == 0)
ret = -EFAULT;
break;
}
ret += sizeof(n);
total_len -= sizeof(n);
pipe->note_loss = false;
}
#endif
if (!pipe_empty(head, tail)) {
struct pipe_buffer *buf = pipe_buf(pipe, tail);
size_t chars = buf->len;
size_t written;
int error;
if (chars > total_len) {
if (buf->flags & PIPE_BUF_FLAG_WHOLE) {
if (ret == 0)
ret = -ENOBUFS;
break;
}
chars = total_len;
}
error = pipe_buf_confirm(pipe, buf);
if (error) {
if (!ret)
ret = error;
break;
}
written = copy_page_to_iter(buf->page, buf->offset, chars, to);
if (unlikely(written < chars)) {
if (!ret)
ret = -EFAULT;
break;
}
ret += chars;
buf->offset += chars;
buf->len -= chars;
/* Was it a packet buffer? Clean up and exit */
if (buf->flags & PIPE_BUF_FLAG_PACKET) {
total_len = chars;
buf->len = 0;
}
if (!buf->len) {
wake_writer |= pipe_full(head, tail, pipe->max_usage);
tail = pipe_update_tail(pipe, buf, tail);
}
total_len -= chars;
if (!total_len)
break; /* common path: read succeeded */
if (!pipe_empty(head, tail)) /* More to do? */
continue;
}
if (!pipe->writers)
break;
if (ret)
break;
if ((filp->f_flags & O_NONBLOCK) ||
(iocb->ki_flags & IOCB_NOWAIT)) {
ret = -EAGAIN;
break;
}
mutex_unlock(&pipe->mutex);
/*
* We only get here if we didn't actually read anything.
*
* But because we didn't read anything, at this point we can
* just return directly with -ERESTARTSYS if we're interrupted,
* since we've done any required wakeups and there's no need
* to mark anything accessed. And we've dropped the lock.
*/
if (wait_event_interruptible_exclusive(pipe->rd_wait, pipe_readable(pipe)) < 0)
return -ERESTARTSYS;
wake_next_reader = true;
mutex_lock(&pipe->mutex);
}
if (pipe_is_empty(pipe))
wake_next_reader = false;
mutex_unlock(&pipe->mutex);
if (wake_writer)
wake_up_interruptible_sync_poll(&pipe->wr_wait, EPOLLOUT | EPOLLWRNORM);
if (wake_next_reader)
wake_up_interruptible_sync_poll(&pipe->rd_wait, EPOLLIN | EPOLLRDNORM);
kill_fasync(&pipe->fasync_writers, SIGIO, POLL_OUT);
return ret;
}
static ssize_t
fifo_pipe_read(struct kiocb *iocb, struct iov_iter *to)
{
int ret = anon_pipe_read(iocb, to);
if (ret > 0)
file_accessed(iocb->ki_filp);
return ret;
}
static inline int is_packetized(struct file *file)
{
return (file->f_flags & O_DIRECT) != 0;
}
/* Done while waiting without holding the pipe lock - thus the READ_ONCE() */
static inline bool pipe_writable(const struct pipe_inode_info *pipe)
{
union pipe_index idx = { .head_tail = READ_ONCE(pipe->head_tail) };
unsigned int max_usage = READ_ONCE(pipe->max_usage);
return !pipe_full(idx.head, idx.tail, max_usage) ||
!READ_ONCE(pipe->readers);
}
static ssize_t
anon_pipe_write(struct kiocb *iocb, struct iov_iter *from)
{
struct file *filp = iocb->ki_filp;
struct pipe_inode_info *pipe = filp->private_data;
unsigned int head;
ssize_t ret = 0;
size_t total_len = iov_iter_count(from);
ssize_t chars;
bool was_empty = false;
bool wake_next_writer = false;
/*
* Reject writing to watch queue pipes before the point where we lock
* the pipe.
* Otherwise, lockdep would be unhappy if the caller already has another
* pipe locked.
* If we had to support locking a normal pipe and a notification pipe at
* the same time, we could set up lockdep annotations for that, but
* since we don't actually need that, it's simpler to just bail here.
*/
if (pipe_has_watch_queue(pipe))
return -EXDEV;
/* Null write succeeds. */
if (unlikely(total_len == 0))
return 0;
mutex_lock(&pipe->mutex);
if (!pipe->readers) {
if ((iocb->ki_flags & IOCB_NOSIGNAL) == 0)
send_sig(SIGPIPE, current, 0);
ret = -EPIPE;
goto out;
}
/*
* If it wasn't empty we try to merge new data into
* the last buffer.
*
* That naturally merges small writes, but it also
* page-aligns the rest of the writes for large writes
* spanning multiple pages.
*/
head = pipe->head;
was_empty = pipe_empty(head, pipe->tail);
chars = total_len & (PAGE_SIZE-1);
if (chars && !was_empty) {
struct pipe_buffer *buf = pipe_buf(pipe, head - 1);
int offset = buf->offset + buf->len;
if ((buf->flags & PIPE_BUF_FLAG_CAN_MERGE) &&
offset + chars <= PAGE_SIZE) {
ret = pipe_buf_confirm(pipe, buf);
if (ret)
goto out;
ret = copy_page_from_iter(buf->page, offset, chars, from);
if (unlikely(ret < chars)) {
ret = -EFAULT;
goto out;
}
buf->len += ret;
if (!iov_iter_count(from))
goto out;
}
}
for (;;) {
if (!pipe->readers) {
if ((iocb->ki_flags & IOCB_NOSIGNAL) == 0)
send_sig(SIGPIPE, current, 0);
if (!ret)
ret = -EPIPE;
break;
}
head = pipe->head;
if (!pipe_full(head, pipe->tail, pipe->max_usage)) {
struct pipe_buffer *buf;
struct page *page;
int copied;
page = anon_pipe_get_page(pipe);
if (unlikely(!page)) {
if (!ret)
ret = -ENOMEM;
break;
}
copied = copy_page_from_iter(page, 0, PAGE_SIZE, from);
if (unlikely(copied < PAGE_SIZE && iov_iter_count(from))) {
anon_pipe_put_page(pipe, page);
if (!ret)
ret = -EFAULT;
break;
}
pipe->head = head + 1;
/* Insert it into the buffer array */
buf = pipe_buf(pipe, head);
buf->page = page;
buf->ops = &anon_pipe_buf_ops;
buf->offset = 0;
if (is_packetized(filp))
buf->flags = PIPE_BUF_FLAG_PACKET;
else
buf->flags = PIPE_BUF_FLAG_CAN_MERGE;
buf->len = copied;
ret += copied;
if (!iov_iter_count(from))
break;
continue;
}
/* Wait for buffer space to become available. */
if ((filp->f_flags & O_NONBLOCK) ||
(iocb->ki_flags & IOCB_NOWAIT)) {
if (!ret)
ret = -EAGAIN;
break;
}
if (signal_pending(current)) {
if (!ret)
ret = -ERESTARTSYS;
break;
}
/*
* We're going to release the pipe lock and wait for more
* space. We wake up any readers if necessary, and then
* after waiting we need to re-check whether the pipe
* become empty while we dropped the lock.
*/
mutex_unlock(&pipe->mutex);
if (was_empty)
wake_up_interruptible_sync_poll(&pipe->rd_wait, EPOLLIN | EPOLLRDNORM);
kill_fasync(&pipe->fasync_readers, SIGIO, POLL_IN);
wait_event_interruptible_exclusive(pipe->wr_wait, pipe_writable(pipe));
mutex_lock(&pipe->mutex);
was_empty = pipe_is_empty(pipe);
wake_next_writer = true;
}
out:
if (pipe_is_full(pipe))
wake_next_writer = false;
mutex_unlock(&pipe->mutex);
/*
* If we do do a wakeup event, we do a 'sync' wakeup, because we
* want the reader to start processing things asap, rather than
* leave the data pending.
*
* This is particularly important for small writes, because of
* how (for example) the GNU make jobserver uses small writes to
* wake up pending jobs
*
* Epoll nonsensically wants a wakeup whether the pipe
* was already empty or not.
*/
if (was_empty || pipe->poll_usage)
wake_up_interruptible_sync_poll(&pipe->rd_wait, EPOLLIN | EPOLLRDNORM);
kill_fasync(&pipe->fasync_readers, SIGIO, POLL_IN);
if (wake_next_writer)
wake_up_interruptible_sync_poll(&pipe->wr_wait, EPOLLOUT | EPOLLWRNORM);
return ret;
}
static ssize_t
fifo_pipe_write(struct kiocb *iocb, struct iov_iter *from)
{
int ret = anon_pipe_write(iocb, from);
if (ret > 0) {
struct file *filp = iocb->ki_filp;
if (sb_start_write_trylock(file_inode(filp)->i_sb)) {
int err = file_update_time(filp);
if (err)
ret = err;
sb_end_write(file_inode(filp)->i_sb);
}
}
return ret;
}
static long pipe_ioctl(struct file *filp, unsigned int cmd, unsigned long arg)
{
struct pipe_inode_info *pipe = filp->private_data;
unsigned int count, head, tail;
switch (cmd) {
case FIONREAD:
mutex_lock(&pipe->mutex);
count = 0;
head = pipe->head;
tail = pipe->tail;
while (!pipe_empty(head, tail)) {
count += pipe_buf(pipe, tail)->len;
tail++;
}
mutex_unlock(&pipe->mutex);
return put_user(count, (int __user *)arg);
#ifdef CONFIG_WATCH_QUEUE
case IOC_WATCH_QUEUE_SET_SIZE: {
int ret;
mutex_lock(&pipe->mutex);
ret = watch_queue_set_size(pipe, arg);
mutex_unlock(&pipe->mutex);
return ret;
}
case IOC_WATCH_QUEUE_SET_FILTER:
return watch_queue_set_filter(
pipe, (struct watch_notification_filter __user *)arg);
#endif
default:
return -ENOIOCTLCMD;
}
}
/* No kernel lock held - fine */
static __poll_t
pipe_poll(struct file *filp, poll_table *wait)
{
__poll_t mask;
struct pipe_inode_info *pipe = filp->private_data;
union pipe_index idx;
/* Epoll has some historical nasty semantics, this enables them */
WRITE_ONCE(pipe->poll_usage, true);
/*
* Reading pipe state only -- no need for acquiring the semaphore.
*
* But because this is racy, the code has to add the
* entry to the poll table _first_ ..
*/
if (filp->f_mode & FMODE_READ)
poll_wait(filp, &pipe->rd_wait, wait);
if (filp->f_mode & FMODE_WRITE)
poll_wait(filp, &pipe->wr_wait, wait);
/*
* .. and only then can you do the racy tests. That way,
* if something changes and you got it wrong, the poll
* table entry will wake you up and fix it.
*/
idx.head_tail = READ_ONCE(pipe->head_tail);
mask = 0;
if (filp->f_mode & FMODE_READ) {
if (!pipe_empty(idx.head, idx.tail))
mask |= EPOLLIN | EPOLLRDNORM;
if (!pipe->writers && filp->f_pipe != pipe->w_counter)
mask |= EPOLLHUP;
}
if (filp->f_mode & FMODE_WRITE) {
if (!pipe_full(idx.head, idx.tail, pipe->max_usage))
mask |= EPOLLOUT | EPOLLWRNORM;
/*
* Most Unices do not set EPOLLERR for FIFOs but on Linux they
* behave exactly like pipes for poll().
*/
if (!pipe->readers)
mask |= EPOLLERR;
}
return mask;
}
static void put_pipe_info(struct inode *inode, struct pipe_inode_info *pipe)
{
int kill = 0;
spin_lock(&inode->i_lock);
if (!--pipe->files) {
inode->i_pipe = NULL;
kill = 1;
}
spin_unlock(&inode->i_lock);
if (kill)
free_pipe_info(pipe);
}
static int
pipe_release(struct inode *inode, struct file *file)
{
struct pipe_inode_info *pipe = file->private_data;
mutex_lock(&pipe->mutex);
if (file->f_mode & FMODE_READ)
pipe->readers--;
if (file->f_mode & FMODE_WRITE)
pipe->writers--;
/* Was that the last reader or writer, but not the other side? */
if (!pipe->readers != !pipe->writers) {
wake_up_interruptible_all(&pipe->rd_wait);
wake_up_interruptible_all(&pipe->wr_wait);
kill_fasync(&pipe->fasync_readers, SIGIO, POLL_IN);
kill_fasync(&pipe->fasync_writers, SIGIO, POLL_OUT);
}
mutex_unlock(&pipe->mutex);
put_pipe_info(inode, pipe);
return 0;
}
static int
pipe_fasync(int fd, struct file *filp, int on)
{
struct pipe_inode_info *pipe = filp->private_data;
int retval = 0;
mutex_lock(&pipe->mutex);
if (filp->f_mode & FMODE_READ)
retval = fasync_helper(fd, filp, on, &pipe->fasync_readers);
if ((filp->f_mode & FMODE_WRITE) && retval >= 0) {
retval = fasync_helper(fd, filp, on, &pipe->fasync_writers);
if (retval < 0 && (filp->f_mode & FMODE_READ))
/* this can happen only if on == T */
fasync_helper(-1, filp, 0, &pipe->fasync_readers);
}
mutex_unlock(&pipe->mutex);
return retval;
}
unsigned long account_pipe_buffers(struct user_struct *user,
unsigned long old, unsigned long new)
{
return atomic_long_add_return(new - old, &user->pipe_bufs);
}
bool too_many_pipe_buffers_soft(unsigned long user_bufs)
{
unsigned long soft_limit = READ_ONCE(pipe_user_pages_soft);
return soft_limit && user_bufs > soft_limit;
}
bool too_many_pipe_buffers_hard(unsigned long user_bufs)
{
unsigned long hard_limit = READ_ONCE(pipe_user_pages_hard);
return hard_limit && user_bufs > hard_limit;
}
bool pipe_is_unprivileged_user(void)
{
return !capable(CAP_SYS_RESOURCE) && !capable(CAP_SYS_ADMIN);
}
struct pipe_inode_info *alloc_pipe_info(void)
{
struct pipe_inode_info *pipe;
unsigned long pipe_bufs = PIPE_DEF_BUFFERS;
struct user_struct *user = get_current_user();
unsigned long user_bufs;
unsigned int max_size = READ_ONCE(pipe_max_size);
pipe = kzalloc_obj(struct pipe_inode_info, GFP_KERNEL_ACCOUNT);
if (pipe == NULL)
goto out_free_uid;
if (pipe_bufs * PAGE_SIZE > max_size && !capable(CAP_SYS_RESOURCE))
pipe_bufs = max_size >> PAGE_SHIFT;
user_bufs = account_pipe_buffers(user, 0, pipe_bufs);
if (too_many_pipe_buffers_soft(user_bufs) && pipe_is_unprivileged_user()) {
user_bufs = account_pipe_buffers(user, pipe_bufs, PIPE_MIN_DEF_BUFFERS);
pipe_bufs = PIPE_MIN_DEF_BUFFERS;
}
if (too_many_pipe_buffers_hard(user_bufs) && pipe_is_unprivileged_user())
goto out_revert_acct;
pipe->bufs = kzalloc_objs(struct pipe_buffer, pipe_bufs,
GFP_KERNEL_ACCOUNT);
if (pipe->bufs) {
init_waitqueue_head(&pipe->rd_wait);
init_waitqueue_head(&pipe->wr_wait);
pipe->r_counter = pipe->w_counter = 1;
pipe->max_usage = pipe_bufs;
pipe->ring_size = pipe_bufs;
pipe->nr_accounted = pipe_bufs;
pipe->user = user;
mutex_init(&pipe->mutex);
lock_set_cmp_fn(&pipe->mutex, pipe_lock_cmp_fn, NULL);
return pipe;
}
out_revert_acct:
(void) account_pipe_buffers(user, pipe_bufs, 0);
kfree(pipe);
out_free_uid:
free_uid(user);
return NULL;
}
void free_pipe_info(struct pipe_inode_info *pipe)
{
unsigned int i;
#ifdef CONFIG_WATCH_QUEUE
if (pipe->watch_queue)
watch_queue_clear(pipe->watch_queue);
#endif
(void) account_pipe_buffers(pipe->user, pipe->nr_accounted, 0);
free_uid(pipe->user);
for (i = 0; i < pipe->ring_size; i++) {
struct pipe_buffer *buf = pipe->bufs + i;
if (buf->ops)
pipe_buf_release(pipe, buf);
}
#ifdef CONFIG_WATCH_QUEUE
if (pipe->watch_queue)
put_watch_queue(pipe->watch_queue);
#endif
for (i = 0; i < ARRAY_SIZE(pipe->tmp_page); i++) {
if (pipe->tmp_page[i])
__free_page(pipe->tmp_page[i]);
}
kfree(pipe->bufs);
kfree(pipe);
}
static struct vfsmount *pipe_mnt __ro_after_init;
/*
* pipefs_dname() is called from d_path().
*/
static char *pipefs_dname(struct dentry *dentry, char *buffer, int buflen)
{
return dynamic_dname(buffer, buflen, "pipe:[%lu]",
d_inode(dentry)->i_ino);
}
static const struct dentry_operations pipefs_dentry_operations = {
.d_dname = pipefs_dname,
};
static const struct file_operations pipeanon_fops;
static struct inode * get_pipe_inode(void)
{
struct inode *inode = new_inode_pseudo(pipe_mnt->mnt_sb);
struct pipe_inode_info *pipe;
if (!inode)
goto fail_inode;
inode->i_ino = get_next_ino();
pipe = alloc_pipe_info();
if (!pipe)
goto fail_iput;
inode->i_pipe = pipe;
pipe->files = 2;
pipe->readers = pipe->writers = 1;
inode->i_fop = &pipeanon_fops;
/*
* Mark the inode dirty from the very beginning,
* that way it will never be moved to the dirty
* list because "mark_inode_dirty()" will think
* that it already _is_ on the dirty list.
*/
inode_state_assign_raw(inode, I_DIRTY);
inode->i_mode = S_IFIFO | S_IRUSR | S_IWUSR;
inode->i_uid = current_fsuid();
inode->i_gid = current_fsgid();
simple_inode_init_ts(inode);
return inode;
fail_iput:
iput(inode);
fail_inode:
return NULL;
}
int create_pipe_files(struct file **res, int flags)
{
struct inode *inode = get_pipe_inode();
struct file *f;
int error;
if (!inode)
return -ENFILE;
if (flags & O_NOTIFICATION_PIPE) {
error = watch_queue_init(inode->i_pipe);
if (error) {
free_pipe_info(inode->i_pipe);
iput(inode);
return error;
}
}
f = alloc_file_pseudo(inode, pipe_mnt, "",
O_WRONLY | (flags & (O_NONBLOCK | O_DIRECT)),
&pipeanon_fops);
if (IS_ERR(f)) {
free_pipe_info(inode->i_pipe);
iput(inode);
return PTR_ERR(f);
}
f->private_data = inode->i_pipe;
f->f_pipe = 0;
res[0] = alloc_file_clone(f, O_RDONLY | (flags & O_NONBLOCK),
&pipeanon_fops);
if (IS_ERR(res[0])) {
put_pipe_info(inode, inode->i_pipe);
fput(f);
return PTR_ERR(res[0]);
}
res[0]->private_data = inode->i_pipe;
res[0]->f_pipe = 0;
res[1] = f;
stream_open(inode, res[0]);
stream_open(inode, res[1]);
/* pipe groks IOCB_NOWAIT */
res[0]->f_mode |= FMODE_NOWAIT;
res[1]->f_mode |= FMODE_NOWAIT;
/*
* Disable permission and pre-content events, but enable legacy
* inotify events for legacy users.
*/
file_set_fsnotify_mode(res[0], FMODE_NONOTIFY_PERM);
file_set_fsnotify_mode(res[1], FMODE_NONOTIFY_PERM);
return 0;
}
static int __do_pipe_flags(int *fd, struct file **files, int flags)
{
int error;
int fdw, fdr;
if (flags & ~(O_CLOEXEC | O_NONBLOCK | O_DIRECT | O_NOTIFICATION_PIPE))
return -EINVAL;
error = create_pipe_files(files, flags);
if (error)
return error;
error = get_unused_fd_flags(flags);
if (error < 0)
goto err_read_pipe;
fdr = error;
error = get_unused_fd_flags(flags);
if (error < 0)
goto err_fdr;
fdw = error;
audit_fd_pair(fdr, fdw);
fd[0] = fdr;
fd[1] = fdw;
return 0;
err_fdr:
put_unused_fd(fdr);
err_read_pipe:
fput(files[0]);
fput(files[1]);
return error;
}
int do_pipe_flags(int *fd, int flags)
{
struct file *files[2];
int error = __do_pipe_flags(fd, files, flags);
if (!error) {
fd_install(fd[0], files[0]);
fd_install(fd[1], files[1]);
}
return error;
}
/*
* sys_pipe() is the normal C calling standard for creating
* a pipe. It's not the way Unix traditionally does this, though.
*/
static int do_pipe2(int __user *fildes, int flags)
{
struct file *files[2];
int fd[2];
int error;
error = __do_pipe_flags(fd, files, flags);
if (!error) {
if (unlikely(copy_to_user(fildes, fd, sizeof(fd)))) {
fput(files[0]);
fput(files[1]);
put_unused_fd(fd[0]);
put_unused_fd(fd[1]);
error = -EFAULT;
} else {
fd_install(fd[0], files[0]);
fd_install(fd[1], files[1]);
}
}
return error;
}
SYSCALL_DEFINE2(pipe2, int __user *, fildes, int, flags)
{
return do_pipe2(fildes, flags);
}
SYSCALL_DEFINE1(pipe, int __user *, fildes)
{
return do_pipe2(fildes, 0);
}
/*
* This is the stupid "wait for pipe to be readable or writable"
* model.
*
* See pipe_read/write() for the proper kind of exclusive wait,
* but that requires that we wake up any other readers/writers
* if we then do not end up reading everything (ie the whole
* "wake_next_reader/writer" logic in pipe_read/write()).
*/
void pipe_wait_readable(struct pipe_inode_info *pipe)
{
pipe_unlock(pipe);
wait_event_interruptible(pipe->rd_wait, pipe_readable(pipe));
pipe_lock(pipe);
}
void pipe_wait_writable(struct pipe_inode_info *pipe)
{
pipe_unlock(pipe);
wait_event_interruptible(pipe->wr_wait, pipe_writable(pipe));
pipe_lock(pipe);
}
/*
* This depends on both the wait (here) and the wakeup (wake_up_partner)
* holding the pipe lock, so "*cnt" is stable and we know a wakeup cannot
* race with the count check and waitqueue prep.
*
* Normally in order to avoid races, you'd do the prepare_to_wait() first,
* then check the condition you're waiting for, and only then sleep. But
* because of the pipe lock, we can check the condition before being on
* the wait queue.
*
* We use the 'rd_wait' waitqueue for pipe partner waiting.
*/
static int wait_for_partner(struct pipe_inode_info *pipe, unsigned int *cnt)
{
DEFINE_WAIT(rdwait);
int cur = *cnt;
while (cur == *cnt) {
prepare_to_wait(&pipe->rd_wait, &rdwait, TASK_INTERRUPTIBLE);
pipe_unlock(pipe);
schedule();
finish_wait(&pipe->rd_wait, &rdwait);
pipe_lock(pipe);
if (signal_pending(current))
break;
}
return cur == *cnt ? -ERESTARTSYS : 0;
}
static void wake_up_partner(struct pipe_inode_info *pipe)
{
wake_up_interruptible_all(&pipe->rd_wait);
}
static int fifo_open(struct inode *inode, struct file *filp)
{
bool is_pipe = inode->i_fop == &pipeanon_fops;
struct pipe_inode_info *pipe;
int ret;
filp->f_pipe = 0;
spin_lock(&inode->i_lock);
if (inode->i_pipe) {
pipe = inode->i_pipe;
pipe->files++;
spin_unlock(&inode->i_lock);
} else {
spin_unlock(&inode->i_lock);
pipe = alloc_pipe_info();
if (!pipe)
return -ENOMEM;
pipe->files = 1;
spin_lock(&inode->i_lock);
if (unlikely(inode->i_pipe)) {
inode->i_pipe->files++;
spin_unlock(&inode->i_lock);
free_pipe_info(pipe);
pipe = inode->i_pipe;
} else {
inode->i_pipe = pipe;
spin_unlock(&inode->i_lock);
}
}
filp->private_data = pipe;
/* OK, we have a pipe and it's pinned down */
mutex_lock(&pipe->mutex);
/* We can only do regular read/write on fifos */
stream_open(inode, filp);
switch (filp->f_mode & (FMODE_READ | FMODE_WRITE)) {
case FMODE_READ:
/*
* O_RDONLY
* POSIX.1 says that O_NONBLOCK means return with the FIFO
* opened, even when there is no process writing the FIFO.
*/
pipe->r_counter++;
if (pipe->readers++ == 0)
wake_up_partner(pipe);
if (!is_pipe && !pipe->writers) {
if ((filp->f_flags & O_NONBLOCK)) {
/* suppress EPOLLHUP until we have
* seen a writer */
filp->f_pipe = pipe->w_counter;
} else {
if (wait_for_partner(pipe, &pipe->w_counter))
goto err_rd;
}
}
break;
case FMODE_WRITE:
/*
* O_WRONLY
* POSIX.1 says that O_NONBLOCK means return -1 with
* errno=ENXIO when there is no process reading the FIFO.
*/
ret = -ENXIO;
if (!is_pipe && (filp->f_flags & O_NONBLOCK) && !pipe->readers)
goto err;
pipe->w_counter++;
if (!pipe->writers++)
wake_up_partner(pipe);
if (!is_pipe && !pipe->readers) {
if (wait_for_partner(pipe, &pipe->r_counter))
goto err_wr;
}
break;
case FMODE_READ | FMODE_WRITE:
/*
* O_RDWR
* POSIX.1 leaves this case "undefined" when O_NONBLOCK is set.
* This implementation will NEVER block on a O_RDWR open, since
* the process can at least talk to itself.
*/
pipe->readers++;
pipe->writers++;
pipe->r_counter++;
pipe->w_counter++;
if (pipe->readers == 1 || pipe->writers == 1)
wake_up_partner(pipe);
break;
default:
ret = -EINVAL;
goto err;
}
/* Ok! */
mutex_unlock(&pipe->mutex);
return 0;
err_rd:
if (!--pipe->readers)
wake_up_interruptible(&pipe->wr_wait);
ret = -ERESTARTSYS;
goto err;
err_wr:
if (!--pipe->writers)
wake_up_interruptible_all(&pipe->rd_wait);
ret = -ERESTARTSYS;
goto err;
err:
mutex_unlock(&pipe->mutex);
put_pipe_info(inode, pipe);
return ret;
}
const struct file_operations pipefifo_fops = {
.open = fifo_open,
.read_iter = fifo_pipe_read,
.write_iter = fifo_pipe_write,
.poll = pipe_poll,
.unlocked_ioctl = pipe_ioctl,
.release = pipe_release,
.fasync = pipe_fasync,
.splice_write = iter_file_splice_write,
};
static const struct file_operations pipeanon_fops = {
.open = fifo_open,
.read_iter = anon_pipe_read,
.write_iter = anon_pipe_write,
.poll = pipe_poll,
.unlocked_ioctl = pipe_ioctl,
.release = pipe_release,
.fasync = pipe_fasync,
.splice_write = iter_file_splice_write,
};
/*
* Currently we rely on the pipe array holding a power-of-2 number
* of pages. Returns 0 on error.
*/
unsigned int round_pipe_size(unsigned int size)
{
if (size > (1U << 31))
return 0;
/* Minimum pipe size, as required by POSIX */
if (size < PAGE_SIZE)
return PAGE_SIZE;
return roundup_pow_of_two(size);
}
/*
* Resize the pipe ring to a number of slots.
*
* Note the pipe can be reduced in capacity, but only if the current
* occupancy doesn't exceed nr_slots; if it does, EBUSY will be
* returned instead.
*/
int pipe_resize_ring(struct pipe_inode_info *pipe, unsigned int nr_slots)
{
struct pipe_buffer *bufs;
unsigned int head, tail, mask, n;
/* nr_slots larger than limits of pipe->{head,tail} */
if (unlikely(nr_slots > (pipe_index_t)-1u))
return -EINVAL;
bufs = kzalloc_objs(*bufs, nr_slots, GFP_KERNEL_ACCOUNT | __GFP_NOWARN);
if (unlikely(!bufs))
return -ENOMEM;
spin_lock_irq(&pipe->rd_wait.lock);
mask = pipe->ring_size - 1;
head = pipe->head;
tail = pipe->tail;
n = pipe_occupancy(head, tail);
if (nr_slots < n) {
spin_unlock_irq(&pipe->rd_wait.lock);
kfree(bufs);
return -EBUSY;
}
/*
* The pipe array wraps around, so just start the new one at zero
* and adjust the indices.
*/
if (n > 0) {
unsigned int h = head & mask;
unsigned int t = tail & mask;
if (h > t) {
memcpy(bufs, pipe->bufs + t,
n * sizeof(struct pipe_buffer));
} else {
unsigned int tsize = pipe->ring_size - t;
if (h > 0)
memcpy(bufs + tsize, pipe->bufs,
h * sizeof(struct pipe_buffer));
memcpy(bufs, pipe->bufs + t,
tsize * sizeof(struct pipe_buffer));
}
}
head = n;
tail = 0;
kfree(pipe->bufs);
pipe->bufs = bufs;
pipe->ring_size = nr_slots;
if (pipe->max_usage > nr_slots)
pipe->max_usage = nr_slots;
pipe->tail = tail;
pipe->head = head;
if (!pipe_has_watch_queue(pipe)) {
pipe->max_usage = nr_slots;
pipe->nr_accounted = nr_slots;
}
spin_unlock_irq(&pipe->rd_wait.lock);
/* This might have made more room for writers */
wake_up_interruptible(&pipe->wr_wait);
return 0;
}
/*
* Allocate a new array of pipe buffers and copy the info over. Returns the
* pipe size if successful, or return -ERROR on error.
*/
static long pipe_set_size(struct pipe_inode_info *pipe, unsigned int arg)
{
unsigned long user_bufs;
unsigned int nr_slots, size;
long ret = 0;
if (pipe_has_watch_queue(pipe))
return -EBUSY;
size = round_pipe_size(arg);
nr_slots = size >> PAGE_SHIFT;
if (!nr_slots)
return -EINVAL;
/*
* If trying to increase the pipe capacity, check that an
* unprivileged user is not trying to exceed various limits
* (soft limit check here, hard limit check just below).
* Decreasing the pipe capacity is always permitted, even
* if the user is currently over a limit.
*/
if (nr_slots > pipe->max_usage &&
size > pipe_max_size && !capable(CAP_SYS_RESOURCE))
return -EPERM;
user_bufs = account_pipe_buffers(pipe->user, pipe->nr_accounted, nr_slots);
if (nr_slots > pipe->max_usage &&
(too_many_pipe_buffers_hard(user_bufs) ||
too_many_pipe_buffers_soft(user_bufs)) &&
pipe_is_unprivileged_user()) {
ret = -EPERM;
goto out_revert_acct;
}
ret = pipe_resize_ring(pipe, nr_slots);
if (ret < 0)
goto out_revert_acct;
return pipe->max_usage * PAGE_SIZE;
out_revert_acct:
(void) account_pipe_buffers(pipe->user, nr_slots, pipe->nr_accounted);
return ret;
}
/*
* Note that i_pipe and i_cdev share the same location, so checking ->i_pipe is
* not enough to verify that this is a pipe.
*/
struct pipe_inode_info *get_pipe_info(struct file *file, bool for_splice)
{
struct pipe_inode_info *pipe = file->private_data;
if (!pipe)
return NULL;
if (file->f_op != &pipefifo_fops && file->f_op != &pipeanon_fops)
return NULL;
if (for_splice && pipe_has_watch_queue(pipe))
return NULL;
return pipe;
}
long pipe_fcntl(struct file *file, unsigned int cmd, unsigned int arg)
{
struct pipe_inode_info *pipe;
long ret;
pipe = get_pipe_info(file, false);
if (!pipe)
return -EBADF;
mutex_lock(&pipe->mutex);
switch (cmd) {
case F_SETPIPE_SZ:
ret = pipe_set_size(pipe, arg);
break;
case F_GETPIPE_SZ:
ret = pipe->max_usage * PAGE_SIZE;
break;
default:
ret = -EINVAL;
break;
}
mutex_unlock(&pipe->mutex);
return ret;
}
static const struct super_operations pipefs_ops = {
.destroy_inode = free_inode_nonrcu,
.statfs = simple_statfs,
};
/*
* pipefs should _never_ be mounted by userland - too much of security hassle,
* no real gain from having the whole file system mounted. So we don't need
* any operations on the root directory. However, we need a non-trivial
* d_name - pipe: will go nicely and kill the special-casing in procfs.
*/
static int pipefs_init_fs_context(struct fs_context *fc)
{
struct pseudo_fs_context *ctx = init_pseudo(fc, PIPEFS_MAGIC);
if (!ctx)
return -ENOMEM;
ctx->ops = &pipefs_ops;
ctx->dops = &pipefs_dentry_operations;
return 0;
}
static struct file_system_type pipe_fs_type = {
.name = "pipefs",
.init_fs_context = pipefs_init_fs_context,
.kill_sb = kill_anon_super,
};
#ifdef CONFIG_SYSCTL
static ulong round_pipe_size_ul(ulong size)
{
return round_pipe_size(size);
}
static int u2k_pipe_maxsz(const ulong *u_ptr, uint *k_ptr)
{
return proc_uint_u2k_conv_uop(u_ptr, k_ptr, round_pipe_size_ul);
}
static int do_proc_uint_conv_pipe_maxsz(ulong *u_ptr, uint *k_ptr,
int dir, const struct ctl_table *table)
{
return proc_uint_conv(u_ptr, k_ptr, dir, table, true,
u2k_pipe_maxsz,
proc_uint_k2u_conv);
}
static int proc_dopipe_max_size(const struct ctl_table *table, int write,
void *buffer, size_t *lenp, loff_t *ppos)
{
return proc_douintvec_conv(table, write, buffer, lenp, ppos,
do_proc_uint_conv_pipe_maxsz);
}
static const struct ctl_table fs_pipe_sysctls[] = {
{
.procname = "pipe-max-size",
.data = &pipe_max_size,
.maxlen = sizeof(pipe_max_size),
.mode = 0644,
.proc_handler = proc_dopipe_max_size,
.extra1 = SYSCTL_ONE,
},
{
.procname = "pipe-user-pages-hard",
.data = &pipe_user_pages_hard,
.maxlen = sizeof(pipe_user_pages_hard),
.mode = 0644,
.proc_handler = proc_doulongvec_minmax,
},
{
.procname = "pipe-user-pages-soft",
.data = &pipe_user_pages_soft,
.maxlen = sizeof(pipe_user_pages_soft),
.mode = 0644,
.proc_handler = proc_doulongvec_minmax,
},
};
#endif
static int __init init_pipe_fs(void)
{
int err = register_filesystem(&pipe_fs_type);
if (!err) {
pipe_mnt = kern_mount(&pipe_fs_type);
if (IS_ERR(pipe_mnt)) {
err = PTR_ERR(pipe_mnt);
unregister_filesystem(&pipe_fs_type);
}
}
#ifdef CONFIG_SYSCTL
register_sysctl_init("fs", fs_pipe_sysctls);
#endif
return err;
}
fs_initcall(init_pipe_fs);
]
|
| 23/2 |
2026/04/06 21:06 |
llm |
expert |
0m
Model:
Tokens:
input: 44505
output: 774
thoughts: 1914
Thoughts: **Benign Race Condition in Pipe Locking**
Okay, I've dug into this code, specifically the `pipe_lock` and `pipe_unlock` functions, and the reported data race. The core issue revolves around the `pipe->files` counter and its interaction with the mutex. Let's break down my thought process:
First, I confirmed the implementation: the functions indeed check `pipe->files` before acquiring/releasing the mutex, which seems like a basic safety check. Also, `put_pipe_info` decrements this counter, while `fifo_open` increments it.
The race boils down to a read (in `pipe_lock`) and a write (in `put_pipe_info`) on `pipe->files`. KCSAN flagged a transition from 2 to 1. Given that we are in the context of file operations, `pipe->files` is guaranteed to be at least 1. This means the mutex is still taken in *both* cases of the `if` check, whether it's 2 or 1.
So far, it looks harmless, the mutex protects the critical section. I then checked if there are other usages of `pipe->files`, and there aren't.
Now, I considered where `pipe->files` is initialized: `get_pipe_inode` sets it to 2 because an anonymous pipe has two file descriptors (read and write). For FIFOs, it starts at 1. The counter tracks the number of open file descriptors associated with the pipe.
My gut feeling is that this race is benign.
I started to think about compiler optimizations. If the compiler saw `if (pipe->files)`, it *might* assume stability if there were no writes. However, `mutex_lock` acts as a barrier, preventing any such assumptions. Moreover, `pipe_lock` itself is a function (even if inlined), and the `mutex_lock` call inside it *is* a barrier.
So I am thinking it's safe.
I had the thought that the warning is still not good, and so using `READ_ONCE` or `data_race()` would eliminate the warning. Since these are exported symbols, they are used by other modules.
Let's look at that check again... why does `pipe_lock` check `pipe->files` at all? I concluded that it's probably intended to prevent locking a pipe that's being destroyed. I re-evaluated that and decided the check does not provide any real protection against use-after-free.
I found a discussion on this point - the check was put in to handle a race where a pipe is closed while someone is waiting on it. But `wait_event_interruptible` should wake up, and is woken up, when the pipe is closed. `pipe_release` wakes up `rd_wait` and `wr_wait`.
If woken up because the pipe is closed, the thread may find `pipe->files == 0`. *But*, if it is woken up because the pipe is closed, it means that `pipe_release` was called. But `pipe_release` is only called when the last reference is gone. So the thread calling `pipe_wait_writable` must still have a reference.
This is robust even during process exit. Even when an application is killed, system calls finish, and file references are held until return from the entry code.
After this extensive review, I'm confident that the race is, in fact, benign.
Let's confirm: the race occurs on `pipe->files`. When `pipe_lock` is called from the context holding a file reference, the value is always > 0. The report shows 2 -> 1, both are > 0, and the outcome of the `if` is the same. Therefore, the race is harmless.
Now, what about `pipe_unlock`, could it see `pipe->files` as 0? No, if it was 1 in `pipe_lock` and became 0 before `pipe_unlock`, then `pipe_release` must have taken the mutex. Therefore `pipe_release` cannot finish, and `pipe->files` is stable and > 0. So the only race is *before* the mutex is taken in `pipe_lock`, but the value is >=1 and therefore benign.
I checked all callers and they all hold a file reference. An exception: `fifo_open` calls `put_pipe_info` on error, but does not call `pipe_lock`. Wait, `fifo_open` *does* call `pipe_lock` *after* incrementing `pipe->files` in `wait_for_partner`, so at least 1.
The evidence is all pointing to this race being benign.
As a fix, I would recommend
The KCSAN report identifies a data race on `pipe->files` between `pipe_release()` (via `put_pipe_info()`) and `pipe_wait_writable()` (via `pipe_lock()`).
### Analysis
1. **Variable:** `pipe->files` in `struct pipe_inode_info`. This field tracks the number of `struct file` objects currently referring to the pipe. It is documented as being protected by `inode->i_lock`.
2. **Writer (Thread A):** `pipe_release()` is called when a file descriptor for the pipe is closed. It calls `put_pipe_info()`, which decrements `pipe->files` while holding `inode->i_lock`.
3. **Reader (Thread B):** `pipe_wait_writable()` is called during a `vmsplice` operation. It calls `pipe_lock()`, which reads `pipe->files` without holding `inode->i_lock` to decide whether to acquire the pipe's mutex.
4. **Context:** In the reported race, `pipe->files` changed from `2` to `1`. This means there were at least two files referring to the pipe. One was being closed (Thread A), while another was being used for a syscall (Thread B).
5. **Safety:**
* As long as a thread is performing a syscall on a pipe (like `vmsplice`), it holds a reference to the `struct file` (via `fdget()` or similar). This ensures that `pipe->files` is at least 1.
* Even if other threads close their file descriptors for the same pipe, they can only decrement `pipe->files` down to 1 (the reference held by the active syscall).
* In `pipe_lock()`, the check `if (pipe->files)` will therefore always see a non-zero value (either 2 or 1 in this case).
* Since both 2 and 1 are non-zero, `pipe_lock()` will proceed to call `mutex_lock(&pipe->mutex)`, which is the correct behavior.
* The same logic applies to `pipe_unlock()`. Once the lock is held, `pipe->files` cannot drop to 0 because `pipe_release()` would need to acquire the same mutex to decrement the reader/writer counts before calling `put_pipe_info()`.
### Conclusion
The data race is **BENIGN**. The check `if (pipe->files)` in `pipe_lock()` and `pipe_unlock()` is a "safety" check that is structurally tolerant to concurrent updates. In all valid calling contexts (where the caller holds a file reference), the value of `pipe->files` is guaranteed to be non-zero, making the outcome of the race irrelevant to the logic.
The race exists because `pipe_lock()` accesses `pipe->files` without the `i_lock` that technically protects it, but this is a known pattern in the pipe implementation to avoid excessive locking overhead.
### Recommended Annotations
To silence the KCSAN warning, the accesses to `pipe->files` in `pipe_lock()` and `pipe_unlock()` should be wrapped in `READ_ONCE()` or `data_race()`.
```c
void pipe_lock(struct pipe_inode_info *pipe)
{
if (data_race(pipe->files))
mutex_lock(&pipe->mutex);
}
```
|
| 24/2 |
2026/04/06 21:07 |
tool |
set-results |
0m
Args: map[Benign:true Confident:true]
Results: map[Benign:true Confident:true]
|
| 25/2 |
2026/04/06 21:07 |
llm |
expert |
0m
Model:
Tokens:
input: 48235
output: 557
thoughts: 282
Thoughts: **My Analysis of the Pipe Files Race Condition**
Okay, so the user wants a succinct summary of my analysis, which I've already formulated in the preceding sections. My aim here is to provide that distilled conclusion following established protocol. I've broken down the race condition involving `pipe_release` and `pipe_wait_writable` and their interaction with `pipe->files`.
My summary, in a nutshell, boils down to this: This is a **BENIGN** data race.
Here's the rationale. `pipe->files` is essentially a reference count. The code tracks open file descriptors. The real synchronization concern revolves around the fact that `pipe_release` holds the `inode->i_lock` (through `put_pipe_info`), while `pipe_wait_writable` (through `pipe_lock`) does *not*. However, `pipe_lock` *does* check `pipe->files` before acquiring a mutex.
Crucially, any system call context where we're seeing this happen, like `vmsplice`, means there's at least *one* file descriptor open. A reported race of (2 -> 1) indicates another file was closed. In *either* scenarioβ2 open files or 1β`pipe_lock` sees a non-zero count, and proceeds to take the mutex. Therefore, the code *always* uses the mutex if needed. Even though the value may change and cause a data race, the mutex guards the critical section regardless. The mutex ensures that the access to any shared data (like the pipe's internal buffers) is properly synchronized. The data race can be safely ignored.
Ultimately, although there's a potential race, the intended behavior and safety of the locking mechanism are maintained. I'd recommend, for clarity, perhaps employing `data_race()` or `READ_ONCE()` to handle the `pipe->files` read in `pipe_lock` and the write in `pipe_unlock` for future clarity and potential analysis tools.
|
| 26/1 |
2026/04/06 21:07 |
action |
format-explanation |
0m
Results: map[Explanation:- **Race Summary:** `pipe_release` vs `pipe_wait_writable` on `pipe->files`.
- **Synchronization context:** `pipe_release` (via `put_pipe_info`) decrements
`pipe->files` under `inode->i_lock`. `pipe_wait_writable` (via `pipe_lock`)
reads `pipe->files` without holding `inode->i_lock` to determine if it should
acquire `pipe->mutex`.
- **Final Classification:** **BENIGN**
### Reasoning
The race is on the `pipe->files` field, which tracks the number of `struct file`
objects referring to a pipe.
1. **Contextual Guarantee:** In the reported race, `pipe_wait_writable` is
called from `vmsplice_to_pipe`. This syscall holds a reference to the pipe's
file descriptor (via `fdget`). This ensures that `pipe->files` is at least 1 as
long as the syscall is active.
2. **Safe Transition:** The KCSAN report shows `pipe->files` changing from `2`
to `1`. This occurs because another thread closed a different file descriptor
for the same pipe.
3. **Logical Tolerance:** In `pipe_lock(pipe)`, the code checks `if
(pipe->files)`. Since the current thread holds a file reference, `pipe->files`
will be at least 1. Whether the read sees `2` or `1`, the result of the `if`
condition is `true`, and `mutex_lock(&pipe->mutex)` is correctly called.
4. **No Livelock/UAF:** The check in `pipe_lock` is a "belt and suspenders"
safety measure. If `pipe->files` were 0, it would mean the pipe is being
destroyed, but this is impossible while a thread holds a valid file reference.
Even if the compiler were to re-read the value, it would remain non-zero.
The race is structurally tolerant to stale or concurrent updates because the
outcome of the branch (`true`) is invariant in the context of an active file
operation.
### Recommended Annotations
The accesses to `pipe->files` in `pipe_lock()` and `pipe_unlock()` should be
marked with `data_race()` or `READ_ONCE()` to acknowledge the lockless read and
silence KCSAN.
```c
void pipe_lock(struct pipe_inode_info *pipe)
{
if (data_race(pipe->files))
mutex_lock(&pipe->mutex);
}
```]
|