| 0/0 |
2026/05/17 11:07 |
flow |
assessment-security |
36m
Results: map[DenialOfService:true Explanation:**Analysis of the Bug**
The bug report describes a KASAN slab-use-after-free in `__bpf_trace_run` when
accessing `link->link.prog`. The object being accessed is a `struct
bpf_raw_tp_link` which was allocated during `bpf_raw_tracepoint_open` and freed
via `bpf_link_release`.
**Root Cause:**
When a raw tracepoint BPF link is released, `bpf_link_free()` schedules the
deallocation of the `struct bpf_raw_tp_link` object using `call_rcu()` (which
waits for a normal RCU grace period). However, tracepoints (like
`trace_mm_page_alloc` in the report) are protected by SRCU (`tracepoint_srcu`)
via `guard(srcu_fast_notrace)(&tracepoint_srcu)`.
Because SRCU read-side critical sections can be preempted and can outlive normal
RCU grace periods, the normal RCU grace period can completeβand the `link`
object can be freedβwhile the tracepoint is still executing. When the tracepoint
code resumes and accesses `link->link.prog`, it hits a use-after-free.
**Exploitability:**
This is a classic use-after-free vulnerability. An attacker who can trigger the
bug can potentially reallocate the freed `struct bpf_raw_tp_link` object with
controlled data. By controlling the `prog` pointer, the attacker can hijack the
control flow when `bpf_prog_run(prog, args)` is called, leading to arbitrary
kernel code execution. It can also trivially cause a kernel panic (Denial of
Service).
**Privilege Requirements:**
To trigger this bug, a user must be able to load a BPF program and attach it to
a raw tracepoint.
1. **Unprivileged Processes:** Unprivileged users cannot load the required BPF
program types. `bpf_raw_tp_link_attach()` strictly requires programs of type
`BPF_PROG_TYPE_RAW_TRACEPOINT`, `BPF_PROG_TYPE_TRACING`, `BPF_PROG_TYPE_EXT`, or
`BPF_PROG_TYPE_LSM`. Loading any of these program types requires `CAP_BPF` or
`CAP_SYS_ADMIN` in the initial user namespace. (Even if
`unprivileged_bpf_disabled=0`, unprivileged users are restricted to
`SOCKET_FILTER` and `CGROUP_SKB`, which are rejected by
`bpf_raw_tp_link_attach()`).
2. **User Namespaces:** The capability checks for loading BPF programs
(`bpf_capable()`) check against the initial user namespace. While the kernel
supports BPF tokens to delegate capabilities to user namespaces, creating a BPF
token requires mounting `bpffs` with delegation options (e.g.,
`delegate_progs`). Parsing these delegation mount options explicitly requires
`capable(CAP_SYS_ADMIN)` in the initial user namespace. Thus, an unprivileged
user cannot set up a user namespace to bypass these restrictions.
**Other Triggers:**
This bug is strictly local and specific to the BPF subsystem. It cannot be
triggered via network traffic, remote interactions, malicious filesystems, USB
peripherals, or from within a VM guest.
**Conclusion:**
The bug is a highly exploitable memory corruption issue that can lead to
privilege escalation and denial of service. However, it acts as a
post-compromise/lockdown-bypass vulnerability rather than a primary entry point,
as it strictly requires `CAP_BPF` or `CAP_SYS_ADMIN` privileges to be triggered. Exploitable:true FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
|
| 1/1 |
2026/05/17 11:07 |
action |
syz-repro-to-c-repro |
0m
Results:
|
| 2/1 |
2026/05/17 11:07 |
action |
kernel-checkouter |
1m
Results: map[KernelSrc:/app/workdir/cache/src/0ace5a22bf58e4c37f72d8e2ed6b5cc1da24d6d1]
|
| 3/1 |
2026/05/17 11:08 |
action |
kernel-builder |
14m
Results: map[KernelObj:/app/workdir/cache/build/0a82e5a4cad5f01fcdea81a7856329efb097675d]
|
| 4/1 |
2026/05/17 11:23 |
action |
codesearch-prepare |
15m
Results: map[Index:codesearch-index]
|
| 5/1 |
2026/05/17 11:38 |
agent |
expert |
6m
Model:
Results: map[DenialOfService:true Exploitable:true FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
Instruction:
You are an experienced Linux kernel security engineer. Your task is to analyze given kernel bug report
and determine its security impact based on the following dimensions.
Use the provided tools to examine the source code, check for capability checks (e.g., capable(), ns_capable()),
and understand the nature of the bug. Analyze the given kernel build and configuration.
You can check the kernel config by grepping ".config" file; you can check kernel cmdline by greeping
".config" file for "CONFIG_CMDLINE=". Assume sysctl parameters have default values.
But analyze for the corresponding production build w/o debugging tools enabled (like KASAN, KMSAN, UBSAN).
Don't make assumptions; verify them with source code access. Try different strategies when analyzing the bug:
- think of ways in which the vulnerable code is unreachable
- or the other way around: try to come up with different ideas of how an unprivileged user can reach the bug
If still unsure err on the side of the bug being non-exploitable/not-accessible.
In the final reply, provide a reasoning for your assessment.
Analysis dimensions:
* Exploitable:
Determine if the bug can result in memory corruption or elevated privileges.
Memory safety issues are almost always exploitable (KASAN or UBSAN reports for use-after-free, out-of-bounds;
refcounting issues, corrupted lists, etc). When kernel is crashing on a completly wild pointer access
(e.g. user-space address, or non-canonical address, but not on NULL or address corresponding to KASAN shadow
for NULL address), including both data accesses and control tranfers, that's also usually implies possibility
of exploitation. Such reports usually say "unable to handle kernel paging request".
Uses of uninitialized values detected by KMSAN may be exploitable b/c attacker frequently can affect uninit
values with spraying techniques. However, for these exploitabability depends on how exactly the uninit value
is used in the code, and what it affects.
Think of what happens after the bug is triggered. Some bugs cause kernel panic and halt execution,
they are harder to exploit. For example, BUG reports halts the kernel. However, WARNING reports don't halt
execution in production builds. Debug bug detection tools (like KASAN, KMSAN, KCSAN, UBSAN) are also not enabled
in production builds, so attacker can freely exploit these bugs w/o being detected by these tools.
If you see an integer overflow, think how the overflowed value used later (if it's used as allocation size,
or an array index). If you see an out-of-bounds read, think if it's followed by an out-of-bounds write as well.
Some KCSAN data-races may be exploitable by skilled attackers as well. Think what data structures got corrupted
as the result of data races and how. However, note that kernel has lots of "benign" data races that don't lead
to any runtime misbehavior at all.
* Denial Of Service:
Determine if the bug can result in denial-of-service. Most bugs can, since they cause system crash,
hangs, deadlocks, or resource leaks. This is mostly applicable to WARNING bugs that won't cause system crash
in production. For these think what will be consequences of the violation of the kernel assumptions flagged
by the WARNING. In some cases the unexpected condition is also properly handled by the normal control flow
(e.g. with "if (WARN_ON(...))"), these won't cause denial-of-service. If the condition is not handled,
then it may or may not cause denial-of-service.
* Accessible From Unprivileged Processes:
Determine if the bug can be reached from a typical (non-root) user process that does NOT have any special capabilities
(like CAP_SYS_ADMIN, CAP_NET_ADMIN, CAP_NET_RAW, CAP_PERFMON) or access to device nodes restricted to root.
Assume that unprivileged_bpf_disabled=1, that is eBPF loading is not accessible. However, cBPF (classical BPF)
is still accessible to non-root processes.
Assume that user namespaces are not accessible, that is, the process cannot get the mentioned capabilities even
within a new user namespace (checked by ns_capable() function in the kernel sources).
* Accessible From User Namespaces:
Determine if the bug can be reached within a user-namespace where the process has all capabilities
(including CAP_SYS_ADMIN, CAP_NET_ADMIN, CAP_NET_RAW, CAP_PERFMON). Such capabilities are checked with ns_capable()
function in the kernel sources.
* VM Guest Trigger:
Determine if the bug can be triggered from the context of a typical KVM guest (e.g., set up by a QEMU VMM).
Consider accesses to standard Linux host paravirtualized features (virtio-blk, virtio-net, etc.),
and handling of VM exits in the KVM code.
* VM Host Trigger in The Confidetial Computing Context:
Determine if the bug can be triggered in a confidential computing guest kernel from the context of a KVM host.
Consider access to standard Linux guest paravirtualized features (virtio-blk, virtio-net, etc.).
* Ethernet Network Trigger:
Determine if the bug can be triggered by processing ingress network Ethernet traffic, either directly (network stack)
or via drivers exposed to network data.
* Other Remote Trigger:
Determine if the bug can be triggered by processing remote traffic other than Ethernet (Wifi, Bluetooth, NFC, etc).
* Peripheral Trigger:
Determine if the bug can be triggered via an untrusted peripheral device that can be physically plugged
into a system, such as a USB device or a niche hardware driver handling external hardware inputs.
This is particularly important for mobile and desktop environments where users can plug in unknown devices.
* Malicious Filesystem Trigger:
Determine if the bug can be triggered by the kernel mounting and parsing a malicious filesystem image.
This is highly critical for Desktop and Mobile environments where external media or downloaded images
might be auto-mounted.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt:
The kernel bug report is:
==================================================================
BUG: KASAN: slab-use-after-free in __bpf_trace_run kernel/trace/bpf_trace.c:2075 [inline]
BUG: KASAN: slab-use-after-free in bpf_trace_run4+0xe6/0x850 kernel/trace/bpf_trace.c:2131
Read of size 8 at addr ffff8880361a0318 by task udevd/5299
CPU: 0 UID: 0 PID: 5299 Comm: udevd Not tainted syzkaller #0 PREEMPT(full)
Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
Call Trace:
<TASK>
dump_stack_lvl+0xe8/0x150 lib/dump_stack.c:120
print_address_description mm/kasan/report.c:378 [inline]
print_report+0xba/0x230 mm/kasan/report.c:482
kasan_report+0x117/0x150 mm/kasan/report.c:595
__bpf_trace_run kernel/trace/bpf_trace.c:2075 [inline]
bpf_trace_run4+0xe6/0x850 kernel/trace/bpf_trace.c:2131
__traceiter_mm_page_alloc+0x3d/0x60 include/trace/events/kmem.h:180
__do_trace_mm_page_alloc include/trace/events/kmem.h:180 [inline]
trace_mm_page_alloc+0x149/0x180 include/trace/events/kmem.h:180
__alloc_frozen_pages_noprof+0x1de/0x380 mm/page_alloc.c:5272
alloc_pages_mpol+0x232/0x4a0 mm/mempolicy.c:2484
folio_alloc_mpol_noprof+0x39/0x70 mm/mempolicy.c:2503
swap_cache_alloc_folio+0xd5/0x240 mm/swap_state.c:571
swap_cluster_readahead+0x369/0x690 mm/swap_state.c:749
swapin_readahead+0x196/0xc50 mm/swap_state.c:924
do_swap_page+0x56f/0x5a20 mm/memory.c:4802
handle_pte_fault mm/memory.c:6320 [inline]
__handle_mm_fault mm/memory.c:6455 [inline]
handle_mm_fault+0x12d2/0x3310 mm/memory.c:6624
do_user_addr_fault+0xa73/0x1340 arch/x86/mm/fault.c:1334
handle_page_fault arch/x86/mm/fault.c:1474 [inline]
exc_page_fault+0x6a/0xc0 arch/x86/mm/fault.c:1527
asm_exc_page_fault+0x26/0x30 arch/x86/include/asm/idtentry.h:618
RIP: 0033:0x55f04dab55f0
Code: c0 0f 85 0e 19 00 00 4c 8b 73 18 c7 44 24 28 00 00 00 00 49 89 dc 4c 8d 3d e9 95 02 00 4c 89 34 24 66 0f 1f 84 00 00 00 00 00 <41> 0f b6 1e 80 fb 35 0f 87 e3 01 00 00 0f b6 c3 49 63 04 87 4c 01
RSP: 002b:00007fff48a349c0 EFLAGS: 00010206
RAX: 000055f051ea3770 RBX: 0000000000000034 RCX: 0000000000000063
RDX: 0000000000000381 RSI: 000055f051eb5b50 RDI: 000055f051ecfdae
RBP: 000055f052096b80 R08: 000055f04daf2100 R09: 000055f04daf2140
R10: 0000000000000000 R11: 0000000000000000 R12: 000055f051eb5270
R13: 000055f051eb02c0 R14: 000055f051ea4574 R15: 000055f04dadebcc
</TASK>
Allocated by task 5326:
kasan_save_stack mm/kasan/common.c:57 [inline]
kasan_save_track+0x3e/0x80 mm/kasan/common.c:78
poison_kmalloc_redzone mm/kasan/common.c:398 [inline]
__kasan_kmalloc+0x93/0xb0 mm/kasan/common.c:415
kasan_kmalloc include/linux/kasan.h:263 [inline]
__kmalloc_cache_noprof+0x31c/0x660 mm/slub.c:5383
kmalloc_noprof include/linux/slab.h:950 [inline]
kzalloc_noprof include/linux/slab.h:1188 [inline]
bpf_raw_tp_link_attach+0x278/0x700 kernel/bpf/syscall.c:4264
bpf_raw_tracepoint_open+0x1b2/0x220 kernel/bpf/syscall.c:4312
__sys_bpf+0x846/0x950 kernel/bpf/syscall.c:6270
__do_sys_bpf kernel/bpf/syscall.c:6341 [inline]
__se_sys_bpf kernel/bpf/syscall.c:6339 [inline]
__x64_sys_bpf+0x7c/0x90 kernel/bpf/syscall.c:6339
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x14d/0xf80 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
Freed by task 15:
kasan_save_stack mm/kasan/common.c:57 [inline]
kasan_save_track+0x3e/0x80 mm/kasan/common.c:78
kasan_save_free_info+0x46/0x50 mm/kasan/generic.c:584
poison_slab_object mm/kasan/common.c:253 [inline]
__kasan_slab_free+0x5c/0x80 mm/kasan/common.c:285
kasan_slab_free include/linux/kasan.h:235 [inline]
slab_free_hook mm/slub.c:2692 [inline]
slab_free mm/slub.c:6168 [inline]
kfree+0x1c1/0x630 mm/slub.c:6486
rcu_do_batch kernel/rcu/tree.c:2617 [inline]
rcu_core+0x7cd/0x1070 kernel/rcu/tree.c:2869
handle_softirqs+0x22a/0x870 kernel/softirq.c:622
run_ksoftirqd+0x36/0x60 kernel/softirq.c:1063
smpboot_thread_fn+0x541/0xa50 kernel/smpboot.c:160
kthread+0x388/0x470 kernel/kthread.c:436
ret_from_fork+0x51e/0xb90 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
Last potentially related work creation:
kasan_save_stack+0x3e/0x60 mm/kasan/common.c:57
kasan_record_aux_stack+0xbd/0xd0 mm/kasan/generic.c:556
__call_rcu_common kernel/rcu/tree.c:3131 [inline]
call_rcu+0xee/0x890 kernel/rcu/tree.c:3251
bpf_link_put_direct kernel/bpf/syscall.c:3323 [inline]
bpf_link_release+0x6b/0x80 kernel/bpf/syscall.c:3330
__fput+0x44f/0xa70 fs/file_table.c:469
task_work_run+0x1d9/0x270 kernel/task_work.c:233
exit_task_work include/linux/task_work.h:40 [inline]
do_exit+0x70f/0x23c0 kernel/exit.c:976
do_group_exit+0x21b/0x2d0 kernel/exit.c:1118
get_signal+0x1284/0x1330 kernel/signal.c:3034
arch_do_signal_or_restart+0xbc/0x830 arch/x86/kernel/signal.c:337
__exit_to_user_mode_loop kernel/entry/common.c:64 [inline]
exit_to_user_mode_loop kernel/entry/common.c:98 [inline]
__exit_to_user_mode_prepare include/linux/irq-entry-common.h:226 [inline]
irqentry_exit_to_user_mode_prepare include/linux/irq-entry-common.h:270 [inline]
irqentry_exit_to_user_mode include/linux/irq-entry-common.h:339 [inline]
irqentry_exit+0x176/0x620 kernel/entry/common.c:219
asm_exc_page_fault+0x26/0x30 arch/x86/include/asm/idtentry.h:618
The buggy address belongs to the object at ffff8880361a0300
which belongs to the cache kmalloc-192 of size 192
The buggy address is located 24 bytes inside of
freed 192-byte region [ffff8880361a0300, ffff8880361a03c0)
The buggy address belongs to the physical page:
page: refcount:0 mapcount:0 mapping:0000000000000000 index:0xffff8880361a0400 pfn:0x361a0
flags: 0x4fff00000000200(workingset|node=1|zone=1|lastcpupid=0x7ff)
page_type: f5(slab)
raw: 04fff00000000200 ffff88801ac413c0 ffffea0000d7ce90 ffffea0000e14e90
raw: ffff8880361a0400 000000080010000f 00000000f5000000 0000000000000000
page dumped because: kasan: bad access detected
page_owner tracks the page as allocated
page last allocated via order 0, migratetype Unmovable, gfp_mask 0xd2cc0(GFP_KERNEL|__GFP_NOWARN|__GFP_NORETRY|__GFP_COMP|__GFP_NOMEMALLOC), pid 922, tgid 922 (kworker/0:3), ts 22003490640, free_ts 22003061349
set_page_owner include/linux/page_owner.h:32 [inline]
post_alloc_hook+0x231/0x280 mm/page_alloc.c:1889
prep_new_page mm/page_alloc.c:1897 [inline]
get_page_from_freelist+0x24dc/0x2580 mm/page_alloc.c:3962
__alloc_frozen_pages_noprof+0x18d/0x380 mm/page_alloc.c:5250
alloc_slab_page mm/slub.c:3296 [inline]
allocate_slab+0x77/0x660 mm/slub.c:3485
new_slab mm/slub.c:3543 [inline]
refill_objects+0x331/0x3c0 mm/slub.c:7178
__pcs_replace_empty_main+0x2f9/0x5e0 mm/slub.c:-1
alloc_from_pcs mm/slub.c:4720 [inline]
slab_alloc_node mm/slub.c:4854 [inline]
__do_kmalloc_node mm/slub.c:5262 [inline]
__kmalloc_noprof+0x474/0x760 mm/slub.c:5275
kmalloc_noprof include/linux/slab.h:954 [inline]
virtio_gpu_array_alloc+0x26/0xc0 drivers/gpu/drm/virtio/virtgpu_gem.c:170
virtio_gpu_update_dumb_bo drivers/gpu/drm/virtio/virtgpu_plane.c:171 [inline]
virtio_gpu_primary_plane_update+0x38d/0x13a0 drivers/gpu/drm/virtio/virtgpu_plane.c:265
drm_atomic_helper_commit_planes+0x60f/0xec0 drivers/gpu/drm/drm_atomic_helper.c:3038
drm_atomic_helper_commit_tail+0x5f/0x500 drivers/gpu/drm/drm_atomic_helper.c:1989
commit_tail+0x29a/0x3a0 drivers/gpu/drm/drm_atomic_helper.c:2074
drm_atomic_helper_commit+0xa6e/0xb10 drivers/gpu/drm/drm_atomic_helper.c:2312
drm_atomic_commit+0x246/0x2b0 drivers/gpu/drm/drm_atomic.c:1775
drm_atomic_helper_dirtyfb+0xdec/0xf80 drivers/gpu/drm/drm_damage_helper.c:183
drm_fbdev_shmem_helper_fb_dirty+0x160/0x2d0 drivers/gpu/drm/drm_fbdev_shmem.c:117
page last free pid 70 tgid 70 stack trace:
reset_page_owner include/linux/page_owner.h:25 [inline]
__free_pages_prepare mm/page_alloc.c:1433 [inline]
__free_frozen_pages+0xc2b/0xdb0 mm/page_alloc.c:2978
___free_pages_bulk mm/kasan/shadow.c:333 [inline]
__kasan_populate_vmalloc_do mm/kasan/shadow.c:385 [inline]
__kasan_populate_vmalloc+0x137/0x1d0 mm/kasan/shadow.c:424
kasan_populate_vmalloc include/linux/kasan.h:580 [inline]
alloc_vmap_area+0xd73/0x14b0 mm/vmalloc.c:2129
__get_vm_area_node+0x1f8/0x300 mm/vmalloc.c:3232
__vmalloc_node_range_noprof+0x372/0x1730 mm/vmalloc.c:4024
__vmalloc_node_noprof+0xc2/0x100 mm/vmalloc.c:4124
alloc_thread_stack_node kernel/fork.c:355 [inline]
dup_task_struct+0x228/0x9a0 kernel/fork.c:924
copy_process+0x508/0x3cf0 kernel/fork.c:2050
kernel_clone+0x248/0x8e0 kernel/fork.c:2654
user_mode_thread+0x110/0x180 kernel/fork.c:2730
call_usermodehelper_exec_work+0x5c/0x230 kernel/umh.c:171
process_one_work kernel/workqueue.c:3275 [inline]
process_scheduled_works+0xb02/0x1830 kernel/workqueue.c:3358
worker_thread+0xa50/0xfc0 kernel/workqueue.c:3439
kthread+0x388/0x470 kernel/kthread.c:436
ret_from_fork+0x51e/0xb90 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
Memory state around the buggy address:
ffff8880361a0200: fa fb fb fb fb fb fb fb fb fb fb fb fb fb fb fb
ffff8880361a0280: fb fb fb fb fb fb fb fb fc fc fc fc fc fc fc fc
>ffff8880361a0300: fa fb fb fb fb fb fb fb fb fb fb fb fb fb fb fb
^
ffff8880361a0380: fb fb fb fb fb fb fb fb fc fc fc fc fc fc fc fc
ffff8880361a0400: fa fb fb fb fb fb fb fb fb fb fb fb fb fb fb fb
==================================================================
Reply: **Analysis of the Bug**
The bug report describes a KASAN slab-use-after-free in `__bpf_trace_run` when accessing `link->link.prog`. The object being accessed is a `struct bpf_raw_tp_link` which was allocated during `bpf_raw_tracepoint_open` and freed via `bpf_link_release`.
**Root Cause:**
When a raw tracepoint BPF link is released, `bpf_link_free()` schedules the deallocation of the `struct bpf_raw_tp_link` object using `call_rcu()` (which waits for a normal RCU grace period). However, tracepoints (like `trace_mm_page_alloc` in the report) are protected by SRCU (`tracepoint_srcu`) via `guard(srcu_fast_notrace)(&tracepoint_srcu)`.
Because SRCU read-side critical sections can be preempted and can outlive normal RCU grace periods, the normal RCU grace period can completeβand the `link` object can be freedβwhile the tracepoint is still executing. When the tracepoint code resumes and accesses `link->link.prog`, it hits a use-after-free.
**Exploitability:**
This is a classic use-after-free vulnerability. An attacker who can trigger the bug can potentially reallocate the freed `struct bpf_raw_tp_link` object with controlled data. By controlling the `prog` pointer, the attacker can hijack the control flow when `bpf_prog_run(prog, args)` is called, leading to arbitrary kernel code execution. It can also trivially cause a kernel panic (Denial of Service).
**Privilege Requirements:**
To trigger this bug, a user must be able to load a BPF program and attach it to a raw tracepoint.
1. **Unprivileged Processes:** Unprivileged users cannot load the required BPF program types. `bpf_raw_tp_link_attach()` strictly requires programs of type `BPF_PROG_TYPE_RAW_TRACEPOINT`, `BPF_PROG_TYPE_TRACING`, `BPF_PROG_TYPE_EXT`, or `BPF_PROG_TYPE_LSM`. Loading any of these program types requires `CAP_BPF` or `CAP_SYS_ADMIN` in the initial user namespace. (Even if `unprivileged_bpf_disabled=0`, unprivileged users are restricted to `SOCKET_FILTER` and `CGROUP_SKB`, which are rejected by `bpf_raw_tp_link_attach()`).
2. **User Namespaces:** The capability checks for loading BPF programs (`bpf_capable()`) check against the initial user namespace. While the kernel supports BPF tokens to delegate capabilities to user namespaces, creating a BPF token requires mounting `bpffs` with delegation options (e.g., `delegate_progs`). Parsing these delegation mount options explicitly requires `capable(CAP_SYS_ADMIN)` in the initial user namespace. Thus, an unprivileged user cannot set up a user namespace to bypass these restrictions.
**Other Triggers:**
This bug is strictly local and specific to the BPF subsystem. It cannot be triggered via network traffic, remote interactions, malicious filesystems, USB peripherals, or from within a VM guest.
**Conclusion:**
The bug is a highly exploitable memory corruption issue that can lead to privilege escalation and denial of service. However, it acts as a post-compromise/lockdown-bypass vulnerability rather than a primary entry point, as it strictly requires `CAP_BPF` or `CAP_SYS_ADMIN` privileges to be triggered.
|
| 6/2 |
2026/05/17 11:38 |
llm |
expert |
0m
Model:
Tokens:
input: 9346
output: 46
thoughts: 408
Thoughts: **KASAN Slab-Use-After-Free Investigation**
Okay, so I've got a KASAN report, specifically a slab-use-after-free, and it's pointing to either `__bpf_trace_run` at `kernel/trace/bpf_trace.c:2075` or `bpf_trace_run4+0xe6/0x850` at `kernel/trace/bpf_trace.c:2131`. The critical read, size 8 bytes, is happening at `ffff8880361a0318`. Right, let's trace the object's lifecycle. It looks like the object allocation originated from `bpf_raw_tp_link_attach`, which then calls `bpf_raw_tracepoint_open`, which finally calls `__sys_bpf`. And the deallocation seems to be triggered by `rcu_do_batch`, eventually leading to `kfree` after a sequence that includes `bpf_link_release`, `bpf_link_put_direct`, and `call_rcu`.
So the key functions here are `bpf_trace_run4` and `__bpf_trace_run`. Hmm, the tracepoint involved is `trace_mm_page_alloc`. Let's confirm with the call trace they've provided: I'm seeing `__bpf_trace_run`, then `bpf_trace_run4`, followed by `__traceiter_mm_page_alloc`, `__do_trace_mm_page_alloc`, and finally `trace_mm_page_alloc`. Makes sense.
Alright, time to dive into `kernel/trace/bpf_trace.c` and see what's going on within those functions, and how the interaction with BPF and the RCU mechanism could be causing this. This should be an interesting one.
|
| 7/2 |
2026/05/17 11:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/bpf_trace.c IncludeLines:true Name:__bpf_trace_run]
Results: map[SourceCode:2072: static __always_inline
2073: void __bpf_trace_run(struct bpf_raw_tp_link *link, u64 *args)
2074: {
2075: struct bpf_prog *prog = link->link.prog;
2076: struct bpf_run_ctx *old_run_ctx;
2077: struct bpf_trace_run_ctx run_ctx;
2078:
2079: rcu_read_lock_dont_migrate();
2080: if (unlikely(!bpf_prog_get_recursion_context(prog))) {
2081: bpf_prog_inc_misses_counter(prog);
2082: goto out;
2083: }
2084:
2085: run_ctx.bpf_cookie = link->cookie;
2086: old_run_ctx = bpf_set_run_ctx(&run_ctx.run_ctx);
2087:
2088: (void) bpf_prog_run(prog, args);
2089:
2090: bpf_reset_run_ctx(old_run_ctx);
2091: out:
2092: bpf_prog_put_recursion_context(prog);
2093: rcu_read_unlock_migrate();
2094: }
SourceFile:kernel/trace/bpf_trace.c]
|
| 8/2 |
2026/05/17 11:38 |
llm |
expert |
0m
Model:
Tokens:
input: 10624
output: 47
thoughts: 53
Thoughts: **Analyzing the BPF Raw Tracepoint Link Attachment**
Okay, so I'm looking at this `link->link.prog` read, which tells me we're dealing with a BPF raw tracepoint link. The object in question is a `struct bpf_raw_tp_link`. Now, to truly understand how this link functions, I need to dive into the core attachment logic. My first stop has to be the `bpf_raw_tp_link_attach` function. I'm going to find that in `kernel/bpf/syscall.c`. This is where the magic happens, where the raw tracepoint link actually gets connected. I need to understand what's happening under the hood when a program attaches to this link. I'll need to dissect that function line by line to determine exactly how it is attaching and what checks or processes it is using.
|
| 9/2 |
2026/05/17 11:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_raw_tp_link_attach]
Results: map[SourceCode:4223: static int bpf_raw_tp_link_attach(struct bpf_prog *prog,
4224: const char __user *user_tp_name, u64 cookie,
4225: enum bpf_attach_type attach_type)
4226: {
4227: struct bpf_link_primer link_primer;
4228: struct bpf_raw_tp_link *link;
4229: struct bpf_raw_event_map *btp;
4230: const char *tp_name;
4231: char buf[128];
4232: int err;
4233:
4234: switch (prog->type) {
4235: case BPF_PROG_TYPE_TRACING:
4236: case BPF_PROG_TYPE_EXT:
4237: case BPF_PROG_TYPE_LSM:
4238: if (user_tp_name)
4239: /* The attach point for this category of programs
4240: * should be specified via btf_id during program load.
4241: */
4242: return -EINVAL;
4243: if (prog->type == BPF_PROG_TYPE_TRACING &&
4244: prog->expected_attach_type == BPF_TRACE_RAW_TP) {
4245: tp_name = prog->aux->attach_func_name;
4246: break;
4247: }
4248: return bpf_tracing_prog_attach(prog, 0, 0, 0, attach_type);
4249: case BPF_PROG_TYPE_RAW_TRACEPOINT:
4250: case BPF_PROG_TYPE_RAW_TRACEPOINT_WRITABLE:
4251: if (strncpy_from_user(buf, user_tp_name, sizeof(buf) - 1) < 0)
4252: return -EFAULT;
4253: buf[sizeof(buf) - 1] = 0;
4254: tp_name = buf;
4255: break;
4256: default:
4257: return -EINVAL;
4258: }
4259:
4260: btp = bpf_get_raw_tracepoint(tp_name);
4261: if (!btp)
4262: return -ENOENT;
4263:
4264: link = kzalloc_obj(*link, GFP_USER);
4265: if (!link) {
4266: err = -ENOMEM;
4267: goto out_put_btp;
4268: }
4269: bpf_link_init_sleepable(&link->link, BPF_LINK_TYPE_RAW_TRACEPOINT,
4270: &bpf_raw_tp_link_lops, prog, attach_type,
4271: tracepoint_is_faultable(btp->tp));
4272: link->btp = btp;
4273: link->cookie = cookie;
4274:
4275: err = bpf_link_prime(&link->link, &link_primer);
4276: if (err) {
4277: kfree(link);
4278: goto out_put_btp;
4279: }
4280:
4281: err = bpf_probe_register(link->btp, link);
4282: if (err) {
4283: bpf_link_cleanup(&link_primer);
4284: goto out_put_btp;
4285: }
4286:
4287: return bpf_link_settle(&link_primer);
4288:
4289: out_put_btp:
4290: bpf_put_raw_tracepoint(btp);
4291: return err;
4292: }
SourceFile:kernel/bpf/syscall.c]
|
| 10/2 |
2026/05/17 11:38 |
llm |
expert |
0m
Model:
Tokens:
input: 12034
output: 48
thoughts: 116
Thoughts: **My Analysis of the BPF Raw Tracepoint Link Attachment and Release**
Okay, so let's break this down. First, the `bpf_raw_tp_link_attach` function, it seems like the initial step here is allocating the `link` object. The allocation uses `kzalloc_obj`, which is pretty standard for kernel memory allocation, ensuring initialization to zero. Makes sense for a new object.
Next, it's registered with `bpf_probe_register(link->btp, link)`. This probably involves associating the tracepoint with the BPF program, establishing the core connection. I'd need to dig deeper into `bpf_probe_register` to confirm exactly how that's done, but that's my working hypothesis.
Now, here's where it gets interesting β the deallocation process. The link is eventually freed through a chain of function calls: `bpf_link_release` -> `bpf_link_put_direct` -> `call_rcu`. So it utilizes the RCU mechanism to ensure that existing readers of the link can finish their work.
But wait, hold on a second! `bpf_link_release` doesn't just jump directly into that RCU deallocation. It first calls `link->ops->release(link)`. That's key. It's calling a release operation specific to this particular type of BPF link, this is crucial for releasing any resources that are specific to the raw tracepoint link. I need to understand that `ops->release` function.
Therefore, let's go check what is in `bpf_raw_tp_link_lops`, because that's where I bet the specific implementation is held. I need to be sure what resources must be released specifically for this kind of link, before it is freed.
|
| 11/2 |
2026/05/17 11:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_raw_tp_link_lops]
Results: map[SourceCode:3850: static const struct bpf_link_ops bpf_raw_tp_link_lops = {
3851: .release = bpf_raw_tp_link_release,
3852: .dealloc_deferred = bpf_raw_tp_link_dealloc,
3853: .show_fdinfo = bpf_raw_tp_link_show_fdinfo,
3854: .fill_link_info = bpf_raw_tp_link_fill_link_info,
3855: };
SourceFile:kernel/bpf/syscall.c]
|
| 12/2 |
2026/05/17 11:38 |
llm |
expert |
0m
Model:
Tokens:
input: 12762
output: 47
thoughts: 30
Thoughts: Okay, here's my interpretation of those thoughts, presented as a personal thought process, as requested:
**Diving into BPF Tracepoint Link Management**
Alright, time to dig into the nitty-gritty of raw tracepoint linking. Specifically, I need to take a closer look at the interaction between `bpf_raw_tp_link_release` and `bpf_raw_tp_link_dealloc`. Understanding their relationship is crucial for ensuring proper resource management within the BPF system. I need to figure out the exact responsibilities of each function, and how they contribute to a clean and stable BPF environment. Are they independent, or does one call the other? What locking is involved? What crucial data structures do they interact with? Is there any chance of memory leaks or dangling pointers if used improperly? I need a clear mental model of their behavior to avoid any pitfalls down the line. Let's see how they work together to manage the lifecycle of these tracepoint links.
|
| 13/2 |
2026/05/17 11:39 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_raw_tp_link_release]
Results: map[SourceCode:3779: static void bpf_raw_tp_link_release(struct bpf_link *link)
3780: {
3781: struct bpf_raw_tp_link *raw_tp =
3782: container_of(link, struct bpf_raw_tp_link, link);
3783:
3784: bpf_probe_unregister(raw_tp->btp, raw_tp);
3785: bpf_put_raw_tracepoint(raw_tp->btp);
3786: }
SourceFile:kernel/bpf/syscall.c]
|
| 14/2 |
2026/05/17 11:39 |
llm |
expert |
0m
Model:
Tokens:
input: 13224
output: 48
thoughts: 14
Thoughts: Okay, here's my attempt at summarizing that thought process, keeping in mind I'm speaking to an expert:
**Deallocating a Tracepoint Link: My Internal Checklist**
Alright, so I'm thinking about `bpf_raw_tp_link_dealloc`. Naturally, I'm going through my mental checklist. This isn't just a simple deallocation; this is about untangling a tracepoint connection within the BPF framework. First, I need to ensure that the resources associated with that specific link are properly released. I'm talking about the data structures that glue the tracepoint to the BPF program, the internal bookkeeping, the link itself, all of it.
I'm making sure there are no dangling pointers or orphaned references. Memory leaks are my enemy here. I have to be absolutely certain that any allocated memory associated with this link is freed and that all relevant internal state is reset or cleared. Think about the potential for race conditions; any concurrent access to those resources needs to be handled elegantly during this deallocation.
I'm considering potential error conditions. What happens if the link isn't valid in the first place? What if a dependent resource can't be freed? I have to account for these possibilities and gracefully handle any failures, likely through some robust error reporting or logging mechanism.
Beyond the immediate deallocation, I'm also considering the broader context. How does this removal affect the overall BPF system? Are there any dependencies that need to be updated or informed about this change? Am I triggering any cleanup routines or callbacks? It's about maintaining the integrity and consistency of the entire system as this link is dismantled. Effectively, I'm thinking like a kernel developer here.
Finally, I'm visualizing the critical sections, lock usage, and how to minimize latency here. The whole point of BPF is performance, so a slow deallocation could cripple other concurrent processes. I'm focusing on making this process as efficient as possible.
|
| 15/2 |
2026/05/17 11:39 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_raw_tp_link_dealloc]
Results: map[SourceCode:3788: static void bpf_raw_tp_link_dealloc(struct bpf_link *link)
3789: {
3790: struct bpf_raw_tp_link *raw_tp =
3791: container_of(link, struct bpf_raw_tp_link, link);
3792:
3793: kfree(raw_tp);
3794: }
SourceFile:kernel/bpf/syscall.c]
|
| 16/2 |
2026/05/17 11:39 |
llm |
expert |
0m
Model:
Tokens:
input: 13842
output: 45
thoughts: 134
Thoughts: **Deconstructing the BPF Link Release Process**
Okay, so I'm tracing this BPF link release, and it's calling `bpf_link_release`, which immediately calls `bpf_raw_tp_link_release`. I know what that does β it unregisters the probe, good. But then things get interesting. `bpf_link_put_direct` gets invoked. That's where I need to pay close attention. It looks like `bpf_link_put_direct` then calls `call_rcu` with `&link->rcu` and `bpf_link_free` as the callback. Ah, got it, so the actual freeing is deferred. Then, `bpf_link_free` gets called, and *that* calls `link->ops->dealloc_deferred(link)`. I'm starting to see the pattern here. The use of deferred deallocation is key to minimizing disruption. Wait a minute... looking closer at `bpf_link_put_direct`, it looks like `call_rcu` is *only* used if `link->ops->dealloc_deferred` is actually set. That's good - that's the check, and now I need to confirm the conditions for the use of the `call_rcu` mechanism. Gotta go check that function more closely.
|
| 17/2 |
2026/05/17 11:39 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_link_put_direct]
Results: map[SourceCode:3319: static void bpf_link_put_direct(struct bpf_link *link)
3320: {
3321: if (!atomic64_dec_and_test(&link->refcnt))
3322: return;
3323: bpf_link_free(link);
3324: }
SourceFile:kernel/bpf/syscall.c]
|
| 18/2 |
2026/05/17 11:39 |
llm |
expert |
0m
Model:
Tokens:
input: 14433
output: 43
thoughts: 26
Thoughts: Here's my thought process, summarized as if I were considering this myself:
**Investigating a Direct Call to `bpf_link_free`**
Hmm, so `bpf_link_free` is called directly, is it? That's unusual, and immediately raises some flags. I need to understand why this direct call is happening. My initial thought is to jump right into the implementation of `bpf_link_free` itself. Let's dig into the code and see exactly what's going on within that function. I need to understand its context, its dependencies, and what potential implications this direct call might have. I want to look for any potential for resource leaks, unexpected behavior, or security vulnerabilities that might arise from this direct invocation. My expert intuition tells me to proceed with caution and perform a thorough analysis.
|
| 19/2 |
2026/05/17 11:39 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_link_free]
Results: map[SourceCode:3272: /* bpf_link_free is guaranteed to be called from process context */
3273: static void bpf_link_free(struct bpf_link *link)
3274: {
3275: const struct bpf_link_ops *ops = link->ops;
3276:
3277: bpf_link_free_id(link->id);
3278: /* detach BPF program, clean up used resources */
3279: if (link->prog)
3280: ops->release(link);
3281: if (ops->dealloc_deferred) {
3282: /* Schedule BPF link deallocation, which will only then
3283: * trigger putting BPF program refcount.
3284: * If underlying BPF program is sleepable or BPF link's target
3285: * attach hookpoint is sleepable or otherwise requires RCU GPs
3286: * to ensure link and its underlying BPF program is not
3287: * reachable anymore, we need to first wait for RCU tasks
3288: * trace sync, and then go through "classic" RCU grace period
3289: */
3290: if (link->sleepable || (link->prog && link->prog->sleepable))
3291: call_rcu_tasks_trace(&link->rcu, bpf_link_defer_dealloc_mult_rcu_gp);
3292: else
3293: call_rcu(&link->rcu, bpf_link_defer_dealloc_rcu_gp);
3294: } else if (ops->dealloc) {
3295: bpf_link_dealloc(link);
3296: }
3297: }
SourceFile:kernel/bpf/syscall.c]
|
| 20/2 |
2026/05/17 11:39 |
llm |
expert |
0m
Model:
Tokens:
input: 15174
output: 45
thoughts: 98
Thoughts: **Deconstructing BPF Link Deallocation**
Okay, so I'm tracing the deallocation of a BPF link here. It looks like `bpf_link_free` is the key function; it initiates the process by calling `call_rcu(&link->rcu, bpf_link_defer_dealloc_rcu_gp)`. That's standard RCU-based deallocation, which is good β allows for grace periods. Then `bpf_link_defer_dealloc_rcu_gp` gets called, and that in turn calls `bpf_link_dealloc(link)`. Makes perfect sense.
Now, a red flag⦠I need to confirm whether `bpf_probe_unregister` is somehow part of the link deallocation process, and if so, how. Let me quickly double-check the code for `bpf_probe_unregister`. I need to ensure the probe is unregistered correctly within this whole cleanup sequence. Just want to make sure the resources are fully cleaned up.
|
| 21/2 |
2026/05/17 11:39 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/bpf_trace.c IncludeLines:true Name:bpf_probe_unregister]
Results: map[SourceCode:2159: int bpf_probe_unregister(struct bpf_raw_event_map *btp, struct bpf_raw_tp_link *link)
2160: {
2161: return tracepoint_probe_unregister(btp->tp, (void *)btp->bpf_func, link);
2162: }
SourceFile:kernel/trace/bpf_trace.c]
|
| 22/2 |
2026/05/17 11:39 |
llm |
expert |
0m
Model:
Tokens:
input: 15657
output: 41
thoughts: 26
Thoughts: Here's my take on those thoughts, tailored for an expert audience:
**Unregistering Tracepoint Probes**
Okay, so we're looking at `tracepoint_probe_unregister`. Seems pretty straightforward β it's the function responsible for, well, *unregistering* a tracepoint probe. That's the core function. I'll need to double check the details of `tracepoint_probe_unregister` to ensure the unregistering process is complete, that all related resources are freed, and that any associated side effects are handled correctly. We're talking about kernel internals here, so I want to make sure we're clean before moving on.
|
| 23/2 |
2026/05/17 11:39 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/tracepoint.c IncludeLines:true Name:tracepoint_probe_unregister]
Results: map[SourceCode: 497: /**
498: * tracepoint_probe_unregister - Disconnect a probe from a tracepoint
499: * @tp: tracepoint
500: * @probe: probe function pointer
501: * @data: tracepoint data
502: *
503: * Returns 0 if ok, error value on error.
504: */
505: int tracepoint_probe_unregister(struct tracepoint *tp, void *probe, void *data)
506: {
507: struct tracepoint_func tp_func;
508: int ret;
509:
510: mutex_lock(&tracepoints_mutex);
511: tp_func.func = probe;
512: tp_func.data = data;
513: ret = tracepoint_remove_func(tp, &tp_func);
514: mutex_unlock(&tracepoints_mutex);
515: return ret;
516: }
SourceFile:kernel/tracepoint.c]
|
| 24/2 |
2026/05/17 11:39 |
llm |
expert |
0m
Model:
Tokens:
input: 16174
output: 49
thoughts: 264
Thoughts: **Analyzing Potential Use-After-Free in Tracepoint Removal**
Okay, so `tracepoint_remove_func` seems to straightforwardly remove the function, but I need to be absolutely sure about how it handles concurrency, which is a key consideration. Does it wait for the RCU grace period before returning? No, it just removes the function from the array and updates the tracepoint pointer. That's a crucial detail. The old array gets freed with `call_rcu(&old->rcu, rcu_free_old_probes)`.
So, the tracepoint *could* be called concurrently with this removal process, meaning it might still be executing while the function is being de-registered. If that happens, it's going to use the old array, which contains that `link` pointer. Hmm. The `link` itself is also freed using `call_rcu`, with `bpf_link_defer_dealloc_rcu_gp`.
Wait a minute, both the old array *and* the `link` are freed via `call_rcu`! This is where things get interesting. If the `call_rcu` for the `link` finishes before the tracepoint completes its execution, then we're staring at a classic use-after-free vulnerability. That's a significant concern and needs careful examination.
But, hang on, the tracepoint execution should be protected by either `rcu_read_lock_sched()` or `rcu_read_lock()`. Does that provide the necessary synchronization? To truly understand this, I should examine a concrete example. Let's look at `trace_mm_page_alloc`. That's a good place to start. I'll check `__do_trace_mm_page_alloc` to see what is going on.
|
| 25/2 |
2026/05/17 11:40 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:include/trace/events/kmem.h IncludeLines:true Name:__do_trace_mm_page_alloc]
Results: map[SourceCode: 180: TRACE_EVENT(mm_page_alloc,
SourceFile:include/trace/events/kmem.h]
|
| 26/2 |
2026/05/17 11:40 |
llm |
expert |
0m
Model:
Tokens:
input: 16923
output: 19
thoughts: 43
Thoughts: Here's my thought process, summarized as you requested:
**Deconstructing `__do_trace_mm_page_alloc`**
Okay, so `__do_trace_mm_page_alloc` is clearly being generated by a `TRACE_EVENT`. That's the starting point. I need to understand how the `TRACE_EVENT` macro actually translates into this specific function name, `__do_trace_##name`. Let's dig into the macro definition itself. I need to trace the macro expansion to see how the preprocessor is handling the `##` operator and the `name` argument, and how that's leading to this specific symbol being created. I have a hunch it's a concatenation issue, of course, but I need to confirm the exact mechanics to understand the code generation completely. Understanding this will be key to understanding the underlying tracing mechanism and what's being captured. I'll need to look for `TRACE_EVENT`'s definition and then trace the flow from there. I expect it's going to be a straightforward macro substitution, but I have to be sure. I need to verify how `name` is passed and used within the macro to produce the `__do_trace_` prefix. That's the critical piece.
|
| 27/2 |
2026/05/17 11:40 |
tool |
grepper |
0m
Args: map[Expression:__do_trace_]
Results: map[Output:include/linux/tracepoint.h=158=static inline struct tracepoint *tracepoint_ptr_deref(tracepoint_ptr_t *p)
--
include/linux/tracepoint.h-221- { \
include/linux/tracepoint.h:222: __do_trace_##name(args); \
include/linux/tracepoint.h-223- }
--
include/linux/tracepoint.h-280- __DECLARE_TRACE_COMMON(name, PARAMS(proto), PARAMS(args), PARAMS(data_proto)) \
include/linux/tracepoint.h:281: static inline void __do_trace_##name(proto) \
include/linux/tracepoint.h-282- { \
--
include/linux/tracepoint.h-291- if (static_branch_unlikely(&__tracepoint_##name.key)) \
include/linux/tracepoint.h:292: __do_trace_##name(args); \
include/linux/tracepoint.h-293- if (IS_ENABLED(CONFIG_LOCKDEP) && (cond)) { \
--
include/linux/tracepoint.h-300- __DECLARE_TRACE_COMMON(name, PARAMS(proto), PARAMS(args), PARAMS(data_proto)) \
include/linux/tracepoint.h:301: static inline void __do_trace_##name(proto) \
include/linux/tracepoint.h-302- { \
--
include/linux/tracepoint.h-310- if (static_branch_unlikely(&__tracepoint_##name.key)) \
include/linux/tracepoint.h:311: __do_trace_##name(args); \
include/linux/tracepoint.h-312- if (IS_ENABLED(CONFIG_LOCKDEP)) { \
]
|
| 28/2 |
2026/05/17 11:40 |
llm |
expert |
0m
Model:
Tokens:
input: 17738
output: 41
thoughts: 21
Thoughts: Here's a summary of my thought process:
**Diving into Tracepoint Implementation**
Okay, I need to investigate something related to tracepoints. My immediate instinct is to jump into the core definition. Let's start with `include/linux/tracepoint.h`, specifically around line 281. That's usually where the heart of a structure's definition or a key macro resides, right? It's a good starting point for understanding how the basic tracepoint infrastructure is set up. From there, I can likely trace dependencies and follow the breadcrumbs as needed. This will give me a solid foundation for the next stage of my work.
|
| 29/2 |
2026/05/17 11:40 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:include/linux/tracepoint.h IncludeLines:true Name:__DECLARE_TRACE]
Results: map[SourceCode: SourceFile:]
|
| 30/2 |
2026/05/17 11:40 |
llm |
expert |
0m
Model:
Tokens:
input: 17959
output: 23
thoughts: 14
Thoughts: Okay, let's dive into this. I'm going to start by examining `include/linux/tracepoint.h`. This is where the core definitions for the Linux kernel's tracepoint infrastructure reside. My expectation is that this file will define the fundamental data structures and macros that allow developers to instrument their code with tracing points. I'm looking for the building blocks: what's the general layout? How are tracepoints declared, registered, and accessed? I'll be paying close attention to the mechanisms for data capture β how are arguments handled, what types are supported? And, crucially, how does the kernel manage the overhead associated with tracing? I want to see if the design prioritizes performance, and if so, how. Are there any clever techniques to minimize the impact on running code when tracepoints are disabled? I'll also be searching for any hints on how these tracepoints integrate with higher-level tracing frameworks like `ftrace` or `perf`. Essentially, I'm hoping this will illuminate the very foundation upon which more sophisticated tracing tools are built.
|
| 31/2 |
2026/05/17 11:40 |
tool |
read-file |
0m
Args: map[File:include/linux/tracepoint.h]
Results: map[Contents:/* SPDX-License-Identifier: GPL-2.0-only */
#ifndef _LINUX_TRACEPOINT_H
#define _LINUX_TRACEPOINT_H
/*
* Kernel Tracepoint API.
*
* See Documentation/trace/tracepoints.rst.
*
* Copyright (C) 2008-2014 Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
*
* Heavily inspired from the Linux Kernel Markers.
*/
#include <linux/smp.h>
#include <linux/srcu.h>
#include <linux/errno.h>
#include <linux/types.h>
#include <linux/rcupdate.h>
#include <linux/rcupdate_trace.h>
#include <linux/tracepoint-defs.h>
#include <linux/static_call.h>
struct module;
struct tracepoint;
struct notifier_block;
struct trace_eval_map {
const char *system;
const char *eval_string;
unsigned long eval_value;
};
#define TRACEPOINT_DEFAULT_PRIO 10
extern int
tracepoint_probe_register(struct tracepoint *tp, void *probe, void *data);
extern int
tracepoint_probe_register_prio(struct tracepoint *tp, void *probe, void *data,
int prio);
extern int
tracepoint_probe_register_prio_may_exist(struct tracepoint *tp, void *probe, void *data,
int prio);
extern int
tracepoint_probe_unregister(struct tracepoint *tp, void *probe, void *data);
static inline int
tracepoint_probe_register_may_exist(struct tracepoint *tp, void *probe,
void *data)
{
return tracepoint_probe_register_prio_may_exist(tp, probe, data,
TRACEPOINT_DEFAULT_PRIO);
}
extern void
for_each_kernel_tracepoint(void (*fct)(struct tracepoint *tp, void *priv),
void *priv);
#ifdef CONFIG_MODULES
struct tp_module {
struct list_head list;
struct module *mod;
};
bool trace_module_has_bad_taint(struct module *mod);
extern int register_tracepoint_module_notifier(struct notifier_block *nb);
extern int unregister_tracepoint_module_notifier(struct notifier_block *nb);
void for_each_module_tracepoint(void (*fct)(struct tracepoint *,
struct module *, void *),
void *priv);
void for_each_tracepoint_in_module(struct module *,
void (*fct)(struct tracepoint *,
struct module *, void *),
void *priv);
#else
static inline bool trace_module_has_bad_taint(struct module *mod)
{
return false;
}
static inline
int register_tracepoint_module_notifier(struct notifier_block *nb)
{
return 0;
}
static inline
int unregister_tracepoint_module_notifier(struct notifier_block *nb)
{
return 0;
}
static inline
void for_each_module_tracepoint(void (*fct)(struct tracepoint *,
struct module *, void *),
void *priv)
{
}
static inline
void for_each_tracepoint_in_module(struct module *mod,
void (*fct)(struct tracepoint *,
struct module *, void *),
void *priv)
{
}
#endif /* CONFIG_MODULES */
/*
* tracepoint_synchronize_unregister must be called between the last tracepoint
* probe unregistration and the end of module exit to make sure there is no
* caller executing a probe when it is freed.
*
* An alternative is to use the following for batch reclaim associated
* with a given tracepoint:
*
* - tracepoint_is_faultable() == false: call_srcu()
* - tracepoint_is_faultable() == true: call_rcu_tasks_trace()
*/
#ifdef CONFIG_TRACEPOINTS
extern struct srcu_struct tracepoint_srcu;
static inline void tracepoint_synchronize_unregister(void)
{
synchronize_rcu_tasks_trace();
synchronize_srcu(&tracepoint_srcu);
}
static inline bool tracepoint_is_faultable(struct tracepoint *tp)
{
return tp->ext && tp->ext->faultable;
}
#else
static inline void tracepoint_synchronize_unregister(void)
{ }
static inline bool tracepoint_is_faultable(struct tracepoint *tp)
{
return false;
}
#endif
#ifdef CONFIG_HAVE_SYSCALL_TRACEPOINTS
extern int syscall_regfunc(void);
extern void syscall_unregfunc(void);
#endif /* CONFIG_HAVE_SYSCALL_TRACEPOINTS */
#ifndef PARAMS
#define PARAMS(args...) args
#endif
#define TRACE_DEFINE_ENUM(x)
#define TRACE_DEFINE_SIZEOF(x)
#ifdef CONFIG_HAVE_ARCH_PREL32_RELOCATIONS
static inline struct tracepoint *tracepoint_ptr_deref(tracepoint_ptr_t *p)
{
return offset_to_ptr(p);
}
#define __TRACEPOINT_ENTRY(name) \
asm(" .section \"__tracepoints_ptrs\", \"a\" \n" \
" .balign 4 \n" \
" .long __tracepoint_" #name " - . \n" \
" .previous \n")
#else
static inline struct tracepoint *tracepoint_ptr_deref(tracepoint_ptr_t *p)
{
return *p;
}
#define __TRACEPOINT_ENTRY(name) \
static tracepoint_ptr_t __tracepoint_ptr_##name __used \
__section("__tracepoints_ptrs") = &__tracepoint_##name
#endif
#endif /* _LINUX_TRACEPOINT_H */
/*
* Note: we keep the TRACE_EVENT and DECLARE_TRACE outside the include
* file ifdef protection.
* This is due to the way trace events work. If a file includes two
* trace event headers under one "CREATE_TRACE_POINTS" the first include
* will override the TRACE_EVENT and break the second include.
*/
#ifndef DECLARE_TRACE
#define TP_PROTO(args...) args
#define TP_ARGS(args...) args
#define TP_CONDITION(args...) args
/*
* Individual subsystem my have a separate configuration to
* enable their tracepoints. By default, this file will create
* the tracepoints if CONFIG_TRACEPOINTS is defined. If a subsystem
* wants to be able to disable its tracepoints from being created
* it can define NOTRACE before including the tracepoint headers.
*/
#if defined(CONFIG_TRACEPOINTS) && !defined(NOTRACE)
#define TRACEPOINTS_ENABLED
#endif
#ifdef TRACEPOINTS_ENABLED
#ifdef CONFIG_HAVE_STATIC_CALL
#define __DO_TRACE_CALL(name, args) \
do { \
struct tracepoint_func *it_func_ptr; \
void *__data; \
it_func_ptr = \
rcu_dereference_raw((&__tracepoint_##name)->funcs); \
if (it_func_ptr) { \
__data = (it_func_ptr)->data; \
static_call(tp_func_##name)(__data, args); \
} \
} while (0)
#else
#define __DO_TRACE_CALL(name, args) __traceiter_##name(NULL, args)
#endif /* CONFIG_HAVE_STATIC_CALL */
/*
* Declare an exported function that Rust code can call to trigger this
* tracepoint. This function does not include the static branch; that is done
* in Rust to avoid a function call when the tracepoint is disabled.
*/
#define DEFINE_RUST_DO_TRACE(name, proto, args)
#define __DEFINE_RUST_DO_TRACE(name, proto, args) \
notrace void rust_do_trace_##name(proto) \
{ \
__do_trace_##name(args); \
}
/*
* When a tracepoint is used, it's name is added to the __tracepoint_check
* section. This section is only used at build time to make sure all
* defined tracepoints are used. It is discarded after the build.
*/
# define TRACEPOINT_CHECK(name) \
static const char __used __section("__tracepoint_check") \
__trace_check_##name[] = #name;
/*
* Make sure the alignment of the structure in the __tracepoints section will
* not add unwanted padding between the beginning of the section and the
* structure. Force alignment to the same alignment as the section start.
*
* When lockdep is enabled, we make sure to always test if RCU is
* "watching" regardless if the tracepoint is enabled or not. Tracepoints
* require RCU to be active, and it should always warn at the tracepoint
* site if it is not watching, as it will need to be active when the
* tracepoint is enabled.
*/
#define __DECLARE_TRACE_COMMON(name, proto, args, data_proto) \
extern int __traceiter_##name(data_proto); \
DECLARE_STATIC_CALL(tp_func_##name, __traceiter_##name); \
extern struct tracepoint __tracepoint_##name; \
extern void rust_do_trace_##name(proto); \
static inline int \
register_trace_##name(void (*probe)(data_proto), void *data) \
{ \
return tracepoint_probe_register(&__tracepoint_##name, \
(void *)probe, data); \
} \
static inline int \
register_trace_prio_##name(void (*probe)(data_proto), void *data,\
int prio) \
{ \
return tracepoint_probe_register_prio(&__tracepoint_##name, \
(void *)probe, data, prio); \
} \
static inline int \
unregister_trace_##name(void (*probe)(data_proto), void *data) \
{ \
return tracepoint_probe_unregister(&__tracepoint_##name,\
(void *)probe, data); \
} \
static inline void \
check_trace_callback_type_##name(void (*cb)(data_proto)) \
{ \
} \
static inline bool \
trace_##name##_enabled(void) \
{ \
return static_branch_unlikely(&__tracepoint_##name.key);\
}
#define __DECLARE_TRACE(name, proto, args, cond, data_proto) \
__DECLARE_TRACE_COMMON(name, PARAMS(proto), PARAMS(args), PARAMS(data_proto)) \
static inline void __do_trace_##name(proto) \
{ \
TRACEPOINT_CHECK(name) \
if (cond) { \
guard(srcu_fast_notrace)(&tracepoint_srcu); \
__DO_TRACE_CALL(name, TP_ARGS(args)); \
} \
} \
static inline void trace_##name(proto) \
{ \
if (static_branch_unlikely(&__tracepoint_##name.key)) \
__do_trace_##name(args); \
if (IS_ENABLED(CONFIG_LOCKDEP) && (cond)) { \
WARN_ONCE(!rcu_is_watching(), \
"RCU not watching for tracepoint"); \
} \
}
#define __DECLARE_TRACE_SYSCALL(name, proto, args, data_proto) \
__DECLARE_TRACE_COMMON(name, PARAMS(proto), PARAMS(args), PARAMS(data_proto)) \
static inline void __do_trace_##name(proto) \
{ \
TRACEPOINT_CHECK(name) \
guard(rcu_tasks_trace)(); \
__DO_TRACE_CALL(name, TP_ARGS(args)); \
} \
static inline void trace_##name(proto) \
{ \
might_fault(); \
if (static_branch_unlikely(&__tracepoint_##name.key)) \
__do_trace_##name(args); \
if (IS_ENABLED(CONFIG_LOCKDEP)) { \
WARN_ONCE(!rcu_is_watching(), \
"RCU not watching for tracepoint"); \
} \
}
/*
* We have no guarantee that gcc and the linker won't up-align the tracepoint
* structures, so we create an array of pointers that will be used for iteration
* on the tracepoints.
*
* it_func[0] is never NULL because there is at least one element in the array
* when the array itself is non NULL.
*/
#define __DEFINE_TRACE_EXT(_name, _ext, proto, args) \
static const char __tpstrtab_##_name[] \
__section("__tracepoints_strings") = #_name; \
extern struct static_call_key STATIC_CALL_KEY(tp_func_##_name); \
int __traceiter_##_name(void *__data, proto); \
void __probestub_##_name(void *__data, proto); \
struct tracepoint __tracepoint_##_name __used \
__section("__tracepoints") = { \
.name = __tpstrtab_##_name, \
.key = STATIC_KEY_FALSE_INIT, \
.static_call_key = &STATIC_CALL_KEY(tp_func_##_name), \
.static_call_tramp = STATIC_CALL_TRAMP_ADDR(tp_func_##_name), \
.iterator = &__traceiter_##_name, \
.probestub = &__probestub_##_name, \
.funcs = NULL, \
.ext = _ext, \
}; \
__TRACEPOINT_ENTRY(_name); \
int __traceiter_##_name(void *__data, proto) \
{ \
struct tracepoint_func *it_func_ptr; \
void *it_func; \
\
it_func_ptr = \
rcu_dereference_raw((&__tracepoint_##_name)->funcs); \
if (it_func_ptr) { \
do { \
it_func = READ_ONCE((it_func_ptr)->func); \
__data = (it_func_ptr)->data; \
((void(*)(void *, proto))(it_func))(__data, args); \
} while ((++it_func_ptr)->func); \
} \
return 0; \
} \
void __probestub_##_name(void *__data, proto) \
{ \
} \
DEFINE_STATIC_CALL(tp_func_##_name, __traceiter_##_name); \
DEFINE_RUST_DO_TRACE(_name, TP_PROTO(proto), TP_ARGS(args))
#define DEFINE_TRACE_FN(_name, _reg, _unreg, _proto, _args) \
static struct tracepoint_ext __tracepoint_ext_##_name = { \
.regfunc = _reg, \
.unregfunc = _unreg, \
.faultable = false, \
}; \
__DEFINE_TRACE_EXT(_name, &__tracepoint_ext_##_name, PARAMS(_proto), PARAMS(_args));
#define DEFINE_TRACE_SYSCALL(_name, _reg, _unreg, _proto, _args) \
static struct tracepoint_ext __tracepoint_ext_##_name = { \
.regfunc = _reg, \
.unregfunc = _unreg, \
.faultable = true, \
}; \
__DEFINE_TRACE_EXT(_name, &__tracepoint_ext_##_name, PARAMS(_proto), PARAMS(_args));
#define DEFINE_TRACE(_name, _proto, _args) \
__DEFINE_TRACE_EXT(_name, NULL, PARAMS(_proto), PARAMS(_args));
#define EXPORT_TRACEPOINT_SYMBOL_GPL(name) \
TRACEPOINT_CHECK(name) \
EXPORT_SYMBOL_GPL(__tracepoint_##name); \
EXPORT_SYMBOL_GPL(__traceiter_##name); \
EXPORT_STATIC_CALL_GPL(tp_func_##name)
#define EXPORT_TRACEPOINT_SYMBOL(name) \
TRACEPOINT_CHECK(name) \
EXPORT_SYMBOL(__tracepoint_##name); \
EXPORT_SYMBOL(__traceiter_##name); \
EXPORT_STATIC_CALL(tp_func_##name)
#else /* !TRACEPOINTS_ENABLED */
#define __DECLARE_TRACE_COMMON(name, proto, args, data_proto) \
static inline void trace_##name(proto) \
{ } \
static inline int \
register_trace_##name(void (*probe)(data_proto), \
void *data) \
{ \
return -ENOSYS; \
} \
static inline int \
unregister_trace_##name(void (*probe)(data_proto), \
void *data) \
{ \
return -ENOSYS; \
} \
static inline void check_trace_callback_type_##name(void (*cb)(data_proto)) \
{ \
} \
static inline bool \
trace_##name##_enabled(void) \
{ \
return false; \
}
#define __DECLARE_TRACE(name, proto, args, cond, data_proto) \
__DECLARE_TRACE_COMMON(name, PARAMS(proto), PARAMS(args), PARAMS(data_proto))
#define __DECLARE_TRACE_SYSCALL(name, proto, args, data_proto) \
__DECLARE_TRACE_COMMON(name, PARAMS(proto), PARAMS(args), PARAMS(data_proto))
#define DEFINE_TRACE_FN(name, reg, unreg, proto, args)
#define DEFINE_TRACE_SYSCALL(name, reg, unreg, proto, args)
#define DEFINE_TRACE(name, proto, args)
#define EXPORT_TRACEPOINT_SYMBOL_GPL(name)
#define EXPORT_TRACEPOINT_SYMBOL(name)
#endif /* TRACEPOINTS_ENABLED */
#ifdef CONFIG_TRACING
/**
* tracepoint_string - register constant persistent string to trace system
* @str - a constant persistent string that will be referenced in tracepoints
*
* If constant strings are being used in tracepoints, it is faster and
* more efficient to just save the pointer to the string and reference
* that with a printf "%s" instead of saving the string in the ring buffer
* and wasting space and time.
*
* The problem with the above approach is that userspace tools that read
* the binary output of the trace buffers do not have access to the string.
* Instead they just show the address of the string which is not very
* useful to users.
*
* With tracepoint_string(), the string will be registered to the tracing
* system and exported to userspace via the debugfs/tracing/printk_formats
* file that maps the string address to the string text. This way userspace
* tools that read the binary buffers have a way to map the pointers to
* the ASCII strings they represent.
*
* The @str used must be a constant string and persistent as it would not
* make sense to show a string that no longer exists. But it is still fine
* to be used with modules, because when modules are unloaded, if they
* had tracepoints, the ring buffers are cleared too. As long as the string
* does not change during the life of the module, it is fine to use
* tracepoint_string() within a module.
*/
#define tracepoint_string(str) \
({ \
static const char *___tp_str __tracepoint_string = str; \
___tp_str; \
})
#define __tracepoint_string __used __section("__tracepoint_str")
#else
/*
* tracepoint_string() is used to save the string address for userspace
* tracing tools. When tracing isn't configured, there's no need to save
* anything.
*/
# define tracepoint_string(str) str
# define __tracepoint_string
#endif
#define DECLARE_TRACE(name, proto, args) \
__DECLARE_TRACE(name##_tp, PARAMS(proto), PARAMS(args), \
cpu_online(raw_smp_processor_id()), \
PARAMS(void *__data, proto))
#define DECLARE_TRACE_CONDITION(name, proto, args, cond) \
__DECLARE_TRACE(name##_tp, PARAMS(proto), PARAMS(args), \
cpu_online(raw_smp_processor_id()) && (PARAMS(cond)), \
PARAMS(void *__data, proto))
#define DECLARE_TRACE_SYSCALL(name, proto, args) \
__DECLARE_TRACE_SYSCALL(name##_tp, PARAMS(proto), PARAMS(args), \
PARAMS(void *__data, proto))
#define DECLARE_TRACE_EVENT(name, proto, args) \
__DECLARE_TRACE(name, PARAMS(proto), PARAMS(args), \
cpu_online(raw_smp_processor_id()), \
PARAMS(void *__data, proto))
#define DECLARE_TRACE_EVENT_CONDITION(name, proto, args, cond) \
__DECLARE_TRACE(name, PARAMS(proto), PARAMS(args), \
cpu_online(raw_smp_processor_id()) && (PARAMS(cond)), \
PARAMS(void *__data, proto))
#define DECLARE_TRACE_EVENT_SYSCALL(name, proto, args) \
__DECLARE_TRACE_SYSCALL(name, PARAMS(proto), PARAMS(args), \
PARAMS(void *__data, proto))
#define TRACE_EVENT_FLAGS(event, flag)
#define TRACE_EVENT_PERF_PERM(event, expr...)
#endif /* DECLARE_TRACE */
#ifndef TRACE_EVENT
/*
* For use with the TRACE_EVENT macro:
*
* We define a tracepoint, its arguments, its printk format
* and its 'fast binary record' layout.
*
* Firstly, name your tracepoint via TRACE_EVENT(name : the
* 'subsystem_event' notation is fine.
*
* Think about this whole construct as the
* 'trace_sched_switch() function' from now on.
*
*
* TRACE_EVENT(sched_switch,
*
* *
* * A function has a regular function arguments
* * prototype, declare it via TP_PROTO():
* *
*
* TP_PROTO(struct rq *rq, struct task_struct *prev,
* struct task_struct *next),
*
* *
* * Define the call signature of the 'function'.
* * (Design sidenote: we use this instead of a
* * TP_PROTO1/TP_PROTO2/TP_PROTO3 ugliness.)
* *
*
* TP_ARGS(rq, prev, next),
*
* *
* * Fast binary tracing: define the trace record via
* * TP_STRUCT__entry(). You can think about it like a
* * regular C structure local variable definition.
* *
* * This is how the trace record is structured and will
* * be saved into the ring buffer. These are the fields
* * that will be exposed to user-space in
* * /sys/kernel/tracing/events/<*>/format.
* *
* * The declared 'local variable' is called '__entry'
* *
* * __field(pid_t, prev_pid) is equivalent to a standard declaration:
* *
* * pid_t prev_pid;
* *
* * __array(char, prev_comm, TASK_COMM_LEN) is equivalent to:
* *
* * char prev_comm[TASK_COMM_LEN];
* *
*
* TP_STRUCT__entry(
* __array( char, prev_comm, TASK_COMM_LEN )
* __field( pid_t, prev_pid )
* __field( int, prev_prio )
* __array( char, next_comm, TASK_COMM_LEN )
* __field( pid_t, next_pid )
* __field( int, next_prio )
* ),
*
* *
* * Assign the entry into the trace record, by embedding
* * a full C statement block into TP_fast_assign(). You
* * can refer to the trace record as '__entry' -
* * otherwise you can put arbitrary C code in here.
* *
* * Note: this C code will execute every time a trace event
* * happens, on an active tracepoint.
* *
*
* TP_fast_assign(
* memcpy(__entry->next_comm, next->comm, TASK_COMM_LEN);
* __entry->prev_pid = prev->pid;
* __entry->prev_prio = prev->prio;
* memcpy(__entry->prev_comm, prev->comm, TASK_COMM_LEN);
* __entry->next_pid = next->pid;
* __entry->next_prio = next->prio;
* ),
*
* *
* * Formatted output of a trace record via TP_printk().
* * This is how the tracepoint will appear under ftrace
* * plugins that make use of this tracepoint.
* *
* * (raw-binary tracing wont actually perform this step.)
* *
*
* TP_printk("task %s:%d [%d] ==> %s:%d [%d]",
* __entry->prev_comm, __entry->prev_pid, __entry->prev_prio,
* __entry->next_comm, __entry->next_pid, __entry->next_prio),
*
* );
*
* This macro construct is thus used for the regular printk format
* tracing setup, it is used to construct a function pointer based
* tracepoint callback (this is used by programmatic plugins and
* can also by used by generic instrumentation like SystemTap), and
* it is also used to expose a structured trace record in
* /sys/kernel/tracing/events/.
*
* A set of (un)registration functions can be passed to the variant
* TRACE_EVENT_FN to perform any (un)registration work.
*/
#define DECLARE_EVENT_CLASS(name, proto, args, tstruct, assign, print)
#define DEFINE_EVENT(template, name, proto, args) \
DECLARE_TRACE_EVENT(name, PARAMS(proto), PARAMS(args))
#define DEFINE_EVENT_FN(template, name, proto, args, reg, unreg)\
DECLARE_TRACE_EVENT(name, PARAMS(proto), PARAMS(args))
#define DEFINE_EVENT_PRINT(template, name, proto, args, print) \
DECLARE_TRACE_EVENT(name, PARAMS(proto), PARAMS(args))
#define DEFINE_EVENT_CONDITION(template, name, proto, \
args, cond) \
DECLARE_TRACE_EVENT_CONDITION(name, PARAMS(proto), \
PARAMS(args), PARAMS(cond))
#define TRACE_EVENT(name, proto, args, struct, assign, print) \
DECLARE_TRACE_EVENT(name, PARAMS(proto), PARAMS(args))
#define TRACE_EVENT_FN(name, proto, args, struct, \
assign, print, reg, unreg) \
DECLARE_TRACE_EVENT(name, PARAMS(proto), PARAMS(args))
#define TRACE_EVENT_FN_COND(name, proto, args, cond, struct, \
assign, print, reg, unreg) \
DECLARE_TRACE_EVENT_CONDITION(name, PARAMS(proto), \
PARAMS(args), PARAMS(cond))
#define TRACE_EVENT_CONDITION(name, proto, args, cond, \
struct, assign, print) \
DECLARE_TRACE_EVENT_CONDITION(name, PARAMS(proto), \
PARAMS(args), PARAMS(cond))
#define TRACE_EVENT_SYSCALL(name, proto, args, struct, assign, \
print, reg, unreg) \
DECLARE_TRACE_EVENT_SYSCALL(name, PARAMS(proto), PARAMS(args))
#define TRACE_EVENT_FLAGS(event, flag)
#define TRACE_EVENT_PERF_PERM(event, expr...)
#define DECLARE_EVENT_NOP(name, proto, args) \
static inline void trace_##name(proto) \
{ } \
static inline bool trace_##name##_enabled(void) \
{ \
return false; \
}
#define TRACE_EVENT_NOP(name, proto, args, struct, assign, print) \
DECLARE_EVENT_NOP(name, PARAMS(proto), PARAMS(args))
#define DECLARE_EVENT_CLASS_NOP(name, proto, args, tstruct, assign, print)
#define DEFINE_EVENT_NOP(template, name, proto, args) \
DECLARE_EVENT_NOP(name, PARAMS(proto), PARAMS(args))
#endif /* ifdef TRACE_EVENT (see note above) */
]
|
| 32/2 |
2026/05/17 11:40 |
llm |
expert |
0m
Model:
Tokens:
input: 25636
output: 40
thoughts: 584
Thoughts: **Investigating the Concurrency Implications of Tracepoints and RCU**
Okay, so I'm digging into how tracepoints are implemented, specifically in relation to concurrency and RCU. I'm looking at `include/linux/tracepoint.h` and the macro `__DECLARE_TRACE`. My immediate thought is, "Ah, `guard(srcu_fast_notrace)(&tracepoint_srcu)` is used here for protection." Then, I recalled that `tracepoint_is_faultable()` is involved in choosing between `tracepoint_srcu` and `rcu_tasks_trace`, I had to quickly reassess, no, it seems `__DECLARE_TRACE` consistently uses `srcu_fast_notrace` for all tracepoints defined with it. However, I know there's a variation of `__DECLARE_TRACE`, `__DECLARE_TRACE_SYSCALL` that utilizes `guard(rcu_tasks_trace)()`.
I took `trace_mm_page_alloc` as an example which is defined through `TRACE_EVENT`, which ultimately leverages `__DECLARE_TRACE`, meaning it also uses `srcu_fast_notrace` for its synchronization. This is where things get interesting. I started comparing, when considering the `bpf_link_free` function, which, depending on the `link->sleepable` flag, uses either `call_rcu_tasks_trace` or `call_rcu`. Specifically, if `link->sleepable` is false, it uses standard `call_rcu`.
This is where the alarm bells started ringing. If the tracepoint, like `trace_mm_page_alloc`, is using `srcu_fast_notrace` (which is SRCU) and then something else like `bpf_link_free` uses `call_rcu` (which is standard RCU), a race condition is possible! Specifically, if the tracepoint's SRCU read-side critical section overlaps with `call_rcu`, then `call_rcu` *could* complete before the SRCU read-side critical section is done. This means the RCU read side of `call_rcu` could see data that is in an inconsistent state, causing unexpected issues or data corruption, as the `srcu_fast_notrace` is used to guard. I need to investigate the use and potential conflicts related to `tracepoint_srcu`.
|
| 33/2 |
2026/05/17 11:41 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/tracepoint.c IncludeLines:true Name:tracepoint_srcu]
Results: map[SourceCode: 41: DEFINE_SRCU_FAST(tracepoint_srcu);
SourceFile:kernel/tracepoint.c]
|
| 34/2 |
2026/05/17 11:41 |
llm |
expert |
0m
Model:
Tokens:
input: 26825
output: 46
thoughts: 802
Thoughts: **Use-After-Free Vulnerability in Tracepoint Handling**
Okay, let's break this down. `tracepoint_srcu` is indeed an SRCU, and `tracepoint_synchronize_unregister` calls `synchronize_srcu`, as expected. However, the problem isn't there. `bpf_link_free` doesn't call `tracepoint_synchronize_unregister`; it just calls `call_rcu`, which is a key observation. The tracepoint's unregistration path is through `bpf_probe_unregister`, which calls `tracepoint_probe_unregister`. Critically, that function *doesn't* wait for SRCU synchronization. It just returns. Then `bpf_raw_tp_link_release` returns. Then `bpf_link_put_direct` calls `bpf_link_free`, and that calls `call_rcu` with `bpf_link_defer_dealloc_rcu_gp`. So, the `link` is freed after a *normal* RCU grace period.
The crucial point is that the tracepoint itself is protected by `tracepoint_srcu`! This means that a tracepoint *might* still be executing and accessing the `link` after the link is freed. I need to remember when the `guard(srcu_fast_notrace)` was added to `__DECLARE_TRACE` in `include/linux/tracepoint.h`. I'll need to check the `git log` or `git blame`. Actually, the bug report shows the `link` being freed by `rcu_do_batch` which is normal RCU.
The tracepoint is running in `__do_trace_mm_page_alloc`, and this is protected by SRCU. But, *if* `__do_trace_mm_page_alloc` is protected by SRCU, it can be preempted. This allows a *normal* RCU grace period to complete while the tracepoint is preempted! So, after a preempted tracepoint, `link` can be freed. The tracepoint then resumes, and accesses `link->link.prog`, which is a use-after-free, confirming the root cause!
Now, to be thorough, I should verify if `bpf_raw_tp_link_attach` requires any privileges. It's called from `bpf_raw_tracepoint_open` which is itself called from `__sys_bpf`. I'll need to check `bpf_raw_tracepoint_open` in `kernel/bpf/syscall.c` to be completely sure.
|
| 35/2 |
2026/05/17 11:41 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_raw_tracepoint_open]
Results: map[SourceCode:4296: static int bpf_raw_tracepoint_open(const union bpf_attr *attr)
4297: {
4298: struct bpf_prog *prog;
4299: void __user *tp_name;
4300: __u64 cookie;
4301: int fd;
4302:
4303: if (CHECK_ATTR(BPF_RAW_TRACEPOINT_OPEN))
4304: return -EINVAL;
4305:
4306: prog = bpf_prog_get(attr->raw_tracepoint.prog_fd);
4307: if (IS_ERR(prog))
4308: return PTR_ERR(prog);
4309:
4310: tp_name = u64_to_user_ptr(attr->raw_tracepoint.name);
4311: cookie = attr->raw_tracepoint.cookie;
4312: fd = bpf_raw_tp_link_attach(prog, tp_name, cookie, prog->expected_attach_type);
4313: if (fd < 0)
4314: bpf_prog_put(prog);
4315: return fd;
4316: }
SourceFile:kernel/bpf/syscall.c]
|
| 36/2 |
2026/05/17 11:41 |
llm |
expert |
0m
Model:
Tokens:
input: 28603
output: 42
thoughts: 27
Thoughts: Okay, here's my take on those thoughts, framed as an expert's internal monologue:
**Investigating `bpf_raw_tracepoint_open` Privilege Requirements**
Hmm, let's see, the question is whether `bpf_raw_tracepoint_open` needs elevated privileges to function. That's a good one, and the most logical place to start is the actual system call itself β `__sys_bpf`. I need to understand the underlying mechanism and the checks performed there, which will tell me if any privilege checks are happening. By examining `__sys_bpf`, I can trace the path of execution and the arguments passed to it. This will reveal the access control involved. That should give me a solid answer about `bpf_raw_tracepoint_open`.
|
| 37/2 |
2026/05/17 11:41 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:__sys_bpf]
Results: map[SourceCode:6189: static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size)
6190: {
6191: union bpf_attr attr;
6192: int err;
6193:
6194: err = bpf_check_uarg_tail_zero(uattr, sizeof(attr), size);
6195: if (err)
6196: return err;
6197: size = min_t(u32, size, sizeof(attr));
6198:
6199: /* copy attributes from user space, may be less than sizeof(bpf_attr) */
6200: memset(&attr, 0, sizeof(attr));
6201: if (copy_from_bpfptr(&attr, uattr, size) != 0)
6202: return -EFAULT;
6203:
6204: err = security_bpf(cmd, &attr, size, uattr.is_kernel);
6205: if (err < 0)
6206: return err;
6207:
6208: switch (cmd) {
6209: case BPF_MAP_CREATE:
6210: err = map_create(&attr, uattr);
6211: break;
6212: case BPF_MAP_LOOKUP_ELEM:
6213: err = map_lookup_elem(&attr);
6214: break;
6215: case BPF_MAP_UPDATE_ELEM:
6216: err = map_update_elem(&attr, uattr);
6217: break;
6218: case BPF_MAP_DELETE_ELEM:
6219: err = map_delete_elem(&attr, uattr);
6220: break;
6221: case BPF_MAP_GET_NEXT_KEY:
6222: err = map_get_next_key(&attr);
6223: break;
6224: case BPF_MAP_FREEZE:
6225: err = map_freeze(&attr);
6226: break;
6227: case BPF_PROG_LOAD:
6228: err = bpf_prog_load(&attr, uattr, size);
6229: break;
6230: case BPF_OBJ_PIN:
6231: err = bpf_obj_pin(&attr);
6232: break;
6233: case BPF_OBJ_GET:
6234: err = bpf_obj_get(&attr);
6235: break;
6236: case BPF_PROG_ATTACH:
6237: err = bpf_prog_attach(&attr);
6238: break;
6239: case BPF_PROG_DETACH:
6240: err = bpf_prog_detach(&attr);
6241: break;
6242: case BPF_PROG_QUERY:
6243: err = bpf_prog_query(&attr, uattr.user);
6244: break;
6245: case BPF_PROG_TEST_RUN:
6246: err = bpf_prog_test_run(&attr, uattr.user);
6247: break;
6248: case BPF_PROG_GET_NEXT_ID:
6249: err = bpf_obj_get_next_id(&attr, uattr.user,
6250: &prog_idr, &prog_idr_lock);
6251: break;
6252: case BPF_MAP_GET_NEXT_ID:
6253: err = bpf_obj_get_next_id(&attr, uattr.user,
6254: &map_idr, &map_idr_lock);
6255: break;
6256: case BPF_BTF_GET_NEXT_ID:
6257: err = bpf_obj_get_next_id(&attr, uattr.user,
6258: &btf_idr, &btf_idr_lock);
6259: break;
6260: case BPF_PROG_GET_FD_BY_ID:
6261: err = bpf_prog_get_fd_by_id(&attr);
6262: break;
6263: case BPF_MAP_GET_FD_BY_ID:
6264: err = bpf_map_get_fd_by_id(&attr);
6265: break;
6266: case BPF_OBJ_GET_INFO_BY_FD:
6267: err = bpf_obj_get_info_by_fd(&attr, uattr.user);
6268: break;
6269: case BPF_RAW_TRACEPOINT_OPEN:
6270: err = bpf_raw_tracepoint_open(&attr);
6271: break;
6272: case BPF_BTF_LOAD:
6273: err = bpf_btf_load(&attr, uattr, size);
6274: break;
6275: case BPF_BTF_GET_FD_BY_ID:
6276: err = bpf_btf_get_fd_by_id(&attr);
6277: break;
6278: case BPF_TASK_FD_QUERY:
6279: err = bpf_task_fd_query(&attr, uattr.user);
6280: break;
6281: case BPF_MAP_LOOKUP_AND_DELETE_ELEM:
6282: err = map_lookup_and_delete_elem(&attr);
6283: break;
6284: case BPF_MAP_LOOKUP_BATCH:
6285: err = bpf_map_do_batch(&attr, uattr.user, BPF_MAP_LOOKUP_BATCH);
6286: break;
6287: case BPF_MAP_LOOKUP_AND_DELETE_BATCH:
6288: err = bpf_map_do_batch(&attr, uattr.user,
6289: BPF_MAP_LOOKUP_AND_DELETE_BATCH);
6290: break;
6291: case BPF_MAP_UPDATE_BATCH:
6292: err = bpf_map_do_batch(&attr, uattr.user, BPF_MAP_UPDATE_BATCH);
6293: break;
6294: case BPF_MAP_DELETE_BATCH:
6295: err = bpf_map_do_batch(&attr, uattr.user, BPF_MAP_DELETE_BATCH);
6296: break;
6297: case BPF_LINK_CREATE:
6298: err = link_create(&attr, uattr);
6299: break;
6300: case BPF_LINK_UPDATE:
6301: err = link_update(&attr);
6302: break;
6303: case BPF_LINK_GET_FD_BY_ID:
6304: err = bpf_link_get_fd_by_id(&attr);
6305: break;
6306: case BPF_LINK_GET_NEXT_ID:
6307: err = bpf_obj_get_next_id(&attr, uattr.user,
6308: &link_idr, &link_idr_lock);
6309: break;
6310: case BPF_ENABLE_STATS:
6311: err = bpf_enable_stats(&attr);
6312: break;
6313: case BPF_ITER_CREATE:
6314: err = bpf_iter_create(&attr);
6315: break;
6316: case BPF_LINK_DETACH:
6317: err = link_detach(&attr);
6318: break;
6319: case BPF_PROG_BIND_MAP:
6320: err = bpf_prog_bind_map(&attr);
6321: break;
6322: case BPF_TOKEN_CREATE:
6323: err = token_create(&attr);
6324: break;
6325: case BPF_PROG_STREAM_READ_BY_FD:
6326: err = prog_stream_read(&attr);
6327: break;
6328: case BPF_PROG_ASSOC_STRUCT_OPS:
6329: err = prog_assoc_struct_ops(&attr);
6330: break;
6331: default:
6332: err = -EINVAL;
6333: break;
6334: }
6335:
6336: return err;
6337: }
SourceFile:kernel/bpf/syscall.c]
|
| 38/2 |
2026/05/17 11:41 |
llm |
expert |
0m
Model:
Tokens:
input: 31278
output: 43
thoughts: 128
Thoughts: **Investigating BPF Raw Tracepoint Security**
Okay, so I'm wondering if `bpf_raw_tracepoint_open` has any privilege restrictions, or if it's accessible to regular users. The obvious next step is to examine the kernel code. Let's look at `bpf_prog_get` and `bpf_raw_tracepoint_open` directly, specifically checking for any use of `capable()`. Hmmm, it looks like `bpf_raw_tracepoint_open` *doesn't* have an explicit `capable` check. That's unexpected, potentially a security oversight. But wait... `bpf_prog_get` is for retrieving *existing* programs. Ah, it's starting to click. I bet the real restriction lies in the *loading* of the program itself. So, to load a BPF program of the type `BPF_PROG_TYPE_RAW_TRACEPOINT`, you almost certainly need elevated privileges. I need to confirm this by looking at the `bpf_prog_load` function. That's the key.
|
| 39/2 |
2026/05/17 11:41 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_prog_load]
Results: map[SourceCode:2871: static int bpf_prog_load(union bpf_attr *attr, bpfptr_t uattr, u32 uattr_size)
2872: {
2873: enum bpf_prog_type type = attr->prog_type;
2874: struct bpf_prog *prog, *dst_prog = NULL;
2875: struct btf *attach_btf = NULL;
2876: struct bpf_token *token = NULL;
2877: bool bpf_cap;
2878: int err;
2879: char license[128];
2880:
2881: if (CHECK_ATTR(BPF_PROG_LOAD))
2882: return -EINVAL;
2883:
2884: if (attr->prog_flags & ~(BPF_F_STRICT_ALIGNMENT |
2885: BPF_F_ANY_ALIGNMENT |
2886: BPF_F_TEST_STATE_FREQ |
2887: BPF_F_SLEEPABLE |
2888: BPF_F_TEST_RND_HI32 |
2889: BPF_F_XDP_HAS_FRAGS |
2890: BPF_F_XDP_DEV_BOUND_ONLY |
2891: BPF_F_TEST_REG_INVARIANTS |
2892: BPF_F_TOKEN_FD))
2893: return -EINVAL;
2894:
2895: bpf_prog_load_fixup_attach_type(attr);
2896:
2897: if (attr->prog_flags & BPF_F_TOKEN_FD) {
2898: token = bpf_token_get_from_fd(attr->prog_token_fd);
2899: if (IS_ERR(token))
2900: return PTR_ERR(token);
2901: /* if current token doesn't grant prog loading permissions,
2902: * then we can't use this token, so ignore it and rely on
2903: * system-wide capabilities checks
2904: */
2905: if (!bpf_token_allow_cmd(token, BPF_PROG_LOAD) ||
2906: !bpf_token_allow_prog_type(token, attr->prog_type,
2907: attr->expected_attach_type)) {
2908: bpf_token_put(token);
2909: token = NULL;
2910: }
2911: }
2912:
2913: bpf_cap = bpf_token_capable(token, CAP_BPF);
2914: err = -EPERM;
2915:
2916: if (!IS_ENABLED(CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS) &&
2917: (attr->prog_flags & BPF_F_ANY_ALIGNMENT) &&
2918: !bpf_cap)
2919: goto put_token;
2920:
2921: /* Intent here is for unprivileged_bpf_disabled to block BPF program
2922: * creation for unprivileged users; other actions depend
2923: * on fd availability and access to bpffs, so are dependent on
2924: * object creation success. Even with unprivileged BPF disabled,
2925: * capability checks are still carried out for these
2926: * and other operations.
2927: */
2928: if (sysctl_unprivileged_bpf_disabled && !bpf_cap)
2929: goto put_token;
2930:
2931: if (attr->insn_cnt == 0 ||
2932: attr->insn_cnt > (bpf_cap ? BPF_COMPLEXITY_LIMIT_INSNS : BPF_MAXINSNS)) {
2933: err = -E2BIG;
2934: goto put_token;
2935: }
2936: if (type != BPF_PROG_TYPE_SOCKET_FILTER &&
2937: type != BPF_PROG_TYPE_CGROUP_SKB &&
2938: !bpf_cap)
2939: goto put_token;
2940:
2941: if (is_net_admin_prog_type(type) && !bpf_token_capable(token, CAP_NET_ADMIN))
2942: goto put_token;
2943: if (is_perfmon_prog_type(type) && !bpf_token_capable(token, CAP_PERFMON))
2944: goto put_token;
2945:
2946: /* attach_prog_fd/attach_btf_obj_fd can specify fd of either bpf_prog
2947: * or btf, we need to check which one it is
2948: */
2949: if (attr->attach_prog_fd) {
2950: dst_prog = bpf_prog_get(attr->attach_prog_fd);
2951: if (IS_ERR(dst_prog)) {
2952: dst_prog = NULL;
2953: attach_btf = btf_get_by_fd(attr->attach_btf_obj_fd);
2954: if (IS_ERR(attach_btf)) {
2955: err = -EINVAL;
2956: goto put_token;
2957: }
2958: if (!btf_is_kernel(attach_btf)) {
2959: /* attaching through specifying bpf_prog's BTF
2960: * objects directly might be supported eventually
2961: */
2962: btf_put(attach_btf);
2963: err = -ENOTSUPP;
2964: goto put_token;
2965: }
2966: }
2967: } else if (attr->attach_btf_id) {
2968: /* fall back to vmlinux BTF, if BTF type ID is specified */
2969: attach_btf = bpf_get_btf_vmlinux();
2970: if (IS_ERR(attach_btf)) {
2971: err = PTR_ERR(attach_btf);
2972: goto put_token;
2973: }
2974: if (!attach_btf) {
2975: err = -EINVAL;
2976: goto put_token;
2977: }
2978: btf_get(attach_btf);
2979: }
2980:
2981: if (bpf_prog_load_check_attach(type, attr->expected_attach_type,
2982: attach_btf, attr->attach_btf_id,
2983: dst_prog)) {
2984: if (dst_prog)
2985: bpf_prog_put(dst_prog);
2986: if (attach_btf)
2987: btf_put(attach_btf);
2988: err = -EINVAL;
2989: goto put_token;
2990: }
2991:
2992: /* plain bpf_prog allocation */
2993: prog = bpf_prog_alloc(bpf_prog_size(attr->insn_cnt), GFP_USER);
2994: if (!prog) {
2995: if (dst_prog)
2996: bpf_prog_put(dst_prog);
2997: if (attach_btf)
2998: btf_put(attach_btf);
2999: err = -EINVAL;
3000: goto put_token;
3001: }
3002:
3003: prog->expected_attach_type = attr->expected_attach_type;
3004: prog->sleepable = !!(attr->prog_flags & BPF_F_SLEEPABLE);
3005: prog->aux->attach_btf = attach_btf;
3006: prog->aux->attach_btf_id = attr->attach_btf_id;
3007: prog->aux->dst_prog = dst_prog;
3008: prog->aux->dev_bound = !!attr->prog_ifindex;
3009: prog->aux->xdp_has_frags = attr->prog_flags & BPF_F_XDP_HAS_FRAGS;
3010:
3011: /* move token into prog->aux, reuse taken refcnt */
3012: prog->aux->token = token;
3013: token = NULL;
3014:
3015: prog->aux->user = get_current_user();
3016: prog->len = attr->insn_cnt;
3017:
3018: err = -EFAULT;
3019: if (copy_from_bpfptr(prog->insns,
3020: make_bpfptr(attr->insns, uattr.is_kernel),
3021: bpf_prog_insn_size(prog)) != 0)
3022: goto free_prog;
3023: /* copy eBPF program license from user space */
3024: if (strncpy_from_bpfptr(license,
3025: make_bpfptr(attr->license, uattr.is_kernel),
3026: sizeof(license) - 1) < 0)
3027: goto free_prog;
3028: license[sizeof(license) - 1] = 0;
3029:
3030: /* eBPF programs must be GPL compatible to use GPL-ed functions */
3031: prog->gpl_compatible = license_is_gpl_compatible(license) ? 1 : 0;
3032:
3033: if (attr->signature) {
3034: err = bpf_prog_verify_signature(prog, attr, uattr.is_kernel);
3035: if (err)
3036: goto free_prog;
3037: }
3038:
3039: prog->orig_prog = NULL;
3040: prog->jited = 0;
3041:
3042: atomic64_set(&prog->aux->refcnt, 1);
3043:
3044: if (bpf_prog_is_dev_bound(prog->aux)) {
3045: err = bpf_prog_dev_bound_init(prog, attr);
3046: if (err)
3047: goto free_prog;
3048: }
3049:
3050: if (type == BPF_PROG_TYPE_EXT && dst_prog &&
3051: bpf_prog_is_dev_bound(dst_prog->aux)) {
3052: err = bpf_prog_dev_bound_inherit(prog, dst_prog);
3053: if (err)
3054: goto free_prog;
3055: }
3056:
3057: /*
3058: * Bookkeeping for managing the program attachment chain.
3059: *
3060: * It might be tempting to set attach_tracing_prog flag at the attachment
3061: * time, but this will not prevent from loading bunch of tracing prog
3062: * first, then attach them one to another.
3063: *
3064: * The flag attach_tracing_prog is set for the whole program lifecycle, and
3065: * doesn't have to be cleared in bpf_tracing_link_release, since tracing
3066: * programs cannot change attachment target.
3067: */
3068: if (type == BPF_PROG_TYPE_TRACING && dst_prog &&
3069: dst_prog->type == BPF_PROG_TYPE_TRACING) {
3070: prog->aux->attach_tracing_prog = true;
3071: }
3072:
3073: /* find program type: socket_filter vs tracing_filter */
3074: err = find_prog_type(type, prog);
3075: if (err < 0)
3076: goto free_prog;
3077:
3078: prog->aux->load_time = ktime_get_boottime_ns();
3079: err = bpf_obj_name_cpy(prog->aux->name, attr->prog_name,
3080: sizeof(attr->prog_name));
3081: if (err < 0)
3082: goto free_prog;
3083:
3084: err = security_bpf_prog_load(prog, attr, token, uattr.is_kernel);
3085: if (err)
3086: goto free_prog_sec;
3087:
3088: /* run eBPF verifier */
3089: err = bpf_check(&prog, attr, uattr, uattr_size);
3090: if (err < 0)
3091: goto free_used_maps;
3092:
3093: prog = bpf_prog_select_runtime(prog, &err);
3094: if (err < 0)
3095: goto free_used_maps;
3096:
3097: err = bpf_prog_mark_insn_arrays_ready(prog);
3098: if (err < 0)
3099: goto free_used_maps;
3100:
3101: err = bpf_prog_alloc_id(prog);
3102: if (err)
3103: goto free_used_maps;
3104:
3105: /* Upon success of bpf_prog_alloc_id(), the BPF prog is
3106: * effectively publicly exposed. However, retrieving via
3107: * bpf_prog_get_fd_by_id() will take another reference,
3108: * therefore it cannot be gone underneath us.
3109: *
3110: * Only for the time /after/ successful bpf_prog_new_fd()
3111: * and before returning to userspace, we might just hold
3112: * one reference and any parallel close on that fd could
3113: * rip everything out. Hence, below notifications must
3114: * happen before bpf_prog_new_fd().
3115: *
3116: * Also, any failure handling from this point onwards must
3117: * be using bpf_prog_put() given the program is exposed.
3118: */
3119: bpf_prog_kallsyms_add(prog);
3120: perf_event_bpf_event(prog, PERF_BPF_EVENT_PROG_LOAD, 0);
3121: bpf_audit_prog(prog, BPF_AUDIT_LOAD);
3122:
3123: err = bpf_prog_new_fd(prog);
3124: if (err < 0)
3125: bpf_prog_put(prog);
3126: return err;
3127:
3128: free_used_maps:
3129: /* In case we have subprogs, we need to wait for a grace
3130: * period before we can tear down JIT memory since symbols
3131: * are already exposed under kallsyms.
3132: */
3133: __bpf_prog_put_noref(prog, prog->aux->real_func_cnt);
3134: return err;
3135:
3136: free_prog_sec:
3137: security_bpf_prog_free(prog);
3138: free_prog:
3139: free_uid(prog->aux->user);
3140: if (prog->aux->attach_btf)
3141: btf_put(prog->aux->attach_btf);
3142: bpf_prog_free(prog);
3143: put_token:
3144: bpf_token_put(token);
3145: return err;
3146: }
SourceFile:kernel/bpf/syscall.c]
|
| 40/2 |
2026/05/17 11:41 |
llm |
expert |
0m
Model:
Tokens:
input: 36226
output: 43
thoughts: 294
Thoughts: **Vulnerability Exploration in BPF Program Loading**
Okay, so I'm digging into this `bpf_prog_load` function, and immediately I see this check related to `bpf_cap`. It seems like certain BPF program types require the `CAP_BPF` capability. Specifically, if the program type isn't `BPF_PROG_TYPE_SOCKET_FILTER` or `BPF_PROG_TYPE_CGROUP_SKB`, and `bpf_cap` is false, it jumps out. Now, `bpf_cap` is just `bpf_token_capable(token, CAP_BPF)`, so it checks if the token has `CAP_BPF`. This means something like `BPF_PROG_TYPE_RAW_TRACEPOINT` needs that capability.
But wait a minute, how does this interact with `bpf_raw_tracepoint_open`? Isn't there a potential bypass here? Can someone somehow call `bpf_raw_tracepoint_open` without having `CAP_BPF`? Let me re-examine the `bpf_raw_tracepoint_open` function call. Aha! It calls `bpf_prog_get(attr->raw_tracepoint.prog_fd)`.
This is interesting. So, if a privileged user *loads* the BPF program, giving it `CAP_BPF`, and then passes the file descriptor (`prog_fd`) to an *unprivileged* user, the unprivileged user *could* potentially call `bpf_raw_tracepoint_open`. My initial question is if `bpf_raw_tracepoint_open` itself does a `capable(CAP_BPF)` check? No, it doesn't. But, perhaps `bpf_prog_get` does permission checks? Let me go and check what `bpf_prog_get` does.
|
| 41/2 |
2026/05/17 11:41 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_prog_get]
Results: map[SourceCode:2609: struct bpf_prog *bpf_prog_get(u32 ufd)
2610: {
2611: return __bpf_prog_get(ufd, NULL, false);
2612: }
SourceFile:kernel/bpf/syscall.c]
|
| 42/2 |
2026/05/17 11:41 |
llm |
expert |
0m
Model:
Tokens:
input: 37068
output: 44
thoughts: 13
Thoughts: Here's my take on those thoughts, tailored for an expert audience:
**Diving into `__bpf_prog_get`**
Okay, time to get my hands dirty. My mind immediately zeroes in on `__bpf_prog_get`. I'm already mentally walking through the potential scenarios where this function comes into play. Is it related to the core BPF program structure retrieval? Is it involved in reference counting? Security checks? Memory management? I need to understand the exact context to properly assess its significance. I'm thinking about the implications for program loading, unloading, and lifecycle management within the BPF subsystem. I want to see how this interacts with the `bpf_prog` structure, any potential locks, and any associated error handling. The nuances of its implementation will be crucial for understanding performance bottlenecks and security vulnerabilities. Let's see how it fits within the larger picture of BPF program management within the kernel. I need to understand this to have a solid grasp on this function.
|
| 43/2 |
2026/05/17 11:41 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:__bpf_prog_get]
Results: map[SourceCode:2590: static struct bpf_prog *__bpf_prog_get(u32 ufd, enum bpf_prog_type *attach_type,
2591: bool attach_drv)
2592: {
2593: CLASS(fd, f)(ufd);
2594: struct bpf_prog *prog;
2595:
2596: if (fd_empty(f))
2597: return ERR_PTR(-EBADF);
2598: if (fd_file(f)->f_op != &bpf_prog_fops)
2599: return ERR_PTR(-EINVAL);
2600:
2601: prog = fd_file(f)->private_data;
2602: if (!bpf_prog_get_ok(prog, attach_type, attach_drv))
2603: return ERR_PTR(-EINVAL);
2604:
2605: bpf_prog_inc(prog);
2606: return prog;
2607: }
SourceFile:kernel/bpf/syscall.c]
|
| 44/2 |
2026/05/17 11:41 |
llm |
expert |
0m
Model:
Tokens:
input: 37649
output: 47
thoughts: 104
Thoughts: **Investigating Potential BPF Security Vulnerability**
Okay, so I'm digging into a potential security hole related to BPF and raw tracepoints. It looks like `bpf_raw_tracepoint_open` doesn't enforce the `CAP_BPF` capability check. That's immediately interesting. However, I know an unprivileged user can't *directly* load a `BPF_PROG_TYPE_RAW_TRACEPOINT` program. The question that's nagging at me is: can someone circumvent this restriction? Can they maybe load a `BPF_PROG_TYPE_SOCKET_FILTER` program, and then attach it somehow using `bpf_raw_tracepoint_open`? That would be a serious problem. It could open a new avenue for an unprivileged process to inject BPF code, which would be bypassing the intended security boundaries. I need to get into `bpf_raw_tp_link_attach` to confirm whether this possibility exists or not. Time to dive deeper into the code.
|
| 45/2 |
2026/05/17 11:41 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_raw_tp_link_attach]
Results: map[SourceCode:4223: static int bpf_raw_tp_link_attach(struct bpf_prog *prog,
4224: const char __user *user_tp_name, u64 cookie,
4225: enum bpf_attach_type attach_type)
4226: {
4227: struct bpf_link_primer link_primer;
4228: struct bpf_raw_tp_link *link;
4229: struct bpf_raw_event_map *btp;
4230: const char *tp_name;
4231: char buf[128];
4232: int err;
4233:
4234: switch (prog->type) {
4235: case BPF_PROG_TYPE_TRACING:
4236: case BPF_PROG_TYPE_EXT:
4237: case BPF_PROG_TYPE_LSM:
4238: if (user_tp_name)
4239: /* The attach point for this category of programs
4240: * should be specified via btf_id during program load.
4241: */
4242: return -EINVAL;
4243: if (prog->type == BPF_PROG_TYPE_TRACING &&
4244: prog->expected_attach_type == BPF_TRACE_RAW_TP) {
4245: tp_name = prog->aux->attach_func_name;
4246: break;
4247: }
4248: return bpf_tracing_prog_attach(prog, 0, 0, 0, attach_type);
4249: case BPF_PROG_TYPE_RAW_TRACEPOINT:
4250: case BPF_PROG_TYPE_RAW_TRACEPOINT_WRITABLE:
4251: if (strncpy_from_user(buf, user_tp_name, sizeof(buf) - 1) < 0)
4252: return -EFAULT;
4253: buf[sizeof(buf) - 1] = 0;
4254: tp_name = buf;
4255: break;
4256: default:
4257: return -EINVAL;
4258: }
4259:
4260: btp = bpf_get_raw_tracepoint(tp_name);
4261: if (!btp)
4262: return -ENOENT;
4263:
4264: link = kzalloc_obj(*link, GFP_USER);
4265: if (!link) {
4266: err = -ENOMEM;
4267: goto out_put_btp;
4268: }
4269: bpf_link_init_sleepable(&link->link, BPF_LINK_TYPE_RAW_TRACEPOINT,
4270: &bpf_raw_tp_link_lops, prog, attach_type,
4271: tracepoint_is_faultable(btp->tp));
4272: link->btp = btp;
4273: link->cookie = cookie;
4274:
4275: err = bpf_link_prime(&link->link, &link_primer);
4276: if (err) {
4277: kfree(link);
4278: goto out_put_btp;
4279: }
4280:
4281: err = bpf_probe_register(link->btp, link);
4282: if (err) {
4283: bpf_link_cleanup(&link_primer);
4284: goto out_put_btp;
4285: }
4286:
4287: return bpf_link_settle(&link_primer);
4288:
4289: out_put_btp:
4290: bpf_put_raw_tracepoint(btp);
4291: return err;
4292: }
SourceFile:kernel/bpf/syscall.c]
|
| 46/2 |
2026/05/17 11:41 |
llm |
expert |
0m
Model:
Tokens:
input: 39144
output: 43
thoughts: 332
Thoughts: **My BPF Security Analysis**
Okay, so I'm digging into `bpf_raw_tp_link_attach` and it's looking like this function is pretty locked down. It only plays nice with `BPF_PROG_TYPE_TRACING`, `BPF_PROG_TYPE_EXT`, `BPF_PROG_TYPE_LSM`, `BPF_PROG_TYPE_RAW_TRACEPOINT`, and `BPF_PROG_TYPE_RAW_TRACEPOINT_WRITABLE` programs. And all of those, critically, demand `CAP_BPF` or `CAP_SYS_ADMIN` to even *load* them. Hang on... `BPF_PROG_TYPE_TRACING` *also* requires `CAP_BPF`? I need to double-check that `bpf_prog_load` logic.
Right, let me pull up the code for `bpf_prog_load`. Ah, here it is:
```c
if (type != BPF_PROG_TYPE_SOCKET_FILTER &&
type != BPF_PROG_TYPE_CGROUP_SKB &&
!bpf_cap)
goto put_token;
```
Yup, it's confirmed. Unless you're dealing with `SOCKET_FILTER` or `CGROUP_SKB` programs, you *need* `CAP_BPF` (assuming, of course, that `unprivileged_bpf_disabled` is set to 0. We're assuming it's 1 here, though, so this is critical.) Since we're assuming that `unprivileged_bpf_disabled=1`, that means absolutely *no* BPF programs can be loaded by unprivileged users. This is not going to be accessible from an unprivileged process. It's a hard barrier.
Now, what about user namespaces? I need to remember how capabilities interact with those. The capability check for `CAP_BPF` has to be against the initial user namespace. And I should probably look at what `bpf_token_capable` does, too. It's time to check what `bpf_token_capable` does.
|
| 47/2 |
2026/05/17 11:41 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/token.c IncludeLines:true Name:bpf_token_capable]
Results: map[SourceCode: 17: bool bpf_token_capable(const struct bpf_token *token, int cap)
18: {
19: struct user_namespace *userns;
20:
21: /* BPF token allows ns_capable() level of capabilities */
22: userns = token ? token->userns : &init_user_ns;
23: if (!bpf_ns_capable(userns, cap))
24: return false;
25: if (token && security_bpf_token_capable(token, cap) < 0)
26: return false;
27: return true;
28: }
SourceFile:kernel/bpf/token.c]
|
| 48/2 |
2026/05/17 11:41 |
llm |
expert |
0m
Model:
Tokens:
input: 40211
output: 43
thoughts: 65
Thoughts: Here's my analysis of the situation:
**BPF Token Security Conundrum**
Okay, so the `bpf_token_capable` function is the gatekeeper here. I'm seeing it leverages `token->userns` when a token is actually *present*, which is the expected behavior. But, what happens if no token is provided? Aha, it defaults to `init_user_ns`. That's a critical point!
This opens up a pretty significant avenue for privilege escalation. If a user crafts a BPF token within their own user namespace, *they should* then be able to load BPF programs, which is the functionality we have been looking for.
I need to confirm this understanding, so let's deep dive into `bpf_token_create` and see precisely how these tokens are being generated and handled. The creation process is what matters most here. This is where the root cause probably lays.
|
| 49/2 |
2026/05/17 11:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/token.c IncludeLines:true Name:bpf_token_create]
Results: map[SourceCode: 112: int bpf_token_create(union bpf_attr *attr)
113: {
114: struct bpf_token *token __free(kfree) = NULL;
115: struct bpf_mount_opts *mnt_opts;
116: struct user_namespace *userns;
117: struct inode *inode;
118: CLASS(fd, f)(attr->token_create.bpffs_fd);
119: struct path path;
120: struct super_block *sb;
121: umode_t mode;
122: int err;
123:
124: if (fd_empty(f))
125: return -EBADF;
126:
127: path = fd_file(f)->f_path;
128: sb = path.dentry->d_sb;
129:
130: if (path.dentry != sb->s_root)
131: return -EINVAL;
132: if (sb->s_op != &bpf_super_ops)
133: return -EINVAL;
134: err = path_permission(&path, MAY_ACCESS);
135: if (err)
136: return err;
137:
138: userns = sb->s_user_ns;
139: /*
140: * Enforce that creators of BPF tokens are in the same user
141: * namespace as the BPF FS instance. This makes reasoning about
142: * permissions a lot easier and we can always relax this later.
143: */
144: if (current_user_ns() != userns)
145: return -EPERM;
146: if (!ns_capable(userns, CAP_BPF))
147: return -EPERM;
148:
149: /* Creating BPF token in init_user_ns doesn't make much sense. */
150: if (current_user_ns() == &init_user_ns)
151: return -EOPNOTSUPP;
152:
153: mnt_opts = sb->s_fs_info;
154: if (mnt_opts->delegate_cmds == 0 &&
155: mnt_opts->delegate_maps == 0 &&
156: mnt_opts->delegate_progs == 0 &&
157: mnt_opts->delegate_attachs == 0)
158: return -ENOENT; /* no BPF token delegation is set up */
159:
160: mode = S_IFREG | ((S_IRUSR | S_IWUSR) & ~current_umask());
161: inode = bpf_get_inode(sb, NULL, mode);
162: if (IS_ERR(inode))
163: return PTR_ERR(inode);
164:
165: inode->i_op = &bpf_token_iops;
166: inode->i_fop = &bpf_token_fops;
167: clear_nlink(inode); /* make sure it is unlinked */
168:
169: FD_PREPARE(fdf, O_CLOEXEC,
170: alloc_file_pseudo(inode, path.mnt, BPF_TOKEN_INODE_NAME,
171: O_RDWR, &bpf_token_fops));
172: if (fdf.err)
173: return fdf.err;
174:
175: token = kzalloc_obj(*token, GFP_USER);
176: if (!token)
177: return -ENOMEM;
178:
179: atomic64_set(&token->refcnt, 1);
180:
181: /* remember bpffs owning userns for future ns_capable() checks. */
182: token->userns = userns;
183: token->allowed_cmds = mnt_opts->delegate_cmds;
184: token->allowed_maps = mnt_opts->delegate_maps;
185: token->allowed_progs = mnt_opts->delegate_progs;
186: token->allowed_attachs = mnt_opts->delegate_attachs;
187:
188: err = security_bpf_token_create(token, attr, &path);
189: if (err)
190: return err;
191:
192: get_user_ns(token->userns);
193: fd_prepare_file(fdf)->private_data = no_free_ptr(token);
194: return fd_publish(fdf);
195: }
SourceFile:kernel/bpf/token.c]
|
| 50/2 |
2026/05/17 11:42 |
llm |
expert |
0m
Model:
Tokens:
input: 41897
output: 41
thoughts: 61
Thoughts: **My BPF Token Creation Thoughts**
Okay, so I need to get this token creation process streamlined. The documentation says I need to mount `bpffs`, but I also see this delegation stuff β `delegate_cmds`, `delegate_progs`, all that. Hmm, interesting. The real question is, can an unprivileged user even *do* this mount in a user namespace? I'm pretty sure I remember there being some nuances with BPF and delegation related to namespace capabilities. I need to be *absolutely* certain about this for security and usability reasons.
Right, let's dive into those `bpffs` mount options. Time to verify what's allowed. I need to refresh my memory on what options are available and, more importantly, whether there's any way to restrict or fine-tune user capabilities for this mount point. I'll need to double check any potential security implications around allowing this, but if I can get this working within a user namespace it'll be a massive win.
|
| 51/2 |
2026/05/17 11:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/inode.c IncludeLines:true Name:bpf_init]
Results: map[SourceCode:1083: static int __init bpf_init(void)
1084: {
1085: int ret;
1086:
1087: ret = sysfs_create_mount_point(fs_kobj, "bpf");
1088: if (ret)
1089: return ret;
1090:
1091: ret = register_filesystem(&bpf_fs_type);
1092: if (ret)
1093: sysfs_remove_mount_point(fs_kobj, "bpf");
1094:
1095: return ret;
1096: }
SourceFile:kernel/bpf/inode.c]
|
| 52/2 |
2026/05/17 11:42 |
llm |
expert |
0m
Model:
Tokens:
input: 42422
output: 43
thoughts: 12
Thoughts: Okay, here's my attempt at summarizing that thought process, keeping in mind I'm speaking to a fellow expert:
**Investigating `bpf_fs_type`**
Alright, time to dive in. *bpf_fs_type*... that rings a bell. My initial thought is to quickly verify that, I need to check `bpf_fs_type` and see exactly what's going on there. I'm wondering if there are any specific aspects related to that variable I need to consider for this particular analysis. I'll need to go beyond a cursory glance, and really understand what's happening with it.
|
| 53/2 |
2026/05/17 11:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/inode.c IncludeLines:true Name:bpf_fs_type]
Results: map[SourceCode:1074: static struct file_system_type bpf_fs_type = {
1075: .owner = THIS_MODULE,
1076: .name = "bpf",
1077: .init_fs_context = bpf_init_fs_context,
1078: .parameters = bpf_fs_parameters,
1079: .kill_sb = bpf_kill_super,
1080: .fs_flags = FS_USERNS_MOUNT,
1081: };
SourceFile:kernel/bpf/inode.c]
|
| 54/2 |
2026/05/17 11:42 |
llm |
expert |
0m
Model:
Tokens:
input: 42786
output: 43
thoughts: 92
Thoughts: Okay, here's my interpretation of that thought process, from an expert's perspective:
**Investigating a Potential Unprivileged BPF Security Bypass**
Right, `FS_USERNS_MOUNT` is enabled, which immediately gets my attention. That means we're dealing with a system that potentially allows unprivileged users to create their own user namespaces. The interesting bit here is the potential to mount `bpffs` with delegation. This is an immediate red flag. My mental gears are already turning β if someone can do *that*, then they could quite possibly create a token and subsequently try to load a `BPF_PROG_TYPE_RAW_TRACEPOINT` program. The implications are significant, as raw tracepoints can be incredibly powerful, potentially allowing observation or even modification of kernel behavior.
Now, before I jump to any conclusions, I need to verify a critical point. I have to look into `bpf_parse_param` code. The key question is whether delegation options, when mounting `bpffs`, require elevated privileges in the initial user namespace. If they do, this whole line of thinking is a dead end. If not... well, that's where things get interesting, and I'll need to dig much deeper into the security implications. Time to start debugging!
|
| 55/2 |
2026/05/17 11:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/inode.c IncludeLines:true Name:bpf_parse_param]
Results: map[SourceCode: 804: static int bpf_parse_param(struct fs_context *fc, struct fs_parameter *param)
805: {
806: struct bpf_mount_opts *opts = fc->s_fs_info;
807: struct fs_parse_result result;
808: kuid_t uid;
809: kgid_t gid;
810: int opt, err;
811:
812: opt = fs_parse(fc, bpf_fs_parameters, param, &result);
813: if (opt < 0) {
814: /* We might like to report bad mount options here, but
815: * traditionally we've ignored all mount options, so we'd
816: * better continue to ignore non-existing options for bpf.
817: */
818: if (opt == -ENOPARAM) {
819: opt = vfs_parse_fs_param_source(fc, param);
820: if (opt != -ENOPARAM)
821: return opt;
822:
823: return 0;
824: }
825:
826: if (opt < 0)
827: return opt;
828: }
829:
830: switch (opt) {
831: case OPT_UID:
832: uid = make_kuid(current_user_ns(), result.uint_32);
833: if (!uid_valid(uid))
834: goto bad_value;
835:
836: /*
837: * The requested uid must be representable in the
838: * filesystem's idmapping.
839: */
840: if (!kuid_has_mapping(fc->user_ns, uid))
841: goto bad_value;
842:
843: opts->uid = uid;
844: break;
845: case OPT_GID:
846: gid = make_kgid(current_user_ns(), result.uint_32);
847: if (!gid_valid(gid))
848: goto bad_value;
849:
850: /*
851: * The requested gid must be representable in the
852: * filesystem's idmapping.
853: */
854: if (!kgid_has_mapping(fc->user_ns, gid))
855: goto bad_value;
856:
857: opts->gid = gid;
858: break;
859: case OPT_MODE:
860: opts->mode = result.uint_32 & S_IALLUGO;
861: break;
862: case OPT_DELEGATE_CMDS:
863: case OPT_DELEGATE_MAPS:
864: case OPT_DELEGATE_PROGS:
865: case OPT_DELEGATE_ATTACHS: {
866: struct bpffs_btf_enums info;
867: const struct btf_type *enum_t;
868: const char *enum_pfx;
869: u64 *delegate_msk, msk = 0;
870: char *p, *str;
871: int val;
872:
873: /* ignore errors, fallback to hex */
874: (void)find_bpffs_btf_enums(&info);
875:
876: switch (opt) {
877: case OPT_DELEGATE_CMDS:
878: delegate_msk = &opts->delegate_cmds;
879: enum_t = info.cmd_t;
880: enum_pfx = "BPF_";
881: break;
882: case OPT_DELEGATE_MAPS:
883: delegate_msk = &opts->delegate_maps;
884: enum_t = info.map_t;
885: enum_pfx = "BPF_MAP_TYPE_";
886: break;
887: case OPT_DELEGATE_PROGS:
888: delegate_msk = &opts->delegate_progs;
889: enum_t = info.prog_t;
890: enum_pfx = "BPF_PROG_TYPE_";
891: break;
892: case OPT_DELEGATE_ATTACHS:
893: delegate_msk = &opts->delegate_attachs;
894: enum_t = info.attach_t;
895: enum_pfx = "BPF_";
896: break;
897: default:
898: return -EINVAL;
899: }
900:
901: str = param->string;
902: while ((p = strsep(&str, ":"))) {
903: if (strcmp(p, "any") == 0) {
904: msk |= ~0ULL;
905: } else if (find_btf_enum_const(info.btf, enum_t, enum_pfx, p, &val)) {
906: msk |= 1ULL << val;
907: } else {
908: err = kstrtou64(p, 0, &msk);
909: if (err)
910: return err;
911: }
912: }
913:
914: /* Setting delegation mount options requires privileges */
915: if (msk && !capable(CAP_SYS_ADMIN))
916: return -EPERM;
917:
918: *delegate_msk |= msk;
919: break;
920: }
921: default:
922: /* ignore unknown mount options */
923: break;
924: }
925:
926: return 0;
927: bad_value:
928: return invalfc(fc, "Bad value for '%s'", param->key);
929: }
SourceFile:kernel/bpf/inode.c]
|
| 56/2 |
2026/05/17 11:42 |
llm |
expert |
0m
Model:
Tokens:
input: 44998
output: 40
thoughts: 330
Thoughts: **Analyzing the Capability Requirements for BPF Program Loading**
Okay, so this `if (msk && !capable(CAP_SYS_ADMIN))` check is crucial. It means `CAP_SYS_ADMIN` in the *initial* user namespace is needed if we're using `capable()`, not `ns_capable()`. That's a key distinction. This means an unprivileged user can't mount `bpffs` with delegation options, even within a user namespace, which effectively blocks their path to creating a token for `BPF_PROG_TYPE_RAW_TRACEPOINT` loading.
But wait, what if `unprivileged_bpf_disabled` is *not* 1? Oh, right, the instructions *assume* that it *is* 1 β that unprivileged users can't load BPF programs. And the instructions also assume user namespaces are off-limits, no capabilities available there either, checked by `ns_capable()`.
However, now I am rereading the instructions and they are stating: *"Accessible From User Namespaces: Determine if the bug can be reached within a user-namespace where the process has all capabilities (including CAP_SYS_ADMIN, CAP_NET_ADMIN, CAP_NET_RAW, CAP_PERFMON). Such capabilities are checked with ns_capable() function in the kernel sources."*
So, if the exploit *demands* `capable(CAP_SYS_ADMIN)`, then it's *not* reachable from within a user namespace. Excellent. So, it is the initial namespace! I need to re-think this now.
And, `bpf_prog_load` uses `bpf_capable()`. Right, I need to go and look at `bpf_capable()` to understand this. Time to get into the kernel source...
|
| 57/2 |
2026/05/17 11:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:include/linux/capability.h IncludeLines:true Name:bpf_capable]
Results: map[SourceCode: 195: static inline bool bpf_capable(void)
196: {
197: return capable(CAP_BPF) || capable(CAP_SYS_ADMIN);
198: }
SourceFile:include/linux/capability.h]
|
| 58/2 |
2026/05/17 11:42 |
llm |
expert |
0m
Model:
Tokens:
input: 45839
output: 45
thoughts: 248
Thoughts: **My Analysis of BPF Raw Tracepoint Program Loading**
Okay, so `bpf_capable()` relies on `capable()`, which, as I understand it, means it's checking against the initial user namespace. This makes it a non-starter for user namespaces; `bpf_capable()` will always return false in a user namespace. Therefore, my initial thought is that a process within a user namespace simply cannot load `BPF_PROG_TYPE_RAW_TRACEPOINT` programs without some kind of token. And, as I recall, creating this token requires mounting `bpffs` with delegation, which, naturally, also requires `capable(CAP_SYS_ADMIN)` in the *initial* user namespace. Consequently, loading raw tracepoint programs should be completely inaccessible from within a user namespace.
But wait, I need to be thorough. Is there *any* other conceivable attack vector for this bug? My focus here is on the bug in `bpf_link_free`, specifically when `link->sleepable` is false. This leads to a call to `call_rcu(&link->rcu, bpf_link_defer_dealloc_rcu_gp)`. I know that this particular situation arises with `BPF_LINK_TYPE_RAW_TRACEPOINT`. But are there any other BPF link types that:
1. Use `bpf_link_free`
2. Have `link->sleepable == false`
3. Are attached to a tracepoint?
Hmm, let me think... could `bpf_tracing_prog_attach` be a potential candidate?
|
| 59/2 |
2026/05/17 11:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_tracing_prog_attach]
Results: map[SourceCode:3566: static int bpf_tracing_prog_attach(struct bpf_prog *prog,
3567: int tgt_prog_fd,
3568: u32 btf_id,
3569: u64 bpf_cookie,
3570: enum bpf_attach_type attach_type)
3571: {
3572: struct bpf_link_primer link_primer;
3573: struct bpf_prog *tgt_prog = NULL;
3574: struct bpf_trampoline *tr = NULL;
3575: struct bpf_tracing_link *link;
3576: u64 key = 0;
3577: int err;
3578:
3579: switch (prog->type) {
3580: case BPF_PROG_TYPE_TRACING:
3581: if (prog->expected_attach_type != BPF_TRACE_FENTRY &&
3582: prog->expected_attach_type != BPF_TRACE_FEXIT &&
3583: prog->expected_attach_type != BPF_TRACE_FSESSION &&
3584: prog->expected_attach_type != BPF_MODIFY_RETURN) {
3585: err = -EINVAL;
3586: goto out_put_prog;
3587: }
3588: break;
3589: case BPF_PROG_TYPE_EXT:
3590: if (prog->expected_attach_type != 0) {
3591: err = -EINVAL;
3592: goto out_put_prog;
3593: }
3594: break;
3595: case BPF_PROG_TYPE_LSM:
3596: if (prog->expected_attach_type != BPF_LSM_MAC) {
3597: err = -EINVAL;
3598: goto out_put_prog;
3599: }
3600: break;
3601: default:
3602: err = -EINVAL;
3603: goto out_put_prog;
3604: }
3605:
3606: if (!!tgt_prog_fd != !!btf_id) {
3607: err = -EINVAL;
3608: goto out_put_prog;
3609: }
3610:
3611: if (tgt_prog_fd) {
3612: /*
3613: * For now we only allow new targets for BPF_PROG_TYPE_EXT. If this
3614: * part would be changed to implement the same for
3615: * BPF_PROG_TYPE_TRACING, do not forget to update the way how
3616: * attach_tracing_prog flag is set.
3617: */
3618: if (prog->type != BPF_PROG_TYPE_EXT) {
3619: err = -EINVAL;
3620: goto out_put_prog;
3621: }
3622:
3623: tgt_prog = bpf_prog_get(tgt_prog_fd);
3624: if (IS_ERR(tgt_prog)) {
3625: err = PTR_ERR(tgt_prog);
3626: tgt_prog = NULL;
3627: goto out_put_prog;
3628: }
3629:
3630: key = bpf_trampoline_compute_key(tgt_prog, NULL, btf_id);
3631: }
3632:
3633: if (prog->expected_attach_type == BPF_TRACE_FSESSION) {
3634: struct bpf_fsession_link *fslink;
3635:
3636: fslink = kzalloc_obj(*fslink, GFP_USER);
3637: if (fslink) {
3638: bpf_link_init(&fslink->fexit.link, BPF_LINK_TYPE_TRACING,
3639: &bpf_tracing_link_lops, prog, attach_type);
3640: fslink->fexit.cookie = bpf_cookie;
3641: link = &fslink->link;
3642: } else {
3643: link = NULL;
3644: }
3645: } else {
3646: link = kzalloc_obj(*link, GFP_USER);
3647: }
3648: if (!link) {
3649: err = -ENOMEM;
3650: goto out_put_prog;
3651: }
3652: bpf_link_init(&link->link.link, BPF_LINK_TYPE_TRACING,
3653: &bpf_tracing_link_lops, prog, attach_type);
3654:
3655: link->link.cookie = bpf_cookie;
3656:
3657: mutex_lock(&prog->aux->dst_mutex);
3658:
3659: /* There are a few possible cases here:
3660: *
3661: * - if prog->aux->dst_trampoline is set, the program was just loaded
3662: * and not yet attached to anything, so we can use the values stored
3663: * in prog->aux
3664: *
3665: * - if prog->aux->dst_trampoline is NULL, the program has already been
3666: * attached to a target and its initial target was cleared (below)
3667: *
3668: * - if tgt_prog != NULL, the caller specified tgt_prog_fd +
3669: * target_btf_id using the link_create API.
3670: *
3671: * - if tgt_prog == NULL when this function was called using the old
3672: * raw_tracepoint_open API, and we need a target from prog->aux
3673: *
3674: * - if prog->aux->dst_trampoline and tgt_prog is NULL, the program
3675: * was detached and is going for re-attachment.
3676: *
3677: * - if prog->aux->dst_trampoline is NULL and tgt_prog and prog->aux->attach_btf
3678: * are NULL, then program was already attached and user did not provide
3679: * tgt_prog_fd so we have no way to find out or create trampoline
3680: */
3681: if (!prog->aux->dst_trampoline && !tgt_prog) {
3682: /*
3683: * Allow re-attach for TRACING and LSM programs. If it's
3684: * currently linked, bpf_trampoline_link_prog will fail.
3685: * EXT programs need to specify tgt_prog_fd, so they
3686: * re-attach in separate code path.
3687: */
3688: if (prog->type != BPF_PROG_TYPE_TRACING &&
3689: prog->type != BPF_PROG_TYPE_LSM) {
3690: err = -EINVAL;
3691: goto out_unlock;
3692: }
3693: /* We can allow re-attach only if we have valid attach_btf. */
3694: if (!prog->aux->attach_btf) {
3695: err = -EINVAL;
3696: goto out_unlock;
3697: }
3698: btf_id = prog->aux->attach_btf_id;
3699: key = bpf_trampoline_compute_key(NULL, prog->aux->attach_btf, btf_id);
3700: }
3701:
3702: if (!prog->aux->dst_trampoline ||
3703: (key && key != prog->aux->dst_trampoline->key)) {
3704: /* If there is no saved target, or the specified target is
3705: * different from the destination specified at load time, we
3706: * need a new trampoline and a check for compatibility
3707: */
3708: struct bpf_attach_target_info tgt_info = {};
3709:
3710: err = bpf_check_attach_target(NULL, prog, tgt_prog, btf_id,
3711: &tgt_info);
3712: if (err)
3713: goto out_unlock;
3714:
3715: if (tgt_info.tgt_mod) {
3716: module_put(prog->aux->mod);
3717: prog->aux->mod = tgt_info.tgt_mod;
3718: }
3719:
3720: tr = bpf_trampoline_get(key, &tgt_info);
3721: if (!tr) {
3722: err = -ENOMEM;
3723: goto out_unlock;
3724: }
3725: } else {
3726: /* The caller didn't specify a target, or the target was the
3727: * same as the destination supplied during program load. This
3728: * means we can reuse the trampoline and reference from program
3729: * load time, and there is no need to allocate a new one. This
3730: * can only happen once for any program, as the saved values in
3731: * prog->aux are cleared below.
3732: */
3733: tr = prog->aux->dst_trampoline;
3734: tgt_prog = prog->aux->dst_prog;
3735: }
3736:
3737: err = bpf_link_prime(&link->link.link, &link_primer);
3738: if (err)
3739: goto out_unlock;
3740:
3741: err = bpf_trampoline_link_prog(&link->link, tr, tgt_prog);
3742: if (err) {
3743: bpf_link_cleanup(&link_primer);
3744: link = NULL;
3745: goto out_unlock;
3746: }
3747:
3748: link->tgt_prog = tgt_prog;
3749: link->trampoline = tr;
3750:
3751: /* Always clear the trampoline and target prog from prog->aux to make
3752: * sure the original attach destination is not kept alive after a
3753: * program is (re-)attached to another target.
3754: */
3755: if (prog->aux->dst_prog &&
3756: (tgt_prog_fd || tr != prog->aux->dst_trampoline))
3757: /* got extra prog ref from syscall, or attaching to different prog */
3758: bpf_prog_put(prog->aux->dst_prog);
3759: if (prog->aux->dst_trampoline && tr != prog->aux->dst_trampoline)
3760: /* we allocated a new trampoline, so free the old one */
3761: bpf_trampoline_put(prog->aux->dst_trampoline);
3762:
3763: prog->aux->dst_prog = NULL;
3764: prog->aux->dst_trampoline = NULL;
3765: mutex_unlock(&prog->aux->dst_mutex);
3766:
3767: return bpf_link_settle(&link_primer);
3768: out_unlock:
3769: if (tr && tr != prog->aux->dst_trampoline)
3770: bpf_trampoline_put(tr);
3771: mutex_unlock(&prog->aux->dst_mutex);
3772: kfree(link);
3773: out_put_prog:
3774: if (tgt_prog_fd && tgt_prog)
3775: bpf_prog_put(tgt_prog);
3776: return err;
3777: }
SourceFile:kernel/bpf/syscall.c]
|
| 60/2 |
2026/05/17 11:42 |
llm |
expert |
0m
Model:
Tokens:
input: 49861
output: 46
thoughts: 38
Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience, presented in the first person:
**Investigating the BPF Tracepoint Attachment Mechanism**
Alright, so `bpf_tracing_prog_attach` is clearly the entry point for attaching BPF programs to tracing events. That means itβs responsible for the heavy lifting of linking everything together. And I immediately see that it's relying on `bpf_tracing_link_lops`. That's the key: `bpf_tracing_link_lops`. Time to drill down into *that* function. I need to understand exactly what `bpf_tracing_link_lops` is doing. I'm assuming it handles the underlying link operations, but I want to verify the implementation. This is likely where the real magic (or potential gotchas) reside when it comes to attaching the program to the tracepoint. I'll need to dissect the code and follow the logic within `bpf_tracing_link_lops` to confirm my expectations about the overall process. I also need to verify that it is not doing anything I would consider unexpected or sub-optimal.
|
| 61/2 |
2026/05/17 11:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_tracing_link_lops]
Results: map[SourceCode:3559: static const struct bpf_link_ops bpf_tracing_link_lops = {
3560: .release = bpf_tracing_link_release,
3561: .dealloc = bpf_tracing_link_dealloc,
3562: .show_fdinfo = bpf_tracing_link_show_fdinfo,
3563: .fill_link_info = bpf_tracing_link_fill_link_info,
3564: };
SourceFile:kernel/bpf/syscall.c]
|
| 62/2 |
2026/05/17 11:42 |
llm |
expert |
0m
Model:
Tokens:
input: 50348
output: 45
thoughts: 323
Thoughts: **Deallocation Flow in BPF Tracing Links**
Okay, so I'm looking at this deallocation process for BPF tracing links, and something seems off. The code uses `.dealloc` instead of `.dealloc_deferred`, and that's immediately a flag. If `.dealloc_deferred` isn't set, `bpf_link_free` just jumps straight to `bpf_link_dealloc(link)`.
Let me confirm my understanding with the `bpf_link_free` snippet:
```c
if (ops->dealloc_deferred) {
if (link->sleepable || (link->prog && link->prog->sleepable))
call_rcu_tasks_trace(&link->rcu, bpf_link_defer_dealloc_mult_rcu_gp);
else
call_rcu(&link->rcu, bpf_link_defer_dealloc_rcu_gp);
} else if (ops->dealloc) {
bpf_link_dealloc(link);
}
```
Yeah, that's what I thought. If `.dealloc_deferred` is *not* set, it's a direct call to `bpf_link_dealloc(link)`. Which in turn, calls the link's `ops->dealloc(link)`. Now, I recall that `bpf_link_put_direct` calls `bpf_link_free` directly. So `bpf_tracing_link_dealloc` is invoked immediately, with no RCU grace period!
That's a bit surprising. Let me make sure I'm not missing something with the program release. I need to know about the RCU implications, since `bpf_tracing_link_release` is the one unlinking the program from the trampoline. Okay, time to check `bpf_tracing_link_release`.
|
| 63/2 |
2026/05/17 11:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_tracing_link_release]
Results: map[SourceCode:3500: static void bpf_tracing_link_release(struct bpf_link *link)
3501: {
3502: struct bpf_tracing_link *tr_link =
3503: container_of(link, struct bpf_tracing_link, link.link);
3504:
3505: WARN_ON_ONCE(bpf_trampoline_unlink_prog(&tr_link->link,
3506: tr_link->trampoline,
3507: tr_link->tgt_prog));
3508:
3509: bpf_trampoline_put(tr_link->trampoline);
3510:
3511: /* tgt_prog is NULL if target is a kernel function */
3512: if (tr_link->tgt_prog)
3513: bpf_prog_put(tr_link->tgt_prog);
3514: }
SourceFile:kernel/bpf/syscall.c]
|
| 64/2 |
2026/05/17 11:42 |
llm |
expert |
0m
Model:
Tokens:
input: 51414
output: 47
thoughts: 40
Thoughts: Okay, here's my take on those thoughts, framed as my own:
**Unlinking BPF Programs: The RCU Question**
Alright, so `bpf_trampoline_unlink_prog` is tasked with unlinking the BPF program. Makes sense. But the real question burning in my mind is: *does it properly handle the Resource, Control, and Update (RCU) synchronization that's crucial in these kernel operations?* We're dealing with concurrent access, and potential dangling pointers are my bread and butter. I can't just assume it handles it correctly; I need to verify. So, let's dive into the `bpf_trampoline_unlink_prog` code itself and see how the deletion is actually implemented. Time to put the debugging hat on and analyze!
|
| 65/2 |
2026/05/17 11:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/trampoline.c IncludeLines:true Name:bpf_trampoline_unlink_prog]
Results: map[SourceCode: 883: /* bpf_trampoline_unlink_prog() should never fail. */
884: int bpf_trampoline_unlink_prog(struct bpf_tramp_link *link,
885: struct bpf_trampoline *tr,
886: struct bpf_prog *tgt_prog)
887: {
888: int err;
889:
890: mutex_lock(&tr->mutex);
891: err = __bpf_trampoline_unlink_prog(link, tr, tgt_prog);
892: mutex_unlock(&tr->mutex);
893: return err;
894: }
SourceFile:kernel/bpf/trampoline.c]
|
| 66/2 |
2026/05/17 11:42 |
llm |
expert |
0m
Model:
Tokens:
input: 51890
output: 48
thoughts: 16
Thoughts: Okay, let's dive into this. Here's my thought process on checking `__bpf_trampoline_unlink_prog`, described as if they are my own:
**Analyzing the BPF Trampoline Unlink Function**
Alright, I need to get a handle on `__bpf_trampoline_unlink_prog`. This likely pertains to removing BPF programs that have been linked, presumably using trampolines. My first thought is to understand the context. What's the surrounding code doing? Is this related to cleanup, error handling, or perhaps some specific BPF functionality? I need to know where this function is called from, and what data it operates on.
I need to understand the underlying data structures, probably the trampoline itself. I'll need to know:
* How the programs are linked, i.e., what's connecting the BPF program to the target (the trampoline).
* What resources are being released during the unlinking process. Is memory being freed? Are reference counts being decremented? Are there any locks involved?
* Potential side effects. Does this operation affect other linked programs or related data?
I'll be looking for potential issues:
* Race conditions. If multiple threads are manipulating the same trampoline/program, are there proper synchronization mechanisms in place?
* Memory leaks. If resources aren't freed correctly, we'll have problems.
* Invalid pointer dereferences. I'll check how this is handled, as well.
* Security vulnerabilities. Is there a way to misuse the unlinking process to achieve unintended effects?
Ultimately, my goal is to verify that this function correctly and safely removes BPF programs linked via trampolines, ensuring proper resource management and preventing any potential security flaws. I'll need to analyze the code, trace execution paths, and potentially write tests to ensure robustness.
|
| 67/2 |
2026/05/17 11:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/trampoline.c IncludeLines:true Name:__bpf_trampoline_unlink_prog]
Results: map[SourceCode: 853: static int __bpf_trampoline_unlink_prog(struct bpf_tramp_link *link,
854: struct bpf_trampoline *tr,
855: struct bpf_prog *tgt_prog)
856: {
857: enum bpf_tramp_prog_type kind;
858: int err;
859:
860: kind = bpf_attach_type_to_tramp(link->link.prog);
861: if (kind == BPF_TRAMP_REPLACE) {
862: WARN_ON_ONCE(!tr->extension_prog);
863: err = bpf_arch_text_poke(tr->func.addr, BPF_MOD_JUMP,
864: BPF_MOD_NOP,
865: tr->extension_prog->bpf_func, NULL);
866: tr->extension_prog = NULL;
867: guard(mutex)(&tgt_prog->aux->ext_mutex);
868: tgt_prog->aux->is_extended = false;
869: return err;
870: } else if (kind == BPF_TRAMP_FSESSION) {
871: struct bpf_fsession_link *fslink =
872: container_of(link, struct bpf_fsession_link, link.link);
873:
874: hlist_del_init(&fslink->fexit.tramp_hlist);
875: tr->progs_cnt[BPF_TRAMP_FEXIT]--;
876: kind = BPF_TRAMP_FENTRY;
877: }
878: hlist_del_init(&link->tramp_hlist);
879: tr->progs_cnt[kind]--;
880: return bpf_trampoline_update(tr, true /* lock_direct_mutex */);
881: }
SourceFile:kernel/bpf/trampoline.c]
|
| 68/2 |
2026/05/17 11:42 |
llm |
expert |
0m
Model:
Tokens:
input: 52915
output: 45
thoughts: 35
Thoughts: Okay, here's my interpretation of that thought process, from an expert's perspective:
**Trampoline Update and Concurrency Considerations**
Right, so `bpf_trampoline_update`. It's job is to update the trampoline, and the question that immediately springs to mind, given the nature of BPF and kernel modifications, is: does this function incorporate any mechanism to handle concurrent access and prevent data races? Specifically, does it take into account Read-Copy-Update (RCU) synchronization? My experience dictates that in this kind of kernel update scenario, RCU is often the go-to approach to ensure that readers of the trampoline data continue to function correctly while the update is in progress. Therefore, before moving on, the first thing I need to do is to dive into the source code for `bpf_trampoline_update` and confirm whether or not it's using RCU or some other synchronization method. I need to understand how it protects against concurrent access.
|
| 69/2 |
2026/05/17 11:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/trampoline.c IncludeLines:true Name:bpf_trampoline_update]
Results: map[SourceCode: 607: static int bpf_trampoline_update(struct bpf_trampoline *tr, bool lock_direct_mutex)
608: {
609: struct bpf_tramp_image *im;
610: struct bpf_tramp_links *tlinks;
611: u32 orig_flags = tr->flags;
612: bool ip_arg = false;
613: int err, total, size;
614:
615: tlinks = bpf_trampoline_get_progs(tr, &total, &ip_arg);
616: if (IS_ERR(tlinks))
617: return PTR_ERR(tlinks);
618:
619: if (total == 0) {
620: err = unregister_fentry(tr, orig_flags, tr->cur_image->image);
621: bpf_tramp_image_put(tr->cur_image);
622: tr->cur_image = NULL;
623: goto out;
624: }
625:
626: /* clear all bits except SHARE_IPMODIFY and TAIL_CALL_CTX */
627: tr->flags &= (BPF_TRAMP_F_SHARE_IPMODIFY | BPF_TRAMP_F_TAIL_CALL_CTX);
628:
629: if (tlinks[BPF_TRAMP_FEXIT].nr_links ||
630: tlinks[BPF_TRAMP_MODIFY_RETURN].nr_links) {
631: /* NOTE: BPF_TRAMP_F_RESTORE_REGS and BPF_TRAMP_F_SKIP_FRAME
632: * should not be set together.
633: */
634: tr->flags |= BPF_TRAMP_F_CALL_ORIG | BPF_TRAMP_F_SKIP_FRAME;
635: } else {
636: tr->flags |= BPF_TRAMP_F_RESTORE_REGS;
637: }
638:
639: if (ip_arg)
640: tr->flags |= BPF_TRAMP_F_IP_ARG;
641:
642: #ifdef CONFIG_DYNAMIC_FTRACE_WITH_DIRECT_CALLS
643: again:
644: if (tr->flags & BPF_TRAMP_F_CALL_ORIG) {
645: if (tr->flags & BPF_TRAMP_F_SHARE_IPMODIFY) {
646: /* The BPF_TRAMP_F_SKIP_FRAME can be cleared in the
647: * first try, reset it in the second try.
648: */
649: tr->flags |= BPF_TRAMP_F_ORIG_STACK | BPF_TRAMP_F_SKIP_FRAME;
650: } else if (IS_ENABLED(CONFIG_DYNAMIC_FTRACE_WITH_JMP)) {
651: /* Use "jmp" instead of "call" for the trampoline
652: * in the origin call case, and we don't need to
653: * skip the frame.
654: */
655: tr->flags &= ~BPF_TRAMP_F_SKIP_FRAME;
656: }
657: }
658: #endif
659:
660: size = arch_bpf_trampoline_size(&tr->func.model, tr->flags,
661: tlinks, tr->func.addr);
662: if (size < 0) {
663: err = size;
664: goto out;
665: }
666:
667: if (size > PAGE_SIZE) {
668: err = -E2BIG;
669: goto out;
670: }
671:
672: im = bpf_tramp_image_alloc(tr->key, size);
673: if (IS_ERR(im)) {
674: err = PTR_ERR(im);
675: goto out;
676: }
677:
678: err = arch_prepare_bpf_trampoline(im, im->image, im->image + size,
679: &tr->func.model, tr->flags, tlinks,
680: tr->func.addr);
681: if (err < 0)
682: goto out_free;
683:
684: err = arch_protect_bpf_trampoline(im->image, im->size);
685: if (err)
686: goto out_free;
687:
688: WARN_ON(tr->cur_image && total == 0);
689: if (tr->cur_image)
690: /* progs already running at this address */
691: err = modify_fentry(tr, orig_flags, tr->cur_image->image,
692: im->image, lock_direct_mutex);
693: else
694: /* first time registering */
695: err = register_fentry(tr, im->image);
696:
697: #ifdef CONFIG_DYNAMIC_FTRACE_WITH_DIRECT_CALLS
698: if (err == -EAGAIN) {
699: /* -EAGAIN from bpf_tramp_ftrace_ops_func. Now
700: * BPF_TRAMP_F_SHARE_IPMODIFY is set, we can generate the
701: * trampoline again, and retry register.
702: */
703: bpf_tramp_image_free(im);
704: goto again;
705: }
706: #endif
707: if (err)
708: goto out_free;
709:
710: if (tr->cur_image)
711: bpf_tramp_image_put(tr->cur_image);
712: tr->cur_image = im;
713: out:
714: /* If any error happens, restore previous flags */
715: if (err)
716: tr->flags = orig_flags;
717: kfree(tlinks);
718: return err;
719:
720: out_free:
721: bpf_tramp_image_free(im);
722: goto out;
723: }
SourceFile:kernel/bpf/trampoline.c]
|
| 70/2 |
2026/05/17 11:42 |
llm |
expert |
0m
Model:
Tokens:
input: 55089
output: 42
thoughts: 31
Thoughts: Okay, here's my take on those thoughts, framed as an expert's internal monologue:
**Unpacking `unregister_fentry`**
Right, so `unregister_fentry` is the function responsible for⦠well, unregistering the fentry. Seems straightforward enough. But, being the seasoned pro I am, my mind immediately jumps to the critical details: *does it handle RCU correctly*? That's the real crux of this operation. We're dealing with live code modifications, and in a multi-threaded environment like the kernel, RCU (Read-Copy-Update) is absolutely essential to prevent data corruption and ensure thread safety. Any removal of the fentry must be properly synchronized to avoid race conditions. Time to delve into the code itself and see how it is handled in `unregister_fentry`.
|
| 71/2 |
2026/05/17 11:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/trampoline.c IncludeLines:true Name:unregister_fentry]
Results: map[SourceCode: 389: static int unregister_fentry(struct bpf_trampoline *tr, u32 orig_flags,
390: void *old_addr)
391: {
392: int ret;
393:
394: if (tr->func.ftrace_managed)
395: ret = direct_ops_del(tr, old_addr);
396: else
397: ret = bpf_trampoline_update_fentry(tr, orig_flags, old_addr, NULL);
398:
399: return ret;
400: }
SourceFile:kernel/bpf/trampoline.c]
|
| 72/2 |
2026/05/17 11:42 |
llm |
expert |
0m
Model:
Tokens:
input: 55542
output: 46
thoughts: 1059
Thoughts: **Analysis of a BPF Use-After-Free Vulnerability**
Okay, let's break this down. My initial thought process is centered around how `bpf_trampoline_update_fentry` interacts with memory management. It's clear that it uses `bpf_arch_text_poke`, which in turn calls `text_poke_bp`. The use of `text_poke_bp` and its IPI synchronization is key, but the lack of RCU grace period protection raised a red flag immediately. Then, I observed that the trampoline image is freed with `call_rcu_tasks_trace`, while the `link` object is immediately freed, directly calling `bpf_link_dealloc`. This seems problematic since the trampoline could still be executing.
Thankfully, `bpf_tracing_link_release` calls `bpf_trampoline_unlink_prog` which updates the trampoline, and the old image is freed safely. Also, the trampoline itself accesses the `bpf_prog` object and not the `link` object. The `bpf_prog` object is also protected by `call_rcu_tasks_trace`. So, the immediate free of the `link` object *should* be fine in the general case.
However, the problem surfaced when considering tracepoints. Tracepoints directly hold a pointer to the `link` object in the tracepoint array! This is where the vulnerability lies and I can see why `bpf_raw_tp_link_release` is different. This release unregisters the probe, which will remove the `link` pointer from the array, but a concurrent execution might still be reading it.
This led to the realization that the `.dealloc_deferred = bpf_raw_tp_link_dealloc` is used, so the `bpf_link_free` calls `call_rcu` instead. So, the tracepoint is protected by an RCU grace period. But it's an ordinary RCU grace period, and the tracepoint itself is protected by SRCU. This is where the bug arises because a normal RCU grace period is not enough to cover it. The root cause is that SRCU is active and we need to wait for SRCU grace periods instead of RCU grace periods!
This bug is specific to `BPF_PROG_TYPE_RAW_TRACEPOINT` and `BPF_PROG_TYPE_RAW_TRACEPOINT_WRITABLE` programs, both of which require `CAP_BPF` or `CAP_SYS_ADMIN` in the initial user namespace. The good news is unprivileged users cannot load these programs, and thus it's inaccessible from unprivileged processes or user namespaces.
Now, the question is, is it exploitable? Yes, it's a use-after-free of `struct bpf_raw_tp_link`. A malicious actor could free the `link` object, reallocate that memory, and control the `prog` pointer. Then, calling `bpf_prog_run(prog, args)` becomes an arbitrary code execution primitive. It's definitely exploitable for privilege escalation β for instance, bypassing lockdown restrictions, although `CAP_BPF` is already very powerful on its own. It's also exploitable as a denial-of-service, leading to kernel crashes.
The bug is not accessible from VM guests or hosts in confidential computing, nor via network or file system-based attacks. The attacker needs local code execution with `CAP_BPF` to load and attach a BPF program.
As a final check, I considered whether unprivileged users could attach tracepoints via `perf_event_open`. They *could* do so if `perf_event_paranoid` is set to -1 or lower, but normally this is set to 2. Even then, the vulnerability cannot be triggered because `perf_event_open` uses `perf_event_attach_bpf_prog` instead of `bpf_raw_tp_link_attach`.
|
| 73/2 |
2026/05/17 11:43 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_event_attach_bpf_prog]
Results: map[SourceCode:1920: int perf_event_attach_bpf_prog(struct perf_event *event,
1921: struct bpf_prog *prog,
1922: u64 bpf_cookie)
1923: {
1924: struct bpf_prog_array *old_array;
1925: struct bpf_prog_array *new_array;
1926: int ret = -EEXIST;
1927:
1928: /*
1929: * Kprobe override only works if they are on the function entry,
1930: * and only if they are on the opt-in list.
1931: */
1932: if (prog->kprobe_override &&
1933: (!trace_kprobe_on_func_entry(event->tp_event) ||
1934: !trace_kprobe_error_injectable(event->tp_event)))
1935: return -EINVAL;
1936:
1937: mutex_lock(&bpf_event_mutex);
1938:
1939: if (event->prog)
1940: goto unlock;
1941:
1942: old_array = bpf_event_rcu_dereference(event->tp_event->prog_array);
1943: if (old_array &&
1944: bpf_prog_array_length(old_array) >= BPF_TRACE_MAX_PROGS) {
1945: ret = -E2BIG;
1946: goto unlock;
1947: }
1948:
1949: ret = bpf_prog_array_copy(old_array, NULL, prog, bpf_cookie, &new_array);
1950: if (ret < 0)
1951: goto unlock;
1952:
1953: /* set the new array to event->tp_event and set event->prog */
1954: event->prog = prog;
1955: event->bpf_cookie = bpf_cookie;
1956: rcu_assign_pointer(event->tp_event->prog_array, new_array);
1957: bpf_prog_array_free_sleepable(old_array);
1958:
1959: unlock:
1960: mutex_unlock(&bpf_event_mutex);
1961: return ret;
1962: }
SourceFile:kernel/trace/bpf_trace.c]
|
| 74/2 |
2026/05/17 11:43 |
llm |
expert |
0m
Model:
Tokens:
input: 58198
output: 49
thoughts: 77
Thoughts: Okay, let's see...
**Deconstructing the BPF Perf Event Attachment**
Alright, let's break this down. Where is this `perf_event_attach_bpf_prog` function? Ah, yes, it's buried in `kernel/trace/bpf_trace.c`. So, that clearly handles the core attachment of a BPF program to a perf event. I need to remember that. The whole thing, in the context of tracing, hinges on that. Now, how does the perf event tracepoint actually run the attached BPF program? It looks like `perf_trace_run_bpf_submit` handles the actual execution. Okay, *that's* the key. I need to drill down into `perf_trace_run_bpf_submit` next to understand the execution pathway within this context. That's where the rubber meets the road. Time to dive in.
|
| 75/2 |
2026/05/17 11:43 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/bpf_trace.c IncludeLines:true Name:perf_trace_run_bpf_submit]
Results: map[SourceCode:11325: void perf_trace_run_bpf_submit(void *raw_data, int size, int rctx,
11326: struct trace_event_call *call, u64 count,
11327: struct pt_regs *regs, struct hlist_head *head,
11328: struct task_struct *task)
11329: {
11330: if (bpf_prog_array_valid(call)) {
11331: *(struct pt_regs **)raw_data = regs;
11332: if (!trace_call_bpf(call, raw_data) || hlist_empty(head)) {
11333: perf_swevent_put_recursion_context(rctx);
11334: return;
11335: }
11336: }
11337: perf_tp_event(call->event.type, count, raw_data, size, regs, head,
11338: rctx, task);
11339: }
SourceFile:kernel/events/core.c]
|
| 76/2 |
2026/05/17 11:43 |
llm |
expert |
0m
Model:
Tokens:
input: 58830
output: 29
thoughts: 492
Thoughts: **Analyzing BPF Program Loading and Raw Tracepoint Links**
Okay, let's break this down. My initial thought was that `trace_call_bpf` might utilize `struct bpf_raw_tp_link` directly when executing a BPF program. However, a quick review reveals it uses `struct bpf_prog_array` instead. So the crux of this bug is squarely around the usage of `struct bpf_raw_tp_link`, specifically in the context of `bpf_raw_tracepoint_open`. And as I recall, `bpf_raw_tracepoint_open` is restricted by `CAP_BPF` since it only accepts `BPF_PROG_TYPE_RAW_TRACEPOINT` and related types, which need elevated privileges.
Now, a more interesting question comes to mind: what about other BPF program types? Can an unprivileged user load, say, a `BPF_PROG_TYPE_TRACING` program? No, `bpf_prog_load` appropriately checks `bpf_cap` for that and other types such as `BPF_PROG_TYPE_EXT` and `BPF_PROG_TYPE_LSM`. I recall those checks.
So, it seems that only `BPF_PROG_TYPE_SOCKET_FILTER` and `BPF_PROG_TYPE_CGROUP_SKB` programs can be loaded by an unprivileged user, assuming `unprivileged_bpf_disabled` is set to 0.
But then, the crucial part: `bpf_raw_tp_link_attach` specifically rejects these types! Looking at the code, it explicitly disallows `BPF_PROG_TYPE_TRACING`, `BPF_PROG_TYPE_EXT`, `BPF_PROG_TYPE_LSM`, `BPF_PROG_TYPE_RAW_TRACEPOINT`, and `BPF_PROG_TYPE_RAW_TRACEPOINT_WRITABLE`.
Therefore, it seems that unprivileged users cannot use `bpf_raw_tp_link_attach` under any current configuration.
Let me explore if there are any other ways to create a `bpf_raw_tp_link`, and I will start by checking references to `bpf_raw_tp_link_lops`.
|
| 77/2 |
2026/05/17 11:43 |
tool |
codesearch-find-references |
0m
Args: map[Name:bpf_raw_tp_link_lops]
Results: map[References:<nil> TruncatedOutput:false]
|
| 78/2 |
2026/05/17 11:43 |
llm |
expert |
0m
Model:
Tokens:
input: 59869
output: 29
thoughts: 29
Thoughts: Here's my thought process, summarized as you requested:
**Navigating the Labyrinth of BPF References**
Okay, so I'm trying to track down all the uses of `bpf_raw_tp_link_lops`. This feels like a task where I'm going to want to make sure I am thorough. I need to make sure I don't miss anything. Let's see... Rather than going for a more manual approach, I should leverage the tools at my disposal. Right, the best way to do that is probably to jump straight to using `codesearch-find-references`. That'll be the fastest and most reliable approach for finding everything. Let's go ahead and use that specifically with the name `bpf_raw_tp_link_lops`. That should give me a complete, or close to complete, list of all the places this structure is used. Now to get started.
|
| 79/2 |
2026/05/17 11:43 |
tool |
codesearch-find-references |
0m
Args: map[Name:bpf_raw_tp_link_lops]
Results: map[References:<nil> TruncatedOutput:false]
|
| 80/2 |
2026/05/17 11:43 |
llm |
expert |
0m
Model:
Tokens:
input: 60140
output: 28
thoughts: 23
Thoughts: Okay, here's my interpretation of those thoughts, tailored for an expert audience, and delivered as a first-person summary:
**Initial Assessment and Alternative Approach**
Right, so the initial query returned nada, huh? Hmm, that's a bit perplexing. Given my experience, I was expecting some results. Clearly, the approach I took initially isn't yielding the desired information. I'm going to shift gears. Instead of relying on that, I'll bypass the usual avenues and try `bpf_raw_tp_link_attach`. It feels like there's a specific angle I'm not hitting with the first command, and this feels like it will get me closer. Let's see if this gets me a better perspective.
|
| 81/2 |
2026/05/17 11:43 |
tool |
codesearch-find-references |
0m
Args: map[Name:bpf_raw_tp_link_attach]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:bpf_raw_tracepoint_open SourceFile:kernel/bpf/syscall.c SourceLine:4312] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:link_create SourceFile:kernel/bpf/syscall.c SourceLine:5725]] TruncatedOutput:false]
|
| 82/2 |
2026/05/17 11:43 |
llm |
expert |
0m
Model:
Tokens:
input: 60472
output: 40
thoughts: 31
Thoughts: Okay, here's my take on those thoughts, framed as an expert's internal monologue:
**Diving into the BPF Link Creation**
Right, so `link_create`... interesting. I see it's also responsible for calling `bpf_raw_tp_link_attach`. Makes sense, the creation process needs to handle attaching the BPF program, obviously. Let's delve into that `link_create` function itself. I want to understand the specifics of how the links are constructed and how that attachment is managed. I'll need to analyze the code flow to grasp the crucial aspects. It's time to dig in and confirm my prior understanding.
|
| 83/2 |
2026/05/17 11:43 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:link_create]
Results: map[SourceCode:5681: static int link_create(union bpf_attr *attr, bpfptr_t uattr)
5682: {
5683: struct bpf_prog *prog;
5684: int ret;
5685:
5686: if (CHECK_ATTR(BPF_LINK_CREATE))
5687: return -EINVAL;
5688:
5689: if (attr->link_create.attach_type == BPF_STRUCT_OPS)
5690: return bpf_struct_ops_link_create(attr);
5691:
5692: prog = bpf_prog_get(attr->link_create.prog_fd);
5693: if (IS_ERR(prog))
5694: return PTR_ERR(prog);
5695:
5696: ret = bpf_prog_attach_check_attach_type(prog,
5697: attr->link_create.attach_type);
5698: if (ret)
5699: goto out;
5700:
5701: switch (prog->type) {
5702: case BPF_PROG_TYPE_CGROUP_SKB:
5703: case BPF_PROG_TYPE_CGROUP_SOCK:
5704: case BPF_PROG_TYPE_CGROUP_SOCK_ADDR:
5705: case BPF_PROG_TYPE_SOCK_OPS:
5706: case BPF_PROG_TYPE_CGROUP_DEVICE:
5707: case BPF_PROG_TYPE_CGROUP_SYSCTL:
5708: case BPF_PROG_TYPE_CGROUP_SOCKOPT:
5709: ret = cgroup_bpf_link_attach(attr, prog);
5710: break;
5711: case BPF_PROG_TYPE_EXT:
5712: ret = bpf_tracing_prog_attach(prog,
5713: attr->link_create.target_fd,
5714: attr->link_create.target_btf_id,
5715: attr->link_create.tracing.cookie,
5716: attr->link_create.attach_type);
5717: break;
5718: case BPF_PROG_TYPE_LSM:
5719: case BPF_PROG_TYPE_TRACING:
5720: if (attr->link_create.attach_type != prog->expected_attach_type) {
5721: ret = -EINVAL;
5722: goto out;
5723: }
5724: if (prog->expected_attach_type == BPF_TRACE_RAW_TP)
5725: ret = bpf_raw_tp_link_attach(prog, NULL, attr->link_create.tracing.cookie,
5726: attr->link_create.attach_type);
5727: else if (prog->expected_attach_type == BPF_TRACE_ITER)
5728: ret = bpf_iter_link_attach(attr, uattr, prog);
5729: else if (prog->expected_attach_type == BPF_LSM_CGROUP)
5730: ret = cgroup_bpf_link_attach(attr, prog);
5731: else
5732: ret = bpf_tracing_prog_attach(prog,
5733: attr->link_create.target_fd,
5734: attr->link_create.target_btf_id,
5735: attr->link_create.tracing.cookie,
5736: attr->link_create.attach_type);
5737: break;
5738: case BPF_PROG_TYPE_FLOW_DISSECTOR:
5739: case BPF_PROG_TYPE_SK_LOOKUP:
5740: ret = netns_bpf_link_create(attr, prog);
5741: break;
5742: case BPF_PROG_TYPE_SK_MSG:
5743: case BPF_PROG_TYPE_SK_SKB:
5744: ret = sock_map_link_create(attr, prog);
5745: break;
5746: #ifdef CONFIG_NET
5747: case BPF_PROG_TYPE_XDP:
5748: ret = bpf_xdp_link_attach(attr, prog);
5749: break;
5750: case BPF_PROG_TYPE_SCHED_CLS:
5751: if (attr->link_create.attach_type == BPF_TCX_INGRESS ||
5752: attr->link_create.attach_type == BPF_TCX_EGRESS)
5753: ret = tcx_link_attach(attr, prog);
5754: else
5755: ret = netkit_link_attach(attr, prog);
5756: break;
5757: case BPF_PROG_TYPE_NETFILTER:
5758: ret = bpf_nf_link_attach(attr, prog);
5759: break;
5760: #endif
5761: case BPF_PROG_TYPE_PERF_EVENT:
5762: case BPF_PROG_TYPE_TRACEPOINT:
5763: ret = bpf_perf_link_attach(attr, prog);
5764: break;
5765: case BPF_PROG_TYPE_KPROBE:
5766: if (attr->link_create.attach_type == BPF_PERF_EVENT)
5767: ret = bpf_perf_link_attach(attr, prog);
5768: else if (attr->link_create.attach_type == BPF_TRACE_KPROBE_MULTI ||
5769: attr->link_create.attach_type == BPF_TRACE_KPROBE_SESSION)
5770: ret = bpf_kprobe_multi_link_attach(attr, prog);
5771: else if (attr->link_create.attach_type == BPF_TRACE_UPROBE_MULTI ||
5772: attr->link_create.attach_type == BPF_TRACE_UPROBE_SESSION)
5773: ret = bpf_uprobe_multi_link_attach(attr, prog);
5774: break;
5775: default:
5776: ret = -EINVAL;
5777: }
5778:
5779: out:
5780: if (ret < 0)
5781: bpf_prog_put(prog);
5782: return ret;
5783: }
SourceFile:kernel/bpf/syscall.c]
|
| 84/2 |
2026/05/17 11:43 |
llm |
expert |
0m
Model:
Tokens:
input: 62535
output: 45
thoughts: 141
Thoughts: **Examining BPF Program Attachments for Unprivileged Users**
Okay, so I'm digging into these BPF program attachment points again. The `link_create` function initially caught my eye. It seems like when dealing with `BPF_PROG_TYPE_TRACING` and `BPF_TRACE_RAW_TP`, it's calling `bpf_raw_tp_link_attach`. Right. And as I recall, `BPF_PROG_TYPE_TRACING` *still* needs `CAP_BPF` to actually load. Dammit. So that route is a dead end for an unprivileged user.
Alright, let's pivot. What about `BPF_PROG_TYPE_TRACEPOINT`? That's a different beast, right? I see `link_create` calls `bpf_perf_link_attach` when working with a `TRACEPOINT`. Okay, this is interesting. I've got to investigate `bpf_perf_link_attach`. Let's see what we can find there.
|
| 85/2 |
2026/05/17 11:43 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_perf_link_attach]
Results: map[SourceCode:4171: static int bpf_perf_link_attach(const union bpf_attr *attr, struct bpf_prog *prog)
4172: {
4173: struct bpf_link_primer link_primer;
4174: struct bpf_perf_link *link;
4175: struct perf_event *event;
4176: struct file *perf_file;
4177: int err;
4178:
4179: if (attr->link_create.flags)
4180: return -EINVAL;
4181:
4182: perf_file = perf_event_get(attr->link_create.target_fd);
4183: if (IS_ERR(perf_file))
4184: return PTR_ERR(perf_file);
4185:
4186: link = kzalloc_obj(*link, GFP_USER);
4187: if (!link) {
4188: err = -ENOMEM;
4189: goto out_put_file;
4190: }
4191: bpf_link_init(&link->link, BPF_LINK_TYPE_PERF_EVENT, &bpf_perf_link_lops, prog,
4192: attr->link_create.attach_type);
4193: link->perf_file = perf_file;
4194:
4195: err = bpf_link_prime(&link->link, &link_primer);
4196: if (err) {
4197: kfree(link);
4198: goto out_put_file;
4199: }
4200:
4201: event = perf_file->private_data;
4202: err = perf_event_set_bpf_prog(event, prog, attr->link_create.perf_event.bpf_cookie);
4203: if (err) {
4204: bpf_link_cleanup(&link_primer);
4205: goto out_put_file;
4206: }
4207: /* perf_event_set_bpf_prog() doesn't take its own refcnt on prog */
4208: bpf_prog_inc(prog);
4209:
4210: return bpf_link_settle(&link_primer);
4211:
4212: out_put_file:
4213: fput(perf_file);
4214: return err;
4215: }
SourceFile:kernel/bpf/syscall.c]
|
| 86/2 |
2026/05/17 11:43 |
llm |
expert |
0m
Model:
Tokens:
input: 63669
output: 46
thoughts: 38
Thoughts: Here's my analysis of the provided information, as if these were my internal thoughts:
**Dissecting the BPF Perf Link Attachment**
Okay, so we're dealing with `bpf_perf_link_attach`. My first instinct is to see how it's implemented and what it leverages. The text hints that it utilizes `bpf_perf_link_lops`. That's the key here. It's likely that `bpf_perf_link_lops` holds the operations or function pointers that define the specifics of how this link is created and managed. My next step is to immediately dive into `bpf_perf_link_lops` to understand the low-level details of the underlying mechanics. We must discover how the perf link is functioning.
|
| 87/2 |
2026/05/17 11:43 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_perf_link_lops]
Results: map[SourceCode:4164: static const struct bpf_link_ops bpf_perf_link_lops = {
4165: .release = bpf_perf_link_release,
4166: .dealloc = bpf_perf_link_dealloc,
4167: .fill_link_info = bpf_perf_link_fill_link_info,
4168: .show_fdinfo = bpf_perf_link_show_fdinfo,
4169: };
SourceFile:kernel/bpf/syscall.c]
|
| 88/2 |
2026/05/17 11:43 |
llm |
expert |
0m
Model:
Tokens:
input: 64077
output: 46
thoughts: 63
Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience:
**Immediate Link Deallocation Concern**
Right, so I see it's using `.dealloc` instead of `.dealloc_deferred`. That's a significant detail. This means the `link` object is being freed immediately, which, in the context of BPF program loading and performance monitoring, could be... well, problematic if not handled with absolute care.
My immediate follow-up question is: Is `perf_event_set_bpf_prog` responsible for keeping the `link` pointer alive? Does it save a copy, or does it implicitly *rely* on the link's continued existence? I need to know precisely how that function handles the underlying BPF link.
Okay, time to dive into `perf_event_set_bpf_prog`. Let's confirm how the link pointer is managed within that call.
|
| 89/2 |
2026/05/17 11:43 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_event_set_bpf_prog]
Results: map[SourceCode:11668: int perf_event_set_bpf_prog(struct perf_event *event,
11669: struct bpf_prog *prog,
11670: u64 bpf_cookie)
11671: {
11672: struct perf_event_context *ctx;
11673: int ret;
11674:
11675: ctx = perf_event_ctx_lock(event);
11676: ret = __perf_event_set_bpf_prog(event, prog, bpf_cookie);
11677: perf_event_ctx_unlock(event, ctx);
11678:
11679: return ret;
11680: }
SourceFile:kernel/events/core.c]
|
| 90/2 |
2026/05/17 11:43 |
llm |
expert |
0m
Model:
Tokens:
input: 64618
output: 47
thoughts: 17
Thoughts: Here's my attempt at summarizing the thought process, tailored for an expert audience and written in the first person:
**Diving into `__perf_event_set_bpf_prog`**
Okay, let's get right into it. My focus is now squarely on the `__perf_event_set_bpf_prog` function. I need to understand its role in the bigger picture, particularly how it interacts with the BPF machinery within the performance event subsystem. I want to quickly refresh my memory on the core aspects: What is this function doing? What data structures does it manipulate? What are the key locking considerations? Are there any potential gotchas or race conditions that I should be aware of, given the concurrent nature of these performance events? I will have to start by tracing the function's call sites and dependencies, and then dive into the code itself.
|
| 91/2 |
2026/05/17 11:43 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:__perf_event_set_bpf_prog]
Results: map[SourceCode:11621: static int __perf_event_set_bpf_prog(struct perf_event *event,
11622: struct bpf_prog *prog,
11623: u64 bpf_cookie)
11624: {
11625: bool is_kprobe, is_uprobe, is_tracepoint, is_syscall_tp;
11626:
11627: if (event->state <= PERF_EVENT_STATE_REVOKED)
11628: return -ENODEV;
11629:
11630: if (!perf_event_is_tracing(event))
11631: return perf_event_set_bpf_handler(event, prog, bpf_cookie);
11632:
11633: is_kprobe = event->tp_event->flags & TRACE_EVENT_FL_KPROBE;
11634: is_uprobe = event->tp_event->flags & TRACE_EVENT_FL_UPROBE;
11635: is_tracepoint = event->tp_event->flags & TRACE_EVENT_FL_TRACEPOINT;
11636: is_syscall_tp = is_syscall_trace_event(event->tp_event);
11637: if (!is_kprobe && !is_uprobe && !is_tracepoint && !is_syscall_tp)
11638: /* bpf programs can only be attached to u/kprobe or tracepoint */
11639: return -EINVAL;
11640:
11641: if (((is_kprobe || is_uprobe) && prog->type != BPF_PROG_TYPE_KPROBE) ||
11642: (is_tracepoint && prog->type != BPF_PROG_TYPE_TRACEPOINT) ||
11643: (is_syscall_tp && prog->type != BPF_PROG_TYPE_TRACEPOINT))
11644: return -EINVAL;
11645:
11646: if (prog->type == BPF_PROG_TYPE_KPROBE && prog->sleepable && !is_uprobe)
11647: /* only uprobe programs are allowed to be sleepable */
11648: return -EINVAL;
11649:
11650: /* Kprobe override only works for kprobes, not uprobes. */
11651: if (prog->kprobe_override && !is_kprobe)
11652: return -EINVAL;
11653:
11654: /* Writing to context allowed only for uprobes. */
11655: if (prog->aux->kprobe_write_ctx && !is_uprobe)
11656: return -EINVAL;
11657:
11658: if (is_tracepoint || is_syscall_tp) {
11659: int off = trace_event_get_offsets(event->tp_event);
11660:
11661: if (prog->aux->max_ctx_offset > off)
11662: return -EACCES;
11663: }
11664:
11665: return perf_event_attach_bpf_prog(event, prog, bpf_cookie);
11666: }
SourceFile:kernel/events/core.c]
|
| 92/2 |
2026/05/17 11:43 |
llm |
expert |
0m
Model:
Tokens:
input: 65792
output: 24
thoughts: 939
Thoughts: **Analysis of a BPF Use-After-Free Vulnerability**
Okay, let's break this down. My initial thought was that the problem was related to how `perf_event_attach_bpf_prog` was saving the `prog` pointer, not the `link` pointer. This made me think `bpf_perf_link_lops` wouldn't need a `.dealloc_deferred` function because the link object itself isn't directly accessed by the tracepoint. The code seemed to suggest that only `bpf_raw_tp_link_lops` required it. The reason being that `bpf_probe_register` registers the `link` object as the data pointer for the tracepoint.
Then, I went back to re-examine `bpf_raw_tp_link_attach`. Sure enough, the function `bpf_probe_register` registers the `link` pointer, as I suspected:
```c
err = bpf_probe_register(link->btp, link);
```
And in the function `__bpf_trace_run`, the `link` object is passed as the first argument, so `__bpf_trace_run` accesses `link->link.prog`:
```c
void __bpf_trace_run(struct bpf_raw_tp_link *link, u64 *args)
```
This confirmed my suspicion. The problem is that the `link` is freed with `call_rcu` in `bpf_link_free`, not `call_srcu`, and that's the core of the use-after-free vulnerability.
So, the vulnerability is only reachable via `bpf_raw_tp_link_attach`. This function requires a BPF program type of `BPF_PROG_TYPE_RAW_TRACEPOINT`, `BPF_PROG_TYPE_RAW_TRACEPOINT_WRITABLE`, `BPF_PROG_TYPE_TRACING`, `BPF_PROG_TYPE_EXT`, or `BPF_PROG_TYPE_LSM`. Loading any of these requires `CAP_BPF` or `CAP_SYS_ADMIN`. And, as we know, unprivileged users cannot load those program types, nor can they create the necessary token to do so within a user namespace.
Let's do a quick summary:
* **Exploitable:** Yes, it's a use-after-free vulnerability, where the attacker can control the contents of the `prog` pointer, and can overwrite it to execute arbitrary code.
* **Denial of Service:** Yes, use-after-free leading to a kernel crash.
* **Accessible from Unprivileged Processes:** No, requires `CAP_BPF`.
* **Accessible from User Namespaces:** No, `CAP_BPF` is checked against the initial user namespace.
* **VM Guest Trigger:** No.
* **VM Host Trigger in Confidential Computing:** No.
* **Ethernet Network Trigger:** No.
* **Other Remote Trigger:** No.
* **Peripheral Trigger:** No.
* **Malicious Filesystem Trigger:** No.
But, I need to be sure. Is there another way to reach `bpf_raw_tp_link_attach`? What about `bpf_raw_tracepoint_open` and a `BPF_PROG_TYPE_SOCKET_FILTER` program?
```c
switch (prog->type) {
case BPF_PROG_TYPE_TRACING:
case BPF_PROG_TYPE_EXT:
case BPF_PROG_TYPE_LSM:
...
case BPF_PROG_TYPE_RAW_TRACEPOINT:
case BPF_PROG_TYPE_RAW_TRACEPOINT_WRITABLE:
...
default:
return -EINVAL;
}
```
Okay, it returns `-EINVAL` for `BPF_PROG_TYPE_SOCKET_FILTER`. So, even if an unprivileged user can load a `SOCKET_FILTER`, they can't use it to trigger this vulnerability.
Now, let me check other paths to `bpf_probe_register` again to be absolutely certain...
|
| 93/2 |
2026/05/17 11:44 |
tool |
codesearch-find-references |
0m
Args: map[Name:bpf_probe_register]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:bpf_raw_tp_link_attach SourceFile:kernel/bpf/syscall.c SourceLine:4281]] TruncatedOutput:false]
|
| 94/2 |
2026/05/17 11:44 |
llm |
expert |
0m
Model:
Tokens:
input: 67749
output: 66
thoughts: 363
Thoughts: **My Assessment of the BPF Raw Tracepoint Vulnerability**
Okay, let's break this down. My analysis points to a very specific vulnerability. It hinges solely on the `bpf_raw_tp_link_attach` function calling `bpf_probe_register`. This immediately narrows the scope considerably. The exploit's accessibility is effectively gated by `CAP_BPF` or `CAP_SYS_ADMIN` privileges. This is crucial; it means unprivileged users are completely out of the picture, and user namespaces are also protected.
To be absolutely certain, I'm verifying if `unprivileged_bpf_disabled` being set to zero makes a difference. Can an unprivileged user load `BPF_PROG_TYPE_RAW_TRACEPOINT`? The code specifically checks `bpf_cap` during `bpf_prog_load`, particularly for `BPF_PROG_TYPE_RAW_TRACEPOINT`. The check happens here:
```c
if (type != BPF_PROG_TYPE_SOCKET_FILTER &&
type != BPF_PROG_TYPE_CGROUP_SKB &&
!bpf_cap)
goto put_token;
```
It's clear, even with `unprivileged_bpf_disabled` set to zero, no dice. Unprivileged users can't load `RAW_TRACEPOINT` programs. Furthermore, `bpf_raw_tp_link_attach` prevents them from attaching `SOCKET_FILTER` or `CGROUP_SKB` to raw tracepoints.
Therefore, this is unequivocally *not* accessible to unprivileged users. I've confirmed that.
Let's summarize the key dimensions:
* Exploitable: **True**
* DenialOfService: **True**
* Unprivileged: **False**
* UserNamespace: **False**
* VMGuestTrigger: **False**
* VMHostTrigger: **False**
* NetworkTrigger: **False**
* RemoteTrigger: **False**
* PeripheralTrigger: **False**
* FilesystemTrigger: **False**
I'm now ready to use `set-results` and submit my analysis with this comprehensive justification.
|
| 95/2 |
2026/05/17 11:44 |
tool |
set-results |
0m
Args: map[DenialOfService:true Exploitable:true FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
Results: map[DenialOfService:true Exploitable:true FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
|
| 96/2 |
2026/05/17 11:44 |
llm |
expert |
0m
Model:
Tokens:
input: 68729
output: 793
thoughts: 363
Thoughts: **My Analysis of the BPF Tracepoint Use-After-Free Vulnerability**
Okay, the analysis is complete, and it's time to craft the response. The core issue, as I understand it, is a classic use-after-free vulnerability within the `__bpf_trace_run` function. Specifically, it involves the `struct bpf_raw_tp_link` structure.
The root cause hinges on a timing mismatch between RCU and SRCU. `bpf_link_free` employs `call_rcu` for freeing the `link` object. Meanwhile, the tracepoint execution itself, specifically within a function like `__do_trace_mm_page_alloc`, is protected by an SRCU critical section implemented via `tracepoint_srcu`. This is where the problem arises. Because SRCU read-side critical sections can be preempted and, crucially, outlive normal RCU grace periods, it's entirely possible for the `link` object to be freed *while* the tracepoint is still executing, leaving a dangling pointer. This results in the dreaded use-after-free when the code attempts to access `link->link.prog`.
From a practical perspective, this is exploitable. It could potentially be leveraged for privilege escalation, provided the attacker already possesses `CAP_BPF`, but they might be aiming to bypass security restrictions like lockdown or execute arbitrary kernel code. Alternatively, the vulnerability can definitely be weaponized for denial of service β a kernel crash.
However, and this is important, this isn't a vulnerability accessible to just *anyone*. Creating a raw tracepoint link via `bpf_raw_tp_link_attach` requires a BPF program of type `BPF_PROG_TYPE_RAW_TRACEPOINT` (or similar tracing types), which demands `CAP_BPF` or `CAP_SYS_ADMIN` privileges to load. Additionally, this is not an issue accessible from user namespaces because `CAP_BPF` is checked against the initial user namespace. Even creating a BPF token for delegating capabilities is not feasible because mounting `bpffs` with delegation options requires `CAP_SYS_ADMIN` in the initial user namespace.
Essentially, this narrows the attack surface significantly. It's not a remotely exploitable vulnerability via network, or other external vectors. No remote, peripheral, or filesystem vectors are in play, and it's also not something that can be triggered from within VM guests.
|
| 97/1 |
2026/05/17 11:44 |
action |
format-explanation |
0m
Results: map[Explanation:**Analysis of the Bug**
The bug report describes a KASAN slab-use-after-free in `__bpf_trace_run` when
accessing `link->link.prog`. The object being accessed is a `struct
bpf_raw_tp_link` which was allocated during `bpf_raw_tracepoint_open` and freed
via `bpf_link_release`.
**Root Cause:**
When a raw tracepoint BPF link is released, `bpf_link_free()` schedules the
deallocation of the `struct bpf_raw_tp_link` object using `call_rcu()` (which
waits for a normal RCU grace period). However, tracepoints (like
`trace_mm_page_alloc` in the report) are protected by SRCU (`tracepoint_srcu`)
via `guard(srcu_fast_notrace)(&tracepoint_srcu)`.
Because SRCU read-side critical sections can be preempted and can outlive normal
RCU grace periods, the normal RCU grace period can completeβand the `link`
object can be freedβwhile the tracepoint is still executing. When the tracepoint
code resumes and accesses `link->link.prog`, it hits a use-after-free.
**Exploitability:**
This is a classic use-after-free vulnerability. An attacker who can trigger the
bug can potentially reallocate the freed `struct bpf_raw_tp_link` object with
controlled data. By controlling the `prog` pointer, the attacker can hijack the
control flow when `bpf_prog_run(prog, args)` is called, leading to arbitrary
kernel code execution. It can also trivially cause a kernel panic (Denial of
Service).
**Privilege Requirements:**
To trigger this bug, a user must be able to load a BPF program and attach it to
a raw tracepoint.
1. **Unprivileged Processes:** Unprivileged users cannot load the required BPF
program types. `bpf_raw_tp_link_attach()` strictly requires programs of type
`BPF_PROG_TYPE_RAW_TRACEPOINT`, `BPF_PROG_TYPE_TRACING`, `BPF_PROG_TYPE_EXT`, or
`BPF_PROG_TYPE_LSM`. Loading any of these program types requires `CAP_BPF` or
`CAP_SYS_ADMIN` in the initial user namespace. (Even if
`unprivileged_bpf_disabled=0`, unprivileged users are restricted to
`SOCKET_FILTER` and `CGROUP_SKB`, which are rejected by
`bpf_raw_tp_link_attach()`).
2. **User Namespaces:** The capability checks for loading BPF programs
(`bpf_capable()`) check against the initial user namespace. While the kernel
supports BPF tokens to delegate capabilities to user namespaces, creating a BPF
token requires mounting `bpffs` with delegation options (e.g.,
`delegate_progs`). Parsing these delegation mount options explicitly requires
`capable(CAP_SYS_ADMIN)` in the initial user namespace. Thus, an unprivileged
user cannot set up a user namespace to bypass these restrictions.
**Other Triggers:**
This bug is strictly local and specific to the BPF subsystem. It cannot be
triggered via network traffic, remote interactions, malicious filesystems, USB
peripherals, or from within a VM guest.
**Conclusion:**
The bug is a highly exploitable memory corruption issue that can lead to
privilege escalation and denial of service. However, it acts as a
post-compromise/lockdown-bypass vulnerability rather than a primary entry point,
as it strictly requires `CAP_BPF` or `CAP_SYS_ADMIN` privileges to be triggered.]
|