| 0/0 |
2026/05/18 09:38 |
flow |
assessment-security |
9m
Results: map[DenialOfService:true Explanation:The bug report indicates a WARNING in `perf_event_throttle_group` triggered by
`lockdep_assert_event_ctx(leader)`.
### Root Cause Analysis
The `perf_event_throttle_group` function iterates over the `event->sibling_list`
of a perf event group using the `for_each_sibling_event` macro. This macro
includes a lockdep assertion that checks if `ctx->mutex` is held whenever
hardirqs are enabled.
For hardware perf events, overflows are handled in NMI or hardware interrupt
context where hardirqs are disabled, making it safe to iterate the list without
locks (as IPIs are blocked). However, for **software events** (e.g.,
`PERF_COUNT_SW_PAGE_FAULTS`) and **tracepoints**, overflows are handled
synchronously in the context of the triggering instruction. In these cases,
`perf_event_throttle_group` can be reached with **hardirqs enabled**.
Because hardirqs are enabled and no locks (`ctx->lock` or `ctx->mutex`) are
held, the iteration over `sibling_list` is vulnerable to concurrent
modifications via Inter-Processor Interrupts (IPIs).
### Impact: Denial of Service
If a malicious user creates a group of software events and triggers overflows
rapidly (e.g., by causing page faults in a loop), `perf_event_throttle_group`
will be invoked. Concurrently, the user can close the file descriptor of a
sibling event from another thread.
Closing the file descriptor invokes `perf_release` ->
`perf_event_release_kernel` -> `perf_remove_from_context`, which sends an IPI to
the CPU where the event is active. The IPI handler
(`__perf_remove_from_context`) will interrupt `perf_event_throttle_group`,
acquire `ctx->lock`, and call `perf_group_detach`.
`perf_group_detach` removes the sibling from the list using
`list_del_init(&sibling->sibling_list)`. This reinitializes the sibling's list
pointers to point to itself. When the IPI returns, `perf_event_throttle_group`
resumes its `list_for_each_entry` iteration. Because the `next` pointer now
points to the element itself, the loop will **infinite loop**, causing a hard
lockup on the CPU.
### Exploitability
This bug is **not exploitable for privilege escalation or memory corruption**.
Although it is a concurrent list modification, the `perf_event` structures are
freed via RCU (`call_rcu`). Since the software event overflow handler (e.g.,
`do_perf_sw_event`) runs inside an RCU read-side critical section
(`rcu_read_lock()`), the CPU getting stuck in the infinite loop will prevent the
RCU grace period from ever ending. Consequently, the memory is never freed,
precluding any Use-After-Free (UAF) scenarios. It strictly results in a Denial
of Service (DoS).
### Reachability
The bug is reachable by **unprivileged users** (and within user namespaces).
While unprivileged users cannot create tracepoint events under the default
`sysctl_perf_event_paranoid = 2` configuration, they are fully permitted to
create and group `PERF_TYPE_SOFTWARE` events (like page faults or context
switches) which are sufficient to trigger the vulnerable code path. Exploitable:false FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:true UserNamespace:true VMGuestTrigger:false VMHostTrigger:false]
|
| 1/1 |
2026/05/18 09:38 |
action |
syz-repro-to-c-repro |
0m
Results: map[SimplifiedCRepro:// autogenerated by syzkaller (https://github.com/google/syzkaller)
#define _GNU_SOURCE
#include <endian.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/mount.h>
#include <sys/syscall.h>
#include <sys/types.h>
#include <unistd.h>
#define BITMASK(bf_off,bf_len) (((1ull << (bf_len)) - 1) << (bf_off))
#define STORE_BY_BITMASK(type,htobe,addr,val,bf_off,bf_len) *(type*)(addr) = htobe((htobe(*(type*)(addr)) & ~BITMASK((bf_off), (bf_len))) | (((type)(val) << (bf_off)) & BITMASK((bf_off), (bf_len))))
uint64_t r[1] = {0xffffffffffffffff};
int main(void)
{
syscall(__NR_mmap, /*addr=*/0x1ffffffff000ul, /*len=*/0x1000ul, /*prot=*/0ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
syscall(__NR_mmap, /*addr=*/0x200000000000ul, /*len=*/0x1000000ul, /*prot=PROT_WRITE|PROT_READ|PROT_EXEC*/7ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
syscall(__NR_mmap, /*addr=*/0x200001000000ul, /*len=*/0x1000ul, /*prot=*/0ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
const char* reason;
(void)reason;
intptr_t res = 0;
if (write(1, "executing program\n", sizeof("executing program\n") - 1)) {}
// perf_event_open arguments: [
// attr: ptr[in, perf_event_attr] {
// perf_event_attr {
// type: perf_event_type = 0x2 (4 bytes)
// size: len = 0x80 (4 bytes)
// config0: int8 = 0xeb (1 bytes)
// config1: int8 = 0x0 (1 bytes)
// config2: int8 = 0x0 (1 bytes)
// config3: int8 = 0x0 (1 bytes)
// config4: const = 0x0 (4 bytes)
// sample_freq: int64 = 0x1 (8 bytes)
// sample_type: perf_sample_type = 0x4000 (8 bytes)
// read_format: perf_read_format = 0x2 (8 bytes)
// disabled: int64 = 0x0 (0 bytes)
// inherit: int64 = 0x0 (0 bytes)
// pinned: int64 = 0x0 (0 bytes)
// exclusive: int64 = 0x0 (0 bytes)
// exclude_user: int64 = 0x0 (0 bytes)
// exclude_kernel: int64 = 0x0 (0 bytes)
// exclude_hv: int64 = 0x0 (0 bytes)
// exclude_idle: int64 = 0x0 (0 bytes)
// mmap: int64 = 0x0 (0 bytes)
// comm: int64 = 0x0 (0 bytes)
// freq: int64 = 0x0 (0 bytes)
// inherit_stat: int64 = 0x0 (0 bytes)
// enable_on_exec: int64 = 0x0 (0 bytes)
// task: int64 = 0x0 (0 bytes)
// watermark: int64 = 0x0 (0 bytes)
// precise_ip: int64 = 0x0 (0 bytes)
// mmap_data: int64 = 0x0 (0 bytes)
// sample_id_all: int64 = 0x0 (0 bytes)
// exclude_host: int64 = 0x0 (0 bytes)
// exclude_guest: int64 = 0x0 (0 bytes)
// exclude_callchain_kernel: int64 = 0x0 (0 bytes)
// exclude_callchain_user: int64 = 0x0 (0 bytes)
// mmap2: int64 = 0x0 (0 bytes)
// comm_exec: int64 = 0x0 (0 bytes)
// use_clockid: int64 = 0x0 (0 bytes)
// context_switch: int64 = 0x0 (0 bytes)
// write_backward: int64 = 0x0 (0 bytes)
// namespaces: int64 = 0x0 (0 bytes)
// ksymbol: int64 = 0x0 (0 bytes)
// bpf_event: int64 = 0x0 (0 bytes)
// aux_output: int64 = 0x0 (0 bytes)
// cgroup: int64 = 0x0 (0 bytes)
// text_poke: int64 = 0x0 (0 bytes)
// build_id: int64 = 0x0 (0 bytes)
// inherit_thread: int64 = 0x0 (0 bytes)
// remove_on_exec: int64 = 0x0 (0 bytes)
// sigtrap: int64 = 0x0 (0 bytes)
// __reserved_1: const = 0x0 (8 bytes)
// wakeup_events: int32 = 0x0 (4 bytes)
// bp_type: perf_bp_type = 0x0 (4 bytes)
// bp_config: union perf_bp_config {
// perf_bp: perf_bp {
// bp_addr: nil
// bp_len: perf_bp_lens = 0xb (8 bytes)
// }
// }
// branch_sample_type: perf_branch_sample_type = 0x0 (8 bytes)
// sample_regs_user: int64 = 0x0 (8 bytes)
// sample_stack_user: int32 = 0x0 (4 bytes)
// clockid: clock_type = 0x3 (4 bytes)
// sample_regs_intr: int64 = 0x3 (8 bytes)
// aux_watermark: int32 = 0x0 (4 bytes)
// sample_max_stack: int16 = 0x0 (2 bytes)
// __reserved_2: const = 0x0 (2 bytes)
// aux_sample_size: int32 = 0x0 (4 bytes)
// __reserved_3: const = 0x0 (4 bytes)
// sig_data: int64 = 0x0 (8 bytes)
// }
// }
// pid: pid (resource)
// cpu: intptr = 0xffffffffffffffff (8 bytes)
// group: fd_perf (resource)
// flags: perf_flags = 0x0 (8 bytes)
// ]
// returns fd_perf
*(uint32_t*)0x200000000040 = 2;
*(uint32_t*)0x200000000044 = 0x80;
*(uint8_t*)0x200000000048 = 0xeb;
*(uint8_t*)0x200000000049 = 0;
*(uint8_t*)0x20000000004a = 0;
*(uint8_t*)0x20000000004b = 0;
*(uint32_t*)0x20000000004c = 0;
*(uint64_t*)0x200000000050 = 1;
*(uint64_t*)0x200000000058 = 0x4000;
*(uint64_t*)0x200000000060 = 2;
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 0, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 1, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 2, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 3, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 4, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 5, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 6, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 7, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 8, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 9, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 10, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 11, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 12, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 13, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 14, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 15, 2);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 17, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 18, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 19, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 20, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 21, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 22, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 23, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 24, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 25, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 26, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 27, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 28, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 29, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 30, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 31, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 32, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 33, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 34, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 35, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 36, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 37, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 38, 26);
*(uint32_t*)0x200000000070 = 0;
*(uint32_t*)0x200000000074 = 0;
*(uint64_t*)0x200000000078 = 0;
*(uint64_t*)0x200000000080 = 0xb;
*(uint64_t*)0x200000000088 = 0;
*(uint64_t*)0x200000000090 = 0;
*(uint32_t*)0x200000000098 = 0;
*(uint32_t*)0x20000000009c = 3;
*(uint64_t*)0x2000000000a0 = 3;
*(uint32_t*)0x2000000000a8 = 0;
*(uint16_t*)0x2000000000ac = 0;
*(uint16_t*)0x2000000000ae = 0;
*(uint32_t*)0x2000000000b0 = 0;
*(uint32_t*)0x2000000000b4 = 0;
*(uint64_t*)0x2000000000b8 = 0;
res = syscall(__NR_perf_event_open, /*attr=*/0x200000000040ul, /*pid=*/0, /*cpu=*/(intptr_t)-1, /*group=*/(intptr_t)-1, /*flags=*/0ul);
if (res != -1)
r[0] = res;
// ioctl$PERF_EVENT_IOC_PERIOD arguments: [
// fd: fd_perf (resource)
// cmd: const = 0x40082404 (4 bytes)
// period: ptr[in, int64] {
// int64 = 0x2 (8 bytes)
// }
// ]
*(uint64_t*)0x200000000300 = 2;
syscall(__NR_ioctl, /*fd=*/r[0], /*cmd=*/0x40082404, /*period=*/0x200000000300ul);
return 0;
}
]
|
| 2/1 |
2026/05/18 09:38 |
action |
kernel-checkouter |
0m
Results: map[KernelSrc:/app/workdir/cache/src/87e8e956d1b362a2ace926b3022ad0a492ea02a8]
|
| 3/1 |
2026/05/18 09:38 |
action |
kernel-builder |
0m
Results: map[KernelObj:/app/workdir/cache/build/24adcabfa4f1d827ee879ddd0fbecd6386ab0cba]
|
| 4/1 |
2026/05/18 09:38 |
action |
codesearch-prepare |
0m
Results: map[Index:codesearch-index]
|
| 5/1 |
2026/05/18 09:38 |
agent |
expert |
8m
Model:
Results: map[DenialOfService:true Exploitable:false FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:true UserNamespace:true VMGuestTrigger:false VMHostTrigger:false]
Instruction:
You are an experienced Linux kernel security engineer. Your task is to analyze given kernel bug report
and determine its security impact based on the following dimensions.
Use the provided tools to examine the source code, check for capability checks (e.g., capable(), ns_capable()),
and understand the nature of the bug. Analyze the given kernel build and configuration.
You can check the kernel config by grepping ".config" file; you can check kernel cmdline by greeping
".config" file for "CONFIG_CMDLINE=". Assume sysctl parameters have default values.
But analyze for the corresponding production build w/o debugging tools enabled (like KASAN, KMSAN, UBSAN).
Don't make assumptions; verify them with source code access. Try different strategies when analyzing the bug:
- think of ways in which the vulnerable code is unreachable
- or the other way around: try to come up with different ideas of how an unprivileged user can reach the bug
If still unsure err on the side of the bug being non-exploitable/not-accessible.
In the final reply, provide a reasoning for your assessment.
Analysis dimensions:
* Exploitable:
Determine if the bug can result in memory corruption or elevated privileges.
Memory safety issues are almost always exploitable (KASAN or UBSAN reports for use-after-free, out-of-bounds;
refcounting issues, corrupted lists, etc). When kernel is crashing on a completly wild pointer access
(e.g. user-space address, or non-canonical address, but not on NULL or address corresponding to KASAN shadow
for NULL address), including both data accesses and control tranfers, that's also usually implies possibility
of exploitation. Such reports usually say "unable to handle kernel paging request".
Uses of uninitialized values detected by KMSAN may be exploitable b/c attacker frequently can affect uninit
values with spraying techniques. However, for these exploitabability depends on how exactly the uninit value
is used in the code, and what it affects.
Think of what happens after the bug is triggered. Some bugs cause kernel panic and halt execution,
they are harder to exploit. For example, BUG reports halts the kernel. However, WARNING reports don't halt
execution in production builds. Debug bug detection tools (like KASAN, KMSAN, KCSAN, UBSAN) are also not enabled
in production builds, so attacker can freely exploit these bugs w/o being detected by these tools.
If you see an integer overflow, think how the overflowed value used later (if it's used as allocation size,
or an array index). If you see an out-of-bounds read, think if it's followed by an out-of-bounds write as well.
Some KCSAN data-races may be exploitable by skilled attackers as well. Think what data structures got corrupted
as the result of data races and how. However, note that kernel has lots of "benign" data races that don't lead
to any runtime misbehavior at all.
* Denial Of Service:
Determine if the bug can result in denial-of-service. Most bugs can, since they cause system crash,
hangs, deadlocks, or resource leaks. This is mostly applicable to WARNING bugs that won't cause system crash
in production. For these think what will be consequences of the violation of the kernel assumptions flagged
by the WARNING. In some cases the unexpected condition is also properly handled by the normal control flow
(e.g. with "if (WARN_ON(...))"), these won't cause denial-of-service. If the condition is not handled,
then it may or may not cause denial-of-service.
* Accessible From Unprivileged Processes:
Determine if the bug can be reached from a typical (non-root) user process that does NOT have any special capabilities
(like CAP_SYS_ADMIN, CAP_NET_ADMIN, CAP_NET_RAW, CAP_PERFMON) or access to device nodes restricted to root.
Assume that unprivileged_bpf_disabled=1, that is eBPF loading is not accessible. However, cBPF (classical BPF)
is still accessible to non-root processes.
Assume that user namespaces are not accessible, that is, the process cannot get the mentioned capabilities even
within a new user namespace (checked by ns_capable() function in the kernel sources).
* Accessible From User Namespaces:
Determine if the bug can be reached within a user-namespace where the process has all capabilities
(including CAP_SYS_ADMIN, CAP_NET_ADMIN, CAP_NET_RAW, CAP_PERFMON). Such capabilities are checked with ns_capable()
function in the kernel sources.
* VM Guest Trigger:
Determine if the bug can be triggered from the context of a typical KVM guest (e.g., set up by a QEMU VMM).
Consider accesses to standard Linux host paravirtualized features (virtio-blk, virtio-net, etc.),
and handling of VM exits in the KVM code.
* VM Host Trigger in The Confidetial Computing Context:
Determine if the bug can be triggered in a confidential computing guest kernel from the context of a KVM host.
Consider access to standard Linux guest paravirtualized features (virtio-blk, virtio-net, etc.).
* Ethernet Network Trigger:
Determine if the bug can be triggered by processing ingress network Ethernet traffic, either directly (network stack)
or via drivers exposed to network data.
* Other Remote Trigger:
Determine if the bug can be triggered by processing remote traffic other than Ethernet (Wifi, Bluetooth, NFC, etc).
* Peripheral Trigger:
Determine if the bug can be triggered via an untrusted peripheral device that can be physically plugged
into a system, such as a USB device or a niche hardware driver handling external hardware inputs.
This is particularly important for mobile and desktop environments where users can plug in unknown devices.
* Malicious Filesystem Trigger:
Determine if the bug can be triggered by the kernel mounting and parsing a malicious filesystem image.
This is highly critical for Desktop and Mobile environments where external media or downloaded images
might be auto-mounted.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt:
The kernel bug report is:
------------[ cut here ]------------
WARNING: CPU: 1 PID: 7301 at kernel/events/core.c:2703 perf_event_throttle_group+0x375/0x3d0 kernel/events/core.c:2703
Modules linked in:
CPU: 1 UID: 0 PID: 7301 Comm: syz.0.1188 Not tainted syzkaller #0 PREEMPT(full)
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 10/25/2025
RIP: 0010:perf_event_throttle_group+0x375/0x3d0 kernel/events/core.c:2703
Code: 97 32 00 e9 ec fe ff ff e8 08 83 cc ff eb 05 e8 01 83 cc ff 5b 41 5c 41 5d 41 5e 41 5f 5d e9 d2 41 56 09 cc e8 ec 82 cc ff 90 <0f> 0b 90 e9 54 fe ff ff 44 89 f1 80 e1 07 80 c1 03 38 c1 0f 8c c9
RSP: 0018:ffffc900033b7490 EFLAGS: 00010293
RAX: ffffffff81f383d4 RBX: ffff888079cf1900 RCX: ffff8880203ddac0
RDX: 0000000000000000 RSI: 0000000000000000 RDI: 0000000000000001
RBP: 0000000000000000 R08: ffffc900033b7137 R09: 0000000000000000
R10: ffffc900033b7120 R11: fffff52000676e27 R12: ffff888079cf1b18
R13: dffffc0000000000 R14: ffff888079cf1b38 R15: ffffffff8dff1ab8
FS: 000055556f9df500(0000) GS:ffff888126238000(0000) knlGS:0000000000000000
CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 0000001b33c63fff CR3: 0000000071556000 CR4: 00000000003526f0
Call Trace:
<TASK>
__perf_event_account_interrupt+0x23c/0x250 kernel/events/core.c:10180
__perf_event_overflow+0x105/0xe70 kernel/events/core.c:10327
perf_swevent_overflow kernel/events/core.c:10467 [inline]
perf_swevent_event+0x530/0x5e0 kernel/events/core.c:10505
perf_tp_event+0x4f6/0x1380 kernel/events/core.c:11012
perf_trace_run_bpf_submit+0xee/0x170 kernel/events/core.c:10936
do_perf_trace_lock_acquire include/trace/events/lock.h:24 [inline]
perf_trace_lock_acquire+0x335/0x410 include/trace/events/lock.h:24
__do_trace_lock_acquire include/trace/events/lock.h:24 [inline]
trace_lock_acquire include/trace/events/lock.h:24 [inline]
lock_acquire+0x311/0x360 kernel/locking/lockdep.c:5831
rcu_lock_acquire include/linux/rcupdate.h:331 [inline]
rcu_read_lock include/linux/rcupdate.h:867 [inline]
class_rcu_constructor include/linux/rcupdate.h:1195 [inline]
futex_hash+0x5d/0x2d0 kernel/futex/core.c:308
class_hb_constructor kernel/futex/futex.h:240 [inline]
futex_wake+0x161/0x560 kernel/futex/waitwake.c:172
do_futex+0x395/0x420 kernel/futex/syscalls.c:135
__do_sys_futex kernel/futex/syscalls.c:207 [inline]
__se_sys_futex+0x36f/0x400 kernel/futex/syscalls.c:188
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0xfa/0xfa0 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7f1b5b78f6c9
Code: ff ff c3 66 2e 0f 1f 84 00 00 00 00 00 0f 1f 40 00 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 c7 c1 a8 ff ff ff f7 d8 64 89 01 48
RSP: 002b:00007ffd23d3b728 EFLAGS: 00000246 ORIG_RAX: 00000000000000ca
RAX: ffffffffffffffda RBX: 00007f1b5b9e5fa8 RCX: 00007f1b5b78f6c9
RDX: 00000000000f4240 RSI: 0000000000000081 RDI: 00007f1b5b9e5fac
RBP: 0000000000000000 R08: 3fffffffffffffff R09: 0000000223d3ba1f
R10: 0000000000000000 R11: 0000000000000246 R12: 00007f1b5b9e5fac
R13: 00007f1b5b9e5fa0 R14: 0000000000000843 R15: 0000000000000003
</TASK>
It is reproducible with the followint program.
Keep in mind that it may lack the precise threading, sandboxing, and some arguments of a working reproducer.
But it should give an idea of the involved syscalls.
// autogenerated by syzkaller (https://github.com/google/syzkaller)
#define _GNU_SOURCE
#include <endian.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/mount.h>
#include <sys/syscall.h>
#include <sys/types.h>
#include <unistd.h>
#define BITMASK(bf_off,bf_len) (((1ull << (bf_len)) - 1) << (bf_off))
#define STORE_BY_BITMASK(type,htobe,addr,val,bf_off,bf_len) *(type*)(addr) = htobe((htobe(*(type*)(addr)) & ~BITMASK((bf_off), (bf_len))) | (((type)(val) << (bf_off)) & BITMASK((bf_off), (bf_len))))
uint64_t r[1] = {0xffffffffffffffff};
int main(void)
{
syscall(__NR_mmap, /*addr=*/0x1ffffffff000ul, /*len=*/0x1000ul, /*prot=*/0ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
syscall(__NR_mmap, /*addr=*/0x200000000000ul, /*len=*/0x1000000ul, /*prot=PROT_WRITE|PROT_READ|PROT_EXEC*/7ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
syscall(__NR_mmap, /*addr=*/0x200001000000ul, /*len=*/0x1000ul, /*prot=*/0ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
const char* reason;
(void)reason;
intptr_t res = 0;
if (write(1, "executing program\n", sizeof("executing program\n") - 1)) {}
// perf_event_open arguments: [
// attr: ptr[in, perf_event_attr] {
// perf_event_attr {
// type: perf_event_type = 0x2 (4 bytes)
// size: len = 0x80 (4 bytes)
// config0: int8 = 0xeb (1 bytes)
// config1: int8 = 0x0 (1 bytes)
// config2: int8 = 0x0 (1 bytes)
// config3: int8 = 0x0 (1 bytes)
// config4: const = 0x0 (4 bytes)
// sample_freq: int64 = 0x1 (8 bytes)
// sample_type: perf_sample_type = 0x4000 (8 bytes)
// read_format: perf_read_format = 0x2 (8 bytes)
// disabled: int64 = 0x0 (0 bytes)
// inherit: int64 = 0x0 (0 bytes)
// pinned: int64 = 0x0 (0 bytes)
// exclusive: int64 = 0x0 (0 bytes)
// exclude_user: int64 = 0x0 (0 bytes)
// exclude_kernel: int64 = 0x0 (0 bytes)
// exclude_hv: int64 = 0x0 (0 bytes)
// exclude_idle: int64 = 0x0 (0 bytes)
// mmap: int64 = 0x0 (0 bytes)
// comm: int64 = 0x0 (0 bytes)
// freq: int64 = 0x0 (0 bytes)
// inherit_stat: int64 = 0x0 (0 bytes)
// enable_on_exec: int64 = 0x0 (0 bytes)
// task: int64 = 0x0 (0 bytes)
// watermark: int64 = 0x0 (0 bytes)
// precise_ip: int64 = 0x0 (0 bytes)
// mmap_data: int64 = 0x0 (0 bytes)
// sample_id_all: int64 = 0x0 (0 bytes)
// exclude_host: int64 = 0x0 (0 bytes)
// exclude_guest: int64 = 0x0 (0 bytes)
// exclude_callchain_kernel: int64 = 0x0 (0 bytes)
// exclude_callchain_user: int64 = 0x0 (0 bytes)
// mmap2: int64 = 0x0 (0 bytes)
// comm_exec: int64 = 0x0 (0 bytes)
// use_clockid: int64 = 0x0 (0 bytes)
// context_switch: int64 = 0x0 (0 bytes)
// write_backward: int64 = 0x0 (0 bytes)
// namespaces: int64 = 0x0 (0 bytes)
// ksymbol: int64 = 0x0 (0 bytes)
// bpf_event: int64 = 0x0 (0 bytes)
// aux_output: int64 = 0x0 (0 bytes)
// cgroup: int64 = 0x0 (0 bytes)
// text_poke: int64 = 0x0 (0 bytes)
// build_id: int64 = 0x0 (0 bytes)
// inherit_thread: int64 = 0x0 (0 bytes)
// remove_on_exec: int64 = 0x0 (0 bytes)
// sigtrap: int64 = 0x0 (0 bytes)
// __reserved_1: const = 0x0 (8 bytes)
// wakeup_events: int32 = 0x0 (4 bytes)
// bp_type: perf_bp_type = 0x0 (4 bytes)
// bp_config: union perf_bp_config {
// perf_bp: perf_bp {
// bp_addr: nil
// bp_len: perf_bp_lens = 0xb (8 bytes)
// }
// }
// branch_sample_type: perf_branch_sample_type = 0x0 (8 bytes)
// sample_regs_user: int64 = 0x0 (8 bytes)
// sample_stack_user: int32 = 0x0 (4 bytes)
// clockid: clock_type = 0x3 (4 bytes)
// sample_regs_intr: int64 = 0x3 (8 bytes)
// aux_watermark: int32 = 0x0 (4 bytes)
// sample_max_stack: int16 = 0x0 (2 bytes)
// __reserved_2: const = 0x0 (2 bytes)
// aux_sample_size: int32 = 0x0 (4 bytes)
// __reserved_3: const = 0x0 (4 bytes)
// sig_data: int64 = 0x0 (8 bytes)
// }
// }
// pid: pid (resource)
// cpu: intptr = 0xffffffffffffffff (8 bytes)
// group: fd_perf (resource)
// flags: perf_flags = 0x0 (8 bytes)
// ]
// returns fd_perf
*(uint32_t*)0x200000000040 = 2;
*(uint32_t*)0x200000000044 = 0x80;
*(uint8_t*)0x200000000048 = 0xeb;
*(uint8_t*)0x200000000049 = 0;
*(uint8_t*)0x20000000004a = 0;
*(uint8_t*)0x20000000004b = 0;
*(uint32_t*)0x20000000004c = 0;
*(uint64_t*)0x200000000050 = 1;
*(uint64_t*)0x200000000058 = 0x4000;
*(uint64_t*)0x200000000060 = 2;
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 0, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 1, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 2, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 3, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 4, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 5, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 6, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 7, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 8, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 9, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 10, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 11, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 12, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 13, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 14, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 15, 2);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 17, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 18, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 19, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 20, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 21, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 22, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 23, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 24, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 25, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 26, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 27, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 28, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 29, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 30, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 31, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 32, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 33, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 34, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 35, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 36, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 37, 1);
STORE_BY_BITMASK(uint64_t, , 0x200000000068, 0, 38, 26);
*(uint32_t*)0x200000000070 = 0;
*(uint32_t*)0x200000000074 = 0;
*(uint64_t*)0x200000000078 = 0;
*(uint64_t*)0x200000000080 = 0xb;
*(uint64_t*)0x200000000088 = 0;
*(uint64_t*)0x200000000090 = 0;
*(uint32_t*)0x200000000098 = 0;
*(uint32_t*)0x20000000009c = 3;
*(uint64_t*)0x2000000000a0 = 3;
*(uint32_t*)0x2000000000a8 = 0;
*(uint16_t*)0x2000000000ac = 0;
*(uint16_t*)0x2000000000ae = 0;
*(uint32_t*)0x2000000000b0 = 0;
*(uint32_t*)0x2000000000b4 = 0;
*(uint64_t*)0x2000000000b8 = 0;
res = syscall(__NR_perf_event_open, /*attr=*/0x200000000040ul, /*pid=*/0, /*cpu=*/(intptr_t)-1, /*group=*/(intptr_t)-1, /*flags=*/0ul);
if (res != -1)
r[0] = res;
// ioctl$PERF_EVENT_IOC_PERIOD arguments: [
// fd: fd_perf (resource)
// cmd: const = 0x40082404 (4 bytes)
// period: ptr[in, int64] {
// int64 = 0x2 (8 bytes)
// }
// ]
*(uint64_t*)0x200000000300 = 2;
syscall(__NR_ioctl, /*fd=*/r[0], /*cmd=*/0x40082404, /*period=*/0x200000000300ul);
return 0;
}
Reply: The bug report indicates a WARNING in `perf_event_throttle_group` triggered by `lockdep_assert_event_ctx(leader)`.
### Root Cause Analysis
The `perf_event_throttle_group` function iterates over the `event->sibling_list` of a perf event group using the `for_each_sibling_event` macro. This macro includes a lockdep assertion that checks if `ctx->mutex` is held whenever hardirqs are enabled.
For hardware perf events, overflows are handled in NMI or hardware interrupt context where hardirqs are disabled, making it safe to iterate the list without locks (as IPIs are blocked). However, for **software events** (e.g., `PERF_COUNT_SW_PAGE_FAULTS`) and **tracepoints**, overflows are handled synchronously in the context of the triggering instruction. In these cases, `perf_event_throttle_group` can be reached with **hardirqs enabled**.
Because hardirqs are enabled and no locks (`ctx->lock` or `ctx->mutex`) are held, the iteration over `sibling_list` is vulnerable to concurrent modifications via Inter-Processor Interrupts (IPIs).
### Impact: Denial of Service
If a malicious user creates a group of software events and triggers overflows rapidly (e.g., by causing page faults in a loop), `perf_event_throttle_group` will be invoked. Concurrently, the user can close the file descriptor of a sibling event from another thread.
Closing the file descriptor invokes `perf_release` -> `perf_event_release_kernel` -> `perf_remove_from_context`, which sends an IPI to the CPU where the event is active. The IPI handler (`__perf_remove_from_context`) will interrupt `perf_event_throttle_group`, acquire `ctx->lock`, and call `perf_group_detach`.
`perf_group_detach` removes the sibling from the list using `list_del_init(&sibling->sibling_list)`. This reinitializes the sibling's list pointers to point to itself. When the IPI returns, `perf_event_throttle_group` resumes its `list_for_each_entry` iteration. Because the `next` pointer now points to the element itself, the loop will **infinite loop**, causing a hard lockup on the CPU.
### Exploitability
This bug is **not exploitable for privilege escalation or memory corruption**.
Although it is a concurrent list modification, the `perf_event` structures are freed via RCU (`call_rcu`). Since the software event overflow handler (e.g., `do_perf_sw_event`) runs inside an RCU read-side critical section (`rcu_read_lock()`), the CPU getting stuck in the infinite loop will prevent the RCU grace period from ever ending. Consequently, the memory is never freed, precluding any Use-After-Free (UAF) scenarios. It strictly results in a Denial of Service (DoS).
### Reachability
The bug is reachable by **unprivileged users** (and within user namespaces). While unprivileged users cannot create tracepoint events under the default `sysctl_perf_event_paranoid = 2` configuration, they are fully permitted to create and group `PERF_TYPE_SOFTWARE` events (like page faults or context switches) which are sufficient to trigger the vulnerable code path.
|
| 6/2 |
2026/05/18 09:38 |
llm |
expert |
0m
Model:
Tokens:
input: 10786
output: 43
thoughts: 86
Thoughts: Here's my analysis of the situation:
**Investigating a Kernel Performance Event Throttle Warning**
Okay, so I've got a `WARNING` flag showing up within `perf_event_throttle_group`, and the offending line is smack-dab in `kernel/events/core.c:2703`. My initial thought is to figure out the root cause. It looks like `syzkaller` is triggering this with a `perf_event_open` call followed by an `ioctl` for `PERF_EVENT_IOC_PERIOD`. That's my starting point. Now I need to dive into the code around that line in `core.c` and understand the conditions that lead to this warning being generated. I need to see *why* `syzkaller`'s interaction is causing a problem with the performance event throttling. I'll be looking for possible race conditions, unexpected event configurations, or perhaps a bug in how periods are being handled within the context of throttle groups. It's time to dig into the details and find out exactly what's going on.
|
| 7/2 |
2026/05/18 09:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_event_throttle_group]
Results: map[SourceCode:2698: static void perf_event_throttle_group(struct perf_event *event)
2699: {
2700: struct perf_event *sibling, *leader = event->group_leader;
2701:
2702: perf_event_throttle(leader);
2703: for_each_sibling_event(sibling, leader)
2704: perf_event_throttle(sibling);
2705: }
SourceFile:kernel/events/core.c]
|
| 8/2 |
2026/05/18 09:38 |
llm |
expert |
0m
Model:
Tokens:
input: 11296
output: 43
thoughts: 46
Thoughts: Okay, let's see, the report points to line 2703. Hmm, that's where we find the call to `for_each_sibling_event(sibling, leader)`. Right, I need to jog my memory on what that function actually *does*. Let me quickly re-familiarize myself with the purpose and implementation of `for_each_sibling_event`. I'm going to quickly recall its definition and internal logic to fully grasp its impact in this context. It's time to refresh the details of that specific function call.
|
| 9/2 |
2026/05/18 09:38 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:for_each_sibling_event]
Results: map[SourceCode: SourceFile:]
|
| 10/2 |
2026/05/18 09:38 |
llm |
expert |
0m
Model:
Tokens:
input: 11524
output: 21
thoughts: 16
Thoughts: Here's my thought process, as if I were thinking it:
**Digging into the Code with grep**
Alright, I need to track down where `for_each_sibling_event` is being used. Given my experience, grep is definitely the tool for this. It's quick, powerful, and perfect for navigating a codebase. Let's fire it up and see where this function call pops up. I'm expecting a few hits, but it will get me on my way. This should be a pretty straightforward search. Time to get some results and start digging!
|
| 11/2 |
2026/05/18 09:38 |
tool |
grepper |
0m
Args: map[Expression:for_each_sibling_event]
Results: map[Output:.clang-format=71=ForEachMacros:
--
.clang-format-449- - 'for_each_sgtable_sg'
.clang-format:450: - 'for_each_sibling_event'
.clang-format-451- - 'for_each_sta_active_link'
--
arch/alpha/kernel/perf_event.c=340=static int collect_events(struct perf_event *group, int max_count,
--
arch/alpha/kernel/perf_event.c-353- }
arch/alpha/kernel/perf_event.c:354: for_each_sibling_event(pe, group) {
arch/alpha/kernel/perf_event.c-355- if (!is_software_event(pe) && pe->state != PERF_EVENT_STATE_OFF) {
--
arch/arm/mach-imx/mmdc.c=256=static bool mmdc_pmu_group_is_valid(struct perf_event *event)
--
arch/arm/mach-imx/mmdc.c-269-
arch/arm/mach-imx/mmdc.c:270: for_each_sibling_event(sibling, leader) {
arch/arm/mach-imx/mmdc.c-271- if (!mmdc_pmu_group_event_is_valid(sibling, pmu, &counter_mask))
--
arch/arm/mm/cache-l2x0-pmu.c=272=static bool l2x0_pmu_group_is_valid(struct perf_event *event)
--
arch/arm/mm/cache-l2x0-pmu.c-283-
arch/arm/mm/cache-l2x0-pmu.c:284: for_each_sibling_event(sibling, leader) {
arch/arm/mm/cache-l2x0-pmu.c-285- if (sibling->pmu == pmu)
--
arch/loongarch/kernel/perf_event.c=629=static int validate_group(struct perf_event *event)
--
arch/loongarch/kernel/perf_event.c-638-
arch/loongarch/kernel/perf_event.c:639: for_each_sibling_event(sibling, leader) {
arch/loongarch/kernel/perf_event.c-640- if (loongarch_pmu_alloc_counter(&fake_cpuc, &sibling->hw) < 0)
--
arch/mips/kernel/perf_event_mipsxx.c=760=static int validate_group(struct perf_event *event)
--
arch/mips/kernel/perf_event_mipsxx.c-769-
arch/mips/kernel/perf_event_mipsxx.c:770: for_each_sibling_event(sibling, leader) {
arch/mips/kernel/perf_event_mipsxx.c-771- if (mipsxx_pmu_alloc_counter(&fake_cpuc, &sibling->hw) < 0)
--
arch/powerpc/perf/core-book3s.c=1590=static int collect_events(struct perf_event *group, int max_count,
--
arch/powerpc/perf/core-book3s.c-1603- }
arch/powerpc/perf/core-book3s.c:1604: for_each_sibling_event(event, group) {
arch/powerpc/perf/core-book3s.c-1605- if (event->pmu->task_ctx_nr == perf_hw_context &&
--
arch/powerpc/perf/core-fsl-emb.c=251=static int collect_events(struct perf_event *group, int max_count,
--
arch/powerpc/perf/core-fsl-emb.c-262- }
arch/powerpc/perf/core-fsl-emb.c:263: for_each_sibling_event(event, group) {
arch/powerpc/perf/core-fsl-emb.c-264- if (!is_software_event(event) &&
--
arch/sparc/kernel/perf_event.c=1345=static int collect_events(struct perf_event *group, int max_count,
--
arch/sparc/kernel/perf_event.c-1358- }
arch/sparc/kernel/perf_event.c:1359: for_each_sibling_event(event, group) {
arch/sparc/kernel/perf_event.c-1360- if (!is_software_event(event) &&
--
arch/x86/events/amd/ibs.c=255=static int validate_group(struct perf_event *event)
--
arch/x86/events/amd/ibs.c-264-
arch/x86/events/amd/ibs.c:265: for_each_sibling_event(sibling, event->group_leader) {
arch/x86/events/amd/ibs.c-266- if (sibling->pmu == event->pmu)
--
arch/x86/events/core.c=1173=static int collect_events(struct cpu_hw_events *cpuc, struct perf_event *leader, bool dogrp)
--
arch/x86/events/core.c-1212-
arch/x86/events/core.c:1213: for_each_sibling_event(event, leader) {
arch/x86/events/core.c-1214- if (!is_x86_event(event) || event->state <= PERF_EVENT_STATE_OFF)
--
arch/x86/events/core.c=2409=static int validate_group(struct perf_event *event)
--
arch/x86/events/core.c-2424-
arch/x86/events/core.c:2425: for_each_sibling_event(sibling, leader) {
arch/x86/events/core.c-2426- if (!is_x86_event(sibling))
--
arch/x86/events/intel/core.c=4222=static int intel_pmu_hw_config(struct perf_event *event)
--
arch/x86/events/intel/core.c-4288-
arch/x86/events/intel/core.c:4289: for_each_sibling_event(sibling, leader) {
arch/x86/events/intel/core.c-4290- if (branch_sample_call_stack(sibling))
--
arch/x86/events/intel/core.c-4399- if (leader->nr_siblings) {
arch/x86/events/intel/core.c:4400: for_each_sibling_event(sibling, leader) {
arch/x86/events/intel/core.c-4401- if (!is_x86_event(sibling)) {
--
arch/x86/events/intel/core.c-4427- if (leader->nr_siblings) {
arch/x86/events/intel/core.c:4428: for_each_sibling_event(sibling, leader)
arch/x86/events/intel/core.c-4429- intel_pmu_set_acr_caused_constr(sibling, idx++, cause_mask);
--
arch/x86/events/intel/core.c-4521- if (!is_mem_loads_aux_event(leader)) {
arch/x86/events/intel/core.c:4522: for_each_sibling_event(sibling, leader) {
arch/x86/events/intel/core.c-4523- if (is_mem_loads_aux_event(sibling))
--
arch/x86/events/intel/lbr.c=949=static void intel_pmu_lbr_counters_reorder(struct cpu_hw_events *cpuc,
--
arch/x86/events/intel/lbr.c-959-
arch/x86/events/intel/lbr.c:960: for_each_sibling_event(sibling, leader) {
arch/x86/events/intel/lbr.c-961- if (!branch_sample_counters(sibling))
--
arch/x86/events/intel/uncore.c=385=uncore_collect_events(struct intel_uncore_box *box, struct perf_event *leader,
--
arch/x86/events/intel/uncore.c-407-
arch/x86/events/intel/uncore.c:408: for_each_sibling_event(event, leader) {
arch/x86/events/intel/uncore.c-409- if (!is_box_event(box, event) ||
--
drivers/dma/idxd/perfmon.c=55=static int perfmon_collect_events(struct idxd_pmu *idxd_pmu,
--
drivers/dma/idxd/perfmon.c-76-
drivers/dma/idxd/perfmon.c:77: for_each_sibling_event(event, leader) {
drivers/dma/idxd/perfmon.c-78- if (!is_idxd_event(idxd_pmu, event) ||
--
drivers/iommu/intel/perfmon.c=257=static int iommu_pmu_validate_group(struct perf_event *event)
--
drivers/iommu/intel/perfmon.c-266- */
drivers/iommu/intel/perfmon.c:267: for_each_sibling_event(sibling, event->group_leader) {
drivers/iommu/intel/perfmon.c-268- if (!is_iommu_pmu_event(iommu_pmu, sibling) ||
--
drivers/perf/alibaba_uncore_drw_pmu.c=525=static int ali_drw_pmu_event_init(struct perf_event *event)
--
drivers/perf/alibaba_uncore_drw_pmu.c-556-
drivers/perf/alibaba_uncore_drw_pmu.c:557: for_each_sibling_event(sibling, event->group_leader) {
drivers/perf/alibaba_uncore_drw_pmu.c-558- if (sibling != event && !is_software_event(sibling)) {
--
drivers/perf/arm-cci.c=1239=static int validate_group(struct perf_event *event)
--
drivers/perf/arm-cci.c-1255-
drivers/perf/arm-cci.c:1256: for_each_sibling_event(sibling, leader) {
drivers/perf/arm-cci.c-1257- if (!validate_event(event->pmu, &fake_pmu, sibling))
--
drivers/perf/arm-ccn.c=704=static int arm_ccn_pmu_event_init(struct perf_event *event)
--
drivers/perf/arm-ccn.c-825-
drivers/perf/arm-ccn.c:826: for_each_sibling_event(sibling, event->group_leader) {
drivers/perf/arm-ccn.c-827- if (sibling->pmu != event->pmu &&
--
drivers/perf/arm-cmn.c=1687=static int arm_cmn_validate_group(struct arm_cmn *cmn, struct perf_event *event)
--
drivers/perf/arm-cmn.c-1707-
drivers/perf/arm-cmn.c:1708: for_each_sibling_event(sibling, leader)
drivers/perf/arm-cmn.c-1709- arm_cmn_val_add_event(cmn, val, sibling);
--
drivers/perf/arm-ni.c=292=static int arm_ni_validate_group(struct perf_event *event)
--
drivers/perf/arm-ni.c-303-
drivers/perf/arm-ni.c:304: for_each_sibling_event(sibling, leader) {
drivers/perf/arm-ni.c-305- if (!arm_ni_val_count_event(sibling, &val))
--
drivers/perf/arm_cspmu/arm_cspmu.c=578=static bool arm_cspmu_validate_group(struct perf_event *event)
--
drivers/perf/arm_cspmu/arm_cspmu.c-590-
drivers/perf/arm_cspmu/arm_cspmu.c:591: for_each_sibling_event(sibling, leader) {
drivers/perf/arm_cspmu/arm_cspmu.c-592- if (!arm_cspmu_validate_event(event->pmu, &fake_hw_events,
--
drivers/perf/arm_dmc620_pmu.c=512=static int dmc620_pmu_event_init(struct perf_event *event)
--
drivers/perf/arm_dmc620_pmu.c-556-
drivers/perf/arm_dmc620_pmu.c:557: for_each_sibling_event(sibling, event->group_leader) {
drivers/perf/arm_dmc620_pmu.c-558- if (sibling != event &&
--
drivers/perf/arm_dsu_pmu.c=507=static bool dsu_pmu_validate_group(struct perf_event *event)
--
drivers/perf/arm_dsu_pmu.c-517- return false;
drivers/perf/arm_dsu_pmu.c:518: for_each_sibling_event(sibling, leader) {
drivers/perf/arm_dsu_pmu.c-519- if (!dsu_pmu_validate_event(event->pmu, &fake_hw, sibling))
--
drivers/perf/arm_pmu.c=400=validate_group(struct perf_event *event)
--
drivers/perf/arm_pmu.c-416-
drivers/perf/arm_pmu.c:417: for_each_sibling_event(sibling, leader) {
drivers/perf/arm_pmu.c-418- if (!validate_event(event->pmu, &fake_pmu, sibling))
--
drivers/perf/arm_smmuv3_pmu.c=395=static int smmu_pmu_event_init(struct perf_event *event)
--
drivers/perf/arm_smmuv3_pmu.c-444-
drivers/perf/arm_smmuv3_pmu.c:445: for_each_sibling_event(sibling, event->group_leader) {
drivers/perf/arm_smmuv3_pmu.c-446- if (is_software_event(sibling))
--
drivers/perf/dwc_pcie_pmu.c=401=static int dwc_pcie_pmu_validate_group(struct perf_event *event)
--
drivers/perf/dwc_pcie_pmu.c-414-
drivers/perf/dwc_pcie_pmu.c:415: for_each_sibling_event(sibling, leader) {
drivers/perf/dwc_pcie_pmu.c-416- type = DWC_PCIE_EVENT_TYPE(sibling);
--
drivers/perf/dwc_pcie_pmu.c=432=static int dwc_pcie_pmu_event_init(struct perf_event *event)
--
drivers/perf/dwc_pcie_pmu.c-449-
drivers/perf/dwc_pcie_pmu.c:450: for_each_sibling_event(sibling, event->group_leader) {
drivers/perf/dwc_pcie_pmu.c-451- if (sibling->pmu != event->pmu && !is_software_event(sibling))
--
drivers/perf/fsl_imx8_ddr_perf.c=395=static int ddr_perf_event_init(struct perf_event *event)
--
drivers/perf/fsl_imx8_ddr_perf.c-423- return -EINVAL;
drivers/perf/fsl_imx8_ddr_perf.c:424: for_each_sibling_event(sibling, event->group_leader) {
drivers/perf/fsl_imx8_ddr_perf.c-425- if (!ddr_perf_filters_compatible(event, sibling))
--
drivers/perf/fsl_imx8_ddr_perf.c-429-
drivers/perf/fsl_imx8_ddr_perf.c:430: for_each_sibling_event(sibling, event->group_leader) {
drivers/perf/fsl_imx8_ddr_perf.c-431- if (sibling->pmu != event->pmu &&
--
drivers/perf/fsl_imx9_ddr_perf.c=557=static int ddr_perf_event_init(struct perf_event *event)
--
drivers/perf/fsl_imx9_ddr_perf.c-582-
drivers/perf/fsl_imx9_ddr_perf.c:583: for_each_sibling_event(sibling, event->group_leader) {
drivers/perf/fsl_imx9_ddr_perf.c-584- if (sibling->pmu != event->pmu &&
--
drivers/perf/fujitsu_uncore_pmu.c=169=static bool fujitsu_uncore_validate_event_group(struct perf_event *event)
--
drivers/perf/fujitsu_uncore_pmu.c-181-
drivers/perf/fujitsu_uncore_pmu.c:182: for_each_sibling_event(sibling, leader) {
drivers/perf/fujitsu_uncore_pmu.c-183- if (sibling->pmu == event->pmu)
--
drivers/perf/hisilicon/hisi_pcie_pmu.c=334=static bool hisi_pcie_pmu_validate_event_group(struct perf_event *event)
--
drivers/perf/hisilicon/hisi_pcie_pmu.c-349-
drivers/perf/hisilicon/hisi_pcie_pmu.c:350: for_each_sibling_event(sibling, event->group_leader) {
drivers/perf/hisilicon/hisi_pcie_pmu.c-351- if (is_software_event(sibling))
--
drivers/perf/hisilicon/hisi_uncore_pmu.c=97=static bool hisi_validate_event_group(struct perf_event *event)
--
drivers/perf/hisilicon/hisi_uncore_pmu.c-116-
drivers/perf/hisilicon/hisi_uncore_pmu.c:117: for_each_sibling_event(sibling, event->group_leader) {
drivers/perf/hisilicon/hisi_uncore_pmu.c-118- if (is_software_event(sibling))
--
drivers/perf/hisilicon/hns3_pmu.c=1054=static bool hns3_pmu_validate_event_group(struct perf_event *event)
--
drivers/perf/hisilicon/hns3_pmu.c-1069-
drivers/perf/hisilicon/hns3_pmu.c:1070: for_each_sibling_event(sibling, event->group_leader) {
drivers/perf/hisilicon/hns3_pmu.c-1071- if (is_software_event(sibling))
--
drivers/perf/marvell_pem_pmu.c=189=static int pem_perf_event_init(struct perf_event *event)
--
drivers/perf/marvell_pem_pmu.c-213-
drivers/perf/marvell_pem_pmu.c:214: for_each_sibling_event(sibling, event->group_leader) {
drivers/perf/marvell_pem_pmu.c-215- if (sibling->pmu != event->pmu &&
--
drivers/perf/qcom_l2_pmu.c=438=static int l2_cache_event_init(struct perf_event *event)
--
drivers/perf/qcom_l2_pmu.c-478-
drivers/perf/qcom_l2_pmu.c:479: for_each_sibling_event(sibling, event->group_leader) {
drivers/perf/qcom_l2_pmu.c-480- if (sibling->pmu != event->pmu &&
--
drivers/perf/qcom_l2_pmu.c-515-
drivers/perf/qcom_l2_pmu.c:516: for_each_sibling_event(sibling, event->group_leader) {
drivers/perf/qcom_l2_pmu.c-517- if ((sibling != event) &&
--
drivers/perf/qcom_l3_pmu.c=451=static bool qcom_l3_cache__validate_event_group(struct perf_event *event)
--
drivers/perf/qcom_l3_pmu.c-462-
drivers/perf/qcom_l3_pmu.c:463: for_each_sibling_event(sibling, leader) {
drivers/perf/qcom_l3_pmu.c-464- if (is_software_event(sibling))
--
drivers/perf/starfive_starlink_pmu.c=344=static bool starlink_pmu_validate_event_group(struct perf_event *event)
--
drivers/perf/starfive_starlink_pmu.c-357-
drivers/perf/starfive_starlink_pmu.c:358: for_each_sibling_event(sibling, leader) {
drivers/perf/starfive_starlink_pmu.c-359- if (sibling->pmu != event->pmu && !is_software_event(sibling))
--
drivers/perf/thunderx2_pmu.c=539=static bool tx2_uncore_validate_event_group(struct perf_event *event,
--
drivers/perf/thunderx2_pmu.c-550-
drivers/perf/thunderx2_pmu.c:551: for_each_sibling_event(sibling, leader) {
drivers/perf/thunderx2_pmu.c-552- if (!tx2_uncore_validate_event(event->pmu, sibling, &counters))
--
drivers/perf/xgene_pmu.c=876=static int xgene_perf_event_init(struct perf_event *event)
--
drivers/perf/xgene_pmu.c-923-
drivers/perf/xgene_pmu.c:924: for_each_sibling_event(sibling, event->group_leader) {
drivers/perf/xgene_pmu.c-925- if (sibling->pmu != event->pmu &&
--
include/linux/perf_event.h=736=struct pmu_event_list {
--
include/linux/perf_event.h-755-
include/linux/perf_event.h:756:#define for_each_sibling_event(sibling, event) \
include/linux/perf_event.h-757- lockdep_assert_event_ctx(event); \
--
kernel/events/core.c=739=static void perf_event_update_sibling_time(struct perf_event *leader)
--
kernel/events/core.c-742-
kernel/events/core.c:743: for_each_sibling_event(sibling, leader)
kernel/events/core.c-744- perf_event_update_time(sibling);
--
kernel/events/core.c=2047=static bool perf_event_validate_size(struct perf_event *event)
--
kernel/events/core.c-2061- * after the size has been validated, but we cannot safely use
kernel/events/core.c:2062: * for_each_sibling_event() until group_leader->ctx is set. A new group
kernel/events/core.c-2063- * leader cannot have any siblings yet, so we can safely skip checking
--
kernel/events/core.c-2068-
kernel/events/core.c:2069: for_each_sibling_event(sibling, group_leader) {
kernel/events/core.c-2070- if (__perf_event_read_size(sibling->attr.read_format,
--
kernel/events/core.c=2078=static void perf_group_attach(struct perf_event *event)
--
kernel/events/core.c-2105-
kernel/events/core.c:2106: for_each_sibling_event(pos, group_leader)
kernel/events/core.c-2107- perf_event__header_size(pos);
--
kernel/events/core.c=2162=static void perf_put_aux_event(struct perf_event *event)
--
kernel/events/core.c-2180- */
kernel/events/core.c:2181: for_each_sibling_event(iter, event) {
kernel/events/core.c-2182- if (iter->aux_event != event)
--
kernel/events/core.c=2251=static void perf_group_detach(struct perf_event *event)
--
kernel/events/core.c-2312-out:
kernel/events/core.c:2313: for_each_sibling_event(tmp, leader)
kernel/events/core.c-2314- perf_event__header_size(tmp);
--
kernel/events/core.c=2408=group_sched_out(struct perf_event *group_event, struct perf_event_context *ctx)
--
kernel/events/core.c-2421- */
kernel/events/core.c:2422: for_each_sibling_event(event, group_event)
kernel/events/core.c-2423- event_sched_out(event, ctx);
--
kernel/events/core.c=2689=static void perf_event_unthrottle_group(struct perf_event *event, bool skip_start_event)
--
kernel/events/core.c-2693- perf_event_unthrottle(leader, skip_start_event ? leader != event : true);
kernel/events/core.c:2694: for_each_sibling_event(sibling, leader)
kernel/events/core.c-2695- perf_event_unthrottle(sibling, skip_start_event ? sibling != event : true);
--
kernel/events/core.c=2698=static void perf_event_throttle_group(struct perf_event *event)
--
kernel/events/core.c-2702- perf_event_throttle(leader);
kernel/events/core.c:2703: for_each_sibling_event(sibling, leader)
kernel/events/core.c-2704- perf_event_throttle(sibling);
--
kernel/events/core.c=2765=group_sched_in(struct perf_event *group_event, struct perf_event_context *ctx)
--
kernel/events/core.c-2780- */
kernel/events/core.c:2781: for_each_sibling_event(event, group_event) {
kernel/events/core.c-2782- if (event_sched_in(event, ctx)) {
--
kernel/events/core.c-2796- */
kernel/events/core.c:2797: for_each_sibling_event(event, group_event) {
kernel/events/core.c-2798- if (event == partial_group)
--
kernel/events/core.c=3986=static inline void group_update_userpage(struct perf_event *group_event)
--
kernel/events/core.c-3992-
kernel/events/core.c:3993: for_each_sibling_event(event, group_event)
kernel/events/core.c-3994- event_update_userpage(event);
--
kernel/events/core.c=4669=static void __perf_event_read(void *info)
--
kernel/events/core.c-4706-
kernel/events/core.c:4707: for_each_sibling_event(sub, event)
kernel/events/core.c-4708- perf_pmu_read(sub);
--
kernel/events/core.c=5922=static int __perf_read_group_add(struct perf_event *leader,
--
kernel/events/core.c-5987-
kernel/events/core.c:5988: for_each_sibling_event(sub, leader) {
kernel/events/core.c-5989- values[n++] += perf_event_count(sub, false);
--
kernel/events/core.c=6204=static void perf_event_for_each(struct perf_event *event,
--
kernel/events/core.c-6214- perf_event_for_each_child(event, func);
kernel/events/core.c:6215: for_each_sibling_event(sibling, event)
kernel/events/core.c-6216- perf_event_for_each_child(sibling, func);
--
kernel/events/core.c=7789=static void perf_output_read_group(struct perf_output_handle *handle,
--
kernel/events/core.c-7824-
kernel/events/core.c:7825: for_each_sibling_event(sub, leader) {
kernel/events/core.c-7826- n = 0;
--
kernel/events/core.c=10961=static void perf_tp_event_target_task(u64 count, void *record,
--
kernel/events/core.c-10972- __perf_tp_event_target_task(count, record, regs, data, raw, event);
kernel/events/core.c:10973: for_each_sibling_event(sibling, event)
kernel/events/core.c-10974- __perf_tp_event_target_task(count, record, regs, data, raw, sibling);
--
kernel/events/core.c-10978- __perf_tp_event_target_task(count, record, regs, data, raw, event);
kernel/events/core.c:10979: for_each_sibling_event(sibling, event)
kernel/events/core.c-10980- __perf_tp_event_target_task(count, record, regs, data, raw, sibling);
--
kernel/events/core.c=13387=SYSCALL_DEFINE5(perf_event_open,
--
kernel/events/core.c-13700-
kernel/events/core.c:13701: for_each_sibling_event(sibling, group_leader) {
kernel/events/core.c-13702- perf_remove_from_context(sibling, 0);
--
kernel/events/core.c-13715- */
kernel/events/core.c:13716: for_each_sibling_event(sibling, group_leader) {
kernel/events/core.c-13717- sibling->pmu_ctx = pmu_ctx;
--
kernel/events/core.c=13898=static void __perf_pmu_remove(struct perf_event_context *ctx,
--
kernel/events/core.c-13909-
kernel/events/core.c:13910: for_each_sibling_event(sibling, event) {
kernel/events/core.c-13911- perf_remove_from_context(sibling, 0);
--
kernel/events/core.c=14397=static int inherit_group(struct perf_event *parent_event,
--
kernel/events/core.c-14415- */
kernel/events/core.c:14416: for_each_sibling_event(sub, parent_event) {
kernel/events/core.c-14417- child_ctr = inherit_event(sub, parent, parent_ctx,
]
|
| 12/2 |
2026/05/18 09:38 |
llm |
expert |
0m
Model:
Tokens:
input: 20133
output: 45
thoughts: 27
Thoughts: Here's my take on those thoughts:
**Dissecting the `for_each_sibling_event` Macro**
Okay, so I'm looking at `for_each_sibling_event`... right, it's a macro, and judging by the include path, `include/linux/perf_event.h`, it's clearly related to the performance events subsystem in the kernel. No surprise there. I'm already anticipating what it does: likely iterating through a set of performance event structures that are related somehow, probably sharing the same performance counter or something similar. I need to dig into the macro definition itself to see the specifics of the iteration and exactly what constitutes a "sibling" in this context. Is it using the standard `list_for_each_entry` pattern or something more specialized for this perf event data structure? Knowing how it works is key to understanding the control flow and optimization potential around performance monitoring within the kernel.
|
| 13/2 |
2026/05/18 09:38 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:include/linux/perf_event.h IncludeLines:true Name:for_each_sibling_event]
Results: map[SourceCode: SourceFile:]
|
| 14/2 |
2026/05/18 09:38 |
llm |
expert |
0m
Model:
Tokens:
input: 20424
output: 23
thoughts: 24
Thoughts: Here's my attempt at summarizing the thought process, tailored for an expert audience:
**Diving into the Perf Event Header**
Okay, let's get down to brass tacks. I'm going to pull up `include/linux/perf_event.h` and focus in on line 756 or thereabouts. I'm probably interested in some aspect of the perf event subsystem and I suspect that line contains something germane to what I'm currently working on, probably relating to the architecture's implementation or the API of the perf_events. I'm likely chasing a particular data structure or macro definition β perhaps related to event sampling, configuration, or perhaps even an interaction with the hardware performance counters themselves. My mental models are already engaged. I'm anticipating to quickly identify the relevant structure and related implementation to quickly understand the current state or find an issue.
|
| 15/2 |
2026/05/18 09:38 |
tool |
grepper |
0m
Args: map[Expression:#define for_each_sibling_event]
Results: map[Output:include/linux/perf_event.h=736=struct pmu_event_list {
--
include/linux/perf_event.h-755-
include/linux/perf_event.h:756:#define for_each_sibling_event(sibling, event) \
include/linux/perf_event.h-757- lockdep_assert_event_ctx(event); \
]
|
| 16/2 |
2026/05/18 09:38 |
llm |
expert |
0m
Model:
Tokens:
input: 20766
output: 22
thoughts: 33
Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience:
**Analyzing `perf_event.h` for `for_each_sibling_event`**
Right, so I'm thinking, let's get down to brass tacks. I need to understand how the sibling event iteration is implemented within the perf subsystem. My immediate instinct is to crack open the source code, specifically `include/linux/perf_event.h`. Itβs the obvious starting point given the context. And to do that efficiently, I'll leverage `read-file` directly. This gets me past any potential build system or environment vagaries. Once I've got that header file in hand, I can search for the macro definition for `for_each_sibling_event`. I'm expecting to find a pretty straightforward loop construct, likely built on some internal data structure related to the event's linkage. I'll need to pay close attention to the iteration variable's type and what it references. Ultimately, this will give me the low-down on how the sibling relationships are traversed and accessed, which is crucial for any higher-level analysis or modification of the event's logic.
|
| 17/2 |
2026/05/18 09:38 |
tool |
grepper |
0m
Args: map[Expression:define for_each_sibling_event]
Results: map[Output:include/linux/perf_event.h=736=struct pmu_event_list {
--
include/linux/perf_event.h-755-
include/linux/perf_event.h:756:#define for_each_sibling_event(sibling, event) \
include/linux/perf_event.h-757- lockdep_assert_event_ctx(event); \
]
|
| 18/2 |
2026/05/18 09:38 |
llm |
expert |
0m
Model:
Tokens:
input: 21184
output: 24
thoughts: 48
Thoughts: Okay, let me think this through. Hmm, there's a problem, and I need to come up with a solution quickly.
**Quick Thinking and Troubleshooting a Code Search**
Initially, I considered using `grepper` and thought, "Wait, can I just leverage `grepper` with a lookahead argument, like `-A 10`?" But then, I recalled that `grepper`'s implementation probably doesn't support that particular feature, or perhaps it has a different syntax for lookaround. It's a common oversight in these utility programs. Okay, so that idea's a dead end.
Right, so to make this work, I'll need to read the content of the file and perform a local search using `grep`. I'll use the `-A` option locally. That'll allow me to access the lines around the search results and do what I need. Simple and effective; problem solved.
|
| 19/2 |
2026/05/18 09:38 |
tool |
read-file |
0m
Args: map[File:include/linux/perf_event.h]
Results: map[Contents:/*
* Performance events:
*
* Copyright (C) 2008-2009, Thomas Gleixner <tglx@linutronix.de>
* Copyright (C) 2008-2011, Red Hat, Inc., Ingo Molnar
* Copyright (C) 2008-2011, Red Hat, Inc., Peter Zijlstra
*
* Data type definitions, declarations, prototypes.
*
* Started by: Thomas Gleixner and Ingo Molnar
*
* For licencing details see kernel-base/COPYING
*/
#ifndef _LINUX_PERF_EVENT_H
#define _LINUX_PERF_EVENT_H
#include <uapi/linux/perf_event.h>
#include <uapi/linux/bpf_perf_event.h>
/*
* Kernel-internal data types and definitions:
*/
#ifdef CONFIG_PERF_EVENTS
# include <asm/perf_event.h>
# include <asm/local64.h>
#endif
#ifdef CONFIG_HAVE_HW_BREAKPOINT
# include <linux/rhashtable-types.h>
# include <asm/hw_breakpoint.h>
#endif
#include <linux/list.h>
#include <linux/mutex.h>
#include <linux/rculist.h>
#include <linux/rcupdate.h>
#include <linux/spinlock.h>
#include <linux/hrtimer.h>
#include <linux/fs.h>
#include <linux/pid_namespace.h>
#include <linux/workqueue.h>
#include <linux/ftrace.h>
#include <linux/cpu.h>
#include <linux/irq_work.h>
#include <linux/static_key.h>
#include <linux/jump_label_ratelimit.h>
#include <linux/atomic.h>
#include <linux/sysfs.h>
#include <linux/perf_regs.h>
#include <linux/cgroup.h>
#include <linux/refcount.h>
#include <linux/security.h>
#include <linux/static_call.h>
#include <linux/lockdep.h>
#include <asm/local.h>
struct perf_callchain_entry {
u64 nr;
u64 ip[]; /* /proc/sys/kernel/perf_event_max_stack */
};
struct perf_callchain_entry_ctx {
struct perf_callchain_entry *entry;
u32 max_stack;
u32 nr;
short contexts;
bool contexts_maxed;
};
typedef unsigned long (*perf_copy_f)(void *dst, const void *src,
unsigned long off, unsigned long len);
struct perf_raw_frag {
union {
struct perf_raw_frag *next;
unsigned long pad;
};
perf_copy_f copy;
void *data;
u32 size;
} __packed;
struct perf_raw_record {
struct perf_raw_frag frag;
u32 size;
};
static __always_inline bool perf_raw_frag_last(const struct perf_raw_frag *frag)
{
return frag->pad < sizeof(u64);
}
/*
* branch stack layout:
* nr: number of taken branches stored in entries[]
* hw_idx: The low level index of raw branch records
* for the most recent branch.
* -1ULL means invalid/unknown.
*
* Note that nr can vary from sample to sample
* branches (to, from) are stored from most recent
* to least recent, i.e., entries[0] contains the most
* recent branch.
* The entries[] is an abstraction of raw branch records,
* which may not be stored in age order in HW, e.g. Intel LBR.
* The hw_idx is to expose the low level index of raw
* branch record for the most recent branch aka entries[0].
* The hw_idx index is between -1 (unknown) and max depth,
* which can be retrieved in /sys/devices/cpu/caps/branches.
* For the architectures whose raw branch records are
* already stored in age order, the hw_idx should be 0.
*/
struct perf_branch_stack {
u64 nr;
u64 hw_idx;
struct perf_branch_entry entries[];
};
struct task_struct;
/*
* extra PMU register associated with an event
*/
struct hw_perf_event_extra {
u64 config; /* register value */
unsigned int reg; /* register address or index */
int alloc; /* extra register already allocated */
int idx; /* index in shared_regs->regs[] */
};
/**
* hw_perf_event::flag values
*
* PERF_EVENT_FLAG_ARCH bits are reserved for architecture-specific
* usage.
*/
#define PERF_EVENT_FLAG_ARCH 0x0fffffff
#define PERF_EVENT_FLAG_USER_READ_CNT 0x80000000
static_assert((PERF_EVENT_FLAG_USER_READ_CNT & PERF_EVENT_FLAG_ARCH) == 0);
/**
* struct hw_perf_event - performance event hardware details:
*/
struct hw_perf_event {
#ifdef CONFIG_PERF_EVENTS
union {
struct { /* hardware */
u64 config;
u64 config1;
u64 last_tag;
u64 dyn_constraint;
unsigned long config_base;
unsigned long event_base;
int event_base_rdpmc;
int idx;
int last_cpu;
int flags;
struct hw_perf_event_extra extra_reg;
struct hw_perf_event_extra branch_reg;
};
struct { /* aux / Intel-PT */
u64 aux_config;
/*
* For AUX area events, aux_paused cannot be a state
* flag because it can be updated asynchronously to
* state.
*/
unsigned int aux_paused;
};
struct { /* software */
struct hrtimer hrtimer;
};
struct { /* tracepoint */
/* for tp_event->class */
struct list_head tp_list;
};
struct { /* amd_power */
u64 pwr_acc;
u64 ptsc;
};
#ifdef CONFIG_HAVE_HW_BREAKPOINT
struct { /* breakpoint */
/*
* Crufty hack to avoid the chicken and egg
* problem hw_breakpoint has with context
* creation and event initalization.
*/
struct arch_hw_breakpoint info;
struct rhlist_head bp_list;
};
#endif
struct { /* amd_iommu */
u8 iommu_bank;
u8 iommu_cntr;
u16 padding;
u64 conf;
u64 conf1;
};
};
/*
* If the event is a per task event, this will point to the task in
* question. See the comment in perf_event_alloc().
*/
struct task_struct *target;
/*
* PMU would store hardware filter configuration
* here.
*/
void *addr_filters;
/* Last sync'ed generation of filters */
unsigned long addr_filters_gen;
/*
* hw_perf_event::state flags; used to track the PERF_EF_* state.
*/
/* the counter is stopped */
#define PERF_HES_STOPPED 0x01
/* event->count up-to-date */
#define PERF_HES_UPTODATE 0x02
#define PERF_HES_ARCH 0x04
int state;
/*
* The last observed hardware counter value, updated with a
* local64_cmpxchg() such that pmu::read() can be called nested.
*/
local64_t prev_count;
/*
* The period to start the next sample with.
*/
u64 sample_period;
union {
struct { /* Sampling */
/*
* The period we started this sample with.
*/
u64 last_period;
/*
* However much is left of the current period;
* note that this is a full 64bit value and
* allows for generation of periods longer
* than hardware might allow.
*/
local64_t period_left;
};
struct { /* Topdown events counting for context switch */
u64 saved_metric;
u64 saved_slots;
};
};
/*
* State for throttling the event, see __perf_event_overflow() and
* perf_adjust_freq_unthr_context().
*/
u64 interrupts_seq;
u64 interrupts;
/*
* State for freq target events, see __perf_event_overflow() and
* perf_adjust_freq_unthr_context().
*/
u64 freq_time_stamp;
u64 freq_count_stamp;
#endif /* CONFIG_PERF_EVENTS */
};
struct perf_event;
struct perf_event_pmu_context;
/*
* Common implementation detail of pmu::{start,commit,cancel}_txn
*/
/* txn to add/schedule event on PMU */
#define PERF_PMU_TXN_ADD 0x1
/* txn to read event group from PMU */
#define PERF_PMU_TXN_READ 0x2
/**
* pmu::capabilities flags
*/
#define PERF_PMU_CAP_NO_INTERRUPT 0x0001
#define PERF_PMU_CAP_NO_NMI 0x0002
#define PERF_PMU_CAP_AUX_NO_SG 0x0004
#define PERF_PMU_CAP_EXTENDED_REGS 0x0008
#define PERF_PMU_CAP_EXCLUSIVE 0x0010
#define PERF_PMU_CAP_ITRACE 0x0020
#define PERF_PMU_CAP_NO_EXCLUDE 0x0040
#define PERF_PMU_CAP_AUX_OUTPUT 0x0080
#define PERF_PMU_CAP_EXTENDED_HW_TYPE 0x0100
#define PERF_PMU_CAP_AUX_PAUSE 0x0200
#define PERF_PMU_CAP_AUX_PREFER_LARGE 0x0400
/**
* pmu::scope
*/
enum perf_pmu_scope {
PERF_PMU_SCOPE_NONE = 0,
PERF_PMU_SCOPE_CORE,
PERF_PMU_SCOPE_DIE,
PERF_PMU_SCOPE_CLUSTER,
PERF_PMU_SCOPE_PKG,
PERF_PMU_SCOPE_SYS_WIDE,
PERF_PMU_MAX_SCOPE,
};
struct perf_output_handle;
#define PMU_NULL_DEV ((void *)(~0UL))
/**
* struct pmu - generic performance monitoring unit
*/
struct pmu {
struct list_head entry;
spinlock_t events_lock;
struct list_head events;
struct module *module;
struct device *dev;
struct device *parent;
const struct attribute_group **attr_groups;
const struct attribute_group **attr_update;
const char *name;
int type;
/*
* various common per-pmu feature flags
*/
int capabilities;
/*
* PMU scope
*/
unsigned int scope;
struct perf_cpu_pmu_context * __percpu *cpu_pmu_context;
atomic_t exclusive_cnt; /* < 0: cpu; > 0: tsk */
int task_ctx_nr;
int hrtimer_interval_ms;
/* number of address filters this PMU can do */
unsigned int nr_addr_filters;
/*
* Fully disable/enable this PMU, can be used to protect from the PMI
* as well as for lazy/batch writing of the MSRs.
*/
void (*pmu_enable) (struct pmu *pmu); /* optional */
void (*pmu_disable) (struct pmu *pmu); /* optional */
/*
* Try and initialize the event for this PMU.
*
* Returns:
* -ENOENT -- @event is not for this PMU
*
* -ENODEV -- @event is for this PMU but PMU not present
* -EBUSY -- @event is for this PMU but PMU temporarily unavailable
* -EINVAL -- @event is for this PMU but @event is not valid
* -EOPNOTSUPP -- @event is for this PMU, @event is valid, but not supported
* -EACCES -- @event is for this PMU, @event is valid, but no privileges
*
* 0 -- @event is for this PMU and valid
*
* Other error return values are allowed.
*/
int (*event_init) (struct perf_event *event);
/*
* Notification that the event was mapped or unmapped. Called
* in the context of the mapping task.
*/
void (*event_mapped) (struct perf_event *event, struct mm_struct *mm); /* optional */
void (*event_unmapped) (struct perf_event *event, struct mm_struct *mm); /* optional */
/*
* Flags for ->add()/->del()/ ->start()/->stop(). There are
* matching hw_perf_event::state flags.
*/
/* start the counter when adding */
#define PERF_EF_START 0x01
/* reload the counter when starting */
#define PERF_EF_RELOAD 0x02
/* update the counter when stopping */
#define PERF_EF_UPDATE 0x04
/* AUX area event, pause tracing */
#define PERF_EF_PAUSE 0x08
/* AUX area event, resume tracing */
#define PERF_EF_RESUME 0x10
/*
* Adds/Removes a counter to/from the PMU, can be done inside a
* transaction, see the ->*_txn() methods.
*
* The add/del callbacks will reserve all hardware resources required
* to service the event, this includes any counter constraint
* scheduling etc.
*
* Called with IRQs disabled and the PMU disabled on the CPU the event
* is on.
*
* ->add() called without PERF_EF_START should result in the same state
* as ->add() followed by ->stop().
*
* ->del() must always PERF_EF_UPDATE stop an event. If it calls
* ->stop() that must deal with already being stopped without
* PERF_EF_UPDATE.
*/
int (*add) (struct perf_event *event, int flags);
void (*del) (struct perf_event *event, int flags);
/*
* Starts/Stops a counter present on the PMU.
*
* The PMI handler should stop the counter when perf_event_overflow()
* returns !0. ->start() will be used to continue.
*
* Also used to change the sample period.
*
* Called with IRQs disabled and the PMU disabled on the CPU the event
* is on -- will be called from NMI context with the PMU generates
* NMIs.
*
* ->stop() with PERF_EF_UPDATE will read the counter and update
* period/count values like ->read() would.
*
* ->start() with PERF_EF_RELOAD will reprogram the counter
* value, must be preceded by a ->stop() with PERF_EF_UPDATE.
*
* ->stop() with PERF_EF_PAUSE will stop as simply as possible. Will not
* overlap another ->stop() with PERF_EF_PAUSE nor ->start() with
* PERF_EF_RESUME.
*
* ->start() with PERF_EF_RESUME will start as simply as possible but
* only if the counter is not otherwise stopped. Will not overlap
* another ->start() with PERF_EF_RESUME nor ->stop() with
* PERF_EF_PAUSE.
*
* Notably, PERF_EF_PAUSE/PERF_EF_RESUME *can* be concurrent with other
* ->stop()/->start() invocations, just not itself.
*/
void (*start) (struct perf_event *event, int flags);
void (*stop) (struct perf_event *event, int flags);
/*
* Updates the counter value of the event.
*
* For sampling capable PMUs this will also update the software period
* hw_perf_event::period_left field.
*/
void (*read) (struct perf_event *event);
/*
* Group events scheduling is treated as a transaction, add
* group events as a whole and perform one schedulability test.
* If the test fails, roll back the whole group
*
* Start the transaction, after this ->add() doesn't need to
* do schedulability tests.
*
* Optional.
*/
void (*start_txn) (struct pmu *pmu, unsigned int txn_flags);
/*
* If ->start_txn() disabled the ->add() schedulability test
* then ->commit_txn() is required to perform one. On success
* the transaction is closed. On error the transaction is kept
* open until ->cancel_txn() is called.
*
* Optional.
*/
int (*commit_txn) (struct pmu *pmu);
/*
* Will cancel the transaction, assumes ->del() is called
* for each successful ->add() during the transaction.
*
* Optional.
*/
void (*cancel_txn) (struct pmu *pmu);
/*
* Will return the value for perf_event_mmap_page::index for this event,
* if no implementation is provided it will default to 0 (see
* perf_event_idx_default).
*/
int (*event_idx) (struct perf_event *event); /*optional */
/*
* context-switches callback
*/
void (*sched_task) (struct perf_event_pmu_context *pmu_ctx,
struct task_struct *task, bool sched_in);
/*
* Kmem cache of PMU specific data
*/
struct kmem_cache *task_ctx_cache;
/*
* Set up pmu-private data structures for an AUX area
*/
void *(*setup_aux) (struct perf_event *event, void **pages,
int nr_pages, bool overwrite);
/* optional */
/*
* Free pmu-private AUX data structures
*/
void (*free_aux) (void *aux); /* optional */
/*
* Take a snapshot of the AUX buffer without touching the event
* state, so that preempting ->start()/->stop() callbacks does
* not interfere with their logic. Called in PMI context.
*
* Returns the size of AUX data copied to the output handle.
*
* Optional.
*/
long (*snapshot_aux) (struct perf_event *event,
struct perf_output_handle *handle,
unsigned long size);
/*
* Validate address range filters: make sure the HW supports the
* requested configuration and number of filters; return 0 if the
* supplied filters are valid, -errno otherwise.
*
* Runs in the context of the ioctl()ing process and is not serialized
* with the rest of the PMU callbacks.
*/
int (*addr_filters_validate) (struct list_head *filters);
/* optional */
/*
* Synchronize address range filter configuration:
* translate hw-agnostic filters into hardware configuration in
* event::hw::addr_filters.
*
* Runs as a part of filter sync sequence that is done in ->start()
* callback by calling perf_event_addr_filters_sync().
*
* May (and should) traverse event::addr_filters::list, for which its
* caller provides necessary serialization.
*/
void (*addr_filters_sync) (struct perf_event *event);
/* optional */
/*
* Check if event can be used for aux_output purposes for
* events of this PMU.
*
* Runs from perf_event_open(). Should return 0 for "no match"
* or non-zero for "match".
*/
int (*aux_output_match) (struct perf_event *event);
/* optional */
/*
* Skip programming this PMU on the given CPU. Typically needed for
* big.LITTLE things.
*/
bool (*filter) (struct pmu *pmu, int cpu); /* optional */
/*
* Check period value for PERF_EVENT_IOC_PERIOD ioctl.
*/
int (*check_period) (struct perf_event *event, u64 value); /* optional */
};
enum perf_addr_filter_action_t {
PERF_ADDR_FILTER_ACTION_STOP = 0,
PERF_ADDR_FILTER_ACTION_START,
PERF_ADDR_FILTER_ACTION_FILTER,
};
/**
* struct perf_addr_filter - address range filter definition
* @entry: event's filter list linkage
* @path: object file's path for file-based filters
* @offset: filter range offset
* @size: filter range size (size==0 means single address trigger)
* @action: filter/start/stop
*
* This is a hardware-agnostic filter configuration as specified by the user.
*/
struct perf_addr_filter {
struct list_head entry;
struct path path;
unsigned long offset;
unsigned long size;
enum perf_addr_filter_action_t action;
};
/**
* struct perf_addr_filters_head - container for address range filters
* @list: list of filters for this event
* @lock: spinlock that serializes accesses to the @list and event's
* (and its children's) filter generations.
* @nr_file_filters: number of file-based filters
*
* A child event will use parent's @list (and therefore @lock), so they are
* bundled together; see perf_event_addr_filters().
*/
struct perf_addr_filters_head {
struct list_head list;
raw_spinlock_t lock;
unsigned int nr_file_filters;
};
struct perf_addr_filter_range {
unsigned long start;
unsigned long size;
};
/*
* The normal states are:
*
* ACTIVE --.
* ^ |
* | |
* sched_{in,out}() |
* | |
* v |
* ,---> INACTIVE --+ <-.
* | | |
* | {dis,en}able()
* sched_in() | |
* | OFF <--' --+
* | |
* `---> ERROR ------'
*
* That is:
*
* sched_in: INACTIVE -> {ACTIVE,ERROR}
* sched_out: ACTIVE -> INACTIVE
* disable: {ACTIVE,INACTIVE} -> OFF
* enable: {OFF,ERROR} -> INACTIVE
*
* Where {OFF,ERROR} are disabled states.
*
* Then we have the {EXIT,REVOKED,DEAD} states which are various shades of
* defunct events:
*
* - EXIT means task that the even was assigned to died, but child events
* still live, and further children can still be created. But the event
* itself will never be active again. It can only transition to
* {REVOKED,DEAD};
*
* - REVOKED means the PMU the event was associated with is gone; all
* functionality is stopped but the event is still alive. Can only
* transition to DEAD;
*
* - DEAD event really is DYING tearing down state and freeing bits.
*
*/
enum perf_event_state {
PERF_EVENT_STATE_DEAD = -5,
PERF_EVENT_STATE_REVOKED = -4, /* pmu gone, must not touch */
PERF_EVENT_STATE_EXIT = -3, /* task died, still inherit */
PERF_EVENT_STATE_ERROR = -2, /* scheduling error, can enable */
PERF_EVENT_STATE_OFF = -1,
PERF_EVENT_STATE_INACTIVE = 0,
PERF_EVENT_STATE_ACTIVE = 1,
};
struct file;
struct perf_sample_data;
typedef void (*perf_overflow_handler_t)(struct perf_event *,
struct perf_sample_data *,
struct pt_regs *regs);
/*
* Event capabilities. For event_caps and groups caps.
*
* PERF_EV_CAP_SOFTWARE: Is a software event.
* PERF_EV_CAP_READ_ACTIVE_PKG: A CPU event (or cgroup event) that can be read
* from any CPU in the package where it is active.
* PERF_EV_CAP_SIBLING: An event with this flag must be a group sibling and
* cannot be a group leader. If an event with this flag is detached from the
* group it is scheduled out and moved into an unrecoverable ERROR state.
* PERF_EV_CAP_READ_SCOPE: A CPU event that can be read from any CPU of the
* PMU scope where it is active.
*/
#define PERF_EV_CAP_SOFTWARE BIT(0)
#define PERF_EV_CAP_READ_ACTIVE_PKG BIT(1)
#define PERF_EV_CAP_SIBLING BIT(2)
#define PERF_EV_CAP_READ_SCOPE BIT(3)
#define SWEVENT_HLIST_BITS 8
#define SWEVENT_HLIST_SIZE (1 << SWEVENT_HLIST_BITS)
struct swevent_hlist {
struct hlist_head heads[SWEVENT_HLIST_SIZE];
struct rcu_head rcu_head;
};
#define PERF_ATTACH_CONTEXT 0x0001
#define PERF_ATTACH_GROUP 0x0002
#define PERF_ATTACH_TASK 0x0004
#define PERF_ATTACH_TASK_DATA 0x0008
#define PERF_ATTACH_GLOBAL_DATA 0x0010
#define PERF_ATTACH_SCHED_CB 0x0020
#define PERF_ATTACH_CHILD 0x0040
#define PERF_ATTACH_EXCLUSIVE 0x0080
#define PERF_ATTACH_CALLCHAIN 0x0100
#define PERF_ATTACH_ITRACE 0x0200
struct bpf_prog;
struct perf_cgroup;
struct perf_buffer;
struct pmu_event_list {
raw_spinlock_t lock;
struct list_head list;
};
/*
* event->sibling_list is modified whole holding both ctx->lock and ctx->mutex
* as such iteration must hold either lock. However, since ctx->lock is an IRQ
* safe lock, and is only held by the CPU doing the modification, having IRQs
* disabled is sufficient since it will hold-off the IPIs.
*/
#ifdef CONFIG_PROVE_LOCKING
# define lockdep_assert_event_ctx(event) \
WARN_ON_ONCE(__lockdep_enabled && \
(this_cpu_read(hardirqs_enabled) && \
lockdep_is_held(&(event)->ctx->mutex) != LOCK_STATE_HELD))
#else
# define lockdep_assert_event_ctx(event)
#endif
#define for_each_sibling_event(sibling, event) \
lockdep_assert_event_ctx(event); \
if ((event)->group_leader == (event)) \
list_for_each_entry((sibling), &(event)->sibling_list, sibling_list)
/**
* struct perf_event - performance event kernel representation:
*/
struct perf_event {
#ifdef CONFIG_PERF_EVENTS
/*
* entry onto perf_event_context::event_list;
* modifications require ctx->lock
* RCU safe iterations.
*/
struct list_head event_entry;
/*
* Locked for modification by both ctx->mutex and ctx->lock; holding
* either sufficies for read.
*/
struct list_head sibling_list;
struct list_head active_list;
/*
* Node on the pinned or flexible tree located at the event context;
*/
struct rb_node group_node;
u64 group_index;
/*
* We need storage to track the entries in perf_pmu_migrate_context; we
* cannot use the event_entry because of RCU and we want to keep the
* group in tact which avoids us using the other two entries.
*/
struct list_head migrate_entry;
struct hlist_node hlist_entry;
struct list_head active_entry;
int nr_siblings;
/* Not serialized. Only written during event initialization. */
int event_caps;
/* The cumulative AND of all event_caps for events in this group. */
int group_caps;
unsigned int group_generation;
struct perf_event *group_leader;
/*
* event->pmu will always point to pmu in which this event belongs.
* Whereas event->pmu_ctx->pmu may point to other pmu when group of
* different pmu events is created.
*/
struct pmu *pmu;
void *pmu_private;
enum perf_event_state state;
unsigned int attach_state;
local64_t count;
atomic64_t child_count;
/*
* These are the total time in nanoseconds that the event
* has been enabled (i.e. eligible to run, and the task has
* been scheduled in, if this is a per-task event)
* and running (scheduled onto the CPU), respectively.
*/
u64 total_time_enabled;
u64 total_time_running;
u64 tstamp;
struct perf_event_attr attr;
u16 header_size;
u16 id_header_size;
u16 read_size;
struct hw_perf_event hw;
struct perf_event_context *ctx;
/*
* event->pmu_ctx points to perf_event_pmu_context in which the event
* is added. This pmu_ctx can be of other pmu for sw event when that
* sw event is part of a group which also contains non-sw events.
*/
struct perf_event_pmu_context *pmu_ctx;
atomic_long_t refcount;
/*
* These accumulate total time (in nanoseconds) that children
* events have been enabled and running, respectively.
*/
atomic64_t child_total_time_enabled;
atomic64_t child_total_time_running;
/*
* Protect attach/detach and child_list:
*/
struct mutex child_mutex;
struct list_head child_list;
struct perf_event *parent;
int oncpu;
int cpu;
struct list_head owner_entry;
struct task_struct *owner;
/* mmap bits */
struct mutex mmap_mutex;
refcount_t mmap_count;
struct perf_buffer *rb;
struct list_head rb_entry;
unsigned long rcu_batches;
int rcu_pending;
/* poll related */
wait_queue_head_t waitq;
struct fasync_struct *fasync;
/* delayed work for NMIs and such */
unsigned int pending_wakeup;
unsigned int pending_kill;
unsigned int pending_disable;
unsigned long pending_addr; /* SIGTRAP */
struct irq_work pending_irq;
struct irq_work pending_disable_irq;
struct callback_head pending_task;
unsigned int pending_work;
atomic_t event_limit;
/* address range filters */
struct perf_addr_filters_head addr_filters;
/* vma address array for file-based filders */
struct perf_addr_filter_range *addr_filter_ranges;
unsigned long addr_filters_gen;
/* for aux_output events */
struct perf_event *aux_event;
void (*destroy)(struct perf_event *);
struct rcu_head rcu_head;
struct pid_namespace *ns;
u64 id;
atomic64_t lost_samples;
u64 (*clock)(void);
perf_overflow_handler_t overflow_handler;
void *overflow_handler_context;
struct bpf_prog *prog;
u64 bpf_cookie;
#ifdef CONFIG_EVENT_TRACING
struct trace_event_call *tp_event;
struct event_filter *filter;
# ifdef CONFIG_FUNCTION_TRACER
struct ftrace_ops ftrace_ops;
# endif
#endif
#ifdef CONFIG_CGROUP_PERF
struct perf_cgroup *cgrp; /* cgroup event is attach to */
#endif
#ifdef CONFIG_SECURITY
void *security;
#endif
struct list_head sb_list;
struct list_head pmu_list;
/*
* Certain events gets forwarded to another pmu internally by over-
* writing kernel copy of event->attr.type without user being aware
* of it. event->orig_type contains original 'type' requested by
* user.
*/
u32 orig_type;
#endif /* CONFIG_PERF_EVENTS */
};
/*
* ,-----------------------[1:n]------------------------.
* V V
* perf_event_context <-[1:n]-> perf_event_pmu_context <-[1:n]- perf_event
* | |
* `--[n:1]-> pmu <-[1:n]--'
*
*
* struct perf_event_pmu_context lifetime is refcount based and RCU freed
* (similar to perf_event_context). Locking is as if it were a member of
* perf_event_context; specifically:
*
* modification, both: ctx->mutex && ctx->lock
* reading, either: ctx->mutex || ctx->lock
*
* There is one exception to this; namely put_pmu_ctx() isn't always called
* with ctx->mutex held; this means that as long as we can guarantee the epc
* has events the above rules hold.
*
* Specificially, sys_perf_event_open()'s group_leader case depends on
* ctx->mutex pinning the configuration. Since we hold a reference on
* group_leader (through the filedesc) it can't go away, therefore it's
* associated pmu_ctx must exist and cannot change due to ctx->mutex.
*
* perf_event holds a refcount on perf_event_context
* perf_event holds a refcount on perf_event_pmu_context
*/
struct perf_event_pmu_context {
struct pmu *pmu;
struct perf_event_context *ctx;
struct list_head pmu_ctx_entry;
struct list_head pinned_active;
struct list_head flexible_active;
/* Used to identify the per-cpu perf_event_pmu_context */
unsigned int embedded : 1;
unsigned int nr_events;
unsigned int nr_cgroups;
unsigned int nr_freq;
atomic_t refcount; /* event <-> epc */
struct rcu_head rcu_head;
/*
* Set when one or more (plausibly active) event can't be scheduled
* due to pmu overcommit or pmu constraints, except tolerant to
* events not necessary to be active due to scheduling constraints,
* such as cgroups.
*/
int rotate_necessary;
};
static inline bool perf_pmu_ctx_is_active(struct perf_event_pmu_context *epc)
{
return !list_empty(&epc->flexible_active) || !list_empty(&epc->pinned_active);
}
struct perf_event_groups {
struct rb_root tree;
u64 index;
};
/**
* struct perf_event_context - event context structure
*
* Used as a container for task events and CPU events as well:
*/
struct perf_event_context {
/*
* Protect the states of the events in the list,
* nr_active, and the list:
*/
raw_spinlock_t lock;
/*
* Protect the list of events. Locking either mutex or lock
* is sufficient to ensure the list doesn't change; to change
* the list you need to lock both the mutex and the spinlock.
*/
struct mutex mutex;
struct list_head pmu_ctx_list;
struct perf_event_groups pinned_groups;
struct perf_event_groups flexible_groups;
struct list_head event_list;
int nr_events;
int nr_user;
int is_active;
int nr_stat;
int nr_freq;
int rotate_disable;
refcount_t refcount; /* event <-> ctx */
struct task_struct *task;
/*
* Context clock, runs when context enabled.
*/
u64 time;
u64 timestamp;
u64 timeoffset;
/*
* These fields let us detect when two contexts have both
* been cloned (inherited) from a common ancestor.
*/
struct perf_event_context *parent_ctx;
u64 parent_gen;
u64 generation;
int pin_count;
#ifdef CONFIG_CGROUP_PERF
int nr_cgroups; /* cgroup evts */
#endif
struct rcu_head rcu_head;
/*
* The count of events for which using the switch-out fast path
* should be avoided.
*
* Sum (event->pending_work + events with
* (attr->inherit && (attr->sample_type & PERF_SAMPLE_READ)))
*
* The SIGTRAP is targeted at ctx->task, as such it won't do changing
* that until the signal is delivered.
*/
local_t nr_no_switch_fast;
};
/**
* struct perf_ctx_data - PMU specific data for a task
* @rcu_head: To avoid the race on free PMU specific data
* @refcount: To track users
* @global: To track system-wide users
* @ctx_cache: Kmem cache of PMU specific data
* @data: PMU specific data
*
* Currently, the struct is only used in Intel LBR call stack mode to
* save/restore the call stack of a task on context switches.
*
* The rcu_head is used to prevent the race on free the data.
* The data only be allocated when Intel LBR call stack mode is enabled.
* The data will be freed when the mode is disabled.
* The content of the data will only be accessed in context switch, which
* should be protected by rcu_read_lock().
*
* Because of the alignment requirement of Intel Arch LBR, the Kmem cache
* is used to allocate the PMU specific data. The ctx_cache is to track
* the Kmem cache.
*
* Careful: Struct perf_ctx_data is added as a pointer in struct task_struct.
* When system-wide Intel LBR call stack mode is enabled, a buffer with
* constant size will be allocated for each task.
* Also, system memory consumption can further grow when the size of
* struct perf_ctx_data enlarges.
*/
struct perf_ctx_data {
struct rcu_head rcu_head;
refcount_t refcount;
int global;
struct kmem_cache *ctx_cache;
void *data;
};
struct perf_cpu_pmu_context {
struct perf_event_pmu_context epc;
struct perf_event_pmu_context *task_epc;
struct list_head sched_cb_entry;
int sched_cb_usage;
int active_oncpu;
int exclusive;
int pmu_disable_count;
raw_spinlock_t hrtimer_lock;
struct hrtimer hrtimer;
ktime_t hrtimer_interval;
unsigned int hrtimer_active;
};
/**
* struct perf_event_cpu_context - per cpu event context structure
*/
struct perf_cpu_context {
struct perf_event_context ctx;
struct perf_event_context *task_ctx;
int online;
#ifdef CONFIG_CGROUP_PERF
struct perf_cgroup *cgrp;
#endif
/*
* Per-CPU storage for iterators used in visit_groups_merge. The default
* storage is of size 2 to hold the CPU and any CPU event iterators.
*/
int heap_size;
struct perf_event **heap;
struct perf_event *heap_default[2];
};
struct perf_output_handle {
struct perf_event *event;
struct perf_buffer *rb;
unsigned long wakeup;
unsigned long size;
union {
u64 flags; /* perf_output*() */
u64 aux_flags; /* perf_aux_output*() */
struct {
u64 skip_read : 1;
};
};
union {
void *addr;
unsigned long head;
};
int page;
};
struct bpf_perf_event_data_kern {
bpf_user_pt_regs_t *regs;
struct perf_sample_data *data;
struct perf_event *event;
};
#ifdef CONFIG_CGROUP_PERF
/*
* perf_cgroup_info keeps track of time_enabled for a cgroup.
* This is a per-cpu dynamically allocated data structure.
*/
struct perf_cgroup_info {
u64 time;
u64 timestamp;
u64 timeoffset;
int active;
};
struct perf_cgroup {
struct cgroup_subsys_state css;
struct perf_cgroup_info __percpu *info;
};
/*
* Must ensure cgroup is pinned (css_get) before calling
* this function. In other words, we cannot call this function
* if there is no cgroup event for the current CPU context.
*/
static inline struct perf_cgroup *
perf_cgroup_from_task(struct task_struct *task, struct perf_event_context *ctx)
{
return container_of(task_css_check(task, perf_event_cgrp_id,
ctx ? lockdep_is_held(&ctx->lock)
: true),
struct perf_cgroup, css);
}
#endif /* CONFIG_CGROUP_PERF */
#ifdef CONFIG_PERF_EVENTS
extern struct perf_event_context *perf_cpu_task_ctx(void);
extern void *perf_aux_output_begin(struct perf_output_handle *handle,
struct perf_event *event);
extern void perf_aux_output_end(struct perf_output_handle *handle,
unsigned long size);
extern int perf_aux_output_skip(struct perf_output_handle *handle,
unsigned long size);
extern void *perf_get_aux(struct perf_output_handle *handle);
extern void perf_aux_output_flag(struct perf_output_handle *handle, u64 flags);
extern void perf_event_itrace_started(struct perf_event *event);
extern int perf_pmu_register(struct pmu *pmu, const char *name, int type);
extern int perf_pmu_unregister(struct pmu *pmu);
extern void __perf_event_task_sched_in(struct task_struct *prev,
struct task_struct *task);
extern void __perf_event_task_sched_out(struct task_struct *prev,
struct task_struct *next);
extern int perf_event_init_task(struct task_struct *child, u64 clone_flags);
extern void perf_event_exit_task(struct task_struct *child);
extern void perf_event_free_task(struct task_struct *task);
extern void perf_event_delayed_put(struct task_struct *task);
extern struct file *perf_event_get(unsigned int fd);
extern const struct perf_event *perf_get_event(struct file *file);
extern const struct perf_event_attr *perf_event_attrs(struct perf_event *event);
extern void perf_event_print_debug(void);
extern void perf_pmu_disable(struct pmu *pmu);
extern void perf_pmu_enable(struct pmu *pmu);
extern void perf_sched_cb_dec(struct pmu *pmu);
extern void perf_sched_cb_inc(struct pmu *pmu);
extern int perf_event_task_disable(void);
extern int perf_event_task_enable(void);
extern void perf_pmu_resched(struct pmu *pmu);
extern int perf_event_refresh(struct perf_event *event, int refresh);
extern void perf_event_update_userpage(struct perf_event *event);
extern int perf_event_release_kernel(struct perf_event *event);
extern struct perf_event *
perf_event_create_kernel_counter(struct perf_event_attr *attr,
int cpu,
struct task_struct *task,
perf_overflow_handler_t callback,
void *context);
extern void perf_pmu_migrate_context(struct pmu *pmu,
int src_cpu, int dst_cpu);
extern int perf_event_read_local(struct perf_event *event, u64 *value,
u64 *enabled, u64 *running);
extern u64 perf_event_read_value(struct perf_event *event,
u64 *enabled, u64 *running);
extern struct perf_callchain_entry *perf_callchain(struct perf_event *event, struct pt_regs *regs);
static inline bool branch_sample_no_flags(const struct perf_event *event)
{
return event->attr.branch_sample_type & PERF_SAMPLE_BRANCH_NO_FLAGS;
}
static inline bool branch_sample_no_cycles(const struct perf_event *event)
{
return event->attr.branch_sample_type & PERF_SAMPLE_BRANCH_NO_CYCLES;
}
static inline bool branch_sample_type(const struct perf_event *event)
{
return event->attr.branch_sample_type & PERF_SAMPLE_BRANCH_TYPE_SAVE;
}
static inline bool branch_sample_hw_index(const struct perf_event *event)
{
return event->attr.branch_sample_type & PERF_SAMPLE_BRANCH_HW_INDEX;
}
static inline bool branch_sample_priv(const struct perf_event *event)
{
return event->attr.branch_sample_type & PERF_SAMPLE_BRANCH_PRIV_SAVE;
}
static inline bool branch_sample_counters(const struct perf_event *event)
{
return event->attr.branch_sample_type & PERF_SAMPLE_BRANCH_COUNTERS;
}
static inline bool branch_sample_call_stack(const struct perf_event *event)
{
return event->attr.branch_sample_type & PERF_SAMPLE_BRANCH_CALL_STACK;
}
struct perf_sample_data {
/*
* Fields set by perf_sample_data_init() unconditionally,
* group so as to minimize the cachelines touched.
*/
u64 sample_flags;
u64 period;
u64 dyn_size;
/*
* Fields commonly set by __perf_event_header__init_id(),
* group so as to minimize the cachelines touched.
*/
u64 type;
struct {
u32 pid;
u32 tid;
} tid_entry;
u64 time;
u64 id;
struct {
u32 cpu;
u32 reserved;
} cpu_entry;
/*
* The other fields, optionally {set,used} by
* perf_{prepare,output}_sample().
*/
u64 ip;
struct perf_callchain_entry *callchain;
struct perf_raw_record *raw;
struct perf_branch_stack *br_stack;
u64 *br_stack_cntr;
union perf_sample_weight weight;
union perf_mem_data_src data_src;
u64 txn;
struct perf_regs regs_user;
struct perf_regs regs_intr;
u64 stack_user_size;
u64 stream_id;
u64 cgroup;
u64 addr;
u64 phys_addr;
u64 data_page_size;
u64 code_page_size;
u64 aux_size;
} ____cacheline_aligned;
/* default value for data source */
#define PERF_MEM_NA (PERF_MEM_S(OP, NA) |\
PERF_MEM_S(LVL, NA) |\
PERF_MEM_S(SNOOP, NA) |\
PERF_MEM_S(LOCK, NA) |\
PERF_MEM_S(TLB, NA) |\
PERF_MEM_S(LVLNUM, NA))
static inline void perf_sample_data_init(struct perf_sample_data *data,
u64 addr, u64 period)
{
/* remaining struct members initialized in perf_prepare_sample() */
data->sample_flags = PERF_SAMPLE_PERIOD;
data->period = period;
data->dyn_size = 0;
if (addr) {
data->addr = addr;
data->sample_flags |= PERF_SAMPLE_ADDR;
}
}
static inline void perf_sample_save_callchain(struct perf_sample_data *data,
struct perf_event *event,
struct pt_regs *regs)
{
int size = 1;
if (!(event->attr.sample_type & PERF_SAMPLE_CALLCHAIN))
return;
if (WARN_ON_ONCE(data->sample_flags & PERF_SAMPLE_CALLCHAIN))
return;
data->callchain = perf_callchain(event, regs);
size += data->callchain->nr;
data->dyn_size += size * sizeof(u64);
data->sample_flags |= PERF_SAMPLE_CALLCHAIN;
}
static inline void perf_sample_save_raw_data(struct perf_sample_data *data,
struct perf_event *event,
struct perf_raw_record *raw)
{
struct perf_raw_frag *frag = &raw->frag;
u32 sum = 0;
int size;
if (!(event->attr.sample_type & PERF_SAMPLE_RAW))
return;
if (WARN_ON_ONCE(data->sample_flags & PERF_SAMPLE_RAW))
return;
do {
sum += frag->size;
if (perf_raw_frag_last(frag))
break;
frag = frag->next;
} while (1);
size = round_up(sum + sizeof(u32), sizeof(u64));
raw->size = size - sizeof(u32);
frag->pad = raw->size - sum;
data->raw = raw;
data->dyn_size += size;
data->sample_flags |= PERF_SAMPLE_RAW;
}
static inline bool has_branch_stack(struct perf_event *event)
{
return event->attr.sample_type & PERF_SAMPLE_BRANCH_STACK;
}
static inline void perf_sample_save_brstack(struct perf_sample_data *data,
struct perf_event *event,
struct perf_branch_stack *brs,
u64 *brs_cntr)
{
int size = sizeof(u64); /* nr */
if (!has_branch_stack(event))
return;
if (WARN_ON_ONCE(data->sample_flags & PERF_SAMPLE_BRANCH_STACK))
return;
if (branch_sample_hw_index(event))
size += sizeof(u64);
brs->nr = min_t(u16, event->attr.sample_max_stack, brs->nr);
size += brs->nr * sizeof(struct perf_branch_entry);
/*
* The extension space for counters is appended after the
* struct perf_branch_stack. It is used to store the occurrences
* of events of each branch.
*/
if (brs_cntr)
size += brs->nr * sizeof(u64);
data->br_stack = brs;
data->br_stack_cntr = brs_cntr;
data->dyn_size += size;
data->sample_flags |= PERF_SAMPLE_BRANCH_STACK;
}
static inline u32 perf_sample_data_size(struct perf_sample_data *data,
struct perf_event *event)
{
u32 size = sizeof(struct perf_event_header);
size += event->header_size + event->id_header_size;
size += data->dyn_size;
return size;
}
/*
* Clear all bitfields in the perf_branch_entry.
* The to and from fields are not cleared because they are
* systematically modified by caller.
*/
static inline void perf_clear_branch_entry_bitfields(struct perf_branch_entry *br)
{
br->mispred = 0;
br->predicted = 0;
br->in_tx = 0;
br->abort = 0;
br->cycles = 0;
br->type = 0;
br->spec = PERF_BR_SPEC_NA;
br->reserved = 0;
}
extern void perf_output_sample(struct perf_output_handle *handle,
struct perf_event_header *header,
struct perf_sample_data *data,
struct perf_event *event);
extern void perf_prepare_sample(struct perf_sample_data *data,
struct perf_event *event,
struct pt_regs *regs);
extern void perf_prepare_header(struct perf_event_header *header,
struct perf_sample_data *data,
struct perf_event *event,
struct pt_regs *regs);
extern int perf_event_overflow(struct perf_event *event,
struct perf_sample_data *data,
struct pt_regs *regs);
extern void perf_event_output_forward(struct perf_event *event,
struct perf_sample_data *data,
struct pt_regs *regs);
extern void perf_event_output_backward(struct perf_event *event,
struct perf_sample_data *data,
struct pt_regs *regs);
extern int perf_event_output(struct perf_event *event,
struct perf_sample_data *data,
struct pt_regs *regs);
static inline bool
is_default_overflow_handler(struct perf_event *event)
{
perf_overflow_handler_t overflow_handler = event->overflow_handler;
if (likely(overflow_handler == perf_event_output_forward))
return true;
if (unlikely(overflow_handler == perf_event_output_backward))
return true;
return false;
}
extern void
perf_event_header__init_id(struct perf_event_header *header,
struct perf_sample_data *data,
struct perf_event *event);
extern void
perf_event__output_id_sample(struct perf_event *event,
struct perf_output_handle *handle,
struct perf_sample_data *sample);
extern void
perf_log_lost_samples(struct perf_event *event, u64 lost);
static inline bool event_has_any_exclude_flag(struct perf_event *event)
{
struct perf_event_attr *attr = &event->attr;
return attr->exclude_idle || attr->exclude_user ||
attr->exclude_kernel || attr->exclude_hv ||
attr->exclude_guest || attr->exclude_host;
}
static inline bool is_sampling_event(struct perf_event *event)
{
return event->attr.sample_period != 0;
}
/*
* Return 1 for a software event, 0 for a hardware event
*/
static inline int is_software_event(struct perf_event *event)
{
return event->event_caps & PERF_EV_CAP_SOFTWARE;
}
/*
* Return 1 for event in sw context, 0 for event in hw context
*/
static inline int in_software_context(struct perf_event *event)
{
return event->pmu_ctx->pmu->task_ctx_nr == perf_sw_context;
}
static inline int is_exclusive_pmu(struct pmu *pmu)
{
return pmu->capabilities & PERF_PMU_CAP_EXCLUSIVE;
}
extern struct static_key perf_swevent_enabled[PERF_COUNT_SW_MAX];
extern void ___perf_sw_event(u32, u64, struct pt_regs *, u64);
extern void __perf_sw_event(u32, u64, struct pt_regs *, u64);
#ifndef perf_arch_fetch_caller_regs
static inline void perf_arch_fetch_caller_regs(struct pt_regs *regs, unsigned long ip) { }
#endif
/*
* When generating a perf sample in-line, instead of from an interrupt /
* exception, we lack a pt_regs. This is typically used from software events
* like: SW_CONTEXT_SWITCHES, SW_MIGRATIONS and the tie-in with tracepoints.
*
* We typically don't need a full set, but (for x86) do require:
* - ip for PERF_SAMPLE_IP
* - cs for user_mode() tests
* - sp for PERF_SAMPLE_CALLCHAIN
* - eflags for MISC bits and CALLCHAIN (see: perf_hw_regs())
*
* NOTE: assumes @regs is otherwise already 0 filled; this is important for
* things like PERF_SAMPLE_REGS_INTR.
*/
static inline void perf_fetch_caller_regs(struct pt_regs *regs)
{
perf_arch_fetch_caller_regs(regs, CALLER_ADDR0);
}
static __always_inline void
perf_sw_event(u32 event_id, u64 nr, struct pt_regs *regs, u64 addr)
{
if (static_key_false(&perf_swevent_enabled[event_id]))
__perf_sw_event(event_id, nr, regs, addr);
}
DECLARE_PER_CPU(struct pt_regs, __perf_regs[4]);
/*
* 'Special' version for the scheduler, it hard assumes no recursion,
* which is guaranteed by us not actually scheduling inside other swevents
* because those disable preemption.
*/
static __always_inline void __perf_sw_event_sched(u32 event_id, u64 nr, u64 addr)
{
struct pt_regs *regs = this_cpu_ptr(&__perf_regs[0]);
perf_fetch_caller_regs(regs);
___perf_sw_event(event_id, nr, regs, addr);
}
extern struct static_key_false perf_sched_events;
static __always_inline bool __perf_sw_enabled(int swevt)
{
return static_key_false(&perf_swevent_enabled[swevt]);
}
static inline void perf_event_task_migrate(struct task_struct *task)
{
if (__perf_sw_enabled(PERF_COUNT_SW_CPU_MIGRATIONS))
task->sched_migrated = 1;
}
static inline void perf_event_task_sched_in(struct task_struct *prev,
struct task_struct *task)
{
if (static_branch_unlikely(&perf_sched_events))
__perf_event_task_sched_in(prev, task);
if (__perf_sw_enabled(PERF_COUNT_SW_CPU_MIGRATIONS) &&
task->sched_migrated) {
__perf_sw_event_sched(PERF_COUNT_SW_CPU_MIGRATIONS, 1, 0);
task->sched_migrated = 0;
}
}
static inline void perf_event_task_sched_out(struct task_struct *prev,
struct task_struct *next)
{
if (__perf_sw_enabled(PERF_COUNT_SW_CONTEXT_SWITCHES))
__perf_sw_event_sched(PERF_COUNT_SW_CONTEXT_SWITCHES, 1, 0);
#ifdef CONFIG_CGROUP_PERF
if (__perf_sw_enabled(PERF_COUNT_SW_CGROUP_SWITCHES) &&
perf_cgroup_from_task(prev, NULL) !=
perf_cgroup_from_task(next, NULL))
__perf_sw_event_sched(PERF_COUNT_SW_CGROUP_SWITCHES, 1, 0);
#endif
if (static_branch_unlikely(&perf_sched_events))
__perf_event_task_sched_out(prev, next);
}
extern void perf_event_mmap(struct vm_area_struct *vma);
extern void perf_event_ksymbol(u16 ksym_type, u64 addr, u32 len,
bool unregister, const char *sym);
extern void perf_event_bpf_event(struct bpf_prog *prog,
enum perf_bpf_event_type type,
u16 flags);
#define PERF_GUEST_ACTIVE 0x01
#define PERF_GUEST_USER 0x02
struct perf_guest_info_callbacks {
unsigned int (*state)(void);
unsigned long (*get_ip)(void);
unsigned int (*handle_intel_pt_intr)(void);
};
#ifdef CONFIG_GUEST_PERF_EVENTS
extern struct perf_guest_info_callbacks __rcu *perf_guest_cbs;
DECLARE_STATIC_CALL(__perf_guest_state, *perf_guest_cbs->state);
DECLARE_STATIC_CALL(__perf_guest_get_ip, *perf_guest_cbs->get_ip);
DECLARE_STATIC_CALL(__perf_guest_handle_intel_pt_intr, *perf_guest_cbs->handle_intel_pt_intr);
static inline unsigned int perf_guest_state(void)
{
return static_call(__perf_guest_state)();
}
static inline unsigned long perf_guest_get_ip(void)
{
return static_call(__perf_guest_get_ip)();
}
static inline unsigned int perf_guest_handle_intel_pt_intr(void)
{
return static_call(__perf_guest_handle_intel_pt_intr)();
}
extern void perf_register_guest_info_callbacks(struct perf_guest_info_callbacks *cbs);
extern void perf_unregister_guest_info_callbacks(struct perf_guest_info_callbacks *cbs);
#else /* !CONFIG_GUEST_PERF_EVENTS: */
static inline unsigned int perf_guest_state(void) { return 0; }
static inline unsigned long perf_guest_get_ip(void) { return 0; }
static inline unsigned int perf_guest_handle_intel_pt_intr(void) { return 0; }
#endif /* !CONFIG_GUEST_PERF_EVENTS */
extern void perf_event_exec(void);
extern void perf_event_comm(struct task_struct *tsk, bool exec);
extern void perf_event_namespaces(struct task_struct *tsk);
extern void perf_event_fork(struct task_struct *tsk);
extern void perf_event_text_poke(const void *addr,
const void *old_bytes, size_t old_len,
const void *new_bytes, size_t new_len);
/* Callchains */
DECLARE_PER_CPU(struct perf_callchain_entry, perf_callchain_entry);
extern void perf_callchain_user(struct perf_callchain_entry_ctx *entry, struct pt_regs *regs);
extern void perf_callchain_kernel(struct perf_callchain_entry_ctx *entry, struct pt_regs *regs);
extern struct perf_callchain_entry *
get_perf_callchain(struct pt_regs *regs, bool kernel, bool user,
u32 max_stack, bool crosstask, bool add_mark);
extern int get_callchain_buffers(int max_stack);
extern void put_callchain_buffers(void);
extern struct perf_callchain_entry *get_callchain_entry(int *rctx);
extern void put_callchain_entry(int rctx);
extern int sysctl_perf_event_max_stack;
extern int sysctl_perf_event_max_contexts_per_stack;
static inline int perf_callchain_store_context(struct perf_callchain_entry_ctx *ctx, u64 ip)
{
if (ctx->contexts < sysctl_perf_event_max_contexts_per_stack) {
struct perf_callchain_entry *entry = ctx->entry;
entry->ip[entry->nr++] = ip;
++ctx->contexts;
return 0;
} else {
ctx->contexts_maxed = true;
return -1; /* no more room, stop walking the stack */
}
}
static inline int perf_callchain_store(struct perf_callchain_entry_ctx *ctx, u64 ip)
{
if (ctx->nr < ctx->max_stack && !ctx->contexts_maxed) {
struct perf_callchain_entry *entry = ctx->entry;
entry->ip[entry->nr++] = ip;
++ctx->nr;
return 0;
} else {
return -1; /* no more room, stop walking the stack */
}
}
extern int sysctl_perf_event_paranoid;
extern int sysctl_perf_event_sample_rate;
extern void perf_sample_event_took(u64 sample_len_ns);
/* Access to perf_event_open(2) syscall. */
#define PERF_SECURITY_OPEN 0
/* Finer grained perf_event_open(2) access control. */
#define PERF_SECURITY_CPU 1
#define PERF_SECURITY_KERNEL 2
#define PERF_SECURITY_TRACEPOINT 3
static inline int perf_is_paranoid(void)
{
return sysctl_perf_event_paranoid > -1;
}
extern int perf_allow_kernel(void);
static inline int perf_allow_cpu(void)
{
if (sysctl_perf_event_paranoid > 0 && !perfmon_capable())
return -EACCES;
return security_perf_event_open(PERF_SECURITY_CPU);
}
static inline int perf_allow_tracepoint(void)
{
if (sysctl_perf_event_paranoid > -1 && !perfmon_capable())
return -EPERM;
return security_perf_event_open(PERF_SECURITY_TRACEPOINT);
}
extern int perf_exclude_event(struct perf_event *event, struct pt_regs *regs);
extern void perf_event_init(void);
extern void perf_tp_event(u16 event_type, u64 count, void *record,
int entry_size, struct pt_regs *regs,
struct hlist_head *head, int rctx,
struct task_struct *task);
extern void perf_bp_event(struct perf_event *event, void *data);
extern unsigned long perf_misc_flags(struct perf_event *event, struct pt_regs *regs);
extern unsigned long perf_instruction_pointer(struct perf_event *event,
struct pt_regs *regs);
#ifndef perf_arch_misc_flags
# define perf_arch_misc_flags(regs) \
(user_mode(regs) ? PERF_RECORD_MISC_USER : PERF_RECORD_MISC_KERNEL)
# define perf_arch_instruction_pointer(regs) instruction_pointer(regs)
#endif
#ifndef perf_arch_bpf_user_pt_regs
# define perf_arch_bpf_user_pt_regs(regs) regs
#endif
#ifndef perf_arch_guest_misc_flags
static inline unsigned long perf_arch_guest_misc_flags(struct pt_regs *regs)
{
unsigned long guest_state = perf_guest_state();
if (!(guest_state & PERF_GUEST_ACTIVE))
return 0;
if (guest_state & PERF_GUEST_USER)
return PERF_RECORD_MISC_GUEST_USER;
else
return PERF_RECORD_MISC_GUEST_KERNEL;
}
# define perf_arch_guest_misc_flags(regs) perf_arch_guest_misc_flags(regs)
#endif
static inline bool needs_branch_stack(struct perf_event *event)
{
return event->attr.branch_sample_type != 0;
}
static inline bool has_aux(struct perf_event *event)
{
return event->pmu && event->pmu->setup_aux;
}
static inline bool has_aux_action(struct perf_event *event)
{
return event->attr.aux_sample_size ||
event->attr.aux_pause ||
event->attr.aux_resume;
}
static inline bool is_write_backward(struct perf_event *event)
{
return !!event->attr.write_backward;
}
static inline bool has_addr_filter(struct perf_event *event)
{
return event->pmu->nr_addr_filters;
}
/*
* An inherited event uses parent's filters
*/
static inline struct perf_addr_filters_head *
perf_event_addr_filters(struct perf_event *event)
{
struct perf_addr_filters_head *ifh = &event->addr_filters;
if (event->parent)
ifh = &event->parent->addr_filters;
return ifh;
}
static inline struct fasync_struct **perf_event_fasync(struct perf_event *event)
{
/* Only the parent has fasync state */
if (event->parent)
event = event->parent;
return &event->fasync;
}
extern void perf_event_addr_filters_sync(struct perf_event *event);
extern void perf_report_aux_output_id(struct perf_event *event, u64 hw_id);
extern int perf_output_begin(struct perf_output_handle *handle,
struct perf_sample_data *data,
struct perf_event *event, unsigned int size);
extern int perf_output_begin_forward(struct perf_output_handle *handle,
struct perf_sample_data *data,
struct perf_event *event,
unsigned int size);
extern int perf_output_begin_backward(struct perf_output_handle *handle,
struct perf_sample_data *data,
struct perf_event *event,
unsigned int size);
extern void perf_output_end(struct perf_output_handle *handle);
extern unsigned int perf_output_copy(struct perf_output_handle *handle,
const void *buf, unsigned int len);
extern unsigned int perf_output_skip(struct perf_output_handle *handle,
unsigned int len);
extern long perf_output_copy_aux(struct perf_output_handle *aux_handle,
struct perf_output_handle *handle,
unsigned long from, unsigned long to);
extern int perf_swevent_get_recursion_context(void);
extern void perf_swevent_put_recursion_context(int rctx);
extern u64 perf_swevent_set_period(struct perf_event *event);
extern void perf_event_enable(struct perf_event *event);
extern void perf_event_disable(struct perf_event *event);
extern void perf_event_disable_local(struct perf_event *event);
extern void perf_event_disable_inatomic(struct perf_event *event);
extern void perf_event_task_tick(void);
extern int perf_event_account_interrupt(struct perf_event *event);
extern int perf_event_period(struct perf_event *event, u64 value);
extern u64 perf_event_pause(struct perf_event *event, bool reset);
#else /* !CONFIG_PERF_EVENTS: */
static inline void *
perf_aux_output_begin(struct perf_output_handle *handle,
struct perf_event *event) { return NULL; }
static inline void
perf_aux_output_end(struct perf_output_handle *handle, unsigned long size)
{ }
static inline int
perf_aux_output_skip(struct perf_output_handle *handle,
unsigned long size) { return -EINVAL; }
static inline void *
perf_get_aux(struct perf_output_handle *handle) { return NULL; }
static inline void
perf_event_task_migrate(struct task_struct *task) { }
static inline void
perf_event_task_sched_in(struct task_struct *prev,
struct task_struct *task) { }
static inline void
perf_event_task_sched_out(struct task_struct *prev,
struct task_struct *next) { }
static inline int perf_event_init_task(struct task_struct *child,
u64 clone_flags) { return 0; }
static inline void perf_event_exit_task(struct task_struct *child) { }
static inline void perf_event_free_task(struct task_struct *task) { }
static inline void perf_event_delayed_put(struct task_struct *task) { }
static inline struct file *perf_event_get(unsigned int fd) { return ERR_PTR(-EINVAL); }
static inline const struct perf_event *perf_get_event(struct file *file)
{
return ERR_PTR(-EINVAL);
}
static inline const struct perf_event_attr *perf_event_attrs(struct perf_event *event)
{
return ERR_PTR(-EINVAL);
}
static inline int perf_event_read_local(struct perf_event *event, u64 *value,
u64 *enabled, u64 *running)
{
return -EINVAL;
}
static inline void perf_event_print_debug(void) { }
static inline int perf_event_task_disable(void) { return -EINVAL; }
static inline int perf_event_task_enable(void) { return -EINVAL; }
static inline int perf_event_refresh(struct perf_event *event, int refresh)
{
return -EINVAL;
}
static inline void
perf_sw_event(u32 event_id, u64 nr, struct pt_regs *regs, u64 addr) { }
static inline void
perf_bp_event(struct perf_event *event, void *data) { }
static inline void perf_event_mmap(struct vm_area_struct *vma) { }
typedef int (perf_ksymbol_get_name_f)(char *name, int name_len, void *data);
static inline void perf_event_ksymbol(u16 ksym_type, u64 addr, u32 len,
bool unregister, const char *sym) { }
static inline void perf_event_bpf_event(struct bpf_prog *prog,
enum perf_bpf_event_type type,
u16 flags) { }
static inline void perf_event_exec(void) { }
static inline void perf_event_comm(struct task_struct *tsk, bool exec) { }
static inline void perf_event_namespaces(struct task_struct *tsk) { }
static inline void perf_event_fork(struct task_struct *tsk) { }
static inline void perf_event_text_poke(const void *addr,
const void *old_bytes,
size_t old_len,
const void *new_bytes,
size_t new_len) { }
static inline void perf_event_init(void) { }
static inline int perf_swevent_get_recursion_context(void) { return -1; }
static inline void perf_swevent_put_recursion_context(int rctx) { }
static inline u64 perf_swevent_set_period(struct perf_event *event) { return 0; }
static inline void perf_event_enable(struct perf_event *event) { }
static inline void perf_event_disable(struct perf_event *event) { }
static inline int __perf_event_disable(void *info) { return -1; }
static inline void perf_event_task_tick(void) { }
static inline int perf_event_release_kernel(struct perf_event *event) { return 0; }
static inline int
perf_event_period(struct perf_event *event, u64 value) { return -EINVAL; }
static inline u64
perf_event_pause(struct perf_event *event, bool reset) { return 0; }
static inline int
perf_exclude_event(struct perf_event *event, struct pt_regs *regs) { return 0; }
#endif /* !CONFIG_PERF_EVENTS */
#if defined(CONFIG_PERF_EVENTS) && defined(CONFIG_CPU_SUP_INTEL)
extern void perf_restore_debug_store(void);
#else
static inline void perf_restore_debug_store(void) { }
#endif
#define perf_output_put(handle, x) perf_output_copy((handle), &(x), sizeof(x))
struct perf_pmu_events_attr {
struct device_attribute attr;
u64 id;
const char *event_str;
};
struct perf_pmu_events_ht_attr {
struct device_attribute attr;
u64 id;
const char *event_str_ht;
const char *event_str_noht;
};
struct perf_pmu_events_hybrid_attr {
struct device_attribute attr;
u64 id;
const char *event_str;
u64 pmu_type;
};
struct perf_pmu_format_hybrid_attr {
struct device_attribute attr;
u64 pmu_type;
};
ssize_t perf_event_sysfs_show(struct device *dev, struct device_attribute *attr,
char *page);
#define PMU_EVENT_ATTR(_name, _var, _id, _show) \
static struct perf_pmu_events_attr _var = { \
.attr = __ATTR(_name, 0444, _show, NULL), \
.id = _id, \
};
#define PMU_EVENT_ATTR_STRING(_name, _var, _str) \
static struct perf_pmu_events_attr _var = { \
.attr = __ATTR(_name, 0444, perf_event_sysfs_show, NULL), \
.id = 0, \
.event_str = _str, \
};
#define PMU_EVENT_ATTR_ID(_name, _show, _id) \
(&((struct perf_pmu_events_attr[]) { \
{ .attr = __ATTR(_name, 0444, _show, NULL), \
.id = _id, } \
})[0].attr.attr)
#define PMU_FORMAT_ATTR_SHOW(_name, _format) \
static ssize_t \
_name##_show(struct device *dev, \
struct device_attribute *attr, \
char *page) \
{ \
BUILD_BUG_ON(sizeof(_format) >= PAGE_SIZE); \
return sprintf(page, _format "\n"); \
} \
#define PMU_FORMAT_ATTR(_name, _format) \
PMU_FORMAT_ATTR_SHOW(_name, _format) \
\
static struct device_attribute format_attr_##_name = __ATTR_RO(_name)
/* Performance counter hotplug functions */
#ifdef CONFIG_PERF_EVENTS
extern int perf_event_init_cpu(unsigned int cpu);
extern int perf_event_exit_cpu(unsigned int cpu);
#else
# define perf_event_init_cpu NULL
# define perf_event_exit_cpu NULL
#endif
extern void arch_perf_update_userpage(struct perf_event *event,
struct perf_event_mmap_page *userpg,
u64 now);
/*
* Snapshot branch stack on software events.
*
* Branch stack can be very useful in understanding software events. For
* example, when a long function, e.g. sys_perf_event_open, returns an
* errno, it is not obvious why the function failed. Branch stack could
* provide very helpful information in this type of scenarios.
*
* On software event, it is necessary to stop the hardware branch recorder
* fast. Otherwise, the hardware register/buffer will be flushed with
* entries of the triggering event. Therefore, static call is used to
* stop the hardware recorder.
*/
/*
* cnt is the number of entries allocated for entries.
* Return number of entries copied to .
*/
typedef int (perf_snapshot_branch_stack_t)(struct perf_branch_entry *entries,
unsigned int cnt);
DECLARE_STATIC_CALL(perf_snapshot_branch_stack, perf_snapshot_branch_stack_t);
#ifndef PERF_NEEDS_LOPWR_CB
static inline void perf_lopwr_cb(bool mode)
{
}
#endif
#endif /* _LINUX_PERF_EVENT_H */
]
|
| 20/2 |
2026/05/18 09:38 |
llm |
expert |
0m
Model:
Tokens:
input: 42108
output: 44
thoughts: 697
Thoughts: **Analyzing the perf_event_throttle_group Lockdep Warning**
Okay, so I'm looking at this warning that's popping up, specifically the one triggered at `kernel/events/core.c:2703`. That line is using the `for_each_sibling_event` macro, which is defined to iterate through a list of sibling events. I see that the macro includes `lockdep_assert_event_ctx(event)` as its first operation. Digging into that macro, which is only active if `CONFIG_PROVE_LOCKING` is enabled, I'm checking that hard IRQs are *not* enabled when a context mutex is not held.
The warning is triggered by `lockdep_assert_event_ctx(leader)`, where `leader` is the `group_leader` of the event. I need to figure out why this assertion is failing. The code path that leads here is `perf_event_throttle_group`, which then calls `for_each_sibling_event`. `perf_event_throttle_group` is called by `__perf_event_account_interrupt`. This function is invoked by `__perf_event_overflow`. Further exploration leads to `perf_swevent_overflow`, then `perf_swevent_event`, then `perf_tp_event`, and finally `perf_trace_run_bpf_submit`.
Now, if I think about the context of execution, the key is `perf_trace_run_bpf_submit`. I was initially expecting this to be an interrupt context, where hard IRQs would be disabled. However, this is a tracepoint event, and tracepoints can be triggered from virtually anywhere in the kernel. This means it *is* possible for a tracepoint event to fire with hard IRQs enabled. If hard IRQs are enabled, and the `ctx->mutex` isn't held, then the `lockdep_assert_event_ctx` macro *will* warn. That's the likely scenario. I need to dig further into the contexts that can invoke `perf_trace_run_bpf_submit` to confirm. I should also check `__perf_event_account_interrupt` to gain a better understanding of the lock usage in that section.
|
| 21/2 |
2026/05/18 09:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:__perf_event_account_interrupt]
Results: map[SourceCode:10162: static int
10163: __perf_event_account_interrupt(struct perf_event *event, int throttle)
10164: {
10165: struct hw_perf_event *hwc = &event->hw;
10166: int ret = 0;
10167: u64 seq;
10168:
10169: seq = __this_cpu_read(perf_throttled_seq);
10170: if (seq != hwc->interrupts_seq) {
10171: hwc->interrupts_seq = seq;
10172: hwc->interrupts = 1;
10173: } else {
10174: hwc->interrupts++;
10175: }
10176:
10177: if (unlikely(throttle && hwc->interrupts >= max_samples_per_tick)) {
10178: __this_cpu_inc(perf_throttled_count);
10179: tick_dep_set_cpu(smp_processor_id(), TICK_DEP_BIT_PERF_EVENTS);
10180: perf_event_throttle_group(event);
10181: ret = 1;
10182: }
10183:
10184: if (event->attr.freq) {
10185: u64 now = perf_clock();
10186: s64 delta = now - hwc->freq_time_stamp;
10187:
10188: hwc->freq_time_stamp = now;
10189:
10190: if (delta > 0 && delta < 2*TICK_NSEC)
10191: perf_adjust_period(event, delta, hwc->last_period, true);
10192: }
10193:
10194: return ret;
10195: }
SourceFile:kernel/events/core.c]
|
| 22/2 |
2026/05/18 09:38 |
llm |
expert |
0m
Model:
Tokens:
input: 43905
output: 42
thoughts: 373
Thoughts: **Analyzing the Potential Lock Depletion in Performance Events**
Okay, so I'm tracing through this `perf_event_throttle_group` function. It gets called when the hardware counter's interrupts exceed a certain threshold, specifically when `hwc->interrupts` crosses `max_samples_per_tick`. I'm seeing that it iterates over sibling events using `for_each_sibling_event(sibling, leader)`. Ah, but before that loop even starts, there's a `lockdep_assert_event_ctx(event)` check.
Let's dissect this `lockdep_assert_event_ctx` macro... I see it uses the following code.
```c
# define lockdep_assert_event_ctx(event) \
WARN_ON_ONCE(__lockdep_enabled && \
(this_cpu_read(hardirqs_enabled) && \
lockdep_is_held(&(event)->ctx->mutex) != LOCK_STATE_HELD))
```
Basically, if lockdep is enabled, and if hard IRQs are enabled on this CPU at the time of the check ( `this_cpu_read(hardirqs_enabled)` is true ), the code expects that the `ctx->mutex` associated with the event is already held. If it isn't, there's a warning. My concern here is the context that calls `perf_event_throttle_group`. It's `__perf_event_account_interrupt`, which gets called by `__perf_event_overflow`.
Now, if `__perf_event_overflow` is triggered by a software event (like a tracepoint), hard IRQs *might* be enabled at that moment. So, if we end up in the software event path, we're in trouble. Then we call `perf_event_throttle_group` from this context, which requires that `ctx->mutex` be held when hard IRQs are enabled!
But does `perf_swevent_overflow` disable IRQs? I need to check the code for `perf_swevent_overflow` now. It's crucial for the mutex to be held when hard IRQs are enabled.
|
| 23/2 |
2026/05/18 09:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_swevent_overflow]
Results: map[SourceCode:10453: static void perf_swevent_overflow(struct perf_event *event, u64 overflow,
10454: struct perf_sample_data *data,
10455: struct pt_regs *regs)
10456: {
10457: struct hw_perf_event *hwc = &event->hw;
10458: int throttle = 0;
10459:
10460: if (!overflow)
10461: overflow = perf_swevent_set_period(event);
10462:
10463: if (hwc->interrupts == MAX_INTERRUPTS)
10464: return;
10465:
10466: for (; overflow; overflow--) {
10467: if (__perf_event_overflow(event, throttle,
10468: data, regs)) {
10469: /*
10470: * We inhibit the overflow from happening when
10471: * hwc->interrupts == MAX_INTERRUPTS.
10472: */
10473: break;
10474: }
10475: throttle = 1;
10476: }
10477: }
SourceFile:kernel/events/core.c]
|
| 24/2 |
2026/05/18 09:38 |
llm |
expert |
0m
Model:
Tokens:
input: 45177
output: 42
thoughts: 12
Thoughts: Okay, here's my interpretation of that thought process, assuming I'm an expert in this field, and rendered in the first person:
**Investigating `perf_swevent_event`**
Right, let's take a look at `perf_swevent_event`. I'm going to start there. This suggests we're dealing with performance monitoring, likely related to software events tracked by `perf`. My initial reaction is to think about the various contexts where this specific event might be critical. What's the context? What application are we troubleshooting or optimizing? Knowing the environment is key. Am I looking at something related to context switches, scheduling, synchronization primitives, or perhaps even application-level counters exposed through the perf infrastructure? Understanding the specific software event this refers to is critical to go any further. I need to refresh myself on how this event is being utilized in the current setup. I'm going to dig into the documentation and relevant code to get the lay of the land before making any assumptions.
|
| 25/2 |
2026/05/18 09:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_swevent_event]
Results: map[SourceCode:10479: static void perf_swevent_event(struct perf_event *event, u64 nr,
10480: struct perf_sample_data *data,
10481: struct pt_regs *regs)
10482: {
10483: struct hw_perf_event *hwc = &event->hw;
10484:
10485: local64_add(nr, &event->count);
10486:
10487: if (!regs)
10488: return;
10489:
10490: if (!is_sampling_event(event))
10491: return;
10492:
10493: if ((event->attr.sample_type & PERF_SAMPLE_PERIOD) && !event->attr.freq) {
10494: data->period = nr;
10495: return perf_swevent_overflow(event, 1, data, regs);
10496: } else
10497: data->period = event->hw.last_period;
10498:
10499: if (nr == 1 && hwc->sample_period == 1 && !event->attr.freq)
10500: return perf_swevent_overflow(event, 1, data, regs);
10501:
10502: if (local64_add_negative(nr, &hwc->period_left))
10503: return;
10504:
10505: perf_swevent_overflow(event, 0, data, regs);
10506: }
SourceFile:kernel/events/core.c]
|
| 26/2 |
2026/05/18 09:38 |
llm |
expert |
0m
Model:
Tokens:
input: 45925
output: 41
thoughts: 11
Thoughts: Okay, here's my interpretation of that thought process, assuming I'm deep in my domain of performance analysis:
**Diving into `perf_tp_event`**
Right, let's get right down to it. My initial thought is focused on a specific performance metric - `perf_tp_event`. I'm assuming that this is a key component, perhaps a specific event related to a tracepoint or performance monitoring event. Time to dive in. I need to understand what this thing is doing, how it's behaving, and what insights it can provide. This is likely the entry point to a whole chain of thought and analysis. Let's see what's what.
|
| 27/2 |
2026/05/18 09:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_tp_event]
Results: map[SourceCode:10984: void perf_tp_event(u16 event_type, u64 count, void *record, int entry_size,
10985: struct pt_regs *regs, struct hlist_head *head, int rctx,
10986: struct task_struct *task)
10987: {
10988: struct perf_sample_data data;
10989: struct perf_event *event;
10990:
10991: struct perf_raw_record raw = {
10992: .frag = {
10993: .size = entry_size,
10994: .data = record,
10995: },
10996: };
10997:
10998: perf_trace_buf_update(record, event_type);
10999:
11000: hlist_for_each_entry_rcu(event, head, hlist_entry) {
11001: if (perf_tp_event_match(event, &raw, regs)) {
11002: /*
11003: * Here use the same on-stack perf_sample_data,
11004: * some members in data are event-specific and
11005: * need to be re-computed for different sweveents.
11006: * Re-initialize data->sample_flags safely to avoid
11007: * the problem that next event skips preparing data
11008: * because data->sample_flags is set.
11009: */
11010: perf_sample_data_init(&data, 0, 0);
11011: perf_sample_save_raw_data(&data, event, &raw);
11012: perf_swevent_event(event, count, &data, regs);
11013: }
11014: }
11015:
11016: /*
11017: * If we got specified a target task, also iterate its context and
11018: * deliver this event there too.
11019: */
11020: if (task && task != current) {
11021: struct perf_event_context *ctx;
11022:
11023: rcu_read_lock();
11024: ctx = rcu_dereference(task->perf_event_ctxp);
11025: if (!ctx)
11026: goto unlock;
11027:
11028: raw_spin_lock(&ctx->lock);
11029: perf_tp_event_target_task(count, record, regs, &data, &raw, ctx);
11030: raw_spin_unlock(&ctx->lock);
11031: unlock:
11032: rcu_read_unlock();
11033: }
11034:
11035: perf_swevent_put_recursion_context(rctx);
11036: }
SourceFile:kernel/events/core.c]
|
| 28/2 |
2026/05/18 09:38 |
llm |
expert |
0m
Model:
Tokens:
input: 46989
output: 46
thoughts: 593
Thoughts: **Analyzing a Potential Lockdep False Positive in Perf Events**
Okay, let's break this down. It looks like `perf_tp_event` is being called from `perf_trace_run_bpf_submit`, which in turn gets called from `perf_trace_lock_acquire`. And `perf_trace_lock_acquire` is, of course, a tracepoint for lockdep. The stack trace clearly shows this chain of events. So, essentially, `rcu_read_lock` leads to `lock_acquire`, which eventually triggers this `perf_tp_event`.
The interesting part is that when we get to `perf_tp_event`, hard interrupts are *enabled*. This is because `rcu_read_lock` doesn't disable them. This means `this_cpu_read(hardirqs_enabled)` is returning true at that point in the execution. Now, the `perf_event_throttle_group` code path calls `for_each_sibling_event`, which in turn calls `lockdep_assert_event_ctx(event)`.
Here's where things get suspicious. `lockdep_assert_event_ctx` *requires* `ctx->mutex` to be held if `hardirqs_enabled` is true. But the code path up to `perf_tp_event` does *not* hold this mutex! This looks like a potential false positive from Lockdep. We really need to confirm if this is the case or if there is something I'm missing. Time to delve into the implementation of `lockdep_assert_event_ctx`. Let's see what's actually going on in there.
|
| 29/2 |
2026/05/18 09:38 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:include/linux/perf_event.h IncludeLines:true Name:lockdep_assert_event_ctx]
Results: map[SourceCode: SourceFile:]
|
| 30/2 |
2026/05/18 09:38 |
llm |
expert |
0m
Model:
Tokens:
input: 48004
output: 42
thoughts: 539
Thoughts: **Analyzing `lockdep_assert_event_ctx` and `sibling_list` in the context of `perf_event_throttle_group`**
Okay, I've just reviewed the `lockdep_assert_event_ctx` macro, and it's pretty straightforward. Under `CONFIG_PROVE_LOCKING`, it asserts that if hard IRQs are enabled, the `ctx->mutex` must be held. This is where things get interesting in relation to `perf_event_throttle_group`, which gets called from `__perf_event_account_interrupt`. My immediate question is whether iterating `sibling_list` within `perf_event_throttle_group` is safe without holding `ctx->mutex` if hard IRQs are enabled.
The comments for `sibling_list` are useful. They say iteration needs *either* `ctx->mutex` or `ctx->lock`. Initially, I thought that since `ctx->lock` is IRQ safe, disabling hard IRQs would be sufficient. I considered that since the CPU modifying the list holds the `ctx->lock` and disables IRQs, then an IPI couldn't modify the list on the same CPU, so we're safe.
However, I've got to think through this a bit more. We only disable IRQs on the *local* CPU. `ctx->lock` is a raw spinlock. We're not actually holding the spinlock if we just disable IRQs! That means another CPU could potentially modify the list while we're iterating. This seemed like a potential race.
Ah, but the list is only ever modified by the CPU or task that owns the context. That seems safe in and of itself.
But, back to my original question about `perf_event_throttle_group`. It's called from `__perf_event_account_interrupt`, but I'm wrong! It actually gets called from `__perf_event_overflow`. Hmm. Now, where does *that* leave us?
|
| 31/2 |
2026/05/18 09:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:__perf_event_overflow]
Results: map[SourceCode:10313: static int __perf_event_overflow(struct perf_event *event,
10314: int throttle, struct perf_sample_data *data,
10315: struct pt_regs *regs)
10316: {
10317: int events = atomic_read(&event->event_limit);
10318: int ret = 0;
10319:
10320: /*
10321: * Non-sampling counters might still use the PMI to fold short
10322: * hardware counters, ignore those.
10323: */
10324: if (unlikely(!is_sampling_event(event)))
10325: return 0;
10326:
10327: ret = __perf_event_account_interrupt(event, throttle);
10328:
10329: if (event->attr.aux_pause)
10330: perf_event_aux_pause(event->aux_event, true);
10331:
10332: if (event->prog && event->prog->type == BPF_PROG_TYPE_PERF_EVENT &&
10333: !bpf_overflow_handler(event, data, regs))
10334: goto out;
10335:
10336: /*
10337: * XXX event_limit might not quite work as expected on inherited
10338: * events
10339: */
10340:
10341: event->pending_kill = POLL_IN;
10342: if (events && atomic_dec_and_test(&event->event_limit)) {
10343: ret = 1;
10344: event->pending_kill = POLL_HUP;
10345: perf_event_disable_inatomic(event);
10346: event->pmu->stop(event, 0);
10347: }
10348:
10349: if (event->attr.sigtrap) {
10350: /*
10351: * The desired behaviour of sigtrap vs invalid samples is a bit
10352: * tricky; on the one hand, one should not loose the SIGTRAP if
10353: * it is the first event, on the other hand, we should also not
10354: * trigger the WARN or override the data address.
10355: */
10356: bool valid_sample = sample_is_allowed(event, regs);
10357: unsigned int pending_id = 1;
10358: enum task_work_notify_mode notify_mode;
10359:
10360: if (regs)
10361: pending_id = hash32_ptr((void *)instruction_pointer(regs)) ?: 1;
10362:
10363: notify_mode = in_nmi() ? TWA_NMI_CURRENT : TWA_RESUME;
10364:
10365: if (!event->pending_work &&
10366: !task_work_add(current, &event->pending_task, notify_mode)) {
10367: event->pending_work = pending_id;
10368: local_inc(&event->ctx->nr_no_switch_fast);
10369: WARN_ON_ONCE(!atomic_long_inc_not_zero(&event->refcount));
10370:
10371: event->pending_addr = 0;
10372: if (valid_sample && (data->sample_flags & PERF_SAMPLE_ADDR))
10373: event->pending_addr = data->addr;
10374:
10375: } else if (event->attr.exclude_kernel && valid_sample) {
10376: /*
10377: * Should not be able to return to user space without
10378: * consuming pending_work; with exceptions:
10379: *
10380: * 1. Where !exclude_kernel, events can overflow again
10381: * in the kernel without returning to user space.
10382: *
10383: * 2. Events that can overflow again before the IRQ-
10384: * work without user space progress (e.g. hrtimer).
10385: * To approximate progress (with false negatives),
10386: * check 32-bit hash of the current IP.
10387: */
10388: WARN_ON_ONCE(event->pending_work != pending_id);
10389: }
10390: }
10391:
10392: READ_ONCE(event->overflow_handler)(event, data, regs);
10393:
10394: if (*perf_event_fasync(event) && event->pending_kill) {
10395: event->pending_wakeup = 1;
10396: irq_work_queue(&event->pending_irq);
10397: }
10398: out:
10399: if (event->attr.aux_resume)
10400: perf_event_aux_pause(event->aux_event, false);
10401:
10402: return ret;
10403: }
SourceFile:kernel/events/core.c]
|
| 32/2 |
2026/05/18 09:38 |
llm |
expert |
0m
Model:
Tokens:
input: 50536
output: 46
thoughts: 1290
Thoughts: **A Deep Dive into a Potential Data Race in the Perf Event System**
Okay, so here's what I've got. I'm tracing the execution path from `__perf_event_overflow` back through the call chain. It goes like this: `__perf_event_overflow` is called by `perf_swevent_overflow`, which is called by `perf_swevent_event`, which in turn is called by `perf_tp_event`. `perf_tp_event` is called by `perf_trace_run_bpf_submit`, which then is called by `perf_trace_lock_acquire`, and onward down the line to `rcu_read_lock`.
The crucial point: `rcu_read_lock` *doesn't* disable interrupts (IRQs). This means `perf_tp_event` is being called with IRQs enabled. This is where things get interesting. From `perf_tp_event`, the flow goes through `perf_swevent_event`, which calls `perf_swevent_overflow` then `__perf_event_overflow`, which triggers `__perf_event_account_interrupt`. This calls `perf_event_throttle_group`.
Within `perf_event_throttle_group`, we're iterating over `sibling_list` using `for_each_sibling_event`, which in turn calls `lockdep_assert_event_ctx`. And that function flags a warning if `hardirqs_enabled` is true, but `ctx->mutex` is *not* held. That warning seems like a false positive. We're not holding `ctx->mutex`, so the warning seems accurate; however, the comment says, essentially, that since `ctx->lock` is IRQ-safe, we only need IRQs disabled to protect `sibling_list`.
But here's the kicker: If IRQs are *not* disabled, then an IPI can interrupt us and modify the list! We're in `rcu_read_lock`, so we don't hold either `ctx->lock` or `ctx->mutex`, and IRQs are enabled! An IPI could arrive, grab `ctx->lock`, and modify `sibling_list` while we're iterating it! Bingo. This looks like a data race: a use-after-free or list corruption scenario, potentially leading to a crash.
Now, let's look at `perf_tp_event`. It uses `hlist_for_each_entry_rcu` to iterate the `head`, meaning the `event` is RCU-protected. But crucially, the `event->sibling_list` is *not* protected by RCU, it's supposed to be protected by either `ctx->lock` or `ctx->mutex`, or by disabling IRQs. And `perf_tp_event` *does* not disable IRQs. So, when the code goes to `perf_swevent_event`, which doesn't disable IRQs either, `perf_event_throttle_group` is getting called *with interrupts enabled* and it iterates `sibling_list` without any locks! That's the core of the issue.
Can `sibling_list` actually be modified by another CPU? Yes! `perf_group_detach` or `perf_group_attach` modify `sibling_list` by holding `ctx->mutex` and `ctx->lock`. If we're *not* holding either of those locks, and IRQs are enabled, we have a race condition.
The comment claims that disabling IRQs is sufficient, because `ctx->lock` is an IRQ-safe lock and is only held by the CPU that's modifying the list. That's true, because `ctx->lock` is per-CPU or per-task. If we disable IRQs, we prevent IPIs from interrupting us. *But here*, IRQs are *not* disabled! An IPI could interrupt `perf_event_throttle_group` and modify `sibling_list` while we are iterating it! It's a real bug: `perf_event_throttle_group` can crash.
So, why doesn't `perf_tp_event` disable IRQs? Well, tracepoints usually disable preemption, but *not* IRQs. But does `perf_trace_run_bpf_submit` disable preemption? Let's check...
|
| 33/2 |
2026/05/18 09:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_trace_run_bpf_submit]
Results: map[SourceCode:10924: void perf_trace_run_bpf_submit(void *raw_data, int size, int rctx,
10925: struct trace_event_call *call, u64 count,
10926: struct pt_regs *regs, struct hlist_head *head,
10927: struct task_struct *task)
10928: {
10929: if (bpf_prog_array_valid(call)) {
10930: *(struct pt_regs **)raw_data = regs;
10931: if (!trace_call_bpf(call, raw_data) || hlist_empty(head)) {
10932: perf_swevent_put_recursion_context(rctx);
10933: return;
10934: }
10935: }
10936: perf_tp_event(call->event.type, count, raw_data, size, regs, head,
10937: rctx, task);
10938: }
SourceFile:kernel/events/core.c]
|
| 34/2 |
2026/05/18 09:38 |
llm |
expert |
0m
Model:
Tokens:
input: 53131
output: 41
thoughts: 1396
Thoughts: **The Perils of Concurrent List Iteration in Perf Events**
Okay, so here's the breakdown of what I'm seeing. It appears there's a serious concurrency bug lurking within the Perf event system, specifically related to how sibling lists are handled during throttling. `perf_swevent_get_recursion_context` disables preemption, but critically, it *doesn't* disable interrupts. Now, the `perf_event_throttle_group` function iterates the `sibling_list`, but it can be called with interrupts *enabled*. This is a significant problem.
My initial thought was that it was only tracepoint events, but no, software events are also vulnerable. `perf_sw_event` triggers `___perf_sw_event`, which disables preemption but not interrupts. So, we've got a window where `__perf_event_account_interrupt`, which is indirectly called through `__perf_event_overflow` *can* be entered with interrupts enabled.
The crucial piece is `throttle`. `throttle` becomes true after the first overflow in the `perf_swevent_overflow` function in a loop. Here's where the overflow logic kicks in. The important bit is when the overflow count is greater than one. The first call to `perf_swevent_overflow` always has throttle set to 0. It then calls `perf_swevent_set_period`, which could result in further overflows. The second overflow and subsequent overflows will have throttle set to 1. This means, if a lot of periods have passed, we can call `perf_event_throttle_group` with `hwc->interrupts` increasing every time `__perf_event_account_interrupt` is called, potentially reaching `max_samples_per_tick`.
So, what's exploitable here? If we can trigger `perf_event_throttle_group` with interrupts enabled, we can create a race condition. It looks like it is possible!
The critical code paths that modify `sibling_list` are within `perf_group_attach` and `perf_group_detach`, which are invoked when perf events are created or destroyed. These operations triggered by another CPU generate an IPI, which then calls either `__perf_install_in_context` or `__perf_remove_from_context`. These IPI handlers *hold `ctx->lock`*, and this is how they modify the `sibling_list`.
Here's the race: if our CPU is inside `perf_event_throttle_group` iterating `sibling_list`, and *an IPI arrives*, the IPI will interrupt us (because interrupts are enabled!). The IPI handler, after getting the lock will modify `sibling_list`. When the IPI returns, `perf_event_throttle_group` *continues iterating `sibling_list`*, but now the list has been modified *out from under it*.
This is a textbook list corruption/use-after-free vulnerability! `list_for_each_entry` is *not* safe for concurrent modification. If the current list entry gets unlinked and potentially freed while we're iterating, `sibling->sibling_list.next` (or previous) will point to freed memory or `LIST_POISON1`. The next iteration dereferences this pointer and we get a crash (page fault). Worst case, if memory is reallocated, it could lead to arbitrary code execution or memory corruption. Although the perf_event memory is freed via RCU, there are still timing windows where this can occur. It looks like both `list_del` and `list_del_init` are used.
This is a real problem. I'm going to check `perf_group_detach` next.
|
| 35/2 |
2026/05/18 09:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_group_detach]
Results: map[SourceCode:2251: static void perf_group_detach(struct perf_event *event)
2252: {
2253: struct perf_event *leader = event->group_leader;
2254: struct perf_event *sibling, *tmp;
2255: struct perf_event_context *ctx = event->ctx;
2256:
2257: lockdep_assert_held(&ctx->lock);
2258:
2259: /*
2260: * We can have double detach due to exit/hot-unplug + close.
2261: */
2262: if (!(event->attach_state & PERF_ATTACH_GROUP))
2263: return;
2264:
2265: event->attach_state &= ~PERF_ATTACH_GROUP;
2266:
2267: perf_put_aux_event(event);
2268:
2269: /*
2270: * If this is a sibling, remove it from its group.
2271: */
2272: if (leader != event) {
2273: list_del_init(&event->sibling_list);
2274: event->group_leader->nr_siblings--;
2275: event->group_leader->group_generation++;
2276: goto out;
2277: }
2278:
2279: /*
2280: * If this was a group event with sibling events then
2281: * upgrade the siblings to singleton events by adding them
2282: * to whatever list we are on.
2283: */
2284: list_for_each_entry_safe(sibling, tmp, &event->sibling_list, sibling_list) {
2285:
2286: /*
2287: * Events that have PERF_EV_CAP_SIBLING require being part of
2288: * a group and cannot exist on their own, schedule them out
2289: * and move them into the ERROR state. Also see
2290: * _perf_event_enable(), it will not be able to recover this
2291: * ERROR state.
2292: */
2293: if (sibling->event_caps & PERF_EV_CAP_SIBLING)
2294: __event_disable(sibling, ctx, PERF_EVENT_STATE_ERROR);
2295:
2296: sibling->group_leader = sibling;
2297: list_del_init(&sibling->sibling_list);
2298:
2299: /* Inherit group flags from the previous leader */
2300: sibling->group_caps = event->group_caps;
2301:
2302: if (sibling->attach_state & PERF_ATTACH_CONTEXT) {
2303: add_event_to_groups(sibling, event->ctx);
2304:
2305: if (sibling->state == PERF_EVENT_STATE_ACTIVE)
2306: list_add_tail(&sibling->active_list, get_event_list(sibling));
2307: }
2308:
2309: WARN_ON_ONCE(sibling->ctx != event->ctx);
2310: }
2311:
2312: out:
2313: for_each_sibling_event(tmp, leader)
2314: perf_event__header_size(tmp);
2315:
2316: perf_event__header_size(leader);
2317: }
SourceFile:kernel/events/core.c]
|
| 36/2 |
2026/05/18 09:38 |
llm |
expert |
0m
Model:
Tokens:
input: 56361
output: 41
thoughts: 388
Thoughts: **Analyzing a Kernel Infinite Loop Vulnerability**
Okay, so I'm looking at this code and it's using `list_del_init`. That immediately caught my eye. I know that means `sibling->sibling_list.next` is being set to itself. I see a potential problem now. If `perf_event_throttle_group` is iterating through `sibling_list` and, say, an Interruptible Processor Interrupt (IPI) comes along and detaches the current `sibling`, the list is reinitialized. The `sibling_list` of `sibling` is now pointing to itself again.
When the IPI returns, `perf_event_throttle_group` carries on. But now, it continues to the next element, which *is* `sibling` itself! That's a classic infinite loop scenario.
Wait, let me double-check how `list_for_each_entry` is implemented. Okay, yeah, here it is:
```c
#define list_for_each_entry(pos, head, member) \
for (pos = list_first_entry(head, typeof(*pos), member); \
!list_entry_is_head(pos, head, member); \
pos = list_next_entry(pos, member))
```
So, if `pos->member.next` points to `pos->member`, like we expect because of `list_del_init`, `list_next_entry(pos, member)` *will* return `pos`. Then `list_entry_is_head(pos, head, member)` will check if `&pos->member == head`. Since `pos` isn't the head (the head is `leader->sibling_list`), that condition is going to be false. We're trapped! Infinite loop confirmed!
Okay, so an infinite loop in the kernel with interrupts enabled? That's bad. Wait, an infinite loop *will* hang the CPU. That's a classic Denial of Service. Now, is there any way this can lead to privilege escalation? If it's *just* an infinite loop, then it's a DoS.
But wait a minute, I need to consider all cases. What happens if `perf_group_attach` is also called...?
|
| 37/2 |
2026/05/18 09:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_group_attach]
Results: map[SourceCode:2078: static void perf_group_attach(struct perf_event *event)
2079: {
2080: struct perf_event *group_leader = event->group_leader, *pos;
2081:
2082: lockdep_assert_held(&event->ctx->lock);
2083:
2084: /*
2085: * We can have double attach due to group movement (move_group) in
2086: * perf_event_open().
2087: */
2088: if (event->attach_state & PERF_ATTACH_GROUP)
2089: return;
2090:
2091: event->attach_state |= PERF_ATTACH_GROUP;
2092:
2093: if (group_leader == event)
2094: return;
2095:
2096: WARN_ON_ONCE(group_leader->ctx != event->ctx);
2097:
2098: group_leader->group_caps &= event->event_caps;
2099:
2100: list_add_tail(&event->sibling_list, &group_leader->sibling_list);
2101: group_leader->nr_siblings++;
2102: group_leader->group_generation++;
2103:
2104: perf_event__header_size(group_leader);
2105:
2106: for_each_sibling_event(pos, group_leader)
2107: perf_event__header_size(pos);
2108: }
SourceFile:kernel/events/core.c]
|
| 38/2 |
2026/05/18 09:39 |
llm |
expert |
0m
Model:
Tokens:
input: 57757
output: 42
thoughts: 1077
Thoughts: **Analyzing a Potential Race Condition in the Perf Event System**
Okay, let's break this down. My initial thought was: `list_add_tail` is used, so if we're iterating through a list, we might see a new element added, or we might miss it β it's not going to corrupt the iteration itself. But then the red flag goes up: what if the event is detached and freed while we're iterating? That memory goes back to the slab allocator, right? But wait, `perf_event` uses RCU for freeing. So, during that RCU grace period, the memory is *still* valid.
I need to trace the call stack: `perf_event_throttle_group` is called from `perf_tp_event`, which is called from `perf_trace_run_bpf_submit`, which is called from `trace_lock_acquire`, and then `lock_acquire`, and finally inside an `rcu_read_lock` critical section. Hold on, `rcu_read_lock` is *held*! This means we're in an RCU read-side critical section. Therefore, the `perf_event` cannot be freed while we're in `perf_event_throttle_group`.
So it's an infinite loop... which points to a Denial of Service.
I've got to double-check something. Can `perf_event_throttle_group` be called with interrupts enabled? Yes, we've established that `perf_tp_event` is called with them enabled. I've checked all the call paths, the `perf_swevent_event`, `perf_swevent_overflow`, `__perf_event_overflow`, `__perf_event_account_interrupt`, and `perf_event_throttle_group`, and they do not disable IRQs. So we're iterating `sibling_list` with IRQs enabled. Aha! `lockdep_assert_event_ctx` catches this because `hardirqs_enabled` is true and `ctx->mutex` isn't held.
Wait, are there other ways `perf_event_throttle_group` gets called? What about hardware events? Hardware events trigger NMIs or hardware IRQs. In those cases, hardirqs are disabled! `this_cpu_read(hardirqs_enabled)` will return false. And `lockdep_assert_event_ctx` will not warn. The warning is *specifically* due to software events or tracepoints overflowing with IRQs enabled.
If `perf_event_throttle_group` is called with IRQs enabled, is that a bug? Yes, because it iterates `sibling_list` without holding `ctx->lock` or `ctx->mutex` or disabling IRQs. Hmm, if we disabled IRQs, would that be safe? Disabling IRQs prevents IPIs on the *current* CPU. But what about another CPU modifying `sibling_list`? That `ctx->lock` is either per-CPU or per-task. Another CPU can't modify the memory without sending an IPI to this CPU if it's a CPU context, and can't modify a task-based context without sending an IPI to the task's CPU. So, disabling IRQs *would* prevent concurrent modifications.
The problem? IRQs are *NOT* disabled. An IPI can interrupt us and modify `sibling_list`. This is a real bug!
Okay, can we trigger this warning from unprivileged user space? The syzkaller reproducer uses `perf_event_open` to create a software event. Looking at the `perf_event_attr`, we see: `type = 2` (PERF_TYPE_TRACEPOINT). Can an unprivileged user *create* a tracepoint event? That's the next check.
|
| 39/2 |
2026/05/18 09:39 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_allow_tracepoint]
Results: map[SourceCode:1787: static inline int perf_allow_tracepoint(void)
1788: {
1789: if (sysctl_perf_event_paranoid > -1 && !perfmon_capable())
1790: return -EPERM;
1791:
1792: return security_perf_event_open(PERF_SECURITY_TRACEPOINT);
1793: }
SourceFile:include/linux/perf_event.h]
|
| 40/2 |
2026/05/18 09:39 |
llm |
expert |
0m
Model:
Tokens:
input: 59835
output: 41
thoughts: 192
Thoughts: **Analyzing Tracepoint and Software Events with Default Sysctl Values**
Okay, let's break this down. My initial thought is that `sysctl_perf_event_paranoid` defaults to 2, which is critical. This means that to use `perf_allow_tracepoint`, I need `perfmon_capable()`. And that, in turn, hinges on either `CAP_PERFMON` or `CAP_SYS_ADMIN`. Since the prompt explicitly states a non-root user *without* special capabilities, the user definitely won't have either of those, correct? Therefore, given `sysctl_perf_event_paranoid` at its default value of 2, a regular user **cannot** create tracepoint events.
But, hang on, there's `PERF_TYPE_SOFTWARE`, which is type 1. Does that open a different avenue? Can an unprivileged user create software events? I need to check how `perf_event_open` handles `PERF_TYPE_SOFTWARE` to be sure. This feels like the key piece to the puzzle, so I'll need to dig into that specific code path.
|
| 41/2 |
2026/05/18 09:39 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_allow_cpu]
Results: map[SourceCode:1779: static inline int perf_allow_cpu(void)
1780: {
1781: if (sysctl_perf_event_paranoid > 0 && !perfmon_capable())
1782: return -EACCES;
1783:
1784: return security_perf_event_open(PERF_SECURITY_CPU);
1785: }
SourceFile:include/linux/perf_event.h]
|
| 42/2 |
2026/05/18 09:39 |
llm |
expert |
0m
Model:
Tokens:
input: 60449
output: 43
thoughts: 41
Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience:
**Investigating `perf_allow_cpu` and Event Types**
Hmm, a quick clarification is needed. `perf_allow_cpu` is, of course, specifically designed for managing CPU-related performance events, isn't it? My focus shifted from hardware and the kernel, to more general considerations of software.
Right, so that leads me to question: what about events related to `PERF_TYPE_SOFTWARE`? I seem to recall different software event types (such as `PERF_COUNT_SW_CPU_CLOCK` etc), that I may be able to use in conjunction with this. I'd like to get the data for software events, and ensure it functions correctly.
Okay, time to dive deeper. I'm going to consult the implementation details of the `perf_event_open` system call. The precise mechanics of how the kernel handles different event types, how it sets up the necessary data structures, and how it performs filtering will be critical here. I need to be precise on how to utilize `perf_event_open`, and how it allows or restricts the allocation of software-based events.
|
| 43/2 |
2026/05/18 09:39 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:sys_perf_event_open]
Results: map[SourceCode: SourceFile:]
|
| 44/2 |
2026/05/18 09:39 |
llm |
expert |
0m
Model:
Tokens:
input: 60808
output: 26
thoughts: 20
Thoughts: Okay, here's my interpretation of that:
**Diving into `perf_event_open`**
Alright, let's break this down. Seeing `SYSCALL_DEFINE5(perf_event_open, ...)` immediately tells me we're looking at the core definition of the `perf_event_open` system call within the kernel. It's the point of entry, the fundamental gateway for userspace applications to interact with the performance monitoring infrastructure. The `SYSCALL_DEFINE5` macro is crucial, letting us know this syscall takes five arguments. I need to remember the arguments it takes and their order; the specific signature is critical to how it interfaces with the performance counter hardware and associated kernel subsystems. This also signals that the next thing to examine is the implementation of that system call to see how it operates, and what data it utilizes or requires.
|
| 45/2 |
2026/05/18 09:39 |
tool |
grepper |
0m
Args: map[Expression:SYSCALL_DEFINE5\(perf_event_open]
Results: map[Output:kernel/events/core.c=13348=perf_check_permission(struct perf_event_attr *attr, struct task_struct *task)
--
kernel/events/core.c-13386- */
kernel/events/core.c:13387:SYSCALL_DEFINE5(perf_event_open,
kernel/events/core.c-13388- struct perf_event_attr __user *, attr_uptr,
]
|
| 46/2 |
2026/05/18 09:39 |
llm |
expert |
0m
Model:
Tokens:
input: 61161
output: 41
thoughts: 11
Thoughts: Okay, here's my attempt at summarizing that thought process, tailored for an expert audience:
**Analyzing `perf_check_permission`**
Alright, let's dive right in and check out the `perf_check_permission` function. My initial thought process focuses on understanding its role within the broader performance monitoring context. I need to quickly recall the typical permissioning models involved in this kind of system, considering the likely use cases for performance analysis. Are we talking about user-level permissions, group permissions, or something more granular based on specific events or counters? The function's name suggests it's the gatekeeper, so I'm immediately probing for how finely tuned this check is, and what resources it's protecting. Is it a central check, or are there multiple layers involved? How does it handle potential race conditions or concurrent access scenarios? And of course, the big questions β what are the potential performance bottlenecks introduced by this check? How often is it called, and what's the overhead involved in executing it? I need to analyze its code in detail, paying close attention to any external dependencies and potential impact on the overall performance monitoring infrastructure. I'm also considering how it interacts with the underlying kernel interfaces and other performance related tools. I need to be prepared to identify areas for optimization, whether that's caching results, improving the algorithm, or reducing the frequency of the check itself.
|
| 47/2 |
2026/05/18 09:39 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_check_permission]
Results: map[SourceCode:13347: static bool
13348: perf_check_permission(struct perf_event_attr *attr, struct task_struct *task)
13349: {
13350: unsigned int ptrace_mode = PTRACE_MODE_READ_REALCREDS;
13351: bool is_capable = perfmon_capable();
13352:
13353: if (attr->sigtrap) {
13354: /*
13355: * perf_event_attr::sigtrap sends signals to the other task.
13356: * Require the current task to also have CAP_KILL.
13357: */
13358: rcu_read_lock();
13359: is_capable &= ns_capable(__task_cred(task)->user_ns, CAP_KILL);
13360: rcu_read_unlock();
13361:
13362: /*
13363: * If the required capabilities aren't available, checks for
13364: * ptrace permissions: upgrade to ATTACH, since sending signals
13365: * can effectively change the target task.
13366: */
13367: ptrace_mode = PTRACE_MODE_ATTACH_REALCREDS;
13368: }
13369:
13370: /*
13371: * Preserve ptrace permission check for backwards compatibility. The
13372: * ptrace check also includes checks that the current task and other
13373: * task have matching uids, and is therefore not done here explicitly.
13374: */
13375: return is_capable || ptrace_may_access(task, ptrace_mode);
13376: }
SourceFile:kernel/events/core.c]
|
| 48/2 |
2026/05/18 09:39 |
llm |
expert |
0m
Model:
Tokens:
input: 62016
output: 23
thoughts: 55
Thoughts: Okay, let me think this through.
**Permission Check Deep Dive**
Hmm, `perf_check_permission`... right, it's just verifying attachment to the `task`. That's a key first step, but it doesn't give the whole picture, does it? We're missing something crucial. What about the system-wide paranoia level set by `sysctl_perf_event_paranoid`? That's definitely a factor. We can't forget that!
Alright, let's trace the execution path. Let's dig into the `perf_event_open` function, starting around line 13387. I'll need to refresh my memory on how these checks and restrictions interact, and exactly where `sysctl_perf_event_paranoid` is evaluated. I'm going to have to see how it's being used.
|
| 49/2 |
2026/05/18 09:39 |
tool |
grepper |
0m
Args: map[Expression:sysctl_perf_event_paranoid]
Results: map[Output:include/linux/perf_event.h=1746=static inline int perf_callchain_store(struct perf_callchain_entry_ctx *ctx, u64 ip)
--
include/linux/perf_event.h-1758-
include/linux/perf_event.h:1759:extern int sysctl_perf_event_paranoid;
include/linux/perf_event.h-1760-extern int sysctl_perf_event_sample_rate;
--
include/linux/perf_event.h=1772=static inline int perf_is_paranoid(void)
include/linux/perf_event.h-1773-{
include/linux/perf_event.h:1774: return sysctl_perf_event_paranoid > -1;
include/linux/perf_event.h-1775-}
--
include/linux/perf_event.h=1779=static inline int perf_allow_cpu(void)
include/linux/perf_event.h-1780-{
include/linux/perf_event.h:1781: if (sysctl_perf_event_paranoid > 0 && !perfmon_capable())
include/linux/perf_event.h-1782- return -EACCES;
--
include/linux/perf_event.h=1787=static inline int perf_allow_tracepoint(void)
include/linux/perf_event.h-1788-{
include/linux/perf_event.h:1789: if (sysctl_perf_event_paranoid > -1 && !perfmon_capable())
include/linux/perf_event.h-1790- return -EPERM;
--
kernel/events/core.c=458=static struct kmem_cache *perf_event_cache;
--
kernel/events/core.c-466- */
kernel/events/core.c:467:int sysctl_perf_event_paranoid __read_mostly = 2;
kernel/events/core.c-468-
--
kernel/events/core.c=544=static const struct ctl_table events_core_sysctl_table[] = {
--
kernel/events/core.c-550- .procname = "perf_event_paranoid",
kernel/events/core.c:551: .data = &sysctl_perf_event_paranoid,
kernel/events/core.c:552: .maxlen = sizeof(sysctl_perf_event_paranoid),
kernel/events/core.c-553- .mode = 0644,
--
kernel/events/core.c=14261=int perf_allow_kernel(void)
kernel/events/core.c-14262-{
kernel/events/core.c:14263: if (sysctl_perf_event_paranoid > 1 && !perfmon_capable())
kernel/events/core.c-14264- return -EACCES;
--
kernel/ksyms_common.c=9=static inline int kallsyms_for_perf(void)
--
kernel/ksyms_common.c-11-#ifdef CONFIG_PERF_EVENTS
kernel/ksyms_common.c:12: extern int sysctl_perf_event_paranoid;
kernel/ksyms_common.c-13-
kernel/ksyms_common.c:14: if (sysctl_perf_event_paranoid <= 1)
kernel/ksyms_common.c-15- return 1;
--
kernel/ksyms_common.c-22- * kernel profiling and are explicitly not paranoid (so kptr_restrict
kernel/ksyms_common.c:23: * is clear, and sysctl_perf_event_paranoid isn't set).
kernel/ksyms_common.c-24- *
]
|
| 50/2 |
2026/05/18 09:39 |
llm |
expert |
0m
Model:
Tokens:
input: 63271
output: 19
thoughts: 82
Thoughts: **Analyzing `perf_allow_kernel` in the Context of `perf_event_open`**
Okay, so I'm focusing on the interaction of `perf_allow_kernel` with the broader performance monitoring architecture. It looks like `perf_allow_kernel` hinges on a `sysctl` value, specifically `sysctl_perf_event_paranoid`. Right, so if `sysctl_perf_event_paranoid` is set to 2, then `perf_allow_kernel` is going to return an `-EACCES` error, meaning access is denied to unprivileged users, if they attempt something that requires kernel-level access through the performance monitoring infrastructure. That makes perfect sense for system security and control, but I need to understand its exact application.
The critical question now is, where does `perf_allow_kernel` come into play within the `perf_event_open` system call? I need to trace its usage within that function to fully grasp how security permissions are enforced and how it affects the ability of unprivileged users to create and manage performance monitoring events, particularly those involving kernel data. I need to see the conditional branches within `perf_event_open` that call this function.
|
| 51/2 |
2026/05/18 09:39 |
tool |
grepper |
0m
Args: map[Expression:perf_allow_kernel]
Results: map[Output:arch/x86/events/amd/ibs.c=1226=static int perf_ibs_handle_irq(struct perf_ibs *perf_ibs, struct pt_regs *iregs)
--
arch/x86/events/amd/ibs.c-1346- if ((event->attr.sample_type & PERF_SAMPLE_RAW) &&
arch/x86/events/amd/ibs.c:1347: perf_allow_kernel()) {
arch/x86/events/amd/ibs.c-1348- perf_ibs_phyaddr_clear(perf_ibs, &ibs_data);
--
arch/x86/events/intel/bts.c=564=static int bts_event_init(struct perf_event *event)
--
arch/x86/events/intel/bts.c-577- if (event->attr.exclude_kernel) {
arch/x86/events/intel/bts.c:578: ret = perf_allow_kernel();
arch/x86/events/intel/bts.c-579- if (ret)
--
drivers/perf/arm_spe_pmu.c-43- * Cache if the event is allowed to trace Context information.
drivers/perf/arm_spe_pmu.c:44: * This allows us to perform the check, i.e, perf_allow_kernel(),
drivers/perf/arm_spe_pmu.c-45- * in the context of the event owner, once, during the event_init().
--
drivers/perf/arm_spe_pmu.c=51=static void set_spe_event_has_cx(struct perf_event *event)
drivers/perf/arm_spe_pmu.c-52-{
drivers/perf/arm_spe_pmu.c:53: if (IS_ENABLED(CONFIG_PID_IN_CONTEXTIDR) && !perf_allow_kernel())
drivers/perf/arm_spe_pmu.c-54- event->hw.flags |= SPE_PMU_HW_FLAGS_CX;
--
drivers/perf/arm_spe_pmu.c=774=static int arm_spe_pmu_event_init(struct perf_event *event)
--
drivers/perf/arm_spe_pmu.c-840- if (reg & (PMSCR_EL1_PA | PMSCR_EL1_PCT))
drivers/perf/arm_spe_pmu.c:841: return perf_allow_kernel();
drivers/perf/arm_spe_pmu.c-842-
--
include/linux/perf_event.h=1772=static inline int perf_is_paranoid(void)
--
include/linux/perf_event.h-1776-
include/linux/perf_event.h:1777:extern int perf_allow_kernel(void);
include/linux/perf_event.h-1778-
--
kernel/events/core.c=13087=static int perf_copy_attr(struct perf_event_attr __user *uattr,
--
kernel/events/core.c-13153- if (mask & PERF_SAMPLE_BRANCH_PERM_PLM) {
kernel/events/core.c:13154: ret = perf_allow_kernel();
kernel/events/core.c-13155- if (ret)
--
kernel/events/core.c=13387=SYSCALL_DEFINE5(perf_event_open,
--
kernel/events/core.c-13418- if (!attr.exclude_kernel) {
kernel/events/core.c:13419: err = perf_allow_kernel();
kernel/events/core.c-13420- if (err)
--
kernel/events/core.c-13438- if ((attr.sample_type & PERF_SAMPLE_PHYS_ADDR)) {
kernel/events/core.c:13439: err = perf_allow_kernel();
kernel/events/core.c-13440- if (err)
--
kernel/events/core.c=14253=const struct perf_event_attr *perf_event_attrs(struct perf_event *event)
--
kernel/events/core.c-14260-
kernel/events/core.c:14261:int perf_allow_kernel(void)
kernel/events/core.c-14262-{
--
kernel/events/core.c-14267-}
kernel/events/core.c:14268:EXPORT_SYMBOL_GPL(perf_allow_kernel);
kernel/events/core.c-14269-
]
|
| 52/2 |
2026/05/18 09:39 |
llm |
expert |
0m
Model:
Tokens:
input: 64820
output: 19
thoughts: 84
Thoughts: **Diving into the Perf Event Permissions**
Okay, so I'm looking at how `perf_event_open` handles kernel access restrictions. It's pretty straightforward: if the `attr.exclude_kernel` flag is *not* set, the code hits `perf_allow_kernel()`. Makes perfect sense. But if `attr.exclude_kernel` *is* set, then the check is bypassed, right? That's the basic logic, I get it.
But it just raised another question in my mind...what's the deal with `perf_allow_cpu()` and `perf_allow_tracepoint()`? Are they handled the same way? Are there similar conditional checks that determine their usage or not based on some other attribute flags? I should go digging in the code and see where those functions get called. I really need to understand the full picture of the permission model to avoid any unexpected behavior. I need to trace the call paths and see exactly how they're used.
|
| 53/2 |
2026/05/18 09:39 |
tool |
grepper |
0m
Args: map[Expression:perf_allow_cpu]
Results: map[Output:arch/x86/events/intel/core.c=4222=static int intel_pmu_hw_config(struct perf_event *event)
--
arch/x86/events/intel/core.c-4536-
arch/x86/events/intel/core.c:4537: ret = perf_allow_cpu();
arch/x86/events/intel/core.c-4538- if (ret)
--
arch/x86/events/intel/p4.c=751=static int p4_validate_raw_event(struct perf_event *event)
--
arch/x86/events/intel/p4.c-780- if (p4_ht_active() && p4_event_bind_map[v].shared) {
arch/x86/events/intel/p4.c:781: v = perf_allow_cpu();
arch/x86/events/intel/p4.c-782- if (v)
--
include/linux/perf_event.h=1777=extern int perf_allow_kernel(void);
include/linux/perf_event.h-1778-
include/linux/perf_event.h:1779:static inline int perf_allow_cpu(void)
include/linux/perf_event.h-1780-{
--
kernel/events/core.c=4956=find_get_context(struct task_struct *task, struct perf_event *event)
--
kernel/events/core.c-4964- /* Must be root to operate on a CPU event: */
kernel/events/core.c:4965: err = perf_allow_cpu();
kernel/events/core.c-4966- if (err)
]
|
| 54/2 |
2026/05/18 09:39 |
llm |
expert |
0m
Model:
Tokens:
input: 65594
output: 20
thoughts: 58
Thoughts: Okay, here's my take on those thoughts, framed as my own reasoning:
**Analyzing Perf Event Security**
Alright, let's break this down. My initial observation is that `perf_allow_cpu` is being checked within `find_get_context`. The specific case that jumps out is when the `task` pointer is NULL, which signifies a CPU-level event. This has some interesting security implications. Essentially, the code is allowing unprivileged users to generate task events, but only those tied to the CPU. The question that immediately arises is... what about `perf_allow_tracepoint`? Does it have similar limitations, or can a user potentially exploit a similar avenue to trigger a privileged event in a less controlled manner? I need to analyze its implementation in context to understand the potential attack surface. It looks like more investigation is warranted here.
|
| 55/2 |
2026/05/18 09:39 |
tool |
grepper |
0m
Args: map[Expression:perf_allow_tracepoint]
Results: map[Output:include/linux/perf_event.h=1779=static inline int perf_allow_cpu(void)
--
include/linux/perf_event.h-1786-
include/linux/perf_event.h:1787:static inline int perf_allow_tracepoint(void)
include/linux/perf_event.h-1788-{
--
kernel/trace/trace_event_perf.c=27=static int perf_trace_event_perm(struct trace_event_call *tp_event,
--
kernel/trace/trace_event_perf.c-51- if (ftrace_event_is_function(tp_event)) {
kernel/trace/trace_event_perf.c:52: ret = perf_allow_tracepoint();
kernel/trace/trace_event_perf.c-53- if (ret)
--
kernel/trace/trace_event_perf.c-88- */
kernel/trace/trace_event_perf.c:89: ret = perf_allow_tracepoint();
kernel/trace/trace_event_perf.c-90- if (ret)
]
|
| 56/2 |
2026/05/18 09:39 |
llm |
expert |
0m
Model:
Tokens:
input: 66140
output: 42
thoughts: 287
Thoughts: **My Analysis of Tracepoint and Software Events**
Okay, so I'm looking at how tracing works within the kernel, specifically around user permissions. It seems `perf_allow_tracepoint` is a key check within `perf_trace_event_perm`. That function, in turn, is called during the creation of a tracepoint event. This means, by default, with `sysctl_perf_event_paranoid` at its default value of 2, unprivileged users are blocked from creating these tracepoint events. That's straightforward.
But...what about software events? Can those be created by unprivileged users? I need to clarify that. Ah, yes, `PERF_TYPE_SOFTWARE` events *are* allowed, particularly for task events.
Now, here's the critical question: Can these software events, even the ones unprivileged users can create, end up triggering the `perf_event_throttle_group` mechanism when interrupts are enabled? I need to understand the execution path here. Let's dig into an example like `PERF_COUNT_SW_PAGE_FAULTS`.
Okay, page faults. Page faults happen within the context of the instruction that caused them. If the fault itself occurs with interrupts enabled, then the page fault handler *itself* runs with interrupts enabled. Wait... is that right? Let me confirm. Yes, that's correct, at least on x86, page faults are handled with interrupts enabled after some initial setup.
Right, so where does that actually trigger the perf infrastructure? Let's check `do_user_addr_fault`. Yes, there it is! `do_user_addr_fault` calls `perf_sw_event(PERF_COUNT_SW_PAGE_FAULTS, 1, regs, address)`. And that call then leads to `__perf_sw_event`, which is where the core perf infrastructure interaction happens. That's my key focus.
|
| 57/2 |
2026/05/18 09:39 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:__perf_sw_event]
Results: map[SourceCode:10637: void __perf_sw_event(u32 event_id, u64 nr, struct pt_regs *regs, u64 addr)
10638: {
10639: int rctx;
10640:
10641: preempt_disable_notrace();
10642: rctx = perf_swevent_get_recursion_context();
10643: if (unlikely(rctx < 0))
10644: goto fail;
10645:
10646: ___perf_sw_event(event_id, nr, regs, addr);
10647:
10648: perf_swevent_put_recursion_context(rctx);
10649: fail:
10650: preempt_enable_notrace();
10651: }
SourceFile:kernel/events/core.c]
|
| 58/2 |
2026/05/18 09:39 |
llm |
expert |
0m
Model:
Tokens:
input: 67143
output: 42
thoughts: 9
Thoughts: Okay, here's my attempt at summarizing the thought process, tailored for an expert audience, and written in the first person:
**Analyzing `___perf_sw_event` and Its Implications**
Alright, let's dissect this. When I see `___perf_sw_event`, my mind immediately jumps to the performance monitoring realm, specifically the user-space software event tracing mechanisms. My initial thought is to quickly ascertain what triggers this function, and if it's related to any specific, defined perf events. Is this a custom event, or one of the standard ones? What parameters does it accept, and how is it related to the perf data buffers? I need to understand the data's format.
My next consideration pivots to potential performance implications. Is this a critical path function? What are the overheads associated with its execution? Are there any potential bottlenecks introduced by these event hooks? I'll need to think about lock contention if any of these events are occurring within a heavily contended section. I would immediately start thinking about how to estimate the frequency of the event's use and determine its runtime cost. What are the calling sequences that end in this event? I'll need to consider how to minimize the impact, especially if this event is being triggered frequently, or from a high-frequency interrupt context.
I'm thinking about the bigger picture. Is this part of a profiling setup, or is it intended for general runtime performance monitoring? Are these events being collected for user-space analysis, kernel-space analysis, or both? I need to look into how the data is being consumed and if this impacts the system's performance.
Moreover, are there any interactions with other performance monitoring subsystems? Does it leverage or interfere with other tracing tools such as `ftrace` or `bcc`? Iβll need to explore that for potential interferences. Is this a system-wide instrumentation, or targeted at a specific application or subsystem? The scope is critical.
Finally, I'll need to quickly consider the security implications. Are there any potential information leaks through these events, and any chances that these events might be triggered maliciously? Access controls will need to be checked. Overall, my mind is racing, trying to form a mental model of how `___perf_sw_event` fits into the broader performance landscape of this system. I want to build a test environment to run my hypotheses!
|
| 59/2 |
2026/05/18 09:39 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:___perf_sw_event]
Results: map[SourceCode:10626: void ___perf_sw_event(u32 event_id, u64 nr, struct pt_regs *regs, u64 addr)
10627: {
10628: struct perf_sample_data data;
10629:
10630: if (WARN_ON_ONCE(!regs))
10631: return;
10632:
10633: perf_sample_data_init(&data, addr, 0);
10634: do_perf_sw_event(PERF_TYPE_SOFTWARE, event_id, nr, &data, regs);
10635: }
SourceFile:kernel/events/core.c]
|
| 60/2 |
2026/05/18 09:39 |
llm |
expert |
0m
Model:
Tokens:
input: 67895
output: 43
thoughts: 10
Thoughts: Okay, here's my attempt at summarizing the thought process, tailored for an expert audience, and written in the first person:
**Analyzing the `do_perf_sw_event` Function**
Alright, let's break down this `do_perf_sw_event` function. My initial thought process revolves around understanding its context within the broader performance monitoring infrastructure. I'm immediately looking for its role in handling software performance events. What specific events is it designed to manage? Is it a central dispatch point, or more specialized?
I'm mentally mapping the function's inputs. What data structures are passed in? `perf_event` is a given, obviously, but what crucial metadata is included? I'm anticipating that context is key: CPU context, interrupt state, perhaps a timestamp, and any relevant data from the associated hardware counters, if applicable, since this is software, there may be no relevant hardware counters.
My next step is to quickly identify the core logic. I want to pinpoint the crucial parts that differentiate this from hardware event handling. What's the control flow? What conditional checks are involved? Are there any obvious race conditions that need careful consideration, since the function is probably handling interrupts. How are events processed β is there a direct callback mechanism, or some sort of queuing or processing stage? What data is being updated or tracked in this function?
I'm drilling down further, I need to know how these events are actually *recorded* or *reported*. Is there a write operation to shared memory buffers for user-space access? Are there statistical aggregations taking place? Is there potentially thread synchronization happening here? If it's related to any specific software instrumentation, I need to understand how the data relates to the program's code, or the kernel code where this function is called.
Finally, I'm thinking about potential performance implications. Is this a hot path, or relatively infrequent? Are there any potential bottlenecks introduced by this code? Are there any opportunities for optimization, such as using atomic operations, lock-free data structures, or batching event processing? I'm already formulating questions in my mind to investigate. I'm going to want to check the architectureβs performance, especially considering where the function is called.
|
| 61/2 |
2026/05/18 09:39 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:do_perf_sw_event]
Results: map[SourceCode:10591: static void do_perf_sw_event(enum perf_type_id type, u32 event_id,
10592: u64 nr,
10593: struct perf_sample_data *data,
10594: struct pt_regs *regs)
10595: {
10596: struct swevent_htable *swhash = this_cpu_ptr(&swevent_htable);
10597: struct perf_event *event;
10598: struct hlist_head *head;
10599:
10600: rcu_read_lock();
10601: head = find_swevent_head_rcu(swhash, type, event_id);
10602: if (!head)
10603: goto end;
10604:
10605: hlist_for_each_entry_rcu(event, head, hlist_entry) {
10606: if (perf_swevent_match(event, type, event_id, data, regs))
10607: perf_swevent_event(event, nr, data, regs);
10608: }
10609: end:
10610: rcu_read_unlock();
10611: }
SourceFile:kernel/events/core.c]
|
| 62/2 |
2026/05/18 09:39 |
llm |
expert |
0m
Model:
Tokens:
input: 68804
output: 40
thoughts: 289
Thoughts: **Analyzing the Software Event Flow**
Okay, so I'm tracing the path of a performance software event. Right, `do_perf_sw_event` is the starting point, iterating through the `head` with RCU, and then it calls `perf_swevent_event`. `perf_swevent_event` then cascades into `perf_swevent_overflow`, leading to `__perf_event_overflow`. That in turn calls `__perf_event_account_interrupt`, finally ending up at `perf_event_throttle_group`.
Now, the crucial point: `__perf_sw_event` only disables preemption, *not* interrupts. This means, if a page fault occurs *while interrupts are enabled*, `perf_event_throttle_group` will be invoked with interrupts still enabled. And here's the kicker: an unprivileged user can indeed create a software event specifically for page faults (`PERF_COUNT_SW_PAGE_FAULTS`). They can trigger these page faults quite easily by playing with unmapped memory using read/write or `mmap`.
Hold on... the syzkaller reproducer I'm looking at, uses `type: perf_event_type = 0x2` which translates to `PERF_TYPE_TRACEPOINT`. That's unexpected. How did syzkaller, which I thought runs as an unprivileged process, manage to create a tracepoint event? Then I recalled that syzkaller operates within a user namespace, and it's granted `CAP_SYS_ADMIN` inside that namespace. I have to verify if `perf_allow_tracepoint` permits the creation of tracepoints within a user namespace. Let me check the `perfmon_capable()` function to determine the exact conditions.
|
| 63/2 |
2026/05/18 09:39 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:include/linux/security.h IncludeLines:true Name:perfmon_capable]
Results: map[SourceCode: 190: static inline bool perfmon_capable(void)
191: {
192: return capable(CAP_PERFMON) || capable(CAP_SYS_ADMIN);
193: }
SourceFile:include/linux/capability.h]
|
| 64/2 |
2026/05/18 09:39 |
llm |
expert |
0m
Model:
Tokens:
input: 69595
output: 41
thoughts: 2320
Thoughts: **My Analysis of the Perf Event Throttle Vulnerability**
Okay, here's what I've got. It seems like the core issue revolves around how `perf_event_throttle_group` interacts with concurrent event removal within the kernel, leading to a denial-of-service condition.
First, I see that I don't have the `CAP_PERFMON` capability, so creating events is tricky. But wait, I recall that I can create software events. The key is in how software events are handled within the `perf_event_open` syscall. There is no access check for `PERF_TYPE_SOFTWARE`, so I can create those unprivileged. That's my entry point.
Next, I can set up a group of software events. I will make a leader event (e.g., page faults) and a sibling event. The plan is to trigger the leader event to overflow and then to have its overflow trigger `perf_event_throttle_group` which iterates through the sibling list.
Here's where it gets interesting: I'll use a page fault to trigger the loop. Concurrently, from another thread, I'll close the sibling's file descriptor. This causes the sibling to be removed from the list while `perf_event_throttle_group` is iterating through it. This will call `perf_group_detach`, which is where the problem is.
The race condition causes `list_del_init` to be called on a sibling in the middle of iteration within `perf_event_throttle_group`. This modification of the linked list causes `perf_event_throttle_group` to get stuck in an infinite loop due to a modified `next` pointer.
The infinite loop is in the kernel, and it can hang an entire CPU. Although it's an infinite loop, and I'm iterating `sibling_list`, it is in `perf_event_throttle_group`, which has `perf_event_throttle`. This calls `tick_dep_set_cpu`, but it won't actually be an infinite loop, as `perf_event_throttle` has a check for `MAX_INTERRUPTS`. But the CPU will be stuck in a tight loop.
However, I'm stuck with the impact being just a Denial of Service, not privilege escalation, since the memory is protected by RCU, so I can't trigger a use-after-free.
I need to confirm that `rcu_read_lock` is held, in `do_perf_sw_event`, but I see that's the case. Also, in tracepoints, the `hlist_for_each_entry_rcu` is NOT protected by a lock, so that is an area of concern to note.
So, in summary, I can create a group of software events, trigger the leader, and concurrently close a sibling to cause an infinite loop in the kernel. This hangs the CPU and creates a Denial of Service. It's not exploitable for privilege escalation, just DoS.
|
| 65/2 |
2026/05/18 09:39 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_tp_event]
Results: map[SourceCode:10984: void perf_tp_event(u16 event_type, u64 count, void *record, int entry_size,
10985: struct pt_regs *regs, struct hlist_head *head, int rctx,
10986: struct task_struct *task)
10987: {
10988: struct perf_sample_data data;
10989: struct perf_event *event;
10990:
10991: struct perf_raw_record raw = {
10992: .frag = {
10993: .size = entry_size,
10994: .data = record,
10995: },
10996: };
10997:
10998: perf_trace_buf_update(record, event_type);
10999:
11000: hlist_for_each_entry_rcu(event, head, hlist_entry) {
11001: if (perf_tp_event_match(event, &raw, regs)) {
11002: /*
11003: * Here use the same on-stack perf_sample_data,
11004: * some members in data are event-specific and
11005: * need to be re-computed for different sweveents.
11006: * Re-initialize data->sample_flags safely to avoid
11007: * the problem that next event skips preparing data
11008: * because data->sample_flags is set.
11009: */
11010: perf_sample_data_init(&data, 0, 0);
11011: perf_sample_save_raw_data(&data, event, &raw);
11012: perf_swevent_event(event, count, &data, regs);
11013: }
11014: }
11015:
11016: /*
11017: * If we got specified a target task, also iterate its context and
11018: * deliver this event there too.
11019: */
11020: if (task && task != current) {
11021: struct perf_event_context *ctx;
11022:
11023: rcu_read_lock();
11024: ctx = rcu_dereference(task->perf_event_ctxp);
11025: if (!ctx)
11026: goto unlock;
11027:
11028: raw_spin_lock(&ctx->lock);
11029: perf_tp_event_target_task(count, record, regs, &data, &raw, ctx);
11030: raw_spin_unlock(&ctx->lock);
11031: unlock:
11032: rcu_read_unlock();
11033: }
11034:
11035: perf_swevent_put_recursion_context(rctx);
11036: }
SourceFile:kernel/events/core.c]
|
| 66/2 |
2026/05/18 09:39 |
llm |
expert |
6m
Model:
Tokens:
input: 73466
output: 18
thoughts: 91
Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience:
**Analyzing the RCU Usage in `perf_tp_event`**
Hmm, so `perf_tp_event` is leveraging `hlist_for_each_entry_rcu`. But hold on... I'm not seeing an explicit `rcu_read_lock()` call around it. That immediately raises a flag. Why is this designed in this way? Ah, right. It hits me β because tracepoints, by their very nature, are invoked with an RCU read lock *already* established. That's the key. Tracepoints utilize either `rcu_read_lock_sched()` or the standard `rcu_read_lock()`. Therefore, the RCU protection is implicitly present at the point of the iteration. Just to be absolutely sure though, let's just confirm how this is set up by checking the `__DO_TRACE` macros again. It is always good to refresh this sort of memory.
|
| 67/2 |
2026/05/18 09:46 |
tool |
grepper |
0m
Args: map[Expression:__DO_TRACE]
Results: map[Output:include/linux/tracepoint.h=157=static inline struct tracepoint *tracepoint_ptr_deref(tracepoint_ptr_t *p)
--
include/linux/tracepoint.h-196-#ifdef CONFIG_HAVE_STATIC_CALL
include/linux/tracepoint.h:197:#define __DO_TRACE_CALL(name, args) \
include/linux/tracepoint.h-198- do { \
--
include/linux/tracepoint.h-208-#else
include/linux/tracepoint.h:209:#define __DO_TRACE_CALL(name, args) __traceiter_##name(NULL, args)
include/linux/tracepoint.h-210-#endif /* CONFIG_HAVE_STATIC_CALL */
--
include/linux/tracepoint.h-274- guard(preempt_notrace)(); \
include/linux/tracepoint.h:275: __DO_TRACE_CALL(name, TP_ARGS(args)); \
include/linux/tracepoint.h-276- } \
--
include/linux/tracepoint.h-292- guard(rcu_tasks_trace)(); \
include/linux/tracepoint.h:293: __DO_TRACE_CALL(name, TP_ARGS(args)); \
include/linux/tracepoint.h-294- } \
--
kernel/tracepoint.c=277=static int tracepoint_add_func(struct tracepoint *tp,
--
kernel/tracepoint.c-300- * that the new probe callbacks array is consistent before setting
kernel/tracepoint.c:301: * a pointer to it. This array is referenced by __DO_TRACE from
kernel/tracepoint.c-302- * include/linux/tracepoint.h using rcu_dereference_sched().
]
|
| 68/2 |
2026/05/18 09:46 |
llm |
expert |
0m
Model:
Tokens:
input: 74264
output: 41
thoughts: 1526
Thoughts: **Kernel Perf Event Infinite Loop Vulnerability**
Okay, here's what I've got. Tracepoints use `rcu_read_lock_sched_notrace()` or a similar mechanism, or `preempt_disable_notrace()`, which effectively acts as an RCU read-side critical section. And `perf_event` is freed with `call_rcu()`. But hold on, `call_rcu()` waits for `rcu_read_lock()` and not `rcu_read_lock_sched()`! In modern kernels, however, `call_rcu()` and `call_rcu_sched()` are consolidated. This means that `preempt_disable()` is sufficient to block `call_rcu()` from running. So, the memory is *not* being freed, leading to an infinite loop.
This is a kernel infinite loop, a denial-of-service (DoS) vulnerability. Can this be triggered by unprivileged processes? Yes, it can. Unprivileged processes can create software events like `PERF_COUNT_SW_PAGE_FAULTS` and then trigger them. Does that software event, specifically, trigger `perf_event_throttle_group`? I think it does... Let's see... `__perf_event_account_interrupt` is called from `__perf_event_overflow` which is called from `perf_swevent_overflow` which is called from `perf_swevent_event` which is called from `do_perf_sw_event` which is called from `___perf_sw_event` which is called from `__perf_sw_event` which is called from `perf_sw_event`. `perf_sw_event` is called from the page fault handler! Bingo! It's reachable from unprivileged processes, and therefore reachable from user namespaces.
Now, is this a VM-guest trigger, host trigger, Ethernet, or peripheral trigger, or a malicious file system trigger? No, it's just a perf event bug, triggered by local syscalls.
Let's double-check the `perf_event_throttle_group`. In `__perf_event_account_interrupt`, I see:
```c
if (unlikely(throttle && hwc->interrupts >= max_samples_per_tick)) {
__this_cpu_inc(perf_throttled_count);
tick_dep_set_cpu(smp_processor_id(), TICK_DEP_BIT_PERF_EVENTS);
perf_event_throttle_group(event);
ret = 1;
}
```
Yep, it's called. Hold on, `max_samples_per_tick` is a sysctl. Its default value is `sysctl_perf_event_sample_rate`, which is 100000. So `max_samples_per_tick` is `100000 / HZ`. If `HZ` is 1000, then `max_samples_per_tick` is 100. That means we need 100 page faults in one tick (1ms). Very easy to achieve. But wait, `hwc->interrupts` is reset every tick? Let's check:
```c
seq = __this_cpu_read(perf_throttled_seq);
if (seq != hwc->interrupts_seq) {
hwc->interrupts_seq = seq;
hwc->interrupts = 1;
} else {
hwc->interrupts++;
}
```
Yes, `perf_throttled_seq` is incremented every tick in `perf_event_task_tick`. So, we need 100 page faults in one tick. This is incredibly easy.
Is there a catch? What if the event is NOT a group leader? No. `perf_event_throttle_group` throttles the leader *and* all its siblings. It iterates `leader->sibling_list`.
Now, the WARNING is: `WARN_ON_ONCE(__lockdep_enabled && (this_cpu_read(hardirqs_enabled) && lockdep_is_held(&(event)->ctx->mutex) != LOCK_STATE_HELD))`. If `lockdep_enabled` is false (production build), there's no warning. However, the data race / list corruption *still* exists. If we trigger the data race, we get the infinite loop. So, it's a DoS in production builds.
Finally, can this lead to privilege escalation? Let me think... `list_del_init` is called, `next` points to itself. `list_add_
|
| 69/2 |
2026/05/18 09:46 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:list_del_event]
Results: map[SourceCode:2110: /*
2111: * Remove an event from the lists for its context.
2112: * Must be called with ctx->mutex and ctx->lock held.
2113: */
2114: static void
2115: list_del_event(struct perf_event *event, struct perf_event_context *ctx)
2116: {
2117: WARN_ON_ONCE(event->ctx != ctx);
2118: lockdep_assert_held(&ctx->lock);
2119:
2120: /*
2121: * We can have double detach due to exit/hot-unplug + close.
2122: */
2123: if (!(event->attach_state & PERF_ATTACH_CONTEXT))
2124: return;
2125:
2126: event->attach_state &= ~PERF_ATTACH_CONTEXT;
2127:
2128: ctx->nr_events--;
2129: if (event->hw.flags & PERF_EVENT_FLAG_USER_READ_CNT)
2130: ctx->nr_user--;
2131: if (event->attr.inherit_stat)
2132: ctx->nr_stat--;
2133: if (has_inherit_and_sample_read(&event->attr))
2134: local_dec(&ctx->nr_no_switch_fast);
2135:
2136: list_del_rcu(&event->event_entry);
2137:
2138: if (event->group_leader == event)
2139: del_event_from_groups(event, ctx);
2140:
2141: ctx->generation++;
2142: event->pmu_ctx->nr_events--;
2143: }
SourceFile:kernel/events/core.c]
|
| 70/2 |
2026/05/18 09:46 |
llm |
expert |
0m
Model:
Tokens:
input: 77380
output: 43
thoughts: 464
Thoughts: **My Analysis of the Perf Event Throttle Group Vulnerability**
Okay, here's what I've figured out. It appears we have a denial-of-service vulnerability in the `perf_event_throttle_group` function. It's an infinite loop, not an exploitable memory corruption, thank goodness.
Let me break down my thought process: Initially, I saw that `list_del_event` removes from `event_entry` and `groups`, but crucially, *not* `sibling_list`. `sibling_list` is handled by `perf_group_attach` and `perf_group_detach`. These functions appear to only modify `sibling_list` with `list_del_init` and `list_add_tail`. The scary part was the `list_del_init` call β thatβs the one I need to focus on. It looks like this is what's causing the loop. The `list_add_tail` is safe. I've double-checked that there is no `LIST_POISON1` dereference. It seems to just create an endless loop. I can safely conclude: itβs a DoS, not an exploit.
Now let's delve into *why* this infinite loop happens. It seems that the bug resides in the `perf_event_throttle_group` function. This code iterates through the `sibling_list` without proper locking. No `ctx->lock` or `ctx->mutex` is held while IRQs are enabled, which is a major red flag.
The trigger appears to be software events (like page faults) or tracepoints overflowing and reaching `max_samples_per_tick`. While this code is iterating, another CPU can close the file descriptor of a sibling event which triggers `perf_group_detach`.
`perf_group_detach`, in turn, sends an IPI to the CPU where the event is active. The IPI's purpose is to interrupt the process. This interrupts `perf_event_throttle_group` because, as I said, IRQs are enabled. The IPI handler then makes a `list_del_init` call on the sibling.
The crucial point is that when the IPI handler returns, `perf_event_throttle_group` resumes its iteration, but now the `next` pointer within the iterator points to the sibling itself! This causes the infinite loop and, ultimately, a hard lockup β a denial-of-service condition.
The good news (relatively speaking) is that unprivileged processes can trigger this because they can create software events. However, and this is important, it's *not* exploitable for privilege escalation. The memory isn't freed (it's RCU protected), and we are using `list_del_init` instead of `list_del`. Therefore, the kernel doesn't give up control.
Just to be absolutely certain, I need to double-check that `perf_group_detach` is indeed called from the IPI handler. Yes, I'm confident that when a user closes the file descriptor, `perf_release` is called, which then calls `perf_event_release_kernel`. The IPI is responsible for this call.
|
| 71/2 |
2026/05/18 09:46 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_event_release_kernel]
Results: map[SourceCode:5763: /*
5764: * Kill an event dead; while event:refcount will preserve the event
5765: * object, it will not preserve its functionality. Once the last 'user'
5766: * gives up the object, we'll destroy the thing.
5767: */
5768: int perf_event_release_kernel(struct perf_event *event)
5769: {
5770: struct perf_event_context *ctx = event->ctx;
5771: struct perf_event *child, *tmp;
5772:
5773: /*
5774: * If we got here through err_alloc: free_event(event); we will not
5775: * have attached to a context yet.
5776: */
5777: if (!ctx) {
5778: WARN_ON_ONCE(event->attach_state &
5779: (PERF_ATTACH_CONTEXT|PERF_ATTACH_GROUP));
5780: goto no_ctx;
5781: }
5782:
5783: if (!is_kernel_event(event))
5784: perf_remove_from_owner(event);
5785:
5786: ctx = perf_event_ctx_lock(event);
5787: WARN_ON_ONCE(ctx->parent_ctx);
5788:
5789: /*
5790: * Mark this event as STATE_DEAD, there is no external reference to it
5791: * anymore.
5792: *
5793: * Anybody acquiring event->child_mutex after the below loop _must_
5794: * also see this, most importantly inherit_event() which will avoid
5795: * placing more children on the list.
5796: *
5797: * Thus this guarantees that we will in fact observe and kill _ALL_
5798: * child events.
5799: */
5800: if (event->state > PERF_EVENT_STATE_REVOKED) {
5801: perf_remove_from_context(event, DETACH_GROUP|DETACH_DEAD);
5802: } else {
5803: event->state = PERF_EVENT_STATE_DEAD;
5804: }
5805:
5806: perf_event_ctx_unlock(event, ctx);
5807:
5808: again:
5809: mutex_lock(&event->child_mutex);
5810: list_for_each_entry(child, &event->child_list, child_list) {
5811: /*
5812: * Cannot change, child events are not migrated, see the
5813: * comment with perf_event_ctx_lock_nested().
5814: */
5815: ctx = READ_ONCE(child->ctx);
5816: /*
5817: * Since child_mutex nests inside ctx::mutex, we must jump
5818: * through hoops. We start by grabbing a reference on the ctx.
5819: *
5820: * Since the event cannot get freed while we hold the
5821: * child_mutex, the context must also exist and have a !0
5822: * reference count.
5823: */
5824: get_ctx(ctx);
5825:
5826: /*
5827: * Now that we have a ctx ref, we can drop child_mutex, and
5828: * acquire ctx::mutex without fear of it going away. Then we
5829: * can re-acquire child_mutex.
5830: */
5831: mutex_unlock(&event->child_mutex);
5832: mutex_lock(&ctx->mutex);
5833: mutex_lock(&event->child_mutex);
5834:
5835: /*
5836: * Now that we hold ctx::mutex and child_mutex, revalidate our
5837: * state, if child is still the first entry, it didn't get freed
5838: * and we can continue doing so.
5839: */
5840: tmp = list_first_entry_or_null(&event->child_list,
5841: struct perf_event, child_list);
5842: if (tmp == child) {
5843: perf_remove_from_context(child, DETACH_GROUP | DETACH_CHILD);
5844: } else {
5845: child = NULL;
5846: }
5847:
5848: mutex_unlock(&event->child_mutex);
5849: mutex_unlock(&ctx->mutex);
5850:
5851: if (child) {
5852: /* Last reference unless ->pending_task work is pending */
5853: put_event(child);
5854: }
5855: put_ctx(ctx);
5856:
5857: goto again;
5858: }
5859: mutex_unlock(&event->child_mutex);
5860:
5861: no_ctx:
5862: /*
5863: * Last reference unless ->pending_task work is pending on this event
5864: * or any of its children.
5865: */
5866: put_event(event);
5867: return 0;
5868: }
SourceFile:kernel/events/core.c]
|
| 72/2 |
2026/05/18 09:46 |
llm |
expert |
0m
Model:
Tokens:
input: 80109
output: 43
thoughts: 32
Thoughts: Okay, here's my interpretation of that thought process, from an expert's perspective:
**Deconstructing `perf_event_release_kernel`'s Implications**
Right, so `perf_event_release_kernel` is being called, which immediately tells me we're cleaning up a kernel-space performance event. The critical thing here is what that function *does*. It's calling `perf_remove_from_context`, and that's the core of the cleanup operation. The flags passed, `DETACH_GROUP | DETACH_DEAD`, immediately jump out. `DETACH_GROUP` means we're breaking the association of the event with its performance monitoring group, and that's essential for a clean shutdown. It's about disentangling the event from its logical container. But `DETACH_DEAD` is even more significant; it means the event is considered "dead," as in, it's not going to be used anymore, and it needs to be completely removed. This implies it's either finished its intended duration, or being prematurely shut down, or has encountered some sort of irrecoverable error, meaning it must be detached and completely removed from any ongoing context. It gives me a good idea of what is happening when an event is done with its lifecycle.
|
| 73/2 |
2026/05/18 09:46 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_remove_from_context]
Results: map[SourceCode:2536: /*
2537: * Remove the event from a task's (or a CPU's) list of events.
2538: *
2539: * If event->ctx is a cloned context, callers must make sure that
2540: * every task struct that event->ctx->task could possibly point to
2541: * remains valid. This is OK when called from perf_release since
2542: * that only calls us on the top-level context, which can't be a clone.
2543: * When called from perf_event_exit_task, it's OK because the
2544: * context has been detached from its task.
2545: */
2546: static void perf_remove_from_context(struct perf_event *event, unsigned long flags)
2547: {
2548: struct perf_event_context *ctx = event->ctx;
2549:
2550: lockdep_assert_held(&ctx->mutex);
2551:
2552: /*
2553: * Because of perf_event_exit_task(), perf_remove_from_context() ought
2554: * to work in the face of TASK_TOMBSTONE, unlike every other
2555: * event_function_call() user.
2556: */
2557: raw_spin_lock_irq(&ctx->lock);
2558: if (!ctx->is_active) {
2559: __perf_remove_from_context(event, this_cpu_ptr(&perf_cpu_context),
2560: ctx, (void *)flags);
2561: raw_spin_unlock_irq(&ctx->lock);
2562: return;
2563: }
2564: raw_spin_unlock_irq(&ctx->lock);
2565:
2566: event_function_call(event, __perf_remove_from_context, (void *)flags);
2567: }
SourceFile:kernel/events/core.c]
|
| 74/2 |
2026/05/18 09:46 |
llm |
expert |
0m
Model:
Tokens:
input: 81021
output: 44
thoughts: 36
Thoughts: Here's a summary of my understanding:
**Analyzing the Event Handling Flow**
Okay, so I'm tracing how events are managed within the system, specifically when they are removed. I'm focusing on the call to `event_function_call`. This function is pivotal; it triggers an Inter-Processor Interrupt (IPI), and this IPI is directed to the CPU where the event in question is currently active. The key part is that the IPI handler, the code that actually responds to this interrupt, is `__perf_remove_from_context`. This clearly tells me that the removal of the event from its context (likely the CPU's performance monitoring context) is handled asynchronously by another CPU. This design is interesting for concurrency and avoiding lock contention during the removal process. I need to dig deeper into `__perf_remove_from_context` now to see what steps are taken during the removal.
|
| 75/2 |
2026/05/18 09:46 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:__perf_remove_from_context]
Results: map[SourceCode:2471: /*
2472: * Cross CPU call to remove a performance event
2473: *
2474: * We disable the event on the hardware level first. After that we
2475: * remove it from the context list.
2476: */
2477: static void
2478: __perf_remove_from_context(struct perf_event *event,
2479: struct perf_cpu_context *cpuctx,
2480: struct perf_event_context *ctx,
2481: void *info)
2482: {
2483: struct perf_event_pmu_context *pmu_ctx = event->pmu_ctx;
2484: enum perf_event_state state = PERF_EVENT_STATE_OFF;
2485: unsigned long flags = (unsigned long)info;
2486:
2487: ctx_time_update(cpuctx, ctx);
2488:
2489: /*
2490: * Ensure event_sched_out() switches to OFF, at the very least
2491: * this avoids raising perf_pending_task() at this time.
2492: */
2493: if (flags & DETACH_EXIT)
2494: state = PERF_EVENT_STATE_EXIT;
2495: if (flags & DETACH_REVOKE)
2496: state = PERF_EVENT_STATE_REVOKED;
2497: if (flags & DETACH_DEAD)
2498: state = PERF_EVENT_STATE_DEAD;
2499:
2500: event_sched_out(event, ctx);
2501:
2502: if (event->state > PERF_EVENT_STATE_OFF)
2503: perf_cgroup_event_disable(event, ctx);
2504:
2505: perf_event_set_state(event, min(event->state, state));
2506:
2507: if (flags & DETACH_GROUP)
2508: perf_group_detach(event);
2509: if (flags & DETACH_CHILD)
2510: perf_child_detach(event);
2511: list_del_event(event, ctx);
2512:
2513: if (!pmu_ctx->nr_events) {
2514: pmu_ctx->rotate_necessary = 0;
2515:
2516: if (ctx->task && ctx->is_active) {
2517: struct perf_cpu_pmu_context *cpc = this_cpc(pmu_ctx->pmu);
2518:
2519: WARN_ON_ONCE(cpc->task_epc && cpc->task_epc != pmu_ctx);
2520: cpc->task_epc = NULL;
2521: }
2522: }
2523:
2524: if (!ctx->nr_events && ctx->is_active) {
2525: if (ctx == &cpuctx->ctx)
2526: update_cgrp_time_from_cpuctx(cpuctx, true);
2527:
2528: ctx->is_active = 0;
2529: if (ctx->task) {
2530: WARN_ON_ONCE(cpuctx->task_ctx != ctx);
2531: cpuctx->task_ctx = NULL;
2532: }
2533: }
2534: }
SourceFile:kernel/events/core.c]
|
| 76/2 |
2026/05/18 09:46 |
llm |
expert |
0m
Model:
Tokens:
input: 82293
output: 66
thoughts: 1040
Thoughts: **Analysis of the Perf Event Group Detachment Denial of Service**
Okay, let's break this down. My initial thought was: `__perf_remove_from_context` calls `perf_group_detach`, which in turn calls `list_del_init` on the `sibling_list` of an event. This `__perf_remove_from_context` function is called from `event_function_call`, which runs on a CPU with IRQs disabled and the context lock held. Now, the key is `perf_event_throttle_group` which, from what I can see, *runs with IRQs enabled*. This sets up a classic IPI (Inter-Processor Interrupt) race.
So, here's what happens: `perf_event_throttle_group` is running, an IPI interrupts it. The IPI handler acquires the context lock, and then executes `__perf_remove_from_context`. This detaches the event from the group via `list_del_init`. The IPI then returns, and `perf_event_throttle_group` resumes. Because the list entry was removed out from under it by the IPI, the list iteration gets into an infinite loop. **That's a solid Denial of Service (DoS) right there.**
My next thought was around reachability and exploitability. This is *reachable from unprivileged processes* because software events are easily created and triggered. It's not limited to user namespaces, although it's accessible through them as well. The good news (from a security perspective) is that it's *not exploitable for privilege escalation* because memory isn't freed. It's an RCU-protected `list_del_init`, so no double-free or use-after-free.
So I quickly filled in the results: Exploitability: False; DenialOfService: True; Unprivileged: True; UserNamespace: True; and the rest are False.
But then I thought "wait, is there any chance of memory corruption?" I remembered `put_event` being called after `perf_remove_from_context`. `put_event` decrements a reference count, and if it hits zero, it calls `_free_event`. `_free_event` calls `call_rcu` on `event->rcu_head` with `free_event_rcu`. Since `do_perf_sw_event` (which calls `perf_swevent_overflow` and eventually `perf_event_throttle_group`) gets stuck in a loop inside `rcu_read_lock()`, the RCU grace period *never* ends. Therefore, `free_event_rcu` is never called, and the memory is never freed. It's a hard lockup, not a use-after-free.
Then I realized, what about tracepoints? Tracepoints can also call `perf_event_throttle_group` -- and that means `rcu_read_lock_sched_notrace()` or `rcu_read_lock_sched()` is used! So, preemption is disabled within that function, leading to another infinite loop. Again, the grace period never completes, and the memory is never freed.
Alright, what else could trigger `perf_event_throttle_group`? Hardware events. I considered hardware events triggering NMIs or hardware IRQs. In those cases, hardirqs are disabled, meaning an IPI cannot interrupt `perf_event_throttle_group`. The IPI is pending until `perf_event_throttle_group` finishes, which means for hardware events, no race condition exists! **The race condition *only* exists for software events and tracepoints** because they can call `perf_event_throttle_group` with IRQs enabled. And since unprivileged users can create software events, it makes this scenario easily triggerable.
Finally, I considered what would happen if a user creates both hardware and software events in the same group. If the software event is the leader, the group overflows, and we trigger the race. If the hardware event is the leader, the group triggers from NMI/IRQ with IRQs disabled and no race. It's still easily reproducible with just software events.
I went back and confirmed that `perf_swevent_overflow` can, indeed, call `perf_event_throttle_group` via a fairly straightforward function call path from `do_perf_sw_event`. `do_perf_sw_event` calls `perf_swevent_event`, which calls `perf_swevent_overflow`. `perf_swevent_overflow` then calls `__perf_event_overflow`, which calls `__perf_event_account_interrupt`, which finally calls `perf_event_throttle_group`.
And with that, I'm ready to
|
| 77/2 |
2026/05/18 09:47 |
tool |
set-results |
0m
Args: map[DenialOfService:true Exploitable:false FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:true UserNamespace:true VMGuestTrigger:false VMHostTrigger:false]
Results: map[DenialOfService:true Exploitable:false FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:true UserNamespace:true VMGuestTrigger:false VMHostTrigger:false]
|
| 78/2 |
2026/05/18 09:47 |
llm |
expert |
0m
Model:
Tokens:
input: 84488
output: 733
thoughts: 440
Thoughts: **Final Thoughts on the Perf Event Throttle Group Vulnerability**
Okay, the analysis is complete, and it's time to craft the final response. Hereβs how I see this:
The crux of the matter is the `perf_event_throttle_group` function. It's the culprit. The warning flags a serious issue: it's iterating over `event->sibling_list` without the necessary synchronization. Specifically, it's missing the protection of either `ctx->mutex` or `ctx->lock`, and it does this while hard interrupt requests are enabled. This is a recipe for disaster.
This unsynchronized iteration is triggered by software events, like page faults, and by tracepoints. These are the "hooks" that can lead to overflows, which in turn call `perf_event_throttle_group` with interrupts active. The key here is the enabled interrupts. Because IRQs are enabled during the critical section, a potential Inter-Processor Interrupt (IPI) can interrupt the list iteration, and that's precisely what happens.
Now, consider this: If another thread closes the file descriptor associated with a sibling event concurrently, this triggers a cascade of events: `perf_release` calls `perf_event_release_kernel` which then calls `perf_remove_from_context`. This chain of events sends an IPI to the CPU where the "offending" event is currently active.
That IPI handler, `__perf_remove_from_context`, acquires `ctx->lock` and then calls `perf_group_detach`. This function removes the sibling event from the list. The list removal happens using `list_del_init`, which is important for the race condition.
When the IPI handler returns, `perf_event_throttle_group` resumes its iteration from where it left off. Because `list_del_init` modifies the list entry to point to itself, the `list_for_each_entry` loop in `perf_event_throttle_group` now enters an infinite loop. This, in turn, causes a hard lockup on the CPU. The CPU is stuck in the loop. This constitutes a Denial of Service, bringing the system to a standstill.
Importantly, this is not an exploit for privilege escalation. The `perf_event` memory is protected by RCU (Read-Copy Update). The CPU is stuck inside an RCU read-side critical section, like `rcu_read_lock` in `do_perf_sw_event`. Because the CPU is stuck, the RCU grace period never ends. So, the memory never gets freed. Therefore, there's no opportunity for a use-after-free situation, as the memory remains valid within that critical section.
The vulnerability's impact is that an unprivileged user can trigger this. They need to create a group of software events, let's say `PERF_COUNT_SW_PAGE_FAULTS`, set a low sampling period, and then trigger a lot of page faults in a loop. Simultaneously, they concurrently close the file descriptor of a sibling event. This controlled concurrent interaction results in the race condition and, ultimately, the system lockup.
|
| 79/1 |
2026/05/18 09:47 |
action |
format-explanation |
0m
Results: map[Explanation:The bug report indicates a WARNING in `perf_event_throttle_group` triggered by
`lockdep_assert_event_ctx(leader)`.
### Root Cause Analysis
The `perf_event_throttle_group` function iterates over the `event->sibling_list`
of a perf event group using the `for_each_sibling_event` macro. This macro
includes a lockdep assertion that checks if `ctx->mutex` is held whenever
hardirqs are enabled.
For hardware perf events, overflows are handled in NMI or hardware interrupt
context where hardirqs are disabled, making it safe to iterate the list without
locks (as IPIs are blocked). However, for **software events** (e.g.,
`PERF_COUNT_SW_PAGE_FAULTS`) and **tracepoints**, overflows are handled
synchronously in the context of the triggering instruction. In these cases,
`perf_event_throttle_group` can be reached with **hardirqs enabled**.
Because hardirqs are enabled and no locks (`ctx->lock` or `ctx->mutex`) are
held, the iteration over `sibling_list` is vulnerable to concurrent
modifications via Inter-Processor Interrupts (IPIs).
### Impact: Denial of Service
If a malicious user creates a group of software events and triggers overflows
rapidly (e.g., by causing page faults in a loop), `perf_event_throttle_group`
will be invoked. Concurrently, the user can close the file descriptor of a
sibling event from another thread.
Closing the file descriptor invokes `perf_release` ->
`perf_event_release_kernel` -> `perf_remove_from_context`, which sends an IPI to
the CPU where the event is active. The IPI handler
(`__perf_remove_from_context`) will interrupt `perf_event_throttle_group`,
acquire `ctx->lock`, and call `perf_group_detach`.
`perf_group_detach` removes the sibling from the list using
`list_del_init(&sibling->sibling_list)`. This reinitializes the sibling's list
pointers to point to itself. When the IPI returns, `perf_event_throttle_group`
resumes its `list_for_each_entry` iteration. Because the `next` pointer now
points to the element itself, the loop will **infinite loop**, causing a hard
lockup on the CPU.
### Exploitability
This bug is **not exploitable for privilege escalation or memory corruption**.
Although it is a concurrent list modification, the `perf_event` structures are
freed via RCU (`call_rcu`). Since the software event overflow handler (e.g.,
`do_perf_sw_event`) runs inside an RCU read-side critical section
(`rcu_read_lock()`), the CPU getting stuck in the infinite loop will prevent the
RCU grace period from ever ending. Consequently, the memory is never freed,
precluding any Use-After-Free (UAF) scenarios. It strictly results in a Denial
of Service (DoS).
### Reachability
The bug is reachable by **unprivileged users** (and within user namespaces).
While unprivileged users cannot create tracepoint events under the default
`sysctl_perf_event_paranoid = 2` configuration, they are fully permitted to
create and group `PERF_TYPE_SOFTWARE` events (like page faults or context
switches) which are sufficient to trigger the vulnerable code path.]
|