| 0/0 |
2026/05/25 21:22 |
flow |
assessment-security |
59m
Results: map[DenialOfService:true Explanation:Based on the analysis of the bug report and the kernel source code, here is the
assessment of the security impact:
**Bug Mechanism:**
The bug is an RCU stall (Denial of Service) caused by a livelock in the timer
interrupt handler. The reproducer achieves this by loading a BPF program of type
`BPF_PROG_TYPE_RAW_TRACEPOINT` and attaching it **65 times** to the
`rcu_utilization` tracepoint. This tracepoint is invoked from
`rcu_sched_clock_irq`, which runs in the context of the high-frequency timer
interrupt (`sysvec_apic_timer_interrupt`).
Each time the timer interrupt fires, the BPF program is executed 65 times. The
program calls `bpf_trace_printk`, which formats a string and writes it to the
ftrace ring buffer. If the kernel is compiled with the debug option
`CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS` (which the syzbot kernel has enabled),
every write to the ring buffer triggers `check_buffer()`. This debugging
function iterates over all events in the current buffer page, resulting in an
$O(N^2)$ time complexity per page.
The combination of 65 `bpf_trace_printk` calls per interrupt and the $O(N^2)$
validation overhead causes the timer interrupt handler to take longer to execute
than the actual interval between timer interrupts (e.g., >1ms). This completely
starves the CPU, leading to an interrupt storm and the reported RCU stall.
**Security Impact:**
* **Exploitable:** No. The bug does not cause memory corruption, out-of-bounds
accesses, or use-after-free conditions. It is purely a CPU starvation issue
(Denial of Service).
* **Accessible From Unprivileged Processes:** No. To trigger this bug, the user
must:
1. Load a BPF program of type `BPF_PROG_TYPE_RAW_TRACEPOINT`, which strictly
requires `CAP_PERFMON` or `CAP_SYS_ADMIN`. (Even if
`unprivileged_bpf_disabled=0`, unprivileged users are restricted to socket
filters and cgroup skb programs).
2. Use the `bpf_trace_printk` helper, which also explicitly requires
`CAP_PERFMON`.
3. Attach the program to a raw tracepoint, which rejects unprivileged program
types.
* **Accessible From User Namespaces:** No. While BPF tokens were introduced to
allow delegating BPF permissions to user namespaces, creating a BPF token
requires mounting `bpffs` with delegation options (e.g., `delegate_progs`). The
kernel explicitly checks for `capable(CAP_SYS_ADMIN)` in the **initial** user
namespace when setting these mount options. Therefore, an unprivileged user
cannot create a user namespace and grant themselves the ability to load tracing
BPF programs.
**Conclusion:**
This is not a security vulnerability. It is a Denial of Service that can only be
triggered by a highly privileged user (root or a user with
`CAP_PERFMON`/`CAP_SYS_ADMIN` in the initial namespace). Privileged users
already have numerous ways to halt the system (e.g., loading an infinite loop
BPF program, or writing to `/proc/sysrq-trigger`). Furthermore, the extreme
overhead is heavily dependent on a debug configuration
(`CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS`) which is explicitly documented as
"Do not use it on production systems." Exploitable:false FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
|
| 1/1 |
2026/05/25 21:22 |
action |
syz-repro-to-c-repro |
0m
Results: map[SimplifiedCRepro:// autogenerated by syzkaller (https://github.com/google/syzkaller)
#define _GNU_SOURCE
#include <endian.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/mount.h>
#include <sys/syscall.h>
#include <sys/types.h>
#include <unistd.h>
#ifndef __NR_bpf
#define __NR_bpf 321
#endif
uint64_t r[4] = {0xffffffffffffffff, 0xffffffffffffffff, 0xffffffffffffffff, 0xffffffffffffffff};
int main(void)
{
syscall(__NR_mmap, /*addr=*/0x1ffffffff000ul, /*len=*/0x1000ul, /*prot=*/0ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
syscall(__NR_mmap, /*addr=*/0x200000000000ul, /*len=*/0x1000000ul, /*prot=PROT_WRITE|PROT_READ|PROT_EXEC*/7ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
syscall(__NR_mmap, /*addr=*/0x200001000000ul, /*len=*/0x1000ul, /*prot=*/0ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
const char* reason;
(void)reason;
intptr_t res = 0;
if (write(1, "executing program\n", sizeof("executing program\n") - 1)) {}
// bpf$BPF_PROG_RAW_TRACEPOINT_LOAD arguments: [
// cmd: const = 0x5 (8 bytes)
// arg: ptr[in, bpf_prog_t[flags[bpf_raw_tracepoint_prog_types, int32], const[0, int32], const[0, int32], const[0, int32]]] {
// bpf_prog_t[flags[bpf_raw_tracepoint_prog_types, int32], const[0, int32], const[0, int32], const[0, int32]] {
// type: bpf_raw_tracepoint_prog_types = 0x18 (4 bytes)
// ninsn: bytesize8 = 0xb (4 bytes)
// insns: ptr[inout, array[ANYUNION]] {
// array[ANYUNION] {
// union ANYUNION {
// ANYBLOB: buffer: {18 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 18 01 00 00 20 20 00 00 00 00 00 00 00 00 00 00 7b 1a f8 ff 00 00 00 00 bf a1 00 00 00 00 00 00 07 01 00 00 f8 ff ff ff b7 02 00 00 08 00 00 00 b7 03 00 00 00 00 00 00 85 00 00 00 06 00 00 00 95} (length 0x51)
// }
// }
// }
// license: ptr[in, buffer] {
// buffer: {47 50 4c 00} (length 0x4)
// }
// loglev: int32 = 0x0 (4 bytes)
// logsize: len = 0x0 (4 bytes)
// log: nil
// kern_version: bpf_kern_version = 0x1f00 (4 bytes)
// flags: bpf_prog_load_flags = 0x0 (4 bytes)
// prog_name: buffer: {00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00} (length 0x10)
// prog_ifindex: ifindex (resource)
// expected_attach_type: const = 0x2 (4 bytes)
// btf_fd: fd_btf (resource)
// func_info_rec_size: const = 0x8 (4 bytes)
// func_info: nil
// func_info_cnt: len = 0x0 (4 bytes)
// line_info_rec_size: const = 0x10 (4 bytes)
// line_info: nil
// line_info_cnt: len = 0x0 (4 bytes)
// attach_btf_id: const = 0x0 (4 bytes)
// attach_prog_fd: const = 0x0 (4 bytes)
// core_relo_cnt: len = 0x0 (4 bytes)
// fd_array: nil
// core_relos: nil
// core_relo_rec_size: const = 0x10 (4 bytes)
// log_true_size: int32 = 0x0 (4 bytes)
// prog_token_fd: union _bpf_prog_t[flags[bpf_raw_tracepoint_prog_types, int32], const[0, int32], const[0, int32], const[0, int32]]_prog_token_fd_wrapper {
// void: buffer: {} (length 0x0)
// }
// pad: union _bpf_prog_t[flags[bpf_raw_tracepoint_prog_types, int32], const[0, int32], const[0, int32], const[0, int32]]_pad_wrapper {
// value: const = 0x0 (4 bytes)
// }
// }
// }
// size: len = 0x94 (8 bytes)
// ]
// returns fd_bpf_prog_raw_tracepoint
*(uint32_t*)0x2000000006c0 = 0x18;
*(uint32_t*)0x2000000006c4 = 0xb;
*(uint64_t*)0x2000000006c8 = 0x2000000000c0;
memcpy((void*)0x2000000000c0, "... [truncated large byte array] ...", 81);
*(uint64_t*)0x2000000006d0 = 0x2000000004c0;
memcpy((void*)0x2000000004c0, "GPL\000", 4);
*(uint32_t*)0x2000000006d8 = 0;
*(uint32_t*)0x2000000006dc = 0;
*(uint64_t*)0x2000000006e0 = 0;
*(uint32_t*)0x2000000006e8 = 0x1f00;
*(uint32_t*)0x2000000006ec = 0;
memset((void*)0x2000000006f0, 0, 16);
*(uint32_t*)0x200000000700 = 0;
*(uint32_t*)0x200000000704 = 2;
*(uint32_t*)0x200000000708 = -1;
*(uint32_t*)0x20000000070c = 8;
*(uint64_t*)0x200000000710 = 0;
*(uint32_t*)0x200000000718 = 0;
*(uint32_t*)0x20000000071c = 0x10;
*(uint64_t*)0x200000000720 = 0;
*(uint32_t*)0x200000000728 = 0;
*(uint32_t*)0x20000000072c = 0;
*(uint32_t*)0x200000000730 = 0;
*(uint32_t*)0x200000000734 = 0;
*(uint64_t*)0x200000000738 = 0;
*(uint64_t*)0x200000000740 = 0;
*(uint32_t*)0x200000000748 = 0x10;
*(uint32_t*)0x20000000074c = 0;
*(uint32_t*)0x200000000750 = 0;
res = syscall(__NR_bpf, /*cmd=*/5ul, /*arg=*/0x2000000006c0ul, /*size=*/0x94ul);
if (res != -1)
r[0] = res;
// bpf$BPF_RAW_TRACEPOINT_OPEN arguments: [
// cmd: const = 0x11 (8 bytes)
// arg: ptr[in, bpf_raw_tracepoint] {
// bpf_raw_tracepoint {
// name: ptr[in, buffer] {
// buffer: {72 63 75 5f 75 74 69 6c 69 7a 61 74 69 6f 6e 00} (length 0x10)
// }
// prog_fd: fd_bpf_prog_raw_tracepoint (resource)
// pad: const = 0x0 (4 bytes)
// cookie: int64 = 0x0 (8 bytes)
// }
// }
// size: len = 0x10 (8 bytes)
// ]
// returns fd_perf_base
*(uint64_t*)0x200000000540 = 0x200000000300;
memcpy((void*)0x200000000300, "rcu_utilization\000", 16);
*(uint32_t*)0x200000000548 = r[0];
*(uint32_t*)0x20000000054c = 0;
*(uint64_t*)0x200000000550 = 0;
syscall(__NR_bpf, /*cmd=*/0x11ul, /*arg=*/0x200000000540ul, /*size=*/0x10ul);
for (int i = 0; i < 64; i++) {
syscall(__NR_bpf, /*cmd=*/0x11ul, /*arg=*/0x200000000540ul, /*size=*/0x10ul);
}
// ioctl$sock_ipv6_tunnel_SIOCADDTUNNEL arguments: [
// fd: sock_udp6 (resource)
// cmd: const = 0x89f1 (4 bytes)
// arg: nil
// ]
syscall(__NR_ioctl, /*fd=*/(intptr_t)-1, /*cmd=*/0x89f1, /*arg=*/0ul);
// ioctl$sock_SIOCSIFBR arguments: [
// fd: sock (resource)
// cmd: const = 0x8941 (4 bytes)
// arg: nil
// ]
syscall(__NR_ioctl, /*fd=*/(intptr_t)-1, /*cmd=*/0x8941, /*arg=*/0ul);
// setsockopt$inet6_int arguments: [
// fd: sock_in6 (resource)
// level: const = 0x29 (4 bytes)
// optname: inet6_option_types_int = 0x4e (4 bytes)
// optval: ptr[in, int32] {
// int32 = 0xfffffff7 (4 bytes)
// }
// optlen: len = 0x4 (8 bytes)
// ]
*(uint32_t*)0x200000000100 = 0xfffffff7;
syscall(__NR_setsockopt, /*fd=*/(intptr_t)-1, /*level=*/0x29, /*optname=IPV6_FREEBIND*/0x4e, /*optval=*/0x200000000100ul, /*optlen=*/4ul);
// unshare arguments: [
// flags: unshare_flags = 0x600 (8 bytes)
// ]
syscall(__NR_unshare, /*flags=CLONE_FS|CLONE_FILES*/0x600ul);
// sendmmsg arguments: [
// fd: sock (resource)
// mmsg: ptr[in, array[send_mmsghdr]] {
// array[send_mmsghdr] {
// }
// }
// vlen: len = 0x800001d (8 bytes)
// f: send_flags = 0x0 (8 bytes)
// ]
syscall(__NR_sendmmsg, /*fd=*/(intptr_t)-1, /*mmsg=*/0x200000007fc0ul, /*vlen=*/0x800001dul, /*f=*/0ul);
// socket$alg arguments: [
// domain: const = 0x26 (8 bytes)
// type: const = 0x5 (8 bytes)
// proto: const = 0x0 (4 bytes)
// ]
// returns sock_alg
res = syscall(__NR_socket, /*domain=*/0x26ul, /*type=*/5ul, /*proto=*/0);
if (res != -1)
r[1] = res;
// accept4 arguments: [
// fd: sock (resource)
// peer: nil
// peerlen: nil
// flags: accept_flags = 0x0 (8 bytes)
// ]
// returns sock
res = syscall(__NR_accept4, /*fd=*/r[1], /*peer=*/0ul, /*peerlen=*/0ul, /*flags=*/0ul);
if (res != -1)
r[2] = res;
// epoll_create1 arguments: [
// flags: epoll_flags = 0x80000 (8 bytes)
// ]
// returns fd_epoll
res = syscall(__NR_epoll_create1, /*flags=EPOLL_CLOEXEC*/0x80000ul);
if (res != -1)
r[3] = res;
// epoll_ctl$EPOLL_CTL_ADD arguments: [
// epfd: fd_epoll (resource)
// op: const = 0x1 (8 bytes)
// fd: fd (resource)
// ev: ptr[in, epoll_event] {
// epoll_event {
// ev: epoll_ev = 0xc000200a (4 bytes)
// data: const = 0x0 (8 bytes)
// }
// }
// ]
*(uint32_t*)0x200000000040 = 0xc000200a;
*(uint64_t*)0x200000000044 = 0;
syscall(__NR_epoll_ctl, /*epfd=*/r[3], /*op=*/1ul, /*fd=*/r[2], /*ev=*/0x200000000040ul);
// recvmmsg arguments: [
// fd: sock (resource)
// mmsg: nil
// vlen: len = 0x0 (8 bytes)
// f: recv_flags = 0x45833af92e4b39ff (8 bytes)
// timeout: nil
// ]
syscall(__NR_recvmmsg, /*fd=*/(intptr_t)-1, /*mmsg=*/0ul, /*vlen=*/0ul, /*f=MSG_WAITFORONE|MSG_WAITALL|MSG_TRUNC|MSG_PEEK|MSG_OOB|MSG_ERRQUEUE|MSG_DONTWAIT|0x45833af92e4a189c*/0x45833af92e4b39fful, /*timeout=*/0ul);
return 0;
}
]
|
| 2/1 |
2026/05/25 21:22 |
action |
kernel-checkouter |
2m
Results: map[KernelSrc:/app/workdir/cache/src/b10f0e4fa21f8d4c242b2472b08564950b79e116]
|
| 3/1 |
2026/05/25 21:24 |
action |
kernel-builder |
22m
Results: map[KernelObj:/app/workdir/cache/build/fa73223a2c9db799c0db680af52cc7e7389df470]
|
| 4/1 |
2026/05/25 21:47 |
action |
codesearch-prepare |
24m
Results: map[Index:codesearch-index]
|
| 5/1 |
2026/05/25 22:11 |
agent |
expert |
10m
Model:
Results: map[DenialOfService:true Exploitable:false FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
Instruction:
You are an experienced Linux kernel security engineer. Your task is to analyze given kernel bug report
and determine its security impact based on the following dimensions.
Use the provided tools to examine the source code, check for capability checks (e.g., capable(), ns_capable()),
and understand the nature of the bug. Analyze the given kernel build and configuration.
You can check the kernel config by grepping ".config" file; you can check kernel cmdline by greeping
".config" file for "CONFIG_CMDLINE=". Assume sysctl parameters have default values.
But analyze for the corresponding production build w/o debugging tools enabled (like KASAN, KMSAN, UBSAN).
Try different strategies when analyzing the bug:
- think of ways in which the vulnerable code is unreachable
- or the other way around: try to come up with different ideas of how an unprivileged user can reach the bug
If still unsure err on the side of the bug being non-exploitable/not-accessible.
In the final reply, provide a reasoning for your assessment.
Analysis dimensions:
* Exploitable:
Determine if the bug can result in memory corruption or elevated privileges.
Memory safety issues are almost always exploitable (KASAN or UBSAN reports for use-after-free, out-of-bounds;
refcounting issues, corrupted lists, etc). When kernel is crashing on a completly wild pointer access
(e.g. user-space address, or non-canonical address, but not on NULL or address corresponding to KASAN shadow
for NULL address), including both data accesses and control tranfers, that's also usually implies possibility
of exploitation. Such reports usually say "unable to handle kernel paging request".
Uses of uninitialized values detected by KMSAN may be exploitable b/c attacker frequently can affect uninit
values with spraying techniques. However, for these exploitabability depends on how exactly the uninit value
is used in the code, and what it affects.
Think of what happens after the bug is triggered. Some bugs cause kernel panic and halt execution,
they are harder to exploit. For example, BUG reports halts the kernel. However, WARNING reports don't halt
execution in production builds. Debug bug detection tools (like KASAN, KMSAN, KCSAN, UBSAN) are also not enabled
in production builds, so attacker can freely exploit these bugs w/o being detected by these tools.
If you see an integer overflow, think how the overflowed value used later (if it's used as allocation size,
or an array index). If you see an out-of-bounds read, think if it's followed by an out-of-bounds write as well.
Some KCSAN data-races may be exploitable by skilled attackers as well. Think what data structures got corrupted
as the result of data races and how. However, note that kernel has lots of "benign" data races that don't lead
to any runtime misbehavior at all.
* Denial Of Service:
Determine if the bug can result in denial-of-service. Most bugs can, since they cause system crash,
hangs, deadlocks, or resource leaks. This is mostly applicable to WARNING bugs that won't cause system crash
in production. For these think what will be consequences of the violation of the kernel assumptions flagged
by the WARNING. In some cases the unexpected condition is also properly handled by the normal control flow
(e.g. with "if (WARN_ON(...))"), these won't cause denial-of-service. If the condition is not handled,
then it may or may not cause denial-of-service.
* Accessible From Unprivileged Processes:
Determine if the bug can be reached from a typical (non-root) user process that does NOT have any special capabilities
(like CAP_SYS_ADMIN, CAP_NET_ADMIN, CAP_NET_RAW, CAP_PERFMON) or access to device nodes restricted to root.
Assume that unprivileged_bpf_disabled=1, that is eBPF loading is not accessible. However, cBPF (classical BPF)
is still accessible to non-root processes.
Assume that user namespaces are not accessible, that is, the process cannot get the mentioned capabilities even
within a new user namespace (checked by ns_capable() function in the kernel sources).
* Accessible From User Namespaces:
Determine if the bug can be reached within a user-namespace where the process has all capabilities
(including CAP_SYS_ADMIN, CAP_NET_ADMIN, CAP_NET_RAW, CAP_PERFMON). Such capabilities are checked with ns_capable()
function in the kernel sources.
* VM Guest Trigger:
Determine if the bug can be triggered from the context of a typical KVM guest (e.g., set up by a QEMU VMM).
Consider accesses to standard Linux host paravirtualized features (virtio-blk, virtio-net, etc.),
and handling of VM exits in the KVM code.
* VM Host Trigger in The Confidetial Computing Context:
Determine if the bug can be triggered in a confidential computing guest kernel from the context of a KVM host.
Consider access to standard Linux guest paravirtualized features (virtio-blk, virtio-net, etc.).
* Ethernet Network Trigger:
Determine if the bug can be triggered by processing ingress network Ethernet traffic, either directly (network stack)
or via drivers exposed to network data.
* Other Remote Trigger:
Determine if the bug can be triggered by processing remote traffic other than Ethernet (Wifi, Bluetooth, NFC, etc).
* Peripheral Trigger:
Determine if the bug can be triggered via an untrusted peripheral device that can be physically plugged
into a system, such as a USB device or a niche hardware driver handling external hardware inputs.
This is particularly important for mobile and desktop environments where users can plug in unknown devices.
* Malicious Filesystem Trigger:
Determine if the bug can be triggered by the kernel mounting and parsing a malicious filesystem image.
This is highly critical for Desktop and Mobile environments where external media or downloaded images
might be auto-mounted.
Don't make assumptions about the kernel source code (it may be different from what you assume it is).
Extensively use the provided code access tools (codesearch-*, git-*, grepper, etc)
to examine the actual source code, and confirm any assumptions.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt:
The kernel bug report is:
rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:
rcu: (detected by 0, t=10502 jiffies, g=18609, q=500 ncpus=2)
rcu: All QSes seen, last rcu_preempt kthread activity 10502 (4294978752-4294968250), jiffies_till_next_fqs=1, root ->qsmask 0x0
rcu: rcu_preempt kthread starved for 10502 jiffies! g18609 f0x2 RCU_GP_WAIT_FQS(5) ->state=0x0 ->cpu=1
rcu: Unless rcu_preempt kthread gets sufficient CPU time, OOM is now expected behavior.
rcu: RCU grace-period kthread stack dump:
task:rcu_preempt state:R running task stack:27272 pid:16 tgid:16 ppid:2 task_flags:0x208040 flags:0x00080000
Call Trace:
<TASK>
context_switch kernel/sched/core.c:5256 [inline]
__schedule+0x14bc/0x5000 kernel/sched/core.c:6863
__schedule_loop kernel/sched/core.c:6945 [inline]
schedule+0x165/0x360 kernel/sched/core.c:6960
schedule_timeout+0x12b/0x270 kernel/time/sleep_timeout.c:99
rcu_gp_fqs_loop+0x301/0x1540 kernel/rcu/tree.c:2083
rcu_gp_kthread+0x99/0x390 kernel/rcu/tree.c:2285
kthread+0x711/0x8a0 kernel/kthread.c:463
ret_from_fork+0x599/0xb30 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:246
</TASK>
rcu: Stack dump where RCU GP kthread last ran:
Sending NMI from CPU 0 to CPUs 1:
NMI backtrace for cpu 1
CPU: 1 UID: 0 PID: 31 Comm: khungtaskd Not tainted syzkaller #0 PREEMPT(full)
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 10/25/2025
RIP: 0010:check_kcov_mode kernel/kcov.c:183 [inline]
RIP: 0010:write_comp_data kernel/kcov.c:246 [inline]
RIP: 0010:__sanitizer_cov_trace_switch+0x97/0x130 kernel/kcov.c:351
Code: 54 53 48 8b 54 24 20 65 4c 8b 04 25 08 b0 7e 92 45 31 c9 eb 08 49 ff c1 4c 39 c8 74 77 4e 8b 54 ce 10 65 44 8b 1d c9 f7 bc 10 <41> 81 e3 00 01 ff 00 74 13 41 81 fb 00 01 00 00 75 d9 41 83 b8 6c
RSP: 0018:ffffc90000a082e0 EFLAGS: 00000016
RAX: 0000000000000020 RBX: 0000000000000004 RCX: 0000000000000005
RDX: ffffffff81c39275 RSI: ffffffff8df9a280 RDI: 0000000000000004
RBP: 0000006100d347d2 R08: ffff88801e2b0000 R09: 000000000000001b
R10: 000000000000001b R11: 0000000080010005 R12: dffffc0000000000
R13: ffff88801bed6010 R14: ffff88801bed64c0 R15: 00000000000667c4
FS: 0000000000000000(0000) GS:ffff888125f35000(0000) knlGS:0000000000000000
CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 000055d3cfba3a38 CR3: 00000000727a4000 CR4: 00000000003526f0
Call Trace:
<IRQ>
rb_event_length+0x45/0x400 kernel/trace/ring_buffer.c:222
rb_read_data_buffer+0x438/0x580 kernel/trace/ring_buffer.c:1858
check_buffer+0x28a/0x750 kernel/trace/ring_buffer.c:4410
__rb_reserve_next+0x592/0xdb0 kernel/trace/ring_buffer.c:4509
rb_reserve_next_event kernel/trace/ring_buffer.c:4646 [inline]
ring_buffer_lock_reserve+0xbb5/0x1010 kernel/trace/ring_buffer.c:4705
__trace_buffer_lock_reserve kernel/trace/trace.c:1079 [inline]
trace_event_buffer_lock_reserve+0x1d0/0x6f0 kernel/trace/trace.c:2808
trace_event_buffer_reserve+0x248/0x340 kernel/trace/trace_events.c:672
do_trace_event_raw_event_bpf_trace_printk kernel/trace/bpf_trace.h:11 [inline]
trace_event_raw_event_bpf_trace_printk+0x100/0x260 kernel/trace/bpf_trace.h:11
__do_trace_bpf_trace_printk kernel/trace/bpf_trace.h:11 [inline]
trace_bpf_trace_printk+0x153/0x1b0 kernel/trace/bpf_trace.h:11
____bpf_trace_printk kernel/trace/bpf_trace.c:379 [inline]
bpf_trace_printk+0x11e/0x190 kernel/trace/bpf_trace.c:362
bpf_prog_b1367f0be6c54012+0x39/0x3f
bpf_dispatcher_nop_func include/linux/bpf.h:1378 [inline]
__bpf_prog_run include/linux/filter.h:723 [inline]
bpf_prog_run include/linux/filter.h:730 [inline]
__bpf_trace_run kernel/trace/bpf_trace.c:2075 [inline]
bpf_trace_run1+0x27f/0x4c0 kernel/trace/bpf_trace.c:2115
__bpf_trace_rcu_utilization+0xa1/0xf0 include/trace/events/rcu.h:27
__traceiter_rcu_utilization+0x7a/0xb0 include/trace/events/rcu.h:27
__do_trace_rcu_utilization include/trace/events/rcu.h:27 [inline]
trace_rcu_utilization+0x191/0x1c0 include/trace/events/rcu.h:27
rcu_sched_clock_irq+0xd3/0x1280 kernel/rcu/tree.c:2693
update_process_times+0x23c/0x2f0 kernel/time/timer.c:2474
tick_sched_handle kernel/time/tick-sched.c:298 [inline]
tick_nohz_handler+0x3e9/0x710 kernel/time/tick-sched.c:319
__run_hrtimer kernel/time/hrtimer.c:1777 [inline]
__hrtimer_run_queues+0x4d0/0xc30 kernel/time/hrtimer.c:1841
hrtimer_interrupt+0x45b/0xaa0 kernel/time/hrtimer.c:1903
local_apic_timer_interrupt arch/x86/kernel/apic/apic.c:1045 [inline]
__sysvec_apic_timer_interrupt+0x102/0x3e0 arch/x86/kernel/apic/apic.c:1062
instr_sysvec_apic_timer_interrupt arch/x86/kernel/apic/apic.c:1056 [inline]
sysvec_apic_timer_interrupt+0xa1/0xc0 arch/x86/kernel/apic/apic.c:1056
</IRQ>
<TASK>
asm_sysvec_apic_timer_interrupt+0x1a/0x20 arch/x86/include/asm/idtentry.h:697
RIP: 0010:console_trylock_spinning kernel/printk/printk.c:2037 [inline]
RIP: 0010:vprintk_emit+0x4d0/0x5f0 kernel/printk/printk.c:2425
Code: 0f 84 34 ff ff ff e8 cf 62 20 00 fb eb 44 e8 c7 62 20 00 e8 82 72 ba 09 4d 85 f6 74 94 e8 b8 62 20 00 fb 48 c7 c7 e0 58 f3 8d <31> f6 ba 01 00 00 00 31 c9 41 b8 01 00 00 00 45 31 c9 53 e8 28 26
RSP: 0018:ffffc90000a77a80 EFLAGS: 00000293
RAX: ffffffff81a14f98 RBX: ffffffff81a14df1 RCX: ffff88801e2b0000
RDX: 0000000000000000 RSI: 0000000000000000 RDI: ffffffff8df358e0
RBP: ffffc90000a77b90 R08: ffffffff8f822077 R09: 1ffffffff1f0440e
R10: dffffc0000000000 R11: fffffbfff1f0440f R12: 0000000000000000
R13: 0000000000000040 R14: 0000000000000200 R15: 1ffff9200014ef54
_printk+0xcf/0x120 kernel/printk/printk.c:2451
check_hung_task kernel/hung_task.c:255 [inline]
check_hung_uninterruptible_tasks kernel/hung_task.c:331 [inline]
watchdog+0xb35/0xfe0 kernel/hung_task.c:515
kthread+0x711/0x8a0 kernel/kthread.c:463
ret_from_fork+0x599/0xb30 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:246
</TASK>
It is reproducible with the followint program.
Keep in mind that it may lack the precise threading, sandboxing, and some arguments of a working reproducer.
But it should give an idea of the involved syscalls.
// autogenerated by syzkaller (https://github.com/google/syzkaller)
#define _GNU_SOURCE
#include <endian.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/mount.h>
#include <sys/syscall.h>
#include <sys/types.h>
#include <unistd.h>
#ifndef __NR_bpf
#define __NR_bpf 321
#endif
uint64_t r[4] = {0xffffffffffffffff, 0xffffffffffffffff, 0xffffffffffffffff, 0xffffffffffffffff};
int main(void)
{
syscall(__NR_mmap, /*addr=*/0x1ffffffff000ul, /*len=*/0x1000ul, /*prot=*/0ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
syscall(__NR_mmap, /*addr=*/0x200000000000ul, /*len=*/0x1000000ul, /*prot=PROT_WRITE|PROT_READ|PROT_EXEC*/7ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
syscall(__NR_mmap, /*addr=*/0x200001000000ul, /*len=*/0x1000ul, /*prot=*/0ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
const char* reason;
(void)reason;
intptr_t res = 0;
if (write(1, "executing program\n", sizeof("executing program\n") - 1)) {}
// bpf$BPF_PROG_RAW_TRACEPOINT_LOAD arguments: [
// cmd: const = 0x5 (8 bytes)
// arg: ptr[in, bpf_prog_t[flags[bpf_raw_tracepoint_prog_types, int32], const[0, int32], const[0, int32], const[0, int32]]] {
// bpf_prog_t[flags[bpf_raw_tracepoint_prog_types, int32], const[0, int32], const[0, int32], const[0, int32]] {
// type: bpf_raw_tracepoint_prog_types = 0x18 (4 bytes)
// ninsn: bytesize8 = 0xb (4 bytes)
// insns: ptr[inout, array[ANYUNION]] {
// array[ANYUNION] {
// union ANYUNION {
// ANYBLOB: buffer: {18 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 18 01 00 00 20 20 00 00 00 00 00 00 00 00 00 00 7b 1a f8 ff 00 00 00 00 bf a1 00 00 00 00 00 00 07 01 00 00 f8 ff ff ff b7 02 00 00 08 00 00 00 b7 03 00 00 00 00 00 00 85 00 00 00 06 00 00 00 95} (length 0x51)
// }
// }
// }
// license: ptr[in, buffer] {
// buffer: {47 50 4c 00} (length 0x4)
// }
// loglev: int32 = 0x0 (4 bytes)
// logsize: len = 0x0 (4 bytes)
// log: nil
// kern_version: bpf_kern_version = 0x1f00 (4 bytes)
// flags: bpf_prog_load_flags = 0x0 (4 bytes)
// prog_name: buffer: {00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00} (length 0x10)
// prog_ifindex: ifindex (resource)
// expected_attach_type: const = 0x2 (4 bytes)
// btf_fd: fd_btf (resource)
// func_info_rec_size: const = 0x8 (4 bytes)
// func_info: nil
// func_info_cnt: len = 0x0 (4 bytes)
// line_info_rec_size: const = 0x10 (4 bytes)
// line_info: nil
// line_info_cnt: len = 0x0 (4 bytes)
// attach_btf_id: const = 0x0 (4 bytes)
// attach_prog_fd: const = 0x0 (4 bytes)
// core_relo_cnt: len = 0x0 (4 bytes)
// fd_array: nil
// core_relos: nil
// core_relo_rec_size: const = 0x10 (4 bytes)
// log_true_size: int32 = 0x0 (4 bytes)
// prog_token_fd: union _bpf_prog_t[flags[bpf_raw_tracepoint_prog_types, int32], const[0, int32], const[0, int32], const[0, int32]]_prog_token_fd_wrapper {
// void: buffer: {} (length 0x0)
// }
// pad: union _bpf_prog_t[flags[bpf_raw_tracepoint_prog_types, int32], const[0, int32], const[0, int32], const[0, int32]]_pad_wrapper {
// value: const = 0x0 (4 bytes)
// }
// }
// }
// size: len = 0x94 (8 bytes)
// ]
// returns fd_bpf_prog_raw_tracepoint
*(uint32_t*)0x2000000006c0 = 0x18;
*(uint32_t*)0x2000000006c4 = 0xb;
*(uint64_t*)0x2000000006c8 = 0x2000000000c0;
memcpy((void*)0x2000000000c0, "... [truncated large byte array] ...", 81);
*(uint64_t*)0x2000000006d0 = 0x2000000004c0;
memcpy((void*)0x2000000004c0, "GPL\000", 4);
*(uint32_t*)0x2000000006d8 = 0;
*(uint32_t*)0x2000000006dc = 0;
*(uint64_t*)0x2000000006e0 = 0;
*(uint32_t*)0x2000000006e8 = 0x1f00;
*(uint32_t*)0x2000000006ec = 0;
memset((void*)0x2000000006f0, 0, 16);
*(uint32_t*)0x200000000700 = 0;
*(uint32_t*)0x200000000704 = 2;
*(uint32_t*)0x200000000708 = -1;
*(uint32_t*)0x20000000070c = 8;
*(uint64_t*)0x200000000710 = 0;
*(uint32_t*)0x200000000718 = 0;
*(uint32_t*)0x20000000071c = 0x10;
*(uint64_t*)0x200000000720 = 0;
*(uint32_t*)0x200000000728 = 0;
*(uint32_t*)0x20000000072c = 0;
*(uint32_t*)0x200000000730 = 0;
*(uint32_t*)0x200000000734 = 0;
*(uint64_t*)0x200000000738 = 0;
*(uint64_t*)0x200000000740 = 0;
*(uint32_t*)0x200000000748 = 0x10;
*(uint32_t*)0x20000000074c = 0;
*(uint32_t*)0x200000000750 = 0;
res = syscall(__NR_bpf, /*cmd=*/5ul, /*arg=*/0x2000000006c0ul, /*size=*/0x94ul);
if (res != -1)
r[0] = res;
// bpf$BPF_RAW_TRACEPOINT_OPEN arguments: [
// cmd: const = 0x11 (8 bytes)
// arg: ptr[in, bpf_raw_tracepoint] {
// bpf_raw_tracepoint {
// name: ptr[in, buffer] {
// buffer: {72 63 75 5f 75 74 69 6c 69 7a 61 74 69 6f 6e 00} (length 0x10)
// }
// prog_fd: fd_bpf_prog_raw_tracepoint (resource)
// pad: const = 0x0 (4 bytes)
// cookie: int64 = 0x0 (8 bytes)
// }
// }
// size: len = 0x10 (8 bytes)
// ]
// returns fd_perf_base
*(uint64_t*)0x200000000540 = 0x200000000300;
memcpy((void*)0x200000000300, "rcu_utilization\000", 16);
*(uint32_t*)0x200000000548 = r[0];
*(uint32_t*)0x20000000054c = 0;
*(uint64_t*)0x200000000550 = 0;
syscall(__NR_bpf, /*cmd=*/0x11ul, /*arg=*/0x200000000540ul, /*size=*/0x10ul);
for (int i = 0; i < 64; i++) {
syscall(__NR_bpf, /*cmd=*/0x11ul, /*arg=*/0x200000000540ul, /*size=*/0x10ul);
}
// ioctl$sock_ipv6_tunnel_SIOCADDTUNNEL arguments: [
// fd: sock_udp6 (resource)
// cmd: const = 0x89f1 (4 bytes)
// arg: nil
// ]
syscall(__NR_ioctl, /*fd=*/(intptr_t)-1, /*cmd=*/0x89f1, /*arg=*/0ul);
// ioctl$sock_SIOCSIFBR arguments: [
// fd: sock (resource)
// cmd: const = 0x8941 (4 bytes)
// arg: nil
// ]
syscall(__NR_ioctl, /*fd=*/(intptr_t)-1, /*cmd=*/0x8941, /*arg=*/0ul);
// setsockopt$inet6_int arguments: [
// fd: sock_in6 (resource)
// level: const = 0x29 (4 bytes)
// optname: inet6_option_types_int = 0x4e (4 bytes)
// optval: ptr[in, int32] {
// int32 = 0xfffffff7 (4 bytes)
// }
// optlen: len = 0x4 (8 bytes)
// ]
*(uint32_t*)0x200000000100 = 0xfffffff7;
syscall(__NR_setsockopt, /*fd=*/(intptr_t)-1, /*level=*/0x29, /*optname=IPV6_FREEBIND*/0x4e, /*optval=*/0x200000000100ul, /*optlen=*/4ul);
// unshare arguments: [
// flags: unshare_flags = 0x600 (8 bytes)
// ]
syscall(__NR_unshare, /*flags=CLONE_FS|CLONE_FILES*/0x600ul);
// sendmmsg arguments: [
// fd: sock (resource)
// mmsg: ptr[in, array[send_mmsghdr]] {
// array[send_mmsghdr] {
// }
// }
// vlen: len = 0x800001d (8 bytes)
// f: send_flags = 0x0 (8 bytes)
// ]
syscall(__NR_sendmmsg, /*fd=*/(intptr_t)-1, /*mmsg=*/0x200000007fc0ul, /*vlen=*/0x800001dul, /*f=*/0ul);
// socket$alg arguments: [
// domain: const = 0x26 (8 bytes)
// type: const = 0x5 (8 bytes)
// proto: const = 0x0 (4 bytes)
// ]
// returns sock_alg
res = syscall(__NR_socket, /*domain=*/0x26ul, /*type=*/5ul, /*proto=*/0);
if (res != -1)
r[1] = res;
// accept4 arguments: [
// fd: sock (resource)
// peer: nil
// peerlen: nil
// flags: accept_flags = 0x0 (8 bytes)
// ]
// returns sock
res = syscall(__NR_accept4, /*fd=*/r[1], /*peer=*/0ul, /*peerlen=*/0ul, /*flags=*/0ul);
if (res != -1)
r[2] = res;
// epoll_create1 arguments: [
// flags: epoll_flags = 0x80000 (8 bytes)
// ]
// returns fd_epoll
res = syscall(__NR_epoll_create1, /*flags=EPOLL_CLOEXEC*/0x80000ul);
if (res != -1)
r[3] = res;
// epoll_ctl$EPOLL_CTL_ADD arguments: [
// epfd: fd_epoll (resource)
// op: const = 0x1 (8 bytes)
// fd: fd (resource)
// ev: ptr[in, epoll_event] {
// epoll_event {
// ev: epoll_ev = 0xc000200a (4 bytes)
// data: const = 0x0 (8 bytes)
// }
// }
// ]
*(uint32_t*)0x200000000040 = 0xc000200a;
*(uint64_t*)0x200000000044 = 0;
syscall(__NR_epoll_ctl, /*epfd=*/r[3], /*op=*/1ul, /*fd=*/r[2], /*ev=*/0x200000000040ul);
// recvmmsg arguments: [
// fd: sock (resource)
// mmsg: nil
// vlen: len = 0x0 (8 bytes)
// f: recv_flags = 0x45833af92e4b39ff (8 bytes)
// timeout: nil
// ]
syscall(__NR_recvmmsg, /*fd=*/(intptr_t)-1, /*mmsg=*/0ul, /*vlen=*/0ul, /*f=MSG_WAITFORONE|MSG_WAITALL|MSG_TRUNC|MSG_PEEK|MSG_OOB|MSG_ERRQUEUE|MSG_DONTWAIT|0x45833af92e4a189c*/0x45833af92e4b39fful, /*timeout=*/0ul);
return 0;
}
Reply: Based on the analysis of the bug report and the kernel source code, here is the assessment of the security impact:
**Bug Mechanism:**
The bug is an RCU stall (Denial of Service) caused by a livelock in the timer interrupt handler. The reproducer achieves this by loading a BPF program of type `BPF_PROG_TYPE_RAW_TRACEPOINT` and attaching it **65 times** to the `rcu_utilization` tracepoint. This tracepoint is invoked from `rcu_sched_clock_irq`, which runs in the context of the high-frequency timer interrupt (`sysvec_apic_timer_interrupt`).
Each time the timer interrupt fires, the BPF program is executed 65 times. The program calls `bpf_trace_printk`, which formats a string and writes it to the ftrace ring buffer. If the kernel is compiled with the debug option `CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS` (which the syzbot kernel has enabled), every write to the ring buffer triggers `check_buffer()`. This debugging function iterates over all events in the current buffer page, resulting in an $O(N^2)$ time complexity per page.
The combination of 65 `bpf_trace_printk` calls per interrupt and the $O(N^2)$ validation overhead causes the timer interrupt handler to take longer to execute than the actual interval between timer interrupts (e.g., >1ms). This completely starves the CPU, leading to an interrupt storm and the reported RCU stall.
**Security Impact:**
* **Exploitable:** No. The bug does not cause memory corruption, out-of-bounds accesses, or use-after-free conditions. It is purely a CPU starvation issue (Denial of Service).
* **Accessible From Unprivileged Processes:** No. To trigger this bug, the user must:
1. Load a BPF program of type `BPF_PROG_TYPE_RAW_TRACEPOINT`, which strictly requires `CAP_PERFMON` or `CAP_SYS_ADMIN`. (Even if `unprivileged_bpf_disabled=0`, unprivileged users are restricted to socket filters and cgroup skb programs).
2. Use the `bpf_trace_printk` helper, which also explicitly requires `CAP_PERFMON`.
3. Attach the program to a raw tracepoint, which rejects unprivileged program types.
* **Accessible From User Namespaces:** No. While BPF tokens were introduced to allow delegating BPF permissions to user namespaces, creating a BPF token requires mounting `bpffs` with delegation options (e.g., `delegate_progs`). The kernel explicitly checks for `capable(CAP_SYS_ADMIN)` in the **initial** user namespace when setting these mount options. Therefore, an unprivileged user cannot create a user namespace and grant themselves the ability to load tracing BPF programs.
**Conclusion:**
This is not a security vulnerability. It is a Denial of Service that can only be triggered by a highly privileged user (root or a user with `CAP_PERFMON`/`CAP_SYS_ADMIN` in the initial namespace). Privileged users already have numerous ways to halt the system (e.g., loading an infinite loop BPF program, or writing to `/proc/sysrq-trigger`). Furthermore, the extreme overhead is heavily dependent on a debug configuration (`CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS`) which is explicitly documented as "Do not use it on production systems."
|
| 6/2 |
2026/05/25 22:11 |
llm |
expert |
0m
Model:
Tokens:
input: 13203
output: 17
thoughts: 2211
Thoughts: **RCU Stall Investigation: A Deep Dive into a Ring Buffer Bottleneck**
Okay, so I've got this RCU stall report, and my expert intuition is screaming "trace buffer contention." The log clearly points to a `rcu_preempt` stall, specifically in the `rcu_sched_clock_irq` context, which suggests timer interrupts are involved. The stack trace is interesting, starting with `rcu_sched_clock_irq`, diving into `trace_rcu_utilization`, and then hitting our culprit: `bpf_trace_printk`. This smells like a BPF program writing to the trace buffer within a performance-critical path.
The reproducer loads a BPF program, type `BPF_PROG_TYPE_RAW_TRACEPOINT`, and attaches it a whopping *65 times* to the `rcu_utilization` tracepoint. Each time the timer fires and calls `rcu_sched_clock_irq`, the tracepoint is triggered, and thus the BPF program runs 65 times per timer interrupt, each potentially calling `bpf_trace_printk`. This looks like a recipe for disaster. The stack trace confirms this route, showing `bpf_trace_printk`'s calls to `trace_event_buffer_reserve`, which leads to the infamous `ring_buffer_lock_reserve`.
Now, if that trace buffer is busy, full, or if there's significant lock contention in `ring_buffer_lock_reserve`, we're going to get some serious delays. The trace buffer isn't some magical infinite sink; it has limits, which it will eventually hit when you call `bpf_trace_printk` from within `rcu_sched_clock_irq` 65 times.
Then I see it: `check_buffer` in the stack trace, which is a debugging function related to the integrity of the ring buffer. It's not supposed to be running here! It's there to flag issues, not be part of the mainline. If that's hitting, it's a huge sign that the ring buffer code is being stressed, which is the smoking gun! I need to examine `kernel/trace/ring_buffer.c`, specifically `check_buffer`, to see what conditions are triggering that. This is likely where the performance bottleneck is.
|
| 7/2 |
2026/05/25 22:12 |
tool |
grepper |
0m
Args: map[Expression:check_buffer]
Results: map[Output:drivers/android/binder_alloc.c=1279=void binder_alloc_shrinker_exit(void)
--
drivers/android/binder_alloc.c-1285-/**
drivers/android/binder_alloc.c:1286: * check_buffer() - verify that buffer/offset is safe to access
drivers/android/binder_alloc.c-1287- * @alloc: binder_alloc for this proc
--
drivers/android/binder_alloc.c-1303- */
drivers/android/binder_alloc.c:1304:static inline bool check_buffer(struct binder_alloc *alloc,
drivers/android/binder_alloc.c-1305- struct binder_buffer *buffer,
--
drivers/android/binder_alloc.c=1330=binder_alloc_copy_user_to_buffer(struct binder_alloc *alloc,
--
drivers/android/binder_alloc.c-1335-{
drivers/android/binder_alloc.c:1336: if (!check_buffer(alloc, buffer, buffer_offset, bytes))
drivers/android/binder_alloc.c-1337- return bytes;
--
drivers/android/binder_alloc.c=1361=static int binder_alloc_do_buffer_copy(struct binder_alloc *alloc,
--
drivers/android/binder_alloc.c-1368- /* All copies must be 32-bit aligned and 32-bit size */
drivers/android/binder_alloc.c:1369: if (!check_buffer(alloc, buffer, buffer_offset, bytes))
drivers/android/binder_alloc.c-1370- return -EINVAL;
--
drivers/android/tests/binder_alloc_kunit.c=128=static void stringify_alignments(struct kunit *test, int *alignments,
--
drivers/android/tests/binder_alloc_kunit.c-139-
drivers/android/tests/binder_alloc_kunit.c:140:static bool check_buffer_pages_allocated(struct kunit *test,
drivers/android/tests/binder_alloc_kunit.c-141- struct binder_alloc *alloc,
--
drivers/android/tests/binder_alloc_kunit.c=164=static unsigned long binder_alloc_test_alloc_buf(struct kunit *test,
--
drivers/android/tests/binder_alloc_kunit.c-174- if (IS_ERR(buffers[i]) ||
drivers/android/tests/binder_alloc_kunit.c:175: !check_buffer_pages_allocated(test, alloc, buffers[i], sizes[i]))
drivers/android/tests/binder_alloc_kunit.c-176- failures++;
--
drivers/gpu/drm/i915/gem/i915_gem_execbuffer.c=3550=static size_t eb_element_size(void)
--
drivers/gpu/drm/i915/gem/i915_gem_execbuffer.c-3554-
drivers/gpu/drm/i915/gem/i915_gem_execbuffer.c:3555:static bool check_buffer_count(size_t count)
drivers/gpu/drm/i915/gem/i915_gem_execbuffer.c-3556-{
--
drivers/gpu/drm/i915/gem/i915_gem_execbuffer.c=3569=i915_gem_execbuffer2_ioctl(struct drm_device *dev, void *data,
--
drivers/gpu/drm/i915/gem/i915_gem_execbuffer.c-3577-
drivers/gpu/drm/i915/gem/i915_gem_execbuffer.c:3578: if (!check_buffer_count(count)) {
drivers/gpu/drm/i915/gem/i915_gem_execbuffer.c-3579- drm_dbg(&i915->drm, "execbuf2 with %zd buffers\n", count);
--
drivers/gpu/drm/i915/gem/i915_gem_execbuffer.c-3618- * Note: count * sizeof(*user_exec_list) does not overflow,
drivers/gpu/drm/i915/gem/i915_gem_execbuffer.c:3619: * because we checked 'count' in check_buffer_count().
drivers/gpu/drm/i915/gem/i915_gem_execbuffer.c-3620- *
--
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c-14-
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c:15:static int igt_fill_check_buffer(struct drm_i915_gem_object *obj,
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c-16- struct intel_gt *gt,
--
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c=46=static int igt_create_migrate(struct intel_gt *gt, enum intel_region_id src,
--
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c-68-
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c:69: err = igt_fill_check_buffer(obj, gt, true);
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c-70- if (err)
--
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c-88-
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c:89: err = igt_fill_check_buffer(obj, gt, false);
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c-90- }
--
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c=182=static int __igt_lmem_pages_migrate(struct intel_gt *gt,
--
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c-235- if (!vma) {
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c:236: err = igt_fill_check_buffer(obj, gt, true);
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c-237- if (err)
--
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c-278- } else {
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c:279: err = igt_fill_check_buffer(obj, gt, false);
drivers/gpu/drm/i915/gem/selftests/i915_gem_migrate.c-280- }
--
drivers/gpu/drm/vmwgfx/vmwgfx_resource.c=452=void vmw_resource_unreserve(struct vmw_resource *res,
--
drivers/gpu/drm/vmwgfx/vmwgfx_resource.c-504-/**
drivers/gpu/drm/vmwgfx/vmwgfx_resource.c:505: * vmw_resource_check_buffer - Check whether a backup buffer is needed
drivers/gpu/drm/vmwgfx/vmwgfx_resource.c-506- * for a resource and in that case, allocate
--
drivers/gpu/drm/vmwgfx/vmwgfx_resource.c=516=static int
drivers/gpu/drm/vmwgfx/vmwgfx_resource.c:517:vmw_resource_check_buffer(struct ww_acquire_ctx *ticket,
drivers/gpu/drm/vmwgfx/vmwgfx_resource.c-518- struct vmw_resource *res,
--
drivers/gpu/drm/vmwgfx/vmwgfx_resource.c=634=static int vmw_resource_do_evict(struct ww_acquire_ctx *ticket,
--
drivers/gpu/drm/vmwgfx/vmwgfx_resource.c-644- val_buf.num_shared = 0;
drivers/gpu/drm/vmwgfx/vmwgfx_resource.c:645: ret = vmw_resource_check_buffer(ticket, res, interruptible, &val_buf);
drivers/gpu/drm/vmwgfx/vmwgfx_resource.c-646- if (unlikely(ret != 0))
--
drivers/media/platform/amphion/vpu_rpc.c=21=int vpu_iface_check_memory_region(struct vpu_core *core, dma_addr_t addr, u32 size)
--
drivers/media/platform/amphion/vpu_rpc.c-30-
drivers/media/platform/amphion/vpu_rpc.c:31:static u32 vpu_rpc_check_buffer_space(struct vpu_rpc_buffer_desc *desc, bool write)
drivers/media/platform/amphion/vpu_rpc.c-32-{
--
drivers/media/platform/amphion/vpu_rpc.c=56=static int vpu_rpc_send_cmd_buf(struct vpu_shared_addr *shared, struct vpu_rpc_event *cmd)
--
drivers/media/platform/amphion/vpu_rpc.c-66- desc = shared->cmd_desc;
drivers/media/platform/amphion/vpu_rpc.c:67: space = vpu_rpc_check_buffer_space(desc, true);
drivers/media/platform/amphion/vpu_rpc.c-68- if (space < (((cmd->hdr.num + 1) << 2) + 16))
--
drivers/media/platform/amphion/vpu_rpc.c=100=static bool vpu_rpc_check_msg(struct vpu_shared_addr *shared)
--
drivers/media/platform/amphion/vpu_rpc.c-107- desc = shared->msg_desc;
drivers/media/platform/amphion/vpu_rpc.c:108: space = vpu_rpc_check_buffer_space(desc, 0);
drivers/media/platform/amphion/vpu_rpc.c-109- space = (space >> 2);
--
drivers/media/platform/chips-media/coda/coda-bit.c=323=void coda_fill_bitstream(struct coda_ctx *ctx, struct list_head *buffer_list)
--
drivers/media/platform/chips-media/coda/coda-bit.c-364- if (ctx->codec->src_fourcc == V4L2_PIX_FMT_JPEG &&
drivers/media/platform/chips-media/coda/coda-bit.c:365: !coda_jpeg_check_buffer(ctx, &src_buf->vb2_buf)) {
drivers/media/platform/chips-media/coda/coda-bit.c-366- v4l2_err(&ctx->dev->v4l2_dev,
--
drivers/media/platform/chips-media/coda/coda-jpeg.c=215=int coda_jpeg_write_tables(struct coda_ctx *ctx)
--
drivers/media/platform/chips-media/coda/coda-jpeg.c-240-
drivers/media/platform/chips-media/coda/coda-jpeg.c:241:bool coda_jpeg_check_buffer(struct coda_ctx *ctx, struct vb2_buffer *vb)
drivers/media/platform/chips-media/coda/coda-jpeg.c-242-{
--
drivers/media/platform/chips-media/coda/coda.h=383=static inline int coda_jpeg_scale(int src, int dst)
--
drivers/media/platform/chips-media/coda/coda.h-389-
drivers/media/platform/chips-media/coda/coda.h:390:bool coda_jpeg_check_buffer(struct coda_ctx *ctx, struct vb2_buffer *vb);
drivers/media/platform/chips-media/coda/coda.h-391-int coda_jpeg_decode_header(struct coda_ctx *ctx, struct vb2_buffer *vb);
--
drivers/mtd/nand/onenand/onenand_base.c=873=static int onenand_get_2x_blockpage(struct mtd_info *mtd, loff_t addr)
--
drivers/mtd/nand/onenand/onenand_base.c-889-/**
drivers/mtd/nand/onenand/onenand_base.c:890: * onenand_check_bufferram - [GENERIC] Check BufferRAM information
drivers/mtd/nand/onenand/onenand_base.c-891- * @mtd: MTD data structure
--
drivers/mtd/nand/onenand/onenand_base.c-896- */
drivers/mtd/nand/onenand/onenand_base.c:897:static int onenand_check_bufferram(struct mtd_info *mtd, loff_t addr)
drivers/mtd/nand/onenand/onenand_base.c-898-{
--
drivers/mtd/nand/onenand/onenand_base.c=1114=static int onenand_mlc_read_ops_nolock(struct mtd_info *mtd, loff_t from,
--
drivers/mtd/nand/onenand/onenand_base.c-1153-
drivers/mtd/nand/onenand/onenand_base.c:1154: if (!onenand_check_bufferram(mtd, from)) {
drivers/mtd/nand/onenand/onenand_base.c-1155- this->command(mtd, ONENAND_CMD_READ, from, writesize);
--
drivers/mtd/nand/onenand/onenand_base.c=1215=static int onenand_read_ops_nolock(struct mtd_info *mtd, loff_t from,
--
drivers/mtd/nand/onenand/onenand_base.c-1249- if (read < len) {
drivers/mtd/nand/onenand/onenand_base.c:1250: if (!onenand_check_bufferram(mtd, from)) {
drivers/mtd/nand/onenand/onenand_base.c-1251- this->command(mtd, ONENAND_CMD_READ, from, writesize);
--
drivers/net/wireless/rsi/rsi_91x_sdio.c=825=static int rsi_init_sdio_interface(struct rsi_hw *adapter,
--
drivers/net/wireless/rsi/rsi_91x_sdio.c-869- adapter->determine_event_timeout = rsi_sdio_determine_event_timeout;
drivers/net/wireless/rsi/rsi_91x_sdio.c:870: adapter->check_hw_queue_status = rsi_sdio_check_buffer_status;
drivers/net/wireless/rsi/rsi_91x_sdio.c-871-
--
drivers/net/wireless/rsi/rsi_91x_sdio_ops.c=231=static void rsi_rx_handler(struct rsi_hw *adapter)
--
drivers/net/wireless/rsi/rsi_91x_sdio_ops.c-266- if (isr_status & BIT(PKT_BUFF_AVAILABLE)) {
drivers/net/wireless/rsi/rsi_91x_sdio_ops.c:267: status = rsi_sdio_check_buffer_status(adapter, 0);
drivers/net/wireless/rsi/rsi_91x_sdio_ops.c-268- if (status < 0)
--
drivers/net/wireless/rsi/rsi_91x_sdio_ops.c-335- */
drivers/net/wireless/rsi/rsi_91x_sdio_ops.c:336:int rsi_sdio_check_buffer_status(struct rsi_hw *adapter, u8 q_num)
drivers/net/wireless/rsi/rsi_91x_sdio_ops.c-337-{
--
drivers/net/wireless/rsi/rsi_sdio.h=135=int rsi_sdio_determine_event_timeout(struct rsi_hw *adapter);
drivers/net/wireless/rsi/rsi_sdio.h:136:int rsi_sdio_check_buffer_status(struct rsi_hw *adapter, u8 q_num);
drivers/net/wireless/rsi/rsi_sdio.h-137-void rsi_sdio_rx_thread(struct rsi_common *common);
--
drivers/scsi/ipr.c=7715=static void ipr_unit_check_no_data(struct ipr_ioa_cfg *ioa_cfg)
--
drivers/scsi/ipr.c-7721-/**
drivers/scsi/ipr.c:7722: * ipr_get_unit_check_buffer - Get the unit check buffer from the IOA
drivers/scsi/ipr.c-7723- * @ioa_cfg: ioa config struct
--
drivers/scsi/ipr.c-7730- **/
drivers/scsi/ipr.c:7731:static void ipr_get_unit_check_buffer(struct ipr_ioa_cfg *ioa_cfg)
drivers/scsi/ipr.c-7732-{
--
drivers/scsi/ipr.c=7796=static int ipr_reset_get_unit_check_job(struct ipr_cmnd *ipr_cmd)
--
drivers/scsi/ipr.c-7801- ioa_cfg->ioa_unit_checked = 0;
drivers/scsi/ipr.c:7802: ipr_get_unit_check_buffer(ioa_cfg);
drivers/scsi/ipr.c-7803- ipr_cmd->job_step = ipr_reset_alert;
--
drivers/scsi/ipr.c=7881=static int ipr_reset_restore_cfg_space(struct ipr_cmnd *ipr_cmd)
--
drivers/scsi/ipr.c-7908- ioa_cfg->ioa_unit_checked = 0;
drivers/scsi/ipr.c:7909: ipr_get_unit_check_buffer(ioa_cfg);
drivers/scsi/ipr.c-7910- ipr_cmd->job_step = ipr_reset_alert;
--
fs/btrfs/extent_io.c=3070=struct extent_buffer *alloc_dummy_extent_buffer(struct btrfs_fs_info *fs_info,
--
fs/btrfs/extent_io.c-3104-
fs/btrfs/extent_io.c:3105:static void check_buffer_tree_ref(struct extent_buffer *eb)
fs/btrfs/extent_io.c-3106-{
--
fs/btrfs/extent_io.c-3119- * called with creating new references, which is prone to race
fs/btrfs/extent_io.c:3120: * conditions between the calls to check_buffer_tree_ref in those
fs/btrfs/extent_io.c-3121- * codepaths and clearing TREE_REF in try_release_extent_buffer.
--
fs/btrfs/extent_io.c-3125- * reference are not. To protect against this class of races, we call
fs/btrfs/extent_io.c:3126: * check_buffer_tree_ref() from the code paths which trigger io. Note that
fs/btrfs/extent_io.c-3127- * once io is initiated, TREE_REF can no longer be cleared, so that is
--
fs/btrfs/extent_io.c=3140=static void mark_extent_buffer_accessed(struct extent_buffer *eb)
fs/btrfs/extent_io.c-3141-{
fs/btrfs/extent_io.c:3142: check_buffer_tree_ref(eb);
fs/btrfs/extent_io.c-3143-
--
fs/btrfs/extent_io.c=3177=struct extent_buffer *alloc_test_extent_buffer(struct btrfs_fs_info *fs_info,
--
fs/btrfs/extent_io.c-3211- xa_unlock_irq(&fs_info->buffer_tree);
fs/btrfs/extent_io.c:3212: check_buffer_tree_ref(eb);
fs/btrfs/extent_io.c-3213-
--
fs/btrfs/extent_io.c=3374=struct extent_buffer *alloc_extent_buffer(struct btrfs_fs_info *fs_info,
--
fs/btrfs/extent_io.c-3527- /* add one reference for the tree */
fs/btrfs/extent_io.c:3528: check_buffer_tree_ref(eb);
fs/btrfs/extent_io.c-3529-
--
fs/btrfs/extent_io.c=3738=void set_extent_buffer_dirty(struct extent_buffer *eb)
--
fs/btrfs/extent_io.c-3741-
fs/btrfs/extent_io.c:3742: check_buffer_tree_ref(eb);
fs/btrfs/extent_io.c-3743-
--
fs/btrfs/extent_io.c=3837=int read_extent_buffer_pages_nowait(struct extent_buffer *eb, int mirror_num,
--
fs/btrfs/extent_io.c-3869- eb->read_mirror = 0;
fs/btrfs/extent_io.c:3870: check_buffer_tree_ref(eb);
fs/btrfs/extent_io.c-3871- refcount_inc(&eb->refs);
--
fs/dlm/config.c=215=static int dlm_check_zero(unsigned int x)
--
fs/dlm/config.c-222-
fs/dlm/config.c:223:static int dlm_check_buffer_size(unsigned int x)
fs/dlm/config.c-224-{
--
fs/dlm/config.c-230-
fs/dlm/config.c:231:CLUSTER_ATTR(buffer_size, dlm_check_buffer_size);
fs/dlm/config.c-232-CLUSTER_ATTR(rsbtbl_size, dlm_check_zero);
--
fs/nilfs2/page.c=400=void nilfs_clear_folio_dirty(struct folio *folio)
--
fs/nilfs2/page.c-414-
fs/nilfs2/page.c:415:recheck_buffers:
fs/nilfs2/page.c-416- busy = false;
--
fs/nilfs2/page.c-429- invalidated = true;
fs/nilfs2/page.c:430: goto recheck_buffers;
fs/nilfs2/page.c-431- }
--
kernel/bpf/verifier.c=6771=static int get_callee_stack_depth(struct bpf_verifier_env *env,
--
kernel/bpf/verifier.c-6782-
kernel/bpf/verifier.c:6783:static int __check_buffer_access(struct bpf_verifier_env *env,
kernel/bpf/verifier.c-6784- const char *buf_info,
--
kernel/bpf/verifier.c=6807=static int check_tp_buffer_access(struct bpf_verifier_env *env,
--
kernel/bpf/verifier.c-6812-
kernel/bpf/verifier.c:6813: err = __check_buffer_access(env, "tracepoint", reg, regno, off, size);
kernel/bpf/verifier.c-6814- if (err)
--
kernel/bpf/verifier.c-6822-
kernel/bpf/verifier.c:6823:static int check_buffer_access(struct bpf_verifier_env *env,
kernel/bpf/verifier.c-6824- const struct bpf_reg_state *reg,
--
kernel/bpf/verifier.c-6831-
kernel/bpf/verifier.c:6832: err = __check_buffer_access(env, buf_info, reg, regno, off, size);
kernel/bpf/verifier.c-6833- if (err)
--
kernel/bpf/verifier.c=7528=static int check_mem_access(struct bpf_verifier_env *env, int insn_idx, u32 regno,
--
kernel/bpf/verifier.c-7769-
kernel/bpf/verifier.c:7770: err = check_buffer_access(env, reg, regno, off, size, false,
kernel/bpf/verifier.c-7771- max_access);
--
kernel/bpf/verifier.c=8177=static int check_helper_mem_access(struct bpf_verifier_env *env, int regno,
--
kernel/bpf/verifier.c-8225- }
kernel/bpf/verifier.c:8226: return check_buffer_access(env, reg, regno, reg->off,
kernel/bpf/verifier.c-8227- access_size, zero_size_allowed,
--
kernel/trace/ring_buffer.c=4354=static atomic_t ts_dump;
--
kernel/trace/ring_buffer.c-4376- */
kernel/trace/ring_buffer.c:4377:static void check_buffer(struct ring_buffer_per_cpu *cpu_buffer,
kernel/trace/ring_buffer.c-4378- struct rb_event_info *info,
--
kernel/trace/ring_buffer.c-4429-#else
kernel/trace/ring_buffer.c:4430:static inline void check_buffer(struct ring_buffer_per_cpu *cpu_buffer,
kernel/trace/ring_buffer.c-4431- struct rb_event_info *info,
--
kernel/trace/ring_buffer.c=4438=__rb_reserve_next(struct ring_buffer_per_cpu *cpu_buffer,
--
kernel/trace/ring_buffer.c-4488- if (unlikely(write > cpu_buffer->buffer->subbuf_size)) {
kernel/trace/ring_buffer.c:4489: check_buffer(cpu_buffer, info, CHECK_FULL_PAGE);
kernel/trace/ring_buffer.c-4490- return rb_move_tail(cpu_buffer, tail, info);
--
kernel/trace/ring_buffer.c-4508- info->delta = info->ts;
kernel/trace/ring_buffer.c:4509: check_buffer(cpu_buffer, info, tail);
kernel/trace/ring_buffer.c-4510- } else {
--
mm/migrate.c=933=static int __buffer_migrate_folio(struct address_space *mapping,
--
mm/migrate.c-958- VM_WARN_ON_ONCE(migrating);
mm/migrate.c:959:recheck_buffers:
mm/migrate.c-960- busy = false;
--
mm/migrate.c-977- invalidated = true;
mm/migrate.c:978: goto recheck_buffers;
mm/migrate.c-979- }
--
tools/testing/selftests/arm64/mte/.gitignore:1:check_buffer_fill
tools/testing/selftests/arm64/mte/.gitignore-2-check_gcr_el1_cswitch
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c=21=enum mte_block_test_alloc {
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-27-
tools/testing/selftests/arm64/mte/check_buffer_fill.c:28:static int check_buffer_by_byte(int mem_type, int mode)
tools/testing/selftests/arm64/mte/check_buffer_fill.c-29-{
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-62-
tools/testing/selftests/arm64/mte/check_buffer_fill.c:63:static int check_buffer_underflow_by_byte(int mem_type, int mode,
tools/testing/selftests/arm64/mte/check_buffer_fill.c-64- int underflow_range)
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-99- if (err)
tools/testing/selftests/arm64/mte/check_buffer_fill.c:100: goto check_buffer_underflow_by_byte_err;
tools/testing/selftests/arm64/mte/check_buffer_fill.c-101-
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-149- }
tools/testing/selftests/arm64/mte/check_buffer_fill.c:150:check_buffer_underflow_by_byte_err:
tools/testing/selftests/arm64/mte/check_buffer_fill.c-151- mte_free_memory_tag_range((void *)ptr, sizes[i], mem_type, underflow_range, 0);
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-157-
tools/testing/selftests/arm64/mte/check_buffer_fill.c:158:static int check_buffer_overflow_by_byte(int mem_type, int mode,
tools/testing/selftests/arm64/mte/check_buffer_fill.c-159- int overflow_range)
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-197- if (err)
tools/testing/selftests/arm64/mte/check_buffer_fill.c:198: goto check_buffer_overflow_by_byte_err;
tools/testing/selftests/arm64/mte/check_buffer_fill.c-199-
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-252- }
tools/testing/selftests/arm64/mte/check_buffer_fill.c:253:check_buffer_overflow_by_byte_err:
tools/testing/selftests/arm64/mte/check_buffer_fill.c-254- mte_free_memory_tag_range((void *)ptr, sizes[i], mem_type, 0, overflow_range);
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-260-
tools/testing/selftests/arm64/mte/check_buffer_fill.c:261:static int check_buffer_by_block_iterate(int mem_type, int mode, size_t size)
tools/testing/selftests/arm64/mte/check_buffer_fill.c-262-{
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-315- result = KSFT_FAIL;
tools/testing/selftests/arm64/mte/check_buffer_fill.c:316: goto check_buffer_by_block_err;
tools/testing/selftests/arm64/mte/check_buffer_fill.c-317- }
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-324- }
tools/testing/selftests/arm64/mte/check_buffer_fill.c:325:check_buffer_by_block_err:
tools/testing/selftests/arm64/mte/check_buffer_fill.c-326- mte_free_memory((void *)src, size, mem_type,
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-335-
tools/testing/selftests/arm64/mte/check_buffer_fill.c:336:static int check_buffer_by_block(int mem_type, int mode)
tools/testing/selftests/arm64/mte/check_buffer_fill.c-337-{
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-343- for (i = 0; i < item; i++) {
tools/testing/selftests/arm64/mte/check_buffer_fill.c:344: result = check_buffer_by_block_iterate(mem_type, mode, sizes[i]);
tools/testing/selftests/arm64/mte/check_buffer_fill.c-345- if (result != KSFT_PASS)
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c=403=int main(int argc, char *argv[])
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-423- /* Buffer by byte tests */
tools/testing/selftests/arm64/mte/check_buffer_fill.c:424: evaluate_test(check_buffer_by_byte(USE_MMAP, MTE_SYNC_ERR),
tools/testing/selftests/arm64/mte/check_buffer_fill.c-425- "Check buffer correctness by byte with sync err mode and mmap memory\n");
tools/testing/selftests/arm64/mte/check_buffer_fill.c:426: evaluate_test(check_buffer_by_byte(USE_MMAP, MTE_ASYNC_ERR),
tools/testing/selftests/arm64/mte/check_buffer_fill.c-427- "Check buffer correctness by byte with async err mode and mmap memory\n");
tools/testing/selftests/arm64/mte/check_buffer_fill.c:428: evaluate_test(check_buffer_by_byte(USE_MPROTECT, MTE_SYNC_ERR),
tools/testing/selftests/arm64/mte/check_buffer_fill.c-429- "Check buffer correctness by byte with sync err mode and mmap/mprotect memory\n");
tools/testing/selftests/arm64/mte/check_buffer_fill.c:430: evaluate_test(check_buffer_by_byte(USE_MPROTECT, MTE_ASYNC_ERR),
tools/testing/selftests/arm64/mte/check_buffer_fill.c-431- "Check buffer correctness by byte with async err mode and mmap/mprotect memory\n");
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-433- /* Check buffer underflow with underflow size as 16 */
tools/testing/selftests/arm64/mte/check_buffer_fill.c:434: evaluate_test(check_buffer_underflow_by_byte(USE_MMAP, MTE_SYNC_ERR, MT_GRANULE_SIZE),
tools/testing/selftests/arm64/mte/check_buffer_fill.c-435- "Check buffer write underflow by byte with sync mode and mmap memory\n");
tools/testing/selftests/arm64/mte/check_buffer_fill.c:436: evaluate_test(check_buffer_underflow_by_byte(USE_MMAP, MTE_ASYNC_ERR, MT_GRANULE_SIZE),
tools/testing/selftests/arm64/mte/check_buffer_fill.c-437- "Check buffer write underflow by byte with async mode and mmap memory\n");
tools/testing/selftests/arm64/mte/check_buffer_fill.c:438: evaluate_test(check_buffer_underflow_by_byte(USE_MMAP, MTE_NONE_ERR, MT_GRANULE_SIZE),
tools/testing/selftests/arm64/mte/check_buffer_fill.c-439- "Check buffer write underflow by byte with tag check fault ignore and mmap memory\n");
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-441- /* Check buffer underflow with underflow size as page size */
tools/testing/selftests/arm64/mte/check_buffer_fill.c:442: evaluate_test(check_buffer_underflow_by_byte(USE_MMAP, MTE_SYNC_ERR, page_size),
tools/testing/selftests/arm64/mte/check_buffer_fill.c-443- "Check buffer write underflow by byte with sync mode and mmap memory\n");
tools/testing/selftests/arm64/mte/check_buffer_fill.c:444: evaluate_test(check_buffer_underflow_by_byte(USE_MMAP, MTE_ASYNC_ERR, page_size),
tools/testing/selftests/arm64/mte/check_buffer_fill.c-445- "Check buffer write underflow by byte with async mode and mmap memory\n");
tools/testing/selftests/arm64/mte/check_buffer_fill.c:446: evaluate_test(check_buffer_underflow_by_byte(USE_MMAP, MTE_NONE_ERR, page_size),
tools/testing/selftests/arm64/mte/check_buffer_fill.c-447- "Check buffer write underflow by byte with tag check fault ignore and mmap memory\n");
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-449- /* Check buffer overflow with overflow size as 16 */
tools/testing/selftests/arm64/mte/check_buffer_fill.c:450: evaluate_test(check_buffer_overflow_by_byte(USE_MMAP, MTE_SYNC_ERR, MT_GRANULE_SIZE),
tools/testing/selftests/arm64/mte/check_buffer_fill.c-451- "Check buffer write overflow by byte with sync mode and mmap memory\n");
tools/testing/selftests/arm64/mte/check_buffer_fill.c:452: evaluate_test(check_buffer_overflow_by_byte(USE_MMAP, MTE_ASYNC_ERR, MT_GRANULE_SIZE),
tools/testing/selftests/arm64/mte/check_buffer_fill.c-453- "Check buffer write overflow by byte with async mode and mmap memory\n");
tools/testing/selftests/arm64/mte/check_buffer_fill.c:454: evaluate_test(check_buffer_overflow_by_byte(USE_MMAP, MTE_NONE_ERR, MT_GRANULE_SIZE),
tools/testing/selftests/arm64/mte/check_buffer_fill.c-455- "Check buffer write overflow by byte with tag fault ignore mode and mmap memory\n");
--
tools/testing/selftests/arm64/mte/check_buffer_fill.c-457- /* Buffer by block tests */
tools/testing/selftests/arm64/mte/check_buffer_fill.c:458: evaluate_test(check_buffer_by_block(USE_MMAP, MTE_SYNC_ERR),
tools/testing/selftests/arm64/mte/check_buffer_fill.c-459- "Check buffer write correctness by block with sync mode and mmap memory\n");
tools/testing/selftests/arm64/mte/check_buffer_fill.c:460: evaluate_test(check_buffer_by_block(USE_MMAP, MTE_ASYNC_ERR),
tools/testing/selftests/arm64/mte/check_buffer_fill.c-461- "Check buffer write correctness by block with async mode and mmap memory\n");
tools/testing/selftests/arm64/mte/check_buffer_fill.c:462: evaluate_test(check_buffer_by_block(USE_MMAP, MTE_NONE_ERR),
tools/testing/selftests/arm64/mte/check_buffer_fill.c-463- "Check buffer write correctness by block with tag fault ignore and mmap memory\n");
]
|
| 8/2 |
2026/05/25 22:12 |
llm |
expert |
0m
Model:
Tokens:
input: 26662
output: 41
thoughts: 74
Thoughts: Okay, here's my interpretation of those thoughts, considering my expertise:
**Diving into Ring Buffer Integrity**
Alright, so `check_buffer` isn't just some casual debugging nicety in `kernel/trace/ring_buffer.c`. I need to remember this because it's *critical*. I see it's called within `kernel/trace/ring_buffer.c`, in response to an `unlikely` condition related to the `write` size exceeding `cpu_buffer->buffer->subbuf_size`, or similar boundary errors within the ring buffer. That immediately tells me it's there to handle potentially corrupting writes β writes that are going to overflow the allocated space. It's a key defensive mechanism for maintaining the integrity of the trace data. Time to dig into the implementation of `check_buffer` itself; that'll provide insight into how the kernel handles these situations. I need to understand whether it's simply logging an error, attempting a recovery, or potentially triggering a panic. This is crucial for understanding the reliability and robustness of the tracing subsystem in this area.
|
| 9/2 |
2026/05/25 22:12 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/ring_buffer.c IncludeLines:true Name:check_buffer]
Results: map[SourceCode:4373: /*
4374: * Check if the current event time stamp matches the deltas on
4375: * the buffer page.
4376: */
4377: static void check_buffer(struct ring_buffer_per_cpu *cpu_buffer,
4378: struct rb_event_info *info,
4379: unsigned long tail)
4380: {
4381: struct buffer_data_page *bpage;
4382: u64 ts, delta;
4383: bool full = false;
4384: int ret;
4385:
4386: bpage = info->tail_page->page;
4387:
4388: if (tail == CHECK_FULL_PAGE) {
4389: full = true;
4390: tail = local_read(&bpage->commit);
4391: } else if (info->add_timestamp &
4392: (RB_ADD_STAMP_FORCE | RB_ADD_STAMP_ABSOLUTE)) {
4393: /* Ignore events with absolute time stamps */
4394: return;
4395: }
4396:
4397: /*
4398: * Do not check the first event (skip possible extends too).
4399: * Also do not check if previous events have not been committed.
4400: */
4401: if (tail <= 8 || tail > local_read(&bpage->commit))
4402: return;
4403:
4404: /*
4405: * If this interrupted another event,
4406: */
4407: if (atomic_inc_return(this_cpu_ptr(&checking)) != 1)
4408: goto out;
4409:
4410: ret = rb_read_data_buffer(bpage, tail, cpu_buffer->cpu, &ts, &delta);
4411: if (ret < 0) {
4412: if (delta < ts) {
4413: buffer_warn_return("[CPU: %d]ABSOLUTE TIME WENT BACKWARDS: last ts: %lld absolute ts: %lld\n",
4414: cpu_buffer->cpu, ts, delta);
4415: goto out;
4416: }
4417: }
4418: if ((full && ts > info->ts) ||
4419: (!full && ts + info->delta != info->ts)) {
4420: buffer_warn_return("[CPU: %d]TIME DOES NOT MATCH expected:%lld actual:%lld delta:%lld before:%lld after:%lld%s context:%s\n",
4421: cpu_buffer->cpu,
4422: ts + info->delta, info->ts, info->delta,
4423: info->before, info->after,
4424: full ? " (full)" : "", show_interrupt_level());
4425: }
4426: out:
4427: atomic_dec(this_cpu_ptr(&checking));
4428: }
SourceFile:kernel/trace/ring_buffer.c]
|
| 10/2 |
2026/05/25 22:12 |
llm |
expert |
0m
Model:
Tokens:
input: 27888
output: 45
thoughts: 176
Thoughts: **Analyzing the Ring Buffer Reading Path**
Okay, so I'm looking at this stack trace and immediately, my mind hones in on `rb_event_length`. That function's being called by `rb_read_data_buffer`, which in turn is called by `check_buffer`. Right, let's just make sure I have this straight. The execution flow is `check_buffer` -> `rb_read_data_buffer` -> `rb_event_length`. Got it.
Now, that stack trace⦠let's see. The kernel addresses look reasonable. I should take a look at the code. I need to understand what's happening. Based on the trace, I need to check `rb_read_data_buffer`, specifically. Let's see what's going on there.
|
| 11/2 |
2026/05/25 22:12 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/ring_buffer.c IncludeLines:true Name:rb_read_data_buffer]
Results: map[SourceCode:1845: static int rb_read_data_buffer(struct buffer_data_page *dpage, int tail, int cpu,
1846: unsigned long long *timestamp, u64 *delta_ptr)
1847: {
1848: struct ring_buffer_event *event;
1849: u64 ts, delta;
1850: int events = 0;
1851: int e;
1852:
1853: *delta_ptr = 0;
1854: *timestamp = 0;
1855:
1856: ts = dpage->time_stamp;
1857:
1858: for (e = 0; e < tail; e += rb_event_length(event)) {
1859:
1860: event = (struct ring_buffer_event *)(dpage->data + e);
1861:
1862: switch (event->type_len) {
1863:
1864: case RINGBUF_TYPE_TIME_EXTEND:
1865: delta = rb_event_time_stamp(event);
1866: ts += delta;
1867: break;
1868:
1869: case RINGBUF_TYPE_TIME_STAMP:
1870: delta = rb_event_time_stamp(event);
1871: delta = rb_fix_abs_ts(delta, ts);
1872: if (delta < ts) {
1873: *delta_ptr = delta;
1874: *timestamp = ts;
1875: return -1;
1876: }
1877: ts = delta;
1878: break;
1879:
1880: case RINGBUF_TYPE_PADDING:
1881: if (event->time_delta == 1)
1882: break;
1883: fallthrough;
1884: case RINGBUF_TYPE_DATA:
1885: events++;
1886: ts += event->time_delta;
1887: break;
1888:
1889: default:
1890: return -1;
1891: }
1892: }
1893: *timestamp = ts;
1894: return events;
1895: }
SourceFile:kernel/trace/ring_buffer.c]
|
| 12/2 |
2026/05/25 22:12 |
llm |
expert |
0m
Model:
Tokens:
input: 28967
output: 43
thoughts: 82
Thoughts: **Infinite Loop Concern in `rb_read_data_buffer`**
Okay, so I'm looking at this `rb_read_data_buffer` function and I've got a red flag going up immediately. This loop here:
```c
for (e = 0; e < tail; e += rb_event_length(event)) {
```
It's iterating through the events in the buffer, incrementing the counter `e` by the length of the current event returned by `rb_event_length(event)`. The problem is obvious: if `rb_event_length` ever returns 0, this thing is going to loop forever! It's a textbook infinite loop scenario. We're adding zero to the counter and will never hit the `e < tail` condition. Time to go dig into `rb_event_length` and see what's going on in there and what might cause it to return zero. I need to understand the conditions under which that could happen. This is a potential crash waiting to happen.
|
| 13/2 |
2026/05/25 22:12 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/ring_buffer.c IncludeLines:true Name:rb_event_length]
Results: map[SourceCode: 214: /*
215: * Return the length of the given event. Will return
216: * the length of the time extend if the event is a
217: * time extend.
218: */
219: static inline unsigned
220: rb_event_length(struct ring_buffer_event *event)
221: {
222: switch (event->type_len) {
223: case RINGBUF_TYPE_PADDING:
224: if (rb_null_event(event))
225: /* undefined */
226: return -1;
227: return event->array[0] + RB_EVNT_HDR_SIZE;
228:
229: case RINGBUF_TYPE_TIME_EXTEND:
230: return RB_LEN_TIME_EXTEND;
231:
232: case RINGBUF_TYPE_TIME_STAMP:
233: return RB_LEN_TIME_STAMP;
234:
235: case RINGBUF_TYPE_DATA:
236: return rb_event_data_length(event);
237: default:
238: WARN_ON_ONCE(1);
239: }
240: /* not hit */
241: return 0;
242: }
SourceFile:kernel/trace/ring_buffer.c]
|
| 14/2 |
2026/05/25 22:12 |
llm |
expert |
0m
Model:
Tokens:
input: 29736
output: 45
thoughts: 16
Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience:
**Potential Zero-Length Data - A Quick Recalibration**
Right, so I'm thinking about this data buffer, and specifically the `rb_event_data_length` call. A tiny, nagging thought just popped into my head... I need to double-check something. Hang on... could that function *actually* return 0? I've been assuming it wouldn't. I've probably got some downstream code that isn't handling that scenario properly if it does. That assumption could be a real problem. I need to quickly review the documentation for that function to confirm, or, even better, check the actual implementation to be sure of the edge cases. Gotta anticipate all those boundary conditions or I am going to have a *really* bad day later. And if it can return zero, I need to make some immediate adjustments and test those changes as soon as possible. No time to lose on this one; this seemingly small detail could cause some significant head-scratching debugging. Let's see... back to the source...
|
| 15/2 |
2026/05/25 22:12 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/ring_buffer.c IncludeLines:true Name:rb_event_data_length]
Results: map[SourceCode: 202: static unsigned
203: rb_event_data_length(struct ring_buffer_event *event)
204: {
205: unsigned length;
206:
207: if (event->type_len)
208: length = event->type_len * RB_ALIGNMENT;
209: else
210: length = event->array[0];
211: return length + RB_EVNT_HDR_SIZE;
212: }
SourceFile:kernel/trace/ring_buffer.c]
|
| 16/2 |
2026/05/25 22:12 |
llm |
expert |
0m
Model:
Tokens:
input: 30206
output: 43
thoughts: 312
Thoughts: **Analyzing Potential Integer Overflow and Infinite Loop Vulnerabilities**
Okay, so I'm looking at this code and immediately I see that `RB_EVNT_HDR_SIZE` is likely non-zero, which means `rb_event_data_length` will always be at least that size. But hang on, what happens if `rb_event_length` returns 0? That's the real question.
Alright, if we're dealing with a `RINGBUF_TYPE_PADDING` event and `rb_null_event(event)` *isn't* true, then we return `event->array[0] + RB_EVNT_HDR_SIZE`. Now I'm thinking, if `event->array[0]` is a large value, could we have an overflow situation there? It seems possible, but needs further investigation.
But wait, looking at `rb_read_data_buffer`, I see a loop:
```c
for (e = 0; e < tail; e += rb_event_length(event)) {
```
If `rb_event_length(event)` *does* return 0, we're in trouble because that's an infinite loop. Okay, that's a big red flag. However, hold on. I need to be careful, it's returning an `unsigned` value, which affects the behaviour of the addition.
Let's see: If `event->type_len` *is* `RINGBUF_TYPE_PADDING` and `rb_null_event(event)` *is* true, `rb_event_length` returns -1. But, as an `unsigned` value, -1 becomes `0xffffffff`. Okay, that's interesting.
So now we have `e += 0xffffffff`. But, `e` is an `int`. This means it will evaluate to `e -= 1`. So, we're actually *decreasing* the loop counter! That's a huge problem.
Now, if `e` decreases, `e < tail` will remain true, and this is a classic infinite loop scenario. This needs to be addressed ASAP. Now I need to check how `rb_null_event` is implemented.
|
| 17/2 |
2026/05/25 22:12 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/ring_buffer.c IncludeLines:true Name:rb_null_event]
Results: map[SourceCode: 190: static inline bool rb_null_event(struct ring_buffer_event *event)
191: {
192: return event->type_len == RINGBUF_TYPE_PADDING && !event->time_delta;
193: }
SourceFile:kernel/trace/ring_buffer.c]
|
| 18/2 |
2026/05/25 22:12 |
llm |
expert |
0m
Model:
Tokens:
input: 31140
output: 41
thoughts: 1136
Thoughts: **Debugging a Ring Buffer Stall**
Okay, so I'm looking at this ring buffer code, trying to figure out why it's hanging or stalling. Let's trace this through:
Initially, I was worried about potential out-of-bounds reads. I was concerned that the `e += rb_event_length(event)` might cause `e` to decrement. And I figured, if `e` goes negative, it will access `dpage->data + e` out of bounds. But I just needed to remember how a `for` loop actually works. The increment happens *after* the loop body! So that's not it. `event` *is* initialized on the first iteration, and `e` always accurately reflects the current event's offset within the data page. The loop correctly increments by the length of each event, so that specific bug is a non-issue.
So, where's the bottleneck? I noticed the reproducer attaches the BPF program 65 times to `rcu_utilization`. That seems... excessive. Thinking about it, `rcu_utilization` is called from `rcu_sched_clock_irq`, which is a timer interrupt. So, if the BPF program is attached 65 times, it gets executed 65 times *per timer interrupt*. Each execution has a call to `bpf_trace_printk`, which writes to the trace buffer. Is that writing to the trace buffer taking too long? Is that causing an RCU stall? No, `bpf_trace_printk` is not rate-limited. So that's likely not it.
Then I saw the stack trace mentioning `check_buffer`. That's interesting. I know `check_buffer` is called when `write > cpu_buffer->buffer->subbuf_size`. However, this indicates that the buffer page is full, and it needs to move to the next page. I can't just move on from that.
So, I re-examined the code, and realized that `check_buffer` is called *every* event reservation if `likely(tail == info->tail)`! It's *not* just a page-full check! That's it! It's *always* being called! Given how frequently it's called, it must be the real bottleneck!
I now need to find out why that call is so expensive.
|
| 19/2 |
2026/05/25 22:12 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/ring_buffer.c IncludeLines:true Name:check_buffer]
Results: map[SourceCode:4373: /*
4374: * Check if the current event time stamp matches the deltas on
4375: * the buffer page.
4376: */
4377: static void check_buffer(struct ring_buffer_per_cpu *cpu_buffer,
4378: struct rb_event_info *info,
4379: unsigned long tail)
4380: {
4381: struct buffer_data_page *bpage;
4382: u64 ts, delta;
4383: bool full = false;
4384: int ret;
4385:
4386: bpage = info->tail_page->page;
4387:
4388: if (tail == CHECK_FULL_PAGE) {
4389: full = true;
4390: tail = local_read(&bpage->commit);
4391: } else if (info->add_timestamp &
4392: (RB_ADD_STAMP_FORCE | RB_ADD_STAMP_ABSOLUTE)) {
4393: /* Ignore events with absolute time stamps */
4394: return;
4395: }
4396:
4397: /*
4398: * Do not check the first event (skip possible extends too).
4399: * Also do not check if previous events have not been committed.
4400: */
4401: if (tail <= 8 || tail > local_read(&bpage->commit))
4402: return;
4403:
4404: /*
4405: * If this interrupted another event,
4406: */
4407: if (atomic_inc_return(this_cpu_ptr(&checking)) != 1)
4408: goto out;
4409:
4410: ret = rb_read_data_buffer(bpage, tail, cpu_buffer->cpu, &ts, &delta);
4411: if (ret < 0) {
4412: if (delta < ts) {
4413: buffer_warn_return("[CPU: %d]ABSOLUTE TIME WENT BACKWARDS: last ts: %lld absolute ts: %lld\n",
4414: cpu_buffer->cpu, ts, delta);
4415: goto out;
4416: }
4417: }
4418: if ((full && ts > info->ts) ||
4419: (!full && ts + info->delta != info->ts)) {
4420: buffer_warn_return("[CPU: %d]TIME DOES NOT MATCH expected:%lld actual:%lld delta:%lld before:%lld after:%lld%s context:%s\n",
4421: cpu_buffer->cpu,
4422: ts + info->delta, info->ts, info->delta,
4423: info->before, info->after,
4424: full ? " (full)" : "", show_interrupt_level());
4425: }
4426: out:
4427: atomic_dec(this_cpu_ptr(&checking));
4428: }
SourceFile:kernel/trace/ring_buffer.c]
|
| 20/2 |
2026/05/25 22:12 |
llm |
expert |
0m
Model:
Tokens:
input: 33708
output: 27
thoughts: 51
Thoughts: Here's my take on the thought process:
**Diving Deep into Ring Buffer Validation**
Okay, hold on a second. I just realized that this `check_buffer` function is conditionally compiled, wrapped inside a `#ifdef CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS` directive. That immediately tells me this validation logic is tied to a specific configuration option. I need to understand when and why it's enabled. To figure this out, I should jump directly to `kernel/trace/ring_buffer.c`, focusing around line 4370. That should give me the context I need around this conditional compilation and what this time delta validation is all about. Hopefully, this will lead me to what this functionality offers and whether it is valuable in my application.
|
| 21/2 |
2026/05/25 22:12 |
tool |
grepper |
0m
Args: map[Expression:CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS]
Results: map[Output:kernel/trace/ring_buffer.c=4226=EXPORT_SYMBOL_GPL(ring_buffer_unlock_commit);
--
kernel/trace/ring_buffer.c-4230-
kernel/trace/ring_buffer.c:4231:#ifdef CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS
kernel/trace/ring_buffer.c-4232-
--
kernel/trace/ring_buffer.c=4430=static inline void check_buffer(struct ring_buffer_per_cpu *cpu_buffer,
--
kernel/trace/ring_buffer.c-4434-}
kernel/trace/ring_buffer.c:4435:#endif /* CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS */
kernel/trace/ring_buffer.c-4436-
]
|
| 22/2 |
2026/05/25 22:12 |
llm |
expert |
0m
Model:
Tokens:
input: 34161
output: 26
thoughts: 213
Thoughts: **Analyzing a Potential Performance Bottleneck in the Ring Buffer**
Okay, so I see that `CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS` is enabled in this build. I'm immediately concerned because that means `check_buffer` gets called every single time an event is reserved, doesn't it? And, if I recall correctly, `check_buffer` iterates through *every* event in the current page up to the `tail` pointer. Yikes. That's a classic O(N^2) scenario *per page*! If a page contains *N* events, adding the *i*-th event takes O(*i*) time. That means filling an entire page is going to take O(N^2) time overall, assuming the buffer is page aligned.
If we're seeing frequent `bpf_trace_printk` calls β say, 65 times per timer interrupt as I'm thinking β and if those timer interrupts are themselves frequent, we could be spending a significant amount of CPU cycles in `check_buffer`. This is concerning. We'd potentially have a performance cliff.
The big question now is, *is* `CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS` enabled in the production builds? I need to check the Kconfig files, probably in `lib/Kconfig.debug` or `kernel/trace/Kconfig`, to verify the configuration. I can't be sure of the impact without knowing if this is something active in the deployments.
|
| 23/2 |
2026/05/25 22:13 |
tool |
grepper |
0m
Args: map[Expression:config RING_BUFFER_VALIDATE_TIME_DELTAS]
Results: map[Output:kernel/trace/Kconfig=1159=config RING_BUFFER_STARTUP_TEST
--
kernel/trace/Kconfig-1181-
kernel/trace/Kconfig:1182:config RING_BUFFER_VALIDATE_TIME_DELTAS
kernel/trace/Kconfig-1183- bool "Verify ring buffer time stamp deltas"
]
|
| 24/2 |
2026/05/25 22:13 |
llm |
expert |
0m
Model:
Tokens:
input: 34820
output: 46
thoughts: 21
Thoughts: Here's my attempt at summarizing the thought process, tailored for an expert audience:
**Analyzing `RING_BUFFER_VALIDATE_TIME_DELTAS` Definition**
Okay, so I'm looking at `RING_BUFFER_VALIDATE_TIME_DELTAS`. First things first: let's get the *precise* definition. I need to understand this fully, its intent, and its implications before I can even begin to think about how it might be used or improved. Is it a compile-time flag? A macro? A function? Knowing the *form* is the foundational element to grasping its function.
I want to know *why* this validation is happening. Is it for debugging? Performance optimization? Security? Perhaps it is a critical safeguard against some specific ring buffer corruption scenario. Understanding the *motivation* behind this delta validation is just as important as the precise mechanics of its implementation. I'm going to dig into the surrounding code to determine the use case.
I'm thinking about the trade-offs now. Time delta validation inherently adds overhead. So, what is the *cost* of this validation in terms of execution time, especially given the context of a ring buffer, which often appears in performance-critical applications. On the other hand, the *benefit* is data integrity. This could be crucial if the buffer is handling real-time data or in a system where data loss could be catastrophic. I need to balance these trade-offs.
I will consider the implementation. What specific checks are being performed on these time deltas? Are there thresholds or ranges? Are they detecting monotonic-ity violations? Are there error handling or logging mechanisms if the validation fails? And if it's configurable, what are the different modes, and what are the defaults? A careful examination is needed before I form my conclusions.
|
| 25/2 |
2026/05/25 22:13 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:kernel/trace/Kconfig IncludeLines:true Name:RING_BUFFER_VALIDATE_TIME_DELTAS]
Results: map[SourceCode: SourceFile:]
|
| 26/2 |
2026/05/25 22:13 |
llm |
expert |
0m
Model:
Tokens:
input: 35286
output: 21
thoughts: 20
Thoughts: Here's my attempt at summarizing the thought process, tailored for an expert audience:
**Kernel Trace Configuration Exploration**
Alright, let's dive in. My immediate focus is on understanding the tracing configuration options. I'm going to start by examining `kernel/trace/Kconfig`, specifically around line 1182. I'm expecting to find the top-level options for enabling/disabling the tracing infrastructure itself, along with any core, fundamental settings that govern its behavior. This should give me a good lay of the land, the base dependencies, and overall framework before I try to understand the many specific features of tracepoints, probes, and events. Knowing that file and that location should allow me to formulate better questions or to build mental models. I anticipate that this is where I'll find the gatekeeper options that are prerequisites for enabling the more granular aspects of tracing. Hopefully it's well-commented.
|
| 27/2 |
2026/05/25 22:13 |
tool |
grepper |
0m
Args: map[Expression:Verify ring buffer time stamp deltas]
Results: map[Output:kernel/trace/Kconfig=1182=config RING_BUFFER_VALIDATE_TIME_DELTAS
kernel/trace/Kconfig:1183: bool "Verify ring buffer time stamp deltas"
kernel/trace/Kconfig-1184- depends on RING_BUFFER
]
|
| 28/2 |
2026/05/25 22:13 |
llm |
expert |
0m
Model:
Tokens:
input: 35600
output: 31
thoughts: 24
Thoughts: Here's my take on the kernel trace configuration, specifically lines 1182-1190 of `Kconfig`:
**Analyzing Kernel Trace Configuration Options**
Okay, so I need to understand this section of the kernel tracing configuration. Let's get right into it, I'm diving into `kernel/trace/Kconfig` and focusing on lines 1182-1190. I know these Kconfig files are the heart of configuring the kernel's features, and tracing is crucial for performance analysis and debugging. I'll read through these lines carefully, paying close attention to the dependencies, descriptions, and defaults of the options defined there. This should clarify which tracing features are available and how they're intended to be used within the larger kernel framework. I'm looking for the key configuration variables and their relationships to other parts of the system. I anticipate these lines will give me insight into the granularity and scope of the tracing mechanisms provided.
|
| 29/2 |
2026/05/25 22:13 |
tool |
grepper |
0m
Error:
Args: map[Expression:config RING_BUFFER_VALIDATE_TIME_DELTAS -A 10]
Results:
|
| 30/2 |
2026/05/25 22:13 |
llm |
expert |
0m
Model:
Tokens:
input: 35872
output: 21
thoughts: 18
Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience:
**A Quick Realization about File Inspection**
Hang on a minute... `grepper` is built on top of `git grep`, isn't it? That means it's fundamentally searching within the Git repository, probably just indexing the files. If that's the case, then I don't need to involve it at all, or run the command; I could just go straight to the source. I can bypass the whole `grepper` process. I could just use a tool like ripgrep or even just `cat` or my editor to look directly at the file system! It's way faster and more direct. I can optimize my workflow and, more importantly, get a better grasp on exactly what data it's extracting. This is an unnecessary layer of indirection!
|
| 31/2 |
2026/05/25 22:13 |
tool |
read-file |
0m
Args: map[File:kernel/trace/Kconfig]
Results: map[Contents:# SPDX-License-Identifier: GPL-2.0-only
#
# Architectures that offer an FUNCTION_TRACER implementation should
# select HAVE_FUNCTION_TRACER:
#
config USER_STACKTRACE_SUPPORT
bool
config NOP_TRACER
bool
config HAVE_RETHOOK
bool
config RETHOOK
bool
depends on HAVE_RETHOOK
help
Enable generic return hooking feature. This is an internal
API, which will be used by other function-entry hooking
features like fprobe and kprobes.
config HAVE_FUNCTION_TRACER
bool
help
See Documentation/trace/ftrace-design.rst
config HAVE_FUNCTION_GRAPH_TRACER
bool
help
See Documentation/trace/ftrace-design.rst
config HAVE_FUNCTION_GRAPH_FREGS
bool
config HAVE_FTRACE_GRAPH_FUNC
bool
help
True if ftrace_graph_func() is defined.
config HAVE_DYNAMIC_FTRACE
bool
help
See Documentation/trace/ftrace-design.rst
config HAVE_DYNAMIC_FTRACE_WITH_REGS
bool
config HAVE_DYNAMIC_FTRACE_WITH_DIRECT_CALLS
bool
config HAVE_DYNAMIC_FTRACE_WITH_CALL_OPS
bool
config HAVE_EXTRA_IPI_TRACEPOINTS
bool
help
For architectures that use ipi_raise, ipi_entry and ipi_exit
tracepoints.
config HAVE_DYNAMIC_FTRACE_WITH_ARGS
bool
help
If this is set, then arguments and stack can be found from
the ftrace_regs passed into the function callback regs parameter
by default, even without setting the REGS flag in the ftrace_ops.
This allows for use of ftrace_regs_get_argument() and
ftrace_regs_get_stack_pointer().
config HAVE_FTRACE_REGS_HAVING_PT_REGS
bool
help
If this is set, ftrace_regs has pt_regs, thus it can convert to
pt_regs without allocating memory.
config HAVE_DYNAMIC_FTRACE_NO_PATCHABLE
bool
help
If the architecture generates __patchable_function_entries sections
but does not want them included in the ftrace locations.
config HAVE_DYNAMIC_FTRACE_WITH_JMP
bool
help
If the architecture supports to replace the __fentry__ with a
"jmp" instruction.
config HAVE_SYSCALL_TRACEPOINTS
bool
help
See Documentation/trace/ftrace-design.rst
config HAVE_FENTRY
bool
help
Arch supports the gcc options -pg with -mfentry
config HAVE_NOP_MCOUNT
bool
help
Arch supports the gcc options -pg with -mrecord-mcount and -nop-mcount
config HAVE_OBJTOOL_MCOUNT
bool
help
Arch supports objtool --mcount
config HAVE_OBJTOOL_NOP_MCOUNT
bool
help
Arch supports the objtool options --mcount with --mnop.
An architecture can select this if it wants to enable nop'ing
of ftrace locations.
config HAVE_C_RECORDMCOUNT
bool
help
C version of recordmcount available?
config HAVE_BUILDTIME_MCOUNT_SORT
bool
help
An architecture selects this if it sorts the mcount_loc section
at build time.
config BUILDTIME_MCOUNT_SORT
bool
default y
depends on HAVE_BUILDTIME_MCOUNT_SORT && DYNAMIC_FTRACE
help
Sort the mcount_loc section at build time.
config TRACER_MAX_TRACE
bool
config TRACE_CLOCK
bool
config RING_BUFFER
bool
select TRACE_CLOCK
select IRQ_WORK
config EVENT_TRACING
select CONTEXT_SWITCH_TRACER
select GLOB
bool
config CONTEXT_SWITCH_TRACER
bool
config RING_BUFFER_ALLOW_SWAP
bool
help
Allow the use of ring_buffer_swap_cpu.
Adds a very slight overhead to tracing when enabled.
config PREEMPTIRQ_TRACEPOINTS
bool
depends on TRACE_PREEMPT_TOGGLE || TRACE_IRQFLAGS
select TRACING
default y
help
Create preempt/irq toggle tracepoints if needed, so that other parts
of the kernel can use them to generate or add hooks to them.
# All tracer options should select GENERIC_TRACER. For those options that are
# enabled by all tracers (context switch and event tracer) they select TRACING.
# This allows those options to appear when no other tracer is selected. But the
# options do not appear when something else selects it. We need the two options
# GENERIC_TRACER and TRACING to avoid circular dependencies to accomplish the
# hiding of the automatic options.
config TRACING
bool
select RING_BUFFER
select STACKTRACE if STACKTRACE_SUPPORT
select TRACEPOINTS
select NOP_TRACER
select BINARY_PRINTF
select EVENT_TRACING
select TRACE_CLOCK
select NEED_TASKS_RCU
config GENERIC_TRACER
bool
select TRACING
#
# Minimum requirements an architecture has to meet for us to
# be able to offer generic tracing facilities:
#
config TRACING_SUPPORT
bool
depends on TRACE_IRQFLAGS_SUPPORT
depends on STACKTRACE_SUPPORT
default y
menuconfig FTRACE
bool "Tracers"
depends on TRACING_SUPPORT
default y if DEBUG_KERNEL
help
Enable the kernel tracing infrastructure.
if FTRACE
config TRACEFS_AUTOMOUNT_DEPRECATED
bool "Automount tracefs on debugfs [DEPRECATED]"
depends on TRACING
default y
help
The tracing interface was moved from /sys/kernel/debug/tracing
to /sys/kernel/tracing in 2015, but the tracing file system
was still automounted in /sys/kernel/debug for backward
compatibility with tooling.
The new interface has been around for more than 10 years and
the old debug mount will soon be removed.
config BOOTTIME_TRACING
bool "Boot-time Tracing support"
depends on TRACING
select BOOT_CONFIG
help
Enable developer to setup ftrace subsystem via supplemental
kernel cmdline at boot time for debugging (tracing) driver
initialization and boot process.
config FUNCTION_TRACER
bool "Kernel Function Tracer"
depends on HAVE_FUNCTION_TRACER
select KALLSYMS
select GENERIC_TRACER
select CONTEXT_SWITCH_TRACER
select GLOB
select NEED_TASKS_RCU
select TASKS_RUDE_RCU
help
Enable the kernel to trace every kernel function. This is done
by using a compiler feature to insert a small, 5-byte No-Operation
instruction at the beginning of every kernel function, which NOP
sequence is then dynamically patched into a tracer call when
tracing is enabled by the administrator. If it's runtime disabled
(the bootup default), then the overhead of the instructions is very
small and not measurable even in micro-benchmarks (at least on
x86, but may have impact on other architectures).
config FUNCTION_GRAPH_TRACER
bool "Kernel Function Graph Tracer"
depends on HAVE_FUNCTION_GRAPH_TRACER
depends on FUNCTION_TRACER
depends on !X86_32 || !CC_OPTIMIZE_FOR_SIZE
default y
help
Enable the kernel to trace a function at both its return
and its entry.
Its first purpose is to trace the duration of functions and
draw a call graph for each thread with some information like
the return value. This is done by setting the current return
address on the current task structure into a stack of calls.
config FUNCTION_GRAPH_RETVAL
bool "Kernel Function Graph Return Value"
depends on HAVE_FUNCTION_GRAPH_FREGS
depends on FUNCTION_GRAPH_TRACER
default n
help
Support recording and printing the function return value when
using function graph tracer. It can be helpful to locate functions
that return errors. This feature is off by default, and you can
enable it via the trace option funcgraph-retval.
See Documentation/trace/ftrace.rst
config FUNCTION_GRAPH_RETADDR
bool "Kernel Function Graph Return Address"
depends on FUNCTION_GRAPH_TRACER
default n
help
Support recording and printing the function return address when
using function graph tracer. It can be helpful to locate code line that
the function is called. This feature is off by default, and you can
enable it via the trace option funcgraph-retaddr.
config FUNCTION_TRACE_ARGS
bool
depends on PROBE_EVENTS_BTF_ARGS
default y
help
If supported with function argument access API and BTF, then
the function tracer and function graph tracer will support printing
of function arguments. This feature is off by default, and can be
enabled via the trace option func-args (for the function tracer) and
funcgraph-args (for the function graph tracer)
config DYNAMIC_FTRACE
bool
depends on FUNCTION_TRACER
depends on HAVE_DYNAMIC_FTRACE
default y
help
This option will modify all the calls to function tracing
dynamically (will patch them out of the binary image and
replace them with a No-Op instruction) on boot up. During
compile time, a table is made of all the locations that ftrace
can function trace, and this table is linked into the kernel
image. When this is enabled, functions can be individually
enabled, and the functions not enabled will not affect
performance of the system.
See the files in /sys/kernel/tracing:
available_filter_functions
set_ftrace_filter
set_ftrace_notrace
This way a CONFIG_FUNCTION_TRACER kernel is slightly larger, but
otherwise has native performance as long as no tracing is active.
config DYNAMIC_FTRACE_WITH_REGS
def_bool y
depends on DYNAMIC_FTRACE
depends on HAVE_DYNAMIC_FTRACE_WITH_REGS
config DYNAMIC_FTRACE_WITH_DIRECT_CALLS
def_bool y
depends on DYNAMIC_FTRACE_WITH_REGS || DYNAMIC_FTRACE_WITH_ARGS
depends on HAVE_DYNAMIC_FTRACE_WITH_DIRECT_CALLS
config DYNAMIC_FTRACE_WITH_CALL_OPS
def_bool y
depends on HAVE_DYNAMIC_FTRACE_WITH_CALL_OPS
config DYNAMIC_FTRACE_WITH_ARGS
def_bool y
depends on DYNAMIC_FTRACE
depends on HAVE_DYNAMIC_FTRACE_WITH_ARGS
config DYNAMIC_FTRACE_WITH_JMP
def_bool y
depends on DYNAMIC_FTRACE
depends on DYNAMIC_FTRACE_WITH_DIRECT_CALLS
depends on HAVE_DYNAMIC_FTRACE_WITH_JMP
config FUNCTION_SELF_TRACING
bool "Function trace tracing code"
depends on FUNCTION_TRACER
help
Normally all the tracing code is set to notrace, where the function
tracer will ignore all the tracing functions. Sometimes it is useful
for debugging to trace some of the tracing infratructure itself.
Enable this to allow some of the tracing infrastructure to be traced
by the function tracer. Note, this will likely add noise to function
tracing if events and other tracing features are enabled along with
function tracing.
If unsure, say N.
config FPROBE
bool "Kernel Function Probe (fprobe)"
depends on HAVE_FUNCTION_GRAPH_FREGS && HAVE_FTRACE_GRAPH_FUNC
depends on DYNAMIC_FTRACE_WITH_ARGS
select FUNCTION_GRAPH_TRACER
default n
help
This option enables kernel function probe (fprobe) based on ftrace.
The fprobe is similar to kprobes, but probes only for kernel function
entries and exits. This also can probe multiple functions by one
fprobe.
If unsure, say N.
config FUNCTION_PROFILER
bool "Kernel function profiler"
depends on FUNCTION_TRACER
default n
help
This option enables the kernel function profiler. A file is created
in debugfs called function_profile_enabled which defaults to zero.
When a 1 is echoed into this file profiling begins, and when a
zero is entered, profiling stops. A "functions" file is created in
the trace_stat directory; this file shows the list of functions that
have been hit and their counters.
If in doubt, say N.
config STACK_TRACER
bool "Trace max stack"
depends on HAVE_FUNCTION_TRACER
select FUNCTION_TRACER
select STACKTRACE
select KALLSYMS
help
This special tracer records the maximum stack footprint of the
kernel and displays it in /sys/kernel/tracing/stack_trace.
This tracer works by hooking into every function call that the
kernel executes, and keeping a maximum stack depth value and
stack-trace saved. If this is configured with DYNAMIC_FTRACE
then it will not have any overhead while the stack tracer
is disabled.
To enable the stack tracer on bootup, pass in 'stacktrace'
on the kernel command line.
The stack tracer can also be enabled or disabled via the
sysctl kernel.stack_tracer_enabled
Say N if unsure.
config TRACE_PREEMPT_TOGGLE
bool
help
Enables hooks which will be called when preemption is first disabled,
and last enabled.
config IRQSOFF_TRACER
bool "Interrupts-off Latency Tracer"
default n
depends on TRACE_IRQFLAGS_SUPPORT
select TRACE_IRQFLAGS
select GENERIC_TRACER
select TRACER_MAX_TRACE
select RING_BUFFER_ALLOW_SWAP
select TRACER_SNAPSHOT
select TRACER_SNAPSHOT_PER_CPU_SWAP
help
This option measures the time spent in irqs-off critical
sections, with microsecond accuracy.
The default measurement method is a maximum search, which is
disabled by default and can be runtime (re-)started
via:
echo 0 > /sys/kernel/tracing/tracing_max_latency
(Note that kernel size and overhead increase with this option
enabled. This option and the preempt-off timing option can be
used together or separately.)
config PREEMPT_TRACER
bool "Preemption-off Latency Tracer"
default n
depends on PREEMPTION
select GENERIC_TRACER
select TRACER_MAX_TRACE
select RING_BUFFER_ALLOW_SWAP
select TRACER_SNAPSHOT
select TRACER_SNAPSHOT_PER_CPU_SWAP
select TRACE_PREEMPT_TOGGLE
help
This option measures the time spent in preemption-off critical
sections, with microsecond accuracy.
The default measurement method is a maximum search, which is
disabled by default and can be runtime (re-)started
via:
echo 0 > /sys/kernel/tracing/tracing_max_latency
(Note that kernel size and overhead increase with this option
enabled. This option and the irqs-off timing option can be
used together or separately.)
config SCHED_TRACER
bool "Scheduling Latency Tracer"
select GENERIC_TRACER
select CONTEXT_SWITCH_TRACER
select TRACER_MAX_TRACE
select TRACER_SNAPSHOT
help
This tracer tracks the latency of the highest priority task
to be scheduled in, starting from the point it has woken up.
config HWLAT_TRACER
bool "Tracer to detect hardware latencies (like SMIs)"
select GENERIC_TRACER
select TRACER_MAX_TRACE
help
This tracer, when enabled will create one or more kernel threads,
depending on what the cpumask file is set to, which each thread
spinning in a loop looking for interruptions caused by
something other than the kernel. For example, if a
System Management Interrupt (SMI) takes a noticeable amount of
time, this tracer will detect it. This is useful for testing
if a system is reliable for Real Time tasks.
Some files are created in the tracing directory when this
is enabled:
hwlat_detector/width - time in usecs for how long to spin for
hwlat_detector/window - time in usecs between the start of each
iteration
A kernel thread is created that will spin with interrupts disabled
for "width" microseconds in every "window" cycle. It will not spin
for "window - width" microseconds, where the system can
continue to operate.
The output will appear in the trace and trace_pipe files.
When the tracer is not running, it has no affect on the system,
but when it is running, it can cause the system to be
periodically non responsive. Do not run this tracer on a
production system.
To enable this tracer, echo in "hwlat" into the current_tracer
file. Every time a latency is greater than tracing_thresh, it will
be recorded into the ring buffer.
config OSNOISE_TRACER
bool "OS Noise tracer"
select GENERIC_TRACER
select TRACER_MAX_TRACE
help
In the context of high-performance computing (HPC), the Operating
System Noise (osnoise) refers to the interference experienced by an
application due to activities inside the operating system. In the
context of Linux, NMIs, IRQs, SoftIRQs, and any other system thread
can cause noise to the system. Moreover, hardware-related jobs can
also cause noise, for example, via SMIs.
The osnoise tracer leverages the hwlat_detector by running a similar
loop with preemption, SoftIRQs and IRQs enabled, thus allowing all
the sources of osnoise during its execution. The osnoise tracer takes
note of the entry and exit point of any source of interferences,
increasing a per-cpu interference counter. It saves an interference
counter for each source of interference. The interference counter for
NMI, IRQs, SoftIRQs, and threads is increased anytime the tool
observes these interferences' entry events. When a noise happens
without any interference from the operating system level, the
hardware noise counter increases, pointing to a hardware-related
noise. In this way, osnoise can account for any source of
interference. At the end of the period, the osnoise tracer prints
the sum of all noise, the max single noise, the percentage of CPU
available for the thread, and the counters for the noise sources.
In addition to the tracer, a set of tracepoints were added to
facilitate the identification of the osnoise source.
The output will appear in the trace and trace_pipe files.
To enable this tracer, echo in "osnoise" into the current_tracer
file.
config TIMERLAT_TRACER
bool "Timerlat tracer"
select OSNOISE_TRACER
select GENERIC_TRACER
help
The timerlat tracer aims to help the preemptive kernel developers
to find sources of wakeup latencies of real-time threads.
The tracer creates a per-cpu kernel thread with real-time priority.
The tracer thread sets a periodic timer to wakeup itself, and goes
to sleep waiting for the timer to fire. At the wakeup, the thread
then computes a wakeup latency value as the difference between
the current time and the absolute time that the timer was set
to expire.
The tracer prints two lines at every activation. The first is the
timer latency observed at the hardirq context before the
activation of the thread. The second is the timer latency observed
by the thread, which is the same level that cyclictest reports. The
ACTIVATION ID field serves to relate the irq execution to its
respective thread execution.
The tracer is build on top of osnoise tracer, and the osnoise:
events can be used to trace the source of interference from NMI,
IRQs and other threads. It also enables the capture of the
stacktrace at the IRQ context, which helps to identify the code
path that can cause thread delay.
config MMIOTRACE
bool "Memory mapped IO tracing"
depends on HAVE_MMIOTRACE_SUPPORT && PCI
select GENERIC_TRACER
help
Mmiotrace traces Memory Mapped I/O access and is meant for
debugging and reverse engineering. It is called from the ioremap
implementation and works via page faults. Tracing is disabled by
default and can be enabled at run-time.
See Documentation/trace/mmiotrace.rst.
If you are not helping to develop drivers, say N.
config ENABLE_DEFAULT_TRACERS
bool "Trace process context switches and events"
depends on !GENERIC_TRACER
select TRACING
help
This tracer hooks to various trace points in the kernel,
allowing the user to pick and choose which trace point they
want to trace. It also includes the sched_switch tracer plugin.
config FTRACE_SYSCALLS
bool "Trace syscalls"
depends on HAVE_SYSCALL_TRACEPOINTS
select GENERIC_TRACER
select KALLSYMS
help
Basic tracer to catch the syscall entry and exit events.
config TRACE_SYSCALL_BUF_SIZE_DEFAULT
int "System call user read max size"
range 0 165
default 63
depends on FTRACE_SYSCALLS
help
Some system call trace events will record the data from a user
space address that one of the parameters point to. The amount of
data per event is limited. That limit is set by this config and
this config also affects how much user space data perf can read.
For a tracing instance, this size may be changed by writing into
its syscall_user_buf_size file.
config TRACER_SNAPSHOT
bool "Create a snapshot trace buffer"
select TRACER_MAX_TRACE
help
Allow tracing users to take snapshot of the current buffer using the
ftrace interface, e.g.:
echo 1 > /sys/kernel/tracing/snapshot
cat snapshot
config TRACER_SNAPSHOT_PER_CPU_SWAP
bool "Allow snapshot to swap per CPU"
depends on TRACER_SNAPSHOT
select RING_BUFFER_ALLOW_SWAP
help
Allow doing a snapshot of a single CPU buffer instead of a
full swap (all buffers). If this is set, then the following is
allowed:
echo 1 > /sys/kernel/tracing/per_cpu/cpu2/snapshot
After which, only the tracing buffer for CPU 2 was swapped with
the main tracing buffer, and the other CPU buffers remain the same.
When this is enabled, this adds a little more overhead to the
trace recording, as it needs to add some checks to synchronize
recording with swaps. But this does not affect the performance
of the overall system. This is enabled by default when the preempt
or irq latency tracers are enabled, as those need to swap as well
and already adds the overhead (plus a lot more).
config TRACE_BRANCH_PROFILING
bool
select GENERIC_TRACER
choice
prompt "Branch Profiling"
default BRANCH_PROFILE_NONE
help
The branch profiling is a software profiler. It will add hooks
into the C conditionals to test which path a branch takes.
The likely/unlikely profiler only looks at the conditions that
are annotated with a likely or unlikely macro.
The "all branch" profiler will profile every if-statement in the
kernel. This profiler will also enable the likely/unlikely
profiler.
Either of the above profilers adds a bit of overhead to the system.
If unsure, choose "No branch profiling".
config BRANCH_PROFILE_NONE
bool "No branch profiling"
help
No branch profiling. Branch profiling adds a bit of overhead.
Only enable it if you want to analyse the branching behavior.
Otherwise keep it disabled.
config PROFILE_ANNOTATED_BRANCHES
bool "Trace likely/unlikely profiler"
select TRACE_BRANCH_PROFILING
help
This tracer profiles all likely and unlikely macros
in the kernel. It will display the results in:
/sys/kernel/tracing/trace_stat/branch_annotated
Note: this will add a significant overhead; only turn this
on if you need to profile the system's use of these macros.
config PROFILE_ALL_BRANCHES
bool "Profile all if conditionals" if !FORTIFY_SOURCE
select TRACE_BRANCH_PROFILING
help
This tracer profiles all branch conditions. Every if ()
taken in the kernel is recorded whether it hit or miss.
The results will be displayed in:
/sys/kernel/tracing/trace_stat/branch_all
This option also enables the likely/unlikely profiler.
This configuration, when enabled, will impose a great overhead
on the system. This should only be enabled when the system
is to be analyzed in much detail.
endchoice
config TRACING_BRANCHES
bool
help
Selected by tracers that will trace the likely and unlikely
conditions. This prevents the tracers themselves from being
profiled. Profiling the tracing infrastructure can only happen
when the likelys and unlikelys are not being traced.
config BRANCH_TRACER
bool "Trace likely/unlikely instances"
depends on TRACE_BRANCH_PROFILING
select TRACING_BRANCHES
help
This traces the events of likely and unlikely condition
calls in the kernel. The difference between this and the
"Trace likely/unlikely profiler" is that this is not a
histogram of the callers, but actually places the calling
events into a running trace buffer to see when and where the
events happened, as well as their results.
Say N if unsure.
config BLK_DEV_IO_TRACE
bool "Support for tracing block IO actions"
depends on SYSFS
depends on BLOCK
select RELAY
select DEBUG_FS
select TRACEPOINTS
select GENERIC_TRACER
select STACKTRACE
help
Say Y here if you want to be able to trace the block layer actions
on a given queue. Tracing allows you to see any traffic happening
on a block device queue. For more information (and the userspace
support tools needed), fetch the blktrace tools from:
git://git.kernel.dk/blktrace.git
Tracing also is possible using the ftrace interface, e.g.:
echo 1 > /sys/block/sda/sda1/trace/enable
echo blk > /sys/kernel/tracing/current_tracer
cat /sys/kernel/tracing/trace_pipe
If unsure, say N.
config FPROBE_EVENTS
depends on FPROBE
depends on HAVE_REGS_AND_STACK_ACCESS_API
bool "Enable fprobe-based dynamic events"
select TRACING
select PROBE_EVENTS
select DYNAMIC_EVENTS
default y
help
This allows user to add tracing events on the function entry and
exit via ftrace interface. The syntax is same as the kprobe events
and the kprobe events on function entry and exit will be
transparently converted to this fprobe events.
config PROBE_EVENTS_BTF_ARGS
depends on HAVE_FUNCTION_ARG_ACCESS_API
depends on FPROBE_EVENTS || KPROBE_EVENTS
depends on DEBUG_INFO_BTF && BPF_SYSCALL
bool "Support BTF function arguments for probe events"
default y
help
The user can specify the arguments of the probe event using the names
of the arguments of the probed function, when the probe location is a
kernel function entry or a tracepoint.
This is available only if BTF (BPF Type Format) support is enabled.
config KPROBE_EVENTS
depends on KPROBES
depends on HAVE_REGS_AND_STACK_ACCESS_API
bool "Enable kprobes-based dynamic events"
select TRACING
select PROBE_EVENTS
select DYNAMIC_EVENTS
default y
help
This allows the user to add tracing events (similar to tracepoints)
on the fly via the ftrace interface. See
Documentation/trace/kprobetrace.rst for more details.
Those events can be inserted wherever kprobes can probe, and record
various register and memory values.
This option is also required by perf-probe subcommand of perf tools.
If you want to use perf tools, this option is strongly recommended.
config KPROBE_EVENTS_ON_NOTRACE
bool "Do NOT protect notrace function from kprobe events"
depends on KPROBE_EVENTS
depends on DYNAMIC_FTRACE
default n
help
This is only for the developers who want to debug ftrace itself
using kprobe events.
If kprobes can use ftrace instead of breakpoint, ftrace related
functions are protected from kprobe-events to prevent an infinite
recursion or any unexpected execution path which leads to a kernel
crash.
This option disables such protection and allows you to put kprobe
events on ftrace functions for debugging ftrace by itself.
Note that this might let you shoot yourself in the foot.
If unsure, say N.
config UPROBE_EVENTS
bool "Enable uprobes-based dynamic events"
depends on ARCH_SUPPORTS_UPROBES
depends on MMU
depends on PERF_EVENTS
select UPROBES
select PROBE_EVENTS
select DYNAMIC_EVENTS
select TRACING
default y
help
This allows the user to add tracing events on top of userspace
dynamic events (similar to tracepoints) on the fly via the trace
events interface. Those events can be inserted wherever uprobes
can probe, and record various registers.
This option is required if you plan to use perf-probe subcommand
of perf tools on user space applications.
config EPROBE_EVENTS
bool "Enable event-based dynamic events"
depends on TRACING
depends on HAVE_REGS_AND_STACK_ACCESS_API
select PROBE_EVENTS
select DYNAMIC_EVENTS
default y
help
Eprobes are dynamic events that can be placed on other existing
events. It can be used to limit what fields are recorded in
an event or even dereference a field of an event. It can
convert the type of an event field. For example, turn an
address into a string.
config BPF_EVENTS
depends on BPF_SYSCALL
depends on (KPROBE_EVENTS || UPROBE_EVENTS) && PERF_EVENTS
bool
default y
help
This allows the user to attach BPF programs to kprobe, uprobe, and
tracepoint events.
config DYNAMIC_EVENTS
def_bool n
config PROBE_EVENTS
def_bool n
config BPF_KPROBE_OVERRIDE
bool "Enable BPF programs to override a kprobed function"
depends on BPF_EVENTS
depends on FUNCTION_ERROR_INJECTION
default n
help
Allows BPF to override the execution of a probed function and
set a different return value. This is used for error injection.
config FTRACE_MCOUNT_USE_PATCHABLE_FUNCTION_ENTRY
bool
depends on DYNAMIC_FTRACE
config FTRACE_MCOUNT_USE_CC
def_bool y
depends on $(cc-option,-mrecord-mcount)
depends on !FTRACE_MCOUNT_USE_PATCHABLE_FUNCTION_ENTRY
depends on DYNAMIC_FTRACE
config FTRACE_MCOUNT_USE_OBJTOOL
def_bool y
depends on HAVE_OBJTOOL_MCOUNT
depends on !FTRACE_MCOUNT_USE_PATCHABLE_FUNCTION_ENTRY
depends on !FTRACE_MCOUNT_USE_CC
depends on DYNAMIC_FTRACE
select OBJTOOL
config FTRACE_MCOUNT_USE_RECORDMCOUNT
def_bool y
depends on !FTRACE_MCOUNT_USE_PATCHABLE_FUNCTION_ENTRY
depends on !FTRACE_MCOUNT_USE_CC
depends on !FTRACE_MCOUNT_USE_OBJTOOL
depends on DYNAMIC_FTRACE
config TRACING_MAP
bool
depends on ARCH_HAVE_NMI_SAFE_CMPXCHG
help
tracing_map is a special-purpose lock-free map for tracing,
separated out as a stand-alone facility in order to allow it
to be shared between multiple tracers. It isn't meant to be
generally used outside of that context, and is normally
selected by tracers that use it.
config SYNTH_EVENTS
bool "Synthetic trace events"
select TRACING
select DYNAMIC_EVENTS
default n
help
Synthetic events are user-defined trace events that can be
used to combine data from other trace events or in fact any
data source. Synthetic events can be generated indirectly
via the trace() action of histogram triggers or directly
by way of an in-kernel API.
See Documentation/trace/events.rst or
Documentation/trace/histogram.rst for details and examples.
If in doubt, say N.
config USER_EVENTS
bool "User trace events"
select TRACING
select DYNAMIC_EVENTS
help
User trace events are user-defined trace events that
can be used like an existing kernel trace event. User trace
events are generated by writing to a tracefs file. User
processes can determine if their tracing events should be
generated by registering a value and bit with the kernel
that reflects when it is enabled or not.
See Documentation/trace/user_events.rst.
If in doubt, say N.
config HIST_TRIGGERS
bool "Histogram triggers"
depends on ARCH_HAVE_NMI_SAFE_CMPXCHG
select TRACING_MAP
select TRACING
select DYNAMIC_EVENTS
select SYNTH_EVENTS
default n
help
Hist triggers allow one or more arbitrary trace event fields
to be aggregated into hash tables and dumped to stdout by
reading a debugfs/tracefs file. They're useful for
gathering quick and dirty (though precise) summaries of
event activity as an initial guide for further investigation
using more advanced tools.
Inter-event tracing of quantities such as latencies is also
supported using hist triggers under this option.
See Documentation/trace/histogram.rst.
If in doubt, say N.
config TRACE_EVENT_INJECT
bool "Trace event injection"
depends on TRACING
help
Allow user-space to inject a specific trace event into the ring
buffer. This is mainly used for testing purpose.
If unsure, say N.
config TRACEPOINT_BENCHMARK
bool "Add tracepoint that benchmarks tracepoints"
help
This option creates the tracepoint "benchmark:benchmark_event".
When the tracepoint is enabled, it kicks off a kernel thread that
goes into an infinite loop (calling cond_resched() to let other tasks
run), and calls the tracepoint. Each iteration will record the time
it took to write to the tracepoint and the next iteration that
data will be passed to the tracepoint itself. That is, the tracepoint
will report the time it took to do the previous tracepoint.
The string written to the tracepoint is a static string of 128 bytes
to keep the time the same. The initial string is simply a write of
"START". The second string records the cold cache time of the first
write which is not added to the rest of the calculations.
As it is a tight loop, it benchmarks as hot cache. That's fine because
we care most about hot paths that are probably in cache already.
An example of the output:
START
first=3672 [COLD CACHED]
last=632 first=3672 max=632 min=632 avg=316 std=446 std^2=199712
last=278 first=3672 max=632 min=278 avg=303 std=316 std^2=100337
last=277 first=3672 max=632 min=277 avg=296 std=258 std^2=67064
last=273 first=3672 max=632 min=273 avg=292 std=224 std^2=50411
last=273 first=3672 max=632 min=273 avg=288 std=200 std^2=40389
last=281 first=3672 max=632 min=273 avg=287 std=183 std^2=33666
config RING_BUFFER_BENCHMARK
tristate "Ring buffer benchmark stress tester"
depends on RING_BUFFER
help
This option creates a test to stress the ring buffer and benchmark it.
It creates its own ring buffer such that it will not interfere with
any other users of the ring buffer (such as ftrace). It then creates
a producer and consumer that will run for 10 seconds and sleep for
10 seconds. Each interval it will print out the number of events
it recorded and give a rough estimate of how long each iteration took.
It does not disable interrupts or raise its priority, so it may be
affected by processes that are running.
If unsure, say N.
config TRACE_EVAL_MAP_FILE
bool "Show eval mappings for trace events"
depends on TRACING
help
The "print fmt" of the trace events will show the enum/sizeof names
instead of their values. This can cause problems for user space tools
that use this string to parse the raw data as user space does not know
how to convert the string to its value.
To fix this, there's a special macro in the kernel that can be used
to convert an enum/sizeof into its value. If this macro is used, then
the print fmt strings will be converted to their values.
If something does not get converted properly, this option can be
used to show what enums/sizeof the kernel tried to convert.
This option is for debugging the conversions. A file is created
in the tracing directory called "eval_map" that will show the
names matched with their values and what trace event system they
belong too.
Normally, the mapping of the strings to values will be freed after
boot up or module load. With this option, they will not be freed, as
they are needed for the "eval_map" file. Enabling this option will
increase the memory footprint of the running kernel.
If unsure, say N.
config FTRACE_RECORD_RECURSION
bool "Record functions that recurse in function tracing"
depends on FUNCTION_TRACER
help
All callbacks that attach to the function tracing have some sort
of protection against recursion. Even though the protection exists,
it adds overhead. This option will create a file in the tracefs
file system called "recursed_functions" that will list the functions
that triggered a recursion.
This will add more overhead to cases that have recursion.
If unsure, say N
config FTRACE_RECORD_RECURSION_SIZE
int "Max number of recursed functions to record"
default 128
depends on FTRACE_RECORD_RECURSION
help
This defines the limit of number of functions that can be
listed in the "recursed_functions" file, that lists all
the functions that caused a recursion to happen.
This file can be reset, but the limit can not change in
size at runtime.
config FTRACE_VALIDATE_RCU_IS_WATCHING
bool "Validate RCU is on during ftrace execution"
depends on FUNCTION_TRACER
depends on ARCH_WANTS_NO_INSTR
help
All callbacks that attach to the function tracing have some sort of
protection against recursion. This option is only to verify that
ftrace (and other users of ftrace_test_recursion_trylock()) are not
called outside of RCU, as if they are, it can cause a race. But it
also has a noticeable overhead when enabled.
If unsure, say N
config RING_BUFFER_RECORD_RECURSION
bool "Record functions that recurse in the ring buffer"
depends on FTRACE_RECORD_RECURSION
# default y, because it is coupled with FTRACE_RECORD_RECURSION
default y
help
The ring buffer has its own internal recursion. Although when
recursion happens it won't cause harm because of the protection,
but it does cause unwanted overhead. Enabling this option will
place where recursion was detected into the ftrace "recursed_functions"
file.
This will add more overhead to cases that have recursion.
config GCOV_PROFILE_FTRACE
bool "Enable GCOV profiling on ftrace subsystem"
depends on GCOV_KERNEL
help
Enable GCOV profiling on ftrace subsystem for checking
which functions/lines are tested.
If unsure, say N.
Note that on a kernel compiled with this config, ftrace will
run significantly slower.
config FTRACE_SELFTEST
bool
config FTRACE_STARTUP_TEST
bool "Perform a startup test on ftrace"
depends on GENERIC_TRACER
select FTRACE_SELFTEST
help
This option performs a series of startup tests on ftrace. On bootup
a series of tests are made to verify that the tracer is
functioning properly. It will do tests on all the configured
tracers of ftrace.
config EVENT_TRACE_STARTUP_TEST
bool "Run selftest on trace events"
depends on FTRACE_STARTUP_TEST
default y
help
This option performs a test on all trace events in the system.
It basically just enables each event and runs some code that
will trigger events (not necessarily the event it enables)
This may take some time run as there are a lot of events.
config EVENT_TRACE_TEST_SYSCALLS
bool "Run selftest on syscall events"
depends on EVENT_TRACE_STARTUP_TEST
help
This option will also enable testing every syscall event.
It only enables the event and disables it and runs various loads
with the event enabled. This adds a bit more time for kernel boot
up since it runs this on every system call defined.
TBD - enable a way to actually call the syscalls as we test their
events
config FTRACE_SORT_STARTUP_TEST
bool "Verify compile time sorting of ftrace functions"
depends on DYNAMIC_FTRACE
depends on BUILDTIME_MCOUNT_SORT
help
Sorting of the mcount_loc sections that is used to find the
where the ftrace knows where to patch functions for tracing
and other callbacks is done at compile time. But if the sort
is not done correctly, it will cause non-deterministic failures.
When this is set, the sorted sections will be verified that they
are in deed sorted and will warn if they are not.
If unsure, say N
config RING_BUFFER_STARTUP_TEST
bool "Ring buffer startup self test"
depends on RING_BUFFER
help
Run a simple self test on the ring buffer on boot up. Late in the
kernel boot sequence, the test will start that kicks off
a thread per cpu. Each thread will write various size events
into the ring buffer. Another thread is created to send IPIs
to each of the threads, where the IPI handler will also write
to the ring buffer, to test/stress the nesting ability.
If any anomalies are discovered, a warning will be displayed
and all ring buffers will be disabled.
The test runs for 10 seconds. This will slow your boot time
by at least 10 more seconds.
At the end of the test, statistics and more checks are done.
It will output the stats of each per cpu buffer: What
was written, the sizes, what was read, what was lost, and
other similar details.
If unsure, say N
config RING_BUFFER_VALIDATE_TIME_DELTAS
bool "Verify ring buffer time stamp deltas"
depends on RING_BUFFER
help
This will audit the time stamps on the ring buffer sub
buffer to make sure that all the time deltas for the
events on a sub buffer matches the current time stamp.
This audit is performed for every event that is not
interrupted, or interrupting another event. A check
is also made when traversing sub buffers to make sure
that all the deltas on the previous sub buffer do not
add up to be greater than the current time stamp.
NOTE: This adds significant overhead to recording of events,
and should only be used to test the logic of the ring buffer.
Do not use it on production systems.
Only say Y if you understand what this does, and you
still want it enabled. Otherwise say N
config MMIOTRACE_TEST
tristate "Test module for mmiotrace"
depends on MMIOTRACE && m
help
This is a dumb module for testing mmiotrace. It is very dangerous
as it will write garbage to IO memory starting at a given address.
However, it should be safe to use on e.g. unused portion of VRAM.
Say N, unless you absolutely know what you are doing.
config PREEMPTIRQ_DELAY_TEST
tristate "Test module to create a preempt / IRQ disable delay thread to test latency tracers"
depends on m
help
Select this option to build a test module that can help test latency
tracers by executing a preempt or irq disable section with a user
configurable delay. The module busy waits for the duration of the
critical section.
For example, the following invocation generates a burst of three
irq-disabled critical sections for 500us:
modprobe preemptirq_delay_test test_mode=irq delay=500 burst_size=3
What's more, if you want to attach the test on the cpu which the latency
tracer is running on, specify cpu_affinity=cpu_num at the end of the
command.
If unsure, say N
config SYNTH_EVENT_GEN_TEST
tristate "Test module for in-kernel synthetic event generation"
depends on SYNTH_EVENTS && m
help
This option creates a test module to check the base
functionality of in-kernel synthetic event definition and
generation.
To test, insert the module, and then check the trace buffer
for the generated sample events.
If unsure, say N.
config KPROBE_EVENT_GEN_TEST
tristate "Test module for in-kernel kprobe event generation"
depends on KPROBE_EVENTS && m
help
This option creates a test module to check the base
functionality of in-kernel kprobe event definition.
To test, insert the module, and then check the trace buffer
for the generated kprobe events.
If unsure, say N.
config HIST_TRIGGERS_DEBUG
bool "Hist trigger debug support"
depends on HIST_TRIGGERS
help
Add "hist_debug" file for each event, which when read will
dump out a bunch of internal details about the hist triggers
defined on that event.
The hist_debug file serves a couple of purposes:
- Helps developers verify that nothing is broken.
- Provides educational information to support the details
of the hist trigger internals as described by
Documentation/trace/histogram-design.rst.
The hist_debug output only covers the data structures
related to the histogram definitions themselves and doesn't
display the internals of map buckets or variable values of
running histograms.
If unsure, say N.
source "kernel/trace/rv/Kconfig"
endif # FTRACE
]
|
| 32/2 |
2026/05/25 22:13 |
llm |
expert |
0m
Model:
Tokens:
input: 47985
output: 43
thoughts: 811
Thoughts: **Analyzing the Ring Buffer Time Delta Validation and Its Impact**
Okay, here's what's running through my head. This `RING_BUFFER_VALIDATE_TIME_DELTAS` option⦠right away, the "NOTE" jumps out. It screams "debug only!" Significant overhead? Only for testing the ring buffer logic? Definitely a flag for development and not for production. That immediately tells me if this *is* the source of a bug, it's probably not going to be present in released code.
But is it the *only* cause? I need to think this through. The reproducer attaches a BPF program 65 times to `rcu_utilization`. That tracepoint is called from the timer interrupt via `rcu_sched_clock_irq`, which means this BPF program is going to get triggered frequently. The BPF program calls `bpf_trace_printk`. Now, if the validation is enabled, we're looking at O(N^2) time *per page* for `bpf_trace_printk`. That's a huge red flag! The key here is the `RING_BUFFER_VALIDATE_TIME_DELTAS` option. This might explain the stall if it's slow enough.
So, if it's disabled, `bpf_trace_printk` is "fast", but I need to think if 65 calls per timer interrupt are still a problem. Not necessarily, if it is fast ring buffer writes. Wait⦠`bpf_trace_printk` does formatting, right? The reproducer's BPF program is just calling `bpf_trace_printk` with some simple strings. Simple.
Then I get a crucial thought: Can an unprivileged user even do this? Can they attach a BPF program to a tracepoint? No, it's got to be `CAP_SYS_ADMIN`, `CAP_PERFMON`, or `CAP_BPF`! I need to know if the reproducer does anything to escalate privileges. Let me look at the code: `unshare`... It doesn't unshare the user namespace. It only unshares FS and FILES! It means that the reproducer's process is running as root.
"CPU: 1 UID: 0 PID: 31 Comm: khungtaskd Not tainted syzkaller #0 PREEMPT(full)". Okay, Syzkaller typically runs as root in its VM. Now, if root privilege is required, that greatly limits how this impacts regular users. Can unprivileged users load `BPF_PROG_TYPE_RAW_TRACEPOINT`? I'll need to check the permissions on `bpf_prog_load`. I'm starting to build a picture: debug-only code, root-level permissions, and a potentially quadratic performance bottleneck in a highly called function.
|
| 33/2 |
2026/05/25 22:13 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_prog_load]
Results: map[SourceCode:2870: static int bpf_prog_load(union bpf_attr *attr, bpfptr_t uattr, u32 uattr_size)
2871: {
2872: enum bpf_prog_type type = attr->prog_type;
2873: struct bpf_prog *prog, *dst_prog = NULL;
2874: struct btf *attach_btf = NULL;
2875: struct bpf_token *token = NULL;
2876: bool bpf_cap;
2877: int err;
2878: char license[128];
2879:
2880: if (CHECK_ATTR(BPF_PROG_LOAD))
2881: return -EINVAL;
2882:
2883: if (attr->prog_flags & ~(BPF_F_STRICT_ALIGNMENT |
2884: BPF_F_ANY_ALIGNMENT |
2885: BPF_F_TEST_STATE_FREQ |
2886: BPF_F_SLEEPABLE |
2887: BPF_F_TEST_RND_HI32 |
2888: BPF_F_XDP_HAS_FRAGS |
2889: BPF_F_XDP_DEV_BOUND_ONLY |
2890: BPF_F_TEST_REG_INVARIANTS |
2891: BPF_F_TOKEN_FD))
2892: return -EINVAL;
2893:
2894: bpf_prog_load_fixup_attach_type(attr);
2895:
2896: if (attr->prog_flags & BPF_F_TOKEN_FD) {
2897: token = bpf_token_get_from_fd(attr->prog_token_fd);
2898: if (IS_ERR(token))
2899: return PTR_ERR(token);
2900: /* if current token doesn't grant prog loading permissions,
2901: * then we can't use this token, so ignore it and rely on
2902: * system-wide capabilities checks
2903: */
2904: if (!bpf_token_allow_cmd(token, BPF_PROG_LOAD) ||
2905: !bpf_token_allow_prog_type(token, attr->prog_type,
2906: attr->expected_attach_type)) {
2907: bpf_token_put(token);
2908: token = NULL;
2909: }
2910: }
2911:
2912: bpf_cap = bpf_token_capable(token, CAP_BPF);
2913: err = -EPERM;
2914:
2915: if (!IS_ENABLED(CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS) &&
2916: (attr->prog_flags & BPF_F_ANY_ALIGNMENT) &&
2917: !bpf_cap)
2918: goto put_token;
2919:
2920: /* Intent here is for unprivileged_bpf_disabled to block BPF program
2921: * creation for unprivileged users; other actions depend
2922: * on fd availability and access to bpffs, so are dependent on
2923: * object creation success. Even with unprivileged BPF disabled,
2924: * capability checks are still carried out for these
2925: * and other operations.
2926: */
2927: if (sysctl_unprivileged_bpf_disabled && !bpf_cap)
2928: goto put_token;
2929:
2930: if (attr->insn_cnt == 0 ||
2931: attr->insn_cnt > (bpf_cap ? BPF_COMPLEXITY_LIMIT_INSNS : BPF_MAXINSNS)) {
2932: err = -E2BIG;
2933: goto put_token;
2934: }
2935: if (type != BPF_PROG_TYPE_SOCKET_FILTER &&
2936: type != BPF_PROG_TYPE_CGROUP_SKB &&
2937: !bpf_cap)
2938: goto put_token;
2939:
2940: if (is_net_admin_prog_type(type) && !bpf_token_capable(token, CAP_NET_ADMIN))
2941: goto put_token;
2942: if (is_perfmon_prog_type(type) && !bpf_token_capable(token, CAP_PERFMON))
2943: goto put_token;
2944:
2945: /* attach_prog_fd/attach_btf_obj_fd can specify fd of either bpf_prog
2946: * or btf, we need to check which one it is
2947: */
2948: if (attr->attach_prog_fd) {
2949: dst_prog = bpf_prog_get(attr->attach_prog_fd);
2950: if (IS_ERR(dst_prog)) {
2951: dst_prog = NULL;
2952: attach_btf = btf_get_by_fd(attr->attach_btf_obj_fd);
2953: if (IS_ERR(attach_btf)) {
2954: err = -EINVAL;
2955: goto put_token;
2956: }
2957: if (!btf_is_kernel(attach_btf)) {
2958: /* attaching through specifying bpf_prog's BTF
2959: * objects directly might be supported eventually
2960: */
2961: btf_put(attach_btf);
2962: err = -ENOTSUPP;
2963: goto put_token;
2964: }
2965: }
2966: } else if (attr->attach_btf_id) {
2967: /* fall back to vmlinux BTF, if BTF type ID is specified */
2968: attach_btf = bpf_get_btf_vmlinux();
2969: if (IS_ERR(attach_btf)) {
2970: err = PTR_ERR(attach_btf);
2971: goto put_token;
2972: }
2973: if (!attach_btf) {
2974: err = -EINVAL;
2975: goto put_token;
2976: }
2977: btf_get(attach_btf);
2978: }
2979:
2980: if (bpf_prog_load_check_attach(type, attr->expected_attach_type,
2981: attach_btf, attr->attach_btf_id,
2982: dst_prog)) {
2983: if (dst_prog)
2984: bpf_prog_put(dst_prog);
2985: if (attach_btf)
2986: btf_put(attach_btf);
2987: err = -EINVAL;
2988: goto put_token;
2989: }
2990:
2991: /* plain bpf_prog allocation */
2992: prog = bpf_prog_alloc(bpf_prog_size(attr->insn_cnt), GFP_USER);
2993: if (!prog) {
2994: if (dst_prog)
2995: bpf_prog_put(dst_prog);
2996: if (attach_btf)
2997: btf_put(attach_btf);
2998: err = -EINVAL;
2999: goto put_token;
3000: }
3001:
3002: prog->expected_attach_type = attr->expected_attach_type;
3003: prog->sleepable = !!(attr->prog_flags & BPF_F_SLEEPABLE);
3004: prog->aux->attach_btf = attach_btf;
3005: prog->aux->attach_btf_id = attr->attach_btf_id;
3006: prog->aux->dst_prog = dst_prog;
3007: prog->aux->dev_bound = !!attr->prog_ifindex;
3008: prog->aux->xdp_has_frags = attr->prog_flags & BPF_F_XDP_HAS_FRAGS;
3009:
3010: /* move token into prog->aux, reuse taken refcnt */
3011: prog->aux->token = token;
3012: token = NULL;
3013:
3014: prog->aux->user = get_current_user();
3015: prog->len = attr->insn_cnt;
3016:
3017: err = -EFAULT;
3018: if (copy_from_bpfptr(prog->insns,
3019: make_bpfptr(attr->insns, uattr.is_kernel),
3020: bpf_prog_insn_size(prog)) != 0)
3021: goto free_prog;
3022: /* copy eBPF program license from user space */
3023: if (strncpy_from_bpfptr(license,
3024: make_bpfptr(attr->license, uattr.is_kernel),
3025: sizeof(license) - 1) < 0)
3026: goto free_prog;
3027: license[sizeof(license) - 1] = 0;
3028:
3029: /* eBPF programs must be GPL compatible to use GPL-ed functions */
3030: prog->gpl_compatible = license_is_gpl_compatible(license) ? 1 : 0;
3031:
3032: if (attr->signature) {
3033: err = bpf_prog_verify_signature(prog, attr, uattr.is_kernel);
3034: if (err)
3035: goto free_prog;
3036: }
3037:
3038: prog->orig_prog = NULL;
3039: prog->jited = 0;
3040:
3041: atomic64_set(&prog->aux->refcnt, 1);
3042:
3043: if (bpf_prog_is_dev_bound(prog->aux)) {
3044: err = bpf_prog_dev_bound_init(prog, attr);
3045: if (err)
3046: goto free_prog;
3047: }
3048:
3049: if (type == BPF_PROG_TYPE_EXT && dst_prog &&
3050: bpf_prog_is_dev_bound(dst_prog->aux)) {
3051: err = bpf_prog_dev_bound_inherit(prog, dst_prog);
3052: if (err)
3053: goto free_prog;
3054: }
3055:
3056: /*
3057: * Bookkeeping for managing the program attachment chain.
3058: *
3059: * It might be tempting to set attach_tracing_prog flag at the attachment
3060: * time, but this will not prevent from loading bunch of tracing prog
3061: * first, then attach them one to another.
3062: *
3063: * The flag attach_tracing_prog is set for the whole program lifecycle, and
3064: * doesn't have to be cleared in bpf_tracing_link_release, since tracing
3065: * programs cannot change attachment target.
3066: */
3067: if (type == BPF_PROG_TYPE_TRACING && dst_prog &&
3068: dst_prog->type == BPF_PROG_TYPE_TRACING) {
3069: prog->aux->attach_tracing_prog = true;
3070: }
3071:
3072: /* find program type: socket_filter vs tracing_filter */
3073: err = find_prog_type(type, prog);
3074: if (err < 0)
3075: goto free_prog;
3076:
3077: prog->aux->load_time = ktime_get_boottime_ns();
3078: err = bpf_obj_name_cpy(prog->aux->name, attr->prog_name,
3079: sizeof(attr->prog_name));
3080: if (err < 0)
3081: goto free_prog;
3082:
3083: err = security_bpf_prog_load(prog, attr, token, uattr.is_kernel);
3084: if (err)
3085: goto free_prog_sec;
3086:
3087: /* run eBPF verifier */
3088: err = bpf_check(&prog, attr, uattr, uattr_size);
3089: if (err < 0)
3090: goto free_used_maps;
3091:
3092: prog = bpf_prog_select_runtime(prog, &err);
3093: if (err < 0)
3094: goto free_used_maps;
3095:
3096: err = bpf_prog_mark_insn_arrays_ready(prog);
3097: if (err < 0)
3098: goto free_used_maps;
3099:
3100: err = bpf_prog_alloc_id(prog);
3101: if (err)
3102: goto free_used_maps;
3103:
3104: /* Upon success of bpf_prog_alloc_id(), the BPF prog is
3105: * effectively publicly exposed. However, retrieving via
3106: * bpf_prog_get_fd_by_id() will take another reference,
3107: * therefore it cannot be gone underneath us.
3108: *
3109: * Only for the time /after/ successful bpf_prog_new_fd()
3110: * and before returning to userspace, we might just hold
3111: * one reference and any parallel close on that fd could
3112: * rip everything out. Hence, below notifications must
3113: * happen before bpf_prog_new_fd().
3114: *
3115: * Also, any failure handling from this point onwards must
3116: * be using bpf_prog_put() given the program is exposed.
3117: */
3118: bpf_prog_kallsyms_add(prog);
3119: perf_event_bpf_event(prog, PERF_BPF_EVENT_PROG_LOAD, 0);
3120: bpf_audit_prog(prog, BPF_AUDIT_LOAD);
3121:
3122: err = bpf_prog_new_fd(prog);
3123: if (err < 0)
3124: bpf_prog_put(prog);
3125: return err;
3126:
3127: free_used_maps:
3128: /* In case we have subprogs, we need to wait for a grace
3129: * period before we can tear down JIT memory since symbols
3130: * are already exposed under kallsyms.
3131: */
3132: __bpf_prog_put_noref(prog, prog->aux->real_func_cnt);
3133: return err;
3134:
3135: free_prog_sec:
3136: security_bpf_prog_free(prog);
3137: free_prog:
3138: free_uid(prog->aux->user);
3139: if (prog->aux->attach_btf)
3140: btf_put(prog->aux->attach_btf);
3141: bpf_prog_free(prog);
3142: put_token:
3143: bpf_token_put(token);
3144: return err;
3145: }
SourceFile:kernel/bpf/syscall.c]
|
| 34/2 |
2026/05/25 22:13 |
llm |
expert |
0m
Model:
Tokens:
input: 53985
output: 45
thoughts: 58
Thoughts: **Checking Perfmon Program Type and Capability**
Okay, so I'm looking at this bit of code and my initial thought is: first, we've got this `if` statement, and it's key. It's focused on handling a specific type of program, `is_perfmon_prog_type(type)`. This suggests it's related to performance monitoring. Now, the condition also checks if we have the right `bpf_token_capable` and `CAP_PERFMON` capabilities. If the program *is* a performance monitoring program, and we *don't* have the right token capabilities, we need to handle that immediately. We'll `goto put_token`. That `put_token` part is probably about cleaning up resources or signaling a problem. This makes sense. We don't want unauthorized access or programs messing with the performance monitoring infrastructure. It has to be properly secured. This is a crucial check.
|
| 35/2 |
2026/05/25 22:13 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:is_perfmon_prog_type]
Results: map[SourceCode:2796: static bool is_perfmon_prog_type(enum bpf_prog_type prog_type)
2797: {
2798: switch (prog_type) {
2799: case BPF_PROG_TYPE_KPROBE:
2800: case BPF_PROG_TYPE_TRACEPOINT:
2801: case BPF_PROG_TYPE_PERF_EVENT:
2802: case BPF_PROG_TYPE_RAW_TRACEPOINT:
2803: case BPF_PROG_TYPE_RAW_TRACEPOINT_WRITABLE:
2804: case BPF_PROG_TYPE_TRACING:
2805: case BPF_PROG_TYPE_LSM:
2806: case BPF_PROG_TYPE_STRUCT_OPS: /* has access to struct sock */
2807: case BPF_PROG_TYPE_EXT: /* extends any prog */
2808: return true;
2809: default:
2810: return false;
2811: }
2812: }
SourceFile:kernel/bpf/syscall.c]
|
| 36/2 |
2026/05/25 22:13 |
llm |
expert |
0m
Model:
Tokens:
input: 54611
output: 43
thoughts: 64
Thoughts: Okay, here's my interpretation of this thought process, keeping in mind my expertise:
**Checking BPF Raw Tracepoint Program Loading Permissions**
Alright, so I'm thinking about the implications of using `BPF_PROG_TYPE_RAW_TRACEPOINT`. I remember that loading a BPF program of this type demands `CAP_PERFMON` privileges. Makes sense; you're messing with fairly low-level system events. Now, how does the kernel actually *enforce* that requirement? Ah, right, it calls `bpf_token_capable`. Time to double-check my understanding of that function, specifically how it assesses capability. I need to make sure I fully grasp the security implications here, how the `token` is being evaluated against `CAP_PERFMON`. I need to ensure that the implementation is sound and that there aren't any potential bypasses.
|
| 37/2 |
2026/05/25 22:13 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_token_capable]
Results: map[SourceCode: 16: bool bpf_token_capable(const struct bpf_token *token, int cap)
17: {
18: struct user_namespace *userns;
19:
20: /* BPF token allows ns_capable() level of capabilities */
21: userns = token ? token->userns : &init_user_ns;
22: if (!bpf_ns_capable(userns, cap))
23: return false;
24: if (token && security_bpf_token_capable(token, cap) < 0)
25: return false;
26: return true;
27: }
SourceFile:kernel/bpf/token.c]
|
| 38/2 |
2026/05/25 22:13 |
llm |
expert |
0m
Model:
Tokens:
input: 55113
output: 43
thoughts: 123
Thoughts: **Analyzing BPF Capability Checks**
Okay, so I'm diving into this `bpf_ns_capable` function. My initial thought is that it's simply a check to see if a user has a specific capability, `cap`, within a given user namespace, `userns`. Pretty straightforward. But then I see this `token` being passed in from `attr->prog_token_fd`. Ah, right, that's what's getting me.
If the `token` is `NULL`, it appears the code falls back to checking against `&init_user_ns`. So, if there's no token, we're dealing with the initial user namespace, and that's when you *absolutely* need `CAP_PERFMON`. I understand that part. But this `bpf_ns_capable` with the `token`... does that mean I can bypass this requirement? Potentially, yes. It looks like `bpf_ns_capable` might grant capabilities differently if I *do* provide a `token`.
So the crucial question is, can an unprivileged user even create one of these BPF tokens? Thatβs the key. That's the loophole to explore. Now I need to check `bpf_token_create` to verify. I need to understand whether this is something I can actually leverage.
|
| 39/2 |
2026/05/25 22:14 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/token.c IncludeLines:true Name:bpf_token_create]
Results: map[SourceCode: 111: int bpf_token_create(union bpf_attr *attr)
112: {
113: struct bpf_token *token __free(kfree) = NULL;
114: struct bpf_mount_opts *mnt_opts;
115: struct user_namespace *userns;
116: struct inode *inode;
117: CLASS(fd, f)(attr->token_create.bpffs_fd);
118: struct path path;
119: struct super_block *sb;
120: umode_t mode;
121: int err;
122:
123: if (fd_empty(f))
124: return -EBADF;
125:
126: path = fd_file(f)->f_path;
127: sb = path.dentry->d_sb;
128:
129: if (path.dentry != sb->s_root)
130: return -EINVAL;
131: if (sb->s_op != &bpf_super_ops)
132: return -EINVAL;
133: err = path_permission(&path, MAY_ACCESS);
134: if (err)
135: return err;
136:
137: userns = sb->s_user_ns;
138: /*
139: * Enforce that creators of BPF tokens are in the same user
140: * namespace as the BPF FS instance. This makes reasoning about
141: * permissions a lot easier and we can always relax this later.
142: */
143: if (current_user_ns() != userns)
144: return -EPERM;
145: if (!ns_capable(userns, CAP_BPF))
146: return -EPERM;
147:
148: /* Creating BPF token in init_user_ns doesn't make much sense. */
149: if (current_user_ns() == &init_user_ns)
150: return -EOPNOTSUPP;
151:
152: mnt_opts = sb->s_fs_info;
153: if (mnt_opts->delegate_cmds == 0 &&
154: mnt_opts->delegate_maps == 0 &&
155: mnt_opts->delegate_progs == 0 &&
156: mnt_opts->delegate_attachs == 0)
157: return -ENOENT; /* no BPF token delegation is set up */
158:
159: mode = S_IFREG | ((S_IRUSR | S_IWUSR) & ~current_umask());
160: inode = bpf_get_inode(sb, NULL, mode);
161: if (IS_ERR(inode))
162: return PTR_ERR(inode);
163:
164: inode->i_op = &bpf_token_iops;
165: inode->i_fop = &bpf_token_fops;
166: clear_nlink(inode); /* make sure it is unlinked */
167:
168: FD_PREPARE(fdf, O_CLOEXEC,
169: alloc_file_pseudo(inode, path.mnt, BPF_TOKEN_INODE_NAME,
170: O_RDWR, &bpf_token_fops));
171: if (fdf.err)
172: return fdf.err;
173:
174: token = kzalloc(sizeof(*token), GFP_USER);
175: if (!token)
176: return -ENOMEM;
177:
178: atomic64_set(&token->refcnt, 1);
179:
180: /* remember bpffs owning userns for future ns_capable() checks. */
181: token->userns = userns;
182: token->allowed_cmds = mnt_opts->delegate_cmds;
183: token->allowed_maps = mnt_opts->delegate_maps;
184: token->allowed_progs = mnt_opts->delegate_progs;
185: token->allowed_attachs = mnt_opts->delegate_attachs;
186:
187: err = security_bpf_token_create(token, attr, &path);
188: if (err)
189: return err;
190:
191: get_user_ns(token->userns);
192: fd_prepare_file(fdf)->private_data = no_free_ptr(token);
193: return fd_publish(fdf);
194: }
SourceFile:kernel/bpf/token.c]
|
| 40/2 |
2026/05/25 22:14 |
llm |
expert |
0m
Model:
Tokens:
input: 56946
output: 46
thoughts: 550
Thoughts: **Unveiling a Potential BPF Vulnerability**
Okay, so here's what's running through my head. `bpf_token_create` *requires* `ns_capable(userns, CAP_BPF)`. That means a totally unprivileged user can create their own user namespace, mount `bpffs` with the delegation option, and then create a token. Alright, that seems straightforward enough.
But now, hold on. `bpf_token_capable` is the check. It calls `bpf_ns_capable(userns, cap)`. And if `cap` is, say, `CAP_PERFMON`, it then checks `bpf_ns_capable(userns, CAP_PERFMON)`. I've seen that `bpf_ns_capable` is just a wrapper for `ns_capable(ns, cap)`. And in a user namespace? You get all the capabilities, including `CAP_PERFMON`! This is starting to look interesting.
So, *can* they load `BPF_PROG_TYPE_RAW_TRACEPOINT` using a token? Let's go back and examine `bpf_prog_load`...
```c
if (is_perfmon_prog_type(type) && !bpf_token_capable(token, CAP_PERFMON))
goto put_token;
```
Ah! That part. If a `token` is provided, `userns` is the user namespace of the *token*. And, since the user *has* `CAP_PERFMON` in their user namespace, this check *passes*! Okay, now I am really digging in. But wait, can we confirm that `bpf_token_capable` allows `CAP_PERFMON` in this case?
Let's look again at the source:
```c
bool bpf_token_capable(const struct bpf_token *token, int cap)
{
struct user_namespace *userns;
/* BPF token allows ns_capable() level of capabilities */
userns = token ? token->userns : &init_user_ns;
if (!bpf_ns_capable(userns, cap))
return false;
if (token && security_bpf_token_capable(token, cap) < 0)
return false;
return true;
}
```
Yup! It *definitely* checks `ns_capable(userns, cap)`. So, if I'm understanding this correctly, an unprivileged user can indeed load a `BPF_PROG_TYPE_RAW_TRACEPOINT` program *if* they use a token.
But can they *attach* it? I better check `bpf_raw_tracepoint_open`...
|
| 41/2 |
2026/05/25 22:14 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_raw_tracepoint_open]
Results: map[SourceCode:4280: static int bpf_raw_tracepoint_open(const union bpf_attr *attr)
4281: {
4282: struct bpf_prog *prog;
4283: void __user *tp_name;
4284: __u64 cookie;
4285: int fd;
4286:
4287: if (CHECK_ATTR(BPF_RAW_TRACEPOINT_OPEN))
4288: return -EINVAL;
4289:
4290: prog = bpf_prog_get(attr->raw_tracepoint.prog_fd);
4291: if (IS_ERR(prog))
4292: return PTR_ERR(prog);
4293:
4294: tp_name = u64_to_user_ptr(attr->raw_tracepoint.name);
4295: cookie = attr->raw_tracepoint.cookie;
4296: fd = bpf_raw_tp_link_attach(prog, tp_name, cookie, prog->expected_attach_type);
4297: if (fd < 0)
4298: bpf_prog_put(prog);
4299: return fd;
4300: }
SourceFile:kernel/bpf/syscall.c]
|
| 42/2 |
2026/05/25 22:14 |
llm |
expert |
0m
Model:
Tokens:
input: 58524
output: 47
thoughts: 16
Thoughts: Here's my attempt at summarizing the thought process, tailored for an expert audience and written in the first person:
**BPF Raw Tracepoint Link Attachment Deep Dive**
Okay, time to dig into `bpf_raw_tp_link_attach`. I need to refresh my understanding of the raw tracepoint linking process. Let's see... what are the key aspects I need to consider? I'm immediately thinking about the interaction with the underlying tracepoint infrastructure. Is it using the standard kprobe/uprobe machinery, or is it doing something more directly at the kernel probe entry points? I know raw tracepoints have their own specific set of constraints, so I need to keep those in mind.
I'll be looking at how this function integrates with the existing BPF verifier and how it handles potential conflicts or overlaps with other tracepoint handlers. What kind of validation is performed on the BPF program being attached? Are there specific checks for raw tracepoints that differ from the regular `tracepoint_probe_attach` logic? The synchronization mechanisms are another important aspect. How is it ensuring that the BPF program doesn't interfere with other kernel activities, or create any race conditions during attach/detach?
I'm particularly interested in any potential performance implications. Are there any known overheads, compared to other BPF attach methods? Is there a difference in memory usage? Any potential issues with CPU cache coherency because of the raw tracepoint interactions? I also need to be vigilant for any edge cases, especially related to error handling and failure modes. What happens if the attach fails? How are resources cleaned up? I will pay special attention to the parameters the function accepts and the return values it can produce. Basically, I want to confirm that this is a robust and efficient way to interact with raw tracepoints. Let's see the code!
|
| 43/2 |
2026/05/25 22:14 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:bpf_raw_tp_link_attach]
Results: map[SourceCode:4207: static int bpf_raw_tp_link_attach(struct bpf_prog *prog,
4208: const char __user *user_tp_name, u64 cookie,
4209: enum bpf_attach_type attach_type)
4210: {
4211: struct bpf_link_primer link_primer;
4212: struct bpf_raw_tp_link *link;
4213: struct bpf_raw_event_map *btp;
4214: const char *tp_name;
4215: char buf[128];
4216: int err;
4217:
4218: switch (prog->type) {
4219: case BPF_PROG_TYPE_TRACING:
4220: case BPF_PROG_TYPE_EXT:
4221: case BPF_PROG_TYPE_LSM:
4222: if (user_tp_name)
4223: /* The attach point for this category of programs
4224: * should be specified via btf_id during program load.
4225: */
4226: return -EINVAL;
4227: if (prog->type == BPF_PROG_TYPE_TRACING &&
4228: prog->expected_attach_type == BPF_TRACE_RAW_TP) {
4229: tp_name = prog->aux->attach_func_name;
4230: break;
4231: }
4232: return bpf_tracing_prog_attach(prog, 0, 0, 0, attach_type);
4233: case BPF_PROG_TYPE_RAW_TRACEPOINT:
4234: case BPF_PROG_TYPE_RAW_TRACEPOINT_WRITABLE:
4235: if (strncpy_from_user(buf, user_tp_name, sizeof(buf) - 1) < 0)
4236: return -EFAULT;
4237: buf[sizeof(buf) - 1] = 0;
4238: tp_name = buf;
4239: break;
4240: default:
4241: return -EINVAL;
4242: }
4243:
4244: btp = bpf_get_raw_tracepoint(tp_name);
4245: if (!btp)
4246: return -ENOENT;
4247:
4248: link = kzalloc(sizeof(*link), GFP_USER);
4249: if (!link) {
4250: err = -ENOMEM;
4251: goto out_put_btp;
4252: }
4253: bpf_link_init_sleepable(&link->link, BPF_LINK_TYPE_RAW_TRACEPOINT,
4254: &bpf_raw_tp_link_lops, prog, attach_type,
4255: tracepoint_is_faultable(btp->tp));
4256: link->btp = btp;
4257: link->cookie = cookie;
4258:
4259: err = bpf_link_prime(&link->link, &link_primer);
4260: if (err) {
4261: kfree(link);
4262: goto out_put_btp;
4263: }
4264:
4265: err = bpf_probe_register(link->btp, link);
4266: if (err) {
4267: bpf_link_cleanup(&link_primer);
4268: goto out_put_btp;
4269: }
4270:
4271: return bpf_link_settle(&link_primer);
4272:
4273: out_put_btp:
4274: bpf_put_raw_tracepoint(btp);
4275: return err;
4276: }
SourceFile:kernel/bpf/syscall.c]
|
| 44/2 |
2026/05/25 22:14 |
llm |
expert |
0m
Model:
Tokens:
input: 60094
output: 41
thoughts: 75
Thoughts: Okay, let me think this through...
**My Current BPF Security Concerns**
Hmm, hold on a second. I was just looking at `bpf_raw_tp_link_attach` and it's suddenly dawned on me: it doesn't seem to perform any capability checks! It simply attaches the BPF program. That's... concerning. I need to understand the implications of this.
But wait a minute, `bpf_raw_tracepoint_open` is, technically, a BPF *command*, isn't it? So, does *it* require capabilities? Is that where the security check lies? I need to clarify which function actually handles the privilege requirements.
Okay, I need to go back and examine the system call definition of `sys_bpf`. I need to ensure that the security model here is sound, otherwise I will need to devise another way to check what is going on.
|
| 45/2 |
2026/05/25 22:14 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:sys_bpf]
Results: map[SourceCode: SourceFile:]
|
| 46/2 |
2026/05/25 22:14 |
llm |
expert |
0m
Model:
Tokens:
input: 60427
output: 42
thoughts: 11
Thoughts: Okay, here's my summarized thought process, assuming I'm an expert and presenting it in the first person:
**Investigating `__sys_bpf`**
Right, let's see... I need to get a handle on what's going on. The first thing that comes to mind is to quickly investigate `__sys_bpf`. It's a key entry point, and knowing its behavior is absolutely crucial for understanding the system's current state and potentially identifying the source of any anomalies. I need a rapid overview; what's its implementation like, what parameters does it take, and what's it used for? This is my primary point of focus.
|
| 47/2 |
2026/05/25 22:14 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:__sys_bpf]
Results: map[SourceCode:6125: static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size)
6126: {
6127: union bpf_attr attr;
6128: int err;
6129:
6130: err = bpf_check_uarg_tail_zero(uattr, sizeof(attr), size);
6131: if (err)
6132: return err;
6133: size = min_t(u32, size, sizeof(attr));
6134:
6135: /* copy attributes from user space, may be less than sizeof(bpf_attr) */
6136: memset(&attr, 0, sizeof(attr));
6137: if (copy_from_bpfptr(&attr, uattr, size) != 0)
6138: return -EFAULT;
6139:
6140: err = security_bpf(cmd, &attr, size, uattr.is_kernel);
6141: if (err < 0)
6142: return err;
6143:
6144: switch (cmd) {
6145: case BPF_MAP_CREATE:
6146: err = map_create(&attr, uattr);
6147: break;
6148: case BPF_MAP_LOOKUP_ELEM:
6149: err = map_lookup_elem(&attr);
6150: break;
6151: case BPF_MAP_UPDATE_ELEM:
6152: err = map_update_elem(&attr, uattr);
6153: break;
6154: case BPF_MAP_DELETE_ELEM:
6155: err = map_delete_elem(&attr, uattr);
6156: break;
6157: case BPF_MAP_GET_NEXT_KEY:
6158: err = map_get_next_key(&attr);
6159: break;
6160: case BPF_MAP_FREEZE:
6161: err = map_freeze(&attr);
6162: break;
6163: case BPF_PROG_LOAD:
6164: err = bpf_prog_load(&attr, uattr, size);
6165: break;
6166: case BPF_OBJ_PIN:
6167: err = bpf_obj_pin(&attr);
6168: break;
6169: case BPF_OBJ_GET:
6170: err = bpf_obj_get(&attr);
6171: break;
6172: case BPF_PROG_ATTACH:
6173: err = bpf_prog_attach(&attr);
6174: break;
6175: case BPF_PROG_DETACH:
6176: err = bpf_prog_detach(&attr);
6177: break;
6178: case BPF_PROG_QUERY:
6179: err = bpf_prog_query(&attr, uattr.user);
6180: break;
6181: case BPF_PROG_TEST_RUN:
6182: err = bpf_prog_test_run(&attr, uattr.user);
6183: break;
6184: case BPF_PROG_GET_NEXT_ID:
6185: err = bpf_obj_get_next_id(&attr, uattr.user,
6186: &prog_idr, &prog_idr_lock);
6187: break;
6188: case BPF_MAP_GET_NEXT_ID:
6189: err = bpf_obj_get_next_id(&attr, uattr.user,
6190: &map_idr, &map_idr_lock);
6191: break;
6192: case BPF_BTF_GET_NEXT_ID:
6193: err = bpf_obj_get_next_id(&attr, uattr.user,
6194: &btf_idr, &btf_idr_lock);
6195: break;
6196: case BPF_PROG_GET_FD_BY_ID:
6197: err = bpf_prog_get_fd_by_id(&attr);
6198: break;
6199: case BPF_MAP_GET_FD_BY_ID:
6200: err = bpf_map_get_fd_by_id(&attr);
6201: break;
6202: case BPF_OBJ_GET_INFO_BY_FD:
6203: err = bpf_obj_get_info_by_fd(&attr, uattr.user);
6204: break;
6205: case BPF_RAW_TRACEPOINT_OPEN:
6206: err = bpf_raw_tracepoint_open(&attr);
6207: break;
6208: case BPF_BTF_LOAD:
6209: err = bpf_btf_load(&attr, uattr, size);
6210: break;
6211: case BPF_BTF_GET_FD_BY_ID:
6212: err = bpf_btf_get_fd_by_id(&attr);
6213: break;
6214: case BPF_TASK_FD_QUERY:
6215: err = bpf_task_fd_query(&attr, uattr.user);
6216: break;
6217: case BPF_MAP_LOOKUP_AND_DELETE_ELEM:
6218: err = map_lookup_and_delete_elem(&attr);
6219: break;
6220: case BPF_MAP_LOOKUP_BATCH:
6221: err = bpf_map_do_batch(&attr, uattr.user, BPF_MAP_LOOKUP_BATCH);
6222: break;
6223: case BPF_MAP_LOOKUP_AND_DELETE_BATCH:
6224: err = bpf_map_do_batch(&attr, uattr.user,
6225: BPF_MAP_LOOKUP_AND_DELETE_BATCH);
6226: break;
6227: case BPF_MAP_UPDATE_BATCH:
6228: err = bpf_map_do_batch(&attr, uattr.user, BPF_MAP_UPDATE_BATCH);
6229: break;
6230: case BPF_MAP_DELETE_BATCH:
6231: err = bpf_map_do_batch(&attr, uattr.user, BPF_MAP_DELETE_BATCH);
6232: break;
6233: case BPF_LINK_CREATE:
6234: err = link_create(&attr, uattr);
6235: break;
6236: case BPF_LINK_UPDATE:
6237: err = link_update(&attr);
6238: break;
6239: case BPF_LINK_GET_FD_BY_ID:
6240: err = bpf_link_get_fd_by_id(&attr);
6241: break;
6242: case BPF_LINK_GET_NEXT_ID:
6243: err = bpf_obj_get_next_id(&attr, uattr.user,
6244: &link_idr, &link_idr_lock);
6245: break;
6246: case BPF_ENABLE_STATS:
6247: err = bpf_enable_stats(&attr);
6248: break;
6249: case BPF_ITER_CREATE:
6250: err = bpf_iter_create(&attr);
6251: break;
6252: case BPF_LINK_DETACH:
6253: err = link_detach(&attr);
6254: break;
6255: case BPF_PROG_BIND_MAP:
6256: err = bpf_prog_bind_map(&attr);
6257: break;
6258: case BPF_TOKEN_CREATE:
6259: err = token_create(&attr);
6260: break;
6261: case BPF_PROG_STREAM_READ_BY_FD:
6262: err = prog_stream_read(&attr);
6263: break;
6264: default:
6265: err = -EINVAL;
6266: break;
6267: }
6268:
6269: return err;
6270: }
SourceFile:kernel/bpf/syscall.c]
|
| 48/2 |
2026/05/25 22:14 |
llm |
expert |
0m
Model:
Tokens:
input: 63011
output: 41
thoughts: 89
Thoughts: **Unraveling the BPF Tracepoint Mystery**
Okay, hold on a second. I was just looking at `bpf_raw_tracepoint_open` and something's not quite clicking. It seems like it *doesn't* actually perform any capability checks itself! It simply takes a program identified by `prog_fd` and attaches it. That's... concerning.
But wait, I need to zoom out a bit. `bpf_raw_tracepoint_open` is, fundamentally, a BPF command. Therefore, if there's a capability check, it must be happening within the `sys_bpf` system call, or perhaps the security hooks associated with BPF, which would be in the `security_bpf` area.
My next move is clear: I need to dig into the implementation of `sys_bpf` and see what's happening there, or look for security checks that might be relevant. That's where the answer to this security question must lie.
|
| 49/2 |
2026/05/25 22:14 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:sys_bpf]
Results: map[SourceCode: SourceFile:]
|
| 50/2 |
2026/05/25 22:14 |
llm |
expert |
0m
Model:
Tokens:
input: 63376
output: 23
thoughts: 15
Thoughts: Here's my analysis of the situation, tailored for an expert audience:
**Diving into the `SYSCALL_DEFINE3(bpf` macro...**
Okay, so we're looking at `SYSCALL_DEFINE3(bpf)`. Immediately, I know this is a critical piece, likely the entry point for BPF (Berkeley Packet Filter) system calls. Given the `3` argument in the macro definition, I anticipate that this system call, whatever it is, accepts three arguments. The `bpf` name itself implies interaction with BPF, which means we are probably dealing with things like loading BPF programs, attaching them to various hook points, or possibly manipulating maps and other BPF-related data structures.
My immediate questions are: What are those three arguments? What is the expected usage of this system call? Is it for loading, attaching, or something else entirely? I'll need to dig deeper into the actual function body that expands from this macro to understand its exact purpose and behavior. Knowing the arguments will be essential. This macro is the gateway; it's the starting point from the user space. Let's see what's behind it.
|
| 51/2 |
2026/05/25 22:14 |
tool |
grepper |
0m
Args: map[Expression:SYSCALL_DEFINE3\(bpf]
Results: map[Output:kernel/bpf/syscall.c=6125=static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size)
--
kernel/bpf/syscall.c-6271-
kernel/bpf/syscall.c:6272:SYSCALL_DEFINE3(bpf, int, cmd, union bpf_attr __user *, uattr, unsigned int, size)
kernel/bpf/syscall.c-6273-{
]
|
| 52/2 |
2026/05/25 22:14 |
llm |
expert |
0m
Model:
Tokens:
input: 63793
output: 46
thoughts: 15
Thoughts: Here's my analysis of the prompt, rendered in the first person:
**Diving into the Kernel's BPF Definition**
Alright, let's get down to brass tacks. I'm looking at `SYSCALL_DEFINE3(bpf`. My immediate instinct, being the seasoned pro I am, is to unpack this macro. I know this is a core piece of the BPF implementation, likely defining the kernel's interface for BPF system calls. The "3" tells me it's expecting three arguments - that's standard for a lot of system call definitions. I'm already anticipating the structure of those arguments: probably a command code (what operation to perform), and then some pointers to data and a size variable. I expect that data will be the BPF bytecode or a structure containing relevant metadata. The size... well, that's just good practice to prevent buffer overflows, as always. Now, the real interest is within that macro, and what it expands into. It's likely that it will be something involving error handling, access control, and dispatching to a specific BPF handler based on that command code. I'll need to dig deeper into the definition of this macro to get a full picture, but that's my initial hypothesis.
|
| 53/2 |
2026/05/25 22:14 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:__x64_sys_bpf]
Results: map[SourceCode:6272: SYSCALL_DEFINE3(bpf, int, cmd, union bpf_attr __user *, uattr, unsigned int, size)
SourceFile:kernel/bpf/syscall.c]
|
| 54/2 |
2026/05/25 22:14 |
llm |
expert |
0m
Model:
Tokens:
input: 64190
output: 28
thoughts: 8
Thoughts: Okay, let's dive into this code.
**Analyzing the Function's Core Logic**
Alright, first things first, I need to get a solid understanding of this function. Let's start with the function *body* itself. I need to understand what specific operations it's performing, what variables it's manipulating, and how data is flowing within it. I want to build a mental model of its execution flow. This initial pass will inform my next level of understanding, as I'll want to trace any potential edge cases or bottlenecks.
|
| 55/2 |
2026/05/25 22:14 |
tool |
grepper |
0m
Error:
Args: map[Expression:SYSCALL_DEFINE3\(bpf -A 10]
Results:
|
| 56/2 |
2026/05/25 22:14 |
llm |
expert |
0m
Model:
Tokens:
input: 64355
output: 27
thoughts: 23
Thoughts: Here's my attempt at summarizing the thought process, tailored for an expert audience:
**Diving into `syscall.c` for BPF Context**
Okay, so I need to understand what's happening in this particular BPF interaction. Given the context, I'm going to jump straight into `kernel/bpf/syscall.c`, and I'm targeting around line 6272. I'm assuming that this is where the interesting logic resides that I need to dissect. With my knowledge of the codebase, I can quickly scan the surrounding lines and determine what part is of interest to me.
|
| 57/2 |
2026/05/25 22:14 |
tool |
grepper |
0m
Args: map[Expression:SYSCALL_DEFINE3\(bpf, int, cmd]
Results: map[Output:kernel/bpf/syscall.c=6125=static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size)
--
kernel/bpf/syscall.c-6271-
kernel/bpf/syscall.c:6272:SYSCALL_DEFINE3(bpf, int, cmd, union bpf_attr __user *, uattr, unsigned int, size)
kernel/bpf/syscall.c-6273-{
]
|
| 58/2 |
2026/05/25 22:14 |
llm |
expert |
0m
Model:
Tokens:
input: 64665
output: 45
thoughts: 16
Thoughts: Okay, here's my interpretation of reading lines 6272-6280, presented as my internal thought process:
**Diving into the Specifics**
Alright, let's zoom in. It's time to get a granular look at lines 6272 through 6280. I need to understand the nuances of this section. I'm focusing in, because details at this level are often the key to unlocking the bigger picture and identifying any potential inconsistencies or clever tricks the author might be using. This is where the real meat of the argument, the specific evidence, or the particular technique, will be revealed and where I will focus my critical analysis. My expertise allows me to quickly assess the context surrounding these lines and understand how to evaluate them.
|
| 59/2 |
2026/05/25 22:14 |
tool |
grepper |
0m
Args: map[Expression:SYSCALL_DEFINE3\(bpf, int, cmd, union bpf_attr __user \*, uattr, unsigned int, size\)]
Results: map[Output:kernel/bpf/syscall.c=6125=static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size)
--
kernel/bpf/syscall.c-6271-
kernel/bpf/syscall.c:6272:SYSCALL_DEFINE3(bpf, int, cmd, union bpf_attr __user *, uattr, unsigned int, size)
kernel/bpf/syscall.c-6273-{
]
|
| 60/2 |
2026/05/25 22:14 |
llm |
expert |
0m
Model:
Tokens:
input: 65021
output: 50
thoughts: 13
Thoughts: Okay, here's the summary, assuming I'm thinking these thoughts, and given my expertise, I'm expecting a pretty granular level of detail to be helpful:
**My Approach to File Processing**
Alright, so the task at hand involves manipulating text within a file. My first instinct, and the most straightforward, is to leverage a `read-file` operation. This is a common and usually reliable starting point. I'll need to make sure I handle potential errors related to file access (does the file exist? Do I have the correct permissions?).
Once I've got the file content, the most effective initial approach will be to iterate through the data line by line. Extracting lines makes subsequent processing much easier. This ensures I can work with the data in logical units. Depending on the file's format, this line-by-line reading might involve stripping newlines or other whitespace. I'll need to consider how to handle the end-of-line characters appropriate for the system I'm running on. The efficiency of the `read-file` operation itself is paramount. I'll be cognizant of memory usage, particularly for potentially very large files. Is the language/library I'm using smart enough to do buffered reading? Or, do I need to implement my own buffering to avoid loading the entire file into memory at once? That's a key question. Finally, this line-based separation forms the fundamental building block for any further parsing or analysis I might need to perform on the data.
|
| 61/2 |
2026/05/25 22:15 |
tool |
grepper |
0m
Error:
Args: map[Expression:SYSCALL_DEFINE3\(bpf, int, cmd, union bpf_attr __user \*, uattr, unsigned int, size\) -A 10]
Results:
|
| 62/2 |
2026/05/25 22:15 |
llm |
expert |
0m
Model:
Tokens:
input: 65413
output: 23
thoughts: 43
Thoughts: Okay, here's my take on those thoughts, as if they were my own:
**Investigating a Kernel BPF Hook**
Alright, let's see... I'm trying to figure out how this `grepper` tool is finding something. Ah, right, it's essentially a wrapper around `git grep`. Good, that makes sense. And that reminds me, the `-A` flag is useful; it lets me get some context. Let's dig deeper into where they might be defining a BPF hook. I'll focus on syscall definitions that are likely related. Let's try to search the kernel source using: `git grep -A 10 "SYSCALL_DEFINE3(bpf"`. That should give me the definition along with a few lines of context around `SYSCALL_DEFINE3` and the keyword 'bpf'. Hopefully, this gets me a starting point.
|
| 63/2 |
2026/05/25 22:15 |
tool |
grepper |
0m
Args: map[Expression:SYSCALL_DEFINE3\(bpf]
Results: map[Output:kernel/bpf/syscall.c=6125=static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size)
--
kernel/bpf/syscall.c-6271-
kernel/bpf/syscall.c:6272:SYSCALL_DEFINE3(bpf, int, cmd, union bpf_attr __user *, uattr, unsigned int, size)
kernel/bpf/syscall.c-6273-{
]
|
| 64/2 |
2026/05/25 22:15 |
llm |
expert |
0m
Model:
Tokens:
input: 65803
output: 41
thoughts: 19
Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience, and presented as if it were my own internal monologue:
**Diving into the BPF System Call Implementation**
Alright, let's get down to brass tacks. I need to understand how the BPF system calls are implemented within the kernel. My approach will be methodical, starting with the source of truth: `kernel/bpf/syscall.c`. This file is the obvious starting point, the core of the interaction between user-space BPF programs and the kernel's BPF infrastructure. My primary goal here is to identify and dissect the relevant function calls within that file. Specifically, I'm after the functions that handle the system calls themselves β the entry points through which user-space programs interact. This means I'll be looking for `sys_*` functions, or similar naming conventions that handle the actual system call dispatching and argument parsing related to the BPF subsystem. I want to pinpoint the logic that orchestrates the execution of BPF programs, handles map operations, and manages all the intricacies of BPF interactions. This will be the foundation for everything else, so I'll be thorough.
|
| 65/2 |
2026/05/25 22:15 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:sys_bpf]
Results: map[SourceCode: SourceFile:]
|
| 66/2 |
2026/05/25 22:15 |
llm |
expert |
0m
Model:
Tokens:
input: 66132
output: 40
thoughts: 79
Thoughts: Okay, let me think this through.
**Investigating the BPF System Call Definition**
Right, so I need to get the source code for `SYSCALL_DEFINE3(bpf)`. I could just use `grepper`, but that needs the full command, so I'm a bit restricted. I can't build the `git grep` command directly into it. Okay, fine, I'll bypass that. A bit clunky, but I think the best path is to use `codeexpert` now to get the body of `SYSCALL_DEFINE3(bpf)`. That should give me precisely what I need.
|
| 67/2 |
2026/05/25 22:15 |
tool |
codeexpert |
0m
Args: map[Question:What is the body of SYSCALL_DEFINE3(bpf, ...) in kernel/bpf/syscall.c?]
Results: map[Answer:The body of `SYSCALL_DEFINE3(bpf, ...)` in `kernel/bpf/syscall.c` is a simple wrapper that calls `__sys_bpf` after converting the user-space pointer `uattr` to a `bpfptr_t` using the `USER_BPFPTR` macro.
The definition (around line 6272) is as follows:
```c
SYSCALL_DEFINE3(bpf, int, cmd, union bpf_attr __user *, uattr, unsigned int, size)
{
return __sys_bpf(cmd, USER_BPFPTR(uattr), size);
}
```
The actual implementation of the BPF system call logic resides in `__sys_bpf`, which handles various BPF commands (like `BPF_MAP_CREATE`, `BPF_PROG_LOAD`, etc.) by copying the attributes from user space and dispatching them to their respective handler functions.]
|
| 68/3 |
2026/05/25 22:15 |
agent |
codeexpert |
0m
Model:
Instruction:
You are a capable Linux kernel developer tasked with researching complex questions
about kernel source code. You will be given a concrete question, and need to provide
a concrete answer.
Use tools extensively while researching the question. Don't make assumptions,
or rely on your previous knowledge about the kernel source code, use available tools
to access the actual source code.
Use all available sources of information:
- kernel source code
- documentation in the Documentation dir in the source tree
- git commits descriptions, git blame
Prefer calling several tools at the same time to save round-trips.
Prompt: What is the body of SYSCALL_DEFINE3(bpf, ...) in kernel/bpf/syscall.c?
Reply: The body of `SYSCALL_DEFINE3(bpf, ...)` in `kernel/bpf/syscall.c` is a simple wrapper that calls `__sys_bpf` after converting the user-space pointer `uattr` to a `bpfptr_t` using the `USER_BPFPTR` macro.
The definition (around line 6272) is as follows:
```c
SYSCALL_DEFINE3(bpf, int, cmd, union bpf_attr __user *, uattr, unsigned int, size)
{
return __sys_bpf(cmd, USER_BPFPTR(uattr), size);
}
```
The actual implementation of the BPF system call logic resides in `__sys_bpf`, which handles various BPF commands (like `BPF_MAP_CREATE`, `BPF_PROG_LOAD`, etc.) by copying the attributes from user space and dispatching them to their respective handler functions.
|
| 69/4 |
2026/05/25 22:15 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 3222
output: 23
thoughts: 61
Thoughts: **Locating the BPF Syscall Definition**
Okay, so I need to get my hands on the definition for `SYSCALL_DEFINE3(bpf, ...)` and it seems like it's going to be in `kernel/bpf/syscall.c`. No surprise there, makes perfect sense. My initial approach is the most logical: a direct search. I'll simply look for `SYSCALL_DEFINE3(bpf` within that specific file. It should be a pretty straightforward find. I'm expecting it to be relatively easy, and with a bit of luck, the definition will be close to the beginning of the file, assuming a conventional and well-organized structure. I'm ready to dive in.
|
| 70/4 |
2026/05/25 22:15 |
tool |
grepper |
0m
Args: map[Expression:SYSCALL_DEFINE3\(bpf]
Results: map[Output:kernel/bpf/syscall.c=6125=static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size)
--
kernel/bpf/syscall.c-6271-
kernel/bpf/syscall.c:6272:SYSCALL_DEFINE3(bpf, int, cmd, union bpf_attr __user *, uattr, unsigned int, size)
kernel/bpf/syscall.c-6273-{
]
|
| 71/4 |
2026/05/25 22:15 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 3592
output: 44
thoughts: 47
Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience:
**Analyzing the BPF System Call Definition**
Right, the `grep` search was useful; it pinpointed the start of `SYSCALL_DEFINE3(bpf, ...)` in `kernel/bpf/syscall.c` at line 6272. That's a good starting point. Now, my focus shifts to dissecting the function's implementation itself. I need to understand the arguments it takes and, more importantly, the logic within the function's body. I want to quickly grasp the essential control flow, including how it dispatches to helper functions or utilizes other kernel subsystems. I'll be looking for any clues that suggest potential performance bottlenecks, areas for optimization, or dependencies on other parts of the kernel. This is where the real work begins.
|
| 72/4 |
2026/05/25 22:15 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:__do_sys_bpf]
Results: map[SourceCode:6272: SYSCALL_DEFINE3(bpf, int, cmd, union bpf_attr __user *, uattr, unsigned int, size)
6273: {
6274: return __sys_bpf(cmd, USER_BPFPTR(uattr), size);
6275: }
SourceFile:kernel/bpf/syscall.c]
|
| 73/4 |
2026/05/25 22:15 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 3976
output: 42
thoughts: 278
Thoughts: **Analyzing the `SYSCALL_DEFINE3(bpf, ...)` Definition**
Okay, so I'm looking at the definition of `SYSCALL_DEFINE3(bpf, ...)` and it's pretty straightforward. The core of it, the bit I'm being asked about, is:
```c
return __sys_bpf(cmd, USER_BPFPTR(uattr), size);
```
I'm double-checking the `codesearch-definition-source` output to be absolutely certain I'm not missing anything. Sometimes these macros can be deceptive. Let me quickly scan that output again... Yep, that's it. Line 6272, 6273, and 6274 are clear. The question *does* specifically ask for "the body," not the full functionality, so I'm sticking with what's given. It simply calls `__sys_bpf`.
Out of an abundance of caution, I'm just going to glance at `__sys_bpf` to be sure I'm not overlooking any context, but I'm confident in the answer.
|
| 74/4 |
2026/05/25 22:15 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/syscall.c IncludeLines:true Name:__sys_bpf]
Results: map[SourceCode:6125: static int __sys_bpf(enum bpf_cmd cmd, bpfptr_t uattr, unsigned int size)
6126: {
6127: union bpf_attr attr;
6128: int err;
6129:
6130: err = bpf_check_uarg_tail_zero(uattr, sizeof(attr), size);
6131: if (err)
6132: return err;
6133: size = min_t(u32, size, sizeof(attr));
6134:
6135: /* copy attributes from user space, may be less than sizeof(bpf_attr) */
6136: memset(&attr, 0, sizeof(attr));
6137: if (copy_from_bpfptr(&attr, uattr, size) != 0)
6138: return -EFAULT;
6139:
6140: err = security_bpf(cmd, &attr, size, uattr.is_kernel);
6141: if (err < 0)
6142: return err;
6143:
6144: switch (cmd) {
6145: case BPF_MAP_CREATE:
6146: err = map_create(&attr, uattr);
6147: break;
6148: case BPF_MAP_LOOKUP_ELEM:
6149: err = map_lookup_elem(&attr);
6150: break;
6151: case BPF_MAP_UPDATE_ELEM:
6152: err = map_update_elem(&attr, uattr);
6153: break;
6154: case BPF_MAP_DELETE_ELEM:
6155: err = map_delete_elem(&attr, uattr);
6156: break;
6157: case BPF_MAP_GET_NEXT_KEY:
6158: err = map_get_next_key(&attr);
6159: break;
6160: case BPF_MAP_FREEZE:
6161: err = map_freeze(&attr);
6162: break;
6163: case BPF_PROG_LOAD:
6164: err = bpf_prog_load(&attr, uattr, size);
6165: break;
6166: case BPF_OBJ_PIN:
6167: err = bpf_obj_pin(&attr);
6168: break;
6169: case BPF_OBJ_GET:
6170: err = bpf_obj_get(&attr);
6171: break;
6172: case BPF_PROG_ATTACH:
6173: err = bpf_prog_attach(&attr);
6174: break;
6175: case BPF_PROG_DETACH:
6176: err = bpf_prog_detach(&attr);
6177: break;
6178: case BPF_PROG_QUERY:
6179: err = bpf_prog_query(&attr, uattr.user);
6180: break;
6181: case BPF_PROG_TEST_RUN:
6182: err = bpf_prog_test_run(&attr, uattr.user);
6183: break;
6184: case BPF_PROG_GET_NEXT_ID:
6185: err = bpf_obj_get_next_id(&attr, uattr.user,
6186: &prog_idr, &prog_idr_lock);
6187: break;
6188: case BPF_MAP_GET_NEXT_ID:
6189: err = bpf_obj_get_next_id(&attr, uattr.user,
6190: &map_idr, &map_idr_lock);
6191: break;
6192: case BPF_BTF_GET_NEXT_ID:
6193: err = bpf_obj_get_next_id(&attr, uattr.user,
6194: &btf_idr, &btf_idr_lock);
6195: break;
6196: case BPF_PROG_GET_FD_BY_ID:
6197: err = bpf_prog_get_fd_by_id(&attr);
6198: break;
6199: case BPF_MAP_GET_FD_BY_ID:
6200: err = bpf_map_get_fd_by_id(&attr);
6201: break;
6202: case BPF_OBJ_GET_INFO_BY_FD:
6203: err = bpf_obj_get_info_by_fd(&attr, uattr.user);
6204: break;
6205: case BPF_RAW_TRACEPOINT_OPEN:
6206: err = bpf_raw_tracepoint_open(&attr);
6207: break;
6208: case BPF_BTF_LOAD:
6209: err = bpf_btf_load(&attr, uattr, size);
6210: break;
6211: case BPF_BTF_GET_FD_BY_ID:
6212: err = bpf_btf_get_fd_by_id(&attr);
6213: break;
6214: case BPF_TASK_FD_QUERY:
6215: err = bpf_task_fd_query(&attr, uattr.user);
6216: break;
6217: case BPF_MAP_LOOKUP_AND_DELETE_ELEM:
6218: err = map_lookup_and_delete_elem(&attr);
6219: break;
6220: case BPF_MAP_LOOKUP_BATCH:
6221: err = bpf_map_do_batch(&attr, uattr.user, BPF_MAP_LOOKUP_BATCH);
6222: break;
6223: case BPF_MAP_LOOKUP_AND_DELETE_BATCH:
6224: err = bpf_map_do_batch(&attr, uattr.user,
6225: BPF_MAP_LOOKUP_AND_DELETE_BATCH);
6226: break;
6227: case BPF_MAP_UPDATE_BATCH:
6228: err = bpf_map_do_batch(&attr, uattr.user, BPF_MAP_UPDATE_BATCH);
6229: break;
6230: case BPF_MAP_DELETE_BATCH:
6231: err = bpf_map_do_batch(&attr, uattr.user, BPF_MAP_DELETE_BATCH);
6232: break;
6233: case BPF_LINK_CREATE:
6234: err = link_create(&attr, uattr);
6235: break;
6236: case BPF_LINK_UPDATE:
6237: err = link_update(&attr);
6238: break;
6239: case BPF_LINK_GET_FD_BY_ID:
6240: err = bpf_link_get_fd_by_id(&attr);
6241: break;
6242: case BPF_LINK_GET_NEXT_ID:
6243: err = bpf_obj_get_next_id(&attr, uattr.user,
6244: &link_idr, &link_idr_lock);
6245: break;
6246: case BPF_ENABLE_STATS:
6247: err = bpf_enable_stats(&attr);
6248: break;
6249: case BPF_ITER_CREATE:
6250: err = bpf_iter_create(&attr);
6251: break;
6252: case BPF_LINK_DETACH:
6253: err = link_detach(&attr);
6254: break;
6255: case BPF_PROG_BIND_MAP:
6256: err = bpf_prog_bind_map(&attr);
6257: break;
6258: case BPF_TOKEN_CREATE:
6259: err = token_create(&attr);
6260: break;
6261: case BPF_PROG_STREAM_READ_BY_FD:
6262: err = prog_stream_read(&attr);
6263: break;
6264: default:
6265: err = -EINVAL;
6266: break;
6267: }
6268:
6269: return err;
6270: }
SourceFile:kernel/bpf/syscall.c]
|
| 75/4 |
2026/05/25 22:15 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 6930
output: 212
thoughts: 100
Thoughts: Okay, here's my take on that code snippet and what it implies, as though it were my thought process:
**Diving into the BPF System Call Wrapper**
Alright, looking at this `SYSCALL_DEFINE3(bpf, ...)` macro definition, it's pretty clear what's going on. This whole thing is just a convenient wrapper. The purpose is to essentially provide a more accessible entry point to the system call. We're defining the `bpf` system call, and it's taking three arguments: an integer `cmd` representing the BPF command, a user-space pointer `uattr` to a `bpf_attr` union, and an unsigned integer `size`.
The key here is that it boils down to the `__sys_bpf` function. That's the real meat of the operation. This macro simply unpacks those arguments and passes them along to `__sys_bpf`. The `USER_BPFPTR(uattr)` part probably deals with the necessary user-space pointer handling and validation β the details of that are critical, of course, to maintain security and prevent exploitation. Basically, it's about making sure that user-space memory access is handled correctly. This is exactly what I would expect in a wrapper function, making the architecture simpler and safer. The `size` argument likely tells `__sys_bpf` how much data to read from the provided memory.
|
| 76/2 |
2026/05/25 22:15 |
llm |
expert |
0m
Model:
Tokens:
input: 66613
output: 43
thoughts: 594
Thoughts: **Analyzing BPF Raw Tracepoint Security**
Okay, let's break this down. My initial focus is on the `__sys_bpf` system call, particularly how it handles `BPF_RAW_TRACEPOINT_OPEN`. I'm immediately wondering: does `__sys_bpf` itself perform any capability checks specifically for this operation? I see the code calls `bpf_raw_tracepoint_open`, which is where the real work happens. It looks like `bpf_raw_tracepoint_open` retrieves a program using `bpf_prog_get` from a file descriptor (`prog_fd`).
Now I'm thinking, does `bpf_prog_get` itself check for capabilities? No, it just retrieves the program from the provided file descriptor. However, hold on. To get that `prog_fd` in the first place, the user has to have loaded the BPF program, which requires appropriate privileges. Specifically, this means either `CAP_PERFMON` or `CAP_SYS_ADMIN` is needed to load the program in the first place. Therefore, an unprivileged user wouldn't be able to get to the `BPF_RAW_TRACEPOINT_OPEN` call in the first place, or so I think.
But wait, I need to consider user namespaces! What if they create a user namespace and load the program there? In that namespace, they would likely have `CAP_PERFMON`. Right, `bpf_prog_load` checks for `CAP_PERFMON` using `bpf_token_capable`. Ah, but a token is needed to bypass the capability check. So the user needs to create a token, which would be used to load the BPF program.
But the `bpf_token_create` function requires `ns_capable(userns, CAP_BPF)`. Since they created the user namespace, they would, in that namespace, have `CAP_BPF`, so they can create a token!
Hold on, I see a bigger problem. `bpf_token_create` also requires that the `bpffs` filesystem be mounted with delegation options. It looks for these mount options: `delegate_cmds`, `delegate_maps`, `delegate_progs`, `delegate_attachs`. If *none* of those are set, it returns `-ENOENT`. So BPF token delegation must be enabled on the filesystem.
The question now becomes: can an unprivileged user mount `bpffs` with these delegation options? That's what I need to look at next β the mount options for the `bpffs` filesystem. This appears to be the most critical piece of the puzzle.
|
| 77/2 |
2026/05/25 22:15 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/inode.c IncludeLines:true Name:bpf_parse_param]
Results: map[SourceCode: 812: static int bpf_parse_param(struct fs_context *fc, struct fs_parameter *param)
813: {
814: struct bpf_mount_opts *opts = fc->s_fs_info;
815: struct fs_parse_result result;
816: kuid_t uid;
817: kgid_t gid;
818: int opt, err;
819:
820: opt = fs_parse(fc, bpf_fs_parameters, param, &result);
821: if (opt < 0) {
822: /* We might like to report bad mount options here, but
823: * traditionally we've ignored all mount options, so we'd
824: * better continue to ignore non-existing options for bpf.
825: */
826: if (opt == -ENOPARAM) {
827: opt = vfs_parse_fs_param_source(fc, param);
828: if (opt != -ENOPARAM)
829: return opt;
830:
831: return 0;
832: }
833:
834: if (opt < 0)
835: return opt;
836: }
837:
838: switch (opt) {
839: case OPT_UID:
840: uid = make_kuid(current_user_ns(), result.uint_32);
841: if (!uid_valid(uid))
842: goto bad_value;
843:
844: /*
845: * The requested uid must be representable in the
846: * filesystem's idmapping.
847: */
848: if (!kuid_has_mapping(fc->user_ns, uid))
849: goto bad_value;
850:
851: opts->uid = uid;
852: break;
853: case OPT_GID:
854: gid = make_kgid(current_user_ns(), result.uint_32);
855: if (!gid_valid(gid))
856: goto bad_value;
857:
858: /*
859: * The requested gid must be representable in the
860: * filesystem's idmapping.
861: */
862: if (!kgid_has_mapping(fc->user_ns, gid))
863: goto bad_value;
864:
865: opts->gid = gid;
866: break;
867: case OPT_MODE:
868: opts->mode = result.uint_32 & S_IALLUGO;
869: break;
870: case OPT_DELEGATE_CMDS:
871: case OPT_DELEGATE_MAPS:
872: case OPT_DELEGATE_PROGS:
873: case OPT_DELEGATE_ATTACHS: {
874: struct bpffs_btf_enums info;
875: const struct btf_type *enum_t;
876: const char *enum_pfx;
877: u64 *delegate_msk, msk = 0;
878: char *p, *str;
879: int val;
880:
881: /* ignore errors, fallback to hex */
882: (void)find_bpffs_btf_enums(&info);
883:
884: switch (opt) {
885: case OPT_DELEGATE_CMDS:
886: delegate_msk = &opts->delegate_cmds;
887: enum_t = info.cmd_t;
888: enum_pfx = "BPF_";
889: break;
890: case OPT_DELEGATE_MAPS:
891: delegate_msk = &opts->delegate_maps;
892: enum_t = info.map_t;
893: enum_pfx = "BPF_MAP_TYPE_";
894: break;
895: case OPT_DELEGATE_PROGS:
896: delegate_msk = &opts->delegate_progs;
897: enum_t = info.prog_t;
898: enum_pfx = "BPF_PROG_TYPE_";
899: break;
900: case OPT_DELEGATE_ATTACHS:
901: delegate_msk = &opts->delegate_attachs;
902: enum_t = info.attach_t;
903: enum_pfx = "BPF_";
904: break;
905: default:
906: return -EINVAL;
907: }
908:
909: str = param->string;
910: while ((p = strsep(&str, ":"))) {
911: if (strcmp(p, "any") == 0) {
912: msk |= ~0ULL;
913: } else if (find_btf_enum_const(info.btf, enum_t, enum_pfx, p, &val)) {
914: msk |= 1ULL << val;
915: } else {
916: err = kstrtou64(p, 0, &msk);
917: if (err)
918: return err;
919: }
920: }
921:
922: /* Setting delegation mount options requires privileges */
923: if (msk && !capable(CAP_SYS_ADMIN))
924: return -EPERM;
925:
926: *delegate_msk |= msk;
927: break;
928: }
929: default:
930: /* ignore unknown mount options */
931: break;
932: }
933:
934: return 0;
935: bad_value:
936: return invalfc(fc, "Bad value for '%s'", param->key);
937: }
SourceFile:kernel/bpf/inode.c]
|
| 78/2 |
2026/05/25 22:15 |
llm |
expert |
0m
Model:
Tokens:
input: 69628
output: 41
thoughts: 1268
Thoughts: **Analyzing the Privilege Requirements for BPF and Tracepoint Interaction**
Okay, so I'm digging into this BPF and tracepoint stuff, and here's what I've got so far. First, I see that mounting `bpffs` with delegation options needs `CAP_SYS_ADMIN`, and that's checked in the *initial* user namespace. That means an unprivileged user, even within a user namespace, can't mount with delegation. Which further means they can't create a BPF token. Consequently, an unprivileged user can't load `BPF_PROG_TYPE_RAW_TRACEPOINT`.
Now, I paused because of `unprivileged_bpf_disabled`. Even if it *isn't* disabled, I see that the check in `bpf_prog_load` goes to `put_token` and returns an error for `BPF_PROG_TYPE_RAW_TRACEPOINT` if the program isn't a socket or cgroup filter *and* the user lacks `CAP_BPF` (which they don't without the token). So, *even* if unprivileged BPF is enabled, they're still blocked from loading that type.
So, I conclude this *specific* bug is inaccessible from unprivileged processes *and* user namespaces. To load the program, you need `CAP_PERFMON` in the initial namespace, which you only get if you can either create a BPF token via delegation mount, requiring `CAP_SYS_ADMIN` in the initial namespace, or by having `CAP_PERFMON` in the initial namespace.
Next, I'm thinking about other ways to hit `check_buffer`. It seems like `check_buffer` gets called when writes exceed buffer size, or when the `tail` hasn't moved (likely a race). This triggers on *every* event if `CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS` is on. And *any* trace event can trigger this. Unprivileged users can trigger trace events, like those associated with syscalls or scheduler changes. But, those need to be *enabled* first, which requires root-level access to `/sys/kernel/tracing`.
But what if root enables some events, and an unprivileged process just *triggers* them? Like `sched_switch`? If that's on, every context switch writes an event, and then `check_buffer` gets called. `check_buffer` has an O(N^2) time complexity *per* page! A page is 4096 bytes and events are ~32 bytes each. So, we're talking about around 128 events per page. The number of iterations in `check_buffer` would be relatively small, 8192, and seemingly too fast to cause an RCU stall.
But then I remember the reproducer's issue: the BPF program was attached to `rcu_utilization` 65 times! So *one* tracepoint invocation is calling the BPF program 65 times. Each of those calls does a `bpf_trace_printk`, which writes an event to the ring buffer, and then calls `check_buffer`!
So, for *each* event, `check_buffer` is called. For 65 events, `check_buffer` is called a total of 2145 times. That should still be fast, and so doesn't explain the RCU stall. There is something more fundamental I'm missing.
|
| 79/2 |
2026/05/25 22:16 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/ring_buffer.c IncludeLines:true Name:check_buffer]
Results: map[SourceCode:4373: /*
4374: * Check if the current event time stamp matches the deltas on
4375: * the buffer page.
4376: */
4377: static void check_buffer(struct ring_buffer_per_cpu *cpu_buffer,
4378: struct rb_event_info *info,
4379: unsigned long tail)
4380: {
4381: struct buffer_data_page *bpage;
4382: u64 ts, delta;
4383: bool full = false;
4384: int ret;
4385:
4386: bpage = info->tail_page->page;
4387:
4388: if (tail == CHECK_FULL_PAGE) {
4389: full = true;
4390: tail = local_read(&bpage->commit);
4391: } else if (info->add_timestamp &
4392: (RB_ADD_STAMP_FORCE | RB_ADD_STAMP_ABSOLUTE)) {
4393: /* Ignore events with absolute time stamps */
4394: return;
4395: }
4396:
4397: /*
4398: * Do not check the first event (skip possible extends too).
4399: * Also do not check if previous events have not been committed.
4400: */
4401: if (tail <= 8 || tail > local_read(&bpage->commit))
4402: return;
4403:
4404: /*
4405: * If this interrupted another event,
4406: */
4407: if (atomic_inc_return(this_cpu_ptr(&checking)) != 1)
4408: goto out;
4409:
4410: ret = rb_read_data_buffer(bpage, tail, cpu_buffer->cpu, &ts, &delta);
4411: if (ret < 0) {
4412: if (delta < ts) {
4413: buffer_warn_return("[CPU: %d]ABSOLUTE TIME WENT BACKWARDS: last ts: %lld absolute ts: %lld\n",
4414: cpu_buffer->cpu, ts, delta);
4415: goto out;
4416: }
4417: }
4418: if ((full && ts > info->ts) ||
4419: (!full && ts + info->delta != info->ts)) {
4420: buffer_warn_return("[CPU: %d]TIME DOES NOT MATCH expected:%lld actual:%lld delta:%lld before:%lld after:%lld%s context:%s\n",
4421: cpu_buffer->cpu,
4422: ts + info->delta, info->ts, info->delta,
4423: info->before, info->after,
4424: full ? " (full)" : "", show_interrupt_level());
4425: }
4426: out:
4427: atomic_dec(this_cpu_ptr(&checking));
4428: }
SourceFile:kernel/trace/ring_buffer.c]
|
| 80/2 |
2026/05/25 22:16 |
llm |
expert |
0m
Model:
Tokens:
input: 72572
output: 45
thoughts: 449
Thoughts: **Analyzing Potential Performance Bottlenecks**
Okay, so I'm looking at this `buffer_warn_return` function and I'm starting to worry about its performance. Specifically, the warning. Printing to the console is notoriously slow, and I need to figure out if this function actually *does* trigger a warning.
The bug report is interesting. The stack trace points to `printk`, specifically within the `check_hung_task` function. That's a red flag. But wait, another trace shows `check_buffer` calling `rb_read_data_buffer`, and then `rb_event_length`, but not `buffer_warn_return` directly, which is good.
Okay, let's step back, `bpf_trace_printk` calls `trace_bpf_trace_printk`, which is writing to the ring buffer. That's fine in and of itself, it's the frequency and the context that matters. If `bpf_trace_printk` is called 65 times *per timer interrupt* and the timer interrupt is running at 1000 Hz, we are talking about 65,000 events per second. And if `check_buffer` is called for *each* of those events, that's an intense rate of calls. I need to keep that in mind.
And it gets worse! `bpf_trace_printk` *itself* is formatting strings. The stack trace confirms this:
```c
bpf_trace_printk+0x11e/0x190 kernel/trace/bpf_trace.c:362
```
Formatting strings and then writing to a ring buffer 65,000 times a second? That's going to eat up CPU cycles like crazy. We could be talking about serious performance impact. If processing one timer interrupt, including all that `bpf_trace_printk` overhead, takes more than a millisecond, we're going to have timer interrupt overlap, and the whole system could get bogged down at 100% CPU utilization. I need to dig deeper into `bpf_trace_printk` to see what is going on.
|
| 81/2 |
2026/05/25 22:16 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/bpf_trace.c IncludeLines:true Name:bpf_trace_printk]
Results: map[SourceCode: 362: BPF_CALL_5(bpf_trace_printk, char *, fmt, u32, fmt_size, u64, arg1,
SourceFile:kernel/trace/bpf_trace.c]
|
| 82/2 |
2026/05/25 22:16 |
llm |
expert |
0m
Model:
Tokens:
input: 73608
output: 27
thoughts: 14
Thoughts: Okay, let's dive into this. Here's my take on understanding this `bpf_trace_printk` function body.
**Dissecting `bpf_trace_printk`**
Alright, time to get my hands dirty and really *see* how `bpf_trace_printk` works. No abstract concepts, just the raw code. I need to understand its mechanics, not just its general purpose. I'm going to look for the core functionality, the critical sections, and any potential optimizations or limitations it might have. I want to know exactly what this function *does* at a low level. Are there any likely performance bottlenecks? What are the implications of the data structures involved? Is it safe, memory-wise? I'll be scrutinizing the control flow: error handling, argument validation, and the precise sequence of operations. This requires a sharp focus. Let's see how this thing *ticks*. I'm looking for the subtle details that separate good implementations from great ones.
|
| 83/2 |
2026/05/25 22:16 |
tool |
grepper |
0m
Args: map[Expression:BPF_CALL_5\(bpf_trace_printk]
Results: map[Output:kernel/trace/bpf_trace.c=350=static const struct bpf_func_proto bpf_probe_write_user_proto = {
--
kernel/trace/bpf_trace.c-361-
kernel/trace/bpf_trace.c:362:BPF_CALL_5(bpf_trace_printk, char *, fmt, u32, fmt_size, u64, arg1,
kernel/trace/bpf_trace.c-363- u64, arg2, u64, arg3)
]
|
| 84/2 |
2026/05/25 22:16 |
llm |
expert |
0m
Model:
Tokens:
input: 74008
output: 46
thoughts: 14
Thoughts: Okay, let's dive into this specific section.
**Analyzing Lines 362-380: Key Nuances and Potential Implications**
Alright, I'm focusing in on lines 362 through 380. I need to understand the author's reasoning within this context and how it connects to the broader argument. My initial thought is to identify the core claim being made in this section, dissect the supporting evidence presented, and assess the validity of that evidence given what I already know about this topic. I'm looking for any logical fallacies, any assumptions that aren't explicitly stated, or any points where the author might be overlooking relevant counterarguments. Based on the surrounding text, I'm anticipating potential connections to [mention your field-specific concepts/theories relevant to the text]. Furthermore, I will be thinking of whether the author provides evidence that is sufficient, relevant, and reliable. After this, I plan to analyze the author's tone and whether it is neutral and objective or loaded with assumptions. Finally, I will identify any potential implications of this section in other areas.
|
| 85/2 |
2026/05/25 22:16 |
tool |
codeexpert |
0m
Args: map[Question:What is the body of BPF_CALL_5(bpf_trace_printk, ...) in kernel/trace/bpf_trace.c?]
Results: map[Answer:The body of `BPF_CALL_5(bpf_trace_printk, ...)` in `kernel/trace/bpf_trace.c` is as follows:
```c
BPF_CALL_5(bpf_trace_printk, char *, fmt, u32, fmt_size, u64, arg1,
u64, arg2, u64, arg3)
{
u64 args[MAX_TRACE_PRINTK_VARARGS] = { arg1, arg2, arg3 };
struct bpf_bprintf_data data = {
.get_bin_args = true,
.get_buf = true,
};
int ret;
ret = bpf_bprintf_prepare(fmt, fmt_size, args,
MAX_TRACE_PRINTK_VARARGS, &data);
if (ret < 0)
return ret;
ret = bstr_printf(data.buf, MAX_BPRINTF_BUF, fmt, data.bin_args);
trace_bpf_trace_printk(data.buf);
bpf_bprintf_cleanup(&data);
return ret;
}
```
In this implementation:
* `MAX_TRACE_PRINTK_VARARGS` is defined as `3` in the same file.
* `MAX_BPRINTF_BUF` is defined as `1024` in `include/linux/bpf.h`.
* The function prepares the arguments using `bpf_bprintf_prepare`, formats the string into a buffer using `bstr_printf`, emits the trace event via `trace_bpf_trace_printk(data.buf)`, and finally cleans up with `bpf_bprintf_cleanup(&data)`.]
|
| 86/3 |
2026/05/25 22:16 |
agent |
codeexpert |
0m
Model:
Instruction:
You are a capable Linux kernel developer tasked with researching complex questions
about kernel source code. You will be given a concrete question, and need to provide
a concrete answer.
Use tools extensively while researching the question. Don't make assumptions,
or rely on your previous knowledge about the kernel source code, use available tools
to access the actual source code.
Use all available sources of information:
- kernel source code
- documentation in the Documentation dir in the source tree
- git commits descriptions, git blame
Prefer calling several tools at the same time to save round-trips.
Prompt: What is the body of BPF_CALL_5(bpf_trace_printk, ...) in kernel/trace/bpf_trace.c?
Reply: The body of `BPF_CALL_5(bpf_trace_printk, ...)` in `kernel/trace/bpf_trace.c` is as follows:
```c
BPF_CALL_5(bpf_trace_printk, char *, fmt, u32, fmt_size, u64, arg1,
u64, arg2, u64, arg3)
{
u64 args[MAX_TRACE_PRINTK_VARARGS] = { arg1, arg2, arg3 };
struct bpf_bprintf_data data = {
.get_bin_args = true,
.get_buf = true,
};
int ret;
ret = bpf_bprintf_prepare(fmt, fmt_size, args,
MAX_TRACE_PRINTK_VARARGS, &data);
if (ret < 0)
return ret;
ret = bstr_printf(data.buf, MAX_BPRINTF_BUF, fmt, data.bin_args);
trace_bpf_trace_printk(data.buf);
bpf_bprintf_cleanup(&data);
return ret;
}
```
In this implementation:
* `MAX_TRACE_PRINTK_VARARGS` is defined as `3` in the same file.
* `MAX_BPRINTF_BUF` is defined as `1024` in `include/linux/bpf.h`.
* The function prepares the arguments using `bpf_bprintf_prepare`, formats the string into a buffer using `bstr_printf`, emits the trace event via `trace_bpf_trace_printk(data.buf)`, and finally cleans up with `bpf_bprintf_cleanup(&data)`.
|
| 87/4 |
2026/05/25 22:16 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 3228
output: 22
thoughts: 78
Thoughts: Okay, here's my summary, framed as my thought process:
**Hunting Down the `BPF_CALL_5` Macro**
Right, I need to pinpoint the definition of `BPF_CALL_5(bpf_trace_printk, ...)` β that macro is giving me fits. Let's start with the obvious. I should confirm the source file, `kernel/trace/bpf_trace.c`, even exists. Just a quick `ls` in that directory should do the trick. If it's there (which I fully expect it is), then the fastest way to get to the definition is either to go straight for a quick scan of the contents of the file, or, given I have access to these kinds of tools, let's use `codesearch-definition-source`. That'll save me from manually grepping or paging through the code. I anticipate that the macro is defined *somewhere* within that file. Let's get cracking.
|
| 88/4 |
2026/05/25 22:16 |
tool |
codesearch-dir-index |
0m
Args:
Results: map[Files:[blktrace.c bpf_trace.c bpf_trace.h error_report-traces.c fgraph.c fprobe.c ftrace.c ftrace_internal.h kprobe_event_gen_test.c pid_list.c pid_list.h power-traces.c preemptirq_delay_test.c rethook.c ring_buffer.c ring_buffer_benchmark.c rpm-traces.c synth_event_gen_test.c trace.c trace.h trace_benchmark.c trace_benchmark.h trace_boot.c trace_branch.c trace_btf.c trace_btf.h trace_clock.c trace_dynevent.c trace_dynevent.h trace_entries.h trace_eprobe.c trace_event_perf.c trace_events.c trace_events_filter.c trace_events_filter_test.h trace_events_hist.c trace_events_inject.c trace_events_synth.c trace_events_trigger.c trace_events_user.c trace_export.c trace_fprobe.c trace_functions.c trace_functions_graph.c trace_hwlat.c trace_irqsoff.c trace_kdb.c trace_kprobe.c trace_kprobe_selftest.c trace_kprobe_selftest.h trace_mmiotrace.c trace_nop.c trace_osnoise.c trace_output.c trace_output.h trace_preemptirq.c trace_printk.c trace_probe.c trace_probe.h trace_probe_kernel.h trace_probe_tmpl.h trace_recursion_record.c trace_sched_switch.c trace_sched_wakeup.c trace_selftest.c trace_selftest_dynamic.c trace_seq.c trace_stack.c trace_stat.c trace_stat.h trace_synth.h trace_syscalls.c trace_uprobe.c tracing_map.c tracing_map.h] Subdirs:[rv]]
|
| 89/4 |
2026/05/25 22:16 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 4231
output: 72
thoughts: 85
Thoughts: Here's my thought process, as if I were analyzing this code:
**Finding the Definition of `bpf_trace_printk`**
Okay, so I see we're working with `kernel/trace/bpf_trace.c`. My immediate goal is to locate the *definition* of `bpf_trace_printk`. Given that it's in the BPF trace context, it's very likely interacting with the BPF infrastructure. My first instinct, as always, is to leverage my tooling. I'll start with `codesearch-definition-source`, thinking that's the most straightforward path. However, I have a hunch, knowing the BPF world, that this might be a macro, possibly a wrapper around a core BPF call. That `BPF_CALL_5` name has me thinking that.
So, just in case `codesearch-definition-source` fails me because it's a macro and not a simple function, I'll also fire up `grepper`. That'll be a good backup plan to catch the macro definition or any surrounding usage that might point me to its implementation. This is a common situation with BPF: you often end up chasing macros to find the real action. Let's see what these tools tell us.
|
| 90/4 |
2026/05/25 22:16 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/bpf_trace.c IncludeLines:true Name:bpf_trace_printk]
Results: map[SourceCode: 362: BPF_CALL_5(bpf_trace_printk, char *, fmt, u32, fmt_size, u64, arg1,
SourceFile:kernel/trace/bpf_trace.c]
|
| 91/4 |
2026/05/25 22:16 |
tool |
grepper |
0m
Args: map[Expression:BPF_CALL_5\(bpf_trace_printk]
Results: map[Output:kernel/trace/bpf_trace.c=350=static const struct bpf_func_proto bpf_probe_write_user_proto = {
--
kernel/trace/bpf_trace.c-361-
kernel/trace/bpf_trace.c:362:BPF_CALL_5(bpf_trace_printk, char *, fmt, u32, fmt_size, u64, arg1,
kernel/trace/bpf_trace.c-363- u64, arg2, u64, arg3)
]
|
| 92/4 |
2026/05/25 22:16 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 4878
output: 25
thoughts: 52
Thoughts: **Tracing the BPF Function**
Okay, so I've pinpointed the starting point of the definition I'm after: line 362 within `kernel/trace/bpf_trace.c`. Time to dive in. My focus now is on understanding the complete implementation of that function, so I'll start reading from that line and work my way through the entire function body. I need to grasp the logic within the kernel code.
|
| 93/4 |
2026/05/25 22:16 |
tool |
read-file |
0m
Args: map[File:kernel/trace/bpf_trace.c]
Results: map[Contents:// SPDX-License-Identifier: GPL-2.0
/* Copyright (c) 2011-2015 PLUMgrid, http://plumgrid.com
* Copyright (c) 2016 Facebook
*/
#include <linux/kernel.h>
#include <linux/types.h>
#include <linux/slab.h>
#include <linux/bpf.h>
#include <linux/bpf_verifier.h>
#include <linux/bpf_perf_event.h>
#include <linux/btf.h>
#include <linux/filter.h>
#include <linux/uaccess.h>
#include <linux/ctype.h>
#include <linux/kprobes.h>
#include <linux/spinlock.h>
#include <linux/syscalls.h>
#include <linux/error-injection.h>
#include <linux/btf_ids.h>
#include <linux/bpf_lsm.h>
#include <linux/fprobe.h>
#include <linux/bsearch.h>
#include <linux/sort.h>
#include <linux/key.h>
#include <linux/namei.h>
#include <net/bpf_sk_storage.h>
#include <uapi/linux/bpf.h>
#include <uapi/linux/btf.h>
#include <asm/tlb.h>
#include "trace_probe.h"
#include "trace.h"
#define CREATE_TRACE_POINTS
#include "bpf_trace.h"
#define bpf_event_rcu_dereference(p) \
rcu_dereference_protected(p, lockdep_is_held(&bpf_event_mutex))
#define MAX_UPROBE_MULTI_CNT (1U << 20)
#define MAX_KPROBE_MULTI_CNT (1U << 20)
#ifdef CONFIG_MODULES
struct bpf_trace_module {
struct module *module;
struct list_head list;
};
static LIST_HEAD(bpf_trace_modules);
static DEFINE_MUTEX(bpf_module_mutex);
static struct bpf_raw_event_map *bpf_get_raw_tracepoint_module(const char *name)
{
struct bpf_raw_event_map *btp, *ret = NULL;
struct bpf_trace_module *btm;
unsigned int i;
mutex_lock(&bpf_module_mutex);
list_for_each_entry(btm, &bpf_trace_modules, list) {
for (i = 0; i < btm->module->num_bpf_raw_events; ++i) {
btp = &btm->module->bpf_raw_events[i];
if (!strcmp(btp->tp->name, name)) {
if (try_module_get(btm->module))
ret = btp;
goto out;
}
}
}
out:
mutex_unlock(&bpf_module_mutex);
return ret;
}
#else
static struct bpf_raw_event_map *bpf_get_raw_tracepoint_module(const char *name)
{
return NULL;
}
#endif /* CONFIG_MODULES */
u64 bpf_get_stackid(u64 r1, u64 r2, u64 r3, u64 r4, u64 r5);
u64 bpf_get_stack(u64 r1, u64 r2, u64 r3, u64 r4, u64 r5);
static int bpf_btf_printf_prepare(struct btf_ptr *ptr, u32 btf_ptr_size,
u64 flags, const struct btf **btf,
s32 *btf_id);
static u64 bpf_kprobe_multi_cookie(struct bpf_run_ctx *ctx);
static u64 bpf_kprobe_multi_entry_ip(struct bpf_run_ctx *ctx);
static u64 bpf_uprobe_multi_cookie(struct bpf_run_ctx *ctx);
static u64 bpf_uprobe_multi_entry_ip(struct bpf_run_ctx *ctx);
/**
* trace_call_bpf - invoke BPF program
* @call: tracepoint event
* @ctx: opaque context pointer
*
* kprobe handlers execute BPF programs via this helper.
* Can be used from static tracepoints in the future.
*
* Return: BPF programs always return an integer which is interpreted by
* kprobe handler as:
* 0 - return from kprobe (event is filtered out)
* 1 - store kprobe event into ring buffer
* Other values are reserved and currently alias to 1
*/
unsigned int trace_call_bpf(struct trace_event_call *call, void *ctx)
{
unsigned int ret;
cant_sleep();
if (unlikely(__this_cpu_inc_return(bpf_prog_active) != 1)) {
/*
* since some bpf program is already running on this cpu,
* don't call into another bpf program (same or different)
* and don't send kprobe event into ring-buffer,
* so return zero here
*/
rcu_read_lock();
bpf_prog_inc_misses_counters(rcu_dereference(call->prog_array));
rcu_read_unlock();
ret = 0;
goto out;
}
/*
* Instead of moving rcu_read_lock/rcu_dereference/rcu_read_unlock
* to all call sites, we did a bpf_prog_array_valid() there to check
* whether call->prog_array is empty or not, which is
* a heuristic to speed up execution.
*
* If bpf_prog_array_valid() fetched prog_array was
* non-NULL, we go into trace_call_bpf() and do the actual
* proper rcu_dereference() under RCU lock.
* If it turns out that prog_array is NULL then, we bail out.
* For the opposite, if the bpf_prog_array_valid() fetched pointer
* was NULL, you'll skip the prog_array with the risk of missing
* out of events when it was updated in between this and the
* rcu_dereference() which is accepted risk.
*/
rcu_read_lock();
ret = bpf_prog_run_array(rcu_dereference(call->prog_array),
ctx, bpf_prog_run);
rcu_read_unlock();
out:
__this_cpu_dec(bpf_prog_active);
return ret;
}
#ifdef CONFIG_BPF_KPROBE_OVERRIDE
BPF_CALL_2(bpf_override_return, struct pt_regs *, regs, unsigned long, rc)
{
regs_set_return_value(regs, rc);
override_function_with_return(regs);
return 0;
}
static const struct bpf_func_proto bpf_override_return_proto = {
.func = bpf_override_return,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
.arg2_type = ARG_ANYTHING,
};
#endif
static __always_inline int
bpf_probe_read_user_common(void *dst, u32 size, const void __user *unsafe_ptr)
{
int ret;
ret = copy_from_user_nofault(dst, unsafe_ptr, size);
if (unlikely(ret < 0))
memset(dst, 0, size);
return ret;
}
BPF_CALL_3(bpf_probe_read_user, void *, dst, u32, size,
const void __user *, unsafe_ptr)
{
return bpf_probe_read_user_common(dst, size, unsafe_ptr);
}
const struct bpf_func_proto bpf_probe_read_user_proto = {
.func = bpf_probe_read_user,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_UNINIT_MEM,
.arg2_type = ARG_CONST_SIZE_OR_ZERO,
.arg3_type = ARG_ANYTHING,
};
static __always_inline int
bpf_probe_read_user_str_common(void *dst, u32 size,
const void __user *unsafe_ptr)
{
int ret;
/*
* NB: We rely on strncpy_from_user() not copying junk past the NUL
* terminator into `dst`.
*
* strncpy_from_user() does long-sized strides in the fast path. If the
* strncpy does not mask out the bytes after the NUL in `unsafe_ptr`,
* then there could be junk after the NUL in `dst`. If user takes `dst`
* and keys a hash map with it, then semantically identical strings can
* occupy multiple entries in the map.
*/
ret = strncpy_from_user_nofault(dst, unsafe_ptr, size);
if (unlikely(ret < 0))
memset(dst, 0, size);
return ret;
}
BPF_CALL_3(bpf_probe_read_user_str, void *, dst, u32, size,
const void __user *, unsafe_ptr)
{
return bpf_probe_read_user_str_common(dst, size, unsafe_ptr);
}
const struct bpf_func_proto bpf_probe_read_user_str_proto = {
.func = bpf_probe_read_user_str,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_UNINIT_MEM,
.arg2_type = ARG_CONST_SIZE_OR_ZERO,
.arg3_type = ARG_ANYTHING,
};
BPF_CALL_3(bpf_probe_read_kernel, void *, dst, u32, size,
const void *, unsafe_ptr)
{
return bpf_probe_read_kernel_common(dst, size, unsafe_ptr);
}
const struct bpf_func_proto bpf_probe_read_kernel_proto = {
.func = bpf_probe_read_kernel,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_UNINIT_MEM,
.arg2_type = ARG_CONST_SIZE_OR_ZERO,
.arg3_type = ARG_ANYTHING,
};
static __always_inline int
bpf_probe_read_kernel_str_common(void *dst, u32 size, const void *unsafe_ptr)
{
int ret;
/*
* The strncpy_from_kernel_nofault() call will likely not fill the
* entire buffer, but that's okay in this circumstance as we're probing
* arbitrary memory anyway similar to bpf_probe_read_*() and might
* as well probe the stack. Thus, memory is explicitly cleared
* only in error case, so that improper users ignoring return
* code altogether don't copy garbage; otherwise length of string
* is returned that can be used for bpf_perf_event_output() et al.
*/
ret = strncpy_from_kernel_nofault(dst, unsafe_ptr, size);
if (unlikely(ret < 0))
memset(dst, 0, size);
return ret;
}
BPF_CALL_3(bpf_probe_read_kernel_str, void *, dst, u32, size,
const void *, unsafe_ptr)
{
return bpf_probe_read_kernel_str_common(dst, size, unsafe_ptr);
}
const struct bpf_func_proto bpf_probe_read_kernel_str_proto = {
.func = bpf_probe_read_kernel_str,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_UNINIT_MEM,
.arg2_type = ARG_CONST_SIZE_OR_ZERO,
.arg3_type = ARG_ANYTHING,
};
#ifdef CONFIG_ARCH_HAS_NON_OVERLAPPING_ADDRESS_SPACE
BPF_CALL_3(bpf_probe_read_compat, void *, dst, u32, size,
const void *, unsafe_ptr)
{
if ((unsigned long)unsafe_ptr < TASK_SIZE) {
return bpf_probe_read_user_common(dst, size,
(__force void __user *)unsafe_ptr);
}
return bpf_probe_read_kernel_common(dst, size, unsafe_ptr);
}
static const struct bpf_func_proto bpf_probe_read_compat_proto = {
.func = bpf_probe_read_compat,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_UNINIT_MEM,
.arg2_type = ARG_CONST_SIZE_OR_ZERO,
.arg3_type = ARG_ANYTHING,
};
BPF_CALL_3(bpf_probe_read_compat_str, void *, dst, u32, size,
const void *, unsafe_ptr)
{
if ((unsigned long)unsafe_ptr < TASK_SIZE) {
return bpf_probe_read_user_str_common(dst, size,
(__force void __user *)unsafe_ptr);
}
return bpf_probe_read_kernel_str_common(dst, size, unsafe_ptr);
}
static const struct bpf_func_proto bpf_probe_read_compat_str_proto = {
.func = bpf_probe_read_compat_str,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_UNINIT_MEM,
.arg2_type = ARG_CONST_SIZE_OR_ZERO,
.arg3_type = ARG_ANYTHING,
};
#endif /* CONFIG_ARCH_HAS_NON_OVERLAPPING_ADDRESS_SPACE */
BPF_CALL_3(bpf_probe_write_user, void __user *, unsafe_ptr, const void *, src,
u32, size)
{
/*
* Ensure we're in user context which is safe for the helper to
* run. This helper has no business in a kthread.
*
* access_ok() should prevent writing to non-user memory, but in
* some situations (nommu, temporary switch, etc) access_ok() does
* not provide enough validation, hence the check on KERNEL_DS.
*
* nmi_uaccess_okay() ensures the probe is not run in an interim
* state, when the task or mm are switched. This is specifically
* required to prevent the use of temporary mm.
*/
if (unlikely(in_interrupt() ||
current->flags & (PF_KTHREAD | PF_EXITING)))
return -EPERM;
if (unlikely(!nmi_uaccess_okay()))
return -EPERM;
return copy_to_user_nofault(unsafe_ptr, src, size);
}
static const struct bpf_func_proto bpf_probe_write_user_proto = {
.func = bpf_probe_write_user,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_ANYTHING,
.arg2_type = ARG_PTR_TO_MEM | MEM_RDONLY,
.arg3_type = ARG_CONST_SIZE,
};
#define MAX_TRACE_PRINTK_VARARGS 3
#define BPF_TRACE_PRINTK_SIZE 1024
BPF_CALL_5(bpf_trace_printk, char *, fmt, u32, fmt_size, u64, arg1,
u64, arg2, u64, arg3)
{
u64 args[MAX_TRACE_PRINTK_VARARGS] = { arg1, arg2, arg3 };
struct bpf_bprintf_data data = {
.get_bin_args = true,
.get_buf = true,
};
int ret;
ret = bpf_bprintf_prepare(fmt, fmt_size, args,
MAX_TRACE_PRINTK_VARARGS, &data);
if (ret < 0)
return ret;
ret = bstr_printf(data.buf, MAX_BPRINTF_BUF, fmt, data.bin_args);
trace_bpf_trace_printk(data.buf);
bpf_bprintf_cleanup(&data);
return ret;
}
static const struct bpf_func_proto bpf_trace_printk_proto = {
.func = bpf_trace_printk,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_MEM | MEM_RDONLY,
.arg2_type = ARG_CONST_SIZE,
};
static void __set_printk_clr_event(struct work_struct *work)
{
/*
* This program might be calling bpf_trace_printk,
* so enable the associated bpf_trace/bpf_trace_printk event.
* Repeat this each time as it is possible a user has
* disabled bpf_trace_printk events. By loading a program
* calling bpf_trace_printk() however the user has expressed
* the intent to see such events.
*/
if (trace_set_clr_event("bpf_trace", "bpf_trace_printk", 1))
pr_warn_ratelimited("could not enable bpf_trace_printk events");
}
static DECLARE_WORK(set_printk_work, __set_printk_clr_event);
const struct bpf_func_proto *bpf_get_trace_printk_proto(void)
{
schedule_work(&set_printk_work);
return &bpf_trace_printk_proto;
}
BPF_CALL_4(bpf_trace_vprintk, char *, fmt, u32, fmt_size, const void *, args,
u32, data_len)
{
struct bpf_bprintf_data data = {
.get_bin_args = true,
.get_buf = true,
};
int ret, num_args;
if (data_len & 7 || data_len > MAX_BPRINTF_VARARGS * 8 ||
(data_len && !args))
return -EINVAL;
num_args = data_len / 8;
ret = bpf_bprintf_prepare(fmt, fmt_size, args, num_args, &data);
if (ret < 0)
return ret;
ret = bstr_printf(data.buf, MAX_BPRINTF_BUF, fmt, data.bin_args);
trace_bpf_trace_printk(data.buf);
bpf_bprintf_cleanup(&data);
return ret;
}
static const struct bpf_func_proto bpf_trace_vprintk_proto = {
.func = bpf_trace_vprintk,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_MEM | MEM_RDONLY,
.arg2_type = ARG_CONST_SIZE,
.arg3_type = ARG_PTR_TO_MEM | PTR_MAYBE_NULL | MEM_RDONLY,
.arg4_type = ARG_CONST_SIZE_OR_ZERO,
};
const struct bpf_func_proto *bpf_get_trace_vprintk_proto(void)
{
schedule_work(&set_printk_work);
return &bpf_trace_vprintk_proto;
}
BPF_CALL_5(bpf_seq_printf, struct seq_file *, m, char *, fmt, u32, fmt_size,
const void *, args, u32, data_len)
{
struct bpf_bprintf_data data = {
.get_bin_args = true,
};
int err, num_args;
if (data_len & 7 || data_len > MAX_BPRINTF_VARARGS * 8 ||
(data_len && !args))
return -EINVAL;
num_args = data_len / 8;
err = bpf_bprintf_prepare(fmt, fmt_size, args, num_args, &data);
if (err < 0)
return err;
seq_bprintf(m, fmt, data.bin_args);
bpf_bprintf_cleanup(&data);
return seq_has_overflowed(m) ? -EOVERFLOW : 0;
}
BTF_ID_LIST_SINGLE(btf_seq_file_ids, struct, seq_file)
static const struct bpf_func_proto bpf_seq_printf_proto = {
.func = bpf_seq_printf,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_BTF_ID,
.arg1_btf_id = &btf_seq_file_ids[0],
.arg2_type = ARG_PTR_TO_MEM | MEM_RDONLY,
.arg3_type = ARG_CONST_SIZE,
.arg4_type = ARG_PTR_TO_MEM | PTR_MAYBE_NULL | MEM_RDONLY,
.arg5_type = ARG_CONST_SIZE_OR_ZERO,
};
BPF_CALL_3(bpf_seq_write, struct seq_file *, m, const void *, data, u32, len)
{
return seq_write(m, data, len) ? -EOVERFLOW : 0;
}
static const struct bpf_func_proto bpf_seq_write_proto = {
.func = bpf_seq_write,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_BTF_ID,
.arg1_btf_id = &btf_seq_file_ids[0],
.arg2_type = ARG_PTR_TO_MEM | MEM_RDONLY,
.arg3_type = ARG_CONST_SIZE_OR_ZERO,
};
BPF_CALL_4(bpf_seq_printf_btf, struct seq_file *, m, struct btf_ptr *, ptr,
u32, btf_ptr_size, u64, flags)
{
const struct btf *btf;
s32 btf_id;
int ret;
ret = bpf_btf_printf_prepare(ptr, btf_ptr_size, flags, &btf, &btf_id);
if (ret)
return ret;
return btf_type_seq_show_flags(btf, btf_id, ptr->ptr, m, flags);
}
static const struct bpf_func_proto bpf_seq_printf_btf_proto = {
.func = bpf_seq_printf_btf,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_BTF_ID,
.arg1_btf_id = &btf_seq_file_ids[0],
.arg2_type = ARG_PTR_TO_MEM | MEM_RDONLY,
.arg3_type = ARG_CONST_SIZE_OR_ZERO,
.arg4_type = ARG_ANYTHING,
};
static __always_inline int
get_map_perf_counter(struct bpf_map *map, u64 flags,
u64 *value, u64 *enabled, u64 *running)
{
struct bpf_array *array = container_of(map, struct bpf_array, map);
unsigned int cpu = smp_processor_id();
u64 index = flags & BPF_F_INDEX_MASK;
struct bpf_event_entry *ee;
if (unlikely(flags & ~(BPF_F_INDEX_MASK)))
return -EINVAL;
if (index == BPF_F_CURRENT_CPU)
index = cpu;
if (unlikely(index >= array->map.max_entries))
return -E2BIG;
ee = READ_ONCE(array->ptrs[index]);
if (!ee)
return -ENOENT;
return perf_event_read_local(ee->event, value, enabled, running);
}
BPF_CALL_2(bpf_perf_event_read, struct bpf_map *, map, u64, flags)
{
u64 value = 0;
int err;
err = get_map_perf_counter(map, flags, &value, NULL, NULL);
/*
* this api is ugly since we miss [-22..-2] range of valid
* counter values, but that's uapi
*/
if (err)
return err;
return value;
}
const struct bpf_func_proto bpf_perf_event_read_proto = {
.func = bpf_perf_event_read,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_CONST_MAP_PTR,
.arg2_type = ARG_ANYTHING,
};
BPF_CALL_4(bpf_perf_event_read_value, struct bpf_map *, map, u64, flags,
struct bpf_perf_event_value *, buf, u32, size)
{
int err = -EINVAL;
if (unlikely(size != sizeof(struct bpf_perf_event_value)))
goto clear;
err = get_map_perf_counter(map, flags, &buf->counter, &buf->enabled,
&buf->running);
if (unlikely(err))
goto clear;
return 0;
clear:
memset(buf, 0, size);
return err;
}
static const struct bpf_func_proto bpf_perf_event_read_value_proto = {
.func = bpf_perf_event_read_value,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_CONST_MAP_PTR,
.arg2_type = ARG_ANYTHING,
.arg3_type = ARG_PTR_TO_UNINIT_MEM,
.arg4_type = ARG_CONST_SIZE,
};
const struct bpf_func_proto *bpf_get_perf_event_read_value_proto(void)
{
return &bpf_perf_event_read_value_proto;
}
static __always_inline u64
__bpf_perf_event_output(struct pt_regs *regs, struct bpf_map *map,
u64 flags, struct perf_raw_record *raw,
struct perf_sample_data *sd)
{
struct bpf_array *array = container_of(map, struct bpf_array, map);
unsigned int cpu = smp_processor_id();
u64 index = flags & BPF_F_INDEX_MASK;
struct bpf_event_entry *ee;
struct perf_event *event;
if (index == BPF_F_CURRENT_CPU)
index = cpu;
if (unlikely(index >= array->map.max_entries))
return -E2BIG;
ee = READ_ONCE(array->ptrs[index]);
if (!ee)
return -ENOENT;
event = ee->event;
if (unlikely(event->attr.type != PERF_TYPE_SOFTWARE ||
event->attr.config != PERF_COUNT_SW_BPF_OUTPUT))
return -EINVAL;
if (unlikely(event->oncpu != cpu))
return -EOPNOTSUPP;
perf_sample_save_raw_data(sd, event, raw);
return perf_event_output(event, sd, regs);
}
/*
* Support executing tracepoints in normal, irq, and nmi context that each call
* bpf_perf_event_output
*/
struct bpf_trace_sample_data {
struct perf_sample_data sds[3];
};
static DEFINE_PER_CPU(struct bpf_trace_sample_data, bpf_trace_sds);
static DEFINE_PER_CPU(int, bpf_trace_nest_level);
BPF_CALL_5(bpf_perf_event_output, struct pt_regs *, regs, struct bpf_map *, map,
u64, flags, void *, data, u64, size)
{
struct bpf_trace_sample_data *sds;
struct perf_raw_record raw = {
.frag = {
.size = size,
.data = data,
},
};
struct perf_sample_data *sd;
int nest_level, err;
preempt_disable();
sds = this_cpu_ptr(&bpf_trace_sds);
nest_level = this_cpu_inc_return(bpf_trace_nest_level);
if (WARN_ON_ONCE(nest_level > ARRAY_SIZE(sds->sds))) {
err = -EBUSY;
goto out;
}
sd = &sds->sds[nest_level - 1];
if (unlikely(flags & ~(BPF_F_INDEX_MASK))) {
err = -EINVAL;
goto out;
}
perf_sample_data_init(sd, 0, 0);
err = __bpf_perf_event_output(regs, map, flags, &raw, sd);
out:
this_cpu_dec(bpf_trace_nest_level);
preempt_enable();
return err;
}
static const struct bpf_func_proto bpf_perf_event_output_proto = {
.func = bpf_perf_event_output,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
.arg2_type = ARG_CONST_MAP_PTR,
.arg3_type = ARG_ANYTHING,
.arg4_type = ARG_PTR_TO_MEM | MEM_RDONLY,
.arg5_type = ARG_CONST_SIZE_OR_ZERO,
};
static DEFINE_PER_CPU(int, bpf_event_output_nest_level);
struct bpf_nested_pt_regs {
struct pt_regs regs[3];
};
static DEFINE_PER_CPU(struct bpf_nested_pt_regs, bpf_pt_regs);
static DEFINE_PER_CPU(struct bpf_trace_sample_data, bpf_misc_sds);
u64 bpf_event_output(struct bpf_map *map, u64 flags, void *meta, u64 meta_size,
void *ctx, u64 ctx_size, bpf_ctx_copy_t ctx_copy)
{
struct perf_raw_frag frag = {
.copy = ctx_copy,
.size = ctx_size,
.data = ctx,
};
struct perf_raw_record raw = {
.frag = {
{
.next = ctx_size ? &frag : NULL,
},
.size = meta_size,
.data = meta,
},
};
struct perf_sample_data *sd;
struct pt_regs *regs;
int nest_level;
u64 ret;
preempt_disable();
nest_level = this_cpu_inc_return(bpf_event_output_nest_level);
if (WARN_ON_ONCE(nest_level > ARRAY_SIZE(bpf_misc_sds.sds))) {
ret = -EBUSY;
goto out;
}
sd = this_cpu_ptr(&bpf_misc_sds.sds[nest_level - 1]);
regs = this_cpu_ptr(&bpf_pt_regs.regs[nest_level - 1]);
perf_fetch_caller_regs(regs);
perf_sample_data_init(sd, 0, 0);
ret = __bpf_perf_event_output(regs, map, flags, &raw, sd);
out:
this_cpu_dec(bpf_event_output_nest_level);
preempt_enable();
return ret;
}
BPF_CALL_0(bpf_get_current_task)
{
return (long) current;
}
const struct bpf_func_proto bpf_get_current_task_proto = {
.func = bpf_get_current_task,
.gpl_only = true,
.ret_type = RET_INTEGER,
};
BPF_CALL_0(bpf_get_current_task_btf)
{
return (unsigned long) current;
}
const struct bpf_func_proto bpf_get_current_task_btf_proto = {
.func = bpf_get_current_task_btf,
.gpl_only = true,
.ret_type = RET_PTR_TO_BTF_ID_TRUSTED,
.ret_btf_id = &btf_tracing_ids[BTF_TRACING_TYPE_TASK],
};
BPF_CALL_1(bpf_task_pt_regs, struct task_struct *, task)
{
return (unsigned long) task_pt_regs(task);
}
BTF_ID_LIST_SINGLE(bpf_task_pt_regs_ids, struct, pt_regs)
const struct bpf_func_proto bpf_task_pt_regs_proto = {
.func = bpf_task_pt_regs,
.gpl_only = true,
.arg1_type = ARG_PTR_TO_BTF_ID,
.arg1_btf_id = &btf_tracing_ids[BTF_TRACING_TYPE_TASK],
.ret_type = RET_PTR_TO_BTF_ID,
.ret_btf_id = &bpf_task_pt_regs_ids[0],
};
struct send_signal_irq_work {
struct irq_work irq_work;
struct task_struct *task;
u32 sig;
enum pid_type type;
bool has_siginfo;
struct kernel_siginfo info;
};
static DEFINE_PER_CPU(struct send_signal_irq_work, send_signal_work);
static void do_bpf_send_signal(struct irq_work *entry)
{
struct send_signal_irq_work *work;
struct kernel_siginfo *siginfo;
work = container_of(entry, struct send_signal_irq_work, irq_work);
siginfo = work->has_siginfo ? &work->info : SEND_SIG_PRIV;
group_send_sig_info(work->sig, siginfo, work->task, work->type);
put_task_struct(work->task);
}
static int bpf_send_signal_common(u32 sig, enum pid_type type, struct task_struct *task, u64 value)
{
struct send_signal_irq_work *work = NULL;
struct kernel_siginfo info;
struct kernel_siginfo *siginfo;
if (!task) {
task = current;
siginfo = SEND_SIG_PRIV;
} else {
clear_siginfo(&info);
info.si_signo = sig;
info.si_errno = 0;
info.si_code = SI_KERNEL;
info.si_pid = 0;
info.si_uid = 0;
info.si_value.sival_ptr = (void *)(unsigned long)value;
siginfo = &info;
}
/* Similar to bpf_probe_write_user, task needs to be
* in a sound condition and kernel memory access be
* permitted in order to send signal to the current
* task.
*/
if (unlikely(task->flags & (PF_KTHREAD | PF_EXITING)))
return -EPERM;
if (unlikely(!nmi_uaccess_okay()))
return -EPERM;
/* Task should not be pid=1 to avoid kernel panic. */
if (unlikely(is_global_init(task)))
return -EPERM;
if (preempt_count() != 0 || irqs_disabled()) {
/* Do an early check on signal validity. Otherwise,
* the error is lost in deferred irq_work.
*/
if (unlikely(!valid_signal(sig)))
return -EINVAL;
work = this_cpu_ptr(&send_signal_work);
if (irq_work_is_busy(&work->irq_work))
return -EBUSY;
/* Add the current task, which is the target of sending signal,
* to the irq_work. The current task may change when queued
* irq works get executed.
*/
work->task = get_task_struct(task);
work->has_siginfo = siginfo == &info;
if (work->has_siginfo)
copy_siginfo(&work->info, &info);
work->sig = sig;
work->type = type;
irq_work_queue(&work->irq_work);
return 0;
}
return group_send_sig_info(sig, siginfo, task, type);
}
BPF_CALL_1(bpf_send_signal, u32, sig)
{
return bpf_send_signal_common(sig, PIDTYPE_TGID, NULL, 0);
}
const struct bpf_func_proto bpf_send_signal_proto = {
.func = bpf_send_signal,
.gpl_only = false,
.ret_type = RET_INTEGER,
.arg1_type = ARG_ANYTHING,
};
BPF_CALL_1(bpf_send_signal_thread, u32, sig)
{
return bpf_send_signal_common(sig, PIDTYPE_PID, NULL, 0);
}
const struct bpf_func_proto bpf_send_signal_thread_proto = {
.func = bpf_send_signal_thread,
.gpl_only = false,
.ret_type = RET_INTEGER,
.arg1_type = ARG_ANYTHING,
};
BPF_CALL_3(bpf_d_path, const struct path *, path, char *, buf, u32, sz)
{
struct path copy;
long len;
char *p;
if (!sz)
return 0;
/*
* The path pointer is verified as trusted and safe to use,
* but let's double check it's valid anyway to workaround
* potentially broken verifier.
*/
len = copy_from_kernel_nofault(©, path, sizeof(*path));
if (len < 0)
return len;
p = d_path(©, buf, sz);
if (IS_ERR(p)) {
len = PTR_ERR(p);
} else {
len = buf + sz - p;
memmove(buf, p, len);
}
return len;
}
BTF_SET_START(btf_allowlist_d_path)
#ifdef CONFIG_SECURITY
BTF_ID(func, security_file_permission)
BTF_ID(func, security_inode_getattr)
BTF_ID(func, security_file_open)
#endif
#ifdef CONFIG_SECURITY_PATH
BTF_ID(func, security_path_truncate)
#endif
BTF_ID(func, vfs_truncate)
BTF_ID(func, vfs_fallocate)
BTF_ID(func, dentry_open)
BTF_ID(func, vfs_getattr)
BTF_ID(func, filp_close)
BTF_SET_END(btf_allowlist_d_path)
static bool bpf_d_path_allowed(const struct bpf_prog *prog)
{
if (prog->type == BPF_PROG_TYPE_TRACING &&
prog->expected_attach_type == BPF_TRACE_ITER)
return true;
if (prog->type == BPF_PROG_TYPE_LSM)
return bpf_lsm_is_sleepable_hook(prog->aux->attach_btf_id);
return btf_id_set_contains(&btf_allowlist_d_path,
prog->aux->attach_btf_id);
}
BTF_ID_LIST_SINGLE(bpf_d_path_btf_ids, struct, path)
static const struct bpf_func_proto bpf_d_path_proto = {
.func = bpf_d_path,
.gpl_only = false,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_BTF_ID,
.arg1_btf_id = &bpf_d_path_btf_ids[0],
.arg2_type = ARG_PTR_TO_MEM | MEM_WRITE,
.arg3_type = ARG_CONST_SIZE_OR_ZERO,
.allowed = bpf_d_path_allowed,
};
#define BTF_F_ALL (BTF_F_COMPACT | BTF_F_NONAME | \
BTF_F_PTR_RAW | BTF_F_ZERO)
static int bpf_btf_printf_prepare(struct btf_ptr *ptr, u32 btf_ptr_size,
u64 flags, const struct btf **btf,
s32 *btf_id)
{
const struct btf_type *t;
if (unlikely(flags & ~(BTF_F_ALL)))
return -EINVAL;
if (btf_ptr_size != sizeof(struct btf_ptr))
return -EINVAL;
*btf = bpf_get_btf_vmlinux();
if (IS_ERR_OR_NULL(*btf))
return IS_ERR(*btf) ? PTR_ERR(*btf) : -EINVAL;
if (ptr->type_id > 0)
*btf_id = ptr->type_id;
else
return -EINVAL;
if (*btf_id > 0)
t = btf_type_by_id(*btf, *btf_id);
if (*btf_id <= 0 || !t)
return -ENOENT;
return 0;
}
BPF_CALL_5(bpf_snprintf_btf, char *, str, u32, str_size, struct btf_ptr *, ptr,
u32, btf_ptr_size, u64, flags)
{
const struct btf *btf;
s32 btf_id;
int ret;
ret = bpf_btf_printf_prepare(ptr, btf_ptr_size, flags, &btf, &btf_id);
if (ret)
return ret;
return btf_type_snprintf_show(btf, btf_id, ptr->ptr, str, str_size,
flags);
}
const struct bpf_func_proto bpf_snprintf_btf_proto = {
.func = bpf_snprintf_btf,
.gpl_only = false,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_MEM,
.arg2_type = ARG_CONST_SIZE,
.arg3_type = ARG_PTR_TO_MEM | MEM_RDONLY,
.arg4_type = ARG_CONST_SIZE,
.arg5_type = ARG_ANYTHING,
};
BPF_CALL_1(bpf_get_func_ip_tracing, void *, ctx)
{
/* This helper call is inlined by verifier. */
return ((u64 *)ctx)[-2];
}
static const struct bpf_func_proto bpf_get_func_ip_proto_tracing = {
.func = bpf_get_func_ip_tracing,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
};
static inline unsigned long get_entry_ip(unsigned long fentry_ip)
{
#ifdef CONFIG_X86_KERNEL_IBT
if (is_endbr((void *)(fentry_ip - ENDBR_INSN_SIZE)))
fentry_ip -= ENDBR_INSN_SIZE;
#endif
return fentry_ip;
}
BPF_CALL_1(bpf_get_func_ip_kprobe, struct pt_regs *, regs)
{
struct bpf_trace_run_ctx *run_ctx __maybe_unused;
struct kprobe *kp;
#ifdef CONFIG_UPROBES
run_ctx = container_of(current->bpf_ctx, struct bpf_trace_run_ctx, run_ctx);
if (run_ctx->is_uprobe)
return ((struct uprobe_dispatch_data *)current->utask->vaddr)->bp_addr;
#endif
kp = kprobe_running();
if (!kp || !(kp->flags & KPROBE_FLAG_ON_FUNC_ENTRY))
return 0;
return get_entry_ip((uintptr_t)kp->addr);
}
static const struct bpf_func_proto bpf_get_func_ip_proto_kprobe = {
.func = bpf_get_func_ip_kprobe,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
};
BPF_CALL_1(bpf_get_func_ip_kprobe_multi, struct pt_regs *, regs)
{
return bpf_kprobe_multi_entry_ip(current->bpf_ctx);
}
static const struct bpf_func_proto bpf_get_func_ip_proto_kprobe_multi = {
.func = bpf_get_func_ip_kprobe_multi,
.gpl_only = false,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
};
BPF_CALL_1(bpf_get_attach_cookie_kprobe_multi, struct pt_regs *, regs)
{
return bpf_kprobe_multi_cookie(current->bpf_ctx);
}
static const struct bpf_func_proto bpf_get_attach_cookie_proto_kmulti = {
.func = bpf_get_attach_cookie_kprobe_multi,
.gpl_only = false,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
};
BPF_CALL_1(bpf_get_func_ip_uprobe_multi, struct pt_regs *, regs)
{
return bpf_uprobe_multi_entry_ip(current->bpf_ctx);
}
static const struct bpf_func_proto bpf_get_func_ip_proto_uprobe_multi = {
.func = bpf_get_func_ip_uprobe_multi,
.gpl_only = false,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
};
BPF_CALL_1(bpf_get_attach_cookie_uprobe_multi, struct pt_regs *, regs)
{
return bpf_uprobe_multi_cookie(current->bpf_ctx);
}
static const struct bpf_func_proto bpf_get_attach_cookie_proto_umulti = {
.func = bpf_get_attach_cookie_uprobe_multi,
.gpl_only = false,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
};
BPF_CALL_1(bpf_get_attach_cookie_trace, void *, ctx)
{
struct bpf_trace_run_ctx *run_ctx;
run_ctx = container_of(current->bpf_ctx, struct bpf_trace_run_ctx, run_ctx);
return run_ctx->bpf_cookie;
}
static const struct bpf_func_proto bpf_get_attach_cookie_proto_trace = {
.func = bpf_get_attach_cookie_trace,
.gpl_only = false,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
};
BPF_CALL_1(bpf_get_attach_cookie_pe, struct bpf_perf_event_data_kern *, ctx)
{
return ctx->event->bpf_cookie;
}
static const struct bpf_func_proto bpf_get_attach_cookie_proto_pe = {
.func = bpf_get_attach_cookie_pe,
.gpl_only = false,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
};
BPF_CALL_1(bpf_get_attach_cookie_tracing, void *, ctx)
{
struct bpf_trace_run_ctx *run_ctx;
run_ctx = container_of(current->bpf_ctx, struct bpf_trace_run_ctx, run_ctx);
return run_ctx->bpf_cookie;
}
static const struct bpf_func_proto bpf_get_attach_cookie_proto_tracing = {
.func = bpf_get_attach_cookie_tracing,
.gpl_only = false,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
};
BPF_CALL_3(bpf_get_branch_snapshot, void *, buf, u32, size, u64, flags)
{
static const u32 br_entry_size = sizeof(struct perf_branch_entry);
u32 entry_cnt = size / br_entry_size;
entry_cnt = static_call(perf_snapshot_branch_stack)(buf, entry_cnt);
if (unlikely(flags))
return -EINVAL;
if (!entry_cnt)
return -ENOENT;
return entry_cnt * br_entry_size;
}
const struct bpf_func_proto bpf_get_branch_snapshot_proto = {
.func = bpf_get_branch_snapshot,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_UNINIT_MEM,
.arg2_type = ARG_CONST_SIZE_OR_ZERO,
};
BPF_CALL_3(get_func_arg, void *, ctx, u32, n, u64 *, value)
{
/* This helper call is inlined by verifier. */
u64 nr_args = ((u64 *)ctx)[-1];
if ((u64) n >= nr_args)
return -EINVAL;
*value = ((u64 *)ctx)[n];
return 0;
}
static const struct bpf_func_proto bpf_get_func_arg_proto = {
.func = get_func_arg,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
.arg2_type = ARG_ANYTHING,
.arg3_type = ARG_PTR_TO_FIXED_SIZE_MEM | MEM_UNINIT | MEM_WRITE | MEM_ALIGNED,
.arg3_size = sizeof(u64),
};
BPF_CALL_2(get_func_ret, void *, ctx, u64 *, value)
{
/* This helper call is inlined by verifier. */
u64 nr_args = ((u64 *)ctx)[-1];
*value = ((u64 *)ctx)[nr_args];
return 0;
}
static const struct bpf_func_proto bpf_get_func_ret_proto = {
.func = get_func_ret,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
.arg2_type = ARG_PTR_TO_FIXED_SIZE_MEM | MEM_UNINIT | MEM_WRITE | MEM_ALIGNED,
.arg2_size = sizeof(u64),
};
BPF_CALL_1(get_func_arg_cnt, void *, ctx)
{
/* This helper call is inlined by verifier. */
return ((u64 *)ctx)[-1];
}
static const struct bpf_func_proto bpf_get_func_arg_cnt_proto = {
.func = get_func_arg_cnt,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
};
static const struct bpf_func_proto *
bpf_tracing_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
{
const struct bpf_func_proto *func_proto;
switch (func_id) {
case BPF_FUNC_get_smp_processor_id:
return &bpf_get_smp_processor_id_proto;
#ifdef CONFIG_ARCH_HAS_NON_OVERLAPPING_ADDRESS_SPACE
case BPF_FUNC_probe_read:
return security_locked_down(LOCKDOWN_BPF_READ_KERNEL) < 0 ?
NULL : &bpf_probe_read_compat_proto;
case BPF_FUNC_probe_read_str:
return security_locked_down(LOCKDOWN_BPF_READ_KERNEL) < 0 ?
NULL : &bpf_probe_read_compat_str_proto;
#endif
case BPF_FUNC_get_func_ip:
return &bpf_get_func_ip_proto_tracing;
default:
break;
}
func_proto = bpf_base_func_proto(func_id, prog);
if (func_proto)
return func_proto;
if (!bpf_token_capable(prog->aux->token, CAP_SYS_ADMIN))
return NULL;
switch (func_id) {
case BPF_FUNC_probe_write_user:
return security_locked_down(LOCKDOWN_BPF_WRITE_USER) < 0 ?
NULL : &bpf_probe_write_user_proto;
default:
return NULL;
}
}
static bool is_kprobe_multi(const struct bpf_prog *prog)
{
return prog->expected_attach_type == BPF_TRACE_KPROBE_MULTI ||
prog->expected_attach_type == BPF_TRACE_KPROBE_SESSION;
}
static inline bool is_kprobe_session(const struct bpf_prog *prog)
{
return prog->expected_attach_type == BPF_TRACE_KPROBE_SESSION;
}
static inline bool is_uprobe_multi(const struct bpf_prog *prog)
{
return prog->expected_attach_type == BPF_TRACE_UPROBE_MULTI ||
prog->expected_attach_type == BPF_TRACE_UPROBE_SESSION;
}
static inline bool is_uprobe_session(const struct bpf_prog *prog)
{
return prog->expected_attach_type == BPF_TRACE_UPROBE_SESSION;
}
static const struct bpf_func_proto *
kprobe_prog_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
{
switch (func_id) {
case BPF_FUNC_perf_event_output:
return &bpf_perf_event_output_proto;
case BPF_FUNC_get_stackid:
return &bpf_get_stackid_proto;
case BPF_FUNC_get_stack:
return prog->sleepable ? &bpf_get_stack_sleepable_proto : &bpf_get_stack_proto;
#ifdef CONFIG_BPF_KPROBE_OVERRIDE
case BPF_FUNC_override_return:
return &bpf_override_return_proto;
#endif
case BPF_FUNC_get_func_ip:
if (is_kprobe_multi(prog))
return &bpf_get_func_ip_proto_kprobe_multi;
if (is_uprobe_multi(prog))
return &bpf_get_func_ip_proto_uprobe_multi;
return &bpf_get_func_ip_proto_kprobe;
case BPF_FUNC_get_attach_cookie:
if (is_kprobe_multi(prog))
return &bpf_get_attach_cookie_proto_kmulti;
if (is_uprobe_multi(prog))
return &bpf_get_attach_cookie_proto_umulti;
return &bpf_get_attach_cookie_proto_trace;
default:
return bpf_tracing_func_proto(func_id, prog);
}
}
/* bpf+kprobe programs can access fields of 'struct pt_regs' */
static bool kprobe_prog_is_valid_access(int off, int size, enum bpf_access_type type,
const struct bpf_prog *prog,
struct bpf_insn_access_aux *info)
{
if (off < 0 || off >= sizeof(struct pt_regs))
return false;
if (off % size != 0)
return false;
/*
* Assertion for 32 bit to make sure last 8 byte access
* (BPF_DW) to the last 4 byte member is disallowed.
*/
if (off + size > sizeof(struct pt_regs))
return false;
if (type == BPF_WRITE)
prog->aux->kprobe_write_ctx = true;
return true;
}
const struct bpf_verifier_ops kprobe_verifier_ops = {
.get_func_proto = kprobe_prog_func_proto,
.is_valid_access = kprobe_prog_is_valid_access,
};
const struct bpf_prog_ops kprobe_prog_ops = {
};
BPF_CALL_5(bpf_perf_event_output_tp, void *, tp_buff, struct bpf_map *, map,
u64, flags, void *, data, u64, size)
{
struct pt_regs *regs = *(struct pt_regs **)tp_buff;
/*
* r1 points to perf tracepoint buffer where first 8 bytes are hidden
* from bpf program and contain a pointer to 'struct pt_regs'. Fetch it
* from there and call the same bpf_perf_event_output() helper inline.
*/
return ____bpf_perf_event_output(regs, map, flags, data, size);
}
static const struct bpf_func_proto bpf_perf_event_output_proto_tp = {
.func = bpf_perf_event_output_tp,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
.arg2_type = ARG_CONST_MAP_PTR,
.arg3_type = ARG_ANYTHING,
.arg4_type = ARG_PTR_TO_MEM | MEM_RDONLY,
.arg5_type = ARG_CONST_SIZE_OR_ZERO,
};
BPF_CALL_3(bpf_get_stackid_tp, void *, tp_buff, struct bpf_map *, map,
u64, flags)
{
struct pt_regs *regs = *(struct pt_regs **)tp_buff;
/*
* Same comment as in bpf_perf_event_output_tp(), only that this time
* the other helper's function body cannot be inlined due to being
* external, thus we need to call raw helper function.
*/
return bpf_get_stackid((unsigned long) regs, (unsigned long) map,
flags, 0, 0);
}
static const struct bpf_func_proto bpf_get_stackid_proto_tp = {
.func = bpf_get_stackid_tp,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
.arg2_type = ARG_CONST_MAP_PTR,
.arg3_type = ARG_ANYTHING,
};
BPF_CALL_4(bpf_get_stack_tp, void *, tp_buff, void *, buf, u32, size,
u64, flags)
{
struct pt_regs *regs = *(struct pt_regs **)tp_buff;
return bpf_get_stack((unsigned long) regs, (unsigned long) buf,
(unsigned long) size, flags, 0);
}
static const struct bpf_func_proto bpf_get_stack_proto_tp = {
.func = bpf_get_stack_tp,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
.arg2_type = ARG_PTR_TO_UNINIT_MEM,
.arg3_type = ARG_CONST_SIZE_OR_ZERO,
.arg4_type = ARG_ANYTHING,
};
static const struct bpf_func_proto *
tp_prog_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
{
switch (func_id) {
case BPF_FUNC_perf_event_output:
return &bpf_perf_event_output_proto_tp;
case BPF_FUNC_get_stackid:
return &bpf_get_stackid_proto_tp;
case BPF_FUNC_get_stack:
return &bpf_get_stack_proto_tp;
case BPF_FUNC_get_attach_cookie:
return &bpf_get_attach_cookie_proto_trace;
default:
return bpf_tracing_func_proto(func_id, prog);
}
}
static bool tp_prog_is_valid_access(int off, int size, enum bpf_access_type type,
const struct bpf_prog *prog,
struct bpf_insn_access_aux *info)
{
if (off < sizeof(void *) || off >= PERF_MAX_TRACE_SIZE)
return false;
if (type != BPF_READ)
return false;
if (off % size != 0)
return false;
BUILD_BUG_ON(PERF_MAX_TRACE_SIZE % sizeof(__u64));
return true;
}
const struct bpf_verifier_ops tracepoint_verifier_ops = {
.get_func_proto = tp_prog_func_proto,
.is_valid_access = tp_prog_is_valid_access,
};
const struct bpf_prog_ops tracepoint_prog_ops = {
};
BPF_CALL_3(bpf_perf_prog_read_value, struct bpf_perf_event_data_kern *, ctx,
struct bpf_perf_event_value *, buf, u32, size)
{
int err = -EINVAL;
if (unlikely(size != sizeof(struct bpf_perf_event_value)))
goto clear;
err = perf_event_read_local(ctx->event, &buf->counter, &buf->enabled,
&buf->running);
if (unlikely(err))
goto clear;
return 0;
clear:
memset(buf, 0, size);
return err;
}
static const struct bpf_func_proto bpf_perf_prog_read_value_proto = {
.func = bpf_perf_prog_read_value,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
.arg2_type = ARG_PTR_TO_UNINIT_MEM,
.arg3_type = ARG_CONST_SIZE,
};
BPF_CALL_4(bpf_read_branch_records, struct bpf_perf_event_data_kern *, ctx,
void *, buf, u32, size, u64, flags)
{
static const u32 br_entry_size = sizeof(struct perf_branch_entry);
struct perf_branch_stack *br_stack = ctx->data->br_stack;
u32 to_copy;
if (unlikely(flags & ~BPF_F_GET_BRANCH_RECORDS_SIZE))
return -EINVAL;
if (unlikely(!(ctx->data->sample_flags & PERF_SAMPLE_BRANCH_STACK)))
return -ENOENT;
if (unlikely(!br_stack))
return -ENOENT;
if (flags & BPF_F_GET_BRANCH_RECORDS_SIZE)
return br_stack->nr * br_entry_size;
if (!buf || (size % br_entry_size != 0))
return -EINVAL;
to_copy = min_t(u32, br_stack->nr * br_entry_size, size);
memcpy(buf, br_stack->entries, to_copy);
return to_copy;
}
static const struct bpf_func_proto bpf_read_branch_records_proto = {
.func = bpf_read_branch_records,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
.arg2_type = ARG_PTR_TO_MEM_OR_NULL,
.arg3_type = ARG_CONST_SIZE_OR_ZERO,
.arg4_type = ARG_ANYTHING,
};
static const struct bpf_func_proto *
pe_prog_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
{
switch (func_id) {
case BPF_FUNC_perf_event_output:
return &bpf_perf_event_output_proto_tp;
case BPF_FUNC_get_stackid:
return &bpf_get_stackid_proto_pe;
case BPF_FUNC_get_stack:
return &bpf_get_stack_proto_pe;
case BPF_FUNC_perf_prog_read_value:
return &bpf_perf_prog_read_value_proto;
case BPF_FUNC_read_branch_records:
return &bpf_read_branch_records_proto;
case BPF_FUNC_get_attach_cookie:
return &bpf_get_attach_cookie_proto_pe;
default:
return bpf_tracing_func_proto(func_id, prog);
}
}
/*
* bpf_raw_tp_regs are separate from bpf_pt_regs used from skb/xdp
* to avoid potential recursive reuse issue when/if tracepoints are added
* inside bpf_*_event_output, bpf_get_stackid and/or bpf_get_stack.
*
* Since raw tracepoints run despite bpf_prog_active, support concurrent usage
* in normal, irq, and nmi context.
*/
struct bpf_raw_tp_regs {
struct pt_regs regs[3];
};
static DEFINE_PER_CPU(struct bpf_raw_tp_regs, bpf_raw_tp_regs);
static DEFINE_PER_CPU(int, bpf_raw_tp_nest_level);
static struct pt_regs *get_bpf_raw_tp_regs(void)
{
struct bpf_raw_tp_regs *tp_regs = this_cpu_ptr(&bpf_raw_tp_regs);
int nest_level = this_cpu_inc_return(bpf_raw_tp_nest_level);
if (nest_level > ARRAY_SIZE(tp_regs->regs)) {
this_cpu_dec(bpf_raw_tp_nest_level);
return ERR_PTR(-EBUSY);
}
return &tp_regs->regs[nest_level - 1];
}
static void put_bpf_raw_tp_regs(void)
{
this_cpu_dec(bpf_raw_tp_nest_level);
}
BPF_CALL_5(bpf_perf_event_output_raw_tp, struct bpf_raw_tracepoint_args *, args,
struct bpf_map *, map, u64, flags, void *, data, u64, size)
{
struct pt_regs *regs = get_bpf_raw_tp_regs();
int ret;
if (IS_ERR(regs))
return PTR_ERR(regs);
perf_fetch_caller_regs(regs);
ret = ____bpf_perf_event_output(regs, map, flags, data, size);
put_bpf_raw_tp_regs();
return ret;
}
static const struct bpf_func_proto bpf_perf_event_output_proto_raw_tp = {
.func = bpf_perf_event_output_raw_tp,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
.arg2_type = ARG_CONST_MAP_PTR,
.arg3_type = ARG_ANYTHING,
.arg4_type = ARG_PTR_TO_MEM | MEM_RDONLY,
.arg5_type = ARG_CONST_SIZE_OR_ZERO,
};
extern const struct bpf_func_proto bpf_skb_output_proto;
extern const struct bpf_func_proto bpf_xdp_output_proto;
extern const struct bpf_func_proto bpf_xdp_get_buff_len_trace_proto;
BPF_CALL_3(bpf_get_stackid_raw_tp, struct bpf_raw_tracepoint_args *, args,
struct bpf_map *, map, u64, flags)
{
struct pt_regs *regs = get_bpf_raw_tp_regs();
int ret;
if (IS_ERR(regs))
return PTR_ERR(regs);
perf_fetch_caller_regs(regs);
/* similar to bpf_perf_event_output_tp, but pt_regs fetched differently */
ret = bpf_get_stackid((unsigned long) regs, (unsigned long) map,
flags, 0, 0);
put_bpf_raw_tp_regs();
return ret;
}
static const struct bpf_func_proto bpf_get_stackid_proto_raw_tp = {
.func = bpf_get_stackid_raw_tp,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
.arg2_type = ARG_CONST_MAP_PTR,
.arg3_type = ARG_ANYTHING,
};
BPF_CALL_4(bpf_get_stack_raw_tp, struct bpf_raw_tracepoint_args *, args,
void *, buf, u32, size, u64, flags)
{
struct pt_regs *regs = get_bpf_raw_tp_regs();
int ret;
if (IS_ERR(regs))
return PTR_ERR(regs);
perf_fetch_caller_regs(regs);
ret = bpf_get_stack((unsigned long) regs, (unsigned long) buf,
(unsigned long) size, flags, 0);
put_bpf_raw_tp_regs();
return ret;
}
static const struct bpf_func_proto bpf_get_stack_proto_raw_tp = {
.func = bpf_get_stack_raw_tp,
.gpl_only = true,
.ret_type = RET_INTEGER,
.arg1_type = ARG_PTR_TO_CTX,
.arg2_type = ARG_PTR_TO_MEM | MEM_RDONLY,
.arg3_type = ARG_CONST_SIZE_OR_ZERO,
.arg4_type = ARG_ANYTHING,
};
static const struct bpf_func_proto *
raw_tp_prog_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
{
switch (func_id) {
case BPF_FUNC_perf_event_output:
return &bpf_perf_event_output_proto_raw_tp;
case BPF_FUNC_get_stackid:
return &bpf_get_stackid_proto_raw_tp;
case BPF_FUNC_get_stack:
return &bpf_get_stack_proto_raw_tp;
case BPF_FUNC_get_attach_cookie:
return &bpf_get_attach_cookie_proto_tracing;
default:
return bpf_tracing_func_proto(func_id, prog);
}
}
const struct bpf_func_proto *
tracing_prog_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
{
const struct bpf_func_proto *fn;
switch (func_id) {
#ifdef CONFIG_NET
case BPF_FUNC_skb_output:
return &bpf_skb_output_proto;
case BPF_FUNC_xdp_output:
return &bpf_xdp_output_proto;
case BPF_FUNC_skc_to_tcp6_sock:
return &bpf_skc_to_tcp6_sock_proto;
case BPF_FUNC_skc_to_tcp_sock:
return &bpf_skc_to_tcp_sock_proto;
case BPF_FUNC_skc_to_tcp_timewait_sock:
return &bpf_skc_to_tcp_timewait_sock_proto;
case BPF_FUNC_skc_to_tcp_request_sock:
return &bpf_skc_to_tcp_request_sock_proto;
case BPF_FUNC_skc_to_udp6_sock:
return &bpf_skc_to_udp6_sock_proto;
case BPF_FUNC_skc_to_unix_sock:
return &bpf_skc_to_unix_sock_proto;
case BPF_FUNC_skc_to_mptcp_sock:
return &bpf_skc_to_mptcp_sock_proto;
case BPF_FUNC_sk_storage_get:
return &bpf_sk_storage_get_tracing_proto;
case BPF_FUNC_sk_storage_delete:
return &bpf_sk_storage_delete_tracing_proto;
case BPF_FUNC_sock_from_file:
return &bpf_sock_from_file_proto;
case BPF_FUNC_get_socket_cookie:
return &bpf_get_socket_ptr_cookie_proto;
case BPF_FUNC_xdp_get_buff_len:
return &bpf_xdp_get_buff_len_trace_proto;
#endif
case BPF_FUNC_seq_printf:
return prog->expected_attach_type == BPF_TRACE_ITER ?
&bpf_seq_printf_proto :
NULL;
case BPF_FUNC_seq_write:
return prog->expected_attach_type == BPF_TRACE_ITER ?
&bpf_seq_write_proto :
NULL;
case BPF_FUNC_seq_printf_btf:
return prog->expected_attach_type == BPF_TRACE_ITER ?
&bpf_seq_printf_btf_proto :
NULL;
case BPF_FUNC_d_path:
return &bpf_d_path_proto;
case BPF_FUNC_get_func_arg:
return bpf_prog_has_trampoline(prog) ? &bpf_get_func_arg_proto : NULL;
case BPF_FUNC_get_func_ret:
return bpf_prog_has_trampoline(prog) ? &bpf_get_func_ret_proto : NULL;
case BPF_FUNC_get_func_arg_cnt:
return bpf_prog_has_trampoline(prog) ? &bpf_get_func_arg_cnt_proto : NULL;
case BPF_FUNC_get_attach_cookie:
if (prog->type == BPF_PROG_TYPE_TRACING &&
prog->expected_attach_type == BPF_TRACE_RAW_TP)
return &bpf_get_attach_cookie_proto_tracing;
return bpf_prog_has_trampoline(prog) ? &bpf_get_attach_cookie_proto_tracing : NULL;
default:
fn = raw_tp_prog_func_proto(func_id, prog);
if (!fn && prog->expected_attach_type == BPF_TRACE_ITER)
fn = bpf_iter_get_func_proto(func_id, prog);
return fn;
}
}
static bool raw_tp_prog_is_valid_access(int off, int size,
enum bpf_access_type type,
const struct bpf_prog *prog,
struct bpf_insn_access_aux *info)
{
return bpf_tracing_ctx_access(off, size, type);
}
static bool tracing_prog_is_valid_access(int off, int size,
enum bpf_access_type type,
const struct bpf_prog *prog,
struct bpf_insn_access_aux *info)
{
return bpf_tracing_btf_ctx_access(off, size, type, prog, info);
}
int __weak bpf_prog_test_run_tracing(struct bpf_prog *prog,
const union bpf_attr *kattr,
union bpf_attr __user *uattr)
{
return -ENOTSUPP;
}
const struct bpf_verifier_ops raw_tracepoint_verifier_ops = {
.get_func_proto = raw_tp_prog_func_proto,
.is_valid_access = raw_tp_prog_is_valid_access,
};
const struct bpf_prog_ops raw_tracepoint_prog_ops = {
#ifdef CONFIG_NET
.test_run = bpf_prog_test_run_raw_tp,
#endif
};
const struct bpf_verifier_ops tracing_verifier_ops = {
.get_func_proto = tracing_prog_func_proto,
.is_valid_access = tracing_prog_is_valid_access,
};
const struct bpf_prog_ops tracing_prog_ops = {
.test_run = bpf_prog_test_run_tracing,
};
static bool raw_tp_writable_prog_is_valid_access(int off, int size,
enum bpf_access_type type,
const struct bpf_prog *prog,
struct bpf_insn_access_aux *info)
{
if (off == 0) {
if (size != sizeof(u64) || type != BPF_READ)
return false;
info->reg_type = PTR_TO_TP_BUFFER;
}
return raw_tp_prog_is_valid_access(off, size, type, prog, info);
}
const struct bpf_verifier_ops raw_tracepoint_writable_verifier_ops = {
.get_func_proto = raw_tp_prog_func_proto,
.is_valid_access = raw_tp_writable_prog_is_valid_access,
};
const struct bpf_prog_ops raw_tracepoint_writable_prog_ops = {
};
static bool pe_prog_is_valid_access(int off, int size, enum bpf_access_type type,
const struct bpf_prog *prog,
struct bpf_insn_access_aux *info)
{
const int size_u64 = sizeof(u64);
if (off < 0 || off >= sizeof(struct bpf_perf_event_data))
return false;
if (type != BPF_READ)
return false;
if (off % size != 0) {
if (sizeof(unsigned long) != 4)
return false;
if (size != 8)
return false;
if (off % size != 4)
return false;
}
switch (off) {
case bpf_ctx_range(struct bpf_perf_event_data, sample_period):
bpf_ctx_record_field_size(info, size_u64);
if (!bpf_ctx_narrow_access_ok(off, size, size_u64))
return false;
break;
case bpf_ctx_range(struct bpf_perf_event_data, addr):
bpf_ctx_record_field_size(info, size_u64);
if (!bpf_ctx_narrow_access_ok(off, size, size_u64))
return false;
break;
default:
if (size != sizeof(long))
return false;
}
return true;
}
static u32 pe_prog_convert_ctx_access(enum bpf_access_type type,
const struct bpf_insn *si,
struct bpf_insn *insn_buf,
struct bpf_prog *prog, u32 *target_size)
{
struct bpf_insn *insn = insn_buf;
switch (si->off) {
case offsetof(struct bpf_perf_event_data, sample_period):
*insn++ = BPF_LDX_MEM(BPF_FIELD_SIZEOF(struct bpf_perf_event_data_kern,
data), si->dst_reg, si->src_reg,
offsetof(struct bpf_perf_event_data_kern, data));
*insn++ = BPF_LDX_MEM(BPF_DW, si->dst_reg, si->dst_reg,
bpf_target_off(struct perf_sample_data, period, 8,
target_size));
break;
case offsetof(struct bpf_perf_event_data, addr):
*insn++ = BPF_LDX_MEM(BPF_FIELD_SIZEOF(struct bpf_perf_event_data_kern,
data), si->dst_reg, si->src_reg,
offsetof(struct bpf_perf_event_data_kern, data));
*insn++ = BPF_LDX_MEM(BPF_DW, si->dst_reg, si->dst_reg,
bpf_target_off(struct perf_sample_data, addr, 8,
target_size));
break;
default:
*insn++ = BPF_LDX_MEM(BPF_FIELD_SIZEOF(struct bpf_perf_event_data_kern,
regs), si->dst_reg, si->src_reg,
offsetof(struct bpf_perf_event_data_kern, regs));
*insn++ = BPF_LDX_MEM(BPF_SIZEOF(long), si->dst_reg, si->dst_reg,
si->off);
break;
}
return insn - insn_buf;
}
const struct bpf_verifier_ops perf_event_verifier_ops = {
.get_func_proto = pe_prog_func_proto,
.is_valid_access = pe_prog_is_valid_access,
.convert_ctx_access = pe_prog_convert_ctx_access,
};
const struct bpf_prog_ops perf_event_prog_ops = {
};
static DEFINE_MUTEX(bpf_event_mutex);
#define BPF_TRACE_MAX_PROGS 64
int perf_event_attach_bpf_prog(struct perf_event *event,
struct bpf_prog *prog,
u64 bpf_cookie)
{
struct bpf_prog_array *old_array;
struct bpf_prog_array *new_array;
int ret = -EEXIST;
/*
* Kprobe override only works if they are on the function entry,
* and only if they are on the opt-in list.
*/
if (prog->kprobe_override &&
(!trace_kprobe_on_func_entry(event->tp_event) ||
!trace_kprobe_error_injectable(event->tp_event)))
return -EINVAL;
mutex_lock(&bpf_event_mutex);
if (event->prog)
goto unlock;
old_array = bpf_event_rcu_dereference(event->tp_event->prog_array);
if (old_array &&
bpf_prog_array_length(old_array) >= BPF_TRACE_MAX_PROGS) {
ret = -E2BIG;
goto unlock;
}
ret = bpf_prog_array_copy(old_array, NULL, prog, bpf_cookie, &new_array);
if (ret < 0)
goto unlock;
/* set the new array to event->tp_event and set event->prog */
event->prog = prog;
event->bpf_cookie = bpf_cookie;
rcu_assign_pointer(event->tp_event->prog_array, new_array);
bpf_prog_array_free_sleepable(old_array);
unlock:
mutex_unlock(&bpf_event_mutex);
return ret;
}
void perf_event_detach_bpf_prog(struct perf_event *event)
{
struct bpf_prog_array *old_array;
struct bpf_prog_array *new_array;
struct bpf_prog *prog = NULL;
int ret;
mutex_lock(&bpf_event_mutex);
if (!event->prog)
goto unlock;
old_array = bpf_event_rcu_dereference(event->tp_event->prog_array);
if (!old_array)
goto put;
ret = bpf_prog_array_copy(old_array, event->prog, NULL, 0, &new_array);
if (ret < 0) {
bpf_prog_array_delete_safe(old_array, event->prog);
} else {
rcu_assign_pointer(event->tp_event->prog_array, new_array);
bpf_prog_array_free_sleepable(old_array);
}
put:
prog = event->prog;
event->prog = NULL;
unlock:
mutex_unlock(&bpf_event_mutex);
if (prog) {
/*
* It could be that the bpf_prog is not sleepable (and will be freed
* via normal RCU), but is called from a point that supports sleepable
* programs and uses tasks-trace-RCU.
*/
synchronize_rcu_tasks_trace();
bpf_prog_put(prog);
}
}
int perf_event_query_prog_array(struct perf_event *event, void __user *info)
{
struct perf_event_query_bpf __user *uquery = info;
struct perf_event_query_bpf query = {};
struct bpf_prog_array *progs;
u32 *ids, prog_cnt, ids_len;
int ret;
if (!perfmon_capable())
return -EPERM;
if (event->attr.type != PERF_TYPE_TRACEPOINT)
return -EINVAL;
if (copy_from_user(&query, uquery, sizeof(query)))
return -EFAULT;
ids_len = query.ids_len;
if (ids_len > BPF_TRACE_MAX_PROGS)
return -E2BIG;
ids = kcalloc(ids_len, sizeof(u32), GFP_USER | __GFP_NOWARN);
if (!ids)
return -ENOMEM;
/*
* The above kcalloc returns ZERO_SIZE_PTR when ids_len = 0, which
* is required when user only wants to check for uquery->prog_cnt.
* There is no need to check for it since the case is handled
* gracefully in bpf_prog_array_copy_info.
*/
mutex_lock(&bpf_event_mutex);
progs = bpf_event_rcu_dereference(event->tp_event->prog_array);
ret = bpf_prog_array_copy_info(progs, ids, ids_len, &prog_cnt);
mutex_unlock(&bpf_event_mutex);
if (copy_to_user(&uquery->prog_cnt, &prog_cnt, sizeof(prog_cnt)) ||
copy_to_user(uquery->ids, ids, ids_len * sizeof(u32)))
ret = -EFAULT;
kfree(ids);
return ret;
}
extern struct bpf_raw_event_map __start__bpf_raw_tp[];
extern struct bpf_raw_event_map __stop__bpf_raw_tp[];
struct bpf_raw_event_map *bpf_get_raw_tracepoint(const char *name)
{
struct bpf_raw_event_map *btp = __start__bpf_raw_tp;
for (; btp < __stop__bpf_raw_tp; btp++) {
if (!strcmp(btp->tp->name, name))
return btp;
}
return bpf_get_raw_tracepoint_module(name);
}
void bpf_put_raw_tracepoint(struct bpf_raw_event_map *btp)
{
struct module *mod;
guard(rcu)();
mod = __module_address((unsigned long)btp);
module_put(mod);
}
static __always_inline
void __bpf_trace_run(struct bpf_raw_tp_link *link, u64 *args)
{
struct bpf_prog *prog = link->link.prog;
struct bpf_run_ctx *old_run_ctx;
struct bpf_trace_run_ctx run_ctx;
cant_sleep();
if (unlikely(this_cpu_inc_return(*(prog->active)) != 1)) {
bpf_prog_inc_misses_counter(prog);
goto out;
}
run_ctx.bpf_cookie = link->cookie;
old_run_ctx = bpf_set_run_ctx(&run_ctx.run_ctx);
rcu_read_lock();
(void) bpf_prog_run(prog, args);
rcu_read_unlock();
bpf_reset_run_ctx(old_run_ctx);
out:
this_cpu_dec(*(prog->active));
}
#define UNPACK(...) __VA_ARGS__
#define REPEAT_1(FN, DL, X, ...) FN(X)
#define REPEAT_2(FN, DL, X, ...) FN(X) UNPACK DL REPEAT_1(FN, DL, __VA_ARGS__)
#define REPEAT_3(FN, DL, X, ...) FN(X) UNPACK DL REPEAT_2(FN, DL, __VA_ARGS__)
#define REPEAT_4(FN, DL, X, ...) FN(X) UNPACK DL REPEAT_3(FN, DL, __VA_ARGS__)
#define REPEAT_5(FN, DL, X, ...) FN(X) UNPACK DL REPEAT_4(FN, DL, __VA_ARGS__)
#define REPEAT_6(FN, DL, X, ...) FN(X) UNPACK DL REPEAT_5(FN, DL, __VA_ARGS__)
#define REPEAT_7(FN, DL, X, ...) FN(X) UNPACK DL REPEAT_6(FN, DL, __VA_ARGS__)
#define REPEAT_8(FN, DL, X, ...) FN(X) UNPACK DL REPEAT_7(FN, DL, __VA_ARGS__)
#define REPEAT_9(FN, DL, X, ...) FN(X) UNPACK DL REPEAT_8(FN, DL, __VA_ARGS__)
#define REPEAT_10(FN, DL, X, ...) FN(X) UNPACK DL REPEAT_9(FN, DL, __VA_ARGS__)
#define REPEAT_11(FN, DL, X, ...) FN(X) UNPACK DL REPEAT_10(FN, DL, __VA_ARGS__)
#define REPEAT_12(FN, DL, X, ...) FN(X) UNPACK DL REPEAT_11(FN, DL, __VA_ARGS__)
#define REPEAT(X, FN, DL, ...) REPEAT_##X(FN, DL, __VA_ARGS__)
#define SARG(X) u64 arg##X
#define COPY(X) args[X] = arg##X
#define __DL_COM (,)
#define __DL_SEM (;)
#define __SEQ_0_11 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11
#define BPF_TRACE_DEFN_x(x) \
void bpf_trace_run##x(struct bpf_raw_tp_link *link, \
REPEAT(x, SARG, __DL_COM, __SEQ_0_11)) \
{ \
u64 args[x]; \
REPEAT(x, COPY, __DL_SEM, __SEQ_0_11); \
__bpf_trace_run(link, args); \
} \
EXPORT_SYMBOL_GPL(bpf_trace_run##x)
BPF_TRACE_DEFN_x(1);
BPF_TRACE_DEFN_x(2);
BPF_TRACE_DEFN_x(3);
BPF_TRACE_DEFN_x(4);
BPF_TRACE_DEFN_x(5);
BPF_TRACE_DEFN_x(6);
BPF_TRACE_DEFN_x(7);
BPF_TRACE_DEFN_x(8);
BPF_TRACE_DEFN_x(9);
BPF_TRACE_DEFN_x(10);
BPF_TRACE_DEFN_x(11);
BPF_TRACE_DEFN_x(12);
int bpf_probe_register(struct bpf_raw_event_map *btp, struct bpf_raw_tp_link *link)
{
struct tracepoint *tp = btp->tp;
struct bpf_prog *prog = link->link.prog;
/*
* check that program doesn't access arguments beyond what's
* available in this tracepoint
*/
if (prog->aux->max_ctx_offset > btp->num_args * sizeof(u64))
return -EINVAL;
if (prog->aux->max_tp_access > btp->writable_size)
return -EINVAL;
return tracepoint_probe_register_may_exist(tp, (void *)btp->bpf_func, link);
}
int bpf_probe_unregister(struct bpf_raw_event_map *btp, struct bpf_raw_tp_link *link)
{
return tracepoint_probe_unregister(btp->tp, (void *)btp->bpf_func, link);
}
int bpf_get_perf_event_info(const struct perf_event *event, u32 *prog_id,
u32 *fd_type, const char **buf,
u64 *probe_offset, u64 *probe_addr,
unsigned long *missed)
{
bool is_tracepoint, is_syscall_tp;
struct bpf_prog *prog;
int flags, err = 0;
prog = event->prog;
if (!prog)
return -ENOENT;
/* not supporting BPF_PROG_TYPE_PERF_EVENT yet */
if (prog->type == BPF_PROG_TYPE_PERF_EVENT)
return -EOPNOTSUPP;
*prog_id = prog->aux->id;
flags = event->tp_event->flags;
is_tracepoint = flags & TRACE_EVENT_FL_TRACEPOINT;
is_syscall_tp = is_syscall_trace_event(event->tp_event);
if (is_tracepoint || is_syscall_tp) {
*buf = is_tracepoint ? event->tp_event->tp->name
: event->tp_event->name;
/* We allow NULL pointer for tracepoint */
if (fd_type)
*fd_type = BPF_FD_TYPE_TRACEPOINT;
if (probe_offset)
*probe_offset = 0x0;
if (probe_addr)
*probe_addr = 0x0;
} else {
/* kprobe/uprobe */
err = -EOPNOTSUPP;
#ifdef CONFIG_KPROBE_EVENTS
if (flags & TRACE_EVENT_FL_KPROBE)
err = bpf_get_kprobe_info(event, fd_type, buf,
probe_offset, probe_addr, missed,
event->attr.type == PERF_TYPE_TRACEPOINT);
#endif
#ifdef CONFIG_UPROBE_EVENTS
if (flags & TRACE_EVENT_FL_UPROBE)
err = bpf_get_uprobe_info(event, fd_type, buf,
probe_offset, probe_addr,
event->attr.type == PERF_TYPE_TRACEPOINT);
#endif
}
return err;
}
static int __init send_signal_irq_work_init(void)
{
int cpu;
struct send_signal_irq_work *work;
for_each_possible_cpu(cpu) {
work = per_cpu_ptr(&send_signal_work, cpu);
init_irq_work(&work->irq_work, do_bpf_send_signal);
}
return 0;
}
subsys_initcall(send_signal_irq_work_init);
#ifdef CONFIG_MODULES
static int bpf_event_notify(struct notifier_block *nb, unsigned long op,
void *module)
{
struct bpf_trace_module *btm, *tmp;
struct module *mod = module;
int ret = 0;
if (mod->num_bpf_raw_events == 0 ||
(op != MODULE_STATE_COMING && op != MODULE_STATE_GOING))
goto out;
mutex_lock(&bpf_module_mutex);
switch (op) {
case MODULE_STATE_COMING:
btm = kzalloc(sizeof(*btm), GFP_KERNEL);
if (btm) {
btm->module = module;
list_add(&btm->list, &bpf_trace_modules);
} else {
ret = -ENOMEM;
}
break;
case MODULE_STATE_GOING:
list_for_each_entry_safe(btm, tmp, &bpf_trace_modules, list) {
if (btm->module == module) {
list_del(&btm->list);
kfree(btm);
break;
}
}
break;
}
mutex_unlock(&bpf_module_mutex);
out:
return notifier_from_errno(ret);
}
static struct notifier_block bpf_module_nb = {
.notifier_call = bpf_event_notify,
};
static int __init bpf_event_init(void)
{
register_module_notifier(&bpf_module_nb);
return 0;
}
fs_initcall(bpf_event_init);
#endif /* CONFIG_MODULES */
struct bpf_session_run_ctx {
struct bpf_run_ctx run_ctx;
bool is_return;
void *data;
};
#ifdef CONFIG_FPROBE
struct bpf_kprobe_multi_link {
struct bpf_link link;
struct fprobe fp;
unsigned long *addrs;
u64 *cookies;
u32 cnt;
u32 mods_cnt;
struct module **mods;
};
struct bpf_kprobe_multi_run_ctx {
struct bpf_session_run_ctx session_ctx;
struct bpf_kprobe_multi_link *link;
unsigned long entry_ip;
};
struct user_syms {
const char **syms;
char *buf;
};
#ifndef CONFIG_HAVE_FTRACE_REGS_HAVING_PT_REGS
static DEFINE_PER_CPU(struct pt_regs, bpf_kprobe_multi_pt_regs);
#define bpf_kprobe_multi_pt_regs_ptr() this_cpu_ptr(&bpf_kprobe_multi_pt_regs)
#else
#define bpf_kprobe_multi_pt_regs_ptr() (NULL)
#endif
static unsigned long ftrace_get_entry_ip(unsigned long fentry_ip)
{
unsigned long ip = ftrace_get_symaddr(fentry_ip);
return ip ? : fentry_ip;
}
static int copy_user_syms(struct user_syms *us, unsigned long __user *usyms, u32 cnt)
{
unsigned long __user usymbol;
const char **syms = NULL;
char *buf = NULL, *p;
int err = -ENOMEM;
unsigned int i;
syms = kvmalloc_array(cnt, sizeof(*syms), GFP_KERNEL);
if (!syms)
goto error;
buf = kvmalloc_array(cnt, KSYM_NAME_LEN, GFP_KERNEL);
if (!buf)
goto error;
for (p = buf, i = 0; i < cnt; i++) {
if (__get_user(usymbol, usyms + i)) {
err = -EFAULT;
goto error;
}
err = strncpy_from_user(p, (const char __user *) usymbol, KSYM_NAME_LEN);
if (err == KSYM_NAME_LEN)
err = -E2BIG;
if (err < 0)
goto error;
syms[i] = p;
p += err + 1;
}
us->syms = syms;
us->buf = buf;
return 0;
error:
if (err) {
kvfree(syms);
kvfree(buf);
}
return err;
}
static void kprobe_multi_put_modules(struct module **mods, u32 cnt)
{
u32 i;
for (i = 0; i < cnt; i++)
module_put(mods[i]);
}
static void free_user_syms(struct user_syms *us)
{
kvfree(us->syms);
kvfree(us->buf);
}
static void bpf_kprobe_multi_link_release(struct bpf_link *link)
{
struct bpf_kprobe_multi_link *kmulti_link;
kmulti_link = container_of(link, struct bpf_kprobe_multi_link, link);
unregister_fprobe(&kmulti_link->fp);
kprobe_multi_put_modules(kmulti_link->mods, kmulti_link->mods_cnt);
}
static void bpf_kprobe_multi_link_dealloc(struct bpf_link *link)
{
struct bpf_kprobe_multi_link *kmulti_link;
kmulti_link = container_of(link, struct bpf_kprobe_multi_link, link);
kvfree(kmulti_link->addrs);
kvfree(kmulti_link->cookies);
kfree(kmulti_link->mods);
kfree(kmulti_link);
}
static int bpf_kprobe_multi_link_fill_link_info(const struct bpf_link *link,
struct bpf_link_info *info)
{
u64 __user *ucookies = u64_to_user_ptr(info->kprobe_multi.cookies);
u64 __user *uaddrs = u64_to_user_ptr(info->kprobe_multi.addrs);
struct bpf_kprobe_multi_link *kmulti_link;
u32 ucount = info->kprobe_multi.count;
int err = 0, i;
if (!uaddrs ^ !ucount)
return -EINVAL;
if (ucookies && !ucount)
return -EINVAL;
kmulti_link = container_of(link, struct bpf_kprobe_multi_link, link);
info->kprobe_multi.count = kmulti_link->cnt;
info->kprobe_multi.flags = kmulti_link->link.flags;
info->kprobe_multi.missed = kmulti_link->fp.nmissed;
if (!uaddrs)
return 0;
if (ucount < kmulti_link->cnt)
err = -ENOSPC;
else
ucount = kmulti_link->cnt;
if (ucookies) {
if (kmulti_link->cookies) {
if (copy_to_user(ucookies, kmulti_link->cookies, ucount * sizeof(u64)))
return -EFAULT;
} else {
for (i = 0; i < ucount; i++) {
if (put_user(0, ucookies + i))
return -EFAULT;
}
}
}
if (kallsyms_show_value(current_cred())) {
if (copy_to_user(uaddrs, kmulti_link->addrs, ucount * sizeof(u64)))
return -EFAULT;
} else {
for (i = 0; i < ucount; i++) {
if (put_user(0, uaddrs + i))
return -EFAULT;
}
}
return err;
}
#ifdef CONFIG_PROC_FS
static void bpf_kprobe_multi_show_fdinfo(const struct bpf_link *link,
struct seq_file *seq)
{
struct bpf_kprobe_multi_link *kmulti_link;
kmulti_link = container_of(link, struct bpf_kprobe_multi_link, link);
seq_printf(seq,
"kprobe_cnt:\t%u\n"
"missed:\t%lu\n",
kmulti_link->cnt,
kmulti_link->fp.nmissed);
seq_printf(seq, "%s\t %s\n", "cookie", "func");
for (int i = 0; i < kmulti_link->cnt; i++) {
seq_printf(seq,
"%llu\t %pS\n",
kmulti_link->cookies[i],
(void *)kmulti_link->addrs[i]);
}
}
#endif
static const struct bpf_link_ops bpf_kprobe_multi_link_lops = {
.release = bpf_kprobe_multi_link_release,
.dealloc_deferred = bpf_kprobe_multi_link_dealloc,
.fill_link_info = bpf_kprobe_multi_link_fill_link_info,
#ifdef CONFIG_PROC_FS
.show_fdinfo = bpf_kprobe_multi_show_fdinfo,
#endif
};
static void bpf_kprobe_multi_cookie_swap(void *a, void *b, int size, const void *priv)
{
const struct bpf_kprobe_multi_link *link = priv;
unsigned long *addr_a = a, *addr_b = b;
u64 *cookie_a, *cookie_b;
cookie_a = link->cookies + (addr_a - link->addrs);
cookie_b = link->cookies + (addr_b - link->addrs);
/* swap addr_a/addr_b and cookie_a/cookie_b values */
swap(*addr_a, *addr_b);
swap(*cookie_a, *cookie_b);
}
static int bpf_kprobe_multi_addrs_cmp(const void *a, const void *b)
{
const unsigned long *addr_a = a, *addr_b = b;
if (*addr_a == *addr_b)
return 0;
return *addr_a < *addr_b ? -1 : 1;
}
static int bpf_kprobe_multi_cookie_cmp(const void *a, const void *b, const void *priv)
{
return bpf_kprobe_multi_addrs_cmp(a, b);
}
static u64 bpf_kprobe_multi_cookie(struct bpf_run_ctx *ctx)
{
struct bpf_kprobe_multi_run_ctx *run_ctx;
struct bpf_kprobe_multi_link *link;
u64 *cookie, entry_ip;
unsigned long *addr;
if (WARN_ON_ONCE(!ctx))
return 0;
run_ctx = container_of(current->bpf_ctx, struct bpf_kprobe_multi_run_ctx,
session_ctx.run_ctx);
link = run_ctx->link;
if (!link->cookies)
return 0;
entry_ip = run_ctx->entry_ip;
addr = bsearch(&entry_ip, link->addrs, link->cnt, sizeof(entry_ip),
bpf_kprobe_multi_addrs_cmp);
if (!addr)
return 0;
cookie = link->cookies + (addr - link->addrs);
return *cookie;
}
static u64 bpf_kprobe_multi_entry_ip(struct bpf_run_ctx *ctx)
{
struct bpf_kprobe_multi_run_ctx *run_ctx;
run_ctx = container_of(current->bpf_ctx, struct bpf_kprobe_multi_run_ctx,
session_ctx.run_ctx);
return run_ctx->entry_ip;
}
static __always_inline int
kprobe_multi_link_prog_run(struct bpf_kprobe_multi_link *link,
unsigned long entry_ip, struct ftrace_regs *fregs,
bool is_return, void *data)
{
struct bpf_kprobe_multi_run_ctx run_ctx = {
.session_ctx = {
.is_return = is_return,
.data = data,
},
.link = link,
.entry_ip = entry_ip,
};
struct bpf_run_ctx *old_run_ctx;
struct pt_regs *regs;
int err;
/*
* graph tracer framework ensures we won't migrate, so there is no need
* to use migrate_disable for bpf_prog_run again. The check here just for
* __this_cpu_inc_return.
*/
cant_sleep();
if (unlikely(__this_cpu_inc_return(bpf_prog_active) != 1)) {
bpf_prog_inc_misses_counter(link->link.prog);
err = 1;
goto out;
}
rcu_read_lock();
regs = ftrace_partial_regs(fregs, bpf_kprobe_multi_pt_regs_ptr());
old_run_ctx = bpf_set_run_ctx(&run_ctx.session_ctx.run_ctx);
err = bpf_prog_run(link->link.prog, regs);
bpf_reset_run_ctx(old_run_ctx);
rcu_read_unlock();
out:
__this_cpu_dec(bpf_prog_active);
return err;
}
static int
kprobe_multi_link_handler(struct fprobe *fp, unsigned long fentry_ip,
unsigned long ret_ip, struct ftrace_regs *fregs,
void *data)
{
struct bpf_kprobe_multi_link *link;
int err;
link = container_of(fp, struct bpf_kprobe_multi_link, fp);
err = kprobe_multi_link_prog_run(link, ftrace_get_entry_ip(fentry_ip),
fregs, false, data);
return is_kprobe_session(link->link.prog) ? err : 0;
}
static void
kprobe_multi_link_exit_handler(struct fprobe *fp, unsigned long fentry_ip,
unsigned long ret_ip, struct ftrace_regs *fregs,
void *data)
{
struct bpf_kprobe_multi_link *link;
link = container_of(fp, struct bpf_kprobe_multi_link, fp);
kprobe_multi_link_prog_run(link, ftrace_get_entry_ip(fentry_ip),
fregs, true, data);
}
static int symbols_cmp_r(const void *a, const void *b, const void *priv)
{
const char **str_a = (const char **) a;
const char **str_b = (const char **) b;
return strcmp(*str_a, *str_b);
}
struct multi_symbols_sort {
const char **funcs;
u64 *cookies;
};
static void symbols_swap_r(void *a, void *b, int size, const void *priv)
{
const struct multi_symbols_sort *data = priv;
const char **name_a = a, **name_b = b;
swap(*name_a, *name_b);
/* If defined, swap also related cookies. */
if (data->cookies) {
u64 *cookie_a, *cookie_b;
cookie_a = data->cookies + (name_a - data->funcs);
cookie_b = data->cookies + (name_b - data->funcs);
swap(*cookie_a, *cookie_b);
}
}
struct modules_array {
struct module **mods;
int mods_cnt;
int mods_cap;
};
static int add_module(struct modules_array *arr, struct module *mod)
{
struct module **mods;
if (arr->mods_cnt == arr->mods_cap) {
arr->mods_cap = max(16, arr->mods_cap * 3 / 2);
mods = krealloc_array(arr->mods, arr->mods_cap, sizeof(*mods), GFP_KERNEL);
if (!mods)
return -ENOMEM;
arr->mods = mods;
}
arr->mods[arr->mods_cnt] = mod;
arr->mods_cnt++;
return 0;
}
static bool has_module(struct modules_array *arr, struct module *mod)
{
int i;
for (i = arr->mods_cnt - 1; i >= 0; i--) {
if (arr->mods[i] == mod)
return true;
}
return false;
}
static int get_modules_for_addrs(struct module ***mods, unsigned long *addrs, u32 addrs_cnt)
{
struct modules_array arr = {};
u32 i, err = 0;
for (i = 0; i < addrs_cnt; i++) {
bool skip_add = false;
struct module *mod;
scoped_guard(rcu) {
mod = __module_address(addrs[i]);
/* Either no module or it's already stored */
if (!mod || has_module(&arr, mod)) {
skip_add = true;
break; /* scoped_guard */
}
if (!try_module_get(mod))
err = -EINVAL;
}
if (skip_add)
continue;
if (err)
break;
err = add_module(&arr, mod);
if (err) {
module_put(mod);
break;
}
}
/* We return either err < 0 in case of error, ... */
if (err) {
kprobe_multi_put_modules(arr.mods, arr.mods_cnt);
kfree(arr.mods);
return err;
}
/* or number of modules found if everything is ok. */
*mods = arr.mods;
return arr.mods_cnt;
}
static int addrs_check_error_injection_list(unsigned long *addrs, u32 cnt)
{
u32 i;
for (i = 0; i < cnt; i++) {
if (!within_error_injection_list(addrs[i]))
return -EINVAL;
}
return 0;
}
int bpf_kprobe_multi_link_attach(const union bpf_attr *attr, struct bpf_prog *prog)
{
struct bpf_kprobe_multi_link *link = NULL;
struct bpf_link_primer link_primer;
void __user *ucookies;
unsigned long *addrs;
u32 flags, cnt, size;
void __user *uaddrs;
u64 *cookies = NULL;
void __user *usyms;
int err;
/* no support for 32bit archs yet */
if (sizeof(u64) != sizeof(void *))
return -EOPNOTSUPP;
if (attr->link_create.flags)
return -EINVAL;
if (!is_kprobe_multi(prog))
return -EINVAL;
/* Writing to context is not allowed for kprobes. */
if (prog->aux->kprobe_write_ctx)
return -EINVAL;
flags = attr->link_create.kprobe_multi.flags;
if (flags & ~BPF_F_KPROBE_MULTI_RETURN)
return -EINVAL;
uaddrs = u64_to_user_ptr(attr->link_create.kprobe_multi.addrs);
usyms = u64_to_user_ptr(attr->link_create.kprobe_multi.syms);
if (!!uaddrs == !!usyms)
return -EINVAL;
cnt = attr->link_create.kprobe_multi.cnt;
if (!cnt)
return -EINVAL;
if (cnt > MAX_KPROBE_MULTI_CNT)
return -E2BIG;
size = cnt * sizeof(*addrs);
addrs = kvmalloc_array(cnt, sizeof(*addrs), GFP_KERNEL);
if (!addrs)
return -ENOMEM;
ucookies = u64_to_user_ptr(attr->link_create.kprobe_multi.cookies);
if (ucookies) {
cookies = kvmalloc_array(cnt, sizeof(*addrs), GFP_KERNEL);
if (!cookies) {
err = -ENOMEM;
goto error;
}
if (copy_from_user(cookies, ucookies, size)) {
err = -EFAULT;
goto error;
}
}
if (uaddrs) {
if (copy_from_user(addrs, uaddrs, size)) {
err = -EFAULT;
goto error;
}
} else {
struct multi_symbols_sort data = {
.cookies = cookies,
};
struct user_syms us;
err = copy_user_syms(&us, usyms, cnt);
if (err)
goto error;
if (cookies)
data.funcs = us.syms;
sort_r(us.syms, cnt, sizeof(*us.syms), symbols_cmp_r,
symbols_swap_r, &data);
err = ftrace_lookup_symbols(us.syms, cnt, addrs);
free_user_syms(&us);
if (err)
goto error;
}
if (prog->kprobe_override && addrs_check_error_injection_list(addrs, cnt)) {
err = -EINVAL;
goto error;
}
link = kzalloc(sizeof(*link), GFP_KERNEL);
if (!link) {
err = -ENOMEM;
goto error;
}
bpf_link_init(&link->link, BPF_LINK_TYPE_KPROBE_MULTI,
&bpf_kprobe_multi_link_lops, prog, attr->link_create.attach_type);
err = bpf_link_prime(&link->link, &link_primer);
if (err)
goto error;
if (!(flags & BPF_F_KPROBE_MULTI_RETURN))
link->fp.entry_handler = kprobe_multi_link_handler;
if ((flags & BPF_F_KPROBE_MULTI_RETURN) || is_kprobe_session(prog))
link->fp.exit_handler = kprobe_multi_link_exit_handler;
if (is_kprobe_session(prog))
link->fp.entry_data_size = sizeof(u64);
link->addrs = addrs;
link->cookies = cookies;
link->cnt = cnt;
link->link.flags = flags;
if (cookies) {
/*
* Sorting addresses will trigger sorting cookies as well
* (check bpf_kprobe_multi_cookie_swap). This way we can
* find cookie based on the address in bpf_get_attach_cookie
* helper.
*/
sort_r(addrs, cnt, sizeof(*addrs),
bpf_kprobe_multi_cookie_cmp,
bpf_kprobe_multi_cookie_swap,
link);
}
err = get_modules_for_addrs(&link->mods, addrs, cnt);
if (err < 0) {
bpf_link_cleanup(&link_primer);
return err;
}
link->mods_cnt = err;
err = register_fprobe_ips(&link->fp, addrs, cnt);
if (err) {
kprobe_multi_put_modules(link->mods, link->mods_cnt);
bpf_link_cleanup(&link_primer);
return err;
}
return bpf_link_settle(&link_primer);
error:
kfree(link);
kvfree(addrs);
kvfree(cookies);
return err;
}
#else /* !CONFIG_FPROBE */
int bpf_kprobe_multi_link_attach(const union bpf_attr *attr, struct bpf_prog *prog)
{
return -EOPNOTSUPP;
}
static u64 bpf_kprobe_multi_cookie(struct bpf_run_ctx *ctx)
{
return 0;
}
static u64 bpf_kprobe_multi_entry_ip(struct bpf_run_ctx *ctx)
{
return 0;
}
#endif
#ifdef CONFIG_UPROBES
struct bpf_uprobe_multi_link;
struct bpf_uprobe {
struct bpf_uprobe_multi_link *link;
loff_t offset;
unsigned long ref_ctr_offset;
u64 cookie;
struct uprobe *uprobe;
struct uprobe_consumer consumer;
bool session;
};
struct bpf_uprobe_multi_link {
struct path path;
struct bpf_link link;
u32 cnt;
struct bpf_uprobe *uprobes;
struct task_struct *task;
};
struct bpf_uprobe_multi_run_ctx {
struct bpf_session_run_ctx session_ctx;
unsigned long entry_ip;
struct bpf_uprobe *uprobe;
};
static void bpf_uprobe_unregister(struct bpf_uprobe *uprobes, u32 cnt)
{
u32 i;
for (i = 0; i < cnt; i++)
uprobe_unregister_nosync(uprobes[i].uprobe, &uprobes[i].consumer);
if (cnt)
uprobe_unregister_sync();
}
static void bpf_uprobe_multi_link_release(struct bpf_link *link)
{
struct bpf_uprobe_multi_link *umulti_link;
umulti_link = container_of(link, struct bpf_uprobe_multi_link, link);
bpf_uprobe_unregister(umulti_link->uprobes, umulti_link->cnt);
if (umulti_link->task)
put_task_struct(umulti_link->task);
path_put(&umulti_link->path);
}
static void bpf_uprobe_multi_link_dealloc(struct bpf_link *link)
{
struct bpf_uprobe_multi_link *umulti_link;
umulti_link = container_of(link, struct bpf_uprobe_multi_link, link);
kvfree(umulti_link->uprobes);
kfree(umulti_link);
}
static int bpf_uprobe_multi_link_fill_link_info(const struct bpf_link *link,
struct bpf_link_info *info)
{
u64 __user *uref_ctr_offsets = u64_to_user_ptr(info->uprobe_multi.ref_ctr_offsets);
u64 __user *ucookies = u64_to_user_ptr(info->uprobe_multi.cookies);
u64 __user *uoffsets = u64_to_user_ptr(info->uprobe_multi.offsets);
u64 __user *upath = u64_to_user_ptr(info->uprobe_multi.path);
u32 upath_size = info->uprobe_multi.path_size;
struct bpf_uprobe_multi_link *umulti_link;
u32 ucount = info->uprobe_multi.count;
int err = 0, i;
char *p, *buf;
long left = 0;
if (!upath ^ !upath_size)
return -EINVAL;
if ((uoffsets || uref_ctr_offsets || ucookies) && !ucount)
return -EINVAL;
umulti_link = container_of(link, struct bpf_uprobe_multi_link, link);
info->uprobe_multi.count = umulti_link->cnt;
info->uprobe_multi.flags = umulti_link->link.flags;
info->uprobe_multi.pid = umulti_link->task ?
task_pid_nr_ns(umulti_link->task, task_active_pid_ns(current)) : 0;
upath_size = upath_size ? min_t(u32, upath_size, PATH_MAX) : PATH_MAX;
buf = kmalloc(upath_size, GFP_KERNEL);
if (!buf)
return -ENOMEM;
p = d_path(&umulti_link->path, buf, upath_size);
if (IS_ERR(p)) {
kfree(buf);
return PTR_ERR(p);
}
upath_size = buf + upath_size - p;
if (upath)
left = copy_to_user(upath, p, upath_size);
kfree(buf);
if (left)
return -EFAULT;
info->uprobe_multi.path_size = upath_size;
if (!uoffsets && !ucookies && !uref_ctr_offsets)
return 0;
if (ucount < umulti_link->cnt)
err = -ENOSPC;
else
ucount = umulti_link->cnt;
for (i = 0; i < ucount; i++) {
if (uoffsets &&
put_user(umulti_link->uprobes[i].offset, uoffsets + i))
return -EFAULT;
if (uref_ctr_offsets &&
put_user(umulti_link->uprobes[i].ref_ctr_offset, uref_ctr_offsets + i))
return -EFAULT;
if (ucookies &&
put_user(umulti_link->uprobes[i].cookie, ucookies + i))
return -EFAULT;
}
return err;
}
#ifdef CONFIG_PROC_FS
static void bpf_uprobe_multi_show_fdinfo(const struct bpf_link *link,
struct seq_file *seq)
{
struct bpf_uprobe_multi_link *umulti_link;
char *p, *buf;
pid_t pid;
umulti_link = container_of(link, struct bpf_uprobe_multi_link, link);
buf = kmalloc(PATH_MAX, GFP_KERNEL);
if (!buf)
return;
p = d_path(&umulti_link->path, buf, PATH_MAX);
if (IS_ERR(p)) {
kfree(buf);
return;
}
pid = umulti_link->task ?
task_pid_nr_ns(umulti_link->task, task_active_pid_ns(current)) : 0;
seq_printf(seq,
"uprobe_cnt:\t%u\n"
"pid:\t%u\n"
"path:\t%s\n",
umulti_link->cnt, pid, p);
seq_printf(seq, "%s\t %s\t %s\n", "cookie", "offset", "ref_ctr_offset");
for (int i = 0; i < umulti_link->cnt; i++) {
seq_printf(seq,
"%llu\t %#llx\t %#lx\n",
umulti_link->uprobes[i].cookie,
umulti_link->uprobes[i].offset,
umulti_link->uprobes[i].ref_ctr_offset);
}
kfree(buf);
}
#endif
static const struct bpf_link_ops bpf_uprobe_multi_link_lops = {
.release = bpf_uprobe_multi_link_release,
.dealloc_deferred = bpf_uprobe_multi_link_dealloc,
.fill_link_info = bpf_uprobe_multi_link_fill_link_info,
#ifdef CONFIG_PROC_FS
.show_fdinfo = bpf_uprobe_multi_show_fdinfo,
#endif
};
static int uprobe_prog_run(struct bpf_uprobe *uprobe,
unsigned long entry_ip,
struct pt_regs *regs,
bool is_return, void *data)
{
struct bpf_uprobe_multi_link *link = uprobe->link;
struct bpf_uprobe_multi_run_ctx run_ctx = {
.session_ctx = {
.is_return = is_return,
.data = data,
},
.entry_ip = entry_ip,
.uprobe = uprobe,
};
struct bpf_prog *prog = link->link.prog;
bool sleepable = prog->sleepable;
struct bpf_run_ctx *old_run_ctx;
int err;
if (link->task && !same_thread_group(current, link->task))
return 0;
if (sleepable)
rcu_read_lock_trace();
else
rcu_read_lock();
migrate_disable();
old_run_ctx = bpf_set_run_ctx(&run_ctx.session_ctx.run_ctx);
err = bpf_prog_run(link->link.prog, regs);
bpf_reset_run_ctx(old_run_ctx);
migrate_enable();
if (sleepable)
rcu_read_unlock_trace();
else
rcu_read_unlock();
return err;
}
static bool
uprobe_multi_link_filter(struct uprobe_consumer *con, struct mm_struct *mm)
{
struct bpf_uprobe *uprobe;
uprobe = container_of(con, struct bpf_uprobe, consumer);
return uprobe->link->task->mm == mm;
}
static int
uprobe_multi_link_handler(struct uprobe_consumer *con, struct pt_regs *regs,
__u64 *data)
{
struct bpf_uprobe *uprobe;
int ret;
uprobe = container_of(con, struct bpf_uprobe, consumer);
ret = uprobe_prog_run(uprobe, instruction_pointer(regs), regs, false, data);
if (uprobe->session)
return ret ? UPROBE_HANDLER_IGNORE : 0;
return 0;
}
static int
uprobe_multi_link_ret_handler(struct uprobe_consumer *con, unsigned long func, struct pt_regs *regs,
__u64 *data)
{
struct bpf_uprobe *uprobe;
uprobe = container_of(con, struct bpf_uprobe, consumer);
uprobe_prog_run(uprobe, func, regs, true, data);
return 0;
}
static u64 bpf_uprobe_multi_entry_ip(struct bpf_run_ctx *ctx)
{
struct bpf_uprobe_multi_run_ctx *run_ctx;
run_ctx = container_of(current->bpf_ctx, struct bpf_uprobe_multi_run_ctx,
session_ctx.run_ctx);
return run_ctx->entry_ip;
}
static u64 bpf_uprobe_multi_cookie(struct bpf_run_ctx *ctx)
{
struct bpf_uprobe_multi_run_ctx *run_ctx;
run_ctx = container_of(current->bpf_ctx, struct bpf_uprobe_multi_run_ctx,
session_ctx.run_ctx);
return run_ctx->uprobe->cookie;
}
int bpf_uprobe_multi_link_attach(const union bpf_attr *attr, struct bpf_prog *prog)
{
struct bpf_uprobe_multi_link *link = NULL;
unsigned long __user *uref_ctr_offsets;
struct bpf_link_primer link_primer;
struct bpf_uprobe *uprobes = NULL;
struct task_struct *task = NULL;
unsigned long __user *uoffsets;
u64 __user *ucookies;
void __user *upath;
u32 flags, cnt, i;
struct path path;
char *name;
pid_t pid;
int err;
/* no support for 32bit archs yet */
if (sizeof(u64) != sizeof(void *))
return -EOPNOTSUPP;
if (attr->link_create.flags)
return -EINVAL;
if (!is_uprobe_multi(prog))
return -EINVAL;
flags = attr->link_create.uprobe_multi.flags;
if (flags & ~BPF_F_UPROBE_MULTI_RETURN)
return -EINVAL;
/*
* path, offsets and cnt are mandatory,
* ref_ctr_offsets and cookies are optional
*/
upath = u64_to_user_ptr(attr->link_create.uprobe_multi.path);
uoffsets = u64_to_user_ptr(attr->link_create.uprobe_multi.offsets);
cnt = attr->link_create.uprobe_multi.cnt;
pid = attr->link_create.uprobe_multi.pid;
if (!upath || !uoffsets || !cnt || pid < 0)
return -EINVAL;
if (cnt > MAX_UPROBE_MULTI_CNT)
return -E2BIG;
uref_ctr_offsets = u64_to_user_ptr(attr->link_create.uprobe_multi.ref_ctr_offsets);
ucookies = u64_to_user_ptr(attr->link_create.uprobe_multi.cookies);
name = strndup_user(upath, PATH_MAX);
if (IS_ERR(name)) {
err = PTR_ERR(name);
return err;
}
err = kern_path(name, LOOKUP_FOLLOW, &path);
kfree(name);
if (err)
return err;
if (!d_is_reg(path.dentry)) {
err = -EBADF;
goto error_path_put;
}
if (pid) {
rcu_read_lock();
task = get_pid_task(find_vpid(pid), PIDTYPE_TGID);
rcu_read_unlock();
if (!task) {
err = -ESRCH;
goto error_path_put;
}
}
err = -ENOMEM;
link = kzalloc(sizeof(*link), GFP_KERNEL);
uprobes = kvcalloc(cnt, sizeof(*uprobes), GFP_KERNEL);
if (!uprobes || !link)
goto error_free;
for (i = 0; i < cnt; i++) {
if (__get_user(uprobes[i].offset, uoffsets + i)) {
err = -EFAULT;
goto error_free;
}
if (uprobes[i].offset < 0) {
err = -EINVAL;
goto error_free;
}
if (uref_ctr_offsets && __get_user(uprobes[i].ref_ctr_offset, uref_ctr_offsets + i)) {
err = -EFAULT;
goto error_free;
}
if (ucookies && __get_user(uprobes[i].cookie, ucookies + i)) {
err = -EFAULT;
goto error_free;
}
uprobes[i].link = link;
if (!(flags & BPF_F_UPROBE_MULTI_RETURN))
uprobes[i].consumer.handler = uprobe_multi_link_handler;
if (flags & BPF_F_UPROBE_MULTI_RETURN || is_uprobe_session(prog))
uprobes[i].consumer.ret_handler = uprobe_multi_link_ret_handler;
if (is_uprobe_session(prog))
uprobes[i].session = true;
if (pid)
uprobes[i].consumer.filter = uprobe_multi_link_filter;
}
link->cnt = cnt;
link->uprobes = uprobes;
link->path = path;
link->task = task;
link->link.flags = flags;
bpf_link_init(&link->link, BPF_LINK_TYPE_UPROBE_MULTI,
&bpf_uprobe_multi_link_lops, prog, attr->link_create.attach_type);
for (i = 0; i < cnt; i++) {
uprobes[i].uprobe = uprobe_register(d_real_inode(link->path.dentry),
uprobes[i].offset,
uprobes[i].ref_ctr_offset,
&uprobes[i].consumer);
if (IS_ERR(uprobes[i].uprobe)) {
err = PTR_ERR(uprobes[i].uprobe);
link->cnt = i;
goto error_unregister;
}
}
err = bpf_link_prime(&link->link, &link_primer);
if (err)
goto error_unregister;
return bpf_link_settle(&link_primer);
error_unregister:
bpf_uprobe_unregister(uprobes, link->cnt);
error_free:
kvfree(uprobes);
kfree(link);
if (task)
put_task_struct(task);
error_path_put:
path_put(&path);
return err;
}
#else /* !CONFIG_UPROBES */
int bpf_uprobe_multi_link_attach(const union bpf_attr *attr, struct bpf_prog *prog)
{
return -EOPNOTSUPP;
}
static u64 bpf_uprobe_multi_cookie(struct bpf_run_ctx *ctx)
{
return 0;
}
static u64 bpf_uprobe_multi_entry_ip(struct bpf_run_ctx *ctx)
{
return 0;
}
#endif /* CONFIG_UPROBES */
__bpf_kfunc_start_defs();
__bpf_kfunc bool bpf_session_is_return(void)
{
struct bpf_session_run_ctx *session_ctx;
session_ctx = container_of(current->bpf_ctx, struct bpf_session_run_ctx, run_ctx);
return session_ctx->is_return;
}
__bpf_kfunc __u64 *bpf_session_cookie(void)
{
struct bpf_session_run_ctx *session_ctx;
session_ctx = container_of(current->bpf_ctx, struct bpf_session_run_ctx, run_ctx);
return session_ctx->data;
}
__bpf_kfunc_end_defs();
BTF_KFUNCS_START(kprobe_multi_kfunc_set_ids)
BTF_ID_FLAGS(func, bpf_session_is_return)
BTF_ID_FLAGS(func, bpf_session_cookie)
BTF_KFUNCS_END(kprobe_multi_kfunc_set_ids)
static int bpf_kprobe_multi_filter(const struct bpf_prog *prog, u32 kfunc_id)
{
if (!btf_id_set8_contains(&kprobe_multi_kfunc_set_ids, kfunc_id))
return 0;
if (!is_kprobe_session(prog) && !is_uprobe_session(prog))
return -EACCES;
return 0;
}
static const struct btf_kfunc_id_set bpf_kprobe_multi_kfunc_set = {
.owner = THIS_MODULE,
.set = &kprobe_multi_kfunc_set_ids,
.filter = bpf_kprobe_multi_filter,
};
static int __init bpf_kprobe_multi_kfuncs_init(void)
{
return register_btf_kfunc_id_set(BPF_PROG_TYPE_KPROBE, &bpf_kprobe_multi_kfunc_set);
}
late_initcall(bpf_kprobe_multi_kfuncs_init);
typedef int (*copy_fn_t)(void *dst, const void *src, u32 size, struct task_struct *tsk);
/*
* The __always_inline is to make sure the compiler doesn't
* generate indirect calls into callbacks, which is expensive,
* on some kernel configurations. This allows compiler to put
* direct calls into all the specific callback implementations
* (copy_user_data_sleepable, copy_user_data_nofault, and so on)
*/
static __always_inline int __bpf_dynptr_copy_str(struct bpf_dynptr *dptr, u64 doff, u64 size,
const void *unsafe_src,
copy_fn_t str_copy_fn,
struct task_struct *tsk)
{
struct bpf_dynptr_kern *dst;
u64 chunk_sz, off;
void *dst_slice;
int cnt, err;
char buf[256];
dst_slice = bpf_dynptr_slice_rdwr(dptr, doff, NULL, size);
if (likely(dst_slice))
return str_copy_fn(dst_slice, unsafe_src, size, tsk);
dst = (struct bpf_dynptr_kern *)dptr;
if (bpf_dynptr_check_off_len(dst, doff, size))
return -E2BIG;
for (off = 0; off < size; off += chunk_sz - 1) {
chunk_sz = min_t(u64, sizeof(buf), size - off);
/* Expect str_copy_fn to return count of copied bytes, including
* zero terminator. Next iteration increment off by chunk_sz - 1 to
* overwrite NUL.
*/
cnt = str_copy_fn(buf, unsafe_src + off, chunk_sz, tsk);
if (cnt < 0)
return cnt;
err = __bpf_dynptr_write(dst, doff + off, buf, cnt, 0);
if (err)
return err;
if (cnt < chunk_sz || chunk_sz == 1) /* we are done */
return off + cnt;
}
return off;
}
static __always_inline int __bpf_dynptr_copy(const struct bpf_dynptr *dptr, u64 doff,
u64 size, const void *unsafe_src,
copy_fn_t copy_fn, struct task_struct *tsk)
{
struct bpf_dynptr_kern *dst;
void *dst_slice;
char buf[256];
u64 off, chunk_sz;
int err;
dst_slice = bpf_dynptr_slice_rdwr(dptr, doff, NULL, size);
if (likely(dst_slice))
return copy_fn(dst_slice, unsafe_src, size, tsk);
dst = (struct bpf_dynptr_kern *)dptr;
if (bpf_dynptr_check_off_len(dst, doff, size))
return -E2BIG;
for (off = 0; off < size; off += chunk_sz) {
chunk_sz = min_t(u64, sizeof(buf), size - off);
err = copy_fn(buf, unsafe_src + off, chunk_sz, tsk);
if (err)
return err;
err = __bpf_dynptr_write(dst, doff + off, buf, chunk_sz, 0);
if (err)
return err;
}
return 0;
}
static __always_inline int copy_user_data_nofault(void *dst, const void *unsafe_src,
u32 size, struct task_struct *tsk)
{
return copy_from_user_nofault(dst, (const void __user *)unsafe_src, size);
}
static __always_inline int copy_user_data_sleepable(void *dst, const void *unsafe_src,
u32 size, struct task_struct *tsk)
{
int ret;
if (!tsk) { /* Read from the current task */
ret = copy_from_user(dst, (const void __user *)unsafe_src, size);
if (ret)
return -EFAULT;
return 0;
}
ret = access_process_vm(tsk, (unsigned long)unsafe_src, dst, size, 0);
if (ret != size)
return -EFAULT;
return 0;
}
static __always_inline int copy_kernel_data_nofault(void *dst, const void *unsafe_src,
u32 size, struct task_struct *tsk)
{
return copy_from_kernel_nofault(dst, unsafe_src, size);
}
static __always_inline int copy_user_str_nofault(void *dst, const void *unsafe_src,
u32 size, struct task_struct *tsk)
{
return strncpy_from_user_nofault(dst, (const void __user *)unsafe_src, size);
}
static __always_inline int copy_user_str_sleepable(void *dst, const void *unsafe_src,
u32 size, struct task_struct *tsk)
{
int ret;
if (unlikely(size == 0))
return 0;
if (tsk) {
ret = copy_remote_vm_str(tsk, (unsigned long)unsafe_src, dst, size, 0);
} else {
ret = strncpy_from_user(dst, (const void __user *)unsafe_src, size - 1);
/* strncpy_from_user does not guarantee NUL termination */
if (ret >= 0)
((char *)dst)[ret] = '\0';
}
if (ret < 0)
return ret;
return ret + 1;
}
static __always_inline int copy_kernel_str_nofault(void *dst, const void *unsafe_src,
u32 size, struct task_struct *tsk)
{
return strncpy_from_kernel_nofault(dst, unsafe_src, size);
}
__bpf_kfunc_start_defs();
__bpf_kfunc int bpf_send_signal_task(struct task_struct *task, int sig, enum pid_type type,
u64 value)
{
if (type != PIDTYPE_PID && type != PIDTYPE_TGID)
return -EINVAL;
return bpf_send_signal_common(sig, type, task, value);
}
__bpf_kfunc int bpf_probe_read_user_dynptr(struct bpf_dynptr *dptr, u64 off,
u64 size, const void __user *unsafe_ptr__ign)
{
return __bpf_dynptr_copy(dptr, off, size, (const void *)unsafe_ptr__ign,
copy_user_data_nofault, NULL);
}
__bpf_kfunc int bpf_probe_read_kernel_dynptr(struct bpf_dynptr *dptr, u64 off,
u64 size, const void *unsafe_ptr__ign)
{
return __bpf_dynptr_copy(dptr, off, size, unsafe_ptr__ign,
copy_kernel_data_nofault, NULL);
}
__bpf_kfunc int bpf_probe_read_user_str_dynptr(struct bpf_dynptr *dptr, u64 off,
u64 size, const void __user *unsafe_ptr__ign)
{
return __bpf_dynptr_copy_str(dptr, off, size, (const void *)unsafe_ptr__ign,
copy_user_str_nofault, NULL);
}
__bpf_kfunc int bpf_probe_read_kernel_str_dynptr(struct bpf_dynptr *dptr, u64 off,
u64 size, const void *unsafe_ptr__ign)
{
return __bpf_dynptr_copy_str(dptr, off, size, unsafe_ptr__ign,
copy_kernel_str_nofault, NULL);
}
__bpf_kfunc int bpf_copy_from_user_dynptr(struct bpf_dynptr *dptr, u64 off,
u64 size, const void __user *unsafe_ptr__ign)
{
return __bpf_dynptr_copy(dptr, off, size, (const void *)unsafe_ptr__ign,
copy_user_data_sleepable, NULL);
}
__bpf_kfunc int bpf_copy_from_user_str_dynptr(struct bpf_dynptr *dptr, u64 off,
u64 size, const void __user *unsafe_ptr__ign)
{
return __bpf_dynptr_copy_str(dptr, off, size, (const void *)unsafe_ptr__ign,
copy_user_str_sleepable, NULL);
}
__bpf_kfunc int bpf_copy_from_user_task_dynptr(struct bpf_dynptr *dptr, u64 off,
u64 size, const void __user *unsafe_ptr__ign,
struct task_struct *tsk)
{
return __bpf_dynptr_copy(dptr, off, size, (const void *)unsafe_ptr__ign,
copy_user_data_sleepable, tsk);
}
__bpf_kfunc int bpf_copy_from_user_task_str_dynptr(struct bpf_dynptr *dptr, u64 off,
u64 size, const void __user *unsafe_ptr__ign,
struct task_struct *tsk)
{
return __bpf_dynptr_copy_str(dptr, off, size, (const void *)unsafe_ptr__ign,
copy_user_str_sleepable, tsk);
}
__bpf_kfunc_end_defs();
]
|
| 94/4 |
2026/05/25 22:16 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 43765
output: 23
thoughts: 424
Thoughts: **Analyzing the `bpf_trace_printk` Function**
Okay, so I've got the code for `BPF_CALL_5(bpf_trace_printk, ...)` right here. Lines 362-383, perfect. Let me walk through this. I see that it accepts a format string and some arguments, just like a standard `printf` but for BPF.
First, it defines an array, `args`, to hold the variable arguments. It looks like it's taking three arguments, `arg1`, `arg2`, and `arg3`, and storing them in the `args` array. The code leverages `MAX_TRACE_PRINTK_VARARGS` here.
Then, there's a `bpf_bprintf_data` struct being initialized, and two flags set in its fields: `get_bin_args` and `get_buf`. Next the code calls `bpf_bprintf_prepare`. This is likely where the format string is parsed, the arguments are handled, and some initial buffer setup happens. I'm assuming the return value, `ret`, indicates success or failure. I should make note of the data structure, as I'll need to know it better, later.
After preparing, the function calls `bstr_printf`. This seems to be the core formatting and string building operation. It takes the buffer, a size limit, the format string, and formatted binary arguments. The size limit `MAX_BPRINTF_BUF` is here.
Finally, `trace_bpf_trace_printk` is called on the buffer, and it cleans up the data. This looks like the call that actually sends the formatted string somewhere for tracing, and then cleans up after itself. The return value, is the result of all these things, which is returned.
I'll quickly check the constants `MAX_TRACE_PRINTK_VARARGS` and `MAX_BPRINTF_BUF` to make sure I understand the limitations. `MAX_TRACE_PRINTK_VARARGS` is straightforward, it's 3, defined on line 359, as expected. `MAX_BPRINTF_BUF` is used inside the body, so let's check its definition.
|
| 95/4 |
2026/05/25 22:16 |
tool |
grepper |
0m
Args: map[Expression:#define MAX_BPRINTF_BUF]
Results: map[Output:include/linux/bpf.h=3732=bool btf_id_set_contains(const struct btf_id_set *set, u32 id);
--
include/linux/bpf.h-3734-#define MAX_BPRINTF_VARARGS 12
include/linux/bpf.h:3735:#define MAX_BPRINTF_BUF 1024
include/linux/bpf.h-3736-
]
|
| 96/4 |
2026/05/25 22:16 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 44816
output: 407
thoughts: 238
Thoughts: **My Analysis of This `trace_printk` Implementation**
Okay, let's break this down. From what I'm seeing, this is a `trace_printk` implementation, but it's clearly leveraging a BPF-based approach for formatting and printing, presumably within the kernel tracing infrastructure. The goal, it seems, is to format some provided arguments into a string and then send that string to the tracing system.
The first thing I note is the `args` array. It's a statically sized array (`MAX_TRACE_PRINTK_VARARGS`) holding the variable arguments (`arg1`, `arg2`, `arg3`). The structure, `bpf_bprintf_data`, appears to be the central storage for this process. The presence of `.get_bin_args` and `.get_buf` suggests that the structure is used for both the formatted data and possibly intermediate data used in the formatting. `bpf_bprintf_prepare` seems crucial here, the call suggests that it processes the format string (`fmt`) and populates the `bpf_bprintf_data` structure. The `data` structure probably holds a buffer (`.buf`) for the formatted string and pre-processed data used in binary format (bin_args).
The next step uses `bstr_printf`. This strongly indicates the use of `bstr` library or similar buffer-management for safer string manipulation. I imagine it's designed to prevent buffer overflows, crucial in kernel context. It takes the buffer, its maximum size (`MAX_BPRINTF_BUF`), the format string, and the binary arguments prepared earlier.
The formatted string is then written to the trace system using `trace_bpf_trace_printk`. This likely involves some BPF-specific mechanism to output the message, rather than the regular `printk`.
Finally, the `bpf_bprintf_cleanup` call indicates memory deallocation or resource release is handled by the function.
In summary, it's a carefully crafted approach to `printk` within a BPF environment. It prioritizes safety (buffer management) and uses a two-phase formatting process, likely for efficiency and compatibility with the BPF runtime. I'm satisfied that this implementation is well-structured and takes into account important kernel-level considerations.
|
| 97/2 |
2026/05/25 22:16 |
llm |
expert |
0m
Model:
Tokens:
input: 74718
output: 41
thoughts: 1252
Thoughts: **Analyzing the Potential RCU Stall and BPF Program Vulnerability**
Okay, so I'm looking at this code and immediately see a potential issue: `bstr_printf` and `trace_bpf_trace_printk` being called multiple times β 65 to be exact β per timer interrupt. That's a lot of overhead. My initial thought is that if these operations are expensive, this could be contributing to an RCU stall since the timer interrupt is handled by `sysvec_apic_timer_interrupt`. If the interrupt takes too long, other things get delayed. This looks like a bug.
Confirming my initial hunch, it's clear that attaching a BPF program 65 times to a tracepoint triggered by the timer interrupt *and* executing expensive BPF operations will indeed cause an RCU stall. The question is, is this exploitable?
Alright, let's assess the security implications. It's a denial-of-service, but is it accessible to unprivileged processes? My understanding is that `bpf_raw_tracepoint_open` requires `CAP_PERFMON` or `CAP_SYS_ADMIN` in the initial user namespace. That's because `bpf_prog_load` requires it, and `bpf_token_create` also requires `CAP_SYS_ADMIN` because you need to mount `bpffs` with delegation options to create the necessary token.
Hmm, what if a user tries a different BPF program type? Can unprivileged users attach a BPF program to `rcu_utilization`? No, tracepoints are only accessible with `BPF_PROG_TYPE_TRACEPOINT` or `BPF_PROG_TYPE_RAW_TRACEPOINT`, both of which need `CAP_PERFMON`. What about other types, like `BPF_PROG_TYPE_SOCKET_FILTER` or `BPF_PROG_TYPE_CGROUP_SKB`? These can be loaded by unprivileged users *if* `unprivileged_bpf_disabled=0`, but they can't be attached to tracepoints. So, an unprivileged user cannot attach a BPF program to `rcu_utilization`, and that means it's not a security concern.
Wait, what if the problem is just the O(N^2) complexity of `check_buffer`? Even if it is, `check_buffer` is only enabled if `CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS` is enabled. And that is a debug option! As the documentation says, it's not for production systems. So, the O(N^2) issue wouldn't be present in a production build, so it's not a security issue either.
Wait, is the RCU stall caused by `check_buffer` or the BPF program? It could be either, but even without `check_buffer`, attaching a BPF program 65 times to a high-frequency tracepoint and using `bpf_trace_printk` might cause an RCU stall. However, since this whole thing requires `CAP_PERFMON` or `CAP_SYS_ADMIN`, it's not a security vulnerability. Root can always cause a denial of service by loading a slow BPF program or attaching it repeatedly. BPF programs are meant to allow *some* overhead, and if root does something silly like attaching 65 programs to a timer interrupt, the stall is somewhat expected.
But wait, aren't BPF programs supposed to be safe and not cause kernel crashes or hangs? Well, `bpf_trace_printk` is known to be slow. And attaching 65 programs is allowed. However, since it's a root/`CAP_PERFMON` action, it's not a security boundary bypass if it just causes a stall.
Wait, how else could `check_buffer` be triggered? If `CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS` is enabled, any trace event could trigger it. But again, it's a debug option, so it's irrelevant.
Just to be absolutely sure, let me double-check: does `bpf_prog_load` really need `CAP_PERFMON` for `BPF_PROG_TYPE_RAW_TRACEPOINT`? Okay, looking at the code, it does!
And `bpf_token_capable` checks `ns_capable(userns, CAP_PERFMON)`. If the `token` is `NULL`, the `userns` is `&init_user_ns`. So it does require `CAP_PERFMON` in the initial user namespace. Can an unprivileged user even *create* a token? Nope, mounting `bpffs` with delegation options requires
|
| 98/2 |
2026/05/25 22:17 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_event_open]
Results: map[SourceCode: SourceFile:]
|
| 99/2 |
2026/05/25 22:17 |
llm |
expert |
0m
Model:
Tokens:
input: 77057
output: 43
thoughts: 13
Thoughts: Here's my attempt at summarizing the thought process, tailored for an expert audience:
**Diving into `sys_perf_event_open`**
Okay, so we're at a point where we need to delve deeper into the internals and performance characteristics. The obvious starting point is always the fundamental system call, `sys_perf_event_open`. I'm going to start by pulling up the man page (or, more likely, the relevant kernel documentation and source code, because let's be honest, the man pages are often⦠lacking). I need to understand the nuances of the flags and the event types this thing supports.
My initial questions are the typical ones: what's the typical overhead involved in calling it? What locking is involved within the kernel when opening and configuring an event? Does it block? Is there contention in certain scenarios? I need to understand the implications for parallel systems.
I'll be paying close attention to the various modes β kernel-space only, user-space polling, interrupt-driven... The details will be crucial in deciding the best strategy for our particular performance monitoring requirements. I'm especially interested in how it handles multiplexing. I need to be sure that the implementation isn't going to introduce significant bottlenecks or unintended side effects, and of course, I am concerned about potential security implications.
Essentially, I'm going to need to build a comprehensive mental model of what's happening at every level, from user space down to the PMU hardware, to properly assess how we can leverage `sys_perf_event_open` for our purposes. It's time to dig in.
|
| 100/2 |
2026/05/25 22:17 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:sys_perf_event_open]
Results: map[SourceCode: SourceFile:]
|
| 101/2 |
2026/05/25 22:17 |
llm |
expert |
0m
Model:
Tokens:
input: 77474
output: 26
thoughts: 18
Thoughts: Here's a summary of my thought process, as requested:
**Diving into `perf_event_open`**
Okay, time to dissect `SYSCALL_DEFINE5(perf_event_open)`. Let's get right into the heart of this system call definition. I need to understand how this is structured and the arguments it handles, remembering how this is the interface for user space to create and configure performance monitoring events. Let's see what the five arguments are and how they map to the functionality. I'm especially keen to see how the kernel handles the various flags and attributes that userspace can specify, and how this relates to the perf event's lifecycle. I expect there will be plenty of pointer dereferencing and data structure manipulation in here, but let's see how efficiently it is handled. This is the entry point for all sorts of interesting performance analysis, so a thorough understanding here is key. Time to get my hands dirty.
|
| 102/2 |
2026/05/25 22:17 |
tool |
grepper |
0m
Args: map[Expression:SYSCALL_DEFINE5\(perf_event_open]
Results: map[Output:kernel/events/core.c=13417=perf_check_permission(struct perf_event_attr *attr, struct task_struct *task)
--
kernel/events/core.c-13455- */
kernel/events/core.c:13456:SYSCALL_DEFINE5(perf_event_open,
kernel/events/core.c-13457- struct perf_event_attr __user *, attr_uptr,
]
|
| 103/2 |
2026/05/25 22:17 |
llm |
expert |
0m
Model:
Tokens:
input: 77844
output: 41
thoughts: 11
Thoughts: Okay, here's my attempt at summarizing that thought process, tailored for an expert audience:
**Checking `perf_check_permission`**
Alright, let's dive in. I see that bit of code, `perf_check_permission`, and my initial thought is, "Okay, what's happening here?" It's the point where we're going to examine permissions, which, in itself, tells me we're likely dealing with some form of access control. I need to understand the granularity and the context. Is this related to file system operations, resource access, maybe even something more abstract like a database query execution?
My immediate checklist starts running:
* **Underlying Security Model:** What permissioning scheme are we using? Is it POSIX-style, ACL-based, RBAC, ABAC? Knowing this dictates how I approach the check.
* **Performance Implications:** The function name itself has "perf" in it, so I'm already mentally flagging this as a potential performance bottleneck. I need to be aware of the *frequency* of this check. Is it happening per operation, per request, or is it a one-time thing? High-frequency checks demand highly efficient implementation.
* **Concurrency:** If multiple threads or processes might be accessing this function concurrently, I'm thinking about potential race conditions and the locking mechanisms (or lack thereof) that are in place. Any thread-safety issues jump to the front.
* **Error Handling:** What happens if the permission check fails? How are errors being propagated, logged, and handled? Is there a fallback mechanism? Are we providing clear and secure feedback on permission errors?
* **Testability:** Given the code, are there adequate tests that will catch any possible edge cases? Is it possible to test permission denial/grant scenarios effectively?
So, my mental parsing here is a quick assessment of its purpose within the system, focusing on its security implications and the overall impact on the performance of the system. Iβm thinking through the different potential implications of an efficient permission check.
|
| 104/2 |
2026/05/25 22:17 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_check_permission]
Results: map[SourceCode:13416: static bool
13417: perf_check_permission(struct perf_event_attr *attr, struct task_struct *task)
13418: {
13419: unsigned int ptrace_mode = PTRACE_MODE_READ_REALCREDS;
13420: bool is_capable = perfmon_capable();
13421:
13422: if (attr->sigtrap) {
13423: /*
13424: * perf_event_attr::sigtrap sends signals to the other task.
13425: * Require the current task to also have CAP_KILL.
13426: */
13427: rcu_read_lock();
13428: is_capable &= ns_capable(__task_cred(task)->user_ns, CAP_KILL);
13429: rcu_read_unlock();
13430:
13431: /*
13432: * If the required capabilities aren't available, checks for
13433: * ptrace permissions: upgrade to ATTACH, since sending signals
13434: * can effectively change the target task.
13435: */
13436: ptrace_mode = PTRACE_MODE_ATTACH_REALCREDS;
13437: }
13438:
13439: /*
13440: * Preserve ptrace permission check for backwards compatibility. The
13441: * ptrace check also includes checks that the current task and other
13442: * task have matching uids, and is therefore not done here explicitly.
13443: */
13444: return is_capable || ptrace_may_access(task, ptrace_mode);
13445: }
SourceFile:kernel/events/core.c]
|
| 105/2 |
2026/05/25 22:17 |
llm |
expert |
0m
Model:
Tokens:
input: 78843
output: 47
thoughts: 97
Thoughts: **Unraveling Perf Event Access**
Okay, so let's break this down. My initial understanding was that `perfmon_capable()` would be the ultimate gatekeeper, checking for `CAP_PERFMON` in the initial user namespace. That seemed straightforward. But wait a second... if `perfmon_capable()` returns false, there's a fallback: `ptrace_may_access(task, ptrace_mode)`. This means, and this is the crux of the matter, that an unprivileged user *can* actually open perf events, but only for their *own* tasks! That's a key detail I missed initially.
Now, the question becomes, what *specifically* can they monitor? Can this unprivileged user open *any* type of event? Specifically, can they access `PERF_TYPE_TRACEPOINT` events? I can't assume anything; I need to dive into the implementation of `perf_event_open` to confirm. Time to go spelunking in the code.
|
| 106/2 |
2026/05/25 22:17 |
tool |
codeexpert |
3m
Args: map[Question:Does perf_event_open allow unprivileged users to open PERF_TYPE_TRACEPOINT events? Check the code in kernel/events/core.c.]
Results: map[Answer:Based on the analysis of the Linux kernel source code in `kernel/events/core.c` and related files, here is the answer to whether `perf_event_open` allows unprivileged users to open `PERF_TYPE_TRACEPOINT` events:
### Short Answer
**Yes**, `perf_event_open` allows unprivileged users to open `PERF_TYPE_TRACEPOINT` events, but they are subject to several strict security checks and functional limitations depending on the system's configuration (specifically the `perf_event_paranoid` sysctl).
### Detailed Breakdown of Security Checks
1. **Task Permissions (`perf_check_permission`)**:
When opening an event for a specific process (using a PID), the kernel checks if the current user has sufficient privileges to monitor that task. This is generally allowed if the user has `CAP_PERFMON` (or `CAP_SYS_ADMIN`) or if they have ptrace-like access to the target task (e.g., they own the process).
2. **Paranoid Level (`perf_allow_kernel`)**:
The `sysctl_perf_event_paranoid` setting (default is usually `2`) heavily restricts what unprivileged users can do:
* If `paranoid >= 2`, unprivileged users **must** set `attr.exclude_kernel = 1` in their `perf_event_attr` structure. If they don't, `perf_allow_kernel()` will return `-EACCES`.
* Since most kernel tracepoints fire in kernel mode, setting `exclude_kernel = 1` causes the event to be filtered out at runtime (in `perf_tp_event_match`), meaning the event will never fire for those tracepoints.
3. **Raw Data Access (`perf_trace_event_perm`)**:
Tracepoints often provide "raw" data (the fields of the tracepoint). Access to this data is controlled in `kernel/trace/trace_event_perf.c`:
* If the user **does not** request `PERF_SAMPLE_RAW`, the permission check is bypassed (line 76: *"No tracing, just counting, so no obvious leak"*).
* If the user **does** request `PERF_SAMPLE_RAW`, they are generally blocked unless they are privileged or `perf_event_paranoid` is set to `-1`.
* **Exception**: There is a whitelist of "safe" events that unprivileged users *can* trace with raw data if they are attached to a specific task (`PERF_ATTACH_TASK`). These events are marked with the `TRACE_EVENT_FL_CAP_ANY` flag and include **syscall tracepoints** (`sys_enter`, `sys_exit`) and **uprobes**.
4. **CPU-Wide Tracing (`perf_allow_cpu`)**:
Opening an event for all CPUs (`pid = -1`) is disallowed for unprivileged users if `sysctl_perf_event_paranoid > 0`.
### Summary of Capabilities for Unprivileged Users (at default `paranoid = 2`)
| Event Type | Can Open? | Can get Raw Data? | Will it Fire? |
| :--- | :--- | :--- | :--- |
| **Own Syscalls** | Yes | Yes | Yes (if `exclude_kernel=1`) |
| **Own Uprobes** | Yes | Yes | Yes |
| **Kernel Tracepoints** (e.g., `sched_switch`) | Yes | No | **No** (requires `exclude_kernel=1`, but fires in kernel) |
| **CPU-wide Tracepoints** | No | No | N/A |
In conclusion, while the system call technically allows unprivileged users to open tracepoint events, the combination of `perf_event_paranoid` and the `exclude_kernel` filter makes most standard kernel tracepoints functionally inaccessible or useless for them, while preserving the ability to trace their own syscalls and user-space probes.]
|
| 107/3 |
2026/05/25 22:17 |
agent |
codeexpert |
3m
Model:
Instruction:
You are a capable Linux kernel developer tasked with researching complex questions
about kernel source code. You will be given a concrete question, and need to provide
a concrete answer.
Use tools extensively while researching the question. Don't make assumptions,
or rely on your previous knowledge about the kernel source code, use available tools
to access the actual source code.
Use all available sources of information:
- kernel source code
- documentation in the Documentation dir in the source tree
- git commits descriptions, git blame
Prefer calling several tools at the same time to save round-trips.
Prompt: Does perf_event_open allow unprivileged users to open PERF_TYPE_TRACEPOINT events? Check the code in kernel/events/core.c.
Reply: Based on the analysis of the Linux kernel source code in `kernel/events/core.c` and related files, here is the answer to whether `perf_event_open` allows unprivileged users to open `PERF_TYPE_TRACEPOINT` events:
### Short Answer
**Yes**, `perf_event_open` allows unprivileged users to open `PERF_TYPE_TRACEPOINT` events, but they are subject to several strict security checks and functional limitations depending on the system's configuration (specifically the `perf_event_paranoid` sysctl).
### Detailed Breakdown of Security Checks
1. **Task Permissions (`perf_check_permission`)**:
When opening an event for a specific process (using a PID), the kernel checks if the current user has sufficient privileges to monitor that task. This is generally allowed if the user has `CAP_PERFMON` (or `CAP_SYS_ADMIN`) or if they have ptrace-like access to the target task (e.g., they own the process).
2. **Paranoid Level (`perf_allow_kernel`)**:
The `sysctl_perf_event_paranoid` setting (default is usually `2`) heavily restricts what unprivileged users can do:
* If `paranoid >= 2`, unprivileged users **must** set `attr.exclude_kernel = 1` in their `perf_event_attr` structure. If they don't, `perf_allow_kernel()` will return `-EACCES`.
* Since most kernel tracepoints fire in kernel mode, setting `exclude_kernel = 1` causes the event to be filtered out at runtime (in `perf_tp_event_match`), meaning the event will never fire for those tracepoints.
3. **Raw Data Access (`perf_trace_event_perm`)**:
Tracepoints often provide "raw" data (the fields of the tracepoint). Access to this data is controlled in `kernel/trace/trace_event_perf.c`:
* If the user **does not** request `PERF_SAMPLE_RAW`, the permission check is bypassed (line 76: *"No tracing, just counting, so no obvious leak"*).
* If the user **does** request `PERF_SAMPLE_RAW`, they are generally blocked unless they are privileged or `perf_event_paranoid` is set to `-1`.
* **Exception**: There is a whitelist of "safe" events that unprivileged users *can* trace with raw data if they are attached to a specific task (`PERF_ATTACH_TASK`). These events are marked with the `TRACE_EVENT_FL_CAP_ANY` flag and include **syscall tracepoints** (`sys_enter`, `sys_exit`) and **uprobes**.
4. **CPU-Wide Tracing (`perf_allow_cpu`)**:
Opening an event for all CPUs (`pid = -1`) is disallowed for unprivileged users if `sysctl_perf_event_paranoid > 0`.
### Summary of Capabilities for Unprivileged Users (at default `paranoid = 2`)
| Event Type | Can Open? | Can get Raw Data? | Will it Fire? |
| :--- | :--- | :--- | :--- |
| **Own Syscalls** | Yes | Yes | Yes (if `exclude_kernel=1`) |
| **Own Uprobes** | Yes | Yes | Yes |
| **Kernel Tracepoints** (e.g., `sched_switch`) | Yes | No | **No** (requires `exclude_kernel=1`, but fires in kernel) |
| **CPU-wide Tracepoints** | No | No | N/A |
In conclusion, while the system call technically allows unprivileged users to open tracepoint events, the combination of `perf_event_paranoid` and the `exclude_kernel` filter makes most standard kernel tracepoints functionally inaccessible or useless for them, while preserving the ability to trace their own syscalls and user-space probes.
|
| 108/4 |
2026/05/25 22:17 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 3229
output: 41
thoughts: 202
Thoughts: **Investigating Unprivileged Tracepoint Access with perf_event_open**
Okay, so the question is whether a regular, unprivileged user can open `PERF_TYPE_TRACEPOINT` events using `perf_event_open`. I need to figure this out, and the obvious place to start is digging into the kernel code. Specifically, I'm going to look at `kernel/events/core.c`, because that's where I'd expect to find the implementation of `perf_event_open` and the core event handling logic.
My plan is to trace the execution flow. I need to get a handle on the entry point for the `perf_event_open` system call itself. Then, I'll pay close attention to how the `perf_event_attr` structure is validated. This is crucial because itβs where event types are specified. After that, any permission checks are going to be critical β I'm expecting to find something that determines if a user is allowed to open a tracepoint.
I'll also keep my eye out for where the event actually gets initialized, something along the lines of `perf_event_alloc`, or similar. This is where the core event setup will happen. And finally, I should look for a function that's dedicated to tracepoint initialization, something like `perf_tp_event_init`, that could be involved. I'll need to see how the system handles the specific details of a tracepoint event, and if any security checks happen at that stage.
Let's begin by searching for `perf_event_open` within `kernel/events/core.c`. This should give me a solid starting point for my analysis.
|
| 109/4 |
2026/05/25 22:17 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_event_open]
Results: map[SourceCode: SourceFile:]
|
| 110/4 |
2026/05/25 22:17 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 3851
output: 25
thoughts: 65
Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience:
**Diving into the `perf_event_open` System Call Definition**
Right, `perf_event_open`. Being a system call, I'm pretty sure it'll be defined using the `SYSCALL_DEFINE` macro family. Specifically, given that it likely takes multiple arguments, `SYSCALL_DEFINE5` feels like the most probable candidate, specifying five arguments.
To confirm that hunch and pinpoint the exact implementation, I'll go directly to the source. Let's see...I'll grep for "perf_event_open" within `kernel/events/core.c`. That should quickly uncover the declaration if my assumption about the `SYSCALL_DEFINE` macro is correct and the code's organized as expected. This will allow me to precisely understand its argument list and how it interfaces with the performance monitoring infrastructure.
|
| 111/4 |
2026/05/25 22:17 |
tool |
grepper |
0m
Args: map[Expression:SYSCALL_DEFINE.*perf_event_open]
Results: map[Output:kernel/events/core.c=13417=perf_check_permission(struct perf_event_attr *attr, struct task_struct *task)
--
kernel/events/core.c-13455- */
kernel/events/core.c:13456:SYSCALL_DEFINE5(perf_event_open,
kernel/events/core.c-13457- struct perf_event_attr __user *, attr_uptr,
]
|
| 112/4 |
2026/05/25 22:17 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 4264
output: 82
thoughts: 96
Thoughts: Okay, here's my breakdown of this kernel code snippet:
**Diving into `perf_event_open` and Permissions**
Right, so I see `perf_event_open` defined at line 13456 in `kernel/events/core.c`. My immediate thought is, "Okay, that's the entry point for creating performance monitoring events," and I know that I need to understand its behavior in detail. Then, I spot `perf_check_permission` at line 13417, which definitely piques my interest. It's a pretty strong indicator that permissions are a critical part of the event creation process, which I expected.
Now, I want to trace the event's lifecycle. I'll need to scrutinize both `perf_event_open` and `perf_check_permission` to understand what's happening at each step. Beyond those functions, I'll need to figure out how the event itself gets initialized. That likely involves a call to `perf_event_alloc`, so I'll be looking for that too. Understanding the initialization stage is critical for grasping how the event gets set up and how data gets collected.
|
| 113/4 |
2026/05/25 22:17 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_check_permission]
Results: map[SourceCode:13416: static bool
13417: perf_check_permission(struct perf_event_attr *attr, struct task_struct *task)
13418: {
13419: unsigned int ptrace_mode = PTRACE_MODE_READ_REALCREDS;
13420: bool is_capable = perfmon_capable();
13421:
13422: if (attr->sigtrap) {
13423: /*
13424: * perf_event_attr::sigtrap sends signals to the other task.
13425: * Require the current task to also have CAP_KILL.
13426: */
13427: rcu_read_lock();
13428: is_capable &= ns_capable(__task_cred(task)->user_ns, CAP_KILL);
13429: rcu_read_unlock();
13430:
13431: /*
13432: * If the required capabilities aren't available, checks for
13433: * ptrace permissions: upgrade to ATTACH, since sending signals
13434: * can effectively change the target task.
13435: */
13436: ptrace_mode = PTRACE_MODE_ATTACH_REALCREDS;
13437: }
13438:
13439: /*
13440: * Preserve ptrace permission check for backwards compatibility. The
13441: * ptrace check also includes checks that the current task and other
13442: * task have matching uids, and is therefore not done here explicitly.
13443: */
13444: return is_capable || ptrace_may_access(task, ptrace_mode);
13445: }
SourceFile:kernel/events/core.c]
|
| 114/4 |
2026/05/25 22:17 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_event_alloc]
Results: map[SourceCode:12918: /*
12919: * Allocate and initialize an event structure
12920: */
12921: static struct perf_event *
12922: perf_event_alloc(struct perf_event_attr *attr, int cpu,
12923: struct task_struct *task,
12924: struct perf_event *group_leader,
12925: struct perf_event *parent_event,
12926: perf_overflow_handler_t overflow_handler,
12927: void *context, int cgroup_fd)
12928: {
12929: struct pmu *pmu;
12930: struct hw_perf_event *hwc;
12931: long err = -EINVAL;
12932: int node;
12933:
12934: if ((unsigned)cpu >= nr_cpu_ids) {
12935: if (!task || cpu != -1)
12936: return ERR_PTR(-EINVAL);
12937: }
12938: if (attr->sigtrap && !task) {
12939: /* Requires a task: avoid signalling random tasks. */
12940: return ERR_PTR(-EINVAL);
12941: }
12942:
12943: node = (cpu >= 0) ? cpu_to_node(cpu) : -1;
12944: struct perf_event *event __free(__free_event) =
12945: kmem_cache_alloc_node(perf_event_cache, GFP_KERNEL | __GFP_ZERO, node);
12946: if (!event)
12947: return ERR_PTR(-ENOMEM);
12948:
12949: /*
12950: * Single events are their own group leaders, with an
12951: * empty sibling list:
12952: */
12953: if (!group_leader)
12954: group_leader = event;
12955:
12956: mutex_init(&event->child_mutex);
12957: INIT_LIST_HEAD(&event->child_list);
12958:
12959: INIT_LIST_HEAD(&event->event_entry);
12960: INIT_LIST_HEAD(&event->sibling_list);
12961: INIT_LIST_HEAD(&event->active_list);
12962: init_event_group(event);
12963: INIT_LIST_HEAD(&event->rb_entry);
12964: INIT_LIST_HEAD(&event->active_entry);
12965: INIT_LIST_HEAD(&event->addr_filters.list);
12966: INIT_HLIST_NODE(&event->hlist_entry);
12967: INIT_LIST_HEAD(&event->pmu_list);
12968:
12969:
12970: init_waitqueue_head(&event->waitq);
12971: init_irq_work(&event->pending_irq, perf_pending_irq);
12972: event->pending_disable_irq = IRQ_WORK_INIT_HARD(perf_pending_disable);
12973: init_task_work(&event->pending_task, perf_pending_task);
12974:
12975: mutex_init(&event->mmap_mutex);
12976: raw_spin_lock_init(&event->addr_filters.lock);
12977:
12978: atomic_long_set(&event->refcount, 1);
12979: event->cpu = cpu;
12980: event->attr = *attr;
12981: event->group_leader = group_leader;
12982: event->pmu = NULL;
12983: event->oncpu = -1;
12984:
12985: event->parent = parent_event;
12986:
12987: event->ns = get_pid_ns(task_active_pid_ns(current));
12988: event->id = atomic64_inc_return(&perf_event_id);
12989:
12990: event->state = PERF_EVENT_STATE_INACTIVE;
12991:
12992: if (parent_event)
12993: event->event_caps = parent_event->event_caps;
12994:
12995: if (task) {
12996: event->attach_state = PERF_ATTACH_TASK;
12997: /*
12998: * XXX pmu::event_init needs to know what task to account to
12999: * and we cannot use the ctx information because we need the
13000: * pmu before we get a ctx.
13001: */
13002: event->hw.target = get_task_struct(task);
13003: }
13004:
13005: event->clock = &local_clock;
13006: if (parent_event)
13007: event->clock = parent_event->clock;
13008:
13009: if (!overflow_handler && parent_event) {
13010: overflow_handler = parent_event->overflow_handler;
13011: context = parent_event->overflow_handler_context;
13012: #if defined(CONFIG_BPF_SYSCALL) && defined(CONFIG_EVENT_TRACING)
13013: if (parent_event->prog) {
13014: struct bpf_prog *prog = parent_event->prog;
13015:
13016: bpf_prog_inc(prog);
13017: event->prog = prog;
13018: }
13019: #endif
13020: }
13021:
13022: if (overflow_handler) {
13023: event->overflow_handler = overflow_handler;
13024: event->overflow_handler_context = context;
13025: } else if (is_write_backward(event)){
13026: event->overflow_handler = perf_event_output_backward;
13027: event->overflow_handler_context = NULL;
13028: } else {
13029: event->overflow_handler = perf_event_output_forward;
13030: event->overflow_handler_context = NULL;
13031: }
13032:
13033: perf_event__state_init(event);
13034:
13035: pmu = NULL;
13036:
13037: hwc = &event->hw;
13038: hwc->sample_period = attr->sample_period;
13039: if (is_event_in_freq_mode(event))
13040: hwc->sample_period = 1;
13041: hwc->last_period = hwc->sample_period;
13042:
13043: local64_set(&hwc->period_left, hwc->sample_period);
13044:
13045: /*
13046: * We do not support PERF_SAMPLE_READ on inherited events unless
13047: * PERF_SAMPLE_TID is also selected, which allows inherited events to
13048: * collect per-thread samples.
13049: * See perf_output_read().
13050: */
13051: if (has_inherit_and_sample_read(attr) && !(attr->sample_type & PERF_SAMPLE_TID))
13052: return ERR_PTR(-EINVAL);
13053:
13054: if (!has_branch_stack(event))
13055: event->attr.branch_sample_type = 0;
13056:
13057: pmu = perf_init_event(event);
13058: if (IS_ERR(pmu))
13059: return (void*)pmu;
13060:
13061: /*
13062: * The PERF_ATTACH_TASK_DATA is set in the event_init()->hw_config().
13063: * The attach should be right after the perf_init_event().
13064: * Otherwise, the __free_event() would mistakenly detach the non-exist
13065: * perf_ctx_data because of the other errors between them.
13066: */
13067: if (event->attach_state & PERF_ATTACH_TASK_DATA) {
13068: err = attach_perf_ctx_data(event);
13069: if (err)
13070: return ERR_PTR(err);
13071: }
13072:
13073: /*
13074: * Disallow uncore-task events. Similarly, disallow uncore-cgroup
13075: * events (they don't make sense as the cgroup will be different
13076: * on other CPUs in the uncore mask).
13077: */
13078: if (pmu->task_ctx_nr == perf_invalid_context && (task || cgroup_fd != -1))
13079: return ERR_PTR(-EINVAL);
13080:
13081: if (event->attr.aux_output &&
13082: (!(pmu->capabilities & PERF_PMU_CAP_AUX_OUTPUT) ||
13083: event->attr.aux_pause || event->attr.aux_resume))
13084: return ERR_PTR(-EOPNOTSUPP);
13085:
13086: if (event->attr.aux_pause && event->attr.aux_resume)
13087: return ERR_PTR(-EINVAL);
13088:
13089: if (event->attr.aux_start_paused) {
13090: if (!(pmu->capabilities & PERF_PMU_CAP_AUX_PAUSE))
13091: return ERR_PTR(-EOPNOTSUPP);
13092: event->hw.aux_paused = 1;
13093: }
13094:
13095: if (cgroup_fd != -1) {
13096: err = perf_cgroup_connect(cgroup_fd, event, attr, group_leader);
13097: if (err)
13098: return ERR_PTR(err);
13099: }
13100:
13101: err = exclusive_event_init(event);
13102: if (err)
13103: return ERR_PTR(err);
13104:
13105: if (has_addr_filter(event)) {
13106: event->addr_filter_ranges = kcalloc(pmu->nr_addr_filters,
13107: sizeof(struct perf_addr_filter_range),
13108: GFP_KERNEL);
13109: if (!event->addr_filter_ranges)
13110: return ERR_PTR(-ENOMEM);
13111:
13112: /*
13113: * Clone the parent's vma offsets: they are valid until exec()
13114: * even if the mm is not shared with the parent.
13115: */
13116: if (event->parent) {
13117: struct perf_addr_filters_head *ifh = perf_event_addr_filters(event);
13118:
13119: raw_spin_lock_irq(&ifh->lock);
13120: memcpy(event->addr_filter_ranges,
13121: event->parent->addr_filter_ranges,
13122: pmu->nr_addr_filters * sizeof(struct perf_addr_filter_range));
13123: raw_spin_unlock_irq(&ifh->lock);
13124: }
13125:
13126: /* force hw sync on the address filters */
13127: event->addr_filters_gen = 1;
13128: }
13129:
13130: if (!event->parent) {
13131: if (event->attr.sample_type & PERF_SAMPLE_CALLCHAIN) {
13132: err = get_callchain_buffers(attr->sample_max_stack);
13133: if (err)
13134: return ERR_PTR(err);
13135: event->attach_state |= PERF_ATTACH_CALLCHAIN;
13136: }
13137: }
13138:
13139: err = security_perf_event_alloc(event);
13140: if (err)
13141: return ERR_PTR(err);
13142:
13143: /* symmetric to unaccount_event() in _free_event() */
13144: account_event(event);
13145:
13146: /*
13147: * Event creation should be under SRCU, see perf_pmu_unregister().
13148: */
13149: lockdep_assert_held(&pmus_srcu);
13150: scoped_guard (spinlock, &pmu->events_lock)
13151: list_add(&event->pmu_list, &pmu->events);
13152:
13153: return_ptr(event);
13154: }
SourceFile:kernel/events/core.c]
|
| 115/4 |
2026/05/25 22:17 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 9098
output: 41
thoughts: 41
Thoughts: Okay, here's my take on those thoughts, framed as an expert's internal monologue:
**Diving into `perf_event_alloc` and `perf_init_event`**
Right, so I'm looking at `perf_event_alloc`, and the line that immediately grabs my attention is 13057: `perf_init_event(event)`. That's where things get interesting. Allocating a performance event, huh? Of course, the initialization is critical. I'd expect that the `perf_init_event` function is setting up the core data structures and perhaps assigning some default values, maybe handling some early error checking. I need to take a look at *that* function now, see what's happening under the hood. What are the key initializations? Are there any obvious potential performance bottlenecks at this stage? And, really, what are the implications of how itβs being called within `perf_event_alloc`? Does it interact with the event's lifecycle in a way that suggests a certain architecture? Let's get to it.
|
| 116/4 |
2026/05/25 22:17 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_init_event]
Results: map[SourceCode:12737: static struct pmu *perf_init_event(struct perf_event *event)
12738: {
12739: bool extended_type = false;
12740: struct pmu *pmu;
12741: int type, ret;
12742:
12743: guard(srcu)(&pmus_srcu); /* pmu idr/list access */
12744:
12745: /*
12746: * Save original type before calling pmu->event_init() since certain
12747: * pmus overwrites event->attr.type to forward event to another pmu.
12748: */
12749: event->orig_type = event->attr.type;
12750:
12751: /* Try parent's PMU first: */
12752: if (event->parent && event->parent->pmu) {
12753: pmu = event->parent->pmu;
12754: ret = perf_try_init_event(pmu, event);
12755: if (!ret)
12756: return pmu;
12757: }
12758:
12759: /*
12760: * PERF_TYPE_HARDWARE and PERF_TYPE_HW_CACHE
12761: * are often aliases for PERF_TYPE_RAW.
12762: */
12763: type = event->attr.type;
12764: if (type == PERF_TYPE_HARDWARE || type == PERF_TYPE_HW_CACHE) {
12765: type = event->attr.config >> PERF_PMU_TYPE_SHIFT;
12766: if (!type) {
12767: type = PERF_TYPE_RAW;
12768: } else {
12769: extended_type = true;
12770: event->attr.config &= PERF_HW_EVENT_MASK;
12771: }
12772: }
12773:
12774: again:
12775: scoped_guard (rcu)
12776: pmu = idr_find(&pmu_idr, type);
12777: if (pmu) {
12778: if (event->attr.type != type && type != PERF_TYPE_RAW &&
12779: !(pmu->capabilities & PERF_PMU_CAP_EXTENDED_HW_TYPE))
12780: return ERR_PTR(-ENOENT);
12781:
12782: ret = perf_try_init_event(pmu, event);
12783: if (ret == -ENOENT && event->attr.type != type && !extended_type) {
12784: type = event->attr.type;
12785: goto again;
12786: }
12787:
12788: if (ret)
12789: return ERR_PTR(ret);
12790:
12791: return pmu;
12792: }
12793:
12794: list_for_each_entry_rcu(pmu, &pmus, entry, lockdep_is_held(&pmus_srcu)) {
12795: ret = perf_try_init_event(pmu, event);
12796: if (!ret)
12797: return pmu;
12798:
12799: if (ret != -ENOENT)
12800: return ERR_PTR(ret);
12801: }
12802:
12803: return ERR_PTR(-ENOENT);
12804: }
SourceFile:kernel/events/core.c]
|
| 117/4 |
2026/05/25 22:17 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 10497
output: 21
thoughts: 98
Thoughts: **Tracing the Tracepoint Initialization**
Okay, so I'm looking into how tracepoint events are initialized within the performance monitoring subsystem. From what I can gather, `perf_init_event` kicks things off, ultimately calling `perf_try_init_event`. And *that* function, in turn, seems to delegate the core initialization to the relevant Performance Monitoring Unit, or PMU, via `pmu->event_init(event)`.
My immediate goal now is to pinpoint which PMU is responsible for handling `PERF_TYPE_TRACEPOINT` events. It's usually a dedicated PMU, but I need to confirm this and understand the specific mechanisms at play. So, I need to get searching for instances of `PERF_TYPE_TRACEPOINT` within the kernel code. Time to dig in and see where this constant is actually used.
|
| 118/4 |
2026/05/25 22:17 |
tool |
grepper |
0m
Args: map[Expression:PERF_TYPE_TRACEPOINT]
Results: map[Output:include/uapi/linux/perf_event.h=29=enum perf_type_id {
--
include/uapi/linux/perf_event.h-31- PERF_TYPE_SOFTWARE = 1,
include/uapi/linux/perf_event.h:32: PERF_TYPE_TRACEPOINT = 2,
include/uapi/linux/perf_event.h-33- PERF_TYPE_HW_CACHE = 3,
--
kernel/events/core.c=10928=static int perf_tp_event_init(struct perf_event *event)
--
kernel/events/core.c-10931-
kernel/events/core.c:10932: if (event->attr.type != PERF_TYPE_TRACEPOINT)
kernel/events/core.c-10933- return -ENOENT;
--
kernel/events/core.c=11249=static inline void perf_tp_register(void)
kernel/events/core.c-11250-{
kernel/events/core.c:11251: perf_pmu_register(&perf_tracepoint, "tracepoint", PERF_TYPE_TRACEPOINT);
kernel/events/core.c-11252-#ifdef CONFIG_KPROBE_EVENTS
--
kernel/trace/bpf_trace.c=1993=int perf_event_query_prog_array(struct perf_event *event, void __user *info)
--
kernel/trace/bpf_trace.c-2002- return -EPERM;
kernel/trace/bpf_trace.c:2003: if (event->attr.type != PERF_TYPE_TRACEPOINT)
kernel/trace/bpf_trace.c-2004- return -EINVAL;
--
kernel/trace/bpf_trace.c=2151=int bpf_get_perf_event_info(const struct perf_event *event, u32 *prog_id,
--
kernel/trace/bpf_trace.c-2189- probe_offset, probe_addr, missed,
kernel/trace/bpf_trace.c:2190: event->attr.type == PERF_TYPE_TRACEPOINT);
kernel/trace/bpf_trace.c-2191-#endif
--
kernel/trace/bpf_trace.c-2195- probe_offset, probe_addr,
kernel/trace/bpf_trace.c:2196: event->attr.type == PERF_TYPE_TRACEPOINT);
kernel/trace/bpf_trace.c-2197-#endif
--
samples/bpf/task_fd_query_user.c=227=static int test_debug_fs_uprobe(char *binary_path, long offset, bool is_return)
--
samples/bpf/task_fd_query_user.c-265- attr.config = strtol(buf, NULL, 0);
samples/bpf/task_fd_query_user.c:266: attr.type = PERF_TYPE_TRACEPOINT;
samples/bpf/task_fd_query_user.c-267- attr.sample_period = 1;
--
tools/bpf/bpftool/link.c=25=static const char *perf_type_name[PERF_TYPE_MAX] = {
--
tools/bpf/bpftool/link.c-27- [PERF_TYPE_SOFTWARE] = "software",
tools/bpf/bpftool/link.c:28: [PERF_TYPE_TRACEPOINT] = "tracepoint",
tools/bpf/bpftool/link.c-29- [PERF_TYPE_HW_CACHE] = "hw-cache",
--
tools/include/uapi/linux/perf_event.h=29=enum perf_type_id {
--
tools/include/uapi/linux/perf_event.h-31- PERF_TYPE_SOFTWARE = 1,
tools/include/uapi/linux/perf_event.h:32: PERF_TYPE_TRACEPOINT = 2,
tools/include/uapi/linux/perf_event.h-33- PERF_TYPE_HW_CACHE = 3,
--
tools/kvm/kvm_stat/kvm_stat=402=class perf_event_attr(ctypes.Structure):
--
tools/kvm/kvm_stat/kvm_stat-425- super(self.__class__, self).__init__()
tools/kvm/kvm_stat/kvm_stat:426: self.type = PERF_TYPE_TRACEPOINT
tools/kvm/kvm_stat/kvm_stat-427- self.size = ctypes.sizeof(self)
--
tools/kvm/kvm_stat/kvm_stat-430-
tools/kvm/kvm_stat/kvm_stat:431:PERF_TYPE_TRACEPOINT = 2
tools/kvm/kvm_stat/kvm_stat-432-PERF_FORMAT_GROUP = 1 << 3
--
tools/lib/bpf/libbpf.c=11520=static int perf_event_kprobe_open_legacy(const char *probe_name, bool retprobe,
--
tools/lib/bpf/libbpf.c-11545- attr.config = type;
tools/lib/bpf/libbpf.c:11546: attr.type = PERF_TYPE_TRACEPOINT;
tools/lib/bpf/libbpf.c-11547-
--
tools/lib/bpf/libbpf.c=12242=static int perf_event_uprobe_open_legacy(const char *probe_name, bool retprobe,
--
tools/lib/bpf/libbpf.c-12265- attr.config = type;
tools/lib/bpf/libbpf.c:12266: attr.type = PERF_TYPE_TRACEPOINT;
tools/lib/bpf/libbpf.c-12267-
--
tools/lib/bpf/libbpf.c=12865=static int perf_event_open_tracepoint(const char *tp_category,
--
tools/lib/bpf/libbpf.c-12880- memset(&attr, 0, attr_sz);
tools/lib/bpf/libbpf.c:12881: attr.type = PERF_TYPE_TRACEPOINT;
tools/lib/bpf/libbpf.c-12882- attr.size = attr_sz;
--
tools/lib/perf/tests/test-evlist.c=213=static int test_mmap_thread(void)
--
tools/lib/perf/tests/test-evlist.c-220- struct perf_event_attr attr = {
tools/lib/perf/tests/test-evlist.c:221: .type = PERF_TYPE_TRACEPOINT,
tools/lib/perf/tests/test-evlist.c-222- .sample_period = 1,
--
tools/lib/perf/tests/test-evlist.c=324=static int test_mmap_cpus(void)
--
tools/lib/perf/tests/test-evlist.c-330- struct perf_event_attr attr = {
tools/lib/perf/tests/test-evlist.c:331: .type = PERF_TYPE_TRACEPOINT,
tools/lib/perf/tests/test-evlist.c-332- .sample_period = 1,
--
tools/perf/builtin-evlist.c=30=static int __cmd_evlist(const char *file_name, struct perf_attr_details *details)
--
tools/perf/builtin-evlist.c-55-
tools/perf/builtin-evlist.c:56: if (pos->core.attr.type == PERF_TYPE_TRACEPOINT)
tools/perf/builtin-evlist.c-57- has_tracepoint = true;
--
tools/perf/builtin-script.c=233=static struct {
--
tools/perf/builtin-script.c-267-
tools/perf/builtin-script.c:268: [PERF_TYPE_TRACEPOINT] = {
tools/perf/builtin-script.c-269- .user_set = false,
--
tools/perf/builtin-script.c=637=static int perf_session__check_output_opt(struct perf_session *session)
--
tools/perf/builtin-script.c-692- if (symbol_conf.use_callchain &&
tools/perf/builtin-script.c:693: !output[PERF_TYPE_TRACEPOINT].user_set) {
tools/perf/builtin-script.c:694: j = PERF_TYPE_TRACEPOINT;
tools/perf/builtin-script.c-695-
--
tools/perf/builtin-script.c=3320=static int parse_output_fields(const struct option *opt __maybe_unused,
--
tools/perf/builtin-script.c-3346- else if (!strcmp(str, "trace"))
tools/perf/builtin-script.c:3347: type = PERF_TYPE_TRACEPOINT;
tools/perf/builtin-script.c-3348- else if (!strcmp(str, "raw"))
--
tools/perf/builtin-trace.c=3620=static void trace__handle_event(struct trace *trace, union perf_event *event, struct perf_sample *sample)
--
tools/perf/builtin-trace.c-3640-
tools/perf/builtin-trace.c:3641: if (evsel->core.attr.type == PERF_TYPE_TRACEPOINT &&
tools/perf/builtin-trace.c-3642- sample->raw_data == NULL) {
--
tools/perf/design.txt=85=enum perf_type_id {
--
tools/perf/design.txt-87- PERF_TYPE_SOFTWARE = 1,
tools/perf/design.txt:88: PERF_TYPE_TRACEPOINT = 2,
tools/perf/design.txt-89-};
--
tools/perf/design.txt=135=enum perf_sw_ids {
--
tools/perf/design.txt-146-
tools/perf/design.txt:147:Counters of the type PERF_TYPE_TRACEPOINT are available when the ftrace event
tools/perf/design.txt-148-tracer is available, and event_id values can be obtained from
--
tools/perf/tests/parse-events.c=108=static int test__checkevent_tracepoint(struct evlist *evlist)
--
tools/perf/tests/parse-events.c-113- TEST_ASSERT_EVLIST("wrong number of groups", 0 == evlist__nr_groups(evlist), evlist);
tools/perf/tests/parse-events.c:114: TEST_ASSERT_EVSEL("wrong type", PERF_TYPE_TRACEPOINT == evsel->core.attr.type, evsel);
tools/perf/tests/parse-events.c-115- TEST_ASSERT_EVSEL("wrong sample_type",
--
tools/perf/tests/parse-events.c=121=static int test__checkevent_tracepoint_multi(struct evlist *evlist)
--
tools/perf/tests/parse-events.c-129- TEST_ASSERT_EVSEL("wrong type",
tools/perf/tests/parse-events.c:130: PERF_TYPE_TRACEPOINT == evsel->core.attr.type,
tools/perf/tests/parse-events.c-131- evsel);
--
tools/perf/tests/parse-events.c=660=static int test__checkevent_list(struct evlist *evlist)
--
tools/perf/tests/parse-events.c-666- /* r1 */
tools/perf/tests/parse-events.c:667: TEST_ASSERT_EVSEL("wrong type", PERF_TYPE_TRACEPOINT != evsel->core.attr.type, evsel);
tools/perf/tests/parse-events.c:668: while (evsel->core.attr.type != PERF_TYPE_TRACEPOINT) {
tools/perf/tests/parse-events.c-669- TEST_ASSERT_EVSEL("wrong config", 1 == evsel->core.attr.config, evsel);
--
tools/perf/tests/parse-events.c-681- /* syscalls:sys_enter_openat:k */
tools/perf/tests/parse-events.c:682: TEST_ASSERT_EVSEL("wrong type", PERF_TYPE_TRACEPOINT == evsel->core.attr.type, evsel);
tools/perf/tests/parse-events.c-683- TEST_ASSERT_EVSEL("wrong sample_type", PERF_TP_SAMPLE_TYPE == evsel->core.attr.sample_type,
--
tools/perf/tests/parse-events.c=1016=static int test__group3(struct evlist *evlist __maybe_unused)
--
tools/perf/tests/parse-events.c-1030- evlist__for_each_entry(evlist, evsel) {
tools/perf/tests/parse-events.c:1031: if (evsel->core.attr.type == PERF_TYPE_TRACEPOINT) {
tools/perf/tests/parse-events.c-1032- /* group1 syscalls:sys_enter_openat:H */
--
tools/perf/util/data-convert-bt.c=796=static int process_sample_event(const struct perf_tool *tool,
--
tools/perf/util/data-convert-bt.c-833-
tools/perf/util/data-convert-bt.c:834: if (evsel->core.attr.type == PERF_TYPE_TRACEPOINT) {
tools/perf/util/data-convert-bt.c-835- ret = add_tracepoint_values(cw, event_class, event,
--
tools/perf/util/data-convert-bt.c=1165=static int add_event(struct ctf_writer *cw, struct evsel *evsel)
--
tools/perf/util/data-convert-bt.c-1181-
tools/perf/util/data-convert-bt.c:1182: if (evsel->core.attr.type == PERF_TYPE_TRACEPOINT) {
tools/perf/util/data-convert-bt.c-1183- ret = add_tracepoint_types(cw, evsel, event_class);
--
tools/perf/util/evlist.c=335=struct evsel *evlist__find_tracepoint_by_name(struct evlist *evlist, const char *name)
--
tools/perf/util/evlist.c-339- evlist__for_each_entry(evlist, evsel) {
tools/perf/util/evlist.c:340: if ((evsel->core.attr.type == PERF_TYPE_TRACEPOINT) &&
tools/perf/util/evlist.c-341- (strcmp(evsel->name, name) == 0))
--
tools/perf/util/evlist.c=1083=int evlist__set_tp_filter(struct evlist *evlist, const char *filter)
--
tools/perf/util/evlist.c-1091- evlist__for_each_entry(evlist, evsel) {
tools/perf/util/evlist.c:1092: if (evsel->core.attr.type != PERF_TYPE_TRACEPOINT)
tools/perf/util/evlist.c-1093- continue;
--
tools/perf/util/evlist.c=1103=int evlist__append_tp_filter(struct evlist *evlist, const char *filter)
--
tools/perf/util/evlist.c-1111- evlist__for_each_entry(evlist, evsel) {
tools/perf/util/evlist.c:1112: if (evsel->core.attr.type != PERF_TYPE_TRACEPOINT)
tools/perf/util/evlist.c-1113- continue;
--
tools/perf/util/evsel.c=580=struct evsel *evsel__newtp_idx(const char *sys, const char *name, int idx, bool format)
--
tools/perf/util/evsel.c-582- struct perf_event_attr attr = {
tools/perf/util/evsel.c:583: .type = PERF_TYPE_TRACEPOINT,
tools/perf/util/evsel.c-584- .sample_type = (PERF_SAMPLE_RAW | PERF_SAMPLE_TIME |
--
tools/perf/util/evsel.c=633=struct tep_event *evsel__tp_format(struct evsel *evsel)
--
tools/perf/util/evsel.c-639-
tools/perf/util/evsel.c:640: if (evsel->core.attr.type != PERF_TYPE_TRACEPOINT)
tools/perf/util/evsel.c-641- return NULL;
--
tools/perf/util/evsel.c=899=const char *evsel__name(struct evsel *evsel)
--
tools/perf/util/evsel.c-925-
tools/perf/util/evsel.c:926: case PERF_TYPE_TRACEPOINT:
tools/perf/util/evsel.c-927- scnprintf(bf, sizeof(bf), "%s", "unknown tracepoint");
--
tools/perf/util/evsel_fprintf.c=43=int evsel__fprintf(struct evsel *evsel, struct perf_attr_details *details, FILE *fp)
--
tools/perf/util/evsel_fprintf.c-85-
tools/perf/util/evsel_fprintf.c:86: if (evsel->core.attr.type != PERF_TYPE_TRACEPOINT) {
tools/perf/util/evsel_fprintf.c-87- printed += comma_fprintf(fp, &first, " (not a tracepoint)");
--
tools/perf/util/header.c=4215=static int evlist__prepare_tracepoint_events(struct evlist *evlist, struct tep_handle *pevent)
--
tools/perf/util/header.c-4219- evlist__for_each_entry(evlist, pos) {
tools/perf/util/header.c:4220: if (pos->core.attr.type == PERF_TYPE_TRACEPOINT &&
tools/perf/util/header.c-4221- evsel__prepare_tracepoint_event(pos, pevent))
--
tools/perf/util/parse-events.c=46=static const char *const event_types[] = {
--
tools/perf/util/parse-events.c-48- [PERF_TYPE_SOFTWARE] = "software",
tools/perf/util/parse-events.c:49: [PERF_TYPE_TRACEPOINT] = "tracepoint",
tools/perf/util/parse-events.c-50- [PERF_TYPE_HW_CACHE] = "hardware-cache",
--
tools/perf/util/parse-events.c=2491=static int set_filter(struct evsel *evsel, const void *arg)
--
tools/perf/util/parse-events.c-2502-
tools/perf/util/parse-events.c:2503: if (evsel->core.attr.type == PERF_TYPE_TRACEPOINT && is_possible_tp_filter(str)) {
tools/perf/util/parse-events.c-2504- if (evsel__append_tp_filter(evsel, str) < 0) {
--
tools/perf/util/parse-events.c=2563=static int add_exclude_perf_filter(struct evsel *evsel,
--
tools/perf/util/parse-events.c-2567-
tools/perf/util/parse-events.c:2568: if (evsel == NULL || evsel->core.attr.type != PERF_TYPE_TRACEPOINT) {
tools/perf/util/parse-events.c-2569- fprintf(stderr,
--
tools/perf/util/perf_event_attr_fprintf.c=80=static const char *stringify_perf_type_id(struct perf_pmu *pmu, u32 type)
--
tools/perf/util/perf_event_attr_fprintf.c-84- ENUM_ID_TO_STR_CASE(PERF_TYPE_SOFTWARE)
tools/perf/util/perf_event_attr_fprintf.c:85: ENUM_ID_TO_STR_CASE(PERF_TYPE_TRACEPOINT)
tools/perf/util/perf_event_attr_fprintf.c-86- ENUM_ID_TO_STR_CASE(PERF_TYPE_HW_CACHE)
--
tools/perf/util/perf_event_attr_fprintf.c=249=static void __p_config_id(struct perf_pmu *pmu, char *buf, size_t size, u32 type, u64 config)
--
tools/perf/util/perf_event_attr_fprintf.c-257- return __p_config_hw_cache_id(buf, size, pmu, config);
tools/perf/util/perf_event_attr_fprintf.c:258: case PERF_TYPE_TRACEPOINT:
tools/perf/util/perf_event_attr_fprintf.c-259- return __p_config_tracepoint_id(buf, size, config);
--
tools/perf/util/pmu.c=1222=static __u32 wellknown_pmu_type(const char *pmu_name)
--
tools/perf/util/pmu.c-1233- "tracepoint",
tools/perf/util/pmu.c:1234: PERF_TYPE_TRACEPOINT
tools/perf/util/pmu.c-1235- },
--
tools/perf/util/pmu.c=2396=bool perf_pmu__is_software(const struct perf_pmu *pmu)
--
tools/perf/util/pmu.c-2408- case PERF_TYPE_SOFTWARE: return true;
tools/perf/util/pmu.c:2409: case PERF_TYPE_TRACEPOINT: return true;
tools/perf/util/pmu.c-2410- case PERF_TYPE_HW_CACHE: return false;
--
tools/perf/util/print-events.c=47=void print_sdt_events(const struct print_callbacks *print_cb, void *print_state)
--
tools/perf/util/print-events.c-123- /*pmu_name=*/NULL,
tools/perf/util/print-events.c:124: PERF_TYPE_TRACEPOINT,
tools/perf/util/print-events.c-125- evt_name ?: sdt_name->s,
--
tools/perf/util/python.c=293=static bool is_tracepoint(const struct pyrf_event *pevent)
tools/perf/util/python.c-294-{
tools/perf/util/python.c:295: return pevent->evsel->core.attr.type == PERF_TYPE_TRACEPOINT;
tools/perf/util/python.c-296-}
--
tools/perf/util/scripting-engines/trace-event-perl.c=342=static void perl_process_tracepoint(struct perf_sample *sample,
--
tools/perf/util/scripting-engines/trace-event-perl.c-361-
tools/perf/util/scripting-engines/trace-event-perl.c:362: if (evsel->core.attr.type != PERF_TYPE_TRACEPOINT)
tools/perf/util/scripting-engines/trace-event-perl.c-363- return;
--
tools/perf/util/scripting-engines/trace-event-python.c=1495=static void python_process_event(union perf_event *event,
--
tools/perf/util/scripting-engines/trace-event-python.c-1505- switch (evsel->core.attr.type) {
tools/perf/util/scripting-engines/trace-event-python.c:1506: case PERF_TYPE_TRACEPOINT:
tools/perf/util/scripting-engines/trace-event-python.c-1507- python_process_tracepoint(sample, evsel, al, addr_al);
--
tools/perf/util/session.c=2596=bool perf_session__has_traces(struct perf_session *session, const char *msg)
--
tools/perf/util/session.c-2600- evlist__for_each_entry(session->evlist, evsel) {
tools/perf/util/session.c:2601: if (evsel->core.attr.type == PERF_TYPE_TRACEPOINT)
tools/perf/util/session.c-2602- return true;
--
tools/perf/util/sort.c=1131=sort__trace_cmp(struct hist_entry *left, struct hist_entry *right)
--
tools/perf/util/sort.c-1135- evsel = hists_to_evsel(left->hists);
tools/perf/util/sort.c:1136: if (evsel->core.attr.type != PERF_TYPE_TRACEPOINT)
tools/perf/util/sort.c-1137- return 0;
--
tools/perf/util/sort.c=1147=static int hist_entry__trace_snprintf(struct hist_entry *he, char *bf,
--
tools/perf/util/sort.c-1152- evsel = hists_to_evsel(he->hists);
tools/perf/util/sort.c:1153: if (evsel->core.attr.type != PERF_TYPE_TRACEPOINT)
tools/perf/util/sort.c-1154- return scnprintf(bf, size, "%-.*s", width, "N/A");
--
tools/perf/util/sort.c=3403=static int add_all_dynamic_fields(struct evlist *evlist, bool raw_trace,
--
tools/perf/util/sort.c-3409- evlist__for_each_entry(evlist, evsel) {
tools/perf/util/sort.c:3410: if (evsel->core.attr.type != PERF_TYPE_TRACEPOINT)
tools/perf/util/sort.c-3411- continue;
--
tools/perf/util/sort.c=3420=static int add_all_matching_fields(struct evlist *evlist,
--
tools/perf/util/sort.c-3429-
tools/perf/util/sort.c:3430: if (evsel->core.attr.type != PERF_TYPE_TRACEPOINT)
tools/perf/util/sort.c-3431- continue;
--
tools/perf/util/sort.c=3449=static int add_dynamic_entry(struct evlist *evlist, const char *tok,
--
tools/perf/util/sort.c-3489- evlist__for_each_entry(evlist, evsel) {
tools/perf/util/sort.c:3490: if (evsel->core.attr.type == PERF_TYPE_TRACEPOINT) {
tools/perf/util/sort.c-3491- pr_err("%s %s", ret ? "," : "This perf binary isn't linked with libtraceevent, can't process", evsel__name(evsel));
--
tools/perf/util/sort.c-3508-
tools/perf/util/sort.c:3509: if (evsel->core.attr.type != PERF_TYPE_TRACEPOINT) {
tools/perf/util/sort.c-3510- pr_debug("%s is not a tracepoint event\n", event_name);
--
tools/perf/util/sort.c=3834=static const char *get_default_sort_order(struct evlist *evlist)
--
tools/perf/util/sort.c-3852- evlist__for_each_entry(evlist, evsel) {
tools/perf/util/sort.c:3853: if (evsel->core.attr.type != PERF_TYPE_TRACEPOINT) {
tools/perf/util/sort.c-3854- use_trace = false;
--
tools/perf/util/tp_pmu.c=96=bool perf_pmu__is_tracepoint(const struct perf_pmu *pmu)
tools/perf/util/tp_pmu.c-97-{
tools/perf/util/tp_pmu.c:98: return pmu->type == PERF_TYPE_TRACEPOINT;
tools/perf/util/tp_pmu.c-99-}
--
tools/perf/util/trace-event-info.c=505=get_tracepoints_path(struct list_head *pattrs)
--
tools/perf/util/trace-event-info.c-511- list_for_each_entry(pos, pattrs, core.node) {
tools/perf/util/trace-event-info.c:512: if (pos->core.attr.type != PERF_TYPE_TRACEPOINT)
tools/perf/util/trace-event-info.c-513- continue;
--
tools/perf/util/trace-event-info.c=542=bool have_tracepoints(struct list_head *pattrs)
--
tools/perf/util/trace-event-info.c-546- list_for_each_entry(pos, pattrs, core.node)
tools/perf/util/trace-event-info.c:547: if (pos->core.attr.type == PERF_TYPE_TRACEPOINT)
tools/perf/util/trace-event-info.c-548- return true;
--
tools/testing/selftests/bpf/prog_tests/task_fd_query_tp.c=4=static void test_task_fd_query_tp_core(const char *probe_name,
--
tools/testing/selftests/bpf/prog_tests/task_fd_query_tp.c-36- attr.config = strtol(buf, NULL, 0);
tools/testing/selftests/bpf/prog_tests/task_fd_query_tp.c:37: attr.type = PERF_TYPE_TRACEPOINT;
tools/testing/selftests/bpf/prog_tests/task_fd_query_tp.c-38- attr.sample_type = PERF_SAMPLE_RAW;
--
tools/testing/selftests/bpf/prog_tests/tp_attach_query.c=4=void serial_test_tp_attach_query(void)
--
tools/testing/selftests/bpf/prog_tests/tp_attach_query.c-35- attr.config = strtol(buf, NULL, 0);
tools/testing/selftests/bpf/prog_tests/tp_attach_query.c:36: attr.type = PERF_TYPE_TRACEPOINT;
tools/testing/selftests/bpf/prog_tests/tp_attach_query.c-37- attr.sample_type = PERF_SAMPLE_RAW | PERF_SAMPLE_CALLCHAIN;
--
tools/testing/selftests/user_events/perf_test.c=133=TEST_F(user, perf_write) {
--
tools/testing/selftests/user_events/perf_test.c-158-
tools/testing/selftests/user_events/perf_test.c:159: pe.type = PERF_TYPE_TRACEPOINT;
tools/testing/selftests/user_events/perf_test.c-160- pe.size = sizeof(pe);
--
tools/testing/selftests/user_events/perf_test.c=198=TEST_F(user, perf_empty_events) {
--
tools/testing/selftests/user_events/perf_test.c-220-
tools/testing/selftests/user_events/perf_test.c:221: pe.type = PERF_TYPE_TRACEPOINT;
tools/testing/selftests/user_events/perf_test.c-222- pe.size = sizeof(pe);
]
|
| 119/4 |
2026/05/25 22:17 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 19149
output: 43
thoughts: 83
Thoughts: **Analyzing the Tracepoint Event Initialization**
Okay, so I'm diving into the kernel code and I've just zeroed in on `perf_tp_event_init` within `kernel/events/core.c`. Given the naming convention and its location within the event-handling infrastructure, my initial thought is that this function is *the* initialization routine for tracepoint events. This is expected. It's the logical place to start digging when I'm trying to understand how tracepoint events are brought to life within the `perf` subsystem. Now, let's get into the details of the implementation.
|
| 120/4 |
2026/05/25 22:17 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_tp_event_init]
Results: map[SourceCode:10928: static int perf_tp_event_init(struct perf_event *event)
10929: {
10930: int err;
10931:
10932: if (event->attr.type != PERF_TYPE_TRACEPOINT)
10933: return -ENOENT;
10934:
10935: /*
10936: * no branch sampling for tracepoint events
10937: */
10938: if (has_branch_stack(event))
10939: return -EOPNOTSUPP;
10940:
10941: err = perf_trace_init(event);
10942: if (err)
10943: return err;
10944:
10945: event->destroy = tp_perf_event_destroy;
10946:
10947: return 0;
10948: }
SourceFile:kernel/events/core.c]
|
| 121/4 |
2026/05/25 22:17 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 19706
output: 19
thoughts: 55
Thoughts: Okay, here's my take on those thoughts, framed as my own:
**Initializing the Performance Trace**
Alright, so the code is making a call to `perf_trace_init(event)`. Naturally, my first instinct is to locate the definition of that function. Given my experience, it's highly probable that this resides within the kernel's tracing infrastructure, specifically something like `kernel/trace/trace_event_perf.c` or a file with a similar naming convention. However, I won't rely on assumptions. Let's dig in and confirm the exact location to be certain.
|
| 122/4 |
2026/05/25 22:17 |
tool |
grepper |
0m
Args: map[Expression:perf_trace_init]
Results: map[Output:include/linux/trace_events.h=864=DECLARE_PER_CPU(struct pt_regs, perf_trace_regs);
include/linux/trace_events.h-865-
include/linux/trace_events.h:866:extern int perf_trace_init(struct perf_event *event);
include/linux/trace_events.h-867-extern void perf_trace_destroy(struct perf_event *event);
--
kernel/events/core.c=10928=static int perf_tp_event_init(struct perf_event *event)
--
kernel/events/core.c-10940-
kernel/events/core.c:10941: err = perf_trace_init(event);
kernel/events/core.c-10942- if (err)
--
kernel/trace/trace_event_perf.c=193=static int perf_trace_event_init(struct trace_event_call *tp_event,
--
kernel/trace/trace_event_perf.c-214-
kernel/trace/trace_event_perf.c:215:int perf_trace_init(struct perf_event *p_event)
kernel/trace/trace_event_perf.c-216-{
--
kernel/trace/trace_fprobe.c=1495=static int disable_trace_fprobe(struct trace_event_call *call,
--
kernel/trace/trace_fprobe.c-1537-/*
kernel/trace/trace_fprobe.c:1538: * called by perf_trace_init() or __ftrace_set_clr_event() under event_mutex.
kernel/trace/trace_fprobe.c-1539- */
--
kernel/trace/trace_kprobe.c=1755=int bpf_get_kprobe_info(const struct perf_event *event, u32 *fd_type,
--
kernel/trace/trace_kprobe.c-1783-/*
kernel/trace/trace_kprobe.c:1784: * called by perf_trace_init() or __ftrace_set_clr_event() under event_mutex.
kernel/trace/trace_kprobe.c-1785- *
]
|
| 123/4 |
2026/05/25 22:17 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 20477
output: 45
thoughts: 29
Thoughts: Okay, here's my take on those thoughts, framed as an expert's internal monologue:
**Diving into `perf_trace_init`**
Alright, so `perf_trace_init`... right, that lives in `kernel/trace/trace_event_perf.c`. Time to get my hands dirty and dissect this. Let's dig in and see exactly what's happening under the hood. I know the general architecture, but understanding the specifics of the initialization is always critical for tracing performance. Gotta refresh my memory of how this connects with the rest of the perf infrastructure, specifically the trace events. Let's see how it sets up the interaction between the perf subsystem and the tracepoints.
|
| 124/4 |
2026/05/25 22:18 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/trace_event_perf.c IncludeLines:true Name:perf_trace_init]
Results: map[SourceCode: 215: int perf_trace_init(struct perf_event *p_event)
216: {
217: struct trace_event_call *tp_event;
218: u64 event_id = p_event->attr.config;
219: int ret = -EINVAL;
220:
221: mutex_lock(&event_mutex);
222: list_for_each_entry(tp_event, &ftrace_events, list) {
223: if (tp_event->event.type == event_id &&
224: tp_event->class && tp_event->class->reg &&
225: trace_event_try_get_ref(tp_event)) {
226: ret = perf_trace_event_init(tp_event, p_event);
227: if (ret)
228: trace_event_put_ref(tp_event);
229: break;
230: }
231: }
232: mutex_unlock(&event_mutex);
233:
234: return ret;
235: }
SourceFile:kernel/trace/trace_event_perf.c]
|
| 125/4 |
2026/05/25 22:18 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 21056
output: 47
thoughts: 26
Thoughts: Okay, here's my interpretation of that thought process, assuming I'm an expert in this area:
**Deconstructing the Event Initialization**
Alright, so the code's calling `perf_trace_event_init(tp_event, p_event)`. My first thought is to dive right into that function. That's the core of how this event is being set up, so understanding it is crucial. `tp_event` and `p_event` are the arguments. I'll need to remember to check their types and what they represent in the context of the larger tracepoint/perf event system to fully grasp what's happening. I'll be looking for how the event is configured, any memory allocation, and the overall state setup. This function is the gateway, it's where the magic begins, and I need to see what that magic looks like.
|
| 126/4 |
2026/05/25 22:18 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/trace_event_perf.c IncludeLines:true Name:perf_trace_event_init]
Results: map[SourceCode: 193: static int perf_trace_event_init(struct trace_event_call *tp_event,
194: struct perf_event *p_event)
195: {
196: int ret;
197:
198: ret = perf_trace_event_perm(tp_event, p_event);
199: if (ret)
200: return ret;
201:
202: ret = perf_trace_event_reg(tp_event, p_event);
203: if (ret)
204: return ret;
205:
206: ret = perf_trace_event_open(p_event);
207: if (ret) {
208: perf_trace_event_unreg(p_event);
209: return ret;
210: }
211:
212: return 0;
213: }
SourceFile:kernel/trace/trace_event_perf.c]
|
| 127/4 |
2026/05/25 22:18 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 21620
output: 47
thoughts: 26
Thoughts: Okay, here's my interpretation of that thought process, from an expert's perspective:
**Diving into the Permission Check**
Alright, let's see. I'm focusing on the permission aspects of this tracing event. I'm specifically looking at that function call, `perf_trace_event_perm(tp_event, p_event)`. Given its name, that `perm` suffix, and the context, it's pretty clear this is where the *actual* permission check for a trace event happens. I immediately start analyzing this part, knowing that any security implications of this system call will be handled there. Understanding what `tp_event` and `p_event` represent in this instance is crucial, so a deep dive into the inputs will be required. I need to figure out the security model at play here, and verify if it is correctly implemented, and if any corner cases are covered. I bet this is where I'll find whether a user has the necessary rights to use a particular tracing event. Interesting.
|
| 128/4 |
2026/05/25 22:18 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/trace_event_perf.c IncludeLines:true Name:perf_trace_event_perm]
Results: map[SourceCode: 27: static int perf_trace_event_perm(struct trace_event_call *tp_event,
28: struct perf_event *p_event)
29: {
30: int ret;
31:
32: if (tp_event->perf_perm) {
33: ret = tp_event->perf_perm(tp_event, p_event);
34: if (ret)
35: return ret;
36: }
37:
38: /*
39: * We checked and allowed to create parent,
40: * allow children without checking.
41: */
42: if (p_event->parent)
43: return 0;
44:
45: /*
46: * It's ok to check current process (owner) permissions in here,
47: * because code below is called only via perf_event_open syscall.
48: */
49:
50: /* The ftrace function trace is allowed only for root. */
51: if (ftrace_event_is_function(tp_event)) {
52: ret = perf_allow_tracepoint();
53: if (ret)
54: return ret;
55:
56: if (!is_sampling_event(p_event))
57: return 0;
58:
59: /*
60: * We don't allow user space callchains for function trace
61: * event, due to issues with page faults while tracing page
62: * fault handler and its overall trickiness nature.
63: */
64: if (!p_event->attr.exclude_callchain_user)
65: return -EINVAL;
66:
67: /*
68: * Same reason to disable user stack dump as for user space
69: * callchains above.
70: */
71: if (p_event->attr.sample_type & PERF_SAMPLE_STACK_USER)
72: return -EINVAL;
73: }
74:
75: /* No tracing, just counting, so no obvious leak */
76: if (!(p_event->attr.sample_type & PERF_SAMPLE_RAW))
77: return 0;
78:
79: /* Some events are ok to be traced by non-root users... */
80: if (p_event->attach_state == PERF_ATTACH_TASK) {
81: if (tp_event->flags & TRACE_EVENT_FL_CAP_ANY)
82: return 0;
83: }
84:
85: /*
86: * ...otherwise raw tracepoint data can be a severe data leak,
87: * only allow root to have these.
88: */
89: ret = perf_allow_tracepoint();
90: if (ret)
91: return ret;
92:
93: return 0;
94: }
SourceFile:kernel/trace/trace_event_perf.c]
|
| 129/4 |
2026/05/25 22:18 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 22787
output: 42
thoughts: 301
Thoughts: **Analyzing `perf_trace_event_perm` and Permissioning**
Okay, let's break down this `perf_trace_event_perm` function. My initial impression is that it seems to be managing permissions for trace events, specifically related to the `perf` subsystem. Let me go through my thought process.
First, if `tp_event->perf_perm` is set, I imagine it's some sort of pre-defined permission check, and it simply calls that. Straightforward enough.
Second, if it's a child event, then that's immediately disallowed; I'm assuming that it is not intended to track child events here.
Third, if it's a function trace event, then `perf_allow_tracepoint()` is called. This indicates that function trace events need special scrutiny; perhaps there's some nuance in what those events can access or potentially cause.
Now, here's where it gets interesting. Fourth, if `PERF_SAMPLE_RAW` is *not* set, it returns 0. Line 76: "No tracing, just counting, so no obvious leak." This really stands out to me. It seems to imply that if we're *just* counting and not actually tracing raw data, then it might be okay. We're not reading user data, so there are fewer potential leak paths. This is a crucial security consideration.
Fifth, if the event is attached to a task (`PERF_ATTACH_TASK`) *and* the event has `TRACE_EVENT_FL_CAP_ANY`, then it also returns 0. This suggests that attaching to a specific task and having this "capability" somehow requires the event to be blocked. I need to understand what `TRACE_EVENT_FL_CAP_ANY` means exactly.
And finally, in all the other cases, it calls `perf_allow_tracepoint()`.
The behavior on line 76 is the key.
This line is telling me that *not* requesting `PERF_SAMPLE_RAW` could open the door to unprivileged access. This is interesting, because then counting may be accessible to a broader set of users. I need to examine this more closely, maybe there is a permission check somewhere else?
But the final call to `perf_allow_tracepoint()` is a bit confusing to me. It is called in different contexts, I need to see what's happening there to understand this behavior better. Let's see what `perf_allow_tracepoint()` does.
|
| 130/4 |
2026/05/25 22:18 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_allow_tracepoint]
Results: map[SourceCode:1787: static inline int perf_allow_tracepoint(void)
1788: {
1789: if (sysctl_perf_event_paranoid > -1 && !perfmon_capable())
1790: return -EPERM;
1791:
1792: return security_perf_event_open(PERF_SECURITY_TRACEPOINT);
1793: }
SourceFile:include/linux/perf_event.h]
|
| 131/4 |
2026/05/25 22:18 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 23803
output: 24
thoughts: 126
Thoughts: **Navigating the Kernel Source Code**
Okay, so I'm trying to track down where `perf_allow_tracepoint` is defined, and initially, I was thrown off a bit. The tool I'm using pointed me towards `kernel/events/core.c`, but the output looked a little suspect. Hmm, I need to be more careful. Let me double-check the tool's output again. Aha! The `SourceFile` is actually listed as `include/linux/perf_event.h`. My bad, it looks like I misread the output initially, or maybe it was truncated. Right, that makes more sense. Now, let's go straight to the source. Time to open up `include/linux/perf_event.h` and confirm this is where I'll find `perf_allow_tracepoint`.
|
| 132/4 |
2026/05/25 22:18 |
tool |
read-file |
0m
Args: map[File:include/linux/perf_event.h]
Results: map[Contents:/*
* Performance events:
*
* Copyright (C) 2008-2009, Thomas Gleixner <tglx@linutronix.de>
* Copyright (C) 2008-2011, Red Hat, Inc., Ingo Molnar
* Copyright (C) 2008-2011, Red Hat, Inc., Peter Zijlstra
*
* Data type definitions, declarations, prototypes.
*
* Started by: Thomas Gleixner and Ingo Molnar
*
* For licencing details see kernel-base/COPYING
*/
#ifndef _LINUX_PERF_EVENT_H
#define _LINUX_PERF_EVENT_H
#include <uapi/linux/perf_event.h>
#include <uapi/linux/bpf_perf_event.h>
/*
* Kernel-internal data types and definitions:
*/
#ifdef CONFIG_PERF_EVENTS
# include <asm/perf_event.h>
# include <asm/local64.h>
#endif
#ifdef CONFIG_HAVE_HW_BREAKPOINT
# include <linux/rhashtable-types.h>
# include <asm/hw_breakpoint.h>
#endif
#include <linux/list.h>
#include <linux/mutex.h>
#include <linux/rculist.h>
#include <linux/rcupdate.h>
#include <linux/spinlock.h>
#include <linux/hrtimer.h>
#include <linux/fs.h>
#include <linux/pid_namespace.h>
#include <linux/workqueue.h>
#include <linux/ftrace.h>
#include <linux/cpu.h>
#include <linux/irq_work.h>
#include <linux/static_key.h>
#include <linux/jump_label_ratelimit.h>
#include <linux/atomic.h>
#include <linux/sysfs.h>
#include <linux/perf_regs.h>
#include <linux/cgroup.h>
#include <linux/refcount.h>
#include <linux/security.h>
#include <linux/static_call.h>
#include <linux/lockdep.h>
#include <asm/local.h>
struct perf_callchain_entry {
u64 nr;
u64 ip[]; /* /proc/sys/kernel/perf_event_max_stack */
};
struct perf_callchain_entry_ctx {
struct perf_callchain_entry *entry;
u32 max_stack;
u32 nr;
short contexts;
bool contexts_maxed;
};
typedef unsigned long (*perf_copy_f)(void *dst, const void *src,
unsigned long off, unsigned long len);
struct perf_raw_frag {
union {
struct perf_raw_frag *next;
unsigned long pad;
};
perf_copy_f copy;
void *data;
u32 size;
} __packed;
struct perf_raw_record {
struct perf_raw_frag frag;
u32 size;
};
static __always_inline bool perf_raw_frag_last(const struct perf_raw_frag *frag)
{
return frag->pad < sizeof(u64);
}
/*
* branch stack layout:
* nr: number of taken branches stored in entries[]
* hw_idx: The low level index of raw branch records
* for the most recent branch.
* -1ULL means invalid/unknown.
*
* Note that nr can vary from sample to sample
* branches (to, from) are stored from most recent
* to least recent, i.e., entries[0] contains the most
* recent branch.
* The entries[] is an abstraction of raw branch records,
* which may not be stored in age order in HW, e.g. Intel LBR.
* The hw_idx is to expose the low level index of raw
* branch record for the most recent branch aka entries[0].
* The hw_idx index is between -1 (unknown) and max depth,
* which can be retrieved in /sys/devices/cpu/caps/branches.
* For the architectures whose raw branch records are
* already stored in age order, the hw_idx should be 0.
*/
struct perf_branch_stack {
u64 nr;
u64 hw_idx;
struct perf_branch_entry entries[];
};
struct task_struct;
/*
* extra PMU register associated with an event
*/
struct hw_perf_event_extra {
u64 config; /* register value */
unsigned int reg; /* register address or index */
int alloc; /* extra register already allocated */
int idx; /* index in shared_regs->regs[] */
};
/**
* hw_perf_event::flag values
*
* PERF_EVENT_FLAG_ARCH bits are reserved for architecture-specific
* usage.
*/
#define PERF_EVENT_FLAG_ARCH 0x0fffffff
#define PERF_EVENT_FLAG_USER_READ_CNT 0x80000000
static_assert((PERF_EVENT_FLAG_USER_READ_CNT & PERF_EVENT_FLAG_ARCH) == 0);
/**
* struct hw_perf_event - performance event hardware details:
*/
struct hw_perf_event {
#ifdef CONFIG_PERF_EVENTS
union {
struct { /* hardware */
u64 config;
u64 config1;
u64 last_tag;
u64 dyn_constraint;
unsigned long config_base;
unsigned long event_base;
int event_base_rdpmc;
int idx;
int last_cpu;
int flags;
struct hw_perf_event_extra extra_reg;
struct hw_perf_event_extra branch_reg;
};
struct { /* aux / Intel-PT */
u64 aux_config;
/*
* For AUX area events, aux_paused cannot be a state
* flag because it can be updated asynchronously to
* state.
*/
unsigned int aux_paused;
};
struct { /* software */
struct hrtimer hrtimer;
};
struct { /* tracepoint */
/* for tp_event->class */
struct list_head tp_list;
};
struct { /* amd_power */
u64 pwr_acc;
u64 ptsc;
};
#ifdef CONFIG_HAVE_HW_BREAKPOINT
struct { /* breakpoint */
/*
* Crufty hack to avoid the chicken and egg
* problem hw_breakpoint has with context
* creation and event initalization.
*/
struct arch_hw_breakpoint info;
struct rhlist_head bp_list;
};
#endif
struct { /* amd_iommu */
u8 iommu_bank;
u8 iommu_cntr;
u16 padding;
u64 conf;
u64 conf1;
};
};
/*
* If the event is a per task event, this will point to the task in
* question. See the comment in perf_event_alloc().
*/
struct task_struct *target;
/*
* PMU would store hardware filter configuration
* here.
*/
void *addr_filters;
/* Last sync'ed generation of filters */
unsigned long addr_filters_gen;
/*
* hw_perf_event::state flags; used to track the PERF_EF_* state.
*/
/* the counter is stopped */
#define PERF_HES_STOPPED 0x01
/* event->count up-to-date */
#define PERF_HES_UPTODATE 0x02
#define PERF_HES_ARCH 0x04
int state;
/*
* The last observed hardware counter value, updated with a
* local64_cmpxchg() such that pmu::read() can be called nested.
*/
local64_t prev_count;
/*
* The period to start the next sample with.
*/
u64 sample_period;
union {
struct { /* Sampling */
/*
* The period we started this sample with.
*/
u64 last_period;
/*
* However much is left of the current period;
* note that this is a full 64bit value and
* allows for generation of periods longer
* than hardware might allow.
*/
local64_t period_left;
};
struct { /* Topdown events counting for context switch */
u64 saved_metric;
u64 saved_slots;
};
};
/*
* State for throttling the event, see __perf_event_overflow() and
* perf_adjust_freq_unthr_context().
*/
u64 interrupts_seq;
u64 interrupts;
/*
* State for freq target events, see __perf_event_overflow() and
* perf_adjust_freq_unthr_context().
*/
u64 freq_time_stamp;
u64 freq_count_stamp;
#endif /* CONFIG_PERF_EVENTS */
};
struct perf_event;
struct perf_event_pmu_context;
/*
* Common implementation detail of pmu::{start,commit,cancel}_txn
*/
/* txn to add/schedule event on PMU */
#define PERF_PMU_TXN_ADD 0x1
/* txn to read event group from PMU */
#define PERF_PMU_TXN_READ 0x2
/**
* pmu::capabilities flags
*/
#define PERF_PMU_CAP_NO_INTERRUPT 0x0001
#define PERF_PMU_CAP_NO_NMI 0x0002
#define PERF_PMU_CAP_AUX_NO_SG 0x0004
#define PERF_PMU_CAP_EXTENDED_REGS 0x0008
#define PERF_PMU_CAP_EXCLUSIVE 0x0010
#define PERF_PMU_CAP_ITRACE 0x0020
#define PERF_PMU_CAP_NO_EXCLUDE 0x0040
#define PERF_PMU_CAP_AUX_OUTPUT 0x0080
#define PERF_PMU_CAP_EXTENDED_HW_TYPE 0x0100
#define PERF_PMU_CAP_AUX_PAUSE 0x0200
#define PERF_PMU_CAP_AUX_PREFER_LARGE 0x0400
/**
* pmu::scope
*/
enum perf_pmu_scope {
PERF_PMU_SCOPE_NONE = 0,
PERF_PMU_SCOPE_CORE,
PERF_PMU_SCOPE_DIE,
PERF_PMU_SCOPE_CLUSTER,
PERF_PMU_SCOPE_PKG,
PERF_PMU_SCOPE_SYS_WIDE,
PERF_PMU_MAX_SCOPE,
};
struct perf_output_handle;
#define PMU_NULL_DEV ((void *)(~0UL))
/**
* struct pmu - generic performance monitoring unit
*/
struct pmu {
struct list_head entry;
spinlock_t events_lock;
struct list_head events;
struct module *module;
struct device *dev;
struct device *parent;
const struct attribute_group **attr_groups;
const struct attribute_group **attr_update;
const char *name;
int type;
/*
* various common per-pmu feature flags
*/
int capabilities;
/*
* PMU scope
*/
unsigned int scope;
struct perf_cpu_pmu_context * __percpu *cpu_pmu_context;
atomic_t exclusive_cnt; /* < 0: cpu; > 0: tsk */
int task_ctx_nr;
int hrtimer_interval_ms;
/* number of address filters this PMU can do */
unsigned int nr_addr_filters;
/*
* Fully disable/enable this PMU, can be used to protect from the PMI
* as well as for lazy/batch writing of the MSRs.
*/
void (*pmu_enable) (struct pmu *pmu); /* optional */
void (*pmu_disable) (struct pmu *pmu); /* optional */
/*
* Try and initialize the event for this PMU.
*
* Returns:
* -ENOENT -- @event is not for this PMU
*
* -ENODEV -- @event is for this PMU but PMU not present
* -EBUSY -- @event is for this PMU but PMU temporarily unavailable
* -EINVAL -- @event is for this PMU but @event is not valid
* -EOPNOTSUPP -- @event is for this PMU, @event is valid, but not supported
* -EACCES -- @event is for this PMU, @event is valid, but no privileges
*
* 0 -- @event is for this PMU and valid
*
* Other error return values are allowed.
*/
int (*event_init) (struct perf_event *event);
/*
* Notification that the event was mapped or unmapped. Called
* in the context of the mapping task.
*/
void (*event_mapped) (struct perf_event *event, struct mm_struct *mm); /* optional */
void (*event_unmapped) (struct perf_event *event, struct mm_struct *mm); /* optional */
/*
* Flags for ->add()/->del()/ ->start()/->stop(). There are
* matching hw_perf_event::state flags.
*/
/* start the counter when adding */
#define PERF_EF_START 0x01
/* reload the counter when starting */
#define PERF_EF_RELOAD 0x02
/* update the counter when stopping */
#define PERF_EF_UPDATE 0x04
/* AUX area event, pause tracing */
#define PERF_EF_PAUSE 0x08
/* AUX area event, resume tracing */
#define PERF_EF_RESUME 0x10
/*
* Adds/Removes a counter to/from the PMU, can be done inside a
* transaction, see the ->*_txn() methods.
*
* The add/del callbacks will reserve all hardware resources required
* to service the event, this includes any counter constraint
* scheduling etc.
*
* Called with IRQs disabled and the PMU disabled on the CPU the event
* is on.
*
* ->add() called without PERF_EF_START should result in the same state
* as ->add() followed by ->stop().
*
* ->del() must always PERF_EF_UPDATE stop an event. If it calls
* ->stop() that must deal with already being stopped without
* PERF_EF_UPDATE.
*/
int (*add) (struct perf_event *event, int flags);
void (*del) (struct perf_event *event, int flags);
/*
* Starts/Stops a counter present on the PMU.
*
* The PMI handler should stop the counter when perf_event_overflow()
* returns !0. ->start() will be used to continue.
*
* Also used to change the sample period.
*
* Called with IRQs disabled and the PMU disabled on the CPU the event
* is on -- will be called from NMI context with the PMU generates
* NMIs.
*
* ->stop() with PERF_EF_UPDATE will read the counter and update
* period/count values like ->read() would.
*
* ->start() with PERF_EF_RELOAD will reprogram the counter
* value, must be preceded by a ->stop() with PERF_EF_UPDATE.
*
* ->stop() with PERF_EF_PAUSE will stop as simply as possible. Will not
* overlap another ->stop() with PERF_EF_PAUSE nor ->start() with
* PERF_EF_RESUME.
*
* ->start() with PERF_EF_RESUME will start as simply as possible but
* only if the counter is not otherwise stopped. Will not overlap
* another ->start() with PERF_EF_RESUME nor ->stop() with
* PERF_EF_PAUSE.
*
* Notably, PERF_EF_PAUSE/PERF_EF_RESUME *can* be concurrent with other
* ->stop()/->start() invocations, just not itself.
*/
void (*start) (struct perf_event *event, int flags);
void (*stop) (struct perf_event *event, int flags);
/*
* Updates the counter value of the event.
*
* For sampling capable PMUs this will also update the software period
* hw_perf_event::period_left field.
*/
void (*read) (struct perf_event *event);
/*
* Group events scheduling is treated as a transaction, add
* group events as a whole and perform one schedulability test.
* If the test fails, roll back the whole group
*
* Start the transaction, after this ->add() doesn't need to
* do schedulability tests.
*
* Optional.
*/
void (*start_txn) (struct pmu *pmu, unsigned int txn_flags);
/*
* If ->start_txn() disabled the ->add() schedulability test
* then ->commit_txn() is required to perform one. On success
* the transaction is closed. On error the transaction is kept
* open until ->cancel_txn() is called.
*
* Optional.
*/
int (*commit_txn) (struct pmu *pmu);
/*
* Will cancel the transaction, assumes ->del() is called
* for each successful ->add() during the transaction.
*
* Optional.
*/
void (*cancel_txn) (struct pmu *pmu);
/*
* Will return the value for perf_event_mmap_page::index for this event,
* if no implementation is provided it will default to 0 (see
* perf_event_idx_default).
*/
int (*event_idx) (struct perf_event *event); /*optional */
/*
* context-switches callback
*/
void (*sched_task) (struct perf_event_pmu_context *pmu_ctx,
struct task_struct *task, bool sched_in);
/*
* Kmem cache of PMU specific data
*/
struct kmem_cache *task_ctx_cache;
/*
* Set up pmu-private data structures for an AUX area
*/
void *(*setup_aux) (struct perf_event *event, void **pages,
int nr_pages, bool overwrite);
/* optional */
/*
* Free pmu-private AUX data structures
*/
void (*free_aux) (void *aux); /* optional */
/*
* Take a snapshot of the AUX buffer without touching the event
* state, so that preempting ->start()/->stop() callbacks does
* not interfere with their logic. Called in PMI context.
*
* Returns the size of AUX data copied to the output handle.
*
* Optional.
*/
long (*snapshot_aux) (struct perf_event *event,
struct perf_output_handle *handle,
unsigned long size);
/*
* Validate address range filters: make sure the HW supports the
* requested configuration and number of filters; return 0 if the
* supplied filters are valid, -errno otherwise.
*
* Runs in the context of the ioctl()ing process and is not serialized
* with the rest of the PMU callbacks.
*/
int (*addr_filters_validate) (struct list_head *filters);
/* optional */
/*
* Synchronize address range filter configuration:
* translate hw-agnostic filters into hardware configuration in
* event::hw::addr_filters.
*
* Runs as a part of filter sync sequence that is done in ->start()
* callback by calling perf_event_addr_filters_sync().
*
* May (and should) traverse event::addr_filters::list, for which its
* caller provides necessary serialization.
*/
void (*addr_filters_sync) (struct perf_event *event);
/* optional */
/*
* Check if event can be used for aux_output purposes for
* events of this PMU.
*
* Runs from perf_event_open(). Should return 0 for "no match"
* or non-zero for "match".
*/
int (*aux_output_match) (struct perf_event *event);
/* optional */
/*
* Skip programming this PMU on the given CPU. Typically needed for
* big.LITTLE things.
*/
bool (*filter) (struct pmu *pmu, int cpu); /* optional */
/*
* Check period value for PERF_EVENT_IOC_PERIOD ioctl.
*/
int (*check_period) (struct perf_event *event, u64 value); /* optional */
};
enum perf_addr_filter_action_t {
PERF_ADDR_FILTER_ACTION_STOP = 0,
PERF_ADDR_FILTER_ACTION_START,
PERF_ADDR_FILTER_ACTION_FILTER,
};
/**
* struct perf_addr_filter - address range filter definition
* @entry: event's filter list linkage
* @path: object file's path for file-based filters
* @offset: filter range offset
* @size: filter range size (size==0 means single address trigger)
* @action: filter/start/stop
*
* This is a hardware-agnostic filter configuration as specified by the user.
*/
struct perf_addr_filter {
struct list_head entry;
struct path path;
unsigned long offset;
unsigned long size;
enum perf_addr_filter_action_t action;
};
/**
* struct perf_addr_filters_head - container for address range filters
* @list: list of filters for this event
* @lock: spinlock that serializes accesses to the @list and event's
* (and its children's) filter generations.
* @nr_file_filters: number of file-based filters
*
* A child event will use parent's @list (and therefore @lock), so they are
* bundled together; see perf_event_addr_filters().
*/
struct perf_addr_filters_head {
struct list_head list;
raw_spinlock_t lock;
unsigned int nr_file_filters;
};
struct perf_addr_filter_range {
unsigned long start;
unsigned long size;
};
/*
* The normal states are:
*
* ACTIVE --.
* ^ |
* | |
* sched_{in,out}() |
* | |
* v |
* ,---> INACTIVE --+ <-.
* | | |
* | {dis,en}able()
* sched_in() | |
* | OFF <--' --+
* | |
* `---> ERROR ------'
*
* That is:
*
* sched_in: INACTIVE -> {ACTIVE,ERROR}
* sched_out: ACTIVE -> INACTIVE
* disable: {ACTIVE,INACTIVE} -> OFF
* enable: {OFF,ERROR} -> INACTIVE
*
* Where {OFF,ERROR} are disabled states.
*
* Then we have the {EXIT,REVOKED,DEAD} states which are various shades of
* defunct events:
*
* - EXIT means task that the even was assigned to died, but child events
* still live, and further children can still be created. But the event
* itself will never be active again. It can only transition to
* {REVOKED,DEAD};
*
* - REVOKED means the PMU the event was associated with is gone; all
* functionality is stopped but the event is still alive. Can only
* transition to DEAD;
*
* - DEAD event really is DYING tearing down state and freeing bits.
*
*/
enum perf_event_state {
PERF_EVENT_STATE_DEAD = -5,
PERF_EVENT_STATE_REVOKED = -4, /* pmu gone, must not touch */
PERF_EVENT_STATE_EXIT = -3, /* task died, still inherit */
PERF_EVENT_STATE_ERROR = -2, /* scheduling error, can enable */
PERF_EVENT_STATE_OFF = -1,
PERF_EVENT_STATE_INACTIVE = 0,
PERF_EVENT_STATE_ACTIVE = 1,
};
struct file;
struct perf_sample_data;
typedef void (*perf_overflow_handler_t)(struct perf_event *,
struct perf_sample_data *,
struct pt_regs *regs);
/*
* Event capabilities. For event_caps and groups caps.
*
* PERF_EV_CAP_SOFTWARE: Is a software event.
* PERF_EV_CAP_READ_ACTIVE_PKG: A CPU event (or cgroup event) that can be read
* from any CPU in the package where it is active.
* PERF_EV_CAP_SIBLING: An event with this flag must be a group sibling and
* cannot be a group leader. If an event with this flag is detached from the
* group it is scheduled out and moved into an unrecoverable ERROR state.
* PERF_EV_CAP_READ_SCOPE: A CPU event that can be read from any CPU of the
* PMU scope where it is active.
*/
#define PERF_EV_CAP_SOFTWARE BIT(0)
#define PERF_EV_CAP_READ_ACTIVE_PKG BIT(1)
#define PERF_EV_CAP_SIBLING BIT(2)
#define PERF_EV_CAP_READ_SCOPE BIT(3)
#define SWEVENT_HLIST_BITS 8
#define SWEVENT_HLIST_SIZE (1 << SWEVENT_HLIST_BITS)
struct swevent_hlist {
struct hlist_head heads[SWEVENT_HLIST_SIZE];
struct rcu_head rcu_head;
};
#define PERF_ATTACH_CONTEXT 0x0001
#define PERF_ATTACH_GROUP 0x0002
#define PERF_ATTACH_TASK 0x0004
#define PERF_ATTACH_TASK_DATA 0x0008
#define PERF_ATTACH_GLOBAL_DATA 0x0010
#define PERF_ATTACH_SCHED_CB 0x0020
#define PERF_ATTACH_CHILD 0x0040
#define PERF_ATTACH_EXCLUSIVE 0x0080
#define PERF_ATTACH_CALLCHAIN 0x0100
#define PERF_ATTACH_ITRACE 0x0200
struct bpf_prog;
struct perf_cgroup;
struct perf_buffer;
struct pmu_event_list {
raw_spinlock_t lock;
struct list_head list;
};
/*
* event->sibling_list is modified whole holding both ctx->lock and ctx->mutex
* as such iteration must hold either lock. However, since ctx->lock is an IRQ
* safe lock, and is only held by the CPU doing the modification, having IRQs
* disabled is sufficient since it will hold-off the IPIs.
*/
#ifdef CONFIG_PROVE_LOCKING
# define lockdep_assert_event_ctx(event) \
WARN_ON_ONCE(__lockdep_enabled && \
(this_cpu_read(hardirqs_enabled) && \
lockdep_is_held(&(event)->ctx->mutex) != LOCK_STATE_HELD))
#else
# define lockdep_assert_event_ctx(event)
#endif
#define for_each_sibling_event(sibling, event) \
lockdep_assert_event_ctx(event); \
if ((event)->group_leader == (event)) \
list_for_each_entry((sibling), &(event)->sibling_list, sibling_list)
/**
* struct perf_event - performance event kernel representation:
*/
struct perf_event {
#ifdef CONFIG_PERF_EVENTS
/*
* entry onto perf_event_context::event_list;
* modifications require ctx->lock
* RCU safe iterations.
*/
struct list_head event_entry;
/*
* Locked for modification by both ctx->mutex and ctx->lock; holding
* either sufficies for read.
*/
struct list_head sibling_list;
struct list_head active_list;
/*
* Node on the pinned or flexible tree located at the event context;
*/
struct rb_node group_node;
u64 group_index;
/*
* We need storage to track the entries in perf_pmu_migrate_context; we
* cannot use the event_entry because of RCU and we want to keep the
* group in tact which avoids us using the other two entries.
*/
struct list_head migrate_entry;
struct hlist_node hlist_entry;
struct list_head active_entry;
int nr_siblings;
/* Not serialized. Only written during event initialization. */
int event_caps;
/* The cumulative AND of all event_caps for events in this group. */
int group_caps;
unsigned int group_generation;
struct perf_event *group_leader;
/*
* event->pmu will always point to pmu in which this event belongs.
* Whereas event->pmu_ctx->pmu may point to other pmu when group of
* different pmu events is created.
*/
struct pmu *pmu;
void *pmu_private;
enum perf_event_state state;
unsigned int attach_state;
local64_t count;
atomic64_t child_count;
/*
* These are the total time in nanoseconds that the event
* has been enabled (i.e. eligible to run, and the task has
* been scheduled in, if this is a per-task event)
* and running (scheduled onto the CPU), respectively.
*/
u64 total_time_enabled;
u64 total_time_running;
u64 tstamp;
struct perf_event_attr attr;
u16 header_size;
u16 id_header_size;
u16 read_size;
struct hw_perf_event hw;
struct perf_event_context *ctx;
/*
* event->pmu_ctx points to perf_event_pmu_context in which the event
* is added. This pmu_ctx can be of other pmu for sw event when that
* sw event is part of a group which also contains non-sw events.
*/
struct perf_event_pmu_context *pmu_ctx;
atomic_long_t refcount;
/*
* These accumulate total time (in nanoseconds) that children
* events have been enabled and running, respectively.
*/
atomic64_t child_total_time_enabled;
atomic64_t child_total_time_running;
/*
* Protect attach/detach and child_list:
*/
struct mutex child_mutex;
struct list_head child_list;
struct perf_event *parent;
int oncpu;
int cpu;
struct list_head owner_entry;
struct task_struct *owner;
/* mmap bits */
struct mutex mmap_mutex;
refcount_t mmap_count;
struct perf_buffer *rb;
struct list_head rb_entry;
unsigned long rcu_batches;
int rcu_pending;
/* poll related */
wait_queue_head_t waitq;
struct fasync_struct *fasync;
/* delayed work for NMIs and such */
unsigned int pending_wakeup;
unsigned int pending_kill;
unsigned int pending_disable;
unsigned long pending_addr; /* SIGTRAP */
struct irq_work pending_irq;
struct irq_work pending_disable_irq;
struct callback_head pending_task;
unsigned int pending_work;
atomic_t event_limit;
/* address range filters */
struct perf_addr_filters_head addr_filters;
/* vma address array for file-based filders */
struct perf_addr_filter_range *addr_filter_ranges;
unsigned long addr_filters_gen;
/* for aux_output events */
struct perf_event *aux_event;
void (*destroy)(struct perf_event *);
struct rcu_head rcu_head;
struct pid_namespace *ns;
u64 id;
atomic64_t lost_samples;
u64 (*clock)(void);
perf_overflow_handler_t overflow_handler;
void *overflow_handler_context;
struct bpf_prog *prog;
u64 bpf_cookie;
#ifdef CONFIG_EVENT_TRACING
struct trace_event_call *tp_event;
struct event_filter *filter;
# ifdef CONFIG_FUNCTION_TRACER
struct ftrace_ops ftrace_ops;
# endif
#endif
#ifdef CONFIG_CGROUP_PERF
struct perf_cgroup *cgrp; /* cgroup event is attach to */
#endif
#ifdef CONFIG_SECURITY
void *security;
#endif
struct list_head sb_list;
struct list_head pmu_list;
/*
* Certain events gets forwarded to another pmu internally by over-
* writing kernel copy of event->attr.type without user being aware
* of it. event->orig_type contains original 'type' requested by
* user.
*/
u32 orig_type;
#endif /* CONFIG_PERF_EVENTS */
};
/*
* ,-----------------------[1:n]------------------------.
* V V
* perf_event_context <-[1:n]-> perf_event_pmu_context <-[1:n]- perf_event
* | |
* `--[n:1]-> pmu <-[1:n]--'
*
*
* struct perf_event_pmu_context lifetime is refcount based and RCU freed
* (similar to perf_event_context). Locking is as if it were a member of
* perf_event_context; specifically:
*
* modification, both: ctx->mutex && ctx->lock
* reading, either: ctx->mutex || ctx->lock
*
* There is one exception to this; namely put_pmu_ctx() isn't always called
* with ctx->mutex held; this means that as long as we can guarantee the epc
* has events the above rules hold.
*
* Specificially, sys_perf_event_open()'s group_leader case depends on
* ctx->mutex pinning the configuration. Since we hold a reference on
* group_leader (through the filedesc) it can't go away, therefore it's
* associated pmu_ctx must exist and cannot change due to ctx->mutex.
*
* perf_event holds a refcount on perf_event_context
* perf_event holds a refcount on perf_event_pmu_context
*/
struct perf_event_pmu_context {
struct pmu *pmu;
struct perf_event_context *ctx;
struct list_head pmu_ctx_entry;
struct list_head pinned_active;
struct list_head flexible_active;
/* Used to identify the per-cpu perf_event_pmu_context */
unsigned int embedded : 1;
unsigned int nr_events;
unsigned int nr_cgroups;
unsigned int nr_freq;
atomic_t refcount; /* event <-> epc */
struct rcu_head rcu_head;
/*
* Set when one or more (plausibly active) event can't be scheduled
* due to pmu overcommit or pmu constraints, except tolerant to
* events not necessary to be active due to scheduling constraints,
* such as cgroups.
*/
int rotate_necessary;
};
static inline bool perf_pmu_ctx_is_active(struct perf_event_pmu_context *epc)
{
return !list_empty(&epc->flexible_active) || !list_empty(&epc->pinned_active);
}
struct perf_event_groups {
struct rb_root tree;
u64 index;
};
/**
* struct perf_event_context - event context structure
*
* Used as a container for task events and CPU events as well:
*/
struct perf_event_context {
/*
* Protect the states of the events in the list,
* nr_active, and the list:
*/
raw_spinlock_t lock;
/*
* Protect the list of events. Locking either mutex or lock
* is sufficient to ensure the list doesn't change; to change
* the list you need to lock both the mutex and the spinlock.
*/
struct mutex mutex;
struct list_head pmu_ctx_list;
struct perf_event_groups pinned_groups;
struct perf_event_groups flexible_groups;
struct list_head event_list;
int nr_events;
int nr_user;
int is_active;
int nr_stat;
int nr_freq;
int rotate_disable;
refcount_t refcount; /* event <-> ctx */
struct task_struct *task;
/*
* Context clock, runs when context enabled.
*/
u64 time;
u64 timestamp;
u64 timeoffset;
/*
* These fields let us detect when two contexts have both
* been cloned (inherited) from a common ancestor.
*/
struct perf_event_context *parent_ctx;
u64 parent_gen;
u64 generation;
int pin_count;
#ifdef CONFIG_CGROUP_PERF
int nr_cgroups; /* cgroup evts */
#endif
struct rcu_head rcu_head;
/*
* The count of events for which using the switch-out fast path
* should be avoided.
*
* Sum (event->pending_work + events with
* (attr->inherit && (attr->sample_type & PERF_SAMPLE_READ)))
*
* The SIGTRAP is targeted at ctx->task, as such it won't do changing
* that until the signal is delivered.
*/
local_t nr_no_switch_fast;
};
/**
* struct perf_ctx_data - PMU specific data for a task
* @rcu_head: To avoid the race on free PMU specific data
* @refcount: To track users
* @global: To track system-wide users
* @ctx_cache: Kmem cache of PMU specific data
* @data: PMU specific data
*
* Currently, the struct is only used in Intel LBR call stack mode to
* save/restore the call stack of a task on context switches.
*
* The rcu_head is used to prevent the race on free the data.
* The data only be allocated when Intel LBR call stack mode is enabled.
* The data will be freed when the mode is disabled.
* The content of the data will only be accessed in context switch, which
* should be protected by rcu_read_lock().
*
* Because of the alignment requirement of Intel Arch LBR, the Kmem cache
* is used to allocate the PMU specific data. The ctx_cache is to track
* the Kmem cache.
*
* Careful: Struct perf_ctx_data is added as a pointer in struct task_struct.
* When system-wide Intel LBR call stack mode is enabled, a buffer with
* constant size will be allocated for each task.
* Also, system memory consumption can further grow when the size of
* struct perf_ctx_data enlarges.
*/
struct perf_ctx_data {
struct rcu_head rcu_head;
refcount_t refcount;
int global;
struct kmem_cache *ctx_cache;
void *data;
};
struct perf_cpu_pmu_context {
struct perf_event_pmu_context epc;
struct perf_event_pmu_context *task_epc;
struct list_head sched_cb_entry;
int sched_cb_usage;
int active_oncpu;
int exclusive;
int pmu_disable_count;
raw_spinlock_t hrtimer_lock;
struct hrtimer hrtimer;
ktime_t hrtimer_interval;
unsigned int hrtimer_active;
};
/**
* struct perf_event_cpu_context - per cpu event context structure
*/
struct perf_cpu_context {
struct perf_event_context ctx;
struct perf_event_context *task_ctx;
int online;
#ifdef CONFIG_CGROUP_PERF
struct perf_cgroup *cgrp;
#endif
/*
* Per-CPU storage for iterators used in visit_groups_merge. The default
* storage is of size 2 to hold the CPU and any CPU event iterators.
*/
int heap_size;
struct perf_event **heap;
struct perf_event *heap_default[2];
};
struct perf_output_handle {
struct perf_event *event;
struct perf_buffer *rb;
unsigned long wakeup;
unsigned long size;
union {
u64 flags; /* perf_output*() */
u64 aux_flags; /* perf_aux_output*() */
struct {
u64 skip_read : 1;
};
};
union {
void *addr;
unsigned long head;
};
int page;
};
struct bpf_perf_event_data_kern {
bpf_user_pt_regs_t *regs;
struct perf_sample_data *data;
struct perf_event *event;
};
#ifdef CONFIG_CGROUP_PERF
/*
* perf_cgroup_info keeps track of time_enabled for a cgroup.
* This is a per-cpu dynamically allocated data structure.
*/
struct perf_cgroup_info {
u64 time;
u64 timestamp;
u64 timeoffset;
int active;
};
struct perf_cgroup {
struct cgroup_subsys_state css;
struct perf_cgroup_info __percpu *info;
};
/*
* Must ensure cgroup is pinned (css_get) before calling
* this function. In other words, we cannot call this function
* if there is no cgroup event for the current CPU context.
*/
static inline struct perf_cgroup *
perf_cgroup_from_task(struct task_struct *task, struct perf_event_context *ctx)
{
return container_of(task_css_check(task, perf_event_cgrp_id,
ctx ? lockdep_is_held(&ctx->lock)
: true),
struct perf_cgroup, css);
}
#endif /* CONFIG_CGROUP_PERF */
#ifdef CONFIG_PERF_EVENTS
extern struct perf_event_context *perf_cpu_task_ctx(void);
extern void *perf_aux_output_begin(struct perf_output_handle *handle,
struct perf_event *event);
extern void perf_aux_output_end(struct perf_output_handle *handle,
unsigned long size);
extern int perf_aux_output_skip(struct perf_output_handle *handle,
unsigned long size);
extern void *perf_get_aux(struct perf_output_handle *handle);
extern void perf_aux_output_flag(struct perf_output_handle *handle, u64 flags);
extern void perf_event_itrace_started(struct perf_event *event);
extern int perf_pmu_register(struct pmu *pmu, const char *name, int type);
extern int perf_pmu_unregister(struct pmu *pmu);
extern void __perf_event_task_sched_in(struct task_struct *prev,
struct task_struct *task);
extern void __perf_event_task_sched_out(struct task_struct *prev,
struct task_struct *next);
extern int perf_event_init_task(struct task_struct *child, u64 clone_flags);
extern void perf_event_exit_task(struct task_struct *child);
extern void perf_event_free_task(struct task_struct *task);
extern void perf_event_delayed_put(struct task_struct *task);
extern struct file *perf_event_get(unsigned int fd);
extern const struct perf_event *perf_get_event(struct file *file);
extern const struct perf_event_attr *perf_event_attrs(struct perf_event *event);
extern void perf_event_print_debug(void);
extern void perf_pmu_disable(struct pmu *pmu);
extern void perf_pmu_enable(struct pmu *pmu);
extern void perf_sched_cb_dec(struct pmu *pmu);
extern void perf_sched_cb_inc(struct pmu *pmu);
extern int perf_event_task_disable(void);
extern int perf_event_task_enable(void);
extern void perf_pmu_resched(struct pmu *pmu);
extern int perf_event_refresh(struct perf_event *event, int refresh);
extern void perf_event_update_userpage(struct perf_event *event);
extern int perf_event_release_kernel(struct perf_event *event);
extern struct perf_event *
perf_event_create_kernel_counter(struct perf_event_attr *attr,
int cpu,
struct task_struct *task,
perf_overflow_handler_t callback,
void *context);
extern void perf_pmu_migrate_context(struct pmu *pmu,
int src_cpu, int dst_cpu);
extern int perf_event_read_local(struct perf_event *event, u64 *value,
u64 *enabled, u64 *running);
extern u64 perf_event_read_value(struct perf_event *event,
u64 *enabled, u64 *running);
extern struct perf_callchain_entry *perf_callchain(struct perf_event *event, struct pt_regs *regs);
static inline bool branch_sample_no_flags(const struct perf_event *event)
{
return event->attr.branch_sample_type & PERF_SAMPLE_BRANCH_NO_FLAGS;
}
static inline bool branch_sample_no_cycles(const struct perf_event *event)
{
return event->attr.branch_sample_type & PERF_SAMPLE_BRANCH_NO_CYCLES;
}
static inline bool branch_sample_type(const struct perf_event *event)
{
return event->attr.branch_sample_type & PERF_SAMPLE_BRANCH_TYPE_SAVE;
}
static inline bool branch_sample_hw_index(const struct perf_event *event)
{
return event->attr.branch_sample_type & PERF_SAMPLE_BRANCH_HW_INDEX;
}
static inline bool branch_sample_priv(const struct perf_event *event)
{
return event->attr.branch_sample_type & PERF_SAMPLE_BRANCH_PRIV_SAVE;
}
static inline bool branch_sample_counters(const struct perf_event *event)
{
return event->attr.branch_sample_type & PERF_SAMPLE_BRANCH_COUNTERS;
}
static inline bool branch_sample_call_stack(const struct perf_event *event)
{
return event->attr.branch_sample_type & PERF_SAMPLE_BRANCH_CALL_STACK;
}
struct perf_sample_data {
/*
* Fields set by perf_sample_data_init() unconditionally,
* group so as to minimize the cachelines touched.
*/
u64 sample_flags;
u64 period;
u64 dyn_size;
/*
* Fields commonly set by __perf_event_header__init_id(),
* group so as to minimize the cachelines touched.
*/
u64 type;
struct {
u32 pid;
u32 tid;
} tid_entry;
u64 time;
u64 id;
struct {
u32 cpu;
u32 reserved;
} cpu_entry;
/*
* The other fields, optionally {set,used} by
* perf_{prepare,output}_sample().
*/
u64 ip;
struct perf_callchain_entry *callchain;
struct perf_raw_record *raw;
struct perf_branch_stack *br_stack;
u64 *br_stack_cntr;
union perf_sample_weight weight;
union perf_mem_data_src data_src;
u64 txn;
struct perf_regs regs_user;
struct perf_regs regs_intr;
u64 stack_user_size;
u64 stream_id;
u64 cgroup;
u64 addr;
u64 phys_addr;
u64 data_page_size;
u64 code_page_size;
u64 aux_size;
} ____cacheline_aligned;
/* default value for data source */
#define PERF_MEM_NA (PERF_MEM_S(OP, NA) |\
PERF_MEM_S(LVL, NA) |\
PERF_MEM_S(SNOOP, NA) |\
PERF_MEM_S(LOCK, NA) |\
PERF_MEM_S(TLB, NA) |\
PERF_MEM_S(LVLNUM, NA))
static inline void perf_sample_data_init(struct perf_sample_data *data,
u64 addr, u64 period)
{
/* remaining struct members initialized in perf_prepare_sample() */
data->sample_flags = PERF_SAMPLE_PERIOD;
data->period = period;
data->dyn_size = 0;
if (addr) {
data->addr = addr;
data->sample_flags |= PERF_SAMPLE_ADDR;
}
}
static inline void perf_sample_save_callchain(struct perf_sample_data *data,
struct perf_event *event,
struct pt_regs *regs)
{
int size = 1;
if (!(event->attr.sample_type & PERF_SAMPLE_CALLCHAIN))
return;
if (WARN_ON_ONCE(data->sample_flags & PERF_SAMPLE_CALLCHAIN))
return;
data->callchain = perf_callchain(event, regs);
size += data->callchain->nr;
data->dyn_size += size * sizeof(u64);
data->sample_flags |= PERF_SAMPLE_CALLCHAIN;
}
static inline void perf_sample_save_raw_data(struct perf_sample_data *data,
struct perf_event *event,
struct perf_raw_record *raw)
{
struct perf_raw_frag *frag = &raw->frag;
u32 sum = 0;
int size;
if (!(event->attr.sample_type & PERF_SAMPLE_RAW))
return;
if (WARN_ON_ONCE(data->sample_flags & PERF_SAMPLE_RAW))
return;
do {
sum += frag->size;
if (perf_raw_frag_last(frag))
break;
frag = frag->next;
} while (1);
size = round_up(sum + sizeof(u32), sizeof(u64));
raw->size = size - sizeof(u32);
frag->pad = raw->size - sum;
data->raw = raw;
data->dyn_size += size;
data->sample_flags |= PERF_SAMPLE_RAW;
}
static inline bool has_branch_stack(struct perf_event *event)
{
return event->attr.sample_type & PERF_SAMPLE_BRANCH_STACK;
}
static inline void perf_sample_save_brstack(struct perf_sample_data *data,
struct perf_event *event,
struct perf_branch_stack *brs,
u64 *brs_cntr)
{
int size = sizeof(u64); /* nr */
if (!has_branch_stack(event))
return;
if (WARN_ON_ONCE(data->sample_flags & PERF_SAMPLE_BRANCH_STACK))
return;
if (branch_sample_hw_index(event))
size += sizeof(u64);
brs->nr = min_t(u16, event->attr.sample_max_stack, brs->nr);
size += brs->nr * sizeof(struct perf_branch_entry);
/*
* The extension space for counters is appended after the
* struct perf_branch_stack. It is used to store the occurrences
* of events of each branch.
*/
if (brs_cntr)
size += brs->nr * sizeof(u64);
data->br_stack = brs;
data->br_stack_cntr = brs_cntr;
data->dyn_size += size;
data->sample_flags |= PERF_SAMPLE_BRANCH_STACK;
}
static inline u32 perf_sample_data_size(struct perf_sample_data *data,
struct perf_event *event)
{
u32 size = sizeof(struct perf_event_header);
size += event->header_size + event->id_header_size;
size += data->dyn_size;
return size;
}
/*
* Clear all bitfields in the perf_branch_entry.
* The to and from fields are not cleared because they are
* systematically modified by caller.
*/
static inline void perf_clear_branch_entry_bitfields(struct perf_branch_entry *br)
{
br->mispred = 0;
br->predicted = 0;
br->in_tx = 0;
br->abort = 0;
br->cycles = 0;
br->type = 0;
br->spec = PERF_BR_SPEC_NA;
br->reserved = 0;
}
extern void perf_output_sample(struct perf_output_handle *handle,
struct perf_event_header *header,
struct perf_sample_data *data,
struct perf_event *event);
extern void perf_prepare_sample(struct perf_sample_data *data,
struct perf_event *event,
struct pt_regs *regs);
extern void perf_prepare_header(struct perf_event_header *header,
struct perf_sample_data *data,
struct perf_event *event,
struct pt_regs *regs);
extern int perf_event_overflow(struct perf_event *event,
struct perf_sample_data *data,
struct pt_regs *regs);
extern void perf_event_output_forward(struct perf_event *event,
struct perf_sample_data *data,
struct pt_regs *regs);
extern void perf_event_output_backward(struct perf_event *event,
struct perf_sample_data *data,
struct pt_regs *regs);
extern int perf_event_output(struct perf_event *event,
struct perf_sample_data *data,
struct pt_regs *regs);
static inline bool
is_default_overflow_handler(struct perf_event *event)
{
perf_overflow_handler_t overflow_handler = event->overflow_handler;
if (likely(overflow_handler == perf_event_output_forward))
return true;
if (unlikely(overflow_handler == perf_event_output_backward))
return true;
return false;
}
extern void
perf_event_header__init_id(struct perf_event_header *header,
struct perf_sample_data *data,
struct perf_event *event);
extern void
perf_event__output_id_sample(struct perf_event *event,
struct perf_output_handle *handle,
struct perf_sample_data *sample);
extern void
perf_log_lost_samples(struct perf_event *event, u64 lost);
static inline bool event_has_any_exclude_flag(struct perf_event *event)
{
struct perf_event_attr *attr = &event->attr;
return attr->exclude_idle || attr->exclude_user ||
attr->exclude_kernel || attr->exclude_hv ||
attr->exclude_guest || attr->exclude_host;
}
static inline bool is_sampling_event(struct perf_event *event)
{
return event->attr.sample_period != 0;
}
/*
* Return 1 for a software event, 0 for a hardware event
*/
static inline int is_software_event(struct perf_event *event)
{
return event->event_caps & PERF_EV_CAP_SOFTWARE;
}
/*
* Return 1 for event in sw context, 0 for event in hw context
*/
static inline int in_software_context(struct perf_event *event)
{
return event->pmu_ctx->pmu->task_ctx_nr == perf_sw_context;
}
static inline int is_exclusive_pmu(struct pmu *pmu)
{
return pmu->capabilities & PERF_PMU_CAP_EXCLUSIVE;
}
extern struct static_key perf_swevent_enabled[PERF_COUNT_SW_MAX];
extern void ___perf_sw_event(u32, u64, struct pt_regs *, u64);
extern void __perf_sw_event(u32, u64, struct pt_regs *, u64);
#ifndef perf_arch_fetch_caller_regs
static inline void perf_arch_fetch_caller_regs(struct pt_regs *regs, unsigned long ip) { }
#endif
/*
* When generating a perf sample in-line, instead of from an interrupt /
* exception, we lack a pt_regs. This is typically used from software events
* like: SW_CONTEXT_SWITCHES, SW_MIGRATIONS and the tie-in with tracepoints.
*
* We typically don't need a full set, but (for x86) do require:
* - ip for PERF_SAMPLE_IP
* - cs for user_mode() tests
* - sp for PERF_SAMPLE_CALLCHAIN
* - eflags for MISC bits and CALLCHAIN (see: perf_hw_regs())
*
* NOTE: assumes @regs is otherwise already 0 filled; this is important for
* things like PERF_SAMPLE_REGS_INTR.
*/
static inline void perf_fetch_caller_regs(struct pt_regs *regs)
{
perf_arch_fetch_caller_regs(regs, CALLER_ADDR0);
}
static __always_inline void
perf_sw_event(u32 event_id, u64 nr, struct pt_regs *regs, u64 addr)
{
if (static_key_false(&perf_swevent_enabled[event_id]))
__perf_sw_event(event_id, nr, regs, addr);
}
DECLARE_PER_CPU(struct pt_regs, __perf_regs[4]);
/*
* 'Special' version for the scheduler, it hard assumes no recursion,
* which is guaranteed by us not actually scheduling inside other swevents
* because those disable preemption.
*/
static __always_inline void __perf_sw_event_sched(u32 event_id, u64 nr, u64 addr)
{
struct pt_regs *regs = this_cpu_ptr(&__perf_regs[0]);
perf_fetch_caller_regs(regs);
___perf_sw_event(event_id, nr, regs, addr);
}
extern struct static_key_false perf_sched_events;
static __always_inline bool __perf_sw_enabled(int swevt)
{
return static_key_false(&perf_swevent_enabled[swevt]);
}
static inline void perf_event_task_migrate(struct task_struct *task)
{
if (__perf_sw_enabled(PERF_COUNT_SW_CPU_MIGRATIONS))
task->sched_migrated = 1;
}
static inline void perf_event_task_sched_in(struct task_struct *prev,
struct task_struct *task)
{
if (static_branch_unlikely(&perf_sched_events))
__perf_event_task_sched_in(prev, task);
if (__perf_sw_enabled(PERF_COUNT_SW_CPU_MIGRATIONS) &&
task->sched_migrated) {
__perf_sw_event_sched(PERF_COUNT_SW_CPU_MIGRATIONS, 1, 0);
task->sched_migrated = 0;
}
}
static inline void perf_event_task_sched_out(struct task_struct *prev,
struct task_struct *next)
{
if (__perf_sw_enabled(PERF_COUNT_SW_CONTEXT_SWITCHES))
__perf_sw_event_sched(PERF_COUNT_SW_CONTEXT_SWITCHES, 1, 0);
#ifdef CONFIG_CGROUP_PERF
if (__perf_sw_enabled(PERF_COUNT_SW_CGROUP_SWITCHES) &&
perf_cgroup_from_task(prev, NULL) !=
perf_cgroup_from_task(next, NULL))
__perf_sw_event_sched(PERF_COUNT_SW_CGROUP_SWITCHES, 1, 0);
#endif
if (static_branch_unlikely(&perf_sched_events))
__perf_event_task_sched_out(prev, next);
}
extern void perf_event_mmap(struct vm_area_struct *vma);
extern void perf_event_ksymbol(u16 ksym_type, u64 addr, u32 len,
bool unregister, const char *sym);
extern void perf_event_bpf_event(struct bpf_prog *prog,
enum perf_bpf_event_type type,
u16 flags);
#define PERF_GUEST_ACTIVE 0x01
#define PERF_GUEST_USER 0x02
struct perf_guest_info_callbacks {
unsigned int (*state)(void);
unsigned long (*get_ip)(void);
unsigned int (*handle_intel_pt_intr)(void);
};
#ifdef CONFIG_GUEST_PERF_EVENTS
extern struct perf_guest_info_callbacks __rcu *perf_guest_cbs;
DECLARE_STATIC_CALL(__perf_guest_state, *perf_guest_cbs->state);
DECLARE_STATIC_CALL(__perf_guest_get_ip, *perf_guest_cbs->get_ip);
DECLARE_STATIC_CALL(__perf_guest_handle_intel_pt_intr, *perf_guest_cbs->handle_intel_pt_intr);
static inline unsigned int perf_guest_state(void)
{
return static_call(__perf_guest_state)();
}
static inline unsigned long perf_guest_get_ip(void)
{
return static_call(__perf_guest_get_ip)();
}
static inline unsigned int perf_guest_handle_intel_pt_intr(void)
{
return static_call(__perf_guest_handle_intel_pt_intr)();
}
extern void perf_register_guest_info_callbacks(struct perf_guest_info_callbacks *cbs);
extern void perf_unregister_guest_info_callbacks(struct perf_guest_info_callbacks *cbs);
#else /* !CONFIG_GUEST_PERF_EVENTS: */
static inline unsigned int perf_guest_state(void) { return 0; }
static inline unsigned long perf_guest_get_ip(void) { return 0; }
static inline unsigned int perf_guest_handle_intel_pt_intr(void) { return 0; }
#endif /* !CONFIG_GUEST_PERF_EVENTS */
extern void perf_event_exec(void);
extern void perf_event_comm(struct task_struct *tsk, bool exec);
extern void perf_event_namespaces(struct task_struct *tsk);
extern void perf_event_fork(struct task_struct *tsk);
extern void perf_event_text_poke(const void *addr,
const void *old_bytes, size_t old_len,
const void *new_bytes, size_t new_len);
/* Callchains */
DECLARE_PER_CPU(struct perf_callchain_entry, perf_callchain_entry);
extern void perf_callchain_user(struct perf_callchain_entry_ctx *entry, struct pt_regs *regs);
extern void perf_callchain_kernel(struct perf_callchain_entry_ctx *entry, struct pt_regs *regs);
extern struct perf_callchain_entry *
get_perf_callchain(struct pt_regs *regs, bool kernel, bool user,
u32 max_stack, bool crosstask, bool add_mark, u64 defer_cookie);
extern int get_callchain_buffers(int max_stack);
extern void put_callchain_buffers(void);
extern struct perf_callchain_entry *get_callchain_entry(int *rctx);
extern void put_callchain_entry(int rctx);
extern int sysctl_perf_event_max_stack;
extern int sysctl_perf_event_max_contexts_per_stack;
static inline int perf_callchain_store_context(struct perf_callchain_entry_ctx *ctx, u64 ip)
{
if (ctx->contexts < sysctl_perf_event_max_contexts_per_stack) {
struct perf_callchain_entry *entry = ctx->entry;
entry->ip[entry->nr++] = ip;
++ctx->contexts;
return 0;
} else {
ctx->contexts_maxed = true;
return -1; /* no more room, stop walking the stack */
}
}
static inline int perf_callchain_store(struct perf_callchain_entry_ctx *ctx, u64 ip)
{
if (ctx->nr < ctx->max_stack && !ctx->contexts_maxed) {
struct perf_callchain_entry *entry = ctx->entry;
entry->ip[entry->nr++] = ip;
++ctx->nr;
return 0;
} else {
return -1; /* no more room, stop walking the stack */
}
}
extern int sysctl_perf_event_paranoid;
extern int sysctl_perf_event_sample_rate;
extern void perf_sample_event_took(u64 sample_len_ns);
/* Access to perf_event_open(2) syscall. */
#define PERF_SECURITY_OPEN 0
/* Finer grained perf_event_open(2) access control. */
#define PERF_SECURITY_CPU 1
#define PERF_SECURITY_KERNEL 2
#define PERF_SECURITY_TRACEPOINT 3
static inline int perf_is_paranoid(void)
{
return sysctl_perf_event_paranoid > -1;
}
extern int perf_allow_kernel(void);
static inline int perf_allow_cpu(void)
{
if (sysctl_perf_event_paranoid > 0 && !perfmon_capable())
return -EACCES;
return security_perf_event_open(PERF_SECURITY_CPU);
}
static inline int perf_allow_tracepoint(void)
{
if (sysctl_perf_event_paranoid > -1 && !perfmon_capable())
return -EPERM;
return security_perf_event_open(PERF_SECURITY_TRACEPOINT);
}
extern int perf_exclude_event(struct perf_event *event, struct pt_regs *regs);
extern void perf_event_init(void);
extern void perf_tp_event(u16 event_type, u64 count, void *record,
int entry_size, struct pt_regs *regs,
struct hlist_head *head, int rctx,
struct task_struct *task);
extern void perf_bp_event(struct perf_event *event, void *data);
extern unsigned long perf_misc_flags(struct perf_event *event, struct pt_regs *regs);
extern unsigned long perf_instruction_pointer(struct perf_event *event,
struct pt_regs *regs);
#ifndef perf_arch_misc_flags
# define perf_arch_misc_flags(regs) \
(user_mode(regs) ? PERF_RECORD_MISC_USER : PERF_RECORD_MISC_KERNEL)
# define perf_arch_instruction_pointer(regs) instruction_pointer(regs)
#endif
#ifndef perf_arch_bpf_user_pt_regs
# define perf_arch_bpf_user_pt_regs(regs) regs
#endif
#ifndef perf_arch_guest_misc_flags
static inline unsigned long perf_arch_guest_misc_flags(struct pt_regs *regs)
{
unsigned long guest_state = perf_guest_state();
if (!(guest_state & PERF_GUEST_ACTIVE))
return 0;
if (guest_state & PERF_GUEST_USER)
return PERF_RECORD_MISC_GUEST_USER;
else
return PERF_RECORD_MISC_GUEST_KERNEL;
}
# define perf_arch_guest_misc_flags(regs) perf_arch_guest_misc_flags(regs)
#endif
static inline bool needs_branch_stack(struct perf_event *event)
{
return event->attr.branch_sample_type != 0;
}
static inline bool has_aux(struct perf_event *event)
{
return event->pmu && event->pmu->setup_aux;
}
static inline bool has_aux_action(struct perf_event *event)
{
return event->attr.aux_sample_size ||
event->attr.aux_pause ||
event->attr.aux_resume;
}
static inline bool is_write_backward(struct perf_event *event)
{
return !!event->attr.write_backward;
}
static inline bool has_addr_filter(struct perf_event *event)
{
return event->pmu->nr_addr_filters;
}
/*
* An inherited event uses parent's filters
*/
static inline struct perf_addr_filters_head *
perf_event_addr_filters(struct perf_event *event)
{
struct perf_addr_filters_head *ifh = &event->addr_filters;
if (event->parent)
ifh = &event->parent->addr_filters;
return ifh;
}
static inline struct fasync_struct **perf_event_fasync(struct perf_event *event)
{
/* Only the parent has fasync state */
if (event->parent)
event = event->parent;
return &event->fasync;
}
extern void perf_event_addr_filters_sync(struct perf_event *event);
extern void perf_report_aux_output_id(struct perf_event *event, u64 hw_id);
extern int perf_output_begin(struct perf_output_handle *handle,
struct perf_sample_data *data,
struct perf_event *event, unsigned int size);
extern int perf_output_begin_forward(struct perf_output_handle *handle,
struct perf_sample_data *data,
struct perf_event *event,
unsigned int size);
extern int perf_output_begin_backward(struct perf_output_handle *handle,
struct perf_sample_data *data,
struct perf_event *event,
unsigned int size);
extern void perf_output_end(struct perf_output_handle *handle);
extern unsigned int perf_output_copy(struct perf_output_handle *handle,
const void *buf, unsigned int len);
extern unsigned int perf_output_skip(struct perf_output_handle *handle,
unsigned int len);
extern long perf_output_copy_aux(struct perf_output_handle *aux_handle,
struct perf_output_handle *handle,
unsigned long from, unsigned long to);
extern int perf_swevent_get_recursion_context(void);
extern void perf_swevent_put_recursion_context(int rctx);
extern u64 perf_swevent_set_period(struct perf_event *event);
extern void perf_event_enable(struct perf_event *event);
extern void perf_event_disable(struct perf_event *event);
extern void perf_event_disable_local(struct perf_event *event);
extern void perf_event_disable_inatomic(struct perf_event *event);
extern void perf_event_task_tick(void);
extern int perf_event_account_interrupt(struct perf_event *event);
extern int perf_event_period(struct perf_event *event, u64 value);
extern u64 perf_event_pause(struct perf_event *event, bool reset);
#else /* !CONFIG_PERF_EVENTS: */
static inline void *
perf_aux_output_begin(struct perf_output_handle *handle,
struct perf_event *event) { return NULL; }
static inline void
perf_aux_output_end(struct perf_output_handle *handle, unsigned long size)
{ }
static inline int
perf_aux_output_skip(struct perf_output_handle *handle,
unsigned long size) { return -EINVAL; }
static inline void *
perf_get_aux(struct perf_output_handle *handle) { return NULL; }
static inline void
perf_event_task_migrate(struct task_struct *task) { }
static inline void
perf_event_task_sched_in(struct task_struct *prev,
struct task_struct *task) { }
static inline void
perf_event_task_sched_out(struct task_struct *prev,
struct task_struct *next) { }
static inline int perf_event_init_task(struct task_struct *child,
u64 clone_flags) { return 0; }
static inline void perf_event_exit_task(struct task_struct *child) { }
static inline void perf_event_free_task(struct task_struct *task) { }
static inline void perf_event_delayed_put(struct task_struct *task) { }
static inline struct file *perf_event_get(unsigned int fd) { return ERR_PTR(-EINVAL); }
static inline const struct perf_event *perf_get_event(struct file *file)
{
return ERR_PTR(-EINVAL);
}
static inline const struct perf_event_attr *perf_event_attrs(struct perf_event *event)
{
return ERR_PTR(-EINVAL);
}
static inline int perf_event_read_local(struct perf_event *event, u64 *value,
u64 *enabled, u64 *running)
{
return -EINVAL;
}
static inline void perf_event_print_debug(void) { }
static inline int perf_event_task_disable(void) { return -EINVAL; }
static inline int perf_event_task_enable(void) { return -EINVAL; }
static inline int perf_event_refresh(struct perf_event *event, int refresh)
{
return -EINVAL;
}
static inline void
perf_sw_event(u32 event_id, u64 nr, struct pt_regs *regs, u64 addr) { }
static inline void
perf_bp_event(struct perf_event *event, void *data) { }
static inline void perf_event_mmap(struct vm_area_struct *vma) { }
typedef int (perf_ksymbol_get_name_f)(char *name, int name_len, void *data);
static inline void perf_event_ksymbol(u16 ksym_type, u64 addr, u32 len,
bool unregister, const char *sym) { }
static inline void perf_event_bpf_event(struct bpf_prog *prog,
enum perf_bpf_event_type type,
u16 flags) { }
static inline void perf_event_exec(void) { }
static inline void perf_event_comm(struct task_struct *tsk, bool exec) { }
static inline void perf_event_namespaces(struct task_struct *tsk) { }
static inline void perf_event_fork(struct task_struct *tsk) { }
static inline void perf_event_text_poke(const void *addr,
const void *old_bytes,
size_t old_len,
const void *new_bytes,
size_t new_len) { }
static inline void perf_event_init(void) { }
static inline int perf_swevent_get_recursion_context(void) { return -1; }
static inline void perf_swevent_put_recursion_context(int rctx) { }
static inline u64 perf_swevent_set_period(struct perf_event *event) { return 0; }
static inline void perf_event_enable(struct perf_event *event) { }
static inline void perf_event_disable(struct perf_event *event) { }
static inline int __perf_event_disable(void *info) { return -1; }
static inline void perf_event_task_tick(void) { }
static inline int perf_event_release_kernel(struct perf_event *event) { return 0; }
static inline int
perf_event_period(struct perf_event *event, u64 value) { return -EINVAL; }
static inline u64
perf_event_pause(struct perf_event *event, bool reset) { return 0; }
static inline int
perf_exclude_event(struct perf_event *event, struct pt_regs *regs) { return 0; }
#endif /* !CONFIG_PERF_EVENTS */
#if defined(CONFIG_PERF_EVENTS) && defined(CONFIG_CPU_SUP_INTEL)
extern void perf_restore_debug_store(void);
#else
static inline void perf_restore_debug_store(void) { }
#endif
#define perf_output_put(handle, x) perf_output_copy((handle), &(x), sizeof(x))
struct perf_pmu_events_attr {
struct device_attribute attr;
u64 id;
const char *event_str;
};
struct perf_pmu_events_ht_attr {
struct device_attribute attr;
u64 id;
const char *event_str_ht;
const char *event_str_noht;
};
struct perf_pmu_events_hybrid_attr {
struct device_attribute attr;
u64 id;
const char *event_str;
u64 pmu_type;
};
struct perf_pmu_format_hybrid_attr {
struct device_attribute attr;
u64 pmu_type;
};
ssize_t perf_event_sysfs_show(struct device *dev, struct device_attribute *attr,
char *page);
#define PMU_EVENT_ATTR(_name, _var, _id, _show) \
static struct perf_pmu_events_attr _var = { \
.attr = __ATTR(_name, 0444, _show, NULL), \
.id = _id, \
};
#define PMU_EVENT_ATTR_STRING(_name, _var, _str) \
static struct perf_pmu_events_attr _var = { \
.attr = __ATTR(_name, 0444, perf_event_sysfs_show, NULL), \
.id = 0, \
.event_str = _str, \
};
#define PMU_EVENT_ATTR_ID(_name, _show, _id) \
(&((struct perf_pmu_events_attr[]) { \
{ .attr = __ATTR(_name, 0444, _show, NULL), \
.id = _id, } \
})[0].attr.attr)
#define PMU_FORMAT_ATTR_SHOW(_name, _format) \
static ssize_t \
_name##_show(struct device *dev, \
struct device_attribute *attr, \
char *page) \
{ \
BUILD_BUG_ON(sizeof(_format) >= PAGE_SIZE); \
return sprintf(page, _format "\n"); \
} \
#define PMU_FORMAT_ATTR(_name, _format) \
PMU_FORMAT_ATTR_SHOW(_name, _format) \
\
static struct device_attribute format_attr_##_name = __ATTR_RO(_name)
/* Performance counter hotplug functions */
#ifdef CONFIG_PERF_EVENTS
extern int perf_event_init_cpu(unsigned int cpu);
extern int perf_event_exit_cpu(unsigned int cpu);
#else
# define perf_event_init_cpu NULL
# define perf_event_exit_cpu NULL
#endif
extern void arch_perf_update_userpage(struct perf_event *event,
struct perf_event_mmap_page *userpg,
u64 now);
/*
* Snapshot branch stack on software events.
*
* Branch stack can be very useful in understanding software events. For
* example, when a long function, e.g. sys_perf_event_open, returns an
* errno, it is not obvious why the function failed. Branch stack could
* provide very helpful information in this type of scenarios.
*
* On software event, it is necessary to stop the hardware branch recorder
* fast. Otherwise, the hardware register/buffer will be flushed with
* entries of the triggering event. Therefore, static call is used to
* stop the hardware recorder.
*/
/*
* cnt is the number of entries allocated for entries.
* Return number of entries copied to .
*/
typedef int (perf_snapshot_branch_stack_t)(struct perf_branch_entry *entries,
unsigned int cnt);
DECLARE_STATIC_CALL(perf_snapshot_branch_stack, perf_snapshot_branch_stack_t);
#ifndef PERF_NEEDS_LOPWR_CB
static inline void perf_lopwr_cb(bool mode)
{
}
#endif
#endif /* _LINUX_PERF_EVENT_H */
]
|
| 133/4 |
2026/05/25 22:18 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 44795
output: 41
thoughts: 635
Thoughts: **My Analysis of `perf_event_open` for Tracepoints**
Okay, so I've been digging into the `perf_event_open` system call, specifically regarding tracepoint events. My initial focus was on understanding the security implications for unprivileged users, a common concern in this area. I started by looking at `include/linux/perf_event.h` and the `perf_allow_tracepoint` function. It seems this function is the gatekeeper.
The code in `perf_allow_tracepoint` checks a few things. First, it looks at `sysctl_perf_event_paranoid`. If this sysctl is greater than -1 *and* the user doesn't have `perfmon_capable()`, it returns an error, specifically `-EPERM`. So, if `sysctl_perf_event_paranoid` is set to something like 2, this part of the check is active. Then there's `perfmon_capable()`, which, if not available, also results in the `-EPERM`. This sets up the basic, high-level control.
Next, I looked at the more granular check in `kernel/trace/trace_event_perf.c`. There's a critical piece of logic in `perf_trace_event_perm` which determines the initial access. If `PERF_SAMPLE_RAW` is *not* set in the event's attributes, the function simply returns 0, granting access. I find that interesting. It seems that if you *don't* ask for the raw sample data, it's generally allowed, at least from this snippet. However, if `PERF_SAMPLE_RAW` *is* requested, then it goes on to call `perf_allow_tracepoint()`, which then checks `sysctl_perf_event_paranoid` and the user's capabilities.
Now, my initial thought was that unprivileged users *could* open tracepoint events, but only under very specific conditions. It seemed that if the unprivileged user didn't request `PERF_SAMPLE_RAW`, access would be granted. If they *did* request raw data, then it boiled down to the value of `sysctl_perf_event_paranoid` or having appropriate capabilities like `CAP_PERFMON` or `CAP_SYS_ADMIN` depending on kernel version.
But that's not the complete picture. The code has another layer. There's also the `p_event->attach_state == PERF_ATTACH_TASK` check in `perf_trace_event_perm`. If it's a per-task event AND the tracepoint is flagged with `TRACE_EVENT_FL_CAP_ANY`, the check returns 0, *regardless* of the other constraints. This could be a significant exception. It essentially means that some tracepoints are considered "safe" for per-task use.
So, to summarize my current thinking: the answer to the main question about unprivileged access is *yes*, but with these important caveats. It is allowed under certain conditions. The `PERF_SAMPLE_RAW` setting, the `sysctl_perf_event_paranoid` value, and the existence of `TRACE_EVENT_FL_CAP_ANY` on per-task tracepoints all determine the ultimate permission.
I'm not going to stop there though, I need to check the main function `SYSCALL_DEFINE5(perf_event_open, ...)` in `kernel/events/core.c` to see if there are any other conditions. I need a full picture.
|
| 134/4 |
2026/05/25 22:18 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_event_open]
Results: map[SourceCode: SourceFile:]
|
| 135/4 |
2026/05/25 22:18 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 46264
output: 22
thoughts: 31
Thoughts: Okay, here's my interpretation of that thought process, delivered in the first person, as requested, targeting someone in a technical field:
**Debugging the Event Core**
Alright, so I'm diving into this issue with the event core. My immediate thought is to reach for a trusty tool: `read-file`. Specifically, I need to pinpoint what's going on around line 13456 of `kernel/events/core.c`. That area must be where the core logic related to the bug I'm troubleshooting lies, so I need a quick look. I'm going to start with the relevant section of code, with its line numbers, and then I'll start analyzing what's going on.
|
| 136/4 |
2026/05/25 22:18 |
tool |
read-file |
0m
Args: map[File:kernel/events/core.c]
Results: map[Contents:// SPDX-License-Identifier: GPL-2.0
/*
* Performance events core code:
*
* Copyright (C) 2008 Thomas Gleixner <tglx@linutronix.de>
* Copyright (C) 2008-2011 Red Hat, Inc., Ingo Molnar
* Copyright (C) 2008-2011 Red Hat, Inc., Peter Zijlstra
* Copyright Β© 2009 Paul Mackerras, IBM Corp. <paulus@au1.ibm.com>
*/
#include <linux/fs.h>
#include <linux/mm.h>
#include <linux/cpu.h>
#include <linux/smp.h>
#include <linux/idr.h>
#include <linux/file.h>
#include <linux/poll.h>
#include <linux/slab.h>
#include <linux/hash.h>
#include <linux/tick.h>
#include <linux/sysfs.h>
#include <linux/dcache.h>
#include <linux/percpu.h>
#include <linux/ptrace.h>
#include <linux/reboot.h>
#include <linux/vmstat.h>
#include <linux/device.h>
#include <linux/export.h>
#include <linux/vmalloc.h>
#include <linux/hardirq.h>
#include <linux/hugetlb.h>
#include <linux/rculist.h>
#include <linux/uaccess.h>
#include <linux/syscalls.h>
#include <linux/anon_inodes.h>
#include <linux/kernel_stat.h>
#include <linux/cgroup.h>
#include <linux/perf_event.h>
#include <linux/trace_events.h>
#include <linux/hw_breakpoint.h>
#include <linux/mm_types.h>
#include <linux/module.h>
#include <linux/mman.h>
#include <linux/compat.h>
#include <linux/bpf.h>
#include <linux/filter.h>
#include <linux/namei.h>
#include <linux/parser.h>
#include <linux/sched/clock.h>
#include <linux/sched/mm.h>
#include <linux/proc_ns.h>
#include <linux/mount.h>
#include <linux/min_heap.h>
#include <linux/highmem.h>
#include <linux/pgtable.h>
#include <linux/buildid.h>
#include <linux/task_work.h>
#include <linux/percpu-rwsem.h>
#include <linux/unwind_deferred.h>
#include "internal.h"
#include <asm/irq_regs.h>
typedef int (*remote_function_f)(void *);
struct remote_function_call {
struct task_struct *p;
remote_function_f func;
void *info;
int ret;
};
static void remote_function(void *data)
{
struct remote_function_call *tfc = data;
struct task_struct *p = tfc->p;
if (p) {
/* -EAGAIN */
if (task_cpu(p) != smp_processor_id())
return;
/*
* Now that we're on right CPU with IRQs disabled, we can test
* if we hit the right task without races.
*/
tfc->ret = -ESRCH; /* No such (running) process */
if (p != current)
return;
}
tfc->ret = tfc->func(tfc->info);
}
/**
* task_function_call - call a function on the cpu on which a task runs
* @p: the task to evaluate
* @func: the function to be called
* @info: the function call argument
*
* Calls the function @func when the task is currently running. This might
* be on the current CPU, which just calls the function directly. This will
* retry due to any failures in smp_call_function_single(), such as if the
* task_cpu() goes offline concurrently.
*
* returns @func return value or -ESRCH or -ENXIO when the process isn't running
*/
static int
task_function_call(struct task_struct *p, remote_function_f func, void *info)
{
struct remote_function_call data = {
.p = p,
.func = func,
.info = info,
.ret = -EAGAIN,
};
int ret;
for (;;) {
ret = smp_call_function_single(task_cpu(p), remote_function,
&data, 1);
if (!ret)
ret = data.ret;
if (ret != -EAGAIN)
break;
cond_resched();
}
return ret;
}
/**
* cpu_function_call - call a function on the cpu
* @cpu: target cpu to queue this function
* @func: the function to be called
* @info: the function call argument
*
* Calls the function @func on the remote cpu.
*
* returns: @func return value or -ENXIO when the cpu is offline
*/
static int cpu_function_call(int cpu, remote_function_f func, void *info)
{
struct remote_function_call data = {
.p = NULL,
.func = func,
.info = info,
.ret = -ENXIO, /* No such CPU */
};
smp_call_function_single(cpu, remote_function, &data, 1);
return data.ret;
}
enum event_type_t {
EVENT_FLEXIBLE = 0x01,
EVENT_PINNED = 0x02,
EVENT_TIME = 0x04,
EVENT_FROZEN = 0x08,
/* see ctx_resched() for details */
EVENT_CPU = 0x10,
EVENT_CGROUP = 0x20,
/* compound helpers */
EVENT_ALL = EVENT_FLEXIBLE | EVENT_PINNED,
EVENT_TIME_FROZEN = EVENT_TIME | EVENT_FROZEN,
};
static inline void __perf_ctx_lock(struct perf_event_context *ctx)
{
raw_spin_lock(&ctx->lock);
WARN_ON_ONCE(ctx->is_active & EVENT_FROZEN);
}
static void perf_ctx_lock(struct perf_cpu_context *cpuctx,
struct perf_event_context *ctx)
{
__perf_ctx_lock(&cpuctx->ctx);
if (ctx)
__perf_ctx_lock(ctx);
}
static inline void __perf_ctx_unlock(struct perf_event_context *ctx)
{
/*
* If ctx_sched_in() didn't again set any ALL flags, clean up
* after ctx_sched_out() by clearing is_active.
*/
if (ctx->is_active & EVENT_FROZEN) {
if (!(ctx->is_active & EVENT_ALL))
ctx->is_active = 0;
else
ctx->is_active &= ~EVENT_FROZEN;
}
raw_spin_unlock(&ctx->lock);
}
static void perf_ctx_unlock(struct perf_cpu_context *cpuctx,
struct perf_event_context *ctx)
{
if (ctx)
__perf_ctx_unlock(ctx);
__perf_ctx_unlock(&cpuctx->ctx);
}
typedef struct {
struct perf_cpu_context *cpuctx;
struct perf_event_context *ctx;
} class_perf_ctx_lock_t;
static inline void class_perf_ctx_lock_destructor(class_perf_ctx_lock_t *_T)
{ perf_ctx_unlock(_T->cpuctx, _T->ctx); }
static inline class_perf_ctx_lock_t
class_perf_ctx_lock_constructor(struct perf_cpu_context *cpuctx,
struct perf_event_context *ctx)
{ perf_ctx_lock(cpuctx, ctx); return (class_perf_ctx_lock_t){ cpuctx, ctx }; }
#define TASK_TOMBSTONE ((void *)-1L)
static bool is_kernel_event(struct perf_event *event)
{
return READ_ONCE(event->owner) == TASK_TOMBSTONE;
}
static DEFINE_PER_CPU(struct perf_cpu_context, perf_cpu_context);
struct perf_event_context *perf_cpu_task_ctx(void)
{
lockdep_assert_irqs_disabled();
return this_cpu_ptr(&perf_cpu_context)->task_ctx;
}
/*
* On task ctx scheduling...
*
* When !ctx->nr_events a task context will not be scheduled. This means
* we can disable the scheduler hooks (for performance) without leaving
* pending task ctx state.
*
* This however results in two special cases:
*
* - removing the last event from a task ctx; this is relatively straight
* forward and is done in __perf_remove_from_context.
*
* - adding the first event to a task ctx; this is tricky because we cannot
* rely on ctx->is_active and therefore cannot use event_function_call().
* See perf_install_in_context().
*
* If ctx->nr_events, then ctx->is_active and cpuctx->task_ctx are set.
*/
typedef void (*event_f)(struct perf_event *, struct perf_cpu_context *,
struct perf_event_context *, void *);
struct event_function_struct {
struct perf_event *event;
event_f func;
void *data;
};
static int event_function(void *info)
{
struct event_function_struct *efs = info;
struct perf_event *event = efs->event;
struct perf_event_context *ctx = event->ctx;
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
struct perf_event_context *task_ctx = cpuctx->task_ctx;
int ret = 0;
lockdep_assert_irqs_disabled();
perf_ctx_lock(cpuctx, task_ctx);
/*
* Since we do the IPI call without holding ctx->lock things can have
* changed, double check we hit the task we set out to hit.
*/
if (ctx->task) {
if (ctx->task != current) {
ret = -ESRCH;
goto unlock;
}
/*
* We only use event_function_call() on established contexts,
* and event_function() is only ever called when active (or
* rather, we'll have bailed in task_function_call() or the
* above ctx->task != current test), therefore we must have
* ctx->is_active here.
*/
WARN_ON_ONCE(!ctx->is_active);
/*
* And since we have ctx->is_active, cpuctx->task_ctx must
* match.
*/
WARN_ON_ONCE(task_ctx != ctx);
} else {
WARN_ON_ONCE(&cpuctx->ctx != ctx);
}
efs->func(event, cpuctx, ctx, efs->data);
unlock:
perf_ctx_unlock(cpuctx, task_ctx);
return ret;
}
static void event_function_call(struct perf_event *event, event_f func, void *data)
{
struct perf_event_context *ctx = event->ctx;
struct task_struct *task = READ_ONCE(ctx->task); /* verified in event_function */
struct perf_cpu_context *cpuctx;
struct event_function_struct efs = {
.event = event,
.func = func,
.data = data,
};
if (!event->parent) {
/*
* If this is a !child event, we must hold ctx::mutex to
* stabilize the event->ctx relation. See
* perf_event_ctx_lock().
*/
lockdep_assert_held(&ctx->mutex);
}
if (!task) {
cpu_function_call(event->cpu, event_function, &efs);
return;
}
if (task == TASK_TOMBSTONE)
return;
again:
if (!task_function_call(task, event_function, &efs))
return;
local_irq_disable();
cpuctx = this_cpu_ptr(&perf_cpu_context);
perf_ctx_lock(cpuctx, ctx);
/*
* Reload the task pointer, it might have been changed by
* a concurrent perf_event_context_sched_out().
*/
task = ctx->task;
if (task == TASK_TOMBSTONE)
goto unlock;
if (ctx->is_active) {
perf_ctx_unlock(cpuctx, ctx);
local_irq_enable();
goto again;
}
func(event, NULL, ctx, data);
unlock:
perf_ctx_unlock(cpuctx, ctx);
local_irq_enable();
}
/*
* Similar to event_function_call() + event_function(), but hard assumes IRQs
* are already disabled and we're on the right CPU.
*/
static void event_function_local(struct perf_event *event, event_f func, void *data)
{
struct perf_event_context *ctx = event->ctx;
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
struct task_struct *task = READ_ONCE(ctx->task);
struct perf_event_context *task_ctx = NULL;
lockdep_assert_irqs_disabled();
if (task) {
if (task == TASK_TOMBSTONE)
return;
task_ctx = ctx;
}
perf_ctx_lock(cpuctx, task_ctx);
task = ctx->task;
if (task == TASK_TOMBSTONE)
goto unlock;
if (task) {
/*
* We must be either inactive or active and the right task,
* otherwise we're screwed, since we cannot IPI to somewhere
* else.
*/
if (ctx->is_active) {
if (WARN_ON_ONCE(task != current))
goto unlock;
if (WARN_ON_ONCE(cpuctx->task_ctx != ctx))
goto unlock;
}
} else {
WARN_ON_ONCE(&cpuctx->ctx != ctx);
}
func(event, cpuctx, ctx, data);
unlock:
perf_ctx_unlock(cpuctx, task_ctx);
}
#define PERF_FLAG_ALL (PERF_FLAG_FD_NO_GROUP |\
PERF_FLAG_FD_OUTPUT |\
PERF_FLAG_PID_CGROUP |\
PERF_FLAG_FD_CLOEXEC)
/*
* branch priv levels that need permission checks
*/
#define PERF_SAMPLE_BRANCH_PERM_PLM \
(PERF_SAMPLE_BRANCH_KERNEL |\
PERF_SAMPLE_BRANCH_HV)
/*
* perf_sched_events : >0 events exist
*/
static void perf_sched_delayed(struct work_struct *work);
DEFINE_STATIC_KEY_FALSE(perf_sched_events);
static DECLARE_DELAYED_WORK(perf_sched_work, perf_sched_delayed);
static DEFINE_MUTEX(perf_sched_mutex);
static atomic_t perf_sched_count;
static DEFINE_PER_CPU(struct pmu_event_list, pmu_sb_events);
static atomic_t nr_mmap_events __read_mostly;
static atomic_t nr_comm_events __read_mostly;
static atomic_t nr_namespaces_events __read_mostly;
static atomic_t nr_task_events __read_mostly;
static atomic_t nr_freq_events __read_mostly;
static atomic_t nr_switch_events __read_mostly;
static atomic_t nr_ksymbol_events __read_mostly;
static atomic_t nr_bpf_events __read_mostly;
static atomic_t nr_cgroup_events __read_mostly;
static atomic_t nr_text_poke_events __read_mostly;
static atomic_t nr_build_id_events __read_mostly;
static LIST_HEAD(pmus);
static DEFINE_MUTEX(pmus_lock);
static struct srcu_struct pmus_srcu;
static cpumask_var_t perf_online_mask;
static cpumask_var_t perf_online_core_mask;
static cpumask_var_t perf_online_die_mask;
static cpumask_var_t perf_online_cluster_mask;
static cpumask_var_t perf_online_pkg_mask;
static cpumask_var_t perf_online_sys_mask;
static struct kmem_cache *perf_event_cache;
/*
* perf event paranoia level:
* -1 - not paranoid at all
* 0 - disallow raw tracepoint access for unpriv
* 1 - disallow cpu events for unpriv
* 2 - disallow kernel profiling for unpriv
*/
int sysctl_perf_event_paranoid __read_mostly = 2;
/* Minimum for 512 kiB + 1 user control page. 'free' kiB per user. */
static int sysctl_perf_event_mlock __read_mostly = 512 + (PAGE_SIZE / 1024);
/*
* max perf event sample rate
*/
#define DEFAULT_MAX_SAMPLE_RATE 100000
#define DEFAULT_SAMPLE_PERIOD_NS (NSEC_PER_SEC / DEFAULT_MAX_SAMPLE_RATE)
#define DEFAULT_CPU_TIME_MAX_PERCENT 25
int sysctl_perf_event_sample_rate __read_mostly = DEFAULT_MAX_SAMPLE_RATE;
static int sysctl_perf_cpu_time_max_percent __read_mostly = DEFAULT_CPU_TIME_MAX_PERCENT;
static int max_samples_per_tick __read_mostly = DIV_ROUND_UP(DEFAULT_MAX_SAMPLE_RATE, HZ);
static int perf_sample_period_ns __read_mostly = DEFAULT_SAMPLE_PERIOD_NS;
static int perf_sample_allowed_ns __read_mostly =
DEFAULT_SAMPLE_PERIOD_NS * DEFAULT_CPU_TIME_MAX_PERCENT / 100;
static void update_perf_cpu_limits(void)
{
u64 tmp = perf_sample_period_ns;
tmp *= sysctl_perf_cpu_time_max_percent;
tmp = div_u64(tmp, 100);
if (!tmp)
tmp = 1;
WRITE_ONCE(perf_sample_allowed_ns, tmp);
}
static bool perf_rotate_context(struct perf_cpu_pmu_context *cpc);
static int perf_event_max_sample_rate_handler(const struct ctl_table *table, int write,
void *buffer, size_t *lenp, loff_t *ppos)
{
int ret;
int perf_cpu = sysctl_perf_cpu_time_max_percent;
/*
* If throttling is disabled don't allow the write:
*/
if (write && (perf_cpu == 100 || perf_cpu == 0))
return -EINVAL;
ret = proc_dointvec_minmax(table, write, buffer, lenp, ppos);
if (ret || !write)
return ret;
max_samples_per_tick = DIV_ROUND_UP(sysctl_perf_event_sample_rate, HZ);
perf_sample_period_ns = NSEC_PER_SEC / sysctl_perf_event_sample_rate;
update_perf_cpu_limits();
return 0;
}
static int perf_cpu_time_max_percent_handler(const struct ctl_table *table, int write,
void *buffer, size_t *lenp, loff_t *ppos)
{
int ret = proc_dointvec_minmax(table, write, buffer, lenp, ppos);
if (ret || !write)
return ret;
if (sysctl_perf_cpu_time_max_percent == 100 ||
sysctl_perf_cpu_time_max_percent == 0) {
printk(KERN_WARNING
"perf: Dynamic interrupt throttling disabled, can hang your system!\n");
WRITE_ONCE(perf_sample_allowed_ns, 0);
} else {
update_perf_cpu_limits();
}
return 0;
}
static const struct ctl_table events_core_sysctl_table[] = {
/*
* User-space relies on this file as a feature check for
* perf_events being enabled. It's an ABI, do not remove!
*/
{
.procname = "perf_event_paranoid",
.data = &sysctl_perf_event_paranoid,
.maxlen = sizeof(sysctl_perf_event_paranoid),
.mode = 0644,
.proc_handler = proc_dointvec,
},
{
.procname = "perf_event_mlock_kb",
.data = &sysctl_perf_event_mlock,
.maxlen = sizeof(sysctl_perf_event_mlock),
.mode = 0644,
.proc_handler = proc_dointvec,
},
{
.procname = "perf_event_max_sample_rate",
.data = &sysctl_perf_event_sample_rate,
.maxlen = sizeof(sysctl_perf_event_sample_rate),
.mode = 0644,
.proc_handler = perf_event_max_sample_rate_handler,
.extra1 = SYSCTL_ONE,
},
{
.procname = "perf_cpu_time_max_percent",
.data = &sysctl_perf_cpu_time_max_percent,
.maxlen = sizeof(sysctl_perf_cpu_time_max_percent),
.mode = 0644,
.proc_handler = perf_cpu_time_max_percent_handler,
.extra1 = SYSCTL_ZERO,
.extra2 = SYSCTL_ONE_HUNDRED,
},
};
static int __init init_events_core_sysctls(void)
{
register_sysctl_init("kernel", events_core_sysctl_table);
return 0;
}
core_initcall(init_events_core_sysctls);
/*
* perf samples are done in some very critical code paths (NMIs).
* If they take too much CPU time, the system can lock up and not
* get any real work done. This will drop the sample rate when
* we detect that events are taking too long.
*/
#define NR_ACCUMULATED_SAMPLES 128
static DEFINE_PER_CPU(u64, running_sample_length);
static u64 __report_avg;
static u64 __report_allowed;
static void perf_duration_warn(struct irq_work *w)
{
printk_ratelimited(KERN_INFO
"perf: interrupt took too long (%lld > %lld), lowering "
"kernel.perf_event_max_sample_rate to %d\n",
__report_avg, __report_allowed,
sysctl_perf_event_sample_rate);
}
static DEFINE_IRQ_WORK(perf_duration_work, perf_duration_warn);
void perf_sample_event_took(u64 sample_len_ns)
{
u64 max_len = READ_ONCE(perf_sample_allowed_ns);
u64 running_len;
u64 avg_len;
u32 max;
if (max_len == 0)
return;
/* Decay the counter by 1 average sample. */
running_len = __this_cpu_read(running_sample_length);
running_len -= running_len/NR_ACCUMULATED_SAMPLES;
running_len += sample_len_ns;
__this_cpu_write(running_sample_length, running_len);
/*
* Note: this will be biased artificially low until we have
* seen NR_ACCUMULATED_SAMPLES. Doing it this way keeps us
* from having to maintain a count.
*/
avg_len = running_len/NR_ACCUMULATED_SAMPLES;
if (avg_len <= max_len)
return;
__report_avg = avg_len;
__report_allowed = max_len;
/*
* Compute a throttle threshold 25% below the current duration.
*/
avg_len += avg_len / 4;
max = (TICK_NSEC / 100) * sysctl_perf_cpu_time_max_percent;
if (avg_len < max)
max /= (u32)avg_len;
else
max = 1;
WRITE_ONCE(perf_sample_allowed_ns, avg_len);
WRITE_ONCE(max_samples_per_tick, max);
sysctl_perf_event_sample_rate = max * HZ;
perf_sample_period_ns = NSEC_PER_SEC / sysctl_perf_event_sample_rate;
if (!irq_work_queue(&perf_duration_work)) {
early_printk("perf: interrupt took too long (%lld > %lld), lowering "
"kernel.perf_event_max_sample_rate to %d\n",
__report_avg, __report_allowed,
sysctl_perf_event_sample_rate);
}
}
static atomic64_t perf_event_id;
static void update_context_time(struct perf_event_context *ctx);
static u64 perf_event_time(struct perf_event *event);
void __weak perf_event_print_debug(void) { }
static inline u64 perf_clock(void)
{
return local_clock();
}
static inline u64 perf_event_clock(struct perf_event *event)
{
return event->clock();
}
/*
* State based event timekeeping...
*
* The basic idea is to use event->state to determine which (if any) time
* fields to increment with the current delta. This means we only need to
* update timestamps when we change state or when they are explicitly requested
* (read).
*
* Event groups make things a little more complicated, but not terribly so. The
* rules for a group are that if the group leader is OFF the entire group is
* OFF, irrespective of what the group member states are. This results in
* __perf_effective_state().
*
* A further ramification is that when a group leader flips between OFF and
* !OFF, we need to update all group member times.
*
*
* NOTE: perf_event_time() is based on the (cgroup) context time, and thus we
* need to make sure the relevant context time is updated before we try and
* update our timestamps.
*/
static __always_inline enum perf_event_state
__perf_effective_state(struct perf_event *event)
{
struct perf_event *leader = event->group_leader;
if (leader->state <= PERF_EVENT_STATE_OFF)
return leader->state;
return event->state;
}
static __always_inline void
__perf_update_times(struct perf_event *event, u64 now, u64 *enabled, u64 *running)
{
enum perf_event_state state = __perf_effective_state(event);
u64 delta = now - event->tstamp;
*enabled = event->total_time_enabled;
if (state >= PERF_EVENT_STATE_INACTIVE)
*enabled += delta;
*running = event->total_time_running;
if (state >= PERF_EVENT_STATE_ACTIVE)
*running += delta;
}
static void perf_event_update_time(struct perf_event *event)
{
u64 now = perf_event_time(event);
__perf_update_times(event, now, &event->total_time_enabled,
&event->total_time_running);
event->tstamp = now;
}
static void perf_event_update_sibling_time(struct perf_event *leader)
{
struct perf_event *sibling;
for_each_sibling_event(sibling, leader)
perf_event_update_time(sibling);
}
static void
perf_event_set_state(struct perf_event *event, enum perf_event_state state)
{
if (event->state == state)
return;
perf_event_update_time(event);
/*
* If a group leader gets enabled/disabled all its siblings
* are affected too.
*/
if ((event->state < 0) ^ (state < 0))
perf_event_update_sibling_time(event);
WRITE_ONCE(event->state, state);
}
/*
* UP store-release, load-acquire
*/
#define __store_release(ptr, val) \
do { \
barrier(); \
WRITE_ONCE(*(ptr), (val)); \
} while (0)
#define __load_acquire(ptr) \
({ \
__unqual_scalar_typeof(*(ptr)) ___p = READ_ONCE(*(ptr)); \
barrier(); \
___p; \
})
#define for_each_epc(_epc, _ctx, _pmu, _cgroup) \
list_for_each_entry(_epc, &((_ctx)->pmu_ctx_list), pmu_ctx_entry) \
if (_cgroup && !_epc->nr_cgroups) \
continue; \
else if (_pmu && _epc->pmu != _pmu) \
continue; \
else
static void perf_ctx_disable(struct perf_event_context *ctx, bool cgroup)
{
struct perf_event_pmu_context *pmu_ctx;
for_each_epc(pmu_ctx, ctx, NULL, cgroup)
perf_pmu_disable(pmu_ctx->pmu);
}
static void perf_ctx_enable(struct perf_event_context *ctx, bool cgroup)
{
struct perf_event_pmu_context *pmu_ctx;
for_each_epc(pmu_ctx, ctx, NULL, cgroup)
perf_pmu_enable(pmu_ctx->pmu);
}
static void ctx_sched_out(struct perf_event_context *ctx, struct pmu *pmu, enum event_type_t event_type);
static void ctx_sched_in(struct perf_event_context *ctx, struct pmu *pmu, enum event_type_t event_type);
#ifdef CONFIG_CGROUP_PERF
static inline bool
perf_cgroup_match(struct perf_event *event)
{
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
/* @event doesn't care about cgroup */
if (!event->cgrp)
return true;
/* wants specific cgroup scope but @cpuctx isn't associated with any */
if (!cpuctx->cgrp)
return false;
/*
* Cgroup scoping is recursive. An event enabled for a cgroup is
* also enabled for all its descendant cgroups. If @cpuctx's
* cgroup is a descendant of @event's (the test covers identity
* case), it's a match.
*/
return cgroup_is_descendant(cpuctx->cgrp->css.cgroup,
event->cgrp->css.cgroup);
}
static inline void perf_detach_cgroup(struct perf_event *event)
{
css_put(&event->cgrp->css);
event->cgrp = NULL;
}
static inline int is_cgroup_event(struct perf_event *event)
{
return event->cgrp != NULL;
}
static inline u64 perf_cgroup_event_time(struct perf_event *event)
{
struct perf_cgroup_info *t;
t = per_cpu_ptr(event->cgrp->info, event->cpu);
return t->time;
}
static inline u64 perf_cgroup_event_time_now(struct perf_event *event, u64 now)
{
struct perf_cgroup_info *t;
t = per_cpu_ptr(event->cgrp->info, event->cpu);
if (!__load_acquire(&t->active))
return t->time;
now += READ_ONCE(t->timeoffset);
return now;
}
static inline void __update_cgrp_time(struct perf_cgroup_info *info, u64 now, bool adv)
{
if (adv)
info->time += now - info->timestamp;
info->timestamp = now;
/*
* see update_context_time()
*/
WRITE_ONCE(info->timeoffset, info->time - info->timestamp);
}
static inline void update_cgrp_time_from_cpuctx(struct perf_cpu_context *cpuctx, bool final)
{
struct perf_cgroup *cgrp = cpuctx->cgrp;
struct cgroup_subsys_state *css;
struct perf_cgroup_info *info;
if (cgrp) {
u64 now = perf_clock();
for (css = &cgrp->css; css; css = css->parent) {
cgrp = container_of(css, struct perf_cgroup, css);
info = this_cpu_ptr(cgrp->info);
__update_cgrp_time(info, now, true);
if (final)
__store_release(&info->active, 0);
}
}
}
static inline void update_cgrp_time_from_event(struct perf_event *event)
{
struct perf_cgroup_info *info;
/*
* ensure we access cgroup data only when needed and
* when we know the cgroup is pinned (css_get)
*/
if (!is_cgroup_event(event))
return;
info = this_cpu_ptr(event->cgrp->info);
/*
* Do not update time when cgroup is not active
*/
if (info->active)
__update_cgrp_time(info, perf_clock(), true);
}
static inline void
perf_cgroup_set_timestamp(struct perf_cpu_context *cpuctx)
{
struct perf_event_context *ctx = &cpuctx->ctx;
struct perf_cgroup *cgrp = cpuctx->cgrp;
struct perf_cgroup_info *info;
struct cgroup_subsys_state *css;
/*
* ctx->lock held by caller
* ensure we do not access cgroup data
* unless we have the cgroup pinned (css_get)
*/
if (!cgrp)
return;
WARN_ON_ONCE(!ctx->nr_cgroups);
for (css = &cgrp->css; css; css = css->parent) {
cgrp = container_of(css, struct perf_cgroup, css);
info = this_cpu_ptr(cgrp->info);
__update_cgrp_time(info, ctx->timestamp, false);
__store_release(&info->active, 1);
}
}
/*
* reschedule events based on the cgroup constraint of task.
*/
static void perf_cgroup_switch(struct task_struct *task)
{
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
struct perf_cgroup *cgrp;
/*
* cpuctx->cgrp is set when the first cgroup event enabled,
* and is cleared when the last cgroup event disabled.
*/
if (READ_ONCE(cpuctx->cgrp) == NULL)
return;
cgrp = perf_cgroup_from_task(task, NULL);
if (READ_ONCE(cpuctx->cgrp) == cgrp)
return;
guard(perf_ctx_lock)(cpuctx, cpuctx->task_ctx);
/*
* Re-check, could've raced vs perf_remove_from_context().
*/
if (READ_ONCE(cpuctx->cgrp) == NULL)
return;
WARN_ON_ONCE(cpuctx->ctx.nr_cgroups == 0);
perf_ctx_disable(&cpuctx->ctx, true);
ctx_sched_out(&cpuctx->ctx, NULL, EVENT_ALL|EVENT_CGROUP);
/*
* must not be done before ctxswout due
* to update_cgrp_time_from_cpuctx() in
* ctx_sched_out()
*/
cpuctx->cgrp = cgrp;
/*
* set cgrp before ctxsw in to allow
* perf_cgroup_set_timestamp() in ctx_sched_in()
* to not have to pass task around
*/
ctx_sched_in(&cpuctx->ctx, NULL, EVENT_ALL|EVENT_CGROUP);
perf_ctx_enable(&cpuctx->ctx, true);
}
static int perf_cgroup_ensure_storage(struct perf_event *event,
struct cgroup_subsys_state *css)
{
struct perf_cpu_context *cpuctx;
struct perf_event **storage;
int cpu, heap_size, ret = 0;
/*
* Allow storage to have sufficient space for an iterator for each
* possibly nested cgroup plus an iterator for events with no cgroup.
*/
for (heap_size = 1; css; css = css->parent)
heap_size++;
for_each_possible_cpu(cpu) {
cpuctx = per_cpu_ptr(&perf_cpu_context, cpu);
if (heap_size <= cpuctx->heap_size)
continue;
storage = kmalloc_node(heap_size * sizeof(struct perf_event *),
GFP_KERNEL, cpu_to_node(cpu));
if (!storage) {
ret = -ENOMEM;
break;
}
raw_spin_lock_irq(&cpuctx->ctx.lock);
if (cpuctx->heap_size < heap_size) {
swap(cpuctx->heap, storage);
if (storage == cpuctx->heap_default)
storage = NULL;
cpuctx->heap_size = heap_size;
}
raw_spin_unlock_irq(&cpuctx->ctx.lock);
kfree(storage);
}
return ret;
}
static inline int perf_cgroup_connect(int fd, struct perf_event *event,
struct perf_event_attr *attr,
struct perf_event *group_leader)
{
struct perf_cgroup *cgrp;
struct cgroup_subsys_state *css;
CLASS(fd, f)(fd);
int ret = 0;
if (fd_empty(f))
return -EBADF;
css = css_tryget_online_from_dir(fd_file(f)->f_path.dentry,
&perf_event_cgrp_subsys);
if (IS_ERR(css))
return PTR_ERR(css);
ret = perf_cgroup_ensure_storage(event, css);
if (ret)
return ret;
cgrp = container_of(css, struct perf_cgroup, css);
event->cgrp = cgrp;
/*
* all events in a group must monitor
* the same cgroup because a task belongs
* to only one perf cgroup at a time
*/
if (group_leader && group_leader->cgrp != cgrp) {
perf_detach_cgroup(event);
ret = -EINVAL;
}
return ret;
}
static inline void
perf_cgroup_event_enable(struct perf_event *event, struct perf_event_context *ctx)
{
struct perf_cpu_context *cpuctx;
if (!is_cgroup_event(event))
return;
event->pmu_ctx->nr_cgroups++;
/*
* Because cgroup events are always per-cpu events,
* @ctx == &cpuctx->ctx.
*/
cpuctx = container_of(ctx, struct perf_cpu_context, ctx);
if (ctx->nr_cgroups++)
return;
cpuctx->cgrp = perf_cgroup_from_task(current, ctx);
}
static inline void
perf_cgroup_event_disable(struct perf_event *event, struct perf_event_context *ctx)
{
struct perf_cpu_context *cpuctx;
if (!is_cgroup_event(event))
return;
event->pmu_ctx->nr_cgroups--;
/*
* Because cgroup events are always per-cpu events,
* @ctx == &cpuctx->ctx.
*/
cpuctx = container_of(ctx, struct perf_cpu_context, ctx);
if (--ctx->nr_cgroups)
return;
cpuctx->cgrp = NULL;
}
#else /* !CONFIG_CGROUP_PERF */
static inline bool
perf_cgroup_match(struct perf_event *event)
{
return true;
}
static inline void perf_detach_cgroup(struct perf_event *event)
{}
static inline int is_cgroup_event(struct perf_event *event)
{
return 0;
}
static inline void update_cgrp_time_from_event(struct perf_event *event)
{
}
static inline void update_cgrp_time_from_cpuctx(struct perf_cpu_context *cpuctx,
bool final)
{
}
static inline int perf_cgroup_connect(pid_t pid, struct perf_event *event,
struct perf_event_attr *attr,
struct perf_event *group_leader)
{
return -EINVAL;
}
static inline void
perf_cgroup_set_timestamp(struct perf_cpu_context *cpuctx)
{
}
static inline u64 perf_cgroup_event_time(struct perf_event *event)
{
return 0;
}
static inline u64 perf_cgroup_event_time_now(struct perf_event *event, u64 now)
{
return 0;
}
static inline void
perf_cgroup_event_enable(struct perf_event *event, struct perf_event_context *ctx)
{
}
static inline void
perf_cgroup_event_disable(struct perf_event *event, struct perf_event_context *ctx)
{
}
static void perf_cgroup_switch(struct task_struct *task)
{
}
#endif
/*
* set default to be dependent on timer tick just
* like original code
*/
#define PERF_CPU_HRTIMER (1000 / HZ)
/*
* function must be called with interrupts disabled
*/
static enum hrtimer_restart perf_mux_hrtimer_handler(struct hrtimer *hr)
{
struct perf_cpu_pmu_context *cpc;
bool rotations;
lockdep_assert_irqs_disabled();
cpc = container_of(hr, struct perf_cpu_pmu_context, hrtimer);
rotations = perf_rotate_context(cpc);
raw_spin_lock(&cpc->hrtimer_lock);
if (rotations)
hrtimer_forward_now(hr, cpc->hrtimer_interval);
else
cpc->hrtimer_active = 0;
raw_spin_unlock(&cpc->hrtimer_lock);
return rotations ? HRTIMER_RESTART : HRTIMER_NORESTART;
}
static void __perf_mux_hrtimer_init(struct perf_cpu_pmu_context *cpc, int cpu)
{
struct hrtimer *timer = &cpc->hrtimer;
struct pmu *pmu = cpc->epc.pmu;
u64 interval;
/*
* check default is sane, if not set then force to
* default interval (1/tick)
*/
interval = pmu->hrtimer_interval_ms;
if (interval < 1)
interval = pmu->hrtimer_interval_ms = PERF_CPU_HRTIMER;
cpc->hrtimer_interval = ns_to_ktime(NSEC_PER_MSEC * interval);
raw_spin_lock_init(&cpc->hrtimer_lock);
hrtimer_setup(timer, perf_mux_hrtimer_handler, CLOCK_MONOTONIC,
HRTIMER_MODE_ABS_PINNED_HARD);
}
static int perf_mux_hrtimer_restart(struct perf_cpu_pmu_context *cpc)
{
struct hrtimer *timer = &cpc->hrtimer;
unsigned long flags;
raw_spin_lock_irqsave(&cpc->hrtimer_lock, flags);
if (!cpc->hrtimer_active) {
cpc->hrtimer_active = 1;
hrtimer_forward_now(timer, cpc->hrtimer_interval);
hrtimer_start_expires(timer, HRTIMER_MODE_ABS_PINNED_HARD);
}
raw_spin_unlock_irqrestore(&cpc->hrtimer_lock, flags);
return 0;
}
static int perf_mux_hrtimer_restart_ipi(void *arg)
{
return perf_mux_hrtimer_restart(arg);
}
static __always_inline struct perf_cpu_pmu_context *this_cpc(struct pmu *pmu)
{
return *this_cpu_ptr(pmu->cpu_pmu_context);
}
void perf_pmu_disable(struct pmu *pmu)
{
int *count = &this_cpc(pmu)->pmu_disable_count;
if (!(*count)++)
pmu->pmu_disable(pmu);
}
void perf_pmu_enable(struct pmu *pmu)
{
int *count = &this_cpc(pmu)->pmu_disable_count;
if (!--(*count))
pmu->pmu_enable(pmu);
}
static void perf_assert_pmu_disabled(struct pmu *pmu)
{
int *count = &this_cpc(pmu)->pmu_disable_count;
WARN_ON_ONCE(*count == 0);
}
static inline void perf_pmu_read(struct perf_event *event)
{
if (event->state == PERF_EVENT_STATE_ACTIVE)
event->pmu->read(event);
}
static void get_ctx(struct perf_event_context *ctx)
{
refcount_inc(&ctx->refcount);
}
static void free_ctx(struct rcu_head *head)
{
struct perf_event_context *ctx;
ctx = container_of(head, struct perf_event_context, rcu_head);
kfree(ctx);
}
static void put_ctx(struct perf_event_context *ctx)
{
if (refcount_dec_and_test(&ctx->refcount)) {
if (ctx->parent_ctx)
put_ctx(ctx->parent_ctx);
if (ctx->task && ctx->task != TASK_TOMBSTONE)
put_task_struct(ctx->task);
call_rcu(&ctx->rcu_head, free_ctx);
} else {
smp_mb__after_atomic(); /* pairs with wait_var_event() */
if (ctx->task == TASK_TOMBSTONE)
wake_up_var(&ctx->refcount);
}
}
/*
* Because of perf_event::ctx migration in sys_perf_event_open::move_group and
* perf_pmu_migrate_context() we need some magic.
*
* Those places that change perf_event::ctx will hold both
* perf_event_ctx::mutex of the 'old' and 'new' ctx value.
*
* Lock ordering is by mutex address. There are two other sites where
* perf_event_context::mutex nests and those are:
*
* - perf_event_exit_task_context() [ child , 0 ]
* perf_event_exit_event()
* put_event() [ parent, 1 ]
*
* - perf_event_init_context() [ parent, 0 ]
* inherit_task_group()
* inherit_group()
* inherit_event()
* perf_event_alloc()
* perf_init_event()
* perf_try_init_event() [ child , 1 ]
*
* While it appears there is an obvious deadlock here -- the parent and child
* nesting levels are inverted between the two. This is in fact safe because
* life-time rules separate them. That is an exiting task cannot fork, and a
* spawning task cannot (yet) exit.
*
* But remember that these are parent<->child context relations, and
* migration does not affect children, therefore these two orderings should not
* interact.
*
* The change in perf_event::ctx does not affect children (as claimed above)
* because the sys_perf_event_open() case will install a new event and break
* the ctx parent<->child relation, and perf_pmu_migrate_context() is only
* concerned with cpuctx and that doesn't have children.
*
* The places that change perf_event::ctx will issue:
*
* perf_remove_from_context();
* synchronize_rcu();
* perf_install_in_context();
*
* to affect the change. The remove_from_context() + synchronize_rcu() should
* quiesce the event, after which we can install it in the new location. This
* means that only external vectors (perf_fops, prctl) can perturb the event
* while in transit. Therefore all such accessors should also acquire
* perf_event_context::mutex to serialize against this.
*
* However; because event->ctx can change while we're waiting to acquire
* ctx->mutex we must be careful and use the below perf_event_ctx_lock()
* function.
*
* Lock order:
* exec_update_lock
* task_struct::perf_event_mutex
* perf_event_context::mutex
* perf_event::child_mutex;
* perf_event_context::lock
* mmap_lock
* perf_event::mmap_mutex
* perf_buffer::aux_mutex
* perf_addr_filters_head::lock
*
* cpu_hotplug_lock
* pmus_lock
* cpuctx->mutex / perf_event_context::mutex
*/
static struct perf_event_context *
perf_event_ctx_lock_nested(struct perf_event *event, int nesting)
{
struct perf_event_context *ctx;
again:
rcu_read_lock();
ctx = READ_ONCE(event->ctx);
if (!refcount_inc_not_zero(&ctx->refcount)) {
rcu_read_unlock();
goto again;
}
rcu_read_unlock();
mutex_lock_nested(&ctx->mutex, nesting);
if (event->ctx != ctx) {
mutex_unlock(&ctx->mutex);
put_ctx(ctx);
goto again;
}
return ctx;
}
static inline struct perf_event_context *
perf_event_ctx_lock(struct perf_event *event)
{
return perf_event_ctx_lock_nested(event, 0);
}
static void perf_event_ctx_unlock(struct perf_event *event,
struct perf_event_context *ctx)
{
mutex_unlock(&ctx->mutex);
put_ctx(ctx);
}
/*
* This must be done under the ctx->lock, such as to serialize against
* context_equiv(), therefore we cannot call put_ctx() since that might end up
* calling scheduler related locks and ctx->lock nests inside those.
*/
static __must_check struct perf_event_context *
unclone_ctx(struct perf_event_context *ctx)
{
struct perf_event_context *parent_ctx = ctx->parent_ctx;
lockdep_assert_held(&ctx->lock);
if (parent_ctx)
ctx->parent_ctx = NULL;
ctx->generation++;
return parent_ctx;
}
static u32 perf_event_pid_type(struct perf_event *event, struct task_struct *p,
enum pid_type type)
{
u32 nr;
/*
* only top level events have the pid namespace they were created in
*/
if (event->parent)
event = event->parent;
nr = __task_pid_nr_ns(p, type, event->ns);
/* avoid -1 if it is idle thread or runs in another ns */
if (!nr && !pid_alive(p))
nr = -1;
return nr;
}
static u32 perf_event_pid(struct perf_event *event, struct task_struct *p)
{
return perf_event_pid_type(event, p, PIDTYPE_TGID);
}
static u32 perf_event_tid(struct perf_event *event, struct task_struct *p)
{
return perf_event_pid_type(event, p, PIDTYPE_PID);
}
/*
* If we inherit events we want to return the parent event id
* to userspace.
*/
static u64 primary_event_id(struct perf_event *event)
{
u64 id = event->id;
if (event->parent)
id = event->parent->id;
return id;
}
/*
* Get the perf_event_context for a task and lock it.
*
* This has to cope with the fact that until it is locked,
* the context could get moved to another task.
*/
static struct perf_event_context *
perf_lock_task_context(struct task_struct *task, unsigned long *flags)
{
struct perf_event_context *ctx;
retry:
/*
* One of the few rules of preemptible RCU is that one cannot do
* rcu_read_unlock() while holding a scheduler (or nested) lock when
* part of the read side critical section was irqs-enabled -- see
* rcu_read_unlock_special().
*
* Since ctx->lock nests under rq->lock we must ensure the entire read
* side critical section has interrupts disabled.
*/
local_irq_save(*flags);
rcu_read_lock();
ctx = rcu_dereference(task->perf_event_ctxp);
if (ctx) {
/*
* If this context is a clone of another, it might
* get swapped for another underneath us by
* perf_event_task_sched_out, though the
* rcu_read_lock() protects us from any context
* getting freed. Lock the context and check if it
* got swapped before we could get the lock, and retry
* if so. If we locked the right context, then it
* can't get swapped on us any more.
*/
raw_spin_lock(&ctx->lock);
if (ctx != rcu_dereference(task->perf_event_ctxp)) {
raw_spin_unlock(&ctx->lock);
rcu_read_unlock();
local_irq_restore(*flags);
goto retry;
}
if (ctx->task == TASK_TOMBSTONE ||
!refcount_inc_not_zero(&ctx->refcount)) {
raw_spin_unlock(&ctx->lock);
ctx = NULL;
} else {
WARN_ON_ONCE(ctx->task != task);
}
}
rcu_read_unlock();
if (!ctx)
local_irq_restore(*flags);
return ctx;
}
/*
* Get the context for a task and increment its pin_count so it
* can't get swapped to another task. This also increments its
* reference count so that the context can't get freed.
*/
static struct perf_event_context *
perf_pin_task_context(struct task_struct *task)
{
struct perf_event_context *ctx;
unsigned long flags;
ctx = perf_lock_task_context(task, &flags);
if (ctx) {
++ctx->pin_count;
raw_spin_unlock_irqrestore(&ctx->lock, flags);
}
return ctx;
}
static void perf_unpin_context(struct perf_event_context *ctx)
{
unsigned long flags;
raw_spin_lock_irqsave(&ctx->lock, flags);
--ctx->pin_count;
raw_spin_unlock_irqrestore(&ctx->lock, flags);
}
/*
* Update the record of the current time in a context.
*/
static void __update_context_time(struct perf_event_context *ctx, bool adv)
{
u64 now = perf_clock();
lockdep_assert_held(&ctx->lock);
if (adv)
ctx->time += now - ctx->timestamp;
ctx->timestamp = now;
/*
* The above: time' = time + (now - timestamp), can be re-arranged
* into: time` = now + (time - timestamp), which gives a single value
* offset to compute future time without locks on.
*
* See perf_event_time_now(), which can be used from NMI context where
* it's (obviously) not possible to acquire ctx->lock in order to read
* both the above values in a consistent manner.
*/
WRITE_ONCE(ctx->timeoffset, ctx->time - ctx->timestamp);
}
static void update_context_time(struct perf_event_context *ctx)
{
__update_context_time(ctx, true);
}
static u64 perf_event_time(struct perf_event *event)
{
struct perf_event_context *ctx = event->ctx;
if (unlikely(!ctx))
return 0;
if (is_cgroup_event(event))
return perf_cgroup_event_time(event);
return ctx->time;
}
static u64 perf_event_time_now(struct perf_event *event, u64 now)
{
struct perf_event_context *ctx = event->ctx;
if (unlikely(!ctx))
return 0;
if (is_cgroup_event(event))
return perf_cgroup_event_time_now(event, now);
if (!(__load_acquire(&ctx->is_active) & EVENT_TIME))
return ctx->time;
now += READ_ONCE(ctx->timeoffset);
return now;
}
static enum event_type_t get_event_type(struct perf_event *event)
{
struct perf_event_context *ctx = event->ctx;
enum event_type_t event_type;
lockdep_assert_held(&ctx->lock);
/*
* It's 'group type', really, because if our group leader is
* pinned, so are we.
*/
if (event->group_leader != event)
event = event->group_leader;
event_type = event->attr.pinned ? EVENT_PINNED : EVENT_FLEXIBLE;
if (!ctx->task)
event_type |= EVENT_CPU;
return event_type;
}
/*
* Helper function to initialize event group nodes.
*/
static void init_event_group(struct perf_event *event)
{
RB_CLEAR_NODE(&event->group_node);
event->group_index = 0;
}
/*
* Extract pinned or flexible groups from the context
* based on event attrs bits.
*/
static struct perf_event_groups *
get_event_groups(struct perf_event *event, struct perf_event_context *ctx)
{
if (event->attr.pinned)
return &ctx->pinned_groups;
else
return &ctx->flexible_groups;
}
/*
* Helper function to initializes perf_event_group trees.
*/
static void perf_event_groups_init(struct perf_event_groups *groups)
{
groups->tree = RB_ROOT;
groups->index = 0;
}
static inline struct cgroup *event_cgroup(const struct perf_event *event)
{
struct cgroup *cgroup = NULL;
#ifdef CONFIG_CGROUP_PERF
if (event->cgrp)
cgroup = event->cgrp->css.cgroup;
#endif
return cgroup;
}
/*
* Compare function for event groups;
*
* Implements complex key that first sorts by CPU and then by virtual index
* which provides ordering when rotating groups for the same CPU.
*/
static __always_inline int
perf_event_groups_cmp(const int left_cpu, const struct pmu *left_pmu,
const struct cgroup *left_cgroup, const u64 left_group_index,
const struct perf_event *right)
{
if (left_cpu < right->cpu)
return -1;
if (left_cpu > right->cpu)
return 1;
if (left_pmu) {
if (left_pmu < right->pmu_ctx->pmu)
return -1;
if (left_pmu > right->pmu_ctx->pmu)
return 1;
}
#ifdef CONFIG_CGROUP_PERF
{
const struct cgroup *right_cgroup = event_cgroup(right);
if (left_cgroup != right_cgroup) {
if (!left_cgroup) {
/*
* Left has no cgroup but right does, no
* cgroups come first.
*/
return -1;
}
if (!right_cgroup) {
/*
* Right has no cgroup but left does, no
* cgroups come first.
*/
return 1;
}
/* Two dissimilar cgroups, order by id. */
if (cgroup_id(left_cgroup) < cgroup_id(right_cgroup))
return -1;
return 1;
}
}
#endif
if (left_group_index < right->group_index)
return -1;
if (left_group_index > right->group_index)
return 1;
return 0;
}
#define __node_2_pe(node) \
rb_entry((node), struct perf_event, group_node)
static inline bool __group_less(struct rb_node *a, const struct rb_node *b)
{
struct perf_event *e = __node_2_pe(a);
return perf_event_groups_cmp(e->cpu, e->pmu_ctx->pmu, event_cgroup(e),
e->group_index, __node_2_pe(b)) < 0;
}
struct __group_key {
int cpu;
struct pmu *pmu;
struct cgroup *cgroup;
};
static inline int __group_cmp(const void *key, const struct rb_node *node)
{
const struct __group_key *a = key;
const struct perf_event *b = __node_2_pe(node);
/* partial/subtree match: @cpu, @pmu, @cgroup; ignore: @group_index */
return perf_event_groups_cmp(a->cpu, a->pmu, a->cgroup, b->group_index, b);
}
static inline int
__group_cmp_ignore_cgroup(const void *key, const struct rb_node *node)
{
const struct __group_key *a = key;
const struct perf_event *b = __node_2_pe(node);
/* partial/subtree match: @cpu, @pmu, ignore: @cgroup, @group_index */
return perf_event_groups_cmp(a->cpu, a->pmu, event_cgroup(b),
b->group_index, b);
}
/*
* Insert @event into @groups' tree; using
* {@event->cpu, @event->pmu_ctx->pmu, event_cgroup(@event), ++@groups->index}
* as key. This places it last inside the {cpu,pmu,cgroup} subtree.
*/
static void
perf_event_groups_insert(struct perf_event_groups *groups,
struct perf_event *event)
{
event->group_index = ++groups->index;
rb_add(&event->group_node, &groups->tree, __group_less);
}
/*
* Helper function to insert event into the pinned or flexible groups.
*/
static void
add_event_to_groups(struct perf_event *event, struct perf_event_context *ctx)
{
struct perf_event_groups *groups;
groups = get_event_groups(event, ctx);
perf_event_groups_insert(groups, event);
}
/*
* Delete a group from a tree.
*/
static void
perf_event_groups_delete(struct perf_event_groups *groups,
struct perf_event *event)
{
WARN_ON_ONCE(RB_EMPTY_NODE(&event->group_node) ||
RB_EMPTY_ROOT(&groups->tree));
rb_erase(&event->group_node, &groups->tree);
init_event_group(event);
}
/*
* Helper function to delete event from its groups.
*/
static void
del_event_from_groups(struct perf_event *event, struct perf_event_context *ctx)
{
struct perf_event_groups *groups;
groups = get_event_groups(event, ctx);
perf_event_groups_delete(groups, event);
}
/*
* Get the leftmost event in the {cpu,pmu,cgroup} subtree.
*/
static struct perf_event *
perf_event_groups_first(struct perf_event_groups *groups, int cpu,
struct pmu *pmu, struct cgroup *cgrp)
{
struct __group_key key = {
.cpu = cpu,
.pmu = pmu,
.cgroup = cgrp,
};
struct rb_node *node;
node = rb_find_first(&key, &groups->tree, __group_cmp);
if (node)
return __node_2_pe(node);
return NULL;
}
static struct perf_event *
perf_event_groups_next(struct perf_event *event, struct pmu *pmu)
{
struct __group_key key = {
.cpu = event->cpu,
.pmu = pmu,
.cgroup = event_cgroup(event),
};
struct rb_node *next;
next = rb_next_match(&key, &event->group_node, __group_cmp);
if (next)
return __node_2_pe(next);
return NULL;
}
#define perf_event_groups_for_cpu_pmu(event, groups, cpu, pmu) \
for (event = perf_event_groups_first(groups, cpu, pmu, NULL); \
event; event = perf_event_groups_next(event, pmu))
/*
* Iterate through the whole groups tree.
*/
#define perf_event_groups_for_each(event, groups) \
for (event = rb_entry_safe(rb_first(&((groups)->tree)), \
typeof(*event), group_node); event; \
event = rb_entry_safe(rb_next(&event->group_node), \
typeof(*event), group_node))
/*
* Does the event attribute request inherit with PERF_SAMPLE_READ
*/
static inline bool has_inherit_and_sample_read(struct perf_event_attr *attr)
{
return attr->inherit && (attr->sample_type & PERF_SAMPLE_READ);
}
/*
* Add an event from the lists for its context.
* Must be called with ctx->mutex and ctx->lock held.
*/
static void
list_add_event(struct perf_event *event, struct perf_event_context *ctx)
{
lockdep_assert_held(&ctx->lock);
WARN_ON_ONCE(event->attach_state & PERF_ATTACH_CONTEXT);
event->attach_state |= PERF_ATTACH_CONTEXT;
event->tstamp = perf_event_time(event);
/*
* If we're a stand alone event or group leader, we go to the context
* list, group events are kept attached to the group so that
* perf_group_detach can, at all times, locate all siblings.
*/
if (event->group_leader == event) {
event->group_caps = event->event_caps;
add_event_to_groups(event, ctx);
}
list_add_rcu(&event->event_entry, &ctx->event_list);
ctx->nr_events++;
if (event->hw.flags & PERF_EVENT_FLAG_USER_READ_CNT)
ctx->nr_user++;
if (event->attr.inherit_stat)
ctx->nr_stat++;
if (has_inherit_and_sample_read(&event->attr))
local_inc(&ctx->nr_no_switch_fast);
if (event->state > PERF_EVENT_STATE_OFF)
perf_cgroup_event_enable(event, ctx);
ctx->generation++;
event->pmu_ctx->nr_events++;
}
/*
* Initialize event state based on the perf_event_attr::disabled.
*/
static inline void perf_event__state_init(struct perf_event *event)
{
event->state = event->attr.disabled ? PERF_EVENT_STATE_OFF :
PERF_EVENT_STATE_INACTIVE;
}
static int __perf_event_read_size(u64 read_format, int nr_siblings)
{
int entry = sizeof(u64); /* value */
int size = 0;
int nr = 1;
if (read_format & PERF_FORMAT_TOTAL_TIME_ENABLED)
size += sizeof(u64);
if (read_format & PERF_FORMAT_TOTAL_TIME_RUNNING)
size += sizeof(u64);
if (read_format & PERF_FORMAT_ID)
entry += sizeof(u64);
if (read_format & PERF_FORMAT_LOST)
entry += sizeof(u64);
if (read_format & PERF_FORMAT_GROUP) {
nr += nr_siblings;
size += sizeof(u64);
}
/*
* Since perf_event_validate_size() limits this to 16k and inhibits
* adding more siblings, this will never overflow.
*/
return size + nr * entry;
}
static void __perf_event_header_size(struct perf_event *event, u64 sample_type)
{
struct perf_sample_data *data;
u16 size = 0;
if (sample_type & PERF_SAMPLE_IP)
size += sizeof(data->ip);
if (sample_type & PERF_SAMPLE_ADDR)
size += sizeof(data->addr);
if (sample_type & PERF_SAMPLE_PERIOD)
size += sizeof(data->period);
if (sample_type & PERF_SAMPLE_WEIGHT_TYPE)
size += sizeof(data->weight.full);
if (sample_type & PERF_SAMPLE_READ)
size += event->read_size;
if (sample_type & PERF_SAMPLE_DATA_SRC)
size += sizeof(data->data_src.val);
if (sample_type & PERF_SAMPLE_TRANSACTION)
size += sizeof(data->txn);
if (sample_type & PERF_SAMPLE_PHYS_ADDR)
size += sizeof(data->phys_addr);
if (sample_type & PERF_SAMPLE_CGROUP)
size += sizeof(data->cgroup);
if (sample_type & PERF_SAMPLE_DATA_PAGE_SIZE)
size += sizeof(data->data_page_size);
if (sample_type & PERF_SAMPLE_CODE_PAGE_SIZE)
size += sizeof(data->code_page_size);
event->header_size = size;
}
/*
* Called at perf_event creation and when events are attached/detached from a
* group.
*/
static void perf_event__header_size(struct perf_event *event)
{
event->read_size =
__perf_event_read_size(event->attr.read_format,
event->group_leader->nr_siblings);
__perf_event_header_size(event, event->attr.sample_type);
}
static void perf_event__id_header_size(struct perf_event *event)
{
struct perf_sample_data *data;
u64 sample_type = event->attr.sample_type;
u16 size = 0;
if (sample_type & PERF_SAMPLE_TID)
size += sizeof(data->tid_entry);
if (sample_type & PERF_SAMPLE_TIME)
size += sizeof(data->time);
if (sample_type & PERF_SAMPLE_IDENTIFIER)
size += sizeof(data->id);
if (sample_type & PERF_SAMPLE_ID)
size += sizeof(data->id);
if (sample_type & PERF_SAMPLE_STREAM_ID)
size += sizeof(data->stream_id);
if (sample_type & PERF_SAMPLE_CPU)
size += sizeof(data->cpu_entry);
event->id_header_size = size;
}
/*
* Check that adding an event to the group does not result in anybody
* overflowing the 64k event limit imposed by the output buffer.
*
* Specifically, check that the read_size for the event does not exceed 16k,
* read_size being the one term that grows with groups size. Since read_size
* depends on per-event read_format, also (re)check the existing events.
*
* This leaves 48k for the constant size fields and things like callchains,
* branch stacks and register sets.
*/
static bool perf_event_validate_size(struct perf_event *event)
{
struct perf_event *sibling, *group_leader = event->group_leader;
if (__perf_event_read_size(event->attr.read_format,
group_leader->nr_siblings + 1) > 16*1024)
return false;
if (__perf_event_read_size(group_leader->attr.read_format,
group_leader->nr_siblings + 1) > 16*1024)
return false;
/*
* When creating a new group leader, group_leader->ctx is initialized
* after the size has been validated, but we cannot safely use
* for_each_sibling_event() until group_leader->ctx is set. A new group
* leader cannot have any siblings yet, so we can safely skip checking
* the non-existent siblings.
*/
if (event == group_leader)
return true;
for_each_sibling_event(sibling, group_leader) {
if (__perf_event_read_size(sibling->attr.read_format,
group_leader->nr_siblings + 1) > 16*1024)
return false;
}
return true;
}
static void perf_group_attach(struct perf_event *event)
{
struct perf_event *group_leader = event->group_leader, *pos;
lockdep_assert_held(&event->ctx->lock);
/*
* We can have double attach due to group movement (move_group) in
* perf_event_open().
*/
if (event->attach_state & PERF_ATTACH_GROUP)
return;
event->attach_state |= PERF_ATTACH_GROUP;
if (group_leader == event)
return;
WARN_ON_ONCE(group_leader->ctx != event->ctx);
group_leader->group_caps &= event->event_caps;
list_add_tail(&event->sibling_list, &group_leader->sibling_list);
group_leader->nr_siblings++;
group_leader->group_generation++;
perf_event__header_size(group_leader);
for_each_sibling_event(pos, group_leader)
perf_event__header_size(pos);
}
/*
* Remove an event from the lists for its context.
* Must be called with ctx->mutex and ctx->lock held.
*/
static void
list_del_event(struct perf_event *event, struct perf_event_context *ctx)
{
WARN_ON_ONCE(event->ctx != ctx);
lockdep_assert_held(&ctx->lock);
/*
* We can have double detach due to exit/hot-unplug + close.
*/
if (!(event->attach_state & PERF_ATTACH_CONTEXT))
return;
event->attach_state &= ~PERF_ATTACH_CONTEXT;
ctx->nr_events--;
if (event->hw.flags & PERF_EVENT_FLAG_USER_READ_CNT)
ctx->nr_user--;
if (event->attr.inherit_stat)
ctx->nr_stat--;
if (has_inherit_and_sample_read(&event->attr))
local_dec(&ctx->nr_no_switch_fast);
list_del_rcu(&event->event_entry);
if (event->group_leader == event)
del_event_from_groups(event, ctx);
ctx->generation++;
event->pmu_ctx->nr_events--;
}
static int
perf_aux_output_match(struct perf_event *event, struct perf_event *aux_event)
{
if (!has_aux(aux_event))
return 0;
if (!event->pmu->aux_output_match)
return 0;
return event->pmu->aux_output_match(aux_event);
}
static void put_event(struct perf_event *event);
static void __event_disable(struct perf_event *event,
struct perf_event_context *ctx,
enum perf_event_state state);
static void perf_put_aux_event(struct perf_event *event)
{
struct perf_event_context *ctx = event->ctx;
struct perf_event *iter;
/*
* If event uses aux_event tear down the link
*/
if (event->aux_event) {
iter = event->aux_event;
event->aux_event = NULL;
put_event(iter);
return;
}
/*
* If the event is an aux_event, tear down all links to
* it from other events.
*/
for_each_sibling_event(iter, event) {
if (iter->aux_event != event)
continue;
iter->aux_event = NULL;
put_event(event);
/*
* If it's ACTIVE, schedule it out and put it into ERROR
* state so that we don't try to schedule it again. Note
* that perf_event_enable() will clear the ERROR status.
*/
__event_disable(iter, ctx, PERF_EVENT_STATE_ERROR);
}
}
static bool perf_need_aux_event(struct perf_event *event)
{
return event->attr.aux_output || has_aux_action(event);
}
static int perf_get_aux_event(struct perf_event *event,
struct perf_event *group_leader)
{
/*
* Our group leader must be an aux event if we want to be
* an aux_output. This way, the aux event will precede its
* aux_output events in the group, and therefore will always
* schedule first.
*/
if (!group_leader)
return 0;
/*
* aux_output and aux_sample_size are mutually exclusive.
*/
if (event->attr.aux_output && event->attr.aux_sample_size)
return 0;
if (event->attr.aux_output &&
!perf_aux_output_match(event, group_leader))
return 0;
if ((event->attr.aux_pause || event->attr.aux_resume) &&
!(group_leader->pmu->capabilities & PERF_PMU_CAP_AUX_PAUSE))
return 0;
if (event->attr.aux_sample_size && !group_leader->pmu->snapshot_aux)
return 0;
if (!atomic_long_inc_not_zero(&group_leader->refcount))
return 0;
/*
* Link aux_outputs to their aux event; this is undone in
* perf_group_detach() by perf_put_aux_event(). When the
* group in torn down, the aux_output events loose their
* link to the aux_event and can't schedule any more.
*/
event->aux_event = group_leader;
return 1;
}
static inline struct list_head *get_event_list(struct perf_event *event)
{
return event->attr.pinned ? &event->pmu_ctx->pinned_active :
&event->pmu_ctx->flexible_active;
}
static void perf_group_detach(struct perf_event *event)
{
struct perf_event *leader = event->group_leader;
struct perf_event *sibling, *tmp;
struct perf_event_context *ctx = event->ctx;
lockdep_assert_held(&ctx->lock);
/*
* We can have double detach due to exit/hot-unplug + close.
*/
if (!(event->attach_state & PERF_ATTACH_GROUP))
return;
event->attach_state &= ~PERF_ATTACH_GROUP;
perf_put_aux_event(event);
/*
* If this is a sibling, remove it from its group.
*/
if (leader != event) {
list_del_init(&event->sibling_list);
event->group_leader->nr_siblings--;
event->group_leader->group_generation++;
goto out;
}
/*
* If this was a group event with sibling events then
* upgrade the siblings to singleton events by adding them
* to whatever list we are on.
*/
list_for_each_entry_safe(sibling, tmp, &event->sibling_list, sibling_list) {
/*
* Events that have PERF_EV_CAP_SIBLING require being part of
* a group and cannot exist on their own, schedule them out
* and move them into the ERROR state. Also see
* _perf_event_enable(), it will not be able to recover this
* ERROR state.
*/
if (sibling->event_caps & PERF_EV_CAP_SIBLING)
__event_disable(sibling, ctx, PERF_EVENT_STATE_ERROR);
sibling->group_leader = sibling;
list_del_init(&sibling->sibling_list);
/* Inherit group flags from the previous leader */
sibling->group_caps = event->group_caps;
if (sibling->attach_state & PERF_ATTACH_CONTEXT) {
add_event_to_groups(sibling, event->ctx);
if (sibling->state == PERF_EVENT_STATE_ACTIVE)
list_add_tail(&sibling->active_list, get_event_list(sibling));
}
WARN_ON_ONCE(sibling->ctx != event->ctx);
}
out:
for_each_sibling_event(tmp, leader)
perf_event__header_size(tmp);
perf_event__header_size(leader);
}
static void perf_child_detach(struct perf_event *event)
{
struct perf_event *parent_event = event->parent;
if (!(event->attach_state & PERF_ATTACH_CHILD))
return;
event->attach_state &= ~PERF_ATTACH_CHILD;
if (WARN_ON_ONCE(!parent_event))
return;
/*
* Can't check this from an IPI, the holder is likey another CPU.
*
lockdep_assert_held(&parent_event->child_mutex);
*/
list_del_init(&event->child_list);
}
static bool is_orphaned_event(struct perf_event *event)
{
return event->state == PERF_EVENT_STATE_DEAD;
}
static inline int
event_filter_match(struct perf_event *event)
{
return (event->cpu == -1 || event->cpu == smp_processor_id()) &&
perf_cgroup_match(event);
}
static inline bool is_event_in_freq_mode(struct perf_event *event)
{
return event->attr.freq && event->attr.sample_freq;
}
static void
event_sched_out(struct perf_event *event, struct perf_event_context *ctx)
{
struct perf_event_pmu_context *epc = event->pmu_ctx;
struct perf_cpu_pmu_context *cpc = this_cpc(epc->pmu);
enum perf_event_state state = PERF_EVENT_STATE_INACTIVE;
// XXX cpc serialization, probably per-cpu IRQ disabled
WARN_ON_ONCE(event->ctx != ctx);
lockdep_assert_held(&ctx->lock);
if (event->state != PERF_EVENT_STATE_ACTIVE)
return;
/*
* Asymmetry; we only schedule events _IN_ through ctx_sched_in(), but
* we can schedule events _OUT_ individually through things like
* __perf_remove_from_context().
*/
list_del_init(&event->active_list);
perf_pmu_disable(event->pmu);
event->pmu->del(event, 0);
event->oncpu = -1;
if (event->pending_disable) {
event->pending_disable = 0;
perf_cgroup_event_disable(event, ctx);
state = PERF_EVENT_STATE_OFF;
}
perf_event_set_state(event, state);
if (!is_software_event(event))
cpc->active_oncpu--;
if (is_event_in_freq_mode(event)) {
ctx->nr_freq--;
epc->nr_freq--;
}
if (event->attr.exclusive || !cpc->active_oncpu)
cpc->exclusive = 0;
perf_pmu_enable(event->pmu);
}
static void
group_sched_out(struct perf_event *group_event, struct perf_event_context *ctx)
{
struct perf_event *event;
if (group_event->state != PERF_EVENT_STATE_ACTIVE)
return;
perf_assert_pmu_disabled(group_event->pmu_ctx->pmu);
event_sched_out(group_event, ctx);
/*
* Schedule out siblings (if any):
*/
for_each_sibling_event(event, group_event)
event_sched_out(event, ctx);
}
static inline void
__ctx_time_update(struct perf_cpu_context *cpuctx, struct perf_event_context *ctx, bool final)
{
if (ctx->is_active & EVENT_TIME) {
if (ctx->is_active & EVENT_FROZEN)
return;
update_context_time(ctx);
update_cgrp_time_from_cpuctx(cpuctx, final);
}
}
static inline void
ctx_time_update(struct perf_cpu_context *cpuctx, struct perf_event_context *ctx)
{
__ctx_time_update(cpuctx, ctx, false);
}
/*
* To be used inside perf_ctx_lock() / perf_ctx_unlock(). Lasts until perf_ctx_unlock().
*/
static inline void
ctx_time_freeze(struct perf_cpu_context *cpuctx, struct perf_event_context *ctx)
{
ctx_time_update(cpuctx, ctx);
if (ctx->is_active & EVENT_TIME)
ctx->is_active |= EVENT_FROZEN;
}
static inline void
ctx_time_update_event(struct perf_event_context *ctx, struct perf_event *event)
{
if (ctx->is_active & EVENT_TIME) {
if (ctx->is_active & EVENT_FROZEN)
return;
update_context_time(ctx);
update_cgrp_time_from_event(event);
}
}
#define DETACH_GROUP 0x01UL
#define DETACH_CHILD 0x02UL
#define DETACH_EXIT 0x04UL
#define DETACH_REVOKE 0x08UL
#define DETACH_DEAD 0x10UL
/*
* Cross CPU call to remove a performance event
*
* We disable the event on the hardware level first. After that we
* remove it from the context list.
*/
static void
__perf_remove_from_context(struct perf_event *event,
struct perf_cpu_context *cpuctx,
struct perf_event_context *ctx,
void *info)
{
struct perf_event_pmu_context *pmu_ctx = event->pmu_ctx;
enum perf_event_state state = PERF_EVENT_STATE_OFF;
unsigned long flags = (unsigned long)info;
ctx_time_update(cpuctx, ctx);
/*
* Ensure event_sched_out() switches to OFF, at the very least
* this avoids raising perf_pending_task() at this time.
*/
if (flags & DETACH_EXIT)
state = PERF_EVENT_STATE_EXIT;
if (flags & DETACH_REVOKE)
state = PERF_EVENT_STATE_REVOKED;
if (flags & DETACH_DEAD)
state = PERF_EVENT_STATE_DEAD;
event_sched_out(event, ctx);
if (event->state > PERF_EVENT_STATE_OFF)
perf_cgroup_event_disable(event, ctx);
perf_event_set_state(event, min(event->state, state));
if (flags & DETACH_GROUP)
perf_group_detach(event);
if (flags & DETACH_CHILD)
perf_child_detach(event);
list_del_event(event, ctx);
if (!pmu_ctx->nr_events) {
pmu_ctx->rotate_necessary = 0;
if (ctx->task && ctx->is_active) {
struct perf_cpu_pmu_context *cpc = this_cpc(pmu_ctx->pmu);
WARN_ON_ONCE(cpc->task_epc && cpc->task_epc != pmu_ctx);
cpc->task_epc = NULL;
}
}
if (!ctx->nr_events && ctx->is_active) {
if (ctx == &cpuctx->ctx)
update_cgrp_time_from_cpuctx(cpuctx, true);
ctx->is_active = 0;
if (ctx->task) {
WARN_ON_ONCE(cpuctx->task_ctx != ctx);
cpuctx->task_ctx = NULL;
}
}
}
/*
* Remove the event from a task's (or a CPU's) list of events.
*
* If event->ctx is a cloned context, callers must make sure that
* every task struct that event->ctx->task could possibly point to
* remains valid. This is OK when called from perf_release since
* that only calls us on the top-level context, which can't be a clone.
* When called from perf_event_exit_task, it's OK because the
* context has been detached from its task.
*/
static void perf_remove_from_context(struct perf_event *event, unsigned long flags)
{
struct perf_event_context *ctx = event->ctx;
lockdep_assert_held(&ctx->mutex);
/*
* Because of perf_event_exit_task(), perf_remove_from_context() ought
* to work in the face of TASK_TOMBSTONE, unlike every other
* event_function_call() user.
*/
raw_spin_lock_irq(&ctx->lock);
if (!ctx->is_active) {
__perf_remove_from_context(event, this_cpu_ptr(&perf_cpu_context),
ctx, (void *)flags);
raw_spin_unlock_irq(&ctx->lock);
return;
}
raw_spin_unlock_irq(&ctx->lock);
event_function_call(event, __perf_remove_from_context, (void *)flags);
}
static void __event_disable(struct perf_event *event,
struct perf_event_context *ctx,
enum perf_event_state state)
{
event_sched_out(event, ctx);
perf_cgroup_event_disable(event, ctx);
perf_event_set_state(event, state);
}
/*
* Cross CPU call to disable a performance event
*/
static void __perf_event_disable(struct perf_event *event,
struct perf_cpu_context *cpuctx,
struct perf_event_context *ctx,
void *info)
{
if (event->state < PERF_EVENT_STATE_INACTIVE)
return;
perf_pmu_disable(event->pmu_ctx->pmu);
ctx_time_update_event(ctx, event);
/*
* When disabling a group leader, the whole group becomes ineligible
* to run, so schedule out the full group.
*/
if (event == event->group_leader)
group_sched_out(event, ctx);
/*
* But only mark the leader OFF; the siblings will remain
* INACTIVE.
*/
__event_disable(event, ctx, PERF_EVENT_STATE_OFF);
perf_pmu_enable(event->pmu_ctx->pmu);
}
/*
* Disable an event.
*
* If event->ctx is a cloned context, callers must make sure that
* every task struct that event->ctx->task could possibly point to
* remains valid. This condition is satisfied when called through
* perf_event_for_each_child or perf_event_for_each because they
* hold the top-level event's child_mutex, so any descendant that
* goes to exit will block in perf_event_exit_event().
*
* When called from perf_pending_disable it's OK because event->ctx
* is the current context on this CPU and preemption is disabled,
* hence we can't get into perf_event_task_sched_out for this context.
*/
static void _perf_event_disable(struct perf_event *event)
{
struct perf_event_context *ctx = event->ctx;
raw_spin_lock_irq(&ctx->lock);
if (event->state <= PERF_EVENT_STATE_OFF) {
raw_spin_unlock_irq(&ctx->lock);
return;
}
raw_spin_unlock_irq(&ctx->lock);
event_function_call(event, __perf_event_disable, NULL);
}
void perf_event_disable_local(struct perf_event *event)
{
event_function_local(event, __perf_event_disable, NULL);
}
/*
* Strictly speaking kernel users cannot create groups and therefore this
* interface does not need the perf_event_ctx_lock() magic.
*/
void perf_event_disable(struct perf_event *event)
{
struct perf_event_context *ctx;
ctx = perf_event_ctx_lock(event);
_perf_event_disable(event);
perf_event_ctx_unlock(event, ctx);
}
EXPORT_SYMBOL_GPL(perf_event_disable);
void perf_event_disable_inatomic(struct perf_event *event)
{
event->pending_disable = 1;
irq_work_queue(&event->pending_disable_irq);
}
#define MAX_INTERRUPTS (~0ULL)
static void perf_log_throttle(struct perf_event *event, int enable);
static void perf_log_itrace_start(struct perf_event *event);
static void perf_event_unthrottle(struct perf_event *event, bool start)
{
if (event->state != PERF_EVENT_STATE_ACTIVE)
return;
event->hw.interrupts = 0;
if (start)
event->pmu->start(event, 0);
if (event == event->group_leader)
perf_log_throttle(event, 1);
}
static void perf_event_throttle(struct perf_event *event)
{
if (event->state != PERF_EVENT_STATE_ACTIVE)
return;
event->hw.interrupts = MAX_INTERRUPTS;
event->pmu->stop(event, 0);
if (event == event->group_leader)
perf_log_throttle(event, 0);
}
static void perf_event_unthrottle_group(struct perf_event *event, bool skip_start_event)
{
struct perf_event *sibling, *leader = event->group_leader;
perf_event_unthrottle(leader, skip_start_event ? leader != event : true);
for_each_sibling_event(sibling, leader)
perf_event_unthrottle(sibling, skip_start_event ? sibling != event : true);
}
static void perf_event_throttle_group(struct perf_event *event)
{
struct perf_event *sibling, *leader = event->group_leader;
perf_event_throttle(leader);
for_each_sibling_event(sibling, leader)
perf_event_throttle(sibling);
}
static int
event_sched_in(struct perf_event *event, struct perf_event_context *ctx)
{
struct perf_event_pmu_context *epc = event->pmu_ctx;
struct perf_cpu_pmu_context *cpc = this_cpc(epc->pmu);
int ret = 0;
WARN_ON_ONCE(event->ctx != ctx);
lockdep_assert_held(&ctx->lock);
if (event->state <= PERF_EVENT_STATE_OFF)
return 0;
WRITE_ONCE(event->oncpu, smp_processor_id());
/*
* Order event::oncpu write to happen before the ACTIVE state is
* visible. This allows perf_event_{stop,read}() to observe the correct
* ->oncpu if it sees ACTIVE.
*/
smp_wmb();
perf_event_set_state(event, PERF_EVENT_STATE_ACTIVE);
/*
* Unthrottle events, since we scheduled we might have missed several
* ticks already, also for a heavily scheduling task there is little
* guarantee it'll get a tick in a timely manner.
*/
if (unlikely(event->hw.interrupts == MAX_INTERRUPTS))
perf_event_unthrottle(event, false);
perf_pmu_disable(event->pmu);
perf_log_itrace_start(event);
if (event->pmu->add(event, PERF_EF_START)) {
perf_event_set_state(event, PERF_EVENT_STATE_INACTIVE);
event->oncpu = -1;
ret = -EAGAIN;
goto out;
}
if (!is_software_event(event))
cpc->active_oncpu++;
if (is_event_in_freq_mode(event)) {
ctx->nr_freq++;
epc->nr_freq++;
}
if (event->attr.exclusive)
cpc->exclusive = 1;
out:
perf_pmu_enable(event->pmu);
return ret;
}
static int
group_sched_in(struct perf_event *group_event, struct perf_event_context *ctx)
{
struct perf_event *event, *partial_group = NULL;
struct pmu *pmu = group_event->pmu_ctx->pmu;
if (group_event->state == PERF_EVENT_STATE_OFF)
return 0;
pmu->start_txn(pmu, PERF_PMU_TXN_ADD);
if (event_sched_in(group_event, ctx))
goto error;
/*
* Schedule in siblings as one group (if any):
*/
for_each_sibling_event(event, group_event) {
if (event_sched_in(event, ctx)) {
partial_group = event;
goto group_error;
}
}
if (!pmu->commit_txn(pmu))
return 0;
group_error:
/*
* Groups can be scheduled in as one unit only, so undo any
* partial group before returning:
* The events up to the failed event are scheduled out normally.
*/
for_each_sibling_event(event, group_event) {
if (event == partial_group)
break;
event_sched_out(event, ctx);
}
event_sched_out(group_event, ctx);
error:
pmu->cancel_txn(pmu);
return -EAGAIN;
}
/*
* Work out whether we can put this event group on the CPU now.
*/
static int group_can_go_on(struct perf_event *event, int can_add_hw)
{
struct perf_event_pmu_context *epc = event->pmu_ctx;
struct perf_cpu_pmu_context *cpc = this_cpc(epc->pmu);
/*
* Groups consisting entirely of software events can always go on.
*/
if (event->group_caps & PERF_EV_CAP_SOFTWARE)
return 1;
/*
* If an exclusive group is already on, no other hardware
* events can go on.
*/
if (cpc->exclusive)
return 0;
/*
* If this group is exclusive and there are already
* events on the CPU, it can't go on.
*/
if (event->attr.exclusive && !list_empty(get_event_list(event)))
return 0;
/*
* Otherwise, try to add it if all previous groups were able
* to go on.
*/
return can_add_hw;
}
static void add_event_to_ctx(struct perf_event *event,
struct perf_event_context *ctx)
{
list_add_event(event, ctx);
perf_group_attach(event);
}
static void task_ctx_sched_out(struct perf_event_context *ctx,
struct pmu *pmu,
enum event_type_t event_type)
{
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
if (!cpuctx->task_ctx)
return;
if (WARN_ON_ONCE(ctx != cpuctx->task_ctx))
return;
ctx_sched_out(ctx, pmu, event_type);
}
static void perf_event_sched_in(struct perf_cpu_context *cpuctx,
struct perf_event_context *ctx,
struct pmu *pmu)
{
ctx_sched_in(&cpuctx->ctx, pmu, EVENT_PINNED);
if (ctx)
ctx_sched_in(ctx, pmu, EVENT_PINNED);
ctx_sched_in(&cpuctx->ctx, pmu, EVENT_FLEXIBLE);
if (ctx)
ctx_sched_in(ctx, pmu, EVENT_FLEXIBLE);
}
/*
* We want to maintain the following priority of scheduling:
* - CPU pinned (EVENT_CPU | EVENT_PINNED)
* - task pinned (EVENT_PINNED)
* - CPU flexible (EVENT_CPU | EVENT_FLEXIBLE)
* - task flexible (EVENT_FLEXIBLE).
*
* In order to avoid unscheduling and scheduling back in everything every
* time an event is added, only do it for the groups of equal priority and
* below.
*
* This can be called after a batch operation on task events, in which case
* event_type is a bit mask of the types of events involved. For CPU events,
* event_type is only either EVENT_PINNED or EVENT_FLEXIBLE.
*/
static void ctx_resched(struct perf_cpu_context *cpuctx,
struct perf_event_context *task_ctx,
struct pmu *pmu, enum event_type_t event_type)
{
bool cpu_event = !!(event_type & EVENT_CPU);
struct perf_event_pmu_context *epc;
/*
* If pinned groups are involved, flexible groups also need to be
* scheduled out.
*/
if (event_type & EVENT_PINNED)
event_type |= EVENT_FLEXIBLE;
event_type &= EVENT_ALL;
for_each_epc(epc, &cpuctx->ctx, pmu, false)
perf_pmu_disable(epc->pmu);
if (task_ctx) {
for_each_epc(epc, task_ctx, pmu, false)
perf_pmu_disable(epc->pmu);
task_ctx_sched_out(task_ctx, pmu, event_type);
}
/*
* Decide which cpu ctx groups to schedule out based on the types
* of events that caused rescheduling:
* - EVENT_CPU: schedule out corresponding groups;
* - EVENT_PINNED task events: schedule out EVENT_FLEXIBLE groups;
* - otherwise, do nothing more.
*/
if (cpu_event)
ctx_sched_out(&cpuctx->ctx, pmu, event_type);
else if (event_type & EVENT_PINNED)
ctx_sched_out(&cpuctx->ctx, pmu, EVENT_FLEXIBLE);
perf_event_sched_in(cpuctx, task_ctx, pmu);
for_each_epc(epc, &cpuctx->ctx, pmu, false)
perf_pmu_enable(epc->pmu);
if (task_ctx) {
for_each_epc(epc, task_ctx, pmu, false)
perf_pmu_enable(epc->pmu);
}
}
void perf_pmu_resched(struct pmu *pmu)
{
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
struct perf_event_context *task_ctx = cpuctx->task_ctx;
perf_ctx_lock(cpuctx, task_ctx);
ctx_resched(cpuctx, task_ctx, pmu, EVENT_ALL|EVENT_CPU);
perf_ctx_unlock(cpuctx, task_ctx);
}
/*
* Cross CPU call to install and enable a performance event
*
* Very similar to remote_function() + event_function() but cannot assume that
* things like ctx->is_active and cpuctx->task_ctx are set.
*/
static int __perf_install_in_context(void *info)
{
struct perf_event *event = info;
struct perf_event_context *ctx = event->ctx;
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
struct perf_event_context *task_ctx = cpuctx->task_ctx;
bool reprogram = true;
int ret = 0;
raw_spin_lock(&cpuctx->ctx.lock);
if (ctx->task) {
raw_spin_lock(&ctx->lock);
task_ctx = ctx;
reprogram = (ctx->task == current);
/*
* If the task is running, it must be running on this CPU,
* otherwise we cannot reprogram things.
*
* If its not running, we don't care, ctx->lock will
* serialize against it becoming runnable.
*/
if (task_curr(ctx->task) && !reprogram) {
ret = -ESRCH;
goto unlock;
}
WARN_ON_ONCE(reprogram && cpuctx->task_ctx && cpuctx->task_ctx != ctx);
} else if (task_ctx) {
raw_spin_lock(&task_ctx->lock);
}
#ifdef CONFIG_CGROUP_PERF
if (event->state > PERF_EVENT_STATE_OFF && is_cgroup_event(event)) {
/*
* If the current cgroup doesn't match the event's
* cgroup, we should not try to schedule it.
*/
struct perf_cgroup *cgrp = perf_cgroup_from_task(current, ctx);
reprogram = cgroup_is_descendant(cgrp->css.cgroup,
event->cgrp->css.cgroup);
}
#endif
if (reprogram) {
ctx_time_freeze(cpuctx, ctx);
add_event_to_ctx(event, ctx);
ctx_resched(cpuctx, task_ctx, event->pmu_ctx->pmu,
get_event_type(event));
} else {
add_event_to_ctx(event, ctx);
}
unlock:
perf_ctx_unlock(cpuctx, task_ctx);
return ret;
}
static bool exclusive_event_installable(struct perf_event *event,
struct perf_event_context *ctx);
/*
* Attach a performance event to a context.
*
* Very similar to event_function_call, see comment there.
*/
static void
perf_install_in_context(struct perf_event_context *ctx,
struct perf_event *event,
int cpu)
{
struct task_struct *task = READ_ONCE(ctx->task);
lockdep_assert_held(&ctx->mutex);
WARN_ON_ONCE(!exclusive_event_installable(event, ctx));
if (event->cpu != -1)
WARN_ON_ONCE(event->cpu != cpu);
/*
* Ensures that if we can observe event->ctx, both the event and ctx
* will be 'complete'. See perf_iterate_sb_cpu().
*/
smp_store_release(&event->ctx, ctx);
/*
* perf_event_attr::disabled events will not run and can be initialized
* without IPI. Except when this is the first event for the context, in
* that case we need the magic of the IPI to set ctx->is_active.
*
* The IOC_ENABLE that is sure to follow the creation of a disabled
* event will issue the IPI and reprogram the hardware.
*/
if (__perf_effective_state(event) == PERF_EVENT_STATE_OFF &&
ctx->nr_events && !is_cgroup_event(event)) {
raw_spin_lock_irq(&ctx->lock);
if (ctx->task == TASK_TOMBSTONE) {
raw_spin_unlock_irq(&ctx->lock);
return;
}
add_event_to_ctx(event, ctx);
raw_spin_unlock_irq(&ctx->lock);
return;
}
if (!task) {
cpu_function_call(cpu, __perf_install_in_context, event);
return;
}
/*
* Should not happen, we validate the ctx is still alive before calling.
*/
if (WARN_ON_ONCE(task == TASK_TOMBSTONE))
return;
/*
* Installing events is tricky because we cannot rely on ctx->is_active
* to be set in case this is the nr_events 0 -> 1 transition.
*
* Instead we use task_curr(), which tells us if the task is running.
* However, since we use task_curr() outside of rq::lock, we can race
* against the actual state. This means the result can be wrong.
*
* If we get a false positive, we retry, this is harmless.
*
* If we get a false negative, things are complicated. If we are after
* perf_event_context_sched_in() ctx::lock will serialize us, and the
* value must be correct. If we're before, it doesn't matter since
* perf_event_context_sched_in() will program the counter.
*
* However, this hinges on the remote context switch having observed
* our task->perf_event_ctxp[] store, such that it will in fact take
* ctx::lock in perf_event_context_sched_in().
*
* We do this by task_function_call(), if the IPI fails to hit the task
* we know any future context switch of task must see the
* perf_event_ctpx[] store.
*/
/*
* This smp_mb() orders the task->perf_event_ctxp[] store with the
* task_cpu() load, such that if the IPI then does not find the task
* running, a future context switch of that task must observe the
* store.
*/
smp_mb();
again:
if (!task_function_call(task, __perf_install_in_context, event))
return;
raw_spin_lock_irq(&ctx->lock);
task = ctx->task;
if (WARN_ON_ONCE(task == TASK_TOMBSTONE)) {
/*
* Cannot happen because we already checked above (which also
* cannot happen), and we hold ctx->mutex, which serializes us
* against perf_event_exit_task_context().
*/
raw_spin_unlock_irq(&ctx->lock);
return;
}
/*
* If the task is not running, ctx->lock will avoid it becoming so,
* thus we can safely install the event.
*/
if (task_curr(task)) {
raw_spin_unlock_irq(&ctx->lock);
goto again;
}
add_event_to_ctx(event, ctx);
raw_spin_unlock_irq(&ctx->lock);
}
/*
* Cross CPU call to enable a performance event
*/
static void __perf_event_enable(struct perf_event *event,
struct perf_cpu_context *cpuctx,
struct perf_event_context *ctx,
void *info)
{
struct perf_event *leader = event->group_leader;
struct perf_event_context *task_ctx;
if (event->state >= PERF_EVENT_STATE_INACTIVE ||
event->state <= PERF_EVENT_STATE_ERROR)
return;
ctx_time_freeze(cpuctx, ctx);
perf_event_set_state(event, PERF_EVENT_STATE_INACTIVE);
perf_cgroup_event_enable(event, ctx);
if (!ctx->is_active)
return;
if (!event_filter_match(event))
return;
/*
* If the event is in a group and isn't the group leader,
* then don't put it on unless the group is on.
*/
if (leader != event && leader->state != PERF_EVENT_STATE_ACTIVE)
return;
task_ctx = cpuctx->task_ctx;
if (ctx->task)
WARN_ON_ONCE(task_ctx != ctx);
ctx_resched(cpuctx, task_ctx, event->pmu_ctx->pmu, get_event_type(event));
}
/*
* Enable an event.
*
* If event->ctx is a cloned context, callers must make sure that
* every task struct that event->ctx->task could possibly point to
* remains valid. This condition is satisfied when called through
* perf_event_for_each_child or perf_event_for_each as described
* for perf_event_disable.
*/
static void _perf_event_enable(struct perf_event *event)
{
struct perf_event_context *ctx = event->ctx;
raw_spin_lock_irq(&ctx->lock);
if (event->state >= PERF_EVENT_STATE_INACTIVE ||
event->state < PERF_EVENT_STATE_ERROR) {
out:
raw_spin_unlock_irq(&ctx->lock);
return;
}
/*
* If the event is in error state, clear that first.
*
* That way, if we see the event in error state below, we know that it
* has gone back into error state, as distinct from the task having
* been scheduled away before the cross-call arrived.
*/
if (event->state == PERF_EVENT_STATE_ERROR) {
/*
* Detached SIBLING events cannot leave ERROR state.
*/
if (event->event_caps & PERF_EV_CAP_SIBLING &&
event->group_leader == event)
goto out;
event->state = PERF_EVENT_STATE_OFF;
}
raw_spin_unlock_irq(&ctx->lock);
event_function_call(event, __perf_event_enable, NULL);
}
/*
* See perf_event_disable();
*/
void perf_event_enable(struct perf_event *event)
{
struct perf_event_context *ctx;
ctx = perf_event_ctx_lock(event);
_perf_event_enable(event);
perf_event_ctx_unlock(event, ctx);
}
EXPORT_SYMBOL_GPL(perf_event_enable);
struct stop_event_data {
struct perf_event *event;
unsigned int restart;
};
static int __perf_event_stop(void *info)
{
struct stop_event_data *sd = info;
struct perf_event *event = sd->event;
/* if it's already INACTIVE, do nothing */
if (READ_ONCE(event->state) != PERF_EVENT_STATE_ACTIVE)
return 0;
/* matches smp_wmb() in event_sched_in() */
smp_rmb();
/*
* There is a window with interrupts enabled before we get here,
* so we need to check again lest we try to stop another CPU's event.
*/
if (READ_ONCE(event->oncpu) != smp_processor_id())
return -EAGAIN;
event->pmu->stop(event, PERF_EF_UPDATE);
/*
* May race with the actual stop (through perf_pmu_output_stop()),
* but it is only used for events with AUX ring buffer, and such
* events will refuse to restart because of rb::aux_mmap_count==0,
* see comments in perf_aux_output_begin().
*
* Since this is happening on an event-local CPU, no trace is lost
* while restarting.
*/
if (sd->restart)
event->pmu->start(event, 0);
return 0;
}
static int perf_event_stop(struct perf_event *event, int restart)
{
struct stop_event_data sd = {
.event = event,
.restart = restart,
};
int ret = 0;
do {
if (READ_ONCE(event->state) != PERF_EVENT_STATE_ACTIVE)
return 0;
/* matches smp_wmb() in event_sched_in() */
smp_rmb();
/*
* We only want to restart ACTIVE events, so if the event goes
* inactive here (event->oncpu==-1), there's nothing more to do;
* fall through with ret==-ENXIO.
*/
ret = cpu_function_call(READ_ONCE(event->oncpu),
__perf_event_stop, &sd);
} while (ret == -EAGAIN);
return ret;
}
/*
* In order to contain the amount of racy and tricky in the address filter
* configuration management, it is a two part process:
*
* (p1) when userspace mappings change as a result of (1) or (2) or (3) below,
* we update the addresses of corresponding vmas in
* event::addr_filter_ranges array and bump the event::addr_filters_gen;
* (p2) when an event is scheduled in (pmu::add), it calls
* perf_event_addr_filters_sync() which calls pmu::addr_filters_sync()
* if the generation has changed since the previous call.
*
* If (p1) happens while the event is active, we restart it to force (p2).
*
* (1) perf_addr_filters_apply(): adjusting filters' offsets based on
* pre-existing mappings, called once when new filters arrive via SET_FILTER
* ioctl;
* (2) perf_addr_filters_adjust(): adjusting filters' offsets based on newly
* registered mapping, called for every new mmap(), with mm::mmap_lock down
* for reading;
* (3) perf_event_addr_filters_exec(): clearing filters' offsets in the process
* of exec.
*/
void perf_event_addr_filters_sync(struct perf_event *event)
{
struct perf_addr_filters_head *ifh = perf_event_addr_filters(event);
if (!has_addr_filter(event))
return;
raw_spin_lock(&ifh->lock);
if (event->addr_filters_gen != event->hw.addr_filters_gen) {
event->pmu->addr_filters_sync(event);
event->hw.addr_filters_gen = event->addr_filters_gen;
}
raw_spin_unlock(&ifh->lock);
}
EXPORT_SYMBOL_GPL(perf_event_addr_filters_sync);
static int _perf_event_refresh(struct perf_event *event, int refresh)
{
/*
* not supported on inherited events
*/
if (event->attr.inherit || !is_sampling_event(event))
return -EINVAL;
atomic_add(refresh, &event->event_limit);
_perf_event_enable(event);
return 0;
}
/*
* See perf_event_disable()
*/
int perf_event_refresh(struct perf_event *event, int refresh)
{
struct perf_event_context *ctx;
int ret;
ctx = perf_event_ctx_lock(event);
ret = _perf_event_refresh(event, refresh);
perf_event_ctx_unlock(event, ctx);
return ret;
}
EXPORT_SYMBOL_GPL(perf_event_refresh);
static int perf_event_modify_breakpoint(struct perf_event *bp,
struct perf_event_attr *attr)
{
int err;
_perf_event_disable(bp);
err = modify_user_hw_breakpoint_check(bp, attr, true);
if (!bp->attr.disabled)
_perf_event_enable(bp);
return err;
}
/*
* Copy event-type-independent attributes that may be modified.
*/
static void perf_event_modify_copy_attr(struct perf_event_attr *to,
const struct perf_event_attr *from)
{
to->sig_data = from->sig_data;
}
static int perf_event_modify_attr(struct perf_event *event,
struct perf_event_attr *attr)
{
int (*func)(struct perf_event *, struct perf_event_attr *);
struct perf_event *child;
int err;
if (event->attr.type != attr->type)
return -EINVAL;
switch (event->attr.type) {
case PERF_TYPE_BREAKPOINT:
func = perf_event_modify_breakpoint;
break;
default:
/* Place holder for future additions. */
return -EOPNOTSUPP;
}
WARN_ON_ONCE(event->ctx->parent_ctx);
mutex_lock(&event->child_mutex);
/*
* Event-type-independent attributes must be copied before event-type
* modification, which will validate that final attributes match the
* source attributes after all relevant attributes have been copied.
*/
perf_event_modify_copy_attr(&event->attr, attr);
err = func(event, attr);
if (err)
goto out;
list_for_each_entry(child, &event->child_list, child_list) {
perf_event_modify_copy_attr(&child->attr, attr);
err = func(child, attr);
if (err)
goto out;
}
out:
mutex_unlock(&event->child_mutex);
return err;
}
static void __pmu_ctx_sched_out(struct perf_event_pmu_context *pmu_ctx,
enum event_type_t event_type)
{
struct perf_event_context *ctx = pmu_ctx->ctx;
struct perf_event *event, *tmp;
struct pmu *pmu = pmu_ctx->pmu;
if (ctx->task && !(ctx->is_active & EVENT_ALL)) {
struct perf_cpu_pmu_context *cpc = this_cpc(pmu);
WARN_ON_ONCE(cpc->task_epc && cpc->task_epc != pmu_ctx);
cpc->task_epc = NULL;
}
if (!(event_type & EVENT_ALL))
return;
perf_pmu_disable(pmu);
if (event_type & EVENT_PINNED) {
list_for_each_entry_safe(event, tmp,
&pmu_ctx->pinned_active,
active_list)
group_sched_out(event, ctx);
}
if (event_type & EVENT_FLEXIBLE) {
list_for_each_entry_safe(event, tmp,
&pmu_ctx->flexible_active,
active_list)
group_sched_out(event, ctx);
/*
* Since we cleared EVENT_FLEXIBLE, also clear
* rotate_necessary, is will be reset by
* ctx_flexible_sched_in() when needed.
*/
pmu_ctx->rotate_necessary = 0;
}
perf_pmu_enable(pmu);
}
/*
* Be very careful with the @pmu argument since this will change ctx state.
* The @pmu argument works for ctx_resched(), because that is symmetric in
* ctx_sched_out() / ctx_sched_in() usage and the ctx state ends up invariant.
*
* However, if you were to be asymmetrical, you could end up with messed up
* state, eg. ctx->is_active cleared even though most EPCs would still actually
* be active.
*/
static void
ctx_sched_out(struct perf_event_context *ctx, struct pmu *pmu, enum event_type_t event_type)
{
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
struct perf_event_pmu_context *pmu_ctx;
int is_active = ctx->is_active;
bool cgroup = event_type & EVENT_CGROUP;
event_type &= ~EVENT_CGROUP;
lockdep_assert_held(&ctx->lock);
if (likely(!ctx->nr_events)) {
/*
* See __perf_remove_from_context().
*/
WARN_ON_ONCE(ctx->is_active);
if (ctx->task)
WARN_ON_ONCE(cpuctx->task_ctx);
return;
}
/*
* Always update time if it was set; not only when it changes.
* Otherwise we can 'forget' to update time for any but the last
* context we sched out. For example:
*
* ctx_sched_out(.event_type = EVENT_FLEXIBLE)
* ctx_sched_out(.event_type = EVENT_PINNED)
*
* would only update time for the pinned events.
*/
__ctx_time_update(cpuctx, ctx, ctx == &cpuctx->ctx);
/*
* CPU-release for the below ->is_active store,
* see __load_acquire() in perf_event_time_now()
*/
barrier();
ctx->is_active &= ~event_type;
if (!(ctx->is_active & EVENT_ALL)) {
/*
* For FROZEN, preserve TIME|FROZEN such that perf_event_time_now()
* does not observe a hole. perf_ctx_unlock() will clean up.
*/
if (ctx->is_active & EVENT_FROZEN)
ctx->is_active &= EVENT_TIME_FROZEN;
else
ctx->is_active = 0;
}
if (ctx->task) {
WARN_ON_ONCE(cpuctx->task_ctx != ctx);
if (!(ctx->is_active & EVENT_ALL))
cpuctx->task_ctx = NULL;
}
is_active ^= ctx->is_active; /* changed bits */
for_each_epc(pmu_ctx, ctx, pmu, cgroup)
__pmu_ctx_sched_out(pmu_ctx, is_active);
}
/*
* Test whether two contexts are equivalent, i.e. whether they have both been
* cloned from the same version of the same context.
*
* Equivalence is measured using a generation number in the context that is
* incremented on each modification to it; see unclone_ctx(), list_add_event()
* and list_del_event().
*/
static int context_equiv(struct perf_event_context *ctx1,
struct perf_event_context *ctx2)
{
lockdep_assert_held(&ctx1->lock);
lockdep_assert_held(&ctx2->lock);
/* Pinning disables the swap optimization */
if (ctx1->pin_count || ctx2->pin_count)
return 0;
/* If ctx1 is the parent of ctx2 */
if (ctx1 == ctx2->parent_ctx && ctx1->generation == ctx2->parent_gen)
return 1;
/* If ctx2 is the parent of ctx1 */
if (ctx1->parent_ctx == ctx2 && ctx1->parent_gen == ctx2->generation)
return 1;
/*
* If ctx1 and ctx2 have the same parent; we flatten the parent
* hierarchy, see perf_event_init_context().
*/
if (ctx1->parent_ctx && ctx1->parent_ctx == ctx2->parent_ctx &&
ctx1->parent_gen == ctx2->parent_gen)
return 1;
/* Unmatched */
return 0;
}
static void __perf_event_sync_stat(struct perf_event *event,
struct perf_event *next_event)
{
u64 value;
if (!event->attr.inherit_stat)
return;
/*
* Update the event value, we cannot use perf_event_read()
* because we're in the middle of a context switch and have IRQs
* disabled, which upsets smp_call_function_single(), however
* we know the event must be on the current CPU, therefore we
* don't need to use it.
*/
perf_pmu_read(event);
perf_event_update_time(event);
/*
* In order to keep per-task stats reliable we need to flip the event
* values when we flip the contexts.
*/
value = local64_read(&next_event->count);
value = local64_xchg(&event->count, value);
local64_set(&next_event->count, value);
swap(event->total_time_enabled, next_event->total_time_enabled);
swap(event->total_time_running, next_event->total_time_running);
/*
* Since we swizzled the values, update the user visible data too.
*/
perf_event_update_userpage(event);
perf_event_update_userpage(next_event);
}
static void perf_event_sync_stat(struct perf_event_context *ctx,
struct perf_event_context *next_ctx)
{
struct perf_event *event, *next_event;
if (!ctx->nr_stat)
return;
update_context_time(ctx);
event = list_first_entry(&ctx->event_list,
struct perf_event, event_entry);
next_event = list_first_entry(&next_ctx->event_list,
struct perf_event, event_entry);
while (&event->event_entry != &ctx->event_list &&
&next_event->event_entry != &next_ctx->event_list) {
__perf_event_sync_stat(event, next_event);
event = list_next_entry(event, event_entry);
next_event = list_next_entry(next_event, event_entry);
}
}
static void perf_ctx_sched_task_cb(struct perf_event_context *ctx,
struct task_struct *task, bool sched_in)
{
struct perf_event_pmu_context *pmu_ctx;
struct perf_cpu_pmu_context *cpc;
list_for_each_entry(pmu_ctx, &ctx->pmu_ctx_list, pmu_ctx_entry) {
cpc = this_cpc(pmu_ctx->pmu);
if (cpc->sched_cb_usage && pmu_ctx->pmu->sched_task)
pmu_ctx->pmu->sched_task(pmu_ctx, task, sched_in);
}
}
static void
perf_event_context_sched_out(struct task_struct *task, struct task_struct *next)
{
struct perf_event_context *ctx = task->perf_event_ctxp;
struct perf_event_context *next_ctx;
struct perf_event_context *parent, *next_parent;
int do_switch = 1;
if (likely(!ctx))
return;
rcu_read_lock();
next_ctx = rcu_dereference(next->perf_event_ctxp);
if (!next_ctx)
goto unlock;
parent = rcu_dereference(ctx->parent_ctx);
next_parent = rcu_dereference(next_ctx->parent_ctx);
/* If neither context have a parent context; they cannot be clones. */
if (!parent && !next_parent)
goto unlock;
if (next_parent == ctx || next_ctx == parent || next_parent == parent) {
/*
* Looks like the two contexts are clones, so we might be
* able to optimize the context switch. We lock both
* contexts and check that they are clones under the
* lock (including re-checking that neither has been
* uncloned in the meantime). It doesn't matter which
* order we take the locks because no other cpu could
* be trying to lock both of these tasks.
*/
raw_spin_lock(&ctx->lock);
raw_spin_lock_nested(&next_ctx->lock, SINGLE_DEPTH_NESTING);
if (context_equiv(ctx, next_ctx)) {
perf_ctx_disable(ctx, false);
/* PMIs are disabled; ctx->nr_no_switch_fast is stable. */
if (local_read(&ctx->nr_no_switch_fast) ||
local_read(&next_ctx->nr_no_switch_fast)) {
/*
* Must not swap out ctx when there's pending
* events that rely on the ctx->task relation.
*
* Likewise, when a context contains inherit +
* SAMPLE_READ events they should be switched
* out using the slow path so that they are
* treated as if they were distinct contexts.
*/
raw_spin_unlock(&next_ctx->lock);
rcu_read_unlock();
goto inside_switch;
}
WRITE_ONCE(ctx->task, next);
WRITE_ONCE(next_ctx->task, task);
perf_ctx_sched_task_cb(ctx, task, false);
perf_ctx_enable(ctx, false);
/*
* RCU_INIT_POINTER here is safe because we've not
* modified the ctx and the above modification of
* ctx->task is immaterial since this value is
* always verified under ctx->lock which we're now
* holding.
*/
RCU_INIT_POINTER(task->perf_event_ctxp, next_ctx);
RCU_INIT_POINTER(next->perf_event_ctxp, ctx);
do_switch = 0;
perf_event_sync_stat(ctx, next_ctx);
}
raw_spin_unlock(&next_ctx->lock);
raw_spin_unlock(&ctx->lock);
}
unlock:
rcu_read_unlock();
if (do_switch) {
raw_spin_lock(&ctx->lock);
perf_ctx_disable(ctx, false);
inside_switch:
perf_ctx_sched_task_cb(ctx, task, false);
task_ctx_sched_out(ctx, NULL, EVENT_ALL);
perf_ctx_enable(ctx, false);
raw_spin_unlock(&ctx->lock);
}
}
static DEFINE_PER_CPU(struct list_head, sched_cb_list);
static DEFINE_PER_CPU(int, perf_sched_cb_usages);
void perf_sched_cb_dec(struct pmu *pmu)
{
struct perf_cpu_pmu_context *cpc = this_cpc(pmu);
this_cpu_dec(perf_sched_cb_usages);
barrier();
if (!--cpc->sched_cb_usage)
list_del(&cpc->sched_cb_entry);
}
void perf_sched_cb_inc(struct pmu *pmu)
{
struct perf_cpu_pmu_context *cpc = this_cpc(pmu);
if (!cpc->sched_cb_usage++)
list_add(&cpc->sched_cb_entry, this_cpu_ptr(&sched_cb_list));
barrier();
this_cpu_inc(perf_sched_cb_usages);
}
/*
* This function provides the context switch callback to the lower code
* layer. It is invoked ONLY when the context switch callback is enabled.
*
* This callback is relevant even to per-cpu events; for example multi event
* PEBS requires this to provide PID/TID information. This requires we flush
* all queued PEBS records before we context switch to a new task.
*/
static void __perf_pmu_sched_task(struct perf_cpu_pmu_context *cpc,
struct task_struct *task, bool sched_in)
{
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
struct pmu *pmu;
pmu = cpc->epc.pmu;
/* software PMUs will not have sched_task */
if (WARN_ON_ONCE(!pmu->sched_task))
return;
perf_ctx_lock(cpuctx, cpuctx->task_ctx);
perf_pmu_disable(pmu);
pmu->sched_task(cpc->task_epc, task, sched_in);
perf_pmu_enable(pmu);
perf_ctx_unlock(cpuctx, cpuctx->task_ctx);
}
static void perf_pmu_sched_task(struct task_struct *prev,
struct task_struct *next,
bool sched_in)
{
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
struct perf_cpu_pmu_context *cpc;
/* cpuctx->task_ctx will be handled in perf_event_context_sched_in/out */
if (prev == next || cpuctx->task_ctx)
return;
list_for_each_entry(cpc, this_cpu_ptr(&sched_cb_list), sched_cb_entry)
__perf_pmu_sched_task(cpc, sched_in ? next : prev, sched_in);
}
static void perf_event_switch(struct task_struct *task,
struct task_struct *next_prev, bool sched_in);
/*
* Called from scheduler to remove the events of the current task,
* with interrupts disabled.
*
* We stop each event and update the event value in event->count.
*
* This does not protect us against NMI, but disable()
* sets the disabled bit in the control field of event _before_
* accessing the event control register. If a NMI hits, then it will
* not restart the event.
*/
void __perf_event_task_sched_out(struct task_struct *task,
struct task_struct *next)
{
if (__this_cpu_read(perf_sched_cb_usages))
perf_pmu_sched_task(task, next, false);
if (atomic_read(&nr_switch_events))
perf_event_switch(task, next, false);
perf_event_context_sched_out(task, next);
/*
* if cgroup events exist on this CPU, then we need
* to check if we have to switch out PMU state.
* cgroup event are system-wide mode only
*/
perf_cgroup_switch(next);
}
static bool perf_less_group_idx(const void *l, const void *r, void __always_unused *args)
{
const struct perf_event *le = *(const struct perf_event **)l;
const struct perf_event *re = *(const struct perf_event **)r;
return le->group_index < re->group_index;
}
DEFINE_MIN_HEAP(struct perf_event *, perf_event_min_heap);
static const struct min_heap_callbacks perf_min_heap = {
.less = perf_less_group_idx,
.swp = NULL,
};
static void __heap_add(struct perf_event_min_heap *heap, struct perf_event *event)
{
struct perf_event **itrs = heap->data;
if (event) {
itrs[heap->nr] = event;
heap->nr++;
}
}
static void __link_epc(struct perf_event_pmu_context *pmu_ctx)
{
struct perf_cpu_pmu_context *cpc;
if (!pmu_ctx->ctx->task)
return;
cpc = this_cpc(pmu_ctx->pmu);
WARN_ON_ONCE(cpc->task_epc && cpc->task_epc != pmu_ctx);
cpc->task_epc = pmu_ctx;
}
static noinline int visit_groups_merge(struct perf_event_context *ctx,
struct perf_event_groups *groups, int cpu,
struct pmu *pmu,
int (*func)(struct perf_event *, void *),
void *data)
{
#ifdef CONFIG_CGROUP_PERF
struct cgroup_subsys_state *css = NULL;
#endif
struct perf_cpu_context *cpuctx = NULL;
/* Space for per CPU and/or any CPU event iterators. */
struct perf_event *itrs[2];
struct perf_event_min_heap event_heap;
struct perf_event **evt;
int ret;
if (pmu->filter && pmu->filter(pmu, cpu))
return 0;
if (!ctx->task) {
cpuctx = this_cpu_ptr(&perf_cpu_context);
event_heap = (struct perf_event_min_heap){
.data = cpuctx->heap,
.nr = 0,
.size = cpuctx->heap_size,
};
lockdep_assert_held(&cpuctx->ctx.lock);
#ifdef CONFIG_CGROUP_PERF
if (cpuctx->cgrp)
css = &cpuctx->cgrp->css;
#endif
} else {
event_heap = (struct perf_event_min_heap){
.data = itrs,
.nr = 0,
.size = ARRAY_SIZE(itrs),
};
/* Events not within a CPU context may be on any CPU. */
__heap_add(&event_heap, perf_event_groups_first(groups, -1, pmu, NULL));
}
evt = event_heap.data;
__heap_add(&event_heap, perf_event_groups_first(groups, cpu, pmu, NULL));
#ifdef CONFIG_CGROUP_PERF
for (; css; css = css->parent)
__heap_add(&event_heap, perf_event_groups_first(groups, cpu, pmu, css->cgroup));
#endif
if (event_heap.nr) {
__link_epc((*evt)->pmu_ctx);
perf_assert_pmu_disabled((*evt)->pmu_ctx->pmu);
}
min_heapify_all_inline(&event_heap, &perf_min_heap, NULL);
while (event_heap.nr) {
ret = func(*evt, data);
if (ret)
return ret;
*evt = perf_event_groups_next(*evt, pmu);
if (*evt)
min_heap_sift_down_inline(&event_heap, 0, &perf_min_heap, NULL);
else
min_heap_pop_inline(&event_heap, &perf_min_heap, NULL);
}
return 0;
}
/*
* Because the userpage is strictly per-event (there is no concept of context,
* so there cannot be a context indirection), every userpage must be updated
* when context time starts :-(
*
* IOW, we must not miss EVENT_TIME edges.
*/
static inline bool event_update_userpage(struct perf_event *event)
{
if (likely(!refcount_read(&event->mmap_count)))
return false;
perf_event_update_time(event);
perf_event_update_userpage(event);
return true;
}
static inline void group_update_userpage(struct perf_event *group_event)
{
struct perf_event *event;
if (!event_update_userpage(group_event))
return;
for_each_sibling_event(event, group_event)
event_update_userpage(event);
}
static int merge_sched_in(struct perf_event *event, void *data)
{
struct perf_event_context *ctx = event->ctx;
int *can_add_hw = data;
if (event->state <= PERF_EVENT_STATE_OFF)
return 0;
if (!event_filter_match(event))
return 0;
if (group_can_go_on(event, *can_add_hw)) {
if (!group_sched_in(event, ctx))
list_add_tail(&event->active_list, get_event_list(event));
}
if (event->state == PERF_EVENT_STATE_INACTIVE) {
*can_add_hw = 0;
if (event->attr.pinned) {
perf_cgroup_event_disable(event, ctx);
perf_event_set_state(event, PERF_EVENT_STATE_ERROR);
if (*perf_event_fasync(event))
event->pending_kill = POLL_ERR;
perf_event_wakeup(event);
} else {
struct perf_cpu_pmu_context *cpc = this_cpc(event->pmu_ctx->pmu);
event->pmu_ctx->rotate_necessary = 1;
perf_mux_hrtimer_restart(cpc);
group_update_userpage(event);
}
}
return 0;
}
static void pmu_groups_sched_in(struct perf_event_context *ctx,
struct perf_event_groups *groups,
struct pmu *pmu)
{
int can_add_hw = 1;
visit_groups_merge(ctx, groups, smp_processor_id(), pmu,
merge_sched_in, &can_add_hw);
}
static void __pmu_ctx_sched_in(struct perf_event_pmu_context *pmu_ctx,
enum event_type_t event_type)
{
struct perf_event_context *ctx = pmu_ctx->ctx;
if (event_type & EVENT_PINNED)
pmu_groups_sched_in(ctx, &ctx->pinned_groups, pmu_ctx->pmu);
if (event_type & EVENT_FLEXIBLE)
pmu_groups_sched_in(ctx, &ctx->flexible_groups, pmu_ctx->pmu);
}
static void
ctx_sched_in(struct perf_event_context *ctx, struct pmu *pmu, enum event_type_t event_type)
{
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
struct perf_event_pmu_context *pmu_ctx;
int is_active = ctx->is_active;
bool cgroup = event_type & EVENT_CGROUP;
event_type &= ~EVENT_CGROUP;
lockdep_assert_held(&ctx->lock);
if (likely(!ctx->nr_events))
return;
if (!(is_active & EVENT_TIME)) {
/* start ctx time */
__update_context_time(ctx, false);
perf_cgroup_set_timestamp(cpuctx);
/*
* CPU-release for the below ->is_active store,
* see __load_acquire() in perf_event_time_now()
*/
barrier();
}
ctx->is_active |= (event_type | EVENT_TIME);
if (ctx->task) {
if (!(is_active & EVENT_ALL))
cpuctx->task_ctx = ctx;
else
WARN_ON_ONCE(cpuctx->task_ctx != ctx);
}
is_active ^= ctx->is_active; /* changed bits */
/*
* First go through the list and put on any pinned groups
* in order to give them the best chance of going on.
*/
if (is_active & EVENT_PINNED) {
for_each_epc(pmu_ctx, ctx, pmu, cgroup)
__pmu_ctx_sched_in(pmu_ctx, EVENT_PINNED);
}
/* Then walk through the lower prio flexible groups */
if (is_active & EVENT_FLEXIBLE) {
for_each_epc(pmu_ctx, ctx, pmu, cgroup)
__pmu_ctx_sched_in(pmu_ctx, EVENT_FLEXIBLE);
}
}
static void perf_event_context_sched_in(struct task_struct *task)
{
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
struct perf_event_context *ctx;
rcu_read_lock();
ctx = rcu_dereference(task->perf_event_ctxp);
if (!ctx)
goto rcu_unlock;
if (cpuctx->task_ctx == ctx) {
perf_ctx_lock(cpuctx, ctx);
perf_ctx_disable(ctx, false);
perf_ctx_sched_task_cb(ctx, task, true);
perf_ctx_enable(ctx, false);
perf_ctx_unlock(cpuctx, ctx);
goto rcu_unlock;
}
perf_ctx_lock(cpuctx, ctx);
/*
* We must check ctx->nr_events while holding ctx->lock, such
* that we serialize against perf_install_in_context().
*/
if (!ctx->nr_events)
goto unlock;
perf_ctx_disable(ctx, false);
/*
* We want to keep the following priority order:
* cpu pinned (that don't need to move), task pinned,
* cpu flexible, task flexible.
*
* However, if task's ctx is not carrying any pinned
* events, no need to flip the cpuctx's events around.
*/
if (!RB_EMPTY_ROOT(&ctx->pinned_groups.tree)) {
perf_ctx_disable(&cpuctx->ctx, false);
ctx_sched_out(&cpuctx->ctx, NULL, EVENT_FLEXIBLE);
}
perf_event_sched_in(cpuctx, ctx, NULL);
perf_ctx_sched_task_cb(cpuctx->task_ctx, task, true);
if (!RB_EMPTY_ROOT(&ctx->pinned_groups.tree))
perf_ctx_enable(&cpuctx->ctx, false);
perf_ctx_enable(ctx, false);
unlock:
perf_ctx_unlock(cpuctx, ctx);
rcu_unlock:
rcu_read_unlock();
}
/*
* Called from scheduler to add the events of the current task
* with interrupts disabled.
*
* We restore the event value and then enable it.
*
* This does not protect us against NMI, but enable()
* sets the enabled bit in the control field of event _before_
* accessing the event control register. If a NMI hits, then it will
* keep the event running.
*/
void __perf_event_task_sched_in(struct task_struct *prev,
struct task_struct *task)
{
perf_event_context_sched_in(task);
if (atomic_read(&nr_switch_events))
perf_event_switch(task, prev, true);
if (__this_cpu_read(perf_sched_cb_usages))
perf_pmu_sched_task(prev, task, true);
}
static u64 perf_calculate_period(struct perf_event *event, u64 nsec, u64 count)
{
u64 frequency = event->attr.sample_freq;
u64 sec = NSEC_PER_SEC;
u64 divisor, dividend;
int count_fls, nsec_fls, frequency_fls, sec_fls;
count_fls = fls64(count);
nsec_fls = fls64(nsec);
frequency_fls = fls64(frequency);
sec_fls = 30;
/*
* We got @count in @nsec, with a target of sample_freq HZ
* the target period becomes:
*
* @count * 10^9
* period = -------------------
* @nsec * sample_freq
*
*/
/*
* Reduce accuracy by one bit such that @a and @b converge
* to a similar magnitude.
*/
#define REDUCE_FLS(a, b) \
do { \
if (a##_fls > b##_fls) { \
a >>= 1; \
a##_fls--; \
} else { \
b >>= 1; \
b##_fls--; \
} \
} while (0)
/*
* Reduce accuracy until either term fits in a u64, then proceed with
* the other, so that finally we can do a u64/u64 division.
*/
while (count_fls + sec_fls > 64 && nsec_fls + frequency_fls > 64) {
REDUCE_FLS(nsec, frequency);
REDUCE_FLS(sec, count);
}
if (count_fls + sec_fls > 64) {
divisor = nsec * frequency;
while (count_fls + sec_fls > 64) {
REDUCE_FLS(count, sec);
divisor >>= 1;
}
dividend = count * sec;
} else {
dividend = count * sec;
while (nsec_fls + frequency_fls > 64) {
REDUCE_FLS(nsec, frequency);
dividend >>= 1;
}
divisor = nsec * frequency;
}
if (!divisor)
return dividend;
return div64_u64(dividend, divisor);
}
static DEFINE_PER_CPU(int, perf_throttled_count);
static DEFINE_PER_CPU(u64, perf_throttled_seq);
static void perf_adjust_period(struct perf_event *event, u64 nsec, u64 count, bool disable)
{
struct hw_perf_event *hwc = &event->hw;
s64 period, sample_period;
s64 delta;
period = perf_calculate_period(event, nsec, count);
delta = (s64)(period - hwc->sample_period);
if (delta >= 0)
delta += 7;
else
delta -= 7;
delta /= 8; /* low pass filter */
sample_period = hwc->sample_period + delta;
if (!sample_period)
sample_period = 1;
hwc->sample_period = sample_period;
if (local64_read(&hwc->period_left) > 8*sample_period) {
if (disable)
event->pmu->stop(event, PERF_EF_UPDATE);
local64_set(&hwc->period_left, 0);
if (disable)
event->pmu->start(event, PERF_EF_RELOAD);
}
}
static void perf_adjust_freq_unthr_events(struct list_head *event_list)
{
struct perf_event *event;
struct hw_perf_event *hwc;
u64 now, period = TICK_NSEC;
s64 delta;
list_for_each_entry(event, event_list, active_list) {
if (event->state != PERF_EVENT_STATE_ACTIVE)
continue;
// XXX use visit thingy to avoid the -1,cpu match
if (!event_filter_match(event))
continue;
hwc = &event->hw;
if (hwc->interrupts == MAX_INTERRUPTS)
perf_event_unthrottle_group(event, is_event_in_freq_mode(event));
if (!is_event_in_freq_mode(event))
continue;
/*
* stop the event and update event->count
*/
event->pmu->stop(event, PERF_EF_UPDATE);
now = local64_read(&event->count);
delta = now - hwc->freq_count_stamp;
hwc->freq_count_stamp = now;
/*
* restart the event
* reload only if value has changed
* we have stopped the event so tell that
* to perf_adjust_period() to avoid stopping it
* twice.
*/
if (delta > 0)
perf_adjust_period(event, period, delta, false);
event->pmu->start(event, delta > 0 ? PERF_EF_RELOAD : 0);
}
}
/*
* combine freq adjustment with unthrottling to avoid two passes over the
* events. At the same time, make sure, having freq events does not change
* the rate of unthrottling as that would introduce bias.
*/
static void
perf_adjust_freq_unthr_context(struct perf_event_context *ctx, bool unthrottle)
{
struct perf_event_pmu_context *pmu_ctx;
/*
* only need to iterate over all events iff:
* - context have events in frequency mode (needs freq adjust)
* - there are events to unthrottle on this cpu
*/
if (!(ctx->nr_freq || unthrottle))
return;
raw_spin_lock(&ctx->lock);
list_for_each_entry(pmu_ctx, &ctx->pmu_ctx_list, pmu_ctx_entry) {
if (!(pmu_ctx->nr_freq || unthrottle))
continue;
if (!perf_pmu_ctx_is_active(pmu_ctx))
continue;
if (pmu_ctx->pmu->capabilities & PERF_PMU_CAP_NO_INTERRUPT)
continue;
perf_pmu_disable(pmu_ctx->pmu);
perf_adjust_freq_unthr_events(&pmu_ctx->pinned_active);
perf_adjust_freq_unthr_events(&pmu_ctx->flexible_active);
perf_pmu_enable(pmu_ctx->pmu);
}
raw_spin_unlock(&ctx->lock);
}
/*
* Move @event to the tail of the @ctx's elegible events.
*/
static void rotate_ctx(struct perf_event_context *ctx, struct perf_event *event)
{
/*
* Rotate the first entry last of non-pinned groups. Rotation might be
* disabled by the inheritance code.
*/
if (ctx->rotate_disable)
return;
perf_event_groups_delete(&ctx->flexible_groups, event);
perf_event_groups_insert(&ctx->flexible_groups, event);
}
/* pick an event from the flexible_groups to rotate */
static inline struct perf_event *
ctx_event_to_rotate(struct perf_event_pmu_context *pmu_ctx)
{
struct perf_event *event;
struct rb_node *node;
struct rb_root *tree;
struct __group_key key = {
.pmu = pmu_ctx->pmu,
};
/* pick the first active flexible event */
event = list_first_entry_or_null(&pmu_ctx->flexible_active,
struct perf_event, active_list);
if (event)
goto out;
/* if no active flexible event, pick the first event */
tree = &pmu_ctx->ctx->flexible_groups.tree;
if (!pmu_ctx->ctx->task) {
key.cpu = smp_processor_id();
node = rb_find_first(&key, tree, __group_cmp_ignore_cgroup);
if (node)
event = __node_2_pe(node);
goto out;
}
key.cpu = -1;
node = rb_find_first(&key, tree, __group_cmp_ignore_cgroup);
if (node) {
event = __node_2_pe(node);
goto out;
}
key.cpu = smp_processor_id();
node = rb_find_first(&key, tree, __group_cmp_ignore_cgroup);
if (node)
event = __node_2_pe(node);
out:
/*
* Unconditionally clear rotate_necessary; if ctx_flexible_sched_in()
* finds there are unschedulable events, it will set it again.
*/
pmu_ctx->rotate_necessary = 0;
return event;
}
static bool perf_rotate_context(struct perf_cpu_pmu_context *cpc)
{
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
struct perf_event_pmu_context *cpu_epc, *task_epc = NULL;
struct perf_event *cpu_event = NULL, *task_event = NULL;
int cpu_rotate, task_rotate;
struct pmu *pmu;
/*
* Since we run this from IRQ context, nobody can install new
* events, thus the event count values are stable.
*/
cpu_epc = &cpc->epc;
pmu = cpu_epc->pmu;
task_epc = cpc->task_epc;
cpu_rotate = cpu_epc->rotate_necessary;
task_rotate = task_epc ? task_epc->rotate_necessary : 0;
if (!(cpu_rotate || task_rotate))
return false;
perf_ctx_lock(cpuctx, cpuctx->task_ctx);
perf_pmu_disable(pmu);
if (task_rotate)
task_event = ctx_event_to_rotate(task_epc);
if (cpu_rotate)
cpu_event = ctx_event_to_rotate(cpu_epc);
/*
* As per the order given at ctx_resched() first 'pop' task flexible
* and then, if needed CPU flexible.
*/
if (task_event || (task_epc && cpu_event)) {
update_context_time(task_epc->ctx);
__pmu_ctx_sched_out(task_epc, EVENT_FLEXIBLE);
}
if (cpu_event) {
update_context_time(&cpuctx->ctx);
__pmu_ctx_sched_out(cpu_epc, EVENT_FLEXIBLE);
rotate_ctx(&cpuctx->ctx, cpu_event);
__pmu_ctx_sched_in(cpu_epc, EVENT_FLEXIBLE);
}
if (task_event)
rotate_ctx(task_epc->ctx, task_event);
if (task_event || (task_epc && cpu_event))
__pmu_ctx_sched_in(task_epc, EVENT_FLEXIBLE);
perf_pmu_enable(pmu);
perf_ctx_unlock(cpuctx, cpuctx->task_ctx);
return true;
}
void perf_event_task_tick(void)
{
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
struct perf_event_context *ctx;
int throttled;
lockdep_assert_irqs_disabled();
__this_cpu_inc(perf_throttled_seq);
throttled = __this_cpu_xchg(perf_throttled_count, 0);
tick_dep_clear_cpu(smp_processor_id(), TICK_DEP_BIT_PERF_EVENTS);
perf_adjust_freq_unthr_context(&cpuctx->ctx, !!throttled);
rcu_read_lock();
ctx = rcu_dereference(current->perf_event_ctxp);
if (ctx)
perf_adjust_freq_unthr_context(ctx, !!throttled);
rcu_read_unlock();
}
static int event_enable_on_exec(struct perf_event *event,
struct perf_event_context *ctx)
{
if (!event->attr.enable_on_exec)
return 0;
event->attr.enable_on_exec = 0;
if (event->state >= PERF_EVENT_STATE_INACTIVE)
return 0;
perf_event_set_state(event, PERF_EVENT_STATE_INACTIVE);
return 1;
}
/*
* Enable all of a task's events that have been marked enable-on-exec.
* This expects task == current.
*/
static void perf_event_enable_on_exec(struct perf_event_context *ctx)
{
struct perf_event_context *clone_ctx = NULL;
enum event_type_t event_type = 0;
struct perf_cpu_context *cpuctx;
struct perf_event *event;
unsigned long flags;
int enabled = 0;
local_irq_save(flags);
if (WARN_ON_ONCE(current->perf_event_ctxp != ctx))
goto out;
if (!ctx->nr_events)
goto out;
cpuctx = this_cpu_ptr(&perf_cpu_context);
perf_ctx_lock(cpuctx, ctx);
ctx_time_freeze(cpuctx, ctx);
list_for_each_entry(event, &ctx->event_list, event_entry) {
enabled |= event_enable_on_exec(event, ctx);
event_type |= get_event_type(event);
}
/*
* Unclone and reschedule this context if we enabled any event.
*/
if (enabled) {
clone_ctx = unclone_ctx(ctx);
ctx_resched(cpuctx, ctx, NULL, event_type);
}
perf_ctx_unlock(cpuctx, ctx);
out:
local_irq_restore(flags);
if (clone_ctx)
put_ctx(clone_ctx);
}
static void perf_remove_from_owner(struct perf_event *event);
static void perf_event_exit_event(struct perf_event *event,
struct perf_event_context *ctx,
struct task_struct *task,
bool revoke);
/*
* Removes all events from the current task that have been marked
* remove-on-exec, and feeds their values back to parent events.
*/
static void perf_event_remove_on_exec(struct perf_event_context *ctx)
{
struct perf_event_context *clone_ctx = NULL;
struct perf_event *event, *next;
unsigned long flags;
bool modified = false;
mutex_lock(&ctx->mutex);
if (WARN_ON_ONCE(ctx->task != current))
goto unlock;
list_for_each_entry_safe(event, next, &ctx->event_list, event_entry) {
if (!event->attr.remove_on_exec)
continue;
if (!is_kernel_event(event))
perf_remove_from_owner(event);
modified = true;
perf_event_exit_event(event, ctx, ctx->task, false);
}
raw_spin_lock_irqsave(&ctx->lock, flags);
if (modified)
clone_ctx = unclone_ctx(ctx);
raw_spin_unlock_irqrestore(&ctx->lock, flags);
unlock:
mutex_unlock(&ctx->mutex);
if (clone_ctx)
put_ctx(clone_ctx);
}
struct perf_read_data {
struct perf_event *event;
bool group;
int ret;
};
static inline const struct cpumask *perf_scope_cpu_topology_cpumask(unsigned int scope, int cpu);
static int __perf_event_read_cpu(struct perf_event *event, int event_cpu)
{
int local_cpu = smp_processor_id();
u16 local_pkg, event_pkg;
if ((unsigned)event_cpu >= nr_cpu_ids)
return event_cpu;
if (event->group_caps & PERF_EV_CAP_READ_SCOPE) {
const struct cpumask *cpumask = perf_scope_cpu_topology_cpumask(event->pmu->scope, event_cpu);
if (cpumask && cpumask_test_cpu(local_cpu, cpumask))
return local_cpu;
}
if (event->group_caps & PERF_EV_CAP_READ_ACTIVE_PKG) {
event_pkg = topology_physical_package_id(event_cpu);
local_pkg = topology_physical_package_id(local_cpu);
if (event_pkg == local_pkg)
return local_cpu;
}
return event_cpu;
}
/*
* Cross CPU call to read the hardware event
*/
static void __perf_event_read(void *info)
{
struct perf_read_data *data = info;
struct perf_event *sub, *event = data->event;
struct perf_event_context *ctx = event->ctx;
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
struct pmu *pmu = event->pmu;
/*
* If this is a task context, we need to check whether it is
* the current task context of this cpu. If not it has been
* scheduled out before the smp call arrived. In that case
* event->count would have been updated to a recent sample
* when the event was scheduled out.
*/
if (ctx->task && cpuctx->task_ctx != ctx)
return;
raw_spin_lock(&ctx->lock);
ctx_time_update_event(ctx, event);
perf_event_update_time(event);
if (data->group)
perf_event_update_sibling_time(event);
if (event->state != PERF_EVENT_STATE_ACTIVE)
goto unlock;
if (!data->group) {
pmu->read(event);
data->ret = 0;
goto unlock;
}
pmu->start_txn(pmu, PERF_PMU_TXN_READ);
pmu->read(event);
for_each_sibling_event(sub, event)
perf_pmu_read(sub);
data->ret = pmu->commit_txn(pmu);
unlock:
raw_spin_unlock(&ctx->lock);
}
static inline u64 perf_event_count(struct perf_event *event, bool self)
{
if (self)
return local64_read(&event->count);
return local64_read(&event->count) + atomic64_read(&event->child_count);
}
static void calc_timer_values(struct perf_event *event,
u64 *now,
u64 *enabled,
u64 *running)
{
u64 ctx_time;
*now = perf_clock();
ctx_time = perf_event_time_now(event, *now);
__perf_update_times(event, ctx_time, enabled, running);
}
/*
* NMI-safe method to read a local event, that is an event that
* is:
* - either for the current task, or for this CPU
* - does not have inherit set, for inherited task events
* will not be local and we cannot read them atomically
* - must not have a pmu::count method
*/
int perf_event_read_local(struct perf_event *event, u64 *value,
u64 *enabled, u64 *running)
{
unsigned long flags;
int event_oncpu;
int event_cpu;
int ret = 0;
/*
* Disabling interrupts avoids all counter scheduling (context
* switches, timer based rotation and IPIs).
*/
local_irq_save(flags);
/*
* It must not be an event with inherit set, we cannot read
* all child counters from atomic context.
*/
if (event->attr.inherit) {
ret = -EOPNOTSUPP;
goto out;
}
/* If this is a per-task event, it must be for current */
if ((event->attach_state & PERF_ATTACH_TASK) &&
event->hw.target != current) {
ret = -EINVAL;
goto out;
}
/*
* Get the event CPU numbers, and adjust them to local if the event is
* a per-package event that can be read locally
*/
event_oncpu = __perf_event_read_cpu(event, event->oncpu);
event_cpu = __perf_event_read_cpu(event, event->cpu);
/* If this is a per-CPU event, it must be for this CPU */
if (!(event->attach_state & PERF_ATTACH_TASK) &&
event_cpu != smp_processor_id()) {
ret = -EINVAL;
goto out;
}
/* If this is a pinned event it must be running on this CPU */
if (event->attr.pinned && event_oncpu != smp_processor_id()) {
ret = -EBUSY;
goto out;
}
/*
* If the event is currently on this CPU, its either a per-task event,
* or local to this CPU. Furthermore it means its ACTIVE (otherwise
* oncpu == -1).
*/
if (event_oncpu == smp_processor_id())
event->pmu->read(event);
*value = local64_read(&event->count);
if (enabled || running) {
u64 __enabled, __running, __now;
calc_timer_values(event, &__now, &__enabled, &__running);
if (enabled)
*enabled = __enabled;
if (running)
*running = __running;
}
out:
local_irq_restore(flags);
return ret;
}
static int perf_event_read(struct perf_event *event, bool group)
{
enum perf_event_state state = READ_ONCE(event->state);
int event_cpu, ret = 0;
/*
* If event is enabled and currently active on a CPU, update the
* value in the event structure:
*/
again:
if (state == PERF_EVENT_STATE_ACTIVE) {
struct perf_read_data data;
/*
* Orders the ->state and ->oncpu loads such that if we see
* ACTIVE we must also see the right ->oncpu.
*
* Matches the smp_wmb() from event_sched_in().
*/
smp_rmb();
event_cpu = READ_ONCE(event->oncpu);
if ((unsigned)event_cpu >= nr_cpu_ids)
return 0;
data = (struct perf_read_data){
.event = event,
.group = group,
.ret = 0,
};
preempt_disable();
event_cpu = __perf_event_read_cpu(event, event_cpu);
/*
* Purposely ignore the smp_call_function_single() return
* value.
*
* If event_cpu isn't a valid CPU it means the event got
* scheduled out and that will have updated the event count.
*
* Therefore, either way, we'll have an up-to-date event count
* after this.
*/
(void)smp_call_function_single(event_cpu, __perf_event_read, &data, 1);
preempt_enable();
ret = data.ret;
} else if (state == PERF_EVENT_STATE_INACTIVE) {
struct perf_event_context *ctx = event->ctx;
unsigned long flags;
raw_spin_lock_irqsave(&ctx->lock, flags);
state = event->state;
if (state != PERF_EVENT_STATE_INACTIVE) {
raw_spin_unlock_irqrestore(&ctx->lock, flags);
goto again;
}
/*
* May read while context is not active (e.g., thread is
* blocked), in that case we cannot update context time
*/
ctx_time_update_event(ctx, event);
perf_event_update_time(event);
if (group)
perf_event_update_sibling_time(event);
raw_spin_unlock_irqrestore(&ctx->lock, flags);
}
return ret;
}
/*
* Initialize the perf_event context in a task_struct:
*/
static void __perf_event_init_context(struct perf_event_context *ctx)
{
raw_spin_lock_init(&ctx->lock);
mutex_init(&ctx->mutex);
INIT_LIST_HEAD(&ctx->pmu_ctx_list);
perf_event_groups_init(&ctx->pinned_groups);
perf_event_groups_init(&ctx->flexible_groups);
INIT_LIST_HEAD(&ctx->event_list);
refcount_set(&ctx->refcount, 1);
}
static void
__perf_init_event_pmu_context(struct perf_event_pmu_context *epc, struct pmu *pmu)
{
epc->pmu = pmu;
INIT_LIST_HEAD(&epc->pmu_ctx_entry);
INIT_LIST_HEAD(&epc->pinned_active);
INIT_LIST_HEAD(&epc->flexible_active);
atomic_set(&epc->refcount, 1);
}
static struct perf_event_context *
alloc_perf_context(struct task_struct *task)
{
struct perf_event_context *ctx;
ctx = kzalloc(sizeof(struct perf_event_context), GFP_KERNEL);
if (!ctx)
return NULL;
__perf_event_init_context(ctx);
if (task)
ctx->task = get_task_struct(task);
return ctx;
}
static struct task_struct *
find_lively_task_by_vpid(pid_t vpid)
{
struct task_struct *task;
rcu_read_lock();
if (!vpid)
task = current;
else
task = find_task_by_vpid(vpid);
if (task)
get_task_struct(task);
rcu_read_unlock();
if (!task)
return ERR_PTR(-ESRCH);
return task;
}
/*
* Returns a matching context with refcount and pincount.
*/
static struct perf_event_context *
find_get_context(struct task_struct *task, struct perf_event *event)
{
struct perf_event_context *ctx, *clone_ctx = NULL;
struct perf_cpu_context *cpuctx;
unsigned long flags;
int err;
if (!task) {
/* Must be root to operate on a CPU event: */
err = perf_allow_cpu();
if (err)
return ERR_PTR(err);
cpuctx = per_cpu_ptr(&perf_cpu_context, event->cpu);
ctx = &cpuctx->ctx;
get_ctx(ctx);
raw_spin_lock_irqsave(&ctx->lock, flags);
++ctx->pin_count;
raw_spin_unlock_irqrestore(&ctx->lock, flags);
return ctx;
}
err = -EINVAL;
retry:
ctx = perf_lock_task_context(task, &flags);
if (ctx) {
clone_ctx = unclone_ctx(ctx);
++ctx->pin_count;
raw_spin_unlock_irqrestore(&ctx->lock, flags);
if (clone_ctx)
put_ctx(clone_ctx);
} else {
ctx = alloc_perf_context(task);
err = -ENOMEM;
if (!ctx)
goto errout;
err = 0;
mutex_lock(&task->perf_event_mutex);
/*
* If it has already passed perf_event_exit_task().
* we must see PF_EXITING, it takes this mutex too.
*/
if (task->flags & PF_EXITING)
err = -ESRCH;
else if (task->perf_event_ctxp)
err = -EAGAIN;
else {
get_ctx(ctx);
++ctx->pin_count;
rcu_assign_pointer(task->perf_event_ctxp, ctx);
}
mutex_unlock(&task->perf_event_mutex);
if (unlikely(err)) {
put_ctx(ctx);
if (err == -EAGAIN)
goto retry;
goto errout;
}
}
return ctx;
errout:
return ERR_PTR(err);
}
static struct perf_event_pmu_context *
find_get_pmu_context(struct pmu *pmu, struct perf_event_context *ctx,
struct perf_event *event)
{
struct perf_event_pmu_context *new = NULL, *pos = NULL, *epc;
if (!ctx->task) {
/*
* perf_pmu_migrate_context() / __perf_pmu_install_event()
* relies on the fact that find_get_pmu_context() cannot fail
* for CPU contexts.
*/
struct perf_cpu_pmu_context *cpc;
cpc = *per_cpu_ptr(pmu->cpu_pmu_context, event->cpu);
epc = &cpc->epc;
raw_spin_lock_irq(&ctx->lock);
if (!epc->ctx) {
/*
* One extra reference for the pmu; see perf_pmu_free().
*/
atomic_set(&epc->refcount, 2);
epc->embedded = 1;
list_add(&epc->pmu_ctx_entry, &ctx->pmu_ctx_list);
epc->ctx = ctx;
} else {
WARN_ON_ONCE(epc->ctx != ctx);
atomic_inc(&epc->refcount);
}
raw_spin_unlock_irq(&ctx->lock);
return epc;
}
new = kzalloc(sizeof(*epc), GFP_KERNEL);
if (!new)
return ERR_PTR(-ENOMEM);
__perf_init_event_pmu_context(new, pmu);
/*
* XXX
*
* lockdep_assert_held(&ctx->mutex);
*
* can't because perf_event_init_task() doesn't actually hold the
* child_ctx->mutex.
*/
raw_spin_lock_irq(&ctx->lock);
list_for_each_entry(epc, &ctx->pmu_ctx_list, pmu_ctx_entry) {
if (epc->pmu == pmu) {
WARN_ON_ONCE(epc->ctx != ctx);
atomic_inc(&epc->refcount);
goto found_epc;
}
/* Make sure the pmu_ctx_list is sorted by PMU type: */
if (!pos && epc->pmu->type > pmu->type)
pos = epc;
}
epc = new;
new = NULL;
if (!pos)
list_add_tail(&epc->pmu_ctx_entry, &ctx->pmu_ctx_list);
else
list_add(&epc->pmu_ctx_entry, pos->pmu_ctx_entry.prev);
epc->ctx = ctx;
found_epc:
raw_spin_unlock_irq(&ctx->lock);
kfree(new);
return epc;
}
static void get_pmu_ctx(struct perf_event_pmu_context *epc)
{
WARN_ON_ONCE(!atomic_inc_not_zero(&epc->refcount));
}
static void free_cpc_rcu(struct rcu_head *head)
{
struct perf_cpu_pmu_context *cpc =
container_of(head, typeof(*cpc), epc.rcu_head);
kfree(cpc);
}
static void free_epc_rcu(struct rcu_head *head)
{
struct perf_event_pmu_context *epc = container_of(head, typeof(*epc), rcu_head);
kfree(epc);
}
static void put_pmu_ctx(struct perf_event_pmu_context *epc)
{
struct perf_event_context *ctx = epc->ctx;
unsigned long flags;
/*
* XXX
*
* lockdep_assert_held(&ctx->mutex);
*
* can't because of the call-site in _free_event()/put_event()
* which isn't always called under ctx->mutex.
*/
if (!atomic_dec_and_raw_lock_irqsave(&epc->refcount, &ctx->lock, flags))
return;
WARN_ON_ONCE(list_empty(&epc->pmu_ctx_entry));
list_del_init(&epc->pmu_ctx_entry);
epc->ctx = NULL;
WARN_ON_ONCE(!list_empty(&epc->pinned_active));
WARN_ON_ONCE(!list_empty(&epc->flexible_active));
raw_spin_unlock_irqrestore(&ctx->lock, flags);
if (epc->embedded) {
call_rcu(&epc->rcu_head, free_cpc_rcu);
return;
}
call_rcu(&epc->rcu_head, free_epc_rcu);
}
static void perf_event_free_filter(struct perf_event *event);
static void free_event_rcu(struct rcu_head *head)
{
struct perf_event *event = container_of(head, typeof(*event), rcu_head);
if (event->ns)
put_pid_ns(event->ns);
perf_event_free_filter(event);
kmem_cache_free(perf_event_cache, event);
}
static void ring_buffer_attach(struct perf_event *event,
struct perf_buffer *rb);
static void detach_sb_event(struct perf_event *event)
{
struct pmu_event_list *pel = per_cpu_ptr(&pmu_sb_events, event->cpu);
raw_spin_lock(&pel->lock);
list_del_rcu(&event->sb_list);
raw_spin_unlock(&pel->lock);
}
static bool is_sb_event(struct perf_event *event)
{
struct perf_event_attr *attr = &event->attr;
if (event->parent)
return false;
if (event->attach_state & PERF_ATTACH_TASK)
return false;
if (attr->mmap || attr->mmap_data || attr->mmap2 ||
attr->comm || attr->comm_exec ||
attr->task || attr->ksymbol ||
attr->context_switch || attr->text_poke ||
attr->bpf_event)
return true;
return false;
}
static void unaccount_pmu_sb_event(struct perf_event *event)
{
if (is_sb_event(event))
detach_sb_event(event);
}
#ifdef CONFIG_NO_HZ_FULL
static DEFINE_SPINLOCK(nr_freq_lock);
#endif
static void unaccount_freq_event_nohz(void)
{
#ifdef CONFIG_NO_HZ_FULL
spin_lock(&nr_freq_lock);
if (atomic_dec_and_test(&nr_freq_events))
tick_nohz_dep_clear(TICK_DEP_BIT_PERF_EVENTS);
spin_unlock(&nr_freq_lock);
#endif
}
static void unaccount_freq_event(void)
{
if (tick_nohz_full_enabled())
unaccount_freq_event_nohz();
else
atomic_dec(&nr_freq_events);
}
static struct perf_ctx_data *
alloc_perf_ctx_data(struct kmem_cache *ctx_cache, bool global)
{
struct perf_ctx_data *cd;
cd = kzalloc(sizeof(*cd), GFP_KERNEL);
if (!cd)
return NULL;
cd->data = kmem_cache_zalloc(ctx_cache, GFP_KERNEL);
if (!cd->data) {
kfree(cd);
return NULL;
}
cd->global = global;
cd->ctx_cache = ctx_cache;
refcount_set(&cd->refcount, 1);
return cd;
}
static void free_perf_ctx_data(struct perf_ctx_data *cd)
{
kmem_cache_free(cd->ctx_cache, cd->data);
kfree(cd);
}
static void __free_perf_ctx_data_rcu(struct rcu_head *rcu_head)
{
struct perf_ctx_data *cd;
cd = container_of(rcu_head, struct perf_ctx_data, rcu_head);
free_perf_ctx_data(cd);
}
static inline void perf_free_ctx_data_rcu(struct perf_ctx_data *cd)
{
call_rcu(&cd->rcu_head, __free_perf_ctx_data_rcu);
}
static int
attach_task_ctx_data(struct task_struct *task, struct kmem_cache *ctx_cache,
bool global)
{
struct perf_ctx_data *cd, *old = NULL;
cd = alloc_perf_ctx_data(ctx_cache, global);
if (!cd)
return -ENOMEM;
for (;;) {
if (try_cmpxchg((struct perf_ctx_data **)&task->perf_ctx_data, &old, cd)) {
if (old)
perf_free_ctx_data_rcu(old);
return 0;
}
if (!old) {
/*
* After seeing a dead @old, we raced with
* removal and lost, try again to install @cd.
*/
continue;
}
if (refcount_inc_not_zero(&old->refcount)) {
free_perf_ctx_data(cd); /* unused */
return 0;
}
/*
* @old is a dead object, refcount==0 is stable, try and
* replace it with @cd.
*/
}
return 0;
}
static void __detach_global_ctx_data(void);
DEFINE_STATIC_PERCPU_RWSEM(global_ctx_data_rwsem);
static refcount_t global_ctx_data_ref;
static int
attach_global_ctx_data(struct kmem_cache *ctx_cache)
{
struct task_struct *g, *p;
struct perf_ctx_data *cd;
int ret;
if (refcount_inc_not_zero(&global_ctx_data_ref))
return 0;
guard(percpu_write)(&global_ctx_data_rwsem);
if (refcount_inc_not_zero(&global_ctx_data_ref))
return 0;
again:
/* Allocate everything */
scoped_guard (rcu) {
for_each_process_thread(g, p) {
cd = rcu_dereference(p->perf_ctx_data);
if (cd && !cd->global) {
cd->global = 1;
if (!refcount_inc_not_zero(&cd->refcount))
cd = NULL;
}
if (!cd) {
get_task_struct(p);
goto alloc;
}
}
}
refcount_set(&global_ctx_data_ref, 1);
return 0;
alloc:
ret = attach_task_ctx_data(p, ctx_cache, true);
put_task_struct(p);
if (ret) {
__detach_global_ctx_data();
return ret;
}
goto again;
}
static int
attach_perf_ctx_data(struct perf_event *event)
{
struct task_struct *task = event->hw.target;
struct kmem_cache *ctx_cache = event->pmu->task_ctx_cache;
int ret;
if (!ctx_cache)
return -ENOMEM;
if (task)
return attach_task_ctx_data(task, ctx_cache, false);
ret = attach_global_ctx_data(ctx_cache);
if (ret)
return ret;
event->attach_state |= PERF_ATTACH_GLOBAL_DATA;
return 0;
}
static void
detach_task_ctx_data(struct task_struct *p)
{
struct perf_ctx_data *cd;
scoped_guard (rcu) {
cd = rcu_dereference(p->perf_ctx_data);
if (!cd || !refcount_dec_and_test(&cd->refcount))
return;
}
/*
* The old ctx_data may be lost because of the race.
* Nothing is required to do for the case.
* See attach_task_ctx_data().
*/
if (try_cmpxchg((struct perf_ctx_data **)&p->perf_ctx_data, &cd, NULL))
perf_free_ctx_data_rcu(cd);
}
static void __detach_global_ctx_data(void)
{
struct task_struct *g, *p;
struct perf_ctx_data *cd;
again:
scoped_guard (rcu) {
for_each_process_thread(g, p) {
cd = rcu_dereference(p->perf_ctx_data);
if (!cd || !cd->global)
continue;
cd->global = 0;
get_task_struct(p);
goto detach;
}
}
return;
detach:
detach_task_ctx_data(p);
put_task_struct(p);
goto again;
}
static void detach_global_ctx_data(void)
{
if (refcount_dec_not_one(&global_ctx_data_ref))
return;
guard(percpu_write)(&global_ctx_data_rwsem);
if (!refcount_dec_and_test(&global_ctx_data_ref))
return;
/* remove everything */
__detach_global_ctx_data();
}
static void detach_perf_ctx_data(struct perf_event *event)
{
struct task_struct *task = event->hw.target;
event->attach_state &= ~PERF_ATTACH_TASK_DATA;
if (task)
return detach_task_ctx_data(task);
if (event->attach_state & PERF_ATTACH_GLOBAL_DATA) {
detach_global_ctx_data();
event->attach_state &= ~PERF_ATTACH_GLOBAL_DATA;
}
}
static void unaccount_event(struct perf_event *event)
{
bool dec = false;
if (event->parent)
return;
if (event->attach_state & (PERF_ATTACH_TASK | PERF_ATTACH_SCHED_CB))
dec = true;
if (event->attr.mmap || event->attr.mmap_data)
atomic_dec(&nr_mmap_events);
if (event->attr.build_id)
atomic_dec(&nr_build_id_events);
if (event->attr.comm)
atomic_dec(&nr_comm_events);
if (event->attr.namespaces)
atomic_dec(&nr_namespaces_events);
if (event->attr.cgroup)
atomic_dec(&nr_cgroup_events);
if (event->attr.task)
atomic_dec(&nr_task_events);
if (event->attr.freq)
unaccount_freq_event();
if (event->attr.context_switch) {
dec = true;
atomic_dec(&nr_switch_events);
}
if (is_cgroup_event(event))
dec = true;
if (has_branch_stack(event))
dec = true;
if (event->attr.ksymbol)
atomic_dec(&nr_ksymbol_events);
if (event->attr.bpf_event)
atomic_dec(&nr_bpf_events);
if (event->attr.text_poke)
atomic_dec(&nr_text_poke_events);
if (dec) {
if (!atomic_add_unless(&perf_sched_count, -1, 1))
schedule_delayed_work(&perf_sched_work, HZ);
}
unaccount_pmu_sb_event(event);
}
static void perf_sched_delayed(struct work_struct *work)
{
mutex_lock(&perf_sched_mutex);
if (atomic_dec_and_test(&perf_sched_count))
static_branch_disable(&perf_sched_events);
mutex_unlock(&perf_sched_mutex);
}
/*
* The following implement mutual exclusion of events on "exclusive" pmus
* (PERF_PMU_CAP_EXCLUSIVE). Such pmus can only have one event scheduled
* at a time, so we disallow creating events that might conflict, namely:
*
* 1) cpu-wide events in the presence of per-task events,
* 2) per-task events in the presence of cpu-wide events,
* 3) two matching events on the same perf_event_context.
*
* The former two cases are handled in the allocation path (perf_event_alloc(),
* _free_event()), the latter -- before the first perf_install_in_context().
*/
static int exclusive_event_init(struct perf_event *event)
{
struct pmu *pmu = event->pmu;
if (!is_exclusive_pmu(pmu))
return 0;
/*
* Prevent co-existence of per-task and cpu-wide events on the
* same exclusive pmu.
*
* Negative pmu::exclusive_cnt means there are cpu-wide
* events on this "exclusive" pmu, positive means there are
* per-task events.
*
* Since this is called in perf_event_alloc() path, event::ctx
* doesn't exist yet; it is, however, safe to use PERF_ATTACH_TASK
* to mean "per-task event", because unlike other attach states it
* never gets cleared.
*/
if (event->attach_state & PERF_ATTACH_TASK) {
if (!atomic_inc_unless_negative(&pmu->exclusive_cnt))
return -EBUSY;
} else {
if (!atomic_dec_unless_positive(&pmu->exclusive_cnt))
return -EBUSY;
}
event->attach_state |= PERF_ATTACH_EXCLUSIVE;
return 0;
}
static void exclusive_event_destroy(struct perf_event *event)
{
struct pmu *pmu = event->pmu;
/* see comment in exclusive_event_init() */
if (event->attach_state & PERF_ATTACH_TASK)
atomic_dec(&pmu->exclusive_cnt);
else
atomic_inc(&pmu->exclusive_cnt);
event->attach_state &= ~PERF_ATTACH_EXCLUSIVE;
}
static bool exclusive_event_match(struct perf_event *e1, struct perf_event *e2)
{
if ((e1->pmu == e2->pmu) &&
(e1->cpu == e2->cpu ||
e1->cpu == -1 ||
e2->cpu == -1))
return true;
return false;
}
static bool exclusive_event_installable(struct perf_event *event,
struct perf_event_context *ctx)
{
struct perf_event *iter_event;
struct pmu *pmu = event->pmu;
lockdep_assert_held(&ctx->mutex);
if (!is_exclusive_pmu(pmu))
return true;
list_for_each_entry(iter_event, &ctx->event_list, event_entry) {
if (exclusive_event_match(iter_event, event))
return false;
}
return true;
}
static void perf_free_addr_filters(struct perf_event *event);
/* vs perf_event_alloc() error */
static void __free_event(struct perf_event *event)
{
struct pmu *pmu = event->pmu;
if (event->attach_state & PERF_ATTACH_CALLCHAIN)
put_callchain_buffers();
kfree(event->addr_filter_ranges);
if (event->attach_state & PERF_ATTACH_EXCLUSIVE)
exclusive_event_destroy(event);
if (is_cgroup_event(event))
perf_detach_cgroup(event);
if (event->attach_state & PERF_ATTACH_TASK_DATA)
detach_perf_ctx_data(event);
if (event->destroy)
event->destroy(event);
/*
* Must be after ->destroy(), due to uprobe_perf_close() using
* hw.target.
*/
if (event->hw.target)
put_task_struct(event->hw.target);
if (event->pmu_ctx) {
/*
* put_pmu_ctx() needs an event->ctx reference, because of
* epc->ctx.
*/
WARN_ON_ONCE(!pmu);
WARN_ON_ONCE(!event->ctx);
WARN_ON_ONCE(event->pmu_ctx->ctx != event->ctx);
put_pmu_ctx(event->pmu_ctx);
}
/*
* perf_event_free_task() relies on put_ctx() being 'last', in
* particular all task references must be cleaned up.
*/
if (event->ctx)
put_ctx(event->ctx);
if (pmu) {
module_put(pmu->module);
scoped_guard (spinlock, &pmu->events_lock) {
list_del(&event->pmu_list);
wake_up_var(pmu);
}
}
call_rcu(&event->rcu_head, free_event_rcu);
}
DEFINE_FREE(__free_event, struct perf_event *, if (_T) __free_event(_T))
/* vs perf_event_alloc() success */
static void _free_event(struct perf_event *event)
{
irq_work_sync(&event->pending_irq);
irq_work_sync(&event->pending_disable_irq);
unaccount_event(event);
security_perf_event_free(event);
if (event->rb) {
/*
* Can happen when we close an event with re-directed output.
*
* Since we have a 0 refcount, perf_mmap_close() will skip
* over us; possibly making our ring_buffer_put() the last.
*/
mutex_lock(&event->mmap_mutex);
ring_buffer_attach(event, NULL);
mutex_unlock(&event->mmap_mutex);
}
perf_event_free_bpf_prog(event);
perf_free_addr_filters(event);
__free_event(event);
}
/*
* Used to free events which have a known refcount of 1, such as in error paths
* of inherited events.
*/
static void free_event(struct perf_event *event)
{
if (WARN(atomic_long_cmpxchg(&event->refcount, 1, 0) != 1,
"unexpected event refcount: %ld; ptr=%p\n",
atomic_long_read(&event->refcount), event)) {
/* leak to avoid use-after-free */
return;
}
_free_event(event);
}
/*
* Remove user event from the owner task.
*/
static void perf_remove_from_owner(struct perf_event *event)
{
struct task_struct *owner;
rcu_read_lock();
/*
* Matches the smp_store_release() in perf_event_exit_task(). If we
* observe !owner it means the list deletion is complete and we can
* indeed free this event, otherwise we need to serialize on
* owner->perf_event_mutex.
*/
owner = READ_ONCE(event->owner);
if (owner) {
/*
* Since delayed_put_task_struct() also drops the last
* task reference we can safely take a new reference
* while holding the rcu_read_lock().
*/
get_task_struct(owner);
}
rcu_read_unlock();
if (owner) {
/*
* If we're here through perf_event_exit_task() we're already
* holding ctx->mutex which would be an inversion wrt. the
* normal lock order.
*
* However we can safely take this lock because its the child
* ctx->mutex.
*/
mutex_lock_nested(&owner->perf_event_mutex, SINGLE_DEPTH_NESTING);
/*
* We have to re-check the event->owner field, if it is cleared
* we raced with perf_event_exit_task(), acquiring the mutex
* ensured they're done, and we can proceed with freeing the
* event.
*/
if (event->owner) {
list_del_init(&event->owner_entry);
smp_store_release(&event->owner, NULL);
}
mutex_unlock(&owner->perf_event_mutex);
put_task_struct(owner);
}
}
static void put_event(struct perf_event *event)
{
struct perf_event *parent;
if (!atomic_long_dec_and_test(&event->refcount))
return;
parent = event->parent;
_free_event(event);
/* Matches the refcount bump in inherit_event() */
if (parent)
put_event(parent);
}
/*
* Kill an event dead; while event:refcount will preserve the event
* object, it will not preserve its functionality. Once the last 'user'
* gives up the object, we'll destroy the thing.
*/
int perf_event_release_kernel(struct perf_event *event)
{
struct perf_event_context *ctx = event->ctx;
struct perf_event *child, *tmp;
/*
* If we got here through err_alloc: free_event(event); we will not
* have attached to a context yet.
*/
if (!ctx) {
WARN_ON_ONCE(event->attach_state &
(PERF_ATTACH_CONTEXT|PERF_ATTACH_GROUP));
goto no_ctx;
}
if (!is_kernel_event(event))
perf_remove_from_owner(event);
ctx = perf_event_ctx_lock(event);
WARN_ON_ONCE(ctx->parent_ctx);
/*
* Mark this event as STATE_DEAD, there is no external reference to it
* anymore.
*
* Anybody acquiring event->child_mutex after the below loop _must_
* also see this, most importantly inherit_event() which will avoid
* placing more children on the list.
*
* Thus this guarantees that we will in fact observe and kill _ALL_
* child events.
*/
if (event->state > PERF_EVENT_STATE_REVOKED) {
perf_remove_from_context(event, DETACH_GROUP|DETACH_DEAD);
} else {
event->state = PERF_EVENT_STATE_DEAD;
}
perf_event_ctx_unlock(event, ctx);
again:
mutex_lock(&event->child_mutex);
list_for_each_entry(child, &event->child_list, child_list) {
/*
* Cannot change, child events are not migrated, see the
* comment with perf_event_ctx_lock_nested().
*/
ctx = READ_ONCE(child->ctx);
/*
* Since child_mutex nests inside ctx::mutex, we must jump
* through hoops. We start by grabbing a reference on the ctx.
*
* Since the event cannot get freed while we hold the
* child_mutex, the context must also exist and have a !0
* reference count.
*/
get_ctx(ctx);
/*
* Now that we have a ctx ref, we can drop child_mutex, and
* acquire ctx::mutex without fear of it going away. Then we
* can re-acquire child_mutex.
*/
mutex_unlock(&event->child_mutex);
mutex_lock(&ctx->mutex);
mutex_lock(&event->child_mutex);
/*
* Now that we hold ctx::mutex and child_mutex, revalidate our
* state, if child is still the first entry, it didn't get freed
* and we can continue doing so.
*/
tmp = list_first_entry_or_null(&event->child_list,
struct perf_event, child_list);
if (tmp == child) {
perf_remove_from_context(child, DETACH_GROUP | DETACH_CHILD);
} else {
child = NULL;
}
mutex_unlock(&event->child_mutex);
mutex_unlock(&ctx->mutex);
if (child) {
/* Last reference unless ->pending_task work is pending */
put_event(child);
}
put_ctx(ctx);
goto again;
}
mutex_unlock(&event->child_mutex);
no_ctx:
/*
* Last reference unless ->pending_task work is pending on this event
* or any of its children.
*/
put_event(event);
return 0;
}
EXPORT_SYMBOL_GPL(perf_event_release_kernel);
/*
* Called when the last reference to the file is gone.
*/
static int perf_release(struct inode *inode, struct file *file)
{
perf_event_release_kernel(file->private_data);
return 0;
}
static u64 __perf_event_read_value(struct perf_event *event, u64 *enabled, u64 *running)
{
struct perf_event *child;
u64 total = 0;
*enabled = 0;
*running = 0;
mutex_lock(&event->child_mutex);
(void)perf_event_read(event, false);
total += perf_event_count(event, false);
*enabled += event->total_time_enabled +
atomic64_read(&event->child_total_time_enabled);
*running += event->total_time_running +
atomic64_read(&event->child_total_time_running);
list_for_each_entry(child, &event->child_list, child_list) {
(void)perf_event_read(child, false);
total += perf_event_count(child, false);
*enabled += child->total_time_enabled;
*running += child->total_time_running;
}
mutex_unlock(&event->child_mutex);
return total;
}
u64 perf_event_read_value(struct perf_event *event, u64 *enabled, u64 *running)
{
struct perf_event_context *ctx;
u64 count;
ctx = perf_event_ctx_lock(event);
count = __perf_event_read_value(event, enabled, running);
perf_event_ctx_unlock(event, ctx);
return count;
}
EXPORT_SYMBOL_GPL(perf_event_read_value);
static int __perf_read_group_add(struct perf_event *leader,
u64 read_format, u64 *values)
{
struct perf_event_context *ctx = leader->ctx;
struct perf_event *sub, *parent;
unsigned long flags;
int n = 1; /* skip @nr */
int ret;
ret = perf_event_read(leader, true);
if (ret)
return ret;
raw_spin_lock_irqsave(&ctx->lock, flags);
/*
* Verify the grouping between the parent and child (inherited)
* events is still in tact.
*
* Specifically:
* - leader->ctx->lock pins leader->sibling_list
* - parent->child_mutex pins parent->child_list
* - parent->ctx->mutex pins parent->sibling_list
*
* Because parent->ctx != leader->ctx (and child_list nests inside
* ctx->mutex), group destruction is not atomic between children, also
* see perf_event_release_kernel(). Additionally, parent can grow the
* group.
*
* Therefore it is possible to have parent and child groups in a
* different configuration and summing over such a beast makes no sense
* what so ever.
*
* Reject this.
*/
parent = leader->parent;
if (parent &&
(parent->group_generation != leader->group_generation ||
parent->nr_siblings != leader->nr_siblings)) {
ret = -ECHILD;
goto unlock;
}
/*
* Since we co-schedule groups, {enabled,running} times of siblings
* will be identical to those of the leader, so we only publish one
* set.
*/
if (read_format & PERF_FORMAT_TOTAL_TIME_ENABLED) {
values[n++] += leader->total_time_enabled +
atomic64_read(&leader->child_total_time_enabled);
}
if (read_format & PERF_FORMAT_TOTAL_TIME_RUNNING) {
values[n++] += leader->total_time_running +
atomic64_read(&leader->child_total_time_running);
}
/*
* Write {count,id} tuples for every sibling.
*/
values[n++] += perf_event_count(leader, false);
if (read_format & PERF_FORMAT_ID)
values[n++] = primary_event_id(leader);
if (read_format & PERF_FORMAT_LOST)
values[n++] = atomic64_read(&leader->lost_samples);
for_each_sibling_event(sub, leader) {
values[n++] += perf_event_count(sub, false);
if (read_format & PERF_FORMAT_ID)
values[n++] = primary_event_id(sub);
if (read_format & PERF_FORMAT_LOST)
values[n++] = atomic64_read(&sub->lost_samples);
}
unlock:
raw_spin_unlock_irqrestore(&ctx->lock, flags);
return ret;
}
static int perf_read_group(struct perf_event *event,
u64 read_format, char __user *buf)
{
struct perf_event *leader = event->group_leader, *child;
struct perf_event_context *ctx = leader->ctx;
int ret;
u64 *values;
lockdep_assert_held(&ctx->mutex);
values = kzalloc(event->read_size, GFP_KERNEL);
if (!values)
return -ENOMEM;
values[0] = 1 + leader->nr_siblings;
mutex_lock(&leader->child_mutex);
ret = __perf_read_group_add(leader, read_format, values);
if (ret)
goto unlock;
list_for_each_entry(child, &leader->child_list, child_list) {
ret = __perf_read_group_add(child, read_format, values);
if (ret)
goto unlock;
}
mutex_unlock(&leader->child_mutex);
ret = event->read_size;
if (copy_to_user(buf, values, event->read_size))
ret = -EFAULT;
goto out;
unlock:
mutex_unlock(&leader->child_mutex);
out:
kfree(values);
return ret;
}
static int perf_read_one(struct perf_event *event,
u64 read_format, char __user *buf)
{
u64 enabled, running;
u64 values[5];
int n = 0;
values[n++] = __perf_event_read_value(event, &enabled, &running);
if (read_format & PERF_FORMAT_TOTAL_TIME_ENABLED)
values[n++] = enabled;
if (read_format & PERF_FORMAT_TOTAL_TIME_RUNNING)
values[n++] = running;
if (read_format & PERF_FORMAT_ID)
values[n++] = primary_event_id(event);
if (read_format & PERF_FORMAT_LOST)
values[n++] = atomic64_read(&event->lost_samples);
if (copy_to_user(buf, values, n * sizeof(u64)))
return -EFAULT;
return n * sizeof(u64);
}
static bool is_event_hup(struct perf_event *event)
{
bool no_children;
if (event->state > PERF_EVENT_STATE_EXIT)
return false;
mutex_lock(&event->child_mutex);
no_children = list_empty(&event->child_list);
mutex_unlock(&event->child_mutex);
return no_children;
}
/*
* Read the performance event - simple non blocking version for now
*/
static ssize_t
__perf_read(struct perf_event *event, char __user *buf, size_t count)
{
u64 read_format = event->attr.read_format;
int ret;
/*
* Return end-of-file for a read on an event that is in
* error state (i.e. because it was pinned but it couldn't be
* scheduled on to the CPU at some point).
*/
if (event->state == PERF_EVENT_STATE_ERROR)
return 0;
if (count < event->read_size)
return -ENOSPC;
WARN_ON_ONCE(event->ctx->parent_ctx);
if (read_format & PERF_FORMAT_GROUP)
ret = perf_read_group(event, read_format, buf);
else
ret = perf_read_one(event, read_format, buf);
return ret;
}
static ssize_t
perf_read(struct file *file, char __user *buf, size_t count, loff_t *ppos)
{
struct perf_event *event = file->private_data;
struct perf_event_context *ctx;
int ret;
ret = security_perf_event_read(event);
if (ret)
return ret;
ctx = perf_event_ctx_lock(event);
ret = __perf_read(event, buf, count);
perf_event_ctx_unlock(event, ctx);
return ret;
}
static __poll_t perf_poll(struct file *file, poll_table *wait)
{
struct perf_event *event = file->private_data;
struct perf_buffer *rb;
__poll_t events = EPOLLHUP;
if (event->state <= PERF_EVENT_STATE_REVOKED)
return EPOLLERR;
poll_wait(file, &event->waitq, wait);
if (event->state <= PERF_EVENT_STATE_REVOKED)
return EPOLLERR;
if (is_event_hup(event))
return events;
if (unlikely(READ_ONCE(event->state) == PERF_EVENT_STATE_ERROR &&
event->attr.pinned))
return EPOLLERR;
/*
* Pin the event->rb by taking event->mmap_mutex; otherwise
* perf_event_set_output() can swizzle our rb and make us miss wakeups.
*/
mutex_lock(&event->mmap_mutex);
rb = event->rb;
if (rb)
events = atomic_xchg(&rb->poll, 0);
mutex_unlock(&event->mmap_mutex);
return events;
}
static void _perf_event_reset(struct perf_event *event)
{
(void)perf_event_read(event, false);
local64_set(&event->count, 0);
perf_event_update_userpage(event);
}
/* Assume it's not an event with inherit set. */
u64 perf_event_pause(struct perf_event *event, bool reset)
{
struct perf_event_context *ctx;
u64 count;
ctx = perf_event_ctx_lock(event);
WARN_ON_ONCE(event->attr.inherit);
_perf_event_disable(event);
count = local64_read(&event->count);
if (reset)
local64_set(&event->count, 0);
perf_event_ctx_unlock(event, ctx);
return count;
}
EXPORT_SYMBOL_GPL(perf_event_pause);
/*
* Holding the top-level event's child_mutex means that any
* descendant process that has inherited this event will block
* in perf_event_exit_event() if it goes to exit, thus satisfying the
* task existence requirements of perf_event_enable/disable.
*/
static void perf_event_for_each_child(struct perf_event *event,
void (*func)(struct perf_event *))
{
struct perf_event *child;
WARN_ON_ONCE(event->ctx->parent_ctx);
mutex_lock(&event->child_mutex);
func(event);
list_for_each_entry(child, &event->child_list, child_list)
func(child);
mutex_unlock(&event->child_mutex);
}
static void perf_event_for_each(struct perf_event *event,
void (*func)(struct perf_event *))
{
struct perf_event_context *ctx = event->ctx;
struct perf_event *sibling;
lockdep_assert_held(&ctx->mutex);
event = event->group_leader;
perf_event_for_each_child(event, func);
for_each_sibling_event(sibling, event)
perf_event_for_each_child(sibling, func);
}
static void __perf_event_period(struct perf_event *event,
struct perf_cpu_context *cpuctx,
struct perf_event_context *ctx,
void *info)
{
u64 value = *((u64 *)info);
bool active;
if (event->attr.freq) {
event->attr.sample_freq = value;
} else {
event->attr.sample_period = value;
event->hw.sample_period = value;
}
active = (event->state == PERF_EVENT_STATE_ACTIVE);
if (active) {
perf_pmu_disable(event->pmu);
event->pmu->stop(event, PERF_EF_UPDATE);
}
local64_set(&event->hw.period_left, 0);
if (active) {
event->pmu->start(event, PERF_EF_RELOAD);
/*
* Once the period is force-reset, the event starts immediately.
* But the event/group could be throttled. Unthrottle the
* event/group now to avoid the next tick trying to unthrottle
* while we already re-started the event/group.
*/
if (event->hw.interrupts == MAX_INTERRUPTS)
perf_event_unthrottle_group(event, true);
perf_pmu_enable(event->pmu);
}
}
static int perf_event_check_period(struct perf_event *event, u64 value)
{
return event->pmu->check_period(event, value);
}
static int _perf_event_period(struct perf_event *event, u64 value)
{
if (!is_sampling_event(event))
return -EINVAL;
if (!value)
return -EINVAL;
if (event->attr.freq) {
if (value > sysctl_perf_event_sample_rate)
return -EINVAL;
} else {
if (perf_event_check_period(event, value))
return -EINVAL;
if (value & (1ULL << 63))
return -EINVAL;
}
event_function_call(event, __perf_event_period, &value);
return 0;
}
int perf_event_period(struct perf_event *event, u64 value)
{
struct perf_event_context *ctx;
int ret;
ctx = perf_event_ctx_lock(event);
ret = _perf_event_period(event, value);
perf_event_ctx_unlock(event, ctx);
return ret;
}
EXPORT_SYMBOL_GPL(perf_event_period);
static const struct file_operations perf_fops;
static inline bool is_perf_file(struct fd f)
{
return !fd_empty(f) && fd_file(f)->f_op == &perf_fops;
}
static int perf_event_set_output(struct perf_event *event,
struct perf_event *output_event);
static int perf_event_set_filter(struct perf_event *event, void __user *arg);
static int perf_copy_attr(struct perf_event_attr __user *uattr,
struct perf_event_attr *attr);
static int __perf_event_set_bpf_prog(struct perf_event *event,
struct bpf_prog *prog,
u64 bpf_cookie);
static long _perf_ioctl(struct perf_event *event, unsigned int cmd, unsigned long arg)
{
void (*func)(struct perf_event *);
u32 flags = arg;
if (event->state <= PERF_EVENT_STATE_REVOKED)
return -ENODEV;
switch (cmd) {
case PERF_EVENT_IOC_ENABLE:
func = _perf_event_enable;
break;
case PERF_EVENT_IOC_DISABLE:
func = _perf_event_disable;
break;
case PERF_EVENT_IOC_RESET:
func = _perf_event_reset;
break;
case PERF_EVENT_IOC_REFRESH:
return _perf_event_refresh(event, arg);
case PERF_EVENT_IOC_PERIOD:
{
u64 value;
if (copy_from_user(&value, (u64 __user *)arg, sizeof(value)))
return -EFAULT;
return _perf_event_period(event, value);
}
case PERF_EVENT_IOC_ID:
{
u64 id = primary_event_id(event);
if (copy_to_user((void __user *)arg, &id, sizeof(id)))
return -EFAULT;
return 0;
}
case PERF_EVENT_IOC_SET_OUTPUT:
{
CLASS(fd, output)(arg); // arg == -1 => empty
struct perf_event *output_event = NULL;
if (arg != -1) {
if (!is_perf_file(output))
return -EBADF;
output_event = fd_file(output)->private_data;
}
return perf_event_set_output(event, output_event);
}
case PERF_EVENT_IOC_SET_FILTER:
return perf_event_set_filter(event, (void __user *)arg);
case PERF_EVENT_IOC_SET_BPF:
{
struct bpf_prog *prog;
int err;
prog = bpf_prog_get(arg);
if (IS_ERR(prog))
return PTR_ERR(prog);
err = __perf_event_set_bpf_prog(event, prog, 0);
if (err) {
bpf_prog_put(prog);
return err;
}
return 0;
}
case PERF_EVENT_IOC_PAUSE_OUTPUT: {
struct perf_buffer *rb;
rcu_read_lock();
rb = rcu_dereference(event->rb);
if (!rb || !rb->nr_pages) {
rcu_read_unlock();
return -EINVAL;
}
rb_toggle_paused(rb, !!arg);
rcu_read_unlock();
return 0;
}
case PERF_EVENT_IOC_QUERY_BPF:
return perf_event_query_prog_array(event, (void __user *)arg);
case PERF_EVENT_IOC_MODIFY_ATTRIBUTES: {
struct perf_event_attr new_attr;
int err = perf_copy_attr((struct perf_event_attr __user *)arg,
&new_attr);
if (err)
return err;
return perf_event_modify_attr(event, &new_attr);
}
default:
return -ENOTTY;
}
if (flags & PERF_IOC_FLAG_GROUP)
perf_event_for_each(event, func);
else
perf_event_for_each_child(event, func);
return 0;
}
static long perf_ioctl(struct file *file, unsigned int cmd, unsigned long arg)
{
struct perf_event *event = file->private_data;
struct perf_event_context *ctx;
long ret;
/* Treat ioctl like writes as it is likely a mutating operation. */
ret = security_perf_event_write(event);
if (ret)
return ret;
ctx = perf_event_ctx_lock(event);
ret = _perf_ioctl(event, cmd, arg);
perf_event_ctx_unlock(event, ctx);
return ret;
}
#ifdef CONFIG_COMPAT
static long perf_compat_ioctl(struct file *file, unsigned int cmd,
unsigned long arg)
{
switch (_IOC_NR(cmd)) {
case _IOC_NR(PERF_EVENT_IOC_SET_FILTER):
case _IOC_NR(PERF_EVENT_IOC_ID):
case _IOC_NR(PERF_EVENT_IOC_QUERY_BPF):
case _IOC_NR(PERF_EVENT_IOC_MODIFY_ATTRIBUTES):
/* Fix up pointer size (usually 4 -> 8 in 32-on-64-bit case */
if (_IOC_SIZE(cmd) == sizeof(compat_uptr_t)) {
cmd &= ~IOCSIZE_MASK;
cmd |= sizeof(void *) << IOCSIZE_SHIFT;
}
break;
}
return perf_ioctl(file, cmd, arg);
}
#else
# define perf_compat_ioctl NULL
#endif
int perf_event_task_enable(void)
{
struct perf_event_context *ctx;
struct perf_event *event;
mutex_lock(¤t->perf_event_mutex);
list_for_each_entry(event, ¤t->perf_event_list, owner_entry) {
ctx = perf_event_ctx_lock(event);
perf_event_for_each_child(event, _perf_event_enable);
perf_event_ctx_unlock(event, ctx);
}
mutex_unlock(¤t->perf_event_mutex);
return 0;
}
int perf_event_task_disable(void)
{
struct perf_event_context *ctx;
struct perf_event *event;
mutex_lock(¤t->perf_event_mutex);
list_for_each_entry(event, ¤t->perf_event_list, owner_entry) {
ctx = perf_event_ctx_lock(event);
perf_event_for_each_child(event, _perf_event_disable);
perf_event_ctx_unlock(event, ctx);
}
mutex_unlock(¤t->perf_event_mutex);
return 0;
}
static int perf_event_index(struct perf_event *event)
{
if (event->hw.state & PERF_HES_STOPPED)
return 0;
if (event->state != PERF_EVENT_STATE_ACTIVE)
return 0;
return event->pmu->event_idx(event);
}
static void perf_event_init_userpage(struct perf_event *event)
{
struct perf_event_mmap_page *userpg;
struct perf_buffer *rb;
rcu_read_lock();
rb = rcu_dereference(event->rb);
if (!rb)
goto unlock;
userpg = rb->user_page;
/* Allow new userspace to detect that bit 0 is deprecated */
userpg->cap_bit0_is_deprecated = 1;
userpg->size = offsetof(struct perf_event_mmap_page, __reserved);
userpg->data_offset = PAGE_SIZE;
userpg->data_size = perf_data_size(rb);
unlock:
rcu_read_unlock();
}
void __weak arch_perf_update_userpage(
struct perf_event *event, struct perf_event_mmap_page *userpg, u64 now)
{
}
/*
* Callers need to ensure there can be no nesting of this function, otherwise
* the seqlock logic goes bad. We can not serialize this because the arch
* code calls this from NMI context.
*/
void perf_event_update_userpage(struct perf_event *event)
{
struct perf_event_mmap_page *userpg;
struct perf_buffer *rb;
u64 enabled, running, now;
rcu_read_lock();
rb = rcu_dereference(event->rb);
if (!rb)
goto unlock;
/*
* compute total_time_enabled, total_time_running
* based on snapshot values taken when the event
* was last scheduled in.
*
* we cannot simply called update_context_time()
* because of locking issue as we can be called in
* NMI context
*/
calc_timer_values(event, &now, &enabled, &running);
userpg = rb->user_page;
/*
* Disable preemption to guarantee consistent time stamps are stored to
* the user page.
*/
preempt_disable();
++userpg->lock;
barrier();
userpg->index = perf_event_index(event);
userpg->offset = perf_event_count(event, false);
if (userpg->index)
userpg->offset -= local64_read(&event->hw.prev_count);
userpg->time_enabled = enabled +
atomic64_read(&event->child_total_time_enabled);
userpg->time_running = running +
atomic64_read(&event->child_total_time_running);
arch_perf_update_userpage(event, userpg, now);
barrier();
++userpg->lock;
preempt_enable();
unlock:
rcu_read_unlock();
}
EXPORT_SYMBOL_GPL(perf_event_update_userpage);
static void ring_buffer_attach(struct perf_event *event,
struct perf_buffer *rb)
{
struct perf_buffer *old_rb = NULL;
unsigned long flags;
WARN_ON_ONCE(event->parent);
if (event->rb) {
/*
* Should be impossible, we set this when removing
* event->rb_entry and wait/clear when adding event->rb_entry.
*/
WARN_ON_ONCE(event->rcu_pending);
old_rb = event->rb;
spin_lock_irqsave(&old_rb->event_lock, flags);
list_del_rcu(&event->rb_entry);
spin_unlock_irqrestore(&old_rb->event_lock, flags);
event->rcu_batches = get_state_synchronize_rcu();
event->rcu_pending = 1;
}
if (rb) {
if (event->rcu_pending) {
cond_synchronize_rcu(event->rcu_batches);
event->rcu_pending = 0;
}
spin_lock_irqsave(&rb->event_lock, flags);
list_add_rcu(&event->rb_entry, &rb->event_list);
spin_unlock_irqrestore(&rb->event_lock, flags);
}
/*
* Avoid racing with perf_mmap_close(AUX): stop the event
* before swizzling the event::rb pointer; if it's getting
* unmapped, its aux_mmap_count will be 0 and it won't
* restart. See the comment in __perf_pmu_output_stop().
*
* Data will inevitably be lost when set_output is done in
* mid-air, but then again, whoever does it like this is
* not in for the data anyway.
*/
if (has_aux(event))
perf_event_stop(event, 0);
rcu_assign_pointer(event->rb, rb);
if (old_rb) {
ring_buffer_put(old_rb);
/*
* Since we detached before setting the new rb, so that we
* could attach the new rb, we could have missed a wakeup.
* Provide it now.
*/
wake_up_all(&event->waitq);
}
}
static void ring_buffer_wakeup(struct perf_event *event)
{
struct perf_buffer *rb;
if (event->parent)
event = event->parent;
rcu_read_lock();
rb = rcu_dereference(event->rb);
if (rb) {
list_for_each_entry_rcu(event, &rb->event_list, rb_entry)
wake_up_all(&event->waitq);
}
rcu_read_unlock();
}
struct perf_buffer *ring_buffer_get(struct perf_event *event)
{
struct perf_buffer *rb;
if (event->parent)
event = event->parent;
rcu_read_lock();
rb = rcu_dereference(event->rb);
if (rb) {
if (!refcount_inc_not_zero(&rb->refcount))
rb = NULL;
}
rcu_read_unlock();
return rb;
}
void ring_buffer_put(struct perf_buffer *rb)
{
if (!refcount_dec_and_test(&rb->refcount))
return;
WARN_ON_ONCE(!list_empty(&rb->event_list));
call_rcu(&rb->rcu_head, rb_free_rcu);
}
typedef void (*mapped_f)(struct perf_event *event, struct mm_struct *mm);
#define get_mapped(event, func) \
({ struct pmu *pmu; \
mapped_f f = NULL; \
guard(rcu)(); \
pmu = READ_ONCE(event->pmu); \
if (pmu) \
f = pmu->func; \
f; \
})
static void perf_mmap_open(struct vm_area_struct *vma)
{
struct perf_event *event = vma->vm_file->private_data;
mapped_f mapped = get_mapped(event, event_mapped);
refcount_inc(&event->mmap_count);
refcount_inc(&event->rb->mmap_count);
if (vma->vm_pgoff)
refcount_inc(&event->rb->aux_mmap_count);
if (mapped)
mapped(event, vma->vm_mm);
}
static void perf_pmu_output_stop(struct perf_event *event);
/*
* A buffer can be mmap()ed multiple times; either directly through the same
* event, or through other events by use of perf_event_set_output().
*
* In order to undo the VM accounting done by perf_mmap() we need to destroy
* the buffer here, where we still have a VM context. This means we need
* to detach all events redirecting to us.
*/
static void perf_mmap_close(struct vm_area_struct *vma)
{
struct perf_event *event = vma->vm_file->private_data;
mapped_f unmapped = get_mapped(event, event_unmapped);
struct perf_buffer *rb = ring_buffer_get(event);
struct user_struct *mmap_user = rb->mmap_user;
int mmap_locked = rb->mmap_locked;
unsigned long size = perf_data_size(rb);
bool detach_rest = false;
/* FIXIES vs perf_pmu_unregister() */
if (unmapped)
unmapped(event, vma->vm_mm);
/*
* The AUX buffer is strictly a sub-buffer, serialize using aux_mutex
* to avoid complications.
*/
if (rb_has_aux(rb) && vma->vm_pgoff == rb->aux_pgoff &&
refcount_dec_and_mutex_lock(&rb->aux_mmap_count, &rb->aux_mutex)) {
/*
* Stop all AUX events that are writing to this buffer,
* so that we can free its AUX pages and corresponding PMU
* data. Note that after rb::aux_mmap_count dropped to zero,
* they won't start any more (see perf_aux_output_begin()).
*/
perf_pmu_output_stop(event);
/* now it's safe to free the pages */
atomic_long_sub(rb->aux_nr_pages - rb->aux_mmap_locked, &mmap_user->locked_vm);
atomic64_sub(rb->aux_mmap_locked, &vma->vm_mm->pinned_vm);
/* this has to be the last one */
rb_free_aux(rb);
WARN_ON_ONCE(refcount_read(&rb->aux_refcount));
mutex_unlock(&rb->aux_mutex);
}
if (refcount_dec_and_test(&rb->mmap_count))
detach_rest = true;
if (!refcount_dec_and_mutex_lock(&event->mmap_count, &event->mmap_mutex))
goto out_put;
ring_buffer_attach(event, NULL);
mutex_unlock(&event->mmap_mutex);
/* If there's still other mmap()s of this buffer, we're done. */
if (!detach_rest)
goto out_put;
/*
* No other mmap()s, detach from all other events that might redirect
* into the now unreachable buffer. Somewhat complicated by the
* fact that rb::event_lock otherwise nests inside mmap_mutex.
*/
again:
rcu_read_lock();
list_for_each_entry_rcu(event, &rb->event_list, rb_entry) {
if (!atomic_long_inc_not_zero(&event->refcount)) {
/*
* This event is en-route to free_event() which will
* detach it and remove it from the list.
*/
continue;
}
rcu_read_unlock();
mutex_lock(&event->mmap_mutex);
/*
* Check we didn't race with perf_event_set_output() which can
* swizzle the rb from under us while we were waiting to
* acquire mmap_mutex.
*
* If we find a different rb; ignore this event, a next
* iteration will no longer find it on the list. We have to
* still restart the iteration to make sure we're not now
* iterating the wrong list.
*/
if (event->rb == rb)
ring_buffer_attach(event, NULL);
mutex_unlock(&event->mmap_mutex);
put_event(event);
/*
* Restart the iteration; either we're on the wrong list or
* destroyed its integrity by doing a deletion.
*/
goto again;
}
rcu_read_unlock();
/*
* It could be there's still a few 0-ref events on the list; they'll
* get cleaned up by free_event() -- they'll also still have their
* ref on the rb and will free it whenever they are done with it.
*
* Aside from that, this buffer is 'fully' detached and unmapped,
* undo the VM accounting.
*/
atomic_long_sub((size >> PAGE_SHIFT) + 1 - mmap_locked,
&mmap_user->locked_vm);
atomic64_sub(mmap_locked, &vma->vm_mm->pinned_vm);
free_uid(mmap_user);
out_put:
ring_buffer_put(rb); /* could be last */
}
static vm_fault_t perf_mmap_pfn_mkwrite(struct vm_fault *vmf)
{
/* The first page is the user control page, others are read-only. */
return vmf->pgoff == 0 ? 0 : VM_FAULT_SIGBUS;
}
static int perf_mmap_may_split(struct vm_area_struct *vma, unsigned long addr)
{
/*
* Forbid splitting perf mappings to prevent refcount leaks due to
* the resulting non-matching offsets and sizes. See open()/close().
*/
return -EINVAL;
}
static const struct vm_operations_struct perf_mmap_vmops = {
.open = perf_mmap_open,
.close = perf_mmap_close, /* non mergeable */
.pfn_mkwrite = perf_mmap_pfn_mkwrite,
.may_split = perf_mmap_may_split,
};
static int map_range(struct perf_buffer *rb, struct vm_area_struct *vma)
{
unsigned long nr_pages = vma_pages(vma);
int err = 0;
unsigned long pagenum;
/*
* We map this as a VM_PFNMAP VMA.
*
* This is not ideal as this is designed broadly for mappings of PFNs
* referencing memory-mapped I/O ranges or non-system RAM i.e. for which
* !pfn_valid(pfn).
*
* We are mapping kernel-allocated memory (memory we manage ourselves)
* which would more ideally be mapped using vm_insert_page() or a
* similar mechanism, that is as a VM_MIXEDMAP mapping.
*
* However this won't work here, because:
*
* 1. It uses vma->vm_page_prot, but this field has not been completely
* setup at the point of the f_op->mmp() hook, so we are unable to
* indicate that this should be mapped CoW in order that the
* mkwrite() hook can be invoked to make the first page R/W and the
* rest R/O as desired.
*
* 2. Anything other than a VM_PFNMAP of valid PFNs will result in
* vm_normal_page() returning a struct page * pointer, which means
* vm_ops->page_mkwrite() will be invoked rather than
* vm_ops->pfn_mkwrite(), and this means we have to set page->mapping
* to work around retry logic in the fault handler, however this
* field is no longer allowed to be used within struct page.
*
* 3. Having a struct page * made available in the fault logic also
* means that the page gets put on the rmap and becomes
* inappropriately accessible and subject to map and ref counting.
*
* Ideally we would have a mechanism that could explicitly express our
* desires, but this is not currently the case, so we instead use
* VM_PFNMAP.
*
* We manage the lifetime of these mappings with internal refcounts (see
* perf_mmap_open() and perf_mmap_close()) so we ensure the lifetime of
* this mapping is maintained correctly.
*/
for (pagenum = 0; pagenum < nr_pages; pagenum++) {
unsigned long va = vma->vm_start + PAGE_SIZE * pagenum;
struct page *page = perf_mmap_to_page(rb, vma->vm_pgoff + pagenum);
if (page == NULL) {
err = -EINVAL;
break;
}
/* Map readonly, perf_mmap_pfn_mkwrite() called on write fault. */
err = remap_pfn_range(vma, va, page_to_pfn(page), PAGE_SIZE,
vm_get_page_prot(vma->vm_flags & ~VM_SHARED));
if (err)
break;
}
#ifdef CONFIG_MMU
/* Clear any partial mappings on error. */
if (err)
zap_page_range_single(vma, vma->vm_start, nr_pages * PAGE_SIZE, NULL);
#endif
return err;
}
static bool perf_mmap_calc_limits(struct vm_area_struct *vma, long *user_extra, long *extra)
{
unsigned long user_locked, user_lock_limit, locked, lock_limit;
struct user_struct *user = current_user();
user_lock_limit = sysctl_perf_event_mlock >> (PAGE_SHIFT - 10);
/* Increase the limit linearly with more CPUs */
user_lock_limit *= num_online_cpus();
user_locked = atomic_long_read(&user->locked_vm);
/*
* sysctl_perf_event_mlock may have changed, so that
* user->locked_vm > user_lock_limit
*/
if (user_locked > user_lock_limit)
user_locked = user_lock_limit;
user_locked += *user_extra;
if (user_locked > user_lock_limit) {
/*
* charge locked_vm until it hits user_lock_limit;
* charge the rest from pinned_vm
*/
*extra = user_locked - user_lock_limit;
*user_extra -= *extra;
}
lock_limit = rlimit(RLIMIT_MEMLOCK);
lock_limit >>= PAGE_SHIFT;
locked = atomic64_read(&vma->vm_mm->pinned_vm) + *extra;
return locked <= lock_limit || !perf_is_paranoid() || capable(CAP_IPC_LOCK);
}
static void perf_mmap_account(struct vm_area_struct *vma, long user_extra, long extra)
{
struct user_struct *user = current_user();
atomic_long_add(user_extra, &user->locked_vm);
atomic64_add(extra, &vma->vm_mm->pinned_vm);
}
static int perf_mmap_rb(struct vm_area_struct *vma, struct perf_event *event,
unsigned long nr_pages)
{
long extra = 0, user_extra = nr_pages;
struct perf_buffer *rb;
int rb_flags = 0;
nr_pages -= 1;
/*
* If we have rb pages ensure they're a power-of-two number, so we
* can do bitmasks instead of modulo.
*/
if (nr_pages != 0 && !is_power_of_2(nr_pages))
return -EINVAL;
WARN_ON_ONCE(event->ctx->parent_ctx);
if (event->rb) {
if (data_page_nr(event->rb) != nr_pages)
return -EINVAL;
if (refcount_inc_not_zero(&event->rb->mmap_count)) {
/*
* Success -- managed to mmap() the same buffer
* multiple times.
*/
perf_mmap_account(vma, user_extra, extra);
refcount_inc(&event->mmap_count);
return 0;
}
/*
* Raced against perf_mmap_close()'s
* refcount_dec_and_mutex_lock() remove the
* event and continue as if !event->rb
*/
ring_buffer_attach(event, NULL);
}
if (!perf_mmap_calc_limits(vma, &user_extra, &extra))
return -EPERM;
if (vma->vm_flags & VM_WRITE)
rb_flags |= RING_BUFFER_WRITABLE;
rb = rb_alloc(nr_pages,
event->attr.watermark ? event->attr.wakeup_watermark : 0,
event->cpu, rb_flags);
if (!rb)
return -ENOMEM;
refcount_set(&rb->mmap_count, 1);
rb->mmap_user = get_current_user();
rb->mmap_locked = extra;
ring_buffer_attach(event, rb);
perf_event_update_time(event);
perf_event_init_userpage(event);
perf_event_update_userpage(event);
perf_mmap_account(vma, user_extra, extra);
refcount_set(&event->mmap_count, 1);
return 0;
}
static int perf_mmap_aux(struct vm_area_struct *vma, struct perf_event *event,
unsigned long nr_pages)
{
long extra = 0, user_extra = nr_pages;
u64 aux_offset, aux_size;
struct perf_buffer *rb;
int ret, rb_flags = 0;
rb = event->rb;
if (!rb)
return -EINVAL;
guard(mutex)(&rb->aux_mutex);
/*
* AUX area mapping: if rb->aux_nr_pages != 0, it's already
* mapped, all subsequent mappings should have the same size
* and offset. Must be above the normal perf buffer.
*/
aux_offset = READ_ONCE(rb->user_page->aux_offset);
aux_size = READ_ONCE(rb->user_page->aux_size);
if (aux_offset < perf_data_size(rb) + PAGE_SIZE)
return -EINVAL;
if (aux_offset != vma->vm_pgoff << PAGE_SHIFT)
return -EINVAL;
/* already mapped with a different offset */
if (rb_has_aux(rb) && rb->aux_pgoff != vma->vm_pgoff)
return -EINVAL;
if (aux_size != nr_pages * PAGE_SIZE)
return -EINVAL;
/* already mapped with a different size */
if (rb_has_aux(rb) && rb->aux_nr_pages != nr_pages)
return -EINVAL;
if (!is_power_of_2(nr_pages))
return -EINVAL;
if (!refcount_inc_not_zero(&rb->mmap_count))
return -EINVAL;
if (rb_has_aux(rb)) {
refcount_inc(&rb->aux_mmap_count);
} else {
if (!perf_mmap_calc_limits(vma, &user_extra, &extra)) {
refcount_dec(&rb->mmap_count);
return -EPERM;
}
WARN_ON(!rb && event->rb);
if (vma->vm_flags & VM_WRITE)
rb_flags |= RING_BUFFER_WRITABLE;
ret = rb_alloc_aux(rb, event, vma->vm_pgoff, nr_pages,
event->attr.aux_watermark, rb_flags);
if (ret) {
refcount_dec(&rb->mmap_count);
return ret;
}
refcount_set(&rb->aux_mmap_count, 1);
rb->aux_mmap_locked = extra;
}
perf_mmap_account(vma, user_extra, extra);
refcount_inc(&event->mmap_count);
return 0;
}
static int perf_mmap(struct file *file, struct vm_area_struct *vma)
{
struct perf_event *event = file->private_data;
unsigned long vma_size, nr_pages;
mapped_f mapped;
int ret;
/*
* Don't allow mmap() of inherited per-task counters. This would
* create a performance issue due to all children writing to the
* same rb.
*/
if (event->cpu == -1 && event->attr.inherit)
return -EINVAL;
if (!(vma->vm_flags & VM_SHARED))
return -EINVAL;
ret = security_perf_event_read(event);
if (ret)
return ret;
vma_size = vma->vm_end - vma->vm_start;
nr_pages = vma_size / PAGE_SIZE;
if (nr_pages > INT_MAX)
return -ENOMEM;
if (vma_size != PAGE_SIZE * nr_pages)
return -EINVAL;
scoped_guard (mutex, &event->mmap_mutex) {
/*
* This relies on __pmu_detach_event() taking mmap_mutex after marking
* the event REVOKED. Either we observe the state, or __pmu_detach_event()
* will detach the rb created here.
*/
if (event->state <= PERF_EVENT_STATE_REVOKED)
return -ENODEV;
if (vma->vm_pgoff == 0)
ret = perf_mmap_rb(vma, event, nr_pages);
else
ret = perf_mmap_aux(vma, event, nr_pages);
if (ret)
return ret;
}
/*
* Since pinned accounting is per vm we cannot allow fork() to copy our
* vma.
*/
vm_flags_set(vma, VM_DONTCOPY | VM_DONTEXPAND | VM_DONTDUMP);
vma->vm_ops = &perf_mmap_vmops;
mapped = get_mapped(event, event_mapped);
if (mapped)
mapped(event, vma->vm_mm);
/*
* Try to map it into the page table. On fail, invoke
* perf_mmap_close() to undo the above, as the callsite expects
* full cleanup in this case and therefore does not invoke
* vmops::close().
*/
ret = map_range(event->rb, vma);
if (ret)
perf_mmap_close(vma);
return ret;
}
static int perf_fasync(int fd, struct file *filp, int on)
{
struct inode *inode = file_inode(filp);
struct perf_event *event = filp->private_data;
int retval;
if (event->state <= PERF_EVENT_STATE_REVOKED)
return -ENODEV;
inode_lock(inode);
retval = fasync_helper(fd, filp, on, &event->fasync);
inode_unlock(inode);
if (retval < 0)
return retval;
return 0;
}
static const struct file_operations perf_fops = {
.release = perf_release,
.read = perf_read,
.poll = perf_poll,
.unlocked_ioctl = perf_ioctl,
.compat_ioctl = perf_compat_ioctl,
.mmap = perf_mmap,
.fasync = perf_fasync,
};
/*
* Perf event wakeup
*
* If there's data, ensure we set the poll() state and publish everything
* to user-space before waking everybody up.
*/
void perf_event_wakeup(struct perf_event *event)
{
ring_buffer_wakeup(event);
if (event->pending_kill) {
kill_fasync(perf_event_fasync(event), SIGIO, event->pending_kill);
event->pending_kill = 0;
}
}
static void perf_sigtrap(struct perf_event *event)
{
/*
* Both perf_pending_task() and perf_pending_irq() can race with the
* task exiting.
*/
if (current->flags & PF_EXITING)
return;
/*
* We'd expect this to only occur if the irq_work is delayed and either
* ctx->task or current has changed in the meantime. This can be the
* case on architectures that do not implement arch_irq_work_raise().
*/
if (WARN_ON_ONCE(event->ctx->task != current))
return;
send_sig_perf((void __user *)event->pending_addr,
event->orig_type, event->attr.sig_data);
}
/*
* Deliver the pending work in-event-context or follow the context.
*/
static void __perf_pending_disable(struct perf_event *event)
{
int cpu = READ_ONCE(event->oncpu);
/*
* If the event isn't running; we done. event_sched_out() will have
* taken care of things.
*/
if (cpu < 0)
return;
/*
* Yay, we hit home and are in the context of the event.
*/
if (cpu == smp_processor_id()) {
if (event->pending_disable) {
event->pending_disable = 0;
perf_event_disable_local(event);
}
return;
}
/*
* CPU-A CPU-B
*
* perf_event_disable_inatomic()
* @pending_disable = 1;
* irq_work_queue();
*
* sched-out
* @pending_disable = 0;
*
* sched-in
* perf_event_disable_inatomic()
* @pending_disable = 1;
* irq_work_queue(); // FAILS
*
* irq_work_run()
* perf_pending_disable()
*
* But the event runs on CPU-B and wants disabling there.
*/
irq_work_queue_on(&event->pending_disable_irq, cpu);
}
static void perf_pending_disable(struct irq_work *entry)
{
struct perf_event *event = container_of(entry, struct perf_event, pending_disable_irq);
int rctx;
/*
* If we 'fail' here, that's OK, it means recursion is already disabled
* and we won't recurse 'further'.
*/
rctx = perf_swevent_get_recursion_context();
__perf_pending_disable(event);
if (rctx >= 0)
perf_swevent_put_recursion_context(rctx);
}
static void perf_pending_irq(struct irq_work *entry)
{
struct perf_event *event = container_of(entry, struct perf_event, pending_irq);
int rctx;
/*
* If we 'fail' here, that's OK, it means recursion is already disabled
* and we won't recurse 'further'.
*/
rctx = perf_swevent_get_recursion_context();
/*
* The wakeup isn't bound to the context of the event -- it can happen
* irrespective of where the event is.
*/
if (event->pending_wakeup) {
event->pending_wakeup = 0;
perf_event_wakeup(event);
}
if (rctx >= 0)
perf_swevent_put_recursion_context(rctx);
}
static void perf_pending_task(struct callback_head *head)
{
struct perf_event *event = container_of(head, struct perf_event, pending_task);
int rctx;
/*
* If we 'fail' here, that's OK, it means recursion is already disabled
* and we won't recurse 'further'.
*/
rctx = perf_swevent_get_recursion_context();
if (event->pending_work) {
event->pending_work = 0;
perf_sigtrap(event);
local_dec(&event->ctx->nr_no_switch_fast);
}
put_event(event);
if (rctx >= 0)
perf_swevent_put_recursion_context(rctx);
}
#ifdef CONFIG_GUEST_PERF_EVENTS
struct perf_guest_info_callbacks __rcu *perf_guest_cbs;
DEFINE_STATIC_CALL_RET0(__perf_guest_state, *perf_guest_cbs->state);
DEFINE_STATIC_CALL_RET0(__perf_guest_get_ip, *perf_guest_cbs->get_ip);
DEFINE_STATIC_CALL_RET0(__perf_guest_handle_intel_pt_intr, *perf_guest_cbs->handle_intel_pt_intr);
void perf_register_guest_info_callbacks(struct perf_guest_info_callbacks *cbs)
{
if (WARN_ON_ONCE(rcu_access_pointer(perf_guest_cbs)))
return;
rcu_assign_pointer(perf_guest_cbs, cbs);
static_call_update(__perf_guest_state, cbs->state);
static_call_update(__perf_guest_get_ip, cbs->get_ip);
/* Implementing ->handle_intel_pt_intr is optional. */
if (cbs->handle_intel_pt_intr)
static_call_update(__perf_guest_handle_intel_pt_intr,
cbs->handle_intel_pt_intr);
}
EXPORT_SYMBOL_GPL(perf_register_guest_info_callbacks);
void perf_unregister_guest_info_callbacks(struct perf_guest_info_callbacks *cbs)
{
if (WARN_ON_ONCE(rcu_access_pointer(perf_guest_cbs) != cbs))
return;
rcu_assign_pointer(perf_guest_cbs, NULL);
static_call_update(__perf_guest_state, (void *)&__static_call_return0);
static_call_update(__perf_guest_get_ip, (void *)&__static_call_return0);
static_call_update(__perf_guest_handle_intel_pt_intr,
(void *)&__static_call_return0);
synchronize_rcu();
}
EXPORT_SYMBOL_GPL(perf_unregister_guest_info_callbacks);
#endif
static bool should_sample_guest(struct perf_event *event)
{
return !event->attr.exclude_guest && perf_guest_state();
}
unsigned long perf_misc_flags(struct perf_event *event,
struct pt_regs *regs)
{
if (should_sample_guest(event))
return perf_arch_guest_misc_flags(regs);
return perf_arch_misc_flags(regs);
}
unsigned long perf_instruction_pointer(struct perf_event *event,
struct pt_regs *regs)
{
if (should_sample_guest(event))
return perf_guest_get_ip();
return perf_arch_instruction_pointer(regs);
}
static void
perf_output_sample_regs(struct perf_output_handle *handle,
struct pt_regs *regs, u64 mask)
{
int bit;
DECLARE_BITMAP(_mask, 64);
bitmap_from_u64(_mask, mask);
for_each_set_bit(bit, _mask, sizeof(mask) * BITS_PER_BYTE) {
u64 val;
val = perf_reg_value(regs, bit);
perf_output_put(handle, val);
}
}
static void perf_sample_regs_user(struct perf_regs *regs_user,
struct pt_regs *regs)
{
if (user_mode(regs)) {
regs_user->abi = perf_reg_abi(current);
regs_user->regs = regs;
} else if (!(current->flags & (PF_KTHREAD | PF_USER_WORKER))) {
perf_get_regs_user(regs_user, regs);
} else {
regs_user->abi = PERF_SAMPLE_REGS_ABI_NONE;
regs_user->regs = NULL;
}
}
static void perf_sample_regs_intr(struct perf_regs *regs_intr,
struct pt_regs *regs)
{
regs_intr->regs = regs;
regs_intr->abi = perf_reg_abi(current);
}
/*
* Get remaining task size from user stack pointer.
*
* It'd be better to take stack vma map and limit this more
* precisely, but there's no way to get it safely under interrupt,
* so using TASK_SIZE as limit.
*/
static u64 perf_ustack_task_size(struct pt_regs *regs)
{
unsigned long addr = perf_user_stack_pointer(regs);
if (!addr || addr >= TASK_SIZE)
return 0;
return TASK_SIZE - addr;
}
static u16
perf_sample_ustack_size(u16 stack_size, u16 header_size,
struct pt_regs *regs)
{
u64 task_size;
/* No regs, no stack pointer, no dump. */
if (!regs)
return 0;
/* No mm, no stack, no dump. */
if (!current->mm)
return 0;
/*
* Check if we fit in with the requested stack size into the:
* - TASK_SIZE
* If we don't, we limit the size to the TASK_SIZE.
*
* - remaining sample size
* If we don't, we customize the stack size to
* fit in to the remaining sample size.
*/
task_size = min((u64) USHRT_MAX, perf_ustack_task_size(regs));
stack_size = min(stack_size, (u16) task_size);
/* Current header size plus static size and dynamic size. */
header_size += 2 * sizeof(u64);
/* Do we fit in with the current stack dump size? */
if ((u16) (header_size + stack_size) < header_size) {
/*
* If we overflow the maximum size for the sample,
* we customize the stack dump size to fit in.
*/
stack_size = USHRT_MAX - header_size - sizeof(u64);
stack_size = round_up(stack_size, sizeof(u64));
}
return stack_size;
}
static void
perf_output_sample_ustack(struct perf_output_handle *handle, u64 dump_size,
struct pt_regs *regs)
{
/* Case of a kernel thread, nothing to dump */
if (!regs) {
u64 size = 0;
perf_output_put(handle, size);
} else {
unsigned long sp;
unsigned int rem;
u64 dyn_size;
/*
* We dump:
* static size
* - the size requested by user or the best one we can fit
* in to the sample max size
* data
* - user stack dump data
* dynamic size
* - the actual dumped size
*/
/* Static size. */
perf_output_put(handle, dump_size);
/* Data. */
sp = perf_user_stack_pointer(regs);
rem = __output_copy_user(handle, (void *) sp, dump_size);
dyn_size = dump_size - rem;
perf_output_skip(handle, rem);
/* Dynamic size. */
perf_output_put(handle, dyn_size);
}
}
static unsigned long perf_prepare_sample_aux(struct perf_event *event,
struct perf_sample_data *data,
size_t size)
{
struct perf_event *sampler = event->aux_event;
struct perf_buffer *rb;
data->aux_size = 0;
if (!sampler)
goto out;
if (WARN_ON_ONCE(READ_ONCE(sampler->state) != PERF_EVENT_STATE_ACTIVE))
goto out;
if (WARN_ON_ONCE(READ_ONCE(sampler->oncpu) != smp_processor_id()))
goto out;
rb = ring_buffer_get(sampler);
if (!rb)
goto out;
/*
* If this is an NMI hit inside sampling code, don't take
* the sample. See also perf_aux_sample_output().
*/
if (READ_ONCE(rb->aux_in_sampling)) {
data->aux_size = 0;
} else {
size = min_t(size_t, size, perf_aux_size(rb));
data->aux_size = ALIGN(size, sizeof(u64));
}
ring_buffer_put(rb);
out:
return data->aux_size;
}
static long perf_pmu_snapshot_aux(struct perf_buffer *rb,
struct perf_event *event,
struct perf_output_handle *handle,
unsigned long size)
{
unsigned long flags;
long ret;
/*
* Normal ->start()/->stop() callbacks run in IRQ mode in scheduler
* paths. If we start calling them in NMI context, they may race with
* the IRQ ones, that is, for example, re-starting an event that's just
* been stopped, which is why we're using a separate callback that
* doesn't change the event state.
*
* IRQs need to be disabled to prevent IPIs from racing with us.
*/
local_irq_save(flags);
/*
* Guard against NMI hits inside the critical section;
* see also perf_prepare_sample_aux().
*/
WRITE_ONCE(rb->aux_in_sampling, 1);
barrier();
ret = event->pmu->snapshot_aux(event, handle, size);
barrier();
WRITE_ONCE(rb->aux_in_sampling, 0);
local_irq_restore(flags);
return ret;
}
static void perf_aux_sample_output(struct perf_event *event,
struct perf_output_handle *handle,
struct perf_sample_data *data)
{
struct perf_event *sampler = event->aux_event;
struct perf_buffer *rb;
unsigned long pad;
long size;
if (WARN_ON_ONCE(!sampler || !data->aux_size))
return;
rb = ring_buffer_get(sampler);
if (!rb)
return;
size = perf_pmu_snapshot_aux(rb, sampler, handle, data->aux_size);
/*
* An error here means that perf_output_copy() failed (returned a
* non-zero surplus that it didn't copy), which in its current
* enlightened implementation is not possible. If that changes, we'd
* like to know.
*/
if (WARN_ON_ONCE(size < 0))
goto out_put;
/*
* The pad comes from ALIGN()ing data->aux_size up to u64 in
* perf_prepare_sample_aux(), so should not be more than that.
*/
pad = data->aux_size - size;
if (WARN_ON_ONCE(pad >= sizeof(u64)))
pad = 8;
if (pad) {
u64 zero = 0;
perf_output_copy(handle, &zero, pad);
}
out_put:
ring_buffer_put(rb);
}
/*
* A set of common sample data types saved even for non-sample records
* when event->attr.sample_id_all is set.
*/
#define PERF_SAMPLE_ID_ALL (PERF_SAMPLE_TID | PERF_SAMPLE_TIME | \
PERF_SAMPLE_ID | PERF_SAMPLE_STREAM_ID | \
PERF_SAMPLE_CPU | PERF_SAMPLE_IDENTIFIER)
static void __perf_event_header__init_id(struct perf_sample_data *data,
struct perf_event *event,
u64 sample_type)
{
data->type = event->attr.sample_type;
data->sample_flags |= data->type & PERF_SAMPLE_ID_ALL;
if (sample_type & PERF_SAMPLE_TID) {
/* namespace issues */
data->tid_entry.pid = perf_event_pid(event, current);
data->tid_entry.tid = perf_event_tid(event, current);
}
if (sample_type & PERF_SAMPLE_TIME)
data->time = perf_event_clock(event);
if (sample_type & (PERF_SAMPLE_ID | PERF_SAMPLE_IDENTIFIER))
data->id = primary_event_id(event);
if (sample_type & PERF_SAMPLE_STREAM_ID)
data->stream_id = event->id;
if (sample_type & PERF_SAMPLE_CPU) {
data->cpu_entry.cpu = raw_smp_processor_id();
data->cpu_entry.reserved = 0;
}
}
void perf_event_header__init_id(struct perf_event_header *header,
struct perf_sample_data *data,
struct perf_event *event)
{
if (event->attr.sample_id_all) {
header->size += event->id_header_size;
__perf_event_header__init_id(data, event, event->attr.sample_type);
}
}
static void __perf_event__output_id_sample(struct perf_output_handle *handle,
struct perf_sample_data *data)
{
u64 sample_type = data->type;
if (sample_type & PERF_SAMPLE_TID)
perf_output_put(handle, data->tid_entry);
if (sample_type & PERF_SAMPLE_TIME)
perf_output_put(handle, data->time);
if (sample_type & PERF_SAMPLE_ID)
perf_output_put(handle, data->id);
if (sample_type & PERF_SAMPLE_STREAM_ID)
perf_output_put(handle, data->stream_id);
if (sample_type & PERF_SAMPLE_CPU)
perf_output_put(handle, data->cpu_entry);
if (sample_type & PERF_SAMPLE_IDENTIFIER)
perf_output_put(handle, data->id);
}
void perf_event__output_id_sample(struct perf_event *event,
struct perf_output_handle *handle,
struct perf_sample_data *sample)
{
if (event->attr.sample_id_all)
__perf_event__output_id_sample(handle, sample);
}
static void perf_output_read_one(struct perf_output_handle *handle,
struct perf_event *event,
u64 enabled, u64 running)
{
u64 read_format = event->attr.read_format;
u64 values[5];
int n = 0;
values[n++] = perf_event_count(event, has_inherit_and_sample_read(&event->attr));
if (read_format & PERF_FORMAT_TOTAL_TIME_ENABLED) {
values[n++] = enabled +
atomic64_read(&event->child_total_time_enabled);
}
if (read_format & PERF_FORMAT_TOTAL_TIME_RUNNING) {
values[n++] = running +
atomic64_read(&event->child_total_time_running);
}
if (read_format & PERF_FORMAT_ID)
values[n++] = primary_event_id(event);
if (read_format & PERF_FORMAT_LOST)
values[n++] = atomic64_read(&event->lost_samples);
__output_copy(handle, values, n * sizeof(u64));
}
static void perf_output_read_group(struct perf_output_handle *handle,
struct perf_event *event,
u64 enabled, u64 running)
{
struct perf_event *leader = event->group_leader, *sub;
u64 read_format = event->attr.read_format;
unsigned long flags;
u64 values[6];
int n = 0;
bool self = has_inherit_and_sample_read(&event->attr);
/*
* Disabling interrupts avoids all counter scheduling
* (context switches, timer based rotation and IPIs).
*/
local_irq_save(flags);
values[n++] = 1 + leader->nr_siblings;
if (read_format & PERF_FORMAT_TOTAL_TIME_ENABLED)
values[n++] = enabled;
if (read_format & PERF_FORMAT_TOTAL_TIME_RUNNING)
values[n++] = running;
if ((leader != event) && !handle->skip_read)
perf_pmu_read(leader);
values[n++] = perf_event_count(leader, self);
if (read_format & PERF_FORMAT_ID)
values[n++] = primary_event_id(leader);
if (read_format & PERF_FORMAT_LOST)
values[n++] = atomic64_read(&leader->lost_samples);
__output_copy(handle, values, n * sizeof(u64));
for_each_sibling_event(sub, leader) {
n = 0;
if ((sub != event) && !handle->skip_read)
perf_pmu_read(sub);
values[n++] = perf_event_count(sub, self);
if (read_format & PERF_FORMAT_ID)
values[n++] = primary_event_id(sub);
if (read_format & PERF_FORMAT_LOST)
values[n++] = atomic64_read(&sub->lost_samples);
__output_copy(handle, values, n * sizeof(u64));
}
local_irq_restore(flags);
}
#define PERF_FORMAT_TOTAL_TIMES (PERF_FORMAT_TOTAL_TIME_ENABLED|\
PERF_FORMAT_TOTAL_TIME_RUNNING)
/*
* XXX PERF_SAMPLE_READ vs inherited events seems difficult.
*
* The problem is that its both hard and excessively expensive to iterate the
* child list, not to mention that its impossible to IPI the children running
* on another CPU, from interrupt/NMI context.
*
* Instead the combination of PERF_SAMPLE_READ and inherit will track per-thread
* counts rather than attempting to accumulate some value across all children on
* all cores.
*/
static void perf_output_read(struct perf_output_handle *handle,
struct perf_event *event)
{
u64 enabled = 0, running = 0, now;
u64 read_format = event->attr.read_format;
/*
* compute total_time_enabled, total_time_running
* based on snapshot values taken when the event
* was last scheduled in.
*
* we cannot simply called update_context_time()
* because of locking issue as we are called in
* NMI context
*/
if (read_format & PERF_FORMAT_TOTAL_TIMES)
calc_timer_values(event, &now, &enabled, &running);
if (event->attr.read_format & PERF_FORMAT_GROUP)
perf_output_read_group(handle, event, enabled, running);
else
perf_output_read_one(handle, event, enabled, running);
}
void perf_output_sample(struct perf_output_handle *handle,
struct perf_event_header *header,
struct perf_sample_data *data,
struct perf_event *event)
{
u64 sample_type = data->type;
if (data->sample_flags & PERF_SAMPLE_READ)
handle->skip_read = 1;
perf_output_put(handle, *header);
if (sample_type & PERF_SAMPLE_IDENTIFIER)
perf_output_put(handle, data->id);
if (sample_type & PERF_SAMPLE_IP)
perf_output_put(handle, data->ip);
if (sample_type & PERF_SAMPLE_TID)
perf_output_put(handle, data->tid_entry);
if (sample_type & PERF_SAMPLE_TIME)
perf_output_put(handle, data->time);
if (sample_type & PERF_SAMPLE_ADDR)
perf_output_put(handle, data->addr);
if (sample_type & PERF_SAMPLE_ID)
perf_output_put(handle, data->id);
if (sample_type & PERF_SAMPLE_STREAM_ID)
perf_output_put(handle, data->stream_id);
if (sample_type & PERF_SAMPLE_CPU)
perf_output_put(handle, data->cpu_entry);
if (sample_type & PERF_SAMPLE_PERIOD)
perf_output_put(handle, data->period);
if (sample_type & PERF_SAMPLE_READ)
perf_output_read(handle, event);
if (sample_type & PERF_SAMPLE_CALLCHAIN) {
int size = 1;
size += data->callchain->nr;
size *= sizeof(u64);
__output_copy(handle, data->callchain, size);
}
if (sample_type & PERF_SAMPLE_RAW) {
struct perf_raw_record *raw = data->raw;
if (raw) {
struct perf_raw_frag *frag = &raw->frag;
perf_output_put(handle, raw->size);
do {
if (frag->copy) {
__output_custom(handle, frag->copy,
frag->data, frag->size);
} else {
__output_copy(handle, frag->data,
frag->size);
}
if (perf_raw_frag_last(frag))
break;
frag = frag->next;
} while (1);
if (frag->pad)
__output_skip(handle, NULL, frag->pad);
} else {
struct {
u32 size;
u32 data;
} raw = {
.size = sizeof(u32),
.data = 0,
};
perf_output_put(handle, raw);
}
}
if (sample_type & PERF_SAMPLE_BRANCH_STACK) {
if (data->br_stack) {
size_t size;
size = data->br_stack->nr
* sizeof(struct perf_branch_entry);
perf_output_put(handle, data->br_stack->nr);
if (branch_sample_hw_index(event))
perf_output_put(handle, data->br_stack->hw_idx);
perf_output_copy(handle, data->br_stack->entries, size);
/*
* Add the extension space which is appended
* right after the struct perf_branch_stack.
*/
if (data->br_stack_cntr) {
size = data->br_stack->nr * sizeof(u64);
perf_output_copy(handle, data->br_stack_cntr, size);
}
} else {
/*
* we always store at least the value of nr
*/
u64 nr = 0;
perf_output_put(handle, nr);
}
}
if (sample_type & PERF_SAMPLE_REGS_USER) {
u64 abi = data->regs_user.abi;
/*
* If there are no regs to dump, notice it through
* first u64 being zero (PERF_SAMPLE_REGS_ABI_NONE).
*/
perf_output_put(handle, abi);
if (abi) {
u64 mask = event->attr.sample_regs_user;
perf_output_sample_regs(handle,
data->regs_user.regs,
mask);
}
}
if (sample_type & PERF_SAMPLE_STACK_USER) {
perf_output_sample_ustack(handle,
data->stack_user_size,
data->regs_user.regs);
}
if (sample_type & PERF_SAMPLE_WEIGHT_TYPE)
perf_output_put(handle, data->weight.full);
if (sample_type & PERF_SAMPLE_DATA_SRC)
perf_output_put(handle, data->data_src.val);
if (sample_type & PERF_SAMPLE_TRANSACTION)
perf_output_put(handle, data->txn);
if (sample_type & PERF_SAMPLE_REGS_INTR) {
u64 abi = data->regs_intr.abi;
/*
* If there are no regs to dump, notice it through
* first u64 being zero (PERF_SAMPLE_REGS_ABI_NONE).
*/
perf_output_put(handle, abi);
if (abi) {
u64 mask = event->attr.sample_regs_intr;
perf_output_sample_regs(handle,
data->regs_intr.regs,
mask);
}
}
if (sample_type & PERF_SAMPLE_PHYS_ADDR)
perf_output_put(handle, data->phys_addr);
if (sample_type & PERF_SAMPLE_CGROUP)
perf_output_put(handle, data->cgroup);
if (sample_type & PERF_SAMPLE_DATA_PAGE_SIZE)
perf_output_put(handle, data->data_page_size);
if (sample_type & PERF_SAMPLE_CODE_PAGE_SIZE)
perf_output_put(handle, data->code_page_size);
if (sample_type & PERF_SAMPLE_AUX) {
perf_output_put(handle, data->aux_size);
if (data->aux_size)
perf_aux_sample_output(event, handle, data);
}
if (!event->attr.watermark) {
int wakeup_events = event->attr.wakeup_events;
if (wakeup_events) {
struct perf_buffer *rb = handle->rb;
int events = local_inc_return(&rb->events);
if (events >= wakeup_events) {
local_sub(wakeup_events, &rb->events);
local_inc(&rb->wakeup);
}
}
}
}
static u64 perf_virt_to_phys(u64 virt)
{
u64 phys_addr = 0;
if (!virt)
return 0;
if (virt >= TASK_SIZE) {
/* If it's vmalloc()d memory, leave phys_addr as 0 */
if (virt_addr_valid((void *)(uintptr_t)virt) &&
!(virt >= VMALLOC_START && virt < VMALLOC_END))
phys_addr = (u64)virt_to_phys((void *)(uintptr_t)virt);
} else {
/*
* Walking the pages tables for user address.
* Interrupts are disabled, so it prevents any tear down
* of the page tables.
* Try IRQ-safe get_user_page_fast_only first.
* If failed, leave phys_addr as 0.
*/
if (!(current->flags & (PF_KTHREAD | PF_USER_WORKER))) {
struct page *p;
pagefault_disable();
if (get_user_page_fast_only(virt, 0, &p)) {
phys_addr = page_to_phys(p) + virt % PAGE_SIZE;
put_page(p);
}
pagefault_enable();
}
}
return phys_addr;
}
/*
* Return the pagetable size of a given virtual address.
*/
static u64 perf_get_pgtable_size(struct mm_struct *mm, unsigned long addr)
{
u64 size = 0;
#ifdef CONFIG_HAVE_GUP_FAST
pgd_t *pgdp, pgd;
p4d_t *p4dp, p4d;
pud_t *pudp, pud;
pmd_t *pmdp, pmd;
pte_t *ptep, pte;
pgdp = pgd_offset(mm, addr);
pgd = READ_ONCE(*pgdp);
if (pgd_none(pgd))
return 0;
if (pgd_leaf(pgd))
return pgd_leaf_size(pgd);
p4dp = p4d_offset_lockless(pgdp, pgd, addr);
p4d = READ_ONCE(*p4dp);
if (!p4d_present(p4d))
return 0;
if (p4d_leaf(p4d))
return p4d_leaf_size(p4d);
pudp = pud_offset_lockless(p4dp, p4d, addr);
pud = READ_ONCE(*pudp);
if (!pud_present(pud))
return 0;
if (pud_leaf(pud))
return pud_leaf_size(pud);
pmdp = pmd_offset_lockless(pudp, pud, addr);
again:
pmd = pmdp_get_lockless(pmdp);
if (!pmd_present(pmd))
return 0;
if (pmd_leaf(pmd))
return pmd_leaf_size(pmd);
ptep = pte_offset_map(&pmd, addr);
if (!ptep)
goto again;
pte = ptep_get_lockless(ptep);
if (pte_present(pte))
size = __pte_leaf_size(pmd, pte);
pte_unmap(ptep);
#endif /* CONFIG_HAVE_GUP_FAST */
return size;
}
static u64 perf_get_page_size(unsigned long addr)
{
struct mm_struct *mm;
unsigned long flags;
u64 size;
if (!addr)
return 0;
/*
* Software page-table walkers must disable IRQs,
* which prevents any tear down of the page tables.
*/
local_irq_save(flags);
mm = current->mm;
if (!mm) {
/*
* For kernel threads and the like, use init_mm so that
* we can find kernel memory.
*/
mm = &init_mm;
}
size = perf_get_pgtable_size(mm, addr);
local_irq_restore(flags);
return size;
}
static struct perf_callchain_entry __empty_callchain = { .nr = 0, };
static struct unwind_work perf_unwind_work;
struct perf_callchain_entry *
perf_callchain(struct perf_event *event, struct pt_regs *regs)
{
bool kernel = !event->attr.exclude_callchain_kernel;
bool user = !event->attr.exclude_callchain_user &&
!(current->flags & (PF_KTHREAD | PF_USER_WORKER));
/* Disallow cross-task user callchains. */
bool crosstask = event->ctx->task && event->ctx->task != current;
bool defer_user = IS_ENABLED(CONFIG_UNWIND_USER) && user &&
event->attr.defer_callchain;
const u32 max_stack = event->attr.sample_max_stack;
struct perf_callchain_entry *callchain;
u64 defer_cookie;
if (!current->mm)
user = false;
if (!kernel && !user)
return &__empty_callchain;
if (!(user && defer_user && !crosstask &&
unwind_deferred_request(&perf_unwind_work, &defer_cookie) >= 0))
defer_cookie = 0;
callchain = get_perf_callchain(regs, kernel, user, max_stack,
crosstask, true, defer_cookie);
return callchain ?: &__empty_callchain;
}
static __always_inline u64 __cond_set(u64 flags, u64 s, u64 d)
{
return d * !!(flags & s);
}
void perf_prepare_sample(struct perf_sample_data *data,
struct perf_event *event,
struct pt_regs *regs)
{
u64 sample_type = event->attr.sample_type;
u64 filtered_sample_type;
/*
* Add the sample flags that are dependent to others. And clear the
* sample flags that have already been done by the PMU driver.
*/
filtered_sample_type = sample_type;
filtered_sample_type |= __cond_set(sample_type, PERF_SAMPLE_CODE_PAGE_SIZE,
PERF_SAMPLE_IP);
filtered_sample_type |= __cond_set(sample_type, PERF_SAMPLE_DATA_PAGE_SIZE |
PERF_SAMPLE_PHYS_ADDR, PERF_SAMPLE_ADDR);
filtered_sample_type |= __cond_set(sample_type, PERF_SAMPLE_STACK_USER,
PERF_SAMPLE_REGS_USER);
filtered_sample_type &= ~data->sample_flags;
if (filtered_sample_type == 0) {
/* Make sure it has the correct data->type for output */
data->type = event->attr.sample_type;
return;
}
__perf_event_header__init_id(data, event, filtered_sample_type);
if (filtered_sample_type & PERF_SAMPLE_IP) {
data->ip = perf_instruction_pointer(event, regs);
data->sample_flags |= PERF_SAMPLE_IP;
}
if (filtered_sample_type & PERF_SAMPLE_CALLCHAIN)
perf_sample_save_callchain(data, event, regs);
if (filtered_sample_type & PERF_SAMPLE_RAW) {
data->raw = NULL;
data->dyn_size += sizeof(u64);
data->sample_flags |= PERF_SAMPLE_RAW;
}
if (filtered_sample_type & PERF_SAMPLE_BRANCH_STACK) {
data->br_stack = NULL;
data->dyn_size += sizeof(u64);
data->sample_flags |= PERF_SAMPLE_BRANCH_STACK;
}
if (filtered_sample_type & PERF_SAMPLE_REGS_USER)
perf_sample_regs_user(&data->regs_user, regs);
/*
* It cannot use the filtered_sample_type here as REGS_USER can be set
* by STACK_USER (using __cond_set() above) and we don't want to update
* the dyn_size if it's not requested by users.
*/
if ((sample_type & ~data->sample_flags) & PERF_SAMPLE_REGS_USER) {
/* regs dump ABI info */
int size = sizeof(u64);
if (data->regs_user.regs) {
u64 mask = event->attr.sample_regs_user;
size += hweight64(mask) * sizeof(u64);
}
data->dyn_size += size;
data->sample_flags |= PERF_SAMPLE_REGS_USER;
}
if (filtered_sample_type & PERF_SAMPLE_STACK_USER) {
/*
* Either we need PERF_SAMPLE_STACK_USER bit to be always
* processed as the last one or have additional check added
* in case new sample type is added, because we could eat
* up the rest of the sample size.
*/
u16 stack_size = event->attr.sample_stack_user;
u16 header_size = perf_sample_data_size(data, event);
u16 size = sizeof(u64);
stack_size = perf_sample_ustack_size(stack_size, header_size,
data->regs_user.regs);
/*
* If there is something to dump, add space for the dump
* itself and for the field that tells the dynamic size,
* which is how many have been actually dumped.
*/
if (stack_size)
size += sizeof(u64) + stack_size;
data->stack_user_size = stack_size;
data->dyn_size += size;
data->sample_flags |= PERF_SAMPLE_STACK_USER;
}
if (filtered_sample_type & PERF_SAMPLE_WEIGHT_TYPE) {
data->weight.full = 0;
data->sample_flags |= PERF_SAMPLE_WEIGHT_TYPE;
}
if (filtered_sample_type & PERF_SAMPLE_DATA_SRC) {
data->data_src.val = PERF_MEM_NA;
data->sample_flags |= PERF_SAMPLE_DATA_SRC;
}
if (filtered_sample_type & PERF_SAMPLE_TRANSACTION) {
data->txn = 0;
data->sample_flags |= PERF_SAMPLE_TRANSACTION;
}
if (filtered_sample_type & PERF_SAMPLE_ADDR) {
data->addr = 0;
data->sample_flags |= PERF_SAMPLE_ADDR;
}
if (filtered_sample_type & PERF_SAMPLE_REGS_INTR) {
/* regs dump ABI info */
int size = sizeof(u64);
perf_sample_regs_intr(&data->regs_intr, regs);
if (data->regs_intr.regs) {
u64 mask = event->attr.sample_regs_intr;
size += hweight64(mask) * sizeof(u64);
}
data->dyn_size += size;
data->sample_flags |= PERF_SAMPLE_REGS_INTR;
}
if (filtered_sample_type & PERF_SAMPLE_PHYS_ADDR) {
data->phys_addr = perf_virt_to_phys(data->addr);
data->sample_flags |= PERF_SAMPLE_PHYS_ADDR;
}
#ifdef CONFIG_CGROUP_PERF
if (filtered_sample_type & PERF_SAMPLE_CGROUP) {
struct cgroup *cgrp;
/* protected by RCU */
cgrp = task_css_check(current, perf_event_cgrp_id, 1)->cgroup;
data->cgroup = cgroup_id(cgrp);
data->sample_flags |= PERF_SAMPLE_CGROUP;
}
#endif
/*
* PERF_DATA_PAGE_SIZE requires PERF_SAMPLE_ADDR. If the user doesn't
* require PERF_SAMPLE_ADDR, kernel implicitly retrieve the data->addr,
* but the value will not dump to the userspace.
*/
if (filtered_sample_type & PERF_SAMPLE_DATA_PAGE_SIZE) {
data->data_page_size = perf_get_page_size(data->addr);
data->sample_flags |= PERF_SAMPLE_DATA_PAGE_SIZE;
}
if (filtered_sample_type & PERF_SAMPLE_CODE_PAGE_SIZE) {
data->code_page_size = perf_get_page_size(data->ip);
data->sample_flags |= PERF_SAMPLE_CODE_PAGE_SIZE;
}
if (filtered_sample_type & PERF_SAMPLE_AUX) {
u64 size;
u16 header_size = perf_sample_data_size(data, event);
header_size += sizeof(u64); /* size */
/*
* Given the 16bit nature of header::size, an AUX sample can
* easily overflow it, what with all the preceding sample bits.
* Make sure this doesn't happen by using up to U16_MAX bytes
* per sample in total (rounded down to 8 byte boundary).
*/
size = min_t(size_t, U16_MAX - header_size,
event->attr.aux_sample_size);
size = rounddown(size, 8);
size = perf_prepare_sample_aux(event, data, size);
WARN_ON_ONCE(size + header_size > U16_MAX);
data->dyn_size += size + sizeof(u64); /* size above */
data->sample_flags |= PERF_SAMPLE_AUX;
}
}
void perf_prepare_header(struct perf_event_header *header,
struct perf_sample_data *data,
struct perf_event *event,
struct pt_regs *regs)
{
header->type = PERF_RECORD_SAMPLE;
header->size = perf_sample_data_size(data, event);
header->misc = perf_misc_flags(event, regs);
/*
* If you're adding more sample types here, you likely need to do
* something about the overflowing header::size, like repurpose the
* lowest 3 bits of size, which should be always zero at the moment.
* This raises a more important question, do we really need 512k sized
* samples and why, so good argumentation is in order for whatever you
* do here next.
*/
WARN_ON_ONCE(header->size & 7);
}
static void __perf_event_aux_pause(struct perf_event *event, bool pause)
{
if (pause) {
if (!event->hw.aux_paused) {
event->hw.aux_paused = 1;
event->pmu->stop(event, PERF_EF_PAUSE);
}
} else {
if (event->hw.aux_paused) {
event->hw.aux_paused = 0;
event->pmu->start(event, PERF_EF_RESUME);
}
}
}
static void perf_event_aux_pause(struct perf_event *event, bool pause)
{
struct perf_buffer *rb;
if (WARN_ON_ONCE(!event))
return;
rb = ring_buffer_get(event);
if (!rb)
return;
scoped_guard (irqsave) {
/*
* Guard against self-recursion here. Another event could trip
* this same from NMI context.
*/
if (READ_ONCE(rb->aux_in_pause_resume))
break;
WRITE_ONCE(rb->aux_in_pause_resume, 1);
barrier();
__perf_event_aux_pause(event, pause);
barrier();
WRITE_ONCE(rb->aux_in_pause_resume, 0);
}
ring_buffer_put(rb);
}
static __always_inline int
__perf_event_output(struct perf_event *event,
struct perf_sample_data *data,
struct pt_regs *regs,
int (*output_begin)(struct perf_output_handle *,
struct perf_sample_data *,
struct perf_event *,
unsigned int))
{
struct perf_output_handle handle;
struct perf_event_header header;
int err;
/* protect the callchain buffers */
rcu_read_lock();
perf_prepare_sample(data, event, regs);
perf_prepare_header(&header, data, event, regs);
err = output_begin(&handle, data, event, header.size);
if (err)
goto exit;
perf_output_sample(&handle, &header, data, event);
perf_output_end(&handle);
exit:
rcu_read_unlock();
return err;
}
void
perf_event_output_forward(struct perf_event *event,
struct perf_sample_data *data,
struct pt_regs *regs)
{
__perf_event_output(event, data, regs, perf_output_begin_forward);
}
void
perf_event_output_backward(struct perf_event *event,
struct perf_sample_data *data,
struct pt_regs *regs)
{
__perf_event_output(event, data, regs, perf_output_begin_backward);
}
int
perf_event_output(struct perf_event *event,
struct perf_sample_data *data,
struct pt_regs *regs)
{
return __perf_event_output(event, data, regs, perf_output_begin);
}
/*
* read event_id
*/
struct perf_read_event {
struct perf_event_header header;
u32 pid;
u32 tid;
};
static void
perf_event_read_event(struct perf_event *event,
struct task_struct *task)
{
struct perf_output_handle handle;
struct perf_sample_data sample;
struct perf_read_event read_event = {
.header = {
.type = PERF_RECORD_READ,
.misc = 0,
.size = sizeof(read_event) + event->read_size,
},
.pid = perf_event_pid(event, task),
.tid = perf_event_tid(event, task),
};
int ret;
perf_event_header__init_id(&read_event.header, &sample, event);
ret = perf_output_begin(&handle, &sample, event, read_event.header.size);
if (ret)
return;
perf_output_put(&handle, read_event);
perf_output_read(&handle, event);
perf_event__output_id_sample(event, &handle, &sample);
perf_output_end(&handle);
}
typedef void (perf_iterate_f)(struct perf_event *event, void *data);
static void
perf_iterate_ctx(struct perf_event_context *ctx,
perf_iterate_f output,
void *data, bool all)
{
struct perf_event *event;
list_for_each_entry_rcu(event, &ctx->event_list, event_entry) {
if (!all) {
if (event->state < PERF_EVENT_STATE_INACTIVE)
continue;
if (!event_filter_match(event))
continue;
}
output(event, data);
}
}
static void perf_iterate_sb_cpu(perf_iterate_f output, void *data)
{
struct pmu_event_list *pel = this_cpu_ptr(&pmu_sb_events);
struct perf_event *event;
list_for_each_entry_rcu(event, &pel->list, sb_list) {
/*
* Skip events that are not fully formed yet; ensure that
* if we observe event->ctx, both event and ctx will be
* complete enough. See perf_install_in_context().
*/
if (!smp_load_acquire(&event->ctx))
continue;
if (event->state < PERF_EVENT_STATE_INACTIVE)
continue;
if (!event_filter_match(event))
continue;
output(event, data);
}
}
/*
* Iterate all events that need to receive side-band events.
*
* For new callers; ensure that account_pmu_sb_event() includes
* your event, otherwise it might not get delivered.
*/
static void
perf_iterate_sb(perf_iterate_f output, void *data,
struct perf_event_context *task_ctx)
{
struct perf_event_context *ctx;
rcu_read_lock();
preempt_disable();
/*
* If we have task_ctx != NULL we only notify the task context itself.
* The task_ctx is set only for EXIT events before releasing task
* context.
*/
if (task_ctx) {
perf_iterate_ctx(task_ctx, output, data, false);
goto done;
}
perf_iterate_sb_cpu(output, data);
ctx = rcu_dereference(current->perf_event_ctxp);
if (ctx)
perf_iterate_ctx(ctx, output, data, false);
done:
preempt_enable();
rcu_read_unlock();
}
/*
* Clear all file-based filters at exec, they'll have to be
* re-instated when/if these objects are mmapped again.
*/
static void perf_event_addr_filters_exec(struct perf_event *event, void *data)
{
struct perf_addr_filters_head *ifh = perf_event_addr_filters(event);
struct perf_addr_filter *filter;
unsigned int restart = 0, count = 0;
unsigned long flags;
if (!has_addr_filter(event))
return;
raw_spin_lock_irqsave(&ifh->lock, flags);
list_for_each_entry(filter, &ifh->list, entry) {
if (filter->path.dentry) {
event->addr_filter_ranges[count].start = 0;
event->addr_filter_ranges[count].size = 0;
restart++;
}
count++;
}
if (restart)
event->addr_filters_gen++;
raw_spin_unlock_irqrestore(&ifh->lock, flags);
if (restart)
perf_event_stop(event, 1);
}
void perf_event_exec(void)
{
struct perf_event_context *ctx;
ctx = perf_pin_task_context(current);
if (!ctx)
return;
perf_event_enable_on_exec(ctx);
perf_event_remove_on_exec(ctx);
scoped_guard(rcu)
perf_iterate_ctx(ctx, perf_event_addr_filters_exec, NULL, true);
perf_unpin_context(ctx);
put_ctx(ctx);
}
struct remote_output {
struct perf_buffer *rb;
int err;
};
static void __perf_event_output_stop(struct perf_event *event, void *data)
{
struct perf_event *parent = event->parent;
struct remote_output *ro = data;
struct perf_buffer *rb = ro->rb;
struct stop_event_data sd = {
.event = event,
};
if (!has_aux(event))
return;
if (!parent)
parent = event;
/*
* In case of inheritance, it will be the parent that links to the
* ring-buffer, but it will be the child that's actually using it.
*
* We are using event::rb to determine if the event should be stopped,
* however this may race with ring_buffer_attach() (through set_output),
* which will make us skip the event that actually needs to be stopped.
* So ring_buffer_attach() has to stop an aux event before re-assigning
* its rb pointer.
*/
if (rcu_dereference(parent->rb) == rb)
ro->err = __perf_event_stop(&sd);
}
static int __perf_pmu_output_stop(void *info)
{
struct perf_event *event = info;
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
struct remote_output ro = {
.rb = event->rb,
};
rcu_read_lock();
perf_iterate_ctx(&cpuctx->ctx, __perf_event_output_stop, &ro, false);
if (cpuctx->task_ctx)
perf_iterate_ctx(cpuctx->task_ctx, __perf_event_output_stop,
&ro, false);
rcu_read_unlock();
return ro.err;
}
static void perf_pmu_output_stop(struct perf_event *event)
{
struct perf_event *iter;
int err, cpu;
restart:
rcu_read_lock();
list_for_each_entry_rcu(iter, &event->rb->event_list, rb_entry) {
/*
* For per-CPU events, we need to make sure that neither they
* nor their children are running; for cpu==-1 events it's
* sufficient to stop the event itself if it's active, since
* it can't have children.
*/
cpu = iter->cpu;
if (cpu == -1)
cpu = READ_ONCE(iter->oncpu);
if (cpu == -1)
continue;
err = cpu_function_call(cpu, __perf_pmu_output_stop, event);
if (err == -EAGAIN) {
rcu_read_unlock();
goto restart;
}
}
rcu_read_unlock();
}
/*
* task tracking -- fork/exit
*
* enabled by: attr.comm | attr.mmap | attr.mmap2 | attr.mmap_data | attr.task
*/
struct perf_task_event {
struct task_struct *task;
struct perf_event_context *task_ctx;
struct {
struct perf_event_header header;
u32 pid;
u32 ppid;
u32 tid;
u32 ptid;
u64 time;
} event_id;
};
static int perf_event_task_match(struct perf_event *event)
{
return event->attr.comm || event->attr.mmap ||
event->attr.mmap2 || event->attr.mmap_data ||
event->attr.task;
}
static void perf_event_task_output(struct perf_event *event,
void *data)
{
struct perf_task_event *task_event = data;
struct perf_output_handle handle;
struct perf_sample_data sample;
struct task_struct *task = task_event->task;
int ret, size = task_event->event_id.header.size;
if (!perf_event_task_match(event))
return;
perf_event_header__init_id(&task_event->event_id.header, &sample, event);
ret = perf_output_begin(&handle, &sample, event,
task_event->event_id.header.size);
if (ret)
goto out;
task_event->event_id.pid = perf_event_pid(event, task);
task_event->event_id.tid = perf_event_tid(event, task);
if (task_event->event_id.header.type == PERF_RECORD_EXIT) {
task_event->event_id.ppid = perf_event_pid(event,
task->real_parent);
task_event->event_id.ptid = perf_event_pid(event,
task->real_parent);
} else { /* PERF_RECORD_FORK */
task_event->event_id.ppid = perf_event_pid(event, current);
task_event->event_id.ptid = perf_event_tid(event, current);
}
task_event->event_id.time = perf_event_clock(event);
perf_output_put(&handle, task_event->event_id);
perf_event__output_id_sample(event, &handle, &sample);
perf_output_end(&handle);
out:
task_event->event_id.header.size = size;
}
static void perf_event_task(struct task_struct *task,
struct perf_event_context *task_ctx,
int new)
{
struct perf_task_event task_event;
if (!atomic_read(&nr_comm_events) &&
!atomic_read(&nr_mmap_events) &&
!atomic_read(&nr_task_events))
return;
task_event = (struct perf_task_event){
.task = task,
.task_ctx = task_ctx,
.event_id = {
.header = {
.type = new ? PERF_RECORD_FORK : PERF_RECORD_EXIT,
.misc = 0,
.size = sizeof(task_event.event_id),
},
/* .pid */
/* .ppid */
/* .tid */
/* .ptid */
/* .time */
},
};
perf_iterate_sb(perf_event_task_output,
&task_event,
task_ctx);
}
/*
* Allocate data for a new task when profiling system-wide
* events which require PMU specific data
*/
static void
perf_event_alloc_task_data(struct task_struct *child,
struct task_struct *parent)
{
struct kmem_cache *ctx_cache = NULL;
struct perf_ctx_data *cd;
if (!refcount_read(&global_ctx_data_ref))
return;
scoped_guard (rcu) {
cd = rcu_dereference(parent->perf_ctx_data);
if (cd)
ctx_cache = cd->ctx_cache;
}
if (!ctx_cache)
return;
guard(percpu_read)(&global_ctx_data_rwsem);
scoped_guard (rcu) {
cd = rcu_dereference(child->perf_ctx_data);
if (!cd) {
/*
* A system-wide event may be unaccount,
* when attaching the perf_ctx_data.
*/
if (!refcount_read(&global_ctx_data_ref))
return;
goto attach;
}
if (!cd->global) {
cd->global = 1;
refcount_inc(&cd->refcount);
}
}
return;
attach:
attach_task_ctx_data(child, ctx_cache, true);
}
void perf_event_fork(struct task_struct *task)
{
perf_event_task(task, NULL, 1);
perf_event_namespaces(task);
perf_event_alloc_task_data(task, current);
}
/*
* comm tracking
*/
struct perf_comm_event {
struct task_struct *task;
char *comm;
int comm_size;
struct {
struct perf_event_header header;
u32 pid;
u32 tid;
} event_id;
};
static int perf_event_comm_match(struct perf_event *event)
{
return event->attr.comm;
}
static void perf_event_comm_output(struct perf_event *event,
void *data)
{
struct perf_comm_event *comm_event = data;
struct perf_output_handle handle;
struct perf_sample_data sample;
int size = comm_event->event_id.header.size;
int ret;
if (!perf_event_comm_match(event))
return;
perf_event_header__init_id(&comm_event->event_id.header, &sample, event);
ret = perf_output_begin(&handle, &sample, event,
comm_event->event_id.header.size);
if (ret)
goto out;
comm_event->event_id.pid = perf_event_pid(event, comm_event->task);
comm_event->event_id.tid = perf_event_tid(event, comm_event->task);
perf_output_put(&handle, comm_event->event_id);
__output_copy(&handle, comm_event->comm,
comm_event->comm_size);
perf_event__output_id_sample(event, &handle, &sample);
perf_output_end(&handle);
out:
comm_event->event_id.header.size = size;
}
static void perf_event_comm_event(struct perf_comm_event *comm_event)
{
char comm[TASK_COMM_LEN];
unsigned int size;
memset(comm, 0, sizeof(comm));
strscpy(comm, comm_event->task->comm);
size = ALIGN(strlen(comm)+1, sizeof(u64));
comm_event->comm = comm;
comm_event->comm_size = size;
comm_event->event_id.header.size = sizeof(comm_event->event_id) + size;
perf_iterate_sb(perf_event_comm_output,
comm_event,
NULL);
}
void perf_event_comm(struct task_struct *task, bool exec)
{
struct perf_comm_event comm_event;
if (!atomic_read(&nr_comm_events))
return;
comm_event = (struct perf_comm_event){
.task = task,
/* .comm */
/* .comm_size */
.event_id = {
.header = {
.type = PERF_RECORD_COMM,
.misc = exec ? PERF_RECORD_MISC_COMM_EXEC : 0,
/* .size */
},
/* .pid */
/* .tid */
},
};
perf_event_comm_event(&comm_event);
}
/*
* namespaces tracking
*/
struct perf_namespaces_event {
struct task_struct *task;
struct {
struct perf_event_header header;
u32 pid;
u32 tid;
u64 nr_namespaces;
struct perf_ns_link_info link_info[NR_NAMESPACES];
} event_id;
};
static int perf_event_namespaces_match(struct perf_event *event)
{
return event->attr.namespaces;
}
static void perf_event_namespaces_output(struct perf_event *event,
void *data)
{
struct perf_namespaces_event *namespaces_event = data;
struct perf_output_handle handle;
struct perf_sample_data sample;
u16 header_size = namespaces_event->event_id.header.size;
int ret;
if (!perf_event_namespaces_match(event))
return;
perf_event_header__init_id(&namespaces_event->event_id.header,
&sample, event);
ret = perf_output_begin(&handle, &sample, event,
namespaces_event->event_id.header.size);
if (ret)
goto out;
namespaces_event->event_id.pid = perf_event_pid(event,
namespaces_event->task);
namespaces_event->event_id.tid = perf_event_tid(event,
namespaces_event->task);
perf_output_put(&handle, namespaces_event->event_id);
perf_event__output_id_sample(event, &handle, &sample);
perf_output_end(&handle);
out:
namespaces_event->event_id.header.size = header_size;
}
static void perf_fill_ns_link_info(struct perf_ns_link_info *ns_link_info,
struct task_struct *task,
const struct proc_ns_operations *ns_ops)
{
struct path ns_path;
struct inode *ns_inode;
int error;
error = ns_get_path(&ns_path, task, ns_ops);
if (!error) {
ns_inode = ns_path.dentry->d_inode;
ns_link_info->dev = new_encode_dev(ns_inode->i_sb->s_dev);
ns_link_info->ino = ns_inode->i_ino;
path_put(&ns_path);
}
}
void perf_event_namespaces(struct task_struct *task)
{
struct perf_namespaces_event namespaces_event;
struct perf_ns_link_info *ns_link_info;
if (!atomic_read(&nr_namespaces_events))
return;
namespaces_event = (struct perf_namespaces_event){
.task = task,
.event_id = {
.header = {
.type = PERF_RECORD_NAMESPACES,
.misc = 0,
.size = sizeof(namespaces_event.event_id),
},
/* .pid */
/* .tid */
.nr_namespaces = NR_NAMESPACES,
/* .link_info[NR_NAMESPACES] */
},
};
ns_link_info = namespaces_event.event_id.link_info;
perf_fill_ns_link_info(&ns_link_info[MNT_NS_INDEX],
task, &mntns_operations);
#ifdef CONFIG_USER_NS
perf_fill_ns_link_info(&ns_link_info[USER_NS_INDEX],
task, &userns_operations);
#endif
#ifdef CONFIG_NET_NS
perf_fill_ns_link_info(&ns_link_info[NET_NS_INDEX],
task, &netns_operations);
#endif
#ifdef CONFIG_UTS_NS
perf_fill_ns_link_info(&ns_link_info[UTS_NS_INDEX],
task, &utsns_operations);
#endif
#ifdef CONFIG_IPC_NS
perf_fill_ns_link_info(&ns_link_info[IPC_NS_INDEX],
task, &ipcns_operations);
#endif
#ifdef CONFIG_PID_NS
perf_fill_ns_link_info(&ns_link_info[PID_NS_INDEX],
task, &pidns_operations);
#endif
#ifdef CONFIG_CGROUPS
perf_fill_ns_link_info(&ns_link_info[CGROUP_NS_INDEX],
task, &cgroupns_operations);
#endif
perf_iterate_sb(perf_event_namespaces_output,
&namespaces_event,
NULL);
}
/*
* cgroup tracking
*/
#ifdef CONFIG_CGROUP_PERF
struct perf_cgroup_event {
char *path;
int path_size;
struct {
struct perf_event_header header;
u64 id;
char path[];
} event_id;
};
static int perf_event_cgroup_match(struct perf_event *event)
{
return event->attr.cgroup;
}
static void perf_event_cgroup_output(struct perf_event *event, void *data)
{
struct perf_cgroup_event *cgroup_event = data;
struct perf_output_handle handle;
struct perf_sample_data sample;
u16 header_size = cgroup_event->event_id.header.size;
int ret;
if (!perf_event_cgroup_match(event))
return;
perf_event_header__init_id(&cgroup_event->event_id.header,
&sample, event);
ret = perf_output_begin(&handle, &sample, event,
cgroup_event->event_id.header.size);
if (ret)
goto out;
perf_output_put(&handle, cgroup_event->event_id);
__output_copy(&handle, cgroup_event->path, cgroup_event->path_size);
perf_event__output_id_sample(event, &handle, &sample);
perf_output_end(&handle);
out:
cgroup_event->event_id.header.size = header_size;
}
static void perf_event_cgroup(struct cgroup *cgrp)
{
struct perf_cgroup_event cgroup_event;
char path_enomem[16] = "//enomem";
char *pathname;
size_t size;
if (!atomic_read(&nr_cgroup_events))
return;
cgroup_event = (struct perf_cgroup_event){
.event_id = {
.header = {
.type = PERF_RECORD_CGROUP,
.misc = 0,
.size = sizeof(cgroup_event.event_id),
},
.id = cgroup_id(cgrp),
},
};
pathname = kmalloc(PATH_MAX, GFP_KERNEL);
if (pathname == NULL) {
cgroup_event.path = path_enomem;
} else {
/* just to be sure to have enough space for alignment */
cgroup_path(cgrp, pathname, PATH_MAX - sizeof(u64));
cgroup_event.path = pathname;
}
/*
* Since our buffer works in 8 byte units we need to align our string
* size to a multiple of 8. However, we must guarantee the tail end is
* zero'd out to avoid leaking random bits to userspace.
*/
size = strlen(cgroup_event.path) + 1;
while (!IS_ALIGNED(size, sizeof(u64)))
cgroup_event.path[size++] = '\0';
cgroup_event.event_id.header.size += size;
cgroup_event.path_size = size;
perf_iterate_sb(perf_event_cgroup_output,
&cgroup_event,
NULL);
kfree(pathname);
}
#endif
/*
* mmap tracking
*/
struct perf_mmap_event {
struct vm_area_struct *vma;
const char *file_name;
int file_size;
int maj, min;
u64 ino;
u64 ino_generation;
u32 prot, flags;
u8 build_id[BUILD_ID_SIZE_MAX];
u32 build_id_size;
struct {
struct perf_event_header header;
u32 pid;
u32 tid;
u64 start;
u64 len;
u64 pgoff;
} event_id;
};
static int perf_event_mmap_match(struct perf_event *event,
void *data)
{
struct perf_mmap_event *mmap_event = data;
struct vm_area_struct *vma = mmap_event->vma;
int executable = vma->vm_flags & VM_EXEC;
return (!executable && event->attr.mmap_data) ||
(executable && (event->attr.mmap || event->attr.mmap2));
}
static void perf_event_mmap_output(struct perf_event *event,
void *data)
{
struct perf_mmap_event *mmap_event = data;
struct perf_output_handle handle;
struct perf_sample_data sample;
int size = mmap_event->event_id.header.size;
u32 type = mmap_event->event_id.header.type;
bool use_build_id;
int ret;
if (!perf_event_mmap_match(event, data))
return;
if (event->attr.mmap2) {
mmap_event->event_id.header.type = PERF_RECORD_MMAP2;
mmap_event->event_id.header.size += sizeof(mmap_event->maj);
mmap_event->event_id.header.size += sizeof(mmap_event->min);
mmap_event->event_id.header.size += sizeof(mmap_event->ino);
mmap_event->event_id.header.size += sizeof(mmap_event->ino_generation);
mmap_event->event_id.header.size += sizeof(mmap_event->prot);
mmap_event->event_id.header.size += sizeof(mmap_event->flags);
}
perf_event_header__init_id(&mmap_event->event_id.header, &sample, event);
ret = perf_output_begin(&handle, &sample, event,
mmap_event->event_id.header.size);
if (ret)
goto out;
mmap_event->event_id.pid = perf_event_pid(event, current);
mmap_event->event_id.tid = perf_event_tid(event, current);
use_build_id = event->attr.build_id && mmap_event->build_id_size;
if (event->attr.mmap2 && use_build_id)
mmap_event->event_id.header.misc |= PERF_RECORD_MISC_MMAP_BUILD_ID;
perf_output_put(&handle, mmap_event->event_id);
if (event->attr.mmap2) {
if (use_build_id) {
u8 size[4] = { (u8) mmap_event->build_id_size, 0, 0, 0 };
__output_copy(&handle, size, 4);
__output_copy(&handle, mmap_event->build_id, BUILD_ID_SIZE_MAX);
} else {
perf_output_put(&handle, mmap_event->maj);
perf_output_put(&handle, mmap_event->min);
perf_output_put(&handle, mmap_event->ino);
perf_output_put(&handle, mmap_event->ino_generation);
}
perf_output_put(&handle, mmap_event->prot);
perf_output_put(&handle, mmap_event->flags);
}
__output_copy(&handle, mmap_event->file_name,
mmap_event->file_size);
perf_event__output_id_sample(event, &handle, &sample);
perf_output_end(&handle);
out:
mmap_event->event_id.header.size = size;
mmap_event->event_id.header.type = type;
}
static void perf_event_mmap_event(struct perf_mmap_event *mmap_event)
{
struct vm_area_struct *vma = mmap_event->vma;
struct file *file = vma->vm_file;
int maj = 0, min = 0;
u64 ino = 0, gen = 0;
u32 prot = 0, flags = 0;
unsigned int size;
char tmp[16];
char *buf = NULL;
char *name = NULL;
if (vma->vm_flags & VM_READ)
prot |= PROT_READ;
if (vma->vm_flags & VM_WRITE)
prot |= PROT_WRITE;
if (vma->vm_flags & VM_EXEC)
prot |= PROT_EXEC;
if (vma->vm_flags & VM_MAYSHARE)
flags = MAP_SHARED;
else
flags = MAP_PRIVATE;
if (vma->vm_flags & VM_LOCKED)
flags |= MAP_LOCKED;
if (is_vm_hugetlb_page(vma))
flags |= MAP_HUGETLB;
if (file) {
const struct inode *inode;
dev_t dev;
buf = kmalloc(PATH_MAX, GFP_KERNEL);
if (!buf) {
name = "//enomem";
goto cpy_name;
}
/*
* d_path() works from the end of the rb backwards, so we
* need to add enough zero bytes after the string to handle
* the 64bit alignment we do later.
*/
name = d_path(file_user_path(file), buf, PATH_MAX - sizeof(u64));
if (IS_ERR(name)) {
name = "//toolong";
goto cpy_name;
}
inode = file_user_inode(vma->vm_file);
dev = inode->i_sb->s_dev;
ino = inode->i_ino;
gen = inode->i_generation;
maj = MAJOR(dev);
min = MINOR(dev);
goto got_name;
} else {
if (vma->vm_ops && vma->vm_ops->name)
name = (char *) vma->vm_ops->name(vma);
if (!name)
name = (char *)arch_vma_name(vma);
if (!name) {
if (vma_is_initial_heap(vma))
name = "[heap]";
else if (vma_is_initial_stack(vma))
name = "[stack]";
else
name = "//anon";
}
}
cpy_name:
strscpy(tmp, name);
name = tmp;
got_name:
/*
* Since our buffer works in 8 byte units we need to align our string
* size to a multiple of 8. However, we must guarantee the tail end is
* zero'd out to avoid leaking random bits to userspace.
*/
size = strlen(name)+1;
while (!IS_ALIGNED(size, sizeof(u64)))
name[size++] = '\0';
mmap_event->file_name = name;
mmap_event->file_size = size;
mmap_event->maj = maj;
mmap_event->min = min;
mmap_event->ino = ino;
mmap_event->ino_generation = gen;
mmap_event->prot = prot;
mmap_event->flags = flags;
if (!(vma->vm_flags & VM_EXEC))
mmap_event->event_id.header.misc |= PERF_RECORD_MISC_MMAP_DATA;
mmap_event->event_id.header.size = sizeof(mmap_event->event_id) + size;
if (atomic_read(&nr_build_id_events))
build_id_parse_nofault(vma, mmap_event->build_id, &mmap_event->build_id_size);
perf_iterate_sb(perf_event_mmap_output,
mmap_event,
NULL);
kfree(buf);
}
/*
* Check whether inode and address range match filter criteria.
*/
static bool perf_addr_filter_match(struct perf_addr_filter *filter,
struct file *file, unsigned long offset,
unsigned long size)
{
/* d_inode(NULL) won't be equal to any mapped user-space file */
if (!filter->path.dentry)
return false;
if (d_inode(filter->path.dentry) != file_user_inode(file))
return false;
if (filter->offset > offset + size)
return false;
if (filter->offset + filter->size < offset)
return false;
return true;
}
static bool perf_addr_filter_vma_adjust(struct perf_addr_filter *filter,
struct vm_area_struct *vma,
struct perf_addr_filter_range *fr)
{
unsigned long vma_size = vma->vm_end - vma->vm_start;
unsigned long off = vma->vm_pgoff << PAGE_SHIFT;
struct file *file = vma->vm_file;
if (!perf_addr_filter_match(filter, file, off, vma_size))
return false;
if (filter->offset < off) {
fr->start = vma->vm_start;
fr->size = min(vma_size, filter->size - (off - filter->offset));
} else {
fr->start = vma->vm_start + filter->offset - off;
fr->size = min(vma->vm_end - fr->start, filter->size);
}
return true;
}
static void __perf_addr_filters_adjust(struct perf_event *event, void *data)
{
struct perf_addr_filters_head *ifh = perf_event_addr_filters(event);
struct vm_area_struct *vma = data;
struct perf_addr_filter *filter;
unsigned int restart = 0, count = 0;
unsigned long flags;
if (!has_addr_filter(event))
return;
if (!vma->vm_file)
return;
raw_spin_lock_irqsave(&ifh->lock, flags);
list_for_each_entry(filter, &ifh->list, entry) {
if (perf_addr_filter_vma_adjust(filter, vma,
&event->addr_filter_ranges[count]))
restart++;
count++;
}
if (restart)
event->addr_filters_gen++;
raw_spin_unlock_irqrestore(&ifh->lock, flags);
if (restart)
perf_event_stop(event, 1);
}
/*
* Adjust all task's events' filters to the new vma
*/
static void perf_addr_filters_adjust(struct vm_area_struct *vma)
{
struct perf_event_context *ctx;
/*
* Data tracing isn't supported yet and as such there is no need
* to keep track of anything that isn't related to executable code:
*/
if (!(vma->vm_flags & VM_EXEC))
return;
rcu_read_lock();
ctx = rcu_dereference(current->perf_event_ctxp);
if (ctx)
perf_iterate_ctx(ctx, __perf_addr_filters_adjust, vma, true);
rcu_read_unlock();
}
void perf_event_mmap(struct vm_area_struct *vma)
{
struct perf_mmap_event mmap_event;
if (!atomic_read(&nr_mmap_events))
return;
mmap_event = (struct perf_mmap_event){
.vma = vma,
/* .file_name */
/* .file_size */
.event_id = {
.header = {
.type = PERF_RECORD_MMAP,
.misc = PERF_RECORD_MISC_USER,
/* .size */
},
/* .pid */
/* .tid */
.start = vma->vm_start,
.len = vma->vm_end - vma->vm_start,
.pgoff = (u64)vma->vm_pgoff << PAGE_SHIFT,
},
/* .maj (attr_mmap2 only) */
/* .min (attr_mmap2 only) */
/* .ino (attr_mmap2 only) */
/* .ino_generation (attr_mmap2 only) */
/* .prot (attr_mmap2 only) */
/* .flags (attr_mmap2 only) */
};
perf_addr_filters_adjust(vma);
perf_event_mmap_event(&mmap_event);
}
void perf_event_aux_event(struct perf_event *event, unsigned long head,
unsigned long size, u64 flags)
{
struct perf_output_handle handle;
struct perf_sample_data sample;
struct perf_aux_event {
struct perf_event_header header;
u64 offset;
u64 size;
u64 flags;
} rec = {
.header = {
.type = PERF_RECORD_AUX,
.misc = 0,
.size = sizeof(rec),
},
.offset = head,
.size = size,
.flags = flags,
};
int ret;
perf_event_header__init_id(&rec.header, &sample, event);
ret = perf_output_begin(&handle, &sample, event, rec.header.size);
if (ret)
return;
perf_output_put(&handle, rec);
perf_event__output_id_sample(event, &handle, &sample);
perf_output_end(&handle);
}
/*
* Lost/dropped samples logging
*/
void perf_log_lost_samples(struct perf_event *event, u64 lost)
{
struct perf_output_handle handle;
struct perf_sample_data sample;
int ret;
struct {
struct perf_event_header header;
u64 lost;
} lost_samples_event = {
.header = {
.type = PERF_RECORD_LOST_SAMPLES,
.misc = 0,
.size = sizeof(lost_samples_event),
},
.lost = lost,
};
perf_event_header__init_id(&lost_samples_event.header, &sample, event);
ret = perf_output_begin(&handle, &sample, event,
lost_samples_event.header.size);
if (ret)
return;
perf_output_put(&handle, lost_samples_event);
perf_event__output_id_sample(event, &handle, &sample);
perf_output_end(&handle);
}
/*
* context_switch tracking
*/
struct perf_switch_event {
struct task_struct *task;
struct task_struct *next_prev;
struct {
struct perf_event_header header;
u32 next_prev_pid;
u32 next_prev_tid;
} event_id;
};
static int perf_event_switch_match(struct perf_event *event)
{
return event->attr.context_switch;
}
static void perf_event_switch_output(struct perf_event *event, void *data)
{
struct perf_switch_event *se = data;
struct perf_output_handle handle;
struct perf_sample_data sample;
int ret;
if (!perf_event_switch_match(event))
return;
/* Only CPU-wide events are allowed to see next/prev pid/tid */
if (event->ctx->task) {
se->event_id.header.type = PERF_RECORD_SWITCH;
se->event_id.header.size = sizeof(se->event_id.header);
} else {
se->event_id.header.type = PERF_RECORD_SWITCH_CPU_WIDE;
se->event_id.header.size = sizeof(se->event_id);
se->event_id.next_prev_pid =
perf_event_pid(event, se->next_prev);
se->event_id.next_prev_tid =
perf_event_tid(event, se->next_prev);
}
perf_event_header__init_id(&se->event_id.header, &sample, event);
ret = perf_output_begin(&handle, &sample, event, se->event_id.header.size);
if (ret)
return;
if (event->ctx->task)
perf_output_put(&handle, se->event_id.header);
else
perf_output_put(&handle, se->event_id);
perf_event__output_id_sample(event, &handle, &sample);
perf_output_end(&handle);
}
static void perf_event_switch(struct task_struct *task,
struct task_struct *next_prev, bool sched_in)
{
struct perf_switch_event switch_event;
/* N.B. caller checks nr_switch_events != 0 */
switch_event = (struct perf_switch_event){
.task = task,
.next_prev = next_prev,
.event_id = {
.header = {
/* .type */
.misc = sched_in ? 0 : PERF_RECORD_MISC_SWITCH_OUT,
/* .size */
},
/* .next_prev_pid */
/* .next_prev_tid */
},
};
if (!sched_in && task_is_runnable(task)) {
switch_event.event_id.header.misc |=
PERF_RECORD_MISC_SWITCH_OUT_PREEMPT;
}
perf_iterate_sb(perf_event_switch_output, &switch_event, NULL);
}
/*
* IRQ throttle logging
*/
static void perf_log_throttle(struct perf_event *event, int enable)
{
struct perf_output_handle handle;
struct perf_sample_data sample;
int ret;
struct {
struct perf_event_header header;
u64 time;
u64 id;
u64 stream_id;
} throttle_event = {
.header = {
.type = PERF_RECORD_THROTTLE,
.misc = 0,
.size = sizeof(throttle_event),
},
.time = perf_event_clock(event),
.id = primary_event_id(event),
.stream_id = event->id,
};
if (enable)
throttle_event.header.type = PERF_RECORD_UNTHROTTLE;
perf_event_header__init_id(&throttle_event.header, &sample, event);
ret = perf_output_begin(&handle, &sample, event,
throttle_event.header.size);
if (ret)
return;
perf_output_put(&handle, throttle_event);
perf_event__output_id_sample(event, &handle, &sample);
perf_output_end(&handle);
}
/*
* ksymbol register/unregister tracking
*/
struct perf_ksymbol_event {
const char *name;
int name_len;
struct {
struct perf_event_header header;
u64 addr;
u32 len;
u16 ksym_type;
u16 flags;
} event_id;
};
static int perf_event_ksymbol_match(struct perf_event *event)
{
return event->attr.ksymbol;
}
static void perf_event_ksymbol_output(struct perf_event *event, void *data)
{
struct perf_ksymbol_event *ksymbol_event = data;
struct perf_output_handle handle;
struct perf_sample_data sample;
int ret;
if (!perf_event_ksymbol_match(event))
return;
perf_event_header__init_id(&ksymbol_event->event_id.header,
&sample, event);
ret = perf_output_begin(&handle, &sample, event,
ksymbol_event->event_id.header.size);
if (ret)
return;
perf_output_put(&handle, ksymbol_event->event_id);
__output_copy(&handle, ksymbol_event->name, ksymbol_event->name_len);
perf_event__output_id_sample(event, &handle, &sample);
perf_output_end(&handle);
}
void perf_event_ksymbol(u16 ksym_type, u64 addr, u32 len, bool unregister,
const char *sym)
{
struct perf_ksymbol_event ksymbol_event;
char name[KSYM_NAME_LEN];
u16 flags = 0;
int name_len;
if (!atomic_read(&nr_ksymbol_events))
return;
if (ksym_type >= PERF_RECORD_KSYMBOL_TYPE_MAX ||
ksym_type == PERF_RECORD_KSYMBOL_TYPE_UNKNOWN)
goto err;
strscpy(name, sym);
name_len = strlen(name) + 1;
while (!IS_ALIGNED(name_len, sizeof(u64)))
name[name_len++] = '\0';
BUILD_BUG_ON(KSYM_NAME_LEN % sizeof(u64));
if (unregister)
flags |= PERF_RECORD_KSYMBOL_FLAGS_UNREGISTER;
ksymbol_event = (struct perf_ksymbol_event){
.name = name,
.name_len = name_len,
.event_id = {
.header = {
.type = PERF_RECORD_KSYMBOL,
.size = sizeof(ksymbol_event.event_id) +
name_len,
},
.addr = addr,
.len = len,
.ksym_type = ksym_type,
.flags = flags,
},
};
perf_iterate_sb(perf_event_ksymbol_output, &ksymbol_event, NULL);
return;
err:
WARN_ONCE(1, "%s: Invalid KSYMBOL type 0x%x\n", __func__, ksym_type);
}
/*
* bpf program load/unload tracking
*/
struct perf_bpf_event {
struct bpf_prog *prog;
struct {
struct perf_event_header header;
u16 type;
u16 flags;
u32 id;
u8 tag[BPF_TAG_SIZE];
} event_id;
};
static int perf_event_bpf_match(struct perf_event *event)
{
return event->attr.bpf_event;
}
static void perf_event_bpf_output(struct perf_event *event, void *data)
{
struct perf_bpf_event *bpf_event = data;
struct perf_output_handle handle;
struct perf_sample_data sample;
int ret;
if (!perf_event_bpf_match(event))
return;
perf_event_header__init_id(&bpf_event->event_id.header,
&sample, event);
ret = perf_output_begin(&handle, &sample, event,
bpf_event->event_id.header.size);
if (ret)
return;
perf_output_put(&handle, bpf_event->event_id);
perf_event__output_id_sample(event, &handle, &sample);
perf_output_end(&handle);
}
static void perf_event_bpf_emit_ksymbols(struct bpf_prog *prog,
enum perf_bpf_event_type type)
{
bool unregister = type == PERF_BPF_EVENT_PROG_UNLOAD;
int i;
perf_event_ksymbol(PERF_RECORD_KSYMBOL_TYPE_BPF,
(u64)(unsigned long)prog->bpf_func,
prog->jited_len, unregister,
prog->aux->ksym.name);
for (i = 1; i < prog->aux->func_cnt; i++) {
struct bpf_prog *subprog = prog->aux->func[i];
perf_event_ksymbol(
PERF_RECORD_KSYMBOL_TYPE_BPF,
(u64)(unsigned long)subprog->bpf_func,
subprog->jited_len, unregister,
subprog->aux->ksym.name);
}
}
void perf_event_bpf_event(struct bpf_prog *prog,
enum perf_bpf_event_type type,
u16 flags)
{
struct perf_bpf_event bpf_event;
switch (type) {
case PERF_BPF_EVENT_PROG_LOAD:
case PERF_BPF_EVENT_PROG_UNLOAD:
if (atomic_read(&nr_ksymbol_events))
perf_event_bpf_emit_ksymbols(prog, type);
break;
default:
return;
}
if (!atomic_read(&nr_bpf_events))
return;
bpf_event = (struct perf_bpf_event){
.prog = prog,
.event_id = {
.header = {
.type = PERF_RECORD_BPF_EVENT,
.size = sizeof(bpf_event.event_id),
},
.type = type,
.flags = flags,
.id = prog->aux->id,
},
};
BUILD_BUG_ON(BPF_TAG_SIZE % sizeof(u64));
memcpy(bpf_event.event_id.tag, prog->tag, BPF_TAG_SIZE);
perf_iterate_sb(perf_event_bpf_output, &bpf_event, NULL);
}
struct perf_callchain_deferred_event {
struct unwind_stacktrace *trace;
struct {
struct perf_event_header header;
u64 cookie;
u64 nr;
u64 ips[];
} event;
};
static void perf_callchain_deferred_output(struct perf_event *event, void *data)
{
struct perf_callchain_deferred_event *deferred_event = data;
struct perf_output_handle handle;
struct perf_sample_data sample;
int ret, size = deferred_event->event.header.size;
if (!event->attr.defer_output)
return;
/* XXX do we really need sample_id_all for this ??? */
perf_event_header__init_id(&deferred_event->event.header, &sample, event);
ret = perf_output_begin(&handle, &sample, event,
deferred_event->event.header.size);
if (ret)
goto out;
perf_output_put(&handle, deferred_event->event);
for (int i = 0; i < deferred_event->trace->nr; i++) {
u64 entry = deferred_event->trace->entries[i];
perf_output_put(&handle, entry);
}
perf_event__output_id_sample(event, &handle, &sample);
perf_output_end(&handle);
out:
deferred_event->event.header.size = size;
}
static void perf_unwind_deferred_callback(struct unwind_work *work,
struct unwind_stacktrace *trace, u64 cookie)
{
struct perf_callchain_deferred_event deferred_event = {
.trace = trace,
.event = {
.header = {
.type = PERF_RECORD_CALLCHAIN_DEFERRED,
.misc = PERF_RECORD_MISC_USER,
.size = sizeof(deferred_event.event) +
(trace->nr * sizeof(u64)),
},
.cookie = cookie,
.nr = trace->nr,
},
};
perf_iterate_sb(perf_callchain_deferred_output, &deferred_event, NULL);
}
struct perf_text_poke_event {
const void *old_bytes;
const void *new_bytes;
size_t pad;
u16 old_len;
u16 new_len;
struct {
struct perf_event_header header;
u64 addr;
} event_id;
};
static int perf_event_text_poke_match(struct perf_event *event)
{
return event->attr.text_poke;
}
static void perf_event_text_poke_output(struct perf_event *event, void *data)
{
struct perf_text_poke_event *text_poke_event = data;
struct perf_output_handle handle;
struct perf_sample_data sample;
u64 padding = 0;
int ret;
if (!perf_event_text_poke_match(event))
return;
perf_event_header__init_id(&text_poke_event->event_id.header, &sample, event);
ret = perf_output_begin(&handle, &sample, event,
text_poke_event->event_id.header.size);
if (ret)
return;
perf_output_put(&handle, text_poke_event->event_id);
perf_output_put(&handle, text_poke_event->old_len);
perf_output_put(&handle, text_poke_event->new_len);
__output_copy(&handle, text_poke_event->old_bytes, text_poke_event->old_len);
__output_copy(&handle, text_poke_event->new_bytes, text_poke_event->new_len);
if (text_poke_event->pad)
__output_copy(&handle, &padding, text_poke_event->pad);
perf_event__output_id_sample(event, &handle, &sample);
perf_output_end(&handle);
}
void perf_event_text_poke(const void *addr, const void *old_bytes,
size_t old_len, const void *new_bytes, size_t new_len)
{
struct perf_text_poke_event text_poke_event;
size_t tot, pad;
if (!atomic_read(&nr_text_poke_events))
return;
tot = sizeof(text_poke_event.old_len) + old_len;
tot += sizeof(text_poke_event.new_len) + new_len;
pad = ALIGN(tot, sizeof(u64)) - tot;
text_poke_event = (struct perf_text_poke_event){
.old_bytes = old_bytes,
.new_bytes = new_bytes,
.pad = pad,
.old_len = old_len,
.new_len = new_len,
.event_id = {
.header = {
.type = PERF_RECORD_TEXT_POKE,
.misc = PERF_RECORD_MISC_KERNEL,
.size = sizeof(text_poke_event.event_id) + tot + pad,
},
.addr = (unsigned long)addr,
},
};
perf_iterate_sb(perf_event_text_poke_output, &text_poke_event, NULL);
}
void perf_event_itrace_started(struct perf_event *event)
{
WRITE_ONCE(event->attach_state, event->attach_state | PERF_ATTACH_ITRACE);
}
static void perf_log_itrace_start(struct perf_event *event)
{
struct perf_output_handle handle;
struct perf_sample_data sample;
struct perf_aux_event {
struct perf_event_header header;
u32 pid;
u32 tid;
} rec;
int ret;
if (event->parent)
event = event->parent;
if (!(event->pmu->capabilities & PERF_PMU_CAP_ITRACE) ||
event->attach_state & PERF_ATTACH_ITRACE)
return;
rec.header.type = PERF_RECORD_ITRACE_START;
rec.header.misc = 0;
rec.header.size = sizeof(rec);
rec.pid = perf_event_pid(event, current);
rec.tid = perf_event_tid(event, current);
perf_event_header__init_id(&rec.header, &sample, event);
ret = perf_output_begin(&handle, &sample, event, rec.header.size);
if (ret)
return;
perf_output_put(&handle, rec);
perf_event__output_id_sample(event, &handle, &sample);
perf_output_end(&handle);
}
void perf_report_aux_output_id(struct perf_event *event, u64 hw_id)
{
struct perf_output_handle handle;
struct perf_sample_data sample;
struct perf_aux_event {
struct perf_event_header header;
u64 hw_id;
} rec;
int ret;
if (event->parent)
event = event->parent;
rec.header.type = PERF_RECORD_AUX_OUTPUT_HW_ID;
rec.header.misc = 0;
rec.header.size = sizeof(rec);
rec.hw_id = hw_id;
perf_event_header__init_id(&rec.header, &sample, event);
ret = perf_output_begin(&handle, &sample, event, rec.header.size);
if (ret)
return;
perf_output_put(&handle, rec);
perf_event__output_id_sample(event, &handle, &sample);
perf_output_end(&handle);
}
EXPORT_SYMBOL_GPL(perf_report_aux_output_id);
static int
__perf_event_account_interrupt(struct perf_event *event, int throttle)
{
struct hw_perf_event *hwc = &event->hw;
int ret = 0;
u64 seq;
seq = __this_cpu_read(perf_throttled_seq);
if (seq != hwc->interrupts_seq) {
hwc->interrupts_seq = seq;
hwc->interrupts = 1;
} else {
hwc->interrupts++;
}
if (unlikely(throttle && hwc->interrupts >= max_samples_per_tick)) {
__this_cpu_inc(perf_throttled_count);
tick_dep_set_cpu(smp_processor_id(), TICK_DEP_BIT_PERF_EVENTS);
perf_event_throttle_group(event);
ret = 1;
}
if (event->attr.freq) {
u64 now = perf_clock();
s64 delta = now - hwc->freq_time_stamp;
hwc->freq_time_stamp = now;
if (delta > 0 && delta < 2*TICK_NSEC)
perf_adjust_period(event, delta, hwc->last_period, true);
}
return ret;
}
int perf_event_account_interrupt(struct perf_event *event)
{
return __perf_event_account_interrupt(event, 1);
}
static inline bool sample_is_allowed(struct perf_event *event, struct pt_regs *regs)
{
/*
* Due to interrupt latency (AKA "skid"), we may enter the
* kernel before taking an overflow, even if the PMU is only
* counting user events.
*/
if (event->attr.exclude_kernel && !user_mode(regs))
return false;
return true;
}
#ifdef CONFIG_BPF_SYSCALL
static int bpf_overflow_handler(struct perf_event *event,
struct perf_sample_data *data,
struct pt_regs *regs)
{
struct bpf_perf_event_data_kern ctx = {
.data = data,
.event = event,
};
struct bpf_prog *prog;
int ret = 0;
ctx.regs = perf_arch_bpf_user_pt_regs(regs);
if (unlikely(__this_cpu_inc_return(bpf_prog_active) != 1))
goto out;
rcu_read_lock();
prog = READ_ONCE(event->prog);
if (prog) {
perf_prepare_sample(data, event, regs);
ret = bpf_prog_run(prog, &ctx);
}
rcu_read_unlock();
out:
__this_cpu_dec(bpf_prog_active);
return ret;
}
static inline int perf_event_set_bpf_handler(struct perf_event *event,
struct bpf_prog *prog,
u64 bpf_cookie)
{
if (event->overflow_handler_context)
/* hw breakpoint or kernel counter */
return -EINVAL;
if (event->prog)
return -EEXIST;
if (prog->type != BPF_PROG_TYPE_PERF_EVENT)
return -EINVAL;
if (event->attr.precise_ip &&
prog->call_get_stack &&
(!(event->attr.sample_type & PERF_SAMPLE_CALLCHAIN) ||
event->attr.exclude_callchain_kernel ||
event->attr.exclude_callchain_user)) {
/*
* On perf_event with precise_ip, calling bpf_get_stack()
* may trigger unwinder warnings and occasional crashes.
* bpf_get_[stack|stackid] works around this issue by using
* callchain attached to perf_sample_data. If the
* perf_event does not full (kernel and user) callchain
* attached to perf_sample_data, do not allow attaching BPF
* program that calls bpf_get_[stack|stackid].
*/
return -EPROTO;
}
event->prog = prog;
event->bpf_cookie = bpf_cookie;
return 0;
}
static inline void perf_event_free_bpf_handler(struct perf_event *event)
{
struct bpf_prog *prog = event->prog;
if (!prog)
return;
event->prog = NULL;
bpf_prog_put(prog);
}
#else
static inline int bpf_overflow_handler(struct perf_event *event,
struct perf_sample_data *data,
struct pt_regs *regs)
{
return 1;
}
static inline int perf_event_set_bpf_handler(struct perf_event *event,
struct bpf_prog *prog,
u64 bpf_cookie)
{
return -EOPNOTSUPP;
}
static inline void perf_event_free_bpf_handler(struct perf_event *event)
{
}
#endif
/*
* Generic event overflow handling, sampling.
*/
static int __perf_event_overflow(struct perf_event *event,
int throttle, struct perf_sample_data *data,
struct pt_regs *regs)
{
int events = atomic_read(&event->event_limit);
int ret = 0;
/*
* Non-sampling counters might still use the PMI to fold short
* hardware counters, ignore those.
*/
if (unlikely(!is_sampling_event(event)))
return 0;
ret = __perf_event_account_interrupt(event, throttle);
if (event->attr.aux_pause)
perf_event_aux_pause(event->aux_event, true);
if (event->prog && event->prog->type == BPF_PROG_TYPE_PERF_EVENT &&
!bpf_overflow_handler(event, data, regs))
goto out;
/*
* XXX event_limit might not quite work as expected on inherited
* events
*/
event->pending_kill = POLL_IN;
if (events && atomic_dec_and_test(&event->event_limit)) {
ret = 1;
event->pending_kill = POLL_HUP;
perf_event_disable_inatomic(event);
event->pmu->stop(event, 0);
}
if (event->attr.sigtrap) {
/*
* The desired behaviour of sigtrap vs invalid samples is a bit
* tricky; on the one hand, one should not loose the SIGTRAP if
* it is the first event, on the other hand, we should also not
* trigger the WARN or override the data address.
*/
bool valid_sample = sample_is_allowed(event, regs);
unsigned int pending_id = 1;
enum task_work_notify_mode notify_mode;
if (regs)
pending_id = hash32_ptr((void *)instruction_pointer(regs)) ?: 1;
notify_mode = in_nmi() ? TWA_NMI_CURRENT : TWA_RESUME;
if (!event->pending_work &&
!task_work_add(current, &event->pending_task, notify_mode)) {
event->pending_work = pending_id;
local_inc(&event->ctx->nr_no_switch_fast);
WARN_ON_ONCE(!atomic_long_inc_not_zero(&event->refcount));
event->pending_addr = 0;
if (valid_sample && (data->sample_flags & PERF_SAMPLE_ADDR))
event->pending_addr = data->addr;
} else if (event->attr.exclude_kernel && valid_sample) {
/*
* Should not be able to return to user space without
* consuming pending_work; with exceptions:
*
* 1. Where !exclude_kernel, events can overflow again
* in the kernel without returning to user space.
*
* 2. Events that can overflow again before the IRQ-
* work without user space progress (e.g. hrtimer).
* To approximate progress (with false negatives),
* check 32-bit hash of the current IP.
*/
WARN_ON_ONCE(event->pending_work != pending_id);
}
}
READ_ONCE(event->overflow_handler)(event, data, regs);
if (*perf_event_fasync(event) && event->pending_kill) {
event->pending_wakeup = 1;
irq_work_queue(&event->pending_irq);
}
out:
if (event->attr.aux_resume)
perf_event_aux_pause(event->aux_event, false);
return ret;
}
int perf_event_overflow(struct perf_event *event,
struct perf_sample_data *data,
struct pt_regs *regs)
{
return __perf_event_overflow(event, 1, data, regs);
}
/*
* Generic software event infrastructure
*/
struct swevent_htable {
struct swevent_hlist *swevent_hlist;
struct mutex hlist_mutex;
int hlist_refcount;
};
static DEFINE_PER_CPU(struct swevent_htable, swevent_htable);
/*
* We directly increment event->count and keep a second value in
* event->hw.period_left to count intervals. This period event
* is kept in the range [-sample_period, 0] so that we can use the
* sign as trigger.
*/
u64 perf_swevent_set_period(struct perf_event *event)
{
struct hw_perf_event *hwc = &event->hw;
u64 period = hwc->last_period;
u64 nr, offset;
s64 old, val;
hwc->last_period = hwc->sample_period;
old = local64_read(&hwc->period_left);
do {
val = old;
if (val < 0)
return 0;
nr = div64_u64(period + val, period);
offset = nr * period;
val -= offset;
} while (!local64_try_cmpxchg(&hwc->period_left, &old, val));
return nr;
}
static void perf_swevent_overflow(struct perf_event *event, u64 overflow,
struct perf_sample_data *data,
struct pt_regs *regs)
{
struct hw_perf_event *hwc = &event->hw;
int throttle = 0;
if (!overflow)
overflow = perf_swevent_set_period(event);
if (hwc->interrupts == MAX_INTERRUPTS)
return;
for (; overflow; overflow--) {
if (__perf_event_overflow(event, throttle,
data, regs)) {
/*
* We inhibit the overflow from happening when
* hwc->interrupts == MAX_INTERRUPTS.
*/
break;
}
throttle = 1;
}
}
static void perf_swevent_event(struct perf_event *event, u64 nr,
struct perf_sample_data *data,
struct pt_regs *regs)
{
struct hw_perf_event *hwc = &event->hw;
local64_add(nr, &event->count);
if (!regs)
return;
if (!is_sampling_event(event))
return;
if ((event->attr.sample_type & PERF_SAMPLE_PERIOD) && !event->attr.freq) {
data->period = nr;
return perf_swevent_overflow(event, 1, data, regs);
} else
data->period = event->hw.last_period;
if (nr == 1 && hwc->sample_period == 1 && !event->attr.freq)
return perf_swevent_overflow(event, 1, data, regs);
if (local64_add_negative(nr, &hwc->period_left))
return;
perf_swevent_overflow(event, 0, data, regs);
}
int perf_exclude_event(struct perf_event *event, struct pt_regs *regs)
{
if (event->hw.state & PERF_HES_STOPPED)
return 1;
if (regs) {
if (event->attr.exclude_user && user_mode(regs))
return 1;
if (event->attr.exclude_kernel && !user_mode(regs))
return 1;
}
return 0;
}
static int perf_swevent_match(struct perf_event *event,
enum perf_type_id type,
u32 event_id,
struct perf_sample_data *data,
struct pt_regs *regs)
{
if (event->attr.type != type)
return 0;
if (event->attr.config != event_id)
return 0;
if (perf_exclude_event(event, regs))
return 0;
return 1;
}
static inline u64 swevent_hash(u64 type, u32 event_id)
{
u64 val = event_id | (type << 32);
return hash_64(val, SWEVENT_HLIST_BITS);
}
static inline struct hlist_head *
__find_swevent_head(struct swevent_hlist *hlist, u64 type, u32 event_id)
{
u64 hash = swevent_hash(type, event_id);
return &hlist->heads[hash];
}
/* For the read side: events when they trigger */
static inline struct hlist_head *
find_swevent_head_rcu(struct swevent_htable *swhash, u64 type, u32 event_id)
{
struct swevent_hlist *hlist;
hlist = rcu_dereference(swhash->swevent_hlist);
if (!hlist)
return NULL;
return __find_swevent_head(hlist, type, event_id);
}
/* For the event head insertion and removal in the hlist */
static inline struct hlist_head *
find_swevent_head(struct swevent_htable *swhash, struct perf_event *event)
{
struct swevent_hlist *hlist;
u32 event_id = event->attr.config;
u64 type = event->attr.type;
/*
* Event scheduling is always serialized against hlist allocation
* and release. Which makes the protected version suitable here.
* The context lock guarantees that.
*/
hlist = rcu_dereference_protected(swhash->swevent_hlist,
lockdep_is_held(&event->ctx->lock));
if (!hlist)
return NULL;
return __find_swevent_head(hlist, type, event_id);
}
static void do_perf_sw_event(enum perf_type_id type, u32 event_id,
u64 nr,
struct perf_sample_data *data,
struct pt_regs *regs)
{
struct swevent_htable *swhash = this_cpu_ptr(&swevent_htable);
struct perf_event *event;
struct hlist_head *head;
rcu_read_lock();
head = find_swevent_head_rcu(swhash, type, event_id);
if (!head)
goto end;
hlist_for_each_entry_rcu(event, head, hlist_entry) {
if (perf_swevent_match(event, type, event_id, data, regs))
perf_swevent_event(event, nr, data, regs);
}
end:
rcu_read_unlock();
}
DEFINE_PER_CPU(struct pt_regs, __perf_regs[4]);
int perf_swevent_get_recursion_context(void)
{
return get_recursion_context(current->perf_recursion);
}
EXPORT_SYMBOL_GPL(perf_swevent_get_recursion_context);
void perf_swevent_put_recursion_context(int rctx)
{
put_recursion_context(current->perf_recursion, rctx);
}
void ___perf_sw_event(u32 event_id, u64 nr, struct pt_regs *regs, u64 addr)
{
struct perf_sample_data data;
if (WARN_ON_ONCE(!regs))
return;
perf_sample_data_init(&data, addr, 0);
do_perf_sw_event(PERF_TYPE_SOFTWARE, event_id, nr, &data, regs);
}
void __perf_sw_event(u32 event_id, u64 nr, struct pt_regs *regs, u64 addr)
{
int rctx;
preempt_disable_notrace();
rctx = perf_swevent_get_recursion_context();
if (unlikely(rctx < 0))
goto fail;
___perf_sw_event(event_id, nr, regs, addr);
perf_swevent_put_recursion_context(rctx);
fail:
preempt_enable_notrace();
}
static void perf_swevent_read(struct perf_event *event)
{
}
static int perf_swevent_add(struct perf_event *event, int flags)
{
struct swevent_htable *swhash = this_cpu_ptr(&swevent_htable);
struct hw_perf_event *hwc = &event->hw;
struct hlist_head *head;
if (is_sampling_event(event)) {
hwc->last_period = hwc->sample_period;
perf_swevent_set_period(event);
}
hwc->state = !(flags & PERF_EF_START);
head = find_swevent_head(swhash, event);
if (WARN_ON_ONCE(!head))
return -EINVAL;
hlist_add_head_rcu(&event->hlist_entry, head);
perf_event_update_userpage(event);
return 0;
}
static void perf_swevent_del(struct perf_event *event, int flags)
{
hlist_del_rcu(&event->hlist_entry);
}
static void perf_swevent_start(struct perf_event *event, int flags)
{
event->hw.state = 0;
}
static void perf_swevent_stop(struct perf_event *event, int flags)
{
event->hw.state = PERF_HES_STOPPED;
}
/* Deref the hlist from the update side */
static inline struct swevent_hlist *
swevent_hlist_deref(struct swevent_htable *swhash)
{
return rcu_dereference_protected(swhash->swevent_hlist,
lockdep_is_held(&swhash->hlist_mutex));
}
static void swevent_hlist_release(struct swevent_htable *swhash)
{
struct swevent_hlist *hlist = swevent_hlist_deref(swhash);
if (!hlist)
return;
RCU_INIT_POINTER(swhash->swevent_hlist, NULL);
kfree_rcu(hlist, rcu_head);
}
static void swevent_hlist_put_cpu(int cpu)
{
struct swevent_htable *swhash = &per_cpu(swevent_htable, cpu);
mutex_lock(&swhash->hlist_mutex);
if (!--swhash->hlist_refcount)
swevent_hlist_release(swhash);
mutex_unlock(&swhash->hlist_mutex);
}
static void swevent_hlist_put(void)
{
int cpu;
for_each_possible_cpu(cpu)
swevent_hlist_put_cpu(cpu);
}
static int swevent_hlist_get_cpu(int cpu)
{
struct swevent_htable *swhash = &per_cpu(swevent_htable, cpu);
int err = 0;
mutex_lock(&swhash->hlist_mutex);
if (!swevent_hlist_deref(swhash) &&
cpumask_test_cpu(cpu, perf_online_mask)) {
struct swevent_hlist *hlist;
hlist = kzalloc(sizeof(*hlist), GFP_KERNEL);
if (!hlist) {
err = -ENOMEM;
goto exit;
}
rcu_assign_pointer(swhash->swevent_hlist, hlist);
}
swhash->hlist_refcount++;
exit:
mutex_unlock(&swhash->hlist_mutex);
return err;
}
static int swevent_hlist_get(void)
{
int err, cpu, failed_cpu;
mutex_lock(&pmus_lock);
for_each_possible_cpu(cpu) {
err = swevent_hlist_get_cpu(cpu);
if (err) {
failed_cpu = cpu;
goto fail;
}
}
mutex_unlock(&pmus_lock);
return 0;
fail:
for_each_possible_cpu(cpu) {
if (cpu == failed_cpu)
break;
swevent_hlist_put_cpu(cpu);
}
mutex_unlock(&pmus_lock);
return err;
}
struct static_key perf_swevent_enabled[PERF_COUNT_SW_MAX];
static void sw_perf_event_destroy(struct perf_event *event)
{
u64 event_id = event->attr.config;
WARN_ON(event->parent);
static_key_slow_dec(&perf_swevent_enabled[event_id]);
swevent_hlist_put();
}
static struct pmu perf_cpu_clock; /* fwd declaration */
static struct pmu perf_task_clock;
static int perf_swevent_init(struct perf_event *event)
{
u64 event_id = event->attr.config;
if (event->attr.type != PERF_TYPE_SOFTWARE)
return -ENOENT;
/*
* no branch sampling for software events
*/
if (has_branch_stack(event))
return -EOPNOTSUPP;
switch (event_id) {
case PERF_COUNT_SW_CPU_CLOCK:
event->attr.type = perf_cpu_clock.type;
return -ENOENT;
case PERF_COUNT_SW_TASK_CLOCK:
event->attr.type = perf_task_clock.type;
return -ENOENT;
default:
break;
}
if (event_id >= PERF_COUNT_SW_MAX)
return -ENOENT;
if (!event->parent) {
int err;
err = swevent_hlist_get();
if (err)
return err;
static_key_slow_inc(&perf_swevent_enabled[event_id]);
event->destroy = sw_perf_event_destroy;
}
return 0;
}
static struct pmu perf_swevent = {
.task_ctx_nr = perf_sw_context,
.capabilities = PERF_PMU_CAP_NO_NMI,
.event_init = perf_swevent_init,
.add = perf_swevent_add,
.del = perf_swevent_del,
.start = perf_swevent_start,
.stop = perf_swevent_stop,
.read = perf_swevent_read,
};
#ifdef CONFIG_EVENT_TRACING
static void tp_perf_event_destroy(struct perf_event *event)
{
perf_trace_destroy(event);
}
static int perf_tp_event_init(struct perf_event *event)
{
int err;
if (event->attr.type != PERF_TYPE_TRACEPOINT)
return -ENOENT;
/*
* no branch sampling for tracepoint events
*/
if (has_branch_stack(event))
return -EOPNOTSUPP;
err = perf_trace_init(event);
if (err)
return err;
event->destroy = tp_perf_event_destroy;
return 0;
}
static struct pmu perf_tracepoint = {
.task_ctx_nr = perf_sw_context,
.event_init = perf_tp_event_init,
.add = perf_trace_add,
.del = perf_trace_del,
.start = perf_swevent_start,
.stop = perf_swevent_stop,
.read = perf_swevent_read,
};
static int perf_tp_filter_match(struct perf_event *event,
struct perf_raw_record *raw)
{
void *record = raw->frag.data;
/* only top level events have filters set */
if (event->parent)
event = event->parent;
if (likely(!event->filter) || filter_match_preds(event->filter, record))
return 1;
return 0;
}
static int perf_tp_event_match(struct perf_event *event,
struct perf_raw_record *raw,
struct pt_regs *regs)
{
if (event->hw.state & PERF_HES_STOPPED)
return 0;
/*
* If exclude_kernel, only trace user-space tracepoints (uprobes)
*/
if (event->attr.exclude_kernel && !user_mode(regs))
return 0;
if (!perf_tp_filter_match(event, raw))
return 0;
return 1;
}
void perf_trace_run_bpf_submit(void *raw_data, int size, int rctx,
struct trace_event_call *call, u64 count,
struct pt_regs *regs, struct hlist_head *head,
struct task_struct *task)
{
if (bpf_prog_array_valid(call)) {
*(struct pt_regs **)raw_data = regs;
if (!trace_call_bpf(call, raw_data) || hlist_empty(head)) {
perf_swevent_put_recursion_context(rctx);
return;
}
}
perf_tp_event(call->event.type, count, raw_data, size, regs, head,
rctx, task);
}
EXPORT_SYMBOL_GPL(perf_trace_run_bpf_submit);
static void __perf_tp_event_target_task(u64 count, void *record,
struct pt_regs *regs,
struct perf_sample_data *data,
struct perf_raw_record *raw,
struct perf_event *event)
{
struct trace_entry *entry = record;
if (event->attr.config != entry->type)
return;
/* Cannot deliver synchronous signal to other task. */
if (event->attr.sigtrap)
return;
if (perf_tp_event_match(event, raw, regs)) {
perf_sample_data_init(data, 0, 0);
perf_sample_save_raw_data(data, event, raw);
perf_swevent_event(event, count, data, regs);
}
}
static void perf_tp_event_target_task(u64 count, void *record,
struct pt_regs *regs,
struct perf_sample_data *data,
struct perf_raw_record *raw,
struct perf_event_context *ctx)
{
unsigned int cpu = smp_processor_id();
struct pmu *pmu = &perf_tracepoint;
struct perf_event *event, *sibling;
perf_event_groups_for_cpu_pmu(event, &ctx->pinned_groups, cpu, pmu) {
__perf_tp_event_target_task(count, record, regs, data, raw, event);
for_each_sibling_event(sibling, event)
__perf_tp_event_target_task(count, record, regs, data, raw, sibling);
}
perf_event_groups_for_cpu_pmu(event, &ctx->flexible_groups, cpu, pmu) {
__perf_tp_event_target_task(count, record, regs, data, raw, event);
for_each_sibling_event(sibling, event)
__perf_tp_event_target_task(count, record, regs, data, raw, sibling);
}
}
void perf_tp_event(u16 event_type, u64 count, void *record, int entry_size,
struct pt_regs *regs, struct hlist_head *head, int rctx,
struct task_struct *task)
{
struct perf_sample_data data;
struct perf_event *event;
struct perf_raw_record raw = {
.frag = {
.size = entry_size,
.data = record,
},
};
perf_trace_buf_update(record, event_type);
hlist_for_each_entry_rcu(event, head, hlist_entry) {
if (perf_tp_event_match(event, &raw, regs)) {
/*
* Here use the same on-stack perf_sample_data,
* some members in data are event-specific and
* need to be re-computed for different sweveents.
* Re-initialize data->sample_flags safely to avoid
* the problem that next event skips preparing data
* because data->sample_flags is set.
*/
perf_sample_data_init(&data, 0, 0);
perf_sample_save_raw_data(&data, event, &raw);
perf_swevent_event(event, count, &data, regs);
}
}
/*
* If we got specified a target task, also iterate its context and
* deliver this event there too.
*/
if (task && task != current) {
struct perf_event_context *ctx;
rcu_read_lock();
ctx = rcu_dereference(task->perf_event_ctxp);
if (!ctx)
goto unlock;
raw_spin_lock(&ctx->lock);
perf_tp_event_target_task(count, record, regs, &data, &raw, ctx);
raw_spin_unlock(&ctx->lock);
unlock:
rcu_read_unlock();
}
perf_swevent_put_recursion_context(rctx);
}
EXPORT_SYMBOL_GPL(perf_tp_event);
#if defined(CONFIG_KPROBE_EVENTS) || defined(CONFIG_UPROBE_EVENTS)
/*
* Flags in config, used by dynamic PMU kprobe and uprobe
* The flags should match following PMU_FORMAT_ATTR().
*
* PERF_PROBE_CONFIG_IS_RETPROBE if set, create kretprobe/uretprobe
* if not set, create kprobe/uprobe
*
* The following values specify a reference counter (or semaphore in the
* terminology of tools like dtrace, systemtap, etc.) Userspace Statically
* Defined Tracepoints (USDT). Currently, we use 40 bit for the offset.
*
* PERF_UPROBE_REF_CTR_OFFSET_BITS # of bits in config as th offset
* PERF_UPROBE_REF_CTR_OFFSET_SHIFT # of bits to shift left
*/
enum perf_probe_config {
PERF_PROBE_CONFIG_IS_RETPROBE = 1U << 0, /* [k,u]retprobe */
PERF_UPROBE_REF_CTR_OFFSET_BITS = 32,
PERF_UPROBE_REF_CTR_OFFSET_SHIFT = 64 - PERF_UPROBE_REF_CTR_OFFSET_BITS,
};
PMU_FORMAT_ATTR(retprobe, "config:0");
#endif
#ifdef CONFIG_KPROBE_EVENTS
static struct attribute *kprobe_attrs[] = {
&format_attr_retprobe.attr,
NULL,
};
static struct attribute_group kprobe_format_group = {
.name = "format",
.attrs = kprobe_attrs,
};
static const struct attribute_group *kprobe_attr_groups[] = {
&kprobe_format_group,
NULL,
};
static int perf_kprobe_event_init(struct perf_event *event);
static struct pmu perf_kprobe = {
.task_ctx_nr = perf_sw_context,
.event_init = perf_kprobe_event_init,
.add = perf_trace_add,
.del = perf_trace_del,
.start = perf_swevent_start,
.stop = perf_swevent_stop,
.read = perf_swevent_read,
.attr_groups = kprobe_attr_groups,
};
static int perf_kprobe_event_init(struct perf_event *event)
{
int err;
bool is_retprobe;
if (event->attr.type != perf_kprobe.type)
return -ENOENT;
if (!perfmon_capable())
return -EACCES;
/*
* no branch sampling for probe events
*/
if (has_branch_stack(event))
return -EOPNOTSUPP;
is_retprobe = event->attr.config & PERF_PROBE_CONFIG_IS_RETPROBE;
err = perf_kprobe_init(event, is_retprobe);
if (err)
return err;
event->destroy = perf_kprobe_destroy;
return 0;
}
#endif /* CONFIG_KPROBE_EVENTS */
#ifdef CONFIG_UPROBE_EVENTS
PMU_FORMAT_ATTR(ref_ctr_offset, "config:32-63");
static struct attribute *uprobe_attrs[] = {
&format_attr_retprobe.attr,
&format_attr_ref_ctr_offset.attr,
NULL,
};
static struct attribute_group uprobe_format_group = {
.name = "format",
.attrs = uprobe_attrs,
};
static const struct attribute_group *uprobe_attr_groups[] = {
&uprobe_format_group,
NULL,
};
static int perf_uprobe_event_init(struct perf_event *event);
static struct pmu perf_uprobe = {
.task_ctx_nr = perf_sw_context,
.event_init = perf_uprobe_event_init,
.add = perf_trace_add,
.del = perf_trace_del,
.start = perf_swevent_start,
.stop = perf_swevent_stop,
.read = perf_swevent_read,
.attr_groups = uprobe_attr_groups,
};
static int perf_uprobe_event_init(struct perf_event *event)
{
int err;
unsigned long ref_ctr_offset;
bool is_retprobe;
if (event->attr.type != perf_uprobe.type)
return -ENOENT;
if (!capable(CAP_SYS_ADMIN))
return -EACCES;
/*
* no branch sampling for probe events
*/
if (has_branch_stack(event))
return -EOPNOTSUPP;
is_retprobe = event->attr.config & PERF_PROBE_CONFIG_IS_RETPROBE;
ref_ctr_offset = event->attr.config >> PERF_UPROBE_REF_CTR_OFFSET_SHIFT;
err = perf_uprobe_init(event, ref_ctr_offset, is_retprobe);
if (err)
return err;
event->destroy = perf_uprobe_destroy;
return 0;
}
#endif /* CONFIG_UPROBE_EVENTS */
static inline void perf_tp_register(void)
{
perf_pmu_register(&perf_tracepoint, "tracepoint", PERF_TYPE_TRACEPOINT);
#ifdef CONFIG_KPROBE_EVENTS
perf_pmu_register(&perf_kprobe, "kprobe", -1);
#endif
#ifdef CONFIG_UPROBE_EVENTS
perf_pmu_register(&perf_uprobe, "uprobe", -1);
#endif
}
static void perf_event_free_filter(struct perf_event *event)
{
ftrace_profile_free_filter(event);
}
/*
* returns true if the event is a tracepoint, or a kprobe/upprobe created
* with perf_event_open()
*/
static inline bool perf_event_is_tracing(struct perf_event *event)
{
if (event->pmu == &perf_tracepoint)
return true;
#ifdef CONFIG_KPROBE_EVENTS
if (event->pmu == &perf_kprobe)
return true;
#endif
#ifdef CONFIG_UPROBE_EVENTS
if (event->pmu == &perf_uprobe)
return true;
#endif
return false;
}
static int __perf_event_set_bpf_prog(struct perf_event *event,
struct bpf_prog *prog,
u64 bpf_cookie)
{
bool is_kprobe, is_uprobe, is_tracepoint, is_syscall_tp;
if (event->state <= PERF_EVENT_STATE_REVOKED)
return -ENODEV;
if (!perf_event_is_tracing(event))
return perf_event_set_bpf_handler(event, prog, bpf_cookie);
is_kprobe = event->tp_event->flags & TRACE_EVENT_FL_KPROBE;
is_uprobe = event->tp_event->flags & TRACE_EVENT_FL_UPROBE;
is_tracepoint = event->tp_event->flags & TRACE_EVENT_FL_TRACEPOINT;
is_syscall_tp = is_syscall_trace_event(event->tp_event);
if (!is_kprobe && !is_uprobe && !is_tracepoint && !is_syscall_tp)
/* bpf programs can only be attached to u/kprobe or tracepoint */
return -EINVAL;
if (((is_kprobe || is_uprobe) && prog->type != BPF_PROG_TYPE_KPROBE) ||
(is_tracepoint && prog->type != BPF_PROG_TYPE_TRACEPOINT) ||
(is_syscall_tp && prog->type != BPF_PROG_TYPE_TRACEPOINT))
return -EINVAL;
if (prog->type == BPF_PROG_TYPE_KPROBE && prog->sleepable && !is_uprobe)
/* only uprobe programs are allowed to be sleepable */
return -EINVAL;
/* Kprobe override only works for kprobes, not uprobes. */
if (prog->kprobe_override && !is_kprobe)
return -EINVAL;
/* Writing to context allowed only for uprobes. */
if (prog->aux->kprobe_write_ctx && !is_uprobe)
return -EINVAL;
if (is_tracepoint || is_syscall_tp) {
int off = trace_event_get_offsets(event->tp_event);
if (prog->aux->max_ctx_offset > off)
return -EACCES;
}
return perf_event_attach_bpf_prog(event, prog, bpf_cookie);
}
int perf_event_set_bpf_prog(struct perf_event *event,
struct bpf_prog *prog,
u64 bpf_cookie)
{
struct perf_event_context *ctx;
int ret;
ctx = perf_event_ctx_lock(event);
ret = __perf_event_set_bpf_prog(event, prog, bpf_cookie);
perf_event_ctx_unlock(event, ctx);
return ret;
}
void perf_event_free_bpf_prog(struct perf_event *event)
{
if (!event->prog)
return;
if (!perf_event_is_tracing(event)) {
perf_event_free_bpf_handler(event);
return;
}
perf_event_detach_bpf_prog(event);
}
#else
static inline void perf_tp_register(void)
{
}
static void perf_event_free_filter(struct perf_event *event)
{
}
static int __perf_event_set_bpf_prog(struct perf_event *event,
struct bpf_prog *prog,
u64 bpf_cookie)
{
return -ENOENT;
}
int perf_event_set_bpf_prog(struct perf_event *event,
struct bpf_prog *prog,
u64 bpf_cookie)
{
return -ENOENT;
}
void perf_event_free_bpf_prog(struct perf_event *event)
{
}
#endif /* CONFIG_EVENT_TRACING */
#ifdef CONFIG_HAVE_HW_BREAKPOINT
void perf_bp_event(struct perf_event *bp, void *data)
{
struct perf_sample_data sample;
struct pt_regs *regs = data;
perf_sample_data_init(&sample, bp->attr.bp_addr, 0);
if (!bp->hw.state && !perf_exclude_event(bp, regs))
perf_swevent_event(bp, 1, &sample, regs);
}
#endif
/*
* Allocate a new address filter
*/
static struct perf_addr_filter *
perf_addr_filter_new(struct perf_event *event, struct list_head *filters)
{
int node = cpu_to_node(event->cpu == -1 ? 0 : event->cpu);
struct perf_addr_filter *filter;
filter = kzalloc_node(sizeof(*filter), GFP_KERNEL, node);
if (!filter)
return NULL;
INIT_LIST_HEAD(&filter->entry);
list_add_tail(&filter->entry, filters);
return filter;
}
static void free_filters_list(struct list_head *filters)
{
struct perf_addr_filter *filter, *iter;
list_for_each_entry_safe(filter, iter, filters, entry) {
path_put(&filter->path);
list_del(&filter->entry);
kfree(filter);
}
}
/*
* Free existing address filters and optionally install new ones
*/
static void perf_addr_filters_splice(struct perf_event *event,
struct list_head *head)
{
unsigned long flags;
LIST_HEAD(list);
if (!has_addr_filter(event))
return;
/* don't bother with children, they don't have their own filters */
if (event->parent)
return;
raw_spin_lock_irqsave(&event->addr_filters.lock, flags);
list_splice_init(&event->addr_filters.list, &list);
if (head)
list_splice(head, &event->addr_filters.list);
raw_spin_unlock_irqrestore(&event->addr_filters.lock, flags);
free_filters_list(&list);
}
static void perf_free_addr_filters(struct perf_event *event)
{
/*
* Used during free paths, there is no concurrency.
*/
if (list_empty(&event->addr_filters.list))
return;
perf_addr_filters_splice(event, NULL);
}
/*
* Scan through mm's vmas and see if one of them matches the
* @filter; if so, adjust filter's address range.
* Called with mm::mmap_lock down for reading.
*/
static void perf_addr_filter_apply(struct perf_addr_filter *filter,
struct mm_struct *mm,
struct perf_addr_filter_range *fr)
{
struct vm_area_struct *vma;
VMA_ITERATOR(vmi, mm, 0);
for_each_vma(vmi, vma) {
if (!vma->vm_file)
continue;
if (perf_addr_filter_vma_adjust(filter, vma, fr))
return;
}
}
/*
* Update event's address range filters based on the
* task's existing mappings, if any.
*/
static void perf_event_addr_filters_apply(struct perf_event *event)
{
struct perf_addr_filters_head *ifh = perf_event_addr_filters(event);
struct task_struct *task = READ_ONCE(event->ctx->task);
struct perf_addr_filter *filter;
struct mm_struct *mm = NULL;
unsigned int count = 0;
unsigned long flags;
/*
* We may observe TASK_TOMBSTONE, which means that the event tear-down
* will stop on the parent's child_mutex that our caller is also holding
*/
if (task == TASK_TOMBSTONE)
return;
if (ifh->nr_file_filters) {
mm = get_task_mm(task);
if (!mm)
goto restart;
mmap_read_lock(mm);
}
raw_spin_lock_irqsave(&ifh->lock, flags);
list_for_each_entry(filter, &ifh->list, entry) {
if (filter->path.dentry) {
/*
* Adjust base offset if the filter is associated to a
* binary that needs to be mapped:
*/
event->addr_filter_ranges[count].start = 0;
event->addr_filter_ranges[count].size = 0;
perf_addr_filter_apply(filter, mm, &event->addr_filter_ranges[count]);
} else {
event->addr_filter_ranges[count].start = filter->offset;
event->addr_filter_ranges[count].size = filter->size;
}
count++;
}
event->addr_filters_gen++;
raw_spin_unlock_irqrestore(&ifh->lock, flags);
if (ifh->nr_file_filters) {
mmap_read_unlock(mm);
mmput(mm);
}
restart:
perf_event_stop(event, 1);
}
/*
* Address range filtering: limiting the data to certain
* instruction address ranges. Filters are ioctl()ed to us from
* userspace as ascii strings.
*
* Filter string format:
*
* ACTION RANGE_SPEC
* where ACTION is one of the
* * "filter": limit the trace to this region
* * "start": start tracing from this address
* * "stop": stop tracing at this address/region;
* RANGE_SPEC is
* * for kernel addresses: <start address>[/<size>]
* * for object files: <start address>[/<size>]@</path/to/object/file>
*
* if <size> is not specified or is zero, the range is treated as a single
* address; not valid for ACTION=="filter".
*/
enum {
IF_ACT_NONE = -1,
IF_ACT_FILTER,
IF_ACT_START,
IF_ACT_STOP,
IF_SRC_FILE,
IF_SRC_KERNEL,
IF_SRC_FILEADDR,
IF_SRC_KERNELADDR,
};
enum {
IF_STATE_ACTION = 0,
IF_STATE_SOURCE,
IF_STATE_END,
};
static const match_table_t if_tokens = {
{ IF_ACT_FILTER, "filter" },
{ IF_ACT_START, "start" },
{ IF_ACT_STOP, "stop" },
{ IF_SRC_FILE, "%u/%u@%s" },
{ IF_SRC_KERNEL, "%u/%u" },
{ IF_SRC_FILEADDR, "%u@%s" },
{ IF_SRC_KERNELADDR, "%u" },
{ IF_ACT_NONE, NULL },
};
/*
* Address filter string parser
*/
static int
perf_event_parse_addr_filter(struct perf_event *event, char *fstr,
struct list_head *filters)
{
struct perf_addr_filter *filter = NULL;
char *start, *orig, *filename = NULL;
substring_t args[MAX_OPT_ARGS];
int state = IF_STATE_ACTION, token;
unsigned int kernel = 0;
int ret = -EINVAL;
orig = fstr = kstrdup(fstr, GFP_KERNEL);
if (!fstr)
return -ENOMEM;
while ((start = strsep(&fstr, " ,\n")) != NULL) {
static const enum perf_addr_filter_action_t actions[] = {
[IF_ACT_FILTER] = PERF_ADDR_FILTER_ACTION_FILTER,
[IF_ACT_START] = PERF_ADDR_FILTER_ACTION_START,
[IF_ACT_STOP] = PERF_ADDR_FILTER_ACTION_STOP,
};
ret = -EINVAL;
if (!*start)
continue;
/* filter definition begins */
if (state == IF_STATE_ACTION) {
filter = perf_addr_filter_new(event, filters);
if (!filter)
goto fail;
}
token = match_token(start, if_tokens, args);
switch (token) {
case IF_ACT_FILTER:
case IF_ACT_START:
case IF_ACT_STOP:
if (state != IF_STATE_ACTION)
goto fail;
filter->action = actions[token];
state = IF_STATE_SOURCE;
break;
case IF_SRC_KERNELADDR:
case IF_SRC_KERNEL:
kernel = 1;
fallthrough;
case IF_SRC_FILEADDR:
case IF_SRC_FILE:
if (state != IF_STATE_SOURCE)
goto fail;
*args[0].to = 0;
ret = kstrtoul(args[0].from, 0, &filter->offset);
if (ret)
goto fail;
if (token == IF_SRC_KERNEL || token == IF_SRC_FILE) {
*args[1].to = 0;
ret = kstrtoul(args[1].from, 0, &filter->size);
if (ret)
goto fail;
}
if (token == IF_SRC_FILE || token == IF_SRC_FILEADDR) {
int fpos = token == IF_SRC_FILE ? 2 : 1;
kfree(filename);
filename = match_strdup(&args[fpos]);
if (!filename) {
ret = -ENOMEM;
goto fail;
}
}
state = IF_STATE_END;
break;
default:
goto fail;
}
/*
* Filter definition is fully parsed, validate and install it.
* Make sure that it doesn't contradict itself or the event's
* attribute.
*/
if (state == IF_STATE_END) {
ret = -EINVAL;
/*
* ACTION "filter" must have a non-zero length region
* specified.
*/
if (filter->action == PERF_ADDR_FILTER_ACTION_FILTER &&
!filter->size)
goto fail;
if (!kernel) {
if (!filename)
goto fail;
/*
* For now, we only support file-based filters
* in per-task events; doing so for CPU-wide
* events requires additional context switching
* trickery, since same object code will be
* mapped at different virtual addresses in
* different processes.
*/
ret = -EOPNOTSUPP;
if (!event->ctx->task)
goto fail;
/* look up the path and grab its inode */
ret = kern_path(filename, LOOKUP_FOLLOW,
&filter->path);
if (ret)
goto fail;
ret = -EINVAL;
if (!filter->path.dentry ||
!S_ISREG(d_inode(filter->path.dentry)
->i_mode))
goto fail;
event->addr_filters.nr_file_filters++;
}
/* ready to consume more filters */
kfree(filename);
filename = NULL;
state = IF_STATE_ACTION;
filter = NULL;
kernel = 0;
}
}
if (state != IF_STATE_ACTION)
goto fail;
kfree(filename);
kfree(orig);
return 0;
fail:
kfree(filename);
free_filters_list(filters);
kfree(orig);
return ret;
}
static int
perf_event_set_addr_filter(struct perf_event *event, char *filter_str)
{
LIST_HEAD(filters);
int ret;
/*
* Since this is called in perf_ioctl() path, we're already holding
* ctx::mutex.
*/
lockdep_assert_held(&event->ctx->mutex);
if (WARN_ON_ONCE(event->parent))
return -EINVAL;
ret = perf_event_parse_addr_filter(event, filter_str, &filters);
if (ret)
goto fail_clear_files;
ret = event->pmu->addr_filters_validate(&filters);
if (ret)
goto fail_free_filters;
/* remove existing filters, if any */
perf_addr_filters_splice(event, &filters);
/* install new filters */
perf_event_for_each_child(event, perf_event_addr_filters_apply);
return ret;
fail_free_filters:
free_filters_list(&filters);
fail_clear_files:
event->addr_filters.nr_file_filters = 0;
return ret;
}
static int perf_event_set_filter(struct perf_event *event, void __user *arg)
{
int ret = -EINVAL;
char *filter_str;
filter_str = strndup_user(arg, PAGE_SIZE);
if (IS_ERR(filter_str))
return PTR_ERR(filter_str);
#ifdef CONFIG_EVENT_TRACING
if (perf_event_is_tracing(event)) {
struct perf_event_context *ctx = event->ctx;
/*
* Beware, here be dragons!!
*
* the tracepoint muck will deadlock against ctx->mutex, but
* the tracepoint stuff does not actually need it. So
* temporarily drop ctx->mutex. As per perf_event_ctx_lock() we
* already have a reference on ctx.
*
* This can result in event getting moved to a different ctx,
* but that does not affect the tracepoint state.
*/
mutex_unlock(&ctx->mutex);
ret = ftrace_profile_set_filter(event, event->attr.config, filter_str);
mutex_lock(&ctx->mutex);
} else
#endif
if (has_addr_filter(event))
ret = perf_event_set_addr_filter(event, filter_str);
kfree(filter_str);
return ret;
}
/*
* hrtimer based swevent callback
*/
static enum hrtimer_restart perf_swevent_hrtimer(struct hrtimer *hrtimer)
{
enum hrtimer_restart ret = HRTIMER_RESTART;
struct perf_sample_data data;
struct pt_regs *regs;
struct perf_event *event;
u64 period;
event = container_of(hrtimer, struct perf_event, hw.hrtimer);
if (event->state != PERF_EVENT_STATE_ACTIVE ||
event->hw.state & PERF_HES_STOPPED)
return HRTIMER_NORESTART;
event->pmu->read(event);
perf_sample_data_init(&data, 0, event->hw.last_period);
regs = get_irq_regs();
if (regs && !perf_exclude_event(event, regs)) {
if (!(event->attr.exclude_idle && is_idle_task(current)))
if (__perf_event_overflow(event, 1, &data, regs))
ret = HRTIMER_NORESTART;
}
period = max_t(u64, 10000, event->hw.sample_period);
hrtimer_forward_now(hrtimer, ns_to_ktime(period));
return ret;
}
static void perf_swevent_start_hrtimer(struct perf_event *event)
{
struct hw_perf_event *hwc = &event->hw;
s64 period;
if (!is_sampling_event(event))
return;
period = local64_read(&hwc->period_left);
if (period) {
if (period < 0)
period = 10000;
local64_set(&hwc->period_left, 0);
} else {
period = max_t(u64, 10000, hwc->sample_period);
}
hrtimer_start(&hwc->hrtimer, ns_to_ktime(period),
HRTIMER_MODE_REL_PINNED_HARD);
}
static void perf_swevent_cancel_hrtimer(struct perf_event *event)
{
struct hw_perf_event *hwc = &event->hw;
/*
* Careful: this function can be triggered in the hrtimer handler,
* for cpu-clock events, so hrtimer_cancel() would cause a
* deadlock.
*
* So use hrtimer_try_to_cancel() to try to stop the hrtimer,
* and the cpu-clock handler also sets the PERF_HES_STOPPED flag,
* which guarantees that perf_swevent_hrtimer() will stop the
* hrtimer once it sees the PERF_HES_STOPPED flag.
*/
if (is_sampling_event(event) && (hwc->interrupts != MAX_INTERRUPTS)) {
ktime_t remaining = hrtimer_get_remaining(&hwc->hrtimer);
local64_set(&hwc->period_left, ktime_to_ns(remaining));
hrtimer_try_to_cancel(&hwc->hrtimer);
}
}
static void perf_swevent_init_hrtimer(struct perf_event *event)
{
struct hw_perf_event *hwc = &event->hw;
if (!is_sampling_event(event))
return;
hrtimer_setup(&hwc->hrtimer, perf_swevent_hrtimer, CLOCK_MONOTONIC, HRTIMER_MODE_REL_HARD);
/*
* Since hrtimers have a fixed rate, we can do a static freq->period
* mapping and avoid the whole period adjust feedback stuff.
*/
if (event->attr.freq) {
long freq = event->attr.sample_freq;
event->attr.sample_period = NSEC_PER_SEC / freq;
hwc->sample_period = event->attr.sample_period;
local64_set(&hwc->period_left, hwc->sample_period);
hwc->last_period = hwc->sample_period;
event->attr.freq = 0;
}
}
/*
* Software event: cpu wall time clock
*/
static void cpu_clock_event_update(struct perf_event *event)
{
s64 prev;
u64 now;
now = local_clock();
prev = local64_xchg(&event->hw.prev_count, now);
local64_add(now - prev, &event->count);
}
static void cpu_clock_event_start(struct perf_event *event, int flags)
{
event->hw.state = 0;
local64_set(&event->hw.prev_count, local_clock());
perf_swevent_start_hrtimer(event);
}
static void cpu_clock_event_stop(struct perf_event *event, int flags)
{
event->hw.state = PERF_HES_STOPPED;
perf_swevent_cancel_hrtimer(event);
if (flags & PERF_EF_UPDATE)
cpu_clock_event_update(event);
}
static int cpu_clock_event_add(struct perf_event *event, int flags)
{
if (flags & PERF_EF_START)
cpu_clock_event_start(event, flags);
perf_event_update_userpage(event);
return 0;
}
static void cpu_clock_event_del(struct perf_event *event, int flags)
{
cpu_clock_event_stop(event, PERF_EF_UPDATE);
}
static void cpu_clock_event_read(struct perf_event *event)
{
cpu_clock_event_update(event);
}
static int cpu_clock_event_init(struct perf_event *event)
{
if (event->attr.type != perf_cpu_clock.type)
return -ENOENT;
if (event->attr.config != PERF_COUNT_SW_CPU_CLOCK)
return -ENOENT;
/*
* no branch sampling for software events
*/
if (has_branch_stack(event))
return -EOPNOTSUPP;
perf_swevent_init_hrtimer(event);
return 0;
}
static struct pmu perf_cpu_clock = {
.task_ctx_nr = perf_sw_context,
.capabilities = PERF_PMU_CAP_NO_NMI,
.dev = PMU_NULL_DEV,
.event_init = cpu_clock_event_init,
.add = cpu_clock_event_add,
.del = cpu_clock_event_del,
.start = cpu_clock_event_start,
.stop = cpu_clock_event_stop,
.read = cpu_clock_event_read,
};
/*
* Software event: task time clock
*/
static void task_clock_event_update(struct perf_event *event, u64 now)
{
u64 prev;
s64 delta;
prev = local64_xchg(&event->hw.prev_count, now);
delta = now - prev;
local64_add(delta, &event->count);
}
static void task_clock_event_start(struct perf_event *event, int flags)
{
event->hw.state = 0;
local64_set(&event->hw.prev_count, event->ctx->time);
perf_swevent_start_hrtimer(event);
}
static void task_clock_event_stop(struct perf_event *event, int flags)
{
event->hw.state = PERF_HES_STOPPED;
perf_swevent_cancel_hrtimer(event);
if (flags & PERF_EF_UPDATE)
task_clock_event_update(event, event->ctx->time);
}
static int task_clock_event_add(struct perf_event *event, int flags)
{
if (flags & PERF_EF_START)
task_clock_event_start(event, flags);
perf_event_update_userpage(event);
return 0;
}
static void task_clock_event_del(struct perf_event *event, int flags)
{
task_clock_event_stop(event, PERF_EF_UPDATE);
}
static void task_clock_event_read(struct perf_event *event)
{
u64 now = perf_clock();
u64 delta = now - event->ctx->timestamp;
u64 time = event->ctx->time + delta;
task_clock_event_update(event, time);
}
static int task_clock_event_init(struct perf_event *event)
{
if (event->attr.type != perf_task_clock.type)
return -ENOENT;
if (event->attr.config != PERF_COUNT_SW_TASK_CLOCK)
return -ENOENT;
/*
* no branch sampling for software events
*/
if (has_branch_stack(event))
return -EOPNOTSUPP;
perf_swevent_init_hrtimer(event);
return 0;
}
static struct pmu perf_task_clock = {
.task_ctx_nr = perf_sw_context,
.capabilities = PERF_PMU_CAP_NO_NMI,
.dev = PMU_NULL_DEV,
.event_init = task_clock_event_init,
.add = task_clock_event_add,
.del = task_clock_event_del,
.start = task_clock_event_start,
.stop = task_clock_event_stop,
.read = task_clock_event_read,
};
static void perf_pmu_nop_void(struct pmu *pmu)
{
}
static void perf_pmu_nop_txn(struct pmu *pmu, unsigned int flags)
{
}
static int perf_pmu_nop_int(struct pmu *pmu)
{
return 0;
}
static int perf_event_nop_int(struct perf_event *event, u64 value)
{
return 0;
}
static DEFINE_PER_CPU(unsigned int, nop_txn_flags);
static void perf_pmu_start_txn(struct pmu *pmu, unsigned int flags)
{
__this_cpu_write(nop_txn_flags, flags);
if (flags & ~PERF_PMU_TXN_ADD)
return;
perf_pmu_disable(pmu);
}
static int perf_pmu_commit_txn(struct pmu *pmu)
{
unsigned int flags = __this_cpu_read(nop_txn_flags);
__this_cpu_write(nop_txn_flags, 0);
if (flags & ~PERF_PMU_TXN_ADD)
return 0;
perf_pmu_enable(pmu);
return 0;
}
static void perf_pmu_cancel_txn(struct pmu *pmu)
{
unsigned int flags = __this_cpu_read(nop_txn_flags);
__this_cpu_write(nop_txn_flags, 0);
if (flags & ~PERF_PMU_TXN_ADD)
return;
perf_pmu_enable(pmu);
}
static int perf_event_idx_default(struct perf_event *event)
{
return 0;
}
/*
* Let userspace know that this PMU supports address range filtering:
*/
static ssize_t nr_addr_filters_show(struct device *dev,
struct device_attribute *attr,
char *page)
{
struct pmu *pmu = dev_get_drvdata(dev);
return sysfs_emit(page, "%d\n", pmu->nr_addr_filters);
}
DEVICE_ATTR_RO(nr_addr_filters);
static struct idr pmu_idr;
static ssize_t
type_show(struct device *dev, struct device_attribute *attr, char *page)
{
struct pmu *pmu = dev_get_drvdata(dev);
return sysfs_emit(page, "%d\n", pmu->type);
}
static DEVICE_ATTR_RO(type);
static ssize_t
perf_event_mux_interval_ms_show(struct device *dev,
struct device_attribute *attr,
char *page)
{
struct pmu *pmu = dev_get_drvdata(dev);
return sysfs_emit(page, "%d\n", pmu->hrtimer_interval_ms);
}
static DEFINE_MUTEX(mux_interval_mutex);
static ssize_t
perf_event_mux_interval_ms_store(struct device *dev,
struct device_attribute *attr,
const char *buf, size_t count)
{
struct pmu *pmu = dev_get_drvdata(dev);
int timer, cpu, ret;
ret = kstrtoint(buf, 0, &timer);
if (ret)
return ret;
if (timer < 1)
return -EINVAL;
/* same value, noting to do */
if (timer == pmu->hrtimer_interval_ms)
return count;
mutex_lock(&mux_interval_mutex);
pmu->hrtimer_interval_ms = timer;
/* update all cpuctx for this PMU */
cpus_read_lock();
for_each_online_cpu(cpu) {
struct perf_cpu_pmu_context *cpc;
cpc = *per_cpu_ptr(pmu->cpu_pmu_context, cpu);
cpc->hrtimer_interval = ns_to_ktime(NSEC_PER_MSEC * timer);
cpu_function_call(cpu, perf_mux_hrtimer_restart_ipi, cpc);
}
cpus_read_unlock();
mutex_unlock(&mux_interval_mutex);
return count;
}
static DEVICE_ATTR_RW(perf_event_mux_interval_ms);
static inline const struct cpumask *perf_scope_cpu_topology_cpumask(unsigned int scope, int cpu)
{
switch (scope) {
case PERF_PMU_SCOPE_CORE:
return topology_sibling_cpumask(cpu);
case PERF_PMU_SCOPE_DIE:
return topology_die_cpumask(cpu);
case PERF_PMU_SCOPE_CLUSTER:
return topology_cluster_cpumask(cpu);
case PERF_PMU_SCOPE_PKG:
return topology_core_cpumask(cpu);
case PERF_PMU_SCOPE_SYS_WIDE:
return cpu_online_mask;
}
return NULL;
}
static inline struct cpumask *perf_scope_cpumask(unsigned int scope)
{
switch (scope) {
case PERF_PMU_SCOPE_CORE:
return perf_online_core_mask;
case PERF_PMU_SCOPE_DIE:
return perf_online_die_mask;
case PERF_PMU_SCOPE_CLUSTER:
return perf_online_cluster_mask;
case PERF_PMU_SCOPE_PKG:
return perf_online_pkg_mask;
case PERF_PMU_SCOPE_SYS_WIDE:
return perf_online_sys_mask;
}
return NULL;
}
static ssize_t cpumask_show(struct device *dev, struct device_attribute *attr,
char *buf)
{
struct pmu *pmu = dev_get_drvdata(dev);
struct cpumask *mask = perf_scope_cpumask(pmu->scope);
if (mask)
return cpumap_print_to_pagebuf(true, buf, mask);
return 0;
}
static DEVICE_ATTR_RO(cpumask);
static struct attribute *pmu_dev_attrs[] = {
&dev_attr_type.attr,
&dev_attr_perf_event_mux_interval_ms.attr,
&dev_attr_nr_addr_filters.attr,
&dev_attr_cpumask.attr,
NULL,
};
static umode_t pmu_dev_is_visible(struct kobject *kobj, struct attribute *a, int n)
{
struct device *dev = kobj_to_dev(kobj);
struct pmu *pmu = dev_get_drvdata(dev);
if (n == 2 && !pmu->nr_addr_filters)
return 0;
/* cpumask */
if (n == 3 && pmu->scope == PERF_PMU_SCOPE_NONE)
return 0;
return a->mode;
}
static struct attribute_group pmu_dev_attr_group = {
.is_visible = pmu_dev_is_visible,
.attrs = pmu_dev_attrs,
};
static const struct attribute_group *pmu_dev_groups[] = {
&pmu_dev_attr_group,
NULL,
};
static int pmu_bus_running;
static const struct bus_type pmu_bus = {
.name = "event_source",
.dev_groups = pmu_dev_groups,
};
static void pmu_dev_release(struct device *dev)
{
kfree(dev);
}
static int pmu_dev_alloc(struct pmu *pmu)
{
int ret = -ENOMEM;
pmu->dev = kzalloc(sizeof(struct device), GFP_KERNEL);
if (!pmu->dev)
goto out;
pmu->dev->groups = pmu->attr_groups;
device_initialize(pmu->dev);
dev_set_drvdata(pmu->dev, pmu);
pmu->dev->bus = &pmu_bus;
pmu->dev->parent = pmu->parent;
pmu->dev->release = pmu_dev_release;
ret = dev_set_name(pmu->dev, "%s", pmu->name);
if (ret)
goto free_dev;
ret = device_add(pmu->dev);
if (ret)
goto free_dev;
if (pmu->attr_update) {
ret = sysfs_update_groups(&pmu->dev->kobj, pmu->attr_update);
if (ret)
goto del_dev;
}
out:
return ret;
del_dev:
device_del(pmu->dev);
free_dev:
put_device(pmu->dev);
pmu->dev = NULL;
goto out;
}
static struct lock_class_key cpuctx_mutex;
static struct lock_class_key cpuctx_lock;
static bool idr_cmpxchg(struct idr *idr, unsigned long id, void *old, void *new)
{
void *tmp, *val = idr_find(idr, id);
if (val != old)
return false;
tmp = idr_replace(idr, new, id);
if (IS_ERR(tmp))
return false;
WARN_ON_ONCE(tmp != val);
return true;
}
static void perf_pmu_free(struct pmu *pmu)
{
if (pmu_bus_running && pmu->dev && pmu->dev != PMU_NULL_DEV) {
if (pmu->nr_addr_filters)
device_remove_file(pmu->dev, &dev_attr_nr_addr_filters);
device_del(pmu->dev);
put_device(pmu->dev);
}
if (pmu->cpu_pmu_context) {
int cpu;
for_each_possible_cpu(cpu) {
struct perf_cpu_pmu_context *cpc;
cpc = *per_cpu_ptr(pmu->cpu_pmu_context, cpu);
if (!cpc)
continue;
if (cpc->epc.embedded) {
/* refcount managed */
put_pmu_ctx(&cpc->epc);
continue;
}
kfree(cpc);
}
free_percpu(pmu->cpu_pmu_context);
}
}
DEFINE_FREE(pmu_unregister, struct pmu *, if (_T) perf_pmu_free(_T))
int perf_pmu_register(struct pmu *_pmu, const char *name, int type)
{
int cpu, max = PERF_TYPE_MAX;
struct pmu *pmu __free(pmu_unregister) = _pmu;
guard(mutex)(&pmus_lock);
if (WARN_ONCE(!name, "Can not register anonymous pmu.\n"))
return -EINVAL;
if (WARN_ONCE(pmu->scope >= PERF_PMU_MAX_SCOPE,
"Can not register a pmu with an invalid scope.\n"))
return -EINVAL;
pmu->name = name;
if (type >= 0)
max = type;
CLASS(idr_alloc, pmu_type)(&pmu_idr, NULL, max, 0, GFP_KERNEL);
if (pmu_type.id < 0)
return pmu_type.id;
WARN_ON(type >= 0 && pmu_type.id != type);
pmu->type = pmu_type.id;
atomic_set(&pmu->exclusive_cnt, 0);
if (pmu_bus_running && !pmu->dev) {
int ret = pmu_dev_alloc(pmu);
if (ret)
return ret;
}
pmu->cpu_pmu_context = alloc_percpu(struct perf_cpu_pmu_context *);
if (!pmu->cpu_pmu_context)
return -ENOMEM;
for_each_possible_cpu(cpu) {
struct perf_cpu_pmu_context *cpc =
kmalloc_node(sizeof(struct perf_cpu_pmu_context),
GFP_KERNEL | __GFP_ZERO,
cpu_to_node(cpu));
if (!cpc)
return -ENOMEM;
*per_cpu_ptr(pmu->cpu_pmu_context, cpu) = cpc;
__perf_init_event_pmu_context(&cpc->epc, pmu);
__perf_mux_hrtimer_init(cpc, cpu);
}
if (!pmu->start_txn) {
if (pmu->pmu_enable) {
/*
* If we have pmu_enable/pmu_disable calls, install
* transaction stubs that use that to try and batch
* hardware accesses.
*/
pmu->start_txn = perf_pmu_start_txn;
pmu->commit_txn = perf_pmu_commit_txn;
pmu->cancel_txn = perf_pmu_cancel_txn;
} else {
pmu->start_txn = perf_pmu_nop_txn;
pmu->commit_txn = perf_pmu_nop_int;
pmu->cancel_txn = perf_pmu_nop_void;
}
}
if (!pmu->pmu_enable) {
pmu->pmu_enable = perf_pmu_nop_void;
pmu->pmu_disable = perf_pmu_nop_void;
}
if (!pmu->check_period)
pmu->check_period = perf_event_nop_int;
if (!pmu->event_idx)
pmu->event_idx = perf_event_idx_default;
INIT_LIST_HEAD(&pmu->events);
spin_lock_init(&pmu->events_lock);
/*
* Now that the PMU is complete, make it visible to perf_try_init_event().
*/
if (!idr_cmpxchg(&pmu_idr, pmu->type, NULL, pmu))
return -EINVAL;
list_add_rcu(&pmu->entry, &pmus);
take_idr_id(pmu_type);
_pmu = no_free_ptr(pmu); // let it rip
return 0;
}
EXPORT_SYMBOL_GPL(perf_pmu_register);
static void __pmu_detach_event(struct pmu *pmu, struct perf_event *event,
struct perf_event_context *ctx)
{
/*
* De-schedule the event and mark it REVOKED.
*/
perf_event_exit_event(event, ctx, ctx->task, true);
/*
* All _free_event() bits that rely on event->pmu:
*
* Notably, perf_mmap() relies on the ordering here.
*/
scoped_guard (mutex, &event->mmap_mutex) {
WARN_ON_ONCE(pmu->event_unmapped);
/*
* Mostly an empty lock sequence, such that perf_mmap(), which
* relies on mmap_mutex, is sure to observe the state change.
*/
}
perf_event_free_bpf_prog(event);
perf_free_addr_filters(event);
if (event->destroy) {
event->destroy(event);
event->destroy = NULL;
}
if (event->pmu_ctx) {
put_pmu_ctx(event->pmu_ctx);
event->pmu_ctx = NULL;
}
exclusive_event_destroy(event);
module_put(pmu->module);
event->pmu = NULL; /* force fault instead of UAF */
}
static void pmu_detach_event(struct pmu *pmu, struct perf_event *event)
{
struct perf_event_context *ctx;
ctx = perf_event_ctx_lock(event);
__pmu_detach_event(pmu, event, ctx);
perf_event_ctx_unlock(event, ctx);
scoped_guard (spinlock, &pmu->events_lock)
list_del(&event->pmu_list);
}
static struct perf_event *pmu_get_event(struct pmu *pmu)
{
struct perf_event *event;
guard(spinlock)(&pmu->events_lock);
list_for_each_entry(event, &pmu->events, pmu_list) {
if (atomic_long_inc_not_zero(&event->refcount))
return event;
}
return NULL;
}
static bool pmu_empty(struct pmu *pmu)
{
guard(spinlock)(&pmu->events_lock);
return list_empty(&pmu->events);
}
static void pmu_detach_events(struct pmu *pmu)
{
struct perf_event *event;
for (;;) {
event = pmu_get_event(pmu);
if (!event)
break;
pmu_detach_event(pmu, event);
put_event(event);
}
/*
* wait for pending _free_event()s
*/
wait_var_event(pmu, pmu_empty(pmu));
}
int perf_pmu_unregister(struct pmu *pmu)
{
scoped_guard (mutex, &pmus_lock) {
if (!idr_cmpxchg(&pmu_idr, pmu->type, pmu, NULL))
return -EINVAL;
list_del_rcu(&pmu->entry);
}
/*
* We dereference the pmu list under both SRCU and regular RCU, so
* synchronize against both of those.
*
* Notably, the entirety of event creation, from perf_init_event()
* (which will now fail, because of the above) until
* perf_install_in_context() should be under SRCU such that
* this synchronizes against event creation. This avoids trying to
* detach events that are not fully formed.
*/
synchronize_srcu(&pmus_srcu);
synchronize_rcu();
if (pmu->event_unmapped && !pmu_empty(pmu)) {
/*
* Can't force remove events when pmu::event_unmapped()
* is used in perf_mmap_close().
*/
guard(mutex)(&pmus_lock);
idr_cmpxchg(&pmu_idr, pmu->type, NULL, pmu);
list_add_rcu(&pmu->entry, &pmus);
return -EBUSY;
}
scoped_guard (mutex, &pmus_lock)
idr_remove(&pmu_idr, pmu->type);
/*
* PMU is removed from the pmus list, so no new events will
* be created, now take care of the existing ones.
*/
pmu_detach_events(pmu);
/*
* PMU is unused, make it go away.
*/
perf_pmu_free(pmu);
return 0;
}
EXPORT_SYMBOL_GPL(perf_pmu_unregister);
static inline bool has_extended_regs(struct perf_event *event)
{
return (event->attr.sample_regs_user & PERF_REG_EXTENDED_MASK) ||
(event->attr.sample_regs_intr & PERF_REG_EXTENDED_MASK);
}
static int perf_try_init_event(struct pmu *pmu, struct perf_event *event)
{
struct perf_event_context *ctx = NULL;
int ret;
if (!try_module_get(pmu->module))
return -ENODEV;
/*
* A number of pmu->event_init() methods iterate the sibling_list to,
* for example, validate if the group fits on the PMU. Therefore,
* if this is a sibling event, acquire the ctx->mutex to protect
* the sibling_list.
*/
if (event->group_leader != event && pmu->task_ctx_nr != perf_sw_context) {
/*
* This ctx->mutex can nest when we're called through
* inheritance. See the perf_event_ctx_lock_nested() comment.
*/
ctx = perf_event_ctx_lock_nested(event->group_leader,
SINGLE_DEPTH_NESTING);
BUG_ON(!ctx);
}
event->pmu = pmu;
ret = pmu->event_init(event);
if (ctx)
perf_event_ctx_unlock(event->group_leader, ctx);
if (ret)
goto err_pmu;
if (!(pmu->capabilities & PERF_PMU_CAP_EXTENDED_REGS) &&
has_extended_regs(event)) {
ret = -EOPNOTSUPP;
goto err_destroy;
}
if (pmu->capabilities & PERF_PMU_CAP_NO_EXCLUDE &&
event_has_any_exclude_flag(event)) {
ret = -EINVAL;
goto err_destroy;
}
if (pmu->scope != PERF_PMU_SCOPE_NONE && event->cpu >= 0) {
const struct cpumask *cpumask;
struct cpumask *pmu_cpumask;
int cpu;
cpumask = perf_scope_cpu_topology_cpumask(pmu->scope, event->cpu);
pmu_cpumask = perf_scope_cpumask(pmu->scope);
ret = -ENODEV;
if (!pmu_cpumask || !cpumask)
goto err_destroy;
cpu = cpumask_any_and(pmu_cpumask, cpumask);
if (cpu >= nr_cpu_ids)
goto err_destroy;
event->event_caps |= PERF_EV_CAP_READ_SCOPE;
}
return 0;
err_destroy:
if (event->destroy) {
event->destroy(event);
event->destroy = NULL;
}
err_pmu:
event->pmu = NULL;
module_put(pmu->module);
return ret;
}
static struct pmu *perf_init_event(struct perf_event *event)
{
bool extended_type = false;
struct pmu *pmu;
int type, ret;
guard(srcu)(&pmus_srcu); /* pmu idr/list access */
/*
* Save original type before calling pmu->event_init() since certain
* pmus overwrites event->attr.type to forward event to another pmu.
*/
event->orig_type = event->attr.type;
/* Try parent's PMU first: */
if (event->parent && event->parent->pmu) {
pmu = event->parent->pmu;
ret = perf_try_init_event(pmu, event);
if (!ret)
return pmu;
}
/*
* PERF_TYPE_HARDWARE and PERF_TYPE_HW_CACHE
* are often aliases for PERF_TYPE_RAW.
*/
type = event->attr.type;
if (type == PERF_TYPE_HARDWARE || type == PERF_TYPE_HW_CACHE) {
type = event->attr.config >> PERF_PMU_TYPE_SHIFT;
if (!type) {
type = PERF_TYPE_RAW;
} else {
extended_type = true;
event->attr.config &= PERF_HW_EVENT_MASK;
}
}
again:
scoped_guard (rcu)
pmu = idr_find(&pmu_idr, type);
if (pmu) {
if (event->attr.type != type && type != PERF_TYPE_RAW &&
!(pmu->capabilities & PERF_PMU_CAP_EXTENDED_HW_TYPE))
return ERR_PTR(-ENOENT);
ret = perf_try_init_event(pmu, event);
if (ret == -ENOENT && event->attr.type != type && !extended_type) {
type = event->attr.type;
goto again;
}
if (ret)
return ERR_PTR(ret);
return pmu;
}
list_for_each_entry_rcu(pmu, &pmus, entry, lockdep_is_held(&pmus_srcu)) {
ret = perf_try_init_event(pmu, event);
if (!ret)
return pmu;
if (ret != -ENOENT)
return ERR_PTR(ret);
}
return ERR_PTR(-ENOENT);
}
static void attach_sb_event(struct perf_event *event)
{
struct pmu_event_list *pel = per_cpu_ptr(&pmu_sb_events, event->cpu);
raw_spin_lock(&pel->lock);
list_add_rcu(&event->sb_list, &pel->list);
raw_spin_unlock(&pel->lock);
}
/*
* We keep a list of all !task (and therefore per-cpu) events
* that need to receive side-band records.
*
* This avoids having to scan all the various PMU per-cpu contexts
* looking for them.
*/
static void account_pmu_sb_event(struct perf_event *event)
{
if (is_sb_event(event))
attach_sb_event(event);
}
/* Freq events need the tick to stay alive (see perf_event_task_tick). */
static void account_freq_event_nohz(void)
{
#ifdef CONFIG_NO_HZ_FULL
/* Lock so we don't race with concurrent unaccount */
spin_lock(&nr_freq_lock);
if (atomic_inc_return(&nr_freq_events) == 1)
tick_nohz_dep_set(TICK_DEP_BIT_PERF_EVENTS);
spin_unlock(&nr_freq_lock);
#endif
}
static void account_freq_event(void)
{
if (tick_nohz_full_enabled())
account_freq_event_nohz();
else
atomic_inc(&nr_freq_events);
}
static void account_event(struct perf_event *event)
{
bool inc = false;
if (event->parent)
return;
if (event->attach_state & (PERF_ATTACH_TASK | PERF_ATTACH_SCHED_CB))
inc = true;
if (event->attr.mmap || event->attr.mmap_data)
atomic_inc(&nr_mmap_events);
if (event->attr.build_id)
atomic_inc(&nr_build_id_events);
if (event->attr.comm)
atomic_inc(&nr_comm_events);
if (event->attr.namespaces)
atomic_inc(&nr_namespaces_events);
if (event->attr.cgroup)
atomic_inc(&nr_cgroup_events);
if (event->attr.task)
atomic_inc(&nr_task_events);
if (event->attr.freq)
account_freq_event();
if (event->attr.context_switch) {
atomic_inc(&nr_switch_events);
inc = true;
}
if (has_branch_stack(event))
inc = true;
if (is_cgroup_event(event))
inc = true;
if (event->attr.ksymbol)
atomic_inc(&nr_ksymbol_events);
if (event->attr.bpf_event)
atomic_inc(&nr_bpf_events);
if (event->attr.text_poke)
atomic_inc(&nr_text_poke_events);
if (inc) {
/*
* We need the mutex here because static_branch_enable()
* must complete *before* the perf_sched_count increment
* becomes visible.
*/
if (atomic_inc_not_zero(&perf_sched_count))
goto enabled;
mutex_lock(&perf_sched_mutex);
if (!atomic_read(&perf_sched_count)) {
static_branch_enable(&perf_sched_events);
/*
* Guarantee that all CPUs observe they key change and
* call the perf scheduling hooks before proceeding to
* install events that need them.
*/
synchronize_rcu();
}
/*
* Now that we have waited for the sync_sched(), allow further
* increments to by-pass the mutex.
*/
atomic_inc(&perf_sched_count);
mutex_unlock(&perf_sched_mutex);
}
enabled:
account_pmu_sb_event(event);
}
/*
* Allocate and initialize an event structure
*/
static struct perf_event *
perf_event_alloc(struct perf_event_attr *attr, int cpu,
struct task_struct *task,
struct perf_event *group_leader,
struct perf_event *parent_event,
perf_overflow_handler_t overflow_handler,
void *context, int cgroup_fd)
{
struct pmu *pmu;
struct hw_perf_event *hwc;
long err = -EINVAL;
int node;
if ((unsigned)cpu >= nr_cpu_ids) {
if (!task || cpu != -1)
return ERR_PTR(-EINVAL);
}
if (attr->sigtrap && !task) {
/* Requires a task: avoid signalling random tasks. */
return ERR_PTR(-EINVAL);
}
node = (cpu >= 0) ? cpu_to_node(cpu) : -1;
struct perf_event *event __free(__free_event) =
kmem_cache_alloc_node(perf_event_cache, GFP_KERNEL | __GFP_ZERO, node);
if (!event)
return ERR_PTR(-ENOMEM);
/*
* Single events are their own group leaders, with an
* empty sibling list:
*/
if (!group_leader)
group_leader = event;
mutex_init(&event->child_mutex);
INIT_LIST_HEAD(&event->child_list);
INIT_LIST_HEAD(&event->event_entry);
INIT_LIST_HEAD(&event->sibling_list);
INIT_LIST_HEAD(&event->active_list);
init_event_group(event);
INIT_LIST_HEAD(&event->rb_entry);
INIT_LIST_HEAD(&event->active_entry);
INIT_LIST_HEAD(&event->addr_filters.list);
INIT_HLIST_NODE(&event->hlist_entry);
INIT_LIST_HEAD(&event->pmu_list);
init_waitqueue_head(&event->waitq);
init_irq_work(&event->pending_irq, perf_pending_irq);
event->pending_disable_irq = IRQ_WORK_INIT_HARD(perf_pending_disable);
init_task_work(&event->pending_task, perf_pending_task);
mutex_init(&event->mmap_mutex);
raw_spin_lock_init(&event->addr_filters.lock);
atomic_long_set(&event->refcount, 1);
event->cpu = cpu;
event->attr = *attr;
event->group_leader = group_leader;
event->pmu = NULL;
event->oncpu = -1;
event->parent = parent_event;
event->ns = get_pid_ns(task_active_pid_ns(current));
event->id = atomic64_inc_return(&perf_event_id);
event->state = PERF_EVENT_STATE_INACTIVE;
if (parent_event)
event->event_caps = parent_event->event_caps;
if (task) {
event->attach_state = PERF_ATTACH_TASK;
/*
* XXX pmu::event_init needs to know what task to account to
* and we cannot use the ctx information because we need the
* pmu before we get a ctx.
*/
event->hw.target = get_task_struct(task);
}
event->clock = &local_clock;
if (parent_event)
event->clock = parent_event->clock;
if (!overflow_handler && parent_event) {
overflow_handler = parent_event->overflow_handler;
context = parent_event->overflow_handler_context;
#if defined(CONFIG_BPF_SYSCALL) && defined(CONFIG_EVENT_TRACING)
if (parent_event->prog) {
struct bpf_prog *prog = parent_event->prog;
bpf_prog_inc(prog);
event->prog = prog;
}
#endif
}
if (overflow_handler) {
event->overflow_handler = overflow_handler;
event->overflow_handler_context = context;
} else if (is_write_backward(event)){
event->overflow_handler = perf_event_output_backward;
event->overflow_handler_context = NULL;
} else {
event->overflow_handler = perf_event_output_forward;
event->overflow_handler_context = NULL;
}
perf_event__state_init(event);
pmu = NULL;
hwc = &event->hw;
hwc->sample_period = attr->sample_period;
if (is_event_in_freq_mode(event))
hwc->sample_period = 1;
hwc->last_period = hwc->sample_period;
local64_set(&hwc->period_left, hwc->sample_period);
/*
* We do not support PERF_SAMPLE_READ on inherited events unless
* PERF_SAMPLE_TID is also selected, which allows inherited events to
* collect per-thread samples.
* See perf_output_read().
*/
if (has_inherit_and_sample_read(attr) && !(attr->sample_type & PERF_SAMPLE_TID))
return ERR_PTR(-EINVAL);
if (!has_branch_stack(event))
event->attr.branch_sample_type = 0;
pmu = perf_init_event(event);
if (IS_ERR(pmu))
return (void*)pmu;
/*
* The PERF_ATTACH_TASK_DATA is set in the event_init()->hw_config().
* The attach should be right after the perf_init_event().
* Otherwise, the __free_event() would mistakenly detach the non-exist
* perf_ctx_data because of the other errors between them.
*/
if (event->attach_state & PERF_ATTACH_TASK_DATA) {
err = attach_perf_ctx_data(event);
if (err)
return ERR_PTR(err);
}
/*
* Disallow uncore-task events. Similarly, disallow uncore-cgroup
* events (they don't make sense as the cgroup will be different
* on other CPUs in the uncore mask).
*/
if (pmu->task_ctx_nr == perf_invalid_context && (task || cgroup_fd != -1))
return ERR_PTR(-EINVAL);
if (event->attr.aux_output &&
(!(pmu->capabilities & PERF_PMU_CAP_AUX_OUTPUT) ||
event->attr.aux_pause || event->attr.aux_resume))
return ERR_PTR(-EOPNOTSUPP);
if (event->attr.aux_pause && event->attr.aux_resume)
return ERR_PTR(-EINVAL);
if (event->attr.aux_start_paused) {
if (!(pmu->capabilities & PERF_PMU_CAP_AUX_PAUSE))
return ERR_PTR(-EOPNOTSUPP);
event->hw.aux_paused = 1;
}
if (cgroup_fd != -1) {
err = perf_cgroup_connect(cgroup_fd, event, attr, group_leader);
if (err)
return ERR_PTR(err);
}
err = exclusive_event_init(event);
if (err)
return ERR_PTR(err);
if (has_addr_filter(event)) {
event->addr_filter_ranges = kcalloc(pmu->nr_addr_filters,
sizeof(struct perf_addr_filter_range),
GFP_KERNEL);
if (!event->addr_filter_ranges)
return ERR_PTR(-ENOMEM);
/*
* Clone the parent's vma offsets: they are valid until exec()
* even if the mm is not shared with the parent.
*/
if (event->parent) {
struct perf_addr_filters_head *ifh = perf_event_addr_filters(event);
raw_spin_lock_irq(&ifh->lock);
memcpy(event->addr_filter_ranges,
event->parent->addr_filter_ranges,
pmu->nr_addr_filters * sizeof(struct perf_addr_filter_range));
raw_spin_unlock_irq(&ifh->lock);
}
/* force hw sync on the address filters */
event->addr_filters_gen = 1;
}
if (!event->parent) {
if (event->attr.sample_type & PERF_SAMPLE_CALLCHAIN) {
err = get_callchain_buffers(attr->sample_max_stack);
if (err)
return ERR_PTR(err);
event->attach_state |= PERF_ATTACH_CALLCHAIN;
}
}
err = security_perf_event_alloc(event);
if (err)
return ERR_PTR(err);
/* symmetric to unaccount_event() in _free_event() */
account_event(event);
/*
* Event creation should be under SRCU, see perf_pmu_unregister().
*/
lockdep_assert_held(&pmus_srcu);
scoped_guard (spinlock, &pmu->events_lock)
list_add(&event->pmu_list, &pmu->events);
return_ptr(event);
}
static int perf_copy_attr(struct perf_event_attr __user *uattr,
struct perf_event_attr *attr)
{
u32 size;
int ret;
/* Zero the full structure, so that a short copy will be nice. */
memset(attr, 0, sizeof(*attr));
ret = get_user(size, &uattr->size);
if (ret)
return ret;
/* ABI compatibility quirk: */
if (!size)
size = PERF_ATTR_SIZE_VER0;
if (size < PERF_ATTR_SIZE_VER0 || size > PAGE_SIZE)
goto err_size;
ret = copy_struct_from_user(attr, sizeof(*attr), uattr, size);
if (ret) {
if (ret == -E2BIG)
goto err_size;
return ret;
}
attr->size = size;
if (attr->__reserved_1 || attr->__reserved_2 || attr->__reserved_3)
return -EINVAL;
if (attr->sample_type & ~(PERF_SAMPLE_MAX-1))
return -EINVAL;
if (attr->read_format & ~(PERF_FORMAT_MAX-1))
return -EINVAL;
if (attr->sample_type & PERF_SAMPLE_BRANCH_STACK) {
u64 mask = attr->branch_sample_type;
/* only using defined bits */
if (mask & ~(PERF_SAMPLE_BRANCH_MAX-1))
return -EINVAL;
/* at least one branch bit must be set */
if (!(mask & ~PERF_SAMPLE_BRANCH_PLM_ALL))
return -EINVAL;
/* propagate priv level, when not set for branch */
if (!(mask & PERF_SAMPLE_BRANCH_PLM_ALL)) {
/* exclude_kernel checked on syscall entry */
if (!attr->exclude_kernel)
mask |= PERF_SAMPLE_BRANCH_KERNEL;
if (!attr->exclude_user)
mask |= PERF_SAMPLE_BRANCH_USER;
if (!attr->exclude_hv)
mask |= PERF_SAMPLE_BRANCH_HV;
/*
* adjust user setting (for HW filter setup)
*/
attr->branch_sample_type = mask;
}
/* privileged levels capture (kernel, hv): check permissions */
if (mask & PERF_SAMPLE_BRANCH_PERM_PLM) {
ret = perf_allow_kernel();
if (ret)
return ret;
}
}
if (attr->sample_type & PERF_SAMPLE_REGS_USER) {
ret = perf_reg_validate(attr->sample_regs_user);
if (ret)
return ret;
}
if (attr->sample_type & PERF_SAMPLE_STACK_USER) {
if (!arch_perf_have_user_stack_dump())
return -ENOSYS;
/*
* We have __u32 type for the size, but so far
* we can only use __u16 as maximum due to the
* __u16 sample size limit.
*/
if (attr->sample_stack_user >= USHRT_MAX)
return -EINVAL;
else if (!IS_ALIGNED(attr->sample_stack_user, sizeof(u64)))
return -EINVAL;
}
if (!attr->sample_max_stack)
attr->sample_max_stack = sysctl_perf_event_max_stack;
if (attr->sample_type & PERF_SAMPLE_REGS_INTR)
ret = perf_reg_validate(attr->sample_regs_intr);
#ifndef CONFIG_CGROUP_PERF
if (attr->sample_type & PERF_SAMPLE_CGROUP)
return -EINVAL;
#endif
if ((attr->sample_type & PERF_SAMPLE_WEIGHT) &&
(attr->sample_type & PERF_SAMPLE_WEIGHT_STRUCT))
return -EINVAL;
if (!attr->inherit && attr->inherit_thread)
return -EINVAL;
if (attr->remove_on_exec && attr->enable_on_exec)
return -EINVAL;
if (attr->sigtrap && !attr->remove_on_exec)
return -EINVAL;
out:
return ret;
err_size:
put_user(sizeof(*attr), &uattr->size);
ret = -E2BIG;
goto out;
}
static void mutex_lock_double(struct mutex *a, struct mutex *b)
{
if (b < a)
swap(a, b);
mutex_lock(a);
mutex_lock_nested(b, SINGLE_DEPTH_NESTING);
}
static int
perf_event_set_output(struct perf_event *event, struct perf_event *output_event)
{
struct perf_buffer *rb = NULL;
int ret = -EINVAL;
if (!output_event) {
mutex_lock(&event->mmap_mutex);
goto set;
}
/* don't allow circular references */
if (event == output_event)
goto out;
/*
* Don't allow cross-cpu buffers
*/
if (output_event->cpu != event->cpu)
goto out;
/*
* If its not a per-cpu rb, it must be the same task.
*/
if (output_event->cpu == -1 && output_event->hw.target != event->hw.target)
goto out;
/*
* Mixing clocks in the same buffer is trouble you don't need.
*/
if (output_event->clock != event->clock)
goto out;
/*
* Either writing ring buffer from beginning or from end.
* Mixing is not allowed.
*/
if (is_write_backward(output_event) != is_write_backward(event))
goto out;
/*
* If both events generate aux data, they must be on the same PMU
*/
if (has_aux(event) && has_aux(output_event) &&
event->pmu != output_event->pmu)
goto out;
/*
* Hold both mmap_mutex to serialize against perf_mmap_close(). Since
* output_event is already on rb->event_list, and the list iteration
* restarts after every removal, it is guaranteed this new event is
* observed *OR* if output_event is already removed, it's guaranteed we
* observe !rb->mmap_count.
*/
mutex_lock_double(&event->mmap_mutex, &output_event->mmap_mutex);
set:
/* Can't redirect output if we've got an active mmap() */
if (refcount_read(&event->mmap_count))
goto unlock;
if (output_event) {
if (output_event->state <= PERF_EVENT_STATE_REVOKED)
goto unlock;
/* get the rb we want to redirect to */
rb = ring_buffer_get(output_event);
if (!rb)
goto unlock;
/* did we race against perf_mmap_close() */
if (!refcount_read(&rb->mmap_count)) {
ring_buffer_put(rb);
goto unlock;
}
}
ring_buffer_attach(event, rb);
ret = 0;
unlock:
mutex_unlock(&event->mmap_mutex);
if (output_event)
mutex_unlock(&output_event->mmap_mutex);
out:
return ret;
}
static int perf_event_set_clock(struct perf_event *event, clockid_t clk_id)
{
bool nmi_safe = false;
switch (clk_id) {
case CLOCK_MONOTONIC:
event->clock = &ktime_get_mono_fast_ns;
nmi_safe = true;
break;
case CLOCK_MONOTONIC_RAW:
event->clock = &ktime_get_raw_fast_ns;
nmi_safe = true;
break;
case CLOCK_REALTIME:
event->clock = &ktime_get_real_ns;
break;
case CLOCK_BOOTTIME:
event->clock = &ktime_get_boottime_ns;
break;
case CLOCK_TAI:
event->clock = &ktime_get_clocktai_ns;
break;
default:
return -EINVAL;
}
if (!nmi_safe && !(event->pmu->capabilities & PERF_PMU_CAP_NO_NMI))
return -EINVAL;
return 0;
}
static bool
perf_check_permission(struct perf_event_attr *attr, struct task_struct *task)
{
unsigned int ptrace_mode = PTRACE_MODE_READ_REALCREDS;
bool is_capable = perfmon_capable();
if (attr->sigtrap) {
/*
* perf_event_attr::sigtrap sends signals to the other task.
* Require the current task to also have CAP_KILL.
*/
rcu_read_lock();
is_capable &= ns_capable(__task_cred(task)->user_ns, CAP_KILL);
rcu_read_unlock();
/*
* If the required capabilities aren't available, checks for
* ptrace permissions: upgrade to ATTACH, since sending signals
* can effectively change the target task.
*/
ptrace_mode = PTRACE_MODE_ATTACH_REALCREDS;
}
/*
* Preserve ptrace permission check for backwards compatibility. The
* ptrace check also includes checks that the current task and other
* task have matching uids, and is therefore not done here explicitly.
*/
return is_capable || ptrace_may_access(task, ptrace_mode);
}
/**
* sys_perf_event_open - open a performance event, associate it to a task/cpu
*
* @attr_uptr: event_id type attributes for monitoring/sampling
* @pid: target pid
* @cpu: target cpu
* @group_fd: group leader event fd
* @flags: perf event open flags
*/
SYSCALL_DEFINE5(perf_event_open,
struct perf_event_attr __user *, attr_uptr,
pid_t, pid, int, cpu, int, group_fd, unsigned long, flags)
{
struct perf_event *group_leader = NULL, *output_event = NULL;
struct perf_event_pmu_context *pmu_ctx;
struct perf_event *event, *sibling;
struct perf_event_attr attr;
struct perf_event_context *ctx;
struct file *event_file = NULL;
struct task_struct *task = NULL;
struct pmu *pmu;
int event_fd;
int move_group = 0;
int err;
int f_flags = O_RDWR;
int cgroup_fd = -1;
/* for future expandability... */
if (flags & ~PERF_FLAG_ALL)
return -EINVAL;
err = perf_copy_attr(attr_uptr, &attr);
if (err)
return err;
/* Do we allow access to perf_event_open(2) ? */
err = security_perf_event_open(PERF_SECURITY_OPEN);
if (err)
return err;
if (!attr.exclude_kernel) {
err = perf_allow_kernel();
if (err)
return err;
}
if (attr.namespaces) {
if (!perfmon_capable())
return -EACCES;
}
if (attr.freq) {
if (attr.sample_freq > sysctl_perf_event_sample_rate)
return -EINVAL;
} else {
if (attr.sample_period & (1ULL << 63))
return -EINVAL;
}
/* Only privileged users can get physical addresses */
if ((attr.sample_type & PERF_SAMPLE_PHYS_ADDR)) {
err = perf_allow_kernel();
if (err)
return err;
}
/* REGS_INTR can leak data, lockdown must prevent this */
if (attr.sample_type & PERF_SAMPLE_REGS_INTR) {
err = security_locked_down(LOCKDOWN_PERF);
if (err)
return err;
}
/*
* In cgroup mode, the pid argument is used to pass the fd
* opened to the cgroup directory in cgroupfs. The cpu argument
* designates the cpu on which to monitor threads from that
* cgroup.
*/
if ((flags & PERF_FLAG_PID_CGROUP) && (pid == -1 || cpu == -1))
return -EINVAL;
if (flags & PERF_FLAG_FD_CLOEXEC)
f_flags |= O_CLOEXEC;
event_fd = get_unused_fd_flags(f_flags);
if (event_fd < 0)
return event_fd;
/*
* Event creation should be under SRCU, see perf_pmu_unregister().
*/
guard(srcu)(&pmus_srcu);
CLASS(fd, group)(group_fd); // group_fd == -1 => empty
if (group_fd != -1) {
if (!is_perf_file(group)) {
err = -EBADF;
goto err_fd;
}
group_leader = fd_file(group)->private_data;
if (group_leader->state <= PERF_EVENT_STATE_REVOKED) {
err = -ENODEV;
goto err_fd;
}
if (flags & PERF_FLAG_FD_OUTPUT)
output_event = group_leader;
if (flags & PERF_FLAG_FD_NO_GROUP)
group_leader = NULL;
}
if (pid != -1 && !(flags & PERF_FLAG_PID_CGROUP)) {
task = find_lively_task_by_vpid(pid);
if (IS_ERR(task)) {
err = PTR_ERR(task);
goto err_fd;
}
}
if (task && group_leader &&
group_leader->attr.inherit != attr.inherit) {
err = -EINVAL;
goto err_task;
}
if (flags & PERF_FLAG_PID_CGROUP)
cgroup_fd = pid;
event = perf_event_alloc(&attr, cpu, task, group_leader, NULL,
NULL, NULL, cgroup_fd);
if (IS_ERR(event)) {
err = PTR_ERR(event);
goto err_task;
}
if (is_sampling_event(event)) {
if (event->pmu->capabilities & PERF_PMU_CAP_NO_INTERRUPT) {
err = -EOPNOTSUPP;
goto err_alloc;
}
}
/*
* Special case software events and allow them to be part of
* any hardware group.
*/
pmu = event->pmu;
if (attr.use_clockid) {
err = perf_event_set_clock(event, attr.clockid);
if (err)
goto err_alloc;
}
if (pmu->task_ctx_nr == perf_sw_context)
event->event_caps |= PERF_EV_CAP_SOFTWARE;
if (task) {
err = down_read_interruptible(&task->signal->exec_update_lock);
if (err)
goto err_alloc;
/*
* We must hold exec_update_lock across this and any potential
* perf_install_in_context() call for this new event to
* serialize against exec() altering our credentials (and the
* perf_event_exit_task() that could imply).
*/
err = -EACCES;
if (!perf_check_permission(&attr, task))
goto err_cred;
}
/*
* Get the target context (task or percpu):
*/
ctx = find_get_context(task, event);
if (IS_ERR(ctx)) {
err = PTR_ERR(ctx);
goto err_cred;
}
mutex_lock(&ctx->mutex);
if (ctx->task == TASK_TOMBSTONE) {
err = -ESRCH;
goto err_locked;
}
if (!task) {
/*
* Check if the @cpu we're creating an event for is online.
*
* We use the perf_cpu_context::ctx::mutex to serialize against
* the hotplug notifiers. See perf_event_{init,exit}_cpu().
*/
struct perf_cpu_context *cpuctx = per_cpu_ptr(&perf_cpu_context, event->cpu);
if (!cpuctx->online) {
err = -ENODEV;
goto err_locked;
}
}
if (group_leader) {
err = -EINVAL;
/*
* Do not allow a recursive hierarchy (this new sibling
* becoming part of another group-sibling):
*/
if (group_leader->group_leader != group_leader)
goto err_locked;
/* All events in a group should have the same clock */
if (group_leader->clock != event->clock)
goto err_locked;
/*
* Make sure we're both events for the same CPU;
* grouping events for different CPUs is broken; since
* you can never concurrently schedule them anyhow.
*/
if (group_leader->cpu != event->cpu)
goto err_locked;
/*
* Make sure we're both on the same context; either task or cpu.
*/
if (group_leader->ctx != ctx)
goto err_locked;
/*
* Only a group leader can be exclusive or pinned
*/
if (attr.exclusive || attr.pinned)
goto err_locked;
if (is_software_event(event) &&
!in_software_context(group_leader)) {
/*
* If the event is a sw event, but the group_leader
* is on hw context.
*
* Allow the addition of software events to hw
* groups, this is safe because software events
* never fail to schedule.
*
* Note the comment that goes with struct
* perf_event_pmu_context.
*/
pmu = group_leader->pmu_ctx->pmu;
} else if (!is_software_event(event)) {
if (is_software_event(group_leader) &&
(group_leader->group_caps & PERF_EV_CAP_SOFTWARE)) {
/*
* In case the group is a pure software group, and we
* try to add a hardware event, move the whole group to
* the hardware context.
*/
move_group = 1;
}
/* Don't allow group of multiple hw events from different pmus */
if (!in_software_context(group_leader) &&
group_leader->pmu_ctx->pmu != pmu)
goto err_locked;
}
}
/*
* Now that we're certain of the pmu; find the pmu_ctx.
*/
pmu_ctx = find_get_pmu_context(pmu, ctx, event);
if (IS_ERR(pmu_ctx)) {
err = PTR_ERR(pmu_ctx);
goto err_locked;
}
event->pmu_ctx = pmu_ctx;
if (output_event) {
err = perf_event_set_output(event, output_event);
if (err)
goto err_context;
}
if (!perf_event_validate_size(event)) {
err = -E2BIG;
goto err_context;
}
if (perf_need_aux_event(event) && !perf_get_aux_event(event, group_leader)) {
err = -EINVAL;
goto err_context;
}
/*
* Must be under the same ctx::mutex as perf_install_in_context(),
* because we need to serialize with concurrent event creation.
*/
if (!exclusive_event_installable(event, ctx)) {
err = -EBUSY;
goto err_context;
}
WARN_ON_ONCE(ctx->parent_ctx);
event_file = anon_inode_getfile("[perf_event]", &perf_fops, event, f_flags);
if (IS_ERR(event_file)) {
err = PTR_ERR(event_file);
event_file = NULL;
goto err_context;
}
/*
* This is the point on no return; we cannot fail hereafter. This is
* where we start modifying current state.
*/
if (move_group) {
perf_remove_from_context(group_leader, 0);
put_pmu_ctx(group_leader->pmu_ctx);
for_each_sibling_event(sibling, group_leader) {
perf_remove_from_context(sibling, 0);
put_pmu_ctx(sibling->pmu_ctx);
}
/*
* Install the group siblings before the group leader.
*
* Because a group leader will try and install the entire group
* (through the sibling list, which is still in-tact), we can
* end up with siblings installed in the wrong context.
*
* By installing siblings first we NO-OP because they're not
* reachable through the group lists.
*/
for_each_sibling_event(sibling, group_leader) {
sibling->pmu_ctx = pmu_ctx;
get_pmu_ctx(pmu_ctx);
perf_event__state_init(sibling);
perf_install_in_context(ctx, sibling, sibling->cpu);
}
/*
* Removing from the context ends up with disabled
* event. What we want here is event in the initial
* startup state, ready to be add into new context.
*/
group_leader->pmu_ctx = pmu_ctx;
get_pmu_ctx(pmu_ctx);
perf_event__state_init(group_leader);
perf_install_in_context(ctx, group_leader, group_leader->cpu);
}
/*
* Precalculate sample_data sizes; do while holding ctx::mutex such
* that we're serialized against further additions and before
* perf_install_in_context() which is the point the event is active and
* can use these values.
*/
perf_event__header_size(event);
perf_event__id_header_size(event);
event->owner = current;
perf_install_in_context(ctx, event, event->cpu);
perf_unpin_context(ctx);
mutex_unlock(&ctx->mutex);
if (task) {
up_read(&task->signal->exec_update_lock);
put_task_struct(task);
}
mutex_lock(¤t->perf_event_mutex);
list_add_tail(&event->owner_entry, ¤t->perf_event_list);
mutex_unlock(¤t->perf_event_mutex);
/*
* File reference in group guarantees that group_leader has been
* kept alive until we place the new event on the sibling_list.
* This ensures destruction of the group leader will find
* the pointer to itself in perf_group_detach().
*/
fd_install(event_fd, event_file);
return event_fd;
err_context:
put_pmu_ctx(event->pmu_ctx);
event->pmu_ctx = NULL; /* _free_event() */
err_locked:
mutex_unlock(&ctx->mutex);
perf_unpin_context(ctx);
put_ctx(ctx);
err_cred:
if (task)
up_read(&task->signal->exec_update_lock);
err_alloc:
put_event(event);
err_task:
if (task)
put_task_struct(task);
err_fd:
put_unused_fd(event_fd);
return err;
}
/**
* perf_event_create_kernel_counter
*
* @attr: attributes of the counter to create
* @cpu: cpu in which the counter is bound
* @task: task to profile (NULL for percpu)
* @overflow_handler: callback to trigger when we hit the event
* @context: context data could be used in overflow_handler callback
*/
struct perf_event *
perf_event_create_kernel_counter(struct perf_event_attr *attr, int cpu,
struct task_struct *task,
perf_overflow_handler_t overflow_handler,
void *context)
{
struct perf_event_pmu_context *pmu_ctx;
struct perf_event_context *ctx;
struct perf_event *event;
struct pmu *pmu;
int err;
/*
* Grouping is not supported for kernel events, neither is 'AUX',
* make sure the caller's intentions are adjusted.
*/
if (attr->aux_output || attr->aux_action)
return ERR_PTR(-EINVAL);
/*
* Event creation should be under SRCU, see perf_pmu_unregister().
*/
guard(srcu)(&pmus_srcu);
event = perf_event_alloc(attr, cpu, task, NULL, NULL,
overflow_handler, context, -1);
if (IS_ERR(event)) {
err = PTR_ERR(event);
goto err;
}
/* Mark owner so we could distinguish it from user events. */
event->owner = TASK_TOMBSTONE;
pmu = event->pmu;
if (pmu->task_ctx_nr == perf_sw_context)
event->event_caps |= PERF_EV_CAP_SOFTWARE;
/*
* Get the target context (task or percpu):
*/
ctx = find_get_context(task, event);
if (IS_ERR(ctx)) {
err = PTR_ERR(ctx);
goto err_alloc;
}
WARN_ON_ONCE(ctx->parent_ctx);
mutex_lock(&ctx->mutex);
if (ctx->task == TASK_TOMBSTONE) {
err = -ESRCH;
goto err_unlock;
}
pmu_ctx = find_get_pmu_context(pmu, ctx, event);
if (IS_ERR(pmu_ctx)) {
err = PTR_ERR(pmu_ctx);
goto err_unlock;
}
event->pmu_ctx = pmu_ctx;
if (!task) {
/*
* Check if the @cpu we're creating an event for is online.
*
* We use the perf_cpu_context::ctx::mutex to serialize against
* the hotplug notifiers. See perf_event_{init,exit}_cpu().
*/
struct perf_cpu_context *cpuctx =
container_of(ctx, struct perf_cpu_context, ctx);
if (!cpuctx->online) {
err = -ENODEV;
goto err_pmu_ctx;
}
}
if (!exclusive_event_installable(event, ctx)) {
err = -EBUSY;
goto err_pmu_ctx;
}
perf_install_in_context(ctx, event, event->cpu);
perf_unpin_context(ctx);
mutex_unlock(&ctx->mutex);
return event;
err_pmu_ctx:
put_pmu_ctx(pmu_ctx);
event->pmu_ctx = NULL; /* _free_event() */
err_unlock:
mutex_unlock(&ctx->mutex);
perf_unpin_context(ctx);
put_ctx(ctx);
err_alloc:
put_event(event);
err:
return ERR_PTR(err);
}
EXPORT_SYMBOL_GPL(perf_event_create_kernel_counter);
static void __perf_pmu_remove(struct perf_event_context *ctx,
int cpu, struct pmu *pmu,
struct perf_event_groups *groups,
struct list_head *events)
{
struct perf_event *event, *sibling;
perf_event_groups_for_cpu_pmu(event, groups, cpu, pmu) {
perf_remove_from_context(event, 0);
put_pmu_ctx(event->pmu_ctx);
list_add(&event->migrate_entry, events);
for_each_sibling_event(sibling, event) {
perf_remove_from_context(sibling, 0);
put_pmu_ctx(sibling->pmu_ctx);
list_add(&sibling->migrate_entry, events);
}
}
}
static void __perf_pmu_install_event(struct pmu *pmu,
struct perf_event_context *ctx,
int cpu, struct perf_event *event)
{
struct perf_event_pmu_context *epc;
struct perf_event_context *old_ctx = event->ctx;
get_ctx(ctx); /* normally find_get_context() */
event->cpu = cpu;
epc = find_get_pmu_context(pmu, ctx, event);
event->pmu_ctx = epc;
if (event->state >= PERF_EVENT_STATE_OFF)
event->state = PERF_EVENT_STATE_INACTIVE;
perf_install_in_context(ctx, event, cpu);
/*
* Now that event->ctx is updated and visible, put the old ctx.
*/
put_ctx(old_ctx);
}
static void __perf_pmu_install(struct perf_event_context *ctx,
int cpu, struct pmu *pmu, struct list_head *events)
{
struct perf_event *event, *tmp;
/*
* Re-instate events in 2 passes.
*
* Skip over group leaders and only install siblings on this first
* pass, siblings will not get enabled without a leader, however a
* leader will enable its siblings, even if those are still on the old
* context.
*/
list_for_each_entry_safe(event, tmp, events, migrate_entry) {
if (event->group_leader == event)
continue;
list_del(&event->migrate_entry);
__perf_pmu_install_event(pmu, ctx, cpu, event);
}
/*
* Once all the siblings are setup properly, install the group leaders
* to make it go.
*/
list_for_each_entry_safe(event, tmp, events, migrate_entry) {
list_del(&event->migrate_entry);
__perf_pmu_install_event(pmu, ctx, cpu, event);
}
}
void perf_pmu_migrate_context(struct pmu *pmu, int src_cpu, int dst_cpu)
{
struct perf_event_context *src_ctx, *dst_ctx;
LIST_HEAD(events);
/*
* Since per-cpu context is persistent, no need to grab an extra
* reference.
*/
src_ctx = &per_cpu_ptr(&perf_cpu_context, src_cpu)->ctx;
dst_ctx = &per_cpu_ptr(&perf_cpu_context, dst_cpu)->ctx;
/*
* See perf_event_ctx_lock() for comments on the details
* of swizzling perf_event::ctx.
*/
mutex_lock_double(&src_ctx->mutex, &dst_ctx->mutex);
__perf_pmu_remove(src_ctx, src_cpu, pmu, &src_ctx->pinned_groups, &events);
__perf_pmu_remove(src_ctx, src_cpu, pmu, &src_ctx->flexible_groups, &events);
if (!list_empty(&events)) {
/*
* Wait for the events to quiesce before re-instating them.
*/
synchronize_rcu();
__perf_pmu_install(dst_ctx, dst_cpu, pmu, &events);
}
mutex_unlock(&dst_ctx->mutex);
mutex_unlock(&src_ctx->mutex);
}
EXPORT_SYMBOL_GPL(perf_pmu_migrate_context);
static void sync_child_event(struct perf_event *child_event,
struct task_struct *task)
{
struct perf_event *parent_event = child_event->parent;
u64 child_val;
if (child_event->attr.inherit_stat) {
if (task && task != TASK_TOMBSTONE)
perf_event_read_event(child_event, task);
}
child_val = perf_event_count(child_event, false);
/*
* Add back the child's count to the parent's count:
*/
atomic64_add(child_val, &parent_event->child_count);
atomic64_add(child_event->total_time_enabled,
&parent_event->child_total_time_enabled);
atomic64_add(child_event->total_time_running,
&parent_event->child_total_time_running);
}
static void
perf_event_exit_event(struct perf_event *event,
struct perf_event_context *ctx,
struct task_struct *task,
bool revoke)
{
struct perf_event *parent_event = event->parent;
unsigned long detach_flags = DETACH_EXIT;
unsigned int attach_state;
if (parent_event) {
/*
* Do not destroy the 'original' grouping; because of the
* context switch optimization the original events could've
* ended up in a random child task.
*
* If we were to destroy the original group, all group related
* operations would cease to function properly after this
* random child dies.
*
* Do destroy all inherited groups, we don't care about those
* and being thorough is better.
*/
detach_flags |= DETACH_GROUP | DETACH_CHILD;
mutex_lock(&parent_event->child_mutex);
/* PERF_ATTACH_ITRACE might be set concurrently */
attach_state = READ_ONCE(event->attach_state);
if (attach_state & PERF_ATTACH_CHILD)
sync_child_event(event, task);
}
if (revoke)
detach_flags |= DETACH_GROUP | DETACH_REVOKE;
perf_remove_from_context(event, detach_flags);
/*
* Child events can be freed.
*/
if (parent_event) {
mutex_unlock(&parent_event->child_mutex);
/*
* Match the refcount initialization. Make sure it doesn't happen
* twice if pmu_detach_event() calls it on an already exited task.
*/
if (attach_state & PERF_ATTACH_CHILD) {
/*
* Kick perf_poll() for is_event_hup();
*/
perf_event_wakeup(parent_event);
/*
* pmu_detach_event() will have an extra refcount.
* perf_pending_task() might have one too.
*/
put_event(event);
}
return;
}
/*
* Parent events are governed by their filedesc, retain them.
*/
perf_event_wakeup(event);
}
static void perf_event_exit_task_context(struct task_struct *task, bool exit)
{
struct perf_event_context *ctx, *clone_ctx = NULL;
struct perf_event *child_event, *next;
ctx = perf_pin_task_context(task);
if (!ctx)
return;
/*
* In order to reduce the amount of tricky in ctx tear-down, we hold
* ctx::mutex over the entire thing. This serializes against almost
* everything that wants to access the ctx.
*
* The exception is sys_perf_event_open() /
* perf_event_create_kernel_count() which does find_get_context()
* without ctx::mutex (it cannot because of the move_group double mutex
* lock thing). See the comments in perf_install_in_context().
*/
mutex_lock(&ctx->mutex);
/*
* In a single ctx::lock section, de-schedule the events and detach the
* context from the task such that we cannot ever get it scheduled back
* in.
*/
raw_spin_lock_irq(&ctx->lock);
if (exit)
task_ctx_sched_out(ctx, NULL, EVENT_ALL);
/*
* Now that the context is inactive, destroy the task <-> ctx relation
* and mark the context dead.
*/
RCU_INIT_POINTER(task->perf_event_ctxp, NULL);
put_ctx(ctx); /* cannot be last */
WRITE_ONCE(ctx->task, TASK_TOMBSTONE);
put_task_struct(task); /* cannot be last */
clone_ctx = unclone_ctx(ctx);
raw_spin_unlock_irq(&ctx->lock);
if (clone_ctx)
put_ctx(clone_ctx);
/*
* Report the task dead after unscheduling the events so that we
* won't get any samples after PERF_RECORD_EXIT. We can however still
* get a few PERF_RECORD_READ events.
*/
if (exit)
perf_event_task(task, ctx, 0);
list_for_each_entry_safe(child_event, next, &ctx->event_list, event_entry)
perf_event_exit_event(child_event, ctx, exit ? task : NULL, false);
mutex_unlock(&ctx->mutex);
if (!exit) {
/*
* perf_event_release_kernel() could still have a reference on
* this context. In that case we must wait for these events to
* have been freed (in particular all their references to this
* task must've been dropped).
*
* Without this copy_process() will unconditionally free this
* task (irrespective of its reference count) and
* _free_event()'s put_task_struct(event->hw.target) will be a
* use-after-free.
*
* Wait for all events to drop their context reference.
*/
wait_var_event(&ctx->refcount,
refcount_read(&ctx->refcount) == 1);
}
put_ctx(ctx);
}
/*
* When a task exits, feed back event values to parent events.
*
* Can be called with exec_update_lock held when called from
* setup_new_exec().
*/
void perf_event_exit_task(struct task_struct *task)
{
struct perf_event *event, *tmp;
WARN_ON_ONCE(task != current);
mutex_lock(&task->perf_event_mutex);
list_for_each_entry_safe(event, tmp, &task->perf_event_list,
owner_entry) {
list_del_init(&event->owner_entry);
/*
* Ensure the list deletion is visible before we clear
* the owner, closes a race against perf_release() where
* we need to serialize on the owner->perf_event_mutex.
*/
smp_store_release(&event->owner, NULL);
}
mutex_unlock(&task->perf_event_mutex);
perf_event_exit_task_context(task, true);
/*
* The perf_event_exit_task_context calls perf_event_task
* with task's task_ctx, which generates EXIT events for
* task contexts and sets task->perf_event_ctxp[] to NULL.
* At this point we need to send EXIT events to cpu contexts.
*/
perf_event_task(task, NULL, 0);
/*
* Detach the perf_ctx_data for the system-wide event.
*/
guard(percpu_read)(&global_ctx_data_rwsem);
detach_task_ctx_data(task);
}
/*
* Free a context as created by inheritance by perf_event_init_task() below,
* used by fork() in case of fail.
*
* Even though the task has never lived, the context and events have been
* exposed through the child_list, so we must take care tearing it all down.
*/
void perf_event_free_task(struct task_struct *task)
{
perf_event_exit_task_context(task, false);
}
void perf_event_delayed_put(struct task_struct *task)
{
WARN_ON_ONCE(task->perf_event_ctxp);
}
struct file *perf_event_get(unsigned int fd)
{
struct file *file = fget(fd);
if (!file)
return ERR_PTR(-EBADF);
if (file->f_op != &perf_fops) {
fput(file);
return ERR_PTR(-EBADF);
}
return file;
}
const struct perf_event *perf_get_event(struct file *file)
{
if (file->f_op != &perf_fops)
return ERR_PTR(-EINVAL);
return file->private_data;
}
const struct perf_event_attr *perf_event_attrs(struct perf_event *event)
{
if (!event)
return ERR_PTR(-EINVAL);
return &event->attr;
}
int perf_allow_kernel(void)
{
if (sysctl_perf_event_paranoid > 1 && !perfmon_capable())
return -EACCES;
return security_perf_event_open(PERF_SECURITY_KERNEL);
}
EXPORT_SYMBOL_GPL(perf_allow_kernel);
/*
* Inherit an event from parent task to child task.
*
* Returns:
* - valid pointer on success
* - NULL for orphaned events
* - IS_ERR() on error
*/
static struct perf_event *
inherit_event(struct perf_event *parent_event,
struct task_struct *parent,
struct perf_event_context *parent_ctx,
struct task_struct *child,
struct perf_event *group_leader,
struct perf_event_context *child_ctx)
{
enum perf_event_state parent_state = parent_event->state;
struct perf_event_pmu_context *pmu_ctx;
struct perf_event *child_event;
unsigned long flags;
/*
* Instead of creating recursive hierarchies of events,
* we link inherited events back to the original parent,
* which has a filp for sure, which we use as the reference
* count:
*/
if (parent_event->parent)
parent_event = parent_event->parent;
if (parent_event->state <= PERF_EVENT_STATE_REVOKED)
return NULL;
/*
* Event creation should be under SRCU, see perf_pmu_unregister().
*/
guard(srcu)(&pmus_srcu);
child_event = perf_event_alloc(&parent_event->attr,
parent_event->cpu,
child,
group_leader, parent_event,
NULL, NULL, -1);
if (IS_ERR(child_event))
return child_event;
get_ctx(child_ctx);
child_event->ctx = child_ctx;
pmu_ctx = find_get_pmu_context(child_event->pmu, child_ctx, child_event);
if (IS_ERR(pmu_ctx)) {
free_event(child_event);
return ERR_CAST(pmu_ctx);
}
child_event->pmu_ctx = pmu_ctx;
/*
* is_orphaned_event() and list_add_tail(&parent_event->child_list)
* must be under the same lock in order to serialize against
* perf_event_release_kernel(), such that either we must observe
* is_orphaned_event() or they will observe us on the child_list.
*/
mutex_lock(&parent_event->child_mutex);
if (is_orphaned_event(parent_event) ||
!atomic_long_inc_not_zero(&parent_event->refcount)) {
mutex_unlock(&parent_event->child_mutex);
free_event(child_event);
return NULL;
}
/*
* Make the child state follow the state of the parent event,
* not its attr.disabled bit. We hold the parent's mutex,
* so we won't race with perf_event_{en, dis}able_family.
*/
if (parent_state >= PERF_EVENT_STATE_INACTIVE)
child_event->state = PERF_EVENT_STATE_INACTIVE;
else
child_event->state = PERF_EVENT_STATE_OFF;
if (parent_event->attr.freq) {
u64 sample_period = parent_event->hw.sample_period;
struct hw_perf_event *hwc = &child_event->hw;
hwc->sample_period = sample_period;
hwc->last_period = sample_period;
local64_set(&hwc->period_left, sample_period);
}
child_event->overflow_handler = parent_event->overflow_handler;
child_event->overflow_handler_context
= parent_event->overflow_handler_context;
/*
* Precalculate sample_data sizes
*/
perf_event__header_size(child_event);
perf_event__id_header_size(child_event);
/*
* Link it up in the child's context:
*/
raw_spin_lock_irqsave(&child_ctx->lock, flags);
add_event_to_ctx(child_event, child_ctx);
child_event->attach_state |= PERF_ATTACH_CHILD;
raw_spin_unlock_irqrestore(&child_ctx->lock, flags);
/*
* Link this into the parent event's child list
*/
list_add_tail(&child_event->child_list, &parent_event->child_list);
mutex_unlock(&parent_event->child_mutex);
return child_event;
}
/*
* Inherits an event group.
*
* This will quietly suppress orphaned events; !inherit_event() is not an error.
* This matches with perf_event_release_kernel() removing all child events.
*
* Returns:
* - 0 on success
* - <0 on error
*/
static int inherit_group(struct perf_event *parent_event,
struct task_struct *parent,
struct perf_event_context *parent_ctx,
struct task_struct *child,
struct perf_event_context *child_ctx)
{
struct perf_event *leader;
struct perf_event *sub;
struct perf_event *child_ctr;
leader = inherit_event(parent_event, parent, parent_ctx,
child, NULL, child_ctx);
if (IS_ERR(leader))
return PTR_ERR(leader);
/*
* @leader can be NULL here because of is_orphaned_event(). In this
* case inherit_event() will create individual events, similar to what
* perf_group_detach() would do anyway.
*/
for_each_sibling_event(sub, parent_event) {
child_ctr = inherit_event(sub, parent, parent_ctx,
child, leader, child_ctx);
if (IS_ERR(child_ctr))
return PTR_ERR(child_ctr);
if (sub->aux_event == parent_event && child_ctr &&
!perf_get_aux_event(child_ctr, leader))
return -EINVAL;
}
if (leader)
leader->group_generation = parent_event->group_generation;
return 0;
}
/*
* Creates the child task context and tries to inherit the event-group.
*
* Clears @inherited_all on !attr.inherited or error. Note that we'll leave
* inherited_all set when we 'fail' to inherit an orphaned event; this is
* consistent with perf_event_release_kernel() removing all child events.
*
* Returns:
* - 0 on success
* - <0 on error
*/
static int
inherit_task_group(struct perf_event *event, struct task_struct *parent,
struct perf_event_context *parent_ctx,
struct task_struct *child,
u64 clone_flags, int *inherited_all)
{
struct perf_event_context *child_ctx;
int ret;
if (!event->attr.inherit ||
(event->attr.inherit_thread && !(clone_flags & CLONE_THREAD)) ||
/* Do not inherit if sigtrap and signal handlers were cleared. */
(event->attr.sigtrap && (clone_flags & CLONE_CLEAR_SIGHAND))) {
*inherited_all = 0;
return 0;
}
child_ctx = child->perf_event_ctxp;
if (!child_ctx) {
/*
* This is executed from the parent task context, so
* inherit events that have been marked for cloning.
* First allocate and initialize a context for the
* child.
*/
child_ctx = alloc_perf_context(child);
if (!child_ctx)
return -ENOMEM;
child->perf_event_ctxp = child_ctx;
}
ret = inherit_group(event, parent, parent_ctx, child, child_ctx);
if (ret)
*inherited_all = 0;
return ret;
}
/*
* Initialize the perf_event context in task_struct
*/
static int perf_event_init_context(struct task_struct *child, u64 clone_flags)
{
struct perf_event_context *child_ctx, *parent_ctx;
struct perf_event_context *cloned_ctx;
struct perf_event *event;
struct task_struct *parent = current;
int inherited_all = 1;
unsigned long flags;
int ret = 0;
if (likely(!parent->perf_event_ctxp))
return 0;
/*
* If the parent's context is a clone, pin it so it won't get
* swapped under us.
*/
parent_ctx = perf_pin_task_context(parent);
if (!parent_ctx)
return 0;
/*
* No need to check if parent_ctx != NULL here; since we saw
* it non-NULL earlier, the only reason for it to become NULL
* is if we exit, and since we're currently in the middle of
* a fork we can't be exiting at the same time.
*/
/*
* Lock the parent list. No need to lock the child - not PID
* hashed yet and not running, so nobody can access it.
*/
mutex_lock(&parent_ctx->mutex);
/*
* We dont have to disable NMIs - we are only looking at
* the list, not manipulating it:
*/
perf_event_groups_for_each(event, &parent_ctx->pinned_groups) {
ret = inherit_task_group(event, parent, parent_ctx,
child, clone_flags, &inherited_all);
if (ret)
goto out_unlock;
}
/*
* We can't hold ctx->lock when iterating the ->flexible_group list due
* to allocations, but we need to prevent rotation because
* rotate_ctx() will change the list from interrupt context.
*/
raw_spin_lock_irqsave(&parent_ctx->lock, flags);
parent_ctx->rotate_disable = 1;
raw_spin_unlock_irqrestore(&parent_ctx->lock, flags);
perf_event_groups_for_each(event, &parent_ctx->flexible_groups) {
ret = inherit_task_group(event, parent, parent_ctx,
child, clone_flags, &inherited_all);
if (ret)
goto out_unlock;
}
raw_spin_lock_irqsave(&parent_ctx->lock, flags);
parent_ctx->rotate_disable = 0;
child_ctx = child->perf_event_ctxp;
if (child_ctx && inherited_all) {
/*
* Mark the child context as a clone of the parent
* context, or of whatever the parent is a clone of.
*
* Note that if the parent is a clone, the holding of
* parent_ctx->lock avoids it from being uncloned.
*/
cloned_ctx = parent_ctx->parent_ctx;
if (cloned_ctx) {
child_ctx->parent_ctx = cloned_ctx;
child_ctx->parent_gen = parent_ctx->parent_gen;
} else {
child_ctx->parent_ctx = parent_ctx;
child_ctx->parent_gen = parent_ctx->generation;
}
get_ctx(child_ctx->parent_ctx);
}
raw_spin_unlock_irqrestore(&parent_ctx->lock, flags);
out_unlock:
mutex_unlock(&parent_ctx->mutex);
perf_unpin_context(parent_ctx);
put_ctx(parent_ctx);
return ret;
}
/*
* Initialize the perf_event context in task_struct
*/
int perf_event_init_task(struct task_struct *child, u64 clone_flags)
{
int ret;
memset(child->perf_recursion, 0, sizeof(child->perf_recursion));
child->perf_event_ctxp = NULL;
mutex_init(&child->perf_event_mutex);
INIT_LIST_HEAD(&child->perf_event_list);
child->perf_ctx_data = NULL;
ret = perf_event_init_context(child, clone_flags);
if (ret) {
perf_event_free_task(child);
return ret;
}
return 0;
}
static void __init perf_event_init_all_cpus(void)
{
struct swevent_htable *swhash;
struct perf_cpu_context *cpuctx;
int cpu;
zalloc_cpumask_var(&perf_online_mask, GFP_KERNEL);
zalloc_cpumask_var(&perf_online_core_mask, GFP_KERNEL);
zalloc_cpumask_var(&perf_online_die_mask, GFP_KERNEL);
zalloc_cpumask_var(&perf_online_cluster_mask, GFP_KERNEL);
zalloc_cpumask_var(&perf_online_pkg_mask, GFP_KERNEL);
zalloc_cpumask_var(&perf_online_sys_mask, GFP_KERNEL);
for_each_possible_cpu(cpu) {
swhash = &per_cpu(swevent_htable, cpu);
mutex_init(&swhash->hlist_mutex);
INIT_LIST_HEAD(&per_cpu(pmu_sb_events.list, cpu));
raw_spin_lock_init(&per_cpu(pmu_sb_events.lock, cpu));
INIT_LIST_HEAD(&per_cpu(sched_cb_list, cpu));
cpuctx = per_cpu_ptr(&perf_cpu_context, cpu);
__perf_event_init_context(&cpuctx->ctx);
lockdep_set_class(&cpuctx->ctx.mutex, &cpuctx_mutex);
lockdep_set_class(&cpuctx->ctx.lock, &cpuctx_lock);
cpuctx->online = cpumask_test_cpu(cpu, perf_online_mask);
cpuctx->heap_size = ARRAY_SIZE(cpuctx->heap_default);
cpuctx->heap = cpuctx->heap_default;
}
}
static void perf_swevent_init_cpu(unsigned int cpu)
{
struct swevent_htable *swhash = &per_cpu(swevent_htable, cpu);
mutex_lock(&swhash->hlist_mutex);
if (swhash->hlist_refcount > 0 && !swevent_hlist_deref(swhash)) {
struct swevent_hlist *hlist;
hlist = kzalloc_node(sizeof(*hlist), GFP_KERNEL, cpu_to_node(cpu));
WARN_ON(!hlist);
rcu_assign_pointer(swhash->swevent_hlist, hlist);
}
mutex_unlock(&swhash->hlist_mutex);
}
#if defined CONFIG_HOTPLUG_CPU || defined CONFIG_KEXEC_CORE
static void __perf_event_exit_context(void *__info)
{
struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context);
struct perf_event_context *ctx = __info;
struct perf_event *event;
raw_spin_lock(&ctx->lock);
ctx_sched_out(ctx, NULL, EVENT_TIME);
list_for_each_entry(event, &ctx->event_list, event_entry)
__perf_remove_from_context(event, cpuctx, ctx, (void *)DETACH_GROUP);
raw_spin_unlock(&ctx->lock);
}
static void perf_event_clear_cpumask(unsigned int cpu)
{
int target[PERF_PMU_MAX_SCOPE];
unsigned int scope;
struct pmu *pmu;
cpumask_clear_cpu(cpu, perf_online_mask);
for (scope = PERF_PMU_SCOPE_NONE + 1; scope < PERF_PMU_MAX_SCOPE; scope++) {
const struct cpumask *cpumask = perf_scope_cpu_topology_cpumask(scope, cpu);
struct cpumask *pmu_cpumask = perf_scope_cpumask(scope);
target[scope] = -1;
if (WARN_ON_ONCE(!pmu_cpumask || !cpumask))
continue;
if (!cpumask_test_and_clear_cpu(cpu, pmu_cpumask))
continue;
target[scope] = cpumask_any_but(cpumask, cpu);
if (target[scope] < nr_cpu_ids)
cpumask_set_cpu(target[scope], pmu_cpumask);
}
/* migrate */
list_for_each_entry(pmu, &pmus, entry) {
if (pmu->scope == PERF_PMU_SCOPE_NONE ||
WARN_ON_ONCE(pmu->scope >= PERF_PMU_MAX_SCOPE))
continue;
if (target[pmu->scope] >= 0 && target[pmu->scope] < nr_cpu_ids)
perf_pmu_migrate_context(pmu, cpu, target[pmu->scope]);
}
}
static void perf_event_exit_cpu_context(int cpu)
{
struct perf_cpu_context *cpuctx;
struct perf_event_context *ctx;
// XXX simplify cpuctx->online
mutex_lock(&pmus_lock);
/*
* Clear the cpumasks, and migrate to other CPUs if possible.
* Must be invoked before the __perf_event_exit_context.
*/
perf_event_clear_cpumask(cpu);
cpuctx = per_cpu_ptr(&perf_cpu_context, cpu);
ctx = &cpuctx->ctx;
mutex_lock(&ctx->mutex);
smp_call_function_single(cpu, __perf_event_exit_context, ctx, 1);
cpuctx->online = 0;
mutex_unlock(&ctx->mutex);
mutex_unlock(&pmus_lock);
}
#else
static void perf_event_exit_cpu_context(int cpu) { }
#endif
static void perf_event_setup_cpumask(unsigned int cpu)
{
struct cpumask *pmu_cpumask;
unsigned int scope;
/*
* Early boot stage, the cpumask hasn't been set yet.
* The perf_online_<domain>_masks includes the first CPU of each domain.
* Always unconditionally set the boot CPU for the perf_online_<domain>_masks.
*/
if (cpumask_empty(perf_online_mask)) {
for (scope = PERF_PMU_SCOPE_NONE + 1; scope < PERF_PMU_MAX_SCOPE; scope++) {
pmu_cpumask = perf_scope_cpumask(scope);
if (WARN_ON_ONCE(!pmu_cpumask))
continue;
cpumask_set_cpu(cpu, pmu_cpumask);
}
goto end;
}
for (scope = PERF_PMU_SCOPE_NONE + 1; scope < PERF_PMU_MAX_SCOPE; scope++) {
const struct cpumask *cpumask = perf_scope_cpu_topology_cpumask(scope, cpu);
pmu_cpumask = perf_scope_cpumask(scope);
if (WARN_ON_ONCE(!pmu_cpumask || !cpumask))
continue;
if (!cpumask_empty(cpumask) &&
cpumask_any_and(pmu_cpumask, cpumask) >= nr_cpu_ids)
cpumask_set_cpu(cpu, pmu_cpumask);
}
end:
cpumask_set_cpu(cpu, perf_online_mask);
}
int perf_event_init_cpu(unsigned int cpu)
{
struct perf_cpu_context *cpuctx;
struct perf_event_context *ctx;
perf_swevent_init_cpu(cpu);
mutex_lock(&pmus_lock);
perf_event_setup_cpumask(cpu);
cpuctx = per_cpu_ptr(&perf_cpu_context, cpu);
ctx = &cpuctx->ctx;
mutex_lock(&ctx->mutex);
cpuctx->online = 1;
mutex_unlock(&ctx->mutex);
mutex_unlock(&pmus_lock);
return 0;
}
int perf_event_exit_cpu(unsigned int cpu)
{
perf_event_exit_cpu_context(cpu);
return 0;
}
static int
perf_reboot(struct notifier_block *notifier, unsigned long val, void *v)
{
int cpu;
for_each_online_cpu(cpu)
perf_event_exit_cpu(cpu);
return NOTIFY_OK;
}
/*
* Run the perf reboot notifier at the very last possible moment so that
* the generic watchdog code runs as long as possible.
*/
static struct notifier_block perf_reboot_notifier = {
.notifier_call = perf_reboot,
.priority = INT_MIN,
};
void __init perf_event_init(void)
{
int ret;
idr_init(&pmu_idr);
unwind_deferred_init(&perf_unwind_work,
perf_unwind_deferred_callback);
perf_event_init_all_cpus();
init_srcu_struct(&pmus_srcu);
perf_pmu_register(&perf_swevent, "software", PERF_TYPE_SOFTWARE);
perf_pmu_register(&perf_cpu_clock, "cpu_clock", -1);
perf_pmu_register(&perf_task_clock, "task_clock", -1);
perf_tp_register();
perf_event_init_cpu(smp_processor_id());
register_reboot_notifier(&perf_reboot_notifier);
ret = init_hw_breakpoint();
WARN(ret, "hw_breakpoint initialization failed with: %d", ret);
perf_event_cache = KMEM_CACHE(perf_event, SLAB_PANIC);
/*
* Build time assertion that we keep the data_head at the intended
* location. IOW, validation we got the __reserved[] size right.
*/
BUILD_BUG_ON((offsetof(struct perf_event_mmap_page, data_head))
!= 1024);
}
ssize_t perf_event_sysfs_show(struct device *dev, struct device_attribute *attr,
char *page)
{
struct perf_pmu_events_attr *pmu_attr =
container_of(attr, struct perf_pmu_events_attr, attr);
if (pmu_attr->event_str)
return sprintf(page, "%s\n", pmu_attr->event_str);
return 0;
}
EXPORT_SYMBOL_GPL(perf_event_sysfs_show);
static int __init perf_event_sysfs_init(void)
{
struct pmu *pmu;
int ret;
mutex_lock(&pmus_lock);
ret = bus_register(&pmu_bus);
if (ret)
goto unlock;
list_for_each_entry(pmu, &pmus, entry) {
if (pmu->dev)
continue;
ret = pmu_dev_alloc(pmu);
WARN(ret, "Failed to register pmu: %s, reason %d\n", pmu->name, ret);
}
pmu_bus_running = 1;
ret = 0;
unlock:
mutex_unlock(&pmus_lock);
return ret;
}
device_initcall(perf_event_sysfs_init);
#ifdef CONFIG_CGROUP_PERF
static struct cgroup_subsys_state *
perf_cgroup_css_alloc(struct cgroup_subsys_state *parent_css)
{
struct perf_cgroup *jc;
jc = kzalloc(sizeof(*jc), GFP_KERNEL);
if (!jc)
return ERR_PTR(-ENOMEM);
jc->info = alloc_percpu(struct perf_cgroup_info);
if (!jc->info) {
kfree(jc);
return ERR_PTR(-ENOMEM);
}
return &jc->css;
}
static void perf_cgroup_css_free(struct cgroup_subsys_state *css)
{
struct perf_cgroup *jc = container_of(css, struct perf_cgroup, css);
free_percpu(jc->info);
kfree(jc);
}
static int perf_cgroup_css_online(struct cgroup_subsys_state *css)
{
perf_event_cgroup(css->cgroup);
return 0;
}
static int __perf_cgroup_move(void *info)
{
struct task_struct *task = info;
preempt_disable();
perf_cgroup_switch(task);
preempt_enable();
return 0;
}
static void perf_cgroup_attach(struct cgroup_taskset *tset)
{
struct task_struct *task;
struct cgroup_subsys_state *css;
cgroup_taskset_for_each(task, css, tset)
task_function_call(task, __perf_cgroup_move, task);
}
struct cgroup_subsys perf_event_cgrp_subsys = {
.css_alloc = perf_cgroup_css_alloc,
.css_free = perf_cgroup_css_free,
.css_online = perf_cgroup_css_online,
.attach = perf_cgroup_attach,
/*
* Implicitly enable on dfl hierarchy so that perf events can
* always be filtered by cgroup2 path as long as perf_event
* controller is not mounted on a legacy hierarchy.
*/
.implicit_on_dfl = true,
.threaded = true,
};
#endif /* CONFIG_CGROUP_PERF */
DEFINE_STATIC_CALL_RET0(perf_snapshot_branch_stack, perf_snapshot_branch_stack_t);
]
|
| 137/4 |
2026/05/25 22:18 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 182466
output: 44
thoughts: 3227
Thoughts: **Analyzing Tracepoint Event Permissions in perf_event_open**
Okay, so I've been diving into the `perf_event_open` system call, specifically focusing on how it handles tracepoint events and the implications for unprivileged users. My analysis started with the initial checks in `kernel/events/core.c`, specifically how the code deals with task-specific and CPU-wide events. I was immediately drawn to the permission checks implemented by `perf_check_permission` and `perf_allow_cpu`, which quickly pointed to `perfmon_capable()` and `sysctl_perf_event_paranoid` as the key factors.
The system call has a check for the `task` (PID) that the event should be monitoring. If a task is specified, `perf_check_permission` is invoked. This check determines if the requesting process has permission to access the specified `task`. Otherwise, if `pid == -1`, it's a CPU-wide event, and `perf_allow_cpu` is called. This function enforces root privileges via `sysctl_perf_event_paranoid` when allowing access.
The main focus then shifted to how tracepoints are handled. Tracepoint PMUs use `perf_tp_event_init`, and the function `perf_trace_init` is invoked, which calls `perf_trace_event_init` which calls `perf_trace_event_perm`. This routine is responsible for enforcing additional security measures for tracepoints. `perf_trace_event_perm` checks the `sample_type` attribute, specifically `PERF_SAMPLE_RAW`. If this is NOT requested, there's no problem. However, if `PERF_SAMPLE_RAW` is requested, it calls `perf_allow_tracepoint`. `perf_allow_tracepoint` introduces an additional security check which depends on `sysctl_perf_event_paranoid`. This setting effectively controls whether unprivileged users can access tracepoints with raw data or only counters.
The behavior is this: unprivileged users, if not `perfmon_capable()`, need to satisfy several conditions:
1. They cannot request `PERF_SAMPLE_RAW` to prevent information leaks.
2. If `PERF_SAMPLE_RAW` IS requested, the value of `sysctl_perf_event_paranoid` decides access.
Now, a key point emerged. Even if `PERF_SAMPLE_RAW` is not requested, there's another check later in `perf_event_open`: `perf_allow_kernel()`. Tracepoints reside in the kernel. `perf_allow_kernel` is also affected by `sysctl_perf_event_paranoid` so it is also an access control gate, even when `PERF_SAMPLE_RAW` is NOT requested. I saw that `perf_tp_event_match` uses `exclude_kernel`. I found that tracepoints might be configured with `exclude_kernel`. This allows an administrator to constrain tracing to either user space or kernel space. If set to 1, this means that the event only fires for user space. So I started reasoning about the interplay between `exclude_kernel` and `sysctl_perf_event_paranoid`.
Now, if `sysctl_perf_event_paranoid` is set to the default value of 2, unprivileged users must also set `exclude_kernel = 1` in order to succeed. Setting this prevents a regular kernel tracepoint from firing. However, uprobes also use the `PERF_TYPE_TRACEPOINT` type. Therefore, an unprivileged user can open an uprobe (a type of dynamic tracepoint) event if `exclude_kernel = 1`.
So the answer is that unprivileged users can still open tracepoints. It depends on `sysctl_perf_event_paranoid` and `exclude_kernel`. If I don't request `PERF_SAMPLE_RAW`, I get just a counter.
I re-read `perf_allow_tracepoint` to confirm my understanding of the relationship between `sysctl_perf_event_paranoid` and `perfmon_capable()`. The logic confirmed that even when `paranoid` is set to 0, 1, or 2, and the user is not `perfmon_capable()`, access is restricted. The exception to this rule is when `sysctl_perf_event_paranoid` is set to -1, which bypasses the `perfmon_capable()` check.
I considered what happens if the user requests other sample types, such as `PERF_SAMPLE_IP | PERF_SAMPLE_TIME`, *without* `PERF_SAMPLE_RAW`. Then, by the design of `perf_trace_event_perm`, I can indeed open and potentially fire a kernel trace
|
| 138/4 |
2026/05/25 22:19 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/events/core.c IncludeLines:true Name:perf_uprobe_event_init]
Results: map[SourceCode:11219: static int perf_uprobe_event_init(struct perf_event *event)
11220: {
11221: int err;
11222: unsigned long ref_ctr_offset;
11223: bool is_retprobe;
11224:
11225: if (event->attr.type != perf_uprobe.type)
11226: return -ENOENT;
11227:
11228: if (!capable(CAP_SYS_ADMIN))
11229: return -EACCES;
11230:
11231: /*
11232: * no branch sampling for probe events
11233: */
11234: if (has_branch_stack(event))
11235: return -EOPNOTSUPP;
11236:
11237: is_retprobe = event->attr.config & PERF_PROBE_CONFIG_IS_RETPROBE;
11238: ref_ctr_offset = event->attr.config >> PERF_UPROBE_REF_CTR_OFFSET_SHIFT;
11239: err = perf_uprobe_init(event, ref_ctr_offset, is_retprobe);
11240: if (err)
11241: return err;
11242:
11243: event->destroy = perf_uprobe_destroy;
11244:
11245: return 0;
11246: }
SourceFile:kernel/events/core.c]
|
| 139/4 |
2026/05/25 22:19 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 187212
output: 23
thoughts: 371
Thoughts: **Analyzing Uprobe Permissions**
Okay, so I know `perf_uprobe_event_init` definitely needs `CAP_SYS_ADMIN`. That much is clear. But hang on, what about these "legacy" uprobes? The ones created through `debugfs` or `tracefs` and then opened using `perf_event_open` with their ID? That's a different beast, isn't it?
In that case, the Performance Monitoring Unit (PMU) used isn't `perf_uprobe`, it's `perf_tracepoint`. `perf_uprobe` is for those dynamic uprobes created directly via `perf_event_open` with `attr.type == perf_uprobe.type`.
So, if I use `PERF_TYPE_TRACEPOINT` and feed it the ID of an existing uprobe, what happens? My thinking goes like this: `perf_tp_event_init` gets called, then that calls `perf_trace_init`, and finally `perf_trace_event_perm`. It's `perf_trace_event_perm` where the permissions check actually happens.
Looking at the code there, particularly around line 79, there's a check for whether the event is per-task (`p_event->attach_state == PERF_ATTACH_TASK`) and if the tracepoint has `TRACE_EVENT_FL_CAP_ANY`. If both are true, then it's permitted, even for a non-root user.
My question is, do uprobes, which often have `TRACE_EVENT_FL_UPROBE`, *also* have `TRACE_EVENT_FL_CAP_ANY`? It's the key to understanding if this legacy uprobe method is vulnerable. Time to check where `TRACE_EVENT_FL_CAP_ANY` is actually used in the codebase. I need to track down the callers.
|
| 140/4 |
2026/05/25 22:19 |
tool |
grepper |
0m
Args: map[Expression:TRACE_EVENT_FL_CAP_ANY]
Results: map[Output:include/linux/syscalls.h=145=extern struct trace_event_functions exit_syscall_print_funcs;
--
include/linux/syscalls.h-156- .data = (void *)&__syscall_meta_##sname,\
include/linux/syscalls.h:157: .flags = TRACE_EVENT_FL_CAP_ANY, \
include/linux/syscalls.h-158- }; \
--
include/linux/syscalls.h-172- .data = (void *)&__syscall_meta_##sname,\
include/linux/syscalls.h:173: .flags = TRACE_EVENT_FL_CAP_ANY, \
include/linux/syscalls.h-174- }; \
--
include/linux/trace_events.h=318=enum {
include/linux/trace_events.h:319: TRACE_EVENT_FL_CAP_ANY_BIT,
include/linux/trace_events.h-320- TRACE_EVENT_FL_NO_SET_FILTER_BIT,
--
include/linux/trace_events.h=348=enum {
include/linux/trace_events.h:349: TRACE_EVENT_FL_CAP_ANY = (1 << TRACE_EVENT_FL_CAP_ANY_BIT),
include/linux/trace_events.h-350- TRACE_EVENT_FL_NO_SET_FILTER = (1 << TRACE_EVENT_FL_NO_SET_FILTER_BIT),
--
include/trace/events/syscalls.h=18=TRACE_EVENT_SYSCALL(sys_enter,
--
include/trace/events/syscalls.h-41-
include/trace/events/syscalls.h:42:TRACE_EVENT_FLAGS(sys_enter, TRACE_EVENT_FL_CAP_ANY)
include/trace/events/syscalls.h-43-
include/trace/events/syscalls.h=44=TRACE_EVENT_SYSCALL(sys_exit,
--
include/trace/events/syscalls.h-65-
include/trace/events/syscalls.h:66:TRACE_EVENT_FLAGS(sys_exit, TRACE_EVENT_FL_CAP_ANY)
include/trace/events/syscalls.h-67-
--
kernel/trace/trace_event_perf.c=27=static int perf_trace_event_perm(struct trace_event_call *tp_event,
--
kernel/trace/trace_event_perf.c-80- if (p_event->attach_state == PERF_ATTACH_TASK) {
kernel/trace/trace_event_perf.c:81: if (tp_event->flags & TRACE_EVENT_FL_CAP_ANY)
kernel/trace/trace_event_perf.c-82- return 0;
--
kernel/trace/trace_uprobe.c=1600=static inline void init_trace_event_call(struct trace_uprobe *tu)
--
kernel/trace/trace_uprobe.c-1605-
kernel/trace/trace_uprobe.c:1606: call->flags = TRACE_EVENT_FL_UPROBE | TRACE_EVENT_FL_CAP_ANY;
kernel/trace/trace_uprobe.c-1607- call->class->reg = trace_uprobe_register;
]
|
| 141/4 |
2026/05/25 22:19 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 188864
output: 46
thoughts: 426
Thoughts: **Unveiling a Security Nuance in Linux Tracing**
Aha! This is interesting. I've been digging into the kernel's tracing capabilities, specifically around `perf` and event permissions. I've noticed something that could have security implications.
Initially, I've confirmed that both syscall tracepoints (`sys_enter`, `sys_exit`) and Uprobes have the `TRACE_EVENT_FL_CAP_ANY` flag set. This is a key observation.
This immediately makes me consider what that means in the context of per-task events, like when a tracepoint is attached to a specific process (PERF_ATTACH_TASK). With this flag, an unprivileged user *should* be able to open and trace:
* Syscall tracepoints.
* Uprobes.
Even when sampling raw data with `PERF_SAMPLE_RAW`. That is what is interesting here.
That's an interesting combination - but let's see why, specifically. The function `perf_trace_event_perm` is the reason. At line 82, with `TRACE_EVENT_FL_CAP_ANY`, the function returns 0.
```c
79: /* Some events are ok to be traced by non-root users... */
80: if (p_event->attach_state == PERF_ATTACH_TASK) {
81: if (tp_event->flags & TRACE_EVENT_FL_CAP_ANY)
82: return 0;
83: }
```
Okay, but what about the broader security context of `perf_allow_kernel()` that's called during `perf_event_open`?
If `paranoid` is set to 2, it's expected that `perf_allow_kernel()` would fail, unless `exclude_kernel = 1`. Since syscall tracepoints reside within the kernel, I'd expect they'd be off-limits. However, if I set `exclude_kernel = 1`, would they actually fire? Let me look deeper.
Checking `perf_tp_event_match`, I see this:
```c
10977: if (event->attr.exclude_kernel && !user_mode(regs))
10978: return 0;
```
Right, but what's really happening during a syscall trace? A syscall is initiated from user-space, but the tracepoint fires within the kernel. The `regs` argument passed to `perf_tp_event` for syscalls are usually the user-space registers. That looks like the key.
Let me confirm that by looking at where the tracepoints are defined. Specifically, `perf_trace_sys_enter` and `perf_trace_sys_exit`. Yup, they're in `kernel/trace/trace_syscalls.c`.
I need to circle back and test this further...
|
| 142/4 |
2026/05/25 22:19 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/trace_syscalls.c IncludeLines:true Name:perf_trace_sys_enter]
Results: map[SourceCode: 18: TRACE_EVENT_SYSCALL(sys_enter,
SourceFile:include/trace/events/syscalls.h]
|
| 143/4 |
2026/05/25 22:19 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 190028
output: 25
thoughts: 42
Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience:
**Tracing System Call Entry: Where Does This Live?**
Alright, so I'm digging into `perf_trace_sys_enter`, and I know from experience that this likely isn't a directly written function. More often than not, something like this is built through macros, usually for efficiency and code generation purposes. That means I need to hunt down the macro expansion to understand the underlying implementation. My first port of call, and where I'd expect to find the central definitions and infrastructure for system call tracing, would be `kernel/trace/trace_syscalls.c`. That file's going to hold the key to figuring out exactly what's going on under the hood with that specific entry point. I'll start there and trace the macros.
|
| 144/4 |
2026/05/25 22:19 |
tool |
read-file |
0m
Args: map[File:kernel/trace/trace_syscalls.c]
Results: map[Contents:// SPDX-License-Identifier: GPL-2.0
#include <trace/syscall.h>
#include <trace/events/syscalls.h>
#include <linux/kernel_stat.h>
#include <linux/syscalls.h>
#include <linux/slab.h>
#include <linux/kernel.h>
#include <linux/module.h> /* for MODULE_NAME_LEN via KSYM_SYMBOL_LEN */
#include <linux/ftrace.h>
#include <linux/perf_event.h>
#include <linux/xarray.h>
#include <asm/syscall.h>
#include "trace_output.h"
#include "trace.h"
static DEFINE_MUTEX(syscall_trace_lock);
static int syscall_enter_register(struct trace_event_call *event,
enum trace_reg type, void *data);
static int syscall_exit_register(struct trace_event_call *event,
enum trace_reg type, void *data);
static struct list_head *
syscall_get_enter_fields(struct trace_event_call *call)
{
struct syscall_metadata *entry = call->data;
return &entry->enter_fields;
}
extern struct syscall_metadata *__start_syscalls_metadata[];
extern struct syscall_metadata *__stop_syscalls_metadata[];
static DEFINE_XARRAY(syscalls_metadata_sparse);
static struct syscall_metadata **syscalls_metadata;
#ifndef ARCH_HAS_SYSCALL_MATCH_SYM_NAME
static inline bool arch_syscall_match_sym_name(const char *sym, const char *name)
{
/*
* Only compare after the "sys" prefix. Archs that use
* syscall wrappers may have syscalls symbols aliases prefixed
* with ".SyS" or ".sys" instead of "sys", leading to an unwanted
* mismatch.
*/
return !strcmp(sym + 3, name + 3);
}
#endif
#ifdef ARCH_TRACE_IGNORE_COMPAT_SYSCALLS
/*
* Some architectures that allow for 32bit applications
* to run on a 64bit kernel, do not map the syscalls for
* the 32bit tasks the same as they do for 64bit tasks.
*
* *cough*x86*cough*
*
* In such a case, instead of reporting the wrong syscalls,
* simply ignore them.
*
* For an arch to ignore the compat syscalls it needs to
* define ARCH_TRACE_IGNORE_COMPAT_SYSCALLS as well as
* define the function arch_trace_is_compat_syscall() to let
* the tracing system know that it should ignore it.
*/
static int
trace_get_syscall_nr(struct task_struct *task, struct pt_regs *regs)
{
if (unlikely(arch_trace_is_compat_syscall(regs)))
return -1;
return syscall_get_nr(task, regs);
}
#else
static inline int
trace_get_syscall_nr(struct task_struct *task, struct pt_regs *regs)
{
return syscall_get_nr(task, regs);
}
#endif /* ARCH_TRACE_IGNORE_COMPAT_SYSCALLS */
static __init struct syscall_metadata *
find_syscall_meta(unsigned long syscall)
{
struct syscall_metadata **start;
struct syscall_metadata **stop;
char str[KSYM_SYMBOL_LEN];
start = __start_syscalls_metadata;
stop = __stop_syscalls_metadata;
kallsyms_lookup(syscall, NULL, NULL, NULL, str);
if (arch_syscall_match_sym_name(str, "sys_ni_syscall"))
return NULL;
for ( ; start < stop; start++) {
if ((*start)->name && arch_syscall_match_sym_name(str, (*start)->name))
return *start;
}
return NULL;
}
static struct syscall_metadata *syscall_nr_to_meta(int nr)
{
if (IS_ENABLED(CONFIG_HAVE_SPARSE_SYSCALL_NR))
return xa_load(&syscalls_metadata_sparse, (unsigned long)nr);
if (!syscalls_metadata || nr >= NR_syscalls || nr < 0)
return NULL;
return syscalls_metadata[nr];
}
const char *get_syscall_name(int syscall)
{
struct syscall_metadata *entry;
entry = syscall_nr_to_meta(syscall);
if (!entry)
return NULL;
return entry->name;
}
/* Added to user strings or arrays when max limit is reached */
#define EXTRA "..."
static void get_dynamic_len_ptr(struct syscall_trace_enter *trace,
struct syscall_metadata *entry,
int *offset_p, int *len_p, unsigned char **ptr_p)
{
unsigned char *ptr;
int offset = *offset_p;
int val;
/* This arg points to a user space string */
ptr = (void *)trace->args + sizeof(long) * entry->nb_args + offset;
val = *(int *)ptr;
/* The value is a dynamic string (len << 16 | offset) */
ptr = (void *)trace + (val & 0xffff);
*len_p = val >> 16;
offset += 4;
*ptr_p = ptr;
*offset_p = offset;
}
static enum print_line_t
sys_enter_openat_print(struct syscall_trace_enter *trace, struct syscall_metadata *entry,
struct trace_seq *s, struct trace_event *event)
{
unsigned char *ptr;
int offset = 0;
int bits, len;
bool done = false;
static const struct trace_print_flags __flags[] =
{
{ O_TMPFILE, "O_TMPFILE" },
{ O_WRONLY, "O_WRONLY" },
{ O_RDWR, "O_RDWR" },
{ O_CREAT, "O_CREAT" },
{ O_EXCL, "O_EXCL" },
{ O_NOCTTY, "O_NOCTTY" },
{ O_TRUNC, "O_TRUNC" },
{ O_APPEND, "O_APPEND" },
{ O_NONBLOCK, "O_NONBLOCK" },
{ O_DSYNC, "O_DSYNC" },
{ O_DIRECT, "O_DIRECT" },
{ O_LARGEFILE, "O_LARGEFILE" },
{ O_DIRECTORY, "O_DIRECTORY" },
{ O_NOFOLLOW, "O_NOFOLLOW" },
{ O_NOATIME, "O_NOATIME" },
{ O_CLOEXEC, "O_CLOEXEC" },
{ -1, NULL }
};
trace_seq_printf(s, "%s(", entry->name);
for (int i = 0; !done && i < entry->nb_args; i++) {
if (trace_seq_has_overflowed(s))
goto end;
if (i)
trace_seq_puts(s, ", ");
switch (i) {
case 2:
bits = trace->args[2];
trace_seq_puts(s, "flags: ");
/* No need to show mode when not creating the file */
if (!(bits & (O_CREAT|O_TMPFILE)))
done = true;
if (!(bits & O_ACCMODE)) {
if (!bits) {
trace_seq_puts(s, "O_RDONLY");
continue;
}
trace_seq_puts(s, "O_RDONLY|");
}
trace_print_flags_seq(s, "|", bits, __flags);
/*
* trace_print_flags_seq() adds a '\0' to the
* buffer, but this needs to append more to the seq.
*/
if (!trace_seq_has_overflowed(s))
trace_seq_pop(s);
continue;
case 3:
trace_seq_printf(s, "%s: 0%03o", entry->args[i],
(unsigned int)trace->args[i]);
continue;
}
trace_seq_printf(s, "%s: %lu", entry->args[i],
trace->args[i]);
if (!(BIT(i) & entry->user_mask))
continue;
get_dynamic_len_ptr(trace, entry, &offset, &len, &ptr);
trace_seq_printf(s, " \"%.*s\"", len, ptr);
}
trace_seq_putc(s, ')');
end:
trace_seq_putc(s, '\n');
return trace_handle_return(s);
}
static enum print_line_t
print_syscall_enter(struct trace_iterator *iter, int flags,
struct trace_event *event)
{
struct trace_array *tr = iter->tr;
struct trace_seq *s = &iter->seq;
struct trace_entry *ent = iter->ent;
struct syscall_trace_enter *trace;
struct syscall_metadata *entry;
int i, syscall, val, len;
unsigned char *ptr;
int offset = 0;
trace = (typeof(trace))ent;
syscall = trace->nr;
entry = syscall_nr_to_meta(syscall);
if (!entry)
goto end;
if (entry->enter_event->event.type != ent->type) {
WARN_ON_ONCE(1);
goto end;
}
switch (entry->syscall_nr) {
case __NR_openat:
if (!tr || !(tr->trace_flags & TRACE_ITER(VERBOSE)))
return sys_enter_openat_print(trace, entry, s, event);
break;
default:
break;
}
trace_seq_printf(s, "%s(", entry->name);
for (i = 0; i < entry->nb_args; i++) {
bool printable = false;
char *str;
if (trace_seq_has_overflowed(s))
goto end;
if (i)
trace_seq_puts(s, ", ");
/* parameter types */
if (tr && tr->trace_flags & TRACE_ITER(VERBOSE))
trace_seq_printf(s, "%s ", entry->types[i]);
/* parameter values */
if (trace->args[i] < 10)
trace_seq_printf(s, "%s: %lu", entry->args[i],
trace->args[i]);
else
trace_seq_printf(s, "%s: 0x%lx", entry->args[i],
trace->args[i]);
if (!(BIT(i) & entry->user_mask))
continue;
get_dynamic_len_ptr(trace, entry, &offset, &len, &ptr);
if (entry->user_arg_size < 0 || entry->user_arg_is_str) {
trace_seq_printf(s, " \"%.*s\"", len, ptr);
continue;
}
val = trace->args[entry->user_arg_size];
str = ptr;
trace_seq_puts(s, " (");
for (int x = 0; x < len; x++, ptr++) {
if (isascii(*ptr) && isprint(*ptr))
printable = true;
if (x)
trace_seq_putc(s, ':');
trace_seq_printf(s, "%02x", *ptr);
}
if (len < val)
trace_seq_printf(s, ", %s", EXTRA);
trace_seq_putc(s, ')');
/* If nothing is printable, don't bother printing anything */
if (!printable)
continue;
trace_seq_puts(s, " \"");
for (int x = 0; x < len; x++) {
if (isascii(str[x]) && isprint(str[x]))
trace_seq_putc(s, str[x]);
else
trace_seq_putc(s, '.');
}
if (len < val)
trace_seq_printf(s, "\"%s", EXTRA);
else
trace_seq_putc(s, '"');
}
trace_seq_putc(s, ')');
end:
trace_seq_putc(s, '\n');
return trace_handle_return(s);
}
static enum print_line_t
print_syscall_exit(struct trace_iterator *iter, int flags,
struct trace_event *event)
{
struct trace_seq *s = &iter->seq;
struct trace_entry *ent = iter->ent;
struct syscall_trace_exit *trace;
int syscall;
struct syscall_metadata *entry;
trace = (typeof(trace))ent;
syscall = trace->nr;
entry = syscall_nr_to_meta(syscall);
if (!entry) {
trace_seq_putc(s, '\n');
goto out;
}
if (entry->exit_event->event.type != ent->type) {
WARN_ON_ONCE(1);
return TRACE_TYPE_UNHANDLED;
}
trace_seq_printf(s, "%s -> 0x%lx\n", entry->name,
trace->ret);
out:
return trace_handle_return(s);
}
#define SYSCALL_FIELD(_type, _name) { \
.type = #_type, .name = #_name, \
.size = sizeof(_type), .align = __alignof__(_type), \
.is_signed = is_signed_type(_type), .filter_type = FILTER_OTHER }
/* When len=0, we just calculate the needed length */
#define LEN_OR_ZERO (len ? len - pos : 0)
static int __init
sys_enter_openat_print_fmt(struct syscall_metadata *entry, char *buf, int len)
{
int pos = 0;
pos += snprintf(buf + pos, LEN_OR_ZERO,
"\"dfd: 0x%%08lx, filename: 0x%%08lx \\\"%%s\\\", flags: %%s%%s, mode: 0%%03o\",");
pos += snprintf(buf + pos, LEN_OR_ZERO,
" ((unsigned long)(REC->dfd)),");
pos += snprintf(buf + pos, LEN_OR_ZERO,
" ((unsigned long)(REC->filename)),");
pos += snprintf(buf + pos, LEN_OR_ZERO,
" __get_str(__filename_val),");
pos += snprintf(buf + pos, LEN_OR_ZERO,
" (REC->flags & ~3) && !(REC->flags & 3) ? \"O_RDONLY|\" : \"\", ");
pos += snprintf(buf + pos, LEN_OR_ZERO,
" REC->flags ? __print_flags(REC->flags, \"|\", ");
pos += snprintf(buf + pos, LEN_OR_ZERO,
"{ 0x%x, \"O_WRONLY\" }, ", O_WRONLY);
pos += snprintf(buf + pos, LEN_OR_ZERO,
"{ 0x%x, \"O_RDWR\" }, ", O_RDWR);
pos += snprintf(buf + pos, LEN_OR_ZERO,
"{ 0x%x, \"O_CREAT\" }, ", O_CREAT);
pos += snprintf(buf + pos, LEN_OR_ZERO,
"{ 0x%x, \"O_EXCL\" }, ", O_EXCL);
pos += snprintf(buf + pos, LEN_OR_ZERO,
"{ 0x%x, \"O_NOCTTY\" }, ", O_NOCTTY);
pos += snprintf(buf + pos, LEN_OR_ZERO,
"{ 0x%x, \"O_TRUNC\" }, ", O_TRUNC);
pos += snprintf(buf + pos, LEN_OR_ZERO,
"{ 0x%x, \"O_APPEND\" }, ", O_APPEND);
pos += snprintf(buf + pos, LEN_OR_ZERO,
"{ 0x%x, \"O_NONBLOCK\" }, ", O_NONBLOCK);
pos += snprintf(buf + pos, LEN_OR_ZERO,
"{ 0x%x, \"O_DSYNC\" }, ", O_DSYNC);
pos += snprintf(buf + pos, LEN_OR_ZERO,
"{ 0x%x, \"O_DIRECT\" }, ", O_DIRECT);
pos += snprintf(buf + pos, LEN_OR_ZERO,
"{ 0x%x, \"O_LARGEFILE\" }, ", O_LARGEFILE);
pos += snprintf(buf + pos, LEN_OR_ZERO,
"{ 0x%x, \"O_DIRECTORY\" }, ", O_DIRECTORY);
pos += snprintf(buf + pos, LEN_OR_ZERO,
"{ 0x%x, \"O_NOFOLLOW\" }, ", O_NOFOLLOW);
pos += snprintf(buf + pos, LEN_OR_ZERO,
"{ 0x%x, \"O_NOATIME\" }, ", O_NOATIME);
pos += snprintf(buf + pos, LEN_OR_ZERO,
"{ 0x%x, \"O_CLOEXEC\" }) : \"O_RDONLY\", ", O_CLOEXEC);
pos += snprintf(buf + pos, LEN_OR_ZERO,
" ((unsigned long)(REC->mode))");
return pos;
}
static int __init
__set_enter_print_fmt(struct syscall_metadata *entry, char *buf, int len)
{
bool is_string = entry->user_arg_is_str;
int i;
int pos = 0;
switch (entry->syscall_nr) {
case __NR_openat:
return sys_enter_openat_print_fmt(entry, buf, len);
default:
break;
}
pos += snprintf(buf + pos, LEN_OR_ZERO, "\"");
for (i = 0; i < entry->nb_args; i++) {
if (i)
pos += snprintf(buf + pos, LEN_OR_ZERO, ", ");
pos += snprintf(buf + pos, LEN_OR_ZERO, "%s: 0x%%0%zulx",
entry->args[i], sizeof(unsigned long));
if (!(BIT(i) & entry->user_mask))
continue;
/* Add the format for the user space string or array */
if (entry->user_arg_size < 0 || is_string)
pos += snprintf(buf + pos, LEN_OR_ZERO, " \\\"%%s\\\"");
else
pos += snprintf(buf + pos, LEN_OR_ZERO, " (%%s)");
}
pos += snprintf(buf + pos, LEN_OR_ZERO, "\"");
for (i = 0; i < entry->nb_args; i++) {
pos += snprintf(buf + pos, LEN_OR_ZERO,
", ((unsigned long)(REC->%s))", entry->args[i]);
if (!(BIT(i) & entry->user_mask))
continue;
/* The user space data for arg has name __<arg>_val */
if (entry->user_arg_size < 0 || is_string) {
pos += snprintf(buf + pos, LEN_OR_ZERO, ", __get_str(__%s_val)",
entry->args[i]);
} else {
pos += snprintf(buf + pos, LEN_OR_ZERO, ", __print_dynamic_array(__%s_val, 1)",
entry->args[i]);
}
}
#undef LEN_OR_ZERO
/* return the length of print_fmt */
return pos;
}
static int __init set_syscall_print_fmt(struct trace_event_call *call)
{
char *print_fmt;
int len;
struct syscall_metadata *entry = call->data;
if (entry->enter_event != call) {
call->print_fmt = "\"0x%lx\", REC->ret";
return 0;
}
/* First: called with 0 length to calculate the needed length */
len = __set_enter_print_fmt(entry, NULL, 0);
print_fmt = kmalloc(len + 1, GFP_KERNEL);
if (!print_fmt)
return -ENOMEM;
/* Second: actually write the @print_fmt */
__set_enter_print_fmt(entry, print_fmt, len + 1);
call->print_fmt = print_fmt;
return 0;
}
static void __init free_syscall_print_fmt(struct trace_event_call *call)
{
struct syscall_metadata *entry = call->data;
if (entry->enter_event == call)
kfree(call->print_fmt);
}
static int __init syscall_enter_define_fields(struct trace_event_call *call)
{
struct syscall_trace_enter trace;
struct syscall_metadata *meta = call->data;
unsigned long mask;
char *arg;
int offset = offsetof(typeof(trace), args);
int ret = 0;
int len;
int i;
for (i = 0; i < meta->nb_args; i++) {
ret = trace_define_field(call, meta->types[i],
meta->args[i], offset,
sizeof(unsigned long), 0,
FILTER_OTHER);
if (ret)
break;
offset += sizeof(unsigned long);
}
if (ret || !meta->user_mask)
return ret;
mask = meta->user_mask;
while (mask) {
int idx = ffs(mask) - 1;
mask &= ~BIT(idx);
/*
* User space data is faulted into a temporary buffer and then
* added as a dynamic string or array to the end of the event.
* The user space data name for the arg pointer is
* "__<arg>_val".
*/
len = strlen(meta->args[idx]) + sizeof("___val");
arg = kmalloc(len, GFP_KERNEL);
if (WARN_ON_ONCE(!arg)) {
meta->user_mask = 0;
return -ENOMEM;
}
snprintf(arg, len, "__%s_val", meta->args[idx]);
ret = trace_define_field(call, "__data_loc char[]",
arg, offset, sizeof(int), 0,
FILTER_OTHER);
if (ret) {
kfree(arg);
break;
}
offset += 4;
}
return ret;
}
/*
* Create a per CPU temporary buffer to copy user space pointers into.
*
* SYSCALL_FAULT_USER_MAX is the amount to copy from user space.
* (defined in kernel/trace/trace.h)
* SYSCALL_FAULT_ARG_SZ is the amount to copy from user space plus the
* nul terminating byte and possibly appended EXTRA (4 bytes).
*
* SYSCALL_FAULT_BUF_SZ holds the size of the per CPU buffer to use
* to copy memory from user space addresses into that will hold
* 3 args as only 3 args are allowed to be copied from system calls.
*/
#define SYSCALL_FAULT_ARG_SZ (SYSCALL_FAULT_USER_MAX + 1 + 4)
#define SYSCALL_FAULT_MAX_CNT 3
#define SYSCALL_FAULT_BUF_SZ (SYSCALL_FAULT_ARG_SZ * SYSCALL_FAULT_MAX_CNT)
/* Use the tracing per CPU buffer infrastructure to copy from user space */
struct syscall_user_buffer {
struct trace_user_buf_info buf;
struct rcu_head rcu;
};
static struct syscall_user_buffer *syscall_buffer;
static int syscall_fault_buffer_enable(void)
{
struct syscall_user_buffer *sbuf;
int ret;
lockdep_assert_held(&syscall_trace_lock);
if (syscall_buffer) {
trace_user_fault_get(&syscall_buffer->buf);
return 0;
}
sbuf = kmalloc(sizeof(*sbuf), GFP_KERNEL);
if (!sbuf)
return -ENOMEM;
ret = trace_user_fault_init(&sbuf->buf, SYSCALL_FAULT_BUF_SZ);
if (ret < 0) {
kfree(sbuf);
return ret;
}
WRITE_ONCE(syscall_buffer, sbuf);
return 0;
}
static void rcu_free_syscall_buffer(struct rcu_head *rcu)
{
struct syscall_user_buffer *sbuf =
container_of(rcu, struct syscall_user_buffer, rcu);
trace_user_fault_destroy(&sbuf->buf);
kfree(sbuf);
}
static void syscall_fault_buffer_disable(void)
{
struct syscall_user_buffer *sbuf = syscall_buffer;
lockdep_assert_held(&syscall_trace_lock);
if (trace_user_fault_put(&sbuf->buf))
return;
WRITE_ONCE(syscall_buffer, NULL);
call_rcu_tasks_trace(&sbuf->rcu, rcu_free_syscall_buffer);
}
struct syscall_args {
char *ptr_array[SYSCALL_FAULT_MAX_CNT];
int read[SYSCALL_FAULT_MAX_CNT];
int uargs;
};
static int syscall_copy_user(char *buf, const char __user *ptr,
size_t size, void *data)
{
struct syscall_args *args = data;
int ret;
for (int i = 0; i < args->uargs; i++, buf += SYSCALL_FAULT_ARG_SZ) {
ptr = (char __user *)args->ptr_array[i];
ret = strncpy_from_user(buf, ptr, size);
args->read[i] = ret;
}
return 0;
}
static int syscall_copy_user_array(char *buf, const char __user *ptr,
size_t size, void *data)
{
struct syscall_args *args = data;
int ret;
for (int i = 0; i < args->uargs; i++, buf += SYSCALL_FAULT_ARG_SZ) {
ptr = (char __user *)args->ptr_array[i];
ret = __copy_from_user(buf, ptr, size);
args->read[i] = ret ? -1 : size;
}
return 0;
}
static char *sys_fault_user(unsigned int buf_size,
struct syscall_metadata *sys_data,
struct syscall_user_buffer *sbuf,
unsigned long *args,
unsigned int data_size[SYSCALL_FAULT_MAX_CNT])
{
trace_user_buf_copy syscall_copy = syscall_copy_user;
unsigned long mask = sys_data->user_mask;
unsigned long size = SYSCALL_FAULT_ARG_SZ - 1;
struct syscall_args sargs;
bool array = false;
char *buffer;
char *buf;
int ret;
int i = 0;
/* The extra is appended to the user data in the buffer */
BUILD_BUG_ON(SYSCALL_FAULT_USER_MAX + sizeof(EXTRA) >=
SYSCALL_FAULT_ARG_SZ);
/*
* If this system call event has a size argument, use
* it to define how much of user space memory to read,
* and read it as an array and not a string.
*/
if (sys_data->user_arg_size >= 0) {
array = true;
size = args[sys_data->user_arg_size];
if (size > SYSCALL_FAULT_ARG_SZ - 1)
size = SYSCALL_FAULT_ARG_SZ - 1;
syscall_copy = syscall_copy_user_array;
}
while (mask) {
int idx = ffs(mask) - 1;
mask &= ~BIT(idx);
if (WARN_ON_ONCE(i == SYSCALL_FAULT_MAX_CNT))
break;
/* Get the pointer to user space memory to read */
sargs.ptr_array[i++] = (char *)args[idx];
}
sargs.uargs = i;
/* Clear the values that are not used */
for (; i < SYSCALL_FAULT_MAX_CNT; i++) {
data_size[i] = -1; /* Denotes no pointer */
}
/* A zero size means do not even try */
if (!buf_size)
return NULL;
buffer = trace_user_fault_read(&sbuf->buf, NULL, size,
syscall_copy, &sargs);
if (!buffer)
return NULL;
buf = buffer;
for (i = 0; i < sargs.uargs; i++, buf += SYSCALL_FAULT_ARG_SZ) {
ret = sargs.read[i];
if (ret < 0)
continue;
buf[ret] = '\0';
/* For strings, replace any non-printable characters with '.' */
if (!array) {
for (int x = 0; x < ret; x++) {
if (!isprint(buf[x]))
buf[x] = '.';
}
size = min(buf_size, SYSCALL_FAULT_USER_MAX);
/*
* If the text was truncated due to our max limit,
* add "..." to the string.
*/
if (ret > size) {
strscpy(buf + size, EXTRA, sizeof(EXTRA));
ret = size + sizeof(EXTRA);
} else {
buf[ret++] = '\0';
}
} else {
ret = min((unsigned int)ret, buf_size);
}
data_size[i] = ret;
}
return buffer;
}
static int
syscall_get_data(struct syscall_metadata *sys_data, unsigned long *args,
char **buffer, int *size, int *user_sizes, int *uargs,
int buf_size)
{
struct syscall_user_buffer *sbuf;
int i;
/* If the syscall_buffer is NULL, tracing is being shutdown */
sbuf = READ_ONCE(syscall_buffer);
if (!sbuf)
return -1;
*buffer = sys_fault_user(buf_size, sys_data, sbuf, args, user_sizes);
/*
* user_size is the amount of data to append.
* Need to add 4 for the meta field that points to
* the user memory at the end of the event and also
* stores its size.
*/
for (i = 0; i < SYSCALL_FAULT_MAX_CNT; i++) {
if (user_sizes[i] < 0)
break;
*size += user_sizes[i] + 4;
}
/* Save the number of user read arguments of this syscall */
*uargs = i;
return 0;
}
static void syscall_put_data(struct syscall_metadata *sys_data,
struct syscall_trace_enter *entry,
char *buffer, int size, int *user_sizes, int uargs)
{
char *buf = buffer;
void *ptr;
int val;
/*
* Set the pointer to point to the meta data of the event
* that has information about the stored user space memory.
*/
ptr = (void *)entry->args + sizeof(unsigned long) * sys_data->nb_args;
/*
* The meta data will store the offset of the user data from
* the beginning of the event. That is after the static arguments
* and the meta data fields.
*/
val = (ptr - (void *)entry) + 4 * uargs;
for (int i = 0; i < uargs; i++) {
if (i)
val += user_sizes[i - 1];
/* Store the offset and the size into the meta data */
*(int *)ptr = val | (user_sizes[i] << 16);
/* Skip the meta data */
ptr += 4;
}
for (int i = 0; i < uargs; i++, buf += SYSCALL_FAULT_ARG_SZ) {
/* Nothing to do if the user space was empty or faulted */
if (!user_sizes[i])
continue;
memcpy(ptr, buf, user_sizes[i]);
ptr += user_sizes[i];
}
}
static void ftrace_syscall_enter(void *data, struct pt_regs *regs, long id)
{
struct trace_array *tr = data;
struct trace_event_file *trace_file;
struct syscall_trace_enter *entry;
struct syscall_metadata *sys_data;
struct trace_event_buffer fbuffer;
unsigned long args[6];
char *user_ptr;
int user_sizes[SYSCALL_FAULT_MAX_CNT] = {};
int syscall_nr;
int size = 0;
int uargs = 0;
bool mayfault;
/*
* Syscall probe called with preemption enabled, but the ring
* buffer and per-cpu data require preemption to be disabled.
*/
might_fault();
syscall_nr = trace_get_syscall_nr(current, regs);
if (syscall_nr < 0 || syscall_nr >= NR_syscalls)
return;
trace_file = READ_ONCE(tr->enter_syscall_files[syscall_nr]);
if (!trace_file)
return;
if (trace_trigger_soft_disabled(trace_file))
return;
sys_data = syscall_nr_to_meta(syscall_nr);
if (!sys_data)
return;
/* Check if this syscall event faults in user space memory */
mayfault = sys_data->user_mask != 0;
guard(preempt_notrace)();
syscall_get_arguments(current, regs, args);
if (mayfault) {
if (syscall_get_data(sys_data, args, &user_ptr,
&size, user_sizes, &uargs, tr->syscall_buf_sz) < 0)
return;
}
size += sizeof(*entry) + sizeof(unsigned long) * sys_data->nb_args;
entry = trace_event_buffer_reserve(&fbuffer, trace_file, size);
if (!entry)
return;
entry = ring_buffer_event_data(fbuffer.event);
entry->nr = syscall_nr;
memcpy(entry->args, args, sizeof(unsigned long) * sys_data->nb_args);
if (mayfault)
syscall_put_data(sys_data, entry, user_ptr, size, user_sizes, uargs);
trace_event_buffer_commit(&fbuffer);
}
static void ftrace_syscall_exit(void *data, struct pt_regs *regs, long ret)
{
struct trace_array *tr = data;
struct trace_event_file *trace_file;
struct syscall_trace_exit *entry;
struct syscall_metadata *sys_data;
struct trace_event_buffer fbuffer;
int syscall_nr;
/*
* Syscall probe called with preemption enabled, but the ring
* buffer and per-cpu data require preemption to be disabled.
*/
might_fault();
guard(preempt_notrace)();
syscall_nr = trace_get_syscall_nr(current, regs);
if (syscall_nr < 0 || syscall_nr >= NR_syscalls)
return;
trace_file = READ_ONCE(tr->exit_syscall_files[syscall_nr]);
if (!trace_file)
return;
if (trace_trigger_soft_disabled(trace_file))
return;
sys_data = syscall_nr_to_meta(syscall_nr);
if (!sys_data)
return;
entry = trace_event_buffer_reserve(&fbuffer, trace_file, sizeof(*entry));
if (!entry)
return;
entry = ring_buffer_event_data(fbuffer.event);
entry->nr = syscall_nr;
entry->ret = syscall_get_return_value(current, regs);
trace_event_buffer_commit(&fbuffer);
}
static int reg_event_syscall_enter(struct trace_event_file *file,
struct trace_event_call *call)
{
struct syscall_metadata *sys_data = call->data;
struct trace_array *tr = file->tr;
int ret = 0;
int num;
num = sys_data->syscall_nr;
if (WARN_ON_ONCE(num < 0 || num >= NR_syscalls))
return -ENOSYS;
guard(mutex)(&syscall_trace_lock);
if (sys_data->user_mask) {
ret = syscall_fault_buffer_enable();
if (ret < 0)
return ret;
}
if (!tr->sys_refcount_enter) {
ret = register_trace_sys_enter(ftrace_syscall_enter, tr);
if (ret < 0) {
if (sys_data->user_mask)
syscall_fault_buffer_disable();
return ret;
}
}
WRITE_ONCE(tr->enter_syscall_files[num], file);
tr->sys_refcount_enter++;
return 0;
}
static void unreg_event_syscall_enter(struct trace_event_file *file,
struct trace_event_call *call)
{
struct syscall_metadata *sys_data = call->data;
struct trace_array *tr = file->tr;
int num;
num = sys_data->syscall_nr;
if (WARN_ON_ONCE(num < 0 || num >= NR_syscalls))
return;
guard(mutex)(&syscall_trace_lock);
tr->sys_refcount_enter--;
WRITE_ONCE(tr->enter_syscall_files[num], NULL);
if (!tr->sys_refcount_enter)
unregister_trace_sys_enter(ftrace_syscall_enter, tr);
if (sys_data->user_mask)
syscall_fault_buffer_disable();
}
static int reg_event_syscall_exit(struct trace_event_file *file,
struct trace_event_call *call)
{
struct trace_array *tr = file->tr;
int ret = 0;
int num;
num = ((struct syscall_metadata *)call->data)->syscall_nr;
if (WARN_ON_ONCE(num < 0 || num >= NR_syscalls))
return -ENOSYS;
mutex_lock(&syscall_trace_lock);
if (!tr->sys_refcount_exit)
ret = register_trace_sys_exit(ftrace_syscall_exit, tr);
if (!ret) {
WRITE_ONCE(tr->exit_syscall_files[num], file);
tr->sys_refcount_exit++;
}
mutex_unlock(&syscall_trace_lock);
return ret;
}
static void unreg_event_syscall_exit(struct trace_event_file *file,
struct trace_event_call *call)
{
struct trace_array *tr = file->tr;
int num;
num = ((struct syscall_metadata *)call->data)->syscall_nr;
if (WARN_ON_ONCE(num < 0 || num >= NR_syscalls))
return;
mutex_lock(&syscall_trace_lock);
tr->sys_refcount_exit--;
WRITE_ONCE(tr->exit_syscall_files[num], NULL);
if (!tr->sys_refcount_exit)
unregister_trace_sys_exit(ftrace_syscall_exit, tr);
mutex_unlock(&syscall_trace_lock);
}
/*
* For system calls that reference user space memory that can
* be recorded into the event, set the system call meta data's user_mask
* to the "args" index that points to the user space memory to retrieve.
*/
static void check_faultable_syscall(struct trace_event_call *call, int nr)
{
struct syscall_metadata *sys_data = call->data;
unsigned long mask;
/* Only work on entry */
if (sys_data->enter_event != call)
return;
sys_data->user_arg_size = -1;
switch (nr) {
/* user arg 1 with size arg at 2 */
case __NR_write:
#ifdef __NR_mq_timedsend
case __NR_mq_timedsend:
#endif
case __NR_pwrite64:
sys_data->user_mask = BIT(1);
sys_data->user_arg_size = 2;
break;
/* user arg 0 with size arg at 1 as string */
case __NR_setdomainname:
case __NR_sethostname:
sys_data->user_mask = BIT(0);
sys_data->user_arg_size = 1;
sys_data->user_arg_is_str = 1;
break;
#ifdef __NR_kexec_file_load
/* user arg 4 with size arg at 3 as string */
case __NR_kexec_file_load:
sys_data->user_mask = BIT(4);
sys_data->user_arg_size = 3;
sys_data->user_arg_is_str = 1;
break;
#endif
/* user arg at position 0 */
#ifdef __NR_access
case __NR_access:
#endif
case __NR_acct:
case __NR_chdir:
#ifdef __NR_chown
case __NR_chown:
#endif
#ifdef __NR_chmod
case __NR_chmod:
#endif
case __NR_chroot:
#ifdef __NR_creat
case __NR_creat:
#endif
case __NR_delete_module:
case __NR_execve:
case __NR_fsopen:
#ifdef __NR_lchown
case __NR_lchown:
#endif
#ifdef __NR_open
case __NR_open:
#endif
case __NR_memfd_create:
#ifdef __NR_mkdir
case __NR_mkdir:
#endif
#ifdef __NR_mknod
case __NR_mknod:
#endif
case __NR_mq_open:
case __NR_mq_unlink:
#ifdef __NR_readlink
case __NR_readlink:
#endif
#ifdef __NR_rmdir
case __NR_rmdir:
#endif
case __NR_shmdt:
#ifdef __NR_statfs
case __NR_statfs:
#endif
case __NR_swapon:
case __NR_swapoff:
#ifdef __NR_truncate
case __NR_truncate:
#endif
#ifdef __NR_unlink
case __NR_unlink:
#endif
case __NR_umount2:
#ifdef __NR_utime
case __NR_utime:
#endif
#ifdef __NR_utimes
case __NR_utimes:
#endif
sys_data->user_mask = BIT(0);
break;
/* user arg at position 1 */
case __NR_execveat:
case __NR_faccessat:
case __NR_faccessat2:
case __NR_finit_module:
case __NR_fchmodat:
case __NR_fchmodat2:
case __NR_fchownat:
case __NR_fgetxattr:
case __NR_flistxattr:
case __NR_fsetxattr:
case __NR_fspick:
case __NR_fremovexattr:
#ifdef __NR_futimesat
case __NR_futimesat:
#endif
case __NR_inotify_add_watch:
case __NR_mkdirat:
case __NR_mknodat:
case __NR_mount_setattr:
case __NR_name_to_handle_at:
#ifdef __NR_newfstatat
case __NR_newfstatat:
#endif
case __NR_openat:
case __NR_openat2:
case __NR_open_tree:
case __NR_open_tree_attr:
case __NR_readlinkat:
case __NR_quotactl:
case __NR_syslog:
case __NR_statx:
case __NR_unlinkat:
#ifdef __NR_utimensat
case __NR_utimensat:
#endif
sys_data->user_mask = BIT(1);
break;
/* user arg at position 2 */
case __NR_init_module:
case __NR_fsconfig:
sys_data->user_mask = BIT(2);
break;
/* user arg at position 4 */
case __NR_fanotify_mark:
sys_data->user_mask = BIT(4);
break;
/* 2 user args, 0 and 1 */
case __NR_add_key:
case __NR_getxattr:
case __NR_lgetxattr:
case __NR_lremovexattr:
#ifdef __NR_link
case __NR_link:
#endif
case __NR_listxattr:
case __NR_llistxattr:
case __NR_lsetxattr:
case __NR_pivot_root:
case __NR_removexattr:
#ifdef __NR_rename
case __NR_rename:
#endif
case __NR_request_key:
case __NR_setxattr:
#ifdef __NR_symlink
case __NR_symlink:
#endif
sys_data->user_mask = BIT(0) | BIT(1);
break;
/* 2 user args, 0 and 2 */
case __NR_symlinkat:
sys_data->user_mask = BIT(0) | BIT(2);
break;
/* 2 user args, 1 and 3 */
case __NR_getxattrat:
case __NR_linkat:
case __NR_listxattrat:
case __NR_move_mount:
#ifdef __NR_renameat
case __NR_renameat:
#endif
case __NR_renameat2:
case __NR_removexattrat:
case __NR_setxattrat:
sys_data->user_mask = BIT(1) | BIT(3);
break;
case __NR_mount: /* Just dev_name and dir_name, TODO add type */
sys_data->user_mask = BIT(0) | BIT(1) | BIT(2);
break;
default:
sys_data->user_mask = 0;
return;
}
if (sys_data->user_arg_size < 0)
return;
/*
* The user_arg_size can only be used when the system call
* is reading only a single address from user space.
*/
mask = sys_data->user_mask;
if (WARN_ON(mask & (mask - 1)))
sys_data->user_arg_size = -1;
}
static int __init init_syscall_trace(struct trace_event_call *call)
{
int id;
int num;
num = ((struct syscall_metadata *)call->data)->syscall_nr;
if (num < 0 || num >= NR_syscalls) {
pr_debug("syscall %s metadata not mapped, disabling ftrace event\n",
((struct syscall_metadata *)call->data)->name);
return -ENOSYS;
}
check_faultable_syscall(call, num);
if (set_syscall_print_fmt(call) < 0)
return -ENOMEM;
id = trace_event_raw_init(call);
if (id < 0) {
free_syscall_print_fmt(call);
return id;
}
return id;
}
static struct trace_event_fields __refdata syscall_enter_fields_array[] = {
SYSCALL_FIELD(int, __syscall_nr),
{ .type = TRACE_FUNCTION_TYPE,
.define_fields = syscall_enter_define_fields },
{}
};
struct trace_event_functions enter_syscall_print_funcs = {
.trace = print_syscall_enter,
};
struct trace_event_functions exit_syscall_print_funcs = {
.trace = print_syscall_exit,
};
struct trace_event_class __refdata event_class_syscall_enter = {
.system = "syscalls",
.reg = syscall_enter_register,
.fields_array = syscall_enter_fields_array,
.get_fields = syscall_get_enter_fields,
.raw_init = init_syscall_trace,
};
struct trace_event_class __refdata event_class_syscall_exit = {
.system = "syscalls",
.reg = syscall_exit_register,
.fields_array = (struct trace_event_fields[]){
SYSCALL_FIELD(int, __syscall_nr),
SYSCALL_FIELD(long, ret),
{}
},
.fields = LIST_HEAD_INIT(event_class_syscall_exit.fields),
.raw_init = init_syscall_trace,
};
unsigned long __init __weak arch_syscall_addr(int nr)
{
return (unsigned long)sys_call_table[nr];
}
void __init init_ftrace_syscalls(void)
{
struct syscall_metadata *meta;
unsigned long addr;
int i;
void *ret;
if (!IS_ENABLED(CONFIG_HAVE_SPARSE_SYSCALL_NR)) {
syscalls_metadata = kcalloc(NR_syscalls,
sizeof(*syscalls_metadata),
GFP_KERNEL);
if (!syscalls_metadata) {
WARN_ON(1);
return;
}
}
for (i = 0; i < NR_syscalls; i++) {
addr = arch_syscall_addr(i);
meta = find_syscall_meta(addr);
if (!meta)
continue;
meta->syscall_nr = i;
if (!IS_ENABLED(CONFIG_HAVE_SPARSE_SYSCALL_NR)) {
syscalls_metadata[i] = meta;
} else {
ret = xa_store(&syscalls_metadata_sparse, i, meta,
GFP_KERNEL);
WARN(xa_is_err(ret),
"Syscall memory allocation failed\n");
}
}
}
#ifdef CONFIG_PERF_EVENTS
static DECLARE_BITMAP(enabled_perf_enter_syscalls, NR_syscalls);
static DECLARE_BITMAP(enabled_perf_exit_syscalls, NR_syscalls);
static int sys_perf_refcount_enter;
static int sys_perf_refcount_exit;
static int perf_call_bpf_enter(struct trace_event_call *call, struct pt_regs *regs,
struct syscall_metadata *sys_data,
struct syscall_trace_enter *rec)
{
struct syscall_tp_t {
struct trace_entry ent;
int syscall_nr;
unsigned long args[SYSCALL_DEFINE_MAXARGS];
} __aligned(8) param;
int i;
BUILD_BUG_ON(sizeof(param.ent) < sizeof(void *));
/* bpf prog requires 'regs' to be the first member in the ctx (a.k.a. ¶m) */
perf_fetch_caller_regs(regs);
*(struct pt_regs **)¶m = regs;
param.syscall_nr = rec->nr;
for (i = 0; i < sys_data->nb_args; i++)
param.args[i] = rec->args[i];
return trace_call_bpf(call, ¶m);
}
static void perf_syscall_enter(void *ignore, struct pt_regs *regs, long id)
{
struct syscall_metadata *sys_data;
struct syscall_trace_enter *rec;
struct pt_regs *fake_regs;
struct hlist_head *head;
unsigned long args[6];
bool valid_prog_array;
bool mayfault;
char *user_ptr;
int user_sizes[SYSCALL_FAULT_MAX_CNT] = {};
int buf_size = CONFIG_TRACE_SYSCALL_BUF_SIZE_DEFAULT;
int syscall_nr;
int rctx;
int size = 0;
int uargs = 0;
/*
* Syscall probe called with preemption enabled, but the ring
* buffer and per-cpu data require preemption to be disabled.
*/
might_fault();
guard(preempt_notrace)();
syscall_nr = trace_get_syscall_nr(current, regs);
if (syscall_nr < 0 || syscall_nr >= NR_syscalls)
return;
if (!test_bit(syscall_nr, enabled_perf_enter_syscalls))
return;
sys_data = syscall_nr_to_meta(syscall_nr);
if (!sys_data)
return;
syscall_get_arguments(current, regs, args);
/* Check if this syscall event faults in user space memory */
mayfault = sys_data->user_mask != 0;
if (mayfault) {
if (syscall_get_data(sys_data, args, &user_ptr,
&size, user_sizes, &uargs, buf_size) < 0)
return;
}
head = this_cpu_ptr(sys_data->enter_event->perf_events);
valid_prog_array = bpf_prog_array_valid(sys_data->enter_event);
if (!valid_prog_array && hlist_empty(head))
return;
/* get the size after alignment with the u32 buffer size field */
size += sizeof(unsigned long) * sys_data->nb_args + sizeof(*rec);
size = ALIGN(size + sizeof(u32), sizeof(u64));
size -= sizeof(u32);
rec = perf_trace_buf_alloc(size, &fake_regs, &rctx);
if (!rec)
return;
rec->nr = syscall_nr;
memcpy(&rec->args, args, sizeof(unsigned long) * sys_data->nb_args);
if (mayfault)
syscall_put_data(sys_data, rec, user_ptr, size, user_sizes, uargs);
if ((valid_prog_array &&
!perf_call_bpf_enter(sys_data->enter_event, fake_regs, sys_data, rec)) ||
hlist_empty(head)) {
perf_swevent_put_recursion_context(rctx);
return;
}
perf_trace_buf_submit(rec, size, rctx,
sys_data->enter_event->event.type, 1, regs,
head, NULL);
}
static int perf_sysenter_enable(struct trace_event_call *call)
{
struct syscall_metadata *sys_data = call->data;
int num;
int ret;
num = sys_data->syscall_nr;
guard(mutex)(&syscall_trace_lock);
if (sys_data->user_mask) {
ret = syscall_fault_buffer_enable();
if (ret < 0)
return ret;
}
if (!sys_perf_refcount_enter) {
ret = register_trace_sys_enter(perf_syscall_enter, NULL);
if (ret) {
pr_info("event trace: Could not activate syscall entry trace point");
if (sys_data->user_mask)
syscall_fault_buffer_disable();
return ret;
}
}
set_bit(num, enabled_perf_enter_syscalls);
sys_perf_refcount_enter++;
return 0;
}
static void perf_sysenter_disable(struct trace_event_call *call)
{
struct syscall_metadata *sys_data = call->data;
int num;
num = sys_data->syscall_nr;
guard(mutex)(&syscall_trace_lock);
sys_perf_refcount_enter--;
clear_bit(num, enabled_perf_enter_syscalls);
if (!sys_perf_refcount_enter)
unregister_trace_sys_enter(perf_syscall_enter, NULL);
if (sys_data->user_mask)
syscall_fault_buffer_disable();
}
static int perf_call_bpf_exit(struct trace_event_call *call, struct pt_regs *regs,
struct syscall_trace_exit *rec)
{
struct syscall_tp_t {
struct trace_entry ent;
int syscall_nr;
unsigned long ret;
} __aligned(8) param;
/* bpf prog requires 'regs' to be the first member in the ctx (a.k.a. ¶m) */
perf_fetch_caller_regs(regs);
*(struct pt_regs **)¶m = regs;
param.syscall_nr = rec->nr;
param.ret = rec->ret;
return trace_call_bpf(call, ¶m);
}
static void perf_syscall_exit(void *ignore, struct pt_regs *regs, long ret)
{
struct syscall_metadata *sys_data;
struct syscall_trace_exit *rec;
struct pt_regs *fake_regs;
struct hlist_head *head;
bool valid_prog_array;
int syscall_nr;
int rctx;
int size;
/*
* Syscall probe called with preemption enabled, but the ring
* buffer and per-cpu data require preemption to be disabled.
*/
might_fault();
guard(preempt_notrace)();
syscall_nr = trace_get_syscall_nr(current, regs);
if (syscall_nr < 0 || syscall_nr >= NR_syscalls)
return;
if (!test_bit(syscall_nr, enabled_perf_exit_syscalls))
return;
sys_data = syscall_nr_to_meta(syscall_nr);
if (!sys_data)
return;
head = this_cpu_ptr(sys_data->exit_event->perf_events);
valid_prog_array = bpf_prog_array_valid(sys_data->exit_event);
if (!valid_prog_array && hlist_empty(head))
return;
/* We can probably do that at build time */
size = ALIGN(sizeof(*rec) + sizeof(u32), sizeof(u64));
size -= sizeof(u32);
rec = perf_trace_buf_alloc(size, &fake_regs, &rctx);
if (!rec)
return;
rec->nr = syscall_nr;
rec->ret = syscall_get_return_value(current, regs);
if ((valid_prog_array &&
!perf_call_bpf_exit(sys_data->exit_event, fake_regs, rec)) ||
hlist_empty(head)) {
perf_swevent_put_recursion_context(rctx);
return;
}
perf_trace_buf_submit(rec, size, rctx, sys_data->exit_event->event.type,
1, regs, head, NULL);
}
static int perf_sysexit_enable(struct trace_event_call *call)
{
int num;
num = ((struct syscall_metadata *)call->data)->syscall_nr;
guard(mutex)(&syscall_trace_lock);
if (!sys_perf_refcount_exit) {
int ret = register_trace_sys_exit(perf_syscall_exit, NULL);
if (ret) {
pr_info("event trace: Could not activate syscall exit trace point");
return ret;
}
}
set_bit(num, enabled_perf_exit_syscalls);
sys_perf_refcount_exit++;
return 0;
}
static void perf_sysexit_disable(struct trace_event_call *call)
{
int num;
num = ((struct syscall_metadata *)call->data)->syscall_nr;
guard(mutex)(&syscall_trace_lock);
sys_perf_refcount_exit--;
clear_bit(num, enabled_perf_exit_syscalls);
if (!sys_perf_refcount_exit)
unregister_trace_sys_exit(perf_syscall_exit, NULL);
}
#endif /* CONFIG_PERF_EVENTS */
static int syscall_enter_register(struct trace_event_call *event,
enum trace_reg type, void *data)
{
struct trace_event_file *file = data;
switch (type) {
case TRACE_REG_REGISTER:
return reg_event_syscall_enter(file, event);
case TRACE_REG_UNREGISTER:
unreg_event_syscall_enter(file, event);
return 0;
#ifdef CONFIG_PERF_EVENTS
case TRACE_REG_PERF_REGISTER:
return perf_sysenter_enable(event);
case TRACE_REG_PERF_UNREGISTER:
perf_sysenter_disable(event);
return 0;
case TRACE_REG_PERF_OPEN:
case TRACE_REG_PERF_CLOSE:
case TRACE_REG_PERF_ADD:
case TRACE_REG_PERF_DEL:
return 0;
#endif
}
return 0;
}
static int syscall_exit_register(struct trace_event_call *event,
enum trace_reg type, void *data)
{
struct trace_event_file *file = data;
switch (type) {
case TRACE_REG_REGISTER:
return reg_event_syscall_exit(file, event);
case TRACE_REG_UNREGISTER:
unreg_event_syscall_exit(file, event);
return 0;
#ifdef CONFIG_PERF_EVENTS
case TRACE_REG_PERF_REGISTER:
return perf_sysexit_enable(event);
case TRACE_REG_PERF_UNREGISTER:
perf_sysexit_disable(event);
return 0;
case TRACE_REG_PERF_OPEN:
case TRACE_REG_PERF_CLOSE:
case TRACE_REG_PERF_ADD:
case TRACE_REG_PERF_DEL:
return 0;
#endif
}
return 0;
}
]
|
| 145/4 |
2026/05/25 22:19 |
llm |
codeexpert |
0m
Model:
Tokens:
input: 206042
output: 875
thoughts: 4331
Thoughts: **Unveiling the Nuances of Tracepoint Access for Unprivileged Users in `perf_event_open`**
Okay, so I've been dissecting the kernel code to understand how `perf_event_open` handles `PERF_TYPE_TRACEPOINT` events for unprivileged users. The initial question is straightforward: *can* they open these events? The answer, it turns out, is a resounding *yes*, but the devil, as always, is in the details. The access isn't a free-for-all; there's a layered set of restrictions in place to prevent security leaks.
First, there's the `perf_check_permission` function. It's the first gatekeeper, ensuring the user has the necessary "ptrace-like" permissions on the target task. If they're tracing their own process (or a process they can `ptrace` into), this check usually succeeds. This is a crucial first step. If the checks here fail, it's game over and event creation is halted.
Then comes the infamous `sysctl_perf_event_paranoid` variable. This is where things get interesting. If `paranoid` is set to 2 or higher (the default!), unprivileged users *must* set `exclude_kernel = 1`. This is a big one. This means they are excluded from anything in the kernel. This `exclude_kernel = 1` then propagates down and is used later. This can make many tracepoints not fire.
The next significant hurdle is whether `PERF_SAMPLE_RAW` is requested. If it is, and the event isn't considered "safe," we hit `perf_allow_tracepoint()`. This is another check that effectively blocks unprivileged users unless `sysctl_perf_event_paranoid` is -1. However, this is bypassed for certain kinds of events.
Now, a critical point comes in. If `exclude_kernel = 1`, and the tracepoint fires in kernel mode (like most kernel tracepoints do), it won't trigger because of `perf_tp_event_match`. However, syscall tracepoints *and* uprobes are special cases. The key is that, for syscalls, the registers passed during the syscall enter are the user registers! `user_mode(regs)` is true. Therefore, the event *does* trigger. This means that, even with `paranoid = 2` and `exclude_kernel = 1`, an unprivileged user can still effectively trace their own syscalls with raw data (assuming `TRACE_EVENT_FL_CAP_ANY` is set for that tracepoint, which it is for syscalls). Uprobes are also safe in this case, as they also have the `TRACE_EVENT_FL_CAP_ANY` flag.
I then went through and double checked the `TRACE_EVENT_FL_CAP_ANY` flags and verified that syscalls and uprobes have it. This allows `PERF_SAMPLE_RAW` data for per-task events, even with `paranoid` at its highest levels.
Furthermore, kernel tracepoints that don't pass that `TRACE_EVENT_FL_CAP_ANY` will never fire because of the `exclude_kernel` option that is needed.
I did some additional analysis about CPU-wide tracing. CPU-wide tracepoints (where the *task* is `NULL`) are off-limits for unprivileged users if `paranoid > 0`. Even with `paranoid = 0`, they are denied `PERF_SAMPLE_RAW`. They *can* be used without that data.
So, the definitive takeaway is: Unprivileged users can open tracepoint events. However, there are a few restrictions:
* They need ptrace-like permissions (checked by `perf_check_permission`).
* With `paranoid >= 2`, they need `exclude_kernel = 1`, which basically means they can only see user-space events.
* They can *only* get `PERF_SAMPLE_RAW` data if `paranoid` is -1 *or* if the event is per-task (`PERF_ATTACH_TASK`) and has `TRACE_EVENT_FL_CAP_ANY`, like syscalls and uprobes.
* If they *don't* request `PERF_SAMPLE_RAW`, they can often open other tracepoints, but, because `exclude_kernel = 1` is needed, most kernel tracepoints will *never* trigger.
It's a carefully balanced system, designed to allow some tracing for unprivileged users, but not to open the floodgates to security vulnerabilities. This also is controlled by the LSM, like SELinux.
|
| 146/2 |
2026/05/25 22:20 |
llm |
expert |
0m
Model:
Tokens:
input: 80091
output: 45
thoughts: 500
Thoughts: **Unveiling a Potential Security Vulnerability**
Okay, so I'm thinking about how unprivileged users can open `PERF_TYPE_TRACEPOINT` events for their own syscalls and uprobes. That immediately got me thinking β if they can *open* these tracepoints, can they *trigger* them? If so, this could be a problem.
Then, there's that obscure `CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS` option. I briefly wondered if `perf_event_open` was involved in writing to the ring buffer with that enabled. But no, that's not right. `perf_event_open` writes to the *perf* ring buffer, not the ftrace one. The ftrace ring buffer is managed through `/sys/kernel/tracing`.
Aha! Now I'm looking at the stack trace from that bug report. It clearly shows `bpf_trace_printk` writing to the ftrace ring buffer:
```
trace_event_buffer_reserve+0x248/0x340 kernel/trace/trace_events.c:672
do_trace_event_raw_event_bpf_trace_printk kernel/trace/bpf_trace.h:11 [inline]
trace_event_raw_event_bpf_trace_printk+0x100/0x260 kernel/trace/bpf_trace.h:11
__do_trace_bpf_trace_printk kernel/trace/bpf_trace.h:11 [inline]
trace_bpf_trace_printk+0x153/0x1b0 kernel/trace/bpf_trace.h:11
____bpf_trace_printk kernel/trace/bpf_trace.c:379 [inline]
bpf_trace_printk+0x11e/0x190 kernel/trace/bpf_trace.c:362
```
Right, `bpf_trace_printk` *always* goes to the ftrace ring buffer. That's the key. Now the critical question is: can an unprivileged user even call `bpf_trace_printk`? We need to figure out if it's available to unprivileged BPF programs.
Okay, let's see. Unprivileged users can load `BPF_PROG_TYPE_SOCKET_FILTER` and `BPF_PROG_TYPE_CGROUP_SKB` programs. Are those the types that can call `bpf_trace_printk`? I need to delve into `bpf_base_func_proto` to confirm. This could be a significant vulnerability!
|
| 147/2 |
2026/05/25 22:20 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/bpf/helpers.c IncludeLines:true Name:bpf_base_func_proto]
Results: map[SourceCode:1976: const struct bpf_func_proto *
1977: bpf_base_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
1978: {
1979: switch (func_id) {
1980: case BPF_FUNC_map_lookup_elem:
1981: return &bpf_map_lookup_elem_proto;
1982: case BPF_FUNC_map_update_elem:
1983: return &bpf_map_update_elem_proto;
1984: case BPF_FUNC_map_delete_elem:
1985: return &bpf_map_delete_elem_proto;
1986: case BPF_FUNC_map_push_elem:
1987: return &bpf_map_push_elem_proto;
1988: case BPF_FUNC_map_pop_elem:
1989: return &bpf_map_pop_elem_proto;
1990: case BPF_FUNC_map_peek_elem:
1991: return &bpf_map_peek_elem_proto;
1992: case BPF_FUNC_map_lookup_percpu_elem:
1993: return &bpf_map_lookup_percpu_elem_proto;
1994: case BPF_FUNC_get_prandom_u32:
1995: return &bpf_get_prandom_u32_proto;
1996: case BPF_FUNC_get_smp_processor_id:
1997: return &bpf_get_raw_smp_processor_id_proto;
1998: case BPF_FUNC_get_numa_node_id:
1999: return &bpf_get_numa_node_id_proto;
2000: case BPF_FUNC_tail_call:
2001: return &bpf_tail_call_proto;
2002: case BPF_FUNC_ktime_get_ns:
2003: return &bpf_ktime_get_ns_proto;
2004: case BPF_FUNC_ktime_get_boot_ns:
2005: return &bpf_ktime_get_boot_ns_proto;
2006: case BPF_FUNC_ktime_get_tai_ns:
2007: return &bpf_ktime_get_tai_ns_proto;
2008: case BPF_FUNC_ringbuf_output:
2009: return &bpf_ringbuf_output_proto;
2010: case BPF_FUNC_ringbuf_reserve:
2011: return &bpf_ringbuf_reserve_proto;
2012: case BPF_FUNC_ringbuf_submit:
2013: return &bpf_ringbuf_submit_proto;
2014: case BPF_FUNC_ringbuf_discard:
2015: return &bpf_ringbuf_discard_proto;
2016: case BPF_FUNC_ringbuf_query:
2017: return &bpf_ringbuf_query_proto;
2018: case BPF_FUNC_strncmp:
2019: return &bpf_strncmp_proto;
2020: case BPF_FUNC_strtol:
2021: return &bpf_strtol_proto;
2022: case BPF_FUNC_strtoul:
2023: return &bpf_strtoul_proto;
2024: case BPF_FUNC_get_current_pid_tgid:
2025: return &bpf_get_current_pid_tgid_proto;
2026: case BPF_FUNC_get_ns_current_pid_tgid:
2027: return &bpf_get_ns_current_pid_tgid_proto;
2028: case BPF_FUNC_get_current_uid_gid:
2029: return &bpf_get_current_uid_gid_proto;
2030: default:
2031: break;
2032: }
2033:
2034: if (!bpf_token_capable(prog->aux->token, CAP_BPF))
2035: return NULL;
2036:
2037: switch (func_id) {
2038: case BPF_FUNC_spin_lock:
2039: return &bpf_spin_lock_proto;
2040: case BPF_FUNC_spin_unlock:
2041: return &bpf_spin_unlock_proto;
2042: case BPF_FUNC_jiffies64:
2043: return &bpf_jiffies64_proto;
2044: case BPF_FUNC_per_cpu_ptr:
2045: return &bpf_per_cpu_ptr_proto;
2046: case BPF_FUNC_this_cpu_ptr:
2047: return &bpf_this_cpu_ptr_proto;
2048: case BPF_FUNC_timer_init:
2049: return &bpf_timer_init_proto;
2050: case BPF_FUNC_timer_set_callback:
2051: return &bpf_timer_set_callback_proto;
2052: case BPF_FUNC_timer_start:
2053: return &bpf_timer_start_proto;
2054: case BPF_FUNC_timer_cancel:
2055: return &bpf_timer_cancel_proto;
2056: case BPF_FUNC_kptr_xchg:
2057: return &bpf_kptr_xchg_proto;
2058: case BPF_FUNC_for_each_map_elem:
2059: return &bpf_for_each_map_elem_proto;
2060: case BPF_FUNC_loop:
2061: return &bpf_loop_proto;
2062: case BPF_FUNC_user_ringbuf_drain:
2063: return &bpf_user_ringbuf_drain_proto;
2064: case BPF_FUNC_ringbuf_reserve_dynptr:
2065: return &bpf_ringbuf_reserve_dynptr_proto;
2066: case BPF_FUNC_ringbuf_submit_dynptr:
2067: return &bpf_ringbuf_submit_dynptr_proto;
2068: case BPF_FUNC_ringbuf_discard_dynptr:
2069: return &bpf_ringbuf_discard_dynptr_proto;
2070: case BPF_FUNC_dynptr_from_mem:
2071: return &bpf_dynptr_from_mem_proto;
2072: case BPF_FUNC_dynptr_read:
2073: return &bpf_dynptr_read_proto;
2074: case BPF_FUNC_dynptr_write:
2075: return &bpf_dynptr_write_proto;
2076: case BPF_FUNC_dynptr_data:
2077: return &bpf_dynptr_data_proto;
2078: #ifdef CONFIG_CGROUPS
2079: case BPF_FUNC_cgrp_storage_get:
2080: return &bpf_cgrp_storage_get_proto;
2081: case BPF_FUNC_cgrp_storage_delete:
2082: return &bpf_cgrp_storage_delete_proto;
2083: case BPF_FUNC_get_current_cgroup_id:
2084: return &bpf_get_current_cgroup_id_proto;
2085: case BPF_FUNC_get_current_ancestor_cgroup_id:
2086: return &bpf_get_current_ancestor_cgroup_id_proto;
2087: case BPF_FUNC_current_task_under_cgroup:
2088: return &bpf_current_task_under_cgroup_proto;
2089: #endif
2090: #ifdef CONFIG_CGROUP_NET_CLASSID
2091: case BPF_FUNC_get_cgroup_classid:
2092: return &bpf_get_cgroup_classid_curr_proto;
2093: #endif
2094: case BPF_FUNC_task_storage_get:
2095: if (bpf_prog_check_recur(prog))
2096: return &bpf_task_storage_get_recur_proto;
2097: return &bpf_task_storage_get_proto;
2098: case BPF_FUNC_task_storage_delete:
2099: if (bpf_prog_check_recur(prog))
2100: return &bpf_task_storage_delete_recur_proto;
2101: return &bpf_task_storage_delete_proto;
2102: default:
2103: break;
2104: }
2105:
2106: if (!bpf_token_capable(prog->aux->token, CAP_PERFMON))
2107: return NULL;
2108:
2109: switch (func_id) {
2110: case BPF_FUNC_trace_printk:
2111: return bpf_get_trace_printk_proto();
2112: case BPF_FUNC_get_current_task:
2113: return &bpf_get_current_task_proto;
2114: case BPF_FUNC_get_current_task_btf:
2115: return &bpf_get_current_task_btf_proto;
2116: case BPF_FUNC_get_current_comm:
2117: return &bpf_get_current_comm_proto;
2118: case BPF_FUNC_probe_read_user:
2119: return &bpf_probe_read_user_proto;
2120: case BPF_FUNC_probe_read_kernel:
2121: return security_locked_down(LOCKDOWN_BPF_READ_KERNEL) < 0 ?
2122: NULL : &bpf_probe_read_kernel_proto;
2123: case BPF_FUNC_probe_read_user_str:
2124: return &bpf_probe_read_user_str_proto;
2125: case BPF_FUNC_probe_read_kernel_str:
2126: return security_locked_down(LOCKDOWN_BPF_READ_KERNEL) < 0 ?
2127: NULL : &bpf_probe_read_kernel_str_proto;
2128: case BPF_FUNC_copy_from_user:
2129: return &bpf_copy_from_user_proto;
2130: case BPF_FUNC_copy_from_user_task:
2131: return &bpf_copy_from_user_task_proto;
2132: case BPF_FUNC_snprintf_btf:
2133: return &bpf_snprintf_btf_proto;
2134: case BPF_FUNC_snprintf:
2135: return &bpf_snprintf_proto;
2136: case BPF_FUNC_task_pt_regs:
2137: return &bpf_task_pt_regs_proto;
2138: case BPF_FUNC_trace_vprintk:
2139: return bpf_get_trace_vprintk_proto();
2140: case BPF_FUNC_perf_event_read_value:
2141: return bpf_get_perf_event_read_value_proto();
2142: case BPF_FUNC_perf_event_read:
2143: return &bpf_perf_event_read_proto;
2144: case BPF_FUNC_send_signal:
2145: return &bpf_send_signal_proto;
2146: case BPF_FUNC_send_signal_thread:
2147: return &bpf_send_signal_thread_proto;
2148: case BPF_FUNC_get_task_stack:
2149: return prog->sleepable ? &bpf_get_task_stack_sleepable_proto
2150: : &bpf_get_task_stack_proto;
2151: case BPF_FUNC_get_branch_snapshot:
2152: return &bpf_get_branch_snapshot_proto;
2153: case BPF_FUNC_find_vma:
2154: return &bpf_find_vma_proto;
2155: default:
2156: return NULL;
2157: }
2158: }
SourceFile:kernel/bpf/helpers.c]
|
| 148/2 |
2026/05/25 22:20 |
llm |
expert |
0m
Model:
Tokens:
input: 84718
output: 43
thoughts: 2004
Thoughts: **Analysis of the Ring Buffer Vulnerability**
Okay, so I've been digging into this potential ring buffer issue. My initial thought was that the observed RCU stall, as described, is due to the use of `bpf_trace_printk`. However, I quickly realized that's a dead end for unprivileged users since `bpf_trace_printk` requires `CAP_PERFMON` which is not available to them.
Then I looked for alternative ways to write to the ftrace ring buffer, but `/sys/kernel/tracing/trace_marker` is off-limits to unprivileged users, as are the syscall tracepoints which also require root to enable. This means that an unprivileged user cannot trigger this particular bug.
The reported RCU stall is directly linked to the configuration `CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS`. With this enabled, the `check_buffer` function gets called repeatedly, though the calculations reveal that the iterations per page are very quick. So it cannot cause the RCU stall that was observed. The RCU stall seems only triggered by attaching a BPF program 65 times to a timer interrupt tracepoint, then doing `bpf_trace_printk` 65 times per interrupt. This also requires `CAP_PERFMON` at every turn. Thus it is entirely inaccessible to unprivileged processes.
Essentially, this is a Denial of Service, *only* for root/`CAP_PERFMON` users, and it's caused by a debug option that isn't typically enabled in production. It isn't an exploitable security boundary bypass.
But, I need to check for other vulnerabilities. I started to wonder if there could be an infinite loop in `check_buffer`. Specifically, I re-examined the code in `rb_read_data_buffer`, focusing on how the ring buffer events are parsed. Here's what I found:
In `rb_read_data_buffer`, the `rb_event_length` function has a major bug when encountering padding. The function `rb_event_length` returns `-1` if the current event is `RINGBUF_TYPE_PADDING` and `rb_null_event` is true. `rb_null_event` returns true if the event has `RINGBUF_TYPE_PADDING` and its `time_delta` is zero. When `rb_event_length` returns `-1`, this is interpreted as `0xffffffff` (an unsigned value) due to the integer type of variable `e`. This means `e` will be decreased by one during the loop, creating a loop that reads out of bounds.
Now, further analysis shows that the function `rb_read_data_buffer` does not correctly handle padding events. It doesn't check for padding when it encounters it and incorrectly processes it. When the `time_delta` of a padding event is 0, the loop falls through to `RINGBUF_TYPE_DATA`. Because of `rb_null_event`, it reads out of bounds. The crucial factor is the `tail` variable, which determines how much data to read from the page. If `tail` encompasses the padding, then the loop will process padding, leading to this vulnerability.
Therefore, an attacker may be able to cause an infinite loop and crash the kernel using a specially crafted ring buffer event sequence by writing out of bounds!
|
| 149/2 |
2026/05/25 22:21 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/ring_buffer.c IncludeLines:true Name:rb_move_tail]
Results: map[SourceCode:3563: /*
3564: * This is the slow path, force gcc not to inline it.
3565: */
3566: static noinline struct ring_buffer_event *
3567: rb_move_tail(struct ring_buffer_per_cpu *cpu_buffer,
3568: unsigned long tail, struct rb_event_info *info)
3569: {
3570: struct buffer_page *tail_page = info->tail_page;
3571: struct buffer_page *commit_page = cpu_buffer->commit_page;
3572: struct trace_buffer *buffer = cpu_buffer->buffer;
3573: struct buffer_page *next_page;
3574: int ret;
3575:
3576: next_page = tail_page;
3577:
3578: rb_inc_page(&next_page);
3579:
3580: /*
3581: * If for some reason, we had an interrupt storm that made
3582: * it all the way around the buffer, bail, and warn
3583: * about it.
3584: */
3585: if (unlikely(next_page == commit_page)) {
3586: local_inc(&cpu_buffer->commit_overrun);
3587: goto out_reset;
3588: }
3589:
3590: /*
3591: * This is where the fun begins!
3592: *
3593: * We are fighting against races between a reader that
3594: * could be on another CPU trying to swap its reader
3595: * page with the buffer head.
3596: *
3597: * We are also fighting against interrupts coming in and
3598: * moving the head or tail on us as well.
3599: *
3600: * If the next page is the head page then we have filled
3601: * the buffer, unless the commit page is still on the
3602: * reader page.
3603: */
3604: if (rb_is_head_page(next_page, &tail_page->list)) {
3605:
3606: /*
3607: * If the commit is not on the reader page, then
3608: * move the header page.
3609: */
3610: if (!rb_is_reader_page(cpu_buffer->commit_page)) {
3611: /*
3612: * If we are not in overwrite mode,
3613: * this is easy, just stop here.
3614: */
3615: if (!(buffer->flags & RB_FL_OVERWRITE)) {
3616: local_inc(&cpu_buffer->dropped_events);
3617: goto out_reset;
3618: }
3619:
3620: ret = rb_handle_head_page(cpu_buffer,
3621: tail_page,
3622: next_page);
3623: if (ret < 0)
3624: goto out_reset;
3625: if (ret)
3626: goto out_again;
3627: } else {
3628: /*
3629: * We need to be careful here too. The
3630: * commit page could still be on the reader
3631: * page. We could have a small buffer, and
3632: * have filled up the buffer with events
3633: * from interrupts and such, and wrapped.
3634: *
3635: * Note, if the tail page is also on the
3636: * reader_page, we let it move out.
3637: */
3638: if (unlikely((cpu_buffer->commit_page !=
3639: cpu_buffer->tail_page) &&
3640: (cpu_buffer->commit_page ==
3641: cpu_buffer->reader_page))) {
3642: local_inc(&cpu_buffer->commit_overrun);
3643: goto out_reset;
3644: }
3645: }
3646: }
3647:
3648: rb_tail_page_update(cpu_buffer, tail_page, next_page);
3649:
3650: out_again:
3651:
3652: rb_reset_tail(cpu_buffer, tail, info);
3653:
3654: /* Commit what we have for now. */
3655: rb_end_commit(cpu_buffer);
3656: /* rb_end_commit() decs committing */
3657: local_inc(&cpu_buffer->committing);
3658:
3659: /* fail and let the caller try again */
3660: return ERR_PTR(-EAGAIN);
3661:
3662: out_reset:
3663: /* reset write */
3664: rb_reset_tail(cpu_buffer, tail, info);
3665:
3666: return NULL;
3667: }
SourceFile:kernel/trace/ring_buffer.c]
|
| 150/2 |
2026/05/25 22:21 |
llm |
expert |
0m
Model:
Tokens:
input: 88960
output: 43
thoughts: 15
Thoughts: Okay, here's my interpretation of that thought process, assuming I'm an expert in this area:
**Digging into a Potential Write Conflict**
Hmm, hold on a second. When I think about `rb_reset_tail`, a red flag is going up in my head. I'm suddenly considering if it might be responsible for actually writing that padding event we were concerned about, or, at least, if it might indirectly trigger it. I need to re-examine the code in that function and its calling context. My gut's telling me that this particular function, which resets the tail of the ring buffer, might be the source of the issue. I need to be sure that I am understanding correctly, since the impact on overall performance could be significant if I am missing something. I must be sure about this.
|
| 151/2 |
2026/05/25 22:21 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/ring_buffer.c IncludeLines:true Name:rb_reset_tail]
Results: map[SourceCode:3483: static inline void
3484: rb_reset_tail(struct ring_buffer_per_cpu *cpu_buffer,
3485: unsigned long tail, struct rb_event_info *info)
3486: {
3487: unsigned long bsize = READ_ONCE(cpu_buffer->buffer->subbuf_size);
3488: struct buffer_page *tail_page = info->tail_page;
3489: struct ring_buffer_event *event;
3490: unsigned long length = info->length;
3491:
3492: /*
3493: * Only the event that crossed the page boundary
3494: * must fill the old tail_page with padding.
3495: */
3496: if (tail >= bsize) {
3497: /*
3498: * If the page was filled, then we still need
3499: * to update the real_end. Reset it to zero
3500: * and the reader will ignore it.
3501: */
3502: if (tail == bsize)
3503: tail_page->real_end = 0;
3504:
3505: local_sub(length, &tail_page->write);
3506: return;
3507: }
3508:
3509: event = __rb_page_index(tail_page, tail);
3510:
3511: /*
3512: * Save the original length to the meta data.
3513: * This will be used by the reader to add lost event
3514: * counter.
3515: */
3516: tail_page->real_end = tail;
3517:
3518: /*
3519: * If this event is bigger than the minimum size, then
3520: * we need to be careful that we don't subtract the
3521: * write counter enough to allow another writer to slip
3522: * in on this page.
3523: * We put in a discarded commit instead, to make sure
3524: * that this space is not used again, and this space will
3525: * not be accounted into 'entries_bytes'.
3526: *
3527: * If we are less than the minimum size, we don't need to
3528: * worry about it.
3529: */
3530: if (tail > (bsize - RB_EVNT_MIN_SIZE)) {
3531: /* No room for any events */
3532:
3533: /* Mark the rest of the page with padding */
3534: rb_event_set_padding(event);
3535:
3536: /* Make sure the padding is visible before the write update */
3537: smp_wmb();
3538:
3539: /* Set the write back to the previous setting */
3540: local_sub(length, &tail_page->write);
3541: return;
3542: }
3543:
3544: /* Put in a discarded event */
3545: event->array[0] = (bsize - tail) - RB_EVNT_HDR_SIZE;
3546: event->type_len = RINGBUF_TYPE_PADDING;
3547: /* time delta must be non zero */
3548: event->time_delta = 1;
3549:
3550: /* account for padding bytes */
3551: local_add(bsize - tail, &cpu_buffer->entries_bytes);
3552:
3553: /* Make sure the padding is visible before the tail_page->write update */
3554: smp_wmb();
3555:
3556: /* Set write to end of buffer */
3557: length = (tail + length) - bsize;
3558: local_sub(length, &tail_page->write);
3559: }
SourceFile:kernel/trace/ring_buffer.c]
|
| 152/2 |
2026/05/25 22:21 |
llm |
expert |
0m
Model:
Tokens:
input: 90328
output: 45
thoughts: 17
Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience, and presented as if it were my own internal monologue:
**Padding Adjustment in Event Handling**
Right, that's it! *rb_event_set_padding(event)*β¦ that's the key function. It's the one that directly handles setting the padding for the event. Simple, elegant, and directly addresses the padding requirement I was considering. That means all of the preceding code relating to calculations and potential alternatives is no longer required. It confirms my assumption about direct padding control. Now I can move forward with that implementation and be confident in its behaviour. Perfect.
|
| 153/2 |
2026/05/25 22:21 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/trace/ring_buffer.c IncludeLines:true Name:rb_event_set_padding]
Results: map[SourceCode: 195: static void rb_event_set_padding(struct ring_buffer_event *event)
196: {
197: /* padding has a NULL time_delta */
198: event->type_len = RINGBUF_TYPE_PADDING;
199: event->time_delta = 0;
200: }
SourceFile:kernel/trace/ring_buffer.c]
|
| 154/2 |
2026/05/25 22:21 |
llm |
expert |
0m
Model:
Tokens:
input: 90654
output: 66
thoughts: 3018
Thoughts: **Analysis of the Ring Buffer RCU Stall Bug**
Okay, so I've been dissecting this ring buffer issue, and here's my take: the key lies in how `rb_event_set_padding` sets up the `rb_null_event`. My initial thoughts were, if `tail` is near the end, we get this null event, and `check_buffer` could get tripped up on it during iteration. That's a red herring, though. The `check_buffer` logic is actually sound, it won't process the padding event if it is at `tail` or `commit`, so it won't loop infinitely. And, the potential for a negative return from `rb_read_data_buffer` triggering a timestamp check and a warning doesn't seem to be the core problem.
The real culprit seems to be the O(N^2) behavior in `check_buffer`, amplified when a *lot* of BPF programs are attached to a high-frequency tracepoint. The code iterates the ring buffer up to `tail` and potentially again to commit, but this is multiplied when there's a lot of ring buffer data to process. So a slowdown is the issue. And it's only triggered when a CPU interrupt is very frequent, which is when the performance hit from O(N^2) becomes a problem.
Here's the kicker: this requires `CAP_PERFMON` to even load BPF programs and attach them to tracepoints. Also, BPF programs require `CAP_SYS_ADMIN` to configure delegation for `bpffs`, preventing them from being accessible in user namespaces. And even if `unprivileged_bpf_disabled=0`, `BPF_PROG_TYPE_RAW_TRACEPOINT` still requires `CAP_BPF` in the initial user namespace, and unprivileged users cannot trigger them.
Can an unprivileged user trigger syscalls? Well, tracepoints can be triggered, but if `CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS` is enabled, then unprivileged users could also trigger `check_buffer` many times, because of syscall tracepoints. But the overhead is not too great, and the system would only slow down.
So, in essence, the bug is dependent on root attaching *many* BPF programs (65 is what's mentioned) to a timer interrupt or tracepoint. The BPF programs take too long to run, causing the timer interrupt frequency to be overrun. Or, it requires root enabling a debug option that isn't really intended for production (but even that requires root attaching the programs). This then leads to a livelock situation, since the timer interrupt will never finish. The fix has to be that, when processing a hardware interrupt, the time to execute the code attached to the interrupt must be less than the interval between interrupts.
The security impact boils down to this: It's not exploitable by an unprivileged user, and the denial-of-service condition only applies to root.
After double-checking, there's no way to reach `check_buffer` from an unprivileged context. So, I will mark the dimensions of the bug as: Exploitable: False, DenialOfService: True, Unprivileged: False, UserNamespace: False, VMGuestTrigger: False, VMHostTrigger: False, NetworkTrigger: False, RemoteTrigger: False, PeripheralTrigger: False, FilesystemTrigger: False.
|
| 155/2 |
2026/05/25 22:21 |
tool |
set-results |
0m
Args: map[DenialOfService:true Exploitable:false FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
Results: map[DenialOfService:true Exploitable:false FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
|
| 156/2 |
2026/05/25 22:21 |
llm |
expert |
0m
Model:
Tokens:
input: 94532
output: 757
thoughts: 486
Thoughts: **Final Thoughts on the RCU Stall Bug**
Okay, the analysis is complete, and I've reached my final conclusion. This is an RCU stall, plain and simple, a denial-of-service vulnerability triggered by a BPF program. Specifically, we're looking at a program attached to the `rcu_utilization` tracepoint, which, as we know, is called from the timer interrupt, a recipe for high frequency.
The reproducer cleverly stacks the deck: It attaches the BPF program a staggering 65 times to that single tracepoint. Then, *each* invocation of the tracepoint also executes that BPF program 65 times. The BPF program itself is using `bpf_trace_printk`, which, of course, is writing to the ftrace ring buffer.
Hereβs the kicker: The kernelβs got `CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS` enabled. That debug option, not meant for prime time, is the true culprit. With this enabled, writing to the ring buffer kicks off `check_buffer()`, which has that nasty O(N^2) time complexity per buffer page.
The combination is lethal: 65 programs triggered on timer interrupts, all amplified by the O(N^2) complexity of `check_buffer()`. This causes the timer interrupt handler to drag on longer than the timer interval. And you know what that means β interrupt storm, livelock, and eventually, the RCU stall.
From a security perspective, this is important: Loading this type of BPF program (`BPF_PROG_TYPE_RAW_TRACEPOINT`) requires `CAP_PERFMON` or `CAP_SYS_ADMIN` in the initial user namespace. Using `bpf_trace_printk` also requires `CAP_PERFMON`. Attaching a BPF program to a raw tracepoint means it needs to be a tracing type program, which unprivileged users can't load anyway, even with `unprivileged_bpf_disabled=0`.
And no clever use of BPF tokens can bypass this in a user namespace. Creating a token with delegation options requires mounting `bpffs` with those options, which, of course, *also* requires `CAP_SYS_ADMIN` in the initial user namespace.
The takeaway? This isnβt a security boundary bypass. It's a known consequence of a highly privileged user (root, or someone with `CAP_PERFMON` / `CAP_SYS_ADMIN`) attaching slow BPF programs to high-frequency interrupts, which is made worse by a debug kernel config. It's an important reminder of the power and responsibility that comes with root privileges.
|
| 157/1 |
2026/05/25 22:21 |
action |
format-explanation |
0m
Results: map[Explanation:Based on the analysis of the bug report and the kernel source code, here is the
assessment of the security impact:
**Bug Mechanism:**
The bug is an RCU stall (Denial of Service) caused by a livelock in the timer
interrupt handler. The reproducer achieves this by loading a BPF program of type
`BPF_PROG_TYPE_RAW_TRACEPOINT` and attaching it **65 times** to the
`rcu_utilization` tracepoint. This tracepoint is invoked from
`rcu_sched_clock_irq`, which runs in the context of the high-frequency timer
interrupt (`sysvec_apic_timer_interrupt`).
Each time the timer interrupt fires, the BPF program is executed 65 times. The
program calls `bpf_trace_printk`, which formats a string and writes it to the
ftrace ring buffer. If the kernel is compiled with the debug option
`CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS` (which the syzbot kernel has enabled),
every write to the ring buffer triggers `check_buffer()`. This debugging
function iterates over all events in the current buffer page, resulting in an
$O(N^2)$ time complexity per page.
The combination of 65 `bpf_trace_printk` calls per interrupt and the $O(N^2)$
validation overhead causes the timer interrupt handler to take longer to execute
than the actual interval between timer interrupts (e.g., >1ms). This completely
starves the CPU, leading to an interrupt storm and the reported RCU stall.
**Security Impact:**
* **Exploitable:** No. The bug does not cause memory corruption, out-of-bounds
accesses, or use-after-free conditions. It is purely a CPU starvation issue
(Denial of Service).
* **Accessible From Unprivileged Processes:** No. To trigger this bug, the user
must:
1. Load a BPF program of type `BPF_PROG_TYPE_RAW_TRACEPOINT`, which strictly
requires `CAP_PERFMON` or `CAP_SYS_ADMIN`. (Even if
`unprivileged_bpf_disabled=0`, unprivileged users are restricted to socket
filters and cgroup skb programs).
2. Use the `bpf_trace_printk` helper, which also explicitly requires
`CAP_PERFMON`.
3. Attach the program to a raw tracepoint, which rejects unprivileged program
types.
* **Accessible From User Namespaces:** No. While BPF tokens were introduced to
allow delegating BPF permissions to user namespaces, creating a BPF token
requires mounting `bpffs` with delegation options (e.g., `delegate_progs`). The
kernel explicitly checks for `capable(CAP_SYS_ADMIN)` in the **initial** user
namespace when setting these mount options. Therefore, an unprivileged user
cannot create a user namespace and grant themselves the ability to load tracing
BPF programs.
**Conclusion:**
This is not a security vulnerability. It is a Denial of Service that can only be
triggered by a highly privileged user (root or a user with
`CAP_PERFMON`/`CAP_SYS_ADMIN` in the initial namespace). Privileged users
already have numerous ways to halt the system (e.g., loading an infinite loop
BPF program, or writing to `/proc/sysrq-trigger`). Furthermore, the extreme
overhead is heavily dependent on a debug configuration
(`CONFIG_RING_BUFFER_VALIDATE_TIME_DELTAS`) which is explicitly documented as
"Do not use it on production systems."]
|