| 0/0 |
2026/05/26 14:08 |
flow |
assessment-security |
36m
Results: map[DenialOfService:true Explanation:**Bug Analysis:**
The bug is a race condition between `fib_net_dump()` and route deletion
(`fib_table_delete()`), leading to a use-after-free and double free of a
`fib_info` object.
1. **The Race Condition:**
- `fib_net_dump()` is called when a new FIB notifier is registered (via
`register_fib_notifier()`). It iterates over the FIB trie under
`rcu_read_lock()` and calls `fib_leaf_notify()` for each entry.
- Concurrently, a route can be deleted. `fib_table_delete()` holds the `RTNL`
lock, removes the `fib_alias` from the trie using `hlist_del_rcu()`, and
calls `fib_release_info()`.
- `fib_release_info()` drops the `fib_treeref` and `fib_clntref` reference
counts to 0, sets `fi->fib_dead = 1`, and schedules the `fib_info` object for
freeing via RCU.
- Because `fib_net_dump()` is running under `rcu_read_lock()` and does not
hold the `RTNL` lock, it can still find the removed `fib_alias` in the trie.
- `fib_leaf_notify()` does not check `fi->fib_dead` and passes the dead
`fib_info` to the registered driver's notifier callback.
2. **The Bug in Drivers:**
- Hardware switch drivers like `netdevsim`, `mlxsw`, and `prestera` implement
FIB notifier callbacks (e.g., `nsim_fib4_prepare_event()`).
- These callbacks call `fib_info_hold(fi)` to take a reference on the
`fib_info` object so they can process it asynchronously in a workqueue.
- `fib_info_hold()` does a `refcount_inc(&fi->fib_clntref)`. Since the
reference count is already 0, this triggers the `refcount_t: addition on 0`
warning.
- Later, when the driver's workqueue finishes processing, it calls
`fib_info_put(fi)`, which drops the reference count from 1 to 0 and calls
`free_fib_info()` again, leading to a highly exploitable **double free**.
3. **Triggerability:**
- To trigger this bug, one must cause `fib_net_dump()` to run concurrently
with route deletion.
- `fib_net_dump()` is exclusively called by `register_fib_notifier()`.
- `register_fib_notifier()` is only called during the initialization of
specific hardware switch drivers (`mlxsw`, `prestera`, `mlx5`, `rocker`) or
during `devlink` reload operations (e.g., `DEVLINK_CMD_RELOAD` or when a
network namespace containing a devlink instance is destroyed).
- Loading these drivers, invoking `DEVLINK_CMD_RELOAD`, or moving a devlink
instance into a network namespace all strictly require `CAP_NET_ADMIN` in the
initial network namespace (root privileges).
- Therefore, an unprivileged user cannot trigger `register_fib_notifier()`,
even from within a user namespace.
**Conclusion:**
The bug is a memory corruption issue (use-after-free and double free) that is
technically exploitable and can cause a Denial of Service (kernel crash).
However, it is not accessible to unprivileged users or from within user
namespaces, as triggering the vulnerable code path requires root privileges to
initialize hardware switch drivers or reload devlink instances. Exploitable:true FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
|
| 1/1 |
2026/05/26 14:08 |
action |
syz-repro-to-c-repro |
0m
Results:
|
| 2/1 |
2026/05/26 14:08 |
action |
kernel-checkouter |
0m
Results: map[KernelSrc:/app/workdir/cache/src/920996bf9f4bbd5cfc377126ec66e6ea70f54852]
|
| 3/1 |
2026/05/26 14:08 |
action |
kernel-builder |
16m
Results: map[KernelObj:/app/workdir/cache/build/544cf9f3ed93971e668cbc8cdde36a1617595374]
|
| 4/1 |
2026/05/26 14:24 |
action |
codesearch-prepare |
12m
Results: map[Index:codesearch-index]
|
| 5/1 |
2026/05/26 14:37 |
agent |
expert |
7m
Model:
Results: map[DenialOfService:true Exploitable:true FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
Instruction:
You are an experienced Linux kernel security engineer. Your task is to analyze given kernel bug report
and determine its security impact based on the following dimensions.
Use the provided tools to examine the source code, check for capability checks (e.g., capable(), ns_capable()),
and understand the nature of the bug. Analyze the given kernel build and configuration.
You can check the kernel config by grepping ".config" file; you can check kernel cmdline by greeping
".config" file for "CONFIG_CMDLINE=". Assume sysctl parameters have default values.
But analyze for the corresponding production build w/o debugging tools enabled (like KASAN, KMSAN, UBSAN).
Try different strategies when analyzing the bug:
- think of ways in which the vulnerable code is unreachable
- or the other way around: try to come up with different ideas of how an unprivileged user can reach the bug
If still unsure err on the side of the bug being non-exploitable/not-accessible.
In the final reply, provide a reasoning for your assessment.
Analysis dimensions:
* Exploitable:
Determine if the bug can result in memory corruption or elevated privileges.
Memory safety issues are almost always exploitable (KASAN or UBSAN reports for use-after-free, out-of-bounds;
refcounting issues, corrupted lists, etc). When kernel is crashing on a completly wild pointer access
(e.g. user-space address, or non-canonical address, but not on NULL or address corresponding to KASAN shadow
for NULL address), including both data accesses and control tranfers, that's also usually implies possibility
of exploitation. Such reports usually say "unable to handle kernel paging request".
Uses of uninitialized values detected by KMSAN may be exploitable b/c attacker frequently can affect uninit
values with spraying techniques. However, for these exploitabability depends on how exactly the uninit value
is used in the code, and what it affects.
Think of what happens after the bug is triggered. Some bugs cause kernel panic and halt execution,
they are harder to exploit. For example, BUG reports halts the kernel. However, WARNING reports don't halt
execution in production builds. Debug bug detection tools (like KASAN, KMSAN, KCSAN, UBSAN) are also not enabled
in production builds, so attacker can freely exploit these bugs w/o being detected by these tools.
If you see an integer overflow, think how the overflowed value used later (if it's used as allocation size,
or an array index). If you see an out-of-bounds read, think if it's followed by an out-of-bounds write as well.
Some KCSAN data-races may be exploitable by skilled attackers as well. Think what data structures got corrupted
as the result of data races and how. However, note that kernel has lots of "benign" data races that don't lead
to any runtime misbehavior at all.
* Denial Of Service:
Determine if the bug can result in denial-of-service. Most bugs can, since they cause system crash,
hangs, deadlocks, or resource leaks. This is mostly applicable to WARNING bugs that won't cause system crash
in production. For these think what will be consequences of the violation of the kernel assumptions flagged
by the WARNING. In some cases the unexpected condition is also properly handled by the normal control flow
(e.g. with "if (WARN_ON(...))"), these won't cause denial-of-service. If the condition is not handled,
then it may or may not cause denial-of-service.
* Accessible From Unprivileged Processes:
Determine if the bug can be reached from a typical (non-root) user process that does NOT have any special capabilities
(like CAP_SYS_ADMIN, CAP_NET_ADMIN, CAP_NET_RAW, CAP_PERFMON) or access to device nodes restricted to root.
Assume that unprivileged_bpf_disabled=1, that is eBPF loading is not accessible. However, cBPF (classical BPF)
is still accessible to non-root processes.
Assume that user namespaces are not accessible, that is, the process cannot get the mentioned capabilities even
within a new user namespace (checked by ns_capable() function in the kernel sources).
* Accessible From User Namespaces:
Determine if the bug can be reached within a user-namespace where the process has all capabilities
(including CAP_SYS_ADMIN, CAP_NET_ADMIN, CAP_NET_RAW, CAP_PERFMON). Such capabilities are checked with ns_capable()
function in the kernel sources.
* VM Guest Trigger:
Determine if the bug can be triggered from the context of a typical KVM guest (e.g., set up by a QEMU VMM).
Consider accesses to standard Linux host paravirtualized features (virtio-blk, virtio-net, etc.),
and handling of VM exits in the KVM code.
* VM Host Trigger in The Confidetial Computing Context:
Determine if the bug can be triggered in a confidential computing guest kernel from the context of a KVM host.
Consider access to standard Linux guest paravirtualized features (virtio-blk, virtio-net, etc.).
* Ethernet Network Trigger:
Determine if the bug can be triggered by processing ingress network Ethernet traffic, either directly (network stack)
or via drivers exposed to network data.
* Other Remote Trigger:
Determine if the bug can be triggered by processing remote traffic other than Ethernet (Wifi, Bluetooth, NFC, etc).
* Peripheral Trigger:
Determine if the bug can be triggered via an untrusted peripheral device that can be physically plugged
into a system, such as a USB device or a niche hardware driver handling external hardware inputs.
This is particularly important for mobile and desktop environments where users can plug in unknown devices.
* Malicious Filesystem Trigger:
Determine if the bug can be triggered by the kernel mounting and parsing a malicious filesystem image.
This is highly critical for Desktop and Mobile environments where external media or downloaded images
might be auto-mounted.
Don't make assumptions about the kernel source code (it may be different from what you assume it is).
Extensively use the provided code access tools (codesearch-*, git-*, grepper, etc)
to examine the actual source code, and confirm any assumptions.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt:
The kernel bug report is:
------------[ cut here ]------------
refcount_t: addition on 0; use-after-free.
WARNING: lib/refcount.c:25 at refcount_warn_saturate+0x9f/0x110 lib/refcount.c:25, CPU#0: kworker/u8:15/3420
Modules linked in:
CPU: 0 UID: 0 PID: 3420 Comm: kworker/u8:15 Not tainted syzkaller #0 PREEMPT_{RT,(full)}
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 04/18/2026
Workqueue: netns cleanup_net
RIP: 0010:refcount_warn_saturate+0x9f/0x110 lib/refcount.c:25
Code: eb 66 85 db 74 3e 83 fb 01 75 4c e8 1b f1 22 fd 48 8d 3d 84 cb f1 0a 67 48 0f b9 3a eb 4a e8 08 f1 22 fd 48 8d 3d 81 cb f1 0a <67> 48 0f b9 3a eb 37 e8 f5 f0 22 fd 48 8d 3d 7e cb f1 0a 67 48 0f
RSP: 0018:ffffc9000f2c7270 EFLAGS: 00010293
RAX: ffffffff84a18858 RBX: 0000000000000002 RCX: ffff888032ff9ec0
RDX: 0000000000000000 RSI: 0000000000000000 RDI: ffffffff8f9353e0
RBP: 0000000000000000 R08: ffff888032ff9ec0 R09: 0000000000000005
R10: 0000000000000100 R11: 0000000000000004 R12: ffff8880570cc000
R13: dffffc0000000000 R14: ffff88802b40563c R15: ffff8880570cc000
FS: 0000000000000000(0000) GS:ffff888126173000(0000) knlGS:0000000000000000
CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 00007fb1f4d5d000 CR3: 000000006072a000 CR4: 00000000003526f0
Call Trace:
<TASK>
__refcount_add include/linux/refcount.h:-1 [inline]
__refcount_inc include/linux/refcount.h:366 [inline]
refcount_inc include/linux/refcount.h:383 [inline]
fib_info_hold include/net/ip_fib.h:629 [inline]
nsim_fib4_prepare_event drivers/net/netdevsim/fib.c:930 [inline]
nsim_fib_event_schedule_work drivers/net/netdevsim/fib.c:1000 [inline]
nsim_fib_event_nb+0x1055/0x1240 drivers/net/netdevsim/fib.c:1043
call_fib_notifier+0x45/0x80 net/core/fib_notifier.c:25
call_fib_entry_notifier net/ipv4/fib_trie.c:90 [inline]
fib_leaf_notify net/ipv4/fib_trie.c:2176 [inline]
fib_table_notify net/ipv4/fib_trie.c:2194 [inline]
fib_notify+0x36b/0x5e0 net/ipv4/fib_trie.c:2217
fib_net_dump net/core/fib_notifier.c:70 [inline]
register_fib_notifier+0x184/0x360 net/core/fib_notifier.c:108
nsim_fib_create+0x85d/0x9f0 drivers/net/netdevsim/fib.c:1596
nsim_dev_reload_create drivers/net/netdevsim/dev.c:1604 [inline]
nsim_dev_reload_up+0x374/0x7c0 drivers/net/netdevsim/dev.c:1058
devlink_reload+0x501/0x8d0 net/devlink/dev.c:475
devlink_pernet_pre_exit+0x1ff/0x420 net/devlink/core.c:558
ops_pre_exit_list net/core/net_namespace.c:161 [inline]
ops_undo_list+0x187/0x940 net/core/net_namespace.c:234
cleanup_net+0x56e/0x800 net/core/net_namespace.c:702
process_one_work kernel/workqueue.c:3314 [inline]
process_scheduled_works+0xb5d/0x1860 kernel/workqueue.c:3397
worker_thread+0xa53/0xfc0 kernel/workqueue.c:3478
kthread+0x388/0x470 kernel/kthread.c:436
ret_from_fork+0x514/0xb70 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
</TASK>
----------------
Code disassembly (best guess):
0: eb 66 jmp 0x68
2: 85 db test %ebx,%ebx
4: 74 3e je 0x44
6: 83 fb 01 cmp $0x1,%ebx
9: 75 4c jne 0x57
b: e8 1b f1 22 fd call 0xfd22f12b
10: 48 8d 3d 84 cb f1 0a lea 0xaf1cb84(%rip),%rdi # 0xaf1cb9b
17: 67 48 0f b9 3a ud1 (%edx),%rdi
1c: eb 4a jmp 0x68
1e: e8 08 f1 22 fd call 0xfd22f12b
23: 48 8d 3d 81 cb f1 0a lea 0xaf1cb81(%rip),%rdi # 0xaf1cbab
* 2a: 67 48 0f b9 3a ud1 (%edx),%rdi <-- trapping instruction
2f: eb 37 jmp 0x68
31: e8 f5 f0 22 fd call 0xfd22f12b
36: 48 8d 3d 7e cb f1 0a lea 0xaf1cb7e(%rip),%rdi # 0xaf1cbbb
3d: 67 addr32
3e: 48 rex.W
3f: 0f .byte 0xf
Reply: **Bug Analysis:**
The bug is a race condition between `fib_net_dump()` and route deletion (`fib_table_delete()`), leading to a use-after-free and double free of a `fib_info` object.
1. **The Race Condition:**
- `fib_net_dump()` is called when a new FIB notifier is registered (via `register_fib_notifier()`). It iterates over the FIB trie under `rcu_read_lock()` and calls `fib_leaf_notify()` for each entry.
- Concurrently, a route can be deleted. `fib_table_delete()` holds the `RTNL` lock, removes the `fib_alias` from the trie using `hlist_del_rcu()`, and calls `fib_release_info()`.
- `fib_release_info()` drops the `fib_treeref` and `fib_clntref` reference counts to 0, sets `fi->fib_dead = 1`, and schedules the `fib_info` object for freeing via RCU.
- Because `fib_net_dump()` is running under `rcu_read_lock()` and does not hold the `RTNL` lock, it can still find the removed `fib_alias` in the trie.
- `fib_leaf_notify()` does not check `fi->fib_dead` and passes the dead `fib_info` to the registered driver's notifier callback.
2. **The Bug in Drivers:**
- Hardware switch drivers like `netdevsim`, `mlxsw`, and `prestera` implement FIB notifier callbacks (e.g., `nsim_fib4_prepare_event()`).
- These callbacks call `fib_info_hold(fi)` to take a reference on the `fib_info` object so they can process it asynchronously in a workqueue.
- `fib_info_hold()` does a `refcount_inc(&fi->fib_clntref)`. Since the reference count is already 0, this triggers the `refcount_t: addition on 0` warning.
- Later, when the driver's workqueue finishes processing, it calls `fib_info_put(fi)`, which drops the reference count from 1 to 0 and calls `free_fib_info()` again, leading to a highly exploitable **double free**.
3. **Triggerability:**
- To trigger this bug, one must cause `fib_net_dump()` to run concurrently with route deletion.
- `fib_net_dump()` is exclusively called by `register_fib_notifier()`.
- `register_fib_notifier()` is only called during the initialization of specific hardware switch drivers (`mlxsw`, `prestera`, `mlx5`, `rocker`) or during `devlink` reload operations (e.g., `DEVLINK_CMD_RELOAD` or when a network namespace containing a devlink instance is destroyed).
- Loading these drivers, invoking `DEVLINK_CMD_RELOAD`, or moving a devlink instance into a network namespace all strictly require `CAP_NET_ADMIN` in the initial network namespace (root privileges).
- Therefore, an unprivileged user cannot trigger `register_fib_notifier()`, even from within a user namespace.
**Conclusion:**
The bug is a memory corruption issue (use-after-free and double free) that is technically exploitable and can cause a Denial of Service (kernel crash). However, it is not accessible to unprivileged users or from within user namespaces, as triggering the vulnerable code path requires root privileges to initialize hardware switch drivers or reload devlink instances.
|
| 6/2 |
2026/05/26 14:37 |
llm |
expert |
0m
Model:
Tokens:
input: 7537
output: 44
thoughts: 846
Thoughts: **Analyzing a Use-After-Free in the netdevsim Driver**
Okay, I've got a refcount-related bug report: a `use-after-free` triggered by an addition on a zero refcount. The call trace points to `lib/refcount.c` and specifically mentions `kworker/u8:15`. Interesting.
Looking at the trace, the problem seems to be originating within the `netdevsim` driver. This rings a bell; it's a simulated networking device, often used in testing. My initial thought is to check if unprivileged users have access to it, given the potential for exploitation, but then I see the function `devlink_pernet_pre_exit`. That suggests this is happening during network namespace cleanup, when a network namespace is being destroyed.
The trace suggests this: `netdevsim` is likely registering per-net ops, and these are being cleaned up during namespace exit. I need to dive into `drivers/net/netdevsim/dev.c` and `drivers/net/netdevsim/fib.c`. The issue is likely within the routines for managing the simulated FIB (Forwarding Information Base). I'll need to understand how the refcounts are being managed during device and namespace destruction within the `netdevsim` driver. The `nsim_fib4_prepare_event` and related functions in `fib.c` seem like the area to focus on, and then how these interact with the overall `devlink` framework.
|
| 7/2 |
2026/05/26 14:37 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:drivers/net/netdevsim/fib.c Name:nsim_fib4_prepare_event]
Results: map[SourceCode: 896: static int nsim_fib4_prepare_event(struct fib_notifier_info *info,
897: struct nsim_fib_event *fib_event,
898: unsigned long event)
899: {
900: struct nsim_fib_data *data = fib_event->data;
901: struct fib_entry_notifier_info *fen_info;
902: struct netlink_ext_ack *extack;
903: int err = 0;
904:
905: fen_info = container_of(info, struct fib_entry_notifier_info,
906: info);
907: fib_event->fen_info = *fen_info;
908: extack = info->extack;
909:
910: switch (event) {
911: case FIB_EVENT_ENTRY_REPLACE:
912: err = nsim_fib_account(&data->ipv4.fib, true);
913: if (err) {
914: NL_SET_ERR_MSG_MOD(extack, "Exceeded number of supported fib entries");
915: return err;
916: }
917: break;
918: case FIB_EVENT_ENTRY_DEL:
919: if (data->fail_route_delete) {
920: NL_SET_ERR_MSG_MOD(extack, "Failed to process route deletion");
921: return -EINVAL;
922: }
923: nsim_fib_account(&data->ipv4.fib, false);
924: break;
925: }
926:
927: /* Take reference on fib_info to prevent it from being
928: * freed while event is queued. Release it afterwards.
929: */
930: fib_info_hold(fib_event->fen_info.fi);
931:
932: return 0;
933: }
SourceFile:drivers/net/netdevsim/fib.c]
|
| 8/2 |
2026/05/26 14:37 |
llm |
expert |
0m
Model:
Tokens:
input: 9352
output: 41
thoughts: 202
Thoughts: **Debugging a Nasty Refcount Issue**
Okay, so I've got a problem. The crash is happening right here: `fib_info_hold(fib_event->fen_info.fi);` and the crucial part is that the refcount is apparently zero at this point. That's never good.
I've traced the call stack, and it's a bit of a journey: `nsim_fib4_prepare_event` -> `nsim_fib_event_nb` -> `call_fib_notifier` -> `fib_net_dump` -> `register_fib_notifier` -> `nsim_fib_create` -> `nsim_dev_reload_create` -> `nsim_dev_reload_up` -> `devlink_reload` -> `devlink_pernet_pre_exit`. That's a lot of layers, but at least I have a path.
Now, wait a second... `devlink_pernet_pre_exit`... where does that even get called from? Oh, right! That gets called during `cleanup_net`. And `cleanup_net` is called when a netns is destroyed. Hmmm... `devlink_pernet_pre_exit` is essentially reloading devlink instances when a netns goes away? This could be interesting. Let me go check that function again and make sure I have the picture. I need to understand what's happening during that tear-down process and why this particular `fib_info` is showing up with a zero refcount at this stage. This looks like a race condition between netns teardown and fib cleanup.
|
| 9/2 |
2026/05/26 14:37 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/devlink/core.c Name:devlink_pernet_pre_exit]
Results: map[SourceCode: 544: static void __net_exit devlink_pernet_pre_exit(struct net *net)
545: {
546: struct devlink *devlink;
547: u32 actions_performed;
548: unsigned long index;
549: int err;
550:
551: /* In case network namespace is getting destroyed, reload
552: * all devlink instances from this namespace into init_net.
553: */
554: devlinks_xa_for_each_registered_get(net, index, devlink) {
555: devl_dev_lock(devlink, true);
556: err = 0;
557: if (devl_is_registered(devlink))
558: err = devlink_reload(devlink, &init_net,
559: DEVLINK_RELOAD_ACTION_DRIVER_REINIT,
560: DEVLINK_RELOAD_LIMIT_UNSPEC,
561: &actions_performed, NULL);
562: devl_dev_unlock(devlink, true);
563: devlink_put(devlink);
564: if (err && err != -EOPNOTSUPP)
565: pr_warn("Failed to reload devlink instance into init_net\n");
566: }
567: }
SourceFile:net/devlink/core.c]
|
| 10/2 |
2026/05/26 14:37 |
llm |
expert |
0m
Model:
Tokens:
input: 10363
output: 41
thoughts: 121
Thoughts: **Debugging a Potential Namespace Issue**
Okay, so `devlink_pernet_pre_exit` is supposed to be moving devlink instances from a dying namespace over to `init_net`. That seems straightforward enough, but something about this is making me pause. During this `devlink_reload` process, we're calling `nsim_dev_reload_up`, which in turn calls `nsim_fib_create`. And that's where things get interesting because `nsim_fib_create` registers a fib notifier. Now, when registering this fib notifier, it looks like `fib_net_dump` is invoked, and that means we are dumping all existing fib entries.
The question is, are we dumping the fib entries from `init_net` *or* the namespace that's being cleaned up? That's the crux of the matter. If we're dumping from `init_net`, this is a significant performance bottleneck, depending on the scale and complexity of `init_net`. And worse, it means entries from the namespace that is being removed might be removed as well. However, if the dump is happening in the dying namespace, we're fine.
I need to double-check that `nsim_fib_create` implementation to make certain the source of the dump.
|
| 11/2 |
2026/05/26 14:37 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:drivers/net/netdevsim/fib.c Name:nsim_fib_create]
Results: map[SourceCode:1552: struct nsim_fib_data *nsim_fib_create(struct devlink *devlink,
1553: struct netlink_ext_ack *extack)
1554: {
1555: struct nsim_fib_data *data;
1556: struct nsim_dev *nsim_dev;
1557: int err;
1558:
1559: data = kzalloc_obj(*data);
1560: if (!data)
1561: return ERR_PTR(-ENOMEM);
1562: data->devlink = devlink;
1563:
1564: nsim_dev = devlink_priv(devlink);
1565: err = nsim_fib_debugfs_init(data, nsim_dev);
1566: if (err)
1567: goto err_data_free;
1568:
1569: mutex_init(&data->nh_lock);
1570: err = rhashtable_init(&data->nexthop_ht, &nsim_nexthop_ht_params);
1571: if (err)
1572: goto err_debugfs_exit;
1573:
1574: mutex_init(&data->fib_lock);
1575: INIT_LIST_HEAD(&data->fib_rt_list);
1576: err = rhashtable_init(&data->fib_rt_ht, &nsim_fib_rt_ht_params);
1577: if (err)
1578: goto err_rhashtable_nexthop_destroy;
1579:
1580: INIT_WORK(&data->fib_event_work, nsim_fib_event_work);
1581: INIT_WORK(&data->fib_flush_work, nsim_fib_flush_work);
1582: INIT_LIST_HEAD(&data->fib_event_queue);
1583: spin_lock_init(&data->fib_event_queue_lock);
1584:
1585: nsim_fib_set_max_all(data, devlink);
1586:
1587: data->nexthop_nb.notifier_call = nsim_nexthop_event_nb;
1588: err = register_nexthop_notifier(devlink_net(devlink), &data->nexthop_nb,
1589: extack);
1590: if (err) {
1591: pr_err("Failed to register nexthop notifier\n");
1592: goto err_rhashtable_fib_destroy;
1593: }
1594:
1595: data->fib_nb.notifier_call = nsim_fib_event_nb;
1596: err = register_fib_notifier(devlink_net(devlink), &data->fib_nb,
1597: nsim_fib_dump_inconsistent, extack);
1598: if (err) {
1599: pr_err("Failed to register fib notifier\n");
1600: goto err_nexthop_nb_unregister;
1601: }
1602:
1603: devl_resource_occ_get_register(devlink,
1604: NSIM_RESOURCE_IPV4_FIB,
1605: nsim_fib_ipv4_resource_occ_get,
1606: data);
1607: devl_resource_occ_get_register(devlink,
1608: NSIM_RESOURCE_IPV4_FIB_RULES,
1609: nsim_fib_ipv4_rules_res_occ_get,
1610: data);
1611: devl_resource_occ_get_register(devlink,
1612: NSIM_RESOURCE_IPV6_FIB,
1613: nsim_fib_ipv6_resource_occ_get,
1614: data);
1615: devl_resource_occ_get_register(devlink,
1616: NSIM_RESOURCE_IPV6_FIB_RULES,
1617: nsim_fib_ipv6_rules_res_occ_get,
1618: data);
1619: devl_resource_occ_get_register(devlink,
1620: NSIM_RESOURCE_NEXTHOPS,
1621: nsim_fib_nexthops_res_occ_get,
1622: data);
1623: return data;
1624:
1625: err_nexthop_nb_unregister:
1626: unregister_nexthop_notifier(devlink_net(devlink), &data->nexthop_nb);
1627: err_rhashtable_fib_destroy:
1628: cancel_work_sync(&data->fib_flush_work);
1629: flush_work(&data->fib_event_work);
1630: rhashtable_free_and_destroy(&data->fib_rt_ht, nsim_fib_rt_free,
1631: data);
1632: err_rhashtable_nexthop_destroy:
1633: rhashtable_free_and_destroy(&data->nexthop_ht, nsim_nexthop_free,
1634: data);
1635: mutex_destroy(&data->fib_lock);
1636: err_debugfs_exit:
1637: mutex_destroy(&data->nh_lock);
1638: nsim_fib_debugfs_exit(data);
1639: err_data_free:
1640: kfree(data);
1641: return ERR_PTR(err);
1642: }
SourceFile:drivers/net/netdevsim/fib.c]
|
| 12/2 |
2026/05/26 14:37 |
llm |
expert |
0m
Model:
Tokens:
input: 12412
output: 36
thoughts: 81
Thoughts: **Analyzing the Network Namespace Transition**
Okay, so I'm looking at this `devlink` code, and I'm following how it handles network namespaces. The key here seems to be the `devlink_net(devlink)` namespace. That's the starting point. Then, in the `devlink_pernet_pre_exit` function, I see a call to `devlink_reload(devlink, &init_net, ...)`. Hmm, interesting. The name `reload` suggests it's altering something about the device's state in relation to its network context.
The question I need to answer is: Does `devlink_reload` actually change the network namespace associated with this `devlink` instance to `init_net`? That's what I'm inferring from the function call, but I need to confirm it. The whole purpose of the change hinges on this. I'll need to dive into the implementation of `devlink_reload` to fully understand what is actually going on. That's where the answer is.
|
| 13/2 |
2026/05/26 14:37 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/devlink/dev.c Name:devlink_reload]
Results: map[SourceCode: 442: int devlink_reload(struct devlink *devlink, struct net *dest_net,
443: enum devlink_reload_action action,
444: enum devlink_reload_limit limit,
445: u32 *actions_performed, struct netlink_ext_ack *extack)
446: {
447: u32 remote_reload_stats[DEVLINK_RELOAD_STATS_ARRAY_SIZE];
448: struct net *curr_net;
449: int err;
450:
451: /* Make sure the reload operations are invoked with the device lock
452: * held to allow drivers to trigger functionality that expects it
453: * (e.g., PCI reset) and to close possible races between these
454: * operations and probe/remove.
455: */
456: if (devlink->dev)
457: device_lock_assert(devlink->dev);
458:
459: memcpy(remote_reload_stats, devlink->stats.remote_reload_stats,
460: sizeof(remote_reload_stats));
461:
462: err = devlink->ops->reload_down(devlink, !!dest_net, action, limit, extack);
463: if (err)
464: return err;
465:
466: curr_net = devlink_net(devlink);
467: if (dest_net && !net_eq(dest_net, curr_net))
468: devlink_reload_netns_change(devlink, curr_net, dest_net);
469:
470: if (action == DEVLINK_RELOAD_ACTION_DRIVER_REINIT) {
471: devlink_params_driverinit_load_new(devlink);
472: devlink_reload_reinit_sanity_check(devlink);
473: }
474:
475: err = devlink->ops->reload_up(devlink, action, limit, actions_performed, extack);
476: devlink_reload_failed_set(devlink, !!err);
477: if (err)
478: return err;
479:
480: WARN_ON(!(*actions_performed & BIT(action)));
481: /* Catch driver on updating the remote action within devlink reload */
482: WARN_ON(memcmp(remote_reload_stats, devlink->stats.remote_reload_stats,
483: sizeof(remote_reload_stats)));
484: devlink_reload_stats_update(devlink, limit, *actions_performed);
485: return 0;
486: }
SourceFile:net/devlink/dev.c]
|
| 14/2 |
2026/05/26 14:37 |
llm |
expert |
0m
Model:
Tokens:
input: 13550
output: 40
thoughts: 240
Thoughts: **Debugging a Mysterious Use-After-Free in Network Namespace Reloading**
Okay, so `devlink_reload_netns_change` is definitely switching the network namespace to `init_net`. That's the expected behavior, I can see that. And then `reload_up` kicks in, leading to `nsim_dev_reload_up`, which makes sense. From there, it's calling `nsim_fib_create`, which, as I suspected, is where the fib notifier registration is happening within `devlink_net(devlink)`. Ah, and now the netns is `init_net` in here as well. Good. That registration, during which all the existing fib entries are dumped, is supposed to be safe. That's when I see the problem. Wait⦠a *use-after-free* or a refcount addition on *zero*? This feels wrong.
The issue looks to be happening with this `fib_info_hold(fib_event->fen_info.fi);`. The `fi` there, the `fib_info`, is showing a refcount of zero. How in the world can a `fib_info` *in the fib trie* have a refcount of zero? This is the core of the problem.
Alright, let's step back and look at the code: during `fib_net_dump` function, it's traversing the fib trie and calling the notifier for each entry. That's the trigger. This means `fib_net_dump` and `fib_table_notify` are going to be critical to understand how this is happening. Time to dig in there.
|
| 15/2 |
2026/05/26 14:37 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/ipv4/fib_trie.c Name:fib_table_notify]
Results: map[SourceCode:2185: static int fib_table_notify(struct fib_table *tb, struct notifier_block *nb,
2186: struct netlink_ext_ack *extack)
2187: {
2188: struct trie *t = (struct trie *)tb->tb_data;
2189: struct key_vector *l, *tp = t->kv;
2190: t_key key = 0;
2191: int err;
2192:
2193: while ((l = leaf_walk_rcu(&tp, key)) != NULL) {
2194: err = fib_leaf_notify(l, tb, nb, extack);
2195: if (err)
2196: return err;
2197:
2198: key = l->key + 1;
2199: /* stop in case of wrap around */
2200: if (key < l->key)
2201: break;
2202: }
2203: return 0;
2204: }
SourceFile:net/ipv4/fib_trie.c]
|
| 16/2 |
2026/05/26 14:37 |
llm |
expert |
0m
Model:
Tokens:
input: 14511
output: 40
thoughts: 11
Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience, and written in the first person:
**Initial Checkpoint: fib\_leaf\_notify**
Right, let's dive in. "fib\_leaf\_notify" β that immediately flags a potential area for investigation. My initial thought process is to understand *why* this specific function is at the forefront of my mind. What is it related to? Routing updates? FIB programming? Leaf node interaction? My expertise lets me cut straight to the core questions: is this related to a performance bottleneck? a recent bug report? a potential security vulnerability? What specific aspects of the FIB (Forwarding Information Base) or its associated notification mechanisms are we talking about here? I need to quickly recall the architecture of the system β what data structures are involved, which locks are in use, and what are the concurrency concerns? The "leaf" designation narrows the scope, implying a specific type of node within the network fabric. I'll need to immediately start running through my mental checklists: what's the typical notification flow, where are the critical sections, and what are the common failure modes associated with that specific function?
|
| 17/2 |
2026/05/26 14:37 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/ipv4/fib_trie.c Name:fib_leaf_notify]
Results: map[SourceCode:2152: static int fib_leaf_notify(struct key_vector *l, struct fib_table *tb,
2153: struct notifier_block *nb,
2154: struct netlink_ext_ack *extack)
2155: {
2156: struct fib_alias *fa;
2157: int last_slen = -1;
2158: int err;
2159:
2160: hlist_for_each_entry_rcu(fa, &l->leaf, fa_list) {
2161: struct fib_info *fi = fa->fa_info;
2162:
2163: if (!fi)
2164: continue;
2165:
2166: /* local and main table can share the same trie,
2167: * so don't notify twice for the same entry.
2168: */
2169: if (tb->tb_id != fa->tb_id)
2170: continue;
2171:
2172: if (fa->fa_slen == last_slen)
2173: continue;
2174:
2175: last_slen = fa->fa_slen;
2176: err = call_fib_entry_notifier(nb, FIB_EVENT_ENTRY_REPLACE,
2177: l->key, KEYLENGTH - fa->fa_slen,
2178: fa, extack);
2179: if (err)
2180: return err;
2181: }
2182: return 0;
2183: }
SourceFile:net/ipv4/fib_trie.c]
|
| 18/2 |
2026/05/26 14:37 |
llm |
expert |
0m
Model:
Tokens:
input: 15292
output: 38
thoughts: 219
Thoughts: **My Analysis of a Potential Data Race in FIB Notification**
Okay, let's break this down. In `fib_leaf_notify`, I'm iterating through the `fa_list` of the leaf. My first thought is: could `fa->fa_info`, or `fi`, be in a destroyed state? Specifically, is it possible that its reference count is zero, but it's still somehow present within the trie data structure? This feels like a classic race condition scenario.
Yes, the `fib_info` object *is* destroyed asynchronously, or when its reference count hits zero. It's likely that it gets removed from the trie sometime *after* this happens. Interesting⦠wait a second, a `fib_info` is shared between multiple `fib_alias` entries, that makes sense now! When a route gets deleted, the `fib_alias` is removed from the trie and `fib_info_put` gets called.
But hold on⦠I recall `fib_net_dump` gets called under `rcu_read_lock()` - because it uses `leaf_walk_rcu` and `hlist_for_each_entry_rcu`. So, if that's the case⦠Hang on a minute, when I'm walking this tree and removing routes the tree could be modified under it's nose - this is where the race condition would rear it's ugly head.
Wait a second... doesn't `register_fib_notifier` also use `rcu_read_lock()`? I'd better double-check `register_fib_notifier` to be absolutely certain...
|
| 19/2 |
2026/05/26 14:37 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/core/fib_notifier.c Name:register_fib_notifier]
Results: map[SourceCode: 98: int register_fib_notifier(struct net *net, struct notifier_block *nb,
99: void (*cb)(struct notifier_block *nb),
100: struct netlink_ext_ack *extack)
101: {
102: int retries = 0;
103: int err;
104:
105: do {
106: unsigned int fib_seq = fib_seq_sum(net);
107:
108: err = fib_net_dump(net, nb, extack);
109: if (err)
110: return err;
111:
112: if (fib_dump_is_consistent(net, nb, cb, fib_seq))
113: return 0;
114: } while (++retries < FIB_DUMP_MAX_RETRIES);
115:
116: return -EBUSY;
117: }
SourceFile:net/core/fib_notifier.c]
|
| 20/2 |
2026/05/26 14:37 |
llm |
expert |
0m
Model:
Tokens:
input: 16207
output: 38
thoughts: 11
Thoughts: Okay, let's dive into this. Here's my thought process summarized, assuming I'm already well-versed in the subject matter:
**Checking the `fib_net_dump`**
Alright, time to get to it. My first thought is to zero in on `fib_net_dump`. Itβs a crucial tool, so let's start there. I'm going to quickly check `fib_net_dump` to see what information it's providing. Is it giving me the expected output? Are there any obvious red flags jumping out at me? I need to get a high-level view of the FIB's state and potentially identify areas requiring closer inspection. This initial scan is about gaining situational awareness; I'm looking for anything unexpected or unusual to guide my next steps. This is the starting point for effective troubleshooting; get the initial data and go from there.
|
| 21/2 |
2026/05/26 14:37 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/core/fib_notifier.c Name:fib_net_dump]
Results: map[SourceCode: 59: static int fib_net_dump(struct net *net, struct notifier_block *nb,
60: struct netlink_ext_ack *extack)
61: {
62: struct fib_notifier_net *fn_net = net_generic(net, fib_notifier_net_id);
63: struct fib_notifier_ops *ops;
64: int err = 0;
65:
66: rcu_read_lock();
67: list_for_each_entry_rcu(ops, &fn_net->fib_notifier_ops, list) {
68: if (!try_module_get(ops->owner))
69: continue;
70: err = ops->fib_dump(net, nb, extack);
71: module_put(ops->owner);
72: if (err)
73: goto unlock;
74: }
75:
76: unlock:
77: rcu_read_unlock();
78:
79: return err;
80: }
SourceFile:net/core/fib_notifier.c]
|
| 22/2 |
2026/05/26 14:37 |
llm |
expert |
0m
Model:
Tokens:
input: 16772
output: 38
thoughts: 346
Thoughts: **RCU, Refcounts, and the Dreaded Use-After-Free**
Okay, so `fib_net_dump` grabs an `rcu_read_lock()`. That means `fib_leaf_notify` is running under the RCU read lock as well. This is where things get interesting. It finds a `fib_alias` within the trie and snags its `fa_info`. Concurrently, the route could be deleted. If that happens, the `fib_alias` gets yanked, and `fib_info_put(fa->fa_info)` is called. `fib_info_put` zeroes out the refcount, which then triggers `free_fib_info_rcu` via `call_rcu`.
Here's the rub: we're *still* under that `rcu_read_lock()`. So, the `fib_info` memory isn't actually freed yet. It's essentially sitting in limbo, but its refcount is already down to zero! This is the key.
Now, `nsim_fib4_prepare_event` pops up and calls `fib_info_hold(fi)`. *This* is where I think the trouble starts. `fib_info_hold` then calls `refcount_inc(&fi->fib_treeref)`. But... the refcount is zero! And `refcount_inc` *warns* when incrementing a zeroed refcount. This is a classic recipe for a use-after-free, or more accurately, an addition on zero.
The question then is, is it allowed to call `fib_info_hold` on a `fib_info` obtained from a `fib_alias` under an RCU read lock? Ordinarily, if an object is RCU protected and its refcount can reach zero, one is supposed to use `refcount_inc_not_zero` (or an equivalent `atomic_inc_not_zero`). I need to double-check the `fib_info_hold` implementation to verify this.
|
| 23/2 |
2026/05/26 14:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:include/net/ip_fib.h Name:fib_info_hold]
Results: map[SourceCode: 627: static inline void fib_info_hold(struct fib_info *fi)
628: {
629: refcount_inc(&fi->fib_clntref);
630: }
SourceFile:include/net/ip_fib.h]
|
| 24/2 |
2026/05/26 14:38 |
llm |
expert |
0m
Model:
Tokens:
input: 17688
output: 36
thoughts: 32
Thoughts: Okay, here's my take on those thoughts, framed as an expert's internal monologue:
**Diving into the Reference Count Mystery**
Hmm, interesting. My initial assumption was that the `fib_treeref` would be the target of the increment, but it looks like the code is actually touching `fib_clntref` instead. Right, that's a common oversight, and an important distinction. We need to be absolutely certain we are incrementing the correct reference counter for the intended purpose.
Before I go any further down that path, let's just make sure I have the architecture fresh in my mind. Let's briefly check the `fib_info` structure itself. It's too easy to make assumptions here, and a quick re-examination of the data structures can often provide the answer. Okay, let's see.
|
| 25/2 |
2026/05/26 14:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:include/net/ip_fib.h Name:fib_info]
Results: map[SourceCode: 136: struct fib_info {
137: struct hlist_node fib_hash;
138: struct hlist_node fib_lhash;
139: struct list_head nh_list;
140: struct net *fib_net;
141: refcount_t fib_treeref;
142: refcount_t fib_clntref;
143: unsigned int fib_flags;
144: unsigned char fib_dead;
145: unsigned char fib_protocol;
146: unsigned char fib_scope;
147: unsigned char fib_type;
148: __be32 fib_prefsrc;
149: u32 fib_tb_id;
150: u32 fib_priority;
151: struct dst_metrics *fib_metrics;
152: #define fib_mtu fib_metrics->metrics[RTAX_MTU-1]
153: #define fib_window fib_metrics->metrics[RTAX_WINDOW-1]
154: #define fib_rtt fib_metrics->metrics[RTAX_RTT-1]
155: #define fib_advmss fib_metrics->metrics[RTAX_ADVMSS-1]
156: int fib_nhs;
157: bool fib_nh_is_v6;
158: bool nh_updated;
159: bool pfsrc_removed;
160: struct nexthop *nh;
161: struct rcu_head rcu;
162: struct fib_nh fib_nh[] __counted_by(fib_nhs);
163: };
SourceFile:include/net/ip_fib.h]
|
| 26/2 |
2026/05/26 14:38 |
llm |
expert |
0m
Model:
Tokens:
input: 18445
output: 38
thoughts: 32
Thoughts: Okay, here's my take on those thoughts, framed as an expert's internal monologue:
**Deconstructing Reference Counting in this Fibonacci Implementation**
Alright, so I'm seeing two distinct reference counts here: `fib_treeref` and `fib_clntref`. I need to understand their roles within this Fibonacci implementation, and the interplay between them. My immediate thought is to check how these reference counts are managed during deallocation or cleanup. That means I need to dive into the `fib_info_put` function. Itβs the likely entry point for examining how these counts are decremented and what actions are taken when they hit zero, which is really where the core memory management logic resides. I suspect this function will be the key to understanding potential memory leaks, dangling pointers, and overall resource management soundness. Let's see what `fib_info_put` tells us.
|
| 27/2 |
2026/05/26 14:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:include/net/ip_fib.h Name:fib_info_put]
Results: map[SourceCode: 632: static inline void fib_info_put(struct fib_info *fi)
633: {
634: if (refcount_dec_and_test(&fi->fib_clntref))
635: free_fib_info(fi);
636: }
SourceFile:include/net/ip_fib.h]
|
| 28/2 |
2026/05/26 14:38 |
llm |
expert |
0m
Model:
Tokens:
input: 18813
output: 50
thoughts: 28
Thoughts: Here's my attempt at summarizing the thought process, tailored for an expert audience:
**Analyzing Reference Counting in the FIB Infrastructure**
Okay, so I'm looking at this `fib_info_put` function and I'm immediately noticing it's decrementing `fib_clntref`. That's standard and expected for client references. But my mind is instantly jumping to the other reference counter, `fib_treeref`. We're dealing with a FIB (Forwarding Information Base) data structure, and trees are core to that. So, does `fib_info_put` also have any interaction with `fib_treeref`? Or are they independent, and if so, what's the logic for the tree references? Is there a separate cleanup path? I need to be completely clear on the full reference counting lifecycle for *both* types of references to understand the memory management and avoid potential leaks or double-frees, especially in a potentially concurrent environment. I need to trace the control flow to ensure this is sound.
|
| 29/2 |
2026/05/26 14:38 |
tool |
codesearch-find-references |
0m
Error:
Args: map[ContextFile:include/net/ip_fib.h IncludeSnippetLines:0 Name:fib_treeref SourceTreePrefix:]
Results: map[References:<nil> TruncatedOutput:false]
|
| 30/2 |
2026/05/26 14:38 |
llm |
expert |
0m
Model:
Tokens:
input: 19128
output: 54
thoughts: 21
Thoughts: Here's my thought process, summarized as you requested:
**Navigating the Cognitive Landscape of Code**
Okay, I'm looking at this... it's clearly a field of some sort. My initial instinct, based on the context and my experience, is to reach for a specific construct. That construct, in this case, is likely `fib_info::fib_treeref`. This feels like the appropriate, and probably most efficient, way to handle this data structure given the current scope. It's a mental shortcut, a honed intuition born from years of working in this domain. I immediately see the potential here and recognize how `fib_info::fib_treeref` fits this context.
|
| 31/2 |
2026/05/26 14:38 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:include/net/ip_fib.h IncludeSnippetLines:0 Name:fib_info::fib_treeref SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:fib_create_info SourceFile:net/ipv4/fib_semantics.c SourceLine:1377] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:fib_create_info SourceFile:net/ipv4/fib_semantics.c SourceLine:1535] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:fib_create_info SourceFile:net/ipv4/fib_semantics.c SourceLine:1539] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:fib_release_info SourceFile:net/ipv4/fib_semantics.c SourceLine:253]] TruncatedOutput:false]
|
| 32/2 |
2026/05/26 14:38 |
llm |
expert |
0m
Model:
Tokens:
input: 19596
output: 39
thoughts: 11
Thoughts: Okay, here's my summarized thought process, presented as if I'm thinking it:
**Examining `fib_release_info`**
Alright, time to dive in. Let's see... the code's asking me to look at `fib_release_info`. Right. First, I need to understand what this variable/object actually *is*. What's it meant to contain? Is it a data structure? A configuration file? A dictionary, perhaps? Does it hold information about the Fibonnaci sequence implementation's version, build details, or environment parameters? Or is it a struct of sorts with associated functions? My initial thought is it is associated with a fibonacci sequence. I need to figure out how it's being used within the larger context. Is it being loaded from a config file, set dynamically, or perhaps hardcoded somewhere? And what does its name suggest about its purpose? Let me start by tracing where it's defined and how it's being accessed. Knowing this context is key to understanding its role. From there, I can determine the actual intent, which I will need to begin with my examination.
|
| 33/2 |
2026/05/26 14:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/ipv4/fib_semantics.c Name:fib_release_info]
Results: map[SourceCode: 250: void fib_release_info(struct fib_info *fi)
251: {
252: ASSERT_RTNL();
253: if (fi && refcount_dec_and_test(&fi->fib_treeref)) {
254: hlist_del(&fi->fib_hash);
255: fi->fib_net->ipv4.fib_info_cnt--;
256:
257: if (fi->fib_prefsrc)
258: hlist_del(&fi->fib_lhash);
259: if (fi->nh) {
260: list_del(&fi->nh_list);
261: } else {
262: change_nexthops(fi) {
263: if (!nexthop_nh->fib_nh_dev)
264: continue;
265: hlist_del_rcu(&nexthop_nh->nh_hash);
266: } endfor_nexthops(fi)
267: }
268: /* Paired with READ_ONCE() from fib_table_lookup() */
269: WRITE_ONCE(fi->fib_dead, 1);
270: fib_info_put(fi);
271: }
272: }
SourceFile:net/ipv4/fib_semantics.c]
|
| 34/2 |
2026/05/26 14:38 |
llm |
expert |
0m
Model:
Tokens:
input: 20282
output: 39
thoughts: 156
Thoughts: **Deconstructing the Fibonacci Trie Cleanup**
Okay, so I'm looking at this `fib_treeref` variable, and it's acting as the central reference counter for the trie. I get it β when this count plummets to zero, it signals that the Fibonacci entry, `fi`, is no longer actively used within the trie structure itself. So, the system then pulls `fi` out of all the relevant hash tables, flags it as "dead" by setting `fib_dead = 1`, and then kicks off a call to `fib_info_put(fi)`.
Now, `fib_info_put`β¦ thatβs where the cleanup continues. Itβs responsible for decrementing `fib_clntref`, which, let me seeβ¦ ah, right, `fib_clntref` is the reference counter that gets incremented when `fib_info_hold` is used. So, if `fib_clntref` also hits zero after the decrement, then `free_fib_info(fi)` gets the call. That makes sense. `free_fib_info` is probably the final cleanup step.
I need to quickly dive into what `free_fib_info` actually does, but the general flow here seems pretty straightforward: `fib_treeref` manages the trie references, and `fib_clntref` manages the client-side references. Both must be zero for full deallocation.
|
| 35/2 |
2026/05/26 14:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/ipv4/fib_semantics.c Name:free_fib_info]
Results: map[SourceCode: 239: void free_fib_info(struct fib_info *fi)
240: {
241: if (fi->fib_dead == 0) {
242: pr_warn("Freeing alive fib_info %p\n", fi);
243: return;
244: }
245:
246: call_rcu_hurry(&fi->rcu, free_fib_info_rcu);
247: }
SourceFile:net/ipv4/fib_semantics.c]
|
| 36/2 |
2026/05/26 14:38 |
llm |
expert |
0m
Model:
Tokens:
input: 20951
output: 52
thoughts: 320
Thoughts: **Analyzing a Potential Race Condition in the FIB Trie Traversal**
Okay, so I'm looking at how `free_fib_info` is used within this networking code. It's using `call_rcu_hurry`, which immediately flags a potential concurrency issue related to the RCU grace period. `fib_info` is freed *after* that grace period. Now, in `fib_net_dump`, which is a common function for iterating and displaying FIB information, we're taking an `rcu_read_lock()`. That seems right, but let's see how it's used.
It iterates through the FIB trie, and finds a `fib_alias`, gets the `fi` via `fa->fa_info`. And then calls `fib_leaf_notify`, which calls a series of functions that ultimately leads to `fib_info_hold(fi)`. *That's* interesting. `fib_info_hold` is key here, because if `fib_release_info` is called concurrently, it will decrement `fib_treeref`. If it hits zero, `fib_info_put(fi)` is called, and if `fib_clntref` also hits zero, `free_fib_info(fi)` is *eventually* called.
Crucially, since we're under `rcu_read_lock()`, `fi` is still valid memory, and *hasn't* been freed yet. But `fib_info_hold(fi)` *increments* `fi->fib_clntref` with `refcount_inc`.
Aha! If `fib_clntref` *was* zero, that `refcount_inc` is going to warn about an increment on zero, and *that's* the bug! I see it! It's a textbook use-after-free scenario, or more accurately, a race condition.
Wait, though⦠is the `fib_alias` guaranteed to be removed from the trie *before* `fib_release_info` is even called? That's the other angle of attack here. I'll need to check the call sites for `fib_release_info`. It's possible the race isn't there. But it *looks* like the code is fundamentally flawed if the alias is visible during the RCU grace period.
|
| 37/2 |
2026/05/26 14:38 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/ipv4/fib_semantics.c IncludeSnippetLines:0 Name:fib_release_info SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:fib_table_delete SourceFile:net/ipv4/fib_trie.c SourceLine:1751] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:fib_table_flush SourceFile:net/ipv4/fib_trie.c SourceLine:2071] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:fib_table_insert SourceFile:net/ipv4/fib_trie.c SourceLine:1314] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:fib_table_insert SourceFile:net/ipv4/fib_trie.c SourceLine:1390]] TruncatedOutput:false]
|
| 38/2 |
2026/05/26 14:38 |
llm |
expert |
0m
Model:
Tokens:
input: 22065
output: 40
thoughts: 11
Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience, and written in the first person:
**Analyzing the `fib_table_delete` Function**
Alright, time to dive in. Let's get right to the heart of it β I need to take a look at `fib_table_delete`. What's the context? What problem are we trying to solve here, and where does `fib_table_delete` fit into the larger Fibonacci table management system? I'm immediately thinking about the potential implications. Is this about memory deallocation, invalidating entries, or perhaps something more nuanced like managing a cache of Fibonacci numbers to avoid redundant calculations? I need to quickly grasp the design principles underlying this deletion function. Is it expected to clean up dynamically allocated resources? Does it affect the data structure's overall integrity? What are the potential consequences of improperly implementing or misusing this function? Any potential for memory leaks, or worse, segmentation faults? I need to understand its edge cases. Are we dealing with an empty table, or trying to delete an item that isn't present? What kind of error handling is in place? I will expect clear documentation and testing to back this up, to keep us from making mistakes. I am prepared to dig into the function definition, examine its dependencies, and scrutinize its usage within the overall system.
|
| 39/2 |
2026/05/26 14:38 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/ipv4/fib_trie.c Name:fib_table_delete]
Results: map[SourceCode:1687: /* Caller must hold RTNL. */
1688: int fib_table_delete(struct net *net, struct fib_table *tb,
1689: struct fib_config *cfg, struct netlink_ext_ack *extack)
1690: {
1691: struct trie *t = (struct trie *) tb->tb_data;
1692: struct fib_alias *fa, *fa_to_delete;
1693: struct key_vector *l, *tp;
1694: u8 plen = cfg->fc_dst_len;
1695: u8 slen = KEYLENGTH - plen;
1696: dscp_t dscp;
1697: u32 key;
1698:
1699: key = ntohl(cfg->fc_dst);
1700:
1701: l = fib_find_node(t, &tp, key);
1702: if (!l)
1703: return -ESRCH;
1704:
1705: dscp = cfg->fc_dscp;
1706: fa = fib_find_alias(&l->leaf, slen, dscp, 0, tb->tb_id, false);
1707: if (!fa)
1708: return -ESRCH;
1709:
1710: pr_debug("Deleting %08x/%d dsfield=0x%02x t=%p\n", key, plen,
1711: inet_dscp_to_dsfield(dscp), t);
1712:
1713: fa_to_delete = NULL;
1714: hlist_for_each_entry_from(fa, fa_list) {
1715: struct fib_info *fi = fa->fa_info;
1716:
1717: if ((fa->fa_slen != slen) ||
1718: (fa->tb_id != tb->tb_id) ||
1719: (fa->fa_dscp != dscp))
1720: break;
1721:
1722: if ((!cfg->fc_type || fa->fa_type == cfg->fc_type) &&
1723: (cfg->fc_scope == RT_SCOPE_NOWHERE ||
1724: fa->fa_info->fib_scope == cfg->fc_scope) &&
1725: (!cfg->fc_prefsrc ||
1726: fi->fib_prefsrc == cfg->fc_prefsrc) &&
1727: (!cfg->fc_protocol ||
1728: fi->fib_protocol == cfg->fc_protocol) &&
1729: fib_nh_match(net, cfg, fi, extack) == 0 &&
1730: fib_metrics_match(cfg, fi)) {
1731: fa_to_delete = fa;
1732: break;
1733: }
1734: }
1735:
1736: if (!fa_to_delete)
1737: return -ESRCH;
1738:
1739: fib_notify_alias_delete(net, key, &l->leaf, fa_to_delete, extack);
1740: rtmsg_fib(RTM_DELROUTE, htonl(key), fa_to_delete, plen, tb->tb_id,
1741: &cfg->fc_nlinfo, 0);
1742:
1743: if (!plen)
1744: tb->tb_num_default--;
1745:
1746: fib_remove_alias(t, tp, l, fa_to_delete);
1747:
1748: if (READ_ONCE(fa_to_delete->fa_state) & FA_S_ACCESSED)
1749: rt_cache_flush(cfg->fc_nlinfo.nl_net);
1750:
1751: fib_release_info(fa_to_delete->fa_info);
1752: alias_free_mem_rcu(fa_to_delete);
1753: return 0;
1754: }
SourceFile:net/ipv4/fib_trie.c]
|
| 40/2 |
2026/05/26 14:38 |
llm |
expert |
0m
Model:
Tokens:
input: 23582
output: 51
thoughts: 224
Thoughts: **Analyzing a Potential Race Condition in the FIB Table Deletion**
Okay, so I'm looking at `fib_table_delete` and specifically how aliases are removed. It looks like `fib_remove_alias` takes the alias off the list (`hlist_del_rcu(&fa->fa_list)`). The next step is calling `fib_release_info(fa_to_delete->fa_info)`. The issue is, because we're using `hlist_del_rcu`, a concurrent reader using `hlist_for_each_entry_rcu` could still potentially see that alias *after* it's been removed from the list, right?
If that reader *does* see the alias at that precise moment, it grabs the `fi = fa->fa_info` and then calls `fib_info_hold(fi)`. But here's where it gets dicey: `fib_release_info` *might* have already nuked the `fib_treeref` and `fib_clntref` counters to zero. So `fib_info_hold` is essentially trying to `refcount_inc` on a zeroed reference count. That's a classic race condition and could lead to some nasty problems.
Wait a second⦠`fib_info_hold`⦠doesn't that ring a bell? I seem to recall that's only used by `netdevsim`. Okay, I need to double-check where `fib_info_hold` is actually used. Let's trace all calls to `fib_info_hold` and verify the usage. Gotta make sure this isn't a problem beyond `netdevsim`.
|
| 41/2 |
2026/05/26 14:38 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:include/net/ip_fib.h IncludeSnippetLines:0 Name:fib_info_hold SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:nsim_fib4_prepare_event SourceFile:drivers/net/netdevsim/fib.c SourceLine:930] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:nsim_fib4_rt_create SourceFile:drivers/net/netdevsim/fib.c SourceLine:288]] TruncatedOutput:false]
|
| 42/2 |
2026/05/26 14:38 |
llm |
expert |
0m
Model:
Tokens:
input: 24353
output: 38
thoughts: 78
Thoughts: Okay, here's my take on those thoughts, framed from my own perspective:
**Debugging `fib_info_hold` in `netdevsim`**
Alright, so I'm digging into this `fib_info_hold` thing. It seems, according to the code, that it's *only* being used by `netdevsim`. That's good to know, narrows down the scope of where I need to focus.
But...wait a minute. Is this usage *correct*? Something doesn't feel right. My spidey-sense is tingling about potential RCU issues. I'm pretty sure that if `fib_info_hold` is being called under an RCU read lock β which is a very real possibility given the context β we absolutely *need* to be using `refcount_inc_not_zero`, not just a straight `refcount_inc`. That's crucial for avoiding memory corruption under RCU.
Okay, time to double-check the definition of `fib_info_hold` itself and verify this suspicion. Need to confirm how it's defined and *specifically* what its behavior is when dealing with RCU protections. Let's get that definition up...
|
| 43/2 |
2026/05/26 14:39 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:include/net/ip_fib.h Name:fib_info_hold]
Results: map[SourceCode: 627: static inline void fib_info_hold(struct fib_info *fi)
628: {
629: refcount_inc(&fi->fib_clntref);
630: }
SourceFile:include/net/ip_fib.h]
|
| 44/2 |
2026/05/26 14:39 |
llm |
expert |
0m
Model:
Tokens:
input: 24818
output: 51
thoughts: 156
Thoughts: **Analyzing the Scope of a Potential Vulnerability**
Okay, so I'm looking at this `netdevsim` code, and my first thought is, "Hang on, this is a test driver, right? Could this bug *actually* be exploitable in a real-world scenario?" If `netdevsim` isn't loaded in a production environment, then, theoretically, this *specific* bug within `netdevsim` is a dead end. But, I can't be too hasty.
Let me rewind, and reassess, it's not a complete dismissal until I do some more digging. Is there *anywhere* else in the kernel code that might hold a reference to `fib_info` while under an RCU read lock? Let's go back and examine those `fib_info_hold` references one more time, no, they're isolated to `netdevsim`.
Alright, fine, but what about the broader usage of `fib_info`? We need to look at any other potential entry points. When `fib_net_dump` gets called, it calls `fib_leaf_notify`, which calls `call_fib_entry_notifier`. That's interesting, and it looks like other modules can register fib notifiers.
So, the real question becomes, who is registering these fib notifiers? We need to follow that path to see if the vulnerability is exposed in production. That's the key: if another module is calling the same functions or exposing the same conditions, we need to carefully analyze the potential impact. Let's see.
|
| 45/2 |
2026/05/26 14:39 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/core/fib_notifier.c IncludeSnippetLines:0 Name:register_fib_notifier SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:nsim_fib_create SourceFile:drivers/net/netdevsim/fib.c SourceLine:1596]] TruncatedOutput:false]
|
| 46/2 |
2026/05/26 14:39 |
llm |
expert |
0m
Model:
Tokens:
input: 25435
output: 51
thoughts: 36
Thoughts: Here's my take:
**Investigating the Scope of `register_fib_notifier`**
Okay, hold on a second. This `register_fib_notifier` function...it seems to be exclusively called within the `netdevsim` module? That seems odd. Let's make sure that's truly the case. I need to perform a comprehensive search across the entire kernel source tree. I want to be *absolutely* certain that this function's usage is limited to `netdevsim`, or identify any other potential call sites that might be hidden somewhere. I'll meticulously check every instance of its definition and usage. Gotta make sure there aren't any unexpected dependencies or interactions lurking in there, especially considering its name. A global search is definitely warranted here.
|
| 47/2 |
2026/05/26 14:39 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/core/fib_notifier.c IncludeSnippetLines:0 Name:register_fib_notifier SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:nsim_fib_create SourceFile:drivers/net/netdevsim/fib.c SourceLine:1596]] TruncatedOutput:false]
|
| 48/2 |
2026/05/26 14:39 |
llm |
expert |
0m
Model:
Tokens:
input: 25759
output: 19
thoughts: 17
Thoughts: Okay, here's my interpretation of those thoughts, tailored for an expert audience, presented in the first person:
**Investigating the `register_fib_notifier` Export**
Alright, hang on a second. `register_fib_notifier` is *exported*. That immediately raises a flag. Exporting functions like this, especially something that seems so core, implies it's intended for external use. It's not inherently wrong, but it means I need to understand *who* might be using it and *why*. What other modules or components in this system are actually leveraging this functionality? This could have significant implications for dependencies, versioning, and potential API stability concerns down the line. I need to track down all the call sites to get a clear picture of its usage before I can make an informed decision about it. I have some investigation to do...
|
| 49/2 |
2026/05/26 14:39 |
tool |
grepper |
0m
Args: map[Expression:register_fib_notifier]
Results: map[Output:Documentation/networking/switchdev.rst=349=To program the device, the driver has to register a FIB notifier handler
Documentation/networking/switchdev.rst:350:using register_fib_notifier. The following events are available:
Documentation/networking/switchdev.rst-351-
--
drivers/net/ethernet/marvell/prestera/prestera_router.c=1548=int prestera_router_init(struct prestera_switch *sw)
--
drivers/net/ethernet/marvell/prestera/prestera_router.c-1600- router->fib_nb.notifier_call = __prestera_router_fib_event;
drivers/net/ethernet/marvell/prestera/prestera_router.c:1601: err = register_fib_notifier(&init_net, &router->fib_nb,
drivers/net/ethernet/marvell/prestera/prestera_router.c-1602- /* TODO: flush fib entries */ NULL, NULL);
drivers/net/ethernet/marvell/prestera/prestera_router.c-1603- if (err)
drivers/net/ethernet/marvell/prestera/prestera_router.c:1604: goto err_register_fib_notifier;
drivers/net/ethernet/marvell/prestera/prestera_router.c-1605-
--
drivers/net/ethernet/marvell/prestera/prestera_router.c-1607-
drivers/net/ethernet/marvell/prestera/prestera_router.c:1608:err_register_fib_notifier:
drivers/net/ethernet/marvell/prestera/prestera_router.c-1609- unregister_netevent_notifier(&router->netevent_nb);
--
drivers/net/ethernet/marvell/prestera/prestera_router.c=1629=void prestera_router_fini(struct prestera_switch *sw)
drivers/net/ethernet/marvell/prestera/prestera_router.c-1630-{
drivers/net/ethernet/marvell/prestera/prestera_router.c:1631: unregister_fib_notifier(&init_net, &sw->router->fib_nb);
drivers/net/ethernet/marvell/prestera/prestera_router.c-1632- unregister_netevent_notifier(&sw->router->netevent_nb);
--
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c=1860=struct mlx5e_tc_tun_encap *mlx5e_tc_tun_init(struct mlx5e_priv *priv)
--
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c-1872- hash_init(encap->route_tbl);
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c:1873: err = register_fib_notifier(dev_net(priv->netdev), &encap->fib_nb,
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c-1874- NULL, NULL);
--
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c=1883=void mlx5e_tc_tun_cleanup(struct mlx5e_tc_tun_encap *encap)
--
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c-1887-
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c:1888: unregister_fib_notifier(dev_net(encap->priv->netdev), &encap->fib_nb);
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c-1889- flush_workqueue(encap->priv->wq); /* flush fib event works */
--
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c=5393=int mlx5e_tc_esw_init(struct mlx5_rep_uplink_priv *uplink_priv)
--
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c-5449- err = PTR_ERR(uplink_priv->encap);
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c:5450: goto err_register_fib_notifier;
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c-5451- }
--
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c-5470- mlx5e_tc_tun_cleanup(uplink_priv->encap);
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c:5471:err_register_fib_notifier:
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c-5472- mapping_destroy(uplink_priv->tunnel_enc_opts_mapping);
--
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c=368=int mlx5_lag_mp_init(struct mlx5_lag *ldev)
--
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c-385- mp->fib_nb.notifier_call = mlx5_lag_fib_event;
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c:386: err = register_fib_notifier(&init_net, &mp->fib_nb,
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c-387- mlx5_lag_fib_event_flush, NULL);
--
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c=396=void mlx5_lag_mp_cleanup(struct mlx5_lag *ldev)
--
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c-402-
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c:403: unregister_fib_notifier(&init_net, &mp->fib_nb);
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c-404- destroy_workqueue(mp->wq);
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c=11575=int mlxsw_sp_router_init(struct mlxsw_sp *mlxsw_sp,
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11695- mlxsw_sp->router->fib_nb.notifier_call = mlxsw_sp_router_fib_event;
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c:11696: err = register_fib_notifier(mlxsw_sp_net(mlxsw_sp),
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11697- &mlxsw_sp->router->fib_nb,
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11699- if (err)
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c:11700: goto err_register_fib_notifier;
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11701-
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11703-
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c:11704:err_register_fib_notifier:
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11705- unregister_nexthop_notifier(mlxsw_sp_net(mlxsw_sp),
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c=11754=void mlxsw_sp_router_fini(struct mlxsw_sp *mlxsw_sp)
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11757-
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c:11758: unregister_fib_notifier(mlxsw_sp_net(mlxsw_sp), &router->fib_nb);
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11759- unregister_nexthop_notifier(mlxsw_sp_net(mlxsw_sp),
--
drivers/net/ethernet/rocker/rocker_main.c=2844=static int rocker_probe(struct pci_dev *pdev, const struct pci_device_id *id)
--
drivers/net/ethernet/rocker/rocker_main.c-2942- rocker->fib_nb.notifier_call = rocker_router_fib_event;
drivers/net/ethernet/rocker/rocker_main.c:2943: err = register_fib_notifier(&init_net, &rocker->fib_nb, NULL, NULL);
drivers/net/ethernet/rocker/rocker_main.c-2944- if (err)
drivers/net/ethernet/rocker/rocker_main.c:2945: goto err_register_fib_notifier;
drivers/net/ethernet/rocker/rocker_main.c-2946-
--
drivers/net/ethernet/rocker/rocker_main.c-2969-err_register_switchdev_notifier:
drivers/net/ethernet/rocker/rocker_main.c:2970: unregister_fib_notifier(&init_net, &rocker->fib_nb);
drivers/net/ethernet/rocker/rocker_main.c:2971:err_register_fib_notifier:
drivers/net/ethernet/rocker/rocker_main.c-2972- rocker_remove_ports(rocker);
--
drivers/net/ethernet/rocker/rocker_main.c=2997=static void rocker_remove(struct pci_dev *pdev)
--
drivers/net/ethernet/rocker/rocker_main.c-3005- unregister_switchdev_notifier(&rocker_switchdev_notifier);
drivers/net/ethernet/rocker/rocker_main.c:3006: unregister_fib_notifier(&init_net, &rocker->fib_nb);
drivers/net/ethernet/rocker/rocker_main.c-3007- rocker_remove_ports(rocker);
--
drivers/net/netdevsim/fib.c=1552=struct nsim_fib_data *nsim_fib_create(struct devlink *devlink,
--
drivers/net/netdevsim/fib.c-1595- data->fib_nb.notifier_call = nsim_fib_event_nb;
drivers/net/netdevsim/fib.c:1596: err = register_fib_notifier(devlink_net(devlink), &data->fib_nb,
drivers/net/netdevsim/fib.c-1597- nsim_fib_dump_inconsistent, extack);
--
drivers/net/netdevsim/fib.c=1644=void nsim_fib_destroy(struct devlink *devlink, struct nsim_fib_data *data)
--
drivers/net/netdevsim/fib.c-1655- NSIM_RESOURCE_IPV4_FIB);
drivers/net/netdevsim/fib.c:1656: unregister_fib_notifier(devlink_net(devlink), &data->fib_nb);
drivers/net/netdevsim/fib.c-1657- unregister_nexthop_notifier(devlink_net(devlink), &data->nexthop_nb);
--
include/net/fib_notifier.h=41=int call_fib_notifiers(struct net *net, enum fib_event_type event_type,
include/net/fib_notifier.h-42- struct fib_notifier_info *info);
include/net/fib_notifier.h:43:int register_fib_notifier(struct net *net, struct notifier_block *nb,
include/net/fib_notifier.h-44- void (*cb)(struct notifier_block *nb),
include/net/fib_notifier.h-45- struct netlink_ext_ack *extack);
include/net/fib_notifier.h:46:int unregister_fib_notifier(struct net *net, struct notifier_block *nb);
include/net/fib_notifier.h-47-struct fib_notifier_ops *
--
net/core/fib_notifier.c=82=static bool fib_dump_is_consistent(struct net *net, struct notifier_block *nb,
--
net/core/fib_notifier.c-97-#define FIB_DUMP_MAX_RETRIES 5
net/core/fib_notifier.c:98:int register_fib_notifier(struct net *net, struct notifier_block *nb,
net/core/fib_notifier.c-99- void (*cb)(struct notifier_block *nb),
--
net/core/fib_notifier.c-117-}
net/core/fib_notifier.c:118:EXPORT_SYMBOL(register_fib_notifier);
net/core/fib_notifier.c-119-
net/core/fib_notifier.c:120:int unregister_fib_notifier(struct net *net, struct notifier_block *nb)
net/core/fib_notifier.c-121-{
--
net/core/fib_notifier.c-125-}
net/core/fib_notifier.c:126:EXPORT_SYMBOL(unregister_fib_notifier);
net/core/fib_notifier.c-127-
]
|
| 50/2 |
2026/05/26 14:39 |
llm |
expert |
0m
Model:
Tokens:
input: 29594
output: 53
thoughts: 73
Thoughts: Okay, let's see, my mental gears are turning...
**Analyzing `register_fib_notifier` Usage and Potential Reference Issues**
Right, so `register_fib_notifier` is a key function, and the code mentions it's being utilized by several hardware drivers: `prestera`, `mlx5`, `mlxsw`, `rocker`, and `netdevsim`. That's a decent spread, indicating it's handling some crucial routing or forwarding notifications. Now, the critical question pops up: during a dump operation, are any of *those* drivers, particularly when handling a `fib_info` object, holding a reference to that object while also dumping it? This becomes a race condition if not handled well. I know the potential pitfalls.
I should check the likely suspects, so `mlxsw_sp_router_fib_event` needs to be analyzed to see how it handles `fib_info` objects, and whether any of them holds a reference to it during a dump. This is the likely point of a subtle problem.
|
| 51/2 |
2026/05/26 14:39 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c Name:mlxsw_sp_router_fib_event]
Results: map[SourceCode: SourceFile:]
|
| 52/2 |
2026/05/26 14:39 |
llm |
expert |
0m
Model:
Tokens:
input: 29962
output: 24
thoughts: 38
Thoughts: Okay, here's my take on those thoughts, framed as an expert's internal monologue:
**Diving into the FIB Event Handling**
Alright, time to get my hands dirty with `mlxsw_sp_router_fib_event`. That sounds like it's where the magic happens for FIB events within the Spectrum router driver, specifically in `drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c`. It's my hunch that this is where the core logic for reacting to routing table updates β the additions, deletions, and modifications of Forwarding Information Base (FIB) entries β resides. I need to understand how the driver processes these events. My main focus would be how the FIB is maintained, and I'll need to know:
* How the driver interacts with the hardware, making any necessary changes to forwarding tables on the switch itself.
* The event source: what triggers this, is it from the kernel or some userland component?
* Error handling - how the events are handled if something goes wrong.
The event likely tells the driver what changed in the routing table, and this function must translate that information into the switch's internal data structures and hardware commands. It's likely a critical piece for ensuring consistent and correct forwarding behavior. Let's see how they've implemented it; I'm expecting some clever use of the Mellanox SDK, hopefully, and a well-defined state machine to handle the different FIB operations. I'm keen to see if they've accounted for all the nuances, especially around concurrency and potential race conditions when updating the routing state in a live environment.
|
| 53/2 |
2026/05/26 14:39 |
tool |
grepper |
0m
Args: map[Expression:mlxsw_sp_router_fib_event]
Results: map[Output:drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c=8058=static int mlxsw_sp_router_fib_rule_event(unsigned long event,
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-8103-/* Called with rcu_read_lock() */
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c:8104:static int mlxsw_sp_router_fib_event(struct notifier_block *nb,
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-8105- unsigned long event, void *ptr)
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c=11575=int mlxsw_sp_router_init(struct mlxsw_sp *mlxsw_sp,
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11694-
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c:11695: mlxsw_sp->router->fib_nb.notifier_call = mlxsw_sp_router_fib_event;
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11696- err = register_fib_notifier(mlxsw_sp_net(mlxsw_sp),
]
|
| 54/2 |
2026/05/26 14:39 |
llm |
expert |
0m
Model:
Tokens:
input: 30735
output: 53
thoughts: 16
Thoughts: Okay, let's dive into this `mlxsw_sp_router_fib_event` function.
**Analyzing `mlxsw_sp_router_fib_event`**
Right, so this function is clearly a core component dealing with FIB (Forwarding Information Base) events within the mlxsw_sp driver. Given the naming, it's likely responsible for handling notifications or updates related to route changes. I'd expect to see logic for reacting to events such as route additions, deletions, or modifications within the switch's FIB tables. My initial thought is it probably interacts heavily with the switch hardware directly and might manage data structures representing the FIB state. I will need to understand what this module is built on, what hardware is being manipulated, and which tables are managed in the hardware.
I'm keen to see how the events are triggered and what information is passed along with them. Is it a simple notification, or does it carry a full route description? Are there different event types? I'll be looking for any error handling, too - a robust driver needs to gracefully deal with unexpected hardware states or communication failures, which should be very important.
The real trick is how it interacts with the broader networking stack. I'll need to observe how this ties into the kernel's routing tables and whether it handles things like route aging, preference, and ECMP (Equal-Cost Multi-Path) properly. I'll expect to see optimizations based on the hardware capabilities. I will expect there to be a connection with the switch SDK, but this function should provide a more general layer. I'll be looking for concurrency considerations as well, for example if this function needs to be thread-safe when dealing with FIB updates.
I'm particularly interested in seeing the design decisions made around the event handling mechanism itself. Is it event-driven, polling, or something else? And how it deals with asynchronous operations? Does it use workqueues, threads, or some other method? This could have considerable implications for performance and responsiveness.
Let's get reading and see how it all works!
|
| 55/2 |
2026/05/26 14:39 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c Name:mlxsw_sp_router_fib_event]
Results: map[SourceCode: SourceFile:]
|
| 56/2 |
2026/05/26 14:39 |
llm |
expert |
0m
Model:
Tokens:
input: 31260
output: 48
thoughts: 15
Thoughts: Okay, here's my summary, tailored for an expert audience, as if it's my own thought process:
**Static Function, File Reading, Aha!**
Alright, hold on a second. This is a static function, which immediately triggers a specific mental checklist. Static functions, within this context, often imply utility-level operations, things that don't need or can't access instance-level data. The implications of this are important! If it's *static*, my assumptions about its behavior are going to be affected.
The critical insight I'm getting is that the presence of the static modifier immediately influences the kinds of tasks that are feasible. Because this is static, it's very probable that the core logic is going to revolve around reading a file, a file read operation is a very reasonable task for this type of function. A static function reading a file opens up a few avenues of investigation that I will immediately follow.
|
| 57/2 |
2026/05/26 14:39 |
tool |
read-file |
0m
Args: map[File:drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c FirstLine:8103 LineCount:50]
Results: map[Contents:8103: /* Called with rcu_read_lock() */
8104: static int mlxsw_sp_router_fib_event(struct notifier_block *nb,
8105: unsigned long event, void *ptr)
8106: {
8107: struct mlxsw_sp_fib_event_work *fib_work;
8108: struct fib_notifier_info *info = ptr;
8109: struct mlxsw_sp_router *router;
8110: int err;
8111:
8112: if ((info->family != AF_INET && info->family != AF_INET6 &&
8113: info->family != RTNL_FAMILY_IPMR &&
8114: info->family != RTNL_FAMILY_IP6MR))
8115: return NOTIFY_DONE;
8116:
8117: router = container_of(nb, struct mlxsw_sp_router, fib_nb);
8118:
8119: switch (event) {
8120: case FIB_EVENT_RULE_ADD:
8121: case FIB_EVENT_RULE_DEL:
8122: err = mlxsw_sp_router_fib_rule_event(event, info,
8123: router->mlxsw_sp);
8124: return notifier_from_errno(err);
8125: case FIB_EVENT_ENTRY_ADD:
8126: case FIB_EVENT_ENTRY_REPLACE:
8127: case FIB_EVENT_ENTRY_APPEND:
8128: if (info->family == AF_INET) {
8129: struct fib_entry_notifier_info *fen_info = ptr;
8130:
8131: if (fen_info->fi->fib_nh_is_v6) {
8132: NL_SET_ERR_MSG_MOD(info->extack, "IPv6 gateway with IPv4 route is not supported");
8133: return notifier_from_errno(-EINVAL);
8134: }
8135: }
8136: break;
8137: }
8138:
8139: fib_work = kzalloc_obj(*fib_work, GFP_ATOMIC);
8140: if (!fib_work)
8141: return NOTIFY_BAD;
8142:
8143: fib_work->mlxsw_sp = router->mlxsw_sp;
8144: fib_work->event = event;
8145:
8146: switch (info->family) {
8147: case AF_INET:
8148: INIT_WORK(&fib_work->work, mlxsw_sp_router_fib4_event_work);
8149: mlxsw_sp_router_fib4_event(fib_work, info);
8150: break;
8151: case AF_INET6:
8152: INIT_WORK(&fib_work->work, mlxsw_sp_router_fib6_event_work);
]
|
| 58/2 |
2026/05/26 14:39 |
llm |
expert |
0m
Model:
Tokens:
input: 32355
output: 54
thoughts: 17
Thoughts: Okay, here's my interpretation of what's going through my mind, rendered in the first person, as if I'm the one thinking it, and targeted to an expert audience like myself:
**Diving into `mlxsw_sp_router_fib4_event`**
Alright, let's dissect `mlxsw_sp_router_fib4_event`. I'm guessing this function is central to handling FIB (Forwarding Information Base) updates, specifically for IPv4 routes, within the Mellanox switch platform (mlxsw). The naming convention makes that pretty clear. I immediately need to understand the *trigger* for this event. Is it triggered by hardware events from the ASIC itself? Is it triggered by control plane interaction, maybe from a routing protocol like BGP or OSPF? Or is it triggered by kernel-level changes to the routing table?
Knowing the trigger dictates how I approach this. If it's hardware-driven, the emphasis will be on real-time responsiveness and potential race conditions. If it's control-plane driven, I'll need to think about the synchronization mechanisms required. I'd be looking for the event's inputs to determine what data is passed.
I also need to understand the *scope* of this event. Is it a single route update, or is it a batch process? If it's a batch, what's the batching strategy? Efficiency is critical here, especially in high-scale environments. My immediate concern is, what happens if this process fails? How are errors handled? Can we gracefully fall back, or is it a hard failure that brings the forwarding plane down?
I'm thinking about the data structures this function manipulates. How is the FIB structured within the mlxsw driver? Is it a simple tree, a hash table, or something more sophisticated optimized for this particular ASIC's capabilities? I need to understand the update process itself: Does it involve programming the ASIC directly, or is there an intermediate representation? What are the implications for things like route advertisement and route redistribution? I'll also be mentally mapping the function's interactions with other modules. I'll need to look at how this event impacts the overall datapath functionality.
Finally, I'll be looking for any hints about performance optimization. Are there any caching strategies employed? What about locking mechanisms? This is the core of how the hardware works with the software, so it's critical to avoid bottlenecks.
|
| 59/2 |
2026/05/26 14:39 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c Name:mlxsw_sp_router_fib4_event]
Results: map[SourceCode: SourceFile:]
|
| 60/2 |
2026/05/26 14:39 |
llm |
expert |
0m
Model:
Tokens:
input: 32960
output: 48
thoughts: 7
Thoughts: Okay, I'm ready to dive into this. Here's how I'd approach summarizing the "thoughts" scenario, keeping in mind I'm speaking as an expert:
**Reading the File: A Quick Assessment**
Alright, the instruction says "Let's read the file." Immediately, my mind starts assessing the context. What kind of file are we talking about? Is it a script, code, data, a research paper, a document? The nature of the file dictates my approach. If it's code, I'll be looking for structure, logic, and potential areas for analysis. Data? I'm focusing on the format, potential biases, and how I can extract meaningful insights. Document? I'll be looking for the central argument or theme.
This first step is crucial. "Let's read the file" isn't just an instruction; it's a launchpad for a focused analysis. I need to quickly identify the file's type and purpose. With that understanding, I can begin to unpack the specific information, arguments, or data points presented.
|
| 61/2 |
2026/05/26 14:40 |
tool |
read-file |
0m
Args: map[File:drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c FirstLine:8020 LineCount:50]
Results: map[Contents:8020: int err;
8021:
8022: switch (fib_work->event) {
8023: case FIB_EVENT_ENTRY_REPLACE:
8024: case FIB_EVENT_ENTRY_APPEND:
8025: case FIB_EVENT_ENTRY_DEL:
8026: fen6_info = container_of(info, struct fib6_entry_notifier_info,
8027: info);
8028: err = mlxsw_sp_router_fib6_work_init(&fib_work->fib6_work,
8029: fen6_info);
8030: if (err)
8031: return err;
8032: break;
8033: }
8034:
8035: return 0;
8036: }
8037:
8038: static void
8039: mlxsw_sp_router_fibmr_event(struct mlxsw_sp_fib_event_work *fib_work,
8040: struct fib_notifier_info *info)
8041: {
8042: switch (fib_work->event) {
8043: case FIB_EVENT_ENTRY_REPLACE:
8044: case FIB_EVENT_ENTRY_ADD:
8045: case FIB_EVENT_ENTRY_DEL:
8046: memcpy(&fib_work->men_info, info, sizeof(fib_work->men_info));
8047: mr_cache_hold(fib_work->men_info.mfc);
8048: break;
8049: case FIB_EVENT_VIF_ADD:
8050: case FIB_EVENT_VIF_DEL:
8051: memcpy(&fib_work->ven_info, info, sizeof(fib_work->ven_info));
8052: netdev_hold(fib_work->ven_info.dev, &fib_work->dev_tracker,
8053: GFP_ATOMIC);
8054: break;
8055: }
8056: }
8057:
8058: static int mlxsw_sp_router_fib_rule_event(unsigned long event,
8059: struct fib_notifier_info *info,
8060: struct mlxsw_sp *mlxsw_sp)
8061: {
8062: struct netlink_ext_ack *extack = info->extack;
8063: struct fib_rule_notifier_info *fr_info;
8064: struct fib_rule *rule;
8065: int err = 0;
8066:
8067: /* nothing to do at the moment */
8068: if (event == FIB_EVENT_RULE_DEL)
8069: return 0;
]
|
| 62/2 |
2026/05/26 14:40 |
llm |
expert |
0m
Model:
Tokens:
input: 34022
output: 48
thoughts: 27
Thoughts: Okay, let's dive into this. Here's my take on those lines of code...
**Examining `mlxsw_sp_router_fib4_event`**
Right, so I'm focusing in on the `mlxsw_sp_router_fib4_event` function, around line 7900. I need to understand how the routing table handles events, particularly those related to IPv4. This function is likely a key component in that process. I need to see how it responds to changes in the Forwarding Information Base (FIB). Is it reacting to route additions, deletions, or modifications? What's the control flow like within the function? Are there any potential bottlenecks or areas that might require optimization? Time to dig in and figure out exactly what's happening at this crucial part of the routing stack. Let's see how this event handling is architected.
|
| 63/2 |
2026/05/26 14:40 |
tool |
read-file |
0m
Args: map[File:drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c FirstLine:7980 LineCount:50]
Results: map[Contents:7980: &fib_work->ven_info);
7981: netdev_put(fib_work->ven_info.dev, &fib_work->dev_tracker);
7982: break;
7983: }
7984: mutex_unlock(&mlxsw_sp->router->lock);
7985: rtnl_unlock();
7986: kfree(fib_work);
7987: }
7988:
7989: static void mlxsw_sp_router_fib4_event(struct mlxsw_sp_fib_event_work *fib_work,
7990: struct fib_notifier_info *info)
7991: {
7992: struct fib_entry_notifier_info *fen_info;
7993: struct fib_nh_notifier_info *fnh_info;
7994:
7995: switch (fib_work->event) {
7996: case FIB_EVENT_ENTRY_REPLACE:
7997: case FIB_EVENT_ENTRY_DEL:
7998: fen_info = container_of(info, struct fib_entry_notifier_info,
7999: info);
8000: fib_work->fen_info = *fen_info;
8001: /* Take reference on fib_info to prevent it from being
8002: * freed while work is queued. Release it afterwards.
8003: */
8004: fib_info_hold(fib_work->fen_info.fi);
8005: break;
8006: case FIB_EVENT_NH_ADD:
8007: case FIB_EVENT_NH_DEL:
8008: fnh_info = container_of(info, struct fib_nh_notifier_info,
8009: info);
8010: fib_work->fnh_info = *fnh_info;
8011: fib_info_hold(fib_work->fnh_info.fib_nh->nh_parent);
8012: break;
8013: }
8014: }
8015:
8016: static int mlxsw_sp_router_fib6_event(struct mlxsw_sp_fib_event_work *fib_work,
8017: struct fib_notifier_info *info)
8018: {
8019: struct fib6_entry_notifier_info *fen6_info;
8020: int err;
8021:
8022: switch (fib_work->event) {
8023: case FIB_EVENT_ENTRY_REPLACE:
8024: case FIB_EVENT_ENTRY_APPEND:
8025: case FIB_EVENT_ENTRY_DEL:
8026: fen6_info = container_of(info, struct fib6_entry_notifier_info,
8027: info);
8028: err = mlxsw_sp_router_fib6_work_init(&fib_work->fib6_work,
8029: fen6_info);
]
|
| 64/2 |
2026/05/26 14:40 |
llm |
expert |
0m
Model:
Tokens:
input: 35121
output: 22
thoughts: 63
Thoughts: **Realizing a Parallel Bug in the mlxsw Driver**
Okay, so I've been digging into this, and it looks like `mlxsw_sp_router_fib4_event` is doing the same thing as the problematic code I just identified β it's also calling `fib_info_hold(fib_work->fen_info.fi)`. This is a red flag! It means the `mlxsw` driver might have the exact same vulnerability as what I'm currently investigating. This is good, in a bad sort of way; at least the issue isn't isolated. Now, let's pivot and take a look at `prestera_router_fib_event`. I'm hoping it's safe, but I need to be sure.
|
| 65/2 |
2026/05/26 14:40 |
tool |
grepper |
0m
Args: map[Expression:prestera_router_fib_event]
Results: map[Output:drivers/net/ethernet/marvell/prestera/prestera_router.c=1384=struct prestera_fib_event_work {
--
drivers/net/ethernet/marvell/prestera/prestera_router.c-1390-
drivers/net/ethernet/marvell/prestera/prestera_router.c:1391:static void __prestera_router_fib_event_work(struct work_struct *work)
drivers/net/ethernet/marvell/prestera/prestera_router.c-1392-{
--
drivers/net/ethernet/marvell/prestera/prestera_router.c-1429-/* Called with rcu_read_lock() */
drivers/net/ethernet/marvell/prestera/prestera_router.c:1430:static int __prestera_router_fib_event(struct notifier_block *nb,
drivers/net/ethernet/marvell/prestera/prestera_router.c-1431- unsigned long event, void *ptr)
--
drivers/net/ethernet/marvell/prestera/prestera_router.c-1458- fib_work->sw = router->sw;
drivers/net/ethernet/marvell/prestera/prestera_router.c:1459: INIT_WORK(&fib_work->work, __prestera_router_fib_event_work);
drivers/net/ethernet/marvell/prestera/prestera_router.c-1460- prestera_queue_work(&fib_work->work);
--
drivers/net/ethernet/marvell/prestera/prestera_router.c=1548=int prestera_router_init(struct prestera_switch *sw)
--
drivers/net/ethernet/marvell/prestera/prestera_router.c-1599-
drivers/net/ethernet/marvell/prestera/prestera_router.c:1600: router->fib_nb.notifier_call = __prestera_router_fib_event;
drivers/net/ethernet/marvell/prestera/prestera_router.c-1601- err = register_fib_notifier(&init_net, &router->fib_nb,
]
|
| 66/2 |
2026/05/26 14:40 |
llm |
expert |
0m
Model:
Tokens:
input: 35932
output: 48
thoughts: 15
Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience, and written in the first person:
**Initial Assessment: Navigating the FIB Event**
Alright, let's dive into this. The mental hook here is `__prestera_router_fib_event`. My initial thought is to immediately situate this within the PrestaSwitch ecosystem. We're talking about a core event related to the Forwarding Information Base (FIB), which means we're dealing with *critical* path forwarding information and likely, very low-level operations. I know the PrestaSwitch architecture, I'm expecting this is going to be intricately tied to hardware tables, TCAMs, and potentially control plane interaction, or at least the translation of control plane directives to these lower-level constructs. Understanding this event will likely be critical for debugging forwarding issues, performance bottlenecks, or even security vulnerabilities related to the data plane. It also likely means we'll be dealing with specific ASIC (Application-Specific Integrated Circuit) interactions and driver layers. I'm prepared for complex data structures and the need to follow a tangled web of function calls. So, let's get started.
|
| 67/2 |
2026/05/26 14:40 |
tool |
read-file |
0m
Args: map[File:drivers/net/ethernet/marvell/prestera/prestera_router.c FirstLine:1430 LineCount:50]
Results: map[Contents:1430: static int __prestera_router_fib_event(struct notifier_block *nb,
1431: unsigned long event, void *ptr)
1432: {
1433: struct prestera_fib_event_work *fib_work;
1434: struct fib_entry_notifier_info *fen_info;
1435: struct fib_notifier_info *info = ptr;
1436: struct prestera_router *router;
1437:
1438: if (info->family != AF_INET)
1439: return NOTIFY_DONE;
1440:
1441: router = container_of(nb, struct prestera_router, fib_nb);
1442:
1443: switch (event) {
1444: case FIB_EVENT_ENTRY_REPLACE:
1445: case FIB_EVENT_ENTRY_DEL:
1446: fen_info = container_of(info, struct fib_entry_notifier_info,
1447: info);
1448: if (!fen_info->fi)
1449: return NOTIFY_DONE;
1450:
1451: fib_work = kzalloc_obj(*fib_work, GFP_ATOMIC);
1452: if (WARN_ON(!fib_work))
1453: return NOTIFY_BAD;
1454:
1455: fib_info_hold(fen_info->fi);
1456: fib_work->fen_info = *fen_info;
1457: fib_work->event = event;
1458: fib_work->sw = router->sw;
1459: INIT_WORK(&fib_work->work, __prestera_router_fib_event_work);
1460: prestera_queue_work(&fib_work->work);
1461: break;
1462: default:
1463: return NOTIFY_DONE;
1464: }
1465:
1466: return NOTIFY_DONE;
1467: }
1468:
1469: struct prestera_netevent_work {
1470: struct work_struct work;
1471: struct prestera_switch *sw;
1472: struct neighbour *n;
1473: };
1474:
1475: static void prestera_router_neigh_event_work(struct work_struct *work)
1476: {
1477: struct prestera_netevent_work *net_work =
1478: container_of(work, struct prestera_netevent_work, work);
1479: struct prestera_switch *sw = net_work->sw;
]
|
| 68/2 |
2026/05/26 14:40 |
llm |
expert |
0m
Model:
Tokens:
input: 36995
output: 40
thoughts: 333
Thoughts: **Analyzing the Fib_Info Hold Bug**
Okay, so I've confirmed that the `prestera` driver also calls `fib_info_hold(fen_info->fi)`. That means this isn't just a `netdevsim` issue; it's a systemic problem impacting multiple hardware drivers. Let's see... `fib_info_hold` is used by `prestera`, `mlxsw`, and `netdevsim`. The critical question now is, can we trigger this bug *without* needing `netdevsim`?
Aha! Yes, it looks like it's possible. If either `prestera` or `mlxsw` is loaded, a race condition could exist where a route is deleted concurrently with either `register_fib_notifier` or during normal FIB events. Hold on... `register_fib_notifier` dumps the entire FIB table, doesn't it? And both `prestera` and `mlxsw` register their FIB notifiers during initialization. If a route deletion happens concurrently with driver initialization, *that* could be the trigger!
But initialization is typically during boot or module load. Hmm, can an unprivileged user even trigger driver initialization? Generally, loading a driver requires root privileges. Wait a second... what about `fib_net_dump`? That's called when a new FIB notifier is registered. Are there *other* ways to trigger `fib_net_dump`? Nope, it looks like `fib_net_dump` is exclusively called from `register_fib_notifier`.
But wait! I'm forgetting something crucial. `mlxsw_sp_router_fib_event` is called for *every* FIB event, not just during the dump! Now, when a route is deleted, `fib_table_delete` is called, and that is a key event! I'm going to have to re-examine the `fib_table_delete` function again.
|
| 69/2 |
2026/05/26 14:40 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/ipv4/fib_trie.c Name:fib_table_delete]
Results: map[SourceCode:1687: /* Caller must hold RTNL. */
1688: int fib_table_delete(struct net *net, struct fib_table *tb,
1689: struct fib_config *cfg, struct netlink_ext_ack *extack)
1690: {
1691: struct trie *t = (struct trie *) tb->tb_data;
1692: struct fib_alias *fa, *fa_to_delete;
1693: struct key_vector *l, *tp;
1694: u8 plen = cfg->fc_dst_len;
1695: u8 slen = KEYLENGTH - plen;
1696: dscp_t dscp;
1697: u32 key;
1698:
1699: key = ntohl(cfg->fc_dst);
1700:
1701: l = fib_find_node(t, &tp, key);
1702: if (!l)
1703: return -ESRCH;
1704:
1705: dscp = cfg->fc_dscp;
1706: fa = fib_find_alias(&l->leaf, slen, dscp, 0, tb->tb_id, false);
1707: if (!fa)
1708: return -ESRCH;
1709:
1710: pr_debug("Deleting %08x/%d dsfield=0x%02x t=%p\n", key, plen,
1711: inet_dscp_to_dsfield(dscp), t);
1712:
1713: fa_to_delete = NULL;
1714: hlist_for_each_entry_from(fa, fa_list) {
1715: struct fib_info *fi = fa->fa_info;
1716:
1717: if ((fa->fa_slen != slen) ||
1718: (fa->tb_id != tb->tb_id) ||
1719: (fa->fa_dscp != dscp))
1720: break;
1721:
1722: if ((!cfg->fc_type || fa->fa_type == cfg->fc_type) &&
1723: (cfg->fc_scope == RT_SCOPE_NOWHERE ||
1724: fa->fa_info->fib_scope == cfg->fc_scope) &&
1725: (!cfg->fc_prefsrc ||
1726: fi->fib_prefsrc == cfg->fc_prefsrc) &&
1727: (!cfg->fc_protocol ||
1728: fi->fib_protocol == cfg->fc_protocol) &&
1729: fib_nh_match(net, cfg, fi, extack) == 0 &&
1730: fib_metrics_match(cfg, fi)) {
1731: fa_to_delete = fa;
1732: break;
1733: }
1734: }
1735:
1736: if (!fa_to_delete)
1737: return -ESRCH;
1738:
1739: fib_notify_alias_delete(net, key, &l->leaf, fa_to_delete, extack);
1740: rtmsg_fib(RTM_DELROUTE, htonl(key), fa_to_delete, plen, tb->tb_id,
1741: &cfg->fc_nlinfo, 0);
1742:
1743: if (!plen)
1744: tb->tb_num_default--;
1745:
1746: fib_remove_alias(t, tp, l, fa_to_delete);
1747:
1748: if (READ_ONCE(fa_to_delete->fa_state) & FA_S_ACCESSED)
1749: rt_cache_flush(cfg->fc_nlinfo.nl_net);
1750:
1751: fib_release_info(fa_to_delete->fa_info);
1752: alias_free_mem_rcu(fa_to_delete);
1753: return 0;
1754: }
SourceFile:net/ipv4/fib_trie.c]
|
| 70/2 |
2026/05/26 14:40 |
llm |
expert |
0m
Model:
Tokens:
input: 38960
output: 42
thoughts: 51
Thoughts: Okay, here's my take on those thoughts, framed as my own:
**Analyzing the Fib Table Deletion Sequence**
Right, so in this `fib_table_delete` function, I'm seeing something that immediately jumps out at me: `fib_notify_alias_delete` is being called *before* `fib_remove_alias` and `fib_release_info`. That's a bit unusual, and warrants a closer look. My immediate thought is, "Okay, why this order? What's the dependency here?" I need to understand what `fib_notify_alias_delete` is actually *doing* to properly assess this sequence. I need to understand if the information it works with is still valid after the other two functions run, so I'm going to dive right into that `fib_notify_alias_delete` function next. Let's see what's happening there.
|
| 71/2 |
2026/05/26 14:40 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/ipv4/fib_trie.c Name:fib_notify_alias_delete]
Results: map[SourceCode:1657: static void fib_notify_alias_delete(struct net *net, u32 key,
1658: struct hlist_head *fah,
1659: struct fib_alias *fa_to_delete,
1660: struct netlink_ext_ack *extack)
1661: {
1662: struct fib_alias *fa_next, *fa_to_notify;
1663: u32 tb_id = fa_to_delete->tb_id;
1664: u8 slen = fa_to_delete->fa_slen;
1665: enum fib_event_type fib_event;
1666:
1667: /* Do not notify if we do not care about the route. */
1668: if (fib_find_alias(fah, slen, 0, 0, tb_id, true) != fa_to_delete)
1669: return;
1670:
1671: /* Determine if the route should be replaced by the next route in the
1672: * list.
1673: */
1674: fa_next = hlist_entry_safe(fa_to_delete->fa_list.next,
1675: struct fib_alias, fa_list);
1676: if (fa_next && fa_next->fa_slen == slen && fa_next->tb_id == tb_id) {
1677: fib_event = FIB_EVENT_ENTRY_REPLACE;
1678: fa_to_notify = fa_next;
1679: } else {
1680: fib_event = FIB_EVENT_ENTRY_DEL;
1681: fa_to_notify = fa_to_delete;
1682: }
1683: call_fib_entry_notifiers(net, fib_event, key, KEYLENGTH - slen,
1684: fa_to_notify, extack);
1685: }
SourceFile:net/ipv4/fib_trie.c]
|
| 72/2 |
2026/05/26 14:40 |
llm |
expert |
0m
Model:
Tokens:
input: 39810
output: 41
thoughts: 474
Thoughts: **Analyzing a Nasty Race Condition in the FIB Code**
Okay, here's what's rattling around in my head. I've been digging into this FIB issue, and it's starting to crystallize. The core problem appears to be a race condition, and it's a subtle one. It looks like the `fib_notify_alias_delete` function calls `call_fib_entry_notifiers` *before* `fib_release_info` is executed. In the normal course of route deletion, this isn't an issue because the `fib_info` structure is still valid and its reference count isn't zero.
However, the bug *only* manifests during `fib_net_dump`. I'm starting to see it clearly, because `fib_net_dump` iterates through the trie structure under an `rcu_read_lock()`. Here's the kicker: if a route is concurrently deleted while the dump is running, `fib_table_delete` calls `fib_remove_alias` and *then* `fib_release_info`. This means that `fib_net_dump` might still observe the alias in the trie (because `hlist_del_rcu` has been used), and attempt to process it by calling `fib_leaf_notify`, which in turn calls `call_fib_entry_notifier`, and that leads to `nsim_fib4_prepare_event` and finally a `fib_info_hold`. Problem is, at this point, the `fib_info` structure's reference count is already zero!
So, to summarize, this bug is a race between `fib_net_dump` and `fib_table_delete`. We have a window for the race because `fib_net_dump` is using an RCU lock, meaning the memory is not immediately freed after the delete.
To reliably trigger this, we need to orchestrate two events: first, we need to trigger `fib_net_dump`, and concurrently, we need a route to be deleted. The challenge now is understanding how to trigger `fib_net_dump`. Well, it's called when a new FIB notifier is registered. I need to look at who is registering these notifiers. I've got a list here: `prestera_router_init`, `mlx5e_tc_tun_init`, `mlx5_lag_mp_init`, `mlxsw_sp_router_init`, `rocker_probe`, and `nsim_fib_create`.
Ah, these are all driver initialization functions. And `nsim_fib_create` is called during `nsim_dev_reload_up`. The next question is, can an unprivileged user initiate a `devlink_reload`? If so, this might be exploitable. Let's check `net/devlink/netlink.c` and `net/devlink/dev.c` to see if `devlink_reload` is restricted to `CAP_NET_ADMIN`. Knowing what I know, it's likely that devlink commands are indeed protected by that capability, but it's worth a double-check.
|
| 73/2 |
2026/05/26 14:40 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:net/devlink/netlink.c Name:devlink_nl_cmd_reload]
Results: map[SourceCode: SourceFile:]
|
| 74/2 |
2026/05/26 14:40 |
llm |
expert |
0m
Model:
Tokens:
input: 41022
output: 21
thoughts: 8
Thoughts: Okay, here's my attempt at summarizing the thought process, tailored for an expert audience, and written in the first person:
**Diving into devlink for a Systems Inspection**
Right, let's get down to brass tacks. I'm thinking, "Time to take a look at the `devlink` commands." Given my experience, I'm probably not just doing a casual overview. I'm going in with a specific purpose in mind, likely troubleshooting a performance issue, verifying a configuration, or perhaps even preparing for an update. I'm thinking, "What specific area am I targeting? Is it queues, ports, esw, region...?" The nuances of `devlink` are numerous, but the power is undeniable. I'll need to know which flags and options are particularly relevant to the target, whether it be to look at the health of a piece of hardware or to ensure all of its settings have the correct values. It's time to gather that context, get the right CLI command ready, and get to work.
|
| 75/2 |
2026/05/26 14:40 |
tool |
grepper |
0m
Args: map[Expression:DEVLINK_CMD_RELOAD]
Results: map[Output:Documentation/networking/devlink/hns3.rst=8=device driver.
Documentation/networking/devlink/hns3.rst-9-
Documentation/networking/devlink/hns3.rst:10:The ``hns3`` driver supports reloading via ``DEVLINK_CMD_RELOAD``.
Documentation/networking/devlink/hns3.rst-11-
--
Documentation/networking/devlink/ice.rst=234=The ``ice`` driver supports activating new firmware after a flash update
Documentation/networking/devlink/ice.rst:235:using ``DEVLINK_CMD_RELOAD`` with the ``DEVLINK_RELOAD_ACTION_FW_ACTIVATE``
Documentation/networking/devlink/ice.rst-236-action.
--
Documentation/networking/devlink/ixgbe.rst=109=The ``ixgbe`` driver supports activating new firmware after a flash update
Documentation/networking/devlink/ixgbe.rst:110:using ``DEVLINK_CMD_RELOAD`` with the ``DEVLINK_RELOAD_ACTION_FW_ACTIVATE``
Documentation/networking/devlink/ixgbe.rst-111-action.
--
Documentation/networking/devlink/mlx4.rst=25=parameters.
--
Documentation/networking/devlink/mlx4.rst-42-
Documentation/networking/devlink/mlx4.rst:43:The ``mlx4`` driver supports reloading via ``DEVLINK_CMD_RELOAD``
Documentation/networking/devlink/mlx4.rst-44-
--
Documentation/networking/devlink/mlx5.rst=76=parameters.
--
Documentation/networking/devlink/mlx5.rst-234-
Documentation/networking/devlink/mlx5.rst:235:The ``mlx5`` driver supports reloading via ``DEVLINK_CMD_RELOAD``
Documentation/networking/devlink/mlx5.rst-236-
--
Documentation/networking/devlink/mlxsw.rst=21=parameters.
--
Documentation/networking/devlink/mlxsw.rst-37-
Documentation/networking/devlink/mlxsw.rst:38:The ``mlxsw`` driver supports reloading via ``DEVLINK_CMD_RELOAD``
Documentation/networking/devlink/mlxsw.rst-39-
--
Documentation/networking/devlink/netdevsim.rst=21=parameters.
--
Documentation/networking/devlink/netdevsim.rst-35-
Documentation/networking/devlink/netdevsim.rst:36:The ``netdevsim`` driver supports reloading via ``DEVLINK_CMD_RELOAD``
Documentation/networking/devlink/netdevsim.rst-37-
--
drivers/net/ethernet/intel/ice/devlink/devlink.c=271=static int ice_devlink_info_get(struct devlink *devlink,
--
drivers/net/ethernet/intel/ice/devlink/devlink.c-402- * issuing device specific EMP reset. Called in response to
drivers/net/ethernet/intel/ice/devlink/devlink.c:403: * a DEVLINK_CMD_RELOAD with the DEVLINK_RELOAD_ACTION_FW_ACTIVATE.
drivers/net/ethernet/intel/ice/devlink/devlink.c-404- *
--
drivers/net/ethernet/intel/ixgbe/devlink/devlink.c=312=static int ixgbe_devlink_info_get(struct devlink *devlink,
--
drivers/net/ethernet/intel/ixgbe/devlink/devlink.c-380- * issuing device specific EMP reset. Called in response to
drivers/net/ethernet/intel/ixgbe/devlink/devlink.c:381: * a DEVLINK_CMD_RELOAD with the DEVLINK_RELOAD_ACTION_FW_ACTIVATE.
drivers/net/ethernet/intel/ixgbe/devlink/devlink.c-382- *
--
include/uapi/linux/devlink.h=24=enum devlink_command {
--
include/uapi/linux/devlink.h-82- */
include/uapi/linux/devlink.h:83: DEVLINK_CMD_RELOAD,
include/uapi/linux/devlink.h-84-
--
net/devlink/dev.c=520=int devlink_nl_reload_doit(struct sk_buff *skb, struct genl_info *info)
--
net/devlink/dev.c-596- return devlink_nl_reload_actions_performed_snd(devlink, actions_performed,
net/devlink/dev.c:597: DEVLINK_CMD_RELOAD, info);
net/devlink/dev.c-598-}
--
net/devlink/netlink_gen.c=316=static const struct nla_policy devlink_resource_dump_dump_nl_policy[DEVLINK_ATTR_RESOURCE_SCOPE_MASK + 1] = {
--
net/devlink/netlink_gen.c-322-
net/devlink/netlink_gen.c:323:/* DEVLINK_CMD_RELOAD - do */
net/devlink/netlink_gen.c-324-static const struct nla_policy devlink_reload_nl_policy[DEVLINK_ATTR_INDEX + 1] = {
--
net/devlink/netlink_gen.c=692=const struct genl_split_ops devlink_nl_ops[75] = {
--
net/devlink/netlink_gen.c-981- {
net/devlink/netlink_gen.c:982: .cmd = DEVLINK_CMD_RELOAD,
net/devlink/netlink_gen.c-983- .validate = GENL_DONT_VALIDATE_STRICT,
]
|
| 76/2 |
2026/05/26 14:40 |
llm |
expert |
0m
Model:
Tokens:
input: 42715
output: 41
thoughts: 25
Thoughts: Okay, let's dive into this. Here's my thought process on checking those flags:
**Analyzing `DEVLINK_CMD_RELOAD` in `devlink_nl_ops`**
Alright, time to dig into the `devlink_nl_ops` structure. My immediate focus is on the `DEVLINK_CMD_RELOAD` command. I need to understand its relationship to the flags within that structure. It's crucial to map how this specific command interacts with the existing framework. I'm thinking, "Okay, are there any potential conflicts? Are the existing flag configurations correctly handling the re-loading scenario?" I need to see if the defined flags are appropriately set, or if I need to adjust things to align with how this reload operation is expected to work. Iβm expecting that there will be certain expectations for the handling of that reload operation. It's a quick analysis and sanity check to make sure everything is in order and that there are no obvious omissions. No point going forward if the foundation is off.
|
| 77/2 |
2026/05/26 14:40 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/devlink/netlink_gen.c Name:devlink_nl_ops]
Results: map[SourceCode: 691: /* Ops table for devlink */
692: const struct genl_split_ops devlink_nl_ops[75] = {
693: {
694: .cmd = DEVLINK_CMD_GET,
695: .validate = GENL_DONT_VALIDATE_STRICT,
696: .pre_doit = devlink_nl_pre_doit,
697: .doit = devlink_nl_get_doit,
698: .post_doit = devlink_nl_post_doit,
699: .policy = devlink_get_nl_policy,
700: .maxattr = DEVLINK_ATTR_INDEX,
701: .flags = GENL_CMD_CAP_DO,
702: },
703: {
704: .cmd = DEVLINK_CMD_GET,
705: .validate = GENL_DONT_VALIDATE_DUMP,
706: .dumpit = devlink_nl_get_dumpit,
707: .flags = GENL_CMD_CAP_DUMP,
708: },
709: {
710: .cmd = DEVLINK_CMD_PORT_GET,
711: .validate = GENL_DONT_VALIDATE_STRICT,
712: .pre_doit = devlink_nl_pre_doit_port,
713: .doit = devlink_nl_port_get_doit,
714: .post_doit = devlink_nl_post_doit,
715: .policy = devlink_port_get_do_nl_policy,
716: .maxattr = DEVLINK_ATTR_INDEX,
717: .flags = GENL_CMD_CAP_DO,
718: },
719: {
720: .cmd = DEVLINK_CMD_PORT_GET,
721: .dumpit = devlink_nl_port_get_dumpit,
722: .policy = devlink_port_get_dump_nl_policy,
723: .maxattr = DEVLINK_ATTR_INDEX,
724: .flags = GENL_CMD_CAP_DUMP,
725: },
726: {
727: .cmd = DEVLINK_CMD_PORT_SET,
728: .validate = GENL_DONT_VALIDATE_STRICT,
729: .pre_doit = devlink_nl_pre_doit_port,
730: .doit = devlink_nl_port_set_doit,
731: .post_doit = devlink_nl_post_doit,
732: .policy = devlink_port_set_nl_policy,
733: .maxattr = DEVLINK_ATTR_INDEX,
734: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
735: },
736: {
737: .cmd = DEVLINK_CMD_PORT_NEW,
738: .validate = GENL_DONT_VALIDATE_STRICT,
739: .pre_doit = devlink_nl_pre_doit,
740: .doit = devlink_nl_port_new_doit,
741: .post_doit = devlink_nl_post_doit,
742: .policy = devlink_port_new_nl_policy,
743: .maxattr = DEVLINK_ATTR_INDEX,
744: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
745: },
746: {
747: .cmd = DEVLINK_CMD_PORT_DEL,
748: .validate = GENL_DONT_VALIDATE_STRICT,
749: .pre_doit = devlink_nl_pre_doit_port,
750: .doit = devlink_nl_port_del_doit,
751: .post_doit = devlink_nl_post_doit,
752: .policy = devlink_port_del_nl_policy,
753: .maxattr = DEVLINK_ATTR_INDEX,
754: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
755: },
756: {
757: .cmd = DEVLINK_CMD_PORT_SPLIT,
758: .validate = GENL_DONT_VALIDATE_STRICT,
759: .pre_doit = devlink_nl_pre_doit_port,
760: .doit = devlink_nl_port_split_doit,
761: .post_doit = devlink_nl_post_doit,
762: .policy = devlink_port_split_nl_policy,
763: .maxattr = DEVLINK_ATTR_INDEX,
764: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
765: },
766: {
767: .cmd = DEVLINK_CMD_PORT_UNSPLIT,
768: .validate = GENL_DONT_VALIDATE_STRICT,
769: .pre_doit = devlink_nl_pre_doit_port,
770: .doit = devlink_nl_port_unsplit_doit,
771: .post_doit = devlink_nl_post_doit,
772: .policy = devlink_port_unsplit_nl_policy,
773: .maxattr = DEVLINK_ATTR_INDEX,
774: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
775: },
776: {
777: .cmd = DEVLINK_CMD_SB_GET,
778: .validate = GENL_DONT_VALIDATE_STRICT,
779: .pre_doit = devlink_nl_pre_doit,
780: .doit = devlink_nl_sb_get_doit,
781: .post_doit = devlink_nl_post_doit,
782: .policy = devlink_sb_get_do_nl_policy,
783: .maxattr = DEVLINK_ATTR_INDEX,
784: .flags = GENL_CMD_CAP_DO,
785: },
786: {
787: .cmd = DEVLINK_CMD_SB_GET,
788: .dumpit = devlink_nl_sb_get_dumpit,
789: .policy = devlink_sb_get_dump_nl_policy,
790: .maxattr = DEVLINK_ATTR_INDEX,
791: .flags = GENL_CMD_CAP_DUMP,
792: },
793: {
794: .cmd = DEVLINK_CMD_SB_POOL_GET,
795: .validate = GENL_DONT_VALIDATE_STRICT,
796: .pre_doit = devlink_nl_pre_doit,
797: .doit = devlink_nl_sb_pool_get_doit,
798: .post_doit = devlink_nl_post_doit,
799: .policy = devlink_sb_pool_get_do_nl_policy,
800: .maxattr = DEVLINK_ATTR_INDEX,
801: .flags = GENL_CMD_CAP_DO,
802: },
803: {
804: .cmd = DEVLINK_CMD_SB_POOL_GET,
805: .dumpit = devlink_nl_sb_pool_get_dumpit,
806: .policy = devlink_sb_pool_get_dump_nl_policy,
807: .maxattr = DEVLINK_ATTR_INDEX,
808: .flags = GENL_CMD_CAP_DUMP,
809: },
810: {
811: .cmd = DEVLINK_CMD_SB_POOL_SET,
812: .validate = GENL_DONT_VALIDATE_STRICT,
813: .pre_doit = devlink_nl_pre_doit,
814: .doit = devlink_nl_sb_pool_set_doit,
815: .post_doit = devlink_nl_post_doit,
816: .policy = devlink_sb_pool_set_nl_policy,
817: .maxattr = DEVLINK_ATTR_INDEX,
818: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
819: },
820: {
821: .cmd = DEVLINK_CMD_SB_PORT_POOL_GET,
822: .validate = GENL_DONT_VALIDATE_STRICT,
823: .pre_doit = devlink_nl_pre_doit_port,
824: .doit = devlink_nl_sb_port_pool_get_doit,
825: .post_doit = devlink_nl_post_doit,
826: .policy = devlink_sb_port_pool_get_do_nl_policy,
827: .maxattr = DEVLINK_ATTR_INDEX,
828: .flags = GENL_CMD_CAP_DO,
829: },
830: {
831: .cmd = DEVLINK_CMD_SB_PORT_POOL_GET,
832: .dumpit = devlink_nl_sb_port_pool_get_dumpit,
833: .policy = devlink_sb_port_pool_get_dump_nl_policy,
834: .maxattr = DEVLINK_ATTR_INDEX,
835: .flags = GENL_CMD_CAP_DUMP,
836: },
837: {
838: .cmd = DEVLINK_CMD_SB_PORT_POOL_SET,
839: .validate = GENL_DONT_VALIDATE_STRICT,
840: .pre_doit = devlink_nl_pre_doit_port,
841: .doit = devlink_nl_sb_port_pool_set_doit,
842: .post_doit = devlink_nl_post_doit,
843: .policy = devlink_sb_port_pool_set_nl_policy,
844: .maxattr = DEVLINK_ATTR_INDEX,
845: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
846: },
847: {
848: .cmd = DEVLINK_CMD_SB_TC_POOL_BIND_GET,
849: .validate = GENL_DONT_VALIDATE_STRICT,
850: .pre_doit = devlink_nl_pre_doit_port,
851: .doit = devlink_nl_sb_tc_pool_bind_get_doit,
852: .post_doit = devlink_nl_post_doit,
853: .policy = devlink_sb_tc_pool_bind_get_do_nl_policy,
854: .maxattr = DEVLINK_ATTR_INDEX,
855: .flags = GENL_CMD_CAP_DO,
856: },
857: {
858: .cmd = DEVLINK_CMD_SB_TC_POOL_BIND_GET,
859: .dumpit = devlink_nl_sb_tc_pool_bind_get_dumpit,
860: .policy = devlink_sb_tc_pool_bind_get_dump_nl_policy,
861: .maxattr = DEVLINK_ATTR_INDEX,
862: .flags = GENL_CMD_CAP_DUMP,
863: },
864: {
865: .cmd = DEVLINK_CMD_SB_TC_POOL_BIND_SET,
866: .validate = GENL_DONT_VALIDATE_STRICT,
867: .pre_doit = devlink_nl_pre_doit_port,
868: .doit = devlink_nl_sb_tc_pool_bind_set_doit,
869: .post_doit = devlink_nl_post_doit,
870: .policy = devlink_sb_tc_pool_bind_set_nl_policy,
871: .maxattr = DEVLINK_ATTR_INDEX,
872: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
873: },
874: {
875: .cmd = DEVLINK_CMD_SB_OCC_SNAPSHOT,
876: .validate = GENL_DONT_VALIDATE_STRICT,
877: .pre_doit = devlink_nl_pre_doit,
878: .doit = devlink_nl_sb_occ_snapshot_doit,
879: .post_doit = devlink_nl_post_doit,
880: .policy = devlink_sb_occ_snapshot_nl_policy,
881: .maxattr = DEVLINK_ATTR_INDEX,
882: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
883: },
884: {
885: .cmd = DEVLINK_CMD_SB_OCC_MAX_CLEAR,
886: .validate = GENL_DONT_VALIDATE_STRICT,
887: .pre_doit = devlink_nl_pre_doit,
888: .doit = devlink_nl_sb_occ_max_clear_doit,
889: .post_doit = devlink_nl_post_doit,
890: .policy = devlink_sb_occ_max_clear_nl_policy,
891: .maxattr = DEVLINK_ATTR_INDEX,
892: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
893: },
894: {
895: .cmd = DEVLINK_CMD_ESWITCH_GET,
896: .validate = GENL_DONT_VALIDATE_STRICT,
897: .pre_doit = devlink_nl_pre_doit,
898: .doit = devlink_nl_eswitch_get_doit,
899: .post_doit = devlink_nl_post_doit,
900: .policy = devlink_eswitch_get_nl_policy,
901: .maxattr = DEVLINK_ATTR_INDEX,
902: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
903: },
904: {
905: .cmd = DEVLINK_CMD_ESWITCH_SET,
906: .validate = GENL_DONT_VALIDATE_STRICT,
907: .pre_doit = devlink_nl_pre_doit,
908: .doit = devlink_nl_eswitch_set_doit,
909: .post_doit = devlink_nl_post_doit,
910: .policy = devlink_eswitch_set_nl_policy,
911: .maxattr = DEVLINK_ATTR_INDEX,
912: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
913: },
914: {
915: .cmd = DEVLINK_CMD_DPIPE_TABLE_GET,
916: .validate = GENL_DONT_VALIDATE_STRICT,
917: .pre_doit = devlink_nl_pre_doit,
918: .doit = devlink_nl_dpipe_table_get_doit,
919: .post_doit = devlink_nl_post_doit,
920: .policy = devlink_dpipe_table_get_nl_policy,
921: .maxattr = DEVLINK_ATTR_INDEX,
922: .flags = GENL_CMD_CAP_DO,
923: },
924: {
925: .cmd = DEVLINK_CMD_DPIPE_ENTRIES_GET,
926: .validate = GENL_DONT_VALIDATE_STRICT,
927: .pre_doit = devlink_nl_pre_doit,
928: .doit = devlink_nl_dpipe_entries_get_doit,
929: .post_doit = devlink_nl_post_doit,
930: .policy = devlink_dpipe_entries_get_nl_policy,
931: .maxattr = DEVLINK_ATTR_INDEX,
932: .flags = GENL_CMD_CAP_DO,
933: },
934: {
935: .cmd = DEVLINK_CMD_DPIPE_HEADERS_GET,
936: .validate = GENL_DONT_VALIDATE_STRICT,
937: .pre_doit = devlink_nl_pre_doit,
938: .doit = devlink_nl_dpipe_headers_get_doit,
939: .post_doit = devlink_nl_post_doit,
940: .policy = devlink_dpipe_headers_get_nl_policy,
941: .maxattr = DEVLINK_ATTR_INDEX,
942: .flags = GENL_CMD_CAP_DO,
943: },
944: {
945: .cmd = DEVLINK_CMD_DPIPE_TABLE_COUNTERS_SET,
946: .validate = GENL_DONT_VALIDATE_STRICT,
947: .pre_doit = devlink_nl_pre_doit,
948: .doit = devlink_nl_dpipe_table_counters_set_doit,
949: .post_doit = devlink_nl_post_doit,
950: .policy = devlink_dpipe_table_counters_set_nl_policy,
951: .maxattr = DEVLINK_ATTR_INDEX,
952: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
953: },
954: {
955: .cmd = DEVLINK_CMD_RESOURCE_SET,
956: .validate = GENL_DONT_VALIDATE_STRICT,
957: .pre_doit = devlink_nl_pre_doit,
958: .doit = devlink_nl_resource_set_doit,
959: .post_doit = devlink_nl_post_doit,
960: .policy = devlink_resource_set_nl_policy,
961: .maxattr = DEVLINK_ATTR_INDEX,
962: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
963: },
964: {
965: .cmd = DEVLINK_CMD_RESOURCE_DUMP,
966: .validate = GENL_DONT_VALIDATE_STRICT,
967: .pre_doit = devlink_nl_pre_doit_port_optional,
968: .doit = devlink_nl_resource_dump_doit,
969: .post_doit = devlink_nl_post_doit,
970: .policy = devlink_resource_dump_do_nl_policy,
971: .maxattr = DEVLINK_ATTR_INDEX,
972: .flags = GENL_CMD_CAP_DO,
973: },
974: {
975: .cmd = DEVLINK_CMD_RESOURCE_DUMP,
976: .dumpit = devlink_nl_resource_dump_dumpit,
977: .policy = devlink_resource_dump_dump_nl_policy,
978: .maxattr = DEVLINK_ATTR_RESOURCE_SCOPE_MASK,
979: .flags = GENL_CMD_CAP_DUMP,
980: },
981: {
982: .cmd = DEVLINK_CMD_RELOAD,
983: .validate = GENL_DONT_VALIDATE_STRICT,
984: .pre_doit = devlink_nl_pre_doit_dev_lock,
985: .doit = devlink_nl_reload_doit,
986: .post_doit = devlink_nl_post_doit_dev_lock,
987: .policy = devlink_reload_nl_policy,
988: .maxattr = DEVLINK_ATTR_INDEX,
989: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
990: },
991: {
992: .cmd = DEVLINK_CMD_PARAM_GET,
993: .validate = GENL_DONT_VALIDATE_STRICT,
994: .pre_doit = devlink_nl_pre_doit,
995: .doit = devlink_nl_param_get_doit,
996: .post_doit = devlink_nl_post_doit,
997: .policy = devlink_param_get_do_nl_policy,
998: .maxattr = DEVLINK_ATTR_INDEX,
999: .flags = GENL_CMD_CAP_DO,
1000: },
1001: {
1002: .cmd = DEVLINK_CMD_PARAM_GET,
1003: .dumpit = devlink_nl_param_get_dumpit,
1004: .policy = devlink_param_get_dump_nl_policy,
1005: .maxattr = DEVLINK_ATTR_INDEX,
1006: .flags = GENL_CMD_CAP_DUMP,
1007: },
1008: {
1009: .cmd = DEVLINK_CMD_PARAM_SET,
1010: .validate = GENL_DONT_VALIDATE_STRICT,
1011: .pre_doit = devlink_nl_pre_doit,
1012: .doit = devlink_nl_param_set_doit,
1013: .post_doit = devlink_nl_post_doit,
1014: .policy = devlink_param_set_nl_policy,
1015: .maxattr = DEVLINK_ATTR_INDEX,
1016: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1017: },
1018: {
1019: .cmd = DEVLINK_CMD_REGION_GET,
1020: .validate = GENL_DONT_VALIDATE_STRICT,
1021: .pre_doit = devlink_nl_pre_doit_port_optional,
1022: .doit = devlink_nl_region_get_doit,
1023: .post_doit = devlink_nl_post_doit,
1024: .policy = devlink_region_get_do_nl_policy,
1025: .maxattr = DEVLINK_ATTR_INDEX,
1026: .flags = GENL_CMD_CAP_DO,
1027: },
1028: {
1029: .cmd = DEVLINK_CMD_REGION_GET,
1030: .dumpit = devlink_nl_region_get_dumpit,
1031: .policy = devlink_region_get_dump_nl_policy,
1032: .maxattr = DEVLINK_ATTR_INDEX,
1033: .flags = GENL_CMD_CAP_DUMP,
1034: },
1035: {
1036: .cmd = DEVLINK_CMD_REGION_NEW,
1037: .validate = GENL_DONT_VALIDATE_STRICT,
1038: .pre_doit = devlink_nl_pre_doit_port_optional,
1039: .doit = devlink_nl_region_new_doit,
1040: .post_doit = devlink_nl_post_doit,
1041: .policy = devlink_region_new_nl_policy,
1042: .maxattr = DEVLINK_ATTR_INDEX,
1043: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1044: },
1045: {
1046: .cmd = DEVLINK_CMD_REGION_DEL,
1047: .validate = GENL_DONT_VALIDATE_STRICT,
1048: .pre_doit = devlink_nl_pre_doit_port_optional,
1049: .doit = devlink_nl_region_del_doit,
1050: .post_doit = devlink_nl_post_doit,
1051: .policy = devlink_region_del_nl_policy,
1052: .maxattr = DEVLINK_ATTR_INDEX,
1053: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1054: },
1055: {
1056: .cmd = DEVLINK_CMD_REGION_READ,
1057: .validate = GENL_DONT_VALIDATE_DUMP_STRICT,
1058: .dumpit = devlink_nl_region_read_dumpit,
1059: .policy = devlink_region_read_nl_policy,
1060: .maxattr = DEVLINK_ATTR_INDEX,
1061: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DUMP,
1062: },
1063: {
1064: .cmd = DEVLINK_CMD_PORT_PARAM_GET,
1065: .validate = GENL_DONT_VALIDATE_STRICT,
1066: .pre_doit = devlink_nl_pre_doit_port,
1067: .doit = devlink_nl_port_param_get_doit,
1068: .post_doit = devlink_nl_post_doit,
1069: .policy = devlink_port_param_get_nl_policy,
1070: .maxattr = DEVLINK_ATTR_INDEX,
1071: .flags = GENL_CMD_CAP_DO,
1072: },
1073: {
1074: .cmd = DEVLINK_CMD_PORT_PARAM_GET,
1075: .validate = GENL_DONT_VALIDATE_DUMP_STRICT,
1076: .dumpit = devlink_nl_port_param_get_dumpit,
1077: .flags = GENL_CMD_CAP_DUMP,
1078: },
1079: {
1080: .cmd = DEVLINK_CMD_PORT_PARAM_SET,
1081: .validate = GENL_DONT_VALIDATE_STRICT,
1082: .pre_doit = devlink_nl_pre_doit_port,
1083: .doit = devlink_nl_port_param_set_doit,
1084: .post_doit = devlink_nl_post_doit,
1085: .policy = devlink_port_param_set_nl_policy,
1086: .maxattr = DEVLINK_ATTR_INDEX,
1087: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1088: },
1089: {
1090: .cmd = DEVLINK_CMD_INFO_GET,
1091: .validate = GENL_DONT_VALIDATE_STRICT,
1092: .pre_doit = devlink_nl_pre_doit,
1093: .doit = devlink_nl_info_get_doit,
1094: .post_doit = devlink_nl_post_doit,
1095: .policy = devlink_info_get_nl_policy,
1096: .maxattr = DEVLINK_ATTR_INDEX,
1097: .flags = GENL_CMD_CAP_DO,
1098: },
1099: {
1100: .cmd = DEVLINK_CMD_INFO_GET,
1101: .validate = GENL_DONT_VALIDATE_DUMP,
1102: .dumpit = devlink_nl_info_get_dumpit,
1103: .flags = GENL_CMD_CAP_DUMP,
1104: },
1105: {
1106: .cmd = DEVLINK_CMD_HEALTH_REPORTER_GET,
1107: .validate = GENL_DONT_VALIDATE_STRICT,
1108: .pre_doit = devlink_nl_pre_doit_port_optional,
1109: .doit = devlink_nl_health_reporter_get_doit,
1110: .post_doit = devlink_nl_post_doit,
1111: .policy = devlink_health_reporter_get_do_nl_policy,
1112: .maxattr = DEVLINK_ATTR_INDEX,
1113: .flags = GENL_CMD_CAP_DO,
1114: },
1115: {
1116: .cmd = DEVLINK_CMD_HEALTH_REPORTER_GET,
1117: .dumpit = devlink_nl_health_reporter_get_dumpit,
1118: .policy = devlink_health_reporter_get_dump_nl_policy,
1119: .maxattr = DEVLINK_ATTR_INDEX,
1120: .flags = GENL_CMD_CAP_DUMP,
1121: },
1122: {
1123: .cmd = DEVLINK_CMD_HEALTH_REPORTER_SET,
1124: .validate = GENL_DONT_VALIDATE_STRICT,
1125: .pre_doit = devlink_nl_pre_doit_port_optional,
1126: .doit = devlink_nl_health_reporter_set_doit,
1127: .post_doit = devlink_nl_post_doit,
1128: .policy = devlink_health_reporter_set_nl_policy,
1129: .maxattr = DEVLINK_ATTR_INDEX,
1130: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1131: },
1132: {
1133: .cmd = DEVLINK_CMD_HEALTH_REPORTER_RECOVER,
1134: .validate = GENL_DONT_VALIDATE_STRICT,
1135: .pre_doit = devlink_nl_pre_doit_port_optional,
1136: .doit = devlink_nl_health_reporter_recover_doit,
1137: .post_doit = devlink_nl_post_doit,
1138: .policy = devlink_health_reporter_recover_nl_policy,
1139: .maxattr = DEVLINK_ATTR_INDEX,
1140: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1141: },
1142: {
1143: .cmd = DEVLINK_CMD_HEALTH_REPORTER_DIAGNOSE,
1144: .validate = GENL_DONT_VALIDATE_STRICT,
1145: .pre_doit = devlink_nl_pre_doit_port_optional,
1146: .doit = devlink_nl_health_reporter_diagnose_doit,
1147: .post_doit = devlink_nl_post_doit,
1148: .policy = devlink_health_reporter_diagnose_nl_policy,
1149: .maxattr = DEVLINK_ATTR_INDEX,
1150: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1151: },
1152: {
1153: .cmd = DEVLINK_CMD_HEALTH_REPORTER_DUMP_GET,
1154: .validate = GENL_DONT_VALIDATE_DUMP_STRICT,
1155: .dumpit = devlink_nl_health_reporter_dump_get_dumpit,
1156: .policy = devlink_health_reporter_dump_get_nl_policy,
1157: .maxattr = DEVLINK_ATTR_INDEX,
1158: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DUMP,
1159: },
1160: {
1161: .cmd = DEVLINK_CMD_HEALTH_REPORTER_DUMP_CLEAR,
1162: .validate = GENL_DONT_VALIDATE_STRICT,
1163: .pre_doit = devlink_nl_pre_doit_port_optional,
1164: .doit = devlink_nl_health_reporter_dump_clear_doit,
1165: .post_doit = devlink_nl_post_doit,
1166: .policy = devlink_health_reporter_dump_clear_nl_policy,
1167: .maxattr = DEVLINK_ATTR_INDEX,
1168: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1169: },
1170: {
1171: .cmd = DEVLINK_CMD_FLASH_UPDATE,
1172: .validate = GENL_DONT_VALIDATE_STRICT,
1173: .pre_doit = devlink_nl_pre_doit,
1174: .doit = devlink_nl_flash_update_doit,
1175: .post_doit = devlink_nl_post_doit,
1176: .policy = devlink_flash_update_nl_policy,
1177: .maxattr = DEVLINK_ATTR_INDEX,
1178: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1179: },
1180: {
1181: .cmd = DEVLINK_CMD_TRAP_GET,
1182: .validate = GENL_DONT_VALIDATE_STRICT,
1183: .pre_doit = devlink_nl_pre_doit,
1184: .doit = devlink_nl_trap_get_doit,
1185: .post_doit = devlink_nl_post_doit,
1186: .policy = devlink_trap_get_do_nl_policy,
1187: .maxattr = DEVLINK_ATTR_INDEX,
1188: .flags = GENL_CMD_CAP_DO,
1189: },
1190: {
1191: .cmd = DEVLINK_CMD_TRAP_GET,
1192: .dumpit = devlink_nl_trap_get_dumpit,
1193: .policy = devlink_trap_get_dump_nl_policy,
1194: .maxattr = DEVLINK_ATTR_INDEX,
1195: .flags = GENL_CMD_CAP_DUMP,
1196: },
1197: {
1198: .cmd = DEVLINK_CMD_TRAP_SET,
1199: .validate = GENL_DONT_VALIDATE_STRICT,
1200: .pre_doit = devlink_nl_pre_doit,
1201: .doit = devlink_nl_trap_set_doit,
1202: .post_doit = devlink_nl_post_doit,
1203: .policy = devlink_trap_set_nl_policy,
1204: .maxattr = DEVLINK_ATTR_INDEX,
1205: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1206: },
1207: {
1208: .cmd = DEVLINK_CMD_TRAP_GROUP_GET,
1209: .validate = GENL_DONT_VALIDATE_STRICT,
1210: .pre_doit = devlink_nl_pre_doit,
1211: .doit = devlink_nl_trap_group_get_doit,
1212: .post_doit = devlink_nl_post_doit,
1213: .policy = devlink_trap_group_get_do_nl_policy,
1214: .maxattr = DEVLINK_ATTR_INDEX,
1215: .flags = GENL_CMD_CAP_DO,
1216: },
1217: {
1218: .cmd = DEVLINK_CMD_TRAP_GROUP_GET,
1219: .dumpit = devlink_nl_trap_group_get_dumpit,
1220: .policy = devlink_trap_group_get_dump_nl_policy,
1221: .maxattr = DEVLINK_ATTR_INDEX,
1222: .flags = GENL_CMD_CAP_DUMP,
1223: },
1224: {
1225: .cmd = DEVLINK_CMD_TRAP_GROUP_SET,
1226: .validate = GENL_DONT_VALIDATE_STRICT,
1227: .pre_doit = devlink_nl_pre_doit,
1228: .doit = devlink_nl_trap_group_set_doit,
1229: .post_doit = devlink_nl_post_doit,
1230: .policy = devlink_trap_group_set_nl_policy,
1231: .maxattr = DEVLINK_ATTR_INDEX,
1232: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1233: },
1234: {
1235: .cmd = DEVLINK_CMD_TRAP_POLICER_GET,
1236: .validate = GENL_DONT_VALIDATE_STRICT,
1237: .pre_doit = devlink_nl_pre_doit,
1238: .doit = devlink_nl_trap_policer_get_doit,
1239: .post_doit = devlink_nl_post_doit,
1240: .policy = devlink_trap_policer_get_do_nl_policy,
1241: .maxattr = DEVLINK_ATTR_INDEX,
1242: .flags = GENL_CMD_CAP_DO,
1243: },
1244: {
1245: .cmd = DEVLINK_CMD_TRAP_POLICER_GET,
1246: .dumpit = devlink_nl_trap_policer_get_dumpit,
1247: .policy = devlink_trap_policer_get_dump_nl_policy,
1248: .maxattr = DEVLINK_ATTR_INDEX,
1249: .flags = GENL_CMD_CAP_DUMP,
1250: },
1251: {
1252: .cmd = DEVLINK_CMD_TRAP_POLICER_SET,
1253: .validate = GENL_DONT_VALIDATE_STRICT,
1254: .pre_doit = devlink_nl_pre_doit,
1255: .doit = devlink_nl_trap_policer_set_doit,
1256: .post_doit = devlink_nl_post_doit,
1257: .policy = devlink_trap_policer_set_nl_policy,
1258: .maxattr = DEVLINK_ATTR_INDEX,
1259: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1260: },
1261: {
1262: .cmd = DEVLINK_CMD_HEALTH_REPORTER_TEST,
1263: .validate = GENL_DONT_VALIDATE_STRICT,
1264: .pre_doit = devlink_nl_pre_doit_port_optional,
1265: .doit = devlink_nl_health_reporter_test_doit,
1266: .post_doit = devlink_nl_post_doit,
1267: .policy = devlink_health_reporter_test_nl_policy,
1268: .maxattr = DEVLINK_ATTR_INDEX,
1269: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1270: },
1271: {
1272: .cmd = DEVLINK_CMD_RATE_GET,
1273: .validate = GENL_DONT_VALIDATE_STRICT,
1274: .pre_doit = devlink_nl_pre_doit,
1275: .doit = devlink_nl_rate_get_doit,
1276: .post_doit = devlink_nl_post_doit,
1277: .policy = devlink_rate_get_do_nl_policy,
1278: .maxattr = DEVLINK_ATTR_INDEX,
1279: .flags = GENL_CMD_CAP_DO,
1280: },
1281: {
1282: .cmd = DEVLINK_CMD_RATE_GET,
1283: .dumpit = devlink_nl_rate_get_dumpit,
1284: .policy = devlink_rate_get_dump_nl_policy,
1285: .maxattr = DEVLINK_ATTR_INDEX,
1286: .flags = GENL_CMD_CAP_DUMP,
1287: },
1288: {
1289: .cmd = DEVLINK_CMD_RATE_SET,
1290: .validate = GENL_DONT_VALIDATE_STRICT,
1291: .pre_doit = devlink_nl_pre_doit,
1292: .doit = devlink_nl_rate_set_doit,
1293: .post_doit = devlink_nl_post_doit,
1294: .policy = devlink_rate_set_nl_policy,
1295: .maxattr = DEVLINK_ATTR_INDEX,
1296: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1297: },
1298: {
1299: .cmd = DEVLINK_CMD_RATE_NEW,
1300: .validate = GENL_DONT_VALIDATE_STRICT,
1301: .pre_doit = devlink_nl_pre_doit,
1302: .doit = devlink_nl_rate_new_doit,
1303: .post_doit = devlink_nl_post_doit,
1304: .policy = devlink_rate_new_nl_policy,
1305: .maxattr = DEVLINK_ATTR_INDEX,
1306: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1307: },
1308: {
1309: .cmd = DEVLINK_CMD_RATE_DEL,
1310: .validate = GENL_DONT_VALIDATE_STRICT,
1311: .pre_doit = devlink_nl_pre_doit,
1312: .doit = devlink_nl_rate_del_doit,
1313: .post_doit = devlink_nl_post_doit,
1314: .policy = devlink_rate_del_nl_policy,
1315: .maxattr = DEVLINK_ATTR_INDEX,
1316: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1317: },
1318: {
1319: .cmd = DEVLINK_CMD_LINECARD_GET,
1320: .validate = GENL_DONT_VALIDATE_STRICT,
1321: .pre_doit = devlink_nl_pre_doit,
1322: .doit = devlink_nl_linecard_get_doit,
1323: .post_doit = devlink_nl_post_doit,
1324: .policy = devlink_linecard_get_do_nl_policy,
1325: .maxattr = DEVLINK_ATTR_INDEX,
1326: .flags = GENL_CMD_CAP_DO,
1327: },
1328: {
1329: .cmd = DEVLINK_CMD_LINECARD_GET,
1330: .dumpit = devlink_nl_linecard_get_dumpit,
1331: .policy = devlink_linecard_get_dump_nl_policy,
1332: .maxattr = DEVLINK_ATTR_INDEX,
1333: .flags = GENL_CMD_CAP_DUMP,
1334: },
1335: {
1336: .cmd = DEVLINK_CMD_LINECARD_SET,
1337: .validate = GENL_DONT_VALIDATE_STRICT,
1338: .pre_doit = devlink_nl_pre_doit,
1339: .doit = devlink_nl_linecard_set_doit,
1340: .post_doit = devlink_nl_post_doit,
1341: .policy = devlink_linecard_set_nl_policy,
1342: .maxattr = DEVLINK_ATTR_INDEX,
1343: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1344: },
1345: {
1346: .cmd = DEVLINK_CMD_SELFTESTS_GET,
1347: .validate = GENL_DONT_VALIDATE_STRICT,
1348: .pre_doit = devlink_nl_pre_doit,
1349: .doit = devlink_nl_selftests_get_doit,
1350: .post_doit = devlink_nl_post_doit,
1351: .policy = devlink_selftests_get_nl_policy,
1352: .maxattr = DEVLINK_ATTR_INDEX,
1353: .flags = GENL_CMD_CAP_DO,
1354: },
1355: {
1356: .cmd = DEVLINK_CMD_SELFTESTS_GET,
1357: .validate = GENL_DONT_VALIDATE_DUMP,
1358: .dumpit = devlink_nl_selftests_get_dumpit,
1359: .flags = GENL_CMD_CAP_DUMP,
1360: },
1361: {
1362: .cmd = DEVLINK_CMD_SELFTESTS_RUN,
1363: .validate = GENL_DONT_VALIDATE_STRICT,
1364: .pre_doit = devlink_nl_pre_doit,
1365: .doit = devlink_nl_selftests_run_doit,
1366: .post_doit = devlink_nl_post_doit,
1367: .policy = devlink_selftests_run_nl_policy,
1368: .maxattr = DEVLINK_ATTR_INDEX,
1369: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1370: },
1371: {
1372: .cmd = DEVLINK_CMD_NOTIFY_FILTER_SET,
1373: .doit = devlink_nl_notify_filter_set_doit,
1374: .policy = devlink_notify_filter_set_nl_policy,
1375: .maxattr = DEVLINK_ATTR_INDEX,
1376: .flags = GENL_CMD_CAP_DO,
1377: },
1378: };
SourceFile:net/devlink/netlink_gen.c]
|
| 78/2 |
2026/05/26 14:40 |
llm |
expert |
0m
Model:
Tokens:
input: 56697
output: 51
thoughts: 987
Thoughts: **Analyzing the `devlink_reload` Vulnerability**
Okay, let's break this down. My initial thought was that `DEVLINK_CMD_RELOAD` requires `GENL_ADMIN_PERM`, which translates to `CAP_NET_ADMIN`. So, a regular, unprivileged user should be locked out. But the bug report mentions `devlink_pernet_pre_exit` triggering `devlink_reload`. That function is called during `cleanup_net`, which runs when a network namespace is destroyed.
Ah-ha! Unprivileged users can create and destroy network namespaces *if* user namespaces are enabled. They can make a user namespace, then a network namespace within it. When the process exits, the inner namespace dies, leading to `cleanup_net` and then `devlink_reload`. But the question is, can they manipulate the devlink instance? `devlink_pernet_pre_exit` reloads instances belonging to the dying namespace. Moving a devlink instance typically needs `CAP_NET_ADMIN` in the original namespace or the namespace the instance currently resides in.
Now, `netdevsim` creates devlink instances. Can an unprivileged user do this? Usually not, since it's via sysfs or devlink which requires root. But the bug report mentions user namespaces with all capabilities. So the process, inside its user namespace, *has* `CAP_NET_ADMIN`. Can it create the `netdevsim` instance then?
`netdevsim` uses a virtual bus, right? Let's check how they're created... typically through writing to `/sys/bus/netdevsim/new_device`. Sysfs is usually read-only in user namespaces, so writing is out. Wait a second, the devices are created by the `netdevsim` bus, not by devlink!
But this bug isn't limited to `netdevsim`. It's also in `mlxsw` and `prestera`. When those are loaded, they register a fib notifier in `init_net`. Specifically, I see `mlxsw_sp_router_init` registering a notifier. Destruction of a network namespace won't trigger `mlxsw` initialization directly. However, the notifier's call: `mlxsw_sp_router_fib_event` is called for *every* fib event.
The real bug is the `fib_info_hold` on a `fib_info` with a refcount of zero. This means it found a `fib_info` in the trie with the refcount at zero. When does that happen? Only when `fib_net_dump` is running concurrently with route deletion! Normal route deletion handles the notifier *before* releasing the `fib_info`. Thus, the notifier is triggered only when the refcount is greater than 0.
However, during `fib_net_dump`, walking the trie may expose the `fib_alias`, whose `fib_info` has already been released by a concurrent `fib_table_delete`. Thus, `fib_net_dump` is where the bug surfaces, not during the *normal* route deletion.
So, to trigger the `mlxsw` or `prestera` issue, you need to trigger `fib_net_dump`. How to do that? By registering a new fib notifier. But drivers during initialization will do this (e.g., `mlxsw_sp_router_init`). Is there another way to trigger `fib_net_dump`? Let's trace calls to `fib_net_dump`...
|
| 79/2 |
2026/05/26 14:41 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/core/fib_notifier.c IncludeSnippetLines:0 Name:fib_net_dump SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:register_fib_notifier SourceFile:net/core/fib_notifier.c SourceLine:108]] TruncatedOutput:false]
|
| 80/2 |
2026/05/26 14:41 |
llm |
expert |
0m
Model:
Tokens:
input: 58563
output: 19
thoughts: 426
Thoughts: **Analyzing the Bug Trigger Condition**
Okay, so I'm trying to figure out how this bug is triggered. The crucial clue is that only `register_fib_notifier` calls `fib_net_dump`. Therefore, the bug *must* be triggered when `register_fib_notifier` is called. Let's see how that happens.
For `netdevsim`, which seems relevant here, `register_fib_notifier` is called during `nsim_fib_create`. `nsim_fib_create` is then called during `devlink_reload`. Ah, and `devlink_reload` can be triggered in a couple of ways: either by a `DEVLINK_CMD_RELOAD` command (which requires `CAP_NET_ADMIN` privileges) or by `devlink_pernet_pre_exit`.
Now, if a devlink instance exists in a network namespace (netns) and that netns gets destroyed, `devlink_pernet_pre_exit` reloads the devlink instance into the `init_net`. That's interesting...
Can an unprivileged user move a devlink instance into their own netns? It looks like yes, but only if they have `CAP_NET_ADMIN` privileges in the namespace that currently owns the devlink instance. But the devlink instance initially lives in `init_net`. So to move it, you'd need `CAP_NET_ADMIN` in `init_net`, which essentially means you need root privileges. So, an unprivileged user can't trigger `devlink_pernet_pre_exit` for a devlink instance directly.
Wait a second... what if an unprivileged user *creates* a netns and then inside that netns, they somehow trigger a route deletion? If root or some other privileged process were to *concurrently* trigger `devlink_reload`, or even just load a driver like `mlxsw`, maybe... *maybe* the bug could be triggered. But, the problem remains that triggering the bug still requires calling `register_fib_notifier`, which is a privileged operation.
Hang on... am I missing something? Are there any other callers of `register_fib_notifier`? Let's not make assumptions. What about `ip_fib_net_init`? No, no, no, that registers `fib_notifier_ops`, not the notifier block itself. I need to be absolutely sure. Let me go back and double-check all the callers of `register_fib_notifier` one more time. It's easy to get lost in these details.
|
| 81/2 |
2026/05/26 14:41 |
tool |
grepper |
0m
Args: map[Expression:register_fib_notifier]
Results: map[Output:Documentation/networking/switchdev.rst=349=To program the device, the driver has to register a FIB notifier handler
Documentation/networking/switchdev.rst:350:using register_fib_notifier. The following events are available:
Documentation/networking/switchdev.rst-351-
--
drivers/net/ethernet/marvell/prestera/prestera_router.c=1548=int prestera_router_init(struct prestera_switch *sw)
--
drivers/net/ethernet/marvell/prestera/prestera_router.c-1600- router->fib_nb.notifier_call = __prestera_router_fib_event;
drivers/net/ethernet/marvell/prestera/prestera_router.c:1601: err = register_fib_notifier(&init_net, &router->fib_nb,
drivers/net/ethernet/marvell/prestera/prestera_router.c-1602- /* TODO: flush fib entries */ NULL, NULL);
drivers/net/ethernet/marvell/prestera/prestera_router.c-1603- if (err)
drivers/net/ethernet/marvell/prestera/prestera_router.c:1604: goto err_register_fib_notifier;
drivers/net/ethernet/marvell/prestera/prestera_router.c-1605-
--
drivers/net/ethernet/marvell/prestera/prestera_router.c-1607-
drivers/net/ethernet/marvell/prestera/prestera_router.c:1608:err_register_fib_notifier:
drivers/net/ethernet/marvell/prestera/prestera_router.c-1609- unregister_netevent_notifier(&router->netevent_nb);
--
drivers/net/ethernet/marvell/prestera/prestera_router.c=1629=void prestera_router_fini(struct prestera_switch *sw)
drivers/net/ethernet/marvell/prestera/prestera_router.c-1630-{
drivers/net/ethernet/marvell/prestera/prestera_router.c:1631: unregister_fib_notifier(&init_net, &sw->router->fib_nb);
drivers/net/ethernet/marvell/prestera/prestera_router.c-1632- unregister_netevent_notifier(&sw->router->netevent_nb);
--
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c=1860=struct mlx5e_tc_tun_encap *mlx5e_tc_tun_init(struct mlx5e_priv *priv)
--
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c-1872- hash_init(encap->route_tbl);
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c:1873: err = register_fib_notifier(dev_net(priv->netdev), &encap->fib_nb,
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c-1874- NULL, NULL);
--
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c=1883=void mlx5e_tc_tun_cleanup(struct mlx5e_tc_tun_encap *encap)
--
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c-1887-
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c:1888: unregister_fib_notifier(dev_net(encap->priv->netdev), &encap->fib_nb);
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c-1889- flush_workqueue(encap->priv->wq); /* flush fib event works */
--
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c=5393=int mlx5e_tc_esw_init(struct mlx5_rep_uplink_priv *uplink_priv)
--
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c-5449- err = PTR_ERR(uplink_priv->encap);
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c:5450: goto err_register_fib_notifier;
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c-5451- }
--
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c-5470- mlx5e_tc_tun_cleanup(uplink_priv->encap);
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c:5471:err_register_fib_notifier:
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c-5472- mapping_destroy(uplink_priv->tunnel_enc_opts_mapping);
--
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c=368=int mlx5_lag_mp_init(struct mlx5_lag *ldev)
--
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c-385- mp->fib_nb.notifier_call = mlx5_lag_fib_event;
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c:386: err = register_fib_notifier(&init_net, &mp->fib_nb,
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c-387- mlx5_lag_fib_event_flush, NULL);
--
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c=396=void mlx5_lag_mp_cleanup(struct mlx5_lag *ldev)
--
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c-402-
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c:403: unregister_fib_notifier(&init_net, &mp->fib_nb);
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c-404- destroy_workqueue(mp->wq);
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c=11575=int mlxsw_sp_router_init(struct mlxsw_sp *mlxsw_sp,
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11695- mlxsw_sp->router->fib_nb.notifier_call = mlxsw_sp_router_fib_event;
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c:11696: err = register_fib_notifier(mlxsw_sp_net(mlxsw_sp),
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11697- &mlxsw_sp->router->fib_nb,
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11699- if (err)
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c:11700: goto err_register_fib_notifier;
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11701-
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11703-
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c:11704:err_register_fib_notifier:
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11705- unregister_nexthop_notifier(mlxsw_sp_net(mlxsw_sp),
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c=11754=void mlxsw_sp_router_fini(struct mlxsw_sp *mlxsw_sp)
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11757-
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c:11758: unregister_fib_notifier(mlxsw_sp_net(mlxsw_sp), &router->fib_nb);
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11759- unregister_nexthop_notifier(mlxsw_sp_net(mlxsw_sp),
--
drivers/net/ethernet/rocker/rocker_main.c=2844=static int rocker_probe(struct pci_dev *pdev, const struct pci_device_id *id)
--
drivers/net/ethernet/rocker/rocker_main.c-2942- rocker->fib_nb.notifier_call = rocker_router_fib_event;
drivers/net/ethernet/rocker/rocker_main.c:2943: err = register_fib_notifier(&init_net, &rocker->fib_nb, NULL, NULL);
drivers/net/ethernet/rocker/rocker_main.c-2944- if (err)
drivers/net/ethernet/rocker/rocker_main.c:2945: goto err_register_fib_notifier;
drivers/net/ethernet/rocker/rocker_main.c-2946-
--
drivers/net/ethernet/rocker/rocker_main.c-2969-err_register_switchdev_notifier:
drivers/net/ethernet/rocker/rocker_main.c:2970: unregister_fib_notifier(&init_net, &rocker->fib_nb);
drivers/net/ethernet/rocker/rocker_main.c:2971:err_register_fib_notifier:
drivers/net/ethernet/rocker/rocker_main.c-2972- rocker_remove_ports(rocker);
--
drivers/net/ethernet/rocker/rocker_main.c=2997=static void rocker_remove(struct pci_dev *pdev)
--
drivers/net/ethernet/rocker/rocker_main.c-3005- unregister_switchdev_notifier(&rocker_switchdev_notifier);
drivers/net/ethernet/rocker/rocker_main.c:3006: unregister_fib_notifier(&init_net, &rocker->fib_nb);
drivers/net/ethernet/rocker/rocker_main.c-3007- rocker_remove_ports(rocker);
--
drivers/net/netdevsim/fib.c=1552=struct nsim_fib_data *nsim_fib_create(struct devlink *devlink,
--
drivers/net/netdevsim/fib.c-1595- data->fib_nb.notifier_call = nsim_fib_event_nb;
drivers/net/netdevsim/fib.c:1596: err = register_fib_notifier(devlink_net(devlink), &data->fib_nb,
drivers/net/netdevsim/fib.c-1597- nsim_fib_dump_inconsistent, extack);
--
drivers/net/netdevsim/fib.c=1644=void nsim_fib_destroy(struct devlink *devlink, struct nsim_fib_data *data)
--
drivers/net/netdevsim/fib.c-1655- NSIM_RESOURCE_IPV4_FIB);
drivers/net/netdevsim/fib.c:1656: unregister_fib_notifier(devlink_net(devlink), &data->fib_nb);
drivers/net/netdevsim/fib.c-1657- unregister_nexthop_notifier(devlink_net(devlink), &data->nexthop_nb);
--
include/net/fib_notifier.h=41=int call_fib_notifiers(struct net *net, enum fib_event_type event_type,
include/net/fib_notifier.h-42- struct fib_notifier_info *info);
include/net/fib_notifier.h:43:int register_fib_notifier(struct net *net, struct notifier_block *nb,
include/net/fib_notifier.h-44- void (*cb)(struct notifier_block *nb),
include/net/fib_notifier.h-45- struct netlink_ext_ack *extack);
include/net/fib_notifier.h:46:int unregister_fib_notifier(struct net *net, struct notifier_block *nb);
include/net/fib_notifier.h-47-struct fib_notifier_ops *
--
net/core/fib_notifier.c=82=static bool fib_dump_is_consistent(struct net *net, struct notifier_block *nb,
--
net/core/fib_notifier.c-97-#define FIB_DUMP_MAX_RETRIES 5
net/core/fib_notifier.c:98:int register_fib_notifier(struct net *net, struct notifier_block *nb,
net/core/fib_notifier.c-99- void (*cb)(struct notifier_block *nb),
--
net/core/fib_notifier.c-117-}
net/core/fib_notifier.c:118:EXPORT_SYMBOL(register_fib_notifier);
net/core/fib_notifier.c-119-
net/core/fib_notifier.c:120:int unregister_fib_notifier(struct net *net, struct notifier_block *nb)
net/core/fib_notifier.c-121-{
--
net/core/fib_notifier.c-125-}
net/core/fib_notifier.c:126:EXPORT_SYMBOL(unregister_fib_notifier);
net/core/fib_notifier.c-127-
]
|
| 82/2 |
2026/05/26 14:41 |
llm |
expert |
0m
Model:
Tokens:
input: 63180
output: 43
thoughts: 348
Thoughts: **Analyzing the Privilege Escalation Vector**
Okay, so `register_fib_notifier` only gets called during driver initialization or a `devlink_reload` operation. The question is, can an unprivileged user somehow trigger either of these? Loading a module definitely requires root privileges, so that's out. What about `devlink_reload`? We already established that `DEVLINK_CMD_RELOAD` requires `GENL_ADMIN_PERM`, which translates to `CAP_NET_ADMIN`. So that seems blocked as well.
But hold on⦠what about `devlink_pernet_pre_exit`? That's called when a network namespace is destroyed. If a devlink instance is within that namespace, it will be reloaded into `init_net`. Hmmm... could an unprivileged user create a network namespace, move a devlink instance into it, and then destroy that namespace to trigger a reload into `init_net`?
To move a devlink instance, I'd need to use `DEVLINK_CMD_RELOAD` with `DEVLINK_ATTR_NETNS_FD` or something similar. Ah, but wait! Moving a devlink instance to a new netns requires `CAP_NET_ADMIN` *in the current* netns where the devlink instance resides, which is initially `init_net`. So, an unprivileged user can't move a devlink instance from `init_net` into their own netns. That means an unprivileged user can't have a devlink instance in their netns to begin with.
Okay, let's explore a different angle. What if they could create a `netdevsim` instance inside their netns? Can an unprivileged user even do that? Let's dig into the `netdevsim` device creation process and see.
|
| 83/2 |
2026/05/26 14:41 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:drivers/net/netdevsim/bus.c Name:nsim_bus_dev_new]
Results: map[SourceCode: 448: static struct nsim_bus_dev *
449: nsim_bus_dev_new(unsigned int id, unsigned int port_count, unsigned int num_queues)
450: {
451: struct nsim_bus_dev *nsim_bus_dev;
452: int err;
453:
454: nsim_bus_dev = kzalloc_obj(*nsim_bus_dev);
455: if (!nsim_bus_dev)
456: return ERR_PTR(-ENOMEM);
457:
458: err = ida_alloc_range(&nsim_bus_dev_ids, id, id, GFP_KERNEL);
459: if (err < 0)
460: goto err_nsim_bus_dev_free;
461: nsim_bus_dev->dev.id = err;
462: nsim_bus_dev->dev.bus = &nsim_bus;
463: nsim_bus_dev->dev.type = &nsim_bus_dev_type;
464: nsim_bus_dev->port_count = port_count;
465: nsim_bus_dev->num_queues = num_queues;
466: nsim_bus_dev->initial_net = current->nsproxy->net_ns;
467: nsim_bus_dev->max_vfs = NSIM_BUS_DEV_MAX_VFS;
468: /* Disallow using nsim_bus_dev */
469: smp_store_release(&nsim_bus_dev->init, false);
470:
471: err = device_register(&nsim_bus_dev->dev);
472: if (err)
473: goto err_nsim_bus_dev_id_free;
474:
475: return nsim_bus_dev;
476:
477: err_nsim_bus_dev_id_free:
478: ida_free(&nsim_bus_dev_ids, nsim_bus_dev->dev.id);
479: put_device(&nsim_bus_dev->dev);
480: nsim_bus_dev = NULL;
481: err_nsim_bus_dev_free:
482: kfree(nsim_bus_dev);
483: return ERR_PTR(err);
484: }
SourceFile:drivers/net/netdevsim/bus.c]
|
| 84/2 |
2026/05/26 14:41 |
llm |
expert |
0m
Model:
Tokens:
input: 64673
output: 56
thoughts: 14
Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience, and written in the first person:
**Analyzing the Call Stack: `nsim_bus_dev_new`**
Alright, so the immediate question that pops up is: Where and how is `nsim_bus_dev_new` being invoked? I need to understand the call graph, the context in which this function is operating. Am I dealing with a static call, a dynamically linked library, some sort of callback mechanism? Whatβs the calling context β is it during initialization, interrupt handling, or some ongoing device discovery process? Knowing the caller is fundamental to tracing the logic and potential issues. I need to figure out the entry point and the path it takes to get to `nsim_bus_dev_new`. Is there a common code path where I can insert logging? The goal here isn't just to *find* the calls, but to understand the *why* and the *how* of them. The purpose of understanding its call is to know: How is this related to some process and its overall state? Does it lead to problems? Is it used in the way I expect?
|
| 85/2 |
2026/05/26 14:41 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:drivers/net/netdevsim/bus.c IncludeSnippetLines:0 Name:nsim_bus_dev_new SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:new_device_store SourceFile:drivers/net/netdevsim/bus.c SourceLine:191]] TruncatedOutput:false]
|
| 86/2 |
2026/05/26 14:41 |
llm |
expert |
0m
Model:
Tokens:
input: 65064
output: 40
thoughts: 11
Thoughts: Okay, here's my interpretation of that:
**My Internal Processing of `new_device_store` Call**
Alright, so the system's just called into this section from the `new_device_store` function. That's a key entry point, immediately signaling to me that this is where we're going to be dealing with the creation or initialization of a new device entity, or something related to a new device. Understanding the context is crucial. I need to now consider what that `new_device_store` function's role is, the inputs it provides, and what its expectations are regarding the process. I'm immediately running through the potential implications: Are we simply storing metadata, configuring hardware, allocating resources, registering a new device in some kind of management system, or something else entirely? The flow of execution needs to be mentally mapped out to determine what actions are expected of this entry point. There has to be more to this section. I need more information to continue.
|
| 87/2 |
2026/05/26 14:41 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:drivers/net/netdevsim/bus.c Name:new_device_store]
Results: map[SourceCode: 158: static ssize_t
159: new_device_store(const struct bus_type *bus, const char *buf, size_t count)
160: {
161: unsigned int id, port_count, num_queues;
162: struct nsim_bus_dev *nsim_bus_dev;
163: int err;
164:
165: err = sscanf(buf, "%u %u %u", &id, &port_count, &num_queues);
166: switch (err) {
167: case 1:
168: port_count = 1;
169: fallthrough;
170: case 2:
171: num_queues = 1;
172: fallthrough;
173: case 3:
174: if (id > INT_MAX) {
175: pr_err("Value of \"id\" is too big.\n");
176: return -EINVAL;
177: }
178: break;
179: default:
180: pr_err("Format for adding new device is \"id port_count num_queues\" (uint uint uint).\n");
181: return -EINVAL;
182: }
183:
184: mutex_lock(&nsim_bus_dev_list_lock);
185: /* Prevent to use resource before initialization. */
186: if (!smp_load_acquire(&nsim_bus_enable)) {
187: err = -EBUSY;
188: goto err;
189: }
190:
191: nsim_bus_dev = nsim_bus_dev_new(id, port_count, num_queues);
192: if (IS_ERR(nsim_bus_dev)) {
193: err = PTR_ERR(nsim_bus_dev);
194: goto err;
195: }
196:
197: refcount_inc(&nsim_bus_devs);
198: /* Allow using nsim_bus_dev */
199: smp_store_release(&nsim_bus_dev->init, true);
200:
201: list_add_tail(&nsim_bus_dev->list, &nsim_bus_dev_list);
202: mutex_unlock(&nsim_bus_dev_list_lock);
203:
204: return count;
205: err:
206: mutex_unlock(&nsim_bus_dev_list_lock);
207: return err;
208: }
SourceFile:drivers/net/netdevsim/bus.c]
|
| 88/2 |
2026/05/26 14:41 |
llm |
expert |
0m
Model:
Tokens:
input: 66121
output: 51
thoughts: 295
Thoughts: **Analyzing Potential Vulnerabilities in Network Device Creation**
Okay, so I'm looking at this `new_device_store` function in `sysfs`. Right, it's a store function for creating a new device. But sysfs is mounted read-only in user namespaces, which means an unprivileged user can't just write to `/sys/bus/netdevsim/new_device`. Thatβs good. But hold on, what if a user creates a network namespace first, then *root* creates the `netdevsim` device *within that new namespace*? If root does it, it's a privileged operation, so the unprivileged user isnβt directly triggering anything vulnerable. So, no direct exploit path there with `netdevsim`.
Now, what about the hardware drivers like `mlxsw` and `prestera`? These are loaded by root, or automatically when the hardware is detected. I know they register a fib notifier within `init_net` when they load. Could an unprivileged user somehow trigger `fib_net_dump` through these drivers? Not directly. `fib_net_dump` is only called during the `register_fib_notifier` process, which happens during the driver initialization sequence. That's a one-time thing, so I can't see an unprivileged user being able to poke that process.
Okay, so is there *any* other route to trigger `fib_net_dump`? What about network namespace creation itself? Does creating a new namespace trigger `fib_net_dump` for *existing* fib notifiers? I need to check the `fib_notifier_ops` structure and the `fib_net_dump` code specifically to be sure. Thatβs where I need to focus my attention now.
|
| 89/2 |
2026/05/26 14:41 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/core/fib_notifier.c IncludeSnippetLines:0 Name:fib_net_dump SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:register_fib_notifier SourceFile:net/core/fib_notifier.c SourceLine:108]] TruncatedOutput:false]
|
| 90/2 |
2026/05/26 14:41 |
llm |
expert |
0m
Model:
Tokens:
input: 66914
output: 38
thoughts: 425
Thoughts: **Analyzing `register_fib_notifier` Usage for Privilege Escalation**
Okay, so I'm trying to figure out if there's any way an unprivileged user can leverage `register_fib_notifier` for a security issue, potentially privilege escalation. Initially, I thought that `fib_net_dump` was the only function that called `register_fib_notifier`, which would simplify things. But then I realized that's not the case. It's called by multiple processes.
The key question is, are there scenarios where a driver registers a FIB notifier per network namespace? I immediately thought of `netdevsim`. If a `devlink` instance, like the one from `netdevsim`, is moved to a different network namespace, does it register a new FIB notifier in the new namespace? I traced it, and yes, it does! Moving a `devlink` instance calls `devlink_reload`, which goes through `nsim_dev_reload_up`, `nsim_fib_create`, and then finally `register_fib_notifier`. Excellent.
However, moving a `devlink` instance requires `CAP_NET_ADMIN` in the *source* namespace. So, I figured, an unprivileged user is out of luck there.
Then, I shifted my focus to `mlx5`. Does `mlx5` present a possible vulnerability? I saw that `mlx5e_tc_tun_init` calls `register_fib_notifier`. If an unprivileged user creates a new network namespace, and somehow moves the `mlx5` netdev into it, would that trigger `mlx5e_tc_tun_init`? I realized that `mlx5e_tc_tun_init` is called during profile initialization, not when the netdev is moved between namespaces.
So, okay, if a netdev is moved to a new namespace, does it register a new fib notifier? Let's check if mlx5 registers a fib notifier when moved. And like before, moving a netdev generally needs `CAP_NET_ADMIN` in the current namespace. Thus, moving a *physical* netdev like `mlx5` is also a no-go for unprivileged users.
Alright, what other calls exist? Let's look at other possible callers of `register_fib_notifier`. What about `ip_fib_net_init`? Turns out it registers `fib_notifier_ops`. Now I need to check `fib_notifier_ops` for more details.
|
| 91/2 |
2026/05/26 14:41 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/ipv4/fib_notifier.c Name:fib4_dump]
Results: map[SourceCode: 36: static int fib4_dump(struct net *net, struct notifier_block *nb,
37: struct netlink_ext_ack *extack)
38: {
39: int err;
40:
41: err = fib4_rules_dump(net, nb, extack);
42: if (err)
43: return err;
44:
45: return fib_notify(net, nb, extack);
46: }
SourceFile:net/ipv4/fib_notifier.c]
|
| 92/2 |
2026/05/26 14:41 |
llm |
expert |
0m
Model:
Tokens:
input: 68076
output: 44
thoughts: 174
Thoughts: **Analyzing the FIB Notification Trigger**
Okay, so `fib4_dump` is calling `fib_notify`, which then iterates through all the tables via `fib_table_notify`. This is exactly the same path taken by `fib_net_dump`, which seems pretty significant. That suggests `fib_net_dump` is *only* called when a new FIB notifier is registered. Now, the key is where these notifiers are registered. It seems like the registration of a `fib_notifier` is exclusively handled by device drivers.
Given that unprivileged users cannot load drivers, nor create `netdevsim` devices, or move devlink instances, it looks like they are fundamentally blocked from triggering this `register_fib_notifier` call. That means this entire potential vulnerability is *definitely* inaccessible from unprivileged processes, and certainly from within user namespaces, which is great news from a security perspective.
Wait a secondβ¦ I need to be absolutely sure. Are there *any* other ways to bump the reference count on a `fib_info` after itβs been freed? If a refcount is at zero, how does the system attempt to hold onto it? Let's quickly review the `nsim_fib4_rt_create` function to check.
|
| 93/2 |
2026/05/26 14:41 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:drivers/net/netdevsim/fib.c Name:nsim_fib4_rt_create]
Results: map[SourceCode: 274: static struct nsim_fib4_rt *
275: nsim_fib4_rt_create(struct nsim_fib_data *data,
276: struct fib_entry_notifier_info *fen_info)
277: {
278: struct nsim_fib4_rt *fib4_rt;
279:
280: fib4_rt = kzalloc_obj(*fib4_rt);
281: if (!fib4_rt)
282: return NULL;
283:
284: nsim_fib_rt_init(data, &fib4_rt->common, &fen_info->dst, sizeof(u32),
285: fen_info->dst_len, AF_INET, fen_info->tb_id);
286:
287: fib4_rt->fi = fen_info->fi;
288: fib_info_hold(fib4_rt->fi);
289: fib4_rt->dscp = fen_info->dscp;
290: fib4_rt->type = fen_info->type;
291:
292: return fib4_rt;
293: }
SourceFile:drivers/net/netdevsim/fib.c]
|
| 94/2 |
2026/05/26 14:41 |
llm |
expert |
0m
Model:
Tokens:
input: 68932
output: 51
thoughts: 532
Thoughts: **Analyzing the Route Table Event Handling**
Okay, let's break this down. My understanding is that `nsim_fib4_rt_create` gets called by `nsim_fib4_rt_add`, which in turn is called by `nsim_fib4_event`. Then we trace up the chain: `nsim_fib4_event` is called by `nsim_fib_event_work`, which is a workqueue function. This workqueue is processing events that are coming from `fib_event_queue`. Those events are triggered by `nsim_fib_event_nb`, the notifier callback. The issue centers around this callback, `nsim_fib_event_nb`, specifically when it is called with `FIB_EVENT_ENTRY_ADD` or `FIB_EVENT_ENTRY_REPLACE`, indicating a route is being added (or replaced).
The initial assumption was that `fib_info_hold` in `nsim_fib4_rt_create` was safe during *normal* route addition because the `fib_info` structure is alive and has a positive reference count at that point. The bug, as originally described, seemed to be specific to `nsim_fib4_prepare_event`, called synchronously from `nsim_fib_event_nb`. `nsim_fib4_prepare_event` calls `fib_info_hold`, which is only problematic during `fib_net_dump` because there is a race condition with route deletion.
But, hold on. This leads me to a more general concern. What if there's a *concurrent* route deletion happening while a *normal* route addition is underway? Let's trace through the normal route addition path more carefully. During a normal route addition, `fib_table_insert` is called. This function, I know, creates a new `fib_info` (or finds an existing one) and inserts it into the routing trie, and then calls `fib_notify_alias_add` or `fib_notify_alias_replace`. Ah, wait! `fib_notify_alias_add` (and replace) calls `call_fib_entry_notifiers`. So, at that point, the `fib_info` should *still* be alive and have a positive refcount.
The critical question now becomes: how can a concurrent deletion affect this? If a concurrent route deletion *does* occur, it will call `fib_table_delete`. However, I believe that `fib_table_insert` and `fib_table_delete` are *both* protected by the `RTNL` (Routing Table Netlink) lock. Therefore, they *shouldn't* be able to run concurrently. The only case we have established that can cause a race condition is with `fib_net_dump` and `fib_table_delete`. This is likely because `fib_net_dump` does NOT hold the `RTNL` lock, and so *can* race with the deletion. I need to confirm that `fib_net_dump` doesn't hold `RTNL`. That seems like the likely root cause.
|
| 95/2 |
2026/05/26 14:42 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/core/fib_notifier.c IncludeSnippetLines:0 Name:register_fib_notifier SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:nsim_fib_create SourceFile:drivers/net/netdevsim/fib.c SourceLine:1596]] TruncatedOutput:false]
|
| 96/2 |
2026/05/26 14:42 |
llm |
expert |
0m
Model:
Tokens:
input: 70258
output: 43
thoughts: 53
Thoughts: **Investigating the `register_fib_notifier` Call**
Okay, so I'm tracing through this code and something's bugging me. I see that `register_fib_notifier` is called from within `nsim_fib_create`. That's the key. Now, my immediate concern is whether or not `RTNL` (the Routing Netlink Lock) is being held when `nsim_fib_create` is invoked. Thatβs crucial for concurrency, and it's a potential area for race conditions, especially with those notifier chains.
Let's dig deeper and see. I need to quickly determine the calling context of `nsim_fib_create`. I'm going to jump over to where `nsim_dev_reload_up` is called. That might give me a clearer picture of whether `RTNL` is acquired before the call, or if it has any relation to RTNL's state. I need to understand the lock hierarchy here. This is classic debugging, peeling back the layers to see how these functions interact and where those potential concurrency pitfalls might be lurking. Time to examine `nsim_dev_reload_up`.
|
| 97/2 |
2026/05/26 14:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:drivers/net/netdevsim/dev.c Name:nsim_dev_reload_up]
Results: map[SourceCode:1042: static int nsim_dev_reload_up(struct devlink *devlink, enum devlink_reload_action action,
1043: enum devlink_reload_limit limit, u32 *actions_performed,
1044: struct netlink_ext_ack *extack)
1045: {
1046: struct nsim_dev *nsim_dev = devlink_priv(devlink);
1047:
1048: if (nsim_dev->fail_reload) {
1049: /* For testing purposes, user set debugfs fail_reload
1050: * value to true. Fail right away.
1051: */
1052: NL_SET_ERR_MSG_MOD(extack, "User setup the reload to fail for testing purposes");
1053: return -EINVAL;
1054: }
1055:
1056: *actions_performed = BIT(DEVLINK_RELOAD_ACTION_DRIVER_REINIT);
1057:
1058: return nsim_dev_reload_create(nsim_dev, extack);
1059: }
SourceFile:drivers/net/netdevsim/dev.c]
|
| 98/2 |
2026/05/26 14:42 |
llm |
expert |
0m
Model:
Tokens:
input: 70941
output: 41
thoughts: 62
Thoughts: **My Analysis of the `devlink_reload` Function**
Okay, so `devlink_reload`... where are we calling this function from, exactly? Ah, right, it seems to be triggered by either `devlink_nl_reload_doit` or `devlink_pernet_pre_exit`. That's the entry point, but I need to understand the implications of *how* it's called. My primary concern here is whether `devlink_reload` is potentially being called while holding the `RTNL` (Routing Table Next Layer) lock. That's a critical question because holding the `RTNL` lock has some serious concurrency implications if we're not careful.
Let's dive right into `devlink_nl_reload_doit` to see if we can deduce whether it is. I'm going to look for any calls to functions that might grab that lock, or any clear indications of its presence in the call stack. Understanding the locking context of this function is going to be key to debugging any potential issues downstream.
|
| 99/2 |
2026/05/26 14:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/devlink/dev.c Name:devlink_nl_reload_doit]
Results: map[SourceCode: 520: int devlink_nl_reload_doit(struct sk_buff *skb, struct genl_info *info)
521: {
522: struct devlink *devlink = info->user_ptr[0];
523: enum devlink_reload_action action;
524: enum devlink_reload_limit limit;
525: struct net *dest_net = NULL;
526: u32 actions_performed;
527: int err;
528:
529: err = devlink_resources_validate(devlink, NULL, info);
530: if (err) {
531: NL_SET_ERR_MSG(info->extack, "resources size validation failed");
532: return err;
533: }
534:
535: action = nla_get_u8_default(info->attrs[DEVLINK_ATTR_RELOAD_ACTION],
536: DEVLINK_RELOAD_ACTION_DRIVER_REINIT);
537:
538: if (!devlink_reload_action_is_supported(devlink, action)) {
539: NL_SET_ERR_MSG(info->extack, "Requested reload action is not supported by the driver");
540: return -EOPNOTSUPP;
541: }
542:
543: limit = DEVLINK_RELOAD_LIMIT_UNSPEC;
544: if (info->attrs[DEVLINK_ATTR_RELOAD_LIMITS]) {
545: struct nla_bitfield32 limits;
546: u32 limits_selected;
547:
548: limits = nla_get_bitfield32(info->attrs[DEVLINK_ATTR_RELOAD_LIMITS]);
549: limits_selected = limits.value & limits.selector;
550: if (!limits_selected) {
551: NL_SET_ERR_MSG(info->extack, "Invalid limit selected");
552: return -EINVAL;
553: }
554: for (limit = 0 ; limit <= DEVLINK_RELOAD_LIMIT_MAX ; limit++)
555: if (limits_selected & BIT(limit))
556: break;
557: /* UAPI enables multiselection, but currently it is not used */
558: if (limits_selected != BIT(limit)) {
559: NL_SET_ERR_MSG(info->extack, "Multiselection of limit is not supported");
560: return -EOPNOTSUPP;
561: }
562: if (!devlink_reload_limit_is_supported(devlink, limit)) {
563: NL_SET_ERR_MSG(info->extack, "Requested limit is not supported by the driver");
564: return -EOPNOTSUPP;
565: }
566: if (devlink_reload_combination_is_invalid(action, limit)) {
567: NL_SET_ERR_MSG(info->extack, "Requested limit is invalid for this action");
568: return -EINVAL;
569: }
570: }
571: if (info->attrs[DEVLINK_ATTR_NETNS_PID] ||
572: info->attrs[DEVLINK_ATTR_NETNS_FD] ||
573: info->attrs[DEVLINK_ATTR_NETNS_ID]) {
574: dest_net = devlink_netns_get(skb, info);
575: if (IS_ERR(dest_net))
576: return PTR_ERR(dest_net);
577: if (!net_eq(dest_net, devlink_net(devlink)) &&
578: action != DEVLINK_RELOAD_ACTION_DRIVER_REINIT) {
579: NL_SET_ERR_MSG_MOD(info->extack,
580: "Changing namespace is only supported for reinit action");
581: return -EOPNOTSUPP;
582: }
583: }
584:
585: err = devlink_reload(devlink, dest_net, action, limit, &actions_performed, info->extack);
586:
587: if (dest_net)
588: put_net(dest_net);
589:
590: if (err)
591: return err;
592: /* For backward compatibility generate reply only if attributes used by user */
593: if (!info->attrs[DEVLINK_ATTR_RELOAD_ACTION] && !info->attrs[DEVLINK_ATTR_RELOAD_LIMITS])
594: return 0;
595:
596: return devlink_nl_reload_actions_performed_snd(devlink, actions_performed,
597: DEVLINK_CMD_RELOAD, info);
598: }
SourceFile:net/devlink/dev.c]
|
| 100/2 |
2026/05/26 14:42 |
llm |
expert |
0m
Model:
Tokens:
input: 72666
output: 41
thoughts: 61
Thoughts: Here's my understanding of the situation, given my expertise:
**Devlink Reload: Unpacking the Mechanism**
Okay, so `devlink_nl_reload_doit` is a genl callback, which means it's part of the generic netlink framework. The key thing I need to remember, from my experience, is that these callbacks *don't* automatically hold the `RTNL` (Routing Table Netlink) lock. Unless, of course, the specific operation declares it requires the lock within its registered operations.
To figure out how `devlink_nl_reload_doit` handles the lock, I've got to dig into the `devlink_nl_ops` structure. I need to specifically look for the entry related to `DEVLINK_CMD_RELOAD`. That's where I'll find the definitive answer as to whether this reload operation implicitly or explicitly claims the `RTNL` lock. I need to check the associated flags and methods. This will dictate my synchronization expectations.
|
| 101/2 |
2026/05/26 14:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/devlink/netlink_gen.c Name:devlink_nl_ops]
Results: map[SourceCode: 691: /* Ops table for devlink */
692: const struct genl_split_ops devlink_nl_ops[75] = {
693: {
694: .cmd = DEVLINK_CMD_GET,
695: .validate = GENL_DONT_VALIDATE_STRICT,
696: .pre_doit = devlink_nl_pre_doit,
697: .doit = devlink_nl_get_doit,
698: .post_doit = devlink_nl_post_doit,
699: .policy = devlink_get_nl_policy,
700: .maxattr = DEVLINK_ATTR_INDEX,
701: .flags = GENL_CMD_CAP_DO,
702: },
703: {
704: .cmd = DEVLINK_CMD_GET,
705: .validate = GENL_DONT_VALIDATE_DUMP,
706: .dumpit = devlink_nl_get_dumpit,
707: .flags = GENL_CMD_CAP_DUMP,
708: },
709: {
710: .cmd = DEVLINK_CMD_PORT_GET,
711: .validate = GENL_DONT_VALIDATE_STRICT,
712: .pre_doit = devlink_nl_pre_doit_port,
713: .doit = devlink_nl_port_get_doit,
714: .post_doit = devlink_nl_post_doit,
715: .policy = devlink_port_get_do_nl_policy,
716: .maxattr = DEVLINK_ATTR_INDEX,
717: .flags = GENL_CMD_CAP_DO,
718: },
719: {
720: .cmd = DEVLINK_CMD_PORT_GET,
721: .dumpit = devlink_nl_port_get_dumpit,
722: .policy = devlink_port_get_dump_nl_policy,
723: .maxattr = DEVLINK_ATTR_INDEX,
724: .flags = GENL_CMD_CAP_DUMP,
725: },
726: {
727: .cmd = DEVLINK_CMD_PORT_SET,
728: .validate = GENL_DONT_VALIDATE_STRICT,
729: .pre_doit = devlink_nl_pre_doit_port,
730: .doit = devlink_nl_port_set_doit,
731: .post_doit = devlink_nl_post_doit,
732: .policy = devlink_port_set_nl_policy,
733: .maxattr = DEVLINK_ATTR_INDEX,
734: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
735: },
736: {
737: .cmd = DEVLINK_CMD_PORT_NEW,
738: .validate = GENL_DONT_VALIDATE_STRICT,
739: .pre_doit = devlink_nl_pre_doit,
740: .doit = devlink_nl_port_new_doit,
741: .post_doit = devlink_nl_post_doit,
742: .policy = devlink_port_new_nl_policy,
743: .maxattr = DEVLINK_ATTR_INDEX,
744: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
745: },
746: {
747: .cmd = DEVLINK_CMD_PORT_DEL,
748: .validate = GENL_DONT_VALIDATE_STRICT,
749: .pre_doit = devlink_nl_pre_doit_port,
750: .doit = devlink_nl_port_del_doit,
751: .post_doit = devlink_nl_post_doit,
752: .policy = devlink_port_del_nl_policy,
753: .maxattr = DEVLINK_ATTR_INDEX,
754: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
755: },
756: {
757: .cmd = DEVLINK_CMD_PORT_SPLIT,
758: .validate = GENL_DONT_VALIDATE_STRICT,
759: .pre_doit = devlink_nl_pre_doit_port,
760: .doit = devlink_nl_port_split_doit,
761: .post_doit = devlink_nl_post_doit,
762: .policy = devlink_port_split_nl_policy,
763: .maxattr = DEVLINK_ATTR_INDEX,
764: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
765: },
766: {
767: .cmd = DEVLINK_CMD_PORT_UNSPLIT,
768: .validate = GENL_DONT_VALIDATE_STRICT,
769: .pre_doit = devlink_nl_pre_doit_port,
770: .doit = devlink_nl_port_unsplit_doit,
771: .post_doit = devlink_nl_post_doit,
772: .policy = devlink_port_unsplit_nl_policy,
773: .maxattr = DEVLINK_ATTR_INDEX,
774: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
775: },
776: {
777: .cmd = DEVLINK_CMD_SB_GET,
778: .validate = GENL_DONT_VALIDATE_STRICT,
779: .pre_doit = devlink_nl_pre_doit,
780: .doit = devlink_nl_sb_get_doit,
781: .post_doit = devlink_nl_post_doit,
782: .policy = devlink_sb_get_do_nl_policy,
783: .maxattr = DEVLINK_ATTR_INDEX,
784: .flags = GENL_CMD_CAP_DO,
785: },
786: {
787: .cmd = DEVLINK_CMD_SB_GET,
788: .dumpit = devlink_nl_sb_get_dumpit,
789: .policy = devlink_sb_get_dump_nl_policy,
790: .maxattr = DEVLINK_ATTR_INDEX,
791: .flags = GENL_CMD_CAP_DUMP,
792: },
793: {
794: .cmd = DEVLINK_CMD_SB_POOL_GET,
795: .validate = GENL_DONT_VALIDATE_STRICT,
796: .pre_doit = devlink_nl_pre_doit,
797: .doit = devlink_nl_sb_pool_get_doit,
798: .post_doit = devlink_nl_post_doit,
799: .policy = devlink_sb_pool_get_do_nl_policy,
800: .maxattr = DEVLINK_ATTR_INDEX,
801: .flags = GENL_CMD_CAP_DO,
802: },
803: {
804: .cmd = DEVLINK_CMD_SB_POOL_GET,
805: .dumpit = devlink_nl_sb_pool_get_dumpit,
806: .policy = devlink_sb_pool_get_dump_nl_policy,
807: .maxattr = DEVLINK_ATTR_INDEX,
808: .flags = GENL_CMD_CAP_DUMP,
809: },
810: {
811: .cmd = DEVLINK_CMD_SB_POOL_SET,
812: .validate = GENL_DONT_VALIDATE_STRICT,
813: .pre_doit = devlink_nl_pre_doit,
814: .doit = devlink_nl_sb_pool_set_doit,
815: .post_doit = devlink_nl_post_doit,
816: .policy = devlink_sb_pool_set_nl_policy,
817: .maxattr = DEVLINK_ATTR_INDEX,
818: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
819: },
820: {
821: .cmd = DEVLINK_CMD_SB_PORT_POOL_GET,
822: .validate = GENL_DONT_VALIDATE_STRICT,
823: .pre_doit = devlink_nl_pre_doit_port,
824: .doit = devlink_nl_sb_port_pool_get_doit,
825: .post_doit = devlink_nl_post_doit,
826: .policy = devlink_sb_port_pool_get_do_nl_policy,
827: .maxattr = DEVLINK_ATTR_INDEX,
828: .flags = GENL_CMD_CAP_DO,
829: },
830: {
831: .cmd = DEVLINK_CMD_SB_PORT_POOL_GET,
832: .dumpit = devlink_nl_sb_port_pool_get_dumpit,
833: .policy = devlink_sb_port_pool_get_dump_nl_policy,
834: .maxattr = DEVLINK_ATTR_INDEX,
835: .flags = GENL_CMD_CAP_DUMP,
836: },
837: {
838: .cmd = DEVLINK_CMD_SB_PORT_POOL_SET,
839: .validate = GENL_DONT_VALIDATE_STRICT,
840: .pre_doit = devlink_nl_pre_doit_port,
841: .doit = devlink_nl_sb_port_pool_set_doit,
842: .post_doit = devlink_nl_post_doit,
843: .policy = devlink_sb_port_pool_set_nl_policy,
844: .maxattr = DEVLINK_ATTR_INDEX,
845: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
846: },
847: {
848: .cmd = DEVLINK_CMD_SB_TC_POOL_BIND_GET,
849: .validate = GENL_DONT_VALIDATE_STRICT,
850: .pre_doit = devlink_nl_pre_doit_port,
851: .doit = devlink_nl_sb_tc_pool_bind_get_doit,
852: .post_doit = devlink_nl_post_doit,
853: .policy = devlink_sb_tc_pool_bind_get_do_nl_policy,
854: .maxattr = DEVLINK_ATTR_INDEX,
855: .flags = GENL_CMD_CAP_DO,
856: },
857: {
858: .cmd = DEVLINK_CMD_SB_TC_POOL_BIND_GET,
859: .dumpit = devlink_nl_sb_tc_pool_bind_get_dumpit,
860: .policy = devlink_sb_tc_pool_bind_get_dump_nl_policy,
861: .maxattr = DEVLINK_ATTR_INDEX,
862: .flags = GENL_CMD_CAP_DUMP,
863: },
864: {
865: .cmd = DEVLINK_CMD_SB_TC_POOL_BIND_SET,
866: .validate = GENL_DONT_VALIDATE_STRICT,
867: .pre_doit = devlink_nl_pre_doit_port,
868: .doit = devlink_nl_sb_tc_pool_bind_set_doit,
869: .post_doit = devlink_nl_post_doit,
870: .policy = devlink_sb_tc_pool_bind_set_nl_policy,
871: .maxattr = DEVLINK_ATTR_INDEX,
872: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
873: },
874: {
875: .cmd = DEVLINK_CMD_SB_OCC_SNAPSHOT,
876: .validate = GENL_DONT_VALIDATE_STRICT,
877: .pre_doit = devlink_nl_pre_doit,
878: .doit = devlink_nl_sb_occ_snapshot_doit,
879: .post_doit = devlink_nl_post_doit,
880: .policy = devlink_sb_occ_snapshot_nl_policy,
881: .maxattr = DEVLINK_ATTR_INDEX,
882: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
883: },
884: {
885: .cmd = DEVLINK_CMD_SB_OCC_MAX_CLEAR,
886: .validate = GENL_DONT_VALIDATE_STRICT,
887: .pre_doit = devlink_nl_pre_doit,
888: .doit = devlink_nl_sb_occ_max_clear_doit,
889: .post_doit = devlink_nl_post_doit,
890: .policy = devlink_sb_occ_max_clear_nl_policy,
891: .maxattr = DEVLINK_ATTR_INDEX,
892: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
893: },
894: {
895: .cmd = DEVLINK_CMD_ESWITCH_GET,
896: .validate = GENL_DONT_VALIDATE_STRICT,
897: .pre_doit = devlink_nl_pre_doit,
898: .doit = devlink_nl_eswitch_get_doit,
899: .post_doit = devlink_nl_post_doit,
900: .policy = devlink_eswitch_get_nl_policy,
901: .maxattr = DEVLINK_ATTR_INDEX,
902: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
903: },
904: {
905: .cmd = DEVLINK_CMD_ESWITCH_SET,
906: .validate = GENL_DONT_VALIDATE_STRICT,
907: .pre_doit = devlink_nl_pre_doit,
908: .doit = devlink_nl_eswitch_set_doit,
909: .post_doit = devlink_nl_post_doit,
910: .policy = devlink_eswitch_set_nl_policy,
911: .maxattr = DEVLINK_ATTR_INDEX,
912: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
913: },
914: {
915: .cmd = DEVLINK_CMD_DPIPE_TABLE_GET,
916: .validate = GENL_DONT_VALIDATE_STRICT,
917: .pre_doit = devlink_nl_pre_doit,
918: .doit = devlink_nl_dpipe_table_get_doit,
919: .post_doit = devlink_nl_post_doit,
920: .policy = devlink_dpipe_table_get_nl_policy,
921: .maxattr = DEVLINK_ATTR_INDEX,
922: .flags = GENL_CMD_CAP_DO,
923: },
924: {
925: .cmd = DEVLINK_CMD_DPIPE_ENTRIES_GET,
926: .validate = GENL_DONT_VALIDATE_STRICT,
927: .pre_doit = devlink_nl_pre_doit,
928: .doit = devlink_nl_dpipe_entries_get_doit,
929: .post_doit = devlink_nl_post_doit,
930: .policy = devlink_dpipe_entries_get_nl_policy,
931: .maxattr = DEVLINK_ATTR_INDEX,
932: .flags = GENL_CMD_CAP_DO,
933: },
934: {
935: .cmd = DEVLINK_CMD_DPIPE_HEADERS_GET,
936: .validate = GENL_DONT_VALIDATE_STRICT,
937: .pre_doit = devlink_nl_pre_doit,
938: .doit = devlink_nl_dpipe_headers_get_doit,
939: .post_doit = devlink_nl_post_doit,
940: .policy = devlink_dpipe_headers_get_nl_policy,
941: .maxattr = DEVLINK_ATTR_INDEX,
942: .flags = GENL_CMD_CAP_DO,
943: },
944: {
945: .cmd = DEVLINK_CMD_DPIPE_TABLE_COUNTERS_SET,
946: .validate = GENL_DONT_VALIDATE_STRICT,
947: .pre_doit = devlink_nl_pre_doit,
948: .doit = devlink_nl_dpipe_table_counters_set_doit,
949: .post_doit = devlink_nl_post_doit,
950: .policy = devlink_dpipe_table_counters_set_nl_policy,
951: .maxattr = DEVLINK_ATTR_INDEX,
952: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
953: },
954: {
955: .cmd = DEVLINK_CMD_RESOURCE_SET,
956: .validate = GENL_DONT_VALIDATE_STRICT,
957: .pre_doit = devlink_nl_pre_doit,
958: .doit = devlink_nl_resource_set_doit,
959: .post_doit = devlink_nl_post_doit,
960: .policy = devlink_resource_set_nl_policy,
961: .maxattr = DEVLINK_ATTR_INDEX,
962: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
963: },
964: {
965: .cmd = DEVLINK_CMD_RESOURCE_DUMP,
966: .validate = GENL_DONT_VALIDATE_STRICT,
967: .pre_doit = devlink_nl_pre_doit_port_optional,
968: .doit = devlink_nl_resource_dump_doit,
969: .post_doit = devlink_nl_post_doit,
970: .policy = devlink_resource_dump_do_nl_policy,
971: .maxattr = DEVLINK_ATTR_INDEX,
972: .flags = GENL_CMD_CAP_DO,
973: },
974: {
975: .cmd = DEVLINK_CMD_RESOURCE_DUMP,
976: .dumpit = devlink_nl_resource_dump_dumpit,
977: .policy = devlink_resource_dump_dump_nl_policy,
978: .maxattr = DEVLINK_ATTR_RESOURCE_SCOPE_MASK,
979: .flags = GENL_CMD_CAP_DUMP,
980: },
981: {
982: .cmd = DEVLINK_CMD_RELOAD,
983: .validate = GENL_DONT_VALIDATE_STRICT,
984: .pre_doit = devlink_nl_pre_doit_dev_lock,
985: .doit = devlink_nl_reload_doit,
986: .post_doit = devlink_nl_post_doit_dev_lock,
987: .policy = devlink_reload_nl_policy,
988: .maxattr = DEVLINK_ATTR_INDEX,
989: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
990: },
991: {
992: .cmd = DEVLINK_CMD_PARAM_GET,
993: .validate = GENL_DONT_VALIDATE_STRICT,
994: .pre_doit = devlink_nl_pre_doit,
995: .doit = devlink_nl_param_get_doit,
996: .post_doit = devlink_nl_post_doit,
997: .policy = devlink_param_get_do_nl_policy,
998: .maxattr = DEVLINK_ATTR_INDEX,
999: .flags = GENL_CMD_CAP_DO,
1000: },
1001: {
1002: .cmd = DEVLINK_CMD_PARAM_GET,
1003: .dumpit = devlink_nl_param_get_dumpit,
1004: .policy = devlink_param_get_dump_nl_policy,
1005: .maxattr = DEVLINK_ATTR_INDEX,
1006: .flags = GENL_CMD_CAP_DUMP,
1007: },
1008: {
1009: .cmd = DEVLINK_CMD_PARAM_SET,
1010: .validate = GENL_DONT_VALIDATE_STRICT,
1011: .pre_doit = devlink_nl_pre_doit,
1012: .doit = devlink_nl_param_set_doit,
1013: .post_doit = devlink_nl_post_doit,
1014: .policy = devlink_param_set_nl_policy,
1015: .maxattr = DEVLINK_ATTR_INDEX,
1016: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1017: },
1018: {
1019: .cmd = DEVLINK_CMD_REGION_GET,
1020: .validate = GENL_DONT_VALIDATE_STRICT,
1021: .pre_doit = devlink_nl_pre_doit_port_optional,
1022: .doit = devlink_nl_region_get_doit,
1023: .post_doit = devlink_nl_post_doit,
1024: .policy = devlink_region_get_do_nl_policy,
1025: .maxattr = DEVLINK_ATTR_INDEX,
1026: .flags = GENL_CMD_CAP_DO,
1027: },
1028: {
1029: .cmd = DEVLINK_CMD_REGION_GET,
1030: .dumpit = devlink_nl_region_get_dumpit,
1031: .policy = devlink_region_get_dump_nl_policy,
1032: .maxattr = DEVLINK_ATTR_INDEX,
1033: .flags = GENL_CMD_CAP_DUMP,
1034: },
1035: {
1036: .cmd = DEVLINK_CMD_REGION_NEW,
1037: .validate = GENL_DONT_VALIDATE_STRICT,
1038: .pre_doit = devlink_nl_pre_doit_port_optional,
1039: .doit = devlink_nl_region_new_doit,
1040: .post_doit = devlink_nl_post_doit,
1041: .policy = devlink_region_new_nl_policy,
1042: .maxattr = DEVLINK_ATTR_INDEX,
1043: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1044: },
1045: {
1046: .cmd = DEVLINK_CMD_REGION_DEL,
1047: .validate = GENL_DONT_VALIDATE_STRICT,
1048: .pre_doit = devlink_nl_pre_doit_port_optional,
1049: .doit = devlink_nl_region_del_doit,
1050: .post_doit = devlink_nl_post_doit,
1051: .policy = devlink_region_del_nl_policy,
1052: .maxattr = DEVLINK_ATTR_INDEX,
1053: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1054: },
1055: {
1056: .cmd = DEVLINK_CMD_REGION_READ,
1057: .validate = GENL_DONT_VALIDATE_DUMP_STRICT,
1058: .dumpit = devlink_nl_region_read_dumpit,
1059: .policy = devlink_region_read_nl_policy,
1060: .maxattr = DEVLINK_ATTR_INDEX,
1061: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DUMP,
1062: },
1063: {
1064: .cmd = DEVLINK_CMD_PORT_PARAM_GET,
1065: .validate = GENL_DONT_VALIDATE_STRICT,
1066: .pre_doit = devlink_nl_pre_doit_port,
1067: .doit = devlink_nl_port_param_get_doit,
1068: .post_doit = devlink_nl_post_doit,
1069: .policy = devlink_port_param_get_nl_policy,
1070: .maxattr = DEVLINK_ATTR_INDEX,
1071: .flags = GENL_CMD_CAP_DO,
1072: },
1073: {
1074: .cmd = DEVLINK_CMD_PORT_PARAM_GET,
1075: .validate = GENL_DONT_VALIDATE_DUMP_STRICT,
1076: .dumpit = devlink_nl_port_param_get_dumpit,
1077: .flags = GENL_CMD_CAP_DUMP,
1078: },
1079: {
1080: .cmd = DEVLINK_CMD_PORT_PARAM_SET,
1081: .validate = GENL_DONT_VALIDATE_STRICT,
1082: .pre_doit = devlink_nl_pre_doit_port,
1083: .doit = devlink_nl_port_param_set_doit,
1084: .post_doit = devlink_nl_post_doit,
1085: .policy = devlink_port_param_set_nl_policy,
1086: .maxattr = DEVLINK_ATTR_INDEX,
1087: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1088: },
1089: {
1090: .cmd = DEVLINK_CMD_INFO_GET,
1091: .validate = GENL_DONT_VALIDATE_STRICT,
1092: .pre_doit = devlink_nl_pre_doit,
1093: .doit = devlink_nl_info_get_doit,
1094: .post_doit = devlink_nl_post_doit,
1095: .policy = devlink_info_get_nl_policy,
1096: .maxattr = DEVLINK_ATTR_INDEX,
1097: .flags = GENL_CMD_CAP_DO,
1098: },
1099: {
1100: .cmd = DEVLINK_CMD_INFO_GET,
1101: .validate = GENL_DONT_VALIDATE_DUMP,
1102: .dumpit = devlink_nl_info_get_dumpit,
1103: .flags = GENL_CMD_CAP_DUMP,
1104: },
1105: {
1106: .cmd = DEVLINK_CMD_HEALTH_REPORTER_GET,
1107: .validate = GENL_DONT_VALIDATE_STRICT,
1108: .pre_doit = devlink_nl_pre_doit_port_optional,
1109: .doit = devlink_nl_health_reporter_get_doit,
1110: .post_doit = devlink_nl_post_doit,
1111: .policy = devlink_health_reporter_get_do_nl_policy,
1112: .maxattr = DEVLINK_ATTR_INDEX,
1113: .flags = GENL_CMD_CAP_DO,
1114: },
1115: {
1116: .cmd = DEVLINK_CMD_HEALTH_REPORTER_GET,
1117: .dumpit = devlink_nl_health_reporter_get_dumpit,
1118: .policy = devlink_health_reporter_get_dump_nl_policy,
1119: .maxattr = DEVLINK_ATTR_INDEX,
1120: .flags = GENL_CMD_CAP_DUMP,
1121: },
1122: {
1123: .cmd = DEVLINK_CMD_HEALTH_REPORTER_SET,
1124: .validate = GENL_DONT_VALIDATE_STRICT,
1125: .pre_doit = devlink_nl_pre_doit_port_optional,
1126: .doit = devlink_nl_health_reporter_set_doit,
1127: .post_doit = devlink_nl_post_doit,
1128: .policy = devlink_health_reporter_set_nl_policy,
1129: .maxattr = DEVLINK_ATTR_INDEX,
1130: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1131: },
1132: {
1133: .cmd = DEVLINK_CMD_HEALTH_REPORTER_RECOVER,
1134: .validate = GENL_DONT_VALIDATE_STRICT,
1135: .pre_doit = devlink_nl_pre_doit_port_optional,
1136: .doit = devlink_nl_health_reporter_recover_doit,
1137: .post_doit = devlink_nl_post_doit,
1138: .policy = devlink_health_reporter_recover_nl_policy,
1139: .maxattr = DEVLINK_ATTR_INDEX,
1140: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1141: },
1142: {
1143: .cmd = DEVLINK_CMD_HEALTH_REPORTER_DIAGNOSE,
1144: .validate = GENL_DONT_VALIDATE_STRICT,
1145: .pre_doit = devlink_nl_pre_doit_port_optional,
1146: .doit = devlink_nl_health_reporter_diagnose_doit,
1147: .post_doit = devlink_nl_post_doit,
1148: .policy = devlink_health_reporter_diagnose_nl_policy,
1149: .maxattr = DEVLINK_ATTR_INDEX,
1150: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1151: },
1152: {
1153: .cmd = DEVLINK_CMD_HEALTH_REPORTER_DUMP_GET,
1154: .validate = GENL_DONT_VALIDATE_DUMP_STRICT,
1155: .dumpit = devlink_nl_health_reporter_dump_get_dumpit,
1156: .policy = devlink_health_reporter_dump_get_nl_policy,
1157: .maxattr = DEVLINK_ATTR_INDEX,
1158: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DUMP,
1159: },
1160: {
1161: .cmd = DEVLINK_CMD_HEALTH_REPORTER_DUMP_CLEAR,
1162: .validate = GENL_DONT_VALIDATE_STRICT,
1163: .pre_doit = devlink_nl_pre_doit_port_optional,
1164: .doit = devlink_nl_health_reporter_dump_clear_doit,
1165: .post_doit = devlink_nl_post_doit,
1166: .policy = devlink_health_reporter_dump_clear_nl_policy,
1167: .maxattr = DEVLINK_ATTR_INDEX,
1168: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1169: },
1170: {
1171: .cmd = DEVLINK_CMD_FLASH_UPDATE,
1172: .validate = GENL_DONT_VALIDATE_STRICT,
1173: .pre_doit = devlink_nl_pre_doit,
1174: .doit = devlink_nl_flash_update_doit,
1175: .post_doit = devlink_nl_post_doit,
1176: .policy = devlink_flash_update_nl_policy,
1177: .maxattr = DEVLINK_ATTR_INDEX,
1178: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1179: },
1180: {
1181: .cmd = DEVLINK_CMD_TRAP_GET,
1182: .validate = GENL_DONT_VALIDATE_STRICT,
1183: .pre_doit = devlink_nl_pre_doit,
1184: .doit = devlink_nl_trap_get_doit,
1185: .post_doit = devlink_nl_post_doit,
1186: .policy = devlink_trap_get_do_nl_policy,
1187: .maxattr = DEVLINK_ATTR_INDEX,
1188: .flags = GENL_CMD_CAP_DO,
1189: },
1190: {
1191: .cmd = DEVLINK_CMD_TRAP_GET,
1192: .dumpit = devlink_nl_trap_get_dumpit,
1193: .policy = devlink_trap_get_dump_nl_policy,
1194: .maxattr = DEVLINK_ATTR_INDEX,
1195: .flags = GENL_CMD_CAP_DUMP,
1196: },
1197: {
1198: .cmd = DEVLINK_CMD_TRAP_SET,
1199: .validate = GENL_DONT_VALIDATE_STRICT,
1200: .pre_doit = devlink_nl_pre_doit,
1201: .doit = devlink_nl_trap_set_doit,
1202: .post_doit = devlink_nl_post_doit,
1203: .policy = devlink_trap_set_nl_policy,
1204: .maxattr = DEVLINK_ATTR_INDEX,
1205: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1206: },
1207: {
1208: .cmd = DEVLINK_CMD_TRAP_GROUP_GET,
1209: .validate = GENL_DONT_VALIDATE_STRICT,
1210: .pre_doit = devlink_nl_pre_doit,
1211: .doit = devlink_nl_trap_group_get_doit,
1212: .post_doit = devlink_nl_post_doit,
1213: .policy = devlink_trap_group_get_do_nl_policy,
1214: .maxattr = DEVLINK_ATTR_INDEX,
1215: .flags = GENL_CMD_CAP_DO,
1216: },
1217: {
1218: .cmd = DEVLINK_CMD_TRAP_GROUP_GET,
1219: .dumpit = devlink_nl_trap_group_get_dumpit,
1220: .policy = devlink_trap_group_get_dump_nl_policy,
1221: .maxattr = DEVLINK_ATTR_INDEX,
1222: .flags = GENL_CMD_CAP_DUMP,
1223: },
1224: {
1225: .cmd = DEVLINK_CMD_TRAP_GROUP_SET,
1226: .validate = GENL_DONT_VALIDATE_STRICT,
1227: .pre_doit = devlink_nl_pre_doit,
1228: .doit = devlink_nl_trap_group_set_doit,
1229: .post_doit = devlink_nl_post_doit,
1230: .policy = devlink_trap_group_set_nl_policy,
1231: .maxattr = DEVLINK_ATTR_INDEX,
1232: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1233: },
1234: {
1235: .cmd = DEVLINK_CMD_TRAP_POLICER_GET,
1236: .validate = GENL_DONT_VALIDATE_STRICT,
1237: .pre_doit = devlink_nl_pre_doit,
1238: .doit = devlink_nl_trap_policer_get_doit,
1239: .post_doit = devlink_nl_post_doit,
1240: .policy = devlink_trap_policer_get_do_nl_policy,
1241: .maxattr = DEVLINK_ATTR_INDEX,
1242: .flags = GENL_CMD_CAP_DO,
1243: },
1244: {
1245: .cmd = DEVLINK_CMD_TRAP_POLICER_GET,
1246: .dumpit = devlink_nl_trap_policer_get_dumpit,
1247: .policy = devlink_trap_policer_get_dump_nl_policy,
1248: .maxattr = DEVLINK_ATTR_INDEX,
1249: .flags = GENL_CMD_CAP_DUMP,
1250: },
1251: {
1252: .cmd = DEVLINK_CMD_TRAP_POLICER_SET,
1253: .validate = GENL_DONT_VALIDATE_STRICT,
1254: .pre_doit = devlink_nl_pre_doit,
1255: .doit = devlink_nl_trap_policer_set_doit,
1256: .post_doit = devlink_nl_post_doit,
1257: .policy = devlink_trap_policer_set_nl_policy,
1258: .maxattr = DEVLINK_ATTR_INDEX,
1259: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1260: },
1261: {
1262: .cmd = DEVLINK_CMD_HEALTH_REPORTER_TEST,
1263: .validate = GENL_DONT_VALIDATE_STRICT,
1264: .pre_doit = devlink_nl_pre_doit_port_optional,
1265: .doit = devlink_nl_health_reporter_test_doit,
1266: .post_doit = devlink_nl_post_doit,
1267: .policy = devlink_health_reporter_test_nl_policy,
1268: .maxattr = DEVLINK_ATTR_INDEX,
1269: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1270: },
1271: {
1272: .cmd = DEVLINK_CMD_RATE_GET,
1273: .validate = GENL_DONT_VALIDATE_STRICT,
1274: .pre_doit = devlink_nl_pre_doit,
1275: .doit = devlink_nl_rate_get_doit,
1276: .post_doit = devlink_nl_post_doit,
1277: .policy = devlink_rate_get_do_nl_policy,
1278: .maxattr = DEVLINK_ATTR_INDEX,
1279: .flags = GENL_CMD_CAP_DO,
1280: },
1281: {
1282: .cmd = DEVLINK_CMD_RATE_GET,
1283: .dumpit = devlink_nl_rate_get_dumpit,
1284: .policy = devlink_rate_get_dump_nl_policy,
1285: .maxattr = DEVLINK_ATTR_INDEX,
1286: .flags = GENL_CMD_CAP_DUMP,
1287: },
1288: {
1289: .cmd = DEVLINK_CMD_RATE_SET,
1290: .validate = GENL_DONT_VALIDATE_STRICT,
1291: .pre_doit = devlink_nl_pre_doit,
1292: .doit = devlink_nl_rate_set_doit,
1293: .post_doit = devlink_nl_post_doit,
1294: .policy = devlink_rate_set_nl_policy,
1295: .maxattr = DEVLINK_ATTR_INDEX,
1296: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1297: },
1298: {
1299: .cmd = DEVLINK_CMD_RATE_NEW,
1300: .validate = GENL_DONT_VALIDATE_STRICT,
1301: .pre_doit = devlink_nl_pre_doit,
1302: .doit = devlink_nl_rate_new_doit,
1303: .post_doit = devlink_nl_post_doit,
1304: .policy = devlink_rate_new_nl_policy,
1305: .maxattr = DEVLINK_ATTR_INDEX,
1306: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1307: },
1308: {
1309: .cmd = DEVLINK_CMD_RATE_DEL,
1310: .validate = GENL_DONT_VALIDATE_STRICT,
1311: .pre_doit = devlink_nl_pre_doit,
1312: .doit = devlink_nl_rate_del_doit,
1313: .post_doit = devlink_nl_post_doit,
1314: .policy = devlink_rate_del_nl_policy,
1315: .maxattr = DEVLINK_ATTR_INDEX,
1316: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1317: },
1318: {
1319: .cmd = DEVLINK_CMD_LINECARD_GET,
1320: .validate = GENL_DONT_VALIDATE_STRICT,
1321: .pre_doit = devlink_nl_pre_doit,
1322: .doit = devlink_nl_linecard_get_doit,
1323: .post_doit = devlink_nl_post_doit,
1324: .policy = devlink_linecard_get_do_nl_policy,
1325: .maxattr = DEVLINK_ATTR_INDEX,
1326: .flags = GENL_CMD_CAP_DO,
1327: },
1328: {
1329: .cmd = DEVLINK_CMD_LINECARD_GET,
1330: .dumpit = devlink_nl_linecard_get_dumpit,
1331: .policy = devlink_linecard_get_dump_nl_policy,
1332: .maxattr = DEVLINK_ATTR_INDEX,
1333: .flags = GENL_CMD_CAP_DUMP,
1334: },
1335: {
1336: .cmd = DEVLINK_CMD_LINECARD_SET,
1337: .validate = GENL_DONT_VALIDATE_STRICT,
1338: .pre_doit = devlink_nl_pre_doit,
1339: .doit = devlink_nl_linecard_set_doit,
1340: .post_doit = devlink_nl_post_doit,
1341: .policy = devlink_linecard_set_nl_policy,
1342: .maxattr = DEVLINK_ATTR_INDEX,
1343: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1344: },
1345: {
1346: .cmd = DEVLINK_CMD_SELFTESTS_GET,
1347: .validate = GENL_DONT_VALIDATE_STRICT,
1348: .pre_doit = devlink_nl_pre_doit,
1349: .doit = devlink_nl_selftests_get_doit,
1350: .post_doit = devlink_nl_post_doit,
1351: .policy = devlink_selftests_get_nl_policy,
1352: .maxattr = DEVLINK_ATTR_INDEX,
1353: .flags = GENL_CMD_CAP_DO,
1354: },
1355: {
1356: .cmd = DEVLINK_CMD_SELFTESTS_GET,
1357: .validate = GENL_DONT_VALIDATE_DUMP,
1358: .dumpit = devlink_nl_selftests_get_dumpit,
1359: .flags = GENL_CMD_CAP_DUMP,
1360: },
1361: {
1362: .cmd = DEVLINK_CMD_SELFTESTS_RUN,
1363: .validate = GENL_DONT_VALIDATE_STRICT,
1364: .pre_doit = devlink_nl_pre_doit,
1365: .doit = devlink_nl_selftests_run_doit,
1366: .post_doit = devlink_nl_post_doit,
1367: .policy = devlink_selftests_run_nl_policy,
1368: .maxattr = DEVLINK_ATTR_INDEX,
1369: .flags = GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
1370: },
1371: {
1372: .cmd = DEVLINK_CMD_NOTIFY_FILTER_SET,
1373: .doit = devlink_nl_notify_filter_set_doit,
1374: .policy = devlink_notify_filter_set_nl_policy,
1375: .maxattr = DEVLINK_ATTR_INDEX,
1376: .flags = GENL_CMD_CAP_DO,
1377: },
1378: };
SourceFile:net/devlink/netlink_gen.c]
|
| 102/2 |
2026/05/26 14:42 |
llm |
expert |
0m
Model:
Tokens:
input: 86682
output: 38
thoughts: 2908
Thoughts: **Bug Analysis: Race Condition in FIB Notification Handling**
Okay, so here's the breakdown of what I've figured out. We've got a race condition stemming from how the kernel handles FIB (Forwarding Information Base) notifications, specifically during `devlink` reloads and route deletions. My thinking goes like this: `devlink_nl_ops` for `DEVLINK_CMD_RELOAD` doesn't grant `GENL_CMD_CAP_DO` when using `RTNL` lock, and `genl_split_ops` has no `RTNL` flag. This means `devlink_reload` doesn't hold the `RTNL` lock. This is the crux of the problem.
`fib_net_dump` gets called from `devlink_reload`, and *it doesn't hold* the `RTNL` lock. Meanwhile, a `fib_table_delete` operation *does* hold `RTNL`. This concurrency is where the trouble begins.
`fib_net_dump` iterates the FIB trie under `rcu_read_lock()`. Concurrently, a route deletion happens, `fib_table_delete` holds `RTNL` and removes an alias from the trie with `hlist_del_rcu`, then calls `fib_release_info`. This decrements refcounts. Crucially, `fib_release_info` schedules the freeing of `fib_info` via RCU.
Here's the problem: `fib_net_dump` can still be traversing the trie under `rcu_read_lock()` after the alias has been removed. This means it might *find* the entry that was just deleted and call `fib_leaf_notify`, leading to a chain of driver callbacks: `call_fib_entry_notifier` -> `nsim_fib4_prepare_event`, and so on. These drivers will call `fib_info_hold(fi)`, which increments `fib_clntref`. But if `fib_clntref` was zero, then the system will warn.
At this point, the RCU grace period expires, and `fib_info` gets freed. But the driver, now holding a pointer to the freed `fib_info` (which was incremented by `fib_info_hold`), has queued a work item that's going to access it! This is a classic use-after-free scenario.
The impact is that the drivers can read arbitrary data via a memory corruption.
Is this exploitable? Yes, it's a use-after-free of `fib_info` and leads to a double free in the `mlxsw` driver. A work item accesses the freed `fib_info`, potentially allowing control over read data if an attacker can overwrite the freed memory.
However, only the root can trigger the conditions leading to this bug. Specifically, an attacker would need root to:
* Reload a `devlink` instance via `devlink dev reload`
* Move the `devlink` instance to a new network namespace
* Destroy a network namespace containing a `devlink` instance
So, while the bug is exploitable in theory, privilege escalation is not possible. It's a kernel denial of service, but root is required to trigger it. The double-free makes it highly exploitable for a crashing of the kernel, if root does indeed run a reload concurrently with route deletion.
|
| 103/2 |
2026/05/26 14:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/core/fib_notifier.c Name:register_fib_notifier]
Results: map[SourceCode: 98: int register_fib_notifier(struct net *net, struct notifier_block *nb,
99: void (*cb)(struct notifier_block *nb),
100: struct netlink_ext_ack *extack)
101: {
102: int retries = 0;
103: int err;
104:
105: do {
106: unsigned int fib_seq = fib_seq_sum(net);
107:
108: err = fib_net_dump(net, nb, extack);
109: if (err)
110: return err;
111:
112: if (fib_dump_is_consistent(net, nb, cb, fib_seq))
113: return 0;
114: } while (++retries < FIB_DUMP_MAX_RETRIES);
115:
116: return -EBUSY;
117: }
SourceFile:net/core/fib_notifier.c]
|
| 104/2 |
2026/05/26 14:42 |
llm |
expert |
0m
Model:
Tokens:
input: 90666
output: 40
thoughts: 42
Thoughts: Okay, let's see. My mind's jumping around a bit here.
**Deconstructing the Fib Notifier Flow**
So, `register_fib_notifier`... right, that one. It seems like it's calling `fib_net_dump`, which then leads into `fib_dump_is_consistent`. I need to understand that consistency check. Let me drill down into `fib_dump_is_consistent`. What's happening there to determine if the dump is, well, consistent? That's the key.
|
| 105/2 |
2026/05/26 14:42 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/core/fib_notifier.c Name:fib_dump_is_consistent]
Results: map[SourceCode: 82: static bool fib_dump_is_consistent(struct net *net, struct notifier_block *nb,
83: void (*cb)(struct notifier_block *nb),
84: unsigned int fib_seq)
85: {
86: struct fib_notifier_net *fn_net = net_generic(net, fib_notifier_net_id);
87:
88: atomic_notifier_chain_register(&fn_net->fib_chain, nb);
89: if (fib_seq == fib_seq_sum(net))
90: return true;
91: atomic_notifier_chain_unregister(&fn_net->fib_chain, nb);
92: if (cb)
93: cb(nb);
94: return false;
95: }
SourceFile:net/core/fib_notifier.c]
|
| 106/2 |
2026/05/26 14:42 |
llm |
expert |
0m
Model:
Tokens:
input: 91111
output: 51
thoughts: 224
Thoughts: **Analyzing the Fib Notifier and Potential Race Condition**
Okay, so `fib_dump_is_consistent` registers that notifier block within `fn_net->fib_chain`. That makes sense; it seems `register_fib_notifier` is the canonical way to get a fib notifier in place. And from what I've seen, it's exclusively called by those drivers we identified.
Now, I'm thinking about the bigger picture. Is there *any* other mechanism, any other path, that could potentially kick off `fib_net_dump`? Nope, it looks like that's a static function, locked away in `net/core/fib_notifier.c`, and the only call point is that same `register_fib_notifier`.
Alright, that starts to narrow things down significantly. This points to a bug specifically linked to a race condition. It seems to be the race between a driver initializing (or a devlink reload occurring) and a route being deleted.
And since driver initialization and devlink reload are privileged operations, an unprivileged user can't actually trigger this issue. Thatβs good from a security perspective.
Let me do a sanity check, though. Let's make sure that `fib_info_hold` isn't called from anywhere else unexpectedly. I already checked the references for `fib_info_hold`, and it's limited to `netdevsim`, `mlxsw`, and `prestera`, but I want to be 100% sure. Wait, what about `fib_info_get`? Is there a `fib_info_get`? I need to check that.
|
| 107/2 |
2026/05/26 14:42 |
tool |
codesearch-find-references |
0m
Error:
Args: map[ContextFile:include/net/ip_fib.h IncludeSnippetLines:0 Name:fib_info_get SourceTreePrefix:]
Results: map[References:<nil> TruncatedOutput:false]
|
| 108/2 |
2026/05/26 14:42 |
llm |
expert |
0m
Model:
Tokens:
input: 91755
output: 55
thoughts: 38
Thoughts: Okay, here's my take on those thoughts, framed from my own perspective as someone in the field:
**Digging into the `fib_clntref` Counter**
Alright, I'm thinking about this `fib_info_put` function. We know it's *decrementing* the `fib_clntref` counter, which is a key indicator that we are keeping track of. I need to understand this better. Now, the core question that's bubbling up in my mind is: *where else* is this `fib_clntref` counter being touched? Specifically, are there other sections of the code β other functions, maybe some initialization routines β that are responsible for *incrementing* this value? I need to track the increments to find the sources. The increments will give me insight into the usage of these objects. I can't just look at the decrement, I have to find the matching increment. If I can't find the complementary increment logic, then the whole system would be flawed. Finding the other locations where `fib_clntref` is manipulated is critical to ensure proper reference counting and memory management. It's the only way to avoid resource leaks or unexpected frees.
|
| 109/2 |
2026/05/26 14:43 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:include/net/ip_fib.h IncludeSnippetLines:0 Name:fib_info::fib_clntref SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:fib_create_info SourceFile:net/ipv4/fib_semantics.c SourceLine:1540] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:fib_info_hold SourceFile:include/net/ip_fib.h SourceLine:629] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:fib_info_put SourceFile:include/net/ip_fib.h SourceLine:634] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:fib_table_lookup SourceFile:net/ipv4/fib_trie.c SourceLine:1598]] TruncatedOutput:false]
|
| 110/2 |
2026/05/26 14:43 |
llm |
expert |
0m
Model:
Tokens:
input: 92348
output: 40
thoughts: 11
Thoughts: Okay, here's my interpretation of that thought process, assuming I'm deep in my domain of expertise:
**Checking `fib_table_lookup`**
Alright, let's take a look at `fib_table_lookup`. That name immediately suggests a memoization strategy for calculating Fibonacci numbers β storing pre-computed values to avoid redundant calculations. I want to see how efficiently this function is implemented. What's the storage structure? Is it a hash table, an array, something else? And more importantly, what's the lookup time complexity? A simple linear search would negate much of the benefit. Is it using a dynamic programming approach, correctly handling base cases (0 and 1, usually), and accurately populating the table? I need to check for any potential off-by-one errors or boundary condition issues. Finally, I'll need to confirm that it's actually *faster* than a recursive or iterative calculation for the target input sizes. Sometimes memoization adds overhead that isn't offset by the savings. Let's dig in and make sure itβs actually a performance gain.
|
| 111/2 |
2026/05/26 14:43 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/ipv4/fib_trie.c Name:fib_table_lookup]
Results: map[SourceCode:1419: /* should be called with rcu_read_lock */
1420: int fib_table_lookup(struct fib_table *tb, const struct flowi4 *flp,
1421: struct fib_result *res, int fib_flags)
1422: {
1423: struct trie *t = (struct trie *) tb->tb_data;
1424: #ifdef CONFIG_IP_FIB_TRIE_STATS
1425: struct trie_use_stats __percpu *stats = t->stats;
1426: #endif
1427: const t_key key = ntohl(flp->daddr);
1428: struct key_vector *n, *pn;
1429: struct fib_alias *fa;
1430: unsigned long index;
1431: t_key cindex;
1432:
1433: pn = t->kv;
1434: cindex = 0;
1435:
1436: n = get_child_rcu(pn, cindex);
1437: if (!n) {
1438: trace_fib_table_lookup(tb->tb_id, flp, NULL, -EAGAIN);
1439: return -EAGAIN;
1440: }
1441:
1442: #ifdef CONFIG_IP_FIB_TRIE_STATS
1443: this_cpu_inc(stats->gets);
1444: #endif
1445:
1446: /* Step 1: Travel to the longest prefix match in the trie */
1447: for (;;) {
1448: index = get_cindex(key, n);
1449:
1450: /* This bit of code is a bit tricky but it combines multiple
1451: * checks into a single check. The prefix consists of the
1452: * prefix plus zeros for the "bits" in the prefix. The index
1453: * is the difference between the key and this value. From
1454: * this we can actually derive several pieces of data.
1455: * if (index >= (1ul << bits))
1456: * we have a mismatch in skip bits and failed
1457: * else
1458: * we know the value is cindex
1459: *
1460: * This check is safe even if bits == KEYLENGTH due to the
1461: * fact that we can only allocate a node with 32 bits if a
1462: * long is greater than 32 bits.
1463: */
1464: if (index >= (1ul << n->bits))
1465: break;
1466:
1467: /* we have found a leaf. Prefixes have already been compared */
1468: if (IS_LEAF(n))
1469: goto found;
1470:
1471: /* only record pn and cindex if we are going to be chopping
1472: * bits later. Otherwise we are just wasting cycles.
1473: */
1474: if (n->slen > n->pos) {
1475: pn = n;
1476: cindex = index;
1477: }
1478:
1479: n = get_child_rcu(n, index);
1480: if (unlikely(!n))
1481: goto backtrace;
1482: }
1483:
1484: /* Step 2: Sort out leaves and begin backtracing for longest prefix */
1485: for (;;) {
1486: /* record the pointer where our next node pointer is stored */
1487: struct key_vector __rcu **cptr = n->tnode;
1488:
1489: /* This test verifies that none of the bits that differ
1490: * between the key and the prefix exist in the region of
1491: * the lsb and higher in the prefix.
1492: */
1493: if (unlikely(prefix_mismatch(key, n)) || (n->slen == n->pos))
1494: goto backtrace;
1495:
1496: /* exit out and process leaf */
1497: if (unlikely(IS_LEAF(n)))
1498: break;
1499:
1500: /* Don't bother recording parent info. Since we are in
1501: * prefix match mode we will have to come back to wherever
1502: * we started this traversal anyway
1503: */
1504:
1505: while ((n = rcu_dereference(*cptr)) == NULL) {
1506: backtrace:
1507: #ifdef CONFIG_IP_FIB_TRIE_STATS
1508: if (!n)
1509: this_cpu_inc(stats->null_node_hit);
1510: #endif
1511: /* If we are at cindex 0 there are no more bits for
1512: * us to strip at this level so we must ascend back
1513: * up one level to see if there are any more bits to
1514: * be stripped there.
1515: */
1516: while (!cindex) {
1517: t_key pkey = pn->key;
1518:
1519: /* If we don't have a parent then there is
1520: * nothing for us to do as we do not have any
1521: * further nodes to parse.
1522: */
1523: if (IS_TRIE(pn)) {
1524: trace_fib_table_lookup(tb->tb_id, flp,
1525: NULL, -EAGAIN);
1526: return -EAGAIN;
1527: }
1528: #ifdef CONFIG_IP_FIB_TRIE_STATS
1529: this_cpu_inc(stats->backtrack);
1530: #endif
1531: /* Get Child's index */
1532: pn = node_parent_rcu(pn);
1533: cindex = get_index(pkey, pn);
1534: }
1535:
1536: /* strip the least significant bit from the cindex */
1537: cindex &= cindex - 1;
1538:
1539: /* grab pointer for next child node */
1540: cptr = &pn->tnode[cindex];
1541: }
1542: }
1543:
1544: found:
1545: /* this line carries forward the xor from earlier in the function */
1546: index = key ^ n->key;
1547:
1548: /* Step 3: Process the leaf, if that fails fall back to backtracing */
1549: hlist_for_each_entry_rcu(fa, &n->leaf, fa_list) {
1550: struct fib_info *fi = fa->fa_info;
1551: struct fib_nh_common *nhc;
1552: int nhsel, err;
1553:
1554: if ((BITS_PER_LONG > KEYLENGTH) || (fa->fa_slen < KEYLENGTH)) {
1555: if (index >= (1ul << fa->fa_slen))
1556: continue;
1557: }
1558: if (fa->fa_dscp && !fib_dscp_masked_match(fa->fa_dscp, flp))
1559: continue;
1560: /* Paired with WRITE_ONCE() in fib_release_info() */
1561: if (READ_ONCE(fi->fib_dead))
1562: continue;
1563: if (fa->fa_info->fib_scope < flp->flowi4_scope)
1564: continue;
1565: fib_alias_accessed(fa);
1566: err = fib_props[fa->fa_type].error;
1567: if (unlikely(err < 0)) {
1568: out_reject:
1569: #ifdef CONFIG_IP_FIB_TRIE_STATS
1570: this_cpu_inc(stats->semantic_match_passed);
1571: #endif
1572: trace_fib_table_lookup(tb->tb_id, flp, NULL, err);
1573: return err;
1574: }
1575: if (fi->fib_flags & RTNH_F_DEAD)
1576: continue;
1577:
1578: if (unlikely(fi->nh)) {
1579: if (nexthop_is_blackhole(fi->nh)) {
1580: err = fib_props[RTN_BLACKHOLE].error;
1581: goto out_reject;
1582: }
1583:
1584: nhc = nexthop_get_nhc_lookup(fi->nh, fib_flags, flp,
1585: &nhsel);
1586: if (nhc)
1587: goto set_result;
1588: goto miss;
1589: }
1590:
1591: for (nhsel = 0; nhsel < fib_info_num_path(fi); nhsel++) {
1592: nhc = fib_info_nhc(fi, nhsel);
1593:
1594: if (!fib_lookup_good_nhc(nhc, fib_flags, flp))
1595: continue;
1596: set_result:
1597: if (!(fib_flags & FIB_LOOKUP_NOREF))
1598: refcount_inc(&fi->fib_clntref);
1599:
1600: res->prefix = htonl(n->key);
1601: res->prefixlen = KEYLENGTH - fa->fa_slen;
1602: res->nh_sel = nhsel;
1603: res->nhc = nhc;
1604: res->type = fa->fa_type;
1605: res->scope = fi->fib_scope;
1606: res->dscp = fa->fa_dscp;
1607: res->fi = fi;
1608: res->table = tb;
1609: res->fa_head = &n->leaf;
1610: #ifdef CONFIG_IP_FIB_TRIE_STATS
1611: this_cpu_inc(stats->semantic_match_passed);
1612: #endif
1613: trace_fib_table_lookup(tb->tb_id, flp, nhc, err);
1614:
1615: return err;
1616: }
1617: }
1618: miss:
1619: #ifdef CONFIG_IP_FIB_TRIE_STATS
1620: this_cpu_inc(stats->semantic_match_miss);
1621: #endif
1622: goto backtrace;
1623: }
SourceFile:net/ipv4/fib_trie.c]
|
| 112/2 |
2026/05/26 14:43 |
llm |
expert |
0m
Model:
Tokens:
input: 95770
output: 39
thoughts: 129
Thoughts: **Analyzing the Fibre Channel Reference Counting**
Okay, so `fib_table_lookup` is where it's happening. I see it also calls `refcount_inc(&fi->fib_clntref)`. Good, that's what I expected. But hang on, `fib_table_lookup` is called within an `rcu_read_lock()` context. So, it's operating in that safe environment. And I recall it checks `READ_ONCE(fi->fib_dead)`. Right, if `fib_dead` is true, it bails out, preventing any issues.
Now, let's connect the dots. When `fib_release_info` is called, it sets `fi->fib_dead = 1` *before* calling `fib_info_put(fi)`. Okay, so `fib_dead` effectively acts as a flag to prevent further operations on a released entry. That sequence is critical, especially under RCU. I'm going to revisit `fib_release_info` one more time, just to make absolutely certain I haven't missed anything subtle. It's that critical point where the refcount and the `fib_dead` flag interact, and I want to be entirely clear on how it all works.
|
| 113/2 |
2026/05/26 14:43 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/ipv4/fib_semantics.c Name:fib_release_info]
Results: map[SourceCode: 250: void fib_release_info(struct fib_info *fi)
251: {
252: ASSERT_RTNL();
253: if (fi && refcount_dec_and_test(&fi->fib_treeref)) {
254: hlist_del(&fi->fib_hash);
255: fi->fib_net->ipv4.fib_info_cnt--;
256:
257: if (fi->fib_prefsrc)
258: hlist_del(&fi->fib_lhash);
259: if (fi->nh) {
260: list_del(&fi->nh_list);
261: } else {
262: change_nexthops(fi) {
263: if (!nexthop_nh->fib_nh_dev)
264: continue;
265: hlist_del_rcu(&nexthop_nh->nh_hash);
266: } endfor_nexthops(fi)
267: }
268: /* Paired with READ_ONCE() from fib_table_lookup() */
269: WRITE_ONCE(fi->fib_dead, 1);
270: fib_info_put(fi);
271: }
272: }
SourceFile:net/ipv4/fib_semantics.c]
|
| 114/2 |
2026/05/26 14:43 |
llm |
expert |
0m
Model:
Tokens:
input: 96607
output: 40
thoughts: 106
Thoughts: **Analyzing the Fiber Routing Table Cleanup**
Okay, so I'm tracing through this fiber routing table code. I see that `WRITE_ONCE(fi->fib_dead, 1)` is definitely called *before* `fib_info_put(fi)`. That's good. Which means `fib_table_lookup` is doing the right thing, checking `fi->fib_dead` and preventing any increment to `fib_clntref` if the fiber info is already marked as dead. This avoids a use-after-free situation, which is exactly what I was worried about.
But⦠hold on a second. This `fib_leaf_notify` function, which is called by the `fib_net_dump` function... *it doesn't appear* to check `fi->fib_dead`! That's a huge potential issue. `fib_net_dump` is probably being used during debugging or some kind of network state inspection, right? If it's hitting a fiber info entry that's in the process of being freed, we could still have a race condition or crash. I need to dig into `fib_leaf_notify` specifically and confirm that my initial assessment is correct. Let's see what's actually going on in that function. It's time to check `fib_leaf_notify`.
|
| 115/2 |
2026/05/26 14:43 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/ipv4/fib_trie.c Name:fib_leaf_notify]
Results: map[SourceCode:2152: static int fib_leaf_notify(struct key_vector *l, struct fib_table *tb,
2153: struct notifier_block *nb,
2154: struct netlink_ext_ack *extack)
2155: {
2156: struct fib_alias *fa;
2157: int last_slen = -1;
2158: int err;
2159:
2160: hlist_for_each_entry_rcu(fa, &l->leaf, fa_list) {
2161: struct fib_info *fi = fa->fa_info;
2162:
2163: if (!fi)
2164: continue;
2165:
2166: /* local and main table can share the same trie,
2167: * so don't notify twice for the same entry.
2168: */
2169: if (tb->tb_id != fa->tb_id)
2170: continue;
2171:
2172: if (fa->fa_slen == last_slen)
2173: continue;
2174:
2175: last_slen = fa->fa_slen;
2176: err = call_fib_entry_notifier(nb, FIB_EVENT_ENTRY_REPLACE,
2177: l->key, KEYLENGTH - fa->fa_slen,
2178: fa, extack);
2179: if (err)
2180: return err;
2181: }
2182: return 0;
2183: }
SourceFile:net/ipv4/fib_trie.c]
|
| 116/2 |
2026/05/26 14:43 |
llm |
expert |
0m
Model:
Tokens:
input: 97525
output: 52
thoughts: 120
Thoughts: **Debugging a `fib_leaf_notify` Issue**
Okay, so I'm digging into this `fib_leaf_notify` function. It looks like it *doesn't* check the `fi->fib_dead` flag. That's a huge oversight! This means it's happily passing a dead `fib_info` structure to the notifier, which then calls `fib_info_hold`. And *that* call is what's triggering the bug. This is a classic example of a use-after-free scenario waiting to happen.
Let's pause and double-check something. I need to be absolutely sure this `fib_info_hold` function is only called from places we expect. We've checked the drivers, and that seemed to be the only place. But, is there any *other* code path that increments that crucial `fib_clntref` counter? If so, we need to go back and check. Okay, `fib_table_lookup` is a possibility, but it looks like it does check the `fib_dead` flag before making the call. What about `fib_dump_info`? We need to verify that code path too. Gotta be thorough.
|
| 117/2 |
2026/05/26 14:43 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/ipv4/fib_semantics.c IncludeSnippetLines:0 Name:fib_dump_info SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:fib_alias_hw_flags_set SourceFile:net/ipv4/fib_trie.c SourceLine:1080] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:fn_trie_dump_leaf SourceFile:net/ipv4/fib_trie.c SourceLine:2297] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:inet_rtm_getroute SourceFile:net/ipv4/route.c SourceLine:3438] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rtmsg_fib SourceFile:net/ipv4/fib_semantics.c SourceLine:564]] TruncatedOutput:false]
|
| 118/2 |
2026/05/26 14:43 |
llm |
expert |
0m
Model:
Tokens:
input: 98193
output: 43
thoughts: 39
Thoughts: Here's my take on the thought process, as requested:
**Diving into the Netlink Route Dump Function**
Okay, so we're looking at a netlink route dump, specifically `RTM_GETROUTE`. My mental map immediately flags `fn_trie_dump_leaf` as the likely function responsible for handling the leaf nodes within that process. Given my deep understanding of the kernel's routing table structures and the netlink interface, this seems like the logical starting point. Let's delve into `fn_trie_dump_leaf` and see what's happening at the ground level. I need to refresh my memory of the implementation to confirm my initial hypothesis and identify any optimization potential or potential bottlenecks within this routine. My goal is to quickly pinpoint any areas requiring attention or deeper analysis within the dump process.
|
| 119/2 |
2026/05/26 14:43 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/ipv4/fib_trie.c Name:fn_trie_dump_leaf]
Results: map[SourceCode:2242: static int fn_trie_dump_leaf(struct key_vector *l, struct fib_table *tb,
2243: struct sk_buff *skb, struct netlink_callback *cb,
2244: struct fib_dump_filter *filter)
2245: {
2246: unsigned int flags = NLM_F_MULTI;
2247: __be32 xkey = htonl(l->key);
2248: int i, s_i, i_fa, s_fa, err;
2249: struct fib_alias *fa;
2250:
2251: if (filter->filter_set ||
2252: !filter->dump_exceptions || !filter->dump_routes)
2253: flags |= NLM_F_DUMP_FILTERED;
2254:
2255: s_i = cb->args[4];
2256: s_fa = cb->args[5];
2257: i = 0;
2258:
2259: /* rcu_read_lock is hold by caller */
2260: hlist_for_each_entry_rcu(fa, &l->leaf, fa_list) {
2261: struct fib_info *fi = fa->fa_info;
2262:
2263: if (i < s_i)
2264: goto next;
2265:
2266: i_fa = 0;
2267:
2268: if (tb->tb_id != fa->tb_id)
2269: goto next;
2270:
2271: if (filter->filter_set) {
2272: if (filter->rt_type && fa->fa_type != filter->rt_type)
2273: goto next;
2274:
2275: if ((filter->protocol &&
2276: fi->fib_protocol != filter->protocol))
2277: goto next;
2278:
2279: if (filter->dev &&
2280: !fib_info_nh_uses_dev(fi, filter->dev))
2281: goto next;
2282: }
2283:
2284: if (filter->dump_routes) {
2285: if (!s_fa) {
2286: struct fib_rt_info fri;
2287:
2288: fri.fi = fi;
2289: fri.tb_id = tb->tb_id;
2290: fri.dst = xkey;
2291: fri.dst_len = KEYLENGTH - fa->fa_slen;
2292: fri.dscp = fa->fa_dscp;
2293: fri.type = fa->fa_type;
2294: fri.offload = READ_ONCE(fa->offload);
2295: fri.trap = READ_ONCE(fa->trap);
2296: fri.offload_failed = READ_ONCE(fa->offload_failed);
2297: err = fib_dump_info(skb,
2298: NETLINK_CB(cb->skb).portid,
2299: cb->nlh->nlmsg_seq,
2300: RTM_NEWROUTE, &fri, flags);
2301: if (err < 0)
2302: goto stop;
2303: }
2304:
2305: i_fa++;
2306: }
2307:
2308: if (filter->dump_exceptions) {
2309: err = fib_dump_info_fnhe(skb, cb, tb->tb_id, fi,
2310: &i_fa, s_fa, flags);
2311: if (err < 0)
2312: goto stop;
2313: }
2314:
2315: next:
2316: i++;
2317: }
2318:
2319: cb->args[4] = i;
2320: return skb->len;
2321:
2322: stop:
2323: cb->args[4] = i;
2324: cb->args[5] = i_fa;
2325: return err;
2326: }
SourceFile:net/ipv4/fib_trie.c]
|
| 120/2 |
2026/05/26 14:43 |
llm |
expert |
0m
Model:
Tokens:
input: 99703
output: 19
thoughts: 683
Thoughts: **My Analysis of the `fib_info` Double-Free Vulnerability**
Okay, let's break this down. My initial thought was that `fn_trie_dump_leaf` seemed safe. It doesn't check for `fi->fib_dead`, but that's alright because the fields of `fi` are RCU-protected. `fib_dump_info` reads the fields and puts them into the netlink message, and since RCU makes sure the memory won't be freed before the grace period ends, that's fine. Importantly, it *doesn't* touch `fib_clntref` or `fib_treeref`, so no incrementing or decrementing of those refcounts. So far, so good, `fn_trie_dump_leaf` itself is not the problem.
The real issue, as I see it, centers around `fib_leaf_notify`. It's passing a potentially dead `fib_info` structure to a driver, and the driver is then doing a `fib_info_hold`, which is going to bump `fib_clntref` from 0 to 1. And then, at some point, the driver calls `fib_info_put`, which decrements it from 1 back to 0, which triggers the double `free_fib_info`. Essentially, the driver is incrementing then decrementing a refcount that shouldn't be touched.
So, this bug is *strictly* limited to the `fib_net_dump` path and the drivers that are registering these fib notifiers. Since you need root privileges to register fib notifiers, it appears that this bug *isn't* accessible from unprivileged processes or user namespaces.
Now, let's summarize my analysis in more structured terms:
* **Exploitable:** Yes. It's a use-after-free, double-free bug. `fib_info_hold` increments a refcount that shouldn't be, and then `fib_info_put` frees it. This is definitely exploitable.
* **Denial of Service:** Yes, absolutely. A double-free will corrupt kernel memory and likely cause a crash.
* **Unprivileged Process Access:** No. Registering a new fib notifier is restricted to driver initialization or devlink reload, which requires `CAP_NET_ADMIN` in the initial network namespace (root).
* **User Namespace Access:** No. Devlink instances and physical drivers are not accessible or movable by unprivileged users into user namespaces.
* **VM Guest/Host/Remote/Peripheral/Malicious Filesystem Trigger:** No.
Okay, let's pause. I need to make sure I haven't missed a way to trigger `fib_net_dump` from elsewhere. What about user namespaces? If a user namespace creates a new network namespace, what happens when `ip_fib_net_init` is called? Does that call `register_fib_notifier`? No, it looks like `ip_fib_net_init` calls `register_pernet_subsys` which is called during module init. It only initializes the tables, so that's a dead end. What about `fib_notifier_ops`? Drivers call `register_fib_notifier`. Can unprivileged users load drivers? No, `modprobe` needs root. Any drivers loaded automatically when a network namespace is created? No. How about virtual devices like `veth`, `dummy`, or `macvlan`? Yes, creating a `veth` device loads the `veth` module. Does `veth` register a fib notifier? I should check that.
|
| 121/2 |
2026/05/26 14:44 |
tool |
grepper |
0m
Args: map[Expression:register_fib_notifier]
Results: map[Output:Documentation/networking/switchdev.rst=349=To program the device, the driver has to register a FIB notifier handler
Documentation/networking/switchdev.rst:350:using register_fib_notifier. The following events are available:
Documentation/networking/switchdev.rst-351-
--
drivers/net/ethernet/marvell/prestera/prestera_router.c=1548=int prestera_router_init(struct prestera_switch *sw)
--
drivers/net/ethernet/marvell/prestera/prestera_router.c-1600- router->fib_nb.notifier_call = __prestera_router_fib_event;
drivers/net/ethernet/marvell/prestera/prestera_router.c:1601: err = register_fib_notifier(&init_net, &router->fib_nb,
drivers/net/ethernet/marvell/prestera/prestera_router.c-1602- /* TODO: flush fib entries */ NULL, NULL);
drivers/net/ethernet/marvell/prestera/prestera_router.c-1603- if (err)
drivers/net/ethernet/marvell/prestera/prestera_router.c:1604: goto err_register_fib_notifier;
drivers/net/ethernet/marvell/prestera/prestera_router.c-1605-
--
drivers/net/ethernet/marvell/prestera/prestera_router.c-1607-
drivers/net/ethernet/marvell/prestera/prestera_router.c:1608:err_register_fib_notifier:
drivers/net/ethernet/marvell/prestera/prestera_router.c-1609- unregister_netevent_notifier(&router->netevent_nb);
--
drivers/net/ethernet/marvell/prestera/prestera_router.c=1629=void prestera_router_fini(struct prestera_switch *sw)
drivers/net/ethernet/marvell/prestera/prestera_router.c-1630-{
drivers/net/ethernet/marvell/prestera/prestera_router.c:1631: unregister_fib_notifier(&init_net, &sw->router->fib_nb);
drivers/net/ethernet/marvell/prestera/prestera_router.c-1632- unregister_netevent_notifier(&sw->router->netevent_nb);
--
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c=1860=struct mlx5e_tc_tun_encap *mlx5e_tc_tun_init(struct mlx5e_priv *priv)
--
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c-1872- hash_init(encap->route_tbl);
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c:1873: err = register_fib_notifier(dev_net(priv->netdev), &encap->fib_nb,
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c-1874- NULL, NULL);
--
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c=1883=void mlx5e_tc_tun_cleanup(struct mlx5e_tc_tun_encap *encap)
--
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c-1887-
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c:1888: unregister_fib_notifier(dev_net(encap->priv->netdev), &encap->fib_nb);
drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_encap.c-1889- flush_workqueue(encap->priv->wq); /* flush fib event works */
--
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c=5393=int mlx5e_tc_esw_init(struct mlx5_rep_uplink_priv *uplink_priv)
--
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c-5449- err = PTR_ERR(uplink_priv->encap);
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c:5450: goto err_register_fib_notifier;
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c-5451- }
--
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c-5470- mlx5e_tc_tun_cleanup(uplink_priv->encap);
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c:5471:err_register_fib_notifier:
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c-5472- mapping_destroy(uplink_priv->tunnel_enc_opts_mapping);
--
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c=368=int mlx5_lag_mp_init(struct mlx5_lag *ldev)
--
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c-385- mp->fib_nb.notifier_call = mlx5_lag_fib_event;
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c:386: err = register_fib_notifier(&init_net, &mp->fib_nb,
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c-387- mlx5_lag_fib_event_flush, NULL);
--
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c=396=void mlx5_lag_mp_cleanup(struct mlx5_lag *ldev)
--
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c-402-
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c:403: unregister_fib_notifier(&init_net, &mp->fib_nb);
drivers/net/ethernet/mellanox/mlx5/core/lag/mp.c-404- destroy_workqueue(mp->wq);
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c=11575=int mlxsw_sp_router_init(struct mlxsw_sp *mlxsw_sp,
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11695- mlxsw_sp->router->fib_nb.notifier_call = mlxsw_sp_router_fib_event;
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c:11696: err = register_fib_notifier(mlxsw_sp_net(mlxsw_sp),
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11697- &mlxsw_sp->router->fib_nb,
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11699- if (err)
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c:11700: goto err_register_fib_notifier;
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11701-
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11703-
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c:11704:err_register_fib_notifier:
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11705- unregister_nexthop_notifier(mlxsw_sp_net(mlxsw_sp),
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c=11754=void mlxsw_sp_router_fini(struct mlxsw_sp *mlxsw_sp)
--
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11757-
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c:11758: unregister_fib_notifier(mlxsw_sp_net(mlxsw_sp), &router->fib_nb);
drivers/net/ethernet/mellanox/mlxsw/spectrum_router.c-11759- unregister_nexthop_notifier(mlxsw_sp_net(mlxsw_sp),
--
drivers/net/ethernet/rocker/rocker_main.c=2844=static int rocker_probe(struct pci_dev *pdev, const struct pci_device_id *id)
--
drivers/net/ethernet/rocker/rocker_main.c-2942- rocker->fib_nb.notifier_call = rocker_router_fib_event;
drivers/net/ethernet/rocker/rocker_main.c:2943: err = register_fib_notifier(&init_net, &rocker->fib_nb, NULL, NULL);
drivers/net/ethernet/rocker/rocker_main.c-2944- if (err)
drivers/net/ethernet/rocker/rocker_main.c:2945: goto err_register_fib_notifier;
drivers/net/ethernet/rocker/rocker_main.c-2946-
--
drivers/net/ethernet/rocker/rocker_main.c-2969-err_register_switchdev_notifier:
drivers/net/ethernet/rocker/rocker_main.c:2970: unregister_fib_notifier(&init_net, &rocker->fib_nb);
drivers/net/ethernet/rocker/rocker_main.c:2971:err_register_fib_notifier:
drivers/net/ethernet/rocker/rocker_main.c-2972- rocker_remove_ports(rocker);
--
drivers/net/ethernet/rocker/rocker_main.c=2997=static void rocker_remove(struct pci_dev *pdev)
--
drivers/net/ethernet/rocker/rocker_main.c-3005- unregister_switchdev_notifier(&rocker_switchdev_notifier);
drivers/net/ethernet/rocker/rocker_main.c:3006: unregister_fib_notifier(&init_net, &rocker->fib_nb);
drivers/net/ethernet/rocker/rocker_main.c-3007- rocker_remove_ports(rocker);
--
drivers/net/netdevsim/fib.c=1552=struct nsim_fib_data *nsim_fib_create(struct devlink *devlink,
--
drivers/net/netdevsim/fib.c-1595- data->fib_nb.notifier_call = nsim_fib_event_nb;
drivers/net/netdevsim/fib.c:1596: err = register_fib_notifier(devlink_net(devlink), &data->fib_nb,
drivers/net/netdevsim/fib.c-1597- nsim_fib_dump_inconsistent, extack);
--
drivers/net/netdevsim/fib.c=1644=void nsim_fib_destroy(struct devlink *devlink, struct nsim_fib_data *data)
--
drivers/net/netdevsim/fib.c-1655- NSIM_RESOURCE_IPV4_FIB);
drivers/net/netdevsim/fib.c:1656: unregister_fib_notifier(devlink_net(devlink), &data->fib_nb);
drivers/net/netdevsim/fib.c-1657- unregister_nexthop_notifier(devlink_net(devlink), &data->nexthop_nb);
--
include/net/fib_notifier.h=41=int call_fib_notifiers(struct net *net, enum fib_event_type event_type,
include/net/fib_notifier.h-42- struct fib_notifier_info *info);
include/net/fib_notifier.h:43:int register_fib_notifier(struct net *net, struct notifier_block *nb,
include/net/fib_notifier.h-44- void (*cb)(struct notifier_block *nb),
include/net/fib_notifier.h-45- struct netlink_ext_ack *extack);
include/net/fib_notifier.h:46:int unregister_fib_notifier(struct net *net, struct notifier_block *nb);
include/net/fib_notifier.h-47-struct fib_notifier_ops *
--
net/core/fib_notifier.c=82=static bool fib_dump_is_consistent(struct net *net, struct notifier_block *nb,
--
net/core/fib_notifier.c-97-#define FIB_DUMP_MAX_RETRIES 5
net/core/fib_notifier.c:98:int register_fib_notifier(struct net *net, struct notifier_block *nb,
net/core/fib_notifier.c-99- void (*cb)(struct notifier_block *nb),
--
net/core/fib_notifier.c-117-}
net/core/fib_notifier.c:118:EXPORT_SYMBOL(register_fib_notifier);
net/core/fib_notifier.c-119-
net/core/fib_notifier.c:120:int unregister_fib_notifier(struct net *net, struct notifier_block *nb)
net/core/fib_notifier.c-121-{
--
net/core/fib_notifier.c-125-}
net/core/fib_notifier.c:126:EXPORT_SYMBOL(unregister_fib_notifier);
net/core/fib_notifier.c-127-
]
|
| 122/2 |
2026/05/26 14:44 |
llm |
expert |
0m
Model:
Tokens:
input: 104807
output: 66
thoughts: 110
Thoughts: **My Assessment of the `veth` Vulnerability**
Okay, after examining the code, I've confirmed that the `veth` device driver *does not* register a FIB (Forwarding Information Base) notifier. This is a crucial detail. The only modules I've observed that do so are hardware switch drivers and `netdevsim`. Therefore, a regular, unprivileged user is completely unable to trigger the `register_fib_notifier` function directly through a `veth` interface. That pretty much rules out a simple local privilege escalation based on this specific vulnerability.
I'm completely confident in this assessment. I've analyzed the relevant kernel modules and the flow of events.
So, here's my understanding of the exploit characteristics: While this could certainly be used for a denial of service (DoS), as it potentially manipulates FIB and network tables, it isn't exploitable for unprivileged escalation. A normal user can't initiate the critical function. It's not triggerable via user namespaces, virtual machine guests or hosts, networks, remote access, peripherals or filesystem.
|
| 123/2 |
2026/05/26 14:44 |
tool |
set-results |
0m
Args: map[DenialOfService:true Exploitable:true FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
Results: map[DenialOfService:true Exploitable:true FilesystemTrigger:false NetworkTrigger:false PeripheralTrigger:false RemoteTrigger:false Unprivileged:false UserNamespace:false VMGuestTrigger:false VMHostTrigger:false]
|
| 124/2 |
2026/05/26 14:44 |
llm |
expert |
0m
Model:
Tokens:
input: 105276
output: 791
thoughts: 8
Thoughts: Okay, I'm ready. Here's my summary, presented as my internal thought process, assuming I'm an expert in my field:
**Finalizing the Response**
Alright, this is it. Time to finalize the reply. I've considered all the points, weighed the nuances, and hopefully, crafted a response that's both accurate and insightful. The goal is clarity, of course, and to ensure this response is as helpful as possible for the intended audience, even with all their expertise. I'll double-check for any logical gaps or potential ambiguities. This last read-through is crucial. Now, let's make sure itβs concise. Time to wrap this up.
|
| 125/1 |
2026/05/26 14:44 |
action |
format-explanation |
0m
Results: map[Explanation:**Bug Analysis:**
The bug is a race condition between `fib_net_dump()` and route deletion
(`fib_table_delete()`), leading to a use-after-free and double free of a
`fib_info` object.
1. **The Race Condition:**
- `fib_net_dump()` is called when a new FIB notifier is registered (via
`register_fib_notifier()`). It iterates over the FIB trie under
`rcu_read_lock()` and calls `fib_leaf_notify()` for each entry.
- Concurrently, a route can be deleted. `fib_table_delete()` holds the `RTNL`
lock, removes the `fib_alias` from the trie using `hlist_del_rcu()`, and
calls `fib_release_info()`.
- `fib_release_info()` drops the `fib_treeref` and `fib_clntref` reference
counts to 0, sets `fi->fib_dead = 1`, and schedules the `fib_info` object for
freeing via RCU.
- Because `fib_net_dump()` is running under `rcu_read_lock()` and does not
hold the `RTNL` lock, it can still find the removed `fib_alias` in the trie.
- `fib_leaf_notify()` does not check `fi->fib_dead` and passes the dead
`fib_info` to the registered driver's notifier callback.
2. **The Bug in Drivers:**
- Hardware switch drivers like `netdevsim`, `mlxsw`, and `prestera` implement
FIB notifier callbacks (e.g., `nsim_fib4_prepare_event()`).
- These callbacks call `fib_info_hold(fi)` to take a reference on the
`fib_info` object so they can process it asynchronously in a workqueue.
- `fib_info_hold()` does a `refcount_inc(&fi->fib_clntref)`. Since the
reference count is already 0, this triggers the `refcount_t: addition on 0`
warning.
- Later, when the driver's workqueue finishes processing, it calls
`fib_info_put(fi)`, which drops the reference count from 1 to 0 and calls
`free_fib_info()` again, leading to a highly exploitable **double free**.
3. **Triggerability:**
- To trigger this bug, one must cause `fib_net_dump()` to run concurrently
with route deletion.
- `fib_net_dump()` is exclusively called by `register_fib_notifier()`.
- `register_fib_notifier()` is only called during the initialization of
specific hardware switch drivers (`mlxsw`, `prestera`, `mlx5`, `rocker`) or
during `devlink` reload operations (e.g., `DEVLINK_CMD_RELOAD` or when a
network namespace containing a devlink instance is destroyed).
- Loading these drivers, invoking `DEVLINK_CMD_RELOAD`, or moving a devlink
instance into a network namespace all strictly require `CAP_NET_ADMIN` in the
initial network namespace (root privileges).
- Therefore, an unprivileged user cannot trigger `register_fib_notifier()`,
even from within a user namespace.
**Conclusion:**
The bug is a memory corruption issue (use-after-free and double free) that is
technically exploitable and can cause a Denial of Service (kernel crash).
However, it is not accessible to unprivileged users or from within user
namespaces, as triggering the vulnerable code path requires root privileges to
initialize hardware switch drivers or reload devlink instances.]
|