| 0/0 |
2026/03/10 00:37 |
flow |
repro |
19m
Error:
|
| 1/1 |
2026/03/10 00:37 |
action |
kernel-checkouter |
0m
Results: map[KernelSrc:/usr/local/google/home/dvyukov/syzkaller/agent/workdir/cache/src/859542951139fae2ad13d9c5018e94599c0eb588]
|
| 2/1 |
2026/03/10 00:38 |
action |
kernel-builder |
5m
Results: map[KernelObj:/usr/local/google/home/dvyukov/syzkaller/agent/workdir/cache/build/6725fe165b0d885ff218a5bc780f19236c05503b]
|
| 3/1 |
2026/03/10 00:43 |
action |
codesearch-prepare |
6m
Results: map[Index:codesearch-index]
|
| 4/1 |
2026/03/10 00:50 |
agent |
crash-repro-finder |
6m
Model:
Error:
Instruction:
You are an expert in the Linux kernel fuzzing. Your goal is to write a syzkaller program to trigger a specific bug.
In the final output provide only the syz program that triggers the bug, and could be executed directly,
without backticks.
Don't make assumptions about the kernel source code, use the provided codesearch tools
to examine the kernel code instead.
Document about syzkaller program syntax:
===
# Program syntax
Syzkaller uses a compact domain-specific language (DSL) for programs
to log executed programs, test its code, and persist programs in the
corpus. This page provides a brief description of the corresponding
syntax. Some useful information can also be found in the
[existing examples](/sys/linux/test) and in the program
[deserialization code](/prog/encoding.go).
Together with execution options, the DSL provides everything that
syz-executor needs to run a program.
For example, consider the program:
```
r0 = syz_open_dev$loop(&(0x7f00000011c0), 0x0, 0x0)
r1 = openat$6lowpan_control(0xffffffffffffff9c, &(0x7f00000000c0), 0x2, 0x0)
ioctl$LOOP_SET_FD(r0, 0x4c00, r1)
```
Each line in this program describes a particular syscall invocation,
with the first two calls saving the result in temporary variables `r0`
and `r1`, which are passed to the third call.
```
line = assignment | call
assignment = variable " = " call
call = syscall-name "(" [arg ["," arg]*] ")" ["(" [call-prop ["," call-prop*] ")"]
arg = "nil" | "AUTO" | const-arg | resource-arg | result-arg | pointer-arg | string-arg | struct-arg | array-arg | union-arg
const-arg = "0x" hex-integer
resource-arg = variable ["/" hex-integer] ["+" hex-integer]
result-arg = "<" variable "=>" arg
pointer-arg = "&" pointer-arg-addr ["=ANY"] "=" arg
pointer-arg-addr = "AUTO" | "(" pointer-addr ["/" region-size] ")"
string-arg = "'" escaped-string "'" | "\"" escaped-string "\"" | "\"$" escaped-string "\""
struct-arg = "{" [arg ["," arg]*] "}"
array-arg = "[" [arg ["," arg]*] "]"
union-arg = "@" field-name ["=" arg]
call-prop = prop-name ": " prop-value
variable = "r" dec-integer
pointer-addr = hex-integer
region-size = hex-integer
```
Programs may also contain blank lines and comments.
```
# Obtain a file handle
r0 = openat(0xffffffffffffff9c, &AUTO='./file1\x00', 0x42, 0x1ff)
# Perform a write operation
write(r0, &AUTO="01010101", 0x4)
```
### Memory management
Memory management is performed by syzkaller itself. It will allocate
virtual memory regions of the necessary size and set the final values
of pointer arguments.
By using the `AUTO` keyword, programs can give syzkaller the full
control over storing the data. This may be convenient e.g. when a
parameter must be passed by reference, but the exact location of its
value is not of particular importance.
```
r1 = syz_genetlink_get_family_id$nl80211(&AUTO='nl80211\x00', 0xffffffffffffffff)
ioctl$sock_SIOCGIFINDEX_80211(r0, 0x8933, &AUTO={'wlan0\x00', <r2=>0x0})
```
Alternatively, some data can be "anchored" to specific addresses. It
may be especially important when a memory region must be shared
between multiple calls. In this case, pointer addresses must be given
at the 0x7f0000000000 offset. Before the actual execution, syzkaller
will adjust pointers to the start of the actual mmap'ed region.
### Call properties
Call properties specify extra information about how a specific call
must be executed. Each call within a program has its own set of call
properties. If no properties are provided, syzkaller takes the default
ones.
Currently, syzkaller supports the following call properties.
#### Fault injection
Syntax: `fail_nth: N`.
It takes an integer (base 10) argument `N`. If the argument is
non-negative, a fault will be injected into the `N`-th occasion.
```
r0 = openat$6lowpan_control(0xffffffffffffff9c, &(0x7f00000000c0), 0x2, 0x0)
ioctl$LOOP_SET_FD(r0, 0x4c00, r0) (fail_nth: 5)
```
#### Async
Syntax: `async`.
Instructs `syz-executor` not to wait until the call completes and
to proceed immediately to the next call.
```
r0 = openat(0xffffffffffffff9c, &AUTO='./file1\x00', 0x42, 0x1ff)
write(r0, &AUTO="01010101", 0x4) (async)
read(r0, &AUTO=""/4, 0x4)
close(r0)
```
When setting `async` flags be aware of the following considerations:
* Such programs should only be executed in threaded mode (i.e. `-threaded`
flag must be passed to `syz-executor`.
* Each `async` call is executed in a separate thread and there's a
limited number of available threads (`kMaxThreads = 16`).
* If an `async` call produces a resource, keep in mind that some other call
might take it as input and `syz-executor` will just pass 0 if the resource-
producing call has not finished by that time.
===
Document about syzlang system call descriptions syntax:
===
# Syscall description language
aka `syzlang` (`[siːzˈlæŋg]`)
Pseudo-formal grammar of syscall description:
```
syscallname "(" [arg ["," arg]*] ")" [type] ["(" attribute* ")"]
arg = argname type
argname = identifier
type = typename [ "[" type-options "]" ]
typename = "const" | "intN" | "intptr" | "flags" | "array" | "ptr" |
"string" | "filename" | "glob" | "len" |
"bytesize" | "bytesizeN" | "bitsize" | "vma" | "proc" |
"compressed_image"
type-options = [type-opt ["," type-opt]]
```
common type-options include:
```
"opt" - the argument is optional (like mmap fd argument, or accept peer argument)
```
rest of the type-options are type-specific:
```
"const": integer constant, type-options:
value, underlying type (one of "intN", "intptr")
"intN"/"intptr": an integer without a particular meaning, type-options:
either an optional range of values (e.g. "5:10", or "100:200")
or a reference to flags description (see below),
or a single value
optionally followed by an alignment parameter if using a range
"flags": a set of values, type-options:
reference to flags description (see below), underlying int type (e.g. "int32")
"array": a variable/fixed-length array, type-options:
type of elements, optional size (fixed "5", or ranged "5:10", boundaries inclusive)
"ptr"/"ptr64": a pointer to an object, type-options:
direction (in/out/inout); type of the object
ptr64 has size of 8 bytes regardless of target pointer size
"string": a zero-terminated memory buffer (no pointer indirection implied), type-options:
either a string value in quotes for constant strings (e.g. "foo" or `deadbeef` for hex literal),
or a reference to string flags (special value `filename` produces file names),
optionally followed by a buffer size (string values will be padded with \x00 to that size)
"stringnoz": a non-zero-terminated memory buffer (no pointer indirection implied), type-options:
either a string value in quotes for constant strings (e.g. "foo" or `deadbeef` for hex literal),
or a reference to string flags,
"glob": glob pattern to match on the target files, type-options:
a pattern string in quotes (syntax: https://golang.org/pkg/path/filepath/#Match)
(e.g. "/sys/" or "/sys/**/*"),
or include exclude glob too (e.g. "/sys/**/*:-/sys/power/state")
"fmt": a string representation of an integer (not zero-terminated), type-options:
format (one of "dec", "hex", "oct") and the value (a resource, int, flags or proc)
the resulting data is always fixed-size (formatted as "%020llu", "0x%016llx" or "%023llo", respectively)
"len": length of another field (for array it is number of elements), type-options:
argname of the object
"bytesize": similar to "len", but always denotes the size in bytes, type-options:
argname of the object
"bitsize": similar to "len", but always denotes the size in bits, type-options:
argname of the object
"offsetof": offset of the field from the beginning of the parent struct, type-options:
field
"vma"/"vma64": a pointer to a set of pages (used as input for mmap/munmap/mremap/madvise), type-options:
optional number of pages (e.g. vma[7]), or a range of pages (e.g. vma[2-4])
vma64 has size of 8 bytes regardless of target pointer size
"proc": per process int (see description below), type-options:
value range start, how many values per process, underlying type
"compressed_image": zlib-compressed disk image
syscalls accepting compressed images must be marked with `no_generate`
and `no_minimize` call attributes. if the content of the decompressed image
can be checked by a `fsck`-like command, use the `fsck` syscall attribute
"text": machine code of the specified type, type-options:
text type (x86_real, x86_16, x86_32, x86_64, arm64)
"void": type with static size 0
mostly useful inside of templates and varlen unions, can't be syscall argument
```
flags/len/flags also have trailing underlying type type-option when used in structs/unions/pointers.
Flags are described as:
```
flagname = const ["," const]*
```
or for string flags as:
```
flagname = "\"" literal "\"" ["," "\"" literal "\""]*
```
Call attributes are:
```
"disabled": the call will not be used in fuzzing; useful to temporary disable some calls
or prohibit particular argument combinations.
"timeout[N]": additional execution timeout (in ms) for the call on top of some default value.
"prog_timeout[N]": additional execution timeout (in ms) for the whole program if it contains this call;
if a program contains several such calls, the max value is used.
"ignore_return": ignore return value of this syscall in fallback feedback; need to be used for calls
that don't return fixed error codes but rather something else (e.g. the current time).
"breaks_returns": ignore return values of all subsequent calls in the program in fallback feedback (can't be trusted).
"no_generate": do not try to generate this syscall, i.e. use only seed descriptions to produce it.
"no_minimize": do not modify instances of this syscall when trying to minimize a crashing program.
"no_squash": do not attempt to pass squashed arguments to this syscall.
Without that, the fuzzer will sometimes attempt to replace complex structures with arrays of bytes,
possibly triggering interesting mutations, but also making programs hard to reason about.
"fsck": the content of the compressed buffer argument for this syscall is a file system and the
string argument is a fsck-like command that will be called to verify the filesystem.
"remote_cover": wait longer to collect remote coverage for this call.
"kfuzz_test": the call is a kfuzztest target.
"snapshot": the call is enabled by default only in snapshot fuzzing mode, but can also be enabled in
the non-snasphot mode when listed in "enable_syscalls" with its full name (as opposed to a wildcard match).
It can also always be disabled via "disable_syscalls".
The attribute is generally used to mark calls that are not safe to execute in non-snapshot mode
(can lead to false positives, or lost connections to test machines.
```
## Ints
`int8`, `int16`, `int32` and `int64` denote an integer of the corresponding size.
`intptr` denotes a pointer-sized integer, i.e. C `long` type.
By appending `be` suffix (e.g. `int16be`) integers become big-endian.
It's possible to specify a range of values for an integer in the format of `int32[0:100]` or `int32[0:4096, 512]` for a 512-aligned int.
Integers can also take a reference to flags description or a value as its first type-option.
In that case, the alignment parameter is not supported.
To denote a bitfield of size N use `int64:N`.
It's possible to use these various kinds of ints as base types for `const`, `flags`, `len` and `proc`.
```
example_struct {
f0 int8 # random 1-byte integer
f1 const[0x42, int16be] # const 2-byte integer with value 0x4200 (big-endian 0x42)
f2 int32[0:100] # random 4-byte integer with values from 0 to 100 inclusive
f3 int32[1:10, 2] # random 4-byte integer with values {1, 3, 5, 7, 9}
f4 int64:20 # random 20-bit bitfield
f5 int8[10] # const 1-byte integer with value 10
f6 int32[flagname] # random 4-byte integer from the set of values referenced by flagname
}
```
## Structs
Structs are described as:
```
structname "{" "\n"
(fieldname type ("(" fieldattribute* ")")? (if[expression])? "\n")+
"}" ("[" attribute* "]")?
```
Fields can have attributes specified in parentheses after the field, independent
of their type. `in/out/inout` attribute specify per-field direction, for example:
```
foo {
field0 const[1, int32] (in)
field1 int32 (inout)
field2 fd (out)
}
```
You may specify conditions that determine whether a field will be included:
```
foo {
field0 int32
field1 int32 (if[value[field0] == 0x1])
}
```
See [the corresponding section](syscall_descriptions_syntax.md#conditional-fields)
for more details.
`out_overlay` attribute allows to have separate input and output layouts for the struct.
Fields before the `out_overlay` field are input, fields starting from `out_overlay` are output.
Input and output fields overlap in memory (both start from the beginning of the struct in memory).
For example:
```
foo {
in0 const[1, int32]
in1 flags[bar, int8]
in2 ptr[in, string]
out0 fd (out_overlay)
out1 int32
}
```
Structs can have attributes specified in square brackets after the struct.
Attributes are:
- `packed`: the struct does not have paddings between fields and has alignment 1; this is similar to GNU C `__attribute__((packed))`; struct alignment can be overridden with `align` attribute
- `align[N]`: the struct has alignment N and padded up to multiple of `N`; contents of the padding are unspecified (though, frequently are zeros); similar to GNU C `__attribute__((aligned(N)))`
- `size[N]`: the struct is padded up to the specified size `N`; contents of the padding are unspecified (though, frequently are zeros)
## Unions
Unions are described as:
```
unionname "[" "\n"
(fieldname type (if[expression])? "\n")+
"]" ("[" attribute* "]")?
```
During fuzzing, syzkaller randomly picks one of the union options.
You may also specify conditions that determine whether the corresponding
option may or may not be selected, depending on values of other fields. See
[the corresponding section](syscall_descriptions_syntax.md#conditional-fields)
for more details.
Unions can have attributes specified in square brackets after the union.
Attributes are:
- `varlen`: union size is the size of the particular chosen option (not statically known); without this attribute unions are statically sized as maximum of all options (similar to C unions)
- `size[N]`: the union is padded up to the specified size `N`; contents of the padding are unspecified (though, frequently are zeros)
## Resources
Resources represent values that need to be passed from output of one syscall to input of another syscall. For example, `close` syscall requires an input value (fd) previously returned by `open` or `pipe` syscall. To achieve this, `fd` is declared as a resource. This is a way of modelling dependencies between syscalls, as defining a syscall as the producer of a resource and another syscall as the consumer defines a loose sense of ordering between them. Resources are described as:
```
"resource" identifier "[" underlying_type "]" [ ":" const ("," const)* ]
```
`underlying_type` is either one of `int8`, `int16`, `int32`, `int64`, `intptr` or another resource (which models inheritance, for example, a socket is a subtype of fd). The optional set of constants represent resource special values, for example, `0xffffffffffffffff` (-1) for "no fd", or `AT_FDCWD` for "the current dir". Special values are used once in a while as resource values. If no special values specified, special value of `0` is used. Resources can then be used as types, for example:
```
resource fd[int32]: 0xffffffffffffffff, AT_FDCWD, 1000000
resource sock[fd]
resource sock_unix[sock]
socket(...) sock
accept(fd sock, ...) sock
listen(fd sock, backlog int32)
```
Resources don't have to be necessarily returned by a syscall. They can be used as any other data type. For example:
```
resource my_resource[int32]
request_producer(..., arg ptr[out, my_resource])
request_consumer(..., arg ptr[inout, test_struct])
test_struct {
...
attr my_resource
}
```
For more complex producer/consumer scenarios, field attributes can be utilized.
For example:
```
resource my_resource_1[int32]
resource my_resource_2[int32]
request_produce1_consume2(..., arg ptr[inout, test_struct])
test_struct {
...
field0 my_resource_1 (out)
field1 my_resource_2 (in)
}
```
Each resource type must be "produced" (used as an output) by at least one syscall
(outside of unions and optional pointers) and "consumed" (used as an input)
by at least one syscall.
## Type Aliases
Complex types that are often repeated can be given short type aliases using the
following syntax:
```
type identifier underlying_type
```
For example:
```
type signalno int32[0:65]
type net_port proc[20000, 4, int16be]
```
Then, type alias can be used instead of the underlying type in any contexts.
Underlying type needs to be described as if it's a struct field, that is,
with the base type if it's required. However, type alias can be used as syscall
arguments as well. Underlying types are currently restricted to integer types,
`ptr`, `ptr64`, `const`, `flags` and `proc` types.
There are some builtin type aliases:
```
type bool8 int8[0:1]
type bool16 int16[0:1]
type bool32 int32[0:1]
type bool64 int64[0:1]
type boolptr intptr[0:1]
type fileoff[BASE] BASE
type filename string[filename]
type buffer[DIR] ptr[DIR, array[int8]]
```
## Type Templates
Type templates can be declared as follows:
```
type buffer[DIR] ptr[DIR, array[int8]]
type fileoff[BASE] BASE
type nlattr[TYPE, PAYLOAD] {
nla_len len[parent, int16]
nla_type const[TYPE, int16]
payload PAYLOAD
} [align_4]
```
and later used as follows:
```
syscall(a buffer[in], b fileoff[int64], c ptr[in, nlattr[FOO, int32]])
```
There is builtin type template `optional` defined as:
```
type optional[T] [
val T
void void
] [varlen]
```
## Length
You can specify length of a particular field in struct or a named argument by
using `len`, `bytesize` and `bitsize` types, for example:
```
write(fd fd, buf ptr[in, array[int8]], count len[buf])
sock_fprog {
len len[filter, int16]
filter ptr[in, array[sock_filter]]
}
```
If `len`'s argument is a pointer, then the length of the pointee argument is used.
To denote the length of a field in N-byte words use `bytesizeN`, possible values
for N are 1, 2, 4 and 8.
To denote the length of the parent struct, you can use `len[parent, int8]`.
To denote the length of the higher level parent when structs are embedded into
one another, you can specify the type name of the particular parent:
```
s1 {
f0 len[s2] # length of s2
}
s2 {
f0 s1
f1 array[int32]
f2 len[parent, int32]
}
```
`len` argument can also be a path expression which allows more complex
addressing. Path expressions are similar to C field references, but also allow
referencing parent and sibling elements. A special reference `syscall` used
in the beginning of the path allows to refer directly to the syscall arguments.
For example:
```
s1 {
a ptr[in, s2]
b ptr[in, s3]
c array[int8]
}
s2 {
d array[int8]
}
s3 {
# This refers to the array c in the parent s1.
e len[s1:c, int32]
# This refers to the array d in the sibling s2.
f len[s1:a:d, int32]
# This refers to the array k in the child s4.
g len[i:j, int32]
# This refers to syscall argument l.
h len[syscall:l, int32]
i ptr[in, s4]
}
s4 {
j array[int8]
}
foo(k ptr[in, s1], l ptr[in, array[int8]])
```
## Proc
The `proc` type can be used to denote per process integers.
The idea is to have a separate range of values for each executor, so they don't interfere.
The simplest example is a port number.
The `proc[20000, 4, int16be]` type means that we want to generate an `int16be`
integer starting from `20000` and assign `4` values for each process.
As a result the executor number `n` will get values in the `[20000 + n * 4, 20000 + (n + 1) * 4)` range.
## Integer Constants
Integer constants can be specified as decimal literals, as `0x`-prefixed
hex literals, as `'`-surrounded char literals, or as symbolic constants
extracted from kernel headers or defined by `define` directives. For example:
```
foo(a const[10], b const[-10])
foo(a const[0xabcd])
foo(a int8['a':'z'])
foo(a const[PATH_MAX])
foo(a int32[PATH_MAX])
foo(a ptr[in, array[int8, MY_PATH_MAX]])
define MY_PATH_MAX PATH_MAX + 2
```
## Conditional fields
### In structures
In syzlang, it's possible to specify a condition for every struct field that
determines whether the field should be included or omitted:
```
header_fields {
magic const[0xabcd, int16]
haveInteger int8
} [packed]
packet {
header header_fields
integer int64 (if[value[header:haveInteger] == 0x1])
body array[int8]
} [packed]
some_call(a ptr[in, packet])
```
In this example, the `packet` structure will include the field `integer` only
if `header.haveInteger == 1`. In memory, `packet` will have the following
layout:
| header.magic = 0xabcd | header.haveInteger = 0x1 | integer | body |
| --------------------- | ------------------------ | ------- | ---- |
That corresponds to e.g. the following program:
```
some_call(&AUTO={{AUTO, 0x1}, @value=0xabcd, []})
```
If `header.haveInteger` is not `1`, syzkaller will just pretend that the field
`integer` does not exist.
```
some_call(&AUTO={{AUTO, 0x0}, @void, []})
```
| header.magic = 0xabcd | header.haveInteger = 0x0 | body |
| --------------------- | ------------------------ | ---- |
Every conditional field is assumed to be of variable length and so is the struct
to which this field belongs.
When a variable length field appears in the middle of a structure, the structure
must be marked with `[packed].`
Conditions on bitfields are prohibited:
```
struct {
f0 int
f1 int:3 (if[value[f0] == 0x1]) # It will not compile.
}
```
But you may reference bitfields in your conditions:
```
struct {
f0 int:1
f1 int:7
f2 int (if[value[f0] == value[f1]])
} [packed]
```
### In unions
Let's consider the following example.
```
struct {
type int
body alternatives
}
alternatives [
int int64 (if[value[struct:type] == 0x1])
arr array[int64, 5] (if[value[struct:type] == 0x2])
default int32
] [varlen]
some_call(a ptr[in, struct])
```
In this case, the union option will be selected depending on the value of the
`type` field. For example, if `type` is `0x1`, then it can be either `int` or
`default`:
```
some_call(&AUTO={0x1, @int=0x123})
some_call(&AUTO={0x1, @default=0x123})
```
If `type` is `0x2`, it can be either `arr` or `default`.
If `type` is neither `0x1` nor `0x2`, syzkaller may only select `default`:
```
some_call(&AUTO={0x0, @default=0xabcd})
```
To ensure that a union can always be constructed, the last union field **must always
have no condition**.
Thus, the following definition would fail to compile:
```
alternatives [
int int64 (if[value[struct:type] == 0x1])
arr array[int64, 5] (if[value[struct:type] == 0x1])
] [varlen]
```
During prog mutation and generation syzkaller will select a random union field
whose condition is satisfied.
### Expression syntax
Currently, only `==`, `!=`, `&` and `||` operators are supported. However, the
functionality was designed in such a way that adding more operators is easy.
Feel free to file a GitHub issue or write us an email in case it's needed.
Expressions are evaluated as `int64` values. If the final result of an
expression is not 0, it's assumed to be satisfied.
If you want to reference a field's value, you can do it via
`value[path:to:field]`, which is similar to the `len[]` argument.
```
sub_struct {
f0 int
# Reference a field in a parent struct.
f1 int (if[value[struct:f2]]) # Same as if[value[struct:f2] != 0].
}
struct {
f2 int
f3 sub_struct
f4 int (if[value[f2] == 0x2]) # Reference a sibling field.
f5 int (if[value[f3:f0] == 0x1]) # Reference a nested field.
f6 int (if[value[f3:f0] == 0x1 || value[f3:f0] == 0x2]) # Reference a nested field which either equals to 0x1 or 0x2.
} [packed]
call(a ptr[in, struct])
```
The referenced field must be of integer type and there must be no
conditional fields in the path to it. For example, the following
descriptions will not compile.
```
struct {
f0 int
f1 int (if[value[f0] == 0x1])
f2 int (if[value[f1] == 0x1])
}
```
You may also reference constants in expressions:
```
struct {
f0 int
f1 int
f2 int (if[value[f0] & SOME_CONST == OTHER_CONST])
}
```
## Meta
Description files can also contain `meta` directives that specify meta-information for the whole file.
```
meta noextract
```
Tells `make extract` to not extract constants for this file.
Though, `syz-extract` can still be invoked manually on this file.
```
meta arches["arch1", "arch2"]
```
Restricts this file only to the given set of architectures.
`make extract` and `make generate` will not use it on other architectures.
## Misc
Description files also contain `include` directives that refer to Linux kernel header files,
`incdir` directives that refer to custom Linux kernel header directories
and `define` directives that define symbolic constant values.
The syzkaller executor defines some [pseudo system calls](./pseudo_syscalls.md)
that can be used as any other syscall in a description file. These pseudo
system calls expand to literal C code and can perform user-defined
custom actions. You can find some examples in
[executor/common_linux.h](../executor/common_linux.h).
Also see [tips](syscall_descriptions.md#tips) on writing good descriptions.
===
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt:
Bug title: INFO: task hung in rxrpc_destroy_all_calls
The bug report to reproduce:
INFO: task kworker/u8:22:6285 blocked for more than 144 seconds.
Tainted: G L syzkaller #0
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
task:kworker/u8:22 state:D stack:20960 pid:6285 tgid:6285 ppid:2 task_flags:0x4208160 flags:0x00080000
Workqueue: netns cleanup_net
Call Trace:
<TASK>
context_switch kernel/sched/core.c:5295 [inline]
__schedule+0x14fb/0x52c0 kernel/sched/core.c:6907
__schedule_loop kernel/sched/core.c:6989 [inline]
schedule+0x164/0x360 kernel/sched/core.c:7004
rxrpc_destroy_all_calls+0x57e/0x6a0 net/rxrpc/call_object.c:757
rxrpc_exit_net+0x6f/0xc0 net/rxrpc/net_ns.c:112
ops_exit_list net/core/net_namespace.c:199 [inline]
ops_undo_list+0x49f/0x940 net/core/net_namespace.c:252
cleanup_net+0x56e/0x800 net/core/net_namespace.c:704
process_one_work kernel/workqueue.c:3275 [inline]
process_scheduled_works+0xb02/0x1830 kernel/workqueue.c:3358
worker_thread+0xa50/0xfc0 kernel/workqueue.c:3439
kthread+0x388/0x470 kernel/kthread.c:467
ret_from_fork+0x51e/0xb90 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
</TASK>
Showing all locks held in the system:
4 locks held by pr/legacy/17:
1 lock held by khungtaskd/38:
#0: ffffffff8ddcd780 (rcu_read_lock){....}-{1:3}, at: rcu_lock_acquire include/linux/rcupdate.h:312 [inline]
#0: ffffffff8ddcd780 (rcu_read_lock){....}-{1:3}, at: rcu_read_lock include/linux/rcupdate.h:850 [inline]
#0: ffffffff8ddcd780 (rcu_read_lock){....}-{1:3}, at: debug_show_all_locks+0x2e/0x180 kernel/locking/lockdep.c:6775
2 locks held by getty/5556:
#0: ffff8880370370a0 (&tty->ldisc_sem){++++}-{0:0}, at: tty_ldisc_ref_wait+0x25/0x70 drivers/tty/tty_ldisc.c:243
#1: ffffc90003e7e2e0 (&ldata->atomic_read_lock){+.+.}-{4:4}, at: n_tty_read+0x462/0x13c0 drivers/tty/n_tty.c:2211
6 locks held by kworker/1:4/5852:
3 locks held by kworker/1:7/5882:
3 locks held by kworker/u8:20/6283:
3 locks held by kworker/u8:22/6285:
#0: ffff88801aed9138 ((wq_completion)netns){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3250 [inline]
#0: ffff88801aed9138 ((wq_completion)netns){+.+.}-{0:0}, at: process_scheduled_works+0x9ea/0x1830 kernel/workqueue.c:3358
#1: ffffc90003dd7c40 (net_cleanup_work){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3251 [inline]
#1: ffffc90003dd7c40 (net_cleanup_work){+.+.}-{0:0}, at: process_scheduled_works+0xa25/0x1830 kernel/workqueue.c:3358
#2: ffffffff8f14b940 (pernet_ops_rwsem){++++}-{4:4}, at: cleanup_net+0xf4/0x800 net/core/net_namespace.c:675
10 locks held by syz-executor/9410:
4 locks held by udevd/18380:
#0: ffff888039337b80 (&p->lock){+.+.}-{4:4}, at: seq_read_iter+0xb8/0xe20 fs/seq_file.c:183
#1: ffff8880393f4878 (&of->mutex#2){+.+.}-{4:4}, at: kernfs_seq_start+0x5c/0x420 fs/kernfs/file.c:172
#2: ffff88804f2e0788 (kn->active#31){++++}-{0:0}, at: kernfs_get_active_of fs/kernfs/file.c:80 [inline]
#2: ffff88804f2e0788 (kn->active#31){++++}-{0:0}, at: kernfs_seq_start+0xb2/0x420 fs/kernfs/file.c:173
#3: ffff888063386188 (&dev->mutex){....}-{4:4}, at: device_lock_interruptible include/linux/device.h:900 [inline]
#3: ffff888063386188 (&dev->mutex){....}-{4:4}, at: manufacturer_show+0x26/0xa0 drivers/usb/core/sysfs.c:142
1 lock held by syz-executor/18477:
4 locks held by syz.2.4290/18524:
5 locks held by syz.3.4295/18544:
=============================================
NMI backtrace for cpu 1
CPU: 1 UID: 0 PID: 38 Comm: khungtaskd Tainted: G L syzkaller #0 PREEMPT_{RT,(full)}
Tainted: [L]=SOFTLOCKUP
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 02/12/2026
Call Trace:
<TASK>
dump_stack_lvl+0xe8/0x150 lib/dump_stack.c:120
nmi_cpu_backtrace+0x274/0x2d0 lib/nmi_backtrace.c:113
nmi_trigger_cpumask_backtrace+0x17a/0x300 lib/nmi_backtrace.c:62
trigger_all_cpu_backtrace include/linux/nmi.h:161 [inline]
__sys_info lib/sys_info.c:157 [inline]
sys_info+0x135/0x170 lib/sys_info.c:165
check_hung_uninterruptible_tasks kernel/hung_task.c:346 [inline]
watchdog+0xfd9/0x1030 kernel/hung_task.c:515
kthread+0x388/0x470 kernel/kthread.c:467
ret_from_fork+0x51e/0xb90 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
</TASK>
Sending NMI from CPU 1 to CPUs 0:
NMI backtrace for cpu 0
CPU: 0 UID: 0 PID: 18544 Comm: syz.3.4295 Tainted: G L syzkaller #0 PREEMPT_{RT,(full)}
Tainted: [L]=SOFTLOCKUP
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 02/12/2026
RIP: 0010:check_kcov_mode kernel/kcov.c:183 [inline]
RIP: 0010:__sanitizer_cov_trace_pc+0x17/0x80 kernel/kcov.c:217
Code: cc cc cc 90 90 90 90 90 90 90 90 90 90 90 90 90 90 90 90 f3 0f 1e fa 48 8b 04 24 65 48 8b 0d c8 ac 89 10 65 8b 35 e9 ac 89 10 <81> e6 00 00 ff 00 ba 00 01 00 00 23 91 64 0b 00 00 89 d7 09 f7 74
RSP: 0018:ffffc9000639f968 EFLAGS: 00000282
RAX: ffffffff849be5ed RBX: 0000000000000008 RCX: ffff88803d470000
RDX: 0000000000000000 RSI: 0000000080000000 RDI: 0000000000000000
RBP: 0000000000000000 R08: ffffc9000639fba0 R09: ffffc9000639fcf0
R10: ffffc9000639fd30 R11: fffff52000c73fa8 R12: 0000000000000000
R13: 0000000000000000 R14: ffffc9000639fba0 R15: dffffc0000000000
FS: 00007f6ef91456c0(0000) GS:ffff888126340000(0000) knlGS:0000000000000000
CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 00002000000c2030 CR3: 0000000033a5a000 CR4: 00000000003526f0
DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7: 0000000000000600
Call Trace:
<TASK>
__import_iovec+0x3d/0x7e0 lib/iov_iter.c:1367
import_iovec+0x73/0xa0 lib/iov_iter.c:1440
copy_msghdr_from_user net/socket.c:2529 [inline]
recvmsg_copy_msghdr net/socket.c:2778 [inline]
___sys_recvmsg+0x4bd/0x590 net/socket.c:2850
do_recvmmsg+0x33a/0x800 net/socket.c:2949
__sys_recvmmsg net/socket.c:3023 [inline]
__do_sys_recvmmsg net/socket.c:3046 [inline]
__se_sys_recvmmsg net/socket.c:3039 [inline]
__x64_sys_recvmmsg+0x198/0x250 net/socket.c:3039
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x14d/0xf80 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7f6efaf0c799
Code: ff c3 66 2e 0f 1f 84 00 00 00 00 00 0f 1f 44 00 00 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 c7 c1 e8 ff ff ff f7 d8 64 89 01 48
RSP: 002b:00007f6ef9145028 EFLAGS: 00000246 ORIG_RAX: 000000000000012b
RAX: ffffffffffffffda RBX: 00007f6efb186090 RCX: 00007f6efaf0c799
RDX: 0000000000010106 RSI: 00002000000000c0 RDI: 0000000000000003
RBP: 00007f6efafa2bd9 R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000002 R11: 0000000000000246 R12: 0000000000000000
R13: 00007f6efb186128 R14: 00007f6efb186090 R15: 00007ffd940e99c8
</TASK>
The list of existing description files:
acpi_thermal_rel.txt
aio.txt
auto.txt
binfmt.txt
bpf.txt
bpf_prog.txt
bpf_trace.txt
cgroup.txt
damon.txt
dev_ashmem.txt
dev_bifrost.txt
dev_binder.txt
dev_binderfs.txt
dev_block.txt
dev_bsg.txt
dev_bus_usb.txt
dev_camx.txt
dev_cdrom.txt
dev_cec.txt
dev_char_usb.txt
dev_comedi.txt
dev_dma_heap.txt
dev_dri.txt
dev_dsp.txt
dev_dvb_demux.txt
dev_dvb_dvr.txt
dev_dvb_frontend.txt
dev_fb.txt
dev_floppy.txt
dev_hidraw.txt
dev_i2c.txt
dev_i915.txt
dev_img_rogue.txt
dev_infiniband_rdma.txt
dev_infiniband_rdma_cm.txt
dev_input.txt
dev_iommu.txt
dev_kvm.txt
dev_kvm_amd64.txt
dev_kvm_arm64.txt
dev_kvm_extra.txt
dev_kvm_riscv64.txt
dev_loop.txt
dev_mali.txt
dev_media.txt
dev_msm.txt
dev_msr.txt
dev_nbd.txt
dev_net_tun.txt
dev_panthor.txt
dev_ppp.txt
dev_ptmx.txt
dev_ptp.txt
dev_qat_adf_ctl.txt
dev_qrtr_tun.txt
dev_random.txt
dev_rfkill.txt
dev_rtc.txt
dev_sequencer.txt
dev_sg.txt
dev_snapshot.txt
dev_snd_control.txt
dev_snd_hw.txt
dev_snd_midi.txt
dev_snd_pcm.txt
dev_snd_seq.txt
dev_snd_timer.txt
dev_sr.txt
dev_sw_sync.txt
dev_tlk_device.txt
dev_trusty.txt
dev_udmabuf.txt
dev_uhid.txt
dev_uinput.txt
dev_usb_hiddev.txt
dev_usbmon.txt
dev_userio.txt
dev_vfio.txt
dev_vga_arbiter.txt
dev_vhci.txt
dev_video4linux.txt
dev_video4linux_vim2m.txt
dev_virtual_nci.txt
dev_vtpm.txt
fanotify.txt
filesystem.txt
fs_9p.txt
fs_fuse.txt
fs_incfs.txt
fs_ioctl.txt
fs_ioctl_autofs.txt
fs_ioctl_btrfs.txt
fs_ioctl_ext4.txt
fs_ioctl_f2fs.txt
fs_ioctl_fat.txt
fs_ioctl_fscrypt.txt
fs_ioctl_fsverity.txt
fs_ioctl_xfs.txt
futex.txt
hafnium.txt
inotify.txt
io_uring.txt
ipc.txt
key.txt
kfuzztest.txt
l2cap.txt
landlock.txt
lsm.txt
namespaces.txt
net_80211.txt
netfilter.txt
netfilter_arp.txt
netfilter_bridge.txt
netfilter_ipv4.txt
netfilter_ipv6.txt
netfilter_ipvs.txt
netfilter_targets.txt
pagemap_ioctl.txt
perf.txt
prctl.txt
quotactl.txt
seccomp.txt
security_apparmor.txt
security_selinux.txt
security_smack.txt
socket.txt
socket_alg.txt
socket_ax25.txt
socket_bluetooth.txt
socket_caif.txt
socket_can.txt
socket_ieee802154.txt
socket_inet.txt
socket_inet6.txt
socket_inet_dccp.txt
socket_inet_icmp.txt
socket_inet_igmp.txt
socket_inet_l2tp.txt
socket_inet_sctp.txt
socket_inet_tcp.txt
socket_inet_udp.txt
socket_ip_tunnel.txt
socket_isdn.txt
socket_kcm.txt
socket_key.txt
socket_llc.txt
socket_netlink.txt
socket_netlink_audit.txt
socket_netlink_crypto.txt
socket_netlink_generic.txt
socket_netlink_generic_80211.txt
socket_netlink_generic_batadv.txt
socket_netlink_generic_devlink.txt
socket_netlink_generic_ethtool.txt
socket_netlink_generic_fou.txt
socket_netlink_generic_gtp.txt
socket_netlink_generic_mptcp.txt
socket_netlink_generic_net_dm.txt
socket_netlink_generic_netlabel.txt
socket_netlink_generic_nfc.txt
socket_netlink_generic_seg6.txt
socket_netlink_generic_smc.txt
socket_netlink_generic_team.txt
socket_netlink_generic_wireguard.txt
socket_netlink_netfilter.txt
socket_netlink_netfilter_acct.txt
socket_netlink_netfilter_conntrack.txt
socket_netlink_netfilter_conntrack_exp.txt
socket_netlink_netfilter_conntrack_helper.txt
socket_netlink_netfilter_ipset.txt
socket_netlink_netfilter_log.txt
socket_netlink_netfilter_nft_compat.txt
socket_netlink_netfilter_nftables.txt
socket_netlink_netfilter_osf.txt
socket_netlink_netfilter_queue.txt
socket_netlink_netfilter_timeout.txt
socket_netlink_rdma.txt
socket_netlink_route.txt
socket_netlink_route_sched.txt
socket_netlink_route_sched_retired.txt
socket_netlink_sock_diag.txt
socket_netlink_xfrm.txt
socket_netrom.txt
socket_nfc.txt
socket_nvme_of_tcp.txt
socket_packet.txt
socket_phonet.txt
socket_pppox.txt
socket_qipcrtr.txt
socket_rds.txt
socket_rose.txt
socket_rxrpc.txt
socket_tipc.txt
socket_tipc_netlink.txt
socket_unix.txt
socket_vnet.txt
socket_x25.txt
socket_xdp.txt
sys.txt
uffd.txt
usbip.txt
vmw_vmci.txt
vnet.txt
vnet_mptcp.txt
vusb.txt
watch_queue.txt
xattr.txt
|
| 5/2 |
2026/03/10 00:50 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 16208
output: 47
thoughts: 179
Thoughts: **Debugging a Network Namespace Hang**
Okay, so I've got this bug report, "INFO: task hung in rxrpc_destroy_all_calls." That's not ideal, and it's pointing towards a problem when a network namespace is being destroyed. The stack trace is pretty clear: we're deep within the RXRPC subsystem, specifically in `rxrpc_destroy_all_calls`. We've got `rxrpc_exit_net`, then the generic `ops_exit_list` and `ops_undo_list` which are common in net cleanup. Finally, we're in `cleanup_net`, so the flow makes sense: we're cleaning up the network namespace.
Wait a minute, there's another stack trace. This one shows a task in `__import_iovec`, and that goes down to `copy_msghdr_from_user` and `recvmsg_copy_msghdr`, eventually ending up in `___sys_recvmsg` and then `do_recvmmsg`, and ultimately in `__sys_recvmmsg`. That's *another* task, and it's also in `do_recvmmsg`? Huh.
Okay, let's step back and focus on `rxrpc_destroy_all_calls` and figure out what might be causing the hang. I need to get into the code now and see what that function is doing during network namespace destruction.
|
| 6/2 |
2026/03/10 00:50 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_object.c IncludeLines:true Name:rxrpc_destroy_all_calls]
Results: map[SourceCode: 721: /*
722: * Make sure that all calls are gone from a network namespace. To reach this
723: * point, any open UDP sockets in that namespace must have been closed, so any
724: * outstanding calls cannot be doing I/O.
725: */
726: void rxrpc_destroy_all_calls(struct rxrpc_net *rxnet)
727: {
728: struct rxrpc_call *call;
729:
730: _enter("");
731:
732: if (!list_empty(&rxnet->calls)) {
733: spin_lock(&rxnet->call_lock);
734:
735: while (!list_empty(&rxnet->calls)) {
736: call = list_entry(rxnet->calls.next,
737: struct rxrpc_call, link);
738: _debug("Zapping call %p", call);
739:
740: rxrpc_see_call(call, rxrpc_call_see_zap);
741: list_del_init(&call->link);
742:
743: pr_err("Call %p still in use (%d,%s,%lx,%lx)!\n",
744: call, refcount_read(&call->ref),
745: rxrpc_call_states[__rxrpc_call_state(call)],
746: call->flags, call->events);
747:
748: spin_unlock(&rxnet->call_lock);
749: cond_resched();
750: spin_lock(&rxnet->call_lock);
751: }
752:
753: spin_unlock(&rxnet->call_lock);
754: }
755:
756: atomic_dec(&rxnet->nr_calls);
757: wait_var_event(&rxnet->nr_calls, !atomic_read(&rxnet->nr_calls));
758: }
SourceFile:net/rxrpc/call_object.c]
|
| 7/2 |
2026/03/10 00:50 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 17351
output: 60
thoughts: 65
Thoughts: **Debugging a Persistent Hang**
Okay, so I'm stuck. The program's hanging right here: `wait_var_event(&rxnet->nr_calls, !atomic_read(&rxnet->nr_calls));`. That means `rxnet->nr_calls` isn't decrementing to zero, and the wait is never satisfied. I need to figure out where that counter, `rxnet->nr_calls`, is being modified. Time to track down all the read and write accesses to that variable and understand its lifecycle within the system. Let's see if we can identify where things are going wrong.
|
| 8/2 |
2026/03/10 00:50 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/rxrpc/call_object.c IncludeSnippetLines:1 Name:rxrpc_net::nr_calls SourceTreePrefix:net/rxrpc/]
Results: map[References:[map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:rxrpc_alloc_call SourceFile:net/rxrpc/call_object.c SourceLine:181 SourceSnippet: 180: call->rtt_avail = RXRPC_CALL_RTT_AVAIL_MASK;
181: atomic_inc(&rxnet->nr_calls);
182: return call;
] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:rxrpc_destroy_all_calls SourceFile:net/rxrpc/call_object.c SourceLine:756 SourceSnippet: 755:
756: atomic_dec(&rxnet->nr_calls);
757: wait_var_event(&rxnet->nr_calls, !atomic_read(&rxnet->nr_calls));
] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:rxrpc_destroy_all_calls SourceFile:net/rxrpc/call_object.c SourceLine:757 SourceSnippet: 756: atomic_dec(&rxnet->nr_calls);
757: wait_var_event(&rxnet->nr_calls, !atomic_read(&rxnet->nr_calls));
758: }
] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:rxrpc_destroy_all_calls SourceFile:net/rxrpc/call_object.c SourceLine:757 SourceSnippet: 756: atomic_dec(&rxnet->nr_calls);
757: wait_var_event(&rxnet->nr_calls, !atomic_read(&rxnet->nr_calls));
758: }
] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:rxrpc_destroy_all_calls SourceFile:net/rxrpc/call_object.c SourceLine:757 SourceSnippet: 756: atomic_dec(&rxnet->nr_calls);
757: wait_var_event(&rxnet->nr_calls, !atomic_read(&rxnet->nr_calls));
758: }
] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:rxrpc_destroy_all_calls SourceFile:net/rxrpc/call_object.c SourceLine:757 SourceSnippet: 756: atomic_dec(&rxnet->nr_calls);
757: wait_var_event(&rxnet->nr_calls, !atomic_read(&rxnet->nr_calls));
758: }
] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:rxrpc_init_net SourceFile:net/rxrpc/net_ns.c SourceLine:45 SourceSnippet: 44: spin_lock_init(&rxnet->call_lock);
45: atomic_set(&rxnet->nr_calls, 1);
46:
] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:rxrpc_rcu_free_call SourceFile:net/rxrpc/call_object.c SourceLine:676 SourceSnippet: 675: kmem_cache_free(rxrpc_call_jar, call);
676: if (atomic_dec_and_test(&rxnet->nr_calls))
677: wake_up_var(&rxnet->nr_calls);
] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:rxrpc_rcu_free_call SourceFile:net/rxrpc/call_object.c SourceLine:677 SourceSnippet: 676: if (atomic_dec_and_test(&rxnet->nr_calls))
677: wake_up_var(&rxnet->nr_calls);
678: }
]] TruncatedOutput:false]
|
| 9/2 |
2026/03/10 00:50 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 18717
output: 43
thoughts: 314
Thoughts: **Analyzing Potential Reference Leaks in RxRPC Calls**
Okay, so `nr_calls` is the key counter here. It starts at 1, gets bumped up when a new call object is created (`rxrpc_alloc_call`), and should go back down when the call is released (`rxrpc_rcu_free_call`, `rxrpc_destroy_all_calls`). The hang in `rxrpc_destroy_all_calls` is a classic signal of a leaked `rxrpc_call` object – a reference that's not being properly decremented.
Now, how could that happen? We're freeing them through `rxrpc_rcu_free_call`, which is called as a callback (`call_rcu`) from `rxrpc_put_call`. That's the normal path for reference counting, so a leak has to be related to reference counting.
Let's switch gears and examine the other stack trace. It looks like we have a `recvmmsg` syscall involved. Interesting. The call stack points to `__import_iovec`, which means we're dealing with a copy from userspace. A stalled userspace operation, say a userfaultfd or a slow read from userspace, could be the culprit. Does the kernel hold a reference to the `rxrpc_call` object in this scenario? This is especially suspect because the `recvmmsg` is occurring on an AF_RXRPC socket.
I'm starting to think this is not a coincidence and we need to look closely at `rxrpc_recvmsg` to see what's happening during a `recvmmsg` operation.
|
| 10/2 |
2026/03/10 00:50 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/recvmsg.c IncludeLines:true Name:rxrpc_recvmsg]
Results: map[SourceCode: 361: /*
362: * Receive a message from an RxRPC socket
363: * - we need to be careful about two or more threads calling recvmsg
364: * simultaneously
365: */
366: int rxrpc_recvmsg(struct socket *sock, struct msghdr *msg, size_t len,
367: int flags)
368: {
369: struct rxrpc_call *call;
370: struct rxrpc_sock *rx = rxrpc_sk(sock->sk);
371: struct list_head *l;
372: unsigned int call_debug_id = 0;
373: size_t copied = 0;
374: long timeo;
375: int ret;
376:
377: DEFINE_WAIT(wait);
378:
379: trace_rxrpc_recvmsg(0, rxrpc_recvmsg_enter, 0);
380:
381: if (flags & (MSG_OOB | MSG_TRUNC))
382: return -EOPNOTSUPP;
383:
384: timeo = sock_rcvtimeo(&rx->sk, flags & MSG_DONTWAIT);
385:
386: try_again:
387: lock_sock(&rx->sk);
388:
389: /* Return immediately if a client socket has no outstanding calls */
390: if (RB_EMPTY_ROOT(&rx->calls) &&
391: list_empty(&rx->recvmsg_q) &&
392: skb_queue_empty_lockless(&rx->recvmsg_oobq) &&
393: rx->sk.sk_state != RXRPC_SERVER_LISTENING) {
394: release_sock(&rx->sk);
395: return -EAGAIN;
396: }
397:
398: if (list_empty(&rx->recvmsg_q)) {
399: ret = -EWOULDBLOCK;
400: if (timeo == 0) {
401: call = NULL;
402: goto error_no_call;
403: }
404:
405: release_sock(&rx->sk);
406:
407: /* Wait for something to happen */
408: prepare_to_wait_exclusive(sk_sleep(&rx->sk), &wait,
409: TASK_INTERRUPTIBLE);
410: ret = sock_error(&rx->sk);
411: if (ret)
412: goto wait_error;
413:
414: if (list_empty(&rx->recvmsg_q) &&
415: skb_queue_empty_lockless(&rx->recvmsg_oobq)) {
416: if (signal_pending(current))
417: goto wait_interrupted;
418: trace_rxrpc_recvmsg(0, rxrpc_recvmsg_wait, 0);
419: timeo = schedule_timeout(timeo);
420: }
421: finish_wait(sk_sleep(&rx->sk), &wait);
422: goto try_again;
423: }
424:
425: /* Deal with OOB messages before we consider getting normal data. */
426: if (!skb_queue_empty_lockless(&rx->recvmsg_oobq)) {
427: ret = rxrpc_recvmsg_oob(sock, msg, flags);
428: release_sock(&rx->sk);
429: if (ret == -EAGAIN)
430: goto try_again;
431: goto error_no_call;
432: }
433:
434: /* Find the next call and dequeue it if we're not just peeking. If we
435: * do dequeue it, that comes with a ref that we will need to release.
436: * We also want to weed out calls that got requeued whilst we were
437: * shovelling data out.
438: */
439: spin_lock_irq(&rx->recvmsg_lock);
440: l = rx->recvmsg_q.next;
441: call = list_entry(l, struct rxrpc_call, recvmsg_link);
442:
443: if (!rxrpc_call_is_complete(call) &&
444: skb_queue_empty(&call->recvmsg_queue) &&
445: skb_queue_empty(&rx->recvmsg_oobq)) {
446: list_del_init(&call->recvmsg_link);
447: spin_unlock_irq(&rx->recvmsg_lock);
448: release_sock(&rx->sk);
449: trace_rxrpc_recvmsg(call->debug_id, rxrpc_recvmsg_unqueue, 0);
450: rxrpc_put_call(call, rxrpc_call_put_recvmsg);
451: goto try_again;
452: }
453:
454: rxrpc_see_call(call, rxrpc_call_see_recvmsg);
455: if (test_bit(RXRPC_CALL_RELEASED, &call->flags)) {
456: rxrpc_see_call(call, rxrpc_call_see_already_released);
457: list_del_init(&call->recvmsg_link);
458: spin_unlock_irq(&rx->recvmsg_lock);
459: release_sock(&rx->sk);
460: trace_rxrpc_recvmsg(call->debug_id, rxrpc_recvmsg_unqueue, 0);
461: rxrpc_put_call(call, rxrpc_call_put_recvmsg);
462: goto try_again;
463: }
464: if (!(flags & MSG_PEEK))
465: list_del_init(&call->recvmsg_link);
466: else
467: rxrpc_get_call(call, rxrpc_call_get_recvmsg);
468: spin_unlock_irq(&rx->recvmsg_lock);
469:
470: call_debug_id = call->debug_id;
471: trace_rxrpc_recvmsg(call_debug_id, rxrpc_recvmsg_dequeue, 0);
472:
473: /* We're going to drop the socket lock, so we need to lock the call
474: * against interference by sendmsg.
475: */
476: if (!mutex_trylock(&call->user_mutex)) {
477: ret = -EWOULDBLOCK;
478: if (flags & MSG_DONTWAIT)
479: goto error_requeue_call;
480: ret = -ERESTARTSYS;
481: if (mutex_lock_interruptible(&call->user_mutex) < 0)
482: goto error_requeue_call;
483: }
484:
485: release_sock(&rx->sk);
486:
487: if (test_bit(RXRPC_CALL_RELEASED, &call->flags)) {
488: rxrpc_see_call(call, rxrpc_call_see_already_released);
489: mutex_unlock(&call->user_mutex);
490: if (!(flags & MSG_PEEK))
491: rxrpc_put_call(call, rxrpc_call_put_recvmsg);
492: goto try_again;
493: }
494:
495: ret = rxrpc_recvmsg_user_id(call, msg, flags);
496: if (ret < 0)
497: goto error_unlock_call;
498:
499: if (msg->msg_name && call->peer) {
500: size_t len = sizeof(call->dest_srx);
501:
502: memcpy(msg->msg_name, &call->dest_srx, len);
503: msg->msg_namelen = len;
504: }
505:
506: ret = rxrpc_recvmsg_data(sock, call, msg, &msg->msg_iter, len,
507: flags, &copied);
508: if (ret == -EAGAIN)
509: ret = 0;
510: if (ret == -EIO)
511: goto call_failed;
512: if (ret < 0)
513: goto error_unlock_call;
514:
515: if (rxrpc_call_is_complete(call) &&
516: skb_queue_empty(&call->recvmsg_queue))
517: goto call_complete;
518: if (rxrpc_call_has_failed(call))
519: goto call_failed;
520:
521: if (!(flags & MSG_PEEK) &&
522: !skb_queue_empty(&call->recvmsg_queue))
523: rxrpc_notify_socket(call);
524: goto not_yet_complete;
525:
526: call_failed:
527: rxrpc_purge_queue(&call->recvmsg_queue);
528: call_complete:
529: ret = rxrpc_recvmsg_term(call, msg);
530: if (ret < 0)
531: goto error_unlock_call;
532: if (!(flags & MSG_PEEK))
533: rxrpc_release_call(rx, call);
534: msg->msg_flags |= MSG_EOR;
535: ret = 1;
536:
537: not_yet_complete:
538: if (ret == 0)
539: msg->msg_flags |= MSG_MORE;
540: else
541: msg->msg_flags &= ~MSG_MORE;
542: ret = copied;
543:
544: error_unlock_call:
545: mutex_unlock(&call->user_mutex);
546: rxrpc_put_call(call, rxrpc_call_put_recvmsg);
547: trace_rxrpc_recvmsg(call_debug_id, rxrpc_recvmsg_return, ret);
548: return ret;
549:
550: error_requeue_call:
551: if (!(flags & MSG_PEEK)) {
552: spin_lock_irq(&rx->recvmsg_lock);
553: if (list_empty(&call->recvmsg_link)) {
554: list_add(&call->recvmsg_link, &rx->recvmsg_q);
555: rxrpc_see_call(call, rxrpc_call_see_recvmsg_requeue);
556: spin_unlock_irq(&rx->recvmsg_lock);
557: } else if (list_is_first(&call->recvmsg_link, &rx->recvmsg_q)) {
558: spin_unlock_irq(&rx->recvmsg_lock);
559: rxrpc_put_call(call, rxrpc_call_see_recvmsg_requeue_first);
560: } else {
561: list_move(&call->recvmsg_link, &rx->recvmsg_q);
562: spin_unlock_irq(&rx->recvmsg_lock);
563: rxrpc_put_call(call, rxrpc_call_see_recvmsg_requeue_move);
564: }
565: trace_rxrpc_recvmsg(call_debug_id, rxrpc_recvmsg_requeue, 0);
566: } else {
567: rxrpc_put_call(call, rxrpc_call_put_recvmsg_peek_nowait);
568: }
569: error_no_call:
570: release_sock(&rx->sk);
571: error_trace:
572: trace_rxrpc_recvmsg(call_debug_id, rxrpc_recvmsg_return, ret);
573: return ret;
574:
575: wait_interrupted:
576: ret = sock_intr_errno(timeo);
577: wait_error:
578: finish_wait(sk_sleep(&rx->sk), &wait);
579: call = NULL;
580: goto error_trace;
581: }
SourceFile:net/rxrpc/recvmsg.c]
|
| 11/2 |
2026/03/10 00:50 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 23064
output: 45
thoughts: 265
Thoughts: **My Analysis of a Potential RXRPC Race Condition**
Okay, so I'm looking at this `rxrpc_recvmsg` function, and my initial thought is, "Wait a minute..." It seems that when `MSG_PEEK` is set, `rxrpc_recvmsg` *gets* the `call` using `rxrpc_get_call`, but when it isn't set, it removes the `call` from the `recvmsg_q`, which also implies it has a reference to the call. Then, it dives into `rxrpc_recvmsg_data` which is responsible for copying the data into user space.
Now, here's where it gets interesting. If `rxrpc_recvmsg_data` blocks during that `copy_to_user` operation—say, due to a userfaultfd—it's still holding onto that `call` reference.
The critical issue is this: If the network namespace is in the process of being torn down while `rxrpc_recvmsg_data` is blocked, `rxrpc_destroy_all_calls` will be called, and it'll wait for the `nr_calls` count to reach zero. The problem is that `nr_calls` won't be decremented until the `recvmsg` operation completes, which happens after the `copy_to_user` call completes (and `rxrpc_put_call` is invoked).
This seems like a deadlock waiting to happen. `rxrpc_destroy_all_calls` waits for the `recvmsg` call to finish, but the `recvmsg` call can't finish until the destroy operation is finished.
The comment in `rxrpc_destroy_all_calls` states that "any open UDP sockets in that namespace must have been closed, so any outstanding calls cannot be doing I/O." I get that, but what about AF_RXRPC sockets? Are they closed *before* `rxrpc_destroy_all_calls` is invoked? This seems like the key question. I need to track down how `rxrpc_exit_net` handles AF_RXRPC sockets to be sure. This might be a subtle but critical race condition. Let me check `rxrpc_exit_net` to confirm.
|
| 12/2 |
2026/03/10 00:50 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/net_ns.c IncludeLines:true Name:rxrpc_exit_net]
Results: map[SourceCode: 100: /*
101: * Clean up a per-network namespace record.
102: */
103: static __net_exit void rxrpc_exit_net(struct net *net)
104: {
105: struct rxrpc_net *rxnet = rxrpc_net(net);
106:
107: rxnet->live = false;
108: timer_delete_sync(&rxnet->peer_keepalive_timer);
109: cancel_work_sync(&rxnet->peer_keepalive_work);
110: /* Remove the timer again as the worker may have restarted it. */
111: timer_delete_sync(&rxnet->peer_keepalive_timer);
112: rxrpc_destroy_all_calls(rxnet);
113: rxrpc_destroy_all_connections(rxnet);
114: rxrpc_destroy_all_peers(rxnet);
115: rxrpc_destroy_all_locals(rxnet);
116: proc_remove(rxnet->proc_net);
117: }
SourceFile:net/rxrpc/net_ns.c]
|
| 13/2 |
2026/03/10 00:50 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 24190
output: 44
thoughts: 222
Thoughts: **My Line of Reasoning Regarding Network Namespace Destruction and AF_RXRPC Sockets**
Okay, so I'm focusing on `rxrpc_exit_net` here. The key is that it's triggered during the destruction of a network namespace. My immediate thought is, "Okay, when a namespace goes away, are all the sockets within it automatically shut down?" That's a reasonable assumption, but then I remember, no, sockets are closed when their file descriptors are closed.
But then there's this connection – if a socket *is* open, it holds a reference, a crucial reference, to the network namespace via `sk->sk_net`. This is the crux. Which means, the network namespace can't be destroyed until all of its associated sockets are closed! This is fundamental.
I have to be sure about how that socket reference happens. I realize `sk->sk_net` is the pointer in the socket, but how does that pointer get set? Ah, yes, `sk_alloc` uses `sock_net_set(sk, net)`. This means `sk_alloc` takes a reference on the network namespace when it calls `sock_net_set`. I need to be sure that it does; so `get_net(net)`.
So, in the case of an open AF_RXRPC socket, I conclude that the network namespace can't be destroyed as long as that socket persists.
But now, what's a corner case? Could it be in a different network namespace? Could a socket be in one namespace, but the `rxnet` is from another? Let me check `rxrpc_create`. I think that's where I need to look.
|
| 14/2 |
2026/03/10 00:50 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/af_rxrpc.c IncludeLines:true Name:rxrpc_create]
Results: map[SourceCode: 801: /*
802: * create an RxRPC socket
803: */
804: static int rxrpc_create(struct net *net, struct socket *sock, int protocol,
805: int kern)
806: {
807: struct rxrpc_net *rxnet;
808: struct rxrpc_sock *rx;
809: struct sock *sk;
810:
811: _enter("%p,%d", sock, protocol);
812:
813: /* we support transport protocol UDP/UDP6 only */
814: if (protocol != PF_INET &&
815: IS_ENABLED(CONFIG_AF_RXRPC_IPV6) && protocol != PF_INET6)
816: return -EPROTONOSUPPORT;
817:
818: if (sock->type != SOCK_DGRAM)
819: return -ESOCKTNOSUPPORT;
820:
821: sock->ops = &rxrpc_rpc_ops;
822: sock->state = SS_UNCONNECTED;
823:
824: sk = sk_alloc(net, PF_RXRPC, GFP_KERNEL, &rxrpc_proto, kern);
825: if (!sk)
826: return -ENOMEM;
827:
828: sock_init_data(sock, sk);
829: sock_set_flag(sk, SOCK_RCU_FREE);
830: sk->sk_state = RXRPC_UNBOUND;
831: sk->sk_write_space = rxrpc_write_space;
832: sk->sk_max_ack_backlog = 0;
833: sk->sk_destruct = rxrpc_sock_destructor;
834:
835: rx = rxrpc_sk(sk);
836: rx->family = protocol;
837: rx->calls = RB_ROOT;
838:
839: spin_lock_init(&rx->incoming_lock);
840: skb_queue_head_init(&rx->recvmsg_oobq);
841: rx->pending_oobq = RB_ROOT;
842: INIT_LIST_HEAD(&rx->sock_calls);
843: INIT_LIST_HEAD(&rx->to_be_accepted);
844: INIT_LIST_HEAD(&rx->recvmsg_q);
845: spin_lock_init(&rx->recvmsg_lock);
846: rwlock_init(&rx->call_lock);
847: memset(&rx->srx, 0, sizeof(rx->srx));
848:
849: rxnet = rxrpc_net(sock_net(&rx->sk));
850: timer_reduce(&rxnet->peer_keepalive_timer, jiffies + 1);
851:
852: _leave(" = 0 [%p]", rx);
853: return 0;
854: }
SourceFile:net/rxrpc/af_rxrpc.c]
|
| 15/2 |
2026/03/10 00:50 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 25706
output: 44
thoughts: 203
Thoughts: **My Line of Reasoning Regarding the RxRPC Socket Destruction**
Okay, let's break this down. `rxrpc_create` is responsible for setting up a socket within the `net` namespace. If that socket is active, the `net` instance itself can't be destroyed – that's fundamental. Now, I'm seeing that `rxrpc_destroy_all_calls` gets triggered when the `net` is destroyed. The implication here is that the destruction of `net` signifies that *all* sockets within that namespace have been closed.
But wait, if all those sockets are closed, how can `rxrpc_recvmsg` *possibly* still be running? `rxrpc_recvmsg` is, of course, called on a specific socket. If that socket's closed, the underlying `recvmsg` call, by definition, shouldn't be happening. Ah, right. `recvmsg` is ultimately called on a file descriptor! Okay, when a file descriptor is closed, `fput` is called, which then triggers `sock_release`. `sock_release` then calls `sock->ops->release`, which translates to `rxrpc_release` in our case.
But here's the kicker: `rxrpc_release` doesn't appear to wait for a potentially pending `recvmsg` call to complete! This seems like a critical oversight. I need to dive into the implementation of `rxrpc_release` to confirm my suspicion and see if there's a race condition there.
|
| 16/2 |
2026/03/10 00:50 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/af_rxrpc.c IncludeLines:true Name:rxrpc_release]
Results: map[SourceCode: 982: /*
983: * release an RxRPC BSD socket on close() or equivalent
984: */
985: static int rxrpc_release(struct socket *sock)
986: {
987: struct sock *sk = sock->sk;
988:
989: _enter("%p{%p}", sock, sk);
990:
991: if (!sk)
992: return 0;
993:
994: sock->sk = NULL;
995:
996: return rxrpc_release_sock(sk);
997: }
SourceFile:net/rxrpc/af_rxrpc.c]
|
| 17/2 |
2026/03/10 00:50 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 26492
output: 45
thoughts: 210
Thoughts: **My Line of Reasoning**
Okay, here's what's running through my mind. I'm focusing on a potential resource leak, particularly around file descriptors and network namespaces. I'm starting with `fput` and its role in releasing a file descriptor. I know `fput` is triggered when the last reference to a file descriptor is dropped. Now, I'm thinking about `recvmsg`. I know that if `recvmsg` is actively running, it's holding a reference to the file descriptor, right? And `fdget` would have been used to increment this reference count. So, `fput` won't execute until `recvmsg` returns.
Therefore, if `fput` *doesn't* get called, then `rxrpc_release` won't be called, and consequently, the socket won't be closed. If the socket remains open, the network namespace might be prevented from being destroyed.
But, a nagging doubt creeps in: is this really accurate? I need to consider a potential edge case. What if the socket is in a different network namespace? Can a socket's creation and its `sk_net` attribute reside in different namespaces? No, that doesn't seem right. `sk_alloc` uses the `net` struct.
Now, I'm thinking about `rxrpc_local`. This function creates a UDP socket and the thought has led me to re-examine the socket creation process. I should probably re-examine `rxrpc_open_socket` to be sure.
|
| 18/2 |
2026/03/10 00:50 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/local_object.c IncludeLines:true Name:rxrpc_open_socket]
Results: map[SourceCode: 158: /*
159: * create the local socket
160: * - must be called with rxrpc_local_mutex locked
161: */
162: static int rxrpc_open_socket(struct rxrpc_local *local, struct net *net)
163: {
164: struct udp_tunnel_sock_cfg tuncfg = {NULL};
165: struct sockaddr_rxrpc *srx = &local->srx;
166: struct udp_port_cfg udp_conf = {0};
167: struct task_struct *io_thread;
168: struct sock *usk;
169: int ret;
170:
171: _enter("%p{%d,%d}",
172: local, srx->transport_type, srx->transport.family);
173:
174: udp_conf.family = srx->transport.family;
175: udp_conf.use_udp_checksums = true;
176: if (udp_conf.family == AF_INET) {
177: udp_conf.local_ip = srx->transport.sin.sin_addr;
178: udp_conf.local_udp_port = srx->transport.sin.sin_port;
179: #if IS_ENABLED(CONFIG_AF_RXRPC_IPV6)
180: } else {
181: udp_conf.local_ip6 = srx->transport.sin6.sin6_addr;
182: udp_conf.local_udp_port = srx->transport.sin6.sin6_port;
183: udp_conf.use_udp6_tx_checksums = true;
184: udp_conf.use_udp6_rx_checksums = true;
185: #endif
186: }
187: ret = udp_sock_create(net, &udp_conf, &local->socket);
188: if (ret < 0) {
189: _leave(" = %d [socket]", ret);
190: return ret;
191: }
192:
193: tuncfg.encap_type = UDP_ENCAP_RXRPC;
194: tuncfg.encap_rcv = rxrpc_encap_rcv;
195: tuncfg.encap_err_rcv = rxrpc_encap_err_rcv;
196: tuncfg.sk_user_data = local;
197: setup_udp_tunnel_sock(net, local->socket, &tuncfg);
198:
199: /* set the socket up */
200: usk = local->socket->sk;
201: usk->sk_error_report = rxrpc_error_report;
202:
203: switch (srx->transport.family) {
204: case AF_INET6:
205: /* we want to receive ICMPv6 errors */
206: ip6_sock_set_recverr(usk);
207:
208: /* Fall through and set IPv4 options too otherwise we don't get
209: * errors from IPv4 packets sent through the IPv6 socket.
210: */
211: fallthrough;
212: case AF_INET:
213: /* we want to receive ICMP errors */
214: ip_sock_set_recverr(usk);
215:
216: /* we want to set the don't fragment bit */
217: rxrpc_local_dont_fragment(local, true);
218: break;
219:
220: default:
221: BUG();
222: }
223:
224: io_thread = kthread_run(rxrpc_io_thread, local,
225: "krxrpcio/%u", ntohs(udp_conf.local_udp_port));
226: if (IS_ERR(io_thread)) {
227: ret = PTR_ERR(io_thread);
228: goto error_sock;
229: }
230:
231: wait_for_completion(&local->io_thread_ready);
232: WRITE_ONCE(local->io_thread, io_thread);
233: _leave(" = 0");
234: return 0;
235:
236: error_sock:
237: kernel_sock_shutdown(local->socket, SHUT_RDWR);
238: local->socket->sk->sk_user_data = NULL;
239: sock_release(local->socket);
240: local->socket = NULL;
241: return ret;
242: }
SourceFile:net/rxrpc/local_object.c]
|
| 19/2 |
2026/03/10 00:50 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 28461
output: 46
thoughts: 224
Thoughts: **Investigating the Network Namespace Dependency in RxRPC Sockets**
Okay, so I'm tracing the network namespace usage with these RxRPC sockets. My initial thought is: `rxrpc_open_socket` is responsible for creating a UDP socket, right? And that's happening within the `net` context. But the `net` struct is passed in from `rxrpc_lookup_local`. So, where exactly is `rxrpc_lookup_local` getting this `net` structure? Ah, it's retrieved using `rxnet = rxrpc_net(sock_net(&rx->sk))`. That means the UDP socket is going to be bound to the same network namespace as the main AF_RXRPC socket.
This implies that as long as the AF_RXRPC socket is open, the corresponding network namespace is protected from deletion, which makes sense. But, here's a wrinkle. What happens if I close the AF_RXRPC socket, but somehow keep `rxrpc_local` alive? That would present a problem.
Now, how could I keep `rxrpc_local` around after the socket is closed? It looks like `rxrpc_local` is retained by either `rxrpc_call` or `rxrpc_peer`. Let's say `rxrpc_call` is kept alive by a lingering `recvmsg`. Then, `rxrpc_local` would be kept alive as well, even if the socket has been closed. However, if the AF_RXRPC socket is closed, `rxrpc_release_sock` is called. So, I need to check `rxrpc_release_sock` to understand how the system prevents any lingering issues from this scenario.
|
| 20/2 |
2026/03/10 00:51 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/af_rxrpc.c IncludeLines:true Name:rxrpc_release_sock]
Results: map[SourceCode: 926: /*
927: * release an RxRPC socket
928: */
929: static int rxrpc_release_sock(struct sock *sk)
930: {
931: struct rxrpc_sock *rx = rxrpc_sk(sk);
932:
933: _enter("%p{%d,%d}", sk, sk->sk_state, refcount_read(&sk->sk_refcnt));
934:
935: /* declare the socket closed for business */
936: sock_orphan(sk);
937: sk->sk_shutdown = SHUTDOWN_MASK;
938:
939: /* We want to kill off all connections from a service socket
940: * as fast as possible because we can't share these; client
941: * sockets, on the other hand, can share an endpoint.
942: */
943: switch (sk->sk_state) {
944: case RXRPC_SERVER_BOUND:
945: case RXRPC_SERVER_BOUND2:
946: case RXRPC_SERVER_LISTENING:
947: case RXRPC_SERVER_LISTEN_DISABLED:
948: rx->local->service_closed = true;
949: break;
950: }
951:
952: spin_lock_irq(&rx->recvmsg_lock);
953: sk->sk_state = RXRPC_CLOSE;
954: spin_unlock_irq(&rx->recvmsg_lock);
955:
956: if (rx->local && rx->local->service == rx) {
957: write_lock(&rx->local->services_lock);
958: rx->local->service = NULL;
959: write_unlock(&rx->local->services_lock);
960: }
961:
962: /* try to flush out this socket */
963: rxrpc_discard_prealloc(rx);
964: rxrpc_release_calls_on_socket(rx);
965: flush_workqueue(rxrpc_workqueue);
966: rxrpc_purge_oob_queue(sk);
967: rxrpc_purge_queue(&sk->sk_receive_queue);
968:
969: rxrpc_unuse_local(rx->local, rxrpc_local_unuse_release_sock);
970: rxrpc_put_local(rx->local, rxrpc_local_put_release_sock);
971: rx->local = NULL;
972: key_put(rx->key);
973: rx->key = NULL;
974: key_put(rx->securities);
975: rx->securities = NULL;
976: sock_put(sk);
977:
978: _leave(" = 0");
979: return 0;
980: }
SourceFile:net/rxrpc/af_rxrpc.c]
|
| 21/2 |
2026/03/10 00:51 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 29960
output: 49
thoughts: 43
Thoughts: Here's my thought process, as requested:
**Deconstructing Socket Release in RxRPC**
Okay, so I'm looking at `rxrpc_release_sock`, and the critical thing here is that it immediately calls `rxrpc_release_calls_on_socket(rx)`. My first instinct is to dive into the implementation of that function. That's where the real action is, undoubtedly. I need to understand what "releasing calls" actually entails within the context of the socket release. Is this cleaning up data structures? Signaling something? Closing connections? I need to understand this to determine if there are any race conditions, resources that aren't properly freed, or potential leakage issues within this release process. Let's get right to it and check that function!
|
| 22/2 |
2026/03/10 00:51 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_object.c IncludeLines:true Name:rxrpc_release_calls_on_socket]
Results: map[SourceCode: 603: /*
604: * release all the calls associated with a socket
605: */
606: void rxrpc_release_calls_on_socket(struct rxrpc_sock *rx)
607: {
608: struct rxrpc_call *call;
609:
610: _enter("%p", rx);
611:
612: while (!list_empty(&rx->to_be_accepted)) {
613: call = list_entry(rx->to_be_accepted.next,
614: struct rxrpc_call, accept_link);
615: list_del(&call->accept_link);
616: rxrpc_propose_abort(call, RX_CALL_DEAD, -ECONNRESET,
617: rxrpc_abort_call_sock_release_tba);
618: rxrpc_put_call(call, rxrpc_call_put_release_sock_tba);
619: }
620:
621: while (!list_empty(&rx->sock_calls)) {
622: call = list_entry(rx->sock_calls.next,
623: struct rxrpc_call, sock_link);
624: rxrpc_get_call(call, rxrpc_call_get_release_sock);
625: rxrpc_propose_abort(call, RX_CALL_DEAD, -ECONNRESET,
626: rxrpc_abort_call_sock_release);
627: rxrpc_release_call(rx, call);
628: rxrpc_put_call(call, rxrpc_call_put_release_sock);
629: }
630:
631: while ((call = list_first_entry_or_null(&rx->recvmsg_q,
632: struct rxrpc_call, recvmsg_link))) {
633: list_del_init(&call->recvmsg_link);
634: rxrpc_put_call(call, rxrpc_call_put_release_recvmsg_q);
635: }
636:
637: _leave("");
638: }
SourceFile:net/rxrpc/call_object.c]
|
| 23/2 |
2026/03/10 00:51 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 30856
output: 37
thoughts: 282
Thoughts: **My Current Thinking on the RXRPC Call Release Issue**
Okay, let's break this down. My initial thought was that `rxrpc_release_calls_on_socket` was the culprit, removing calls from `recvmsg_q` and calling `rxrpc_put_call`. But that can't be right because if `recvmsg` is running, it should have already pulled the call off the `recvmsg_q`. Right? My understanding is that `recvmsg` handles that removal and maintains a reference to the call.
Then the execution flow gets interesting. `recvmsg` removes the call and blocks within `rxrpc_recvmsg_data` doing a `copy_to_user`. This is where things could go sideways. If the socket closes while that `copy_to_user` is happening, `rxrpc_release_sock` gets called, right? But `rxrpc_release_sock` is called from `rxrpc_release`, which is called from `sock_release`, which in turn is called from `fput`. The critical question is whether `fput` can be called while `recvmsg` is still running, and I don't think it can. `recvmsg` holds a reference to the file descriptor, so `fput` should be blocked.
So, `rxrpc_release_sock` isn't called while `recvmsg` is blocking, which seems correct. So how can we get into a state where the network namespace is destroyed while a `recvmsg` call is in progress? If the namespace is gone, all references should be gone. But the file descriptor *should* hold a reference to the socket, which *should* hold a reference to the network namespace. That reference should prevent destruction... unless I'm missing something fundamental.
Now, wait a second... does the file descriptor *really* hold a reference to the network namespace? I feel like I'm missing something here, it's not immediately obvious. I need to refresh my memory on the socket allocation. I'm going to have to check `sock_alloc` and look at reference counting here. I need a deeper dive.
|
| 24/2 |
2026/03/10 00:51 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/socket.c IncludeLines:true Name:sock_alloc]
Results: map[SourceCode: 632: struct socket *sock_alloc(void)
633: {
634: struct inode *inode;
635: struct socket *sock;
636:
637: inode = new_inode_pseudo(sock_mnt->mnt_sb);
638: if (!inode)
639: return NULL;
640:
641: sock = SOCKET_I(inode);
642:
643: inode->i_ino = get_next_ino();
644: inode->i_mode = S_IFSOCK | S_IRWXUGO;
645: inode->i_uid = current_fsuid();
646: inode->i_gid = current_fsgid();
647: inode->i_op = &sockfs_inode_ops;
648:
649: return sock;
650: }
SourceFile:net/socket.c]
|
| 25/2 |
2026/03/10 00:51 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 31915
output: 39
thoughts: 32
Thoughts: Okay, here's my interpretation of those thoughts, presented as a personal thought process, as requested:
**Network Namespace Conundrum**
Hmm, hold on a second. I'm reviewing this code, specifically around the socket allocation. I was under the impression that the function `sock_alloc` would require a reference to the network namespace. That would make perfect sense, given that sockets are inherently tied to a network context. But, no, wait... that's not what I'm seeing here. It seems like `sock_alloc` doesn't actually take a *reference* to the network namespace directly. Interesting. So, if `sock_alloc` *doesn't* take it, that means the namespace has to be passed elsewhere, somehow. I wonder if there is some other function or structure that holds the association... Ah! Of course. `sk_alloc` does! That's the key. That's where the network namespace reference *does* live. It has to. Otherwise, the socket can't exist within the desired network context. Right, so I was initially mistaken, but I see where I got confused now.
|
| 26/2 |
2026/03/10 00:51 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/core/sock.c IncludeLines:true Name:sk_alloc]
Results: map[SourceCode:2288: /**
2289: * sk_alloc - All socket objects are allocated here
2290: * @net: the applicable net namespace
2291: * @family: protocol family
2292: * @priority: for allocation (%GFP_KERNEL, %GFP_ATOMIC, etc)
2293: * @prot: struct proto associated with this new sock instance
2294: * @kern: is this to be a kernel socket?
2295: */
2296: struct sock *sk_alloc(struct net *net, int family, gfp_t priority,
2297: struct proto *prot, int kern)
2298: {
2299: struct sock *sk;
2300:
2301: sk = sk_prot_alloc(prot, priority | __GFP_ZERO, family);
2302: if (sk) {
2303: sk->sk_family = family;
2304: /*
2305: * See comment in struct sock definition to understand
2306: * why we need sk_prot_creator -acme
2307: */
2308: sk->sk_prot = sk->sk_prot_creator = prot;
2309:
2310: if (READ_ONCE(net->core.sysctl_bypass_prot_mem))
2311: sk->sk_bypass_prot_mem = 1;
2312:
2313: sk->sk_kern_sock = kern;
2314: sock_lock_init(sk);
2315:
2316: sk->sk_net_refcnt = kern ? 0 : 1;
2317: if (likely(sk->sk_net_refcnt)) {
2318: get_net_track(net, &sk->ns_tracker, priority);
2319: sock_inuse_add(net, 1);
2320: } else {
2321: net_passive_inc(net);
2322: __netns_tracker_alloc(net, &sk->ns_tracker,
2323: false, priority);
2324: }
2325:
2326: sock_net_set(sk, net);
2327: refcount_set(&sk->sk_wmem_alloc, SK_WMEM_ALLOC_BIAS);
2328:
2329: mem_cgroup_sk_alloc(sk);
2330: cgroup_sk_alloc(&sk->sk_cgrp_data);
2331: sock_update_classid(&sk->sk_cgrp_data);
2332: sock_update_netprioidx(&sk->sk_cgrp_data);
2333: sk_tx_queue_clear(sk);
2334: }
2335:
2336: return sk;
2337: }
SourceFile:net/core/sock.c]
|
| 27/2 |
2026/03/10 00:51 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 33045
output: 45
thoughts: 357
Thoughts: **Debugging a Network Namespace Reference Count Issue**
Okay, let me think this through. `sk_alloc` is tripping me up. If `kern` is zero, which seems to be the case for user-space sockets, it takes a reference to the network namespace. But if `kern` is one, it doesn't. Now, the `AF_RXRPC` socket, is that a kernel socket? No, it's created in user space with the standard `socket()` call, so `kern` is definitely zero. Therefore, `sk_alloc` *does* take a reference to the network namespace.
Now, if the network namespace is being destroyed, it means its reference count has dropped to zero, right? That should mean all sockets with `sk_net_refcnt == 1` must have already been freed. But wait, if the socket is freed, how can `recvmsg` even be *running* on it? That doesn't make any sense. So, `recvmsg` *is* running, which means the socket is NOT freed. That’s the key.
So, if the socket isn’t freed, how can the network namespace be destroyed? Aha! Maybe the socket is in a *different* network namespace than the one being destroyed. Okay, let's see. Thinking about `rxrpc_create`, it uses `sock_net(&rx->sk)` to get the network namespace, `rxnet`. So, if the socket lives in `net1`, then `rxnet` is `rxrpc_net(net1)`. Then `net1` *can't* be destroyed until the socket is freed.
But wait, what if the socket is created in `net1`, but it's using a `rxrpc_local` from a DIFFERENT network namespace, say `net2`? Is that possible? That's the core question now. I need to look at `rxrpc_lookup_local` to see if that's even a valid scenario. I need to check the code there to see if this is a possibility.
|
| 28/2 |
2026/03/10 00:51 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/local_object.c IncludeLines:true Name:rxrpc_lookup_local]
Results: map[SourceCode: 244: /*
245: * Look up or create a new local endpoint using the specified local address.
246: */
247: struct rxrpc_local *rxrpc_lookup_local(struct net *net,
248: const struct sockaddr_rxrpc *srx)
249: {
250: struct rxrpc_local *local;
251: struct rxrpc_net *rxnet = rxrpc_net(net);
252: struct hlist_node *cursor;
253: long diff;
254: int ret;
255:
256: _enter("{%d,%d,%pISp}",
257: srx->transport_type, srx->transport.family, &srx->transport);
258:
259: mutex_lock(&rxnet->local_mutex);
260:
261: hlist_for_each(cursor, &rxnet->local_endpoints) {
262: local = hlist_entry(cursor, struct rxrpc_local, link);
263:
264: diff = rxrpc_local_cmp_key(local, srx);
265: if (diff != 0)
266: continue;
267:
268: /* Services aren't allowed to share transport sockets, so
269: * reject that here. It is possible that the object is dying -
270: * but it may also still have the local transport address that
271: * we want bound.
272: */
273: if (srx->srx_service) {
274: local = NULL;
275: goto addr_in_use;
276: }
277:
278: /* Found a match. We want to replace a dying object.
279: * Attempting to bind the transport socket may still fail if
280: * we're attempting to use a local address that the dying
281: * object is still using.
282: */
283: if (!rxrpc_use_local(local, rxrpc_local_use_lookup))
284: break;
285:
286: goto found;
287: }
288:
289: local = rxrpc_alloc_local(net, srx);
290: if (!local)
291: goto nomem;
292:
293: ret = rxrpc_open_socket(local, net);
294: if (ret < 0)
295: goto sock_error;
296:
297: if (cursor) {
298: hlist_replace_rcu(cursor, &local->link);
299: cursor->pprev = NULL;
300: } else {
301: hlist_add_head_rcu(&local->link, &rxnet->local_endpoints);
302: }
303:
304: found:
305: mutex_unlock(&rxnet->local_mutex);
306: _leave(" = %p", local);
307: return local;
308:
309: nomem:
310: ret = -ENOMEM;
311: sock_error:
312: mutex_unlock(&rxnet->local_mutex);
313: if (local)
314: call_rcu(&local->rcu, rxrpc_local_rcu);
315: _leave(" = %d", ret);
316: return ERR_PTR(ret);
317:
318: addr_in_use:
319: mutex_unlock(&rxnet->local_mutex);
320: _leave(" = -EADDRINUSE");
321: return ERR_PTR(-EADDRINUSE);
322: }
SourceFile:net/rxrpc/local_object.c]
|
| 29/2 |
2026/03/10 00:51 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 35030
output: 53
thoughts: 307
Thoughts: **Network Namespace Conundrum**
Okay, let's think this through. `rxrpc_lookup_local` relies on `net`, which originates from either `rxrpc_bind` or `rxrpc_connect`. That `net` value is derived from `sock_net(&rx->sk)`. Right, so the local address lookup is *always* tied to the network namespace of the AF_RXRPC socket itself. That means, as long as the AF_RXRPC socket is open, the network namespace *shouldn't* be destructible.
So, how can `rxrpc_destroy_all_calls` ever be invoked if a `recvmsg` call is active? That feels fundamentally wrong. Oh, wait, is the socket *actually* an AF_RXRPC socket within `recvmsg`? Yes, of course, `rxrpc_recvmsg` is the function being called.
But hold on! `rxrpc_recvmsg` operates on a `struct socket *sock`. What if that socket's file descriptor has been passed to another process? Or what if the socket itself was created by a kernel thread? Hmm... What if it was created by `kernel_accept` or `kernel_bind`? No, RxRPC doesn't support `accept`, I remember that.
Wait, could the socket *be* an AF_RXRPC socket that's fundamentally a kernel socket? If it's a kernel socket, then `sk_alloc` gets called with `kern=1`. And that means `sk_net_refcnt` is initialized to 0! Who even *creates* kernel AF_RXRPC sockets? I'd better check `sock_create_kern`. That's the next step.
|
| 30/2 |
2026/03/10 00:51 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/rxrpc/af_rxrpc.c IncludeSnippetLines:1 Name:sock_create_kern SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:afs_open_socket SourceFile:fs/afs/rxrpc.c SourceLine:56 SourceSnippet: 55:
56: ret = sock_create_kern(net->net, AF_RXRPC, SOCK_DGRAM, PF_INET6, &socket);
57: if (ret < 0)
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:ceph_tcp_connect SourceFile:net/ceph/messenger.c SourceLine:448 SourceSnippet: 447: noio_flag = memalloc_noio_save();
448: ret = sock_create_kern(read_pnet(&con->msgr->net), ss.ss_family,
449: SOCK_STREAM, IPPROTO_TCP, &sock);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:coredump_sock_connect SourceFile:fs/coredump.c SourceLine:687 SourceSnippet: 686:
687: retval = sock_create_kern(&init_net, AF_UNIX, SOCK_STREAM, 0, &socket);
688: if (retval < 0)
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:create_socket SourceFile:fs/smb/server/transport_tcp.c SourceLine:462 SourceSnippet: 461:
462: ret = sock_create_kern(current->nsproxy->net_ns, PF_INET6, SOCK_STREAM,
463: IPPROTO_TCP, &ksmbd_socket);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:create_socket SourceFile:fs/smb/server/transport_tcp.c SourceLine:467 SourceSnippet: 466: pr_err("Can't create socket for ipv6, fallback to ipv4: %d\n", ret);
467: ret = sock_create_kern(current->nsproxy->net_ns, PF_INET,
468: SOCK_STREAM, IPPROTO_TCP, &ksmbd_socket);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:dlm_connect SourceFile:fs/dlm/lowcomms.c SourceLine:1584 SourceSnippet:1583: /* Create a socket to communicate with */
1584: result = sock_create_kern(&init_net, dlm_local_addr[0].ss_family,
1585: SOCK_STREAM, dlm_proto_ops->proto, &sock);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:dlm_listen_for_all SourceFile:fs/dlm/lowcomms.c SourceLine:1765 SourceSnippet:1764:
1765: result = sock_create_kern(&init_net, dlm_local_addr[0].ss_family,
1766: SOCK_STREAM, dlm_proto_ops->proto, &sock);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:generic_ip_connect SourceFile:fs/smb/client/connect.c SourceLine:3359 SourceSnippet:3358:
3359: rc = sock_create_kern(net, sfamily, SOCK_STREAM,
3360: IPPROTO_TCP, &server->ssocket);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:inet_ctl_sock_create SourceFile:net/ipv4/af_inet.c SourceLine:1638 SourceSnippet:1637: struct socket *sock;
1638: int rc = sock_create_kern(net, family, type, protocol, &sock);
1639:
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:l2tp_tunnel_sock_create SourceFile:net/l2tp/l2tp_core.c SourceLine:1499 SourceSnippet:1498:
1499: err = sock_create_kern(net, AF_INET6, SOCK_DGRAM,
1500: IPPROTO_L2TP, &sock);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:l2tp_tunnel_sock_create SourceFile:net/l2tp/l2tp_core.c SourceLine:1527 SourceSnippet:1526:
1527: err = sock_create_kern(net, AF_INET, SOCK_DGRAM,
1528: IPPROTO_L2TP, &sock);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:make_receive_sock SourceFile:net/netfilter/ipvs/ip_vs_sync.c SourceLine:1529 SourceSnippet:1528: /* First create a socket */
1529: result = sock_create_kern(ipvs->net, ipvs->bcfg.mcast_af, SOCK_DGRAM,
1530: IPPROTO_UDP, &sock);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:make_send_sock SourceFile:net/netfilter/ipvs/ip_vs_sync.c SourceLine:1472 SourceSnippet:1471: /* First create a socket */
1472: result = sock_create_kern(ipvs->net, ipvs->mcfg.mcast_af, SOCK_DGRAM,
1473: IPPROTO_UDP, &sock);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:mptcp_pm_nl_create_listen_socket SourceFile:net/mptcp/pm_kernel.c SourceLine:839 SourceSnippet: 838:
839: err = sock_create_kern(sock_net(sk), entry->addr.family,
840: SOCK_STREAM, IPPROTO_MPTCP, &entry->lsk);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:mptcp_subflow_create_socket SourceFile:net/mptcp/subflow.c SourceLine:1790 SourceSnippet:1789:
1790: err = sock_create_kern(net, family, SOCK_STREAM, IPPROTO_TCP, &sf);
1791: if (err)
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:nvme_tcp_alloc_queue SourceFile:drivers/nvme/host/tcp.c SourceLine:1785 SourceSnippet:1784:
1785: ret = sock_create_kern(current->nsproxy->net_ns,
1786: ctrl->addr.ss_family, SOCK_STREAM,
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:qmi_sock_create SourceFile:drivers/soc/qcom/qmi_interface.c SourceLine:591 SourceSnippet: 590:
591: ret = sock_create_kern(&init_net, AF_QIPCRTR, SOCK_DGRAM,
592: PF_QIPCRTR, &sock);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:qrtr_ns_init SourceFile:net/qrtr/ns.c SourceLine:695 SourceSnippet: 694:
695: ret = sock_create_kern(&init_net, AF_QIPCRTR, SOCK_DGRAM,
696: PF_QIPCRTR, &qrtr_ns.sock);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rds_tcp_conn_path_connect SourceFile:net/rds/tcp_connect.c SourceLine:125 SourceSnippet: 124: if (ipv6_addr_v4mapped(&conn->c_laddr)) {
125: ret = sock_create_kern(rds_conn_net(conn), PF_INET,
126: SOCK_STREAM, IPPROTO_TCP, &sock);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rds_tcp_conn_path_connect SourceFile:net/rds/tcp_connect.c SourceLine:129 SourceSnippet: 128: } else {
129: ret = sock_create_kern(rds_conn_net(conn), PF_INET6,
130: SOCK_STREAM, IPPROTO_TCP, &sock);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rds_tcp_listen_init SourceFile:net/rds/tcp_listen.c SourceLine:391 SourceSnippet: 390:
391: ret = sock_create_kern(net, isv6 ? PF_INET6 : PF_INET, SOCK_STREAM,
392: IPPROTO_TCP, &sock);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rfcomm_l2sock_create SourceFile:net/bluetooth/rfcomm/core.c SourceLine:203 SourceSnippet: 202:
203: err = sock_create_kern(&init_net, PF_BLUETOOTH, SOCK_SEQPACKET, BTPROTO_L2CAP, sock);
204: if (!err) {
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rxe_qp_init_req SourceFile:drivers/infiniband/sw/rxe/rxe_qp.c SourceLine:292 SourceSnippet: 291:
292: err = sock_create_kern(&init_net, AF_INET, SOCK_DGRAM, 0, &qp->sk);
293: if (err < 0)
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:smc_create_clcsk SourceFile:net/smc/af_smc.c SourceLine:3347 SourceSnippet:3346:
3347: rc = sock_create_kern(net, family, SOCK_STREAM, IPPROTO_TCP,
3348: &smc->clcsock);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:tipc_topsrv_create_listener SourceFile:net/tipc/topsrv.c SourceLine:518 SourceSnippet: 517:
518: rc = sock_create_kern(srv->net, AF_TIPC, SOCK_SEQPACKET, 0, &lsock);
519: if (rc < 0)
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:udp_sock_create4 SourceFile:net/ipv4/udp_tunnel_core.c SourceLine:19 SourceSnippet: 18:
19: err = sock_create_kern(net, AF_INET, SOCK_DGRAM, 0, &sock);
20: if (err < 0)
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:udp_sock_create6 SourceFile:net/ipv6/ip6_udp_tunnel.c SourceLine:24 SourceSnippet: 23:
24: err = sock_create_kern(net, AF_INET6, SOCK_DGRAM, 0, &sock);
25: if (err < 0)
]] TruncatedOutput:false]
|
| 31/2 |
2026/03/10 00:51 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 39016
output: 41
thoughts: 188
Thoughts: **A Kernel Socket Conundrum**
Okay, so I'm tracing the AFS socket creation and something's bugging me. It looks like `afs_open_socket` calls `sock_create_kern` with `AF_RXRPC`—so AFS is making its own kernel-level socket! I know that means the `sk_net_refcnt` is initialized to zero. That's a red flag. If it's zero, then the network namespace could be destroyed even while the AFS socket itself remains open.
And then, if the namespace gets ripped out from under it, `rxrpc_exit_net` gets called, which in turn calls `rxrpc_destroy_all_calls`. But wait a minute... if the AFS socket is still open and has active calls, isn't that a serious problem? I mean, calls could be in flight, data could be in buffers...it's not pretty.
This makes me wonder if AFS is taking care of closing its sockets gracefully when the network namespace is destroyed. AFS *is* a kernel module, right? Does it have its own cleanup routine? Let's check `afs_net_exit` and see if there's any socket closure in that function. It's time to dig deeper!
|
| 32/2 |
2026/03/10 00:51 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/afs/main.c IncludeLines:true Name:afs_net_exit]
Results: map[SourceCode: 138: /*
139: * Clean up and destroy an AFS network namespace record.
140: */
141: static void __net_exit afs_net_exit(struct net *net_ns)
142: {
143: struct afs_net *net = afs_net(net_ns);
144:
145: net->live = false;
146: afs_fs_probe_cleanup(net);
147: afs_cell_purge(net);
148: afs_wait_for_servers(net);
149: afs_close_socket(net);
150: afs_proc_cleanup(net);
151: afs_put_sysnames(net->sysnames);
152: idr_destroy(&net->cells_dyn_ino);
153: kfree_rcu(rcu_access_pointer(net->address_prefs), rcu);
154: }
SourceFile:fs/afs/main.c]
|
| 33/2 |
2026/03/10 00:51 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 39821
output: 49
thoughts: 89
Thoughts: Okay, let me think this through.
**Network Namespace Shutdown Sequencing**
Right, I'm tracing the network cleanup path, specifically how AFS sockets are handled during network namespace destruction. So, `afs_net_exit` calls `afs_close_socket(net)`. That much is clear; `afs_close_socket` is the function responsible for closing the socket. But now I need to be more precise about the sequence of events. Both `afs_net_exit` and `rxrpc_exit_net` are triggered during network namespace teardown. The key question now becomes, which one gets called *first*? Knowing that order is absolutely critical to avoid any race conditions or resource leaks. Okay, let me consult `afs_net_ops` and `rxrpc_net_ops` to understand the net operations. That should give me the precise sequence I need.
|
| 34/2 |
2026/03/10 00:51 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:fs/afs/main.c IncludeSnippetLines:1 Name:afs_net_ops SourceTreePrefix:]
Results: map[References:<nil> TruncatedOutput:false]
|
| 35/2 |
2026/03/10 00:51 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 40166
output: 39
thoughts: 26
Thoughts: Here's my thought process, summarized as you requested:
**Investigating a Kernel Subsystem**
Okay, first things first. I need to get my bearings. I'm going to jump into the kernel source and search for a specific function I'm familiar with – `register_pernet_subsys`. I remember it's related to registering per-network namespace subsystems, so it's a good place to start my investigation.
Specifically, I'll target the `fs/afs/main.c` file. This tells me I'm looking at the AFS file system code, which makes sense given the context of my current task. This file seems like the most likely location for the relevant functionality, but I might need to trace calls or other related files if I don't find it there. I need to understand what this subsystem does and how it's being used. Let's see what we can find.
|
| 36/2 |
2026/03/10 00:51 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/afs/main.c IncludeLines:true Name:afs_init]
Results: map[SourceCode: 163: /*
164: * initialise the AFS client FS module
165: */
166: static int __init afs_init(void)
167: {
168: int ret = -ENOMEM;
169:
170: printk(KERN_INFO "kAFS: Red Hat AFS client v0.1 registering.\n");
171:
172: afs_wq = alloc_workqueue("afs", WQ_PERCPU, 0);
173: if (!afs_wq)
174: goto error_afs_wq;
175: afs_async_calls = alloc_workqueue("kafsd", WQ_MEM_RECLAIM | WQ_UNBOUND, 0);
176: if (!afs_async_calls)
177: goto error_async;
178: afs_lock_manager = alloc_workqueue("kafs_lockd", WQ_MEM_RECLAIM | WQ_PERCPU, 0);
179: if (!afs_lock_manager)
180: goto error_lockmgr;
181:
182: ret = register_pernet_device(&afs_net_ops);
183: if (ret < 0)
184: goto error_net;
185:
186: /* register the filesystems */
187: ret = afs_fs_init();
188: if (ret < 0)
189: goto error_fs;
190:
191: afs_proc_symlink = proc_symlink("fs/afs", NULL, "../self/net/afs");
192: if (!afs_proc_symlink) {
193: ret = -ENOMEM;
194: goto error_proc;
195: }
196:
197: return ret;
198:
199: error_proc:
200: afs_fs_exit();
201: error_fs:
202: unregister_pernet_device(&afs_net_ops);
203: error_net:
204: destroy_workqueue(afs_lock_manager);
205: error_lockmgr:
206: destroy_workqueue(afs_async_calls);
207: error_async:
208: destroy_workqueue(afs_wq);
209: error_afs_wq:
210: rcu_barrier();
211: printk(KERN_ERR "kAFS: failed to register: %d\n", ret);
212: return ret;
213: }
SourceFile:fs/afs/main.c]
|
| 37/2 |
2026/03/10 00:51 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 41203
output: 45
thoughts: 30
Thoughts: Okay, here's my take on those thoughts, framed as if they're my own:
**Network Device Registration – A Quick Check**
Right, so `afs_net_ops` is definitely taken care of. That's registered using `register_pernet_device`, as expected. Good. But that immediately raises a question in my mind: what about `rxrpc_net_ops`? Did I remember to include that in the registration process as well? It's easy to overlook these details, especially when dealing with multiple network stacks. Gotta make sure both are properly initialized and available. I need to quickly verify that `rxrpc` is also correctly integrated into the per-net device setup. It's a fundamental step that I can't afford to overlook.
|
| 38/2 |
2026/03/10 00:51 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/net_ns.c IncludeLines:true Name:rxrpc_net_ops]
Results: map[SourceCode: 119: struct pernet_operations rxrpc_net_ops = {
120: .init = rxrpc_init_net,
121: .exit = rxrpc_exit_net,
122: .id = &rxrpc_net_id,
123: .size = sizeof(struct rxrpc_net),
124: };
SourceFile:net/rxrpc/net_ns.c]
|
| 39/2 |
2026/03/10 00:51 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 41577
output: 53
thoughts: 12
Thoughts: Here's my thought process, summarized as you requested:
**Tracking Down `rxrpc_net_ops` Registration**
Okay, so where *is* `rxrpc_net_ops` actually getting registered within the kernel? That's the key question. I need to figure out the initialization path for the RxRPC network operations structure. I'm going to have to dive into the kernel source to trace it. My initial thought is to try a quick search to locate the definition first, and then trace its uses. I’ll start with a straightforward `grep` for the identifier itself within the kernel source tree. Depending on the results, I might need to refine the search. If it's a macro or used in a particular context, I'll adjust the search to accommodate that. Once I locate the definition and the initial use, I'll then focus on the call chain leading up to its registration, probably looking for calls related to network initialization or module loading, or anything related to struct net_ops. This will likely involve cross-referencing with other network-related structures and initialization sequences. I'm prepared to use `cscope`, `ctags`, or whatever other tools I have to navigate the code efficiently. I'll need to pay close attention to any conditional compilation or platform-specific logic, as that could influence where and how `rxrpc_net_ops` gets registered. This sounds like it could take a bit, but it should be fairly straightforward to track down with some code spelunking.
|
| 40/2 |
2026/03/10 00:52 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/rxrpc/net_ns.c IncludeSnippetLines:1 Name:rxrpc_net_ops SourceTreePrefix:]
Results: map[References:<nil> TruncatedOutput:false]
|
| 41/2 |
2026/03/10 00:52 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 41980
output: 46
thoughts: 27
Thoughts: Here's my attempt at summarizing the thought process, tailored for an expert audience:
**Investigating `rxrpc_net_ops` in `net/rxrpc/af_rxrpc.c`**
Okay, so I need to understand how RxRPC networking operations are handled, specifically the structure `rxrpc_net_ops`. My immediate thought is to delve into the core source, `net/rxrpc/af_rxrpc.c`. This file should contain the definition and, more importantly, the usage of that `rxrpc_net_ops` structure. I'm expecting to find its registration, probably with `proto_register()` or something similar, and then its linkage within the address family setup. I also want to understand how it's integrated with other aspects of the RxRPC stack. Are there any custom netlink messages that are going through this? What about socket options? I'll be looking for references to this structure within the socket creation, bind, listen, connect, accept and send/receive paths of the address family. This is the starting point to trace the networking initialization and interaction with the protocol. This initial search should reveal the significant hooks and data structures to understand.
|
| 42/2 |
2026/03/10 00:52 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/af_rxrpc.c IncludeLines:true Name:af_rxrpc_init]
Results: map[SourceCode:1035: /*
1036: * initialise and register the RxRPC protocol
1037: */
1038: static int __init af_rxrpc_init(void)
1039: {
1040: int ret = -1;
1041:
1042: BUILD_BUG_ON(sizeof(struct rxrpc_skb_priv) > sizeof_field(struct sk_buff, cb));
1043:
1044: ret = -ENOMEM;
1045: rxrpc_gen_version_string();
1046: rxrpc_call_jar = kmem_cache_create(
1047: "rxrpc_call_jar", sizeof(struct rxrpc_call), 0,
1048: SLAB_HWCACHE_ALIGN, NULL);
1049: if (!rxrpc_call_jar) {
1050: pr_notice("Failed to allocate call jar\n");
1051: goto error_call_jar;
1052: }
1053:
1054: rxrpc_workqueue = alloc_ordered_workqueue("krxrpcd", WQ_HIGHPRI | WQ_MEM_RECLAIM);
1055: if (!rxrpc_workqueue) {
1056: pr_notice("Failed to allocate work queue\n");
1057: goto error_work_queue;
1058: }
1059:
1060: ret = rxrpc_init_security();
1061: if (ret < 0) {
1062: pr_crit("Cannot initialise security\n");
1063: goto error_security;
1064: }
1065:
1066: ret = register_pernet_device(&rxrpc_net_ops);
1067: if (ret)
1068: goto error_pernet;
1069:
1070: ret = proto_register(&rxrpc_proto, 1);
1071: if (ret < 0) {
1072: pr_crit("Cannot register protocol\n");
1073: goto error_proto;
1074: }
1075:
1076: ret = sock_register(&rxrpc_family_ops);
1077: if (ret < 0) {
1078: pr_crit("Cannot register socket family\n");
1079: goto error_sock;
1080: }
1081:
1082: ret = register_key_type(&key_type_rxrpc);
1083: if (ret < 0) {
1084: pr_crit("Cannot register client key type\n");
1085: goto error_key_type;
1086: }
1087:
1088: ret = register_key_type(&key_type_rxrpc_s);
1089: if (ret < 0) {
1090: pr_crit("Cannot register server key type\n");
1091: goto error_key_type_s;
1092: }
1093:
1094: ret = rxrpc_sysctl_init();
1095: if (ret < 0) {
1096: pr_crit("Cannot register sysctls\n");
1097: goto error_sysctls;
1098: }
1099:
1100: return 0;
1101:
1102: error_sysctls:
1103: unregister_key_type(&key_type_rxrpc_s);
1104: error_key_type_s:
1105: unregister_key_type(&key_type_rxrpc);
1106: error_key_type:
1107: sock_unregister(PF_RXRPC);
1108: error_sock:
1109: proto_unregister(&rxrpc_proto);
1110: error_proto:
1111: unregister_pernet_device(&rxrpc_net_ops);
1112: error_pernet:
1113: rxrpc_exit_security();
1114: error_security:
1115: destroy_workqueue(rxrpc_workqueue);
1116: error_work_queue:
1117: kmem_cache_destroy(rxrpc_call_jar);
1118: error_call_jar:
1119: return ret;
1120: }
SourceFile:net/rxrpc/af_rxrpc.c]
|
| 43/2 |
2026/03/10 00:52 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 43562
output: 39
thoughts: 501
Thoughts: **Analyzing Network Namespace Destruction with AF_RXRPC and AFS**
Okay, so I'm looking at how `afs_net_ops` and `rxrpc_net_ops` are registered, specifically with `register_pernet_device`. I know the kernel calls the `exit` routines in the reverse order of the `init` routines. The key here is that `afs` is a separate module and could be loaded *after* `rxrpc`. If that happens, then `afs`'s `init` runs *after* `rxrpc`'s, but its `exit` gets called *before* `rxrpc`'s `exit`. That means `afs_net_exit` runs before `rxrpc_exit_net`.
`afs_net_exit` calls `afs_close_socket(net)`. Right, and `afs_close_socket` closes the AFS socket. Which means `sock_release` is called, and that in turn calls `rxrpc_release`, then `rxrpc_release` calls `rxrpc_release_sock`. That last function removes queued calls from `recvmsg_q` and puts them.
But here's the kicker: what if `recvmsg` is *running* on that AFS socket when it's closed? Is that even possible? I'm seeing a stack trace: `__import_iovec` -> `copy_msghdr_from_user` -> `recvmsg_copy_msghdr` -> `___sys_recvmsg` -> `do_recvmmsg` -> `__sys_recvmmsg`. That's a *user-space* task calling `recvmmsg`! Hold on, can user space even call `recvmmsg` on an AFS socket? No! AFS sockets are kernel sockets, they don't have file descriptors!
Wait, if there's no file descriptor, user space can't call `recvmmsg` on it. Therefore, the socket in the stack trace *must* be a user-space AF_RXRPC socket! Bingo. If it's a user-space AF_RXRPC socket, then its `sk_net_refcnt` is 1! That means the network namespace *cannot* be destroyed while the socket is open. This is key. The namespace has a reference.
So, if the socket is open, how can the network namespace be destroyed? This seems fundamentally impossible. This must be the core issue. I need to re-read `sk_alloc` and look closer at the socket's lifecycle to understand how the refcount works.
|
| 44/2 |
2026/03/10 00:52 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/core/sock.c IncludeLines:true Name:sk_alloc]
Results: map[SourceCode:2288: /**
2289: * sk_alloc - All socket objects are allocated here
2290: * @net: the applicable net namespace
2291: * @family: protocol family
2292: * @priority: for allocation (%GFP_KERNEL, %GFP_ATOMIC, etc)
2293: * @prot: struct proto associated with this new sock instance
2294: * @kern: is this to be a kernel socket?
2295: */
2296: struct sock *sk_alloc(struct net *net, int family, gfp_t priority,
2297: struct proto *prot, int kern)
2298: {
2299: struct sock *sk;
2300:
2301: sk = sk_prot_alloc(prot, priority | __GFP_ZERO, family);
2302: if (sk) {
2303: sk->sk_family = family;
2304: /*
2305: * See comment in struct sock definition to understand
2306: * why we need sk_prot_creator -acme
2307: */
2308: sk->sk_prot = sk->sk_prot_creator = prot;
2309:
2310: if (READ_ONCE(net->core.sysctl_bypass_prot_mem))
2311: sk->sk_bypass_prot_mem = 1;
2312:
2313: sk->sk_kern_sock = kern;
2314: sock_lock_init(sk);
2315:
2316: sk->sk_net_refcnt = kern ? 0 : 1;
2317: if (likely(sk->sk_net_refcnt)) {
2318: get_net_track(net, &sk->ns_tracker, priority);
2319: sock_inuse_add(net, 1);
2320: } else {
2321: net_passive_inc(net);
2322: __netns_tracker_alloc(net, &sk->ns_tracker,
2323: false, priority);
2324: }
2325:
2326: sock_net_set(sk, net);
2327: refcount_set(&sk->sk_wmem_alloc, SK_WMEM_ALLOC_BIAS);
2328:
2329: mem_cgroup_sk_alloc(sk);
2330: cgroup_sk_alloc(&sk->sk_cgrp_data);
2331: sock_update_classid(&sk->sk_cgrp_data);
2332: sock_update_netprioidx(&sk->sk_cgrp_data);
2333: sk_tx_queue_clear(sk);
2334: }
2335:
2336: return sk;
2337: }
SourceFile:net/core/sock.c]
|
| 45/2 |
2026/03/10 00:52 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 45485
output: 45
thoughts: 487
Thoughts: **Decoding This Network Namespace Conundrum**
Okay, let's break this down. My initial thought is: if `sk_net_refcnt` is 1, `get_net_track` gets invoked. `get_net_track`, in turn, is responsible for bumping up the reference count of the network namespace. That means, logically, the network namespace should be locked down until `sk_free` is called. And `sk_free`, I know, is triggered when the socket's reference count hits zero. The real question is, when does the socket's ref count drop to zero? `sock_release` calls `sock_put`, which is what releases a socket. But, ah, `recvmsg` has a hold on the file descriptor, which in turn holds a reference to the socket! So, crucially, `sock_release` *doesn't* get called until `recvmsg` completes. That, in turn, means `sk_free` won't happen. And no `sk_free`? Then the network namespace ref count isn't decremented, meaning it can't be destroyed.
So how the heck is `rxrpc_exit_net` even running in this scenario? Wait a minute... Could `rxrpc_exit_net` be running on a *different* network namespace? That's it! Let's say we've got `net1` and `net2`. `rxrpc_exit_net` is executing for `net1`. But what if `recvmsg` is happening on a socket inside `net2`? But no, if `recvmsg` is happening on a socket in `net2`, its calls belong to `net2`. Why would `rxrpc_destroy_all_calls` in `net1` be waiting for a call in `net2`?
Wait, `rxrpc_destroy_all_calls` is waiting for `rxnet->nr_calls` to reach zero. `rxnet` is actually `rxrpc_net(net1)`. So, it's waiting for calls that are specifically in `net1` to be freed. If `recvmsg` *is* running on a socket in `net1`, then it's impossible to destroy `net1`.
Wait. Could it be possible that `recvmsg` is running on a socket in `net2`, but the underlying call is *associated* with `net1`? That makes no sense. How can a call be associated with a different network namespace than the one the socket is in? Hmm, I need to check how `rxrpc_alloc_call` works.
|
| 46/2 |
2026/03/10 00:52 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_object.c IncludeLines:true Name:rxrpc_alloc_call]
Results: map[SourceCode: 118: /*
119: * allocate a new call
120: */
121: struct rxrpc_call *rxrpc_alloc_call(struct rxrpc_sock *rx, gfp_t gfp,
122: unsigned int debug_id)
123: {
124: struct rxrpc_call *call;
125: struct rxrpc_net *rxnet = rxrpc_net(sock_net(&rx->sk));
126:
127: call = kmem_cache_zalloc(rxrpc_call_jar, gfp);
128: if (!call)
129: return NULL;
130:
131: mutex_init(&call->user_mutex);
132:
133: /* Prevent lockdep reporting a deadlock false positive between the afs
134: * filesystem and sys_sendmsg() via the mmap sem.
135: */
136: if (rx->sk.sk_kern_sock)
137: lockdep_set_class(&call->user_mutex,
138: &rxrpc_call_user_mutex_lock_class_key);
139:
140: timer_setup(&call->timer, rxrpc_call_timer_expired, 0);
141: INIT_WORK(&call->destroyer, rxrpc_destroy_call);
142: INIT_LIST_HEAD(&call->link);
143: INIT_LIST_HEAD(&call->wait_link);
144: INIT_LIST_HEAD(&call->accept_link);
145: INIT_LIST_HEAD(&call->recvmsg_link);
146: INIT_LIST_HEAD(&call->sock_link);
147: INIT_LIST_HEAD(&call->attend_link);
148: skb_queue_head_init(&call->recvmsg_queue);
149: skb_queue_head_init(&call->rx_queue);
150: skb_queue_head_init(&call->rx_oos_queue);
151: init_waitqueue_head(&call->waitq);
152: spin_lock_init(&call->notify_lock);
153: refcount_set(&call->ref, 1);
154: call->debug_id = debug_id;
155: call->tx_total_len = -1;
156: call->tx_jumbo_max = 1;
157: call->next_rx_timo = 20 * HZ;
158: call->next_req_timo = 1 * HZ;
159: call->ackr_window = 1;
160: call->ackr_wtop = 1;
161: call->delay_ack_at = KTIME_MAX;
162: call->rack_timo_at = KTIME_MAX;
163: call->ping_at = KTIME_MAX;
164: call->keepalive_at = KTIME_MAX;
165: call->expect_rx_by = KTIME_MAX;
166: call->expect_req_by = KTIME_MAX;
167: call->expect_term_by = KTIME_MAX;
168:
169: memset(&call->sock_node, 0xed, sizeof(call->sock_node));
170:
171: call->rx_winsize = rxrpc_rx_window_size;
172: call->tx_winsize = 16;
173:
174: call->cong_cwnd = RXRPC_MIN_CWND;
175: call->cong_ssthresh = RXRPC_TX_MAX_WINDOW;
176:
177: rxrpc_call_init_rtt(call);
178:
179: call->rxnet = rxnet;
180: call->rtt_avail = RXRPC_CALL_RTT_AVAIL_MASK;
181: atomic_inc(&rxnet->nr_calls);
182: return call;
183: }
SourceFile:net/rxrpc/call_object.c]
|
| 47/2 |
2026/03/10 00:52 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 47767
output: 55
thoughts: 1546
Thoughts: **Debugging a Persistent RXRPC Network Namespace Cleanup Hang**
Okay, let's break this down. My initial thought was, "Aha! `rxrpc_alloc_call` grabs `rxnet` from the socket's network namespace (`sock_net(&rx->sk)`). So the call and socket are definitely in the same namespace." This implies that if the call is in `net1`, the socket *must* be in `net1`, and thus, `net1` *shouldn't* be destroyed while the call exists. But, that's wrong because I know I see `rxrpc_exit_net` running for `net1`!
My next thought was "maybe it is a kernel socket!". If it's a kernel socket, `sk_net_refcnt` is 0, so the namespace *can* be destroyed even if it's open, but then there's no file descriptor, so how could user space call `recvmmsg`? That's not possible! Okay, well a user space task can't make the socket and give it to a kernel module, because `sk_alloc` sets `sk_net_refcnt` based on the `kern` argument.
I explored the scenario where `recvmmsg` is called on a user socket with `sk_net_refcnt` of 0, but there's no path to make this happen. And then, I considered a TUN/TAP device or something similar, and ruled it out because `rxrpc_recvmsg` is only called if it's an AF_RXRPC socket. So, an AF_RXRPC socket accessed via a file descriptor *must* be a user socket, which *should* hold a network namespace reference. This would still mean the namespace can't be destroyed, so how is `rxrpc_exit_net` running?
I went back to the stack traces to find answers. Here, the trace showed `cleanup_net` hanging in `rxrpc_destroy_all_calls`, which is what's blocking everything. But there's a second stack trace in the log, showing a softlockup in `__import_iovec`. I initially thought the two were connected. However, `__import_iovec` happens before `rxrpc_recvmsg`, meaning `rxrpc_recvmsg` isn't called yet and the lockup isn't in my code.
Then I realized something: "Tainted: [L]=SOFTLOCKUP" and "NMI backtrace for cpu 0". The softlockup is NOT directly related to the `rxrpc` issue! The NMI was triggered by the watchdog *because* `cleanup_net` had been hanging for 144 seconds! The second stack trace is just what CPU 0 happened to be doing when the watchdog kicked in.
Thus, it is clear that the real issue is that `cleanup_net` is hung inside `rxrpc_destroy_all_calls`. It's time to figure out why *that* function is hanging.
|
| 48/2 |
2026/03/10 00:52 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/rxrpc/call_object.c IncludeSnippetLines:1 Name:rxrpc_destroy_all_calls SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rxrpc_exit_net SourceFile:net/rxrpc/net_ns.c SourceLine:112 SourceSnippet: 111: timer_delete_sync(&rxnet->peer_keepalive_timer);
112: rxrpc_destroy_all_calls(rxnet);
113: rxrpc_destroy_all_connections(rxnet);
]] TruncatedOutput:false]
|
| 49/2 |
2026/03/10 00:52 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 50168
output: 53
thoughts: 120
Thoughts: Okay, here's my thinking:
**Debugging an rxrpc_destroy_all_calls Hang**
Alright, the debugger's showing that `rxrpc_destroy_all_calls` is hanging. My initial suspicion is that the issue stems from `rxnet->nr_calls` not decrementing to zero as it should. So, I have to figure out why this counter isn't going down. The fundamental problem must be that some `rxrpc_call` isn't being freed. Logically, if a call isn't freed, it means its reference count isn't reaching zero. Therefore, I need to pinpoint the locations where `rxrpc_call` references are being acquired and released. I'll need to examine the `rxrpc_get_call` and `rxrpc_put_call` functions. These are the critical points of call lifecycle management. That's my next step.
|
| 50/2 |
2026/03/10 00:52 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/rxrpc/call_object.c IncludeSnippetLines:1 Name:rxrpc_get_call SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rxrpc_connect_call SourceFile:net/rxrpc/call_object.c SourceLine:303 SourceSnippet: 302: trace_rxrpc_client(NULL, -1, rxrpc_client_queue_new_call);
303: rxrpc_get_call(call, rxrpc_call_get_io_thread);
304: spin_lock_irq(&local->client_call_lock);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rxrpc_find_call_by_user_ID SourceFile:net/rxrpc/call_object.c SourceLine:112 SourceSnippet: 111: found_extant_call:
112: rxrpc_get_call(call, rxrpc_call_get_sendmsg);
113: read_unlock(&rx->call_lock);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rxrpc_incoming_call SourceFile:net/rxrpc/call_object.c SourceLine:475 SourceSnippet: 474:
475: rxrpc_get_call(call, rxrpc_call_get_io_thread);
476:
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rxrpc_new_client_call SourceFile:net/rxrpc/call_object.c SourceLine:380 SourceSnippet: 379: __set_bit(RXRPC_CALL_HAS_USERID, &call->flags);
380: rxrpc_get_call(call, rxrpc_call_get_userid);
381: rb_link_node(&call->sock_node, parent, pp);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rxrpc_notify_socket SourceFile:net/rxrpc/recvmsg.c SourceLine:49 SourceSnippet: 48: if (list_empty(&call->recvmsg_link)) {
49: rxrpc_get_call(call, rxrpc_call_get_notify_socket);
50: list_add_tail(&call->recvmsg_link, &rx->recvmsg_q);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rxrpc_recvmsg SourceFile:net/rxrpc/recvmsg.c SourceLine:467 SourceSnippet: 466: else
467: rxrpc_get_call(call, rxrpc_call_get_recvmsg);
468: spin_unlock_irq(&rx->recvmsg_lock);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rxrpc_release_calls_on_socket SourceFile:net/rxrpc/call_object.c SourceLine:624 SourceSnippet: 623: struct rxrpc_call, sock_link);
624: rxrpc_get_call(call, rxrpc_call_get_release_sock);
625: rxrpc_propose_abort(call, RX_CALL_DEAD, -ECONNRESET,
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rxrpc_service_prealloc_one SourceFile:net/rxrpc/call_accept.c SourceLine:127 SourceSnippet: 126: rx->app_ops->user_attach_call) {
127: rxrpc_get_call(call, rxrpc_call_get_kernel_service);
128: rx->app_ops->user_attach_call(call, user_call_ID);
] map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rxrpc_service_prealloc_one SourceFile:net/rxrpc/call_accept.c SourceLine:131 SourceSnippet: 130:
131: rxrpc_get_call(call, rxrpc_call_get_userid);
132: rb_link_node(&call->sock_node, parent, pp);
]] TruncatedOutput:false]
|
| 51/2 |
2026/03/10 00:52 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 51683
output: 47
thoughts: 330
Thoughts: **Analyzing RxRPC Call Lifecycle and Socket Closure**
Okay, so I'm tracing the reference counting and call management within RxRPC, trying to understand a potential leak or improper handling. I'm focusing on the interaction between calls, sockets, and how they're released. First, `rxrpc_alloc_call` sets the initial reference count (`call->ref`) to 1. That's the starting point. Then, `rxrpc_release_call` decrements it. Makes sense. This routine gets called when a call is done processing in `rxrpc_recvmsg`, or when the socket itself is being shut down via `rxrpc_release_calls_on_socket`.
Right, let's look at what happens when the socket closes. `rxrpc_release_calls_on_socket` is the key. The code iterates through the `rx->sock_calls` list. It appears that for each call currently associated with the socket, the routine grabs a reference with `rxrpc_get_call`, proposes an abort with `rxrpc_propose_abort`, releases it with `rxrpc_release_call`, and then puts it with `rxrpc_put_call`. That `rxrpc_release_call` call looks like the crucial step – it removes the call from `rx->sock_calls` *and* decrements the reference count initially set.
The part that worries me: what if a call *isn't* in `rx->sock_calls` when the socket closes? Are we going to attempt to double-free, or worse, miss releasing a call entirely? We might leak a reference, leading to a memory leak. I need to be sure that the socket closure logic handles *all* call states correctly. Time to check when and how a call is added to that list. I think I will now check the `rxrpc_new_client_call` function to understand the linkage.
|
| 52/2 |
2026/03/10 00:52 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_object.c IncludeLines:true Name:rxrpc_new_client_call]
Results: map[SourceCode: 315: /*
316: * Set up a call for the given parameters.
317: * - Called with the socket lock held, which it must release.
318: * - If it returns a call, the call's lock will need releasing by the caller.
319: */
320: struct rxrpc_call *rxrpc_new_client_call(struct rxrpc_sock *rx,
321: struct rxrpc_conn_parameters *cp,
322: struct rxrpc_call_params *p,
323: gfp_t gfp,
324: unsigned int debug_id)
325: __releases(&rx->sk.sk_lock)
326: __acquires(&call->user_mutex)
327: {
328: struct rxrpc_call *call, *xcall;
329: struct rxrpc_net *rxnet;
330: struct semaphore *limiter;
331: struct rb_node *parent, **pp;
332: int ret;
333:
334: _enter("%p,%lx", rx, p->user_call_ID);
335:
336: if (WARN_ON_ONCE(!cp->peer)) {
337: release_sock(&rx->sk);
338: return ERR_PTR(-EIO);
339: }
340:
341: limiter = rxrpc_get_call_slot(p, gfp);
342: if (!limiter) {
343: release_sock(&rx->sk);
344: return ERR_PTR(-ERESTARTSYS);
345: }
346:
347: call = rxrpc_alloc_client_call(rx, cp, p, gfp, debug_id);
348: if (IS_ERR(call)) {
349: release_sock(&rx->sk);
350: up(limiter);
351: _leave(" = %ld", PTR_ERR(call));
352: return call;
353: }
354:
355: /* We need to protect a partially set up call against the user as we
356: * will be acting outside the socket lock.
357: */
358: mutex_lock(&call->user_mutex);
359:
360: /* Publish the call, even though it is incompletely set up as yet */
361: write_lock(&rx->call_lock);
362:
363: pp = &rx->calls.rb_node;
364: parent = NULL;
365: while (*pp) {
366: parent = *pp;
367: xcall = rb_entry(parent, struct rxrpc_call, sock_node);
368:
369: if (p->user_call_ID < xcall->user_call_ID)
370: pp = &(*pp)->rb_left;
371: else if (p->user_call_ID > xcall->user_call_ID)
372: pp = &(*pp)->rb_right;
373: else
374: goto error_dup_user_ID;
375: }
376:
377: rcu_assign_pointer(call->socket, rx);
378: call->user_call_ID = p->user_call_ID;
379: __set_bit(RXRPC_CALL_HAS_USERID, &call->flags);
380: rxrpc_get_call(call, rxrpc_call_get_userid);
381: rb_link_node(&call->sock_node, parent, pp);
382: rb_insert_color(&call->sock_node, &rx->calls);
383: list_add(&call->sock_link, &rx->sock_calls);
384:
385: write_unlock(&rx->call_lock);
386:
387: rxnet = call->rxnet;
388: spin_lock(&rxnet->call_lock);
389: list_add_tail_rcu(&call->link, &rxnet->calls);
390: spin_unlock(&rxnet->call_lock);
391:
392: /* From this point on, the call is protected by its own lock. */
393: release_sock(&rx->sk);
394:
395: /* Set up or get a connection record and set the protocol parameters,
396: * including channel number and call ID.
397: */
398: ret = rxrpc_connect_call(call, gfp);
399: if (ret < 0)
400: goto error_attached_to_socket;
401:
402: _leave(" = %p [new]", call);
403: return call;
404:
405: /* We unexpectedly found the user ID in the list after taking
406: * the call_lock. This shouldn't happen unless the user races
407: * with itself and tries to add the same user ID twice at the
408: * same time in different threads.
409: */
410: error_dup_user_ID:
411: write_unlock(&rx->call_lock);
412: release_sock(&rx->sk);
413: rxrpc_prefail_call(call, RXRPC_CALL_LOCAL_ERROR, -EEXIST);
414: trace_rxrpc_call(call->debug_id, refcount_read(&call->ref), 0,
415: rxrpc_call_see_userid_exists);
416: mutex_unlock(&call->user_mutex);
417: rxrpc_put_call(call, rxrpc_call_put_userid_exists);
418: _leave(" = -EEXIST");
419: return ERR_PTR(-EEXIST);
420:
421: /* We got an error, but the call is attached to the socket and is in
422: * need of release. However, we might now race with recvmsg() when it
423: * completion notifies the socket. Return 0 from sys_sendmsg() and
424: * leave the error to recvmsg() to deal with.
425: */
426: error_attached_to_socket:
427: trace_rxrpc_call(call->debug_id, refcount_read(&call->ref), ret,
428: rxrpc_call_see_connect_failed);
429: rxrpc_set_call_completion(call, RXRPC_CALL_LOCAL_ERROR, 0, ret);
430: _leave(" = c=%08x [err]", call->debug_id);
431: return call;
432: }
SourceFile:net/rxrpc/call_object.c]
|
| 53/2 |
2026/03/10 00:52 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 54425
output: 47
thoughts: 140
Thoughts: **Initial Client Call Reference Count Investigation**
Okay, so I'm tracing the reference count management for these RxRPC client calls. I see that `rxrpc_new_client_call` calls `rxrpc_alloc_client_call`, which makes sense. Digging deeper, `rxrpc_alloc_client_call` in turn calls `rxrpc_alloc_call`. Right, `rxrpc_alloc_call` is where the initial reference is set - it sets `call->ref` to 1. That tracks.
Now, `rxrpc_new_client_call` calls `rxrpc_get_call` to add the call to `rx->sock_calls`, specifically using `rxrpc_call_get_userid`. And I know that `rxrpc_get_call` increments the reference count. So, if `call->ref` starts at 1 and then `rxrpc_get_call` is called, `call->ref` should now be 2.
This raises a question in my mind: If the initial reference count is 1 and another reference is acquired, who releases that original reference? I need to verify that. I'm going to revisit `rxrpc_alloc_client_call` to refresh my memory on the initial reference's ownership.
|
| 54/2 |
2026/03/10 00:53 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_object.c IncludeLines:true Name:rxrpc_alloc_client_call]
Results: map[SourceCode: 185: /*
186: * Allocate a new client call.
187: */
188: static struct rxrpc_call *rxrpc_alloc_client_call(struct rxrpc_sock *rx,
189: struct rxrpc_conn_parameters *cp,
190: struct rxrpc_call_params *p,
191: gfp_t gfp,
192: unsigned int debug_id)
193: {
194: struct rxrpc_call *call;
195: ktime_t now;
196: int ret;
197:
198: _enter("");
199:
200: call = rxrpc_alloc_call(rx, gfp, debug_id);
201: if (!call)
202: return ERR_PTR(-ENOMEM);
203: now = ktime_get_real();
204: call->acks_latest_ts = now;
205: call->cong_tstamp = now;
206: call->dest_srx = cp->peer->srx;
207: call->dest_srx.srx_service = cp->service_id;
208: call->interruptibility = p->interruptibility;
209: call->tx_total_len = p->tx_total_len;
210: call->key = key_get(cp->key);
211: call->peer = rxrpc_get_peer(cp->peer, rxrpc_peer_get_call);
212: call->local = rxrpc_get_local(cp->local, rxrpc_local_get_call);
213: call->security_level = cp->security_level;
214: if (p->kernel)
215: __set_bit(RXRPC_CALL_KERNEL, &call->flags);
216: if (cp->upgrade)
217: __set_bit(RXRPC_CALL_UPGRADE, &call->flags);
218: if (cp->exclusive)
219: __set_bit(RXRPC_CALL_EXCLUSIVE, &call->flags);
220:
221: if (p->timeouts.normal)
222: call->next_rx_timo = umin(p->timeouts.normal, 1);
223: if (p->timeouts.idle)
224: call->next_req_timo = umin(p->timeouts.idle, 1);
225: if (p->timeouts.hard)
226: call->hard_timo = p->timeouts.hard;
227:
228: ret = rxrpc_init_client_call_security(call);
229: if (ret < 0) {
230: rxrpc_prefail_call(call, RXRPC_CALL_LOCAL_ERROR, ret);
231: rxrpc_put_call(call, rxrpc_call_put_discard_error);
232: return ERR_PTR(ret);
233: }
234:
235: rxrpc_set_call_state(call, RXRPC_CALL_CLIENT_AWAIT_CONN);
236:
237: trace_rxrpc_call(call->debug_id, refcount_read(&call->ref),
238: p->user_call_ID, rxrpc_call_new_client);
239:
240: _leave(" = %p", call);
241: return call;
242: }
SourceFile:net/rxrpc/call_object.c]
|
| 55/2 |
2026/03/10 00:53 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 55915
output: 44
thoughts: 154
Thoughts: **Analyzing Client Call Reference Counting**
Okay, let's break this down. First, `rxrpc_alloc_client_call` hands me back a client call with a `ref` count of 1. Seems straightforward. But then, `rxrpc_new_client_call` uses `rxrpc_get_call` which I know increments that `ref` count, and it adds it to `rx->sock_calls`. So now we're at a `ref` count of 2. Right?
`rxrpc_new_client_call` then *returns* this call to `rxrpc_send_data`. My concern is where the reference count goes from here. `rxrpc_send_data` takes that call and, after using it, calls `rxrpc_put_call`. That's the crucial step. So *`rxrpc_send_data` is dropping a reference*! But wait, does `rxrpc_send_data` drop the *initial* reference or the one it gained from being sent to it from `rxrpc_new_client_call`?
Hold on... I need to double-check this. Let me just open up `rxrpc_send_data` and see exactly what's going on.
|
| 56/2 |
2026/03/10 00:53 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/sendmsg.c IncludeLines:true Name:rxrpc_send_data]
Results: map[SourceCode: 315: /*
316: * send data through a socket
317: * - must be called in process context
318: * - The caller holds the call user access mutex, but not the socket lock.
319: */
320: static int rxrpc_send_data(struct rxrpc_sock *rx,
321: struct rxrpc_call *call,
322: struct msghdr *msg, size_t len,
323: rxrpc_notify_end_tx_t notify_end_tx,
324: bool *_dropped_lock)
325: {
326: struct rxrpc_txbuf *txb;
327: struct sock *sk = &rx->sk;
328: enum rxrpc_call_state state;
329: long timeo;
330: bool more = msg->msg_flags & MSG_MORE;
331: int ret, copied = 0;
332:
333: if (test_bit(RXRPC_CALL_TX_NO_MORE, &call->flags)) {
334: trace_rxrpc_abort(call->debug_id, rxrpc_sendmsg_late_send,
335: call->cid, call->call_id, call->rx_consumed,
336: 0, -EPROTO);
337: return -EPROTO;
338: }
339:
340: timeo = sock_sndtimeo(sk, msg->msg_flags & MSG_DONTWAIT);
341:
342: ret = rxrpc_wait_to_be_connected(call, &timeo);
343: if (ret < 0)
344: return ret;
345:
346: if (call->conn->state == RXRPC_CONN_CLIENT_UNSECURED) {
347: ret = rxrpc_init_client_conn_security(call->conn);
348: if (ret < 0)
349: return ret;
350: }
351:
352: /* this should be in poll */
353: sk_clear_bit(SOCKWQ_ASYNC_NOSPACE, sk);
354:
355: reload:
356: txb = call->tx_pending;
357: call->tx_pending = NULL;
358: if (txb)
359: rxrpc_see_txbuf(txb, rxrpc_txbuf_see_send_more);
360:
361: ret = -EPIPE;
362: if (sk->sk_shutdown & SEND_SHUTDOWN)
363: goto maybe_error;
364: state = rxrpc_call_state(call);
365: ret = -ESHUTDOWN;
366: if (state >= RXRPC_CALL_COMPLETE)
367: goto maybe_error;
368: ret = -EPROTO;
369: if (state != RXRPC_CALL_CLIENT_SEND_REQUEST &&
370: state != RXRPC_CALL_SERVER_ACK_REQUEST &&
371: state != RXRPC_CALL_SERVER_SEND_REPLY) {
372: /* Request phase complete for this client call */
373: trace_rxrpc_abort(call->debug_id, rxrpc_sendmsg_late_send,
374: call->cid, call->call_id, call->rx_consumed,
375: 0, -EPROTO);
376: goto maybe_error;
377: }
378:
379: ret = -EMSGSIZE;
380: if (call->tx_total_len != -1) {
381: if (len - copied > call->tx_total_len)
382: goto maybe_error;
383: if (!more && len - copied != call->tx_total_len)
384: goto maybe_error;
385: }
386:
387: do {
388: if (!txb) {
389: size_t remain;
390:
391: _debug("alloc");
392:
393: if (!rxrpc_check_tx_space(call, NULL))
394: goto wait_for_space;
395:
396: /* See if we need to begin/extend the Tx queue. */
397: if (!call->send_queue || !((call->send_top + 1) & RXRPC_TXQ_MASK)) {
398: ret = rxrpc_alloc_txqueue(sk, call);
399: if (ret < 0)
400: goto maybe_error;
401: }
402:
403: /* Work out the maximum size of a packet. Assume that
404: * the security header is going to be in the padded
405: * region (enc blocksize), but the trailer is not.
406: */
407: remain = more ? INT_MAX : msg_data_left(msg);
408: txb = call->conn->security->alloc_txbuf(call, remain, sk->sk_allocation);
409: if (!txb) {
410: ret = -ENOMEM;
411: goto maybe_error;
412: }
413: }
414:
415: _debug("append");
416:
417: /* append next segment of data to the current buffer */
418: if (msg_data_left(msg) > 0) {
419: size_t copy = umin(txb->space, msg_data_left(msg));
420:
421: _debug("add %zu", copy);
422: if (!copy_from_iter_full(txb->data + txb->offset,
423: copy, &msg->msg_iter))
424: goto efault;
425: _debug("added");
426: txb->space -= copy;
427: txb->len += copy;
428: txb->offset += copy;
429: copied += copy;
430: if (call->tx_total_len != -1)
431: call->tx_total_len -= copy;
432: }
433:
434: /* check for the far side aborting the call or a network error
435: * occurring */
436: if (rxrpc_call_is_complete(call))
437: goto call_terminated;
438:
439: /* add the packet to the send queue if it's now full */
440: if (!txb->space ||
441: (msg_data_left(msg) == 0 && !more)) {
442: if (msg_data_left(msg) == 0 && !more)
443: txb->flags |= RXRPC_LAST_PACKET;
444:
445: ret = call->security->secure_packet(call, txb);
446: if (ret < 0)
447: goto out;
448: rxrpc_queue_packet(rx, call, txb, notify_end_tx);
449: txb = NULL;
450: }
451: } while (msg_data_left(msg) > 0);
452:
453: success:
454: ret = copied;
455: if (rxrpc_call_is_complete(call) &&
456: call->error < 0)
457: ret = call->error;
458: out:
459: call->tx_pending = txb;
460: _leave(" = %d", ret);
461: return ret;
462:
463: call_terminated:
464: rxrpc_put_txbuf(txb, rxrpc_txbuf_put_send_aborted);
465: _leave(" = %d", call->error);
466: return call->error;
467:
468: maybe_error:
469: if (copied)
470: goto success;
471: goto out;
472:
473: efault:
474: ret = -EFAULT;
475: goto out;
476:
477: wait_for_space:
478: ret = -EAGAIN;
479: if (msg->msg_flags & MSG_DONTWAIT)
480: goto maybe_error;
481: mutex_unlock(&call->user_mutex);
482: *_dropped_lock = true;
483: ret = rxrpc_wait_for_tx_window(rx, call, &timeo,
484: msg->msg_flags & MSG_WAITALL);
485: if (ret < 0)
486: goto maybe_error;
487: if (call->interruptibility == RXRPC_INTERRUPTIBLE) {
488: if (mutex_lock_interruptible(&call->user_mutex) < 0) {
489: ret = sock_intr_errno(timeo);
490: goto maybe_error;
491: }
492: } else {
493: mutex_lock(&call->user_mutex);
494: }
495: *_dropped_lock = false;
496: goto reload;
497: }
SourceFile:net/rxrpc/sendmsg.c]
|
| 57/2 |
2026/03/10 00:53 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 59154
output: 45
thoughts: 39
Thoughts: Okay, let's see. My mind is immediately drawn to a potential point of concern.
**Concern Regarding Initial Reference in `rxrpc_send_data`**
Hold on a second, I was just reviewing the `rxrpc_send_data` function, and it's suddenly dawned on me: I don't see anything explicitly dropping that *initial* reference. This is a crucial area. If the function isn't responsible for decrementing the reference count at the beginning, I need to know *where* that initial reference is getting dropped. It’s an easy mistake that can have significant implications. Now, I have to trace it and confirm whether there is a reference leak. My gut tells me this is an area where subtle bugs can hide, especially with complex data structures. The most likely candidate function is `rxrpc_do_sendmsg`. That's where I need to focus my attention next. I'll check that one.
|
| 58/2 |
2026/03/10 00:53 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/sendmsg.c IncludeLines:true Name:rxrpc_do_sendmsg]
Results: map[SourceCode: 654: /*
655: * send a message forming part of a client call through an RxRPC socket
656: * - caller holds the socket locked
657: * - the socket may be either a client socket or a server socket
658: */
659: int rxrpc_do_sendmsg(struct rxrpc_sock *rx, struct msghdr *msg, size_t len)
660: {
661: struct rxrpc_call *call;
662: bool dropped_lock = false;
663: int ret;
664:
665: struct rxrpc_send_params p = {
666: .call.tx_total_len = -1,
667: .call.user_call_ID = 0,
668: .call.nr_timeouts = 0,
669: .call.interruptibility = RXRPC_INTERRUPTIBLE,
670: .abort_code = 0,
671: .command = RXRPC_CMD_SEND_DATA,
672: .exclusive = false,
673: .upgrade = false,
674: };
675:
676: _enter("");
677:
678: ret = rxrpc_sendmsg_cmsg(msg, &p);
679: if (ret < 0)
680: goto error_release_sock;
681:
682: if (p.command == RXRPC_CMD_CHARGE_ACCEPT) {
683: ret = -EINVAL;
684: if (rx->sk.sk_state != RXRPC_SERVER_LISTENING)
685: goto error_release_sock;
686: ret = rxrpc_user_charge_accept(rx, p.call.user_call_ID);
687: goto error_release_sock;
688: }
689:
690: call = rxrpc_find_call_by_user_ID(rx, p.call.user_call_ID);
691: if (!call) {
692: ret = -EBADSLT;
693: if (p.command != RXRPC_CMD_SEND_DATA)
694: goto error_release_sock;
695: call = rxrpc_new_client_call_for_sendmsg(rx, msg, &p);
696: /* The socket is now unlocked... */
697: if (IS_ERR(call))
698: return PTR_ERR(call);
699: /* ... and we have the call lock. */
700: p.call.nr_timeouts = 0;
701: ret = 0;
702: if (rxrpc_call_is_complete(call))
703: goto out_put_unlock;
704: } else {
705: switch (rxrpc_call_state(call)) {
706: case RXRPC_CALL_CLIENT_AWAIT_CONN:
707: case RXRPC_CALL_SERVER_RECV_REQUEST:
708: if (p.command == RXRPC_CMD_SEND_ABORT)
709: break;
710: fallthrough;
711: case RXRPC_CALL_UNINITIALISED:
712: case RXRPC_CALL_SERVER_PREALLOC:
713: rxrpc_put_call(call, rxrpc_call_put_sendmsg);
714: ret = -EBUSY;
715: goto error_release_sock;
716: default:
717: break;
718: }
719:
720: ret = mutex_lock_interruptible(&call->user_mutex);
721: release_sock(&rx->sk);
722: if (ret < 0) {
723: ret = -ERESTARTSYS;
724: goto error_put;
725: }
726:
727: if (p.call.tx_total_len != -1) {
728: ret = -EINVAL;
729: if (call->tx_total_len != -1 ||
730: call->tx_pending ||
731: call->tx_top != 0)
732: goto out_put_unlock;
733: call->tx_total_len = p.call.tx_total_len;
734: }
735: }
736:
737: switch (p.call.nr_timeouts) {
738: case 3:
739: WRITE_ONCE(call->next_rx_timo, p.call.timeouts.normal);
740: fallthrough;
741: case 2:
742: WRITE_ONCE(call->next_req_timo, p.call.timeouts.idle);
743: fallthrough;
744: case 1:
745: if (p.call.timeouts.hard > 0) {
746: ktime_t delay = ms_to_ktime(p.call.timeouts.hard * MSEC_PER_SEC);
747:
748: WRITE_ONCE(call->expect_term_by,
749: ktime_add(p.call.timeouts.hard,
750: ktime_get_real()));
751: trace_rxrpc_timer_set(call, delay, rxrpc_timer_trace_hard);
752: rxrpc_poke_call(call, rxrpc_call_poke_set_timeout);
753:
754: }
755: break;
756: }
757:
758: if (rxrpc_call_is_complete(call)) {
759: /* it's too late for this call */
760: ret = -ESHUTDOWN;
761: goto out_put_unlock;
762: }
763:
764: switch (p.command) {
765: case RXRPC_CMD_SEND_ABORT:
766: rxrpc_propose_abort(call, p.abort_code, -ECONNABORTED,
767: rxrpc_abort_call_sendmsg);
768: ret = 0;
769: break;
770: case RXRPC_CMD_SEND_DATA:
771: ret = rxrpc_send_data(rx, call, msg, len, NULL, &dropped_lock);
772: break;
773: default:
774: ret = -EINVAL;
775: break;
776: }
777:
778: out_put_unlock:
779: if (!dropped_lock)
780: mutex_unlock(&call->user_mutex);
781: error_put:
782: rxrpc_put_call(call, rxrpc_call_put_sendmsg);
783: _leave(" = %d", ret);
784: return ret;
785:
786: error_release_sock:
787: release_sock(&rx->sk);
788: return ret;
789: }
SourceFile:net/rxrpc/sendmsg.c]
|
| 59/2 |
2026/03/10 00:53 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 61530
output: 45
thoughts: 315
Thoughts: **Debugging a Potential RXRPC Call Leak**
Okay, let's break this down. I'm looking at `rxrpc_do_sendmsg`, and it looks like it creates a new client call using `rxrpc_new_client_call_for_sendmsg`. That call comes back with a reference count of 2. One reference is held by the socket, the other is for the caller. But, right at the end of `rxrpc_do_sendmsg`, we have `rxrpc_put_call(call, rxrpc_call_put_sendmsg)`. This drops the caller's reference.
So, the remaining reference *should* be the one in `rx->sock_calls`. Then, when the socket is closed, `rxrpc_release_calls_on_socket` *should* drop that reference, and the call's reference count *should* hit zero, which *should* free the call and decrement `rxnet->nr_calls`. But something is going wrong. `rxrpc_destroy_all_calls` is hanging, which indicates a reference count that's *not* reaching zero.
So, where is the leak? I need to find where another reference is being held. Let's see... `rxrpc_get_call` is key here. Aha! `rxrpc_connect_call` calls `rxrpc_get_call(call, rxrpc_call_get_io_thread)`. It looks like `rxrpc_connect_call` is grabbing a reference for the I/O thread. Alright, I need to check `rxrpc_connect_call` to see if that reference is being properly released.
|
| 60/2 |
2026/03/10 00:53 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_object.c IncludeLines:true Name:rxrpc_connect_call]
Results: map[SourceCode: 286: /*
287: * Start the process of connecting a call. We obtain a peer and a connection
288: * bundle, but the actual association of a call with a connection is offloaded
289: * to the I/O thread to simplify locking.
290: */
291: static int rxrpc_connect_call(struct rxrpc_call *call, gfp_t gfp)
292: {
293: struct rxrpc_local *local = call->local;
294: int ret = -ENOMEM;
295:
296: _enter("{%d,%lx},", call->debug_id, call->user_call_ID);
297:
298: ret = rxrpc_look_up_bundle(call, gfp);
299: if (ret < 0)
300: goto error;
301:
302: trace_rxrpc_client(NULL, -1, rxrpc_client_queue_new_call);
303: rxrpc_get_call(call, rxrpc_call_get_io_thread);
304: spin_lock_irq(&local->client_call_lock);
305: list_add_tail(&call->wait_link, &local->new_client_calls);
306: spin_unlock_irq(&local->client_call_lock);
307: rxrpc_wake_up_io_thread(local);
308: return 0;
309:
310: error:
311: __set_bit(RXRPC_CALL_DISCONNECTED, &call->flags);
312: return ret;
313: }
SourceFile:net/rxrpc/call_object.c]
|
| 61/2 |
2026/03/10 00:53 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 62754
output: 45
thoughts: 67
Thoughts: Here's my understanding, as if I were thinking it:
**Decoding the RxRPC Connection Flow**
Okay, so I'm tracing the connection process and, if I'm reading this right, `rxrpc_connect_call` is grabbing a reference and then stashing the call object into `local->new_client_calls`. That suggests it's essentially queueing up new client calls for processing. Then, it triggers a wake-up signal for the I/O thread. Makes sense; got to let that thread know there's work to be done.
So, the next step is the I/O thread, which is likely polling that `local->new_client_calls` queue and working through each of the entries. Let's dig into that `rxrpc_io_thread` function to confirm my thinking and flesh out the details of what's happening on that end. I need to understand how those queued calls are actually handled by the I/O thread.
|
| 62/2 |
2026/03/10 00:53 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/io_thread.c IncludeLines:true Name:rxrpc_io_thread]
Results: map[SourceCode: 446: /*
447: * I/O and event handling thread.
448: */
449: int rxrpc_io_thread(void *data)
450: {
451: struct rxrpc_connection *conn;
452: struct sk_buff_head rx_queue;
453: struct rxrpc_local *local = data;
454: struct rxrpc_call *call;
455: struct sk_buff *skb;
456: #ifdef CONFIG_AF_RXRPC_INJECT_RX_DELAY
457: ktime_t now;
458: #endif
459: bool should_stop;
460: LIST_HEAD(conn_attend_q);
461: LIST_HEAD(call_attend_q);
462:
463: complete(&local->io_thread_ready);
464:
465: skb_queue_head_init(&rx_queue);
466:
467: set_user_nice(current, MIN_NICE);
468:
469: for (;;) {
470: rxrpc_inc_stat(local->rxnet, stat_io_loop);
471:
472: /* Inject a delay into packets if requested. */
473: #ifdef CONFIG_AF_RXRPC_INJECT_RX_DELAY
474: now = ktime_get_real();
475: while ((skb = skb_peek(&local->rx_delay_queue))) {
476: if (ktime_before(now, skb->tstamp))
477: break;
478: skb = skb_dequeue(&local->rx_delay_queue);
479: skb_queue_tail(&local->rx_queue, skb);
480: }
481: #endif
482:
483: if (!skb_queue_empty(&local->rx_queue)) {
484: spin_lock_irq(&local->rx_queue.lock);
485: skb_queue_splice_tail_init(&local->rx_queue, &rx_queue);
486: spin_unlock_irq(&local->rx_queue.lock);
487: trace_rxrpc_iothread_rx(local, skb_queue_len(&rx_queue));
488: }
489:
490: /* Distribute packets and errors. */
491: while ((skb = __skb_dequeue(&rx_queue))) {
492: struct rxrpc_skb_priv *sp = rxrpc_skb(skb);
493: switch (skb->mark) {
494: case RXRPC_SKB_MARK_PACKET:
495: skb->priority = 0;
496: if (!rxrpc_input_packet(local, &skb))
497: rxrpc_reject_packet(local, skb);
498: trace_rxrpc_rx_done(skb->mark, skb->priority);
499: rxrpc_free_skb(skb, rxrpc_skb_put_input);
500: break;
501: case RXRPC_SKB_MARK_ERROR:
502: rxrpc_input_error(local, skb);
503: rxrpc_free_skb(skb, rxrpc_skb_put_error_report);
504: break;
505: case RXRPC_SKB_MARK_SERVICE_CONN_SECURED:
506: rxrpc_input_conn_event(sp->poke_conn, skb);
507: rxrpc_put_connection(sp->poke_conn, rxrpc_conn_put_poke);
508: rxrpc_free_skb(skb, rxrpc_skb_put_conn_secured);
509: break;
510: default:
511: WARN_ON_ONCE(1);
512: rxrpc_free_skb(skb, rxrpc_skb_put_unknown);
513: break;
514: }
515: }
516:
517: /* Deal with connections that want immediate attention. */
518: if (!list_empty_careful(&local->conn_attend_q)) {
519: spin_lock_irq(&local->lock);
520: list_splice_tail_init(&local->conn_attend_q, &conn_attend_q);
521: spin_unlock_irq(&local->lock);
522: }
523:
524: while ((conn = list_first_entry_or_null(&conn_attend_q,
525: struct rxrpc_connection,
526: attend_link))) {
527: spin_lock_irq(&local->lock);
528: list_del_init(&conn->attend_link);
529: spin_unlock_irq(&local->lock);
530: rxrpc_input_conn_event(conn, NULL);
531: rxrpc_put_connection(conn, rxrpc_conn_put_poke);
532: }
533:
534: if (test_and_clear_bit(RXRPC_CLIENT_CONN_REAP_TIMER,
535: &local->client_conn_flags))
536: rxrpc_discard_expired_client_conns(local);
537:
538: /* Deal with calls that want immediate attention. */
539: spin_lock_irq(&local->lock);
540: list_splice_tail_init(&local->call_attend_q, &call_attend_q);
541: spin_unlock_irq(&local->lock);
542:
543: while ((call = list_first_entry_or_null(&call_attend_q,
544: struct rxrpc_call,
545: attend_link))) {
546: spin_lock_irq(&local->lock);
547: list_del_init(&call->attend_link);
548: spin_unlock_irq(&local->lock);
549: trace_rxrpc_call_poked(call);
550: rxrpc_input_call_event(call);
551: rxrpc_put_call(call, rxrpc_call_put_poke);
552: }
553:
554: if (!list_empty(&local->new_client_calls))
555: rxrpc_connect_client_calls(local);
556:
557: set_current_state(TASK_INTERRUPTIBLE);
558: should_stop = kthread_should_stop();
559: if (!skb_queue_empty(&local->rx_queue) ||
560: !list_empty(&local->call_attend_q) ||
561: !list_empty(&local->conn_attend_q) ||
562: !list_empty(&local->new_client_calls) ||
563: test_bit(RXRPC_CLIENT_CONN_REAP_TIMER,
564: &local->client_conn_flags)) {
565: __set_current_state(TASK_RUNNING);
566: continue;
567: }
568:
569: if (should_stop)
570: break;
571:
572: #ifdef CONFIG_AF_RXRPC_INJECT_RX_DELAY
573: skb = skb_peek(&local->rx_delay_queue);
574: if (skb) {
575: unsigned long timeout;
576: ktime_t tstamp = skb->tstamp;
577: ktime_t now = ktime_get_real();
578: s64 delay_ns = ktime_to_ns(ktime_sub(tstamp, now));
579:
580: if (delay_ns <= 0) {
581: __set_current_state(TASK_RUNNING);
582: continue;
583: }
584:
585: timeout = nsecs_to_jiffies(delay_ns);
586: timeout = umax(timeout, 1);
587: schedule_timeout(timeout);
588: __set_current_state(TASK_RUNNING);
589: continue;
590: }
591: #endif
592:
593: schedule();
594: }
595:
596: __set_current_state(TASK_RUNNING);
597: rxrpc_see_local(local, rxrpc_local_stop);
598: rxrpc_destroy_local(local);
599: WRITE_ONCE(local->io_thread, NULL);
600: rxrpc_see_local(local, rxrpc_local_stopped);
601: return 0;
602: }
SourceFile:net/rxrpc/io_thread.c]
|
| 63/2 |
2026/03/10 00:53 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 65667
output: 47
thoughts: 50
Thoughts: Okay, let's see...
**Analyzing Client Connection Handling in RxRPC**
Hmm, so the `rxrpc_io_thread` is responsible for dealing with `local->new_client_calls`. And it's doing that by calling `rxrpc_connect_client_calls(local)`. Right, that's the key function. I need to understand how *that* works. Time to dive into `rxrpc_connect_client_calls` and figure out exactly what it's doing with these new client calls. I'm assuming it handles the connection establishment, but I want to verify the exact process. Let's dig in.
|
| 64/2 |
2026/03/10 00:53 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/conn_client.c IncludeLines:true Name:rxrpc_connect_client_calls]
Results: map[SourceCode: 505: /*
506: * Connect waiting channels (called from the I/O thread).
507: */
508: void rxrpc_connect_client_calls(struct rxrpc_local *local)
509: {
510: struct rxrpc_call *call;
511: LIST_HEAD(new_client_calls);
512:
513: spin_lock_irq(&local->client_call_lock);
514: list_splice_tail_init(&local->new_client_calls, &new_client_calls);
515: spin_unlock_irq(&local->client_call_lock);
516:
517: while ((call = list_first_entry_or_null(&new_client_calls,
518: struct rxrpc_call, wait_link))) {
519: struct rxrpc_bundle *bundle = call->bundle;
520:
521: list_move_tail(&call->wait_link, &bundle->waiting_calls);
522: rxrpc_see_call(call, rxrpc_call_see_waiting_call);
523:
524: if (rxrpc_bundle_has_space(bundle))
525: rxrpc_activate_channels(bundle);
526: }
527: }
SourceFile:net/rxrpc/conn_client.c]
|
| 65/2 |
2026/03/10 00:53 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 66308
output: 45
thoughts: 51
Thoughts: Okay, let me think this through.
**Analyzing RxRPC Client Connection and Call Management**
Right, so `rxrpc_connect_client_calls` is the function that handles putting calls onto the connection. Hold on... it *moves* the call to `bundle->waiting_calls`. That's important. It doesn't actually decrement the reference count or release the call object. So, the reference isn't dropped there. Okay, so where *is* the reference dropped? The crucial part is to figure out the lifecycle of these call objects. It has to be elsewhere. I need to trace this further. Let's dig into `rxrpc_activate_channels` and see if we can find where the call objects' resources are actually freed up. I bet the call lifecycle continues there.
|
| 66/2 |
2026/03/10 00:53 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/conn_client.c IncludeLines:true Name:rxrpc_activate_channels]
Results: map[SourceCode: 467: /*
468: * Assign channels and callNumbers to waiting calls.
469: */
470: static void rxrpc_activate_channels(struct rxrpc_bundle *bundle)
471: {
472: struct rxrpc_connection *conn;
473: unsigned long avail, mask;
474: unsigned int channel, slot;
475:
476: trace_rxrpc_client(NULL, -1, rxrpc_client_activate_chans);
477:
478: if (bundle->try_upgrade)
479: mask = 1;
480: else
481: mask = ULONG_MAX;
482:
483: while (!list_empty(&bundle->waiting_calls)) {
484: avail = bundle->avail_chans & mask;
485: if (!avail)
486: break;
487: channel = __ffs(avail);
488: clear_bit(channel, &bundle->avail_chans);
489:
490: slot = channel / RXRPC_MAXCALLS;
491: conn = bundle->conns[slot];
492: if (!conn)
493: break;
494:
495: if (bundle->try_upgrade)
496: set_bit(RXRPC_CONN_PROBING_FOR_UPGRADE, &conn->flags);
497: rxrpc_unidle_conn(conn);
498:
499: channel &= (RXRPC_MAXCALLS - 1);
500: conn->act_chans |= 1 << channel;
501: rxrpc_activate_one_channel(conn, channel);
502: }
503: }
SourceFile:net/rxrpc/conn_client.c]
|
| 67/2 |
2026/03/10 00:53 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 67119
output: 47
thoughts: 19
Thoughts: Okay, here's my interpretation of that thought process, assuming I'm an expert in this field, presented as a first-person summary:
**The Channel Assignment Revelation**
Alright, let me think this through... I've been focusing on the overall request handling, and now, finally, it clicks. The crucial part of how we manage the calls and channel assignments has been staring me in the face, yet, I missed it. The process is being driven by `rxrpc_activate_one_channel`. This function is the one assigning a call to a specific channel. That's the core of the connection between the incoming requests and the available resources! It seems obvious now. I can now proceed.
|
| 68/2 |
2026/03/10 00:53 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/conn_client.c IncludeLines:true Name:rxrpc_activate_one_channel]
Results: map[SourceCode: 407: /*
408: * Assign a channel to the call at the front of the queue and wake the call up.
409: * We don't increment the callNumber counter until this number has been exposed
410: * to the world.
411: */
412: static void rxrpc_activate_one_channel(struct rxrpc_connection *conn,
413: unsigned int channel)
414: {
415: struct rxrpc_channel *chan = &conn->channels[channel];
416: struct rxrpc_bundle *bundle = conn->bundle;
417: struct rxrpc_call *call = list_entry(bundle->waiting_calls.next,
418: struct rxrpc_call, wait_link);
419: u32 call_id = chan->call_counter + 1;
420:
421: _enter("C=%x,%u", conn->debug_id, channel);
422:
423: list_del_init(&call->wait_link);
424:
425: trace_rxrpc_client(conn, channel, rxrpc_client_chan_activate);
426:
427: /* Cancel the final ACK on the previous call if it hasn't been sent yet
428: * as the DATA packet will implicitly ACK it.
429: */
430: clear_bit(RXRPC_CONN_FINAL_ACK_0 + channel, &conn->flags);
431: clear_bit(conn->bundle_shift + channel, &bundle->avail_chans);
432:
433: rxrpc_see_call(call, rxrpc_call_see_activate_client);
434: call->conn = rxrpc_get_connection(conn, rxrpc_conn_get_activate_call);
435: call->cid = conn->proto.cid | channel;
436: call->call_id = call_id;
437: call->dest_srx.srx_service = conn->service_id;
438: call->cong_ssthresh = call->peer->cong_ssthresh;
439: if (call->cong_cwnd >= call->cong_ssthresh)
440: call->cong_ca_state = RXRPC_CA_CONGEST_AVOIDANCE;
441: else
442: call->cong_ca_state = RXRPC_CA_SLOW_START;
443:
444: chan->call_id = call_id;
445: chan->call_debug_id = call->debug_id;
446: chan->call = call;
447:
448: rxrpc_see_call(call, rxrpc_call_see_connected);
449: trace_rxrpc_connect_call(call);
450: call->tx_last_sent = ktime_get_real();
451: rxrpc_start_call_timer(call);
452: rxrpc_set_call_state(call, RXRPC_CALL_CLIENT_SEND_REQUEST);
453: wake_up(&call->waitq);
454: }
SourceFile:net/rxrpc/conn_client.c]
|
| 69/2 |
2026/03/10 00:53 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 68242
output: 47
thoughts: 51
Thoughts: Okay, here's my take on those thoughts, framed as my own mental processing:
**Digging into RXRPC Channel Activation and Call Management**
Right, so I'm looking at `rxrpc_activate_one_channel`. The key takeaway is that this function *assigns* the call to `chan->call`. It's a pretty critical operation. But the way I read this, it *doesn't* drop any reference count. That raises a red flag in my mind. Where and when is that crucial reference count actually *decremented* to free up the call structure? I'm betting that's handled elsewhere.
The natural next place to check would be `rxrpc_disconnect_client_call`. That function's primary role is in disconnecting a client call, so it's a likely candidate for reference management. Time to dive into that function and make sure that call deallocation is handled there. I will need to verify the expected reference count handling within `rxrpc_disconnect_client_call`. This is important, it will confirm my understanding of the reference count.
|
| 70/2 |
2026/03/10 00:53 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/conn_client.c IncludeLines:true Name:rxrpc_disconnect_client_call]
Results: map[SourceCode: 570: /*
571: * Disconnect a client call.
572: */
573: void rxrpc_disconnect_client_call(struct rxrpc_bundle *bundle, struct rxrpc_call *call)
574: {
575: struct rxrpc_connection *conn;
576: struct rxrpc_channel *chan = NULL;
577: struct rxrpc_local *local = bundle->local;
578: unsigned int channel;
579: bool may_reuse;
580: u32 cid;
581:
582: _enter("c=%x", call->debug_id);
583:
584: /* Calls that have never actually been assigned a channel can simply be
585: * discarded.
586: */
587: conn = call->conn;
588: if (!conn) {
589: _debug("call is waiting");
590: ASSERTCMP(call->call_id, ==, 0);
591: ASSERT(!test_bit(RXRPC_CALL_EXPOSED, &call->flags));
592: /* May still be on ->new_client_calls. */
593: spin_lock_irq(&local->client_call_lock);
594: list_del_init(&call->wait_link);
595: spin_unlock_irq(&local->client_call_lock);
596: return;
597: }
598:
599: cid = call->cid;
600: channel = cid & RXRPC_CHANNELMASK;
601: chan = &conn->channels[channel];
602: trace_rxrpc_client(conn, channel, rxrpc_client_chan_disconnect);
603:
604: if (WARN_ON(chan->call != call))
605: return;
606:
607: may_reuse = rxrpc_may_reuse_conn(conn);
608:
609: /* If a client call was exposed to the world, we save the result for
610: * retransmission.
611: *
612: * We use a barrier here so that the call number and abort code can be
613: * read without needing to take a lock.
614: *
615: * TODO: Make the incoming packet handler check this and handle
616: * terminal retransmission without requiring access to the call.
617: */
618: if (test_bit(RXRPC_CALL_EXPOSED, &call->flags)) {
619: _debug("exposed %u,%u", call->call_id, call->abort_code);
620: __rxrpc_disconnect_call(conn, call);
621:
622: if (test_and_clear_bit(RXRPC_CONN_PROBING_FOR_UPGRADE, &conn->flags)) {
623: trace_rxrpc_client(conn, channel, rxrpc_client_to_active);
624: bundle->try_upgrade = false;
625: if (may_reuse)
626: rxrpc_activate_channels(bundle);
627: }
628: }
629:
630: /* See if we can pass the channel directly to another call. */
631: if (may_reuse && !list_empty(&bundle->waiting_calls)) {
632: trace_rxrpc_client(conn, channel, rxrpc_client_chan_pass);
633: rxrpc_activate_one_channel(conn, channel);
634: return;
635: }
636:
637: /* Schedule the final ACK to be transmitted in a short while so that it
638: * can be skipped if we find a follow-on call. The first DATA packet
639: * of the follow on call will implicitly ACK this call.
640: */
641: if (call->completion == RXRPC_CALL_SUCCEEDED &&
642: test_bit(RXRPC_CALL_EXPOSED, &call->flags)) {
643: unsigned long final_ack_at = jiffies + 2;
644:
645: chan->final_ack_at = final_ack_at;
646: smp_wmb(); /* vs rxrpc_process_delayed_final_acks() */
647: set_bit(RXRPC_CONN_FINAL_ACK_0 + channel, &conn->flags);
648: rxrpc_reduce_conn_timer(conn, final_ack_at);
649: }
650:
651: /* Deactivate the channel. */
652: chan->call = NULL;
653: set_bit(conn->bundle_shift + channel, &conn->bundle->avail_chans);
654: conn->act_chans &= ~(1 << channel);
655:
656: /* If no channels remain active, then put the connection on the idle
657: * list for a short while. Give it a ref to stop it going away if it
658: * becomes unbundled.
659: */
660: if (!conn->act_chans) {
661: trace_rxrpc_client(conn, channel, rxrpc_client_to_idle);
662: conn->idle_timestamp = jiffies;
663:
664: rxrpc_get_connection(conn, rxrpc_conn_get_idle);
665: list_move_tail(&conn->cache_link, &local->idle_client_conns);
666:
667: rxrpc_set_client_reap_timer(local);
668: }
669: }
SourceFile:net/rxrpc/conn_client.c]
|
| 71/2 |
2026/03/10 00:53 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 70222
output: 57
thoughts: 71
Thoughts: **My Concerns Regarding Reference Counting in RXRPC**
Okay, hold on a second. I just realized something potentially problematic in the RXRPC code. I was under the impression that `rxrpc_disconnect_client_call` was responsible for dropping a reference, but I've just confirmed it *doesn't*. That's a red flag. Now, looking a bit closer, I remember that `rxrpc_connect_call` takes a reference using `rxrpc_call_get_io_thread`.
So, the critical question becomes: where is that initial reference actually *released*? This is a textbook example of a potential memory leak if we're not careful. I need to track down the corresponding release function, which I'm betting is either `rxrpc_call_put_io_thread` or a very similar function with a related name and semantic meaning. I need to confirm that this happens correctly in all the code paths where it's acquired. This warrants a deeper dive, now.
|
| 72/2 |
2026/03/10 00:53 |
tool |
codesearch-find-references |
0m
Error:
Args: map[ContextFile:net/rxrpc/call_object.c IncludeSnippetLines:1 Name:rxrpc_call_get_io_thread SourceTreePrefix:]
Results: map[References:<nil> TruncatedOutput:false]
|
| 73/2 |
2026/03/10 00:53 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 70582
output: 21
thoughts: 59
Thoughts: Okay, let's see...
**Deconstructing the `rxrpc` IO Thread Issue**
Hold on a second... `rxrpc_call_get_io_thread` is an *enum value*? That's a bit unexpected. I need to understand how this is used and where it's released. Given that, let's switch gears and go spelunking in `net/rxrpc/call_object.c`. I'll search for related functions, specifically `rxrpc_call_put_io_thread`, and even broader, anything that starts with `rxrpc_call_put_`, to get a complete picture of the associated resources management. That should clarify the intended flow and any potential areas of concern.
|
| 74/2 |
2026/03/10 00:53 |
tool |
grepper |
0m
Args: map[Expression:rxrpc_call_put_]
Results: map[Output:include/trace/events/rxrpc.h-317- EM(rxrpc_call_new_prealloc_service, "NEW prealloc") \
include/trace/events/rxrpc.h:318: EM(rxrpc_call_put_discard_prealloc, "PUT disc-pre") \
include/trace/events/rxrpc.h:319: EM(rxrpc_call_put_discard_error, "PUT disc-err") \
include/trace/events/rxrpc.h:320: EM(rxrpc_call_put_io_thread, "PUT iothread") \
include/trace/events/rxrpc.h:321: EM(rxrpc_call_put_input, "PUT input ") \
include/trace/events/rxrpc.h:322: EM(rxrpc_call_put_kernel, "PUT kernel ") \
include/trace/events/rxrpc.h:323: EM(rxrpc_call_put_poke, "PUT poke ") \
include/trace/events/rxrpc.h:324: EM(rxrpc_call_put_recvmsg, "PUT recvmsg ") \
include/trace/events/rxrpc.h:325: EM(rxrpc_call_put_recvmsg_peek_nowait, "PUT peek-nwt") \
include/trace/events/rxrpc.h:326: EM(rxrpc_call_put_release_recvmsg_q, "PUT rls-rcmq") \
include/trace/events/rxrpc.h:327: EM(rxrpc_call_put_release_sock, "PUT rls-sock") \
include/trace/events/rxrpc.h:328: EM(rxrpc_call_put_release_sock_tba, "PUT rls-sk-a") \
include/trace/events/rxrpc.h:329: EM(rxrpc_call_put_sendmsg, "PUT sendmsg ") \
include/trace/events/rxrpc.h:330: EM(rxrpc_call_put_userid_exists, "PUT u-exists") \
include/trace/events/rxrpc.h:331: EM(rxrpc_call_put_userid, "PUT user-id ") \
include/trace/events/rxrpc.h-332- EM(rxrpc_call_see_accept, "SEE accept ") \
--
net/rxrpc/af_rxrpc.c=436=void rxrpc_kernel_put_call(struct socket *sock, struct rxrpc_call *call)
net/rxrpc/af_rxrpc.c-437-{
net/rxrpc/af_rxrpc.c:438: rxrpc_put_call(call, rxrpc_call_put_kernel);
net/rxrpc/af_rxrpc.c-439-}
--
net/rxrpc/call_accept.c=179=void rxrpc_discard_prealloc(struct rxrpc_sock *rx)
--
net/rxrpc/call_accept.c-229- call->notify_rx = rxrpc_dummy_notify;
net/rxrpc/call_accept.c:230: rxrpc_put_call(call, rxrpc_call_put_kernel);
net/rxrpc/call_accept.c-231- }
--
net/rxrpc/call_accept.c-233- rxrpc_release_call(rx, call);
net/rxrpc/call_accept.c:234: rxrpc_put_call(call, rxrpc_call_put_discard_prealloc);
net/rxrpc/call_accept.c-235- tail = (tail + 1) & (size - 1);
--
net/rxrpc/call_accept.c=335=bool rxrpc_new_incoming_call(struct rxrpc_local *local,
--
net/rxrpc/call_accept.c-419- rxrpc_queue_rx_call_packet(call, skb);
net/rxrpc/call_accept.c:420: rxrpc_put_call(call, rxrpc_call_put_input);
net/rxrpc/call_accept.c-421- return true;
--
net/rxrpc/call_object.c=188=static struct rxrpc_call *rxrpc_alloc_client_call(struct rxrpc_sock *rx,
--
net/rxrpc/call_object.c-230- rxrpc_prefail_call(call, RXRPC_CALL_LOCAL_ERROR, ret);
net/rxrpc/call_object.c:231: rxrpc_put_call(call, rxrpc_call_put_discard_error);
net/rxrpc/call_object.c-232- return ERR_PTR(ret);
--
net/rxrpc/call_object.c=320=struct rxrpc_call *rxrpc_new_client_call(struct rxrpc_sock *rx,
--
net/rxrpc/call_object.c-416- mutex_unlock(&call->user_mutex);
net/rxrpc/call_object.c:417: rxrpc_put_call(call, rxrpc_call_put_userid_exists);
net/rxrpc/call_object.c-418- _leave(" = -EEXIST");
--
net/rxrpc/call_object.c=561=void rxrpc_release_call(struct rxrpc_sock *rx, struct rxrpc_call *call)
--
net/rxrpc/call_object.c-597- if (putu)
net/rxrpc/call_object.c:598: rxrpc_put_call(call, rxrpc_call_put_userid);
net/rxrpc/call_object.c-599-
--
net/rxrpc/call_object.c=606=void rxrpc_release_calls_on_socket(struct rxrpc_sock *rx)
--
net/rxrpc/call_object.c-617- rxrpc_abort_call_sock_release_tba);
net/rxrpc/call_object.c:618: rxrpc_put_call(call, rxrpc_call_put_release_sock_tba);
net/rxrpc/call_object.c-619- }
--
net/rxrpc/call_object.c-627- rxrpc_release_call(rx, call);
net/rxrpc/call_object.c:628: rxrpc_put_call(call, rxrpc_call_put_release_sock);
net/rxrpc/call_object.c-629- }
--
net/rxrpc/call_object.c-633- list_del_init(&call->recvmsg_link);
net/rxrpc/call_object.c:634: rxrpc_put_call(call, rxrpc_call_put_release_recvmsg_q);
net/rxrpc/call_object.c-635- }
--
net/rxrpc/conn_object.c=191=void rxrpc_disconnect_call(struct rxrpc_call *call)
--
net/rxrpc/conn_object.c-215-
net/rxrpc/conn_object.c:216: rxrpc_put_call(call, rxrpc_call_put_io_thread);
net/rxrpc/conn_object.c-217-}
--
net/rxrpc/io_thread.c=347=static int rxrpc_input_packet_on_conn(struct rxrpc_connection *conn,
--
net/rxrpc/io_thread.c-421- if (rxrpc_to_client(sp)) {
net/rxrpc/io_thread.c:422: rxrpc_put_call(call, rxrpc_call_put_input);
net/rxrpc/io_thread.c-423- return rxrpc_protocol_error(skb,
--
net/rxrpc/io_thread.c-428- rxrpc_implicit_end_call(call, skb);
net/rxrpc/io_thread.c:429: rxrpc_put_call(call, rxrpc_call_put_input);
net/rxrpc/io_thread.c-430- call = NULL;
--
net/rxrpc/io_thread.c-441- rxrpc_queue_rx_call_packet(call, skb);
net/rxrpc/io_thread.c:442: rxrpc_put_call(call, rxrpc_call_put_input);
net/rxrpc/io_thread.c-443- return true;
--
net/rxrpc/io_thread.c=449=int rxrpc_io_thread(void *data)
--
net/rxrpc/io_thread.c-550- rxrpc_input_call_event(call);
net/rxrpc/io_thread.c:551: rxrpc_put_call(call, rxrpc_call_put_poke);
net/rxrpc/io_thread.c-552- }
--
net/rxrpc/recvmsg.c=366=int rxrpc_recvmsg(struct socket *sock, struct msghdr *msg, size_t len,
--
net/rxrpc/recvmsg.c-449- trace_rxrpc_recvmsg(call->debug_id, rxrpc_recvmsg_unqueue, 0);
net/rxrpc/recvmsg.c:450: rxrpc_put_call(call, rxrpc_call_put_recvmsg);
net/rxrpc/recvmsg.c-451- goto try_again;
--
net/rxrpc/recvmsg.c-460- trace_rxrpc_recvmsg(call->debug_id, rxrpc_recvmsg_unqueue, 0);
net/rxrpc/recvmsg.c:461: rxrpc_put_call(call, rxrpc_call_put_recvmsg);
net/rxrpc/recvmsg.c-462- goto try_again;
--
net/rxrpc/recvmsg.c-490- if (!(flags & MSG_PEEK))
net/rxrpc/recvmsg.c:491: rxrpc_put_call(call, rxrpc_call_put_recvmsg);
net/rxrpc/recvmsg.c-492- goto try_again;
--
net/rxrpc/recvmsg.c-545- mutex_unlock(&call->user_mutex);
net/rxrpc/recvmsg.c:546: rxrpc_put_call(call, rxrpc_call_put_recvmsg);
net/rxrpc/recvmsg.c-547- trace_rxrpc_recvmsg(call_debug_id, rxrpc_recvmsg_return, ret);
--
net/rxrpc/recvmsg.c-566- } else {
net/rxrpc/recvmsg.c:567: rxrpc_put_call(call, rxrpc_call_put_recvmsg_peek_nowait);
net/rxrpc/recvmsg.c-568- }
--
net/rxrpc/sendmsg.c=659=int rxrpc_do_sendmsg(struct rxrpc_sock *rx, struct msghdr *msg, size_t len)
--
net/rxrpc/sendmsg.c-712- case RXRPC_CALL_SERVER_PREALLOC:
net/rxrpc/sendmsg.c:713: rxrpc_put_call(call, rxrpc_call_put_sendmsg);
net/rxrpc/sendmsg.c-714- ret = -EBUSY;
--
net/rxrpc/sendmsg.c-781-error_put:
net/rxrpc/sendmsg.c:782: rxrpc_put_call(call, rxrpc_call_put_sendmsg);
net/rxrpc/sendmsg.c-783- _leave(" = %d", ret);
]
|
| 75/2 |
2026/03/10 00:53 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 73842
output: 45
thoughts: 46
Thoughts: Okay, here's my take on those thoughts, framed as an expert's internal monologue:
**Dissecting the Disconnect Process**
Right, so `rxrpc_disconnect_call` is the key function here. Immediately, my mental map flags it: it calls `rxrpc_put_call` with a specific argument, `rxrpc_call_put_io_thread`. Okay, that's interesting. I know that `rxrpc_put_call` is responsible for… well, putting the call object. But the specific "put" operation in this context is what piques my interest. It strongly suggests a particular interaction with the I/O thread, likely dealing with freeing resources or signaling something. Time to dig into `rxrpc_disconnect_call` and understand the implications of this specific `rxrpc_call_put_io_thread` usage. What state is being cleaned up? Is it a graceful shutdown? A forced termination? My instinct says this is likely linked to handling the asynchronous aspects of the connection termination.
|
| 76/2 |
2026/03/10 00:54 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/conn_object.c IncludeLines:true Name:rxrpc_disconnect_call]
Results: map[SourceCode: 187: /*
188: * Disconnect a call and clear any channel it occupies when that call
189: * terminates.
190: */
191: void rxrpc_disconnect_call(struct rxrpc_call *call)
192: {
193: struct rxrpc_connection *conn = call->conn;
194:
195: set_bit(RXRPC_CALL_DISCONNECTED, &call->flags);
196: rxrpc_see_call(call, rxrpc_call_see_disconnected);
197:
198: call->peer->cong_ssthresh = call->cong_ssthresh;
199:
200: if (!hlist_unhashed(&call->error_link)) {
201: spin_lock_irq(&call->peer->lock);
202: hlist_del_init(&call->error_link);
203: spin_unlock_irq(&call->peer->lock);
204: }
205:
206: if (rxrpc_is_client_call(call)) {
207: rxrpc_disconnect_client_call(call->bundle, call);
208: } else {
209: __rxrpc_disconnect_call(conn, call);
210: conn->idle_timestamp = jiffies;
211: if (atomic_dec_and_test(&conn->active))
212: rxrpc_set_service_reap_timer(conn->rxnet,
213: jiffies + rxrpc_connection_expiry * HZ);
214: }
215:
216: rxrpc_put_call(call, rxrpc_call_put_io_thread);
217: }
SourceFile:net/rxrpc/conn_object.c]
|
| 77/2 |
2026/03/10 00:54 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 74681
output: 45
thoughts: 59
Thoughts: Here's my thought process, as if I were analyzing this code:
**Dissecting the Call Disconnect Procedure**
Okay, so I see `rxrpc_disconnect_call` is calling `rxrpc_put_call` with `rxrpc_call_put_io_thread`. That's interesting, it seems we're releasing the call after disconnection, and it's being done from the I/O thread, which makes sense. I need to understand when this `rxrpc_disconnect_call` function is actually triggered. The critical path of the call release is here.
To answer that, I should look at where `rxrpc_disconnect_call` is used. A likely place to start would be `rxrpc_destroy_call`. It could be happening when the call object is cleaned up.
|
| 78/2 |
2026/03/10 00:54 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_object.c IncludeLines:true Name:rxrpc_destroy_call]
Results: map[SourceCode: 680: /*
681: * Final call destruction - but must be done in process context.
682: */
683: static void rxrpc_destroy_call(struct work_struct *work)
684: {
685: struct rxrpc_call *call = container_of(work, struct rxrpc_call, destroyer);
686:
687: timer_delete_sync(&call->timer);
688:
689: rxrpc_cleanup_tx_buffers(call);
690: rxrpc_cleanup_rx_buffers(call);
691: rxrpc_put_txbuf(call->tx_pending, rxrpc_txbuf_put_cleaned);
692: rxrpc_put_connection(call->conn, rxrpc_conn_put_call);
693: rxrpc_deactivate_bundle(call->bundle);
694: rxrpc_put_bundle(call->bundle, rxrpc_bundle_put_call);
695: rxrpc_put_peer(call->peer, rxrpc_peer_put_call);
696: rxrpc_put_local(call->local, rxrpc_local_put_call);
697: call_rcu(&call->rcu, rxrpc_rcu_free_call);
698: }
SourceFile:net/rxrpc/call_object.c]
|
| 79/2 |
2026/03/10 00:54 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 75348
output: 53
thoughts: 34
Thoughts: Okay, here's my take on those thoughts, framed as my own internal monologue:
**Hmm, a Potential Oversight...**
Hold on a second... just revisiting the `rxrpc_destroy_call` function... I'm a bit concerned that it doesn't seem to call `rxrpc_disconnect_call`. That feels like a potential gap in the cleanup process. We should be sure the connection is fully terminated.
So, the question is, *who* is responsible for calling `rxrpc_disconnect_call`? Is there another function that handles the disconnect phase, and is it always invoked before the call is destroyed? Or is this a subtle bug waiting to bite us? We need to figure out the disconnect side of things before destruction so we don't accidentally leave dangling resources or cause unexpected behaviour. We have a gap to close.
|
| 80/2 |
2026/03/10 00:54 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/rxrpc/conn_object.c IncludeSnippetLines:1 Name:rxrpc_disconnect_call SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rxrpc_input_call_event SourceFile:net/rxrpc/call_event.c SourceLine:474 SourceSnippet: 473: if (!test_bit(RXRPC_CALL_DISCONNECTED, &call->flags))
474: rxrpc_disconnect_call(call);
475: if (call->security)
]] TruncatedOutput:false]
|
| 81/2 |
2026/03/10 00:54 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 75750
output: 46
thoughts: 403
Thoughts: **The Persistent Call Conundrum**
Okay, let's break this down. My initial thought was that `rxrpc_input_call_event` calls `rxrpc_disconnect_call`. That's straightforward, and `rxrpc_input_call_event` is called by `rxrpc_io_thread`, which makes sense. I see that `rxrpc_input_call_event` is triggered when the call is complete, which leads to `rxrpc_disconnect_call` releasing the `rxrpc_call_get_io_thread` reference. But then I started to think, *what if* the call *never* completes?
If it's never complete, that reference to the IO thread is *never* dropped. That means the call is never freed, and if the call is never freed, `rxnet->nr_calls` will never reach zero. This is where it gets nasty because if `rxnet->nr_calls` is never zero, then `rxrpc_destroy_all_calls` gets stuck in a loop.
So the core of the problem is that `rxrpc_destroy_all_calls` is waiting for calls to be freed, but some calls are stuck because they're holding that `rxrpc_call_get_io_thread` reference and never completing. This is crucial because `rxrpc_destroy_all_calls` is called from `rxrpc_exit_net`. I need to remember that when `rxrpc_exit_net` is called, all sockets in the namespace should have already been closed. And when a socket closes, `rxrpc_release_calls_on_socket` is called.
Aha! `rxrpc_release_calls_on_socket` calls `rxrpc_propose_abort(call, RX_CALL_DEAD, -ECONNRESET, rxrpc_abort_call_sock_release)`. `rxrpc_propose_abort` *aborts* the call! If it's aborted, the call becomes complete! This sounds promising. Now I need to take a closer look at `rxrpc_propose_abort` to confirm my suspicions and see if this resolves the deadlock issue I've identified.
|
| 82/2 |
2026/03/10 00:54 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_event.c IncludeLines:true Name:rxrpc_propose_abort]
Results: map[SourceCode: 20: /*
21: * Propose an abort to be made in the I/O thread.
22: */
23: bool rxrpc_propose_abort(struct rxrpc_call *call, s32 abort_code, int error,
24: enum rxrpc_abort_reason why)
25: {
26: _enter("{%d},%d,%d,%u", call->debug_id, abort_code, error, why);
27:
28: if (!call->send_abort && !rxrpc_call_is_complete(call)) {
29: call->send_abort_why = why;
30: call->send_abort_err = error;
31: call->send_abort_seq = 0;
32: trace_rxrpc_abort_call(call, abort_code);
33: /* Request abort locklessly vs rxrpc_input_call_event(). */
34: smp_store_release(&call->send_abort, abort_code);
35: rxrpc_poke_call(call, rxrpc_call_poke_abort);
36: return true;
37: }
38:
39: return false;
40: }
SourceFile:net/rxrpc/sendmsg.c]
|
| 83/2 |
2026/03/10 00:54 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 77036
output: 45
thoughts: 396
Thoughts: **My Line of Reasoning on RxRPC Call Abort**
Okay, let's break this down. So, `rxrpc_propose_abort` sets `call->send_abort` and triggers `rxrpc_poke_call`. `rxrpc_poke_call` adds this call to the `local->call_attend_q` and wakes up the I/O thread. The I/O thread, in turn, calls `rxrpc_input_call_event`. Inside `rxrpc_input_call_event`, if `call->send_abort` is set, the call should be aborted. That seems straightforward.
But, hang on, what if the I/O thread is stopped? If `rxrpc_input_call_event` *isn't* called, the abort won't happen! Okay, when does the I/O thread stop? It's stopped when `rxrpc_local` is destroyed. And when does *that* happen? When its reference count hits zero. But a `rxrpc_call` *holds* a reference to `rxrpc_local`! So, `rxrpc_local` won't be destroyed until all the calls are freed!
Therefore, assuming the I/O thread is still running and not blocked, the I/O thread *should* process the abort and complete the call. But, what if it *is* blocked? Let me see... I've checked the stack traces. There's no sign of the I/O thread there, which is odd. This makes me consider if the I/O thread is NOT even running?
The next question: Why would the I/O thread not be running? Ah, got it! It's because `rxrpc_exit_net` calls `rxrpc_destroy_all_calls` *before* `rxrpc_destroy_all_locals`! That's it! I'll need to check the `rxrpc_exit_net` routine to confirm, but I think I've found the issue.
|
| 84/2 |
2026/03/10 00:54 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/net_ns.c IncludeLines:true Name:rxrpc_exit_net]
Results: map[SourceCode: 100: /*
101: * Clean up a per-network namespace record.
102: */
103: static __net_exit void rxrpc_exit_net(struct net *net)
104: {
105: struct rxrpc_net *rxnet = rxrpc_net(net);
106:
107: rxnet->live = false;
108: timer_delete_sync(&rxnet->peer_keepalive_timer);
109: cancel_work_sync(&rxnet->peer_keepalive_work);
110: /* Remove the timer again as the worker may have restarted it. */
111: timer_delete_sync(&rxnet->peer_keepalive_timer);
112: rxrpc_destroy_all_calls(rxnet);
113: rxrpc_destroy_all_connections(rxnet);
114: rxrpc_destroy_all_peers(rxnet);
115: rxrpc_destroy_all_locals(rxnet);
116: proc_remove(rxnet->proc_net);
117: }
SourceFile:net/rxrpc/net_ns.c]
|
| 85/2 |
2026/03/10 00:54 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 78249
output: 57
thoughts: 1200
Thoughts: **Debugging a Potential RXRPC Hang**
Okay, first things first, `rxrpc_destroy_all_calls` *is* called *before* `rxrpc_destroy_all_locals`. This confirms the initial understanding, so the I/O thread *should* still be running during the `rxrpc_destroy_all_calls` execution. But *why* is `rxrpc_destroy_all_calls` hanging? Let's trace it.
Right, `rxrpc_destroy_all_calls` iterates through `rxnet->calls`:
```c
if (!list_empty(&rxnet->calls)) {
spin_lock(&rxnet->call_lock);
while (!list_empty(&rxnet->calls)) {
call = list_entry(rxnet->calls.next,
struct rxrpc_call, link);
_debug("Zapping call %p", call);
rxrpc_see_call(call, rxrpc_call_see_zap);
list_del_init(&call->link);
pr_err("Call %p still in use (%d,%s,%lx,%lx)!\n",
call, refcount_read(&call->ref),
rxrpc_call_states[__rxrpc_call_state(call)],
call->flags, call->events);
spin_unlock(&rxnet->call_lock);
cond_resched();
spin_lock(&rxnet->call_lock);
}
spin_unlock(&rxnet->call_lock);
}
```
Wait a second, there's a print statement "Call %p still in use"! Is this message in the kernel log? It's not in the bug report, but it might be truncated. Okay, `rxrpc_destroy_all_calls` *removes* the call from the `rxnet->calls` list, but it *doesn't* abort the call. It just prints an error! If it's still "in use", *who* is using it?
Ah, here's a crucial comment: "To reach this point, any open UDP sockets in that namespace must have been closed, so any outstanding calls cannot be doing I/O." Wait, are the UDP sockets closed?
The sockets are created by `rxrpc_open_socket` in `rxrpc_lookup_local` and closed by `rxrpc_destroy_local`. But `rxrpc_destroy_all_locals` is called *after* `rxrpc_destroy_all_calls`! So the UDP sockets are *not* closed yet! The comment is wrong!
So, the comment says "any open UDP sockets in that namespace must have been closed" and the code *currently* closes sockets at a point that is logically incorrect! `rxrpc_destroy_all_locals` is called AFTER `rxrpc_destroy_all_calls`!
Okay, the UDP sockets are closed when the `rxrpc_local` struct is destroyed. When is that? When its reference count reaches 0. And when does *that* happen? When all calls and connections using it are freed! `rxrpc_destroy_all_calls` *waits* for the calls to be freed. This is the crux.
If a call is *not* freed, the `rxrpc_local` struct isn't destroyed, meaning the UDP socket *isn't* closed. If the UDP socket isn't closed, the I/O thread *could* still be running. If the I/O thread *is* still running, why isn't the call completing? The call is not aborted! Why isn't the call aborted? Because `rxrpc_release_calls_on_socket` was *never* called.
Why was `rxrpc_release_calls_on_socket` never called? Because the AF_RXRPC socket was never closed!
Wait, if the AF_RXRPC socket *was* never closed, how can `rxrpc_exit_net` be running at all? Maybe the call isn't associated with the socket? Can a call be created without an AF_RXRPC socket? No, `rxrpc_alloc_call` requires `struct rxrpc_sock *rx`.
Okay, what if the AF_RXRPC socket *is* closed, but the call is still alive? If the socket *is* closed, `rxrpc_release_calls_on_socket` *is* called. It proposes an abort with `rxrpc_propose_abort(call, RX_CALL_DEAD, -ECONNRESET, rxrpc_abort_call_sock_release)` and then calls `rxrpc_
|
| 86/2 |
2026/03/10 00:54 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/rxrpc/call_object.c IncludeSnippetLines:1 Name:rxrpc_release_calls_on_socket SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:calls ReferencingEntityKind:function ReferencingEntityName:rxrpc_release_sock SourceFile:net/rxrpc/af_rxrpc.c SourceLine:964 SourceSnippet: 963: rxrpc_discard_prealloc(rx);
964: rxrpc_release_calls_on_socket(rx);
965: flush_workqueue(rxrpc_workqueue);
]] TruncatedOutput:false]
|
| 87/2 |
2026/03/10 00:54 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 80666
output: 47
thoughts: 47
Thoughts: Okay, let me think this through.
**Analyzing Client Call Tracking**
Hmm, so I'm looking at this code and something's nagging at me. It's about this `rx->sock_calls` data structure. The usual assumption is that all the client calls are in this structure, right? But wait, what if that's not *always* true? What if there are scenarios where a call isn't there, and if that's the case, I need to know why.
So, when would a client call *not* be in `rx->sock_calls`? I need to understand the lifecycle of a call more deeply to answer that. I'll need to trace the call's creation and potential removal from the tracking structure.
Okay, the first thing that comes to mind is that it could be related to `rxrpc_new_client_call`. Let me go check that function and its interactions. Understanding that should help me work through the call creation process, figure out what states a client call can be in, and find potential paths to figure out *when* it's not. I'll need to go through it carefully.
|
| 88/2 |
2026/03/10 00:54 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_object.c IncludeLines:true Name:rxrpc_new_client_call]
Results: map[SourceCode: 315: /*
316: * Set up a call for the given parameters.
317: * - Called with the socket lock held, which it must release.
318: * - If it returns a call, the call's lock will need releasing by the caller.
319: */
320: struct rxrpc_call *rxrpc_new_client_call(struct rxrpc_sock *rx,
321: struct rxrpc_conn_parameters *cp,
322: struct rxrpc_call_params *p,
323: gfp_t gfp,
324: unsigned int debug_id)
325: __releases(&rx->sk.sk_lock)
326: __acquires(&call->user_mutex)
327: {
328: struct rxrpc_call *call, *xcall;
329: struct rxrpc_net *rxnet;
330: struct semaphore *limiter;
331: struct rb_node *parent, **pp;
332: int ret;
333:
334: _enter("%p,%lx", rx, p->user_call_ID);
335:
336: if (WARN_ON_ONCE(!cp->peer)) {
337: release_sock(&rx->sk);
338: return ERR_PTR(-EIO);
339: }
340:
341: limiter = rxrpc_get_call_slot(p, gfp);
342: if (!limiter) {
343: release_sock(&rx->sk);
344: return ERR_PTR(-ERESTARTSYS);
345: }
346:
347: call = rxrpc_alloc_client_call(rx, cp, p, gfp, debug_id);
348: if (IS_ERR(call)) {
349: release_sock(&rx->sk);
350: up(limiter);
351: _leave(" = %ld", PTR_ERR(call));
352: return call;
353: }
354:
355: /* We need to protect a partially set up call against the user as we
356: * will be acting outside the socket lock.
357: */
358: mutex_lock(&call->user_mutex);
359:
360: /* Publish the call, even though it is incompletely set up as yet */
361: write_lock(&rx->call_lock);
362:
363: pp = &rx->calls.rb_node;
364: parent = NULL;
365: while (*pp) {
366: parent = *pp;
367: xcall = rb_entry(parent, struct rxrpc_call, sock_node);
368:
369: if (p->user_call_ID < xcall->user_call_ID)
370: pp = &(*pp)->rb_left;
371: else if (p->user_call_ID > xcall->user_call_ID)
372: pp = &(*pp)->rb_right;
373: else
374: goto error_dup_user_ID;
375: }
376:
377: rcu_assign_pointer(call->socket, rx);
378: call->user_call_ID = p->user_call_ID;
379: __set_bit(RXRPC_CALL_HAS_USERID, &call->flags);
380: rxrpc_get_call(call, rxrpc_call_get_userid);
381: rb_link_node(&call->sock_node, parent, pp);
382: rb_insert_color(&call->sock_node, &rx->calls);
383: list_add(&call->sock_link, &rx->sock_calls);
384:
385: write_unlock(&rx->call_lock);
386:
387: rxnet = call->rxnet;
388: spin_lock(&rxnet->call_lock);
389: list_add_tail_rcu(&call->link, &rxnet->calls);
390: spin_unlock(&rxnet->call_lock);
391:
392: /* From this point on, the call is protected by its own lock. */
393: release_sock(&rx->sk);
394:
395: /* Set up or get a connection record and set the protocol parameters,
396: * including channel number and call ID.
397: */
398: ret = rxrpc_connect_call(call, gfp);
399: if (ret < 0)
400: goto error_attached_to_socket;
401:
402: _leave(" = %p [new]", call);
403: return call;
404:
405: /* We unexpectedly found the user ID in the list after taking
406: * the call_lock. This shouldn't happen unless the user races
407: * with itself and tries to add the same user ID twice at the
408: * same time in different threads.
409: */
410: error_dup_user_ID:
411: write_unlock(&rx->call_lock);
412: release_sock(&rx->sk);
413: rxrpc_prefail_call(call, RXRPC_CALL_LOCAL_ERROR, -EEXIST);
414: trace_rxrpc_call(call->debug_id, refcount_read(&call->ref), 0,
415: rxrpc_call_see_userid_exists);
416: mutex_unlock(&call->user_mutex);
417: rxrpc_put_call(call, rxrpc_call_put_userid_exists);
418: _leave(" = -EEXIST");
419: return ERR_PTR(-EEXIST);
420:
421: /* We got an error, but the call is attached to the socket and is in
422: * need of release. However, we might now race with recvmsg() when it
423: * completion notifies the socket. Return 0 from sys_sendmsg() and
424: * leave the error to recvmsg() to deal with.
425: */
426: error_attached_to_socket:
427: trace_rxrpc_call(call->debug_id, refcount_read(&call->ref), ret,
428: rxrpc_call_see_connect_failed);
429: rxrpc_set_call_completion(call, RXRPC_CALL_LOCAL_ERROR, 0, ret);
430: _leave(" = c=%08x [err]", call->debug_id);
431: return call;
432: }
SourceFile:net/rxrpc/call_object.c]
|
| 89/2 |
2026/03/10 00:54 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 82962
output: 47
thoughts: 280
Thoughts: **My Analysis of Potential Call Leaks in RXRPC**
Okay, let me think this through. If `rxrpc_connect_call` fails, that sets things up to go to `error_attached_to_socket`. Then, the completion state of the call gets set to `RXRPC_CALL_LOCAL_ERROR`, and, importantly, the function *returns* the `call` pointer. Then, `rxrpc_do_sendmsg` gets involved and calls `rxrpc_put_call(call, rxrpc_call_put_sendmsg)`. So, the call is still hanging around in both `rx->sock_calls` and `rxnet->calls`. But, the call is considered complete, meaning `recvmsg` *should* pick it up eventually and free it. If `recvmsg` never gets called, `rxrpc_release_calls_on_socket` *should* clean it up, so it *shouldn't* leak. That seems right.
But, hang on... what if `rxrpc_new_client_call` succeeds, but then something goes wrong in `rxrpc_send_data`? If `rxrpc_send_data` returns an error, it passes the error back up. The `rxrpc_do_sendmsg` function will then call `rxrpc_put_call`, and, again, the `call` is left in `rx->sock_calls`. If nobody does a `recvmsg` call, `rxrpc_release_calls_on_socket` will get called eventually, so again, no leak.
Hold on, I also need to think about the *server* side of this. What if this is a server call? I should really check `rxrpc_new_incoming_call` too. It's important to consider all the cases.
|
| 90/2 |
2026/03/10 00:54 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_accept.c IncludeLines:true Name:rxrpc_new_incoming_call]
Results: map[SourceCode: 324: /*
325: * Set up a new incoming call. Called from the I/O thread.
326: *
327: * If this is for a kernel service, when we allocate the call, it will have
328: * three refs on it: (1) the kernel service, (2) the user_call_ID tree, (3) the
329: * retainer ref obtained from the backlog buffer. Prealloc calls for userspace
330: * services only have the ref from the backlog buffer.
331: *
332: * If we want to report an error, we mark the skb with the packet type and
333: * abort code and return false.
334: */
335: bool rxrpc_new_incoming_call(struct rxrpc_local *local,
336: struct rxrpc_peer *peer,
337: struct rxrpc_connection *conn,
338: struct sockaddr_rxrpc *peer_srx,
339: struct sk_buff *skb)
340: {
341: const struct rxrpc_security *sec = NULL;
342: struct rxrpc_skb_priv *sp = rxrpc_skb(skb);
343: struct rxrpc_call *call = NULL;
344: struct rxrpc_sock *rx;
345:
346: _enter("");
347:
348: /* Don't set up a call for anything other than a DATA packet. */
349: if (sp->hdr.type != RXRPC_PACKET_TYPE_DATA)
350: return rxrpc_protocol_error(skb, rxrpc_eproto_no_service_call);
351:
352: read_lock_irq(&local->services_lock);
353:
354: /* Weed out packets to services we're not offering. Packets that would
355: * begin a call are explicitly rejected and the rest are just
356: * discarded.
357: */
358: rx = local->service;
359: if (!rx || (sp->hdr.serviceId != rx->srx.srx_service &&
360: sp->hdr.serviceId != rx->second_service)
361: ) {
362: if (sp->hdr.type == RXRPC_PACKET_TYPE_DATA &&
363: sp->hdr.seq == 1)
364: goto unsupported_service;
365: goto discard;
366: }
367:
368: if (!conn) {
369: sec = rxrpc_get_incoming_security(rx, skb);
370: if (!sec)
371: goto unsupported_security;
372: }
373:
374: spin_lock(&rx->incoming_lock);
375: if (rx->sk.sk_state == RXRPC_SERVER_LISTEN_DISABLED ||
376: rx->sk.sk_state == RXRPC_CLOSE) {
377: rxrpc_direct_conn_abort(skb, rxrpc_abort_shut_down,
378: RX_INVALID_OPERATION, -ESHUTDOWN);
379: goto no_call;
380: }
381:
382: call = rxrpc_alloc_incoming_call(rx, local, peer, conn, sec, peer_srx,
383: skb);
384: if (!call) {
385: skb->mark = RXRPC_SKB_MARK_REJECT_BUSY;
386: goto no_call;
387: }
388:
389: trace_rxrpc_receive(call, rxrpc_receive_incoming,
390: sp->hdr.serial, sp->hdr.seq);
391:
392: /* Make the call live. */
393: rxrpc_incoming_call(rx, call, skb);
394: conn = call->conn;
395:
396: if (rx->app_ops &&
397: rx->app_ops->notify_new_call)
398: rx->app_ops->notify_new_call(&rx->sk, call, call->user_call_ID);
399:
400: spin_lock(&conn->state_lock);
401: if (conn->state == RXRPC_CONN_SERVICE_UNSECURED) {
402: conn->state = RXRPC_CONN_SERVICE_CHALLENGING;
403: set_bit(RXRPC_CONN_EV_CHALLENGE, &call->conn->events);
404: rxrpc_queue_conn(call->conn, rxrpc_conn_queue_challenge);
405: }
406: spin_unlock(&conn->state_lock);
407:
408: spin_unlock(&rx->incoming_lock);
409: read_unlock_irq(&local->services_lock);
410: rxrpc_assess_MTU_size(local, call->peer);
411:
412: if (hlist_unhashed(&call->error_link)) {
413: spin_lock_irq(&call->peer->lock);
414: hlist_add_head(&call->error_link, &call->peer->error_targets);
415: spin_unlock_irq(&call->peer->lock);
416: }
417:
418: _leave(" = %p{%d}", call, call->debug_id);
419: rxrpc_queue_rx_call_packet(call, skb);
420: rxrpc_put_call(call, rxrpc_call_put_input);
421: return true;
422:
423: unsupported_service:
424: read_unlock_irq(&local->services_lock);
425: return rxrpc_direct_conn_abort(skb, rxrpc_abort_service_not_offered,
426: RX_INVALID_OPERATION, -EOPNOTSUPP);
427: unsupported_security:
428: read_unlock_irq(&local->services_lock);
429: return rxrpc_direct_conn_abort(skb, rxrpc_abort_service_not_offered,
430: RX_INVALID_OPERATION, -EKEYREJECTED);
431: no_call:
432: spin_unlock(&rx->incoming_lock);
433: read_unlock_irq(&local->services_lock);
434: _leave(" = f [%u]", skb->mark);
435: return false;
436: discard:
437: read_unlock_irq(&local->services_lock);
438: return true;
439: }
SourceFile:net/rxrpc/call_accept.c]
|
| 91/2 |
2026/03/10 00:54 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 85648
output: 47
thoughts: 74
Thoughts: **Deconstructing the Incoming Call Flow**
Okay, so `rxrpc_new_incoming_call` kicks things off. First, it calls `rxrpc_alloc_incoming_call` to get a call structure, right? Then we move on to the actual handling with `rxrpc_incoming_call`. Finally, the input side gets put down via `rxrpc_put_call(call, rxrpc_call_put_input)`. That last step is interesting, so let's dig into `rxrpc_alloc_incoming_call` and see what's really happening. I want to understand what's being setup at each point.
|
| 92/2 |
2026/03/10 00:55 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_accept.c IncludeLines:true Name:rxrpc_alloc_incoming_call]
Results: map[SourceCode: 241: /*
242: * Allocate a new incoming call from the prealloc pool, along with a connection
243: * and a peer as necessary.
244: */
245: static struct rxrpc_call *rxrpc_alloc_incoming_call(struct rxrpc_sock *rx,
246: struct rxrpc_local *local,
247: struct rxrpc_peer *peer,
248: struct rxrpc_connection *conn,
249: const struct rxrpc_security *sec,
250: struct sockaddr_rxrpc *peer_srx,
251: struct sk_buff *skb)
252: {
253: struct rxrpc_backlog *b = rx->backlog;
254: struct rxrpc_call *call;
255: unsigned short call_head, conn_head, peer_head;
256: unsigned short call_tail, conn_tail, peer_tail;
257: unsigned short call_count, conn_count;
258:
259: if (!b)
260: return NULL;
261:
262: /* #calls >= #conns >= #peers must hold true. */
263: call_head = smp_load_acquire(&b->call_backlog_head);
264: call_tail = b->call_backlog_tail;
265: call_count = CIRC_CNT(call_head, call_tail, RXRPC_BACKLOG_MAX);
266: conn_head = smp_load_acquire(&b->conn_backlog_head);
267: conn_tail = b->conn_backlog_tail;
268: conn_count = CIRC_CNT(conn_head, conn_tail, RXRPC_BACKLOG_MAX);
269: ASSERTCMP(conn_count, >=, call_count);
270: peer_head = smp_load_acquire(&b->peer_backlog_head);
271: peer_tail = b->peer_backlog_tail;
272: ASSERTCMP(CIRC_CNT(peer_head, peer_tail, RXRPC_BACKLOG_MAX), >=,
273: conn_count);
274:
275: if (call_count == 0)
276: return NULL;
277:
278: if (!conn) {
279: if (peer && !rxrpc_get_peer_maybe(peer, rxrpc_peer_get_service_conn))
280: peer = NULL;
281: if (!peer) {
282: peer = b->peer_backlog[peer_tail];
283: peer->srx = *peer_srx;
284: b->peer_backlog[peer_tail] = NULL;
285: smp_store_release(&b->peer_backlog_tail,
286: (peer_tail + 1) &
287: (RXRPC_BACKLOG_MAX - 1));
288:
289: rxrpc_new_incoming_peer(local, peer);
290: }
291:
292: /* Now allocate and set up the connection */
293: conn = b->conn_backlog[conn_tail];
294: b->conn_backlog[conn_tail] = NULL;
295: smp_store_release(&b->conn_backlog_tail,
296: (conn_tail + 1) & (RXRPC_BACKLOG_MAX - 1));
297: conn->local = rxrpc_get_local(local, rxrpc_local_get_prealloc_conn);
298: conn->peer = peer;
299: rxrpc_see_connection(conn, rxrpc_conn_see_new_service_conn);
300: rxrpc_new_incoming_connection(rx, conn, sec, skb);
301: } else {
302: rxrpc_get_connection(conn, rxrpc_conn_get_service_conn);
303: atomic_inc(&conn->active);
304: }
305:
306: /* And now we can allocate and set up a new call */
307: call = b->call_backlog[call_tail];
308: b->call_backlog[call_tail] = NULL;
309: smp_store_release(&b->call_backlog_tail,
310: (call_tail + 1) & (RXRPC_BACKLOG_MAX - 1));
311:
312: rxrpc_see_call(call, rxrpc_call_see_accept);
313: call->local = rxrpc_get_local(conn->local, rxrpc_local_get_call);
314: call->conn = conn;
315: call->security = conn->security;
316: call->security_ix = conn->security_ix;
317: call->peer = rxrpc_get_peer(conn->peer, rxrpc_peer_get_accept);
318: call->dest_srx = peer->srx;
319: call->cong_ssthresh = call->peer->cong_ssthresh;
320: call->tx_last_sent = ktime_get_real();
321: return call;
322: }
SourceFile:net/rxrpc/call_accept.c]
|
| 93/2 |
2026/03/10 00:55 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 87443
output: 48
thoughts: 58
Thoughts: Here's my summary, as the thought process unfolds:
**Digging into RXRPC Call Allocation**
Okay, so I'm tracing the path of an incoming RXRPC call. The initial clue is that `rxrpc_alloc_incoming_call` grabs a call structure from `b->call_backlog`. Alright, seems straightforward enough. But *where* did this call structure come from in the first place? Ah, I remember seeing `rxrpc_service_prealloc_one` mentioned, which likely relates to this pre-allocation of call structures. That's the key. I need to take a look at the code for `rxrpc_service_prealloc_one`. Let's see how that part of the system operates. That will explain the pre-allocation aspect of the calls that are now in the `call_backlog`. I'm hoping it's something that explains when and why these calls are pre-allocated and how that informs the overall incoming call process.
|
| 94/2 |
2026/03/10 00:55 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_accept.c IncludeLines:true Name:rxrpc_service_prealloc_one]
Results: map[SourceCode: 30: /*
31: * Preallocate a single service call, connection and peer and, if possible,
32: * give them a user ID and attach the user's side of the ID to them.
33: */
34: static int rxrpc_service_prealloc_one(struct rxrpc_sock *rx,
35: struct rxrpc_backlog *b,
36: rxrpc_notify_rx_t notify_rx,
37: unsigned long user_call_ID, gfp_t gfp,
38: unsigned int debug_id)
39: {
40: struct rxrpc_call *call, *xcall;
41: struct rxrpc_net *rxnet = rxrpc_net(sock_net(&rx->sk));
42: struct rb_node *parent, **pp;
43: int max, tmp;
44: unsigned int size = RXRPC_BACKLOG_MAX;
45: unsigned int head, tail, call_head, call_tail;
46:
47: max = rx->sk.sk_max_ack_backlog;
48: tmp = rx->sk.sk_ack_backlog;
49: if (tmp >= max) {
50: _leave(" = -ENOBUFS [full %u]", max);
51: return -ENOBUFS;
52: }
53: max -= tmp;
54:
55: /* We don't need more conns and peers than we have calls, but on the
56: * other hand, we shouldn't ever use more peers than conns or conns
57: * than calls.
58: */
59: call_head = b->call_backlog_head;
60: call_tail = READ_ONCE(b->call_backlog_tail);
61: tmp = CIRC_CNT(call_head, call_tail, size);
62: if (tmp >= max) {
63: _leave(" = -ENOBUFS [enough %u]", tmp);
64: return -ENOBUFS;
65: }
66: max = tmp + 1;
67:
68: head = b->peer_backlog_head;
69: tail = READ_ONCE(b->peer_backlog_tail);
70: if (CIRC_CNT(head, tail, size) < max) {
71: struct rxrpc_peer *peer;
72:
73: peer = rxrpc_alloc_peer(rx->local, gfp, rxrpc_peer_new_prealloc);
74: if (!peer)
75: return -ENOMEM;
76: b->peer_backlog[head] = peer;
77: smp_store_release(&b->peer_backlog_head,
78: (head + 1) & (size - 1));
79: }
80:
81: head = b->conn_backlog_head;
82: tail = READ_ONCE(b->conn_backlog_tail);
83: if (CIRC_CNT(head, tail, size) < max) {
84: struct rxrpc_connection *conn;
85:
86: conn = rxrpc_prealloc_service_connection(rxnet, gfp);
87: if (!conn)
88: return -ENOMEM;
89: b->conn_backlog[head] = conn;
90: smp_store_release(&b->conn_backlog_head,
91: (head + 1) & (size - 1));
92: }
93:
94: /* Now it gets complicated, because calls get registered with the
95: * socket here, with a user ID preassigned by the user.
96: */
97: call = rxrpc_alloc_call(rx, gfp, debug_id);
98: if (!call)
99: return -ENOMEM;
100: call->flags |= (1 << RXRPC_CALL_IS_SERVICE);
101: rxrpc_set_call_state(call, RXRPC_CALL_SERVER_PREALLOC);
102: __set_bit(RXRPC_CALL_EV_INITIAL_PING, &call->events);
103:
104: trace_rxrpc_call(call->debug_id, refcount_read(&call->ref),
105: user_call_ID, rxrpc_call_new_prealloc_service);
106:
107: write_lock(&rx->call_lock);
108:
109: /* Check the user ID isn't already in use */
110: pp = &rx->calls.rb_node;
111: parent = NULL;
112: while (*pp) {
113: parent = *pp;
114: xcall = rb_entry(parent, struct rxrpc_call, sock_node);
115: if (user_call_ID < xcall->user_call_ID)
116: pp = &(*pp)->rb_left;
117: else if (user_call_ID > xcall->user_call_ID)
118: pp = &(*pp)->rb_right;
119: else
120: goto id_in_use;
121: }
122:
123: call->user_call_ID = user_call_ID;
124: call->notify_rx = notify_rx;
125: if (rx->app_ops &&
126: rx->app_ops->user_attach_call) {
127: rxrpc_get_call(call, rxrpc_call_get_kernel_service);
128: rx->app_ops->user_attach_call(call, user_call_ID);
129: }
130:
131: rxrpc_get_call(call, rxrpc_call_get_userid);
132: rb_link_node(&call->sock_node, parent, pp);
133: rb_insert_color(&call->sock_node, &rx->calls);
134: set_bit(RXRPC_CALL_HAS_USERID, &call->flags);
135:
136: list_add(&call->sock_link, &rx->sock_calls);
137:
138: write_unlock(&rx->call_lock);
139:
140: rxnet = call->rxnet;
141: spin_lock(&rxnet->call_lock);
142: list_add_tail_rcu(&call->link, &rxnet->calls);
143: spin_unlock(&rxnet->call_lock);
144:
145: b->call_backlog[call_head] = call;
146: smp_store_release(&b->call_backlog_head, (call_head + 1) & (size - 1));
147: _leave(" = 0 [%d -> %lx]", call->debug_id, user_call_ID);
148: return 0;
149:
150: id_in_use:
151: write_unlock(&rx->call_lock);
152: rxrpc_prefail_call(call, RXRPC_CALL_LOCAL_ERROR, -EBADSLT);
153: rxrpc_cleanup_call(call);
154: _leave(" = -EBADSLT");
155: return -EBADSLT;
156: }
SourceFile:net/rxrpc/call_accept.c]
|
| 95/2 |
2026/03/10 00:55 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 89886
output: 46
thoughts: 137
Thoughts: **Analyzing the Allocation and Disposal of RXRPC Calls**
Okay, so `rxrpc_service_prealloc_one`… it allocates a call, sticks it in both `rx->sock_calls` and `b->call_backlog`. That seems standard enough. But hold on, the call's `ref` count is initialized to 2? One from `rxrpc_alloc_call` as expected, but the other from `rxrpc_get_call` using `rxrpc_call_get_userid`? Interesting... I need to remember that.
And wait a second. If the call's residing in `b->call_backlog`, it already has an initial reference, doesn't it? Something's adding a reference that needs to be accounted for.
Right, socket closure triggers `rxrpc_discard_prealloc`. That's where it all gets cleaned up. I better go over `rxrpc_discard_prealloc` and make absolutely sure I understand the reference counting and disposal path when a preallocated call is no longer needed. This looks like a potential point for a subtle race or memory leak, so let's delve deeper into this function now.
|
| 96/2 |
2026/03/10 00:55 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_accept.c IncludeLines:true Name:rxrpc_discard_prealloc]
Results: map[SourceCode: 176: /*
177: * Discard the preallocation on a service.
178: */
179: void rxrpc_discard_prealloc(struct rxrpc_sock *rx)
180: {
181: struct rxrpc_backlog *b = rx->backlog;
182: struct rxrpc_net *rxnet = rxrpc_net(sock_net(&rx->sk));
183: unsigned int size = RXRPC_BACKLOG_MAX, head, tail;
184:
185: if (!b)
186: return;
187: rx->backlog = NULL;
188:
189: /* Make sure that there aren't any incoming calls in progress before we
190: * clear the preallocation buffers.
191: */
192: spin_lock_irq(&rx->incoming_lock);
193: spin_unlock_irq(&rx->incoming_lock);
194:
195: head = b->peer_backlog_head;
196: tail = b->peer_backlog_tail;
197: while (CIRC_CNT(head, tail, size) > 0) {
198: struct rxrpc_peer *peer = b->peer_backlog[tail];
199: rxrpc_put_local(peer->local, rxrpc_local_put_prealloc_peer);
200: kfree(peer);
201: tail = (tail + 1) & (size - 1);
202: }
203:
204: head = b->conn_backlog_head;
205: tail = b->conn_backlog_tail;
206: while (CIRC_CNT(head, tail, size) > 0) {
207: struct rxrpc_connection *conn = b->conn_backlog[tail];
208: write_lock(&rxnet->conn_lock);
209: list_del(&conn->link);
210: list_del(&conn->proc_link);
211: write_unlock(&rxnet->conn_lock);
212: kfree(conn);
213: if (atomic_dec_and_test(&rxnet->nr_conns))
214: wake_up_var(&rxnet->nr_conns);
215: tail = (tail + 1) & (size - 1);
216: }
217:
218: head = b->call_backlog_head;
219: tail = b->call_backlog_tail;
220: while (CIRC_CNT(head, tail, size) > 0) {
221: struct rxrpc_call *call = b->call_backlog[tail];
222: rxrpc_see_call(call, rxrpc_call_see_discard);
223: rcu_assign_pointer(call->socket, rx);
224: if (rx->app_ops &&
225: rx->app_ops->discard_new_call) {
226: _debug("discard %lx", call->user_call_ID);
227: rx->app_ops->discard_new_call(call, call->user_call_ID);
228: if (call->notify_rx)
229: call->notify_rx = rxrpc_dummy_notify;
230: rxrpc_put_call(call, rxrpc_call_put_kernel);
231: }
232: rxrpc_call_completed(call);
233: rxrpc_release_call(rx, call);
234: rxrpc_put_call(call, rxrpc_call_put_discard_prealloc);
235: tail = (tail + 1) & (size - 1);
236: }
237:
238: kfree(b);
239: }
SourceFile:net/rxrpc/call_accept.c]
|
| 97/2 |
2026/03/10 00:55 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 91444
output: 43
thoughts: 440
Thoughts: **Analyzing the Reference Counting in RXRPC Call Handling**
Okay, let's break this down. My initial concern was centered around how `rxrpc_discard_prealloc` cleans up pre-allocated calls. It calls `rxrpc_release_call` and then `rxrpc_put_call`. I see that `rxrpc_release_call` removes the call from `rx->sock_calls` and crucially *drops the initial reference*. Then `rxrpc_put_call` drops the reference that was acquired by `rxrpc_get_call(call, rxrpc_call_get_userid)`. This means the call's reference count hits zero, and it is freed – the standard behavior. This is correct, as I expected.
Now, I need to consider what happens if `rxrpc_new_incoming_call` succeeds. This is where it gets a little more complex. I see that it eventually calls `rxrpc_put_call(call, rxrpc_call_put_input)`. But where does the initial reference come from? Ah, it calls `rxrpc_alloc_incoming_call` which gets the call from `b->call_backlog`. At this point, the call has a `ref` count of 2.
Next, `rxrpc_new_incoming_call` calls `rxrpc_incoming_call(rx, call, skb)`. Then, and this is the crux of the question, it calls `rxrpc_put_call(call, rxrpc_call_put_input)`. This `rxrpc_put_call` function needs clarification; which reference is it dropping? My analysis leads me to believe that it's the *initial* reference taken from `b->call_backlog`. If that is true, that leaves the call with a ref of 1. That remaining reference comes from `rxrpc_get_call(call, rxrpc_call_get_userid)`.
This remaining reference, I believe, will be dropped by `rxrpc_release_call` either when the socket is closed or when `recvmsg` has finished with the call. Since `rxrpc_recvmsg` *does* call `rxrpc_release_call` when a call is complete, this sounds correct. It would be wise to double-check the implementation of `rxrpc_recvmsg` to confirm.
|
| 98/2 |
2026/03/10 00:55 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/recvmsg.c IncludeLines:true Name:rxrpc_recvmsg]
Results: map[SourceCode: 361: /*
362: * Receive a message from an RxRPC socket
363: * - we need to be careful about two or more threads calling recvmsg
364: * simultaneously
365: */
366: int rxrpc_recvmsg(struct socket *sock, struct msghdr *msg, size_t len,
367: int flags)
368: {
369: struct rxrpc_call *call;
370: struct rxrpc_sock *rx = rxrpc_sk(sock->sk);
371: struct list_head *l;
372: unsigned int call_debug_id = 0;
373: size_t copied = 0;
374: long timeo;
375: int ret;
376:
377: DEFINE_WAIT(wait);
378:
379: trace_rxrpc_recvmsg(0, rxrpc_recvmsg_enter, 0);
380:
381: if (flags & (MSG_OOB | MSG_TRUNC))
382: return -EOPNOTSUPP;
383:
384: timeo = sock_rcvtimeo(&rx->sk, flags & MSG_DONTWAIT);
385:
386: try_again:
387: lock_sock(&rx->sk);
388:
389: /* Return immediately if a client socket has no outstanding calls */
390: if (RB_EMPTY_ROOT(&rx->calls) &&
391: list_empty(&rx->recvmsg_q) &&
392: skb_queue_empty_lockless(&rx->recvmsg_oobq) &&
393: rx->sk.sk_state != RXRPC_SERVER_LISTENING) {
394: release_sock(&rx->sk);
395: return -EAGAIN;
396: }
397:
398: if (list_empty(&rx->recvmsg_q)) {
399: ret = -EWOULDBLOCK;
400: if (timeo == 0) {
401: call = NULL;
402: goto error_no_call;
403: }
404:
405: release_sock(&rx->sk);
406:
407: /* Wait for something to happen */
408: prepare_to_wait_exclusive(sk_sleep(&rx->sk), &wait,
409: TASK_INTERRUPTIBLE);
410: ret = sock_error(&rx->sk);
411: if (ret)
412: goto wait_error;
413:
414: if (list_empty(&rx->recvmsg_q) &&
415: skb_queue_empty_lockless(&rx->recvmsg_oobq)) {
416: if (signal_pending(current))
417: goto wait_interrupted;
418: trace_rxrpc_recvmsg(0, rxrpc_recvmsg_wait, 0);
419: timeo = schedule_timeout(timeo);
420: }
421: finish_wait(sk_sleep(&rx->sk), &wait);
422: goto try_again;
423: }
424:
425: /* Deal with OOB messages before we consider getting normal data. */
426: if (!skb_queue_empty_lockless(&rx->recvmsg_oobq)) {
427: ret = rxrpc_recvmsg_oob(sock, msg, flags);
428: release_sock(&rx->sk);
429: if (ret == -EAGAIN)
430: goto try_again;
431: goto error_no_call;
432: }
433:
434: /* Find the next call and dequeue it if we're not just peeking. If we
435: * do dequeue it, that comes with a ref that we will need to release.
436: * We also want to weed out calls that got requeued whilst we were
437: * shovelling data out.
438: */
439: spin_lock_irq(&rx->recvmsg_lock);
440: l = rx->recvmsg_q.next;
441: call = list_entry(l, struct rxrpc_call, recvmsg_link);
442:
443: if (!rxrpc_call_is_complete(call) &&
444: skb_queue_empty(&call->recvmsg_queue) &&
445: skb_queue_empty(&rx->recvmsg_oobq)) {
446: list_del_init(&call->recvmsg_link);
447: spin_unlock_irq(&rx->recvmsg_lock);
448: release_sock(&rx->sk);
449: trace_rxrpc_recvmsg(call->debug_id, rxrpc_recvmsg_unqueue, 0);
450: rxrpc_put_call(call, rxrpc_call_put_recvmsg);
451: goto try_again;
452: }
453:
454: rxrpc_see_call(call, rxrpc_call_see_recvmsg);
455: if (test_bit(RXRPC_CALL_RELEASED, &call->flags)) {
456: rxrpc_see_call(call, rxrpc_call_see_already_released);
457: list_del_init(&call->recvmsg_link);
458: spin_unlock_irq(&rx->recvmsg_lock);
459: release_sock(&rx->sk);
460: trace_rxrpc_recvmsg(call->debug_id, rxrpc_recvmsg_unqueue, 0);
461: rxrpc_put_call(call, rxrpc_call_put_recvmsg);
462: goto try_again;
463: }
464: if (!(flags & MSG_PEEK))
465: list_del_init(&call->recvmsg_link);
466: else
467: rxrpc_get_call(call, rxrpc_call_get_recvmsg);
468: spin_unlock_irq(&rx->recvmsg_lock);
469:
470: call_debug_id = call->debug_id;
471: trace_rxrpc_recvmsg(call_debug_id, rxrpc_recvmsg_dequeue, 0);
472:
473: /* We're going to drop the socket lock, so we need to lock the call
474: * against interference by sendmsg.
475: */
476: if (!mutex_trylock(&call->user_mutex)) {
477: ret = -EWOULDBLOCK;
478: if (flags & MSG_DONTWAIT)
479: goto error_requeue_call;
480: ret = -ERESTARTSYS;
481: if (mutex_lock_interruptible(&call->user_mutex) < 0)
482: goto error_requeue_call;
483: }
484:
485: release_sock(&rx->sk);
486:
487: if (test_bit(RXRPC_CALL_RELEASED, &call->flags)) {
488: rxrpc_see_call(call, rxrpc_call_see_already_released);
489: mutex_unlock(&call->user_mutex);
490: if (!(flags & MSG_PEEK))
491: rxrpc_put_call(call, rxrpc_call_put_recvmsg);
492: goto try_again;
493: }
494:
495: ret = rxrpc_recvmsg_user_id(call, msg, flags);
496: if (ret < 0)
497: goto error_unlock_call;
498:
499: if (msg->msg_name && call->peer) {
500: size_t len = sizeof(call->dest_srx);
501:
502: memcpy(msg->msg_name, &call->dest_srx, len);
503: msg->msg_namelen = len;
504: }
505:
506: ret = rxrpc_recvmsg_data(sock, call, msg, &msg->msg_iter, len,
507: flags, &copied);
508: if (ret == -EAGAIN)
509: ret = 0;
510: if (ret == -EIO)
511: goto call_failed;
512: if (ret < 0)
513: goto error_unlock_call;
514:
515: if (rxrpc_call_is_complete(call) &&
516: skb_queue_empty(&call->recvmsg_queue))
517: goto call_complete;
518: if (rxrpc_call_has_failed(call))
519: goto call_failed;
520:
521: if (!(flags & MSG_PEEK) &&
522: !skb_queue_empty(&call->recvmsg_queue))
523: rxrpc_notify_socket(call);
524: goto not_yet_complete;
525:
526: call_failed:
527: rxrpc_purge_queue(&call->recvmsg_queue);
528: call_complete:
529: ret = rxrpc_recvmsg_term(call, msg);
530: if (ret < 0)
531: goto error_unlock_call;
532: if (!(flags & MSG_PEEK))
533: rxrpc_release_call(rx, call);
534: msg->msg_flags |= MSG_EOR;
535: ret = 1;
536:
537: not_yet_complete:
538: if (ret == 0)
539: msg->msg_flags |= MSG_MORE;
540: else
541: msg->msg_flags &= ~MSG_MORE;
542: ret = copied;
543:
544: error_unlock_call:
545: mutex_unlock(&call->user_mutex);
546: rxrpc_put_call(call, rxrpc_call_put_recvmsg);
547: trace_rxrpc_recvmsg(call_debug_id, rxrpc_recvmsg_return, ret);
548: return ret;
549:
550: error_requeue_call:
551: if (!(flags & MSG_PEEK)) {
552: spin_lock_irq(&rx->recvmsg_lock);
553: if (list_empty(&call->recvmsg_link)) {
554: list_add(&call->recvmsg_link, &rx->recvmsg_q);
555: rxrpc_see_call(call, rxrpc_call_see_recvmsg_requeue);
556: spin_unlock_irq(&rx->recvmsg_lock);
557: } else if (list_is_first(&call->recvmsg_link, &rx->recvmsg_q)) {
558: spin_unlock_irq(&rx->recvmsg_lock);
559: rxrpc_put_call(call, rxrpc_call_see_recvmsg_requeue_first);
560: } else {
561: list_move(&call->recvmsg_link, &rx->recvmsg_q);
562: spin_unlock_irq(&rx->recvmsg_lock);
563: rxrpc_put_call(call, rxrpc_call_see_recvmsg_requeue_move);
564: }
565: trace_rxrpc_recvmsg(call_debug_id, rxrpc_recvmsg_requeue, 0);
566: } else {
567: rxrpc_put_call(call, rxrpc_call_put_recvmsg_peek_nowait);
568: }
569: error_no_call:
570: release_sock(&rx->sk);
571: error_trace:
572: trace_rxrpc_recvmsg(call_debug_id, rxrpc_recvmsg_return, ret);
573: return ret;
574:
575: wait_interrupted:
576: ret = sock_intr_errno(timeo);
577: wait_error:
578: finish_wait(sk_sleep(&rx->sk), &wait);
579: call = NULL;
580: goto error_trace;
581: }
SourceFile:net/rxrpc/recvmsg.c]
|
| 99/2 |
2026/03/10 00:55 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 96075
output: 48
thoughts: 621
Thoughts: **RX-RPC Call Lifecycle Deep Dive**
Okay, let me think this through, really focusing on these call reference counts and where they're getting decremented. So, `rxrpc_recvmsg`... if a call is complete, *and* the `MSG_PEEK` flag isn't set, then `rxrpc_recvmsg` calls `rxrpc_release_call`. Right, `rxrpc_release_call` *removes* the call from `rx->sock_calls` and that drops the initial reference to the call. Okay, that's step one.
Then, `rxrpc_recvmsg` hits `rxrpc_put_call(call, rxrpc_call_put_recvmsg)` at `error_unlock_call`. Before we get to that, wait, `rxrpc_recvmsg` also removes the call from `rx->recvmsg_q` at the *beginning* of the function. Ah, I see! If it's removed, it's taking *over* the reference previously held by `rx->recvmsg_q`. And *then* it drops that reference at the end of the function. So, that's a decrement of one reference count.
Okay, so putting it together. `rxrpc_release_call` drops the initial reference and then `rxrpc_put_call` drops the `recvmsg_q` reference - that gets the call's ref count down to zero and frees it.
But hold on... What if `recvmsg` is *never* called? The call just sits there in both `rx->recvmsg_q` and `rx->sock_calls`. Then, when the socket closes, `rxrpc_release_calls_on_socket` kicks in. That function removes the call from `rx->sock_calls` (and calls `rxrpc_release_call` to drop that initial reference). Then, it removes the call from `rx->recvmsg_q` and calls `rxrpc_put_call` (which drops the `recvmsg_q` reference). Bingo, again, ref count hits zero, and the call is freed.
What if the call isn't in `rx->recvmsg_q`? If it’s not in there, it doesn’t have the `recvmsg_q` reference. In this case, `rxrpc_release_calls_on_socket` just drops the initial reference. The call's refcount then hits zero.
Now, let me think about server and client calls. What if the call is a *SERVER* call and it’s *not* in `rx->sock_calls`? When *would* a server call be added to `rx->sock_calls`? Ah, `rxrpc_service_prealloc_one`! So, yes, it *is* in `rx->sock_calls`.
And what about a *CLIENT* call that’s *not* in `rx->sock_calls`? Well, `rxrpc_new_client_call` puts it in `rx->sock_calls`. So it *is* there as well.
Okay, one last thought - what about a *KERNEL* call? These are created by `rxrpc_kernel_begin_call`. Got it.
|
| 100/2 |
2026/03/10 00:55 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/af_rxrpc.c IncludeLines:true Name:rxrpc_kernel_begin_call]
Results: map[SourceCode: 313: /**
314: * rxrpc_kernel_begin_call - Allow a kernel service to begin a call
315: * @sock: The socket on which to make the call
316: * @peer: The peer to contact
317: * @key: The security context to use (defaults to socket setting)
318: * @user_call_ID: The ID to use
319: * @tx_total_len: Total length of data to transmit during the call (or -1)
320: * @hard_timeout: The maximum lifespan of the call in sec
321: * @gfp: The allocation constraints
322: * @notify_rx: Where to send notifications instead of socket queue
323: * @service_id: The ID of the service to contact
324: * @upgrade: Request service upgrade for call
325: * @interruptibility: The call is interruptible, or can be canceled.
326: * @debug_id: The debug ID for tracing to be assigned to the call
327: *
328: * Allow a kernel service to begin a call on the nominated socket. This just
329: * sets up all the internal tracking structures and allocates connection and
330: * call IDs as appropriate.
331: *
332: * The default socket destination address and security may be overridden by
333: * supplying @srx and @key.
334: *
335: * Return: The new call or an error code.
336: */
337: struct rxrpc_call *rxrpc_kernel_begin_call(struct socket *sock,
338: struct rxrpc_peer *peer,
339: struct key *key,
340: unsigned long user_call_ID,
341: s64 tx_total_len,
342: u32 hard_timeout,
343: gfp_t gfp,
344: rxrpc_notify_rx_t notify_rx,
345: u16 service_id,
346: bool upgrade,
347: enum rxrpc_interruptibility interruptibility,
348: unsigned int debug_id)
349: {
350: struct rxrpc_conn_parameters cp;
351: struct rxrpc_call_params p;
352: struct rxrpc_call *call;
353: struct rxrpc_sock *rx = rxrpc_sk(sock->sk);
354:
355: _enter(",,%x,%lx", key_serial(key), user_call_ID);
356:
357: if (WARN_ON_ONCE(peer->local != rx->local))
358: return ERR_PTR(-EIO);
359:
360: lock_sock(&rx->sk);
361:
362: if (!key)
363: key = rx->key;
364: if (key && !key->payload.data[0])
365: key = NULL; /* a no-security key */
366:
367: memset(&p, 0, sizeof(p));
368: p.user_call_ID = user_call_ID;
369: p.tx_total_len = tx_total_len;
370: p.interruptibility = interruptibility;
371: p.kernel = true;
372: p.timeouts.hard = hard_timeout;
373:
374: memset(&cp, 0, sizeof(cp));
375: cp.local = rx->local;
376: cp.peer = peer;
377: cp.key = key;
378: cp.security_level = rx->min_sec_level;
379: cp.exclusive = false;
380: cp.upgrade = upgrade;
381: cp.service_id = service_id;
382: call = rxrpc_new_client_call(rx, &cp, &p, gfp, debug_id);
383: /* The socket has been unlocked. */
384: if (!IS_ERR(call)) {
385: call->notify_rx = notify_rx;
386: mutex_unlock(&call->user_mutex);
387: }
388:
389: _leave(" = %p", call);
390: return call;
391: }
SourceFile:net/rxrpc/af_rxrpc.c]
|
| 101/2 |
2026/03/10 00:55 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 98747
output: 42
thoughts: 389
Thoughts: **My Analysis of RXRPC Kernel Call Lifecycle**
Okay, let me break this down. `rxrpc_kernel_begin_call` is calling `rxrpc_new_client_call`. That creates a call with a reference count of 2. It then *unlocks* `call->user_mutex` and returns the call to the caller. Crucially, the *caller* now holds a reference to the call. The anticipated cleanup happens when the caller finally calls `rxrpc_kernel_put_call`. That function, in turn, calls `rxrpc_put_call` with the `rxrpc_call_put_kernel` flag. This drops the caller's reference. The remaining reference is now held within `rx->sock_calls`, which is typical socket lifecycle. This last reference is taken care of by `rxrpc_release_calls_on_socket` which happens when the socket closes. So the call should be safely freed because the ref count hits zero.
But here's where the problem arises. What if the caller *never* calls `rxrpc_kernel_put_call`? Specifically, if AFS is the caller it calls `rxrpc_kernel_put_call` when the call is complete. Suppose the network namespace gets destroyed? `afs_net_exit` calls `afs_close_socket`, which, I know, closes the socket. Consequently, `rxrpc_release_calls_on_socket` gets called, which aborts the call and drops the initial reference held in `rx->sock_calls`. But if AFS *still* holds a reference, the reference count will be 1! Therefore, the call isn't freed. If the call isn't freed, then `rxnet->nr_calls` doesn't reach zero. If `rxnet->nr_calls` doesn't reach zero, `rxrpc_destroy_all_calls` hangs. So, I need to check to see if AFS is actually holding a reference to the call after the socket closes. I will need to verify the lifecycle of the call in `afs_make_call`.
|
| 102/2 |
2026/03/10 00:55 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/afs/rxrpc.c IncludeLines:true Name:afs_make_call]
Results: map[SourceCode: 333: /*
334: * Initiate a call and synchronously queue up the parameters for dispatch. Any
335: * error is stored into the call struct, which the caller must check for.
336: */
337: void afs_make_call(struct afs_call *call, gfp_t gfp)
338: {
339: struct rxrpc_call *rxcall;
340: struct msghdr msg;
341: struct kvec iov[1];
342: size_t len;
343: s64 tx_total_len;
344: int ret;
345:
346: _enter(",{%pISp+%u},", rxrpc_kernel_remote_addr(call->peer), call->service_id);
347:
348: ASSERT(call->type != NULL);
349: ASSERT(call->type->name != NULL);
350:
351: _debug("____MAKE %p{%s,%x} [%d]____",
352: call, call->type->name, key_serial(call->key),
353: atomic_read(&call->net->nr_outstanding_calls));
354:
355: trace_afs_make_call(call);
356:
357: /* Work out the length we're going to transmit. This is awkward for
358: * calls such as FS.StoreData where there's an extra injection of data
359: * after the initial fixed part.
360: */
361: tx_total_len = call->request_size;
362: if (call->write_iter)
363: tx_total_len += iov_iter_count(call->write_iter);
364:
365: /* If the call is going to be asynchronous, we need an extra ref for
366: * the call to hold itself so the caller need not hang on to its ref.
367: */
368: if (call->async) {
369: afs_get_call(call, afs_call_trace_get);
370: call->drop_ref = true;
371: }
372:
373: /* create a call */
374: rxcall = rxrpc_kernel_begin_call(call->net->socket, call->peer, call->key,
375: (unsigned long)call,
376: tx_total_len,
377: call->max_lifespan,
378: gfp,
379: (call->async ?
380: afs_wake_up_async_call :
381: afs_wake_up_call_waiter),
382: call->service_id,
383: call->upgrade,
384: (call->intr ? RXRPC_PREINTERRUPTIBLE :
385: RXRPC_UNINTERRUPTIBLE),
386: call->debug_id);
387: if (IS_ERR(rxcall)) {
388: ret = PTR_ERR(rxcall);
389: call->error = ret;
390: goto error_kill_call;
391: }
392:
393: call->rxcall = rxcall;
394: call->issue_time = ktime_get_real();
395:
396: /* send the request */
397: iov[0].iov_base = call->request;
398: iov[0].iov_len = call->request_size;
399:
400: msg.msg_name = NULL;
401: msg.msg_namelen = 0;
402: iov_iter_kvec(&msg.msg_iter, ITER_SOURCE, iov, 1, call->request_size);
403: msg.msg_control = NULL;
404: msg.msg_controllen = 0;
405: msg.msg_flags = MSG_WAITALL | (call->write_iter ? MSG_MORE : 0);
406:
407: ret = rxrpc_kernel_send_data(call->net->socket, rxcall,
408: &msg, call->request_size,
409: afs_notify_end_request_tx);
410: if (ret < 0)
411: goto error_do_abort;
412:
413: if (call->write_iter) {
414: msg.msg_iter = *call->write_iter;
415: msg.msg_flags &= ~MSG_MORE;
416: trace_afs_send_data(call, &msg);
417:
418: ret = rxrpc_kernel_send_data(call->net->socket,
419: call->rxcall, &msg,
420: iov_iter_count(&msg.msg_iter),
421: afs_notify_end_request_tx);
422: *call->write_iter = msg.msg_iter;
423:
424: trace_afs_sent_data(call, &msg, ret);
425: if (ret < 0)
426: goto error_do_abort;
427: }
428:
429: /* Note that at this point, we may have received the reply or an abort
430: * - and an asynchronous call may already have completed.
431: *
432: * afs_wait_for_call_to_complete(call)
433: * must be called to synchronously clean up.
434: */
435: return;
436:
437: error_do_abort:
438: if (ret != -ECONNABORTED)
439: rxrpc_kernel_abort_call(call->net->socket, rxcall,
440: RX_USER_ABORT, ret,
441: afs_abort_send_data_error);
442: if (call->async) {
443: afs_see_call(call, afs_call_trace_async_abort);
444: return;
445: }
446:
447: if (ret == -ECONNABORTED) {
448: len = 0;
449: iov_iter_kvec(&msg.msg_iter, ITER_DEST, NULL, 0, 0);
450: rxrpc_kernel_recv_data(call->net->socket, rxcall,
451: &msg.msg_iter, &len, false,
452: &call->abort_code, &call->service_id);
453: call->responded = true;
454: }
455: call->error = ret;
456: trace_afs_call_done(call);
457: error_kill_call:
458: if (call->async)
459: afs_see_call(call, afs_call_trace_async_kill);
460: if (call->type->immediate_cancel)
461: call->type->immediate_cancel(call);
462:
463: /* We need to dispose of the extra ref we grabbed for an async call.
464: * The call, however, might be queued on afs_async_calls and we need to
465: * make sure we don't get any more notifications that might requeue it.
466: */
467: if (call->rxcall)
468: rxrpc_kernel_shutdown_call(call->net->socket, call->rxcall);
469: if (call->async) {
470: if (cancel_work_sync(&call->async_work))
471: afs_put_call(call);
472: afs_set_call_complete(call, ret, 0);
473: }
474:
475: call->error = ret;
476: call->state = AFS_CALL_COMPLETE;
477: _leave(" = %d", ret);
478: }
SourceFile:fs/afs/rxrpc.c]
|
| 103/2 |
2026/03/10 00:55 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 101990
output: 43
thoughts: 308
Thoughts: **AFS Call Lifecycle and Network Namespace Destruction**
Okay, let me trace this AFS call lifecycle. It looks like `afs_make_call` kicks things off, and that calls `rxrpc_kernel_begin_call`. Then we're calling `rxrpc_kernel_send_data`. Right, if it's asynchronous, it immediately returns. When it actually completes, `afs_wake_up_async_call` gets called. This schedules `call->async_work`, which means `afs_process_async_call` runs later. `afs_process_async_call` then calls `afs_deliver_to_call`. After that, we hit `afs_put_call(call)`. And this in turn calls `afs_call_destructor`. Aha, the crucial bit: `afs_call_destructor` ends up calling `rxrpc_kernel_put_call(call->net->socket, call->rxcall)`. So, AFS is dropping its reference at the point the call is freed, got it!
Now, the network namespace destruction scenario is where it gets interesting. I see that if the network namespace is destroyed, `afs_net_exit` gets called. That makes sense. `afs_net_exit` then calls `afs_close_socket(net)`. Okay, so the socket is closed. But what about all those asynchronous calls that might still be outstanding? That’s what I’m most interested in.
Ah, here it is: `afs_net_exit` calls `afs_wait_for_servers(net)` *before* `afs_close_socket(net)`. So, that *should* be the mechanism to handle outstanding calls. But the question is: what does `afs_wait_for_servers` actually *do*? Does it wait for all those calls to complete? I should verify that. Time to check `afs_wait_for_servers`.
|
| 104/2 |
2026/03/10 00:55 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/afs/server.c IncludeLines:true Name:afs_wait_for_servers]
Results: map[SourceCode: 526: /*
527: * Wait for outstanding servers.
528: */
529: void afs_wait_for_servers(struct afs_net *net)
530: {
531: _enter("");
532:
533: atomic_dec(&net->servers_outstanding);
534: wait_var_event(&net->servers_outstanding,
535: !atomic_read(&net->servers_outstanding));
536: _leave("");
537: }
SourceFile:fs/afs/server.c]
|
| 105/2 |
2026/03/10 00:55 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 102937
output: 50
thoughts: 380
Thoughts: **A Deep Dive into AFS Server Destruction and Potential Hangs**
Okay, let's break this down. My initial thought is that `afs_wait_for_servers` blocks until `net->servers_outstanding` hits zero. But *how* does that happen? It's when all the servers are destroyed, right? Wait, hold on... are all the calls also destroyed when all servers are? No! Calls *hold* references to servers!
Aha! So, if a call holds a reference, the server *won't* be destroyed until the call itself is. That means `afs_wait_for_servers` is really waiting for *all* calls to be destroyed! That makes sense.
If `afs_wait_for_servers` waits on all calls, then `afs_close_socket` *won't* be called until all calls are destroyed. If all calls are destroyed, `rxrpc_kernel_put_call` must've been called for *every* call. Therefore, `rxrpc_destroy_all_calls` should *not* hang.
But... what if a call *isn't* destroyed? Why would a call linger? It's waiting for a reply from the server! If the network namespace is being destroyed, the network interfaces are down. The server will *never* reply! So, the call will just... wait... until a timeout expires.
Hmm, what *is* the timeout? `afs_make_call` sets the `hard_timeout` to `call->max_lifespan`. What's `max_lifespan` usually set to? A few minutes, I'd guess.
But the bug report says, "INFO: task kworker/u8:22:6285 blocked for more than 144 seconds." That's 2.4 minutes! If the timeout is larger than that, the task *will* be flagged as hung! The real question is: What is the actual timeout configured for AFS calls? I need to get to the bottom of this!
|
| 106/2 |
2026/03/10 00:55 |
tool |
codesearch-find-references |
0m
Error:
Args: map[ContextFile:fs/afs/rxrpc.c IncludeSnippetLines:1 Name:max_lifespan SourceTreePrefix:]
Results: map[References:<nil> TruncatedOutput:false]
|
| 107/2 |
2026/03/10 00:55 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 103834
output: 19
thoughts: 19
Thoughts: Okay, here's my summarized thought process, presented as if I were thinking it:
**Digging into the Filesystem**
Alright, I need to find where the `max_lifespan` variable is defined or used within the filesystem, specifically the AFS component. Let's start the search within the `fs/afs/` directory. That's usually the right place to begin poking around for related configurations. Hopefully, a quick grep or find will turn it up without too much fuss. I’m expecting to find it in a config file, maybe a header file, or possibly even used within some kernel module code if it's deeply tied to AFS operation. Time to get to work!
|
| 108/2 |
2026/03/10 00:56 |
tool |
grepper |
0m
Args: map[Expression:max_lifespan]
Results: map[Output:fs/afs/fsclient.c=1749=bool afs_fs_get_capabilities(struct afs_net *net, struct afs_server *server,
--
fs/afs/fsclient.c-1769- call->async = true;
fs/afs/fsclient.c:1770: call->max_lifespan = AFS_PROBE_MAX_LIFESPAN;
fs/afs/fsclient.c-1771-
--
fs/afs/internal.h=129=struct afs_call {
--
fs/afs/internal.h-167- unsigned long long remaining; /* How much is left to receive */
fs/afs/internal.h:168: unsigned int max_lifespan; /* Maximum lifespan in secs to set if not 0 */
fs/afs/internal.h-169- unsigned request_size; /* size of request data */
--
fs/afs/rxrpc.c=337=void afs_make_call(struct afs_call *call, gfp_t gfp)
--
fs/afs/rxrpc.c-376- tx_total_len,
fs/afs/rxrpc.c:377: call->max_lifespan,
fs/afs/rxrpc.c-378- gfp,
--
fs/afs/vlclient.c=110=struct afs_vldb_entry *afs_vl_get_entry_by_name_u(struct afs_vl_cursor *vc,
--
fs/afs/vlclient.c-137- call->ret_vldb = entry;
fs/afs/vlclient.c:138: call->max_lifespan = AFS_VL_MAX_LIFESPAN;
fs/afs/vlclient.c-139- call->peer = rxrpc_kernel_get_peer(vc->alist->addrs[vc->addr_index].peer);
--
fs/afs/vlclient.c=258=struct afs_addr_list *afs_vl_get_addrs_u(struct afs_vl_cursor *vc,
--
fs/afs/vlclient.c-278- call->ret_alist = NULL;
fs/afs/vlclient.c:279: call->max_lifespan = AFS_VL_MAX_LIFESPAN;
fs/afs/vlclient.c-280- call->peer = rxrpc_kernel_get_peer(vc->alist->addrs[vc->addr_index].peer);
--
fs/afs/vlclient.c=385=struct afs_call *afs_vl_get_capabilities(struct afs_net *net,
--
fs/afs/vlclient.c-409- call->async = true;
fs/afs/vlclient.c:410: call->max_lifespan = AFS_PROBE_MAX_LIFESPAN;
fs/afs/vlclient.c-411-
--
fs/afs/vlclient.c=634=struct afs_addr_list *afs_yfsvl_get_endpoints(struct afs_vl_cursor *vc,
--
fs/afs/vlclient.c-651- call->ret_alist = NULL;
fs/afs/vlclient.c:652: call->max_lifespan = AFS_VL_MAX_LIFESPAN;
fs/afs/vlclient.c-653- call->peer = rxrpc_kernel_get_peer(vc->alist->addrs[vc->addr_index].peer);
--
fs/afs/vlclient.c=756=char *afs_yfsvl_get_cell_name(struct afs_vl_cursor *vc)
--
fs/afs/vlclient.c-770- call->ret_str = NULL;
fs/afs/vlclient.c:771: call->max_lifespan = AFS_VL_MAX_LIFESPAN;
fs/afs/vlclient.c-772- call->peer = rxrpc_kernel_get_peer(vc->alist->addrs[vc->addr_index].peer);
]
|
| 109/2 |
2026/03/10 00:56 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 105073
output: 25
thoughts: 37
Thoughts: Okay, here's my interpretation of those thoughts, presented as a first-person summary, keeping in mind my expertise:
**Checking Lifespan Parameters in AFS**
Alright, hang on a second. I'm seeing `AFS_PROBE_MAX_LIFESPAN` and `AFS_VL_MAX_LIFESPAN` mentioned here. That immediately triggers a mental checklist. We're talking about the lifespan of probes and volume location information in AFS. Given the performance implications of stale data and the overhead of constant updates, lifespan parameters are crucial.
My first thought is, what are the current values set to? I need to know the numerical values associated with these constants. Are they optimized for our environment? Are they set conservatively, perhaps impacting availability? Or are they set too aggressively, potentially leading to unnecessary load on the system? I need to see those values directly to begin to assess their impact and if further tuning is needed. Gotta get those numbers.
|
| 110/2 |
2026/03/10 00:56 |
tool |
grepper |
0m
Args: map[Expression:AFS_PROBE_MAX_LIFESPAN]
Results: map[Output:fs/afs/afs.h-22-#define AFS_VL_MAX_LIFESPAN 120
fs/afs/afs.h:23:#define AFS_PROBE_MAX_LIFESPAN 30
fs/afs/afs.h-24-
--
fs/afs/fsclient.c=1749=bool afs_fs_get_capabilities(struct afs_net *net, struct afs_server *server,
--
fs/afs/fsclient.c-1769- call->async = true;
fs/afs/fsclient.c:1770: call->max_lifespan = AFS_PROBE_MAX_LIFESPAN;
fs/afs/fsclient.c-1771-
--
fs/afs/vlclient.c=385=struct afs_call *afs_vl_get_capabilities(struct afs_net *net,
--
fs/afs/vlclient.c-409- call->async = true;
fs/afs/vlclient.c:410: call->max_lifespan = AFS_PROBE_MAX_LIFESPAN;
fs/afs/vlclient.c-411-
]
|
| 111/2 |
2026/03/10 00:56 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 105650
output: 48
thoughts: 115
Thoughts: Here's my thought process, summarized as you requested:
**Timeout Conundrum**
Okay, so I'm looking at this `AFS_VL_MAX_LIFESPAN` which is defined as 120 seconds. Right, so if we have a call that’s waiting for a reply, and we hit that 120-second mark, we should expect a timeout, naturally. But, hold on, the task hung timeout is set to 144 seconds! That’s strange… the call *should* timeout before we hit the task hung condition, shouldn’t it? This feels… off.
Now I’m wondering, what happens if `max_lifespan` – which I believe is related to `AFS_VL_MAX_LIFESPAN` – is actually zero? What’s the behavior in that edge case? If it’s zero, what’s *actually* the timeout mechanism at play? We need to figure that out because zero could mean "infinite", "immediate" or some other implementation detail. I need to understand what this means more precisely.
Alright, I need to check `rxrpc_kernel_begin_call` immediately. That’s where this thing gets started, and it should reveal what’s going on with the timeout calculation in this context.
|
| 112/2 |
2026/03/10 00:56 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/af_rxrpc.c IncludeLines:true Name:rxrpc_kernel_begin_call]
Results: map[SourceCode: 313: /**
314: * rxrpc_kernel_begin_call - Allow a kernel service to begin a call
315: * @sock: The socket on which to make the call
316: * @peer: The peer to contact
317: * @key: The security context to use (defaults to socket setting)
318: * @user_call_ID: The ID to use
319: * @tx_total_len: Total length of data to transmit during the call (or -1)
320: * @hard_timeout: The maximum lifespan of the call in sec
321: * @gfp: The allocation constraints
322: * @notify_rx: Where to send notifications instead of socket queue
323: * @service_id: The ID of the service to contact
324: * @upgrade: Request service upgrade for call
325: * @interruptibility: The call is interruptible, or can be canceled.
326: * @debug_id: The debug ID for tracing to be assigned to the call
327: *
328: * Allow a kernel service to begin a call on the nominated socket. This just
329: * sets up all the internal tracking structures and allocates connection and
330: * call IDs as appropriate.
331: *
332: * The default socket destination address and security may be overridden by
333: * supplying @srx and @key.
334: *
335: * Return: The new call or an error code.
336: */
337: struct rxrpc_call *rxrpc_kernel_begin_call(struct socket *sock,
338: struct rxrpc_peer *peer,
339: struct key *key,
340: unsigned long user_call_ID,
341: s64 tx_total_len,
342: u32 hard_timeout,
343: gfp_t gfp,
344: rxrpc_notify_rx_t notify_rx,
345: u16 service_id,
346: bool upgrade,
347: enum rxrpc_interruptibility interruptibility,
348: unsigned int debug_id)
349: {
350: struct rxrpc_conn_parameters cp;
351: struct rxrpc_call_params p;
352: struct rxrpc_call *call;
353: struct rxrpc_sock *rx = rxrpc_sk(sock->sk);
354:
355: _enter(",,%x,%lx", key_serial(key), user_call_ID);
356:
357: if (WARN_ON_ONCE(peer->local != rx->local))
358: return ERR_PTR(-EIO);
359:
360: lock_sock(&rx->sk);
361:
362: if (!key)
363: key = rx->key;
364: if (key && !key->payload.data[0])
365: key = NULL; /* a no-security key */
366:
367: memset(&p, 0, sizeof(p));
368: p.user_call_ID = user_call_ID;
369: p.tx_total_len = tx_total_len;
370: p.interruptibility = interruptibility;
371: p.kernel = true;
372: p.timeouts.hard = hard_timeout;
373:
374: memset(&cp, 0, sizeof(cp));
375: cp.local = rx->local;
376: cp.peer = peer;
377: cp.key = key;
378: cp.security_level = rx->min_sec_level;
379: cp.exclusive = false;
380: cp.upgrade = upgrade;
381: cp.service_id = service_id;
382: call = rxrpc_new_client_call(rx, &cp, &p, gfp, debug_id);
383: /* The socket has been unlocked. */
384: if (!IS_ERR(call)) {
385: call->notify_rx = notify_rx;
386: mutex_unlock(&call->user_mutex);
387: }
388:
389: _leave(" = %p", call);
390: return call;
391: }
SourceFile:net/rxrpc/af_rxrpc.c]
|
| 113/2 |
2026/03/10 00:56 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 107393
output: 47
thoughts: 32
Thoughts: Okay, here's my interpretation of those thoughts, tailored for an expert audience, presented in the first person:
**Decoding the Timeout Logic**
Alright, hang on a second. We're looking at `hard_timeout`, and it's zero. Now, my immediate question is: *what* does that actually mean for the effective timeout value? Is zero a valid setting, or is it interpreted as something else, like "no timeout" or a default? It's not immediately obvious from context.
I need to dig into the code and clarify the behavior. I'm going to jump straight into `rxrpc_alloc_client_call`. That function's where the client call structure gets initialized, and that's the place that I expect will handle interpreting the `hard_timeout` value. I'm guessing that's where any default timeout values would be set, or where the special handling for a zero value would occur. Time to verify my suspicion and understand this behavior completely. It's crucial for understanding the client-side interaction here.
|
| 114/2 |
2026/03/10 00:56 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_object.c IncludeLines:true Name:rxrpc_alloc_client_call]
Results: map[SourceCode: 185: /*
186: * Allocate a new client call.
187: */
188: static struct rxrpc_call *rxrpc_alloc_client_call(struct rxrpc_sock *rx,
189: struct rxrpc_conn_parameters *cp,
190: struct rxrpc_call_params *p,
191: gfp_t gfp,
192: unsigned int debug_id)
193: {
194: struct rxrpc_call *call;
195: ktime_t now;
196: int ret;
197:
198: _enter("");
199:
200: call = rxrpc_alloc_call(rx, gfp, debug_id);
201: if (!call)
202: return ERR_PTR(-ENOMEM);
203: now = ktime_get_real();
204: call->acks_latest_ts = now;
205: call->cong_tstamp = now;
206: call->dest_srx = cp->peer->srx;
207: call->dest_srx.srx_service = cp->service_id;
208: call->interruptibility = p->interruptibility;
209: call->tx_total_len = p->tx_total_len;
210: call->key = key_get(cp->key);
211: call->peer = rxrpc_get_peer(cp->peer, rxrpc_peer_get_call);
212: call->local = rxrpc_get_local(cp->local, rxrpc_local_get_call);
213: call->security_level = cp->security_level;
214: if (p->kernel)
215: __set_bit(RXRPC_CALL_KERNEL, &call->flags);
216: if (cp->upgrade)
217: __set_bit(RXRPC_CALL_UPGRADE, &call->flags);
218: if (cp->exclusive)
219: __set_bit(RXRPC_CALL_EXCLUSIVE, &call->flags);
220:
221: if (p->timeouts.normal)
222: call->next_rx_timo = umin(p->timeouts.normal, 1);
223: if (p->timeouts.idle)
224: call->next_req_timo = umin(p->timeouts.idle, 1);
225: if (p->timeouts.hard)
226: call->hard_timo = p->timeouts.hard;
227:
228: ret = rxrpc_init_client_call_security(call);
229: if (ret < 0) {
230: rxrpc_prefail_call(call, RXRPC_CALL_LOCAL_ERROR, ret);
231: rxrpc_put_call(call, rxrpc_call_put_discard_error);
232: return ERR_PTR(ret);
233: }
234:
235: rxrpc_set_call_state(call, RXRPC_CALL_CLIENT_AWAIT_CONN);
236:
237: trace_rxrpc_call(call->debug_id, refcount_read(&call->ref),
238: p->user_call_ID, rxrpc_call_new_client);
239:
240: _leave(" = %p", call);
241: return call;
242: }
SourceFile:net/rxrpc/call_object.c]
|
| 115/2 |
2026/03/10 00:56 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 108721
output: 47
thoughts: 60
Thoughts: Okay, let me think this through.
**Analyzing Timeout Logic in the Code**
Hmm, so I see this condition: `p->timeouts.hard` being zero directly implies `call->hard_timo` is also zero. That makes sense, but the implications... If `call->hard_timo` is indeed zero, then what's governing the actual timeout behavior? I need to understand what happens when there's no "hard" timeout set. Where is the fallback or alternative timeout mechanism defined?
Right, the next logical step is to trace how the timeout is *actually* started. I need to refresh my memory on `rxrpc_start_call_timer`. I should probably go back and examine that function's implementation to see how the timeout value is ultimately determined and how it interacts with these variables I'm looking at. Let's see how that timer is initialized and what happens when `hard_timo` is zero.
|
| 116/2 |
2026/03/10 00:56 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_event.c IncludeLines:true Name:rxrpc_start_call_timer]
Results: map[SourceCode: 244: /*
245: * Initiate the call ack/resend/expiry timer.
246: */
247: void rxrpc_start_call_timer(struct rxrpc_call *call)
248: {
249: if (call->hard_timo) {
250: ktime_t delay = ms_to_ktime(call->hard_timo * 1000);
251:
252: call->expect_term_by = ktime_add(ktime_get_real(), delay);
253: trace_rxrpc_timer_set(call, delay, rxrpc_timer_trace_hard);
254: }
255: call->timer.expires = jiffies;
256: }
SourceFile:net/rxrpc/call_object.c]
|
| 117/2 |
2026/03/10 00:56 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 109277
output: 47
thoughts: 108
Thoughts: **Analyzing the RxRPC Call Timeout Mechanism**
Okay, let's break this down. My initial observation is key: if `call->hard_timo` is zero, then `call->expect_term_by` is set to `KTIME_MAX`. That's a clever way to signal *no* hard timeout. Makes sense. But that begs the question: what about the other timeouts? I've got to look at those.
I see that `call->next_rx_timo` is 20 times `HZ`, and `call->next_req_timo` is 1 times `HZ`. Interesting... so there are clearly other timers at play here. These are the "soft" timers, if you will. The Rx timer is 20 times the jiffies, which makes sense. The Request timer is only 1 jiffy, which seems quite brief. I wonder when that timer is activated. I had better go check `rxrpc_call_timer_expired` to see what is going on there. I'll need to figure out how these timers are used to determine if the call has failed or needs to be retried.
|
| 118/2 |
2026/03/10 00:56 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_event.c IncludeLines:true Name:rxrpc_call_timer_expired]
Results: map[SourceCode: 65: static void rxrpc_call_timer_expired(struct timer_list *t)
66: {
67: struct rxrpc_call *call = timer_container_of(call, t, timer);
68:
69: _enter("%d", call->debug_id);
70:
71: if (!__rxrpc_call_is_complete(call)) {
72: trace_rxrpc_timer_expired(call);
73: rxrpc_poke_call(call, rxrpc_call_poke_timer);
74: }
75: }
SourceFile:net/rxrpc/call_object.c]
|
| 119/2 |
2026/03/10 00:56 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 109875
output: 47
thoughts: 63
Thoughts: **Analyzing the Timer Expiration and Call Event Handling**
Okay, so I'm tracing through this RxRPC code, and I've just hit a crucial point. It looks like when the `rxrpc_call_timer_expired` function gets triggered, it, in turn, calls `rxrpc_poke_call(call, rxrpc_call_poke_timer)`. Right. That makes sense; it's poking the call structure, presumably to indicate the timer's expired.
But then, immediately following that, we have `rxrpc_input_call_event` being called. That's the key. I need to understand the interaction here. I need to see *how* the timer expiration, signaled by the poke, is processed as an *input* event. Let's dig into `rxrpc_input_call_event` now. That's where the magic happens, or perhaps where things break down!
|
| 120/2 |
2026/03/10 00:56 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/call_event.c IncludeLines:true Name:rxrpc_input_call_event]
Results: map[SourceCode: 297: /*
298: * Handle retransmission and deferred ACK/abort generation.
299: */
300: bool rxrpc_input_call_event(struct rxrpc_call *call)
301: {
302: struct sk_buff *skb;
303: ktime_t now, t;
304: bool did_receive = false, saw_ack = false;
305: s32 abort_code;
306:
307: rxrpc_see_call(call, rxrpc_call_see_input);
308:
309: //printk("\n--------------------\n");
310: _enter("{%d,%s,%lx}",
311: call->debug_id, rxrpc_call_states[__rxrpc_call_state(call)],
312: call->events);
313:
314: /* Handle abort request locklessly, vs rxrpc_propose_abort(). */
315: abort_code = smp_load_acquire(&call->send_abort);
316: if (abort_code) {
317: rxrpc_abort_call(call, 0, call->send_abort, call->send_abort_err,
318: call->send_abort_why);
319: goto out;
320: }
321:
322: do {
323: skb = __skb_dequeue(&call->rx_queue);
324: if (skb) {
325: struct rxrpc_skb_priv *sp = rxrpc_skb(skb);
326:
327: if (__rxrpc_call_is_complete(call) ||
328: skb->mark == RXRPC_SKB_MARK_ERROR) {
329: rxrpc_free_skb(skb, rxrpc_skb_put_call_rx);
330: goto out;
331: }
332:
333: saw_ack |= sp->hdr.type == RXRPC_PACKET_TYPE_ACK;
334:
335: rxrpc_input_call_packet(call, skb);
336: rxrpc_free_skb(skb, rxrpc_skb_put_call_rx);
337: did_receive = true;
338: }
339:
340: t = ktime_sub(call->rack_timo_at, ktime_get_real());
341: if (t <= 0) {
342: trace_rxrpc_timer_exp(call, t,
343: rxrpc_timer_trace_rack_off + call->rack_timer_mode);
344: call->rack_timo_at = KTIME_MAX;
345: rxrpc_rack_timer_expired(call, t);
346: }
347:
348: } while (!skb_queue_empty(&call->rx_queue));
349:
350: /* If we see our async-event poke, check for timeout trippage. */
351: now = ktime_get_real();
352: t = ktime_sub(call->expect_rx_by, now);
353: if (t <= 0) {
354: trace_rxrpc_timer_exp(call, t, rxrpc_timer_trace_expect_rx);
355: goto expired;
356: }
357:
358: t = ktime_sub(call->expect_req_by, now);
359: if (t <= 0) {
360: call->expect_req_by = KTIME_MAX;
361: if (__rxrpc_call_state(call) == RXRPC_CALL_SERVER_RECV_REQUEST) {
362: trace_rxrpc_timer_exp(call, t, rxrpc_timer_trace_idle);
363: goto expired;
364: }
365: }
366:
367: t = ktime_sub(READ_ONCE(call->expect_term_by), now);
368: if (t <= 0) {
369: trace_rxrpc_timer_exp(call, t, rxrpc_timer_trace_hard);
370: goto expired;
371: }
372:
373: t = ktime_sub(call->delay_ack_at, now);
374: if (t <= 0) {
375: trace_rxrpc_timer_exp(call, t, rxrpc_timer_trace_delayed_ack);
376: call->delay_ack_at = KTIME_MAX;
377: rxrpc_send_ACK(call, RXRPC_ACK_DELAY, 0,
378: rxrpc_propose_ack_delayed_ack);
379: }
380:
381: t = ktime_sub(call->ping_at, now);
382: if (t <= 0) {
383: trace_rxrpc_timer_exp(call, t, rxrpc_timer_trace_ping);
384: call->ping_at = KTIME_MAX;
385: rxrpc_send_ACK(call, RXRPC_ACK_PING, 0,
386: rxrpc_propose_ack_ping_for_keepalive);
387: }
388:
389: now = ktime_get_real();
390: t = ktime_sub(call->keepalive_at, now);
391: if (t <= 0) {
392: trace_rxrpc_timer_exp(call, t, rxrpc_timer_trace_keepalive);
393: call->keepalive_at = KTIME_MAX;
394: rxrpc_send_ACK(call, RXRPC_ACK_PING, 0,
395: rxrpc_propose_ack_ping_for_keepalive);
396: }
397:
398: if (test_and_clear_bit(RXRPC_CALL_EV_INITIAL_PING, &call->events))
399: rxrpc_send_initial_ping(call);
400:
401: rxrpc_transmit_some_data(call, UINT_MAX, rxrpc_txdata_new_data);
402:
403: if (saw_ack)
404: rxrpc_congestion_degrade(call);
405:
406: if (did_receive &&
407: (__rxrpc_call_state(call) == RXRPC_CALL_CLIENT_SEND_REQUEST ||
408: __rxrpc_call_state(call) == RXRPC_CALL_SERVER_SEND_REPLY)) {
409: t = ktime_sub(call->rack_timo_at, ktime_get_real());
410: trace_rxrpc_rack(call, t);
411: }
412:
413: /* Process events */
414: if (test_and_clear_bit(RXRPC_CALL_EV_ACK_LOST, &call->events))
415: rxrpc_send_ACK(call, RXRPC_ACK_PING, 0,
416: rxrpc_propose_ack_ping_for_lost_ack);
417:
418: if (call->tx_nr_lost > 0 &&
419: __rxrpc_call_state(call) != RXRPC_CALL_CLIENT_RECV_REPLY &&
420: !test_bit(RXRPC_CALL_TX_ALL_ACKED, &call->flags))
421: rxrpc_resend(call);
422:
423: if (test_and_clear_bit(RXRPC_CALL_RX_IS_IDLE, &call->flags))
424: rxrpc_send_ACK(call, RXRPC_ACK_IDLE, 0,
425: rxrpc_propose_ack_rx_idle);
426:
427: if (call->ackr_nr_unacked > 2) {
428: if (call->rtt_count < 3)
429: rxrpc_send_ACK(call, RXRPC_ACK_PING, 0,
430: rxrpc_propose_ack_ping_for_rtt);
431: else if (ktime_before(ktime_add_ms(call->rtt_last_req, 1000),
432: ktime_get_real()))
433: rxrpc_send_ACK(call, RXRPC_ACK_PING, 0,
434: rxrpc_propose_ack_ping_for_old_rtt);
435: else
436: rxrpc_send_ACK(call, RXRPC_ACK_IDLE, 0,
437: rxrpc_propose_ack_input_data);
438: }
439:
440: /* Make sure the timer is restarted */
441: if (!__rxrpc_call_is_complete(call)) {
442: ktime_t next = READ_ONCE(call->expect_term_by), delay;
443:
444: #define set(T) { ktime_t _t = (T); if (ktime_before(_t, next)) next = _t; }
445:
446: set(call->expect_req_by);
447: set(call->expect_rx_by);
448: set(call->delay_ack_at);
449: set(call->rack_timo_at);
450: set(call->keepalive_at);
451: set(call->ping_at);
452:
453: now = ktime_get_real();
454: delay = ktime_sub(next, now);
455: if (delay <= 0) {
456: rxrpc_poke_call(call, rxrpc_call_poke_timer_now);
457: } else {
458: unsigned long nowj = jiffies, delayj, nextj;
459:
460: delayj = umax(nsecs_to_jiffies(delay), 1);
461: nextj = nowj + delayj;
462: if (time_before(nextj, call->timer.expires) ||
463: !timer_pending(&call->timer)) {
464: trace_rxrpc_timer_restart(call, delay, delayj);
465: timer_reduce(&call->timer, nextj);
466: }
467: }
468: }
469:
470: out:
471: if (__rxrpc_call_is_complete(call)) {
472: timer_delete_sync(&call->timer);
473: if (!test_bit(RXRPC_CALL_DISCONNECTED, &call->flags))
474: rxrpc_disconnect_call(call);
475: if (call->security)
476: call->security->free_call_crypto(call);
477: } else {
478: if (did_receive &&
479: call->peer->ackr_adv_pmtud &&
480: call->peer->pmtud_pending)
481: rxrpc_send_probe_for_pmtud(call);
482: }
483: _leave("");
484: return true;
485:
486: expired:
487: if (test_bit(RXRPC_CALL_RX_HEARD, &call->flags) &&
488: (int)call->conn->hi_serial - (int)call->rx_serial > 0) {
489: trace_rxrpc_call_reset(call);
490: rxrpc_abort_call(call, 0, RX_CALL_DEAD, -ECONNRESET,
491: rxrpc_abort_call_reset);
492: } else {
493: rxrpc_abort_call(call, 0, RX_CALL_TIMEOUT, -ETIME,
494: rxrpc_abort_call_timeout);
495: }
496: goto out;
497: }
SourceFile:net/rxrpc/call_event.c]
|
| 121/2 |
2026/03/10 00:56 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 113711
output: 45
thoughts: 170
Thoughts: **Analyzing the Expiration Logic**
Okay, so I'm digging into this code and trying to understand how these timers work. My initial thought is, "Hang on, if `call->expect_term_by` is set to `KTIME_MAX`, that call will *never* expire." That seems pretty straightforward, but I need to make sure I haven't missed something.
Now, what about `call->expect_rx_by`? That's the other timer I need to consider. I know it's updated whenever a packet comes in, but that leaves a big question mark: what happens if no packet *is* received? Where is that initial value set? I quickly recall that `rxrpc_start_call_timer` only sets the termination timer.
I think `rxrpc_alloc_call` is the answer. Ah, yes, `call->expect_rx_by` is initialized to `KTIME_MAX` right there. So, it's set to expire forever if no packets are received. Okay, that makes more sense.
So when does it get set to something else? When we send data? Perhaps it's set in the transmission function? I think it must be in `rxrpc_transmit_some_data`. Let's confirm that.
|
| 122/2 |
2026/03/10 00:56 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/output.c IncludeLines:true Name:rxrpc_transmit_some_data]
Results: map[SourceCode: 259: void rxrpc_transmit_some_data(struct rxrpc_call *call, unsigned int limit,
260: enum rxrpc_txdata_trace trace)
261: {
262: switch (__rxrpc_call_state(call)) {
263: case RXRPC_CALL_SERVER_ACK_REQUEST:
264: if (call->tx_bottom == READ_ONCE(call->send_top))
265: return;
266: rxrpc_begin_service_reply(call);
267: fallthrough;
268:
269: case RXRPC_CALL_SERVER_SEND_REPLY:
270: case RXRPC_CALL_CLIENT_SEND_REQUEST:
271: if (!rxrpc_tx_window_space(call))
272: return;
273: if (call->tx_bottom == READ_ONCE(call->send_top)) {
274: rxrpc_inc_stat(call->rxnet, stat_tx_data_underflow);
275: return;
276: }
277: rxrpc_transmit_fresh_data(call, limit, trace);
278: break;
279: default:
280: return;
281: }
282: }
SourceFile:net/rxrpc/call_event.c]
|
| 123/2 |
2026/03/10 00:56 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 114602
output: 45
thoughts: 103
Thoughts: Okay, let me think this through.
**Analyzing the Timer Logic in rxrpc_transmit_fresh_data**
Alright, I need to get this straight in my head. If `call->expect_rx_by` and `call->expect_term_by` are both set to `KTIME_MAX`, it's screaming "this timer will never expire!". That's a huge red flag. Something isn't right if we're relying on a timer that's effectively disabled.
But, before I jump to conclusions, what about `call->next_rx_timo`? What's its relationship here? Ah, yes, I seem to recall that `call->next_rx_timo` is used to update `call->expect_rx_by`, specifically when we transmit a packet. Okay, so that must be where we're actually setting the expected receive time.
Let me trace this through. Okay, to be sure, the key is understanding how `rxrpc_transmit_fresh_data` works. I need to refresh my memory and step through that function's code again to fully grasp the logic. I need to see how the "send" path interacts with the timing.
|
| 124/2 |
2026/03/10 00:56 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/output.c IncludeLines:true Name:rxrpc_transmit_fresh_data]
Results: map[SourceCode: 191: /*
192: * Transmit some as-yet untransmitted data, to a maximum of the supplied limit.
193: */
194: static void rxrpc_transmit_fresh_data(struct rxrpc_call *call, unsigned int limit,
195: enum rxrpc_txdata_trace trace)
196: {
197: int space = rxrpc_tx_window_space(call);
198:
199: if (!test_bit(RXRPC_CALL_EXPOSED, &call->flags)) {
200: if (call->send_top == call->tx_top)
201: return;
202: rxrpc_expose_client_call(call);
203: }
204:
205: while (space > 0) {
206: struct rxrpc_send_data_req req = {
207: .now = ktime_get_real(),
208: .seq = call->tx_transmitted + 1,
209: .n = 0,
210: .trace = trace,
211: };
212: struct rxrpc_txqueue *tq;
213: struct rxrpc_txbuf *txb;
214: rxrpc_seq_t send_top, seq;
215: int limit = min(space, max(call->peer->pmtud_jumbo, 1));
216:
217: /* Order send_top before the contents of the new txbufs and
218: * txqueue pointers
219: */
220: send_top = smp_load_acquire(&call->send_top);
221: if (call->tx_top == send_top)
222: break;
223:
224: trace_rxrpc_transmit(call, send_top, space);
225:
226: tq = call->tx_qtail;
227: seq = call->tx_top;
228: trace_rxrpc_tq(call, tq, seq, rxrpc_tq_decant);
229:
230: do {
231: int ix;
232:
233: seq++;
234: ix = seq & RXRPC_TXQ_MASK;
235: if (!ix) {
236: tq = tq->next;
237: trace_rxrpc_tq(call, tq, seq, rxrpc_tq_decant_advance);
238: }
239: if (!req.tq)
240: req.tq = tq;
241: txb = tq->bufs[ix];
242: req.n++;
243: if (!txb->jumboable)
244: break;
245: } while (req.n < limit && before(seq, send_top));
246:
247: if (txb->flags & RXRPC_LAST_PACKET) {
248: rxrpc_close_tx_phase(call);
249: tq = NULL;
250: }
251: call->tx_qtail = tq;
252: call->tx_top = seq;
253:
254: space -= req.n;
255: rxrpc_send_data_packet(call, &req);
256: }
257: }
SourceFile:net/rxrpc/call_event.c]
|
| 125/2 |
2026/03/10 00:56 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 116039
output: 45
thoughts: 22
Thoughts: Here's my understanding of the thought process, summarized as you requested:
**Digging into the `rxrpc_send_data_packet` Function**
Okay, so I'm trying to wrap my head around this RxRPC data packet transmission. The code mentions `rxrpc_send_data_packet`, and it's making a specific call that I need to investigate further. It seems critical for managing the expected response behavior. Now, I'm thinking, I need to focus on what happens within `rxrpc_send_data_packet` with that `call->expect_rx_by` field. It's becoming obvious that the function is responsible for more than just sending the data; it’s setting a deadline or a condition on when we anticipate receiving a response. Understanding this deadline is going to be vital for diagnosing timeouts, retransmissions, or potential issues with the connection or the receiving end. I need to understand how `call->expect_rx_by` is being utilized to manage the lifecycle of this data packet.
|
| 126/2 |
2026/03/10 00:56 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:net/rxrpc/output.c IncludeLines:true Name:rxrpc_send_data_packet]
Results: map[SourceCode: 636: /*
637: * Send one or more packets through the transport endpoint
638: */
639: void rxrpc_send_data_packet(struct rxrpc_call *call, struct rxrpc_send_data_req *req)
640: {
641: struct rxrpc_wire_header *whdr;
642: struct rxrpc_connection *conn = call->conn;
643: enum rxrpc_tx_point frag;
644: struct rxrpc_txqueue *tq = req->tq;
645: struct rxrpc_txbuf *txb;
646: struct msghdr msg;
647: rxrpc_seq_t seq = req->seq;
648: size_t len = sizeof(*whdr);
649: bool new_call = test_bit(RXRPC_CALL_BEGAN_RX_TIMER, &call->flags);
650: int ret, stat_ix;
651:
652: _enter("%x,%x-%x", tq->qbase, seq, seq + req->n - 1);
653:
654: whdr = page_frag_alloc(&call->local->tx_alloc, sizeof(*whdr), GFP_NOFS);
655: if (!whdr)
656: return; /* Drop the packet if no memory. */
657:
658: call->local->kvec[0].iov_base = whdr;
659: call->local->kvec[0].iov_len = sizeof(*whdr);
660:
661: stat_ix = umin(req->n, ARRAY_SIZE(call->rxnet->stat_tx_jumbo)) - 1;
662: atomic_inc(&call->rxnet->stat_tx_jumbo[stat_ix]);
663:
664: len += rxrpc_prepare_data_packet(call, req, whdr);
665: txb = tq->bufs[seq & RXRPC_TXQ_MASK];
666:
667: iov_iter_kvec(&msg.msg_iter, WRITE, call->local->kvec, 1 + req->n, len);
668:
669: msg.msg_name = &call->peer->srx.transport;
670: msg.msg_namelen = call->peer->srx.transport_len;
671: msg.msg_control = NULL;
672: msg.msg_controllen = 0;
673: msg.msg_flags = MSG_SPLICE_PAGES;
674:
675: /* Send the packet with the don't fragment bit set unless we think it's
676: * too big or if this is a retransmission.
677: */
678: if (seq == call->tx_transmitted + 1 &&
679: len >= sizeof(struct rxrpc_wire_header) + call->peer->max_data) {
680: rxrpc_local_dont_fragment(conn->local, false);
681: frag = rxrpc_tx_point_call_data_frag;
682: } else {
683: rxrpc_local_dont_fragment(conn->local, true);
684: frag = rxrpc_tx_point_call_data_nofrag;
685: }
686:
687: /* Track what we've attempted to transmit at least once so that the
688: * retransmission algorithm doesn't try to resend what we haven't sent
689: * yet.
690: */
691: if (seq == call->tx_transmitted + 1)
692: call->tx_transmitted = seq + req->n - 1;
693:
694: if (IS_ENABLED(CONFIG_AF_RXRPC_INJECT_LOSS)) {
695: static int lose;
696:
697: if ((lose++ & 7) == 7) {
698: ret = 0;
699: trace_rxrpc_tx_data(call, txb->seq, txb->serial, txb->flags,
700: rxrpc_txdata_inject_loss);
701: rxrpc_peer_mark_tx(conn->peer);
702: goto done;
703: }
704: }
705:
706: /* send the packet by UDP
707: * - returns -EMSGSIZE if UDP would have to fragment the packet
708: * to go out of the interface
709: * - in which case, we'll have processed the ICMP error
710: * message and update the peer record
711: */
712: rxrpc_inc_stat(call->rxnet, stat_tx_data_send);
713: ret = do_udp_sendmsg(conn->local->socket, &msg, len);
714: rxrpc_peer_mark_tx(conn->peer);
715:
716: if (ret == -EMSGSIZE) {
717: rxrpc_inc_stat(call->rxnet, stat_tx_data_send_msgsize);
718: trace_rxrpc_tx_packet(call->debug_id, whdr, frag);
719: ret = 0;
720: } else if (ret < 0) {
721: rxrpc_inc_stat(call->rxnet, stat_tx_data_send_fail);
722: trace_rxrpc_tx_fail(call->debug_id, txb->serial, ret, frag);
723: } else {
724: trace_rxrpc_tx_packet(call->debug_id, whdr, frag);
725: }
726:
727: rxrpc_tx_backoff(call, ret);
728:
729: if (ret < 0) {
730: /* Cancel the call if the initial transmission fails or if we
731: * hit due to network routing issues that aren't going away
732: * anytime soon. The layer above can arrange the
733: * retransmission.
734: */
735: if (new_call ||
736: ret == -ENETUNREACH ||
737: ret == -EHOSTUNREACH ||
738: ret == -ECONNREFUSED)
739: rxrpc_set_call_completion(call, RXRPC_CALL_LOCAL_ERROR,
740: RX_USER_ABORT, ret);
741: }
742:
743: done:
744: _leave(" = %d [%u]", ret, call->peer->max_data);
745: }
SourceFile:net/rxrpc/output.c]
|
| 127/2 |
2026/03/10 00:56 |
llm |
crash-repro-finder |
0m
Model:
Tokens:
input: 118264
output: 55
thoughts: 39
Thoughts: Okay, here's my take on those thoughts, framed as an expert's internal monologue:
**Investigating the `rxrpc_send_data_packet` Function and `call->expect_rx_by`**
Alright, hang on a second. I'm deep in this RX-RPC code, and I've hit a snag. I was under the impression that the `rxrpc_send_data_packet` function, which is, you know, _supposed_ to handle sending our data packets, also had the crucial responsibility of setting the `call->expect_rx_by` field. This field, as I understand it, represents the point in time – or more accurately, the sequence number – by which we expect a response to this specific packet. It's a key part of our flow control and retransmission strategy, right? Without it, how can we reliably know when to time out or retransmit?
But now... I'm looking at the code again, and it doesn't seem to be setting it. *Damn*. Where the hell is this field actually getting populated? I need to track down exactly which function is responsible for setting `call->expect_rx_by`. I have to trace back the code paths carefully, and I may have to check the data flow itself. This could be a critical bug. I need to figure out if it's being set elsewhere, if something's missing entirely, or if there's a misunderstanding of the intended behavior on my end. This is a potential reliability issue, so I need to find it fast.
|
| 128/2 |
2026/03/10 00:57 |
tool |
codesearch-find-references |
0m
Args: map[ContextFile:net/rxrpc/output.c IncludeSnippetLines:1 Name:rxrpc_call::expect_rx_by SourceTreePrefix:]
Results: map[References:[map[ReferenceKind:writes ReferencingEntityKind:function ReferencingEntityName:rxrpc_alloc_call SourceFile:net/rxrpc/call_object.c SourceLine:165 SourceSnippet: 164: call->keepalive_at = KTIME_MAX;
165: call->expect_rx_by = KTIME_MAX;
166: call->expect_req_by = KTIME_MAX;
] map[ReferenceKind:reads ReferencingEntityKind:function ReferencingEntityName:rxrpc_call_seq_show SourceFile:net/rxrpc/proc.c SourceLine:80 SourceSnippet: 79: if (state != RXRPC_CALL_SERVER_PREALLOC)
80: timeout = ktime_ms_delta(READ_ONCE(call->expect_rx_by), ktime_get_real());
81:
] map[ReferenceKind:reads ReferencingEntityKind:function ReferencingEntityName:rxrpc_call_seq_show SourceFile:net/rxrpc/proc.c SourceLine:80 SourceSnippet: 79: if (state != RXRPC_CALL_SERVER_PREALLOC)
80: timeout = ktime_ms_delta(READ_ONCE(call->expect_rx_by), ktime_get_real());
81:
] map[ReferenceKind:reads ReferencingEntityKind:function ReferencingEntityName:rxrpc_call_seq_show SourceFile:net/rxrpc/proc.c SourceLine:80 SourceSnippet: 79: if (state != RXRPC_CALL_SERVER_PREALLOC)
80: timeout = ktime_ms_delta(READ_ONCE(call->expect_rx_by), ktime_get_real());
81:
] map[ReferenceKind:reads ReferencingEntityKind:function ReferencingEntityName:rxrpc_call_seq_show SourceFile:net/rxrpc/proc.c SourceLine:80 SourceSnippet: 79: if (state != RXRPC_CALL_SERVER_PREALLOC)
80: timeout = ktime_ms_delta(READ_ONCE(call->expect_rx_by), ktime_get_real());
81:
] map[ReferenceKind:reads ReferencingEntityKind:function ReferencingEntityName:rxrpc_call_seq_show SourceFile:net/rxrpc/proc.c SourceLine:80 SourceSnippet: 79: if (state != RXRPC_CALL_SERVER_PREALLOC)
80: timeout = ktime_ms_delta(READ_ONCE(call->expect_rx_by), ktime_get_real());
81:
] map[ReferenceKind:reads ReferencingEntityKind:function ReferencingEntityName:rxrpc_call_seq_show SourceFile:net/rxrpc/proc.c SourceLine:80 SourceSnippet: 79: if (state != RXRPC_CALL_SERVER_PREALLOC)
80: timeout = ktime_ms_delta(READ_ONCE(call->expect_rx_by), ktime_get_real());
81:
] map[ReferenceKind:takes-address-of ReferencingEntityKind:function ReferencingEntityName:rxrpc_call_seq_show SourceFile:net/rxrpc/proc.c SourceLine:80 SourceSnippet: 79: if (state != RXRPC_CALL_SERVER_PREALLOC)
80: timeout = ktime_ms_delta(READ_ONCE(call->expect_rx_by), ktime_get_real());
81:
] map[ReferenceKind:reads ReferencingEntityKind:function ReferencingEntityName:rxrpc_input_call_event SourceFile:net/rxrpc/call_event.c SourceLine:352 SourceSnippet: 351: now = ktime_get_real();
352: t = ktime_sub(call->expect_rx_by, now);
353: if (t <= 0) {
] map[ReferenceKind:reads ReferencingEntityKind:function ReferencingEntityName:rxrpc_input_call_event SourceFile:net/rxrpc/call_event.c SourceLine:447 SourceSnippet: 446: set(call->expect_req_by);
447: set(call->expect_rx_by);
448: set(call->delay_ack_at);
] map[ReferenceKind:writes ReferencingEntityKind:function ReferencingEntityName:rxrpc_input_call_packet SourceFile:net/rxrpc/input.c SourceLine:1248 SourceSnippet:1247:
1248: call->expect_rx_by = ktime_add(ktime_get_real(), delay);
1249: trace_rxrpc_timer_set(call, delay, rxrpc_timer_trace_expect_rx);
] map[ReferenceKind:writes ReferencingEntityKind:function ReferencingEntityName:rxrpc_prepare_data_packet SourceFile:net/rxrpc/output.c SourceLine:627 SourceSnippet: 626:
627: call->expect_rx_by = ktime_add(req->now, delay);
628: trace_rxrpc_timer_set(call, delay, rxrpc_timer_trace_expect_rx);
]] TruncatedOutput:false]
|
| 129/2 |
2026/03/10 00:57 |
llm |
crash-repro-finder |
0m
Model:
Error:
|