| 0/0 |
2026/04/21 11:15 |
flow |
assessment-kcsan |
17m
Results: map[Benign:true Confident:true Explanation:The data race occurs between **`ktime_get_coarse_real_ts64_mg()`** and
**`timekeeping_update_from_shadow()`** on the field **`tk->coarse_nsec`** (part
of `struct timekeeper`).
### Synchronization Context
* **Reader:** `ktime_get_coarse_real_ts64_mg()` is a lockless reader that uses
a sequence counter (`tk_core.seq`) via `read_seqcount_begin()` and
`read_seqcount_retry()` to ensure data consistency.
* **Writer:** `timekeeping_update_from_shadow()` is called in interrupt
context (timer tick) and holds the `tk_core.lock` spinlock. It also increments
the sequence counter via `write_seqcount_begin()` and `write_seqcount_end()`
around a `memcpy()` that updates the live timekeeper.
The race is possible because the reader does not hold the spinlock, and both
sides use plain memory accesses (or `memcpy`) for the data itself.
### Final Classification: **BENIGN**
### Reasoning
The race is structurally benign because the algorithm is designed to be tolerant
of stale or torn reads through the use of a **sequence counter (seqlock)**.
1. **Retry Mechanism:** If the reader (`CPU 0`) reads `coarse_nsec` while the
writer (`CPU 1`) is executing the `memcpy()`, the reader might see a partially
updated (torn) value. However, the `read_seqcount_retry()` call at the end of
the loop will detect that a write was in progress (sequence number was odd) or
has completed (sequence number changed), forcing the reader to discard the
result and retry the loop.
2. **Data Safety:** The racing variable is a simple integer representing
nanoseconds. A torn read of this value does not lead to a crash, invalid pointer
dereference, or out-of-bounds access. It is only used to populate a `timespec64`
structure, which is then validated against a "floor" value and potentially
discarded if the retry fails.
3. **Ordering:** The `read_seqcount_begin()` and `read_seqcount_retry()`
primitives include the necessary memory barriers (`smp_load_acquire` and
`smp_rmb`) to prevent the compiler from hoisting the reads out of the protected
loop.
### Recommended Annotations
To silence the KCSAN report, the plain accesses inside the `seqcount` loop
should be wrapped in `data_race()` or converted to `READ_ONCE()`. Specifically,
in `tk_xtime_coarse()`:
```c
static inline struct timespec64 tk_xtime_coarse(const struct timekeeper *tk)
{
struct timespec64 ts;
ts.tv_sec = data_race(tk->xtime_sec);
ts.tv_nsec = data_race(tk->coarse_nsec);
return ts;
}
```
And similarly for the `offs_real` access in `ktime_get_coarse_real_ts64_mg()`.]
|
| 1/1 |
2026/04/21 11:15 |
action |
kernel-checkouter |
0m
Results: map[KernelSrc:/app/workdir/cache/src/a6dd3c7c36afbd009cb1a2e2b3c929bc0d13fb93]
|
| 2/1 |
2026/04/21 11:15 |
action |
kernel-builder |
0m
Results: map[KernelObj:/app/workdir/cache/build/5d2cd1e11e85298745d218bb92edf43ba514519f]
|
| 3/1 |
2026/04/21 11:15 |
action |
codesearch-prepare |
0m
Results: map[Index:codesearch-index]
|
| 4/1 |
2026/04/21 11:15 |
agent |
expert |
17m
Model:
Results: map[Benign:true Confident:true]
Instruction: # KCSAN Data Race Severity Analysis Protocol
You are an expert Linux kernel concurrency engineer analyzing a Kernel
Concurrency Sanitizer (KCSAN) data race report to determine if it is
**BENIGN** or **HARMFUL**.
## 1. CLASSIFICATIONS
### **BENIGN (Truly Benign)**
The logic is sound and structurally tolerant to compiler optimizations or
stale/torn reads.
- **Diagnostics/Stats:** Reads used only for '/proc', '/sys', counters, or
'pr_info'.
- **Heuristic Hints:** A "hint" flag where an old value only causes a
slightly delayed update or a sub-optimal but safe fast-path.
- **Single-Writer Flag Updates:** A single writer updating flags where the
concurrent read is a simple bitwise check (e.g., 'flags & MASK'). These are
historically tolerated, assuming neither "Fused Accesses" nor "Ordering
Violations" are relevant in this context.
- **Marked Reloads:** A load feeding into a 'cmpxchg()' loop or checked
against a later 'READ_ONCE()' reload.
- **Safe Overwrites:** Writing the same value already present.
### **HARMFUL (Logic Bug or Marking Required)**
The race causes incorrect behavior due to a synchronization failure or
because missing annotations allow the compiler to break the algorithm.
**Marking Required for Correctness:**
The algorithm is logically sound but requires annotations ('READ_ONCE()',
'WRITE_ONCE()', 'smp_load_acquire()', 'smp_store_release()', etc.) to be safe.
- **Fused Accesses:** The compiler might merge accesses or hoist a load out
of a loop, breaking polling/wait loops (livelocks).
- **Torn Accesses:** A large access (e.g., 64-bit on 32-bit arch) might be
split into multiple non-atomic accesses. Note that 'READ_ONCE()' does **not**
guarantee atomicity for 64-bit variables on 32-bit architectures.
- **Ordering Violations:** The race breaks a "happens-before" relationship
(requires primitives with implied or explicit memory barriers).
**Logic Bugs:**
A fundamental synchronization failure. Marking accesses will **not** fix it;
the logic itself must change.
- **Pointers/Lifecycle:** The racing variable is a pointer being dereferenced
or a refcount governing object lifecycle (Use-After-Free risk).
- **Control Flow:** The variable guards a critical section, memory allocation,
or hardware command.
- **Bitfields:** Concurrent writes to different bits in the same word.
Compilers often use non-atomic read-modify-write sequences, meaning a
write to 'bit_A' can "clobber" a concurrent write to 'bit_B'. However,
do not blindly assume all bitfield accesses are harmful; you must prove
that a concurrent write actually clobbers another in a way that breaks
logic.
- **Complex Structures:** Races on shared lists, trees, or hashmaps.
- **Lossy Updates:** Concurrent plain RMW operations (e.g., 'var++') on
non-diagnostic variables where every increment must be preserved.
- **State Machines:** Races allowing a state machine to bypass transitions
or enter an invalid state.
- **Adjacent Unsynchronized Operations:** Consider races happening at the
same time. For example, if both threads execute 'struct->has_elements = true;
list_add(node, &struct->list);', the race on 'has_elements' implies an
adjacent race on 'list_head', which is HARMFUL.
## 2. RESEARCH & ANALYSIS WORKFLOW
1. **Locate the Race:** Find the exact variables and functions in the stack
traces. **Use codesearch tools to read the actual source code and
confirm all assumptions.** Do not speculate about hypothetical compiler
behaviors or theoretical dangers (e.g., dismissing something as
"fundamentally unsafe") without tracing the actual data flow to a crash.
2. **Contextualize:** Identify held locks, RCU sections, or interrupt
contexts. Explain why the race is possible (e.g., "Thread A holds 'lock',
but Thread B is a lockless reader").
3. **Data Flow:** Follow the racing variable into its subsequent uses. If
the reader sees a stale or torn value, what branch is taken? Does it lead
to a pointer dereference or 'BUG_ON()'? **You must prove harm via this
data flow.**
4. **Failure Sequencing (if HARMFUL):** Construct a two-column execution trace
demonstrating the failure.
*Example format:*
CPU0 CPU1
function_a()
// lockless read
ptr = global_ptr
<PREEMPT>
function_b()
lock(&my_lock)
global_ptr = NULL
unlock(&my_lock)
kfree(ptr)
if (ptr)
*ptr = 1; // -> Use-After-Free!
## 3. OUTPUT FORMAT
- **Race Summary:** '[Function A]' vs '[Function B]' on '[Variable/Field]'.
- **Synchronization context:** List held locks and explain the lack of mutual
exclusion.
- **Final Classification:** **BENIGN** or **HARMFUL**.
**If BENIGN:**
- **Reasoning:** Briefly explain why the race is structurally tolerant to stale
or torn reads.
- **Recommended Annotations:** Suggest appropriate annotations (e.g.,
'READ_ONCE()', 'data_race()').
**If HARMFUL:**
- **Failure Reasoning:** Explain and provide a two-column interleaving showing
exactly how the race leads to a failure such as corruption or kernel crash.
- **Recommended Fix (ONLY if fix is trivial):** Suggest a structural fix (e.g.,
"Hold 'mapping->i_pages' lock", "Convert to 'atomic_t'") or required memory
ordering annotations (e.g., "Wrap in 'READ_ONCE()'", "Use
'smp_load_acquire()'").
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt:
The data race report is:
==================================================================
BUG: KCSAN: data-race in ktime_get_coarse_real_ts64_mg / timekeeping_update_from_shadow
write to 0xffffffff895d6488 of 304 bytes by interrupt on cpu 1:
timekeeping_update_from_shadow+0x40d/0x440 kernel/time/timekeeping.c:829
__timekeeping_advance+0xa5d/0xc10 kernel/time/timekeeping.c:2532
timekeeping_advance kernel/time/timekeeping.c:2540 [inline]
update_wall_time+0x21/0x50 kernel/time/timekeeping.c:2550
tick_do_update_jiffies64+0x169/0x1c0 kernel/time/tick-sched.c:149
tick_sched_do_timer kernel/time/tick-sched.c:253 [inline]
tick_nohz_handler+0x8d/0x3d0 kernel/time/tick-sched.c:312
__run_hrtimer kernel/time/hrtimer.c:1930 [inline]
__hrtimer_run_queues+0x276/0x4f0 kernel/time/hrtimer.c:1994
hrtimer_interrupt+0x261/0x850 kernel/time/hrtimer.c:2113
local_apic_timer_interrupt arch/x86/kernel/apic/apic.c:1050 [inline]
__sysvec_apic_timer_interrupt+0x5f/0x1c0 arch/x86/kernel/apic/apic.c:1067
instr_sysvec_apic_timer_interrupt arch/x86/kernel/apic/apic.c:1061 [inline]
sysvec_apic_timer_interrupt+0x6f/0x80 arch/x86/kernel/apic/apic.c:1061
asm_sysvec_apic_timer_interrupt+0x1a/0x20 arch/x86/include/asm/idtentry.h:697
arch_atomic64_read arch/x86/include/asm/atomic64_64.h:15 [inline]
raw_atomic64_read include/linux/atomic/atomic-arch-fallback.h:2583 [inline]
raw_atomic_long_read include/linux/atomic/atomic-long.h:38 [inline]
atomic_long_read include/linux/atomic/atomic-instrumented.h:3189 [inline]
find_watchpoint kernel/kcsan/core.c:131 [inline]
check_access kernel/kcsan/core.c:737 [inline]
__tsan_read4+0x2c/0x190 kernel/kcsan/core.c:1024
ns_of_pid include/linux/pid.h:153 [inline]
task_active_pid_ns kernel/pid.c:577 [inline]
pid_vnr+0x3b/0x100 kernel/pid.c:556
scm_set_cred include/net/scm.h:73 [inline]
unix_skb_to_scm net/unix/af_unix.c:2002 [inline]
__unix_dgram_recvmsg+0x5be/0x870 net/unix/af_unix.c:2628
unix_dgram_recvmsg+0x7e/0x90 net/unix/af_unix.c:2676
sock_recvmsg_nosec+0xc2/0xf0 net/socket.c:1137
____sys_recvmsg+0x26f/0x280 net/socket.c:2916
___sys_recvmsg+0x11f/0x3b0 net/socket.c:2960
do_recvmmsg+0x1ef/0x560 net/socket.c:3055
__sys_recvmmsg net/socket.c:3129 [inline]
__do_sys_recvmmsg net/socket.c:3152 [inline]
__se_sys_recvmmsg net/socket.c:3145 [inline]
__x64_sys_recvmmsg+0xe5/0x170 net/socket.c:3145
x64_sys_call+0x80f/0x3020 arch/x86/include/generated/asm/syscalls_64.h:300
do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline]
do_syscall_64+0x12c/0x3b0 arch/x86/entry/syscall_64.c:94
entry_SYSCALL_64_after_hwframe+0x77/0x7f
read to 0xffffffff895d64f8 of 4 bytes by task 12 on cpu 0:
tk_xtime_coarse kernel/time/timekeeping.c:212 [inline]
ktime_get_coarse_real_ts64_mg+0x89/0x1a0 kernel/time/timekeeping.c:2606
current_time+0x41/0x240 fs/inode.c:2370
need_cmtime_update fs/inode.c:2392 [inline]
file_update_time_flags+0x83/0x2e0 fs/inode.c:2413
file_update_time+0x17/0x20 fs/inode.c:2443
shmem_file_write_iter+0x9c/0xf0 mm/shmem.c:3475
lo_rw_aio+0x67d/0x730 drivers/block/loop.c:-1
do_req_filebacked drivers/block/loop.c:-1 [inline]
loop_handle_cmd drivers/block/loop.c:1925 [inline]
loop_process_work+0x56c/0xac0 drivers/block/loop.c:1960
loop_workfn+0x31/0x40 drivers/block/loop.c:1984
process_one_work kernel/workqueue.c:3302 [inline]
process_scheduled_works+0x4f0/0x9c0 kernel/workqueue.c:3385
worker_thread+0x58a/0x780 kernel/workqueue.c:3466
kthread+0x22a/0x280 kernel/kthread.c:436
ret_from_fork+0x146/0x330 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
value changed: 0x31af5fed -> 0x3247f66d
Reported by Kernel Concurrency Sanitizer on:
CPU: 0 UID: 0 PID: 12 Comm: kworker/u8:0 Not tainted syzkaller #0 PREEMPT(full)
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 03/18/2026
Workqueue: loop1 loop_workfn
==================================================================
Reply: The data race occurs between **`ktime_get_coarse_real_ts64_mg()`** and **`timekeeping_update_from_shadow()`** on the field **`tk->coarse_nsec`** (part of `struct timekeeper`).
### Synchronization Context
* **Reader:** `ktime_get_coarse_real_ts64_mg()` is a lockless reader that uses a sequence counter (`tk_core.seq`) via `read_seqcount_begin()` and `read_seqcount_retry()` to ensure data consistency.
* **Writer:** `timekeeping_update_from_shadow()` is called in interrupt context (timer tick) and holds the `tk_core.lock` spinlock. It also increments the sequence counter via `write_seqcount_begin()` and `write_seqcount_end()` around a `memcpy()` that updates the live timekeeper.
The race is possible because the reader does not hold the spinlock, and both sides use plain memory accesses (or `memcpy`) for the data itself.
### Final Classification: **BENIGN**
### Reasoning
The race is structurally benign because the algorithm is designed to be tolerant of stale or torn reads through the use of a **sequence counter (seqlock)**.
1. **Retry Mechanism:** If the reader (`CPU 0`) reads `coarse_nsec` while the writer (`CPU 1`) is executing the `memcpy()`, the reader might see a partially updated (torn) value. However, the `read_seqcount_retry()` call at the end of the loop will detect that a write was in progress (sequence number was odd) or has completed (sequence number changed), forcing the reader to discard the result and retry the loop.
2. **Data Safety:** The racing variable is a simple integer representing nanoseconds. A torn read of this value does not lead to a crash, invalid pointer dereference, or out-of-bounds access. It is only used to populate a `timespec64` structure, which is then validated against a "floor" value and potentially discarded if the retry fails.
3. **Ordering:** The `read_seqcount_begin()` and `read_seqcount_retry()` primitives include the necessary memory barriers (`smp_load_acquire` and `smp_rmb`) to prevent the compiler from hoisting the reads out of the protected loop.
### Recommended Annotations
To silence the KCSAN report, the plain accesses inside the `seqcount` loop should be wrapped in `data_race()` or converted to `READ_ONCE()`. Specifically, in `tk_xtime_coarse()`:
```c
static inline struct timespec64 tk_xtime_coarse(const struct timekeeper *tk)
{
struct timespec64 ts;
ts.tv_sec = data_race(tk->xtime_sec);
ts.tv_nsec = data_race(tk->coarse_nsec);
return ts;
}
```
And similarly for the `offs_real` access in `ktime_get_coarse_real_ts64_mg()`.
|
| 5/2 |
2026/04/21 11:15 |
llm |
expert |
0m
Model:
Tokens:
input: 5865
output: 22
thoughts: 151
Thoughts: **Analyzing the Data Race Report**
Okay, so the report flags a nasty data race – that always grabs my attention. It looks like we have a conflict between `ktime_get_coarse_real_ts64_mg` reading some data and `timekeeping_update_from_shadow` writing to it. The report specifies the memory addresses involved, `0xffffffff895d6488` for the write (304 bytes!) and `0xffffffff895d64f8` for the read (only 4 bytes, interesting).
The read happens within `tk_xtime_coarse`, which is conveniently inlined within the problematic `ktime_get_coarse_real_ts64_mg` function. This simplifies the investigation to a degree, as we know exactly where the read is occurring. Now, let's dive into the code and see what variables are being accessed and how these two functions interact at the memory addresses specified. We need to understand the memory layout and the synchronization mechanisms (or lack thereof) around this critical section. Let's see if this is a known issue, or if we are looking at something new.
|
| 6/2 |
2026/04/21 11:15 |
tool |
codesearch-dir-index |
0m
Args:
Results: map[Files:[alarmtimer.c clockevents.c clocksource-wdtest.c clocksource.c hrtimer.c itimer.c jiffies.c namespace.c namespace_internal.h namespace_vdso.c ntp.c ntp_internal.h posix-clock.c posix-cpu-timers.c posix-stubs.c posix-timers.c posix-timers.h sched_clock.c sleep_timeout.c test_udelay.c tick-broadcast-hrtimer.c tick-broadcast.c tick-common.c tick-internal.h tick-legacy.c tick-oneshot.c tick-sched.c tick-sched.h time.c time_test.c timeconv.c timecounter.c timekeeping.c timekeeping.h timekeeping_debug.c timekeeping_internal.h timer.c timer_list.c timer_migration.c timer_migration.h vsyscall.c] Subdirs:<nil>]
|
| 7/2 |
2026/04/21 11:15 |
llm |
expert |
0m
Model:
Tokens:
input: 6632
output: 52
thoughts: 50
Thoughts: Here's my analysis of the provided thought process, tailored for an expert audience:
**Diving into the Kernel Timekeeping Mechanisms**
Alright, let's get down to brass tacks. I'm focusing in on `kernel/time/timekeeping.c`. My first port of call, naturally, will be dissecting `ktime_get_coarse_real_ts64_mg` and then, by extension, `tk_xtime_coarse`. I expect that `ktime_get_coarse_real_ts64_mg` is using `tk_xtime_coarse` to get a fast but probably less precise time, likely suitable for userland applications that don't need absolute microsecond accuracy. I'll need to figure out what the MG in the name stands for; maybe machine-guarded or something similar that indicates some sort of concurrency handling. I'm keen to see if it's using the standard `xtime` variable or if there's any clever trickery or optimizations happening, perhaps utilizing a hardware timer or some other optimization to improve the performance of this "coarse" time retrieval. I'm also interested to know how coarse the time is; if it is useful for the system.
|
| 8/2 |
2026/04/21 11:15 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/time/timekeeping.c IncludeLines:true Name:ktime_get_coarse_real_ts64_mg]
Results: map[SourceCode:2588: /**
2589: * ktime_get_coarse_real_ts64_mg - return latter of coarse grained time or floor
2590: * @ts: timespec64 to be filled
2591: *
2592: * Fetch the global mg_floor value, convert it to realtime and compare it
2593: * to the current coarse-grained time. Fill @ts with whichever is
2594: * latest. Note that this is a filesystem-specific interface and should be
2595: * avoided outside of that context.
2596: */
2597: void ktime_get_coarse_real_ts64_mg(struct timespec64 *ts)
2598: {
2599: struct timekeeper *tk = &tk_core.timekeeper;
2600: u64 floor = atomic64_read(&mg_floor);
2601: ktime_t f_real, offset, coarse;
2602: unsigned int seq;
2603:
2604: do {
2605: seq = read_seqcount_begin(&tk_core.seq);
2606: *ts = tk_xtime_coarse(tk);
2607: offset = tk_core.timekeeper.offs_real;
2608: } while (read_seqcount_retry(&tk_core.seq, seq));
2609:
2610: coarse = timespec64_to_ktime(*ts);
2611: f_real = ktime_add(floor, offset);
2612: if (ktime_after(f_real, coarse))
2613: *ts = ktime_to_timespec64(f_real);
2614: }
SourceFile:kernel/time/timekeeping.c]
|
| 9/2 |
2026/04/21 11:15 |
llm |
expert |
0m
Model:
Tokens:
input: 7501
output: 43
thoughts: 9
Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience, and written in the first person:
**My Current Focus: Understanding `tk_xtime_coarse`**
Right now, my mind is completely absorbed by `tk_xtime_coarse`. It's a key function, and I'm dissecting its implications. I know its purpose, but I'm thinking beyond that now. What are the subtle trade-offs in its design? How does it balance accuracy with efficiency? I'm considering the context of its use within the broader kernel timing infrastructure. What's the relationship with other timekeeping mechanisms? Is it a crucial component in maintaining system responsiveness and stability, or does it serve a more specialized purpose?
I'm digging deeper into how it integrates with the hardware's clock sources and how any potential discrepancies are handled. I'm trying to visualize the various scenarios where it gets called and the impact of the return value. Are there any edge cases I might be overlooking? And most importantly, I'm thinking about whether it could be optimized in my specific area of interest. Any subtle improvements to it could translate to significant performance gains, especially in time-sensitive tasks. Essentially, I'm trying to picture the complete picture that this component paints, and what it means for the system as a whole.
|
| 10/2 |
2026/04/21 11:15 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/time/timekeeping.c IncludeLines:true Name:tk_xtime_coarse]
Results: map[SourceCode: 207: static inline struct timespec64 tk_xtime_coarse(const struct timekeeper *tk)
208: {
209: struct timespec64 ts;
210:
211: ts.tv_sec = tk->xtime_sec;
212: ts.tv_nsec = tk->coarse_nsec;
213: return ts;
214: }
SourceFile:kernel/time/timekeeping.c]
|
| 11/2 |
2026/04/21 11:15 |
llm |
expert |
0m
Model:
Tokens:
input: 7978
output: 34
thoughts: 129
Thoughts: **Analyzing a Potential Race Condition**
Okay, so we've got a timing issue here, a classic race scenario. My gut's telling me there's contention related to `tk->coarse_nsec`. We're dealing with a write and a read, and the address offsets are raising some red flags. The write is happening at `0xffffffff895d6488`, and it's a 304-byte operation. The read is targeting `0xffffffff895d64f8`, and it's just a 4-byte access.
Now, the crucial calculation: the difference between the read and write addresses. Let's see... `0xffffffff895d64f8 - 0xffffffff895d6488` equals 0x70, which is 112 in decimal. That 0x70 offset is the key because my immediate suspicion is that it corresponds to the location of `tk->coarse_nsec` within the `struct timekeeper`.
Now, the real test is confirming that suspicion. I need to refresh my memory of the `struct timekeeper` layout. Let's dig into the header file and confirm that `tk->coarse_nsec` is indeed at that offset. If it is, we've likely pinpointed the source of the race. We need to figure out why the read and write are happening concurrently and how to synchronize them.
|
| 12/2 |
2026/04/21 11:15 |
tool |
codesearch-struct-layout |
0m
Args: map[ContextFile:kernel/time/timekeeping.c Name:timekeeper]
Results: map[Fields:[map[Name:tkr_mono OffsetBits:0 SizeBits:448] map[Name:xtime_sec OffsetBits:448 SizeBits:64] map[Name:ktime_sec OffsetBits:512 SizeBits:64] map[Name:wall_to_monotonic OffsetBits:576 SizeBits:128] map[Name:offs_real OffsetBits:704 SizeBits:64] map[Name:offs_boot OffsetBits:768 SizeBits:64] map[Name: OffsetBits:832 SizeBits:64] map[Name:coarse_nsec OffsetBits:896 SizeBits:32] map[Name:id OffsetBits:928 SizeBits:32] map[Name:tkr_raw OffsetBits:960 SizeBits:448] map[Name:raw_sec OffsetBits:1408 SizeBits:64] map[Name:cs_id OffsetBits:1472 SizeBits:32] map[Name:cs_ns_to_cyc_mult OffsetBits:1504 SizeBits:32] map[Name:cs_ns_to_cyc_shift OffsetBits:1536 SizeBits:32] map[Name:cs_ns_to_cyc_maxns OffsetBits:1600 SizeBits:64] map[Name:clock_was_set_seq OffsetBits:1664 SizeBits:32] map[Name:cs_was_changed_seq OffsetBits:1696 SizeBits:8] map[Name:clock_valid OffsetBits:1704 SizeBits:8] map[Name: OffsetBits:1728 SizeBits:128] map[Name:cycle_interval OffsetBits:1856 SizeBits:64] map[Name:xtime_interval OffsetBits:1920 SizeBits:64] map[Name:xtime_remainder OffsetBits:1984 SizeBits:64] map[Name:raw_interval OffsetBits:2048 SizeBits:64] map[Name:next_leap_ktime OffsetBits:2112 SizeBits:64] map[Name:ntp_tick OffsetBits:2176 SizeBits:64] map[Name:ntp_error OffsetBits:2240 SizeBits:64] map[Name:ntp_error_shift OffsetBits:2304 SizeBits:32] map[Name:ntp_err_mult OffsetBits:2336 SizeBits:32] map[Name:skip_second_overflow OffsetBits:2368 SizeBits:32] map[Name:tai_offset OffsetBits:2400 SizeBits:32]]]
|
| 13/2 |
2026/04/21 11:15 |
llm |
expert |
0m
Model:
Tokens:
input: 9159
output: 40
thoughts: 82
Thoughts: **Analyzing a Timekeeping Write**
Okay, so I'm looking at this memory write, and the first thing that jumps out is that `coarse_nsec` is sitting pretty at an offset of 896 bits, which is exactly 112 bytes. That's a good starting point. Now, the write itself covers 304 bytes, beginning at address `0xffffffff895d6488`. Given the address, that looks highly probable to be the beginning of `tk_core.timekeeper`. It's time to check that out and confirm my suspicions about that address. I need to understand what's being modified within that 304-byte region to truly understand the write's implications.
|
| 14/2 |
2026/04/21 11:15 |
tool |
codesearch-definition-source |
0m
Error:
Args: map[ContextFile:kernel/time/timekeeping.c IncludeLines:true Name:tk_core]
Results: map[SourceCode: SourceFile:]
|
| 15/2 |
2026/04/21 11:15 |
llm |
expert |
0m
Model:
Tokens:
input: 9460
output: 27
thoughts: 22
Thoughts: Okay, here's my summary, taking the perspective of someone familiar with code and debugging:
**Debugging Thoughts: Potential Static Variable**
Alright, I'm getting a hunch about this issue. The behavior I'm seeing is pointing me in a direction, and I suspect it could be related to a static variable. They can be tricky little devils! The next logical step is to confirm that suspicion. Time to get a little more specific. I need to zero in on where this `tk_core` thing is happening. Let me quickly examine the relevant file content. I'll search for occurrences of `tk_core` and see what comes up. That should give me the information I need to confirm the issue.
|
| 16/2 |
2026/04/21 11:15 |
tool |
codesearch-file-index |
0m
Args: map[SourceFile:kernel/time/timekeeping.c]
Results: map[Entities:[map[Kind:function Name:__do_adjtimex] map[Kind:function Name:__ktime_get_fast_ns] map[Kind:function Name:__ktime_get_real_seconds] map[Kind:function Name:__timekeeping_advance] map[Kind:function Name:__timekeeping_inject_offset] map[Kind:function Name:__timekeeping_inject_sleeptime] map[Kind:function Name:__timekeeping_set_tai_offset] map[Kind:function Name:accumulate_nsecs_to_secs] map[Kind:function Name:adjust_historical_crosststamp] map[Kind:function Name:change_clocksource] map[Kind:function Name:clocksource_disable_inline_read] map[Kind:function Name:clocksource_enable_inline_read] map[Kind:function Name:convert_base_to_cs] map[Kind:function Name:convert_clock] map[Kind:function Name:convert_cs_to_base] map[Kind:function Name:convert_ns_to_cs] map[Kind:function Name:delta_to_ns_safe] map[Kind:function Name:do_adjtimex] map[Kind:function Name:do_settimeofday64] map[Kind:function Name:do_timer] map[Kind:function Name:dummy_clock_read] map[Kind:function Name:get_device_system_crosststamp] map[Kind:function Name:getboottime64] map[Kind:function Name:halt_fast_timekeeper] map[Kind:function Name:ktime_expiry_to_cycles] map[Kind:function Name:ktime_get] map[Kind:function Name:ktime_get_boot_fast_ns] map[Kind:function Name:ktime_get_clock_ts64] map[Kind:function Name:ktime_get_coarse_real_ts64] map[Kind:function Name:ktime_get_coarse_real_ts64_mg] map[Kind:function Name:ktime_get_coarse_ts64] map[Kind:function Name:ktime_get_coarse_with_offset] map[Kind:function Name:ktime_get_mono_fast_ns] map[Kind:function Name:ktime_get_ntp_seconds] map[Kind:function Name:ktime_get_raw] map[Kind:function Name:ktime_get_raw_fast_ns] map[Kind:function Name:ktime_get_raw_ts64] map[Kind:function Name:ktime_get_real_fast_ns] map[Kind:function Name:ktime_get_real_seconds] map[Kind:function Name:ktime_get_real_ts64] map[Kind:function Name:ktime_get_real_ts64_mg] map[Kind:function Name:ktime_get_resolution_ns] map[Kind:function Name:ktime_get_seconds] map[Kind:function Name:ktime_get_snapshot] map[Kind:function Name:ktime_get_tai_fast_ns] map[Kind:function Name:ktime_get_ts64] map[Kind:function Name:ktime_get_update_offsets_now] map[Kind:function Name:ktime_get_with_offset] map[Kind:function Name:ktime_mono_to_any] map[Kind:function Name:ktime_real_to_base_clock] map[Kind:function Name:logarithmic_accumulation] map[Kind:function Name:pvclock_gtod_register_notifier] map[Kind:function Name:pvclock_gtod_unregister_notifier] map[Kind:function Name:random_get_entropy_fallback] map[Kind:function Name:read_persistent_clock64] map[Kind:function Name:read_persistent_wall_and_boot_offset] map[Kind:function Name:scale64_check_overflow] map[Kind:function Name:timekeeper_is_core_tk] map[Kind:function Name:timekeeper_lock_irqsave] map[Kind:function Name:timekeeper_unlock_irqrestore] map[Kind:function Name:timekeeping_adjust] map[Kind:function Name:timekeeping_advance] map[Kind:function Name:timekeeping_apply_adjustment] map[Kind:function Name:timekeeping_clocksource_has_base] map[Kind:function Name:timekeeping_cycles_to_ns] map[Kind:function Name:timekeeping_forward_now] map[Kind:function Name:timekeeping_get_ns] map[Kind:function Name:timekeeping_init] map[Kind:function Name:timekeeping_init_ops] map[Kind:function Name:timekeeping_inject_offset] map[Kind:function Name:timekeeping_max_deferment] map[Kind:function Name:timekeeping_notify] map[Kind:function Name:timekeeping_restore_shadow] map[Kind:function Name:timekeeping_resume] map[Kind:function Name:timekeeping_suspend] map[Kind:function Name:timekeeping_syscore_resume] map[Kind:function Name:timekeeping_syscore_suspend] map[Kind:function Name:timekeeping_update_from_shadow] map[Kind:function Name:timekeeping_valid_for_hres] map[Kind:function Name:timekeeping_validate_timex] map[Kind:function Name:timekeeping_warp_clock] map[Kind:function Name:timestamp_in_interval] map[Kind:function Name:tk_aux_advance] map[Kind:function Name:tk_aux_setup] map[Kind:function Name:tk_aux_update_clocksource] map[Kind:function Name:tk_clock_read] map[Kind:function Name:tk_get_aux_ts64] map[Kind:function Name:tk_is_aux] map[Kind:function Name:tk_normalize_xtime] map[Kind:function Name:tk_set_wall_to_mono] map[Kind:function Name:tk_set_xtime] map[Kind:function Name:tk_setup_internals] map[Kind:function Name:tk_update_aux_offs] map[Kind:function Name:tk_update_coarse_nsecs] map[Kind:function Name:tk_update_ktime_data] map[Kind:function Name:tk_update_leap_state] map[Kind:function Name:tk_update_leap_state_all] map[Kind:function Name:tk_update_ns_to_cyc] map[Kind:function Name:tk_update_sleep_time] map[Kind:function Name:tk_xtime] map[Kind:function Name:tk_xtime_add] map[Kind:function Name:tk_xtime_coarse] map[Kind:function Name:tkd_basic_setup] map[Kind:function Name:update_fast_timekeeper] map[Kind:function Name:update_pvclock_gtod] map[Kind:function Name:update_wall_time] map[Kind:struct Name:adjtimex_result] map[Kind:struct Name:tk_data] map[Kind:struct Name:tk_fast] map[Kind:global_variable Name:__UNIQUE_ID_addressable_do_settimeofday64_574] map[Kind:global_variable Name:__UNIQUE_ID_addressable_get_device_system_crosststamp_569] map[Kind:global_variable Name:__UNIQUE_ID_addressable_getboottime64_597] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_543] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_boot_fast_ns_517] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_clock_ts64_582] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_coarse_real_ts64_598] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_coarse_ts64_599] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_coarse_with_offset_552] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_mono_fast_ns_515] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_raw_555] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_raw_fast_ns_516] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_raw_ts64_579] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_real_fast_ns_519] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_real_seconds_562] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_real_ts64_540] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_resolution_ns_546] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_seconds_561] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_snapshot_565] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_tai_fast_ns_518] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_ts64_558] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_get_with_offset_549] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_mono_to_any_554] map[Kind:global_variable Name:__UNIQUE_ID_addressable_ktime_real_to_base_clock_568] map[Kind:global_variable Name:__UNIQUE_ID_addressable_pvclock_gtod_register_notifier_524] map[Kind:global_variable Name:__UNIQUE_ID_addressable_pvclock_gtod_unregister_notifier_527] map[Kind:global_variable Name:__UNIQUE_ID_addressable_random_get_entropy_fallback_601] map[Kind:global_variable Name:__UNIQUE_ID_addressable_timekeeping_clocksource_has_base_571] map[Kind:global_variable Name:__UNIQUE_ID_addressable_timekeeping_init_ops_586] map[Kind:global_variable Name:__gendwarfksyms_ptr_do_settimeofday64] map[Kind:global_variable Name:__gendwarfksyms_ptr_get_device_system_crosststamp] map[Kind:global_variable Name:__gendwarfksyms_ptr_getboottime64] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_boot_fast_ns] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_clock_ts64] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_coarse_real_ts64] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_coarse_ts64] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_coarse_with_offset] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_mono_fast_ns] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_raw] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_raw_fast_ns] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_raw_ts64] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_real_fast_ns] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_real_seconds] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_real_ts64] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_resolution_ns] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_seconds] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_snapshot] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_tai_fast_ns] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_ts64] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_get_with_offset] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_mono_to_any] map[Kind:global_variable Name:__gendwarfksyms_ptr_ktime_real_to_base_clock] map[Kind:global_variable Name:__gendwarfksyms_ptr_pvclock_gtod_register_notifier] map[Kind:global_variable Name:__gendwarfksyms_ptr_pvclock_gtod_unregister_notifier] map[Kind:global_variable Name:__gendwarfksyms_ptr_random_get_entropy_fallback] map[Kind:global_variable Name:__gendwarfksyms_ptr_timekeeping_clocksource_has_base] map[Kind:global_variable Name:clocksource_read_inlined] map[Kind:global_variable Name:dummy_clock] map[Kind:global_variable Name:offsets] map[Kind:global_variable Name:pvclock_gtod_chain] map[Kind:global_variable Name:timekeeping_syscore] map[Kind:global_variable Name:timekeeping_syscore_ops] map[Kind:global_variable Name:tk_fast_mono] map[Kind:global_variable Name:tk_fast_raw] map[Kind:enum Name:timekeeping_adv_mode]]]
|
| 17/2 |
2026/04/21 11:15 |
llm |
expert |
0m
Model:
Tokens:
input: 13655
output: 17
thoughts: 45
Thoughts: Okay, here's my take on those thoughts, framed as an expert's internal monologue:
**Investigating a Potential Dependency Issue**
Alright, something's not quite right here. I'm scanning through this list, and I'm not spotting `tk_core`. That immediately jumps out because I'm seeing it referenced within `ktime_get_coarse_real_ts64_mg`. It's definitely being *used* somewhere, which means it needs to be defined. This implies a potential missing dependency or an incorrect include path, and it's time to dig into where `tk_core` is declared. I need to track down its definition to understand the relationships here. Let's get to it.
|
| 18/2 |
2026/04/21 11:15 |
tool |
grepper |
0m
Args:
Results: map[Output:
Full output is too long, showing 500 out of 558 lines.
Use more precise expression if possible.
[arch/x86/kvm/x86.c=9976=static DECLARE_WORK(pvclock_gtod_work, pvclock_gtod_update_fn);
--
arch/x86/kvm/x86.c-9978-/*
arch/x86/kvm/x86.c:9979: * Indirection to move queue_work() out of the tk_core.seq write held
arch/x86/kvm/x86.c-9980- * region to prevent possible deadlocks against time accessors which
--
arch/x86/kvm/x86.c=9993=static int pvclock_gtod_notify(struct notifier_block *nb, unsigned long unused,
--
arch/x86/kvm/x86.c-10003- * TSC based clocksource. Delegate queue_work() to irq_work as
arch/x86/kvm/x86.c:10004: * this is invoked with tk_core.seq write held.
arch/x86/kvm/x86.c-10005- */
--
drivers/bluetooth/btmtk.c=43=struct btmtk_section_map {
--
drivers/bluetooth/btmtk.c-61-
drivers/bluetooth/btmtk.c:62:static void btmtk_coredump(struct hci_dev *hdev)
drivers/bluetooth/btmtk.c-63-{
--
drivers/bluetooth/btmtk.c-70-
drivers/bluetooth/btmtk.c:71:static void btmtk_coredump_hdr(struct hci_dev *hdev, struct sk_buff *skb)
drivers/bluetooth/btmtk.c-72-{
--
drivers/bluetooth/btmtk.c-91-
drivers/bluetooth/btmtk.c:92:static void btmtk_coredump_notify(struct hci_dev *hdev, int state)
drivers/bluetooth/btmtk.c-93-{
--
drivers/bluetooth/btmtk.c=402=int btmtk_register_coredump(struct hci_dev *hdev, const char *name,
--
drivers/bluetooth/btmtk.c-413-
drivers/bluetooth/btmtk.c:414: return hci_devcd_register(hdev, btmtk_coredump, btmtk_coredump_hdr,
drivers/bluetooth/btmtk.c:415: btmtk_coredump_notify);
drivers/bluetooth/btmtk.c-416-}
--
drivers/bluetooth/btmtk.h=155=typedef int (*btmtk_reset_sync_func_t)(struct hci_dev *, void *);
drivers/bluetooth/btmtk.h-156-
drivers/bluetooth/btmtk.h:157:struct btmtk_coredump_info {
drivers/bluetooth/btmtk.h-158- const char *driver_name;
--
drivers/bluetooth/btmtk.h=164=struct btmtk_data {
--
drivers/bluetooth/btmtk.h-168- btmtk_reset_sync_func_t reset_sync;
drivers/bluetooth/btmtk.h:169: struct btmtk_coredump_info cd_info;
drivers/bluetooth/btmtk.h-170-
--
drivers/media/pci/cx25821/cx25821-alsa.c=35=do { \
--
drivers/media/pci/cx25821/cx25821-alsa.c-38-} while (0)
drivers/media/pci/cx25821/cx25821-alsa.c:39:#define dprintk_core(level, fmt, arg...) \
drivers/media/pci/cx25821/cx25821-alsa.c-40-do { \
--
kernel/time/timekeeping.c=55=static struct tk_data timekeeper_data[TIMEKEEPERS_MAX];
--
kernel/time/timekeeping.c-57-/* The core timekeeper */
kernel/time/timekeeping.c:58:#define tk_core (timekeeper_data[TIMEKEEPER_CORE])
kernel/time/timekeeping.c-59-
--
kernel/time/timekeeping.c=156=unsigned long timekeeper_lock_irqsave(void)
--
kernel/time/timekeeping.c-159-
kernel/time/timekeeping.c:160: raw_spin_lock_irqsave(&tk_core.lock, flags);
kernel/time/timekeeping.c-161- return flags;
--
kernel/time/timekeeping.c=164=void timekeeper_unlock_irqrestore(unsigned long flags)
kernel/time/timekeeping.c-165-{
kernel/time/timekeeping.c:166: raw_spin_unlock_irqrestore(&tk_core.lock, flags);
kernel/time/timekeeping.c-167-}
--
kernel/time/timekeeping.c=277=static DEFINE_STATIC_KEY_FALSE(clocksource_read_inlined);
--
kernel/time/timekeeping.c-287- * the wrong clocksource is passed to the wrong read function.
kernel/time/timekeeping.c:288: * This isn't necessary to use when holding the tk_core.lock or doing
kernel/time/timekeeping.c-289- * a read of the fast-timekeeper tkrs (which is protected by its own locking
--
kernel/time/timekeeping.c=577=u64 notrace ktime_get_boot_fast_ns(void)
kernel/time/timekeeping.c-578-{
kernel/time/timekeeping.c:579: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-580-
--
kernel/time/timekeeping.c=594=u64 notrace ktime_get_tai_fast_ns(void)
kernel/time/timekeeping.c-595-{
kernel/time/timekeeping.c:596: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-597-
--
kernel/time/timekeeping.c=663=int pvclock_gtod_register_notifier(struct notifier_block *nb)
kernel/time/timekeeping.c-664-{
kernel/time/timekeeping.c:665: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-666- int ret;
kernel/time/timekeeping.c-667-
kernel/time/timekeeping.c:668: guard(raw_spinlock_irqsave)(&tk_core.lock);
kernel/time/timekeeping.c-669- ret = raw_notifier_chain_register(&pvclock_gtod_chain, nb);
--
kernel/time/timekeeping.c=681=int pvclock_gtod_unregister_notifier(struct notifier_block *nb)
kernel/time/timekeeping.c-682-{
kernel/time/timekeeping.c:683: guard(raw_spinlock_irqsave)(&tk_core.lock);
kernel/time/timekeeping.c-684- return raw_notifier_chain_unregister(&pvclock_gtod_chain, nb);
--
kernel/time/timekeeping.c=884=bool ktime_expiry_to_cycles(enum clocksource_ids id, ktime_t expires_ns, u64 *cycles)
kernel/time/timekeeping.c-885-{
kernel/time/timekeeping.c:886: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-887- struct tk_read_base *tkrm = &tk->tkr_mono;
--
kernel/time/timekeeping.c-903- do {
kernel/time/timekeeping.c:904: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-905-
--
kernel/time/timekeeping.c-915-
kernel/time/timekeeping.c:916: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-917-
--
kernel/time/timekeeping.c=934=void ktime_get_real_ts64(struct timespec64 *ts)
kernel/time/timekeeping.c-935-{
kernel/time/timekeeping.c:936: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-937- unsigned int seq;
--
kernel/time/timekeeping.c-942- do {
kernel/time/timekeeping.c:943: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-944-
--
kernel/time/timekeeping.c-947-
kernel/time/timekeeping.c:948: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-949-
--
kernel/time/timekeeping.c=955=ktime_t ktime_get(void)
kernel/time/timekeeping.c-956-{
kernel/time/timekeeping.c:957: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-958- unsigned int seq;
--
kernel/time/timekeeping.c-964- do {
kernel/time/timekeeping.c:965: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-966- base = tk->tkr_mono.base;
--
kernel/time/timekeeping.c-968-
kernel/time/timekeeping.c:969: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-970-
--
kernel/time/timekeeping.c=975=u32 ktime_get_resolution_ns(void)
kernel/time/timekeeping.c-976-{
kernel/time/timekeeping.c:977: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-978- unsigned int seq;
--
kernel/time/timekeeping.c-983- do {
kernel/time/timekeeping.c:984: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-985- nsecs = tk->tkr_mono.mult >> tk->tkr_mono.shift;
kernel/time/timekeeping.c:986: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-987-
--
kernel/time/timekeeping.c=992=static const ktime_t *const offsets[TK_OFFS_MAX] = {
kernel/time/timekeeping.c:993: [TK_OFFS_REAL] = &tk_core.timekeeper.offs_real,
kernel/time/timekeeping.c:994: [TK_OFFS_BOOT] = &tk_core.timekeeper.offs_boot,
kernel/time/timekeeping.c:995: [TK_OFFS_TAI] = &tk_core.timekeeper.offs_tai,
kernel/time/timekeeping.c-996-};
--
kernel/time/timekeeping.c=998=ktime_t ktime_get_with_offset(enum tk_offsets offs)
kernel/time/timekeeping.c-999-{
kernel/time/timekeeping.c:1000: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-1001- const ktime_t *offset = offsets[offs];
--
kernel/time/timekeeping.c-1008- do {
kernel/time/timekeeping.c:1009: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-1010- base = ktime_add(tk->tkr_mono.base, *offset);
--
kernel/time/timekeeping.c-1012-
kernel/time/timekeeping.c:1013: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-1014-
--
kernel/time/timekeeping.c=1020=ktime_t ktime_get_coarse_with_offset(enum tk_offsets offs)
kernel/time/timekeeping.c-1021-{
kernel/time/timekeeping.c:1022: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-1023- const ktime_t *offset = offsets[offs];
--
kernel/time/timekeeping.c-1030- do {
kernel/time/timekeeping.c:1031: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-1032- base = ktime_add(tk->tkr_mono.base, *offset);
--
kernel/time/timekeeping.c-1034-
kernel/time/timekeeping.c:1035: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-1036-
--
kernel/time/timekeeping.c=1046=ktime_t ktime_mono_to_any(ktime_t tmono, enum tk_offsets offs)
--
kernel/time/timekeeping.c-1060- do {
kernel/time/timekeeping.c:1061: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-1062- tconv = ktime_add(tmono, *offset);
kernel/time/timekeeping.c:1063: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-1064-
--
kernel/time/timekeeping.c=1072=ktime_t ktime_get_raw(void)
kernel/time/timekeeping.c-1073-{
kernel/time/timekeeping.c:1074: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-1075- unsigned int seq;
--
kernel/time/timekeeping.c-1079- do {
kernel/time/timekeeping.c:1080: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-1081- base = tk->tkr_raw.base;
--
kernel/time/timekeeping.c-1083-
kernel/time/timekeeping.c:1084: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-1085-
--
kernel/time/timekeeping.c=1098=void ktime_get_ts64(struct timespec64 *ts)
kernel/time/timekeeping.c-1099-{
kernel/time/timekeeping.c:1100: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-1101- struct timespec64 tomono;
--
kernel/time/timekeeping.c-1107- do {
kernel/time/timekeeping.c:1108: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-1109- ts->tv_sec = tk->xtime_sec;
--
kernel/time/timekeeping.c-1112-
kernel/time/timekeeping.c:1113: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-1114-
--
kernel/time/timekeeping.c=1130=time64_t ktime_get_seconds(void)
kernel/time/timekeeping.c-1131-{
kernel/time/timekeeping.c:1132: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-1133-
--
kernel/time/timekeeping.c=1149=time64_t ktime_get_real_seconds(void)
kernel/time/timekeeping.c-1150-{
kernel/time/timekeeping.c:1151: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-1152- time64_t seconds;
--
kernel/time/timekeeping.c-1158- do {
kernel/time/timekeeping.c:1159: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-1160- seconds = tk->xtime_sec;
kernel/time/timekeeping.c-1161-
kernel/time/timekeeping.c:1162: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-1163-
--
kernel/time/timekeeping.c=1178=noinstr time64_t __ktime_get_real_seconds(void)
kernel/time/timekeeping.c-1179-{
kernel/time/timekeeping.c:1180: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-1181-
--
kernel/time/timekeeping.c=1189=void ktime_get_snapshot(struct system_time_snapshot *systime_snapshot)
kernel/time/timekeeping.c-1190-{
kernel/time/timekeeping.c:1191: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-1192- unsigned int seq;
--
kernel/time/timekeeping.c-1202- do {
kernel/time/timekeeping.c:1203: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-1204- now = tk_clock_read(&tk->tkr_mono);
--
kernel/time/timekeeping.c-1208- base_real = ktime_add(tk->tkr_mono.base,
kernel/time/timekeeping.c:1209: tk_core.timekeeper.offs_real);
kernel/time/timekeeping.c-1210- base_boot = ktime_add(tk->tkr_mono.base,
kernel/time/timekeeping.c:1211: tk_core.timekeeper.offs_boot);
kernel/time/timekeeping.c-1212- base_raw = tk->tkr_raw.base;
--
kernel/time/timekeeping.c-1214- nsec_raw = timekeeping_cycles_to_ns(&tk->tkr_raw, now);
kernel/time/timekeeping.c:1215: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-1216-
--
kernel/time/timekeeping.c=1258=static int adjust_historical_crosststamp(struct system_time_snapshot *history,
--
kernel/time/timekeeping.c-1263-{
kernel/time/timekeeping.c:1264: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-1265- u64 corr_raw, corr_real;
--
kernel/time/timekeeping.c=1346=static bool convert_base_to_cs(struct system_counterval_t *scv)
kernel/time/timekeeping.c-1347-{
kernel/time/timekeeping.c:1348: struct clocksource *cs = tk_core.timekeeper.tkr_mono.clock;
kernel/time/timekeeping.c-1349- struct clocksource_base *base;
--
kernel/time/timekeeping.c=1374=static bool convert_cs_to_base(u64 *cycles, enum clocksource_ids base_id)
kernel/time/timekeeping.c-1375-{
kernel/time/timekeeping.c:1376: struct clocksource *cs = tk_core.timekeeper.tkr_mono.clock;
kernel/time/timekeeping.c-1377- struct clocksource_base *base;
--
kernel/time/timekeeping.c=1393=static bool convert_ns_to_cs(u64 *delta)
kernel/time/timekeeping.c-1394-{
kernel/time/timekeeping.c:1395: struct tk_read_base *tkr = &tk_core.timekeeper.tkr_mono;
kernel/time/timekeeping.c-1396-
--
kernel/time/timekeeping.c=1414=bool ktime_real_to_base_clock(ktime_t treal, enum clocksource_ids base_id, u64 *cycles)
kernel/time/timekeeping.c-1415-{
kernel/time/timekeeping.c:1416: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-1417- unsigned int seq;
--
kernel/time/timekeeping.c-1420- do {
kernel/time/timekeeping.c:1421: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-1422- if ((u64)treal < tk->tkr_mono.base_real)
--
kernel/time/timekeeping.c-1429- return false;
kernel/time/timekeeping.c:1430: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-1431-
--
kernel/time/timekeeping.c=1447=int get_device_system_crosststamp(int (*get_time_fn)
--
kernel/time/timekeeping.c-1455- struct system_counterval_t system_counterval = {};
kernel/time/timekeeping.c:1456: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-1457- u64 cycles, now, interval_start;
--
kernel/time/timekeeping.c-1466- do {
kernel/time/timekeeping.c:1467: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-1468- /*
--
kernel/time/timekeeping.c-1501- base_real = ktime_add(tk->tkr_mono.base,
kernel/time/timekeeping.c:1502: tk_core.timekeeper.offs_real);
kernel/time/timekeeping.c-1503- base_raw = tk->tkr_raw.base;
--
kernel/time/timekeeping.c-1506- nsec_raw = timekeeping_cycles_to_ns(&tk->tkr_raw, cycles);
kernel/time/timekeeping.c:1507: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-1508-
--
kernel/time/timekeeping.c=1558=bool timekeeping_clocksource_has_base(enum clocksource_ids id)
--
kernel/time/timekeeping.c-1564- */
kernel/time/timekeeping.c:1565: struct clocksource_base *base = READ_ONCE(tk_core.timekeeper.tkr_mono.clock->base);
kernel/time/timekeeping.c-1566-
--
kernel/time/timekeeping.c=1577=int do_settimeofday64(const struct timespec64 *ts)
--
kernel/time/timekeeping.c-1583-
kernel/time/timekeeping.c:1584: scoped_guard (raw_spinlock_irqsave, &tk_core.lock) {
kernel/time/timekeeping.c:1585: struct timekeeper *tks = &tk_core.shadow_timekeeper;
kernel/time/timekeeping.c-1586-
--
kernel/time/timekeeping.c-1592- if (timespec64_compare(&tks->wall_to_monotonic, &ts_delta) > 0) {
kernel/time/timekeeping.c:1593: timekeeping_restore_shadow(&tk_core);
kernel/time/timekeeping.c-1594- return -EINVAL;
--
kernel/time/timekeeping.c-1598- tk_set_xtime(tks, ts);
kernel/time/timekeeping.c:1599: timekeeping_update_from_shadow(&tk_core, TK_UPDATE_ALL);
kernel/time/timekeeping.c-1600- }
--
kernel/time/timekeeping.c=1665=static int timekeeping_inject_offset(const struct timespec64 *ts)
--
kernel/time/timekeeping.c-1668-
kernel/time/timekeeping.c:1669: scoped_guard (raw_spinlock_irqsave, &tk_core.lock)
kernel/time/timekeeping.c:1670: ret = __timekeeping_inject_offset(&tk_core, ts);
kernel/time/timekeeping.c-1671-
--
kernel/time/timekeeping.c=1726=static int change_clocksource(void *data)
--
kernel/time/timekeeping.c-1743-
kernel/time/timekeeping.c:1744: scoped_guard (raw_spinlock_irqsave, &tk_core.lock) {
kernel/time/timekeeping.c:1745: struct timekeeper *tks = &tk_core.shadow_timekeeper;
kernel/time/timekeeping.c-1746-
--
kernel/time/timekeeping.c-1749- tk_setup_internals(tks, new);
kernel/time/timekeeping.c:1750: timekeeping_update_from_shadow(&tk_core, TK_UPDATE_ALL);
kernel/time/timekeeping.c-1751- }
--
kernel/time/timekeeping.c=1771=int timekeeping_notify(struct clocksource *clock)
kernel/time/timekeeping.c-1772-{
kernel/time/timekeeping.c:1773: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-1774-
--
kernel/time/timekeeping.c=1800=void ktime_get_raw_ts64(struct timespec64 *ts)
kernel/time/timekeeping.c-1801-{
kernel/time/timekeeping.c:1802: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-1803- unsigned int seq;
--
kernel/time/timekeeping.c-1806- do {
kernel/time/timekeeping.c:1807: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-1808- ts->tv_sec = tk->raw_sec;
--
kernel/time/timekeeping.c-1810-
kernel/time/timekeeping.c:1811: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-1812-
--
kernel/time/timekeeping.c=1855=int timekeeping_valid_for_hres(void)
kernel/time/timekeeping.c-1856-{
kernel/time/timekeeping.c:1857: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-1858- unsigned int seq;
--
kernel/time/timekeeping.c-1861- do {
kernel/time/timekeeping.c:1862: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-1863-
--
kernel/time/timekeeping.c-1865-
kernel/time/timekeeping.c:1866: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-1867-
--
kernel/time/timekeeping.c=1874=u64 timekeeping_max_deferment(void)
kernel/time/timekeeping.c-1875-{
kernel/time/timekeeping.c:1876: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-1877- unsigned int seq;
--
kernel/time/timekeeping.c-1880- do {
kernel/time/timekeeping.c:1881: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-1882-
--
kernel/time/timekeeping.c-1884-
kernel/time/timekeeping.c:1885: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-1886-
--
kernel/time/timekeeping.c=1956=void __init timekeeping_init(void)
--
kernel/time/timekeeping.c-1958- struct timespec64 wall_time, boot_offset, wall_to_mono;
kernel/time/timekeeping.c:1959: struct timekeeper *tks = &tk_core.shadow_timekeeper;
kernel/time/timekeeping.c-1960- struct clocksource *clock;
kernel/time/timekeeping.c-1961-
kernel/time/timekeeping.c:1962: tkd_basic_setup(&tk_core, TIMEKEEPER_CORE, true);
kernel/time/timekeeping.c-1963- tk_aux_setup();
--
kernel/time/timekeeping.c-1982-
kernel/time/timekeeping.c:1983: guard(raw_spinlock_irqsave)(&tk_core.lock);
kernel/time/timekeeping.c-1984-
--
kernel/time/timekeeping.c-1996-
kernel/time/timekeeping.c:1997: timekeeping_update_from_shadow(&tk_core, TK_CLOCK_WAS_SET);
kernel/time/timekeeping.c-1998-}
--
kernel/time/timekeeping.c=2073=void timekeeping_inject_sleeptime64(const struct timespec64 *delta)
kernel/time/timekeeping.c-2074-{
kernel/time/timekeeping.c:2075: scoped_guard(raw_spinlock_irqsave, &tk_core.lock) {
kernel/time/timekeeping.c:2076: struct timekeeper *tks = &tk_core.shadow_timekeeper;
kernel/time/timekeeping.c-2077-
--
kernel/time/timekeeping.c-2080- __timekeeping_inject_sleeptime(tks, delta);
kernel/time/timekeeping.c:2081: timekeeping_update_from_shadow(&tk_core, TK_UPDATE_ALL);
kernel/time/timekeeping.c-2082- }
--
kernel/time/timekeeping.c=2092=void timekeeping_resume(void)
kernel/time/timekeeping.c-2093-{
kernel/time/timekeeping.c:2094: struct timekeeper *tks = &tk_core.shadow_timekeeper;
kernel/time/timekeeping.c-2095- struct clocksource *clock = tks->tkr_mono.clock;
--
kernel/time/timekeeping.c-2105-
kernel/time/timekeeping.c:2106: raw_spin_lock_irqsave(&tk_core.lock, flags);
kernel/time/timekeeping.c-2107-
--
kernel/time/timekeeping.c-2140- timekeeping_suspended = 0;
kernel/time/timekeeping.c:2141: timekeeping_update_from_shadow(&tk_core, TK_CLOCK_WAS_SET);
kernel/time/timekeeping.c:2142: raw_spin_unlock_irqrestore(&tk_core.lock, flags);
kernel/time/timekeeping.c-2143-
--
kernel/time/timekeeping.c=2157=int timekeeping_suspend(void)
kernel/time/timekeeping.c-2158-{
kernel/time/timekeeping.c:2159: struct timekeeper *tks = &tk_core.shadow_timekeeper;
kernel/time/timekeeping.c-2160- struct timespec64 delta, delta_delta;
--
kernel/time/timekeeping.c-2177-
kernel/time/timekeeping.c:2178: raw_spin_lock_irqsave(&tk_core.lock, flags);
kernel/time/timekeeping.c-2179- timekeeping_forward_now(tks);
--
kernel/time/timekeeping.c-2212-
kernel/time/timekeeping.c:2213: timekeeping_update_from_shadow(&tk_core, 0);
kernel/time/timekeeping.c-2214- halt_fast_timekeeper(tks);
kernel/time/timekeeping.c:2215: raw_spin_unlock_irqrestore(&tk_core.lock, flags);
kernel/time/timekeeping.c-2216-
--
kernel/time/timekeeping.c=2537=static bool timekeeping_advance(enum timekeeping_adv_mode mode)
kernel/time/timekeeping.c-2538-{
kernel/time/timekeeping.c:2539: guard(raw_spinlock_irqsave)(&tk_core.lock);
kernel/time/timekeeping.c:2540: return __timekeeping_advance(&tk_core, mode);
kernel/time/timekeeping.c-2541-}
--
kernel/time/timekeeping.c=2566=void getboottime64(struct timespec64 *ts)
kernel/time/timekeeping.c-2567-{
kernel/time/timekeeping.c:2568: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-2569- ktime_t t = ktime_sub(tk->offs_real, tk->offs_boot);
--
kernel/time/timekeeping.c=2575=void ktime_get_coarse_real_ts64(struct timespec64 *ts)
kernel/time/timekeeping.c-2576-{
kernel/time/timekeeping.c:2577: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-2578- unsigned int seq;
--
kernel/time/timekeeping.c-2580- do {
kernel/time/timekeeping.c:2581: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-2582-
kernel/time/timekeeping.c-2583- *ts = tk_xtime_coarse(tk);
kernel/time/timekeeping.c:2584: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-2585-}
--
kernel/time/timekeeping.c=2597=void ktime_get_coarse_real_ts64_mg(struct timespec64 *ts)
kernel/time/timekeeping.c-2598-{
kernel/time/timekeeping.c:2599: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-2600- u64 floor = atomic64_read(&mg_floor);
--
kernel/time/timekeeping.c-2604- do {
kernel/time/timekeeping.c:2605: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-2606- *ts = tk_xtime_coarse(tk);
kernel/time/timekeeping.c:2607: offset = tk_core.timekeeper.offs_real;
kernel/time/timekeeping.c:2608: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-2609-
--
kernel/time/timekeeping.c=2635=void ktime_get_real_ts64_mg(struct timespec64 *ts)
kernel/time/timekeeping.c-2636-{
kernel/time/timekeeping.c:2637: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-2638- ktime_t old = atomic64_read(&mg_floor);
--
kernel/time/timekeeping.c-2643- do {
kernel/time/timekeeping.c:2644: seq = read_seqcount_begin(&tk_core.seq);
kernel/time/timekeeping.c-2645-
--
kernel/time/timekeeping.c-2648- nsecs = timekeeping_get_ns(&tk->tkr_mono);
kernel/time/timekeeping.c:2649: offset = tk_core.timekeeper.offs_real;
kernel/time/timekeeping.c:2650: } while (read_seqcount_retry(&tk_core.seq, seq));
kernel/time/timekeeping.c-2651-
--
kernel/time/timekeeping.c=2676=void ktime_get_coarse_ts64(struct timespec64 *ts)
kernel/time/timekeeping.c-2677-{
kernel/time/timekeeping.c:2678: struct timekeeper *tk = &tk_core.timekeeper;
kernel/time/timekeeping.c-2679- struct timespec64 now, mono;
--
kernel/time/timekeeping.c-2682- do {
kernel/time/timekeeping.c:2683: seq = read_seqcount_begin(&tk_core.seq);
]
]
|
| 19/2 |
2026/04/21 11:15 |
llm |
expert |
0m
Model:
Tokens:
input: 23925
output: 40
thoughts: 94
Thoughts: Okay, let's break this down.
**Understanding the Timekeeping Core**
So, the code's defining the `tk_core` – that's the heart of the timekeeping system. It's actually a reference to an element within the `timekeeper_data` array. Right, looking at it, it's defined as `(timekeeper_data[TIMEKEEPER_CORE])`. Makes perfect sense.
Now, the `timekeeper_data` itself? That's an array, specifically an array of `struct tk_data`. So, each element in the array is a structure that presumably holds all the necessary data for a specific timekeeper. That `TIMEKEEPERS_MAX` suggests there could be multiple timekeepers running in parallel, interesting. Let's see what exactly `struct tk_data` encompasses.
|
| 20/2 |
2026/04/21 11:15 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/time/timekeeping.c IncludeLines:true Name:tk_data]
Results: map[SourceCode: 44: /*
45: * The most important data for readout fits into a single 64 byte
46: * cache line.
47: */
48: struct tk_data {
49: seqcount_raw_spinlock_t seq;
50: struct timekeeper timekeeper;
51: struct timekeeper shadow_timekeeper;
52: raw_spinlock_t lock;
53: } ____cacheline_aligned;
SourceFile:kernel/time/timekeeping.c]
|
| 21/2 |
2026/04/21 11:15 |
llm |
expert |
0m
Model:
Tokens:
input: 24397
output: 45
thoughts: 119
Thoughts: Okay, let's break this down.
**Analyzing Timekeeping Synchronization**
Alright, so the core of this seems to be about synchronization between time updates and time reads within the kernel, specifically focusing on `timekeeping_update_from_shadow` and `ktime_get_coarse_real_ts64_mg`. I'm seeing a write operation happening within `timekeeping_update_from_shadow`, and a read operation in `ktime_get_coarse_real_ts64_mg`. My immediate thought is, how are these reads and writes being protected from concurrent access?
It looks like `ktime_get_coarse_real_ts64_mg` is utilizing `read_seqcount_begin(&tk_core.seq)` and `read_seqcount_retry(&tk_core.seq, seq)`. That's smart. That suggests a sequence lock is in play, aiming to protect the `tk_core.timekeeper` structure from being read while it is being updated. The read side obtains a sequence number, performs the read, and then checks if the sequence number has changed. If it has, the read is retried, ensuring data consistency.
The next logical step is to dive into `timekeeping_update_from_shadow` to confirm the write side of the synchronization. I need to examine how the write operation interacts with that sequence lock to ensure proper protection and prevent data corruption during time updates. I'll need to verify that `timekeeping_update_from_shadow` increments or otherwise handles `tk_core.seq` correctly during its updates.
|
| 22/2 |
2026/04/21 11:15 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/time/timekeeping.c IncludeLines:true Name:timekeeping_update_from_shadow]
Results: map[SourceCode: 783: static void timekeeping_update_from_shadow(struct tk_data *tkd, unsigned int action)
784: {
785: struct timekeeper *tk = &tkd->shadow_timekeeper;
786:
787: lockdep_assert_held(&tkd->lock);
788:
789: /*
790: * Block out readers before running the updates below because that
791: * updates VDSO and other time related infrastructure. Not blocking
792: * the readers might let a reader see time going backwards when
793: * reading from the VDSO after the VDSO update and then reading in
794: * the kernel from the timekeeper before that got updated.
795: */
796: write_seqcount_begin(&tkd->seq);
797:
798: if (action & TK_CLEAR_NTP) {
799: tk->ntp_error = 0;
800: ntp_clear(tk->id);
801: }
802:
803: tk_update_leap_state(tk);
804: tk_update_ktime_data(tk);
805: tk->tkr_mono.base_real = tk->tkr_mono.base + tk->offs_real;
806:
807: if (tk->id == TIMEKEEPER_CORE) {
808: tk_update_ns_to_cyc(tk, &tkd->timekeeper);
809: update_vsyscall(tk);
810: update_pvclock_gtod(tk, action & TK_CLOCK_WAS_SET);
811:
812: update_fast_timekeeper(&tk->tkr_mono, &tk_fast_mono);
813: update_fast_timekeeper(&tk->tkr_raw, &tk_fast_raw);
814: } else if (tk_is_aux(tk)) {
815: vdso_time_update_aux(tk);
816: }
817:
818: if (action & TK_CLOCK_WAS_SET)
819: tk->clock_was_set_seq++;
820:
821: /*
822: * Update the real timekeeper.
823: *
824: * We could avoid this memcpy() by switching pointers, but that has
825: * the downside that the reader side does not longer benefit from
826: * the cacheline optimized data layout of the timekeeper and requires
827: * another indirection.
828: */
829: memcpy(&tkd->timekeeper, tk, sizeof(*tk));
830: write_seqcount_end(&tkd->seq);
831: }
SourceFile:kernel/time/timekeeping.c]
|
| 23/2 |
2026/04/21 11:15 |
llm |
expert |
0m
Model:
Tokens:
input: 25715
output: 24
thoughts: 2627
Thoughts: **KCSAN Data Race Investigation: timekeeping_update_from_shadow and ktime_get_coarse_real_ts64_mg**
Okay, so KCSAN is reporting a data race, and I need to figure out why. The scenario involves `timekeeping_update_from_shadow` writing to `tkd->timekeeper` and `ktime_get_coarse_real_ts64_mg` reading from it, both using a `seqcount`. This *should* be safe, as `seqcount` is supposed to synchronize concurrent access, with the reader retrying if a write happens during its read.
The core of the problem lies in these seemingly innocuous lines: `memcpy(&tkd->timekeeper, tk, sizeof(*tk))` in the writer and the reads inside the `read_seqcount_retry` loop in the reader, specifically `tk->coarse_nsec`. Both `read_seqcount_begin` and `write_seqcount_begin` are called before and after their respective accesses.
My initial thought is that if this is correctly implemented with `seqcount_raw_spinlock_t`, then KCSAN should not report a data race. However, KCSAN is reporting a "value changed" race on `coarse_nsec`, which is troubling.
The first instinct is to check the `read_seqcount_begin` implementation. This is where KCSAN's behavior is key. It looks like `read_seqcount_begin` *does* call `kcsan_atomic_next`, marking the region as protected. That should be enough to tell KCSAN to ignore the data race.
I need to confirm that `tk_core.seq` is indeed the `seqcount_raw_spinlock_t`. It is!
Also, I have to remember that `ktime_get_coarse_real_ts64_mg` has a `_mg` suffix, which points to some "multigrain" implications. But it looks identical to `ktime_get_coarse_real_ts64` (the non `_mg` version), so that is not the issue.
Now, I re-examine the reported code: the writer `memcpy`s, and the reader accesses fields of the `timekeeper` struct within the `read_seqcount_retry` loop, including `coarse_nsec`. The offset of the read in the report matches the offset of `coarse_nsec` in `tk_core.timekeeper`.
The key here is that KCSAN *should* be aware of the read-write sequence that happens. `read_seqcount_begin` contains a call to `kcsan_atomic_next`, which *should* instruct KCSAN to ignore the accesses that follow up to `read_seqcount_retry`. But it doesn't.
Now I realize that `read_seqcount_begin` is actually calling `__read_seqcount_begin` which *does* call `kcsan_atomic_next(KCSAN_SEQLOCK_REGION_MAX)`. That's the correct way to tell KCSAN that this is a synchronized region.
So, since it's using the correct primitives (seqcount) *and* `kcsan_atomic_next` is being called, KCSAN *should not* be complaining. It is a bit suspicious that `memcpy` covers the *whole* `timekeeper` struct, so a torn read should be detected and retried. The critical region in the reader covers reading the entire struct. The `seqcount` *should* be functioning correctly.
So, the mystery deepens. The data race *should* be handled correctly by the `seqcount`. I need to dig deeper into the KCSAN report and the surrounding code to understand why KCSAN is flagging this as a data race. It's possible there's an interaction I'm missing, or perhaps a KCSAN bug.
|
| 24/2 |
2026/04/21 11:16 |
tool |
read-file |
0m
Args: map[File:include/linux/seqlock.h]
Results: map[Contents:/* SPDX-License-Identifier: GPL-2.0 */
#ifndef __LINUX_SEQLOCK_H
#define __LINUX_SEQLOCK_H
/*
* seqcount_t / seqlock_t - a reader-writer consistency mechanism with
* lockless readers (read-only retry loops), and no writer starvation.
*
* See Documentation/locking/seqlock.rst
*
* Copyrights:
* - Based on x86_64 vsyscall gettimeofday: Keith Owens, Andrea Arcangeli
* - Sequence counters with associated locks, (C) 2020 Linutronix GmbH
*/
#include <linux/compiler.h>
#include <linux/cleanup.h>
#include <linux/kcsan-checks.h>
#include <linux/lockdep.h>
#include <linux/mutex.h>
#include <linux/preempt.h>
#include <linux/seqlock_types.h>
#include <linux/spinlock.h>
#include <asm/processor.h>
/*
* The seqlock seqcount_t interface does not prescribe a precise sequence of
* read begin/retry/end. For readers, typically there is a call to
* read_seqcount_begin() and read_seqcount_retry(), however, there are more
* esoteric cases which do not follow this pattern.
*
* As a consequence, we take the following best-effort approach for raw usage
* via seqcount_t under KCSAN: upon beginning a seq-reader critical section,
* pessimistically mark the next KCSAN_SEQLOCK_REGION_MAX memory accesses as
* atomics; if there is a matching read_seqcount_retry() call, no following
* memory operations are considered atomic. Usage of the seqlock_t interface
* is not affected.
*/
#define KCSAN_SEQLOCK_REGION_MAX 1000
static inline void __seqcount_init(seqcount_t *s, const char *name,
struct lock_class_key *key)
{
/*
* Make sure we are not reinitializing a held lock:
*/
lockdep_init_map(&s->dep_map, name, key, 0);
s->sequence = 0;
}
#ifdef CONFIG_DEBUG_LOCK_ALLOC
# define SEQCOUNT_DEP_MAP_INIT(lockname) \
.dep_map = { .name = #lockname }
/**
* seqcount_init() - runtime initializer for seqcount_t
* @s: Pointer to the seqcount_t instance
*/
# define seqcount_init(s) \
do { \
static struct lock_class_key __key; \
__seqcount_init((s), #s, &__key); \
} while (0)
static inline void seqcount_lockdep_reader_access(const seqcount_t *s)
{
seqcount_t *l = (seqcount_t *)s;
unsigned long flags;
local_irq_save(flags);
seqcount_acquire_read(&l->dep_map, 0, 0, _RET_IP_);
seqcount_release(&l->dep_map, _RET_IP_);
local_irq_restore(flags);
}
#else
# define SEQCOUNT_DEP_MAP_INIT(lockname)
# define seqcount_init(s) __seqcount_init(s, NULL, NULL)
# define seqcount_lockdep_reader_access(x)
#endif
/**
* SEQCNT_ZERO() - static initializer for seqcount_t
* @name: Name of the seqcount_t instance
*/
#define SEQCNT_ZERO(name) { .sequence = 0, SEQCOUNT_DEP_MAP_INIT(name) }
/*
* Sequence counters with associated locks (seqcount_LOCKNAME_t)
*
* A sequence counter which associates the lock used for writer
* serialization at initialization time. This enables lockdep to validate
* that the write side critical section is properly serialized.
*
* For associated locks which do not implicitly disable preemption,
* preemption protection is enforced in the write side function.
*
* Lockdep is never used in any for the raw write variants.
*
* See Documentation/locking/seqlock.rst
*/
/*
* typedef seqcount_LOCKNAME_t - sequence counter with LOCKNAME associated
* @seqcount: The real sequence counter
* @lock: Pointer to the associated lock
*
* A plain sequence counter with external writer synchronization by
* LOCKNAME @lock. The lock is associated to the sequence counter in the
* static initializer or init function. This enables lockdep to validate
* that the write side critical section is properly serialized.
*
* LOCKNAME: raw_spinlock, spinlock, rwlock or mutex
*/
/*
* seqcount_LOCKNAME_init() - runtime initializer for seqcount_LOCKNAME_t
* @s: Pointer to the seqcount_LOCKNAME_t instance
* @lock: Pointer to the associated lock
*/
#define seqcount_LOCKNAME_init(s, _lock, lockname) \
do { \
seqcount_##lockname##_t *____s = (s); \
seqcount_init(&____s->seqcount); \
__SEQ_LOCK(____s->lock = (_lock)); \
} while (0)
#define seqcount_raw_spinlock_init(s, lock) seqcount_LOCKNAME_init(s, lock, raw_spinlock)
#define seqcount_spinlock_init(s, lock) seqcount_LOCKNAME_init(s, lock, spinlock)
#define seqcount_rwlock_init(s, lock) seqcount_LOCKNAME_init(s, lock, rwlock)
#define seqcount_mutex_init(s, lock) seqcount_LOCKNAME_init(s, lock, mutex)
/*
* SEQCOUNT_LOCKNAME() - Instantiate seqcount_LOCKNAME_t and helpers
* seqprop_LOCKNAME_*() - Property accessors for seqcount_LOCKNAME_t
*
* @lockname: "LOCKNAME" part of seqcount_LOCKNAME_t
* @locktype: LOCKNAME canonical C data type
* @preemptible: preemptibility of above locktype
* @lockbase: prefix for associated lock/unlock
*/
#define SEQCOUNT_LOCKNAME(lockname, locktype, preemptible, lockbase) \
static __always_inline seqcount_t * \
__seqprop_##lockname##_ptr(seqcount_##lockname##_t *s) \
{ \
return &s->seqcount; \
} \
\
static __always_inline const seqcount_t * \
__seqprop_##lockname##_const_ptr(const seqcount_##lockname##_t *s) \
{ \
return &s->seqcount; \
} \
\
static __always_inline unsigned \
__seqprop_##lockname##_sequence(const seqcount_##lockname##_t *s) \
{ \
unsigned seq = smp_load_acquire(&s->seqcount.sequence); \
\
if (!IS_ENABLED(CONFIG_PREEMPT_RT)) \
return seq; \
\
if (preemptible && unlikely(seq & 1)) { \
__SEQ_LOCK(lockbase##_lock(s->lock)); \
__SEQ_LOCK(lockbase##_unlock(s->lock)); \
\
/* \
* Re-read the sequence counter since the (possibly \
* preempted) writer made progress. \
*/ \
seq = smp_load_acquire(&s->seqcount.sequence); \
} \
\
return seq; \
} \
\
static __always_inline bool \
__seqprop_##lockname##_preemptible(const seqcount_##lockname##_t *s) \
{ \
if (!IS_ENABLED(CONFIG_PREEMPT_RT)) \
return preemptible; \
\
/* PREEMPT_RT relies on the above LOCK+UNLOCK */ \
return false; \
} \
\
static __always_inline void \
__seqprop_##lockname##_assert(const seqcount_##lockname##_t *s) \
{ \
__SEQ_LOCK(lockdep_assert_held(s->lock)); \
}
/*
* __seqprop() for seqcount_t
*/
static inline seqcount_t *__seqprop_ptr(seqcount_t *s)
{
return s;
}
static inline const seqcount_t *__seqprop_const_ptr(const seqcount_t *s)
{
return s;
}
static inline unsigned __seqprop_sequence(const seqcount_t *s)
{
return smp_load_acquire(&s->sequence);
}
static inline bool __seqprop_preemptible(const seqcount_t *s)
{
return false;
}
static inline void __seqprop_assert(const seqcount_t *s)
{
lockdep_assert_preemption_disabled();
}
#define __SEQ_RT IS_ENABLED(CONFIG_PREEMPT_RT)
SEQCOUNT_LOCKNAME(raw_spinlock, raw_spinlock_t, false, raw_spin)
SEQCOUNT_LOCKNAME(spinlock, spinlock_t, __SEQ_RT, spin)
SEQCOUNT_LOCKNAME(rwlock, rwlock_t, __SEQ_RT, read)
SEQCOUNT_LOCKNAME(mutex, struct mutex, true, mutex)
#undef SEQCOUNT_LOCKNAME
/*
* SEQCNT_LOCKNAME_ZERO - static initializer for seqcount_LOCKNAME_t
* @name: Name of the seqcount_LOCKNAME_t instance
* @lock: Pointer to the associated LOCKNAME
*/
#define SEQCOUNT_LOCKNAME_ZERO(seq_name, assoc_lock) { \
.seqcount = SEQCNT_ZERO(seq_name.seqcount), \
__SEQ_LOCK(.lock = (assoc_lock)) \
}
#define SEQCNT_RAW_SPINLOCK_ZERO(name, lock) SEQCOUNT_LOCKNAME_ZERO(name, lock)
#define SEQCNT_SPINLOCK_ZERO(name, lock) SEQCOUNT_LOCKNAME_ZERO(name, lock)
#define SEQCNT_RWLOCK_ZERO(name, lock) SEQCOUNT_LOCKNAME_ZERO(name, lock)
#define SEQCNT_MUTEX_ZERO(name, lock) SEQCOUNT_LOCKNAME_ZERO(name, lock)
#define SEQCNT_WW_MUTEX_ZERO(name, lock) SEQCOUNT_LOCKNAME_ZERO(name, lock)
#define __seqprop_case(s, lockname, prop) \
seqcount_##lockname##_t: __seqprop_##lockname##_##prop
#define __seqprop(s, prop) _Generic(*(s), \
seqcount_t: __seqprop_##prop, \
__seqprop_case((s), raw_spinlock, prop), \
__seqprop_case((s), spinlock, prop), \
__seqprop_case((s), rwlock, prop), \
__seqprop_case((s), mutex, prop))
#define seqprop_ptr(s) __seqprop(s, ptr)(s)
#define seqprop_const_ptr(s) __seqprop(s, const_ptr)(s)
#define seqprop_sequence(s) __seqprop(s, sequence)(s)
#define seqprop_preemptible(s) __seqprop(s, preemptible)(s)
#define seqprop_assert(s) __seqprop(s, assert)(s)
/**
* __read_seqcount_begin() - begin a seqcount_t read section
* @s: Pointer to seqcount_t or any of the seqcount_LOCKNAME_t variants
*
* Return: count to be passed to read_seqcount_retry()
*/
#define __read_seqcount_begin(s) \
({ \
unsigned __seq; \
\
while (unlikely((__seq = seqprop_sequence(s)) & 1)) \
cpu_relax(); \
\
kcsan_atomic_next(KCSAN_SEQLOCK_REGION_MAX); \
__seq; \
})
/**
* raw_read_seqcount_begin() - begin a seqcount_t read section w/o lockdep
* @s: Pointer to seqcount_t or any of the seqcount_LOCKNAME_t variants
*
* Return: count to be passed to read_seqcount_retry()
*/
#define raw_read_seqcount_begin(s) __read_seqcount_begin(s)
/**
* read_seqcount_begin() - begin a seqcount_t read critical section
* @s: Pointer to seqcount_t or any of the seqcount_LOCKNAME_t variants
*
* Return: count to be passed to read_seqcount_retry()
*/
#define read_seqcount_begin(s) \
({ \
seqcount_lockdep_reader_access(seqprop_const_ptr(s)); \
raw_read_seqcount_begin(s); \
})
/**
* raw_read_seqcount() - read the raw seqcount_t counter value
* @s: Pointer to seqcount_t or any of the seqcount_LOCKNAME_t variants
*
* raw_read_seqcount opens a read critical section of the given
* seqcount_t, without any lockdep checking, and without checking or
* masking the sequence counter LSB. Calling code is responsible for
* handling that.
*
* Return: count to be passed to read_seqcount_retry()
*/
#define raw_read_seqcount(s) \
({ \
unsigned __seq = seqprop_sequence(s); \
\
kcsan_atomic_next(KCSAN_SEQLOCK_REGION_MAX); \
__seq; \
})
/**
* raw_seqcount_try_begin() - begin a seqcount_t read critical section
* w/o lockdep and w/o counter stabilization
* @s: Pointer to seqcount_t or any of the seqcount_LOCKNAME_t variants
* @start: count to be passed to read_seqcount_retry()
*
* Similar to raw_seqcount_begin(), except it enables eliding the critical
* section entirely if odd, instead of doing the speculation knowing it will
* fail.
*
* Useful when counter stabilization is more or less equivalent to taking
* the lock and there is a slowpath that does that.
*
* If true, start will be set to the (even) sequence count read.
*
* Return: true when a read critical section is started.
*/
#define raw_seqcount_try_begin(s, start) \
({ \
start = raw_read_seqcount(s); \
!(start & 1); \
})
/**
* raw_seqcount_begin() - begin a seqcount_t read critical section w/o
* lockdep and w/o counter stabilization
* @s: Pointer to seqcount_t or any of the seqcount_LOCKNAME_t variants
*
* raw_seqcount_begin opens a read critical section of the given
* seqcount_t. Unlike read_seqcount_begin(), this function will not wait
* for the count to stabilize. If a writer is active when it begins, it
* will fail the read_seqcount_retry() at the end of the read critical
* section instead of stabilizing at the beginning of it.
*
* Use this only in special kernel hot paths where the read section is
* small and has a high probability of success through other external
* means. It will save a single branching instruction.
*
* Return: count to be passed to read_seqcount_retry()
*/
#define raw_seqcount_begin(s) \
({ \
/* \
* If the counter is odd, let read_seqcount_retry() fail \
* by decrementing the counter. \
*/ \
raw_read_seqcount(s) & ~1; \
})
/**
* __read_seqcount_retry() - end a seqcount_t read section w/o barrier
* @s: Pointer to seqcount_t or any of the seqcount_LOCKNAME_t variants
* @start: count, from read_seqcount_begin()
*
* __read_seqcount_retry is like read_seqcount_retry, but has no smp_rmb()
* barrier. Callers should ensure that smp_rmb() or equivalent ordering is
* provided before actually loading any of the variables that are to be
* protected in this critical section.
*
* Use carefully, only in critical code, and comment how the barrier is
* provided.
*
* Return: true if a read section retry is required, else false
*/
#define __read_seqcount_retry(s, start) \
do___read_seqcount_retry(seqprop_const_ptr(s), start)
static inline int do___read_seqcount_retry(const seqcount_t *s, unsigned start)
{
kcsan_atomic_next(0);
return unlikely(READ_ONCE(s->sequence) != start);
}
/**
* read_seqcount_retry() - end a seqcount_t read critical section
* @s: Pointer to seqcount_t or any of the seqcount_LOCKNAME_t variants
* @start: count, from read_seqcount_begin()
*
* read_seqcount_retry closes the read critical section of given
* seqcount_t. If the critical section was invalid, it must be ignored
* (and typically retried).
*
* Return: true if a read section retry is required, else false
*/
#define read_seqcount_retry(s, start) \
do_read_seqcount_retry(seqprop_const_ptr(s), start)
static inline int do_read_seqcount_retry(const seqcount_t *s, unsigned start)
{
smp_rmb();
return do___read_seqcount_retry(s, start);
}
/**
* raw_write_seqcount_begin() - start a seqcount_t write section w/o lockdep
* @s: Pointer to seqcount_t or any of the seqcount_LOCKNAME_t variants
*
* Context: check write_seqcount_begin()
*/
#define raw_write_seqcount_begin(s) \
do { \
if (seqprop_preemptible(s)) \
preempt_disable(); \
\
do_raw_write_seqcount_begin(seqprop_ptr(s)); \
} while (0)
static inline void do_raw_write_seqcount_begin(seqcount_t *s)
{
kcsan_nestable_atomic_begin();
s->sequence++;
smp_wmb();
}
/**
* raw_write_seqcount_end() - end a seqcount_t write section w/o lockdep
* @s: Pointer to seqcount_t or any of the seqcount_LOCKNAME_t variants
*
* Context: check write_seqcount_end()
*/
#define raw_write_seqcount_end(s) \
do { \
do_raw_write_seqcount_end(seqprop_ptr(s)); \
\
if (seqprop_preemptible(s)) \
preempt_enable(); \
} while (0)
static inline void do_raw_write_seqcount_end(seqcount_t *s)
{
smp_wmb();
s->sequence++;
kcsan_nestable_atomic_end();
}
/**
* write_seqcount_begin_nested() - start a seqcount_t write section with
* custom lockdep nesting level
* @s: Pointer to seqcount_t or any of the seqcount_LOCKNAME_t variants
* @subclass: lockdep nesting level
*
* See Documentation/locking/lockdep-design.rst
* Context: check write_seqcount_begin()
*/
#define write_seqcount_begin_nested(s, subclass) \
do { \
seqprop_assert(s); \
\
if (seqprop_preemptible(s)) \
preempt_disable(); \
\
do_write_seqcount_begin_nested(seqprop_ptr(s), subclass); \
} while (0)
static inline void do_write_seqcount_begin_nested(seqcount_t *s, int subclass)
{
seqcount_acquire(&s->dep_map, subclass, 0, _RET_IP_);
do_raw_write_seqcount_begin(s);
}
/**
* write_seqcount_begin() - start a seqcount_t write side critical section
* @s: Pointer to seqcount_t or any of the seqcount_LOCKNAME_t variants
*
* Context: sequence counter write side sections must be serialized and
* non-preemptible. Preemption will be automatically disabled if and
* only if the seqcount write serialization lock is associated, and
* preemptible. If readers can be invoked from hardirq or softirq
* context, interrupts or bottom halves must be respectively disabled.
*/
#define write_seqcount_begin(s) \
do { \
seqprop_assert(s); \
\
if (seqprop_preemptible(s)) \
preempt_disable(); \
\
do_write_seqcount_begin(seqprop_ptr(s)); \
} while (0)
static inline void do_write_seqcount_begin(seqcount_t *s)
{
do_write_seqcount_begin_nested(s, 0);
}
/**
* write_seqcount_end() - end a seqcount_t write side critical section
* @s: Pointer to seqcount_t or any of the seqcount_LOCKNAME_t variants
*
* Context: Preemption will be automatically re-enabled if and only if
* the seqcount write serialization lock is associated, and preemptible.
*/
#define write_seqcount_end(s) \
do { \
do_write_seqcount_end(seqprop_ptr(s)); \
\
if (seqprop_preemptible(s)) \
preempt_enable(); \
} while (0)
static inline void do_write_seqcount_end(seqcount_t *s)
{
seqcount_release(&s->dep_map, _RET_IP_);
do_raw_write_seqcount_end(s);
}
/**
* raw_write_seqcount_barrier() - do a seqcount_t write barrier
* @s: Pointer to seqcount_t or any of the seqcount_LOCKNAME_t variants
*
* This can be used to provide an ordering guarantee instead of the usual
* consistency guarantee. It is one wmb cheaper, because it can collapse
* the two back-to-back wmb()s.
*
* Note that writes surrounding the barrier should be declared atomic (e.g.
* via WRITE_ONCE): a) to ensure the writes become visible to other threads
* atomically, avoiding compiler optimizations; b) to document which writes are
* meant to propagate to the reader critical section. This is necessary because
* neither writes before nor after the barrier are enclosed in a seq-writer
* critical section that would ensure readers are aware of ongoing writes::
*
* seqcount_t seq;
* bool X = true, Y = false;
*
* void read(void)
* {
* bool x, y;
*
* do {
* int s = read_seqcount_begin(&seq);
*
* x = X; y = Y;
*
* } while (read_seqcount_retry(&seq, s));
*
* BUG_ON(!x && !y);
* }
*
* void write(void)
* {
* WRITE_ONCE(Y, true);
*
* raw_write_seqcount_barrier(seq);
*
* WRITE_ONCE(X, false);
* }
*/
#define raw_write_seqcount_barrier(s) \
do_raw_write_seqcount_barrier(seqprop_ptr(s))
static inline void do_raw_write_seqcount_barrier(seqcount_t *s)
{
kcsan_nestable_atomic_begin();
s->sequence++;
smp_wmb();
s->sequence++;
kcsan_nestable_atomic_end();
}
/**
* write_seqcount_invalidate() - invalidate in-progress seqcount_t read
* side operations
* @s: Pointer to seqcount_t or any of the seqcount_LOCKNAME_t variants
*
* After write_seqcount_invalidate, no seqcount_t read side operations
* will complete successfully and see data older than this.
*/
#define write_seqcount_invalidate(s) \
do_write_seqcount_invalidate(seqprop_ptr(s))
static inline void do_write_seqcount_invalidate(seqcount_t *s)
{
smp_wmb();
kcsan_nestable_atomic_begin();
s->sequence+=2;
kcsan_nestable_atomic_end();
}
/*
* Latch sequence counters (seqcount_latch_t)
*
* A sequence counter variant where the counter even/odd value is used to
* switch between two copies of protected data. This allows the read path,
* typically NMIs, to safely interrupt the write side critical section.
*
* As the write sections are fully preemptible, no special handling for
* PREEMPT_RT is needed.
*/
typedef struct {
seqcount_t seqcount;
} seqcount_latch_t;
/**
* SEQCNT_LATCH_ZERO() - static initializer for seqcount_latch_t
* @seq_name: Name of the seqcount_latch_t instance
*/
#define SEQCNT_LATCH_ZERO(seq_name) { \
.seqcount = SEQCNT_ZERO(seq_name.seqcount), \
}
/**
* seqcount_latch_init() - runtime initializer for seqcount_latch_t
* @s: Pointer to the seqcount_latch_t instance
*/
#define seqcount_latch_init(s) seqcount_init(&(s)->seqcount)
/**
* raw_read_seqcount_latch() - pick even/odd latch data copy
* @s: Pointer to seqcount_latch_t
*
* See raw_write_seqcount_latch() for details and a full reader/writer
* usage example.
*
* Return: sequence counter raw value. Use the lowest bit as an index for
* picking which data copy to read. The full counter must then be checked
* with raw_read_seqcount_latch_retry().
*/
static __always_inline unsigned raw_read_seqcount_latch(const seqcount_latch_t *s)
{
/*
* Pairs with the first smp_wmb() in raw_write_seqcount_latch().
* Due to the dependent load, a full smp_rmb() is not needed.
*/
return READ_ONCE(s->seqcount.sequence);
}
/**
* read_seqcount_latch() - pick even/odd latch data copy
* @s: Pointer to seqcount_latch_t
*
* See write_seqcount_latch() for details and a full reader/writer usage
* example.
*
* Return: sequence counter raw value. Use the lowest bit as an index for
* picking which data copy to read. The full counter must then be checked
* with read_seqcount_latch_retry().
*/
static __always_inline unsigned read_seqcount_latch(const seqcount_latch_t *s)
{
kcsan_atomic_next(KCSAN_SEQLOCK_REGION_MAX);
return raw_read_seqcount_latch(s);
}
/**
* raw_read_seqcount_latch_retry() - end a seqcount_latch_t read section
* @s: Pointer to seqcount_latch_t
* @start: count, from raw_read_seqcount_latch()
*
* Return: true if a read section retry is required, else false
*/
static __always_inline int
raw_read_seqcount_latch_retry(const seqcount_latch_t *s, unsigned start)
{
smp_rmb();
return unlikely(READ_ONCE(s->seqcount.sequence) != start);
}
/**
* read_seqcount_latch_retry() - end a seqcount_latch_t read section
* @s: Pointer to seqcount_latch_t
* @start: count, from read_seqcount_latch()
*
* Return: true if a read section retry is required, else false
*/
static __always_inline int
read_seqcount_latch_retry(const seqcount_latch_t *s, unsigned start)
{
kcsan_atomic_next(0);
return raw_read_seqcount_latch_retry(s, start);
}
/**
* raw_write_seqcount_latch() - redirect latch readers to even/odd copy
* @s: Pointer to seqcount_latch_t
*/
static __always_inline void raw_write_seqcount_latch(seqcount_latch_t *s)
{
smp_wmb(); /* prior stores before incrementing "sequence" */
s->seqcount.sequence++;
smp_wmb(); /* increment "sequence" before following stores */
}
/**
* write_seqcount_latch_begin() - redirect latch readers to odd copy
* @s: Pointer to seqcount_latch_t
*
* The latch technique is a multiversion concurrency control method that allows
* queries during non-atomic modifications. If you can guarantee queries never
* interrupt the modification -- e.g. the concurrency is strictly between CPUs
* -- you most likely do not need this.
*
* Where the traditional RCU/lockless data structures rely on atomic
* modifications to ensure queries observe either the old or the new state the
* latch allows the same for non-atomic updates. The trade-off is doubling the
* cost of storage; we have to maintain two copies of the entire data
* structure.
*
* Very simply put: we first modify one copy and then the other. This ensures
* there is always one copy in a stable state, ready to give us an answer.
*
* The basic form is a data structure like::
*
* struct latch_struct {
* seqcount_latch_t seq;
* struct data_struct data[2];
* };
*
* Where a modification, which is assumed to be externally serialized, does the
* following::
*
* void latch_modify(struct latch_struct *latch, ...)
* {
* write_seqcount_latch_begin(&latch->seq);
* modify(latch->data[0], ...);
* write_seqcount_latch(&latch->seq);
* modify(latch->data[1], ...);
* write_seqcount_latch_end(&latch->seq);
* }
*
* The query will have a form like::
*
* struct entry *latch_query(struct latch_struct *latch, ...)
* {
* struct entry *entry;
* unsigned seq, idx;
*
* do {
* seq = read_seqcount_latch(&latch->seq);
*
* idx = seq & 0x01;
* entry = data_query(latch->data[idx], ...);
*
* // This includes needed smp_rmb()
* } while (read_seqcount_latch_retry(&latch->seq, seq));
*
* return entry;
* }
*
* So during the modification, queries are first redirected to data[1]. Then we
* modify data[0]. When that is complete, we redirect queries back to data[0]
* and we can modify data[1].
*
* NOTE:
*
* The non-requirement for atomic modifications does _NOT_ include
* the publishing of new entries in the case where data is a dynamic
* data structure.
*
* An iteration might start in data[0] and get suspended long enough
* to miss an entire modification sequence, once it resumes it might
* observe the new entry.
*
* NOTE2:
*
* When data is a dynamic data structure; one should use regular RCU
* patterns to manage the lifetimes of the objects within.
*/
static __always_inline void write_seqcount_latch_begin(seqcount_latch_t *s)
{
kcsan_nestable_atomic_begin();
raw_write_seqcount_latch(s);
}
/**
* write_seqcount_latch() - redirect latch readers to even copy
* @s: Pointer to seqcount_latch_t
*/
static __always_inline void write_seqcount_latch(seqcount_latch_t *s)
{
raw_write_seqcount_latch(s);
}
/**
* write_seqcount_latch_end() - end a seqcount_latch_t write section
* @s: Pointer to seqcount_latch_t
*
* Marks the end of a seqcount_latch_t writer section, after all copies of the
* latch-protected data have been updated.
*/
static __always_inline void write_seqcount_latch_end(seqcount_latch_t *s)
{
kcsan_nestable_atomic_end();
}
#define __SEQLOCK_UNLOCKED(lockname) \
{ \
.seqcount = SEQCNT_SPINLOCK_ZERO(lockname, &(lockname).lock), \
.lock = __SPIN_LOCK_UNLOCKED(lockname) \
}
/**
* seqlock_init() - dynamic initializer for seqlock_t
* @sl: Pointer to the seqlock_t instance
*/
#define seqlock_init(sl) \
do { \
spin_lock_init(&(sl)->lock); \
seqcount_spinlock_init(&(sl)->seqcount, &(sl)->lock); \
} while (0)
/**
* DEFINE_SEQLOCK(sl) - Define a statically allocated seqlock_t
* @sl: Name of the seqlock_t instance
*/
#define DEFINE_SEQLOCK(sl) \
seqlock_t sl = __SEQLOCK_UNLOCKED(sl)
/**
* read_seqbegin() - start a seqlock_t read side critical section
* @sl: Pointer to seqlock_t
*
* Return: count, to be passed to read_seqretry()
*/
static inline unsigned read_seqbegin(const seqlock_t *sl)
__acquires_shared(sl) __no_context_analysis
{
return read_seqcount_begin(&sl->seqcount);
}
/**
* read_seqretry() - end a seqlock_t read side section
* @sl: Pointer to seqlock_t
* @start: count, from read_seqbegin()
*
* read_seqretry closes the read side critical section of given seqlock_t.
* If the critical section was invalid, it must be ignored (and typically
* retried).
*
* Return: true if a read section retry is required, else false
*/
static inline unsigned read_seqretry(const seqlock_t *sl, unsigned start)
__releases_shared(sl) __no_context_analysis
{
return read_seqcount_retry(&sl->seqcount, start);
}
/*
* For all seqlock_t write side functions, use the internal
* do_write_seqcount_begin() instead of generic write_seqcount_begin().
* This way, no redundant lockdep_assert_held() checks are added.
*/
/**
* write_seqlock() - start a seqlock_t write side critical section
* @sl: Pointer to seqlock_t
*
* write_seqlock opens a write side critical section for the given
* seqlock_t. It also implicitly acquires the spinlock_t embedded inside
* that sequential lock. All seqlock_t write side sections are thus
* automatically serialized and non-preemptible.
*
* Context: if the seqlock_t read section, or other write side critical
* sections, can be invoked from hardirq or softirq contexts, use the
* _irqsave or _bh variants of this function instead.
*/
static inline void write_seqlock(seqlock_t *sl)
__acquires(sl) __no_context_analysis
{
spin_lock(&sl->lock);
do_write_seqcount_begin(&sl->seqcount.seqcount);
}
/**
* write_sequnlock() - end a seqlock_t write side critical section
* @sl: Pointer to seqlock_t
*
* write_sequnlock closes the (serialized and non-preemptible) write side
* critical section of given seqlock_t.
*/
static inline void write_sequnlock(seqlock_t *sl)
__releases(sl) __no_context_analysis
{
do_write_seqcount_end(&sl->seqcount.seqcount);
spin_unlock(&sl->lock);
}
/**
* write_seqlock_bh() - start a softirqs-disabled seqlock_t write section
* @sl: Pointer to seqlock_t
*
* _bh variant of write_seqlock(). Use only if the read side section, or
* other write side sections, can be invoked from softirq contexts.
*/
static inline void write_seqlock_bh(seqlock_t *sl)
__acquires(sl) __no_context_analysis
{
spin_lock_bh(&sl->lock);
do_write_seqcount_begin(&sl->seqcount.seqcount);
}
/**
* write_sequnlock_bh() - end a softirqs-disabled seqlock_t write section
* @sl: Pointer to seqlock_t
*
* write_sequnlock_bh closes the serialized, non-preemptible, and
* softirqs-disabled, seqlock_t write side critical section opened with
* write_seqlock_bh().
*/
static inline void write_sequnlock_bh(seqlock_t *sl)
__releases(sl) __no_context_analysis
{
do_write_seqcount_end(&sl->seqcount.seqcount);
spin_unlock_bh(&sl->lock);
}
/**
* write_seqlock_irq() - start a non-interruptible seqlock_t write section
* @sl: Pointer to seqlock_t
*
* _irq variant of write_seqlock(). Use only if the read side section, or
* other write sections, can be invoked from hardirq contexts.
*/
static inline void write_seqlock_irq(seqlock_t *sl)
__acquires(sl) __no_context_analysis
{
spin_lock_irq(&sl->lock);
do_write_seqcount_begin(&sl->seqcount.seqcount);
}
/**
* write_sequnlock_irq() - end a non-interruptible seqlock_t write section
* @sl: Pointer to seqlock_t
*
* write_sequnlock_irq closes the serialized and non-interruptible
* seqlock_t write side section opened with write_seqlock_irq().
*/
static inline void write_sequnlock_irq(seqlock_t *sl)
__releases(sl) __no_context_analysis
{
do_write_seqcount_end(&sl->seqcount.seqcount);
spin_unlock_irq(&sl->lock);
}
static inline unsigned long __write_seqlock_irqsave(seqlock_t *sl)
__acquires(sl) __no_context_analysis
{
unsigned long flags;
spin_lock_irqsave(&sl->lock, flags);
do_write_seqcount_begin(&sl->seqcount.seqcount);
return flags;
}
/**
* write_seqlock_irqsave() - start a non-interruptible seqlock_t write
* section
* @lock: Pointer to seqlock_t
* @flags: Stack-allocated storage for saving caller's local interrupt
* state, to be passed to write_sequnlock_irqrestore().
*
* _irqsave variant of write_seqlock(). Use it only if the read side
* section, or other write sections, can be invoked from hardirq context.
*/
#define write_seqlock_irqsave(lock, flags) \
do { flags = __write_seqlock_irqsave(lock); } while (0)
/**
* write_sequnlock_irqrestore() - end non-interruptible seqlock_t write
* section
* @sl: Pointer to seqlock_t
* @flags: Caller's saved interrupt state, from write_seqlock_irqsave()
*
* write_sequnlock_irqrestore closes the serialized and non-interruptible
* seqlock_t write section previously opened with write_seqlock_irqsave().
*/
static inline void
write_sequnlock_irqrestore(seqlock_t *sl, unsigned long flags)
__releases(sl) __no_context_analysis
{
do_write_seqcount_end(&sl->seqcount.seqcount);
spin_unlock_irqrestore(&sl->lock, flags);
}
/**
* read_seqlock_excl() - begin a seqlock_t locking reader section
* @sl: Pointer to seqlock_t
*
* read_seqlock_excl opens a seqlock_t locking reader critical section. A
* locking reader exclusively locks out *both* other writers *and* other
* locking readers, but it does not update the embedded sequence number.
*
* Locking readers act like a normal spin_lock()/spin_unlock().
*
* Context: if the seqlock_t write section, *or other read sections*, can
* be invoked from hardirq or softirq contexts, use the _irqsave or _bh
* variant of this function instead.
*
* The opened read section must be closed with read_sequnlock_excl().
*/
static inline void read_seqlock_excl(seqlock_t *sl)
__acquires_shared(sl) __no_context_analysis
{
spin_lock(&sl->lock);
}
/**
* read_sequnlock_excl() - end a seqlock_t locking reader critical section
* @sl: Pointer to seqlock_t
*/
static inline void read_sequnlock_excl(seqlock_t *sl)
__releases_shared(sl) __no_context_analysis
{
spin_unlock(&sl->lock);
}
/**
* read_seqlock_excl_bh() - start a seqlock_t locking reader section with
* softirqs disabled
* @sl: Pointer to seqlock_t
*
* _bh variant of read_seqlock_excl(). Use this variant only if the
* seqlock_t write side section, *or other read sections*, can be invoked
* from softirq contexts.
*/
static inline void read_seqlock_excl_bh(seqlock_t *sl)
__acquires_shared(sl) __no_context_analysis
{
spin_lock_bh(&sl->lock);
}
/**
* read_sequnlock_excl_bh() - stop a seqlock_t softirq-disabled locking
* reader section
* @sl: Pointer to seqlock_t
*/
static inline void read_sequnlock_excl_bh(seqlock_t *sl)
__releases_shared(sl) __no_context_analysis
{
spin_unlock_bh(&sl->lock);
}
/**
* read_seqlock_excl_irq() - start a non-interruptible seqlock_t locking
* reader section
* @sl: Pointer to seqlock_t
*
* _irq variant of read_seqlock_excl(). Use this only if the seqlock_t
* write side section, *or other read sections*, can be invoked from a
* hardirq context.
*/
static inline void read_seqlock_excl_irq(seqlock_t *sl)
__acquires_shared(sl) __no_context_analysis
{
spin_lock_irq(&sl->lock);
}
/**
* read_sequnlock_excl_irq() - end an interrupts-disabled seqlock_t
* locking reader section
* @sl: Pointer to seqlock_t
*/
static inline void read_sequnlock_excl_irq(seqlock_t *sl)
__releases_shared(sl) __no_context_analysis
{
spin_unlock_irq(&sl->lock);
}
static inline unsigned long __read_seqlock_excl_irqsave(seqlock_t *sl)
__acquires_shared(sl) __no_context_analysis
{
unsigned long flags;
spin_lock_irqsave(&sl->lock, flags);
return flags;
}
/**
* read_seqlock_excl_irqsave() - start a non-interruptible seqlock_t
* locking reader section
* @lock: Pointer to seqlock_t
* @flags: Stack-allocated storage for saving caller's local interrupt
* state, to be passed to read_sequnlock_excl_irqrestore().
*
* _irqsave variant of read_seqlock_excl(). Use this only if the seqlock_t
* write side section, *or other read sections*, can be invoked from a
* hardirq context.
*/
#define read_seqlock_excl_irqsave(lock, flags) \
do { flags = __read_seqlock_excl_irqsave(lock); } while (0)
/**
* read_sequnlock_excl_irqrestore() - end non-interruptible seqlock_t
* locking reader section
* @sl: Pointer to seqlock_t
* @flags: Caller saved interrupt state, from read_seqlock_excl_irqsave()
*/
static inline void
read_sequnlock_excl_irqrestore(seqlock_t *sl, unsigned long flags)
__releases_shared(sl) __no_context_analysis
{
spin_unlock_irqrestore(&sl->lock, flags);
}
/**
* read_seqbegin_or_lock() - begin a seqlock_t lockless or locking reader
* @lock: Pointer to seqlock_t
* @seq : Marker and return parameter. If the passed value is even, the
* reader will become a *lockless* seqlock_t reader as in read_seqbegin().
* If the passed value is odd, the reader will become a *locking* reader
* as in read_seqlock_excl(). In the first call to this function, the
* caller *must* initialize and pass an even value to @seq; this way, a
* lockless read can be optimistically tried first.
*
* read_seqbegin_or_lock is an API designed to optimistically try a normal
* lockless seqlock_t read section first. If an odd counter is found, the
* lockless read trial has failed, and the next read iteration transforms
* itself into a full seqlock_t locking reader.
*
* This is typically used to avoid seqlock_t lockless readers starvation
* (too much retry loops) in the case of a sharp spike in write side
* activity.
*
* Context: if the seqlock_t write section, *or other read sections*, can
* be invoked from hardirq or softirq contexts, use the _irqsave or _bh
* variant of this function instead.
*
* Check Documentation/locking/seqlock.rst for template example code.
*
* Return: the encountered sequence counter value, through the @seq
* parameter, which is overloaded as a return parameter. This returned
* value must be checked with need_seqretry(). If the read section need to
* be retried, this returned value must also be passed as the @seq
* parameter of the next read_seqbegin_or_lock() iteration.
*/
static inline void read_seqbegin_or_lock(seqlock_t *lock, int *seq)
__acquires_shared(lock) __no_context_analysis
{
if (!(*seq & 1)) /* Even */
*seq = read_seqbegin(lock);
else /* Odd */
read_seqlock_excl(lock);
}
/**
* need_seqretry() - validate seqlock_t "locking or lockless" read section
* @lock: Pointer to seqlock_t
* @seq: sequence count, from read_seqbegin_or_lock()
*
* Return: true if a read section retry is required, false otherwise
*/
static inline int need_seqretry(seqlock_t *lock, int seq)
__releases_shared(lock) __no_context_analysis
{
return !(seq & 1) && read_seqretry(lock, seq);
}
/**
* done_seqretry() - end seqlock_t "locking or lockless" reader section
* @lock: Pointer to seqlock_t
* @seq: count, from read_seqbegin_or_lock()
*
* done_seqretry finishes the seqlock_t read side critical section started
* with read_seqbegin_or_lock() and validated by need_seqretry().
*/
static inline void done_seqretry(seqlock_t *lock, int seq)
__no_context_analysis
{
if (seq & 1)
read_sequnlock_excl(lock);
}
/**
* read_seqbegin_or_lock_irqsave() - begin a seqlock_t lockless reader, or
* a non-interruptible locking reader
* @lock: Pointer to seqlock_t
* @seq: Marker and return parameter. Check read_seqbegin_or_lock().
*
* This is the _irqsave variant of read_seqbegin_or_lock(). Use it only if
* the seqlock_t write section, *or other read sections*, can be invoked
* from hardirq context.
*
* Note: Interrupts will be disabled only for "locking reader" mode.
*
* Return:
*
* 1. The saved local interrupts state in case of a locking reader, to
* be passed to done_seqretry_irqrestore().
*
* 2. The encountered sequence counter value, returned through @seq
* overloaded as a return parameter. Check read_seqbegin_or_lock().
*/
static inline unsigned long
read_seqbegin_or_lock_irqsave(seqlock_t *lock, int *seq)
__acquires_shared(lock) __no_context_analysis
{
unsigned long flags = 0;
if (!(*seq & 1)) /* Even */
*seq = read_seqbegin(lock);
else /* Odd */
read_seqlock_excl_irqsave(lock, flags);
return flags;
}
/**
* done_seqretry_irqrestore() - end a seqlock_t lockless reader, or a
* non-interruptible locking reader section
* @lock: Pointer to seqlock_t
* @seq: Count, from read_seqbegin_or_lock_irqsave()
* @flags: Caller's saved local interrupt state in case of a locking
* reader, also from read_seqbegin_or_lock_irqsave()
*
* This is the _irqrestore variant of done_seqretry(). The read section
* must've been opened with read_seqbegin_or_lock_irqsave(), and validated
* by need_seqretry().
*/
static inline void
done_seqretry_irqrestore(seqlock_t *lock, int seq, unsigned long flags)
__no_context_analysis
{
if (seq & 1)
read_sequnlock_excl_irqrestore(lock, flags);
}
enum ss_state {
ss_done = 0,
ss_lock,
ss_lock_irqsave,
ss_lockless,
};
struct ss_tmp {
enum ss_state state;
unsigned long data;
spinlock_t *lock;
spinlock_t *lock_irqsave;
};
static __always_inline void __scoped_seqlock_cleanup(struct ss_tmp *sst)
__no_context_analysis
{
if (sst->lock)
spin_unlock(sst->lock);
if (sst->lock_irqsave)
spin_unlock_irqrestore(sst->lock_irqsave, sst->data);
}
extern void __scoped_seqlock_invalid_target(void);
#if (defined(CONFIG_CC_IS_GCC) && CONFIG_GCC_VERSION < 90000) || defined(CONFIG_KASAN)
/*
* For some reason some GCC-8 architectures (nios2, alpha) have trouble
* determining that the ss_done state is impossible in __scoped_seqlock_next()
* below.
*
* Similarly KASAN is known to confuse compilers enough to break this. But we
* don't care about code quality for KASAN builds anyway.
*/
static inline void __scoped_seqlock_bug(void) { }
#else
/*
* Canary for compiler optimization -- if the compiler doesn't realize this is
* an impossible state, it very likely generates sub-optimal code here.
*/
extern void __scoped_seqlock_bug(void);
#endif
static __always_inline void
__scoped_seqlock_next(struct ss_tmp *sst, seqlock_t *lock, enum ss_state target)
__no_context_analysis
{
switch (sst->state) {
case ss_done:
__scoped_seqlock_bug();
return;
case ss_lock:
case ss_lock_irqsave:
sst->state = ss_done;
return;
case ss_lockless:
if (!read_seqretry(lock, sst->data)) {
sst->state = ss_done;
return;
}
break;
}
switch (target) {
case ss_done:
__scoped_seqlock_invalid_target();
return;
case ss_lock:
sst->lock = &lock->lock;
spin_lock(sst->lock);
sst->state = ss_lock;
return;
case ss_lock_irqsave:
sst->lock_irqsave = &lock->lock;
spin_lock_irqsave(sst->lock_irqsave, sst->data);
sst->state = ss_lock_irqsave;
return;
case ss_lockless:
sst->data = read_seqbegin(lock);
return;
}
}
/*
* Context analysis no-op helper to release seqlock at the end of the for-scope;
* the alias analysis of the compiler will recognize that the pointer @s is an
* alias to @_seqlock passed to read_seqbegin(_seqlock) below.
*/
static __always_inline void __scoped_seqlock_cleanup_ctx(struct ss_tmp **s)
__releases_shared(*((seqlock_t **)s)) __no_context_analysis {}
#define __scoped_seqlock_read(_seqlock, _target, _s) \
for (struct ss_tmp _s __cleanup(__scoped_seqlock_cleanup) = \
{ .state = ss_lockless, .data = read_seqbegin(_seqlock) }, \
*__UNIQUE_ID(ctx) __cleanup(__scoped_seqlock_cleanup_ctx) =\
(struct ss_tmp *)_seqlock; \
_s.state != ss_done; \
__scoped_seqlock_next(&_s, _seqlock, _target))
/**
* scoped_seqlock_read() - execute the read-side critical section
* without manual sequence counter handling
* or calls to other helpers
* @_seqlock: pointer to seqlock_t protecting the data
* @_target: an enum ss_state: one of {ss_lock, ss_lock_irqsave, ss_lockless}
* indicating the type of critical read section
*
* Example::
*
* scoped_seqlock_read (&lock, ss_lock) {
* // read-side critical section
* }
*
* Starts with a lockess pass first. If it fails, restarts the critical
* section with the lock held.
*/
#define scoped_seqlock_read(_seqlock, _target) \
__scoped_seqlock_read(_seqlock, _target, __UNIQUE_ID(seqlock))
DEFINE_LOCK_GUARD_1(seqlock_init, seqlock_t, seqlock_init(_T->lock), /* */)
DECLARE_LOCK_GUARD_1_ATTRS(seqlock_init, __acquires(_T), __releases(*(seqlock_t **)_T))
#define class_seqlock_init_constructor(_T) WITH_LOCK_GUARD_1_ATTRS(seqlock_init, _T)
#endif /* __LINUX_SEQLOCK_H */
]
|
| 25/2 |
2026/04/21 11:16 |
llm |
expert |
0m
Model:
Tokens:
input: 43374
output: 42
thoughts: 6501
Thoughts: **KCSAN Race Detection with Sequence Locks in Timekeeping**
Okay, so I'm diving into this KCSAN report. The gist is, KCSAN's flagged a potential data race involving reads and writes to `struct timekeeper` within the kernel's timekeeping mechanism, specifically around `ktime_get_coarse_real_ts64_mg`. The code uses sequence locks (`read_seqcount_begin`, `write_seqcount_begin`) to synchronize access, which *should* prevent these races. My initial thought was KCSAN was being overly sensitive, maybe the atomic region size, `KCSAN_SEQLOCK_REGION_MAX`, was too small, but it's 1000, which should be sufficient for the tiny loop in `ktime_get_coarse_real_ts64_mg`. I'm starting to suspect something more subtle.
The write stack trace points to `timekeeping_update_from_shadow`, which updates the `timekeeper` structure. It calls `memcpy` to copy `shadow_timekeeper` to `tk_core.timekeeper`. It's happening during a timer interrupt. The read stack trace shows the read happening in `ktime_get_coarse_real_ts64_mg`, and it's accessing `tk->coarse_nsec`. `tk` points to `&tk_core.timekeeper`. The read is protected by a sequence lock. The write also uses a sequence lock, so this *should* be safe. The read and write should be protected by the same sequence lock.
But then there's the `tk_core` macro, which expands to `timekeeper_data[TIMEKEEPER_CORE]`. I'm wondering if this array access is throwing KCSAN off. It might not be treating it as an "atomic" access. Even though `kcsan_atomic_next` is called, maybe it's missing something. The read and write addresses align, the write is to the `timekeeper` structure and the read is to `coarse_nsec`.
Then I checked `read_seqcount_begin` and confirmed that it *should* be protecting the access, via `kcsan_atomic_next`. The sizes make sense, the offsets make sense, the sequence lock is *present* on both sides of the access. The access uses `smp_load_acquire` which is correct for a sequence lock. I double checked the sizes of data, `timekeeper` and other structures, and they are consistent.
Then it clicked. In `ktime_get_coarse_real_ts64_mg` the read is `tk_core.timekeeper.offs_real`. Is that the same as the write sequence lock? Let's check the struct. Okay, the write is using the same `tk_core.seq` lock. So the code should be protected. So, this should be safe. But, the write is happening in an interrupt context. The read is not. KCSAN should not report this.
Let me compare with the original `ktime_get_coarse_real_ts64`, which doesn't have the `offs_real` access. Nope. Same structure, same protections. This seems *very* odd.
Now I am suspicious about how KCSAN interacts with the sequence lock in the context of an interrupt, and if the macro expansion of `tk_core` is creating the problem. Both are using the same seqcount, both operations are inside the sequence lock... I need to dig deeper into the KCSAN code and see if there is something specific to interrupts and macro expansion that might be triggering this false positive. Or, if there's a subtle interaction I'm missing.
|
| 26/2 |
2026/04/21 11:16 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/time/timekeeping.c IncludeLines:true Name:update_wall_time]
Results: map[SourceCode:2543: /**
2544: * update_wall_time - Uses the current clocksource to increment the wall time
2545: *
2546: * It also updates the enabled auxiliary clock timekeepers
2547: */
2548: void update_wall_time(void)
2549: {
2550: if (timekeeping_advance(TK_ADV_TICK))
2551: clock_was_set_delayed();
2552: tk_aux_advance();
2553: }
SourceFile:kernel/time/timekeeping.c]
|
| 27/2 |
2026/04/21 11:16 |
llm |
expert |
0m
Model:
Tokens:
input: 50902
output: 52
thoughts: 771
Thoughts: **Analyzing a Potential Race Condition in Timekeeping**
Okay, let's break this down. I'm looking at some code related to timekeeping, specifically the `timekeeping_advance` and `__timekeeping_advance` functions. Nothing immediately jumps out there, just the standard locking mechanism using a raw spinlock.
But then, the `write` stack trace is where things get interesting. I see an interrupt occurring right in the middle of a KCSAN access check, specifically within `pid_vnr`. The stack trace reveals the path: interrupt -> `instr_sysvec_apic_timer_interrupt` -> KCSAN's access check within `pid_vnr`, which is itself called from `scm_set_cred` and `__unix_dgram_recvmsg`. This is a normal timer interrupt, nothing unusual there, but the timing is suspicious.
Then I see it: a change in `coarse_nsec`. That's the key. KCSAN is reporting a change: `0x31af5fed -> 0x3247f66d`. This should *not* happen under normal circumstances. We use `read_seqcount_begin`/`retry` to handle these scenarios, ensuring that if a race *did* occur, it wouldn't be reported.
But, the access is covered by the seqcount. So why am I seeing a race?
The `_mg` suffix on `ktime_get_coarse_real_ts64_mg` is making me pause. This usually indicates something slightly different is happening. Let me check the full implementation of the function `ktime_get_coarse_real_ts64_mg`. I'll dig into that and see if there is any special behavior going on.
|
| 28/2 |
2026/04/21 11:16 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/time/timekeeping.c IncludeLines:true Name:ktime_get_coarse_real_ts64_mg]
Results: map[SourceCode:2588: /**
2589: * ktime_get_coarse_real_ts64_mg - return latter of coarse grained time or floor
2590: * @ts: timespec64 to be filled
2591: *
2592: * Fetch the global mg_floor value, convert it to realtime and compare it
2593: * to the current coarse-grained time. Fill @ts with whichever is
2594: * latest. Note that this is a filesystem-specific interface and should be
2595: * avoided outside of that context.
2596: */
2597: void ktime_get_coarse_real_ts64_mg(struct timespec64 *ts)
2598: {
2599: struct timekeeper *tk = &tk_core.timekeeper;
2600: u64 floor = atomic64_read(&mg_floor);
2601: ktime_t f_real, offset, coarse;
2602: unsigned int seq;
2603:
2604: do {
2605: seq = read_seqcount_begin(&tk_core.seq);
2606: *ts = tk_xtime_coarse(tk);
2607: offset = tk_core.timekeeper.offs_real;
2608: } while (read_seqcount_retry(&tk_core.seq, seq));
2609:
2610: coarse = timespec64_to_ktime(*ts);
2611: f_real = ktime_add(floor, offset);
2612: if (ktime_after(f_real, coarse))
2613: *ts = ktime_to_timespec64(f_real);
2614: }
SourceFile:kernel/time/timekeeping.c]
|
| 29/2 |
2026/04/21 11:16 |
llm |
expert |
0m
Model:
Tokens:
input: 52611
output: 37
thoughts: 2179
Thoughts: **Data Race Analysis with KCSAN and `seqcount`**
Okay, let's break this down. KCSAN is flagging a potential data race, and my initial thought is that it *shouldn't* be happening, because we're using `seqcount` for protection. Let me trace my logic:
The core of the issue seems to be around `ktime_get_coarse_real_ts64_mg` (and by extension `ktime_get_coarse_real_ts64`). In this function I'm seeing an access to `tk_core.timekeeper.offs_real` within a `seqcount` protected region. Specifically, `offset = tk_core.timekeeper.offs_real;` This seems odd, because `tk` is initialized as a pointer to `tk_core.timekeeper`.
The writer is `timekeeping_update_from_shadow` doing a `memcpy` of the entire `struct timekeeper` from a local shadow copy. The key is that `memcpy` overwrites the whole struct and this is protected using a `seqcount`.
I'm verifying that the read is properly annotated: `read_seqcount_begin` is correctly called, which *should* tell KCSAN that the read is atomic. And indeed, `write_seqcount_begin` and `write_seqcount_end` surround the `memcpy`. Inside these functions I see calls to `kcsan_nestable_atomic_begin()` so all of this should be atomic.
My assumption is that the `memcpy` and the reads in the `ktime_get_coarse_real_ts64_mg` function should be marked as atomic by KCSAN when `seqcount` is used. However, it seems KCSAN is reporting a race. Maybe it's not being detected due to the `memcpy` not being a "marked" access.
I double-checked to see that `read_seqcount_begin` is properly defined (it's using the standard header). I also see that `ktime_get_coarse_real_ts64_mg` uses `atomic64_read` outside the `seqcount` loop, but that doesn't seem to be the reported problem.
A key point is that the report says a 4-byte read, but the code does an 8-byte read of `tk_core.timekeeper.offs_real`. However, that access seems to be only within the `seqcount` loop.
I see that `tk_xtime_coarse` is an inline function. This could be an issue, but `kcsan_atomic_next` should be able to handle accesses from inlined functions.
Okay, let's zoom in on what KCSAN is reporting. The report is pointing to a 4-byte read of `coarse_nsec`. And indeed, that's what's happening *inside* the `seqcount` loop, via `tk_xtime_coarse(tk)`.
The crucial thing is that the direct access to `tk->coarse_nsec` is *not* using `READ_ONCE()` or similar. That's typical for `seqcount` readers, because of the retry loop, but KCSAN seems to be missing that context. I need to keep this in mind.
Finally, I remember that `ktime_get_coarse_real_ts64_mg` was recently added. Let me check the non-mg version, too. That should have the same structure.
Wait, I think I see a pattern. It's the `tk_core` expansion, perhaps, which is `timekeeper_data[TIMEKEEPER_CORE]`. The read in `ktime_get_coarse_real_ts64_mg` doesn't *directly* use `tk`, it uses `tk_core.timekeeper.offs_real`. Because of that, the whole `struct timekeeper` is not being directly accessed, which could be the source of the problem. However, both of them are accessing the same thing so I'm missing something here.
KCSAN is flagging a *true* data race. It's benign in that we're using a `seqcount` to handle it, but KCSAN is correctly pointing this out *because* the algorithm is not being properly annotated. So, I see that this is a race condition. It might be a benign one due to the nature of the algorithm itself, so there needs to be a deeper analysis to determine the severity and the correct way to fix it.
|
| 30/2 |
2026/04/21 11:17 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:fs/inode.c IncludeLines:true Name:current_time]
Results: map[SourceCode:2352: /**
2353: * current_time - Return FS time (possibly fine-grained)
2354: * @inode: inode.
2355: *
2356: * Return the current time truncated to the time granularity supported by
2357: * the fs, as suitable for a ctime/mtime change. If the ctime is flagged
2358: * as having been QUERIED, get a fine-grained timestamp, but don't update
2359: * the floor.
2360: *
2361: * For a multigrain inode, this is effectively an estimate of the timestamp
2362: * that a file would receive. An actual update must go through
2363: * inode_set_ctime_current().
2364: */
2365: struct timespec64 current_time(struct inode *inode)
2366: {
2367: struct timespec64 now;
2368: u32 cns;
2369:
2370: ktime_get_coarse_real_ts64_mg(&now);
2371:
2372: if (!is_mgtime(inode))
2373: goto out;
2374:
2375: /* If nothing has queried it, then coarse time is fine */
2376: cns = smp_load_acquire(&inode->i_ctime_nsec);
2377: if (cns & I_CTIME_QUERIED) {
2378: /*
2379: * If there is no apparent change, then get a fine-grained
2380: * timestamp.
2381: */
2382: if (now.tv_nsec == (cns & ~I_CTIME_QUERIED))
2383: ktime_get_real_ts64(&now);
2384: }
2385: out:
2386: return timestamp_truncate(now, inode);
2387: }
SourceFile:fs/inode.c]
|
| 31/2 |
2026/04/21 11:17 |
llm |
expert |
9m
Model:
Tokens:
input: 56388
output: 45
thoughts: 191
Thoughts: **Analyzing the Timekeeping Mechanism**
Okay, here we go. I'm seeing that `ktime_get_coarse_real_ts64_mg` is being employed here. That's the starting point. Now, let's refocus on the write path. Specifically, I'm examining the update process.
I see that `timekeeping_update_from_shadow` is called within `update_wall_time`. Then, `update_wall_time` is called by `tick_do_update_jiffies64`. This happens *every* tick. So, this is potentially a hot path.
The immediate race condition concerns `coarse_nsec`. `coarse_nsec` is updated in `timekeeping_update_from_shadow`, via a `memcpy`. Hold on... before the `memcpy`, it looks like `coarse_nsec` is calculated. It's done inside `timekeeping_update_from_shadow`:
```c
804: tk_update_ktime_data(tk);
```
Okay, now let's dig into `tk_update_ktime_data`. That's where the heart of the matter lies.
|
| 32/2 |
2026/04/21 11:26 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/time/timekeeping.c IncludeLines:true Name:tk_update_ktime_data]
Results: map[SourceCode: 711: /*
712: * Update the ktime_t based scalar nsec members of the timekeeper
713: */
714: static inline void tk_update_ktime_data(struct timekeeper *tk)
715: {
716: u64 seconds;
717: u32 nsec;
718:
719: /*
720: * The xtime based monotonic readout is:
721: * nsec = (xtime_sec + wtm_sec) * 1e9 + wtm_nsec + now();
722: * The ktime based monotonic readout is:
723: * nsec = base_mono + now();
724: * ==> base_mono = (xtime_sec + wtm_sec) * 1e9 + wtm_nsec
725: */
726: seconds = (u64)(tk->xtime_sec + tk->wall_to_monotonic.tv_sec);
727: nsec = (u32) tk->wall_to_monotonic.tv_nsec;
728: tk->tkr_mono.base = ns_to_ktime(seconds * NSEC_PER_SEC + nsec);
729:
730: /*
731: * The sum of the nanoseconds portions of xtime and
732: * wall_to_monotonic can be greater/equal one second. Take
733: * this into account before updating tk->ktime_sec.
734: */
735: nsec += (u32)(tk->tkr_mono.xtime_nsec >> tk->tkr_mono.shift);
736: if (nsec >= NSEC_PER_SEC)
737: seconds++;
738: tk->ktime_sec = seconds;
739:
740: /* Update the monotonic raw base */
741: tk->tkr_raw.base = ns_to_ktime(tk->raw_sec * NSEC_PER_SEC);
742: }
SourceFile:kernel/time/timekeeping.c]
|
| 33/2 |
2026/04/21 11:26 |
llm |
expert |
0m
Model:
Tokens:
input: 57468
output: 46
thoughts: 42
Thoughts: Okay, here's my take on those thoughts, framed as an expert's internal monologue:
**Investigating `coarse_nsec` Updates**
Hmm, I'm going through this code, and something's immediately jumping out at me. I'm looking for where `coarse_nsec` is actually updated, and I'm not seeing it here. That's a red flag. It's a key variable, and if it's not being regularly refreshed, the timing calculations downstream are going to be completely off. So, where *is* it getting its values? I need to backtrack, figure out the data flow. Let me run a quick search... aha! Found it: `tk_update_coarse_nsecs`. Okay, so it looks like that's the function responsible. Now I need to understand the timing of *that* call, and what's going into it. I need to make sure the values being set are accurate and that the update frequency is sufficient for the accuracy requirements of the system.
|
| 34/2 |
2026/04/21 11:26 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/time/timekeeping.c IncludeLines:true Name:tk_update_coarse_nsecs]
Results: map[SourceCode: 216: /*
217: * Update the nanoseconds part for the coarse time keepers. They can't rely
218: * on xtime_nsec because xtime_nsec could be adjusted by a small negative
219: * amount when the multiplication factor of the clock is adjusted, which
220: * could cause the coarse clocks to go slightly backwards. See
221: * timekeeping_apply_adjustment(). Thus we keep a separate copy for the coarse
222: * clockids which only is updated when the clock has been set or we have
223: * accumulated time.
224: */
225: static inline void tk_update_coarse_nsecs(struct timekeeper *tk)
226: {
227: tk->coarse_nsec = tk->tkr_mono.xtime_nsec >> tk->tkr_mono.shift;
228: }
SourceFile:kernel/time/timekeeping.c]
|
| 35/2 |
2026/04/21 11:26 |
llm |
expert |
0m
Model:
Tokens:
input: 58045
output: 46
thoughts: 32
Thoughts: Okay, here's my take on those thoughts, framed as my own:
**Digging into `tk_update_coarse_nsecs`**
Alright, so I'm trying to trace the lifecycle of this `tk_update_coarse_nsecs` function. My initial question is *where* does this thing actually get called? It's not always obvious, and a little digging is necessary. Right, that function has to be called from *somewhere*. My current understanding, based on the codebase I'm immersed in, is that `tk_update_coarse_nsecs` is invoked within `accumulate_nsecs_to_secs`. That clarifies one level of call hierarchy for me. Now, I can use that piece of information to investigate what `accumulate_nsecs_to_secs` does further, and what leads it to call that function.
|
| 36/2 |
2026/04/21 11:26 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/time/timekeeping.c IncludeLines:true Name:accumulate_nsecs_to_secs]
Results: map[SourceCode:2382: /*
2383: * accumulate_nsecs_to_secs - Accumulates nsecs into secs
2384: *
2385: * Helper function that accumulates the nsecs greater than a second
2386: * from the xtime_nsec field to the xtime_secs field.
2387: * It also calls into the NTP code to handle leapsecond processing.
2388: */
2389: static inline unsigned int accumulate_nsecs_to_secs(struct timekeeper *tk)
2390: {
2391: u64 nsecps = (u64)NSEC_PER_SEC << tk->tkr_mono.shift;
2392: unsigned int clock_set = 0;
2393:
2394: while (tk->tkr_mono.xtime_nsec >= nsecps) {
2395: int leap;
2396:
2397: tk->tkr_mono.xtime_nsec -= nsecps;
2398: tk->xtime_sec++;
2399:
2400: /*
2401: * Skip NTP update if this second was accumulated before,
2402: * i.e. xtime_nsec underflowed in timekeeping_adjust()
2403: */
2404: if (unlikely(tk->skip_second_overflow)) {
2405: tk->skip_second_overflow = 0;
2406: continue;
2407: }
2408:
2409: /* Figure out if its a leap sec and apply if needed */
2410: leap = second_overflow(tk->id, tk->xtime_sec);
2411: if (unlikely(leap)) {
2412: struct timespec64 ts;
2413:
2414: tk->xtime_sec += leap;
2415:
2416: ts.tv_sec = leap;
2417: ts.tv_nsec = 0;
2418: tk_set_wall_to_mono(tk,
2419: timespec64_sub(tk->wall_to_monotonic, ts));
2420:
2421: __timekeeping_set_tai_offset(tk, tk->tai_offset - leap);
2422:
2423: clock_set = TK_CLOCK_WAS_SET;
2424: }
2425: }
2426: return clock_set;
2427: }
SourceFile:kernel/time/timekeeping.c]
|
| 37/2 |
2026/04/21 11:26 |
llm |
expert |
0m
Model:
Tokens:
input: 59021
output: 23
thoughts: 17
Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience:
**Where's the Data? A Quick Diagnostic**
Hmm, that's odd. I was expecting to see the output right there, but it's not present. That's not the usual behavior. Let me take a moment to confirm that the expected data is, in fact, absent. This probably indicates an error at the data output stage, but let's assume nothing and work through a quick fault-finding process.
Okay, instead of manually searching the file system, I'll leverage `grep`. This will allow me to quickly search for the expected output string across the dataset and provide a rapid answer for my concerns. This seems the easiest path for such a quick investigation of the data, and will hopefully give me a quick answer on whether the data is present somewhere else, or if the process hasn't generated the needed data.
|
| 38/2 |
2026/04/21 11:26 |
tool |
grepper |
0m
Args: map[Expression:tk_update_coarse_nsecs]
Results: map[Output:kernel/time/timekeeping.c=207=static inline struct timespec64 tk_xtime_coarse(const struct timekeeper *tk)
--
kernel/time/timekeeping.c-224- */
kernel/time/timekeeping.c:225:static inline void tk_update_coarse_nsecs(struct timekeeper *tk)
kernel/time/timekeeping.c-226-{
--
kernel/time/timekeeping.c=230=static void tk_set_xtime(struct timekeeper *tk, const struct timespec64 *ts)
--
kernel/time/timekeeping.c-233- tk->tkr_mono.xtime_nsec = (u64)ts->tv_nsec << tk->tkr_mono.shift;
kernel/time/timekeeping.c:234: tk_update_coarse_nsecs(tk);
kernel/time/timekeeping.c-235-}
--
kernel/time/timekeeping.c=237=static void tk_xtime_add(struct timekeeper *tk, const struct timespec64 *ts)
--
kernel/time/timekeeping.c-241- tk_normalize_xtime(tk);
kernel/time/timekeeping.c:242: tk_update_coarse_nsecs(tk);
kernel/time/timekeeping.c-243-}
--
kernel/time/timekeeping.c=841=static void timekeeping_forward_now(struct timekeeper *tk)
--
kernel/time/timekeeping.c-859- }
kernel/time/timekeeping.c:860: tk_update_coarse_nsecs(tk);
kernel/time/timekeeping.c-861-}
--
kernel/time/timekeeping.c=2476=static bool __timekeeping_advance(struct tk_data *tkd, enum timekeeping_adv_mode mode)
--
kernel/time/timekeeping.c-2529- if (orig_offset != offset)
kernel/time/timekeeping.c:2530: tk_update_coarse_nsecs(tk);
kernel/time/timekeeping.c-2531-
]
|
| 39/2 |
2026/04/21 11:26 |
llm |
expert |
0m
Model:
Tokens:
input: 59799
output: 17
thoughts: 4340
Thoughts: **Analyzing a KCSAN Report: Timekeeping Data Race**
Okay, so I've got a KCSAN report indicating a data race in the kernel's timekeeping mechanism. The report points to a write happening in `timekeeping_update_from_shadow`, specifically the `memcpy` operation updating the real timekeeper from the shadow timekeeper. This write is triggered by an interrupt. The reported race involves a read from `ktime_get_coarse_real_ts64_mg`, and the access is on the `coarse_nsec` field of the timekeeper struct.
Initially, I figured this should be safe because `ktime_get_coarse_real_ts64_mg` uses a `seqcount` to protect its access. The code uses `read_seqcount_begin` and `read_seqcount_retry` surrounding the read. This seems like it should have avoided the race. I then notice the presence of `ktime_get_coarse_real_ts64` and wonder why the race isn't reported there, because it appears almost identical.
Then, I focused on the differences between `ktime_get_coarse_real_ts64_mg` and `ktime_get_coarse_real_ts64`. The critical difference is the presence of an access to `tk_core.timekeeper.offs_real` in `ktime_get_coarse_real_ts64_mg`. However, this is also protected by the `seqcount`. I confirm that the `seqcount` is correctly used, but then the stack trace for the write reveals that it's being triggered in an interrupt context. This explains the race, but not why KCSAN is reporting it despite `read_seqcount_begin`.
Investigating the `seqcount` protection, I confirm that `read_seqcount_begin` should mark the critical section as "atomic" from KCSAN's perspective. The `memcpy` in the `write` side also utilizes `write_seqcount_begin` and `write_seqcount_end`, ensuring the write is protected as well.
The race occurs because KCSAN detects an unprotected read of `coarse_nsec` (at offset 112). While the read is inside a loop guarded by `read_seqcount_begin` and `read_seqcount_retry`, it is NOT marked with a `READ_ONCE`. On the other side, `memcpy` is being done to the same `struct timekeeper`. While `memcpy` is atomic in the sense that the individual writes happen contiguously, KCSAN still reports it due to the lack of explicit marking of the reads and writes.
I suspect KCSAN is flagging the race because of this missing `READ_ONCE` on the read side. While the `seqcount` guarantees data consistency with the retry mechanism, KCSAN still reports it due to the inherent concurrent access without explicit marking (because KCSAN flags concurrent unmarked accesses). Interestingly, `timekeeping.c` is quite sensitive in the kernel, and this type of pattern is often fixed by adding either `READ_ONCE` or `WRITE_ONCE` or `data_race()`.
So, the next step is to examine why `ktime_get_coarse_real_ts64` does *not* report the same race, even though it's structured similarly. It is possible it just hasn't triggered yet, or maybe the specific context in which `ktime_get_coarse_real_ts64_mg` is called is more prone to this race. I would then check to determine if the functions in `timekeeping.c` make use of `data_race()` calls for similar reasons. The fix is probably to add `READ_ONCE` to the access of `tk->coarse_nsec`.
|
| 40/2 |
2026/04/21 11:27 |
tool |
grepper |
0m
Args: map[Expression:data_race]
Results: map[Output:
Full output is too long, showing 500 out of 1267 lines.
Use more precise expression if possible.
[Documentation/dev-tools/checkpatch.rst=456=Comments
--
Documentation/dev-tools/checkpatch.rst-476- **DATA_RACE**
Documentation/dev-tools/checkpatch.rst:477: Applications of data_race() should have a comment so as to document the
Documentation/dev-tools/checkpatch.rst-478- reasoning behind why it was deemed safe.
--
Documentation/dev-tools/kcsan.rst=87=the below options are available:
Documentation/dev-tools/kcsan.rst-88-
Documentation/dev-tools/kcsan.rst:89:* KCSAN understands the ``data_race(expr)`` annotation, which tells KCSAN that
Documentation/dev-tools/kcsan.rst-90- any data races due to accesses in ``expr`` should be ignored and resulting
--
Documentation/dev-tools/kcsan.rst-93-
Documentation/dev-tools/kcsan.rst:94:* Similar to ``data_race(...)``, the type qualifier ``__data_racy`` can be used
Documentation/dev-tools/kcsan.rst-95- to document that all data races due to accesses to a variable are intended
--
Documentation/translations/zh_CN/dev-tools/kcsan.rst=24=KCSAN 提供了几个其他的配置选项来自定义行为(见 ``lib/Kconfig.kcsan`` 中的各自的
--
Documentation/translations/zh_CN/dev-tools/kcsan.rst-87-
Documentation/translations/zh_CN/dev-tools/kcsan.rst:88:* KCSAN 支持使用 ``data_race(expr)`` 注解,这个注解告诉 KCSAN 任何由访问
Documentation/translations/zh_CN/dev-tools/kcsan.rst-89- ``expr`` 所引起的数据竞争都应该被忽略,其产生的行为后果被认为是安全的。请查阅
--
Documentation/translations/zh_CN/dev-tools/kcsan.rst-91-
Documentation/translations/zh_CN/dev-tools/kcsan.rst:92:* 与 ``data_race(...)`` 相似,可以使用类型限定符 ``__data_racy`` 来标记一个变量
Documentation/translations/zh_CN/dev-tools/kcsan.rst-93- ,所有访问该变量而导致的数据竞争都是故意为之并且应该被 KCSAN 忽略::
--
arch/powerpc/kernel/interrupt.c=103=static notrace void check_return_regs_valid(struct pt_regs *regs)
--
arch/powerpc/kernel/interrupt.c-171-
arch/powerpc/kernel/interrupt.c:172: if (!data_race(warned)) {
arch/powerpc/kernel/interrupt.c:173: data_race(warned = true);
arch/powerpc/kernel/interrupt.c-174- printk("%sSRR0 was: %lx should be: %lx\n", h, srr0, regs->nip);
--
arch/x86/mm/pat/set_memory.c=161=static inline void cpa_inc_4k_install(void)
arch/x86/mm/pat/set_memory.c-162-{
arch/x86/mm/pat/set_memory.c:163: data_race(cpa_4k_install++);
arch/x86/mm/pat/set_memory.c-164-}
--
drivers/block/loop.c=1851=static blk_status_t loop_queue_rq(struct blk_mq_hw_ctx *hctx,
--
drivers/block/loop.c-1859-
drivers/block/loop.c:1860: if (data_race(READ_ONCE(lo->lo_state)) != Lo_bound)
drivers/block/loop.c-1861- return BLK_STS_IOERR;
--
drivers/block/loop.c=2190=static int loop_control_get_free(int idx)
--
drivers/block/loop.c-2203- if (lo->idr_visible &&
drivers/block/loop.c:2204: data_race(READ_ONCE(lo->lo_state)) == Lo_unbound)
drivers/block/loop.c-2205- goto found;
--
drivers/block/zloop.c=887=static blk_status_t zloop_queue_rq(struct blk_mq_hw_ctx *hctx,
--
drivers/block/zloop.c-893-
drivers/block/zloop.c:894: if (data_race(READ_ONCE(zlo->state)) == Zlo_deleting)
drivers/block/zloop.c-895- return BLK_STS_IOERR;
--
drivers/net/wireless/ath/ath9k/htc_drv_txrx.c=1135=void ath9k_htc_rxep(void *drv_priv, struct sk_buff *skb,
--
drivers/net/wireless/ath/ath9k/htc_drv_txrx.c-1144- /* Check if ath9k_rx_init() completed. */
drivers/net/wireless/ath/ath9k/htc_drv_txrx.c:1145: if (!data_race(priv->rx.initialized))
drivers/net/wireless/ath/ath9k/htc_drv_txrx.c-1146- goto err;
--
drivers/net/wireless/ath/ath9k/wmi.c=138=void ath9k_wmi_event_tasklet(struct tasklet_struct *t)
--
drivers/net/wireless/ath/ath9k/wmi.c-158- /* Check if ath9k_htc_probe_device() completed. */
drivers/net/wireless/ath/ath9k/wmi.c:159: if (!data_race(priv->initialized)) {
drivers/net/wireless/ath/ath9k/wmi.c-160- kfree_skb(skb);
--
drivers/virtio/virtio_ring.c=3189=bool virtqueue_enable_cb_delayed(struct virtqueue *_vq)
--
drivers/virtio/virtio_ring.c-3193- if (vq->event_triggered)
drivers/virtio/virtio_ring.c:3194: data_race(vq->event_triggered = false);
drivers/virtio/virtio_ring.c-3195-
--
drivers/virtio/virtio_ring.c=3229=irqreturn_t vring_interrupt(int irq, void *_vq)
--
drivers/virtio/virtio_ring.c-3249- if (vq->event)
drivers/virtio/virtio_ring.c:3250: data_race(vq->event_triggered = true);
drivers/virtio/virtio_ring.c-3251-
--
fs/btrfs/block-group.c=1860=static int reclaim_bgs_cmp(void *unused, const struct list_head *a,
--
fs/btrfs/block-group.c-1874- */
fs/btrfs/block-group.c:1875: return data_race(bg1->used > bg2->used);
fs/btrfs/block-group.c-1876-}
--
fs/btrfs/block-rsv.h=108=static inline bool btrfs_block_rsv_full(const struct btrfs_block_rsv *rsv)
fs/btrfs/block-rsv.h-109-{
fs/btrfs/block-rsv.h:110: return data_race(rsv->full);
fs/btrfs/block-rsv.h-111-}
--
fs/btrfs/extent-tree.c=2084=static noinline int __btrfs_run_delayed_refs(struct btrfs_trans_handle *trans,
--
fs/btrfs/extent-tree.c-2101- */
fs/btrfs/extent-tree.c:2102: max_count = data_race(delayed_refs->num_heads_ready);
fs/btrfs/extent-tree.c-2103- min_bytes = U64_MAX;
--
fs/btrfs/file.c=138=int btrfs_drop_extents(struct btrfs_trans_handle *trans,
--
fs/btrfs/file.c-180-
fs/btrfs/file.c:181: if (data_race(args->start >= inode->disk_i_size) && !args->replace_extent)
fs/btrfs/file.c-182- modify_tree = 0;
--
fs/btrfs/space-info.c=1997=u64 btrfs_account_ro_block_groups_free_space(struct btrfs_space_info *sinfo)
--
fs/btrfs/space-info.c-2003- /* It's df, we don't care if it's racy */
fs/btrfs/space-info.c:2004: if (data_race(list_empty(&sinfo->ro_bgs)))
fs/btrfs/space-info.c-2005- return 0;
--
fs/btrfs/tree-log.c=3736=static int inode_logged(const struct btrfs_trans_handle *trans,
--
fs/btrfs/tree-log.c-3747- */
fs/btrfs/tree-log.c:3748: if (data_race(inode->logged_trans) == trans->transid)
fs/btrfs/tree-log.c-3749- return 1;
--
fs/btrfs/tree-log.c=4610=static void fill_inode_item(struct btrfs_trans_handle *trans,
--
fs/btrfs/tree-log.c-4626- * We access ->logged_trans directly since it was already set
fs/btrfs/tree-log.c:4627: * up in the call chain by btrfs_log_inode(), and data_race()
fs/btrfs/tree-log.c-4628- * to avoid false alerts from KCSAN and since it was set already
--
fs/btrfs/tree-log.c-4631- */
fs/btrfs/tree-log.c:4632: ASSERT(data_race(inode->logged_trans) > 0);
fs/btrfs/tree-log.c:4633: if (data_race(inode->logged_trans) < trans->transid)
fs/btrfs/tree-log.c-4634- gen = 0;
--
fs/btrfs/tree-log.c=5457=static int get_inode_size_to_log(struct btrfs_trans_handle *trans,
--
fs/btrfs/tree-log.c-5470- * Our caller called inode_logged(), so logged_trans is up to date.
fs/btrfs/tree-log.c:5471: * Use data_race() to silence any warning from KCSAN. Once logged_trans
fs/btrfs/tree-log.c-5472- * is set, it can only be reset to 0 after inode eviction.
fs/btrfs/tree-log.c-5473- */
fs/btrfs/tree-log.c:5474: if (data_race(inode->logged_trans) == trans->transid) {
fs/btrfs/tree-log.c-5475- ret = btrfs_search_slot(NULL, inode->root->log_root, &key, path, 0, 0);
--
fs/btrfs/zoned.c=3201=void btrfs_show_zoned_stats(struct btrfs_fs_info *fs_info, struct seq_file *seq)
--
fs/btrfs/zoned.c-3222-
fs/btrfs/zoned.c:3223: data_reloc_bg = data_race(fs_info->data_reloc_bg);
fs/btrfs/zoned.c-3224- if (data_reloc_bg)
--
fs/btrfs/zoned.c-3226- data_reloc_bg);
fs/btrfs/zoned.c:3227: treelog_bg = data_race(fs_info->treelog_bg);
fs/btrfs/zoned.c-3228- if (treelog_bg)
--
fs/inode.c=111=static unsigned long get_mg_ctime_updates(void)
--
fs/inode.c-116- for_each_possible_cpu(i)
fs/inode.c:117: sum += data_race(per_cpu(mg_ctime_updates, i));
fs/inode.c-118- return sum;
--
fs/inode.c=121=static unsigned long get_mg_fine_stamps(void)
--
fs/inode.c-126- for_each_possible_cpu(i)
fs/inode.c:127: sum += data_race(per_cpu(mg_fine_stamps, i));
fs/inode.c-128- return sum;
--
fs/inode.c=131=static unsigned long get_mg_ctime_swaps(void)
--
fs/inode.c-136- for_each_possible_cpu(i)
fs/inode.c:137: sum += data_race(per_cpu(mg_ctime_swaps, i));
fs/inode.c-138- return sum;
--
fs/iomap/direct-io.c=401=static int iomap_dio_bio_iter(struct iomap_iter *iter, struct iomap_dio *dio)
--
fs/iomap/direct-io.c-542- */
fs/iomap/direct-io.c:543: if (unlikely(data_race(dio->error)))
fs/iomap/direct-io.c-544- goto out;
--
fs/jbd2/transaction.c=312=static int start_this_handle(journal_t *journal, handle_t *handle,
--
fs/jbd2/transaction.c-343- */
fs/jbd2/transaction.c:344: if (!data_race(journal->j_running_transaction)) {
fs/jbd2/transaction.c-345- /*
--
fs/jbd2/transaction.c=1517=int jbd2_journal_dirty_metadata(handle_t *handle, struct buffer_head *bh)
--
fs/jbd2/transaction.c-1540- */
fs/jbd2/transaction.c:1541: if (data_race(jh->b_transaction != transaction &&
fs/jbd2/transaction.c-1542- jh->b_next_transaction != transaction)) {
--
fs/jbd2/transaction.c-1553- }
fs/jbd2/transaction.c:1554: if (data_race(jh->b_modified == 1)) {
fs/jbd2/transaction.c-1555- /* If it's in our transaction it must be in BJ_Metadata list. */
fs/jbd2/transaction.c:1556: if (data_race(jh->b_transaction == transaction &&
fs/jbd2/transaction.c-1557- jh->b_jlist != BJ_Metadata)) {
--
fs/xfs/xfs_trace.h=1330=DECLARE_EVENT_CLASS(xfs_dquot_class,
--
fs/xfs/xfs_trace.h-1358- __entry->flags = dqp->q_flags;
fs/xfs/xfs_trace.h:1359: __entry->nrefs = data_race(dqp->q_lockref.count);
fs/xfs/xfs_trace.h-1360-
--
include/linux/bpf.h=519=static inline void bpf_long_memcpy(void *dst, const void *src, u32 size)
--
include/linux/bpf.h-525- while (size--)
include/linux/bpf.h:526: data_race(*ldst++ = *lsrc++);
include/linux/bpf.h-527-}
--
include/linux/compiler.h=15=void ftrace_likely_update(struct ftrace_likely_data *f, int val,
--
include/linux/compiler.h-169-/**
include/linux/compiler.h:170: * data_race - mark an expression as containing intentional data races
include/linux/compiler.h-171- *
include/linux/compiler.h:172: * This data_race() macro is useful for situations in which data races
include/linux/compiler.h-173- * should be forgiven. One example is diagnostic code that accesses
--
include/linux/compiler.h-177- * be plain C-language accesses and those in the diagnostic code should
include/linux/compiler.h:178: * use data_race(). This way, KCSAN will complain if buggy lockless
include/linux/compiler.h-179- * accesses to that variable are introduced, even if the buggy accesses
--
include/linux/compiler.h-183- * to tooling that data races here are to be ignored. If the access must
include/linux/compiler.h:184: * be atomic *and* KCSAN should ignore the access, use both data_race()
include/linux/compiler.h:185: * and READ_ONCE(), for example, data_race(READ_ONCE(x)).
include/linux/compiler.h-186- */
include/linux/compiler.h:187:#define data_race(expr) \
include/linux/compiler.h-188-({ \
--
include/linux/console.h=522=static inline short console_srcu_read_flags(const struct console *con)
--
include/linux/console.h-529- */
include/linux/console.h:530: return data_race(READ_ONCE(con->flags));
include/linux/console.h-531-}
--
include/linux/mm.h=3382=static inline void update_hiwater_rss(struct mm_struct *mm)
--
include/linux/mm.h-3385-
include/linux/mm.h:3386: if (data_race(mm->hiwater_rss) < _rss)
include/linux/mm.h:3387: data_race(mm->hiwater_rss = _rss);
include/linux/mm.h-3388-}
--
include/linux/page-flags.h=942=static inline bool page_has_type(const struct page *page)
include/linux/page-flags.h-943-{
include/linux/page-flags.h:944: return page_type_has_type(data_race(page->page_type));
include/linux/page-flags.h-945-}
--
include/linux/page-flags.h=948=static __always_inline bool folio_test_##fname(const struct folio *folio) \
include/linux/page-flags.h-949-{ \
include/linux/page-flags.h:950: return data_race(folio->page.page_type >> 24) == PGTY_##lname; \
include/linux/page-flags.h-951-} \
include/linux/page-flags.h=952=static __always_inline void __folio_set_##fname(struct folio *folio) \
--
include/linux/page-flags.h-955- return; \
include/linux/page-flags.h:956: VM_BUG_ON_FOLIO(data_race(folio->page.page_type) != UINT_MAX, \
include/linux/page-flags.h-957- folio); \
--
include/linux/page-flags.h=970=static __always_inline int Page##uname(const struct page *page) \
include/linux/page-flags.h-971-{ \
include/linux/page-flags.h:972: return data_race(page->page_type >> 24) == PGTY_##lname; \
include/linux/page-flags.h-973-} \
include/linux/page-flags.h=974=static __always_inline void __SetPage##uname(struct page *page) \
--
include/linux/page-flags.h-977- return; \
include/linux/page-flags.h:978: VM_BUG_ON_PAGE(data_race(page->page_type) != UINT_MAX, page); \
include/linux/page-flags.h-979- page->page_type = (unsigned int)PGTY_##lname << 24; \
--
include/linux/ptr_ring.h=49=static inline bool __ptr_ring_full(struct ptr_ring *r)
include/linux/ptr_ring.h-50-{
include/linux/ptr_ring.h:51: return data_race(r->queue[r->producer]);
include/linux/ptr_ring.h-52-}
--
include/linux/ptr_ring.h=104=static inline int __ptr_ring_produce(struct ptr_ring *r, void *ptr)
include/linux/ptr_ring.h-105-{
include/linux/ptr_ring.h:106: if (unlikely(!r->size) || data_race(r->queue[r->producer]))
include/linux/ptr_ring.h-107- return -ENOSPC;
--
include/linux/ptr_ring.h=194=static inline bool __ptr_ring_empty(struct ptr_ring *r)
--
include/linux/ptr_ring.h-196- if (likely(r->size))
include/linux/ptr_ring.h:197: return !data_race(r->queue[READ_ONCE(r->consumer_head)]);
include/linux/ptr_ring.h-198- return true;
--
include/linux/ptr_ring.h=249=static inline void __ptr_ring_zero_tail(struct ptr_ring *r, int consumer_head)
--
include/linux/ptr_ring.h-258- while (likely(head > r->consumer_tail))
include/linux/ptr_ring.h:259: data_race(r->queue[--head] = NULL);
include/linux/ptr_ring.h-260-
--
include/linux/rseq_entry.h=114=static __always_inline bool __rseq_grant_slice_extension(bool work_pending)
--
include/linux/rseq_entry.h-163- /* Store expiry time for arming the timer on the way out */
include/linux/rseq_entry.h:164: curr->rseq.slice.expires = data_race(rseq_slice_ext_nsecs) + ktime_get_mono_fast_ns();
include/linux/rseq_entry.h-165- /*
--
include/linux/rtmutex.h=42=static inline bool rt_mutex_base_is_locked(struct rt_mutex_base *lock)
include/linux/rtmutex.h-43-{
include/linux/rtmutex.h:44: return data_race(READ_ONCE(lock->owner) != NULL);
include/linux/rtmutex.h-45-}
--
include/linux/rtmutex.h=50=static inline struct task_struct *rt_mutex_owner(struct rt_mutex_base *lock)
include/linux/rtmutex.h-51-{
include/linux/rtmutex.h:52: unsigned long owner = (unsigned long) data_race(READ_ONCE(lock->owner));
include/linux/rtmutex.h-53-
--
include/linux/rwsem.h=132=static inline bool rwsem_is_contended(struct rw_semaphore *sem)
include/linux/rwsem.h-133-{
include/linux/rwsem.h:134: return data_race(sem->first_waiter != NULL);
include/linux/rwsem.h-135-}
--
include/linux/srcutiny.h=143=static inline void srcu_torture_stats_print(struct srcu_struct *ssp,
--
include/linux/srcutiny.h-147-
include/linux/srcutiny.h:148: idx = ((data_race(READ_ONCE(ssp->srcu_idx)) + 1) & 0x2) >> 1;
include/linux/srcutiny.h-149- pr_alert("%s%s Tiny SRCU per-CPU(idx=%d): (%hd,%hd) gp: %lu->%lu\n",
include/linux/srcutiny.h-150- tt, tf, idx,
include/linux/srcutiny.h:151: data_race(READ_ONCE(ssp->srcu_lock_nesting[!idx])),
include/linux/srcutiny.h:152: data_race(READ_ONCE(ssp->srcu_lock_nesting[idx])),
include/linux/srcutiny.h:153: data_race(READ_ONCE(ssp->srcu_idx)),
include/linux/srcutiny.h:154: data_race(READ_ONCE(ssp->srcu_idx_max)));
include/linux/srcutiny.h-155-}
--
include/net/sock.h=2531=static inline int sock_error(struct sock *sk)
--
include/net/sock.h-2537- */
include/net/sock.h:2538: if (likely(data_race(!sk->sk_err)))
include/net/sock.h-2539- return 0;
--
io_uring/fdinfo.c=60=static void __io_uring_show_fdinfo(struct io_ring_ctx *ctx, struct seq_file *m)
--
io_uring/fdinfo.c-86- seq_printf(m, "SqTail:\t%u\n", sq_tail);
io_uring/fdinfo.c:87: seq_printf(m, "CachedSqHead:\t%u\n", data_race(ctx->cached_sq_head));
io_uring/fdinfo.c-88- seq_printf(m, "CqMask:\t0x%x\n", cq_mask);
--
io_uring/fdinfo.c-90- seq_printf(m, "CqTail:\t%u\n", cq_tail);
io_uring/fdinfo.c:91: seq_printf(m, "CachedCqTail:\t%u\n", data_race(ctx->cached_cq_tail));
io_uring/fdinfo.c-92- seq_printf(m, "SQEs:\t%u\n", sq_tail - sq_head);
--
io_uring/io_uring.c=2243=static __poll_t io_uring_poll(struct file *file, poll_table *wait)
--
io_uring/io_uring.c-2247-
io_uring/io_uring.c:2248: if (unlikely(!(data_race(ctx->int_flags) & IO_RING_F_POLL_ACTIVATED)))
io_uring/io_uring.c-2249- io_activate_pollwq(ctx);
--
io_uring/io_uring.h=501=static inline void io_commit_cqring_flush(struct io_ring_ctx *ctx)
io_uring/io_uring.h-502-{
io_uring/io_uring.h:503: if (unlikely(data_race(ctx->int_flags) & SHOULD_FLUSH_MASK))
io_uring/io_uring.h-504- __io_commit_cqring_flush(ctx);
--
io_uring/loop.h=20=static inline bool io_has_loop_ops(struct io_ring_ctx *ctx)
io_uring/loop.h-21-{
io_uring/loop.h:22: return data_race(ctx->loop_step);
io_uring/loop.h-23-}
--
io_uring/refs.h=21=static inline bool req_ref_put_and_test_atomic(struct io_kiocb *req)
io_uring/refs.h-22-{
io_uring/refs.h:23: WARN_ON_ONCE(!(data_race(req->flags) & REQ_F_REFCOUNT));
io_uring/refs.h-24- WARN_ON_ONCE(req_ref_zero_or_close_to_overflow(req));
--
io_uring/timeout.c=620=int io_timeout(struct io_kiocb *req, unsigned int issue_flags)
--
io_uring/timeout.c-639-
io_uring/timeout.c:640: tail = data_race(ctx->cached_cq_tail) - atomic_read(&ctx->cq_timeouts);
io_uring/timeout.c-641- timeout->target_seq = tail + off;
--
io_uring/tw.c=170=void io_req_local_work_add(struct io_kiocb *req, unsigned flags)
--
io_uring/tw.c-224- io_ctx_mark_taskrun(ctx);
io_uring/tw.c:225: if (data_race(ctx->int_flags) & IO_RING_F_HAS_EVFD)
io_uring/tw.c-226- io_eventfd_signal(ctx, false);
--
kernel/fork.c=1967=__latent_entropy struct task_struct *copy_process(
--
kernel/fork.c-2140- retval = -EAGAIN;
kernel/fork.c:2141: if (data_race(nr_threads >= max_threads))
kernel/fork.c-2142- goto bad_fork_cleanup_count;
--
kernel/irq/irqdesc.c=1002=static unsigned int kstat_irqs_desc(struct irq_desc *desc, const struct cpumask *cpumask)
--
kernel/irq/irqdesc.c-1009- !irq_is_nmi(desc))
kernel/irq/irqdesc.c:1010: return data_race(desc->tot_count);
kernel/irq/irqdesc.c-1011-
kernel/irq/irqdesc.c-1012- for_each_cpu(cpu, cpumask)
kernel/irq/irqdesc.c:1013: sum += data_race(per_cpu(desc->kstat_irqs->cnt, cpu));
kernel/irq/irqdesc.c-1014- return sum;
--
kernel/kcsan/core.c=527=kcsan_setup_watchpoint(const volatile void *ptr, size_t size, int type, unsigned long ip)
--
kernel/kcsan/core.c-643- */
kernel/kcsan/core.c:644: if (diff && !kcsan_ignore_data_race(size, type, old, new, diff))
kernel/kcsan/core.c-645- value_change = KCSAN_VALUE_CHANGE_TRUE;
--
kernel/kcsan/core.c-667- /*
kernel/kcsan/core.c:668: * No need to increment 'data_races' counter, as the racing
kernel/kcsan/core.c-669- * thread already did.
--
kernel/kcsan/debugfs.c=26=static const char *const counter_names[] = {
--
kernel/kcsan/debugfs.c-28- [KCSAN_COUNTER_SETUP_WATCHPOINTS] = "setup_watchpoints",
kernel/kcsan/debugfs.c:29: [KCSAN_COUNTER_DATA_RACES] = "data_races",
kernel/kcsan/debugfs.c-30- [KCSAN_COUNTER_ASSERT_FAILURES] = "assert_failures",
--
kernel/kcsan/debugfs.c=138=static ssize_t insert_report_filterlist(const char *func)
--
kernel/kcsan/debugfs.c-156- */
kernel/kcsan/debugfs.c:157: if (data_race(report_filterlist.used == report_filterlist.size)) {
kernel/kcsan/debugfs.c-158- new_size = (report_filterlist.size ?: 4) * 2;
--
kernel/kcsan/kcsan_test.c=358=static noinline void test_kernel_write_uninstrumented(void) { test_var++; }
kernel/kcsan/kcsan_test.c-359-
kernel/kcsan/kcsan_test.c:360:static noinline void test_kernel_data_race(void) { data_race(test_var++); }
kernel/kcsan/kcsan_test.c-361-
--
kernel/kcsan/kcsan_test.c=976=static void test_zero_size_access(struct kunit *test)
--
kernel/kcsan/kcsan_test.c-1001-
kernel/kcsan/kcsan_test.c:1002:/* Test the data_race() macro. */
kernel/kcsan/kcsan_test.c-1003-__no_kcsan
kernel/kcsan/kcsan_test.c:1004:static void test_data_race(struct kunit *test)
kernel/kcsan/kcsan_test.c-1005-{
--
kernel/kcsan/kcsan_test.c-1007-
kernel/kcsan/kcsan_test.c:1008: begin_test_checks(test_kernel_data_race, test_kernel_data_race);
kernel/kcsan/kcsan_test.c-1009- do {
--
kernel/kcsan/kcsan_test.c=1428=static struct kunit_case kcsan_test_cases[] = {
--
kernel/kcsan/kcsan_test.c-1441- KCSAN_KUNIT_CASE(test_zero_size_access),
kernel/kcsan/kcsan_test.c:1442: KCSAN_KUNIT_CASE(test_data_race),
kernel/kcsan/kcsan_test.c-1443- KCSAN_KUNIT_CASE(test_data_racy_qualifier),
--
kernel/kcsan/permissive.h=37=static bool
kernel/kcsan/permissive.h:38:kcsan_ignore_data_race(size_t size, int type, u64 old, u64 new, u64 diff)
kernel/kcsan/permissive.h-39-{
--
kernel/locking/locktorture.c=1000=static void __torture_print_stats(char *page,
--
kernel/locking/locktorture.c-1005- int i, n_stress;
kernel/locking/locktorture.c:1006: long max = 0, min = statp ? data_race(statp[0].n_lock_acquired) : 0;
kernel/locking/locktorture.c-1007- long long sum = 0;
--
kernel/locking/locktorture.c-1010- for (i = 0; i < n_stress; i++) {
kernel/locking/locktorture.c:1011: if (data_race(statp[i].n_lock_fail))
kernel/locking/locktorture.c-1012- fail = true;
kernel/locking/locktorture.c:1013: cur = data_race(statp[i].n_lock_acquired);
kernel/locking/locktorture.c-1014- sum += cur;
--
kernel/locking/mutex.c=323=bool ww_mutex_spin_on_owner(struct mutex *lock, struct ww_acquire_ctx *ww_ctx,
--
kernel/locking/mutex.c-357- */
kernel/locking/mutex.c:358: if (waiter && data_race(lock->first_waiter != waiter))
kernel/locking/mutex.c-359- return false;
--
kernel/locking/osq_lock.c=93=bool osq_lock(struct optimistic_spin_queue *lock)
--
kernel/locking/osq_lock.c-163- */
kernel/locking/osq_lock.c:164: if (data_race(prev->next) == node &&
kernel/locking/osq_lock.c-165- cmpxchg(&prev->next, node, NULL) == node)
--
kernel/locking/rtmutex.c=1490=static bool rtmutex_spin_on_owner(struct rt_mutex_base *lock,
--
kernel/locking/rtmutex.c-1516- if (!owner_on_cpu(owner) || need_resched() ||
kernel/locking/rtmutex.c:1517: !data_race(rt_mutex_waiter_is_top_waiter(lock, waiter))) {
kernel/locking/rtmutex.c-1518- res = false;
--
kernel/locking/ww_mutex.h=403=ww_mutex_set_context_fastpath(struct ww_mutex *lock, struct ww_acquire_ctx *ctx)
--
kernel/locking/ww_mutex.h-428- */
kernel/locking/ww_mutex.h:429: has_waiters = data_race(__ww_mutex_has_waiters(&lock->base));
kernel/locking/ww_mutex.h-430- if (likely(!has_waiters))
--
kernel/printk/nbcon.c=980=static bool nbcon_emit_next_record(struct nbcon_write_context *wctxt, bool use_atomic)
--
kernel/printk/nbcon.c-1025- */
kernel/printk/nbcon.c:1026: con_dropped = data_race(READ_ONCE(con->dropped));
kernel/printk/nbcon.c-1027-
--
kernel/rcu/rcutorture.c=2781=rcu_torture_stats_print(void)
--
kernel/rcu/rcutorture.c-2840- pr_cont("barrier: %ld/%ld:%ld ",
kernel/rcu/rcutorture.c:2841: data_race(n_barrier_successes),
kernel/rcu/rcutorture.c:2842: data_race(n_barrier_attempts),
kernel/rcu/rcutorture.c:2843: data_race(n_rcu_torture_barrier_error));
kernel/rcu/rcutorture.c:2844: pr_cont("read-exits: %ld ", data_race(n_read_exits)); // Statistic.
kernel/rcu/rcutorture.c-2845- pr_cont("nocb-toggles: %ld:%ld ",
--
kernel/rcu/srcutree.c=2032=void srcu_torture_stats_print(struct srcu_struct *ssp, char *tt, char *tf)
--
kernel/rcu/srcutree.c-2057- sdp = per_cpu_ptr(ssp->sda, cpu);
kernel/rcu/srcutree.c:2058: u0 = data_race(atomic_long_read(&sdp->srcu_ctrs[!idx].srcu_unlocks));
kernel/rcu/srcutree.c:2059: u1 = data_race(atomic_long_read(&sdp->srcu_ctrs[idx].srcu_unlocks));
kernel/rcu/srcutree.c-2060-
--
kernel/rcu/srcutree.c-2066-
kernel/rcu/srcutree.c:2067: l0 = data_race(atomic_long_read(&sdp->srcu_ctrs[!idx].srcu_locks));
kernel/rcu/srcutree.c:2068: l1 = data_race(atomic_long_read(&sdp->srcu_ctrs[idx].srcu_locks));
kernel/rcu/srcutree.c-2069-
--
kernel/rcu/tasks.h=233=static const char *tasks_gp_state_getname(struct rcu_tasks *rtp)
kernel/rcu/tasks.h-234-{
kernel/rcu/tasks.h:235: int i = data_race(rtp->gp_state); // Let KCSAN detect update races
kernel/rcu/tasks.h-236- int j = READ_ONCE(i); // Prevent the compiler from reading twice
--
kernel/rcu/tasks.h=246=static void cblist_init_generic(struct rcu_tasks *rtp)
--
kernel/rcu/tasks.h-298- pr_info("%s: Setting shift to %d and lim to %d rcu_task_cb_adjust=%d rcu_task_cpu_ids=%d.\n",
kernel/rcu/tasks.h:299: rtp->name, data_race(rtp->percpu_enqueue_shift), data_race(rtp->percpu_enqueue_lim),
kernel/rcu/tasks.h-300- rcu_task_cb_adjust, rcu_task_cpu_ids);
--
kernel/rcu/tasks.h=718=static void show_rcu_tasks_generic_gp_kthread(struct rcu_tasks *rtp, char *s)
--
kernel/rcu/tasks.h-727-
]
]
|
| 41/2 |
2026/04/21 11:27 |
llm |
expert |
2m
Model:
Tokens:
input: 75916
output: 46
thoughts: 1964
Thoughts: **KCSAN False Positive in `ktime_get_coarse_real_ts64_mg`?**
Okay, so I'm digging into a KCSAN report that seems to point to a data race in `kernel/time/timekeeping.c`, but the immediate suspect, `data_race`, isn't directly present there. The traces lead me to `fs/inode.c`, specifically `mg_ctime` related calls, but the real target is `ktime_get_coarse_real_ts64_mg`, which is definitely in `kernel/time/timekeeping.c`.
My initial thought is to slap `data_race()` around the `tk_xtime_coarse` accesses, which *should* silence KCSAN. And it seems safe given the surrounding `seqcount` loop in both `ktime_get_coarse_real_ts64_mg` and `ktime_get_coarse_real_ts64`. However, `tk_xtime_coarse` is also used in `ktime_get_coarse_real_ts64`, so any changes would affect both.
Looking closer at `tk_xtime_coarse`, it's a simple structure access to `tk->xtime_sec` and `tk->coarse_nsec`. Applying `data_race()` there feels like the correct local fix, so to speak.
The puzzle deepens because KCSAN is *only* complaining about the `_mg` variant. I suspect it's triggered by `current_time` called from `file_update_time_flags`, which is in turn called by `shmem_file_write_iter`. It's likely a very hot path under syzkaller.
Aha! There's a real value change reported: `0x31af5fed -> 0x3247f66d`. This isn't a theoretical issue; the values are actually torn. Despite the `seqcount` retry, KCSAN is flagging this as a problem, presumably because it doesn't intrinsically trust the retry mechanism as a benign race. Ideally, KCSAN should recognize `read_seqcount_begin` and similar primitives.
Examining the code, specifically the `ktime_get_coarse_real_ts64_mg` function, I notice `tk` is initialized *outside* the loop: `struct timekeeper *tk = &tk_core.timekeeper;`. `tk_core` is a macro, which expands to `timekeeper_data[TIMEKEEPER_CORE]`. So `tk` is a pointer to the core `timekeeper` data structure. I also see that `offs_real` is accessed inside the loop.
Could there be compiler optimizations playing a part here because of the macro? Expanding `tk_core` seems to yield `timekeeper_data[TIMEKEEPER_CORE]`. And indeed, the trace points to line 2606 where `tk_xtime_coarse` is called within the loop.
I also observe that there's also an `atomic64_read` from `mg_floor` called outside the `do ... while` loop.
Digging further, I trace the access in the read stack trace in line 2606.
The critical issue appears to be inside the loop in `ktime_get_coarse_real_ts64_mg`. I see accesses to `tk_core.timekeeper.offs_real`. So this translates to `timekeeper_data[TIMEKEEPER_CORE].timekeeper.offs_real`. And `tk_core.seq` translates to `timekeeper_data[TIMEKEEPER_CORE].seq`.
If `TIMEKEEPER_CORE` is 0, this simplifies it to `timekeeper_data[0]`. Tracing back, I see `timekeeping_update_from_shadow` takes a `struct tk_data *tkd`, and in `__timekeeping_advance`, this is called with `&tk_core`. So, `tkd` is `&timekeeper_data[0]`. Inside `timekeeping_update_from_shadow`, there is a `memcpy` from `tk` to `tkd->timekeeper`, which is writing to `timekeeper_data[0].timekeeper`. The reader is reading from `timekeeper_data[0].timekeeper`.
The `seqcount` mechanism *should* protect everything.
So why is KCSAN complaining? I suspect `memcpy` is the culprit. KCSAN is often conservative with `memcpy`, treating it as non-atomic even inside `kcsan_nestable_atomic_begin()` because it's a
|
| 42/2 |
2026/04/21 11:30 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/kcsan/core.c IncludeLines:true Name:kcsan_setup_watchpoint]
Results: map[SourceCode: 526: static noinline void
527: kcsan_setup_watchpoint(const volatile void *ptr, size_t size, int type, unsigned long ip)
528: {
529: const bool is_write = (type & KCSAN_ACCESS_WRITE) != 0;
530: const bool is_assert = (type & KCSAN_ACCESS_ASSERT) != 0;
531: atomic_long_t *watchpoint;
532: u64 old, new, diff;
533: enum kcsan_value_change value_change = KCSAN_VALUE_CHANGE_MAYBE;
534: bool interrupt_watcher = kcsan_interrupt_watcher;
535: unsigned long ua_flags = user_access_save();
536: struct kcsan_ctx *ctx = get_ctx();
537: unsigned long access_mask = ctx->access_mask;
538: unsigned long irq_flags = 0;
539: bool is_reorder_access;
540:
541: /*
542: * Always reset kcsan_skip counter in slow-path to avoid underflow; see
543: * should_watch().
544: */
545: reset_kcsan_skip();
546:
547: if (!kcsan_is_enabled(ctx))
548: goto out;
549:
550: /*
551: * Check to-ignore addresses after kcsan_is_enabled(), as we may access
552: * memory that is not yet initialized during early boot.
553: */
554: if (!is_assert && kcsan_ignore_address(ptr))
555: goto out;
556:
557: if (!check_encodable((unsigned long)ptr, size)) {
558: atomic_long_inc(&kcsan_counters[KCSAN_COUNTER_UNENCODABLE_ACCESSES]);
559: goto out;
560: }
561:
562: /*
563: * The local CPU cannot observe reordering of its own accesses, and
564: * therefore we need to take care of 2 cases to avoid false positives:
565: *
566: * 1. Races of the reordered access with interrupts. To avoid, if
567: * the current access is reorder_access, disable interrupts.
568: * 2. Avoid races of scoped accesses from nested interrupts (below).
569: */
570: is_reorder_access = find_reorder_access(ctx, ptr, size, type, ip);
571: if (is_reorder_access)
572: interrupt_watcher = false;
573: /*
574: * Avoid races of scoped accesses from nested interrupts (or scheduler).
575: * Assume setting up a watchpoint for a non-scoped (normal) access that
576: * also conflicts with a current scoped access. In a nested interrupt,
577: * which shares the context, it would check a conflicting scoped access.
578: * To avoid, disable scoped access checking.
579: */
580: ctx->disable_scoped++;
581:
582: /*
583: * Save and restore the IRQ state trace touched by KCSAN, since KCSAN's
584: * runtime is entered for every memory access, and potentially useful
585: * information is lost if dirtied by KCSAN.
586: */
587: kcsan_save_irqtrace(current);
588: if (!interrupt_watcher)
589: local_irq_save(irq_flags);
590:
591: watchpoint = insert_watchpoint((unsigned long)ptr, size, is_write);
592: if (watchpoint == NULL) {
593: /*
594: * Out of capacity: the size of 'watchpoints', and the frequency
595: * with which should_watch() returns true should be tweaked so
596: * that this case happens very rarely.
597: */
598: atomic_long_inc(&kcsan_counters[KCSAN_COUNTER_NO_CAPACITY]);
599: goto out_unlock;
600: }
601:
602: atomic_long_inc(&kcsan_counters[KCSAN_COUNTER_SETUP_WATCHPOINTS]);
603: atomic_long_inc(&kcsan_counters[KCSAN_COUNTER_USED_WATCHPOINTS]);
604:
605: /*
606: * Read the current value, to later check and infer a race if the data
607: * was modified via a non-instrumented access, e.g. from a device.
608: */
609: old = is_reorder_access ? 0 : read_instrumented_memory(ptr, size);
610:
611: /*
612: * Delay this thread, to increase probability of observing a racy
613: * conflicting access.
614: */
615: delay_access(type);
616:
617: /*
618: * Re-read value, and check if it is as expected; if not, we infer a
619: * racy access.
620: */
621: if (!is_reorder_access) {
622: new = read_instrumented_memory(ptr, size);
623: } else {
624: /*
625: * Reordered accesses cannot be used for value change detection,
626: * because the memory location may no longer be accessible and
627: * could result in a fault.
628: */
629: new = 0;
630: access_mask = 0;
631: }
632:
633: diff = old ^ new;
634: if (access_mask)
635: diff &= access_mask;
636:
637: /*
638: * Check if we observed a value change.
639: *
640: * Also check if the data race should be ignored (the rules depend on
641: * non-zero diff); if it is to be ignored, the below rules for
642: * KCSAN_VALUE_CHANGE_MAYBE apply.
643: */
644: if (diff && !kcsan_ignore_data_race(size, type, old, new, diff))
645: value_change = KCSAN_VALUE_CHANGE_TRUE;
646:
647: /* Check if this access raced with another. */
648: if (!consume_watchpoint(watchpoint)) {
649: /*
650: * Depending on the access type, map a value_change of MAYBE to
651: * TRUE (always report) or FALSE (never report).
652: */
653: if (value_change == KCSAN_VALUE_CHANGE_MAYBE) {
654: if (access_mask != 0) {
655: /*
656: * For access with access_mask, we require a
657: * value-change, as it is likely that races on
658: * ~access_mask bits are expected.
659: */
660: value_change = KCSAN_VALUE_CHANGE_FALSE;
661: } else if (size > 8 || is_assert) {
662: /* Always assume a value-change. */
663: value_change = KCSAN_VALUE_CHANGE_TRUE;
664: }
665: }
666:
667: /*
668: * No need to increment 'data_races' counter, as the racing
669: * thread already did.
670: *
671: * Count 'assert_failures' for each failed ASSERT access,
672: * therefore both this thread and the racing thread may
673: * increment this counter.
674: */
675: if (is_assert && value_change == KCSAN_VALUE_CHANGE_TRUE)
676: atomic_long_inc(&kcsan_counters[KCSAN_COUNTER_ASSERT_FAILURES]);
677:
678: kcsan_report_known_origin(ptr, size, type, ip,
679: value_change, watchpoint - watchpoints,
680: old, new, access_mask);
681: } else if (value_change == KCSAN_VALUE_CHANGE_TRUE) {
682: /* Inferring a race, since the value should not have changed. */
683:
684: atomic_long_inc(&kcsan_counters[KCSAN_COUNTER_RACES_UNKNOWN_ORIGIN]);
685: if (is_assert)
686: atomic_long_inc(&kcsan_counters[KCSAN_COUNTER_ASSERT_FAILURES]);
687:
688: if (IS_ENABLED(CONFIG_KCSAN_REPORT_RACE_UNKNOWN_ORIGIN) || is_assert) {
689: kcsan_report_unknown_origin(ptr, size, type, ip,
690: old, new, access_mask);
691: }
692: }
693:
694: /*
695: * Remove watchpoint; must be after reporting, since the slot may be
696: * reused after this point.
697: */
698: remove_watchpoint(watchpoint);
699: atomic_long_dec(&kcsan_counters[KCSAN_COUNTER_USED_WATCHPOINTS]);
700:
701: out_unlock:
702: if (!interrupt_watcher)
703: local_irq_restore(irq_flags);
704: kcsan_restore_irqtrace(current);
705: ctx->disable_scoped--;
706:
707: /*
708: * Reordered accesses cannot be used for value change detection,
709: * therefore never consider for reordering if access_mask is set.
710: * ASSERT_EXCLUSIVE are not real accesses, ignore them as well.
711: */
712: if (!access_mask && !is_assert)
713: set_reorder_access(ctx, ptr, size, type, ip);
714: out:
715: user_access_restore(ua_flags);
716: }
SourceFile:kernel/kcsan/core.c]
|
| 43/2 |
2026/04/21 11:30 |
llm |
expert |
0m
Model:
Tokens:
input: 81979
output: 48
thoughts: 29
Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience:
**Analyzing Potential Data Race Ignore**
Alright, I'm scanning this code, and my eye immediately catches something at line 644: `kcsan_ignore_data_race`. *That* jumps out at me. A `kcsan_ignore_data_race` call... that's a red flag, but a strategically placed one, potentially. It demands immediate investigation. I need to understand why this specific data race is being ignored. Is it a false positive that needs adjustment or a necessary optimization? I need to dig deeper into the context around that call – what's being protected, what's being accessed, what assumptions are being made, and what implications does this ignore have on the system's overall safety and correctness? This warrants a deeper dive.
|
| 44/2 |
2026/04/21 11:30 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/kcsan/permissive.h IncludeLines:true Name:kcsan_ignore_data_race]
Results: map[SourceCode: 34: /*
35: * Data race ignore rules based on access type and value change patterns.
36: */
37: static bool
38: kcsan_ignore_data_race(size_t size, int type, u64 old, u64 new, u64 diff)
39: {
40: if (!IS_ENABLED(CONFIG_KCSAN_PERMISSIVE))
41: return false;
42:
43: /*
44: * Rules here are only for plain read accesses, so that we still report
45: * data races between plain read-write accesses.
46: */
47: if (type || size > sizeof(long))
48: return false;
49:
50: /*
51: * A common pattern is checking/setting just 1 bit in a variable; for
52: * example:
53: *
54: * if (flags & SOME_FLAG) { ... }
55: *
56: * and elsewhere flags is updated concurrently:
57: *
58: * flags |= SOME_OTHER_FLAG; // just 1 bit
59: *
60: * While it is still recommended that such accesses be marked
61: * appropriately, in many cases these types of data races are so common
62: * that marking them all is often unrealistic and left to maintainer
63: * preference.
64: *
65: * The assumption in all cases is that with all known compiler
66: * optimizations (including those that tear accesses), because no more
67: * than 1 bit changed, the plain accesses are safe despite the presence
68: * of data races.
69: *
70: * The rules here will ignore the data races if we observe no more than
71: * 1 bit changed.
72: *
73: * Of course many operations can effecively change just 1 bit, but the
74: * general assuption that data races involving 1-bit changes can be
75: * tolerated still applies.
76: *
77: * And in case a true bug is missed, the bug likely manifests as a
78: * reportable data race elsewhere.
79: */
80: if (hweight64(diff) == 1) {
81: /*
82: * Exception: Report data races where the values look like
83: * ordinary booleans (one of them was 0 and the 0th bit was
84: * changed) More often than not, they come with interesting
85: * memory ordering requirements, so let's report them.
86: */
87: if (!((!old || !new) && diff == 1))
88: return true;
89: }
90:
91: return false;
92: }
SourceFile:kernel/kcsan/permissive.h]
|
| 45/2 |
2026/04/21 11:30 |
llm |
expert |
0m
Model:
Tokens:
input: 83060
output: 41
thoughts: 91
Thoughts: Okay, here's my take on those thoughts, framed from my own perspective:
**Debugging a KCSAN Watchpoint Issue**
Alright, this is interesting. We're in permissive mode, which gives us the opportunity to see what's going on. Ah, right, I'm seeing that `kcsan_setup_watchpoint` gets called for the read access. The specific value being read is `tk->coarse_nsec`. And yes, it's just a regular, plain read – nothing obviously atomic about the access itself.
Now, my initial thought is, based on my understanding of how KCSAN and seqcounts should interact, if `read_seqcount_begin` had been in play – if the "atomic" flag was correctly set – `kcsan_setup_watchpoint` should have been skipped. That's the expected behavior, to avoid false positives on what should be safe, atomic reads. So, the key here is to confirm `should_watch`, this will show us whether or not that flag was appropriately set earlier.
|
| 46/2 |
2026/04/21 11:30 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/kcsan/core.c IncludeLines:true Name:should_watch]
Results: map[SourceCode: 267: static __always_inline bool
268: should_watch(struct kcsan_ctx *ctx, const volatile void *ptr, size_t size, int type)
269: {
270: /*
271: * Never set up watchpoints when memory operations are atomic.
272: *
273: * Need to check this first, before kcsan_skip check below: (1) atomics
274: * should not count towards skipped instructions, and (2) to actually
275: * decrement kcsan_atomic_next for consecutive instruction stream.
276: */
277: if (is_atomic(ctx, ptr, size, type))
278: return false;
279:
280: if (this_cpu_dec_return(kcsan_skip) >= 0)
281: return false;
282:
283: /*
284: * NOTE: If we get here, kcsan_skip must always be reset in slow path
285: * via reset_kcsan_skip() to avoid underflow.
286: */
287:
288: /* this operation should be watched */
289: return true;
290: }
SourceFile:kernel/kcsan/core.c]
|
| 47/2 |
2026/04/21 11:30 |
llm |
expert |
0m
Model:
Tokens:
input: 83804
output: 41
thoughts: 6
Thoughts: Okay, here's my attempt at summarizing the thought process, tailored for an expert audience, and written in the first person:
**Analyzing the `is_atomic` Flag**
Right, let me see. We're talking about an `is_atomic` flag, which suggests some kind of concurrency control or memory management aspect. My immediate thought is, "What kind of data structure or operation are we discussing where atomicity is a critical concern?" Without more context, I'm already mentally running through the usual suspects: shared counters, lock-free data structures, inter-thread communication primitives, maybe even transactional memory.
The presence of this flag strongly implies that whatever's being described can't be safely modified, or accessed concurrently, without some form of synchronization. It's not just about performance; it's about *correctness*. Incorrect handling of an atomic operation can lead to data corruption, race conditions, and all sorts of nasty, unpredictable behavior. I'm thinking about potential consistency models: sequential consistency, relaxed consistency, and how the `is_atomic` flag relates to these. Is this something that's platform-specific? Does this flag map to specific CPU instructions or memory barriers? What's the compiler's role in interpreting this flag?
Essentially, my brain is now asking *why* this flag exists. What are the specific guarantees this flag provides? What are the implications if it's set, and what are the implications if it's *not* set? Are there performance trade-offs? Are there other flags or attributes that this one interacts with? I'm already diving into the nuances of memory models and the possible optimizations and potential pitfalls that arise from the use of atomic operations. I'm practically itching to see the surrounding code!
|
| 48/2 |
2026/04/21 11:30 |
tool |
codesearch-definition-source |
0m
Args: map[ContextFile:kernel/kcsan/core.c IncludeLines:true Name:is_atomic]
Results: map[SourceCode: 229: /* Rules for generic atomic accesses. Called from fast-path. */
230: static __always_inline bool
231: is_atomic(struct kcsan_ctx *ctx, const volatile void *ptr, size_t size, int type)
232: {
233: if (type & KCSAN_ACCESS_ATOMIC)
234: return true;
235:
236: /*
237: * Unless explicitly declared atomic, never consider an assertion access
238: * as atomic. This allows using them also in atomic regions, such as
239: * seqlocks, without implicitly changing their semantics.
240: */
241: if (type & KCSAN_ACCESS_ASSERT)
242: return false;
243:
244: if (IS_ENABLED(CONFIG_KCSAN_ASSUME_PLAIN_WRITES_ATOMIC) &&
245: (type & KCSAN_ACCESS_WRITE) && size <= sizeof(long) &&
246: !(type & KCSAN_ACCESS_COMPOUND) && IS_ALIGNED((unsigned long)ptr, size))
247: return true; /* Assume aligned writes up to word size are atomic. */
248:
249: if (ctx->atomic_next > 0) {
250: /*
251: * Because we do not have separate contexts for nested
252: * interrupts, in case atomic_next is set, we simply assume that
253: * the outer interrupt set atomic_next. In the worst case, we
254: * will conservatively consider operations as atomic. This is a
255: * reasonable trade-off to make, since this case should be
256: * extremely rare; however, even if extremely rare, it could
257: * lead to false positives otherwise.
258: */
259: if ((hardirq_count() >> HARDIRQ_SHIFT) < 2)
260: --ctx->atomic_next; /* in task, or outer interrupt */
261: return true;
262: }
263:
264: return ctx->atomic_nest_count > 0 || ctx->in_flat_atomic;
265: }
SourceFile:kernel/kcsan/core.c]
|
| 49/2 |
2026/04/21 11:30 |
llm |
expert |
1m
Model:
Tokens:
input: 84871
output: 23
thoughts: 2637
Thoughts: **KCSAN False Positive in Timekeeping: A Deep Dive**
Okay, let's break this down. First, the initial confusion: I was trying to understand why KCSAN was reporting a race condition when, logically, it shouldn't. The `read_seqcount_begin` should've made the read atomic and prevented the watchpoint from being set, right? But the report indicated a race between a read and a write (a 304-byte `memcpy` in an interrupt handler). I started retracing my steps. I realized `kcsan_nestable_atomic_begin()` was called in the interrupt, so the write should have been considered atomic too.
My initial assumption was that the KCSAN report was incorrectly flagging accesses *within* atomic regions. But the crucial piece was realizing the write stack trace: it wasn't the interrupt *during* the read, but a race that occurred because the write had set a watchpoint that the read violated. The trace was showing the state of the CPU when the race was *detected*, not the sequence of events.
The key was the instrumentation of `memcpy`. It was being treated as a single, large access. I looked at `kcsan_setup_watchpoint` and the `size > 8` condition, and noticed that KCSAN was setting a watchpoint on the `memcpy` because of its size. But I was missing something: the atomic context was preventing the watchpoint from being set. I saw the `__tsan_read4` call which seemed unrelated; it was the access in `pid_vnr`.
Now I thought maybe `ktime_get_coarse_real_ts64_mg` was missing the `READ_ONCE` for the data it reads. However, this didn't explain the report. Even if `read_seqcount_begin` wasn't used, I still would have seen the watchpoint. I then remembered `seqcount` is a special case.
The key then became the read and write: `tk->coarse_nsec` and `memcpy(&tkd->timekeeper, tk, sizeof(*tk))`. This looked like a *classic* `seqcount`-protected access. `seqcount` *should* be handling the potential for torn reads or writes. My concern was around whether the `coarse_nsec` could still be torn.
Analyzing the race report further, I saw that it was a 4-byte value that changed, which is an atomic operation. The write, which was a `memcpy`, might not be atomic. However, `seqcount` is designed to handle this. So, then, the race seemed benign. But *why* was KCSAN reporting it? Was there some missing annotation? Was the fact that `timekeeping_update_from_shadow` used `memcpy` the problem?
I checked the reads inside `ktime_get_coarse_real_ts64_mg` again: `xtime_sec`, `coarse_nsec` and `offs_real`. All of these are 8-byte or 4-byte reads, and aligned. They are atomic reads on x86-64. The struct itself is cacheline aligned.
The `memcpy` might not be atomic, but `seqcount` should handle the potential for a torn write and retry.
I'm starting to think this is a **false positive**. The report is likely a side-effect of `ktime_get_coarse_real_ts64_mg` (or potentially `_mg` functions) being added. It may be due to missing `data_race()` annotations.
|
| 50/2 |
2026/04/21 11:32 |
tool |
read-file |
0m
Args: map[File:kernel/time/timekeeping.c]
Results: map[Contents:// SPDX-License-Identifier: GPL-2.0
/*
* Kernel timekeeping code and accessor functions. Based on code from
* timer.c, moved in commit 8524070b7982.
*/
#include <linux/audit.h>
#include <linux/clocksource.h>
#include <linux/compiler.h>
#include <linux/jiffies.h>
#include <linux/kobject.h>
#include <linux/module.h>
#include <linux/nmi.h>
#include <linux/pvclock_gtod.h>
#include <linux/random.h>
#include <linux/sched/clock.h>
#include <linux/sched/loadavg.h>
#include <linux/static_key.h>
#include <linux/stop_machine.h>
#include <linux/syscore_ops.h>
#include <linux/tick.h>
#include <linux/time.h>
#include <linux/timex.h>
#include <linux/timekeeper_internal.h>
#include <vdso/auxclock.h>
#include "tick-internal.h"
#include "timekeeping_internal.h"
#include "ntp_internal.h"
#define TK_CLEAR_NTP (1 << 0)
#define TK_CLOCK_WAS_SET (1 << 1)
#define TK_UPDATE_ALL (TK_CLEAR_NTP | TK_CLOCK_WAS_SET)
enum timekeeping_adv_mode {
/* Update timekeeper when a tick has passed */
TK_ADV_TICK,
/* Update timekeeper on a direct frequency change */
TK_ADV_FREQ
};
/*
* The most important data for readout fits into a single 64 byte
* cache line.
*/
struct tk_data {
seqcount_raw_spinlock_t seq;
struct timekeeper timekeeper;
struct timekeeper shadow_timekeeper;
raw_spinlock_t lock;
} ____cacheline_aligned;
static struct tk_data timekeeper_data[TIMEKEEPERS_MAX];
/* The core timekeeper */
#define tk_core (timekeeper_data[TIMEKEEPER_CORE])
#ifdef CONFIG_POSIX_AUX_CLOCKS
static inline bool tk_get_aux_ts64(unsigned int tkid, struct timespec64 *ts)
{
return ktime_get_aux_ts64(CLOCK_AUX + tkid - TIMEKEEPER_AUX_FIRST, ts);
}
static inline bool tk_is_aux(const struct timekeeper *tk)
{
return tk->id >= TIMEKEEPER_AUX_FIRST && tk->id <= TIMEKEEPER_AUX_LAST;
}
#else
static inline bool tk_get_aux_ts64(unsigned int tkid, struct timespec64 *ts)
{
return false;
}
static inline bool tk_is_aux(const struct timekeeper *tk)
{
return false;
}
#endif
static inline void tk_update_aux_offs(struct timekeeper *tk, ktime_t offs)
{
tk->offs_aux = offs;
tk->monotonic_to_aux = ktime_to_timespec64(offs);
}
/* flag for if timekeeping is suspended */
int __read_mostly timekeeping_suspended;
/**
* struct tk_fast - NMI safe timekeeper
* @seq: Sequence counter for protecting updates. The lowest bit
* is the index for the tk_read_base array
* @base: tk_read_base array. Access is indexed by the lowest bit of
* @seq.
*
* See @update_fast_timekeeper() below.
*/
struct tk_fast {
seqcount_latch_t seq;
struct tk_read_base base[2];
};
/* Suspend-time cycles value for halted fast timekeeper. */
static u64 cycles_at_suspend;
static u64 dummy_clock_read(struct clocksource *cs)
{
if (timekeeping_suspended)
return cycles_at_suspend;
return local_clock();
}
static struct clocksource dummy_clock = {
.read = dummy_clock_read,
};
/*
* Boot time initialization which allows local_clock() to be utilized
* during early boot when clocksources are not available. local_clock()
* returns nanoseconds already so no conversion is required, hence mult=1
* and shift=0. When the first proper clocksource is installed then
* the fast time keepers are updated with the correct values.
*/
#define FAST_TK_INIT \
{ \
.clock = &dummy_clock, \
.mask = CLOCKSOURCE_MASK(64), \
.mult = 1, \
.shift = 0, \
}
static struct tk_fast tk_fast_mono ____cacheline_aligned = {
.seq = SEQCNT_LATCH_ZERO(tk_fast_mono.seq),
.base[0] = FAST_TK_INIT,
.base[1] = FAST_TK_INIT,
};
static struct tk_fast tk_fast_raw ____cacheline_aligned = {
.seq = SEQCNT_LATCH_ZERO(tk_fast_raw.seq),
.base[0] = FAST_TK_INIT,
.base[1] = FAST_TK_INIT,
};
#ifdef CONFIG_POSIX_AUX_CLOCKS
static __init void tk_aux_setup(void);
static void tk_aux_update_clocksource(void);
static void tk_aux_advance(void);
#else
static inline void tk_aux_setup(void) { }
static inline void tk_aux_update_clocksource(void) { }
static inline void tk_aux_advance(void) { }
#endif
unsigned long timekeeper_lock_irqsave(void)
{
unsigned long flags;
raw_spin_lock_irqsave(&tk_core.lock, flags);
return flags;
}
void timekeeper_unlock_irqrestore(unsigned long flags)
{
raw_spin_unlock_irqrestore(&tk_core.lock, flags);
}
/*
* Multigrain timestamps require tracking the latest fine-grained timestamp
* that has been issued, and never returning a coarse-grained timestamp that is
* earlier than that value.
*
* mg_floor represents the latest fine-grained time that has been handed out as
* a file timestamp on the system. This is tracked as a monotonic ktime_t, and
* converted to a realtime clock value on an as-needed basis.
*
* Maintaining mg_floor ensures the multigrain interfaces never issue a
* timestamp earlier than one that has been previously issued.
*
* The exception to this rule is when there is a backward realtime clock jump. If
* such an event occurs, a timestamp can appear to be earlier than a previous one.
*/
static __cacheline_aligned_in_smp atomic64_t mg_floor;
static inline void tk_normalize_xtime(struct timekeeper *tk)
{
while (tk->tkr_mono.xtime_nsec >= ((u64)NSEC_PER_SEC << tk->tkr_mono.shift)) {
tk->tkr_mono.xtime_nsec -= (u64)NSEC_PER_SEC << tk->tkr_mono.shift;
tk->xtime_sec++;
}
while (tk->tkr_raw.xtime_nsec >= ((u64)NSEC_PER_SEC << tk->tkr_raw.shift)) {
tk->tkr_raw.xtime_nsec -= (u64)NSEC_PER_SEC << tk->tkr_raw.shift;
tk->raw_sec++;
}
}
static inline struct timespec64 tk_xtime(const struct timekeeper *tk)
{
struct timespec64 ts;
ts.tv_sec = tk->xtime_sec;
ts.tv_nsec = (long)(tk->tkr_mono.xtime_nsec >> tk->tkr_mono.shift);
return ts;
}
static inline struct timespec64 tk_xtime_coarse(const struct timekeeper *tk)
{
struct timespec64 ts;
ts.tv_sec = tk->xtime_sec;
ts.tv_nsec = tk->coarse_nsec;
return ts;
}
/*
* Update the nanoseconds part for the coarse time keepers. They can't rely
* on xtime_nsec because xtime_nsec could be adjusted by a small negative
* amount when the multiplication factor of the clock is adjusted, which
* could cause the coarse clocks to go slightly backwards. See
* timekeeping_apply_adjustment(). Thus we keep a separate copy for the coarse
* clockids which only is updated when the clock has been set or we have
* accumulated time.
*/
static inline void tk_update_coarse_nsecs(struct timekeeper *tk)
{
tk->coarse_nsec = tk->tkr_mono.xtime_nsec >> tk->tkr_mono.shift;
}
static void tk_set_xtime(struct timekeeper *tk, const struct timespec64 *ts)
{
tk->xtime_sec = ts->tv_sec;
tk->tkr_mono.xtime_nsec = (u64)ts->tv_nsec << tk->tkr_mono.shift;
tk_update_coarse_nsecs(tk);
}
static void tk_xtime_add(struct timekeeper *tk, const struct timespec64 *ts)
{
tk->xtime_sec += ts->tv_sec;
tk->tkr_mono.xtime_nsec += (u64)ts->tv_nsec << tk->tkr_mono.shift;
tk_normalize_xtime(tk);
tk_update_coarse_nsecs(tk);
}
static void tk_set_wall_to_mono(struct timekeeper *tk, struct timespec64 wtm)
{
struct timespec64 tmp;
/*
* Verify consistency of: offset_real = -wall_to_monotonic
* before modifying anything
*/
set_normalized_timespec64(&tmp, -tk->wall_to_monotonic.tv_sec,
-tk->wall_to_monotonic.tv_nsec);
WARN_ON_ONCE(tk->offs_real != timespec64_to_ktime(tmp));
tk->wall_to_monotonic = wtm;
set_normalized_timespec64(&tmp, -wtm.tv_sec, -wtm.tv_nsec);
/* Paired with READ_ONCE() in ktime_mono_to_any() */
WRITE_ONCE(tk->offs_real, timespec64_to_ktime(tmp));
WRITE_ONCE(tk->offs_tai, ktime_add(tk->offs_real, ktime_set(tk->tai_offset, 0)));
}
static inline void tk_update_sleep_time(struct timekeeper *tk, ktime_t delta)
{
/* Paired with READ_ONCE() in ktime_mono_to_any() */
WRITE_ONCE(tk->offs_boot, ktime_add(tk->offs_boot, delta));
/*
* Timespec representation for VDSO update to avoid 64bit division
* on every update.
*/
tk->monotonic_to_boot = ktime_to_timespec64(tk->offs_boot);
}
#ifdef CONFIG_ARCH_WANTS_CLOCKSOURCE_READ_INLINE
#include <asm/clock_inlined.h>
static DEFINE_STATIC_KEY_FALSE(clocksource_read_inlined);
/*
* tk_clock_read - atomic clocksource read() helper
*
* This helper is necessary to use in the read paths because, while the
* seqcount ensures we don't return a bad value while structures are updated,
* it doesn't protect from potential crashes. There is the possibility that
* the tkr's clocksource may change between the read reference, and the
* clock reference passed to the read function. This can cause crashes if
* the wrong clocksource is passed to the wrong read function.
* This isn't necessary to use when holding the tk_core.lock or doing
* a read of the fast-timekeeper tkrs (which is protected by its own locking
* and update logic).
*/
static __always_inline u64 tk_clock_read(const struct tk_read_base *tkr)
{
struct clocksource *clock = READ_ONCE(tkr->clock);
if (static_branch_likely(&clocksource_read_inlined))
return arch_inlined_clocksource_read(clock);
return clock->read(clock);
}
static inline void clocksource_disable_inline_read(void)
{
static_branch_disable(&clocksource_read_inlined);
}
static inline void clocksource_enable_inline_read(void)
{
static_branch_enable(&clocksource_read_inlined);
}
#else
static __always_inline u64 tk_clock_read(const struct tk_read_base *tkr)
{
struct clocksource *clock = READ_ONCE(tkr->clock);
return clock->read(clock);
}
static inline void clocksource_disable_inline_read(void) { }
static inline void clocksource_enable_inline_read(void) { }
#endif
/**
* tk_setup_internals - Set up internals to use clocksource clock.
*
* @tk: The target timekeeper to setup.
* @clock: Pointer to clocksource.
*
* Calculates a fixed cycle/nsec interval for a given clocksource/adjustment
* pair and interval request.
*
* Unless you're the timekeeping code, you should not be using this!
*/
static void tk_setup_internals(struct timekeeper *tk, struct clocksource *clock)
{
u64 interval;
u64 tmp, ntpinterval;
struct clocksource *old_clock;
++tk->cs_was_changed_seq;
old_clock = tk->tkr_mono.clock;
tk->tkr_mono.clock = clock;
tk->tkr_mono.mask = clock->mask;
tk->tkr_mono.cycle_last = tk_clock_read(&tk->tkr_mono);
tk->tkr_raw.clock = clock;
tk->tkr_raw.mask = clock->mask;
tk->tkr_raw.cycle_last = tk->tkr_mono.cycle_last;
/* Do the ns -> cycle conversion first, using original mult */
tmp = NTP_INTERVAL_LENGTH;
tmp <<= clock->shift;
ntpinterval = tmp;
tmp += clock->mult/2;
do_div(tmp, clock->mult);
if (tmp == 0)
tmp = 1;
interval = (u64) tmp;
tk->cycle_interval = interval;
/* Go back from cycles -> shifted ns */
tk->xtime_interval = interval * clock->mult;
tk->xtime_remainder = ntpinterval - tk->xtime_interval;
tk->raw_interval = interval * clock->mult;
/* if changing clocks, convert xtime_nsec shift units */
if (old_clock) {
int shift_change = clock->shift - old_clock->shift;
if (shift_change < 0) {
tk->tkr_mono.xtime_nsec >>= -shift_change;
tk->tkr_raw.xtime_nsec >>= -shift_change;
} else {
tk->tkr_mono.xtime_nsec <<= shift_change;
tk->tkr_raw.xtime_nsec <<= shift_change;
}
}
tk->tkr_mono.shift = clock->shift;
tk->tkr_raw.shift = clock->shift;
tk->ntp_error = 0;
tk->ntp_error_shift = NTP_SCALE_SHIFT - clock->shift;
tk->ntp_tick = ntpinterval << tk->ntp_error_shift;
/*
* The timekeeper keeps its own mult values for the currently
* active clocksource. These value will be adjusted via NTP
* to counteract clock drifting.
*/
tk->tkr_mono.mult = clock->mult;
tk->tkr_raw.mult = clock->mult;
tk->ntp_err_mult = 0;
tk->skip_second_overflow = 0;
tk->cs_id = clock->id;
/* Coupled clockevent data */
if (IS_ENABLED(CONFIG_GENERIC_CLOCKEVENTS_COUPLED) &&
clock->flags & CLOCK_SOURCE_HAS_COUPLED_CLOCK_EVENT) {
/*
* Aim for an one hour maximum delta and use KHz to handle
* clocksources with a frequency above 4GHz correctly as
* the frequency argument of clocks_calc_mult_shift() is u32.
*/
clocks_calc_mult_shift(&tk->cs_ns_to_cyc_mult, &tk->cs_ns_to_cyc_shift,
NSEC_PER_MSEC, clock->freq_khz, 3600 * 1000);
/*
* Initialize the conversion limit as the previous clocksource
* might have the same shift/mult pair so the quick check in
* tk_update_ns_to_cyc() fails to update it after a clocksource
* change leaving it effectivly zero.
*/
tk->cs_ns_to_cyc_maxns = div_u64(clock->mask, tk->cs_ns_to_cyc_mult);
}
}
/* Timekeeper helper functions. */
static noinline u64 delta_to_ns_safe(const struct tk_read_base *tkr, u64 delta)
{
return mul_u64_u32_add_u64_shr(delta, tkr->mult, tkr->xtime_nsec, tkr->shift);
}
static __always_inline u64 timekeeping_cycles_to_ns(const struct tk_read_base *tkr, u64 cycles)
{
/* Calculate the delta since the last update_wall_time() */
u64 mask = tkr->mask, delta = (cycles - tkr->cycle_last) & mask;
/*
* This detects both negative motion and the case where the delta
* overflows the multiplication with tkr->mult.
*/
if (unlikely(delta > tkr->clock->max_cycles)) {
/*
* Handle clocksource inconsistency between CPUs to prevent
* time from going backwards by checking for the MSB of the
* mask being set in the delta.
*/
if (delta & ~(mask >> 1))
return tkr->xtime_nsec >> tkr->shift;
return delta_to_ns_safe(tkr, delta);
}
return ((delta * tkr->mult) + tkr->xtime_nsec) >> tkr->shift;
}
static __always_inline u64 timekeeping_get_ns(const struct tk_read_base *tkr)
{
return timekeeping_cycles_to_ns(tkr, tk_clock_read(tkr));
}
/**
* update_fast_timekeeper - Update the fast and NMI safe monotonic timekeeper.
* @tkr: Timekeeping readout base from which we take the update
* @tkf: Pointer to NMI safe timekeeper
*
* We want to use this from any context including NMI and tracing /
* instrumenting the timekeeping code itself.
*
* Employ the latch technique; see @write_seqcount_latch.
*
* So if a NMI hits the update of base[0] then it will use base[1]
* which is still consistent. In the worst case this can result is a
* slightly wrong timestamp (a few nanoseconds). See
* @ktime_get_mono_fast_ns.
*/
static void update_fast_timekeeper(const struct tk_read_base *tkr,
struct tk_fast *tkf)
{
struct tk_read_base *base = tkf->base;
/* Force readers off to base[1] */
write_seqcount_latch_begin(&tkf->seq);
/* Update base[0] */
memcpy(base, tkr, sizeof(*base));
/* Force readers back to base[0] */
write_seqcount_latch(&tkf->seq);
/* Update base[1] */
memcpy(base + 1, base, sizeof(*base));
write_seqcount_latch_end(&tkf->seq);
}
static __always_inline u64 __ktime_get_fast_ns(struct tk_fast *tkf)
{
struct tk_read_base *tkr;
unsigned int seq;
u64 now;
do {
seq = read_seqcount_latch(&tkf->seq);
tkr = tkf->base + (seq & 0x01);
now = ktime_to_ns(tkr->base);
now += timekeeping_get_ns(tkr);
} while (read_seqcount_latch_retry(&tkf->seq, seq));
return now;
}
/**
* ktime_get_mono_fast_ns - Fast NMI safe access to clock monotonic
*
* This timestamp is not guaranteed to be monotonic across an update.
* The timestamp is calculated by:
*
* now = base_mono + clock_delta * slope
*
* So if the update lowers the slope, readers who are forced to the
* not yet updated second array are still using the old steeper slope.
*
* tmono
* ^
* | o n
* | o n
* | u
* | o
* |o
* |12345678---> reader order
*
* o = old slope
* u = update
* n = new slope
*
* So reader 6 will observe time going backwards versus reader 5.
*
* While other CPUs are likely to be able to observe that, the only way
* for a CPU local observation is when an NMI hits in the middle of
* the update. Timestamps taken from that NMI context might be ahead
* of the following timestamps. Callers need to be aware of that and
* deal with it.
*/
u64 notrace ktime_get_mono_fast_ns(void)
{
return __ktime_get_fast_ns(&tk_fast_mono);
}
EXPORT_SYMBOL_GPL(ktime_get_mono_fast_ns);
/**
* ktime_get_raw_fast_ns - Fast NMI safe access to clock monotonic raw
*
* Contrary to ktime_get_mono_fast_ns() this is always correct because the
* conversion factor is not affected by NTP/PTP correction.
*/
u64 notrace ktime_get_raw_fast_ns(void)
{
return __ktime_get_fast_ns(&tk_fast_raw);
}
EXPORT_SYMBOL_GPL(ktime_get_raw_fast_ns);
/**
* ktime_get_boot_fast_ns - NMI safe and fast access to boot clock.
*
* To keep it NMI safe since we're accessing from tracing, we're not using a
* separate timekeeper with updates to monotonic clock and boot offset
* protected with seqcounts. This has the following minor side effects:
*
* (1) Its possible that a timestamp be taken after the boot offset is updated
* but before the timekeeper is updated. If this happens, the new boot offset
* is added to the old timekeeping making the clock appear to update slightly
* earlier:
* CPU 0 CPU 1
* timekeeping_inject_sleeptime64()
* __timekeeping_inject_sleeptime(tk, delta);
* timestamp();
* timekeeping_update_staged(tkd, TK_CLEAR_NTP...);
*
* (2) On 32-bit systems, the 64-bit boot offset (tk->offs_boot) may be
* partially updated. Since the tk->offs_boot update is a rare event, this
* should be a rare occurrence which postprocessing should be able to handle.
*
* The caveats vs. timestamp ordering as documented for ktime_get_mono_fast_ns()
* apply as well.
*/
u64 notrace ktime_get_boot_fast_ns(void)
{
struct timekeeper *tk = &tk_core.timekeeper;
return (ktime_get_mono_fast_ns() + ktime_to_ns(data_race(tk->offs_boot)));
}
EXPORT_SYMBOL_GPL(ktime_get_boot_fast_ns);
/**
* ktime_get_tai_fast_ns - NMI safe and fast access to tai clock.
*
* The same limitations as described for ktime_get_boot_fast_ns() apply. The
* mono time and the TAI offset are not read atomically which may yield wrong
* readouts. However, an update of the TAI offset is an rare event e.g., caused
* by settime or adjtimex with an offset. The user of this function has to deal
* with the possibility of wrong timestamps in post processing.
*/
u64 notrace ktime_get_tai_fast_ns(void)
{
struct timekeeper *tk = &tk_core.timekeeper;
return (ktime_get_mono_fast_ns() + ktime_to_ns(data_race(tk->offs_tai)));
}
EXPORT_SYMBOL_GPL(ktime_get_tai_fast_ns);
/**
* ktime_get_real_fast_ns: - NMI safe and fast access to clock realtime.
*
* See ktime_get_mono_fast_ns() for documentation of the time stamp ordering.
*/
u64 ktime_get_real_fast_ns(void)
{
struct tk_fast *tkf = &tk_fast_mono;
struct tk_read_base *tkr;
u64 baser, delta;
unsigned int seq;
do {
seq = raw_read_seqcount_latch(&tkf->seq);
tkr = tkf->base + (seq & 0x01);
baser = ktime_to_ns(tkr->base_real);
delta = timekeeping_get_ns(tkr);
} while (raw_read_seqcount_latch_retry(&tkf->seq, seq));
return baser + delta;
}
EXPORT_SYMBOL_GPL(ktime_get_real_fast_ns);
/**
* halt_fast_timekeeper - Prevent fast timekeeper from accessing clocksource.
* @tk: Timekeeper to snapshot.
*
* It generally is unsafe to access the clocksource after timekeeping has been
* suspended, so take a snapshot of the readout base of @tk and use it as the
* fast timekeeper's readout base while suspended. It will return the same
* number of cycles every time until timekeeping is resumed at which time the
* proper readout base for the fast timekeeper will be restored automatically.
*/
static void halt_fast_timekeeper(const struct timekeeper *tk)
{
static struct tk_read_base tkr_dummy;
const struct tk_read_base *tkr = &tk->tkr_mono;
memcpy(&tkr_dummy, tkr, sizeof(tkr_dummy));
cycles_at_suspend = tk_clock_read(tkr);
tkr_dummy.clock = &dummy_clock;
tkr_dummy.base_real = tkr->base + tk->offs_real;
update_fast_timekeeper(&tkr_dummy, &tk_fast_mono);
tkr = &tk->tkr_raw;
memcpy(&tkr_dummy, tkr, sizeof(tkr_dummy));
tkr_dummy.clock = &dummy_clock;
update_fast_timekeeper(&tkr_dummy, &tk_fast_raw);
}
static RAW_NOTIFIER_HEAD(pvclock_gtod_chain);
static void update_pvclock_gtod(struct timekeeper *tk, bool was_set)
{
raw_notifier_call_chain(&pvclock_gtod_chain, was_set, tk);
}
/**
* pvclock_gtod_register_notifier - register a pvclock timedata update listener
* @nb: Pointer to the notifier block to register
*/
int pvclock_gtod_register_notifier(struct notifier_block *nb)
{
struct timekeeper *tk = &tk_core.timekeeper;
int ret;
guard(raw_spinlock_irqsave)(&tk_core.lock);
ret = raw_notifier_chain_register(&pvclock_gtod_chain, nb);
update_pvclock_gtod(tk, true);
return ret;
}
EXPORT_SYMBOL_GPL(pvclock_gtod_register_notifier);
/**
* pvclock_gtod_unregister_notifier - unregister a pvclock
* timedata update listener
* @nb: Pointer to the notifier block to unregister
*/
int pvclock_gtod_unregister_notifier(struct notifier_block *nb)
{
guard(raw_spinlock_irqsave)(&tk_core.lock);
return raw_notifier_chain_unregister(&pvclock_gtod_chain, nb);
}
EXPORT_SYMBOL_GPL(pvclock_gtod_unregister_notifier);
/*
* tk_update_leap_state - helper to update the next_leap_ktime
*/
static inline void tk_update_leap_state(struct timekeeper *tk)
{
tk->next_leap_ktime = ntp_get_next_leap(tk->id);
if (tk->next_leap_ktime != KTIME_MAX)
/* Convert to monotonic time */
tk->next_leap_ktime = ktime_sub(tk->next_leap_ktime, tk->offs_real);
}
/*
* Leap state update for both shadow and the real timekeeper
* Separate to spare a full memcpy() of the timekeeper.
*/
static void tk_update_leap_state_all(struct tk_data *tkd)
{
write_seqcount_begin(&tkd->seq);
tk_update_leap_state(&tkd->shadow_timekeeper);
tkd->timekeeper.next_leap_ktime = tkd->shadow_timekeeper.next_leap_ktime;
write_seqcount_end(&tkd->seq);
}
/*
* Update the ktime_t based scalar nsec members of the timekeeper
*/
static inline void tk_update_ktime_data(struct timekeeper *tk)
{
u64 seconds;
u32 nsec;
/*
* The xtime based monotonic readout is:
* nsec = (xtime_sec + wtm_sec) * 1e9 + wtm_nsec + now();
* The ktime based monotonic readout is:
* nsec = base_mono + now();
* ==> base_mono = (xtime_sec + wtm_sec) * 1e9 + wtm_nsec
*/
seconds = (u64)(tk->xtime_sec + tk->wall_to_monotonic.tv_sec);
nsec = (u32) tk->wall_to_monotonic.tv_nsec;
tk->tkr_mono.base = ns_to_ktime(seconds * NSEC_PER_SEC + nsec);
/*
* The sum of the nanoseconds portions of xtime and
* wall_to_monotonic can be greater/equal one second. Take
* this into account before updating tk->ktime_sec.
*/
nsec += (u32)(tk->tkr_mono.xtime_nsec >> tk->tkr_mono.shift);
if (nsec >= NSEC_PER_SEC)
seconds++;
tk->ktime_sec = seconds;
/* Update the monotonic raw base */
tk->tkr_raw.base = ns_to_ktime(tk->raw_sec * NSEC_PER_SEC);
}
static inline void tk_update_ns_to_cyc(struct timekeeper *tks, struct timekeeper *tkc)
{
struct tk_read_base *tkrs = &tks->tkr_mono;
struct tk_read_base *tkrc = &tkc->tkr_mono;
unsigned int shift;
if (!IS_ENABLED(CONFIG_GENERIC_CLOCKEVENTS_COUPLED) ||
!(tkrs->clock->flags & CLOCK_SOURCE_HAS_COUPLED_CLOCK_EVENT))
return;
if (tkrs->mult == tkrc->mult && tkrs->shift == tkrc->shift)
return;
/*
* The conversion math is simple:
*
* CS::MULT (1 << NS_TO_CYC_SHIFT)
* --------------- = ----------------------
* (1 << CS:SHIFT) NS_TO_CYC_MULT
*
* Ergo:
*
* NS_TO_CYC_MULT = (1 << (CS::SHIFT + NS_TO_CYC_SHIFT)) / CS::MULT
*
* NS_TO_CYC_SHIFT has been set up in tk_setup_internals()
*/
shift = tkrs->shift + tks->cs_ns_to_cyc_shift;
tks->cs_ns_to_cyc_mult = (u32)div_u64(1ULL << shift, tkrs->mult);
tks->cs_ns_to_cyc_maxns = div_u64(tkrs->clock->mask, tks->cs_ns_to_cyc_mult);
}
/*
* Restore the shadow timekeeper from the real timekeeper.
*/
static void timekeeping_restore_shadow(struct tk_data *tkd)
{
lockdep_assert_held(&tkd->lock);
memcpy(&tkd->shadow_timekeeper, &tkd->timekeeper, sizeof(tkd->timekeeper));
}
static void timekeeping_update_from_shadow(struct tk_data *tkd, unsigned int action)
{
struct timekeeper *tk = &tkd->shadow_timekeeper;
lockdep_assert_held(&tkd->lock);
/*
* Block out readers before running the updates below because that
* updates VDSO and other time related infrastructure. Not blocking
* the readers might let a reader see time going backwards when
* reading from the VDSO after the VDSO update and then reading in
* the kernel from the timekeeper before that got updated.
*/
write_seqcount_begin(&tkd->seq);
if (action & TK_CLEAR_NTP) {
tk->ntp_error = 0;
ntp_clear(tk->id);
}
tk_update_leap_state(tk);
tk_update_ktime_data(tk);
tk->tkr_mono.base_real = tk->tkr_mono.base + tk->offs_real;
if (tk->id == TIMEKEEPER_CORE) {
tk_update_ns_to_cyc(tk, &tkd->timekeeper);
update_vsyscall(tk);
update_pvclock_gtod(tk, action & TK_CLOCK_WAS_SET);
update_fast_timekeeper(&tk->tkr_mono, &tk_fast_mono);
update_fast_timekeeper(&tk->tkr_raw, &tk_fast_raw);
} else if (tk_is_aux(tk)) {
vdso_time_update_aux(tk);
}
if (action & TK_CLOCK_WAS_SET)
tk->clock_was_set_seq++;
/*
* Update the real timekeeper.
*
* We could avoid this memcpy() by switching pointers, but that has
* the downside that the reader side does not longer benefit from
* the cacheline optimized data layout of the timekeeper and requires
* another indirection.
*/
memcpy(&tkd->timekeeper, tk, sizeof(*tk));
write_seqcount_end(&tkd->seq);
}
/**
* timekeeping_forward_now - update clock to the current time
* @tk: Pointer to the timekeeper to update
*
* Forward the current clock to update its state since the last call to
* update_wall_time(). This is useful before significant clock changes,
* as it avoids having to deal with this time offset explicitly.
*/
static void timekeeping_forward_now(struct timekeeper *tk)
{
u64 cycle_now, delta;
cycle_now = tk_clock_read(&tk->tkr_mono);
delta = clocksource_delta(cycle_now, tk->tkr_mono.cycle_last, tk->tkr_mono.mask,
tk->tkr_mono.clock->max_raw_delta);
tk->tkr_mono.cycle_last = cycle_now;
tk->tkr_raw.cycle_last = cycle_now;
while (delta > 0) {
u64 max = tk->tkr_mono.clock->max_cycles;
u64 incr = delta < max ? delta : max;
tk->tkr_mono.xtime_nsec += incr * tk->tkr_mono.mult;
tk->tkr_raw.xtime_nsec += incr * tk->tkr_raw.mult;
tk_normalize_xtime(tk);
delta -= incr;
}
tk_update_coarse_nsecs(tk);
}
/*
* ktime_expiry_to_cycles - Convert a expiry time to clocksource cycles
* @id: Clocksource ID which is required for validity
* @expires_ns: Absolute CLOCK_MONOTONIC expiry time (nsecs) to be converted
* @cycles: Pointer to storage for corresponding absolute cycles value
*
* Convert a CLOCK_MONOTONIC based absolute expiry time to a cycles value
* based on the correlated clocksource of the clockevent device by using
* the base nanoseconds and cycles values of the last timekeeper update and
* converting the delta between @expires_ns and base nanoseconds to cycles.
*
* This only works for clockevent devices which are using a less than or
* equal comparator against the clocksource.
*
* Utilizing this avoids two clocksource reads for such devices, the
* ktime_get() in clockevents_program_event() to calculate the delta expiry
* value and the readout in the device::set_next_event() callback to
* convert the delta back to a absolute comparator value.
*
* Returns: True if @id matches the current clocksource ID, false otherwise
*/
bool ktime_expiry_to_cycles(enum clocksource_ids id, ktime_t expires_ns, u64 *cycles)
{
struct timekeeper *tk = &tk_core.timekeeper;
struct tk_read_base *tkrm = &tk->tkr_mono;
ktime_t base_ns, delta_ns, max_ns;
u64 base_cycles, delta_cycles;
unsigned int seq;
u32 mult, shift;
/*
* Racy check to avoid the seqcount overhead when ID does not match. If
* the relevant clocksource is installed concurrently, then this will
* just delay the switch over to this mechanism until the next event is
* programmed. If the ID is not matching the clock events code will use
* the regular relative set_next_event() callback as before.
*/
if (data_race(tk->cs_id) != id)
return false;
do {
seq = read_seqcount_begin(&tk_core.seq);
if (tk->cs_id != id)
return false;
base_cycles = tkrm->cycle_last;
base_ns = tkrm->base + (tkrm->xtime_nsec >> tkrm->shift);
mult = tk->cs_ns_to_cyc_mult;
shift = tk->cs_ns_to_cyc_shift;
max_ns = tk->cs_ns_to_cyc_maxns;
} while (read_seqcount_retry(&tk_core.seq, seq));
/* Prevent negative deltas and multiplication overflows */
delta_ns = min(expires_ns - base_ns, max_ns);
delta_ns = max(delta_ns, 0);
/* Convert to cycles */
delta_cycles = ((u64)delta_ns * mult) >> shift;
*cycles = base_cycles + delta_cycles;
return true;
}
/**
* ktime_get_real_ts64 - Returns the time of day in a timespec64.
* @ts: pointer to the timespec to be set
*
* Returns the time of day in a timespec64 (WARN if suspended).
*/
void ktime_get_real_ts64(struct timespec64 *ts)
{
struct timekeeper *tk = &tk_core.timekeeper;
unsigned int seq;
u64 nsecs;
WARN_ON(timekeeping_suspended);
do {
seq = read_seqcount_begin(&tk_core.seq);
ts->tv_sec = tk->xtime_sec;
nsecs = timekeeping_get_ns(&tk->tkr_mono);
} while (read_seqcount_retry(&tk_core.seq, seq));
ts->tv_nsec = 0;
timespec64_add_ns(ts, nsecs);
}
EXPORT_SYMBOL(ktime_get_real_ts64);
ktime_t ktime_get(void)
{
struct timekeeper *tk = &tk_core.timekeeper;
unsigned int seq;
ktime_t base;
u64 nsecs;
WARN_ON(timekeeping_suspended);
do {
seq = read_seqcount_begin(&tk_core.seq);
base = tk->tkr_mono.base;
nsecs = timekeeping_get_ns(&tk->tkr_mono);
} while (read_seqcount_retry(&tk_core.seq, seq));
return ktime_add_ns(base, nsecs);
}
EXPORT_SYMBOL_GPL(ktime_get);
u32 ktime_get_resolution_ns(void)
{
struct timekeeper *tk = &tk_core.timekeeper;
unsigned int seq;
u32 nsecs;
WARN_ON(timekeeping_suspended);
do {
seq = read_seqcount_begin(&tk_core.seq);
nsecs = tk->tkr_mono.mult >> tk->tkr_mono.shift;
} while (read_seqcount_retry(&tk_core.seq, seq));
return nsecs;
}
EXPORT_SYMBOL_GPL(ktime_get_resolution_ns);
static const ktime_t *const offsets[TK_OFFS_MAX] = {
[TK_OFFS_REAL] = &tk_core.timekeeper.offs_real,
[TK_OFFS_BOOT] = &tk_core.timekeeper.offs_boot,
[TK_OFFS_TAI] = &tk_core.timekeeper.offs_tai,
};
ktime_t ktime_get_with_offset(enum tk_offsets offs)
{
struct timekeeper *tk = &tk_core.timekeeper;
const ktime_t *offset = offsets[offs];
unsigned int seq;
ktime_t base;
u64 nsecs;
WARN_ON(timekeeping_suspended);
do {
seq = read_seqcount_begin(&tk_core.seq);
base = ktime_add(tk->tkr_mono.base, *offset);
nsecs = timekeeping_get_ns(&tk->tkr_mono);
} while (read_seqcount_retry(&tk_core.seq, seq));
return ktime_add_ns(base, nsecs);
}
EXPORT_SYMBOL_GPL(ktime_get_with_offset);
ktime_t ktime_get_coarse_with_offset(enum tk_offsets offs)
{
struct timekeeper *tk = &tk_core.timekeeper;
const ktime_t *offset = offsets[offs];
unsigned int seq;
ktime_t base;
u64 nsecs;
WARN_ON(timekeeping_suspended);
do {
seq = read_seqcount_begin(&tk_core.seq);
base = ktime_add(tk->tkr_mono.base, *offset);
nsecs = tk->coarse_nsec;
} while (read_seqcount_retry(&tk_core.seq, seq));
return ktime_add_ns(base, nsecs);
}
EXPORT_SYMBOL_GPL(ktime_get_coarse_with_offset);
/**
* ktime_mono_to_any() - convert monotonic time to any other time
* @tmono: time to convert.
* @offs: which offset to use
*/
ktime_t ktime_mono_to_any(ktime_t tmono, enum tk_offsets offs)
{
const ktime_t *offset = offsets[offs];
unsigned int seq;
ktime_t tconv;
if (IS_ENABLED(CONFIG_64BIT)) {
/*
* Paired with WRITE_ONCE()s in tk_set_wall_to_mono() and
* tk_update_sleep_time().
*/
return ktime_add(tmono, READ_ONCE(*offset));
}
do {
seq = read_seqcount_begin(&tk_core.seq);
tconv = ktime_add(tmono, *offset);
} while (read_seqcount_retry(&tk_core.seq, seq));
return tconv;
}
EXPORT_SYMBOL_GPL(ktime_mono_to_any);
/**
* ktime_get_raw - Returns the raw monotonic time in ktime_t format
*/
ktime_t ktime_get_raw(void)
{
struct timekeeper *tk = &tk_core.timekeeper;
unsigned int seq;
ktime_t base;
u64 nsecs;
do {
seq = read_seqcount_begin(&tk_core.seq);
base = tk->tkr_raw.base;
nsecs = timekeeping_get_ns(&tk->tkr_raw);
} while (read_seqcount_retry(&tk_core.seq, seq));
return ktime_add_ns(base, nsecs);
}
EXPORT_SYMBOL_GPL(ktime_get_raw);
/**
* ktime_get_ts64 - get the monotonic clock in timespec64 format
* @ts: pointer to timespec variable
*
* The function calculates the monotonic clock from the realtime
* clock and the wall_to_monotonic offset and stores the result
* in normalized timespec64 format in the variable pointed to by @ts.
*/
void ktime_get_ts64(struct timespec64 *ts)
{
struct timekeeper *tk = &tk_core.timekeeper;
struct timespec64 tomono;
unsigned int seq;
u64 nsec;
WARN_ON(timekeeping_suspended);
do {
seq = read_seqcount_begin(&tk_core.seq);
ts->tv_sec = tk->xtime_sec;
nsec = timekeeping_get_ns(&tk->tkr_mono);
tomono = tk->wall_to_monotonic;
} while (read_seqcount_retry(&tk_core.seq, seq));
ts->tv_sec += tomono.tv_sec;
ts->tv_nsec = 0;
timespec64_add_ns(ts, nsec + tomono.tv_nsec);
}
EXPORT_SYMBOL_GPL(ktime_get_ts64);
/**
* ktime_get_seconds - Get the seconds portion of CLOCK_MONOTONIC
*
* Returns the seconds portion of CLOCK_MONOTONIC with a single non
* serialized read. tk->ktime_sec is of type 'unsigned long' so this
* works on both 32 and 64 bit systems. On 32 bit systems the readout
* covers ~136 years of uptime which should be enough to prevent
* premature wrap arounds.
*/
time64_t ktime_get_seconds(void)
{
struct timekeeper *tk = &tk_core.timekeeper;
WARN_ON(timekeeping_suspended);
return tk->ktime_sec;
}
EXPORT_SYMBOL_GPL(ktime_get_seconds);
/**
* ktime_get_real_seconds - Get the seconds portion of CLOCK_REALTIME
*
* Returns the wall clock seconds since 1970.
*
* For 64bit systems the fast access to tk->xtime_sec is preserved. On
* 32bit systems the access must be protected with the sequence
* counter to provide "atomic" access to the 64bit tk->xtime_sec
* value.
*/
time64_t ktime_get_real_seconds(void)
{
struct timekeeper *tk = &tk_core.timekeeper;
time64_t seconds;
unsigned int seq;
if (IS_ENABLED(CONFIG_64BIT))
return tk->xtime_sec;
do {
seq = read_seqcount_begin(&tk_core.seq);
seconds = tk->xtime_sec;
} while (read_seqcount_retry(&tk_core.seq, seq));
return seconds;
}
EXPORT_SYMBOL_GPL(ktime_get_real_seconds);
/**
* __ktime_get_real_seconds - Unprotected access to CLOCK_REALTIME seconds
*
* The same as ktime_get_real_seconds() but without the sequence counter
* protection. This function is used in restricted contexts like the x86 MCE
* handler and in KGDB. It's unprotected on 32-bit vs. concurrent half
* completed modification and only to be used for such critical contexts.
*
* Returns: Racy snapshot of the CLOCK_REALTIME seconds value
*/
noinstr time64_t __ktime_get_real_seconds(void)
{
struct timekeeper *tk = &tk_core.timekeeper;
return tk->xtime_sec;
}
/**
* ktime_get_snapshot - snapshots the realtime/monotonic raw clocks with counter
* @systime_snapshot: pointer to struct receiving the system time snapshot
*/
void ktime_get_snapshot(struct system_time_snapshot *systime_snapshot)
{
struct timekeeper *tk = &tk_core.timekeeper;
unsigned int seq;
ktime_t base_raw;
ktime_t base_real;
ktime_t base_boot;
u64 nsec_raw;
u64 nsec_real;
u64 now;
WARN_ON_ONCE(timekeeping_suspended);
do {
seq = read_seqcount_begin(&tk_core.seq);
now = tk_clock_read(&tk->tkr_mono);
systime_snapshot->cs_id = tk->tkr_mono.clock->id;
systime_snapshot->cs_was_changed_seq = tk->cs_was_changed_seq;
systime_snapshot->clock_was_set_seq = tk->clock_was_set_seq;
base_real = ktime_add(tk->tkr_mono.base,
tk_core.timekeeper.offs_real);
base_boot = ktime_add(tk->tkr_mono.base,
tk_core.timekeeper.offs_boot);
base_raw = tk->tkr_raw.base;
nsec_real = timekeeping_cycles_to_ns(&tk->tkr_mono, now);
nsec_raw = timekeeping_cycles_to_ns(&tk->tkr_raw, now);
} while (read_seqcount_retry(&tk_core.seq, seq));
systime_snapshot->cycles = now;
systime_snapshot->real = ktime_add_ns(base_real, nsec_real);
systime_snapshot->boot = ktime_add_ns(base_boot, nsec_real);
systime_snapshot->raw = ktime_add_ns(base_raw, nsec_raw);
}
EXPORT_SYMBOL_GPL(ktime_get_snapshot);
/* Scale base by mult/div checking for overflow */
static int scale64_check_overflow(u64 mult, u64 div, u64 *base)
{
u64 tmp, rem;
tmp = div64_u64_rem(*base, div, &rem);
if (((int)sizeof(u64)*8 - fls64(mult) < fls64(tmp)) ||
((int)sizeof(u64)*8 - fls64(mult) < fls64(rem)))
return -EOVERFLOW;
tmp *= mult;
rem = div64_u64(rem * mult, div);
*base = tmp + rem;
return 0;
}
/**
* adjust_historical_crosststamp - adjust crosstimestamp previous to current interval
* @history: Snapshot representing start of history
* @partial_history_cycles: Cycle offset into history (fractional part)
* @total_history_cycles: Total history length in cycles
* @discontinuity: True indicates clock was set on history period
* @ts: Cross timestamp that should be adjusted using
* partial/total ratio
*
* Helper function used by get_device_system_crosststamp() to correct the
* crosstimestamp corresponding to the start of the current interval to the
* system counter value (timestamp point) provided by the driver. The
* total_history_* quantities are the total history starting at the provided
* reference point and ending at the start of the current interval. The cycle
* count between the driver timestamp point and the start of the current
* interval is partial_history_cycles.
*/
static int adjust_historical_crosststamp(struct system_time_snapshot *history,
u64 partial_history_cycles,
u64 total_history_cycles,
bool discontinuity,
struct system_device_crosststamp *ts)
{
struct timekeeper *tk = &tk_core.timekeeper;
u64 corr_raw, corr_real;
bool interp_forward;
int ret;
if (total_history_cycles == 0 || partial_history_cycles == 0)
return 0;
/* Interpolate shortest distance from beginning or end of history */
interp_forward = partial_history_cycles > total_history_cycles / 2;
partial_history_cycles = interp_forward ?
total_history_cycles - partial_history_cycles :
partial_history_cycles;
/*
* Scale the monotonic raw time delta by:
* partial_history_cycles / total_history_cycles
*/
corr_raw = (u64)ktime_to_ns(
ktime_sub(ts->sys_monoraw, history->raw));
ret = scale64_check_overflow(partial_history_cycles,
total_history_cycles, &corr_raw);
if (ret)
return ret;
/*
* If there is a discontinuity in the history, scale monotonic raw
* correction by:
* mult(real)/mult(raw) yielding the realtime correction
* Otherwise, calculate the realtime correction similar to monotonic
* raw calculation
*/
if (discontinuity) {
corr_real = mul_u64_u32_div
(corr_raw, tk->tkr_mono.mult, tk->tkr_raw.mult);
} else {
corr_real = (u64)ktime_to_ns(
ktime_sub(ts->sys_realtime, history->real));
ret = scale64_check_overflow(partial_history_cycles,
total_history_cycles, &corr_real);
if (ret)
return ret;
}
/* Fixup monotonic raw and real time time values */
if (interp_forward) {
ts->sys_monoraw = ktime_add_ns(history->raw, corr_raw);
ts->sys_realtime = ktime_add_ns(history->real, corr_real);
} else {
ts->sys_monoraw = ktime_sub_ns(ts->sys_monoraw, corr_raw);
ts->sys_realtime = ktime_sub_ns(ts->sys_realtime, corr_real);
}
return 0;
}
/*
* timestamp_in_interval - true if ts is chronologically in [start, end]
*
* True if ts occurs chronologically at or after start, and before or at end.
*/
static bool timestamp_in_interval(u64 start, u64 end, u64 ts)
{
if (ts >= start && ts <= end)
return true;
if (start > end && (ts >= start || ts <= end))
return true;
return false;
}
static bool convert_clock(u64 *val, u32 numerator, u32 denominator)
{
u64 rem, res;
if (!numerator || !denominator)
return false;
res = div64_u64_rem(*val, denominator, &rem) * numerator;
*val = res + div_u64(rem * numerator, denominator);
return true;
}
static bool convert_base_to_cs(struct system_counterval_t *scv)
{
struct clocksource *cs = tk_core.timekeeper.tkr_mono.clock;
struct clocksource_base *base;
u32 num, den;
/* The timestamp was taken from the time keeper clock source */
if (cs->id == scv->cs_id)
return true;
/*
* Check whether cs_id matches the base clock. Prevent the compiler from
* re-evaluating @base as the clocksource might change concurrently.
*/
base = READ_ONCE(cs->base);
if (!base || base->id != scv->cs_id)
return false;
num = scv->use_nsecs ? cs->freq_khz : base->numerator;
den = scv->use_nsecs ? USEC_PER_SEC : base->denominator;
if (!convert_clock(&scv->cycles, num, den))
return false;
scv->cycles += base->offset;
return true;
}
static bool convert_cs_to_base(u64 *cycles, enum clocksource_ids base_id)
{
struct clocksource *cs = tk_core.timekeeper.tkr_mono.clock;
struct clocksource_base *base;
/*
* Check whether base_id matches the base clock. Prevent the compiler from
* re-evaluating @base as the clocksource might change concurrently.
*/
base = READ_ONCE(cs->base);
if (!base || base->id != base_id)
return false;
*cycles -= base->offset;
if (!convert_clock(cycles, base->denominator, base->numerator))
return false;
return true;
}
static bool convert_ns_to_cs(u64 *delta)
{
struct tk_read_base *tkr = &tk_core.timekeeper.tkr_mono;
if (BITS_TO_BYTES(fls64(*delta) + tkr->shift) >= sizeof(*delta))
return false;
*delta = div_u64((*delta << tkr->shift) - tkr->xtime_nsec, tkr->mult);
return true;
}
/**
* ktime_real_to_base_clock() - Convert CLOCK_REALTIME timestamp to a base clock timestamp
* @treal: CLOCK_REALTIME timestamp to convert
* @base_id: base clocksource id
* @cycles: pointer to store the converted base clock timestamp
*
* Converts a supplied, future realtime clock value to the corresponding base clock value.
*
* Return: true if the conversion is successful, false otherwise.
*/
bool ktime_real_to_base_clock(ktime_t treal, enum clocksource_ids base_id, u64 *cycles)
{
struct timekeeper *tk = &tk_core.timekeeper;
unsigned int seq;
u64 delta;
do {
seq = read_seqcount_begin(&tk_core.seq);
if ((u64)treal < tk->tkr_mono.base_real)
return false;
delta = (u64)treal - tk->tkr_mono.base_real;
if (!convert_ns_to_cs(&delta))
return false;
*cycles = tk->tkr_mono.cycle_last + delta;
if (!convert_cs_to_base(cycles, base_id))
return false;
} while (read_seqcount_retry(&tk_core.seq, seq));
return true;
}
EXPORT_SYMBOL_GPL(ktime_real_to_base_clock);
/**
* get_device_system_crosststamp - Synchronously capture system/device timestamp
* @get_time_fn: Callback to get simultaneous device time and
* system counter from the device driver
* @ctx: Context passed to get_time_fn()
* @history_begin: Historical reference point used to interpolate system
* time when counter provided by the driver is before the current interval
* @xtstamp: Receives simultaneously captured system and device time
*
* Reads a timestamp from a device and correlates it to system time
*/
int get_device_system_crosststamp(int (*get_time_fn)
(ktime_t *device_time,
struct system_counterval_t *sys_counterval,
void *ctx),
void *ctx,
struct system_time_snapshot *history_begin,
struct system_device_crosststamp *xtstamp)
{
struct system_counterval_t system_counterval = {};
struct timekeeper *tk = &tk_core.timekeeper;
u64 cycles, now, interval_start;
unsigned int clock_was_set_seq = 0;
ktime_t base_real, base_raw;
u64 nsec_real, nsec_raw;
u8 cs_was_changed_seq;
unsigned int seq;
bool do_interp;
int ret;
do {
seq = read_seqcount_begin(&tk_core.seq);
/*
* Try to synchronously capture device time and a system
* counter value calling back into the device driver
*/
ret = get_time_fn(&xtstamp->device, &system_counterval, ctx);
if (ret)
return ret;
/*
* Verify that the clocksource ID associated with the captured
* system counter value is the same as for the currently
* installed timekeeper clocksource
*/
if (system_counterval.cs_id == CSID_GENERIC ||
!convert_base_to_cs(&system_counterval))
return -ENODEV;
cycles = system_counterval.cycles;
/*
* Check whether the system counter value provided by the
* device driver is on the current timekeeping interval.
*/
now = tk_clock_read(&tk->tkr_mono);
interval_start = tk->tkr_mono.cycle_last;
if (!timestamp_in_interval(interval_start, now, cycles)) {
clock_was_set_seq = tk->clock_was_set_seq;
cs_was_changed_seq = tk->cs_was_changed_seq;
cycles = interval_start;
do_interp = true;
} else {
do_interp = false;
}
base_real = ktime_add(tk->tkr_mono.base,
tk_core.timekeeper.offs_real);
base_raw = tk->tkr_raw.base;
nsec_real = timekeeping_cycles_to_ns(&tk->tkr_mono, cycles);
nsec_raw = timekeeping_cycles_to_ns(&tk->tkr_raw, cycles);
} while (read_seqcount_retry(&tk_core.seq, seq));
xtstamp->sys_realtime = ktime_add_ns(base_real, nsec_real);
xtstamp->sys_monoraw = ktime_add_ns(base_raw, nsec_raw);
/*
* Interpolate if necessary, adjusting back from the start of the
* current interval
*/
if (do_interp) {
u64 partial_history_cycles, total_history_cycles;
bool discontinuity;
/*
* Check that the counter value is not before the provided
* history reference and that the history doesn't cross a
* clocksource change
*/
if (!history_begin ||
!timestamp_in_interval(history_begin->cycles,
cycles, system_counterval.cycles) ||
history_begin->cs_was_changed_seq != cs_was_changed_seq)
return -EINVAL;
partial_history_cycles = cycles - system_counterval.cycles;
total_history_cycles = cycles - history_begin->cycles;
discontinuity =
history_begin->clock_was_set_seq != clock_was_set_seq;
ret = adjust_historical_crosststamp(history_begin,
partial_history_cycles,
total_history_cycles,
discontinuity, xtstamp);
if (ret)
return ret;
}
return 0;
}
EXPORT_SYMBOL_GPL(get_device_system_crosststamp);
/**
* timekeeping_clocksource_has_base - Check whether the current clocksource
* is based on given a base clock
* @id: base clocksource ID
*
* Note: The return value is a snapshot which can become invalid right
* after the function returns.
*
* Return: true if the timekeeper clocksource has a base clock with @id,
* false otherwise
*/
bool timekeeping_clocksource_has_base(enum clocksource_ids id)
{
/*
* This is a snapshot, so no point in using the sequence
* count. Just prevent the compiler from re-evaluating @base as the
* clocksource might change concurrently.
*/
struct clocksource_base *base = READ_ONCE(tk_core.timekeeper.tkr_mono.clock->base);
return base ? base->id == id : false;
}
EXPORT_SYMBOL_GPL(timekeeping_clocksource_has_base);
/**
* do_settimeofday64 - Sets the time of day.
* @ts: pointer to the timespec64 variable containing the new time
*
* Sets the time of day to the new time and update NTP and notify hrtimers
*/
int do_settimeofday64(const struct timespec64 *ts)
{
struct timespec64 ts_delta, xt;
if (!timespec64_valid_settod(ts))
return -EINVAL;
scoped_guard (raw_spinlock_irqsave, &tk_core.lock) {
struct timekeeper *tks = &tk_core.shadow_timekeeper;
timekeeping_forward_now(tks);
xt = tk_xtime(tks);
ts_delta = timespec64_sub(*ts, xt);
if (timespec64_compare(&tks->wall_to_monotonic, &ts_delta) > 0) {
timekeeping_restore_shadow(&tk_core);
return -EINVAL;
}
tk_set_wall_to_mono(tks, timespec64_sub(tks->wall_to_monotonic, ts_delta));
tk_set_xtime(tks, ts);
timekeeping_update_from_shadow(&tk_core, TK_UPDATE_ALL);
}
/* Signal hrtimers about time change */
clock_was_set(CLOCK_SET_WALL);
audit_tk_injoffset(ts_delta);
add_device_randomness(ts, sizeof(*ts));
return 0;
}
EXPORT_SYMBOL(do_settimeofday64);
static inline bool timekeeper_is_core_tk(struct timekeeper *tk)
{
return !IS_ENABLED(CONFIG_POSIX_AUX_CLOCKS) || tk->id == TIMEKEEPER_CORE;
}
/**
* __timekeeping_inject_offset - Adds or subtracts from the current time.
* @tkd: Pointer to the timekeeper to modify
* @ts: Pointer to the timespec variable containing the offset
*
* Adds or subtracts an offset value from the current time.
*/
static int __timekeeping_inject_offset(struct tk_data *tkd, const struct timespec64 *ts)
{
struct timekeeper *tks = &tkd->shadow_timekeeper;
struct timespec64 tmp;
if (ts->tv_nsec < 0 || ts->tv_nsec >= NSEC_PER_SEC)
return -EINVAL;
timekeeping_forward_now(tks);
if (timekeeper_is_core_tk(tks)) {
/* Make sure the proposed value is valid */
tmp = timespec64_add(tk_xtime(tks), *ts);
if (timespec64_compare(&tks->wall_to_monotonic, ts) > 0 ||
!timespec64_valid_settod(&tmp)) {
timekeeping_restore_shadow(tkd);
return -EINVAL;
}
tk_xtime_add(tks, ts);
tk_set_wall_to_mono(tks, timespec64_sub(tks->wall_to_monotonic, *ts));
} else {
struct tk_read_base *tkr_mono = &tks->tkr_mono;
ktime_t now, offs;
/* Get the current time */
now = ktime_add_ns(tkr_mono->base, timekeeping_get_ns(tkr_mono));
/* Add the relative offset change */
offs = ktime_add(tks->offs_aux, timespec64_to_ktime(*ts));
/* Prevent that the resulting time becomes negative */
if (ktime_add(now, offs) < 0) {
timekeeping_restore_shadow(tkd);
return -EINVAL;
}
tk_update_aux_offs(tks, offs);
}
timekeeping_update_from_shadow(tkd, TK_UPDATE_ALL);
return 0;
}
static int timekeeping_inject_offset(const struct timespec64 *ts)
{
int ret;
scoped_guard (raw_spinlock_irqsave, &tk_core.lock)
ret = __timekeeping_inject_offset(&tk_core, ts);
/* Signal hrtimers about time change */
if (!ret)
clock_was_set(CLOCK_SET_WALL);
return ret;
}
/*
* Indicates if there is an offset between the system clock and the hardware
* clock/persistent clock/rtc.
*/
int persistent_clock_is_local;
/*
* Adjust the time obtained from the CMOS to be UTC time instead of
* local time.
*
* This is ugly, but preferable to the alternatives. Otherwise we
* would either need to write a program to do it in /etc/rc (and risk
* confusion if the program gets run more than once; it would also be
* hard to make the program warp the clock precisely n hours) or
* compile in the timezone information into the kernel. Bad, bad....
*
* - TYT, 1992-01-01
*
* The best thing to do is to keep the CMOS clock in universal time (UTC)
* as real UNIX machines always do it. This avoids all headaches about
* daylight saving times and warping kernel clocks.
*/
void timekeeping_warp_clock(void)
{
if (sys_tz.tz_minuteswest != 0) {
struct timespec64 adjust;
persistent_clock_is_local = 1;
adjust.tv_sec = sys_tz.tz_minuteswest * 60;
adjust.tv_nsec = 0;
timekeeping_inject_offset(&adjust);
}
}
/*
* __timekeeping_set_tai_offset - Sets the TAI offset from UTC and monotonic
*/
static void __timekeeping_set_tai_offset(struct timekeeper *tk, s32 tai_offset)
{
tk->tai_offset = tai_offset;
tk->offs_tai = ktime_add(tk->offs_real, ktime_set(tai_offset, 0));
}
/*
* change_clocksource - Swaps clocksources if a new one is available
*
* Accumulates current time interval and initializes new clocksource
*/
static int change_clocksource(void *data)
{
struct clocksource *new = data, *old = NULL;
/*
* If the clocksource is in a module, get a module reference.
* Succeeds for built-in code (owner == NULL) as well. Abort if the
* reference can't be acquired.
*/
if (!try_module_get(new->owner))
return 0;
/* Abort if the device can't be enabled */
if (new->enable && new->enable(new) != 0) {
module_put(new->owner);
return 0;
}
scoped_guard (raw_spinlock_irqsave, &tk_core.lock) {
struct timekeeper *tks = &tk_core.shadow_timekeeper;
timekeeping_forward_now(tks);
old = tks->tkr_mono.clock;
tk_setup_internals(tks, new);
timekeeping_update_from_shadow(&tk_core, TK_UPDATE_ALL);
}
tk_aux_update_clocksource();
if (old) {
if (old->disable)
old->disable(old);
module_put(old->owner);
}
return 0;
}
/**
* timekeeping_notify - Install a new clock source
* @clock: pointer to the clock source
*
* This function is called from clocksource.c after a new, better clock
* source has been registered. The caller holds the clocksource_mutex.
*/
int timekeeping_notify(struct clocksource *clock)
{
struct timekeeper *tk = &tk_core.timekeeper;
if (tk->tkr_mono.clock == clock)
return 0;
/* Disable inlined reads accross the clocksource switch */
clocksource_disable_inline_read();
stop_machine(change_clocksource, clock, NULL);
/*
* If the clocksource has been selected and supports inlined reads
* enable the branch.
*/
if (tk->tkr_mono.clock == clock && clock->flags & CLOCK_SOURCE_CAN_INLINE_READ)
clocksource_enable_inline_read();
tick_clock_notify();
return tk->tkr_mono.clock == clock ? 0 : -1;
}
/**
* ktime_get_raw_ts64 - Returns the raw monotonic time in a timespec
* @ts: pointer to the timespec64 to be set
*
* Returns the raw monotonic time (completely un-modified by ntp)
*/
void ktime_get_raw_ts64(struct timespec64 *ts)
{
struct timekeeper *tk = &tk_core.timekeeper;
unsigned int seq;
u64 nsecs;
do {
seq = read_seqcount_begin(&tk_core.seq);
ts->tv_sec = tk->raw_sec;
nsecs = timekeeping_get_ns(&tk->tkr_raw);
} while (read_seqcount_retry(&tk_core.seq, seq));
ts->tv_nsec = 0;
timespec64_add_ns(ts, nsecs);
}
EXPORT_SYMBOL(ktime_get_raw_ts64);
/**
* ktime_get_clock_ts64 - Returns time of a clock in a timespec
* @id: POSIX clock ID of the clock to read
* @ts: Pointer to the timespec64 to be set
*
* The timestamp is invalidated (@ts->sec is set to -1) if the
* clock @id is not available.
*/
void ktime_get_clock_ts64(clockid_t id, struct timespec64 *ts)
{
/* Invalidate time stamp */
ts->tv_sec = -1;
ts->tv_nsec = 0;
switch (id) {
case CLOCK_REALTIME:
ktime_get_real_ts64(ts);
return;
case CLOCK_MONOTONIC:
ktime_get_ts64(ts);
return;
case CLOCK_MONOTONIC_RAW:
ktime_get_raw_ts64(ts);
return;
case CLOCK_AUX ... CLOCK_AUX_LAST:
if (IS_ENABLED(CONFIG_POSIX_AUX_CLOCKS))
ktime_get_aux_ts64(id, ts);
return;
default:
WARN_ON_ONCE(1);
}
}
EXPORT_SYMBOL_GPL(ktime_get_clock_ts64);
/**
* timekeeping_valid_for_hres - Check if timekeeping is suitable for hres
*/
int timekeeping_valid_for_hres(void)
{
struct timekeeper *tk = &tk_core.timekeeper;
unsigned int seq;
int ret;
do {
seq = read_seqcount_begin(&tk_core.seq);
ret = tk->tkr_mono.clock->flags & CLOCK_SOURCE_VALID_FOR_HRES;
} while (read_seqcount_retry(&tk_core.seq, seq));
return ret;
}
/**
* timekeeping_max_deferment - Returns max time the clocksource can be deferred
*/
u64 timekeeping_max_deferment(void)
{
struct timekeeper *tk = &tk_core.timekeeper;
unsigned int seq;
u64 ret;
do {
seq = read_seqcount_begin(&tk_core.seq);
ret = tk->tkr_mono.clock->max_idle_ns;
} while (read_seqcount_retry(&tk_core.seq, seq));
return ret;
}
/**
* read_persistent_clock64 - Return time from the persistent clock.
* @ts: Pointer to the storage for the readout value
*
* Weak dummy function for arches that do not yet support it.
* Reads the time from the battery backed persistent clock.
* Returns a timespec with tv_sec=0 and tv_nsec=0 if unsupported.
*
* XXX - Do be sure to remove it once all arches implement it.
*/
void __weak read_persistent_clock64(struct timespec64 *ts)
{
ts->tv_sec = 0;
ts->tv_nsec = 0;
}
/**
* read_persistent_wall_and_boot_offset - Read persistent clock, and also offset
* from the boot.
* @wall_time: current time as returned by persistent clock
* @boot_offset: offset that is defined as wall_time - boot_time
*
* Weak dummy function for arches that do not yet support it.
*
* The default function calculates offset based on the current value of
* local_clock(). This way architectures that support sched_clock() but don't
* support dedicated boot time clock will provide the best estimate of the
* boot time.
*/
void __weak __init
read_persistent_wall_and_boot_offset(struct timespec64 *wall_time,
struct timespec64 *boot_offset)
{
read_persistent_clock64(wall_time);
*boot_offset = ns_to_timespec64(local_clock());
}
static __init void tkd_basic_setup(struct tk_data *tkd, enum timekeeper_ids tk_id, bool valid)
{
raw_spin_lock_init(&tkd->lock);
seqcount_raw_spinlock_init(&tkd->seq, &tkd->lock);
tkd->timekeeper.id = tkd->shadow_timekeeper.id = tk_id;
tkd->timekeeper.clock_valid = tkd->shadow_timekeeper.clock_valid = valid;
}
/*
* Flag reflecting whether timekeeping_resume() has injected sleeptime.
*
* The flag starts of false and is only set when a suspend reaches
* timekeeping_suspend(), timekeeping_resume() sets it to false when the
* timekeeper clocksource is not stopping across suspend and has been
* used to update sleep time. If the timekeeper clocksource has stopped
* then the flag stays true and is used by the RTC resume code to decide
* whether sleeptime must be injected and if so the flag gets false then.
*
* If a suspend fails before reaching timekeeping_resume() then the flag
* stays false and prevents erroneous sleeptime injection.
*/
static bool suspend_timing_needed;
/* Flag for if there is a persistent clock on this platform */
static bool persistent_clock_exists;
/*
* timekeeping_init - Initializes the clocksource and common timekeeping values
*/
void __init timekeeping_init(void)
{
struct timespec64 wall_time, boot_offset, wall_to_mono;
struct timekeeper *tks = &tk_core.shadow_timekeeper;
struct clocksource *clock;
tkd_basic_setup(&tk_core, TIMEKEEPER_CORE, true);
tk_aux_setup();
read_persistent_wall_and_boot_offset(&wall_time, &boot_offset);
if (timespec64_valid_settod(&wall_time) &&
timespec64_to_ns(&wall_time) > 0) {
persistent_clock_exists = true;
} else if (timespec64_to_ns(&wall_time) != 0) {
pr_warn("Persistent clock returned invalid value");
wall_time = (struct timespec64){0};
}
if (timespec64_compare(&wall_time, &boot_offset) < 0)
boot_offset = (struct timespec64){0};
/*
* We want set wall_to_mono, so the following is true:
* wall time + wall_to_mono = boot time
*/
wall_to_mono = timespec64_sub(boot_offset, wall_time);
guard(raw_spinlock_irqsave)(&tk_core.lock);
ntp_init();
clock = clocksource_default_clock();
if (clock->enable)
clock->enable(clock);
tk_setup_internals(tks, clock);
tk_set_xtime(tks, &wall_time);
tks->raw_sec = 0;
tk_set_wall_to_mono(tks, wall_to_mono);
timekeeping_update_from_shadow(&tk_core, TK_CLOCK_WAS_SET);
}
/* time in seconds when suspend began for persistent clock */
static struct timespec64 timekeeping_suspend_time;
/**
* __timekeeping_inject_sleeptime - Internal function to add sleep interval
* @tk: Pointer to the timekeeper to be updated
* @delta: Pointer to the delta value in timespec64 format
*
* Takes a timespec offset measuring a suspend interval and properly
* adds the sleep offset to the timekeeping variables.
*/
static void __timekeeping_inject_sleeptime(struct timekeeper *tk,
const struct timespec64 *delta)
{
if (!timespec64_valid_strict(delta)) {
printk_deferred(KERN_WARNING
"__timekeeping_inject_sleeptime: Invalid "
"sleep delta value!\n");
return;
}
tk_xtime_add(tk, delta);
tk_set_wall_to_mono(tk, timespec64_sub(tk->wall_to_monotonic, *delta));
tk_update_sleep_time(tk, timespec64_to_ktime(*delta));
tk_debug_account_sleep_time(delta);
}
#if defined(CONFIG_PM_SLEEP) && defined(CONFIG_RTC_HCTOSYS_DEVICE)
/*
* We have three kinds of time sources to use for sleep time
* injection, the preference order is:
* 1) non-stop clocksource
* 2) persistent clock (ie: RTC accessible when irqs are off)
* 3) RTC
*
* 1) and 2) are used by timekeeping, 3) by RTC subsystem.
* If system has neither 1) nor 2), 3) will be used finally.
*
*
* If timekeeping has injected sleeptime via either 1) or 2),
* 3) becomes needless, so in this case we don't need to call
* rtc_resume(), and this is what timekeeping_rtc_skipresume()
* means.
*/
bool timekeeping_rtc_skipresume(void)
{
return !suspend_timing_needed;
}
/*
* 1) can be determined whether to use or not only when doing
* timekeeping_resume() which is invoked after rtc_suspend(),
* so we can't skip rtc_suspend() surely if system has 1).
*
* But if system has 2), 2) will definitely be used, so in this
* case we don't need to call rtc_suspend(), and this is what
* timekeeping_rtc_skipsuspend() means.
*/
bool timekeeping_rtc_skipsuspend(void)
{
return persistent_clock_exists;
}
/**
* timekeeping_inject_sleeptime64 - Adds suspend interval to timeekeeping values
* @delta: pointer to a timespec64 delta value
*
* This hook is for architectures that cannot support read_persistent_clock64
* because their RTC/persistent clock is only accessible when irqs are enabled.
* and also don't have an effective nonstop clocksource.
*
* This function should only be called by rtc_resume(), and allows
* a suspend offset to be injected into the timekeeping values.
*/
void timekeeping_inject_sleeptime64(const struct timespec64 *delta)
{
scoped_guard(raw_spinlock_irqsave, &tk_core.lock) {
struct timekeeper *tks = &tk_core.shadow_timekeeper;
suspend_timing_needed = false;
timekeeping_forward_now(tks);
__timekeeping_inject_sleeptime(tks, delta);
timekeeping_update_from_shadow(&tk_core, TK_UPDATE_ALL);
}
/* Signal hrtimers about time change */
clock_was_set(CLOCK_SET_WALL | CLOCK_SET_BOOT);
}
#endif
/**
* timekeeping_resume - Resumes the generic timekeeping subsystem.
*/
void timekeeping_resume(void)
{
struct timekeeper *tks = &tk_core.shadow_timekeeper;
struct clocksource *clock = tks->tkr_mono.clock;
struct timespec64 ts_new, ts_delta;
bool inject_sleeptime = false;
u64 cycle_now, nsec;
unsigned long flags;
read_persistent_clock64(&ts_new);
clockevents_resume();
clocksource_resume();
raw_spin_lock_irqsave(&tk_core.lock, flags);
/*
* After system resumes, we need to calculate the suspended time and
* compensate it for the OS time. There are 3 sources that could be
* used: Nonstop clocksource during suspend, persistent clock and rtc
* device.
*
* One specific platform may have 1 or 2 or all of them, and the
* preference will be:
* suspend-nonstop clocksource -> persistent clock -> rtc
* The less preferred source will only be tried if there is no better
* usable source. The rtc part is handled separately in rtc core code.
*/
cycle_now = tk_clock_read(&tks->tkr_mono);
nsec = clocksource_stop_suspend_timing(clock, cycle_now);
if (nsec > 0) {
ts_delta = ns_to_timespec64(nsec);
inject_sleeptime = true;
} else if (timespec64_compare(&ts_new, &timekeeping_suspend_time) > 0) {
ts_delta = timespec64_sub(ts_new, timekeeping_suspend_time);
inject_sleeptime = true;
}
if (inject_sleeptime) {
suspend_timing_needed = false;
__timekeeping_inject_sleeptime(tks, &ts_delta);
}
/* Re-base the last cycle value */
tks->tkr_mono.cycle_last = cycle_now;
tks->tkr_raw.cycle_last = cycle_now;
tks->ntp_error = 0;
timekeeping_suspended = 0;
timekeeping_update_from_shadow(&tk_core, TK_CLOCK_WAS_SET);
raw_spin_unlock_irqrestore(&tk_core.lock, flags);
touch_softlockup_watchdog();
/* Resume the clockevent device(s) and hrtimers */
tick_resume();
/* Notify timerfd as resume is equivalent to clock_was_set() */
timerfd_resume();
}
static void timekeeping_syscore_resume(void *data)
{
timekeeping_resume();
}
int timekeeping_suspend(void)
{
struct timekeeper *tks = &tk_core.shadow_timekeeper;
struct timespec64 delta, delta_delta;
static struct timespec64 old_delta;
struct clocksource *curr_clock;
unsigned long flags;
u64 cycle_now;
read_persistent_clock64(&timekeeping_suspend_time);
/*
* On some systems the persistent_clock can not be detected at
* timekeeping_init by its return value, so if we see a valid
* value returned, update the persistent_clock_exists flag.
*/
if (timekeeping_suspend_time.tv_sec || timekeeping_suspend_time.tv_nsec)
persistent_clock_exists = true;
suspend_timing_needed = true;
raw_spin_lock_irqsave(&tk_core.lock, flags);
timekeeping_forward_now(tks);
timekeeping_suspended = 1;
/*
* Since we've called forward_now, cycle_last stores the value
* just read from the current clocksource. Save this to potentially
* use in suspend timing.
*/
curr_clock = tks->tkr_mono.clock;
cycle_now = tks->tkr_mono.cycle_last;
clocksource_start_suspend_timing(curr_clock, cycle_now);
if (persistent_clock_exists) {
/*
* To avoid drift caused by repeated suspend/resumes,
* which each can add ~1 second drift error,
* try to compensate so the difference in system time
* and persistent_clock time stays close to constant.
*/
delta = timespec64_sub(tk_xtime(tks), timekeeping_suspend_time);
delta_delta = timespec64_sub(delta, old_delta);
if (abs(delta_delta.tv_sec) >= 2) {
/*
* if delta_delta is too large, assume time correction
* has occurred and set old_delta to the current delta.
*/
old_delta = delta;
} else {
/* Otherwise try to adjust old_system to compensate */
timekeeping_suspend_time =
timespec64_add(timekeeping_suspend_time, delta_delta);
}
}
timekeeping_update_from_shadow(&tk_core, 0);
halt_fast_timekeeper(tks);
raw_spin_unlock_irqrestore(&tk_core.lock, flags);
tick_suspend();
clocksource_suspend();
clockevents_suspend();
return 0;
}
static int timekeeping_syscore_suspend(void *data)
{
return timekeeping_suspend();
}
/* sysfs resume/suspend bits for timekeeping */
static const struct syscore_ops timekeeping_syscore_ops = {
.resume = timekeeping_syscore_resume,
.suspend = timekeeping_syscore_suspend,
};
static struct syscore timekeeping_syscore = {
.ops = &timekeeping_syscore_ops,
};
static int __init timekeeping_init_ops(void)
{
register_syscore(&timekeeping_syscore);
return 0;
}
device_initcall(timekeeping_init_ops);
/*
* Apply a multiplier adjustment to the timekeeper
*/
static __always_inline void timekeeping_apply_adjustment(struct timekeeper *tk,
s64 offset,
s32 mult_adj)
{
s64 interval = tk->cycle_interval;
if (mult_adj == 0) {
return;
} else if (mult_adj == -1) {
interval = -interval;
offset = -offset;
} else if (mult_adj != 1) {
interval *= mult_adj;
offset *= mult_adj;
}
/*
* So the following can be confusing.
*
* To keep things simple, lets assume mult_adj == 1 for now.
*
* When mult_adj != 1, remember that the interval and offset values
* have been appropriately scaled so the math is the same.
*
* The basic idea here is that we're increasing the multiplier
* by one, this causes the xtime_interval to be incremented by
* one cycle_interval. This is because:
* xtime_interval = cycle_interval * mult
* So if mult is being incremented by one:
* xtime_interval = cycle_interval * (mult + 1)
* Its the same as:
* xtime_interval = (cycle_interval * mult) + cycle_interval
* Which can be shortened to:
* xtime_interval += cycle_interval
*
* So offset stores the non-accumulated cycles. Thus the current
* time (in shifted nanoseconds) is:
* now = (offset * adj) + xtime_nsec
* Now, even though we're adjusting the clock frequency, we have
* to keep time consistent. In other words, we can't jump back
* in time, and we also want to avoid jumping forward in time.
*
* So given the same offset value, we need the time to be the same
* both before and after the freq adjustment.
* now = (offset * adj_1) + xtime_nsec_1
* now = (offset * adj_2) + xtime_nsec_2
* So:
* (offset * adj_1) + xtime_nsec_1 =
* (offset * adj_2) + xtime_nsec_2
* And we know:
* adj_2 = adj_1 + 1
* So:
* (offset * adj_1) + xtime_nsec_1 =
* (offset * (adj_1+1)) + xtime_nsec_2
* (offset * adj_1) + xtime_nsec_1 =
* (offset * adj_1) + offset + xtime_nsec_2
* Canceling the sides:
* xtime_nsec_1 = offset + xtime_nsec_2
* Which gives us:
* xtime_nsec_2 = xtime_nsec_1 - offset
* Which simplifies to:
* xtime_nsec -= offset
*/
if ((mult_adj > 0) && (tk->tkr_mono.mult + mult_adj < mult_adj)) {
/* NTP adjustment caused clocksource mult overflow */
WARN_ON_ONCE(1);
return;
}
tk->tkr_mono.mult += mult_adj;
tk->xtime_interval += interval;
tk->tkr_mono.xtime_nsec -= offset;
}
/*
* Adjust the timekeeper's multiplier to the correct frequency
* and also to reduce the accumulated error value.
*/
static void timekeeping_adjust(struct timekeeper *tk, s64 offset)
{
u64 ntp_tl = ntp_tick_length(tk->id);
u32 mult;
/*
* Determine the multiplier from the current NTP tick length.
* Avoid expensive division when the tick length doesn't change.
*/
if (likely(tk->ntp_tick == ntp_tl)) {
mult = tk->tkr_mono.mult - tk->ntp_err_mult;
} else {
tk->ntp_tick = ntp_tl;
mult = div64_u64((tk->ntp_tick >> tk->ntp_error_shift) -
tk->xtime_remainder, tk->cycle_interval);
}
/*
* If the clock is behind the NTP time, increase the multiplier by 1
* to catch up with it. If it's ahead and there was a remainder in the
* tick division, the clock will slow down. Otherwise it will stay
* ahead until the tick length changes to a non-divisible value.
*/
tk->ntp_err_mult = tk->ntp_error > 0 ? 1 : 0;
mult += tk->ntp_err_mult;
timekeeping_apply_adjustment(tk, offset, mult - tk->tkr_mono.mult);
if (unlikely(tk->tkr_mono.clock->maxadj &&
(abs(tk->tkr_mono.mult - tk->tkr_mono.clock->mult)
> tk->tkr_mono.clock->maxadj))) {
printk_once(KERN_WARNING
"Adjusting %s more than 11%% (%ld vs %ld)\n",
tk->tkr_mono.clock->name, (long)tk->tkr_mono.mult,
(long)tk->tkr_mono.clock->mult + tk->tkr_mono.clock->maxadj);
}
/*
* It may be possible that when we entered this function, xtime_nsec
* was very small. Further, if we're slightly speeding the clocksource
* in the code above, its possible the required corrective factor to
* xtime_nsec could cause it to underflow.
*
* Now, since we have already accumulated the second and the NTP
* subsystem has been notified via second_overflow(), we need to skip
* the next update.
*/
if (unlikely((s64)tk->tkr_mono.xtime_nsec < 0)) {
tk->tkr_mono.xtime_nsec += (u64)NSEC_PER_SEC <<
tk->tkr_mono.shift;
tk->xtime_sec--;
tk->skip_second_overflow = 1;
}
}
/*
* accumulate_nsecs_to_secs - Accumulates nsecs into secs
*
* Helper function that accumulates the nsecs greater than a second
* from the xtime_nsec field to the xtime_secs field.
* It also calls into the NTP code to handle leapsecond processing.
*/
static inline unsigned int accumulate_nsecs_to_secs(struct timekeeper *tk)
{
u64 nsecps = (u64)NSEC_PER_SEC << tk->tkr_mono.shift;
unsigned int clock_set = 0;
while (tk->tkr_mono.xtime_nsec >= nsecps) {
int leap;
tk->tkr_mono.xtime_nsec -= nsecps;
tk->xtime_sec++;
/*
* Skip NTP update if this second was accumulated before,
* i.e. xtime_nsec underflowed in timekeeping_adjust()
*/
if (unlikely(tk->skip_second_overflow)) {
tk->skip_second_overflow = 0;
continue;
}
/* Figure out if its a leap sec and apply if needed */
leap = second_overflow(tk->id, tk->xtime_sec);
if (unlikely(leap)) {
struct timespec64 ts;
tk->xtime_sec += leap;
ts.tv_sec = leap;
ts.tv_nsec = 0;
tk_set_wall_to_mono(tk,
timespec64_sub(tk->wall_to_monotonic, ts));
__timekeeping_set_tai_offset(tk, tk->tai_offset - leap);
clock_set = TK_CLOCK_WAS_SET;
}
}
return clock_set;
}
/*
* logarithmic_accumulation - shifted accumulation of cycles
*
* This functions accumulates a shifted interval of cycles into
* a shifted interval nanoseconds. Allows for O(log) accumulation
* loop.
*
* Returns the unconsumed cycles.
*/
static u64 logarithmic_accumulation(struct timekeeper *tk, u64 offset,
u32 shift, unsigned int *clock_set)
{
u64 interval = tk->cycle_interval << shift;
u64 snsec_per_sec;
/* If the offset is smaller than a shifted interval, do nothing */
if (offset < interval)
return offset;
/* Accumulate one shifted interval */
offset -= interval;
tk->tkr_mono.cycle_last += interval;
tk->tkr_raw.cycle_last += interval;
tk->tkr_mono.xtime_nsec += tk->xtime_interval << shift;
*clock_set |= accumulate_nsecs_to_secs(tk);
/* Accumulate raw time */
tk->tkr_raw.xtime_nsec += tk->raw_interval << shift;
snsec_per_sec = (u64)NSEC_PER_SEC << tk->tkr_raw.shift;
while (tk->tkr_raw.xtime_nsec >= snsec_per_sec) {
tk->tkr_raw.xtime_nsec -= snsec_per_sec;
tk->raw_sec++;
}
/* Accumulate error between NTP and clock interval */
tk->ntp_error += tk->ntp_tick << shift;
tk->ntp_error -= (tk->xtime_interval + tk->xtime_remainder) <<
(tk->ntp_error_shift + shift);
return offset;
}
/*
* timekeeping_advance - Updates the timekeeper to the current time and
* current NTP tick length
*/
static bool __timekeeping_advance(struct tk_data *tkd, enum timekeeping_adv_mode mode)
{
struct timekeeper *tk = &tkd->shadow_timekeeper;
struct timekeeper *real_tk = &tkd->timekeeper;
unsigned int clock_set = 0;
int shift = 0, maxshift;
u64 offset, orig_offset;
/* Make sure we're fully resumed: */
if (unlikely(timekeeping_suspended))
return false;
offset = clocksource_delta(tk_clock_read(&tk->tkr_mono),
tk->tkr_mono.cycle_last, tk->tkr_mono.mask,
tk->tkr_mono.clock->max_raw_delta);
orig_offset = offset;
/* Check if there's really nothing to do */
if (offset < real_tk->cycle_interval && mode == TK_ADV_TICK)
return false;
/*
* With NO_HZ we may have to accumulate many cycle_intervals
* (think "ticks") worth of time at once. To do this efficiently,
* we calculate the largest doubling multiple of cycle_intervals
* that is smaller than the offset. We then accumulate that
* chunk in one go, and then try to consume the next smaller
* doubled multiple.
*/
shift = ilog2(offset) - ilog2(tk->cycle_interval);
shift = max(0, shift);
/* Bound shift to one less than what overflows tick_length */
maxshift = (64 - (ilog2(ntp_tick_length(tk->id)) + 1)) - 1;
shift = min(shift, maxshift);
while (offset >= tk->cycle_interval) {
offset = logarithmic_accumulation(tk, offset, shift, &clock_set);
if (offset < tk->cycle_interval<<shift)
shift--;
}
/* Adjust the multiplier to correct NTP error */
timekeeping_adjust(tk, offset);
/*
* Finally, make sure that after the rounding
* xtime_nsec isn't larger than NSEC_PER_SEC
*/
clock_set |= accumulate_nsecs_to_secs(tk);
/*
* To avoid inconsistencies caused adjtimex TK_ADV_FREQ calls
* making small negative adjustments to the base xtime_nsec
* value, only update the coarse clocks if we accumulated time
*/
if (orig_offset != offset)
tk_update_coarse_nsecs(tk);
timekeeping_update_from_shadow(tkd, clock_set);
return !!clock_set;
}
static bool timekeeping_advance(enum timekeeping_adv_mode mode)
{
guard(raw_spinlock_irqsave)(&tk_core.lock);
return __timekeeping_advance(&tk_core, mode);
}
/**
* update_wall_time - Uses the current clocksource to increment the wall time
*
* It also updates the enabled auxiliary clock timekeepers
*/
void update_wall_time(void)
{
if (timekeeping_advance(TK_ADV_TICK))
clock_was_set_delayed();
tk_aux_advance();
}
/**
* getboottime64 - Return the real time of system boot.
* @ts: pointer to the timespec64 to be set
*
* Returns the wall-time of boot in a timespec64.
*
* This is based on the wall_to_monotonic offset and the total suspend
* time. Calls to settimeofday will affect the value returned (which
* basically means that however wrong your real time clock is at boot time,
* you get the right time here).
*/
void getboottime64(struct timespec64 *ts)
{
struct timekeeper *tk = &tk_core.timekeeper;
ktime_t t = ktime_sub(tk->offs_real, tk->offs_boot);
*ts = ktime_to_timespec64(t);
}
EXPORT_SYMBOL_GPL(getboottime64);
void ktime_get_coarse_real_ts64(struct timespec64 *ts)
{
struct timekeeper *tk = &tk_core.timekeeper;
unsigned int seq;
do {
seq = read_seqcount_begin(&tk_core.seq);
*ts = tk_xtime_coarse(tk);
} while (read_seqcount_retry(&tk_core.seq, seq));
}
EXPORT_SYMBOL(ktime_get_coarse_real_ts64);
/**
* ktime_get_coarse_real_ts64_mg - return latter of coarse grained time or floor
* @ts: timespec64 to be filled
*
* Fetch the global mg_floor value, convert it to realtime and compare it
* to the current coarse-grained time. Fill @ts with whichever is
* latest. Note that this is a filesystem-specific interface and should be
* avoided outside of that context.
*/
void ktime_get_coarse_real_ts64_mg(struct timespec64 *ts)
{
struct timekeeper *tk = &tk_core.timekeeper;
u64 floor = atomic64_read(&mg_floor);
ktime_t f_real, offset, coarse;
unsigned int seq;
do {
seq = read_seqcount_begin(&tk_core.seq);
*ts = tk_xtime_coarse(tk);
offset = tk_core.timekeeper.offs_real;
} while (read_seqcount_retry(&tk_core.seq, seq));
coarse = timespec64_to_ktime(*ts);
f_real = ktime_add(floor, offset);
if (ktime_after(f_real, coarse))
*ts = ktime_to_timespec64(f_real);
}
/**
* ktime_get_real_ts64_mg - attempt to update floor value and return result
* @ts: pointer to the timespec to be set
*
* Get a monotonic fine-grained time value and attempt to swap it into
* mg_floor. If that succeeds then accept the new floor value. If it fails
* then another task raced in during the interim time and updated the
* floor. Since any update to the floor must be later than the previous
* floor, either outcome is acceptable.
*
* Typically this will be called after calling ktime_get_coarse_real_ts64_mg(),
* and determining that the resulting coarse-grained timestamp did not effect
* a change in ctime. Any more recent floor value would effect a change to
* ctime, so there is no need to retry the atomic64_try_cmpxchg() on failure.
*
* @ts will be filled with the latest floor value, regardless of the outcome of
* the cmpxchg. Note that this is a filesystem specific interface and should be
* avoided outside of that context.
*/
void ktime_get_real_ts64_mg(struct timespec64 *ts)
{
struct timekeeper *tk = &tk_core.timekeeper;
ktime_t old = atomic64_read(&mg_floor);
ktime_t offset, mono;
unsigned int seq;
u64 nsecs;
do {
seq = read_seqcount_begin(&tk_core.seq);
ts->tv_sec = tk->xtime_sec;
mono = tk->tkr_mono.base;
nsecs = timekeeping_get_ns(&tk->tkr_mono);
offset = tk_core.timekeeper.offs_real;
} while (read_seqcount_retry(&tk_core.seq, seq));
mono = ktime_add_ns(mono, nsecs);
/*
* Attempt to update the floor with the new time value. As any
* update must be later then the existing floor, and would effect
* a change to ctime from the perspective of the current task,
* accept the resulting floor value regardless of the outcome of
* the swap.
*/
if (atomic64_try_cmpxchg(&mg_floor, &old, mono)) {
ts->tv_nsec = 0;
timespec64_add_ns(ts, nsecs);
timekeeping_inc_mg_floor_swaps();
} else {
/*
* Another task changed mg_floor since "old" was fetched.
* "old" has been updated with the latest value of "mg_floor".
* That value is newer than the previous floor value, which
* is enough to effect a change to ctime. Accept it.
*/
*ts = ktime_to_timespec64(ktime_add(old, offset));
}
}
void ktime_get_coarse_ts64(struct timespec64 *ts)
{
struct timekeeper *tk = &tk_core.timekeeper;
struct timespec64 now, mono;
unsigned int seq;
do {
seq = read_seqcount_begin(&tk_core.seq);
now = tk_xtime_coarse(tk);
mono = tk->wall_to_monotonic;
} while (read_seqcount_retry(&tk_core.seq, seq));
set_normalized_timespec64(ts, now.tv_sec + mono.tv_sec,
now.tv_nsec + mono.tv_nsec);
}
EXPORT_SYMBOL(ktime_get_coarse_ts64);
/*
* Must hold jiffies_lock
*/
void do_timer(unsigned long ticks)
{
jiffies_64 += ticks;
calc_global_load();
}
/**
* ktime_get_update_offsets_now - hrtimer helper
* @cwsseq: pointer to check and store the clock was set sequence number
* @offs_real: pointer to storage for monotonic -> realtime offset
* @offs_boot: pointer to storage for monotonic -> boottime offset
* @offs_tai: pointer to storage for monotonic -> clock tai offset
*
* Returns current monotonic time and updates the offsets if the
* sequence number in @cwsseq and timekeeper.clock_was_set_seq are
* different.
*
* Called from hrtimer_interrupt() or retrigger_next_event()
*/
ktime_t ktime_get_update_offsets_now(unsigned int *cwsseq, ktime_t *offs_real,
ktime_t *offs_boot, ktime_t *offs_tai)
{
struct timekeeper *tk = &tk_core.timekeeper;
unsigned int seq;
ktime_t base;
u64 nsecs;
do {
seq = read_seqcount_begin(&tk_core.seq);
base = tk->tkr_mono.base;
nsecs = timekeeping_get_ns(&tk->tkr_mono);
base = ktime_add_ns(base, nsecs);
if (*cwsseq != tk->clock_was_set_seq) {
*cwsseq = tk->clock_was_set_seq;
*offs_real = tk->offs_real;
*offs_boot = tk->offs_boot;
*offs_tai = tk->offs_tai;
}
/* Handle leapsecond insertion adjustments */
if (unlikely(base >= tk->next_leap_ktime))
*offs_real = ktime_sub(tk->offs_real, ktime_set(1, 0));
} while (read_seqcount_retry(&tk_core.seq, seq));
return base;
}
/*
* timekeeping_validate_timex - Ensures the timex is ok for use in do_adjtimex
*/
static int timekeeping_validate_timex(const struct __kernel_timex *txc, bool aux_clock)
{
if (txc->modes & ADJ_ADJTIME) {
/* singleshot must not be used with any other mode bits */
if (!(txc->modes & ADJ_OFFSET_SINGLESHOT))
return -EINVAL;
if (!(txc->modes & ADJ_OFFSET_READONLY) &&
!capable(CAP_SYS_TIME))
return -EPERM;
} else {
/* In order to modify anything, you gotta be super-user! */
if (txc->modes && !capable(CAP_SYS_TIME))
return -EPERM;
/*
* if the quartz is off by more than 10% then
* something is VERY wrong!
*/
if (txc->modes & ADJ_TICK &&
(txc->tick < 900000/USER_HZ ||
txc->tick > 1100000/USER_HZ))
return -EINVAL;
}
if (txc->modes & ADJ_SETOFFSET) {
/* In order to inject time, you gotta be super-user! */
if (!capable(CAP_SYS_TIME))
return -EPERM;
/*
* Validate if a timespec/timeval used to inject a time
* offset is valid. Offsets can be positive or negative, so
* we don't check tv_sec. The value of the timeval/timespec
* is the sum of its fields,but *NOTE*:
* The field tv_usec/tv_nsec must always be non-negative and
* we can't have more nanoseconds/microseconds than a second.
*/
if (txc->time.tv_usec < 0)
return -EINVAL;
if (txc->modes & ADJ_NANO) {
if (txc->time.tv_usec >= NSEC_PER_SEC)
return -EINVAL;
} else {
if (txc->time.tv_usec >= USEC_PER_SEC)
return -EINVAL;
}
}
/*
* Check for potential multiplication overflows that can
* only happen on 64-bit systems:
*/
if ((txc->modes & ADJ_FREQUENCY) && (BITS_PER_LONG == 64)) {
if (LLONG_MIN / PPM_SCALE > txc->freq)
return -EINVAL;
if (LLONG_MAX / PPM_SCALE < txc->freq)
return -EINVAL;
}
if (aux_clock) {
/* Auxiliary clocks are similar to TAI and do not have leap seconds */
if (txc->modes & ADJ_STATUS &&
txc->status & (STA_INS | STA_DEL))
return -EINVAL;
/* No TAI offset setting */
if (txc->modes & ADJ_TAI)
return -EINVAL;
/* No PPS support either */
if (txc->modes & ADJ_STATUS &&
txc->status & (STA_PPSFREQ | STA_PPSTIME))
return -EINVAL;
}
return 0;
}
/**
* random_get_entropy_fallback - Returns the raw clock source value,
* used by random.c for platforms with no valid random_get_entropy().
*/
unsigned long random_get_entropy_fallback(void)
{
struct tk_read_base *tkr = &tk_core.timekeeper.tkr_mono;
struct clocksource *clock = READ_ONCE(tkr->clock);
if (unlikely(timekeeping_suspended || !clock))
return 0;
return clock->read(clock);
}
EXPORT_SYMBOL_GPL(random_get_entropy_fallback);
struct adjtimex_result {
struct audit_ntp_data ad;
struct timespec64 delta;
bool clock_set;
};
static int __do_adjtimex(struct tk_data *tkd, struct __kernel_timex *txc,
struct adjtimex_result *result)
{
struct timekeeper *tks = &tkd->shadow_timekeeper;
bool aux_clock = !timekeeper_is_core_tk(tks);
struct timespec64 ts;
s32 orig_tai, tai;
int ret;
/* Validate the data before disabling interrupts */
ret = timekeeping_validate_timex(txc, aux_clock);
if (ret)
return ret;
add_device_randomness(txc, sizeof(*txc));
if (!aux_clock)
ktime_get_real_ts64(&ts);
else
tk_get_aux_ts64(tkd->timekeeper.id, &ts);
add_device_randomness(&ts, sizeof(ts));
guard(raw_spinlock_irqsave)(&tkd->lock);
if (!tks->clock_valid)
return -ENODEV;
if (txc->modes & ADJ_SETOFFSET) {
result->delta.tv_sec = txc->time.tv_sec;
result->delta.tv_nsec = txc->time.tv_usec;
if (!(txc->modes & ADJ_NANO))
result->delta.tv_nsec *= 1000;
ret = __timekeeping_inject_offset(tkd, &result->delta);
if (ret)
return ret;
result->clock_set = true;
}
orig_tai = tai = tks->tai_offset;
ret = ntp_adjtimex(tks->id, txc, &ts, &tai, &result->ad);
if (tai != orig_tai) {
__timekeeping_set_tai_offset(tks, tai);
timekeeping_update_from_shadow(tkd, TK_CLOCK_WAS_SET);
result->clock_set = true;
} else {
tk_update_leap_state_all(tkd);
}
/* Update the multiplier immediately if frequency was set directly */
if (txc->modes & (ADJ_FREQUENCY | ADJ_TICK))
result->clock_set |= __timekeeping_advance(tkd, TK_ADV_FREQ);
return ret;
}
/**
* do_adjtimex() - Accessor function to NTP __do_adjtimex function
* @txc: Pointer to kernel_timex structure containing NTP parameters
*/
int do_adjtimex(struct __kernel_timex *txc)
{
struct adjtimex_result result = { };
int ret;
ret = __do_adjtimex(&tk_core, txc, &result);
if (ret < 0)
return ret;
if (txc->modes & ADJ_SETOFFSET)
audit_tk_injoffset(result.delta);
audit_ntp_log(&result.ad);
if (result.clock_set)
clock_was_set(CLOCK_SET_WALL);
ntp_notify_cmos_timer(result.delta.tv_sec != 0);
return ret;
}
/*
* Invoked from NTP with the time keeper lock held, so lockless access is
* fine.
*/
long ktime_get_ntp_seconds(unsigned int id)
{
return timekeeper_data[id].timekeeper.xtime_sec;
}
#ifdef CONFIG_NTP_PPS
/**
* hardpps() - Accessor function to NTP __hardpps function
* @phase_ts: Pointer to timespec64 structure representing phase timestamp
* @raw_ts: Pointer to timespec64 structure representing raw timestamp
*/
void hardpps(const struct timespec64 *phase_ts, const struct timespec64 *raw_ts)
{
guard(raw_spinlock_irqsave)(&tk_core.lock);
__hardpps(phase_ts, raw_ts);
}
EXPORT_SYMBOL(hardpps);
#endif /* CONFIG_NTP_PPS */
#ifdef CONFIG_POSIX_AUX_CLOCKS
#include "posix-timers.h"
/*
* Bitmap for the activated auxiliary timekeepers to allow lockless quick
* checks in the hot paths without touching extra cache lines. If set, then
* the state of the corresponding timekeeper has to be re-checked under
* timekeeper::lock.
*/
static unsigned long aux_timekeepers;
static inline unsigned int clockid_to_tkid(unsigned int id)
{
return TIMEKEEPER_AUX_FIRST + id - CLOCK_AUX;
}
static inline struct tk_data *aux_get_tk_data(clockid_t id)
{
if (!clockid_aux_valid(id))
return NULL;
return &timekeeper_data[clockid_to_tkid(id)];
}
/* Invoked from timekeeping after a clocksource change */
static void tk_aux_update_clocksource(void)
{
unsigned long active = READ_ONCE(aux_timekeepers);
unsigned int id;
for_each_set_bit(id, &active, BITS_PER_LONG) {
struct tk_data *tkd = &timekeeper_data[id + TIMEKEEPER_AUX_FIRST];
struct timekeeper *tks = &tkd->shadow_timekeeper;
guard(raw_spinlock_irqsave)(&tkd->lock);
if (!tks->clock_valid)
continue;
timekeeping_forward_now(tks);
tk_setup_internals(tks, tk_core.timekeeper.tkr_raw.clock);
timekeeping_update_from_shadow(tkd, TK_UPDATE_ALL);
}
}
static void tk_aux_advance(void)
{
unsigned long active = READ_ONCE(aux_timekeepers);
unsigned int id;
/* Lockless quick check to avoid extra cache lines */
for_each_set_bit(id, &active, BITS_PER_LONG) {
struct tk_data *aux_tkd = &timekeeper_data[id + TIMEKEEPER_AUX_FIRST];
guard(raw_spinlock)(&aux_tkd->lock);
if (aux_tkd->shadow_timekeeper.clock_valid)
__timekeeping_advance(aux_tkd, TK_ADV_TICK);
}
}
/**
* ktime_get_aux - Get time for a AUX clock
* @id: ID of the clock to read (CLOCK_AUX...)
* @kt: Pointer to ktime_t to store the time stamp
*
* Returns: True if the timestamp is valid, false otherwise
*/
bool ktime_get_aux(clockid_t id, ktime_t *kt)
{
struct tk_data *aux_tkd = aux_get_tk_data(id);
struct timekeeper *aux_tk;
unsigned int seq;
ktime_t base;
u64 nsecs;
WARN_ON(timekeeping_suspended);
if (!aux_tkd)
return false;
aux_tk = &aux_tkd->timekeeper;
do {
seq = read_seqcount_begin(&aux_tkd->seq);
if (!aux_tk->clock_valid)
return false;
base = ktime_add(aux_tk->tkr_mono.base, aux_tk->offs_aux);
nsecs = timekeeping_get_ns(&aux_tk->tkr_mono);
} while (read_seqcount_retry(&aux_tkd->seq, seq));
*kt = ktime_add_ns(base, nsecs);
return true;
}
EXPORT_SYMBOL_GPL(ktime_get_aux);
/**
* ktime_get_aux_ts64 - Get time for a AUX clock
* @id: ID of the clock to read (CLOCK_AUX...)
* @ts: Pointer to timespec64 to store the time stamp
*
* Returns: True if the timestamp is valid, false otherwise
*/
bool ktime_get_aux_ts64(clockid_t id, struct timespec64 *ts)
{
ktime_t now;
if (!ktime_get_aux(id, &now))
return false;
*ts = ktime_to_timespec64(now);
return true;
}
EXPORT_SYMBOL_GPL(ktime_get_aux_ts64);
static int aux_get_res(clockid_t id, struct timespec64 *tp)
{
if (!clockid_aux_valid(id))
return -ENODEV;
tp->tv_sec = aux_clock_resolution_ns() / NSEC_PER_SEC;
tp->tv_nsec = aux_clock_resolution_ns() % NSEC_PER_SEC;
return 0;
}
static int aux_get_timespec(clockid_t id, struct timespec64 *tp)
{
return ktime_get_aux_ts64(id, tp) ? 0 : -ENODEV;
}
static int aux_clock_set(const clockid_t id, const struct timespec64 *tnew)
{
struct tk_data *aux_tkd = aux_get_tk_data(id);
struct timekeeper *aux_tks;
ktime_t tnow, nsecs;
if (!timespec64_valid_settod(tnew))
return -EINVAL;
if (!aux_tkd)
return -ENODEV;
aux_tks = &aux_tkd->shadow_timekeeper;
guard(raw_spinlock_irq)(&aux_tkd->lock);
if (!aux_tks->clock_valid)
return -ENODEV;
/* Forward the timekeeper base time */
timekeeping_forward_now(aux_tks);
/*
* Get the updated base time. tkr_mono.base has not been
* updated yet, so do that first. That makes the update
* in timekeeping_update_from_shadow() redundant, but
* that's harmless. After that @tnow can be calculated
* by using tkr_mono::cycle_last, which has been set
* by timekeeping_forward_now().
*/
tk_update_ktime_data(aux_tks);
nsecs = timekeeping_cycles_to_ns(&aux_tks->tkr_mono, aux_tks->tkr_mono.cycle_last);
tnow = ktime_add(aux_tks->tkr_mono.base, nsecs);
/*
* Calculate the new AUX offset as delta to @tnow ("monotonic").
* That avoids all the tk::xtime back and forth conversions as
* xtime ("realtime") is not applicable for auxiliary clocks and
* kept in sync with "monotonic".
*/
tk_update_aux_offs(aux_tks, ktime_sub(timespec64_to_ktime(*tnew), tnow));
timekeeping_update_from_shadow(aux_tkd, TK_UPDATE_ALL);
return 0;
}
static int aux_clock_adj(const clockid_t id, struct __kernel_timex *txc)
{
struct tk_data *aux_tkd = aux_get_tk_data(id);
struct adjtimex_result result = { };
if (!aux_tkd)
return -ENODEV;
/*
* @result is ignored for now as there are neither hrtimers nor a
* RTC related to auxiliary clocks for now.
*/
return __do_adjtimex(aux_tkd, txc, &result);
}
const struct k_clock clock_aux = {
.clock_getres = aux_get_res,
.clock_get_timespec = aux_get_timespec,
.clock_set = aux_clock_set,
.clock_adj = aux_clock_adj,
};
static void aux_clock_enable(clockid_t id)
{
struct tk_read_base *tkr_raw = &tk_core.timekeeper.tkr_raw;
struct tk_data *aux_tkd = aux_get_tk_data(id);
struct timekeeper *aux_tks = &aux_tkd->shadow_timekeeper;
/* Prevent the core timekeeper from changing. */
guard(raw_spinlock_irq)(&tk_core.lock);
/*
* Setup the auxiliary clock assuming that the raw core timekeeper
* clock frequency conversion is close enough. Userspace has to
* adjust for the deviation via clock_adjtime(2).
*/
guard(raw_spinlock_nested)(&aux_tkd->lock);
/* Remove leftovers of a previous registration */
memset(aux_tks, 0, sizeof(*aux_tks));
/* Restore the timekeeper id */
aux_tks->id = aux_tkd->timekeeper.id;
/* Setup the timekeeper based on the current system clocksource */
tk_setup_internals(aux_tks, tkr_raw->clock);
/* Mark it valid and set it live */
aux_tks->clock_valid = true;
timekeeping_update_from_shadow(aux_tkd, TK_UPDATE_ALL);
}
static void aux_clock_disable(clockid_t id)
{
struct tk_data *aux_tkd = aux_get_tk_data(id);
guard(raw_spinlock_irq)(&aux_tkd->lock);
aux_tkd->shadow_timekeeper.clock_valid = false;
timekeeping_update_from_shadow(aux_tkd, TK_UPDATE_ALL);
}
static DEFINE_MUTEX(aux_clock_mutex);
static ssize_t aux_clock_enable_store(struct kobject *kobj, struct kobj_attribute *attr,
const char *buf, size_t count)
{
/* Lazy atoi() as name is "0..7" */
int id = kobj->name[0] & 0x7;
bool enable;
if (!capable(CAP_SYS_TIME))
return -EPERM;
if (kstrtobool(buf, &enable) < 0)
return -EINVAL;
guard(mutex)(&aux_clock_mutex);
if (enable == test_bit(id, &aux_timekeepers))
return count;
if (enable) {
aux_clock_enable(CLOCK_AUX + id);
set_bit(id, &aux_timekeepers);
} else {
aux_clock_disable(CLOCK_AUX + id);
clear_bit(id, &aux_timekeepers);
}
return count;
}
static ssize_t aux_clock_enable_show(struct kobject *kobj, struct kobj_attribute *attr, char *buf)
{
unsigned long active = READ_ONCE(aux_timekeepers);
/* Lazy atoi() as name is "0..7" */
int id = kobj->name[0] & 0x7;
return sysfs_emit(buf, "%d\n", test_bit(id, &active));
}
static struct kobj_attribute aux_clock_enable_attr = __ATTR_RW(aux_clock_enable);
static struct attribute *aux_clock_enable_attrs[] = {
&aux_clock_enable_attr.attr,
NULL
};
static const struct attribute_group aux_clock_enable_attr_group = {
.attrs = aux_clock_enable_attrs,
};
static int __init tk_aux_sysfs_init(void)
{
struct kobject *auxo, *tko = kobject_create_and_add("time", kernel_kobj);
int ret = -ENOMEM;
if (!tko)
return ret;
auxo = kobject_create_and_add("aux_clocks", tko);
if (!auxo)
goto err_clean;
for (int i = 0; i < MAX_AUX_CLOCKS; i++) {
char id[2] = { [0] = '0' + i, };
struct kobject *clk = kobject_create_and_add(id, auxo);
if (!clk) {
ret = -ENOMEM;
goto err_clean;
}
ret = sysfs_create_group(clk, &aux_clock_enable_attr_group);
if (ret)
goto err_clean;
}
return 0;
err_clean:
kobject_put(auxo);
kobject_put(tko);
return ret;
}
late_initcall(tk_aux_sysfs_init);
static __init void tk_aux_setup(void)
{
for (int i = TIMEKEEPER_AUX_FIRST; i <= TIMEKEEPER_AUX_LAST; i++)
tkd_basic_setup(&timekeeper_data[i], i, false);
}
#endif /* CONFIG_POSIX_AUX_CLOCKS */
]
|
| 51/2 |
2026/04/21 11:32 |
llm |
expert |
0m
Model:
Tokens:
input: 121656
output: 691
thoughts: 3407
Thoughts: **Analysis of Data Race in ktime_get_coarse_real_ts64_mg**
Okay, I've dived into this `ktime_get_coarse_real_ts64_mg` function and, let me tell you, it's pretty straightforward once you break it down. I see it's reading `tk_core.timekeeper.offs_real` and accessing `tk->xtime_sec` and `tk->coarse_nsec`. The key here is that `tk` points to `&tk_core.timekeeper`, so accessing it the same. I'm immediately noticing the similarities between `ktime_get_coarse_real_ts64_mg` and the non-mg version, `ktime_get_coarse_real_ts64`. Both read `tk->xtime_sec` and `tk->coarse_nsec`.
The crucial detail, of course, is that these reads aren't protected by `READ_ONCE` or anything like that. However, and this is where it gets interesting, these reads *are* inside a `read_seqcount_begin` and `read_seqcount_retry` loop. That's my immediate thought: seqlock! The data is consistent.
I'm thinking KCSAN has flagged this as a potential data race. My knee-jerk reaction? It's probably benign because of the seqlock. Even if there's a "race," the `read_seqcount_retry` mechanism *should* catch any inconsistencies, and the reader will simply retry until it gets a consistent view. However, KCSAN *is* complaining, and I know why, because the compiler *could* do something wonky, and reorder the reads.
But looking at it again, the barriers of `smp_load_acquire` and `smp_rmb()` in the seqcount are designed to prevent reordering outside of the loop, and the seqlock is designed to deal with the compiler reordering *inside* the loop. The race is **BENIGN**.
The KCSAN report has triggered an analysis, though, and it's time to check what's being updated. In this case, `xtime_sec` and `coarse_nsec`. If those reads were for a pointer, and it caused a crash, that would be bad. The data is consistent.
I see `ktime_get_coarse_real_ts64_mg` is specific to the "multigrain" timestamps used by filesystems. It's comparing the "coarse" time with a floor value to ensure that the coarse time never goes backwards. The `offset` (`tk_core.timekeeper.offs_real`) is read inside the `seqcount` loop and looks correct. This looks consistent with other functions, like `ktime_get_real_ts64`. The same pattern exists, without the protection of `READ_ONCE`. KCSAN must be getting lucky.
I've also checked for torn reads. `coarse_nsec` is 4 bytes and `xtime_sec` is 8 bytes, both aligned for atomic reads. There's no tearing, but even if there was the seqlock would handle it. There are barriers protecting this. Also, the compiler cannot optimize away the loop.
Okay, I've got a read and write race, but it's fine.
The write side is `timekeeping_update_from_shadow`. It looks like this is a classic producer-consumer scenario, with an interrupt calling `timekeeping_update_from_shadow` to update timekeeper variables, and our function reading them. It's a race, but a benign one! The race is between updates from interrupts, and reads from tasks (filesystems).
I see the `write` stack trace has `__tsan_read4` in it. This confirms my idea. `__tsan_read4` does not signify a data race.
I still conclude that this race is **BENIGN**. A key observation is `coarse_nsec` being updated by one tick, which shows that updates happen from the timer interrupt. I am confident that my conclusion is correct.
I am concluding that `data_race()` would be useful, or adding `READ_ONCE`/`WRITE_ONCE()` to `tk->xtime_sec` and `tk->coarse_nsec`.
I am checking if `memcpy` is causing an issue. No.
I'm checking that the function is for multigrain. It is.
I am satisfied that my analysis is correct. The race is in a seqlock and is benign. The race is on the timekeeper.
The data race occurs between `ktime_get_coarse_real_ts64_mg()` (the reader) and `timekeeping_update_from_shadow()` (the writer) on the `coarse_nsec` field of the global `timekeeper` structure.
### Analysis
1. **Reader Context:** `ktime_get_coarse_real_ts64_mg()` is a lockless reader used in the filesystem layer (e.g., `current_time()`) to fetch a coarse-grained timestamp. It uses a sequence counter (`tk_core.seq`) to ensure consistency:
```c
do {
seq = read_seqcount_begin(&tk_core.seq);
*ts = tk_xtime_coarse(tk); // Reads tk->xtime_sec and tk->coarse_nsec
offset = tk_core.timekeeper.offs_real;
} while (read_seqcount_retry(&tk_core.seq, seq));
```
2. **Writer Context:** `timekeeping_update_from_shadow()` is called during the timer interrupt (via `update_wall_time()`) to update the main timekeeper from a shadow copy. It also uses the sequence counter:
```c
write_seqcount_begin(&tkd->seq);
...
memcpy(&tkd->timekeeper, tk, sizeof(*tk)); // Overwrites the whole struct
write_seqcount_end(&tkd->seq);
```
3. **The Race:** The reader performs plain loads of `tk->coarse_nsec` (and other fields) while the writer performs a `memcpy()` to the same memory location. Since these are not marked as atomic or wrapped in `data_race()`, KCSAN detects a data race.
4. **Safety:** The race is structurally benign because it is protected by the `seqcount` mechanism. If the reader observes a torn or stale value due to the concurrent `memcpy()`, the `read_seqcount_retry()` check will fail because the sequence number will have changed (or will be odd during the write). The reader will then retry the operation until it gets a consistent snapshot. The values read are used to populate a `struct timespec64` and do not influence critical control flow or pointer dereferences before the retry check.
### Final Classification: **BENIGN**
The logic is sound and structurally tolerant to stale or torn reads due to the use of sequence counters. The report is a result of missing annotations for KCSAN.
### Recommended Annotations
The accesses in `tk_xtime_coarse()` and `ktime_get_coarse_real_ts64_mg()` should be wrapped in `data_race()` to signal to KCSAN that these concurrent accesses are intentional and handled by the retry loop. Alternatively, `READ_ONCE()` could be used for individual fields, although `data_race()` is often preferred for seqlock-protected regions where the compiler is already prevented from moving code out of the loop by the seqlock barriers.
|
| 52/2 |
2026/04/21 11:33 |
tool |
set-results |
0m
Args: map[Benign:true Confident:true]
Results: map[Benign:true Confident:true]
|
| 53/2 |
2026/04/21 11:33 |
llm |
expert |
0m
Model:
Tokens:
input: 126791
output: 685
thoughts: 41
Thoughts: Here's my thought process on this task:
**Reflecting on the Summarization Task**
Okay, so the user wants a summary of the analysis, specifically presented as *my* thoughts. That's straightforward enough. I already did the heavy lifting – the classification, the reasoning, the justification – in previous stages. Now, my focus is purely on crafting the *presentation* of the final response. My task here is to present the previous analysis in a manner that sounds like my own internal monologue, making it relevant and readable for a user, especially an expert. The user wants the response formatted, so I will take care of the specific formatting requirements. This is a crucial step to ensure the result matches the users expectations.
|
| 54/1 |
2026/04/21 11:33 |
action |
format-explanation |
0m
Results: map[Explanation:The data race occurs between **`ktime_get_coarse_real_ts64_mg()`** and
**`timekeeping_update_from_shadow()`** on the field **`tk->coarse_nsec`** (part
of `struct timekeeper`).
### Synchronization Context
* **Reader:** `ktime_get_coarse_real_ts64_mg()` is a lockless reader that uses
a sequence counter (`tk_core.seq`) via `read_seqcount_begin()` and
`read_seqcount_retry()` to ensure data consistency.
* **Writer:** `timekeeping_update_from_shadow()` is called in interrupt
context (timer tick) and holds the `tk_core.lock` spinlock. It also increments
the sequence counter via `write_seqcount_begin()` and `write_seqcount_end()`
around a `memcpy()` that updates the live timekeeper.
The race is possible because the reader does not hold the spinlock, and both
sides use plain memory accesses (or `memcpy`) for the data itself.
### Final Classification: **BENIGN**
### Reasoning
The race is structurally benign because the algorithm is designed to be tolerant
of stale or torn reads through the use of a **sequence counter (seqlock)**.
1. **Retry Mechanism:** If the reader (`CPU 0`) reads `coarse_nsec` while the
writer (`CPU 1`) is executing the `memcpy()`, the reader might see a partially
updated (torn) value. However, the `read_seqcount_retry()` call at the end of
the loop will detect that a write was in progress (sequence number was odd) or
has completed (sequence number changed), forcing the reader to discard the
result and retry the loop.
2. **Data Safety:** The racing variable is a simple integer representing
nanoseconds. A torn read of this value does not lead to a crash, invalid pointer
dereference, or out-of-bounds access. It is only used to populate a `timespec64`
structure, which is then validated against a "floor" value and potentially
discarded if the retry fails.
3. **Ordering:** The `read_seqcount_begin()` and `read_seqcount_retry()`
primitives include the necessary memory barriers (`smp_load_acquire` and
`smp_rmb`) to prevent the compiler from hoisting the reads out of the protected
loop.
### Recommended Annotations
To silence the KCSAN report, the plain accesses inside the `seqcount` loop
should be wrapped in `data_race()` or converted to `READ_ONCE()`. Specifically,
in `tk_xtime_coarse()`:
```c
static inline struct timespec64 tk_xtime_coarse(const struct timekeeper *tk)
{
struct timespec64 ts;
ts.tv_sec = data_race(tk->xtime_sec);
ts.tv_nsec = data_race(tk->coarse_nsec);
return ts;
}
```
And similarly for the `offs_real` access in `ktime_get_coarse_real_ts64_mg()`.]
|