Skip to content

fix(virtqueue): order descriptor publication against the notification check - #2669

Merged
mkroening merged 1 commit into
hermit-os:mainfrom
stlankes:queue
Aug 21, 2026
Merged

fix(virtqueue): order descriptor publication against the notification check#2669
mkroening merged 1 commit into
hermit-os:mainfrom
stlankes:queue

Conversation

@stlankes

Copy link
Copy Markdown
Contributor

Both rings decide whether to notify the device by reading state the device writes, right after publishing a descriptor, with nothing ordering the two. The preceding barrier orders the descriptor contents against the field that publishes them, not that field against this later load, so a weakly ordered machine can read the flag before our publication is visible.

This causes deadlocks in polling mode, which are prevented by adding memory barriers.

… check

Both rings decide whether to notify the device by reading state the device
writes, right after publishing a descriptor, with nothing ordering the two.
The preceding barrier orders the descriptor contents against the field that
publishes them, not that field against this later load, so a weakly ordered
machine can read the flag before our publication is visible.

This causes deadlocks in polling mode, which are prevented by adding memory
barriers.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Benchmark Results

Details
Benchmark Current: 7184aa0 Previous: 2e23902 Performance Ratio
startup_benchmark Build Time 94.18 s 80.34 s 1.17
startup_benchmark File Size 0.78 MB 0.80 MB 0.98
Startup Time - 1 core 0.75 s (±0.02 s) 0.75 s (±0.02 s) 1.00
Startup Time - 2 cores 0.75 s (±0.02 s) 0.74 s (±0.02 s) 1.02
Startup Time - 4 cores 0.77 s (±0.01 s) 0.74 s (±0.02 s) 1.03
multithreaded_benchmark Build Time 91.99 s 82.11 s 1.12
multithreaded_benchmark File Size 0.88 MB 0.86 MB 1.03
Multithreaded Pi Efficiency - 2 Threads 65.57 % (±7.98 %) 85.89 % (±6.61 %) 0.76
Multithreaded Pi Efficiency - 4 Threads 41.56 % (±2.91 %) 43.43 % (±2.56 %) 0.96
Multithreaded Pi Efficiency - 8 Threads 20.00 % (±1.57 %) 25.76 % (±1.53 %) 0.78
micro_benchmarks Build Time 216.18 s 80.40 s 2.69
micro_benchmarks File Size 0.88 MB 0.86 MB 1.03
Scheduling time - 1 thread 197.73 ticks (±50.00 ticks) 62.65 ticks (±4.06 ticks) 3.16
Scheduling time - 2 threads 107.44 ticks (±32.65 ticks) 34.08 ticks (±4.10 ticks) 3.15
Micro - Time for syscall (getpid) 9.36 ticks (±4.41 ticks) 3.45 ticks (±0.58 ticks) 2.71
Memcpy speed - (built_in) block size 4096 56802.00 MByte/s (±40001.94 MByte/s) 82448.38 MByte/s (±56997.13 MByte/s) 0.69
Memcpy speed - (built_in) block size 1048576 14874.16 MByte/s (±12761.27 MByte/s) 30585.98 MByte/s (±24707.84 MByte/s) 0.49
Memcpy speed - (built_in) block size 16777216 12680.41 MByte/s (±10804.21 MByte/s) 26340.06 MByte/s (±21720.96 MByte/s) 0.48
Memset speed - (built_in) block size 4096 57037.56 MByte/s (±40149.36 MByte/s) 82292.76 MByte/s (±56891.50 MByte/s) 0.69
Memset speed - (built_in) block size 1048576 15231.27 MByte/s (±12965.56 MByte/s) 31323.85 MByte/s (±25145.86 MByte/s) 0.49
Memset speed - (built_in) block size 16777216 13106.68 MByte/s (±11094.44 MByte/s) 27104.68 MByte/s (±22209.94 MByte/s) 0.48
Memcpy speed - (rust) block size 4096 54882.81 MByte/s (±39121.98 MByte/s) 74097.96 MByte/s (±51811.44 MByte/s) 0.74
Memcpy speed - (rust) block size 1048576 14427.31 MByte/s (±12693.24 MByte/s) 30361.60 MByte/s (±24602.37 MByte/s) 0.48
Memcpy speed - (rust) block size 16777216 13138.89 MByte/s (±11445.15 MByte/s) 27625.34 MByte/s (±22806.88 MByte/s) 0.48
Memset speed - (rust) block size 4096 55373.89 MByte/s (±39464.05 MByte/s) 74373.47 MByte/s (±51976.48 MByte/s) 0.74
Memset speed - (rust) block size 1048576 14726.38 MByte/s (±12831.83 MByte/s) 31110.89 MByte/s (±25033.24 MByte/s) 0.47
Memset speed - (rust) block size 16777216 13410.92 MByte/s (±11572.30 MByte/s) 28386.93 MByte/s (±23265.03 MByte/s) 0.47
alloc_benchmarks Build Time 213.34 s 74.76 s 2.85
alloc_benchmarks File Size 0.86 MB 0.87 MB 0.98
Allocations - Allocation success 91.38 % 91.31 % 1.00
Allocations - Deallocation success 100.00 % 100.00 % 1
Allocations - Pre-fail Allocations 61.60 % 61.44 % 1.00
Allocations - Average Allocation time 23456.36 Ticks (±1601.07 Ticks) 5860.58 Ticks (±98.43 Ticks) 4.00
Allocations - Average Allocation time (no fail) 24362.93 Ticks (±2025.26 Ticks) 6554.81 Ticks (±92.86 Ticks) 3.72
Allocations - Average Deallocation time 6864.14 Ticks (±1641.54 Ticks) 1805.01 Ticks (±250.35 Ticks) 3.80
mutex_benchmark Build Time 222.70 s 79.82 s 2.79
mutex_benchmark File Size 0.88 MB 0.86 MB 1.03
Mutex Stress Test Average Time per Iteration - 1 Threads 36.84 ns (±6.97 ns) 12.10 ns (±0.41 ns) 3.04
Mutex Stress Test Average Time per Iteration - 2 Threads 32.94 ns (±8.80 ns) 40.26 ns (±1.68 ns) 0.82

This comment was automatically generated by workflow using github-action-benchmark.

@mkroening mkroening left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me! 👍

@mkroening
mkroening added this pull request to the merge queue Aug 21, 2026
Merged via the queue into hermit-os:main with commit 1ae5e44 Aug 21, 2026
41 of 42 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants