#1149313 mesa-libgallium: radeonsi deadlock (100% CPU) in util_queue_finish() on context destroy, fixed upstream by d1a6be75db94

Package:
mesa-libgallium
Source:
mesa-libgallium
Description:
shared infrastructure for Mesa drivers
Submitter:
Romain Reignier
Date:
2026-09-29 11:51:01 UTC
Severity:
normal
Tags:
#1149313#5
Date:
2026-09-29 11:13:11 UTC
From:
To:
Dear Maintainer,

On trixie, plasmashell (Plasma 6.3.6 Wayland) hung for good with its
main thread at 100% CPU, after I unplugged a USB-C monitor and unlocked
the session. The panel froze and plasmashell stopped answering D-Bus
(kquitapp6 timed out). The only way out was to kill and restart it.

Hardware: ThinkPad T14, AMD Radeon 840M/860M (amdgpu, GFX1152),
kernel 7.1.8 (trixie-backports), Qt 6.8.2+dfsg-9+deb13u2.

This is the upstream Mesa race in util_queue_finish(), fixed by:

  commit d1a6be75db949477b45748c2e93a7199cadec212
  util/u_queue: Fix data race on num_threads during finish.
https://gitlab.freedesktop.org/mesa/mesa/-/commit/d1a6be75db949477b45748c2e93a7199cadec212
  (Closes upstream #13738, MR !36919; first released in 25.3.0,
  not backported to any 25.0.x/25.2.x stable release)

The problem
-----------
util_queue_finish() (src/util/u_queue.c) holds the queue lock while it
allocates one fence per worker thread and sizes a barrier from
queue->num_threads. It then sets create_threads_on_demand = true again
and unlocks. After that, it re-reads queue->num_threads without the
lock for its wait loop. If another thread submits a job in between,
util_queue_add_job() can start an extra worker thread. The wait loop
then reads one fence past the end of the malloc'd array. That memory
is never 0, so do_futex_fence_wait() spins forever
(futex(FUTEX_WAIT_BITSET, 2) returns EAGAIN each time, hence the busy
loop).

Evidence from a core dump of the hung process
(with mesa-libgallium-dbgsym 25.0.7-2+deb13u1)
-----------------------------------------------------------------------
Main thread:

  #1  sys_futex (addr1=0x5628b02b1068, op=9, val1=2, ...) at ../src/util/futex.c:43
  #2  futex_wait (addr=0x5628b02b1068, value=2, timeout=0x0) at ../src/util/futex.c:55
  #3  do_futex_fence_wait (fence=0x5628b02b1068, timeout=false, ...) at ../src/util/u_queue.c:131
  #4  _util_queue_fence_wait (fence=0x5628b02b1068) at ../src/util/u_queue.c:146
  #7  util_queue_finish (queue=0x5628a888c968) at ../src/util/u_queue.c:722
          i = 6
          fences = 0x5628b02b1050
  #8  si_set_debug_callback (ctx=0x5628af84d720, cb=0x0) at ../src/gallium/drivers/radeonsi/si_pipe.c:447
  #9  si_destroy_context (context=0x5628af84d720) at ../src/gallium/drivers/radeonsi/si_pipe.c:197
  #10 tc_destroy (_pipe=0x5628afd04bb0) at ../src/gallium/auxiliary/util/u_threaded_context.c:5183
  #11 st_destroy_context_priv (st=0x5628b1303150, destroy_pipe=true) at ../src/mesa/state_tracker/st_context.c:358
  #12 st_destroy_context (st=0x5628b1303150) at ../src/mesa/state_tracker/st_context.c:976
  #13 dri_destroy_context (ctx=0x5628b033fb50) at ../src/gallium/frontends/dri/dri_context.c:280
  #14 driDestroyContext (ctx=<optimized out>) at ../src/gallium/frontends/dri/dri_util.c:639
  #15 dri2_destroy_context (disp=<optimized out>, ctx=0x5628af787e80) at ../src/egl/drivers/dri2/egl_dri2.c:1307
  #17 eglDestroyContext (dpy=<optimized out>, ctx=0x5628af787e80) at ../src/egl/main/eglapi.c:930
  #18 QtWaylandClient::QWaylandGLContext::~QWaylandGLContext() () at libQt6WaylandEglClientHwIntegration.so.6
  #20 QOpenGLContext::destroy() () at libQt6Gui.so.6
  #22 QRhiGles2InitParams::newFallbackSurface(QSurfaceFormat const&) () at libQt6Gui.so.6
  #23 QSGRhiSupport::maybeCreateOffscreenSurface(QWindow*) () at libQt6Quick.so.6
  #25 QWindow::event(QEvent*) () at libQt6Gui.so.6
  #28 QGuiApplicationPrivate::processExposeEvent(...) () at libQt6Gui.so.6

State in frame #7:

  *queue: name = "plasmashel:sh" (&screen->shader_compiler_queue),
          num_threads = 7, max_threads = 12, num_queued = 0,
          create_threads_on_demand = true
  barrier (pthread_barrier_t) count = 6
  (gdb) x/8wx fences
  0x5628b02b1050: 0x00000000 0x00000000 0x00000000 0x00000000
  0x5628b02b1060: 0x00000000 0x00000000 0x00000091 0x00000000

So the fences array and the barrier were sized for 6 threads, and all
6 fences are signalled (0). But queue->num_threads grew to 7 before
the unlocked wait loop, so it waits on fences[6]. That address is past
the end of the 24-byte allocation, in the next malloc chunk's size
field (0x91 = 0x90 | PREV_INUSE), which will never become 0. All other
worker threads are idle in util_queue_thread_func (u_queue.c:275).

This matches the upstream commit message exactly ("waiting on an
invalid fence[1] value ... a second thread was created").

The patch
---------
The upstream fix is two lines: save num_threads before unlocking and
use the saved value in the wait loop. It applies cleanly to
25.0.7-2+deb13u1 (checked with patch --dry-run on top of the Debian
patch series). A DEP-3 version is attached.

Mesa 26.1.6-1~bpo13+1 (trixie-backports) and unstable already have the
fix. Please consider it for the next trixie point release, since any
radeonsi user with a multithreaded Qt Quick / EGL application can hit
it on context destruction.

The race is timing-dependent, so I can't reproduce it on demand. I
have not yet run a rebuilt 25.0.7 with the patch, so I can't confirm
the fix from my own testing. The analysis above is from the core dump
and the upstream commit.

Thanks!