#1053864 libdrm-amdgpu1: gpu crash on graphics start with Radeon 760M (both sway and gdm3)

Package:
libdrm-amdgpu1
Source:
libdrm-amdgpu1
Description:
Userspace interface to amdgpu-specific kernel DRM services -- runtime
Submitter:
Simon Heath
Date:
2023-12-26 02:12:03 UTC
Severity:
normal
Tags:
#1053864#5
Date:
2023-10-13 04:47:57 UTC
From:
To:
Dear Maintainer,

When GDM3 starts, or when I turn it off and log into the console by hand
and then start sway or another WM, often the graphics mode switch will
hang for a few seconds on an unresponsive black screen, then go back to
a text console for an instant and try again.  This seems to repeat 0-3
times until eventually it works successfully.  Sometimes it works on the
first try, often on the second try, etc.

Once Sway or GDM3 and Xorg have actually started, it *seems* perfectly
stable, as far as I've seen so far.

This is a brand new GPU chipset afaik so graphics bugs are pretty
understandable.

CPU: AMD Ryzen 5 7640U w/ Radeon 760M Graphics
Extended renderer info from `glxinfo`:
    Device: AMD Radeon Graphics (gfx1103_r1, LLVM 16.0.6, DRM 3.54, 6.5.0-1-amd64) (0x15bf)
    Version: 23.2.1

I also see the following errors in dmesg associated with the
apparent-crash-and-restart:

[   26.625039] [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ring sdma0 timeout, signaled seq=23, emitted seq=25
[   26.625482] [drm:amdgpu_job_timedout [amdgpu]] *ERROR* Process information: process  pid 0 thread  pid 0
[   26.625820] amdgpu 0000:c1:00.0: amdgpu: GPU reset begin!
[   26.810595] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
[   26.810761] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
[   26.944169] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
[   26.944310] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
[   27.077693] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
[   27.077834] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
[   27.211163] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
[   27.211303] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
[   27.344634] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
[   27.344776] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
[   27.478028] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
[   27.478175] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
[   27.611499] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
[   27.611640] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
[   27.744960] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
[   27.745097] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
[   27.878425] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
[   27.878564] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
[   27.880086] amdgpu 0000:c1:00.0: amdgpu: MODE2 reset
[   27.909811] amdgpu 0000:c1:00.0: amdgpu: GPU reset succeeded, trying to resume
[   27.910426] [drm] PCIE GART of 512M enabled (table at 0x000000801FD00000).
[   27.910540] amdgpu 0000:c1:00.0: amdgpu: SMU is resuming...
[   27.911480] amdgpu 0000:c1:00.0: amdgpu: SMU is resumed successfully!
[   27.913327] [drm] DMUB hardware initialized: version=0x08000E00
[   27.918776] [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:264
[   27.921376] [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:272
[   27.923969] [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:280
[   27.926566] [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:288
[   27.934650] [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:264
[   27.937248] [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:272
[   27.939841] [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:280
[   27.942439] [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:288
[   28.328853] [drm] kiq ring mec 3 pipe 1 q 0
[   28.331133] [drm] VCN decode and encode initialized successfully(under DPG Mode).
[   28.331252] amdgpu 0000:c1:00.0: [drm:jpeg_v4_0_hw_init [amdgpu]] JPEG decode initialized successfully.
[   28.331965] amdgpu 0000:c1:00.0: amdgpu: ring gfx_0.0.0 uses VM inv eng 0 on hub 0
[   28.331968] amdgpu 0000:c1:00.0: amdgpu: ring comp_1.0.0 uses VM inv eng 1 on hub 0
[   28.331971] amdgpu 0000:c1:00.0: amdgpu: ring comp_1.1.0 uses VM inv eng 4 on hub 0
[   28.331973] amdgpu 0000:c1:00.0: amdgpu: ring comp_1.2.0 uses VM inv eng 6 on hub 0
[   28.331975] amdgpu 0000:c1:00.0: amdgpu: ring comp_1.3.0 uses VM inv eng 7 on hub 0
[   28.331977] amdgpu 0000:c1:00.0: amdgpu: ring comp_1.0.1 uses VM inv eng 8 on hub 0
[   28.331979] amdgpu 0000:c1:00.0: amdgpu: ring comp_1.1.1 uses VM inv eng 9 on hub 0
[   28.331981] amdgpu 0000:c1:00.0: amdgpu: ring comp_1.2.1 uses VM inv eng 10 on hub 0
[   28.331983] amdgpu 0000:c1:00.0: amdgpu: ring comp_1.3.1 uses VM inv eng 11 on hub 0
[   28.331985] amdgpu 0000:c1:00.0: amdgpu: ring sdma0 uses VM inv eng 12 on hub 0
[   28.331987] amdgpu 0000:c1:00.0: amdgpu: ring vcn_unified_0 uses VM inv eng 0 on hub 8
[   28.331990] amdgpu 0000:c1:00.0: amdgpu: ring jpeg_dec uses VM inv eng 1 on hub 8
[   28.331992] amdgpu 0000:c1:00.0: amdgpu: ring mes_kiq_3.1.0 uses VM inv eng 13 on hub 0
[   28.334786] amdgpu 0000:c1:00.0: amdgpu: recover vram bo from shadow start
[   28.334791] amdgpu 0000:c1:00.0: amdgpu: recover vram bo from shadow done
[   28.334933] [drm] Skip scheduling IBs!
[   28.334955] [drm] Skip scheduling IBs!
[   28.334964] [drm] Skip scheduling IBs!
[   28.334971] [drm] Skip scheduling IBs!
[   28.334979] [drm] Skip scheduling IBs!
[   28.334987] [drm] Skip scheduling IBs!
[   28.334995] [drm] Skip scheduling IBs!
[   28.335006] [drm] Skip scheduling IBs!
[   28.335014] [drm] Skip scheduling IBs!
[   28.335070] [drm] Skip scheduling IBs!
[   28.335079] [drm] Skip scheduling IBs!
[   28.335085] [drm] Skip scheduling IBs!
[   28.336265] [drm] ring gfx_32776.1.1 was added
[   28.337256] [drm] ring compute_32776.2.2 was added
[   28.338182] [drm] ring sdma_32776.3.3 was added
[   28.338234] [drm] ring gfx_32776.1.1 ib test pass
[   28.338272] [drm] ring compute_32776.2.2 ib test pass
[   28.338470] [drm] ring sdma_32776.3.3 ib test pass
[   28.339726] amdgpu 0000:c1:00.0: amdgpu: GPU reset(1) succeeded!
[   28.518882] [drm] Skip scheduling IBs!
[   28.518892] [drm] Skip scheduling IBs!
[   28.518897] [drm] Skip scheduling IBs!
[   28.520085] [drm] Skip scheduling IBs!
[   28.521361] [drm] Skip scheduling IBs!
[   28.541083] [drm] Skip scheduling IBs!
[   28.541114] [drm] Skip scheduling IBs!
[   28.541143] [drm] Skip scheduling IBs!
[   28.541159] [drm] Skip scheduling IBs!
[   28.541173] [drm] Skip scheduling IBs!
[   28.541193] [drm] Skip scheduling IBs!
[   28.541215] [drm] Skip scheduling IBs!
[   28.541219] [drm] Skip scheduling IBs!
[   28.541239] [drm] Skip scheduling IBs!


I also get the following errors in dmesg from time to time, but they have no visible impact so far:

[ 1046.269344] [drm:amdgpu_dm_process_dmub_aux_transfer_sync [amdgpu]] *ERROR* wait_for_completion_timeout timeout!
[ 1056.509203] [drm:amdgpu_dm_process_dmub_aux_transfer_sync [amdgpu]] *ERROR* wait_for_completion_timeout timeout!
[ 1066.749132] [drm:amdgpu_dm_process_dmub_aux_transfer_sync [amdgpu]] *ERROR* wait_for_completion_timeout timeout!
[ 1076.988590] [drm:amdgpu_dm_process_dmub_aux_transfer_sync [amdgpu]] *ERROR* wait_for_completion_timeout timeout!
[ 1087.228896] [drm:amdgpu_dm_process_dmub_aux_transfer_sync [amdgpu]] *ERROR* wait_for_completion_timeout timeout!
[ 1094.983205] i2c_hid_acpi i2c-FRMW0005:00: i2c_hid_get_input: incomplete report (7/65535)
[ 1097.468792] [drm:amdgpu_dm_process_dmub_aux_transfer_sync [amdgpu]] *ERROR* wait_for_completion_timeout timeout!
[ 1107.708726] [drm:amdgpu_dm_process_dmub_aux_transfer_sync [amdgpu]] *ERROR* wait_for_completion_timeout timeout!
[ 1117.948141] [drm:amdgpu_dm_process_dmub_aux_transfer_sync [amdgpu]] *ERROR* wait_for_completion_timeout timeout!
[ 1128.188485] [drm:amdgpu_dm_process_dmub_aux_transfer_sync [amdgpu]] *ERROR* wait_for_completion_timeout timeout!
[ 1138.428402] [drm:amdgpu_dm_process_dmub_aux_transfer_sync [amdgpu]] *ERROR* wait_for_completion_timeout timeout!
[ 1148.668306] [drm:amdgpu_dm_process_dmub_aux_transfer_sync [amdgpu]] *ERROR* wait_for_completion_timeout timeout!
[ 1158.908169] [drm:amdgpu_dm_process_dmub_aux_transfer_sync [amdgpu]] *ERROR* wait_for_completion_timeout timeout!
[ 1169.147619] [drm:amdgpu_dm_process_dmub_aux_transfer_sync [amdgpu]] *ERROR* wait_for_completion_timeout timeout!
[ 1179.387933] [drm:amdgpu_dm_process_dmub_aux_transfer_sync [amdgpu]] *ERROR* wait_for_completion_timeout timeout!


Thank you, hopefully this info is useful to someone!

Simon Heath

#1053864#10
Date:
2023-11-08 15:14:20 UTC
From:
To:
Control: tag -1 moreinfo

Those messages are actually from the kernel driver.
Can you test whether the issue is still present with kernel 6.5.8-1 (Testing)
and if so, also try it with 6.5.10-1 from Unstable?

#1053864#17
Date:
2023-11-12 23:29:59 UTC
From:
To:
Oop, my bad.  I was wondering why I hadn't seen it go through on the bug
report...

The issue is still present in apt package linux-image-6.5.0-3 (Kernel
6.5.8-1) , and linux-image-6.5.0-4 (kernel 6.5.10-1). Same messages, as
far as I can see, but here's the dmesg output from the 6.5.10-1 kernel
in case there's something subtly different.

Thanks,
Simon
---- [    7.490078] ucsi_acpi USBC000:00: ucsi_handle_connector_change: GET_CONNECTOR_STATUS failed (-5) [    7.605873] ucsi_acpi USBC000:00: possible UCSI driver bug 1 [    7.605903] ucsi_acpi USBC000:00: ucsi_handle_connector_change: GET_CONNECTOR_STATUS failed (-22) [   13.555707] pipewire[1065]: memfd_create() called without MFD_EXEC or MFD_NOEXEC_SEAL set [   23.808871] [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ring sdma0 timeout, signaled seq=23, emitted seq=25 [   23.809320] [drm:amdgpu_job_timedout [amdgpu]] *ERROR* Process information: process  pid 0 thread  pid 0 [   23.809592] amdgpu 0000:c1:00.0: amdgpu: GPU reset begin! [   23.990678] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3 [   23.990842] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue [   24.124228] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3 [   24.124374] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue [   24.257754] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3 [   24.257918] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue [   24.391326] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3 [   24.391555] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue [   24.525068] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3 [   24.525211] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue [   24.658617] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3 [   24.658758] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue [   24.792155] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3 [   24.792326] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue [   24.925815] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3 [   24.925961] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue [   25.059344] [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3 [   25.059488] [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue [   25.061023] amdgpu 0000:c1:00.0: amdgpu: MODE2 reset [   25.090107] amdgpu 0000:c1:00.0: amdgpu: GPU reset succeeded, trying to resume [   25.090767] [drm] PCIE GART of 512M enabled (table at 0x000000801FD00000). [   25.090889] amdgpu 0000:c1:00.0: amdgpu: SMU is resuming... [   25.092526] amdgpu 0000:c1:00.0: amdgpu: SMU is resumed successfully! [   25.094267] [drm] DMUB hardware initialized: version=0x08000E00 [   25.101834] [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:264 [   25.104428] [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:272 [   25.107025] [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:280 [   25.109617] [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:288 [   25.117187] [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:264 [   25.119782] [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:272 [   25.122380] [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:280 [   25.124993] [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:288 [   25.534004] [drm] kiq ring mec 3 pipe 1 q 0 [   25.536314] [drm] VCN decode and encode initialized successfully(under DPG Mode). [   25.536470] amdgpu 0000:c1:00.0: [drm:jpeg_v4_0_hw_init [amdgpu]] JPEG decode initialized successfully. [   25.537196] amdgpu 0000:c1:00.0: amdgpu: ring gfx_0.0.0 uses VM inv eng 0 on hub 0 [   25.537200] amdgpu 0000:c1:00.0: amdgpu: ring comp_1.0.0 uses VM inv eng 1 on hub 0 [   25.537202] amdgpu 0000:c1:00.0: amdgpu: ring comp_1.1.0 uses VM inv eng 4 on hub 0 [   25.537204] amdgpu 0000:c1:00.0: amdgpu: ring comp_1.2.0 uses VM inv eng 6 on hub 0 [   25.537206] amdgpu 0000:c1:00.0: amdgpu: ring comp_1.3.0 uses VM inv eng 7 on hub 0 [   25.537208] amdgpu 0000:c1:00.0: amdgpu: ring comp_1.0.1 uses VM inv eng 8 on hub 0 [   25.537210] amdgpu 0000:c1:00.0: amdgpu: ring comp_1.1.1 uses VM inv eng 9 on hub 0 [   25.537212] amdgpu 0000:c1:00.0: amdgpu: ring comp_1.2.1 uses VM inv eng 10 on hub 0 [   25.537214] amdgpu 0000:c1:00.0: amdgpu: ring comp_1.3.1 uses VM inv eng 11 on hub 0 [   25.537215] amdgpu 0000:c1:00.0: amdgpu: ring sdma0 uses VM inv eng 12 on hub 0 [   25.537217] amdgpu 0000:c1:00.0: amdgpu: ring vcn_unified_0 uses VM inv eng 0 on hub 8 [   25.537219] amdgpu 0000:c1:00.0: amdgpu: ring jpeg_dec uses VM inv eng 1 on hub 8 [   25.537221] amdgpu 0000:c1:00.0: amdgpu: ring mes_kiq_3.1.0 uses VM inv eng 13 on hub 0 [   25.539979] amdgpu 0000:c1:00.0: amdgpu: recover vram bo from shadow start [   25.539981] amdgpu 0000:c1:00.0: amdgpu: recover vram bo from shadow done [   25.540023] [drm] Skip scheduling IBs! [   25.540030] [drm] Skip scheduling IBs! [   25.540033] [drm] Skip scheduling IBs! [   25.540036] [drm] Skip scheduling IBs! [   25.540039] [drm] Skip scheduling IBs! [   25.540042] [drm] Skip scheduling IBs! [   25.540045] [drm] Skip scheduling IBs! [   25.540047] [drm] Skip scheduling IBs! [   25.540051] [drm] Skip scheduling IBs! [   25.540222] [drm] Skip scheduling IBs! [   25.540228] [drm] Skip scheduling IBs! [   25.540231] [drm] Skip scheduling IBs! [   25.541423] [drm] ring gfx_32776.1.1 was added [   25.542373] [drm] ring compute_32776.2.2 was added [   25.543269] [drm] ring sdma_32776.3.3 was added [   25.543319] [drm] ring gfx_32776.1.1 ib test pass [   25.543347] [drm] ring compute_32776.2.2 ib test pass [   25.543526] [drm] ring sdma_32776.3.3 ib test pass [   25.544782] amdgpu 0000:c1:00.0: amdgpu: GPU reset(1) succeeded! [   25.711149] [drm] Skip scheduling IBs! [   25.711159] [drm] Skip scheduling IBs! [   25.711163] [drm] Skip scheduling IBs! [   25.712209] [drm] Skip scheduling IBs! [   25.713415] [drm] Skip scheduling IBs! [   25.730490] [drm] Skip scheduling IBs! [   25.730526] [drm] Skip scheduling IBs! [   25.730553] [drm] Skip scheduling IBs! [   25.730566] [drm] Skip scheduling IBs! [   25.730578] [drm] Skip scheduling IBs! [   25.730595] [drm] Skip scheduling IBs! [   25.730617] [drm] Skip scheduling IBs! [   25.730620] [drm] Skip scheduling IBs! [   25.730632] [drm] Skip scheduling IBs!
#1053864#22
Date:
2023-12-26 02:02:13 UTC
From:
To:
fwupdmgr has a firmware upgrade for the laptop in question (AMD Framework 13), titled
System Firmware (0.0.3.2 → 0.0.3.3).  The 0.0.3.2 version's issues with graphics
drivers are well documented by the vendor; I would swear that this was already
installed, but apparently not.  On doing this update, everything appears to work
fine.

Welp, now we know I guess.  Thanks.