#1142440 linux-image-amd64: amdgpu Mullins: intermittent hibernation resume hang with page fault in drm_sched_job_arm

Package:
linux-image-amd64
Source:
linux-image-amd64
Description:
Linux for 64-bit PCs (meta-package)
Submitter:
Bernie Osei
Date:
2026-07-20 01:31:01 UTC
Severity:
normal
#1142440#5
Date:
2026-07-19 23:52:11 UTC
From:
To:
Dear Maintainer,

Debian 13 trixie, kernel 6.12.95+deb13-amd64.

Hardware:
AMD A8-6410 Mullins integrated Radeon R4/R5 graphics
PCI ID 1002:9851, subsystem 103c:2268.

This CIK GPU is deliberately using amdgpu rather than the default
radeon driver, with these kernel parameters:

radeon.cik_support=0 amdgpu.cik_support=1 amdgpu.dpm=1

Hibernation uses a 16 GiB NOCOW Btrfs swapfile inside the LUKS2
encrypted root filesystem. The swapfile is active, and its Btrfs
physical resume offset matches the resume_offset on the kernel
command line:

resume=/dev/mapper/sda8_crypt
resume_offset=8278889

The Dracut initramfs contains the resume module and
systemd-hibernate-resume.

One manually initiated hibernation and resume completed successfully,
restoring the full desktop session, including unsaved editor contents.

A subsequent hibernation resume failed intermittently. After entering
the LUKS passphrase, the restored desktop became visible, but the system
froze before the GNOME lock-screen password field appeared. Caps Lock
continued to respond, but the graphical session and Ctrl+Alt+F3 did not.
Several short presses of the power button had no effect, and a forced
power-off was required.

The previous-boot kernel journal showed this sequence after restoration:

amdgpu_ring_test_helper: ring comp_1.0.0 through comp_1.0.7 test
failed (-110)

amdgpu: SRBM_SOFT_RESET=0x00100040

[drm] scheduler comp_1.0.x is not ready, skipping

#PF: error_code(0x0000) - not-present page
RIP: drm_sched_job_arm+0x23/0x60 [gpu_sched]

The call trace included amdgpu_cs_ioctl, drm_ioctl_kernel,
drm_ioctl and amdgpu_drm_ioctl. The page fault occurred more than once.

The journal also contained:

pstore: backend (efi_pstore) writing error (-22)

There was no captured "Kernel panic - not syncing" line. However, a
kernel page fault/oops in the AMDGPU scheduler was recorded and the
machine became unusable.

The same amdgpu_irq_put warning in the DCE 8 suspend path is observed
during both suspend and hibernation:

WARNING at drivers/gpu/drm/amd/amdgpu/amdgpu_irq.c:631 amdgpu_irq_put
dce_v8_0_hw_fini
dce_v8_0_suspend
amdgpu_device_suspend

Ordinary suspend-to-RAM has resumed successfully multiple times,
including after an approximately 18.5-hour suspend. The fatal failure
has so far occurred during hibernation restoration.

I can provide the full previous-boot kernel journal and test an updated
Debian kernel if requested.

#1142440#10
Date:
2026-07-20 01:28:01 UTC
From:
To:
Hello,

I have now tested the immediately preceding Debian kernel on the same
machine:

    6.12.94+deb13-amd64

The hardware and configuration were unchanged:

    AMD A8-6410 Mullins Radeon R4/R5
    PCI ID 1002:9851, subsystem 103c:2268

    radeon.cik_support=0
    amdgpu.cik_support=1
    amdgpu.dpm=1

    resume=/dev/mapper/sda8_crypt
    resume_offset=8278889

The 16 GiB Btrfs hibernation swapfile was active and the 6.12.94
initramfs contained the Dracut resume module.

Results with 6.12.94:

- graphical rendering was normal;
- an ordinary suspend/resume succeeded;
- two hibernation/resume cycles succeeded;
- the restored desktop remained usable;
- no page fault in drm_sched_job_arm was observed during these tests.

However, 6.12.94 still produced the same underlying AMDGPU recovery
errors seen with 6.12.95:

    WARNING at amdgpu_irq_put
    dce_v8_0_hw_fini
    dce_v8_0_suspend

    amdgpu_ring_test_helper:
    ring comp_1.0.0 through comp_1.0.3 test failed (-110)

    amdgpu: SRBM_SOFT_RESET=0x00100040

    [drm] scheduler comp_1.0.x is not ready, skipping

The scheduler-not-ready messages continued repeatedly after the
successful resume.

The observed difference so far is:

- 6.12.94: the AMDGPU compute-ring recovery failure occurred, but the
  system remained usable in the tests performed;
- 6.12.95: the same failure sequence was followed by repeated kernel
  page faults in drm_sched_job_arm and an unusable system requiring a
  forced power-off.

This suggests that the underlying AMDGPU hibernation/resume problem
predates 6.12.95, although 6.12.95 may be more likely to progress into
the fatal GPU scheduler fault.

I have attached the filtered 6.12.94 kernel journal from the successful
suspend and hibernation tests.

Regards,
Bernie Osei



<http://www.wisestamp.com/email-install?utm_source=extension&utm_medium=email&utm_campaign=footer>


*Disclaimer*
This message is intended only for the use of the person(s) ("Intended
Recipient") to whom it is addressed. It may contain information, which is
privileged and confidential. Accordingly any dissemination, distribution,
copying or other use of this message or any of its content by any person
other than the Intended Recipient may constitute a breach of civil or
criminal law and is strictly prohibited. If you are not the Intended
Recipient, please contact the sender as soon as possible.