#1099016 nvidia-legacy-340xx-driver: kernel crash when loading driver on trixie with kernel 6.12.12

#1099016#5
Date:
2025-02-27 09:47:20 UTC
From:
To:
Dear Maintainer,

While trying to load the nvidia-legacy-340xx driver on trixie, the
kernel crashes. Kernel version is 6.12.12.

Full log attached.

févr. 27 10:11:17 debian kernel: ------------[ cut here ]------------
févr. 27 10:11:17 debian kernel: Unpatched return thunk in use. This
should not happen!
févr. 27 10:11:17 debian kernel: WARNING: CPU: 0 PID: 324 at
arch/x86/kernel/cpu/bugs.c:3063 __warn_thunk+0x2a/0x40
févr. 27 10:11:17 debian kernel: Modules linked in: nvidia(POE+)
binfmt_misc drm firewire_sbp2 msr parport_pc ppdev lp parport configfs
efi_pstore nfnetlink efivarfs ip_tables x_tables autofs4 ext4 crc16
mbcache jbd2 crc32c_generic uas usb_storage hid_appleir hid_apple
hid_generic usbhid hid sd_mod ohci_pci ohci_hcd ehci_pci ehci_hcd
firewire_ohci ahci libahci usbcore firewire_core libata forcedeth
scsi_mod crc_itu_t scsi_common i2c_nforce2 usb_common button
févr. 27 10:11:17 debian kernel: CPU: 0 UID: 0 PID: 324 Comm: modprobe
Tainted: P           OE      6.12.12-amd64 #1  Debian 6.12.12-1
févr. 27 10:11:17 debian kernel: Tainted: [P]=PROPRIETARY_MODULE,
[O]=OOT_MODULE, [E]=UNSIGNED_MODULE
févr. 27 10:11:17 debian kernel: Hardware name: Apple Inc.
iMac9,1/Mac-F2218FC8, BIOS     IM91.88Z.008D.B08.0904271717 04/27/09
févr. 27 10:11:17 debian kernel: RIP: 0010:__warn_thunk+0x2a/0x40
févr. 27 10:11:17 debian kernel: Code: 66 0f 1f 00 0f 1f 44 00 00 80 3d
98 a9 dc 01 00 74 05 c3 cc cc cc cc 48 c7 c7 78 c4 13 9a c6 05 83 a9 dc
01 01 e8 d6 2f 06 00 <0f> 0b c3 cc cc cc cc 66 2e 0f 1f 84 00 00 00 00
00 0f 1f 44 00 00
févr. 27 10:11:17 debian kernel: RSP: 0018:ffffa6738044bb00 EFLAGS: 00010282
févr. 27 10:11:17 debian kernel: RAX: 0000000000000000 RBX:
ffffffffc1421920 RCX: 0000000000000027
févr. 27 10:11:17 debian kernel: RDX: ffff934578021788 RSI:
0000000000000001 RDI: ffff934578021780
févr. 27 10:11:17 debian kernel: RBP: ffffa6738044bb50 R08:
0000000000000000 R09: ffffa6738044b980
févr. 27 10:11:17 debian kernel: R10: ffffffff9a8b4348 R11:
0000000000000003 R12: 0000563317d5aa10
févr. 27 10:11:17 debian kernel: R13: ffffa6738044bbd8 R14:
ffff9344f6b21600 R15: ffff934445b3ced8
févr. 27 10:11:17 debian kernel: FS:  00007f615d932640(0000)
GS:ffff934578000000(0000) knlGS:0000000000000000
févr. 27 10:11:17 debian kernel: CS:  0010 DS: 0000 ES: 0000 CR0:
0000000080050033
févr. 27 10:11:17 debian kernel: CR2: 000055ba9df84eb8 CR3:
0000000120c8c000 CR4: 00000000000406f0
févr. 27 10:11:17 debian kernel: Call Trace:
févr. 27 10:11:17 debian kernel:  <TASK>
févr. 27 10:11:17 debian kernel:  ? __warn_thunk+0x2a/0x40
févr. 27 10:11:17 debian kernel:  ? __warn.cold+0x93/0xf6
févr. 27 10:11:17 debian kernel:  ? __warn_thunk+0x2a/0x40
févr. 27 10:11:17 debian kernel:  ? report_bug+0xff/0x140
févr. 27 10:11:17 debian kernel:  ? handle_bug+0x58/0x90
févr. 27 10:11:17 debian kernel:  ? exc_invalid_op+0x17/0x70
févr. 27 10:11:17 debian kernel:  ? asm_exc_invalid_op+0x1a/0x20
févr. 27 10:11:17 debian kernel:  ? nv_drm_init+0x130/0x130 [nvidia]
févr. 27 10:11:17 debian kernel:  ? __warn_thunk+0x2a/0x40
févr. 27 10:11:17 debian kernel:  warn_thunk_thunk+0x1a/0x30
févr. 27 10:11:17 debian kernel:  ? nv_drm_init+0x130/0x130 [nvidia]
févr. 27 10:11:17 debian kernel:  nvidia_init_module+0x4b/0x7e0 [nvidia]
févr. 27 10:11:17 debian kernel:  ? nv_drm_init+0x130/0x130 [nvidia]
févr. 27 10:11:17 debian kernel:  ?
nvidia_frontend_init_module+0x50/0x6e0 [nvidia]
févr. 27 10:11:17 debian kernel:  ? nv_drm_init+0x130/0x130 [nvidia]
févr. 27 10:11:17 debian kernel:  ? do_one_initcall+0x5b/0x310
févr. 27 10:11:17 debian kernel:  ? do_init_module+0x60/0x230
févr. 27 10:11:17 debian kernel:  ? init_module_from_file+0x89/0xe0
févr. 27 10:11:17 debian kernel:  ? idempotent_init_module+0x11e/0x310
févr. 27 10:11:17 debian kernel:  ? __x64_sys_finit_module+0x5e/0xb0
févr. 27 10:11:17 debian kernel:  ? do_syscall_64+0x82/0x190
févr. 27 10:11:17 debian kernel:  ? switch_fpu_return+0x4e/0xd0
févr. 27 10:11:17 debian kernel:  ? syscall_exit_to_user_mode+0x172/0x210
févr. 27 10:11:17 debian kernel:  ? do_syscall_64+0x8e/0x190
févr. 27 10:11:17 debian kernel:  ? __count_memcg_events+0x53/0xf0
févr. 27 10:11:17 debian kernel:  ? count_memcg_events.constprop.0+0x1a/0x30
févr. 27 10:11:17 debian kernel:  ? handle_mm_fault+0x1bb/0x2c0
févr. 27 10:11:17 debian kernel:  ? do_user_addr_fault+0x36c/0x620
févr. 27 10:11:17 debian kernel:  ? exc_page_fault+0x7e/0x180
févr. 27 10:11:17 debian kernel:  ? entry_SYSCALL_64_after_hwframe+0x76/0x7e
févr. 27 10:11:17 debian kernel:  </TASK>
févr. 27 10:11:17 debian kernel: ---[ end trace 0000000000000000 ]---

#1099016#10
Date:
2025-02-28 07:55:26 UTC
From:
To:
Adding build log
#1099016#15
Date:
2025-04-16 08:14:08 UTC
From:
To:
Control severity -1 grave

This is still present with the exact same stack trace with kernel 6.12.21
And debian revision 25 of the NVIDIA driver. (340.108-25)

Btw simply modprobing the driver even on a system without NVIDIA driver gives the exact same error.

Doing so on a virtual box system also has the same outcome.

I'm increasing the severity to grave because it seems a general issue for all. If this is not appropriate please change it back.

Regards
Fab

#1099016#20
Date:
2025-04-16 08:47:44 UTC
From:
To:
Maybe not as grave as I thought
Apparently it's only a warning.

I could be a consequence of this commit which appeard in kernel 6.9

https://github.com/torvalds/linux/commit/
4461438a8405e800f90e0e40409e5f3d07eed381


Le mercredi 16 avril 2025, 10:14:08 CEST Fab Stz a écrit :

#1099016#25
Date:
2025-04-16 08:58:07 UTC
From:
To:
There is nothing we can do about that since we cannot recompile the blob
parts with newer hardening options ...

Does the module still work? I'm trying to backport patches to keep the
module buildable for newer kernels, but I have no way to test it at all ;-)


Andreas

#1099016#30
Date:
2025-04-16 09:47:32 UTC
From:
To:
Last time I tried I thought it failed because of this log, but I would have to check again to be sure. It will take some time until I can access this computer.

Fab

Le 16 avril 2025 10:58:07 GMT+02:00, Andreas Beckmann <anbe@debian.org> a écrit :

#1099016#35
Date:
2025-05-23 07:47:30 UTC
From:
To:
Hello Andreas,

At last I could have a check on that computer.

Actually the driver seems to load fine despite the kernel warning but
the nvidia logo is not displayed and nothing is displayed on screen. But
the computer responds. I can reboot by doing Ctrl+Alt+F1 and then
Ctrl+Alt+Del.

I additionally created a symlink in /usr/lib/xorg/modules/drivers/
nv_drv.so -> nvidia_drv.so because this seems required. Maybe X is now
searching for "nv" instead of "nvidia" ?

So now in that dir I have these symlinks

nv_drv.so -> nvidia_drv.so
nvidia_drv.so -> /etc/alternatives/glx--nvidia_drv.so

However it is still failing. Please find full Xorg.0.log attached

Snippet below:

[    23.691] (EE) LoadModule: Module nv does not have a nvModuleData
data object.
[    23.691] (EE) Failed to load module "nv" (invalid module, 0)
[    24.331] (EE) [drm] Failed to open DRM device for pci:0000:03:00.0: -19
[    24.331] (EE) open /dev/dri/card0: Invalid argument
[    24.331] (EE) open /dev/dri/card0: Invalid argument
[    24.334] (EE) Unable to find a valid framebuffer device
[    24.335] (EE) Screen 0 deleted because of no matching config section.
[    24.335] (EE) Screen 0 deleted because of no matching config section.
[    24.380] (II) Initializing extension MIT-SCREEN-SAVER
[    24.386] (EE) Failed to initialize GLX extension (Compatible NVIDIA
X driver not found)

Maybe there is something problematic with the "nv" vs "nvidia" module name?

Do you have any idea on how to debug further?

Regards
Fab

#1099016#40
Date:
2025-05-23 08:26:16 UTC
From:
To:
On bookworm, Xorg log contains:

[    13.106] (II) LoadModule: "glx"
[    13.117] (II) Loading /usr/lib/xorg/modules/linux/libglx.so
[    13.600] (II) Module glx: vendor="NVIDIA Corporation"
[    13.600] 	compiled for 4.0.2, module version = 1.0.0
[    13.600] 	Module class: X.Org Server Extension
[    13.601] (II) NVIDIA GLX Module  340.108  Wed Dec 11 14:26:50 PST 2019
[    13.602] (II) Applying OutputClass "nvidia" to /dev/dri/card0
[    13.602] 	loading driver: nvidia
[    13.915] (==) Matched nvidia as autoconfigured driver 0
[    13.915] (==) Matched nouveau as autoconfigured driver 1
[    13.915] (==) Matched nv as autoconfigured driver 2
[    13.915] (==) Matched modesetting as autoconfigured driver 3
[    13.915] (==) Matched fbdev as autoconfigured driver 4
[    13.915] (==) Matched vesa as autoconfigured driver 5
[    13.915] (==) Assigned the driver to the xf86ConfigLayout
[    13.915] (II) LoadModule: "nvidia"



While on trixie it contains:

[   332.386] (II) LoadModule: "glx"
[   332.407] (II) Loading /usr/lib/xorg/modules/linux/libglx.so
[   334.514] (II) Module glx: vendor="NVIDIA Corporation"
[   334.514] 	compiled for 4.0.2, module version = 1.0.0
[   334.514] 	Module class: X.Org Server Extension
[   334.523] (II) NVIDIA GLX Module  340.108  Wed Dec 11 14:26:50 PST 2019
[   335.204] (==) Matched nouveau as autoconfigured driver 0
[   335.204] (==) Matched nv as autoconfigured driver 1
[   335.204] (==) Matched modesetting as autoconfigured driver 2
[   335.204] (==) Matched fbdev as autoconfigured driver 3
[   335.204] (==) Matched vesa as autoconfigured driver 4
[   335.204] (==) Assigned the driver to the xf86ConfigLayout
[   335.204] (II) LoadModule: "nouveau"


So  it looks like "nvidia" module/driver is skipped somehow because
trixie lacks these lines.

[    13.602] (II) Applying OutputClass "nvidia" to /dev/dri/card0
[    13.602] 	loading driver: nvidia
[    13.915] (==) Matched nvidia as autoconfigured driver 0


Regards
Fab

Le 23/05/2025 à 09:47, Fab Stz a écrit :

#1099016#45
Date:
2025-05-23 08:38:34 UTC
From:
To:
There is/was a 'nv' driver in Xorg, but it is no longer packaged in
Debian. It was likely superseded by 'nouveau'.

If xorg.conf does not specify a driver, Xorg will try all possibly
fitting driver names for the detected hardware, thus you see the failure
to load nv, but that does not matter.

But you shouldn't rename/symlink the drivers, as that's a source for
major confusion. Please revert.

You might try a minimal xorg.conf (or xorg.conf.d/*.conf snippet)
containing only

         Section "Device"
             Identifier     "My GPU"
             Driver         "nvidia"
         EndSection

to specifically select the proprietary driver and skip autoprobing.
Should help to avoid irrelevant error messages from autoprobed drivers.

But it may well be that the driver is no longer compatible with current
Xorg versions.


Andreas

#1099016#50
Date:
2025-05-23 11:37:41 UTC
From:
To:
Done, thanks for the explanation.
fine when loading only xterm.

Any idea why this xorg.conf is required now while it used to be
autodetected? Is this related to
https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1099016#40 ?

However
- running plasmashell leads to a crash with the nvidia driver. Same for
kwin. There is no crash with "nouveau" though.
- running gnome-session works fine and Gnome shows up without issue with
the nvidia driver.

gdm3 is displayed properly as well, while sddm is not.

plasmashell writes this output:

qt.glx: qglx_findConfig: Failed to finding matching FBConfig for
QSurfaceFormat(version 2.0, options
QFlags<QSurfaceFormat::FormatOption>(ResetNotification), depthBufferSize
-1, redBufferSize 1, greenBufferSize 1, blueBufferSize 1, 
alphaBufferSize -1, stencilBufferSize -1, samples -1, swapBehavior
QSurfaceFormat::SingleBuffer, swapInterval 1, colorSpace QColorSpace(),
profile  QSurfaceFormat::NoProfile)
Could not initialize GLX
Aborted (core dumped)

for kwin it is the same, except with swapInterval "0" instead of "1".


I could launch KDE/Plasma with any of:
export QT_XCB_GL_INTEGRATION=xcb_egl
export QT_XCB_GL_INTEGRATION=none

But graphics are not polished.

KDE/Plasma doesn't work when QT_XCB_GL_INTEGRATION is not set or has:
export QT_XCB_GL_INTEGRATION=xcb_glx

I tried these different values in .config/plasma-workspace/env/xcb.sh

Do you think a rebuild of the nvidia/Opengl driver would fix the issue?
Or is Kde/Plasma requiring features not provided by the nvidia driver?

It seems to work actually, at least partially. :)