#1099016 nvidia-legacy-340xx-driver: kernel crash when loading driver on trixie with kernel 6.12.12 #1099016
- Package:
- nvidia-legacy-340xx-driver
- Source:
- nvidia-legacy-340xx-driver
- Submitter:
- Fab Stz
- Date:
- 2025-05-23 11:39:01 UTC
- Severity:
- normal
Dear Maintainer, While trying to load the nvidia-legacy-340xx driver on trixie, the kernel crashes. Kernel version is 6.12.12. Full log attached. févr. 27 10:11:17 debian kernel: ------------[ cut here ]------------ févr. 27 10:11:17 debian kernel: Unpatched return thunk in use. This should not happen! févr. 27 10:11:17 debian kernel: WARNING: CPU: 0 PID: 324 at arch/x86/kernel/cpu/bugs.c:3063 __warn_thunk+0x2a/0x40 févr. 27 10:11:17 debian kernel: Modules linked in: nvidia(POE+) binfmt_misc drm firewire_sbp2 msr parport_pc ppdev lp parport configfs efi_pstore nfnetlink efivarfs ip_tables x_tables autofs4 ext4 crc16 mbcache jbd2 crc32c_generic uas usb_storage hid_appleir hid_apple hid_generic usbhid hid sd_mod ohci_pci ohci_hcd ehci_pci ehci_hcd firewire_ohci ahci libahci usbcore firewire_core libata forcedeth scsi_mod crc_itu_t scsi_common i2c_nforce2 usb_common button févr. 27 10:11:17 debian kernel: CPU: 0 UID: 0 PID: 324 Comm: modprobe Tainted: P OE 6.12.12-amd64 #1 Debian 6.12.12-1 févr. 27 10:11:17 debian kernel: Tainted: [P]=PROPRIETARY_MODULE, [O]=OOT_MODULE, [E]=UNSIGNED_MODULE févr. 27 10:11:17 debian kernel: Hardware name: Apple Inc. iMac9,1/Mac-F2218FC8, BIOS IM91.88Z.008D.B08.0904271717 04/27/09 févr. 27 10:11:17 debian kernel: RIP: 0010:__warn_thunk+0x2a/0x40 févr. 27 10:11:17 debian kernel: Code: 66 0f 1f 00 0f 1f 44 00 00 80 3d 98 a9 dc 01 00 74 05 c3 cc cc cc cc 48 c7 c7 78 c4 13 9a c6 05 83 a9 dc 01 01 e8 d6 2f 06 00 <0f> 0b c3 cc cc cc cc 66 2e 0f 1f 84 00 00 00 00 00 0f 1f 44 00 00 févr. 27 10:11:17 debian kernel: RSP: 0018:ffffa6738044bb00 EFLAGS: 00010282 févr. 27 10:11:17 debian kernel: RAX: 0000000000000000 RBX: ffffffffc1421920 RCX: 0000000000000027 févr. 27 10:11:17 debian kernel: RDX: ffff934578021788 RSI: 0000000000000001 RDI: ffff934578021780 févr. 27 10:11:17 debian kernel: RBP: ffffa6738044bb50 R08: 0000000000000000 R09: ffffa6738044b980 févr. 27 10:11:17 debian kernel: R10: ffffffff9a8b4348 R11: 0000000000000003 R12: 0000563317d5aa10 févr. 27 10:11:17 debian kernel: R13: ffffa6738044bbd8 R14: ffff9344f6b21600 R15: ffff934445b3ced8 févr. 27 10:11:17 debian kernel: FS: 00007f615d932640(0000) GS:ffff934578000000(0000) knlGS:0000000000000000 févr. 27 10:11:17 debian kernel: CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 févr. 27 10:11:17 debian kernel: CR2: 000055ba9df84eb8 CR3: 0000000120c8c000 CR4: 00000000000406f0 févr. 27 10:11:17 debian kernel: Call Trace: févr. 27 10:11:17 debian kernel: <TASK> févr. 27 10:11:17 debian kernel: ? __warn_thunk+0x2a/0x40 févr. 27 10:11:17 debian kernel: ? __warn.cold+0x93/0xf6 févr. 27 10:11:17 debian kernel: ? __warn_thunk+0x2a/0x40 févr. 27 10:11:17 debian kernel: ? report_bug+0xff/0x140 févr. 27 10:11:17 debian kernel: ? handle_bug+0x58/0x90 févr. 27 10:11:17 debian kernel: ? exc_invalid_op+0x17/0x70 févr. 27 10:11:17 debian kernel: ? asm_exc_invalid_op+0x1a/0x20 févr. 27 10:11:17 debian kernel: ? nv_drm_init+0x130/0x130 [nvidia] févr. 27 10:11:17 debian kernel: ? __warn_thunk+0x2a/0x40 févr. 27 10:11:17 debian kernel: warn_thunk_thunk+0x1a/0x30 févr. 27 10:11:17 debian kernel: ? nv_drm_init+0x130/0x130 [nvidia] févr. 27 10:11:17 debian kernel: nvidia_init_module+0x4b/0x7e0 [nvidia] févr. 27 10:11:17 debian kernel: ? nv_drm_init+0x130/0x130 [nvidia] févr. 27 10:11:17 debian kernel: ? nvidia_frontend_init_module+0x50/0x6e0 [nvidia] févr. 27 10:11:17 debian kernel: ? nv_drm_init+0x130/0x130 [nvidia] févr. 27 10:11:17 debian kernel: ? do_one_initcall+0x5b/0x310 févr. 27 10:11:17 debian kernel: ? do_init_module+0x60/0x230 févr. 27 10:11:17 debian kernel: ? init_module_from_file+0x89/0xe0 févr. 27 10:11:17 debian kernel: ? idempotent_init_module+0x11e/0x310 févr. 27 10:11:17 debian kernel: ? __x64_sys_finit_module+0x5e/0xb0 févr. 27 10:11:17 debian kernel: ? do_syscall_64+0x82/0x190 févr. 27 10:11:17 debian kernel: ? switch_fpu_return+0x4e/0xd0 févr. 27 10:11:17 debian kernel: ? syscall_exit_to_user_mode+0x172/0x210 févr. 27 10:11:17 debian kernel: ? do_syscall_64+0x8e/0x190 févr. 27 10:11:17 debian kernel: ? __count_memcg_events+0x53/0xf0 févr. 27 10:11:17 debian kernel: ? count_memcg_events.constprop.0+0x1a/0x30 févr. 27 10:11:17 debian kernel: ? handle_mm_fault+0x1bb/0x2c0 févr. 27 10:11:17 debian kernel: ? do_user_addr_fault+0x36c/0x620 févr. 27 10:11:17 debian kernel: ? exc_page_fault+0x7e/0x180 févr. 27 10:11:17 debian kernel: ? entry_SYSCALL_64_after_hwframe+0x76/0x7e févr. 27 10:11:17 debian kernel: </TASK> févr. 27 10:11:17 debian kernel: ---[ end trace 0000000000000000 ]---
Adding build log
Control severity -1 grave This is still present with the exact same stack trace with kernel 6.12.21 And debian revision 25 of the NVIDIA driver. (340.108-25) Btw simply modprobing the driver even on a system without NVIDIA driver gives the exact same error. Doing so on a virtual box system also has the same outcome. I'm increasing the severity to grave because it seems a general issue for all. If this is not appropriate please change it back. Regards Fab
Maybe not as grave as I thought Apparently it's only a warning. I could be a consequence of this commit which appeard in kernel 6.9 https://github.com/torvalds/linux/commit/ 4461438a8405e800f90e0e40409e5f3d07eed381 Le mercredi 16 avril 2025, 10:14:08 CEST Fab Stz a écrit :
There is nothing we can do about that since we cannot recompile the blob parts with newer hardening options ... Does the module still work? I'm trying to backport patches to keep the module buildable for newer kernels, but I have no way to test it at all ;-) Andreas
Last time I tried I thought it failed because of this log, but I would have to check again to be sure. It will take some time until I can access this computer. Fab Le 16 avril 2025 10:58:07 GMT+02:00, Andreas Beckmann <anbe@debian.org> a écrit :
Hello Andreas, At last I could have a check on that computer. Actually the driver seems to load fine despite the kernel warning but the nvidia logo is not displayed and nothing is displayed on screen. But the computer responds. I can reboot by doing Ctrl+Alt+F1 and then Ctrl+Alt+Del. I additionally created a symlink in /usr/lib/xorg/modules/drivers/ nv_drv.so -> nvidia_drv.so because this seems required. Maybe X is now searching for "nv" instead of "nvidia" ? So now in that dir I have these symlinks nv_drv.so -> nvidia_drv.so nvidia_drv.so -> /etc/alternatives/glx--nvidia_drv.so However it is still failing. Please find full Xorg.0.log attached Snippet below: [ 23.691] (EE) LoadModule: Module nv does not have a nvModuleData data object. [ 23.691] (EE) Failed to load module "nv" (invalid module, 0) [ 24.331] (EE) [drm] Failed to open DRM device for pci:0000:03:00.0: -19 [ 24.331] (EE) open /dev/dri/card0: Invalid argument [ 24.331] (EE) open /dev/dri/card0: Invalid argument [ 24.334] (EE) Unable to find a valid framebuffer device [ 24.335] (EE) Screen 0 deleted because of no matching config section. [ 24.335] (EE) Screen 0 deleted because of no matching config section. [ 24.380] (II) Initializing extension MIT-SCREEN-SAVER [ 24.386] (EE) Failed to initialize GLX extension (Compatible NVIDIA X driver not found) Maybe there is something problematic with the "nv" vs "nvidia" module name? Do you have any idea on how to debug further? Regards Fab
On bookworm, Xorg log contains: [ 13.106] (II) LoadModule: "glx" [ 13.117] (II) Loading /usr/lib/xorg/modules/linux/libglx.so [ 13.600] (II) Module glx: vendor="NVIDIA Corporation" [ 13.600] compiled for 4.0.2, module version = 1.0.0 [ 13.600] Module class: X.Org Server Extension [ 13.601] (II) NVIDIA GLX Module 340.108 Wed Dec 11 14:26:50 PST 2019 [ 13.602] (II) Applying OutputClass "nvidia" to /dev/dri/card0 [ 13.602] loading driver: nvidia [ 13.915] (==) Matched nvidia as autoconfigured driver 0 [ 13.915] (==) Matched nouveau as autoconfigured driver 1 [ 13.915] (==) Matched nv as autoconfigured driver 2 [ 13.915] (==) Matched modesetting as autoconfigured driver 3 [ 13.915] (==) Matched fbdev as autoconfigured driver 4 [ 13.915] (==) Matched vesa as autoconfigured driver 5 [ 13.915] (==) Assigned the driver to the xf86ConfigLayout [ 13.915] (II) LoadModule: "nvidia" While on trixie it contains: [ 332.386] (II) LoadModule: "glx" [ 332.407] (II) Loading /usr/lib/xorg/modules/linux/libglx.so [ 334.514] (II) Module glx: vendor="NVIDIA Corporation" [ 334.514] compiled for 4.0.2, module version = 1.0.0 [ 334.514] Module class: X.Org Server Extension [ 334.523] (II) NVIDIA GLX Module 340.108 Wed Dec 11 14:26:50 PST 2019 [ 335.204] (==) Matched nouveau as autoconfigured driver 0 [ 335.204] (==) Matched nv as autoconfigured driver 1 [ 335.204] (==) Matched modesetting as autoconfigured driver 2 [ 335.204] (==) Matched fbdev as autoconfigured driver 3 [ 335.204] (==) Matched vesa as autoconfigured driver 4 [ 335.204] (==) Assigned the driver to the xf86ConfigLayout [ 335.204] (II) LoadModule: "nouveau" So it looks like "nvidia" module/driver is skipped somehow because trixie lacks these lines. [ 13.602] (II) Applying OutputClass "nvidia" to /dev/dri/card0 [ 13.602] loading driver: nvidia [ 13.915] (==) Matched nvidia as autoconfigured driver 0 Regards Fab Le 23/05/2025 à 09:47, Fab Stz a écrit :
There is/was a 'nv' driver in Xorg, but it is no longer packaged in
Debian. It was likely superseded by 'nouveau'.
If xorg.conf does not specify a driver, Xorg will try all possibly
fitting driver names for the detected hardware, thus you see the failure
to load nv, but that does not matter.
But you shouldn't rename/symlink the drivers, as that's a source for
major confusion. Please revert.
You might try a minimal xorg.conf (or xorg.conf.d/*.conf snippet)
containing only
Section "Device"
Identifier "My GPU"
Driver "nvidia"
EndSection
to specifically select the proprietary driver and skip autoprobing.
Should help to avoid irrelevant error messages from autoprobed drivers.
But it may well be that the driver is no longer compatible with current
Xorg versions.
Andreas
Done, thanks for the explanation. fine when loading only xterm. Any idea why this xorg.conf is required now while it used to be autodetected? Is this related to https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1099016#40 ? However - running plasmashell leads to a crash with the nvidia driver. Same for kwin. There is no crash with "nouveau" though. - running gnome-session works fine and Gnome shows up without issue with the nvidia driver. gdm3 is displayed properly as well, while sddm is not. plasmashell writes this output: qt.glx: qglx_findConfig: Failed to finding matching FBConfig for QSurfaceFormat(version 2.0, options QFlags<QSurfaceFormat::FormatOption>(ResetNotification), depthBufferSize -1, redBufferSize 1, greenBufferSize 1, blueBufferSize 1, alphaBufferSize -1, stencilBufferSize -1, samples -1, swapBehavior QSurfaceFormat::SingleBuffer, swapInterval 1, colorSpace QColorSpace(), profile QSurfaceFormat::NoProfile) Could not initialize GLX Aborted (core dumped) for kwin it is the same, except with swapInterval "0" instead of "1". I could launch KDE/Plasma with any of: export QT_XCB_GL_INTEGRATION=xcb_egl export QT_XCB_GL_INTEGRATION=none But graphics are not polished. KDE/Plasma doesn't work when QT_XCB_GL_INTEGRATION is not set or has: export QT_XCB_GL_INTEGRATION=xcb_glx I tried these different values in .config/plasma-workspace/env/xcb.sh Do you think a rebuild of the nvidia/Opengl driver would fix the issue? Or is Kde/Plasma requiring features not provided by the nvidia driver? It seems to work actually, at least partially. :)