Dear Maintainer,
*** Reporter, please consider answering these questions, where appropriate ***
* What led up to the situation?
* What exactly did you do (or not do) that was effective (or
ineffective)?
* What was the outcome of this action?
* What outcome did you expect instead?
*** End of the template - remove these template lines ***
* What led up to the situation? sudo apt update sudo apt upgrade I get a new Linux-Image and many more things and the system get unstable I get the same problem on two computers it's almost a normal Debian installation with Mate and after it: sudo apt install wine gnome-disk-utility gparted obs-studio qdirstat curl unrar rar filezilla samba rsync linux-headers-amd64 nvidia-kernel-dkms nvidia-driver firmware-misc-nonfree nvidia-cuda-dev nvidia-cuda-toolkit git scummvm ffmpeg mono-complete mono-runtime make cmake lutris and externally -> Telegram, Discord, Steam, VisualStudio * What exactly did you do (or not do) that was effective (or ineffective)? I use the NvME drive less at least to transfer large files * What was the outcome of this action? a data loss and both computers reboot at any time, it look like / (root) was vanish as nothing work from the CLI, I can't launch any program * What outcome did you expect instead? I expect that / (root) was reload after the NvME get to live again or whatever else happen I found a easy way to fix the problem of the data lost with a partition with a bigger size that the one that I need to fix but isn't the way to fix it, only to recover the files
============================================ NOTES to understand the rest of the message: ============================================ English isn't my natural language Dates are as YYYY-MM-DD all of the partitions I talking are NTFS partitions, is just because if the partition is corrupted I know how to rescue it more easy Type of data corruptions Type 1. the partition can NOT be mount with gnome-disk-utility but it can be mount with the mount command or if it is added to the /mnt/ Type 2. the partition can NOT be mount with gnome-disk-utility either with mount on /mnt/ but if the partition is copy into a file with "dd" or "ddrescue" and you mount it in a loop device you can read the data from it Type 3. the partition can NOT be mount with gnome-disk-utility, mount or the loop device trick and the data is lost unless you used a data rescue software only Type 1 and 2 are present on this problems, so is possible to read it back and isn't data corruption but as normal people can not access it's almost like a data corruption, is needed to use some trick to can read it I didn't have any Type 3 data corruption yet at the end of the first report I wrote a lot of extra info too but as is the first time I made a report I didn't know where to place that info ============================================ The video show the problem, sometimes on the top left corner appear a box where I enter the root password to allow changes 0:40 the Operating system didn't leave me to unmount the partition because ffmpeg was creating a miniature because "caja" was open. it happen the same on two of my physical computers On the computer on the shop I have (the one that create the report, install on 2025-01-09): 6.12.48+deb13-amd64 *1 6.12.57+deb13-amd64 *2 6.12.63+deb13-amd64 *3 6.17.13+deb13-amd64 *4 in use On the computer I have at home: (I install it 2026-11-21) 6.12.41+deb13-amd64 *5 6.12.57+deb13-amd64 6.12.63+deb13-amd64 in use I have a VM on proxmox that don't show that problem with Debian 12 (I create it 2026-01-09) 6.1.0-22-amd64 6.1.0-42-amd64 in use *1 I didn't test it yet *2 it have the problem *3 it have the problem *4 it have the problem, I manually install and this days is the one is use -. It's not too easy to replicate it, only if I move a lot of data between one disk and another, usually if the NvME drivers are used -. mount and unmount any partition does not create this problems and you can see on the video -. if I reboot or power off the computer and a external HDD is mount, the partition get "corrupted" with the Type 2 of corruption (see the start of the post) -. data lost when the computer is turn off by software -. sometimes data lost when the computer crash is Type 1 on internal devices (I don't know on external devices) -. data lost when the computer have a silence crash most of the times is Type 2 on internal an internal devices (when / root disappear) And maybe one strange thing is: It happens most of the times with new NTFS partitions (created with gnome-disk-utility). and I mean new partitions are partitions created recently, with old partitions created way before I move from Debian 12 to Debian 13 look like they are not affected by this bug I can keep the state of that machine frozen in time to can made more test on it and I have some 40, 100, 120 and (8) 160 GB HDDs to can test almost anything if will be needed if someone said me a new test
Hi David, gather more information by attaching a netconsole. You need a second device in your lan to recieve the netconsole messages. Instructions on how to do it are documented in: https://docs.kernel.org/networking/netconsole.html To make things handy once it work to send messaged to the remote host, add (temporary, it can be dropped again after the debug session) a dropin file /etc/default/grub.d/netconsole.cfg containing the required settings: GRUB_CMDLINE_LINUX="$GRUB_CMDLINE_LINUX netconsole=[...]" and then run update-grub2. Can you please try that and then wait to trigger the problem and collect the logs sent via netconsole. Regards, Salvatore
Hi David, Thanks for providning the information, I'm forwarding it to the Debian bug, and exceptionally top-posting here. We will have a look. Regards Salvatore p.s.: please have a look at https://people.kernel.org/tglx/notes-about-netiquette
----- Forwarded message from Z80user <z80user@gmail.com> ----- not sure if that can be related to the incident: I just move like 20 GB from a NvME device to the same device... the system was laggy, it take almost 30 seconds to move the next 16 small files when it finnish surely the cache (disk or system) and during that time I see the message "Ethernet disconected" on the Mate interface and that is what I can see on the output of DMESG (The computer still working isn't crashing) Then I send the "test" to DMESG and I switch to the root user, and I come back to the normal user btw, sorry I can't send much more information but I don't know how to get more. ============= I get too many lines of that type but didn't show too many info with different number than 64 [76635.140106] kauditd_printk_skb: 64 callbacks suppressed It's possible to get it more verbose ? ================= I still have the SSD in this status just if I need to make any other test on the computer in the shop, at home I will try to move it into a VM to do more tests. I create several VMs with Debian 12.13 13.00 13.01 13.02 and 13.03 to see if what files change from one version to another and see the status of my systems I have a sparse NvME so isn't a big problem for me to keep it like that Regards Z80user (David) ================= that is the DMESG 76039.832396] EXT4-fs (nvme0n1p4): mounted filesystem ffbe3bff-6ec6-4dce-8ebf-9a76dc5ef19b r/w with ordered data mode. Quota mode: none. [76262.026088] EXT4-fs (nvme0n1p4): unmounting filesystem ffbe3bff-6ec6-4dce-8ebf-9a76dc5ef19b. [76270.978444] libobs: hotkey [3752]: segfault at e ip 00007fa07f0d844e sp 00007fa057f6c2c0 error 4 in libobs.so.30[e244e,7fa07f01e000+c8000] likely on CPU 8 (core 10, socket 0) [76270.978458] Code: 00 00 00 66 66 2e 0f 1f 84 00 00 00 00 00 66 66 2e 0f 1f 84 00 00 00 00 00 0f 1f 00 0f b6 02 89 c1 83 e0 07 c0 e9 03 83 e1 1f <0f> b6 4c 0b 08 0f a3 c1 40 0f 92 c5 40 84 ed 0f 85 49 ff ff ff 48 [76271.730169] audit: type=1400 audit(1771188917.497:194): apparmor="ALLOWED" operation="capable" class="cap" profile="Xorg" pid=534693 comm="Xorg" capability=12 capname="net_admin" [76271.730423] audit: type=1400 audit(1771188917.497:195): apparmor="ALLOWED" operation="open" class="file" profile="Xorg" name="/etc/nvidia/current/nvidia-drm-outputclass.conf" pid=534693 comm="Xorg" requested_mask="r" denied_mask="r" fsuid=0 ouid=0 [76272.520439] audit: type=1400 audit(1771188918.285:196): apparmor="ALLOWED" operation="open" class="file" profile="Xorg" name="/proc/driver/nvidia/params" pid=534693 comm="Xorg" requested_mask="r" denied_mask="r" fsuid=0 ouid=0 [76272.520471] audit: type=1400 audit(1771188918.285:197): apparmor="ALLOWED" operation="unlink" class="file" profile="Xorg" name="/dev/char/195:255" pid=534693 comm="Xorg" requested_mask="d" denied_mask="d" fsuid=0 ouid=0 [76272.520487] audit: type=1400 audit(1771188918.285:198): apparmor="ALLOWED" operation="symlink" class="file" profile="Xorg" name="/dev/char/195:255" pid=534693 comm="Xorg" requested_mask="c" denied_mask="c" fsuid=0 ouid=0 [76272.520501] audit: type=1400 audit(1771188918.285:199): apparmor="ALLOWED" operation="open" class="file" profile="Xorg" name="/dev/nvidiactl" pid=534693 comm="Xorg" requested_mask="wr" denied_mask="wr" fsuid=0 ouid=0 [76272.562265] audit: type=1400 audit(1771188918.329:200): apparmor="ALLOWED" operation="open" class="file" profile="Xorg" name="/proc/534693/comm" pid=534693 comm="Xorg" requested_mask="r" denied_mask="r" fsuid=0 ouid=0 [76272.619984] audit: type=1400 audit(1771188918.385:201): apparmor="ALLOWED" operation="open" class="file" profile="Xorg" name="/proc/534693/comm" pid=534693 comm="Xorg" requested_mask="r" denied_mask="r" fsuid=0 ouid=0 [76272.620138] audit: type=1400 audit(1771188918.385:202): apparmor="ALLOWED" operation="open" class="file" profile="Xorg" name="/proc/driver/nvidia/params" pid=534693 comm="Xorg" requested_mask="r" denied_mask="r" fsuid=0 ouid=0 [76520.913127] test [76551.982537] EXT4-fs (nvme0n1p4): mounted filesystem ffbe3bff-6ec6-4dce-8ebf-9a76dc5ef19b r/w with ordered data mode. Quota mode: none. [76554.002562] EXT4-fs (nvme0n1p4): unmounting filesystem ffbe3bff-6ec6-4dce-8ebf-9a76dc5ef19b. [76635.140106] kauditd_printk_skb: 64 callbacks suppressed [76635.140110] audit: type=1400 audit(1771189280.905:267): apparmor="ALLOWED" operation="open" class="file" profile="Xorg" name="/etc/nvidia/current/nvidia-drm-outputclass.conf" pid=536610 comm="Xorg" requested_mask="r" denied_mask="r" fsuid=0 ouid=0 [76636.239567] audit: type=1400 audit(1771189282.005:268): apparmor="ALLOWED" operation="open" class="file" profile="Xorg" name="/proc/driver/nvidia/params" pid=536610 comm="Xorg" requested_mask="r" denied_mask="r" fsuid=0 ouid=0 [76636.239600] audit: type=1400 audit(1771189282.005:269): apparmor="ALLOWED" operation="unlink" class="file" profile="Xorg" name="/dev/char/195:255" pid=536610 comm="Xorg" requested_mask="d" denied_mask="d" fsuid=0 ouid=0 [76636.239614] audit: type=1400 audit(1771189282.005:270): apparmor="ALLOWED" operation="symlink" class="file" profile="Xorg" name="/dev/char/195:255" pid=536610 comm="Xorg" requested_mask="c" denied_mask="c" fsuid=0 ouid=0 [76636.239650] audit: type=1400 audit(1771189282.005:271): apparmor="ALLOWED" operation="open" class="file" profile="Xorg" name="/dev/nvidiactl" pid=536610 comm="Xorg" requested_mask="wr" denied_mask="wr" fsuid=0 ouid=0 [76636.290660] audit: type=1400 audit(1771189282.057:272): apparmor="ALLOWED" operation="open" class="file" profile="Xorg" name="/proc/536610/comm" pid=536610 comm="Xorg" requested_mask="r" denied_mask="r" fsuid=0 ouid=0 [76636.336794] audit: type=1400 audit(1771189282.105:273): apparmor="ALLOWED" operation="open" class="file" profile="Xorg" name="/proc/536610/comm" pid=536610 comm="Xorg" requested_mask="r" denied_mask="r" fsuid=0 ouid=0 [76636.336944] audit: type=1400 audit(1771189282.105:274): apparmor="ALLOWED" operation="open" class="file" profile="Xorg" name="/proc/driver/nvidia/params" pid=536610 comm="Xorg" requested_mask="r" denied_mask="r" fsuid=0 ouid=0 [76636.336975] audit: type=1400 audit(1771189282.105:275): apparmor="ALLOWED" operation="unlink" class="file" profile="Xorg" name="/dev/char/195:0" pid=536610 comm="Xorg" requested_mask="d" denied_mask="d" fsuid=0 ouid=0 [76636.336989] audit: type=1400 audit(1771189282.105:276): apparmor="ALLOWED" operation="symlink" class="file" profile="Xorg" name="/dev/char/195:0" pid=536610 comm="Xorg" requested_mask="c" denied_mask="c" fsuid=0 ouid=0----- End forwarded message -----
Hello David, To gather some information if this might be HW failing, can you please provie as well the output of nvme smart-log /dev/nvme0n1 nvme error-log /dev/nvme0n1 please? Again, please include in replies the Debian bug address itself. https://people.kernel.org/tglx/notes-about-netiquette Regards, Salvatore
Hi, I mean to keep 1126671@bugs.debian.org in the recipients. Thanks for the logs, I'm forwarding as a first step to the Debian bug. Regards, Salvatore
Hi, Sorry that we've not responded to your bug report for a while. I'm mostly interested in investigating the first I/O errors on the NVMe drives. I think that later problems, like with the network, are only symptoms of the OS being unable to write to the disk. Are you still seeing this type of I/O error? I'm trying to understand which information comes from which computers. Firstly there is "the computer on the shop I have (the one that create the report, install on 2025-01-09)". In that report I can see: motherboard: Gigabyte A520M DS3H V2 CPU: AMD Ryzen 3000/4000 series NVMe drive 0: Kingston Technology Company, Inc. NV1 NVMe SSD [SM2263XT] (DRAM-less) NVMe drive 1: KIOXIA Corporation NVMe SSD Controller BG5 (DRAM-less) The crash log shows: motherboard: ASUS PRIME X570-PRO CPU: AMD Ryzen 9 3900X NVME drive 0: KIOXIA, unknown model Is that from "the computer I have at home: (I install it 2026-11-21)"? Please check whether the drives showing errors have the latest firmware revision. You can read the current revision from /sys/class/nvme/nvme0/firmware_rev or /sys/class/nvme/nvme1/firmware_rev (depending on the device name). It looks like both those drives have KIOXIA controllers, and firmware updates for KIOXIA drives are available at <https://europe.kioxia.com/es-es/personal/support/download/ssd.html>. If the drives were sold by another manufacturer you'll have to look at that manufacturer's web site. Ben.