#1126671 linux-image-6.17.13+deb13-amd64: Disconect root (path=/) during heavy load on NvME, partial corruption on NTFS.crash soon but not yet

Package:
src:linux
Source:
src:linux
Submitter:
David Gonzalo Rodriguez
Date:
2026-07-26 12:25:02 UTC
Severity:
normal
Tags:
#1126671#5
Date:
2026-01-30 12:00:09 UTC
From:
To:
Dear Maintainer,

*** Reporter, please consider answering these questions, where appropriate ***

   * What led up to the situation?
   * What exactly did you do (or not do) that was effective (or
     ineffective)?
   * What was the outcome of this action?
   * What outcome did you expect instead?

*** End of the template - remove these template lines ***

#1126671#10
Date:
2026-01-30 15:20:23 UTC
From:
To:
   * What led up to the situation?

sudo apt update
sudo apt upgrade
I get a new Linux-Image and many more things and the system get unstable I
get the same problem on two computers

it's almost a normal Debian installation with Mate and after it:
sudo apt install wine gnome-disk-utility gparted obs-studio qdirstat curl
unrar rar filezilla samba rsync linux-headers-amd64 nvidia-kernel-dkms
nvidia-driver firmware-misc-nonfree nvidia-cuda-dev nvidia-cuda-toolkit git
scummvm ffmpeg mono-complete mono-runtime make cmake lutris
and externally -> Telegram, Discord, Steam, VisualStudio

   * What exactly did you do (or not do) that was effective (or
ineffective)?

I use the NvME drive less at least to transfer large files

   * What was the outcome of this action?

a data loss and both computers reboot at any time, it look like / (root)
was vanish as nothing work from the CLI, I can't launch any program

   * What outcome did you expect instead?

I expect that / (root) was reload after the NvME get to live again or
whatever else happen


I found a easy way to fix the problem of the data lost with a partition
with a bigger size that the one that I need to fix but isn't the way to fix
it, only to recover the files

#1126671#17
Date:
2026-02-02 14:59:38 UTC
From:
To:
============================================
NOTES to understand the rest of the message:
============================================
English isn't my natural language
Dates are as YYYY-MM-DD

all of the partitions I talking are NTFS partitions, is just because if the
partition is corrupted I know how to rescue it more easy

Type of data corruptions
Type 1. the partition can NOT be mount with gnome-disk-utility but it can
be mount with the mount command or if it is added to the /mnt/
Type 2. the partition can NOT be mount with gnome-disk-utility either with
mount on /mnt/ but if the partition is copy into a file with "dd" or
"ddrescue" and you mount it in a loop device you can read the data from it
Type 3. the partition can NOT be mount with gnome-disk-utility, mount or
the loop device trick and the data is lost unless you used a data rescue
software

only Type 1 and 2 are present on this problems, so is possible to read it
back and isn't data corruption but as normal people can not access it's
almost like a data corruption, is needed to use some trick to can read it
I didn't have any Type 3 data corruption yet

at the end of the first report I wrote a lot of extra info too but as is
the first time I made a report I didn't know where to place that info
============================================
The video show the problem, sometimes on the top left corner appear a box
where I enter the root password to allow changes
0:40 the Operating system didn't leave me to unmount the partition because
ffmpeg was creating a miniature because "caja" was open.

it happen the same on two of my physical computers
On the computer on the shop I have (the one that create the report, install
on 2025-01-09):
6.12.48+deb13-amd64 *1
6.12.57+deb13-amd64 *2
6.12.63+deb13-amd64 *3
6.17.13+deb13-amd64 *4  in use

On the computer I have at home: (I install it 2026-11-21)
6.12.41+deb13-amd64 *5
6.12.57+deb13-amd64
6.12.63+deb13-amd64 in use

I have a VM on proxmox that don't show that problem with Debian 12 (I
create it 2026-01-09)
6.1.0-22-amd64
6.1.0-42-amd64 in use


*1 I didn't test it yet
*2 it have the problem
*3 it have the problem
*4 it have the problem, I manually install and this days is the one is use


-. It's not too easy to replicate it, only if I move a lot of data between
one disk and another, usually if the NvME drivers are used
-. mount and unmount any partition does not create this problems and you
can see on the video
-. if I reboot or power off the computer and a external HDD is mount, the
partition get "corrupted" with the Type 2 of corruption (see the start of
the post)
-. data lost when the computer is turn off by software
-. sometimes data lost when the computer crash is Type 1 on internal
devices (I don't know on external devices)
-. data lost when the computer have a silence crash most of the times is
Type 2 on internal an internal devices (when / root disappear)

And maybe one strange thing is:
It happens most of the times with new NTFS partitions (created with
gnome-disk-utility). and I mean new partitions are partitions created
recently, with old partitions created way before I move from Debian 12 to
Debian 13 look like they are not affected by this bug

I can keep the state of that machine frozen in time to can made more test
on it and I have some 40, 100, 120 and (8) 160 GB HDDs to can test almost
anything if will be needed if someone said me a new test

#1126671#22
Date:
2026-02-10 20:22:14 UTC
From:
To:
Hi David,
gather more information by attaching a netconsole. You need a second
device in your lan to recieve the netconsole messages.

Instructions on how to do it are documented in:
https://docs.kernel.org/networking/netconsole.html

To make things handy once it work to send messaged to the remote host,
add (temporary, it can be dropped again after the debug session) a
dropin file /etc/default/grub.d/netconsole.cfg containing the required
settings:

GRUB_CMDLINE_LINUX="$GRUB_CMDLINE_LINUX netconsole=[...]"

and then run update-grub2.

Can you please try that and then wait to trigger the problem and
collect the logs sent via netconsole.

Regards,
Salvatore

#1126671#29
Date:
2026-02-17 09:19:36 UTC
From:
To:
Hi David,

Thanks for providning the information, I'm forwarding it to the Debian
bug, and exceptionally top-posting here. We will have a look.

Regards
Salvatore

p.s.: please have a look at https://people.kernel.org/tglx/notes-about-netiquette

#1126671#36
Date:
2026-02-18 15:56:59 UTC
From:
To:
----- Forwarded message from Z80user <z80user@gmail.com> -----

not sure if that can be related to the incident:
I just move like 20 GB from a NvME device to the same device...
the system was laggy, it take almost 30 seconds to move the next 16 small
files
when it finnish surely the cache (disk or system) and during that time I
see the
message "Ethernet disconected" on the Mate interface and that is what I can
see
on the output of DMESG (The computer still working isn't crashing)
Then I send the "test" to DMESG and I switch to the root user, and I come
back
to the normal user

btw, sorry I can't send much more information but I don't know how to get
more.
=============
I get too many lines of that type but didn't show too many info with
different
number than 64
[76635.140106] kauditd_printk_skb: 64 callbacks suppressed

It's possible to get it more verbose ?
=================
I still have the SSD in this status just if I need to make any other test
on the
computer in the shop, at home I will try to move it into a VM to do more
tests.
I create several VMs with Debian 12.13 13.00 13.01 13.02 and 13.03 to see if
what files change from one version to another and see the status of my
systems

I have a sparse NvME so isn't a big problem for me to keep it like that


Regards
Z80user (David)
=================
that is the DMESG

76039.832396] EXT4-fs (nvme0n1p4): mounted filesystem
ffbe3bff-6ec6-4dce-8ebf-9a76dc5ef19b r/w with ordered data mode. Quota
mode: none.
[76262.026088] EXT4-fs (nvme0n1p4): unmounting filesystem
ffbe3bff-6ec6-4dce-8ebf-9a76dc5ef19b.
[76270.978444] libobs: hotkey [3752]: segfault at e ip 00007fa07f0d844e sp
00007fa057f6c2c0 error 4 in libobs.so.30[e244e,7fa07f01e000+c8000] likely
on CPU 8 (core 10, socket 0)
[76270.978458] Code: 00 00 00 66 66 2e 0f 1f 84 00 00 00 00 00 66 66 2e 0f
1f 84 00 00 00 00 00 0f 1f 00 0f b6 02 89 c1 83 e0 07 c0 e9 03 83 e1 1f
<0f> b6 4c 0b 08 0f a3 c1 40 0f 92 c5 40 84 ed 0f 85 49 ff ff ff 48
[76271.730169] audit: type=1400 audit(1771188917.497:194):
apparmor="ALLOWED" operation="capable" class="cap" profile="Xorg"
pid=534693 comm="Xorg" capability=12  capname="net_admin"
[76271.730423] audit: type=1400 audit(1771188917.497:195):
apparmor="ALLOWED" operation="open" class="file" profile="Xorg"
name="/etc/nvidia/current/nvidia-drm-outputclass.conf" pid=534693
comm="Xorg" requested_mask="r" denied_mask="r" fsuid=0 ouid=0
[76272.520439] audit: type=1400 audit(1771188918.285:196):
apparmor="ALLOWED" operation="open" class="file" profile="Xorg"
name="/proc/driver/nvidia/params" pid=534693 comm="Xorg" requested_mask="r"
denied_mask="r" fsuid=0 ouid=0
[76272.520471] audit: type=1400 audit(1771188918.285:197):
apparmor="ALLOWED" operation="unlink" class="file" profile="Xorg"
name="/dev/char/195:255" pid=534693 comm="Xorg" requested_mask="d"
denied_mask="d" fsuid=0 ouid=0
[76272.520487] audit: type=1400 audit(1771188918.285:198):
apparmor="ALLOWED" operation="symlink" class="file" profile="Xorg"
name="/dev/char/195:255" pid=534693 comm="Xorg" requested_mask="c"
denied_mask="c" fsuid=0 ouid=0
[76272.520501] audit: type=1400 audit(1771188918.285:199):
apparmor="ALLOWED" operation="open" class="file" profile="Xorg"
name="/dev/nvidiactl" pid=534693 comm="Xorg" requested_mask="wr"
denied_mask="wr" fsuid=0 ouid=0
[76272.562265] audit: type=1400 audit(1771188918.329:200):
apparmor="ALLOWED" operation="open" class="file" profile="Xorg"
name="/proc/534693/comm" pid=534693 comm="Xorg" requested_mask="r"
denied_mask="r" fsuid=0 ouid=0
[76272.619984] audit: type=1400 audit(1771188918.385:201):
apparmor="ALLOWED" operation="open" class="file" profile="Xorg"
name="/proc/534693/comm" pid=534693 comm="Xorg" requested_mask="r"
denied_mask="r" fsuid=0 ouid=0
[76272.620138] audit: type=1400 audit(1771188918.385:202):
apparmor="ALLOWED" operation="open" class="file" profile="Xorg"
name="/proc/driver/nvidia/params" pid=534693 comm="Xorg" requested_mask="r"
denied_mask="r" fsuid=0 ouid=0
[76520.913127] test
[76551.982537] EXT4-fs (nvme0n1p4): mounted filesystem
ffbe3bff-6ec6-4dce-8ebf-9a76dc5ef19b r/w with ordered data mode. Quota
mode: none.
[76554.002562] EXT4-fs (nvme0n1p4): unmounting filesystem
ffbe3bff-6ec6-4dce-8ebf-9a76dc5ef19b.
[76635.140106] kauditd_printk_skb: 64 callbacks suppressed
[76635.140110] audit: type=1400 audit(1771189280.905:267):
apparmor="ALLOWED" operation="open" class="file" profile="Xorg"
name="/etc/nvidia/current/nvidia-drm-outputclass.conf" pid=536610
comm="Xorg" requested_mask="r" denied_mask="r" fsuid=0 ouid=0
[76636.239567] audit: type=1400 audit(1771189282.005:268):
apparmor="ALLOWED" operation="open" class="file" profile="Xorg"
name="/proc/driver/nvidia/params" pid=536610 comm="Xorg" requested_mask="r"
denied_mask="r" fsuid=0 ouid=0
[76636.239600] audit: type=1400 audit(1771189282.005:269):
apparmor="ALLOWED" operation="unlink" class="file" profile="Xorg"
name="/dev/char/195:255" pid=536610 comm="Xorg" requested_mask="d"
denied_mask="d" fsuid=0 ouid=0
[76636.239614] audit: type=1400 audit(1771189282.005:270):
apparmor="ALLOWED" operation="symlink" class="file" profile="Xorg"
name="/dev/char/195:255" pid=536610 comm="Xorg" requested_mask="c"
denied_mask="c" fsuid=0 ouid=0
[76636.239650] audit: type=1400 audit(1771189282.005:271):
apparmor="ALLOWED" operation="open" class="file" profile="Xorg"
name="/dev/nvidiactl" pid=536610 comm="Xorg" requested_mask="wr"
denied_mask="wr" fsuid=0 ouid=0
[76636.290660] audit: type=1400 audit(1771189282.057:272):
apparmor="ALLOWED" operation="open" class="file" profile="Xorg"
name="/proc/536610/comm" pid=536610 comm="Xorg" requested_mask="r"
denied_mask="r" fsuid=0 ouid=0
[76636.336794] audit: type=1400 audit(1771189282.105:273):
apparmor="ALLOWED" operation="open" class="file" profile="Xorg"
name="/proc/536610/comm" pid=536610 comm="Xorg" requested_mask="r"
denied_mask="r" fsuid=0 ouid=0
[76636.336944] audit: type=1400 audit(1771189282.105:274):
apparmor="ALLOWED" operation="open" class="file" profile="Xorg"
name="/proc/driver/nvidia/params" pid=536610 comm="Xorg" requested_mask="r"
denied_mask="r" fsuid=0 ouid=0
[76636.336975] audit: type=1400 audit(1771189282.105:275):
apparmor="ALLOWED" operation="unlink" class="file" profile="Xorg"
name="/dev/char/195:0" pid=536610 comm="Xorg" requested_mask="d"
denied_mask="d" fsuid=0 ouid=0
[76636.336989] audit: type=1400 audit(1771189282.105:276):
apparmor="ALLOWED" operation="symlink" class="file" profile="Xorg"
name="/dev/char/195:0" pid=536610 comm="Xorg" requested_mask="c"
denied_mask="c" fsuid=0 ouid=0
----- End forwarded message -----
#1126671#41
Date:
2026-02-25 19:48:17 UTC
From:
To:
Hello David,

To gather some information if this might be HW failing, can you please
provie as well the output of

nvme smart-log /dev/nvme0n1
nvme error-log /dev/nvme0n1

please?

Again, please include in replies the Debian bug address itself.
https://people.kernel.org/tglx/notes-about-netiquette

Regards,
Salvatore

#1126671#46
Date:
2026-02-27 13:25:29 UTC
From:
To:
Hi,

I mean to keep 1126671@bugs.debian.org in the recipients.

Thanks for the logs, I'm forwarding as a first step to the Debian bug.

Regards,
Salvatore

#1126671#55
Date:
2026-07-26 12:24:14 UTC
From:
To:
Hi,

Sorry that we've not responded to your bug report for a while.

I'm mostly interested in investigating the first I/O errors on the NVMe
drives.  I think that later problems, like with the network, are only
symptoms of the OS being unable to write to the disk.

Are you still seeing this type of I/O error?

I'm trying to understand which information comes from which computers.
Firstly there is "the computer on the shop I have (the one that create
the report, install on 2025-01-09)".  In that report I can see:

motherboard:  Gigabyte A520M DS3H V2
CPU:          AMD Ryzen 3000/4000 series
NVMe drive 0: Kingston Technology Company, Inc. NV1 NVMe SSD [SM2263XT] (DRAM-less)
NVMe drive 1: KIOXIA Corporation NVMe SSD Controller BG5 (DRAM-less)

The crash log shows:

motherboard:  ASUS PRIME X570-PRO
CPU:          AMD Ryzen 9 3900X
NVME drive 0: KIOXIA, unknown model

Is that from "the computer I have at home: (I install it 2026-11-21)"?

Please check whether the drives showing errors have the latest firmware
revision.  You can read the current revision from
/sys/class/nvme/nvme0/firmware_rev or /sys/class/nvme/nvme1/firmware_rev
(depending on the device name).  It looks like both those drives have
KIOXIA controllers, and firmware updates for KIOXIA drives are available
at <https://europe.kioxia.com/es-es/personal/support/download/ssd.html>.
If the drives were sold by another manufacturer you'll have to look at
that manufacturer's web site.

Ben.