#1116065 linux: kernel oops with rsync on MSI X99A with ntfs3

Package:
src:linux
Source:
src:linux
Submitter:
Antwerpen, G. (Gert) van
Date:
2025-09-29 08:37:02 UTC
Severity:
normal
Tags:
#1116065#5
Date:
2025-09-23 17:59:30 UTC
From:
To:
Summary:
Kernel oops / crash when running rsync on MSI X99A SLI PLUS (MS-7885).
System was stable for years with Debian stable kernel 6.1.x.
The problem only appears after upgrading to Trixie (kernel 6.12 and newer).

System information:
- Machine: MSI X99A SLI PLUS (MS-7885)
- BIOS: 1.D0 (07/15/2016)
- CPU: Intel Xeon (Haswell-E, socket 2011-3)
- RAM: [please fill in, e.g. 64 GB DDR4 ECC/non-ECC]
- Debian version: Trixie (testing)
- Kernel versions tested:
  - 6.1.x (Debian bookworm stable) → works fine
  - 6.12.0-1 (Debian Trixie) → oops/crash
  - 6.16.3-1~bpo13+1 (Debian Trixie backports) → oops/crash

Kernel log excerpt:
BUG: unable to handle page fault for address: ffffffaabfe73e00
#PF: supervisor instruction fetch in kernel mode
#PF: error_code(0x0010) - not-present page
PGD 1de231067 P4D 1de231067 PUD 0
Oops: 0010 [#2] SMP PTI
CPU: 9 UID: 40001 PID: 230866 Comm: rsync Tainted: G      D             6.16.3+deb13-amd64 #1 PREEMPT(lazy)  Debian 6.16.3-1~bpo13+1
Hardware name: MSI MS-7885/X99A SLI PLUS(MS-7885), BIOS 1.D0 07/15/2016
RIP: 0010:0xffffffaabfe73e00
Code: Unable to access opcode bytes at 0xffffffaabfe73dd6.
Call Trace:
 filemap_readahead.isra.0+0x75/0xb0
 filemap_get_pages+0x3ed/0x770
 sock_write_iter+0x18e/0x1a0
 ...
 note: rsync[230866] exited with irqs disabled

(Full logs can be provided if required.)

Steps to reproduce:
1. Run rsync on large data sets (local disk to remote).
2. After some time, system crashes with kernel oops (see logs above).
3. Always reproducible on kernel >= 6.12, never seen on 6.1.

Expected result:
No kernel oops — rsync should run reliably.

Actual result:
Kernel crashes with page fault in kernel mode, requiring system restart.

Additional notes:
- Hardware tested with memtest86+ (no errors).
- No overclocking.
- Issue seems to be a regression introduced in Linux 6.12.
- Possibly related to filesystem or networking modules, but exact trigger unknown.
-- This message may contain information that is not intended for you. If you are not the addressee or if this message was sent to you by mistake, you are requested to inform the sender and delete the message. TNO accepts no liability for the content of this e-mail, for the manner in which you use it and for damage of any kind resulting from the risks inherent to the electronic transmission of messages.

#1116065#14
Date:
2025-09-23 19:01:11 UTC
From:
To:
Hi,
you do not get access to the machine after oops'ing the you might
attach a netconsole to get the relevant logs.

Additionally to the logs ideally you provide all the meta information
collected by running reportbug's bugscripts for the kernel reports.

As for the regreesion itself and identify the breaking commit: Can you
bisect the upstream changes. Ideallally you first can range bit closer
the upstream versions where it is regressing. You can use for that the
snapshot.debian.org service to fetch older linux-image versions. Once
you have  close enough range, then bisect the upstream changes (would
you need help and have instructions to do that?).

YOu might want to drop this when filling a public bugreport ;-)

Regards,
Salvatore

#1116065#21
Date:
2025-09-24 07:44:24 UTC
From:
To:
Hi Gert,

Okay that makes things more difficult.

Unfortunately that is not even the first Ooops. So there should be
more earlier already in the log. Are you able to provide the full
kernel log from once the problem is happening?

If the system is not accessible anymore, then consider attaching a
netconsole so we get logs from as early as possible from boot and then
logged over the network.

The full log would be better, if not, then at least we should get the
first oops and see from there.

This sesm quite an old system, nwer BIOS version seems available still
slightly more recent taht the 07/15/2016 version, so you might
consider updating it.

Regards,
Salvatore

#1116065#26
Date:
2025-09-24 18:40:51 UTC
From:
To:
----- Forwarded message from "Antwerpen, G. (Gert) van" <gert.vanantwerpen@tno.nl> -----

Okay, reproducing is a bit complex because it seems only to happen in a certain combination of things.

But, here is the first Oops for you.

Maybe this helps you a bit.

It seems not to be in "rsync" but in a wine process named "split_lcms_64.exe".

The "rsync" is about 20 minutes later.

Gert



BUG: unable to handle page fault for address: ffffffaabfe73e00

#PF: supervisor instruction fetch in kernel mode

#PF: error_code(0x0010) - not-present page

PGD 1de231067 P4D 1de231067 PUD 0

Oops: Oops: 0010 [#1] SMP PTI

CPU: 5 UID: 40001 PID: 54760 Comm: split_lcms_64.e Not tainted 6.16.3+deb13-amd64 #1 PREEMPT(lazy)  Debian 6.16.3-1~bpo13+1

Hardware name: MSI MS-7885/X99A SLI PLUS(MS-7885), BIOS 1.D0 07/15/2016

RIP: 0010:0xffffffaabfe73e00

Code: Unable to access opcode bytes at 0xffffffaabfe73dd6.

RSP: 0018:ffffcf96c884fa87 EFLAGS: 00010246

RAX: 0000000000000000 RBX: 0000000000010245 RCX: 000000000000000a

RDX: ffff88dad2d13580 RSI: 0000000000000002 RDI: ffff88dad2d13580

RBP: 00000000000714ff R08: 0000000000000001 R09: 0000000000002000

R10: 000000002f701000 R11: 0000000000000011 R12: ffcf96c884fb5800

R13: fff91bdc103a80ff R14: ffff88dcbacb7040 R15: 0000000000000613

FS:  00007f8172721740(0000) GS:ffff88e2b2908000(0000) knlGS:000000007ffc0000

CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033

CR2: ffffffaabfe73dd6 CR3: 000000000724c004 CR4: 00000000003726f0

Call Trace:

<TASK>

? filemap_get_pages+0x13f/0x770

? filemap_read+0xff/0x440

? aa_file_perm+0x121/0x4d0

? vfs_read+0x25d/0x370

? ksys_read+0x6b/0xe0

? do_syscall_64+0x84/0x2f0

? count_memcg_events+0x167/0x1d0

? handle_mm_fault+0x1d7/0x2e0

? do_user_addr_fault+0x2c3/0x7f0

? entry_SYSCALL_64_after_hwframe+0x76/0x7e

</TASK>

Modules linked in: tun overlay ntfs3 snd_seq_dummy snd_hrtimer snd_seq snd_seq_device rfkill qrtr nft_chain_nat xt_MASQUERADE nf_nat xt_addrtype xt_tcpudp xt_conntrack nf_conntrack sunrpc nf_defrag_ipv6 nf_defrag_ipv4 nft_compat nf_tables binfmt_misc nls_ascii nls_cp437 vfat fat intel_rapl_msr intel_rapl_common intel_uncore_frequency intel_uncore_frequency_common x86_pkg_temp_thermal intel_powerclamp coretemp kvm_intel snd_hda_codec_realtek kvm snd_hda_codec_generic snd_hda_scodec_component snd_hda_codec_hdmi snd_hda_intel irqbypass snd_intel_dspcfg ghash_clmulni_intel snd_intel_sdw_acpi sha512_ssse3 sha1_ssse3 snd_hda_codec aesni_intel rapl snd_hda_core intel_cstate snd_hwdep intel_uncore pcspkr snd_pcm mei_me mei snd_timer snd soundcore sg evdev parport_pc ppdev lp parport configfs efi_pstore nfnetlink efivarfs ip_tables x_tables autofs4 crc32c_cryptoapi btrfs blake2b_generic xor raid6_pq hid_generic nouveau usbhid hid drm_gpuvm gpu_sched drm_ttm_helper ttm video drm_exec sd_mod i2c_algo_bit

drm_display_helper cec rc_core drm_client_lib drm_kms_helper ahci xhci_pci xhci_hcd libahci ehci_pci drm ehci_hcd libata iTCO_wdt usbcore intel_pmc_bxt iTCO_vendor_support watchdog e1000e mxm_wmi scsi_mod i2c_i801 lpc_ich i2c_smbus usb_common scsi_common wmi button

CR2: ffffffaabfe73e00
---[ end trace 0000000000000000 ]--- RIP: 0010:0xffffffaabfe73e00 Code: Unable to access opcode bytes at 0xffffffaabfe73dd6. RSP: 0018:ffffcf96c884fa87 EFLAGS: 00010246 RAX: 0000000000000000 RBX: 0000000000010245 RCX: 000000000000000a RDX: ffff88dad2d13580 RSI: 0000000000000002 RDI: ffff88dad2d13580 RBP: 00000000000714ff R08: 0000000000000001 R09: 0000000000002000 R10: 000000002f701000 R11: 0000000000000011 R12: ffcf96c884fb5800 R13: fff91bdc103a80ff R14: ffff88dcbacb7040 R15: 0000000000000613 FS: 00007f8172721740(0000) GS:ffff88e2b2908000(0000) knlGS:000000007ffc0000 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 CR2: ffffffaabfe73dd6 CR3: 000000000724c004 CR4: 00000000003726f0 note: split_lcms_64.e[54760] exited with irqs disabled
----- End forwarded message -----
#1116065#31
Date:
2025-09-26 04:25:40 UTC
From:
To:
Hi Gert,

Yes indeed that is very usefull addition.

It would really be helpful if you can with the ntfs3 module loaded,
triggering the problem and getting to us the *full* kernel log. In
your mail before at least we have now the first OOPS, that is as well
already of help.

If there are some information youe need to anonymize (hostname), that
is not a problem, but getting the log from before the issue triggers,
to just before including the context would be very helpful.

Mentioning ntfs3 slightly rings bell here, need to search for the
earlier repor.

Regards,
Salvatore

p.s.: please keep the Debian bugreport address as well in the loop, so
I'm not the only one getting your reply and people can chime in with
help.

#1116065#36
Date:
2025-09-26 04:48:35 UTC
From:
To:
Hi

here is a very similar report in upstream's bugzilla:
https://bugzilla.kernel.org/show_bug.cgi?format=multiple&id=215460
(though it mentions to be fixed already, or occuring more rarely)

I will drop "regression since 6.12" from the subject, since it is
related to using ntfs3 driver in kernel, which was not available
before, so it's not directly a kernel regression we can bisect from
before situation.

Regards,
Salvatore

#1116065#41
Date:
2025-09-26 06:17:20 UTC
From:
To:
I put the full syslog of the system (from boot until reboot) where the problems happened, into the attached zipfile.
This is based on the trixie-backports kernel (6.16.3+deb13-amd64).
If you want, I can also give you the syslog of the run using the default trixie kernel (6.12.43+deb13-amd64).
Gert

Hi Gert,

Yes indeed that is very usefull addition.

It would really be helpful if you can with the ntfs3 module loaded,
triggering the problem and getting to us the *full* kernel log. In
your mail before at least we have now the first OOPS, that is as well
already of help.

If there are some information youe need to anonymize (hostname), that
is not a problem, but getting the log from before the issue triggers,
to just before including the context would be very helpful.

Mentioning ntfs3 slightly rings bell here, need to search for the
earlier repor.

Regards,
Salvatore

p.s.: please keep the Debian bugreport address as well in the loop, so
I'm not the only one getting your reply and people can chime in with
help.
-- This message may contain information that is not intended for you. If you are not the addressee or if this message was sent to you by mistake, you are requested to inform the sender and delete the message. TNO accepts no liability for the content of this e-mail, for the manner in which you use it and for damage of any kind resulting from the risks inherent to the electronic transmission of messages.

#1116065#48
Date:
2025-09-29 08:35:28 UTC
From:
To:
I tried to reproduce the problem on another computer (same amount of memory, identical files and processes) and was unable to reproduce it.
The only difference between the two setups is that the original is based on SSD's and the replica is using HDD's.
So the problem seems only occur when using SSD's (or, at least the chance is much higher).
GvA
-- This message may contain information that is not intended for you. If you are not the addressee or if this message was sent to you by mistake, you are requested to inform the sender and delete the message. TNO accepts no liability for the content of this e-mail, for the manner in which you use it and for damage of any kind resulting from the risks inherent to the electronic transmission of messages.