#911118 mhddfs hangs for a while on writing

Package:
mhddfs
Source:
mhddfs
Description:
file system for unifying several mount points into one
Submitter:
Marcél Ströhle
Date:
2024-11-17 07:51:01 UTC
Severity:
normal
#911118#5
Date:
2018-10-15 20:40:58 UTC
From:
To:
When I initiate a samba share file transfer to a drive mounted through
MHDDFS, it takes many seconds to start writing, and I observe high MHDDFS
CPU usage. Reading a file works perfectly fine. Local writes are also fine.
The length of the delay seems to scale with the size of the file I want to
write. In the Log, when the client waits for the transfer to start, this
message shows up many times:

mhddfs [2018-10-15 01:03:37]: [140267440191232] mhdd_getxattr: path =
/mnt/hd1/Folder/File-to-Copy.mkv name = security.capability bufsize = 0

Copying the same file to the same location through a direct Samba share is
fine.
This seems to happen with Windows clients (7 & 10 tested), and not with the
other Ubuntu machine I have.

best regards
Marcél

ProblemType: Bug
ApportVersion: 2.20.9-0ubuntu7.4
Architecture: amd64
CurrentDesktop: ubuntu:GNOME
Date: Mon Oct 15 01:24:54 2018
Dependencies:
 adduser 3.116ubuntu1
 apt 1.6.3ubuntu0.1
 apt-utils 1.6.3ubuntu0.1
 ca-certificates 20180409
 debconf 1.5.66
 debconf-i18n 1.5.66
 dpkg 1.19.0.5ubuntu2
 fdisk 2.31.1-0.4ubuntu3.1
 fuse 2.9.7-1ubuntu1
 gcc-8-base 8.2.0-1ubuntu2~18.04
 gpgv 2.2.4-1ubuntu1.1
 libacl1 2.2.52-3build1
 libapt-inst2.0 1.6.3ubuntu0.1
 libapt-pkg5.0 1.6.3ubuntu0.1
 libattr1 1:2.4.47-2build1
 libaudit-common 1:2.8.2-1ubuntu1
 libaudit1 1:2.8.2-1ubuntu1
 libblkid1 2.31.1-0.4ubuntu3.1
 libbz2-1.0 1.0.6-8.1
 libc6 2.27-3ubuntu1
 libcap-ng0 0.7.7-3.1
 libdb5.3 5.3.28-13.1ubuntu1
 libfdisk1 2.31.1-0.4ubuntu3.1
 libffi6 3.2.1-8
 libfuse2 2.9.7-1ubuntu1
 libgcc1 1:8.2.0-1ubuntu2~18.04
 libgcrypt20 1.8.1-4ubuntu1.1
 libgmp10 2:6.1.2+dfsg-2
 libgnutls30 3.5.18-1ubuntu1
 libgpg-error0 1.27-6
 libgpm2 1.20.7-5
 libhogweed4 3.4-1
 libidn2-0 2.0.4-1.1build2
 liblocale-gettext-perl 1.07-3build2
 liblz4-1 0.0~r131-2ubuntu3
 liblzma5 5.2.2-1.3
 libmount1 2.31.1-0.4ubuntu3.1
 libncursesw5 6.1-1ubuntu1.18.04
 libnettle6 3.4-1
 libp11-kit0 0.23.9-2
 libpam-modules 1.1.8-3.6ubuntu2
 libpam-modules-bin 1.1.8-3.6ubuntu2
 libpam0g 1.1.8-3.6ubuntu2
 libpcre3 2:8.39-9
 libseccomp2 2.3.1-2.1ubuntu4
 libselinux1 2.7-2build2
 libsemanage-common 2.7-2build2
 libsemanage1 2.7-2build2
 libsepol1 2.7-1
 libsmartcols1 2.31.1-0.4ubuntu3.1
 libssl1.1 1.1.0g-2ubuntu4.1
 libstdc++6 8.2.0-1ubuntu2~18.04
 libsystemd0 237-3ubuntu10.3
 libtasn1-6 4.13-2
 libtext-charwidth-perl 0.04-7.1
 libtext-iconv-perl 1.7-5build6
 libtext-wrapi18n-perl 0.06-7.1
 libtinfo5 6.1-1ubuntu1.18.04
 libudev1 237-3ubuntu10.3
 libunistring2 0.9.9-0ubuntu1
 libuuid1 2.31.1-0.4ubuntu3.1
 libzstd1 1.3.3+dfsg-2ubuntu1
 mount 2.31.1-0.4ubuntu3.1
 openssl 1.1.0g-2ubuntu4.1
 passwd 1:4.5-1ubuntu1
 perl-base 5.26.1-6ubuntu0.2
 sed 4.4-2
 tar 1.29b-2
 ubuntu-keyring 2018.02.28
 util-linux 2.31.1-0.4ubuntu3.1
 uuid-runtime 2.31.1-0.4ubuntu3.1
 zlib1g 1:1.2.11.dfsg-0ubuntu2
DistroRelease: Ubuntu 18.04
InstallationDate: Installed on 2018-09-28 (16 days ago)
InstallationMedia: Ubuntu 18.04 LTS "Bionic Beaver" - Release amd64
(20180426)
Package: mhddfs 0.1.39+nmu1ubuntu2
PackageArchitecture: amd64
ProcCpuinfoMinimal:
 processor : 7
 vendor_id : GenuineIntel
 cpu family : 6
 model : 58
 model name : Intel(R) Core(TM) i7-3770 CPU @ 3.40GHz
 stepping : 9
 microcode : 0x20
 cpu MHz : 3531.958
 cache size : 8192 KB
 physical id : 0
 siblings : 8
 core id : 3
 cpu cores : 4
 apicid : 7
 initial apicid : 7
 fpu : yes
 fpu_exception : yes
 cpuid level : 13
 wp : yes
 flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat
pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx rdtscp lm
constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc
cpuid aperfmperf pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3
cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic popcnt tsc_deadline_timer aes
xsave avx f16c rdrand lahf_lm cpuid_fault epb pti ssbd ibrs ibpb stibp
tpr_shadow vnmi flexpriority ept vpid fsgsbase smep erms xsaveopt dtherm
ida arat pln pts flush_l1d
 bugs : cpu_meltdown spectre_v1 spectre_v2 spec_store_bypass l1tf
 bogomips : 6800.35
 clflush size : 64
 cache_alignment : 64
 address sizes : 36 bits physical, 48 bits virtual
 power management:
ProcEnviron:
 LC_TIME=de_AT.UTF-8
 LC_MONETARY=de_AT.UTF-8
 TERM=xterm-256color
 PATH=(custom, no user)
 LC_ADDRESS=de_AT.UTF-8
 XDG_RUNTIME_DIR=<set>
 LANG=en_US.UTF-8
 LC_TELEPHONE=de_AT.UTF-8
 LC_NAME=de_AT.UTF-8
 SHELL=/bin/bash
 LC_MEASUREMENT=de_AT.UTF-8
 LC_IDENTIFICATION=de_AT.UTF-8
 LC_NUMERIC=de_AT.UTF-8
 LC_PAPER=de_AT.UTF-8
ProcVersionSignature: Ubuntu 4.15.0-36.39-generic 4.15.18
SourcePackage: mhddfs
Tags:  bionic
Uname: Linux 4.15.0-36-generic x86_64
UpgradeStatus: No upgrade log present (probably fresh install)
_MarkForUpload: True

#911118#10
Date:
2018-10-15 21:57:26 UTC
From:
To:
I must have overlooked one setting in Samba when I was trying to eliminate
its configuration as a source for this issue:
strict allocate = yes
This seems to be what caused the problem. Sorry for the falsely attributed
bug report, but maybe it's interesting information - I still think it's
weird that this only caused an issue with FUSE/MHDDFS, not the other share
I was using.

#911118#21
Date:
2024-11-15 16:39:59 UTC
From:
To:
Control: retitle -1 smbd: please "strict allocate" smarter than 1 byte at a time

I can repro this on samba bookworm, reassigning appropriately.
You only see this because syscalls to your other shares
are infinitely faster (at least 3-20 times) than MHDDFS.

On a Linux smbd mount, "truncate -s 100M 100M" by default (strict allocate = no)
just works and makes a sparse file.

With "strict allocate = yes", I see a busy-loop that looks like this
  mhddfs [2024-11-15 17:17:55] (info): [139671981946560] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104759295
  mhddfs [2024-11-15 17:17:55] (info): [139671963047616] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104763391
  mhddfs [2024-11-15 17:17:55] (info): [139671972497088] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104767487
  mhddfs [2024-11-15 17:17:55] (info): [139671981946560] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104771583
  mhddfs [2024-11-15 17:17:55] (info): [139671963047616] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104775679
  mhddfs [2024-11-15 17:17:55] (info): [139671972497088] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104779775
  mhddfs [2024-11-15 17:17:55] (info): [139671981946560] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104783871
  mhddfs [2024-11-15 17:17:55] (info): [139671963047616] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104787967
  mhddfs [2024-11-15 17:17:55] (info): [139671972497088] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104792063
  mhddfs [2024-11-15 17:17:55] (info): [139671981946560] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104796159
  mhddfs [2024-11-15 17:17:55] (info): [139671963047616] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104800255
  mhddfs [2024-11-15 17:17:55] (info): [139671972497088] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104804351
  mhddfs [2024-11-15 17:17:55] (info): [139671981946560] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104808447
  mhddfs [2024-11-15 17:17:55] (info): [139671963047616] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104812543
  mhddfs [2024-11-15 17:17:55] (info): [139671972497088] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104816639
  mhddfs [2024-11-15 17:17:55] (info): [139671981946560] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104820735
  mhddfs [2024-11-15 17:17:55]: [139671972497088] mhdd_getxattr: path = /home/nabijaczleweli/uwu/mhddfs/100Mc/100M name = security.capability bufsize = 0
  mhddfs [2024-11-15 17:17:55] (info): [139671963047616] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104824831
instead.

And strace -fp "$(pgrep smbd)" shows
  [pid 3565208] pwrite64(28, "\0", 1, 103092223) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103096319) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103100415) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103104511) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103108607) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103112703) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103116799) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103120895) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103124991) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103129087) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103133183) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103137279) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103141375) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103145471) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103149567) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103153663) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103157759) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103161855) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103165951) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103170047) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103174143) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103178239) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103182335) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103186431) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103190527) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103194623) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103198719) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103202815) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103206911) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103211007) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103215103) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103219199) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103223295) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103227391) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103231487) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103235583) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103239679) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103243775) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103247871) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103251967) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103256063) = 1
  [pid 3565208] pwrite64(28, "\0", 1, 103260159) = 1

There /has/ to be some smarter way to do this.

#911118#28
Date:
2024-11-15 19:47:42 UTC
From:
To:
This isn't Samba, perhaps try libc6 or put the bug back to MHDDFS?

I've checked our code, if fallocate() fails, we write 32kb chunks.  But
we won't control how glibc falls back if fallocate is not supported by
the filesystem. 

Andrew Bartlett

#911118#33
Date:
2024-11-15 20:42:56 UTC
From:
To:
Control: clone -1 -2
Control: reassign -2 mhddfs 0.1.39+nmu2
Control: severity -2 wishlist
Control: retitle -2 mhddfs: should supoort fallocate
It also looked like this to me but the falback path wasn't obvious.
(Also, 32k is kinda small IMO but whatever.)
Sure, but since you're already implementing the fallback yourself,
you should probably use fallocate() instead of posix_fallocate().

Under glibc,
  posix_fallocate() does fallocate(2) with the libc fallback if it fails,
  fallocate() just does fallocate(2) and gives you the real error,
so error-detecting after the infallible posix_fallocate() is kinda moot.

Under musl they're both equivalent to fallocate(2).

So while "changing how glibc falls back" is definitely out of scope,
IMO "using a fallible interface to fall back from" could be in scope?

This alone would drop the I/Os issued 8-fold, bumping the buffer
to something more reasonable in 2024 would make it even faster.
I find the belaboured exposition in glibc's posix_fallocate.c convincing,
but MHDDFS should implement fallocate regardless.

#911118#40
Date:
2024-11-16 00:44:38 UTC
From:
To:
These points seem reasonable, but I don't work on Samba day to day any
more (and not on the fileserver either), just still CC'ed on some bugs
and figured I would look into it.

The Samba bugzilla is the place to start a report that those who work
in this area will see (but not act on with any priority, to be clear,
due resourcing constraints) https://bugzilla.samba.org, but to have any
change made to the codebase, I suggest opening a MR
per https://wiki.samba.org/index.php/Contribute

That is the real answer.

Thanks so much for your investigation and I wish you the best 

Andrew Bartlett

#911118#45
Date:
2024-11-16 07:01:54 UTC
From:
To:
Speaking of mhddfs itself, isn't it kinda useless today,
when we had unionfs, aufs, and now have in-kernel overlayfs?

Thanks,

/mjt

#911118#50
Date:
2024-11-16 15:11:53 UTC
From:
To:
Thought so too, but I'm yet to find something that implements something
equivalent to mhddfs (over filesystems; over disks is solved).
All the overlay filesystems are overlays, not stripes, AFAICT.

#911118#57
Date:
2024-11-17 07:49:35 UTC
From:
To:
16.11.2024 18:11, наб wrote:

Hm. Interesting.  Yes, for writes, all other filesystems use
just the top layer, and copy a file from a deeper layer when
it is about to be changed.

I guess overlayfs needs a simple option to turn it into a
stripe-like union where writes goes to the same layer where
the file resides, without the copy-on-write logic.

But apparently this hasn't been done.

Thanks,

/mjt