When I initiate a samba share file transfer to a drive mounted through MHDDFS, it takes many seconds to start writing, and I observe high MHDDFS CPU usage. Reading a file works perfectly fine. Local writes are also fine. The length of the delay seems to scale with the size of the file I want to write. In the Log, when the client waits for the transfer to start, this message shows up many times: mhddfs [2018-10-15 01:03:37]: [140267440191232] mhdd_getxattr: path = /mnt/hd1/Folder/File-to-Copy.mkv name = security.capability bufsize = 0 Copying the same file to the same location through a direct Samba share is fine. This seems to happen with Windows clients (7 & 10 tested), and not with the other Ubuntu machine I have. best regards Marcél ProblemType: Bug ApportVersion: 2.20.9-0ubuntu7.4 Architecture: amd64 CurrentDesktop: ubuntu:GNOME Date: Mon Oct 15 01:24:54 2018 Dependencies: adduser 3.116ubuntu1 apt 1.6.3ubuntu0.1 apt-utils 1.6.3ubuntu0.1 ca-certificates 20180409 debconf 1.5.66 debconf-i18n 1.5.66 dpkg 1.19.0.5ubuntu2 fdisk 2.31.1-0.4ubuntu3.1 fuse 2.9.7-1ubuntu1 gcc-8-base 8.2.0-1ubuntu2~18.04 gpgv 2.2.4-1ubuntu1.1 libacl1 2.2.52-3build1 libapt-inst2.0 1.6.3ubuntu0.1 libapt-pkg5.0 1.6.3ubuntu0.1 libattr1 1:2.4.47-2build1 libaudit-common 1:2.8.2-1ubuntu1 libaudit1 1:2.8.2-1ubuntu1 libblkid1 2.31.1-0.4ubuntu3.1 libbz2-1.0 1.0.6-8.1 libc6 2.27-3ubuntu1 libcap-ng0 0.7.7-3.1 libdb5.3 5.3.28-13.1ubuntu1 libfdisk1 2.31.1-0.4ubuntu3.1 libffi6 3.2.1-8 libfuse2 2.9.7-1ubuntu1 libgcc1 1:8.2.0-1ubuntu2~18.04 libgcrypt20 1.8.1-4ubuntu1.1 libgmp10 2:6.1.2+dfsg-2 libgnutls30 3.5.18-1ubuntu1 libgpg-error0 1.27-6 libgpm2 1.20.7-5 libhogweed4 3.4-1 libidn2-0 2.0.4-1.1build2 liblocale-gettext-perl 1.07-3build2 liblz4-1 0.0~r131-2ubuntu3 liblzma5 5.2.2-1.3 libmount1 2.31.1-0.4ubuntu3.1 libncursesw5 6.1-1ubuntu1.18.04 libnettle6 3.4-1 libp11-kit0 0.23.9-2 libpam-modules 1.1.8-3.6ubuntu2 libpam-modules-bin 1.1.8-3.6ubuntu2 libpam0g 1.1.8-3.6ubuntu2 libpcre3 2:8.39-9 libseccomp2 2.3.1-2.1ubuntu4 libselinux1 2.7-2build2 libsemanage-common 2.7-2build2 libsemanage1 2.7-2build2 libsepol1 2.7-1 libsmartcols1 2.31.1-0.4ubuntu3.1 libssl1.1 1.1.0g-2ubuntu4.1 libstdc++6 8.2.0-1ubuntu2~18.04 libsystemd0 237-3ubuntu10.3 libtasn1-6 4.13-2 libtext-charwidth-perl 0.04-7.1 libtext-iconv-perl 1.7-5build6 libtext-wrapi18n-perl 0.06-7.1 libtinfo5 6.1-1ubuntu1.18.04 libudev1 237-3ubuntu10.3 libunistring2 0.9.9-0ubuntu1 libuuid1 2.31.1-0.4ubuntu3.1 libzstd1 1.3.3+dfsg-2ubuntu1 mount 2.31.1-0.4ubuntu3.1 openssl 1.1.0g-2ubuntu4.1 passwd 1:4.5-1ubuntu1 perl-base 5.26.1-6ubuntu0.2 sed 4.4-2 tar 1.29b-2 ubuntu-keyring 2018.02.28 util-linux 2.31.1-0.4ubuntu3.1 uuid-runtime 2.31.1-0.4ubuntu3.1 zlib1g 1:1.2.11.dfsg-0ubuntu2 DistroRelease: Ubuntu 18.04 InstallationDate: Installed on 2018-09-28 (16 days ago) InstallationMedia: Ubuntu 18.04 LTS "Bionic Beaver" - Release amd64 (20180426) Package: mhddfs 0.1.39+nmu1ubuntu2 PackageArchitecture: amd64 ProcCpuinfoMinimal: processor : 7 vendor_id : GenuineIntel cpu family : 6 model : 58 model name : Intel(R) Core(TM) i7-3770 CPU @ 3.40GHz stepping : 9 microcode : 0x20 cpu MHz : 3531.958 cache size : 8192 KB physical id : 0 siblings : 8 core id : 3 cpu cores : 4 apicid : 7 initial apicid : 7 fpu : yes fpu_exception : yes cpuid level : 13 wp : yes flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc cpuid aperfmperf pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm cpuid_fault epb pti ssbd ibrs ibpb stibp tpr_shadow vnmi flexpriority ept vpid fsgsbase smep erms xsaveopt dtherm ida arat pln pts flush_l1d bugs : cpu_meltdown spectre_v1 spectre_v2 spec_store_bypass l1tf bogomips : 6800.35 clflush size : 64 cache_alignment : 64 address sizes : 36 bits physical, 48 bits virtual power management: ProcEnviron: LC_TIME=de_AT.UTF-8 LC_MONETARY=de_AT.UTF-8 TERM=xterm-256color PATH=(custom, no user) LC_ADDRESS=de_AT.UTF-8 XDG_RUNTIME_DIR=<set> LANG=en_US.UTF-8 LC_TELEPHONE=de_AT.UTF-8 LC_NAME=de_AT.UTF-8 SHELL=/bin/bash LC_MEASUREMENT=de_AT.UTF-8 LC_IDENTIFICATION=de_AT.UTF-8 LC_NUMERIC=de_AT.UTF-8 LC_PAPER=de_AT.UTF-8 ProcVersionSignature: Ubuntu 4.15.0-36.39-generic 4.15.18 SourcePackage: mhddfs Tags: bionic Uname: Linux 4.15.0-36-generic x86_64 UpgradeStatus: No upgrade log present (probably fresh install) _MarkForUpload: True
I must have overlooked one setting in Samba when I was trying to eliminate its configuration as a source for this issue: strict allocate = yes This seems to be what caused the problem. Sorry for the falsely attributed bug report, but maybe it's interesting information - I still think it's weird that this only caused an issue with FUSE/MHDDFS, not the other share I was using.
Control: retitle -1 smbd: please "strict allocate" smarter than 1 byte at a time I can repro this on samba bookworm, reassigning appropriately. You only see this because syscalls to your other shares are infinitely faster (at least 3-20 times) than MHDDFS. On a Linux smbd mount, "truncate -s 100M 100M" by default (strict allocate = no) just works and makes a sparse file. With "strict allocate = yes", I see a busy-loop that looks like this mhddfs [2024-11-15 17:17:55] (info): [139671981946560] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104759295 mhddfs [2024-11-15 17:17:55] (info): [139671963047616] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104763391 mhddfs [2024-11-15 17:17:55] (info): [139671972497088] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104767487 mhddfs [2024-11-15 17:17:55] (info): [139671981946560] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104771583 mhddfs [2024-11-15 17:17:55] (info): [139671963047616] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104775679 mhddfs [2024-11-15 17:17:55] (info): [139671972497088] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104779775 mhddfs [2024-11-15 17:17:55] (info): [139671981946560] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104783871 mhddfs [2024-11-15 17:17:55] (info): [139671963047616] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104787967 mhddfs [2024-11-15 17:17:55] (info): [139671972497088] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104792063 mhddfs [2024-11-15 17:17:55] (info): [139671981946560] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104796159 mhddfs [2024-11-15 17:17:55] (info): [139671963047616] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104800255 mhddfs [2024-11-15 17:17:55] (info): [139671972497088] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104804351 mhddfs [2024-11-15 17:17:55] (info): [139671981946560] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104808447 mhddfs [2024-11-15 17:17:55] (info): [139671963047616] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104812543 mhddfs [2024-11-15 17:17:55] (info): [139671972497088] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104816639 mhddfs [2024-11-15 17:17:55] (info): [139671981946560] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104820735 mhddfs [2024-11-15 17:17:55]: [139671972497088] mhdd_getxattr: path = /home/nabijaczleweli/uwu/mhddfs/100Mc/100M name = security.capability bufsize = 0 mhddfs [2024-11-15 17:17:55] (info): [139671963047616] mhdd_write: /100M, handle = 139671732492784, count = 1, offset = 104824831 instead. And strace -fp "$(pgrep smbd)" shows [pid 3565208] pwrite64(28, "\0", 1, 103092223) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103096319) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103100415) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103104511) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103108607) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103112703) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103116799) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103120895) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103124991) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103129087) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103133183) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103137279) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103141375) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103145471) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103149567) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103153663) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103157759) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103161855) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103165951) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103170047) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103174143) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103178239) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103182335) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103186431) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103190527) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103194623) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103198719) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103202815) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103206911) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103211007) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103215103) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103219199) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103223295) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103227391) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103231487) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103235583) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103239679) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103243775) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103247871) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103251967) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103256063) = 1 [pid 3565208] pwrite64(28, "\0", 1, 103260159) = 1 There /has/ to be some smarter way to do this.
This isn't Samba, perhaps try libc6 or put the bug back to MHDDFS? I've checked our code, if fallocate() fails, we write 32kb chunks. But we won't control how glibc falls back if fallocate is not supported by the filesystem. Andrew Bartlett
Control: clone -1 -2 Control: reassign -2 mhddfs 0.1.39+nmu2 Control: severity -2 wishlist Control: retitle -2 mhddfs: should supoort fallocate It also looked like this to me but the falback path wasn't obvious. (Also, 32k is kinda small IMO but whatever.) Sure, but since you're already implementing the fallback yourself, you should probably use fallocate() instead of posix_fallocate(). Under glibc, posix_fallocate() does fallocate(2) with the libc fallback if it fails, fallocate() just does fallocate(2) and gives you the real error, so error-detecting after the infallible posix_fallocate() is kinda moot. Under musl they're both equivalent to fallocate(2). So while "changing how glibc falls back" is definitely out of scope, IMO "using a fallible interface to fall back from" could be in scope? This alone would drop the I/Os issued 8-fold, bumping the buffer to something more reasonable in 2024 would make it even faster. I find the belaboured exposition in glibc's posix_fallocate.c convincing, but MHDDFS should implement fallocate regardless.
These points seem reasonable, but I don't work on Samba day to day any more (and not on the fileserver either), just still CC'ed on some bugs and figured I would look into it. The Samba bugzilla is the place to start a report that those who work in this area will see (but not act on with any priority, to be clear, due resourcing constraints) https://bugzilla.samba.org, but to have any change made to the codebase, I suggest opening a MR per https://wiki.samba.org/index.php/Contribute That is the real answer. Thanks so much for your investigation and I wish you the best Andrew Bartlett
Speaking of mhddfs itself, isn't it kinda useless today, when we had unionfs, aufs, and now have in-kernel overlayfs? Thanks, /mjt
Thought so too, but I'm yet to find something that implements something equivalent to mhddfs (over filesystems; over disks is solved). All the overlay filesystems are overlays, not stripes, AFAICT.
16.11.2024 18:11, наб wrote: Hm. Interesting. Yes, for writes, all other filesystems use just the top layer, and copy a file from a deeper layer when it is about to be changed. I guess overlayfs needs a simple option to turn it into a stripe-like union where writes goes to the same layer where the file resides, without the copy-on-write logic. But apparently this hasn't been done. Thanks, /mjt