- Package:
- src:libipc-shareable-perl
- Source:
- src:libipc-shareable-perl
- Submitter:
- Paul Gevers
- Date:
- 2026-09-06 12:35:05 UTC
- Severity:
- normal
- Tags:
Hi, I was looking into why one of the loong64 host of ci.d.n was showing a spike is swap (we have 1.8 TB, it was used all) and loss of debci-workers and the culprit seems to be src:libipc-shareable-perl. I ran """ /usr/bin/autopkgtest --no-built-binaries --timeout=30600 --needs-internet=skip --timeout-install=9999 --user debci --apt-upgrade libipc-shareable-perl -- lxc --sudo --name elbrus autopkgtest-unstable-loong64 """ manually on the host and killed it once top reported 70% memory use by perl (of the 127 GB). Swap plummeted from 74 GB to 10 GB. I guess this means there's a memory leak somewhere on loong64. Otherwise, can you please limit the test to a reasonable memory footprint? Paul https://ci.debian.net/munin/ci-worker-loong64-01/ci-worker-loong64-01/memory.html
Hrm, annoying. On amd64 I behaves somewhat sanely, although it also uses lots of shared memory. I have a suspicion which test causes this problem, but that's more guessing. Did you have a chance to see in top, which process/test was the culprit? Also there's a new upstream release, but that's more of a rewrite and I don't dare to guess if it changes the problem. Maybe we should just try, but I don't want to kill the poor loong64 CI hosts … Cheers, gregor
Hi,
No, but I ran it again to capture it. The output halts here:
t/64-nested_segs_untidy.t ..
ok 1 - Initial array seg count ok
ok 2 - After initial aref add, seg count ok
ok 3 - Adding a new aref to an existing element doesn't create a new seg ok
ok 4 - Same with repurposing the aref again
ok 5 - Same with repurposing the aref again with nested
ok 6 - Nested arrays compare ok
ok 7 - Initial href seg count ok
ok 8 - After initial href add, seg count ok
ok 9 - Adding a new href to an existing key doesn't create a new seg ok
ok 10 - Same with repurposing the href again
ok 11 - Adding a new hash inside of existing does bump seg count
ok 12 - Adding a new hash inside of two level existing does bump seg count
ok 13 - Adding a new hash inside of two level existing twice does bump
seg count
ok 14 - Shared memory hash matches test data ok
1..14
ok
root@elbrus:/# ps faux
USER PID %CPU %MEM VSZ RSS TTY STAT START TIME COMMAND
root 616 0.0 0.0 7776 4304 pts/6 Ss 10:05 0:00 /bin/bash
root 623 0.0 0.0 9552 4000 pts/6 R+ 10:08 0:00 \_
ps faux
root 408 0.0 0.0 7488 3744 pts/5 Ss+ 10:04 0:00 bash
-c set -a; [ -r /etc/environment ] && . /etc/environment 2>/dev/nul
root 411 0.0 0.0 12288 4832 pts/5 S+ 10:04 0:00 \_
su -s /bin/bash root -c set -e; exec /tmp/autopkgtest-lxc.3ruobwtl/d
root 415 0.0 0.0 2688 1744 pts/5 S+ 10:04 0:00
\_ /bin/sh /tmp/autopkgtest-lxc.3ruobwtl/downtmp/wrapper.sh --artif
root 427 0.0 0.0 6160 1888 pts/5 S+ 10:04 0:00
\_ tee -a -- /tmp/autopkgtest-lxc.3ruobwtl/downtmp/autodep8-per
root 429 0.0 0.0 6160 1872 pts/5 S+ 10:04 0:00
\_ tee -a -- /tmp/autopkgtest-lxc.3ruobwtl/downtmp/autodep8-per
root 431 0.0 0.0 2688 1728 pts/5 S+ 10:04 0:00
\_ /bin/sh /usr/share/pkg-perl-autopkgtest/runner build-deps
root 438 0.0 0.0 2688 1728 pts/5 S+ 10:04 0:00
\_ /bin/sh /usr/share/pkg-perl-autopkgtest/build-deps.d/smo
root 463 0.1 0.0 22240 11440 pts/5 S+ 10:04 0:00
\_ /usr/bin/perl /usr/bin/prove -I/tmp/autopkgtest-lxc.
root 601 32.8 53.2 1953147760 70989856 pts/5 D+ 10:05 1:12
\_ /usr/bin/perl t/65-seg_size.t
root 1 0.1 0.0 18896 12992 ? Ss 10:04 0:00
/sbin/init
root 43 0.0 0.0 28624 11536 ? Ss 10:04 0:00 \_
/usr/lib/systemd/systemd-journald
message+ 75 0.0 0.0 8784 4160 ? Ss 10:04 0:00 \_
/usr/bin/dbus-daemon --system --address=systemd: --nofork --nopidfil
root 76 0.0 0.0 11456 7184 ? Ss 10:04 0:00 \_
/usr/lib/systemd/systemd-logind
_dhcpcd 90 0.0 0.0 10000 4304 ? Ss 10:04 0:00 \_
dhcpcd: eth0 [ip4] [ip6]
root 91 0.0 0.0 9824 2352 ? S 10:04 0:00 |
\_ dhcpcd: [privileged proxy] eth0 [ip4] [ip6]
_dhcpcd 100 0.0 0.0 9824 2208 ? S 10:04 0:00 |
| \_ dhcpcd: [DHCP6 proxy] fe80::1266:6aff:fe2f:81c4
_dhcpcd 117 0.0 0.0 9824 2208 ? S 10:04 0:00 |
| \_ dhcpcd: [DHCP6 proxy] fc42:5009:ba4b:5ab0:f29:d8d9:95f0:e39f
_dhcpcd 168 0.0 0.0 9824 2144 ? S 10:04 0:00 |
| \_ dhcpcd: [BPF ARP] eth0 10.0.193.234
_dhcpcd 500 0.0 0.0 9824 2080 ? S 10:04 0:00 |
| \_ dhcpcd: [BOOTP proxy] 10.0.193.234
_dhcpcd 92 0.0 0.0 9808 2032 ? S 10:04 0:00 |
\_ dhcpcd: [network proxy] eth0 [ip4] [ip6]
_dhcpcd 93 0.0 0.0 9808 1968 ? S 10:04 0:00 |
\_ dhcpcd: [control proxy] eth0 [ip4] [ip6]
root 132 0.0 0.0 7824 2624 pts/0 Ss+ 10:04 0:00 \_
/usr/sbin/agetty --noreset --noclear --keep-baud 115200,57600,38400,
root 133 0.0 0.0 11408 7792 ? Ss 10:04 0:00 \_
sshd: /usr/sbin/sshd -D [listener] 0 of 10-100 startups
Well, the package is currently on the reject_list on loong64, so don't
worry about that. I can schedule it when you let me know.
Paul
Control: retitle -1 libipc-shareable-perl: t/65-seg_size.t uses insane amount of memory on loong64 Thanks, much appreciated! … … So it's t/65-seg_size.t again. "Again" because this test already has a patch to be skipped on 32bit platforms, which was adopted upstream in the new release (at a slightly different position in the renamed tests but with the same effect). That means that I don't expect any improvement by the new upstream release at first sight. Now we either need to find out what's different on loong64 (staring at the output of `perl -V' didn't show anything obvious to me) or we can skip the test (in general or on loong64). Oh, good to know, and thanks for the offer. Cheers, gregor
Incidentally, the package changed from arch:all to arch:any. And I added salsa-ci tests, and the debrebuild and test-build-any and autopkgtest jobs already failed with: Test Summary Report ------------------- t/47-seg_size.t (Wstat: 9 (Signal: KILL) Tests: 2 Failed: 0) Non-zero wait status: 9 Parse errors: No plan found in TAP output t/99-end.t (Wstat: 512 (exited 2) Tests: 3 Failed: 2) Failed tests: 2-3 Non-zero exit status: 2 Files=64, Tests=1352, 23 wallclock secs ( 0.49 usr 0.10 sys + 10.19 cusr 8.65 csys = 19.43 CPU) Result: FAIL So our friend (t/65-seg_size.t renamed to) t/47-seg_size.t again. And t/99-end.t. The build jobs on amd64/i386 pass. With the same tests. I'm happy to disable t/47-seg_size.t (and have done so in Git) but then we now also have a failing t/99-end.t. Sometimes. Next step: After disabling t/47-seg_size.t we have a green pipeline. I.e. t/99-end.t. also passes. insert-shrug-emoji-here Alright, uploading, let's see if the buildds and ci hosts survive … *** Next day: The build fails on ppc64el https://buildd.debian.org/status/package.php?p=libipc-shareable-perl https://buildd.debian.org/status/fetch.php?pkg=libipc-shareable-perl&arch=ppc64el&ver=1.19-1&stamp=1787773225&raw=0 in _different_ tests: Test Summary Report ------------------- t/72-shm_segments.t (Wstat: 6912 (exited 27) Tests: 47 Failed: 27) Failed tests: 2-6, 10-14, 17-21, 23-27, 29, 31, 33, 35 37, 41, 44 Non-zero exit status: 27 t/74-seg_map.t (Wstat: 512 (exited 2) Tests: 22 Failed: 2) Failed tests: 4, 6 Non-zero exit status: 2 Files=65, Tests=1349, 14 wallclock secs ( 0.29 usr 0.07 sys + 9.31 cusr 1.82 csys = 11.49 CPU) Result: FAIL debci doesn't have 1.19-1 results: https://ci.debian.net/packages/libi/libipc-shareable-perl/ Cheers, gregor
Hi gregor, I've triggered them (as myself) in unstable except on loong64 as that's reject_listed. Our loong64 workers are down at the moment, so out-of-band triggering won't give a result soon. Testing on ppc64el fails as the package is now unknown there (due to the arch:all to arch:any switch with the FTBFS on ppc64el). Should the migration be explicitly blocked, or do you think it's preferable to have the installability step of britney2 try? Paul
ok, thanks! As it doesn't build on ppc64el and breaks the autopkgtests of inetsim, it can't migrate anyway? Ah, maybe that was the last part of your question -- that/if britney should just block migration as per its criteria as usual. Yes, I think this will do … Cheers, gregor
Hi This doesn't block as the package was never built on ppc64el, so the FTBFS isn't a regression (weird consequence of the arch:all to arch:any move). So I was mostly referring to this. Which is "just" new output to stderr, so easy to fix (I don't think the test is very clever (it's also marked superficial, so I don't suspect this output isn't a blocking regression for inetsim). So, once inetsim autopkgtest isn't blocking, should it migrate despite not available anymore on ppc64el in your opinion? Paul
Oops, right. Ah, ok. As I haven't closed #1142815 in the upload (and plan to do so right now as long as we see other issues) … ah -- that's not a blocker either as it already effects older versions. So, yes, I finally see it: please block it for now. (And I hope smeone else will come along and "just fix" this package :)) Cheers, gregor
Hi all,
Okay, I had a stab at this. I looked at IPC-Shareable-1.19/t/47-seg_size.t.
My background is rather in C, and I have never written any Perl, so for
the patch, I will happily leave that to the Perl devs here. Therefore, I
am afraid, no "just fix" solution; but still the cause and a write-up
for a patch.
## The Root Cause
The test has a "beyond RAM limits" block that passes `limit => 0` to
bypass internal module checks, and intentionally requests an
"impossible" 999,999,999,999-byte (~1 TB) segment, expecting the OS to
refuse it.
- On non-Linux systems (macOS, BSD), this works, because they enforce
small, strict default SHMMAX ceilings (e.g., 4 MB to 512 MB).
- On Linux, the default `kernel.shmmax` is effectively unlimited
(`ULONG_MAX - 2**24` - check `/proc/sys/kernel/shmmax`). Therefore,
`shmget()` falls through to the standard overcommit heuristic in
`__vm_enough_memory()`, which only rejects segments that exceed MemTotal
+ SwapTotal.
- On large infrastructure nodes where RAM + swap exceeds 1 TB, the
kernel grants the segment lazily. When `SharedMem.pm` subsequently
performs a `shmread()` to inspect the segment, it forces two massive
physical allocations simultaneously: a 1 TB Perl scalar heap allocation
for the read buffer, and the physical materialisation of the 1 TB
zero-filled shared memory segment via page faults. This results in an
immediate 2 TB memory footprint spike, causing the test runner to be killed.
Paul's numbers confirm this.
Paul reported a virtual size of 1,953,147,760 KiB for `/usr/bin/perl
t/65-seg_size.t` on line 601. When converted, that is exactly
2,000,023,306,240 bytes. Against the 1 TB target request, this maps
almost cleanly to a 2.0000x multiplier. The remaining ~22 MB is also to
be found in Paul's message: `/usr/bin/perl /usr/bin/prove` on line 463
shows a virtual size of 22240 KiB = 22.19 MiB.
## Proposed Fix Architecture
Rather than relying on a hardcoded "impossible" size that modern server
hardware can easily accommodate, the test script (t/47-seg_size.t)
should calculate a dynamic refusal ceiling at runtime:
1. Keep a baseline fallback: Define a static constant of 999999999999
bytes (~1 TB) as the default choice for non-Linux architectures.
2. Handle a lowered Linux shmmax (if defined): On Linux, check
`/proc/sys/kernel/shmmax`. If a system administrator has manually set it
to a value lower than 1 TB, keep the 1 TB test size, as the OS will
already safely reject it.
3. Account for Overcommit Mode 1: Check
`/proc/sys/vm/overcommit_memory`. If it is set to 1 (always overcommit),
the kernel will never reject an allocation at creation time. In this
case, cleanly skip the two test assertions to prevent a guaranteed
out-of-memory crash.
4. Compute a dynamic ceiling via `meminfo`: If overcommit is in standard
mode, parse `/proc/meminfo` to read MemTotal and SwapTotal. Calculate a
new target size equivalent to MemTotal + SwapTotal + 1 GB (adding 1 GB
to safely clear page-alignment rounding bounds). Use this computed size
for the tie operation to guarantee an immediate application-layer or
kernel-level refusal without initialising any real memory blocks.
5. Enforce defensive container fallbacks: Ensure that all file-read
operations on `/proc` files fail gracefully (returning the 1 TB baseline
fallback) if executed inside locked-down sandbox environments, Distrobox
environments, or strict container runtimes where `/proc` access is
restricted.
This runtime logic completely avoids hardcoding architecture-specific
blacklists, cleanly accommodates modern high-memory hardware
configurations, and prevents sudden worker terminations.
Edmund
Thanks for looking into this! It will be helpful if someone tries to fix 47-seg_size.t. 47-seg_size.t is not the current problem, as I have disabled it in the package. The new problem is the one I mentioned in https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1142815#27 : Failures only on ppc64el in two different tests (t/72-shm_segments.t and t/74-seg_map.t). Cheers, gregor
Because this is an RC bug and blocks others behind it, I have create an MR on Salsa. The bug for t-47 is now understood, and I therefore took the liberty to re-enable the test. If you would rather not, I'll be happy to change that again. Upstream links: https://github.com/stevieb9/ipc-shareable/issues/65 https://github.com/stevieb9/ipc-shareable/issues/66 Salsa MR: https://salsa.debian.org/perl-team/modules/packages/libipc-shareable-perl/-/merge_requests/1 I hope this helps! Kind regards, Edmund