#1142815 libipc-shareable-perl: autopkgtest uses insane amount of memory on loong64

#1142815#5
Date:
2026-07-26 15:30:48 UTC
From:
To:
Hi,

I was looking into why one of the loong64 host of ci.d.n was showing a
spike is swap (we have 1.8 TB, it was used all) and loss of
debci-workers and the culprit seems to be src:libipc-shareable-perl. I ran
"""
/usr/bin/autopkgtest --no-built-binaries  --timeout=30600
--needs-internet=skip --timeout-install=9999 --user debci --apt-upgrade 
libipc-shareable-perl -- lxc --sudo --name elbrus
autopkgtest-unstable-loong64
"""
manually on the host and killed it once top reported 70% memory use by
perl (of the 127 GB). Swap plummeted from 74 GB to 10 GB.

I guess this means there's a memory leak somewhere on loong64.
Otherwise, can you please limit the test to a reasonable memory footprint?

Paul

https://ci.debian.net/munin/ci-worker-loong64-01/ci-worker-loong64-01/memory.html

#1142815#10
Date:
2026-08-06 22:07:07 UTC
From:
To:
Hrm, annoying.

On amd64 I behaves somewhat sanely, although it also uses lots of
shared memory.

I have a suspicion which test causes this problem, but that's more
guessing.

Did you have a chance to see in top, which process/test was the
culprit?

Also there's a new upstream release, but that's more of a rewrite
and I don't dare to guess if it changes the problem. Maybe we should
just try, but I don't want to kill the poor loong64 CI hosts …


Cheers,
gregor

#1142815#15
Date:
2026-08-07 10:16:14 UTC
From:
To:
Hi,
No, but I ran it again to capture it. The output halts here:
t/64-nested_segs_untidy.t ..
ok 1 - Initial array seg count ok
ok 2 - After initial aref add, seg count ok
ok 3 - Adding a new aref to an existing element doesn't create a new seg ok
ok 4 - Same with repurposing the aref again
ok 5 - Same with repurposing the aref again with nested
ok 6 - Nested arrays compare ok
ok 7 - Initial href seg count ok
ok 8 - After initial href add, seg count ok
ok 9 - Adding a new href to an existing key doesn't create a new seg ok
ok 10 - Same with repurposing the href again
ok 11 - Adding a new hash inside of existing does bump seg count
ok 12 - Adding a new hash inside of two level existing does bump seg count
ok 13 - Adding a new hash inside of two level existing twice does bump
seg count
ok 14 - Shared memory hash matches test data ok
1..14
ok

root@elbrus:/# ps faux
USER         PID %CPU %MEM    VSZ   RSS TTY      STAT START   TIME COMMAND
root         616  0.0  0.0   7776  4304 pts/6    Ss   10:05   0:00 /bin/bash
root         623  0.0  0.0   9552  4000 pts/6    R+   10:08   0:00  \_
ps faux
root         408  0.0  0.0   7488  3744 pts/5    Ss+  10:04   0:00 bash
-c set -a; [ -r /etc/environment ] && . /etc/environment 2>/dev/nul
root         411  0.0  0.0  12288  4832 pts/5    S+   10:04   0:00  \_
su -s /bin/bash root -c set -e; exec /tmp/autopkgtest-lxc.3ruobwtl/d
root         415  0.0  0.0   2688  1744 pts/5    S+   10:04   0:00
\_ /bin/sh /tmp/autopkgtest-lxc.3ruobwtl/downtmp/wrapper.sh --artif
root         427  0.0  0.0   6160  1888 pts/5    S+   10:04   0:00
    \_ tee -a -- /tmp/autopkgtest-lxc.3ruobwtl/downtmp/autodep8-per
root         429  0.0  0.0   6160  1872 pts/5    S+   10:04   0:00
    \_ tee -a -- /tmp/autopkgtest-lxc.3ruobwtl/downtmp/autodep8-per
root         431  0.0  0.0   2688  1728 pts/5    S+   10:04   0:00
    \_ /bin/sh /usr/share/pkg-perl-autopkgtest/runner build-deps
root         438  0.0  0.0   2688  1728 pts/5    S+   10:04   0:00
        \_ /bin/sh /usr/share/pkg-perl-autopkgtest/build-deps.d/smo
root         463  0.1  0.0  22240 11440 pts/5    S+   10:04   0:00
            \_ /usr/bin/perl /usr/bin/prove -I/tmp/autopkgtest-lxc.
root         601 32.8 53.2 1953147760 70989856 pts/5 D+ 10:05   1:12
                  \_ /usr/bin/perl t/65-seg_size.t
root           1  0.1  0.0  18896 12992 ?        Ss   10:04   0:00
/sbin/init
root          43  0.0  0.0  28624 11536 ?        Ss   10:04   0:00  \_
/usr/lib/systemd/systemd-journald
message+      75  0.0  0.0   8784  4160 ?        Ss   10:04   0:00  \_
/usr/bin/dbus-daemon --system --address=systemd: --nofork --nopidfil
root          76  0.0  0.0  11456  7184 ?        Ss   10:04   0:00  \_
/usr/lib/systemd/systemd-logind
_dhcpcd       90  0.0  0.0  10000  4304 ?        Ss   10:04   0:00  \_
dhcpcd: eth0 [ip4] [ip6]
root          91  0.0  0.0   9824  2352 ?        S    10:04   0:00  |
\_ dhcpcd: [privileged proxy] eth0 [ip4] [ip6]
_dhcpcd      100  0.0  0.0   9824  2208 ?        S    10:04   0:00  |
|   \_ dhcpcd: [DHCP6 proxy] fe80::1266:6aff:fe2f:81c4
_dhcpcd      117  0.0  0.0   9824  2208 ?        S    10:04   0:00  |
|   \_ dhcpcd: [DHCP6 proxy] fc42:5009:ba4b:5ab0:f29:d8d9:95f0:e39f
_dhcpcd      168  0.0  0.0   9824  2144 ?        S    10:04   0:00  |
|   \_ dhcpcd: [BPF ARP] eth0 10.0.193.234
_dhcpcd      500  0.0  0.0   9824  2080 ?        S    10:04   0:00  |
|   \_ dhcpcd: [BOOTP proxy] 10.0.193.234
_dhcpcd       92  0.0  0.0   9808  2032 ?        S    10:04   0:00  |
\_ dhcpcd: [network proxy] eth0 [ip4] [ip6]
_dhcpcd       93  0.0  0.0   9808  1968 ?        S    10:04   0:00  |
\_ dhcpcd: [control proxy] eth0 [ip4] [ip6]
root         132  0.0  0.0   7824  2624 pts/0    Ss+  10:04   0:00  \_
/usr/sbin/agetty --noreset --noclear --keep-baud 115200,57600,38400,
root         133  0.0  0.0  11408  7792 ?        Ss   10:04   0:00  \_
sshd: /usr/sbin/sshd -D [listener] 0 of 10-100 startups


Well, the package is currently on the reject_list on loong64, so don't
worry about that. I can schedule it when you let me know.

Paul

#1142815#20
Date:
2026-08-08 21:12:47 UTC
From:
To:
Control: retitle -1 libipc-shareable-perl: t/65-seg_size.t uses insane amount of memory on loong64

Thanks, much appreciated!
…
…

So it's t/65-seg_size.t again.

"Again" because this test already has a patch to be skipped on 32bit
platforms, which was adopted upstream in the new release (at a
slightly different position in the renamed tests but with the same
effect).

That means that I don't expect any improvement by the new upstream
release at first sight.

Now we either need to find out what's different on loong64 (staring
at the output of `perl -V' didn't show anything obvious to me) or we
can skip the test (in general or on loong64).

Oh, good to know, and thanks for the offer.


Cheers,
gregor

#1142815#27
Date:
2026-08-27 09:12:20 UTC
From:
To:
Incidentally, the package changed from arch:all to arch:any.

And I added salsa-ci tests, and the debrebuild and test-build-any and autopkgtest
jobs already failed with:

Test Summary Report
-------------------
t/47-seg_size.t               (Wstat: 9 (Signal: KILL) Tests: 2 Failed: 0)
   Non-zero wait status: 9
   Parse errors: No plan found in TAP output
t/99-end.t                    (Wstat: 512 (exited 2) Tests: 3 Failed: 2)
   Failed tests:  2-3
   Non-zero exit status: 2
Files=64, Tests=1352, 23 wallclock secs ( 0.49 usr  0.10 sys + 10.19 cusr  8.65 csys = 19.43 CPU)
Result: FAIL

So our friend (t/65-seg_size.t renamed to) t/47-seg_size.t again. And
t/99-end.t.

The build jobs on amd64/i386 pass. With the same tests.


I'm happy to disable t/47-seg_size.t (and have done so in Git) but
then we now also have a failing t/99-end.t. Sometimes.


Next step: After disabling t/47-seg_size.t we have a green pipeline.
I.e. t/99-end.t. also passes. insert-shrug-emoji-here


Alright, uploading, let's see if the buildds and ci hosts survive …

***

Next day:

The build fails on ppc64el
https://buildd.debian.org/status/package.php?p=libipc-shareable-perl
https://buildd.debian.org/status/fetch.php?pkg=libipc-shareable-perl&arch=ppc64el&ver=1.19-1&stamp=1787773225&raw=0
in _different_ tests:

Test Summary Report
-------------------
t/72-shm_segments.t           (Wstat: 6912 (exited 27) Tests: 47 Failed: 27)
   Failed tests:  2-6, 10-14, 17-21, 23-27, 29, 31, 33, 35
                 37, 41, 44
   Non-zero exit status: 27
t/74-seg_map.t                (Wstat: 512 (exited 2) Tests: 22 Failed: 2)
   Failed tests:  4, 6
   Non-zero exit status: 2
Files=65, Tests=1349, 14 wallclock secs ( 0.29 usr  0.07 sys +  9.31 cusr  1.82 csys = 11.49 CPU)
Result: FAIL

debci doesn't have 1.19-1 results:
https://ci.debian.net/packages/libi/libipc-shareable-perl/



Cheers,
gregor

#1142815#32
Date:
2026-08-27 10:30:17 UTC
From:
To:
Hi gregor,


I've triggered them (as myself) in unstable except on loong64 as that's
reject_listed. Our loong64 workers are down at the moment, so
out-of-band triggering won't give a result soon.

Testing on ppc64el fails as the package is now unknown there (due to the
arch:all to arch:any switch with the FTBFS on ppc64el). Should the
migration be explicitly blocked, or do you think it's preferable to have
the installability step of britney2 try?

Paul

#1142815#37
Date:
2026-08-27 11:36:02 UTC
From:
To:
ok, thanks!

As it doesn't build on ppc64el and breaks the autopkgtests of
inetsim, it can't migrate anyway? Ah, maybe that was the last part of
your question -- that/if britney should just block migration as per
its criteria as usual. Yes, I think this will do …


Cheers,
gregor

#1142815#42
Date:
2026-08-27 11:55:41 UTC
From:
To:
Hi


This doesn't block as the package was never built on ppc64el, so the
FTBFS isn't a regression (weird consequence of the arch:all to arch:any
move). So I was mostly referring to this.


Which is "just" new output to stderr, so easy to fix (I don't think the
test is very clever (it's also marked superficial, so I don't suspect
this output isn't a blocking regression for inetsim).


So, once inetsim autopkgtest isn't blocking, should it migrate despite
not available anymore on ppc64el in your opinion?

Paul

#1142815#47
Date:
2026-08-27 12:52:41 UTC
From:
To:
Oops, right.

Ah, ok.

As I haven't closed #1142815 in the upload (and plan to do so right
now as long as we see other issues) … ah -- that's not a blocker
either as it already effects older versions.

So, yes, I finally see it: please block it for now.


(And I hope smeone else will come along and "just fix" this package
:))


Cheers,
gregor

#1142815#52
Date:
2026-08-27 16:23:38 UTC
From:
To:
Hi all,

Okay, I had a stab at this. I looked at IPC-Shareable-1.19/t/47-seg_size.t.

My background is rather in C, and I have never written any Perl, so for
the patch, I will happily leave that to the Perl devs here. Therefore, I
am afraid, no "just fix" solution; but still the cause and a write-up
for a patch.

## The Root Cause

The test has a "beyond RAM limits" block that passes `limit => 0` to
bypass internal module checks, and intentionally requests an
"impossible" 999,999,999,999-byte (~1 TB) segment, expecting the OS to
refuse it.

- On non-Linux systems (macOS, BSD), this works, because they enforce
small, strict default SHMMAX ceilings (e.g., 4 MB to 512 MB).
- On Linux, the default `kernel.shmmax` is effectively unlimited
(`ULONG_MAX - 2**24` - check `/proc/sys/kernel/shmmax`). Therefore,
`shmget()` falls through to the standard overcommit heuristic in
`__vm_enough_memory()`, which only rejects segments that exceed MemTotal
+ SwapTotal.
- On large infrastructure nodes where RAM + swap exceeds 1 TB, the
kernel grants the segment lazily. When `SharedMem.pm` subsequently
performs a `shmread()` to inspect the segment, it forces two massive
physical allocations simultaneously: a 1 TB Perl scalar heap allocation
for the read buffer, and the physical materialisation of the 1 TB
zero-filled shared memory segment via page faults. This results in an
immediate 2 TB memory footprint spike, causing the test runner to be killed.

Paul's numbers confirm this.

Paul reported a virtual size of 1,953,147,760 KiB for `/usr/bin/perl
t/65-seg_size.t` on line 601. When converted, that is exactly
2,000,023,306,240 bytes. Against the 1 TB target request, this maps
almost cleanly to a 2.0000x multiplier. The remaining ~22 MB is also to
be found in Paul's message: `/usr/bin/perl /usr/bin/prove` on line 463
shows a virtual size of 22240 KiB = 22.19 MiB.

## Proposed Fix Architecture

Rather than relying on a hardcoded "impossible" size that modern server
hardware can easily accommodate, the test script (t/47-seg_size.t)
should calculate a dynamic refusal ceiling at runtime:

1. Keep a baseline fallback: Define a static constant of 999999999999
bytes (~1 TB) as the default choice for non-Linux architectures.
2. Handle a lowered Linux shmmax (if defined): On Linux, check
`/proc/sys/kernel/shmmax`. If a system administrator has manually set it
to a value lower than 1 TB, keep the 1 TB test size, as the OS will
already safely reject it.
3. Account for Overcommit Mode 1: Check
`/proc/sys/vm/overcommit_memory`. If it is set to 1 (always overcommit),
the kernel will never reject an allocation at creation time. In this
case, cleanly skip the two test assertions to prevent a guaranteed
out-of-memory crash.
4. Compute a dynamic ceiling via `meminfo`: If overcommit is in standard
mode, parse `/proc/meminfo` to read MemTotal and SwapTotal. Calculate a
new target size equivalent to MemTotal + SwapTotal + 1 GB (adding 1 GB
to safely clear page-alignment rounding bounds). Use this computed size
for the tie operation to guarantee an immediate application-layer or
kernel-level refusal without initialising any real memory blocks.
5. Enforce defensive container fallbacks: Ensure that all file-read
operations on `/proc` files fail gracefully (returning the 1 TB baseline
fallback) if executed inside locked-down sandbox environments, Distrobox
environments, or strict container runtimes where `/proc` access is
restricted.

This runtime logic completely avoids hardcoding architecture-specific
blacklists, cleanly accommodates modern high-memory hardware
configurations, and prevents sudden worker terminations.

     Edmund

#1142815#57
Date:
2026-08-28 09:24:59 UTC
From:
To:
Thanks for looking into this! It will be helpful if someone tries to
fix 47-seg_size.t.

47-seg_size.t is not the current problem, as I have disabled it in
the package. The new problem is the one I mentioned in
https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1142815#27 :
Failures only on ppc64el in two different tests (t/72-shm_segments.t
and t/74-seg_map.t).

Cheers,
gregor

#1142815#68
Date:
2026-09-02 19:42:24 UTC
From:
To:
Because this is an RC bug and blocks others behind it, I have create an
MR on Salsa.

The bug for t-47 is now understood, and I therefore took the liberty to
re-enable the test. If you would rather not, I'll be happy to change
that again.

Upstream links:
https://github.com/stevieb9/ipc-shareable/issues/65
https://github.com/stevieb9/ipc-shareable/issues/66

Salsa MR:
https://salsa.debian.org/perl-team/modules/packages/libipc-shareable-perl/-/merge_requests/1
I hope this helps!

Kind regards,

     Edmund