Bug 9532 - useradd / extrausers chain pseudo, causing stack overflows (segfaults) [s390]
Summary: useradd / extrausers chain pseudo, causing stack overflows (segfaults) [s390]
Status: RESOLVED OBSOLETE
Alias: None
Product: OE-Core
Classification: Build System, Metadata & Runtime
Component: core (show other bugs)
Version: unspecified
Hardware: Other other
: Medium normal
Target Milestone: Future
Assignee: Unassigned
QA Contact:
URL:
Whiteboard:
Depends on:
Blocks:
 
Reported: 2016-04-27 18:16 UTC by Sascha Silbe
Modified: 2020-10-01 08:30 UTC (History)
6 users (show)

See Also:
OS type for building Yocto: ---
Type of Regression: New (Never tested)
Verified:
Documentation change: No (bug/feature does not impact docs)


Attachments

Note You need to log in before you can comment on or make changes to this bug.
Description Sascha Silbe 2016-04-27 18:16:41 UTC
Chaining pseudo instances will cause stack overflows (visible as segfaults) on s390x (and maybe other architectures). This is because the same symbols live both in the LD_PRELOAD'ed shared library as well as in the executable and the linker will happily interleave invocations of non-static functions, leading to different copies of static variables having different values and the code just going wild. I haven't tried figuring out why this happens to work on x86_64. May be differences in linker behaviour or just luck.

The easiest fix is to teach the useradd and extrausers classes to only chain the commands with pseudo if they're not already running inside pseudo (by checking for PSEUDO_PREFIX before setting PSEUDO and allowing for PSEUDO to be empty by using the ${PSEUDO:-} syntax).

It may also be possible to fix pseudo itself to be structured and / or linked in a different way so that there's only a single instance of each symbol. I haven't tried going down that route.

Example error (while installing dbus):

=== Begin ===
DEBUG: SITE files ['endian-big', 'bit-64', 's390x-common', 'common-linux', 'common-glibc', 's390x-linux', 'common']
DEBUG: Executing shell function useradd_sysroot
Running groupadd commands...
NOTE: dbus: Performing groupadd with [--root /home/silbe/poky/build-master/tmp/sysroots/qemus390x -r netdev]
Segmentation fault
WARNING: exit code 1 from a shell command.
ERROR: dbus: groupadd command did not succeed.
ERROR: Function failed: useradd_sysroot (log file is located at /home/silbe/poky/build-master/tmp/work/s390x-poky-linux/dbus/1.10.6-r0/temp/log.do_install.176717)
=== End ===
Comment 1 Ross Burton 2016-04-28 16:20:32 UTC
Peter: can pseudo be fixed to handle this behaviour on S390?
Comment 2 Seebs 2016-04-28 17:07:27 UTC
Interesting question. In theory, pseudo should be trying to detect that it's running under pseudo, and re-execing itself, but if it segfaults before that, it probably can't.

It looks like this can be reproduced without s390 hardware, with qemu? If so, I could try to reproduce it and fix it.

At the very least, we'd almost certainly have to have the version that's fixed so that the client can spawn things correctly, or it'd still blow up sooner or later anyway.

It's sort of surprising that the linker isn't consistent about which versions of things it picks, though.
Comment 3 Sascha Silbe 2016-04-28 18:49:18 UTC
IIRC it segfaulted too early for some re-exec trick. (But it's been a rather lengthy and confusing debug session some time ago, so my memory could be incorrect).

In cause you're trying to reproduce it: This time around I used Debian Jessie. Might also have happened with some Fedora version, but I'm not sure.
Comment 4 Seebs 2016-04-29 20:01:58 UTC
Do I need anything fancy set up to reproduce this, assuming I don't have an actual S390 machine?
Comment 5 Sascha Silbe 2016-05-03 10:55:06 UTC
You can request access to S390 hardware at https://developer.ibm.com/linuxone/ . I had some modifications for building images for S390 applied locally; will try building an x86_64 image with an unmodified poky version now.
Comment 6 Sascha Silbe 2016-05-03 11:14:54 UTC
Building an x86_64 image on s390x with a pristine checkout fails, of course. Obvious in retrospect: First thing, bitbake builds a native toolchain. So to actually encounter this bug, you'd need to add s390 support first... :-/
Comment 7 Sascha Silbe 2016-05-04 14:40:35 UTC
OK, got it reduced to a minimal set of changes for reproducing the
pseudo issue:

- meta/classes/siteinfo.bbclass: s390x is a big-endian 64-bit
  architecture. The GNU target triplet as output by config.guess is
  s390x-ibm-linux-gnu.

- meta/classes/kernel-arch.bbclass: the kernel uses ARCH=s390 for
  s390x. Linux used to support both 31/32-bit machines AKA s390 and
  64-bit machines AKA s390x via a config option, but nowadays only
  64-bit machines are supported.

- meta/recipes-connectivity/openssl/openssl.inc: OpenSSL uses target
  linux64-s390x for s390x

- meta/classes/uninative.bbclass: There's no uninative tarball yet for
  s390x and relocate_sdk.py seems to be broken for big-endian machines
  anyway, so uninative needs to be disabled for BUILD_ARCH=s390x.

After modifying the poky source using the above information, sourcing
oe-init-build-env once to create the default configuration and setting
MACHINE=x86_64 in build/conf/local.conf, the following invocation
reproduces the pseudo segfault:

( . ./oe-init-build-env build && bitbake dbus )
Comment 8 Randy MacLeod 2020-10-01 08:30:33 UTC
Re=open if you are seeing this still. We don't have access to any s390 systems.