Bug 15647 - glibc 2.35 causes meson configure failures
Summary: glibc 2.35 causes meson configure failures
Status: RESOLVED FIXED
Alias: None
Product: OE-Core
Classification: Build System, Metadata & Runtime
Component: core (show other bugs)
Version: 4.0.22
Hardware: x86 x86_64
: High major
Target Milestone: 4.0.26
Assignee: Stephan Wolf
QA Contact:
URL:
Whiteboard:
Depends on:
Blocks:
 
Reported: 2024-11-12 07:38 UTC by Stephan Wolf
Modified: 2025-01-15 16:36 UTC (History)
12 users (show)

See Also:
OS type for building Yocto: ---
Type of Regression: ---
Verified:
Documentation change: Don't know


Attachments
bitbake build log (4.21 KB, text/plain)
2024-11-12 07:38 UTC, Stephan Wolf
no flags Details
glibc patch to revert commit a180e82837517c126b2c083829791d7b1a135511 (5.46 KB, patch)
2024-11-14 16:45 UTC, Stephan Wolf
no flags Details | Diff
meson-log.txt (243.46 KB, text/plain)
2024-11-18 06:15 UTC, Deepesh Varatharajan
no flags Details
qemu bbappend to use vanilla qemu 7.2.15 (311 bytes, text/plain)
2024-11-22 14:42 UTC, Stephan Wolf
no flags Details

Note You need to log in before you can comment on or make changes to this bug.
Description Stephan Wolf 2024-11-12 07:38:04 UTC
Created attachment 5079 [details]
bitbake build log

Hello,

Kirkstone 4.0.22 includes a gcc update from version 11.4 to version 11.5.

This update breaks the build of fwupd. If we switch back to the commit before the update everything is working fine.

We trying to cleanup the sstate-cache by removing duplicates and also to remove the entire content but no success.

Thanks a lot for any help
Comment 1 Ross Burton 2024-11-14 15:47:42 UTC
Quoting from the log:

../fwupd-1.7.6/meson.build:1:0: ERROR: Executables created by c compiler x86_64-kontron-linux-gcc -m64 -march=nehalem -mtune=generic -mfpmath=sse -msse4.2 -fstack-protector-strong -O2 -D_FORTIFY_SOURCE=2 -Wformat -Wformat-security -Werror=format-security --sysroot=/home/stephanw/Repos/kontronos/build/tmp/work/kontron_kbox_a250-kontron-linux/fwupd/1.7.6-r0/recipe-sysroot are not runnable.

So it's not a compiler error, but qemu-user is failing to run the test binary.

It would be useful to have a look at the meson log file itself in the build directory to see if it gives any hints. My hunch is that it's failing on an illegal instruction exception as the new GCC is using instructions that the qemu doesn't support.
Comment 2 Ross Burton 2024-11-14 16:24:25 UTC
Of course if I'm right then every meson package will fail.  Does eg cairo successfully build?
Comment 3 Stephan Wolf 2024-11-14 16:45:24 UTC
Created attachment 5080 [details]
glibc patch to revert commit a180e82837517c126b2c083829791d7b1a135511
Comment 4 Stephan Wolf 2024-11-14 16:47:35 UTC
We found out the commit a180e82837517c126b2c083829791d7b1a135511 of glibc in release/2.35/master branch is the root cause.

https://sourceware.org/git/?p=glibc.git;a=commit;h=a180e82837517c126b2c083829791d7b1a135511

gcc 11.4. was also affected by the error. 

So the bug is related to glibc
Comment 5 Ross Burton 2024-11-14 16:48:47 UTC
I suggest you talk to glibc about whether there are further fixes we should backport instead, as a plain revert is just papering over the issue.
Comment 6 Randy MacLeod 2024-11-15 17:36:22 UTC
Add Sundeep and Deepthi to CC.

Stephan, 

Will you contact glibc upstream or 
would you like Sundeep to take care of that 
once we agree on a simple reproducer.
Comment 7 Randy MacLeod 2024-11-15 17:39:04 UTC
Oops, it's Deepesh who owns the bug now so he'd work with upstream I presume if that is what Stephan wants to do.
Comment 8 Stephan Wolf 2024-11-15 20:41:21 UTC
I don‘t know who is the right responsible but I think the bug is related to glibc. That is what I want to change.
Comment 9 Stephan Wolf 2024-11-15 20:43:53 UTC
(In reply to Ross Burton from comment #2)
> Of course if I'm right then every meson package will fail.  Does eg cairo
> successfully build?

Other meson packages also affected.
Comment 10 Randy MacLeod 2024-11-15 21:26:44 UTC
Stephan,

Okay, my take is that you're happy to have Deepesh (and Sundeep) take care of things. They are trying to reproduce the issue and then come up with a non-bitbake reproducer. If you can help with that, that would be super but if not, we'll figure it out

We'll know more early next week, I expect.
Comment 11 Deepesh Varatharajan 2024-11-18 06:15:41 UTC
Created attachment 5081 [details]
meson-log.txt

Hello,
 
We've tried reproducing the issue on Kirkstone with x86_64-poky-linux-gcc and for us the sanitycheckc_cross.exe is executing successfully. (See the attached meson-log.txt file)
We've one observation that in our case -march=core2 instead of nehalem (We added local.conf with - DEFAULTTUNE = "core2-64", TUNE_FEATURES:append = " nehalem")
Can you let us know we're missing something? or share the exact reproducible steps (it would be good to have a minimal reproducer to share with glibc community).

Regards,
Deepesh
Comment 12 Ross Burton 2024-11-18 10:11:10 UTC
The kontron references and nehalem make be suspicious too.  Stephan: can you replicate with the qemux86-64 MACHINE?  If no, can you share your machine configuration?
Comment 13 Stephan Wolf 2024-11-18 13:07:46 UTC
Hello all,

We are using the following config 

DEFAULTTUNE ?= "corei7-64"
require conf/machine/include/intel-corei7-64-common.inc # from meta-intel layer

results in TUNE_FEATURES="m64 corei7"

Changing it to

DEFAULTTUNE ?= "core2-64"
require conf/machine/include/intel-core2-32-common.inc # from meta-intel layer

results in TUNE_FEATURES="m64 core2"

works as expected. 

I hope this helps

Regards,
Stephan
Comment 14 Ross Burton 2024-11-18 13:10:43 UTC
Feels like the problem is that qemu-user is 'suboptimal' and fails to handle instructions that the new glibc is using.

Try setting QEMU_EXTRAOPTIONS to "-cpu Nehalem" to tell the qemu-user script to pass that option.
Comment 15 Ross Burton 2024-11-18 13:12:36 UTC
If that works then the proper fix is to add that assignment to meta-intel.
Comment 16 Ross Burton 2024-11-18 13:24:24 UTC
Hm I take that back:

meta/conf/machine/include/x86/tune-corei7.inc:QEMU_EXTRAOPTIONS_corei7-64 = " -cpu Nehalem,check=false"
Comment 17 Deepesh Varatharajan 2024-11-19 04:53:07 UTC
Hello,

It seems like it is working fine after following Ross Burton comments or if the issue is still there, Could you please share the exact configuration needed to reproduce the issue.

Regards,
Deepesh
Comment 18 Deepesh Varatharajan 2024-11-20 09:49:04 UTC
Hello,

Just a gentle reminder to provide the reproducer. We are still unable to reproduce the issue on our end.

Thanks,
Deepesh
Comment 19 Stephan Wolf 2024-11-21 11:12:45 UTC
(In reply to Deepesh Varatharajan from comment #17)
> Hello,
> 
> It seems like it is working fine after following Ross Burton comments or if
> the issue is still there, Could you please share the exact configuration
> needed to reproduce the issue.
> 
> Regards,
> Deepesh

The QEMU setting is already set. So it seems this is not the solution as Ross Burton already mentioned. 

What I found is the following:

core2 uses TUNE_FEATURES -march=core2 -mtune=core2 -msse3 -mfpmath=sse

corei7 uses TUNE_FEATURES -march=nehalem -mtune=generic -mfpmath=sse -msse4.2

could it be that -mtune=generic is wrong for nehalem or any thing else?

Regards,
Stephan
Comment 20 Ross Burton 2024-11-21 11:15:57 UTC
Copying in the meta-intel maintainers as the finer details of GCC tune flags on x86 is their speciality.  Anuj, Naveen: can you have a look at this?
Comment 21 Randy MacLeod 2024-11-21 15:59:42 UTC
Anuj, Can you help here?
Comment 22 Ross Burton 2024-11-21 16:02:00 UTC
From the gcc docs:

"""
‘generic’
Produce code optimized for the most common IA32/AMD64/EM64T processors. If you know the CPU on which your code will run, then you should use the corresponding -mtune or -march option instead of -mtune=generic. But, if you do not know exactly what CPU users of your application will have, then you should use this option.

As new processors are deployed in the marketplace, the behavior of this option will change. Therefore, if you upgrade to a newer version of GCC, code generation controlled by this option will change to reflect the processors that are most common at the time that version of GCC is released.

There is no -march=generic option because -march indicates the instruction set the compiler can use, and there is no generic instruction set applicable to all processors. In contrast, -mtune indicates the processor (or, in this case, collection of processors) for which the code is optimized.

"""

So -march=generic _shouldn't_ change the instructions used. Quite possibly a qemu bug where it's missing instruction handling.
Comment 23 Stephan Wolf 2024-11-21 20:05:00 UTC
(In reply to Ross Burton from comment #22)
> From the gcc docs:
> 
> """
> ‘generic’
> Produce code optimized for the most common IA32/AMD64/EM64T processors. If
> you know the CPU on which your code will run, then you should use the
> corresponding -mtune or -march option instead of -mtune=generic. But, if you
> do not know exactly what CPU users of your application will have, then you
> should use this option.
> 
> As new processors are deployed in the marketplace, the behavior of this
> option will change. Therefore, if you upgrade to a newer version of GCC,
> code generation controlled by this option will change to reflect the
> processors that are most common at the time that version of GCC is released.
> 
> There is no -march=generic option because -march indicates the instruction
> set the compiler can use, and there is no generic instruction set applicable
> to all processors. In contrast, -mtune indicates the processor (or, in this
> case, collection of processors) for which the code is optimized.
> 
> """
> 
> So -march=generic _shouldn't_ change the instructions used. Quite possibly a
> qemu bug where it's missing instruction handling.

I'm of the same opinion. 

I had a little more deeper lookmin the core and the stacktrace:

I used -march=nehalem -mtune=nehalem -msse3 -mfpmath=sse 

and the invalid instruction is always a sse4.1 SIMD instruction

0x400181a31a <_dl_check_map_versions+650>       pmaxud %xmm0,%xmm1 

backtrace told me 

#0  0x000000400181a31a in _dl_check_map_versions (map=map@entry=0x400183d2b0, verbose=verbose@entry=1, trace_mode=trace_mode@entry=0) at dl-version.c:227

This should not happen in in one hand and in the other hand qemu with -cpu Nehalem should handle this instruction well.

Regards,
Stephan
Comment 24 Stephan Wolf 2024-11-22 08:49:03 UTC
Hi all,

Comment from myself: gcc with -mtune=nehalem always generates sse4 instructions where possible. If I switch it off with -mno-sse4.1 build is successful.

Regards,
Stephan
Comment 25 Stephan Wolf 2024-11-22 14:42:48 UTC
Created attachment 5085 [details]
qemu bbappend to use vanilla qemu 7.2.15
Comment 26 Stephan Wolf 2024-11-22 14:58:52 UTC
Hi all,

I used the next supported version 7.2.15 of qemu and the build is successful. 

https://bugzilla.yoctoproject.org/attachment.cgi?id=5085

We have five different machines compiling the image. Three of them are used for CI/CD, two cloud hosted AMD machines and one Intel Vmware hosted the other two are developer machines with intel laptop cpu and one amd desktop machine. All of them have the same issue, which tells me that can not be our setup.

Regards,
Stephan
Comment 27 Stephan Wolf 2024-11-27 12:50:29 UTC
(In reply to Ross Burton from comment #2)
> Of course if I'm right then every meson package will fail.  Does eg cairo
> successfully build?

Every meson build was effected. If I remove fwupd temporarily the next meson package was rauc which also failed to build.
Comment 28 Randy MacLeod 2024-12-05 15:51:41 UTC
We cant upversion qemu, it is unfeasible to backport the SSE support, 
disable the SSE support in this tune seems like a bad idea.

We could disable usermode qemu, perhaps.


the only other solution people agreeed on is to crate a meta-ltx-mixin layer with a newer version of qemu.

Ross and Stephan need to decide on a plan.
Comment 29 Enrico Jorns 2024-12-17 20:30:35 UTC
Hi all,

we just ran into this in one of our customer projects, too.
Here it is also a meson-based RAUC build that fails in the same configuration (corei7 tune).

I traced it down to the already-mentioned glibc commit before finding this issue.

Have you already made any plans for how to fix or work around it?

I can also confirm that using '-mno-sse4.1' makes the build succeed for me, too.
(As well as reverting the glibc patch would...)

Regards,
Enrico
Comment 30 Richard Purdie 2024-12-19 16:25:04 UTC
Our options are:

a) Add -mno-sse4.1 to the tune. This is a shame for real hardware but if qemu can't cope with it...

b) We could disable qemu-user mode for meson. I think that would then trigger safe fallbacks. This might be the best option if we can work out a patch to do it.

Upgrading qemu to gain the usermode SSE support is probably too invasive.

Backporting the SSE implementation for usermode qemu is probably also too invasive/risky.

Has anyone tried b)?
Comment 31 Ross Burton 2024-12-19 16:30:04 UTC
<cough> I have a branch...

https://git.yoctoproject.org/poky-contrib/commit/?h=ross/meson&id=9211b8ce0dca7dec9276d2c608ad3ad9eb2c5a0f makes meson emit a warning when it's using a fallback value instead of running code.  At the time that I did that test, the only user was graphene:

https://git.yoctoproject.org/poky-contrib/commit/?h=ross/meson&id=7938ce009dfb8bc101d1476da7db31cbdbe7a3d9

Those commits are from summer 2023 so not ancient and somewhat relevant for kirkstone.
Comment 32 Randy MacLeod 2025-01-09 16:22:00 UTC
Ross to a test build with his branch.
Comment 33 Ross Burton 2025-01-09 17:23:46 UTC
<mind blown>

So I can replicate this with a modified qemux86-64 machine if I build fwupd.  What confuses me is how it gets that far, considering the dependency glib-2.0 should have failed in the same way.

Then I noticed that changing fwupd packages from being machine-specific to tune-specific fixed the issue and now I'm even more confused than before.
Comment 34 Ross Burton 2025-01-09 17:41:06 UTC
qemu.bbclass does this:

  QEMU_OPTIONS[vardeps] += "QEMU_EXTRAOPTIONS_${PACKAGE_ARCH}"

And the tunes do this:

  QEMU_EXTRAOPTIONS_corei7-64 = " -cpu Nehalem,check=false"

But corei7-64 isn't the PACKAGE_ARCH when a recipe sets the package arch to the MACHINE.

I note that Khem discovered this problem three years ago and just worked around it for ppc in qemu.bbclass:

  # Some packages e.g. fwupd sets PACKAGE_ARCH = MACHINE_ARCH and uses meson which
  # needs right options to usermode qemu
  QEMU_EXTRAOPTIONS_qemuppc = " -cpu 7400"
  QEMU_EXTRAOPTIONS_qemuppc64 = " -cpu POWER9"

That's wrong, and this should be fixed properly.  I'm presuming we don't see this in master as newer qemu is more accepting about what instructions it handles.
Comment 35 Ross Burton 2025-01-09 17:54:40 UTC
A horrible workaround is to tell qemu to use the right flags when building for your machine architecture. Try adding this to your local.conf:

QEMU_EXTRAOPTIONS_kontron_kbox_a250 = " -cpu Nehalem,check=false"
Comment 36 Ross Burton 2025-01-10 13:04:36 UTC
https://lore.kernel.org/openembedded-core/20250110130116.894607-1-ross.burton@arm.com/T/#t is the fix for master, and I'll post the backport shortly.
Comment 37 Ross Burton 2025-01-10 13:15:01 UTC
Stephan and Enrico: if you could test https://lore.kernel.org/openembedded-core/20250110131339.924678-1-ross.burton@arm.com/T/#u that would be _much_ appreciated.
Comment 38 Enrico Jorns 2025-01-10 22:07:12 UTC
Thank you Ross for the further investigation!

I can confirm that the patch fixes my RAUC build in kirkstone.
Comment 39 Stephan Wolf 2025-01-14 06:52:59 UTC
It's working for me too.

Thank you all for the effort.

Regards
Comment 40 Ross Burton 2025-01-15 16:36:18 UTC
Fixed in master 414b754a6cbb9cc354b1180efd5c3329568a2537 and backports were sent.