Bug 11648 - python doesn't parse multilib headers well
Summary: python doesn't parse multilib headers well
Status: RESOLVED WONTFIX
Alias: None
Product: OE-Core
Classification: Build System, Metadata & Runtime
Component: devtools / tool chain (show other bugs)
Version: 2.3
Hardware: Other arm
: Medium+ normal
Target Milestone: 4.99
Assignee: Ross Burton
QA Contact:
URL:
Whiteboard:
Depends on:
Blocks:
 
Reported: 2017-06-12 21:11 UTC by Tanner Oakes
Modified: 2019-10-24 22:50 UTC (History)
7 users (show)

See Also:
OS type for building Yocto: ---
Type of Regression: ---
Verified:
Documentation change: No (bug/feature does not impact docs)


Attachments

Note You need to log in before you can comment on or make changes to this bug.
Description Tanner Oakes 2017-06-12 21:11:55 UTC
Python generates some platform code in /usr/lib/plat-linux. It parses some header files using an internal h2py tool. This tool is brain dead simple and doesn't handle #ifdefs. When parsing a header file that has been replaced by a multilib header it generates junk. It doesn't fail to build, but at runtime when using network functions (and probably others) it causes python to crash.

Here is an example of the problem from /usr/lib/plat-linux/IN.py

# Included from bits/wordsize.h
__MHWORDSIZE = 32
__MHWORDSIZE = 64
__MHWORDSIZE = 32
__MHWORDSIZE = __WORDSIZE

What is expected:
# Included from bits/wordsize.h
__WORDSIZE = 32


This affects both python and python3.
Comment 1 brian avery 2017-06-15 14:49:50 UTC
Could you detail the MACHINE/image you built for and an example of one of the python network functions that failed?  It makes replication easier.
Comment 2 Tanner Oakes 2017-06-15 15:10:02 UTC
(In reply to comment #1)
> Could you detail the MACHINE/image you built for and an example of one of
> the python network functions that failed?  It makes replication easier.

I am building for a custom board, but it is similar to an imx6dlsabresd/fsl-image-machine-test.bb

We don't need multilib support because we only are compiling for a 32bit machine, but the mutlilib headers are being created anyways.

In a python script all that is need to fail is

import IN
Comment 3 Juan Manuel Cruz Alcaraz 2017-07-18 16:59:32 UTC
This bug is interesting, there are different levels to approach it.
The root of it is the current implementation for h2py script. It does not recognize #ifdef statements because of clearly technical challenges.

The main challenge is to allow the h2py script to fully understand C macros, which would imply to have a full macro preprocessing step implemented in the script. Nevertheless the following levels describe partial solutions that might help to mitigate the effect that is causing with multilib headers.

Level 1. #ifdef statements are recognized for macros defined as a parameter in the h2py script. This would allow to apply #ifdef to macros defined in the C compiler with the -D parameter. Do not include nested #ifdefs.
Level 2. #ifdef statements are recognized for macros defined inside the same file. Do not include nested #ifdefs.
Level 3. #ifdef statements are recognized for macros defined in files inside a CLIB directory.Do not include nested #ifdefs.
Level 4. Recognize nested #ifdef.
Level 5. Recognize complex macros statements in #ifdefs.

For this particular bug (recognize multilib headers) Level 1 approach should be good enough, since the #ifdef blocks are dependent in architecture definition which is defined as compiler parameter.

The other levels add complexity that do not add value to this particular bug goal. 

Comments?
Comment 4 Juan Manuel Cruz Alcaraz 2017-07-19 17:08:13 UTC
I'm exploring the possibility to use the compiler itself to process macros.

The h2py script is a recursive script that transform every header file given as an argument. It progresses recursively over all the #included files.

The idea is to include a transformation to each file (previous to the recursion call). The preprocessing would call the compiler to transform the file and process the macros applying the precompiler. The resulting preprocessed file would not have #ifdef statements because the compiler took care, for instance with a gcc toolchain:

gcc -E -dD -dI -P -D__arm__ usr/include/bits/syscall.h 

would call the precompiler only and would process all #ifdefs.

The big caveat is that not all toolchains would support this feature or maybe they would support it only partially. An initial implementation would solve the issue for gcc-family toolchains and it should be extended for other toolchains.

Comments?
Comment 5 Alejandro Hernandez 2017-07-19 17:29:47 UTC
(In reply to comment #4)
> I'm exploring the possibility to use the compiler itself to process macros.
> 
> The h2py script is a recursive script that transform every header file given
> as an argument. It progresses recursively over all the #included files.
> 
> The idea is to include a transformation to each file (previous to the
> recursion call). The preprocessing would call the compiler to transform the
> file and process the macros applying the precompiler. The resulting
> preprocessed file would not have #ifdef statements because the compiler took
> care, for instance with a gcc toolchain:
> 
> gcc -E -dD -dI -P -D__arm__ usr/include/bits/syscall.h 
> 
> would call the precompiler only and would process all #ifdefs.
> 
> The big caveat is that not all toolchains would support this feature or
> maybe they would support it only partially. An initial implementation would
> solve the issue for gcc-family toolchains and it should be extended for
> other toolchains.
> 
> Comments?

I completely agree with this, while we could parse the output of the script and get rid of the "junk", letting the compiler take care of it would be a better approach in my opinion, but we first need to evaluate if its worth it, my point being that if its not a lot of junk and its easy to identify what part of the generated file is correct it might just be easier to parse it ourselves and clean it up depending on the build configuration, if it proves to be rather difficult to identify these parts then its likely a better solution would be to use the compiler for this, since we already know its gonna do it correctly.

Also, it might be worth investigating how others solve this problem (there's no need to reinvent the wheel), any distro with multilib builds would have this sort of problem, or perhaps theres something documented on http://bugs.python.org
Comment 6 Juan Manuel Cruz Alcaraz 2017-07-25 19:29:38 UTC
(In reply to comment #5)
> (In reply to comment #4)
> > I'm exploring the possibility to use the compiler itself to process macros.
> > 
> 
> I completely agree with this, while we could parse the output of the script
> and get rid of the "junk", letting the compiler take care of it would be a
> better approach in my opinion, but we first need to evaluate if its worth
> it, my point being that if its not a lot of junk and its easy to identify
> what part of the generated file is correct it might just be easier to parse
> it ourselves and clean it up depending on the build configuration, if it
> proves to be rather difficult to identify these parts then its likely a
> better solution would be to use the compiler for this, since we already know
> its gonna do it correctly.
> 
> Also, it might be worth investigating how others solve this problem (there's
> no need to reinvent the wheel), any distro with multilib builds would have
> this sort of problem, or perhaps theres something documented on
> http://bugs.python.org

I ran some tests using the compiler to preprocess the headers. It does not look a good option. The IN.PY file generated after the preprocessing have important differences from the non-preprocessed IN.PY. It can be a source of incompatibilities from all the different scenarios where the IN.PY file is used. Also depending on the compiler used in cross-compilations the results might diverge in different ways.

I follow your advice and looked for the current status of this issue in the Python rooms and I found the following tickets:

Title: 	Cross compilation fails in regen
http://bugs.python.org/issue28018 

Title: 	fix running regen in cross 
http://bugs.python.org/issue17031 builds

Title: 	h2py.py: search the multiarch include dir if it does exist
http://bugs.python.org/issue17029 

They are aware of different issues with cross-compilation with the regen/h2py.py processes. Issue 28018 follow up this issues. In Python 3.6 and 3.7 regen is not used anymore, so they are not following up these issues in Python 3.5 (last version using regen).
Personally I don't think there is value pursuing to fix the regen scripts since it seems to be a well known restriction for Python 3.5 and below.
I will move this ticket to "Won't do". Let me know if someone feels it is still needed to fix for Python 3.5 and below.
Comment 7 Ross Burton 2019-10-24 22:50:00 UTC
As per last comment, closing.