Bug 4757 - Some build systems have inode count races when building meta-toolchain
Summary: Some build systems have inode count races when building meta-toolchain
Status: RESOLVED WONTFIX
Alias: None
Product: OE-Core
Classification: Build System, Metadata & Runtime
Component: devtools / tool chain (show other bugs)
Version: unspecified
Hardware: x86 Multiple
: Medium normal
Target Milestone: 1.5.1
Assignee: LiRongQing
QA Contact:
URL:
Whiteboard:
: 5287 (view as bug list)
Depends on:
Blocks:
 
Reported: 2013-06-21 06:03 UTC by Zhenhua Luo
Modified: 2014-01-16 16:22 UTC (History)
8 users (show)

See Also:
OS type for building Yocto: ---
Type of Regression: ---
Verified:
Documentation change: No (bug/feature does not impact docs)


Attachments
Log file of meta-toolchain build for qemux86 (582.20 KB, text/plain)
2013-06-21 06:03 UTC, Zhenhua Luo
no flags Details
local.conf file (10.57 KB, text/plain)
2013-06-26 02:46 UTC, Zhenhua Luo
no flags Details
meta-toolchain build log outside of Jenkins (30.97 KB, text/plain)
2013-06-27 02:06 UTC, Zhenhua Luo
no flags Details

Note You need to log in before you can comment on or make changes to this bug.
Description Zhenhua Luo 2013-06-21 06:03:56 UTC
Created attachment 1284 [details]
Log file of meta-toolchain build for qemux86

With the master(master:e04e6dab0afa8f72bf001b0cc106cd13b7694a91) of poky, "bitbake meta-toolchain" failed to build sometimes. 

The issue is not 100% reproducible each time, "bitbake meta-toolchain" can be repeated to reproduce the failure.

During our test, both BB_NUMBER_THREADS and PARALLEL_MAKE are set to 24, we tried the toolchain build for beagleboard, routerstationpro, mpc8315e-rdb, atom-pc and qemux86, the build error is reproduced on beagleboard(the 5th build), mpc8315e-rdb(the 10th build), atom-pc(the 3rd build) and qemux86(the 2nd time). Sometimes the issue happens for the 1st time.  

Following is the partial log, full log is attached(meta-toolchain-fail.log)
| log_check: Using /home/b19537/jenkins/workspace/yocto-upstream-toolchain/label/busy/machine/beagleboard/poky/master/tmp/work/armv7a-vfp-neon-poky-linux-gnueabi/meta-toolchain/1.0-r7/temp/log.do_populate_sdk.567 as logfile
| Logfile is clean
| DEBUG: Shell function populate_sdk_image finished
| DEBUG: SITE files ['endian-little', 'bit-32', 'arm-common', 'common-linux', 'common-glibc', 'arm-linux', 'arm-linux-gnueabi', 'common']
| DEBUG: Executing shell function create_sdk_files
| DEBUG: Shell function create_sdk_files finished
| DEBUG: Executing shell function tar_sdk
| tar: ./sysroots/x86_64-pokysdk-linux/var/lib/rpm/__db.002: file changed as we read it
| DEBUG: Python function do_populate_sdk finished
| ERROR: Function failed: tar_sdk (log file is located at /home/b19537/jenkins/workspace/yocto-upstream-toolchain/label/busy/machine/beagleboard/poky/master/tmp/work/armv7a-vfp-neon-poky-linux-gnueabi/meta-toolchain/1.0-r7/temp/log.do_populate_sdk.567)
Comment 1 Saul Wold 2013-06-24 23:29:53 UTC
Seems like this is a timing issue, but I can't find any way that do_package_sdk is getting paralleled
Comment 2 Laurentiu Palcu 2013-06-25 15:24:52 UTC
I'm testing on our build server with:

BB_NUMBER_THREADS = "40"
PARALLEL_MAKE = "-j 32"

and the build finished just fine 14 times now (and counting). MACHINE is qemux86. Do you have anything else different in your local.conf file besides those 2 variables?

FYI: while writing this it got to build no. 17...
Comment 3 Zhenhua Luo 2013-06-26 02:45:47 UTC
Please see attached local.conf. 

The OS on my build server is CentOS release 5.9.
Comment 4 Zhenhua Luo 2013-06-26 02:46:20 UTC
Created attachment 1291 [details]
local.conf file
Comment 5 Laurentiu Palcu 2013-06-26 11:24:05 UTC
I wasn't able to reproduce this on Ubuntu Server 12.04 with
BB_NUMBER_THREADS = "40"
PARALLEL_MAKE = "-j 32"

After I tried for qemux86 for 25 times, I switched to atom-pc and left it over night. It ran successfully for 728 times...

However, it appears that you're using Jenkins to do builds. Would you please run this oneliner, outside Jenkins environment?

i=0; while [ $? = 0 ]; do ((i = i + 1)); echo "################## Run no. $i ################"; MACHINE=atom-pc bitbake meta-toolchain; done

At least, you do the same test as I do and establish a baseline.

I've read do_populate_sdk() function several times and it doesn't seem to do any parallelism.
Comment 6 Zhenhua Luo 2013-06-26 14:17:02 UTC
I am running the build outside of Jenkins and will post the result when the test finished.
Comment 7 Zhenhua Luo 2013-06-27 02:06:44 UTC
Created attachment 1293 [details]
meta-toolchain build log outside of Jenkins
Comment 8 Zhenhua Luo 2013-06-27 02:07:22 UTC
The error appears in the 6th build, please see attached meta-toolchain-build-outside-Jenkins.log for full log.
Comment 9 Laurentiu Palcu 2013-06-27 15:28:05 UTC
We had an internal discussion about this issue and, apparently, there used to be some issues with pseudo running on older kernels. Something related to fsync syscalls... I'm CCing Peter Sebach, the pseudo maintainer, maybe he remembers what was the problem and can provide a workaround.
Comment 10 Seebs 2013-06-27 17:49:50 UTC
If all that's happening is tar mysteriously failing, and there's something about a file changing as it was being read, the thing we encountered was:

On some kernels (not sure whether older, newer, or what), and some filesystems, Linux will apparently update either file size or file mtime LONG after a file is done being written to. This is, obviously, a severe bug. Unfortunately, it's not a bug we can fix; the workaround is to forcibly fsync things.

In the WR build system, we were using a modified prelink which would fsync files after writing them, and we had to set the environment variable which prevents pseudo from disabling the fsyncs for this to work.
Comment 11 Laurentiu Palcu 2013-07-04 13:54:56 UTC
Zhenhua, any news on this? Were you able to find a workaround base on Peter's suggestions? I think the environment variable Peter was referring to is PSEUDO_ALLOW_FSYNC. Right Peter?
Comment 12 Zhenhua Luo 2013-07-05 01:54:49 UTC
I will give a test with "export PSEUDO_ALLOW_FSYNC=1" and provide the result.
Comment 13 Zhenhua Luo 2013-07-05 03:20:30 UTC
Laurentiu, the issue is reproducible with PSEUDO_ALLOW_FSYNC=1, following is my test command and the issue appears in the second build.

i=0; while [ $? = 0 ]; do ((i = i + 1)); echo "################## Run no. $i ################"; PSEUDO_ALLOW_FSYNC=1 MACHINE=qemux86 bitbake meta-toolchain; done

| log_check: Using /home/b19537/workspace/poky-oss/build-tc/tmp/work/i586-poky-linux/meta-toolchain/1.0-r7/temp/log.do_populate_sdk.4672 as logfile
| Logfile is clean
| DEBUG: Shell function populate_sdk_image finished
| DEBUG: SITE files ['endian-little', 'bit-32', 'ix86-common', 'common-linux', 'common-glibc', 'i586-linux', 'common']
| DEBUG: Executing shell function create_sdk_files
| DEBUG: Shell function create_sdk_files finished
| DEBUG: Executing shell function tar_sdk
| tar: ./sysroots/x86_64-pokysdk-linux/var/lib/rpm/__db.003: file changed as we read it
| DEBUG: Python function do_populate_sdk finished
Comment 14 Seebs 2013-07-08 14:18:19 UTC
I am not quite sure PSEUDO_ALLOW_FSYNC will work when set globally like that -- because I think it may not be in the whitelist of things that bitbake will pass through from the environment.
Comment 15 Laurentiu Palcu 2013-07-08 14:32:32 UTC
If we want an environment variable to be available to tasks, we can do an

export PSEUDO_ALLOW_FSYNC

in local.conf. This way, PSEUDO_ALLOW_FSYNC will be available to all tasks... But I'm not really sure pseudo will use it. It's worth trying though.
Comment 16 Zhenhua Luo 2013-07-09 02:15:59 UTC
(In reply to comment #15)
> If we want an environment variable to be available to tasks, we can do an
> 
> export PSEUDO_ALLOW_FSYNC
> 
> in local.conf. This way, PSEUDO_ALLOW_FSYNC will be available to all
> tasks... But I'm not really sure pseudo will use it. It's worth trying
> though.

I will give a try to export PSEUDO_ALLOW_FSYNC in local.conf.
Comment 17 Zhenhua Luo 2013-07-09 04:52:51 UTC
I added 'export PSEUDO_ALLOW_FSYNC="1"' and 'BB_ENV_EXTRAWHITE_append = " PSEUDO_ALLOW_FSYNC"' in local.conf, the issue appears in the 9th build. 

$ cat conf/local.conf  | grep PSEUDO_ALLOW_FSYNC
export PSEUDO_ALLOW_FSYNC="1"
BB_ENV_EXTRAWHITE_append = " PSEUDO_ALLOW_FSYNC"
$
$ bitbake meta-toolchain -e | grep BB_ENV_EXTRAWHITE
# $BB_ENV_EXTRAWHITE [3 operations]
BB_ENV_EXTRAWHITE="MACHINE DISTRO TCMODE TCLIBC HTTP_PROXY http_proxy HTTPS_PROXY https_proxy FTP_PROXY ftp_proxy FTPS_PROXY ftps_proxy ALL_PROXY all_proxy NO_PROXY no_proxy SSH_AGENT_PID SSH_AUTH_SOCK BB_SRCREV_POLICY SDKMACHINE BB_NUMBER_THREADS BB_NO_NETWORK PARALLEL_MAKE GIT_PROXY_COMMAND SOCKS5_PASSWD SOCKS5_USER SCREENDIR STAMPS_DIR PSEUDO_ALLOW_FSYNC"
#     "${BB_HASHBASE_WHITELIST} DATE TIME SSH_AGENT_PID SSH_AUTH_SOCK PSEUDO_BUILD BB_ENV_EXTRAWHITE DISABLE_SANITY_CHECKS PARALLEL_MAKE BB_NUMBER_THREADS BB_ORIGENV"
BB_HASHCONFIG_WHITELIST="TMPDIR FILE PATH PWD BB_TASKHASH BBPATH DL_DIR SSTATE_DIR THISDIR FILESEXTRAPATHS FILE_DIRNAME HOME LOGNAME SHELL TERM USER FILESPATH STAGING_DIR_HOST STAGING_DIR_TARGET COREBASE PRSERV_HOST PRSERV_DUMPDIR PRSERV_DUMPFILE PRSERV_LOCKDOWN PARALLEL_MAKE CCACHE_DIR EXTERNAL_TOOLCHAIN CCACHE CCACHE_DISABLE LICENSE_PATH DATE TIME SSH_AGENT_PID SSH_AUTH_SOCK PSEUDO_BUILD BB_ENV_EXTRAWHITE DISABLE_SANITY_CHECKS PARALLEL_MAKE BB_NUMBER_THREADS BB_ORIGENV"
Comment 18 Richard Purdie 2013-10-16 12:21:02 UTC
I suspect you need to both allow the fsync, *and* add a sync call before the tar command. You could probably also do soemthing like:

PSEUDO_UNLOAD=1 sync

before the tar. It doesn't change the fact that the underlying kernel is *seriously* broken and I'd rather the build system just said so and refused to run.
Comment 19 Zhenhua Luo 2013-10-16 13:10:39 UTC
(In reply to comment #18)
> I suspect you need to both allow the fsync, *and* add a sync call before the
> tar command. You could probably also do soemthing like:
> 
> PSEUDO_UNLOAD=1 sync
I will try the suggestion. 

> before the tar. It doesn't change the fact that the underlying kernel is
> *seriously* broken and I'd rather the build system just said so and refused
> to run.
The kernel version on my machine is 2.6.18-348.18.1.el5.centos.plus.
Comment 20 Seebs 2013-10-16 18:33:50 UTC
I still want to find out what's actually going on in the fairly rare failures we were seeing. In our case, what was doing it specifically was running the prelinker without fsync, and then every so often tar would claim that a file changed while being written. I would like to modify tar to say specifically what it thinks changed, and reproduce that some day. (The underlying problem we used to have with prelink was that the block count used by "du" and friends would change erratically seconds to minutes after a write was complete. Ugh!)
Comment 21 Zhenhua Luo 2013-10-17 08:30:33 UTC
(In reply to comment #19)
> (In reply to comment #18)
> > I suspect you need to both allow the fsync, *and* add a sync call before the
> > tar command. You could probably also do soemthing like:
> > 
> > PSEUDO_UNLOAD=1 sync
> I will try the suggestion. 
> 
> > before the tar. It doesn't change the fact that the underlying kernel is
> > *seriously* broken and I'd rather the build system just said so and refused
> > to run.
> The kernel version on my machine is 2.6.18-348.18.1.el5.centos.plus.

The issue doesn't appear after add PSEUDO_UNLOAD=1 sync before tar.
Comment 22 Richard Purdie 2013-10-31 17:49:11 UTC
*** Bug 5287 has been marked as a duplicate of this bug. ***
Comment 23 Richard Purdie 2014-01-16 16:22:43 UTC
This isn't something we're going to fix in the public project. Please use a system with a working filesystem. I appreciate they are workaround but they're too nasty to include in the main releases.