Bug 13591 - bitbake may hang when it join()s a live process
Summary: bitbake may hang when it join()s a live process
Status: RESOLVED WORKSFORME
Alias: None
Product: BitBake
Classification: Build System, Metadata & Runtime
Component: bitbake (show other bugs)
Version: 3.1
Hardware: x86 Multiple
: Medium+ normal
Target Milestone: 3.2 M1
Assignee: Unassigned
QA Contact:
URL:
Whiteboard: BACKPORT
Depends on:
Blocks:
 
Reported: 2019-10-10 09:03 UTC by Robert Yang
Modified: 2020-04-23 08:14 UTC (History)
3 users (show)

See Also:
OS type for building Yocto: ---
Type of Regression: ---
Verified:
Documentation change: No (bug/feature does not impact docs)


Attachments
strace -o /tmp/cooker-zombie.log -p 54732 -- Cooker process (202.75 KB, text/plain)
2019-10-10 14:14 UTC, Randy MacLeod
no flags Details

Note You need to log in before you can comment on or make changes to this bug.
Description Robert Yang 2019-10-10 09:03:25 UTC
Reproducer:

$ echo helloworld >> meta/recipes-extended/bash/bash_4.4.18.bb
$ while true; do kill-bb; rm -fr bitbake-cookerdaemon.log tmp/cache/default-glibc/qemux86-64/x86_64/bb_cache.dat* ; bitbake -p; done

It may hang in 10 mins, there are two problems:
* There might be deadlocks when call process.join() if the queue is not NULL,
  so we need cleanup the queue before join() it, but:
* The self.result_queue.get(timeout=0.25) may hang if the queue._wlock is hold
  by SomeOtherProcess, the queue has the following info when it hangs:
  '_wlock': <Lock(owner=SomeOtherProcess)>, so that we may can't clean up the
  queue.

Here is a patch to workaround the issue:

diff --git a/bitbake/lib/bb/cooker.py b/bitbake/lib/bb/cooker.py
index 0607fcc708..e42bcbaedf 100644
--- a/bitbake/lib/bb/cooker.py
+++ b/bitbake/lib/bb/cooker.py
@@ -2082,19 +2082,15 @@ class CookerParser(object):
             for process in self.processes:
                 self.parser_quit.put(None)
 
-        # Cleanup the queue before call process.join(), otherwise there might be
-        # deadlocks.
-        while True:
-            try:
-               self.result_queue.get(timeout=0.25)
-            except queue.Empty:
-                break
-
         for process in self.processes:
             if force:
                 process.join(.1)
                 process.terminate()
             else:
+                # Kill the alive process firstly before join() it to avoid
+                # deadlocks
+                if process.is_alive():
+                    os.kill(process.pid, 9)
                 process.join()
 
         sync = threading.Thread(target=self.bb_cache.sync)
Comment 1 Randy MacLeod 2019-10-10 14:10:10 UTC
We just branched the Zeus release of WR Linux.
I was kicking the tires by doing a world build.
I noticed that the filesystem was close to full so I started removing old builds
and then decided to stop the world build by typing control-c once.
I waited for more than 20 minutes but the build did not terminal and eventually
I ran killall -9 -u <username>

In another project directory, I re-tested and saw similar behaviour. 
The front-end user interface this time seems to be still running but it's been
12+ hours and it's still trying to shutdown:
Waiting for 16 running tasks to finish:
0: perl-5.30.0-r0 do_unpack - 13h38m58s (pid 60610)
1: lib32-perl-5.30.0-r0 do_unpack - 13h38m57s (pid 60664)
2: rrdtool-1.7.2-r0 do_fetch (pid 64609) |                                   <=>                                                                     |
3: lib32-rrdtool-1.7.2-r0 do_fetch - 13h38m32s (pid 64663)
4: openscap-native-1.3.1+gitAUTOINC+c70bc47449-r0 do_fetch (pid 64977) |             <=>                                                             |
5: openssl-native-1.1.1d-r0 do_configure - 13h38m13s (pid 98835)
6: unzip-native-1_6.0-r5 do_compile - 13h38m8s (pid 125085)
7: go-native-1.12.9-r0 do_configure - 13h38m5s (pid 4533)
8: postgresql-11.5-r0 do_unpack - 13h37m57s (pid 11216)
9: lib32-postgresql-11.5-r0 do_unpack - 13h37m57s (pid 11181)
10: m4-native-1.4.18-r0 do_configure - 13h37m49s (pid 21064)
11: lemon-native-3.7.3-r0 do_compile - 13h37m33s (pid 43739)
12: lib32-openscap-1.3.1+gitAUTOINC+c70bc47449-r0 do_fetch - 13h37m28s (pid 50079)
13: openscap-1.3.1+gitAUTOINC+c70bc47449-r0 do_fetch - 13h37m28s (pid 50211)
14: bjam-native-1.71.0-r0 do_compile - 13h37m17s (pid 64077)
15: lzip-native-1.21-r0 do_compile - 13h37m16s (pid 64084)

I'll see if this is reproducible with just oe-core, as well as oe-core+meta-openembedded layers since it's possible that we have a commit in the WR Linux branches that isn't upstream.
Comment 2 Randy MacLeod 2019-10-10 14:14:27 UTC
Created attachment 4577 [details]
strace -o /tmp/cooker-zombie.log -p 54732 -- Cooker process

strace log and here's a partial pstree -p:
systemd(1)─┬─Cooker(54732)─┬─Worker(55794)─┬─bjam-native:com(64077)───run.do_compile.(64181)───bootstrap.sh(64202)───build.sh(64222)───g++(64235)
           │               │               ├─go-native:confi(4533)───run.do_configur(5182)───bash(5225)───dist(7661)─┬─gcc(16416)─┬─as(16420)
           │               │               │                                                                         │            └─cc1(16419)
           │               │               │                                                                         ├─gcc(16910)─┬─as(16915)
           │               │               │                                                                         │            └─cc1(16914)      
           │               │               │                                                                         └─gcc(17121)─┬─as(17126)
           │               │               │                                                                                      └─cc1(17125)
Comment 3 Randy MacLeod 2019-10-10 22:23:24 UTC
We could not reproduce the problem on the same machine in another account or in my account on another machine.
While the symptom may still be present in bitbake/oe-core, the problems that I'm seeing are likely an account/environment or filesystem setting.
Comment 4 Randy MacLeod 2020-04-23 08:14:04 UTC
We are unable to reproduce the error. Reopen if you see it.