Reproducer: $ echo helloworld >> meta/recipes-extended/bash/bash_4.4.18.bb $ while true; do kill-bb; rm -fr bitbake-cookerdaemon.log tmp/cache/default-glibc/qemux86-64/x86_64/bb_cache.dat* ; bitbake -p; done It may hang in 10 mins, there are two problems: * There might be deadlocks when call process.join() if the queue is not NULL, so we need cleanup the queue before join() it, but: * The self.result_queue.get(timeout=0.25) may hang if the queue._wlock is hold by SomeOtherProcess, the queue has the following info when it hangs: '_wlock': <Lock(owner=SomeOtherProcess)>, so that we may can't clean up the queue. Here is a patch to workaround the issue: diff --git a/bitbake/lib/bb/cooker.py b/bitbake/lib/bb/cooker.py index 0607fcc708..e42bcbaedf 100644 --- a/bitbake/lib/bb/cooker.py +++ b/bitbake/lib/bb/cooker.py @@ -2082,19 +2082,15 @@ class CookerParser(object): for process in self.processes: self.parser_quit.put(None) - # Cleanup the queue before call process.join(), otherwise there might be - # deadlocks. - while True: - try: - self.result_queue.get(timeout=0.25) - except queue.Empty: - break - for process in self.processes: if force: process.join(.1) process.terminate() else: + # Kill the alive process firstly before join() it to avoid + # deadlocks + if process.is_alive(): + os.kill(process.pid, 9) process.join() sync = threading.Thread(target=self.bb_cache.sync)
We just branched the Zeus release of WR Linux. I was kicking the tires by doing a world build. I noticed that the filesystem was close to full so I started removing old builds and then decided to stop the world build by typing control-c once. I waited for more than 20 minutes but the build did not terminal and eventually I ran killall -9 -u <username> In another project directory, I re-tested and saw similar behaviour. The front-end user interface this time seems to be still running but it's been 12+ hours and it's still trying to shutdown: Waiting for 16 running tasks to finish: 0: perl-5.30.0-r0 do_unpack - 13h38m58s (pid 60610) 1: lib32-perl-5.30.0-r0 do_unpack - 13h38m57s (pid 60664) 2: rrdtool-1.7.2-r0 do_fetch (pid 64609) | <=> | 3: lib32-rrdtool-1.7.2-r0 do_fetch - 13h38m32s (pid 64663) 4: openscap-native-1.3.1+gitAUTOINC+c70bc47449-r0 do_fetch (pid 64977) | <=> | 5: openssl-native-1.1.1d-r0 do_configure - 13h38m13s (pid 98835) 6: unzip-native-1_6.0-r5 do_compile - 13h38m8s (pid 125085) 7: go-native-1.12.9-r0 do_configure - 13h38m5s (pid 4533) 8: postgresql-11.5-r0 do_unpack - 13h37m57s (pid 11216) 9: lib32-postgresql-11.5-r0 do_unpack - 13h37m57s (pid 11181) 10: m4-native-1.4.18-r0 do_configure - 13h37m49s (pid 21064) 11: lemon-native-3.7.3-r0 do_compile - 13h37m33s (pid 43739) 12: lib32-openscap-1.3.1+gitAUTOINC+c70bc47449-r0 do_fetch - 13h37m28s (pid 50079) 13: openscap-1.3.1+gitAUTOINC+c70bc47449-r0 do_fetch - 13h37m28s (pid 50211) 14: bjam-native-1.71.0-r0 do_compile - 13h37m17s (pid 64077) 15: lzip-native-1.21-r0 do_compile - 13h37m16s (pid 64084) I'll see if this is reproducible with just oe-core, as well as oe-core+meta-openembedded layers since it's possible that we have a commit in the WR Linux branches that isn't upstream.
Created attachment 4577 [details] strace -o /tmp/cooker-zombie.log -p 54732 -- Cooker process strace log and here's a partial pstree -p: systemd(1)─┬─Cooker(54732)─┬─Worker(55794)─┬─bjam-native:com(64077)───run.do_compile.(64181)───bootstrap.sh(64202)───build.sh(64222)───g++(64235) │ │ ├─go-native:confi(4533)───run.do_configur(5182)───bash(5225)───dist(7661)─┬─gcc(16416)─┬─as(16420) │ │ │ │ └─cc1(16419) │ │ │ ├─gcc(16910)─┬─as(16915) │ │ │ │ └─cc1(16914) │ │ │ └─gcc(17121)─┬─as(17126) │ │ │ └─cc1(17125)
We could not reproduce the problem on the same machine in another account or in my account on another machine. While the symptom may still be present in bitbake/oe-core, the problems that I'm seeing are likely an account/environment or filesystem setting.
We are unable to reproduce the error. Reopen if you see it.