Bug 12117 - bitbake servers should be more responsive.
Summary: bitbake servers should be more responsive.
Status: RESOLVED OBSOLETE
Alias: None
Product: BitBake
Classification: Build System, Metadata & Runtime
Component: bitbake (show other bugs)
Version: 2.5
Hardware: x86 Multiple
: Medium enhancement
Target Milestone: 4.99
Assignee: Randy MacLeod
QA Contact:
URL:
Whiteboard:
Depends on:
Blocks:
 
Reported: 2017-09-22 20:00 UTC by Randy MacLeod
Modified: 2020-12-03 16:18 UTC (History)
3 users (show)

See Also:
OS type for building Yocto: ---
Type of Regression: ---
Verified:
Documentation change: Yes (doc changes required)


Attachments

Note You need to log in before you can comment on or make changes to this bug.
Description Randy MacLeod 2017-09-22 20:00:28 UTC
Run the package build stages that bitbake creates and manages at a lower priority (nice, ionice, etc?) than the bitbake UI and server so that we don't need very long timeouts. 

For systems that are just building Yocto projects, separating the control system (bitbake) from the work (package builds, image construction, etc.) using scheduler hints (niceness) and perhaps io scheduler hints (ionice) should be considered. If we made such a change, then builds would likely be slower if other non-bitbake activities were also running on the system but in my experience, that's not the common case. By using nice and ionice, we avoid needing root permissions. This change may increase the 'liveliness' of the build system for interactive users. The only user option should be how 'nice' to be with the build tasks. 

This enhancement was motivated by:
  https://bugzilla.yoctoproject.org/show_bug.cgi?id=12116 
where the bug was worked around by increasing a timeout from 5 seconds to 30.
Comment 1 Richard Purdie 2017-09-25 14:54:31 UTC
Unfortunately I think you misunderstand the issue in 12116. This has nothing to do with kernel scheduling or anything which ionice or nice will help with.

The issue is more likely that there are event handlers run in server context with disappear for non-trivial lengths of time causing the server's main loop to temporarily stall.

When that stalls, if a UI can't talk to the server for X amount of time, it assumes the server has died.

5s was too short for this, 30s is reasonable. The timeout was short due to other UI interaction issues but those were resolved separately so this can be bumped up.

The only thing we could really do here is to try and shorted or limit the amount of work event handlers are allowed to distract the server main loop from.

We could also separate out the event handler execution but since it needs global data store context that is hard to do, problematic and not something I really want to rush into doing (a Low/Future).
Comment 2 Randy MacLeod 2017-09-25 19:06:20 UTC
I did misunderstand the issue in 12116; thanks for explaining what was really going wrong.

I'd like to leave this enhancement open as a low/future as you suggested.

I still find it odd that the build tasks are run at the same priority as the control tasks. In 2.5, I may gather some data on a heavily loaded system to see if making the build tasks nicer improves the UI/server response time.
Comment 3 Randy MacLeod 2018-12-06 17:16:09 UTC
This is the defect that discusses the bitbake control vs work priorities.
To reduce the frequency of being unable to connect to bitbake server, we are now considering increasing the timeout from 30 secondsd to 90 seconds with a warning being issued at 5 seconds. While that will likely prevent the problems from happening on heavily loaded servers, it seems to me that it does not address the root cause of the problem.
Comment 4 Randy MacLeod 2018-12-06 17:17:53 UTC
related to  https://bugzilla.yoctoproject.org/show_bug.cgi?id=13044
Comment 5 Randy MacLeod 2020-12-03 16:18:44 UTC
There are other bugs to cover making the bitbake server handle io asynchronously.