| Summary: | bitbake servers should be more responsive. | ||
|---|---|---|---|
| Product: | [Build System, Metadata & Runtime] BitBake | Reporter: | Randy MacLeod <randy.macleod> |
| Component: | bitbake | Assignee: | Randy MacLeod <randy.macleod> |
| Status: | RESOLVED OBSOLETE | QA Contact: | |
| Severity: | enhancement | ||
| Priority: | Medium | CC: | poky.bs.watcher, poky.watcher, richard.purdie |
| Version: | 2.5 | ||
| Target Milestone: | 4.99 | ||
| Hardware: | x86 | ||
| OS: | Multiple | ||
| Whiteboard: | |||
| OS type for building Yocto: | --- | Type of Regression: | --- |
| Verified: | Documentation change: | Yes (doc changes required) | |
|
Description
Randy MacLeod
2017-09-22 20:00:28 UTC
Unfortunately I think you misunderstand the issue in 12116. This has nothing to do with kernel scheduling or anything which ionice or nice will help with. The issue is more likely that there are event handlers run in server context with disappear for non-trivial lengths of time causing the server's main loop to temporarily stall. When that stalls, if a UI can't talk to the server for X amount of time, it assumes the server has died. 5s was too short for this, 30s is reasonable. The timeout was short due to other UI interaction issues but those were resolved separately so this can be bumped up. The only thing we could really do here is to try and shorted or limit the amount of work event handlers are allowed to distract the server main loop from. We could also separate out the event handler execution but since it needs global data store context that is hard to do, problematic and not something I really want to rush into doing (a Low/Future). I did misunderstand the issue in 12116; thanks for explaining what was really going wrong. I'd like to leave this enhancement open as a low/future as you suggested. I still find it odd that the build tasks are run at the same priority as the control tasks. In 2.5, I may gather some data on a heavily loaded system to see if making the build tasks nicer improves the UI/server response time. This is the defect that discusses the bitbake control vs work priorities. To reduce the frequency of being unable to connect to bitbake server, we are now considering increasing the timeout from 30 secondsd to 90 seconds with a warning being issued at 5 seconds. While that will likely prevent the problems from happening on heavily loaded servers, it seems to me that it does not address the root cause of the problem. There are other bugs to cover making the bitbake server handle io asynchronously. |