It feels like an I/O load issue, work on some sort of `flock`/`dd` timeout is currently done to try to get to the bottom of the cause and verify if it is indeed I/O related. 2 solutions have been suggested in the Mar 25 triage IIRC they were: * cgroups control to give qemu an advantage over builds * nice levels to make qemu higher priority (lowering the builds priority)
We were doing this anyway but this bug can be used to track things.
Most autobuilder workers only have cgroups v1 which doesn't give the IO controls we'd need. We already set ionice/nice levels. The ionice level doesn't affect async writes. We've now added support for running qemu images within tmpfs.
Following the triage call, it seems necessary to: - timestamp the start and end of the image copy to tmpfs - double the qemu start timeout
Many changes including: - moving the qemu rootfs to tmpfs - limiting glibc self-test memory usage - adding smp to qemu image for x86, arm and changing the x86 target machine. and likely more have improve ptests. Closing this generic bug. Open specific issues for additional problems.
As per previous comment.