| Summary: | qemuarm: core-image-sato UI hangs with 6.1 kernel | ||||||
|---|---|---|---|---|---|---|---|
| Product: | [Build System, Metadata & Runtime] OE-Core | Reporter: | Randy MacLeod <randy.macleod> | ||||
| Component: | kernel | Assignee: | Ross Burton <ross.burton> | ||||
| Status: | RESOLVED FIXED | QA Contact: | |||||
| Severity: | normal | ||||||
| Priority: | High | CC: | narpat.mali, tom.zanussi | ||||
| Version: | 4.2 | ||||||
| Target Milestone: | 4.3 M1 | ||||||
| Hardware: | x86 | ||||||
| OS: | Multiple | ||||||
| Whiteboard: | |||||||
| OS type for building Yocto: | --- | Type of Regression: | --- | ||||
| Verified: | Documentation change: | No (bug/feature does not impact docs) | |||||
| Attachments: |
|
||||||
|
Description
Randy MacLeod
2023-04-11 22:06:16 UTC
Randy and Narpat to debug userspace to narrow down the issue. Tested with linux-yocto-dev 6.3.0-rc6-yoctodev-standard+ and the hang doesn't happen. This seems like a kernel regression. Using poky-master: ff633ce7a7ffc0d39b86043caeb63a9ffcef99b8 and include below in conf/local.conf PREFERRED_PROVIDER_virtual/kernel = "linux-yocto-dev" Next step is to compare the boot logs between 6.1.20 & 6.3.0 and sato initialization to understand the how the deadlock is occurring. FYI, sato UI freeze also seen in poky/master = f79046d082 (origin/master, origin/HEAD) cve-exclusions: Document some further linux-yocto CVE statuses for qemuriscv64. Using the head of mickledore branch and qemuarm, I don't see this. I do have a slew of tweaks in my local.conf so I'll do another build with a clean conf. OK I can replicate with sysv. Curious! Yes, all my testing was with sysvinit. I should try xfce with sysvinit since the default for WR Linux is systemd.
For completeness, I confirmed that for qemuriscv64, switching to the 6.3 kernel using:
PREFERRED_PROVIDER_virtual/kernel = "linux-yocto-dev"
also avoids the sato UI hang.
I also switched to mickeldore:
4bb775aecb (HEAD -> mickledore, origin/mickledore)
cve-exclusions: Document some further linux-yocto CVE statuses
and back to the standard kernel and the problem exist there fore me as well.
It's odd that you can't reproduce it since both Narpat and I do see the problem and I even saw it when building on my RPi4 running Wind River Linux and qemu as a guest so it's not just a host issue. For builds on intel, I connected using runqemu publicvnc but when testing on the RPi4/WRL it's local graphics so that takes the rendering out of the picture as a factor.
So a potential lead is that when it works the VNC connection is 32bits-per-pixel rgb888, but when it doesn't the connection is 8bpp rgb222. Which doesn't sounds right _at all_. Found the bug in the kernel, and found the fix. Typically, it was broken from 6.1.13 to 6.1.21, and we were trying to release with 6.1.20. https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux.git/commit/?h=linux-6.1.y&id=1ea3e18e53f2e741b264960a7f829f44110d72d2 is the fix. I suggest we pick that into linux-yocto and let it rebase out when we upgrade past 6.1.20. How did you find the bug/fix? Verified that 6.1.14 broke but 6.1.12 worked. Bisected that further to 6.1.13 is broken too. That reduced it down to a manageable number of commits which I bisected. Thanks to good commit messages in the kernel it was trivial to find the commit in master which fixed the buggy commit. Thanks. That's disappointingly banal! ;-) Fixed with the upgrade to kernel 6.1.25 in f3a422f02113cda2aa7872e350e468bdb447c8c3. |