| Summary: | [kernel] Oops (page fault) in mmc driver on multiple hardware when using mmc-utils | ||
|---|---|---|---|
| Product: | [Build System, Metadata & Runtime] OE-Core | Reporter: | Andreas Ufert <Andreas.Ufert> |
| Component: | kernel | Assignee: | Bruce Ashfield <bruce.ashfield> |
| Status: | RESOLVED FIXED | QA Contact: | |
| Severity: | major | ||
| Priority: | Medium+ | CC: | anuj.mittal, randy.macleod |
| Version: | 5.0 | ||
| Target Milestone: | 5.1 M3 | ||
| Hardware: | x86 | ||
| OS: | x86_64 | ||
| Whiteboard: | backport 5.0 | ||
| OS type for building Yocto: | --- | Type of Regression: | --- |
| Verified: | Documentation change: | No (bug/feature does not impact docs) | |
|
Description
Andreas Ufert
2024-04-26 08:41:03 UTC
The problem was also reported here https://bugzilla.kernel.org/show_bug.cgi?id=218674 a while ago but keeps lingering there untouched, maybe because they don't consider linux-yocto to be a vanilla kernel. This doesn't sound like something I can reproduce on qemu, and I don't have any appropriate x86 hardware available, so hopefully I can get you to try a few things. It would be good to rule out one of our feature/embedded tweaks as the cause of the issues. Would you be willing to build the kernel with the KBRANCH forced to v6.6/base ? The debian reference is useful, and knowing the above would be even more useful to narrow down what might have gone wrong. rebuilt core-image-minimal as given above with these additions: KBRANCH:genericx86-64 = "v6.6/base" SRCREV_machine:genericx86-64 ?= "636203cddfde4e5fc3d092170879aa4aaebaea8e" The SRCREV refers to tag 6.6.21 (not 6.6.28 which is top of this branch). Good news: No more Oopses. None. Even after 45min continuously looping over "mmc extcsd read". However, there's a flaw: "mmc extcsd read" only returns values in (roughly) one out of ten times. Unsuccessful read of extcsd data: # mmc extcsd read /dev/mmcblk1 ============================================= Extended CSD rev 1.0 (MMC 4.0) ============================================= (and nothing more) Successful read of extcsd data: # mmc extcsd read /dev/mmcblk1 ============================================= Extended CSD rev 1.8 (MMC 5.1) ============================================= Card Supported Command sets [S_CMD_SET: 0x01] HPI Features [HPI_FEATURE: 0x01]: implementation based on CMD13 Background operations support [BKOPS_SUPPORT: 0x01] Max Packet Read Cmd [MAX_PACKED_READS: 0x00] Max Packet Write Cmd [MAX_PACKED_WRITES: 0x00] Data TAG support [DATA_TAG_SUPPORT: 0x01] Data TAG Unit Size [TAG_UNIT_SIZE: 0x03] Tag Resources Size [TAG_RES_SIZE: 0x00] [..] (and lots of other data) Additional information: When searching my repository for KBRANCH I noticed that linux-intel from meta-intel layer uses KBRANCH = "6.6/linux" referencing a SRCREV pointing to 6.6.23. So I need to add: We first observed the bug when migrating our actual image to scarthgap which actually doesn't use linux-yocto but linux-intel instead: MACHINE = "intel-corei7-64" PREFERRED_PROVIDER_virtual/kernel = "linux-intel" So the bug affects linux-yocto as well as linux-intel (but *not* linux-yocto with KBRANCH = "v6.6/base" as we learned from my earlier comment). For the sake of my bug report I tried to give a minimal working example which I found in linux-yocto. Add Anuj who may want to know about or even help resolve the bug. Finally getting back to this. Thanks for confirming that v6.6/base has the different behaviour. I don't see anything overly suspicious in v6.6/standard/base, but there are definitely some filesystems (yaffs2, aufs) changes that we carry that might be causing the issue. Since I can't reproduce the problem on qemu, the only viable path that I see is via a bisect, or reverting some of the filesystem additions and re-resting on your setup. Would either method be possible ? All that being said, if linux-intel is also showing the problem, that points away from my suspicions about the fileystems as they shouldn't be present in that repository. Yes, I will gladly build images with whatever kernel configuration or branch you want me to test. I have all three affected boards here with me so on workdays building and testing an image is a no-brainer. Weekends will work too, though there might be some delays before getting back with results. Regarding meta-intel, should I file a report there as well or was adding Anuj sufficient? Good news: With 6.6.25 (current kernel in meta-intel) everything works fine. MMC block device as well as reading mmc extcsd. I can see in 6.6.24 there were some patches affecting mmc (https://cdn.kernel.org/pub/linux/kernel/v6.x/ChangeLog-6.6.24) though I can't really tell if it is related to the bug we were affected by. However, I can cautiously give the all-clear so far. Let's call this "fixed" for now, and revisit if the oops re-appears. |