We have multiple x86_64 based boards currently running kirkstone/4.0. Three of these boards run from a MMC (/dev/mmcblk*). We monitor disk heath using "mmc extcsd read </dev/mmcblk*> | grep EXT_CSD_PRE_EOL_INFO". Everything working fine so far. Now we are about to upgrade to scarthgap/5.0. Here, we are facing a serious problem: The systems runs well *until* the mmc extcsd read command is issued. Then often (not ever) a kernel Oops occurs rendering the mmc device offline. This happens on all three of our MMC based x86_64 (which apart from all have an MMC as main disk differ in various ways: CPU, chipset, ...). Setup/Environment: - hardware: x86_64 based board with MMC as main disk - plain Yocto "hello world" setup with - poky/meta - poky/meta-poky - poky/meta-yocto-bsp - meta-openembedded/meta-oe - MACHINE ?= "genericx86-64" - DISTRO ?= "poky" - WKS_FILE = "directdisk-gpt.wks" - EXTRA_IMAGE_FEATURES += "debug-tweaks" - IMAGE_FEATURES += "allow-empty-password empty-root-password ssh-server-dropbear" - IMAGE_INSTALL:append = " util-linux mmc-utils" Notes: - our systems boot legacy, hence directdisk-gpt.wks - util-linux provides better dmesg - mmc-utils needed for provoking the problem Reproducing the problem: - bitbake core-image-minimal - bmaptool copy core-image-minimal-genericx86-64.rootfs.wic /dev/mmcblk1 - boot it - ssh into it (two parallel terminals) - "dmesg -w" in first terminal - "mmc extcsd read /dev/mmcblk1" in second terminal In case the mmc command doesn't lead to a kernel Oops instantly, simply place it into an endless loop with sleep 1. In all three of our MMC-based x86_64 boards this error occurs: [ 226.274624] BUG: unable to handle page fault for address: ffffffffb53cc621 [ 226.274940] #PF: supervisor write access in kernel mode [ 226.275158] #PF: error_code(0x0003) - permissions violation [ 226.275393] PGD 56a2d067 P4D 56a2d067 PUD 56a2e063 PMD 55c001a1 [ 226.275656] Oops: 0003 [#1] PREEMPT SMP NOPTI [ 226.275837] CPU: 2 PID: 400 Comm: mmc Not tainted 6.6.21-yocto-standard #1 [ 226.276142] Hardware name: Seco 0B03/0B03, BIOS 1.08 12/06/2019 [ 226.276394] RIP: 0010:__mmc_blk_ioctl_cmd+0x207/0x760 [ 226.276615] Code: 8b 3c 24 4c 89 d6 e8 a8 8d fe ff 48 8b 85 68 ff ff ff 4d 85 ff 48 89 43 10 48 8b 85 70 ff ff ff 48 89 43 18 74 1b 48 8b 45 a0 <49> 89 47 10 48 8b 45 a8 49 89 47 18 8b 4d b8 85 c9 0f 85 94 04 00 [ 226.277536] RSP: 0018:ffffae1ec0a1f9e8 EFLAGS: 00010286 [ 226.277754] RAX: 0000000000000000 RBX: ffff9f758160be00 RCX: 0000000000000002 [ 226.278074] RDX: 0000000000000000 RSI: 0000000000000000 RDI: ffff9f7580d99000 [ 226.278387] RBP: ffffae1ec0a1fb98 R08: 0000000000000400 R09: 0000000000000001 [ 226.278708] R10: 0000000000000000 R11: 0000000000000000 R12: ffff9f7580da3000 [ 226.279024] R13: ffffae1ec0a1faf8 R14: 0000000000002710 R15: ffffffffb53cc611 [ 226.279344] FS: 00007f2c6f082740(0000) GS:ffff9f75fbd00000(0000) knlGS:0000000000000000 [ 226.279706] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 [ 226.279948] CR2: ffffffffb53cc621 CR3: 00000001022cc000 CR4: 00000000003506e0 [ 226.280263] Call Trace: [ 226.280336] <TASK> [ 226.280400] ? show_regs+0x69/0x80 [ 226.280525] ? __die+0x28/0x70 [ 226.280630] ? mmc_blk_ioctl_cmd+0xe1/0x150 [ 226.280794] ? page_fault_oops+0x174/0x4b0 [ 226.280954] ? mmc_blk_ioctl_cmd+0xe1/0x150 [ 226.281129] ? search_bpf_extables+0x64/0x90 [ 226.281298] ? __mmc_blk_ioctl_cmd+0x207/0x760 [ 226.281475] ? mmc_blk_ioctl_cmd+0xe1/0x150 [ 226.281638] ? kernelmode_fixup_or_oops+0xa2/0x120 [ 226.281838] ? mmc_blk_ioctl_cmd+0xe1/0x150 [ 226.282000] ? __bad_area_nosemaphore+0x176/0x230 [ 226.282200] ? mmc_blk_ioctl_cmd+0xe1/0x150 [ 226.282371] ? mmc_blk_ioctl_cmd+0xe1/0x150 [ 226.282533] ? bad_area_nosemaphore+0x16/0x20 [ 226.282706] ? do_kern_addr_fault+0x7a/0x90 [ 226.282869] ? exc_page_fault+0xe4/0x150 [ 226.283031] ? asm_exc_page_fault+0x2b/0x30 [ 226.283194] ? mmc_blk_ioctl_cmd+0xd1/0x150 [ 226.283357] ? __mmc_blk_ioctl_cmd+0x207/0x760 [ 226.283534] ? __mmc_blk_ioctl_cmd+0x1e8/0x760 [ 226.283718] ? byt_runtime_resume+0x3e/0x60 [ 226.283886] ? __pfx_mmc_wait_done+0x10/0x10 [ 226.284056] mmc_blk_mq_issue_rq+0x3ca/0xa80 [ 226.284227] mmc_mq_queue_rq+0x14a/0x280 [ 226.284385] blk_mq_dispatch_rq_list+0x1bf/0x780 [ 226.284573] ? sbitmap_find_bit+0x8c/0x150 [ 226.284737] __blk_mq_sched_dispatch_requests+0xb7/0x5f0 [ 226.284970] ? blk_mq_get_tag+0x24e/0x2a0 [ 226.285124] ? __pfx_read_tsc+0x10/0x10 [ 226.285271] ? ktime_get+0x44/0xa0 [ 226.285397] blk_mq_sched_dispatch_requests+0x3b/0x70 [ 226.285613] blk_mq_run_hw_queue+0x103/0x1f0 [ 226.285784] blk_execute_rq+0x112/0x200 [ 226.285929] mmc_blk_ioctl_cmd+0xd1/0x150 [ 226.286098] mmc_blk_ioctl+0x6c/0x110 [ 226.286241] blkdev_ioctl+0xfc/0x280 [ 226.286375] __x64_sys_ioctl+0xa1/0xe0 [ 226.286515] do_syscall_64+0x47/0x90 [ 226.286650] entry_SYSCALL_64_after_hwframe+0x6e/0xd8 [ 226.286860] RIP: 0033:0x7f2c6f17f8fc [ 226.286993] Code: 1e fa 48 8d 44 24 08 48 89 54 24 e0 48 89 44 24 c0 48 8d 44 24 d0 48 89 44 24 c8 b8 10 00 00 00 c7 44 24 b8 10 00 00 00 0f 05 <3d> 00 f0 ff ff 89 c2 77 0b 89 d0 c3 0f 1f 84 00 00 00 00 00 48 8b [ 226.287912] RSP: 002b:00007ffd06c84508 EFLAGS: 00000246 ORIG_RAX: 0000000000000010 [ 226.288254] RAX: ffffffffffffffda RBX: 00007ffd06c84580 RCX: 00007f2c6f17f8fc [ 226.288575] RDX: 00007ffd06c84520 RSI: 00000000c048b300 RDI: 0000000000000003 [ 226.288887] RBP: 00007ffd06c84e84 R08: 0000000000000003 R09: 0000000000000001 [ 226.289199] R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000003 [ 226.289512] R13: 000055a18eac609b R14: 00007ffd06c84e74 R15: 000055a18ead2ed0 [ 226.289826] </TASK> [ 226.289889] Modules linked in: [ 226.289996] CR2: ffffffffb53cc621 [ 226.290116] ---[ end trace 0000000000000000 ]--- [ 226.290301] RIP: 0010:__mmc_blk_ioctl_cmd+0x207/0x760 [ 226.290509] Code: 8b 3c 24 4c 89 d6 e8 a8 8d fe ff 48 8b 85 68 ff ff ff 4d 85 ff 48 89 43 10 48 8b 85 70 ff ff ff 48 89 43 18 74 1b 48 8b 45 a0 <49> 89 47 10 48 8b 45 a8 49 89 47 18 8b 4d b8 85 c9 0f 85 94 04 00 [ 226.291433] RSP: 0018:ffffae1ec0a1f9e8 EFLAGS: 00010286 [ 226.291649] RAX: 0000000000000000 RBX: ffff9f758160be00 RCX: 0000000000000002 [ 226.291969] RDX: 0000000000000000 RSI: 0000000000000000 RDI: ffff9f7580d99000 [ 226.292281] RBP: ffffae1ec0a1fb98 R08: 0000000000000400 R09: 0000000000000001 [ 226.292603] R10: 0000000000000000 R11: 0000000000000000 R12: ffff9f7580da3000 [ 226.292915] R13: ffffae1ec0a1faf8 R14: 0000000000002710 R15: ffffffffb53cc611 [ 226.293235] FS: 00007f2c6f082740(0000) GS:ffff9f75fbd00000(0000) knlGS:0000000000000000 [ 226.293596] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 [ 226.293846] CR2: ffffffffb53cc621 CR3: 00000001022cc000 CR4: 00000000003506e0 [ 226.294159] note: mmc[400] exited with irqs disabled [ 286.361121] udevd[135]: worker [363] /devices/pci0000:00/0000:00:1c.0/mmc_host/mmc1/mmc1:0001/block/mmcblk1/mmcblk1p1 is taking a long time [ 406.481242] udevd[135]: worker [363] /devices/pci0000:00/0000:00:1c.0/mmc_host/mmc1/mmc1:0001/block/mmcblk1/mmcblk1p1 timeout; kill it [ 406.481855] udevd[135]: seq 1845 '/devices/pci0000:00/0000:00:1c.0/mmc_host/mmc1/mmc1:0001/block/mmcblk1/mmcblk1p1' killed Note: We were *not* able to reproduce the problem on the same hardware running Debian 12.5.0 with kernel 6.6.13-1~bpo12+1 from bookworm-backports.
The problem was also reported here https://bugzilla.kernel.org/show_bug.cgi?id=218674 a while ago but keeps lingering there untouched, maybe because they don't consider linux-yocto to be a vanilla kernel.
This doesn't sound like something I can reproduce on qemu, and I don't have any appropriate x86 hardware available, so hopefully I can get you to try a few things. It would be good to rule out one of our feature/embedded tweaks as the cause of the issues. Would you be willing to build the kernel with the KBRANCH forced to v6.6/base ? The debian reference is useful, and knowing the above would be even more useful to narrow down what might have gone wrong.
rebuilt core-image-minimal as given above with these additions: KBRANCH:genericx86-64 = "v6.6/base" SRCREV_machine:genericx86-64 ?= "636203cddfde4e5fc3d092170879aa4aaebaea8e" The SRCREV refers to tag 6.6.21 (not 6.6.28 which is top of this branch). Good news: No more Oopses. None. Even after 45min continuously looping over "mmc extcsd read". However, there's a flaw: "mmc extcsd read" only returns values in (roughly) one out of ten times. Unsuccessful read of extcsd data: # mmc extcsd read /dev/mmcblk1 ============================================= Extended CSD rev 1.0 (MMC 4.0) ============================================= (and nothing more) Successful read of extcsd data: # mmc extcsd read /dev/mmcblk1 ============================================= Extended CSD rev 1.8 (MMC 5.1) ============================================= Card Supported Command sets [S_CMD_SET: 0x01] HPI Features [HPI_FEATURE: 0x01]: implementation based on CMD13 Background operations support [BKOPS_SUPPORT: 0x01] Max Packet Read Cmd [MAX_PACKED_READS: 0x00] Max Packet Write Cmd [MAX_PACKED_WRITES: 0x00] Data TAG support [DATA_TAG_SUPPORT: 0x01] Data TAG Unit Size [TAG_UNIT_SIZE: 0x03] Tag Resources Size [TAG_RES_SIZE: 0x00] [..] (and lots of other data)
Additional information: When searching my repository for KBRANCH I noticed that linux-intel from meta-intel layer uses KBRANCH = "6.6/linux" referencing a SRCREV pointing to 6.6.23. So I need to add: We first observed the bug when migrating our actual image to scarthgap which actually doesn't use linux-yocto but linux-intel instead: MACHINE = "intel-corei7-64" PREFERRED_PROVIDER_virtual/kernel = "linux-intel" So the bug affects linux-yocto as well as linux-intel (but *not* linux-yocto with KBRANCH = "v6.6/base" as we learned from my earlier comment). For the sake of my bug report I tried to give a minimal working example which I found in linux-yocto.
Add Anuj who may want to know about or even help resolve the bug.
Finally getting back to this. Thanks for confirming that v6.6/base has the different behaviour. I don't see anything overly suspicious in v6.6/standard/base, but there are definitely some filesystems (yaffs2, aufs) changes that we carry that might be causing the issue. Since I can't reproduce the problem on qemu, the only viable path that I see is via a bisect, or reverting some of the filesystem additions and re-resting on your setup. Would either method be possible ? All that being said, if linux-intel is also showing the problem, that points away from my suspicions about the fileystems as they shouldn't be present in that repository.
Yes, I will gladly build images with whatever kernel configuration or branch you want me to test. I have all three affected boards here with me so on workdays building and testing an image is a no-brainer. Weekends will work too, though there might be some delays before getting back with results. Regarding meta-intel, should I file a report there as well or was adding Anuj sufficient?
Good news: With 6.6.25 (current kernel in meta-intel) everything works fine. MMC block device as well as reading mmc extcsd. I can see in 6.6.24 there were some patches affecting mmc (https://cdn.kernel.org/pub/linux/kernel/v6.x/ChangeLog-6.6.24) though I can't really tell if it is related to the bug we were affected by. However, I can cautiously give the all-clear so far.
Let's call this "fixed" for now, and revisit if the oops re-appears.