Bug 15474 - [kernel] Oops (page fault) in mmc driver on multiple hardware when using mmc-utils
Summary: [kernel] Oops (page fault) in mmc driver on multiple hardware when using mmc-...
Status: RESOLVED FIXED
Alias: None
Product: OE-Core
Classification: Build System, Metadata & Runtime
Component: kernel (show other bugs)
Version: 5.0
Hardware: x86 x86_64
: Medium+ major
Target Milestone: 5.1 M3
Assignee: Bruce Ashfield
QA Contact:
URL:
Whiteboard: backport 5.0
Depends on:
Blocks:
 
Reported: 2024-04-26 08:41 UTC by Andreas Ufert
Modified: 2024-09-10 13:44 UTC (History)
2 users (show)

See Also:
OS type for building Yocto: ---
Type of Regression: ---
Verified:
Documentation change: No (bug/feature does not impact docs)


Attachments

Note You need to log in before you can comment on or make changes to this bug.
Description Andreas Ufert 2024-04-26 08:41:03 UTC
We have multiple x86_64 based boards currently running kirkstone/4.0. Three of these boards run from a MMC (/dev/mmcblk*). We monitor disk heath using "mmc extcsd read </dev/mmcblk*> | grep EXT_CSD_PRE_EOL_INFO". Everything working fine so far.

Now we are about to upgrade to scarthgap/5.0. Here, we are facing a serious problem: The systems runs well *until* the mmc extcsd read command is issued. Then often (not ever) a kernel Oops occurs rendering the mmc device offline. This happens on all three of our MMC based x86_64 (which apart from all have an MMC as main disk differ in various ways: CPU, chipset, ...).


Setup/Environment:

- hardware: x86_64 based board with MMC as main disk
- plain Yocto "hello world" setup with
  - poky/meta
  - poky/meta-poky
  - poky/meta-yocto-bsp
  - meta-openembedded/meta-oe
  - MACHINE ?= "genericx86-64"
  - DISTRO ?= "poky"
  - WKS_FILE = "directdisk-gpt.wks"
  - EXTRA_IMAGE_FEATURES += "debug-tweaks"
  - IMAGE_FEATURES += "allow-empty-password empty-root-password ssh-server-dropbear"
  - IMAGE_INSTALL:append = " util-linux mmc-utils"


Notes:

  - our systems boot legacy, hence directdisk-gpt.wks
  - util-linux provides better dmesg
  - mmc-utils needed for provoking the problem


Reproducing the problem:

  - bitbake core-image-minimal
  - bmaptool copy core-image-minimal-genericx86-64.rootfs.wic /dev/mmcblk1
  - boot it
  - ssh into it (two parallel terminals)
  - "dmesg -w" in first terminal
  - "mmc extcsd read /dev/mmcblk1" in second terminal


In case the mmc command doesn't lead to a kernel Oops instantly, simply place it into an endless loop with sleep 1.


In all three of our MMC-based x86_64 boards this error occurs:

[  226.274624] BUG: unable to handle page fault for address: ffffffffb53cc621
[  226.274940] #PF: supervisor write access in kernel mode
[  226.275158] #PF: error_code(0x0003) - permissions violation
[  226.275393] PGD 56a2d067 P4D 56a2d067 PUD 56a2e063 PMD 55c001a1
[  226.275656] Oops: 0003 [#1] PREEMPT SMP NOPTI
[  226.275837] CPU: 2 PID: 400 Comm: mmc Not tainted 6.6.21-yocto-standard #1
[  226.276142] Hardware name: Seco 0B03/0B03, BIOS 1.08 12/06/2019
[  226.276394] RIP: 0010:__mmc_blk_ioctl_cmd+0x207/0x760
[  226.276615] Code: 8b 3c 24 4c 89 d6 e8 a8 8d fe ff 48 8b 85 68 ff ff ff 4d 85 ff 48 89 43 10 48 8b 85 70 ff ff ff 48 89 43 18 74 1b 48 8b 45 a0 <49> 89 47 10 48 8b 45 a8 49 89 47 18 8b 4d b8 85 c9 0f 85 94 04 00
[  226.277536] RSP: 0018:ffffae1ec0a1f9e8 EFLAGS: 00010286
[  226.277754] RAX: 0000000000000000 RBX: ffff9f758160be00 RCX: 0000000000000002
[  226.278074] RDX: 0000000000000000 RSI: 0000000000000000 RDI: ffff9f7580d99000
[  226.278387] RBP: ffffae1ec0a1fb98 R08: 0000000000000400 R09: 0000000000000001
[  226.278708] R10: 0000000000000000 R11: 0000000000000000 R12: ffff9f7580da3000
[  226.279024] R13: ffffae1ec0a1faf8 R14: 0000000000002710 R15: ffffffffb53cc611
[  226.279344] FS:  00007f2c6f082740(0000) GS:ffff9f75fbd00000(0000) knlGS:0000000000000000
[  226.279706] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[  226.279948] CR2: ffffffffb53cc621 CR3: 00000001022cc000 CR4: 00000000003506e0
[  226.280263] Call Trace:
[  226.280336]  <TASK>
[  226.280400]  ? show_regs+0x69/0x80
[  226.280525]  ? __die+0x28/0x70
[  226.280630]  ? mmc_blk_ioctl_cmd+0xe1/0x150
[  226.280794]  ? page_fault_oops+0x174/0x4b0
[  226.280954]  ? mmc_blk_ioctl_cmd+0xe1/0x150
[  226.281129]  ? search_bpf_extables+0x64/0x90
[  226.281298]  ? __mmc_blk_ioctl_cmd+0x207/0x760
[  226.281475]  ? mmc_blk_ioctl_cmd+0xe1/0x150
[  226.281638]  ? kernelmode_fixup_or_oops+0xa2/0x120
[  226.281838]  ? mmc_blk_ioctl_cmd+0xe1/0x150
[  226.282000]  ? __bad_area_nosemaphore+0x176/0x230
[  226.282200]  ? mmc_blk_ioctl_cmd+0xe1/0x150
[  226.282371]  ? mmc_blk_ioctl_cmd+0xe1/0x150
[  226.282533]  ? bad_area_nosemaphore+0x16/0x20
[  226.282706]  ? do_kern_addr_fault+0x7a/0x90
[  226.282869]  ? exc_page_fault+0xe4/0x150
[  226.283031]  ? asm_exc_page_fault+0x2b/0x30
[  226.283194]  ? mmc_blk_ioctl_cmd+0xd1/0x150
[  226.283357]  ? __mmc_blk_ioctl_cmd+0x207/0x760
[  226.283534]  ? __mmc_blk_ioctl_cmd+0x1e8/0x760
[  226.283718]  ? byt_runtime_resume+0x3e/0x60
[  226.283886]  ? __pfx_mmc_wait_done+0x10/0x10
[  226.284056]  mmc_blk_mq_issue_rq+0x3ca/0xa80
[  226.284227]  mmc_mq_queue_rq+0x14a/0x280
[  226.284385]  blk_mq_dispatch_rq_list+0x1bf/0x780
[  226.284573]  ? sbitmap_find_bit+0x8c/0x150
[  226.284737]  __blk_mq_sched_dispatch_requests+0xb7/0x5f0
[  226.284970]  ? blk_mq_get_tag+0x24e/0x2a0
[  226.285124]  ? __pfx_read_tsc+0x10/0x10
[  226.285271]  ? ktime_get+0x44/0xa0
[  226.285397]  blk_mq_sched_dispatch_requests+0x3b/0x70
[  226.285613]  blk_mq_run_hw_queue+0x103/0x1f0
[  226.285784]  blk_execute_rq+0x112/0x200
[  226.285929]  mmc_blk_ioctl_cmd+0xd1/0x150
[  226.286098]  mmc_blk_ioctl+0x6c/0x110
[  226.286241]  blkdev_ioctl+0xfc/0x280
[  226.286375]  __x64_sys_ioctl+0xa1/0xe0
[  226.286515]  do_syscall_64+0x47/0x90
[  226.286650]  entry_SYSCALL_64_after_hwframe+0x6e/0xd8
[  226.286860] RIP: 0033:0x7f2c6f17f8fc
[  226.286993] Code: 1e fa 48 8d 44 24 08 48 89 54 24 e0 48 89 44 24 c0 48 8d 44 24 d0 48 89 44 24 c8 b8 10 00 00 00 c7 44 24 b8 10 00 00 00 0f 05 <3d> 00 f0 ff ff 89 c2 77 0b 89 d0 c3 0f 1f 84 00 00 00 00 00 48 8b
[  226.287912] RSP: 002b:00007ffd06c84508 EFLAGS: 00000246 ORIG_RAX: 0000000000000010
[  226.288254] RAX: ffffffffffffffda RBX: 00007ffd06c84580 RCX: 00007f2c6f17f8fc
[  226.288575] RDX: 00007ffd06c84520 RSI: 00000000c048b300 RDI: 0000000000000003
[  226.288887] RBP: 00007ffd06c84e84 R08: 0000000000000003 R09: 0000000000000001
[  226.289199] R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000003
[  226.289512] R13: 000055a18eac609b R14: 00007ffd06c84e74 R15: 000055a18ead2ed0
[  226.289826]  </TASK>
[  226.289889] Modules linked in:
[  226.289996] CR2: ffffffffb53cc621
[  226.290116] ---[ end trace 0000000000000000 ]---
[  226.290301] RIP: 0010:__mmc_blk_ioctl_cmd+0x207/0x760
[  226.290509] Code: 8b 3c 24 4c 89 d6 e8 a8 8d fe ff 48 8b 85 68 ff ff ff 4d 85 ff 48 89 43 10 48 8b 85 70 ff ff ff 48 89 43 18 74 1b 48 8b 45 a0 <49> 89 47 10 48 8b 45 a8 49 89 47 18 8b 4d b8 85 c9 0f 85 94 04 00
[  226.291433] RSP: 0018:ffffae1ec0a1f9e8 EFLAGS: 00010286
[  226.291649] RAX: 0000000000000000 RBX: ffff9f758160be00 RCX: 0000000000000002
[  226.291969] RDX: 0000000000000000 RSI: 0000000000000000 RDI: ffff9f7580d99000
[  226.292281] RBP: ffffae1ec0a1fb98 R08: 0000000000000400 R09: 0000000000000001
[  226.292603] R10: 0000000000000000 R11: 0000000000000000 R12: ffff9f7580da3000
[  226.292915] R13: ffffae1ec0a1faf8 R14: 0000000000002710 R15: ffffffffb53cc611
[  226.293235] FS:  00007f2c6f082740(0000) GS:ffff9f75fbd00000(0000) knlGS:0000000000000000
[  226.293596] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[  226.293846] CR2: ffffffffb53cc621 CR3: 00000001022cc000 CR4: 00000000003506e0
[  226.294159] note: mmc[400] exited with irqs disabled
[  286.361121] udevd[135]: worker [363] /devices/pci0000:00/0000:00:1c.0/mmc_host/mmc1/mmc1:0001/block/mmcblk1/mmcblk1p1 is taking a long time
[  406.481242] udevd[135]: worker [363] /devices/pci0000:00/0000:00:1c.0/mmc_host/mmc1/mmc1:0001/block/mmcblk1/mmcblk1p1 timeout; kill it
[  406.481855] udevd[135]: seq 1845 '/devices/pci0000:00/0000:00:1c.0/mmc_host/mmc1/mmc1:0001/block/mmcblk1/mmcblk1p1' killed


Note: We were *not* able to reproduce the problem on the same hardware running Debian 12.5.0 with kernel 6.6.13-1~bpo12+1 from bookworm-backports.
Comment 1 Andreas Ufert 2024-04-26 08:54:55 UTC
The problem was also reported here https://bugzilla.kernel.org/show_bug.cgi?id=218674 a while ago but keeps lingering there untouched, maybe because they don't consider linux-yocto to be a vanilla kernel.
Comment 2 Bruce Ashfield 2024-04-26 12:49:37 UTC
This doesn't sound like something I can reproduce on qemu, and I don't have any appropriate x86 hardware available, so hopefully I can get you to try a few things.

It would be good to rule out one of our feature/embedded tweaks as the cause of the issues.

Would you be willing to build the kernel with the KBRANCH forced to v6.6/base ?

The debian reference is useful, and knowing the above would be even more useful to narrow down what might have gone wrong.
Comment 3 Andreas Ufert 2024-04-26 17:03:11 UTC
rebuilt core-image-minimal as given above with these additions:

KBRANCH:genericx86-64  = "v6.6/base"
SRCREV_machine:genericx86-64 ?= "636203cddfde4e5fc3d092170879aa4aaebaea8e"

The SRCREV refers to tag 6.6.21 (not 6.6.28 which is top of this branch).

Good news: No more Oopses. None. Even after 45min continuously looping over "mmc extcsd read".

However, there's a flaw: "mmc extcsd read" only returns values in (roughly) one out of ten times.

Unsuccessful read of extcsd data:

# mmc extcsd read /dev/mmcblk1
=============================================
  Extended CSD rev 1.0 (MMC 4.0)
=============================================

(and nothing more)


Successful read of extcsd data:

# mmc extcsd read /dev/mmcblk1
=============================================
  Extended CSD rev 1.8 (MMC 5.1)
=============================================

Card Supported Command sets [S_CMD_SET: 0x01]
HPI Features [HPI_FEATURE: 0x01]: implementation based on CMD13
Background operations support [BKOPS_SUPPORT: 0x01]
Max Packet Read Cmd [MAX_PACKED_READS: 0x00]
Max Packet Write Cmd [MAX_PACKED_WRITES: 0x00]
Data TAG support [DATA_TAG_SUPPORT: 0x01]
Data TAG Unit Size [TAG_UNIT_SIZE: 0x03]
Tag Resources Size [TAG_RES_SIZE: 0x00]
[..]

(and lots of other data)
Comment 4 Andreas Ufert 2024-04-26 17:12:09 UTC
Additional information:

When searching my repository for KBRANCH I noticed that linux-intel from meta-intel layer uses KBRANCH = "6.6/linux" referencing a SRCREV pointing to 6.6.23.

So I need to add: We first observed the bug when migrating our actual image to scarthgap which actually doesn't use linux-yocto but linux-intel instead:

MACHINE = "intel-corei7-64"
PREFERRED_PROVIDER_virtual/kernel = "linux-intel"

So the bug affects linux-yocto as well as linux-intel (but *not* linux-yocto with KBRANCH = "v6.6/base" as we learned from my earlier comment).

For the sake of my bug report I tried to give a minimal working example which I found in linux-yocto.
Comment 5 Randy MacLeod 2024-05-02 14:35:16 UTC
Add Anuj who may want to know about or even help resolve the bug.
Comment 6 Bruce Ashfield 2024-05-02 15:21:49 UTC
Finally getting back to this.

Thanks for confirming that v6.6/base has the different behaviour.

I don't see anything overly suspicious in v6.6/standard/base, but there are definitely some filesystems (yaffs2, aufs) changes that we carry that might be causing the issue.

Since I can't reproduce the problem on qemu, the only viable path that I see is via a bisect, or reverting some of the filesystem additions and re-resting on your setup.

Would either method be possible ?

All that being said, if linux-intel is also showing the problem, that points away from my suspicions about the fileystems as they shouldn't be present in that repository.
Comment 7 Andreas Ufert 2024-05-02 17:47:07 UTC
Yes, I will gladly build images with whatever kernel configuration or branch you want me to test.

I have all three affected boards here with me so on workdays building and testing an image is a no-brainer. Weekends will work too, though there might be some delays before getting back with results.

Regarding meta-intel, should I file a report there as well or was adding Anuj sufficient?
Comment 8 Andreas Ufert 2024-05-08 12:09:42 UTC
Good news: With 6.6.25 (current kernel in meta-intel) everything works fine. MMC block device as well as reading mmc extcsd.

I can see in 6.6.24 there were some patches affecting mmc (https://cdn.kernel.org/pub/linux/kernel/v6.x/ChangeLog-6.6.24) though I can't really tell if it is related to the bug we were affected by.

However, I can cautiously give the all-clear so far.
Comment 9 Bruce Ashfield 2024-09-10 13:44:39 UTC
Let's call this "fixed" for now, and revisit if the oops re-appears.