<?xml version="1.0" encoding="UTF-8" standalone="yes" ?>
<!DOCTYPE bugzilla SYSTEM "https://bugzilla.yoctoproject.org/page.cgi?id=bugzilla.dtd">

<bugzilla version="5.0.6"
          urlbase="https://bugzilla.yoctoproject.org/"
          
          maintainer="it-coreprojects-helpdesk@linuxfoundation.org"
>

    <bug>
          <bug_id>15474</bug_id>
          
          <creation_ts>2024-04-26 08:41:03 +0000</creation_ts>
          <short_desc>[kernel] Oops (page fault) in mmc driver on multiple hardware when using mmc-utils</short_desc>
          <delta_ts>2024-09-10 13:44:39 +0000</delta_ts>
          <reporter_accessible>1</reporter_accessible>
          <cclist_accessible>1</cclist_accessible>
          <classification_id>7</classification_id>
          <classification>Build System, Metadata &amp; Runtime</classification>
          <product>OE-Core</product>
          <component>kernel</component>
          <version>5.0</version>
          <rep_platform>x86</rep_platform>
          <op_sys>x86_64</op_sys>
          <bug_status>RESOLVED</bug_status>
          <resolution>FIXED</resolution>
          
          
          <bug_file_loc></bug_file_loc>
          <status_whiteboard>backport 5.0 </status_whiteboard>
          <keywords></keywords>
          <priority>Medium+</priority>
          <bug_severity>major</bug_severity>
          <target_milestone>5.1 M3</target_milestone>
          
          
          <everconfirmed>1</everconfirmed>
          <reporter name="Andreas Ufert">Andreas.Ufert</reporter>
          <assigned_to name="Bruce Ashfield">bruce.ashfield</assigned_to>
          <cc>anuj.mittal</cc>
    
    <cc>randy.macleod</cc>
          
          
          <cf_os>---</cf_os>
          <cf_regression_type>---</cf_regression_type>
          
          <cf_docchange>No (bug/feature does not impact docs)</cf_docchange>

      

      

      

          <comment_sort_order>oldest_to_newest</comment_sort_order>  
          <long_desc isprivate="0" >
    <commentid>98924</commentid>
    <comment_count>0</comment_count>
    <who name="Andreas Ufert">Andreas.Ufert</who>
    <bug_when>2024-04-26 08:41:03 +0000</bug_when>
    <thetext>We have multiple x86_64 based boards currently running kirkstone/4.0. Three of these boards run from a MMC (/dev/mmcblk*). We monitor disk heath using &quot;mmc extcsd read &lt;/dev/mmcblk*&gt; | grep EXT_CSD_PRE_EOL_INFO&quot;. Everything working fine so far.

Now we are about to upgrade to scarthgap/5.0. Here, we are facing a serious problem: The systems runs well *until* the mmc extcsd read command is issued. Then often (not ever) a kernel Oops occurs rendering the mmc device offline. This happens on all three of our MMC based x86_64 (which apart from all have an MMC as main disk differ in various ways: CPU, chipset, ...).


Setup/Environment:

- hardware: x86_64 based board with MMC as main disk
- plain Yocto &quot;hello world&quot; setup with
  - poky/meta
  - poky/meta-poky
  - poky/meta-yocto-bsp
  - meta-openembedded/meta-oe
  - MACHINE ?= &quot;genericx86-64&quot;
  - DISTRO ?= &quot;poky&quot;
  - WKS_FILE = &quot;directdisk-gpt.wks&quot;
  - EXTRA_IMAGE_FEATURES += &quot;debug-tweaks&quot;
  - IMAGE_FEATURES += &quot;allow-empty-password empty-root-password ssh-server-dropbear&quot;
  - IMAGE_INSTALL:append = &quot; util-linux mmc-utils&quot;


Notes:

  - our systems boot legacy, hence directdisk-gpt.wks
  - util-linux provides better dmesg
  - mmc-utils needed for provoking the problem


Reproducing the problem:

  - bitbake core-image-minimal
  - bmaptool copy core-image-minimal-genericx86-64.rootfs.wic /dev/mmcblk1
  - boot it
  - ssh into it (two parallel terminals)
  - &quot;dmesg -w&quot; in first terminal
  - &quot;mmc extcsd read /dev/mmcblk1&quot; in second terminal


In case the mmc command doesn&apos;t lead to a kernel Oops instantly, simply place it into an endless loop with sleep 1.


In all three of our MMC-based x86_64 boards this error occurs:

[  226.274624] BUG: unable to handle page fault for address: ffffffffb53cc621
[  226.274940] #PF: supervisor write access in kernel mode
[  226.275158] #PF: error_code(0x0003) - permissions violation
[  226.275393] PGD 56a2d067 P4D 56a2d067 PUD 56a2e063 PMD 55c001a1
[  226.275656] Oops: 0003 [#1] PREEMPT SMP NOPTI
[  226.275837] CPU: 2 PID: 400 Comm: mmc Not tainted 6.6.21-yocto-standard #1
[  226.276142] Hardware name: Seco 0B03/0B03, BIOS 1.08 12/06/2019
[  226.276394] RIP: 0010:__mmc_blk_ioctl_cmd+0x207/0x760
[  226.276615] Code: 8b 3c 24 4c 89 d6 e8 a8 8d fe ff 48 8b 85 68 ff ff ff 4d 85 ff 48 89 43 10 48 8b 85 70 ff ff ff 48 89 43 18 74 1b 48 8b 45 a0 &lt;49&gt; 89 47 10 48 8b 45 a8 49 89 47 18 8b 4d b8 85 c9 0f 85 94 04 00
[  226.277536] RSP: 0018:ffffae1ec0a1f9e8 EFLAGS: 00010286
[  226.277754] RAX: 0000000000000000 RBX: ffff9f758160be00 RCX: 0000000000000002
[  226.278074] RDX: 0000000000000000 RSI: 0000000000000000 RDI: ffff9f7580d99000
[  226.278387] RBP: ffffae1ec0a1fb98 R08: 0000000000000400 R09: 0000000000000001
[  226.278708] R10: 0000000000000000 R11: 0000000000000000 R12: ffff9f7580da3000
[  226.279024] R13: ffffae1ec0a1faf8 R14: 0000000000002710 R15: ffffffffb53cc611
[  226.279344] FS:  00007f2c6f082740(0000) GS:ffff9f75fbd00000(0000) knlGS:0000000000000000
[  226.279706] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[  226.279948] CR2: ffffffffb53cc621 CR3: 00000001022cc000 CR4: 00000000003506e0
[  226.280263] Call Trace:
[  226.280336]  &lt;TASK&gt;
[  226.280400]  ? show_regs+0x69/0x80
[  226.280525]  ? __die+0x28/0x70
[  226.280630]  ? mmc_blk_ioctl_cmd+0xe1/0x150
[  226.280794]  ? page_fault_oops+0x174/0x4b0
[  226.280954]  ? mmc_blk_ioctl_cmd+0xe1/0x150
[  226.281129]  ? search_bpf_extables+0x64/0x90
[  226.281298]  ? __mmc_blk_ioctl_cmd+0x207/0x760
[  226.281475]  ? mmc_blk_ioctl_cmd+0xe1/0x150
[  226.281638]  ? kernelmode_fixup_or_oops+0xa2/0x120
[  226.281838]  ? mmc_blk_ioctl_cmd+0xe1/0x150
[  226.282000]  ? __bad_area_nosemaphore+0x176/0x230
[  226.282200]  ? mmc_blk_ioctl_cmd+0xe1/0x150
[  226.282371]  ? mmc_blk_ioctl_cmd+0xe1/0x150
[  226.282533]  ? bad_area_nosemaphore+0x16/0x20
[  226.282706]  ? do_kern_addr_fault+0x7a/0x90
[  226.282869]  ? exc_page_fault+0xe4/0x150
[  226.283031]  ? asm_exc_page_fault+0x2b/0x30
[  226.283194]  ? mmc_blk_ioctl_cmd+0xd1/0x150
[  226.283357]  ? __mmc_blk_ioctl_cmd+0x207/0x760
[  226.283534]  ? __mmc_blk_ioctl_cmd+0x1e8/0x760
[  226.283718]  ? byt_runtime_resume+0x3e/0x60
[  226.283886]  ? __pfx_mmc_wait_done+0x10/0x10
[  226.284056]  mmc_blk_mq_issue_rq+0x3ca/0xa80
[  226.284227]  mmc_mq_queue_rq+0x14a/0x280
[  226.284385]  blk_mq_dispatch_rq_list+0x1bf/0x780
[  226.284573]  ? sbitmap_find_bit+0x8c/0x150
[  226.284737]  __blk_mq_sched_dispatch_requests+0xb7/0x5f0
[  226.284970]  ? blk_mq_get_tag+0x24e/0x2a0
[  226.285124]  ? __pfx_read_tsc+0x10/0x10
[  226.285271]  ? ktime_get+0x44/0xa0
[  226.285397]  blk_mq_sched_dispatch_requests+0x3b/0x70
[  226.285613]  blk_mq_run_hw_queue+0x103/0x1f0
[  226.285784]  blk_execute_rq+0x112/0x200
[  226.285929]  mmc_blk_ioctl_cmd+0xd1/0x150
[  226.286098]  mmc_blk_ioctl+0x6c/0x110
[  226.286241]  blkdev_ioctl+0xfc/0x280
[  226.286375]  __x64_sys_ioctl+0xa1/0xe0
[  226.286515]  do_syscall_64+0x47/0x90
[  226.286650]  entry_SYSCALL_64_after_hwframe+0x6e/0xd8
[  226.286860] RIP: 0033:0x7f2c6f17f8fc
[  226.286993] Code: 1e fa 48 8d 44 24 08 48 89 54 24 e0 48 89 44 24 c0 48 8d 44 24 d0 48 89 44 24 c8 b8 10 00 00 00 c7 44 24 b8 10 00 00 00 0f 05 &lt;3d&gt; 00 f0 ff ff 89 c2 77 0b 89 d0 c3 0f 1f 84 00 00 00 00 00 48 8b
[  226.287912] RSP: 002b:00007ffd06c84508 EFLAGS: 00000246 ORIG_RAX: 0000000000000010
[  226.288254] RAX: ffffffffffffffda RBX: 00007ffd06c84580 RCX: 00007f2c6f17f8fc
[  226.288575] RDX: 00007ffd06c84520 RSI: 00000000c048b300 RDI: 0000000000000003
[  226.288887] RBP: 00007ffd06c84e84 R08: 0000000000000003 R09: 0000000000000001
[  226.289199] R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000003
[  226.289512] R13: 000055a18eac609b R14: 00007ffd06c84e74 R15: 000055a18ead2ed0
[  226.289826]  &lt;/TASK&gt;
[  226.289889] Modules linked in:
[  226.289996] CR2: ffffffffb53cc621
[  226.290116] ---[ end trace 0000000000000000 ]---
[  226.290301] RIP: 0010:__mmc_blk_ioctl_cmd+0x207/0x760
[  226.290509] Code: 8b 3c 24 4c 89 d6 e8 a8 8d fe ff 48 8b 85 68 ff ff ff 4d 85 ff 48 89 43 10 48 8b 85 70 ff ff ff 48 89 43 18 74 1b 48 8b 45 a0 &lt;49&gt; 89 47 10 48 8b 45 a8 49 89 47 18 8b 4d b8 85 c9 0f 85 94 04 00
[  226.291433] RSP: 0018:ffffae1ec0a1f9e8 EFLAGS: 00010286
[  226.291649] RAX: 0000000000000000 RBX: ffff9f758160be00 RCX: 0000000000000002
[  226.291969] RDX: 0000000000000000 RSI: 0000000000000000 RDI: ffff9f7580d99000
[  226.292281] RBP: ffffae1ec0a1fb98 R08: 0000000000000400 R09: 0000000000000001
[  226.292603] R10: 0000000000000000 R11: 0000000000000000 R12: ffff9f7580da3000
[  226.292915] R13: ffffae1ec0a1faf8 R14: 0000000000002710 R15: ffffffffb53cc611
[  226.293235] FS:  00007f2c6f082740(0000) GS:ffff9f75fbd00000(0000) knlGS:0000000000000000
[  226.293596] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[  226.293846] CR2: ffffffffb53cc621 CR3: 00000001022cc000 CR4: 00000000003506e0
[  226.294159] note: mmc[400] exited with irqs disabled
[  286.361121] udevd[135]: worker [363] /devices/pci0000:00/0000:00:1c.0/mmc_host/mmc1/mmc1:0001/block/mmcblk1/mmcblk1p1 is taking a long time
[  406.481242] udevd[135]: worker [363] /devices/pci0000:00/0000:00:1c.0/mmc_host/mmc1/mmc1:0001/block/mmcblk1/mmcblk1p1 timeout; kill it
[  406.481855] udevd[135]: seq 1845 &apos;/devices/pci0000:00/0000:00:1c.0/mmc_host/mmc1/mmc1:0001/block/mmcblk1/mmcblk1p1&apos; killed


Note: We were *not* able to reproduce the problem on the same hardware running Debian 12.5.0 with kernel 6.6.13-1~bpo12+1 from bookworm-backports.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>98925</commentid>
    <comment_count>1</comment_count>
    <who name="Andreas Ufert">Andreas.Ufert</who>
    <bug_when>2024-04-26 08:54:55 +0000</bug_when>
    <thetext>The problem was also reported here https://bugzilla.kernel.org/show_bug.cgi?id=218674 a while ago but keeps lingering there untouched, maybe because they don&apos;t consider linux-yocto to be a vanilla kernel.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>98929</commentid>
    <comment_count>2</comment_count>
    <who name="Bruce Ashfield">bruce.ashfield</who>
    <bug_when>2024-04-26 12:49:37 +0000</bug_when>
    <thetext>This doesn&apos;t sound like something I can reproduce on qemu, and I don&apos;t have any appropriate x86 hardware available, so hopefully I can get you to try a few things.

It would be good to rule out one of our feature/embedded tweaks as the cause of the issues.

Would you be willing to build the kernel with the KBRANCH forced to v6.6/base ?

The debian reference is useful, and knowing the above would be even more useful to narrow down what might have gone wrong.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>98933</commentid>
    <comment_count>3</comment_count>
    <who name="Andreas Ufert">Andreas.Ufert</who>
    <bug_when>2024-04-26 17:03:11 +0000</bug_when>
    <thetext>rebuilt core-image-minimal as given above with these additions:

KBRANCH:genericx86-64  = &quot;v6.6/base&quot;
SRCREV_machine:genericx86-64 ?= &quot;636203cddfde4e5fc3d092170879aa4aaebaea8e&quot;

The SRCREV refers to tag 6.6.21 (not 6.6.28 which is top of this branch).

Good news: No more Oopses. None. Even after 45min continuously looping over &quot;mmc extcsd read&quot;.

However, there&apos;s a flaw: &quot;mmc extcsd read&quot; only returns values in (roughly) one out of ten times.

Unsuccessful read of extcsd data:

# mmc extcsd read /dev/mmcblk1
=============================================
  Extended CSD rev 1.0 (MMC 4.0)
=============================================

(and nothing more)


Successful read of extcsd data:

# mmc extcsd read /dev/mmcblk1
=============================================
  Extended CSD rev 1.8 (MMC 5.1)
=============================================

Card Supported Command sets [S_CMD_SET: 0x01]
HPI Features [HPI_FEATURE: 0x01]: implementation based on CMD13
Background operations support [BKOPS_SUPPORT: 0x01]
Max Packet Read Cmd [MAX_PACKED_READS: 0x00]
Max Packet Write Cmd [MAX_PACKED_WRITES: 0x00]
Data TAG support [DATA_TAG_SUPPORT: 0x01]
Data TAG Unit Size [TAG_UNIT_SIZE: 0x03]
Tag Resources Size [TAG_RES_SIZE: 0x00]
[..]

(and lots of other data)</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>98934</commentid>
    <comment_count>4</comment_count>
    <who name="Andreas Ufert">Andreas.Ufert</who>
    <bug_when>2024-04-26 17:12:09 +0000</bug_when>
    <thetext>Additional information:

When searching my repository for KBRANCH I noticed that linux-intel from meta-intel layer uses KBRANCH = &quot;6.6/linux&quot; referencing a SRCREV pointing to 6.6.23.

So I need to add: We first observed the bug when migrating our actual image to scarthgap which actually doesn&apos;t use linux-yocto but linux-intel instead:

MACHINE = &quot;intel-corei7-64&quot;
PREFERRED_PROVIDER_virtual/kernel = &quot;linux-intel&quot;

So the bug affects linux-yocto as well as linux-intel (but *not* linux-yocto with KBRANCH = &quot;v6.6/base&quot; as we learned from my earlier comment).

For the sake of my bug report I tried to give a minimal working example which I found in linux-yocto.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>98956</commentid>
    <comment_count>5</comment_count>
    <who name="Randy MacLeod">randy.macleod</who>
    <bug_when>2024-05-02 14:35:16 +0000</bug_when>
    <thetext>Add Anuj who may want to know about or even help resolve the bug.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>98963</commentid>
    <comment_count>6</comment_count>
    <who name="Bruce Ashfield">bruce.ashfield</who>
    <bug_when>2024-05-02 15:21:49 +0000</bug_when>
    <thetext>Finally getting back to this.

Thanks for confirming that v6.6/base has the different behaviour.

I don&apos;t see anything overly suspicious in v6.6/standard/base, but there are definitely some filesystems (yaffs2, aufs) changes that we carry that might be causing the issue.

Since I can&apos;t reproduce the problem on qemu, the only viable path that I see is via a bisect, or reverting some of the filesystem additions and re-resting on your setup.

Would either method be possible ?

All that being said, if linux-intel is also showing the problem, that points away from my suspicions about the fileystems as they shouldn&apos;t be present in that repository.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>98964</commentid>
    <comment_count>7</comment_count>
    <who name="Andreas Ufert">Andreas.Ufert</who>
    <bug_when>2024-05-02 17:47:07 +0000</bug_when>
    <thetext>Yes, I will gladly build images with whatever kernel configuration or branch you want me to test.

I have all three affected boards here with me so on workdays building and testing an image is a no-brainer. Weekends will work too, though there might be some delays before getting back with results.

Regarding meta-intel, should I file a report there as well or was adding Anuj sufficient?</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>98977</commentid>
    <comment_count>8</comment_count>
    <who name="Andreas Ufert">Andreas.Ufert</who>
    <bug_when>2024-05-08 12:09:42 +0000</bug_when>
    <thetext>Good news: With 6.6.25 (current kernel in meta-intel) everything works fine. MMC block device as well as reading mmc extcsd.

I can see in 6.6.24 there were some patches affecting mmc (https://cdn.kernel.org/pub/linux/kernel/v6.x/ChangeLog-6.6.24) though I can&apos;t really tell if it is related to the bug we were affected by.

However, I can cautiously give the all-clear so far.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>99629</commentid>
    <comment_count>9</comment_count>
    <who name="Bruce Ashfield">bruce.ashfield</who>
    <bug_when>2024-09-10 13:44:39 +0000</bug_when>
    <thetext>Let&apos;s call this &quot;fixed&quot; for now, and revisit if the oops re-appears.</thetext>
  </long_desc>
      
      

    </bug>

</bugzilla>