<?xml version="1.0" encoding="UTF-8" standalone="yes" ?>
<!DOCTYPE bugzilla SYSTEM "https://bugzilla.yoctoproject.org/page.cgi?id=bugzilla.dtd">

<bugzilla version="5.0.6"
          urlbase="https://bugzilla.yoctoproject.org/"
          
          maintainer="it-coreprojects-helpdesk@linuxfoundation.org"
>

    <bug>
          <bug_id>16217</bug_id>
          
          <creation_ts>2026-03-26 09:55:31 +0000</creation_ts>
          <short_desc>AB-INT: boot hang in qemu (amba chip errors?)</short_desc>
          <delta_ts>2026-07-23 15:10:49 +0000</delta_ts>
          <reporter_accessible>1</reporter_accessible>
          <cclist_accessible>1</cclist_accessible>
          <classification_id>10</classification_id>
          <classification>QA/Testing</classification>
          <product>Runtime Testing</product>
          <component>testimage</component>
          <version>unspecified</version>
          <rep_platform>x86</rep_platform>
          <op_sys>Multiple</op_sys>
          <bug_status>RESOLVED</bug_status>
          <resolution>FIXED</resolution>
          
          
          <bug_file_loc></bug_file_loc>
          <status_whiteboard>AB-INT</status_whiteboard>
          <keywords></keywords>
          <priority>Medium+</priority>
          <bug_severity>normal</bug_severity>
          <target_milestone>6.1 M3</target_milestone>
          
          
          <everconfirmed>1</everconfirmed>
          <reporter name="Mathieu Dubois-Briand">mathieu.dubois-briand</reporter>
          <assigned_to name="Unassigned">unassigned</assigned_to>
          <cc>mikko.rapeli</cc>
    
    <cc>paul</cc>
    
    <cc>Quan.Sun</cc>
    
    <cc>randy.macleod</cc>
    
    <cc>richard.purdie</cc>
    
    <cc>ross.burton</cc>
    
    <cc>yoann.congal</cc>
          
          
          <cf_os>---</cf_os>
          <cf_regression_type>---</cf_regression_type>
          
          <cf_docchange>No (bug/feature does not impact docs)</cf_docchange>

      

      

      

          <comment_sort_order>oldest_to_newest</comment_sort_order>  
          <long_desc isprivate="0" >
    <commentid>104814</commentid>
    <comment_count>0</comment_count>
    <who name="Mathieu Dubois-Briand">mathieu.dubois-briand</who>
    <bug_when>2026-03-26 09:55:31 +0000</bug_when>
    <thetext>We had 3 new qemu boot hangs in the last days. They might be a bit related, but they look different enough to open separate entries. We might merge them later if they really are related.

This one is a kernel freeze, on display initialization. 


file: /srv/pokybuild/yocto-worker/qemuarmv5/build/build/tmp/work/qemuarmv5-poky-linux-gnueabi/core-image-sato/1.0/testimage/qemu_boot_log.20260320162711

[    2.945249] md: If you don&apos;t use raid, use raid=noautodetect
[    2.945298] md: Autodetecting RAID arrays.
[    2.945334] md: autorun ...
[    2.945372] md: ... autorun DONE.
[    3.133443] EXT4-fs (vda): mounted filesystem 88e1cf50-1f9b-4baa-948f-eea96bed87a3 r/w with ordered data mode. Quota mode: disabled.
[    3.133860] VFS: Mounted root (ext4 filesystem) on device 253:0.
[    3.135560] devtmpfs: mounted
[    3.172713] Freeing unused kernel image (initmem) memory: 512K
[    3.172778] Kernel memory protection not selected by kernel config.
[    3.173007] Run /sbin/init as init process

INIT: version 3.14 booting

[   13.289803] drm-clcd-pl111 10120000.display: set up callbacks for Versatile PL110
[   13.291214] amba 10100000.smc: deferred probe pending: (reason unknown)
[   13.291271] amba 10110000.mpmc: deferred probe pending: (reason unknown)
[   13.291290] amba 101e0000.sctl: deferred probe pending: (reason unknown)
[   13.291307] amba 101e1000.watchdog: deferred probe pending: (reason unknown)
[   13.291323] amba 101f0000.sci: deferred probe pending: (reason unknown)
[   13.291339] amba 101f4000.spi: deferred probe pending: (reason unknown)
[   13.291354] amba 1000a000.sci: deferred probe pending: (reason unknown)
[   13.291370] amba 10120000.display: deferred probe pending: (reason unknown)

I also note we had a DRM related issue on ARM a few months/weeks ago, but this is probably unrelated: https://bugzilla.yoctoproject.org/show_bug.cgi?id=16150.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>104819</commentid>
    <comment_count>1</comment_count>
    <who name="Mathieu Dubois-Briand">mathieu.dubois-briand</who>
    <bug_when>2026-03-26 09:56:59 +0000</bug_when>
    <thetext>qemuarmv5 alma8-vk-1 master-next&amp;master completed at 2026-03-20 16:52:58+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3257/steps/16/logs/stdio</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>104828</commentid>
    <comment_count>2</comment_count>
    <who name="Randy MacLeod">randy.macleod</who>
    <bug_when>2026-03-26 14:46:22 +0000</bug_when>
    <thetext>Likely a load-related intermittent arm failure.
It was an armv5 qemu so there&apos;s not much interest in digging into the root cause of issue.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>104857</commentid>
    <comment_count>3</comment_count>
    <who name="Mathieu Dubois-Briand">mathieu.dubois-briand</who>
    <bug_when>2026-03-30 08:50:16 +0000</bug_when>
    <thetext>qemuarmv5 rocky9-vk-2 master completed at 2026-03-30 01:44:03+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3320/steps/15/logs/stdio</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>104865</commentid>
    <comment_count>4</comment_count>
    <who name="Mathieu Dubois-Briand">mathieu.dubois-briand</who>
    <bug_when>2026-03-31 09:03:12 +0000</bug_when>
    <thetext>OK, this one is not amba related, nor armv5 related. But probably is load
related. I&apos;m logging this here, I will probably rename/reopen if it happens
again.

qemuarm64 debian12-vk-2 master completed at 2026-03-27 02:23:32+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/36/builds/3466/steps/15/logs/stdio</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>105054</commentid>
    <comment_count>5</comment_count>
    <who name="Mathieu Dubois-Briand">mathieu.dubois-briand</who>
    <bug_when>2026-04-13 08:30:38 +0000</bug_when>
    <thetext>qemuarmv5 alma8-vk-1 master completed at 2026-04-10 13:47:55+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3401/steps/16/logs/stdio

qemuarmv5 opensuse160-vk-2 master completed at 2026-04-12 01:42:03+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3407/steps/15/logs/stdio

qemuarmv5 ubuntu2204-vk-4 master&amp;master-next completed at 2026-04-12 23:06:38+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3411/steps/15/logs/stdio</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>105250</commentid>
    <comment_count>6</comment_count>
    <who name="Mathieu Dubois-Briand">mathieu.dubois-briand</who>
    <bug_when>2026-04-24 06:54:16 +0000</bug_when>
    <thetext>qemuarmv5 rocky8-vk-1 mathieu/master-next completed at 2026-04-23 07:56:50+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3483/steps/15/logs/stdio</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>105299</commentid>
    <comment_count>7</comment_count>
    <who name="Mathieu Dubois-Briand">mathieu.dubois-briand</who>
    <bug_when>2026-04-30 08:32:14 +0000</bug_when>
    <thetext>qemuarmv5 opensuse156-vk-1 mathieu/master-next&amp;mathieu/master-next-test1 completed at 2026-04-28 19:13:07+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3511/steps/15/logs/stdio</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>105361</commentid>
    <comment_count>8</comment_count>
    <who name="Mathieu Dubois-Briand">mathieu.dubois-briand</who>
    <bug_when>2026-05-06 08:21:26 +0000</bug_when>
    <thetext>qemuarmv5 debian11-vk-1 mathieu/master-next completed at 2026-05-06 05:37:11+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3539/steps/14/logs/stdio</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>105392</commentid>
    <comment_count>9</comment_count>
    <who name="Yoann Congal">yoann.congal</who>
    <bug_when>2026-05-08 06:56:21 +0000</bug_when>
    <thetext>qemuarmv5 opensuse156-vk-1 contrib/ycongal/wrynose-nut completed at 2026-05-08 01:10:53+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3557</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>105577</commentid>
    <comment_count>10</comment_count>
    <who name="Mathieu Dubois-Briand">mathieu.dubois-briand</who>
    <bug_when>2026-05-26 08:56:16 +0000</bug_when>
    <thetext>qemuarmv5 debian12-vk-5 master&amp;master-next completed at 2026-05-20 08:16:59+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3615/steps/15/logs/stdio</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>105688</commentid>
    <comment_count>11</comment_count>
    <who name="Paul Barker">paul</who>
    <bug_when>2026-06-01 12:26:35 +0000</bug_when>
    <thetext>Needs more discussion</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>105755</commentid>
    <comment_count>12</comment_count>
    <who name="Mathieu Dubois-Briand">mathieu.dubois-briand</who>
    <bug_when>2026-06-05 08:15:37 +0000</bug_when>
    <thetext>qemuarmv5 rocky8-vk-1 master&amp;master-next completed at 2026-06-04 23:08:47+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3695/steps/16/logs/stdio</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>105793</commentid>
    <comment_count>13</comment_count>
    <who name="Yoann Congal">yoann.congal</who>
    <bug_when>2026-06-10 07:17:50 +0000</bug_when>
    <thetext>qemuarmv5 fedora43-vk-2 wrynose completed at 2026-06-09 23:15:37+00:00
https://autobuilder.yoctoproject.org/valkyrie/?#/builders/80/builds/3723</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>105808</commentid>
    <comment_count>14</comment_count>
    <who name="Mathieu Dubois-Briand">mathieu.dubois-briand</who>
    <bug_when>2026-06-10 16:42:31 +0000</bug_when>
    <thetext>qemuarmv5 debian11-vk-3 mathieu/master-next completed at 2026-06-10 12:57:03+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3730/steps/14/logs/stdio</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>105963</commentid>
    <comment_count>15</comment_count>
    <who name="Yoann Congal">yoann.congal</who>
    <bug_when>2026-06-20 12:56:52 +0000</bug_when>
    <thetext>qemuarmv5 ubuntu2504-vk-1 wrynose completed at 2026-06-19 21:31:29+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3808/steps/14/logs/stdio</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>105970</commentid>
    <comment_count>16</comment_count>
    <who name="Mathieu Dubois-Briand">mathieu.dubois-briand</who>
    <bug_when>2026-06-22 09:05:19 +0000</bug_when>
    <thetext>qemuarmv5 fedora43-vk-2 stable/2.18-nut&amp;stable/wrynose-nut completed at 2026-06-09 23:15:37+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3723/steps/14/logs/stdio

qemuarmv5 fedora43-vk-2 master completed at 2026-06-22 01:44:47+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3824/steps/15/logs/stdio</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106003</commentid>
    <comment_count>17</comment_count>
    <who name="Mathieu Dubois-Briand">mathieu.dubois-briand</who>
    <bug_when>2026-06-25 10:23:42 +0000</bug_when>
    <thetext>qemuarmv5 opensuse160-vk-1 mathieu/master-next completed at 2026-06-24 08:45:51+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3839/steps/14/logs/stdio</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106164</commentid>
    <comment_count>18</comment_count>
    <who name="Mathieu Dubois-Briand">mathieu.dubois-briand</who>
    <bug_when>2026-07-13 08:11:12 +0000</bug_when>
    <thetext>qemuarmv5 alma9-vk-2 mathieu/master-next completed at 2026-07-12 08:45:31+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3958/steps/15/logs/stdio</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106179</commentid>
    <comment_count>19</comment_count>
    <who name="Mathieu Dubois-Briand">mathieu.dubois-briand</who>
    <bug_when>2026-07-16 07:30:13 +0000</bug_when>
    <thetext>qemuarmv5 opensuse156-vk-1 mathieu/master-next completed at 2026-07-15 20:12:50+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3970/steps/15/logs/stdio</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106209</commentid>
    <comment_count>20</comment_count>
    <who name="Yoann Congal">yoann.congal</who>
    <bug_when>2026-07-18 08:10:48 +0000</bug_when>
    <thetext>qemuarmv5 opensuse160-vk-1 wrynose completed at 2026-07-18 01:03:43+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3988</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106222</commentid>
    <comment_count>21</comment_count>
    <who name="Mathieu Dubois-Briand">mathieu.dubois-briand</who>
    <bug_when>2026-07-20 12:05:14 +0000</bug_when>
    <thetext>qemuarmv5 alma8-vk-1 mathieu/master-next completed at 2026-07-18 14:29:27+00:00
https://autobuilder.yoctoproject.org/valkyrie/#/builders/80/builds/3994/steps/15/logs/stdio</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106240</commentid>
    <comment_count>22</comment_count>
    <who name="Richard Purdie">richard.purdie</who>
    <bug_when>2026-07-21 13:19:01 +0000</bug_when>
    <thetext>Confirmed via https://bugzilla.yoctoproject.org/show_bug.cgi?id=16274 that there is no correlation in this case between cpu microcode version and the failures.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106241</commentid>
    <comment_count>23</comment_count>
    <who name="Richard Purdie">richard.purdie</who>
    <bug_when>2026-07-21 13:22:10 +0000</bug_when>
    <thetext>I did run my parallel image test on amba8-vk-1, the last machine to have failures, starting 100 images with 16 in parallel at once. Of those, 6 failed with the amba error on boot.

I ran a similar test on my local build system and cannot reproduce any. I am suspecting something about the CPU/kernel on the autobuilder workers.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106242</commentid>
    <comment_count>24</comment_count>
    <who name="Richard Purdie">richard.purdie</who>
    <bug_when>2026-07-21 13:40:03 +0000</bug_when>
    <thetext>With 64 started at once, around 3 crashed in amba

I also tried a limit of a max of 4 at once and after about 16 had run, I had 3 with the crash.

This hints the issue isn&apos;t load or memory related.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106243</commentid>
    <comment_count>25</comment_count>
    <who name="Richard Purdie">richard.purdie</who>
    <bug_when>2026-07-21 14:02:10 +0000</bug_when>
    <thetext>I tried copying the failing images from the autobuilder worker to my local system, I can&apos;t get them to fail there. This therefore has to be something about the autobuilder environment...</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106244</commentid>
    <comment_count>26</comment_count>
    <who name="Richard Purdie">richard.purdie</who>
    <bug_when>2026-07-21 14:13:05 +0000</bug_when>
    <thetext>To rule out some kind of native sstate relocation/compiler issue, I tried &quot; &quot;bitbake qemu-system-native -c compile -f&quot;  to force a locally compiled qemu binary. It still shows the amba hang.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106251</commentid>
    <comment_count>27</comment_count>
    <who name="Richard Purdie">richard.purdie</who>
    <bug_when>2026-07-22 08:02:41 +0000</bug_when>
    <thetext>https://valkyrie.yocto.io/pub/shared-failure-data/amba/qemu_boot_log.20260722075206 vs https://valkyrie.yocto.io/pub/shared-failure-data/amba/qemu_boot_log.20260722075206-suc shows the boot with &quot;initcall_debug  ignore_loglevel&quot; on the kernel commandline. It indicates the kernel is doing the same thing with a couple of minor ordering differences and the race/hang looks to be shortly after launching init. The amba deferred probe messages are present on both successful and unsuccessful boots.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106253</commentid>
    <comment_count>28</comment_count>
    <who name="Richard Purdie">richard.purdie</who>
    <bug_when>2026-07-22 08:26:59 +0000</bug_when>
    <thetext>After adding debug to set VERBOSE=very to initscripts in rcS-defaults and then adding echos to banner.sh, the boot is hanging with &quot;echo &gt; $vtmaster&quot; in banner.sh

This suggests something is sometimes going wrong with the console setup.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106256</commentid>
    <comment_count>29</comment_count>
    <who name="Richard Purdie">richard.purdie</who>
    <bug_when>2026-07-22 09:24:32 +0000</bug_when>
    <thetext>Further data points:

On successfully booting images, we see: &quot;warning: FBIOPUT_VSCREENINFO failed, double buffering disabled&quot; (which comes from psplash)

On hanging images, we do not see that message.

Removing psplash from the images means they always boot

Adding a &quot;sleep 5&quot; to the end of the splash init script doesn&apos;t help.

Adding a variant of https://patchwork.yoctoproject.org/project/yocto/patch/x5Mp.1642415168178512138.GgEV@lists.yoctoproject.org/ to psplash to double check we were allowed the double buffer doesn&apos;t seem to help (there were already other double buffering fixes in psplash).

banner.sh has vtdevice of /dev/tty0 and the write to that is the hang.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106257</commentid>
    <comment_count>30</comment_count>
    <who name="Richard Purdie">richard.purdie</who>
    <bug_when>2026-07-22 10:51:49 +0000</bug_when>
    <thetext>psplash is hung in ioctl(ConsoleFd, VT_GETMODE, &amp;vt_mode) in psplash_console_handle_switches() from psplash_console_switch()</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106260</commentid>
    <comment_count>31</comment_count>
    <who name="Richard Purdie">richard.purdie</who>
    <bug_when>2026-07-22 11:29:35 +0000</bug_when>
    <thetext>The VT_GETMODE is basically a lock on console_lock and then a memcpy so it is likely hanging on the lock.

With a &quot;echo t &gt; /proc/sysrq-trigger&quot; added to banner.sh, you get the following boot log for a hang:

https://valkyrie.yocto.io/pub/shared-failure-data/amba/qemu_boot_log.20260722111128

In that log, the following looks interesting to me with regard to some kind of lock issue:

Workqueue: events console_callback
[    9.108829] Call trace: 
[    9.108841]  __schedule from schedule+0x50/0x78
[    9.108879]  schedule from schedule_preempt_disabled+0x2c/0x3c
[    9.108912]  schedule_preempt_disabled from __mutex_lock.constprop.0+0x23c/0x3b8
[    9.108948]  __mutex_lock.constprop.0 from drm_fb_helper_pan_display+0x3c/0x154
[    9.108993]  drm_fb_helper_pan_display from fb_pan_display+0xd0/0x130
[    9.109035]  fb_pan_display from bit_update_start+0x1c/0x38
[    9.109073]  bit_update_start from fbcon_switch+0x2b8/0x3a0
[    9.109115]  fbcon_switch from redraw_screen+0x11c/0x198
[    9.109153]  redraw_screen from complete_change_console+0x3c/0xd4
[    9.109191]  complete_change_console from console_callback+0x5c/0x10c
[    9.109227]  console_callback from process_scheduled_works+0x218/0x368
[    9.109265]  process_scheduled_works from worker_thread+0x210/0x2a4
[    9.109302]  worker_thread from kthread+0x228/0x24c
[    9.109335]  kthread from ret_from_fork+0x14/0x28
[    9.109366] Exception stack(0xd0f41fb0 to 0xd0f41ff8)

[    8.821049] Workqueue: events virtio_gpu_dequeue_ctrl_func
[    8.821076] Call trace: 
[    8.821085]  __schedule from schedule+0x50/0x78
[    8.821116]  schedule from schedule_preempt_disabled+0x2c/0x3c
[    8.821140]  schedule_preempt_disabled from __ww_mutex_lock.constprop.0+0x488/0x624
[    8.821162]  __ww_mutex_lock.constprop.0 from modeset_lock+0xf0/0x100
[    8.821182]  modeset_lock from drm_modeset_lock_all_ctx+0x40/0xb8
[    8.821199]  drm_modeset_lock_all_ctx from drm_client_modeset_probe+0x3b0/0x1154
[    8.821218]  drm_client_modeset_probe from drm_fb_helper_hotplug_event+0xb8/0xdc
[    8.821238]  drm_fb_helper_hotplug_event from drm_client_hotplug+0x58/0x9c
[    8.821257]  drm_client_hotplug from drm_client_dev_hotplug+0x88/0x94
[    8.821275]  drm_client_dev_hotplug from virtio_gpu_dequeue_ctrl_func+0x1c4/0x25c
[    8.821294]  virtio_gpu_dequeue_ctrl_func from process_scheduled_works+0x218/0x368
[    8.821316]  process_scheduled_works from worker_thread+0x210/0x2a4
[    8.821335]  worker_thread from kthread+0x228/0x24c
[    8.821354]  kthread from ret_from_fork+0x14/0x28
[    8.821369] Exception stack(0xd0839fb0 to 0xd0839ff8)

[    8.820150] Workqueue: events drm_fb_helper_damage_work
[    8.820181] Call trace: 
[    8.820188]  __schedule from schedule+0x50/0x78
[    8.820208]  schedule from virtio_gpu_queue_fenced_ctrl_buffer+0x35c/0x480
[    8.820230]  virtio_gpu_queue_fenced_ctrl_buffer from virtio_gpu_cmd_resource_flush+0x98/0xc0
[    8.820253]  virtio_gpu_cmd_resource_flush from virtio_gpu_primary_plane_update+0x3b4/0x3d4
[    8.820273]  virtio_gpu_primary_plane_update from drm_atomic_helper_commit_planes+0x1a4/0x26c
[    8.820294]  drm_atomic_helper_commit_planes from drm_atomic_helper_commit_tail+0x30/0x68
[    8.820313]  drm_atomic_helper_commit_tail from commit_tail+0x144/0x154
[    8.820331]  commit_tail from drm_atomic_helper_commit+0xfc/0x10c
[    8.820349]  drm_atomic_helper_commit from drm_atomic_commit+0xc4/0xf8
[    8.820371]  drm_atomic_commit from drm_atomic_helper_dirtyfb+0x144/0x260
[    8.820391]  drm_atomic_helper_dirtyfb from drm_fbdev_shmem_helper_fb_dirty+0x70/0xd8
[    8.820411]  drm_fbdev_shmem_helper_fb_dirty from drm_fb_helper_damage_work+0xb8/0x208
[    8.820430]  drm_fb_helper_damage_work from process_scheduled_works+0x218/0x368
[    8.820450]  process_scheduled_works from worker_thread+0x210/0x2a4
[    8.820471]  worker_thread from kthread+0x228/0x24c
[    8.820489]  kthread from ret_from_fork+0x14/0x28
[    8.820506] Exception stack(0xd0831fb0 to 0xd0831ff8)

[    9.112321] task:psplash         state:D stack:0     pid:82    tgid:82    ppid:1      task_flags:0x400100 flags:0x00000004
[    9.112354] Call trace: 
[    9.112365]  __schedule from schedule+0x50/0x78
[    9.112399]  schedule from schedule_timeout+0x40/0x124
[    9.112431]  schedule_timeout from __down_common+0x1a8/0x204
[    9.112463]  __down_common from down+0x48/0x7c
[    9.112492]  down from console_lock+0x20/0x48
[    9.112531]  console_lock from vt_ioctl+0x820/0xea0
[    9.112571]  vt_ioctl from tty_ioctl+0x8b0/0x95c
[    9.112610]  tty_ioctl from sys_ioctl+0x694/0x918
[    9.112651]  sys_ioctl from ret_fast_syscall+0x0/0x44
[    9.112680] Exception stack(0xd0fe5fa8 to 0xd0fe5ff0)</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106264</commentid>
    <comment_count>32</comment_count>
    <who name="Richard Purdie">richard.purdie</who>
    <bug_when>2026-07-22 14:42:19 +0000</bug_when>
    <thetext>For fun, I asked Claude about this and the response follows:

This is a genuine deadlock, not just I/O stuck in D. Four tasks are involved, but only two of them form the actual cycle — the other two are victims piling up behind it. The nasty part is that one edge of the cycle runs through a virtqueue wait_event, not a lock, which is why nothing tripped lockdep and you got a silent hang instead of a splat.

Here&apos;s the lock state of each task:

T1 — events/console_callback — console_callback() takes console_lock() at entry, then walks complete_change_console → redraw_screen → fbcon_switch → fb_pan_display → drm_fb_helper_pan_display, which does mutex_lock(&amp;fb_helper-&gt;lock).
Holds: console_lock. Waiting on: fb_helper-&gt;lock.

T2 — events/virtio_gpu_dequeue_ctrl_func — the ctrl-queue drain worker has re-entered DRM via a hotplug: drm_client_dev_hotplug → drm_fb_helper_hotplug_event (which grabs fb_helper-&gt;lock) → drm_client_modeset_probe → drm_modeset_lock_all_ctx → __ww_mutex_lock.
Holds: fb_helper-&gt;lock (and it is the ctrlq drainer). Waiting on: a modeset ww_mutex (connection_mutex / crtc-plane locks).

T3 — events/drm_fb_helper_damage_work — a synchronous fbdev damage commit: drm_fbdev_shmem_helper_fb_dirty → drm_atomic_helper_dirtyfb (acquires the modeset locks via its acquire ctx) → drm_atomic_commit → commit_tail → virtio_gpu_primary_plane_update → virtio_gpu_cmd_resource_flush → virtio_gpu_queue_fenced_ctrl_buffer, sleeping inside virtio_gpu_queue_ctrl_sgs on ctrlq.ack_queue because the ring is full.
Holds: the modeset ww_mutex(es). Waiting on: free space in the ctrl virtqueue.

T4 — psplash (state D) — userspace vt_ioctl → console_lock() → down().
Waiting on: console_lock.

The cycle is T2 ⇄ T3:

T3 holds the modeset locks and is blocked waiting for the ctrl ring to drain.
The only thing that reclaims used ctrl buffers and wakes ctrlq.ack_queue is virtio_gpu_dequeue_ctrl_func — which is T2.
But T2 has been diverted into a hotplug re-probe and is now blocked on the very modeset lock T3 holds. A work_struct can&apos;t run twice concurrently, so re-queuing it on new completions does nothing; it won&apos;t loop back to drain the ring until this invocation returns, and it never will.

So: T3 needs T2 to drain the ring; T2 needs the modeset lock T3 is holding. Neither can move.

T1 and T4 are just collateral. T1 wants fb_helper-&gt;lock, which T2 is holding while deadlocked, so T1 wedges — and because console_callback() grabbed console_lock first, psplash&apos;s VT-switch ioctl (T4) then blocks forever on console_lock. Full chain: T4 → T1 → T2 ⇄ T3.

A couple of things worth calling out:

Lockdep wouldn&apos;t have saved you here even with PROVE_LOCKING on. The T3→T2 edge isn&apos;t a lock acquisition — it&apos;s a wait_event for descriptor space. Lockdep only sees the T2→T3 modeset-lock edge and the fb_helper-&gt;lock ordering; it can&apos;t model &quot;waiting on the worker that would wake me,&quot; so this reads as a plain hung task, not a circular-lock report.

The root cause is architectural in the virtio-gpu path: doing a synchronous DRM modeset re-probe (drm_client_dev_hotplug) inside the ctrl-queue dequeue worker. That worker must stay non-blocking with respect to anything that can wait on ring space, because it&apos;s the ring&apos;s only drain mechanism. The moment it takes a modeset lock that a stuck-on-ring commit holds, you have this loop. The trigger is timing: the ctrl ring filled (damage/resource_flush traffic outrunning the host drain) at the same moment a display-info/hotplug response sent the dequeue worker into a re-probe — classic early-boot collision, which fits your ~8–9s timestamps with psplash still up.

On fixes and mitigations:

The correct fix is kernel-side: the hotplug/re-probe must not run synchronously in virtio_gpu_dequeue_ctrl_func. It should be deferred to a separate work item (the config-changed worker) so the dequeue worker can return and keep reclaiming buffers. If you can move off this exact stable, check whether a newer 6.18.y or mainline already reworked the dequeue→hotplug path — this is exactly the kind of thing that gets a targeted fix, and I&apos;d rather you confirm against the actual changelog than take a commit hash from me that I can&apos;t verify.

Things you can try without a kernel patch, roughly in order of how cleanly they break the cycle:

Disable DRM fbdev emulation (drm_kms_helper.fbdev_emulation=0, or CONFIG_DRM_FBDEV_EMULATION=n). This removes drm_fb_helper_damage_work and the drm_fb_helper_hotplug_event re-probe outright, so both cycle edges disappear. Cost: no fbcon on virtio-gpu.
Drop the boot splash (psplash/plymouth). This only removes T1/T4 — the T2⇄T3 core can still theoretically deadlock — but in practice psplash&apos;s VT switching is a big part of what generates the concurrent damage + console traffic at boot, so it substantially reduces the trigger.
Reduce ring-full pressure: a larger virtqueue and/or making sure the host (QEMU/vhost) drains the ctrl queue promptly makes the wait_event for ring space rarely fire. This is a race-window mitigation, not a real fix — the bug is still latent.

This is 32-bit ARM (the X from Y trace format, ret_fast_syscall, the exception-stack lines) in what looks like a QEMU virtio-gpu guest, so if you&apos;re able to reproduce, the fastest confirmation is to add tracing on virtio_gpu_queue_ctrl_sgs&apos; ring-full wait and on drm_fb_helper_hotplug_event entry and watch them overlap while the ctrl ring sits at zero free descriptors.

If you can paste the l output (CPU backtraces) or a lockdep build&apos;s output, I can confirm the specific modeset lock that&apos;s contended and rule out my one inference — that drm_fb_helper_hotplug_event is holding fb_helper-&gt;lock across drm_client_modeset_probe — but the shape of the cycle doesn&apos;t depend on that detail.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106265</commentid>
    <comment_count>33</comment_count>
    <who name="Richard Purdie">richard.purdie</who>
    <bug_when>2026-07-22 14:45:02 +0000</bug_when>
    <thetext>I&apos;ve tested a kernel with this backport applied and I can&apos;t reproduce the hang with the patch applied:

https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/drivers/gpu/drm/virtio?id=d1b894c5bbb3fee0012bd14356286dc2384e8213</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106277</commentid>
    <comment_count>34</comment_count>
    <who name="Richard Purdie">richard.purdie</who>
    <bug_when>2026-07-22 19:22:09 +0000</bug_when>
    <thetext>Probably duplicates: https://bugzilla.yoctoproject.org/show_bug.cgi?id=16218 https://bugzilla.yoctoproject.org/show_bug.cgi?id=16261</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106287</commentid>
    <comment_count>35</comment_count>
    <who name="Mikko Rapeli">mikko.rapeli</who>
    <bug_when>2026-07-23 09:05:11 +0000</bug_when>
    <thetext>Is it possible to see host machine CPU, memory, IO and networking load when this bug happens in qemu?

I presume these are very high and thus something adds delays to the complex qemu graphics stack which the kernel drivers under qemu can not cope with (could be a bug in kernel for sure). The specific load or combination of different CPU, memory/swap could be one way to trigger this.

I guess this happens on multiple host OSes with different kernel versions. These could also contribute to the timing differences which make this so hard to reproduce anywhere else. I presume the host HW is quite identical, or is it just one physical machine?

I&apos;ve seen severe issues with some virtualization setups where IO was also virtual and actually over network which resulted in annoying and hard to debug bugs in host VMs and their workloads like bitbake builds. I presume qemu machines would have been very unreliable in those setups.</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106304</commentid>
    <comment_count>36</comment_count>
    <who name="Richard Purdie">richard.purdie</who>
    <bug_when>2026-07-23 15:09:42 +0000</bug_when>
    <thetext>*** Bug 16261 has been marked as a duplicate of this bug. ***</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106306</commentid>
    <comment_count>37</comment_count>
    <who name="Richard Purdie">richard.purdie</who>
    <bug_when>2026-07-23 15:10:10 +0000</bug_when>
    <thetext>*** Bug 16218 has been marked as a duplicate of this bug. ***</thetext>
  </long_desc><long_desc isprivate="0" >
    <commentid>106307</commentid>
    <comment_count>38</comment_count>
    <who name="Richard Purdie">richard.purdie</who>
    <bug_when>2026-07-23 15:10:49 +0000</bug_when>
    <thetext>https://git.openembedded.org/openembedded-core/commit/?id=496ab6151ead169d9f4873d1e2a3f5e646c73b07</thetext>
  </long_desc>
      
      

    </bug>

</bugzilla>