Created attachment 4872 [details] dmesg log [Summary] System not able to enter suspend state with the command below : rtcwake -m mem -s 10 [Image source] https://autobuilder.yocto.io/pub/releases/yocto-4.1_M1.rc1/machines/genericx86-64/ [Test target] NUC7 [Steps to reproduce] 1. Boot up the target image on NUC7 2. Open the terminal & run command "rtcwake -m mem -s 10" [Expected] System should go into suspend mode and wakeup from "mem" after 10 seconds. [Actual] System shows: rtcwake: wakeup from "mem" using /dev/rtc0 at Mon Jun 6 14:33:03 2022 rtcwake: write error [Failure rate] 60% (Fail 6 times out of 10 attempts)
Looks like the suspend/resume triggered issues from an earlier failure at boot: [ 5.328534] BUG: unable to handle page fault for address: 0000000000008000 [ 5.328937] #PF: supervisor instruction fetch in kernel mode [ 5.329240] #PF: error_code(0x0010) - not-present page [ 5.329509] PGD 0 P4D 0 [ 5.329605] Oops: 0010 [#1] PREEMPT SMP PTI [ 5.329810] CPU: 3 PID: 133 Comm: psplash Not tainted 5.15.36-yocto-standard #1 [ 5.330224] Hardware name: Intel Corporation NUC7i7BNH/NUC7i7BNB, BIOS BNKBL357.86A.0079.2019.0516.1758 05/16/2019 [ 5.330840] RIP: 0010:0x8000 [ 5.330959] Code: Unable to access opcode bytes at RIP 0x7fd6. [ 5.331274] RSP: 0018:ffffa45e801ffe08 EFLAGS: 00010206 [ 5.331549] RAX: ffffffff8ea1f2c0 RBX: ffff8cc640e04400 RCX: ffff8cc642129d81 [ 5.331953] RDX: 0000000000008000 RSI: 0000000000000001 RDI: ffff8cc640e04400 [ 5.332354] RBP: ffffa45e801ffe20 R08: ffff8cc640d0a3c0 R09: ffff8cc641280348 [ 5.332769] R10: ffffa45e801ffe30 R11: 0000000000000000 R12: ffff8cc640e04410 [ 5.333189] R13: ffff8cc641280348 R14: ffff8cc640d0a3e0 R15: ffff8cc6419bf0c0 [ 5.333608] FS: 0000000000000000(0000) GS:ffff8cc9aed80000(0000) knlGS:0000000000000000 [ 5.334094] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 [ 5.334417] CR2: 0000000000008000 CR3: 00000002a780a003 CR4: 00000000003706e0 [ 5.334837] Call Trace: [ 5.334929] <TASK> [ 5.334997] ? fb_release+0x3c/0x70 [ 5.335164] __fput+0x90/0x250 [ 5.335302] ____fput+0xe/0x10 [ 5.335437] task_work_run+0x64/0xa0 [ 5.335609] do_exit+0x321/0xa00 [ 5.335758] do_group_exit+0x3b/0xa0 [ 5.335931] __x64_sys_exit_group+0x18/0x20 [ 5.336146] do_syscall_64+0x42/0x90 [ 5.336318] entry_SYSCALL_64_after_hwframe+0x44/0xae [ 5.336593] RIP: 0033:0x7fba7fd48431 [ 5.336765] Code: Unable to access opcode bytes at RIP 0x7fba7fd48407. [ 5.337141] RSP: 002b:00007ffda69a7838 EFLAGS: 00000246 ORIG_RAX: 00000000000000e7 [ 5.337592] RAX: ffffffffffffffda RBX: 00007fba7fe5e510 RCX: 00007fba7fd48431 [ 5.338003] RDX: 000000000000003c RSI: 00000000000000e7 RDI: 0000000000000000 [ 5.338405] RBP: 0000000000000000 R08: ffffffffffffff80 R09: 00007ffda69a76f0 [ 5.338806] R10: 0000000000000000 R11: 0000000000000246 R12: 00007fba7fe5e510 [ 5.339208] R13: 0000000000000000 R14: 00007fba7fe5e9e8 R15: 00007fba7fe5ea00 [ 5.339611] </TASK> [ 5.339682] Modules linked in: bnep snd_hda_codec_hdmi snd_hda_codec_realtek snd_hda_codec_generic ledtrig_audio snd_hda_intel i915 snd_intel_dspcfg snd_hda_codec snd_hda_core snd_pcm ttm video snd_timer x86_pkg_temp_thermal backlight [ 5.341006] CR2: 0000000000008000 [ 5.341152] ---[ end trace 21e553179495ff52 ]--- [ 5.341386] RIP: 0010:0x8000 [ 5.341505] Code: Unable to access opcode bytes at RIP 0x7fd6. [ 5.341834] RSP: 0018:ffffa45e801ffe08 EFLAGS: 00010206 [ 5.342120] RAX: ffffffff8ea1f2c0 RBX: ffff8cc640e04400 RCX: ffff8cc642129d81 [ 5.342539] RDX: 0000000000008000 RSI: 0000000000000001 RDI: ffff8cc640e04400 [ 5.342959] RBP: ffffa45e801ffe20 R08: ffff8cc640d0a3c0 R09: ffff8cc641280348 [ 5.343379] R10: ffffa45e801ffe30 R11: 0000000000000000 R12: ffff8cc640e04410 [ 5.343798] R13: ffff8cc641280348 R14: ffff8cc640d0a3e0 R15: ffff8cc6419bf0c0 [ 5.344218] FS: 0000000000000000(0000) GS:ffff8cc9aed80000(0000) knlGS:0000000000000000 [ 5.344703] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 [ 5.345025] CR2: 0000000000008000 CR3: 00000002a780a003 CR4: 00000000003706e0 [ 5.345439] Fixing recursive fault but reboot is needed! Adding Bruce and Anuj, this looks like a kernel+hardware issue.
(in the framebuffer code?)
I noticed: 4f631f9f9d08 fbdev: Prevent possible use-after-free in fb_release() in Bruce's 5.15.43 update which may help this
Sorry for the delay, I had some infrastructure issues. That commit does look like a likely fix for what we are seeing. I obviously can't say for sure, but before we look at it any more, definitely try with the new -stable and see if it recurs.
Jay: Would you be able to test again with Bruce's kernel update applied? The patches are in master-next and testing on the autobuilder atm too.
Sure Richard, I can test it.
Tested with kernel version 5.15.44 and this issue seems to be resolved. Attempts [10/10] - System are able to go into suspend mode and resume without any issue.
Created attachment 4873 [details] dmesg log - v5.15.44
Great, thanks for testing and confirming. The changes merged to master: https://git.yoctoproject.org/poky/commit/?id=6ca79b6e93be29ff7bdf65db23b5b75988f7f1e3 https://git.yoctoproject.org/poky/commit/?id=3abe826c4bad5468bcbaeb5f6e9909b27fc389dd