| Summary: | Linux 3.15+ w/ 0.74 or 0.75 firmware causes kernel panics | ||||||
|---|---|---|---|---|---|---|---|
| Product: | [Hardware Platforms] MinnowBoard MAX Firmware | Reporter: | John 'Warthog9' Hawley <warthog9> | ||||
| Component: | minnowmax-edk2 | Assignee: | John 'Warthog9' Hawley <warthog9> | ||||
| Status: | RESOLVED FIXED | QA Contact: | |||||
| Severity: | critical | ||||||
| Priority: | High | CC: | david.wei, dvhart, ivan.rouzanov, michael.p.krau, mike.wu, shifeix.a.lu, sjolley.yp.pm, warthog9 | ||||
| Version: | 2C A1 | ||||||
| Target Milestone: | Production Release | ||||||
| Hardware: | MinnowBoard Max | ||||||
| OS: | x86_64 | ||||||
| Whiteboard: | |||||||
| OS type for building Yocto: | --- | Type of Regression: | Regression (Used to work) | ||||
| Verified: | Documentation change: | No (bug/feature does not impact docs) | |||||
| Attachments: |
|
||||||
|
Description
John 'Warthog9' Hawley
2014-11-22 00:04:35 UTC
I'd consider this a blocking but to the release of 0.75 (and the associated source) (In reply to comment #1) > I'd consider this a blocking but to the release of 0.75 (and the associated > source) I have to agree. The 3.15 kernel is going on 9 months old at this point (a long time in the Linux world). I'd consider this our top priority of OS and firmware to resolve at this point. Moving to firmware since 3.15+ kernels worked in the past, prior to firmware 0.74. Please keep an OS resource on this as well! The data indicates that the problem is not a wholly Firmware issue, as the fact that the firmware worked fine until the 3.15 release indicates. Now instead of getting into a finger pointing position, (which I am trying to avoid), it seems to me the problem is actually due to a 'collision' between the firmware at 0.74 and the Kernel at 3.15 (something changed in each that precipitates the panic). It would be best if we could attack this problem from both ends and meeting in the middle. I suspect in the end the resulting fix will come on the firmware side (typical in these situations), but approaching the problem from both sides can usually get us to root cause and subsequent fix faster. We have 3 OTC folks looking at this from the Linux kernel side. Sometime near the 3.14-3.15 window, the Linux kernel changed the way EFI runtime services are mapped to work around an issue on larger systems: commit b7b898ae0c0a82489511a1ce1b35f26215e6beb5 Author: Borislav Petkov <bp@suse.de> Date: 2014-01-18 x86/efi: Make efi virtual runtime map passing more robust This change, however, relies on NX support, which is disabled for some reason. Possibly by the MinnowBoard firmware? 20 [ 0.000000] Notice: NX (Execute Disable) protection missing in CPU! Later, during the 3.18 window, a patch was introduced to check for failed mappings, which catches this bug, but disables EFI runtime services. $ git show a5a750a commit a5a750a98fe2812263546cc20badd32bab21ec52 Author: Dave Young <dyoung@redhat.com> Date: 2014-08-14 x86/efi: Clear EFI_RUNTIME_SERVICES if failing to enter virtual mode With this patch applied, the kernel does boot, but we observe the following warnings: 191 [ 0.051145] Error mapping PA 0x71000000 -> VA 0x71000000! 192 [ 0.057209] Error mapping PA 0x71000000 -> VA 0xfffffffeffe00000! 193 [ 0.064052] Error mapping PA 0x78141000 -> VA 0x78141000! 194 [ 0.070113] Error mapping PA 0x78141000 -> VA 0xfffffffeffd41000! 195 [ 0.076951] Error mapping PA 0x78252000 -> VA 0x78252000! 196 [ 0.083010] Error mapping PA 0x78252000 -> VA 0xfffffffefec52000! 197 [ 0.089848] Error mapping PA 0x797ae000 -> VA 0x797ae000! 198 [ 0.095906] Error mapping PA 0x797ae000 -> VA 0xfffffffefe9ae000! 199 [ 0.102744] Error mapping PA 0x79bf8000 -> VA 0x79bf8000! 200 [ 0.108802] Error mapping PA 0x79bf8000 -> VA 0xfffffffefe7f8000! 201 [ 0.115639] Error mapping PA 0x79cf8000 -> VA 0x79cf8000! 202 [ 0.121698] Error mapping PA 0x79cf8000 -> VA 0xfffffffefe4f8000! 203 [ 0.128535] Error mapping PA 0x7a638000 -> VA 0x7a638000! 204 [ 0.134594] Error mapping PA 0x7a638000 -> VA 0xfffffffefda38000! 205 [ 0.141430] Error mapping PA 0xe00f8000 -> VA 0xe00f8000! 206 [ 0.147489] Error mapping PA 0xe00f8000 -> VA 0xfffffffefd8f8000! 207 [ 0.154325] Error mapping PA 0xfed01000 -> VA 0xfed01000! 208 [ 0.160384] Error mapping PA 0xfed01000 -> VA 0xfffffffefd701000! 209 [ 0.167234] Error ident-mapping new memmap (0x77c56000)! Finally, the old mapping behavior can be used to bypass the runtime mapping and boot a 3.15-3.17 Linux kernel by adding the following to the kernel command line (this is not a fix, but it is a useful confirmation of the above): efi=old_map This explains the the behavior of the Linux kernel. There are two pending questions for firmware: 1) What changed between 0.73 and 0.74 which triggered this failed mapping. 2) Why is NX disabled? If NX was disabled between 0.73 and 0.74, this might be easily resolved by reverting that change. One more data point: The same kernel which fails to boot with 0.74 firmware: root@intel-corei7-64:~# uname -r [14:28] <dvhart> 3.17.0bad-yocto-standard-dirty Does boot on 0.73 firmware. Also, NX is active: [14:28] <dvhart> root@intel-corei7-64:~# dmesg | grep NX [14:28] <dvhart> [ 0.000000] NX (Execute Disable) protection: active So, it appears NX was disabled between 0.73 and 0.74 and this breaks the runtime mapping of EFI runtime introduced in the 3.15 and later kernels. Proposed resolution: Please re-enable NX in the firmware. I can confirm that Darren's suggestion about efi=old_map works for a 3.16 Fedora kernel, so this bounds the problem from the kernel side it looks like Created attachment 2258 [details] Clear “XD Bit Disable”(Bit 34 of IA32_MISC_ENABLE(1A0) ) Attached test image gets “XD Bit Disable”(Bit 34 of IA32_MISC_ENABLE(1A0) ) cleared, in order for NX feature to be supported. Tested with Linux version 3.17.3-200.fc20.x86_64 Log: [test@localhost ~]$ dmesg [ 0.000000] Initializing cgroup subsys cpuset [ 0.000000] Initializing cgroup subsys cpu [ 0.000000] Initializing cgroup subsys cpuacct [ 0.000000] Linux version 3.17.3-200.fc20.x86_64 (mockbuild@bkernel02.phx2.fedoraproject.org) (gcc version 4.8.3 20140911 (Red Hat 4.8.3-7) (GCC) ) #1 SMP Fri Nov 14 19:45:42 UTC 2014 [ 0.000000] Command line: BOOT_IMAGE=/vmlinuz-3.17.3-200.fc20.x86_64 root=UUID=765d481d-dfc1-4db3-a2ca-4c2bf198a11c ro vconsole.font=latarcyrheb-sun16 rhgb quiet LANG=en_US.UTF-8 [ 0.000000] e820: BIOS-provided physical RAM map: [ 0.000000] BIOS-e820: [mem 0x0000000000000000-0x000000000008efff] usable [ 0.000000] BIOS-e820: [mem 0x000000000008f000-0x000000000008ffff] ACPI NVS [ 0.000000] BIOS-e820: [mem 0x0000000000090000-0x000000000009dfff] usable [ 0.000000] BIOS-e820: [mem 0x000000000009e000-0x000000000009ffff] reserved [ 0.000000] BIOS-e820: [mem 0x0000000000100000-0x000000001fffffff] usable [ 0.000000] BIOS-e820: [mem 0x0000000020000000-0x00000000200fffff] reserved [ 0.000000] BIOS-e820: [mem 0x0000000020100000-0x0000000079b47fff] usable [ 0.000000] BIOS-e820: [mem 0x0000000079b48000-0x000000007a373fff] reserved [ 0.000000] BIOS-e820: [mem 0x000000007a374000-0x000000007a473fff] ACPI NVS [ 0.000000] BIOS-e820: [mem 0x000000007a474000-0x000000007a4b3fff] ACPI data [ 0.000000] BIOS-e820: [mem 0x000000007a4b4000-0x000000007affffff] usable [ 0.000000] BIOS-e820: [mem 0x00000000e00f8000-0x00000000e00f8fff] reserved [ 0.000000] BIOS-e820: [mem 0x00000000fed01000-0x00000000fed01fff] reserved [ 0.000000] NX (Execute Disable) protection: active [ 0.000000] efi: EFI v2.40 by EDK II [ 0.000000] efi: ACPI=0x7a4b3000 ACPI 2.0=0x7a4b3014 SMBIOS=0x79b6f000 [ 0.000000] efi: mem00: type=7, attr=0xf, range=[0x0000000000000000-0x0000000000001000) (0MB) [ 0.000000] efi: mem01: type=2, attr=0xf, range=[0x0000000000001000-0x0000000000002000) (0MB) [ 0.000000] efi: mem02: type=7, attr=0xf, range=[0x0000000000002000-0x000000000008f000) (0MB) [ 0.000000] efi: mem03: type=10, attr=0xf, range=[0x000000000008f000-0x0000000000090000) (0MB) [ 0.000000] efi: mem04: type=7, attr=0xf, range=[0x0000000000090000-0x000000000009e000) (0MB) [ 0.000000] efi: mem05: type=0, attr=0xf, range=[0x000000000009e000-0x00000000000a0000) (0MB) [ 0.000000] efi: mem06: type=7, attr=0xf, range=[0x0000000000100000-0x0000000001000000) (15MB) Thank you David. I can't get to testing this today, but I can tomorrow. I'll update with the result, but this looks very promising. Thank you! Assign back to John to review the attached BIOS image. (In reply to comment #11) > Assign back to John to review the attached BIOS image. 0.76 resolves this with 3.15 and 3.16. Once 0.76 is released we can close this bug. 0.76 released http://firmware.intel.com/projects/minnowboard-max this bug is complete. |