Bug 6993 - Linux 3.15+ w/ 0.74 or 0.75 firmware causes kernel panics
Summary: Linux 3.15+ w/ 0.74 or 0.75 firmware causes kernel panics
Status: RESOLVED FIXED
Alias: None
Product: MinnowBoard MAX Firmware
Classification: Hardware Platforms
Component: minnowmax-edk2 (show other bugs)
Version: 2C A1
Hardware: MinnowBoard Max x86_64
: High critical
Target Milestone: Production Release
Assignee: John 'Warthog9' Hawley
QA Contact:
URL:
Whiteboard:
Depends on:
Blocks:
 
Reported: 2014-11-22 00:04 UTC by John 'Warthog9' Hawley
Modified: 2015-01-26 22:06 UTC (History)
8 users (show)

See Also:
OS type for building Yocto: ---
Type of Regression: Regression (Used to work)
Verified:
Documentation change: No (bug/feature does not impact docs)


Attachments
Clear “XD Bit Disable”(Bit 34 of IA32_MISC_ENABLE(1A0) ) (2.03 MB, application/x-zip-compressed)
2014-11-26 01:40 UTC, David_Wei
no flags Details

Note You need to log in before you can comment on or make changes to this bug.
Description John 'Warthog9' Hawley 2014-11-22 00:04:35 UTC
Something changed between 0.73 to 0.74 (and has been carried into 0.75) (noting that 0.74 was never released, and 0.75 is currently in final testing and not public yet) that's causing 3.15 or newer kernels to panic on boot up:

---- Boot log follows ----

[    3.594881] rtc_cmos 00:00: rtc core: registered rtc_cmos as rtc0
[    3.604634] rtc_cmos 00:00: alarms up to one month, y3k, 242 bytes nvram
[    3.615242] device-mapper: uevent: version 1.0.3
[    3.623527] device-mapper: ioctl: 4.27.0-ioctl (2013-10-30)
initialised: dm-devel@redhat.com
[    3.636755] EFI Variables Facility v0.08 2004-May-17
[    3.645293] BUG: unable to handle kernel NULL pointer dereference at
          (null)
[    3.657088] IP: [<          (null)>]           (null)
[    3.665724] PGD 0
[    3.670861] Oops: 0010 [#1] SMP
[    3.677329] Modules linked in:
[    3.683544] CPU: 1 PID: 1 Comm: swapper/0 Not tainted
3.16.3-200.fc20.x86_64 #1
[    3.694619] Hardware name: Circuitco Minnowboard Max B3
PLATFORM/MinnowBoard MAX, BIOS MNW2MAX1.X64.0075.R01.1411050927 11/05/2014
[    3.710828] task: ffff88007a830000 ti: ffff88007a82c000 task.ti:
ffff88007a82c000
[    3.722286] RIP: 0010:[<0000000000000000>]  [<          (null)>]
      (null)
[    3.733771] RSP: 0000:ffff88007a82fd78  EFLAGS: 00010046
[    3.738586] usb 1-2: new low-speed USB device number 2 using xhci_hcd
[    3.753086] RAX: ffffffff81ff2380 RBX: ffff88007a992fc0 RCX:
0000000000000000
[    3.764191] RDX: ffff88007a82fdc0 RSI: ffff88003f801400 RDI:
ffff88007a82fdb8
[    3.775290] RBP: ffff88007a82fe28 R08: 0000000000017460 R09:
ffff880075401500
[    3.786374] R10: ffffffff815a4814 R11: ffffea0000fe0800 R12:
ffff88003f801400
[    3.797424] R13: 0000000000000000 R14: ffffffff815a6170 R15:
0000000000000000
[    3.808479] FS:  0000000000000000(0000) GS:ffff880079480000(0000)
knlGS:0000000000000000
[    3.820609] CS:  0010 DS: 0000 ES: 0000 CR0: 000000008005003b
[    3.830092] CR2: 0000000000000000 CR3: 0000000001c11000 CR4:
00000000001007e0
[    3.841159] Stack:
[    3.846450]  ffffffff815a4852 0000000081360c5d ffff8800747ea1f8
0000000000000000
[    3.857945]  ffffffff815a5ce0 ffffffff81cccaa0 0001880000000000
ffffffff81ff2380
[    3.869446]  0000000000000400 0000000000000000 ffff88007a82fdd8
ffffffff81361c6b
[    3.880958] Call Trace:
[    3.886828]  [<ffffffff815a4852>] ? efivar_init+0xb2/0x380
[    3.896148]  [<ffffffff815a5ce0>] ? efivar_update_sysfs_entries+0x70/0x70
[    3.906913]  [<ffffffff81361c6b>] ? kobject_uevent+0xb/0x10
[    3.913218] usb 1-2: New USB device found, idVendor=17ef, idProduct=6009
[    3.913222] usb 1-2: New USB device strings: Mfr=1, Product=2,
SerialNumber=0
[    3.913226] usb 1-2: Product: ThinkPad USB Keyboard with TrackPoint
[    3.913229] usb 1-2: Manufacturer: Lite-On Technology Corp.
[    3.957775]  [<ffffffff81360a46>] ? kset_register+0x56/0x70
[    3.967257]  [<ffffffff815a6170>] ? efivar_create+0x410/0x410
[    3.976905]  [<ffffffff815a61fd>] efivars_sysfs_init+0x8d/0x230
[    3.986719]  [<ffffffff81002144>] do_one_initcall+0xd4/0x210
[    3.996204]  [<ffffffff810ae185>] ? parse_args+0x235/0x450
[    4.005465]  [<ffffffff81d3229d>] kernel_init_freeable+0x190/0x22c
[    4.015508]  [<ffffffff81d319d6>] ? initcall_blacklist+0xb0/0xb0
[    4.025349]  [<ffffffff816f9090>] ? rest_init+0x80/0x80
[    4.034283]  [<ffffffff816f909e>] kernel_init+0xe/0xf0
[    4.043075]  [<ffffffff8170e63c>] ret_from_fork+0x7c/0xb0
[    4.052120]  [<ffffffff816f9090>] ? rest_init+0x80/0x80
[    4.060906] Code:  Bad RIP value.
[    4.067528] RIP  [<          (null)>]           (null)
[    4.076167]  RSP <ffff88007a82fd78>
[    4.082934] CR2: 0000000000000000
[    4.089523] ---[ end trace 012d0b0a2daa505a ]---
[    4.097748] Kernel panic - not syncing: Attempted to kill init!
exitcode=0x00000009
[    4.097748]
[    4.097960] usb 1-2: ep 0x81 - rounding interval to 64 microframes,
ep desc says 80 microframes
[    4.097972] usb 1-2: ep 0x82 - rounding interval to 64 microframes,
ep desc says 80 microframes
[    4.138931] Kernel Offset: 0x0 from 0xffffffff81000000 (relocation
range: 0xffffffff80000000-0xffffffff9fffffff)
[    4.153177] ---[ end Kernel panic - not syncing: Attempted to kill
init! exitcode=0x00000009
[    4.153177]
[    4.169958] ------------[ cut here ]------------
[    4.177960] WARNING: CPU: 1 PID: 1 at arch/x86/kernel/smp.c:124
native_smp_send_reschedule+0x5d/0x60()
[    4.191326] Modules linked in:
[    4.197693] CPU: 1 PID: 1 Comm: swapper/0 Tainted: G      D
3.16.3-200.fc20.x86_64 #1
[    4.210073] Hardware name: Circuitco Minnowboard Max B3
PLATFORM/MinnowBoard MAX, BIOS MNW2MAX1.X64.0075.R01.1411050927 11/05/2014
[    4.226376]  0000000000000000 00000000d0f96c47 ffff880079483d90
ffffffff81707091
[    4.237856]  0000000000000000 ffff880079483dc8 ffffffff8108d0ad
0000000000000000
[    4.249330]  ffff8800794145c0 0000000000000001 0000000000000001
000000000000e608
[    4.260788] Call Trace:
[    4.266594]  <IRQ>  [<ffffffff81707091>] dump_stack+0x45/0x56
[    4.276172]  [<ffffffff8108d0ad>] warn_slowpath_common+0x7d/0xa0
[    4.286020]  [<ffffffff8108d1da>] warn_slowpath_null+0x1a/0x20
[    4.295612]  [<ffffffff8104489d>] native_smp_send_reschedule+0x5d/0x60
[    4.305991]  [<ffffffff810cf0d4>] trigger_load_balance+0x144/0x1b0
[    4.315986]  [<ffffffff810bf3a7>] scheduler_tick+0x97/0xd0
[    4.325182]  [<ffffffff8109b2c0>] update_process_times+0x60/0x70
[    4.334937]  [<ffffffff810fdc85>] tick_sched_handle.isra.17+0x25/0x60
[    4.345186]  [<ffffffff810fdd01>] tick_sched_timer+0x41/0x60
[    4.354512]  [<ffffffff810b3794>] __run_hrtimer+0x74/0x1d0
[    4.363635]  [<ffffffff810fdcc0>] ? tick_sched_handle.isra.17+0x60/0x60
[    4.374037]  [<ffffffff810b3b97>] hrtimer_interrupt+0x107/0x250
[    4.383654]  [<ffffffff81047627>] local_apic_timer_interrupt+0x37/0x60
[    4.393983]  [<ffffffff817114df>] smp_apic_timer_interrupt+0x3f/0x60
[    4.404120]  [<ffffffff8170f5dd>] apic_timer_interrupt+0x6d/0x80
[    4.413875]  <EOI>  [<ffffffff817028f0>] ? panic+0x1cb/0x20c
[    4.423273]  [<ffffffff8108fd87>] do_exit+0xa37/0xa40
[    4.431973]  [<ffffffff81702fe1>] ? printk+0x77/0x8e
[    4.440554]  [<ffffffff8101736c>] oops_end+0x9c/0xe0
[    4.449140]  [<ffffffff8170203d>] no_context+0x2c2/0x2e5
[    4.458118]  [<ffffffff817020d3>] __bad_area_nosemaphore+0x73/0x1ca
[    4.468186]  [<ffffffff8170223d>] bad_area_nosemaphore+0x13/0x15
[    4.477952]  [<ffffffff81059bef>] __do_page_fault+0x9f/0x540
[    4.487333]  [<ffffffff8135e8a2>] ? get_from_free_list+0x42/0x50
[    4.497105]  [<ffffffff8135fa30>] ? ida_get_new_above+0x230/0x2a0
[    4.506991]  [<ffffffff815a6170>] ? efivar_create+0x410/0x410
[    4.516482]  [<ffffffff8105a0b2>] do_page_fault+0x22/0x30
[    4.525537]  [<ffffffff81710688>] page_fault+0x28/0x30
[    4.534255]  [<ffffffff815a6170>] ? efivar_create+0x410/0x410
[    4.543608]  [<ffffffff815a4814>] ? efivar_init+0x74/0x380
[    4.552701]  [<ffffffff815a4852>] ? efivar_init+0xb2/0x380
[    4.561674]  [<ffffffff815a5ce0>] ? efivar_update_sysfs_entries+0x70/0x70
[    4.572047]  [<ffffffff81361c6b>] ? kobject_uevent+0xb/0x10
[    4.581013]  [<ffffffff81360a46>] ? kset_register+0x56/0x70
[    4.589967]  [<ffffffff815a6170>] ? efivar_create+0x410/0x410
[    4.599043]  [<ffffffff815a61fd>] efivars_sysfs_init+0x8d/0x230
[    4.608246]  [<ffffffff81002144>] do_one_initcall+0xd4/0x210
[    4.617124]  [<ffffffff810ae185>] ? parse_args+0x235/0x450
[    4.625698]  [<ffffffff81d3229d>] kernel_init_freeable+0x190/0x22c
[    4.634963]  [<ffffffff81d319d6>] ? initcall_blacklist+0xb0/0xb0
[    4.644021]  [<ffffffff816f9090>] ? rest_init+0x80/0x80
[    4.652177]  [<ffffffff816f909e>] kernel_init+0xe/0xf0
[    4.660170]  [<ffffffff8170e63c>] ret_from_fork+0x7c/0xb0
[    4.668435]  [<ffffffff816f9090>] ? rest_init+0x80/0x80
[    4.676473] ---[ end trace 012d0b0a2daa505b ]---

---- ^ Boot log above ^ ----

The panic is more or less consistent across 3.15 and 3.16, 3.17 it gets longer and more verbose but looks basically like the same thing.

Doing some quick testing I'm showing:

Firmware 0.73:

3.11.10-301.fc20.x86_64	- Boots fine
3.14.4-yocto-standard	- Boots fine
3.15.10-200.fc20.x86_64	- Boots fine
3.16.3-200.fc20.x86_64	- Boots fine
3.17.3-200.fc20.x86_64	- Boots fine

Firmware 0.74:
3.11.10-301.fc20.x86_64	- Boots fine
3.14.4-yocto-standard	- Boots fine
3.15.10-200.fc20.x86_64	- Kernel Panics complaining about EFI variables
3.16.3-200.fc20.x86_64	- Kernel Panics complaining about EFI variables
3.17.3-200.fc20.x86_64	- Kernel Panics complaining about EFI variables

Firmware 0.75:
3.11.10-301.fc20.x86_64	- Boots fine
3.14.4-yocto-standard	- Boots fine
3.15.10-200.fc20.x86_64	- Kernel Panics complaining about EFI variables
3.16.3-200.fc20.x86_64	- Kernel Panics complaining about EFI variables
3.17.3-200.fc20.x86_64	- Kernel Panics complaining about EFI variables

Just my own reading of the trace seems to indicate something is getting set to null in the firmware (EFI variable) and when Linux is attempting to parse it, it can't handle that being null, kernel panic ensues.  Obviously Linux shouldn't be panicing there, but I think that's a parallel problem to a firmware change causing a system to no longer boot.
Comment 1 John 'Warthog9' Hawley 2014-11-22 00:08:00 UTC
I'd consider this a blocking but to the release of 0.75 (and the associated source)
Comment 2 Darren Hart 2014-11-24 16:28:07 UTC
(In reply to comment #1)
> I'd consider this a blocking but to the release of 0.75 (and the associated
> source)

I have to agree. The 3.15 kernel is going on 9 months old at this point (a long time in the Linux world). I'd consider this our top priority of OS and firmware to resolve at this point.
Comment 3 Darren Hart 2014-11-24 16:29:02 UTC
Moving to firmware since 3.15+ kernels worked in the past, prior to firmware 0.74.
Comment 4 Michael Krau 2014-11-24 17:55:13 UTC
Please keep an OS resource on this as well!

The data indicates that the problem is not a wholly Firmware issue, as the fact that the firmware worked fine until the 3.15 release indicates.

Now instead of getting into a finger pointing position, (which I am trying to avoid), it seems to me the problem is actually due to a 'collision' between the firmware at 0.74 and the Kernel at 3.15 (something changed in each that precipitates the panic).  

It would be best if we could attack this problem from both ends and meeting in the middle.  I suspect in the end the resulting fix will come on the firmware side (typical in these situations), but approaching the problem from both sides can usually get us to root cause and subsequent fix faster.
Comment 5 Darren Hart 2014-11-24 18:50:43 UTC
We have 3 OTC folks looking at this from the Linux kernel side.
Comment 6 Darren Hart 2014-11-24 21:56:29 UTC
Sometime near the 3.14-3.15 window, the Linux kernel changed the way EFI runtime services are mapped to work around an issue on larger systems:

commit b7b898ae0c0a82489511a1ce1b35f26215e6beb5
Author: Borislav Petkov <bp@suse.de>
Date:   2014-01-18

    x86/efi: Make efi virtual runtime map passing more robust

This change, however, relies on NX support, which is disabled for some reason. Possibly by the MinnowBoard firmware?

20 [    0.000000] Notice: NX (Execute Disable) protection missing in CPU!

Later, during the 3.18 window, a patch was introduced to check for failed mappings, which catches this bug, but disables EFI runtime services.

$ git show a5a750a
commit a5a750a98fe2812263546cc20badd32bab21ec52
Author: Dave Young <dyoung@redhat.com>
Date:   2014-08-14

    x86/efi: Clear EFI_RUNTIME_SERVICES if failing to enter virtual mode

With this patch applied, the kernel does boot, but we observe the following warnings:

191 [    0.051145] Error mapping PA 0x71000000 -> VA 0x71000000!
192 [    0.057209] Error mapping PA 0x71000000 -> VA 0xfffffffeffe00000!
193 [    0.064052] Error mapping PA 0x78141000 -> VA 0x78141000!
194 [    0.070113] Error mapping PA 0x78141000 -> VA 0xfffffffeffd41000!
195 [    0.076951] Error mapping PA 0x78252000 -> VA 0x78252000!
196 [    0.083010] Error mapping PA 0x78252000 -> VA 0xfffffffefec52000!
197 [    0.089848] Error mapping PA 0x797ae000 -> VA 0x797ae000!
198 [    0.095906] Error mapping PA 0x797ae000 -> VA 0xfffffffefe9ae000!
199 [    0.102744] Error mapping PA 0x79bf8000 -> VA 0x79bf8000!
200 [    0.108802] Error mapping PA 0x79bf8000 -> VA 0xfffffffefe7f8000!
201 [    0.115639] Error mapping PA 0x79cf8000 -> VA 0x79cf8000!
202 [    0.121698] Error mapping PA 0x79cf8000 -> VA 0xfffffffefe4f8000!
203 [    0.128535] Error mapping PA 0x7a638000 -> VA 0x7a638000!
204 [    0.134594] Error mapping PA 0x7a638000 -> VA 0xfffffffefda38000!
205 [    0.141430] Error mapping PA 0xe00f8000 -> VA 0xe00f8000!
206 [    0.147489] Error mapping PA 0xe00f8000 -> VA 0xfffffffefd8f8000!
207 [    0.154325] Error mapping PA 0xfed01000 -> VA 0xfed01000!
208 [    0.160384] Error mapping PA 0xfed01000 -> VA 0xfffffffefd701000!
209 [    0.167234] Error ident-mapping new memmap (0x77c56000)!

Finally, the old mapping behavior can be used to bypass the runtime mapping and boot a 3.15-3.17 Linux kernel by adding the following to the kernel command line (this is not a fix, but it is a useful confirmation of the above):

efi=old_map

This explains the the behavior of the Linux kernel.

There are two pending questions for firmware:

1) What changed between 0.73 and 0.74 which triggered this failed mapping.

2) Why is NX disabled?

If NX was disabled between 0.73 and 0.74, this might be easily resolved by reverting that change.
Comment 7 Darren Hart 2014-11-24 22:31:10 UTC
One more data point:

The same kernel which fails to boot with 0.74 firmware:
root@intel-corei7-64:~# uname -r
[14:28]  <dvhart> 3.17.0bad-yocto-standard-dirty

Does boot on 0.73 firmware. Also, NX is active:

[14:28]  <dvhart> root@intel-corei7-64:~# dmesg | grep NX
[14:28]  <dvhart> [    0.000000] NX (Execute Disable) protection: active

So, it appears NX was disabled between 0.73 and 0.74 and this breaks the runtime mapping of EFI runtime introduced in the 3.15 and later kernels.

Proposed resolution: Please re-enable NX in the firmware.
Comment 8 John 'Warthog9' Hawley 2014-11-25 00:49:18 UTC
I can confirm that Darren's suggestion about efi=old_map works for a 3.16 Fedora kernel, so this bounds the problem from the kernel side it looks like
Comment 9 David_Wei 2014-11-26 01:40:50 UTC
Created attachment 2258 [details]
Clear “XD Bit Disable”(Bit 34 of  IA32_MISC_ENABLE(1A0) )

Attached test image gets “XD Bit Disable”(Bit 34 of  IA32_MISC_ENABLE(1A0) ) cleared, in order for NX feature to be supported.  

Tested with Linux version 3.17.3-200.fc20.x86_64  

Log:
[test@localhost ~]$ dmesg
[    0.000000] Initializing cgroup subsys cpuset
[    0.000000] Initializing cgroup subsys cpu
[    0.000000] Initializing cgroup subsys cpuacct
[    0.000000] Linux version 3.17.3-200.fc20.x86_64 (mockbuild@bkernel02.phx2.fedoraproject.org) (gcc version 4.8.3 20140911 (Red Hat 4.8.3-7) (GCC) ) #1 SMP Fri Nov 14 19:45:42 UTC 2014
[    0.000000] Command line: BOOT_IMAGE=/vmlinuz-3.17.3-200.fc20.x86_64 root=UUID=765d481d-dfc1-4db3-a2ca-4c2bf198a11c ro vconsole.font=latarcyrheb-sun16 rhgb quiet LANG=en_US.UTF-8
[    0.000000] e820: BIOS-provided physical RAM map:
[    0.000000] BIOS-e820: [mem 0x0000000000000000-0x000000000008efff] usable
[    0.000000] BIOS-e820: [mem 0x000000000008f000-0x000000000008ffff] ACPI NVS
[    0.000000] BIOS-e820: [mem 0x0000000000090000-0x000000000009dfff] usable
[    0.000000] BIOS-e820: [mem 0x000000000009e000-0x000000000009ffff] reserved
[    0.000000] BIOS-e820: [mem 0x0000000000100000-0x000000001fffffff] usable
[    0.000000] BIOS-e820: [mem 0x0000000020000000-0x00000000200fffff] reserved
[    0.000000] BIOS-e820: [mem 0x0000000020100000-0x0000000079b47fff] usable
[    0.000000] BIOS-e820: [mem 0x0000000079b48000-0x000000007a373fff] reserved
[    0.000000] BIOS-e820: [mem 0x000000007a374000-0x000000007a473fff] ACPI NVS
[    0.000000] BIOS-e820: [mem 0x000000007a474000-0x000000007a4b3fff] ACPI data
[    0.000000] BIOS-e820: [mem 0x000000007a4b4000-0x000000007affffff] usable
[    0.000000] BIOS-e820: [mem 0x00000000e00f8000-0x00000000e00f8fff] reserved
[    0.000000] BIOS-e820: [mem 0x00000000fed01000-0x00000000fed01fff] reserved
[    0.000000] NX (Execute Disable) protection: active
[    0.000000] efi: EFI v2.40 by EDK II
[    0.000000] efi:  ACPI=0x7a4b3000  ACPI 2.0=0x7a4b3014  SMBIOS=0x79b6f000 
[    0.000000] efi: mem00: type=7, attr=0xf, range=[0x0000000000000000-0x0000000000001000) (0MB)
[    0.000000] efi: mem01: type=2, attr=0xf, range=[0x0000000000001000-0x0000000000002000) (0MB)
[    0.000000] efi: mem02: type=7, attr=0xf, range=[0x0000000000002000-0x000000000008f000) (0MB)
[    0.000000] efi: mem03: type=10, attr=0xf, range=[0x000000000008f000-0x0000000000090000) (0MB)
[    0.000000] efi: mem04: type=7, attr=0xf, range=[0x0000000000090000-0x000000000009e000) (0MB)
[    0.000000] efi: mem05: type=0, attr=0xf, range=[0x000000000009e000-0x00000000000a0000) (0MB)
[    0.000000] efi: mem06: type=7, attr=0xf, range=[0x0000000000100000-0x0000000001000000) (15MB)
Comment 10 Darren Hart 2014-11-26 01:51:18 UTC
Thank you David. I can't get to testing this today, but I can tomorrow. I'll update with the result, but this looks very promising. Thank you!
Comment 11 David_Wei 2014-12-02 07:36:05 UTC
Assign back to John to review the attached BIOS image.
Comment 12 John 'Warthog9' Hawley 2014-12-15 22:50:32 UTC
(In reply to comment #11)
> Assign back to John to review the attached BIOS image.

0.76 resolves this with 3.15 and 3.16.  Once 0.76 is released we can close this bug.
Comment 13 John 'Warthog9' Hawley 2015-01-26 22:06:01 UTC
0.76 released

http://firmware.intel.com/projects/minnowboard-max

this bug is complete.