Bug 5155 - System hangs during booting: BUG: Bad page map in process Xorg pte:0000x310 pmd:3f03e067
Summary: System hangs during booting: BUG: Bad page map in process Xorg pte:0000x310 ...
Status: RESOLVED WONTFIX
Alias: None
Product: BSPs
Classification: Build System, Metadata & Runtime
Component: bsps-meta-intel (show other bugs)
Version: 1.4.2
Hardware: Crownbay x86
: Medium critical
Target Milestone: Future
Assignee: Nuhairi Anuar
QA Contact:
URL:
Whiteboard:
Depends on:
Blocks:
 
Reported: 2013-09-10 07:49 UTC by Maksym Zhelieznyi
Modified: 2014-07-10 05:46 UTC (History)
6 users (show)

See Also:
OS type for building Yocto: ---
Type of Regression: ---
Verified:
Documentation change: Don't know


Attachments
dmesg and lsmod outputs (7.43 KB, application/zip)
2013-09-10 07:49 UTC, Maksym Zhelieznyi
no flags Details
boot logs (1.18 MB, image/jpeg)
2013-10-30 12:16 UTC, Maksym Zhelieznyi
no flags Details
boot logs (1.07 MB, image/jpeg)
2013-10-30 12:17 UTC, Maksym Zhelieznyi
no flags Details
boot logs (1.13 MB, image/jpeg)
2013-10-30 12:17 UTC, Maksym Zhelieznyi
no flags Details
boot logs (1.01 MB, image/jpeg)
2013-10-30 12:19 UTC, Maksym Zhelieznyi
no flags Details

Note You need to log in before you can comment on or make changes to this bug.
Description Maksym Zhelieznyi 2013-09-10 07:49:23 UTC
Created attachment 1482 [details]
dmesg and lsmod outputs

System hangs on boot process on stage when Yocto progress bar almost on the end. Sometimes it's possible to login in text terminal (after switching with Ctrl+alt+f1), but usually system hangs and doesn't respond to anything.
Reproducibility: 1-10%

Occurs with:
- released Yocto Project 1.4.2 - Poky 9.0.2 with released meta-crownbay layer (9.0.0 "Dylan")
- Yocto Project 1.5_M4.rc3 with meta-crownbay layer 1.5_M4.rc3

Test machine - Intel® Atom™ Processor E680T (512K Cache, 1.60 GHz)

Steps to reproduce:
1. Build core-image-sato image
2. Copy it on machine's SSD
3. Boot system


Please see in attached files output of lsmod and dmesg commands (case when it was possible to login)
Comment 1 Darren Hart 2013-09-10 18:26:03 UTC
Can you try booting with psplash=false on the kernel command line? This will prevent the graphical splash screen from appearing. Then please provide the final lines of output at the time of the hang.
Comment 2 Maksym Zhelieznyi 2013-09-11 10:13:36 UTC
(In reply to comment #1)
> Can you try booting with psplash=false on the kernel command line? This will
> prevent the graphical splash screen from appearing. Then please provide the
> final lines of output at the time of the hang.


Below output which you asked about:

X.Org X Server 1.9.3
Release Date: 2010-12-13
X Protocol Version 11, Revision 0
Build Operating System: Linux 3.0.0-32-generic-pae i686
Current Operating System: Linux crownbay 3.8.13-yocto-standard #1 SMP PREEMPT Fri Sep 6 17:32:08 EEST 2013 i686
Kernel command line: ro root=/dev/sda2 video=vesafb vga=0x318 psplash=false BOOT_IMAGE=/boot/bzImage
Build Date 06 September 2013 05:33:40PM
Current version of pitman: 0.30.2
	Before reporting problems, check http://wiki.x.org
	to make sure that you have latest version.
Markers: (--) probed, (**) from config file, (==) default setting,
	(++) from command line, (!!) notice, (II)informational,
	(WW) warning, (EE) error, (NI) not implemented, (??) unknown.
(==) Log file: "/var/log/Xorg.0.log", Time: Fri Sep 6 16:20:00 2013
(==) Using config file: "etc/X11/xorg.conf"
(==) Using system config directory "use/share/X11/xorg.conf.d"
Public key portion is:
ssh-rsa XXXXXXXXXXXXXX…….. root@crownbay
Fingerprint: md5 XX:XX:XX…..
dropbear.
Starting Advanced Configuration and Power Interface daemon: INFO: rcu_preempt detected stalls on CPUs/tasks: {} (detected by 0, t=21002, jiffies, g=749, c=748, q=277)
INFO: Stall ended before state dump start
INFO: rcu_preempt detected stalls on CPUs/tasks: {} (detected by 0, t=84007, jiffies, g=749, c=748, q=529)
INFO: Stall ended before state dump start
INFO: rcu_preempt detected stalls on CPUs/tasks: {} (detected by 0, t=147012, jiffies, g=749, c=748, q=781)
INFO: Stall ended before state dump start
INFO: rcu_preempt detected stalls on CPUs/tasks: {} (detected by 0, t=210017, jiffies, g=749, c=748, q=1033)

Last lines are repeated permanently with different values of 't' and 'q'. Keyboard doesn't response to anything (even num or caps buttons).
Comment 3 Maksym Zhelieznyi 2013-09-11 14:56:57 UTC
(In reply to comment #1)
> Can you try booting with psplash=false on the kernel command line? This will
> prevent the graphical splash screen from appearing. Then please provide the
> final lines of output at the time of the hang.

One more case:

...
INIT: Entering runlevel: 5
Starting system message bus: dbus.
Starting Connection Manager
Starting Xserver
Netfilter messages via NETLINK v0.30.
ip_tables: (C) 2000-2006 Netfilter Core Team
Starting Dropbear SSH server: nf_conntrack version 0.5.0 (15890 buckets, 63560 max)
Will output1024 bit rsa secret key to '/varlib/dropbear/dropbear_rsa_host_key'
Generating key, this may take a while...
IPv6: ADDRCONF(NETDEV_UP): eth0: link is not ready
IPv6: ADDRCONF(NETDEV_CHANGE): eth0: link becomes ready
Public key portion is:
ssh-rsa XXXXXXXX.... root-crownbay
Fingerprint: md5 XX:XX:XX...
dropbear.
Starting Advanced Configuration and Power Interface daemon: acpid.
acpid: starting up

hwclock: can't open '/dev/misc/rtc': No such file or directory
acpid: 1 rule loaded

acpid: waiting for events: event logging is off

Starting syslogd/klogd: done
 * Starting Avahi mDNS/DNS-SD Daemon: avahi-daemon    [fail]
Starting Telephony daemon
Starting Linux NFC daemon
Stopping Bootlog daemon: bootlogd.

Poky Next (Yocto Project Reference Distro) 1.4+snapshot-20130906 crownbay /dev/tty1

=====
And then a prompt to login into tty1. In this case I'm able to login in system as I mentioned before
Comment 4 Nitin Kamble 2013-09-11 17:48:54 UTC
I tried to reproduce this issue on my crownbay hardware. And After 10 tries I could not reproduce the issue.

What hardware are you using?

It is possible that your hardware got damaged, especially the RAM.
Comment 5 Maksym Zhelieznyi 2013-09-11 18:11:20 UTC
(In reply to comment #4)
> I tried to reproduce this issue on my crownbay hardware. And After 10 tries
> I could not reproduce the issue.

It can take big number of attempts to catch this issue. To make it easier I've added script which reboots machine in loop. Even in this case machine can work without the fail for 2-3 hours.

> What hardware are you using?

I use conga-QA6/E680-1G Qseven module with Intel® Atom™ E680 processor with 1.6GHz, 512kB L2 cache and 1GB DDR2 onboard memory. Based on Intel® Platform Controller Hub EG20. 8GB MLC SSD
http://www.congatec.com/en/products/qseven/conga-qa6.html
 
> It is possible that your hardware got damaged, especially the RAM.
I think that possibility that hardware got damaged is low because I reproduce this issue on all machines that available for me (4 units)
Comment 6 Nitin Kamble 2013-09-12 05:01:51 UTC
humm, so you are not using crownbay hardware. This BSP was specifically made for crownbay platform. I think in your case the BIOS or VBIOS may be different from what crownbay has.
Comment 7 Maksym Zhelieznyi 2013-09-13 10:27:29 UTC
(In reply to comment #6)
> humm, so you are not using crownbay hardware. This BSP was specifically made
> for crownbay platform. I think in your case the BIOS or VBIOS may be
> different from what crownbay has.

Here is information about ours BIOS:
    BIOS Version 2.14.1219 American Megatrends. Ver. QTOPR111
    IGB VBIOS Version 1904
if you need additional information let me know please
Comment 8 Nitin Kamble 2013-10-14 16:27:58 UTC
We have 3.10 kernel working with EMGD 1.18 driver now. Once that is upstream, it can be checked if that helps with this issue.
Comment 9 Maksym Zhelieznyi 2013-10-30 12:16:00 UTC
Created attachment 1592 [details]
boot logs
Comment 10 Maksym Zhelieznyi 2013-10-30 12:17:04 UTC
Created attachment 1593 [details]
boot logs
Comment 11 Maksym Zhelieznyi 2013-10-30 12:17:25 UTC
Created attachment 1594 [details]
boot logs
Comment 12 Maksym Zhelieznyi 2013-10-30 12:19:03 UTC
Created attachment 1595 [details]
boot logs
Comment 13 Maksym Zhelieznyi 2013-10-30 12:30:58 UTC
I've verified the system (core-image-sato) with Kernel 3.10 and EMGD v1.18. Issue is still reproduced. 
I've attached 4 screenshots with logs (in kernel command line were set parameters: loglevel=7 and psplash=false).
Sometimes during boot the string about bad swap file is printed:
swap_dup: Bad swap file entry 0000000N
Comment 14 Nitin Kamble 2013-11-06 18:07:46 UTC
This involves the binary EMGD X driver from ISG. As I do not have sources for the driver I am passing this bug to ISG.
Comment 15 Chen Yong 2014-04-11 11:10:33 UTC
Hi,
The EMGD for ATOM Processor E600 series already in maintenance.
Is there still a need to investigate this issue? I may not be able to find someone to help to look reproduce the issue.

Set "Target Milestone" to Future
Comment 16 Nuhairi Anuar 2014-07-10 05:44:38 UTC
The EMGD for ATOM Processor E600 series already in maintenance. There will be no schedule release anymore for EMGD on Intel Atom E600. However, EMGD still support EMGD for Atom E600 through our PAE for critical issues. If you are working on a project with Intel's PAEs support, please contact them and filed the bug through them.
Comment 17 Nuhairi Anuar 2014-07-10 05:46:55 UTC
The EMGD for ATOM Processor E600 series already in maintenance. There will be no schedule release anymore for EMGD on Intel Atom E600. However, EMGD still support EMGD for Atom E600 through our PAE for critical issues. If you are working on a project with Intel's PAEs support, please contact them and filed the bug through them.