Bug 4469

Summary: Shutdown fails for qemuimagetest on danny.
Product: [QA/Testing] Runtime Testing Reporter: Beth Flanagan <elizabeth.flanagan>
Component: generalAssignee: Beth Flanagan <elizabeth.flanagan>
Status: RESOLVED WORKSFORME QA Contact:
Severity: normal    
Priority: Medium+ CC: alexandru.c.georgescu, dvhart, elizabeth.flanagan, ross.burton, sgw
Version: 1.3.1   
Target Milestone: 1.3.2   
Hardware: x86   
OS: Multiple   
URL: http://autobuilder.yoctoproject.org:8011/builders/nightly-qa-extras/builds/8/steps/Running%20Sanity%20Tests/logs/stdio
Whiteboard: (autobuilder)
OS type for building Yocto: --- Type of Regression: ---
Verified: Documentation change: Don't know

Description Beth Flanagan 2013-05-08 22:20:21 UTC
Seeing this on the danny branch with no rhyme or reason:

http://autobuilder.yoctoproject.org:8011/builders/nightly-qa-extras/builds/8/steps/Running%20Sanity%20Tests/logs/stdio

| spawn ssh -o UserKnownHostsFile=/dev/null -o StrictHostKeyChecking=no root@192.168.7.2 /sbin/poweroff
| Warning: Permanently added '192.168.7.2' (RSA) to the list of known hosts.
| root@192.168.7.2's password:
| 	Test_Info: Shutdown Test FAIL
| killing 15188
| killing 14052
| NOTE: 	Test Result for qemux86-64 core-image-sato
| NOTE: 	Testcase       PASS           FAIL           NORESULT
| NOTE: 	SSH            1              0              0
| NOTE: 	SCP            1              0              0
| NOTE: 	zypper_help    1              0              0
| NOTE: 	zypper_search  1              0              0
| NOTE: 	rpm_query      1              0              0
| NOTE: 	connman        1              0              0
| NOTE: 	dmesg          1              0              0
| NOTE: 	shutdown       0              1              0
| DEBUG: Python function do_qemuimagetest_standalone finished
| ERROR: Function failed: Some testcases fail, pls. check test result and test log!!!


Not sure if this is our race condition popping up in danny.
Comment 1 Ross Burton 2013-05-09 11:53:19 UTC
Looks like https://bugzilla.yoctoproject.org/show_bug.cgi?id=3806 isn't resolved then.

This is on ab05, which I can't ssh into for some reason.

CC'ing Darren.  Both qemux86 and x86-64 in this build are without PV.
Comment 2 Ross Burton 2013-05-09 13:30:59 UTC
Screenshot of a shutdown crash in a qemux86 without PV enabled:

http://i.imgur.com/8o5uQsl.png

Working on replicating with a console log now.
Comment 3 Darren Hart 2013-05-09 14:41:24 UTC
Ross, interesting. Can you get me the addr2line of that kernel again?
Comment 4 Saul Wold 2013-05-09 14:52:32 UTC
Ross, check with Halstead to get access to AB05, and check for the build is actually correct with paravirt disabled and that we are not building for sstate somehow.
Comment 5 Ross Burton 2013-05-09 15:55:22 UTC
addr2line says:

native_nmi_stop_other_cpus()
linux/arch/x86/kernel/smp.c:204

grepping the config in the generated kernel package:

# CONFIG_PARAVIRT_GUEST is not set
Comment 6 Ross Burton 2013-05-09 16:11:07 UTC
Full log:

Shutdown: hda
ACPI: Preparing to enter system sleep state S5
Disabling non-boot CPUs ...
Power down.
general protection fault: fffa [#1] PREEMPT SMP 
Modules linked in: iptable_nat nf_nat nf_conntrack_ipv4 nf_defrag_ipv4 nf_conntrack ip_tables x_tables uvesafb cfbfillrect cfbimgblt cfbcopyarea

Pid: 1057, comm: halt Not tainted 3.4.11-yocto-standard #3 Bochs Bochs
EIP: 0060:[<c101c97a>] EFLAGS: 00000246 CPU: 0
EIP is at native_nmi_stop_other_cpus+0xca/0xe0
EAX: 00000000 EBX: 00000246 ECX: fffff000 EDX: 000000ff
ESI: 00000001 EDI: 4ad85ff4 EBP: c7101e78 ESP: c7101e68
 DS: 007b ES: 007b FS: 00d8 GS: 0033 SS: 0068
CR0: 8005003b CR2: 08050564 CR3: 071ec000 CR4: 00000690
DR0: 00000000 DR1: 00000000 DR2: 00000000 DR3: 00000000
DR6: 00000000 DR7: 00000000
Process halt (pid: 1057, ti=c7100000 task=c70b0c60 task.ti=c7100000)
Stack:
 c7101e70 00000000 4321fedc 4ad85ff4 c7101e80 c101bdfc c7101e88 c101bce6
 c7101e90 c101bf4e c7101e9c c1041715 c1832d14 c7101fac c10420a1 c7101eb0
 c105a43b 00000001 c7101ebc c16908d3 c70b0c01 c7101ec4 c168d667 c7101edc
Call Trace:
 [<c101bdfc>] native_machine_shutdown+0x5c/0x90
 [<c101bce6>] native_machine_power_off+0x36/0x50
 [<c101bf4e>] machine_power_off+0xe/0x10
 [<c1041715>] kernel_power_off+0x65/0x70
 [<c10420a1>] sys_reboot+0xf1/0x1b0
 [<c105a43b>] ? get_parent_ip+0xb/0x40
 [<c16908d3>] ? sub_preempt_count+0x43/0xb0
 [<c168d667>] ? _raw_spin_unlock_irqrestore+0x17/0x40
 [<c109ae14>] ? rcu_report_unblock_qs_rnp+0x24/0x80
 [<c109b008>] ? rcu_read_unlock_special+0x198/0x1a0
 [<c109b061>] ? __rcu_read_unlock+0x51/0x60
 [<c1040c9f>] ? sys_kill+0x7f/0x180
 [<c107428d>] ? clockevents_program_event+0x9d/0x140
 [<c1050a1c>] ? hrtimer_interrupt+0x16c/0x280
 [<c110225d>] ? do_sys_open+0x15d/0x1c0
 [<c110225d>] ? do_sys_open+0x15d/0x1c0
 [<c105a43b>] ? get_parent_ip+0xb/0x40
 [<c16908d3>] ? sub_preempt_count+0x43/0xb0
 [<c1034f65>] ? irq_exit+0x65/0xa0
 [<c1693fb9>] ? smp_apic_timer_interrupt+0x59/0x85
 [<c168dd59>] syscall_call+0x7/0xb
Code: 47 ff 74 1d 89 c7 b8 c7 10 00 00 e8 71 1b 2d 00 0f b6 03 f3 0f b8 c0 90 83 f8 01 77 dc 8d 74 26 00 9c 5b fa e8 e8 10 00 00 53 9d <8b> 5d f4 8b 75 f8 8b 7d fc 89 ec 5d c3 89 f6 8d bc 27 00 00 00 
EIP: [<c101c97a>] native_nmi_stop_other_cpus+0xca/0xe0 SS:ESP 0068:c7101e68
---[ end trace c03ab0e63e9de447 ]---
Comment 7 Ross Burton 2013-05-10 13:12:31 UTC
And a trace from qemux86-64:

Power down.
general protection fault: fff2 [#1] PREEMPT SMP 
CPU 0 
Modules linked in: iptable_nat nf_nat nf_conntrack_ipv4 nf_defrag_ipv4 nf_conntrack ip_tables x_tables uvesafb cfbfillrect cfbimgblt cfbcopyarea

Pid: 1076, comm: halt Not tainted 3.4.11-yocto-standard #1 Bochs Bochs
RIP: 0010:[<ffffffff8101b0ea>]  [<ffffffff8101b0ea>] native_nmi_stop_other_cpus+0xea/0x100
RSP: 0018:ffff880006d1bdd8  EFLAGS: 00000246
RAX: ffffffff81c18ce0 RBX: 0000000000000246 RCX: 0000000000000000
RDX: 0000000000000001 RSI: 00000000000000ff RDI: 00000000000000f0
RBP: ffff880006d1bdf8 R08: 0000000000000000 R09: 0a2e6e776f642072
R10: 4320746f6f622d6e R11: 0000000000000000 R12: 0000000000000001
R13: 00007fff40ac8080 R14: 00000000fee1dead R15: 0000000000000001
FS:  00007f86bd4ac700(0000) GS:ffff880007c00000(0000) knlGS:0000000000000000
CS:  0010 DS: 0000 ES: 0000 CR0: 000000008005003b
CR2: 00007f94194ff9b8 CR3: 00000000063a2000 CR4: 00000000000006b0
DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
DR3: 0000000000000000 DR6: 0000000000000000 DR7: 0000000000000000
Process halt (pid: 1076, threadinfo ffff880006d1a000, task ffff8800063face0)
Stack:
 000000004321fedc 000000004321fedc 0000000028121969 00007fff40ac8080
 ffff880006d1be08 ffffffff8101a50c ffff880006d1be18 ffffffff8101a3f6
 ffff880006d1be28 ffffffff8101a58f ffff880006d1be38 ffffffff81045cd2
Call Trace:
 [<ffffffff8101a50c>] native_machine_shutdown+0x5c/0x80
 [<ffffffff8101a3f6>] native_machine_power_off+0x36/0x50
 [<ffffffff8101a58f>] machine_power_off+0xf/0x20
 [<ffffffff81045cd2>] kernel_power_off+0x72/0x80
 [<ffffffff8104677b>] sys_reboot+0x13b/0x200
 [<ffffffff810a8e75>] ? rcu_read_unlock_special+0x1c5/0x1d0
 [<ffffffff810a8ee6>] ? __rcu_read_unlock+0x66/0x70
 [<ffffffff81045100>] ? sys_kill+0xa0/0x1d0
 [<ffffffff8164329d>] ? sub_preempt_count+0x6d/0xd0
 [<ffffffff81115300>] ? kmem_cache_free+0x20/0x120
 [<ffffffff8111ee72>] ? do_sys_open+0x172/0x1e0
 [<ffffffff81646dd2>] system_call_fastpath+0x16/0x1b
Code: 89 c5 bf c7 10 00 00 e8 85 d5 31 00 48 8b 3b 81 e7 ff ff ff 00 f3 48 0f b8 c7 83 f8 01 77 d2 66 90 9c 5b fa e8 08 11 00 00 53 9d <48> 8b 5d e8 4c 8b 65 f0 4c 8b 6d f8 c9 c3 0f 1f 84 00 00 00 00 
RIP  [<ffffffff8101b0ea>] native_nmi_stop_other_cpus+0xea/0x100
 RSP <ffff880006d1bdd8>
---[ end trace 4685912b96066099 ]---

Same location, native_nmi_stop_other_cpus() at linux/arch/x86/kernel/smp.c:204.
Comment 8 Alexandru Georgescu 2013-08-22 06:42:29 UTC
Hi Beth,
I don't know why this landed to me. 
Is this obsolete?

Thanks!
Comment 9 Alexandru Georgescu 2013-08-22 06:57:41 UTC
This appeared before 1.3.1 rc1 and in the RC1 report, all the tests passed (including shutdown): https://wiki.yoctoproject.org/wiki/Full_Pass_Test_Report_for_Yocto_1.3.2_RC1_201305163_Build


Is it safe to mark it resolved worksorme?
Comment 10 Beth Flanagan 2013-08-26 17:42:13 UTC
Marking this resolved worksforme as it seems the issue has been fixed.