Bug 11375 - testimage; runqemu not working when images are downloaded (not built)
Summary: testimage; runqemu not working when images are downloaded (not built)
Status: RESOLVED FIXED
Alias: None
Product: OE-Core
Classification: Build System, Metadata & Runtime
Component: Scripts and Tools (show other bugs)
Version: 2.3
Hardware: x86 Multiple
: High normal
Target Milestone: 2.3 M4
Assignee: brian avery
QA Contact:
URL:
Whiteboard:
Depends on:
Blocks:
 
Reported: 2017-04-18 15:52 UTC by Jose Perez C
Modified: 2017-04-25 17:01 UTC (History)
6 users (show)

See Also:
OS type for building Yocto: ---
Type of Regression: ---
Verified:
Documentation change: No (bug/feature does not impact docs)


Attachments
full log (21.49 KB, application/octet-stream)
2017-04-18 15:52 UTC, Jose Perez C
no flags Details
qemu error.log (21.19 KB, application/octet-stream)
2017-04-18 23:29 UTC, Juan Ramos
no flags Details
log with sed command (45.46 KB, text/plain)
2017-04-19 13:28 UTC, Jose Perez C
no flags Details

Note You need to log in before you can comment on or make changes to this bug.
Description Jose Perez C 2017-04-18 15:52:38 UTC
Created attachment 3731 [details]
full log

commit 1fa1a7f174593e41b8bcf6c2f19565d6da44e991

STEPS: 

1. Download artifacts from AB
  wget:   
   - http://autobuilder.yoctoproject.org/pub/releases/yocto-2.3_M3/machines/qemu/qemux86/core-image-minimal-qemux86.ext4
   - http://autobuilder.yoctoproject.org/pub/releases/yocto-2.3_M3/machines/qemu/qemux86/core-image-minimal-qemux86.manifest
   - http://autobuilder.yoctoproject.org/pub/releases/yocto-2.3_M3/machines/qemu/qemux86/bzImage-qemux86.bin
   -  http://autobuilder.yoctoproject.org/pub/releases/yocto-2.3_M3/machines/qemu/qemux86/core-image-minimal-qemux86.qemuboot.conf
   - http://autobuilder.yoctoproject.org/pub/releases/yocto-2.3_M3/machines/qemu/qemux86/core-image-minimal-qemux86.testdata.json
    
    
2- bitbake core-image-minimal-c testimage

EXPECTED RESULT:
Test are executed correctly 

ACTUAL RESULT:
Traceback (most recent call last):
|   File "/home/jgperezc/Sandbox/poky/scripts/runqemu", line 1235, in <module>
|     ret = main()
|   File "/home/jgperezc/Sandbox/poky/scripts/runqemu", line 1222, in main
|     config.check_and_set()
|   File "/home/jgperezc/Sandbox/poky/scripts/runqemu", line 631, in check_and_set
|     self.validate_paths()
|   File "/home/jgperezc/Sandbox/poky/scripts/runqemu", line 735, in validate_paths
|     tmpdir = self.get('OE_TMPDIR', None)
| TypeError: get() takes 2 positional arguments but 3 were given


Full log attached:
Comment 1 brian avery 2017-04-18 20:26:39 UTC
This is because we now check the information in <image>-qemuboot.conf. If the directories in the qemuboot configuraton file do not exist,  we run bitbake -e to get the actual directory. Since this is already running in a bitbake context, the second bitbake -e fails, then everything fails.There are a couple of workarounds here:
1) sed the 2 files to have the right path for example:
$ sed -i 's:/home/pokybuild/yocto-autobuilder/yocto-worker/nightly-x86/build/build:/big/src/myBugs/RUN/JOSE/build:g' core-image-minimal-qemux86.testdata.json
$ sed -i 's:/home/pokybuild/yocto-autobuilder/yocto-worker/nightly-x86/build/build:/big/src/myBugs/RUN/JOSE/build:g' core-image-minimal-qemux86.qemuboot.conf
This works and should cause no side effects
2) sed the 2 files to have a relative path:
$ sed -i 's:/home/.*/tmp/:tmp/:g' core-image-minimal-qemux86.testdata.json
$ sed -i 's:/home/.*/tmp/:tmp/:g' core-image-minimal-qemux86.qemuboot.conf  
This works for the default testimage test suite but I do not know if it will work or fail for the smart and rpm tests
----
It would be worth knowing if #2 works because, imho, it would be best if Bug 10964 were solved by keeping all the paths relative so that we could move tmp directories easily.
---
A separate alternative (and why I am adding RP, is if we allow bitbake -e to run multiple times against a directory as it is (afaik) a read only operation.
Comment 2 Richard Purdie 2017-04-18 20:31:14 UTC
We do not allow multiple concurrent bitbake -e against a directory since it is not a read only operation (touches cache files).
Comment 3 brian avery 2017-04-18 20:47:59 UTC
Setting to Needinfo so that Jose can do 
$ sed -i 's:/home/.*/tmp/:tmp/:g' core-image-minimal-qemux86.testdata.json
$ sed -i 's:/home/.*/tmp/:tmp/:g' core-image-minimal-qemux86.qemuboot.conf  
and see if the full qa test suites are ok with relative paths. If so, that is probably the better solution to this issue.
Comment 4 Juan Ramos 2017-04-18 22:51:38 UTC
tested using the next steps:

1. Download artifacts from AB
                 
2. put them in build/tmp/deploy/images/qemux86

3. run below comands under tmp/deploy/images/qemux86
   $ sed -i 's:/home/.*/tmp/:tmp/:g' core-image-minimal-qemux86.testdata.json
   $ sed -i 's:/home/.*/tmp/:tmp/:g' core-image-minimal-qemux86.qemuboot.conf  
               
4. Add INHERIT += "testimage" to the conf/local.conf

5. bitbake core-image-minimal-c testimage

ERROR: core-image-minimal-1.0-r0 do_testimage: Couldn't get ip from qemu command line and runqemu output!

Full log attached
Comment 5 brian avery 2017-04-18 22:57:26 UTC
attachment from ip failure?

also please attach the sedded conf files.
Comment 6 Jose Perez C 2017-04-18 23:08:30 UTC
(In reply to comment #5)
> attachment from ip failure?
> 
> also please attach the sedded conf files.

seems like  this one was tested in a different commit... please test it on correct commit
Comment 7 Juan Ramos 2017-04-18 23:29:08 UTC
Created attachment 3736 [details]
qemu error.log
Comment 8 brian avery 2017-04-19 00:39:21 UTC
runqemu - ERROR - Failed to setup tap device. Run runqemu-gen-tapdevs to manually create.

Looks like it was having trouble making the tap device.
Do you have prepopulated tap devs on your host?
Do you have sudo set up to be able to make taps?
---
This should work, does it?
$ runqemu tmp/deploy/images/qemux86/bzImage-qemux86.bin tmp/deploy/images/qemux86/core-image-minimal-qemux86.ext4 nographic 
regardless of whether you have sedded the files.
Comment 9 Jose Perez C 2017-04-19 13:28:40 UTC
Created attachment 3737 [details]
log with sed command

I was able to execute the full testimage suite with the "sed" changes.. there were 2 test cases failed on DNF test cases, all others were executed correcly and PASS. Full log attached 

-  FAIL: test_dnf_install_dependency (dnf.DnfRepoTest)
-  FAIL: test_dnf_reinstall (dnf.DnfRepoTest)
Comment 10 Jose Perez C 2017-04-19 13:33:48 UTC
Back to Brian.. Also adding Alexander to take a lock on the DNF failures.. seems to be related to the run-postinst package instead of current "sed: changes propossed by Brian.
Comment 11 Ross Burton 2017-04-19 13:53:13 UTC
To my eyes those tests are failing because they expect a populated tmp/deploy, which won't exist if you've just downloaded images.  Maybe the test should be bitbaking the bits it wants?
Comment 12 Jose Perez C 2017-04-19 14:06:57 UTC
(In reply to comment #11)
> To my eyes those tests are failing because they expect a populated
> tmp/deploy, which won't exist if you've just downloaded images.  Maybe the
> test should be bitbaking the bits it wants?

Before executing -c testimage I executed:  

 bitbake rpm psplash run-postinsts
Comment 13 brian avery 2017-04-19 19:48:20 UTC
https://patchwork.openembedded.org/series/6422/
Comment 14 Alexander Kanavin 2017-04-20 11:48:35 UTC
(In reply to comment #11)
> To my eyes those tests are failing because they expect a populated
> tmp/deploy, which won't exist if you've just downloaded images.  Maybe the
> test should be bitbaking the bits it wants?

No, it's something else:

======
AssertionError: 1 != 0 : dnf --repofrompath=oe-testimage-repo-noarch,http://192.168.7.7:46099/noarch --repofrompath=oe-testimage-repo-i586,http://192.168.7.7:46099/i586 --repofrompath=oe-testimage-repo-qemux86,http://192.168.7.7:46099/qemux86 --nogpgcheck install -y run-postinsts-dev
Added oe-testimage-repo-noarch repo from http://192.168.7.7:46099/noarch
Added oe-testimage-repo-i586 repo from http://192.168.7.7:46099/i586
Added oe-testimage-repo-qemux86 repo from http://192.168.7.7:46099/qemux86
Failed to synchronize cache for repo 'oe-testimage-repo-noarch', disabling.
Failed to synchronize cache for repo 'oe-testimage-repo-i586', disabling.
Failed to synchronize cache for repo 'oe-testimage-repo-qemux86', disabling.
========

Note that this comes after a successful test which does a very similar installation attempt. Weird timing issue? I'll try to see if I can reproduce the issue locally, but it has never been seen on the autobuilder, so maybe it's something about Jose's machine :-/


Jose, can you run this several times, and check if the failure occurs reliably in the same place, or if it appears/disappears/moves around?
Comment 15 Alexander Kanavin 2017-04-20 11:58:34 UTC
The image where the failures occured is core-image-sato-sdk, not core-image-minimal, so I'd need a precise download link for it though.
Comment 16 Jose Perez C 2017-04-20 14:44:10 UTC
(In reply to comment #15)
> The image where the failures occured is core-image-sato-sdk, not
> core-image-minimal, so I'd need a precise download link for it though.

Is the same place just s/minimal/sato-sdk/g on the links already mentioned
Comment 17 brian avery 2017-04-20 17:46:36 UTC
One additional patch needed for missing dependency: https://patchwork.openembedded.org/series/6443/

Right now, the relative path fix is in rc1. The script at the end of this comment will let you grab the artefacts easily.  

Jose, can you verify that this fixes all the issues except the one you are working with Alex?  If it does, reassign to Alex as my part is done, i think.

Useful (maybe) info:
To workaround the lack of the above patch in rc1, do:
$ echo 'INHERIT+="testimage"' >> conf/local.conf
$ bitbake qemu-helper-native
$ mkdir -p tmp/deploy/images/qemux86/
$ pushd tmp/deploy/images/qemux86/
$ grab.sh
$ popd
$ bitbake core-image-minimal -c testimage

---- grab.sh ----
#!/bin/bash
BASE="http://autobuilder.yoctoproject.org/pub/releases/yocto-2.3.rc1/machines/qemu/qemux86"

LIST="core-image-minimal-qemux86.ext4 core-image-minimal-qemux86.manifest core-image-minimal-qemux86.manifest core-image-minimal-qemux86.qemuboot.conf core-image-minimal-qemux86.testdata.json bzImage-qemux86.bin"
#LIST="core-image-sato-sdk-qemux86.ext4 core-image-sato-sdk-qemux86.manifest core-image-sato-sdk-qemux86.manifest core-image-sato-sdk-qemux86.qemuboot.conf core-image-sato-sdk-qemux86.testdata.json bzImage-qemux86.bin"

for l in $LIST; do
    if [ -e $l ]; then
        echo "already have $l , silly"
    else
        echo "Grabbing $l"
        wget $BASE/$l
    fi
done
------ grab.sh end ----
Comment 18 Jose Perez C 2017-04-20 21:54:37 UTC
(In reply to comment #17)
> One additional patch needed for missing dependency:
> https://patchwork.openembedded.org/series/6443/
> 
> Right now, the relative path fix is in rc1. The script at the end of this
> comment will let you grab the artefacts easily.  
> 
> Jose, can you verify that this fixes all the issues except the one you are
> working with Alex?  If it does, reassign to Alex as my part is done, i think.
> 
> Useful (maybe) info:
> To workaround the lack of the above patch in rc1, do:
> $ echo 'INHERIT+="testimage"' >> conf/local.conf
> $ bitbake qemu-helper-native
> $ mkdir -p tmp/deploy/images/qemux86/
> $ pushd tmp/deploy/images/qemux86/
> $ grab.sh
> $ popd
also bitbake rpm run-postinsts is needed 
> $ bitbake core-image-minimal -c testimage
> 
> ---- grab.sh ----
> #!/bin/bash
> BASE="http://autobuilder.yoctoproject.org/pub/releases/yocto-2.3.rc1/
> machines/qemu/qemux86"
> 
> LIST="core-image-minimal-qemux86.ext4 core-image-minimal-qemux86.manifest
> core-image-minimal-qemux86.manifest core-image-minimal-qemux86.qemuboot.conf
> core-image-minimal-qemux86.testdata.json bzImage-qemux86.bin"
> #LIST="core-image-sato-sdk-qemux86.ext4 core-image-sato-sdk-qemux86.manifest
> core-image-sato-sdk-qemux86.manifest
> core-image-sato-sdk-qemux86.qemuboot.conf
> core-image-sato-sdk-qemux86.testdata.json bzImage-qemux86.bin"
> 
> for l in $LIST; do
>     if [ -e $l ]; then
>         echo "already have $l , silly"
>     else
>         echo "Grabbing $l"
>         wget $BASE/$l
>     fi
> done
> ------ grab.sh end ----

Juts run on core-image-sato-sdk-qemux86 fro 2.3 rc1 on my local machine and everything worked correctly. the DNF errors are just happening on the server will double check what is happening.
Comment 19 Jose Perez C 2017-04-20 21:56:23 UTC
(In reply to comment #14)
> (In reply to comment #11)
> > To my eyes those tests are failing because they expect a populated
> > tmp/deploy, which won't exist if you've just downloaded images.  Maybe the
> > test should be bitbaking the bits it wants?
> 
> No, it's something else:
> 
> ======
> AssertionError: 1 != 0 : dnf
> --repofrompath=oe-testimage-repo-noarch,http://192.168.7.7:46099/noarch
> --repofrompath=oe-testimage-repo-i586,http://192.168.7.7:46099/i586
> --repofrompath=oe-testimage-repo-qemux86,http://192.168.7.7:46099/qemux86
> --nogpgcheck install -y run-postinsts-dev
> Added oe-testimage-repo-noarch repo from http://192.168.7.7:46099/noarch
> Added oe-testimage-repo-i586 repo from http://192.168.7.7:46099/i586
> Added oe-testimage-repo-qemux86 repo from http://192.168.7.7:46099/qemux86
> Failed to synchronize cache for repo 'oe-testimage-repo-noarch', disabling.
> Failed to synchronize cache for repo 'oe-testimage-repo-i586', disabling.
> Failed to synchronize cache for repo 'oe-testimage-repo-qemux86', disabling.
> ========
> 
> Note that this comes after a successful test which does a very similar
> installation attempt. Weird timing issue? I'll try to see if I can reproduce
> the issue locally, but it has never been seen on the autobuilder, so maybe
> it's something about Jose's machine :-/
> 
> 
> Jose, can you run this several times, and check if the failure occurs
> reliably in the same place, or if it appears/disappears/moves around?
This is just happening on the server (OpenSuse 42.2) on my local machine (Debian 8) I did not have any issues when running them.. will check what is happening on the server.
Comment 20 Alexander Kanavin 2017-04-21 05:48:44 UTC
(In reply to comment #19)
> > Jose, can you run this several times, and check if the failure occurs
> > reliably in the same place, or if it appears/disappears/moves around?
> This is just happening on the server (OpenSuse 42.2) on my local machine
> (Debian 8) I did not have any issues when running them.. will check what is
> happening on the server.

By the way, dnf writes additional logs to /var/log/dnf*log, so those might give additional clues.
Comment 21 brian avery 2017-04-21 14:54:05 UTC
Do we understand the Opensuse failure yet? Is it opensuse or that server?
Comment 22 Jose Perez C 2017-04-21 15:05:23 UTC
(In reply to comment #21)
> Do we understand the Opensuse failure yet? Is it opensuse or that server?

Not yet analyzed. but seems like the server settings itself, will check what is happening there.
Comment 23 Jose Perez C 2017-04-21 22:13:33 UTC
(In reply to comment #21)
> Do we understand the Opensuse failure yet? Is it opensuse or that server?

According to the log there is a timeout on the repo created: 

13:10:55 check_finished_transfer_status: Serious error - Curl code (28): Timeout was reached for <url>

already running on GDC AB openSuse 42.2 to check the results there
Comment 24 Alexander Kanavin 2017-04-24 11:47:43 UTC
=====
13:40:10 check_finished_transfer_status: Serious error - Curl code (28): Timeout was reached for http://10.54.69.60:41606/qemux86/repodata/repomd.xml [Connection timed out after 120001 milliseconds]
=====

Is 10.54.69.60 a valid IP address for the host (where the http server that is serving the repositories for the test is run)? Can this ip address be reached from within qemu?
Comment 25 Jose Perez C 2017-04-24 13:06:43 UTC
(In reply to comment #21)
> Do we understand the Opensuse failure yet? Is it opensuse or that server?

I ran it on opensuse 42.2 on local AB and everything was executed correclty.
Comment 26 Jose Perez C 2017-04-24 14:19:41 UTC
(In reply to comment #24)
> =====
> 13:40:10 check_finished_transfer_status: Serious error - Curl code (28):
> Timeout was reached for http://10.54.69.60:41606/qemux86/repodata/repomd.xml
> [Connection timed out after 120001 milliseconds]
> =====
> 
> Is 10.54.69.60 a valid IP address for the host (where the http server that
> is serving the repositories for the test is run)? Can this ip address be
> reached from within qemu?

Yes is valid IP, I ping 10.54.69.60 from within qemu with valid response.
Comment 27 Jose Perez C 2017-04-24 19:58:35 UTC
(In reply to comment #24)
> =====
> 13:40:10 check_finished_transfer_status: Serious error - Curl code (28):
> Timeout was reached for http://10.54.69.60:41606/qemux86/repodata/repomd.xml
> [Connection timed out after 120001 milliseconds]
> =====
> 
> Is 10.54.69.60 a valid IP address for the host (where the http server that
> is serving the repositories for the test is run)? Can this ip address be
> reached from within qemu?

Alex 

I was researching and there was a similar bug reported for DNF + librepo giving this error when there are long timeouts: 

https://bugzilla.redhat.com/show_bug.cgi?id=1272977
Comment 28 Alexander Kanavin 2017-04-25 07:10:36 UTC
(In reply to comment #26)
> (In reply to comment #24)
> > =====
> > 13:40:10 check_finished_transfer_status: Serious error - Curl code (28):
> > Timeout was reached for http://10.54.69.60:41606/qemux86/repodata/repomd.xml
> > [Connection timed out after 120001 milliseconds]
> > =====
> > 
> > Is 10.54.69.60 a valid IP address for the host (where the http server that
> > is serving the repositories for the test is run)? Can this ip address be
> > reached from within qemu?
> 
> Yes is valid IP, I ping 10.54.69.60 from within qemu with valid response.

Are you able to also access the repository server from within qemu from command line? For example, try to download the above file using curl or qemu (the port would vary of course).
Comment 29 Alexander Kanavin 2017-04-25 07:12:31 UTC
(In reply to comment #27)

> Alex 
> 
> I was researching and there was a similar bug reported for DNF + librepo
> giving this error when there are long timeouts: 
> 
> https://bugzilla.redhat.com/show_bug.cgi?id=1272977

Yeah, but this is when the server is unresponsive. Our server is supposed to work :)
Comment 30 Jose Perez C 2017-04-25 13:10:06 UTC
(In reply to comment #29)
> (In reply to comment #27)
> 
> > Alex 
> > 
> > I was researching and there was a similar bug reported for DNF + librepo
> > giving this error when there are long timeouts: 
> > 
> > https://bugzilla.redhat.com/show_bug.cgi?id=1272977
> 
> Yeah, but this is when the server is unresponsive. Our server is supposed to
> work :)

Yes but for some reason (not known yet) on that server is taking reaching the timeout limit. will investigate more on this. Besides of this this bug can be set as resolved as the original problem is already solved.
Comment 31 brian avery 2017-04-25 17:01:46 UTC
Jose, Can you file a separate bug for tracking the dnf failure on the server?

I'll close out this one since the original problem was solved.