| Summary: | First boot postinst actions | ||
|---|---|---|---|
| Product: | [Documentation] Development Manual | Reporter: | Xabier <xabier.marquiegui> |
| Component: | development | Assignee: | Xabier <xabier.marquiegui> |
| Status: | RESOLVED NOTABUG | QA Contact: | |
| Severity: | normal | ||
| Priority: | Undecided | CC: | alex.kanavin, bluelightning, srifenbark |
| Version: | unspecified | ||
| Target Milestone: | --- | ||
| Hardware: | All | ||
| OS: | Multiple | ||
| Whiteboard: | 13 Feb 2017: Set the doc flag to "done." | ||
| OS type for building Yocto: | --- | Type of Regression: | --- |
| Verified: | Documentation change: | No (bug/feature does not impact docs) | |
I'm sorry, I should also mention that the code mentioned above should be included in the pkg_postinst_${PN}_prepend() function when inheriting from systemd.
The things that systemd appends to existing postinst scripts are designed so that they can run both during cross-install and on first boot. The existing solution will simply postpone everything to first boot - where your specific action will run first, then the systemd actions. Your solution has the effect of making the systemd specific things run at cross-install. (and then they will run again at first boot after your custom action). What is the improvement here, why is this better or necessary? Well... I have dedicated a few minutes today to analyzing the problem, and I have a few more observations to throw into this matter.
I might be looking at something the wrong way, but let me describe what I am seeing.
Let's start with a description of the components I am using for my tests:
yocto 2.1 krogoth
systemd 229
Linux amos 4.8.6-fslc+g8d6e2d5 #1 SMP PREEMPT Mon Dec 19 17:12:40 UTC 2016 armv7l GNU/Linux
Test A: Following yocto development manual, I write a recipe with the following content:
...
inherit allarch systemd
...
SYSTEMD_SERVICE_${PN} = "myservice.service"
...
pkg_postinst_${PN}() {
#!/bin/sh
if [ x"$D" = "x" ] ; then
# Native postinst code
echo "HELLO THERE!"
else
exit 1
fi
}
RESULT: System gets stuck on boot showing this message:
[*** ] (1 of 2) A start job is running for... postinsts (2min 12s / no limit)
Test B: Using trap, and adding exit 0 on the other condition (I forgot to put that in my original message as well), my recipe ends up looking like this:
...
inherit allarch systemd
...
SYSTEMD_SERVICE_${PN} = "myservice.service"
...
pkg_postinst_${PN}() {
#!/bin/sh
if [ x"$D" = "x" ] ; then
# Native postinst code
echo "HELLO THERE!"
exit 0
else
trap "exit 1" EXIT
fi
}
RESULT: System boots just fine, and myservice.service is correctly enabled.
With the added exit 0 systemd specific things only run at compile time.
I have also observed that using pkg_postinst_PACKAGENAME_append, or _prepend has no effect on the relative position of my code. Systemd always positions its postinst code after my recipe“s code.
Hi, I will monitor this and see if your solution is final. When so, I will update the manual to reflect the change. Scott The issue here is that your service gets enabled by systemd correctly during cross-install on the host machine, but for some reason the same enabling sequence gets stuck if it gets executed on the target. I still believe there's nothing to be fixed in the manual.
Specifically:
a)
if [ x"$D" = "x" ] ; then
# Native postinst code
echo "HELLO THERE!"
else
exit 1
fi
exit 1 means defer the whole script to first boot, where it gets stuck on execution. You need to investigate this situation further. You can find the actual script that is executed on first boot in /etc/rpm-postinsts and work your way from there.
b)
if [ x"$D" = "x" ] ; then
# Native postinst code
echo "HELLO THERE!"
exit 0
else
trap "exit 1" EXIT
fi
The trap line means "execute 'exit 1' when the script finishes". This really means the script continues to execute during cross-install beyond the 'trap' line, and so systemd service gets enabled at that time. The script is also deferred to first boot, but is a no-op, because you put an 'exit 0' in there.
What happens if you remove the pkg_postinst() segment altogether from the recipe?
I am using opkg as the package manager.
( #v and #^ are just for better readability, I have not placed those lines in my code )
With this on my recipe (almost like b),
#vvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvv
pkg_postinst_${PN}() {
#!/bin/sh
if [ x"$D" = "x" ] ; then
# Native postinst code
echo "HELLO THERE!"
exit 0
else
trap "exit 1" EXIT
fi
}
#^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
, I get this in /var/lib/opkg/info/mypackage.postinst:
#vvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvv
#!/bin/sh
if [ x"$D" = "x" ] ; then
# Native postinst code
echo "HELLO THERE!"
exit 0
else
trap "exit 1" EXIT
fi
OPTS=""
if [ -n "$D" ]; then
OPTS="--root=$D"
fi
if type systemctl >/dev/null 2>/dev/null; then
systemctl $OPTS enable myservice.service
if [ -z "$D" -a "enable" = "enable" ]; then
systemctl restart myservice.service
fi
fi
#^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
, and the system boots just fine.
If I remove pkg_postinst_${PN}() from the recipe I get this in /var/lib/opkg/info/mypackage.postinst:
#vvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvv
#!/bin/sh
OPTS=""
if [ -n "$D" ]; then
OPTS="--root=$D"
fi
if type systemctl >/dev/null 2>/dev/null; then
systemctl $OPTS enable myservice.service
if [ -z "$D" -a "enable" = "enable" ]; then
systemctl restart myservice.service
fi
fi
#^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
, and the system boots just fine.
Now type these two in the command line on the target: systemctl enable myservice.service systemctl restart myservice.service I strongly suspect one of them will hang. I have issued both commands, and they both have returned in a few miliseconds, with a return value of 0. Some more information about my system: BusyBox v1.24.1 (2016-12-12 20:05:45 UTC) multi-call binary. opkg version 0.3.0 Then something else gets in the way during boot - the issue is that executing the postinst script on first boot is causing a freeze somewhere, and your suggested fix merely moves the execution to cross-install time. Please do investigate that. I am also new to systemd, but here's my guess: myservice does live in a custom target mytarget that comes after multi-user target. And my system's default target IS mytarget. Is it possible that postinst gets executed in multi-user target, and if you ask systemd to restart a service that lives in a target that has not been reached, it gets stuck until the system reaches that target? If that is true, if postinst is waiting for systemd to restart the service, and systemd is waiting for postinst to get to the next target, we have a deadlock situation caused by the way postinst is implemented. Does it make any sense? I have no idea either. You can add additional logging to the postinst bits in system class, and see where precisely it gets stuck. Also, I guess systemctl itself has a verbose or debugging mode which you should enable. You are right, what I have been proposing until now is not the right solution. If I were to install this package in a running system that didn't have myservice enabled and running, it would never get automatically enabled, and that is not right.
On the other hand, I'm pretty sure my theory is right. Boot gets stuck in:
systemctl restart myservice.service
If I manually generate a postinst script with the following content, everything works like a charm:
#vvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvvv
#!/bin/sh
if [ x"$D" = "x" ] ; then
# Native postinst code
echo "HELLO THERE!"
else
exit 1
fi
OPTS=""
if [ -n "$D" ]; then
OPTS="--root=$D"
fi
if type systemctl >/dev/null 2>/dev/null; then
systemctl $OPTS enable myservice.service
if [ -z "$D" -a "enable" = "enable" ]; then
systemctl try-restart myservice.service
fi
fi
#^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
So, changing systemctl restart with systemcl try-restart seems to be the right solution (which is code added by systemd bbclass)
How should I preceed? Should I close this bugzilla entry and open a new one for the systemd bbclass?
Yes please. I'm not exactly sure why 'restart' is even needed there - if a unit is enabled, I'd think system will start it when all pre conditions are satisfied. I just update the bug to mark it as resolved / not a bug, since it's not a documentation problem, but one in systemd.bbclass. By the way, the systemd issue is resolved in poky master 2965ccfcdb5a5aefb94e6ec8e24832471f47239f No need to file another bug. Yes, thank you. I have just seen that it's resolved in some branches. Is there any way I can ask for it to be backported to Yocto Krogoth? Yes, sure. Send a patch to the oe-core mailing list (look in list archives for how backports should be annotated). Hi, I know I don't own this bug but I see it was marked as NOT A BUG. I went ahead and set the doc flag to "no". Thanks, Scott That's OK, I didn't know I had to mark that flag. Thank you. :) |
I have found that in cases where you inherit from systemd, you cannot really use the pkg_postinst_${PN} function in your recipes if you want to run code on first boot I think it's because the postinst actions of systemd return 0 An easy fix is to change the documentation of the mega-manual where you say this: pkg_postinst_PACKAGENAME() { if [ x"$D" = "x" ]; then # Actions to carry out on the device go here else exit 1 fi } you could say this: pkg_postinst_PACKAGENAME() { if [ x"$D" = "x" ]; then # Actions to carry out on the device go here else trap "exit 1" EXIT fi } by using trap "exit 1" EXIT instead of using exit 1 you ensure that the code needed to run on first boot will get run