Need a config option which makes it inconvenient or unattractive to use a kernel in an actual product. I'm having various discussions with people who want to supply a kernel, but want to encourage developers to build their own with their product, rather than use the kernel with this option turned on. By default, we would have this config option turned off. Inconvenience in this case should equate to something like a nag screen warning you not to use this kernel for production, but go to www.yoctoproject.org and build your own. Maybe something like a 5 day time bomb which will shut down the kernel.
I've got some ideas for this, have something similar kicking around. Will update the case later.
Here are the initial thoughts on this. Please add comments as appropriate: Time limited kernel images: --------------------------- - trigger: passed into the build as 'delta from build time'. Only set in 'official' builds. Delta is in the format of hh:mm:ss - the feature is multi-arch, and uses very simple mechanisms to be portable and maintainable. - option is a no-op if not enabled, disabled is the default - base build time is stored in the text image, potentially just reusing the already captured kernel build time. The delta is also delta is stored in the text segment of the image. - the variable is defined in such a way that the location can change for BSP specific hooks and non-volatile storage options in the future. - on each boot, a yocto banner and time limited warning is dumped. After expiry, the kernel will not boot at all. - the BSP *must* have a RTC available, the RTC must be initialized. If the RTC shows as reset (i.e. 1970), then the banner is followed by an extra warning and delay. This is to encourage that on boot the RTC be set. - If the RTC is not set, there is little we can do. See below for 'items that are not addressed' - the timeout can be skipped via a boot parameter for truly broken boards with no RTC Items that are not a concern or are not addressed by design: ------------------------------------------------------------ - BSP specific hooks/requirements such as NVRAM, to allow updating or boot counting. - this could be a future extension. - the current implementation is largely arch/board generic - The network is not a requirement, and hence external time servers or contact cannot be required for the time validation. - the relative ease with which you can patch the image to disable the check. - the image build size could be added to the delta via a md5 (or similar technique) to prevent post-build patching of the image. But there's little value in this.
This seems to address the issue Dave described in the bug description. How common is it for the RTC to be unavailable?
Remember the goal: make it easy to boot up a board quickly with an included kernel, but discourage its long-term use in a product without a rebuild. If I understand Bruce's comment "trigger: passed into the build as 'delta from build time'" - your kernel will have an "expiration date" and won't boot beyond that date. Although this might meet the goal, it requires that the RTC be up to date and not reset as noted in Bruce's comments. It also suggests that no matter what you set the expiration date to, you will have somebody who wants to use an old BSP and board with an expired kernel. This is fixable of course with a rebuild, but it seems like a source for complaints. Why not make the trigger to be real time since boot time and do a shutdown?
I'm not sure basing it on build time will address the problem. Unless I'm missing something, this scheme will render the image unbootable 5 days (or whatever the timeout value is) after the image was built, so e.g our BSP images will expire 5 days after we build them? My understanding is that we want the user to be able to grab the image and try it out, but not run it for longer than say 5 days. But the image itself should be usable forever. The user could defeat it by rebooting every 4 days, but that's unattractive enough to render it practically unusable over the long term. So I'm thinking that a delta against uptime would work, and wouldn't need RTC...
I did't read the request as uptime, but I can see that as well. I'd actually do both. Tricking uptime is just as easy as from build time. A kernel shutdown after a timeout of uptime will get a different set of complaints. I suggest the uptime as option 1 and build time as the fallback. With a change to decrease the uptime restriction vs boot failure
I replied earlier from my blackberry, so it wasn't as complete as I liked. The design is essentially the same, with the addition of an uptime limiter (controller). But the nature of the uptime limitation is still a question for me, since the following items need to be considered. - We don't want to inject 'active' events into the kernel. Yes, this is a 'nag' of sorts, but having timers fire, checks during schedule(), decrementer interrupts, etc, all can mess with profiling, caching, powermanagement and any number of things that a 'test driver' may be interested in. Not to mention it can be simply annoying. But on the other hand, we aren't talking about a fine grained timeout, so this may end up being the right thing. - I had read in some context numbers like 90 days, which is more of a past-the release date vs an uptime. 5 days is better if we are talking uptime. - It might be worth adding the check and setting a flag that prevents the forking of new processes, or a cgroup controller that throws new processes into a cage when the uptime expires. We also need to have a way to broadcast in an un-missable way. bootime gets you that, uptime expiration is harder. - Putting a hard (but as I mentioned, avoidable) best before date from the build would be secondary. This factor could be used to decrease the uptime timeout, since you really want to encourage use of a new release, not the older ones. Everything in the original comment stays: the banner, the implementation, but the source of the timing (uptime) changes as the primary trigger.
(In reply to comment #7) > - We don't want to inject 'active' events into the kernel. Yes, this is a 'nag' > of sorts, but having timers fire, checks during schedule(), decrementer > interrupts, etc, all can mess with profiling, caching, powermanagement and any > number of things that a 'test driver' may be interested in. Not to mention it > can be simply annoying. But on the other hand, we aren't talking about a fine > grained timeout, so this may end up being the right thing. Why not just crate a timer at boot - no polling to check the time, etc. Just set a single timer and a handler that safely shuts down.
Request from the embedded hardware group at Intel: Could you add a warning screen to the user every few hours (or however often is appropriate) so that the shutdown does not come as a surprise?
Current request is approx 48 hour uptime, with nag message at approx 2 hr intervals. (note even with that this should likely be configurable in the generic patch.)
A single timer can definitely work. But is pretty easy to evade (again, not that preventing evasion is the primary element here). My issue with warning screens, and periodic updates is that *someone* needs to logged in to see them. They are in the design, but as a runtime only check, they are easy to miss. We of course will log a clear reason for the reboot that can be seen in /var/log/messages .. but it will definitely still be a surprise.
So it seems we need 2 timers - one periodically expiring every 2 hours and one single-shot expiring after 48 hours uptime. Bruce makes a good point - if the user is in X/Sato and we e.g. printk a warning every 2 hours nobody will see it since it only goes to the console. Does that mean we need to also provide a graphical 'popup' for X/sato or something else that the user will see, every 2 hours (and then we need to remove it after another time interval, I guess, to avoid cluttering the screen with warning messages), if the user is in X (we'd also need to detect whether that's the case). Even then, what if the screen is powered down - which means for the user to see it, the warning will need to wake up the screen. Just some additional details to note, and that it's beginning to sounds like a lot of requirements for a trivial warning message.
(In reply to comment #12) > So it seems we need 2 timers - one periodically expiring every 2 hours and one > single-shot expiring after 48 hours uptime. > Agreed > Bruce makes a good point - if the user is in X/Sato and we e.g. printk a > warning every 2 hours nobody will see it since it only goes to the console. > Does that mean we need to also provide a graphical 'popup' for X/sato or I don't think this is appropriate for our intended targets. We aren't creating a system for consumer devices where we expect someone to be at the screen whenever the device is in use. I think the effort to create an X notification would be mostly a waste of time. Due diligence requires us to log the events and warnings, but I don't think we need to go beyond that. We also need to keep this implementation as simple as possible to reduce long term maintenance.
A simple kernel message seems like enough to me. Even if someone is running X, the messages will go to the console if they are using it.. otherwise it should have been documented in some other way. Note, the 2 hr and 48 hr are just examples of what one particular user wants. If this is to be a generic feature then we need to keep that in mind. note, I think you are right, we need two timers.. a nag timer and a reboot timer.... perhaps even the nag message should be configurable? I could easily see someone want to use a nag, but not the reboot.
I agree, we should be able to implement just the two timers, one periodic that just printk()s the contents of say a KERNEL_SELF_DESTRUCT_NAG_MSG every KERNEL_SELF_DESTRUCT_NAG_MSG_INTERVAL seconds and a one-shot that prints a self-destruct message after KERNEL_SELF_DESTRUCT_INTERVAL seconds.
(In reply to comment #15) > I agree, we should be able to implement just the two timers, one periodic that > just printk()s the contents of say a KERNEL_SELF_DESTRUCT_NAG_MSG every > KERNEL_SELF_DESTRUCT_NAG_MSG_INTERVAL seconds and a one-shot that prints a > self-destruct message after KERNEL_SELF_DESTRUCT_INTERVAL seconds. Oh, and of course, self-destructs. ;-)
One problem with all this is that if we just reboot without warning, the user could lose data e.g. if we reboot before dirty pages in the page cache have been flushed to disk, the dirty pages will be lost, and the user not too happy. Is the user expected to realize this possibility or do we need say so explicitly on install?
(In reply to comment #17) > One problem with all this is that if we just reboot without warning, the user > could lose data e.g. if we reboot before dirty pages in the page cache have > been flushed to disk, the dirty pages will be lost, and the user not too happy. > > Is the user expected to realize this possibility or do we need say so > explicitly on install? I mean, we aren't doing it 'without warning', we are printing the warning, but if they don't see it or don't realize that they shouldn't be massively writing to disk when the shutdown happens in 2 hours, etc...
I was thinking it would simply call though the system shutdown/reboot... So anything the kernel normally does to flush the disk cache, turn off devices, etc will happen... however, anything in userspace just goes away -- with the potential for data loss. I assume we can do this, if not, then simply rebooting using the same approach as an OOPS is probably the second best -- and data loss will be expected. (Either way, the option can cause data loss, and that has to be documented...)
Longest bug comment stream ever. If we knew this was going to happen, we would have gone to email. The data loss/surprise angle is the #1 reason why I suggested a boot time check, and still recommend that we keep boot time checks as the fallback option. Any shutdown technique, warning or not, is going to risk data loss or worse corruption. True this is only on a default build, and a single rebuild removes this, but I for one would be annoyed if I didn't see the warning and then bug'd/trap'd or rebooted when I wasn't looking.
Just would like to clarify the requirement (with ECG requests): 1. Default is to have this option off, so that a customized build won't be bothered by this, but special BSPs will have this on. 2. Having warning/reminder message once in a while about the shutdown/reboot/stop-use and the reminder to have a customized build. 3. Stop the use after a period of time (3 days?).
code is in tree. ready for use.