Bug 1007 - There should be a config option for "limited use kernels"
Summary: There should be a config option for "limited use kernels"
Status: RESOLVED FIXED
Alias: None
Product: Kernel
Classification: Yocto Project Subprojects
Component: kernel-configuration (show other bugs)
Version: unspecified
Hardware: All Multiple
: High major
Target Milestone: 1.1 M4
Assignee: Bruce Ashfield
QA Contact:
URL:
Whiteboard: patches are merged to the BSPs. layer...
Depends on:
Blocks:
 
Reported: 2011-04-25 12:37 UTC by Dave Stewart
Modified: 2011-08-08 20:27 UTC (History)
8 users (show)

See Also:
OS type for building Yocto: ---
Type of Regression: ---
Verified:
Documentation change: ---


Attachments

Note You need to log in before you can comment on or make changes to this bug.
Description Dave Stewart 2011-04-25 12:37:35 UTC
Need a config option which makes it inconvenient or unattractive to use a kernel in an actual product.

I'm having various discussions with people who want to supply a kernel, but want to encourage developers to build their own with their product, rather than use the kernel with this option turned on.

By default, we would have this config option turned off.

Inconvenience in this case should equate to something like a nag screen warning you not to use this kernel for production, but go to www.yoctoproject.org and build your own. Maybe something like a 5 day time bomb which will shut down the kernel.
Comment 1 Bruce Ashfield 2011-04-25 20:00:26 UTC
I've got some ideas for this, have something similar kicking around. Will update the case later.
Comment 2 Bruce Ashfield 2011-05-31 13:16:00 UTC
Here are the initial thoughts on this. Please add comments as appropriate:

Time limited kernel images:                                                                               
---------------------------                                                                               
                                                                                                          
 - trigger: passed into the build as 'delta from build time'. Only set                                    
   in 'official' builds. Delta is in the format of hh:mm:ss                                               
                                                                                                          
 - the feature is multi-arch, and uses very simple mechanisms to be                                       
   portable and maintainable.                                                                             
                                                                                                          
 - option is a no-op if not enabled, disabled is the default                                              
                                                                                                          
 - base build time is stored in the text image, potentially just                                          
   reusing the already captured kernel build time. The delta is also                                      
   delta is stored in the text segment of the image.                                                      
                                                                                                          
      - the variable is defined in such a way that the location can                                       
        change for BSP specific hooks and non-volatile storage                                            
        options in the future.                                                                            
                                                                                                          
 - on each boot, a yocto banner and time limited warning is dumped.                                       
   After expiry, the kernel will not boot at all.                                                         
                                                                                                          
 - the BSP *must* have a RTC available, the RTC must be initialized.                                      
   If the RTC shows as reset (i.e. 1970), then the banner is followed                                     
   by an extra warning and delay. This is to encourage that on boot                                       
   the RTC be set.                                                                                        
                                                                                                          
      - If the RTC is not set, there is little we can do. See below                                       
        for 'items that are not addressed'                                                                
      - the timeout can be skipped via a boot parameter for truly                                         
        broken boards with no RTC      

Items that are not a concern or are not addressed by design:                                              
------------------------------------------------------------                                              
                                                                                                          
 - BSP specific hooks/requirements such as NVRAM, to allow updating                                       
   or boot counting.                                                                                      
      - this could be a future extension.                                                                 
      - the current implementation is largely arch/board generic                                          
                                                                                                          
 - The network is not a requirement, and hence external time servers                                      
   or contact cannot be required for the time validation.                                                 
                                                                                                          
 - the relative ease with which you can patch the image to disable                                        
   the check.                                                                                             
     - the image build size could be added to the delta via a md5                                         
       (or similar technique) to prevent post-build patching of                                           
       the image. But there's little value in this.
Comment 3 Darren Hart 2011-05-31 14:00:50 UTC
This seems to address the issue Dave described in the bug description. How common is it for the RTC to be unavailable?
Comment 4 Dave Stewart 2011-05-31 14:32:42 UTC
Remember the goal: make it easy to boot up a board quickly with an included kernel, but discourage its long-term use in a product without a rebuild.

If I understand Bruce's comment "trigger: passed into the build as 'delta from build time'" - your kernel will have an "expiration date" and won't boot beyond that date.

Although this might meet the goal, it requires that the RTC be up to date and not reset as noted in Bruce's comments. It also suggests that no matter what you set the expiration date to, you will have somebody who wants to use an old BSP and board with an expired kernel. This is fixable of course with a rebuild, but it seems like a source for complaints.

Why not make the trigger to be real time since boot time and do a shutdown?
Comment 5 Tom Zanussi 2011-05-31 14:42:46 UTC
I'm not sure basing it on build time will address the problem.  Unless I'm missing something, this scheme will render the image unbootable 5 days (or whatever the timeout value is) after the image was built, so e.g our BSP images will expire 5 days after we build them?

My understanding is that we want the user to be able to grab the image and try it 
out, but not run it for longer than say 5 days.

But the image itself should be usable forever.

The user could defeat it by rebooting every 4 days, but that's unattractive enough to render it practically unusable over the long term.

So I'm thinking that a delta against uptime would work, and wouldn't need RTC...
Comment 6 Bruce Ashfield 2011-05-31 14:55:06 UTC
I did't read the request as uptime, but I can see that as well. I'd actually do both. Tricking uptime is just as easy as from build time. A kernel shutdown after a timeout of uptime will get a different set of complaints. I suggest the uptime as option 1 and build time as the fallback. With a change to decrease the uptime restriction vs boot failure
Comment 7 Bruce Ashfield 2011-05-31 19:25:21 UTC
I replied earlier from my blackberry, so it wasn't as complete as I liked. The design
is essentially the same, with the addition of an uptime limiter (controller). But the 
nature of the uptime limitation is still a question for me, since the following items
need to be considered.

- We don't want to inject 'active' events into the kernel. Yes, this is a 'nag' of sorts,
   but having timers fire, checks during schedule(), decrementer interrupts, etc, all
   can mess with profiling, caching, powermanagement and any number of things
   that a 'test driver' may be interested in. Not to mention it can be simply
   annoying. But on the other hand, we aren't talking about a fine grained timeout,
   so this may end up being the right thing.

- I had read in some context numbers like 90 days, which is more of a past-the
   release date vs an uptime. 5 days is better if we are talking uptime.

- It might be worth adding the check and setting a flag that prevents the forking
   of new processes, or a cgroup controller that throws new processes into a 
   cage when the uptime expires. We also need to have a way to broadcast in an
   un-missable way. bootime gets you that, uptime expiration is harder.

- Putting a hard (but as I mentioned, avoidable) best before date from the build 
   would be secondary. This factor could be used to decrease the uptime timeout,
   since you really want to encourage use of a new release, not the older ones.

Everything in the original comment stays: the banner, the implementation, but 
the source of the timing (uptime) changes as the primary trigger.
Comment 8 Darren Hart 2011-06-03 11:51:00 UTC
(In reply to comment #7)
> - We don't want to inject 'active' events into the kernel. Yes, this is a 'nag' 
>   of sorts, but having timers fire, checks during schedule(), decrementer
>   interrupts, etc, all can mess with profiling, caching, powermanagement and any
>   number of things that a 'test driver' may be interested in. Not to mention it
>   can be simply annoying. But on the other hand, we aren't talking about a fine
>   grained timeout, so this may end up being the right thing.

Why not just crate a timer at boot - no polling to check the time, etc. Just set
a single timer and a handler that safely shuts down.
Comment 9 Julie Fleischer 2011-06-03 13:00:57 UTC
Request from the embedded hardware group at Intel:  Could you add a warning screen to the user every few hours (or however often is appropriate) so that the shutdown does not come as a surprise?
Comment 10 Mark Hatle 2011-06-03 13:50:47 UTC
Current request is approx 48 hour uptime, with nag message at approx 2 hr intervals.

(note even with that this should likely be configurable in the generic patch.)
Comment 11 Bruce Ashfield 2011-06-05 21:56:05 UTC
A single timer can definitely work. But is pretty easy to evade (again, not that
preventing evasion is the primary element here).

My issue with warning screens, and periodic updates is that *someone* needs
to logged in to see them. They are in the design, but as a runtime only check,
they are easy to miss. We of course will log a clear reason for the reboot that
can be seen in /var/log/messages .. but it will definitely still be a surprise.
Comment 12 Tom Zanussi 2011-06-06 07:42:03 UTC
So it seems we need 2 timers - one periodically expiring every 2 hours and one single-shot expiring after 48 hours uptime.

Bruce makes a good point - if the user is in X/Sato and we e.g. printk a warning every 2 hours nobody will see it since it only goes to the console.  Does that mean we need to also provide a graphical 'popup' for X/sato or something else that the user will see, every 2 hours (and then we need to remove it after another time interval, I guess, to avoid cluttering the screen with warning messages), if the user is in X (we'd also need to detect whether that's the case).  Even then, what if the screen is powered down - which means for the user to see it, the warning will need to wake up the screen.  Just some additional details to note, and that it's beginning to sounds like a lot of requirements for a trivial warning message.
Comment 13 Darren Hart 2011-06-06 09:02:28 UTC
(In reply to comment #12)
> So it seems we need 2 timers - one periodically expiring every 2 hours and one
> single-shot expiring after 48 hours uptime.
> 

Agreed

> Bruce makes a good point - if the user is in X/Sato and we e.g. printk a
> warning every 2 hours nobody will see it since it only goes to the console. 
> Does that mean we need to also provide a graphical 'popup' for X/sato or

I don't think this is appropriate for our intended targets. We aren't creating a system for consumer devices where we expect someone to be at the screen whenever the device is in use. I think the effort to create an X notification would be mostly a waste of time. Due diligence requires us to log the events and warnings, but I don't think we need to go beyond that.

We also need to keep this implementation as simple as possible to reduce long term maintenance.
Comment 14 Mark Hatle 2011-06-06 09:05:34 UTC
A simple kernel message seems like enough to me.  Even if someone is running X, the messages will go to the console if they are using it.. otherwise it should have been documented in some other way.

Note, the 2 hr and 48 hr are just examples of what one particular user wants.  If this is to be a generic feature then we need to keep that in mind.  

note, I think you are right, we need two timers.. a nag timer and a reboot timer.... perhaps even the nag message should be configurable?  I could easily see someone want to use a nag, but not the reboot.
Comment 15 Tom Zanussi 2011-06-06 09:30:38 UTC
I agree, we should be able to implement just the two timers, one periodic that just printk()s the contents of say a KERNEL_SELF_DESTRUCT_NAG_MSG every KERNEL_SELF_DESTRUCT_NAG_MSG_INTERVAL seconds and a one-shot that prints a self-destruct message after KERNEL_SELF_DESTRUCT_INTERVAL seconds.
Comment 16 Tom Zanussi 2011-06-06 09:31:47 UTC
(In reply to comment #15)
> I agree, we should be able to implement just the two timers, one periodic that
> just printk()s the contents of say a KERNEL_SELF_DESTRUCT_NAG_MSG every
> KERNEL_SELF_DESTRUCT_NAG_MSG_INTERVAL seconds and a one-shot that prints a
> self-destruct message after KERNEL_SELF_DESTRUCT_INTERVAL seconds.

Oh, and of course, self-destructs. ;-)
Comment 17 Tom Zanussi 2011-06-06 09:45:26 UTC
One problem with all this is that if we just reboot without warning, the user could lose data e.g. if we reboot before dirty pages in the page cache have been flushed to disk, the dirty pages will be lost, and the user not too happy.

Is the user expected to realize this possibility or do we need say so explicitly on install?
Comment 18 Tom Zanussi 2011-06-06 09:48:03 UTC
(In reply to comment #17)
> One problem with all this is that if we just reboot without warning, the user
> could lose data e.g. if we reboot before dirty pages in the page cache have
> been flushed to disk, the dirty pages will be lost, and the user not too happy.
> 
> Is the user expected to realize this possibility or do we need say so
> explicitly on install?

I mean, we aren't doing it 'without warning', we are printing the warning, but if they don't see it or don't realize that they shouldn't be massively writing to disk when the shutdown happens in 2 hours, etc...
Comment 19 Mark Hatle 2011-06-06 09:51:50 UTC
I was thinking it would simply call though the system shutdown/reboot...  So anything the kernel normally does to flush the disk cache, turn off devices, etc will happen... however, anything in userspace just goes away -- with the potential for data loss.

I assume we can do this, if not, then simply rebooting using the same approach as an OOPS is probably the second best -- and data loss will be expected.

(Either way, the option can cause data loss, and that has to be documented...)
Comment 20 Bruce Ashfield 2011-06-06 09:56:55 UTC
Longest bug comment stream ever. If we knew this was going to happen, we would 
have gone to email.

The data loss/surprise angle is the #1 reason why I suggested a boot time check, and 
still recommend that we keep boot time checks as the fallback option. Any shutdown
technique, warning or not, is going to risk data loss or worse corruption. 

True this is only on a default build, and a single rebuild removes this, but I for
one would be annoyed if I didn't see the warning and then bug'd/trap'd or rebooted
when I wasn't looking.
Comment 21 Song Liu 2011-06-06 12:02:14 UTC
Just would like to clarify the requirement (with ECG requests):

1. Default is to have this option off, so that a customized build won't be bothered by this, but special BSPs will have this on.
2. Having warning/reminder message once in a while about the shutdown/reboot/stop-use and the reminder to have a customized build.
3. Stop the use after a period of time (3 days?).
Comment 22 Bruce Ashfield 2011-08-08 20:27:28 UTC
code is in tree. ready for use.