Bug 1761

Summary: 30s timeout on SQLite connections seems arbitrarily large
Product: [Build System, Metadata & Runtime] BitBake Reporter: Joshua Lock - Disabled <josh>
Component: bitbakeAssignee: Lianhao Lu <lianhao.lu>
Status: VERIFIED FIXED QA Contact:
Severity: normal    
Priority: Medium CC: elizabeth.flanagan, poky.bs.watcher, poky.watcher, richard.purdie, sgw, shane.wang
Version: 1.2   
Target Milestone: 1.2   
Hardware: x86   
OS: Multiple   
Whiteboard:
OS type for building Yocto: --- Type of Regression: ---
Verified: Documentation change: ---
Attachments:
Description Flags
mimic bitbake to reproduce the bug none

Description Joshua Lock - Disabled 2011-11-08 12:54:38 UTC
... but when we had it lower (5s) we saw a lot of "OperationalError: database is locked" exceptions on first run.

There's certainly an underlying issue here with a theory that there's an underlying issue in the python->pysqlite->sqlite stack that we trigger with our initial burst of writes to the database.

The 5s timeout ensured the OperationalError exception was raised on each first run on the autobuilder infrastructure, however the issue has proven much more difficult to reproduce on a variety of local setups.

It'd be great to work with upstream on this issue so that we don't have this arbitrarily large timeout value.
Comment 1 Beth Flanagan 2011-11-15 10:52:00 UTC
I think one of the reasons we're seeing this now on the autobuilder infrastructure is due to the addition of an extra build slave per autobuilder. That has slowed things down just enough to expose this issue (which we see in both bernard and master. It is most likely in edison as well.
Comment 2 Richard Purdie 2011-12-18 03:39:10 UTC
We're starting to see this more and more on the autobuilders, e.g. http://autobuilder.pokylinux.org:8010/builders/nightly-x86/builds/286/steps/shell_67/logs/stdio. Reassigning the priority to a medium bug.
Comment 3 Shane Wang 2011-12-18 22:43:19 UTC
Assign to Lianhao to help on it
Comment 4 Lianhao Lu 2011-12-19 02:11:23 UTC
Created attachment 301 [details]
mimic bitbake to reproduce the bug

Using the attached python code (1761.py), we can mimic what happened in bitbake and reproduce this bug easily both on Ubuntu 10.04_x86_64 and openSuse 11.04_x86_64.

sqlite is not designed to work with large amount of concurrent connections at the same time. However, we need to figure out a way to avoid the burst write operations and open connection operations at the same time.
Comment 6 Lianhao Lu 2012-01-18 17:11:22 UTC
Fixed in the following commit by reconnecting in the exception handler to leverage the retry count, and also drop timeout to 5 seconds.

http://git.yoctoproject.org/cgit/cgit.cgi/poky/commit/?id=1fedd166b71a4000116fcc9bd993cb23a7899adb
Comment 7 Joshua Lock - Disabled 2012-04-17 23:33:46 UTC
Verify we don't see the OperationalError issues anymore *and* have a sane timeout