Created attachment 3363 [details] rpm test script Reproduction script attached. When the script is run multiple times concurrently the following error occurs: count 739: Name : bash Relocations: (not relocatable) Version : 4.3.30 Vendor: (none) Release : r0 Build Date: Tue Aug 16 15:42:24 2016 Install Date: Wed Aug 17 08:45:49 2016 Build Host: obsrwr Group : base/shell Source RPM: bash-4.3.30-r0.src.rpm Size : 1048412 License: GPLv3+ Signature : DSA/SHA1, Tue Aug 16 15:42:24 2016, Key ID 5d0492732d8fd654 Packager : Poky <poky@yoctoproject.org> URL : http://tiswww.case.edu/php/chet/bash/bashtop.html Summary : An sh-compatible command language interpreter Architecture: i586 Description : An sh-compatible command language interpreter. count 738: rpmdb: BDB0060 PANIC: fatal region error detected; run recovery ==> rpmdbe_event_notify(0x807aff8, PANIC(0), 0xbffeef8c) app_private (nil) error: db_init:../../rpm-5.4.15/rpmdb/db3.c:1206: dbenv->failchk(-30973): BDB0087 DB_RUNRECOVERY: Fatal error, run database recovery rpmdb: BDB1581 File handles still open at environment close rpmdb: BDB1582 Open file handle: /var/lib/rpm/__db.001 rpmdb: BDB1582 Open file handle: /var/lib/rpm/__db.002 rpmdb: BDB1582 Open file handle: /var/lib/rpm/__db.003 ./run_test.sh: line 17: 2100 Segmentation fault (core dumped) rpm -qi bash Failure at count 738!!! This only occurs when there are this script is run concurrently as root, as normal users cannot append do the /var/lib/rpm/__db.00* files.
Message is usually indicating that something did not close the database cleanly. Looking at the script it doesn't appear to be concurrent access, but serial access. Concurrent access can hit locking issues, which a simple retry of the command is usually enough to resolve the problem. (lock is hit, command errors, other process finishes, clears lock, retry the command and now it succeeds.) Since this appears to be serial, there should be no lock contention. What filesystem is this access on, native filesystem or NFS? We've observed problems with NFS locking in the past. NFS occasionally keeps flocks past when they should have been cleared. This can cause a similar problem to the concurrent access mentioned above. I do not have time to look at this further. If it's consistent reproducible on a regular file system, then it should be possible to find someone to look into the problem and determine why the lock(s) are not being properly cleared.
(In reply to comment #1) > Message is usually indicating that something did not close the database > cleanly. > > Looking at the script it doesn't appear to be concurrent access, but serial > access. > > Concurrent access can hit locking issues, which a simple retry of the > command is usually enough to resolve the problem. (lock is hit, command > errors, other process finishes, clears lock, retry the command and now it > succeeds.) > > Since this appears to be serial, there should be no lock contention. > > What filesystem is this access on, native filesystem or NFS? We've observed > problems with NFS locking in the past. NFS occasionally keeps flocks past > when they should have been cleared. This can cause a similar problem to the > concurrent access mentioned above. > > I do not have time to look at this further. If it's consistent reproducible > on a regular file system, then it should be possible to find someone to look > into the problem and determine why the lock(s) are not being properly > cleared. The problem is reproducible on a regular ext4 filesystem, and the script must be ran multiple times from different terminals at the same time for the bug to reproduce. The problem might indeed be flock related, -or the lack of flocks, perhaps rpm5 should lock the database on access. It doesn't seem to do so in the current implementation.
I could reproduce the issue, the segfault is at level of Berkeley db library, i'll continue to dig into the codebase.
(In reply to comment #3) > I could reproduce the issue, the segfault is at level of Berkeley db > library, i'll continue to dig into the codebase. Thanks for looking into it. It is noteworthy to point out that this only happens when the command is ran as root. Probably because users have only read permissions on the .db files. Perhaps it would be a good idea to do apply an exclusive lock via the flock(2) syscall.
(In reply to comment #4) > (In reply to comment #3) > > I could reproduce the issue, the segfault is at level of Berkeley db > > library, i'll continue to dig into the codebase. > > Thanks for looking into it. It is noteworthy to point out that this only > happens when the command is ran as root. Probably because users have only > read permissions on the .db files. Perhaps it would be a good idea to do > apply an exclusive lock via the flock(2) syscall. Yes, i'm compiling rpm with --with-db-mutex=fcntl to see what happen.
It is worth experimenting with other locking mechanisms in BerkleyDB. One thing to be aware of. The lock format is often architecture dependent and is written into the database at creation time. If you do not use a 'generic' locking method, when it is written into the database by the host system, it will cause the target system (of an incompatible arch) to fail. So any changes you make need to be verified across architectures and endians.
(In reply to comment #6) > It is worth experimenting with other locking mechanisms in BerkleyDB. > > One thing to be aware of. The lock format is often architecture dependent > and is written into the database at creation time. > > If you do not use a 'generic' locking method, when it is written into the > database by the host system, it will cause the target system (of an > incompatible arch) to fail. > > So any changes you make need to be verified across architectures and endians. I think the flock(2) syscall covers the support on archs and endians and this support is at level of rpm5 not in Berkeley db but you are right it needs testing. If everything goes well, i'll test into the GDC Autobuilder.
(In reply to comment #7) > (In reply to comment #6) > > It is worth experimenting with other locking mechanisms in BerkleyDB. > > > > One thing to be aware of. The lock format is often architecture dependent > > and is written into the database at creation time. > > > > If you do not use a 'generic' locking method, when it is written into the > > database by the host system, it will cause the target system (of an > > incompatible arch) to fail. > > > > So any changes you make need to be verified across architectures and endians. > > I think the flock(2) syscall covers the support on archs and endians and > this support is at level of rpm5 not in Berkeley db but you are right it > needs testing. If everything goes well, i'll test into the GDC Autobuilder. It took more time/tries to segfault the counter stops at 274 but the problem persist.
I found that a minor upgrades solves the issue, i sent the patch [1] for review with some notes [2]. [1] http://lists.openembedded.org/pipermail/openembedded-core/2016-September/126992.html [2] http://lists.openembedded.org/pipermail/openembedded-core/2016-September/126993.html
Fixed at, http://git.yoctoproject.org/cgit/cgit.cgi/poky/commit/?id=9393b160733c19af1b1d70cc066b1bddac83bcee