Bug 10157 - rpm: Segmentation fault when running parallel multiple "rpm -qi" commands as root
Summary: rpm: Segmentation fault when running parallel multiple "rpm -qi" commands as ...
Status: RESOLVED FIXED
Alias: None
Product: OE-Core
Classification: Build System, Metadata & Runtime
Component: core (show other bugs)
Version: 2.2
Hardware: All Multiple
: High normal
Target Milestone: 2.2 M4
Assignee: Aníbal Limón
QA Contact:
URL:
Whiteboard:
Depends on:
Blocks:
 
Reported: 2016-08-17 09:11 UTC by Alexandru Moise
Modified: 2016-09-30 20:29 UTC (History)
5 users (show)

See Also:
OS type for building Yocto: ---
Type of Regression: ---
Verified:
Documentation change: No (bug/feature does not impact docs)


Attachments
rpm test script (220 bytes, application/x-shellscript)
2016-08-17 09:11 UTC, Alexandru Moise
no flags Details

Note You need to log in before you can comment on or make changes to this bug.
Description Alexandru Moise 2016-08-17 09:11:42 UTC
Created attachment 3363 [details]
rpm test script

Reproduction script attached.
When the script is run multiple times concurrently the following error occurs:

count 739:
Name        : bash                         Relocations: (not relocatable)
Version     : 4.3.30                            Vendor: (none)
Release     : r0                            Build Date: Tue Aug 16 15:42:24 2016
Install Date: Wed Aug 17 08:45:49 2016      Build Host: obsrwr
Group       : base/shell                    Source RPM: bash-4.3.30-r0.src.rpm
Size        : 1048412                          License: GPLv3+
Signature   : DSA/SHA1, Tue Aug 16 15:42:24 2016, Key ID 5d0492732d8fd654
Packager    : Poky <poky@yoctoproject.org>
URL         : http://tiswww.case.edu/php/chet/bash/bashtop.html
Summary     : An sh-compatible command language interpreter
Architecture: i586
Description :
An sh-compatible command language interpreter.

count 738:
rpmdb: BDB0060 PANIC: fatal region error detected; run recovery
==> rpmdbe_event_notify(0x807aff8, PANIC(0), 0xbffeef8c) app_private (nil)
error: db_init:../../rpm-5.4.15/rpmdb/db3.c:1206: dbenv->failchk(-30973): BDB0087 DB_RUNRECOVERY: Fatal error, run database recovery
rpmdb: BDB1581 File handles still open at environment close
rpmdb: BDB1582 Open file handle: /var/lib/rpm/__db.001
rpmdb: BDB1582 Open file handle: /var/lib/rpm/__db.002
rpmdb: BDB1582 Open file handle: /var/lib/rpm/__db.003
./run_test.sh: line 17:  2100 Segmentation fault      (core dumped) rpm -qi bash
Failure at count 738!!!

This only occurs when there are this script is run concurrently as root, as normal users cannot append do the /var/lib/rpm/__db.00* files.
Comment 1 Mark Hatle 2016-09-21 15:48:51 UTC
Message is usually indicating that something did not close the database cleanly.

Looking at the script it doesn't appear to be concurrent access, but serial access.

Concurrent access can hit locking issues, which a simple retry of the command is usually enough to resolve the problem.  (lock is hit, command errors, other process finishes, clears lock, retry the command and now it succeeds.)

Since this appears to be serial, there should be no lock contention.

What filesystem is this access on, native filesystem or NFS?  We've observed problems with NFS locking in the past.  NFS occasionally keeps flocks past when they should have been cleared.  This can cause a similar problem to the concurrent access mentioned above.

I do not have time to look at this further.  If it's consistent reproducible on a regular file system, then it should be possible to find someone to look into the problem and determine why the lock(s) are not being properly cleared.
Comment 2 Alexandru Moise 2016-09-21 16:40:46 UTC
(In reply to comment #1)
> Message is usually indicating that something did not close the database
> cleanly.
> 
> Looking at the script it doesn't appear to be concurrent access, but serial
> access.
> 
> Concurrent access can hit locking issues, which a simple retry of the
> command is usually enough to resolve the problem.  (lock is hit, command
> errors, other process finishes, clears lock, retry the command and now it
> succeeds.)
> 
> Since this appears to be serial, there should be no lock contention.
> 
> What filesystem is this access on, native filesystem or NFS?  We've observed
> problems with NFS locking in the past.  NFS occasionally keeps flocks past
> when they should have been cleared.  This can cause a similar problem to the
> concurrent access mentioned above.
> 
> I do not have time to look at this further.  If it's consistent reproducible
> on a regular file system, then it should be possible to find someone to look
> into the problem and determine why the lock(s) are not being properly
> cleared.

The problem is reproducible on a regular ext4 filesystem, and the script must be ran multiple times from different terminals at the same time for the bug to reproduce. The problem might indeed be flock related, -or the lack of flocks, perhaps rpm5 should lock the database on access. It doesn't seem to do so in the current implementation.
Comment 3 Aníbal Limón 2016-09-27 18:51:00 UTC
I could reproduce the issue, the segfault is at level of Berkeley db library, i'll continue to dig into the codebase.
Comment 4 Alexandru Moise 2016-09-27 20:25:43 UTC
(In reply to comment #3)
> I could reproduce the issue, the segfault is at level of Berkeley db
> library, i'll continue to dig into the codebase.

Thanks for looking into it. It is noteworthy to point out that this only
happens when the command is ran as root. Probably because users have only
read permissions on the .db files. Perhaps it would be a good idea to do
apply an exclusive lock via the flock(2) syscall.
Comment 5 Aníbal Limón 2016-09-27 20:31:30 UTC
(In reply to comment #4)
> (In reply to comment #3)
> > I could reproduce the issue, the segfault is at level of Berkeley db
> > library, i'll continue to dig into the codebase.
> 
> Thanks for looking into it. It is noteworthy to point out that this only
> happens when the command is ran as root. Probably because users have only
> read permissions on the .db files. Perhaps it would be a good idea to do
> apply an exclusive lock via the flock(2) syscall.

Yes, i'm compiling rpm with --with-db-mutex=fcntl to see what happen.
Comment 6 Mark Hatle 2016-09-27 20:46:37 UTC
It is worth experimenting with other locking mechanisms in BerkleyDB.

One thing to be aware of.  The lock format is often architecture dependent and is written into the database at creation time.

If you do not use a 'generic' locking method, when it is written into the database by the host system, it will cause the target system (of an incompatible arch) to fail.

So any changes you make need to be verified across architectures and endians.
Comment 7 Aníbal Limón 2016-09-27 20:50:28 UTC
(In reply to comment #6)
> It is worth experimenting with other locking mechanisms in BerkleyDB.
> 
> One thing to be aware of.  The lock format is often architecture dependent
> and is written into the database at creation time.
> 
> If you do not use a 'generic' locking method, when it is written into the
> database by the host system, it will cause the target system (of an
> incompatible arch) to fail.
> 
> So any changes you make need to be verified across architectures and endians.

I think the flock(2) syscall covers the support on archs and endians and this support is at level of rpm5 not in Berkeley db but you are right it needs testing. If everything goes well, i'll test into the GDC Autobuilder.
Comment 8 Aníbal Limón 2016-09-27 21:04:18 UTC
(In reply to comment #7)
> (In reply to comment #6)
> > It is worth experimenting with other locking mechanisms in BerkleyDB.
> > 
> > One thing to be aware of.  The lock format is often architecture dependent
> > and is written into the database at creation time.
> > 
> > If you do not use a 'generic' locking method, when it is written into the
> > database by the host system, it will cause the target system (of an
> > incompatible arch) to fail.
> > 
> > So any changes you make need to be verified across architectures and endians.
> 
> I think the flock(2) syscall covers the support on archs and endians and
> this support is at level of rpm5 not in Berkeley db but you are right it
> needs testing. If everything goes well, i'll test into the GDC Autobuilder.

It took more time/tries to segfault the counter stops at 274 but the problem persist.
Comment 9 Aníbal Limón 2016-09-27 23:45:09 UTC
I found that a minor upgrades solves the issue, i sent the patch [1] for review with some notes [2].

[1] http://lists.openembedded.org/pipermail/openembedded-core/2016-September/126992.html
[2] http://lists.openembedded.org/pipermail/openembedded-core/2016-September/126993.html