| Summary: | rpm: Segmentation fault when running parallel multiple "rpm -qi" commands as root | ||||||
|---|---|---|---|---|---|---|---|
| Product: | [Build System, Metadata & Runtime] OE-Core | Reporter: | Alexandru Moise <00moses.alexander00> | ||||
| Component: | core | Assignee: | Aníbal Limón <anibal.limon> | ||||
| Status: | RESOLVED FIXED | QA Contact: | |||||
| Severity: | normal | ||||||
| Priority: | High | CC: | mark.hatle, meta.mr.watcher, meta.watcher, randy.macleod, sgw | ||||
| Version: | 2.2 | ||||||
| Target Milestone: | 2.2 M4 | ||||||
| Hardware: | All | ||||||
| OS: | Multiple | ||||||
| Whiteboard: | |||||||
| OS type for building Yocto: | --- | Type of Regression: | --- | ||||
| Verified: | Documentation change: | No (bug/feature does not impact docs) | |||||
| Attachments: |
|
||||||
|
Description
Alexandru Moise
2016-08-17 09:11:42 UTC
Message is usually indicating that something did not close the database cleanly. Looking at the script it doesn't appear to be concurrent access, but serial access. Concurrent access can hit locking issues, which a simple retry of the command is usually enough to resolve the problem. (lock is hit, command errors, other process finishes, clears lock, retry the command and now it succeeds.) Since this appears to be serial, there should be no lock contention. What filesystem is this access on, native filesystem or NFS? We've observed problems with NFS locking in the past. NFS occasionally keeps flocks past when they should have been cleared. This can cause a similar problem to the concurrent access mentioned above. I do not have time to look at this further. If it's consistent reproducible on a regular file system, then it should be possible to find someone to look into the problem and determine why the lock(s) are not being properly cleared. (In reply to comment #1) > Message is usually indicating that something did not close the database > cleanly. > > Looking at the script it doesn't appear to be concurrent access, but serial > access. > > Concurrent access can hit locking issues, which a simple retry of the > command is usually enough to resolve the problem. (lock is hit, command > errors, other process finishes, clears lock, retry the command and now it > succeeds.) > > Since this appears to be serial, there should be no lock contention. > > What filesystem is this access on, native filesystem or NFS? We've observed > problems with NFS locking in the past. NFS occasionally keeps flocks past > when they should have been cleared. This can cause a similar problem to the > concurrent access mentioned above. > > I do not have time to look at this further. If it's consistent reproducible > on a regular file system, then it should be possible to find someone to look > into the problem and determine why the lock(s) are not being properly > cleared. The problem is reproducible on a regular ext4 filesystem, and the script must be ran multiple times from different terminals at the same time for the bug to reproduce. The problem might indeed be flock related, -or the lack of flocks, perhaps rpm5 should lock the database on access. It doesn't seem to do so in the current implementation. I could reproduce the issue, the segfault is at level of Berkeley db library, i'll continue to dig into the codebase. (In reply to comment #3) > I could reproduce the issue, the segfault is at level of Berkeley db > library, i'll continue to dig into the codebase. Thanks for looking into it. It is noteworthy to point out that this only happens when the command is ran as root. Probably because users have only read permissions on the .db files. Perhaps it would be a good idea to do apply an exclusive lock via the flock(2) syscall. (In reply to comment #4) > (In reply to comment #3) > > I could reproduce the issue, the segfault is at level of Berkeley db > > library, i'll continue to dig into the codebase. > > Thanks for looking into it. It is noteworthy to point out that this only > happens when the command is ran as root. Probably because users have only > read permissions on the .db files. Perhaps it would be a good idea to do > apply an exclusive lock via the flock(2) syscall. Yes, i'm compiling rpm with --with-db-mutex=fcntl to see what happen. It is worth experimenting with other locking mechanisms in BerkleyDB. One thing to be aware of. The lock format is often architecture dependent and is written into the database at creation time. If you do not use a 'generic' locking method, when it is written into the database by the host system, it will cause the target system (of an incompatible arch) to fail. So any changes you make need to be verified across architectures and endians. (In reply to comment #6) > It is worth experimenting with other locking mechanisms in BerkleyDB. > > One thing to be aware of. The lock format is often architecture dependent > and is written into the database at creation time. > > If you do not use a 'generic' locking method, when it is written into the > database by the host system, it will cause the target system (of an > incompatible arch) to fail. > > So any changes you make need to be verified across architectures and endians. I think the flock(2) syscall covers the support on archs and endians and this support is at level of rpm5 not in Berkeley db but you are right it needs testing. If everything goes well, i'll test into the GDC Autobuilder. (In reply to comment #7) > (In reply to comment #6) > > It is worth experimenting with other locking mechanisms in BerkleyDB. > > > > One thing to be aware of. The lock format is often architecture dependent > > and is written into the database at creation time. > > > > If you do not use a 'generic' locking method, when it is written into the > > database by the host system, it will cause the target system (of an > > incompatible arch) to fail. > > > > So any changes you make need to be verified across architectures and endians. > > I think the flock(2) syscall covers the support on archs and endians and > this support is at level of rpm5 not in Berkeley db but you are right it > needs testing. If everything goes well, i'll test into the GDC Autobuilder. It took more time/tries to segfault the counter stops at 274 but the problem persist. I found that a minor upgrades solves the issue, i sent the patch [1] for review with some notes [2]. [1] http://lists.openembedded.org/pipermail/openembedded-core/2016-September/126992.html [2] http://lists.openembedded.org/pipermail/openembedded-core/2016-September/126993.html |