Bug 9268

Summary: bug in relocate_sdk.py
Product: [Build System, Metadata & Runtime] OE-Core Reporter: Juro Bystricky <juro.bystricky>
Component: Scripts and ToolsAssignee: Juro Bystricky <juro.bystricky>
Status: RESOLVED FIXED QA Contact:
Severity: major    
Priority: High CC: maciej.borzecki, ross.burton
Version: 2.1   
Target Milestone: 1.4.5   
Hardware: x86   
OS: Multiple   
Whiteboard:
OS type for building Yocto: --- Type of Regression: ---
Verified: Documentation change: No (bug/feature does not impact docs)
Attachments:
Description Flags
relocate_sdk.py with multiple printouts none

Description Juro Bystricky 2016-03-15 19:47:35 UTC
The script relocate_sdk.py contains a bug(s) that can cause to install an SDK with incorrectly relocated/corrupted files.
If installing the SDK into a non-default location and if the path to the new SDK location is longer than the path of the default SDK location, we will overwrite some data inadvertently. The intent is to write 4096 bytes, but we will in fact end up writing more than that, easily proved by a printout:

...
while (offset + 4096) <= sh_size:
                    path = f.read(4096)
                    new_path = old_prefix.sub(new_prefix, path)             
                    # pad with zeros
                    new_path += b("\0") * (4096 - len(new_path))
                    #print "Changing %s to %s at %s" % (str(path), str(new_path), str(offset))
                    # write it back
                    f.seek(sh_offset + offset)
                    print("Write length: %i" % (len(new_path)))
                    f.write(new_path)
....

This will overwrite/corrupt the beginning of the next entry we want to process.
The reason we do not always observe the problem is (IMHO) because there is another bug, related to file flushing, basically compensating the overwriting bug. 
I believe adding an explicit flush to the code proves my point:
                    # write it back
                    f.seek(sh_offset + offset)
                    print("Write length: %i" % (len(new_path)))
                    f.write(new_path)
                    f.flush()

The flush should make no difference, yet it will guarantee an entry to be overwritten (if the conditions are "right": i.e. new_path longer than 4096). 

Similar code is used in two places.

At first I observed this with a python3 system, but it turns out it has nothing to do with python3, it was also observed with python2 and installing SDK on a NFS mounted drive (most likely due to some differences in fflush)
The remedy would be to ensure we always write 4096 bytes without the assumption that len(new_path) is 4096 bytes.
Comment 1 Juro Bystricky 2016-03-15 20:25:49 UTC
Created attachment 3031 [details]
relocate_sdk.py with multiple printouts
Comment 2 Juro Bystricky 2016-03-28 23:03:37 UTC
A patch that fixes this issue was merged:

http://git.yoctoproject.org/cgit.cgi/poky/commit/?id=c3c793b4286367b7050c8ec92fb90a1a9f85a89a

The patch fixes the relocation of the section ".gccrelocprefix".
That fixes the observed problems.

However, similar fix is still needed for the relocation of the section ".ldsocache".
Comment 3 Juro Bystricky 2016-03-31 00:45:18 UTC
Sent a patch to the mailing list.