Bug 13470

Summary: Autobuilder NFS loses open deleted files
Product: [Infrastructure] AutoBuilder Reporter: Richard Purdie <richard.purdie>
Component: autobuilderAssignee: Michael Halstead <mhalstead>
Status: RESOLVED OBSOLETE QA Contact:
Severity: normal    
Priority: Medium CC: infras.ab.watcher, Infras.watcher, paulg, pidge, randy.macleod
Version: 3.0   
Target Milestone: 5.99   
Hardware: x86   
OS: Multiple   
Whiteboard:
OS type for building Yocto: --- Type of Regression: ---
Verified: Documentation change: No (bug/feature does not impact docs)

Description Richard Purdie 2019-08-08 13:04:08 UTC
If worker A deletes a file which is still being used by worker B, it loses the file rather than it sticking around until its closed.

Documenting this here so we have it recorded somewhere.

pokybuild@debian9-ty-2:/srv/autobuilder/buildhistory/test$ python3 test.py 
['.nfs000000000e48e34100004463', 'test1.py', 'test.py']

then once test.py is started, running test1.py:

[pokybuild@centos7-ty-4 test]$ python3 test1.py 
df69a7bdf92bbdffdae594f12d8efc2aa6c11eb40fa46af36a06c4d8e95b7338
Traceback (most recent call last):
  File "test1.py", line 22, in <module>
    for chunk in iter(lambda: f.read(4096), b""):
  File "test1.py", line 22, in <lambda>
    for chunk in iter(lambda: f.read(4096), b""):
OSError: [Errno 116] Stale file handle

which represents the problem.

[pokybuild@centos7-ty-4 test]$ cat test1.py 
#!/usr/bin/env python3
import subprocess
import time
import os
import hashlib

h = hashlib.sha256()

testdir = "/srv/autobuilder/buildhistory/test"
#subprocess.check_call("cp /usr/bin/zip %s/zip" % testdir, shell=True)
with open(testdir + "/zip", "rb") as f:
    foo = f.read(4096)
    h.update(foo)
    print(h.hexdigest())
    #subprocess.check_call("rm %s/zip" % testdir, shell=True)
    time.sleep(30)
    foo2 = f.read(4096)
    h.update(foo2)
    time.sleep(5)
    foo3 = f.read(4096)
    h.update(foo3)
    for chunk in iter(lambda: f.read(4096), b""):
        h.update(chunk)
    print(h.hexdigest())
    foo3 = f.seek(0)
    foo3 = f.read(4096)
    print(str(os.listdir(testdir)))

#!/usr/bin/env python3
import subprocess
import time
import os

testdir = "/srv/autobuilder/buildhistory/test"
subprocess.check_call("cp /usr/bin/zip %s/zip" % testdir, shell=True)
time.sleep(10)
with open(testdir + "/zip", "rb") as f:
    foo = f.read(4096)
    time.sleep(5)
    subprocess.check_call("rm %s/zip" % testdir, shell=True)
    foo2 = f.read(4096)
    time.sleep(5)
    foo3 = f.read(4096)
    print(str(os.listdir(testdir)))

An example build failure this causes are the WARNING in:

https://autobuilder.yoctoproject.org/typhoon/#/builders/53/builds/847

step1b: WARNING: Logfile for failed setscene task is /home/pokybuild/yocto-worker/qemuarm/build/build/tmp/work/x86_64-linux/python3-setuptools-native/41.0.1-r0/temp/log.do_populate_sysroot_setscene.47244
step1b: WARNING: Setscene task (virtual:native:/home/pokybuild/yocto-worker/qemuarm/build/meta/recipes-devtools/python/python3-setuptools_41.0.1.bb:do_populate_sysroot_setscene) failed with exit code '1' - real task will be run instead

where the setscene log would show a stale file handle.

I'd swear the NFS used to keep the file around until all open handles were closed.
Comment 1 Richard Purdie 2019-08-08 14:50:42 UTC
http://git.yoctoproject.org/cgit.cgi/poky/commit/?id=6c7c0cefd34067311144a1d4c01986fe0a4aef26 was added as a workaround on master
Comment 2 Michael Halstead 2019-11-07 20:02:49 UTC
When the file is deleted on the same host with open handles the client renames the file to .nfsXXXXXX and the handle is maintained on that host. Other clients' file handles go stale when the path changes. I don't think this can be avoided currently.
Comment 3 Randy MacLeod 2022-05-05 15:26:27 UTC
I wonder if this is something that NFS is even capable of...
Comment 4 Randy MacLeod 2023-10-24 14:24:41 UTC
Bulk move from 4.99 or 0.00 to 5.99
Comment 5 Michael Halstead 2025-06-17 20:38:13 UTC
The valkyrie cluster has not experienced these problems. I don't think it is due to hardware changes or updates to NFS though. I suspect code changes in the project avoid this problem. This is solved in all currently supported versions of Yocto Project as far as I know.