3 subject="""comment 5"""
4 date="2022-05-10T16:56:47Z"
6 As well as moving the object file, fsck will need to move any other associated
7 files, including the object lock file. It may as well move the whole
10 Locking is a concern for implementing this in fsck. There
11 would be a race where another process that is locking the object file
12 sees the object file in the old location, so tries to lock it in the old
13 location, but by then the object file has been moved.
15 Experimentally: In v10, moving the object file after it has checked its
16 location in preparation for locking for drop results in it making a
17 separate lock file in the old object directory. That lock file remains after
18 the drop succeeds. In v8/v9, it seems to not create the object
19 file when trying to lock it. (Based on reading the code, I though perhaps
20 it would!) In v8-v10, moving the object directory in the race when it's locking
21 content in place causes the lock to fail; it does not create any lock file
24 So, v10 post drop lock file cleanup is the problem. Or at least one
25 problem, there could be other points in the race than the one I tested
26 that have other behavior. This seems like an ugly race to insert fsck into
27 the middle of; it would be much preferable if fsck could somehow avoid
28 such races when moving the object directory. But how?
30 fsck could lock the object file for drop, and then rather than removeing it,
31 move it to a holding location. Then it could move the object file
32 into the right place the same as `get` does. This should avoid the race.
33 Interrupting fsck at the wrong time would leave the object file in this
34 holding location though. If it used `.git/annex/tmp`, normal commands
35 like `git-annex get` would recover from an interrupted fsck, though
36 they would need to do some work to rehash the tmp file. Re-running
37 fsck would need to also recover from an interrupted fsck.