speed up keys database writes
authorJoey Hess <joeyh@joeyh.name>
Mon, 31 May 2021 18:56:14 +0000 (14:56 -0400)
committerJoey Hess <joeyh@joeyh.name>
Mon, 31 May 2021 19:01:00 +0000 (15:01 -0400)
There seems to be no reason to check the time here. I think it was
inherited from code in Database.Fsck, which does have a reason to commit
every few minutes. Removing that syscall speeds up a git-annex init
in a repo with 100000 annexed files by about 3 seconds.

Sponsored-by: Dartmouth College's Datalad project
Database/Fsck.hs
Database/Keys/SQL.hs
doc/todo/Avoid_lengthy___34__Scanning_for_unlocked_files_...__34__.mdwn
doc/todo/Avoid_lengthy___34__Scanning_for_unlocked_files_...__34__/comment_13_440624874dd3697dd538655765f2b6a2._comment [new file with mode: 0644]

index c1e9841978400d9dd08141cf15cccb6e7bd0a9c1..ab7a14c95e9fd5c02f6f745bd83558fbd7833bf3 100644 (file)
@@ -88,7 +88,9 @@ addDb :: FsckHandle -> Key -> IO ()
 addDb (FsckHandle h _) k = H.queueDb h checkcommit $
        void $ insertUnique $ Fscked k
   where
-       -- commit queue after 1000 files or 5 minutes, whichever comes first
+       -- Commit queue after 1000 changes or 5 minutes, whichever comes first.
+       -- The time based commit allows for an incremental fsck to be
+       -- interrupted and not lose much work.
        checkcommit sz lastcommittime
                | sz > 1000 = return True
                | otherwise = do
index 7d191bfb4c1d35601cd5c94682b9f5b04fc325ff..5aed3db7b44e9cb83c3171cfd7594ca2a2776202 100644 (file)
@@ -27,7 +27,6 @@ import Git.FilePath
 
 import Database.Persist.Sql hiding (Key)
 import Database.Persist.TH
-import Data.Time.Clock
 import Control.Monad
 import Data.Maybe
 
@@ -77,12 +76,8 @@ newtype WriteHandle = WriteHandle H.DbQueue
 queueDb :: SqlPersistM () -> WriteHandle -> IO ()
 queueDb a (WriteHandle h) = H.queueDb h checkcommit a
   where
-       -- commit queue after 1000 changes or 5 minutes, whichever comes first
-       checkcommit sz lastcommittime
-               | sz > 1000 = return True
-               | otherwise = do
-                       now <- getCurrentTime
-                       return $ diffUTCTime now lastcommittime > 300
+       -- commit queue after 1000 changes
+       checkcommit sz _lastcommittime = pure (sz > 1000)
 
 addAssociatedFile :: Key -> TopFilePath -> WriteHandle -> IO ()
 addAssociatedFile k f = queueDb $ do
index fdd7ee79786414918fc4813a6174ac0687674a85..a20f7955ee1232f7aa4d54dac02d6625ecc7900e 100644 (file)
@@ -4,3 +4,6 @@ E.g. following idea came to mind: git-annex could add some flag/beacon file (e.g
 
 [[!meta author=yoh]]
 [[!tag projects/datalad]]
+
+> I think I've improved this all that it can reasonably be sped up,
+> so [[done]]. --[[Joey]]
diff --git a/doc/todo/Avoid_lengthy___34__Scanning_for_unlocked_files_...__34__/comment_13_440624874dd3697dd538655765f2b6a2._comment b/doc/todo/Avoid_lengthy___34__Scanning_for_unlocked_files_...__34__/comment_13_440624874dd3697dd538655765f2b6a2._comment
new file mode 100644 (file)
index 0000000..9c0208f
--- /dev/null
@@ -0,0 +1,18 @@
+[[!comment format=mdwn
+ username="joey"
+ subject="""comment 13"""
+ date="2021-05-31T18:40:59Z"
+ content="""
+There was an unncessary check of the current time per sql insert, removing
+that sped it up by 3 seconds in my benchmark.
+
+Also tried increasing the number of inserts per sqlite transaction from 1k
+to 10k. Memory use increased to 90 mb, but no measurable speed increase.
+
+I don't see much else that can speed up the sqlite part, without going deep
+into the weeds of populating sqlite databases without using sql, or using
+multi-value inserts ([like described here](https://medium.com/@JasonWyatt/squeezing-performance-from-sqlite-insertions-971aff98eef2).
+Both would prevent using persistent to abstract sql away, and would
+only be usable in this case, not speeding up git-annex generally, 
+so not too enthused.
+"""]]