From 0eff5a3f7124eb912d110681517b78b66bf51c09 Mon Sep 17 00:00:00 2001 From: Joey Hess Date: Mon, 14 Jun 2021 11:37:21 -0400 Subject: [PATCH] reproduced --- ..._8693e7e9c800f25cbd274b6781d834d6._comment | 31 +++++++++++++++++++ ..._6d11f6aa4b1a435bdf6d165eb8e6db8a._comment | 17 ++++++++++ 2 files changed, 48 insertions(+) create mode 100644 doc/bugs/significant_performance_regression_impacting_datal/comment_23_8693e7e9c800f25cbd274b6781d834d6._comment create mode 100644 doc/bugs/significant_performance_regression_impacting_datal/comment_24_6d11f6aa4b1a435bdf6d165eb8e6db8a._comment diff --git a/doc/bugs/significant_performance_regression_impacting_datal/comment_23_8693e7e9c800f25cbd274b6781d834d6._comment b/doc/bugs/significant_performance_regression_impacting_datal/comment_23_8693e7e9c800f25cbd274b6781d834d6._comment new file mode 100644 index 0000000000..af4a33a6ec --- /dev/null +++ b/doc/bugs/significant_performance_regression_impacting_datal/comment_23_8693e7e9c800f25cbd274b6781d834d6._comment @@ -0,0 +1,31 @@ +[[!comment format=mdwn + username="joey" + subject="""comment 23""" + date="2021-06-14T14:09:07Z" + content=""" +The file contents all being the same is the crucial thing. On linux, +adding 1000 dup files at a time (all in same directory), I get: + +run 1: 0:08 +run 2: 0:42 +run 3: 1:14 +run 4: 1:46 + +After run 4, adding 1000 files with all different content takes +0:11, so not appreciably slowed down; it only affects adding dups, +and only when there are a *lot* of them. + +This feels like quite an edge case, and also not +really a new problem, since unlocked files would have already +had the same problem before recent changes. + +I thought this might be an innefficiency in sqlite's index, similar to how +hash tables can scale poorly when a lot of things end up in the same +bucket. But disabling the index did not improve performance. + +Aha -- the slowdown is caused by `git-annex add` looking to see what other +annexed files use the same content, so that it can populate any unlocked +files that didn't have the content present before. With all these locked +files now recorded in the db, it has to check each file in turn, and +there's the `O(N^2)` +"""]] diff --git a/doc/bugs/significant_performance_regression_impacting_datal/comment_24_6d11f6aa4b1a435bdf6d165eb8e6db8a._comment b/doc/bugs/significant_performance_regression_impacting_datal/comment_24_6d11f6aa4b1a435bdf6d165eb8e6db8a._comment new file mode 100644 index 0000000000..ff018e693e --- /dev/null +++ b/doc/bugs/significant_performance_regression_impacting_datal/comment_24_6d11f6aa4b1a435bdf6d165eb8e6db8a._comment @@ -0,0 +1,17 @@ +[[!comment format=mdwn + username="joey" + subject="""comment 24""" + date="2021-06-14T15:36:30Z" + content=""" +If the database recorded when files were unlocked or not, that could be +avoided, but tracking that would add a lot of complexity for what is just +an edge case. And probably slow things down generally by some amount due to +the db being larger. + +It seems almost cheating, but it could remember the last few keys it's added, +and avoid trying to populate unlocked files when adding those keys again. +This would slow down the usual case by some tiny amount (eg an IORef access) +but avoid `O(N^2)` in this edge case. Though it wouldn't fix all edge cases, +eg when the files it's adding rotate through X different contents, and X is +larger than the number of keys it remembers. +"""]] -- 2.30.2