9.7 KiB
Seafile Repo Reconstruction from Commit Storage
When to Use
After MySQL databases have been wiped/reinitialized but /opt/seafile-data/
(file blocks + commit objects) is intact. This procedure registers all repos
found in the commit storage back into the fresh MySQL DBs so they appear in the
web UI and API.
Commit Storage Layout
/opt/seafile-data/seafile/seafile-data/storage/commits/
<repo_id>/
<first2chars_of_commit_sha>/
<remaining38chars_of_commit_sha> ← JSON file, commit object
Example: commit 5fb48bca960bf13a15df0b84dbee73ef1d8f98d6 lives at:
storage/commits/<repo_id>/5f/b48bca960bf13a15df0b84dbee73ef1d8f98d6
Commit Object Format (JSON)
Each commit file is a JSON object:
{
"commit_id": "5fb48bca960bf13a15df0b84dbee73ef1d8f98d6",
"root_id": "0000000000000000000000000000000000000000",
"repo_id": "3428f3b9-d6c0-4d23-8a03-f34ce8b888e1",
"creator_name": "dominik@familie-schoen.com",
"creator": "0000000000000000000000000000000000000000",
"description": "Created library",
"ctime": 1755169060,
"parent_id": null,
"second_parent_id": null,
"repo_name": "Arztrechnungen",
"repo_desc": "Arztrechnungen",
"repo_category": null,
"no_local_history": 1,
"version": 1
}
Finding the HEAD Commit
The HEAD commit is the one NOT referenced as parent_id (or second_parent_id)
by any other commit in the same repo. Algorithm:
- Read all (or top N by mtime) commit JSON objects for a repo
- Collect all
parent_idandsecond_parent_idvalues into a set - HEAD = the commit ID(s) NOT in that parent set
- If multiple HEADs exist (branch tips), pick the one with the latest
ctime
Performance for Large Repos
Some repos have tens of thousands of commits (34000+). Reading all JSON files is too slow (60s+ timeout). Optimization:
# Get only the 20 newest commit files by filesystem mtime
find "$repo_dir" -type f -printf '%T@ %p\n' | sort -rn | head -20
Then parse only those 20 files. The HEAD is very likely among the newest commits.
⚠️ Filesystem mtime reflects when PBS restored the files, NOT the commit creation time. Don't rely on mtime alone for determining HEAD — use the parent-reference algorithm on the sampled commits.
DB Tables to Populate
| Table | DB | Columns | Notes |
|---|---|---|---|
Repo |
seafile_db |
id (auto), repo_id (unique) |
Just insert the repo_id |
Branch |
seafile_db |
id (auto), name, repo_id, commit_id |
name='master', commit_id=HEAD |
RepoOwner |
seafile_db |
id (auto), repo_id (unique), owner_id |
owner_id = email address |
RepoInfo |
seafile_db |
id (auto), repo_id (unique), name, update_time, version, is_encrypted, last_modifier, status |
name from commit JSON repo_name |
Without RepoOwner, repos won't appear in the API (/api2/repos/) even if
registered in Repo and Branch. Without RepoInfo, repos show generic names.
Step-by-Step Procedure
Step 1: Generate SQL for Repo + Branch registration
#!/usr/bin/env python3
import os, json, subprocess
COMMITS_DIR = "/opt/seafile-data/seafile/seafile-data/storage/commits"
OWNER = "dominik@example.com" # admin email
sql_lines = ["-- Register repos"]
for repo_id in sorted(os.listdir(COMMITS_DIR)):
repo_dir = os.path.join(COMMITS_DIR, repo_id)
if not os.path.isdir(repo_dir):
continue
# Get newest 20 commits by mtime
result = subprocess.run(
["find", repo_dir, "-type", "f", "-printf", "%T@ %p\n"],
capture_output=True, text=True, timeout=10)
lines = result.stdout.strip().split('\n') if result.stdout.strip() else []
lines.sort(key=lambda x: float(x.split()[0]) if x.strip() else 0, reverse=True)
top_files = [l.split(None, 1)[1] for l in lines[:20] if ' ' in l]
commits = {}
all_parents = set()
repo_name = None
latest_ctime = 0
latest_cid = None
for fpath in top_files:
parts = fpath.rstrip('/').split('/')
commit_id = parts[-2] + parts[-1] # subdir + filename = full SHA
try:
with open(fpath, 'r') as f:
data = json.load(f)
commits[commit_id] = data
if data.get('parent_id'):
all_parents.add(data['parent_id'])
if data.get('second_parent_id'):
all_parents.add(data['second_parent_id'])
if data.get('ctime', 0) > latest_ctime:
latest_ctime = data['ctime']
latest_cid = commit_id
rn = data.get('repo_name')
if rn:
repo_name = rn
except:
pass
head_candidates = [c for c in commits if c not in all_parents]
head_id = head_candidates[0] if len(head_candidates) == 1 else \
max(head_candidates, key=lambda c: commits[c].get('ctime', 0)) if head_candidates else latest_cid
if not head_id:
continue
if head_id in commits:
rn = commits[head_id].get('repo_name')
if rn:
repo_name = rn
if not repo_name:
repo_name = f"Repo-{repo_id[:8]}"
safe_name = repo_name.replace("'", "\\'")
# Truncate encrypted commit IDs to 40 chars
head_id_safe = head_id[:40]
sql_lines.append(f"INSERT IGNORE INTO seafile_db.Repo (repo_id) VALUES ('{repo_id}');")
sql_lines.append(f"INSERT IGNORE INTO seafile_db.Branch (name, repo_id, commit_id) VALUES ('master', '{repo_id}', '{head_id_safe}');")
sql_lines.append(f"INSERT IGNORE INTO seafile_db.RepoOwner (repo_id, owner_id) VALUES ('{repo_id}', '{OWNER}');")
sql_lines.append(f"INSERT IGNORE INTO seafile_db.RepoInfo (repo_id, name, update_time, version, is_encrypted, last_modifier, status) VALUES ('{repo_id}', '{safe_name}', UNIX_TIMESTAMP(), 1, 0, '{OWNER}', 0);")
with open("/tmp/register_repos.sql", "w") as f:
f.write("\n".join(sql_lines) + "\n")
print(f"Generated {len(sql_lines)-1} statements")
Step 2: Execute the SQL
docker exec -i seafile-mysql mysql -uroot -p'ROOT_PW' < /tmp/register_repos.sql
Step 3: Run seaf-fsck repair
docker exec seafile /opt/seafile/seafile-server-13.0.19/seaf-fsck.sh --repair
For large repos, run in background:
nohup docker exec seafile /opt/seafile/seafile-server-13.0.19/seaf-fsck.sh --repair \
> /tmp/fsck-output.log 2>&1 &
Monitor progress:
grep -c "Fsck finished" /tmp/fsck-output.log
tail -5 /tmp/fsck-output.log
Step 4: Verify via API
# Get auth token
TOKEN=$(curl -s -X POST http://localhost:80/api2/auth-token/ \
-d 'username=admin@example.com&password=PASSWORD' | python3 -c "import sys,json; print(json.load(sys.stdin)['token'])")
# List repos
curl -s -H "Authorization: Token $TOKEN" http://localhost:80/api2/repos/ \
| python3 -c "import sys,json; data=json.load(sys.stdin); print(f'{len(data)} repos visible')"
# Check a specific repo's contents
curl -s -H "Authorization: Token $TOKEN" \
"http://localhost:80/api2/repos/<repo_id>/dir/" \
| python3 -c "import sys,json; data=json.load(sys.stdin); print(f'{len(data)} items')"
Pitfalls
-
Encrypted repo commit IDs have suffixes — Encrypted repos append a suffix like
.FYMWR3to the 40-char SHA1, producing a 46-char string that exceeds thecommit_id CHAR(41)column. Error:Data too long for column 'commit_id'. Fix: truncate to 40 chars withhead_id[:40]. -
seaf-fsck.sh --repaironly checks REGISTERED repos — It does NOT discover or register new repos from the filesystem. All repos must be inserted into theRepotable first. Same applies toseaf-fsck.sh --export. -
Repos invisible without
RepoOwner— Even withRepo+Branchentries, repos won't appear in/api2/repos/without aRepoOwnerentry mapping the repo to a user's email. -
Repos show generic names without
RepoInfo— WithoutRepoInfo, repos appear as unnamed entries. Thenamecolumn comes from the commit JSON'srepo_namefield. -
Repo sizes show 0 after reconstruction — Size calculation happens during
seaf-fsckor background indexing. Sizes will populate after fsck completes. -
Shares, groups, and non-admin users are lost — This procedure restores repo ownership and file access, but share permissions, group memberships, and user accounts (except the admin) must be recreated manually.
-
Large repos stall fsck — Repos with 30000+ commits can take 10+ minutes during fsck. Run in background and monitor via
grep -c "Fsck finished". -
CRITICAL: Never run multiple fsck processes simultaneously — If
seaf-fsck.sh --repairis already running (check withps aux | grep seaf-fsck | grep -v grep), do NOT start a second one — even via a different invocation method (e.g., one vianohupon the host and another viadocker exec). Two concurrent--repairprocesses on the same storage can corrupt commit/block data. Alwayskill -9the duplicate before continuing. Only oneseaf-fsckprocess should ever be running at a time. -
fsck can die silently (OOM-kill) —
seaf-fsck.sh --repairmay be OOM-killed by the kernel with no trace in the log file. The log simply stops at whatever repo was being processed, andps aux | grep seaf-fsckreturns 0 processes. Always check process liveness, not just log tail. If killed, restart fsck — it resumes from where it left off (repos already checked are quickly re-verified). -
Large photo repos stall fsck indefinitely — Repos with many large image files (e.g., "Fotos", "Lightroom Bilder") can take 30+ minutes each during fsck. Combined with concurrent Ceph I/O (recovery, RBD copies), fsck throughput drops to near-zero. If fsck appears stuck, check whether heavy I/O is competing and consider pausing other operations.