# Seafile Repo Reconstruction from Commit Storage ## When to Use After MySQL databases have been wiped/reinitialized but `/opt/seafile-data/` (file blocks + commit objects) is intact. This procedure registers all repos found in the commit storage back into the fresh MySQL DBs so they appear in the web UI and API. ## Commit Storage Layout ``` /opt/seafile-data/seafile/seafile-data/storage/commits/ / / ← JSON file, commit object ``` Example: commit `5fb48bca960bf13a15df0b84dbee73ef1d8f98d6` lives at: `storage/commits//5f/b48bca960bf13a15df0b84dbee73ef1d8f98d6` ## Commit Object Format (JSON) Each commit file is a JSON object: ```json { "commit_id": "5fb48bca960bf13a15df0b84dbee73ef1d8f98d6", "root_id": "0000000000000000000000000000000000000000", "repo_id": "3428f3b9-d6c0-4d23-8a03-f34ce8b888e1", "creator_name": "dominik@familie-schoen.com", "creator": "0000000000000000000000000000000000000000", "description": "Created library", "ctime": 1755169060, "parent_id": null, "second_parent_id": null, "repo_name": "Arztrechnungen", "repo_desc": "Arztrechnungen", "repo_category": null, "no_local_history": 1, "version": 1 } ``` ## Finding the HEAD Commit The HEAD commit is the one NOT referenced as `parent_id` (or `second_parent_id`) by any other commit in the same repo. Algorithm: 1. Read all (or top N by mtime) commit JSON objects for a repo 2. Collect all `parent_id` and `second_parent_id` values into a set 3. HEAD = the commit ID(s) NOT in that parent set 4. If multiple HEADs exist (branch tips), pick the one with the latest `ctime` ### Performance for Large Repos Some repos have tens of thousands of commits (34000+). Reading all JSON files is too slow (60s+ timeout). Optimization: ```bash # Get only the 20 newest commit files by filesystem mtime find "$repo_dir" -type f -printf '%T@ %p\n' | sort -rn | head -20 ``` Then parse only those 20 files. The HEAD is very likely among the newest commits. ⚠️ Filesystem mtime reflects when PBS restored the files, NOT the commit creation time. Don't rely on mtime alone for determining HEAD — use the parent-reference algorithm on the sampled commits. ## DB Tables to Populate | Table | DB | Columns | Notes | |-------|-----|---------|-------| | `Repo` | `seafile_db` | `id` (auto), `repo_id` (unique) | Just insert the repo_id | | `Branch` | `seafile_db` | `id` (auto), `name`, `repo_id`, `commit_id` | name='master', commit_id=HEAD | | `RepoOwner` | `seafile_db` | `id` (auto), `repo_id` (unique), `owner_id` | owner_id = email address | | `RepoInfo` | `seafile_db` | `id` (auto), `repo_id` (unique), `name`, `update_time`, `version`, `is_encrypted`, `last_modifier`, `status` | name from commit JSON `repo_name` | Without `RepoOwner`, repos won't appear in the API (`/api2/repos/`) even if registered in `Repo` and `Branch`. Without `RepoInfo`, repos show generic names. ## Step-by-Step Procedure ### Step 1: Generate SQL for Repo + Branch registration ```python #!/usr/bin/env python3 import os, json, subprocess COMMITS_DIR = "/opt/seafile-data/seafile/seafile-data/storage/commits" OWNER = "dominik@example.com" # admin email sql_lines = ["-- Register repos"] for repo_id in sorted(os.listdir(COMMITS_DIR)): repo_dir = os.path.join(COMMITS_DIR, repo_id) if not os.path.isdir(repo_dir): continue # Get newest 20 commits by mtime result = subprocess.run( ["find", repo_dir, "-type", "f", "-printf", "%T@ %p\n"], capture_output=True, text=True, timeout=10) lines = result.stdout.strip().split('\n') if result.stdout.strip() else [] lines.sort(key=lambda x: float(x.split()[0]) if x.strip() else 0, reverse=True) top_files = [l.split(None, 1)[1] for l in lines[:20] if ' ' in l] commits = {} all_parents = set() repo_name = None latest_ctime = 0 latest_cid = None for fpath in top_files: parts = fpath.rstrip('/').split('/') commit_id = parts[-2] + parts[-1] # subdir + filename = full SHA try: with open(fpath, 'r') as f: data = json.load(f) commits[commit_id] = data if data.get('parent_id'): all_parents.add(data['parent_id']) if data.get('second_parent_id'): all_parents.add(data['second_parent_id']) if data.get('ctime', 0) > latest_ctime: latest_ctime = data['ctime'] latest_cid = commit_id rn = data.get('repo_name') if rn: repo_name = rn except: pass head_candidates = [c for c in commits if c not in all_parents] head_id = head_candidates[0] if len(head_candidates) == 1 else \ max(head_candidates, key=lambda c: commits[c].get('ctime', 0)) if head_candidates else latest_cid if not head_id: continue if head_id in commits: rn = commits[head_id].get('repo_name') if rn: repo_name = rn if not repo_name: repo_name = f"Repo-{repo_id[:8]}" safe_name = repo_name.replace("'", "\\'") # Truncate encrypted commit IDs to 40 chars head_id_safe = head_id[:40] sql_lines.append(f"INSERT IGNORE INTO seafile_db.Repo (repo_id) VALUES ('{repo_id}');") sql_lines.append(f"INSERT IGNORE INTO seafile_db.Branch (name, repo_id, commit_id) VALUES ('master', '{repo_id}', '{head_id_safe}');") sql_lines.append(f"INSERT IGNORE INTO seafile_db.RepoOwner (repo_id, owner_id) VALUES ('{repo_id}', '{OWNER}');") sql_lines.append(f"INSERT IGNORE INTO seafile_db.RepoInfo (repo_id, name, update_time, version, is_encrypted, last_modifier, status) VALUES ('{repo_id}', '{safe_name}', UNIX_TIMESTAMP(), 1, 0, '{OWNER}', 0);") with open("/tmp/register_repos.sql", "w") as f: f.write("\n".join(sql_lines) + "\n") print(f"Generated {len(sql_lines)-1} statements") ``` ### Step 2: Execute the SQL ```bash docker exec -i seafile-mysql mysql -uroot -p'ROOT_PW' < /tmp/register_repos.sql ``` ### Step 3: Run seaf-fsck repair ```bash docker exec seafile /opt/seafile/seafile-server-13.0.19/seaf-fsck.sh --repair ``` For large repos, run in background: ```bash nohup docker exec seafile /opt/seafile/seafile-server-13.0.19/seaf-fsck.sh --repair \ > /tmp/fsck-output.log 2>&1 & ``` Monitor progress: ```bash grep -c "Fsck finished" /tmp/fsck-output.log tail -5 /tmp/fsck-output.log ``` ### Step 4: Verify via API ```bash # Get auth token TOKEN=$(curl -s -X POST http://localhost:80/api2/auth-token/ \ -d 'username=admin@example.com&password=PASSWORD' | python3 -c "import sys,json; print(json.load(sys.stdin)['token'])") # List repos curl -s -H "Authorization: Token $TOKEN" http://localhost:80/api2/repos/ \ | python3 -c "import sys,json; data=json.load(sys.stdin); print(f'{len(data)} repos visible')" # Check a specific repo's contents curl -s -H "Authorization: Token $TOKEN" \ "http://localhost:80/api2/repos//dir/" \ | python3 -c "import sys,json; data=json.load(sys.stdin); print(f'{len(data)} items')" ``` ## Pitfalls 1. **Encrypted repo commit IDs have suffixes** — Encrypted repos append a suffix like `.FYMWR3` to the 40-char SHA1, producing a 46-char string that exceeds the `commit_id CHAR(41)` column. Error: `Data too long for column 'commit_id'`. Fix: truncate to 40 chars with `head_id[:40]`. 2. **`seaf-fsck.sh --repair` only checks REGISTERED repos** — It does NOT discover or register new repos from the filesystem. All repos must be inserted into the `Repo` table first. Same applies to `seaf-fsck.sh --export`. 3. **Repos invisible without `RepoOwner`** — Even with `Repo` + `Branch` entries, repos won't appear in `/api2/repos/` without a `RepoOwner` entry mapping the repo to a user's email. 4. **Repos show generic names without `RepoInfo`** — Without `RepoInfo`, repos appear as unnamed entries. The `name` column comes from the commit JSON's `repo_name` field. 5. **Repo sizes show 0 after reconstruction** — Size calculation happens during `seaf-fsck` or background indexing. Sizes will populate after fsck completes. 6. **Shares, groups, and non-admin users are lost** — This procedure restores repo ownership and file access, but share permissions, group memberships, and user accounts (except the admin) must be recreated manually. 7. **Large repos stall fsck** — Repos with 30000+ commits can take 10+ minutes during fsck. Run in background and monitor via `grep -c "Fsck finished"`. 8. **CRITICAL: Never run multiple fsck processes simultaneously** — If `seaf-fsck.sh --repair` is already running (check with `ps aux | grep seaf-fsck | grep -v grep`), do NOT start a second one — even via a different invocation method (e.g., one via `nohup` on the host and another via `docker exec`). Two concurrent `--repair` processes on the same storage can corrupt commit/block data. Always `kill -9` the duplicate before continuing. Only one `seaf-fsck` process should ever be running at a time. 9. **fsck can die silently (OOM-kill)** — `seaf-fsck.sh --repair` may be OOM-killed by the kernel with no trace in the log file. The log simply stops at whatever repo was being processed, and `ps aux | grep seaf-fsck` returns 0 processes. Always check process liveness, not just log tail. If killed, restart fsck — it resumes from where it left off (repos already checked are quickly re-verified). 10. **Large photo repos stall fsck indefinitely** — Repos with many large image files (e.g., "Fotos", "Lightroom Bilder") can take 30+ minutes each during fsck. Combined with concurrent Ceph I/O (recovery, RBD copies), fsck throughput drops to near-zero. If fsck appears stuck, check whether heavy I/O is competing and consider pausing other operations.