Files
hermes-skills/devops/seafile-api/references/seafile-repo-reconstruction.md
T

9.7 KiB

Seafile Repo Reconstruction from Commit Storage

When to Use

After MySQL databases have been wiped/reinitialized but /opt/seafile-data/ (file blocks + commit objects) is intact. This procedure registers all repos found in the commit storage back into the fresh MySQL DBs so they appear in the web UI and API.

Commit Storage Layout

/opt/seafile-data/seafile/seafile-data/storage/commits/
  <repo_id>/
    <first2chars_of_commit_sha>/
      <remaining38chars_of_commit_sha>     ← JSON file, commit object

Example: commit 5fb48bca960bf13a15df0b84dbee73ef1d8f98d6 lives at: storage/commits/<repo_id>/5f/b48bca960bf13a15df0b84dbee73ef1d8f98d6

Commit Object Format (JSON)

Each commit file is a JSON object:

{
  "commit_id": "5fb48bca960bf13a15df0b84dbee73ef1d8f98d6",
  "root_id": "0000000000000000000000000000000000000000",
  "repo_id": "3428f3b9-d6c0-4d23-8a03-f34ce8b888e1",
  "creator_name": "dominik@familie-schoen.com",
  "creator": "0000000000000000000000000000000000000000",
  "description": "Created library",
  "ctime": 1755169060,
  "parent_id": null,
  "second_parent_id": null,
  "repo_name": "Arztrechnungen",
  "repo_desc": "Arztrechnungen",
  "repo_category": null,
  "no_local_history": 1,
  "version": 1
}

Finding the HEAD Commit

The HEAD commit is the one NOT referenced as parent_id (or second_parent_id) by any other commit in the same repo. Algorithm:

  1. Read all (or top N by mtime) commit JSON objects for a repo
  2. Collect all parent_id and second_parent_id values into a set
  3. HEAD = the commit ID(s) NOT in that parent set
  4. If multiple HEADs exist (branch tips), pick the one with the latest ctime

Performance for Large Repos

Some repos have tens of thousands of commits (34000+). Reading all JSON files is too slow (60s+ timeout). Optimization:

# Get only the 20 newest commit files by filesystem mtime
find "$repo_dir" -type f -printf '%T@ %p\n' | sort -rn | head -20

Then parse only those 20 files. The HEAD is very likely among the newest commits.

⚠️ Filesystem mtime reflects when PBS restored the files, NOT the commit creation time. Don't rely on mtime alone for determining HEAD — use the parent-reference algorithm on the sampled commits.

DB Tables to Populate

Table DB Columns Notes
Repo seafile_db id (auto), repo_id (unique) Just insert the repo_id
Branch seafile_db id (auto), name, repo_id, commit_id name='master', commit_id=HEAD
RepoOwner seafile_db id (auto), repo_id (unique), owner_id owner_id = email address
RepoInfo seafile_db id (auto), repo_id (unique), name, update_time, version, is_encrypted, last_modifier, status name from commit JSON repo_name

Without RepoOwner, repos won't appear in the API (/api2/repos/) even if registered in Repo and Branch. Without RepoInfo, repos show generic names.

Step-by-Step Procedure

Step 1: Generate SQL for Repo + Branch registration

#!/usr/bin/env python3
import os, json, subprocess

COMMITS_DIR = "/opt/seafile-data/seafile/seafile-data/storage/commits"
OWNER = "dominik@example.com"  # admin email

sql_lines = ["-- Register repos"]
for repo_id in sorted(os.listdir(COMMITS_DIR)):
    repo_dir = os.path.join(COMMITS_DIR, repo_id)
    if not os.path.isdir(repo_dir):
        continue
    # Get newest 20 commits by mtime
    result = subprocess.run(
        ["find", repo_dir, "-type", "f", "-printf", "%T@ %p\n"],
        capture_output=True, text=True, timeout=10)
    lines = result.stdout.strip().split('\n') if result.stdout.strip() else []
    lines.sort(key=lambda x: float(x.split()[0]) if x.strip() else 0, reverse=True)
    top_files = [l.split(None, 1)[1] for l in lines[:20] if ' ' in l]

    commits = {}
    all_parents = set()
    repo_name = None
    latest_ctime = 0
    latest_cid = None
    for fpath in top_files:
        parts = fpath.rstrip('/').split('/')
        commit_id = parts[-2] + parts[-1]  # subdir + filename = full SHA
        try:
            with open(fpath, 'r') as f:
                data = json.load(f)
            commits[commit_id] = data
            if data.get('parent_id'):
                all_parents.add(data['parent_id'])
            if data.get('second_parent_id'):
                all_parents.add(data['second_parent_id'])
            if data.get('ctime', 0) > latest_ctime:
                latest_ctime = data['ctime']
                latest_cid = commit_id
                rn = data.get('repo_name')
                if rn:
                    repo_name = rn
        except:
            pass

    head_candidates = [c for c in commits if c not in all_parents]
    head_id = head_candidates[0] if len(head_candidates) == 1 else \
              max(head_candidates, key=lambda c: commits[c].get('ctime', 0)) if head_candidates else latest_cid
    if not head_id:
        continue
    if head_id in commits:
        rn = commits[head_id].get('repo_name')
        if rn:
            repo_name = rn
    if not repo_name:
        repo_name = f"Repo-{repo_id[:8]}"
    safe_name = repo_name.replace("'", "\\'")
    # Truncate encrypted commit IDs to 40 chars
    head_id_safe = head_id[:40]

    sql_lines.append(f"INSERT IGNORE INTO seafile_db.Repo (repo_id) VALUES ('{repo_id}');")
    sql_lines.append(f"INSERT IGNORE INTO seafile_db.Branch (name, repo_id, commit_id) VALUES ('master', '{repo_id}', '{head_id_safe}');")
    sql_lines.append(f"INSERT IGNORE INTO seafile_db.RepoOwner (repo_id, owner_id) VALUES ('{repo_id}', '{OWNER}');")
    sql_lines.append(f"INSERT IGNORE INTO seafile_db.RepoInfo (repo_id, name, update_time, version, is_encrypted, last_modifier, status) VALUES ('{repo_id}', '{safe_name}', UNIX_TIMESTAMP(), 1, 0, '{OWNER}', 0);")

with open("/tmp/register_repos.sql", "w") as f:
    f.write("\n".join(sql_lines) + "\n")
print(f"Generated {len(sql_lines)-1} statements")

Step 2: Execute the SQL

docker exec -i seafile-mysql mysql -uroot -p'ROOT_PW' < /tmp/register_repos.sql

Step 3: Run seaf-fsck repair

docker exec seafile /opt/seafile/seafile-server-13.0.19/seaf-fsck.sh --repair

For large repos, run in background:

nohup docker exec seafile /opt/seafile/seafile-server-13.0.19/seaf-fsck.sh --repair \
  > /tmp/fsck-output.log 2>&1 &

Monitor progress:

grep -c "Fsck finished" /tmp/fsck-output.log
tail -5 /tmp/fsck-output.log

Step 4: Verify via API

# Get auth token
TOKEN=$(curl -s -X POST http://localhost:80/api2/auth-token/ \
  -d 'username=admin@example.com&password=PASSWORD' | python3 -c "import sys,json; print(json.load(sys.stdin)['token'])")

# List repos
curl -s -H "Authorization: Token $TOKEN" http://localhost:80/api2/repos/ \
  | python3 -c "import sys,json; data=json.load(sys.stdin); print(f'{len(data)} repos visible')"

# Check a specific repo's contents
curl -s -H "Authorization: Token $TOKEN" \
  "http://localhost:80/api2/repos/<repo_id>/dir/" \
  | python3 -c "import sys,json; data=json.load(sys.stdin); print(f'{len(data)} items')"

Pitfalls

  1. Encrypted repo commit IDs have suffixes — Encrypted repos append a suffix like .FYMWR3 to the 40-char SHA1, producing a 46-char string that exceeds the commit_id CHAR(41) column. Error: Data too long for column 'commit_id'. Fix: truncate to 40 chars with head_id[:40].

  2. seaf-fsck.sh --repair only checks REGISTERED repos — It does NOT discover or register new repos from the filesystem. All repos must be inserted into the Repo table first. Same applies to seaf-fsck.sh --export.

  3. Repos invisible without RepoOwner — Even with Repo + Branch entries, repos won't appear in /api2/repos/ without a RepoOwner entry mapping the repo to a user's email.

  4. Repos show generic names without RepoInfo — Without RepoInfo, repos appear as unnamed entries. The name column comes from the commit JSON's repo_name field.

  5. Repo sizes show 0 after reconstruction — Size calculation happens during seaf-fsck or background indexing. Sizes will populate after fsck completes.

  6. Shares, groups, and non-admin users are lost — This procedure restores repo ownership and file access, but share permissions, group memberships, and user accounts (except the admin) must be recreated manually.

  7. Large repos stall fsck — Repos with 30000+ commits can take 10+ minutes during fsck. Run in background and monitor via grep -c "Fsck finished".

  8. CRITICAL: Never run multiple fsck processes simultaneously — If seaf-fsck.sh --repair is already running (check with ps aux | grep seaf-fsck | grep -v grep), do NOT start a second one — even via a different invocation method (e.g., one via nohup on the host and another via docker exec). Two concurrent --repair processes on the same storage can corrupt commit/block data. Always kill -9 the duplicate before continuing. Only one seaf-fsck process should ever be running at a time.

  9. fsck can die silently (OOM-kill)seaf-fsck.sh --repair may be OOM-killed by the kernel with no trace in the log file. The log simply stops at whatever repo was being processed, and ps aux | grep seaf-fsck returns 0 processes. Always check process liveness, not just log tail. If killed, restart fsck — it resumes from where it left off (repos already checked are quickly re-verified).

  10. Large photo repos stall fsck indefinitely — Repos with many large image files (e.g., "Fotos", "Lightroom Bilder") can take 30+ minutes each during fsck. Combined with concurrent Ceph I/O (recovery, RBD copies), fsck throughput drops to near-zero. If fsck appears stuck, check whether heavy I/O is competing and consider pausing other operations.