248 lines
9.7 KiB
Markdown
248 lines
9.7 KiB
Markdown
# Seafile Repo Reconstruction from Commit Storage
|
|
|
|
## When to Use
|
|
|
|
After MySQL databases have been wiped/reinitialized but `/opt/seafile-data/`
|
|
(file blocks + commit objects) is intact. This procedure registers all repos
|
|
found in the commit storage back into the fresh MySQL DBs so they appear in the
|
|
web UI and API.
|
|
|
|
## Commit Storage Layout
|
|
|
|
```
|
|
/opt/seafile-data/seafile/seafile-data/storage/commits/
|
|
<repo_id>/
|
|
<first2chars_of_commit_sha>/
|
|
<remaining38chars_of_commit_sha> ← JSON file, commit object
|
|
```
|
|
|
|
Example: commit `5fb48bca960bf13a15df0b84dbee73ef1d8f98d6` lives at:
|
|
`storage/commits/<repo_id>/5f/b48bca960bf13a15df0b84dbee73ef1d8f98d6`
|
|
|
|
## Commit Object Format (JSON)
|
|
|
|
Each commit file is a JSON object:
|
|
|
|
```json
|
|
{
|
|
"commit_id": "5fb48bca960bf13a15df0b84dbee73ef1d8f98d6",
|
|
"root_id": "0000000000000000000000000000000000000000",
|
|
"repo_id": "3428f3b9-d6c0-4d23-8a03-f34ce8b888e1",
|
|
"creator_name": "dominik@familie-schoen.com",
|
|
"creator": "0000000000000000000000000000000000000000",
|
|
"description": "Created library",
|
|
"ctime": 1755169060,
|
|
"parent_id": null,
|
|
"second_parent_id": null,
|
|
"repo_name": "Arztrechnungen",
|
|
"repo_desc": "Arztrechnungen",
|
|
"repo_category": null,
|
|
"no_local_history": 1,
|
|
"version": 1
|
|
}
|
|
```
|
|
|
|
## Finding the HEAD Commit
|
|
|
|
The HEAD commit is the one NOT referenced as `parent_id` (or `second_parent_id`)
|
|
by any other commit in the same repo. Algorithm:
|
|
|
|
1. Read all (or top N by mtime) commit JSON objects for a repo
|
|
2. Collect all `parent_id` and `second_parent_id` values into a set
|
|
3. HEAD = the commit ID(s) NOT in that parent set
|
|
4. If multiple HEADs exist (branch tips), pick the one with the latest `ctime`
|
|
|
|
### Performance for Large Repos
|
|
|
|
Some repos have tens of thousands of commits (34000+). Reading all JSON files
|
|
is too slow (60s+ timeout). Optimization:
|
|
|
|
```bash
|
|
# Get only the 20 newest commit files by filesystem mtime
|
|
find "$repo_dir" -type f -printf '%T@ %p\n' | sort -rn | head -20
|
|
```
|
|
|
|
Then parse only those 20 files. The HEAD is very likely among the newest commits.
|
|
|
|
⚠️ Filesystem mtime reflects when PBS restored the files, NOT the commit creation
|
|
time. Don't rely on mtime alone for determining HEAD — use the parent-reference
|
|
algorithm on the sampled commits.
|
|
|
|
## DB Tables to Populate
|
|
|
|
| Table | DB | Columns | Notes |
|
|
|-------|-----|---------|-------|
|
|
| `Repo` | `seafile_db` | `id` (auto), `repo_id` (unique) | Just insert the repo_id |
|
|
| `Branch` | `seafile_db` | `id` (auto), `name`, `repo_id`, `commit_id` | name='master', commit_id=HEAD |
|
|
| `RepoOwner` | `seafile_db` | `id` (auto), `repo_id` (unique), `owner_id` | owner_id = email address |
|
|
| `RepoInfo` | `seafile_db` | `id` (auto), `repo_id` (unique), `name`, `update_time`, `version`, `is_encrypted`, `last_modifier`, `status` | name from commit JSON `repo_name` |
|
|
|
|
Without `RepoOwner`, repos won't appear in the API (`/api2/repos/`) even if
|
|
registered in `Repo` and `Branch`. Without `RepoInfo`, repos show generic names.
|
|
|
|
## Step-by-Step Procedure
|
|
|
|
### Step 1: Generate SQL for Repo + Branch registration
|
|
|
|
```python
|
|
#!/usr/bin/env python3
|
|
import os, json, subprocess
|
|
|
|
COMMITS_DIR = "/opt/seafile-data/seafile/seafile-data/storage/commits"
|
|
OWNER = "dominik@example.com" # admin email
|
|
|
|
sql_lines = ["-- Register repos"]
|
|
for repo_id in sorted(os.listdir(COMMITS_DIR)):
|
|
repo_dir = os.path.join(COMMITS_DIR, repo_id)
|
|
if not os.path.isdir(repo_dir):
|
|
continue
|
|
# Get newest 20 commits by mtime
|
|
result = subprocess.run(
|
|
["find", repo_dir, "-type", "f", "-printf", "%T@ %p\n"],
|
|
capture_output=True, text=True, timeout=10)
|
|
lines = result.stdout.strip().split('\n') if result.stdout.strip() else []
|
|
lines.sort(key=lambda x: float(x.split()[0]) if x.strip() else 0, reverse=True)
|
|
top_files = [l.split(None, 1)[1] for l in lines[:20] if ' ' in l]
|
|
|
|
commits = {}
|
|
all_parents = set()
|
|
repo_name = None
|
|
latest_ctime = 0
|
|
latest_cid = None
|
|
for fpath in top_files:
|
|
parts = fpath.rstrip('/').split('/')
|
|
commit_id = parts[-2] + parts[-1] # subdir + filename = full SHA
|
|
try:
|
|
with open(fpath, 'r') as f:
|
|
data = json.load(f)
|
|
commits[commit_id] = data
|
|
if data.get('parent_id'):
|
|
all_parents.add(data['parent_id'])
|
|
if data.get('second_parent_id'):
|
|
all_parents.add(data['second_parent_id'])
|
|
if data.get('ctime', 0) > latest_ctime:
|
|
latest_ctime = data['ctime']
|
|
latest_cid = commit_id
|
|
rn = data.get('repo_name')
|
|
if rn:
|
|
repo_name = rn
|
|
except:
|
|
pass
|
|
|
|
head_candidates = [c for c in commits if c not in all_parents]
|
|
head_id = head_candidates[0] if len(head_candidates) == 1 else \
|
|
max(head_candidates, key=lambda c: commits[c].get('ctime', 0)) if head_candidates else latest_cid
|
|
if not head_id:
|
|
continue
|
|
if head_id in commits:
|
|
rn = commits[head_id].get('repo_name')
|
|
if rn:
|
|
repo_name = rn
|
|
if not repo_name:
|
|
repo_name = f"Repo-{repo_id[:8]}"
|
|
safe_name = repo_name.replace("'", "\\'")
|
|
# Truncate encrypted commit IDs to 40 chars
|
|
head_id_safe = head_id[:40]
|
|
|
|
sql_lines.append(f"INSERT IGNORE INTO seafile_db.Repo (repo_id) VALUES ('{repo_id}');")
|
|
sql_lines.append(f"INSERT IGNORE INTO seafile_db.Branch (name, repo_id, commit_id) VALUES ('master', '{repo_id}', '{head_id_safe}');")
|
|
sql_lines.append(f"INSERT IGNORE INTO seafile_db.RepoOwner (repo_id, owner_id) VALUES ('{repo_id}', '{OWNER}');")
|
|
sql_lines.append(f"INSERT IGNORE INTO seafile_db.RepoInfo (repo_id, name, update_time, version, is_encrypted, last_modifier, status) VALUES ('{repo_id}', '{safe_name}', UNIX_TIMESTAMP(), 1, 0, '{OWNER}', 0);")
|
|
|
|
with open("/tmp/register_repos.sql", "w") as f:
|
|
f.write("\n".join(sql_lines) + "\n")
|
|
print(f"Generated {len(sql_lines)-1} statements")
|
|
```
|
|
|
|
### Step 2: Execute the SQL
|
|
|
|
```bash
|
|
docker exec -i seafile-mysql mysql -uroot -p'ROOT_PW' < /tmp/register_repos.sql
|
|
```
|
|
|
|
### Step 3: Run seaf-fsck repair
|
|
|
|
```bash
|
|
docker exec seafile /opt/seafile/seafile-server-13.0.19/seaf-fsck.sh --repair
|
|
```
|
|
|
|
For large repos, run in background:
|
|
```bash
|
|
nohup docker exec seafile /opt/seafile/seafile-server-13.0.19/seaf-fsck.sh --repair \
|
|
> /tmp/fsck-output.log 2>&1 &
|
|
```
|
|
|
|
Monitor progress:
|
|
```bash
|
|
grep -c "Fsck finished" /tmp/fsck-output.log
|
|
tail -5 /tmp/fsck-output.log
|
|
```
|
|
|
|
### Step 4: Verify via API
|
|
|
|
```bash
|
|
# Get auth token
|
|
TOKEN=$(curl -s -X POST http://localhost:80/api2/auth-token/ \
|
|
-d 'username=admin@example.com&password=PASSWORD' | python3 -c "import sys,json; print(json.load(sys.stdin)['token'])")
|
|
|
|
# List repos
|
|
curl -s -H "Authorization: Token $TOKEN" http://localhost:80/api2/repos/ \
|
|
| python3 -c "import sys,json; data=json.load(sys.stdin); print(f'{len(data)} repos visible')"
|
|
|
|
# Check a specific repo's contents
|
|
curl -s -H "Authorization: Token $TOKEN" \
|
|
"http://localhost:80/api2/repos/<repo_id>/dir/" \
|
|
| python3 -c "import sys,json; data=json.load(sys.stdin); print(f'{len(data)} items')"
|
|
```
|
|
|
|
## Pitfalls
|
|
|
|
1. **Encrypted repo commit IDs have suffixes** — Encrypted repos append a suffix
|
|
like `.FYMWR3` to the 40-char SHA1, producing a 46-char string that exceeds
|
|
the `commit_id CHAR(41)` column. Error: `Data too long for column 'commit_id'`.
|
|
Fix: truncate to 40 chars with `head_id[:40]`.
|
|
|
|
2. **`seaf-fsck.sh --repair` only checks REGISTERED repos** — It does NOT
|
|
discover or register new repos from the filesystem. All repos must be inserted
|
|
into the `Repo` table first. Same applies to `seaf-fsck.sh --export`.
|
|
|
|
3. **Repos invisible without `RepoOwner`** — Even with `Repo` + `Branch` entries,
|
|
repos won't appear in `/api2/repos/` without a `RepoOwner` entry mapping the
|
|
repo to a user's email.
|
|
|
|
4. **Repos show generic names without `RepoInfo`** — Without `RepoInfo`, repos
|
|
appear as unnamed entries. The `name` column comes from the commit JSON's
|
|
`repo_name` field.
|
|
|
|
5. **Repo sizes show 0 after reconstruction** — Size calculation happens during
|
|
`seaf-fsck` or background indexing. Sizes will populate after fsck completes.
|
|
|
|
6. **Shares, groups, and non-admin users are lost** — This procedure restores
|
|
repo ownership and file access, but share permissions, group memberships,
|
|
and user accounts (except the admin) must be recreated manually.
|
|
|
|
7. **Large repos stall fsck** — Repos with 30000+ commits can take 10+ minutes
|
|
during fsck. Run in background and monitor via `grep -c "Fsck finished"`.
|
|
|
|
8. **CRITICAL: Never run multiple fsck processes simultaneously** — If
|
|
`seaf-fsck.sh --repair` is already running (check with
|
|
`ps aux | grep seaf-fsck | grep -v grep`), do NOT start a second one — even
|
|
via a different invocation method (e.g., one via `nohup` on the host and
|
|
another via `docker exec`). Two concurrent `--repair` processes on the same
|
|
storage can corrupt commit/block data. Always `kill -9` the duplicate before
|
|
continuing. Only one `seaf-fsck` process should ever be running at a time.
|
|
|
|
9. **fsck can die silently (OOM-kill)** — `seaf-fsck.sh --repair` may be
|
|
OOM-killed by the kernel with no trace in the log file. The log simply
|
|
stops at whatever repo was being processed, and `ps aux | grep seaf-fsck`
|
|
returns 0 processes. Always check process liveness, not just log tail.
|
|
If killed, restart fsck — it resumes from where it left off (repos
|
|
already checked are quickly re-verified).
|
|
|
|
10. **Large photo repos stall fsck indefinitely** — Repos with many large
|
|
image files (e.g., "Fotos", "Lightroom Bilder") can take 30+ minutes
|
|
each during fsck. Combined with concurrent Ceph I/O (recovery, RBD
|
|
copies), fsck throughput drops to near-zero. If fsck appears stuck,
|
|
check whether heavy I/O is competing and consider pausing other
|
|
operations.
|