Files

231 lines
7.2 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: optimize-loops
description: "Run metric-driven optimization loops for measurable outcomes. Use when improving search relevance, clustering quality, build performance, prompt quality, or scored behavior through experiments."
version: 1.0.0
author: Hermes Agent (merged from Every Inc compound-engineering ce-optimize)
license: MIT
metadata:
hermes:
tags: [optimization, experimentation, metrics, iteration, compound-engineering]
related_skills: [plan, subagent-driven-development, requesting-code-review]
---
# Iterative Optimization Loops
Run metric-driven iterative optimization. Define a goal, build measurement scaffolding, then run experiments that converge toward the best solution.
**Core principle:** Don't guess what's better — measure it. Every optimization claim needs a number.
## When to Use
- Improving measurable outcomes (build time, latency, relevance scores)
- Tuning prompts, configurations, or algorithms where quality is scored
- Comparing approaches against a metric
- Systematic A/B testing of implementations
**Skip for:** one-off changes, subjective quality without a scoring method, unmeasurable goals.
## The Process
### Phase 0: Setup
#### 0.1 Define the Optimization Target
What are you optimizing? Choose the type:
| Type | When to Use | Examples |
|---|---|---|
| **Hard metric** | Objective, scalar, clear "better" direction | Build time, latency, memory, test pass rate, bundle size |
| **Judge metric** | Quality requires semantic judgment | Search relevance, clustering quality, summarization, UX copy |
**If qualitative:** strongly recommend `judge` type. Hard metrics alone optimize proxy numbers without checking actual quality.
Three-tier approach for judge metrics:
1. **Degenerate gates** (hard, cheap): catch obviously broken solutions
2. **LLM-as-judge** (actual target): sample outputs, score against rubric
3. **Diagnostics** (logged): distribution stats for understanding
#### 0.2 Create the Spec
Define:
- **Goal:** What to optimize and direction (minimize/maximize)
- **Metric:** Exact measurement method
- **Baseline:** Current performance (measure before optimizing!)
- **Constraints:** Budget, time, dependencies
- **Stopping criteria:** Max iterations, target threshold, diminishing returns
Save the spec to `.context/optimize/<spec-name>/spec.yaml`.
### Phase 1: Establish Baseline
Measure the current state BEFORE any changes:
```bash
# Run the metric on the current implementation
python scripts/measure_baseline.py
```
Record the baseline. This is your comparison point.
**Checkpoint:** Write baseline to disk immediately. The conversation is NOT durable storage.
### Phase 2: Generate Hypotheses
Brainstorm 3-5 optimization approaches:
1. [Approach A] — why it might help, predicted impact
2. [Approach B] — why it might help, predicted impact
3. [Approach C] — why it might help, predicted impact
Rank by: expected impact × probability of working ÷ cost to test
### Phase 3: Run Experiments
For each hypothesis:
#### 3.1 Implement the Change
Make the change in an isolated branch or worktree.
#### 3.2 Measure
Run the metric against the changed implementation.
#### 3.3 Record
Write the result to disk IMMEDIATELY:
```yaml
# .context/optimize/<spec-name>/experiment-log.yaml
- experiment: 1
hypothesis: "Cache repeated DB queries"
change: "Added Redis cache layer for getUserById"
metric_value: 245 # ms, was 380ms baseline
improvement: 35.5%
verdict: improved
timestamp: 2026-07-10T18:30:00
```
#### 3.4 Compare
Did it improve over baseline? Over the best so far?
### Phase 4: Iterate
After each batch of experiments:
1. Read the experiment log from disk (not from memory)
2. Identify what worked and what didn't
3. Generate new hypotheses informed by results
4. Run the next batch
**Persistence discipline:**
- Write each result to disk IMMEDIATELY after measurement
- Re-read from disk at every phase boundary
- The experiment log is append-only during Phase 3
- Never present results to the user without writing to disk first
### Phase 5: Final Summary
After stopping criteria are met:
```markdown
## Optimization Results
**Goal:** Reduce API latency
**Baseline:** 380ms average
**Best result:** 165ms (57% improvement)
### Winning Approach
[Cached Redis layer for hot queries + connection pooling]
### Experiment History
| # | Approach | Result | Improvement |
|---|---|---|---|
| 1 | Redis cache | 245ms | 35.5% |
| 2 | Connection pool | 310ms | 18.4% |
| 3 | Cache + Pool | 165ms | 56.6% |
| 4 | Query optimization | 290ms | 23.7% |
### Recommendation
Deploy approach #3 (cache + pool). Estimated impact: 57% latency reduction.
```
Save the final summary to `docs/solutions/optimization-YYYY-MM-DD-<topic>.md` and invoke `compound-learning` to capture insights (which will also sync to Hindsight via Phase 4.5).
## Using delegate_task for Parallel Experiments
For independent optimization approaches, run experiments in parallel:
```python
delegate_task(tasks=[
{
"goal": "Experiment A: Add Redis caching to getUserById. Measure latency.",
"context": "Baseline: 380ms. Implement caching, run benchmark, report result.",
"toolsets": ["terminal", "file"]
},
{
"goal": "Experiment B: Add connection pooling. Measure latency.",
"context": "Baseline: 380ms. Implement pooling, run benchmark, report result.",
"toolsets": ["terminal", "file"]
},
])
```
## LLM-as-Judge Mode
When the metric is qualitative, use a judge prompt:
```python
delegate_task(
goal="""You are a quality judge. Score the following outputs against the rubric.
Rubric:
- Relevance (0-10): How well does the result match the query intent?
- Accuracy (0-10): Is the information correct?
- Completeness (0-10): Does it cover the expected scope?
Score each output independently. Return JSON:
{
"outputs": [
{"id": 1, "relevance": 8, "accuracy": 9, "completeness": 7, "total": 24},
...
]
}
Outputs to score:
[INSERT SAMPLED OUTPUTS]
"""
)
```
## Pitfalls
- **Don't optimize without a baseline** — you can't know if you improved
- **Don't trust in-memory state** — write to disk, re-read at boundaries
- **Don't run too many iterations** — set stopping criteria upfront
- **Don't use hard metrics for qualitative goals** — proxy optimization produces degenerate solutions
- **Don't skip degenerate gates** — "all items in 1 cluster" can score well on a bad metric
- **Don't batch results in memory** — write each result immediately after measurement
## Integration with Other Skills
- **plan** — create an optimization plan first for complex multi-step optimizations
- **subagent-driven-development** — use for implementing each experiment
- **compound-learning** — capture optimization insights after completion
- **requesting-code-review** — review the winning implementation before deploying
## Hermes Agent Integration
- `terminal` — run measurements and benchmarks
- `write_file` / `read_file` — persistence (experiment log, spec)
- `delegate_task` — parallel experiments and LLM-as-judge
- `search_files` — find existing optimization patterns
## Remember
```
Measure before optimizing
Write results to disk immediately
Baseline → Hypothesize → Experiment → Record → Iterate
Use judge metrics for qualitative goals
Stop when diminishing returns
Capture the learning
```