7.2 KiB
name, description, version, author, license, metadata
| name | description | version | author | license | metadata | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| optimize-loops | Run metric-driven optimization loops for measurable outcomes. Use when improving search relevance, clustering quality, build performance, prompt quality, or scored behavior through experiments. | 1.0.0 | Hermes Agent (merged from Every Inc compound-engineering ce-optimize) | MIT |
|
Iterative Optimization Loops
Run metric-driven iterative optimization. Define a goal, build measurement scaffolding, then run experiments that converge toward the best solution.
Core principle: Don't guess what's better — measure it. Every optimization claim needs a number.
When to Use
- Improving measurable outcomes (build time, latency, relevance scores)
- Tuning prompts, configurations, or algorithms where quality is scored
- Comparing approaches against a metric
- Systematic A/B testing of implementations
Skip for: one-off changes, subjective quality without a scoring method, unmeasurable goals.
The Process
Phase 0: Setup
0.1 Define the Optimization Target
What are you optimizing? Choose the type:
| Type | When to Use | Examples |
|---|---|---|
| Hard metric | Objective, scalar, clear "better" direction | Build time, latency, memory, test pass rate, bundle size |
| Judge metric | Quality requires semantic judgment | Search relevance, clustering quality, summarization, UX copy |
If qualitative: strongly recommend judge type. Hard metrics alone optimize proxy numbers without checking actual quality.
Three-tier approach for judge metrics:
- Degenerate gates (hard, cheap): catch obviously broken solutions
- LLM-as-judge (actual target): sample outputs, score against rubric
- Diagnostics (logged): distribution stats for understanding
0.2 Create the Spec
Define:
- Goal: What to optimize and direction (minimize/maximize)
- Metric: Exact measurement method
- Baseline: Current performance (measure before optimizing!)
- Constraints: Budget, time, dependencies
- Stopping criteria: Max iterations, target threshold, diminishing returns
Save the spec to .context/optimize/<spec-name>/spec.yaml.
Phase 1: Establish Baseline
Measure the current state BEFORE any changes:
# Run the metric on the current implementation
python scripts/measure_baseline.py
Record the baseline. This is your comparison point.
Checkpoint: Write baseline to disk immediately. The conversation is NOT durable storage.
Phase 2: Generate Hypotheses
Brainstorm 3-5 optimization approaches:
- [Approach A] — why it might help, predicted impact
- [Approach B] — why it might help, predicted impact
- [Approach C] — why it might help, predicted impact
Rank by: expected impact × probability of working ÷ cost to test
Phase 3: Run Experiments
For each hypothesis:
3.1 Implement the Change
Make the change in an isolated branch or worktree.
3.2 Measure
Run the metric against the changed implementation.
3.3 Record
Write the result to disk IMMEDIATELY:
# .context/optimize/<spec-name>/experiment-log.yaml
- experiment: 1
hypothesis: "Cache repeated DB queries"
change: "Added Redis cache layer for getUserById"
metric_value: 245 # ms, was 380ms baseline
improvement: 35.5%
verdict: improved
timestamp: 2026-07-10T18:30:00
3.4 Compare
Did it improve over baseline? Over the best so far?
Phase 4: Iterate
After each batch of experiments:
- Read the experiment log from disk (not from memory)
- Identify what worked and what didn't
- Generate new hypotheses informed by results
- Run the next batch
Persistence discipline:
- Write each result to disk IMMEDIATELY after measurement
- Re-read from disk at every phase boundary
- The experiment log is append-only during Phase 3
- Never present results to the user without writing to disk first
Phase 5: Final Summary
After stopping criteria are met:
## Optimization Results
**Goal:** Reduce API latency
**Baseline:** 380ms average
**Best result:** 165ms (57% improvement)
### Winning Approach
[Cached Redis layer for hot queries + connection pooling]
### Experiment History
| # | Approach | Result | Improvement |
|---|---|---|---|
| 1 | Redis cache | 245ms | 35.5% |
| 2 | Connection pool | 310ms | 18.4% |
| 3 | Cache + Pool | 165ms | 56.6% |
| 4 | Query optimization | 290ms | 23.7% |
### Recommendation
Deploy approach #3 (cache + pool). Estimated impact: 57% latency reduction.
Save the final summary to docs/solutions/optimization-YYYY-MM-DD-<topic>.md and invoke compound-learning to capture insights (which will also sync to Hindsight via Phase 4.5).
Using delegate_task for Parallel Experiments
For independent optimization approaches, run experiments in parallel:
delegate_task(tasks=[
{
"goal": "Experiment A: Add Redis caching to getUserById. Measure latency.",
"context": "Baseline: 380ms. Implement caching, run benchmark, report result.",
"toolsets": ["terminal", "file"]
},
{
"goal": "Experiment B: Add connection pooling. Measure latency.",
"context": "Baseline: 380ms. Implement pooling, run benchmark, report result.",
"toolsets": ["terminal", "file"]
},
])
LLM-as-Judge Mode
When the metric is qualitative, use a judge prompt:
delegate_task(
goal="""You are a quality judge. Score the following outputs against the rubric.
Rubric:
- Relevance (0-10): How well does the result match the query intent?
- Accuracy (0-10): Is the information correct?
- Completeness (0-10): Does it cover the expected scope?
Score each output independently. Return JSON:
{
"outputs": [
{"id": 1, "relevance": 8, "accuracy": 9, "completeness": 7, "total": 24},
...
]
}
Outputs to score:
[INSERT SAMPLED OUTPUTS]
"""
)
Pitfalls
- Don't optimize without a baseline — you can't know if you improved
- Don't trust in-memory state — write to disk, re-read at boundaries
- Don't run too many iterations — set stopping criteria upfront
- Don't use hard metrics for qualitative goals — proxy optimization produces degenerate solutions
- Don't skip degenerate gates — "all items in 1 cluster" can score well on a bad metric
- Don't batch results in memory — write each result immediately after measurement
Integration with Other Skills
- plan — create an optimization plan first for complex multi-step optimizations
- subagent-driven-development — use for implementing each experiment
- compound-learning — capture optimization insights after completion
- requesting-code-review — review the winning implementation before deploying
Hermes Agent Integration
terminal— run measurements and benchmarkswrite_file/read_file— persistence (experiment log, spec)delegate_task— parallel experiments and LLM-as-judgesearch_files— find existing optimization patterns
Remember
Measure before optimizing
Write results to disk immediately
Baseline → Hypothesize → Experiment → Record → Iterate
Use judge metrics for qualitative goals
Stop when diminishing returns
Capture the learning