Files
hermes-skills/software-development/optimize-loops/SKILL.md
T

7.2 KiB
Raw Blame History

name, description, version, author, license, metadata
name description version author license metadata
optimize-loops Run metric-driven optimization loops for measurable outcomes. Use when improving search relevance, clustering quality, build performance, prompt quality, or scored behavior through experiments. 1.0.0 Hermes Agent (merged from Every Inc compound-engineering ce-optimize) MIT
hermes
tags related_skills
optimization
experimentation
metrics
iteration
compound-engineering
plan
subagent-driven-development
requesting-code-review

Iterative Optimization Loops

Run metric-driven iterative optimization. Define a goal, build measurement scaffolding, then run experiments that converge toward the best solution.

Core principle: Don't guess what's better — measure it. Every optimization claim needs a number.

When to Use

  • Improving measurable outcomes (build time, latency, relevance scores)
  • Tuning prompts, configurations, or algorithms where quality is scored
  • Comparing approaches against a metric
  • Systematic A/B testing of implementations

Skip for: one-off changes, subjective quality without a scoring method, unmeasurable goals.

The Process

Phase 0: Setup

0.1 Define the Optimization Target

What are you optimizing? Choose the type:

Type When to Use Examples
Hard metric Objective, scalar, clear "better" direction Build time, latency, memory, test pass rate, bundle size
Judge metric Quality requires semantic judgment Search relevance, clustering quality, summarization, UX copy

If qualitative: strongly recommend judge type. Hard metrics alone optimize proxy numbers without checking actual quality.

Three-tier approach for judge metrics:

  1. Degenerate gates (hard, cheap): catch obviously broken solutions
  2. LLM-as-judge (actual target): sample outputs, score against rubric
  3. Diagnostics (logged): distribution stats for understanding

0.2 Create the Spec

Define:

  • Goal: What to optimize and direction (minimize/maximize)
  • Metric: Exact measurement method
  • Baseline: Current performance (measure before optimizing!)
  • Constraints: Budget, time, dependencies
  • Stopping criteria: Max iterations, target threshold, diminishing returns

Save the spec to .context/optimize/<spec-name>/spec.yaml.

Phase 1: Establish Baseline

Measure the current state BEFORE any changes:

# Run the metric on the current implementation
python scripts/measure_baseline.py

Record the baseline. This is your comparison point.

Checkpoint: Write baseline to disk immediately. The conversation is NOT durable storage.

Phase 2: Generate Hypotheses

Brainstorm 3-5 optimization approaches:

  1. [Approach A] — why it might help, predicted impact
  2. [Approach B] — why it might help, predicted impact
  3. [Approach C] — why it might help, predicted impact

Rank by: expected impact × probability of working ÷ cost to test

Phase 3: Run Experiments

For each hypothesis:

3.1 Implement the Change

Make the change in an isolated branch or worktree.

3.2 Measure

Run the metric against the changed implementation.

3.3 Record

Write the result to disk IMMEDIATELY:

# .context/optimize/<spec-name>/experiment-log.yaml
- experiment: 1
  hypothesis: "Cache repeated DB queries"
  change: "Added Redis cache layer for getUserById"
  metric_value: 245  # ms, was 380ms baseline
  improvement: 35.5%
  verdict: improved
  timestamp: 2026-07-10T18:30:00

3.4 Compare

Did it improve over baseline? Over the best so far?

Phase 4: Iterate

After each batch of experiments:

  1. Read the experiment log from disk (not from memory)
  2. Identify what worked and what didn't
  3. Generate new hypotheses informed by results
  4. Run the next batch

Persistence discipline:

  • Write each result to disk IMMEDIATELY after measurement
  • Re-read from disk at every phase boundary
  • The experiment log is append-only during Phase 3
  • Never present results to the user without writing to disk first

Phase 5: Final Summary

After stopping criteria are met:

## Optimization Results

**Goal:** Reduce API latency
**Baseline:** 380ms average
**Best result:** 165ms (57% improvement)

### Winning Approach
[Cached Redis layer for hot queries + connection pooling]

### Experiment History
| # | Approach | Result | Improvement |
|---|---|---|---|
| 1 | Redis cache | 245ms | 35.5% |
| 2 | Connection pool | 310ms | 18.4% |
| 3 | Cache + Pool | 165ms | 56.6% |
| 4 | Query optimization | 290ms | 23.7% |

### Recommendation
Deploy approach #3 (cache + pool). Estimated impact: 57% latency reduction.

Save the final summary to docs/solutions/optimization-YYYY-MM-DD-<topic>.md and invoke compound-learning to capture insights (which will also sync to Hindsight via Phase 4.5).

Using delegate_task for Parallel Experiments

For independent optimization approaches, run experiments in parallel:

delegate_task(tasks=[
    {
        "goal": "Experiment A: Add Redis caching to getUserById. Measure latency.",
        "context": "Baseline: 380ms. Implement caching, run benchmark, report result.",
        "toolsets": ["terminal", "file"]
    },
    {
        "goal": "Experiment B: Add connection pooling. Measure latency.",
        "context": "Baseline: 380ms. Implement pooling, run benchmark, report result.",
        "toolsets": ["terminal", "file"]
    },
])

LLM-as-Judge Mode

When the metric is qualitative, use a judge prompt:

delegate_task(
    goal="""You are a quality judge. Score the following outputs against the rubric.

Rubric:
- Relevance (0-10): How well does the result match the query intent?
- Accuracy (0-10): Is the information correct?
- Completeness (0-10): Does it cover the expected scope?

Score each output independently. Return JSON:
{
  "outputs": [
    {"id": 1, "relevance": 8, "accuracy": 9, "completeness": 7, "total": 24},
    ...
  ]
}

Outputs to score:
[INSERT SAMPLED OUTPUTS]
"""
)

Pitfalls

  • Don't optimize without a baseline — you can't know if you improved
  • Don't trust in-memory state — write to disk, re-read at boundaries
  • Don't run too many iterations — set stopping criteria upfront
  • Don't use hard metrics for qualitative goals — proxy optimization produces degenerate solutions
  • Don't skip degenerate gates — "all items in 1 cluster" can score well on a bad metric
  • Don't batch results in memory — write each result immediately after measurement

Integration with Other Skills

  • plan — create an optimization plan first for complex multi-step optimizations
  • subagent-driven-development — use for implementing each experiment
  • compound-learning — capture optimization insights after completion
  • requesting-code-review — review the winning implementation before deploying

Hermes Agent Integration

  • terminal — run measurements and benchmarks
  • write_file / read_file — persistence (experiment log, spec)
  • delegate_task — parallel experiments and LLM-as-judge
  • search_files — find existing optimization patterns

Remember

Measure before optimizing
Write results to disk immediately
Baseline → Hypothesize → Experiment → Record → Iterate
Use judge metrics for qualitative goals
Stop when diminishing returns
Capture the learning