Back to marketplace

dsh-continual-evolve

Agent & Workflow

ZK-Andy/dsh-continual-evolve

A continual self-evolution plugin for DeepSeek Harness with versioned, auditable, rollback-safe harness state refined from session trajectories and a benchmark-driven validation loop.

  • ai-agent
  • deepseek-harness
  • deepseek-harness-plugin
  • dsh
  • dsh-plugin
  • dsh-plugins
  • self-evolving-agents
  • typescript
GitHub Stars
14GitHub
Views
0DSH Plugin Hub
Forks
1GitHub
Open issues
0GitHub Issues
Manifest version
0.3.0dsh-continual-evolve
Latest push
Aug 21, 2026GitHub
License
MITTypeScript
Plugin type
HostRuns in the DSH Host

README

View source

dsh-continual-evolve

中文 | English

awesome · DSH plugin npm CI License: MIT Node Tests

Continual self-evolution for DeepSeek Harness: a versioned, auditable, rollback-safe harness state layer — prompt notes, memories, skills, subagent specs — refined from session trajectories.

The model proposes, the code guarantees. Every mechanical safety property — schema validation, atomic writes, snapshots, versioning, audit trail, acceptance decisions — is enforced in code, never by prompt discipline.

Why

Agents accumulate reusable experience (repeated failures, durable facts, reusable procedures) and forget it next session. This plugin turns that experience into first-class state:

  • Local scope per session; global scope across sessions with merge semantics — plus mechanical promotion guards so only portable, substantial, non-duplicate knowledge reaches global
  • Deterministic rollback: inverse edits generated from applied results — no LLM re-guessing
  • Benchmark loop: candidate refinements are evaluated against frozen cases by a separate scorer before acceptance (rubric encrypted at rest)

How it works

  1. Sediment — the model creates entries via evolve_add, or the automatic review gate proposes them from the session trajectory (turn-interval + compaction checkpoints).
  2. Guard — code-enforced validation: edit schema, blast-radius/scope coherence, and the promotion policy (project-scoped markers, thin content, near-duplicate detection keep the global store clean).
  3. Approve — global writes require explicit human approval; local-fate proposals are consulted before they land.
  4. Apply & inject — atomic apply with snapshot + audit event. Prompt notes and delegation specs inject into the system prompt (capped, relevance-ranked, zero tokens when empty); memories/skills appear as a capped directory index.
  5. Validate & roll back — benchmarks score candidates against frozen cases; rejected candidates roll back deterministically.

Install

# from npm (installs and activates — ships its own bundle patch)
dsh plugin add dsh-continual-evolve

# or from source (first GitHub installs require approving the allowBuilds step)
dsh plugin add ZK-Andy/dsh-continual-evolve

Restart dsh web after installing or updating.

Usage

Commands (in-session):

CommandEffect
/evolvehelp + current local store
/evolve list · history · rollback <id>inspect and revert (add global for the cross-session store)
/evolve plan [msg]run the LLM planner against the store
/evolve wrapupassess this session's local entries: promote / archive / keep
/evolve archive · unarchive · demote <id>hide from injection (data kept, restorable) — demote targets global noise
/evolve failuresaggregated failure classes (gate + benchmark)
/evolve log [tail N] [session <id>]plugin log
/evolve export · import <path>backup / restore a store
/evolve mount · unmount <skillId>hot-mount an executable skill as a live plugin
/evolve goal [objective · done · block]round-driven auto-review goal
/evolve benchmark …case lifecycle, runs, acceptance

Model tools: evolve_list / add / update / delete / rollback.

Injection shape: prompt notes and delegation specs inject with content (≤6/kind × 180 chars, relevance-ranked). Memories and skills appear as a directory index ([kind:id] title, capped at 15 lines with a fold counter) — full text via evolve_list. Empty store = zero injected tokens.

Configuration

KeyDefaultMeaning
baseDirresolved DSH homeroot for the evolve/ stores
autoReviewfalseenable the automatic review gate
reviewIntervalTurns6gate cadence on the turn-interval path
maxReviewInputChars40000trajectory slice handed to the gate
reviewBudgetTokens4096output budget for the gate call
notifyOnAutoReviewtruevisible follow-up notice after an applied gate run
requireGlobalApprovaltrueglobal edits ask for explicit approval
localFatetruegate audits local entries and proposes promote/archive (consulted, never silent)
fateIntervalTurnsfollows reviewIntervalTurnsminimum turns between fate assessments
goalBlockedWrapupTurns3consecutive blocked-goal gate runs trigger one fate assessment (0 disables)
promotionBlockPatternsPOSIX paths, session ids, ~/.dshcontent matching these is project-scoped and never promoted to global
promotionMinChars100whole promotions below this length stay local
injectionDirectoryLines15entry-directory lines per build before folding into a counter
sectionOrder118system-prompt section order
skillsDir<dshHome>/skillswhere skill entries materialize as SKILL.md bundles
rubricKeyauto-generated key fileAES-256-GCM passphrase for benchmark rubrics (DSH_EVOLVE_RUBRIC_KEY overrides)
logToFile / logLevel / logMaxBytestrue / 1 / 5 MiBplugin-owned JSONL file log with rotation
autoRollbackOnRejecttruedeterministic rollback after a benchmark rejection
reviewModelagent's ownoptional cheaper model for the gate ("provider/model")

Example profile patch:

- id: continual-evolve
  config:
    autoReview: true
    reviewIntervalTurns: 6

Development

pnpm install && pnpm build   # deps + tsc -> lib/
pnpm test                    # vitest (527 tests)
pnpm test:coverage           # v8 coverage, thresholds enforced in CI
pnpm lint                    # oxlint src test

Project layout:

├── src/                   # engine, tools, commands, gate, fate, benchmark, usage…
├── test/                  # vitest suites (33 files)
├── lib/                   # build output (tsc)
├── docs/
│   ├── design.md          # full design doc (hardening matrix)
│   ├── FAQ.md             # real failure/fix records
│   ├── gap-analysis.md    # vs prime-agent /refine + penguin-harness
│   ├── experiment-bootstrap.md
│   ├── archive/           # closed point-in-time reports
│   └── research/          # penguin report + prime-agent annotated source
├── examples/README.md     # seed benchmark cases
└── .agents/               # AI collaboration layer (AGENTS.md, skills, ADR notes)

Docs & provenance

License

MIT

Comments

0
Newest first