For goals without natural scalar metric (docs quality, code health, consistency), construct an explicit runnable fitness function plus dual-score (outcome + instrument quality guard) and drive an improvement loop with iterations.jsonl ledger until converge criteria.
/goal GOAL: Complete Fitness Function Dual-Score Improvement Loop for an eval-backed prompt project: For goals without natural scalar metric (docs quality, code health, consistency), construct an explicit runnable fitness function plus dual-score (outcome + instrument quality guard) and drive an improvement loop with iterations.jsonl ledger until converge criteria. CONTEXT: - Before editing, read the nearest AGENTS.md/CLAUDE.md, current issue or PLAN.md, and any failing logs already in the repo. - Inspect prompt files, eval cases, scoring reports, regressions, and failure examples. - Establish a baseline by running or locating evidence for: `./scripts/score.sh --json outputs numeric scores; iterations.jsonl shows progress without instrument gaming; final report matches When to Stop template`. CONSTRAINTS: - Keep the scope limited to this goal; do not expand into unrelated cleanup. - Do not weaken tests, delete assertions, or mask errors to make verification pass. - Respect the repository's AGENTS.md/CLAUDE.md instructions and existing patterns. - Do not delete, weaken, or cherry-pick eval cases to improve the score. - Report representative failures as well as the final score. DONE WHEN: - The implementation or documentation directly satisfies: For goals without natural scalar metric (docs quality, code health, consistency), construct an explicit runnable fitness function plus dual-score (outcome + instrument quality guard) and drive an improvement loop with iterations.jsonl ledger until converge criteria. - The verification command or evidence path succeeds: `./scripts/score.sh --json outputs numeric scores; iterations.jsonl shows progress without instrument gaming; final report matches When to Stop template`. - The final diff is scoped to the relevant files and has no unrelated formatting churn. VERIFY: - Run `./scripts/score.sh --json outputs numeric scores; iterations.jsonl shows progress without instrument gaming; final report matches When to Stop template` or the closest repo-local equivalent if the exact command is not available. - Capture before/after evidence for the behavior, metric, report, or artifact involved. - If verification cannot run locally, stop and report the missing dependency instead of guessing success. OUTPUT: - Summarize changed files, key decisions, verification output, and remaining risks. - Include any follow-up that is required for production rollout or human review. STOP RULES: - Pause if secrets, production access, stakeholder decisions, or destructive data operations are required. - Pause after three failed fix attempts on the same symptom and challenge the root-cause hypothesis. - Do not mark the goal complete until the current repository state has been audited against DONE WHEN.
原始来源: goal-md template
证据摘要: dual-score split + measure/diagnose/act/verify/revert loop + Action Catalog + machine-checkable stopping conditions; source: goal-md template; type: tool-readme; verification: ./scripts/score.sh --json outputs numeric scores; iterations.jsonl shows progress without instrument gaming; final report matches When to Stop template