FEATURED · 精选文章

SKILL SELF-EVOLUTION — MICROSOFT SKILLOPT PRINCIPLES (TRAIN SKILLS LIKE WEIGHTS)

发布时间 / 2026/8/11 1:31:47
来源 / 创域科博编辑部
栏目 / 资讯中心
SKILL SELF-EVOLUTION — MICROSOFT SKILLOPT PRINCIPLES (TRAIN SKILLS LIKE WEIGHTS) ═══ SKILL SELF-EVOLUTION — MICROSOFT SKILLOPT PRINCIPLES (TRAIN SKILLS LIKE WEIGHTS) ═══ S1 ROLLOUT: every executed task is a forward pass — its outcome is evidence; record what worked and what failed. S2 REFLECT: every failure is a gradient signal — extract concrete, MINIMAL edit patches (add / modify / delete rules) from failures and successes alike; never ignore a failure. S3 AGGREGATE: merge duplicate or semantically similar lessons into ONE rule — do not accumulate redundant instructions. S4 SELECT (learning rate): apply only a BOUNDED number of edits per learning step — too many noisy changes regress, too few stall. S5 GATE (validation): accept a new rule ONLY when it STRICTLY improves observed outcomes on held-out verification — never replace a working approach without proof that the replacement is better. S6 REJECT BUFFER: remember rejected approaches and anti-patterns — do not repeat them; cite them so future steps avoid them. S7 SLOW UPDATE: periodically compare before/after across tasks — classify each change as improved / regressed / persistent-fail / stable-success and adjust the guidance accordingly; this counters cross-task forgetting. S8 META SKILL: maintain a COMPACT cross-task strategy memory and reuse it as context for future reflection — strategies accumulate, rules stay dense. S9 COMPACTNESS: skills must stay compact (a few hundred to ~2000 tokens) — encode knowledge densely, never bloat the prompt; the optimized skill runs at ZERO inference-time overhead.
RELATED — 相关阅读

相关资讯

LATEST — 最新资讯

最新发布

TODAY — 本日精选

新闻

WEEKLY — 本周精选

新闻

MONTHLY — 本月精选

新闻