microsoft/SkillOpt

▲ 111 stars today★ 17,790⑂ 1,665

SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.

About microsoft/SkillOpt

microsoft/SkillOpt is an open-source project on GitHub, mainly written in Python. SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits It currently holds 17,790 stars and 1,665 forks with 54 open issues, and was last pushed on 2026-09-05 (repository created 2026-05-08).

Project Overview

Git Homed tracks it on the Today's Trending board, currently at rank #35 with 111 new stars today.

GitHub Repository Details

Repository microsoft/SkillOpt · default branch main · size 24524 KB · watchers 69 · source: GitHub REST API and repository README

README

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

Train agent skills like you train neural networks — with epochs, (mini-)batchsize, learning rates, and validation gates — but without touching model weights.

Project Page Paper Project Video PyPI Python 3.10+ License: MIT

https://github.com/microsoft/SkillOpt/blob/HEAD/microsoft%2FSkillOpt | Trendshift https://github.com/microsoft/SkillOpt/blob/HEAD/microsoft%2FSkillOpt | Trendshift

📖 For installation, data preparation, training/eval commands, configuration, and framework internals, start with the versioned SkillOpt documentation. A concise rendered overview is available in the Documentation & Reproduction Guide, and longer-form engineering analysis appears on the Technical Blog. We also maintain a Changelog for released and unreleased changes.

---

News 🔥🔥🔥

---

Overview

Modern agent skills are usually hand-crafted, generated one-shot by a strong LLM, or evolved through loosely controlled self-revision — none of which behaves like a deep-learning optimizer for the skill itself, and none of which reliably improves over its starting point under feedback.

SkillOpt treats the skill document as the trainable state of a frozen agent, and trains it with the discipline that makes weight-space optimization reproducible. A separate optimizer model turns scored rollouts into bounded add / delete / replace edits on a single skill document; in the default paper-style path, a candidate edit is accepted only when it strictly improves a held-out validation score. A textual learning-rate budget, a rejected-edit buffer, and an epoch-wise slow / meta update make skill training stable while adding zero inference-time model calls at deployment.

The deployed artifact is a compact best_skill.md (typically 300–2,000 tokens) that runs against the unchanged target model. Across six benchmarks, seven target models, and three execution harnesses (direct chat, Codex CLI, Claude Code CLI), SkillOpt is best or tied-best on all 52 evaluated (model, benchmark, harness) cells and on GPT-5.5 lifts the average no-skill accuracy by +23.5 points in direct chat, +24.8 inside the Codex agentic loop, and +19.1 inside Claude Code. Optimized skill artifacts transfer across model scales, between Codex and Claude Code harnesses, and to nearby benchmarks without further optimization.

For the full method, ablations, and per-cell results see the paper; for a visual walkthrough of the loop see the project page; for deeper API / backend / benchmark docs see docs/.

🎬 Demo Video

https://github.com/user-attachments/assets/eb12d3bc-371c-467f-904d-91b61f339ed7

▶ Watch the full demo on YouTube

---

Extensibility & WebUI

Adding a new backend

A backend = a chat / exec target (e.g. openai_chat, claude_chat, qwen_chat, minimax_chat, copilot_chat, openai_compatible, codex_exec, claude_code_exec, cursor_exec, copilot_exec). If a provider implements the OpenAI Chat Completions protocol, try the built-in openai_compatible backend before adding code. See docs/guide/new-backend.md for the full contract. Chat backends add a skillopt/model/_backend.py module; target-only exec backends use the shared harness in codex_harness.py. Both register through common.py, backend_config.py, and skillopt/model/__init__.py.

Adding a new benchmark

A benchmark = a skillopt/envs// package with an adapter, a data loader, a scored rollout helper, a YAML config, and optionally an initial seed skill. See docs/guide/new-benchmark.md for the full contract; the simplest reference is skillopt/envs/searchqa/.

WebUI

Launch the monitoring dashboard (optional):

pip install -e ".[webui]"
python -m skillopt_webui.app

| Flag | Default | Description | |---|---|---| | --port | 7860 | Server port | | --host | 0.0.0.0 | Bind address | | --share | off | Create a public Gradio share link |

The default host listens on every network interface. Use --host 127.0.0.1 for local-only access.

---

Citation

@article{yang2026skillopt,
  title={Skillopt: Executive strategy for self-evolving agent skills},
  author={Yang, Yifan and Gong, Ziyang and Huang, Weiquan and Yang, Qihao and Zhou, Ziwei and Huang, Zisu and Li, Yan and Gao, Xuemei and Dai, Qi and Liu, Bei and others},
  journal={arXiv preprint arXiv:2605.23904},
  year={2026}
}

GitHub Stars & Activity

17,790Stars
1,665Forks
54Open issues
PythonLanguage

GitHub Popularity

GitHub stars17,790
Forks1,665
Open issues54
Primary languagePython
LicenseMIT
Stars gained today111
Created2026-05-08
Last pushed2026-09-05

Trending History

Daily boardrank #35 · ▲ 111 stars

Related GitHub Projects

1

practical-tutorials / project-based-learning

Python★ 285,110⑂ 36,444▲ 215 stars
→
2

rohitg00 / ai-engineering-from-scratch

Python★ 60,276⑂ 10,377▲ 1,329 stars
→
3

debpalash / VoiceStudio

Python★ 43,522⑂ 5,059▲ 3,274 stars
→
4

vectorize-io / hindsight

Python★ 40,775⑂ 5,521▲ 4,413 stars
→
5

topoteretes / cognee

Python★ 31,135⑂ 3,115▲ 103 stars
→
6

alirezarezvani / claude-skills

Python★ 26,762⑂ 3,775▲ 149 stars
→
7

agent0ai / agent-zero

Python★ 19,330⑂ 3,823▲ 22 stars
→
8

0x4m4 / hexstrike-ai

Python★ 12,216⑂ 2,486▲ 56 stars
→

More Trending Repositories