How to Create an Agent Skill

A skill is a folder with a Markdown file in it. That's genuinely the whole format — which is why most guides finish in two minutes and leave you with a skill that doesn't work properly. The hard parts are what goes in the description, what stays out of the main file, and knowing when you wanted something else entirely. If the concept itself is new to you, read agent skills first.

In this guide
  1. Before you start: is a skill the right container?
  2. How to build one, step by step
  3. The description is the whole selection surface
  4. Keep the body to what shapes a decision
  5. Mistakes worth avoiding
  6. When not to write one
  7. FAQ

Before you start: is a skill the right container?

Instructions can live in four places, and they fail differently. Picking wrong is more expensive than writing the skill badly.

Put it inWhen
PromptIt applies to every request. A rule that should never not apply gains nothing from loading conditionally.
SkillIt applies sometimes, the instructions are long, and the agent already has the tools to carry them out.
ToolThe agent needs a new capability, not new guidance. No amount of instruction lets it query your database if nothing exposes the database.
SubagentThe work needs its own context, its own tool permissions, or a result instead of a transcript.

The full distinctions have their own pages: Skill vs Tool, Skill vs Agent, Skill vs Subagent.

How to build one, step by step

  1. Make a folder named in kebab-case — review-pull-request, not ReviewPR. In Claude Code it goes in .claude/skills/ for one project, or ~/.claude/skills/ to have it everywhere.
  2. Create SKILL.md inside it. The filename is case-sensitive and has to be exactly that.
  3. Write the YAML frontmatter. Two fields are required — name and description. Anthropic's rules: the name is lowercase letters, numbers and hyphens, up to 64 characters, and can't contain "anthropic" or "claude"; the description is up to 1024 characters and can't be empty.
  4. Write the body as instructions to follow, not documentation to read. Say what to do, in order, and explain why where the reason affects the judgment.
  5. Move reference material into separate files in the same folder, linked from SKILL.md. These stay on disk until something actually needs them.
  6. Test it by asking for the task in your own words — not by naming the skill. If it doesn't trigger, the description is wrong, and no amount of fixing the body will help.

            review-pull-request/
            ├── SKILL.md           ← only name + description load up front
            └── reference/
                └── checklist.md   ← read only when SKILL.md points at it

            --- SKILL.md ---
            ---
            name: review-pull-request
            description: Review a pull request for correctness
              and test coverage. Use when asked to review a PR
              or check a diff before merging.
            ---

            # Reviewing a pull request
            ...instructions...
            For the checklist, read reference/checklist.md.
            

The description is the whole selection surface

This is the single most important thing on this page. By default only each skill's name and description sit in context — around 100 tokens each in Anthropic's implementation. The body stays on disk until something matches that description.

So the description is not a summary you write afterwards. It's the only thing the agent compares your request against, and anything the body can do that the description doesn't mention is unreachable.

A worked example from this site's own tooling. The skill that writes these pages gained a review path — hand it a finished page instead of authoring a new one. The instructions went in the body. The description still said "Write or edit a page." Asking it to review a page matched nothing, so the capability sat unused until the description named it. The body was never the problem.

Two rules follow. Say what the skill does and when to use it — Anthropic's authoring guidance is explicit about the second half, and it's the half people skip. And write descriptions that exclude each other: once two could plausibly match one request, the agent guesses, and it will sometimes guess wrong.

Keep the body to what shapes a decision

Everything in SKILL.md loads every time the skill fires, so a skill used often pays for its own length on every run. Bundled files cost nothing until read, and a bundled script costs even less — it runs and only its output enters context, never its code.

The useful split isn't by length, it's by role. Guidance that shapes a decision has to be loaded before the decision is made. Reference material consulted mid-task does not.

Measured on the page-writing skill mentioned above: 6,898 words in SKILL.md against 1,886 across three reference files, so about 79% sat in the always-loaded path. Roughly a quarter of that was lookup — file paths, frontmatter keys, markup syntax — needed only after the real decisions were already made. Anthropic suggests keeping the body under 500 lines. Treat that as an alarm rather than a limit: the honest test is whether a section changes what the agent chooses, or only tells it where something lives.

Mistakes worth avoiding

Nesting references. Keep bundled files one level from SKILL.md. A file that points at a file that points at a file tends not to get read all the way down.

Drifting terminology. Pick one word for each thing and keep it. A skill that calls the same object a "page", a "doc" and an "article" makes the agent work out whether those are three things.

Writing rules without reasons. Capitalised MUSTs are brittle — they tell the agent what, never why, so it can't generalise to the case you didn't anticipate. Explaining the reason usually costs one clause and survives situations the rule didn't cover.

Letting status claims go stale. A skill that says "X isn't built yet" keeps saying it long after X ships, because whatever built X didn't touch the skill file. If you write one of those sentences, expect to be wrong about it later.

When not to write one

Skip it for a genuinely one-off task, or an instruction that fits in a sentence you could just type.

And watch for the case that looks like a skill but isn't. If the work needs different tool permissions than whatever runs it — a review that should read but never write — that's a subagent. A skill inherits whatever access its caller already has. That's also why installing one from an untrusted source is a real risk: it can't grant itself new access, but it can direct the access you already granted, which is why Anthropic's guidance is to treat it like installing software and audit the bundled files first.

FAQ

Can a skill make an agent do something it couldn't do before?

Not in terms of access. A skill supplies instructions and bundled files, and it acts entirely through tools the agent already has. It changes competence, not permission — a missing capability calls for a tool, a missing instinct calls for a skill.

How many skills is too many?

There's no hard ceiling, because only each one's name and description stay in context. The practical limit is overlap: once two descriptions could match the same request, the agent is guessing. Prune by distinctness, not by count.

Do I need a special prompt or tool to write a skill?

No. Current Claude models know the format, so asking for a skill produces valid frontmatter and structure without any setup. What they can't do for you is know which parts of your process actually matter — that judgment is the work, and it's why a generated first draft is a starting point rather than a finished skill.