Mehdi Akiki
Published on

How to Write Better Skill Descriptions So Claude Invokes Them Correctly

Authors
  • Mehdi Akiki avatar
    Name
    Mehdi Akiki
    Twitter

Reference

Most skill failures are not caused by bad instructions in the body.

They are caused by a weak description.

That is the line Claude uses to decide whether a skill is relevant enough to load automatically. Anthropic's Claude Code docs are explicit: by default, Claude can load a skill automatically when it is relevant, and the description field is the thing exposed up front for that decision. Anthropic's broader tool-use guidance makes the same point: descriptions are the most important factor in helping Claude choose the right thing at the right time.

If your description is vague, one of two bad outcomes usually happens:

  • Missed invocations: Claude should have loaded the skill, but did not.
  • False positives: Claude loaded the skill when the request was about something else.

This article is about fixing that.

If you want the full hands-on guide to skill files — frontmatter, arguments, allowed-tools, and more — read the companion guide: Claude Code Skills: A Practical Guide to Writing SKILL.md Files.

Contents

  1. The job of a skill description
  2. The most common mistake: writing descriptions like folder names
  3. Bad vs good descriptions
  4. What strong descriptions usually contain
  5. Use the language people actually type
  6. The three failure modes to avoid
  7. A practical formula for writing descriptions
  8. Good trigger language beats clever wording
  9. Put the distinction in the description, not only in the body
  10. Examples you can reuse
  11. When to add "use when"
  12. When to add "do not use for"
  13. Description patterns that usually work well
  14. Description patterns that usually fail
  15. How to test whether your description is good
  16. A simple rewrite process
  17. One subtle point: concise does not mean underspecified
  18. Final rule of thumb

1. The job of a skill description

A skill description is not a label.

It is not metadata for humans only.

It is an invocation trigger.

Claude uses each skill's description to decide when to load it, so descriptions need to be clear about when to use them. That same principle applies directly to skills.

So the right mental model is:

The description tells Claude what this skill does, when it applies, and what makes it different from nearby skills.

That means a good description must do three things well:

  • describe the action
  • describe the context
  • reduce ambiguity

2. The most common mistake: writing descriptions like folder names

A lot of people write descriptions like this:

description: Deployment helper

or:

description: Code review tool

These are weak because they tell Claude almost nothing useful. They do not say what exactly the skill does, when it should be used, what requests should trigger it, or what makes it different from another similar skill.

Anthropic's docs show much stronger examples, like "Deploy the application to production" or "Read files without making changes", where the behavior is concrete and the context is obvious.


3. Bad vs good descriptions

Here is the difference in practice.

Bad

description: Review code

Too broad, no scope, no intent, could apply to linting, security review, architecture review, or style feedback.

Better

description: Review a pull request or recent code changes for security issues, correctness bugs, performance problems, and style violations.

It names the object (pull request or recent code changes), names the task (review), and names the evaluation dimensions (security, correctness, performance, style). Claude now has a much sharper trigger surface.


Bad

description: GitHub helper

This is almost useless.

Better

description: Work on a GitHub issue by number: read the issue, find the relevant code, implement the fix, and prepare the change for review.

It defines the entry point (issue by number), defines the workflow, and filters out unrelated GitHub tasks like PR triage or repo browsing.


Bad

description: Payments context

Weak background knowledge.

Better

description: Context about the legacy payment system. Load when questions involve billing flows, invoicing, Stripe charge-era APIs, or code under src/payments/.

It explicitly frames it as context, gives trigger phrases, names the code area, and narrows the domain. The "load when" style is especially good for non-user-invocable background skills — it makes the trigger boundary much clearer.


4. What strong descriptions usually contain

The best descriptions are usually built from these parts:

1. The verb — what action is happening?

Examples: review, deploy, generate, troubleshoot, migrate, analyze

2. The object — what is the skill acting on?

Examples: pull request, code changes, GitHub issue, API handlers, billing code, deployment target

3. The scope — in what context should this skill activate?

Examples: recent code changes, staging or production, the legacy payments module, a specific handler directory

4. The differentiator — why this skill instead of another nearby one?

Examples: focuses on security and correctness, read-only exploration only, legacy billing context only, generates docs from code not from guesses


5. Use the language people actually type

This part matters more than people think.

Automatic loading is based on relevance, not only explicit slash commands, which means your description needs to match normal conversational prompts well enough to be recognized as relevant.

If people say:

  • "review this PR"
  • "look over these changes"
  • "can you check this for bugs"
  • "deploy this to staging"
  • "generate docs for these handlers"

then descriptions should sound like that world — not like internal taxonomy invented by the author.

Weak phrasing

description: Perform quality assurance validation of software modifications

Better phrasing

description: Review a pull request or recent code changes for bugs, security issues, and performance problems.

6. The three failure modes to avoid

1. Missed invocation

Happens when the description is too vague or too narrow.

description: Work with APIs

A user asks: "can you generate docs for these handlers?"

Claude may not connect that request strongly enough to the skill.

Better:

description: Generate or update API documentation from route handlers by extracting methods, paths, request shapes, response shapes, and error codes.

2. False positives

Happens when the description is broad enough to match too many requests.

description: Fix code problems

This could wrongly fire for lint cleanups, security review, refactoring, build failures, test failures, or performance analysis.

Better:

description: Diagnose and fix failing test cases by reading the test output, tracing the relevant code path, and making the smallest safe correction.

3. Overlapping skills

The silent killer. If you have two skills like:

description: Review code for problems
description: Check pull requests for issues

you created ambiguity. Claude now has to choose between two fuzzy neighbors.

Better split them by purpose:

description: Review a pull request or recent code changes for security issues, correctness bugs, performance problems, and style violations.
description: Summarize a pull request for reviewers by explaining what changed, why it changed, risks, and test coverage.

Now one is analysis, the other is communication. That separation makes invocation cleaner.


7. A practical formula for writing descriptions

Use this template:

description: [Verb] [object] for [goal/outcome]. Use when [common trigger situations]. Do not use for [nearby but different tasks].

You do not always need every part, but this formula forces clarity.

Example:

description: Review a pull request or recent code changes for security issues, correctness bugs, performance problems, and style violations. Use when asked to review code, check a PR, or look for bugs. Do not use for writing PR summaries or generating documentation.

That is strong because it defines what it does, when it applies, and where its boundary ends.


8. Good trigger language beats clever wording

Do not try to sound smart. Try to sound specific.

Too clever

description: A specialized assistant for code hygiene and software excellence

Useful

description: Review a pull request or recent code changes for bugs, security issues, and style problems.

The second one wins easily. Anthropic's examples consistently prefer concrete descriptions over abstract branding-style phrasing.


9. Put the distinction in the description, not only in the body

A common mistake is thinking: "The body explains everything, so the description can stay short."

That is exactly backwards for automatic invocation.

Claude loads descriptions up front and loads the full skill only when relevant. So if the description is weak, the body may never get loaded at all.

That means the most important disambiguation must happen in the description itself — not buried 40 lines down in the skill body.


10. Examples you can reuse

PR review skill

description: Review a pull request or recent code changes for security issues, correctness bugs, performance problems, and style violations.

PR summary skill

description: Summarize a pull request for reviewers by explaining what changed, why it changed, major risks, and how it was tested.

Deploy skill

description: Deploy the application to staging or production after running the required checks and verifying the target environment.

Note: this one should usually also include disable-model-invocation: true because Anthropic recommends restricting automatic invocation for workflows with side effects like deploys.

Legacy system context skill

description: Context about the legacy payment system. Load when questions involve billing flows, invoicing, payment reconciliation, or code in src/payments/.

Docs generation skill

description: Generate or update API documentation from route handlers by extracting methods, paths, request bodies, response shapes, and error codes.

Read-only exploration skill

description: Explore the codebase in read-only mode to find files, trace call paths, and answer questions without making changes.

This pairs naturally with limited allowed-tools, which Anthropic documents as a way to constrain what Claude can do while the skill is active.


11. When to add "use when"

Add a "use when" phrase when the skill could be confused with another one nearby.

For example:

description: Context about the legacy payment system. Use when questions involve billing, invoices, old Stripe charge flows, or payment code under src/payments/.

This is especially useful for:

  • background context skills
  • team-specific domain knowledge
  • monorepos with several similar modules
  • multiple review/debug skills that could overlap

It is less necessary when the action is already extremely concrete.


12. When to add "do not use for"

Use "do not use for" when you have sibling skills that sit close together.

Example:

description: Diagnose and fix failing tests by tracing failures back to the relevant code and making the smallest safe correction. Do not use for refactoring healthy code or reviewing pull requests.

That extra line can reduce accidental matches.


13. Description patterns that usually work well

Action + object + criteria

description: Review a pull request for bugs, security issues, and performance problems.

Action + object + source

description: Generate API documentation from route handlers and code comments.

Context + trigger domain

description: Context about the legacy payment system. Load when working on billing, invoices, or payment reconciliation.

Action + constraints

description: Explore the codebase in read-only mode to find relevant files and trace logic without making changes.

Action + workflow

description: Work on a GitHub issue by reading the issue, locating the relevant code, implementing the fix, and preparing the change for review.

14. Description patterns that usually fail

One-word abstractions

description: Reviewer
description: Deployment
description: Docs
description: Payments

Internal jargon only

description: Runs SDLC quality gates

Generic helper language

description: A helper for various coding tasks

Marketing language

description: A powerful automation assistant for modern engineering teams

These all fail because they do not define trigger boundaries.


15. How to test whether your description is good

Do not guess. Test it.

You can invoke skills directly and Claude can also load them automatically when relevant. That gives you two distinct tests.

Test 1: direct invoke

Use /skill-name and confirm the skill itself behaves correctly.

Test 2: plain-English invoke

Ask for the same thing without the slash command.

Examples:

  • "can you review this PR?"
  • "please generate docs for these handlers"
  • "help me work on GitHub issue 128"
  • "what is going on in the legacy billing code?"

If it does not load, the description may be too vague.

Test 3: neighboring prompts

Try prompts that should not trigger it. If your review skill fires on "write a PR summary," "prepare release notes," or "explain this architecture," the description is probably too broad.

Test 4: conflict test

If you have two similar skills, try a prompt on the border between them and see whether the wrong one fires. If it does, rewrite both descriptions, not just one.


16. A simple rewrite process

When a skill misses or misfires, do this:

Step 1: find the missing trigger phrase

Ask: what did the user say that the description failed to reflect?

Example — user says: "can you look for security problems in this PR?"

Old description:

description: Review code changes

Rewrite:

description: Review a pull request or recent code changes for security issues, correctness bugs, performance problems, and style violations.

Step 2: add scope

Name the object and domain.

Bad:

description: Generate docs

Better:

description: Generate or update API documentation from route handlers and code comments.

Step 3: remove overlap

If two skills collide, differentiate them by output or workflow.

Bad pair:

description: Review code changes
description: Analyze PRs

Better pair:

description: Review a pull request or code changes for bugs, security issues, and performance problems.
description: Summarize a pull request for human reviewers, including scope, risk, and testing notes.

17. One subtle point: concise does not mean underspecified

Recommendations to keep skills clean and structured do not mean descriptions should be tiny at all costs.

Too short:

description: Review code

Better:

description: Review a pull request or recent code changes for security issues, correctness bugs, performance problems, and style violations.

It is longer, but it is still tight. That is the target.


18. Final rule of thumb

A good skill description should answer this question:

If a teammate read only this one line, would they know when Claude should load the skill and when it should not?

If the answer is no, rewrite it.

Because in practice, that one line is doing most of the routing work.


Treat descriptions as part of the skill logic, not as decoration.

The body tells Claude how to do the work.

The description tells Claude when the work belongs here.

If you get the description wrong, the body never even gets a chance.

For the full guide to skill files — structure, frontmatter, arguments, allowed-tools, and common debugging — read the companion article: Claude Code Skills: A Practical Guide to Writing SKILL.md Files.


Keywords

Claude Code skill description, SKILL.md description field, Claude skill invocation, Claude auto-invoke skill, Claude Code custom skills, how to write skill descriptions, Claude tool description best practices, SKILL.md guide, Claude Code workflow automation, Claude Code slash commands, skill description examples, Claude Code missed invocation, Claude skill false positive, Claude Code developer tools