← Verified Skills

Our approach / Evidence before assumptions

A useful skill earns
its place in your context.

Skills package instructions, resources, and tools for a specific job. Their value comes from what they add to that job—not the length of the prompt or the size of a collection.

01 / Know the source.

Start with the publisher and repository. Inspect the skill’s instructions, supporting files, dependencies, and version history. Check what it will read, execute, or send to another service.

Provenance establishes where an artifact came from. It does not establish whether the author is correct, whether the instructions suit your project, or whether the skill will improve an outcome.

Explore publishers ↗

02 / Read the security evidence.

The registry exposes scan findings and verification tiers. Available checks include pattern analysis, blocklist checks, and analysis of suspicious intent. Inspect the actual findings and the version they describe.

A badge is not a guarantee of safety. Checks can miss harmful behavior or flag legitimate workflows. A later version, different tool permissions, or a changed execution environment can change the risk.

Security status and evaluation results answer different questions. A skill can pass a scan and still be irrelevant, incorrect, or less effective than using the model alone.

Open the trust center ↗

Read the scanner’s security guidelines for the checks and reporting process.

03 / Compare the outcome.

Begin with representative tasks and clear success conditions. Compare the same setup with and without the skill. Include tasks where it should activate and tasks where it should stay out of the way.

  • Keep the task, starting files, tools, and permissions consistent.
  • Record the exact skill version and model. Record the harness, effort setting, and provider where available.
  • Grade the resulting work. Use executable checks where possible and inspect subjective judgments.
  • Repeat trials before drawing strong conclusions. Inspect failures as well as passes.
  • Compare elapsed time and available usage alongside quality. Missing telemetry is unknown, not zero.

These are evaluation practices, not a claim that every catalog entry has been benchmarked on every harness. Results apply to the tested conditions. Observed usage from unrelated tasks cannot establish which model or harness is better.

Evaluate in Skill Studio ↗

Keep the toolkit small.

Retain hard-won domain knowledge, reliable scripts, and precise workflows. Revisit broad triggers, generic coaching, duplicate instructions, and elaborate procedures when the model or harness changes.

Prefer project-scoped instructions for project-specific work. Review version changes before updating. Remove a skill when its maintenance and context cost exceed the value it adds.

Manage your installed skills ↗

Grounded in current practice.

OpenAI recommends concise skill descriptions, contextual guidance, and progressive disclosure in Rethinking skills and prompts for GPT-6 Astra.

Anthropic distinguishes the model from the harness and emphasizes outcomes, controlled trials, and appropriate graders in Demystifying evals for AI agents.

The Agent Skills specification defines a portable format. Compatibility with a format does not establish equal behavior across agents.