Decision. For a 3-person team the realistic floor is deterministic, no-model-call checks gated on
.claude/**paths: run the built-inclaude plugin validate --strict[2] on every manifest, JSON-schema-validateplugin.json/marketplace.jsonagainst the now-official SchemaStore schemas [2], lint SKILL.md frontmatter with a sub-second CLI like pulser ⭐ 17 [5] or skill-validator ⭐ 174 [7], and add a “does it load” smoke test by parsing thesystem/initevent fromclaude -p --output-format stream-jsonforplugin_errors[1]. Add at most one cheap model-in-the-loop check (aclaude -pheadless review) only after the deterministic layer is green. Standing up a full eval platform is overkill at this scale, and honestly little is standardized yet — the ecosystem tools are weeks-to-months old and mostly single-maintainer.
The honest state of play (June 2026)
This is an emerging area. Anthropic ships exactly one built-in validator — claude plugin validate — and it is intentionally lenient: unrecognized fields are warnings, not errors, and a plugin with only those warnings still loads at runtime [2]. Everything past that is community tooling, most of it single-maintainer and under a year old. One community audit of 214 published skills claimed 73% had silent structural defects [15] — treat the exact number as marketing, but the direction is real: nobody is linting these files, so cheap deterministic checks have a high catch rate. The good news for a small team: the cheap layer is where almost all the value is, and it needs no model calls, no eval harness, and no spend.
Layer 1 — deterministic checks (no model call, the actual floor)
These run in seconds, cost nothing, and are non-flaky. This is the layer a 3-person team must adopt; the rest is optional.
Built-in: claude plugin validate --strict
Ships with the CLI. Validates plugin.json and marketplace.json structure. Default mode treats unknown fields as warnings and still passes; --strict promotes warnings to errors, which is the documented intent for CI — it catches a misspelled field name or a field left over from another tool’s manifest before publishing [2]. Wrong-typed fields (e.g. keywords as a string instead of an array) fail even without --strict [2].
claude plugin validate ./my-plugin --strict # exit non-zero on any warning
JSON-schema validation of manifests and MCP defs
Official JSON Schemas for plugin.json and marketplace.json are now published on SchemaStore; the manifest’s $schema field is meant to point at https://json.schemastore.org/claude-code-plugin-manifest.json for editor autocomplete and offline validation [2]. The community hesreallyhim/claude-code-json-schema ⭐ 6 repo that pioneered this is now archived precisely because Anthropic’s schemas shipped to SchemaStore [8] — a clean signal that the official path is the one to use. Validate in CI with any standard JSON-schema validator (e.g. ajv-cli, check-jsonschema) pointed at the SchemaStore URL; the same approach covers MCP server entries in .mcp.json.
SKILL.md / agent frontmatter constraints to assert
The skills doc defines the frontmatter contract you can lint against without a model [3]:
| Field | Required | Constraint worth asserting in CI |
|---|---|---|
name |
No | Defaults to directory name; assert it matches dir + has no spaces/path separators [3] |
description |
Recommended | description + when_to_use is truncated at 1,536 chars in the listing — flag overflow [3] |
allowed-tools |
No | Space/comma string or YAML list; assert tool names are real |
disable-model-invocation |
No | boolean; type-check it |
user-invocable |
No | boolean; type-check it |
A plain YAML-parse + required-field + type-check + the 1,536-char cap covers most real defects. The pulser audit’s top finding was exactly this class: unparseable YAML and empty/missing description [6].
“Does it load” smoke test (cheap, one short model call or none)
The headless runner emits a system/init event reporting plugins (loaded) and plugin_errors (each with plugin, type, message, including unsatisfied dependency versions and --plugin-dir load failures). The docs explicitly say: “Use the plugin fields to fail CI when a plugin did not load.” [1] Pattern:
claude --bare -p "ok" --plugin-dir ./my-plugin \
--output-format stream-json --verbose \
| jq -e 'select(.type=="system" and .subtype=="init")
| (.plugin_errors // []) | length == 0'
--bare skips auto-discovery of hooks/skills/MCP/CLAUDE.md so you get the same result on every machine and only what you pass with --plugin-dir/--mcp-config loads [1]. For MCP specifically, --strict-mcp-config loads only the config file you pass and ignores ~/.claude.json, .mcp.json, and managed configs — built for CI [4]. ⚠ Known bug: --strict-mcp-config does not currently override the disabledMcpServers list, so a server disabled in ~/.claude.json stays disabled in CI [4] — pin a clean HOME in the runner to dodge it.
/doctor for local pre-flight
claude doctor checks API connectivity, MCP server health, hook syntax validity, and shell integration, surfacing misconfiguration before a session [13]. It’s interactive-oriented (a developer pre-flight, not a clean CI gate), but worth a mention in the team’s contributor checklist.
Tools comparison (deterministic config/extension linters)
All of these test the artifacts and their config, not model output. Stars are a triage signal — most are very young, single-maintainer projects.
| Tool | ⭐ Stars | What it checks | Runs as | Source |
|---|---|---|---|---|
claude plugin validate --strict (built-in) |
— | plugin.json/marketplace.json structure; strict = warnings-as-errors |
CLI, instant | [2] |
| pulser | ⭐ 17 | 5 checks: YAML parse, required fields, description quality, file structure, cross-refs; ~200ms/40 skills, zero deps | CLI / GitHub Action pulserin/pulser@v1 |
[5] [6] |
| skill-validator | ⭐ 174 | Spec compliance, token counts (o200k_base), link resolution (HEAD), file hygiene, unclosed code fences; optional LLM-judge |
Go CLI / pre-commit / Actions (--emit-annotations, --strict) |
[7] |
| claude-code-json-schema | ⭐ 6 | JSON-schema for manifests — archived, superseded by official SchemaStore schemas | npx github:… / IDE |
[8] |
| agent-skills-lint | ⭐ 1 | Cross-agent skill schema (Claude/Codex/Cursor/…), name+description core + per-flavor rules | npx @swarmclawai/agent-skills-lint |
[9] |
| SkillCheck | — | Validates against the Agent Skills standard; Pro adds anti-slop/WCAG/security | Hosted/web | [16] |
agentskills.io is the cross-tool open standard Claude Code skills follow [3], so validating against it keeps a skill portable to Codex/Cursor/etc. Two takeaways: (1) pick one frontmatter linter — they overlap heavily; skill-validator is the most feature-complete and the only one with built-in CI annotations; pulser is the fastest zero-dep option. (2) For manifests, prefer the built-in validate --strict + SchemaStore combo over any community schema repo.
Layer 2 — cheap model-in-the-loop, runnable per-PR
Only add this once Layer 1 is green; it costs tokens and is non-deterministic. The point is one lightweight headless review, not an eval suite.
-
Headless
claude -pas a project linter.-p/--printruns non-interactively, prints to stdout, and exits with a status code your pipeline can branch on [1]. Pipe the PR diff in (no Bash permission needed) and ask for structured output:git diff origin/main | claude --bare -p \ --append-system-prompt "You review Claude Code skills. Flag skills whose description won't trigger correctly, or that reference missing files." \ --output-format json \ --json-schema '{"type":"object","properties":{"issues":{"type":"array","items":{"type":"string"}}},"required":["issues"]}' \ | jq -e '.structured_output.issues | length == 0'--output-format json+--json-schemareturns schema-conforming output instructured_output, and the payload includestotal_cost_usdso you can budget per run [1]. Cap work with--max-turns. -
anthropics/claude-code-action⭐ 8.1k is the official GitHub Action wrapping the same Agent SDK for PR/issue automation and review [11]. Use it for a@claude reviewstyle pass, not as a gate. -
⚠ Security caveat for any PR-triggered model run. A June 2026 disclosure showed a flaw where a single malicious issue could hijack repositories through the Claude Code GitHub Action [14]. Treat untrusted PR/issue content as a prompt-injection vector: pin action versions, scope
--allowedToolstightly (or run--permission-mode dontAsk), and don’t grant write/secrets to runs triggered by forks.
Concrete GitHub Actions recipes
Deterministic gate (the one every team should copy)
name: Lint extensions
on:
pull_request:
paths: ['.claude/**', '.claude-plugin/**', '.mcp.json']
jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Validate manifests (strict)
run: npx @anthropic-ai/claude-code plugin validate . --strict
- name: Schema-check manifests
run: |
npx ajv-cli validate \
-s https://json.schemastore.org/claude-code-plugin-manifest.json \
-d ".claude-plugin/plugin.json"
- name: Lint skill frontmatter
run: npx pulser eval # or: skill-validator check .claude/skills --strict
The paths: filter is the cost-control lever — skill linting runs only when extension files change, adding ~15s to the pipeline [6]. pulser also ships a one-line Action (uses: pulserin/pulser@v1) with exit 0 = pass, 1 = structural problems, 2 = runtime error [6].
Marketplace-scale precedent
Large community marketplaces already gate manifests this way: the jeremylongshore/claude-code-plugins-plus-skills ⭐ 2.4k marketplace runs claude plugin validate over marketplace.json and every plugin.json in CI so invalid manifests can’t merge [17], and Anthropic’s own claude-plugins-official ⭐ 31k is the reference marketplace.json to diff your schema against [18]. For a starting skill corpus to test tooling against, glebis/claude-skills ⭐ 286 is a clean small collection [19].
The pragmatic minimum bar for a 3-person team
In priority order — stop wherever the budget for maintaining checks runs out:
- CI path filter on
.claude/**+ manifests so checks run only on extension changes [6]. (5 min) claude plugin validate --stricton every manifest [2]. (built-in, free)- One frontmatter linter — pulser or skill-validator — as a required check [5] [7]. (free, sub-second)
- JSON-schema validate manifests/MCP defs against SchemaStore [2]. (free)
- “Does it load” smoke test via
claude -p --output-format stream-json→ assert emptyplugin_errors[1]. (one trivial run) - (Optional, last) One headless
claude -pmodel review of the diff with--json-schemastructured output [1], behind the security guardrails above [14].
Steps 1–5 are deterministic, free, and non-flaky — that is the floor, and it’s enough. Step 6 is the only place you spend tokens, and a small team should keep it to a single advisory check rather than a gate. Anything beyond this (golden datasets, LLM-as-judge rubrics, statistical delta gates) is the eval-platform territory this brief deliberately leaves out — note that skill-validator’s score evaluate LLM-judge mode [7] is the one place a lightweight tool crosses into that space if you ever want it.