Published on

Claude Code for CI and Automation: Headless Workflows, GitHub Actions, and Reusable Skills

Authors
  • Mehdi Akiki avatar
    Name
    Mehdi Akiki
    Twitter

Most developers first meet Claude Code in the terminal.

That is not where it gets strategically interesting.

The real leverage shows up when you stop treating Claude Code as a chat tool and start treating it as automation infrastructure: something that runs non-interactively in CI, operates inside GitHub workflows, and reuses the same project knowledge through CLAUDE.md, hooks, and SKILL.md files. Anthropic now documents this programmatic path explicitly — Claude Code can run non-interactively with claude -p, and the same agent loop is available through the Claude Agent SDK for CLI, Python, and TypeScript.

This is where teams start asking the serious questions:

  • Can we use Claude Code in pull request review?
  • Can it triage issues or propose fixes in CI?
  • Can we keep it safe?
  • Can we make it reusable instead of prompt spaghetti?

Yes — but only if you separate deterministic automation from judgment-heavy work, and only if you are strict about permissions. Anthropic's own docs reflect that split: hooks are for deterministic lifecycle automation, skills are reusable task instructions, and the Agent SDK adds explicit permission controls, approval callbacks, and structured programmatic execution.

Contents

  1. What "headless" means now
  2. Where Claude Code actually fits in CI
  3. GitHub Actions: the practical route most teams start with
  4. Reusable skills are where maintainability comes from
  5. Skills, hooks, and CLAUDE.md do different jobs
  6. Safe patterns for CI automation
  7. Where human approval still matters
  8. A good CI architecture for Claude Code
  9. Use cases that actually make sense
  10. What not to automate first
  11. CLI vs Agent SDK vs GitHub Actions
  12. The professional mindset

1. What "headless" means now

Anthropic's terminology has shifted a bit: the old "headless mode" is now documented as running Claude Code programmatically. The CLI path is still the same idea. You run Claude Code non-interactively with -p / --print, and you can keep using CLI options such as --continue, --allowedTools, and --output-format.

A simple example:

claude -p "Review the changed files and list possible security issues" \
  --allowedTools "Read,Grep,Glob"

That turns Claude Code into something CI can call directly. Anthropic explicitly says the Agent SDK is available as a CLI for scripts and CI/CD, while Python and TypeScript packages are there when you need deeper programmatic control.

So the stack is roughly this:

  • CLI (claude -p) for scripts, CI jobs, and quick automation
  • Agent SDK when you need structured outputs, native message objects, approval callbacks, or custom runtime control
  • GitHub Actions / GitLab CI/CD integrations when you want workflow-native automation on top of that execution model

2. Where Claude Code actually fits in CI

A lot of people get this wrong.

Claude Code is not best used as a blind, all-powerful bot that edits anything on every push. The better pattern is to use it in a few narrow, high-signal places.

PR review assistance

One of the cleanest uses. Claude inspects diffs, looks for security issues, correctness risks, missing tests, or style violations, and leaves a structured review summary. Claude Code's docs position CI as a place for automated review and issue triage, and GitHub Actions support is now part of the official product surface.

Issue-to-PR automation

Stronger than review, but still reasonable if you scope it hard. Anthropic's GitHub Actions docs say Claude Code can respond to @claude mentions in issues or PRs, analyze code, implement features, fix bugs, and open pull requests while following repository standards.

Triage and summarization

Low risk and often high value. Good examples:

  • summarize failing test output
  • cluster flaky test patterns
  • explain what changed in a PR
  • suggest likely owners or affected modules

Deterministic enforcement around AI work

This is where hooks matter. Hooks execute shell commands at defined lifecycle points and give deterministic control over Claude Code's behavior — they are better than "please remember to format" because they enforce formatting, validation, or protection rules mechanically.


3. GitHub Actions: the practical route most teams start with

Anthropic now has an official Claude Code GitHub Actions flow. Their docs describe it as GitHub automation where an @claude mention in a PR or issue can trigger analysis, implementation, or PR creation. It is built on top of the Claude Agent SDK, which means the GitHub integration is not some separate magic product — it sits on the same underlying automation model.

Key points from Anthropic's docs:

  • it follows your project's CLAUDE.md guidance and code patterns automatically
  • your code stays on GitHub's runners
  • it defaults to Sonnet, with Opus 4.6 configurable via the model parameter

That last point gives you a sane default for routine automation and a heavier option for more complex reasoning.

A basic PR review workflow:

name: Claude Code PR Review

on:
  pull_request:
    types: [opened, synchronize]

jobs:
  claude-review:
    runs-on: ubuntu-latest
    permissions:
      contents: read
      pull-requests: write

    steps:
      - name: Checkout
        uses: actions/checkout@v4
        with:
          fetch-depth: 0

      - name: Install Claude Code
        run: npm install -g @anthropic-ai/claude-code

      - name: Get changed files
        id: changed
        run: |
          echo "files=$(git diff --name-only origin/${{ github.base_ref }}...HEAD | tr '\n' ' ')" >> $GITHUB_OUTPUT

      - name: Run Claude review
        id: review
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
        run: |
          claude -p \
            --allowedTools "Read,Grep,Glob" \
            --output-format json \
            "Review these changed files for security issues, correctness bugs, and notable risks: ${{ steps.changed.outputs.files }}. Output a brief summary with severity labels." \
          > review_output.json

      - name: Post review comment
        uses: actions/github-script@v7
        with:
          script: |
            const fs = require('fs');
            const output = JSON.parse(fs.readFileSync('review_output.json', 'utf8'));
            await github.rest.issues.createComment({
              owner: context.repo.owner,
              repo: context.repo.repo,
              issue_number: context.issue.number,
              body: `## Claude Code Review\n\n${output.result}`
            });

This is read-only: Claude analyzes and posts a comment. It never writes to the repository.


4. Reusable skills are where maintainability comes from

Without reusable skills, CI automation degrades into long prompts pasted into YAML files. That does not scale.

Anthropic's skills system exists precisely to avoid that. Skills are reusable Markdown-defined capabilities added through SKILL.md, and Claude can use them when relevant or invoke them directly. Anthropic's docs explicitly call out task-oriented skills such as deployments, commits, or code generation, and recommend disable-model-invocation: true for actions you want invoked explicitly rather than automatically.

For CI, this is a strong pattern. Instead of burying a giant review policy in GitHub Actions YAML, define a reusable skill:

---
name: ci-review
description: Review a pull request diff for security, correctness, and test coverage issues.
disable-model-invocation: true
allowed-tools: Read,Grep,Glob
---

Review the code changes in $ARGUMENTS.

Output format (strict):
## Security
- [SEVERITY] Finding description (file:line)

## Correctness
- [SEVERITY] Finding description (file:line)

## Performance
- [SEVERITY] Finding description (file:line)

If no issues found in a category, write "None found."
Do not ask clarifying questions. Work with what is available.

Why this is better than inline YAML prompts:

  • the workflow stays small
  • the prompt is versioned with the repo
  • engineers can reuse the same logic locally and in CI
  • the standards become visible and editable in one place

That is the difference between "we have AI in CI" and "we have an automation system we can maintain."

For the full skill authoring guide, see Claude Code Skills: A Practical Guide to Writing SKILL.md Files.


5. Skills, hooks, and CLAUDE.md do different jobs

A lot of confusion comes from mixing these together.

LayerPurpose
CLAUDE.mdalways-on project rules and context
SKILL.mdreusable task workflows
hooksdeterministic enforcement at lifecycle events

Use CLAUDE.md for always-on project rules

Anthropic treats CLAUDE.md as project context and instructions. This is where stable repository-wide rules belong: coding conventions, test commands, architecture constraints, or "never touch this directory without approval."

Use skills for reusable task workflows

Skills are for "do this kind of job the same way every time": review a diff, prepare a release note, check migration safety, generate API docs.

Use hooks for deterministic enforcement

Hooks are the part that should not depend on the model remembering. Anthropic says hooks execute shell commands at lifecycle points for deterministic control. Examples: formatting after edits, blocking edits to protected files, notifications, validation.

The rule:

  • Policy / context → CLAUDE.md
  • Reusable workflow → SKILL.md
  • Hard guardrail or always-run automation → hooks

If you respect that split, your automation stays understandable. The CLAUDE.md vs SKILL.md guide goes deeper on this distinction.


6. Safe patterns for CI automation

This is the part people try to skip. Do not.

Anthropic's security docs say Claude Code uses strict read-only permissions by default, requires approval for higher-risk actions like bash execution, and confines writes to the folder where it was started unless explicitly permitted. The Agent SDK documents a full permission evaluation flow: hooks → deny rules → permission mode → allow rules → runtime callback (canUseTool).

Pattern 1: read-only review jobs

Best starting point. Allow only read/search tools.

claude -p "Review the diff for security issues" \
  --allowedTools "Read,Grep,Glob"

Use this for: PR review, architecture checks, docs coverage, change summaries.

Pattern 2: controlled edit jobs on isolated branches

Allow edits, but only in a branch created for automation, and require human review before merge.

claude -p "Fix the failing test in $FILE" \
  --allowedTools "Read,Grep,Glob,Write"

Good for: typo fixes, test additions, issue-to-PR flows.

Pattern 3: hard deny dangerous tools

Anthropic's SDK docs are clear that deny rules are checked before allow rules and even override bypassPermissions. Use this for unrestricted shell use, secrets tools, or infrastructure-adjacent commands.

Pattern 4: human approval for external side effects

If a workflow posts to GitHub, touches Jira, deploys, or modifies production-adjacent infrastructure, keep a human in the loop. The SDK supports approval callbacks and runtime permission decisions — that is the right place to gate risky actions.

On prompt injection

If your CI step includes PR title, PR body, issue text, or commit message content directly in the Claude prompt, a malicious actor can inject instructions. Sanitize or isolate user-controlled content. Never interpolate untrusted text directly into the system prompt. This is the most underappreciated security risk in Claude Code CI integration.


7. Where human approval still matters

This is the line teams need to draw clearly.

Deployments

You can define deployment-like skills — Anthropic uses deployment as a canonical skill example category. But production deploys are exactly the kind of side-effect-heavy operation that should require explicit invocation and, usually, a human checkpoint. Use disable-model-invocation: true.

Cross-system updates

If Claude is touching GitHub, Slack, Jira, internal docs, and cloud infrastructure in one flow, you want approvals and explicit boundaries, not blind execution.

Security-sensitive edits

Authentication, billing, access control, infra, secrets handling. Claude can help. These are still review-first areas.

Large refactors with hidden blast radius

Claude can propose them. It should not silently land them at scale without branch isolation and review.

The rule:

Let Claude automate analysis aggressively. Let it automate edits selectively. Let it automate irreversible side effects only behind explicit approvals.


8. A good CI architecture for Claude Code

A sane architecture usually looks like this:

Layer 1: repository memory

CLAUDE.md for persistent repo rules: how to run tests, coding standards, migration rules, prohibited areas, expected PR format.

Layer 2: reusable skills

A small set of narrow skills: ci-review, test-triage, release-notes, docs-check, migration-safety.

Layer 3: hooks

Deterministic rules: auto-format after edits, block writes to protected files, audit config changes, re-run certain checks.

Layer 4: CI workflows

GitHub Actions or direct claude -p jobs triggered on the right event: PR opened, PR comment with @claude, issue labeled, nightly maintenance job, release branch updates.

Layer 5: human gates

Required review before: merge, deploy, secret-related changes, risky shell execution, external system writes.

That is the grown-up setup. Not vibes, not giant prompts, not one magical bot.


9. Use cases that actually make sense

PR reviewer

Claude inspects only changed files and produces: must-fix issues, likely regressions, missing tests, optional cleanups. Keep it read-only.

Test failure interpreter

When CI fails, Claude summarizes: what failed, most likely root cause, whether failure looks flaky or infra-related, candidate files to inspect next.

Issue-to-draft-PR assistant

Triggered from @claude on an issue. Claude creates a branch, implements a narrow fix, runs tests, and opens a draft PR. Anthropic's GitHub Actions docs explicitly describe this issue and PR mention-based flow.

Release note generator

Claude reads commits or PR labels and drafts release notes in a standard structure.

Docs drift checker

Claude compares changed APIs or flags against docs and reports likely documentation gaps.

Dependency review

Claude reads newly added dependencies, summarizes what they do, their license, and any obvious security flags. Posted as a PR comment.

These are useful because they either reduce human reading time or automate a bounded slice of toil.


10. What not to automate first

Do not start here:

  • "Claude merges approved PRs automatically"
  • "Claude deploys production after tests pass"
  • "Claude rewrites half the repo every night"
  • "Claude has unrestricted bash and network access in CI"

That is how teams create a scary demo instead of a reliable system.

Start with: read-only review, summaries, triage, narrow PR creation behind explicit triggers.

Then expand only after the permission story is clean.


11. CLI vs Agent SDK vs GitHub Actions

Use claude -p when:

  • you want the fastest path
  • your workflow is mostly prompt in, result out
  • shell and CI scripting are enough
  • structured output needs are modest

Anthropic explicitly documents claude -p for non-interactive use with options like --allowedTools and --output-format.

Use the Agent SDK when:

  • you need runtime approval control
  • you want structured programmatic interaction
  • you need native message objects
  • you want Python or TypeScript integration
  • you need callbacks like canUseTool

Anthropic's docs say the SDK adds structured outputs, tool approval callbacks, and native message objects.

Use GitHub Actions when:

  • you want workflow-native triggers
  • you want issue and PR comment integrations
  • your team already lives in GitHub
  • you want repository-scoped automation with familiar CI surfaces

Anthropic's official GitHub Actions docs place Claude directly in that loop.


12. The professional mindset

The amateur instinct is: "Can Claude do everything in CI?"

The professional question is: "Which pieces should be deterministic, which should be agentic, and where do approvals belong?"

Anthropic's docs already imply the right answer:

  • hooks for deterministic lifecycle behavior
  • skills for reusable task instructions
  • permissions and callbacks for tool control
  • CI/CD and GitHub workflows for orchestration
  • project memory for durable repo context

Claude Code is now clearly more than an interactive terminal assistant. Anthropic documents it as a programmable agent system that can run via CLI, Python, TypeScript, GitHub Actions, and GitLab CI/CD while reusing repository context and reusable skills.

The winning strategy is not maximal autonomy. It is disciplined automation:

  • keep repository rules in CLAUDE.md
  • encode repeatable workflows as skills
  • enforce non-negotiables with hooks
  • start with read-only or tightly scoped CI jobs
  • put humans in front of risky side effects

That is how you make Claude Code useful in professional automation without turning your CI pipeline into a slot machine.

For more on the underlying building blocks:


Keywords

Claude Code CI, Claude Code GitHub Actions, Claude Code headless mode, Claude Code automation, Claude Code non-interactive, Claude Code pipeline, Claude Code PR review automation, Claude Code allowedTools, Claude Code print flag, automated code review Claude, Claude Code skills CI, Claude Code security review automation, Claude Code workflow automation, Claude Code Agent SDK, headless Claude Code workflow, Claude Code hooks CI, Claude Code team DevOps

I build and scale reliable production systems. Open to full-time and freelance work with U.S.-based teams that value ownership and execution.

Got something in mind?

Book a Discovery Call