Best agent skills for security review of code, CI and dependencies
The best agent skill for a security review of your code is Sentry's security-review, which reports a vulnerability only after tracing attacker-controlled input to it. For a review of one pull request with a written report, pick Trail of Bits' differential-review. We compared 7 skills from Sentry, Trail of Bits and Firebase on how they keep false alarms out, the evidence each finding needs, what they cover, the agents their publishers document, and license terms. Facts checked on 2026-10-01.
| Name | Best for | Made by | License | Notes | Checked |
|---|---|---|---|---|---|
| Security review | Security review of code with few false alarms | Sentry | CC-BY-SA-4.0 | 1,031 stars | |
| Differential review | Security review of one pull request with a written report | Trail of Bits | CC-BY-SA-4.0 | 7,328 stars | |
| GitHub Actions security review | GitHub Actions workflows open to fork pull requests | Sentry | Apache-2.0 | 1,031 stars | |
| Supply chain risk auditor | Dependency risk on npm, PyPI and Go | Trail of Bits | CC-BY-SA-4.0 | 7,328 stars | |
| Firebase security rules auditor | Firestore and Cloud Storage security rules | Firebase (Google) | Apache-2.0 | 462 stars | |
| Semgrep rule creator | Turning a found bug into a tested Semgrep rule | Trail of Bits | CC-BY-SA-4.0 | 7,328 stars | |
| Code review | General pull request review with security as one check | Sentry | Apache-2.0 | 1,031 stars |
security-review: confirmed vulnerabilities only
security-review is Sentry's agent skill for finding exploitable vulnerabilities in code, for developers who want a security pass without a flood of false alarms. Its strength is the confidence rule: a finding is reported only when a vulnerable pattern meets input an attacker controls, medium-confidence issues are marked as needing verification, and theoretical ones are dropped. The agent must trace data flow across the whole codebase before reporting, and the skill lists what not to flag, such as values from settings or environment variables and auto-escaped React, Vue or Django templates. Topic guides cover injection, XSS, SSRF, crypto and more, with language guides for Python, JavaScript, Go, Rust and Java. Its limit is the other side of that rule: by design it can miss issues it could not confirm, and its examples lean on Django, since Sentry says its skills were written for Sentry employees. getsentry/skills had 1,030 GitHub stars on 2026-10-01; the folder last changed on 2026-09-28. See Security review.
differential-review: one pull request, a written report
differential-review is Trail of Bits' agent skill for a security review of a pull request, commit or diff, for teams that need findings on record. Its strength is evidence: it rates each change by risk rather than size, with authentication, crypto, external calls and removed validation as high risk, runs git blame on removed security code to catch reintroduced bugs, counts callers to measure blast radius, checks test coverage of the changed lines, and always writes a markdown report with line numbers and commits. Review depth scales with the codebase, from reading everything under 20 files to critical paths only above 200. Its limit is weight and scope: seven phases is a lot for a two-line change, and it reviews changes, not a whole codebase. Its optional attacker-modeling subagent ships only in the Claude Code plugin; other agents follow a guide file in the folder instead. trailofbits/skills had 7,320 GitHub stars on 2026-10-01. See Differential review.
gha-security-review: attacks on your GitHub Actions
gha-security-review is Sentry's agent skill for auditing GitHub Actions workflows, for maintainers of repositories that accept pull requests from forks. Its strength is a tight threat model: an outsider without write access who can open fork pull requests, file issues and post comments, so anything that needs write access is out of scope. It checks eight problem classes, among them pull_request_target workflows that run fork code, PR titles or branch names expanded inside run steps, unpinned third-party actions in privileged jobs, and AGENTS.md or CLAUDE.md files that an AI agent in CI could read from a pull request. Every high-confidence finding must name the entry point, payload, execution mechanism, impact and a proof-of-concept sketch. Its limit is that it covers workflows only, never application code. Its attack patterns draw on a 2025 StepSecurity analysis of a real campaign. Apache-2.0; the folder last changed on 2026-09-29. See GitHub Actions security review.
supply-chain-risk-auditor: dependency risk, measured by scripts
supply-chain-risk-auditor is Trail of Bits' agent skill for auditing a project's npm, PyPI and Go dependencies, for teams that want a dependency risk baseline before a release or an audit. Its strength is that two bundled Python scripts do the measuring, not the agent: they check version-matched advisories, archived or abandoned upstreams, how few people can publish an npm package, and install-time scripts, marking each criterion clean, flagged or unassessable with a reason. Trail of Bits built it this way after hand-collected figures proved wrong, such as GitHub contributor counts suggesting five maintainers where npm showed one. Its limit is coverage: it reads package-lock, uv and Go module lockfiles but not yarn.lock, pnpm-lock.yaml or poetry.lock, skips Rust and Maven, and without a signed-in GitHub CLI it gets 60 GitHub requests an hour, so repository checks come back unassessable. It needs uv to run. The folder last changed on 2026-09-16. See Supply chain risk auditor.
firebase-security-rules-auditor: red-team your Firestore rules
firebase-security-rules-auditor is the Firebase team's agent skill for attacking Firestore and Cloud Storage security rules, for apps whose clients read and write Firestore directly. Its strength is a six-point checklist aimed at the bugs that matter there: can a user create a valid document and then update it into an invalid one, do rules trust user-supplied fields such as role or isAdmin, are string and array sizes limited, are types checked, and do hasOnly or diff checks also verify who is writing. The output is JSON with a score from 1, critical, to 5, secure, so you can track it across releases. Its limit is depth: the SKILL.md is about 540 words, with no scripts and no emulator tests, and its examples are about Firestore although Cloud Storage is in scope. It is the only skill here whose publisher documents Gemini CLI. firebase/agent-skills had 461 GitHub stars on 2026-10-01. See Firebase security rules auditor.
semgrep-rule-creator: make a found bug stay found
semgrep-rule-creator is Trail of Bits' agent skill for writing custom Semgrep rules, for teams that found a bug once and want CI to catch every future instance. Its strength is test-first discipline: the agent writes vulnerable and safe test cases before the rule, reads Semgrep's syntax tree dump, iterates until semgrep --test passes, and only then may simplify the patterns. It prefers taint mode for injection bugs, so a rule fires only when untrusted data reaches the dangerous call rather than on every call with constant input. Its limit is that it does not review code itself: it produces one rule and one test file at a time, needs the Semgrep CLI installed, and points to Trail of Bits' separate static-analysis plugin for running existing rulesets. It pairs with security-review or differential-review: they find the issue, this turns it into a check. The folder last changed on 2026-09-16. See Semgrep rule creator.
code-review: security as one item on a wider checklist
code-review is Sentry's agent skill for general pull request review by Sentry's engineering practices, for teams that want security checked alongside everything else. Its strength is breadth in little space: runtime errors, N+1 queries and quadratic loops, side effects, breaking API changes, Django ORM cost, security issues such as injection, XSS, access control gaps and exposed secrets, test quality, and a rule for when a change needs a senior reviewer, such as schema or API contract changes. It also asks for polite, specific feedback and approval when only minor issues remain. Its limit for security work is depth: the whole skill is about 370 words, with no confidence rules and no exploit evidence, so it does not replace security-review. Its examples are Python with Django and TypeScript with React. The folder last changed on 2026-04-19, the oldest of the seven. Apache-2.0. See Code review.
What the licenses let you do
Four of the seven skills carry a share-alike content license. Trail of Bits publishes all its skills under CC-BY-SA-4.0. Sentry's repository is Apache-2.0, but the security-review folder has its own CC-BY-SA-4.0 LICENSE, because its reference material is adapted from the OWASP Cheat Sheet Series. Installing and using these skills unchanged does not put your code under that license. Copying and changing them is allowed, even commercially, with credit, a link to the license, a note of what you changed, and the same license on any changed version you share. Because CC-BY-SA-4.0 is a content license rather than a software license, ask your legal team before bundling one inside a closed product. gha-security-review, code-review and the Firebase auditor are Apache-2.0.
How we compared
We read the SKILL.md and folder of each of the 7 skills on GitHub on 2026-10-01, plus the security-review LICENSE file and the getsentry/skills README. We ranked by five criteria in this order: how the skill keeps false alarms out, the evidence it requires for each finding, what it covers, which agents the publisher documents, and license terms. Stars and last-change dates come from the GitHub API on the same day. We did not run any skill on a real codebase, did not count true or false positives, and did not compare these skills with commercial scanners.
How to choose
- If you want a general security pass on application code, use security-review; use code-review instead only when security is one concern among many.
- If you need a security verdict on one pull request with a record of how it was reached, use differential-review, then have semgrep-rule-creator turn any bug it finds into a CI rule.
- If your repository accepts fork pull requests, uses pull_request_target, or runs an AI agent in CI, run gha-security-review on your workflows.
- If your risk sits in npm, PyPI or Go dependencies with a package-lock, uv or go.mod lockfile, use supply-chain-risk-auditor; for pnpm or Yarn lockfiles it falls back to pinned or latest versions.
- If your app reads and writes Firestore from the client, run firebase-security-rules-auditor before every rules change ships.
Questions people ask
What is the best agent skill for a security review of code?
Sentry's security-review, for a general pass: it reports a vulnerability only when it has traced attacker-controlled input to a vulnerable pattern, and drops theoretical issues. For one pull request with an audit trail, Trail of Bits' differential-review writes a markdown report with line numbers and commits behind each finding.
Can I use the Sentry and Trail of Bits security skills in a commercial project?
Yes. Using them unchanged does not affect your code's license. Trail of Bits' skills and the reference material of Sentry's security-review are CC-BY-SA-4.0: you may copy and change them, even commercially, if you credit the source, link the license, note your changes and share any changed version under the same license. gha-security-review, code-review and the Firebase auditor are Apache-2.0.
Do these security skills work in Codex, Cursor or Gemini CLI?
All seven install into any agent that reads SKILL.md files through the skills CLI. The publishers name different agents: Sentry documents Claude Code, Cursor, Cline and GitHub Copilot; Trail of Bits documents Claude Code, Codex and ChatGPT workspaces; Firebase documents Gemini CLI, Claude Code, Codex, Kimi Code CLI, Cursor, Windsurf and GitHub Copilot.
Sources
- security-review SKILL.md and LICENSE (getsentry/skills)
- gha-security-review SKILL.md (getsentry/skills)
- code-review SKILL.md (getsentry/skills)
- getsentry/skills README
- differential-review SKILL.md (trailofbits/skills)
- supply-chain-risk-auditor SKILL.md (trailofbits/skills)
- semgrep-rule-creator SKILL.md (trailofbits/skills)
- firebase-security-rules-auditor SKILL.md (firebase/agent-skills)