Tools/Security & compliance

Semgrep: static analysis that fits in a pull request

Semgrep parses 30-plus languages and matches YAML patterns in seconds, and the engine is free under LGPL-2.1. Cross-file analysis, the rulesets and the AI triage sit behind paid tiers.

Type
Static analysis with AI rules
Pricing
Free · from $30 per seat

··10 min read

  • SAST
  • Static analysis
  • CI security
  • Rule authoring
  • AppSec
Diagram of the Semgrep scan pipeline, from source files through parsing and rule matching to reported findings

Key takeaways

  • The Semgrep engine is LGPL-2.1 and scans 30-plus languages offline; cross-file analysis needs a proprietary binary and an account.
  • The free tier covers 10 contributors and 10 private repositories, and counts contributors from a rolling 90-day git log.
  • Since 13 December 2024 Semgrep-maintained rules carry the Semgrep Rules License v1.0, which forbids redistribution and service use; OpenGrep is the LGPL fork that answers it.
  • Cross-file analysis silently falls back to single-file after 5 GB of memory or three hours, so a large repository loses depth without failing.
  • AI Autofix costs 20 credits per finding against 20 monthly credits per Teams contributor, which makes autofix the most expensive action on the plan.

Semgrep is a static analyser built around pattern matching: it parses source code into a syntax tree, then reports every place where a YAML rule's pattern matches that tree. The engine is open source under LGPL-2.1, runs on a laptop with no network connection, and finishes a typical repository in seconds. The position taken here is that it is the cheapest credible first SAST tool a small team can put into CI, and that the free tier is a well-built on-ramp rather than a finished product: the analysis that removes most of the noise, the rulesets it ships with and the AI triage all sit behind a login.

It sits between the linters a team already runs and the heavyweight scanners that need a security engineer to feed them. In practice it competes with CodeQL for depth, with SonarQube for the pull-request gate and with Snyk Code for developer ergonomics; the open-source fork OpenGrep competes for the licence. What it replaces most often is a folder of half-maintained grep patterns.

What it is

A command-line scanner plus a hosted platform. The command line does the work: give it a target directory and one or more rulesets, and it returns findings carrying the rule id, severity, message and matched range. The platform, named Semgrep AppSec Platform, adds a dashboard, policy, pull-request comments, Supply Chain, Secrets detection and the AI features. Everything below the platform is the Community Edition.

  • Engine licence LGPL-2.1; latest release 1.179.0, published 1 October 2026
  • 30-plus languages in Community Edition, 35-plus listed for Semgrep Code
  • 3,000-plus community rules, plus registry rulesets by prefix such as p/python
  • Free tier: 10 contributors, 10 private repositories, 60 AI credits a month
  • Teams from $30 per contributor a month; Secrets is $15
  • Contributors counted from the git log over a rolling 90 days
  • Cross-file analysis needs a proprietary binary from semgrep install-semgrep-pro

How it works

Each language has a parser, mostly tree-sitter based, that turns the file into a syntax tree. A rule is a pattern over that tree written with metavariables such as $X and $F, plus optional constraints in patterns, pattern-not, metavariable-regex and taint mode. Matching is structural rather than textual: whitespace, renamed variables and reordered statements do not hide a match, which is why the false-positive rate stays low enough for a pull-request gate.

Semgrep: how a scan turns source into findingsSource files are parsed into a syntax tree, matched against rule patterns, and reported as findings. Rulesets from the registry and the optional Pro engine feed the matching step, and findings leave as terminal output, JSON or SARIF.Semgrep: source in, findings outsemgrep.devEVERY SCANSource filessrc, gitParsertree-sitter ASTRule matchpatterns, taintFindingsid, severityReportjson, SARIFWHAT FEEDS THE MATCHRegistry rulesetsauto, p/pythonPro enginecross-file, login
The rule decides what counts as a finding; the engine only decides where a pattern matches.

Findings carry the rule id, severity and a message, and can be suppressed per line with a nosemgrep comment or per rule with an ignore entry. Output goes to the terminal, to JSON or to SARIF, which is what most dashboards ingest.

The boundary that matters is scope. Community Edition analysis is intraprocedural: it reasons inside a single function. Cross-function and cross-file analysis run on the Pro engine, a separate binary that only activates after semgrep login and semgrep install-semgrep-pro.

Getting started

Install the CLI with pip, pipx, uv or brew; no account is needed to scan with community rules. A rule file is small enough to sit next to the code it protects.

# rules/no-pickle.yaml
rules:
  - id: avoid-pickle-load
    languages: [python]
    severity: WARNING
    message: |
      pickle.load executes whatever is in the stream. Only feed it
      bytes that never left this machine; use json.loads otherwise.
    pattern: pickle.load($STREAM)

Then run semgrep scan --config rules/no-pickle.yaml src/ for one ruleset, semgrep scan --config p/python src/ for a registry prefix, or semgrep --config auto to let the tool choose rulesets from the languages it detects. Add --json or --sarif for machine-readable output, and semgrep ci once a repository is connected to a deployment.

Pricing and the rules licence

The engine is free; the interesting money sits around it. The pricing page as of October 2026 lists three plans, and the free one is capped in a way that only becomes visible at the eleventh contributor.

PlanPriceLimitsWhat it unlocks
Free$010 contributors, 10 private repositoriesCode and Supply Chain, 60 AI credits a month
TeamsFrom $30 per contributor a month500 private repositoriesSSO, role-based access, Secrets at $15, 20 AI credits each
EnterpriseCustomNo limit on repositories or contributorsOn-prem source control, custom CI, 50 AI credits, account manager

Licence is a separate question from price. The engine stays LGPL-2.1, but since 13 December 2024 the rules Semgrep maintains ship under the Semgrep Rules License v1.0, which allows internal, non-competing use and forbids redistribution or offering the rules as a service. Consultants and companies using them internally are inside those terms; a vendor building a competing SAST product on them is not. That change is what produced OpenGrep, an LGPL-2.1 fork maintained by a consortium of security vendors including Aikido, Endor Labs, Jit and Orca Security.

AI features are metered separately in credits, and the costs are published. Comments on a pull request are free, triage costs one credit per finding, and Autofix, which opens a pull request with a fix, costs twenty.

AI actionCreditsWhat it does
AI comments on a pull request0Guidance posted on the PR
AI analysis of a finding1 per findingTriage, remediation guidance, component tagging
AI Autofix20 per findingOpens a pull request that fixes the finding
AI-powered detection scanVariableDepends on scan size and complexity
Agentic Workflows runVariableMulti-step analysis, more tokens than a detection scan

Running it in CI

Semgrep is unusual in that the free tier is genuinely fast, and that changes what is worth gating on. The failure modes in large repositories are about depth and memory rather than about the scan itself.

  • Diff-aware scans analyse only what a pull request touched; cross-file analysis runs on full scans and never on pull-request scans.
  • Cross-file analysis falls back to single-file once a scan passes 5 GB of memory or three hours, so a large repository quietly loses depth instead of failing the build.
  • Monorepo support splits a repository into parts on every plan; distributed scans across several machines start at Teams.
  • Join mode, the only facility in the open engine for joining matches across rules, is documented as experimental and not actively maintained.
  • Rules can be shared through the registry; private rules that keep sensitive rule logic inside the organisation need Teams or above.

Where it shingles

The ceiling arrives sooner than the pricing page suggests. Community Edition only reasons inside one function, so a tainted value crossing a helper is invisible until somebody pays for the Pro engine, and type inference and constant propagation are the same story. The second limit is rule quality: generic patterns without a metavariable constraint produce enough noise that teams learn to ignore the tool, and writing tight rules is a skill a team has to acquire. The third is arithmetic: at $30 per contributor per month on a rolling 90-day count taken from git log, anyone who committed to a private repository in that window is billed, whether or not they are still on the project.

ToolAnalysis depthRule authoringWhat it costs
Semgrep CEIntraprocedural, one function at a timeYAML patterns a reviewer can readFree engine, Semgrep-maintained rules for internal use only
CodeQLWhole-database dataflow across filesQL, a query language with a real learning curveFree for public repositories, GitHub Code Security otherwise
SonarQubeQuality plus baseline security; taint on paid tiersCustom rules need a Java pluginCommunity Build free, paid tiers per instance
Snyk CodeDataflow in a hosted engineVendor-maintained rules, little to write yourselfPaid, priced per developer

The opinionated version: for a team of eight, depth matters less than being read. A ruleset a developer understands and can extend in YAML will be acted on; a whole-database query language nobody has time to learn will not. The moment the threat model is dataflow across a service boundary, Semgrep CE cannot answer the question, and the $30 tier is the cheap option next to migrating the rules to QL.

Verdict

Take it as the default first scanner, with the eyes open about where the paid ceiling sits.

  1. Use the free Community Edition when the goal is a working security gate in CI within a week and nobody on the team writes query languages.
  2. Pay for Teams when single-function analysis produces findings developers argue with, and when SSO, a dashboard and private rules have become requirements rather than wishes.
  3. Do not treat $30 as the price: it is $30 per contributor per month on a 90-day rolling count, and unlimited repositories, on-prem source control and custom CI integrations live in Enterprise.
  4. Choose CodeQL when the question is how data reaches a sink across a compiled codebase, and SonarQube when the gate is about maintainability as much as security.
  5. If the rules have to be redistributed or served to third parties, the engine without the Semgrep-maintained rulesets, or OpenGrep, is the only arrangement that needs no conversation with the vendor.

Sources

  1. Semgrep pricing: Free, Teams and Enterprise
  2. Semgrep Community Edition
  3. Semgrep documentation: cross-file analysis
  4. Semgrep documentation: usage and billing
  5. Semgrep releases: v1.179.0
  6. Semgrep Rules License v1.0
  7. Important updates to Semgrep OSS, 13 December 2024
  8. OpenGrep repository

Frequently asked questions

Is Semgrep free to use in production?

The engine is free under LGPL-2.1 and the CLI runs without an account. The hosted free tier covers up to 10 contributors and 10 private repositories; beyond that Teams starts at $30 per contributor per month. The rules Semgrep maintains are free only for internal, non-competing use.

What is the difference between Community Edition and the Pro engine?

Community Edition analysis is intraprocedural, meaning it reasons inside a single function. The Pro engine adds cross-function analysis, which Semgrep Code runs by default, and optional cross-file analysis. It is a separate binary installed with semgrep install-semgrep-pro after semgrep login.

How accurate is Semgrep?

No precision figure is published, so the honest answer is structural: matching happens against a syntax tree, so a finding is either an exact pattern match or a rule that was written too loosely. Most noise in practice comes from rules without metavariable constraints rather than from the engine.

Does Semgrep send my source code to the cloud?

Run locally or fully in your own CI and the source stays put; only scan metadata is sent to the service. AI features submit the file containing a finding to a model, and Managed Scans clone the repository for the scan and destroy the clone afterwards.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.