How to build Semgrep dashboards in Metabase
Semgrep is a static analysis platform that scans code, dependencies, and secrets with customizable rules, in CI and in the Semgrep AppSec Platform. Metabase is where you turn those security signals into shared, trustworthy dashboards. This guide covers two complementary paths: a lightweight MCP + CLI route that pulls live data with the Semgrep MCP and loads a CSV into Metabase with the Metabase CLI, and a durable pipeline route that syncs Semgrep findings into a database so you can build dashboards anyone can read.
How do you connect Semgrep to Metabase?
Most teams combine both routes: use MCP and CLI uploads for a fast first pass, then move recurring security reporting to a warehouse-backed model.
Live data in, quick analysis out
Pair the Semgrep MCP with the Metabase CLI. Use MCP for live lookups, write a scoped result to CSV, then load it into Metabase as a ready-to-query table and model.
- Quick lookups such as "show me findings by rule and severity"
- Loading a Semgrep export into Metabase in seconds
- Spot-checks and one-off analyses without a warehouse
- Great for exploration, not governed security reporting
- Use read-only, minimally scoped credentials — security tools especially
- CSV uploads are snapshots — refresh or move to the pipeline for history
Durable dashboards with history
Sync Semgrep findings and metadata into a database or warehouse with a connector, custom pipeline, or API, then point Metabase at it.
- Semgrep posture reporting leaders and auditors depend on
- Joining Semgrep data with assets, HR, ticketing, or other security tools
- Long-run trends for findings by rule and severity and new vs. resolved findings
- You own the refresh schedule and the snapshot cadence
- Sync findings, entities, and daily rollups — not the raw event firehose
- Metric definitions must be consistent across tools and teams
What can you analyze from Semgrep data in Metabase?
- Findings by rule and severity — built from findings and the related rules, projects (repos), policies data your sync exposes.
- New vs. resolved findings — built from findings and the related rules, projects (repos), policies data your sync exposes.
- Triage state mix — built from findings and the related rules, projects (repos), policies data your sync exposes.
- Rule noise and ignore rate — built from findings and the related rules, projects (repos), policies data your sync exposes.
- Coverage by repository — built from findings and the related rules, projects (repos), policies data your sync exposes.
Which Semgrep dashboards should you build in Metabase?
Exposure overview
The headline risk picture across the estate.
- Open critical and high findings (number + trend)
- New vs. remediated findings this month (combo)
- SLA compliance by severity (bar)
- Riskiest asset groups (table)
Remediation operations
The queue that needs attention this week.
- Findings past SLA due date (table)
- Median time to remediate by severity (line)
- Assigned findings gone stale (table)
- Reopened findings (number)
Ownership views
Who owns the risk, and who's paying it down.
- Open findings by owner and team (bar)
- Oldest criticals per team (table)
- Fix rate by team (line)
- Exceptions and accepted risk (table)
Trend and posture
The quarter-over-quarter story, in plain numbers.
- Open criticals trend (line)
- Mean age of open criticals (line)
- Scan coverage of the asset inventory (number)
- Remediation velocity by quarter (bar)
How do you use the Semgrep MCP with the Metabase CLI?
Pair the Semgrep MCP with the Metabase CLI for fast, hands-on analysis. MCP is useful for scoped lookups and summarized exports; the Metabase CLI's upload command loads CSV data into Metabase and creates a ready-to-query table and model.
Example workflow
- Ask the MCP server for the open critical and high findings with severity, status, and first-seen dates.
- Export the result as CSV, keeping stable IDs, severities, statuses, owners, and timestamps.
- Run
mb upload csvto load it into Metabase as a table and model, then build questions and dashboards on top.
Be honest about the limits
- MCP lookups are excellent for exploration, not scheduled reporting.
- A CSV upload is a snapshot; refresh it with
mb upload replaceor move to the pipeline for real history. - First-seen and resolved timestamps are required for remediation-time and SLA metrics; daily snapshots are required for open-findings trends.
mb upload csvneeds an uploads database configured under Admin → Settings → Uploads.
How do you set up Semgrep MCP and the Metabase CLI?
Semgrep MCPofficial
- Transport
- Local server over stdio (uvx semgrep-mcp or the semgrep binary)
- Auth
- Optional SEMGREP_APP_TOKEN for AppSec Platform tools
- Best for
- Live scoped lookup and export
Metabase CLIofficial
- Install
npm install -g @metabase/cli- Auth
mb auth login- Load data
mb upload csv --file data.csv- Requires
- An uploads database (Admin → Settings → Uploads)
{
"mcpServers": {
"semgrep": {
"command": "uvx",
"args": ["semgrep-mcp"],
"env": {
"SEMGREP_APP_TOKEN": "your-token"
}
}
}
}The standalone semgrep-mcp package still works but is deprecated — newer Semgrep binaries ship the MCP server built in. A hosted endpoint at https://mcp.semgrep.ai/mcp exists but is explicitly experimental.
# Install the Metabase CLI
npm install -g @metabase/cli
# Log in (opens your browser; requires Metabase v62+)
mb auth login --url https://your-metabase.example.com
# Load a findings export — creates a table AND a model
mb upload csv --file semgrep-findings.csv --collection root
# Refresh that same table later from a new export
mb upload replace <table-id> --file semgrep-findings.csvCan you generate a Semgrep dashboard with AI?
Yes. Use the prompt below with any assistant that can run the Semgrep MCP and the Metabase CLI. It works end to end: if Semgrep tables already exist in Metabase it analyzes those; otherwise it pulls scoped, summarized data over MCP, loads it with mb upload csv, then builds the dashboard and caveats any metric that needs missing history.
Create a polished Metabase dashboard for Semgrep vulnerability management analytics.
Work end to end: get the data into Metabase if it isn't there yet, then build.
Goal: Help security and engineering leaders understand open exposure, remediation velocity, SLA compliance, and the riskiest assets from Semgrep data.
Step 1 — Find or load the data:
- First, check what already exists in Metabase (search for semgrep tables and
models). If durable Semgrep data is already present — synced from a warehouse
or uploaded earlier — use it and skip to Step 2.
- If nothing is there, pull a scoped, summarized export with the Semgrep MCP:
findings, plus rules, projects (repos), policies.
Prefer aggregated or entity-grain views over raw events. Write each result to a
CSV, then load it with the Metabase CLI — run "mb upload csv --file <export>.csv"
so each upload creates a table and a ready-to-query model. Use "mb upload replace
<table-id> --file <export>.csv" to refresh an existing table instead of creating
duplicates.
Step 2 — Inspect before querying:
Do not assume exact table or column names. Inspect available fields, severities,
statuses, timestamps, and whether snapshots or history exist before creating
duration or trend cards.
Important:
- Build on whatever data is present; don't claim Metabase connects natively to
Semgrep — it reads a database or CLI-uploaded tables.
- Never try to load raw security telemetry or event firehoses into Metabase; use
findings, detections, entities, and daily rollups.
- Only compute durations (time to remediate, time to triage, time to detect)
when the required timestamps exist.
- Exclude test environments and suppressed or accepted-risk items from headline
cards, and keep them visible in a labeled register instead.
- A single CSV is a point-in-time snapshot: only build trend cards if there is a
usable date column or multiple periods have been uploaded.
Dashboard title: Semgrep Vulnerability Management Overview
Sections:
1. Executive summary: Open criticals; New findings last 30 days; Remediated last
30 days; SLA compliance; Median time to remediate.
2. Exposure: Open findings by severity and asset group; aging buckets.
3. Remediation: New vs. remediated by week; median remediation days by severity.
4. SLA: Compliance by severity; findings past due with owner.
5. Ownership: Open findings by team; oldest criticals; accepted-risk register.
Filters: Date range, Severity, Status, Team or owner, Environment or asset group.
Output: Build the dashboard if you have permission; otherwise provide the exact
questions, SQL, model definitions, and layout. Include caveats for any metric
that cannot be calculated from the available data.How do you sync Semgrep data into a database or warehouse?
For dashboards that need history and reliability, land Semgrep findings and metadata in a database first, then connect Metabase to that database.
Connector options
- Managed ETL — use a connector when one covers the objects you need.
- Custom pipeline — use the Semgrep API for control over fields, snapshot cadence, and refresh schedule.
- MCP + CSV — use this for quick exploration and one-off slices.
No managed ELT connector. The findings API requires a Team or Enterprise plan — on those tiers, pull findings and projects on a schedule; on the free tier, export CI scan output (JSON/SARIF) into the warehouse instead.
Notes
- Decide the snapshot cadence first (daily is the norm for security data) — posture trends only exist if you build the history.
- Land raw entity tables first, then build clean Metabase models on top.
- Normalize asset, severity (one scale across scanners), status, first-seen, resolved-at, and source-scanner fields.
How should you model Semgrep data in Metabase?
Core tables
| Table | Grain | Key columns |
|---|---|---|
semgrep_findings | one row per finding | id, repo, rule_id, severity, status, first_seen_at, triaged_at, resolved_at |
semgrep_projects | one row per repository | id, name, default_branch, last_scan_at, findings_open |
semgrep_rule_stats_daily | one row per rule per day | rule_id, snapshot_date, findings_new, findings_ignored |
Modeling advice
- Build a clean
vulnerability_findingsmodel with common columns across tools, so multi-source dashboards don't fork definitions. - Separate entity tables (assets, users, controls) from findings, events, and daily snapshot tables.
- Exclude test environments and suppressed or accepted-risk items from headline metrics; keep them in a labeled register instead.
- Use stable IDs for asset, user, and finding joins; display names change.
Which Semgrep metrics should you track in Metabase?
| Metric | Definition | Notes |
|---|---|---|
| Open critical vulnerabilities | Open critical and high findings, deduplicated across scanners. | Count findings, not per-scan rows. |
| Mean time to remediate | Median days from first-seen to resolved, by severity. | Use medians; exclude still-open findings. |
| Vulnerability SLA compliance | Findings closed within their severity SLA over all findings due. | Count overdue open findings against the SLA too. |
| Scan coverage | Assets scanned in the window over the full inventory. | The denominator must come from outside the scanner. |
What SQL powers Semgrep dashboards in Metabase?
These assume a cleaned analytical model in a warehouse (PostgreSQL dialect). Adjust table and column names to match your pipeline.
The exposure table, with age buckets.
SELECT
a.asset_group,
COUNT(*) AS open_critical_findings,
COUNT(*) FILTER (
WHERE f.first_seen_at < CURRENT_DATE - INTERVAL '30 days'
) AS older_than_30d,
COUNT(*) FILTER (
WHERE f.first_seen_at < CURRENT_DATE - INTERVAL '90 days'
) AS older_than_90d
FROM vulnerability_findings f
JOIN assets a ON a.id = f.asset_id
WHERE f.status = 'open'
AND f.severity IN ('critical', 'high')
GROUP BY a.asset_group
ORDER BY open_critical_findings DESC;Remediation velocity from lifecycle timestamps.
SELECT
date_trunc('month', resolved_at) AS month,
severity,
percentile_cont(0.5) WITHIN GROUP (
ORDER BY EXTRACT(EPOCH FROM (resolved_at - first_seen_at)) / 86400
) AS median_days_to_remediate
FROM vulnerability_findings
WHERE resolved_at IS NOT NULL
GROUP BY 1, 2
ORDER BY 1, 2;Findings closed within their SLA window, of all findings due.
SELECT
severity,
COUNT(*) AS findings_due,
COUNT(*) FILTER (
WHERE resolved_at IS NOT NULL AND resolved_at <= sla_due_at
) AS closed_within_sla,
ROUND(
100.0 * COUNT(*) FILTER (
WHERE resolved_at IS NOT NULL AND resolved_at <= sla_due_at
) / NULLIF(COUNT(*), 0), 1
) AS sla_compliance_pct
FROM vulnerability_findings
WHERE sla_due_at < CURRENT_DATE
GROUP BY severity
ORDER BY severity;