Cloud security posture scanner architecture: Cloud Asset Inventory feeding OPA/Rego rules, a common finding schema in BigQuery, Pub/Sub and log-based alerts, and a private dashboard
Two independent sources — the live cloud and the Terraform that describes it — land in one schema, with a lifecycle instead of a log line.

Most posture scanners answer one question well: what's wrong right now. Fewer answer the questions that actually matter to a security team living with the output day after day — is this the same finding I saw yesterday, did the thing I fixed last week actually stay fixed, and does my IaC scanner agree with what's really running in the cloud. Those questions need identity and a lifecycle, not just a fresh list every time a scan runs.

gcp-security-posture-scanner reads a GCP project through Cloud Asset Inventory (both the RESOURCE and IAM_POLICY views, across 8 resource types and 4 IAM-policy types), evaluates 16 CIS-GCP-mapped rules written in Rego, and merges the result into a BigQuery table where every finding has a stable fingerprint and a lifecycle: open, resolved, reopened, with first_seen preserved across all of it. A second source — Checkov over Terraform — lands in the exact same schema, so a change caught before it deploys and a misconfiguration found live in the cloud are one queryable table, not two dashboards you have to mentally merge. It deployed live in europe-west1, and the test suite passed 45 of 45, including a full harden → resolve → regress → reopen drill against real, Terraform-created misconfigurations.

A fingerprint, not a row

Each finding's identity is sha256(source | rule_id | resource | subject)[:16] — deliberately including subject, because one resource can violate one rule in more than one way (two different roles on the same project IAM policy, say). Every scan runs a single MERGE: fingerprints present in this scan are opened or refreshed with first_seen left untouched; fingerprints that were open but are absent from this scan, for the same source and project, are marked resolved. A regression doesn't create a new row — it reopens the same one, with its original discovery date intact. The alternative (delete and reinsert everything each scan) would lose exactly the information a security team cares about: how long was this actually open, and is today's alert new or something we already knew about.

What the org's own guardrails ruled out — and what that's worth knowing

The scanner ships a small lab of deliberately misconfigured Terraform resources so it has something real to find on a fresh project. The first live apply failed twice: the organization's own policies forbid creating a bucket with uniform bucket-level access disabled (HTTP 412) and forbid creating a service-account key at all (HTTP 400). Rather than fight the guardrails, the lab moved to misconfigurations the org actually allows — a public bucket binding, a BigQuery dataset shared with allAuthenticatedUsers, a public Cloud Run service, Editor granted to a service account — and the two blocked rules stayed in the rule set, proven by 33 Rego unit tests instead of a live finding. That failure is itself a small, honest piece of posture information: this organization already prevents those two specific misconfigurations at the platform level, which is worth knowing on its own.

Sources that can't resolve each other's findings

The MERGE that closes a finding is scoped to (source, project) — a clean Checkov run over Terraform can never mark a live Cloud Asset Inventory finding resolved, and vice versa. That matters because the two sources answer genuinely different questions: Checkov catches a mistake before it's deployed; the live scanner catches drift, manual changes, and anything that predates the current Terraform. Conflating them would let a passing CI run silently paper over a real, currently-open misconfiguration in the cloud.

Six decisions, honestly framed

DecisionTrade-off
Rules are Rego, run by the opa binary, not Python or SQLA learning curve, in exchange for one file a security engineer can review without reading the scanner's code
Read the Asset API directly rather than its BigQuery exportScope is exactly the asset types listed in inventory.py — a type not listed is invisible, not clean
Two rules kept unit-test-only rather than droppedHonest about what this org's own policies already prevent, at the cost of two rules with no live evidence
Lifecycle via one MERGE, not delete-and-reinsertMore SQL complexity, in exchange for first_seen and true opened/resolved counts
Checkov severities are a static map in codeCheckov's open-source output usually carries no severity at all
No suppression/accept-risk workflowAn accepted risk stays visibly open rather than silently hidden — deliberately, for now

Three bugs a live run found that a mocked test never would have

A hand-written fixture assumed the BigQuery dataset ACL lived in the RESOURCE view of the asset, the way Compute and Storage do. It doesn't — it's only in the IAM_POLICY view — so GCP-BQ-001 never fired against a fixture built from the public API docs. It was caught by re-capturing the fixture from a real inventory snapshot and diffing the expected rule IDs, which is now exactly what the Python test suite does. Checkov turned up two more: it prints one JSON document per -d directory scanned, which broke a naive json.load with "Extra data"; and it reports the same file under different relative paths depending on which directory was scanned, which briefly turned one real finding into two duplicate rows until the merge started preferring Checkov's own repo-relative path and de-duplicating by fingerprint before staging.

Deployed dashboard showing 4 critical, 8 high, 18 medium and 4 low open findings from both cai-rego and checkov sources, with rule id, evidence, CIS mapping and first-seen date
The private dashboard — no public binding, reached only through gcloud run services proxy — showing real findings from both sources on the live project.

What this doesn't claim to be

Scope is one project and exactly the asset types the code lists — nothing broader is scanned, and nothing outside that list is "confirmed clean," it's simply not looked at. This is configuration posture, not runtime threat detection: it will never see a compromised process on a VM. There's no suppression workflow, so an accepted risk stays open in the dashboard rather than being formally signed off. And GitHub Actions CI here validates, tests and scans — it deliberately doesn't deploy, the same "keep CI out of the blast radius" reasoning the keyless CI/CD build takes further by actually wiring up a live, credential-free deploy path.

Try it yourself

git clone https://github.com/soodrajesh/gcp-security-posture-scanner
cd gcp-security-posture-scanner
gcloud config set project <your-project>
./scripts/up.sh # infra → lab misconfigurations → image → scan → Checkov ingest → live tests → lifecycle drill
./scripts/dashboard.sh # http://localhost:8088, your own identity attached
./scripts/down.sh # delete everything, including the lab

Under €1 for a full build-test-drill-teardown cycle — the scan job and dashboard scale to zero, and BigQuery/Pub/Sub/Cloud Asset usage stays far below free-tier limits.

📢 Have questions or feedback? Drop a comment below or connect with me on Twitter/X@spysood!