Skip to main content

GCP

The GCP scrapers scrapes your GCP account to fetch all the resources & save them as configs.

gcp-scraper.yaml
apiVersion: configs.flanksource.com/v1
kind: ScrapeConfig
metadata:
name: gcp-flanksource
namespace: mc
spec:
gcp:
# An organization on its own scrapes every project beneath it. Add projects to
# narrow it to those that belong to the organization. Listing projects without
# an organization still works, but identities are then tenanted by project.
#- organization: "1234567890"
# projects:
# - workload-prod-eu-02
- project: workload-prod-eu-02
exclude:
- SecurityCenter
#- IAMGroupMembers # disable Google-group expansion (needs Cloud Identity groups.readonly)
#connection: connection://mc/gcloud-flanksource
# IAMPolicy and IAMGroupMembers run by default only when include is empty.
# Once include filters asset types, list the IAM flags explicitly to keep IAM data:
#include:
#- storage.googleapis.com/Bucket
#- container.googleapis.com/Cluster
#- IAMPolicy # RBAC: users/groups/service-accounts -> roles
#- IAMGroupMembers # expand Google group membership
# AuditLogs is opt-in and only runs when listed here:
#- AuditLogs # access history from the BigQuery audit-log dataset
#auditLogs:
#dataset: default._AllLogs
# Project holding the dataset. Required when scraping an organization or
# more than one project, since the dataset lives in exactly one project.
#project: logging-prod
#since: 30d
#excludeMethods:
#- io.k8s.*

Scraper​

FieldDescriptionSchemeRequired
logLevelSpecify the level of logging.string
scheduleSpecify the interval to scrape in cron format. Defaults to every 60 minutes.string
retentionSettings for retaining changes, analysis and scraped itemsRetention
gcpGCP scrape config[]GCP

GCP​

note

Either the connection name or the credentials are required (if Workload Identity is not being used)

FieldDescriptionScheme
auditLogs

Query BigQuery dataset for audit logs

AuditLogs

connection

The connection url to use, mutually exclusive with credentials

Connection

costReporting

Read the Cloud Billing export from BigQuery

CostReporting

credentials

The credentials to use for authentication

EnvVar

endpoint

Custom GCP Endpoint to use

string

exclude

GCP asset types to exclude from scraping

[]string

include

GCP asset types and/or feature flags to scrape. This is a strict allowlist — see Include

[]string

organization

Organization to scrape, as an organization number (1234567890) or a qualified name (organizations/1234567890). Its resource hierarchy and Security Center findings are scraped at the organization root

string

project

GCP Project ID. An alias for a single-entry projects list

string

projects

Narrow the scrape to these projects, as project ids (gcp-proj-1) or qualified names (projects/gcp-proj-1). Empty means every project in the organization

[]string

skipTLSVerify

Skip TLS verification when connecting to GCP

boolean

labels

Labels for each config item.

map[string]string

properties

Custom templatable properties for the scraped config items.

[]ConfigProperty

tags

Tags for each config item. Max allowed: 5

[]ConfigTag

transform

Transform configs after they've been scraped

Transform

You must specify one of organization, projects or project

Include​

include is a strict allowlist. Leave it empty and everything except AuditLogs runs; set it and only what is listed runs. Because it covers both asset types and feature flags, narrowing one dimension turns the other off entirely:

include: [storage.googleapis.com/Bucket] # buckets only, NO IAM/RBAC
include: [IAMPolicy] # IAM/RBAC only, NO assets

List both to filter assets while keeping the rest:

include: [storage.googleapis.com/Bucket, IAMPolicy, IAMGroupMembers]

Asset types come from the GCP supported asset types list. The feature flags are:

FlagDescription
IAMPolicyRBAC access from IAM policy bindings, and the resource hierarchy (organization and folder config items), read in the same pass
IAMGroupMembersExpand Google group membership via the Cloud Identity groups.readonly scope. Disable with exclude: [IAMGroupMembers]
AuditLogsBigQuery audit-log access. Opt-in: it runs only when listed here explicitly

SecurityCenter can be passed to exclude to skip Security Center findings.

Audit Logs​

FieldDescriptionScheme
datasetBigQuery dataset to query audit logs from (e.g., "default._AllLogs")string
projectProject holding the BigQuery dataset. Defaults to the scraped project. An organization-scoped scrape must set this to the project its aggregated log sink writes tostring
sinceTime range to query audit logs (e.g., "24h", "7d", "30d"). Defaults to the last 7 daysstring
userAgentsFilter user agents matching these patternsMatchExpressions
principalEmailsFilter principal emails matching these patternsMatchExpressions
permissionsFilter permissions matching these patternsMatchExpressions
serviceNamesFilter service names matching these patternsMatchExpressions
methodsFilter methods matching these patternsMatchExpressions

Cost Reporting​

Reads the Cloud Billing export from BigQuery. This must be the detailed usage cost export (gcp_billing_export_resource_v1_<BILLING_ACCOUNT_ID>) — the standard export carries no resource column, so every charge would be attributed to its project rather than to the resource that incurred it.

FieldDescriptionScheme
projectProject holding the billing export dataset. Defaults to the scraped projectstring
datasetDataset holding the export table e.g. billing_exportstring
tableThe export table e.g. gcp_billing_export_resource_v1_01ABCD_2345EF_67890Astring
lookbackDaysHow many days of the export to read on each scrape. Values of 0 or less use the 45 day default. BigQuery bills by bytes scanned, so this is the main cost controlint