
·10 min read
Cutting a Managed Prometheus Bill by 70% (While Adding Metrics)
How to attribute a Google Managed Prometheus bill to individual scrape jobs, turn off the GKE default collectors nothing reads, and add app metrics through measured keep-lists — plus the two cardinality bombs that would have multiplied the bill 11×.
Managed Prometheus is easy to switch on, and that's the problem. You pay per sample, GKE enables collectors by default, and the first time anyone looks at the cost is usually when it's already high. Before adding application metrics to a GKE cluster, I wanted the opposite order: attribute every sample, cut what nothing reads, and price every new metric before it ships.
TL;DR
- Google Managed Service for Prometheus (GMP) bills per sample ingested. Every series, at every scrape, costs money.
- 83% of the samples came from cAdvisor, which GKE's managed collection turns on by default. Nothing read them. The free
kubernetes.io/*system metrics already covered container health. - Dropping cAdvisor and kubelet from the cluster's
--monitoringlist removed 86% of the bill in one command, without touching any nodes. - App metrics were then added through keep-lists measured offline before applying. Scraping everything as-is would have cost 4.5× the original bill.
- A check after every rollout caught two cardinality bombs in application code (GPS coordinates as labels, raw request paths as labels) that were on pace for 11× the original bill. Both were contained within the hour, before they reached an invoice.
- Net result: about −71%, with app metrics that didn't exist before.
How GMP Charges You
Self-hosted Prometheus costs you RAM and disk. GMP costs you samples, at roughly $0.05–0.06 per million (it gets cheaper at higher volume).
samples/day = Σ over jobs ( samples_per_scrape × 86400 / scrape_interval )At a 30s scrape interval, that's roughly:
| Active series | Samples / month | Approx. GMP cost / month |
|---|---|---|
| 100k | ~8.6B | ~$500 |
| 1M | ~86B | ~$4,500 |
| 5M | ~430B | ~$20,000 |
Three things follow from that formula:
- Interval matters a lot. 60s costs half of 30s, and for RED and capacity metrics you won't miss the difference.
- Histograms are expensive. Every
_bucket× every label combination is a series. One service I measured was 88.8% bucket lines. - Cardinality is the cost. One label with unbounded values can multiply your bill without anyone deploying anything.
Everything below is shown in percentages and multiples, because they carry over to any cluster size. A 4.5× mistake is a 4.5× mistake at any scale.
Step 1: Find Out What You're Paying For
The first step isn't cutting anything. It's finding out where the samples come from. Also look at Monitoring and Logging as separate SKUs. The billing overview makes it easy to blame the wrong one.
Samples per scrape, per job:
sum by (job) (scrape_samples_post_metric_relabeling)Multiply by scrapes per day and you have a per-job cost. To cross-check against what Google actually billed, use the Cloud Monitoring metric monitoring.googleapis.com/billing/samples_ingested over 7 days (ALIGN_SUM, 86400s alignment).
The breakdown:
cAdvisor: 82.7%
kubelet + JobSet: most of the rest
App metrics: 0%And you can't tune those jobs. GKE's managed scrape configs are labeled addonmanager.kubernetes.io/mode: Reconcile, so any edit you make gets reverted. The only lever is the cluster's component list.
Step 2: Check Who Uses It Before You Cut It
Before turning anything off, check everything that could depend on it:
- Grafana dashboards:
/api/search?type=dash-db, then check what each dashboard queries. Here, none used cAdvisor data. - Alert policies: list them through the Monitoring API and check each condition's
metric.type. The only alert used the freekubernetes.io/*restart count, which this change doesn't touch. - Autoscalers: every KEDA ScaledObject trigger was RabbitMQ. None used Prometheus.
- Log-based metrics: none depended on these.
This step matters. If a cost cut turns off a metric an autoscaler reads, you've created an outage.
Step 3: Drop the Managed Components
gcloud container clusters update <cluster> --region <region> \
--monitoring=SYSTEM,POD,DEPLOYMENT,STATEFULSET,DAEMONSET,STORAGE,HPA,DCGMThe list leaves out CADVISOR, KUBELET, and JOBSET. It's a control-plane change: no nodes recreated, no pods restarted.
The enum gotcha: clusters describe shows the component as SYSTEM_COMPONENTS, but clusters update only accepts SYSTEM. Copy what describe shows and the command fails.
Container CPU, memory, and restarts are still available through the free kubernetes.io/* system metrics. You lose the detailed per-container Prometheus series, which nothing was reading.
What you give up, and how to get it back
This is a real trade-off, so be honest about it. Without cAdvisor you lose the series you'd want in the middle of a hard investigation: CPU throttling (container_cpu_cfs_throttled_periods_total), the working-set vs. cache breakdown, and per-container network and filesystem I/O. The system metrics are enough to spot a memory leak (usage climbing toward the limit, OOM restarts), but not always enough to diagnose one.
Three ways to keep that option without paying for everything all the time:
- Scrape cAdvisor yourself, with a keep-list. GMP can't edit its managed configs, but your own
ClusterNodeMonitoringcan scrape the kubelet's/metrics/cadvisorendpoint, withmetricRelabelingkeeping only the few families you actually debug with:
apiVersion: monitoring.googleapis.com/v1
kind: ClusterNodeMonitoring
metadata:
name: cadvisor-essentials
spec:
endpoints:
- path: /metrics/cadvisor
scheme: https
interval: 60s
metricRelabeling:
- action: keep
sourceLabels: [__name__]
regex: 'container_cpu_cfs_(throttled_)?periods_total|container_memory_working_set_bytes'This is the same structure GKE uses for its own managed gmp-kubelet-cadvisor scraper (kubectl get clusternodemonitoring gmp-kubelet-cadvisor -o yaml on a cluster that has it enabled). The difference is that its keep-list has 19 metric families at a 30s interval, and this one has two at 60s.
- Turn it back on when you need it.
--monitoringis a control-plane setting. AddingCADVISORback takes minutes and no node restarts, so during an incident you can enable it, investigate, and remove it again. kubectl topstill works, because metrics-server doesn't depend on any of this.
Decide this as a team, before the incident, not in the middle of it.
A trap if you write alert rules on those system metrics: kubernetes_io:* metrics have a cluster_name label, not cluster. GMP's ClusterRules injects a cluster= matcher into every selector, so a rule there shows health=ok and can never fire, with no error anywhere. Put those rules in GlobalRules with an explicit cluster_name matcher. More generally, a rule that evaluates without errors isn't proof that it works. Test new rules in staging with a condition you know is true (a threshold set low enough to fire on purpose), confirm the alert reaches its receiver, then set the real threshold.
Step 4: Add App Metrics, Measured First
With the bill down 86%, there was room for the metrics that actually matter: RED metrics for the application services, plus Temporal, Milvus, RabbitMQ, Pulsar, and the embedding servers. "Scrape everything" was the obvious way to bring the cost straight back.
GMP ignores your annotations
prometheus.io/scrape: "true" does nothing on GMP. Only PodMonitoring / ClusterPodMonitoring CRs create targets. Treat the annotations as hints and probe every endpoint before scraping it. Several services annotated scrape: "true" returned 404 on /metrics.
The oldest-replica rule
Where you probe matters as much as what you probe:
api-gateway fresh replica: 87 samples
api-gateway 6-day-old replica: 143,000 seriesLabel cardinality builds up over the life of a process. If you probe whichever pod kubectl get pods lists first, you can underestimate by about 1000×. Always probe the oldest replica.
Scraping every endpoint as-is would have cost 4.5× the bill that had just been cut. That's why the next step exists.
Keep-lists, replayed offline
Rules of thumb:
- Histograms: keep
_sumand_count(enough for average latency and rate). Drop_bucket, and add back a single family's buckets only when you really need a p99. - Python services: drop
python_.*,.*_created, and static process gauges. - RabbitMQ: here the cost comes from the number of objects, not buckets. Drop per-connection and per-channel series. Keep queue depth, throughput, alarms, and raft indices.
- Interval: 60s.
The part that makes this reliable is offline replay. Save the raw /metrics dumps, extract the regex from the actual YAML (kubectl apply --dry-run=client -o json, so you test what will really be applied), anchor it the way Prometheus does (^(?:...)$), and count kept vs dropped series. Also list dead alternatives: branches of the regex that match nothing. Each one is usually a typo, which means a metric you think you're keeping is silently missing.
metricRelabeling:
- {action: drop, sourceLabels: [__name__], regex: '.*_bucket'}
- {action: drop, sourceLabels: [__name__], regex: '(python_.*|.*_created|process_(max_fds|virtual_memory_bytes|start_time_seconds))'}metricRelabeling can see target labels like pod and container, so you can drop per-workload too:
- {action: drop, sourceLabels: [pod, __name__], regex: 'api-gateway-.*;(total_requests_total|request_latency_seconds_.*)'}Step 5: Verify After Apply — and the Cardinality Bombs It Caught
Keep-lists control metric names. Cardinality lives in label values, which grow with traffic and pod age and can't be fully known from a probe. So every rollout ends with the same check: re-run the per-job query. The first time, the app services job came in at 360× the plan, on pace for 11× the original bill. Two culprits, both in application code:
1. GPS coordinates as labels.
location_requests_total{latitude="83.87…",longitude="3.79…"} 1One series per coordinate pair, and they never expire. One service had 268,435 series, a 53 MB /metrics response, which it rebuilt in memory on every scrape.
2. Raw request paths as a label on an internet-facing gateway. Every vulnerability scanner that tried /.env or /wp-config.bak created a new series. That was 5,404 routes on a 6-day-old pod, still growing. That's a cost problem, and it's also a memory-DoS risk: attackers can control your label values.
The per-pod drop rules above contained it within the hour, before it showed up on an invoice. But dropping at scrape time doesn't fix it. The app still creates every series in memory: the 53 MB response was built in the service's own RAM on every scrape, whether GMP kept it or not. Leave the code alone and that memory keeps growing, and the same pattern turns up in the next metric someone adds.
The real fix is in the code, and it's small:
# route: label by the route TEMPLATE, never the raw path
route = request.scope.get("route")
label = route.path if route else "unmatched" # "/api/users/{id}", not "/api/users/8812"
REQUESTS.labels(route=label, method=request.method).inc()
# location: bound the label space — round, geohash, or move it out of metrics
REQUESTS_BY_AREA.labels(geohash=geohash.encode(lat, lon, precision=4)).inc()Unmatched paths (scanners, typos, 404s) all go into one unmatched bucket, so attacker input never becomes a label value. Precise coordinates belong in logs or traces, where high cardinality is normal and you don't pay per series.
That went to the dev team, and a metrics review is now part of how new services get onboarded. The drop rules are the stopgap. The code change is the fix.
Two GMP Gotchas Worth Knowing
basicAuth is half a secret. In PodMonitoring, password is a secretKeyRef but username is a literal string. Commit username: PLACEHOLDER and replace it with sed at apply time. The GMP collector's service account also needs RBAC to read the secret. Scope it to that one secret by name, and apply it before the PodMonitoring:
rules:
- apiGroups: [""]
resources: [secrets]
resourceNames: [<basic-auth-secret>]
verbs: [get, watch]
subjects:
- {kind: ServiceAccount, name: collector, namespace: gmp-system}Snapshot histograms read as zero. GMP stores counters and histograms as cumulative values with a start time and returns the increase. A custom exporter that published an all-time distribution (recomputed each scrape, never incremented) showed up as 0 for every bucket, while gauges from the same scrape were fine. The fix: emit the buckets as a gauge named ..._bucket with an le label. histogram_quantile() only needs le, so it keeps working.
The Result
| Before | After | |
|---|---|---|
| Monthly bill | baseline | about −71% |
| Share of samples from defaults | 83% (cAdvisor) | none |
| App metrics | none | RED metrics on the app services + Temporal, Milvus, RabbitMQ, Pulsar, embeddings |
| Scrape interval | 30s | 60s |
The method holds up as things get added. A recent RabbitMQ dashboard needed 25 more metric names. The offline replay priced the change before it was applied: a fraction of a percent of the bill. It was approved based on that number. That's the habit worth building: price every addition before you apply it.
Lessons
- Attribute before you cut.
sum by (job) (scrape_samples_post_metric_relabeling)is the most useful query for managed Prometheus costs. - Audit the defaults first. Managed platforms turn collectors on for you. Here they were most of the bill.
- Find every consumer first. Dashboards, alerts, autoscalers, log-based metrics. A cost cut that blinds an autoscaler is an outage.
- Know which diagnostics you're giving up. Decide in advance how you'll get them back (a targeted scrape, or a toggle you can flip during an incident).
- Probe the oldest replica. Cardinality builds up over time, and a fresh pod tells you almost nothing.
- Replay keep-lists offline, against the regex from the real YAML, and look for branches that match nothing.
- Unbounded labels are a bug, and on the internet they're user-controlled. Drop them at scrape time to stop the cost, then fix them in code to stop the memory growth.
- Verify after every apply. Planning controls names; only a post-apply check catches label values. That check is what kept an 11× risk to one hour.
Cheap monitoring isn't the goal. The goal is to pay only for metrics someone actually uses.
A note on how this was written: I used AI to help with phrasing and structure. The tone, the opinions, the war stories, the code, and everything else here are mine.