
·9 min read
Whose GPUs Are These? Building a Zero-Footprint GPU Extension for Freelens
Why desktop Kubernetes tools can't answer basic GPU questions, and how I built a Freelens extension that reads the exporters you already run — no DaemonSet, no port-forward, no Prometheus required — plus the MIG and power-attribution bugs I hit along the way.
An 8× A100 server, sliced into 47 MIG instances, 29 GPU pods, about 805 watts at the wall. Someone asks a simple question: "Which namespace is holding all the GPUs, and is any of it actually doing work?" Every desktop Kubernetes tool I had open — k9s, Lens, Freelens — shrugged. So I built the answer into Freelens.
TL;DR
- Desktop Kubernetes tools show you pods, nodes, and CPU/memory. They don't show per-pod GPU usage, per-namespace GPU ownership, or idle VRAM.
- I built a Freelens extension that does — and it adds nothing to the cluster. No DaemonSet, no port-forward, no Prometheus requirement.
- It reads the exporters you already run (dcgm-exporter, per-process exporters) through the kube-apiserver pod-proxy, and falls back to Prometheus/Thanos/VictoriaMetrics/Mimir if it finds one.
- The hard part wasn't the UI. It was attribution: on MIG, DCGM repeats device-level numbers on every slice, and naive summing turned 805 W into 5,033 W.
- First commit to v0.8.1 took about 3.5 weeks. It's now part of the official
freelensapporg as@freelensapp/gpu-extension.
Why This Exists
GPUs are the most expensive thing in most clusters I touch, and they're the least visible thing in the tools I use every day.
Some things that are hard to see:
- "Whose GPUs are these?" There's no per-namespace view. It's an old ask. Open issues in both Lens and Kubeflow requesting per-namespace GPU usage have dozens of upvotes, and there was even a KubeCon talk with that title.
- "GPU %" lies.
DCGM_FI_DEV_GPU_UTILmeasures "a kernel was running", not "the GPU was working hard". A pod at 100% can be using a tenth of the SMs. - Shared GPUs blur everything. On time-sliced GPUs, dcgm-exporter reports device-level numbers, so every pod on the card shows the same values.
- Idle pods hold VRAM. A notebook somebody forgot about last Thursday is sitting on 20 GiB of an A100. That's the first place to look before you buy more GPUs.
- Pending pods, GPU edition. "0/3 nodes are available: 3 Insufficient nvidia.com/gpu" on a MIG-partitioned cluster where nobody advertises
nvidia.com/gpu. The scheduler's message is accurate and doesn't help.
I already had a CLI for this: kubectl-gpugo, on krew since May. But most of my team lives in a GUI, not a terminal. I checked the Freelens catalogue, found no GPU extension, and noticed that Freelens itself already reaches Prometheus through the apiserver proxy. So the plumbing I needed was already there. First commit and first npm publish were on the same day.
The Design Rule: Zero Cluster Footprint
The rule I set at the start: the extension must not need anything installed in the cluster.
That rules out the obvious approaches:
- Your own DaemonSet/agent. Now you need cluster-admin, an image pipeline, and an upgrade story for the thing that's supposed to help you upgrade things.
- Port-forwards. They work, but they're stateful, they leak, and they need a Main-process helper in Freelens. (My RabbitMQ extension does use them. That's a different post.)
- "Install Prometheus first." Lots of GPU clusters, especially dev and on-prem ones, have dcgm-exporter but no Prometheus scraping it.
What every GPU cluster does have: the NVIDIA GPU Operator's dcgm-exporter pods, serving /metrics, and a kube-apiserver that can proxy to any pod you're allowed to reach.
┌───────────────────────────── Freelens (Renderer process only) ─────────────────────────────┐
│ │
│ GpuStore (MobX, 20 s ticker, ref-counted by mounted views) │
│ │ │
│ ├─► GpuScraper.discover() ── podsApi.list() → filter → probe /metrics (cached 60 s)│
│ │ │
│ └─► fetch('/api-kube/api/v1/namespaces/NS/pods/POD:PORT/proxy/metrics') │
│ │ │
└─────────────────────────┼──────────────────────────────────────────────────────────────────┘
▼
kube-apiserver (pod proxy subresource — your RBAC, your kubeconfig)
│
┌───────────────┼──────────────────────┐
▼ ▼ ▼
dcgm-exporter per-process exporter Prometheus / Thanos / VM (fallback, service proxy)The whole extension runs in Freelens's Renderer process. The Main-process entry point is four lines with a comment that says "No main-process behaviour yet." No IPC and no sockets to manage. Every request goes through the same authenticated channel Freelens already uses for kubectl get pods.
How It Works
1. Discovery: find the exporters without being told
On first load, the scraper lists Running pods and keeps the ones whose name, image, or labels contain dcgm, gpu, nvidia, or cuda. For each candidate it picks a port (the prometheus.io/port annotation, then a port named metrics, then the first port), probes /metrics with a 5-second timeout, and classifies the response by content with a small built-in Prometheus text parser:
DCGM_FI_DEV_*→ dcgm-exportergpu_process_memory_bytes{namespace,pod,uuid,pid,gpu}→ per-process exporter (the best source when you have one)- vLLM metrics → an inference server (more on that below)
Discovery is cached for 60 seconds. If no exporter answers, it looks for a Service matching prometheus|thanos-query|vmselect|victoria-metrics|mimir-query|... and queries that through the service proxy instead, undoing the usual relabelling damage (exported_* labels, series stamped with the exporter pod instead of the workload, HA duplicates).
If auto-discovery guesses wrong, you can pin targets per cluster, like gpu-operator/nvidia-dcgm-exporter:9400.
2. Two views of the same GPU: allocation vs. utilization
The core idea: show what the scheduler thinks and what the hardware measures side by side.
| Allocation (scheduler) | Utilization (hardware) | |
|---|---|---|
| Source | node capacity/allocatable, pod limits/requests | DCGM gauges, per-process exporter |
| Counts | nvidia.com/gpu, nvidia.com/gpu.*, nvidia.com/mig-* | util, SM active, tensor, mem BW, VRAM, power |
| Answers | "Who reserved it?" | "Is anyone using it?" |
When those two disagree — a namespace holding 8 slices at 2% util — you've found your waste.
3. Attribution: the actual hard part
Getting the numbers is easy. Putting them on the right pod is the hard part. There are three tiers, in order of trust:
- Per-process exporter present → it knows exactly which PID on which GPU belongs to which pod. Power is split by VRAM share.
- DCGM with pod labels (
--kubernetes, the GPU Operator default) → one row per pod. On MIG the key isgpu:GPU_I_ID, andPROF_GR_ENGINE_ACTIVE × 100stands in for util because MIG slices don't reportGPU_UTIL. - DCGM without pod labels → one row per (node, GPU), with the candidate pods listed. It's honest about not knowing.
One rule covers most of the bugs below: device-level gauges are repeated once per pod, so you take the max, never the sum.
4. Polling without hammering the apiserver
A MobX store with a ref-counted subscribe(): the first mounted view starts a 20-second ticker, and the last one to unmount stops it. Data older than 30 seconds is marked stale. Nodes refresh every 60 seconds, not every tick. Idle history lives in memory only, capped at 6 hours. Nothing is written to the cluster, and only pinned targets are saved to localStorage.
What You Get
A GPU group in the sidebar with eight views: Pods (util, VRAM, power, GPU index like 0:8 for MIG slice 8 on GPU 0, shared ×N / time-sliced badges), Namespaces (the "whose GPUs" page), GPUs (one row per card or slice, with health), Idle & waste (under 5% util and over 256 MiB VRAM, with "idle for"), Allocation (per node, including MIG free per profile and unhealthy devices), Pending (with a "Why" hint, e.g. "you requested nvidia.com/gpu on a MIG-partitioned cluster — request a slice such as nvidia.com/mig-1g.10gb instead"), Inference (vLLM KV cache, queue depth, tokens/s, TTFT — marked saturated at KV ≥ 90% with requests waiting), and Exporters (discovery diagnostics).
Health comes from DCGM: XID errors (79 = "GPU has fallen off the bus"), double-bit ECC, row-remap failures, and throttle-reason bits. A device that doesn't export health fields shows "not exported", not "OK". I'd rather show "I don't know" than a green dot that isn't true.
The Bugs That Taught Me Something
All of these showed up on real hardware: an 8× A100 DGX in mig.strategy=mixed, with 47 slices plus one whole GPU, 29 GPU pods, about 805 W. A redacted capture of it now ships as a test fixture.
The API that existed only in the type definitions. KubeJsonApi.forCluster is in the Freelens typings and missing at runtime. The error was swallowed, so the first build said "No GPU metrics exporter found" on a cluster that had several. The fix was a plain relative fetch('/api-kube' + path). The bigger fix was showing probe errors to the user, which is why the Exporters page exists.
5,033 watts from an 805-watt server. DCGM repeats DCGM_FI_DEV_POWER_USAGE on every MIG slice. Sum per slice and an 8-GPU box draws like a small data center. Fix: count each physical card once.
Then every pod drew a whole card. Once the device total was right, per-pod power was still wrong: each 1g pod was charged about 104 W, and namespaces summed to about 3,000 W against 804 W real. Fix: each slice gets card W × slice_g / Σg on that card — 1g gets 1/7, 3g gets 3/7.
Capacity 1, requested 0, on a 48-device node. In mixed strategy, the node advertises nvidia.com/mig-1g.10gb, nvidia.com/mig-3g.40gb, and so on, not nvidia.com/gpu. I was only counting nvidia.com/gpu. After the fix: capacity 48, requested 29.
GPU 140%. Non-MIG cards emit both GPU_UTIL and PROF_GR_ENGINE_ACTIVE. Adding them double-counts. Now GPU_UTIL wins when both are present.
Time-slicing invented GPUs. nvidia.com/gpu.shared replicas were counted as extra physical devices. The fix excludes *.shared from device counts and adds a badge instead.
The node that was really an exporter pod. DCGM's Hostname label is the exporter pod name unless NODE_NAME is set. Now the node comes from where the exporter pod is actually scheduled.
Being a bad citizen on non-GPU clusters. When the Freelens maintainer reviewed it for adoption, they found I was LISTing pods, services, and nodes every 20 seconds on clusters with no GPUs at all, because an empty discovery result wasn't cached. Also, the keyword gpu matched the node name embedded in static pod names like kube-apiserver-gpu-node-1, so control-plane pods got probed every pass. Fixes: cache negative results for 60 seconds, strip the node name before keyword matching, and only open GPU sections in drawers when the pod or node can actually have a GPU. Measured with a Node drawer open for 65 seconds, the LIST count dropped from 3/3/23 to 0/0/15, the same as with the extension removed.
From Personal Repo to Freelens Org
I published it as @tal-naeh/freelens-gpu-extension, posted a catalogue proposal in the Freelens discussions, and a maintainer offered to adopt it into the freelensapp org. The repo moved with full history and was renamed @freelensapp/gpu-extension. It now ships with the org's release pipeline: provenance, SBOM, checksums, and Playwright integration tests on kind with a fake-GPU fixture. I still maintain it. The old package is deprecated and the old repo is archived with a "moved" notice, because two packages with the same name confuse people trying to install it.
Install: Freelens → Extensions (Cmd/Ctrl+Shift+E) → @freelensapp/gpu-extension. Requires Freelens ≥ 1.8.0.
Lessons
- Use what's already there. The apiserver pod-proxy gives you authenticated, RBAC-scoped access to every exporter in the cluster. You probably don't need your own agent.
- Device metrics aren't pod metrics. On MIG and time-slicing, anything measured per device is repeated per consumer. Max, don't sum, and split power deliberately.
- Show allocation and utilization together. Either one alone misleads. When they disagree, you've usually found the waste.
- "Not exported" beats a false OK. Only show a green status when the data supports it.
- Test on the ugliest hardware you have. Every interesting bug came from a mixed-strategy MIG box, not from kind.
- Being polite on clusters where you have nothing to do is a feature. If an extension spends API calls on a cluster with no GPUs, that's a bug.
- Show the errors. The worst bug was a swallowed exception that looked like "no data".
The GPU was never the hard part. The hard part was being honest about which pod is actually using it.
A note on how this was written: I used AI to help with phrasing and structure. The tone, the opinions, the war stories, the code, and everything else here are mine.