·9 min read

Stop Port-Forwarding to 15672: A Cluster-Native RabbitMQ Console for Freelens

Why I built a RabbitMQ extension for Freelens, how it finds brokers and credentials and tunnels to the Management API from inside the desktop app, and what it takes to make dead-letter replay safe.

Here's the routine. Something's off with a consumer. You open a terminal, kubectl port-forward svc/rabbitmq 15672, run kubectl get secret ... | base64 -d for the password, paste it into the Management UI, find the queue, see 40,000 messages and zero consumers. Then you need to check the same thing in the next environment, so you kill the port-forward, switch context, dig out a different secret, and log in again. I did this often enough that I built a fix into the tool I already had open.


TL;DR

  • The RabbitMQ Management UI is great. Getting to it on Kubernetes isn't: port-forwards, base64'd secrets, one login per environment.
  • I built a Freelens extension that discovers brokers, pulls credentials from Secrets, and tunnels to the Management API itself. You pick the cluster in Freelens and the queues are already there.
  • Architecture: a Main-process engine (discovery, credentials, SPDY port-forward, HTTP client, sessions) and a Renderer UI (React + Freelens components), connected by typed IPC. Secrets never leave Main.
  • It's read-only by default. Writes (publish, purge, delete, replay) sit behind an in-memory Write Mode that's checked on every call in Main, not just hidden in the UI.
  • Dead-letter replay copies, never moves. The DLQ stays the source of truth, so a replay can't lose a message.
  • It went from empty folder to npm in one working day, and was adopted into the official freelensapp org ten days later.

Why This Exists

The Management UI isn't the problem. Everything you have to do before you can use it is:

  • Port-forward by hand. Keep a terminal open, and remember which local port belongs to which environment.
  • Secret archaeology. The Cluster Operator puts credentials in <cluster>-default-user. Bitnami puts them somewhere else. A hand-rolled StatefulSet might use a third place.
  • Log in again for every environment. Five environments means five logins, five browser tabs, and constantly checking that you're looking at QA and not prod.

But Freelens already knows which cluster you're on and already has your kubeconfig. It's missing a way to talk to the broker.

There was a second reason, which only showed up once the first version existed. The interesting RabbitMQ failures are silent: a memory alarm blocking publishers, unacked messages pinning RAM, a backlog that nobody consumes, publishes going nowhere because nothing is bound. Once a broker view sits in the same app as your pods and deployments, you can start pointing at those problems instead of only listing queues.


Architecture: Two Processes, One Rule

Freelens is Electron, so an extension can run code in two places. The Renderer draws the UI. The Main process has Node: sockets, the filesystem, real HTTP. The RabbitMQ extension uses both, and the line between them is a security boundary.

 ┌────────────── Renderer (React 17 + Freelens UI) ──────────────┐
 │  Clusters · Overview · Health · Queues · Exchanges · Conns     │
 │  useResource(key, loader, refreshMs)  → polls via IPC          │
 │  sees: queue stats, usernames, credential *hints* — no secrets │
 └───────────────────────────────┬───────────────────────────────┘
                                 │  typed IPC (src/common/ipc.ts)
                                 │  rabbitmq:discover / queues / purge-queue / replay-messages ...
 ┌───────────────────────────────▼───────────────────────────────┐
 │  Main (Node)                                                   │
 │   discovery.ts      → RabbitmqCluster CRs, Services, workloads │
 │   credentials.ts    → Secrets → Basic auth (+ operator CA)     │
 │   pod-resolver.ts   → one Ready broker pod, named port → number│
 │   port-forward.ts   → 127.0.0.1:random  ⇄  SPDY  ⇄  pod:15672  │
 │   session-manager   → one session per (cluster, target)        │
 │   write-mode gate   → checked on EVERY write, server-side      │
 └───────────────────────────────┬───────────────────────────────┘
                                 ▼
                    kube-apiserver → broker pod → Management API

The one rule: credentials live in Main memory and never cross the IPC boundary. The renderer gets a hint ("from Secret orders-mq-default-user") and a username, and that's all. If someone ever pokes at the renderer, there's no password there to find.

1. Discovery: three ways a broker shows up

RabbitMQ gets deployed in many different ways, so discovery checks several sources in order:

  1. Operator RabbitmqCluster CRs (rabbitmq.com/v1beta1), but only if the CRD exists, so non-operator clusters don't get errors.
  2. Services that look like RabbitMQ: "rabbit" in the name or labels, exposing 15672/15671 next to 5672/5671. It labels each one as Bitnami chart, Helm release, or In-cluster Service, and skips the operator's own metrics Service.
  3. Workloads, where it reads env var names like RABBITMQ_DEFAULT_USER and RABBITMQ_USERNAME to work out which Secret holds the credentials. It records where the credentials are, never the values.

Manual credentials are the last fallback. A cluster can have several brokers, so there's a target selector, and the choice is remembered per Kubernetes cluster.

2. The tunnel

Main picks a Ready broker pod, resolves the named port, and opens a local net.Server on 127.0.0.1 at a random port. Each incoming connection is piped through a SPDY stream using @kubernetes/client-node's PortForward. The Management API client is plain node:http/https with Basic auth, and it pages through /api/queues 500 at a time, capped at 5,000 items.

One trap I didn't expect: Freelens's cluster kubeconfig is a proxy kubeconfig, and it can't port-forward. The tunnel has to use the catalog entity's original kubeconfig path, falling back to the default kubeconfig plus the context name.

3. Sessions

There's one session per (cluster, broker), validated with whoami() when it opens and closed after 5 minutes idle. If the tunnel dies, the session reopens it once automatically. Live pages poll every 5 seconds, and Health every 15.

4. Errors across IPC

Electron keeps only error.message when an error crosses IPC. The type, code, and anything structured are dropped. So RabbitmqError is serialized as JSON inside the message and parsed back on the other side. That's how the UI can tell "credentials rejected, here's a login form" apart from "broker unreachable, here's a retry button".


Safe Writes

A tool that can purge a production queue in one click will eventually purge a production queue. So:

  • Read-only by default. Every session starts that way.
  • Write Mode is a per-session toggle kept only in memory. It needs a confirmation dialog to turn on and resets when the session ends.
  • Main enforces it on every call. Hiding buttons in the renderer would only be a UI convenience, not a guard. assertWriteMode runs before publish, purge, delete, and replay.
  • Destructive actions ask again, showing the object's name.
  • Peek is non-destructive. It's always ack_requeue_true, up to 50 messages, with payloads truncated at 64 KiB. I checked on RabbitMQ 4.3.5 that this does not count against a quorum queue's delivery limit. reject_requeue_true does count, so the difference matters.

Dead-Letter Replay: Copy, Never Move

Replay was the most-requested feature, and it was the one I was most careful with. The extension decodes x-death (reason, source queue, count, timestamps) and lets you search and filter dead letters by reason. Replay works like this:

It publishes copies. The originals stay in the DLQ. Nothing is acked or removed, so the worst case of a bad replay is duplicates, never loss. The trade-off is that there's no idempotency: replay twice and you publish twice. The UI marks replayed messages, and "select all" skips them.

Two destinations, computed in Main from each message's own headers:

ModeExchangeRouting keyWho gets it
failed-queue (default)default exchangex-first-death-queueonly the queue that dead-lettered it
original-exchangeoriginal exchangefirst key from x-deatheveryone bound — fan-out warning shown

Strip the death headers. x-death, x-first-death-*, x-last-death-*, x-delivery-count, x-acquired-count. If you leave a stale x-death on, the broker keeps it as-is, and if the message fails again, cycle detection can drop it silently. In failed-queue mode CC/BCC are stripped too, because a CC header on a default-exchange publish duplicates the message into whatever queue the CC key names. I verified all of this against a real 4.3.5 broker, not just the docs.

Batches are capped at 50 and published one by one. Each message reports routed, unroutable, skipped (for example, a truncated payload), or failed, so one bad message doesn't stop the rest. Only messages visible under the current filter are replayed. I found that bug from a screenshot, where "selected" included rows the filter had hidden.


The Bugs That Taught Me Something

The silent port-forward hang. client-node's PortForward doesn't close the local socket when the pod side closes the tunnel. With HTTP keep-alive, the next request reuses a dead socket and hangs until timeout. The fix was keepAlive: false. That's one handshake per request over a localhost tunnel, which costs nothing.

RabbitMQ 4.x changed the rules. Transient non-exclusive queues are refused, so every test fixture had to become durable. Purge (DELETE .../contents) returns 406 to a strict Accept header, so the client now sends application/json, */*;q=0.1.

"I purged it and it's still full." Management API stats lag about 5 seconds. The queue drawer wasn't polling, so it kept showing the old count. Now the drawer polls like the page does.

Health checks that cried wolf. The first review pass of Health and Clients found 17 bugs. Some examples:

  • Streams flagged as "stuck backlogs". Streams are meant to keep their messages.
  • Prefetch-1 consumers flagged as stuck. Now it takes 5 minutes full with no acks.
  • One node down on a 3-node broker produced a warning per quorum queue, about 165 on a real broker. Those are now folded into a single node-down finding.

Freelens UI quirks. The table header and rows had different padding, so columns didn't line up. A title prop rendered as the cell's content instead of a tooltip. Drawers closed on every refresh, because their state came from the URL, and writing to the URL re-mounted the page. The dropdown rendered behind the sticky header (both were z-index 1). And after an upgrade the extension looked unchanged, because Freelens keeps the old renderer bundle until a full Cmd+Q. That one got a permanent version badge in every page header.

The finding that justified the Clients tab. Grouping connections by workload on a real broker showed that one autoscaler held about 45% of all connections: 120 of 268. Every KEDA ScaledObject trigger using protocol: amqp opens its own connection. The Management UI shows those 120 connections as 120 separate rows, so the pattern only shows up once you group them by workload.


Packaging and the Org Move

It's built with electron-vite into CommonJS bundles. React and the Freelens API are mapped to the host's globals and everything else is bundled, so there are zero runtime dependencies. A smoke script loads the Main bundle at build time to catch "No handler registered" before a user does. Tests run on vitest plus a Docker e2e against rabbitmq:4-management. The org's CI adds Playwright inside Freelens on kind.

It started as @tal-naeh/freelens-rabbitmq-extension, published the same day I started it. I proposed it to the Freelens org that day, it was adopted ten days later with full history, and it now ships as @freelensapp/rabbitmq-extension. Install: Freelens → Extensions (Cmd/Ctrl+Shift+E) → @freelensapp/rabbitmq-extension.


Lessons

  • Fix the steps before the dashboard. The value wasn't a better queue table. It was getting rid of the port-forward, the secret lookup, and the login.
  • Put the security boundary in the process model. Secrets in Main, hints in the Renderer. Write Mode enforced in Main. The UI can be convenient because it isn't what keeps things safe.
  • Replay should copy, not move. If the original stays in the DLQ, the worst case is a duplicate, not a lost message.
  • Test against the real broker version. Half the replay rules (CC duplication, x-death preservation, delivery-limit accounting) only showed up on 4.3.5, not in the docs.
  • Health rules need a review as much as code does. A check that fires 165 times for one dead node teaches people to ignore it.
  • Group by owner. "268 connections" doesn't tell you anything. "One autoscaler holds 120 of them" tells you what to fix.

Most RabbitMQ problems turn out to be visible once you have an easy way to look at the broker. This extension is my way of making that easy.


A note on how this was written: I used AI to help with phrasing and structure. The tone, the opinions, the war stories, the code, and everything else here are mine.