Commit Graph

2 Commits

Author SHA1 Message Date
claude[bot]
10950d286b chore(studio): address review comments on Multigres topology diagram (#49592)
<!-- ccr-slack-attribution -->
_Requested by **Alaister Young** · [Slack
thread](https://supabase.slack.com/archives/C0161K73J1J/p1787738517181409?thread_ts=1787635785.354489&cid=C0161K73J1J)_

Follow-up to #49298, which was squash-merged before @joshenlim's last
review round was addressed. Picking up the review comments here.

## I have read the
[CONTRIBUTING.md](https://github.com/supabase/supabase/blob/master/CONTRIBUTING.md)
file.

YES

## What kind of change does this PR introduce?

Chore — dead code removal and comment corrections. No behavior change.

## What is the current behavior?

Three of @joshenlim's review comments on #49298 are still open on
master:

- [Dead code in
`ha-cluster-cells-query.ts`](https://github.com/supabase/supabase/pull/49298#discussion_r3861029414)
— "seems to be dead code? no one's consuming this file"
- [Dead code in
`ha-cluster-databases-query.ts`](https://github.com/supabase/supabase/pull/49298#discussion_r3861031230)
— "likewise - seems to be dead code"
- [`STATUS_BADGE_VARIANTS`
statuses](https://github.com/supabase/supabase/pull/49298#discussion_r3861095281)
— "just to sanity check these are the only statuses? are there any
failure states? e.g 'Failed'"

Concretely, on master today:

- `apps/studio/data/ha-admin/ha-cluster-cells-query.ts` and
`apps/studio/data/ha-admin/ha-cluster-databases-query.ts` ship query
options that nothing imports. The diagram only reads `/poolers` and
`/gateways`.
- `multipoolerSchema.lifecycleStatus` is an undocumented `z.string()`,
while its neighbours `type` and `servingStatus` both list their expected
proto values in a comment.
- The comment above `HA_POOLER_STATUS_LABELS` claims the labels "Matches
the status vocabulary of the read replica surfaces (getStatusLabel)".
They don't — `getStatusLabel` in `ReadReplicas/ReadReplicas.utils.ts`
also returns `Failed`, `Restarting`, `Resizing` and `Restoring`, none of
which the HA labels have.

## What is the new behavior?

- Deleted both dead query modules and pruned the orphaned `cells` and
`databases` factories from `haAdminKeys`, keeping `poolers` and
`gateways`. Verified by grep that neither file name nor any of their
exported symbols (`haClusterCellsQueryOptions`, `HaClusterCellsData`,
`haClusterDatabasesQueryOptions`, `HaClusterDatabasesData`, the
`*Variables`/`*Error` types) nor `haAdminKeys.cells` /
`haAdminKeys.databases` has a single reference left anywhere outside the
deleted files. No re-export shims left behind. `get-ha-admin.ts` stays —
`poolers` and `gateways` still use it.
- Documented `lifecycleStatus` against the actual enum,
`PoolerLifecycleStatus` in [multigres
`proto/clustermetadata.proto`](https://github.com/multigres/multigres/blob/main/proto/clustermetadata.proto):
`LIFECYCLE_UNKNOWN` (zero value, omitted from JSON) | `STARTING` |
`ACTIVE` | `STOPPING` | `SHUTDOWN` | `QUARANTINED`. `getPoolerStatus`
already maps every member.
- Reworded the `HA_POOLER_STATUS_LABELS` comment to say the labels are a
subset drawn from the read replica vocabulary rather than a match for
it, and noted where the read replica `Failed` lands on the HA side.

**On the `Failed` question:** the answer from the proto is that there is
no dedicated failure member. The terminal states are `QUARANTINED` — the
pooler "has given up trying to become a healthy replica: it cannot
automatically recover to a functioning state (e.g. it could not complete
a pg_rewind, could not restore from backup to start postgres, or fell
irrecoverably behind on replication)", kept alive for forensics — and
`SHUTDOWN`, "durably down". Both already map to `unhealthy` / the
`Unhealthy` warning badge, so the four statuses on
`STATUS_BADGE_VARIANTS` are complete for the enum as it stands. If we'd
rather show `QUARANTINED` as its own `Failed` status with a destructive
badge (matching the read replica surface), that's a small follow-up — a
product/copy call rather than a gap, so not folded in here.

**Not included: [the replication page UX
comment](https://github.com/supabase/supabase/pull/49298#discussion_r3861055216)**
("is there any other content we plan to add here? it feels empty atm...
it's just a repeat of the home page + settings/infrastructure").
@joshenlim flagged that one himself as "UX feedback which can be
addressed separately". It's a product and IA question about what that
page is for, not something to answer with a code change here — leaving
it for @alaister and design.

## Additional context

- Verified locally: `tsc --noEmit` (0 errors), ESLint on the touched
files (clean), Prettier check (clean), and `HaTopology.utils.test.ts` +
`HaInstanceConfiguration.utils.test.ts` (26/26 passing). CI is green as
well.
- Exhaustive grep across the repo (excluding `node_modules`/`.git`/build
output, covering `apps/**` incl. `lite-studio`, `packages/**` and
`e2e/**`) confirmed zero remaining references to the deleted files,
their exported symbols, and the removed key factories.
- No test changes: the diff deletes unreferenced code and edits comments
only, so there's no new behavior to cover. `HaTopology.utils.test.ts`
already pins every `lifecycleStatus` value listed in the new comment.

---
_Generated by [Claude
Code](https://claude.ai/code/session_012StGVQSmPzpGTXrduo9Xyi)_

---------

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Alaister Young <10985857+alaister@users.noreply.github.com>
2026-08-27 08:52:52 +00:00
Alaister Young
928049ce2c [FE-3717] feat(studio): Multigres cluster topology diagram (#49298)
Adds an infrastructure/topology diagram for High Availability
(Multigres) projects showing the real cluster topology — gateway tier,
shard group, and the primary + read replicas inside it — on both the
project homepage and the database/replication page, replacing the
primary-only view and the "Replication unavailable" empty state.

<img width="790" height="541" alt="Screenshot 2026-08-20 at 8 23 42 PM"
src="https://github.com/user-attachments/assets/0bce21e3-2091-4285-84ca-60fdecb10d39"
/>

Addresses
[FE-3717](https://linear.app/supabase/issue/FE-3717/show-replicas-in-replication-diagram).

**Added:**

- `data/ha-admin/` — read-only queries for the mgmt-api
`/ha-admin/v1/{gateways,poolers,cells,databases}` multiadmin passthrough
(ported from `bobbie/ha-stub`, re-authored to `queryOptions`). Responses
are validated with zod at the fetch boundary (all fields optional per
proto3 zero-value omission; enum-shaped fields stay plain strings so new
proto values degrade gracefully); malformed payloads surface through the
diagram's error fallback.
- `HaTopology.utils.ts` — pure topology mapper (+ 26 unit tests): shard
grouping, primary identified via `routingState.role` (deprecated `type`
as fallback) with **failover-safe election** — when the outgoing and
incoming primary briefly both claim `ROUTING_ROLE_PRIMARY`, the highest
routing rule (coordinator term, leader subterm) wins, matching the
multigateway's own election — plus status mapping onto the existing
Healthy / Coming up / Going down / Unhealthy vocabulary, and an AZ
formatter for `id.cell` that degrades to the raw cell name.
- HA diagram nodes/edges: `Multigateway` card, shard group box with
header pill (`Shard 1`, `Automatic failover` + tooltip), `Primary
Database` card styled like the standard diagram's — neutral border,
green icon chip (with the standard CPU / Disk / RAM footer — connections
omitted until their meaning through the multigateway is confirmed),
`Read Replica` cards, and the standard animated replication edges
(status lives on the card badges). Poolers and gateways poll every 30s
without re-running layout (topology projection + structural sharing).
Drag-to-pan works through the shard group box, and the metrics footer's
skeleton matches the loaded row height so the card doesn't shift.
- Accessibility: the failover tooltip trigger is a keyboard-focusable
button, status badges sit in stable `role="status"` live regions, the
region flag is decorative (`alt=""`), and the edge dash/spinner
animations respect `prefers-reduced-motion` (applied to the pipelines
diagram's edges too).
- Fallbacks: `AlertError` ("Failed to retrieve cluster topology") when
either ha-admin query errors, and a "Cluster topology unavailable" empty
state when the topology comes back empty — never a half-rendered
diagram.

**Changed:**

- `InstanceConfiguration` is now topology-source-aware: it branches
internally on `useHighAvailability()`, so both surfaces (homepage
`TopSection` and the replication page) get the right diagram with no new
wiring. The two-pass measured dagre layout moved into a shared
`DiagramFlow`; `nodeTypes`/`edgeTypes` are module-level consts.
- `getEdgeVisual` + the mid-edge icon chip lifted out of
`ReplicationDiagram/Edges.tsx` into
`components/ui/ReactFlow/EdgeVisual.tsx` so both diagrams derive edge
icon + line style from one state object (no behavior change for the
pipelines diagram). The primary card's CPU/Disk/RAM footer is likewise
extracted into a shared `ComputeMetricsFooter`.
- Fixes a latent relayout loop inherited from the region-box pattern:
handing React Flow a freshly created (unmeasured) group node on every
layout pass reset `nodesInitialized`, re-triggering the measured pass
and `fitView` forever — which made the diagram snap back to center and
effectively unpannable. The shared `DiagramFlow` now re-attaches known
measurements to group nodes, which also covers the standard diagram's
region boxes.
- Standard diagram: the API Load Balancer → primary edge is now static —
no data flows over it, the line only indicates a relation.
- `database/replication` page: the HA early-return empty state is
replaced by the diagram under a "High Availability cluster topology"
header. Non-HA projects are untouched.

**Intentional deviations from the mock** (for design review):

1. **No per-replica regions** — alpha replicas are one-per-cell inside a
single region, so the mock's `eu-west-1` / `ap-southeast-1` on sibling
replicas would be false. Availability zone per node, region shown once
on the primary.
2. **"Primary Database", not "Main Database"** — matches the string both
existing diagrams already ship, and the same component now renders both
project types.
3. **No collapse chevron on the shard header** — alpha has exactly one
shard; collapsing it would hide the whole diagram. The group box still
ships; add collapse when `shards.length > 1`.
4. **Failover shown on the shard group, not replica cards** — failover
is a cohort property; per-card badging would assert readiness we can't
verify without a per-pooler `/status` fanout.
5. **Standard node/edge styling reused** (per review) — neutral primary
border + green chip and the default animated edges instead of the mock's
green ring and dashed green arrowed edges, keeping the HA and non-HA
diagrams visually consistent.

**Confirmed against a real local Multigres cluster:** cells are named
`cell-1`/`cell-2`/… (not AZ-shaped — the AZ formatter falls back to the
raw cell name as designed); `GET /platform/projects/{ref}/databases`
returns only the primary row for HA projects; and the `/ha-admin`
passthrough returns **each gateway/pooler record once per cell it fans
out to** — the topology mapper dedupes by id, but worth confirming with
@sbc-bobbie whether the backend should dedupe.

**Known alpha limitation:** node health and the "replicating" edge state
derive from the pooler's *topology record*
(`lifecycleStatus`/`servingStatus`), not a live probe — a pooler that
crashes without publishing a terminal state can read as healthy until
the topology evicts its record, and a serving replica with paused replay
still shows a green edge. This matches the existing replication
diagram's semantics (`ACTIVE_HEALTHY` ⇒ animated edge). Live per-pooler
signals (WAL receiver state, replay position) exist on `GET
/poolers/{cell}/{name}/status` but need a per-pooler fanout —
deliberately deferred, noted on `getPoolerStatus`.

**Still to confirm** (doesn't block review): whether the `/ha-admin`
passthrough is deployed to production or staging-only (if staging-only,
this should get a flag before GA).

## To test

Tested end-to-end locally against a real Multigres project
(standard-project regression pass, HA creation flow, error fallback
against real 500s, and full topology + polling + console checks against
live multiadmin data):

- **HA project homepage**: diagram card shows Multigateway → shard box
(`Shard 1`, count badge, `Automatic failover` tooltip) → green-bordered
Primary Database (region, AZ, size) + Read Replica cards (AZ), dashed
green animated edges to healthy replicas. No flow/map toggle for HA.
- **HA project → Database → Replication**: same diagram under a "High
Availability cluster topology" header; no Destinations section; the old
"Replication unavailable…" state is gone.
- **Error path**: if `/ha-admin/v1/*` fails, both surfaces show "Failed
to retrieve cluster topology" with Contact support — no partial diagram.
- **Standard project regression**: homepage diagram (primary card, flow
⇄ map toggle round-trips), replication page (pipelines diagram +
Destinations) all unchanged; zero requests to `/ha-admin/*`.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
  - Added High Availability topology diagrams to the Replication page.
- Display gateways, primary databases, replicas, shards, statuses,
regions, infrastructure details, and compute metrics.
- Added observability links and live topology updates with loading,
error, and unavailable states.

- **Bug Fixes**
- Improved handling of incomplete infrastructure identities and
unexpected data.
  - Corrected topology layout, node spacing, and visual edge behavior.

- **Accessibility**
- Reduced-motion preferences now disable diagram animations and loading
effects.
  - Improved status announcements for assistive technologies.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Alaister Young <10985857+alaister@users.noreply.github.com>
2026-08-26 16:51:36 +08:00