mirror of
https://github.com/supabase/supabase.git
synced 2026-09-09 03:19:36 +08:00
Adds an infrastructure/topology diagram for High Availability (Multigres) projects showing the real cluster topology — gateway tier, shard group, and the primary + read replicas inside it — on both the project homepage and the database/replication page, replacing the primary-only view and the "Replication unavailable" empty state. <img width="790" height="541" alt="Screenshot 2026-08-20 at 8 23 42 PM" src="https://github.com/user-attachments/assets/0bce21e3-2091-4285-84ca-60fdecb10d39" /> Addresses [FE-3717](https://linear.app/supabase/issue/FE-3717/show-replicas-in-replication-diagram). **Added:** - `data/ha-admin/` — read-only queries for the mgmt-api `/ha-admin/v1/{gateways,poolers,cells,databases}` multiadmin passthrough (ported from `bobbie/ha-stub`, re-authored to `queryOptions`). Responses are validated with zod at the fetch boundary (all fields optional per proto3 zero-value omission; enum-shaped fields stay plain strings so new proto values degrade gracefully); malformed payloads surface through the diagram's error fallback. - `HaTopology.utils.ts` — pure topology mapper (+ 26 unit tests): shard grouping, primary identified via `routingState.role` (deprecated `type` as fallback) with **failover-safe election** — when the outgoing and incoming primary briefly both claim `ROUTING_ROLE_PRIMARY`, the highest routing rule (coordinator term, leader subterm) wins, matching the multigateway's own election — plus status mapping onto the existing Healthy / Coming up / Going down / Unhealthy vocabulary, and an AZ formatter for `id.cell` that degrades to the raw cell name. - HA diagram nodes/edges: `Multigateway` card, shard group box with header pill (`Shard 1`, `Automatic failover` + tooltip), `Primary Database` card styled like the standard diagram's — neutral border, green icon chip (with the standard CPU / Disk / RAM footer — connections omitted until their meaning through the multigateway is confirmed), `Read Replica` cards, and the standard animated replication edges (status lives on the card badges). Poolers and gateways poll every 30s without re-running layout (topology projection + structural sharing). Drag-to-pan works through the shard group box, and the metrics footer's skeleton matches the loaded row height so the card doesn't shift. - Accessibility: the failover tooltip trigger is a keyboard-focusable button, status badges sit in stable `role="status"` live regions, the region flag is decorative (`alt=""`), and the edge dash/spinner animations respect `prefers-reduced-motion` (applied to the pipelines diagram's edges too). - Fallbacks: `AlertError` ("Failed to retrieve cluster topology") when either ha-admin query errors, and a "Cluster topology unavailable" empty state when the topology comes back empty — never a half-rendered diagram. **Changed:** - `InstanceConfiguration` is now topology-source-aware: it branches internally on `useHighAvailability()`, so both surfaces (homepage `TopSection` and the replication page) get the right diagram with no new wiring. The two-pass measured dagre layout moved into a shared `DiagramFlow`; `nodeTypes`/`edgeTypes` are module-level consts. - `getEdgeVisual` + the mid-edge icon chip lifted out of `ReplicationDiagram/Edges.tsx` into `components/ui/ReactFlow/EdgeVisual.tsx` so both diagrams derive edge icon + line style from one state object (no behavior change for the pipelines diagram). The primary card's CPU/Disk/RAM footer is likewise extracted into a shared `ComputeMetricsFooter`. - Fixes a latent relayout loop inherited from the region-box pattern: handing React Flow a freshly created (unmeasured) group node on every layout pass reset `nodesInitialized`, re-triggering the measured pass and `fitView` forever — which made the diagram snap back to center and effectively unpannable. The shared `DiagramFlow` now re-attaches known measurements to group nodes, which also covers the standard diagram's region boxes. - Standard diagram: the API Load Balancer → primary edge is now static — no data flows over it, the line only indicates a relation. - `database/replication` page: the HA early-return empty state is replaced by the diagram under a "High Availability cluster topology" header. Non-HA projects are untouched. **Intentional deviations from the mock** (for design review): 1. **No per-replica regions** — alpha replicas are one-per-cell inside a single region, so the mock's `eu-west-1` / `ap-southeast-1` on sibling replicas would be false. Availability zone per node, region shown once on the primary. 2. **"Primary Database", not "Main Database"** — matches the string both existing diagrams already ship, and the same component now renders both project types. 3. **No collapse chevron on the shard header** — alpha has exactly one shard; collapsing it would hide the whole diagram. The group box still ships; add collapse when `shards.length > 1`. 4. **Failover shown on the shard group, not replica cards** — failover is a cohort property; per-card badging would assert readiness we can't verify without a per-pooler `/status` fanout. 5. **Standard node/edge styling reused** (per review) — neutral primary border + green chip and the default animated edges instead of the mock's green ring and dashed green arrowed edges, keeping the HA and non-HA diagrams visually consistent. **Confirmed against a real local Multigres cluster:** cells are named `cell-1`/`cell-2`/… (not AZ-shaped — the AZ formatter falls back to the raw cell name as designed); `GET /platform/projects/{ref}/databases` returns only the primary row for HA projects; and the `/ha-admin` passthrough returns **each gateway/pooler record once per cell it fans out to** — the topology mapper dedupes by id, but worth confirming with @sbc-bobbie whether the backend should dedupe. **Known alpha limitation:** node health and the "replicating" edge state derive from the pooler's *topology record* (`lifecycleStatus`/`servingStatus`), not a live probe — a pooler that crashes without publishing a terminal state can read as healthy until the topology evicts its record, and a serving replica with paused replay still shows a green edge. This matches the existing replication diagram's semantics (`ACTIVE_HEALTHY` ⇒ animated edge). Live per-pooler signals (WAL receiver state, replay position) exist on `GET /poolers/{cell}/{name}/status` but need a per-pooler fanout — deliberately deferred, noted on `getPoolerStatus`. **Still to confirm** (doesn't block review): whether the `/ha-admin` passthrough is deployed to production or staging-only (if staging-only, this should get a flag before GA). ## To test Tested end-to-end locally against a real Multigres project (standard-project regression pass, HA creation flow, error fallback against real 500s, and full topology + polling + console checks against live multiadmin data): - **HA project homepage**: diagram card shows Multigateway → shard box (`Shard 1`, count badge, `Automatic failover` tooltip) → green-bordered Primary Database (region, AZ, size) + Read Replica cards (AZ), dashed green animated edges to healthy replicas. No flow/map toggle for HA. - **HA project → Database → Replication**: same diagram under a "High Availability cluster topology" header; no Destinations section; the old "Replication unavailable…" state is gone. - **Error path**: if `/ha-admin/v1/*` fails, both surfaces show "Failed to retrieve cluster topology" with Contact support — no partial diagram. - **Standard project regression**: homepage diagram (primary card, flow ⇄ map toggle round-trips), replication page (pipelines diagram + Destinations) all unchanged; zero requests to `/ha-admin/*`. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added High Availability topology diagrams to the Replication page. - Display gateways, primary databases, replicas, shards, statuses, regions, infrastructure details, and compute metrics. - Added observability links and live topology updates with loading, error, and unavailable states. - **Bug Fixes** - Improved handling of incomplete infrastructure identities and unexpected data. - Corrected topology layout, node spacing, and visual edge behavior. - **Accessibility** - Reduced-motion preferences now disable diagram animations and loading effects. - Improved status announcements for assistive technologies. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Alaister Young <10985857+alaister@users.noreply.github.com>
204 lines
6.4 KiB
TypeScript
204 lines
6.4 KiB
TypeScript
import dagre from '@dagrejs/dagre'
|
|
import { Edge, Node, Position } from '@xyflow/react'
|
|
import { groupBy } from 'lodash'
|
|
import { AWS_REGIONS, AWS_REGIONS_KEYS } from 'shared-data'
|
|
|
|
import {
|
|
AVAILABLE_REPLICA_REGIONS,
|
|
AWS_REGIONS_COORDINATES,
|
|
NODE_HEIGHT_FALLBACKS,
|
|
NODE_SEP,
|
|
NODE_WIDTH,
|
|
ReplicaNodeData,
|
|
} from './InstanceConfiguration.constants'
|
|
import type { LoadBalancer } from '@/data/read-replicas/load-balancers-query'
|
|
import type { Database } from '@/data/read-replicas/replicas-query'
|
|
|
|
// [Joshen] Just FYI the nodes generation assumes each project only has one load balancer
|
|
// Will need to change if this eventually becomes otherwise
|
|
|
|
export const generateNodes = ({
|
|
primary,
|
|
replicas,
|
|
loadBalancers,
|
|
}: {
|
|
primary: Database
|
|
replicas: Database[]
|
|
loadBalancers: LoadBalancer[]
|
|
}): Node[] => {
|
|
const position = { x: 0, y: 0 }
|
|
const regions = groupBy(replicas, (d) => {
|
|
const region = AVAILABLE_REPLICA_REGIONS.find((region) => d.region.includes(region.region))
|
|
return region?.key
|
|
})
|
|
|
|
const loadBalancer = loadBalancers.find((x) =>
|
|
x.databases.some((db) => db.identifier === primary.identifier)
|
|
)
|
|
const loadBalancerNode: Node | undefined =
|
|
loadBalancer !== undefined
|
|
? {
|
|
position,
|
|
id: 'load-balancer',
|
|
type: 'LOAD_BALANCER',
|
|
data: {
|
|
numDatabases: loadBalancer.databases.length,
|
|
},
|
|
}
|
|
: undefined
|
|
|
|
// [Joshen] We should be finding from AVAILABLE_REPLICA_REGIONS instead
|
|
// but because the new regions (zurich, stockholm, ohio, paris) dont have
|
|
// coordinates yet in AWS_REGIONS_COORDINATES - we'll need to add them in once
|
|
// they are ready to spin up coordinates for
|
|
const primaryRegion = Object.keys(AWS_REGIONS)
|
|
.map((key) => {
|
|
return {
|
|
key: key as AWS_REGIONS_KEYS,
|
|
name: AWS_REGIONS?.[key as AWS_REGIONS_KEYS].displayName,
|
|
region: AWS_REGIONS?.[key as AWS_REGIONS_KEYS].code,
|
|
coordinates: AWS_REGIONS_COORDINATES[key],
|
|
}
|
|
})
|
|
.find((region) => primary.region.includes(region.region))
|
|
|
|
// [Joshen] Once we have the coordinates for Zurich and Stockholm, we can remove the above
|
|
// and uncomment below for better simplicity
|
|
// const primaryRegion = AVAILABLE_REPLICA_REGIONS.find((region) =>
|
|
// primary.region.includes(region.region)
|
|
// )
|
|
|
|
const primaryNode: Node = {
|
|
position,
|
|
id: primary.identifier,
|
|
type: 'PRIMARY',
|
|
data: {
|
|
id: primary.identifier,
|
|
region: primaryRegion ?? { name: primary.region },
|
|
provider: primary.cloud_provider,
|
|
inserted_at: primary.inserted_at,
|
|
computeSize: primary.size,
|
|
status: primary.status,
|
|
numReplicas: replicas.length,
|
|
numRegions: Object.keys(regions).length,
|
|
hasLoadBalancer: loadBalancer !== undefined,
|
|
},
|
|
}
|
|
|
|
const replicaNodes: Node[] = replicas
|
|
.sort((a, b) => (a.region > b.region ? 1 : -1))
|
|
.map((database) => {
|
|
const region = AVAILABLE_REPLICA_REGIONS.find((region) =>
|
|
database.region.includes(region.region)
|
|
)
|
|
|
|
return {
|
|
position,
|
|
id: database.identifier,
|
|
type: 'READ_REPLICA',
|
|
data: {
|
|
id: database.identifier,
|
|
region,
|
|
provider: database.cloud_provider,
|
|
inserted_at: database.inserted_at,
|
|
computeSize: database.size,
|
|
status: database.status,
|
|
},
|
|
}
|
|
})
|
|
|
|
return [
|
|
...(loadBalancerNode !== undefined ? [loadBalancerNode] : []),
|
|
primaryNode,
|
|
...replicaNodes,
|
|
]
|
|
}
|
|
|
|
export const getDagreNodeHeight = (node: Node) => {
|
|
if (node.measured?.height) return node.measured.height
|
|
return NODE_HEIGHT_FALLBACKS[node.type ?? ''] ?? 100
|
|
}
|
|
|
|
export const getDagreGraphLayout = (
|
|
nodes: Node[],
|
|
edges: Edge[],
|
|
{ ranksep = 60 }: { ranksep?: number } = {}
|
|
) => {
|
|
const dagreGraph = new dagre.graphlib.Graph()
|
|
dagreGraph.setDefaultEdgeLabel(() => ({}))
|
|
dagreGraph.setGraph({ rankdir: 'TB', ranksep, nodesep: NODE_SEP })
|
|
|
|
nodes.forEach((node) => {
|
|
dagreGraph.setNode(node.id, {
|
|
width: NODE_WIDTH / 2,
|
|
height: getDagreNodeHeight(node),
|
|
})
|
|
})
|
|
|
|
edges.forEach((edge) => dagreGraph.setEdge(edge.source, edge.target))
|
|
|
|
dagre.layout(dagreGraph)
|
|
|
|
nodes.forEach((node) => {
|
|
const nodeWithPosition = dagreGraph.node(node.id)
|
|
node.targetPosition = Position.Top
|
|
node.sourcePosition = Position.Bottom
|
|
// We are shifting the dagre node position (anchor=center center) to the top left
|
|
// so it matches the React Flow node anchor point (top left).
|
|
node.position = {
|
|
x: nodeWithPosition.x - nodeWithPosition.width / 2,
|
|
y: nodeWithPosition.y - nodeWithPosition.height / 2,
|
|
}
|
|
|
|
return node
|
|
})
|
|
|
|
return { nodes, edges }
|
|
}
|
|
|
|
/**
|
|
* [Joshen] This is some customized logic to add region nodes as "subflow" as dagre doesn't support
|
|
* subflows, and I didn't want to go down a rabbit hole with the other layout libraries that react-flow
|
|
* supports. Definitely some things to improve in the future
|
|
* - Allow setting max number of nodes per row, so that the chart does not become too horizontally sparse
|
|
* when many many replicas created
|
|
* - Nodes are a bit too spaced out between each other within a region
|
|
*/
|
|
export const addRegionNodes = (nodes: Node[], edges: Edge[]) => {
|
|
const regionNodes: Node[] = []
|
|
const replicaNodes = nodes.filter(
|
|
(node) => node.type === 'READ_REPLICA'
|
|
) as Node<ReplicaNodeData>[]
|
|
|
|
const nodesByRegion = groupBy(replicaNodes, (node) => node.data.region.key)
|
|
Object.entries(nodesByRegion).map(([key, value]) => {
|
|
const region = AVAILABLE_REPLICA_REGIONS.find((r) => r.key === key)
|
|
const nodeXPositions = value.map((x) => x.position.x)
|
|
const nodeYPositions = value.map((x) => x.position.y)
|
|
|
|
const minX = Math.min(...nodeXPositions)
|
|
const maxX = Math.max(...nodeXPositions)
|
|
|
|
const minY = Math.max(...nodeYPositions)
|
|
|
|
const regionNode: Node = {
|
|
id: key,
|
|
position: { x: minX - 10, y: minY - 10 },
|
|
width: maxX - minX + NODE_WIDTH / 2,
|
|
type: 'REGION',
|
|
data: { region, numReplicas: value.length },
|
|
}
|
|
regionNodes.push(regionNode)
|
|
})
|
|
|
|
return { nodes: [...regionNodes, ...nodes], edges }
|
|
}
|
|
|
|
export const formatSeconds = (value: number) => {
|
|
const hours = ~~(value / 3600)
|
|
const minutes = Math.floor((value % 3600) / 60)
|
|
const seconds = Math.floor(value % 60)
|
|
|
|
return `${hours > 0 ? `${hours}h` : ''} ${minutes > 0 ? `${minutes}m` : ''} ${seconds > 0 ? `${seconds}s` : ''}`.trim()
|
|
}
|