mirror of
https://github.com/supabase/supabase.git
synced 2026-09-09 03:19:36 +08:00
Adds an infrastructure/topology diagram for High Availability (Multigres) projects showing the real cluster topology — gateway tier, shard group, and the primary + read replicas inside it — on both the project homepage and the database/replication page, replacing the primary-only view and the "Replication unavailable" empty state. <img width="790" height="541" alt="Screenshot 2026-08-20 at 8 23 42 PM" src="https://github.com/user-attachments/assets/0bce21e3-2091-4285-84ca-60fdecb10d39" /> Addresses [FE-3717](https://linear.app/supabase/issue/FE-3717/show-replicas-in-replication-diagram). **Added:** - `data/ha-admin/` — read-only queries for the mgmt-api `/ha-admin/v1/{gateways,poolers,cells,databases}` multiadmin passthrough (ported from `bobbie/ha-stub`, re-authored to `queryOptions`). Responses are validated with zod at the fetch boundary (all fields optional per proto3 zero-value omission; enum-shaped fields stay plain strings so new proto values degrade gracefully); malformed payloads surface through the diagram's error fallback. - `HaTopology.utils.ts` — pure topology mapper (+ 26 unit tests): shard grouping, primary identified via `routingState.role` (deprecated `type` as fallback) with **failover-safe election** — when the outgoing and incoming primary briefly both claim `ROUTING_ROLE_PRIMARY`, the highest routing rule (coordinator term, leader subterm) wins, matching the multigateway's own election — plus status mapping onto the existing Healthy / Coming up / Going down / Unhealthy vocabulary, and an AZ formatter for `id.cell` that degrades to the raw cell name. - HA diagram nodes/edges: `Multigateway` card, shard group box with header pill (`Shard 1`, `Automatic failover` + tooltip), `Primary Database` card styled like the standard diagram's — neutral border, green icon chip (with the standard CPU / Disk / RAM footer — connections omitted until their meaning through the multigateway is confirmed), `Read Replica` cards, and the standard animated replication edges (status lives on the card badges). Poolers and gateways poll every 30s without re-running layout (topology projection + structural sharing). Drag-to-pan works through the shard group box, and the metrics footer's skeleton matches the loaded row height so the card doesn't shift. - Accessibility: the failover tooltip trigger is a keyboard-focusable button, status badges sit in stable `role="status"` live regions, the region flag is decorative (`alt=""`), and the edge dash/spinner animations respect `prefers-reduced-motion` (applied to the pipelines diagram's edges too). - Fallbacks: `AlertError` ("Failed to retrieve cluster topology") when either ha-admin query errors, and a "Cluster topology unavailable" empty state when the topology comes back empty — never a half-rendered diagram. **Changed:** - `InstanceConfiguration` is now topology-source-aware: it branches internally on `useHighAvailability()`, so both surfaces (homepage `TopSection` and the replication page) get the right diagram with no new wiring. The two-pass measured dagre layout moved into a shared `DiagramFlow`; `nodeTypes`/`edgeTypes` are module-level consts. - `getEdgeVisual` + the mid-edge icon chip lifted out of `ReplicationDiagram/Edges.tsx` into `components/ui/ReactFlow/EdgeVisual.tsx` so both diagrams derive edge icon + line style from one state object (no behavior change for the pipelines diagram). The primary card's CPU/Disk/RAM footer is likewise extracted into a shared `ComputeMetricsFooter`. - Fixes a latent relayout loop inherited from the region-box pattern: handing React Flow a freshly created (unmeasured) group node on every layout pass reset `nodesInitialized`, re-triggering the measured pass and `fitView` forever — which made the diagram snap back to center and effectively unpannable. The shared `DiagramFlow` now re-attaches known measurements to group nodes, which also covers the standard diagram's region boxes. - Standard diagram: the API Load Balancer → primary edge is now static — no data flows over it, the line only indicates a relation. - `database/replication` page: the HA early-return empty state is replaced by the diagram under a "High Availability cluster topology" header. Non-HA projects are untouched. **Intentional deviations from the mock** (for design review): 1. **No per-replica regions** — alpha replicas are one-per-cell inside a single region, so the mock's `eu-west-1` / `ap-southeast-1` on sibling replicas would be false. Availability zone per node, region shown once on the primary. 2. **"Primary Database", not "Main Database"** — matches the string both existing diagrams already ship, and the same component now renders both project types. 3. **No collapse chevron on the shard header** — alpha has exactly one shard; collapsing it would hide the whole diagram. The group box still ships; add collapse when `shards.length > 1`. 4. **Failover shown on the shard group, not replica cards** — failover is a cohort property; per-card badging would assert readiness we can't verify without a per-pooler `/status` fanout. 5. **Standard node/edge styling reused** (per review) — neutral primary border + green chip and the default animated edges instead of the mock's green ring and dashed green arrowed edges, keeping the HA and non-HA diagrams visually consistent. **Confirmed against a real local Multigres cluster:** cells are named `cell-1`/`cell-2`/… (not AZ-shaped — the AZ formatter falls back to the raw cell name as designed); `GET /platform/projects/{ref}/databases` returns only the primary row for HA projects; and the `/ha-admin` passthrough returns **each gateway/pooler record once per cell it fans out to** — the topology mapper dedupes by id, but worth confirming with @sbc-bobbie whether the backend should dedupe. **Known alpha limitation:** node health and the "replicating" edge state derive from the pooler's *topology record* (`lifecycleStatus`/`servingStatus`), not a live probe — a pooler that crashes without publishing a terminal state can read as healthy until the topology evicts its record, and a serving replica with paused replay still shows a green edge. This matches the existing replication diagram's semantics (`ACTIVE_HEALTHY` ⇒ animated edge). Live per-pooler signals (WAL receiver state, replay position) exist on `GET /poolers/{cell}/{name}/status` but need a per-pooler fanout — deliberately deferred, noted on `getPoolerStatus`. **Still to confirm** (doesn't block review): whether the `/ha-admin` passthrough is deployed to production or staging-only (if staging-only, this should get a flag before GA). ## To test Tested end-to-end locally against a real Multigres project (standard-project regression pass, HA creation flow, error fallback against real 500s, and full topology + polling + console checks against live multiadmin data): - **HA project homepage**: diagram card shows Multigateway → shard box (`Shard 1`, count badge, `Automatic failover` tooltip) → green-bordered Primary Database (region, AZ, size) + Read Replica cards (AZ), dashed green animated edges to healthy replicas. No flow/map toggle for HA. - **HA project → Database → Replication**: same diagram under a "High Availability cluster topology" header; no Destinations section; the old "Replication unavailable…" state is gone. - **Error path**: if `/ha-admin/v1/*` fails, both surfaces show "Failed to retrieve cluster topology" with Contact support — no partial diagram. - **Standard project regression**: homepage diagram (primary card, flow ⇄ map toggle round-trips), replication page (pipelines diagram + Destinations) all unchanged; zero requests to `/ha-admin/*`. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added High Availability topology diagrams to the Replication page. - Display gateways, primary databases, replicas, shards, statuses, regions, infrastructure details, and compute metrics. - Added observability links and live topology updates with loading, error, and unavailable states. - **Bug Fixes** - Improved handling of incomplete infrastructure identities and unexpected data. - Corrected topology layout, node spacing, and visual edge behavior. - **Accessibility** - Reduced-motion preferences now disable diagram animations and loading effects. - Improved status announcements for assistive technologies. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Alaister Young <10985857+alaister@users.noreply.github.com>
241 lines
8.1 KiB
TypeScript
241 lines
8.1 KiB
TypeScript
import { Edge, ReactFlowProvider } from '@xyflow/react'
|
|
import { useParams } from 'common'
|
|
import { partition } from 'lodash'
|
|
import { Globe2, Loader2, Network } from 'lucide-react'
|
|
import { useEffect, useMemo, useState } from 'react'
|
|
import { Button } from 'ui'
|
|
|
|
import { DiagramFlow } from './DiagramFlow'
|
|
import { SmoothstepEdge } from './Edge'
|
|
import { HaInstanceConfiguration } from './HaInstanceConfiguration'
|
|
import { addRegionNodes, generateNodes } from './InstanceConfiguration.utils'
|
|
import { LoadBalancerNode, PrimaryNode, RegionNode, ReplicaNode } from './InstanceNode'
|
|
import MapView from './MapView'
|
|
import { REPLICA_STATUS } from '@/components/interfaces/Settings/Infrastructure/ReadReplicas/ReadReplicas.constants'
|
|
import { AlertError } from '@/components/ui/AlertError'
|
|
import { useLoadBalancersQuery } from '@/data/read-replicas/load-balancers-query'
|
|
import { useReadReplicasQuery } from '@/data/read-replicas/replicas-query'
|
|
import {
|
|
ReplicaInitializationStatus,
|
|
useReadReplicasStatusesQuery,
|
|
} from '@/data/read-replicas/replicas-status-query'
|
|
import { useHighAvailability } from '@/hooks/misc/useHighAvailability'
|
|
import { useIsFeatureEnabled } from '@/hooks/misc/useIsFeatureEnabled'
|
|
import { useIsAwsCloudProvider, useSelectedProjectQuery } from '@/hooks/misc/useSelectedProject'
|
|
|
|
const nodeTypes = {
|
|
PRIMARY: PrimaryNode,
|
|
READ_REPLICA: ReplicaNode,
|
|
REGION: RegionNode,
|
|
LOAD_BALANCER: LoadBalancerNode,
|
|
}
|
|
|
|
const edgeTypes = {
|
|
smoothstep: SmoothstepEdge,
|
|
}
|
|
|
|
const InstanceConfigurationUI = () => {
|
|
const { ref: projectRef } = useParams()
|
|
const { isPending: isLoadingProject } = useSelectedProjectQuery()
|
|
|
|
const isAws = useIsAwsCloudProvider()
|
|
const { infrastructureReadReplicas } = useIsFeatureEnabled(['infrastructure:read_replicas'])
|
|
|
|
const [view, setView] = useState<'flow' | 'map'>('flow')
|
|
const [refetchInterval, setRefetchInterval] = useState<number | false>(10000)
|
|
|
|
const {
|
|
data: loadBalancers,
|
|
refetch: refetchLoadBalancers,
|
|
isSuccess: isSuccessLoadBalancers,
|
|
} = useLoadBalancersQuery({ projectRef })
|
|
const {
|
|
data,
|
|
error,
|
|
refetch: refetchReplicas,
|
|
isPending: isLoading,
|
|
isError,
|
|
isSuccess: isSuccessReplicas,
|
|
} = useReadReplicasQuery({ projectRef })
|
|
const [[primary], replicas] = useMemo(
|
|
() => partition(data ?? [], (db) => db.identifier === projectRef),
|
|
[data, projectRef]
|
|
)
|
|
const numReplicas = useMemo(() => data?.length ?? 0, [data])
|
|
|
|
const { data: replicasStatuses, isSuccess: isSuccessReplicasStatuses } =
|
|
useReadReplicasStatusesQuery(
|
|
{ projectRef },
|
|
{
|
|
refetchInterval: refetchInterval,
|
|
refetchOnWindowFocus: false,
|
|
}
|
|
)
|
|
|
|
useEffect(() => {
|
|
if (!isSuccessReplicasStatuses) return
|
|
const refetch = async () => {
|
|
const fixedStatues = [
|
|
REPLICA_STATUS.ACTIVE_HEALTHY,
|
|
REPLICA_STATUS.ACTIVE_UNHEALTHY,
|
|
REPLICA_STATUS.INIT_READ_REPLICA_FAILED,
|
|
]
|
|
const replicasInTransition = replicasStatuses.filter((db) => {
|
|
const { status } = db.replicaInitializationStatus || {}
|
|
return (
|
|
!fixedStatues.includes(db.status) || status === ReplicaInitializationStatus.InProgress
|
|
)
|
|
})
|
|
const hasTransientStatus = replicasInTransition.length > 0
|
|
|
|
// If any replica's status has changed, refetch databases
|
|
if (replicasStatuses.length !== numReplicas) {
|
|
await refetchReplicas()
|
|
setTimeout(() => refetchLoadBalancers(), 2000)
|
|
}
|
|
|
|
// If all replicas are active healthy, stop fetching statuses
|
|
if (!hasTransientStatus) {
|
|
setRefetchInterval(false)
|
|
}
|
|
}
|
|
refetch()
|
|
}, [
|
|
numReplicas,
|
|
isSuccessReplicasStatuses,
|
|
refetchLoadBalancers,
|
|
refetchReplicas,
|
|
replicasStatuses,
|
|
])
|
|
|
|
const nodes = useMemo(
|
|
() =>
|
|
isSuccessReplicas && isSuccessLoadBalancers && primary !== undefined
|
|
? generateNodes({
|
|
primary,
|
|
replicas,
|
|
loadBalancers: loadBalancers ?? [],
|
|
})
|
|
: [],
|
|
[isSuccessReplicas, isSuccessLoadBalancers, primary, replicas, loadBalancers]
|
|
)
|
|
|
|
const edges: Edge[] = useMemo(
|
|
() =>
|
|
isSuccessReplicas && isSuccessLoadBalancers
|
|
? [
|
|
...((loadBalancers ?? []).length > 0
|
|
? [
|
|
{
|
|
id: `load-balancer-${primary.identifier}`,
|
|
source: 'load-balancer',
|
|
target: primary.identifier,
|
|
type: 'smoothstep',
|
|
// Static: no data flows between the load balancer and the
|
|
// database — the line only indicates a relation.
|
|
className: 'cursor-default!',
|
|
},
|
|
]
|
|
: []),
|
|
...replicas.map((database) => {
|
|
return {
|
|
id: `${primary.identifier}-${database.identifier}`,
|
|
source: primary.identifier,
|
|
target: database.identifier,
|
|
type: 'smoothstep',
|
|
animated: true,
|
|
className: 'cursor-default!',
|
|
data: {
|
|
status: database.status,
|
|
identifier: database.identifier,
|
|
connectionString: database.connectionString,
|
|
},
|
|
}
|
|
}),
|
|
]
|
|
: [],
|
|
[isSuccessLoadBalancers, isSuccessReplicas, loadBalancers, primary?.identifier, replicas]
|
|
)
|
|
|
|
return (
|
|
<div className="nowheel h-full">
|
|
<div
|
|
className={`h-full w-full relative ${
|
|
isSuccessReplicas && !isLoadingProject ? '' : 'flex items-center justify-center px-28'
|
|
}`}
|
|
>
|
|
{(isLoading || isLoadingProject) && (
|
|
<div role="status">
|
|
<span className="sr-only">Loading infrastructure...</span>
|
|
<Loader2
|
|
aria-hidden="true"
|
|
className="motion-safe:animate-spin text-foreground-light"
|
|
/>
|
|
</div>
|
|
)}
|
|
{isError && <AlertError error={error} subject="Failed to retrieve replicas" />}
|
|
{isSuccessReplicas && !isLoadingProject && (
|
|
<>
|
|
{infrastructureReadReplicas && (
|
|
<div className="z-10 absolute top-4 right-4 flex items-center justify-center gap-x-2">
|
|
{isAws && (
|
|
<div className="flex items-center justify-center">
|
|
<Button
|
|
variant="default"
|
|
icon={<Network size={15} />}
|
|
className={`rounded-r-none transition ${
|
|
view === 'flow' ? 'opacity-100' : 'opacity-50'
|
|
}`}
|
|
onClick={() => setView('flow')}
|
|
/>
|
|
<Button
|
|
variant="default"
|
|
icon={<Globe2 size={15} />}
|
|
className={`rounded-l-none transition ${
|
|
view === 'map' ? 'opacity-100' : 'opacity-50'
|
|
}`}
|
|
onClick={() => setView('map')}
|
|
/>
|
|
</div>
|
|
)}
|
|
</div>
|
|
)}
|
|
{view === 'flow' ? (
|
|
<DiagramFlow
|
|
nodes={nodes}
|
|
edges={edges}
|
|
nodeTypes={nodeTypes}
|
|
edgeTypes={edgeTypes}
|
|
addGroupNodes={addRegionNodes}
|
|
/>
|
|
) : (
|
|
<MapView />
|
|
)}
|
|
</>
|
|
)}
|
|
</div>
|
|
</div>
|
|
)
|
|
}
|
|
|
|
export const InstanceConfiguration = () => {
|
|
const { isHighAvailability, isPending } = useHighAvailability()
|
|
|
|
// Wait for the project record so an HA project never briefly mounts the
|
|
// standard diagram (and fires its queries) before swapping.
|
|
if (isPending) {
|
|
return (
|
|
<div role="status" className="h-full w-full flex items-center justify-center">
|
|
<span className="sr-only">Loading infrastructure...</span>
|
|
<Loader2 aria-hidden="true" className="motion-safe:animate-spin text-foreground-light" />
|
|
</div>
|
|
)
|
|
}
|
|
|
|
return (
|
|
<ReactFlowProvider>
|
|
{isHighAvailability ? <HaInstanceConfiguration /> : <InstanceConfigurationUI />}
|
|
</ReactFlowProvider>
|
|
)
|
|
}
|