Live Platform Health

Status of every system that touches a request.

Real-time uptime, SLO compliance, and incident history for every dependency in the FunctionFly stack. Includes Postgres, Redis, the orchestrator API, the StateFabric state backend, the runtime fleet, the compute node fleet, Cloudflare R2, and the AI inference service.

What we monitor

Every dependency that affects user-visible behavior. No synthetic "all green" checks — the platform probes each subsystem and reports what it actually sees.

Orchestrator API

Per-route latency, error rates, rate-limit denials, and tenant-rate-limit denials.

Postgres + Redis

Connection-pool saturation, slow-query alerts, cache hit ratio, replication lag (when applicable), and vacuum backlog.

Runtimes

The full runtime fleet — sandbox subprocess pools, watchdog kills, and concurrency-cap usage.

Compute fleet

Active pod count, pod ready time, ingest rate, revocation freshness, and replay verification rate.

Cloud Agents

Wake pipeline, inbox backlog, supervisor concurrency, cross-org pairing denials, and snapshot encryption writes.

Managed AI

Multi-provider ladder, rung failure rate, cache hit ratio, quota denials, BYOK denials, and execution receipts.

Production-launch readiness

What's in general availability today, what's running in shadow mode, and what's queued for staging.

Live /health endpoint with subsystem detaillive: Live
Prometheus /metrics exportlive: Live
Grafana dashboard provisioninglive: Live
Alert ruleslive: Live
Public status pagelive: Live

Status without the marketing

No "all systems operational" rubber stamps. Real probes, real numbers, real incidents when they happen.

Docs reference →