Architecture

The monitoring pipeline, data retention and multi-region design.

Pingi is a Laravel application. The web app never performs checks itself; monitoring runs entirely in queued background work.

Monitoring pipeline

Scheduler ─▶ dispatch due monitors ─▶ region queues (checks-fra, checks-lon, …)
          ─▶ check workers ─▶ HTTP request (SSRF-guarded, timed)
          ─▶ store result ─▶ status state machine ─▶ incidents
          ─▶ notification & webhook queues (never inline in the check)
  • SSRF protection — targets are resolved once, every resolved address is rejected if private, loopback, link-local or otherwise reserved, and the request is pinned to the validated IP so DNS rebinding cannot redirect it.
  • Verification — a failure is re-checked from another (healthy) region before a monitor is marked down. With local execution, or a deployment with a single region, the re-check necessarily runs from the same place.

How checks run

  1. Dispatch (pingi:dispatch-checks, every 10 seconds, one server): due monitors are claimed with a guarded UPDATE that moves next_check_at forward and stamps a claim token, so overlapping dispatchers never queue the same monitor twice. Claimed monitors are grouped per region (rotating through each monitor's regions) into batches of up to 25.
  2. Check (RunCheckBatch on the region's queue): a worker only accepts jobs for the region it is configured as (PINGI_WORKER_REGION). Each monitor's URL — and every redirect — passes the target guard and is requested with its connection pinned to the validated IP. A batch's requests run concurrently; response bodies are capped (a response larger than the cap is judged on its status line and headers, never reported as a failure). A region that keeps receiving work but completes none is reported stale by /pipeline-status and skipped for monitors that have another region.
  3. Record and decide: each outcome is stored in check_results, then the monitor state machine updates the monitor. Reaching the failure threshold requests a verification from a different region; only a confirmed failure opens an incident, and only a confirmed recovery resolves it.
  4. Events: MonitorWentDown, MonitorRecovered and MonitorDegraded are dispatched after commit. RouteAlert turns them into queued channel notifications and webhook deliveries on the alerts queue (never inline in the check), subject to PINGI_NOTIFICATION_DELIVERY and a per-channel hourly rate limit.

Keyword checks. A monitor can require that the body contains (or does not contain) a keyword: case-insensitive, within the first 256 KB, always fetched with GET. A mismatch is a failed check (keyword_mismatch).

Maintenance windows. While a window covering a monitor is active, its checks still run and are recorded, but failures never open an incident (so nothing is alerted). Active and upcoming windows are shown on the status pages that list those monitors.

Plans. Limits are enforced when monitors are created, edited or resumed, and pingi:enforce-plans (hourly, and right after a Stripe subscription change) brings running monitors back within a plan that shrank: intervals below the plan minimum are raised and excess monitors are paused.

Target failures are not infrastructure failures. A DNS failure, timeout, connection error or unexpected status is HttpProbe's normal output — it never throws for a target problem, so every monitor always yields a real check_results row that runs through the state machine like any other result. An actual infrastructure failure (a database error, an exception in the state machine, a crashed worker) is a PHP exception that only happens after the real check already ran; it fails the job (RunCheckBatch is never retried — a missed check is superseded by the next scheduled one, not replayed) without touching the affected monitor's status or failure streak. An infrastructure problem is visible in queue:failed, never as a false "down" on a customer's monitor.

SSL certificates

Every check_interval_hours (6 by default) each active HTTPS monitor's certificate is inspected (pingi:dispatch-ssl-checks, hourly). The host goes through the same target guard as uptime checks and the TLS connection is made to the validated IP, with the hostname sent only as SNI. Pingi then judges the certificate itself: validity dates, whether the chain validates against the trusted CA bundle, and whether its names cover the host (wildcards cover exactly one label). The result is stored in ssl_certificates with where the check ran, and SslInvalid, SslRecovered and SslExpiring fire on transitions. If the worker has no CA bundle, chain validity is recorded as unknown rather than invalid.

Local development

With PINGI_CHECK_EXECUTION=local (the default) every check runs on the local worker and is recorded against the Local development pseudo-region — verifications included. Nothing is attributed to Frankfurt, London, New York or Singapore unless a worker is actually running there.

php artisan schedule:work
php artisan queue:work --queue=checks-local,maintenance,alerts
php artisan pingi:check {monitor-identifier}       # run one uptime check now
php artisan pingi:ssl-check {monitor-identifier}   # inspect one certificate now

Data retention

Data Kept for
Raw check results 14 days
Hourly rollups Per plan retention
Daily rollups Long-term
Incidents Long-term

All timestamps are stored in UTC and statistics periods are UTC hours and UTC days.

Rollups (pingi:rollup hourly at :05, pingi:rollup daily at 00:20, on the maintenance queue) recompute the latest complete periods from raw results. Re-running a period is safe and absorbs late results. Verification checks are excluded (they are triggered by failures and would bias uptime); up and degraded both count as reachable; response-time figures cover reachable checks only; percentiles are exact nearest-rank values, and daily rollups are computed from raw data rather than averaged from hourlies. Downtime is the overlap of the period with the monitor's incidents.

Reading statistics: 24-hour and 7-day figures for a monitor are computed exactly from raw results. 30- and 90-day figures combine daily rollups, today's hourly rollups and the live hour; their percentiles are a check-weighted average and are flagged as approximate. Downtime always comes from incidents. Lists of monitors read uptime from rollups in one query and lag by up to an hour.

Retention (pingi:prune, daily) deletes in small batches with pauses: raw results after 14 days, hourly rollups after the plan's retention, audit logs after the configured period. Daily rollups and incidents are kept.

Tenancy and authorization

Every customer record belongs to an account (workspace). People join accounts through memberships, each with a role:

Role Can
Owner Everything, including billing, deleting the account and managing other owners
Admin Monitors, incidents, status pages, notifications, webhooks, API tokens, team, settings
Member Monitors, incidents, status pages, notifications, their own API tokens
Viewer Read-only

Policies check permissions, never roles directly, and a permission only counts inside the record's own account. Pingi staff (platform admins) use the separate admin panel and have no implicit access inside customer accounts.

API tokens act as their creator within one account. A token stops working the moment it is revoked, expires, or its creator leaves the account, and a token can never do more than its creator's current role allows.

Partitioning plan for check results

check_results is designed to move to MySQL daily RANGE partitioning on checked_at once volume warrants it:

  1. Change the primary key to (id, checked_at) — MySQL requires the partitioning column in every unique key.
  2. PARTITION BY RANGE COLUMNS (checked_at) with one partition per day, plus a pmax catch-all.
  3. A daily job creates tomorrow's partition and drops partitions older than the 14-day raw retention — an instant metadata operation instead of a DELETE over tens of millions of rows.

The table already has no foreign keys (InnoDB cannot partition tables that have them), so this change needs no application code changes.

Platform-admin access to a customer workspace, when it is built, will be a separate admin-panel operation that is explicitly granted and audited — never a bypass in the customer-facing policies.

Data integrity

Every tenant table has account_id with a foreign key to accounts that cascades on delete, and an index led by account_id. Optional references (creator, acknowledging user, region, a monitor's current incident) are set to null when the referenced row goes away, and a plan cannot be deleted while a subscription uses it.

Three tables are deliberately different: check_results, monitor_stats_hourly and monitor_stats_daily have no foreign keys. They are the highest-volume tables, they must stay partitionable (InnoDB cannot partition tables with foreign keys), and inserts on them should not pay for FK checks. Columns elsewhere that point into check_results (incidents.recovery_check_result_id, incident_events.check_result_id) are plain ids for the same reason. The audit log also stores actor ids without foreign keys so the trail outlives deleted users and tokens.

Consequently rows in those tables are removed by background jobs, never by cascades inside a web request: deleting a monitor soft-deletes it (freeing its plan slot immediately), and a queued purge later removes the monitor and its history in batches (pingi:purge-monitors, daily at 04:00, after a configurable grace period). The purge deletes the monitor's raw results and rollups by monitor_id in bounded batches, then removes the monitor row itself, letting foreign keys clean up incidents, certificates and links. The row goes last, so an interrupted or budget-limited run simply continues on the next run. Each purge is recorded as a system audit entry.

Workspace (account) deletion follows the same rule. The owner-only, password-confirmed deletion request itself only ever runs a handful of single, bounded, indexed statements keyed by account_id (deactivate and soft-delete monitors, disable channels/webhooks, revoke API tokens, delete memberships and invitations, soft-delete the account) — never a per-row loop and never a database cascade — so it's safe to run inside the HTTP request. The account row itself is only ever soft-deleted, never hard-deleted: every dependent table's account_id foreign key cascades on delete, so a hard delete would trigger the same uncontrolled cascade this pattern exists to avoid. A daily queued job (pingi:purge-accounts) removes the rest of a soft-deleted workspace's data after a grace period — its monitors (and their check history) are already handled by the existing pingi:purge-monitors pipeline, since deletion soft-deletes them immediately; the account purger's own scope is the smaller, account-level tables (notification channels, webhooks, status pages and their branding files, API tokens, subscriptions, usage records). See RequestAccountDeletion and AccountPurger.

Provisional behaviour (open decisions)

These behave as described today but are not final and may change once decided:

  • Blocked targets. A monitor whose host resolves to a blocked address records a failed uptime check (blocked_target), which feeds the normal incident flow. For SSL, a blocked or unreachable host leaves the last known certificate status unchanged and raises no SSL event. Whether blocked targets should instead raise a separate "misconfigured" alert is undecided.
  • Verification policy. Failures and recoveries are confirmed by one check from another region. If no other region is available (single-region deployments) the same region confirms. Local development always verifies locally, labelled as such.
  • Regional infrastructure. Results are attributed to Frankfurt, London, New York or Singapore only when a worker actually configured as that region runs the check. Deploying those workers (and Redis/Horizon) is outstanding.
  • Plans and pricing. All plan limits, prices and trial terms in config/pingi.php are development placeholders.
  • Hourly statistics retention. Hourly rollups are kept for the plan's retention_days. With the placeholder plans that reaches 730 days, which at 10,000 monitors is roughly 175 million hourly rows. An optional technical cap (PINGI_HOURLY_STATS_MAX_DAYS) exists but is unset; how long hourly detail is kept (versus daily only) is undecided.
  • Deleted-monitor grace period. Deleted monitors are kept, hidden, for PINGI_DELETED_MONITOR_GRACE_DAYS (7 by default) before their history is purged. The final grace period — and whether users get an "undo" within it — is undecided. There is no restore feature yet; one must refuse monitors a purge has already started on.
  • Deleted-workspace grace period. The same applies to a deleted account: kept (soft-deleted, immediately inaccessible) for PINGI_DELETED_ACCOUNT_GRACE_DAYS (7 by default) before its data is purged, with the same undecided final grace period and no restore feature.