Observability
Verifier emits Prometheus metrics; this page covers shipping them into a
working monitoring stack: scraping, the reference Grafana dashboard, the
alert rules, and a runbook for every alert. The reference configs live in
the tree under deploy/observability and are validated by
make observability-lint in CI.
Scraping the verifier
Metrics are served at GET /metrics on the HTTP monitoring API, which is
disabled until the verifier is started with --http-addr:
lota-verifier ... --http-addr 127.0.0.1:8080
The endpoint sits on the reader tier of the monitoring API:
On a loopback bind no key is required; a co-located Prometheus can scrape
127.0.0.1:8080directly.On a non-loopback bind the API fails closed: the verifier refuses to start unless a reader or admin key is configured, and every scrape must present the reader key as a Bearer token (
LOTA_READER_API_KEY; see the deployment page for the key tiers).The monitoring API speaks plain HTTP. Scrape over loopback or a trusted network, or terminate TLS in front of it; never expose the port to an untrusted network.
deploy/observability/prometheus.yml is a complete reference
configuration: it scrapes /metrics with the reader key read from a
root-only credentials file and loads the alert rules from the alerts
directory next to it. On Kubernetes, the Helm chart exposes the same API
through the httpApi values (off by default).
Metrics reference
Every labeled series is zero-filled with its known label values, so alert expressions match from the first scrape and never miss the first event after a verifier restart.
Metric |
Type |
Meaning |
|---|---|---|
|
counter |
Attestation attempts. |
|
counter |
Attestations that verified successfully. |
|
counter |
Failed attestations (parse errors + rejections). |
|
counter |
Rejections by reason: |
|
counter |
Boot-baseline re-anchors by outcome:
|
|
counter |
Protocol-level errors on the attestation port. |
|
histogram |
Attestation verification latency. |
|
gauge |
Outstanding attestation challenges. |
|
gauge |
Registered attestation clients. |
|
gauge |
Currently revoked client AIKs. |
|
gauge |
Currently banned hardware IDs. |
|
gauge |
Consumed nonces retained for replay protection. |
|
gauge |
Loaded PCR policies. |
|
gauge |
Verifier uptime. |
Grafana dashboard
deploy/observability/grafana/lota-verifier-dashboard.json is a reference dashboard importable into a stock Grafana: Dashboards -> New -> Import, upload the JSON, and select the Prometheus datasource that scrapes the verifier. It shows the fleet at a glance (success ratio, clients, revocations, bans, challenge backlog, uptime) over time series for attestation rates, rejections stacked by reason, verification-latency quantiles, re-anchor outcomes, protocol errors, and the replay history. Panel thresholds mirror the alert rules, so a red stat on the dashboard and a firing alert say the same thing.
Alerts and runbooks
deploy/observability/alerts/lota-verifier-alerts.yaml ships 11 rules. Severity encodes the escalation tier:
warning = L2: operator triage. The fleet's integrity verdicts are still sound; investigate during working hours.
critical = L3: integrity or service impact. Act immediately: either a device failed integrity or the verifier cannot render verdicts.
The runbooks below assume shell access to the verifier host and the
monitoring API on 127.0.0.1:8080; on an authenticated deployment add
-H "Authorization: Bearer $(cat /path/to/key)" (reader key for reads,
admin key for mutations).
LotaVerifierDown (critical)
Prometheus cannot scrape the verifier. While it is down no verdicts are rendered and relying parties see stale or missing attestation state.
Check the service and its last words:
systemctl status lota-verifier/journalctl -u lota-verifier -n 100(or the pod logs on Kubernetes).A crash loop after a configuration change: roll the change back first, ask questions later.
If the process is healthy, the path Prometheus uses is broken: firewall, reader key rotated without updating the credentials file, or the monitoring bind address changed.
Escalate to L3 if the verifier does not come back within one restart: collect the journal and the database state before further restarts.
LotaAttestationStall (warning)
Registered clients exist but no attestation arrived for 30 minutes. The fleet went silent: agents down, network path broken, or the attestation port unreachable. A silent fleet is indistinguishable from a compromised one, so do not let this linger.
Confirm the attestation listener is up and reachable from a fleet subnet:
curl -sk https://VERIFIER:8443/should at least open a TLS connection.Pick a known host and check its attest unit:
systemctl status lota-attest.serviceon the device.Check
GET /api/v1/attestations?limit=10for the last records and their timestamps to bound when the fleet went quiet.Escalate if the outage window exceeds the fleet's attestation interval several times over - decide whether stale sessions must be invalidated.
LotaIntegrityLoss (critical)
Device presented PCR or runtime state that does not match its policy or
recorded baseline (pcr_fail or integrity_mismatch). This is the
product's primary tamper signal and the alert the forced drill below
exercises.
Identify the device:
GET /api/v1/attestations?limit=50and filter the failed records; each carries the client ID and failure detail. A PCR failure names every register that mismatched, in index order, so one record holds the whole delta and repeated rounds read identically.Read that client's state:
GET /api/v1/clients/{id}.Decide benign versus hostile. How many registers moved is the first clue: several at once is a reboot into a different boot chain, one alone is a narrower change. Benign causes leave a paper trail: a kernel or bootloader update changes PCRs after a reboot (the re-anchor flow handles it), an agent binary update changes the agent hash pin, an operator changed the kernel command line. No matching change record = treat as hostile.
Hostile: revoke the AIK (
POST /api/v1/clients/{id}/revokewith an actor and a reason) and, on the gaming deployment, consider a hardware ban (POST /api/v1/bans). Preserve the device for forensics; do not fix it back into the fleet.Benign: follow the re-anchor procedure in the bring-up documentation and clear the finding through the review queue, not by loosening policy.
LotaRevokedDeviceActivity (warning)
Revoked or banned device keeps attesting. Expected for a short window right after a revocation (the agent retries on its interval); persistent activity means the remediation never reached the host or someone is replaying its identity.
GET /api/v1/revocations/GET /api/v1/bans- confirm the device is intentionally listed and by whom (GET /api/v1/audit).If the device should have been reprovisioned, verify the host was actually re-enrolled (a re-enrollment issues a fresh AIK certificate; the old identity keeps knocking until the agent state is cleared).
Sustained activity from a banned gaming device is an expected nuisance; sustained activity from a revoked enterprise host is a process failure - chase the owner.
LotaReplayBurst (warning)
More than a handful of nonce failures in 5 minutes. Isolated failures are benign retries after timeouts; a burst is either a replay attempt or a client whose clock is far enough off to keep missing the nonce window.
GET /api/v1/attestations- if the failures concentrate on one client, check that host's clock and network latency first.Failures spread across many clients usually mean a verifier-side problem (nonce store latency); correlate with
lota_verification_duration_secondsand the database.A concentrated burst from one source that is not a registered client is an attack signature: capture the source address from the verifier log and treat it as hostile traffic.
LotaReanchorEscalation (warning)
Device's boot baseline changed in a way the self-service re-anchor
policy would not accept automatically. The re-anchor was refused and the
device keeps failing attestation until an operator acts; this is not the
post-fact review queue (that queue holds lfa re-anchors, which apply
automatically).
Identify the device from the verifier's security log (the escalation is logged with the client ID and the refusal reason).
Corroborate the change: a fleet-wide firmware or kernel rollout produces a wave of escalations with the same measurement delta; a single device with a unique delta deserves the LotaIntegrityLoss treatment.
For a legitimate platform change, force the re-baseline with
POST /api/v1/clients/{clientID}/reanchor(admin key, actor recorded in the audit log); the next attestation re-establishes trust. For anything suspect, revoke or ban instead.
LotaBaselineStoreErrors (critical)
Attestations are being rejected because the verifier cannot read or write its baseline store. The verifier fails closed, so healthy devices are refused while the backend is broken - this is a service outage, not a security event.
Check the database: connectivity, disk space, and (Postgres) whether the instance accepts writes.
The verifier log names the failing operation; a migration mismatch after a partial upgrade also lands here.
Once the store recovers, rejected devices re-attest on their next interval with no operator action.
LotaAttestationFailureRatioHigh (warning)
More than 10% of all attestations failing for 10 minutes. One tampered device cannot move this needle - fleet-wide ratios point at policy, certificate, or infrastructure problems.
The rejections-by-reason dashboard panel says which failure dominates:
sig_failafter a CA change means certificates (did the AIK CA rotate without redistributing?),pcr_failacross the fleet means a policy pushed with wrong pins,baseline_errormeans the store.Correlate the onset with the last policy or infrastructure change and roll it back.
LotaVerifyLatencyP99High (warning)
Verification p99 above one second, sustained. Agents tolerate this but their timeouts are finite; at several seconds the failure ratio starts climbing as a side effect.
Database pressure is the usual cause: check the backend's latency and the verifier host's CPU.
If load grew with the fleet, scale per the deployment page (Verifier deployment topologies) rather than raising timeouts.
LotaPendingChallengesHigh (warning)
More than 1000 challenges issued but never answered. Either a fleet segment stalls mid-handshake (network drops the second leg) or an unauthenticated source is farming challenges.
Compare with
lota_registered_clients: a backlog near the fleet size is a fleet-wide network problem; a backlog far above it is a flood.Challenges expire on their own; the gauge should drain once the cause stops. If it climbs unbounded, capture the source addresses from the verifier log and filter upstream.
LotaConnectionErrorBurst (warning)
Sustained protocol-level errors on the attestation port: a scanner, a protocol-version mismatch after a partial upgrade, or a load balancer health-checking the TLS port with plain HTTP.
The verifier log records the offending source and the parse error.
If the onset matches an agent or verifier rollout, suspect a version mismatch and finish or roll back the rollout.
Forced integrity-loss drill
Run this drill after standing up the stack (and periodically after) to prove the pipeline end to end: a tampered measurement must light the dashboard and page within one scrape interval. It satisfies the "alerts fire on a forced integrity-loss event" acceptance test.
The safe variant injects the mismatch on the verifier side, so no device
is modified: edit a staging verifier's policy to pin an agent hash
that cannot match (for example, flip one hex digit of the pinned value in
agent_hashes), reload, and let one healthy device attest:
The attestation is rejected;
lota_rejections_totalincrements with the corresponding reason.Within one scrape interval the Rejections by reason panel shows the series step, and LotaIntegrityLoss enters pending/firing.
Confirm the page arrives through the deployed notification channel, then restore the correct policy and verify the device's next attestation succeeds.
The full-fidelity variant tampers with a scrap test device instead (edit its kernel command line and reboot, or strip the agent binary's integrity metadata): the device itself fails verification, exercising the same path an attacker would trip. Never run either variant against production policy state without a change record: the drill is indistinguishable from a real event by design, and the on-call rotation should treat it as one until the change record says otherwise.