Detect
Immediate rule packs match supported failure evidence as it arrives. Learned detectors establish service cadence and baselines from history.
Epok watches logs, metrics, traces, infrastructure, and RUM across the systems involved, detects emerging failures, groups the cascade, and separates the trigger from the sustaining cause. Every committed conclusion links to source evidence; when the evidence cannot support one, Epok abstains.
self-service or governed evaluation · shadow mode · no cutover · review the governed plan →
the live demo — this exact screen, no signup →Keep Datadog, New Relic, ClickHouse, or your current search layer for storage and queries. Epok runs the incident loop around that telemetry without becoming part of your production request path.
Immediate rule packs match supported failure evidence as it arrives. Learned detectors establish service cadence and baselines from history.
Related alerts, services, traces, infrastructure signals, and—when present—deploy and user-impact context align into one incident timeline.
Epok separates the initiating trigger from the condition sustaining the failure, then cites the lines, spans, or metrics behind a committed cause.
Slack, PagerDuty, email, or a webhook receives one incident with evidence and, where supported, a reversible response recommendation. Responders execute.
One real incident, worked end to end — a retry storm holding a service down 14 minutes after its trigger cleared. The same incident is open in the live demo.
The request queue is accelerating toward saturation. Epok forecasts the breach, with live ETA and confidence shown, 12 minutes before anything pages.
The counterintuitive right move: shed the retries. And the guard that matters — rollback will NOT recover this. We recommend; you execute.
The verdict separates the trigger (a 40s blip, cleared) from the sustaining cause (the retry policy) — cited, with the last deploy exonerated.
AI assistance is only useful if responders can distinguish evidence from fluency. Epok can abstain and show why the available signals do not support one cause. Use the trial to inspect both cases: committed verdicts and honest "not enough evidence" outcomes.
The input requirement matters as much as the feature name. Here is what each result needs before an SRE should expect it.
Needs: Supported logs, events, or metrics with service and environment context.
Produces: Can match the first supported failure event; no baseline-training window.
Needs: Enough volume or cadence history to establish what is normal for the service.
Produces: Baselines sharpen over the first week; learned detection stays gated while the required history is insufficient.
Needs: Stable service, environment, and timestamps; trace and deploy context improve the links.
Produces: Related alerts and downstream symptoms collapse into one incident chain.
Needs: Enough corroborating logs, spans, metrics, or infrastructure evidence in the incident window.
Produces: Commits with citations when supported; ranks candidates and abstains when it is not.
Needs: A customer identifier in telemetry and, optionally, an account or tier roster.
Produces: Shows affected accounts and tiers beside the technical incident evidence.
Needs: Browser replay instrumentation plus shared trace context between browser and backend.
Produces: Opens the relevant replay at the failure moment from the trace that broke.
The public demo uses synthetic, production-shaped data. Use these surfaces to inspect the actual output, supported coverage, operational limits, and data path before connecting your own telemetry.
Inspect the correlated timeline, cited verdict, candidates, and abstention path.
See what can match immediately, what learns, and the evidence each detector consumes.
Review published rate, concurrency, retention, and safety-ceiling behavior.
Check OTLP, open-shipper, CloudWatch, Loki, Elasticsearch bulk, syslog, and JSON options.
Review isolation, AI handling, subprocessors, retention, identity, and current gaps.
The controlled-evaluation plan is an option for environments that require it—not the minimum process for trying the product. In both paths, current paging remains authoritative and Epok makes no autonomous production change.
An on-call engineer can open a 14-day workspace with any work or personal Google account and no card. Keep every existing destination, dashboard, monitor, and page unchanged.
A platform or security team can approve a representative production boundary, selected fields, expected volume, an ingest-scoped key, and an independent queue, retry policy, and ceiling.
Classify correct, incorrect, abstained, and missed outcomes. Compare alert fanout, time to verified cause, coverage, and responder effort; then expand, coexist, or stop.
When an alert fires, Epok opens one canvas: the drafted root cause, what changed, the cascade timeline, the blast radius, and who's affected — every claim clickable back to the line, span, or metric behind it.

When browser and backend instrumentation preserve shared trace context, a failing trace can open the associated replay at the failure moment. You inspect the user experience and backend path without cross-referencing two timestamps.
Add a roster and emit a customer identifier. Relevant incident windows can then show affected accounts and tiers beside the technical evidence.
Type "why is checkout slow in the last hour." AI translates it to a query, a 42-test validator forces time + limit, then we run it. Query is ground truth; the explanation sits beside it.
When the incident closes, a draft appears — triggering signal, cited evidence chain, matched playbook, and customer-impact rollup, pre-assembled. You edit; you don't author from a blank page.
Bulk-import from Confluence, Notion, GitHub, or Markdown. A citation engine surfaces the best-matching runbook — the specific steps land inside Slack, PagerDuty, and the deep RCA, not a link to a wiki.
Surfaces messages that never appeared in your recent history — the first sign of a fresh failure.
Catches a service that stops logging when it normally logs steadily. Failure with no error — just absence.
Flags volume that jumps, falls, or flatlines against each service's daily and weekly normal.
Groups errors that share a shape, so dozens of variants land as a single alert.
Surfaces crash loops, out-of-memory kills, and unschedulable workloads straight from their logs.
Connects upstream failures, retries, and circuit-breaker trips into the cascade they cause.
The pager is rationed by design. Repeats collapse, cascades arrive as one chain, and severity rides explicit thresholds the product enforces.
Alerts with the same semantic fingerprint collapse into one alert with a fire count.
Correlated signals can arrive as one incident chain: db silent → API refused → frontend 502s.
Repeat fires de-escalate on a widening window. New shapes still get full severity.
Critical / Warning / Info on explicit thresholds the product enforces.
A 5-service app generates a continuous synthetic log stream into a public Epok tenant. Anomaly detection, RCA, and clustering run on it live — Epok working on real-shape data, not a marketing video.
One included-volume meter across logs, metrics, traces, infrastructure, RUM, and replay. Paid-plan overage is $0.20/GB and shown daily; there is no per-host, per-query, or cardinality line.
No. An on-call engineer can run the self-service trial, while a platform or security team can define a governed evaluation. Choose a representative service group, ownership domain, environment, or critical user journey; your current alerts and incident process remain authoritative throughout.
Ingest pauses, your data stays readable, and you add a card when you're ready. Nothing auto-charges.
Self-service signup accepts any personal or work Google account; it is not restricted to Google Workspace. Microsoft login, SAML enterprise SSO, and SCIM are not available today. If your organization cannot use Google, inspect the live demo without signup or request a governed evaluation.
Your plan price and included volume are fixed. Paid-plan overage is $0.20/GB and is shown in your dashboard daily; a documented safety ceiling limits runaway ingest. The trial cannot generate overage or a charge.
No. Log as many unique fields as you want; there is no per-series tax.
No. Detectors run automatically, and you can ask in plain English.
Send logs, metrics, traces, and infrastructure data with OpenTelemetry or an open shipper—there is no proprietary Epok agent for server-side telemetry. Browser RUM and replay use web instrumentation such as OpenTelemetry Web and rrweb.
Yes. Mirror a representative production boundary to Epok through a second OpenTelemetry or open-shipper path, and keep your current alerts unchanged. Compare both systems on the same incident windows before discussing any expansion or cutover.