AI systems / Self-initiated product
From an alert
to an answer.
I turned a manual investigation workflow into an AI triage system: evidence-driven diagnosis, explicit uncertainty, and a clear route back to a human.
Explore the decisions
Manual investigation to automated diagnosis per alert.
Coverage across eight journeys and nine markets.
A separate daily anomaly workflow, built on the same pattern.
One alert, one inspectable verdict
Select a stage to see which part of the reply it produced.PH market · hourly check · previous hour below the lower threshold
- 1Verdict
- Transient, recovered (system) · PH market · 1:00–2:00
Payment routing and failover failure. The primary provider went dark for the hour and failover providers could not cope; recovered by 2:00. - 2Bottom line
- Success per click 13.2% (the alert’s 12.3% is the funnel rate) against a ~55% same-hour, seven-day baseline, more than 10 standard errors below.
- What broke
- Submit to success 27.8% (baseline ~78%), the dominant leak
- Clicks per user 2.74 (baseline ~1.5): heavy retrying
- Primary provider: 1,819 orders to 0; three failover providers at 13–30% success
- 3Scope
- This market only. Sibling markets normal in the same hour (47.5%, 40.8%). Not platform-wide.
- Confidence
- High · posted automatically in the alert thread
- 4Owner · action
- Payments and platform routing. Post-incident ticket: why the primary provider stopped, and why failover ran at 13–30%.
- 1The conclusion comes first, as a named incident class from the playbook.
- 2Evidence is measured against a baseline, not just restated from the alert.
- 3Scope rules out a platform-wide failure before anyone is paged.
- 4The reply ends with who acts next and what to ask.
01 / The process
The signal arrived.
The answer took work.
I joined an eight-person UX diagnostics team monitoring a digital entertainment platform across nine countries. When a conversion metric dropped, an analyst had to query data, check operational context, form a hypothesis, and write a verdict.
Each investigation took 30–50 minutes. Quality depended on who was available. Across time zones, an overnight alert could sit unread for hours. I saw an opportunity to redesign the work I had been hired to do.
I initiated and built
The triage methodology, automation architecture, runbooks, confidence gates, daily anomaly workflow, and recipe audit tooling. The concept began without a formal mandate.
I iterated with the team
My section lead's feedback shaped the verdict format. The AGM challenged runbook structure and cost. A fellow analyst owned alert configuration; backend engineers built a production endpoint.
The product was a decision
The output had to tell someone what happened, what evidence supported it, and what needed attention. A generated paragraph alone would not close the diagnosis gap.
I began by standardizing how to investigate. Then I gave that method to an agent.
02 / The process
Design the boundaries around the model.
The first version was a playbook: five standard queries and a matrix for false alarms, demand shifts, system failures, and recovered incidents. I tested the method on four real cases before wiring it into the pipeline.
The architecture connected a PostHog alert to a Cloudflare mailbox, a poller, a Claude agent, and a verdict in Slack. The important design decisions were what the agent could access, when it could act, and how failure would surface.
The prototype, from signal to delivery
Architecture reconstruction. Analytics access is read-only; the agent’s write action is a reply in the original Slack thread.
Transient, recovered (system) · confidence High
The first live verdict, shown in full above.
Needs human triage
Illustrative: deposit success fell in one market, but provider data for the hour is incomplete. Findings so far are attached; an analyst confirms the cause.
[needs manual review] Deposit success rate · PH market · hourly
Automated triage could not produce a verdict after retries. Please triage this alert by hand.
The local poller is a real limitation: runtime availability affects coverage. Production infrastructure later belonged to backend engineering.
Follow a defined query sequence and decision matrix. This makes diagnoses traceable and repeatable, while accepting that novel failure modes may require changes to the playbook.
Post the verdict under the original alert. Readers get signal and diagnosis together, without another channel or a duplicated notification.
Why not let the model investigate freely?
Consistency was part of the original problem. A written playbook gives humans and the agent a shared, auditable method.
Why not send everything to a person?
It would preserve the bottleneck. Confidence-based routing automates the clearer cases while leaving ambiguous cases open for judgment.
The verdict format was an interaction design problem.
As the system met real use, I revised its output with my section lead's feedback. A teammate needed a usable conclusion with evidence, not a transcript of everything the agent had done. That is why the delivery format and location mattered as much as the investigation itself.
03 / The process
The first live verdict matched an independent triage.
On 22 June at 2:46 AM, a deposit alert fired in a Philippine market: 12.3% success against a 44% threshold. The pipeline had been rejecting live fires since 17 June, so no verdict arrived overnight. After the webhook fix that morning, the alert was replayed and the verdict posted in its thread at 8:50 AM.
The AGM had already triaged the same incident by hand. The conclusions matched: a provider outage, a failover that could not cope, one market affected, already recovered. The fix was proven on three real alerts from that night, an early proof rather than an accuracy benchmark.
PostHog alert · 22 June 2026
The alert fires
Success 12.3% against a 44% threshold. Live fires are still being rejected.
Manual triage by the AGM
A person investigates
An independent conclusion, written without the agent.
AI verdict in the alert thread
The same conclusion
Replayed after the fix, with baseline, scope, owner, and action attached.
Five days of silence changed the design.
The fix addressed the immediate bug: the webhook body omitted a field the Worker required. The lasting change was to treat delivery and failure visibility as part of the experience. Queued alerts, stale work, and failed investigations all need to be observable. I later used the incident in a company-wide AI sharing session.
Running is not the same as working
17–22 June. Every component was up; every live fire was rejected.
- Alert firesPostHog sends the webhook
- Worker rejects it400: the body omits the required alert ID
- Mailbox stays emptyThe poller runs with nothing to pick up
- No verdictFive days, no error in Slack
Staleness watchdog · direct message
A1 triage poller looks down.
1 alert waiting more than 15 min with no verdict.
Transport check · pipeline channel
1 of 6 PostHog fires never reached the triage inbox in the last hour.
PostHog’s own activity log shows these alerts firing, but the inbox has no record of them.
Message text from 22 June and 2 September 2026, condensed. Internal service names replaced.
04 / The process
A repeatable method becomes team infrastructure.
I extended the same playbook-and-agent pattern to the morning anomaly check. It compared metrics, investigated changes, and prepared a report for review. The routine went from about 50 minutes to 10 minutes.
For alert triage, reusable runbooks and recipes expanded coverage across eight user journeys. At that scale, quality needed its own tooling.
From four recipes to roughly 940.
First live triage.
Deposit alerts, three sites.
Five journeys baselined.
A repeatable pattern.
Coverage across
20 site groups.
Eight journeys.
833 configured alerts.
Recipes and configured alerts are different units. These figures describe the coverage fleet, not a count of successful diagnoses.
Scaling exposed a second design problem.
A batch update could propagate a bad rule as easily as a good one. I built an audit catalogue with 46 checks across four severity tiers, covering privacy, routing, evidence, confidence, and consistency. Forty-one runbook documents made the investigation method available beyond a single prompt.
I also created training materials and shared the work company-wide. Backend engineering built a production alert-triage endpoint after the prototype. My contribution was the method and prototype architecture; production infrastructure remained engineering's responsibility.
What the audit caught in 482 rules
Audit runs from 27 August to 13 September 2026. The catalogue grew from 40 checks to 46 as new defect families appeared. About 31 distinct ideas had been hand-copied into 482 rules.
| Defect family | Found | Why it matters |
|---|---|---|
| Verdicts written on empty evidence | 30 recipes | An answer with nothing behind it |
| Rules that break silently on rename | 79 rules | Coverage disappears without an error |
| Verdict words the schema rejects | 21 rules | The reply fails validation |
| Prompt orderings for one method | 54 variants | The same method behaves inconsistently |
| Withdraw recipes that query nothing | 15 of 19 | 12 of them auto-posted at High confidence |
Each audited rule receives a severity, findings, and an audit date. The audit never edits the triage rules themselves.
What I would change
Build the audit tooling before the rapid expansion. Some recipes shipped with query or attribution errors that required remediation.
What still limits the system
The poller depends on my laptop. Some journeys have incomplete evidence coverage, transport can fail, and costs reached about $70/day at roughly 500 alerts before model optimization.
05 / The process
The next question:
when should an agent act?
Building this system led me to design more agents and bots, and to spend more time on the decisions that control their behavior. I am researching and experimenting with Jev as part of that work.
The question I want to investigate is specific: does the available evidence support a verdict, justify another query, or require a person? I want those transitions to be as inspectable as the final output.
Jev · Exploring the decision layer
Personal research · ongoingMy proposed evaluation is to replay reviewed incidents through the existing approach and an experimental decision step, including clear failures, recovered incidents, and missing evidence.
- Supported conclusions: does the evidence actually justify the verdict?
- Appropriate escalation: does the agent recognize when it cannot conclude?
- Useful investigation: does another query reduce uncertainty?
- Practical cost: what happens to time, cost, and unnecessary review?
This is an experiment direction and evaluation plan. I am not claiming measured Jev improvements or a deployed replacement for the existing gates.
The model was one part of the work. Making its decisions usable became the product.