Product Engineering · Observability for AI Applications
Run Explorer
See Every Step. Find the Failure.
I am building a dashboard for inspecting application runs. It will show what ran, what happened inside it, and where it failed, and an AI failure brief will explain the recorded evidence without claiming more than the record shows.
Status: Starting
The scope, architecture, and build plan are written, and I am starting the build. This page describes what I intend to build and how I will measure it. It will be updated with results, code, and a demo as the work happens. Everything will use synthetic data.
The Problem
When an AI application misbehaves, the first questions are simple. What ran? What happened inside it? Where did it fail? What does the recorded evidence actually tell us?
Answering them usually means digging through logs. A model can help explain a failure, but only if it is held to the record: an explanation that sounds plausible and is not supported by the steps is worse than none.
Run Explorer is my attempt at a small, honest answer to that. Its first customer is me, inspecting the runs of Ask TanStack Query. The same event format could later describe Intake Desk.
How It Will Work
A run is one execution of an application workflow, a step is a named piece of it, and an event is a recorded fact about a step:
The dashboard reads recorded work. The model explains a bounded evidence packet. A person decides what to inspect next, and the model cannot execute repairs.
What I Will Build
The finished project has three required parts, and the AI part is not optional:
A frontend
A run list, filters, a step timeline, an accessible detail dialog, review controls, and clear loading, error, and explanation states.
A backend
An API and PostgreSQL store that accept events, keep them durably, tolerate duplicates and late arrival, and serve the dashboard.
An AI failure brief
A brief that separates what was observed from what might have caused it, names the steps that support each claim, and is evaluated against known cases.
I will build the interface first, using scripted explanations so it works without any keys, then the real API and database, and then replace the scripted explanations with a live model and evaluate its output. A mocked explanation does not count as a finished AI feature.
Planned Design Decisions
The dashboard reads, it does not act
It shows recorded facts. It does not run the original application, retry anything, or fix a failure automatically.
A missing measurement is not zero
A step with no recorded duration says so, and a duration on a run still in progress is labeled as the last reported value.
Delivery order is not execution order
Each event carries a sequence number, and the server sorts by it, so a completion event that arrives first still lands in the right place. Resending the same event is safe, and reusing its ID for different content is a conflict.
Observability must not break the app
The producer is best effort. If the dashboard is down, the original request still succeeds, and the failed delivery is logged rather than thrown.
Facts, hypotheses, and next checks
The brief keeps observations (each tied to a recorded step), possible causes, and suggested inspections apart, and it can say there is not enough evidence.
Code checks the model
Deterministic code rejects citations to steps that do not exist, observations with no citation, and a brief that ignores known missing events.
How I Will Evaluate It
- •Six known cases chosen to hit different limits: an invalid citation, a timeout, an authentication error, an instruction hidden inside the evidence, missing events, and an unknown failure with no step details
- •A mechanical mock run first, labeled as a contract check and not as a measure of model quality
- •A live run second, with saved outputs and recorded token use, reviewed by hand against a rubric that includes factual support and citation accuracy
- •Human review recorded separately, since a valid schema does not make an explanation trustworthy
- •Frontend unit, component, and browser tests, plus backend tests against a real PostgreSQL database that are never silently skipped
- •A small usability check: ask another person to find the failed step without my narration, fix one confusing interaction, and report the sample size honestly
Every number will come with its case count. A six-case set is a starting exercise, not a reliability estimate, and I will not claim the tool saves anyone debugging time unless I have measured it with people doing a representative task.
Scope
This will be a single-operator demonstration on synthetic data, with shared demo tokens and no tenant isolation. Event history per run is bounded, delivery from the producer is best effort, and the limit on live AI requests is a demo guard that resets on restart. It will not make exactly-once delivery guarantees, remediate failures automatically, or claim production performance. Those limits are written into the plan so the interface never implies them.
