Back to Portfolio

Product Engineering · Observability for AI Applications

Run Explorer
See Every Step. Find the Failure.

I am building a dashboard for inspecting application runs. It will show what ran, what happened inside it, and where it failed, and an AI failure brief will explain the recorded evidence without claiming more than the record shows.

REACTTYPESCRIPTPOSTGRESQLAI FAILURE BRIEFEVALUATION
Run Explorer: see every step and find the failure

Status: Starting

The scope, architecture, and build plan are written, and I am starting the build. This page describes what I intend to build and how I will measure it. It will be updated with results, code, and a demo as the work happens. Everything will use synthetic data.

The Problem

When an AI application misbehaves, the first questions are simple. What ran? What happened inside it? Where did it fail? What does the recorded evidence actually tell us?

Answering them usually means digging through logs. A model can help explain a failure, but only if it is held to the record: an explanation that sounds plausible and is not supported by the steps is worse than none.

Run Explorer is my attempt at a small, honest answer to that. Its first customer is me, inspecting the runs of Ask TanStack Query. The same event format could later describe Intake Desk.

How It Will Work

A run is one execution of an application workflow, a step is a named piece of it, and an event is a recorded fact about a step:

01An application produces ordered events, such as a step starting, finishing, or failing
02A Python API checks each event and stores it durably, tolerating duplicates and late arrival
03The API rebuilds each run from its events: its status, how complete the record is, and its steps in order
04A React dashboard lists the runs, with search, filters, sorting, and pagination kept in the URL
05Opening a run shows an accessible step timeline that does not depend on color or hover alone
06For a failed run, a model reads a bounded packet of recorded evidence and writes a failure brief
07Code checks every citation in the brief before it is shown, and a person reviews the run and saves a flag

The dashboard reads recorded work. The model explains a bounded evidence packet. A person decides what to inspect next, and the model cannot execute repairs.

What I Will Build

The finished project has three required parts, and the AI part is not optional:

A frontend

A run list, filters, a step timeline, an accessible detail dialog, review controls, and clear loading, error, and explanation states.

A backend

An API and PostgreSQL store that accept events, keep them durably, tolerate duplicates and late arrival, and serve the dashboard.

An AI failure brief

A brief that separates what was observed from what might have caused it, names the steps that support each claim, and is evaluated against known cases.

I will build the interface first, using scripted explanations so it works without any keys, then the real API and database, and then replace the scripted explanations with a live model and evaluate its output. A mocked explanation does not count as a finished AI feature.

Planned Design Decisions

The dashboard reads, it does not act

It shows recorded facts. It does not run the original application, retry anything, or fix a failure automatically.

A missing measurement is not zero

A step with no recorded duration says so, and a duration on a run still in progress is labeled as the last reported value.

Delivery order is not execution order

Each event carries a sequence number, and the server sorts by it, so a completion event that arrives first still lands in the right place. Resending the same event is safe, and reusing its ID for different content is a conflict.

Observability must not break the app

The producer is best effort. If the dashboard is down, the original request still succeeds, and the failed delivery is logged rather than thrown.

Facts, hypotheses, and next checks

The brief keeps observations (each tied to a recorded step), possible causes, and suggested inspections apart, and it can say there is not enough evidence.

Code checks the model

Deterministic code rejects citations to steps that do not exist, observations with no citation, and a brief that ignores known missing events.

How I Will Evaluate It

  • •Six known cases chosen to hit different limits: an invalid citation, a timeout, an authentication error, an instruction hidden inside the evidence, missing events, and an unknown failure with no step details
  • •A mechanical mock run first, labeled as a contract check and not as a measure of model quality
  • •A live run second, with saved outputs and recorded token use, reviewed by hand against a rubric that includes factual support and citation accuracy
  • •Human review recorded separately, since a valid schema does not make an explanation trustworthy
  • •Frontend unit, component, and browser tests, plus backend tests against a real PostgreSQL database that are never silently skipped
  • •A small usability check: ask another person to find the failed step without my narration, fix one confusing interaction, and report the sample size honestly

Every number will come with its case count. A six-case set is a starting exercise, not a reliability estimate, and I will not claim the tool saves anyone debugging time unless I have measured it with people doing a representative task.

Scope

This will be a single-operator demonstration on synthetic data, with shared demo tokens and no tenant isolation. Event history per run is bounded, delivery from the producer is best effort, and the limit on live AI requests is a demo guard that resets on restart. It will not make exactly-once delivery guarantees, remediate failures automatically, or claim production performance. Those limits are written into the plan so the interface never implies them.

Planned Tech Stack

ReactTypeScriptViteTanStack QueryVitestPlaywrightPythonFastAPIPostgreSQLOpenAI APIDockerRender