Back to projects

Inference Autopsy

Trace-First AI System Profiler & Regression Tester

SAMPLE TRACE · REPLAY
5 stages
  1. 01 · workloadrag-longprofile loaded
  2. 02 · retrievaltop-k contextevidence preserved
  3. 03 · inferencestreamingtoken timings captured
  4. 04 · diagnosisContext Blobfishrule matched
  5. 05 · gatecandidate vs baselineready to inspect
The illustrative replay records workload, retrieval, inference, diagnosis, and regression-gate evidence.
TTFT
942ms
ITL P95
71.2ms
STALLS
2
DIAGNOSIS
Context Blobfish
Illustrative sample trace — values mirror the project's public trace schema, not a live benchmark.

Metrics

  • •Published Python package
  • •Public interactive trace explorer
  • •Deterministic and protected live CI regression lanes

Overview

An open-source black-box profiler, workload replayer, and regression tester for OpenAI-compatible LLM endpoints and retrieval workflows. Inference Autopsy records request-level trace evidence, explains latency and quality failures, compares candidate runs against approved baselines, and turns regressions into CI decisions instead of vague performance complaints.

Technologies Used

PythonTyperPydanticHTTPXJSONLPytestNext.jsTypeScriptGitHub Actions

Business Need

AI systems often feel slower or less reliable without leaving enough evidence to explain why. A single average-latency number cannot distinguish first-token delay, slow decoding, tail latency, streaming stalls, rate limits, retrieval regressions, or answer-quality failures. Teams need reproducible traces and explicit gates before they can diagnose or prevent regressions with confidence.

Purpose

To make AI-system performance failures inspectable and repeatable: benchmark a workload, preserve the evidence, diagnose the failure, replay the same shape, and block regressions before promotion.

Key Functionalities

  • •Trace-First Benchmarking: Records request timing, token gaps, workload metadata, errors, and workflow-stage evidence as append-friendly JSONL.
  • •Diagnosis Rules: Converts trace-derived symptoms into memorable failure labels while keeping the supporting metrics visible.
  • •Baseline Diff & Gates: Compares runs and returns distinct exit codes for a passing comparison, a valid regression, or invalid benchmark evidence.
  • •Workload Replay: Replays exact prompts when content is available or regenerates a privacy-preserving workload shape from trace metadata.
  • •Workflow Profiling: Measures retrieval, prompt assembly, model execution, end-to-end latency, recall, and deterministic answer correctness.
  • •Public Evidence Layer: Provides a browser-based trace explorer and bounded live benchmark while preserving the local CLI as the serious tool.

Advantages

  • •Reproducible Evidence: Each result can be traced back to request-level JSONL instead of a dashboard-only aggregate.
  • •Useful CI Semantics: Regression gates can block promotion without treating malformed or incomplete benchmarks as ordinary failures.
  • •Privacy-Aware Replay: Hash-only traces refuse exact replay, while shape replay preserves workload characteristics without exposing prompt text.
  • •System-Level Coverage: Retrieval and answer-quality regressions remain visible even when raw model latency looks healthy.

Key Learnings

  • •A benchmark is only useful when its workload, environment, warm-up behavior, sample count, and trace schema are reproducible.
  • •Percentiles and request-level evidence explain failures that averages hide, especially under concurrency and at the tail.
  • •Deterministic fixture lanes and protected live lanes solve different validation problems and should not be collapsed into one CI job.
  • •Profiling the model is not enough when retrieval, prompt assembly, tools, and answer quality can fail around it.

Challenges Faced

  • •Separating first-byte, first-token, token-gap, and end-to-end timings consistently across streamed and non-streamed responses.
  • •Designing gate expressions and exit codes that distinguish a real regression from invalid or incomplete benchmark evidence.
  • •Making traces useful for replay without silently storing sensitive prompts, answers, or retrieved documents.
  • •Keeping the public demo safe and bounded while ensuring the CLI, static reports, and CI workflows remain the source of truth.

Key Accomplishments

  • •Built and published a Python CLI that benchmarks OpenAI-compatible endpoints, records trace evidence, diagnoses failures, replays workloads, and compares baselines.
  • •Implemented latency, throughput, reliability, retrieval, and answer-quality metrics with percentile summaries and trace-backed diagnosis rules.
  • •Created deterministic fixture regression checks alongside protected live benchmark workflows with validation, artifacts, and promotion-blocking gates.
  • •Shipped an interactive Next.js evidence layer for exploring traces, inspecting token timelines, editing gates, and running bounded live benchmarks.