Inference Autopsy
Trace-First AI System Profiler & Regression Tester
- 01 · workloadrag-longprofile loaded
- 02 · retrievaltop-k contextevidence preserved
- 03 · inferencestreamingtoken timings captured
- 04 · diagnosisContext Blobfishrule matched
- 05 · gatecandidate vs baselineready to inspect
Metrics
- •Published Python package
- •Public interactive trace explorer
- •Deterministic and protected live CI regression lanes
Overview
An open-source black-box profiler, workload replayer, and regression tester for OpenAI-compatible LLM endpoints and retrieval workflows. Inference Autopsy records request-level trace evidence, explains latency and quality failures, compares candidate runs against approved baselines, and turns regressions into CI decisions instead of vague performance complaints.
Technologies Used
Business Need
AI systems often feel slower or less reliable without leaving enough evidence to explain why. A single average-latency number cannot distinguish first-token delay, slow decoding, tail latency, streaming stalls, rate limits, retrieval regressions, or answer-quality failures. Teams need reproducible traces and explicit gates before they can diagnose or prevent regressions with confidence.
Purpose
To make AI-system performance failures inspectable and repeatable: benchmark a workload, preserve the evidence, diagnose the failure, replay the same shape, and block regressions before promotion.
Key Functionalities
- •Trace-First Benchmarking: Records request timing, token gaps, workload metadata, errors, and workflow-stage evidence as append-friendly JSONL.
- •Diagnosis Rules: Converts trace-derived symptoms into memorable failure labels while keeping the supporting metrics visible.
- •Baseline Diff & Gates: Compares runs and returns distinct exit codes for a passing comparison, a valid regression, or invalid benchmark evidence.
- •Workload Replay: Replays exact prompts when content is available or regenerates a privacy-preserving workload shape from trace metadata.
- •Workflow Profiling: Measures retrieval, prompt assembly, model execution, end-to-end latency, recall, and deterministic answer correctness.
- •Public Evidence Layer: Provides a browser-based trace explorer and bounded live benchmark while preserving the local CLI as the serious tool.
Advantages
- •Reproducible Evidence: Each result can be traced back to request-level JSONL instead of a dashboard-only aggregate.
- •Useful CI Semantics: Regression gates can block promotion without treating malformed or incomplete benchmarks as ordinary failures.
- •Privacy-Aware Replay: Hash-only traces refuse exact replay, while shape replay preserves workload characteristics without exposing prompt text.
- •System-Level Coverage: Retrieval and answer-quality regressions remain visible even when raw model latency looks healthy.
Key Learnings
- •A benchmark is only useful when its workload, environment, warm-up behavior, sample count, and trace schema are reproducible.
- •Percentiles and request-level evidence explain failures that averages hide, especially under concurrency and at the tail.
- •Deterministic fixture lanes and protected live lanes solve different validation problems and should not be collapsed into one CI job.
- •Profiling the model is not enough when retrieval, prompt assembly, tools, and answer quality can fail around it.
Challenges Faced
- •Separating first-byte, first-token, token-gap, and end-to-end timings consistently across streamed and non-streamed responses.
- •Designing gate expressions and exit codes that distinguish a real regression from invalid or incomplete benchmark evidence.
- •Making traces useful for replay without silently storing sensitive prompts, answers, or retrieved documents.
- •Keeping the public demo safe and bounded while ensuring the CLI, static reports, and CI workflows remain the source of truth.
Key Accomplishments
- •Built and published a Python CLI that benchmarks OpenAI-compatible endpoints, records trace evidence, diagnoses failures, replays workloads, and compares baselines.
- •Implemented latency, throughput, reliability, retrieval, and answer-quality metrics with percentile summaries and trace-backed diagnosis rules.
- •Created deterministic fixture regression checks alongside protected live benchmark workflows with validation, artifacts, and promotion-blocking gates.
- •Shipped an interactive Next.js evidence layer for exploring traces, inspecting token timelines, editing gates, and running bounded live benchmarks.