An Answer Is the Start. Show Me the Evidence.
Building LaunchPad for the WebMCP Challenge: a research factory where humans and browser agents share a workspace, challenge assumptions, and trace decisions back to evidence.
Things I noticed while building stuff, breaking stuff, and trying to work out why.
Building LaunchPad for the WebMCP Challenge: a research factory where humans and browser agents share a workspace, challenge assumptions, and trace decisions back to evidence.
My first eval set mostly tested the happy path. I had to add the awkward questions before it started telling me anything useful.
I kept tuning prompts when the real problem was retrieval. The model cannot use facts that never made it into the context.
An agent repeating the same tool call is not being persistent. It is usually just being expensive.
“OpenAI-compatible” got me through the first request. The second provider reminded me to test the edges too.
I wanted one useful benchmark and kept adding charts. At some point I had to decide which numbers would actually help me debug a run.
Inference Autopsy started as a stopwatch. The traces became more useful than the average.
AI can write the code quickly. I still need to run it and look at what actually happened.
I sent valid JSON to a Telegram webhook, forgot that valid and useful are different things, and blamed almost everything except my test.
Pocket Nenek became more reliable when I stopped letting the model own the whole conversation.
I wanted an easy place to edit food stories. I got schemas, sync logs, exact headers, and a CMS wearing gridlines.
A tiny Telegram typing indicator sent me through one missing method, one swallowed error, and far too much suspicion.
The hard part was figuring out what the client actually wanted. The bot was the small bridge after that.
I blamed the flashy components. The profiler showed shared state waking up half the page and brought receipts.
One request made a nice demo. Replay and concurrency made it feel like an actual benchmark.
Average latency made the run look healthy. The tail had a different story and, unfortunately, better evidence.
Pocket Nenek got more reliable when AI handled the fuzzy language and normal code kept the keys to the car.
I kept finding failures around the model. That pushed the profiler toward retrieval, tools, and answer quality too.
I made a reading list before writing the profiler. I wanted the metrics to mean something before I put them in a chart.
A first-week reflection on startup pace, React Router, deployment, design handoff, client communication, and learning while juggling Orbital.
Teyvat Translator felt faster after I deliberately made it wait for stable text. Real-time is strange like that.
Natural-language server control is fun. Strict JSON, allowlists, authentication, and refusal are what make it usable.
CARP looked like a QR registration project. The actual work was identity, caregivers, capacity, and refusing duplicates.
A black-box profiler and regression tester for OpenAI-compatible inference endpoints, built to turn slow LLM vibes into useful traces.
Talking to early Hawkerly users made me stop confusing nearby pins with a reason to try somewhere new.
Complicated work needs a plan, tools, and checks. A confident wall of text is not a workflow.
I used git add . for years without questioning it. Learning interactive staging made me realize committing is not just cleanup, it is communication.
Before I let an agent edit a project, I ask what it might break and how I will notice.
These are the things I keep thinking about after I close the editor. Some of them stay open for a while.
Mostly notes to future me. You are welcome to read them too.