Tabs That I Keep Open

Things I noticed while building stuff, breaking stuff, and trying to work out why.

An Answer Is the Start. Show Me the Evidence.

Building LaunchPad for the WebMCP Challenge: a research factory where humans and browser agents share a workspace, challenge assumptions, and trace decisions back to evidence.

#webmcp#ai#engineering#product

My First Eval Set Was Too Nice

My first eval set mostly tested the happy path. I had to add the awkward questions before it started telling me anything useful.

#ai#evals#engineering

RAG Starts With the Documents

I kept tuning prompts when the real problem was retrieval. The model cannot use facts that never made it into the context.

#ai#rag#retrieval

Retries Are Not a Safety Plan

An agent repeating the same tool call is not being persistent. It is usually just being expensive.

#ai#agents#reliability

I Keep Adding One More Metric

I wanted one useful benchmark and kept adding charts. At some point I had to decide which numbers would actually help me debug a run.

#ai#observability#benchmarks

The Bot Needs Boring Rules

Pocket Nenek became more reliable when I stopped letting the model own the whole conversation.

#telegram#ai#automation#engineering

Somehow, Google Sheets Became My CMS

I wanted an easy place to edit food stories. I got schemas, sync logs, exact headers, and a CMS wearing gridlines.

#google-sheets#content#pocket-nenek

The Bot Was Not the Hard Part

The hard part was figuring out what the client actually wanted. The bot was the small bridge after that.

#ai#automation#client-work#product

One Request Was Not a Benchmark

One request made a nice demo. Replay and concurrency made it feel like an actual benchmark.

#ai#engineering#profiling#benchmarks

p99 Ruined My Very Good Average

Average latency made the run look healthy. The tail had a different story and, unfortunately, better evidence.

#ai#latency#inference-autopsy

The Model Was Not the Whole System

I kept finding failures around the model. That pushed the profiler toward retrieval, tools, and answer quality too.

#ai#engineering#profiling#evals

What I Read Before Building the Profiler

I made a reading list before writing the profiler. I wanted the metrics to mean something before I put them in a chart.

#ai#engineering#inference#learning

The QR Code Was the Easy Part

CARP looked like a QR registration project. The actual work was identity, caregivers, capacity, and refusing duplicates.

#databases#product#carp

The Map Was Not the Product

Talking to early Hawkerly users made me stop confusing nearby pins with a reason to try somewhere new.

#product#food-discovery#hawkerly

I Do Not Trust One-Shot Agents

Complicated work needs a plan, tools, and checks. A confident wall of text is not a workflow.

#ai#engineering#agents

What git add -p Taught Me About Committing

I used git add . for years without questioning it. Learning interactive staging made me realize committing is not just cleanup, it is communication.

#git#workflow#lessons

Tabs That I Keep Open

These are the things I keep thinking about after I close the editor. Some of them stay open for a while.

#meta#philosophy#career

Mostly notes to future me. You are welcome to read them too.