An Answer Is the Start. Show Me the Evidence.
An AI answer can sound convincing long before it deserves your confidence. That is the problem I built LaunchPad around for the WebMCP Challenge: if a team is going to act on a recommendation, it should be able to inspect what is holding that recommendation up.
The starting point is deliberately small. You describe a problem. LaunchPad researches it, gathers supporting and contrary evidence, develops a recommendation, and stress-tests it. The output includes the reasoning, the limitations, and a validation plan. It is an idea worth testing, with the assumptions still attached.
The part I find most interesting happens after the first answer. What changes if community posts do not count? What if the research has to be recent, or each feature needs independent corroboration? LaunchPad can compare those policies against the same evidence ledger, show how support changes, and apply or roll back the policy. Sometimes the leader changes. Sometimes it survives the stricter test. Both outcomes tell you something.
That only works because evidence has its own structure. A source produces a finding; findings support an insight; the insight supports a particular component of a candidate solution. The trace view follows that path back from the proposed decision to its source. Citations become something the application can reason about, rather than a bibliography pasted under generated prose.
WebMCP made this an interesting browser project. The person and the browser agent operate the same live workspace through the same TypeScript service. The agent can read the current brief, inspect missing evidence, import findings with provenance, and help move the workflow forward. The page registers the tools relevant to its current stage from a catalog of 22, so the agent sees useful next actions as the work progresses.
I wanted that collaboration to be visible. The voxel factory gives the workflow a physical shape: a problem enters, research moves through the production line, and a blueprint comes out. Its stage and progress come from the application state. The activity log provides the more precise view, with operations and workspace versions that can be inspected after a run.
Giving an agent actions also means deciding when an action should stop. Sensitive evidence review, finalization, and private export use an in-page consent checkpoint tied to the exact records and workspace version. A stale approval cannot quietly authorize a newer decision. Policy previews leave the workspace untouched, and applying a policy preserves the ledger so the comparison can be reversed.
The research pipeline has its own boundaries. Report URLs must appear in the web-search output. AI-generated findings are labeled as citation-linked paraphrases, with the original source available for inspection. Contrary evidence stays visible, and a request that lacks essential details can return clarification questions. A plausible paragraph is not enough to clear the evidence gates.
For the recorded judging journey, the browser completed 26 tool calls from workspace v1 to v19. That run covered policy comparison and rollback, human consent, a four-hop proof trace, and public-safe export. The saved journey records those steps. The separate provider-backed agent evaluations hit rate limits; I kept those failures in the report. A successful recorded walkthrough and broad agent reliability answer different questions.
The product is still a challenge build. Its allowance controls work in the browser, while payment collection and cross-device account persistence remain production work. The curated demo also labels its synthetic private evidence. Those boundaries matter because the central promise is inspectability, and that should apply to how I present the project too.
What I took away from this build is that the useful part of an AI research tool continues after the answer appears. Being able to challenge a source, expose a weak assumption, and see which decisions still stand makes the output much more useful to someone who has to act on it.