I Do Not Trust One-Shot Agents
I do not trust one-shot agents with complicated work, even when the model sounds very confident. A long answer can look finished while quietly skipping the one constraint that mattered.
When I break the work into a plan, the agent has somewhere to put the uncertainty. It can inspect the project, make a small change, run a check, and come back with evidence instead of pretending it understood the entire task at once.
I also think about autonomy as a dial, not a personality trait. A read-only research task can have more freedom than an agent that edits production data, and both should have limits that are visible in the tools they receive.
The tools are part of the design. A narrow tool with a typed input and a useful error is often safer than giving the model a general shell and hoping the prompt keeps it polite.
The loop matters just as much as the model: plan, act, observe, check, and either continue or stop. I want failures to become part of that loop, not something the system smooths over so the transcript looks impressive.
I am less interested in building an agent that can do everything than one that knows when it needs help. A small plan, bounded tools, a budget, and a human review point are not signs that the agent is weak; they are what make it usable.