My First AI Agent
An agent that plans and calls tools to finish a job — plus an honest account of where it lost the plot.
- Aug 2026 – present
- Solo
On this page
The problem
The repetitive part of the job was never the thinking, it was the walking: look up a record, check a service state, cross-reference a log, then write a paragraph about it. That is exactly the shape of work an agent should absorb.
What I built
The goal was modest: an agent that could take a request in plain language, decide which of a small set of tools to call, and come back with an answer and a trace of how it got there.
It took two weeks, and about nine days of that was learning where my assumptions were wrong.
Key decisions
The first version trusted the model too much. The one that worked does three things differently.
- Fewer, narrower tools. Five generous ones collapsed into three with required, typed arguments.
- The loop lives in code. Which tool runs first, what happens on failure, and how many attempts are allowed are decisions I write, not decisions I describe.
- Errors are written for the model. A specific failure message produced a better second attempt than any amount of extra instruction in the prompt.
def lookup_record(record_id: str) -> Record:
"""Fetch one record by its exact ID."""
if not RECORD_ID.fullmatch(record_id):
raise ToolError(
f"'{record_id}' is not a record ID. IDs look like REC-12345."
)
return records.get(record_id)What I learned
It handles the walking reliably and writes a decent first draft of the summary, and a person still checks the answer before it leaves the building. I consider that the correct result rather than an unfinished one — the trace is what makes it checkable.
What broke
Long plans. Given room to plan several steps ahead it became confident and slower, and the extra steps rarely improved the answer. I have not found the rule for how much planning is right, only the limit of what is wrong.
Next
Cap planning depth by default and measure how often the trace is actually read.