Skip to content

My First AI Agent

An agent that plans and calls tools to finish a job — plus an honest account of where it lost the plot.

Status
In progress
Period
Aug 2026 – present
Role
Solo
Stack
  • AI
  • Python
  • LLM
On this page

The problem

The repetitive part of the job was never the thinking, it was the walking: look up a record, check a service state, cross-reference a log, then write a paragraph about it. That is exactly the shape of work an agent should absorb.

What I built

The goal was modest: an agent that could take a request in plain language, decide which of a small set of tools to call, and come back with an answer and a trace of how it got there.

It took two weeks, and about nine days of that was learning where my assumptions were wrong.

Key decisions

The first version trusted the model too much. The one that worked does three things differently.

  • Fewer, narrower tools. Five generous ones collapsed into three with required, typed arguments.
  • The loop lives in code. Which tool runs first, what happens on failure, and how many attempts are allowed are decisions I write, not decisions I describe.
  • Errors are written for the model. A specific failure message produced a better second attempt than any amount of extra instruction in the prompt.
tools.py
def lookup_record(record_id: str) -> Record:
    """Fetch one record by its exact ID."""
    if not RECORD_ID.fullmatch(record_id):
        raise ToolError(
            f"'{record_id}' is not a record ID. IDs look like REC-12345."
        )
    return records.get(record_id)
Errors written for the model: the second attempt improved more from this message than from prompt changes.

What I learned

It handles the walking reliably and writes a decent first draft of the summary, and a person still checks the answer before it leaves the building. I consider that the correct result rather than an unfinished one — the trace is what makes it checkable.

What broke

Long plans. Given room to plan several steps ahead it became confident and slower, and the extra steps rarely improved the answer. I have not found the rule for how much planning is right, only the limit of what is wrong.

Next

Cap planning depth by default and measure how often the trace is actually read.

More builds

Other things I've built

  • Shipped2026

    RabbitMQ Consumer Test Console

    A tool for hitting a message consumer with traffic and watching retries, dead letters and back-pressure in real time.

    Outcome Found two retry storms before they reached production-like load.

    • RabbitMQ
    • Java
    • Spring
    • Observability
  • Shipped2026

    Event-Driven Billing System

    Invoices and payments modelled as events, designed so a retry never charges anyone twice.

    • Kafka
    • Spring
    • PostgreSQL

Next

Building something similar?

I am happy to talk about the boring parts — design, failure modes, and the places I would not trust the first version.

Get in touch