Shipping AI agents as a product engineer
How I think about tool-calling, orchestration, and evals when building agent workflows in production. Not ML research. Product engineering that ships.
Most “AI agent” talk collapses into two piles: research papers about models, or demos that work once in a notebook. Production agents sit in a third pile. They call tools, fail mid-flight, need retries, and have to be cheap enough to run on every customer request.
I am not an ML engineer. I do not train models. I build the systems around them: APIs, orchestration, tool contracts, observability, and the product decisions that decide whether an agent is useful or expensive theater.
What an agent is (in product terms)
An agent is a loop that:
- Takes a goal in natural language or structured input
- Chooses tools (APIs, DBs, message sends, lookups)
- Acts, observes, and decides whether to stop or continue
- Returns a result a human or another system can trust
The model is one component. The product is the loop plus the guardrails.
Tool-calling before clever prompts
The highest leverage work is usually boring:
- Narrow tools with typed inputs. A tool that “does CRM stuff” is a liability. A tool that
create_sms_campaign(audience_id, template_id)is auditable. - Idempotency. Agents retry. If a tool can charge a card or send an SMS twice, you need keys and side-effect design, not a better system prompt.
- Timeouts and budgets. Cap steps, tokens, and wall clock. An unbounded agent is a cost bomb.
- Human-readable traces. When something goes wrong, an engineer should see which tool ran and why, not a blob of chain-of-thought.
At Fello I worked on multi-agent orchestration across communication channels (calling, SMS, AI integrations). The pattern was the same: agents are useful when the tools are solid and the orchestration is explicit.
Orchestration patterns that hold up
Single agent + tools is enough for many products. Add a second agent only when roles are truly different (planner vs executor, or channel specialists).
Fan-out / fan-in helps when you need parallel lookups (fetch profile, fetch history, draft reply). Merge with a deterministic step, not another vague prompt.
State machines beat free-form graphs for anything that touches money or customer messaging. The LLM proposes; the state machine decides what is allowed next.
Evals without a research lab
You do not need a paper to ship. You need:
- A golden set of 20–50 real tasks with expected outcomes
- Checks that are deterministic where possible (tool args, status codes, schema)
- A human spot-check for tone and judgment on a sample
Run evals on every prompt or tool contract change. If you cannot say whether a change made the agent better, you are guessing.
Agents vs embeddings products
I also shipped AniReco, which used embeddings and hybrid ranking. That is a different product shape: retrieve and score, not loop and act. Both are “AI,” but the failure modes differ. Recsys fails by ranking poorly. Agents fail by doing the wrong tool call. Design for the failure mode you actually have.
What I look for in agent work
- Clear ownership of a customer-facing workflow
- Tools that already exist or can be built without waiting on research
- Room to measure cost per successful task
- Teams that want production reliability, not demo theater
If you are hiring for agent systems and need someone who ships the product layer, reach out. The case studies and About page are the packet.