When we build software, we need to define criteria for acceptance, and when it comes to AI agents I believe an agent is only as good as the evals you run on it, so that’s where I start. Hi, my name is Matias Lapolla and I’m currently working at Persiscal building AI native products. I have my little sidehustle entrepreneurship at Senda where I experiment with software. I also wrote about why the description string matters more than the code and what memory is for. More details on my CV and more articles on the blog.
Writing 2026-09-02
An agent that does real work is a small graph: a model node, a tools node, a human approval gate, a recovery node, and a state definition with two fields. That part is an afternoon. Mine worked every time I tried it by hand and failed three runs out of five on one task. When the search tool flaked, it called web_search seven times in a row and then answered nothing. I only saw it because I ran the same task five times and counted. One run tells you how the dice landed, not how the agent behaves. One clause in the system prompt took that case from 2/5 to 5/5, and the same table showed the two honesty cases it cost me. The total moved up 2.5 points and would have congratulated me for shipping it. The same suite picked the model. I ranked eight local models on the same eight cases: the 8B I would have guessed reached for a tool to say hello three times out of three, and a 7B swept the suite four times faster than the 9B.
So You Want to Build AI Agents That Do the Work? Start with Evals
Work AI Engineer · 2024—
- AI Agents Platform 2026—
- Stealth Project 2026—
- GeoScan 2026—
- Starlings Travel 2024—2026
Persiscal
Applied AI in production: an agents platform, a GEO/SEO product, and the enterprise travel system that came before them.
Building
Senda
Building software experiments in the AI era — a system that answers and gets the job done.
trysenda.xyz
Writing 2026-06-08
Your agent has the perfect tool wired up and still answers from memory, calls the wrong one, or quotes your refund policy and then promises the refund — because the model never saw your code, only the description string, and the string didn't tell it any better. The model programs against three strings: the tool name, the description, and each parameter description. Those strings are the API. The execute body is invisible to it. The move that makes it work: write the description for a new hire, not a compiler — encode what it returns, when to reach for it, and the one thing it must NOT conclude from the result (a found: false means "ask the customer," not "invent a status"). A precise description is worth more than a smarter model: Anthropic got SOTA on SWE-bench by refining tool descriptions alone, and a tight contract cut my agent's wrong-tool and hallucinated-state errors to near zero — at zero extra tokens per call.
MCP: Why the Description String Is the Most Important String in Your Harness
Writing 2026-07-03
lastMessages: 20 is where everyone starts and it leaks fast: the customer told you their account tier on turn 3, and by turn 40 the model is asking again — or worse, guessing. Mastra stacks three recall tiers: a recency window (verbatim last N), working memory (a durable per-customer profile the agent rewrites), and semantic recall (embedding search over all past messages). The move that makes it work: scope: 'resource' — one profile and one searchable history per customer, not per chat, so memory survives across conversations. For long threads, Observational Memory runs a three-agent loop (Actor / Observer / Reflector) that compresses history in the background, keeping context flat while conversations run past 200 turns. Part 3 of a series on harness engineering with Mastra. It builds directly on Part 2 — Processors and Guardrails, because in Mastra memory is processors. The default memory setup from every tutorial:
Mastra Memory: The Context Layer That Survives Long Conversations
Work PDF
Curriculum vitae
Want to hire me? Take a look.
Writing
- So You Want to Build AI Agents That Do the Work? Start with Evals 2026-09-02
- Mastra Memory: The Context Layer That Survives Long Conversations 2026-07-03
- MCP: Why the Description String Is the Most Important String in Your Harness 2026-06-08
- Mastra Processors: la capa de guardrails que se ejecuta en cada paso 2026-06-01
Blog