Playbooks · 4 min read
The 30-day playbook, week by week.
From the first call to production: the exact steps, artefacts and sign-offs we use.
Thirty days sounds fast for putting AI agents into production. It is fast. It is also the reason most of our projects ship: a fixed date forces every decision that a pilot would postpone to be taken in the first week.
This is the playbook we follow. Every step ends with something written down and someone’s name next to it.
Week 0: diagnose
It starts with a free call. If there is a fit, we spend the next week inside your operations, most of it with the people who do the work. We sit in on the inbox, the report, the supplier file. We want to see where the hours go and where the money leaks, not how the process is described in a slide.
The artefact is a map of your operations with the value of every agent we would build, and a fixed price. The sign-off is yours: which process goes first. We push for one with volume, clear rules and a number someone already cares about.
Week 1: map
This is the week most pilots skip, and the week that decides whether yours ships.
- Access. Real credentials to the real systems: the CRM, the helpdesk, the ERP, the inbox. Not an export.
- Real cases. A set of past emails, tickets or files, with what a good answer looked like. They become the tests.
- One metric, signed off. First response time, time to spot a new listing, partner hours on admin. One number, owned by the team whose work changes, signed by the person who owns it.
- First approval rules. What the agents may do alone, and what always goes to a person, written in plain language.
If week 1 ends without access and a signed metric, we say so, because it is cheaper for both sides than finding out on day 29.
Week 2: build
Now we build the agents, the rules, the shared memory and the command center, wired into your tools. The agents live where your team already works; nobody gets a new app to learn.
The real cases from week 1 become evals: automated tests that run every answer against what a good answer looked like. Customer-facing answers get a judge. Tone is agreed with the people who own it, often in a single working session with a pile of past replies.
The sign-off at the end of the week is a walkthrough with the team lead, on their own cases.
Week 3: prove
First the evals, on real cases the agents have never seen. Then shadow mode on live work: the agents draft, people send. Every edit a person makes is a signal, and we read them daily.
This is also where the handoff rules are tuned with the team. Which cases go straight to a person, at what threshold, through which channel. The team should be able to change a rule themselves, and in week 3 they practise doing it.
| STEP | ARTEFACT | WHO SIGNS |
|---|---|---|
| Week 0 | Operations map, value per agent, fixed price | You: which process first |
| Week 1 | Access, real cases, first approval rules | The owner of the metric |
| Week 2 | Agents, memory, rules, command center, evals | The team lead, after a walkthrough |
| Week 3 | Eval results, shadow-mode log, tuned handoffs | The team lead: ready to go live |
| Day 30 | Production, monitoring, trained team | Both of us, against the metric |
Day 30: live
On day 30 the agents are in production, on live work, monitored from the command center. The team is trained on it: how to read a run, how to approve, how to pause an agent, how to change a rule.
Production means the agents answer, route or update on their own inside the rules, and the number from week 1 is tracked every day from then on. It does not mean “available on request” or “ready for phase two”.
Why the date holds
Three things keep the date honest. The scope is one process, not the company. The metric is agreed before a line of code, so nobody can move the goalposts later. And we put our money on it: live in 30 days, or your money back.
After day 30 we run the fleet: monitoring, tuning and, when you are ready, the next squad. On our largest project that became one new squad a month, each built on the memory and rules of the ones before it.
Got a pilot gathering dust?
In 30 minutes we’ll tell you what it would take to put it live.
André is the founder and CTO of WizardingCode. Eight years building the software companies run on, now putting agents into production.