Agent Build · 2 weeks · €3,000 fixed
Your first agent, working in two weeks
One task, chosen in the first days with the payback math in the open, then built and running on your own data — an agent doing the work, with a number on how often it gets it right and a list of what it does not.
What determines whether the first agent works
The result depends on choosing the right task, agreeing what a correct answer is, and testing the agent on real inputs.
Testing beyond the demo
A useful build is tested on real inputs beyond the rows selected for a demonstration.
Agreeing what a right answer is
Without a set of real cases and the answers written down in advance, “it works” is an opinion, and so is “it got worse after the last change”.
Choosing the first task
We compare how often each task runs, how long it takes and what it costs. Choosing the first task takes days rather than weeks.
An agent you can run, and a number you can check it against
Two weeks from kickoff to the walkthrough, with a go/no-go on day three and about three hours of your team's time along the way.
- A working agent
- One task, running on your own data at the end of the fortnight, in the mode where the agent drafts and a person approves.
- The number on how often it is right
- Measured on a set of your real cases before handover, not on rows we chose.
- The list of what it gets wrong
- The cases it fails, the ones where it should hand back to a person, and the ones where we would not trust it at all. This is the deliverable that argues against buying more from us, which is why it is on the list.
- The payback math, written down
- Which task we picked and the assumptions behind it, in the open, so you can argue with the arithmetic rather than with our conclusion.
- All of it is yours
- Code, test set and measurements. Carry on with us, your own team, or anyone else.
What the fortnight costs and what happens in it
Fixed price, so there is no scope creep on your bill. For scale: Winder.AI's published rate card puts a two-to-four-week proof of concept at £15,000–£40,000. This is less because it is less — one task rather than a programme, one senior pair rather than a team, and the data you already have rather than a pipeline built to feed it.
- Days 1–3: two one-hour interviews, sample data, and a look at the systems involved. We pick the task and write out the payback assumptions
- Day 3: the go/no-go. If nothing here pays for itself, we say so and there is nothing to pay — not for the three days, not for the interviews. The exit is on day three so both sides can stop before the build starts
- Days 4–6: the test set. Real cases from your work with the correct answers agreed up front
- Days 7–9: the agent, built against that set and run on your data
- Day 10: walkthrough. The agent, its score, and an honest list of what it still gets wrong
Fixed against one task and the data you already have. If the go/no-go on day three comes back no, there is nothing to pay: the fee buys a fortnight that produced a working agent, not the days it took to find out that it would not. If your case needs an integration written from scratch, we say so on the free intro call, before you have paid anything. If a fortnight is a larger first decision than a stranger has earned, start with the €1,500 slice instead: it leaves you something running on one narrow part of the task, and nothing to pay if it does not run. The scope is the test set you sign off on day three; anything outside it is a separate block. Half the fee is invoiced when you say go on day three, half at the final walkthrough.
What is not in the fortnight
- Production hardening and high-availability deployment — the agent runs on your data, not yet under your uptime commitments
- Autonomy without a person approving the output. Draft-and-approve is the mode we hand over in, deliberately
- Integrations that have to be written from scratch against a system we have not seen. We say on the intro call whether your case needs one
- Roles, permissions and audit logs for multiple users
- Monitoring dashboards and alerting on the agent's own behaviour
- A second task. The fortnight is fixed against one
Anything on this list is a separate decision with its own fixed price, agreed before the work starts rather than invoiced afterwards. Most of it is also cheaper to judge once the first agent is running and there is a number to argue from.
Price
€3,000 fixed
2 weeks, kickoff to final walkthrough.
When another route fits better
This service fits one task that runs often enough to matter, data already available in your systems, and someone on your side who can define the right answer for fifty real cases. The following cases need a different first step.
Documenting the process comes first
If the process currently lives in people's heads, your operations lead needs to document one process first. This makes the later build cheaper. We can start when it exists on paper.
A settled specification needs implementation
If the specification is settled and you need implementation capacity, a contractor working to that specification is a better fit. Our fortnight of judgement would add no value.
Ready-made website chat
Website chat is available from established off-the-shelf products at software prices. Choose one of those when it covers the task.
Evidence for an implementation decision
This service produces a running agent, a measurement of its accuracy and a record of its errors. It fits when those results can inform an implementation decision; if the decision is already settled, there is no reason to pay for it.
Agents we have already shipped
Each of these went the same way: one job, real data, and a way of telling whether it was doing the job right.
In production
Catalogue-to-Offer Agent for a Uniform Manufacturer
A production workflow connecting catalogue articles, approved plans, visual checks and saved revisions, with a person choosing the images and accepting the exact files used in the final PDF.
Read the project →In production
Country Explorer: Location Intelligence for Restaurants
A team of agents reads location, footfall and demographic data and writes sourced expansion briefs a human can check against the figures behind them.
Read the project →In production
LetAI: Nutrition Estimation Agent in Production Chat
The agent only reached production chat after its evaluation pipeline was expanded, with new datasets and broader case coverage added before release.
Read the project →
Frequently asked questions
We write down the assumptions behind the payback calculation so you can check the arithmetic. The go/no-go is on day three; if the answer is no, we receive no payment, including no reduced fee for those days. That is why the calculation happens at the start rather than being sold as a separate service. The walkthrough measurement uses your cases. You can also buy the analysis from an independent adviser and bring us the conclusion; we will build against their brief.
The fixed price covers one task, your existing data, and draft-and-approve rather than autonomy. We write those boundaries down before the work starts. Measurement remains part of the deliverable and can show that the agent does not meet the agreed bar.
Real enough to do the work on your data and be measured doing it. It is one task, and it runs where a person still approves the output. What is not in the fortnight: production hardening, wider autonomy, and integrations that have to be written from scratch against a system we have not seen. Those are a separate decision, and we say which of them your case needs before the start rather than at the walkthrough.
Then we say so in the first days, with the math, and you pay nothing at all, including no reduced fee for the days we used. The write-up is still yours: what would have to change for the answer to become yes, the data you would need to start keeping, the volume at which it starts to pay, and the step that would have to be written down first. Reaching that answer takes three hours of your team's time.
Yes, and for a first engagement we prefer it. A slice costs €1,500 and takes 20 hours, and it is not a report: what you get is a working thing on one narrow part of the task, which you run yourself. You put your files in, you get the result out, on your own machine without us. With it comes the number it scores on your real cases, the errors broken down by kind so you can see what a rule would fix and what runs into data you do not have, a set of forty to fifty cases with agreed correct answers that stays yours, and what it would cost to widen this to the whole task. What a slice cannot include in 20 hours is a connection to your mail, CRM or accounting system: it works on an export, and the integration is quoted separately. If no working slice comes out of those hours, you do not pay for them. The fortnight remains available on its own when the task is obviously well defined, but we will not propose it first: for €1,500 you get something running and a number on it, and you decide about the larger figure having seen the work rather than a description of it.
Then we say so on day three and offer a different shape of work, not a discount. A fixed-price outcome requires your own people to agree on what a correct answer looks like. In the first three days, two of your staff label the same cases independently and we compare them. Where they agree, what remains is engineering, and engineering can be priced. Where they half agree, the missing piece is the rule itself, and a rule is written by the hour: the same block of 20 hours, except what it buys is our time and a written finding rather than a promised number. Where they do not agree at all, we do not take the work in any form.
About three hours of your team's time in the first days, a look at your data and systems, and — this is the part that decides the quality — someone who can say what the correct answer is for the cases in the test set. We do not need production access for the fortnight.
The same PhD-trained researchers and engineers who ship our projects. You can see who they are on the team page.
Whatever the number says it should. Usually one of three: widening the agent's autonomy where the measurement supports it, putting your documents behind it so it answers from what the company knows, or keeping the test set running against every model and prompt change so a drop shows up before your users find it. The first two are bought in the same unit: a block of 20 hours at €1,500, with a walkthrough at the end of it. The third is not a block, because a one-off cannot catch a model that changes six weeks from now: it runs monthly at €300 a month, which buys the set re-run on every model and prompt change, the number sent to you each month, and the repair at our expense if it falls more than three points below the threshold. It is also optional, and we say so before invoicing rather than after: the test set and the script that runs it are yours either way, and if someone on your side will press the button once a month, you do not need us for that. How many blocks a given step takes depends on the number the first agent scored, and quoting that total before the number exists is guessing. What we can say is that blocks after the first go further, because the test set is already built.
Want to see who you'd be working with? Meet the team.
No commitment to continue
Request a build
Tell us what the agent would have to do. The first 30-minute call is free: if the task wouldn't pay off, we'll say so before you spend the €3,000.