Writing code got cheap. Knowing it is right did not
By Luis Gallego
I founded Grafica in 2012 and I have designed and built websites for a long time: hundreds of them. For most of those years the scarce thing was typing. A feature cost what it cost because someone had to write it, line by line.
That is over. An AI agent now writes a working first version of most things faster than I can describe them. I use these tools and I would not go back. But there is a wrong conclusion on offer to people who buy websites, and I want to argue against it here.
The wrong conclusion is: code is nearly free now, so development should be nearly free.
What actually got cheap
Typing got cheap. First drafts got cheap. So did boilerplate, the conversion of a design into markup, the tenth variation of a component, and the script that compares two lists.
This is real and it is good. On our Demos page there are thirteen small tools: a clickable site plan, a TV guide, a quote builder, three calculators. A few years ago I doubt a company our size would have built thirteen demos just to show what it can do.
What did not get cheap
Three things, and they were always the expensive part. We just could not see them behind the typing.
Knowing what to build. A calculator is not a form with a formula. It is a decision about what a visitor will do with the number. A lot plan is not an image with hotspots. It is an answer to the questions a buyer would otherwise ask on the phone. No tool decides that for you. It will build the wrong thing, well, and very quickly.
Proving it works. This is the one that surprises people. When code was slow to write, the person who wrote it understood it. Now code arrives finished, plausible and unread.
Here is a small recent example. One of our calculators builds a loan schedule month by month. Running it and looking at the result would prove only that the code agrees with itself. So the check we built computes the same figures another way, with the closed-form formulas for a loan, across sixty cases and 3,825 values. Then we changed one expected value by one cent to make sure the check could fail. It did. That is the moment the calculator became something I would put my name on.
On another of the calculators, the independent check reported 31 mismatches. The calculator was right. The check was wrong. If I had trusted the tool that reported 31 mismatches, I would have broken a correct page to satisfy a broken test.
The same calculator had a defect that had nothing to do with math. With the cursor in a number field, turning the mouse wheel changed the number. A visitor who clicked into a field and scrolled down to read the result would have changed their own input without knowing. No test of the formulas would ever find that. A page has to be used, not only computed.
Answering for it. When a store loses orders in a migration, nobody wants to hear that the tool was confident. Someone has to have checked, and someone has to say «that was ours, and here is the fix». That person’s judgment is what a client is hiring.
One model maker’s own testing points the same way
I do not ask you to take the point from me. In the last week of September 2026, OpenAI launched a new model and, according to reporting by The Wall Street Journal carried by TechCrunch on 29 September, scrapped the release of another after internal testing showed higher levels of deception and a tendency to go ahead with tasks without asking the user for permission.
Read that again as a buyer of software. If that reporting is right, one of the companies building these tools held a model back because, in its own testing, it deceived more and went ahead with tasks without asking. That is not an argument against using AI. It is an argument for never letting it be the only one who checks its own work.
What changed in how I work
I write less and read more. More of the work is now review: is this the right thing, does the evidence support the claim, what would this check print if the thing were broken.
A check counts only if I have seen it fail. A pass means nothing until the same check has been made to fail on purpose in the same run.
The risky acts stay with a person. Passwords, payments, sending, deleting. An agent prepares the work up to that edge and stops.
Every report says what went wrong with the method. Tools are very good at reporting success. I ask for the other section.
What to ask whoever builds your site
You do not need to understand the tools. You need five questions.
- Which parts did a person read, line by line?
- How do you know it works? Show me a check failing on purpose.
- What happens to my passwords and my customers’ data when an AI agent is working on my site?
- If it breaks next month, who answers, and what is the way back?
- What did you decide not to build, and why?
A team that uses AI well will enjoy those questions. A team that is using it to skip the expensive part will change the subject.
What I think happens next
Websites will get cheaper to start and no cheaper to get right. The gap between the two will be where the damage is done: stores that go live with nobody having placed a test order, forms that send nothing, calculators that are wrong by a little.
I would rather work on the right side of that gap. The sites on our Work page are how we try to show it.
Talk to me
Book a 30-minute call. Bring the five questions and ask them of us.