How we use AI at Grafica, and the four things it is not allowed to do
Saying an agency uses AI tells a client little. What matters is how, and where it stops. This is how it works at Grafica, including where it stops.
What AI does here
AI agents read, draft and check. They audit a site, write the first version of a plan, draft reports and client messages, and run the comparisons after a change.
They also work on live systems, production included. That is the part that needs rules.
The order of work: audit, plan, execute
Work on any project goes through three phases, and one phase is closed in writing before the next begins: audit, then plan, then execute. Nothing is created before checking that what it needs is there.
The rule sounds bureaucratic. It exists because of the temptation to skip ahead: quoting before mapping, recommending before measuring. When that temptation appears, the rule is to stop and say so.
The four things AI is not allowed to do
Four kinds of act are reserved to a person, whatever the tool is capable of.
- Credentials. An agent does not type a password or a key. Sign-ins go through a password manager.
- Money and commitments. Prices, payments and promises to a client are made by a person.
- Anything irreversible. Sending a message, deleting something that cannot be restored. A message an agent writes lands in a drafts folder. A person reads it and sends it.
- Decisions that are not technical. For example: whether to take a client, or whether a client may be named in public.
Everything else is delegated, including changes to live sites, under conditions: the way back is written before the change, the host’s own backup is taken first, risky changes are rehearsed on a staging copy, and afterwards the page is read the way a visitor gets it.
A pass only counts if a failure was provoked
This rule applies to every check.
A check that reports «all good» is only evidence if the same check, in the same run, was seen to fail on something that should fail. We ask what the check would print if the thing were broken, and then we break something to find out.
One example. On the mortgage calculator now on our site, a test compared 3,825 computed values across 60 cases against closed-form formulas and found no mismatch. That alone proves little. So one expected value was moved by one cent, and the test was run again. It failed, as it should. Only then did the first result count.
The reverse happens too. A second test, on another of our calculators, reported 31 mismatches. The page was right. The independent solver written to check it was wrong. Trusting the checker would have meant «fixing» a correct page.
A second reader that has not seen the reasoning
Before any text goes to a client, and before a change goes to a live site, a separate reviewer reads it. The reviewer is another AI agent started with a clean slate. It gets the text and the facts that govern it. It never gets the author’s reasoning, so that it judges the text and not the argument for it.
Findings are ranked. The top rank means: carried out as written, this causes a wrong act or a false pass. One of those holds the work until a person decides.
It catches real things. On 5 October 2026, while we were reorganising this blog’s categories, the reviewer pointed out that the plan would have changed the address of all three existing posts. The plan was changed before anything was touched.
This post went through the same review. So did the others in this series. The first pass found problems in every one of them.
Nothing lives only in a chat
Decisions are written into a repository. Nothing lives only in a chat. That is what lets the next session, or a different person, pick the work up without guessing.
Where it has gone wrong
It would be dishonest to describe only the catches.
One report gave a figure from one instrument before the cross-check had run. The figure held, but the order was wrong. Drafts have been written from build notes before the page they described was opened and exercised, and a reviewer found three sentences that rested on old notes. On this blog, a staging rehearsal could not tell pass from fail, because the staging address already redirected before the change, and it was reported as such.
Each of these ends up in a section every report carries: what failed about the method, apart from what failed about the work. We ask for it because a tool that only reports success teaches you nothing.
What this means for a client
You are not paying for AI. You are paying for a result, and for someone who answers for it. Checking every item instead of a sample is, in our view, what the tools make affordable: on the mortgage calculator, 3,825 values, not a handful. The person is why the result can be trusted.
The X-Ray on our Demos page reads a business’s website and ranks where AI and automation fit. Our work is on our Work page.
Ask us how
Book a 30-minute call and ask how these rules would apply to your site.