Agentic SEO Workflows: What Eight Guards Caught That the Agent Did Not
Disclosure
This article is promotional content produced as part of the Semrush Creator Accelerator programme. It contains an affiliate link to Semrush — if you buy through it we may earn a commission at no extra cost to you — and we may also receive a reward from Semrush for publishing it. The pipeline, the failures and the guards described here are our own. Semrush did not supply, review or approve any of it.
We run the SEO for an 883-tool catalogue through an agentic workflow: Claude Code, with MCP connectors to DataForSEO, Bing Webmaster Tools, Google Search Console, Microsoft Clarity and Cloudflare. Research that used to mean five dashboards and a spreadsheet is now a prompt, and the economics are roughly what everyone says they are.
That part worked immediately. This article is about the part that did not.
Three times, the automation reported success while changing nothing. Not a crash, not an error — a clean green result on work it had not done. We only know because we had, by then, built guards that do not trust it. There are eight of them now, they run after every build, and every one exists because something got through.
The trend is right about the bottleneck
Semrush's trend brief on agentic SEO workflows puts the problem in one line: "The challenge isn't automation—it's trusting and validating AI outputs." It names the specific failure mode too — "Raw AI output looks plausible but can be fabricated" — and reports cases of AI fabricating analytics for months before anyone noticed.
Our experience matches that, with one correction we think matters. In our three failures, nothing was fabricated. The agent did not invent a number. It did something quieter and harder to catch: it performed an operation that silently affected nothing, and then truthfully reported that the operation had completed.
That distinction changes what you build. A hallucination check looks for claims that are not true. What we needed was something else — a check that the world actually moved.
Failure one: the guard read yesterday's work
Our data pipeline reads a single large source file and emits small per-page JSON at build time. The integrity checks read the emitted JSON, because that is what the site actually serves.
On 7 September 2026 a fix broke the source file. The build failed. The guard then ran against the output of the previous build, found it perfectly consistent, and passed.
Everything in that sentence is working as designed, and the result is a lie. The guard was not wrong about the data it read. It was reading data that no longer corresponded to anything.
The fix is the first guard in the chain and it does nothing except ask one question: is the generated output newer than its inputs? It runs before any other check, and it exists solely to make that particular comfortable falsehood impossible.
The general form: a validator that reads the artefact of a process cannot tell you whether the process ran. Those are two different questions and it is very easy to believe you have asked the second when you have only asked the first.
Failure two: the replacement that replaced nothing
Our source file is inconsistent in a way that is nobody's fault and entirely predictable. It is CRLF throughout. It stores em dashes two different ways — sometimes as a real character, sometimes as a six-character HTML escape — and both conventions occur inside the same record.
A string replacement written against one convention matches nothing when it meets the other. In JavaScript, String.replace with no match is not an error. It returns the original string. The script completes, writes the file, and reports success.
So we had edits that were applied, verified and recorded — and never made. The guards then passed, correctly, against data nobody had changed.
The fix is a convention rather than a program: every fix script asserts its own replacement count and exits non-zero if it is not exactly one. Of the 79 fix scripts in the repository, 70 refuse to proceed unless they changed precisely what they expected to change. It is three lines per script and it is the only reason this class of failure gets caught at all.
Failure three: two false greens in a row
This is the one worth dwelling on, because the first mistake was ordinary and the second was not.
We wrote a guard to detect a particular bad pattern. During editing, its regular expression was mangled — an escape sequence collapsed, and the pattern silently became one that could not match anything. The guard ran against a file that definitely contained the fault, and reported clean.
That is failure one all over again, and we would have caught it, except that the test we wrote to catch it also failed silently. It was written in the wrong module system for the project, so it errored on load rather than running. The error was swallowed. It reported success too.
Two independent checks, both green, both meaningless. We found it only by doing something we now do routinely: planting the fault deliberately, confirming it was really in the file, and checking that the guard actually went red.
A test that has never failed has not been tested. That is not our insight — watching a new test go red before trusting it is long-standing unit-testing practice, and there is a small literature arguing that a test which has never failed has not earned its place. What we would add is that it matters more for generated code than for hand-written code, because generated code is syntactically confident in a way that broken hand-written code usually is not, and because the same generation step can produce the check and the test for the check in one pass.
A smaller version of the same thing bit us the same week. A syntax check passed a file whose quoting had been mangled into bare identifiers — valid JavaScript, guaranteed to throw on execution. node --check is a syntax validator and it answered the syntax question correctly. Only running the real build caught it.
The guards we actually run
Eight checks run after every build. None of them fetches a vendor, so none can be rate-limited, blocked or defeated by a page that renders its pricing in JavaScript. They are aimed at our own mistakes, which is where the mistakes mostly are.
| Guard | The question it asks |
|---|---|
| generated-fresh | Is the output newer than its inputs? |
| withdrawn | Is a retracted figure gone from every field, not just the edited one? |
| self-contradiction | Does one field call a product free while another charges for it? |
| arithmetic | Do the listing's own numbers multiply out? |
| price-consistency | Is one tier quoted at two different prices? |
| asset-path | Does every image we reference exist on disk? |
| campaign-pages | Are pages under a 90-day commitment still live and unchanged? |
| public-data | Is anything in the public data export that should not be? |
The second one has a history. A fix that corrects a price and leaves the same claim standing in an FAQ answer has happened six times in this repository. That is why the guard reads the whole record rather than the field somebody just edited — the agent reliably fixes what you pointed at, and reliably does not look two fields over.
The blind spot we found this week
The honest part. On 6 October a reader pointed at an empty card on our blog index. One post was rendering with no cover image.
The asset-path guard had passed. It passes every time. It asks whether every referenced image exists on disk, and every referenced image did exist. The broken post referenced no image at all — the frontmatter field was simply absent, and a guard that walks the fields that are present cannot see a field that is not.
"Every asset resolves" and "every post has an asset" are different questions, and we had only ever asked the first.
That is the permanent condition of this work. Each guard encodes a failure we have already had. None of them anticipates the one we have not had yet, and the gaps are invisible precisely because nothing reports them. Eight guards are not a safety net. They are eight specific lessons, and the ninth is out there being learned by somebody right now.
Structured data helps. It is not sufficient.
The trend brief makes a case we largely agree with: that MCP connectors returning native structured data are more trustworthy than AI-generated summaries, and that "the difference between 'data from Semrush' and 'data the AI made up' is the trust layer teams need."
We would extend it rather than repeat it, because we have a case that complicates it.
One of our connected data sources — a third-party SERP endpoint, not Semrush — intermittently returns results for a different query than the one requested. The query gets truncated to its first token somewhere upstream, so a request for a four-word phrase comes back with results for the first word alone. In one session we measured four of five calls corrupt. The same query succeeded on the second attempt, so it is intermittent rather than deterministic.
Every one of those responses was native structured data from a real API. None of it was invented by a model. And all of it was wrong in a way that no hallucination check would ever flag, because there was no hallucination — just a correct-looking payload answering a question nobody asked.
What caught it was a coherence check we wrote ourselves: does the response actually correspond to the query we sent? That check is three lines long, it is not provided by any vendor, and without it we published one wrong verdict that had to be retracted.
So: prefer structured sources, and still validate them. Structured data removes the fabrication risk. It does not remove the staleness risk, the wrong-query risk, or the silent-no-op risk, and those are the three that actually cost us.
Where Semrush sits in this
The reason a trustworthy source layer matters more in an agentic workflow than in a dashboard is simple: nobody looks at the intermediate steps any more. In a dashboard you see the number in context, next to its date and its filter. In an orchestrated workflow the number arrives already folded into a conclusion, and by then there is nothing left to be suspicious of.
Semrush One ships an MCP integration for exactly this position in the stack. Semrush's home page, checked 7 October 2026, describes it as "The world's most powerful traffic, visibility, and market dataset. Inside your AI assistants", invites you to "Connect Semrush MCP to AI", and states the workflow as "Ask AI. Get Semrush data. Bring your live SEO, market, and brand visibility data directly into the chat." Their Creator Accelerator brief adds the claim that matters for this article: that the MCP returns native structured data rather than AI-generated summaries.

We have not used the Semrush MCP. We are a Semrush affiliate, not a customer, and the description above is of Semrush's documentation rather than a review. What we can say from our own pipeline is the narrower claim: the source layer is the part of an agentic workflow that is worth paying for, because it is the only part a guard cannot reconstruct. We can write a check that catches a stale artefact or a no-op edit. We cannot write a check that invents keyword data we never had.
Semrush's pricing page states "Try Semrush free for seven days. Cancel anytime" — checked 7 October 2026. The screenshot below is from 26 September, when that offer sat on Semrush's home page under the headline "Win in search, era after era". That headline now reads "Be found everywhere search happens" — a fair reminder that a dated capture is evidence for its date and nothing after it. We have not taken the trial, so read the terms as Semrush's claim rather than our test.

Semrush One
In an orchestrated workflow nobody inspects the intermediate steps, so the source layer is the one part a guard you write yourself cannot reconstruct.
See Semrush OneIf you are building one of these
Seven things we would tell ourselves at the start. None requires a tool.
- Make the first check "did anything move?" Before any correctness check, confirm the output is newer than its input. Correctness checks on stale artefacts are the most convincing wrong answer available.
- Assert the count on every edit. A replacement that matched nothing must be an error, never a silent pass. This is three lines and it caught more than anything else we wrote.
- Plant the fault. Before you trust a new guard, break something on purpose and watch it go red. A check that has never failed has not been tested.
- Validate the validator in the same runtime it will run in. Our negative test failed open because it was written in the wrong module system and nobody saw the error.
- Check that the response answers the question you asked. Not whether it is well-formed — whether it is about the right query. This is cheap and almost nobody does it.
- Check the whole record, not the field you edited. The agent fixes what you point at. It does not look two fields over, and the contradiction it leaves behind reads as authoritative.
- Write down which failure each guard encodes. Otherwise someone deletes it in eight months during a tidy-up, because it has never fired.
What we actually learned
The agentic part was the easy part, and it is going to keep getting easier. That is not where the work went.
The work went into a question that has no vendor and no dashboard: how do you know the thing you automated actually happened? Our three failures produced no error, no exception and no implausible output. They produced green results, which is the single most expensive kind of wrong, because green is where you stop looking.
Eight guards and 70 assert-counted scripts is what it took to stop believing our own pipeline. We would guess that number only goes up.
Sources and method
| Figure | Source | Date |
|---|---|---|
| 8 integrity guards in the chain | Our own check:integrity script, read from package.json | 2026-10-07 |
| 79 fix scripts, 70 asserting an exact replacement count | Count across the repository's fix scripts | 2026-10-07 |
| Guard passing against the previous build's output | Our own build and guard records | 2026-09-07 |
| Mangled regex plus a negative test that failed open | Our own engineering log for the campaign-page guard | 2026-09-25 |
| 4 of 5 SERP calls returning a truncated query | Our own measured session against a third-party SERP endpoint | 2026-09-08 |
| Six occasions a fix corrected one field and left an FAQ standing | Our own catalogue fix history | 2026-06 → 2026-09 |
| Blog post with no cover image, undetected by the asset guard | Found by a reader; confirmed against all 33 posts | 2026-10-06 |
| Semrush MCP wording, verbatim | Semrush's own home page | 2026-10-07 |
| Screenshot of the "Ask AI. Get Semrush data." panel | Semrush's home page, captured by us | 2026-10-07 |
| "returns native structured data, not AI-generated summaries" | Semrush's Creator Accelerator brief | retrieved 2026-09-22 |
| "Try Semrush free for seven days. Cancel anytime" | Semrush's pricing page | 2026-10-07 |
| Screenshot wording "Win in search, era after era" (headline since changed) | Semrush's home page, captured by us | 2026-09-26 |
Written by Sohail Akhtar. Every failure described here happened to this repository on the date given, and each one is still encoded as a guard you can read.
Some links may be affiliate links. We may earn a small commission at no extra cost to you.
Related Articles
Explore more guides and reviews from our experts.