AI Search Optimisation
Programmatic SEOthat is not thin content
Programmatic SEO works when you have a genuine dataset and each generated page answers a real question better than anything else available. It fails when it is a template with variables swapped in. The difference is entirely in the data and the quality gate, and that is where I spend the time.
- Data first
- Quality gates
- Built to survive an update
Sound familiar
How programmatic goes wrong
We published forty thousand pages and traffic went down.
Most of them have never been crawled.
They are all the same page with a different city name.
Programmatic SEO only works with real data
Programmatic SEO means generating many pages from a structured data source. It works when each page answers a specific question with information that is genuinely useful and hard to find elsewhere. It fails, expensively, when the pages are a template with a variable substituted.
The honest test is simple: if a person landed on one of these pages having searched for exactly that thing, would they be pleased? If the answer is no, generating fifty thousand of them makes your site worse, not bigger.
The data source is the whole project
Most of the work happens before any page exists. You need a dataset where each row supports a page a real person would want: pricing that varies by configuration, comparisons between specific things, availability by location, specifications, aggregated statistics nobody else has assembled.
If your dataset is a list of city names, you do not have a programmatic project. You have a template and a find-and-replace, and search engines have been unimpressed by that for a decade.
Search demand has a shape
Before building, I check the demand curve. Sometimes there are ten thousand viable page-level queries. Sometimes there are four hundred, and generating ten thousand pages means nine thousand six hundred that nobody will ever search for, diluting crawl budget and site quality signals.
Knowing that number early changes the scope, and it is a cheap thing to find out.
How I build a programmatic system
Quality gates before generation
Every page must pass thresholds before it is allowed to exist. Minimum data completeness: if half the fields are empty, the page is not generated. Minimum distinctness: if it is more than a set percentage similar to an existing page, it is not generated. Minimum demand: if nothing suggests anyone searches for this, it does not get a page.
Pages failing the gate either do not exist or are consolidated into a parent page covering the group. That is the mechanism that keeps a programmatic build from becoming a liability.
Templates with real variation
A good template does not fill blanks. It varies structure based on the data. A page with rich data shows comparison tables and detailed breakdowns; a page with sparse data shows a shorter, honest layout rather than padding to hit a word count.
Each page should also contain at least one element that exists only there: a calculated figure, a genuine comparison, an aggregate. That is the thing that makes it worth indexing and, increasingly, worth citing.
Internal linking that is generated too
Ten thousand orphaned pages are ten thousand pages nobody will crawl. Internal linking has to be part of the system: hub pages grouping by meaningful dimensions, related links driven by data similarity rather than randomness, and breadcrumbs that reflect a real hierarchy.
Staged rollout
I never publish the whole set at once. A few hundred pages first, then wait for indexation and engagement data. If those pages are indexed and used, expand. If they are ignored, the quality gate needs tightening, and finding that out at five hundred pages is much cheaper than at fifty thousand.
Monitoring afterwards
Programmatic sets need ongoing attention: how many are indexed, how many receive traffic, which templates underperform, and whether the underlying data has gone stale. Stale data is the silent killer, because a page confidently stating something that stopped being true damages trust more than not having the page.
When I will tell you not to do this
If your data is thin, if search demand does not exist at the page level, if you cannot commit to keeping the data current, or if the site has unresolved technical problems. Adding thousands of pages to a site with crawl issues makes the crawl issues worse, and that is the most common way this ends badly.
What you get
Included in every programmatic build
-
Data source assessment
An honest read on whether your dataset supports pages worth publishing, and how many, before anything is built.
-
Demand curve analysis
How many page-level queries genuinely exist, so the scope matches reality rather than the size of the spreadsheet.
-
Quality gates in the generator
Completeness, distinctness and demand thresholds enforced in code. Pages that fail are consolidated or never created.
-
Generated internal linking
Hubs, data-driven related links and real breadcrumbs, so pages are discoverable rather than orphaned.
The process
Small first, always
-
Assess the data and the demand
Whether the dataset supports useful pages, and how many queries actually exist. This sets the real scope.
-
Build with gates
Template, quality thresholds and generated internal linking, tested on a few hundred pages.
-
Publish in stages
A batch, then indexation and engagement data, then expand or tighten the gate. Never all at once.
Questions
Programmatic SEO questions
Is programmatic SEO risky?
Publishing thousands of near-identical pages is risky. Publishing pages that each answer a real question with real data is not, because that is just a website. The risk lives entirely in the quality gate, which is why I put it in the generator rather than in a review process nobody has time for.
How many pages should we generate?
As many as there is genuine demand and genuine data for, which is usually far fewer than the spreadsheet suggests. I check the demand curve before we scope, and I would rather ship eight hundred good pages than eighty thousand that dilute the site.
Can we use AI to write the pages?
For assembling and phrasing data you already hold, yes, carefully. For inventing content to pad a thin page, no. That is exactly the failure mode this whole approach is designed to avoid, and it is now cheap enough that everyone is doing it, which makes it worthless.
What if the data changes?
The generator reruns and pages update. That is a feature. The risk is stale data nobody notices, so I include freshness monitoring that flags data which has not updated when it should have. A page confidently stating something that stopped being true is worse than no page.
Next step
Tell me what data you have
The dataset decides whether this works. Describe what you hold and I will tell you honestly whether it supports pages worth publishing, and roughly how many.
- Reply within 24h
- I will tell you if the answer is no