One page per question, not one page per search term
The difference between a viable set of pages and bulk content fits in one sentence: one page per real question, not one page per search term. A keyword page comes from a word list — swap the town, the make and an adjective, keep the rest. A question page comes from an inventory and answers what somebody wants to know before they call: which tractor units with a Euro 6 engine are available right now, what they cost, which site they are at, what is documented and what is not. That answer is in no word list. It is in your database, or it is nowhere.
In the vehicle trade the raw material is already there, which is why we talk about those projects first. At Autohaus Brüggemann more than 2,000 vehicles sit in one shared search — on 17 August 2026 it was 2,151 — spread across six sites and filterable by 24 makes. Nord Automobile in Rastede runs a four-digit inventory, 1,058 listings on the same day, filtered through its own AJAX endpoint without a page reload. Eschen Nutzfahrzeuge publishes its inventory in five language versions — German, English, Russian, Polish, Spanish — because trading in heavy commercial vehicles is an export business. All three run the same in-house TYPO3 extension for the vehicle search; every improvement benefits all installations. An inventory whose filter states are part of the address is half the distance. The other half is the page itself, and that is work, not configuration.
The real work is the cut. Five attributes produce tens of thousands of combinations on paper, and most of them return zero or two results and represent no question anybody asks. We select the axes where both things are true: there is search demand, and there are enough well-maintained records to give a complete answer. For every axis a lower limit is agreed — a minimum number of results, a minimum coverage of the mandatory fields. Anything below that is not published. That lower limit is not cosmetic. It is the point where this method differs from its abusive twin.
Where AI helps, and where it produces filler
For enrichment a language model is useful, in four places. It turns a list of attributes into a readable paragraph that states the values in sentences instead of a table. It translates equipment codes and manufacturer abbreviations into plain language a buyer understands. It finds synonyms and question variants under which the same thing is searched for. And it brings the text portion into further languages. How small that portion is can be seen in the setup at Eschen Nutzfahrzeuge: first registration, mileage, engine and images are language-neutral and appear in all five language versions without being entered twice; only what is genuinely text gets translated. That remainder is exactly where a model takes work off your hands. In all four cases it works with values that already exist. It creates none.
It cannot invent substance, and the attempt does not end in weak text but in false statements. Where a record holds no consumption figure, a model will add a plausible one — and for new passenger cars German law (Pkw-EnVKV, the consumer information regulation for new cars) requires figures for fuel consumption, CO2 emissions and the CO2 class in advertising. Here, plausible is the opposite of correct. Whether a record falls under that obligation is decided by the source, not by the template, which is why the values are passed through rather than generated. The same holds outside the mandatory scope for accident-free history, service records, number of previous owners and warranties of any kind. Where there is no data, AI does not produce a page. It produces filler with a liability risk. That is the hard limit, and it coincides with the limit in Google policy: both ask whether something is said at the end that was not there before.
Technically this means a clean separation. Mandatory statements and prices go from the source into the output unchanged, with no model in between — from the authoritative system in custom builds, via OpenImmo for property inventories (the XML exchange format of the German real-estate industry) including the mandatory energy-certificate data, via the mobile.de Search API of your own dealer account, otherwise from an ERP export. The AI enrichment runs during preparation, not at request time, so every passage stays reproducible and can be reviewed. Every generated passage is tied to the field values it came from; if the field is missing, no sentence appears — a gap appears on the review list instead. And the enrichment never decides whether a page is published. The state of the data decides that.
The policy in the original — and the numbers that are not in it
Google calls the abusive case "scaled content abuse" and describes it as generating many pages whose primary purpose is to manipulate search rankings rather than to help users — in the original: "generated for the primary purpose of manipulating search rankings". The examples Google itself lists are concrete: using generative AI or comparable tools to produce many pages with no added value for users; scraping other people content and republishing it with minimal change; stitching material together from several sources without creating genuine value; setting up multiple websites to disguise the scale; filling pages with search terms in text that makes no sense. The term was formally defined in the March 2024 spam update.
The second relevant point is called "doorway abuse", and sets of pages fall into it particularly easily. It refers to pages that target very similar search queries and lead users through intermediate steps that are less useful than the actual destination. A page carrying nothing but a heading, three swapped words and a link to the category page is precisely that — regardless of how it was produced. The test is simple and uncomfortable: if a page contains nothing the destination page would not do better, it is an intermediate step. So in the page sets we build, every page gets its own records, its own values and its own answer — or it does not go live.
And here is the sentence worth the most on this page: the policy contains no number you could measure this against. No percentages, no minimum shares, no thresholds. Figures circulate in SEO blogs — "at least 60 per cent different content", "at least three sources per page", "above 30 per cent of affected URLs a sitewide penalty follows". None of those appear anywhere in Google documentation. Anyone arguing with them is quoting themselves, not Google — and selling a threshold you could clear on paper with one percentage point more. There is no threshold. There is the question of whether a page answers something for somebody, and that question cannot be answered in per cent. Only on the page itself.
The second half: being found and being quoted
A set of pages built for Google alone gives away the other half. More and more questions run through ChatGPT, Perplexity, Microsoft Copilot and the AI overviews in Google Search, and there a paragraph is quoted rather than a page delivered. Four requirements for the text follow from that: the answer in the first sentence instead of after three paragraphs of run-up; sections that still hold true when torn out of context; entities in plain text, meaning company name, location and technical term inside the paragraph rather than in the header; and figures instead of adjectives, because only something concrete can be repeated. The same construction also serves classic search. These are not two pages. It is one.
On top comes the machine-readable layer. The right markup per page type — Product and Offer with price and availability, ItemList for result lists, FAQPage for direct answers, HowTo for procedures, Service for offerings, Organization and LocalBusiness for the company facts, BreadcrumbList for placement. Plus an llms.txt as a curated entry file, and before all of that the most banal check of all: do the bots get through at all. GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot and Bingbot can each be controlled separately in robots.txt, and they do different things: OAI-SearchBot, PerplexityBot and Bingbot fetch content for the answer, while GPTBot and ClaudeBot mainly collect training data. Google-Extended appears in the same file but is not a crawler — that token only controls use in the Gemini applications and has no effect on the AI overviews in Google Search, because those draw on the ordinary search index. And often a firewall or a CDN blocks what robots.txt has long permitted. What that means in detail is on our AEO service page.
Whether we can do this can be checked on this website without asking us. Derived from structured data here are 19 regional pages, 19 technology pages, 14 service pages, 5 sector pages, 26 reference cases and ten landing pages including this one — each of them additionally in an English version, so the sitemap carries more than two hundred addresses (as at 17 August 2026). Along with FAQPage, HowTo, Service, ItemList and BreadcrumbList markup, an llms.txt and a dedicated service page on AEO. This is exactly the technique under discussion here — only with substance per page instead of swapped placeholders. It does not prove that it works in your market. It proves that we operate the method rather than only describing it, and one look at the page source confirms it.
What it requires, what goes wrong, and when we advise against it
Three things are prerequisites, and none of them is technology. First, clean and complete data: rubbish times a thousand is still rubbish, only more expensive. Before the first draft we therefore measure field coverage per attribute and the duplicate rate; that tells us how many pages the inventory can carry. Second, an editorial team that reads samples — not every page, but a dozen regularly, and with the right to reject a template. Third, patience: indexing does not happen on request. We roll out new sets in waves, so that uptake can be observed and the template sharpened before the rest follows.
Advising against it is part of this subject, and the reasons are the same three. Data too thin: if each record holds one image and two attributes, no page emerges that answers a question, and the enrichment only papers over it. A range with no search demand: if nobody searches for these combinations, the set is expensive and ineffective — then a good category page is the right answer. No willingness to maintain it: an inventory that is not kept current will, after six months, produce pages about things that no longer exist. In those cases we say so before you place the order. If a project does not pay for itself, hearing that early is cheaper.