AI landing pages

Landing pages from structured data: enriched by AI, not invented by it

A database with thousands of records can be turned into pages that appear in Google Search and in AI answers — but only if every single page answers a real question in full. Google draws the line to bulk content in its own spam policy, and it draws it by intent and usefulness, not by whether an AI was involved. Sharpness Solutions GmbH in Oldenburg, Lower Saxony, builds these page sets from product data, vehicle stock, property lists and spare-parts catalogues: AI for the enrichment, and a shut-off rule for pages that answer nothing. This website is itself derived from structured data, so you can check the claim instead of believing it. Page updated 17 August 2026. Phone +49 441 21 21 63 0, Mon–Fri 9:00–16:00 CET.

Make an enquiry 0441 21 21 63 0 Mo – Fr, 9:00 – 16:00 Uhr

The situation this question comes from

Enquiries on this subject arrive from two opposite directions. Either a quotation is on the table that promises several thousand pages for a flat fee, and somebody in the company has a bad feeling about it. Or there is a well-maintained inventory — a few thousand vehicles, some tens of thousands of spare-part records, a few hundred properties — and the realisation that none of it reaches search, because all of it sits behind a filter form. Both cases end at the same question: where exactly is the line between a useful set of pages and what Google treats as bulk content.

Your inventory sits behind a filter form

The catalogue is complete, the search works, and yet every combination has the same address. What does not exist as its own URL can neither rank nor be quoted. In our vehicle projects the filter state is therefore part of the address — search[make][], search[model][], search[garage][] — which makes result lists linkable, shareable and usable in campaigns. That is the raw material. It is not yet a page.

Somebody is promising you 5,000 pages for a flat fee

The quotation multiplies the combinations: make times model times location times year, and a large number appears in the offer. What is missing is how many of those combinations actually return results from your inventory, and how many attributes are maintained per result. Without that figure, the page count is an arithmetic exercise, not a plan. So we start with the data audit, not with the page counter.

The first attempt is not in the index

A set of pages already exists, it has been rolled out, and Google never took up most of it — status "crawled, currently not indexed", or not fetched at all. That is neither coincidence nor a penalty. It is an assessment: from Google point of view, these pages are not worth the effort. The reason is visible on the pages themselves, not in robots.txt.

You do not know whether AI text harms you

Somebody in the room has read that Google penalises AI content, and somebody else has read that it makes no difference. Both are half right. The policy is written to be method-neutral: it asks about intent and usefulness. But Google does name generative AI explicitly as an example of the abusive case. The difference is not the tool. The difference is whether a reader learns anything at the end.

One page per question, not one page per search term

The difference between a viable set of pages and bulk content fits in one sentence: one page per real question, not one page per search term. A keyword page comes from a word list — swap the town, the make and an adjective, keep the rest. A question page comes from an inventory and answers what somebody wants to know before they call: which tractor units with a Euro 6 engine are available right now, what they cost, which site they are at, what is documented and what is not. That answer is in no word list. It is in your database, or it is nowhere.

In the vehicle trade the raw material is already there, which is why we talk about those projects first. At Autohaus Brüggemann more than 2,000 vehicles sit in one shared search — on 17 August 2026 it was 2,151 — spread across six sites and filterable by 24 makes. Nord Automobile in Rastede runs a four-digit inventory, 1,058 listings on the same day, filtered through its own AJAX endpoint without a page reload. Eschen Nutzfahrzeuge publishes its inventory in five language versions — German, English, Russian, Polish, Spanish — because trading in heavy commercial vehicles is an export business. All three run the same in-house TYPO3 extension for the vehicle search; every improvement benefits all installations. An inventory whose filter states are part of the address is half the distance. The other half is the page itself, and that is work, not configuration.

The real work is the cut. Five attributes produce tens of thousands of combinations on paper, and most of them return zero or two results and represent no question anybody asks. We select the axes where both things are true: there is search demand, and there are enough well-maintained records to give a complete answer. For every axis a lower limit is agreed — a minimum number of results, a minimum coverage of the mandatory fields. Anything below that is not published. That lower limit is not cosmetic. It is the point where this method differs from its abusive twin.

Where AI helps, and where it produces filler

For enrichment a language model is useful, in four places. It turns a list of attributes into a readable paragraph that states the values in sentences instead of a table. It translates equipment codes and manufacturer abbreviations into plain language a buyer understands. It finds synonyms and question variants under which the same thing is searched for. And it brings the text portion into further languages. How small that portion is can be seen in the setup at Eschen Nutzfahrzeuge: first registration, mileage, engine and images are language-neutral and appear in all five language versions without being entered twice; only what is genuinely text gets translated. That remainder is exactly where a model takes work off your hands. In all four cases it works with values that already exist. It creates none.

It cannot invent substance, and the attempt does not end in weak text but in false statements. Where a record holds no consumption figure, a model will add a plausible one — and for new passenger cars German law (Pkw-EnVKV, the consumer information regulation for new cars) requires figures for fuel consumption, CO2 emissions and the CO2 class in advertising. Here, plausible is the opposite of correct. Whether a record falls under that obligation is decided by the source, not by the template, which is why the values are passed through rather than generated. The same holds outside the mandatory scope for accident-free history, service records, number of previous owners and warranties of any kind. Where there is no data, AI does not produce a page. It produces filler with a liability risk. That is the hard limit, and it coincides with the limit in Google policy: both ask whether something is said at the end that was not there before.

Technically this means a clean separation. Mandatory statements and prices go from the source into the output unchanged, with no model in between — from the authoritative system in custom builds, via OpenImmo for property inventories (the XML exchange format of the German real-estate industry) including the mandatory energy-certificate data, via the mobile.de Search API of your own dealer account, otherwise from an ERP export. The AI enrichment runs during preparation, not at request time, so every passage stays reproducible and can be reviewed. Every generated passage is tied to the field values it came from; if the field is missing, no sentence appears — a gap appears on the review list instead. And the enrichment never decides whether a page is published. The state of the data decides that.

The policy in the original — and the numbers that are not in it

Google calls the abusive case "scaled content abuse" and describes it as generating many pages whose primary purpose is to manipulate search rankings rather than to help users — in the original: "generated for the primary purpose of manipulating search rankings". The examples Google itself lists are concrete: using generative AI or comparable tools to produce many pages with no added value for users; scraping other people content and republishing it with minimal change; stitching material together from several sources without creating genuine value; setting up multiple websites to disguise the scale; filling pages with search terms in text that makes no sense. The term was formally defined in the March 2024 spam update.

The second relevant point is called "doorway abuse", and sets of pages fall into it particularly easily. It refers to pages that target very similar search queries and lead users through intermediate steps that are less useful than the actual destination. A page carrying nothing but a heading, three swapped words and a link to the category page is precisely that — regardless of how it was produced. The test is simple and uncomfortable: if a page contains nothing the destination page would not do better, it is an intermediate step. So in the page sets we build, every page gets its own records, its own values and its own answer — or it does not go live.

And here is the sentence worth the most on this page: the policy contains no number you could measure this against. No percentages, no minimum shares, no thresholds. Figures circulate in SEO blogs — "at least 60 per cent different content", "at least three sources per page", "above 30 per cent of affected URLs a sitewide penalty follows". None of those appear anywhere in Google documentation. Anyone arguing with them is quoting themselves, not Google — and selling a threshold you could clear on paper with one percentage point more. There is no threshold. There is the question of whether a page answers something for somebody, and that question cannot be answered in per cent. Only on the page itself.

The second half: being found and being quoted

A set of pages built for Google alone gives away the other half. More and more questions run through ChatGPT, Perplexity, Microsoft Copilot and the AI overviews in Google Search, and there a paragraph is quoted rather than a page delivered. Four requirements for the text follow from that: the answer in the first sentence instead of after three paragraphs of run-up; sections that still hold true when torn out of context; entities in plain text, meaning company name, location and technical term inside the paragraph rather than in the header; and figures instead of adjectives, because only something concrete can be repeated. The same construction also serves classic search. These are not two pages. It is one.

On top comes the machine-readable layer. The right markup per page type — Product and Offer with price and availability, ItemList for result lists, FAQPage for direct answers, HowTo for procedures, Service for offerings, Organization and LocalBusiness for the company facts, BreadcrumbList for placement. Plus an llms.txt as a curated entry file, and before all of that the most banal check of all: do the bots get through at all. GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot and Bingbot can each be controlled separately in robots.txt, and they do different things: OAI-SearchBot, PerplexityBot and Bingbot fetch content for the answer, while GPTBot and ClaudeBot mainly collect training data. Google-Extended appears in the same file but is not a crawler — that token only controls use in the Gemini applications and has no effect on the AI overviews in Google Search, because those draw on the ordinary search index. And often a firewall or a CDN blocks what robots.txt has long permitted. What that means in detail is on our AEO service page.

Whether we can do this can be checked on this website without asking us. Derived from structured data here are 19 regional pages, 19 technology pages, 14 service pages, 5 sector pages, 26 reference cases and ten landing pages including this one — each of them additionally in an English version, so the sitemap carries more than two hundred addresses (as at 17 August 2026). Along with FAQPage, HowTo, Service, ItemList and BreadcrumbList markup, an llms.txt and a dedicated service page on AEO. This is exactly the technique under discussion here — only with substance per page instead of swapped placeholders. It does not prove that it works in your market. It proves that we operate the method rather than only describing it, and one look at the page source confirms it.

What it requires, what goes wrong, and when we advise against it

Three things are prerequisites, and none of them is technology. First, clean and complete data: rubbish times a thousand is still rubbish, only more expensive. Before the first draft we therefore measure field coverage per attribute and the duplicate rate; that tells us how many pages the inventory can carry. Second, an editorial team that reads samples — not every page, but a dozen regularly, and with the right to reject a template. Third, patience: indexing does not happen on request. We roll out new sets in waves, so that uptake can be observed and the template sharpened before the rest follows.

Advising against it is part of this subject, and the reasons are the same three. Data too thin: if each record holds one image and two attributes, no page emerges that answers a question, and the enrichment only papers over it. A range with no search demand: if nobody searches for these combinations, the set is expensive and ineffective — then a good category page is the right answer. No willingness to maintain it: an inventory that is not kept current will, after six months, produce pages about things that no longer exist. In those cases we say so before you place the order. If a project does not pay for itself, hearing that early is cheaper.

Process

Send us an extract from your data.

A CSV with fifty records says more than a requirements list. You get a written assessment: which questions this inventory can answer, how many pages it can carry, where fields are missing, and roughly what the effort amounts to — including the case where the answer is that a good category page is enough and we advise against the project. A dependable figure follows after that, from the data audit. Sharpness Solutions GmbH, Edewechter Landstraße 161, 26131 Oldenburg, Germany. Phone +49 441 21 21 63 0, Mon–Fri 9:00–16:00 CET, info@sharpness.de.

  1. 01

    Examine the data

    First we look into the data, not into search volumes. How many records are there, which attributes are filled and to what degree, how high is the duplicate rate, which mandatory fields are missing. The result is a sheet of figures showing how many pages this inventory can carry at all. Sometimes the project ends here, and that is the cheapest possible moment for it.

  2. 02

    Find questions instead of counting terms

    We collect the questions actually asked in the market — from search queries, from your own shop or site search, from the sales inbox, and from the words customers use on the phone. That produces a list of questions, not of keywords, and after it the check of which ones your inventory can answer in full.

  3. 03

    Set the page cut and the lower limit

    Now it is decided which axes get a page of their own and which stay inside a filter view. Every axis comes with a lower limit: a minimum number of results and a minimum coverage of the mandatory fields. That limit is put in writing before the first template exists, because later it decides every single publication.

  4. 04

    Build the template and test it on twenty pages

    The page template is built in the existing system — TYPO3, Shopware 6 or WordPress — with the direct answer in the first paragraph, mandatory statements in a fixed place, markup per page type, and the AI enrichment in the preparation step. Acceptance runs on about twenty real records, deliberately including the poorly maintained ones. Testing only the showcase cases tests nothing.

  5. 05

    Roll out in waves

    Publication happens in stages, not all at once: a first wave, then the observation of whether and how quickly it is taken up. Pages below the lower limit stay on the editorial review list. In parallel we check bot access in robots.txt, firewall and CDN, because more page sets fail there than on text quality.

  6. 06

    Measure, re-cut, switch off

    After that it is a maintenance task. We look at which pages are indexed, which were only crawled and which deliver nothing — and switch the latter off or merge them. Along with the life cycle of the inventory: delisted items need a redirect, a 410 or a collection page. If you hand the maintenance over, it runs on the service-level tiers: BASIC 24, STANDARD 8, ADVANCED 4, PREMIUM 2 hours response time, as a monthly flat fee net per project.

Frequently asked questions

What are AI-assisted landing pages?

They are pages generated from structured data — product data, vehicle stock, property lists, spare-parts catalogues — each answering one concrete search question in full. The structure, the values and the mandatory statements come from the database; a language model only handles the enrichment: turning attribute lists into readable paragraphs, translating technical abbreviations, finding question variants, bringing the text portion into further languages. The difference from bulk content is the cut: one page per real question, not one page per search term.

Does Google penalise content created with AI?

No, not because of the tool. Google spam policy is written to be method-neutral: it asks about the intent and the usefulness of a page, not about how it was produced. However, Google does name generative AI explicitly as an example of "scaled content abuse" — namely when it is used to produce many pages with no added value for users. So what matters is whether the page says something that helps somebody. A paragraph enriched by AI about an inventory that genuinely exists is unproblematic. An invented paragraph about nothing is not.

What exactly does "scaled content abuse" mean?

Google uses it to describe generating many pages whose primary purpose is to manipulate search rankings rather than to help users. The examples listed: using generative AI or similar tools to produce many pages with no added value; scraping other people content and republishing it with minimal change; stitching material from several sources together without genuine value; setting up multiple websites to disguise the scale; filling pages with search terms in text that makes no sense. The term was formally defined in the March 2024 spam update.

Is it true that at least 60 per cent of the content has to differ?

No. Google spam policy contains no number you could measure this against — no percentages, no minimum shares, no thresholds. Figures such as "at least 60 per cent different content", "at least three sources per page" or "above 30 per cent of affected URLs a sitewide penalty follows" circulate in SEO blogs but appear nowhere in Google documentation. Anyone arguing with them is quoting themselves. The only testable question is whether a page answers something another page does not already answer better — and that is decided case by case, not by a quota.

What is a doorway page, and when does a set of pages become one?

Google describes "doorway abuse" as pages that target very similar search queries and lead users through intermediate steps that are less useful than the actual destination. A page with a swapped heading, three exchanged words and a link to the category page is exactly that. The test: does the page contain anything the destination page would not do better? If not, it is an intermediate step and does not belong online. So in the sets we build, every page gets its own records, its own values and its own answer.

How many pages can be built from an inventory of 2,000 vehicles?

That figure comes out of the data audit, not out of the combinatorics. On paper five attributes yield tens of thousands of combinations, most of which return zero or two results. Only the axes where two conditions hold at once are viable: there is search demand, and there are enough well-maintained records for a complete answer. For every axis we set a lower limit in advance — a minimum number of results, a minimum coverage of the mandatory fields. Anyone who quotes you a page count without looking at the data has multiplied, not checked.

How long does it take for such pages to appear in Google Search?

That cannot be promised, and anyone who promises it does not know either. Indexing does not happen on request: Google decides per URL whether the effort is worth it, and for new page sets that often means "crawled, currently not indexed" at first. So we roll out in waves rather than all at once — the first wave shows how the template is received and can be sharpened before the rest follows. We monitor it in Google Search Console, page type by page type.

What does a set of pages like this cost?

The effort depends on four things: the condition and completeness of the data, the number of page types, the effort to connect the source, and the number of languages. A set built on a well-maintained vehicle inventory with one page type is a different item from a spare-parts catalogue in five languages with an ERP connection. A dependable figure exists after the data audit, not in the first phone call — so you commission the audit as its own separate step and decide about the rest afterwards. We do not sell flat packages by page count, because the page count is the wrong measure.

When do you advise against this approach?

In three cases, and we say so before you place the order. When the data is too thin: if each record holds one image and two attributes, no page emerges that answers a question — the enrichment only papers over it. When a range has no search demand: if nobody searches for these combinations, the set is expensive and ineffective, and a good category page is the right answer. And when there is no willingness to maintain it: an inventory that is not kept current will, after six months, produce pages about things that no longer exist.

How do the pages get into ChatGPT and Perplexity answers?

By making every section quotable on its own. AI assistants rarely take a whole page; they take a paragraph. So the answer goes in the first sentence and the reasoning after it, and company name, location and technical term sit inside the paragraph rather than in the header, because context is lost when a passage is quoted. Along with that: structured markup per page type, an llms.txt as a curated entry point, and verified access in robots.txt, firewall and CDN. OAI-SearchBot, PerplexityBot and Bingbot fetch content for the answer, GPTBot and ClaudeBot mainly collect training data, and Google-Extended only controls use in the Gemini applications, not the AI overviews in Google Search. Sharpness Solutions GmbH gives no guarantee of being cited.

What happens to a page once the vehicle or product is sold?

This is the question on which page sets in volatile inventories fail, and it belongs settled before the build. For individual items: sold is not deleted — depending on the case, a redirect to the appropriate result list, a clean 410 for items permanently gone, or an overview page showing comparable offers. Pages covering a slice of the inventory stay as long as their lower limit is met, and disappear in an orderly way when it is not. What must not happen: hundreds of dead ends in the index, and pages advertising offers that no longer exist.

Do we need a new CMS for this?

Usually not. We build these sets in the existing system — TYPO3, Shopware 6 or WordPress — because the data is already there or connected from there. In our vehicle projects the search runs on an in-house TYPO3 extension that is the same across several installations, from a three-digit commercial-vehicle inventory to more than two thousand passenger cars. The question is not the CMS. The question is whether the inventory is reachable as a data source and whether the filter states can be turned into separate, linkable addresses. We check both in the audit.

Can you take the inventory directly from mobile.de?

Yes, and the direction decides which interface applies. Reading your own inventory runs through the Search API, also called ad integration: for that you enable Listing Integration in your dealer account and create an API username and password; for registered dealers, access to your own inventory is usually covered by the monthly fee. Creating, changing and deleting ads, uploading, assigning and reordering images, or booking paid extra features — that is the Seller API, and it requires an API account enabled through customer service. Alongside these there is the Insights API for analytics, the Lead API for enquiries, and the Ad Stream, which delivers events server-side over WebSocket. All of them are XML-based. The documentation also states the limits: no parallel requests to the same ad, a maximum image count per ad that depends on the account, and no published rate limits — those are settled with support. For a set of pages, mobile.de is usually not the source but one more output channel: the inventory is kept in your own system.

Who maintains the page set after the roll-out?

You decide, and we put it in writing. Templates, enrichment rules and shut-off criteria are part of the delivery, as is the documentation — so a handover to your own team is possible at any time. If we take on the maintenance, it runs on the service-level tiers: BASIC 24 hours, STANDARD 8 hours, ADVANCED 4 hours, PREMIUM 2 hours response time, as a monthly flat fee net per project. Without such an agreement we handle enquiries in order of arrival within 48 hours during business hours, Mon–Fri 9:00–16:00 CET.

Enquiry

What data do you have?

Describe briefly what is in your database and how many records there are. If you are unsure whether the inventory is sufficient: that is exactly what we check first, and if it is not, we say so before quoting.

  • An answer from someone who knows the system — no phone queue
  • An assessment before the quote, even when it advises against the project
  • Your details are sent to us by email, not into a third-party CRM

Spam protection: Cloudflare Turnstile — no cookies, no tracking.

Call Start a project