404, 410 and Soft 404: Handling Removed Pages Correctly
Do 404s harm rankings, when should you use 410, why soft 404s are dangerous, and should a removed page be redi...
Read
Semantic keyword research is not the collection of individual phrases into a list — it is mapping every question, intent and concept inside a topic. The output is not a 500-row spreadsheet but a structural plan showing which page answers which group of questions.
The difference is easy to state. Classic research answers "how many times a month is this searched?". Semantic research adds a second layer — what does the person asking actually want to know, what do they search before and after, and which entity is the question attached to? The theory behind this sits in the semantic SEO and entity SEO articles.
The output of ordinary keyword research usually looks like this: "seo service", "seo services", "seo service price", "professional seo service", "cheap seo service". Five rows — and most of them are the same entity and very nearly the same intent.
Working from that list creates three problems:
Google's documentation on AI features makes this more pressing: the system uses a query fan-out technique, issuing several related searches behind a single query. A site that covers a topic broadly and deeply matches more of those additional searches.
Before processing a query you have to establish its intent. In practice four types are enough:
There is a simple way to establish intent: search the query on Google and look at the format of the results. If the first ten results are blog posts, Google treats it as an informational query — bringing a sales page to that fight will almost never work.
Stage 1 — The central entity. Write it in one sentence: "This cluster is about technical SEO." The more precise the entity, the easier every following step becomes.
Stage 2 — The entity network. List 20 to 40 concepts connected to it. For technical SEO: robots.txt, sitemaps, canonical tags, indexing, crawl budget, Core Web Vitals, JavaScript rendering, hreflang, 301 redirects, status codes, structured data, mobile usability.
Stage 3 — Harvesting questions. Collect real questions for each concept. Five templates work in almost any topic: what is it · how is it done · why is it needed · which tool checks it · how long or how much.
Stage 4 — Intent sorting. Label every question with one of the four types above.
Stage 5 — The page plan. Group questions that share an intent and require the same answer, then assign each group to one page. One page = one question group.
Paid tools speed things up but are not required. The most valuable data is free:
Search Console → Performance report. Per Google's documentation the report supplies four metrics — clicks, impressions, CTR and average position — and allows filtering across six dimensions: query, page, country, device, search appearance and date. It shows which queries your site already appears for, which is the best possible starting point for finding gaps.
The same document notes two limitations: search results vary "by time, place, device, and recent history of the person searching", so a query in the report may not show your site when you run it yourself; and the newest data is preliminary and may still change.
The search results themselves. The "People also ask" block, autocomplete suggestions and related searches at the foot of the page reflect Google's own understanding of the query space directly.
Customer questions. Sales threads, phone calls, social media comments. This source is almost always overlooked, yet it carries the most accurate wording. Our own SEO questions section is built from exactly this kind of collected question.
This stage turns research into a plan. A simple table is enough:
| Question group | Intent | Page type | Target URL |
|---|---|---|---|
| what technical SEO is, why it matters | Informational | Pillar guide | blog / pillar |
| how to optimise crawl budget | Informational | Supporting post | blog / supporting |
| what a technical SEO audit costs | Transactional | Service page | /en/services/seo-audit |
| choosing a technical SEO agency | Commercial | Service + proof | /en/services/texniki-seo |
| brand name + contact | Navigational | Contact | /en/contact |
The "Target URL" column doubles as your internal linking plan: blog posts should point to the relevant service pages, and service pages should point back to the blog posts that go deeper.
A simplified map of a real topic looks like this:
| Sub-entity | Typical questions | Intent |
|---|---|---|
| Indexing | "why is my page not indexed", "what does crawled - currently not indexed mean" | Informational |
| Speed | "how to speed up a website", "what are Core Web Vitals" | Informational |
| robots.txt | "how to write robots.txt", "how to test robots.txt" | Informational |
| Sitemaps | "how to create sitemap.xml", "how to submit a sitemap" | Informational |
| Hreflang | "hreflang for multilingual sites", "hreflang errors" | Informational |
| Audit | "what is a technical SEO audit", "audit pricing" | Transactional |
Every row is a potential page — but they should not all be written at once. The rule of priority is simple: start with the intent closest to the business goal, then work outward into informational queries.
Cannibalisation is when several of your own pages compete for the same query. Google cannot settle on which to show, and both perform worse.
How to spot it: in Search Console open Performance, filter by the query, then switch to the Pages tab. If two or three URLs appear for one query, you have a problem.
How to fix it:
Once the queries are collected, the biggest risk is stuffing them into the text artificially. Google's spam policies name keyword stuffing explicitly: "repeating the same words or phrases so often that it sounds unnatural".
The right approach:
It matters, but it is not the only criterion. A query searched twenty times a month that leads directly to a sale can be worth more than a generic one searched two thousand times. Use volume to prioritise, not to select.
There is no fixed number. The test is this: if the questions share an intent and a reader would want to read them in sequence on one page, keep them together. If the reader feels "this is now a different topic", split it.
Yes, and query fan-out makes them more important: when the system runs related searches, pages with specific, precise answers gain more chances to match.
There is no norm and there should not be one. If the text reads naturally, it is enough. Instead of counting density, check coverage: have all the subquestions of the topic been answered?
For a mid-sized topic a full map (entity network + questions + page plan) takes an experienced specialist one to two working days. The longest part is not collecting the questions but distributing them across pages correctly.
Do 404s harm rankings, when should you use 410, why soft 404s are dangerous, and should a removed page be redi...
Read
The difference between 301, 302, 307 and 308, Google's strong versus weak signal distinction, redirect methods...
Read
What HTTP status codes are, what the five families mean and how Google interprets them: indexing, canonical si...
Read