Guard the boundary
Confirm source context, inspect assets, and rank sources before extraction.
Modules 1, 2, 3, and 30Upstream semantic evidence
Use 34 modules to guard scope, inspect source assets, extract candidates, classify query intent, read result patterns, compare competitors, rank opportunities, package accepted findings, and route the output into the correct next workflow.
Starting route
Each route names the smallest useful starting module, the recommended sequence, and the point where discovery must stop.
Source Context Check, then Discovery Asset Review when files are present.
Stop conditionStop until audience, offer, allowed topics, blocked topics, protected pages, and the next workflow are clear.
Discovery Asset Review, Source Candidate Ranking, then the smallest relevant extraction module.
Stop conditionHold sources with weak quality, stale scope, unclear labels, or noise risk.
Corpus Scan, NER Pass, Concept Harvest, Frequency Signal Scan, and Placement Signal Scan.
Stop conditionDo not turn the first scan into final entity structure or a topical map.
Query Intent Classification, Query Modifier Scan, Semantic Query Clustering, and Query Treatment Selection.
Stop conditionDo not create pages until page, section, question, merge, anchor, and reject treatments are reviewed.
SERP Entity Harvest, Competitor Entity Harvest, Competitive Coverage Snapshot, SERP Consensus Scan, and SERP Divergence Scan.
Stop conditionUse competitor evidence to discover patterns. Do not copy wording or structure.
Candidate Weighting, Discovery Opportunity Matrix, Entity Universe Package, then Discovery Handoff.
Stop conditionDo not send blocked, uncertain, or unsupported findings into production.
Five phase workflow
The sequence prevents loose keyword, competitor, and source evidence from becoming structure before it has been qualified.
Confirm source context, inspect assets, and rank sources before extraction.
Modules 1, 2, 3, and 30Collect named entities, concepts, signals, modifiers, attributes, and facets, then classify and weight them.
Modules 4 to 12Classify intent, scan modifiers, select treatment, generate and classify queries, cluster meaning, and expand the network.
Modules 13 to 20Harvest entities and patterns, compare coverage, separate consensus from divergence, and test adjacent concepts.
Modules 21 to 29Prioritize opportunities, package the entity universe, identify page type seeds, and create a clean handoff.
Modules 31 to 3434 discovery modules
Every record keeps the purpose, short command, full prompt, return fields, stop rule, and possible routes together.
Confirm the site purpose, audience, offer, allowed topic lanes, blocked topic lanes, protected pages, target workflow, and discovery boundary.
Run Source Context Check on this discovery task.
Review each file, export, URL list, report, sitemap, crawl, search file, analytics file, behavior note, competitor source, or earlier MIRENA output before extraction.
Run Discovery Asset Review on these files.
Scan a page set, draft set, or content export for recurring terms, concepts, phrases, themes, weak signals, overused language, and candidate areas before deeper extraction.
Run Corpus Scan on this content set.
Extract named candidate entities such as people, organizations, products, software, brands, places, known concepts, frameworks, documents, standards, and tools.
Run NER Pass on this asset.
Extract important process terms, decision concepts, category terms, user states, intent concepts, technical ideas, and support concepts that may not appear as named entities.
Run Concept Harvest on this asset.
Normalize mixed candidates into clear types such as person, organization, brand, software, product, feature, location, process, framework, document, category, metric, modifier, format, support concept, or reject.
Run Entity Type Classification on this candidate list.
Weight each candidate by frequency, placement, source quality, source count, intent fit, topic fit, buyer fit, workflow fit, and downstream usefulness.
Run Candidate Weighting on this discovery list.
Measure repeated terms, concepts, phrases, modifiers, entity candidates, and page themes without treating repetition as proof of importance.
Run Frequency Signal Scan on this content set.
Check where candidates appear across titles, headings, openings, body sections, tables, questions, captions, anchors, schema notes, and action sections.
Run Placement Signal Scan on this asset.
Extract modifiers that change audience, product fit, feature angle, location, comparison need, process stage, price sensitivity, problem state, buyer stage, format, or page type.
Run Modifier Harvest on this keyword set.
Collect candidate features, qualities, constraints, use cases, benefits, limitations, categories, criteria, specifications, proof points, and descriptive phrases.
Run Attribute Candidate Harvest on this corpus.
Identify feature, audience, location, comparison, price, problem, process, trust, format, and urgency facets that change user need or page treatment.
Run Facet Intent Extraction on this query set.
Classify each query by primary intent, secondary intent, user stage, likely page type, answer treatment, result format, and source context fit.
Run Query Intent Classification on this query set.
Scan query wording for modifiers that change page type, intent, format, audience, product stage, local need, comparison need, trust need, or next step.
Run Query Modifier Scan on this query set.
Decide whether each query needs a dedicated page, section, question, list, comparison, table, anchor target, link target, template, example, merge, or rejection.
Run Query Treatment Selection on this query list.
Generate possible future queries from accepted candidates, modifiers, user stages, page types, product angles, audience needs, location signals, feature signals, and comparison paths.
Run Synthetic Query Generation on this candidate and modifier set.
Review generated queries for intent, confidence, source context fit, usefulness, risk, and downstream value.
Run Synthetic Query Classification on this generated query list.
Group queries by meaning, intent layer, user job, page type, modifier pattern, and shared concept rather than repeated words alone.
Run Semantic Query Clustering on this query set.
Find hidden user needs implied by the topic, query set, result set, competitors, modifiers, product context, or user stage.
Run Latent Intent Discovery on this topic.
Expand a seed topic into query paths, intent branches, user stages, support questions, comparison paths, feature paths, process paths, and adjacent topics.
Run Query Network Expansion on this seed topic.
Extract repeated candidate entities, dominant terms, attributes, concepts, headings, table topics, question topics, comparison angles, and visible entity signals from top ranking sources.
Run SERP Entity Harvest on these result set competitors.
Review the result set for dominant page types, content formats, repeated sections, result features, answer formats, tables, questions, comparisons, local results, product blocks, and missing angles.
Run SERP Pattern Intake on this query.
Extract candidate entities, attributes, concepts, page formats, proof points, comparison angles, product references, question themes, and support topics from competitor sources.
Run Competitor Entity Harvest on these competitor URLs.
Summarize common coverage, overused coverage, missing coverage, repeated formats, proof patterns, and seeds for useful differentiation.
Run Competitive Coverage Snapshot on this result set.
Identify repeated concepts, claims, expected sections, page formats, examples, definitions, question topics, comparison angles, and answer patterns across the result set.
Run SERP Consensus Scan on this query group.
Find where top pages differ in intent, format, audience, page type, depth, angle, action path, proof, result feature focus, and topic scope.
Run SERP Divergence Scan on this query group.
Review current site coverage, competitor depth, related support areas, repeated entities, query branches, and missing support areas before page planning.
Run Topical Authority Baseline on this topic.
Collect visible structured data types, entity fields, attribute patterns, breadcrumb patterns, sameAs cues, and repeated data fields from competitor sources.
Run Competitive Schema Scan on these competitor pages.
Expand a seed topic into nearby concepts, adjacent categories, related query paths, support concepts, comparisons, processes, features, and explicit rejection candidates.
Run Semantic Neighborhood Expansion on this seed topic.
Rank sources by quality, relevance, freshness risk, evidence strength, extraction value, topic fit, intent fit, and risk of adding noise.
Run Source Candidate Ranking on this discovery set.
Turn accepted candidates, rejected candidates, query paths, modifier groups, result patterns, competitor findings, source notes, and risks into prioritized opportunities.
Run Discovery Opportunity Matrix on this raw discovery output.
Package accepted candidates, rejected candidates, review candidates, query paths, modifier groups, candidate attributes, source notes, result signals, competitor signals, and handoff notes.
Run Entity Universe Package on this discovery output.
Identify signals for likely page archetypes, user jobs, page roles, and downstream routes from intent, modifiers, result patterns, competitor formats, existing pages, user stages, and commercial goals.
Run Page Archetype Seed Discovery on this discovery set.
Route accepted findings, rejected findings, review items, query clusters, modifier groups, result findings, competitor findings, source notes, opportunities, and blocked items into the correct next workflow.
Run Discovery Handoff on this raw discovery output.
Discovery handoff
A useful discovery output names the source, confidence, fit, risk, owner, and next route for every finding.
Accepted entities, concepts, attributes, modifiers, query paths, and result signals with their source evidence.
Repeated terms, adjacent concepts, query branches, sources, and competitor patterns that do not support the approved scope.
Items that need stronger evidence, result validation, source context clarification, or human review.
Impact, effort, confidence, risk, reason, and the recommended workflow for each accepted opportunity.
Protected pages, privacy risks, unsupported claims, schema timing issues, and topics that must not move downstream.
Mapping, briefs, rewriting, entity review, links, information gain, result features, schema cues, hold, review, or reject.
Questions
Raw Semantic Discovery is the upstream workflow that collects and qualifies candidate entities, concepts, modifiers, query paths, intent signals, result patterns, competitor signals, source signals, and opportunity notes before later workflows decide structure or production.
Start with Source Context Check when the project boundary is not fully approved. Use Discovery Asset Review for many files, Corpus Scan for a page set, Query Intent Classification for keyword evidence, or SERP Entity Harvest for competitor evidence.
No. Raw Semantic Discovery collects and qualifies candidates. Entity SEO and Salience later organize identity, attributes, relationships, placement, support, and salience.
No. Discovery gathers signals and rejects noise. Topical mapping turns approved signals into page ownership, hierarchy, roles, routes, overlap controls, and build order.
Yes. Use Query Intent Classification, Query Modifier Scan, Semantic Query Clustering, and Query Treatment Selection before any query becomes a page decision.
Yes. Use competitor and result set modules to collect expected concepts, repeated coverage, missing angles, proof patterns, consensus, and divergence. Competitor evidence should guide discovery rather than dictate wording or structure.
Yes, after Evidence Intake confirms quality, privacy, date range, and interpretation limits. Analytics and behavior evidence can support source ranking, opportunity review, rewrite intake, and route decisions.
Run Discovery Opportunity Matrix, Entity Universe Package, and Discovery Handoff. Route accepted findings into mapping, briefs, rewriting, entity review, internal links, information gain, result feature planning, schema cues after approval, or a hold state.
Next route
Use the Docs library for evidence intake, workflow routing, topical mapping, briefs, rewrites, entity review, information gain, internal links, result formats, and schema cues.
Founder access is €20 per 30 days excluding VAT for one seat and one active MIRENA instance. OpenAI account rules and usage limits remain separate.