Discover the Website
Follow crawlable internal links or process a supplied URL list to build an auditable view of the site.
A practical command centre for crawling websites, finding broken links, auditing redirects, reviewing metadata, diagnosing indexability and planning safer migrations. Learn which configuration to use, which reports matter and how to turn crawl data into a prioritised technical SEO audit.
Start with the technical question, configure the smallest useful crawl and follow the evidence back to the affected URL, its source and the underlying template or rule.
The free version currently crawls up to 500 URLs. Advanced configuration, saved crawls, JavaScript rendering, crawl comparison and other professional features require a licence; confirm current access on the official pricing page.
SEO Spider is desktop crawling software for Windows, macOS and Linux. It discovers URLs and their relationships, requests pages and resources, extracts technical signals and organises findings into tabs, filters, issues, reports and bulk exports. It reveals evidence; an SEO professional still needs to interpret impact and recommend the correct fix.
Follow crawlable internal links or process a supplied URL list to build an auditable view of the site.
Record status codes, directives, canonicals, metadata, headings, content, links, images and more.
Use tabs, filters and issue summaries to narrow thousands of URLs into relevant investigation groups.
Export affected URLs with sources and supporting fields for content, development or migration teams.
Start at a URL and follow eligible links through the configured crawl scope.
Upload or paste an exact URL set for redirects, migrations and controlled checks.
Keep crawl information in RAM for suitable smaller or faster local projects.
Use disk-based storage for larger crawls, saved projects and practical continuity.
SEO Spider shows what its configured crawler can discover and request. Combine it with XML sitemaps, Google Search Console, analytics, server logs and manual checks to find orphaned, indexed or valuable URLs outside the normal link graph.
A technically correct crawl begins with scope. Decide which hostname, subdomains, folders, parameters, resources and rendering method belong in the investigation. An oversized crawl wastes time; an undersized crawl can hide the exact issue you need to diagnose.
Use Spider mode to discover a connected website and List mode when you already have the exact URLs to test.
Control whether the crawler stays on one subdomain or folder and whether linked subdomains, sitemaps and external resources are included.
Choose the appropriate crawler identity when investigating bot-specific output, access rules or parity between search engines and users.
Default behaviour should reflect real crawler access. Custom robots testing can support diagnosis but does not change what search engines are permitted to crawl.
Use the fastest method that reveals the necessary content and links; render JavaScript when critical output depends on client-side execution.
Database storage is generally suited to larger saved crawls on an SSD, while memory storage depends more directly on available RAM.
Reduce threads or requests when a server is fragile, rate-limited or shared with live users; schedule intensive work during suitable periods.
Crawl authorised staging or protected environments with the correct credentials while preventing accidental exposure and external indexing.
Store proven settings and crawl data so recurring audits use comparable scope, filters and evidence instead of an improvised setup.
Write down the issue, affected market and evidence needed before configuring the crawler.
Choose hostname, folders, subdomains, resources, user agent and rendering method.
Crawl representative pages and confirm expected content, links and status information.
Complete the crawl, save the project and document the date and configuration.
Response-code reports reveal whether discovered URLs return content, redirect elsewhere, fail or cannot be reached. The strongest repair workflow uses the Inlinks information to identify every source page, then fixes the source rather than merely recording the destination error.
Locate internal or external URLs returning 4xx responses, including missing pages that still receive links from crawlable content.
Identify pages and resources returning server-side failures and investigate whether the problem is persistent, intermittent or crawl-load related.
Review timeouts, DNS failures, connection refusals, malformed URLs and blocked requests that do not produce an HTTP response.
Find internal links that still point to redirected URLs and replace them with the final preferred destination where appropriate.
Review 302 and 307 responses to ensure the temporary status matches the intended long-term behaviour.
Trace URLs through multiple hops and identify the original sources still linking into an unnecessarily long route.
Detect circular routing where a crawler or visitor cannot reach a final content response.
Use the lower Inlinks tab or bulk export to see exactly which pages, elements and anchor text lead to a problematic URL.
Check outgoing citations, partner pages and referenced resources that no longer resolve or now lead somewhere unsuitable.
Open Response Codes and isolate internal 4xx, 5xx, redirects or no-response URLs.
Find every source page, link position, anchor and template creating the journey.
Restore, redirect, replace or remove according to intent, relevance and value.
Confirm direct 200-status journeys and check that no new chain or broken destination remains.
Hundreds of broken links may come from one header, footer, product template or content component. Group URLs by source and pattern before assigning hundreds of individual fixes.
Screaming Frog extracts essential page elements into separate tabs and filters, making it easier to find missing, duplicated or unusual patterns across thousands of URLs. Treat length warnings as prompts for review rather than universal rules: clarity, uniqueness, intent and usefulness matter more than forcing every element into an arbitrary character count.
Find missing, duplicate, unusually short or long title elements and review whether each one identifies a distinct page purpose.
Isolate absent, repeated or problematic descriptions and decide where a stronger search-result summary could improve relevance and clicks.
Review missing, duplicate or multiple H1s in context and confirm the visible main heading accurately introduces the page.
Use extracted H2s to spot repeated template labels, missing topic structure and pages whose subheadings do not support the intended query.
Filter long, non-ASCII, duplicated or parameter-heavy addresses and investigate whether they reflect a wider architecture or faceting problem.
Locate large image files, missing alt attributes and image links that return errors, then review each case according to purpose.
Use custom extraction or search to identify Open Graph and social-card fields on templates where consistent sharing previews matter.
Find linked PDFs, review response behaviour, file size and available document metadata, and decide whether each resource remains useful.
Group metadata and heading problems by directory, content type or shared value to identify the template causing the largest impact.
Choose a page-element tab and isolate a specific missing, duplicate or unusual condition.
Compare directory, page type and repeated values to expose the common source.
Open representative pages and inspect the live search results before rewriting anything.
Correct the template or priority pages, then verify unique and useful output.
The Indexability columns combine signals found during the crawl, but a crawl alone cannot confirm whether Google has indexed a page. Use these reports to diagnose access, directives and canonicalisation, then verify important URLs with Google Search Console and live search evidence.
Separate potentially indexable pages from URLs excluded by response, directives, canonicalisation or crawler access.
Find URLs or resources prevented from crawling and determine whether the restriction is intentional, inherited or outdated.
Review noindex, nofollow and other directives supplied in HTML or HTTP headers, including conflicting implementations.
Locate missing, multiple, conflicting, non-indexable or redirected canonical targets in HTML and supported HTTP headers.
Review whether indexable standalone pages consistently declare themselves where that is part of the site’s canonical strategy.
Compare submitted URLs with crawl data and identify non-indexable, redirected, broken or missing canonical pages inside the sitemap.
Connect XML sitemap, analytics and Search Console sources to find known URLs absent from the normal internal-link crawl.
Measure how many internal-link steps separate pages from the crawl start and identify priority destinations buried too deeply.
Use URL Inspection and indexing reports to check Google-specific canonical selection, rendering and indexing evidence.
A crawler may access a URL that carries noindex, and an indexable URL may still be excluded by Google for canonical, duplication, quality or demand reasons. Report the precise evidence rather than collapsing every problem into “not indexed”.
Content reports help isolate pages that share identical or similar copy, contain unexpectedly little indexable text or fall outside a site’s topical patterns. They do not judge factual accuracy, originality, expertise or commercial usefulness, so every finding needs direct page and search-intent review.
Identify pages whose analysed content matches exactly and investigate parameters, print versions, faceting or repeated publishing.
Use configured similarity analysis to find pages that overlap heavily without being exact copies.
Filter pages below a useful investigation threshold while recognising that tools, contact pages and focused answers may be intentionally concise.
Include the main content and exclude repeated navigation or footer elements so similarity and word-count analysis reflect the page body.
Inspect recurring phrases across pages or groups to understand vocabulary patterns and possible internal-link opportunities.
Use language checks to find likely mistakes at scale, then review names, brands, technical terms and regional usage manually.
Compare similar pages with the live result set and performance data to determine whether they compete or serve distinct needs.
Advanced similarity workflows can identify closely related pages and topical outliers for further content or migration analysis.
Merge competing pages when one destination can satisfy the shared intent, or strengthen unique value when separate pages are justified.
Two pages can share language and still serve different users; two very different pages can compete for the same query. Combine crawl similarity with intent, rankings, clicks, conversions and direct page review.
JavaScript rendering uses an integrated Chromium environment to execute pages before extracting rendered HTML, links and content. It is essential when meaningful output depends on client-side code, but it is slower and more resource intensive than an HTML crawl. Compare both versions and confirm critical findings with Google’s own tools.
Render pages in headless Chromium before crawling the resulting DOM and collecting JavaScript-dependent elements.
Compare the server response with the executed DOM to see whether key content, metadata, canonicals and links are added, removed or changed.
Identify navigational elements that appear after execution and confirm they use crawlable anchor elements with destination URLs.
Review scripts, styles, APIs and other resources the rendering environment cannot request because of robots rules, errors or access controls.
Inspect the rendered view to confirm visible content, consent overlays, lazy loading and interactive components appear as expected.
Adjust timing only when necessary for slow or delayed applications, then verify the setting does not hide genuine performance problems.
Use relevant crawler identities and viewport settings when checking mobile output, responsive behaviour or bot-specific differences.
Confirm important URLs using URL Inspection, Rich Results Test and other first-party evidence when crawl simulations reveal discrepancies.
Where practical, provide essential content, metadata and crawlable navigation in reliable server output or robust rendering patterns.
Record what the server supplies before executing client-side code.
Run JavaScript mode on representative templates with controlled settings.
Review links, text, directives, canonicals, structured data and status behaviour.
Test priority URLs and repair meaningful parity or access gaps.
SEO Spider can extract hreflang annotations from HTML, HTTP headers and XML sitemaps, then report common implementation problems. The audit must consider complete language and regional clusters: a single page can appear correct while its reciprocal target, canonical or status code breaks the relationship.
Collect alternate-language relationships supplied in page markup, headers or XML sitemaps.
Find pages that reference an alternate URL whose cluster does not link back appropriately.
Identify malformed, unsupported or inconsistent language-region values and compare them with the intended audience.
Locate hreflang targets that redirect, fail, are blocked or return no indexable content response.
Review alternates that canonicalise elsewhere or otherwise send contradictory preferred-URL signals.
Check whether a language-selector or fallback page is declared appropriately where the international experience needs one.
Review every language version, return link, canonical, indexability state and destination response together. International SEO fails at the relationship level even when one page appears correctly marked.
A migration can change URLs, templates, navigation, rendering, canonicals and indexability at the same time. Screaming Frog is most useful when it preserves evidence before launch, tests the proposed destination set and compares the new crawl with the old one. The redirect spreadsheet is only one part of the process.
Crawl and save the existing website before development replaces URLs, internal links or page elements.
Test the new environment with authorised access and a controlled scope while keeping it unavailable to public indexing.
Combine crawl, sitemap, analytics, Search Console and backlink data so important URLs are not lost simply because the link crawl missed them.
Upload the legacy URL set and crawl only those addresses to confirm destination, response and redirect path.
Replace links to redirected legacy URLs with direct links to final canonical destinations.
Compare old and new templates to find changes that were not part of the approved migration plan.
Publish new sitemaps containing only canonical, indexable destination URLs and retain the old URL evidence outside the submitted file.
Use crawl comparison and change detection to isolate URLs, fields and issue counts that changed between saved crawls.
Recrawl priority templates and the mapped URL set immediately after DNS, routing and cache changes settle.
Save crawl data and merge all known URL sources before anything changes.
Audit new templates and validate each changed URL in List mode.
Test priority journeys, redirects, internal links, canonicals and indexability.
Watch Search Console, analytics, logs and rankings while defects are still recoverable.
Internal links can be updated quickly, but old bookmarks, citations, backlinks and search results may continue to request legacy URLs for years. Keep useful migration redirects while the old addresses still carry value or receive legitimate traffic.
The standard tabs cover common technical signals. Paid features extend the crawler into a configurable audit workstation: extract page-specific fields, search source code, join external performance data, schedule repeatable crawls and export evidence for reporting. API quotas and third-party costs remain separate from the Screaming Frog licence.
Collect values from HTML using XPath, CSSPath or regular expressions when standard tabs do not contain the field.
Find pages that contain or omit selected words, phrases, tags or code patterns in raw or rendered output.
Append clicks, impressions, position and inspection data to crawl URLs where the connected property and API permit it.
Join relevant traffic, engagement and conversion metrics to crawled landing pages for prioritisation.
Collect available Lighthouse and field-performance metrics through the API and connect them to crawlable templates.
Connect supported link-data providers to combine crawl architecture with page and domain authority metrics.
Connect supported AI providers to run configured prompts or embedding analysis against selected crawl data.
Run recurring crawls with saved settings and automate exports for agreed monitoring workflows.
Export focused issue sets, inlinks, redirect reports and crawl summaries instead of sending an unfiltered crawl dump.
Crawl links, sitemaps, lists and connected sources to assemble the URL estate.
Join response, directives, canonicals, content, rendering and link relationships.
Append search, analytics, speed and link metrics for business-aware prioritisation.
Export the source, evidence, owner, recommended fix and recrawl result.
Repeatable configurations make the crawler more useful than a long list of one-off warnings. Save the scope and exports for each job, record assumptions, and recrawl the affected sources after implementation.
Spider the canonical production scope, connect sitemaps and Search Console, then group problems by template and impact.
Crawl a representative URL list before and after a release to compare headings, canonicals, directives, schema and links.
Combine crawl depth, inlink count, anchor text and Search Console data to locate valuable pages with weak contextual support.
Group exact, near-duplicate and semantically similar pages, then inspect rankings, intent, links and conversion value.
Use List mode to check that published placements resolve, remain indexable, retain the agreed link and point to the correct destination.
Schedule a controlled list crawl of commercial, editorial and campaign URLs and compare it with the previous result.
Before outreach or placement renewal, verify that candidate and live pages are accessible, relevant, indexable and technically stable. Combine that evidence with editorial quality, traffic and topical fit.
The free version is a genuine crawler for small jobs, but it is restricted to 500 URLs per crawl and excludes many configuration, saving, comparison, integration and automation capabilities. A paid SEO Spider licence is currently £199 per user per year; confirm the live price and currency before purchasing.
| Capability | Free Version | Paid Licence | Practical Meaning |
|---|---|---|---|
| Crawl size | Up to 500 URLs per crawl | No software crawl limit* | Large crawls still depend on machine resources, storage, scope and server tolerance. |
| Core audits | Broken links, metadata, directives, duplicates, hreflang, sitemaps and visualisations | Included | The free version is useful for small sites and controlled URL samples. |
| Configuration | Restricted | Advanced crawl configuration | Paid access is better for precise scope, rendering, extraction and repeatability. |
| Saving and comparison | No saved-crawl workflow | Save, open, compare and detect changes | Essential for migrations, regression testing and recurring audits. |
| JavaScript rendering | Restricted | Included | Needed when important content or links depend on client-side rendering. |
| Custom analysis | Restricted | Custom search, extraction and JavaScript | Useful for site-specific templates, QA and data collection. |
| APIs and automation | Restricted | Integrations, scheduling and reporting | Useful for joining business data and running consistent monitoring. |
| Licence model | Free | Annual licence per user | One licensed user can install on multiple personally used devices, subject to the licence terms. |
Use it for compact websites, samples, quick broken-link checks and learning the core interface.
Use it when crawl size, saved projects, JavaScript, comparison, custom extraction, integrations or scheduling matter.
Licences are assigned per person, with bulk discounts listed for larger purchases on the official pricing page.
The paid software removes the product’s 500-URL cap, but crawl capacity still depends on available memory, database storage, configuration, website size, rendering method and the server being crawled. Define scope before adding hardware.
These products overlap, but they are not interchangeable. Screaming Frog is a highly configurable desktop crawler; Sitebulb emphasises visual technical reporting; Semrush and Ahrefs provide cloud crawlers inside broader research platforms. The strongest choice depends on workflow, scale, collaboration and the data already used by your team.
| Tool | Operating Model | Strongest Use Case | Main Trade-Off |
|---|---|---|---|
| Screaming Frog | Desktop crawler for Windows, macOS and Linux | Deep configuration, migrations, controlled URL lists, extraction and technical investigation | Capacity uses local resources and collaboration needs an agreed file/report workflow. |
| Sitebulb | Desktop and cloud technical auditing options | Visual explanations, prioritised hints and stakeholder-friendly technical reporting | Teams should compare crawl flexibility, limits and reporting needs directly. |
| Semrush Site Audit | Cloud crawler inside a broad marketing platform | Recurring project monitoring connected to keyword, competitor and reporting tools | Less suited to some highly customised desktop-crawl and extraction workflows. |
| Ahrefs Site Audit | Cloud crawler inside an SEO research platform | Technical monitoring alongside backlink and organic-search research | Best value is realised when the wider Ahrefs platform is already part of the workflow. |
Best when you need exact crawl configuration, local data, List mode, migration validation, custom extraction or raw technical evidence.
Best when clear visualisation, audit hints and accessible stakeholder explanations are central to delivery.
Best when cloud project monitoring should sit beside keyword, competitor, backlink and reporting data in one paid platform.
It is especially strong for technical consultants, migrations, QA, URL-list testing and custom extraction. Choose a cloud suite when shared monitoring and connected research are more important, or combine the products when each solves a different stage of the workflow.
These answers clarify the practical limits of the crawler and where additional data or professional judgement is still required.
Yes. The free version can crawl up to 500 URLs per crawl and includes many core checks. Advanced configuration, saved crawls, comparison, integrations and other professional features require a paid licence.
Discovered internal HTML pages and eligible resources can consume the crawl allowance according to configuration. A website with fewer than 500 visible pages may still reach the limit because URLs are not the same as pages.
SEO Spider is installed desktop software for Windows, macOS and Linux. It can connect to cloud APIs and export reports, but the crawl itself runs using the configured machine and storage.
Yes, the licensed version supports JavaScript rendering. Use rendering only where necessary, compare source with rendered output and provide enough time and resources for critical content to appear.
Yes, when you connect other URL sources such as XML sitemaps, analytics and Search Console. A link crawl alone cannot discover a page with no crawlable route from the starting scope.
No crawler can reproduce Google’s complete index. SEO Spider can identify potentially indexable URLs and connect relevant Search Console data, but important cases still need Google-specific verification.
Yes, with suitable scope, database storage, hardware and server-friendly speed settings. Start with representative samples and priority sections before assuming every URL must be crawled in one run.
It is a legitimate professional crawler, but any crawler can create unnecessary server load when configured aggressively. Crawl only authorised sites, set a considerate speed and coordinate large production audits.
No. It discovers and organises evidence. A person still needs to validate the issue, understand business impact, choose the correct repair and confirm the result with a recrawl and real performance data.
GuestPost.UK can combine Screaming Frog evidence with Search Console, analytics, content, architecture and backlink review, then prioritise fixes by risk, reach and commercial value.
Your trusted partner for Guest Posting, Link Building and AI-powered SEO. We help businesses grow authority, rankings and organic traffic with ethical, results-driven strategies.
© 2026. All rights reserved.
WhatsApp us