GuestPost.UK Technical SEO Playbook

Screaming Frog SEO Spider Guide: Technical SEO Audits, Crawling & Site Migrations

A practical command centre for crawling websites, finding broken links, auditing redirects, reviewing metadata, diagnosing indexability and planning safer migrations. Learn which configuration to use, which reports matter and how to turn crawl data into a prioritised technical SEO audit.

500 URLs free300+ issue checksMigration workflowsGuestPost.UK methods
Technical audit map

What do you need to investigate?

Start with the technical question, configure the smallest useful crawl and follow the evidence back to the affected URL, its source and the underlying template or rule.

300+Issues checked

The free version currently crawls up to 500 URLs. Advanced configuration, saved crawls, JavaScript rendering, crawl comparison and other professional features require a licence; confirm current access on the official pricing page.

Crawler overview

What Screaming Frog SEO Spider Actually Does

SEO Spider is desktop crawling software for Windows, macOS and Linux. It discovers URLs and their relationships, requests pages and resources, extracts technical signals and organises findings into tabs, filters, issues, reports and bulk exports. It reveals evidence; an SEO professional still needs to interpret impact and recommend the correct fix.

CRAWL

Discover the Website

Follow crawlable internal links or process a supplied URL list to build an auditable view of the site.

STORE

Collect Technical Signals

Record status codes, directives, canonicals, metadata, headings, content, links, images and more.

FILTER

Isolate the Problem

Use tabs, filters and issue summaries to narrow thousands of URLs into relevant investigation groups.

EXPORT

Build the Action List

Export affected URLs with sources and supporting fields for content, development or migration teams.

Spider Mode

Start at a URL and follow eligible links through the configured crawl scope.

List Mode

Upload or paste an exact URL set for redirects, migrations and controlled checks.

Memory Storage

Keep crawl information in RAM for suitable smaller or faster local projects.

Database Storage

Use disk-based storage for larger crawls, saved projects and practical continuity.

NOTE
A crawler does not see the whole of Google

SEO Spider shows what its configured crawler can discover and request. Combine it with XML sitemaps, Google Search Console, analytics, server logs and manual checks to find orphaned, indexed or valuable URLs outside the normal link graph.

Crawl configuration

Configure the Crawl Before You Press Start

A technically correct crawl begins with scope. Decide which hostname, subdomains, folders, parameters, resources and rendering method belong in the investigation. An oversized crawl wastes time; an undersized crawl can hide the exact issue you need to diagnose.

MODEStarting Point

Choose Spider or List Mode

Use Spider mode to discover a connected website and List mode when you already have the exact URLs to test.

Best for: Matching discovery behaviour to a full audit, redirect check, migration list or targeted validation.
SCOPEBoundaries

Set the Crawl Scope

Control whether the crawler stays on one subdomain or folder and whether linked subdomains, sitemaps and external resources are included.

Best for: Avoiding incomplete audits and preventing unrelated environments from consuming the crawl.
AGENTCrawler Identity

Select the User Agent

Choose the appropriate crawler identity when investigating bot-specific output, access rules or parity between search engines and users.

Best for: Reproducing crawler behaviour without assuming every user agent receives identical HTML.
ROBOTSAccess

Respect or Test Robots Rules

Default behaviour should reflect real crawler access. Custom robots testing can support diagnosis but does not change what search engines are permitted to crawl.

Best for: Separating discovery problems from robots.txt restrictions and safe staging tests.
RENDERJavaScript

Choose HTML or JavaScript Rendering

Use the fastest method that reveals the necessary content and links; render JavaScript when critical output depends on client-side execution.

Best for: Comparing source and rendered experiences on React, Vue, Angular and other dynamic websites.
STORECapacity

Choose the Storage Mode

Database storage is generally suited to larger saved crawls on an SSD, while memory storage depends more directly on available RAM.

Best for: Preventing a large crawl from failing because the machine and storage configuration were overlooked.
SPEEDCourtesy

Control Crawl Speed

Reduce threads or requests when a server is fragile, rate-limited or shared with live users; schedule intensive work during suitable periods.

Best for: Collecting evidence without creating avoidable load or triggering security controls.
AUTHPrivate Sites

Handle Authentication Carefully

Crawl authorised staging or protected environments with the correct credentials while preventing accidental exposure and external indexing.

Best for: Pre-launch audits where the live website cannot yet reveal the new templates and URLs.
SAVERepeatability

Save the Configuration

Store proven settings and crawl data so recurring audits use comparable scope, filters and evidence instead of an improvised setup.

Best for: Monthly monitoring, team handovers and defensible before-and-after comparisons.
01Define the question

Write down the issue, affected market and evidence needed before configuring the crawler.

02Set boundaries

Choose hostname, folders, subdomains, resources, user agent and rendering method.

03Test a small sample

Crawl representative pages and confirm expected content, links and status information.

04Run and preserve

Complete the crawl, save the project and document the date and configuration.

Response codes and redirects

Find Broken Links, Server Errors and Redirect Waste

Response-code reports reveal whether discovered URLs return content, redirect elsewhere, fail or cannot be reached. The strongest repair workflow uses the Inlinks information to identify every source page, then fixes the source rather than merely recording the destination error.

4XXBroken

Client Errors

Locate internal or external URLs returning 4xx responses, including missing pages that still receive links from crawlable content.

Action: Restore, redirect or remove the source link according to the page’s purpose and available replacement.
5XXServer

Server Errors

Identify pages and resources returning server-side failures and investigate whether the problem is persistent, intermittent or crawl-load related.

Action: Confirm with server monitoring and logs, repair the cause and retest under realistic load.
NONEConnection

No Response URLs

Review timeouts, DNS failures, connection refusals, malformed URLs and blocked requests that do not produce an HTTP response.

Action: Retest representative examples and distinguish invalid links from network, firewall or rate-limit behaviour.
301Permanent

Permanent Redirects

Find internal links that still point to redirected URLs and replace them with the final preferred destination where appropriate.

Action: Link directly to the canonical destination and preserve redirects needed for users and old external links.
302Temporary

Temporary Redirects

Review 302 and 307 responses to ensure the temporary status matches the intended long-term behaviour.

Action: Keep genuinely temporary routing; change permanent moves to an appropriate permanent response.
CHAINEfficiency

Redirect Chains

Trace URLs through multiple hops and identify the original sources still linking into an unnecessarily long route.

Action: Update internal links and consolidate safe redirects to the final relevant destination.
LOOPCritical

Redirect Loops

Detect circular routing where a crawler or visitor cannot reach a final content response.

Action: Correct the rule order or mapping, then test every affected entry URL and final destination.
INLINKSources

Inlinks and Source URLs

Use the lower Inlinks tab or bulk export to see exactly which pages, elements and anchor text lead to a problematic URL.

Action: Repair the component or content source so the problem is removed across every affected page.
EXTExternal

Broken External Links

Check outgoing citations, partner pages and referenced resources that no longer resolve or now lead somewhere unsuitable.

Action: Replace with a strong current source, update the destination or remove the link when it no longer helps readers.
01Filter the status

Open Response Codes and isolate internal 4xx, 5xx, redirects or no-response URLs.

02Open the inlinks

Find every source page, link position, anchor and template creating the journey.

03Choose the correct outcome

Restore, redirect, replace or remove according to intent, relevance and value.

04Recrawl the sources

Confirm direct 200-status journeys and check that no new chain or broken destination remains.

AUDIT
Fix causes, not spreadsheet rows

Hundreds of broken links may come from one header, footer, product template or content component. Group URLs by source and pattern before assigning hundreds of individual fixes.

Get a Technical SEO Audit →
Page elements

Audit Titles, Meta Descriptions, Headings and Images at Scale

Screaming Frog extracts essential page elements into separate tabs and filters, making it easier to find missing, duplicated or unusual patterns across thousands of URLs. Treat length warnings as prompts for review rather than universal rules: clarity, uniqueness, intent and usefulness matter more than forcing every element into an arbitrary character count.

TITLESearch Display

Page Titles

Find missing, duplicate, unusually short or long title elements and review whether each one identifies a distinct page purpose.

Best for: Detecting template repetition and pages competing with weak or indistinguishable titles.
DESCSERP Copy

Meta Descriptions

Isolate absent, repeated or problematic descriptions and decide where a stronger search-result summary could improve relevance and clicks.

Best for: Prioritising important landing pages rather than bulk-writing generic descriptions for every crawlable URL.
H1Main Heading

H1 Headings

Review missing, duplicate or multiple H1s in context and confirm the visible main heading accurately introduces the page.

Best for: Finding broken templates and vague headings—not enforcing one technical pattern without inspecting the design.
H2Structure

H2 Headings

Use extracted H2s to spot repeated template labels, missing topic structure and pages whose subheadings do not support the intended query.

Best for: Reviewing large content types before improving individual pages manually.
URLAddress

URL Structure

Filter long, non-ASCII, duplicated or parameter-heavy addresses and investigate whether they reflect a wider architecture or faceting problem.

Best for: Finding inconsistent patterns while preserving useful existing URLs unless a migration is justified.
IMGMedia

Image Size and Alt Text

Locate large image files, missing alt attributes and image links that return errors, then review each case according to purpose.

Best for: Improving accessibility, load efficiency and image understanding without stuffing keywords into alternative text.
SOCIALSharing

Social Metadata

Use custom extraction or search to identify Open Graph and social-card fields on templates where consistent sharing previews matter.

Best for: Checking editorial and campaign pages before promotion across social platforms.
PDFDocuments

PDF Auditing

Find linked PDFs, review response behaviour, file size and available document metadata, and decide whether each resource remains useful.

Best for: Managing guides, reports and downloads that are part of the searchable website estate.
PATTERNTemplates

Pattern-Based Prioritisation

Group metadata and heading problems by directory, content type or shared value to identify the template causing the largest impact.

Best for: Turning thousands of warnings into a few scalable fixes for content and development teams.
01Filter one issue

Choose a page-element tab and isolate a specific missing, duplicate or unusual condition.

02Group by template

Compare directory, page type and repeated values to expose the common source.

03Review intent

Open representative pages and inspect the live search results before rewriting anything.

04Fix and recrawl

Correct the template or priority pages, then verify unique and useful output.

Crawlability and indexability

Understand Why a URL Can or Cannot Enter Search

The Indexability columns combine signals found during the crawl, but a crawl alone cannot confirm whether Google has indexed a page. Use these reports to diagnose access, directives and canonicalisation, then verify important URLs with Google Search Console and live search evidence.

INDEXSummary

Indexability Status

Separate potentially indexable pages from URLs excluded by response, directives, canonicalisation or crawler access.

Best for: Building an investigation queue while remembering that “indexable” does not mean “indexed”.
ROBOTAccess

Robots.txt Blocks

Find URLs or resources prevented from crawling and determine whether the restriction is intentional, inherited or outdated.

Best for: Diagnosing discovery and rendering gaps without confusing crawl blocking with a noindex directive.
NOIDXDirective

Meta Robots and X-Robots-Tag

Review noindex, nofollow and other directives supplied in HTML or HTTP headers, including conflicting implementations.

Best for: Finding important pages unintentionally excluded and low-value pages incorrectly left open.
CANPreferred URL

Canonical Elements

Locate missing, multiple, conflicting, non-indexable or redirected canonical targets in HTML and supported HTTP headers.

Best for: Confirming that duplicate and parameterised pages consistently identify a valid preferred version.
SELFConsistency

Self-Referencing Canonicals

Review whether indexable standalone pages consistently declare themselves where that is part of the site’s canonical strategy.

Best for: Detecting omissions across templates without assuming every non-self-canonical URL is wrong.
SITEMAPDiscovery

XML Sitemap Auditing

Compare submitted URLs with crawl data and identify non-indexable, redirected, broken or missing canonical pages inside the sitemap.

Best for: Keeping sitemaps focused on important, canonical and indexable destination URLs.
ORPHANArchitecture

Orphan Pages

Connect XML sitemap, analytics and Search Console sources to find known URLs absent from the normal internal-link crawl.

Best for: Finding valuable or legacy pages that users and crawlers cannot reach through the website structure.
DEPTHClicks

Crawl Depth

Measure how many internal-link steps separate pages from the crawl start and identify priority destinations buried too deeply.

Best for: Improving access and hierarchy while considering alternate entry points and contextual relevance.
GSCVerification

Search Console Confirmation

Use URL Inspection and indexing reports to check Google-specific canonical selection, rendering and indexing evidence.

Best for: Confirming crawler findings on priority URLs before declaring the issue fixed or unresolved.
RULE
Accessible, indexable and indexed are different states

A crawler may access a URL that carries noindex, and an indexable URL may still be excluded by Google for canonical, duplication, quality or demand reasons. Report the precise evidence rather than collapsing every problem into “not indexed”.

Content quality signals

Find Duplicate, Thin and Overlapping Pages

Content reports help isolate pages that share identical or similar copy, contain unexpectedly little indexable text or fall outside a site’s topical patterns. They do not judge factual accuracy, originality, expertise or commercial usefulness, so every finding needs direct page and search-intent review.

EXACTDuplicates

Exact Duplicate Pages

Identify pages whose analysed content matches exactly and investigate parameters, print versions, faceting or repeated publishing.

Best for: Choosing consolidation, canonicalisation, redirects or intentional differentiation according to purpose.
NEARSimilarity

Near-Duplicate Content

Use configured similarity analysis to find pages that overlap heavily without being exact copies.

Best for: Detecting location, product or topic pages differentiated only by a few substituted words.
LOWWord Count

Low-Content Pages

Filter pages below a useful investigation threshold while recognising that tools, contact pages and focused answers may be intentionally concise.

Best for: Finding unexpected template failures and weak pages—not enforcing one minimum word count site-wide.
AREAExtraction

Content Area Configuration

Include the main content and exclude repeated navigation or footer elements so similarity and word-count analysis reflect the page body.

Best for: Preventing shared templates from distorting duplicate-content and language checks.
GRAMLanguage

N-Gram Analysis

Inspect recurring phrases across pages or groups to understand vocabulary patterns and possible internal-link opportunities.

Best for: Finding repeated language and unlinked topic mentions without turning frequency into a keyword-density target.
SPELLEditorial

Spelling and Grammar

Use language checks to find likely mistakes at scale, then review names, brands, technical terms and regional usage manually.

Best for: Supporting editorial quality control without automatically replacing valid British English or specialist language.
INTENTSERP Review

Search Intent Overlap

Compare similar pages with the live result set and performance data to determine whether they compete or serve distinct needs.

Best for: Separating real cannibalisation from multiple useful pages that target different tasks or funnel stages.
EMBEDSemantic

Semantic Similarity

Advanced similarity workflows can identify closely related pages and topical outliers for further content or migration analysis.

Best for: Prioritising human review across large sites, not allowing an automated similarity score to make publishing decisions.
ACTIONDecision

Consolidate or Differentiate

Merge competing pages when one destination can satisfy the shared intent, or strengthen unique value when separate pages are justified.

Best for: Turning similarity findings into a clearer content architecture and internal-link plan.
CONTENT
Similarity is evidence, not the verdict

Two pages can share language and still serve different users; two very different pages can compete for the same query. Combine crawl similarity with intent, rankings, clicks, conversions and direct page review.

Improve SEO Content →
Rendered websites

Crawl JavaScript Without Mistaking the Simulation for Google

JavaScript rendering uses an integrated Chromium environment to execute pages before extracting rendered HTML, links and content. It is essential when meaningful output depends on client-side code, but it is slower and more resource intensive than an HTML crawl. Compare both versions and confirm critical findings with Google’s own tools.

JSRendering

JavaScript Rendering Mode

Render pages in headless Chromium before crawling the resulting DOM and collecting JavaScript-dependent elements.

Best for: React, Vue, Angular, single-page applications and pages whose source HTML is incomplete.
RAWSource

Raw vs Rendered HTML

Compare the server response with the executed DOM to see whether key content, metadata, canonicals and links are added, removed or changed.

Best for: Finding parity gaps that make crawling and indexing depend entirely on rendering.
LINKDiscovery

JavaScript-Only Links

Identify navigational elements that appear after execution and confirm they use crawlable anchor elements with destination URLs.

Best for: Finding critical pages hidden behind click handlers, buttons or unsupported link patterns.
BLOCKResources

Blocked Resources

Review scripts, styles, APIs and other resources the rendering environment cannot request because of robots rules, errors or access controls.

Best for: Diagnosing incomplete rendering while avoiding unnecessary crawling of every third-party resource.
SHOTVisual QA

Rendered Page and Screenshots

Inspect the rendered view to confirm visible content, consent overlays, lazy loading and interactive components appear as expected.

Best for: Connecting extracted-data differences with what the rendering browser actually displayed.
TIMETiming

AJAX Timeout

Adjust timing only when necessary for slow or delayed applications, then verify the setting does not hide genuine performance problems.

Best for: Pages that have not completed meaningful rendering within the configured wait period.
UAParity

User Agent and Viewport

Use relevant crawler identities and viewport settings when checking mobile output, responsive behaviour or bot-specific differences.

Best for: Controlled comparisons without assuming one desktop render represents every search context.
GOOGLEConfirmation

Google Rendering Checks

Confirm important URLs using URL Inspection, Rich Results Test and other first-party evidence when crawl simulations reveal discrepancies.

Best for: Distinguishing a local rendering configuration issue from a problem Google actually encounters.
FIXResilience

Reduce Rendering Dependency

Where practical, provide essential content, metadata and crawlable navigation in reliable server output or robust rendering patterns.

Best for: Making important SEO signals available consistently to users and multiple search crawlers.
01Crawl the HTML

Record what the server supplies before executing client-side code.

02Render the sample

Run JavaScript mode on representative templates with controlled settings.

03Compare key signals

Review links, text, directives, canonicals, structured data and status behaviour.

04Confirm with Google

Test priority URLs and repair meaningful parity or access gaps.

International SEO

Audit Hreflang and International Page Relationships

SEO Spider can extract hreflang annotations from HTML, HTTP headers and XML sitemaps, then report common implementation problems. The audit must consider complete language and regional clusters: a single page can appear correct while its reciprocal target, canonical or status code breaks the relationship.

LANGAnnotations

Hreflang Discovery

Collect alternate-language relationships supplied in page markup, headers or XML sitemaps.

Best for: Building a site-wide view of declared language and regional alternatives.
RETURNReciprocal

Missing Return Links

Find pages that reference an alternate URL whose cluster does not link back appropriately.

Best for: Repairing incomplete relationships that search engines may ignore.
CODEValidation

Invalid Language Codes

Identify malformed, unsupported or inconsistent language-region values and compare them with the intended audience.

Best for: Correcting syntax without using country codes alone as language identifiers.
200Destinations

Broken or Redirected Alternates

Locate hreflang targets that redirect, fail, are blocked or return no indexable content response.

Best for: Ensuring annotations resolve directly to live, accessible destination pages.
CANCanonical

Canonical Conflicts

Review alternates that canonicalise elsewhere or otherwise send contradictory preferred-URL signals.

Best for: Aligning international clusters with the site’s indexable canonical strategy.
XDEFFallback

X-Default Review

Check whether a language-selector or fallback page is declared appropriately where the international experience needs one.

Best for: Guiding users whose language or region is not represented by a more specific alternate.
INTL
Validate the complete cluster, not one annotation

Review every language version, return link, canonical, indexability state and destination response together. International SEO fails at the relationship level even when one page appears correctly marked.

Open the Hreflang SEO Guide →
Migration control

Protect Rankings During a Website Migration

A migration can change URLs, templates, navigation, rendering, canonicals and indexability at the same time. Screaming Frog is most useful when it preserves evidence before launch, tests the proposed destination set and compares the new crawl with the old one. The redirect spreadsheet is only one part of the process.

BASEBefore Launch

Preserve the Live-Site Baseline

Crawl and save the existing website before development replaces URLs, internal links or page elements.

Check: Status, indexability, titles, headings, canonicals, directives, depth, inlinks and structured data.
STAGEPre-Launch

Crawl the Staging Website

Test the new environment with authorised access and a controlled scope while keeping it unavailable to public indexing.

Check: New templates, navigation, mobile output, rendered content and unintended noindex or robots rules.
MAPURL Mapping

Map Every Valuable Old URL

Combine crawl, sitemap, analytics, Search Console and backlink data so important URLs are not lost simply because the link crawl missed them.

Action: Map each changed URL to the closest useful new destination, not automatically to the homepage.
LISTControlled Test

Validate Redirects in List Mode

Upload the legacy URL set and crawl only those addresses to confirm destination, response and redirect path.

Check: Direct permanent redirects, relevant final URLs, no loops, no chains and no soft-404 destinations.
LINKInternal Routes

Update Internal Links

Replace links to redirected legacy URLs with direct links to final canonical destinations.

Check: Navigation, breadcrumbs, body links, image links, hreflang, canonicals and structured-data URLs.
PARITYSignal Match

Check Technical Parity

Compare old and new templates to find changes that were not part of the approved migration plan.

Check: Content, metadata, canonical targets, directives, language annotations, schema and indexable URL counts.
XMLDiscovery

Replace XML Sitemaps Cleanly

Publish new sitemaps containing only canonical, indexable destination URLs and retain the old URL evidence outside the submitted file.

Check: No redirected, blocked, broken, duplicate or staging URLs remain in submitted sitemaps.
DIFFComparison

Compare Old and New Crawls

Use crawl comparison and change detection to isolate URLs, fields and issue counts that changed between saved crawls.

Check: New, missing and altered URLs rather than manually comparing two unrelated spreadsheets.
LIVELaunch Day

Run a Post-Launch Control Crawl

Recrawl priority templates and the mapped URL set immediately after DNS, routing and cache changes settle.

Check: Server responses, final HTML, rendering, analytics tags, conversion journeys and accidental staging references.
01Capture the old estate

Save crawl data and merge all known URL sources before anything changes.

02Test staging and mappings

Audit new templates and validate each changed URL in List mode.

03Launch and recrawl

Test priority journeys, redirects, internal links, canonicals and indexability.

04Monitor real outcomes

Watch Search Console, analytics, logs and rankings while defects are still recoverable.

RULE
Do not remove redirects because a crawl looks clean

Internal links can be updated quickly, but old bookmarks, citations, backlinks and search results may continue to request legacy URLs for years. Keep useful migration redirects while the old addresses still carry value or receive legitimate traffic.

Advanced analysis

Extract Custom Data, Connect APIs and Automate Repeat Audits

The standard tabs cover common technical signals. Paid features extend the crawler into a configurable audit workstation: extract page-specific fields, search source code, join external performance data, schedule repeatable crawls and export evidence for reporting. API quotas and third-party costs remain separate from the Screaming Frog licence.

XPATHExtraction

Custom Extraction

Collect values from HTML using XPath, CSSPath or regular expressions when standard tabs do not contain the field.

Use it for: Product details, author names, dates, breadcrumbs, stock status, schema fields and template markers.
FINDSource Search

Custom Search

Find pages that contain or omit selected words, phrases, tags or code patterns in raw or rendered output.

Use it for: Analytics tags, old branding, legal copy, internal-link mentions and template QA.
GSCSearch Data

Google Search Console

Append clicks, impressions, position and inspection data to crawl URLs where the connected property and API permit it.

Use it for: Prioritising pages with search demand and finding known URLs outside the internal-link crawl.
GABehaviour

Google Analytics

Join relevant traffic, engagement and conversion metrics to crawled landing pages for prioritisation.

Use it for: Finding valuable orphan pages and separating high-impact defects from unused URL noise.
PSIPerformance

PageSpeed Insights

Collect available Lighthouse and field-performance metrics through the API and connect them to crawlable templates.

Use it for: Grouping performance issues by page type while checking lab and field evidence separately.
LINKAuthority Data

Link-Metric Integrations

Connect supported link-data providers to combine crawl architecture with page and domain authority metrics.

Use it for: Prioritising broken URLs, redirects and orphan pages that have meaningful external links.
AIPrompt Data

AI Prompts and Embeddings

Connect supported AI providers to run configured prompts or embedding analysis against selected crawl data.

Use it for: Classification, language, intent, entities, alt-text assistance and semantic similarity—with manual review.
TIMEAutomation

Scheduling and Command Line

Run recurring crawls with saved settings and automate exports for agreed monitoring workflows.

Use it for: Weekly health checks, release QA, client reporting and large crawls outside business hours.
REPORTDelivery

Bulk Exports and Looker Studio

Export focused issue sets, inlinks, redirect reports and crawl summaries instead of sending an unfiltered crawl dump.

Use it for: Giving developers the source URL, failing URL, evidence and expected outcome needed to act.
Discovery data

Crawl links, sitemaps, lists and connected sources to assemble the URL estate.

Technical data

Join response, directives, canonicals, content, rendering and link relationships.

Performance data

Append search, analytics, speed and link metrics for business-aware prioritisation.

Action data

Export the source, evidence, owner, recommended fix and recrawl result.

Practical playbooks

Six Screaming Frog Workflows Worth Saving

Repeatable configurations make the crawler more useful than a long list of one-off warnings. Save the scope and exports for each job, record assumptions, and recrawl the affected sources after implementation.

01Technical Audit

Full Website Health Crawl

Spider the canonical production scope, connect sitemaps and Search Console, then group problems by template and impact.

Output: An evidence-backed priority list, not an exported catalogue of every available warning.
02Release QA

Template Change Validation

Crawl a representative URL list before and after a release to compare headings, canonicals, directives, schema and links.

Output: A focused change report that separates intentional updates from regression defects.
03Internal Links

Find Pages Needing More Support

Combine crawl depth, inlink count, anchor text and Search Console data to locate valuable pages with weak contextual support.

Output: Relevant source pages and natural anchor opportunities for each priority destination.
04Content Estate

Consolidate Duplicate Topics

Group exact, near-duplicate and semantically similar pages, then inspect rankings, intent, links and conversion value.

Output: A keep, improve, merge, redirect or remove decision supported by evidence.
05Link Campaigns

Audit Guest-Post Placement Pages

Use List mode to check that published placements resolve, remain indexable, retain the agreed link and point to the correct destination.

Output: A placement-quality report covering status, directives, canonical, anchor, rel attribute and destination.
06Monitoring

Recurring Priority-URL Watch

Schedule a controlled list crawl of commercial, editorial and campaign URLs and compare it with the previous result.

Output: Fast alerts for new errors, redirects, noindex, canonical changes, missing content or lost links.
GP
Use crawl data to qualify link-building work

Before outreach or placement renewal, verify that candidate and live pages are accessible, relevant, indexable and technically stable. Combine that evidence with editorial quality, traffic and topical fit.

Explore Link-Building Services →
Free and paid access

Screaming Frog Pricing and Licence Choice

The free version is a genuine crawler for small jobs, but it is restricted to 500 URLs per crawl and excludes many configuration, saving, comparison, integration and automation capabilities. A paid SEO Spider licence is currently £199 per user per year; confirm the live price and currency before purchasing.

CapabilityFree VersionPaid LicencePractical Meaning
Crawl sizeUp to 500 URLs per crawlNo software crawl limit*Large crawls still depend on machine resources, storage, scope and server tolerance.
Core auditsBroken links, metadata, directives, duplicates, hreflang, sitemaps and visualisationsIncludedThe free version is useful for small sites and controlled URL samples.
ConfigurationRestrictedAdvanced crawl configurationPaid access is better for precise scope, rendering, extraction and repeatability.
Saving and comparisonNo saved-crawl workflowSave, open, compare and detect changesEssential for migrations, regression testing and recurring audits.
JavaScript renderingRestrictedIncludedNeeded when important content or links depend on client-side rendering.
Custom analysisRestrictedCustom search, extraction and JavaScriptUseful for site-specific templates, QA and data collection.
APIs and automationRestrictedIntegrations, scheduling and reportingUseful for joining business data and running consistent monitoring.
Licence modelFreeAnnual licence per userOne licensed user can install on multiple personally used devices, subject to the licence terms.
FREEStart Here

Choose Free for Small, Focused Crawls

Use it for compact websites, samples, quick broken-link checks and learning the core interface.

Good fit: A crawl that remains meaningful within 500 discovered URLs.
PAIDProfessional

Choose Paid for Serious Audit Work

Use it when crawl size, saved projects, JavaScript, comparison, custom extraction, integrations or scheduling matter.

Good fit: Consultants, in-house teams, agencies and repeat technical SEO workflows.
TEAMLicensing

Count Users, Not Shared Computers

Licences are assigned per person, with bulk discounts listed for larger purchases on the official pricing page.

Good fit: Teams budgeting access for every person who will operate licensed features.
*
“Unlimited” is not infinite

The paid software removes the product’s 500-URL cap, but crawl capacity still depends on available memory, database storage, configuration, website size, rendering method and the server being crawled. Define scope before adding hardware.

Crawler comparison

Screaming Frog vs Sitebulb, Semrush and Ahrefs

These products overlap, but they are not interchangeable. Screaming Frog is a highly configurable desktop crawler; Sitebulb emphasises visual technical reporting; Semrush and Ahrefs provide cloud crawlers inside broader research platforms. The strongest choice depends on workflow, scale, collaboration and the data already used by your team.

ToolOperating ModelStrongest Use CaseMain Trade-Off
Screaming FrogDesktop crawler for Windows, macOS and LinuxDeep configuration, migrations, controlled URL lists, extraction and technical investigationCapacity uses local resources and collaboration needs an agreed file/report workflow.
SitebulbDesktop and cloud technical auditing optionsVisual explanations, prioritised hints and stakeholder-friendly technical reportingTeams should compare crawl flexibility, limits and reporting needs directly.
Semrush Site AuditCloud crawler inside a broad marketing platformRecurring project monitoring connected to keyword, competitor and reporting toolsLess suited to some highly customised desktop-crawl and extraction workflows.
Ahrefs Site AuditCloud crawler inside an SEO research platformTechnical monitoring alongside backlink and organic-search researchBest value is realised when the wider Ahrefs platform is already part of the workflow.
SFBest Control

Choose Screaming Frog

Best when you need exact crawl configuration, local data, List mode, migration validation, custom extraction or raw technical evidence.

SBBest Visuals

Choose Sitebulb

Best when clear visualisation, audit hints and accessible stakeholder explanations are central to delivery.

VERDICT
Final verdict: use Screaming Frog when investigation control matters most

It is especially strong for technical consultants, migrations, QA, URL-list testing and custom extraction. Choose a cloud suite when shared monitoring and connected research are more important, or combine the products when each solves a different stage of the workflow.

Common questions

Screaming Frog SEO Spider FAQs

These answers clarify the practical limits of the crawler and where additional data or professional judgement is still required.

Is Screaming Frog SEO Spider free?

Yes. The free version can crawl up to 500 URLs per crawl and includes many core checks. Advanced configuration, saved crawls, comparison, integrations and other professional features require a paid licence.

What counts towards the 500-URL limit?

Discovered internal HTML pages and eligible resources can consume the crawl allowance according to configuration. A website with fewer than 500 visible pages may still reach the limit because URLs are not the same as pages.

Is Screaming Frog cloud-based?

SEO Spider is installed desktop software for Windows, macOS and Linux. It can connect to cloud APIs and export reports, but the crawl itself runs using the configured machine and storage.

Can it crawl JavaScript websites?

Yes, the licensed version supports JavaScript rendering. Use rendering only where necessary, compare source with rendered output and provide enough time and resources for critical content to appear.

Can Screaming Frog find orphan pages?

Yes, when you connect other URL sources such as XML sitemaps, analytics and Search Console. A link crawl alone cannot discover a page with no crawlable route from the starting scope.

Does it show what Google has indexed?

No crawler can reproduce Google’s complete index. SEO Spider can identify potentially indexable URLs and connect relevant Search Console data, but important cases still need Google-specific verification.

Can it audit very large websites?

Yes, with suitable scope, database storage, hardware and server-friendly speed settings. Start with representative samples and priority sections before assuming every URL must be crawled in one run.

Is Screaming Frog safe to use?

It is a legitimate professional crawler, but any crawler can create unnecessary server load when configured aggressively. Crawl only authorised sites, set a considerate speed and coordinate large production audits.

Does Screaming Frog fix SEO problems?

No. It discovers and organises evidence. A person still needs to validate the issue, understand business impact, choose the correct repair and confirm the result with a recrawl and real performance data.

SEO
Need the crawl turned into an action plan?

GuestPost.UK can combine Screaming Frog evidence with Search Console, analytics, content, architecture and backlink review, then prioritise fixes by risk, reach and commercial value.

Request a Professional SEO Audit →
guestpost.uk new logo
💙 PayPal
💳 VISA
💳 Mastercard
🏦 Bank
🔒 SSL

© 2026. All rights reserved.

AI
GuestPost AI ConsultantSEO Consultant · Link Building · GEO · Tools
Ask about packages, pricing, SEO tools or a growth plan