Screaming Frog SEO Spider is a desktop crawler, not a hosted SEO dashboard. That distinction explains both its strengths and its failure modes. It gives an operator unusually granular control over discovery, rendering, extraction, storage and export, while asking that operator to manage configuration, hardware, crawl load and sensitive local data.
> Distinctive strength: A desktop operator gets unusually granular control over discovery, rendering, extraction, storage and export with crawl data kept locally. > > Where it stops being an advantage: That control transfers configuration, hardware, crawl-load and sensitive-data responsibility to the operator.
BenPicks has not benchmarked the application hands-on for this profile. The analysis uses current official pricing, FAQ, configuration and general user guides, tutorials and release documentation. The goal is not to repeat a feature list; it is to show what a buyer must configure, measure and protect before trusting a crawl.
The Screaming Frog decision in 60 seconds
| Question | Source-led answer |
|---|
| Best fit | Technical SEO or engineering operator needing configurable local crawls and portable evidence. |
| Poor fit | Team wanting a zero-maintenance hosted dashboard with shared browser access. |
| Free version | 500 URLs per crawl; advanced configuration, saving and many integrations restricted. |
| Paid price | £199, $279 or €245 per user/year; bulk GBP discounts start at five users. |
| Practical scale | Hardware, storage and page complexity govern capacity; five million is a default control, not a promise. |
| Data boundary | Local by default, but APIs, Drive, AI, MCP and custom code can transmit or expose crawl data. |
| Main unknowns | Complete app telemetry lifecycle, centralized enterprise controls and binding support SLA. |
How much does Screaming Frog cost?
One paid licence currently costs £199, $279 or €245 per year. Licences are individual: two operators require two licences. GBP bulk pricing is £189 each for 5–9, £179 for 10–19 and £169 for 20 or more.
That makes the software price unusually transparent. One operator costs £199/year before tax. Five users at £189 cost £945/year. Twenty at £169 cost £3,380/year. Those are direct calculations from the current table, not total-cost estimates.
Total cost also includes suitable SSD/RAM, storage/backup, external API fees, proxy or cloud VM use and technical operator time. AI providers such as OpenAI, Gemini or Anthropic are separate accounts and bills. A low licence price does not make a poorly configured recurring audit cheap.
By default a licence expires back to the restricted free version. Auto-renewal is opt-in at purchase and can be cancelled through subscriptions. There is no monthly tier to compare with SaaS tools.
Is the free version enough?
The free edition crawls up to 500 URLs per crawl and is a valid way to test basic discovery and issue reporting. It cannot save and reopen crawls, and configuration and advanced features are restricted. Use it to learn the interface and audit a deliberately small fixture—not to infer whether a million-URL production crawl will work.
Paid access matters when you need saved configurations, JavaScript rendering, scheduled or command-line runs, crawl comparisons, custom extraction, structured-data checks, external APIs and repeatable evidence. The decision is less “Do I need more than 500 URLs?” and more “Do I need a reproducible audit system rather than a one-off scan?”
How many URLs can it really crawl?
Licensed configuration replaces the free cap with a default five-million-URL control. Official documentation calls this a control rather than a hard ceiling. It estimates that a machine with a 500GB SSD and 16GB RAM can approach ten million URLs in database mode. A cloud tutorial reports 3.1 million URLs on an 8-vCPU, 32GB RAM, 200GB SSD VM in roughly two days.
Those examples are not capacity commitments. URL response time, page size, rendering, assets, custom extraction, external APIs, crawl speed and disk all matter. Start with a representative sample and record URLs per minute, database growth, memory, CPU, target-server errors and estimated completion time.
Set explicit limits for total URLs, depth, query strings, subdomains, redirects, URL length, links per page, page size and path. An unrestricted crawler can discover calendar or faceted-navigation traps, overwhelm a fragile origin or fill local disk long before it produces a useful report.
Database or memory storage?
Database storage is the default and recommended mode for an SSD. It auto-saves crawls, opens large jobs faster and enables comparison, change detection and segments. Retention rules can delete old crawls after a chosen period, while critical projects can be locked.
The warning buyers should not miss: network drives, remotely synchronized folders such as Dropbox or OneDrive, and vault drives are unsupported for the crawl database. File locking, latency or unreliable connections can interfere with the database. Use a verified local filesystem and treat backup/export as a separate step after the crawl is closed.
Memory mode is suited to smaller sites or machines without enough disk. The vendor's rough guide says 8GB RAM can handle a couple hundred thousand URLs, but page complexity changes this. Allocate enough memory while leaving headroom for the OS and browser rendering.
Does JavaScript rendering change the audit?
Yes. Default HTML crawling and rendered crawling answer different questions. Rendering can expose links, canonicals, cookies, structured data and content added by client JavaScript. It also consumes more CPU and memory and can trigger application behavior or third-party requests that a raw HTML crawl never sees.
Run a bounded site section in both modes. Compare discovered URLs, canonical targets, robots directives, content, links and cookies. If results differ, do not merge them into one unexplained issue count. Identify whether the difference reflects intended rendering, delayed interaction, blocked resources or crawler configuration.
Structured-data extraction is also configurable and is not enabled by default for every format. A zero-issue report may mean the extractor was not switched on.
Can it automate recurring audits?
The licensed product supports scheduled jobs and full command-line or headless operation. A schedule can load configuration, authentication and API profiles, run a crawl, export tabs and reports, save to local files or Google Sheets and notify on completion.
Each scheduled crawl launches a new application instance. Overlapping schedules do not politely queue; they can run concurrently and compete for CPU, memory, disk and target-server capacity. Set a maximum duration, leave a buffer between jobs and monitor failed or unfinished runs. A report should not be accepted merely because a scheduled file exists.
Keep the configuration next to the output evidence. Record application version, storage mode, user agent, speed, rendering, scope, exclusions, authentication, APIs and extraction rules. Without that record, a second operator cannot explain why a later crawl disagrees.
What data leaves the machine?
The crawl database and exports are local by default. That does not make every workflow offline. Google Analytics, Search Console, PageSpeed, Majestic, Ahrefs and Moz integrations exchange data with external services. Google Sheets exports put results into Drive. Custom JavaScript can call any API and save files. Direct AI prompts can send selected page data to OpenAI, Gemini or Anthropic; Ollama or another local endpoint can keep processing local if configured that way.
The documentation explicitly says users are responsible for privacy when custom JavaScript sends data and must remove API keys before sharing snippets. Treat every connector as a separate data flow. Define allowed fields, provider, region, retention, credentials, cost and authorization before enabling it.
Authentication profiles store an encrypted form of credentials on disk using AES-256 GCM. Encryption reduces exposure but does not replace device security, least privilege, key rotation or a documented deletion process.
What changes with MCP and AI features?
Recent versions can operate as an MCP server and expose crawl data and actions to an LLM client. Optional tools can run Node scripts, install npm packages and read or write within allowed directories; powerful Node and filesystem capabilities are disabled by default.
Keep them disabled unless the task requires them. If enabled, use a dedicated directory, no production credentials, a restricted test crawl and a review of every tool call. An assistant that can install packages and overwrite files is no longer merely summarizing crawl evidence.
Direct AI prompts can generate alt text, classify intent, extract data or analyze content. Their output remains model-generated. Validate a stratified sample, retain prompts and model settings, and measure false positives and cost. Do not publish generated fixes automatically from a crawl.
Who should shortlist Screaming Frog?
Shortlist it when an operator needs transparent control over crawling, rendering, extraction and export; can maintain local hardware and configurations; and wants evidence that can be rerun or compared. It is especially compelling when a hosted platform's crawl quota or opaque configuration blocks investigation.
Look elsewhere when nontechnical stakeholders need a shared browser dashboard, centralized administration, vendor-managed infrastructure or turnkey collaboration. A hosted crawler may cost more but reduce desktop operations and configuration risk.
Compare choices through AI SEO software, the technical SEO crawler comparison guide, and the enterprise AI SEO buying guide.
A reproducible Screaming Frog evaluation
- Build a bounded fixture with known redirects, errors, robots rules, canonicals, hreflang, structured data, duplicate pages, query traps and JavaScript-only links.
- Record the expected truth before crawling.
- Run the free or paid default HTML crawl with a conservative speed and explicit URL cap.
- Save the configuration and export issue and URL-level evidence.
- Repeat with JavaScript rendering and explain every discovery or directive difference.
- Enable structured data, custom extraction and cookies only when the question requires them.
- Mutate five known properties, rerun and verify comparison/change detection catches all five without unrelated noise.
- Measure time, requests, CPU, RAM and database growth; project the production workload from the observed sample.
- Repeat the saved configuration with a second operator or machine and compare output hashes/counts.
- Test one scheduled headless run with no overlap, bounded output and failure notification.
- Exercise one required API or AI connector using non-sensitive data; document what leaves the device and its cost.
- Export a portable crawl and critical CSV evidence, then restore it in a clean environment.
Pass only when the crawler finds the fixture truth, configuration is reproducible, production load is safe, storage is supported, sensitive data remains within approved flows and a second operator can explain the report. Issue count alone is not an audit verdict.
Official sources