ChatGPT answers, unedited
"we need fast site crawls for a large site, how to compare crawl speed, limits, and cost between platforms?", asked on ChatGPT about once a day from the United States. The last 22 answers, newest first, exactly as given: the raw data behind the SEO Tools live demo.
Week of 24th to 30th Aug 2026 (in progress)
If the priority is fast crawls on a large site, I’d compare platforms on effective pages/sec, maximum crawl size, monthly URL allowance, concurrency, and cost per million pages—not just the advertised “crawl speed.”
Quick comparison
| Platform | Architecture | Published speed / scale | Crawl limits | Cost model | Best fit |
|---|---|---|---|---|---|
| JetOctopus | Cloud | Up to 250 pages/sec | Cloud capacity marketed as unlimited; 1M URLs can take ~1–2 days at medium settings | Usage/plan based | Very large / high-frequency crawls |
| Sitebulb Cloud | Cloud | 300+ URLs/sec at top end | Up to 10M URLs/audit; plans impose monthly URL quotas | Monthly, URL-volume tiers | Large sites + teams |
| Screaming Frog | Desktop | Hardware/server-dependent | 500k default; can push to ~2M with appropriate hardware | $259/year license | Best value for ≤500k–1M and technical control |
| Lumar / Botify / OnCrawl | Cloud | Enterprise-dependent | Generally designed for millions+ | Enterprise/custom | Enterprise SEO, automation, logs |
Sitebulb officially says its Cloud crawler can exceed 300 URLs/sec, while its desktop crawler defaults to 5 URLs/sec and is constrained by your machine and site response time. Its Cloud plans range from 50k URLs/month to 5M+ on Enterprise, with individual-audit limits from 50k to 2.5M+. Sitebulb Sitebulb
JetOctopus publishes a maximum of 250 pages/sec and says 1M URLs can take roughly 1–2 days at a medium five-thread setting, although faster crawling is possible when the target site can tolerate it. Tech SEO Platform Tech SEO Platform
The important metric: cost per million URLs
For your use case, I'd build a benchmark like:
Effective throughput = URLs successfully crawled / elapsed hour
Then:
Cost per 1M URLs = monthly platform cost ÷ (monthly URL allowance / 1M)
And separately track:
- HTML pages/sec
- JS-rendered pages/sec
- % of requested URLs actually completed
- crawl duration for 1M / 5M / 10M URLs
- concurrent crawls allowed
- monthly URL quota
- whether retries/errors consume quota
- storage/retention costs
- API/export costs
- number of users/seats
- scheduling frequency
- rate-limit controls
This matters because a crawler advertised at 250 pages/sec isn't necessarily 5× better than one doing 50 pages/sec. Your origin server, robots.txt, response latency, rendering requirements, and crawler concurrency can become the bottleneck.
How I'd shortlist them
1. JetOctopus — speed/scale first
I'd put this near the top if you're talking about millions of URLs and frequent recrawling. Its cloud architecture removes the laptop/RAM bottleneck, and its published ceiling is 250 pages/sec. It also has native log-file analysis, which is valuable if you're trying to understand Googlebot behavior rather than merely audit HTML. Tech SEO Platform Tech SEO Platform
2. Sitebulb Cloud — strong alternative
Sitebulb Cloud is particularly interesting if you need millions of URLs but also collaborative auditing/reporting. It supports up to 10M URLs per audit and says the Cloud crawler can exceed 300 URLs/sec. The tradeoff is that its pricing is explicitly tied to monthly URL volume. Sitebulb Sitebulb
3. Screaming Frog — best economics when scale isn't extreme
For a site that is, say, 100k–500k URLs, I'd benchmark Screaming Frog very seriously. Its $259/year license is radically cheaper than cloud enterprise crawlers, and you're essentially paying for the software rather than a large recurring crawl allowance. The downside is that your machine becomes part of the infrastructure. Astro SEO Blog
A practical decision rule
- <100k URLs: Screaming Frog is usually the economic winner.
- 100k–500k: Benchmark Screaming Frog against Sitebulb; raw speed may matter less than workflow.
- 500k–1M: Cloud starts becoming much more attractive.
- 1M–5M: I'd focus the benchmark on JetOctopus vs Sitebulb Cloud vs enterprise platforms.
- 5M–10M+: Treat this as a cloud-infrastructure problem rather than a desktop SEO-tool problem.
One particularly useful wrinkle: Sitebulb's Cloud pricing is quota-based, so if you crawl a 2M-page site weekly, you're really buying capacity for ~8M URLs/month—not merely a “2M URL crawler.” Their own example shows that changing crawl frequency can move a customer between plans. Sitebulb
If speed is the #1 requirement
I would run a controlled bake-off rather than trusting vendor numbers.
Use the same 100k–500k representative URL sample on each platform and record:
| Test | What to measure |
|---|---|
| HTML crawl | URLs/sec and total duration |
| High-latency pages | URLs/sec under realistic response times |
| Redirect-heavy section | Completed URLs/sec |
| JS rendering | Rendered pages/sec |
| Deep crawl | Completion rate + duration |
| Concurrent crawls | Aggregate URLs/sec |
| Re-crawl | Time to identify changed URLs |
| API/export | Time from crawl completion to usable dataset |
Then calculate $ / 1M successfully crawled URLs and hours / 1M URLs.
If you tell me roughly how many URLs the site has (1M, 5M, 20M, etc.), how often you need to recrawl it, whether JS rendering is required, and whether you need log-file analysis, I can turn this into a much more concrete platform-by-platform cost/throughput model.
Behind this answer: no web searches, 4 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Week of 17th to 23rd Aug 2026
Answer 2 of 22, 23rd Aug 2026, unknown, mentions 4 tracked brands
If the goal is fast, repeatable crawls of a large site, I’d compare platforms on effective URLs/sec + URL quota + rendering cost, not just headline “crawl speed.”
Quick comparison
| Platform | Crawl speed / scale | Limits & billing model | Cost profile | Best fit |
|---|---|---|---|---|
| Screaming Frog | Local; speed scales with CPU/network. Default 5 threads, configurable by threads or max URLs/sec. Licensed version has a nominal 5M default limit but can go much higher with hardware. Screaming Frog Screaming Frog | Essentially hardware-limited; 5M default, not a hard ceiling | $279/user/year currently; no per-URL crawl credits. Screaming Frog | Cheapest option if you can run your own crawler infrastructure |
| Sitebulb Cloud | Claims 300+ URLs/sec at the high end; cloud removes workstation constraints. Sitebulb Sitebulb | Plans currently range from 50K URLs/month up through 5M+; max audit from 50K to 2.5M+ depending on plan. Sitebulb | Starts around $245/mo for Cloud according to its current comparison material. Sitebulb | Strong price/performance for large SEO crawls |
| Lumar | Advertises up to 450 URLs/sec non-rendered and ~350 rendered. Lumar Lumar | Enterprise/custom packaging | Custom quote. Lumar | Very large sites where crawl throughput is important |
| Oncrawl | Designed for millions of URLs and says it has no artificial speed/site-size limits. Oncrawl - Technical SEO Data | Monthly crawl-token quota; JS/CWV crawling can consume multiple tokens per URL. Oncrawl Help Oncrawl - Technical SEO Data | Custom/plan-based | Sites needing crawl + log-file/data analysis |
| Botify | Enterprise-scale crawler/log analysis rather than a simple URL crawler | Enterprise/custom | Custom | Huge sites where Googlebot crawl behavior and log analysis matter as much as your own crawl speed |
The important catch: URLs/sec isn't comparable by itself
For a 10-million-URL site, I'd benchmark these four things:
- Raw HTML throughput — URLs/sec with JavaScript disabled.
- Rendered throughput — URLs/sec with JS rendering enabled.
- Effective throughput — URLs successfully processed per wall-clock hour, including retries, throttling and errors.
- Cost per million URLs — including the cost of JS rendering, crawl credits, infrastructure and seats.
For example, Sitebulb's published maximum is 300+ URLs/sec, while Lumar publishes up to 450 URLs/sec. Those numbers are useful directional indicators, but they're not an apples-to-apples benchmark because server TTFB, concurrency, page size, redirects and rendering workload can dominate actual throughput. Sitebulb Support Lumar
For your use case, I'd test them like this
Take a representative 1M-URL slice of the site and run:
- HTML-only crawl
- JS-rendered crawl
- Same URL list on every platform
- Same robots/no-robots policy
- Same redirect handling
- Same user agent where possible
- Same concurrency/rate ceiling
- Record: - wall-clock time
- successful URLs
- requests generated
- errors/timeouts/429s
- average and p95 response time
- CPU/RAM/network consumption
- incremental cost
Then calculate:
Throughput = successfully processed URLs ÷ crawl hours
Cost / 1M URLs = total crawl cost ÷ URLs crawled × 1,000,000
That last metric is especially important with Oncrawl-style token systems: 100 URLs isn't necessarily 100 quota units when you turn on JS/CWV features. Oncrawl Help Oncrawl - Technical SEO Data
My shortlist
If I were evaluating this for a large site today:
- Lowest cost / maximum control: Screaming Frog on your own cloud VM. This is surprisingly compelling. Screaming Frog documents a 3.1M-URL crawl taking ~2 days on an 8-vCPU/32GB/200GB SSD GCP VM, costing under £20 in compute charges; you can scale the VM if the origin can handle more concurrency. Screaming Frog
- Best managed price/performance: Sitebulb Cloud. It gives you cloud execution, no project limits, and up to 10M URLs per audit, with published pricing rather than a sales-only quote. Sitebulb Sitebulb
- Maximum managed crawling throughput: Lumar is worth benchmarking, particularly if 100M+ URL workloads or very high concurrency are involved. Its published 450 URL/sec figure is among the clearest high-end throughput claims. Lumar
- Crawl + logs + search-engine behavior: Oncrawl/Botify become more interesting than a pure crawler because you're analyzing how Google/Bing actually crawl the site, not merely how quickly your crawler can fetch it. Oncrawl - Technical SEO Data Botify
One particularly useful baseline: if your site is millions of URLs and you're comfortable operating infrastructure, I'd benchmark Screaming Frog on a cloud VM vs Sitebulb Cloud vs Lumar first. That will tell you whether the premium for managed enterprise crawling is actually buying you meaningful throughput rather than just convenience.
If you give me roughly site size (e.g. 5M/50M/500M URLs), JS-heavy vs mostly HTML, desired recrawl frequency, and whether you need logs, I can turn this into a concrete $/month + expected crawl time comparison for the platforms.
Behind this answer: no web searches, 4 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Answer 3 of 22, 22nd Aug 2026, unknown, mentions 2 tracked brands
For a large site, “crawl speed” is usually not a single metric. You want to compare throughput, throttling behavior, crawl limits, rendering cost, and total cost per million URLs. A good benchmark process looks like this:
1. Define the crawl workload first
Platforms behave very differently depending on what you crawl.
Measure:
| Dimension | Example values |
|---|---|
| URL count | 500k, 5M, 50M URLs |
| Page type | HTML only vs JS-heavy SPA |
| Rendering | No browser vs Chromium rendering |
| Data needed | Links/status codes vs extracted content |
| Frequency | One-time audit vs daily monitoring |
| Concurrency | 10, 100, 1,000+ requests/sec |
A crawler that handles 10M static URLs quickly may be much slower once every page requires JavaScript execution.
2. Compare crawl speed using normalized metrics
Avoid vendor claims like “fast crawler.” Track:
Throughput
- URLs/minute
- URLs/hour
- Requests/second
- Successful pages/hour
Example:
Crawler A:
5,000,000 URLs
6 hours
= 13,888 URLs/min
Crawler B:
5,000,000 URLs
10 hours
= 8,333 URLs/min
Time to first results
Important for large audits:
- How long until first 1,000 URLs are processed?
- Can results stream while crawling?
Scaling curve
Test:
- 10 concurrent workers
- 100 workers
- 1,000 workers
Some systems scale linearly; others hit hidden limits.
3. Compare limits
Build a matrix like:
| Capability | Platform A | Platform B | Platform C |
|---|---|---|---|
| Max URLs/crawl | |||
| Requests/sec limit | |||
| Concurrent crawlers | |||
| Daily/monthly crawl quota | |||
| JS rendering limits | |||
| API rate limits | |||
| Export limits | |||
| Historical storage |
For example, some crawler products publish limits such as millions of requests per crawl and concurrency tiers, while enterprise SEO crawlers often use custom contracts rather than public limits. Crawlspace theStacc
4. Compare cost by effective unit price
Do not compare monthly subscription prices alone.
Calculate:
Cost per million URLs
Formula:
(monthly cost + infrastructure cost) / crawled URLs × 1,000,000
Example:
| Platform | Monthly cost | Crawl volume | Effective cost |
|---|---|---|---|
| A | $500 | 10M URLs | $50 / M URLs |
| B | $5,000 | 100M URLs | $50 / M URLs |
| C | $20,000 | 1B URLs | $20 / M URLs |
Include:
- API calls
- proxy costs
- browser rendering charges
- storage
- engineering time
- retry overhead
Some platforms use page credits, while others meter requests, tokens, or concurrency, making headline prices hard to compare directly. fastCRW
5. Run a controlled benchmark
Use the same:
- domain
- URL list
- crawl depth
- user agent
- concurrency
- rendering settings
- extraction rules
Record:
Start time:
End time:
URLs discovered:
URLs fetched:
HTTP success rate:
Average latency:
Peak requests/sec:
CPU/memory:
Total cost:
For a large production site, run at least:
Test A — Static crawl
- HTML only
- no JS
- 1M URLs
Test B — Realistic crawl
- redirects
- canonicals
- robots handling
- sitemap discovery
Test C — Heavy crawl
- JS rendering
- screenshots
- content extraction
6. Platform categories to compare
Enterprise SEO crawlers
Best for:
- technical SEO audits
- indexability analysis
- scheduled monitoring
Examples include:
- Lumar
- Botify
- Oncrawl
These usually trade lower raw crawl cost for reporting, integrations, and SEO-specific analysis. Enterprise pricing is often custom. theStacc
Developer crawling APIs
Best for:
- building your own pipeline
- extracting content
- feeding search/AI systems
Examples:
- Bright Data
- Oxylabs
- Apify
These often optimize for raw acquisition scale, proxies, rendering, and APIs rather than SEO reports. Context.dev
Self-hosted crawlers
Best for:
- predictable workloads
- maximum control
Examples:
- Apache Nutch
- StormCrawler
- custom Go/Rust crawlers
Costs shift from vendor fees to:
- cloud compute
- storage
- engineering
7. Questions to ask vendors
Ask each vendor:
- “How many URLs/hour can you crawl on a 10M URL site?”
- “Is that with JavaScript rendering enabled?”
- “What happens when we exceed limits?”
- “Are retries counted as pages?”
- “Do redirects and errors consume credits?”
- “Can we control concurrency?”
- “Can you provide crawl logs?”
- “What is the cost at 100M and 1B URLs/month?”
For a large site, the most meaningful benchmark is usually:
URLs/hour + JS rendering capability + cost per million successful pages + operational limits.
A crawler that is twice as fast but costs 5× more may still be cheaper if it reduces infrastructure and engineering work.
Behind this answer: no web searches, 2 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Answer 4 of 22, 21st Aug 2026, unknown, mentions no tracked brands
If the priority is fast, large-site crawls, I’d compare platforms on throughput rather than “pages/minute” marketing numbers. The real bottleneck is usually concurrency × target response time × rendering/anti-bot overhead.
Quick comparison
| Platform | Crawl throughput / limits | Cost model | Best fit |
|---|---|---|---|
| Firecrawl | 50 concurrent browsers on Standard, 100 Growth, 150 Scale; crawl API limits range from 100 → 2,000 requests/min depending on plan | 1 credit/page for crawl; current plans include 100k, 500k, or 1M credits | Fastest path to a managed whole-site crawler |
| Zyte API | 3,000 RPM standard; 10,000 RPM Enterprise; concurrency effectively depends on RPM × target response time | Per successful response; HTTP roughly $0.13–$1.27/1k, browser $1.01–$16.08/1k depending on site difficulty | High-volume scraping where anti-bot + reliability matter |
| Apify | Compute-based rather than a simple fixed crawl concurrency ceiling; you control Actor architecture and parallelism | $0.20/CU Free/Starter, $0.16 Scale, $0.13 Business + monthly plan | Maximum flexibility; good if you're willing to build/tune the crawler |
| Scrapy Cloud | 1 concurrent crawl on free; paid capacity is based on Scrapy Units (1 GB RAM + 1 concurrent crawl) | From $9/unit/month | Best if you already have Scrapy and want to scale your own crawler |
Firecrawl's current published limits are particularly straightforward: Standard has 50 concurrent browsers / 100 crawl requests per minute, Growth 100 / 1,000, and Scale 150 / 2,000. GitHub Firecrawl
Zyte is different: its standard API limit is 3,000 RPM, and concurrency isn't a fixed number. Zyte gives the approximation concurrency ≈ RPM / 60 × average response time. So with 3,000 RPM and a 2-second average response, you're looking at roughly 100 concurrent requests. Zyte
Cost at large scale
For a 1-million-page crawl, the economics can look very different:
- Firecrawl: roughly 1M credits for a basic crawl, before extra costs for rendering/features. Its current Scale plan is 1M credits/month at $599 when billed yearly. Firecrawl GitHub
- Zyte: potentially much cheaper for straightforward HTTP pages—published rates range from $0.13 to $1.27 per 1,000 requests PAYG, depending on site difficulty. Browser rendering ranges from $1.01 to $16.08/1,000. Volume commitments reduce those rates. Zyte #1 Web Scraping Service Zyte
- Apify: you're essentially paying for compute. Current plans put compute at $0.20/CU, falling to $0.16 at Scale and $0.13 at Business, so your crawler implementation has a major effect on the final cost. Apify
- Scrapy Cloud: capacity-based pricing can be attractive if your crawler is efficient; a Scrapy Unit represents 1 GB RAM + 1 concurrent crawl, starting at $9/unit/month on Professional. Zyte #1 Web Scraping Service
The benchmark I'd actually run
Don't benchmark "1,000 URLs." Give every platform the same 100k–1M URL corpus and measure:
- Pages successfully fetched/hour
- P50/P95 page latency
- Effective concurrency
- % failures / retries
- % pages requiring browser rendering
- % blocked / CAPTCHA
- CPU/RAM consumed
- Total cost per 100k successful pages
- Cost per successful page, not cost per request
- Time to finish the entire crawl
For example, if Platform A does 1,000 pages/min at $5/100k but Platform B does 3,000/min at $15/100k, B may be the better choice if finishing the crawl 3× sooner has meaningful value.
My shortlist
For your use case I'd start with:
1. Firecrawl — easiest apples-to-apples test for a large website. Its pricing maps directly to pages and it exposes explicit crawl/concurrency limits. GitHub Firecrawl
2. Zyte API — test this if the site has significant bot protection, JavaScript, geographic variation, or intermittent blocking. Its infrastructure automatically chooses technologies based on the target site, and you pay for successful responses. Zyte
3. Apify — test this if you need custom crawling logic, sophisticated queues, transformations, or want control over the crawler architecture. Apify
4. Scrapy Cloud — particularly compelling if your team already uses Scrapy and wants to scale horizontally without operating the infrastructure yourselves. Zyte #1 Web Scraping Service
One important caveat: don't let raw concurrency determine the winner. A crawler with 150 browser slots can be dramatically slower and more expensive than one doing thousands of lightweight HTTP requests. If most of your target pages are static HTML, HTTP throughput should be your primary benchmark, with browser rendering tested as a separate workload.
If you tell me roughly how many pages the site has (100k / 1M / 10M+), whether JS rendering is needed, and whether you're crawling one domain or many, I can turn this into a concrete cost + expected completion-time comparison for the platforms above.
Behind this answer: no web searches
ChatGPT answered from its own knowledge, without searching the web.
Answer 5 of 22, 20th Aug 2026, gpt-5-6, mentions 2 tracked brands
If crawl speed is the priority for a large site, I’d compare platforms on effective URLs/hour, not the vendor’s maximum requests/sec. The real bottleneck is usually your origin, JS rendering, concurrency, and how the platform meters crawl volume.
Quick comparison
| Platform | Speed / scale | Crawl limits & metering | Cost model | Best fit |
|---|---|---|---|---|
| JetOctopus | Up to 250 pages/sec; says 1M+ pages/day for enterprise | Standard plans are URL-volume based; Enterprise Unlimited removes crawl caps | Volume-based or custom unlimited | Very large sites / fastest cloud crawling |
| Sitebulb Cloud | Hundreds of URLs/sec | Cloud plans have monthly URL/audit limits; Enterprise starts at 5M+ URLs/month | Subscription by capacity | Large sites where usability + crawling matter |
| Oncrawl | Configurable pages/sec; >1 URL/sec requires domain verification | You control max URLs, depth and speed | Typically enterprise/custom | Enterprise SEO + crawl/log analysis |
| Screaming Frog | Desktop/local; speed depends heavily on your machine/network | Paid version has effectively unlimited URLs, constrained by memory/storage | £199/user/year | Cheapest option if you can run crawls yourself |
JetOctopus is the standout if your requirement is specifically "crawl millions of pages as fast as possible." It says its crawler can reach 250 pages/sec, and its enterprise offering advertises 1M+ pages/day, no crawl cap, unlimited simultaneous crawls, and no per-user/domain charges. jetoctopus.com
Sitebulb is also interesting: its cloud product claims hundreds of URLs/sec, while its published server plans range from 1M URLs/month with a 250K-URL audit limit to Enterprise at 5M+ URLs/month and 2.5M+ URLs per audit. sitebulb.com
Oncrawl gives you explicit control over crawl speed and maximum URLs/depth. Importantly, it requires domain verification to crawl faster than 1 URL/sec, and it warns that overly aggressive crawling can overwhelm the origin. help.oncrawl.com
Screaming Frog is a different proposition: it's local rather than cloud-scale. The paid license is £199/year per user and removes the 500-URL limit; the actual maximum is constrained by your machine's memory/storage. www.screamingfrog.co.uk
How I'd benchmark them
For a genuinely large site, don't simply ask vendors "what's your max crawl rate?" Run the same crawl against the same environment:
- Use a representative 1–5M URL corpus, ideally from your sitemap plus known internal URLs.
- Run HTML-only first, then repeat with JavaScript rendering.
- Record:
- URLs successfully fetched
- URLs/minute and URLs/hour
- median and p95 response time
- 429/5xx rate
- bandwidth consumed
- crawl duration
- CPU/RAM impact on your infrastructure
- Test at increasing concurrency: e.g. 10 → 25 → 50 → 100 requests/sec.
- Measure the highest sustainable rate before your origin starts returning errors.
- Calculate:
Effective crawl cost = monthly platform cost ÷ successfully crawled URLs
And separately:
Effective speed = successfully crawled URLs ÷ elapsed time
That exposes an important difference between platforms: 250 pages/sec isn't useful if your site can only safely sustain 40/sec.
My shortlist
If I were evaluating this for a 10M+ page site, I'd start with:
1. JetOctopus — strongest candidate when raw cloud crawling throughput and very large URL volumes dominate. Their enterprise plan explicitly targets 10M/50M/100M+ pages and removes crawl-volume caps. jetoctopus.comjetoctopus.com
2. Sitebulb Cloud — strong alternative if you want high throughput but also put a lot of weight on workflow, reporting, and ease of use. sitebulb.com
3. Oncrawl — worth testing if your crawler needs to be tightly controlled and you want sophisticated enterprise crawl/log analysis. help.oncrawl.com
4. Screaming Frog — excellent value for controlled, engineering-style crawls, but I'd rule it out if the requirement is continuous cloud crawling of tens of millions of URLs. www.screamingfrog.co.uk
If you tell me roughly how many URLs you have (1M / 10M / 100M+), whether JS rendering is required, and how often you want to recrawl, I can turn this into a $/1M URLs + expected crawl time comparison for the platforms.
Behind this answer: 4 searches, 30 results retrieved, 2 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://jetoctopus.com/faq-limits/?utm_source=chatgpt.com | not cited |
| 2 | https://jetoctopus.com/unlimited/?utm_source=chatgpt.com | not cited |
| 3 | https://jetoctopus.com/?utm_source=chatgpt.com | not cited |
| 4 | https://jetoctopus.com/how-to-create-a-new-crawl/?utm_source=chatgpt.com | not cited |
| 5 | https://jetoctopus.com/faqs-2/?utm_source=chatgpt.com | not cited |
| 6 | https://help.oncrawl.com/en/articles/3136175-how-to-modify-crawl-limits-while-crawling?utm_source=chatgpt.com | not cited |
| 7 | https://help.oncrawl.com/en/articles/2565144-why-is-my-crawl-blocked-at-one-url-why-won-t-my-crawl-start?utm_source=chatgpt.com | not cited |
| 8 | https://www.oncrawl.com/technical-seo/oncrawl-desktop-crawler-cloud-based-seo-platform/?utm_source=chatgpt.com | not cited |
| and 22 more results | ||
Tracked brands mentioned
Answer 6 of 22, 19th Aug 2026, gpt-5-6, mentions 3 tracked brands
If the priority is fast, repeatable crawls of a large site, compare platforms on effective throughput, not just the advertised page limit.
| Platform | Crawl-scale model | Speed controls | Published crawl limits | Cost model | Best fit |
|---|---|---|---|---|---|
| Sitebulb Cloud | Cloud infrastructure; large/custom plans can run concurrent crawls | Threads / Chrome instances + URLs/sec cap | Plan-dependent; Cloud is designed for very large crawls | Flat monthly subscription; no crawl credits | Best for frequent large crawls |
| Ahrefs Site Audit | Cloud crawler | Less hands-on speed tuning; “Always-on” audit option | 100k / 500k / 1.5M crawl credits on Lite/Standard/Advanced; up to 5M+ Enterprise; max project sizes 25k–5M | Subscription + crawl credits | Best if you also need Ahrefs' SEO dataset |
| Semrush Site Audit | Cloud crawler | Crawl configuration, page caps, source/masks | 100k / 300k / 1M pages/month for Pro/Guru/Business; per-audit caps of 20k/20k/100k | Subscription with monthly page allowance | Good all-in-one SEO suite |
| Sitebulb Desktop | Your machine | Threads and URLs/sec | Lite: 10k; Pro normally 500k, configurable up to 2M | License | Good for controlled/internal crawls |
Sitebulb explicitly lets you increase threads and set a maximum HTML URLs/sec; actual throughput still depends heavily on TTFB, machine/server resources, and rendering. Its HTML crawler is considerably faster than its Chrome renderer because Chrome has to download and render page resources. support.sitebulb.com
How I'd benchmark them
For a genuinely large site, run the same 100k–500k URL slice through each platform and record:
- Pages/minute — raw crawl throughput.
- Time to first useful report — not merely total crawl completion.
- HTTP requests/minute — important if JS/resources are included.
- JS vs. HTML throughput — test both if your site is JS-heavy.
- Concurrency — how many requests/threads the platform actually sustains.
- 429/5xx rate — fast isn't useful if the crawler overloads your origin.
- Completeness — URLs discovered vs. your known URL set.
- Cost per million URLs — the most useful normalized cost metric.
- Cost of recurring crawls — e.g. 5 × 1M URLs/month.
- Parallel-site capacity — particularly important for agencies/enterprise.
The big pricing distinction
Sitebulb Cloud is unusually attractive for high-frequency crawling because it says Cloud has no crawl credits and uses flat monthly pricing, while allowing multiple large crawls and JS crawling at scale. sitebulb.com
By contrast, Ahrefs and Semrush effectively meter your crawling capacity. Ahrefs currently lists 100k/500k/1.5M monthly Site Audit crawl credits on its first three paid tiers, with project limits ranging from 25k to 250k URLs before Enterprise. ahrefs.com Semrush lists 100k/300k/1M pages per month for Pro/Guru/Business, with separate per-audit limits. www.semrush.com
So, for example, if you need 1M URLs every week, don't compare "$X/month" directly. Calculate:
effective cost / million crawled URLs = monthly platform cost ÷ actual monthly URLs successfully crawled
And separately calculate time per million URLs.
My shortlist
- Maximum crawl throughput + lots of recurring crawls: Sitebulb Cloud.
- Large crawl + strongest SEO research ecosystem: Ahrefs.
- Large crawl + broader marketing/SEO suite: Semrush.
- Maximum control and you have beefy infrastructure: Sitebulb Desktop.
One important caveat: if by “crawl” you mean an API/service that crawls arbitrary URLs at very high concurrency rather than an SEO crawler, I'd compare a different class of products—e.g. crawler APIs/infrastructure rather than Ahrefs/Semrush/Sitebulb. For a large site, that distinction can change the economics dramatically.
Behind this answer: 5 searches, 30 results retrieved, 3 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://ahrefs.com/pricing/?utm_source=chatgpt.com | not cited |
| 2 | https://ahrefs.com/blog/ahrefs-pricing/?utm_source=chatgpt.com | not cited |
| 3 | https://ahrefs.com/faq?utm_source=chatgpt.com | not cited |
| 4 | https://ahrefs.com/pricing?plan=awt&return=GoodbyeSemrush&utm_source=chatgpt.com | not cited |
| 5 | https://docs.ahrefs.com/en/api/docs/limits-consumption?utm_source=chatgpt.com | not cited |
| 6 | https://ahrefs.com/webmaster-tools?utm_source=chatgpt.com | not cited |
| 7 | https://docs.ahrefs.com/en/api/docs/introduction?utm_source=chatgpt.com | not cited |
| 8 | https://help.ahrefs.com/en/articles/1406260-do-you-offer-any-discounts-or-free-trials?utm_source=chatgpt.com | not cited |
| and 22 more results | ||
Tracked brands mentioned
Answer 7 of 22, 18th Aug 2026, gpt-5-6, mentions 3 tracked brands
If crawl speed is the primary requirement, I’d compare the platforms on effective URLs/sec, monthly URL economics, JS-rendering cost, concurrency, and whether the quota is per crawl or shared across the month. Vendor headline speeds are useful, but they are not directly comparable.
Quick comparison
| Platform | Published crawl speed | Crawl limits / quota | Cost model | Best fit |
|---|---|---|---|---|
| Lumar | Up to 450 URLs/s non-rendered, 350 URLs/s rendered | Account-level URL limits; default crawl setup can cap at 100k URLs | Custom enterprise pricing | Maximum raw speed / very large enterprise sites |
| JetOctopus | Up to 250 pages/s | Plans based on monthly crawl-page volume; JS page counts as 2 HTML pages | Published volume-based plans; enterprise no-limit option | High-speed cloud crawling with relatively transparent volume economics |
| Sitebulb Cloud | 300+ URLs/s advertised | Up to 10M URLs/audit on Cloud; no crawl credits | Flat monthly subscription; from $245/mo in current published comparison | Strong price/performance for large recurring crawls |
| Oncrawl | User-configurable; recommends 2–5 URLs/s for most sites, higher speeds available after verification | Monthly crawl-token quota | Monthly quota/token model | Deep SEO analysis + crawl/log/GSC data |
| Screaming Frog | Hardware/network dependent | Paid version effectively unlimited; constrained by local RAM/storage | $279/user/year | Cheapest option if you can throw hardware at it |
Sources: Lumar publishes 450/350 URLs/s benchmarks; JetOctopus says up to 250 pages/s; Sitebulb says Cloud can exceed 300 URLs/s; Oncrawl documents configurable speed; Screaming Frog's paid license is $279/year and its crawl capacity depends on machine memory/storage. www.lumar.io
The important catch: "URLs/sec" isn't the same as throughput
For a large site, I'd benchmark three modes separately:
- HTML-only
- Fetch + parse HTML.
- This is your maximum-throughput baseline.
- Rendered/JS
- Fetch + Chromium/rendering.
- Much more CPU/network intensive.
- JetOctopus, for example, counts a JS page as 2 crawl pages in its pricing model, while Sitebulb says JS crawling doesn't incur an extra charge on Cloud. jetoctopus.comsitebulb.com
- Full audit
- HTML/JS + Core Web Vitals + images/resources + custom extraction + accessibility/etc.
- This is usually the number that matters operationally, and it can be dramatically slower than the headline crawler speed.
Sitebulb explicitly notes that TTFB, thread/Chrome-instance count, rendering, and audit features affect actual throughput; its URL/sec setting is a ceiling, not a guarantee. support.sitebulb.comsitebulb.com
Where the economics get interesting
Sitebulb Cloud is particularly interesting if you're going to run lots of full-site crawls. It advertises no crawl credits, unlimited projects/domains, and up to 10M URLs per audit; its published comparison currently shows Cloud starting at $245/month. sitebulb.com
JetOctopus is more naturally modeled as cost per crawl volume. Its published plans allocate crawl pages monthly, and its pricing page currently shows a 1M-page tier at €383/month billed annually. It also offers an enterprise "no limits" package. jetoctopus.com
Oncrawl needs special care when comparing costs because its quota isn't simply "one URL = one unit." HTML fetch = 1 token, JS rendering = 3, and JS + Core Web Vitals = 5. So a 1M-page JS+CWV crawl can consume 5M tokens. help.oncrawl.com
Screaming Frog is the outlier: the software itself is inexpensive—$279/user/year—but your effective crawl cost is hardware + engineering time + crawl duration. Its paid crawl limit is effectively unlimited, subject to available memory/storage. www.lumar.iowww.screamingfrog.co.uk
What I'd benchmark
For your particular requirement, I'd put the same 1M-URL site through each platform and record:
| Metric | Why it matters |
|---|---|
| Time to first 100k URLs | Startup/discovery overhead |
| Sustained URLs/sec | Actual throughput |
| Time to 1M URLs | The number you'll ultimately care about |
| HTML URLs/sec | Baseline |
| Rendered URLs/sec | JS-heavy site performance |
| Peak concurrent requests | Server impact |
| CPU/RAM consumed | Especially important for desktop crawlers |
| Errors/5xx at each speed | Determines usable rather than theoretical speed |
| Cost for 1M HTML URLs | Basic economics |
| Cost for 1M rendered URLs | Real-world economics |
| Cost for 4 × 1M crawls/month | Recurring monitoring economics |
| Export/API throughput | Can you actually get the data out? |
Most important: don't benchmark at each vendor's maximum setting. Test at 5, 10, 25, 50, 100, and 200 URLs/sec (where supported) and plot crawl time vs. error rate. The fastest platform on paper isn't necessarily the fastest platform your origin can safely handle.
For example, Oncrawl explicitly warns that requests are made sequentially at the configured rate and that excessive speed can overload the server; it recommends 2–5 URLs/sec for most sites. help.oncrawl.comwww.lumar.io Lumar likewise says its infrastructure can sustain very high rates but recommends coordinating the maximum with DevOps because the limiting factor may become your origin infrastructure rather than the crawler. www.lumar.io
My initial shortlist
If this is millions of URLs and speed is genuinely the #1 criterion:
1. Lumar — strongest candidate for absolute throughput.
2. JetOctopus — very strong speed with straightforward crawl-volume economics.
3. Sitebulb Cloud — probably the most compelling price/performance option if its 10M-audit ceiling and feature set fit.
4. Oncrawl — excellent if crawl data needs to be combined deeply with logs/GSC, but model token consumption carefully.
5. Screaming Frog — unbeatable value if you're willing to operate the infrastructure yourself, but not the most convenient choice for continuously crawling millions of URLs.
One caveat: Lumar's pricing is custom, so you can't responsibly calculate its $/million URLs from public pricing alone. Lumar explicitly says its URL allowance and pricing are tailored to the customer's needs. www.lumar.io
If you tell me roughly how many URLs (e.g. 1M/10M/100M), HTML vs JS %, and how many crawls/month, I can turn this into a concrete $/million URLs + hours per crawl + recommended platform comparison.
Behind this answer: 7 searches, 30 results retrieved, 3 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://support.sitebulb.com/en/articles/9853652-crawler-settings?utm_source=chatgpt.com | not cited |
| 2 | https://sitebulb.com/faqs/?utm_source=chatgpt.com | not cited |
| 3 | https://support.sitebulb.com/en/articles/10056982-how-to-crawl-large-websites?utm_source=chatgpt.com | not cited |
| 4 | https://support.sitebulb.com/en/articles/9857604-concurrent-crawling?utm_source=chatgpt.com | not cited |
| 5 | https://sitebulb.com/use-cases/enterprise/?utm_source=chatgpt.com | not cited |
| 6 | https://sitebulb.com/cloud/?utm_source=chatgpt.com | not cited |
| 7 | https://sitebulb.com/?utm_source=chatgpt.com | not cited |
| 8 | https://sitebulb.com/subscriptions/pricing/server/?annual=false&utm_source=chatgpt.com | not cited |
| and 22 more results | ||
Tracked brands mentioned
Answer 8 of 22, 17th Aug 2026, gpt-5-6, mentions 4 tracked brands
If crawl speed at large scale is the main criterion, I’d compare the platforms on sustained URLs/sec, rendered URLs/sec, hard crawl limits, and pricing model—not just the vendor’s headline “fastest” number.
Quick comparison
| Platform | Published crawl speed | Scale / limits | Cost model | Best fit |
|---|---|---|---|---|
| Lumar | Up to 450 URLs/sec; ~350/sec rendered in tests | Designed for hundreds of thousands → millions of URLs; configurable URL/depth limits | Enterprise/custom | Maximum raw cloud crawl throughput |
| Botify | Crawler can exceed 200 URLs/sec; configurable max speed | Millions of pages; SiteCrawler advertises no crawl-budget limits; validated sites can exceed 3 URLs/sec | Enterprise/custom | Large enterprise SEO + analytics |
| Sitebulb Cloud | No comparable headline ceiling published in the material I found | 2.5M+ URLs/audit, 5M+ URLs/month on Enterprise; no crawl credits | From £95/mo, enterprise custom | Strong price/scale balance |
| Screaming Frog | Primarily machine/resource constrained rather than a cloud throughput service | Paid version advertises unlimited URLs, subject to RAM/storage | $279/user/year currently | Cheapest option if you can run your own crawl infrastructure |
Lumar explicitly reports tests up to 450 pages/sec and 350 pages/sec for rendered content. www.lumar.io
Botify lets you set a maximum URLs/sec and dynamically adjust it during a crawl. For validated sites there's no product-enforced crawl-speed ceiling; unvalidated sites are limited to 3 pages/sec. Botify says its crawler can exceed 200 pages/sec, but actual throughput depends heavily on the site's response time. support.botify.com
Sitebulb Cloud's Enterprise plan lists 2.5M+ URLs per audit and 5M+ URLs/month, with concurrent crawling. It emphasizes no crawl-credit limits, including for JavaScript crawling. sitebulb.com
Screaming Frog is different: it's a desktop application, so its practical ceiling is your machine's memory/storage and configuration rather than a vendor-side URL quota. Its paid license is currently $279/year for one user, with volume discounts. www.screamingfrog.co.uk
What I'd benchmark
For a large site, don't compare vendor claims directly. Run the same crawl on each platform and record:
- Raw throughput: URLs/sec sustained over the whole crawl.
- Rendered throughput: URLs/sec with JS rendering enabled.
- Time to 1M URLs: especially useful for operational planning.
- Server impact: average response time, 5xx/429 rate, bandwidth.
- Discovery efficiency: URLs discovered vs actually crawled.
- Concurrency controls: maximum requests, threads, rate limits, per-host throttling.
- Crawl ceilings: URLs/crawl, URLs/month, projects, simultaneous crawls.
- Data extraction limits: HTML storage, custom extraction, exports/API quotas.
- Incremental/recrawl behavior: whether you're paying to recrawl unchanged pages.
- True cost per 1M URLs: subscription + overages + rendering + exports + seats.
The last one is particularly important. “$X/month” isn't a useful comparison if one vendor charges by crawl credits/URLs and another gives you unlimited crawling.
A useful normalization
I'd calculate:
Effective cost / million URLs = total monthly platform cost ÷ millions of URLs successfully crawled
and separately:
Hours / million URLs = 1,000,000 ÷ sustained URLs/sec ÷ 3,600
For example, at 100 URLs/sec, 1M URLs takes about 2.8 hours. At 450 URLs/sec, it's about 37 minutes—assuming the target infrastructure can actually supply pages at that rate.
My shortlist
- Fastest cloud crawling: Lumar — strongest published raw-speed number.
- Enterprise SEO/data ecosystem: Botify — particularly interesting if crawl data needs to feed broader SEO analytics.
- Best transparent economics: Sitebulb Cloud — particularly attractive if you're doing lots of crawls because it explicitly avoids crawl credits.
- Lowest software cost / maximum control: Screaming Frog — excellent if you can provision sufficiently powerful machines and don't need a cloud-native crawling platform.
One caveat: vendor maximum speed ≠ your site's achievable speed. Lumar itself says its crawler is as fast as your infrastructure allows, and Botify recommends increasing speed while watching response time and server errors. support.botify.comhelp.lumar.io
If you give me roughly site size (e.g. 5M/50M/500M URLs), JS rendering yes/no, crawl frequency, and monthly budget, I can turn this into a cost + expected crawl-time model for the 3–4 strongest options.
Behind this answer: 4 searches, 30 results retrieved, 4 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://support.botify.com/en/articles/9108632-botify-subscription-plans?utm_source=chatgpt.com | not cited |
| 2 | https://developers.botify.com/docs/limits?utm_source=chatgpt.com | not cited |
| 3 | https://support.botify.com/en/articles/9108543-modifying-crawl-settings?utm_source=chatgpt.com | not cited |
| 4 | https://www.botify.com/platform/botify-analytics/sitecrawler-feature?utm_source=chatgpt.com | not cited |
| 5 | https://support.botify.com/en/articles/9108542-validating-your-website?utm_source=chatgpt.com | not cited |
| 6 | https://support.botify.com/en/articles/9108541-starting-a-crawl?utm_source=chatgpt.com | not cited |
| 7 | https://support.botify.com/en/articles/9108636-frequently-asked-questions?utm_source=chatgpt.com | not cited |
| 8 | https://old.developers.botify.com/api/rate-limit/?utm_source=chatgpt.com | not cited |
| and 22 more results | ||
Tracked brands mentioned
Week of 10th to 16th Aug 2026
Answer 9 of 22, 16th Aug 2026, gpt-5-6, mentions no tracked brands
If the goal is fast, large-site crawling, I’d compare platforms on pages/minute, per-domain throttling, concurrency, and effective cost per successful page—not just the advertised “requests/sec.”
Quick comparison
| Platform | Throughput / limits | Cost model | Best fit |
|---|---|---|---|
| Firecrawl | Standard: 50 concurrent browsers; Growth: 100; Scale: 150+. Crawl API rate limits are 100/min, 1,000/min, 2,000/min respectively. github.com | 1 credit/page for Crawl; Standard $83/mo for 100k pages, Growth $333 for 500k, Scale $599 for 1M when billed annually. www.firecrawl.dev | Easiest “crawl this site” API; strong default choice |
| Zyte API | Standard: 3,000 RPM; Enterprise: 10,000 RPM. Actual concurrency depends on latency and target-site limits. docs.zyte.comwww.firecrawl.dev | Successful-response pricing. HTTP starts at $0.13/1k requests; browser $1.01/1k, with volume discounts. www.zyte.com | High-throughput HTTP crawling + sophisticated anti-bot handling |
| Apify | Business supports up to 256 concurrent Actor runs; actual crawl speed is determined by the Actor implementation/resources. apify.com | Compute-based: $0.20/CU Starter, $0.16 Scale, $0.13 Business, plus optional proxies/etc. apify.com | Maximum control/custom crawlers |
| Oxylabs Web Scraper API | Up to 50 jobs/s on most paid tiers; Business 100 jobs/s. Rendered jobs are lower: 13/s or 25/s. Domains can be throttled to 1 req/s when success rate falls below 40%. developers.oxylabs.iodocs.zyte.com | Result-based. “Other” sites without JS: $1.15/1k on Micro → $0.75/1k Business; JS: $1.35 → $1.00/1k. oxylabs.io | High-volume scraping where unblocking is important |
| Bright Data Web Scraper API | Advertises unlimited concurrency, with batch/scheduled collection. brightdata.comwww.firecrawl.dev | $1.50/1k records PAYG; $499/mo includes 384k records, then $1.30/1k. brightdata.comwww.firecrawl.dev | Very high concurrency / difficult sites |
| Scrapy Cloud | Starter: 1 concurrent crawl; paid capacity scales by units, where 1 unit = 1 GB RAM + 1 concurrent crawl. Paid crawls have unlimited runtime. www.zyte.comdocs.zyte.com | $9/month per unit. www.zyte.comdocs.zyte.com | If you already have a Scrapy crawler and want managed infrastructure |
What I'd benchmark
For a large site, don't benchmark with a handful of URLs. Take a representative 100k–1M URL sample and run the same workload against each provider.
Measure:
- Wall-clock crawl time
time_to_50%time_to_90%time_to_100%
- Effective throughput
- successful pages/minute
- successful pages/second
- separately measure HTML vs JS-rendered pages
- Success rate
- HTTP 2xx
- 3xx/4xx
- timeout
- CAPTCHA/blocked
- empty/incorrect content
- Cost
total provider cost / successful pages- Include retries, browser usage, proxy traffic, extraction, etc.
- This is much more meaningful than nominal $/1k requests.
- Scaling curve
Run at, say, 10, 25, 50, 100, 250, and 500 concurrent requests. The important question is where throughput stops increasing.
- Per-domain behavior
This matters enormously. A provider might advertise 3,000 RPM globally but deliberately throttle a particular target. Zyte explicitly has per-website and account-website limits, while Oxylabs can reduce a poorly performing domain to 1 request/sec. developers.oxylabs.iodocs.zyte.com
My shortlist
For a large, mostly static site, I'd start with:
Zyte vs Firecrawl vs Apify.
- Firecrawl if you want the least engineering and a straightforward site-crawl abstraction.
- Zyte if raw HTTP throughput, sophisticated proxy/unblocking infrastructure, and successful-response billing matter most.
- Apify if you're willing to own more of the crawler logic in exchange for control over concurrency and resource allocation.
For JS-heavy / anti-bot-heavy sites, add Oxylabs and Bright Data to the benchmark. Their higher apparent cost can be worthwhile if they substantially increase successful-page yield.
One particularly important distinction: Firecrawl's pricing is naturally expressed in pages, whereas Apify charges for compute consumed and Zyte/Oxylabs/Bright Data primarily charge around successful results/requests. www.firecrawl.dev That makes a controlled benchmark essential before extrapolating a million-page crawl.
If you give me site size (e.g. 500k/5M/50M URLs), % JS-rendered, average response size, and desired completion time, I can turn this into a concrete capacity + monthly-cost comparison for the top 3–5 options.
Behind this answer: 7 searches, 30 results retrieved
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://docs.zyte.com/zyte-api/pricing.html?utm_source=chatgpt.com | not cited |
| 2 | https://docs.zyte.com/zyte-api/usage/errors.html?utm_source=chatgpt.com | not cited |
| 3 | https://www.zyte.com/zyte-api/?utm_source=chatgpt.com | not cited |
| 4 | https://docs.zyte.com/zyte-api/migration/scrapingbee/index.html?utm_source=chatgpt.com | not cited |
| 5 | https://docs.zyte.com/zyte-api/usage/reference.html?utm_source=chatgpt.com | not cited |
| 6 | https://www.zyte.com/lp-saas-scraping-api/?utm_source=chatgpt.com | not cited |
| 7 | https://docs.zyte.com/zyte-api/migration/zenrows/index.html?utm_source=chatgpt.com | not cited |
| 8 | https://www.zyte.com/network-intercept/?utm_source=chatgpt.com | not cited |
| and 22 more results | ||
Answer 10 of 22, 15th Aug 2026, gpt-5-6, mentions no tracked brands
For a large-site crawl, I’d compare platforms on pages/minute, concurrency/rate limits, success rate, and effective $/1,000 successful URLs—not just advertised crawl speed.
Quick comparison
| Platform | Cost model | Published limits / scaling | Best fit |
|---|---|---|---|
| Apify Website Content Crawler | Compute-based; roughly $0.20/1k pages HTTP and $0.50–$5/1k browser pages in Apify's published tests | Up to 25/32/128/256 concurrent runs on Free/Starter/Scale/Business; actual crawl throughput depends heavily on actor configuration | Flexible, large crawls where you want control |
| Zyte API | Per successful request; HTTP roughly $0.13–$1.27/1k, browser $1.01–$16.08/1k, depending on site complexity | Standard 3,000 RPM; Enterprise 10,000+ RPM/custom; website-specific limits also apply | High-volume crawling where proxy/anti-bot handling matters |
| Scrapy Cloud | $9/month per compute unit; each unit provides 1 GB RAM and 1 concurrent crawl | Add units to run crawls concurrently; jobs have unlimited runtime on paid units | Running your own Scrapy spiders at predictable infrastructure cost |
Apify's current plans are $29/$199/$999 per month for Starter/Scale/Business, with compute at $0.20/$0.16/$0.13 per CU respectively; concurrency is 32/128/256 runs. apify.com
Zyte's standard API limit is 3,000 RPM, but RPM isn't equivalent to concurrency. Zyte explicitly gives concurrency ≈ RPM / 60 × average response time; for example, 3,000 RPM and a 2-second average response gives roughly 100 concurrent requests. docs.zyte.com
The important part: benchmark effective crawl speed
For your use case, I'd run the same 100k–1M URL corpus through each platform and record:
- Wall-clock time
- Start → 95% URLs completed
- Start → 99% completed
- Pages/minute at steady state
- Successful throughput
successful URLs / minute- Don't count 403s, timeouts, retries, etc. as throughput.
- Concurrency
- Maximum sustainable concurrent requests
- Whether increasing concurrency actually increases throughput
- Point where the target starts throttling you
- Cost
- Total platform bill
- Cost / 1,000 successful URLs
- Cost / 1M successful URLs
- Include browser-rendering and proxy costs
- Reliability
- HTTP success %
- timeout %
- 429/403 %
- retries per URL
- duplicate/skipped URLs
- completeness of the discovered URL set
- Resource consumption
- CPU/RAM per crawler
- Browser vs raw HTTP
- bandwidth
- storage
A useful way to normalize the result
Suppose a platform crawls 1M URLs in 2 hours at $300:
Throughput: ~8,333 successful URLs/minute
Effective cost: $0.30 / 1,000 successful URLs
Cost per 1M: $300
That is much more meaningful than saying "the API supports 3,000 RPM," because a nominal 3,000 RPM limit can be irrelevant if the target site responds slowly or starts rate-limiting you.
My shortlist
If speed is the primary objective: test Zyte API Enterprise and Apify with a highly parallel HTTP crawler first. Zyte explicitly offers higher/custom concurrency and performance assurances at Enterprise scale. www.zyte.com
If raw HTTP is enough: Apify can be extremely inexpensive—their Website Content Crawler documentation estimates about $0.20/1,000 pages for raw HTTP under their baseline assumptions. apify.com
If you need JavaScript rendering: cost changes dramatically. Apify estimates roughly $0.50–$5/1,000 browser pages, while Zyte's browser pricing ranges from about $1.01 to $16.08/1,000 requests depending on target complexity. apify.com
If you already have Scrapy expertise: Scrapy Cloud is worth benchmarking because you're essentially buying scalable crawler compute rather than paying a per-request scraping API. Paid units are $9/month each and provide 1 GB RAM plus one concurrent crawl. docs.zyte.com
One caveat: don't benchmark only against an easy site. For a serious production decision, use at least one fast/unprotected site, one JS-heavy site, and one site with meaningful bot protection. The latter can completely reverse the cost/speed ranking.
If you tell me roughly how many URLs (e.g. 1M, 10M, 100M), HTTP vs browser rendering, and whether you control the target site, I can turn this into a concrete $/1M pages + expected crawl time + concurrency comparison.
Behind this answer: 5 searches, 30 results retrieved
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://www.zyte.com/pricing/?utm_source=chatgpt.com | not cited |
| 2 | https://docs.zyte.com/zyte-api/pricing.html?utm_source=chatgpt.com | not cited |
| 3 | https://www.zyte.com/terms-and-services/?utm_source=chatgpt.com | not cited |
| 4 | https://www.zyte.com/ide/?utm_source=chatgpt.com | not cited |
| 5 | https://www.zyte.com/network-intercept/?utm_source=chatgpt.com | not cited |
| 6 | https://www.zyte.com/ai-web-scraping/?utm_source=chatgpt.com | not cited |
| 7 | https://www.zyte.com/lp-zyte-api/?utm_source=chatgpt.com | not cited |
| 8 | https://docs.zyte.com/zyte-api/usage/reference.html?utm_source=chatgpt.com | not cited |
| and 22 more results | ||
Answer 11 of 22, 14th Aug 2026, gpt-5-6, mentions 3 tracked brands
If your priority is fast crawling of a large site, I’d compare platforms on effective URLs/hour, not the vendor’s headline “URLs/sec.” The latter is usually a ceiling under ideal conditions.
Quick comparison
| Platform | Published crawl speed | Scale / limits | Cost model | Best fit |
|---|---|---|---|---|
| Lumar | Up to 450 URLs/sec non-rendered; 350/sec rendered | Billing includes a set URL allowance; active-project limits are plan-dependent | Custom enterprise pricing | Maximum enterprise-scale speed + sophisticated crawling |
| JetOctopus | Up to 250 pages/sec | Can handle 100M+ URLs; Unlimited plan has no crawl cap and no simultaneous-crawl cap | Site/volume-based; published plans + custom Unlimited | Very large sites where crawl volume is unpredictable |
| Sitebulb Cloud | Hundreds of URLs/sec | Millions of URLs; unlimited projects/domains; no crawl credits | Flat monthly subscription | Strong price/performance for frequent large crawls |
| Screaming Frog | Configurable request rate/threads; fundamentally machine-bound | Paid version defaults to 5M URLs, but can exceed that with sufficient RAM/SSD | $279/user/year | Cheapest option when you can dedicate a powerful machine |
Lumar's published maximum is 450 URLs/sec for normal crawls and 350 URLs/sec for rendered crawls. www.lumar.iohelp.lumar.io JetOctopus advertises up to 250 pages/sec and says it can handle 100M+ URLs. jetoctopus.comjetoctopus.comwww.lumar.io Sitebulb Cloud is explicitly designed for millions of URLs without crawl credits, while Screaming Frog's performance depends heavily on your local hardware. sitebulb.comwww.screamingfrog.co.uk
What I'd benchmark
For a genuinely large site, run the same crawl against the same URL set on each platform and record:
- Time to first 100K URLs
- Time to 1M URLs
- Total URLs/hour
- URLs/hour with JavaScript rendering
- HTTP requests/sec actually hitting your infrastructure
- Peak CPU/RAM/server load
- Completion rate — e.g. 99.8% of URLs successfully processed
- Cost per 1M URLs
- Cost per 1M rendered URLs
- Time spent exporting/processing results
The important distinction is crawler throughput vs. site throughput. Lumar, for example, says it can go as fast as your infrastructure allows, but also warns that excessive crawl rates can overwhelm the site's server. www.lumar.iohelp.lumar.io Screaming Frog similarly lets you specify threads or URLs/sec and recommends agreeing on a crawl rate because higher concurrency can affect server response times. sitebulb.comwww.screamingfrog.co.uk
Cost comparison
Screaming Frog is dramatically cheaper if a local crawl works for you: $279/year for one license, with the paid crawler described as having an unlimited crawl limit subject to machine resources. www.screamingfrog.co.uk A sufficiently provisioned machine can reportedly handle around 10M URLs with 16 GB RAM and a 500 GB SSD, although actual performance depends heavily on configuration and page complexity. sitebulb.comwww.screamingfrog.co.uk
Sitebulb Cloud is interesting if you're crawling frequently. Its model is deliberately flat monthly pricing with no crawl credits, so repeated crawls don't consume an additional crawl allowance. It also says JavaScript crawling at millions of URLs doesn't incur an extra charge. sitebulb.com Its enterprise offering starts at £95/month. sitebulb.com
JetOctopus is attractive when volume gets into the multi-million range. Its published pricing example is $383/month for 1M crawl pages, while its Unlimited enterprise plan removes the crawl cap and supports 10M, 50M, 100M+ pages as often as needed. jetoctopus.com
Lumar is harder to price from public information: it uses tailored proposals, with URL allowances and active-project limits determined according to your requirements. www.lumar.io That makes it harder to compare purely on sticker price, but potentially worthwhile when crawler speed and enterprise workflow matter most.
My shortlist
If you're talking about something like 5–50M URLs, I'd test:
1. JetOctopus — probably the most interesting price/scale combination, particularly if you need repeated crawls.
2. Lumar — strongest candidate if raw cloud crawling speed and enterprise capabilities are the priority.
3. Sitebulb Cloud — particularly compelling if you want predictable costs and lots of recurring crawls.
4. Screaming Frog — benchmark it anyway. If a well-provisioned machine can meet your SLA, its economics are hard for cloud platforms to beat.
One particularly useful metric is:
Effective cost per million successfully crawled URLs = total platform cost ÷ completed millions of URLs
Then calculate it separately for HTML-only and JS-rendered crawls. That will expose cases where a platform looks cheap until rendering, crawl quotas, or repeated crawls enter the equation.
If you tell me roughly how many URLs (e.g. 1M, 10M, 100M), whether JS rendering is required, and how often you need to recrawl, I can turn this into a concrete cost + expected crawl-time model for the platforms above.
Behind this answer: 4 searches, 30 results retrieved, 3 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://help.lumar.io/hc/en-us/articles/23461990979089-How-to-Crawl-a-Website-Accessibility?utm_source=chatgpt.com | not cited |
| 2 | https://help.lumar.io/hc/en-us/articles/23471268960785-How-to-Crawl-a-Website-SEO?utm_source=chatgpt.com | not cited |
| 3 | https://www.lumar.io/product-guides/how-to-crawl/accessibility-crawl-setup/?utm_source=chatgpt.com | not cited |
| 4 | https://www.lumar.io/product-guides/how-to-crawl/how-crawl-website-with-deepcrawl/?utm_source=chatgpt.com | not cited |
| 5 | https://www.lumar.io/product-guides/how-to-crawl/limiting-size-depth-crawl/?utm_source=chatgpt.com | not cited |
| 6 | https://www.lumar.io/product-guides/how-to-crawl/using-crawl-limits-reduce-credit-usage/?utm_source=chatgpt.com | not cited |
| 7 | https://www.lumar.io/product-guides/how-to-crawl/establishing-a-crawling-strategy-for-maximum-efficiency/?utm_source=chatgpt.com | not cited |
| 8 | https://www.lumar.io/spdr/?utm_source=chatgpt.com | not cited |
| and 22 more results | ||
Tracked brands mentioned
Answer 12 of 22, 13th Aug 2026, gpt-5-6, mentions 4 tracked brands
If the goal is fast, repeatable crawls of a large site, I’d compare platforms on throughput + hard limits + all-in crawl cost, not just the advertised “URLs/sec.”
Quick comparison
| Platform | Speed / scale | Crawl limits | Cost model | Best fit |
|---|---|---|---|---|
| Screaming Frog | Hardware-dependent; can run millions of URLs. Default 5 threads, configurable concurrency/URLs-sec. www.screamingfrog.co.uk | Paid version shows 5M default but can exceed it with appropriate hardware; roughly 2M is a recommended range. www.screamingfrog.co.uk | $279/user/year; effectively unlimited URLs, but you supply compute/storage. www.screamingfrog.co.uk | Cheapest if you can run/manage your own crawler |
| Sitebulb Cloud | Advertises 300+ URLs/sec at the high end; cloud resources remove local-machine bottlenecks. sitebulb.com | Designed for millions of URLs; no crawl-credit limits according to its cloud comparison. sitebulb.com | Fixed subscription rather than per-crawl credits; exact plan depends on volume. | Fast cloud crawling without infrastructure work |
| JetOctopus | Advertises up to 250 pages/sec. jetoctopus.com | Plans are based on crawl-page volume; example pricing page shows 1M crawl pages/month or 500K JS pages. jetoctopus.com | Volume subscription; JS pages consume more allowance. | High-volume cloud SEO crawling |
| Enterprise crawlers such as Lumar/Botify/Oncrawl | Generally cloud/distributed and designed for very large sites | Usually negotiated/custom limits | Enterprise/custom pricing | Multi-million/billion URL estates, teams, integrations |
One interesting benchmark for Screaming Frog self-hosted: its own cloud guide reports a 3.1M-URL crawl in about two days on an 8-vCPU/32GB/200GB VM, at deliberately conservative speed, for under £20 of GCP compute in that example. www.screamingfrog.co.uk
The important catch: URLs/sec isn't directly comparable
A vendor saying 250 URLs/sec doesn't necessarily beat one doing 50.
Measure these separately:
- HTML URLs/sec — actual pages successfully fetched.
- Total HTTP requests/sec — critical for server load, especially JS rendering.
- Concurrent connections — how aggressively the crawler hits your origin.
- TTFB-adjusted throughput — otherwise a fast crawler on a fast site looks artificially better.
- JS crawl throughput — headless Chrome can be dramatically slower and generate many more requests. Sitebulb explicitly notes that Chrome crawls fetch page resources, so server request counts can greatly exceed crawled URLs. support.sitebulb.com
- Time to complete 1M URLs — probably your most useful business metric.
- Cost per million URLs — including compute, storage, licenses, and JS multipliers.
- Repeat-crawl economics — especially important if you're crawling daily/weekly.
A good apples-to-apples benchmark
For your site, I'd run each candidate against the same 100k–500k representative URL sample:
- 80% ordinary HTML pages
- 10% redirects/errors
- 10% JS-heavy pages
- Same geographic region
- Same robots.txt behavior
- Same crawl depth and URL discovery rules
- Same concurrency target
- No caching if you're testing origin load
Then record:
URLs discovered
URLs successfully crawled
Wall-clock time
HTML URLs/sec
HTTP requests/sec
Peak concurrency
Average TTFB
5xx / 429 rate
RAM/CPU consumed
Total crawl cost
Cost / 1M URLs
I'd actually make “cost per 1M successfully crawled HTML URLs” the headline metric.
For example, if Platform A advertises 250 URLs/sec but your site throttles it to 40, while Platform B achieves 80 without generating 429s, B is effectively 2× faster for your workload.
My shortlist
For a large but conventional SEO crawl, I'd start with Screaming Frog vs Sitebulb Cloud vs JetOctopus.
- Lowest raw software cost: Screaming Frog.
- Best cloud/no-infrastructure approach: Sitebulb Cloud.
- Very high-volume cloud crawling: JetOctopus.
- Multi-million URL enterprise operation with logs/data integrations: evaluate the enterprise platforms rather than optimizing around a desktop crawler.
And don't benchmark at maximum speed initially. Agree on an origin-safe rate, then increase concurrency until you find the site's actual saturation point. Screaming Frog and Sitebulb both explicitly expose controls for threads/concurrency and URLs-per-second because server response time and server-side limits ultimately constrain throughput. www.screamingfrog.co.uk
If you give me rough site size (e.g. 5M/50M/500M URLs), HTML vs JS percentage, crawl frequency, and whether you need logs/API data, I can turn this into a $/1M URLs + expected completion-time comparison for the platforms that fit.
Behind this answer: 7 searches, 30 results retrieved, 4 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://support.sitebulb.com/en/articles/9853652-crawler-settings?utm_source=chatgpt.com | not cited |
| 2 | https://sitebulb.com/?utm_source=chatgpt.com | not cited |
| 3 | https://sitebulb.com/subscriptions/pricing/server/?annual=false&utm_source=chatgpt.com | not cited |
| 4 | https://sitebulb.com/subscriptions/pricing/index/?annual=false&utm_source=chatgpt.com | not cited |
| 5 | https://sitebulb.com/use-cases/enterprise/?utm_source=chatgpt.com | not cited |
| 6 | https://support.sitebulb.com/en/articles/9857604-concurrent-crawling?utm_source=chatgpt.com | not cited |
| 7 | https://sitebulb.com/cloud/?utm_source=chatgpt.com | not cited |
| 8 | https://sitebulb.com/learn/?utm_source=chatgpt.com | not cited |
| and 22 more results | ||
Tracked brands mentioned
Answer 13 of 22, 12th Aug 2026, gpt-5-6, mentions 4 tracked brands
If the goal is fast, repeatable crawls of a large site, I’d compare platforms on throughput + volume economics, not just the advertised “pages/sec.”
Quick comparison
| Platform | Crawl speed / scale | Crawl limits | Cost model | Best fit |
|---|---|---|---|---|
| Screaming Frog | Desktop/local; speed depends heavily on your CPU, RAM, network and configuration | Free: 500 URLs; paid removes that cap, but machine resources become the practical limit | $279/year for a license currently www.screamingfrog.co.uk | Cheapest option for technical crawling when you can run infrastructure yourself |
| Sitebulb Cloud | Cloud-based; configurable threads + URLs/sec | Up to 10M URLs/audit on Cloud; Desktop typically 500K, configurable higher | Subscription; positioning emphasizes unlimited crawling rather than crawl credits | Strong middle ground for large SEO audits sitebulb.com |
| JetOctopus | Up to 250 pages/sec; says 1M URLs can take ~1–2 days at 5 threads, potentially ~1 hour at high speed | Claims support for 100M+ URLs | Volume-based crawl pricing; example plan shown at 1M pages/month; enterprise unlimited available | Speed/scale leader if your origin can tolerate aggressive crawling jetoctopus.comjetoctopus.com |
| Oncrawl | User-configurable URLs/sec; recommends roughly 2–5 URLs/sec for many sites and allows >10 after domain validation | Configurable max URLs/depth; plan-dependent | Enterprise/quote-oriented | Controlled cloud crawling + SEO analytics help.oncrawl.com |
| Botify | Configurable maximum pages/sec; throttles based on server behavior | Max URLs, depth and subscription limits | Enterprise subscription | Large enterprise sites where crawling is part of a broader SEO data platform support.botify.comsupport.botify.com |
| Lumar | Cloud enterprise crawler | Enterprise-scale; exact limits/package depend on contract | Custom quote | Enterprise technical SEO / website intelligence www.lumar.io |
The important catch: advertised speed ≠ actual crawl speed
For a site with, say, 10M URLs, I'd benchmark:
Effective throughput = URLs successfully fetched / wall-clock hour
rather than accepting a vendor's maximum.
For example, a crawler advertising 250 pages/sec theoretically processes 900,000 URLs/hour, but your actual number could be much lower because of:
- server response time / TTFB
- crawler concurrency
- robots/rate limiting
- redirects
- JS rendering
- asset fetching
- DNS/TLS latency
- crawler-side processing
- intentional throttling when your origin starts returning errors
Sitebulb explicitly notes that crawl speed depends on threads/Chrome instances, URL/sec limits, TTFB and rendering time. support.sitebulb.com Botify likewise says it may intentionally slow down when necessary to avoid overwhelming the server. support.botify.com
How I'd run the bake-off
Give every vendor the same representative 100K–500K URL slice first.
Measure:
- Wall-clock completion time
- Sustained URLs/sec, not peak
- HTTP success rate
- 5xx/429 rate
- Bandwidth consumed
- JS-rendered URLs/sec, separately from HTML
- CPU/RAM consumed if self-hosted
- Data captured per URL
- Export/API throughput
- Cost per 1M URLs
Then extrapolate:
Annual crawl cost = (URLs per crawl × crawls per year ÷ included monthly/annual volume) × effective price
This exposes a big difference between “unlimited crawls” and “X crawl pages included.”
My shortlist
If you're talking about millions to tens of millions of URLs, I'd test these three first:
- JetOctopus — strongest published raw-speed/scale story; it claims up to 250 pages/sec and 100M+ URL capability. jetoctopus.comjetoctopus.com
- Sitebulb Cloud — attractive if you want a polished SEO crawler without relying on a workstation; published limit is 10M URLs/audit. sitebulb.com
- Botify — better candidate if you need enterprise SEO analytics, log data and crawling integrated together rather than simply the fastest crawler. support.botify.com
Screaming Frog is the control/baseline I'd include because its licensing is dramatically cheaper, but you're effectively buying the crawler software and supplying the compute rather than buying a cloud crawling service. www.screamingfrog.co.uk
If you tell me roughly site size (e.g. 5M/50M/100M URLs), crawls per day/week/month, whether JS rendering is required, and whether you need SEO analysis or just raw HTML extraction, I can turn this into a $/1M URLs + expected crawl-time comparison for the platforms.
Behind this answer: 9 searches, 30 results retrieved, 4 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://support.sitebulb.com/en/articles/9853652-crawler-settings?utm_source=chatgpt.com | not cited |
| 2 | https://support.sitebulb.com/en/articles/10056982-how-to-crawl-large-websites?utm_source=chatgpt.com | not cited |
| 3 | https://sitebulb.com/?utm_source=chatgpt.com | not cited |
| 4 | https://sitebulb.com/subscriptions/pricing/index/?annual=false&utm_source=chatgpt.com | not cited |
| 5 | https://sitebulb.com/features/?utm_source=chatgpt.com | not cited |
| 6 | https://support.sitebulb.com/en/articles/9491746-understanding-sitebulb-crawl-speed?utm_source=chatgpt.com | not cited |
| 7 | https://support.sitebulb.com/en/articles/9491717-how-to-crawl-a-url-list?utm_source=chatgpt.com | not cited |
| 8 | https://support.sitebulb.com/en/articles/9853811-include-exclude-urls-settings?utm_source=chatgpt.com | not cited |
| and 22 more results | ||
Tracked brands mentioned
Answer 14 of 22, 11th Aug 2026, gpt-5-6, mentions no tracked brands
If the goal is fast full-site crawling at large scale, I’d compare platforms on effective pages/minute, concurrency, coverage/success rate, and cost per successful page—not just the advertised API rate.
Quick comparison
| Platform | Speed / concurrency | Cost model | Best fit |
|---|---|---|---|
| Firecrawl | 50 concurrent on Standard, 100 Growth, 150 Scale; enterprise custom | Credit/page | Fastest path to managed whole-site crawling |
| Apify | Up to 32/128/256 concurrent Actor runs on Starter/Scale/Business | Compute + proxy/storage/etc. | Maximum control and custom crawlers |
| Zyte | 3,000 RPM standard; 10,000 RPM enterprise | Successful response, target-dependent | High-volume fetching + difficult sites |
| Bright Data | Unlimited concurrency on Web Scraper API | Successful record | Very high concurrency / anti-bot requirements |
Current published limits/prices vary substantially. Firecrawl's Standard plan is 100k pages and 50 concurrent requests for $83/month when billed annually; Growth is 500k pages/100 concurrent for $333; Scale is 1M/150 concurrent for $599. Crawl is charged at 1 credit/page. www.firecrawl.dev
Apify is more infrastructure-oriented: its current plans provide 32, 128, and 256 max concurrent Actor runs on Starter, Scale, and Business, respectively. Compute is $0.20/CU on the lower plans, falling to $0.13/CU on Business, with other usage such as proxies and transfer potentially adding cost. apify.com
Zyte is particularly interesting if you're fetching raw pages rather than needing a complete crawler abstraction. Its standard API limit is 3,000 requests/minute, while enterprise goes to 10,000 RPM. It only charges for successful responses, and pricing depends on the target and whether HTTP or browser rendering is required. docs.zyte.comwww.zyte.com
Bright Data's Web Scraper API advertises unlimited concurrency, with pay-as-you-go pricing of $1.50/1,000 successful records and a $499/month Scale tier including 384k records. brightdata.comdocs.zyte.com
The important part: benchmark them yourself
For a large site, I'd run the same 10k–100k URL corpus through 3 candidates and measure:
- Wall-clock crawl time
pages_completed / elapsed_minutes- Measure p50 and p95 page latency separately.
- Effective throughput
- Don't use maximum advertised concurrency.
- Record actual pages/minute at concurrency 10, 25, 50, 100, etc.
- Find where throughput stops increasing or error rates rise.
- Successful-page cost
total_platform_cost / successfully_crawled_pages- This is much more meaningful than nominal $/1k requests.
- Coverage
% of URLs successfully fetched- Redirects, 403/429s, timeouts, JS failures, duplicate handling.
- Crawler overhead
- Does the service discover links/sitemaps for you?
- Can you constrain paths?
- Can you resume failed jobs?
- Can it stream results instead of waiting for the entire crawl?
- Rendering percentage
- Test both HTTP-only and browser-rendered modes.
- Browser crawling can radically change economics. Zyte, for example, currently lists browser-rendered requests at roughly 8–13× its HTTP pricing depending on target tier. docs.zyte.comwww.zyte.com
My shortlist
If you need the simplest high-speed managed crawler: start with Firecrawl. Its /crawl endpoint handles discovery, rendering and processing, and the company says its infrastructure can scale to millions of URLs/day. www.firecrawl.dev
If you have engineering resources and want maximum control: use Apify. You can tune crawler behavior and concurrency much more deeply, but you need to model compute/proxy/storage costs rather than assuming a simple per-page price. apify.com
If the core problem is fetching huge numbers of pages quickly: I'd benchmark Zyte and Bright Data very seriously. Their request/concurrency models can be better suited to raw acquisition at very high volume. brightdata.comdocs.zyte.com
For example, if you tell me site size (e.g. 1M/10M/100M URLs), % JavaScript-rendered, desired recrawl frequency, and your monthly budget, I can turn this into a concrete 10M-page crawl cost + estimated completion-time comparison across these platforms.
Behind this answer: 9 searches, 30 results retrieved
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://apify.com/apify/web-scraper?utm_source=chatgpt.com | not cited |
| 2 | https://apify.com/mikolabs/website-content-crawler?utm_source=chatgpt.com | not cited |
| 3 | https://help.apify.com/en/articles/5136728-subscribing-to-the-apify-platform?utm_source=chatgpt.com | not cited |
| 4 | https://help.apify.com/en/articles/11079614-reaching-platform-limits?utm_source=chatgpt.com | not cited |
| 5 | https://apify.com/apify/website-content-crawler?fpr=mrbridge&utm_source=chatgpt.com | not cited |
| 6 | https://help.apify.com/en/articles/13671754-navigating-apify-billing?utm_source=chatgpt.com | not cited |
| 7 | https://help.apify.com/en/collections/4082422-pricing-billing?utm_source=chatgpt.com | not cited |
| 8 | https://help.apify.com/en/articles/7959177-subscription-plan-changes-and-proration?utm_source=chatgpt.com | not cited |
| and 22 more results | ||
Answer 15 of 22, 10th Aug 2026, gpt-5-6, mentions 3 tracked brands
For large-site crawling, compare platforms by throughput, crawl controls, resource limits, and total cost of ownership rather than just “pages per minute.” A good evaluation framework looks like this:
1. Define the crawl workload first
Before comparing tools, normalize the test case:
| Variable | Example |
|---|---|
| Site size | 5M URLs |
| Crawl type | Full crawl vs incremental |
| Rendering | HTML only vs JavaScript rendering |
| Assets | HTML only vs images/CSS/JS |
| Depth | Unlimited vs max depth |
| Frequency | Weekly, daily, continuous |
| Output | URLs only vs full SEO data/extractions |
| Concurrency | Allowed requests/sec |
A crawler that is fastest on raw HTML may be much slower once you enable JS rendering, screenshots, extraction rules, or compliance throttling.
2. Compare crawl speed metrics
Measure:
Pages per second (PPS)
Basic throughput:
Pages crawled / elapsed time
Example:
- 1M URLs in 2 hours
- 138 pages/sec
Requests per second (RPS)
More useful when pages contain many resources:
HTTP requests / second
Time to first million URLs
A practical benchmark:
| Platform | Test |
|---|---|
| Start crawl | T=0 |
| Reach 100k URLs | time |
| Reach 1M URLs | time |
| Complete crawl | time |
Sustained throughput
Some tools burst quickly but throttle later. Test a 6–24 hour crawl.
3. Compare crawl limits
Important limits to capture:
| Limit | Why it matters |
|---|---|
| Max URLs/month | Cost driver |
| Concurrent workers | Determines speed |
| Requests/sec controls | Avoids blocking |
| Domain limits | Important for multi-site crawling |
| Storage limits | Large crawls create huge datasets |
| Export/API limits | Can bottleneck pipelines |
| JS rendering quotas | Often expensive |
| Proxy bandwidth | Major cost factor |
4. Cost comparison model
Do not compare only subscription price. Calculate:
Cost per million URLs crawled
Formula:
(total monthly cost) / (millions of URLs crawled)
Example:
| Platform | Monthly cost | Crawl volume | Cost / 1M URLs |
|---|---|---|---|
| A | $500 | 50M URLs | $10 |
| B | $1,500 | 300M URLs | $5 |
| C | $800 | 20M URLs | $40 |
5. Categories of platforms to compare
Enterprise SEO crawlers
Good for technical SEO audits:
- Screaming Frog — strong desktop crawler; large crawls depend on machine resources and configuration.
- Botify — built for enterprise-scale SEO crawling and log analysis.
- Lumar — enterprise technical SEO crawling and monitoring.
Cloud crawler APIs / infrastructure
Good for custom pipelines:
- Apify — actor-based crawling platform with usage-based pricing.
- Scrapy Cloud — managed deployment for Scrapy crawlers.
- Bright Data — web data collection infrastructure with large-scale proxy capabilities.
Build-your-own
For very large sites:
- Scrapy + distributed workers
- Apache Kafka / queues
- Kubernetes workers
- Object storage (S3-compatible)
- Custom URL frontier
This often wins on cost at hundreds of millions of URLs, but requires engineering.
6. Benchmark methodology
Run the same crawl against each platform:
Test site sets
Use:
- 100k URLs
- 1M URLs
- 10M+ URLs
Record:
Start time
End time
URLs discovered
URLs fetched
Errors
Blocked requests
Average latency
CPU/memory usage
Storage generated
Total cost
Example scorecard:
| Metric | Weight |
|---|---|
| Crawl speed | 30% |
| Cost per million URLs | 25% |
| Crawl controls | 15% |
| Data extraction | 10% |
| Reliability | 10% |
| API/export | 10% |
7. Watch for hidden costs
Common surprises:
- JS rendering can multiply costs 5–20×
- Proxy bandwidth becomes the biggest bill
- Storing raw HTML for millions of pages is expensive
- Re-crawling unchanged URLs wastes budget
- Poor URL deduplication causes crawler explosions
- Export pipelines become the bottleneck
Practical recommendation
For a large production site (millions to hundreds of millions of URLs), I would benchmark three approaches:
- Enterprise crawler (fastest time-to-value)
- Cloud crawler API (flexible cost scaling)
- Custom distributed crawler (lowest unit cost at massive scale)
The winner is usually determined by your required crawl frequency and whether you need SEO auditing or a custom data pipeline. A daily 10M URL crawl and a one-time 500M URL crawl have very different optimal choices.
Behind this answer: no web searches, 3 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Week of 3rd to 9th Aug 2026
Answer 16 of 22, 9th Aug 2026, gpt-5-6, mentions 2 tracked brands
For large-site crawling, compare platforms like an infrastructure benchmark, not just “pages per second.” Crawl speed is usually limited by concurrency, politeness rules, rendering requirements, network, and how much processing happens per URL.
A useful comparison framework:
| Dimension | What to measure | Why it matters |
|---|---|---|
| Raw crawl throughput | URLs/minute or pages/hour | How quickly you can finish a full crawl |
| Concurrency limits | Max parallel requests, threads, workers | Determines ceiling speed |
| Rate limits / politeness controls | Requests/sec per domain, delay settings | Prevents blocking and throttling |
| Distributed crawling | Number of crawler nodes/workers | Important for millions+ of URLs |
| JavaScript rendering | Headless browser support, render concurrency | Huge impact on modern sites |
| Queue management | URL frontier size, prioritization, retries | Critical for large sites |
| Memory/storage scaling | URLs stored, crawl history, logs | Matters for recurring crawls |
| API limits | Export/query limits | Affects automation workflows |
| Cost model | Per URL, per crawl, per seat, compute usage | Determines total cost |
Benchmark methodology
Use the same test site or URL set:
1. Define crawl scenarios
Example:
- 100k static HTML pages
- 1M-page ecommerce site
- 100k JS-rendered pages
- 10k pages with heavy images/assets
- Recrawl of changed URLs only
Measure:
- Time to completion
- Successful URLs fetched
- Errors/timeouts
- CPU/RAM/network usage
- Cost per million URLs
A simple metric:
Cost per 1M crawled URLs =
(total crawl cost) / (URLs successfully processed / 1,000,000)
and:
Effective crawl speed =
successful URLs / total crawl time
2. Separate fetch speed from processing speed
Two tools may both claim “1M pages crawled,” but one may spend more time on:
- HTML parsing
- SEO extraction
- screenshots
- JS rendering
- link discovery
- duplicate detection
- indexing
Run two tests:
Fetch-only benchmark
- Download pages
- Minimal parsing
Full pipeline benchmark
- Crawl + extract + store all required data
The second is usually what matters operationally.
3. Compare platform categories
Self-hosted crawlers
Examples:
- Apache Nutch
- Scrapy
Pros:
- Maximum control
- Scale horizontally
- Potentially lowest cost at very large volumes
Cons:
- You own infrastructure
- Need engineering effort
SEO crawling tools
Examples:
- Screaming Frog
- Sitebulb
Pros:
- Easy setup
- Strong reports
Cons:
- Often designed for audits rather than continuous billion-page crawling
Cloud crawler platforms
Examples:
- Apify
- Bright Data
Pros:
- Managed scaling
- Distributed infrastructure
Cons:
- Cost can grow quickly with volume
4. Track hidden costs
The advertised crawl price is often not the full cost.
Include:
Infrastructure
- crawler servers
- proxies
- bandwidth
- storage
- databases
- queues
Operational
- monitoring
- retries
- IP reputation management
- debugging
Engineering
- maintaining parsers
- handling site changes
- scaling failures
Example comparison sheet
| Platform | Crawl speed | Scaling | JS support | Pricing model | Best fit |
|---|---|---|---|---|---|
| Self-hosted crawler | Very high (with engineering) | Excellent | Depends | Infrastructure cost | Large custom crawls |
| SEO crawler app | Medium | Limited | Moderate | License | Audits |
| Managed crawler API | High | Excellent | Usually strong | Usage-based | Fast deployment |
For a large site, I would benchmark with these numbers
- 10M URLs
- 50M URLs
- 500M URLs
Record:
- Maximum sustainable URLs/sec
- Error rate at that speed
- Cost per 10M URLs
- Cost per recurring daily crawl
- Time to recover from failures
- Duplicate/changed URL handling
The winner is usually not the platform with the highest peak speed—it is the one with the lowest cost per successfully processed URL at your required freshness window.
Behind this answer: no web searches, 2 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Answer 17 of 22, 8th Aug 2026, gpt-5-6, mentions 3 tracked brands
If the goal is fast crawling of a large site, I’d compare platforms on actual throughput, scale limits, and cost per million URLs, rather than just the advertised “crawl speed.”
Quick comparison
| Platform | Speed / scale | Crawl limits | Pricing model | Best fit |
|---|---|---|---|---|
| Screaming Frog SEO Spider | Very fast locally; speed depends heavily on your hardware and the site's connection limits | Paid version advertises unlimited*, but practical scale depends on RAM/SSD; SF says ~2M URLs with 4 GB allocated + SSD is a reasonable setup | £199/user/year currently | Lowest-cost high-control crawling |
| Sitebulb | Configurable concurrency; Cloud avoids tying performance to your workstation | Plan-dependent URL limits; Cloud aimed at large-scale crawling | Subscription; Cloud pricing is customized/plan-based | Teams wanting cloud crawling + strong analysis |
| JetOctopus | Advertises up to 250 pages/sec | Packages based on monthly crawl-page volume; enterprise no-limit options | Volume-based subscription | Very large sites / cloud speed |
| Lumar | Enterprise cloud crawler | Enterprise/custom limits | Custom quote | Enterprise crawling, automation and governance |
\*Screaming Frog's “unlimited” is not literally an infinite-performance guarantee: its practical URL capacity depends on allocated memory and storage. Its documentation says the default limit is 5M URLs but that it can crawl substantially more with appropriate hardware. www.screamingfrog.co.ukwww.screamingfrog.co.ukwww.screamingfrog.co.uk
The important speed distinction
Screaming Frog: you control concurrency. Its default is 5 threads, but you can increase threads or specify a maximum URL-request rate. The limiting factor is frequently the site/CDN/server, not the crawler itself. www.screamingfrog.co.ukwww.screamingfrog.co.uk
Sitebulb: similarly lets you control maximum URLs/sec and concurrency. Its HTML crawler is the fastest mode; Chrome/JS crawling is substantially slower because it has to render resources. support.sitebulb.com
JetOctopus: is the standout if your requirement is literally maximum cloud throughput. Its current pricing page advertises crawling up to 250 pages/sec, cloud execution, no domain/user limits, and no crawling-capacity limit. The catch is that your subscription is still based on purchased crawl volume. jetoctopus.com
Cost comparison that matters
For a large site, I'd normalize everything to:
Cost per 1M URLs crawled
For example, suppose you have a 10M-page site and want to crawl it weekly:
- 10M × 4.33 ≈ 43.3M URLs/month
- Then add a separate multiplier if JS rendering is required.
- Compare the resulting monthly/annual platform cost.
This is much more meaningful than comparing monthly subscription prices.
JetOctopus explicitly prices crawl volume and currently shows packages such as 1M crawl pages/month, with JS pages counted differently; it also offers larger/custom volumes and an enterprise no-limit package. jetoctopus.com
Screaming Frog is unusually inexpensive at this scale if you can run it yourself: the current license is £199/year per user, and the vendor recommends SSD/database storage for multi-million-URL crawls. www.screamingfrog.co.ukwww.screamingfrog.co.ukwww.screamingfrog.co.uk
How I'd benchmark them
Don't use the vendors' maximum numbers directly. Run the same 1M-URL crawl against each platform and record:
- URLs/minute
- Total crawl duration
- HTTP requests/sec
- Peak CPU/RAM
- 5xx/429 rate
- Average TTFB
- JS vs non-JS throughput
- Cost per 1M URLs
- Cost per 1M rendered URLs
- Whether multiple crawls can run concurrently
I'd run three tests:
- HTML-only: measure raw crawler throughput.
- JS rendering: measure realistic expensive crawling.
- Throttled: e.g. 10–20 requests/sec, to see how much throughput you can actually use without stressing production.
That last test is important: a crawler capable of 250 pages/sec isn't necessarily useful if your origin/CDN starts returning 429/503 responses at 30 pages/sec.
My shortlist
For a 10M+ URL site, I'd start with:
1. JetOctopus — if maximum cloud throughput and minimizing crawl infrastructure are the priority.
2. Screaming Frog — if you want the lowest cost and can dedicate a well-provisioned machine; its documented multi-million-URL capability is strong. www.screamingfrog.co.ukwww.screamingfrog.co.uk
3. Sitebulb Cloud — if collaboration/reporting and cloud operation matter as much as raw throughput. support.sitebulb.com
4. Lumar — if this is an enterprise procurement where workflow, automation, security, and support justify custom pricing. www.lumar.io
If you give me site size (e.g. 5M/50M/500M URLs), HTML vs JS %, desired crawl frequency, and whether self-hosting is acceptable, I can turn this into a concrete $/month + estimated crawl time comparison.
Behind this answer: 5 searches, 30 results retrieved, 3 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://sitebulb.com/faqs/?utm_source=chatgpt.com | not cited |
| 2 | https://support.sitebulb.com/en/articles/9491746-understanding-sitebulb-crawl-speed?utm_source=chatgpt.com | not cited |
| 3 | https://sitebulb.com/?utm_source=chatgpt.com | not cited |
| 4 | https://support.sitebulb.com/en/articles/10056982-how-to-crawl-large-websites?utm_source=chatgpt.com | not cited |
| 5 | https://sitebulb.com/features/?utm_source=chatgpt.com | not cited |
| 6 | https://sitebulb.com/subscriptions/pricing/index/?annual=false&utm_source=chatgpt.com | not cited |
| 7 | https://sitebulb.com/features/crawl-huge-websites/?utm_source=chatgpt.com | not cited |
| 8 | https://sitebulb.com/cloud/?utm_source=chatgpt.com | not cited |
| and 22 more results | ||
Tracked brands mentioned
Answer 18 of 22, 7th Aug 2026, gpt-5-5, mentions 2 tracked brands
When comparing platforms for fast crawling of a large website (100k–10M+ pages), focus on measurable throughput rather than marketing claims. Most services advertise "fast," but actual performance depends on concurrency limits, rate limiting, JavaScript rendering, and pricing.
Here's a practical comparison framework:
| Factor | Why it matters | Questions to ask |
|---|---|---|
| Crawl speed | Pages per second/minute/hour | What's the maximum sustained crawl rate? |
| Concurrency | Parallel requests | Can I set 100, 500, or 1,000 concurrent requests? |
| Rate limiting | Avoid overwhelming servers | Can I specify requests/sec per domain? |
| JavaScript rendering | Major impact on speed | Can JS be disabled? Is rendering optional? |
| Distributed crawling | Scale to millions of pages | Is crawling spread across multiple workers? |
| Resume support | Large crawls often fail | Can interrupted crawls continue? |
| API throughput | Export results quickly | Are API rate limits documented? |
| Pricing model | Predictable cost | Per page, per crawl, per compute hour, or monthly? |
Typical platform characteristics
| Platform type | Speed | Limits | Cost model | Best for |
|---|---|---|---|---|
| Self-hosted crawlers (Scrapy, StormCrawler) | Very high | Hardware-dependent | Infrastructure only | Millions of pages |
| Commercial cloud crawlers | High | Plan-dependent | Monthly + page quotas | Managed operations |
| Browser-based crawlers | Moderate | Rendering bottleneck | Usually page credits | JS-heavy sites |
| SEO crawlers | High for HTML | Project/page limits | Subscription | Technical SEO audits |
Benchmarks worth running
Test each platform against the same site and measure:
- Time to crawl 100,000 pages
- Average pages/sec
- Peak pages/sec
- Failed requests
- Retry success rate
- CPU/memory usage (self-hosted)
- Total cost
Example:
| Platform | 100k pages | Avg pages/sec | Cost | Notes |
|---|---|---|---|---|
| Platform A | 45 min | 37 | $20 | No JS |
| Platform B | 2 hr | 14 | $35 | JS enabled |
| Platform C | 28 min | 60 | $15 infra | Self-hosted |
Hidden limits to look for
Many vendors don't advertise these:
- Maximum crawl depth
- Maximum URLs per project
- Maximum concurrent crawls
- API request quotas
- Export size limits
- Maximum sitemap size
- Robots.txt overrides
- Per-domain throttling
- Browser rendering quotas
Cost comparison formula
For recurring crawls, estimate:
Monthly cost =
(number of pages × crawl frequency × cost per page)
+ platform subscription
+ rendering charges
+ API overage fees
For example, crawling a 5 million page site weekly:
- 20M pages/month
- If priced at $0.40 per 1,000 pages → about $8,000/month
- If self-hosted on cloud VMs, infrastructure may be substantially lower but requires engineering and operations.
Questions to ask vendors
- What is the maximum sustained crawl rate?
- Can crawl concurrency be configured?
- Are there hard page-count limits?
- How are JavaScript-rendered pages billed?
- Can multiple crawlers run simultaneously?
- Is distributed crawling supported?
- What happens after reaching the monthly quota?
- Are retries counted toward usage?
- Is there an SLA for large crawls?
- Can the platform complete a 10M-page crawl within a specified time window?
If you're evaluating specific platforms (for example, Firecrawl, Crawl4AI, Apify, Bright Data, Diffbot, Screaming Frog, Sitebulb, or custom Scrapy clusters), I can build a side-by-side comparison of their crawl speed, page limits, concurrency controls, and pricing, along with recommendations based on your site's size and JavaScript requirements.
Behind this answer: no web searches, 2 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Answer 19 of 22, 7th Aug 2026, gpt-5-5, mentions 1 tracked brand
When evaluating crawl platforms for a large website (100k–10M+ URLs), raw "pages per second" is only one metric. In practice, crawl throughput depends on rate limiting, JavaScript rendering, concurrency, retries, robots.txt behavior, and the target site's infrastructure.
A practical comparison framework looks like this:
| Factor | Why it matters | How to measure |
|---|---|---|
| Maximum crawl rate | Peak pages/sec | Run the same URL set on each platform |
| Sustained crawl rate | Real-world performance | Measure over 30–60 minutes |
| Concurrent requests | Determines throughput | Check configurable workers/connections |
| JS rendering | Can reduce speed by 10–100× | Compare HTML-only vs rendered crawls |
| Rate limiting | Prevents overwhelming origin | Per-domain requests/sec controls |
| Retry behavior | Affects completion time | Compare retry policies |
| Crawl depth | Deep sites expose bottlenecks | Test identical sitemap or seed list |
| Distributed crawling | Multiple regions/workers | Measure horizontal scaling |
| Cost | Dollars per million pages | Include compute and bandwidth |
Cost comparison
A useful normalized metric is:
Cost per 1 million pages crawled
rather than monthly subscription.
Example template:
| Platform | Monthly cost | Included pages | Extra pages | Cost / 1M pages |
|---|---|---|---|---|
| Platform A | $500 | 5M | $X | $100 |
| Platform B | Usage-based | — | $Y/page | $85 |
| Self-hosted | Infra only | Unlimited | — | Variable |
Performance benchmark
For each platform:
- Crawl the same sitemap
- Disable JS initially
- Same user agent
- Same concurrency (or maximum allowed)
- Same timeout
- Same robots policy
Measure:
- URLs crawled
- URLs/sec
- Average response time
- 95th percentile latency
- Failed URLs
- Retry count
- Total wall-clock time
- CPU/RAM (if self-hosted)
Typical throughput ranges
These are approximate real-world ranges for HTML crawling:
| Scale | Throughput |
|---|---|
| Single-thread crawler | 2–20 pages/sec |
| Moderate cloud crawler | 50–300 pages/sec |
| Distributed crawler | 500–5,000+ pages/sec |
| Large enterprise infrastructure | 10,000+ pages/sec (multiple workers across domains) |
Actual speed is often limited by:
- origin server capacity
- CDN throttling
- robots crawl-delay
- network latency
- JavaScript execution
Important limits to compare
Beyond speed, compare whether the platform supports:
- Maximum concurrent domains
- Maximum URLs per crawl
- Crawl resumption
- Incremental crawling
- Sitemap ingestion
- Change detection
- API-triggered crawls
- Webhook completion notifications
- Proxy rotation
- Geographic crawl locations
- Custom headers/cookies
- Authentication
- Headless browser support
- Export formats (JSON, CSV, Parquet)
Questions to ask vendors
When evaluating commercial platforms, ask for:
- Maximum sustained pages/sec per domain.
- Maximum pages per crawl.
- Whether concurrency is configurable.
- Whether JavaScript rendering is optional.
- Pricing per million pages.
- Rate limiting and politeness controls.
- SLA for crawl completion.
- API rate limits.
- Retry and deduplication behavior.
- Support for distributed workers.
If you're comparing specific platforms (for example, Firecrawl, Crawl4AI, Apify, Bright Data, Diffbot, Zyte, SiteOne, or Screaming Frog), I can also produce a side-by-side matrix covering crawl speed, limits, pricing, scalability, JavaScript support, APIs, and best fit for large-scale crawling.
Behind this answer: no web searches, 1 tracked brand mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Answer 20 of 22, 7th Aug 2026, gpt-5-5, mentions no tracked brands
When comparing crawl platforms for a large website (100k–10M+ pages), it's helpful to separate raw crawl speed from practical throughput. Many vendors advertise requests/second, but the limiting factors are often concurrency policies, JavaScript rendering, rate limits, and pricing.
Here's a framework you can use.
| Metric | Why it matters | How to measure |
|---|---|---|
| Pages/hour | Overall throughput | Crawl a fixed URL list and measure completion time |
| Requests/second | Network speed | Average sustained RPS, not peak |
| Concurrent requests | Primary driver of speed | Maximum parallel fetches allowed |
| JavaScript rendering | Can reduce throughput 10–100× | Compare HTML-only vs JS mode |
| Robots.txt handling | Affects usable speed | Verify configurable behavior |
| Retry/backoff | Impacts large crawls | Measure failures and recovery |
| Incremental crawling | Saves time later | Compare change detection |
| API rate limits | Hidden bottleneck | Read platform limits carefully |
| Cost per million pages | Better than monthly price | Normalize pricing |
Benchmark methodology
Run every platform against the same crawl:
- identical seed URLs
- same maximum depth
- same robots settings
- same JavaScript settings
- same timeout
- same extraction rules
- same region
Measure:
Total pages crawled
Elapsed time
Pages/hour
Successful fetches
Average latency
95th percentile latency
Error rate
Cost
Then calculate
Effective pages/hour
Effective cost =
monthly cost /
pages crawled
Cost per million pages
Time to crawl 1M pages
Example comparison table
| Platform | Crawl speed | JS support | Typical limits | Pricing model |
|---|---|---|---|---|
| Firecrawl | Medium–High | Excellent | Credits/concurrency | Credits |
| Crawl4AI | Depends on infrastructure | Excellent | Your hardware | Free OSS |
| Apify Crawlee | Very high | Excellent | Compute units | Compute usage |
| Bright Data Web Scraper | High | Excellent | Enterprise | Usage |
| Diffbot | Medium | Excellent extraction | API limits | API usage |
| Common Crawl | N/A (dataset) | N/A | Historical only | Free |
What actually determines crawl speed
Usually:
crawl speed =
min(
platform concurrency,
website rate limits,
network latency,
HTML size,
JS rendering time,
storage throughput
)
For example
HTML pages
- 50–300 pages/sec is achievable
JavaScript pages
- 2–20 pages/sec is common
Heavy SPAs
- <1 page/sec can happen
Questions to ask vendors
- Maximum concurrent requests?
- Can concurrency be increased?
- Separate limits for JS rendering?
- Burst vs sustained throughput?
- Per-domain throttling?
- Dedicated workers available?
- Distributed crawling?
- Incremental crawl support?
- Export streaming or only after completion?
- SLA for million-page crawls?
Normalize cost
Rather than comparing monthly plans:
Cost per million pages
Monthly cost /
Monthly page allowance
×
1,000,000
Example
| Platform | Monthly | Included pages | Cost / 1M pages |
|---|---|---|---|
| A | $99 | 500k | $198 |
| B | $300 | 5M | $60 |
| C | Usage | Variable | Calculate from actual usage |
This often reveals that a seemingly more expensive plan has a much lower unit cost at scale.
Recommendations by use case
- Fastest self-hosted: Crawl4AI or Crawlee on your own infrastructure, where you control concurrency and hardware.
- Fastest managed: Apify, Bright Data, or similar platforms with distributed workers and configurable concurrency.
- Best for AI/RAG pipelines: Firecrawl, if its extraction features outweigh raw throughput needs.
- Lowest cost at very large scale: Self-hosted Crawlee or Crawl4AI on cloud instances, especially when crawling millions of pages regularly.
If you're evaluating specific platforms (for example, Firecrawl vs. Apify vs. Crawl4AI vs. Browserbase), I can also build a side-by-side comparison with estimated crawl times and monthly costs for a site of your size (e.g., 500k, 2M, or 10M pages).
Behind this answer: no web searches
ChatGPT answered from its own knowledge, without searching the web.
Answer 21 of 22, 7th Aug 2026, gpt-5-5, mentions 2 tracked brands
When comparing website crawling platforms for a large site (100k–100M+ pages), raw "pages per second" is only one part of the story. The fastest crawler is often the one that can sustain high throughput without overwhelming your origin servers, while still extracting the data you need.
A practical evaluation framework looks like this:
| Metric | Why it matters | How to measure |
|---|---|---|
| Crawl throughput | Pages crawled per second/minute | Crawl a representative section of your site (e.g. 100k URLs) |
| Concurrent requests | Determines maximum speed | Increase concurrency until your server or crawler becomes the bottleneck |
| Crawl limits | Maximum URLs per crawl/project | Check hard platform limits |
| JavaScript rendering | Often reduces speed by 5–20× | Compare HTML-only vs JS rendering |
| Distributed crawling | Uses multiple workers | Check whether scaling is automatic or manual |
| API rate limits | Can bottleneck automation | Measure requests/minute and export limits |
| Cost | Cost per million pages or per crawl hour | Normalize pricing to pages crawled |
Metrics to collect
Run the same crawl configuration on each platform:
- Same seed URLs
- Same crawl depth
- Same robots.txt behavior
- Same JavaScript settings
- Same extraction rules
Then record:
- Total pages discovered
- Pages successfully crawled
- Crawl duration
- Average pages/sec
- Peak pages/sec
- Errors (429, 500, timeouts)
- Memory usage
- Export time
Example comparison table
| Platform | Large crawl support | Typical speed | JS support | Pricing model | Good for |
|---|---|---|---|---|---|
| Screaming Frog | Hundreds of thousands to millions (desktop resources permitting) | High on local hardware | Yes | Per license | Technical SEO |
| Sitebulb | Large enterprise sites | Moderate–high | Yes | License/subscription | SEO audits |
| Apify | Very large, cloud-distributed | Very high | Excellent | Compute usage | Custom crawling |
| Crawl4AI | Depends on deployment | High when self-hosted | Excellent | Open source | AI pipelines |
| Firecrawl | Cloud API | Moderate–high | Excellent | Credits/pages | LLM-ready extraction |
| Bright Data | Enterprise scale | Very high | Excellent | Usage-based | Large-scale web data |
| Scrapy | Virtually unlimited | Very high | Plugins | Open source | Engineering teams |
Normalize cost
Rather than comparing subscription prices, calculate:
- Cost per 100,000 pages
- Cost per 1 million pages
- Cost per GB exported
- Cost per successful page
For example:
| Platform | Monthly price | Pages/month | Effective cost / 1M pages |
|---|---|---|---|
| Platform A | $500 | 5M | $100 |
| Platform B | $900 | 20M | $45 |
| Platform C | Pay-as-you-go | Variable | Depends on workload |
This often reveals that a seemingly more expensive platform is actually cheaper at scale.
Watch for hidden bottlenecks
Many platforms advertise "unlimited crawling," but actual throughput may be constrained by:
- Maximum concurrent requests
- Per-domain rate limiting
- JavaScript rendering queues
- API quotas
- Export size limits
- Maximum crawl duration
- Worker allocation
These limits can dominate performance on very large sites.
Recommended benchmark
If you're evaluating vendors, a reproducible benchmark is:
- 100,000 URLs
- HTML-only crawl
- JavaScript-enabled crawl
- 32 concurrent requests
- Same extraction schema
- Measure:
- total crawl time
- pages/sec
- CPU utilization
- failed pages
- retries
- total cost
This gives a realistic picture of throughput, reliability, and economics.
If your goal is millions of pages, the comparison should focus on cloud-native distributed crawlers (such as Apify, Bright Data, Firecrawl, Crawl4AI, or a custom Scrapy deployment). If you're primarily auditing your own site for SEO, desktop tools like Screaming Frog or Sitebulb may be more appropriate despite lower horizontal scalability.
Behind this answer: no web searches, 2 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Answer 22 of 22, 7th Aug 2026, gpt-5-5, mentions no tracked brands
If you're evaluating platforms for large-scale website crawling, it's useful to compare them across a common set of measurable criteria rather than marketing claims. The key metrics are crawl throughput, scalability, operational limits, and total cost.
Here's a comparison framework you can use.
| Category | What to measure | Why it matters |
|---|---|---|
| Crawl speed | Pages/minute or URLs/second | Primary performance metric |
| Parallelism | Maximum concurrent requests | Determines peak throughput |
| Rate limiting | Requests per host/domain | Prevents overwhelming target sites |
| Maximum crawl size | Pages per crawl and total storage | Large sites can exceed platform limits |
| JavaScript rendering | On/off, concurrent browser limits | JS rendering is often the biggest slowdown |
| Geographic distribution | Number of crawl regions | Useful for latency and regional content |
| Incremental crawling | Changed pages only | Saves time and cost on recurring crawls |
| Scheduling | Continuous vs batch | Important for keeping data fresh |
| API limits | Requests/hour, jobs/day | Can become the bottleneck |
| Export speed | Streaming vs bulk download | Large datasets can take hours to export |
| Failure recovery | Resume interrupted crawls | Critical for multi-day crawls |
| Cost model | Per page, per credit, compute-hour, or subscription | Makes platforms difficult to compare directly |
Normalize everything to cost per million pages
Instead of comparing subscription prices, calculate:
- Cost per 100,000 pages
- Cost per million pages
- Average crawl duration
- Effective pages/hour
For example:
| Platform | Cost | Pages | Effective cost / 1M pages |
|---|---|---|---|
| Platform A | $200 | 500k | $400 |
| Platform B | $500 | 3M | $167 |
| Platform C | $0.002/page | varies | $2,000 |
This often reveals that a seemingly expensive platform is actually cheaper at scale.
Benchmark using the same site
To compare speed fairly:
- Use the identical seed URLs.
- Disable JavaScript unless it's required.
- Use the same robots.txt behavior.
- Use the same crawl depth.
- Exclude media downloads if possible.
- Run each crawl at approximately the same time of day.
Measure:
- Time to first 100k pages
- Time to 500k pages
- Time to completion
- Pages successfully crawled
- HTTP error rate
- Duplicate URL rate
Understand what limits speed
For very large sites, crawl speed is usually constrained by one of these factors:
- Target website rate limits (the biggest limiter)
- Browser rendering (JavaScript)
- DNS and network latency
- API concurrency limits
- Platform-imposed page caps
- Storage and export bottlenecks
Questions to ask each vendor
Ask vendors for concrete numbers rather than qualitative answers:
- What is the maximum crawl throughput (pages/minute)?
- What is the maximum concurrent request count?
- What is the largest production crawl you've supported?
- Is there a hard page limit per crawl?
- How many simultaneous crawl jobs are allowed?
- How are retries handled?
- Is JavaScript rendering billed separately?
- Are there API rate limits?
- Can crawls resume after interruption?
- How are credits calculated?
- Are bandwidth or storage charged separately?
Typical platform trade-offs
While specifics vary, many platforms fall into broad categories:
| Type | Strength | Weakness |
|---|---|---|
| Managed cloud crawler | Easy to use, minimal maintenance | May have page or concurrency limits |
| Serverless crawler | Scales well automatically | Costs can rise quickly with browser rendering |
| Self-hosted crawler | Highest potential throughput and control | Requires infrastructure and operational expertise |
| Enterprise crawling platform | Strong scheduling, compliance, and support | Higher cost and longer procurement process |
Recommended evaluation process
Run a proof of concept using a representative portion of your largest site (for example, 500,000 to 1,000,000 URLs). Record:
- Total crawl duration
- Average pages/minute
- Peak pages/minute
- Success rate
- CPU/browser utilization (if self-hosted)
- Total cost
- Export time
These measurements will give you a more reliable comparison than advertised performance figures.
If you're comparing specific platforms (for example, Firecrawl vs. Crawl4AI vs. Apify vs. Bright Data vs. Scrapy Cloud), I can also build a side-by-side comparison of their documented crawl speed, concurrency limits, pricing models, JavaScript support, and practical throughput for multi-million-page websites.
Behind this answer: no web searches
ChatGPT answered from its own knowledge, without searching the web.
See where your brand stands in AI
Track how AI talks about your brand, free. 40 prompts, no credit card.
Start tracking free
