ChatGPT answers, unedited
"I inherited a messy website, what combo of seo audit tools helps prioritize fixes for performance and crawlability?", asked on ChatGPT about once a day from the United States. The last 22 answers, newest first, exactly as given: the raw data behind the SEO Tools live demo.
Week of 24th to 30th Aug 2026 (in progress)
For an inherited “messy” site, I’d avoid relying on a single SEO score. The best setup is a crawler + Google data + performance profiler + prioritization layer. A good combo:
1. Crawlability & technical structure (find the mess)
Primary crawler:
- Screaming Frog — best for deep technical crawling: broken links, redirect chains, canonicals, indexability, duplicate pages, JavaScript rendering issues, XML sitemap problems, and custom extractions. HighLevel Playbooks Teksyte
- Alternative: Sitebulb — easier-to-read reports and stronger issue prioritization/visualizations if you need to explain fixes to stakeholders. Sitebulb
Run first and export:
- 4xx/5xx URLs
- redirect chains/loops
- orphan pages
- pages blocked by robots/noindex
- duplicate titles/meta/canonicals
- crawl depth
- internal link gaps
- thin or near-duplicate templates
2. Google’s view of the site (ground truth)
Use:
Focus on:
- Pages excluded from indexing
- Crawl stats
- Sitemaps
- Core Web Vitals
- Pages losing impressions/clicks
A crawler tells you what could be wrong; Search Console tells you what Google is actually seeing. Semrush
3. Performance diagnosis
Use:
Prioritize:
- LCP problems (slow main content)
- INP problems (interaction delays)
- CLS issues (layout shifts)
- oversized images
- render-blocking CSS/JS
- excessive third-party scripts
4. Prioritization / “what should we fix first?”
Add one SEO suite:
- Semrush — useful for ongoing site health, rankings, competitors, and automated technical issue tracking. Semrush
- Ahrefs — stronger if the inherited site has backlink/history issues or you need to understand lost authority. HighLevel Playbooks
My recommended stack by situation
Small/medium site (<50k URLs)
- Google Search Console
- Screaming Frog
- PageSpeed Insights
- Semrush or Ahrefs
Large site / ecommerce / many templates
- Screaming Frog (deep crawl)
- Sitebulb (prioritized reporting)
- Google Search Console
- PageSpeed Insights
- Semrush/Ahrefs
- Server log analysis if crawl waste is suspected
How I’d triage the first 2 weeks
- Day 1–2: Crawl the site and create buckets:
- indexing blockers
- performance blockers
- architecture issues
- content/template issues
- Day 3–5: Cross-reference with Search Console:
- Which broken areas actually affect impressions/clicks?
- Week 2: Fix highest-impact items:
- accidental noindex
- bad canonicals
- redirect disasters
- sitemap issues
- slow templates
- internal linking problems
A useful rule: don’t chase 500 warnings. Look for issues affecting many URLs × valuable URLs × Google visibility. That combination usually reveals the real backlog.
Behind this answer: no web searches, 5 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Week of 17th to 23rd Aug 2026
Answer 2 of 22, 23rd Aug 2026, unknown, mentions 3 tracked brands
If you inherited a messy site, I’d use a 3-tool stack, with each tool answering a different question:
- search.google.com — “What is Google actually seeing?”
Start here. Look at Page Indexing, Core Web Vitals, sitemaps, crawl/indexing errors, and URL Inspection. Search Console gives you Google's actual index/crawl perspective rather than just a simulated audit. Google Google Support
- screamingfrog.co.uk — “What is broken across the site?”
This is my first choice for the technical crawl. It finds 4xx/5xxs, redirect chains, orphan/non-indexable URLs, robots directives, canonicals, crawl depth, internal-link problems, duplicate pages, and—if needed—JavaScript-rendered problems. It can also pull Search Console and PageSpeed data into the crawl. Screaming Frog Screaming Frog
- pagespeed.web.dev — “Which performance problems actually matter?”
Use it on representative templates and your important landing pages, rather than obsessing over every URL's score. PSI combines real-user CrUX data with Lighthouse lab diagnostics. PageSpeed Insights Google for Developers
How I'd prioritize the mess
Don't simply fix whatever tool reports the most errors. Build a priority queue roughly like this:
| Priority | Look for | Typical action |
|---|---|---|
| 🔴 P0 | Important pages not indexable | Fix noindex, robots, canonical, status codes |
| 🔴 P0 | 5xxs / major broken URLs | Fix server/app issues |
| 🔴 P0 | Bad redirects / migration leftovers | Collapse chains, restore or redirect URLs |
| 🟠 P1 | Important pages with poor CWV | Fix LCP/INP/CLS bottlenecks |
| 🟠 P1 | Huge crawl waste | Remove parameter traps, duplicate URLs, unnecessary faceting |
| 🟠 P1 | Orphaned valuable pages | Add internal links / sitemap coverage |
| 🟡 P2 | Slow images, JS, CSS | Optimize assets and loading |
| 🟡 P2 | Duplicate titles/meta/content | Consolidate or improve |
| 🟢 P3 | Minor metadata/content warnings | Clean up after structural problems |
For performance specifically, Google currently recommends aiming for LCP ≤2.5s, INP <200ms, and CLS <0.1 for good Core Web Vitals. Google for Developers
Best workflow: Search Console → Screaming Frog → PageSpeed → prioritize by business value × SEO impact × number of affected URLs.
If the site is large (10k+ URLs), I'd add Ahrefs or Semrush for broader backlink/keyword/competitor context—but I wouldn't use either as the primary crawlability diagnostic. The crawler + Search Console combination gives you a much better foundation.
Behind this answer: no web searches, 3 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Answer 3 of 22, 22nd Aug 2026, gpt-5-6, mentions 4 tracked brands
For an inherited “messy” site, I’d avoid relying on a single SEO audit score. The best setup is a stack where each tool answers a different question:
| Goal | Tool | What it uncovers | Why it matters |
|---|---|---|---|
| Find crawl/indexation problems | Screaming Frog SEO Spider | Broken links, redirect chains, orphan pages, bad canonicals, duplicate titles, crawl traps, JS rendering issues | Gives you the raw technical map of the site |
| Prioritize issues by impact | Ahrefs Site Audit or Semrush Site Audit | Severity buckets, trends, recurring issues, technical health tracking | Turns thousands of errors into a fix queue ahrefs.com |
| Validate Google’s view | Google Search Console | Index coverage, crawl stats, manual actions, Core Web Vitals, search queries | Shows what Google actually sees |
| Diagnose speed bottlenecks | Google PageSpeed Insights + Lighthouse | LCP, INP, CLS, render blocking, unused JS/CSS, image problems | Separates “SEO issues” from real performance issues |
| Visualize architecture | Sitebulb | Crawl maps, internal linking problems, priority hints | Helpful when the site structure is chaotic ighenatt.es |
My preferred audit combo
Small/medium site (under ~100k URLs):
- Google Search Console
- Screaming Frog
- PageSpeed Insights
- Ahrefs Site Audit (or Semrush)
Large site / enterprise:
- Screaming Frog or Sitebulb for deep crawling
- Ahrefs/Semrush for ongoing monitoring
- Log file analysis (if you have server access)
- Search Console + analytics data
How I’d prioritize fixes after the first crawl
Don’t start with the longest issue list. Sort by:
Tier 1 — Fix first (can block rankings)
- Pages accidentally
noindexed - Robots.txt blocking important sections
- Bad canonicals
- Massive redirect chains
- 404s on valuable URLs
- Duplicate URL versions being indexed
- Sitemap problems
Tier 2 — Performance wins
- Slow templates affecting thousands of pages
- Huge JavaScript bundles
- Unoptimized images
- Render-blocking CSS/JS
- Poor Core Web Vitals
Tier 3 — Cleanup
- Missing meta descriptions
- Minor heading issues
- Low-impact warnings
- Cosmetic HTML problems
A practical workflow
- Run Screaming Frog crawl
- Export:
- Response codes
- Indexability
- Canonicals
- Page titles/H1s
- Internal links
- Crawl depth
- Export:
- Connect GSC
- Find pages Google cares about that your crawler doesn’t.
- Run PageSpeed on representative templates
- Don’t test 500 URLs; test:
- homepage
- category/listing page
- product/service page
- blog/article page
- Don’t test 500 URLs; test:
- Create a fix backlog
- Columns:
- Issue
- URLs affected
- Traffic impact
- Dev effort
- Priority
- Columns:
A good inherited-site audit usually ends up with 10–20 high-leverage fixes, not a 300-item SEO to-do list. The combination above is designed to find those needles instead of just producing noise. www.teksyte.com
Behind this answer: 1 search, 9 results retrieved, 4 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://netpartners.marketing/seo-audit-tools-2026-ahrefs-semrush-screaming-frog-sitebulb-comparison/?utm_source=chatgpt.com | not cited |
| 2 | https://ahrefs.com/site-audit?r=imc-sat&utm_source=chatgpt.com | not cited |
| 3 | https://designcopy.net/en/best-ai-site-audit-tools/?utm_source=chatgpt.com | not cited |
| 4 | https://lazymetrics.com/blog/best-technical-seo-audit-tools?utm_source=chatgpt.com | not cited |
| 5 | https://worldmetrics.org/best/website-audit-software/?utm_source=chatgpt.com | not cited |
| 6 | https://www.alphonsolabs.com/best-technical-seo-tools-site-audits-2026/?utm_source=chatgpt.com | not cited |
| 7 | https://rankture.com/blog/seo-audit-tools-comparison?utm_source=chatgpt.com | not cited |
| 8 | https://pikaseo.com/articles/best-ai-seo-audit-tools?utm_source=chatgpt.com | not cited |
| and 1 more results | ||
Tracked brands mentioned
Answer 4 of 22, 21st Aug 2026, unknown, mentions 4 tracked brands
For an inherited messy site, I’d avoid relying on one “SEO score” tool. The best setup is a stack where each tool answers a different question: Can Google crawl it? What is broken? What is slow? What matters most? No single platform covers all of those equally well. Ighenatt
A practical combo:
1. Google Search Console — “What is Google actually seeing?”
Use this as your source of truth.
Find:
- Indexed vs excluded pages
- Crawl errors
- Sitemap problems
- Manual actions/security issues
- Core Web Vitals field data
- Pages losing impressions/clicks
Best first step: export:
- Pages indexed
- Not indexed reasons
- Performance queries/pages
- Core Web Vitals issues
Semrush## 2. Screaming Frog SEO Spider — “What is structurally broken?”
This is usually the fastest way to understand a messy inherited site.
Run a full crawl and prioritize:
- 4xx/5xx URLs
- Redirect chains and loops
- Broken internal links
- Duplicate titles/meta descriptions
- Missing canonicals
- Noindex mistakes
- Orphaned pages (when combined with sitemap/GSC data)
- Thin or duplicate content patterns
It’s especially useful because you can inspect URL-level problems instead of just seeing a health score. LazyMetrics
3. PageSpeed Insights + Lighthouse — “What performance fixes are worth doing?”
Use these for:
- Core Web Vitals
- Largest Contentful Paint (LCP)
- Interaction to Next Paint (INP)
- Cumulative Layout Shift (CLS)
- JavaScript/CSS/image bottlenecks
Don’t chase every Lighthouse warning. Prioritize:
- Slow templates affecting many URLs
- Large images
- Heavy JS bundles
- Third-party scripts
- Server response time
Internetzone I## 4. Ahrefs or Semrush — “Which fixes have SEO impact?”
Pick one, not necessarily both.
Use it for:
- Organic traffic trends
- Lost rankings
- Backlink quality
- Broken backlinks
- Competitor comparisons
- Keyword/page prioritization
Their site audit features are useful for ongoing monitoring after the cleanup. HighLevel Playbooks Semrush
5. Optional: log file analysis — “How is Googlebot really crawling?”
If you have server logs, add:
- Screaming Frog Log File Analyzer
- Lumar (formerly DeepCrawl) for enterprise-scale sites
This is valuable for large sites where crawl budget or huge URL inventories are problems. Internetzone I
My recommended audit sequence
- Google Search Console - Identify indexing/crawl problems
- Screaming Frog - Map technical debt
- PageSpeed Insights - Find template-level performance problems
- Ahrefs/Semrush - Tie fixes to traffic and revenue impact
- Create a prioritized backlog
A good prioritization model:
| Priority | Examples |
|---|---|
| P0 — Fix immediately | robots.txt blocking pages, accidental noindex, widespread 5xx errors, broken canonicals |
| P1 — High impact | slow templates, bad internal linking, redirect chains, duplicate indexable pages |
| P2 — Optimization | metadata cleanup, schema improvements, image optimization |
| P3 — Nice-to-have | minor warnings, cosmetic Lighthouse issues |
For most small-to-medium inherited sites, Google Search Console + Screaming Frog + PageSpeed Insights + either Ahrefs or Semrush is the strongest cost/coverage combination. Media Search Group
If you tell me the site size (rough URL count), CMS (WordPress, Shopify, custom, etc.), and whether traffic has dropped, I can suggest a tighter audit stack and order of operations.
Behind this answer: no web searches, 4 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Answer 5 of 22, 20th Aug 2026, gpt-5-6, mentions 3 tracked brands
If you inherited a messy site, I’d use a small stack rather than one “SEO audit” tool. The key is separating discovery (what’s broken) from real-world Google data (what’s actually indexed/slow).
My preferred combo
| Tool | Best for | Use it to prioritize |
|---|---|---|
| Google Search Console | Google’s actual view of the site | Indexing, crawl anomalies, Core Web Vitals, sitemap coverage |
| Screaming Frog SEO Spider | Deep technical crawl | Status codes, redirects, canonicals, robots directives, orphan/duplicate patterns, internal links |
| Ahrefs Site Audit | Ongoing prioritized technical monitoring | Large-scale issue discovery, crawlability, indexability, JS/CSS, Core Web Vitals |
| PageSpeed Insights / Lighthouse | Performance diagnosis | LCP, INP, CLS and the specific resources/code causing slowness |
| Semrush Site Audit (alternative to Ahrefs) | Convenient prioritized backlog | Crawlability + performance + internal linking in a single dashboard |
Ahrefs currently checks 170+ technical/on-page issues, including slow pages, Core Web Vitals, indexability, redirects, JS/CSS, robots and sitemaps. ahrefs.com Semrush similarly groups crawlability, robots.txt, performance, internal linking and other technical findings into prioritized issues. www.semrush.com
How I'd actually use them
1. Start with Search Console.
Don't blindly fix every warning from an SEO crawler. First establish whether Google is actually having trouble:
- Pages excluded from indexing
- Unexpected
noindex - Server/5xx errors
- Sitemap problems
- Crawl/indexing anomalies
- Core Web Vitals failures
2. Run Screaming Frog against the entire site.
This is your forensic crawl. Export the URLs and look particularly for:
- 4xx/5xx pages
- Redirect chains/loops
- Canonicals pointing somewhere unexpected
- Indexable pages with no internal links
- Huge numbers of near-duplicates
- Pagination/faceted-navigation explosions
- Robots/meta-robots conflicts
- Excessive crawl depth
- Bloated URL parameters
3. Run Ahrefs or Semrush—not necessarily both.
I'd choose Ahrefs if the site has substantial SEO/backlink history and you want ongoing monitoring. Its current Always-on Audit can continuously crawl and alert on newly discovered problems. help.ahrefs.com
Choose Semrush if you want a more prescriptive “here's the prioritized backlog” workflow. Its Site Audit explicitly ranks issues by priority and affected-page volume. www.semrush.com
4. Use PageSpeed Insights for the performance rabbit hole.
Your crawler might tell you “these 800 pages are slow.” PSI/Lighthouse helps determine why: oversized images, render-blocking resources, JavaScript, caching, fonts, third-party scripts, etc.
The prioritization rule I'd use
Don't prioritize by the tool's Health Score. Build a backlog roughly like this:
P0 — Can prevent crawling/indexing
- Accidental
noindex - Important URLs blocked by robots.txt
- 5xx/server failures
- Broken canonicals
- Bad redirects
- Sitemap containing non-indexable URLs
P1 — Affects lots of valuable pages
- Slow templates
- Poor Core Web Vitals
- Crawl traps/URL proliferation
- Massive duplicate-page sets
- Broken internal-link architecture
P2 — Important but less existential
- Orphan pages
- Excessive crawl depth
- Missing metadata
- Image optimization
- Minor redirect cleanup
P3 — SEO housekeeping
- Cosmetic warnings
- Low-value metadata tweaks
- Tiny HTML optimizations with negligible user impact
The big trick with an inherited site is to prioritize by affected URLs × business importance × severity, not simply “fix all 1,000 errors.”
If I had to pick the leanest stack: Search Console + Screaming Frog + PageSpeed Insights + Ahrefs. That gives you Google's perspective, a raw technical crawl, performance diagnostics, and ongoing monitoring without drowning in overlapping reports.
Behind this answer: 4 searches, 30 results retrieved, 3 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://help.ahrefs.com/en/articles/1431155-how-do-i-finish-crawling-my-website-faster-in-site-audit?utm_source=chatgpt.com | not cited |
| 2 | https://help.ahrefs.com/en/articles/1402668-troubleshooting-common-issues-with-site-audit-access?utm_source=chatgpt.com | not cited |
| 3 | https://ahrefs.com/academy/how-to-use-ahrefs/site-audit/introduction?utm_source=chatgpt.com | not cited |
| 4 | https://help.ahrefs.com/en/articles/9082329-how-should-i-configure-my-site-audit-settings?utm_source=chatgpt.com | not cited |
| 5 | https://ahrefs.com/seo-audit-tool?utm_source=chatgpt.com | not cited |
| 6 | https://ahrefs.com/website-checker?utm_source=chatgpt.com | not cited |
| 7 | https://help.ahrefs.com/en/collections/87925-faqs-and-troubleshooting?utm_source=chatgpt.com | not cited |
| 8 | https://help.ahrefs.com/en/collections/87920-site-audit?utm_source=chatgpt.com | not cited |
| and 22 more results | ||
Tracked brands mentioned
Answer 6 of 22, 19th Aug 2026, gpt-5-6, mentions 4 tracked brands
For a messy inherited site, I’d avoid relying on a single “SEO score” tool. The best setup is a crawler + Google data + performance tooling + competitive context stack. Most technical SEOs combine tools because crawlability, indexing, and performance are different diagnoses. ighenatt.es
A practical combo:
| Goal | Tool | What it helps prioritize |
|---|---|---|
| Find crawl/indexing problems | Screaming Frog SEO Spider | Broken links, redirect chains, orphan pages, canonical issues, duplicate URLs, thin pages, bad internal linking, JS rendering problems |
| Make sense of technical issues | Sitebulb | Turns crawl data into prioritized “fix this first” recommendations and architecture visuals |
| Google’s actual view of the site | Google Search Console | Index coverage, crawl stats, Core Web Vitals, manual actions, search queries |
| Performance diagnosis | Google Lighthouse + PageSpeed Insights | LCP, CLS, INP, render blocking, JS/CSS issues, image optimization |
| Backlinks + ranking impact | Ahrefs or Semrush | Prioritize fixes based on pages that already have links, rankings, or traffic potential |
My preferred workflow for an inherited mess:
- Start with Google Search Console
- Export:
- Pages not indexed
- Crawl stats
- Core Web Vitals
- Top declining pages
- This tells you what Google is actually struggling with.
- Export:
- Run a full crawl in Screaming Frog
Look first at:
- 5xx errors
- 4xx pages with backlinks
- redirect chains
- canonical mismatches
- duplicate titles/H1s
- pages more than ~3 clicks deep
- huge numbers of parameter URLs
- blocked resources
- Overlay business value
Use Ahrefs/Semrush to ask:
- “Which broken pages have backlinks?”
- “Which slow pages rank?”
- “Which templates affect the most traffic?”
- Run performance checks on templates, not every URL
Usually you only need:
- homepage
- category/listing page
- product/service page
- blog/article page
- checkout/conversion pages
A good prioritization model:
P0 — Fix immediately
- Pages blocked from indexing accidentally
- Sitewide canonical mistakes
- robots.txt/sitemap problems
- server errors
- broken migrations
P1 — High impact
- Slow templates affecting thousands of URLs
- Poor internal linking
- Duplicate content at scale
- Redirect bloat
P2 — Cleanup
- Missing metadata
- Minor schema issues
- Image alt text
- Low-value warnings
If budget is tight, I’d run:
Free/cheap stack
- Google Search Console
- Screaming Frog free version (small sites)
- Lighthouse/PageSpeed Insights
- Ahrefs Webmaster Tools
Professional audit stack
- Screaming Frog + Sitebulb + Ahrefs (or Semrush) + Search Console
That combination gives you enough signal to avoid spending weeks fixing low-impact SEO warnings while missing the few things actually hurting crawlability and performance. ighenatt.esnetpartners.marketing
If you tell me the site type (WordPress, ecommerce, SaaS, local business, etc.) and approximate URL count, I can suggest the crawl settings and audit order.
Behind this answer: 1 search, 9 results retrieved, 4 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://canadacreate.com/technical-seo-audit-stack-b2b-canada-2026/?utm_source=chatgpt.com | not cited |
| 2 | https://ahrefs.com/site-audit?r=imc-sat&utm_source=chatgpt.com | not cited |
| 3 | https://auditweb.site/website-audit-tools/comparison/?utm_source=chatgpt.com | not cited |
| 4 | https://thestacc.com/blog/website-crawler-tools/?utm_source=chatgpt.com | not cited |
| 5 | https://worldmetrics.org/best/site-audit-software/?utm_source=chatgpt.com | not cited |
| 6 | https://worldmetrics.org/best/website-audit-software/?utm_source=chatgpt.com | not cited |
| 7 | https://www.aiocopilot.com/blog/best-seo-audit-tools-comparison-2026?utm_source=chatgpt.com | not cited |
| 8 | https://ranklytics.ai/seo-audit-tool-comparison/?utm_source=chatgpt.com | not cited |
| and 1 more results | ||
Tracked brands mentioned
Answer 7 of 22, 18th Aug 2026, gpt-5-6, mentions 2 tracked brands
For an inherited, messy site, I’d use a stack rather than one “SEO audit” tool. The sweet spot is:
| Tool | Best at | Use it to prioritize |
|---|---|---|
| Google Search Console | Google’s view of crawling/indexing + real search performance | Pages actually failing indexing, crawl anomalies, Core Web Vitals |
| Screaming Frog SEO Spider | Deep technical crawl | 4xx/5xxs, redirect chains, canonicals, orphan/weakly linked pages, robots directives, huge/duplicate pages |
| Ahrefs Site Audit | Prioritized technical audit + monitoring | Turning hundreds of technical findings into errors/warnings and tracking fixes |
| Google PageSpeed Insights | Real-user + lab performance diagnostics | LCP, INP, CLS and the specific resources/code making pages slow |
| Google Lighthouse | Page-level performance/debugging | Developer-level fixes for JS, CSS, images, accessibility and best practices |
The combo I'd actually use
1. GSC = establish what Google is actually experiencing.
Don't start by fixing every warning in a crawler. Look first at indexing, crawl anomalies, sitemap coverage and Core Web Vitals. This prevents you from spending a week fixing technically “bad” URLs that don't matter.
2. Screaming Frog = map the mess.
Run a full crawl and export the important URL-level fields. I'd particularly inspect:
- Status code / redirect chains
- Indexability + indexability reason
- Canonical URL
- Robots directives
- Internal links/inlinks
- Crawl depth
- XML-sitemap inclusion
- Page size
- Response time
- Duplicate titles/H1s
- JavaScript-rendered vs raw HTML differences
This becomes your technical inventory.
3. Ahrefs Site Audit = make the inventory manageable.
Its current Site Audit checks 170+ technical/on-page issues and categorizes them into errors, warnings and notices. It also covers things like slow pages, Core Web Vitals, indexability, redirects, robots.txt, sitemaps, JS/CSS and structured data. ahrefs.com
That's particularly useful on an inherited site because you can establish a baseline and compare subsequent crawls rather than repeatedly starting from scratch.
4. PSI/Lighthouse = diagnose performance, not just score it.
Don't treat a Lighthouse score of 72 vs. 84 as the priority. Look for site-wide patterns: e.g. every template loads 2 MB of JS, product pages have oversized hero images, or a third-party script is hurting INP.
How I'd prioritize the resulting backlog
Use roughly this order:
P0 — Can Google/users not reach important content?
- Important pages returning 4xx/5xx
- Accidental
noindex - robots.txt blocking valuable sections
- Bad canonicals
- Redirect loops/chains
- Important pages absent from internal linking
- Sitemap containing non-indexable URLs
P1 — Important pages are reachable but unnecessarily expensive
- Very slow server response
- Huge HTML/JS/CSS
- Poor LCP/INP/CLS across important templates
- Excessive redirects
- Crawl traps / parameter explosions
- Massive duplicate URL sets
P2 — Quality/architecture cleanup
- Duplicate metadata
- Thin/low-value pages
- Broken internal links
- Orphan pages
- Weak internal linking
- Miscellaneous structured-data and markup issues
The key is to multiply severity × number of affected URLs × business importance. A problem affecting 5,000 revenue-generating URLs beats a prettier but technically worse problem affecting 12 blog posts.
One useful feature of Ahrefs here is that its issue system explicitly distinguishes errors, warnings and notices, while allowing you to customize issue importance for your particular site. ahrefs.com
If you only want three tools: GSC + Screaming Frog + PageSpeed Insights.
If you're going to maintain the site long-term: add Ahrefs Site Audit for recurring monitoring.
If you tell me roughly what kind of site it is (WordPress, Shopify, custom, etc.) and its size (100 / 10k / 1M+ URLs), I can give you a much more specific audit sequence and priority matrix.
Behind this answer: 4 searches, 30 results retrieved, 2 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://ahrefs.com/seo-audit-tool?utm_source=chatgpt.com | not cited |
| 2 | https://ahrefs.com/site-audit?r=imc-sat&utm_source=chatgpt.com | not cited |
| 3 | https://help.ahrefs.com/en/articles/1399529-how-to-use-site-audit-filters-in-page-explorer-and-link-explorer?utm_source=chatgpt.com | not cited |
| 4 | https://help.ahrefs.com/en/articles/9082329-how-should-i-configure-my-site-audit-settings?utm_source=chatgpt.com | not cited |
| 5 | https://ahrefs.com/academy/how-to-use-ahrefs/site-audit/overview?utm_source=chatgpt.com | not cited |
| 6 | https://help.ahrefs.com/en/articles/1424673-what-is-health-score-and-how-is-it-calculated-in-ahrefs-site-audit?utm_source=chatgpt.com | not cited |
| 7 | https://help.ahrefs.com/en/collections/87920-site-audit?utm_source=chatgpt.com | not cited |
| 8 | https://ahrefs.com/webmaster-tools/?utm_source=chatgpt.com | not cited |
| and 22 more results | ||
Tracked brands mentioned
Answer 8 of 22, 17th Aug 2026, gpt-5-6, mentions 5 tracked brands
For an inherited “messy” site, I’d avoid relying on a single SEO score. The best setup is a stack where each tool answers a different question: What is Google seeing? What is technically broken? What is slowing users down? What should we fix first? www.semrush.com
A practical combo:
| Goal | Tool | What it catches | Priority use |
|---|---|---|---|
| Google’s view of the site | Google Search Console | Indexing problems, coverage issues, Core Web Vitals, manual actions, sitemap problems | Start here |
| Deep crawl | Screaming Frog SEO Spider | Broken links, redirect chains, canonicals, duplicate titles, orphan pages, crawl traps | Find technical debt |
| Performance diagnosis | Google PageSpeed Insights + Lighthouse | LCP, INP, CLS, JS/CSS bottlenecks, image issues | Fix speed blockers |
| Broader SEO prioritization | Ahrefs or Semrush | Technical issues + backlinks + rankings + competitors | Decide impact |
| Architecture visualization | Sitebulb (optional) | Crawl paths, internal linking, site structure problems | Useful for big messy sites |
My usual workflow for an inherited site:
- Baseline reality check
- Pull:
- Indexed pages
- Excluded pages
- Crawl errors
- Core Web Vitals
- Top landing pages
- Source: Search Console.
- Pull:
- Run a full crawl
- Crawl with Screaming Frog.
- Export:
- 4xx/5xx URLs
- redirect chains
- duplicate titles/meta
- missing canonicals
- noindex pages
- pages with low internal links
- Match crawl problems against business value
Don’t fix every warning. Prioritize:
- Pages ranking but slow → performance fixes
- Important pages not indexed → crawl/index fixes
- Broken internal links → navigation fixes
- Duplicate/thin pages → consolidation
- Validate performance fixes
Use PageSpeed Insights/Lighthouse on:
- homepage
- top organic landing pages
- templates (product/category/blog pages)
- Monitor after changes
Keep a weekly/monthly crawl and compare:
- number of errors
- indexed pages
- Core Web Vitals
- organic traffic trends
A few stack recommendations by budget:
Free / lean
- Search Console
- Screaming Frog free tier
- PageSpeed Insights
- Lighthouse
Good enough for many small sites. www.mediasearchgroup.com
Professional audit
- Search Console
- Screaming Frog paid
- Ahrefs or Semrush
- PageSpeed Insights
Large site (50k+ URLs)
- Search Console
- Enterprise crawler (e.g., Lumar/OnCrawl)
- Ahrefs/Semrush
- Log file analysis
One thing I’d avoid: starting with an “SEO health score” and chasing 100%. Those tools often surface hundreds of minor warnings. Start with indexability, crawl efficiency, internal linking, and Core Web Vitals—those usually produce the biggest wins on inherited sites. clickraven.comahrefs.com
If you tell me the site size (rough URL count), CMS (WordPress, Shopify, custom, etc.), and whether traffic dropped or it’s just cleanup, I can suggest a tighter audit order.
Behind this answer: 1 search, 8 results retrieved, 5 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://rankture.com/blog/seo-audit-tools-comparison?utm_source=chatgpt.com | not cited |
| 2 | https://rankture.com/blog/free-seo-audit-tools-comparison?utm_source=chatgpt.com | not cited |
| 3 | https://ighenatt.es/en/blog/seo-audit-tools-2026/?utm_source=chatgpt.com | not cited |
| 4 | https://ighenatt.es/en/resources/auditoria-seo/herramientas-auditoria-seo/?utm_source=chatgpt.com | not cited |
| 5 | https://www.cloudways.com/blog/seo-audit-tools-list/?utm_source=chatgpt.com | not cited |
| 6 | https://rankauthority.com/how-to-audit-website-for-seo-a-complete-guide/?utm_source=chatgpt.com | not cited |
| 7 | https://ranklytics.ai/seo-audit-tool-comparison/?utm_source=chatgpt.com | not cited |
| 8 | https://digitalmarketingcdn.com/best-technical-seo-audit-tools/?utm_source=chatgpt.com | not cited |
Tracked brands mentioned
Week of 10th to 16th Aug 2026
Answer 9 of 22, 16th Aug 2026, gpt-5-6, mentions 3 tracked brands
For a messy inherited site, I’d use a small stack where each tool answers a different question, rather than trusting one “SEO health score.”
My preferred combo
| Tool | Use it for | What to prioritize |
|---|---|---|
| Google Search Console | What Google actually sees | Indexing exclusions, crawl stats, Core Web Vitals, sitemap issues |
| Screaming Frog SEO Spider | What exists on the site | 4xx/5xxs, redirect chains, canonicals, noindex, robots.txt, orphan/poorly linked pages, crawl depth |
| PageSpeed Insights / Lighthouse | Why pages are slow | LCP, INP, CLS, TTFB, render-blocking resources, JS/CSS/image problems |
| Semrush Site Audit or Ahrefs Site Audit | Ongoing monitoring + issue aggregation | Recurring technical problems, trends, duplicate/indexability issues |
| Server/log analysis | What search bots actually crawl | Crawl waste, bot traps, excessive parameters, repeated 404s, slow responses |
Screaming Frog is particularly useful for the inherited-site situation because it exposes URL-level technical problems and can integrate Search Console and PageSpeed Insights data into the crawl. It also identifies broken links, redirects, blocked URLs/resources, crawl depth, indexability and internal-linking problems. www.screamingfrog.co.uk
How I'd prioritize the mess
Don't simply fix whatever the audit tool labels "High." Even Screaming Frog explicitly warns that its priority labels are guidance rather than a substitute for site-specific prioritization. www.screamingfrog.co.uk
I'd work in this order:
1. Can Google crawl and index the right pages?
- Accidental
noindex - robots.txt blocking important resources/pages
- 5xx errors
- broken internal links
- bad canonicals
- redirect chains/loops
- XML sitemap containing non-indexable URLs
- huge numbers of low-value parameter URLs
2. Is the site wasting crawl budget?
- Faceted/parameter URLs
- infinite URL combinations
- duplicate pages
- redirect-heavy internal links
- enormous numbers of thin/near-duplicate URLs
- poor internal linking and excessive crawl depth
3. Are important pages actually fast?
Use real-user Core Web Vitals from GSC/CrUX to establish the problem, then PageSpeed/Lighthouse to diagnose it. Don't chase a perfect Lighthouse score if the real-user data doesn't show a meaningful problem.
4. Fix systemic causes before individual URLs.
For example:
18,000 pages have duplicate canonicals → fix the CMS/template generating canonicals.
That's much higher leverage than manually fixing 18,000 URLs.
If you want the leanest setup
I'd start with:
GSC + Screaming Frog + PageSpeed Insights
That gets you surprisingly far without buying a big SEO suite. Current comparisons similarly put those three together as a strong basic technical-audit stack. ighenatt.es Screaming Frog's free version currently crawls up to 500 URLs; larger sites and advanced integrations require a license. www.screamingfrog.co.uk
Then add Semrush or Ahrefs if you need ongoing monitoring, competitive/backlink data, or a more management-friendly dashboard.
One extra tip: if the site is really chaotic, get access to server logs. A crawler tells you what can be crawled; logs tell you what Googlebot is actually spending its time crawling. That's often where the biggest crawlability surprises show up.
Behind this answer: 3 searches, 30 results retrieved, 3 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://www.semrush.com/blog/seo-audit-tools/?msockid=2a02a7961ed961f834cbb16c1fdc604f&utm_source=chatgpt.com | not cited |
| 2 | https://ighenatt.es/en/blog/seo-audit-tools-2026/?utm_source=chatgpt.com | not cited |
| 3 | https://thestacc.com/best/seo-audit-tools/?utm_source=chatgpt.com | not cited |
| 4 | https://www.screamingfrog.co.uk/seo-spider/tutorials/quick-start-guide/?Sid=1&utm_source=chatgpt.com | not cited |
| 5 | https://www.screamingfrog.co.uk/seo-spider/?src=li-leap&utm_source=chatgpt.com | not cited |
| 6 | https://www.screamingfrog.co.uk/seo-spider/?country=207&wmc-currency=USD&utm_source=chatgpt.com | not cited |
| 7 | https://www.screamingfrog.co.uk/seo-spider/?q=templates&utm_source=chatgpt.com | not cited |
| 8 | https://www.screamingfrog.co.uk/seo-spider/user-guide/tabs/?utm_source=chatgpt.com | not cited |
| and 22 more results | ||
Tracked brands mentioned
Answer 10 of 22, 15th Aug 2026, gpt-5-6, mentions 3 tracked brands
For a messy inherited site, I’d use a small stack with clearly separated jobs, rather than trusting one giant “SEO health score.”
My go-to combo
| Tool | Use it for | What to prioritize |
|---|---|---|
| Google Search Console | What Google actually sees | Indexing exclusions, crawl issues, sitemap problems, manual actions, Core Web Vitals |
| Screaming Frog SEO Spider | Deep technical crawl | 4xx/5xxs, redirect chains, canonicals, noindex, orphan-ish pages, duplicate URLs, internal links |
| PageSpeed Insights | Performance | LCP, INP, CLS, render-blocking resources, oversized images, JS/CSS |
| Semrush Site Audit or Ahrefs Site Audit | Ongoing prioritization | Recurring issues, trends, severity, links/traffic context |
Google specifically recommends prioritizing Poor Core Web Vitals issues first, then those affecting the most or most important URLs. Search Console also gives you Google's own crawl/indexing view, so I'd treat it as the source of truth rather than a third-party score. support.google.com
How I'd actually audit the inherited site
1. Start in Search Console
- Look at Pages/Indexing: why aren't URLs indexed?
- Check URL Inspection on important templates.
- Check Core Web Vitals by URL group.
- Look at sitemap status and crawl/indexing anomalies.
2. Crawl the whole site with Screaming Frog
Export the things that can cause real damage:
- 5xx and 4xx URLs
- redirect chains/loops
noindexpages that shouldn't be noindexed- pages blocked by robots.txt
- canonical mismatches
- duplicate URLs
- very deep pages
- pages with few/no internal links
- huge HTML/resource sizes
For a site under ~500 URLs, its free tier can be enough to get started. www.semrush.com
3. Test representative templates in PageSpeed Insights
Don't test only the homepage. Test:
- homepage
- highest-traffic landing page
- category/listing template
- product/article template
- a particularly slow page
Pay attention to field data as well as the lab diagnostics. Google's current Core Web Vitals are LCP, INP and CLS, with recommended “good” thresholds of ≤2.5s, <200ms and <0.1 respectively. developers.google.com
4. Use Semrush/Ahrefs as the prioritization layer
This is where I'd connect technical problems to business impact: a broken page receiving thousands of organic visits deserves attention before a technically ugly page nobody visits.
The prioritization formula I'd use
Don't fix issues in the order the crawler lists them. Rank each finding roughly:
Impact × number/value of affected URLs × likelihood of causing the problem
So I'd generally go:
P0 — Fix immediately
- Important URLs blocked from crawling
noindex/canonical mistakes removing valuable pages from Google- 5xx errors
- broken migration/redirects
- major rendering failures
P1 — High impact
- widespread Core Web Vitals failures
- redirect chains
- enormous JS/image payloads
- bad internal-link architecture
- duplicate URL/indexation explosions
P2 — Cleanup
- missing titles/descriptions
- minor heading issues
- low-value metadata inconsistencies
- small performance opportunities
P3 — Ignore for now
- things that merely make an SEO tool's score prettier without materially affecting crawling, indexing, users, or valuable traffic.
The key is GSC + crawler + performance data, cross-checked rather than averaged. A third-party “78/100 SEO health” score is much less useful than knowing “12,000 URLs are being crawled that should never be indexed, and 40% of organic traffic lands on templates with poor LCP.”
If you tell me roughly how big the site is (e.g. 500, 10k, 1M URLs) and whether it's WordPress/e-commerce/custom, I can suggest the leanest stack and an audit order for that setup.
Behind this answer: 2 searches, 21 results retrieved, 3 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://ighenatt.es/en/blog/seo-audit-tools-2026/?utm_source=chatgpt.com | not cited |
| 2 | https://blog.contentforce.ai/technical-seo-audit/?utm_source=chatgpt.com | not cited |
| 3 | https://thedevbud.com/technical-seo-audit-guide-2026/?utm_source=chatgpt.com | not cited |
| 4 | https://search.google.com/search-console/about?utm_source=chatgpt.com | not cited |
| 5 | https://developers.google.com/search/docs/appearance/page-experience?utm_source=chatgpt.com | not cited |
| 6 | https://developers.google.com/speed/docs/insights/v5/about?authuser=2&hl=en&utm_source=chatgpt.com | not cited |
| 7 | https://support.google.com/webmasters/answer/9205520?hl=en-WS&utm_source=chatgpt.com | not cited |
| 8 | https://developers.google.com/search/blog/2020/05/evaluating-page-experience?utm_source=chatgpt.com | not cited |
| and 13 more results | ||
Tracked brands mentioned
Answer 11 of 22, 14th Aug 2026, gpt-5-6, mentions 3 tracked brands
For a messy inherited site, I’d use a 3-layer combo rather than relying on one “SEO score”:
| Tool | Best for | How I’d use it |
|---|---|---|
| Google Search Console | What Google actually sees | Find indexing/crawl problems, sitemap issues, URL Inspection, and real-user Core Web Vitals. developers.google.com |
| Screaming Frog SEO Spider or Sitebulb | Deep technical crawl | Find broken links, redirect chains, canonicals, robots directives, orphan/duplicate pages, huge pages, JS-rendering problems, etc. Sitebulb is particularly nice for automatically prioritizing findings. sitebulb.com |
| Google PageSpeed Insights | Performance | Validate representative templates using both Lighthouse lab data and real-user CrUX data. Focus on LCP, INP and CLS rather than the raw performance score. developers.google.com |
| Ahrefs Site Audit (optional) | Second opinion + SEO impact | Useful if you already have Ahrefs for links/rankings; its crawler covers technical issues, Core Web Vitals, indexability, redirects, JS/CSS, robots and sitemaps. ahrefs.com |
My preferred workflow
1. Search Console first — establish actual impact.
Look at indexing, sitemap coverage, crawl/indexing anomalies and Core Web Vitals. Don't blindly fix every warning: prioritize issues affecting important pages or lots of URLs. Google explicitly recommends prioritizing Poor CWV issues by affected URL count/importance. developers.google.comsupport.google.com
2. Crawl the whole site — build the technical inventory.
Use Screaming Frog or Sitebulb with JavaScript rendering enabled if the site is JS-heavy. Export the issues rather than treating the crawler's overall score as gospel.
3. PageSpeed on templates, not every URL.
Test representative pages:
- homepage
- main landing/category template
- product/service page
- blog/article template
- worst-performing important page
Then look for template-level causes: oversized images, render-blocking resources, excessive JS, third-party scripts, poor caching, layout shifts, etc. PSI combines lab diagnostics with real-user data, which makes it useful for distinguishing “technically slow” from “actually hurting users.” developers.google.com
4. Turn findings into a priority matrix.
I'd rank fixes roughly:
P0 — indexing/crawl blockers
- accidental
noindex - robots.txt blocking important content
- canonical pointing somewhere wrong
- important pages returning 4xx/5xx
- redirect chains/loops
- huge crawl traps
- important content inaccessible without JS
P1 — widespread performance problems
- poor CWV on major templates
- enormous page/resource weight
- excessive JS
- slow server response
- image problems affecting hundreds/thousands of URLs
P2 — structural SEO problems
- orphan pages
- duplicate URLs
- bad internal linking
- thin/duplicate templates
- sitemap inconsistencies
P3 — cleanup
- missing metadata
- minor redirects
- low-impact warnings
- cosmetic SEO issues
The key is to join crawl data with Search Console data. A broken canonical on one obscure page isn't equivalent to the same problem across your 40,000 highest-value URLs.
If I inherited the site tomorrow, my minimum stack would be Search Console + Screaming Frog + PageSpeed Insights. If you want the crawler itself to help more with prioritization and reporting, I'd swap Screaming Frog for Sitebulb. sitebulb.com
Behind this answer: 2 searches, 19 results retrieved, 3 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://sitebulb.com/product/performance/?utm_source=chatgpt.com | not cited |
| 2 | https://sitebulb.com/product/?utm_source=chatgpt.com | not cited |
| 3 | https://sitebulb.com/?utm_source=chatgpt.com | not cited |
| 4 | https://sitebulb.com/product/single-page-analysis/?utm_source=chatgpt.com | not cited |
| 5 | https://support.google.com/webmasters/answer/10351509?hl=en&utm_source=chatgpt.com | not cited |
| 6 | https://support.google.com/webmasters/answer/9205520?hl=en-EN&utm_source=chatgpt.com | not cited |
| 7 | https://developers.google.com/search/blog/2020/11/search-console-crawl-stats-report?utm_source=chatgpt.com | not cited |
| 8 | https://developers.google.com/speed/docs/insights/v5/get-started?hl=en&utm_source=chatgpt.com | not cited |
| and 11 more results | ||
Tracked brands mentioned
Answer 12 of 22, 13th Aug 2026, gpt-5-6, mentions 3 tracked brands
If you inherited a messy site, I’d use a stack rather than one “SEO audit” tool. The key is separating crawl/indexation problems from real-user performance problems, then using Search Console to validate what Google actually sees.
My preferred combo
- Screaming Frog SEO Spider — your deep technical crawl
Best for finding the shape of the mess: redirect chains, broken links, duplicate URLs, canonicals, orphan-ish structures, pagination, robots directives, status codes, JavaScript-rendered links, etc. It gives you much more granular control than most cloud audits.
- Google Search Console — source of truth for Google
Use it alongside the crawler, especially Page Indexing, Crawl Stats, Sitemaps, and Core Web Vitals. A crawler can tell you that something looks wrong; GSC tells you what Google is actually indexing/crawling.
- PageSpeed Insights + Lighthouse — performance diagnosis
Use this for the performance side rather than treating an SEO tool's “page speed” score as definitive. Look at LCP, INP, CLS, plus the underlying opportunities/diagnostics. Prioritize issues affecting important templates/pages, not every individual URL.
- Ahrefs Site Audit or Semrush Site Audit — prioritization and monitoring
You don't really need both initially. Both turn a huge pile of technical problems into categorized/prioritized issues. Ahrefs currently checks 170+ technical/on-page issues, including Core Web Vitals, slow pages, indexability, redirects, robots, sitemaps, JS/CSS and links. ahrefs.com Semrush similarly provides severity-ranked issues and thematic reports for crawlability, robots.txt, internal linking, performance and Core Web Vitals. www.semrush.com
The workflow I'd use
Screaming Frog → Search Console → PSI/Lighthouse → Ahrefs/Semrush
Then build one prioritized backlog:
| Priority | Look for |
|---|---|
| 🔴 P0 | Important URLs blocked from crawling/indexing, accidental noindex, bad robots.txt, broken canonicalization, server errors, sitemap disasters |
| 🔴 P1 | Large numbers of 3xx/4xx URLs, redirect chains, orphan/high-click-depth pages, duplicate URL variants |
| 🟠 P1 | Slow templates, poor LCP/INP, huge JS/CSS, oversized images, render-blocking resources |
| 🟠 P2 | Internal-link architecture, duplicate/thin pages, unnecessary crawl traps/parameters |
| 🟡 P3 | Metadata, headings, schema warnings and other low-impact hygiene |
Important: don't blindly fix the tool's highest-count issues first. A single accidental noindex on 10,000 product pages can matter vastly more than 50,000 missing meta descriptions.
If you want the leanest setup, I'd go Screaming Frog + Google Search Console + PageSpeed Insights. Add Ahrefs or Semrush when you need better prioritization, historical monitoring, backlink context, or a team-friendly issue backlog. Ahrefs, for example, explicitly categorizes errors/warnings/notices by importance. ahrefs.com
For a genuinely inherited mess, that three-tool core is enough to find and prioritize most performance + crawlability problems without buying five overlapping SEO suites.
Behind this answer: 4 searches, 30 results retrieved, 3 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://help.ahrefs.com/en/articles/2166853-what-is-the-optimal-crawl-speed-i-should-set-in-site-audit?utm_source=chatgpt.com | not cited |
| 2 | https://help.ahrefs.com/en/articles/1431155-how-do-i-finish-crawling-my-website-faster-in-site-audit?utm_source=chatgpt.com | not cited |
| 3 | https://help.ahrefs.com/en/articles/9082329-how-should-i-configure-my-site-audit-settings?utm_source=chatgpt.com | not cited |
| 4 | https://help.ahrefs.com/en/articles/10957674-how-always-on-audit-works?utm_source=chatgpt.com | not cited |
| 5 | https://ahrefs.com/seo/glossary/crawlability?utm_source=chatgpt.com | not cited |
| 6 | https://ahrefs.com/seo-audit-tool?utm_source=chatgpt.com | not cited |
| 7 | https://help.ahrefs.com/en/articles/1399529-how-to-use-site-audit-filters-in-page-explorer-and-link-explorer?utm_source=chatgpt.com | not cited |
| 8 | https://ahrefs.com/academy/how-to-use-ahrefs/site-audit/overview?utm_source=chatgpt.com | not cited |
| and 22 more results | ||
Tracked brands mentioned
Answer 13 of 22, 12th Aug 2026, gpt-5-6, mentions 5 tracked brands
For an inherited “messy” site, I’d avoid relying on one SEO audit score. The best workflow is a stack where each tool answers a different question: Can Google crawl it? Can users load it fast? Which fixes matter most? www.techwench.com
A practical audit stack
| Goal | Tool | What it finds | Priority signal |
|---|---|---|---|
| Google’s actual view | Google Search Console | Indexing issues, crawl errors, pages excluded, sitemap problems, search performance | Highest — Google is telling you what it struggles with |
| Deep crawl / technical archaeology | Screaming Frog SEO Spider | Broken links, redirect chains, canonicals, duplicate titles, orphaned pages, robots directives, JS rendering issues | Highest for inherited sites |
| Performance diagnosis | Google PageSpeed Insights | Core Web Vitals, JS/CSS bottlenecks, image problems, loading issues | Prioritize pages/templates affecting many URLs |
| Real-world speed monitoring | WebPageTest | Waterfalls, TTFB, render behavior, third-party scripts | Finds developer-level causes |
| SEO health dashboard | Ahrefs Site Audit or Semrush Site Audit | Broad technical checks + reporting | Good for organizing issues for stakeholders |
| Server crawl behavior | Log file analyzer (e.g., Screaming Frog Log File Analyser) | What bots actually crawl, wasted crawl budget, ignored pages | Critical for large sites |
The order I’d run them
1. Establish the crawl/indexing baseline
- Connect Google Search Console.
- Export:
- Pages indexed vs excluded
- Crawl stats
- Sitemap status
- “Discovered — currently not indexed”
- “Crawled — currently not indexed”
2. Crawl the whole site
Run Screaming Frog with:
- JavaScript rendering enabled (if the site is a React/Vue/Angular app)
- Crawl canonicals
- Crawl pagination
- Connect Search Console + Analytics data
Look first for:
- 5xx errors
- 404s with internal links
- redirect chains
- noindex pages receiving traffic
- canonical conflicts
- duplicate URL patterns
- huge numbers of low-value URLs
3. Overlay performance
Don’t chase every PageSpeed warning. Prioritize:
- Templates with thousands of URLs
- Pages getting organic traffic
- Pages failing Core Web Vitals
Typical high-impact fixes:
- Remove unnecessary JavaScript
- Optimize oversized images
- Improve server response time (TTFB)
- Fix layout shifts from ads/fonts/images
- Reduce third-party scripts
4. Build a fix priority score
A simple formula:
Priority = SEO impact × affected URLs × ease of fix
Examples:
| Issue | Impact | Fix priority |
|---|---|---|
| 10,000 URLs blocked by robots.txt accidentally | Huge | Immediate |
| 5,000 duplicate title tags | Medium | High |
| One slow blog image | Low | Low |
| 2-second improvement to shared product template | Huge | High |
My “inherited disaster” combo by budget
Free / lean
- Google Search Console
- Screaming Frog free version (up to its crawl limit)
- PageSpeed Insights
- WebPageTest
Small team
- Screaming Frog + Ahrefs or Semrush + Search Console
Enterprise / very large sites
- Screaming Frog or Sitebulb
- Ahrefs/Semrush
- Log file analysis
- Dedicated crawling platforms
The biggest mistake with inherited sites is spending weeks fixing dozens of warnings before finding the few structural problems (bad URL architecture, index bloat, template performance, internal linking failures) that create thousands of issues. lazymetrics.com
If you tell me the site type (WordPress, Shopify, custom app, enterprise CMS) and rough size (hundreds vs millions of URLs), I can suggest a tighter audit sequence.
Behind this answer: 1 search, 9 results retrieved, 5 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://seodiag.com/en/blog/best-technical-seo-audit-tools-2026.html?utm_source=chatgpt.com | not cited |
| 2 | https://rankture.com/blog/seo-audit-tools-comparison?utm_source=chatgpt.com | not cited |
| 3 | https://www.semrush.com/blog/seo-audit-tools/?msockid=2a02a7961ed961f834cbb16c1fdc604f&utm_source=chatgpt.com | not cited |
| 4 | https://clickraven.com/best-seo-audit-tool/?utm_source=chatgpt.com | not cited |
| 5 | https://www.experte.com/seo/seo-audit-tools?utm_source=chatgpt.com | not cited |
| 6 | https://www.fueler.io/blog/best-seo-audit-tools-for-technical-seo?utm_source=chatgpt.com | not cited |
| 7 | https://worldmetrics.org/best/site-audit-software/?utm_source=chatgpt.com | not cited |
| 8 | https://trackseo.pro/blog/website-audit-tools?utm_source=chatgpt.com | not cited |
| and 1 more results | ||
Tracked brands mentioned
Answer 14 of 22, 11th Aug 2026, gpt-5-6, mentions 2 tracked brands
For a messy inherited site, I’d use a four-tool stack, with each tool answering a different question:
| Tool | Best for | What I’d use it for |
|---|---|---|
| Google Search Console | What Google actually sees | Indexing/exclusions, crawl issues, sitemap status, Core Web Vitals, search performance |
| Screaming Frog SEO Spider | Deep technical crawl | 3xx/4xx/5xx, redirect chains, canonicals, robots directives, orphan-ish pages, internal links, duplicate metadata, JS rendering |
| Ahrefs Site Audit | Prioritization + monitoring | Consolidates hundreds of technical issues and shows affected URLs/severity; it currently checks 170+ issue types. ahrefs.com |
| PageSpeed Insights | Performance | Real-user + lab performance, especially LCP, INP and CLS. |
The combo I'd actually run
1. Search Console first → establish what's broken in Google's eyes.
Don't start by fixing every crawler warning. Look at:
- Indexed vs. excluded URLs
- “Crawled — currently not indexed”
- “Discovered — currently not indexed”
- 404/soft-404 patterns
- sitemap coverage
- Core Web Vitals
- Pages receiving impressions but unexpectedly poor clicks
2. Screaming Frog → find the underlying technical causes.
Crawl the site with JavaScript rendering if it's a JS-heavy site. Export the important buckets: status codes, canonicals, robots directives, redirects, internal links, XML sitemap URLs, and page depth. Screaming Frog is particularly useful because you can get down to URL-by-URL evidence rather than relying on an overall “SEO score.” www.techradar.com
3. PageSpeed Insights → prioritize performance separately.
Don't treat the Lighthouse score as your entire performance diagnosis. Pay particular attention to field Core Web Vitals, then use the lab diagnostics to identify causes such as oversized images, excessive JS, render-blocking resources, or slow server response. pagespeedinsights.dev
4. Ahrefs Site Audit → turn the mess into a recurring backlog.
It's useful after the initial forensic crawl because it groups issues, identifies how many URLs are affected, and lets you monitor whether the technical debt is actually shrinking. ahrefs.com
How I'd prioritize the fixes
I'd score findings roughly like this:
P0 — fix immediately
- Important pages blocked by
robots.txt/noindex - Wrong canonicals
- 5xx errors
- Major redirect/indexation problems
- Key pages absent from Google's index
- Severe CWV problems across important templates
P1 — high impact
- Broken internal links
- Large redirect chains
- Bad sitemap URLs
- Huge JS/CSS payloads
- Slow server response
- Important pages buried extremely deep
P2 — cleanup
- Duplicate/missing titles
- Meta descriptions
- Image alt text
- Minor schema issues
- Low-value duplicate URLs
The key is cross-referencing tools rather than blindly following their scores. For example, if Screaming Frog says 20,000 URLs have an issue but Search Console shows only 30 of those URLs matter to Google/search traffic, you probably shouldn't make that your first 20,000-page project.
For a truly chaotic site, my preferred workflow is:
Search Console → Screaming Frog → PageSpeed Insights → Ahrefs → prioritized Jira/Sheets backlog → re-crawl → Search Console validation.
That gives you Google's perspective + raw crawl data + real performance data + ongoing monitoring, without letting any one SEO tool dictate the priorities.
Behind this answer: 4 searches, 30 results retrieved, 2 tracked brands mentioned
Searches ChatGPT ran
Search results ChatGPT retrieved
| # | Page | Used in the answer |
|---|---|---|
| 1 | https://help.ahrefs.com/en/articles/1431155-how-do-i-finish-crawling-my-website-faster-in-site-audit?utm_source=chatgpt.com | not cited |
| 2 | https://help.ahrefs.com/en/articles/2166853-what-is-the-optimal-crawl-speed-i-should-set-in-site-audit?utm_source=chatgpt.com | not cited |
| 3 | https://help.ahrefs.com/en/articles/3119402-how-are-crawl-credits-in-site-audit-spent?utm_source=chatgpt.com | not cited |
| 4 | https://help.ahrefs.com/en/collections/87920-site-audit?utm_source=chatgpt.com | not cited |
| 5 | https://help.ahrefs.com/en/articles/9082329-how-should-i-configure-my-site-audit-settings?utm_source=chatgpt.com | not cited |
| 6 | https://ahrefs.com/site-audit?r=imc-sat&utm_source=chatgpt.com | not cited |
| 7 | https://help.ahrefs.com/en/articles/1608544-why-doesn-t-ahrefs-crawl-or-index-my-website-fully?utm_source=chatgpt.com | not cited |
| 8 | https://help.ahrefs.com/en/articles/1399529-how-to-use-site-audit-filters-in-page-explorer-and-link-explorer?utm_source=chatgpt.com | not cited |
| and 22 more results | ||
Tracked brands mentioned
Answer 15 of 22, 10th Aug 2026, gpt-5-6, mentions 3 tracked brands
For an inherited, messy site, I’d use a combination rather than one “SEO audit” tool:
- Google Search Console — start here for indexing, crawl errors, sitemap coverage, Core Web Vitals, and pages Google actually knows about.
- Screaming Frog SEO Spider — your deep technical crawl: broken links, redirects, canonicals, duplicate/missing metadata, orphan-ish patterns, pagination, hreflang, etc.
- PageSpeed Insights / Lighthouse — diagnose performance and Core Web Vitals at the page level.
- Ahrefs or Semrush — layer on backlinks, organic visibility, keyword cannibalization, and competing pages so you can distinguish technical noise from fixes that could actually affect traffic.
- Chrome DevTools — useful when Lighthouse tells you something is slow but you need to identify the actual JS, CSS, image, or network bottleneck.
The prioritization workflow I'd use
First: crawlability/indexation
- GSC: indexing anomalies and excluded pages
- Screaming Frog: 4xx/5xx, redirect chains, robots directives, canonicals, XML sitemap consistency
- Fix anything preventing important URLs from being crawled/indexed.
Second: performance
- PSI/Lighthouse: LCP, INP, CLS and resource-level diagnostics
- Identify sitewide problems first—bloated JS, render-blocking resources, huge images, poor caching—rather than optimizing individual low-value pages.
Third: SEO impact
- Ahrefs/Semrush + GSC: determine which affected URLs have rankings, impressions, links, or conversions.
- Prioritize fixes that combine high SEO impact × many affected URLs × reasonable implementation effort.
A useful spreadsheet column set is:
Issue → URLs affected → Indexable? → Traffic/impressions → SEO impact → Performance impact → Effort → Priority
If you only want three tools: Search Console + Screaming Frog + PageSpeed Insights. That gives you the strongest foundation for separating “the site is messy” from “this mess is actually hurting Google crawling, indexing, or users.”
Behind this answer: no web searches, 3 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Week of 3rd to 9th Aug 2026
Answer 16 of 22, 9th Aug 2026, gpt-5-6, mentions 3 tracked brands
For an inherited, messy site, I’d use a stack rather than one SEO crawler:
- Google Search Console — start here for actual Google crawl/indexing problems: indexing, crawl stats, sitemap issues, Core Web Vitals, and pages Google is excluding.
- Screaming Frog SEO Spider — your deep technical inventory. Crawl the site and surface broken links, redirect chains, canonical problems, duplicate/near-duplicate pages, orphan-ish patterns, robots directives, pagination, and huge page inventories.
- PageSpeed Insights + Lighthouse — prioritize performance fixes using field data (when available) plus lab diagnostics. Pay particular attention to LCP, INP, CLS, render-blocking resources, JavaScript, image weight, and caching.
- Ahrefs or Semrush — useful for putting technical issues in business context: organic traffic, important landing pages, backlinks, rankings, and competing sites. Their site-audit crawlers are useful, but I wouldn't treat their severity scores as gospel.
- Google Analytics — connect technical problems to actual user/traffic impact. A slow or broken page with meaningful organic traffic should outrank an obscure technical error on a page nobody visits.
How I'd prioritize
Create a simple matrix:
Priority = SEO impact × affected pages × business importance × confidence/ease of fix
Then roughly work in this order:
P0 — blocking
- Important pages accidentally
noindex - Robots.txt blocking critical sections
- Server errors / widespread 5xxs
- Major canonical or redirect mistakes
- Sitemap/indexation failures
- Pages effectively inaccessible to crawlers
P1 — high-impact
- Slow templates affecting lots of organic landing pages
- Poor Core Web Vitals on important URLs
- Large redirect chains
- Broken internal links
- Duplicate/canonicalization problems
- Massive crawl waste from parameters, faceted navigation, or infinite URL variations
P2 — cleanup
- Missing/duplicate titles and descriptions
- Image optimization
- Schema improvements
- Minor heading/metadata inconsistencies
- Low-value duplicate pages
The key is don't blindly fix every warning. A crawler can produce thousands of "issues" that have little SEO consequence. Cross-reference each finding with Search Console + traffic + important URLs/templates before assigning engineering work.
If you tell me roughly what CMS/stack the site uses and its size (e.g. 5k vs. 500k URLs), I can suggest the leanest tool combination and an audit workflow.
Behind this answer: no web searches, 3 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Answer 17 of 22, 8th Aug 2026, gpt-5-6, mentions 3 tracked brands
For an inherited, messy site, I’d use a stack rather than one SEO crawler:
- Google Search Console — crawl/indexing reality: indexing issues, sitemap status, Core Web Vitals, manual actions, and pages Google actually knows about.
- Screaming Frog SEO Spider — technical inventory: broken links, redirects, canonicals, duplicate/missing metadata, orphan-ish patterns, pagination, status codes, and crawl-depth issues.
- PageSpeed Insights + Lighthouse — performance: LCP, INP, CLS, render-blocking resources, image/script problems, and actionable performance diagnostics.
- Ahrefs Site Audit or Semrush Site Audit — prioritization layer: groups technical issues, tracks changes, and adds broader SEO context/backlinks.
- Google Rich Results Test — validate structured data on important templates.
How I'd prioritize the mess
P0 — Can Google crawl/index it?
4xx/5xxerrors- accidental
noindex - robots.txt blocking important URLs
- bad canonicals
- redirect chains/loops
- XML sitemap problems
- duplicate URL variants
P1 — Is it wasting crawl budget or creating index bloat?
- faceted/filter URLs
- parameter duplicates
- huge numbers of thin/duplicate pages
- infinite crawl spaces
- orphaned important pages
- excessive redirects
P2 — Is it slow for users?
- poor LCP/INP/CLS
- oversized images
- excessive JavaScript
- render-blocking CSS/JS
- slow server response/TFFB
P3 — Cleanup
- titles/meta descriptions
- headings
- internal linking
- structured data
- image alt text
Best combo for most sites: Search Console + Screaming Frog + PageSpeed Insights. Add Ahrefs or Semrush when you need ongoing monitoring, backlink context, and a more convenient prioritization dashboard.
The important part is not treating the audit's issue count as the priority. Cross-reference crawler findings with Search Console's indexed/traffic data and the business value of the affected URLs. A single broken canonical on 500 important product pages matters much more than 5,000 missing alt attributes.
Behind this answer: no web searches, 3 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Answer 18 of 22, 7th Aug 2026, gpt-5-5, mentions 4 tracked brands
For a messy inherited site, the best approach is not one SEO crawler—it’s a stack where each tool answers a different question:
- What’s broken? (technical discovery)
- What hurts rankings/users most? (prioritization)
- What should we fix first? (impact vs effort)
A strong combo:
| Need | Tool | Best for | Output you should act on |
|---|---|---|---|
| Full technical crawl | Screaming Frog SEO Spider | Broken links, redirects, canonicals, duplicate pages, metadata, indexability | Export issues by severity |
| Google’s view of the site | Google Search Console | Crawl errors, indexing problems, manual actions, search performance | Pages Google can’t/won’t index |
| Real user performance | Google PageSpeed Insights | Core Web Vitals, JS/CSS/image bottlenecks | LCP, INP, CLS fixes |
| Site architecture + large-scale crawling | Sitebulb | Visual audits, crawl maps, prioritization | “Why this matters” reports |
| Competitive SEO gaps | Semrush or Ahrefs | Backlinks, lost traffic, keyword opportunities | Pages worth saving/improving |
| Structured data validation | Rich Results Test | Schema problems | Fix eligible rich results |
| Log-level crawl analysis (advanced) | Screaming Frog Log File Analyser | What Googlebot actually crawls | Crawl waste and orphan pages |
Practical audit workflow
Phase 1 — Crawl the site
Run Screaming Frog or Sitebulb first.
Create buckets:
Critical
- Pages blocked by robots.txt accidentally
noindexon important pages- Broken internal links
- 5xx errors
- Wrong canonical tags
- Massive redirect chains
High impact
- Duplicate titles/meta descriptions
- Thin pages
- Orphan pages
- Poor internal linking
- Slow templates
Cleanup
- Old URLs
- Low-value archives
- Minor metadata issues
Phase 2 — Confirm with Google data
In Search Console, compare the crawler findings with reality:
Look at:
- Pages indexed vs submitted
- Excluded pages
- Crawl stats
- Pages losing clicks/impressions
- Core Web Vitals report
A common inherited-site surprise: the crawler finds thousands of “problems,” but only 50 pages matter for search traffic.
Phase 3 — Prioritize fixes with an impact matrix
A simple scoring model:
Priority = (Traffic potential × SEO impact × Confidence) ÷ Effort
Examples:
| Issue | Priority |
|---|---|
| 500 important pages blocked by robots.txt | 🔴 Immediate |
| Homepage loads 8 seconds | 🔴 Immediate |
| 20,000 duplicate tag pages indexed | 🔴 High |
| Missing alt text on old blog images | 🟡 Medium |
| Minor title rewrites on pages with no impressions | 🟢 Low |
If I inherited the site tomorrow, my first-week stack would be:
- Google Search Console — understand Google’s problems
- Screaming Frog SEO Spider — full crawl inventory
- Google PageSpeed Insights — performance baseline
- Sitebulb — easier prioritization/reporting
- Ahrefs or Semrush — decide which pages are worth saving
One extra step that often pays off: before deleting or consolidating anything, export all existing URLs, rankings, backlinks, and analytics landing pages. Messy inherited sites often contain valuable pages hidden under poor architecture.
Behind this answer: no web searches, 4 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Answer 19 of 22, 7th Aug 2026, gpt-5-5, mentions 3 tracked brands
If you've inherited a messy website, the biggest challenge usually isn't finding issues—it's figuring out which ones are actually blocking traffic and which can wait. A good stack combines crawler data, performance diagnostics, and search engine data.
A practical combination looks like this:
| Goal | Best tool | Why it matters |
|---|---|---|
| Crawlability | Screaming Frog SEO Spider | Finds broken links, redirects, orphan pages, duplicate titles, canonicals, indexability issues |
| Google indexing | Google Search Console | Shows what Google actually indexed, crawl errors, impressions, and pages losing traffic |
| Performance | PageSpeed Insights | Measures Core Web Vitals and gives actionable optimization recommendations |
| Technical quality | Lighthouse | Audits performance, accessibility, SEO, and best practices |
| Site monitoring | Ahrefs Site Audit or Semrush Site Audit | Prioritizes issues and tracks improvements over time |
My recommended workflow
1. Crawl the whole site first
Run Screaming Frog against every URL.
Look for high-impact problems such as:
- Pages blocked by robots.txt accidentally
- noindex pages that should rank
- Broken internal links (404s)
- Redirect chains
- Canonical mistakes
- Duplicate content
- Orphan pages
- Thin pages
- Missing H1/title tags
This gives you a complete inventory before making changes.
2. Compare against Google's view
Open Search Console and compare.
Focus on:
- Pages discovered but not indexed
- Crawled but not indexed
- Soft 404s
- Duplicate without user-selected canonical
- Server errors
- Sitemap issues
These are often more important than issues a crawler finds because Google is telling you exactly what it's struggling with.
3. Measure performance
Don't test every page.
Instead test:
- Homepage
- Top landing pages
- Highest traffic blog post
- Product/category templates
- Search results pages
Use PageSpeed Insights to identify:
- Largest Contentful Paint (LCP)
- Cumulative Layout Shift (CLS)
- Interaction to Next Paint (INP)
- Render-blocking resources
- Image optimization opportunities
Template-level fixes often improve hundreds of pages at once.
4. Find internal linking problems
Screaming Frog is excellent for:
- Click depth
- Pages with few internal links
- Pages receiving no internal PageRank
- Redirected internal links
Deep pages (5–7 clicks from the homepage) are often good candidates for improved internal linking.
5. Prioritize by impact
Instead of fixing thousands of warnings, sort work into tiers.
Tier 1 (highest ROI)
- Pages accidentally blocked from indexing
- Server errors (5xx)
- Broken internal links
- Incorrect canonicals
- Slow page templates
- XML sitemap issues
Tier 2
- Duplicate metadata
- Missing meta descriptions
- Redirect chains
- Large images
- JavaScript rendering bottlenecks
Tier 3
- Minor title tweaks
- Missing alt text (unless accessibility is a priority)
- Small HTML validation issues
- Cosmetic SEO warnings
If you have a budget
A strong professional stack is:
This combination covers crawling, indexing, backlinks, content opportunities, and performance.
If your budget is $0
You can still get a lot done with:
- Google Search Console
- PageSpeed Insights
- Chrome DevTools Lighthouse
- Free version of Screaming Frog SEO Spider (up to 500 URLs)
That setup is enough to identify most major crawlability and performance issues on a small to medium-sized site.
A useful rule when prioritizing fixes is to score each issue by Impact × Scale × Effort. For example, a single template fix that improves Core Web Vitals across 2,000 pages is usually far more valuable than editing 500 individual page titles. Likewise, resolving an indexing problem affecting an entire section of the site should generally come before metadata clean-up. This approach helps you focus on changes that are most likely to improve search visibility and user experience with the least engineering effort.
Behind this answer: no web searches, 3 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Answer 20 of 22, 7th Aug 2026, gpt-5-5, mentions 3 tracked brands
If you've inherited a messy website, the biggest challenge usually isn't finding issues—it's figuring out which ones actually matter first. A good workflow combines tools that answer different questions:
- Can search engines crawl and index the site?
- Is the site technically healthy?
- Is it fast for users?
- Which problems have the biggest SEO impact?
A practical stack looks like this:
| Purpose | Recommended tool | Why it matters |
|---|---|---|
| Crawl the entire site | Screaming Frog SEO Spider | Finds broken links, redirect chains, duplicate content, missing tags, orphan pages (with analytics integration), crawl depth, etc. |
| Google's view | Google Search Console | Index coverage, crawl stats, Core Web Vitals, sitemap status, pages Google actually sees. |
| Performance | PageSpeed Insights | Real-user Core Web Vitals plus lab performance diagnostics. |
| Site-wide technical audit | Ahrefs Site Audit or Semrush Site Audit | Automatically groups issues by severity and impact. |
| Log analysis (large sites) | Screaming Frog Log File Analyser | Shows what search engine bots actually crawl versus what you expect. |
| Schema validation | Schema Markup Validator | Verifies structured data implementation. |
My preferred workflow
1. Start with Google Search Console
Before crawling anything:
- Indexed vs. excluded pages
- Crawl anomalies
- Core Web Vitals
- Manual actions
- Sitemap errors
This tells you where Google is already struggling.
2. Crawl the site with Screaming Frog
This usually uncovers hundreds or thousands of issues.
Focus on:
- 4xx and 5xx pages
- redirect chains
- duplicate titles
- duplicate H1s
- missing canonicals
- noindex mistakes
- orphan pages
- thin pages
- pages over 3–4 clicks deep
Export everything into Excel or Google Sheets.
3. Compare crawl vs. Search Console
This is where priorities become obvious.
Examples:
- Crawled but not indexed → investigate quality or duplication.
- Indexed but not linked internally → improve architecture.
- Broken URLs receiving backlinks → high priority.
- Pages Google can't reach → fix immediately.
4. Measure templates with PageSpeed Insights
Don't test every page.
Instead test representative templates:
- homepage
- category page
- product page
- article
- landing page
If one product template is slow, hundreds of pages may share the same problem.
5. Run a site audit
Ahrefs or Semrush help rank issues by severity.
Ignore "SEO score."
Instead look for:
- internal linking problems
- crawl waste
- duplicate content
- canonical conflicts
- JavaScript rendering issues
- slow templates
6. Prioritize by impact, not issue count
A simple framework is:
| Priority | Examples |
|---|---|
| Critical (do first) | 5xx errors, robots.txt blocking important pages, accidental noindex, canonical mistakes, broken internal navigation |
| High | Slow templates, redirect chains, orphan pages, poor internal linking, XML sitemap errors |
| Medium | Duplicate titles, missing meta descriptions, oversized images, missing alt text |
| Low | Minor HTML validation issues, tiny CSS savings, low-value warnings |
For very large websites (10k+ pages)
Add:
- server log analysis
- crawl budget analysis
- JavaScript rendering checks
- XML sitemap validation
- indexation ratio (Indexed ÷ Crawlable pages)
These often reveal issues that page-level audits miss.
A simple 80/20 tool stack
If you only want four tools that cover most technical SEO work:
- Google Search Console
- Screaming Frog SEO Spider
- PageSpeed Insights
- Ahrefs Site Audit (or Semrush Site Audit)
That combination gives you:
- Google's perspective on indexing and crawling
- a complete technical crawl
- real-world performance data
- prioritized recommendations across the site
It's a strong balance of depth and efficiency for taking over an existing website and deciding what to fix first.
Behind this answer: no web searches, 3 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Answer 21 of 22, 7th Aug 2026, gpt-5-5, mentions 4 tracked brands
If you've inherited a messy website, the goal isn't to find every issue—it's to identify the small number of fixes that will have the biggest impact on crawling, indexing, and user experience. A combination of complementary tools works much better than relying on a single SEO platform.
Here's a practical stack and what each tool is best at:
| Goal | Tool | Why use it |
|---|---|---|
| Crawl the site like a search engine | Screaming Frog SEO Spider | Finds broken links, redirects, duplicate pages, canonicals, orphan pages, crawl depth, metadata issues |
| Indexing & Google data | Google Search Console | Shows what Google is indexing, crawl errors, Core Web Vitals, sitemap status, queries, coverage |
| Performance | PageSpeed Insights + Lighthouse | Measures Core Web Vitals and identifies render-blocking resources, JS, images, CSS issues |
| Technical monitoring | Ahrefs Site Audit or Semrush Site Audit | Gives an overall technical health score and prioritizes issues |
| Log analysis (advanced) | Screaming Frog Log File Analyser or Splunk/ELK | Reveals how search engine bots actually crawl the site |
| Visual crawl mapping | Sitebulb | Easier than raw crawl data for understanding site architecture |
My recommended workflow
1. Crawl the entire site first
Run Screaming Frog.
Focus on:
- 404 pages
- redirect chains
- redirect loops
- orphan pages
- duplicate titles
- duplicate H1s
- missing canonicals
- noindex mistakes
- pages deeper than 4–5 clicks
- huge image files
- thin content
This becomes your master inventory.
2. Compare against Google Search Console
Look for mismatches:
- Pages Google can't crawl
- Indexed pages that shouldn't exist
- Important pages not indexed
- Soft 404s
- Server errors
- Sitemap problems
If Search Console and Screaming Frog disagree, investigate why.
3. Audit Core Web Vitals
Use Lighthouse/PageSpeed.
Prioritize:
- LCP
- INP
- CLS
Then identify why:
- oversized images
- too much JavaScript
- render-blocking CSS
- unused CSS
- third-party scripts
- slow server response
4. Examine site architecture
This often reveals the biggest wins.
Questions to answer:
- Are important pages buried?
- Are categories linked well?
- Does internal linking support key pages?
- Are orphan pages valuable?
- Is crawl depth too high?
Sitebulb's visualizations make these issues much easier to spot than spreadsheets.
5. Check crawl budget
This matters especially on larger sites (thousands of URLs).
Look for:
- faceted navigation
- parameter URLs
- duplicate archives
- search pages
- infinite calendars
- session IDs
These can waste crawler resources.
6. Review indexing
Common surprises include:
- indexable staging pages
- duplicate HTTP/HTTPS versions
- duplicate www/non-www versions
- pagination issues
- conflicting canonical tags
- accidental noindex directives
Priority order
I usually rank fixes like this:
High impact
- 5xx server errors
- robots.txt blocking important pages
- noindex on valuable pages
- broken canonical tags
- sitemap issues
- redirect loops
- broken internal links
- pages Google can't crawl
Medium impact
- Core Web Vitals failures
- oversized images
- JavaScript rendering problems
- duplicate metadata
- crawl depth
- weak internal linking
Lower impact
- missing alt text
- title length tweaks
- meta description rewrites
- heading hierarchy cleanup
- minor HTML validation issues
These matter, but they're rarely the first place to invest effort.
If I could only choose four tools
- Google Search Console — What Google actually sees and indexes.
- Screaming Frog SEO Spider — The most comprehensive technical crawler for finding structural issues.
- PageSpeed Insights/Lighthouse — Performance diagnostics tied to Core Web Vitals.
- Sitebulb — Excellent for visualizing site architecture and explaining technical issues to stakeholders.
That combination covers nearly everything needed to prioritize performance and crawlability improvements. Once you've identified the highest-impact issues, create a backlog that scores each fix by SEO impact × implementation effort (for example, on a 1–5 scale). This helps you tackle quick, high-value improvements before investing time in lower-impact refinements.
Behind this answer: no web searches, 4 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
Answer 22 of 22, 7th Aug 2026, gpt-5-5, mentions 5 tracked brands
For a messy inherited site, the best approach is not one SEO crawler—it’s a stack where each tool answers a different question:
- What’s broken? (technical audit)
- What matters most? (impact/prioritization)
- What is Google actually seeing? (crawl/index reality)
- What affects users? (performance)
A strong combo:
| Goal | Tool | Best for | What to prioritize from it |
|---|---|---|---|
| Full technical crawl | Screaming Frog SEO Spider | Finding broken links, redirects, duplicate pages, canonicals, thin pages, indexability issues | 5xx errors, accidental noindex, redirect chains, duplicate titles/canonicals |
| Google’s view | Google Search Console | Real crawl/index problems and search performance | Pages excluded from index, crawl stats, manual issues, Core Web Vitals |
| Performance | Google PageSpeed Insights | Core Web Vitals and page speed bottlenecks | LCP, INP, CLS, heavy JS/images, slow templates |
| Technical SEO + scoring | Ahrefs Webmaster Tools or Semrush Site Audit | Issue prioritization, backlinks, organic context | Errors affecting high-value pages first |
| Log-level crawl analysis (advanced) | Screaming Frog Log File Analyser | What bots actually crawl | Wasted crawl budget, ignored important URLs |
| Structured data | Google Rich Results Test | Schema validation | Broken schema on pages that can earn rich results |
| Architecture/content cleanup | Sitebulb | Visual audits and prioritization | Orphan pages, internal linking, crawl paths |
The workflow I’d use on an inherited site
Phase 1 — Establish reality (day 1)
- Connect Google Search Console.
- Export:
- indexed pages
- excluded pages
- sitemap URLs
- top landing pages
- Core Web Vitals issues
- Run a full Screaming Frog crawl.
Phase 2 — Fix “SEO fires” first
Rank issues by:
Impact × confidence × effort
Highest priority:
- Pages blocked from indexing accidentally
- Broken important URLs (404/5xx)
- Wrong canonicals
- Massive duplicate/thin page problems
- Slow templates affecting many URLs
- Bad internal linking
- Sitemap problems
Usually lower priority:
- Missing meta descriptions
- Minor heading issues
- Small image alt gaps
Phase 3 — Performance
Use PageSpeed Insights on:
- homepage
- top organic landing pages
- biggest templates (product/category/blog/etc.)
Don’t optimize one page only—fix the template causing thousands of slow pages.
Phase 4 — Crawl waste cleanup
Look for:
- millions of parameter URLs
- faceted navigation explosions
- redirect chains
- old CMS-generated pages
- orphan content
A practical “minimum stack” if you want to keep costs down
- Google Search Console (free)
- Screaming Frog (free up to 500 URLs; paid for large sites)
- PageSpeed Insights (free)
- Ahrefs Webmaster Tools (free site audit/backlink data)
If the site is very large (100k+ URLs)
Add:
- server log analysis
- Sitebulb or enterprise crawler
- a backlink tool (Ahrefs/Semrush/Majestic)
One thing that helps a lot with inherited sites: make a spreadsheet of all issues mapped to affected URL count + organic traffic/value, rather than fixing the crawler’s “critical” list blindly. A “critical” issue affecting 2 old pages is often less important than a “warning” affecting 50,000 URLs.
Behind this answer: no web searches, 5 tracked brands mentioned
ChatGPT answered from its own knowledge, without searching the web.
Tracked brands mentioned
See where your brand stands in AI
Track how AI talks about your brand, free. 40 prompts, no credit card.
Start tracking free
