Fact-checked 17 July 2026. A content audit is a decision process for an existing content library. It combines a URL inventory, performance data, qualitative review, overlap analysis, business context, and technical state to decide what to keep, refresh, expand, merge, redirect, repurpose, or retire.
Short answer: define the audit objective, preserve a baseline, inventory every relevant URL, join Search Console and analytics evidence, assess intent and information debt, identify duplicated jobs, assign one reversible action per URL, prioritise by impact and confidence, implement in controlled batches, and record a receipt so later performance can be interpreted.
Content inventory, content audit, and SEO audit
| Activity | Core question | Primary output |
|---|---|---|
| Content inventory | What content exists and where? | Structured URL/asset list |
| Content audit | Is each asset still useful, distinct, trustworthy, and worth maintaining? | Keep/refresh/merge/retire decisions |
| SEO audit | What conditions affect search discovery, understanding, and performance? | Prioritised technical, content, and architecture backlog |
The work overlaps, but a crawl is not a content audit. DMT’s SEO audit checklist covers the wider search surface; this guide goes deep on the content portfolio.
Step 1: define the audit objective
Choose the decision the audit must support:
- recover a declining cluster;
- reduce duplicate or cannibalising pages;
- prepare for a migration or redesign;
- find quick-win refreshes;
- improve trust, authorship, and sourcing;
- align the library with a new strategy;
- remove obsolete product or policy guidance;
- build a reliable baseline for future growth.
A traffic-recovery audit and a brand-compliance audit need different evidence. State the scope: whole site, blog, language, folder, template, cluster, or date range.
Step 2: preserve the baseline and rollback path
Before bulk changes, save:
- the current URL inventory and canonical map;
- Search Console query/page exports for a consistent period;
- GA4 organic landing-page and outcome data;
- current titles, excerpts, headings, content, links, status, and dates;
- redirect and sitemap state;
- backups or CMS revisions appropriate to the risk;
- a change manifest with owners and approval.
Deleting or merging content without this evidence makes rollback and post-change analysis needlessly difficult.
Step 3: build the content inventory
Use one row per canonical asset or URL. Recommended fields:
| Group | Fields |
|---|---|
| Identity | URL, canonical, title, ID, status, content type, template |
| Ownership | Author, reviewer, team, business owner |
| Dates | Published, modified, fact-checked, last audited |
| Scope | Audience, reader job, primary topic, intent, cluster, funnel role |
| Structure | Headings, word count as context, categories, tags, media |
| Links | Inbound/outbound internal links, external sources, known backlinks |
| Search state | Indexability, status, sitemap, robots, canonical consistency |
| Performance | GSC clicks, impressions, CTR, position, queries; GA4 sessions and outcomes |
| Decision | Debt, proposed action, priority, owner, approval, receipt |
Authenticated CMS exports often reveal drafts, private posts, and metadata a public sitemap misses. A crawler reveals link paths and rendered conditions. Search Console reveals the pages Google already associates with queries. Reconcile them instead of assuming one source is complete.
Step 4: collect performance evidence
Use a period long enough for the site’s scale and seasonality. Compare equivalent periods when demand changes across the year. Segment by page, query, country, device, and search appearance where useful.
Look for:
- high impressions with weak CTR in context of position;
- positions 4-20 for strategically relevant queries;
- click declines while impressions remain stable;
- impressions and clicks falling together, suggesting demand or visibility change;
- new query families the page only partially answers;
- multiple pages appearing for one query family;
- landing pages with engagement or outcomes disproportionate to search clicks;
- pages with no measurable traffic but strong link, compliance, service, or journey value.
Google’s Search Console documentation explains the query and page performance dimensions. Treat average position and CTR as aggregated diagnostics, not exact universal ranks.
Step 5: assess four kinds of content debt
Information debt
The page omits a mechanism, step, example, constraint, definition, or exception needed to solve the job.
Recency debt
Products, pricing, interfaces, sources, statistics, policies, or advice have materially changed. A visible old date is not recency debt when the facts remain durable; a new date does not remove debt when nothing was reviewed.
Trust debt
Claims lack strong sources, author responsibility is unclear, experience is asserted rather than demonstrated, or the page overstates certainty.
Decision debt
The reader understands the subject but cannot choose the next step because criteria, tradeoffs, risks, or acceptance tests are missing.
Google’s helpful-content self-assessment asks whether a page is original, substantial, trustworthy, produced with care, and satisfying for an intended audience. Use those questions as an editorial challenge, not a mechanical score.
Step 6: review intent and page responsibility
Write one sentence for every important page:
This page helps [reader] [complete job] by providing [distinct value], and it does not own [adjacent job].
If the sentence no longer matches the live content, the page has intent debt. If two URLs have essentially the same sentence, investigate overlap.
Use DMT’s SERP intent evidence sheet to compare current result types and the keyword decision map to assign each cluster to one URL.
Step 7: find cannibalisation and duplication
Potential signals include:
- high similarity in titles, slugs, headings, and full text;
- the same primary intent or focus keyword;
- two URLs alternating for the same valuable query family;
- separate pages created only for synonyms;
- outdated and current versions both indexable;
- category/tag archives duplicating editorial pages;
- internal anchors pointing inconsistently to multiple destinations.
Similar text does not always mean similar intent, and ranking fluctuation does not automatically prove harmful cannibalisation. Review the page jobs, result types, links, conversions, and the best canonical destination before acting.
The content decision matrix
| Action | Use when | Required checks |
|---|---|---|
| Keep | Useful, accurate, distinct, and performing its role | Verify links, metadata, owner, review date |
| Refresh | Intent remains valid but information, trust, recency, or presentation has debt | Preserve strengths; document material changes |
| Expand/reposition | Queries reveal a valuable adjacent job the URL can credibly own | Confirm page type and avoid stealing another URL’s role |
| Merge | Two or more pages divide one job without a useful distinction | Select the strongest destination; consolidate unique value |
| Redirect | A removed URL has a clear, relevant replacement | Map one-to-one where possible; update internal links and sitemap |
| Repurpose | The information is useful but a different format serves it better | Preserve the canonical source and avoid duplicate copies |
| Retire | Obsolete, harmful, unsupported, or no longer strategically useful with no replacement | Check links, outcomes, compliance, and intentional status |
| Noindex/restrict | The page is useful to users but should not appear in public search | Use the correct access/index control; do not block crawling before Google can read noindex |
Google’s canonical guidance explains consolidation signals, while its robots documentation distinguishes crawling from indexing controls. Choose the technical action after the editorial destination is clear.
How to prioritise audit actions
Score each action from 1 to 5:
- Potential impact: valuable demand, audience, journey, or risk affected;
- Evidence confidence: strength of data and diagnosis;
- Strategic fit: contribution to the site’s focused authority;
- Urgency: severity of inaccuracy, decline, or exposure;
- Effort: research, writing, design, development, review, and migration cost;
- Implementation risk: chance and consequence of a wrong change.
Priority = (impact × confidence × fit + urgency) / (effort + risk)
Then apply a sanity check. A legally or factually dangerous page can be urgent even with little traffic. A high-volume page can remain low priority if the site lacks credible expertise.
Write an action brief for every changed URL
| Field | Purpose |
|---|---|
| Current URL and ID | Stable identity |
| Current role | What the page is meant to do |
| Evidence | Performance, query, content, link, or policy signal |
| Debt | Information, recency, trust, decision, or duplication |
| Approved action | Keep, refresh, expand, merge, redirect, repurpose, retire |
| Preserve | Sections, links, rankings, assets, or conversions not to lose |
| Change | Exact content, metadata, link, or technical work |
| Owner/reviewer | Accountability |
| Rollback | Backup, revision, or redirect reversal |
| Acceptance test | What proves implementation is correct |
| Review window | When performance will be assessed |
Implement safely
- Back up the affected content and redirect/canonical state.
- Test the highest-risk template or migration changes outside production when possible.
- Publish small, attributable batches.
- Verify status, canonical, robots, title, content, links, media, and rendered mobile output.
- Update internal links to the preferred destination.
- Update the sitemap and caches as appropriate.
- Record the live URL, date, revision/post ID, and change summary.
- Monitor for crawl, indexing, analytics, or user-experience regressions.
DMT’s internal linking workflow provides a proposal and verification record for the link changes that follow merges and refreshes.
Measure the audit as a cohort
Tag audited URLs by action and implementation date. Compare them with their baseline and, where possible, a reasonable unchanged cohort. Review:
- indexation and canonical state;
- relevant query impressions, clicks, CTR, and position bands;
- organic landing sessions and meaningful outcomes;
- query consolidation after merges;
- broken-link, orphan, and redirect-chain counts;
- reader feedback and support questions;
- remaining refresh and trust debt.
A content audit is successful when the portfolio becomes more useful, coherent, maintainable, and measurable—not merely smaller.
Audit the portfolio at three levels
URL level
Decide whether the individual page fulfils its job. This is where accuracy, intent, sources, query fit, links, and the keep/refresh/merge decision live.
Cluster level
Check whether the subject has a useful pillar, distinct support jobs, coherent internal links, query ownership, and maintenance coverage. A weak URL may be strategically important because it closes a necessary journey; a strong URL may be isolated from the cluster.
Site level
Review primary purpose, content mix, templates, authorship, governance, taxonomies, and resource allocation. A library can contain individually acceptable pages while drifting away from a focused audience.
Connect the cluster review to DMT’s topical authority framework and the site-level decisions to the SEO content strategy. This prevents a spreadsheet of URL scores from replacing portfolio judgement.
Use action cohorts, not one blended total
Track refreshed, merged, redirected, kept, and retired URLs separately. A merged set should be evaluated for query consolidation and destination performance; a refresh should be evaluated against its own baseline; a retired set should be monitored for broken journeys and unexpected losses. Blending all actions can hide whether a particular intervention worked.
Common content audit mistakes
- Deleting every low-traffic page.
- Using word count as the quality score.
- Changing dates without reviewing facts.
- Merging pages without preserving unique value and relevant links.
- Redirecting unrelated URLs to the home page.
- Ignoring PDFs, videos, tools, landing pages, and non-blog assets.
- Auditing performance without intent or business context.
- Making irreversible bulk changes without approval and rollback.
- Finishing the spreadsheet but never implementing or measuring.
Frequently asked questions
How often should a content audit be done?
Review fast-changing and high-value pages continuously or quarterly, and run a broader portfolio audit at least annually. Audit sooner after migrations, strategy changes, unexplained declines, product retirements, or major source changes.
Should pages with zero traffic be deleted?
No. Check whether the page is indexable, useful to an audience, part of a journey, linked externally, required for support or compliance, or serving a niche query. Delete or retire only after a reasoned destination decision.
What is content decay?
Content decay is a useful label for declining relevance, accuracy, visibility, engagement, or outcomes over time. Diagnose the cause: demand, competition, intent, quality, links, technical conditions, and measurement can all change.
Does removing old content improve site quality?
Removing harmful, obsolete, duplicative, or unmaintainable content can improve the library and user experience. Google explicitly warns against adding or removing large amounts of content merely to make a site seem fresh. Make URL-level decisions for people and strategy.
What is the difference between refresh and rewrite?
A refresh preserves the page’s valid intent and strongest value while repaying specific debt. A rewrite replaces most of the approach because the promise, evidence, structure, or page type is no longer fit. Both should protect useful links, URL equity, and reader expectations.
About the author and methodology
Digital Marketer Tayeeb’s content-audit model combines authenticated content inventory, full-text duplicate checks, Search Console and GA4 evidence, information-debt review, explicit approvals, reversible changes, graph refreshes, and live receipts. The purpose is a better decision system, not a dramatic pruning count.