Google Content Effort Algorithm Update: Audit Guide
Google has a way to estimate how much effort went into a page, and that estimate doesn't care whether a person or a model wrote the words. The leaked contentEffort attribute says so directly. Our position: stop asking whether your content "sounds AI." Ask whether a competitor could reproduce it in an afternoon. If they could, it's thin, and it's a liability whoever typed it. The fix is the same for both kinds of thin content. Put things on the page that cost real work to produce, and cut or merge the pages that can't carry that weight.
Below is what the leak shows, what it doesn't, why AI filler and human filler land in the same bucket, and the audit we run to find thin pages before Google's systems find them for you.
What the leak actually shows
In May 2024, internal documentation for Google's Content Warehouse API became public. It lists attribute names and short descriptions for data Google stores about documents and sites. It doesn't include ranking weights, formulas, or any statement about which attributes feed live ranking today.
One attribute in a page-quality module is contentEffort, described in the documentation as an LLM-based effort estimation for article pages. Put simply, Google built a system that uses a language model to judge how much work a page represents.
That description should reset how you think about content. The attribute measures effort. It doesn't measure authorship method, keyword coverage, or word count. A 3,000-word article that restates the top five results scores low on effort. A 700-word piece with original measurements, photos from your own warehouse, and a decision table nobody else has published scores high.
The same documentation includes other attributes that point the same way. originalContentScore relates to originality. Site-level attributes such as siteFocusScore and siteRadius appear to describe how tightly a site sticks to its core topic and how far individual pages drift from it. Together they describe a system that evaluates individual pages and also the pattern of what a site publishes.
Google confirmed the documents were authentic. It also warned against drawing conclusions from information that was out of context, outdated, or incomplete. That warning is fair. We cover the limits of this reading near the end.
There is no single "Google content effort algorithm update"
People search for a "Google content effort algorithm update" as if it shipped on a specific date. It didn't. No announced update carries that name. What exists is a documented signal, plus a public record of Google moving steadily in one direction:
- The helpful content system launched in August 2022. It targeted content written for search engines instead of people.
- Google's February 2023 guidance on AI content said Google rewards quality content however it's produced and treats automation used mainly to manipulate rankings as spam.
- The March 2024 core update folded helpful content signals into core ranking. Alongside it came a spam policy on scaled content abuse, which covers producing large volumes of unoriginal pages "regardless of how it's created." That phrase covers AI, human writers, and templates alike.
- The Search Quality Rater Guidelines have long said the quality of main content depends on "the amount of effort, originality, and talent or skill" behind it. The lowest ratings go to content made with little effort and little originality.
Put the rater guidelines next to the leaked attribute and a pattern appears. Google tells human raters to judge effort, and it built a model-based estimator for the same quality. The rater data trains and checks the machine. You don't need a named update to act on that. The signal is already part of the system.
Why thin AI content and thin human content get the same treatment
"Thin human content" sounds like a contradiction. It shows up constantly in audits. Here are the three patterns we see most:
The freelancer rewrite. A writer gets a keyword, reads the first page of results, and produces a tidy summary. It's human, grammatical, and adds nothing. A model that has seen those same ten pages recognizes the rehash right away.
The template at scale. Location pages, "best X for Y" pages, and product descriptions with only the city or SKU swapped out. Someone wrote the template once. Search engines see hundreds of near-identical documents.
The AI draft with no editor. A model produces a fluent 1,500 words on a topic it knows only from its training data. That training data is the same public web the rehash writer read, so the output lands in the same place.
All three fail for the same reason. The page contains nothing that didn't already exist somewhere else. An effort estimator doesn't need to detect AI. It only needs to notice that the information, structure, and examples match what the rest of the index already holds. That's why "AI detection" is a distraction. Pages written with heavy AI assistance can score well if they carry original material. Pages written entirely by hand can score badly if they don't.
This also explains why "humanizing" AI text fails. Swapping synonyms, varying sentence length, and adding a few casual phrases changes how the text reads. It doesn't add effort. The page still has no new information.
What effort looks like to a machine
If a language model estimates effort, it responds to features it can observe on the page. Nobody outside Google knows the exact inputs. The rater guidelines and the logic of the task point to a short list of features that are expensive to fake:
- Original information. Your own test results, pricing breakdowns, measurements, survey data, or process documentation. Anything a reader can't get from the current top results.
- First-hand evidence. Photos you took, screenshots from real systems, annotated diagrams, video of the process. Stock imagery doesn't count.
- Specificity. Named products, exact settings, real constraints, and the edge case that breaks the standard advice. Vague pages read as low effort because producing them took little effort.
- Decisions, not just descriptions. A page that tells the reader what to choose and why, under which conditions, took more thought than one that lists options.
- Structure that serves the task. Comparison tables, step sequences, calculators, and checklists built for the reader's actual job.
- Accountability. A clear source for the claims and a reason to trust that source on this topic.
Here's an example from our own work. We ran a technical SEO audit that found crawl issues on 12,000 orphaned product URLs. You won't find that fact in anyone else's article, because it came from a real audit. One sentence like that does more for a page's effort profile than three paragraphs explaining in general terms what orphaned pages are.
The e-commerce version of thin content
Thin content hits stores hard because stores generate pages automatically. Every product, variant, filter combination, and category creates a URL. Few of them get the attention a blog post gets.
We've inherited a lot of Magento stores over the years, and the thin-content patterns repeat:
- Boilerplate product copy pasted across hundreds of SKUs with only the model number changed.
- Category pages with a product grid and nothing else, or a block of keyword-stuffed text at the bottom that nobody reads.
- Filter and sort URLs that get indexed and produce thousands of near-duplicate pages.
- Orphaned products that no internal link reaches, which leaves the index full of pages Google can't place in context.
Product pages don't need essays. They need effort that fits the format. That means real photography, accurate specs, sizing or compatibility details your customers actually ask about, answers pulled from your support inbox, and reviews. A category page earns its place with buying guidance that helps someone pick between the products in the grid.
Technical cleanup and content depth are one project. If filter URLs are bloating the index, better copy won't save you. If the crawl is clean but every product page says the same thing, the crawl fix won't save you either. That's why one team owning both the development and the SEO matters. The technical SEO audit guide covers the crawl side in more detail.
How to audit your site's content depth now
This is the process we run. You can do a version of it in-house with a crawler, Google Search Console, and a spreadsheet.
1. Build the complete URL inventory
Pull URLs from four sources: a full site crawl, your XML sitemaps, Search Console's indexed pages report, and your analytics landing pages. Then compare the lists. URLs in the index but missing from your crawl are orphans. URLs in your sitemap that aren't indexed tell you Google already passed on them. The gaps between these lists are often where the thinnest content hides.
2. Attach performance data
For every URL, pull clicks and impressions from Search Console over its full 16-month retention window, plus sessions and conversions from analytics. A longer window keeps seasonal pages from being flagged as dead. Sort by clicks. You'll see a small set of pages carrying the site and a long tail doing nothing.
3. Score each page on effort
Use a simple rubric. Score each question 0, 1, or 2:
- Does this page contain information that doesn't exist on the current top-ranking pages?
- Could a competent competitor reproduce it in under an hour?
- Does it include first-hand evidence such as photos, data, screenshots, or examples from real work?
- Does it answer the follow-up questions a reader would have, or only the headline query?
- Is there a clear, credible source behind the claims?
Reverse-score question 2. A page that's easy to reproduce gets a 0. Anything scoring 4 or below out of 10 is a thin-content candidate. This is our rubric, not Google's, but it maps directly to the effort, originality, and skill criteria in the rater guidelines.
4. Find templated duplication
Run a near-duplicate check across product, category, and location pages. Most enterprise crawlers flag pages above a similarity threshold. Clusters of pages that differ only in a city name or model number are the scaled-content pattern Google's spam policy describes. Treat them as one problem with one fix.
5. Map pages to intent and catch cannibalization
Group URLs by the query they target. When three blog posts compete for the same question, none of them can be the most complete answer. Search Console shows this directly: several URLs trading impressions on the same query.
6. Check site focus
List your content by topic. If a meaningful share of your pages sits outside what your business actually sells or knows, that drift can weaken how Google reads the whole site. An e-commerce store publishing general lifestyle filler to chase traffic is the common version of this.
What to do with pages that fail
Every thin page gets exactly one of four decisions:
- Rebuild. The page targets a query worth winning, and you have material that can make it the best answer. Add the original data, the photos, the decision table, and the specifics. Rewriting the same information in fresher language doesn't count as rebuilding.
- Merge. Several weak pages cover the same intent. Combine them into one strong page and 301-redirect the rest to it.
- Noindex. The page serves users but has no search value. Examples include filtered views, internal utility pages, and thin tag archives. Keep it for visitors and pull it out of the index.
- Remove. The page serves nobody. Return a 410, or a 301 if a closely related page exists and the old URL has links pointing at it.
Resist the urge to rebuild everything. Pruning is often the higher-return move because it concentrates your site's effort profile on the pages that matter. A site with 400 strong pages sends a clearer signal than one with 400 strong pages buried among 3,000 weak ones.
Do the work in order of value. Start with pages that already earn impressions but rank on page two or three. They're closest to paying off, and they show you whether your rebuild approach moves rankings before you apply it across the site.
Where AI fits in a high-effort workflow
None of this means you should stop using AI. It means AI can't supply the part that counts. A model can structure a draft, check coverage against the questions searchers ask, clean up transcripts of your subject-matter experts, and speed up formatting. It can't supply your test results, your photos, your customer questions, or your judgment about which option to recommend.
The workflow that holds up puts the original material first and uses AI to shape it. That means interviewing the person who knows, pulling the data, and taking the photos before anyone drafts a word. Then a human editor checks every claim against the source material. We wrote about building that kind of review into AI output in why your AI should disagree with itself. The principle carries over directly. The model's first pass is where the work starts.
If your team doesn't have the bandwidth to gather original material at the pace you publish, publish less. Twelve pages that took real effort will outperform a hundred pages that didn't. The rater guidelines explain why, and the leaked attribute shows Google built a way to measure it.
Where we could be wrong
The leak shows that contentEffort exists and what it was designed to estimate. It doesn't show how much weight the signal carries, whether it applies beyond article pages, or whether the current system still uses it in the form documented. Google's warning about incomplete context applies here.
What would change our advice? If Google published evidence that effort estimation plays no role in ranking, you'd lose one reason for this audit. You'd keep the others. The rater guidelines, the scaled content abuse policy, and the helpful content signals now inside core ranking all push in the same direction. A page with original information, first-hand evidence, and a clear recommendation also converts better, because it answers the buyer's actual question. We'd run this audit even if the leak had never happened.
Get an audit of your content depth
We run content depth audits alongside technical SEO because the two problems show up on the same URLs. Our process is laid out in Metrix SEO explained. One team handles the inventory, the effort scoring, the prune-and-merge plan, and the Magento or platform changes that make the fixes stick. When the audit turns up pages worth rebuilding, our content SEO services turn them into pages that carry original material and earn their rankings.
Clients stay with us for years because the work moves rankings, revenue, and conversion rate. If you suspect a thin tail is dragging your site down, get an audit and we'll show you which pages to rebuild, merge, and cut first.