How a manual, error-prone negative-keyword process for a six-campaign Performance Max account became a documented, rule-based, self-testing classification system, and what the audit work along the way turned up.
Waxit Car Care sells polishing machines, polishing pads and compounds, pressure washer accessories, and ceramic coating products direct-to-consumer through Shopify. Google Ads runs across six live campaigns, each scoped to a distinct part of the product catalog rather than one broad account, which is precisely what makes search-term hygiene hard to get right at scale.
| 01 | Polishing PMax | Machine polishers, polishing pads, compounds & polish |
| 02 | Pressure Washer Accessories PMax | Hoses, nozzles, guns, lances, adaptors, reels |
| 03 | Core PMax | Prospecting, non-segmented categories |
| 04 | New Customer Core PMax | New-customer acquisition, existing customers excluded |
| 05 | Ceramic Coating PMax | Premium ceramic coating segment |
| 06 | Search, DSA | Dynamic Search Ads, under evaluation for AI Max |
Campaigns owned and independently audited within this engagement, alongside the OnaVibe (sister brand) Friday Launch PMax campaign.
Every owned campaign was audited on a standing weekly cycle rather than reviewed only when a number looked wrong: changelog, performance, budget and bidding, product feed, audience signals, creative, search themes, and negative keywords, in that order, every time. Every recommendation required the business owner's sign-off before anything went live in the account.
Six campaigns sharing one broad product category is a hard boundary to hold by hand, and it was being held by hand: no documented rule set, no record of why a given term was let through or blocked, and no way to check later whether a call was still correct.
Each campaign is scoped to a narrow slice of the catalog: Polishing PMax should never spend on a pressure washer search, and Pressure Washer Accessories should never spend on a pump or machine, since those sit on their own campaign entirely. The boundary between "on-topic" and "in scope for this specific campaign" is easy to get wrong in either direction, and getting it wrong either way has a cost. Too loose, and budget leaks to searches that were never going to convert on that page. Too strict, and the campaign starts blocking its own best-performing, most obvious category terms.
Two findings made the scale of the gap concrete early on: a 180-day pull showed Pressure Washer Accessories PMax was running with zero active negative keywords, and a single unfiltered local-service term cluster, "car wash near me" and its variants, had quietly spent $1,014.62 with zero conversions to show for it. There was no system catching that; there was just whatever anyone happened to notice.
Rather than hand-triaging search terms campaign by campaign, the fix was a documented decision framework applied consistently across 26,640+ search terms spanning both major PMax campaigns over a 180-day window, built so the same call gets made the same way every time, and so every call can be checked later.
Every search term is kept by default. It only becomes a negative if it falls into one of four categories: it names an exclusive competing brand, it names a competitor Waxit doesn't stock, it shows no buying intent (research, comparison, question-format queries), or it's genuinely unrelated to that specific campaign's slice of the catalog, even if it's car-care-adjacent and would belong on a sibling campaign instead.
The single biggest structural change in the framework. The original logic kept a term unless it matched a known competitor, which is unwinnable, since no competitor list is ever complete. The rebuilt logic inverts that: a term is negatived by default unless every token in it positively resolves to real Waxit-vendor vocabulary or genuinely generic category vocabulary, checked against a ground-truth product catalog rather than a keyword blocklist. Positive confirmation required, not absence of a bad signal.
A rule set is only as good as its edge cases. Two examples of how the framework's logic actually resolves a genuinely ambiguous term: one where the old approach was too aggressive, one where it wasn't aggressive enough.
Raw conversion count is a bad filter on its own; a low-volume term can still be a genuine winner. The term "rupes" looked like a marginal performer at a glance, at 4.86 conversions. Measured against the campaign's own benchmark instead, it was beating it on every axis that mattered. It stayed. The framework's standing rule since: judge a term against its own campaign's benchmark, never against raw volume alone.
Every owned campaign was audited on a standing structure: changelog, performance, budget & bidding, product feed, audience signals, creative, search themes, negative keywords, so nothing gets reviewed only when something looks wrong. Three findings from that cadence below.
Account-wide ROAS fell sharply in July, from a 14.25x EOFY peak in late June to 6.37x for the month that followed, the kind of drop that invites a scramble for an account-level explanation. Splitting the number apart first: new-customer acquisition held steady, new customers actually grew as a share of purchases, and the fall was concentrated almost entirely in returning-customer revenue immediately after the EOFY sale ended. Read together, that's a demand-timing event, not an account failure: the right response was to set July as the honest baseline and measure forward from there, not to make reactive changes chasing a number that was never really about execution.
Two calls made early in the engagement were reversed once fresh data came in, on purpose. A bulk negative-keyword recommendation covering DeWalt, Ryobi and Ozito was walked back after a 180-day pull showed those terms actually converting at 40.74x, 14.51x and 18.82x ROAS respectively. A separate claim that pressure washer machines "weren't selling" was retracted once it was confirmed that demand was being correctly captured by a different, dedicated campaign at 17.69x ROAS. The discipline that matters here isn't never being wrong. It's checking against source data before a recommendation ships, and reversing in writing when the data disagrees.
Related services: Google Ads for e-commerce and Google Ads management.
Figures in this case study are drawn directly from working audit files, the search-term classification workbook, and its accompanying regression-test suite and review ledger maintained throughout the engagement. Client and account details are shared with the understanding that this work is being presented as a professional portfolio reference.