A marketing agency’s whole client portfolio kept getting infected, and each time the host cleaned a site, the malware came back within days. The sites weren’t the problem…the keys were. Here’s how I traced a thirty-site compromise back to a single never-revoked account, cut off the access before touching a file, and built the monitoring that’s kept it clean since.
A Cleanup That Kept Not Working
The message came through Upwork. The owner of a small digital marketing agency builds and looks after websites for a portfolio of professional-services firms, around thirty of them, all hosted on a single WP Engine account. For two months his sites had been getting infected with a fake Cloudflare “verify you are human” overlay, the kind that asks visitors to press a key combination and secretly pastes a command into their clipboard. Not only that but Google was showing gambling pages under his clients’ domains and strangers kept appearing as verified owners in Search Console.
He’d done what you’re supposed to do; he reported each infection to the host, and the host’s malware team had scanned and cleaned the affected site. Their backup history for one site literally alternated between “Cleaned Site” and “Possible Infected Version”, back and forth for hours. He’d had the same argument with support several times: they said the malware was on his computer, not the site, and he said the sites were serving it to visitors right now. Both of them were partly right and neither had the full picture.
By the time he wrote to me he was angry, tired, and past the point of wanting another clean. He wanted someone to find out why it wasn’t holding.
Thirty Sites Is a Different Problem
A single infected WordPress site is a bounded problem. You take an image, find the payload, find how it got in, close the door, verify, monitor. I’ve written up several of those here.
Thirty sites on one hosting account is a different shape of problem, and the difference matters more than the number. On one site the question is “what’s on this box” but across an estate it’s “what do all these boxes have in common,” because whatever that is, it’s the thing that lets the attacker come back after each clean. If you treat thirty sites as thirty separate cleanups you’ll be doing the thirty-first next week.
So before I cleaned anything, I needed the whole picture. I asked for three things: SSH access to every environment on the account, an export of the WP Engine portal’s activity log, and the Google Workspace audit logs for the agency’s domain, and then I spent the first day reading rather than removing.
Reading the Host’s Own Logs
WP Engine keeps an activity log for the customer portal: every login, every backup, every cache purge, every SFTP user created, with the IP it came from. Most people never look at it but here it was the single most useful artifact in the whole engagement.
The agency’s legitimate team worked from the United States and India. The log showed a portal user provisioned in mid-March whose 550 logged actions came almost entirely from residential IP space in Pakistan. It had touched the large majority of the installs on the account: backups, cache clears, SFTP provisioning, environment copies, but…nobody on the team worked from Pakistan.
Here’s the part that reframed everything.: that portal user wasn’t forged or planted but was rather created by… the agency owner himself…from his own home IP as a shared login for a pair of offshore contractors he’d hired that month. Their engagement ended in mid-May but their access didn’t. Nobody removed the account and the sites kept getting infected for weeks after the contractors were gone.
Whether the contractors were themselves malicious or their own machines were infected with an info-stealer and someone else was riding their sessions, the evidence couldn’t say, and I’ve been careful never to assert it. For containment though it doesn’t really matter because either way, a platform-wide account with reach into every site had been live and in hostile hands for two months.
So that answered the first question but it didn’t answer the second which was why the host’s cleans kept failing even on sites where the portal account had never done anything obvious.
The Second Key
Every WordPress install on the account had audit logging of some sort: Stream on some, Wordfence on others, an audit-log plugin on a few, etc. Reading across them a pattern showed up on a single evening in late May. On site after site within the same few hours, the same staff administrator account logged in from a VPN exit node and installed things: a fake “Classic Editor” plugin, a must-use loader file, a per-site redirector, a hidden admin user with a name chosen to look like the site.
That staff member was real and her home ISP was a small regional provider in the southern US but the logins that night weren’t from there. Her saved WordPress password had been stolen (most likely by the same kind of clipboard-hijack lure the sites were now serving to their own visitors) and it was being replayed from rotating datacenter IPs. WordPress logins don’t have a second factor unless you add one so a stolen password works from anywhere.
There was a subtler tell too. On several sites the logs showed authenticated admin actions, theme-editor saves and REST calls, with no corresponding login event at all. That’s what session-cookie replay looks like. The attacker wasn’t typing a password; they had a live authenticated session lifted from a browser.
So there were two keys: a platform-wide portal account and a staff admin’s password and sessions. Neither was a WordPress vulnerability or a WP Engine server flaw. The host was right that a device was involved, but the customer was right that the sites were genuinely infected. And…the host’s remediation had never touched either key which is why it kept failing.
Two keys, thirty sites
Both access paths survived the cleanups
Key 01
Hosting portal account
Former contractor access, never revoked
Hosting access
- Provision SFTP users
- Access site environments
- Copy environments
Key 02
WordPress administrator
Stolen password and sessions, still usable after file cleanup
WordPress access
- Install malicious plugins
- Modify themes and settings
- Create hidden administrators
Shared WordPress portfolio
Hosting portal WordPress admin
Around thirty WordPress sites
- Site cleaned
- Surviving access reused
- Malware returns
Four Layers the Scanner Couldn’t See
Once I knew how the attacker was getting in the infections themselves made sense. The visible payload (the fake-Cloudflare overlay) was only the top layer; underneath it on most infected sites were four cooperating layers of persistence and a signature-based file scan addresses only one of them.
The first was the malware kit itself in several strains.: a self-hiding plugin that filtered itself out of the plugin list and its user out of the users list; a footer injector that opened a WebSocket to an attacker-controlled domain and pulled its payload live so there was nothing static on disk to match. On some sites the same injector lived in the database instead: in a header-and-footer plugin’s option, in an Elementor custom-code snippet, in a child theme, etc, and file scans read those sites as clean while they were actively serving malware.
The second was SEO abuse. Turkish gambling doorways on some sites, an Indonesian lottery doorway on another, all being served only to Googlebot and mobile user agents. One site had a cloak in its theme’s functions file that survived several of the host’s cleaning passes (and I’ll admit my own first signature-based pass too). Search Console showed hundreds of thousands of gambling clicks attributed to his clients’ domains.
The third was standing re-entry tooling: web file managers installed as plugins on a dozen sites, a couple of independent web shells a must-use plugin that fetched and ran code from a remote domain on request, web-reachable copies of Adminer, a database console with no logging at all. And on eleven installs there were SFTP users the attacker had created through the portal which no WordPress-level clean would ever see.
The fourth was ownership. The attacker had verified themselves as owners of many of the clients’ domains in Google Search Console using a rotating family of throwaway Gmail accounts. Some verifications were hardcoded meta tags in theme headers; others were DNS TXT records at the registrar. Removing malware doesn’t revoke Search Console ownership so as long as those grants stood, the attacker could see every property’s search data and manipulate its indexing.
There was also a fifth thing I only found on the second pass. A third-party SEO tool had minted WordPress Application Passwords across seven environments, one of them belonging to the compromised staff admin on a production site. Application Passwords are separate REST credentials which survive password resets and rotating secret keys, and nothing short of enumerating and revoking them removes them.
And the backups were poisoned…every restore point WP Engine had taken during the six-week dwell window contained the implant. “Restore from backup” (which the host had offered as a recovery path) would reinfect the site.
Cut the Access First
The order of operations was the whole method. If you clean files while the keys are still in the attacker’s hands then you’re just providing them a tidier site to reinfect, and so the first day of hands-on work touched almost no malware.
Across the whole account · one coordinated window
Revoke shared access
- Hosting & SFTPRemove the contractor portal account and attacker-created SFTP users from every install.
- Trusted credentialsReset the owner’s credentials and compromised staff admin’s password from a trusted machine.
- Google WorkspaceReset sign-in cookies, revoke OAuth tokens, and enforce two-step verification.
- Persistent grantsRevoke stray Application Passwords; remove rogue Search Console owners and their DNS verification records.
Cut off shared access before cleaning files
Repeat for each site · steps 1–6
Clean in a fixed order
- Remove re-entry toolsFile managers, web shells, and Adminer.
- Rotate WordPress keys & saltsInvalidate replayed authentication cookies.
- Reset compromised administratorsReset their credentials and destroy all sessions.
- Remove the payloadsThe malware kit, database injectors, and gambling doorways.
- Restore modified filesUse the exact stock versions of core and theme files.
- Reduce the admin rosterKeep only the one to three people who need access.
Check what the live site actually serves
Repeat for each site · steps 7–8
Verify the live response
- Purge both cachesClear the host’s page cache and object cache.
- Fetch three ways & compareUse cache-busted requests and diff the responses.
- Browser
- Googlebot user agent
- Phone
A Googlebot user agent can reveal user-agent cloaking. It does not reproduce a request from Google’s crawler network—the limitation uncovered later in this investigation.
Step eight is the one people skip. WP Engine fronts every site with a Varnish cache, and a cloak that only serves to crawlers is invisible if you check with a browser and don’t purge first. I fetched every site three ways (cache-busted) before calling it clean.
I worked through the estate one site at a time writing a closeout for each. Twenty-seven production installs in fourteen days…thirteen of them had live malware or cloaks, three had no serving payload but still carried the attacker’s admin user or database residue, and the remaining ten were what I’d call hardening-class: no active infection, but the same open doors as the rest, and on two of them a full site backup sitting in the web root, publicly downloadable, database included.
The Reservoir
A managed host like WP Engine gives every site a staging and a development copy. The agency had around forty of them but nobody had cleaned those because nobody had been infected on those.
But…an environment copy inherits everything: the rogue admins, the backdoors, the SFTP users. There’s no creation event for a backdoor that was copied rather than planted so the audit logs are silent and a staging-to-production push would carry the whole implant back onto a site I’d just cleaned.
I swept all of them, read-only, in a single pass. Most were clean of active malware; a handful still had the contractor account as an administrator.
Two Wrong Turns
I got two things wrong during the engagement, and both taught me something I now do by default.
The first was the staff admin whose password had been stolen. Early on, seeing her account log in from an unfamiliar IP range, I told the client her Google account was compromised as well and had it locked, but…it wasn’t compromised as the range turned out to be her home ISP. Her WordPress password had been stolen; her Google account was fine.
The second was the cloak that survived my first pass. On the site the host had cleaned most often, my initial sweep was signature-based, matching known kit filenames and known payload strings, and it read clean. Three days later the site was serving a Turkish gambling doorway to Googlebot again from a function in its theme that matched none of my signatures. I’d made the same mistake the host had: I’d looked for what I already knew.
That second mistake changed the methodology for the remaining sites and for the nineteen I’d already closed, all of which got a second pass. The new pass didn’t rely on signatures at all. Every core and plugin file was checksummed against the official release. Every theme was grepped for the behaviours a cloak needs: user-agent checks, reverse-DNS lookups, remote fetches, conditional output. Every database store that can inject HTML into a page was read in full. And every site’s Search Console data was pulled, for a reason the next section explains.
The Cloak You Can’t Fetch
One site in the second batch read clean on every check I had, but Search Console said otherwise. The property’s performance report showed more than six hundred thousand clicks over the winter for a casino brand name, landing on URLs the site had never had. The spike had ended in February sowhatever served it had been removed by the host or by an earlier clean before I arrived but nothing I’d run would have found it while it was live, and that bothered me enough to work out why.
The answer was in the host’s logs from the spike window. The doorway hadn’t checked the user agent at all; it had checked the reverse DNS of the requesting IP and served the gambling pages only when the hostname resolved to Google’s crawler network. My Googlebot fetch came from my own IP with a spoofed user agent so to that cloak I was a browser and I got the clean page.
Can you see the spam?
Watch each visitor try the site, or choose one yourself.
1The visit
Visit request
Claims to be
A web browser
2The site checks
Is the postmark
from Google?
No
3What they see
The normal website
A normal visit comes from someone's home or office. The site can tell it isn't Google, so it shows the real website.
How I caught it anyway: Google's own reports showed casino searches and pages this site never had, even though the spam had already stopped by the time I looked.
You can’t fetch your way around a reverse-DNS gate. Google can because Google is Google. So Search Console became a standing part of the method for every site after that: the performance report for queries the client would never rank for, the Users and Permissions page for owners the client didn’t add, and the unused-verification list for tokens that would let a removed owner walk back in. On the same site, that pass also turned up two thousand poisoned rows in the oEmbed cache table, a spam injector from an older incident that no file scan would ever touch. Google’s view of a site is the one check an attacker can’t cloak against and it’s free.
It Held, and Then It Was Tested
After all was said and done, the estate was closed out and every production site had a tripwire on it: a scheduled check for the kit files, the rogue admins, new Application Passwords, and the cloak signatures, alerting on change. The attacker’s last observed activity was a series of attempts to re-validate the stolen staff password from fresh datacenter IPs but the password had been reset anf the attempts failed and stopped.
The client had two open items I couldn’t close for him: the infected staff devices needed reimaging and a few Search Console grants needed removing from accounts I didn’t have owner access to.
In early July one site came back…a new family this time: a fake-Cloudflare CAPTCHA inject that loaded its payload from an attacker domain and, cleverly, only ran on Windows so a crawler, a phone, or a Mac saw a clean page. It had been written straight into a plugin’s database option through a web SQL console the attacker dropped and then deleted, bypassing the WordPress audit trail entirely. I cleaned it in an afternoon, but the next morning the logs showed the attacker logging back in as the same staff member from Tor overnight after her password had been reset for the second time.
That’s not indicative of a password problem but rather a session lifted from a device that had never been cleaned, paired with the host’s convenience “log in to WordPress from the portal” feature. Same root cause now with a second data point behind it. The June closeout hadn’t failed…the thing it had said needed fixing just hadn’t been fixed. So I locked the account, cleaned the site again, and documented the chain for the client with dates.
Building Something That Watches
By late summer the client and I agreed the right arrangement wasn’t another cleanup but rather standing monitoring across the whole portfolio with a small allowance of hands-on time each month and a rule that anything he sent me got triaged for free before any money changed hands. He’s on that retainer now.
The per-site tripwires had a limit I could see clearly by then: they knew the kit’s signatures and nothing else. The July inject had walked straight past them, so I replaced them with something that doesn’t depend on knowing the payload in advance.
The monitor I built takes a structured snapshot of each site every six hours over SSH: the admin roster, the must-use plugins, every PHP file under uploads and in the web root, the database options and templates that can inject into a page, the served homepage as three different user agents, and a checksum of the theme files that run on every request. It keeps a baseline per site and alerts only when the finding set changes so a noisy site doesn’t bury a real one. A second script checks availability and TLS every fifteen minutes. A third watches DNS hourly because the attacker’s Search Console verifications had lived at the registrar and no site-level check would ever see them.
The rule I hold myself to with it is the one the whole engagement taught: never reset a baseline for a finding you haven’t explained. A monitor that can be told “that’s fine” without evidence is a monitor that will eventually be told “that’s fine” about a backdoor.
What the New Baseline Turned Up
Standing monitoring earns its keep in a way a cleanup can’t, and it did so within its first week.
While I was extending what the baseline recorded on each site (adding the full contents of theme files that run on login and the templates that render the footer) one site’s snapshot came back with a login hook that shouldn’t have been there: seventeen unobfuscated lines appended to the theme’s functions file. On every successful login it posted the username, the plaintext password, and the session details to a rented server in a foreign datacenter.
It had been there since mid-July and had gone in two minutes after a hidden administrator account was created through a file manager the attacker had reinstalled. Alongside it were a web shell with its timestamp backdated to 2023 and a dropper in the languages directory guarded by a query string that would recreate the hidden admin and re-append the exfiltrator if anyone removed them. The site’s footer template had been edited in late August to carry a dozen Swiss gambling affiliate links.
I cleaned the site in one coordinated pass, dropper first so it couldn’t undo the rest, then the shell, the hidden admin, the hook, and the footer. Then I rotated every administrator credential across the whole portfolio and destroyed every session and added the exfiltrator’s endpoint and the dropper’s signature to the monitor’s indicator list. The staging copy of the same site had the same implants and got the same treatment a day later, with the agency’s in-progress staging work preserved and every plugin verified byte-for-byte against the vendor’s release.
That find is the argument for monitoring in one paragraph. No one reported that site and nothing was visibly wrong with it; it was found because something was watching the parts of a site that a scanner doesn’t read and the person watching had a baseline to compare against.
What Was Actually Underneath
Strip away the malware families and the site count and the story is short.
A platform-wide hosting account was created for contractors and never revoked when they left. A staff member’s WordPress password and browser sessions were stolen from an infected device. Between them, those two keys reached every site on the account and no amount of file cleaning could change that because cleaning files doesn’t take keys back. The host’s remediation was signature-based file scanning, which is one layer of a five-layer problem, and its backups were snapshots of an infected estate. Everyone involved was partly right and everyone was looking at a different part of the elephant.
What eradicated it was cutting access first, then cleaning each site in a fixed order, then verifying with cache purged and three user agents, then sweeping the staging copies nobody thought about, then watching. What kept it clean was the watching, and the discipline of not calling anything benign until its code had been read.
If you look after many sites and they’ve been “cleaned” more than once, the question worth asking isn’t which scanner to try next but rather what all of those sites have in common that the scanner can’t see.
Stonegate Web Security remediates compromised WordPress estates without taking them offline: finding what every infected site has in common, cutting the access before touching the files, and proving the fix holds with monitoring that doesn’t need to know the next payload in advance. If you run a portfolio of sites and the infections keep coming back, that’s the pattern I specialize in.
Related Reading
-
Case Study: No Backdoor Required
A WordPress site kept serving ClickFix malware after professional cleanups because a stolen administrator login still worked. -
Case Study: The Keys Were Never Taken Back
Two WordPress sites were compromised through administrator access left over from the freelancer who built them. -
Case Study: The Site That Kept Reinfecting Itself
An aged-care provider's site kept reinfecting because database malware, backdoors, and a stolen admin credential survived earlier cleanups. -
WordPress Malware: The Complete Guide for Small Business Owners
How WordPress sites actually get hacked, what malware does once it's in, what cleanup means, and how to tell whether the advice you're getting is sound. Written for owners who run their own site and for owners working with a developer or agency.