If you've tried AI on a real client, you were probably impressed. A writer pastes in the client's service page and a keyword list, asks for a blog post, and gets a draft that's surprisingly good. That first impression is accurate. On one client, today's AI really is that good.
The problem shows up later, and at scale. Agencies that go further start running everything through one chat tool. Writers save client notes in its memory, upload more briefs and style guides, and use it for blog posts, meta titles, schema and ad copy across the whole client list. (Putting client logins and data into a chat tool is its own problem. More on that below.)
Once it's working across dozens of retainer clients, the drafts start to drift:
- A blog post for an immigration-law client drifts into divorce-law topics, because another law client's material is sitting in the same pile.
- One client's banned words and house style turn up in another client's posts.
- A meta title names another client's city or service area.
- LocalBusiness schema goes out with the wrong client's phone number or address.
- An ad headline promises a discount that a different client is running.
Account managers start re-reading every line, and the time savings are gone. The AI didn't get worse. What changed is how much the agency is feeding it.
Why does AI slip once it's used across the whole client list?
An AI model doesn't learn from your agency's use. Vendors release new versions, but the model answering you isn't studying your clients between conversations. When it seems to know something new, like a client's tone or this month's keyword list, that's because the information was handed to it in the moment.
That handed-over material is called context: the briefs, style notes, page copy and history the AI reads before it answers. Context is the only part the agency controls, and it's the part most tools manage worst.
Why does AI mix up clients?
Chat tools remember by piling things up. Every instruction, upload and correction goes into one growing pile, and the AI reads from it every time someone asks for something.
The research is consistent. Language models are most likely to miss details that sit in the middle of a long input,1 and they get less reliable as they're given more to read, even on simple tasks.2
An agency's pile is full of look-alikes: several law firms, several dentists, several roofers, each with near-identical service pages, overlapping keywords and style guides that differ by a few rules. The more alike the material, the easier it is for the AI to grab the wrong piece. Those studies measure long inputs, not agencies, but one client's details showing up in another client's work is what that weakness looks like in one.
Will a newer AI model fix it?
It's tempting to wait for the next release. But today's models already write well from a complete, correct brief. What they don't have is your agency's knowledge: how a post moves from keyword approval to draft to client sign-off, what each client has said yes and no to, and which checks your team runs before anything is published.
An AI model knows how an average agency works. It knows how yours works only from what you give it, and it can use that only if the right piece reaches it at the right moment. A newer model pointed at the same pile makes the same mix-ups.
Does Google penalize AI-written content?
Not for being AI-written. Google's published guidance says it rewards helpful, high-quality content however it's produced. What breaks its spam policies is using automation, AI included, mainly to manipulate rankings.3
That puts the weight back on accuracy. A post with the wrong client's practice area or the wrong city isn't helpful to anyone, whoever wrote it. Keeping each client's information separate is what keeps AI-assisted work useful.
What about client data?
Agencies hold a lot that isn't theirs: client website logins, Search Console and Analytics access, ad accounts, and contracts or NDAs that say what can be shared and with whom. Client passwords never go into a chat tool. For client content and reports, business-tier AI accounts in the agency's own name, with training on your data excluded, are the right starting point. Someone still has to check those terms against what your client agreements allow, and set the rules for what staff paste in.
How do you keep AI from mixing up clients?
The answer is older than the internet: put things where they belong. When an agency's knowledge lives in organized folders, each with a short written note saying what's inside and how to use it, the AI reads only what the current task needs. A blog post for one client is given that client's brief and that client's rules, not the rest of the client list.
A 2026 research paper gave this approach a name and a method. I explain it in plain English in the next post.
What to take from this
- If your AI drafts got worse as you added clients, the cause is the pile of information, not the model.
- Keep clients, their rules and their approved keywords separate from day one. Cleaning up a messy setup later means rebuilding it.
- Ask any AI vendor one question: once we've run fifty clients through this, how does it keep them apart?
- However good the drafts get, a person approves every one before it goes to the client or goes live.
Common questions
Why does ChatGPT mix up my agency's clients?
Chat tools keep everything you give them in one growing pile of memory and uploads. When many clients look alike, the AI can pull a detail from the wrong one. The fix is organizing the agency's information so each task reads only that client's files.
Does Google penalize AI-written content?
Google's guidance says it rewards helpful content however it's produced. Using automation mainly to manipulate rankings is against its spam policies. What matters is whether the page is useful and accurate.
Is it safe to put client logins and data into an AI chat tool?
Never paste client passwords into a chat tool. For content and reports, use business-tier accounts in the agency's name with training excluded, checked against your client contracts and NDAs.
Will a newer AI model fix drafts that got worse?
Usually not. The problem is the information the model is given. A bigger model reading the same disorganized pile makes the same kind of mistakes.
Sources
- Liu, N. F., et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. arxiv.org/abs/2307.03172
- Hong, K., Troynikov, A., & Huber, J. (2025). Context Rot: How Increasing Input Tokens Impacts LLM Performance. Chroma. trychroma.com/research/context-rot
- Google Search Central Blog. (2023, February 8). Google Search's guidance about AI-generated content. developers.google.com