A contemporary office setup featuring computers on desks with modern flat screens.

Photo by Vladimir Srajber on Pexels.

Every AI-generated reply, suggestion and document SimpleWP produces is supposed to come from one place: a single, versioned SimpleWP prompt gateway, rather than prompt wording scattered across the codebase. This is the story of how that gateway went from technically correct to actually complete — the one caller it was never built to serve, the audit that counted what was still outside it, and the missing transport feature that turned out to be the real fix. If you’ve ever described a site to SimpleWP’s chat assistant, every word of that conversation now comes from the same place this story is about.

A Prompt Gateway Meant to Be the One Source of Truth

SimpleWP built a centralized prompt-management service early on, meant to be the one place every AI-generated reply, suggestion or document in the product comes from, rather than having prompt wording scattered across dozens of files. A single source of truth for prompt text sounds like an easy thing to build and a much harder thing to actually finish moving everyone onto. The value of a gateway like this isn’t just tidiness: a versioned prompt can be reviewed, rolled back and reasoned about the same way any other piece of shipped code can, which is much harder to do once the same wording is copy-pasted into three different files under three different names.

Making the Prompt Gateway Slug Mean Something

An earlier pass of work made the gateway’s “slug” identifier — a specific, versioned prompt referenced by name — actually mean something for the first time. Before that pass, the integration existed in the code but had never been wired to run anything at all; the plumbing was there, but nothing flowed through it yet. That same earlier pass explicitly excluded the product’s live, streaming chat assistant from using the gateway at all, because the gateway did not yet support streaming responses over server-sent events. That exclusion was deliberate, scoped and documented — not an oversight, but a conscious decision to ship the gateway for everything it could already serve and come back for the one caller it couldn’t.

What the Exclusion Actually Grew: a Second Prompt Store

Coming back is exactly what a later audit set out to do. That audit found the streaming exclusion had let a second, independent copy of prompt text grow up specifically for the chat assistant, stored locally rather than in the gateway, and that local copy no longer matched the gateway’s own version of similar prompts. The same audit found 17 separate places in the codebase still feeding prompt text directly to an AI model without going through the gateway at all, including four complete prompts that had never been given a database record anywhere — not even in the old, pre-gateway pattern. Seventeen places and a quietly growing second store is less a single gap than a measured starting point: the audit’s real value was turning “the gateway isn’t finished” into a concrete, countable list of exactly what finishing it would mean.

Turning the Prompt Gateway Into a Server-Sent Events Gateway

The fix the team chose was not to special-case the chat assistant forever, carrying two prompt stores side by side indefinitely. It was to add the missing streaming capability to the gateway itself, so the chat assistant could finally be migrated onto the same single source of truth as everything else. Migrating the chat assistant onto the gateway came with one real challenge: its running conversation summary, previously sent to the AI model as one kind of structured message, had to travel as a plain named value instead, because the gateway’s streaming endpoint rejects that particular kind of caller-supplied message outright. The team caught and resolved that format mismatch in its own development environment, as part of the same piece of work, before the cutover ever reached production — covering 2,161 already-summarized sessions in that environment, which is the scale of what the fix had to account for, not a measure of anything going wrong in front of a real user.

Proving Nothing Changed What the Model Reads

Before cutting the live chat assistant over, the team needed proof the new approach wouldn’t silently change what the AI model actually sees — a question of system-prompt versioning as much as of plumbing. They captured the assistant’s exact current system-prompt output byte-for-byte: 20,626 characters for one mode, 11,331 for another. The newly centrally-rendered version was then checked against those captures character for character, in both modes, and matched exactly. That is not a small bar to clear; it means the migration changed where the text came from without changing one character of what it said. A gateway that has proved that kind of parity once is also a gateway that can be trusted with the next version bump, and the one after that, without re-running the same manual comparison by hand every time.

Moving Every Prompt, Including Tool Descriptions

As part of the same consolidation, the team also moved 280 AI tool descriptions and 399 tool-parameter descriptions — the text an AI model reads to decide which action to take and how to call it — out of hardcoded source-code text and into the same central, versioned gateway. The reasoning was simple once stated plainly: a tool’s description is itself prompt text the model reads, not ordinary code, so it belongs wherever the rest of the product’s prompt text lives. Treating hundreds of short tool and parameter descriptions as first-class prompt content, worth the same versioning and review as a user-facing reply, closed a gap that is easy to miss precisely because none of that text was ever meant for a human reader in the first place.

Finishing the Consolidation: Documentation and Cost Reporting

A gateway is only a true single source of truth once everything that reads from the old pattern has been found and pointed at the new one, and that finishing work turned out to have several parts. A small, dedicated automated check was built specifically to catch any future hardcoded prompt text before it could ship again; its first version needed correcting after it was found mis-attributing which line of code actually held the problem it was flagging, and after a length threshold let two genuinely full-sized prompts slip through undetected. Three separate pieces of internal documentation still told engineers to edit prompts in the old, local database table or through a schema-migration change, instructions that would have changed nothing the AI model actually reads anymore; those instructions were corrected, with the old ones kept only as an explicit warning rather than deleted outright, since the old tables still exist and still return a plausible-looking answer. The migration also showed that the platform’s own internal cost-tracking had been reading from the wrong ledger since the chat assistant moved to the gateway — a local usage-log table that had stopped receiving any new rows, so a sum over zero rows was quietly reporting as a real, if wrong, cost figure of zero instead of flagging an error. That was fixed by reading the figure from the gateway’s own ledger instead, which had been recording real spend the entire time. A second parity-checking tool, meant to catch any prompt text growing apart from the product’s actual tool descriptions, had itself been comparing against stale, days-old prompt text the whole time, because it read from the same old local table rather than the gateway — so it had been reporting a clean pass against text no AI model was actually being served.

Closing the Loop on the Prompt Gateway

Finishing the migration meant deleting the chat assistant’s own prompt-assembly function entirely — the one piece of code that used to compose its system prompt from a mix of hardcoded text and database lookups. After this work, the gateway is the only thing in the product that renders AI system-prompt text, for every caller, including the one it was built last for. It’s a companion story to SimpleWP’s own postgres database migration and to the team’s static egress IP relay: a piece of infrastructure that looked finished at the application layer still had loose ends worth tracking down, and tracking them down is what actually finishes the job. Describe your next site to SimpleWP’s chat assistant and every prompt behind the reply comes from the same gateway this story is about.