Prompt Rot: Why AI Prompts Drift and How to Maintain a Prompt Library

A prompt that produced a clean depreciation rollforward in March can produce something noticeably different in September, with no change to the prompt. Nothing in the wording broke. The things around the prompt moved.
I call this prompt rot. It matters more as finance offices move from trying AI occasionally to running recurring work through it: the monthly reconciliation, the variance narrative, the capital additions review. A prompt used once is a prompt. A prompt used every month is a process, and processes need maintenance.
Why Prompts Drift
Three things change underneath a prompt that has not been edited.
- The model changes. AI providers update their models regularly. The same wording can be interpreted differently after an update. A prompt that said "summarize the variances" may now return more detail, less detail, or a different ordering of items.
- The inputs change. A new chart of accounts, a revised tariff, a new FERC order, a different export layout from the billing or work order system. The prompt still describes the old structure, and the model works with what it is given.
- Your requirements change. The capitalization threshold is revised, the board packet format is updated, or the auditors ask for a different level of support. The prompt still reflects last year's version of the task.
The third cause is the one most often missed, because the prompt looks the same as it always did. It is the accounting equivalent of a schedule that still carries last year's rate.
An Illustrative Example
Suppose a prompt reviews capital work orders and flags any item under the capitalization threshold for possible expensing. When the prompt was written, the threshold was $500. The utility's capitalization policy was later revised to $1,000, and nobody updated the prompt.
The prompt keeps running. It still produces a clean table with the same columns. But it now misses every item between $500 and $1,000 that should be flagged under the current policy. Nothing looks wrong, so nobody checks.
Layout changes work the same way. If the billing system export renames a column from "Rate Sched" to "Rate Schedule Code," the model may still infer what the column means. Or it may not, and it may quietly skip it. The output looks complete either way.
Treat Prompts Like Active Accounting Schedules
The fix is borrowed from how we already manage schedules that carry forward period to period. A lead schedule has a preparer, a date, a reviewer, and a tie-out to the general ledger. A prompt that supports recurring work should have the same basic controls.
| Field | What to record |
|---|---|
| Name and purpose | What the prompt does and which close or filing task it supports |
| Owner | One person responsible for keeping it current |
| Version and date | When it was last changed, and what the current version is |
| Last tested | Date of the last known-answer test, and the result |
| Model and tool used | Which model produced the tested results |
| Input source and layout | The report or export the prompt expects, and its column structure |
| Parameters | Thresholds, account ranges, and policy references written into the prompt |
| Change log | What changed, when, and why |
The parameters row is the one that prevents the capitalization threshold problem. If a dollar amount, an account range, or a policy reference appears in the prompt, it should also appear on the record, so that a policy change sends someone to the prompts that depend on it.
The Known-Answer Test
A review of the prompt text alone will not find drift. The reliable check is to run the prompt against an input where you already know the right answer.
- Pick a closed period. Use a month or quarter that has been reconciled, reviewed, and signed off. The result is already established.
- Save the input and the correct output. Keep the export that went into the prompt and the final reviewed result.
- Rerun the prompt on that input. Compare what it returns to the signed-off answer.
- Record the result. Note the date, the model used, and any differences.
As an illustration, suppose the closed-period reconciliation had 94 line items, six of which were true exceptions after review. If the rerun flags 11 items, or flags only four, the prompt has drifted. You have found that out on a closed period, before it touched current-period work.
When to Rerun the Test
- At each quarter-end close, as a routine check
- When the AI provider announces a model update or the tool changes its default model
- When the source system or export layout changes, including new ERP releases and chart of accounts changes
- When a capitalization policy, threshold, tariff, or reporting requirement changes
- Any time the output looks different from what you expect, even if you cannot say why
Most of these triggers are events the finance office already knows about. Add the prompt library to the checklist you already use for those events.
Keep the Library Small and Owned
The library does not need to be large. A working library for a utility or co-op finance office is often ten to fifteen prompts: the recurring tasks that actually run every month or quarter. Prompts used once do not need this level of control.
Every prompt needs one named owner. A shared prompt that nobody maintains has the same problem as a shared spreadsheet that nobody maintains. When something changes, no one is sure whose job it is to update it.
Accountability Does Not Change
Testing a prompt does not replace reviewing its output. The known-answer test tells you whether the prompt still behaves the way it did when you last validated it. It does not tell you whether this month's output is correct. That still requires a reviewer who knows the process and signs off on the result, the same as it would for work prepared by a staff accountant.
Bottom Line
Prompts that support recurring finance work age the same way schedules do. The model changes, the inputs change, and the requirements change. Record a version, a date, an owner, and the parameters for each prompt. Keep a known-answer test for it. Rerun the test at each quarter-end and whenever one of the underlying pieces changes. That takes little time, and it finds the problem on a closed period, before it reaches a filing or a board packet.
Related Articles

