Federal awards require three years of records. Your AI tool may delete them in eighteen months. Neither setting was a decision your organization made.
A program officer asks a reasonable question during a site visit. The outcomes narrative in last year's final report described a change in client retention. Where did that analysis come from?
Two years ago the answer was a spreadsheet, an email thread, and a staff member who remembered doing the work. Today, part of the answer is often a conversation with a chatbot: someone pasted program data into a tool, asked it to identify patterns, and wrote the narrative from what came back.
That conversation is either still retrievable or it is not. Most executives do not know which, and the answer was not determined by anyone at the organization. It was determined by a default setting in a product adopted for reasons that had nothing to do with recordkeeping.
This is not hypothetical. Writing in the Journal of Business and Technical Communication, Andrew Hillen described AI use at the seven-person community mediation nonprofit where he worked as a program director: staff had begun routinely using a chatbot to draft text, and there had been "no training, no formal discussions, and no rules regarding AI use." The same office was writing grant applications, standard operating procedures, and an employee handbook, some of it with AI assistance. That is the ordinary condition of a small nonprofit that adopted a useful tool. The Center for Effective Philanthropy's 2025 survey of nonprofit and foundation leaders found roughly two thirds of nonprofits using AI, and roughly two thirds reporting few or no staff who understand it well.
Disclosure. The Center for Nonprofit AI uses Anthropic's Claude in its research and drafting workflow, including for this article. Claude is one of the four tools examined below. CNAI has no commercial relationship with Anthropic, Google, Microsoft, or OpenAI, and every retention figure here comes from the vendor's own published documentation.
There is a common assumption that AI records sit in a regulatory gap, waiting for someone to write rules for them. For any nonprofit receiving federal funds, that assumption is wrong.
The Uniform Guidance requires recipients to retain all federal award records for three years from the date of submission of their final financial report, with the period extending if litigation, a claim, or an audit begins before those three years expire.
The more consequential provision governs access. Federal agencies, pass-through entities, Inspectors General, the Comptroller General, and their authorized representatives have the right of access to "any records of the recipient or subrecipient pertinent to the Federal award" in order to perform audits, execute site visits, or for any other official use.
Read that scope carefully. The obligation is defined by pertinence to the award. It is not defined by document type, file format, or which system the material lives in. There is no carve-out for a chat interface, and there was never any need to write one. The regulation does not define "Federal award records" either, which means a nonprofit's real defense is never that chat logs are categorically excluded, only that a particular exchange was not pertinent to a particular award. That is a fact-specific argument made after someone has already asked. It is not a policy.
One further detail is widely missed. The same section provides that rights of access "are not limited to the required retention period" but "last as long as the records are retained." Keeping material longer than the rules require does not sit outside the regime; it extends the window during which someone can demand it. That cuts against the instinct to solve this by keeping everything forever.
A third provision closes the loop. Recipients must establish, document, and maintain effective internal control over federal awards providing reasonable assurance the award is being managed in compliance with federal statutes, regulations, and the award's terms. For a federally funded nonprofit, governing how AI is used on award work is therefore not an emerging best practice to adopt when the sector matures. An organization with staff using AI on grant-funded deliverables and no documented control over it does not have a policy gap. It has a control gap, and control gaps are what auditors are trained to find.
Most executives will object that no funder has ever raised this with them, and they are largely right. In the same Center for Effective Philanthropy survey, only 17 percent of nonprofit leaders said they had discussed AI with their funders at all. But the retention obligation attaches to the award, not to whether anyone mentioned it. Funder silence is not permission and it is not a safe harbor. What it means is that no grantee will get advance warning.
Nor is clearer regulation coming. The Office of Management and Budget published a proposed rewrite of these rules in May 2026, with comments closing in July, and as of this writing it has not been finalized. It is the most significant proposed overhaul of federal grant requirements in a decade. It carries the three-year retention rule forward essentially unchanged and introduces no provisions addressing AI-generated material, retention of model interactions, or documentation standards for AI-assisted work. The only mentions of artificial intelligence anywhere in it appear in the preamble's policy justification, not in any operative requirement. Executives waiting for regulatory clarity before acting should understand that the clarity is not being drafted. The existing language is broad enough that regulators appear not to believe new language is needed.
Every AI interaction at a nonprofit is governed by four separate retention clocks. They do not agree, and most organizations are aware of only one.
The vendor's clock is set by the default configuration of whichever tool your staff opened. Depending on product and tier it runs from thirty days to eighteen months to indefinitely. Almost nobody knows what it is set to, and often the organization does not know the tool was opened at all. Writing in Human Service Organizations, Lauri Goldkind, Joy Ming and Alex Fink describe staff who use these tools to speed up tedious work and simply do not mention it, "machine-augmented humans who keep themselves hidden," with adoption running bottom-up in "ad-hoc, unsanctioned ways." When use is undisclosed, nobody chooses a retention setting. The default of a personal account governs by forfeit.
The funder's clock is set by award terms. For federal awards it is three years from submission of the final financial report, understood by whoever manages compliance in finance.
The organization's clock is your records retention schedule, written by someone competent several years ago, silent on prompts and outputs because those categories did not exist when it was drafted.
The legal clock is set by litigation hold, subpoena, or court order. It overrides the other three, can run indefinitely, and nobody thinks about it until the day it applies.
The failure is not that any single clock is wrong. Each is defensible on its own terms. The failure is structural: the shortest clock runs by default, and the longest governs the liability. Technology staff see the vendor's clock, finance sees the funder's, whoever wrote the schedule saw the organization's, and counsel arrives for the legal clock, usually late. Only the executive sees all four, which means only the executive can notice they disagree.
What follows reflects vendor documentation as of August 2026. These settings change, so treat the practice as durable and the numbers as perishable.
Google distinguishes two things that sound alike. The Gemini app, the standalone assistant staff open in a browser or on a phone, offers automatic deletion after three, eighteen, or thirty-six months, or none at all, and the default is eighteen months. Gemini in Workspace, the assistant embedded in Docs and Gmail, is a separate setting with its own options and no documented default. The eighteen-month default belongs to the app, which is the one people open when they want to think something through, and therefore where consequential work happens.
Set that against a three-year federal retention obligation and the conflict is arithmetic rather than interpretive. On a multi-year award the tool deletes material on a rolling eighteen-month cycle while the obligation still has years to run. Nobody chose that. It emerged from a product default meeting a regulation, with no one in the room who could see both.
Microsoft 365 Copilot presents the opposite exposure. Prompts, retrieved data, and responses are stored alongside the organization's other Microsoft 365 content, specifically in a hidden folder in the mailbox of the user who ran the tool, where administrators can manage them through Content search or Purview and apply retention policies. Microsoft publishes no default retention period, so retention is governed entirely by whatever policy the organization has applied. Where none has been applied, material accumulates in user mailboxes with no expiration anyone set.
OpenAI's business products sit between. On ChatGPT Enterprise and Edu, administrators control retention, deleted conversations are removed within thirty days, and an audit log is available through a compliance API. On the API, inputs and outputs may be retained up to thirty days for abuse detection, with zero data retention available only on request for eligible endpoints.
Anthropic's Claude is the third pattern. On Enterprise plans data is retained indefinitely by default unless a Primary Owner or Owner sets a custom period, with a thirty-day minimum. Retention changes are tracked in audit logs, and Enterprise customers can reach conversation content through a compliance API, which is an access layer rather than a discovery product; the eDiscovery capability organizations actually use is built on top of it by third parties.
Four tools, four fundamentally different retention postures, and the difference will usually be invisible to the executive who approved the purchase.
One detail in Claude's design deserves attention beyond the product, because it exposes an assumption buried in all of these tools. Its Enterprise retention period runs from the time of the last message in a conversation, not from when the conversation was created, and expired data cannot be recovered.
That is a sensible way to build a productivity tool. Active work stays available, abandoned threads age out. It is not how any retention obligation works. The federal requirement runs three years from submission of the final financial report, a fixed calendar date set by the award.
Consider the consequence. A staff member has one long exchange in September producing the analysis behind a funder report, then never reopens it. Under an inactivity clock that conversation begins its countdown the day it ends, which is precisely the day it became consequential. A thread reopened weekly to reformat meeting agendas, containing nothing anyone will need, is preserved indefinitely by the same rule.
The tool measures engagement. The obligation measures accountability. Nothing in the product knows the difference, and nothing in it is supposed to.
OpenAI's enterprise documentation states that deleted conversations are removed within thirty days "unless we are legally required to retain them." That final clause is not boilerplate. It is a live condition, and the organization using the product has no control over when it triggers.
The copyright litigation against OpenAI in the Southern District of New York shows how. In May 2025 a magistrate judge directed OpenAI to "preserve and segregate all output log data that would otherwise be deleted on a going forward basis," expressly including data that would have been deleted at a user's request. That obligation ran until it was terminated by stipulation as of late September 2025. For roughly four months, the delete button in a widely used product did not do what its users reasonably believed, for reasons no user could have known.
In January 2026 the court affirmed orders requiring production of the entirety of a de-identified twenty million conversation log sample. The detail worth holding onto is where that number came from: OpenAI proposed the sample itself, as a counter to a demand for one hundred twenty million, and was held to its own proposal.
The lesson has nothing to do with copyright law. A deletion setting is a request to a vendor, not a property of the record. What happens to a conversation is determined by the vendor's obligations, including those arising from litigation your organization is not party to and will not hear about. That cuts both ways: the material you assumed was gone may not be, and the material you assumed was retained may not be either. Neither assumption is a retention position.
Not every AI interaction is a record, and an organization that retains all of them creates a discovery liability worse than the one it was solving. The practical question is which ones matter.
This is not a speculative category. Goldkind and her colleagues describe AI use in human services running from low-stakes donor acknowledgments through to clinical documentation. Clinical documentation is not a draft of the record. It is the record.
Substitution. Did it replace work a person would otherwise have documented? If a staff member would have produced a memo or an analysis your schedule already covers, the interaction that produced it in their place inherits that classification. The tool changed, the record did not.
Consequence. Did it shape a decision, a deliverable, or a representation made to a funder, regulator, board, or client? This is the question that catches grant-funded work, and the one that surprises executives. An AI-drafted outcomes narrative in a funder report is a representation to a funder.
Exposure. Does it contain information the organization would have to produce if asked? Client data, donor data, personnel information, and program records do not change category because they were pasted into a chat window.
Interactions failing all three, and most do, are working material. Brainstorming, rephrasing, and formatting help on documents that never leave the building are transient by nature. The point of the test is not to expand what you keep. It is to stop the decision being made by a product default.
The sector's own best advice points the other way. Goldkind and her colleagues recommend bringing hidden AI use into the open, surfacing and celebrating the staff who found productive uses, and encouraging people to save useful prompts and responses to share. That is good advice, and it carries a consequence nobody has written down: the moment you surface shadow use, you convert a practice you could not have known about into one you demonstrably know about. Knowledge is the trigger condition for a preservation duty. Doing the right management thing creates the obligation the wrong one obscured.
That is not an argument against surfacing it. It is an argument for surfacing it and settling the retention question in the same quarter. It is also worth conceding that a retention regime answers only whether you can produce the record, not whether the decision in it was sound. Billie Sandberg, Rafeel Wasif and Laura Hand argue that accountability and transparency, while necessary, are insufficient on their own, because they locate the problem in an individual or a system and leave the surrounding arrangements untouched. Getting the records right is the floor of this conversation, not the ceiling.
The work is unglamorous and not primarily technical.
Find out what the vendor's clock is set to on every AI tool in use, including the ones nobody formally approved. Compare it against your longest retention obligation, which for most organizations is a funder requirement rather than a statutory one. Amend the retention schedule to name prompts, outputs, and chat logs explicitly, applying the record test rather than inventing a category. Decide, before it matters, which system is the system of record for AI-assisted work on funded deliverables, and require the material land there rather than in a vendor account. Make sure whoever manages litigation hold knows these systems exist.
If you are not sure where your organization stands on the underlying capability, our AI readiness assessment covers governance and data foundations among its five dimensions, and our guide to writing an AI policy is the companion piece to this one.
None of this requires resolving the larger questions about AI the sector is still arguing over. It requires accepting that the tools are already in use, the obligations are already in force, and the gap between them is currently managed by nobody.
The program officer's question is not unreasonable. It is the same question funders have always asked. What has changed is that the answer now depends on a setting in a product, and on whether anyone at your organization thought to look at it.
Sources
Hillen, A. (2024). Exploring artificial intelligence tool use in a nonprofit workplace. Journal of Business and Technical Communication, 38(3), 213-224. doi.org/10.1177/10506519241239661
Goldkind, L., Ming, J., & Fink, A. (2025). AI in the nonprofit human services: Distinguishing between hype, harm, and hope. Human Service Organizations: Management, Leadership & Governance, 49(3), 225-236. doi.org/10.1080/23303131.2024.2427459
Sandberg, B., Wasif, R., & Hand, L. C. (2025). Addressing the promise and peril of AI for nonprofit management through a data feminist pedagogy. Journal of Public Affairs Education, 31(2), 192-212. doi.org/10.1080/15236803.2025.2475589
Center for Effective Philanthropy (2025). AI With Purpose: Nonprofit and Foundation Leaders on Artificial Intelligence. cep.org
Retention requirements for records, 2 C.F.R. § 200.334. ecfr.gov
Access to records, 2 C.F.R. § 200.337. ecfr.gov
Internal controls, 2 C.F.R. § 200.303. ecfr.gov
Office of Management and Budget (2026). Regulation for Federal Financial Assistance, proposed rule. 91 Fed. Reg. 32,198 (May 29, 2026). federalregister.gov
Google (2026). Generative AI in Google Workspace privacy hub. Accessed 14 August 2026. knowledge.workspace.google.com
Microsoft (2026). Data, privacy, and security for Microsoft 365 Copilot. Accessed 14 August 2026. learn.microsoft.com
OpenAI (2026). Enterprise privacy. Accessed 14 August 2026. openai.com
Anthropic (2026). Configure custom data retention controls for Enterprise plans. Accessed 14 August 2026. privacy.claude.com
The New York Times Co. v. Microsoft Corp., No. 1:23-cv-11195 (S.D.N.Y.), ECF 551, preservation order (May 13, 2025); ECF 922, stipulated termination (Oct. 9, 2025).
In re OpenAI, Inc., Copyright Infringement Litigation, No. 1:25-md-3143 (S.D.N.Y.), ECF 1021, order affirming discovery rulings (Jan. 5, 2026).
Images on this page are AI-generated, created by the Center for Nonprofit AI.
A concise monthly briefing on AI strategy for the sector. Articles like this one, delivered when they publish.