Every AI readiness assessment produces a number. Almost none of them will tell you how that number was produced. This page is the full method behind ours: the questions we ask, how answers convert to a score, how the five dimensions are weighted and why, and what the resulting score does and does not predict. We publish it because a score you cannot inspect is a score you cannot argue with, and because we would rather be corrected than trusted by default.
16 questions across 5 dimensions. Every question offers four options, scored 0 to 3 in the order shown. There are no reverse-scored or trick items.
Carries the most because it carries the most in the evidence. RAND's 2024 interview study found leadership and problem framing to be the most common root cause of AI project failure, cited by 84 percent of the 50 industry practitioners interviewed, ahead of data. Shi and Azevedo (2026) found that nonprofits which had never had an internal conversation about these tools were dramatically less likely to support adopting them, an odds ratio of 0.04, with organizational factors alone explaining half the variance in their model.
How often does AI come up in your board or senior leadership conversations?
Is there a named person responsible for AI decisions at your organization?
Does your strategic plan mention AI or data capability?
Consistently the second most-cited cause rather than the first. RAND's practitioners named data problems 30 times out of 50. A 2025 systematic review across 26 studies found poor data quality the most frequently cited barrier within its cluster. We hold this at 20 rather than higher because the one peer-reviewed study of nonprofit AI barriers, Kabra and Saharan (2025), does not place data quality in its top five.
Where does your member, donor, or client data actually live?
If your CEO asked for a key number today, how fast could you produce an accurate answer?
How would you rate the accuracy of your core records?
A Bank for International Settlements study of more than 12,000 European firms found that each additional percentage point of investment allocated to employee training raised the productivity effect of AI adoption by 5.9 percent, the largest complementarity of any input tested. A randomized field experiment published in Organization Science found the same tool made people roughly 40 percent better inside its capability range and 19 percentage points worse outside it, and the difference was whether people understood where that boundary sat.
Are staff already using AI tools like ChatGPT or Claude in their daily work?
Has your organization offered any AI training or guidance to staff?
What is the general mood about AI on your team?
Necessary but not sufficient. A policy does not create capability; its absence caps how much capability an organization can responsibly deploy. This is the weight we hold with the least confidence and the one most open to challenge, which is part of why we publish it.
Do you have an acceptable use policy that covers AI tools?
Your organization likely handles sensitive data. Is there guidance on using it with AI tools?
Who reviews a new AI tool before staff start using it?
Has leadership discussed AI risk, such as privacy, accuracy, or bias?
Scores lightest because current use is mostly an indicator of the other four rather than an independent cause. It earns its place because it is the dimension that reveals whether the other four scores are real. An organization can report a policy and a data system and still have staff pasting constituent records into a consumer chatbot.
Which best describes your organization's AI activity today?
Have you tied any AI effort to a measurable outcome, like hours saved or dollars raised?
Is there budget, even a small amount, set aside for AI tools or training?
Each dimension is scored independently as a percentage of its own maximum, then combined using the weights above. A dimension with four questions is not worth more than a dimension with three; the weight decides its influence, not the question count.
| Dimension | Weight | Questions |
|---|---|---|
| Leadership & Strategy | 30% | 3 |
| Data Foundation | 20% | 3 |
| People & Culture | 25% | 3 |
| Governance & Risk | 15% | 4 |
| Current Use | 10% | 3 |
Weights sum to 100 percent. The reasoning behind each one is stated in the dimension blocks above. Note what this replaces: an unweighted instrument silently weights dimensions by how many questions they happen to contain, which is an accident of drafting rather than a judgment about what matters.
Readiness is a constraint, not an average. An organization with excellent leadership and unusable data is not mostly ready, it is blocked. So the composite cannot sit more than one stage above the lowest-scoring dimension. A profile scoring 100 on four dimensions and 0 on the fifth does not return an 85. It returns a capped score, and the result names the dimension responsible.
This is the part of the method most likely to surprise people, and it is deliberate. Averaging lets an organization compensate for a bottleneck it cannot actually compensate for.
| Stage | Score | What it means |
|---|---|---|
| Observer | 0 to 25 | Your organization is watching AI from the sidelines. That is not a failure, but the gap between observers and everyone else is widening each quarter. The good news: at this stage, small moves like naming an owner and writing a one page policy create outsized progress. |
| Explorer | 26 to 50 | There is real curiosity and scattered activity in your organization, but no structure holding it together. Explorers who add a simple governance layer and one measured pilot typically reach the Builder stage within six months. |
| Builder | 51 to 75 | You have momentum. Leadership is engaged, staff are experimenting, and the basics of governance exist. Your challenge now is converting scattered wins into repeatable workflows with measured outcomes, which is what separates Builders from Operators. |
| Operator | 76 to 100 | AI is becoming part of how your organization actually runs. Your opportunity now is scale and stewardship: deepening measurement, formalizing governance, and sharing what you have learned with the sector. |
They are derived from the best available evidence about what predicts AI outcomes generally. They are not derived from a study of nonprofit AI outcomes, because no such study exists. When our own assessment data can support better weights, we will change them and say so here.
Nothing in the published literature establishes that a higher readiness score causes better AI outcomes in a nonprofit. We are measuring conditions that the evidence associates with success elsewhere. Treat the result as a structured diagnosis, not a forecast.
Every answer is your own assessment of your own organization. Executives tend to rate governance and data quality more favorably than the staff who work with them daily. If you want a sharper reading, have two or three colleagues take it separately and compare.
The instrument is deliberately short enough to finish. That trade buys completion at the cost of nuance, and it means a low score in one dimension should start a conversation rather than settle one.
If you think a weight is wrong, a question is poorly framed, or the evidence has moved, we want to hear it. Write to us through the contact page. Changes to the instrument or the weights will be recorded on this page rather than made quietly.
Instrument version 2, published 9 August 2026. This page is generated directly from the live assessment, so the questions above are always the questions being asked.