What we found
2026 Deep Dive into Capacity for Evaluating Impact
Nonprofit leaders, capacity-builders, and funders are working together to build organizations that are strong enough to know whether their work is making a difference and well supported enough to act on what they learn. The recommendations below follow from the findings and are grouped by these three audiences, each holding distinct leverage to improve the culture and habits of evaluating nonprofit impact. Each list gathers the actions that recurred most across the five state studies.
Recommendations
-
Build a reflection rhythm you can sustain. Short post-program debriefs, a quarterly look at key metrics, and a standing data question in staff meetings all count as evaluation. A simple habit the team will keep does more than an elaborate strategy no one will follow.
Use the data you already have. Before collecting anything new, protect the time to interpret and act on what is already in hand, and develop staff to make sense of data, not just collect it. Interpretation and strategic response, not collection, is often the point of failure.
Share evaluation across the team, the board, participants, and the community. Involve people in parts of the process, including data collection, sensemaking, and shaping recommendations. Without participation, there is little ownership; involving people enhances value for everyone.
Seek out voices that are usually missed. Evaluation is stronger when it includes people whose perspectives are easy to overlook. Build the relationships and methods that let participants and community members shape what gets asked and how.
Shift from counting outputs to evaluating impact. Ask about what has changed for people or places, not only what was delivered. Question whether you are treating compliance as evaluation, and build the capacity to gather and make sense of evidence about the difference you make.
-
Convene peer cohorts and communities of practice. Many leaders named structured time with others doing similar work as the support they most need. Leaders learn as much from peers as from formal training. Build cohorts that can learn together over time.
Teach the move from data to case-making. Most leaders can collect data; far fewer can turn it into a case that donors, funders, and legislators actually hear. Turning data into a compelling case is a skill many leaders would benefit from developing.
Offer right-sized tools and design guidance. Provide simple, tiered guidance matched to an organization's size, mission, and stage, including methods for arts, advocacy, and other relational work that standard program frameworks handle badly, and help leaders choose the few indicators that matter most to their organizations and communities.
Focus on sensemaking and mental models, not just tools. Support reviewing data together, legitimize qualitative and mixedmethods work, and challenge unhelpful assumptions about evaluation and learning.
Support sustained, culturally grounded learning. Favor multi-year capacity-building over one-time training, make evaluation as visible and accessible as fundraising or strategy training, and ensure support for tribal and rural communities is grounded in culture, lived experience, and relationships, not only technical expertise from the profession of program evaluation.
-
Fund evaluation itself. Treat evaluation as a fundable line item, including staff time, training, data systems, and support for making sense of what is collected. This includes funding the coalitions, networks, and shared infrastructure that help organizations evaluate together.
Coordinate and simplify reporting. Move toward shared core metrics, common applications, and fewer, more substantive reports, so grantees can focus on learning from evaluation rather than producing accountability documents.
Close the feedback loop. Share aggregated findings back with the organizations that did the data collecting to make reporting a two-way exchange that helps grantees learn from themselves and from one another rather than a one-way submission.
Broaden what counts as evidence. Accept qualitative, relational, and community-defined evidence alongside required metrics, especially for arts, advocacy, relational, and long-horizon work whose outcomes a standard annual report cannot capture readily.
Make funding relationships trust-based and long-horizon. Co-design metrics with grantees, extend grant terms to match how long outcomes take to unfold, and create space for low-stakes evaluation decoupled from funding decisions. In the current climate, do not penalize organizations that limit data collection to protect vulnerable people.
Key Findings: Deep Dive into Capacity for Evaluating Impact
-
Involving people meaningfully in the evaluation process is a hallmark of the strongest evaluation cultures in the study. Internally, these evaluations involved staff at every level, including the frontline workers who hold the closest relationships with the people a program serves. Those workers usually have the most direct information about how the work is being received and some of the best instincts about how to respond to data in practice. Externally, the strongest cultures also involved participants, partners, and community members in the work of evaluation, not only as a source of data or an audience for a report.
1-
Most evaluations focus on confirming positive data. Grant reports usually highlight what worked. Boards typically read the encouragement of increased participation numbers. Little in this system of confirmation prompts an organization to ask seriously whether its work makes the difference it claims. And when funding hinges on the answer, the most important question about impact can become the one too risky to ask.
2-
What looks like a resource problem is often a literacy problem. The most common challenge is not collecting data. Many organizations already have more data and tools than they use, including CRMs, national benchmarking data, and dashboards. Still, they cannot reliably turn them into decisions or into a compelling case for funders. The gap is in designing good metrics, telling measurement apart from meaning, separating outputs from outcomes, and knowing what not to measure. The data they have are often not the data they need, and the methods they use are ill-suited for evaluating and communicating impact with credibility.
3-
The impact organizations want to claim most is often the most elusive to evaluate. Activity is easy to measure, including meals served, beds filled, participants in programs, diapers distributed, animals treated. The deeper and more durable results are challenging to measure with the tools and timelines most organizations have. Data about changes in a person’s stability, dignity, confidence, civic participation, or long-term trajectory of lifepath are beyond the expertise of many nonprofit leaders. They tend to measure what is visible and immediate, even when they know the work toward impact is why the organization exists.
4-
Some of the constraints leaders described are not organizational. They are failures at the level of funders, systems, and structures that encase nonprofit work: duplicative data collection, funder requirements that conflict across grants, government systems that do not connect, and reporting burdens larger than the grants they attach to. Interviewees report that government and funder databases are often incompatible or outdated. Addressing evaluation systems throughout the sector requires coordination above the level of any single organization.
56-
Evaluation is worthwhile when it becomes a tool for learning rather than a task for reporting. In its strongest form in the interviews, better evaluation produced a clearer case for the work, which brought stronger funding, which built capacity, which produced more impact. Strong evaluation also always resulted in action. That reflection-to-action cycle exists in pockets across the five states. The reporting-driven habit that still dominates most evaluations undermines the value of evaluation. It makes evaluation more work than it is worth.
7-
For evaluation to mean something, the right people must be part of it, and that is harder than it sounds. Standard survey-based instruments can burden or misread vulnerable communities. In rural, frontier, and dispersed places, email-based outreach and city-sized participant expectations do not fit the population. Scheduling an interview with a participant who is on the fringe or hard to reach requires extraordinary effort, the kind many staff have little time or energy for. The organizations that do this well have invested in years of relationship-building and attentiveness to inclusion.
-
Nonprofit leaders across all five states are using data to make the case for their value, but mostly to those who are already convinced. Board members, current donors, and longtime funding partners who understand the premise of the work receive their data with few questions. Others, however, like corporate partners, legislators, unfamiliar funders, and community members who do not share the premise, need a substantial rendering of the value proposition through data. Few leaders in this study know how to make a compelling case for new audiences.
8-
Many leaders described pressure from the broader funding and policy climate. Acute federal and state cuts have destabilized many organizations, resulting in layoffs, lost contracts, and the fear of
more instability. That destabilization undercuts the steady listening and implied trust that good evaluation depends on. At the same time, fear of immigration enforcement and shifting rules around demographic and DEI data are leading some organizations to collect less data, not more.
910-
Interviews show that two common options for evaluation are both inadequate. Third-party evaluation is often the default approach: a consultant runs a study, delivers a report, and leaves. It satisfies a grant requirement but rarely builds insight that outlasts the funding. The organization is no better equipped the following year. A fully in-house evaluation is the other option, but most small and mid-sized organizations cannot resource it. They lack the staff time and technical skills to design and implement consistent impact evaluations.
Evaluating Impact State Reports