illustration of growth chart with bars and lines (Illustration by iStock/Galeanu Mihai)

Governments and funders have invested heavily in evidence to guide public spending and policy decisions over the past two decades. Randomized controlled trials, evidence clearinghouses, and tiered standards promised a way to move beyond ideology and toward what works, particularly in systems under pressure to justify decisions and demonstrate accountability. In many ways, they delivered on that promise, introducing rigor, discipline, and a shared language for accountability into complex public systems.

Yet, despite this investment, outcomes have not improved at the scale many expected. Across workforce, education, health, and human services, governments continue to struggle to translate evidence into lasting improvements in people’s lives. This gap has prompted renewed debate about whether evidence-based policymaking has reached its limits—whether clearing higher bars of proof alone can meaningfully improve results at scale. From my work with state and local governments, the more difficult conclusion is that evidence has been asked to do work it was never designed to do.

What Paying for Outcomes Revealed

A reentry initiative at Rikers Island illustrates the gap between evidence and outcomes directly. New York City implemented a cognitive behavioral therapy program intended to reduce recidivism among young men. The intervention had a strong evidence base. The financing vehicle—the ABLE Social Impact Bond, the first of its kind in the United States—was structured so the city paid only if recidivism fell by 10 percent or more. The financing worked as designed. An independent evaluation found no impact.

What went wrong was not the evidence base but the circumstances in which the program was attempting to create these outcomes. Conditions inside the facility were not what the program required to succeed. Coordination between the city’s correction and social service agencies was fragmented. The way services were actually delivered to young men on Rikers did not match the conditions under which the program had originally been tested. Each of these failures was knowable, but the contract had no mechanism to surface or respond to them.

This pattern is not unique to one project. Across jurisdictions that have invested in outcomes-based contracting—structures in which payment depends on whether programs achieve predefined results—the same dynamic has played out. Governments can change how they pay for results or whether they fund evidence-based programs. That alone does not determine whether results are achieved. The problem runs deeper than any single contract or program structure.

The Limits of the Evidence Base

Evidence can identify interventions that have worked under specific conditions. But it does not account for whether those conditions exist in practice or whether systems are capable of creating them.

The scope of what is captured in the evidence base is narrower than policy designers often assume—and the consequences are significant. Researchers have documented for decades that rigorous evaluations of social programs almost always fail to find meaningful impacts. Of the 13 large randomized controlled trials the federal government has commissioned to evaluate major congressionally authorized programs, 11 found either no significant positive effects or effects that faded shortly after completion. The programs that do clear the bar for strong evidence represent a narrow slice of the decisions governments face every day—yet policy frameworks treat that bar as the organizing principle for how public dollars get spent.

In many areas, no well-established, evidence-based programs exist that match the needs of the population being served. In others, programs that have demonstrated impact in one setting have not produced the same results elsewhere. The Nurse-Family Partnership, for example, a nurse-led home-visiting program for first-time, low-income mothers, showed strong outcomes in its original trials, but community replication has consistently produced smaller effects. Researchers found that nurses in replication sites were not retaining families at the same rate as in controlled settings. Evidence generated in one context does not reliably predict results in another.

When evidence-based programming becomes the primary organizing principle for funding, systems tend to align around compliance with approved models. Funding flows toward programs that meet evidence thresholds. Frontline teams are expected to implement those programs with fidelity. Data is collected to document adherence. Evaluation is used to determine whether predefined outcomes were achieved.

The result is consistency in how programs are administered, but not necessarily improvement in outcomes. Without structured systems for continuous data collection and learning, teams have no clear way to respond if the promised outcomes stall. Data gets reviewed after the fact rather than during implementation. Staff may not have the authority to adjust course even when problems are visible. Differences across populations or settings are treated as deviations from a model rather than signals that something needs to change.

Where Policy Is Heading

Recent federal policy reflects a continued effort to strengthen evidence-based approaches. The Evidence–Based Grantmaking Act (H.R. 7025), introduced in January 2026, would require 15 federal agencies to prioritize grant awards for applicants using evidence-based practices, conduct periodic evaluations during grant terms, and make those results public. It also directs the Office of Management and Budget to define “evidence-based” within one year and requires agencies to operationalize that definition through rulemaking and public comment over the following year.

The intent is right. Public dollars should support strategies with demonstrated results, and the field has spent decades building the infrastructure to identify them.

But strengthening evidence standards is not the same as building systems that can act on them. The legislation clarifies what should be funded and how results should be reported. It does not establish what agencies and grantees are expected to do when outcomes do not improve under real conditions, which is, in fact, the central challenge the field is grappling with.

What It Means to Govern for Outcomes

A different approach is already taking shape in some places. In Lane County, Oregon, local leaders brought together probation, behavioral health, and community providers to address how people with severe mental illness move through the supervision system, as part of a HUD and DOJ-funded Pay for Success Permanent Supportive Housing Demonstration initiative. The county drew on the Housing First evidence base—a well-established model that prioritizes stable housing before addressing other barriers—but did not require fidelity to any single program model. The organizing question was not which program to implement but whether individuals were stabilizing and avoiding reentry into crisis.

To support that goal, partners made three structural changes. A dedicated parole officer was stationed on-site at the housing development, with a caseload capped at roughly 50 participants—replacing the standard model of monthly office check-ins. A unified system case plan was introduced across corrections, housing, and service partners, replacing siloed individual plans and ensuring everyone working with a participant was operating from the same priorities. And partners established monthly continuous improvement meetings where leaders reviewed a shared performance dashboard tracking housing stability and recidivism in real time, adjusting referral protocols and staffing when results diverged from expectations.

The results to date are significant. Of the 231 individuals placed in permanent supportive housing through the initiative, 87 percent have maintained stable housing. The recidivism rate for program participants—defined as reincarceration for a new felony—has fallen to 11 percent, compared to 26 percent for the high-risk reentry population in Lane County overall. For the first time, the county is seeing recidivism rates for individuals assessed as high-risk on par with lower-risk populations.

This approach places different demands on public systems. Access to outcomes-level data becomes essential. Staff need the ability to interpret that data and the authority to act on it. Coordination across agencies becomes central, not optional. Success is defined by whether outcomes improve, not whether program requirements are met.

Changing how government pays for results is not enough if the systems delivering those results are not built to learn and adapt. What’s required is the institutional capacity to learn whether programs are actually improving lives and to act on what they find. Without continuous learning built into how programs operate, even well-supported approaches will continue to fall short of their promise.

If federal policy continues to emphasize evidence, it can also clarify how systems are expected to respond when evidence does not translate into results. Agencies could require that grant applications specify how data will be used during implementation to adjust service delivery, not only at the end of a grant term. Evaluation could be structured to inform decisions while programs are operating. Guidance on evidence standards could address how systems should proceed in areas where evidence is limited or inconclusive. Without that clarity, agencies will continue to fund programs based on evidence thresholds without a shared expectation for how implementation should change when outcomes stall.

The question the Rikers experience surfaces is not whether evidence matters—it does. The question is whether our policy and systems architecture is built to use it in ways that actually improve lives. Outcomes depend on whether systems are structured to coordinate, to adapt, and to respond to what they are learning. Evidence-based programs can inform those efforts. They cannot substitute for them.

Read more stories by Caroline Whistler.