Test, learn, adapt: Designing evaluations for faster learning
This blog is part of a series called “Test, learn, adapt,” in which impact evaluation experts explore research methods and answer common questions that come up in conversations with practitioners. Read the first post here.
When discussing randomized controlled trials (RCTs) with potential partners, two common critiques often come up: that they are too expensive and take too long. A brief look across the vast number of RCTs conducted over the past two decades would seem to support this conclusion—many took years to run and required substantial financial resources.
However, this critique misses an important point. The cost and timeline of some RCTs are not inherent limitations of the RCT design itself, nor are they unique to RCTs as a method. Rather, they largely reflect the research questions these studies typically ask—often about longer‑term outcomes—and the high-quality data required to answer them. This data requirement often includes custom survey data, which tends to be the most expensive component of an RCT.
Focusing on cost and timelines without unpacking what drives them can obscure the range of design choices available within randomized evaluations. Once we separate the method from the question and the data, it becomes clear that RCTs can be designed to be more nimble than is often believed and can generate rigorous evidence faster and at a lower cost.
In this blog post, I unpack what typically drives the cost and timeline of RCTs and show how thoughtful choices about research questions and data can make learning quicker and cheaper without compromising rigor.
How the evaluation question affects the timeline
One of the most important determinants of how long an RCT takes is the question it is designed to answer. Most RCTs highlighted on J‑PAL’s website are impact evaluations that focus on estimating a program’s effects on longer-term outcomes that affect people’s lives and livelihoods, such as employment, earnings, or health. Answering questions about a program’s final impact is relevant when the goal is to assess cost-effectiveness and make decisions about whether to scale the program. But RCTs can also provide earlier indications of whether minimum requirements for impact are met, help identify challenges related to implementation or the underlying program theory, and inform improvements to program delivery.
Assessing early indications of success (or failure)
For a program to have an impact, certain criteria must be met: for a job training program to affect earnings and employment, people must attend and acquire the intended skills; for a handwashing campaign to affect health, people must increase handwashing.
An RCT focused on outcomes that materialize earlier in a program’s theory of change can help policymakers and implementers learn about potential impact—or a lack thereof—early on and inform the best path forward. This might involve revising the program design or implementation if early indications of success are not found, or going ahead with a full-scale evaluation if they are. Testing whether a program is implemented as intended, or whether it affects intermediate outcomes, helps assess whether the assumptions underlying longer-term impacts are likely to hold.
For example, if an evaluation finds that people are not learning from a job training program, implementers can revisit the curriculum or the delivery of the program. This saves both time and money by pivoting earlier in the process compared to waiting to see the effects on income or earnings, which are unlikely to materialize.
Importantly, evaluations focusing on intermediate outcomes are typically used to complement, rather than replace, impact evaluations focusing on longer-term outcomes. This can be done either through pilots—smaller-scale evaluations designed explicitly to test effects on these intermediate outcomes before a full-scale evaluation is rolled out—or by collecting midline data in full-scale evaluations.
Optimizing program delivery through A/B testing
Randomized evaluations can be used not just to ask whether a program works, but how to deliver it most effectively. These questions are especially relevant when a program already exists and implementers are making concrete decisions about outreach, take‑up, participation, or day‑to‑day implementation.
In an A/B test (sometimes referred to as “rapid-fire operational testing” in the context of program implementation), participants are typically randomly assigned to receive slightly different versions of a specific program component. This might involve alternative recruitment messages, different timing or frequency of reminders, or small changes to defaults or incentives, while the rest of the program remains the same. The goal is typically not to compare fundamentally different programs, but to learn which delivery choice performs better along an operational outcome like enrollment, attendance, completion, or engagement.
A/B tests are not a different methodological category from other randomized evaluations. They rely on the same core logic: random assignment creates valid comparison groups, allowing differences in outcomes to be attributed to the design choice being tested rather than to underlying differences between participants. What differs is the question and, often, the outcomes being measured.
Because these tests frequently focus on outcomes that can be observed quickly through program operations, they can generate actionable evidence on a faster timeline and with fewer resources than an evaluation focused on longer-term outcomes, such as earnings or health.
Several organizations have documented how this approach is used in practice, including Youth Impact on Twelve iterative A/B tests to optimize tutoring for scale, Educate! on Maximizing impact through continuous rapid evaluation, and Innovations for Poverty Action (IPA on How Rapid-Fire Testing Can Build Better Financial Products. There are also comprehensive toolkits for A/B testing, such as those by IPA and What Works Hub.
Enabling iterative learning through adaptive designs
Adaptive designs use random assignment iteratively to test, learn, and refine program delivery over multiple rounds by allowing programs to update which variations are tested as evidence accumulates—for example, by dropping clearly underperforming versions and exploring new variations of promising ones. As with A/B testing, the source of credibility remains the same: at each stage, comparisons are grounded in random assignment and a well‑defined counterfactual.
Taken together, these examples illustrate that RCTs can be designed to answer a wide range of longer- and shorter-term questions. Some are aimed at estimating effects on longer-term outcomes; others test whether earlier links in a causal chain are working; and others are designed to optimize delivery choices. These choices about which question to ask have major implications for how quickly an evaluation can generate useful evidence and how resource‑intensive it will be, even though the underlying standard of causal inference through randomization remains unchanged.
How data choices affect costs
In the previous section, I focused on how the questions an evaluation is designed to answer shape how quickly evidence can be generated. Another important part of evaluation design that affects both cost and speed is the data used to answer those questions.
In practice, data collection accounts for a large share of the resources required for most impact evaluations—regardless of the method or the question being asked. In particular, collecting primary survey data often makes up the majority of an evaluation’s overall cost. Survey teams need to be recruited and trained, surveys piloted, and systems put in place to ensure data quality and respondent protection. These practical realities are not unique to randomized evaluations, but they often end up driving the cost of survey‑based impact evaluations, including RCTs.
Reducing the costs of primary data collection
When an evaluation relies on survey data collected specifically for the study, careful choices about how surveys are conducted can sometimes reduce costs substantially. One important dimension is the mode of data collection. In‑person surveys often involve logistically complex fieldwork, particularly when respondents are geographically dispersed and surveyors need to travel to remote areas. While in‑person surveys are often necessary—especially for longer surveys or sensitive topics that require trust and rapport—remote surveys can be a lower‑cost alternative in some settings.
Remote surveys, conducted by phone or through messaging platforms such as WhatsApp, can reduce travel and field logistics costs and make it feasible to reach respondents more quickly. They can be particularly useful for shorter surveys, frequent check‑ins, or time‑sensitive outcomes.
At the same time, remote surveying brings trade‑offs that can affect both data quality and who is reached. Not everyone has reliable access to a phone or data, response rates can differ from in‑person surveys, privacy can be harder to ensure, and some topics may be less appropriate to ask remotely. For these reasons, remote surveying is best viewed as an additional tool that can lower costs in the right context, rather than a universal replacement for in‑person data collection.
Choosing the number and timing of survey rounds
Survey costs also depend heavily on how many rounds of data collection an evaluation includes. Many RCTs are designed with multiple rounds—such as baseline, midline, and endline—to measure change over time, understand dynamics during implementation, or capture both intermediate and long-term outcomes.
Baseline data can bring several benefits. It helps researchers and implementers understand the study sample and the population a program serves. It often informs the randomization process, for example, by ensuring randomized groups are similar on key characteristics. It provides a way to verify that randomization produced comparable groups. In analysis, including baseline data can also improve statistical precision.
At the same time, one advantage of a randomized evaluation is that when random assignment is implemented correctly, the comparison group functions as a valid counterfactual. This means that when randomization can be carried out without first collecting baseline survey data, baseline data is not strictly necessary to justify causal inference in the way that many non‑experimental approaches require.
Similarly, if the main objective is to run a lean evaluation and midline data would not change the course of action, it is possible to design a credible study with only an endline survey. That is, the number and timing of survey rounds should be driven by what information is truly needed to answer the research question, implement randomization well, and interpret results—not by default assumptions about what an RCT “must” include.
Using existing data sources when available
Because surveys are resource‑intensive, costs can often be reduced when the outcomes of interest can be measured using data that already exist. Two sources are especially promising when available and suitable: administrative data and remote sensing data.
Administrative data that is already being collected—such as government records, school administrative systems, clinic records, or program management information systems—can sometimes be used to measure outcomes at a fraction of the cost of primary data collection. When these data are reliable and accessible, they can substantially reduce evaluation costs and allow outcomes to be measured without repeated field visits. Administrative data also often cover large populations, which can be useful for studying heterogeneous effects or for tracking outcomes over longer periods than would be feasible with surveys alone.
Remote sensing data refers to information collected without direct contact with program participants, typically through satellite imagery or other geospatial technologies. These data can be used to measure outcomes related to geography, land use, agricultural activity, infrastructure, environmental conditions, or exposure to shocks. For example, J-PAL affiliated researcher Seema Jayachandran and coauthors evaluated a payments for ecosystem services program in Uganda by tracking changes in tree cover using satellite imagery linked to the GPS coordinates of participants’ land.
When these measures align well with an evaluation’s outcomes of interest, remote sensing can reduce reliance on field‑based surveys and make it possible to measure outcomes across large areas or over time at relatively low cost.
Of course, existing data sources are not a universal substitute for surveys, and they are not always available. Many important outcomes are not captured in administrative systems, and data access, quality, timing, and comparability vary widely across contexts. However, when these data are available and relevant, combining approaches can reduce costs—for example, by using administrative data for core outcomes and baseline data, while relying on targeted survey modules for outcomes that cannot be measured otherwise, and using different data sources to validate or complement one another.
The key point is that data choices—what you measure, how you measure it, and how often you do so—often determine the cost of an evaluation as much as the design itself. More broadly, timelines and costs are closely linked: questions about early outcomes can often be answered more quickly and with existing data, while longer‑term outcomes typically require more time and more resource‑intensive data collection.
“Nimbleness” is a design choice
While the term “nimble evaluations” is sometimes mistaken for a distinct tool, nimbleness is a matter of degree within any evaluation design. It reflects choices about the questions asked, the data used, and the stage of learning an evaluation is meant to support. These choices exist within any rigorous evaluation framework, including randomized evaluations, and they shape what evidence can be generated, how quickly, and how resource‑intensive that process will be.
Seen this way, designing evaluations for faster learning does not come at the cost of rigor. Instead, it requires being clear about what needs to be learned at a given moment and how best to learn it. Finally, nimbler evaluations often do not stand alone; they are frequently paired with more comprehensive impact evaluations to test program designs and delivery choices before a full program is implemented and evaluated at scale.