InteractivearXiv preprint · 2026 · The World Bank
Query2Effect: can a model predict what an experiment would find?
Giuliano Martinelli, Piriyakorn Piriyatamwong, Abelardo Carlos Martinez Lorenzo, Jasmin Baier, Riccardo Orlando, Satvik Garg, Sharif Kazemi, Linxi Wang, Arianna Legovini, Samuel P. Fraiberger
Policymakers ask causal questions long before a trial can answer them. Using thousands of randomised controlled trials, we built a benchmark of 72,000 such questions and showed that first restating the question as a small, structured trial helps a model estimate the effect, especially in unfamiliar domains.
- natural-language queries linked to RCT effect sizes
- 72Knatural-language queries linked to RCT effect sizes
- absolute error versus GPT-5.2 in-domain
- −27%absolute error versus GPT-5.2 in-domain
- absolute error versus GPT-5.2 out of domain
- −34%absolute error versus GPT-5.2 out of domain
- R² even from gold trial descriptions: the task is genuinely hard
- 0.23R² even from gold trial descriptions: the task is genuinely hard
Questions that come before the evidence
“By how much can fertilizer vouchers increase maize yields for smallholder farmers in rural Africa?” Randomised controlled trials answer questions like this by measuring the causal effect of an intervention on an outcome. They are the cornerstone of evidence-based policy, and they are slow and expensive.
NLP for science has mostly looked backwards: retrieving, extracting and synthesising what trials already found. We ask a forward-looking question: can a model predict a continuous effect size, with a confidence interval, from the question alone?
The task
f(query) → (effect, CI low, CI high)
- no treatment or control outcomes observed
- no covariates, no experimental data
- must generalise across interventions and populations
Semantic extrapolation: infer a likely causal magnitude from language alone.
g = 0.997 × 0.25 = 0.249
95% CI [-0.005, 0.503]
Direction
positive
Economic sig.
|g| > 0.1
Statistical
non-significant
The model sees none of this: only the question. It must predict g and both interval bounds, and is scored on the error and on these three readings a policymaker cares about.
Query2Effect: 72,000 questions with known answers
We started from an expert-curated corpus of RCTs, each estimate with its intervention, outcome, standardised effect size and confidence interval, the same kind of structured evidence that ImpactAI extracts. Then we worked backwards: for each estimate, an LLM wrote four questions whose answer is that estimate, ranging from fully specified to vague. Slide through one:
All causal elements explicit
Concrete phrasing
Clear causal intent
Level 0 · I0 · A0 · U0 · Fully specified query explicitly mentioning the intervention and outcome.
“What is the effect of introducing malaria rapid diagnostic tests (mRDTs) in public health centers for diagnosing malaria in children under five in rural Ghana, compared to relying solely on clinical judgment, on the aggregate societal cost per 1000 fever episodes over two years?”
Average length (characters)
Same answer for all four
g = -0.0129 [-0.101, 0.075]
The target never changes; only how much of the trial the question gives away. Vaguer questions carry less of what decides the effect.
RCT estimate 76717
Intervention. mRDTs in public health centres; treatment only after a positive test.
Outcome. Aggregate societal cost (health sector + household) per 1000 fever episodes over 2 years.
g = -0.0129 · [-0.101, 0.075]
- ✓ preserve the meaning of the estimate
- ✓ add no information the trial does not support
- ✓ invent no experimental details
- ✓ one sentence
What is the effect of introducing malaria rapid diagnostic tests (mRDTs) in public health centers for diagnosing malaria in children under five in rural Ghana, compared to relying solely on clinical judgment, on the aggregate societal cost per 1000 fever episodes over two years?
What impact does the introduction of mRDTs for malaria diagnosis in public health centers have on the societal costs associated with managing fever cases?
How does introducing a new diagnostic tool in healthcare settings affect resource utilization efficiency compared to traditional diagnostic methods?
How do diagnostic advancements influence public health economics?
96%
faithful to the RCT
94%
realistic, expert-like
52%
spotted the human (chance: 50)
Training and in-domain testing use health trials, whose interventions and outcomes are more standardised. Education, agriculture, social protection and other sectors form an out-of-domain test set.
Restate the question as a trial, then estimate
A person asked such a question would first work out which intervention and outcome it is about, and only then recall how large such effects tend to be. The Synthetic-RCT pipeline does the same: an LLM rewrites the query as a minimal trial description, without inventing details or numbers, and a fine-tuned regressor predicts the effect from that description.
Query
“Does receiving weekly lectures on complementary therapies improve nursing students’ preparedness for health challenges?”
↓ LLM (GPT-5.2 or open GPT-OSS-20B) writes a synthetic RCT
Synthetic RCT
Intervention: A complementary medicine programme aims to develop nursing interventions through an understanding and exploration of application on various therapies
Outcome: health competency of nursing students, assessed via online questionnaire, reflecting pre-post change.
↓ fine-tuned ModernBERT-large regressor (reads the synthetic RCT)
Effect size, 95% CI
+1.1 [+0.83, +1.42] · Statistically significant positive
Here is the pipeline on real test questions from the paper: three it gets almost exactly right, and three it gets badly wrong. The failures share a pattern: unusually large effects get pulled towards the typical size.
1 · Query
“Do monthly nurse home visits with patient education, self-management coaching, and physician care management lead to improved ADL scores, indicating lesser functional decline, among Medicare beneficiaries with ADL/IADL impairments compared to standard care?”
2 · Synthetic RCT
Intervention: Monthly home visits by registered nurses to Medicare beneficiaries with ADL or IADL impairments: patient education, self-management coaching and coordination with their physicians.
Outcome: Change in ADL score over the study period, compared with standard care.
3 · ModernBERT regressor → effect size
All six: predicted vs gold
Results
Three findings. First, prompted LLMs do poorly: on error metrics they lose even to a baseline that always predicts the average effect. Second, fine-tuning on Query2Effect helps a lot: the best model cuts absolute error by 27% versus GPT-5.2. Third, the structured step pays off out of domain, lowering error by 22% over the same regressor reading raw queries, and by 34% versus GPT-5.2.
The main tables use fully specified questions. As questions get vaguer, every model loses signal, and the pipeline keeps its lead at every level.
A fairer target for vague questions
A level-3 question like “how do diagnostic advancements influence public health economics?” is closer to a meta-analysis than to one trial. Scoring it against the average of similar estimates (same intervention and outcome name; 1,804 of 2,062 questions have one, 2.7 on average) changes the picture.
ModernBERT on queries
one trial 0.083 → averaged 0.178
Synthetic-RCT (GPT-OSS-20B)
one trial 0.131 → averaged 0.220
Table 9 · level-3 queries on Test id. ○ scored against one trial, ● against the average.
In-domain, the two supervised variants are close: both are trained on similar queries and near the ceiling set by gold trial descriptions. The synthetic trial matters when phrasing is unfamiliar, because it normalises the question toward what the model has seen. On AidGrade, which aggregates 600+ development-economics RCTs, the pipeline gets the sign of the effect right about 85% of the time.
Honest caveats
A hard target
Effect sizes are noisy and cluster around similar magnitudes, so R² stays low and the mean is a strong baseline. Correlation and policy-oriented metrics tell more.
Specific queries
The main experiments use fully specified queries; vague ones lack the context an estimate really needs.
English only
The framework is language-agnostic, but only English queries were evaluated.
Where this fits
This work sits next to ImpactAI at the World Bank: the same structured RCT evidence that powers evidence synthesis can also be used to anticipate effects before new trials exist.
The ImpactAI case study →Cite
@article{martinelli2026predicting,
title = "Predicting Causal Effects from Natural Language Queries using Structured Representations",
author = "Martinelli, Giuliano and
Piriyatamwong, Piriyakorn and
Martinez Lorenzo, Abelardo Carlos and
Baier, Jasmin and
Orlando, Riccardo and
Garg, Satvik and
Kazemi, Sharif and
Wang, Linxi and
Legovini, Arianna and
Fraiberger, Samuel P.",
journal = "arXiv preprint arXiv:2605.29631",
year = "2026",
url = "https://arxiv.org/abs/2605.29631"
}Examples from Table 1 and Figure 1; scores from Tables 3–6.