Skip to content
Abelardo Carlos

InteractivearXiv preprint · 2026 · The World Bank

Query2Effect: can a model predict what an experiment would find?

Giuliano Martinelli, Piriyakorn Piriyatamwong, Abelardo Carlos Martinez Lorenzo, Jasmin Baier, Riccardo Orlando, Satvik Garg, Sharif Kazemi, Linxi Wang, Arianna Legovini, Samuel P. Fraiberger

Policymakers ask causal questions long before a trial can answer them. Using thousands of randomised controlled trials, we built a benchmark of 72,000 such questions and showed that first restating the question as a small, structured trial helps a model estimate the effect, especially in unfamiliar domains.

natural-language queries linked to RCT effect sizes
72Knatural-language queries linked to RCT effect sizes
absolute error versus GPT-5.2 in-domain
−27%absolute error versus GPT-5.2 in-domain
absolute error versus GPT-5.2 out of domain
−34%absolute error versus GPT-5.2 out of domain
R² even from gold trial descriptions: the task is genuinely hard
0.23R² even from gold trial descriptions: the task is genuinely hard
01

Questions that come before the evidence

“By how much can fertilizer vouchers increase maize yields for smallholder farmers in rural Africa?” Randomised controlled trials answer questions like this by measuring the causal effect of an intervention on an outcome. They are the cornerstone of evidence-based policy, and they are slow and expensive.

NLP for science has mostly looked backwards: retrieving, extracting and synthesising what trials already found. We ask a forward-looking question: can a model predict a continuous effect size, with a confidence interval, from the question alone?

The task

f(query) → (effect, CI low, CI high)

  • no treatment or control outcomes observed
  • no covariates, no experimental data
  • must generalise across interventions and populations

Semantic extrapolation: infer a likely causal magnitude from language alone.

What the model has to predict Appendix A and D
difference = 0.25 SDcontroltreatedSTANDARDISED EFFECT (HEDGES’ g) · 95% CI|g| ≤ 0.1: not economically meaningful-0.500.511.5your trialMalaria mRDTs (Table 1)Nursing lectures (Figure 1)

g = 0.997 × 0.25 = 0.249
95% CI [-0.005, 0.503]

Direction

positive

Economic sig.

|g| > 0.1

Statistical

non-significant

The model sees none of this: only the question. It must predict g and both interval bounds, and is scored on the error and on these three readings a policymaker cares about.

Hedges’ g = J · (mean treated − mean control) / pooled SD, with J a small-sample correction. The interval uses the usual large-sample standard error; the two real estimates are drawn on the same axis.
02

Query2Effect: 72,000 questions with known answers

We started from an expert-curated corpus of RCTs, each estimate with its intervention, outcome, standardised effect size and confidence interval, the same kind of structured evidence that ImpactAI extracts. Then we worked backwards: for each estimate, an LLM wrote four questions whose answer is that estimate, ranging from fully specified to vague. Slide through one:

One estimate, four questions Tables 1, 2, 3 and 7
L0L1L2L3IAU
Implicitness · I0

All causal elements explicit

Abstraction · A0

Concrete phrasing

Ambiguity · U0

Clear causal intent

Level 0 · I0 · A0 · U0 · Fully specified query explicitly mentioning the intervention and outcome.

“What is the effect of introducing malaria rapid diagnostic tests (mRDTs) in public health centers for diagnosing malaria in children under five in rural Ghana, compared to relying solely on clinical judgment, on the aggregate societal cost per 1000 fever episodes over two years?”

interventionpopulationcomparatoroutcome

Average length (characters)

L0
224
L1
138
L2
108
L3
106

Same answer for all four

g = -0.0129 [-0.101, 0.075]

The target never changes; only how much of the trial the question gives away. Vaguer questions carry less of what decides the effect.

The taxonomy has three dimensions with four levels each; generation uses the four canonical profiles on the diagonal. Element colours are our reading of each query. Lengths: average characters on Test id (Table 3).
Working backwards from the answer Sections 3.3 and 3.4

RCT estimate 76717

Intervention. mRDTs in public health centres; treatment only after a positive test.

Outcome. Aggregate societal cost (health sector + household) per 1000 fever episodes over 2 years.

g = -0.0129 · [-0.101, 0.075]

Gemini 3 Pro × 4 profiles
  • ✓ preserve the meaning of the estimate
  • ✓ add no information the trial does not support
  • ✓ invent no experimental details
  • ✓ one sentence
L0

What is the effect of introducing malaria rapid diagnostic tests (mRDTs) in public health centers for diagnosing malaria in children under five in rural Ghana, compared to relying solely on clinical judgment, on the aggregate societal cost per 1000 fever episodes over two years?

L1

What impact does the introduction of mRDTs for malaria diagnosis in public health centers have on the societal costs associated with managing fever cases?

L2

How does introducing a new diagnostic tool in healthcare settings affect resource utilization efficiency compared to traditional diagnostic methods?

L3

How do diagnostic advancements influence public health economics?

96%

faithful to the RCT

94%

realistic, expert-like

52%

spotted the human (chance: 50)

Generator: Gemini 3 Pro. Validation on 500 queries by two annotators; the blind A/B test used 2,000 pairs.
From trials to a benchmark Section 3.1 and Table 3
7,354

RCTs

74,826

effect estimates

18,031

single intervention–outcome, vs baseline

72,124

queries (× 4 levels, Table 3)

health 13,804 · other sectors 4,227Train41,868Validation5,100Test id8,248Test ood16,908
Estimates from the same RCT never land in two splits. The health sector is used for training and in-domain testing; education, agriculture, social protection and others form the out-of-domain test.

Training and in-domain testing use health trials, whose interventions and outcomes are more standardised. Education, agriculture, social protection and other sectors form an out-of-domain test set.

03

Restate the question as a trial, then estimate

A person asked such a question would first work out which intervention and outcome it is about, and only then recall how large such effects tend to be. The Synthetic-RCT pipeline does the same: an LLM rewrites the query as a minimal trial description, without inventing details or numbers, and a fine-tuned regressor predicts the effect from that description.

Two ways to go from a question to an effect size Figure 1

Query

“Does receiving weekly lectures on complementary therapies improve nursing students’ preparedness for health challenges?”

↓ LLM (GPT-5.2 or open GPT-OSS-20B) writes a synthetic RCT

Synthetic RCT

Intervention: A complementary medicine programme aims to develop nursing interventions through an understanding and exploration of application on various therapies

Outcome: health competency of nursing students, assessed via online questionnaire, reflecting pre-post change.

↓ fine-tuned ModernBERT-large regressor (reads the synthetic RCT)

Effect size, 95% CI

0+1

+1.1 [+0.83, +1.42] · Statistically significant positive

The Synthetic-RCT is a semantic bottleneck: the model first restates the question as an intervention and an outcome, then a separate regressor estimates the effect.

Here is the pipeline on real test questions from the paper: three it gets almost exactly right, and three it gets badly wrong. The failures share a pattern: unusually large effects get pulled towards the typical size.

Six real predictions, best and worst Table 13

1 · Query

“Do monthly nurse home visits with patient education, self-management coaching, and physician care management lead to improved ADL scores, indicating lesser functional decline, among Medicare beneficiaries with ADL/IADL impairments compared to standard care?”

2 · Synthetic RCT

Intervention: Monthly home visits by registered nurses to Medicare beneficiaries with ADL or IADL impairments: patient education, self-management coaching and coordination with their physicians.

Outcome: Change in ADL score over the study period, compared with standard care.

3 · ModernBERT regressor → effect size

-1.5-1-0.500.511.52error +0.000gold 0.204pred 0.204

All six: predicted vs gold

perfectgold →↑ predicted
Synthetic-RCT (GPT-5.2) → ModernBERT on in-domain test queries. Synthetic RCT texts are shortened from the paper. Large effects in either direction are the hardest to predict: the model pulls them towards the typical size.
04

Results

Three findings. First, prompted LLMs do poorly: on error metrics they lose even to a baseline that always predicts the average effect. Second, fine-tuning on Query2Effect helps a lot: the best model cuts absolute error by 27% versus GPT-5.2. Third, the structured step pays off out of domain, lowering error by 22% over the same regressor reading raw queries, and by 34% versus GPT-5.2.

In-domain · health RCTs (Table 4 · Test_id) specific queries
BaselinesPrompted LLMsSupervised (ours)Gold RCT ceiling
mean-effect
0.195
retrieval (BM25)
0.253
Gemini 2.5 Flash
0.354
GPT-OSS 120B
0.416
GPT-5.2
0.237
ModernBERT on queries
0.174
Synthetic-RCT (GPT-OSS-20B)
0.179
Synthetic-RCT (GPT-5.2)
0.172 ★
Gold RCT → ModernBERT
0.157
Mean absolute error: lower is better. Bars start at zero. The best non-oracle score is labelled.

The main tables use fully specified questions. As questions get vaguer, every model loses signal, and the pipeline keeps its lead at every level.

The vaguer the question, the weaker the signal Figure 2, digitised
0.00.10.20.30.4Level 0Level 1Level 2Level 3

A fairer target for vague questions

A level-3 question like “how do diagnostic advancements influence public health economics?” is closer to a meta-analysis than to one trial. Scoring it against the average of similar estimates (same intervention and outcome name; 1,804 of 2,062 questions have one, 2.7 on average) changes the picture.

ModernBERT on queries

one trial 0.083 → averaged 0.178

Synthetic-RCT (GPT-OSS-20B)

one trial 0.131 → averaged 0.220

Table 9 · level-3 queries on Test id. ○ scored against one trial, ● against the average.

Pearson correlation between predicted and gold effects on Test id, per query level. Values read from the paper’s plot (±0.005). The Synthetic-RCT pipeline stays on top at every level.

In-domain, the two supervised variants are close: both are trained on similar queries and near the ceiling set by gold trial descriptions. The synthetic trial matters when phrasing is unfamiliar, because it normalises the question toward what the model has seen. On AidGrade, which aggregates 600+ development-economics RCTs, the pipeline gets the sign of the effect right about 85% of the time.

05

Honest caveats

A hard target

Effect sizes are noisy and cluster around similar magnitudes, so R² stays low and the mean is a strong baseline. Correlation and policy-oriented metrics tell more.

Specific queries

The main experiments use fully specified queries; vague ones lack the context an estimate really needs.

English only

The framework is language-agnostic, but only English queries were evaluated.

Where this fits

This work sits next to ImpactAI at the World Bank: the same structured RCT evidence that powers evidence synthesis can also be used to anticipate effects before new trials exist.

The ImpactAI case study →

Cite

@article{martinelli2026predicting,
    title = "Predicting Causal Effects from Natural Language Queries using Structured Representations",
    author = "Martinelli, Giuliano  and
      Piriyatamwong, Piriyakorn  and
      Martinez Lorenzo, Abelardo Carlos  and
      Baier, Jasmin  and
      Orlando, Riccardo  and
      Garg, Satvik  and
      Kazemi, Sharif  and
      Wang, Linxi  and
      Legovini, Arianna  and
      Fraiberger, Samuel P.",
    journal = "arXiv preprint arXiv:2605.29631",
    year = "2026",
    url = "https://arxiv.org/abs/2605.29631"
}
All publications

Examples from Table 1 and Figure 1; scores from Tables 3–6.