The World Bank · 2024 – present
Fine-tuned encoders: fast, cheap and domain-aware
Small encoders trained for one job each: a retriever and a reader that link text to a development-economics taxonomy in one pass, a relevance classifier, and an effect-size regressor, trained on data we created with LLMs.
What’s inside
Why encoders
One pass, no generation
Fast
An encoder reads the input once and scores: no token-by-token generation, so answers come back almost instantly.
Cheap
Small enough to fine-tune on one GPU and serve in a container; no per-call LLM cost at inference.
Domain-aware
Trained on our taxonomy, it knows that conditional and unconditional transfers are different things.
Linking to a taxonomy
Retrieve, then read
We adapted ReLiK (SapienzaNLP) to development economics: an e5-large-v2 retriever narrows thousands of taxonomy concepts to a handful, and a DeBERTa-v3-large reader picks the right one in context.
Question
What do we know about cash support for familieswith children in school ?
Encode the question
An e5-large-v2 encoder turns the question into a vector, the same space as every concept in the index.
Concept index
→ estimates for this concept → Data Analysis
Conceptual example of the problem, not a recorded prediction of the adapted checkpoint. Candidate names are illustrative.
The retriever
Two encoders, one space
Score: dot product of the two encodings
Training: contrastive loss with in-batch and hard negatives
Concepts are encoded once, offline, into an index (GoldenRetriever); at query time only the question is encoded, and the top concepts are a nearest-neighbour lookup away.
Query: “cash for families if children attend school”
s(q, e⁺) = -0.77 · loss ℒ = 16.014
Softmax over similarities with τ = 0.1, computed live. Training pulls the question towards its concept and pushes the others away, so the correct concept wins even against close siblings like conditional versus unconditional transfers. Toy 2-D vectors; the real ones have 1,024 dimensions.
The reader
Find the mention, choose the concept
One pass: the question and all candidates together
Illustrative probabilities. DeBERTa-v3-large reads everything at once, finds the mention and picks its concept.
Training data
Examples written by an LLM, checked by experts
There was no labelled data for this taxonomy. We used an LLM to write, for each concept, the ways people phrase it and the near misses that belong to a sibling, then trained the encoders on those pairs. Examples below are illustrative.
Taxonomy concept
Conditional cash transfers
Cash paid to households on condition of a behaviour, such as school attendance or health check-ups.
siblings: unconditional transfers · scholarships · school feeding
How users say it
Near misses (a sibling concept)
In context
More encoders
Same recipe, other jobs
Relevance classifier
encoder · (question, estimate) → relevant?
An earlier generation of ImpactAI’s relevance layer: one encoder pass per candidate estimate instead of an LLM call. Later replaced by a configurable Gemini judge.
Effect-size regressor
ModernBERT-large · regression head
Predicts an effect size and both interval bounds from a question or a synthetic trial, trained with mean squared error: 27% lower error than GPT-5.2 in-domain.
The Query2Effect story →ReLiK and GoldenRetriever are by SapienzaNLP. Training dataset, split and final scores of the adapted linker are still being gathered, so none are shown.