Skip to content
Abelardo Carlos
← Projects

The World Bank · 2024 – present

Fine-tuned encoders: fast, cheap and domain-aware

Small encoders trained for one job each: a retriever and a reader that link text to a development-economics taxonomy in one pass, a relevance classifier, and an effect-size regressor, trained on data we created with LLMs.

Why encoders

One pass, no generation

Fast

An encoder reads the input once and scores: no token-by-token generation, so answers come back almost instantly.

Cheap

Small enough to fine-tune on one GPU and serve in a container; no per-call LLM cost at inference.

Domain-aware

Trained on our taxonomy, it knows that conditional and unconditional transfers are different things.

Linking to a taxonomy

Retrieve, then read

We adapted ReLiK (SapienzaNLP) to development economics: an e5-large-v2 retriever narrows thousands of taxonomy concepts to a handful, and a DeBERTa-v3-large reader picks the right one in context.

Question

What do we know about cash support for familieswith children in school ?

Encode the question

An e5-large-v2 encoder turns the question into a vector, the same space as every concept in the index.

Concept index

Conditional cash transfersintervention · sub-type
Unconditional cash transfersintervention · sub-type
Cash transfersintervention · type
Scholarships & fee supportintervention · type
School enrolmentoutcome · type
intervention = "Conditional cash transfers"
→ estimates for this concept → Data Analysis

Conceptual example of the problem, not a recorded prediction of the adapted checkpoint. Candidate names are illustrative.

The retriever

Two encoders, one space

Score: dot product of the two encodings

s(q,e)=Eq(q)⊤Ee(e)

Training: contrastive loss with in-batch and hard negatives

ℒ=−loges(q,e+)/τes(q,e+)/τ+∑e−∈Nes(q,e−)/τ

Concepts are encoded once, offline, into an index (GoldenRetriever); at query time only the question is encoded, and the top concepts are a nearest-neighbour lookup away.

ConditionalUnconditionalScholarshipsSchoolDeworming● query · ● positive concept · ● negatives

Query: “cash for families if children attend school”

Conditional cash transfers
0%
Unconditional cash transfers
0%
Scholarships & fee support
0%
School feeding
50%
Deworming
50%

s(q, e⁺) = -0.77 · loss ℒ = 16.014

Softmax over similarities with τ = 0.1, computed live. Training pulls the question towards its concept and pushes the others away, so the correct concept wins even against close siblings like conditional versus unconditional transfers. Toy 2-D vectors; the real ones have 1,024 dimensions.

The reader

Find the mention, choose the concept

One pass: the question and all candidates together

[CLS]Whatworksforcashforfamiliesifchildrenattendschool?[SEP]Conditional cash transfers[SEP]Unconditional cash transfers[SEP]Scholarships[SEP]
Whatworksforcashforfamiliesifchildrenattendschool?■ span start■ span end“cash … school” → Conditional cash transfers

Illustrative probabilities. DeBERTa-v3-large reads everything at once, finds the mention and picks its concept.

Training data

Examples written by an LLM, checked by experts

There was no labelled data for this taxonomy. We used an LLM to write, for each concept, the ways people phrase it and the near misses that belong to a sibling, then trained the encoders on those pairs. Examples below are illustrative.

Taxonomy concept

Conditional cash transfers

Cash paid to households on condition of a behaviour, such as school attendance or health check-ups.

siblings: unconditional transfers · scholarships · school feeding

LLMwrites examples per concept, and against its siblings

How users say it

“cash if kids go to school”“CCT for enrolment”“paying families for attendance”

Near misses (a sibling concept)

“cash with no strings attached” → Unconditional“school fees waived” → Scholarships & fee support

In context

abstract sentences with the mention marked
→ (question, positive concept, hard negatives) pairs for the retriever, and marked spans for the reader. Experts spot-check samples before training.

More encoders

Same recipe, other jobs

Relevance classifier

encoder · (question, estimate) → relevant?

An earlier generation of ImpactAI’s relevance layer: one encoder pass per candidate estimate instead of an LLM call. Later replaced by a configurable Gemini judge.

Effect-size regressor

ModernBERT-large · regression head

Predicts an effect size and both interval bounds from a question or a synthetic trial, trained with mean squared error: 27% lower error than GPT-5.2 in-domain.

The Query2Effect story →

ReLiK and GoldenRetriever are by SapienzaNLP. Training dataset, split and final scores of the adapted linker are still being gathered, so none are shown.