Sapienza NLP · Babelscape · The World Bank · 2020 – present
Encoder–decoder systems: from meaning graphs to tuned Gemini
Sequence-to-sequence models that write structured output. First, AMR parsers where each paper changed one part of the transformer: the output (CLAP), the attention (aligner), the encoder (LeakDistill), the input (AMRs Assemble!). Then generative models at scale: Gemini fine-tuned on Vertex AI, LLM-built datasets and LLM components in ImpactAI.
What’s inside
Part 1 · Meaning graphs
One transformer, four changes that each paper made
Semantic parsers are encoder–decoders fine-tuned to write a graph (AMR) as a sequence of tokens. Below is that model at the scale we trained it. Pick a paper to see which part it changes: what the decoder writes (CLAP), what the attention learns (the aligner), what the encoder sees during training (LeakDistill), or what goes in when several parsers have already answered (AMRs Assemble!).
The sentence goes through 12 encoder layers once. The decoder then writes the graph one token at a time; in every layer, cross-attention reads the encoder states H to decide the next token.
Input
Output · 0/25 tokens
Size
Training objective
teacher forcing on the linearised gold graph; beam search at inference.
Graph as text
<pointer:4>, which is why it needs thousands of extra vocabulary entries.Example and attention from Figure 3 of the aligner paper (decoder layer 3, head 6).
Part 2 · Generative systems
Beyond graphs: tuning Gemini and putting LLMs to work
The same habits (control the target format, build the training data, fine-tune, measure) carried over to large generative models. At the World Bank we fine-tuned Gemini on Vertex AI to write answers from statistics, and generative models run several steps of ImpactAI.
One training example (illustrative)
{"contents": [
{"role": "user", "parts": [{"text":
"Q: Do conditional cash transfers raise school enrolment?
Stats: 14 studies · pooled g = 0.12 [0.07, 0.17] · I² = 41%"}]},
{"role": "model", "parts": [{"text":
"Across 14 studies, CCTs raised enrolment by a small but
consistent amount (g = 0.12) …"}]}
]}Why tune instead of prompt
- Style and rigour: the answer must report the statistics it was given, with their uncertainty, and nothing else.
- Data we control: real records plus LLM-written examples, reviewed before training.
- Same cloud: data on Cloud Storage, tuning and serving on Vertex AI, next to the rest of ImpactAI’s answer step →
Pipeline as implemented in the ImpactAI summarizer code. The example and its numbers are made up for illustration.
Findings of ACL 2024
MSL: an LLM that learned from annotators
After a parser produced silver graphs, annotators corrected a sample; a fine-tuned LLM propagated those corrections and projected graphs to six more languages, for about $400.
The MSL story →arXiv 2026
Query2Effect: generate structure, then predict
An LLM first writes a structured version of a vague question; a fine-tuned encoder then predicts the effect size. Fine-tuning cuts absolute error by 27–71% against prompted LLMs.
The Query2Effect story →World Bank
LLM components inside ImpactAI
Extraction, linking, relevance judgement, taxonomy naming and the final answer all use generative models, each with its own prompt, schema and evaluation.
ImpactAI, phase by phase →Side project
Maths exercise generator
A teaching tool: prompt templates turn a teacher’s choices into an exercise, and a second call solves it step by step with rendered formulas.
Code on GitHub ↗