Skip to content
Abelardo Carlos
← Projects

Sapienza NLP · Babelscape · The World Bank · 2020 – present

Encoder–decoder systems: from meaning graphs to tuned Gemini

Sequence-to-sequence models that write structured output. First, AMR parsers where each paper changed one part of the transformer: the output (CLAP), the attention (aligner), the encoder (LeakDistill), the input (AMRs Assemble!). Then generative models at scale: Gemini fine-tuned on Vertex AI, LLM-built datasets and LLM components in ImpactAI.

Part 1 · Meaning graphs

One transformer, four changes that each paper made

Semantic parsers are encoder–decoders fine-tuned to write a graph (AMR) as a sequence of tokens. Below is that model at the scale we trained it. Pick a paper to see which part it changes: what the decoder writes (CLAP), what the attention learns (the aligner), what the encoder sees during training (LeakDistill), or what goes in when several parsers have already answered (AMRs Assemble!).

SPRING baselineReference model: nothing changed yet.

The sentence goes through 12 encoder layers once. The decoder then writes the graph one token at a time; in every layer, cross-attention reads the encoder states H to decide the next token.

ENCODER · 12 LAYERSDECODER · 12 LAYERS0self-attentionFFN123456789101116 input tokens → token + position embeddingsthe sentenceHencoder statesThiswasamerchantwhosoldpillsthathadbeeninventedtoquenchthirst.keys, values0masked self-attncross-attentionFFN1234567891011<s> + 0 tokens written so far (shifted right)linear + softmax over |V| (+3,000 pointer & relation tokens)fed back as the next decoder input

Input

Thiswasamerchantwhosoldpillsthathadbeeninventedtoquenchthirst.

Output · 0/25 tokens

Size

BART-large: 12 encoder and 12 decoder layers, 16 attention heads per layer, hidden size 1,024, about 400M parameters.

Training objective

ℒ_nll = − Σₜ log p(yₜ | y₍<ₜ₎, x)
teacher forcing on the linearised gold graph; beam search at inference.

Graph as text

SPRING writes the graph depth-first. Re-entrant nodes are pointer tokens like <pointer:4>, which is why it needs thousands of extra vocabulary entries.

Example and attention from Figure 3 of the aligner paper (decoder layer 3, head 6).

Part 2 · Generative systems

Beyond graphs: tuning Gemini and putting LLMs to work

The same habits (control the target format, build the training data, fine-tune, measure) carried over to large generative models. At the World Bank we fine-tuned Gemini on Vertex AI to write answers from statistics, and generative models run several steps of ImpactAI.

▤Records

question · statistics · answer, from ImpactAI

✦Synthesise

Gemini writes and refines more examples

⫽Split

train / validation

{ }JSONL → GCS

one chat example per line, on Cloud Storage

⚙Vertex AI tuning

supervised fine-tuning job on Gemini

◆Tuned model

answers grounded in the statistics it is given

One training example (illustrative)

{"contents": [
  {"role": "user", "parts": [{"text":
    "Q: Do conditional cash transfers raise school enrolment?
     Stats: 14 studies · pooled g = 0.12 [0.07, 0.17] · I² = 41%"}]},
  {"role": "model", "parts": [{"text":
    "Across 14 studies, CCTs raised enrolment by a small but
     consistent amount (g = 0.12) …"}]}
]}

Why tune instead of prompt

  • Style and rigour: the answer must report the statistics it was given, with their uncertainty, and nothing else.
  • Data we control: real records plus LLM-written examples, reviewed before training.
  • Same cloud: data on Cloud Storage, tuning and serving on Vertex AI, next to the rest of ImpactAI’s answer step →

Pipeline as implemented in the ImpactAI summarizer code. The example and its numbers are made up for illustration.

silver graph + annotator fixes→fine-tuned GPT-3.5
propagaterewrite silver graphs like the annotators
projectwrite the graph for a new language
result168,315 high-quality graphs

Findings of ACL 2024

MSL: an LLM that learned from annotators

After a parser produced silver graphs, annotators corrected a sample; a fine-tuned LLM propagated those corrections and projected graphs to six more languages, for about $400.

The MSL story →
“does free school food help kids learn?”→LLM + encoder
interventionschool feeding
outcomelearning / test scores
populationprimary-school children

arXiv 2026

Query2Effect: generate structure, then predict

An LLM first writes a structured version of a vague question; a fine-tuned encoder then predicts the effect size. Fine-tuning cuts absolute error by 27–71% against prompted LLMs.

The Query2Effect story →
a policy question or a paper passage→Gemini on Vertex AI
entitiesNER and linking to the taxonomy
relationswho, where, for whom, which effect
relevancedoes this study answer the question?

World Bank

LLM components inside ImpactAI

Extraction, linking, relevance judgement, taxonomy naming and the final answer all use generative models, each with its own prompt, schema and evaluation.

ImpactAI, phase by phase →
topic · difficulty · student profile→LLM, two calls
statementan exercise for that student
solutionworked step by step, same context
deliveryStreamlit app on Cloud Run

Side project

Maths exercise generator

A teaching tool: prompt templates turn a teacher’s choices into an exercise, and a second call solves it step by step with rendered formulas.

Code on GitHub ↗