Skip to content
Abelardo Carlos

InteractiveFindings of ACL 2023

LeakDistill: let the parser peek at the graph, then take it away

Pavlo Vasylenko, Pere-Lluís Huguet Cabot*, Abelardo Carlos Martínez Lorenzo*, Roberto Navigli · Sapienza NLP & Babelscape · *equal contribution

A Transformer parser writes an AMR graph without ever seeing one. We showed the encoder the answer's structure during training, through graph adapters built from word alignments, and then distilled that knowledge into a model that needs no graph at all.

SMATCH when the graph is leaked to the encoder (vs 84.6 without)
89.6SMATCH when the graph is leaked to the encoder (vs 84.6 without)
SMATCH on AMR 2.0 with no extra data, the best single model at the time
85.7SMATCH on AMR 2.0 with no extra data, the best single model at the time
graph needed at inference: the adapters are switched off
0graph needed at inference: the adapters are switched off
01

Parsing is writing a graph, blind

Input sentence
Here,itisacountrywithfreedomofspeech:ARG1-of:location:domain:ARG3countrywithlocationisfree-04hereitof##domspeak-01

The input

Our example: Here, it is a country with freedom of speech.

The target

Its AMR graph: country is located here, is it, and is the thing that is free (free-04) with respect to speak-01.

Seq2seq parsing

Parsers such as SPRING fine-tune BART to translate the sentence into a linearised version of the graph. Colours show the alignment between words and graph units.

The catch: the encoder only ever sees words. Everything it learns about the graph's structure must be inferred indirectly, through the decoder's loss.

02

Turning the graph into something the encoder can read

AMR graph
Here,itisacountrywithfreedomofspeech:ARG1-of:location:domain:ARG3countrywithlocationisfree-04hereitof##domspeak-01
conceptPropBank frame

Start from the AMR

The encoder has hidden states for words, not for concepts. To give it graph structure, we need a graph whose nodes are words.

1 · Replace labels with aligned words

Using the word-to-node alignment, every concept and relation takes the label of its word: free-04 → freedom, :ARG1-of → with, :domain → is, :ARG3 → of. But :location has no word in the sentence.

2 · Full WAG

Each edge becomes a node connected to both endpoints, and multi-token words are split into a parent and children: freedom → free + ##dom. The non-aligned location stays as a node with a fresh representation (red). This is the Full Word-Aligned Graph.

3 · Contracted WAG

The alternative: merge non-aligned nodes into their closest parent, so every node is a real token. This is the Contracted WAG: simpler, but it loses the edge label.

03

How the graph gets into the encoder

Why transform the graph at all? The encoder holds one hidden state per input token. AMR nodes such as free-04 or relations such as :ARG1-of have no state there. Once every node is a word, the WAG becomes a graph over the sentence itself: each edge simply says “these two tokens are related in the meaning of the sentence”.

The graph, redrawn on the sentence
Here,itisacountrywithfree##domofspeech.location*Arcs = WAG edges. Tokens with no arc (“,”, “a”, “.”) only keep their own state.
HereHere,,ititisisaacountrycountrywithwithfreefree##dom##domofofspeechspeech..location*location*
13×13 matrix · 9 edges · hover a cell
Left: every WAG edge connects two tokens of the input (or the extra location* node). Right: the matrix the adapter multiplies hidden states by, Â[v][u] = 1/√(d_u·d_v). Values are exact for this sentence.

What does an adapter do with it? Each encoder layer first runs self-attention, where every token reads every other token with learned weights. Then the structural adapter runs a graph convolution: each token is updated only from its WAG neighbours, using the matrix above, passed through GELU and added back to the original state. Press play and follow a token through all 12 layers, with and without the adapters.

Inside the encoder, one layer at a time
Layer 1a · self-attention: “country” reads every tokenHere15%,10%it5%is8%a9%countrywith7%free7%##dom6%of7%speech6%.12%location*

Click a token to follow it. Colours are the 6 values of each hidden state (blue negative, pink positive).

Toy model: 6-dimensional states and random weights, so the numbers are illustrative. The operations are the real ones: self-attention, then GraphConv over the WAG with the coefficients shown, GELU, and the residual connection (Eq. 3–4, Algorithm 1). Both passes share every weight; only the adapters differ.

The same computation, one node at a time. Click any node to see exactly which neighbours feed it and with what weight.

GraphConv on the WAG click a node
Here,itisacountrywithfreedomofspeech0.290.290.35countrywithlocationisfreehereitof##domspeech
GraphConvl(hvl,E)=∑u∈N(v)1dudvWglhul

New state of free · dv = 4

u ∈ N(v)d_u1/√(d_u·d_v)
free (self)40.250
with30.289
of30.289
##dom20.354

Then the adapter adds it back

zvl=Walσ(gvl)+hvl

σ = GELU, no layer normalisation, residual connection (Figure 3).

Degrees and coefficients are computed from this graph, with each node counted as its own neighbour (as in N(v) above).
04

The leak

Put the adapters in and feed them the WAG: this is the Graph Leakage Model. It is cheating on purpose. The WAG is built from the gold graph, so the model is shown the structure of the answer.

That makes it an upper bound: how much could graph structure help, if the encoder had it? With the Full WAG, about five SMATCH points. The Contracted WAG helps much less, so the labels of non-aligned nodes matter.

The problem, of course: at test time there is no gold graph to leak.

Graph Leakage Model (Table 1, AMR 3.0 dev) SMATCH, one axis
SPRING (ours)no graph8391SMATCH · SPRING (ours): 84.584.5Contracted WAGgraph leaked8391SMATCH · Contracted WAG: 86.086.0Full WAGgraph leaked8391SMATCH · Full WAG: 89.689.6
Show the values as a table
SMATCH
SPRING (ours)no graph84.5
Contracted WAGgraph leaked86.0
Full WAGgraph leaked89.6
05

Distilling the leak

First try: knowledge distillation. Use the leaky model as a teacher and a plain BART parser as the student. It does not work: the student scores 83.90, below the baseline.

LeakDistill: self-distillation. One model, two forward passes per training step. The green pass sees the WAG through the adapters; the red pass sees only the sentence. A KL divergence forces both passes to predict the same distribution, so the red pass absorbs what the green pass learned from the graph.

At inference, only the red pass runs. No graph, no adapters, no extra cost.

Green pass: the WAG enters every layer through the adapters.
Here, it is a country with freedom of speech
GREEN PASS · SENTENCE + WAG (TRAINING ONLY)encoder · 12 × (self-attention + FFN)▮ adapter: GraphConv over WAGdecoder12 layersq(next token | …:ARG1-of ( )free-0478%freedom8%speak-016%here5%country3%β·ℒ_leakRED PASS · SENTENCE ONLYencoder · 12 × (self-attention + FFN)▯ adapter skippeddecoder12 layersp(next token | …:ARG1-of ( )free-0435%freedom23%speak-0117%here14%country11%ℒ_nllWAGfrom the gold graphθ · one set of weights, used by both passesα·KL(p ‖ q)0.427ℒ = ℒ_nll + βℒ_leak+ αℒ_KL

iteration

0k / 21k

β (leak loss)

90 (90 → 10)

α (KL)

20

KL(p ‖ q)

0.427

Wiring, loss terms, α = 20 and the β schedule follow the paper (Eq. 5, Section 6.3). The two next-token distributions and how fast they converge are illustrative.

The loss

Parsing loss of the red pass, leak loss of the green pass, and the KL term between them. β starts high, since early in training there is little to distill, and decreases.

ℒLeakDistill=ℒnllD+βℒleak+αℒKL,ℒKL=∑k=0C−1pklogpkqk
Loss weights during training Section 6.3
0501000k7k14k21kβ (leak loss) 90 → 10α (KL) = 20weight · x-axis: training iterations
KD versus LeakDistill (Table 2, AMR 3.0 dev)
SPRING (ours)baseline8287SMATCH · SPRING (ours): 84.584.5KDteacher: Full WAG model8287SMATCH · KD: 83.983.9L_leak + L_nllLeakDistill8287SMATCH · L_leak + L_nll: 84.584.5L_leak + L_KLLeakDistill8287SMATCH · L_leak + L_KL: 85.085.0L_leak + L_nll + L_KLLeakDistill8287SMATCH · L_leak + L_nll + L_KL: 85.085.0
Show the values as a table
SMATCH
SPRING (ours)baseline84.5
KDteacher: Full WAG model83.9
L_leak + L_nllLeakDistill84.5
L_leak + L_KLLeakDistill85.0
L_leak + L_nll + L_KLLeakDistill85.0
LeakDistill variants (Table 5, AMR 3.0 dev)
Contracted WAG8387SMATCH · Contracted WAG: 84.984.9Full WAG8387SMATCH · Full WAG: 85.085.0+ β scheduling8387SMATCH · + β scheduling: 85.185.1+ silver data8387SMATCH · + silver data: 85.385.3+ silver + β8387SMATCH · + silver + β: 85.385.3
Show the values as a table
SMATCH
Contracted WAG84.9
Full WAG85.0
+ β scheduling85.1
+ silver data85.3
+ silver + β85.3
With the adapters switched on (the green pass), the same model reaches 86.09, still well below the 89.58 of the leakage model.
06

Results

On the test sets, LeakDistill is the best single-model parser, and it gets there without extra data. Previous systems relied on up to 200K silver graphs; adding 140K silver pairs to LeakDistill helps only a little.

It is weaker than ATP on reentrancies, negation and especially semantic role labelling, which ATP trains on explicitly. The alignments LeakDistill uses are not counted as extra data, since they are derived from the training set itself.

Test SMATCH (Tables 3 and 4)
SPRINGno extra data8286SMATCH · SPRING: 83.083.0SPRING (ours)no extra data8286SMATCH · SPRING (ours): 83.883.8BiBLno extra data8286SMATCH · BiBL: 83.983.9Ancestorno extra data8286SMATCH · Ancestor: 83.583.5LeakDistillno extra data8286SMATCH · LeakDistill: 84.584.5ATP+40K silver8286SMATCH · ATP: 83.983.9AMRBART+200K silver8286SMATCH · AMRBART: 84.284.2LeakDistill+140K silver8286SMATCH · LeakDistill: 84.684.6
Show the values as a table
SMATCH
SPRINGno extra data83.0
SPRING (ours)no extra data83.8
BiBLno extra data83.9
Ancestorno extra data83.5
LeakDistillno extra data84.5
ATP+40K silver83.9
AMRBART+200K silver84.2
LeakDistill+140K silver84.6
One axis for all rows. Hints give the extra silver data each system uses.
Out of distribution (Table 6)
SPRING6066SMATCH · SPRING: 61.661.6BiBL6066SMATCH · BiBL: 61.161.1ATP6066SMATCH · ATP: 61.261.2AMRBART6066SMATCH · AMRBART: 63.463.4LeakDistill6066SMATCH · LeakDistill: 64.564.5
Show the values as a table
SMATCH
SPRING61.6
BiBL61.1
ATP61.2
AMRBART63.4
LeakDistill64.5

+1.1

SMATCH over AMRBART on BioAMR, the hardest out-of-domain set.

83.5

on AMR 3.0 with BART-base: half a point above SPRING-large, with 140M parameters instead of 400M.

85.9

WWLK on AMR 3.0, the best of all systems on a metric that weighs edge labels.

Limitations and what's next

  • Only tested on AMR parsing; tasks with input–output alignments, such as relation extraction, are a natural next step.
  • Like other seq2seq parsers, performance drops as sentences get longer.
  • Training needs alignments and two forward passes, which adds complexity (inference does not).
Where do alignments come from? Read the cross-attention aligner story →

Cite

@inproceedings{vasylenko-etal-2023-incorporating,
    title = "Incorporating Graph Information in Transformer-based {AMR} Parsing",
    author = "Vasylenko, Pavlo  and
      Huguet Cabot, Pere Llu{\'i}s  and
      Mart{\'i}nez Lorenzo, Abelardo Carlos  and
      Navigli, Roberto",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2023",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.findings-acl.125/",
    doi = "10.18653/v1/2023.findings-acl.125",
    pages = "1995--2011"
}
All publications

Graphs redrawn from Figures 1–2; scores copied from the paper's tables.