InteractiveFindings of ACL 2023
LeakDistill: let the parser peek at the graph, then take it away
Pavlo Vasylenko, Pere-Lluís Huguet Cabot*, Abelardo Carlos Martínez Lorenzo*, Roberto Navigli · Sapienza NLP & Babelscape · *equal contribution
A Transformer parser writes an AMR graph without ever seeing one. We showed the encoder the answer's structure during training, through graph adapters built from word alignments, and then distilled that knowledge into a model that needs no graph at all.
- SMATCH when the graph is leaked to the encoder (vs 84.6 without)
- 89.6SMATCH when the graph is leaked to the encoder (vs 84.6 without)
- SMATCH on AMR 2.0 with no extra data, the best single model at the time
- 85.7SMATCH on AMR 2.0 with no extra data, the best single model at the time
- graph needed at inference: the adapters are switched off
- 0graph needed at inference: the adapters are switched off
Parsing is writing a graph, blind
The input
Our example: Here, it is a country with freedom of speech.
The target
Its AMR graph: country is located here, is it, and is the thing that is free (free-04) with respect to speak-01.
Seq2seq parsing
Parsers such as SPRING fine-tune BART to translate the sentence into a linearised version of the graph. Colours show the alignment between words and graph units.
The catch: the encoder only ever sees words. Everything it learns about the graph's structure must be inferred indirectly, through the decoder's loss.
Turning the graph into something the encoder can read
Start from the AMR
The encoder has hidden states for words, not for concepts. To give it graph structure, we need a graph whose nodes are words.
1 · Replace labels with aligned words
Using the word-to-node alignment, every concept and relation takes the label of its word: free-04 → freedom, :ARG1-of → with, :domain → is, :ARG3 → of. But :location has no word in the sentence.
2 · Full WAG
Each edge becomes a node connected to both endpoints, and multi-token words are split into a parent and children: freedom → free + ##dom. The non-aligned location stays as a node with a fresh representation (red). This is the Full Word-Aligned Graph.
3 · Contracted WAG
The alternative: merge non-aligned nodes into their closest parent, so every node is a real token. This is the Contracted WAG: simpler, but it loses the edge label.
How the graph gets into the encoder
Why transform the graph at all? The encoder holds one hidden state per input token. AMR nodes such as free-04 or relations such as :ARG1-of have no state there. Once every node is a word, the WAG becomes a graph over the sentence itself: each edge simply says “these two tokens are related in the meaning of the sentence”.
What does an adapter do with it? Each encoder layer first runs self-attention, where every token reads every other token with learned weights. Then the structural adapter runs a graph convolution: each token is updated only from its WAG neighbours, using the matrix above, passed through GELU and added back to the original state. Press play and follow a token through all 12 layers, with and without the adapters.
Click a token to follow it. Colours are the 6 values of each hidden state (blue negative, pink positive).
The same computation, one node at a time. Click any node to see exactly which neighbours feed it and with what weight.
New state of free · dv = 4
| u ∈ N(v) | d_u | 1/√(d_u·d_v) |
|---|---|---|
| free (self) | 4 | 0.250 |
| with | 3 | 0.289 |
| of | 3 | 0.289 |
| ##dom | 2 | 0.354 |
Then the adapter adds it back
σ = GELU, no layer normalisation, residual connection (Figure 3).
The leak
Put the adapters in and feed them the WAG: this is the Graph Leakage Model. It is cheating on purpose. The WAG is built from the gold graph, so the model is shown the structure of the answer.
That makes it an upper bound: how much could graph structure help, if the encoder had it? With the Full WAG, about five SMATCH points. The Contracted WAG helps much less, so the labels of non-aligned nodes matter.
The problem, of course: at test time there is no gold graph to leak.
Show the values as a table
| SMATCH | |
|---|---|
| SPRING (ours)no graph | 84.5 |
| Contracted WAGgraph leaked | 86.0 |
| Full WAGgraph leaked | 89.6 |
Distilling the leak
First try: knowledge distillation. Use the leaky model as a teacher and a plain BART parser as the student. It does not work: the student scores 83.90, below the baseline.
LeakDistill: self-distillation. One model, two forward passes per training step. The green pass sees the WAG through the adapters; the red pass sees only the sentence. A KL divergence forces both passes to predict the same distribution, so the red pass absorbs what the green pass learned from the graph.
At inference, only the red pass runs. No graph, no adapters, no extra cost.
iteration
0k / 21k
β (leak loss)
90 (90 → 10)
α (KL)
20
KL(p ‖ q)
0.427
Wiring, loss terms, α = 20 and the β schedule follow the paper (Eq. 5, Section 6.3). The two next-token distributions and how fast they converge are illustrative.
The loss
Parsing loss of the red pass, leak loss of the green pass, and the KL term between them. β starts high, since early in training there is little to distill, and decreases.
Show the values as a table
| SMATCH | |
|---|---|
| SPRING (ours)baseline | 84.5 |
| KDteacher: Full WAG model | 83.9 |
| L_leak + L_nllLeakDistill | 84.5 |
| L_leak + L_KLLeakDistill | 85.0 |
| L_leak + L_nll + L_KLLeakDistill | 85.0 |
Show the values as a table
| SMATCH | |
|---|---|
| Contracted WAG | 84.9 |
| Full WAG | 85.0 |
| + β scheduling | 85.1 |
| + silver data | 85.3 |
| + silver + β | 85.3 |
Results
On the test sets, LeakDistill is the best single-model parser, and it gets there without extra data. Previous systems relied on up to 200K silver graphs; adding 140K silver pairs to LeakDistill helps only a little.
It is weaker than ATP on reentrancies, negation and especially semantic role labelling, which ATP trains on explicitly. The alignments LeakDistill uses are not counted as extra data, since they are derived from the training set itself.
Show the values as a table
| SMATCH | |
|---|---|
| SPRINGno extra data | 83.0 |
| SPRING (ours)no extra data | 83.8 |
| BiBLno extra data | 83.9 |
| Ancestorno extra data | 83.5 |
| LeakDistillno extra data | 84.5 |
| ATP+40K silver | 83.9 |
| AMRBART+200K silver | 84.2 |
| LeakDistill+140K silver | 84.6 |
Show the values as a table
| SMATCH | |
|---|---|
| SPRING | 61.6 |
| BiBL | 61.1 |
| ATP | 61.2 |
| AMRBART | 63.4 |
| LeakDistill | 64.5 |
+1.1
SMATCH over AMRBART on BioAMR, the hardest out-of-domain set.
83.5
on AMR 3.0 with BART-base: half a point above SPRING-large, with 140M parameters instead of 400M.
85.9
WWLK on AMR 3.0, the best of all systems on a metric that weighs edge labels.
Limitations and what's next
- Only tested on AMR parsing; tasks with input–output alignments, such as relation extraction, are a natural next step.
- Like other seq2seq parsers, performance drops as sentences get longer.
- Training needs alignments and two forward passes, which adds complexity (inference does not).
Cite
@inproceedings{vasylenko-etal-2023-incorporating,
title = "Incorporating Graph Information in Transformer-based {AMR} Parsing",
author = "Vasylenko, Pavlo and
Huguet Cabot, Pere Llu{\'i}s and
Mart{\'i}nez Lorenzo, Abelardo Carlos and
Navigli, Roberto",
booktitle = "Findings of the Association for Computational Linguistics: ACL 2023",
month = jul,
year = "2023",
address = "Toronto, Canada",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2023.findings-acl.125/",
doi = "10.18653/v1/2023.findings-acl.125",
pages = "1995--2011"
}Graphs redrawn from Figures 1–2; scores copied from the paper's tables.