InteractiveFindings of ACL 2023
Cross-lingual AMR Aligner: paying attention to cross-attention
Abelardo Carlos Martínez Lorenzo*, Pere-Lluís Huguet Cabot*, Roberto Navigli · Babelscape & Sapienza NLP · *equal contribution
Semantic parsers turn sentences into graphs of meaning. To use those graphs, you also need to know which words each part came from. We found that the parser already knows: it is written in its cross-attention, in any language.
which words does each concept come from?
- correlation of a single untrained head with gold alignment
- 0.635correlation of a single untrained head with gold alignment
- after guiding layer 3, with no loss in parsing quality
- 0.866after guiding layer 3, with no loss in parsing quality
- F1 on average over the next best aligner in DE, ES and IT
- +40F1 on average over the next best aligner in DE, ES and IT
- language-specific rules
- 0language-specific rules
A sentence becomes a graph
scroll to parse ↓
Itisthetimeyouhave wastedforyourrosethatmakesyourrosesoimportant
The input
It is the time you have wasted for your rose that makes your rose so important.
Abstract Meaning Representation
time, waste-01, rose), edges are relations (:ARG0, :beneficiary). A concept mentioned twice, like you or rose, stays one node with extra incoming edges: a reentrancy (dashed).The missing link
time is easy, but waste-01 comes from “have wasted”, and the second “rose” is not a new node: it is the dashed edge. Finding these links is alignment.Alignment
Cross-lingual AMR
Why it matters
Alignments are what downstream applications build on.
Training parsers
Many parsers learn from sentence–graph pairs where each node is tied to its words.
Crossing languages
Alignments carry an English graph onto a sentence in another language, to build data where none exists.
New formalisms and datasets
Projecting graphs to new sentences and languages, as in the MSL dataset, starts from aligned words.
Explaining the output
Knowing which words produced each node lets you check, debug and trust a parse.
How aligners used to work
Itisthetimeyouhave wastedforyourrosethatmakesyourrosesoimportant
✓ rule found a word✕ no match
Rules and statistics
Repetition
rose node plus a reentrant edge. Which mention is which? String rules cannot tell; they need extra heuristics.Spanish
German
The idea: the parser is already looking
State-of-the-art AMR parsers such as SPRING are encoder-decoder Transformers built on BART. The encoder reads the sentence; the decoder writes a linearised graph, one token at a time.
At every step, cross-attention decides how much each input token matters to the token being written: att(Q, K) = softmax(QKᵀ / √dₖ), computed in every layer ℓ and head h.
Hypothesis
When the decoder writes a concept, it should be looking at the words that express it.
If that holds, alignment is a by-product of parsing: no rules, no extra model, and nothing specific to English.
The encoder reads the whole sentence at once. Self-attention lets every word look at every other word, in 12 layers.
Sentence, linearised graph and attention weights from Figure 3 of the paper (SPRING, decoder layer 3, head 6). Encoder self-attention arcs are schematic.
The animation shows one answer per concept. Below, step through every token the decoder writes, relations and pointers included, and watch the same head at work.
Encoder · 12 layers
reads the sentence once
Decoder · 12 layers × 16 heads
writes the graph, token by token
Explore: click any graph token to see where it looks, or click a sentence word to see which graph tokens look at it.
Two things stand out. Without ever seeing an alignment, the head sends both person and merchandise-01 to “merchant”, a match no string rule would make. And pointer tokens such as <pointer:4>, which only refer back to a node already written, carry little meaning: their attention spreads thinly, often towards the final full stop.
Which heads know?
Layer 3, head 6 is not special by accident. A 12-layer decoder with 16 heads per layer has 192 cross-attention maps. For each one, we flattened its matrix and computed Pearson’s r against the gold LEAMR alignment matrix on the validation set. The correlation is clearly positive, and it is not spread evenly: some heads track alignment closely, others hardly at all. We don’t have a full explanation for why, but layer 3 stands out.
Below, the decoder is opened up. Pick a head to see what it reads on the example sentence. Only layer 3, head 6 is the real map from the paper; the others are simulated from their measured r, to show what a high or a low correlation looks like.
Click any head. Rows are layers (bottom = first), columns are heads.
layer 3 · head 6
r ≈ 0.63
real map (Figure 3)Each row is one graph token; its query meets the 16 encoder keys: softmax(q·kᵀ/√d). Dark = where the head looks.
◯ gold alignment. High-r heads put their darkest cell on the circle; low-r heads spread out or drift to the full stop.
Guiding the heads
If some heads already behave like aligners, can we help them? We retrain SPRING with the same hyperparameters and supervise half the heads of layer 3 with silver alignments from LEAMR or ISI. Sparse attention variants did not work, so instead the heads are combined with a learnable scalar mix, as in ELMo, and an extra cross-entropy term joins the parsing loss.
Illustrative training dynamics on the Figure 3 sentence. The loss is the cross-entropy between the layer-3 attention mix and the silver alignment (Section 3.2); the parsing loss is trained at the same time.
Now try it yourself: move the slider to train the heads on a second example and watch the mix weights change.
1 · Mix the heads of one layer
Each head h gets a weight s = softmax(a), with γ and a learned. The model is free to make some heads sparse, like an alignment, while others keep doing other work.
2 · Add an alignment term to the loss
The usual parsing loss, plus a cross-entropy that pushes the mixed attention of each graph token towards the words it is aligned to.
Rows: graph concepts. Columns: subword tokens. Dashed cells: the reference alignment (Table 5, English).
Alignment loss
1.017
Head weights s in layer 3
supervised heads free heads
Illustrative numbers, built by hand to show the mechanism. Real results are in the heatmap above and the tables below.
From attention to alignment
Attention gives a score for every (graph token, sentence token) pair. Six steps turn those scores into alignments. For the final results, the score matrix is the sum of cross-attention over the first four layers.
- 1
Score matrix
Build M ∈ ℝ^(n_d × n_e) from cross-attention weights, one row per graph token and one column per subword.
- 2
Span segmentation
Sum the columns of subwords that form one word (under + stand → understand), and group words into spans.
- 3
Graph segmentation
Sum the rows of tokens that form one semantic unit (<pointer:1> understand -01 → understand-01).
- 4
Map units to spans
Give every graph unit the sentence span with the highest score.
- 5
Special structures
Revise literal subgraphs such as named entities and dates so they align as a block.
- 6
Format
Write the mapping out in the target standard: ISI or LEAMR.
Across languages
For other languages, the same recipe runs on SPRING trained with mBART. We paired the human German, Spanish and Italian translations of the AMR test set with their English graphs. The aligner needs no new rules: it simply reads the attention.
Aligner
hard-02
understand-01
it
cause-01
amr-unknown
so
On this sentence: 4 of 5 reference links recovered, 2 extra.
The Spanish translation asks what is hard, making “Qué” the subject. Our aligner links “it” to “Qué”, which is reasonable but counts as an error against the projected reference.
EN
DE
ES
IT
Non-English references are projected from English with a machine-translation aligner. The guided model was only evaluated in English.
Cross-attention scores 69–72 F1 on German, Spanish and Italian, about 40 points above the next best system on average. It is still lower than English, for two reasons: real linguistic divergence between languages, and errors in the machine-translation projection used to build the non-English references.
Results in English
On the LEAMR benchmark, both variants beat LEAMR, the previous state of the art, on subgraphs and relations. The unguided aligner, which never saw an alignment, is already competitive. Guiding layer 3 gives the best overall scores and does not hurt parsing quality.
Reentrancies remain the hardest case for everyone. On the ISI standard, the guided aligner reaches 95.2 F1 in English.
Subgraph · 1,707 alignments
Relation · 1,263 alignments
Reentrancy · 293 alignments
Duplicate subgraph · 17 alignments
Limitations
- It needs a Transformer parser, and does not transfer to other architectures.
- It assumes the decoder attends to the words most relevant to the next token, which does not always hold.
- Non-English scores are lower, partly because of linguistic divergence and partly because of projected references.
- LEAMR-format output needs an English-specific span segmentation, so cross-lingual results use the ISI format.
Cite
@inproceedings{martinez-lorenzo-etal-2023-cross,
title = "Cross-lingual {AMR} Aligner: Paying Attention to Cross-Attention",
author = "Mart{\'i}nez Lorenzo, Abelardo Carlos and
Huguet Cabot, Pere Llu{\'i}s and
Navigli, Roberto",
booktitle = "Findings of the Association for Computational Linguistics: ACL 2023",
month = jul,
year = "2023",
address = "Toronto, Canada",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2023.findings-acl.109/",
doi = "10.18653/v1/2023.findings-acl.109",
pages = "1726--1742"
}More interactive papers are on the way.