Skip to content
Abelardo Carlos

InteractiveFindings of ACL 2023

Cross-lingual AMR Aligner: paying attention to cross-attention

Abelardo Carlos Martínez Lorenzo*, Pere-Lluís Huguet Cabot*, Roberto Navigli · Babelscape & Sapienza NLP · *equal contribution

Semantic parsers turn sentences into graphs of meaning. To use those graphs, you also need to know which words each part came from. We found that the parser already knows: it is written in its cross-attention, in any language.

:ARG1-of:ARG0-of:ARG0:beneficiary:ARG1:degree:poss:ARG1Itisthetimeyouhave wastedforyourrosethatmakesyourrosesoimportanttime?waste-01?make-02?you?rose?important-01?so?

which words does each concept come from?

correlation of a single untrained head with gold alignment
0.635correlation of a single untrained head with gold alignment
after guiding layer 3, with no loss in parsing quality
0.866after guiding layer 3, with no loss in parsing quality
F1 on average over the next best aligner in DE, ES and IT
+40F1 on average over the next best aligner in DE, ES and IT
language-specific rules
0language-specific rules
01

A sentence becomes a graph

Input sentence

scroll to parse ↓

:ARG1-of:ARG0-of:ARG0:beneficiary:ARG1:degree:poss:ARG1p / timew / waste-01m / make-02y / your / rosei / important-01s / so

Itisthetimeyouhave wastedforyourrosethatmakesyourrosesoimportant

The input

Start with a sentence. This one comes from The Little Prince: It is the time you have wasted for your rose that makes your rose so important.

Abstract Meaning Representation

AMR rewrites the sentence as a graph. Nodes are concepts (time, waste-01, rose), edges are relations (:ARG0, :beneficiary). A concept mentioned twice, like you or rose, stays one node with extra incoming edges: a reentrancy (dashed).

The missing link

The graph has no words in it. So which word does each node come from? time is easy, but waste-01 comes from “have wasted”, and the second “rose” is not a new node: it is the dashed edge. Finding these links is alignment.

Alignment

An alignment says which words each part of the graph comes from. Hover a word or a node to see the pairs.

Cross-lingual AMR

Cross-lingual AMR keeps the English graph and pairs it with a sentence in another language. Same meaning, different words, different order. Switch the language above the graph and watch the colours move.

Why it matters

Alignments are what downstream applications build on.

Training parsers

Many parsers learn from sentence–graph pairs where each node is tied to its words.

Crossing languages

Alignments carry an English graph onto a sentence in another language, to build data where none exists.

New formalisms and datasets

Projecting graphs to new sentences and languages, as in the MSL dataset, starts from aligned words.

Explaining the output

Knowing which words produced each node lets you check, debug and trust a parse.

02

How aligners used to work

AMR graph · Figure 1
:ARG1-of:ARG0-of:ARG0:beneficiary:ARG1:degree:poss:ARG1p / time✓w / waste-01✓m / make-02✓y / you✓r / rose✓i / important-01✓s / so✓

Itisthetimeyouhave wastedforyourrosethatmakesyourrosesoimportant

✓ rule found a word✕ no match

Rules and statistics

Earlier aligners (JAMR, TAMR, ISI, LEAMR) combine hand-written rules with statistics such as Expectation Maximization. JAMR, for instance, applies an ordered list of 14 criteria, including exact and fuzzy string matching. In English, most concepts find their word.

Repetition

But what about repetition? “rose” and “your” appear twice, while the graph has one rose node plus a reentrant edge. Which mention is which? String rules cannot tell; they need extra heuristics.

Spanish

And the rules are written for English. In Spanish, “tiempo” is not time and “perdiste” is not waste. Only look-alike words such as “importante” survive.

German

German fares even worse. On German, Spanish and Italian test sets, rule-based aligners score between 8 and 28 F1. They are, by construction, monolingual.The matcher in the figure is a toy rule for illustration.
03

The idea: the parser is already looking

State-of-the-art AMR parsers such as SPRING are encoder-decoder Transformers built on BART. The encoder reads the sentence; the decoder writes a linearised graph, one token at a time.

At every step, cross-attention decides how much each input token matters to the token being written: att(Q, K) = softmax(QKᵀ / √dₖ), computed in every layer ℓ and head h.

Hypothesis

When the decoder writes a concept, it should be looking at the words that express it.

If that holds, alignment is a by-product of parsing: no rules, no extra model, and nothing specific to English.

The encoder reads the whole sentence at once. Self-attention lets every word look at every other word, in 12 layers.

Thiswasamerchantwhosoldpillsthathadbeeninventedtoquenchthirst.ENCODER · SELF-ATTENTION × 12 LAYERSone vector per wordDECODER · WRITES ONE TOKEN AT A TIME · CROSS-ATTENTION READS THE ENCODER:ARG0:ARG1:ARG0-of:domain:ARG1-of:purpose:ARG1:ARG0sell-01personpillmerchandise-01thisinvent-01quench-01thirst-01

Sentence, linearised graph and attention weights from Figure 3 of the paper (SPRING, decoder layer 3, head 6). Encoder self-attention arcs are schematic.

The animation shows one answer per concept. Below, step through every token the decoder writes, relations and pointers included, and watch the same head at work.

Encoder · 12 layers

reads the sentence once

cross-attention

Decoder · 12 layers × 16 heads

writes the graph, token by token

01/25

Explore: click any graph token to see where it looks, or click a sentence word to see which graph tokens look at it.

<pointer:0>sell-01:ARG0<pointer:1>person:ARG0-of<pointer:2>merchandise-01:domain<pointer:3>this:ARG1<pointer:4>pill:ARG1-of<pointer:5>invent-01:purpose<pointer:6>quench-01:ARG0<pointer:4>:ARG1<pointer:7>thirst-01Thiswasamerchantwhosoldpillsthathadbeeninventedtoquenchthirst.attention weight

Two things stand out. Without ever seeing an alignment, the head sends both person and merchandise-01 to “merchant”, a match no string rule would make. And pointer tokens such as <pointer:4>, which only refer back to a node already written, carry little meaning: their attention spreads thinly, often towards the final full stop.

04

Which heads know?

Layer 3, head 6 is not special by accident. A 12-layer decoder with 16 heads per layer has 192 cross-attention maps. For each one, we flattened its matrix and computed Pearson’s r against the gold LEAMR alignment matrix on the validation set. The correlation is clearly positive, and it is not spread evenly: some heads track alignment closely, others hardly at all. We don’t have a full explanation for why, but layer 3 stands out.

Below, the decoder is opened up. Pick a head to see what it reads on the example sentence. Only layer 3, head 6 is the real map from the paper; the others are simulated from their measured r, to show what a high or a low correlation looks like.

Decoder, opened
colour = how well the head’s attention matches gold alignments (Pearson r)L0layer 0 · head 0 · r ≈ 0.40layer 0 · head 1 · r ≈ 0.09layer 0 · head 2 · r ≈ 0.02layer 0 · head 3 · r ≈ 0.18layer 0 · head 4 · r ≈ 0.09layer 0 · head 5 · r ≈ 0.37layer 0 · head 6 · r ≈ 0.18layer 0 · head 7 · r ≈ 0.06layer 0 · head 8 · r ≈ 0.24layer 0 · head 9 · r ≈ 0.02layer 0 · head 10 · r ≈ 0.03layer 0 · head 11 · r ≈ 0.07layer 0 · head 12 · r ≈ 0.11layer 0 · head 13 · r ≈ 0.10layer 0 · head 14 · r ≈ 0.04layer 0 · head 15 · r ≈ 0.43L1layer 1 · head 0 · r ≈ 0.41layer 1 · head 1 · r ≈ 0.26layer 1 · head 2 · r ≈ 0.40layer 1 · head 3 · r ≈ 0.05layer 1 · head 4 · r ≈ 0.18layer 1 · head 5 · r ≈ 0.34layer 1 · head 6 · r ≈ 0.21layer 1 · head 7 · r ≈ 0.43layer 1 · head 8 · r ≈ 0.49layer 1 · head 9 · r ≈ 0.28layer 1 · head 10 · r ≈ 0.16layer 1 · head 11 · r ≈ 0.37layer 1 · head 12 · r ≈ 0.35layer 1 · head 13 · r ≈ 0.16layer 1 · head 14 · r ≈ 0.53layer 1 · head 15 · r ≈ 0.49L2layer 2 · head 0 · r ≈ 0.46layer 2 · head 1 · r ≈ 0.16layer 2 · head 2 · r ≈ 0.42layer 2 · head 3 · r ≈ 0.42layer 2 · head 4 · r ≈ 0.42layer 2 · head 5 · r ≈ 0.36layer 2 · head 6 · r ≈ 0.35layer 2 · head 7 · r ≈ 0.20layer 2 · head 8 · r ≈ 0.05layer 2 · head 9 · r ≈ 0.41layer 2 · head 10 · r ≈ 0.23layer 2 · head 11 · r ≈ 0.32layer 2 · head 12 · r ≈ 0.26layer 2 · head 13 · r ≈ 0.33layer 2 · head 14 · r ≈ 0.44layer 2 · head 15 · r ≈ 0.21L3layer 3 · head 0 · r ≈ 0.31layer 3 · head 1 · r ≈ 0.30layer 3 · head 2 · r ≈ 0.24layer 3 · head 3 · r ≈ 0.50layer 3 · head 4 · r ≈ 0.32layer 3 · head 5 · r ≈ 0.58layer 3 · head 6 · r ≈ 0.63layer 3 · head 7 · r ≈ 0.44layer 3 · head 8 · r ≈ 0.05layer 3 · head 9 · r ≈ 0.26layer 3 · head 10 · r ≈ 0.07layer 3 · head 11 · r ≈ 0.42layer 3 · head 12 · r ≈ 0.44layer 3 · head 13 · r ≈ 0.11layer 3 · head 14 · r ≈ 0.28layer 3 · head 15 · r ≈ 0.44L4layer 4 · head 0 · r ≈ 0.25layer 4 · head 1 · r ≈ 0.52layer 4 · head 2 · r ≈ 0.41layer 4 · head 3 · r ≈ 0.26layer 4 · head 4 · r ≈ 0.38layer 4 · head 5 · r ≈ 0.17layer 4 · head 6 · r ≈ 0.22layer 4 · head 7 · r ≈ 0.13layer 4 · head 8 · r ≈ 0.28layer 4 · head 9 · r ≈ 0.02layer 4 · head 10 · r ≈ 0.12layer 4 · head 11 · r ≈ 0.04layer 4 · head 12 · r ≈ 0.22layer 4 · head 13 · r ≈ 0.22layer 4 · head 14 · r ≈ 0.32layer 4 · head 15 · r ≈ 0.41L5layer 5 · head 0 · r ≈ 0.17layer 5 · head 1 · r ≈ 0.23layer 5 · head 2 · r ≈ 0.14layer 5 · head 3 · r ≈ 0.04layer 5 · head 4 · r ≈ 0.28layer 5 · head 5 · r ≈ 0.33layer 5 · head 6 · r ≈ 0.24layer 5 · head 7 · r ≈ 0.21layer 5 · head 8 · r ≈ 0.60layer 5 · head 9 · r ≈ 0.33layer 5 · head 10 · r ≈ 0.39layer 5 · head 11 · r ≈ 0.26layer 5 · head 12 · r ≈ 0.11layer 5 · head 13 · r ≈ 0.20layer 5 · head 14 · r ≈ 0.15layer 5 · head 15 · r ≈ 0.12L6layer 6 · head 0 · r ≈ 0.05layer 6 · head 1 · r ≈ 0.10layer 6 · head 2 · r ≈ 0.02layer 6 · head 3 · r ≈ 0.05layer 6 · head 4 · r ≈ 0.03layer 6 · head 5 · r ≈ 0.43layer 6 · head 6 · r ≈ 0.21layer 6 · head 7 · r ≈ 0.49layer 6 · head 8 · r ≈ 0.21layer 6 · head 9 · r ≈ 0.03layer 6 · head 10 · r ≈ 0.04layer 6 · head 11 · r ≈ 0.02layer 6 · head 12 · r ≈ 0.49layer 6 · head 13 · r ≈ 0.29layer 6 · head 14 · r ≈ 0.14layer 6 · head 15 · r ≈ 0.32L7layer 7 · head 0 · r ≈ 0.48layer 7 · head 1 · r ≈ 0.49layer 7 · head 2 · r ≈ 0.60layer 7 · head 3 · r ≈ 0.19layer 7 · head 4 · r ≈ 0.17layer 7 · head 5 · r ≈ 0.27layer 7 · head 6 · r ≈ 0.17layer 7 · head 7 · r ≈ 0.48layer 7 · head 8 · r ≈ 0.50layer 7 · head 9 · r ≈ 0.27layer 7 · head 10 · r ≈ 0.28layer 7 · head 11 · r ≈ 0.32layer 7 · head 12 · r ≈ 0.17layer 7 · head 13 · r ≈ 0.31layer 7 · head 14 · r ≈ 0.21layer 7 · head 15 · r ≈ 0.44L8layer 8 · head 0 · r ≈ 0.53layer 8 · head 1 · r ≈ 0.31layer 8 · head 2 · r ≈ 0.40layer 8 · head 3 · r ≈ 0.33layer 8 · head 4 · r ≈ 0.48layer 8 · head 5 · r ≈ 0.25layer 8 · head 6 · r ≈ 0.43layer 8 · head 7 · r ≈ 0.40layer 8 · head 8 · r ≈ 0.51layer 8 · head 9 · r ≈ 0.40layer 8 · head 10 · r ≈ 0.54layer 8 · head 11 · r ≈ 0.29layer 8 · head 12 · r ≈ 0.18layer 8 · head 13 · r ≈ 0.34layer 8 · head 14 · r ≈ 0.38layer 8 · head 15 · r ≈ 0.23L9layer 9 · head 0 · r ≈ 0.20layer 9 · head 1 · r ≈ 0.34layer 9 · head 2 · r ≈ 0.36layer 9 · head 3 · r ≈ 0.46layer 9 · head 4 · r ≈ 0.37layer 9 · head 5 · r ≈ 0.33layer 9 · head 6 · r ≈ 0.35layer 9 · head 7 · r ≈ 0.26layer 9 · head 8 · r ≈ 0.38layer 9 · head 9 · r ≈ 0.28layer 9 · head 10 · r ≈ 0.42layer 9 · head 11 · r ≈ 0.40layer 9 · head 12 · r ≈ 0.31layer 9 · head 13 · r ≈ 0.19layer 9 · head 14 · r ≈ 0.39layer 9 · head 15 · r ≈ 0.38L10layer 10 · head 0 · r ≈ 0.27layer 10 · head 1 · r ≈ 0.40layer 10 · head 2 · r ≈ 0.21layer 10 · head 3 · r ≈ 0.31layer 10 · head 4 · r ≈ 0.32layer 10 · head 5 · r ≈ 0.36layer 10 · head 6 · r ≈ 0.31layer 10 · head 7 · r ≈ 0.29layer 10 · head 8 · r ≈ 0.20layer 10 · head 9 · r ≈ 0.10layer 10 · head 10 · r ≈ 0.41layer 10 · head 11 · r ≈ 0.34layer 10 · head 12 · r ≈ 0.35layer 10 · head 13 · r ≈ 0.41layer 10 · head 14 · r ≈ 0.18layer 10 · head 15 · r ≈ 0.27L11layer 11 · head 0 · r ≈ 0.14layer 11 · head 1 · r ≈ 0.13layer 11 · head 2 · r ≈ 0.16layer 11 · head 3 · r ≈ 0.15layer 11 · head 4 · r ≈ 0.29layer 11 · head 5 · r ≈ 0.26layer 11 · head 6 · r ≈ 0.28layer 11 · head 7 · r ≈ 0.22layer 11 · head 8 · r ≈ 0.28layer 11 · head 9 · r ≈ 0.21layer 11 · head 10 · r ≈ 0.11layer 11 · head 11 · r ≈ 0.21layer 11 · head 12 · r ≈ 0.25layer 11 · head 13 · r ≈ 0.15layer 11 · head 14 · r ≈ 0.30layer 11 · head 15 · r ≈ 0.270123456789101112131415

Click any head. Rows are layers (bottom = first), columns are heads.

layer 3 · head 6

r ≈ 0.63

real map (Figure 3)

Each row is one graph token; its query meets the 16 encoder keys: softmax(q·kᵀ/√d). Dark = where the head looks.

Thiswasamerchantwhosoldpillsthathadbeeninventedtoquenchthirst.sell-01personpillmerchandise-01thisinvent-01quench-01thirst-01

◯ gold alignment. High-r heads put their darkest cell on the circle; low-r heads spread out or drift to the full stop.

05

Guiding the heads

If some heads already behave like aligners, can we help them? We retrain SPRING with the same hyperparameters and supervise half the heads of layer 3 with silver alignments from LEAMR or ISI. Sparse attention variants did not work, so instead the heads are combined with a learnable scalar mix, as in ELMo, and an extra cross-entropy term joins the parsing loss.

training step 1Forward: the decoder writes person; layer 3 attends over the words.
DECODER WRITESpersonlayer 3 · 16 headspink = 8 supervised heads, mixed by learnt weightsPREDICTED ATTENTION att³(person → word)Thiswasamerchantwhosoldpillsthathadbeeninventedtoquenchthirst.target: silver alignment (ISI / LEAMR)ℒ_align = −log att(merchant)2.96ℒ = ℒ_parse + ℒ_alignℒ_align over steps

Illustrative training dynamics on the Figure 3 sentence. The loss is the cross-entropy between the layer-3 attention mix and the silver alignment (Section 3.2); the parsing loss is trained at the same time.

Now try it yourself: move the slider to train the heads on a second example and watch the mix weights change.

1 · Mix the heads of one layer

Each head h gets a weight s = softmax(a), with γ and a learned. The model is free to make some heads sparse, like an alignment, while others keep doing other work.

attℓ=γ∑h=0H−1shℓatthℓ∈ℝnd×ne

2 · Add an alignment term to the loss

The usual parsing loss, plus a cross-entropy that pushes the mixed attention of each graph token towards the words it is aligned to.

ℒ=−∑j=1ndlogpBART(yj∣y<j,x)−∑j=1∑ialign(i,j)>0nd∑i=1nelog(eattℓ(i,j)∑k=1ndeattℓ(i,k)align(i,j)∑k=1nealign(k,j))
Try it: training with the alignment signal (toy example)
Whyisitsohardtounderstand?hard-02understand-01itcause-01amr-unknownso0.030.080.020.140.570.030.040.030.060.030.030.080.020.090.110.380.210.050.040.210.390.040.060.040.100.050.060.350.130.050.040.090.040.050.040.230.230.070.050.050.060.040.050.040.420.040.050.050.450.250.040.030.030.05

Rows: graph concepts. Columns: subword tokens. Dashed cells: the reference alignment (Table 5, English).

Alignment loss

1.017

Head weights s in layer 3

0
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15

supervised heads free heads

Illustrative numbers, built by hand to show the mechanism. Real results are in the heatmap above and the tables below.

06

From attention to alignment

Attention gives a score for every (graph token, sentence token) pair. Six steps turn those scores into alignments. For the final results, the score matrix is the sum of cross-attention over the first four layers.

  1. 1

    Score matrix

    Build M ∈ ℝ^(n_d × n_e) from cross-attention weights, one row per graph token and one column per subword.

  2. 2

    Span segmentation

    Sum the columns of subwords that form one word (under + stand → understand), and group words into spans.

  3. 3

    Graph segmentation

    Sum the rows of tokens that form one semantic unit (<pointer:1> understand -01 → understand-01).

  4. 4

    Map units to spans

    Give every graph unit the sentence span with the highest score.

  5. 5

    Special structures

    Revise literal subgraphs such as named entities and dates so they align as a block.

  6. 6

    Format

    Write the mapping out in the target standard: ISI or LEAMR.

07

Across languages

For other languages, the same recipe runs on SPRING trained with mBART. We paired the human German, Spanish and Italian translations of the AMR test set with their English graphs. The aligner needs no new rules: it simply reads the attention.

“Why is it so hard to understand?” Figure 4 & Table 5

Aligner

hard-02

understand-01

it

cause-01

amr-unknown

so

Quépredicted
itamr-unknown
reference
cause-01amr-unknown
espredicted—reference—
tanpredicted
cause-01so
reference
so
difícilpredicted
hard-02
reference
hard-02
depredicted—reference—
entenderpredicted
understand-01
reference
understand-01
?predicted—reference—

On this sentence: 4 of 5 reference links recovered, 2 extra.

The Spanish translation asks what is hard, making “Qué” the subject. Our aligner links “it” to “Qué”, which is reasonable but counts as an error against the projected reference.

Alignment F1 by language (ISI standard, Table 3)
Rule / EM alignersCross-attention (ours)

EN

JAMR85.9
TAMR88.1
LEAMR89.0
Ours · unguided94.3
Ours · guided95.2

DE

JAMR12.1
TAMR11.8
LEAMR8.8
Ours · unguided68.8
Ours · guidedn/a

ES

JAMR27.1
TAMR27.5
LEAMR8.5
Ours · unguided72.2
Ours · guidedn/a

IT

JAMR21.9
TAMR21.9
LEAMR9.3
Ours · unguided71.2
Ours · guidedn/a

Non-English references are projected from English with a machine-translation aligner. The guided model was only evaluated in English.

Cross-attention scores 69–72 F1 on German, Spanish and Italian, about 40 points above the next best system on average. It is still lower than English, for two reasons: real linguistic divergence between languages, and errors in the machine-translation projection used to build the non-English references.

08

Results in English

On the LEAMR benchmark, both variants beat LEAMR, the previous state of the art, on subgraphs and relations. The unguided aligner, which never saw an alignment, is already competitive. Guiding layer 3 gives the best overall scores and does not hurt parsing quality.

Reentrancies remain the hardest case for everyone. On the ISI standard, the guided aligner reaches 95.2 F1 in English.

Exact-alignment F1 (LEAMR benchmark, Table 1)
LEAMROurs · unguidedOurs · guided

Subgraph · 1,707 alignments

94.0
94.3
94.5

Relation · 1,263 alignments

85.5
87.4
88.1

Reentrancy · 293 alignments

55.2
44.7
57.0

Duplicate subgraph · 17 alignments

62.5
80.0
75.7

Limitations

  • It needs a Transformer parser, and does not transfer to other architectures.
  • It assumes the decoder attends to the words most relevant to the next token, which does not always hold.
  • Non-English scores are lower, partly because of linguistic divergence and partly because of projected references.
  • LEAMR-format output needs an English-specific span segmentation, so cross-lingual results use the ISI format.

Cite

@inproceedings{martinez-lorenzo-etal-2023-cross,
    title = "Cross-lingual {AMR} Aligner: Paying Attention to Cross-Attention",
    author = "Mart{\'i}nez Lorenzo, Abelardo Carlos  and
      Huguet Cabot, Pere Llu{\'i}s  and
      Navigli, Roberto",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2023",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.findings-acl.109/",
    doi = "10.18653/v1/2023.findings-acl.109",
    pages = "1726--1742"
}
All publications

More interactive papers are on the way.