Skip to content
Abelardo Carlos

InteractiveACL 2023 · Short paper

AMRs Assemble! When a higher score hides a broken graph

Abelardo Carlos Martínez Lorenzo*, Pere-Lluís Huguet Cabot*, Roberto Navigli · Babelscape & Sapienza NLP · *equal contribution

The best AMR parsers were ensembles that merge several predictions into one graph. We found that they climb the leaderboard partly by exploiting the metric, producing graphs that break AMR's own rules. So we built ensemblers that learn to merge, or simply choose, and stay valid.

of graphs merged by Graphene violate AMR constraints
19.7%of graphs merged by Graphene violate AMR constraints
for Assemble!, with the same SMATCH as the best Graphene
0.3%for Assemble!, with the same SMATCH as the best Graphene
SMATCH for Assemble! avg, the best ensemble in the comparison
84.1SMATCH for Assemble! avg, the best ensemble in the comparison
01

Five parsers, five slightly different graphs

Gold AMR Figure 2
reference
:ARG0:ARG1:ARG3:name:op1:op2:ARG0:ARG1:poss:timeschedule-01personpremiere-01date-entitynamemovie"15:00""3:00""Antonio""Banderas"
predicate (frame)conceptconstant

The reference

Antonio Banderas scheduled the premiere of his movie at 3 pm. In the gold AMR, person is reused three times: the one who schedules, the one whose premiere it is, and the owner of the movie.

Prediction 1 · SMATCH 80.0

We trained SPRING with five different seeds. Their graphs disagree in small ways. This one writes premiere as a plain concept instead of the predicate premiere-01, and links the movie with :mod.

Prediction 2 · SMATCH 88.9

This one gets the whole structure right, but reads the time as “3:00” instead of “15:00”.

The merge · SMATCH 85.0

Graphene, a leading ensembling method, merges predictions by voting on nodes and edges. When the votes tie, it can keep both options: two edges from premiere to movie, two times, and :ARG roles on a node that is not a predicate, which AMR forbids.

A broken graph, yet it outscores Prediction 1.

02

Why the metric is fooled

SMATCH turns a graph into a set of triples and measures their overlap with the gold set. It knows nothing about what a valid AMR looks like. Missing a triple costs recall; adding a wrong one only costs some precision. So when in doubt, keeping both options pays off. Switch between the candidates:

The graph as SMATCH sees it: a bag of triples

Gold · 17 triples

  • ✓ (root, :top, schedule)
  • ✓ (schedule, :instance, schedule-01)
  • ✓ (person, :instance, person)
  • ✗ (premiere, :instance, premiere-01)
  • ✓ (date, :instance, date-entity)
  • ✓ (name, :instance, name)
  • ✓ (movie, :instance, movie)
  • ✓ (schedule, :ARG0, person)
  • ✓ (schedule, :ARG1, premiere)
  • ✓ (schedule, :ARG3, date)
  • ✓ (person, :name, name)
  • ✓ (name, :op1, "Antonio")
  • ✓ (name, :op2, "Banderas")
  • ✓ (premiere, :ARG0, person)
  • ✓ (premiere, :ARG1, movie)
  • ✓ (movie, :poss, person)
  • ✓ (date, :time, "15:00")

Not in gold · 4 extra triples

  • + (premiere, :instance, premiere)
  • + (premiere, :poss, person)
  • + (premiere, :mod, movie)
  • + (date, :time, "3:00")
Precision16/20

80.0

Recall16/17

94.1

F1harmonic mean

86.5

Simplified: variables are matched by position and labels must match exactly. The official SMATCH also searches over variable mappings, so its scores (80.0, 88.9, 85.0) differ slightly, but the ranking is the same.

We wrote a checker for four kinds of violation: :ARG relations on non-predicate nodes, :op/:snt relations on predicates, and broken entity and connector structures. On the AMR 3.0 test set, Graphene's merged graphs break these rules 13.7–19.7% of the time.

03

Assemble!: learning to merge

Instead of voting on triples, we train a model to merge. It reads the sentence and several candidate graphs, and writes one graph, just as a parser would. We use LongT5, a seq2seq model built for long inputs, since five linearised graphs plus a sentence make a long sequence.

First we need training data: predictions from parsers that have never seen the sentence.

Building a corpus of predictions Section 2.2
split 1
split 2
split 3
split 4
split 5

train SPRING on 4 folds

predict the 5th

predictions collected

Result: 5 predictions for every sentence, none from a model that saw it in training.

Illustration of the procedure: five models per split, five different splits of the 59,255 AMR 3.0 sentences.

Then two stages. Pre-training extends AMRBART's graph denoising with new tasks where the model rebuilds a graph from several masked copies. Fine-tuning mixes plain parsing with ensembling, shuffling the order and number of candidates every epoch. Pick a task:

What LongT5 is trained on Table 1
  1. Pre-training

  2. Fine-tuning

Sentence plus candidates: the full ensembling task.

Input

<s>x₁x₂…xₘ<g>p1₁…p1ₗ<g>…<g>p5₁…p5ₗ</s>

Output

<s>g₁g₂…gₙ</s>
x = sentence tokens, g = gold graph tokens, p = a SPRING prediction, [mask] = masked span.
04

Or don't merge at all: pick the best one

Generating a whole graph autoregressively is slow. The cheaper idea: every candidate is already a complete graph, so score them and keep the most plausible. Perplexity takes a single forward pass per candidate. If there is an oracle that always picks the best candidate, selection could reach 86.5 SMATCH, 3.4 points above the best single parser.

Select, don't merge Section 2.3

Each of the five parsers scores every candidate; the graph with the lowest average perplexity wins. No ensembler needed.

candidate 1
1.42
candidate 2
1.32
candidate 3
1.22
candidate 4
1.33
candidate 5
1.51

Lower perplexity = the model finds the graph more plausible.

s′=arg⁡mins∈{1,…,l}1l∑j∈{1,…,l}perplexityj(t2ps)
Perplexity values are illustrative. Selection can only return a graph that one of the parsers actually produced, so it cannot invent a structure that breaks AMR's rules.
05

Results

AMR 3.0 test set (Table 2) 1,898 graphs
ModelSMATCHCorrupted graphsTime (s)
Predictions
SPRING₁
83.1
54 · 2.8%
—
SPRING₂
82.7
52 · 2.7%
—
SPRING₃
83.0
73 · 3.8%
—
SPRING₄
82.8
33 · 1.7%
—
SPRING₅
82.6
104 · 5.5%
—
Oracle
Best graph (oracle)
86.5
51 · 2.7%
—
Mergers
Graphene base
83.6
374 · 19.7%
810
Graphene SMATCH
83.8
260 · 13.7%
11,884
Assemble!
83.8
6 · 0.3%
431
Selectors
SMATCH avg
83.7
51 · 2.7%
493
Assemble! zero
83.9
13 · 0.7%
256
Assemble! avg
84.1
22 · 1.2%
635
SMATCH on a zoomed axis (82–87). Time on a log scale; the individual parsers have no ensembling time.

6 vs 374

corrupted graphs: Assemble! matches Graphene SMATCH’s 83.8 with almost no rule violations.

28×

faster than Graphene SMATCH: 431 s versus 11,884 s for the same score.

256 s

for Assemble! zero, the fastest ensemble, which also beats SMATCH-based selection.

A last caveat from the paper. Parsers and ensembles now score around 0.83–0.84 SMATCH, above the reported agreement between human annotators and the consensus (0.83 on newswire, 0.79 on web text). Beyond that point, whether SMATCH still measures parsing quality is an open question, and structural validity deserves a place next to it.

Limitations

  • Only evaluated on AMR parsing; other structured prediction tasks are untested.
  • The constraint checker covers four classes of violation, not every AMR guideline.
  • Large autoregressive ensemblers are still costly; selection is the cheaper path.
A single parser that gets close without ensembling: LeakDistill →

Cite

@inproceedings{martinez-lorenzo-etal-2023-amrs,
    title = "{AMR}s Assemble! Learning to Ensemble with Autoregressive Models for {AMR} Parsing",
    author = "Mart{\'i}nez Lorenzo, Abelardo Carlos  and
      Huguet Cabot, Pere Llu{\'i}s  and
      Navigli, Roberto",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-short.137/",
    doi = "10.18653/v1/2023.acl-short.137",
    pages = "1595--1605"
}
All publications

Graphs redrawn from Figure 2; scores copied from Table 2.