InteractiveACL 2023 · Short paper
AMRs Assemble! When a higher score hides a broken graph
Abelardo Carlos Martínez Lorenzo*, Pere-Lluís Huguet Cabot*, Roberto Navigli · Babelscape & Sapienza NLP · *equal contribution
The best AMR parsers were ensembles that merge several predictions into one graph. We found that they climb the leaderboard partly by exploiting the metric, producing graphs that break AMR's own rules. So we built ensemblers that learn to merge, or simply choose, and stay valid.
- of graphs merged by Graphene violate AMR constraints
- 19.7%of graphs merged by Graphene violate AMR constraints
- for Assemble!, with the same SMATCH as the best Graphene
- 0.3%for Assemble!, with the same SMATCH as the best Graphene
- SMATCH for Assemble! avg, the best ensemble in the comparison
- 84.1SMATCH for Assemble! avg, the best ensemble in the comparison
Five parsers, five slightly different graphs
The reference
Antonio Banderas scheduled the premiere of his movie at 3 pm.
In the gold AMR, person is reused three times: the one who schedules, the one whose premiere it is, and the owner of the movie.
Prediction 1 · SMATCH 80.0
We trained SPRING with five different seeds. Their graphs disagree in small ways. This one writes premiere as a plain concept instead of the predicate premiere-01, and links the movie with :mod.
Prediction 2 · SMATCH 88.9
This one gets the whole structure right, but reads the time as “3:00” instead of “15:00”.
The merge · SMATCH 85.0
Graphene, a leading ensembling method, merges predictions by voting on nodes and edges. When the votes tie, it can keep both options: two edges from premiere to movie, two times, and :ARG roles on a node that is not a predicate, which AMR forbids.
A broken graph, yet it outscores Prediction 1.
Why the metric is fooled
SMATCH turns a graph into a set of triples and measures their overlap with the gold set. It knows nothing about what a valid AMR looks like. Missing a triple costs recall; adding a wrong one only costs some precision. So when in doubt, keeping both options pays off. Switch between the candidates:
Gold · 17 triples
- ✓ (root, :top, schedule)
- ✓ (schedule, :instance, schedule-01)
- ✓ (person, :instance, person)
- ✗ (premiere, :instance, premiere-01)
- ✓ (date, :instance, date-entity)
- ✓ (name, :instance, name)
- ✓ (movie, :instance, movie)
- ✓ (schedule, :ARG0, person)
- ✓ (schedule, :ARG1, premiere)
- ✓ (schedule, :ARG3, date)
- ✓ (person, :name, name)
- ✓ (name, :op1, "Antonio")
- ✓ (name, :op2, "Banderas")
- ✓ (premiere, :ARG0, person)
- ✓ (premiere, :ARG1, movie)
- ✓ (movie, :poss, person)
- ✓ (date, :time, "15:00")
Not in gold · 4 extra triples
- + (premiere, :instance, premiere)
- + (premiere, :poss, person)
- + (premiere, :mod, movie)
- + (date, :time, "3:00")
80.0
94.1
86.5
We wrote a checker for four kinds of violation: :ARG relations on non-predicate nodes, :op/:snt relations on predicates, and broken entity and connector structures. On the AMR 3.0 test set, Graphene's merged graphs break these rules 13.7–19.7% of the time.
Assemble!: learning to merge
Instead of voting on triples, we train a model to merge. It reads the sentence and several candidate graphs, and writes one graph, just as a parser would. We use LongT5, a seq2seq model built for long inputs, since five linearised graphs plus a sentence make a long sequence.
First we need training data: predictions from parsers that have never seen the sentence.
train SPRING on 4 folds
predict the 5th
predictions collected
Result: 5 predictions for every sentence, none from a model that saw it in training.
Then two stages. Pre-training extends AMRBART's graph denoising with new tasks where the model rebuilds a graph from several masked copies. Fine-tuning mixes plain parsing with ensembling, shuffling the order and number of candidates every epoch. Pick a task:
Pre-training
Fine-tuning
Sentence plus candidates: the full ensembling task.
Input
Output
Or don't merge at all: pick the best one
Generating a whole graph autoregressively is slow. The cheaper idea: every candidate is already a complete graph, so score them and keep the most plausible. Perplexity takes a single forward pass per candidate. If there is an oracle that always picks the best candidate, selection could reach 86.5 SMATCH, 3.4 points above the best single parser.
Each of the five parsers scores every candidate; the graph with the lowest average perplexity wins. No ensembler needed.
Lower perplexity = the model finds the graph more plausible.
Results
| Model | SMATCH | Corrupted graphs | Time (s) |
|---|---|---|---|
| Predictions | |||
| SPRING₁ | 83.1 | 54 · 2.8% | — |
| SPRING₂ | 82.7 | 52 · 2.7% | — |
| SPRING₃ | 83.0 | 73 · 3.8% | — |
| SPRING₄ | 82.8 | 33 · 1.7% | — |
| SPRING₅ | 82.6 | 104 · 5.5% | — |
| Oracle | |||
| Best graph (oracle) | 86.5 | 51 · 2.7% | — |
| Mergers | |||
| Graphene base | 83.6 | 374 · 19.7% | 810 |
| Graphene SMATCH | 83.8 | 260 · 13.7% | 11,884 |
| Assemble! | 83.8 | 6 · 0.3% | 431 |
| Selectors | |||
| SMATCH avg | 83.7 | 51 · 2.7% | 493 |
| Assemble! zero | 83.9 | 13 · 0.7% | 256 |
| Assemble! avg | 84.1 | 22 · 1.2% | 635 |
6 vs 374
corrupted graphs: Assemble! matches Graphene SMATCH’s 83.8 with almost no rule violations.
28×
faster than Graphene SMATCH: 431 s versus 11,884 s for the same score.
256 s
for Assemble! zero, the fastest ensemble, which also beats SMATCH-based selection.
A last caveat from the paper. Parsers and ensembles now score around 0.83–0.84 SMATCH, above the reported agreement between human annotators and the consensus (0.83 on newswire, 0.79 on web text). Beyond that point, whether SMATCH still measures parsing quality is an open question, and structural validity deserves a place next to it.
Limitations
- Only evaluated on AMR parsing; other structured prediction tasks are untested.
- The constraint checker covers four classes of violation, not every AMR guideline.
- Large autoregressive ensemblers are still costly; selection is the cheaper path.
Cite
@inproceedings{martinez-lorenzo-etal-2023-amrs,
title = "{AMR}s Assemble! Learning to Ensemble with Autoregressive Models for {AMR} Parsing",
author = "Mart{\'i}nez Lorenzo, Abelardo Carlos and
Huguet Cabot, Pere Llu{\'i}s and
Navigli, Roberto",
booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)",
month = jul,
year = "2023",
address = "Toronto, Canada",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2023.acl-short.137/",
doi = "10.18653/v1/2023.acl-short.137",
pages = "1595--1605"
}Graphs redrawn from Figure 2; scores copied from Table 2.