InteractiveACL 2022 · Long paper · Part 2 of 2
Fully-semantic parsing and generation: building BMR 1.0
Abelardo Carlos Martínez Lorenzo, Marco Maru, Roberto Navigli · Sapienza NLP & Babelscape
A formalism is only an idea until models can learn it. We turned the largest AMR corpus into the first dataset annotated in BMR, trained parsers and generators in four languages, and tested whether a graph of concepts keeps more of the sentence and translates better.
- AMR 3.0 graphs converted into BMR 1.0
- 59,255AMR 3.0 graphs converted into BMR 1.0
- fewer nodes after merging multiwords (936,769 → 828,483)
- −11.6%fewer nodes after merging multiwords (936,769 → 828,483)
- of content nodes linked to a BabelNet synset
- 92%of content nodes linked to a BabelNet synset
- languages for parsing and generation: EN, DE, IT, ES
- 4languages for parsing and generation: EN, DE, IT, ES
Where Part 1 left off
AMR builds graphs out of English lemmas and PropBank frames. Similar inventories exist for other languages, such as AnCora for Spanish or the Chinese PropBank, but they are not linked to each other, so an AMR graph cannot travel between languages.
BMR replaces lemmas with BabelNet synsets and PropBank with VerbAtlas. The open question: where does the training data come from? Annotating tens of thousands of graphs by hand, in every language, is not realistic. So we started from the graphs that already exist.
Part 1 · AAAI 2022
One graph for every language
The position paper: why AMR is not an interlingua, and what BMR changes.
Read Part 1 →
From AMR 3.0 to BMR 1.0
- 0. AMR 3.0
- 1. Relations
- 2. Node merging
- 3. Number · tense · aspect
- 4. Disambiguation
- 5. Languages
0 · Start from AMR 3.0
AMR 3.0 holds 59,255 English sentences with hand-made graphs. Our running example: Students and their parents will take the plane at the last minute.
Note how much the graph spends on English conventions: person :ARG0-of study-01 for students, and a special predicate, have-rel-role-91, just to say “their parents”.
1 · Self-explanatory relations
PropBank frames and numbered arguments are swapped for VerbAtlas using the mapping of Di Fabio et al.: take-01 becomes MOVE_BY_MEANS_OF, :ARG0 becomes :agent, :ARG1 becomes :theme.
The mapping is incomplete, so a linguist mapped the missing verbal predicates by hand, and turned special AMR predicates into new roles (Table 6 of the paper lists 30 BMR roles; other AMR roles are kept).
2 · Node merging
Multiword expressions collapse into one node. Sentences are lemmatised with spaCy, the longest lemma sequences that match a BabelNet entry are found, aligned to the graph, then checked by a linguist (rest of the world exists in BabelNet, but only as a sports team).
Graphs are then collapsed from the leaves up: person + study → student, person + have-rel-role-91 + parent → parent :related student, minute + last → at_the_last_minute. Across the corpus: 936,769 → 828,483 nodes (−11.6%).
3 · Number, tense and aspect
AMR drops grammatical information that carries meaning. Using part-of-speech tags, BMR adds :timing + or :timing - for future or past events, :quantity + for plurals, and :ongoing + for ongoing or habitual actions.
Here: will take gets :timing +, and both students and parents get :quantity +.
4 · Disambiguation
Finally every content node gets a BabelNet synset, via three routes: the VerbAtlas-to-BabelNet mapping for predicates, Wikipedia links for named entities, and ESCHER, a state-of-the-art word sense disambiguation system, for everything else.
plane
→ bn:00001697n
padres (ES)
→ bn:00060643n
92% of content nodes get a synset; 42,549 of 59,255 graphs are fully disambiguated.
One graph, four languages
The result is BMR 1.0. Switch the language above the graph: German, Italian and Spanish translations share the same synsets, the same frame and the same roles. Only the words on the surface change.
AMR 3.0
(t / take-01
:ARG0 (a / and
:op1 (p / person
:ARG0-of (s / study-01))
:op2 (p2 / person
:ARG0-of (h / have-rel-role-91
:ARG1 p
:ARG2 (p3 / parent))))
:ARG1 (p4 / plane)
:time (t2 / minute
:mod (l / last)))BMR 1.0
(t / take / bn:00094732v
:timing +
:agent (a / and
:op1 (s / student / bn:00029806n
:quantity +)
:op2 (p / parent / bn:00060643n
:quantity +
:related s))
:theme (p2 / plane / bn:00001697n)
:timing (t2 / at_the_last_minute
/ bn:00114428r))Does it work?
We compared four versions of the same data. AMR is AMR 3.0 unchanged. AMR+ has every BMR change except synsets (relations, merging, number, tense, aspect). BMR is BMR 1.0. BMR* is BMR with the lemmas removed, pure synsets.
All models are SPRING (BART, or mBART for other languages), with frequent synsets added to the vocabulary. German, Italian and Spanish training data are machine translations paired with the gold graphs; the test set is 1,371 human-translated sentences.
Each row has its own axis: compare dots within a row, not distances across rows.
Show the values as a table
| AMR | AMR+ | BMR | BMR* | |
|---|---|---|---|---|
| BLEU | 44.8 | 49.8 | 50.7 | 45.7 |
| chrF++ | 73.4 | 76.0 | 76.3 | 72.1 |
| METEOR | 42.2 | 43.9 | 44.3 | 42.4 |
| ROUGE-L | 68.2 | 71.7 | 72.8 | 69.7 |
BMR generates the best text on every measure in every language (tied with AMR+ once, on German METEOR). The gap between AMR+ and BMR is the contribution of the synsets alone. Which features matter most? The ablation below adds them one at a time, in English.
Show the values as a table
| score | |
|---|---|
| AMRAMR 3.0 baseline | 44.8 |
| AMR·REL+ self-explanatory relations | 44.9 |
| AMR·NOD+ relations and node merging | 45.5 |
| AMR·NUMnumber | 46.9 |
| AMR·TENtense and aspect | 47.6 |
| AMR·NTnumber, tense and aspect | 49.0 |
| AMR+all of the above | 49.8 |
| BMR+ BabelNet synsets | 50.7 |
Readable relations alone barely move the needle; node merging helps, and number, tense and aspect help most (+4.2 BLEU together). All of them combined (AMR+) beat the baseline by 5.0 BLEU, and adding synsets (BMR) goes further still.
Case study: what the numbers miss
BMR* scores lowest overall, yet its outputs can be better. Generating from AMR loses number (friends) and tense (do), and the reentrant have-rel-role-91 structure confuses whose father it is (my).
BMR* keeps all three, and writes put up with instead of tolerate: a correct synonym that string-matching metrics such as BLEU penalise. The evaluation, not the meaning, is what drops.
AMR
(t / tolerate-01 :polarity - :ARG0 (p / person :ARG0-of (h / have-rel-role-91 :ARG1 (i / i) :ARG2 (f / friend)) :ARG1 (b / behave-01 :ARG0 (p2 / person :ARG0-of (h2 / have-rel-role-91 :ARG1 p :ARG2 (f2 / father)))))
Generated sentence
My friends do not tolerate the behavior of my father.
BMR* (synsets, no lemmas)
(t / bn:00082138v :polarity - :timing - :agent (f / bn:00036538n :related (i / i)) :theme (b / bn:00009656n :related (f2 / bn:00009616n :related f)))
Generated sentence
My friend did not put up with the behaviour of his father.
Where BMR still falls short
Repository
BabelNet covers nouns, verbs, adjectives and adverbs. Conjunctions and ambiguous pronouns such as “anyone” stay as lemmas, and the 500 BabelNet languages are a subset of the ~6,500 spoken in the world.
Disambiguation
ESCHER predicts WordNet senses only, so polysemous multiwords found in BabelNet but not WordNet (“run off at the mouth”) stay undisambiguated: the missing 8%.
Culture
Some concepts exist in one language only. “Espeto”, the Málaga way of grilling freshly caught fish on a skewer, has a synset but no English word: it has to be paraphrased.
In one line
A graph of concepts preserves more of the sentence and is a better bridge between languages, even if it is harder to predict.
Next steps in the paper: a single multilingual model for every language, other cross-lingual tasks such as summarisation, and a formalism with no lexical information at all.
Cite
@inproceedings{martinez-lorenzo-etal-2022-fully,
title = "{F}ully-{S}emantic {P}arsing and {G}eneration: the {B}abel{N}et {M}eaning {R}epresentation",
author = "Mart{\'i}nez Lorenzo, Abelardo Carlos and
Maru, Marco and
Navigli, Roberto",
booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = may,
year = "2022",
address = "Dublin, Ireland",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.acl-long.121/",
doi = "10.18653/v1/2022.acl-long.121",
pages = "1727--1741"
}