Skip to content
Abelardo Carlos

InteractiveACL 2022 · Long paper · Part 2 of 2

Fully-semantic parsing and generation: building BMR 1.0

Abelardo Carlos Martínez Lorenzo, Marco Maru, Roberto Navigli · Sapienza NLP & Babelscape

A formalism is only an idea until models can learn it. We turned the largest AMR corpus into the first dataset annotated in BMR, trained parsers and generators in four languages, and tested whether a graph of concepts keeps more of the sentence and translates better.

AMR 3.0 graphs converted into BMR 1.0
59,255AMR 3.0 graphs converted into BMR 1.0
fewer nodes after merging multiwords (936,769 → 828,483)
−11.6%fewer nodes after merging multiwords (936,769 → 828,483)
of content nodes linked to a BabelNet synset
92%of content nodes linked to a BabelNet synset
languages for parsing and generation: EN, DE, IT, ES
4languages for parsing and generation: EN, DE, IT, ES
01

Where Part 1 left off

AMR builds graphs out of English lemmas and PropBank frames. Similar inventories exist for other languages, such as AnCora for Spanish or the Chinese PropBank, but they are not linked to each other, so an AMR graph cannot travel between languages.

BMR replaces lemmas with BabelNet synsets and PropBank with VerbAtlas. The open question: where does the training data come from? Annotating tens of thousands of graphs by hand, in every language, is not realistic. So we started from the graphs that already exist.

Part 1 · AAAI 2022

One graph for every language

The position paper: why AMR is not an interlingua, and what BMR changes.

Read Part 1 →

02

From AMR 3.0 to BMR 1.0

AMR 3.0 · ACL Figure 1
Studentsandtheirparentswilltaketheplaneatthelastminute.:ARG0:ARG1:time:op1:op2:mod:ARG0-of:ARG0-of:ARG1:ARG2:relatedtake-01andplaneminutepersonpersonlaststudy-01have-rel-role-91parent
  1. 0. AMR 3.0
  2. 1. Relations
  3. 2. Node merging
  4. 3. Number · tense · aspect
  5. 4. Disambiguation
  6. 5. Languages
English lemmaPropBank frameVerbAtlas frameBabelNet synset

0 · Start from AMR 3.0

AMR 3.0 holds 59,255 English sentences with hand-made graphs. Our running example: Students and their parents will take the plane at the last minute.

Note how much the graph spends on English conventions: person :ARG0-of study-01 for students, and a special predicate, have-rel-role-91, just to say “their parents”.

1 · Self-explanatory relations

PropBank frames and numbered arguments are swapped for VerbAtlas using the mapping of Di Fabio et al.: take-01 becomes MOVE_BY_MEANS_OF, :ARG0 becomes :agent, :ARG1 becomes :theme.

The mapping is incomplete, so a linguist mapped the missing verbal predicates by hand, and turned special AMR predicates into new roles (Table 6 of the paper lists 30 BMR roles; other AMR roles are kept).

2 · Node merging

Multiword expressions collapse into one node. Sentences are lemmatised with spaCy, the longest lemma sequences that match a BabelNet entry are found, aligned to the graph, then checked by a linguist (rest of the world exists in BabelNet, but only as a sports team).

Graphs are then collapsed from the leaves up: person + study → student, person + have-rel-role-91 + parent → parent :related student, minute + last → at_the_last_minute. Across the corpus: 936,769 → 828,483 nodes (−11.6%).

3 · Number, tense and aspect

AMR drops grammatical information that carries meaning. Using part-of-speech tags, BMR adds :timing + or :timing - for future or past events, :quantity + for plurals, and :ongoing + for ongoing or habitual actions.

Here: will take gets :timing +, and both students and parents get :quantity +.

4 · Disambiguation

Finally every content node gets a BabelNet synset, via three routes: the VerbAtlas-to-BabelNet mapping for predicates, Wikipedia links for named entities, and ESCHER, a state-of-the-art word sense disambiguation system, for everything else.

plane

airplanegeometric planecarpenter's plane

→ bn:00001697n

padres (ES)

parentsfathers

→ bn:00060643n

92% of content nodes get a synset; 42,549 of 59,255 graphs are fully disambiguated.

One graph, four languages

The result is BMR 1.0. Switch the language above the graph: German, Italian and Spanish translations share the same synsets, the same frame and the same roles. Only the words on the surface change.

The same sentence, as text ACL Appendix A, Figure 4

AMR 3.0

(t / take-01
   :ARG0 (a / and
      :op1 (p / person
         :ARG0-of (s / study-01))
      :op2 (p2 / person
         :ARG0-of (h / have-rel-role-91
            :ARG1 p
            :ARG2 (p3 / parent))))
   :ARG1 (p4 / plane)
   :time (t2 / minute
      :mod (l / last)))

BMR 1.0

(t / take / bn:00094732v
   :timing +
   :agent (a / and
      :op1 (s / student / bn:00029806n
         :quantity +)
      :op2 (p / parent / bn:00060643n
         :quantity +
         :related s))
   :theme (p2 / plane / bn:00001697n)
   :timing (t2 / at_the_last_minute
            / bn:00114428r))
03

Does it work?

We compared four versions of the same data. AMR is AMR 3.0 unchanged. AMR+ has every BMR change except synsets (relations, merging, number, tense, aspect). BMR is BMR 1.0. BMR* is BMR with the lemmas removed, pure synsets.

All models are SPRING (BART, or mBART for other languages), with frequent synsets added to the vocabulary. German, Italian and Spanish training data are machine translations paired with the gold graphs; the test set is 1,371 human-translated sentences.

Graph → Text (Table 1) higher is better
AMRAMR+BMRBMR*
BLEU4352AMR · BLEU: 44.8AMR+ · BLEU: 49.8BMR · BLEU: 50.7BMR* · BLEU: 45.7BMR 50.7chrF++7178AMR · chrF++: 73.4AMR+ · chrF++: 76.0BMR · chrF++: 76.3BMR* · chrF++: 72.1BMR 76.3METEOR4146AMR · METEOR: 42.2AMR+ · METEOR: 43.9BMR · METEOR: 44.3BMR* · METEOR: 42.4BMR 44.3ROUGE-L6774AMR · ROUGE-L: 68.2AMR+ · ROUGE-L: 71.7BMR · ROUGE-L: 72.8BMR* · ROUGE-L: 69.7BMR 72.8

Each row has its own axis: compare dots within a row, not distances across rows.

Show the values as a table
AMRAMR+BMRBMR*
BLEU44.849.850.745.7
chrF++73.476.076.372.1
METEOR42.243.944.342.4
ROUGE-L68.271.772.869.7
Each row is zoomed on its own so that small differences stay visible. Hover, tap or focus a dot for its value.

BMR generates the best text on every measure in every language (tied with AMR+ once, on German METEOR). The gap between AMR+ and BMR is the contribution of the synsets alone. Which features matter most? The ablation below adds them one at a time, in English.

Ablation: what each feature adds (Table 2, English)
AMRAMR 3.0 baseline4352score · AMR: 44.844.8AMR·REL+ self-explanatory relations4352score · AMR·REL: 44.944.9AMR·NOD+ relations and node merging4352score · AMR·NOD: 45.545.5AMR·NUMnumber4352score · AMR·NUM: 46.946.9AMR·TENtense and aspect4352score · AMR·TEN: 47.647.6AMR·NTnumber, tense and aspect4352score · AMR·NT: 49.049.0AMR+all of the above4352score · AMR+: 49.849.8BMR+ BabelNet synsets4352score · BMR: 50.750.7
Show the values as a table
score
AMRAMR 3.0 baseline44.8
AMR·REL+ self-explanatory relations44.9
AMR·NOD+ relations and node merging45.5
AMR·NUMnumber46.9
AMR·TENtense and aspect47.6
AMR·NTnumber, tense and aspect49.0
AMR+all of the above49.8
BMR+ BabelNet synsets50.7
One axis for all rows. Model names follow the paper; number (NUM) and tense/aspect (TEN) are also tested separately, and NT combines them.

Readable relations alone barely move the needle; node merging helps, and number, tense and aspect help most (+4.2 BLEU together). All of them combined (AMR+) beat the baseline by 5.0 BLEU, and adding synsets (BMR) goes further still.

04

Case study: what the numbers miss

BMR* scores lowest overall, yet its outputs can be better. Generating from AMR loses number (friends) and tense (do), and the reentrant have-rel-role-91 structure confuses whose father it is (my).

BMR* keeps all three, and writes put up with instead of tolerate: a correct synonym that string-matching metrics such as BLEU penalise. The evaluation, not the meaning, is what drops.

“My friend did not tolerate his father’s behaviour” ACL Figure 3

AMR

(t / tolerate-01 :polarity -
   :ARG0 (p / person
      :ARG0-of (h / have-rel-role-91
         :ARG1 (i / i)
         :ARG2 (f / friend))
   :ARG1 (b / behave-01
      :ARG0 (p2 / person
         :ARG0-of (h2 / have-rel-role-91
            :ARG1 p
            :ARG2 (f2 / father)))))

Generated sentence

My friends do not tolerate the behavior of my father.

BMR* (synsets, no lemmas)

(t / bn:00082138v :polarity - :timing -
   :agent (f / bn:00036538n
      :related (i / i))
   :theme (b / bn:00009656n
      :related (f2 / bn:00009616n
         :related f)))

Generated sentence

My friend did not put up with the behaviour of his father.

■ number of “friend”■ tense■ whose father■ the predicate
05

Where BMR still falls short

Repository

BabelNet covers nouns, verbs, adjectives and adverbs. Conjunctions and ambiguous pronouns such as “anyone” stay as lemmas, and the 500 BabelNet languages are a subset of the ~6,500 spoken in the world.

Disambiguation

ESCHER predicts WordNet senses only, so polysemous multiwords found in BabelNet but not WordNet (“run off at the mouth”) stay undisambiguated: the missing 8%.

Culture

Some concepts exist in one language only. “Espeto”, the Málaga way of grilling freshly caught fish on a skewer, has a synset but no English word: it has to be paraphrased.

In one line

A graph of concepts preserves more of the sentence and is a better bridge between languages, even if it is harder to predict.

Next steps in the paper: a single multilingual model for every language, other cross-lingual tasks such as summarisation, and a formalism with no lexical information at all.

Cite

@inproceedings{martinez-lorenzo-etal-2022-fully,
    title = "{F}ully-{S}emantic {P}arsing and {G}eneration: the {B}abel{N}et {M}eaning {R}epresentation",
    author = "Mart{\'i}nez Lorenzo, Abelardo Carlos  and
      Maru, Marco  and
      Navigli, Roberto",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.121/",
    doi = "10.18653/v1/2022.acl-long.121",
    pages = "1727--1741"
}