InteractiveAAAI 2022 · Position paper · Part 1 of 2
BabelNet Meaning Representation: one graph for every language
Roberto Navigli, Rexhina Blloshmi, Abelardo Carlos Martínez Lorenzo · Sapienza NLP Group
Semantic parsing turns sentences into graphs that machines can reason over. But the most popular graphs are built from English words and English verb lists. We proposed a representation made of concepts instead of words, so the same graph can stand for a sentence in any language.
- graph shared by every translation of a sentence
- 1graph shared by every translation of a sentence
- languages covered by BabelNet, the concept inventory
- 500languages covered by BabelNet, the concept inventory
- steps from AMR to BMR: concepts, relations, multiwords
- 3steps from AMR to BMR: concepts, relations, multiwords
Why meaning needs a structure
Form is not meaning
Large language models read form. Understanding, the argument goes, needs meaning made explicit: a structure a machine can process and a human can still read.
Take the example from the paper: The student's mouse is on top of the external hard drive.
Semantic parsing
Semantic parsing picks the words up and arranges them into a graph: what is where, what belongs to whom. The sentence stays on top; its content words have moved into the structure.
Abstract Meaning Representation
Many formalisms exist (DRT, EDS, PTG, UCCA, UDS, UMR). The most popular is AMR. It replaces verbs with PropBank frames that carry a sense number (be_located_at.91, hard.04, study.01) and links them with numbered arguments (:ARG1, :ARG2). A student becomes a person who studies.
What the graph is made of
Look closer: every node is either an English word or an English frame. AMR was designed for English sentences, and it is not an interlingua.
Where AMR stops being universal
Words are ambiguous
The node mouse is still a word. A person infers the computer device from hard drive, but the graph also allows an animal sitting on the drive. A word is not a meaning.
English-only rules
AMR uses frames wherever it can, so student becomes person :ARG0-of study.01. That is not quite the same thing (not every person who studies is a student), and the rule does not carry over to other languages.
Multiwords and idioms
External hard drive is split into three nodes, as if its meaning were the sum of its words. For idioms such as miss the boat, composing the words gives the wrong meaning altogether.
Another language, another graph
In Spanish there is no PropBank, so we need a Spanish inventory such as AnCora. Frames do not match (be_located_at.91 vs estar.01, hard.04 vs duro) and relations change (:ARG1-of becomes :mod).
Same meaning, different graph.
BMR: concepts instead of words
Three changes
BMR keeps AMR's shape (a directed, labelled graph) and changes what it is made of, in three steps. Keep scrolling and watch the graph transform.
1 · Concepts
Every node becomes a BabelNet synset, a concept with an identifier. mouse is now bn:00021487n: the computer device, never the animal. Verbs become language-independent VerbAtlas frames: is belongs to STAY-DWELL. Hover a node to see its words in six languages.
2 · Relations and 3 · Multiwords
With VerbAtlas frames come readable, cross-frame roles: :theme and :location instead of :ARG1 and :ARG2. And external hard drive collapses into one synset, bn:21899122n, as does student.
Paraphrases
Because nodes are concepts, a paraphrase such as the student's computer mouse is on the upper side of the external HDD
lands on exactly the same graph.
Every language
Now switch the language (or let it cycle). German, Spanish, Italian, French and Chinese sentences all map to the same five nodes and four edges. The sentence changes; the graph does not.
Two ingredients
Nodes
BabelNet
A multilingual encyclopedic dictionary and semantic network. It groups words into synsets, sets of synonyms in up to 500 languages, integrating WordNet, Wikipedia and more.
Predicates and relations
VerbAtlas
A hand-crafted inventory that clusters verbal concepts into semantically coherent frames, with human-readable roles shared across frames (AGENT, LOCATION, BENEFICIARY) instead of PropBank's numbered, English-specific arguments.
The paper also sketches how to get there automatically: start from AMR 3.0 (59,255 annotated sentences), propagate synsets into the graph with multilingual word sense disambiguation (80–85% accurate in many languages) and entity linking, and swap PropBank frames for VerbAtlas frames using an existing mapping.
That is exactly what we did next. Part 2 builds the dataset and tests it →
Beyond text
A representation made of concepts is not tied to written language at all. The paper closes by imagining BMR as a shared layer across AI:
Machine translation
Parse into the interlingua, generate from it: no bilingual corpora needed.
Question answering
Symbolic questions that retrieve facts from a knowledge base, across languages.
Dialogue
User intent in a language-independent yet human-readable form.
Vision, speech, sound
Concepts, not words, so other modalities can share the same graph.
Knowledge representation
Logical formulas linked to explicit concepts, towards explainable AI.
Honest caveats
- BabelNet covers many languages, but not all of them, so full language independence is still a goal.
- Temporal information and plurality are left out here; nothing prevents adding them (Part 2 does).
- This is a position paper: the evidence comes in Part 2.
Cite
@article{navigli-etal-2022-bmr,
title = "BabelNet Meaning Representation: A Fully Semantic Formalism to Overcome Language Barriers",
author = "Navigli, Roberto and
Blloshmi, Rexhina and
Mart{\'i}nez Lorenzo, Abelardo Carlos",
journal = "Proceedings of the AAAI Conference on Artificial Intelligence",
volume = "36",
number = "11",
pages = "12274--12279",
year = "2022",
doi = "10.1609/aaai.v36i11.21490"
}Figures redrawn from the paper; synset ids and lexicalisations as published.