InteractiveFindings of ACL 2024
The Multilingual Semantic Layer: semantic parsing in eleven languages
Abelardo Carlos Martínez Lorenzo, Pere-Lluís Huguet Cabot, Karim Ghonim, Lu Xu, Hee-Soo Choi, Alberte Fernández Castro, Roberto Navigli
Semantic parsers need annotated graphs, and annotated graphs barely exist outside English. We designed a simpler layer that any language can express in its own words, and built a dataset for eleven languages on an academic budget, with LLMs doing the heavy lifting and humans checking the work.
- languages, including Arabic, Korean and Chinese
- 11languages, including Arabic, Korean and Chinese
- silver graphs parsed from license-free text
- 10M+silver graphs parsed from license-free text
- gold graphs annotated by hand as a benchmark
- 1,100gold graphs annotated by hand as a benchmark
- per graph for the LLM steps, about $400 in total
- $0.0013per graph for the LLM steps, about $400 in total
Semantic parsing has an English problem
Graph formalisms like AMR turn a sentence into concepts and relations that machines can use. But parsers learn from annotated graphs, and annotating is slow: annotators must master the formalism, its repositories and each language's rules.
The result: the big corpora are English, sit behind licences, and efforts in Turkish, Persian, Portuguese or Vietnamese have produced no more than about 200 graphs per language, far too few to train a parser.
DRT
1993
PMB 4.0 · WordNet · English graph
PDT
2003
PDT 3.0 · English graph
UCCA
2013
Wikipedia · English graph
AMR
2013
AMR 3.0 · PropBank · English graph
UMR
2021
UMR 1.0 · PropBank · Non-specific graph
BMR
2022
BMR 1.0 · BabelNet · English graph
One sentence, five formalisms
English AMR
The football player jaywalks across the street to the chiringuito.
AMR uses the PropBank frame jaywalk-00 and spells out a football player as a person who plays football.
Spanish AMR
Now the Spanish translation. Spanish AMR needs a Spanish frame inventory (AnCora): cruzar-01, jugar-02. A different repository, a different graph, and almost no annotated data.
Spanish PMB
The Parallel Meaning Bank takes another route: it projects the English graph onto the translation. The Spanish sentence ends up described by jaywalk, person, street: English words (red).
BMR
BMR simplifies the structure (one node for football player) but also uses the English graph as the interlingua. And jaywalk has no single Spanish word: Spanish says cruzar de manera imprudente.
MSL
MSL uses the sentence's own words. cruzar with :manner imprudente, and futbolista as written. Relations are explicit and need no repository: :agent, :theme, :destination.
Every node is a span of the sentence, so no word sense disambiguation is needed to build it.
Parallel sentences, different graphs
BMR assumed one graph for every translation. Real translations disagree. The same sentence in Catalan mentions the pedestrian crossing; in Korean the jaywalking is one verb; in Chinese the main event is walking towards the bar; in Arabic the crossing itself is called an offence.
So MSL is not an interlingua. It is a layer that lets each language express its own structure, in its own words. Switch languages:
El futbolista del Màlaga CF va travessar el carrer fora del pas de vianants cap al xiringuito.
Building the dataset in seven steps
The recipe starts from AMR 3.0, escapes its licence by parsing free text, fixes the silver graphs by hand, and then teaches an LLM to repeat those fixes and to project graphs into new languages. Press play, or step through it:
AMR 3.0: 59,255 English sentences with gold graphs, all under the LDC licence.
1language
59,255English graphs
For evaluation, native speakers annotated 100 parallel sentences in each of the 11 languages from scratch. In Spanish, where annotators overlapped, their agreement reached 92.34 SMATCH.
Does each step help?
We trained a CLAP parser (mT5-large) on the output of each step. The silver data matches the data it was distilled from; the LLM corrections (MSL_HQ) lift the five corrected languages by 17–24 SMATCH points; projecting to new languages (MSL_HQE) makes Arabic, Korean and Chinese parseable at all.
Show the values as a table
| MSL_AMR | MSL_Silver | MSL_HQ | MSL_HQE | |
|---|---|---|---|---|
| Arabicadded in step 7 | 19.4 | 20.1 | 19.2 | 56.4 |
| Catalanadded in step 7 | 38.4 | 37.4 | 56.3 | 72.4 |
| German | 48.9 | 48.8 | 67.2 | 66.9 |
| English | 54.3 | 55.1 | 72.0 | 71.3 |
| Spanish | 49.7 | 49.3 | 71.9 | 72.9 |
| Koreanadded in step 7 | 26.5 | 27.0 | 35.0 | 56.4 |
| French | 49.0 | 51.5 | 72.3 | 71.9 |
| Galicianadded in step 7 | 42.4 | 41.1 | 57.5 | 69.3 |
| Italian | 46.5 | 47.3 | 71.5 | 71.8 |
| Portugueseadded in step 7 | 40.7 | 41.2 | 58.4 | 72.3 |
| Chineseadded in step 7 | 19.6 | 20.1 | 30.0 | 58.4 |
Spread across languages (SD)
The final dataset brings the gap between languages down to 6.5 SMATCH points. For comparison, a state-of-the-art multilingual AMR parser drops about 9 points from English to German, Spanish or Italian.
Finally, the head-to-head: parse a sentence and generate it back from the graph. The more information a graph keeps, the closer the regenerated sentence. MSL beats AMR by about 19 BLEU and BMR by about 15 on the AMR test set, and holds up out of domain, where AMR and BMR drop.
Show the values as a table
| AMR | BMR | MSL | |
|---|---|---|---|
| German | 21.6 | 27.2 | 41.8 |
| English | 32.4 | 39.0 | 51.4 |
| Spanish | 31.0 | 36.7 | 52.8 |
| Italian | 29.0 | 29.3 | 42.6 |
A layer, not a formalism
Case study · TED
English
“…a raindrop the size of an actual cat or dog when we hear ‘it’s raining cats and dogs’… the dog has to be a small one – a cocker spaniel, or a dachshund…”
Spanish
“…gotas de lluvia del tamaño de un cántaro cuando escuchamos ‘llueve a cántaros’… el cántaro debe ser uno muy pequeño; un botijo, un tarro…”
The idioms are equivalent, but the English text talks about dog breeds and the Spanish about jars. AMR would read the idiom literally, and a single interlingua graph cannot fit both. Languages model ideas; ideas are not the same graph in every language.
Entity Typing
what kind of thing each entity is
Entity Linking
which Wikipedia entity: Màlaga CF
Word Sense Disambiguation
which meaning of each word: BabelNet, WordNet
MSL
who did what to whom, in the sentence’s own words
MSL is deliberately not a complete meaning representation. It extracts relations between concepts and leaves what each concept means to other layers.
Stack word sense disambiguation on top and you get BMR-like graphs; add predicate–argument structures and you get AMR or UMR. Because nodes map one-to-one to sentence spans, MSL could even be parsed with encoder-only models, cheaper than encoder–decoders.
Limitations
- New languages still need a manually annotated test set.
- Two steps relied on a commercial LLM API; the resulting data is released, so they need not be repeated, and open models could replace it.
- SMATCH scores are not comparable across formalisms, so MSL and AMR numbers should not be read side by side.
Cite
@inproceedings{martinez-lorenzo-etal-2024-mitigating,
title = "Mitigating Data Scarcity in Semantic Parsing across Languages with the Multilingual Semantic Layer and its Dataset",
author = "Martinez Lorenzo, Abelardo Carlos and
Huguet Cabot, Pere-Llu{\'i}s and
Ghonim, Karim and
Xu, Lu and
Choi, Hee-Soo and
Fern{\'a}ndez-Castro, Alberte and
Navigli, Roberto",
booktitle = "Findings of the Association for Computational Linguistics: ACL 2024",
month = aug,
year = "2024",
address = "Bangkok, Thailand",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.findings-acl.836/",
doi = "10.18653/v1/2024.findings-acl.836",
pages = "14056--14080"
}Graphs redrawn from Figures 1–3; numbers from Tables 1–4.