Skip to content
Abelardo Carlos

InteractiveFindings of ACL 2024

The Multilingual Semantic Layer: semantic parsing in eleven languages

Abelardo Carlos Martínez Lorenzo, Pere-Lluís Huguet Cabot, Karim Ghonim, Lu Xu, Hee-Soo Choi, Alberte Fernández Castro, Roberto Navigli

Semantic parsers need annotated graphs, and annotated graphs barely exist outside English. We designed a simpler layer that any language can express in its own words, and built a dataset for eleven languages on an academic budget, with LLMs doing the heavy lifting and humans checking the work.

languages, including Arabic, Korean and Chinese
11languages, including Arabic, Korean and Chinese
silver graphs parsed from license-free text
10M+silver graphs parsed from license-free text
gold graphs annotated by hand as a benchmark
1,100gold graphs annotated by hand as a benchmark
per graph for the LLM steps, about $400 in total
$0.0013per graph for the LLM steps, about $400 in total
01

Semantic parsing has an English problem

Graph formalisms like AMR turn a sentence into concepts and relations that machines can use. But parsers learn from annotated graphs, and annotating is slow: annotators must master the formalism, its repositories and each language's rules.

The result: the big corpora are English, sit behind licences, and efforts in Turkish, Persian, Portuguese or Vietnamese have produced no more than about 200 graphs per language, far too few to train a parser.

Semantic parsing corpora (Table 1) sentences in the largest corpus of each formalism

DRT

1993

16,712
6,000

PDT

2003

49,431
49,431

UCCA

2013

7,934
0

AMR

2013

59,255
0

UMR

2021

2,186
1,993

BMR

2022

59,255
0
all sentencesnon-English annotations
“Non-English” counts annotations whose graph was built for a non-English sentence. BMR 1.0 covers five languages, but through the English graph.
02

One sentence, five formalisms

English AMR Figure 1
Thefootballplayerjaywalksacrossthestreettothechiringuito.:ARG0:ARG1:destination:ARG0-of:ARG1jaywalk-00personstreetchiringuitoplay-01football
frame from a repositorylemma

English AMR

The football player jaywalks across the street to the chiringuito. AMR uses the PropBank frame jaywalk-00 and spells out a football player as a person who plays football.

Spanish AMR

Now the Spanish translation. Spanish AMR needs a Spanish frame inventory (AnCora): cruzar-01, jugar-02. A different repository, a different graph, and almost no annotated data.

Spanish PMB

The Parallel Meaning Bank takes another route: it projects the English graph onto the translation. The Spanish sentence ends up described by jaywalk, person, street: English words (red).

BMR

BMR simplifies the structure (one node for football player) but also uses the English graph as the interlingua. And jaywalk has no single Spanish word: Spanish says cruzar de manera imprudente.

MSL

MSL uses the sentence's own words. cruzar with :manner imprudente, and futbolista as written. Relations are explicit and need no repository: :agent, :theme, :destination.

Every node is a span of the sentence, so no word sense disambiguation is needed to build it.

03

Parallel sentences, different graphs

BMR assumed one graph for every translation. Real translations disagree. The same sentence in Catalan mentions the pedestrian crossing; in Korean the jaywalking is one verb; in Chinese the main event is walking towards the bar; in Arabic the crossing itself is called an offence.

So MSL is not an interlingua. It is a layer that lets each language express its own structure, in its own words. Switch languages:

Catalan MSL Figure 3

El futbolista del Màlaga CF va travessar el carrer fora del pas de vianants cap al xiringuito.

:agent:membership:manner:context:theme:targettravessarfutbolistaMàlaga CFforapas de vianantscarrerxiringuito
Catalan says where he crossed: outside the pedestrian crossing (fora del pas de vianants).
04

Building the dataset in seven steps

The recipe starts from AMR 3.0, escapes its licence by parsing free text, fixes the silver graphs by hand, and then teaches an LLM to repeat those fixes and to project graphs into new languages. Press play, or step through it:

Step 0 · AMR 3.0

AMR 3.0: 59,255 English sentences with gold graphs, all under the LDC licence.

ONE EXAMPLE PAIRENThe football player jaywalks across thestreet to the chiringuito.jaywalk-00personstreetchiringuitoENThe football player jaywalks across thestreet to the chiringuito.DEDer Fußballspieler überquert verbotenerweisedie Straße zum Chiringuito.ESEl futbolista cruza la calle de maneraimprudente hacia el chiringuito.FRLe footballeur traverse la rue hors dupassage piéton vers le chiringuito.ITIl calciatore attraversa la strada in modoimprudente verso il chiringuito.jaywalkfootball playerstreetchiringuitoüberquerenFußballspielerStraßeChiringuitocruzarfutbolistacallechiringuitotraverserfootballeurruechiringuitoattraversarecalciatorestradachiringuitoTEDOPUS · parallelOpenSubtitlesOPUS · parallelUbuntuOPUS · parallelBibleOPUS · parallelBooksOPUS · parallelWikipedianon-paralleljaywalkfootball playerstreetchiringuito+ tense · aspect · moodüberquerenFußballspielerStraßeChiringuito+ tense · aspect · moodcruzarfutbolistacallechiringuito+ tense · aspect · moodtraverserfootballeurruechiringuito+ tense · aspect · moodattraversarecalciatorestradachiringuito+ tense · aspect · moodAMR 3.0LDC licence
EN59K
DE—
ES—
FR—
IT—
AR—
CA—
GL—
PT—
ZH—
KO—

1language

59,255English graphs

Figures 4–5 and Table 2. Steps 0–3 follow one example sentence (German, French and Italian translations are ours, for illustration). The two LLM steps cost about $400 in total, $0.0013 per graph; the AMR 3.0 licence alone costs $300.

For evaluation, native speakers annotated 100 parallel sentences in each of the 11 languages from scratch. In Spanish, where annotators overlapped, their agreement reached 92.34 SMATCH.

05

Does each step help?

We trained a CLAP parser (mT5-large) on the output of each step. The silver data matches the data it was distilled from; the LLM corrections (MSL_HQ) lift the five corrected languages by 17–24 SMATCH points; projecting to new languages (MSL_HQE) makes Arabic, Korean and Chinese parseable at all.

Parsing (SMATCH) (Table 3) gold test set, 100 sentences per language
MSL_AMRMSL_SilverMSL_HQMSL_HQE
Arabicadded in step 71874MSL_AMR · Arabic: 19.4MSL_Silver · Arabic: 20.1MSL_HQ · Arabic: 19.2MSL_HQE · Arabic: 56.4MSL_HQE 56.4Catalanadded in step 71874MSL_AMR · Catalan: 38.4MSL_Silver · Catalan: 37.4MSL_HQ · Catalan: 56.3MSL_HQE · Catalan: 72.4MSL_HQE 72.4German1874MSL_AMR · German: 48.9MSL_Silver · German: 48.8MSL_HQ · German: 67.2MSL_HQE · German: 66.9MSL_HQ 67.2English1874MSL_AMR · English: 54.3MSL_Silver · English: 55.1MSL_HQ · English: 72.0MSL_HQE · English: 71.3MSL_HQ 72.0Spanish1874MSL_AMR · Spanish: 49.7MSL_Silver · Spanish: 49.3MSL_HQ · Spanish: 71.9MSL_HQE · Spanish: 72.9MSL_HQE 72.9Koreanadded in step 71874MSL_AMR · Korean: 26.5MSL_Silver · Korean: 27.0MSL_HQ · Korean: 35.0MSL_HQE · Korean: 56.4MSL_HQE 56.4French1874MSL_AMR · French: 49.0MSL_Silver · French: 51.5MSL_HQ · French: 72.3MSL_HQE · French: 71.9MSL_HQ 72.3Galicianadded in step 71874MSL_AMR · Galician: 42.4MSL_Silver · Galician: 41.1MSL_HQ · Galician: 57.5MSL_HQE · Galician: 69.3MSL_HQE 69.3Italian1874MSL_AMR · Italian: 46.5MSL_Silver · Italian: 47.3MSL_HQ · Italian: 71.5MSL_HQE · Italian: 71.8MSL_HQE 71.8Portugueseadded in step 71874MSL_AMR · Portuguese: 40.7MSL_Silver · Portuguese: 41.2MSL_HQ · Portuguese: 58.4MSL_HQE · Portuguese: 72.3MSL_HQE 72.3Chineseadded in step 71874MSL_AMR · Chinese: 19.6MSL_Silver · Chinese: 20.1MSL_HQ · Chinese: 30.0MSL_HQE · Chinese: 58.4MSL_HQE 58.4
Show the values as a table
MSL_AMRMSL_SilverMSL_HQMSL_HQE
Arabicadded in step 719.420.119.256.4
Catalanadded in step 738.437.456.372.4
German48.948.867.266.9
English54.355.172.071.3
Spanish49.749.371.972.9
Koreanadded in step 726.527.035.056.4
French49.051.572.371.9
Galicianadded in step 742.441.157.569.3
Italian46.547.371.571.8
Portugueseadded in step 740.741.258.472.3
Chineseadded in step 719.620.130.058.4

Spread across languages (SD)

MSL_AMR
11.81
MSL_Silver
11.86
MSL_HQ
18.13
MSL_HQE
6.48

The final dataset brings the gap between languages down to 6.5 SMATCH points. For comparison, a state-of-the-art multilingual AMR parser drops about 9 points from English to German, Spanish or Italian.

One axis for all rows. Darker = later annotation step. Hover, tap or focus a dot for its value.

Finally, the head-to-head: parse a sentence and generate it back from the graph. The more information a graph keeps, the closer the regenerated sentence. MSL beats AMR by about 19 BLEU and BMR by about 15 on the AMR test set, and holds up out of domain, where AMR and BMR drop.

Back-translation: sentence → graph → sentence (Table 4) BLEU, one shared cross-lingual model per formalism
AMRBMRMSL
German2054AMR · German: 21.6BMR · German: 27.2MSL · German: 41.8MSL 41.8English2054AMR · English: 32.4BMR · English: 39.0MSL · English: 51.4MSL 51.4Spanish2054AMR · Spanish: 31.0BMR · Spanish: 36.7MSL · Spanish: 52.8MSL 52.8Italian2054AMR · Italian: 29.0BMR · Italian: 29.3MSL · Italian: 42.6MSL 42.6
Show the values as a table
AMRBMRMSL
German21.627.241.8
English32.439.051.4
Spanish31.036.752.8
Italian29.029.342.6
AMR test set (4 translations). One axis for all rows.
06

A layer, not a formalism

Case study · TED

English

“…a raindrop the size of an actual cat or dog when we hear ‘it’s raining cats and dogs’… the dog has to be a small one – a cocker spaniel, or a dachshund…”

Spanish

“…gotas de lluvia del tamaño de un cántaro cuando escuchamos ‘llueve a cántaros’… el cántaro debe ser uno muy pequeño; un botijo, un tarro…”

The idioms are equivalent, but the English text talks about dog breeds and the Spanish about jars. AMR would read the idiom literally, and a single interlingua graph cannot fit both. Languages model ideas; ideas are not the same graph in every language.

Entity Typing

what kind of thing each entity is

Entity Linking

which Wikipedia entity: Màlaga CF

Word Sense Disambiguation

which meaning of each word: BabelNet, WordNet

MSL

who did what to whom, in the sentence’s own words

MSL is deliberately not a complete meaning representation. It extracts relations between concepts and leaves what each concept means to other layers.

Stack word sense disambiguation on top and you get BMR-like graphs; add predicate–argument structures and you get AMR or UMR. Because nodes map one-to-one to sentence spans, MSL could even be parsed with encoder-only models, cheaper than encoder–decoders.

Limitations

  • New languages still need a manually annotated test set.
  • Two steps relied on a commercial LLM API; the resulting data is released, so they need not be repeated, and open models could replace it.
  • SMATCH scores are not comparable across formalisms, so MSL and AMR numbers should not be read side by side.
Where MSL came from: the BMR story →

Cite

@inproceedings{martinez-lorenzo-etal-2024-mitigating,
    title = "Mitigating Data Scarcity in Semantic Parsing across Languages with the Multilingual Semantic Layer and its Dataset",
    author = "Martinez Lorenzo, Abelardo Carlos  and
      Huguet Cabot, Pere-Llu{\'i}s  and
      Ghonim, Karim  and
      Xu, Lu  and
      Choi, Hee-Soo  and
      Fern{\'a}ndez-Castro, Alberte  and
      Navigli, Roberto",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2024",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.findings-acl.836/",
    doi = "10.18653/v1/2024.findings-acl.836",
    pages = "14056--14080"
}
All publications

Graphs redrawn from Figures 1–3; numbers from Tables 1–4.