InteractiveNAACL 2024
MOSAICo: four layers of meaning over Wikipedia, in five languages
Simone Conia, Edoardo Barba, Abelardo Carlos Martínez Lorenzo, Pere-Lluís Huguet Cabot, Riccardo Orlando, Luigi Procopio, Roberto Navigli · all authors contributed equally
Models that use explicit semantics need data labelled with it, and the best datasets are small, English and closed. MOSAICo annotates the same Wikipedia text with word senses, semantic roles, AMR graphs and relations, openly, so that the layers can be studied together.
- silver annotations (sum of Table 2)
- ~900Msilver annotations (sum of Table 2)
- tasks: WSD, SRL, AMR parsing, relation extraction
- 4tasks: WSD, SRL, AMR parsing, relation extraction
- Wikipedia articles per language, all present in all five
- 441KWikipedia articles per language, all present in all five
- GPU hours of annotation, released for free
- 12,470GPU hours of annotation, released for free
Two ways to put meaning into NLP
Implicit
Scale self-supervised language models and let meaning emerge. It works on benchmarks, but it is expensive, opaque, and it is unclear how much the models actually understand.
Explicit
Give models discrete symbols: word senses, predicate structures, graphs. This can reduce parameters and make outputs interpretable. What has been missing is a vast, open dataset annotated with those symbols.
One sentence, four layers
My cat ate the mouse that I bought at the Apple store when I traveled to Seattle last summer.
Plain text: no symbols yet.
How it was built
MOSAICo keeps only Wikipedia articles that exist in all five languages (English, German, Spanish, French, Italian): 441,000 per language, comparable by design. A higher-quality subset, MOSAICo Core, keeps the 17,200 articles marked “good” or “featured” in at least one language. Each layer comes from a state-of-the-art system, several of them from our own group.
- PrepStanza · sentence splitting, tokenisation, lemmas, POS1,050 h
- WSDESCHER · WordNet / BabelNet 5.1 senses, DeBERTa-v3 (mDeBERTa for non-English)1,140 h
- SRLMulti-SRL · span-based, PropBank and VerbAtlas labels at once1,600 h
- SPLeakDistill + CLAP · LeakDistill for English, cross-lingual CLAP (mBART) for the rest5,080 h
- REmREBEL + cRocoDiLe · predicted triplets + Wikipedia links × Wikidata, filtered by an NLI critic3,600 h
For AMR, English uses LeakDistill and the other languages a cross-lingual version of CLAP, trained on translated AMR 3.0.
Can silver compete with gold?
We trained the same systems three ways: on the usual gold data, on a MOSAICo sample of the same size (M-Ref), and on MOSAICo Core (M-Core). Same-size silver trails gold, as expected; the larger Core catches up or overtakes it, and usually generalises better out of domain.
Each row has its own axis: compare dots within a row, not distances across rows. The number on the right is M-Core minus the best other.
Show the values as a table
| gold | M-Ref | M-Core | |
|---|---|---|---|
| ALLEnglish, standard | 81.0 | 79.0 | 82.0 |
| 42DEnglish, rare senses | 54.4 | 51.5 | 56.2 |
| XL-WSD DE | 83.2 | 81.2 | 84.1 |
| XL-WSD ES | 77.5 | 76.8 | 77.8 |
| XL-WSD FR | 84.3 | 83.8 | 84.5 |
| XL-WSD IT | 78.2 | 77.3 | 79.1 |
Takeaway
Trained on MOSAICo Core, ESCHER beats its gold-trained self on every test set and language (77.3 vs 76.4 average). On rare senses, the full corpus goes further still: 58.4 on 42D.
When the layers meet
Four layers on the same text make new questions answerable:
6.7M
A free WSD benchmark: Wiki-WSD
Wikipedia editors link the first mention of a term to its article. Mapped to BabelNet, those links become sense labels: 6.7 million instances, versus 14,166 in XL-WSD.
74%
WSD and SRL disagree more than expected
VerbAtlas frames are clusters of BabelNet synsets, so each verb gets a frame from both systems. They agree only 74% of the time, exposing part-of-speech errors and gaps in the inventories themselves.
95.5%
SRL and AMR mostly agree
Using the cross-lingual aligner, 5.6M AMR predicates line up with SRL predicates. The two systems pick the same PropBank sense 95.5% of the time and the same roles in 92.7% of triplets.
The aligner story →141,128
Facts Wikidata did not have
With WSD, relations can link concepts, not only named entities: nearly half the RE triplets come from it. The English Core alone holds 141,128 facts missing from Wikidata. One, that Euler worked in graph theory, was added to Wikidata months later.
Limitations
- Annotations are silver: quality is shown through downstream results, not a direct manual evaluation.
- Five high-resource languages; low-resource languages are thin on Wikipedia.
- Encyclopedic style and topics can bias models, though out-of-domain results suggest better generalisation.
Cite
@inproceedings{conia-etal-2024-mosaico,
title = "{MOSAIC}o: a Multilingual Open-text Semantically Annotated Interlinked Corpus",
author = "Conia, Simone and
Barba, Edoardo and
Martinez Lorenzo, Abelardo Carlos and
Huguet Cabot, Pere-Llu{\'i}s and
Orlando, Riccardo and
Procopio, Luigi and
Navigli, Roberto",
booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
month = jun,
year = "2024",
address = "Mexico City, Mexico",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.naacl-long.442/",
doi = "10.18653/v1/2024.naacl-long.442",
pages = "7990--8004"
}Figure 1 redrawn; numbers from Tables 1–8.