Skip to content
Abelardo Carlos

InteractiveNAACL 2024

MOSAICo: four layers of meaning over Wikipedia, in five languages

Simone Conia, Edoardo Barba, Abelardo Carlos Martínez Lorenzo, Pere-Lluís Huguet Cabot, Riccardo Orlando, Luigi Procopio, Roberto Navigli · all authors contributed equally

Models that use explicit semantics need data labelled with it, and the best datasets are small, English and closed. MOSAICo annotates the same Wikipedia text with word senses, semantic roles, AMR graphs and relations, openly, so that the layers can be studied together.

silver annotations (sum of Table 2)
~900Msilver annotations (sum of Table 2)
tasks: WSD, SRL, AMR parsing, relation extraction
4tasks: WSD, SRL, AMR parsing, relation extraction
Wikipedia articles per language, all present in all five
441KWikipedia articles per language, all present in all five
GPU hours of annotation, released for free
12,470GPU hours of annotation, released for free
01

Two ways to put meaning into NLP

Implicit

Scale self-supervised language models and let meaning emerge. It works on benchmarks, but it is expensive, opaque, and it is unclear how much the models actually understand.

Explicit

Give models discrete symbols: word senses, predicate structures, graphs. This can reduce parameters and make outputs interpretable. What has been missing is a vast, open dataset annotated with those symbols.

02

One sentence, four layers

One sentence, four layers of meaning Figure 1
manufacturerlocationMycatatethemousethatIboughtattheApplestorewhenItraveledtoSeattlelastsummercat#1eat#1mouse#4buy#1shop#1travel#1last#1summer#1AGENTPATIENTEAT • BITETHEMEAGENTLOCATIONTIMEBUYTIMEAGENTDESTINATIONTIMETRAVEL
:ARG1:ARG0:ARG1-of:location:mod:name:op1:ARG0:time:poss:ARG0:ARG4:name:op1:time:season:mod:modeat-01mousecatbuy-01istorecompanyname"Apple"travel-01date-entitysummeryearlastcityname"Seattle"
WSD Word Sense DisambiguationSRL Semantic Role LabelingSP Semantic Parsing (AMR)RE Relation Extraction

My cat ate the mouse that I bought at the Apple store when I traveled to Seattle last summer. Plain text: no symbols yet.

03

How it was built

MOSAICo keeps only Wikipedia articles that exist in all five languages (English, German, Spanish, French, Italian): 441,000 per language, comparable by design. A higher-quality subset, MOSAICo Core, keeps the 17,200 articles marked “good” or “featured” in at least one language. Each layer comes from a state-of-the-art system, several of them from our own group.

899 million annotations (Table 2) millions, by language and task
EN
342M
DE
148M
ES
99M
FR
161M
IT
149M
WSD · 522.3MSRL · 285.3MSP · 51.7MRE · 39.3M
12,470 GPU hours you don’t have to spend footnote 5, RTX 3090
  • PrepStanza · sentence splitting, tokenisation, lemmas, POS1,050 h
  • WSDESCHER · WordNet / BabelNet 5.1 senses, DeBERTa-v3 (mDeBERTa for non-English)1,140 h
  • SRLMulti-SRL · span-based, PropBank and VerbAtlas labels at once1,600 h
  • SPLeakDistill + CLAP · LeakDistill for English, cross-lingual CLAP (mBART) for the rest5,080 h
  • REmREBEL + cRocoDiLe · predicted triplets + Wikipedia links × Wikidata, filtered by an NLI critic3,600 h

For AMR, English uses LeakDistill and the other languages a cross-lingual version of CLAP, trained on translated AMR 3.0.

04

Can silver compete with gold?

We trained the same systems three ways: on the usual gold data, on a MOSAICo sample of the same size (M-Ref), and on MOSAICo Core (M-Core). Same-size silver trails gold, as expected; the larger Core catches up or overtakes it, and usually generalises better out of domain.

ESCHER trained on gold vs MOSAICo (F1)
goldM-RefM-Core
ALLEnglish, standard7883gold · ALL: 81.0M-Ref · ALL: 79.0M-Core · ALL: 82.0M-Core 82.0+1.042DEnglish, rare senses5058gold · 42D: 54.4M-Ref · 42D: 51.5M-Core · 42D: 56.2M-Core 56.2+1.8XL-WSD DE8086gold · XL-WSD DE: 83.2M-Ref · XL-WSD DE: 81.2M-Core · XL-WSD DE: 84.1M-Core 84.1+0.9XL-WSD ES7579gold · XL-WSD ES: 77.5M-Ref · XL-WSD ES: 76.8M-Core · XL-WSD ES: 77.8M-Core 77.8+0.3XL-WSD FR8286gold · XL-WSD FR: 84.3M-Ref · XL-WSD FR: 83.8M-Core · XL-WSD FR: 84.5M-Core 84.5+0.2XL-WSD IT7681gold · XL-WSD IT: 78.2M-Ref · XL-WSD IT: 77.3M-Core · XL-WSD IT: 79.1M-Core 79.1+0.9

Each row has its own axis: compare dots within a row, not distances across rows. The number on the right is M-Core minus the best other.

Show the values as a table
goldM-RefM-Core
ALLEnglish, standard81.079.082.0
42DEnglish, rare senses54.451.556.2
XL-WSD DE83.281.284.1
XL-WSD ES77.576.877.8
XL-WSD FR84.383.884.5
XL-WSD IT78.277.379.1

Takeaway

Trained on MOSAICo Core, ESCHER beats its gold-trained self on every test set and language (77.3 vs 76.4 average). On rare senses, the full corpus goes further still: 58.4 on 42D.

Benchmarks differ in scale, so each row is zoomed on its own. Hover, tap or focus a dot for its value.
05

When the layers meet

Four layers on the same text make new questions answerable:

WSD

6.7M

A free WSD benchmark: Wiki-WSD

Wikipedia editors link the first mention of a term to its article. Mapped to BabelNet, those links become sense labels: 6.7 million instances, versus 14,166 in XL-WSD.

EN 88.4IT 85.9FR 82.8DE 86.3ES 83.1
WSDSRL

74%

WSD and SRL disagree more than expected

VerbAtlas frames are clusters of BabelNet synsets, so each verb gets a frame from both systems. They agree only 74% of the time, exposing part-of-speech errors and gaps in the inventories themselves.

SRLSP

95.5%

SRL and AMR mostly agree

Using the cross-lingual aligner, 5.6M AMR predicates line up with SRL predicates. The two systems pick the same PropBank sense 95.5% of the time and the same roles in 92.7% of triplets.

The aligner story →
REWSD

141,128

Facts Wikidata did not have

With WSD, relations can link concepts, not only named entities: nearly half the RE triplets come from it. The English Core alone holds 141,128 facts missing from Wikidata. One, that Euler worked in graph theory, was added to Wikidata months later.

Limitations

  • Annotations are silver: quality is shown through downstream results, not a direct manual evaluation.
  • Five high-resource languages; low-resource languages are thin on Wikipedia.
  • Encyclopedic style and topics can bias models, though out-of-domain results suggest better generalisation.
Eleven languages, with human corrections: the MSL story →

Cite

@inproceedings{conia-etal-2024-mosaico,
    title = "{MOSAIC}o: a Multilingual Open-text Semantically Annotated Interlinked Corpus",
    author = "Conia, Simone  and
      Barba, Edoardo  and
      Martinez Lorenzo, Abelardo Carlos  and
      Huguet Cabot, Pere-Llu{\'i}s  and
      Orlando, Riccardo  and
      Procopio, Luigi  and
      Navigli, Roberto",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.442/",
    doi = "10.18653/v1/2024.naacl-long.442",
    pages = "7990--8004"
}
All publications

Figure 1 redrawn; numbers from Tables 1–8.