Skip to content
Abelardo Carlos
← Projects

Research and ImpactAI · 2019 – present

Taxonomies and knowledge graphs

Structuring a field’s knowledge in three layers: concepts (a taxonomy grown from millions of documents with LLMs and experts), relations (typed graphs from semantic parsing, in any language), and completion (missing links found by reasoning and by effect prediction).

Part 1 · Concepts

A taxonomy grown from millions of documents

In production · ImpactAIThe taxonomy loop runs on ImpactAI’s corpus; the points and names in the figure are illustrative.

Start from a domain and every document in it. Information extraction leaves a universe of mentions (interventions, outcomes, populations, countries) each written in its author’s own words. We sort them by type, embed them, try several clustering algorithms, and send large batches to an LLM that partitions and names the groups. Experts review, the borderline points are regrouped, and the rounds continue until the partition stops moving.

Step 1

Start from one domain and everything extracted from its papers: every intervention, outcome, population and setting, as written by each author.

every dot is an element extracted from a paper

Illustrative points and names. Algorithms from the ImpactAI toolkit (HDBSCAN, K-Means, DBSCAN, agglomerative); embeddings from Sentence-Transformers or Vertex AI.

Per dimension

Interventions, outcomes and populations are clustered separately, each with its own granularity.

Experts in the loop

Their corrections to names, definitions and borders steer the next round, until nothing moves.

Kept traceable

Every original mention maps to its canonical concept, so records and concepts stay linked.

Part 2 · Relations

From concepts to an ontology: who did what, for whom, where

Research · BMR and MSL, applied to evidenceThe formalisms are published work; the sentence and graphs in the figure are illustrative.

A taxonomy says what things are; an ontology says how they relate. Semantic parsing turns each sentence into a graph of events and roles, in any language, following the formalisms from my PhD (BMR, and MSL, which keeps each language’s own words under one shared structure). Mapping those roles onto the taxonomy and onto relation types fixed with experts gives typed triples, and the triples of thousands of papers join into one connected graph.

Step 1

Concepts alone are a list. Papers say how they relate: who implemented what, for whom, where, and with what result.

ENThe Ministry of Education implemented a conditional cash transfer for poor rural households in Mexico, and school enrolment rose.

InterventionOutcomePopulationCountry / settingOrganisation

Illustrative sentence. Relation extraction follows the MSL/BMR line of work (semantic graphs in each language’s own words); relation types are defined with domain experts.

Part 3 · Completion

Filling in the missing links

Research · Poderoso and Query2EffectPublished frameworks; the graph, scores and effect in the figure are illustrative.

No graph built from papers is complete. Two things add links: reasoning, where embeddings propose triples and expert rules reject or deduce them in a loop (Poderoso, first designed during my master’s at TU Berlin); and prediction, where an encoder trained on thousands of trials estimates the effect of an intervention on an outcome that has not been studied together (Query2Effect).

Step 1

The graph from Part 2 is never complete: papers report some relations and leave others implicit, or no paper has studied them yet.

targetslocated_intargetsaffectsaffects targets affects located_in (deduced)affects · g = −0.08 [−0.15, −0.01]Conditional cash transfersScholarshipsRural householdsAdolescent girlsMexicoSchool enrolmentEarly marriage

Candidate links

  • CCTs → affects → School enrolment
    proposed
  • CCTs → targets → Adolescent girls
    proposed
  • Mexico → affects → Early marriage
    proposed

Illustrative graph, scores and effect. Embeddings + reasoning follow Poderoso (VLDB workshops 2025); the effect predictor is the Query2Effect encoder (arXiv 2026).