The World Bank · 2024 – present
ImpactAI
Better evidence. Better decisions.
A platform that turns impact-evaluation literature into structured, searchable evidence, and answers policy questions with retrieval, grouping and meta-analysis that stay linked to the studies behind them.
Architecture
Two systems, one evidence base
Offline, literature becomes a structured, normalized base. Online, a question travels through it. Generating text is one step of a longer chain of representation, retrieval and statistics. I planned and led the whole system; the colours show where I also built it myself. Click any step.
What was mine
Architecture, taxonomy, linking, and the team
- Designed the data and AI architecture that connects offline corpus preparation with online analysis.
- Built the orchestrator, question understanding, entity linking, on-the-fly grouping and the store & index layer myself; worked on PDF parsing, information extraction and the taxonomy.
- Fine-tuned the entity-linking bi-encoder and reader, and fine-tuned Gemini on Vertex AI for answers from statistics.
- Used LLMs to generate synthetic data: to build the taxonomy and to create the training data for fine-tuning.
- Built and deployed the services on Google Cloud, including the real-time research flow.
- Lead the AI science team of six: architecture, core backend, code review, evaluation design and cloud budgets.
What the team built with me
ImpactAI is a product of a multidisciplinary team at the World Bank: economists, annotators, engineers and product people. Services such as extraction, retrieval, grouping, analysis and the web app have their own owners; this page explains the whole system and marks my components.
Part 1 · Offline
Building the evidence base
Before anyone asks a question, the literature is collected, read, structured, normalized and indexed. Five steps turn papers into a scientific database.
Collect a broad corpus of development-economics literature from journals, conferences and repositories; a classifier keeps what is relevant and eligible, with its origin recorded.
Part 2 · Online
Answering a question
When a question arrives, an orchestrator runs eight steps: understand it, find and judge the evidence, group it, analyze it, gather context and write a sourced answer.
One Python research flow calls each service in order, validates every hand-off and streams progress to the user.
Part 3 · Platform
Running it on Google Cloud
The whole system at a glance, then how the app talks to the services, how it is deployed, and how the design evolved.
System architecture, simplified. Click the offline or online blocks to jump to their steps.
A React and TypeScript app for asking, following progress and exploring evidence; a Python backend with API contracts, persistence and streaming.
Visuals are conceptual and use illustrative examples: no counts, shares or effect sizes on this page come from the ImpactAI database. The pooling in “Data analysis” is the real formula applied to made-up estimates.
Models
Many models, each with one job
- Embeddings + clusteringorganise concepts and normalise the base
- Adapted ReLiKlink text to the domain taxonomy (first route)
- LLM NER and linkingthe later entity-linking service
- Semantic retrieval + relevancefind and judge evidence per estimate
- Grouping strategiesLLM, embeddings, or both, configurable
- Gemini fine-tuning pipelinequestion + statistics → answer examples, trained on Vertex AI
Evaluation
A separate system, not an afterthought
Evaluation lives outside the runtime: benchmarks, corpus snapshots, retrieval, trace capture, per-component review and exports for inspection. Each question keeps its unit: finding known papers, retrieving estimates, judging relevance, assigning groups and writing a correct answer are different tasks with different metrics, and automatic judges are labelled apart from human review.
Based on public material, the code I worked on and my own account; not a record of the current deployment.