Skip to content
Abelardo Carlos
← Projects

The World Bank · 2024 – present

ImpactAI

Better evidence. Better decisions.

A platform that turns impact-evaluation literature into structured, searchable evidence, and answers policy questions with retrieval, grouping and meta-analysis that stay linked to the studies behind them.

ImpactAI in 2 minutes: one question, from understanding to a sourced summary. An animated walkthrough of the product flow.

Architecture

Two systems, one evidence base

Offline, literature becomes a structured, normalized base. Online, a question travels through it. Generating text is one step of a longer chain of representation, retrieval and statistics. I planned and led the whole system; the colours show where I also built it myself. Click any step.

OFFLINE · BUILD THE EVIDENCE BASE1Data collectionplanned & led2PDF parsingworked on it3Information extractionworked on it4Taxonomyworked on it5Store & indexbuilt it aloneScientific database + evidence and passage indicespapers · estimates · taxonomy · provenanceworked on itONLINE · ANSWER A QUESTION1Orchestratorbuilt it alone2Understandbuilt it alone3Entity linkingbuilt it alone4Retrieve & judgeplanned & led5Group on the flybuilt it alone6Data analysisplanned & led7RAGplanned & led8Final answerplanned & led
My partbuilt it aloneI designed and implemented itworked on itI worked on it with the teamplanned & ledI planned it and led the people who built it

What was mine

Architecture, taxonomy, linking, and the team

  • Designed the data and AI architecture that connects offline corpus preparation with online analysis.
  • Built the orchestrator, question understanding, entity linking, on-the-fly grouping and the store & index layer myself; worked on PDF parsing, information extraction and the taxonomy.
  • Fine-tuned the entity-linking bi-encoder and reader, and fine-tuned Gemini on Vertex AI for answers from statistics.
  • Used LLMs to generate synthetic data: to build the taxonomy and to create the training data for fine-tuning.
  • Built and deployed the services on Google Cloud, including the real-time research flow.
  • Lead the AI science team of six: architecture, core backend, code review, evaluation design and cloud budgets.
The fine-tuned linker in depth →

What the team built with me

ImpactAI is a product of a multidisciplinary team at the World Bank: economists, annotators, engineers and product people. Services such as extraction, retrieval, grouping, analysis and the web app have their own owners; this page explains the whole system and marks my components.

PythonFastAPIReactMySQLPolarsQdrantVertex AIGeminiCloud RunCloud Storage

Part 1 · Offline

Building the evidence base

Before anyone asks a question, the literature is collected, read, structured, normalized and indexed. Five steps turn papers into a scientific database.

planned & led

Collect a broad corpus of development-economics literature from journals, conferences and repositories; a classifier keeps what is relevant and eligible, with its origin recorded.

SOURCESJournalsConferencesWorking papersRepositoriesseen 1 · kept 0 · rejected 0SCRAPERS · APIs› fetch› deduplicate› read metadata1. Domain classifierdevelopment economics?✕ not dev. economics02. Eligibility classifiercausal design + estimates?✕ no impact estimate0CORPUS0each paper keeps its provenancesource · URLretrieved · hashIllustrative stream: venues, years and verdicts are examples, not the production corpus.

Part 2 · Online

Answering a question

When a question arrives, an orchestrator runs eight steps: understand it, find and judge the evidence, group it, analyze it, gather context and write a sourced answer.

built it alone

One Python research flow calls each service in order, validates every hand-off and streams progress to the user.

AppOrchestratorUnderstandRetrievalRelevanceSQLGroupingAnalysisRAGAnswerPOST /research✓ hand-off validated against its contract · pink arrows: progress events streamed to the app

Part 3 · Platform

Running it on Google Cloud

The whole system at a glance, then how the app talks to the services, how it is deployed, and how the design evolved.

SOURCESJournalsWorking papersRepositoriesGOOGLE CLOUDOFFLINE PIPELINE · BUILD THE EVIDENCE BASECollectParse PDFsExtractTaxonomyIndexbatch jobs · LLM extraction and embeddings via Vertex AIEVIDENCE BASECloud SQL · scientific basepapers · estimates · taxonomyCloud Storagevector indices · Parquet · filesONLINE · CLOUD RUN SERVICESAPI +orchestratorUnderstandRetrieve & judgeGroupAnalyzeRAGAnswerreadUsersstaff · researchers · partnersFrontend · Cloud RunReact + TypeScriptCloud SQL · product dataconversations · sources · feedbackVertex AIembeddings · Geminifine-tuning jobs(bi-encoder, Gemini)also used by theonline servicesCloud Build · CI/CDGitHub → deploydeploys every service

System architecture, simplified. Click the offline or online blocks to jump to their steps.

planned & led

A React and TypeScript app for asking, following progress and exploring evidence; a Python backend with API contracts, persistence and streaming.

REACT + TYPESCRIPT APPCash transfers and child malnutritionin Sub-Saharan Africa?UnderstandingRetrieving evidenceJudging relevanceGroupingAnalysingWriting the answerwaitingPOST question1SSE progress3PYTHON BACKENDAPItyped contracts (schemas)auth · conversationsserver-sent eventsOrchestratorcalls the services in order2SERVICES · ONE CONTAINER EACHUnderstandingCloud RunRetrievalCloud RunRelevanceCloud RunGroupingCloud RunData analysisCloud RunRAGCloud RunSynthesisCloud RunProduct databasemessages · sources · feedback41. the app posts the question2. the orchestrator calls each service3. each finished step is streamed back4. answer, sources and feedback are savedSimplified architecture; service names follow the research flow.

Visuals are conceptual and use illustrative examples: no counts, shares or effect sizes on this page come from the ImpactAI database. The pooling in “Data analysis” is the real formula applied to made-up estimates.

Models

Many models, each with one job

  • Embeddings + clusteringorganise concepts and normalise the base
  • Adapted ReLiKlink text to the domain taxonomy (first route)
  • LLM NER and linkingthe later entity-linking service
  • Semantic retrieval + relevancefind and judge evidence per estimate
  • Grouping strategiesLLM, embeddings, or both, configurable
  • Gemini fine-tuning pipelinequestion + statistics → answer examples, trained on Vertex AI

Evaluation

A separate system, not an afterthought

Evaluation lives outside the runtime: benchmarks, corpus snapshots, retrieval, trace capture, per-component review and exports for inspection. Each question keeps its unit: finding known papers, retrieving estimates, judging relevance, assigning groups and writing a correct answer are different tasks with different metrics, and automatic judges are labelled apart from human review.

Papers found
Estimates retrieved
Relevance
Grouping
Answer

Based on public material, the code I worked on and my own account; not a record of the current deployment.