PROJECT / 01Applied AI · Retrieval · Evaluation

A document-grounded RAG system at OeNB

Accurate domain answers depend on retrieving the right source passages before the language model generates a response.

Role
AI Engineer Intern
Context
Oesterreichische Nationalbank (OeNB) · Vienna
Period
Jul 2024 — Mar 2025
INTERACTIVE EVIDENCE

Explore the full pipeline.

A specialized chatbot workflow that grounds answers in an OeNB document collection and measures whether retrieval finds the right evidence.

RAG PIPELINE / SELECT A STAGE Ingest → retrieve → generate

Runtime pipeline

01

Ingest

RETRIEVAL EVALUATION

Coverage and ranking fail differently, so I scored them separately.

SchematicThe eight passages, the ranking, and both scores below are a worked example, sized to make the two metrics legible. The method is the point: benchmark questions with known relevant passages, then recall for coverage and NDCG for ranking.

RECALL4 of 5 relevant passages found
0.80
NDCGRetrieved gain against the best possible order
0.98
THE QUESTION

How can a specialized chatbot answer questions accurately within a defined domain using that domain’s document collection?

HOW I APPROACHED IT
  1. 01

    Collected and processed PDF, TXT, and DOCX sources into a searchable corpus, storing document chunks, embeddings, and source metadata.

  2. 02

    Created benchmark questions with known relevant source passages so retrieval quality could be measured against explicit relevance labels.

  3. 03

    Compared retrieval configurations using recall for relevant-passage coverage and NDCG for ranking quality.

  4. 04

    Passed the strongest retrieved context into Hugging Face and vLLM workflows, then explored prompting, reflection, few-shot examples, and fine-tuning.

The call I made
Chose
Measured retrieval on its own, against labelled passages
Instead of
Judging the system only by its final generated answer
Because
If the right passage is never retrieved, no amount of prompting repairs the answer. Scoring retrieval separately says whether a bad answer is a retrieval failure or a generation failure.
FormatsPDF · TXT · DOCX
Model runtimesHugging Face · vLLM
Retrieval evaluationRecall · NDCG
TOOLS & METHODS
  • Python
  • PostgreSQL
  • Hugging Face
  • vLLM
  • NLTK
  • Stanza