Evaluating Agentic Retrieval Augmented Generation in Open Tender Evaluations

Author: Franz Ottitsch

Supervisor: Julia Neidhardt, Co-Supervisor: Thomas E. Kolb

Abstract

Public tender evaluation requires contracting authorities to assess bidder responses against defined award criteria across long, semi structured documents, a process that remains largely manual, time consuming and costly. Retrieval augmented generation (RAG) systems, particularly agentic variants with iterative retrieval and task decomposition, offer a path to reducing this effort, but no procurement oriented evaluation framework for agentic retrieval augmented generation exists that produces measurable, reproducible scores validated against expert judgement. This thesis develops and empirically validates such a framework for agentic RAG applied to the assessment of bidder responses in public tender procedures. On an anonymised corpus of tender documents, a baseline RAG setup and two agentic configurations are compared in a shared evaluation harness along multiple evaluation dimensions reflecting procurement requirements for verifiability, completeness and auditability. The study tests whether the framework reliably discriminates between configurations, quantifies the effect of underlying language models within a fixed pipeline, and measures the correlation between automated scores and expert human judgement.