This repository contains the dataset, codebase, and evaluation framework for the paper:
"Benchmarking contextual understanding for in-car conversational systems"
Authors: Philipp Habicht, Lev Sorokin, Abdullah Saydemir, Ken E. Friedl, Andrea Stocco.
Note
Note on Implementation: This single, unified repository provides the minimal, verified codebase and end-to-end execution scripts required to reproduce the paper's exact benchmark results. Unlike development or legacy repositories, it eliminates stale dependencies and contains strictly the necessary code, scripts, and datasets needed for lightweight execution.
@misc{habicht2025benchmarking,
title = {Benchmarking Contextual Understanding for In-Car Conversational Systems},
author = {Philipp Habicht and Lev Sorokin and Abdullah Saydemir and Ken Friedl and Andrea Stocco},
journal = {Journal of Systems and Software},
volume = {240},
pages = {112915},
year = {2026},
issn = {0164-1212},
url = {https://www.sciencedirect.com/science/article/pii/S0164121226001482},
google_scholar_id={qjMakFHDy7sC},
code={https://github.com/saydemr/judgebench},
pdf={https://www.sciencedirect.com/science/article/pii/S0164121226001482/pdfft},
}