Author: Matthias Hofmaier
Supervisor: Julia Neidhardt, Co-Supervisor: Thomas E. Kolb
Abstract
Podcasts have become a popular medium in the last decade. The huge amount of data available motivates research on podcast recommender systems to make this data accessible to users. Since interaction-based datasets for podcasts are only available to the large streaming providers, content-based methods are needed to build a recommender system. Building content-based recommender systems is closely related to the field of information retrieval, which is well studied in the podcast domain. However, most of this research examines retrieval based on the textual transcription of the audio file and does not compare the effectiveness of other representations such as metadata or audio features, which represents a gap in the research. Podcasts are often referred to as the auditory counterpart of textual media such as news articles, and using transcriptions also connects these different types of media in the way their content is represented. This similarity in media content motivates research on cross-domain recommender systems, which aim to use information from one source domain, such as podcasts, to generate recommendations in other domains. However, no such system in the podcast domain has been published in research. To address these research gaps, we investigate how to build and evaluate a content-based cross-domain recommender system between podcasts and news articles without the availability of interaction-based data. Furthermore, in this work, we investigate how different attributes used to represent podcasts in a recommender system scenario affect the performance of the system. This is done by creating a manually annotated cross-domain dataset between podcast segments and news articles. Using this dataset, we fit several models, each using a different set of podcast representations, with the goal of generating news article recommendations given a particular podcast segment. Performing an evaluation in terms of ranking quality measures and the beyond-accuracy measure of topical diversity, we find that our approach outperforms four different baseline models in terms of ranking quality. However, we also find that this increase in ranking quality comes at the expense of recommendation diversity. Moreover, contrary to our prior beliefs, we do not observe a particular set of podcast features that outperform all others, but rather a large difference in performance between different podcast shows, which motivates further investigation into the properties of shows that might explain these differences.
