Home | Publications | JTB+26

Pitfalls With Downsampled Data in AutoML for Clustering

MCML Authors

Abstract

Finding the best-suited hyperparameters for applied algorithms is challenging in practice and in scientific evaluation. AutoML provides a reliable way to search for these. Still, a major issue in AutoML is scalability, as repeated evaluations are time-consuming, especially for large datasets. For classification tasks, downsampling has been shown to be effective for handling large datasets. In AutoML for unsupervised ML tasks such as clustering, recent research has also begun to integrate downsampling. However, finding suitable hyperparameters for clustering is quite underexplored and an even bigger challenge. We describe and demonstrate three key issues to address when applying downsampling to this task. Particularly, sampling may alter the instantiation of the underlying distribution, an aspect to which clustering methods are typically sensitive. Furthermore, we provide guidance for practitioners on mitigating some of these problems. Finally, we establish directions for future research to better support AutoML users for clustering.

inproceedings JTB+26


CIKM 2026

35th ACM International Conference on Information and Knowledge Management. Rome, Italy, Nov 07-11, 2026. To be published.
Conference logo
A Conference

Authors

P. Jahn • G. M. Tavares • A. Beer • U. Schlegel • T. Seidl

Research Area

 A3 | Computational Models

BibTeXKey: JTB+26

Back to Top