ClustMe and ClustML data S1 and S2

Gaussian Mixture human-labeled data for clustering design and evaluation

CCBY4.0Introduced 2019-07-10

Code and datasets S1 and S2 used in:

  • Mostafa M. Abbas, Michaël Aupetit, Michael Sedlmair, Halima Bensmail. ClustMe: A Visual Quality Measure for Ranking Monochrome Scatterplots based on Cluster Patterns. Computer Graphics Forum 38(3): 225-236 (2019).
    https://doi.org/10.1111/cgf.13684

  • Hamza, M. M., Ullah, E., Baggag, A., Bensmail, H., Sedlmair, M., & Aupetit, M. ClustML: A Measure of Cluster Pattern Complexity in Scatterplots Learnt from Human-labeled Groupings, to appear in SAGE Information Visualization Journal 23(2) 105-122 (2024). https://doi.org/10.1177/14738716231220536 ARXIV: https://arxiv.org/abs/2106.00599

  • M. Aupetit, M. Sedlmair, M. M. Abbas, A. Baggag and H. Bensmail, Toward Perception-Based Evaluation of Clustering Techniques for Visual Analytics 2019 IEEE Visualization Conference (VIS), Vancouver, BC, Canada, 2019, pp. 141-145. doi: 10.1109/VISUAL.2019.8933620 https://ieeexplore.ieee.org/document/8933620

S1: a set of 1000 scatterplot data generated by 2-component Gaussian Mixture Model with various parameters, together with the 8 GMM parameter values and 34 human judgments about perceived separability in scatterplot image: the task was to answer if they see "one" or "more-than-one" clusters.

S2: a set of 435 pairs of scatterplot data generated by projecting high-dimensional data with various dimensionality reduction techniques, together with 31 human judgments about which of the two images in a pair shows a more complex cluster pattern: right is more complex, left is more complex, or both are equally complex.