Classical clustering assumes that a dataset admits one best partition. Yet many real-world datasets naturally support multiple valid organizations. Images may cluster by color, shape, or semantics; text corpora may cluster by topic, sentiment, stance, or writing style; biological data may cluster by cell type, state, or developmental trajectory. This introduces a unique challenge between output quality and solution variety: each clustering should be meaningful on its own, while the full set of clusterings should be complementary rather than redundant.
This tutorial provides a comprehensive overview of deep multiple clustering, tracing the field's evolution from classical foundations to today's context-aware approaches. We cover deep representation-based methods built on multi-branch autoencoders, multi-head architectures, disentangled latent factors, and subspace strategies; deep multi-view and multi-source methods; and the recent wave of context-aware and user-guided methods driven by large multimodal models that translate natural-language prompts into clustering guidance.
The tutorial closes with empirical comparisons, industrial applications in search, personalization, and recommendation, and a community discussion of open research questions. It is designed for a broad audience and introduces all necessary background from the ground up.
The tutorial is organized into six parts spanning three hours, with Q&A and breaks built in.
| Duration | Presenter | Topic |
|---|---|---|
| 30 min | Jian Pei | Part 1. Introduction and Problem Formulation — motivating examples in vision, text, biology; multi-objective formulation; taxonomy of deep multiple clustering. |
| 5 min | — | Q&A and Break |
| 30 min | Qi Qian | Part 2. Deep Representation-Based Methods — multi-branch autoencoders, multi-head architectures, disentangled latent factors, subspace strategies. |
| 10 min | — | Q&A and Break |
| 30 min | Bangyu Zou | Part 3. Multi-View / Multi-Source Methods — preserving distinct subspace structures across views and sources. |
| 5 min | — | Q&A and Break |
| 30 min | Juhua Hu | Part 4. Context-Aware / User-Guided Methods — natural-language prompts, multimodal proxies, and personalization of clustering. |
| 10 min | — | Q&A and Break |
| 20 min | Huiji Gao | Part 5. Empirical Comparisons, Applications & Future Directions — benchmarks, industry applications, open challenges. |
| 10 min | All | Part 6. Open Discussion and Q&A — community brainstorming on emerging applications and open questions. |
Browse the slide deck below, or download the full PDF for offline viewing.
A compact starting bibliography. A more comprehensive list will accompany the slides.