Logo image
Disentangling Shared and Specific Features in Deep Multimodal Clustering with Contrastive-Complementary Learning
Journal article   Peer reviewed

Disentangling Shared and Specific Features in Deep Multimodal Clustering with Contrastive-Complementary Learning

Don Yates, Hakki Erhan Sevil and Arash Mahyari
Pattern Recognition Letters, Vol.207, pp.274-280
07/07/2026
Web of Science ID: WOS:001824776300001

Metrics

1 Record Views

Abstract

Multi-view representation learning Contrastive and complementary decomposition Representation disentanglement Multi-view clustering Pattern Recognition
Multi-modal representation learning must integrate heterogeneous inputs while preserving both shared semantics and modality-specific complementary cues. Many deep multi-view clustering methods instead assume conditional independence between shared and private latents or enforce strict decorrelation, which can discard informative cross-view dependencies and reduce effectiveness when complementary information is important. We introduce C3U-MVC, a contrastive–complementary framework that learns an optimization-biased separation between aligned common representations and view-specific complementary information without imposing independence or orthogonality constraints. The model combines (i) cross-view contrastive alignment in the shared subspace, (ii) reconstruction from concatenated shared and view-specific latents to preserve complementary signals, and (iii) prototype-conditioned reconstruction with per-instance best-view selection for adaptive assignment. Training proceeds in two stages: contrastive pretraining followed by positives-only prototype refinement. Simple counterexamples show why independence assumptions need not hold in realistic multi-view settings. Experiments on seven benchmarks, including constructed image-pair datasets, true multi-view datasets, and heterogeneous Caltech-101 subsets, show that C3U-MVC achieves state-of-the-art or competitive clustering performance, with the clearest gains appearing when view-dependent complementary information is informative. Ablations further support the roles of complementary reconstruction, latent-capacity allocation, and prototype-guided refinement.

Details

Logo image