C$^2$MOE: Consistency and Complementarity-guided Mixture of Experts for Incomplete Multimodal Emotion Learning
Recent advances in Multimodal Emotion Recognition in Conversations (MERC) highlight its reliance on complete multimodal inputs. However, real-world data often suffer from missing modalities due to transmission errors or user behavior, severely degrading model performance.
Intelligence analysis by Llama

A novel Consistency and Complementarity-guided Mixture of Experts framework, C$2$MOE, is proposed for incomplete multimodal emotion learning. It unifies representation learning and missing modality imputation within a principled information-theoretic framework.
Imagine you're trying to recognize emotions from a conversation, but some of the information is missing. A new way to do this, called C$2$MOE, is better at filling in the missing pieces and making accurate predictions.
Analysis
A Novel Framework for Incomplete Multimodal Emotion Learning
The proposed framework, C$2$MOE, is a novel Consistency and Complementarity-guided Mixture of Experts framework for incomplete multimodal emotion learning. It unifies representation learning and missing modality imputation within a principled information-theoretic framework. Specifically, multimodal knowledge is factorized into consistency and complementarity components via interaction-aware experts. Consistency is captured by maximizing cross-modal predictability, while complementarity is preserved by maximizing conditional entropy between modalities.
Building Upon This Decomposition
Building upon this decomposition, C$2$MOE introduces a dual-branch prediction mechanism for robust imputation under missing modalities. The consistency branch aligns imputed features with the joint distribution by minimizing uncertainty, and the complementarity branch exploits modality-unique cues via entropy maximization. Finally, C$2$MOE employs a learnable reweighting module that dynamically assigns importance scores to each expert's output, yielding a robust and adaptive fusion for imputation.
Experimental Results
Extensive experiments on multiple MERC benchmarks demonstrate that C$2$MOE consistently surpasses state-of-the-art methods across various missing-modality settings, validating its robustness and generalization.
Key points
- C$2$MOE is a novel Consistency and Complementarity-guided Mixture of Experts framework for incomplete multimodal emotion learning.
- It unifies representation learning and missing modality imputation within a principled information-theoretic framework.
- C$2$MOE introduces a dual-branch prediction mechanism for robust imputation under missing modalities.
- Extensive experiments on multiple MERC benchmarks demonstrate that C$2$MOE consistently surpasses state-of-the-art methods across various missing-modality settings.
If C$2$MOE is widely adopted, it could lead to significant improvements in multimodal emotion recognition models, enabling more accurate and robust predictions in real-world scenarios.
However, the success of C$2$MOE also depends on the availability of high-quality training data and the ability to adapt to new and diverse scenarios.



