A real disaster photo paired with a misleading caption can create a convincing false story. In multimodal fake news, an image may be authentic and a sentence may sound plausible—the issue lies in their interaction. Many detectors concatenate the two modalities too early and lose that signal.
IMEX-FND—Interaction-aware Mixture-of-Experts for Fake News Detection—explores how to make those interactions explicit, inspectable, and useful for classification.
The problem with coarse fusion
Typical multimodal methods concatenate text and image features, or pass them through an attention block, before classification. That can hide two important differences:
- Different cases need different evidence. Some false stories are exposed by text-image inconsistency; others by a single modality.
- Decisions become opaque. A classifier may label an item false without indicating whether the text, the image, or their mismatch was decisive.
Five experts for five kinds of interaction
The framework separates cross-modal evidence into five experts:
An instance-level router assigns weights to these experts for each item. Those weights act as a direct explanation of which type of evidence influenced a decision.
Core idea: do not ask one fusion layer to handle every kind of example. Separate the interactions, let the right expert handle the right case, and use routing weights to show why.
A shared representation space
The experts work from a common-dimensional feature space built from a BERT-style text encoder, an MAE-pretrained ViT-B/16 image encoder, and CLIP ViT-B/16 for text-image alignment. Projecting these features into a shared space makes the expert outputs comparable.
Training for robust multimodal reasoning
Multimodal models often rely too heavily on the strongest modality. To reduce that shortcut, the project uses modality replacement during training, losses that distinguish the uniqueness experts, and auxiliary objectives for the synergy expert.
What matters
The central point is not simply to increase model capacity. It is to choose a structure that matches the problem: model decisions should reflect whether the evidence is textual, visual, contradictory, complementary, or redundant.
This article accompanies IMEX-FND: A Traceable Interaction-Aware Mixture-of-Experts Framework for Multimodal Fake News Detection, accepted to WISE 2026. Zijun Wang is the second author. Read the paper ↗ or get in touch to discuss the project.
