Multimodal sarcasm detection (MSD) remains challenging due to the subtle, context-dependent, and often incongruous nature of sarcastic expressions across different modalities. Existing approaches frequently struggle with deep semantic understanding of visual content and lack effective fusion strategies to capture the cross-modal interactions
underlying human sarcasm comprehension. Inspired by neurocognitive findings that sarcasm comprehension flexibly relies on either unimodal or multimodal cues depending on context, we propose the LVLM-Enhanced Two-Stage Multimodal Fusion Framework (LETSMF). LETSMF employs a Large Vision-Language Model (LVLM) to generate context-aware image captions as an enriched visual modality, and integrates featurelevel and decision-level fusion to jointly model intra-modal semantic richness and inter-modal incongruity. Experiments on two public MSD benchmarks demonstrate LETSMF’s effectiveness and strong generalizability, highlighting its potential as a cognitive reasoning component for Agentic AI-driven service systems.