多模态融合的伪装目标检测研究综述
基于RGB图像的单模态伪装目标检测研究
这些文献主要关注在RGB图像中,通过特征增强、多尺度特征融合、边界引导以及上下文与纹理建模等策略来解决伪装目标与背景高度相似的问题,不涉及多模态融合。
- IRFNet: Cognitive-Inspired Iterative Refinement Fusion Network for Camouflaged Object Detection(Guohan Li, Jingxin Wang, Jianming Wei, Zhengyi Xu, 2025, Sensors)
- Camouflaged Object Detection using Multi-Level Feature Cross-Fusion(Tianchi Qiu, Xiuhong Li, Songlin Li, Chenyu Zhou, Kangwei Liu, 2024, 2024 International Joint Conference on Neural Networks (IJCNN))
- Boundary-Guided Fusion of Multi-Level Features Network for Camouflaged Object Detection(Songlin Li, Zhe Li, Boyuan Li, Xiuhong Li, Jiabao Sheng, 2024, 2024 International Joint Conference on Neural Networks (IJCNN))
- Camouflaged object detection via context and texture-aware hierarchical interaction(Zhi Wang, Yangyang Deng, Chenxing Shen, Miaohui Zhang, Xiaoxia Lu, 2026, Scientific Reports)
- Dual-task collaborative network for camouflaged object detection via edge-coarse segmentation map fusion(Bin Jiang, Jinlan Li, Tingting Zhou, K. Zuo, Shidong Xiong, Hanguang Xiao, Gui-bin Bian, 2026, Engineering Applications of Artificial Intelligence)
多模态融合的伪装目标检测与显著性检测
这些文献研究通过引入红外、深度或多模态信息,利用互补优势增强在复杂场景下的检测效果,涵盖了RGB-T、RGB-D等多种模态组合。
- A Dual-Stream Cross-Domain Integration Network for RGB-T Salient Object Detection(Xiaosheng Yu, Xiufei Cheng, Yixiu Liu, Zhigao Zheng, 2025, IEEE Transactions on Consumer Electronics)
- Visible-Infrared Camouflaged Object Detection(Cheng Liu, Zheng Wang, Xinyu Yan, Meijun Sun, Qinghua Hu, 2026, IEEE Transactions on Circuits and Systems for Video Technology)
- Three-Decoder Cross-Modal Interaction Network for Unregistered RGB-T Salient Object Detection(Xin Wen, Jianxun Zhao, Yu He, Haixu Yin, 2025, IEEE Transactions on Instrumentation and Measurement)
- Multi-stream information complementarity network for RGB-D camouflaged object detection(Chenghao Ying, Zhiping Zhou, Kewei Li, Zhaozhong Zhang, Qingshuang Yang, 2025, The Journal of Supercomputing)
跨任务通用框架与高光谱机理分析
这些文献不局限于单一的伪装检测任务,而是从更广泛的视觉通用模型角度(如SOD与COD联合)或从物理层面的光谱特性角度探讨伪装检测的机理与应用。
- VSCode: General Visual Salient and Camouflaged Object Detection with 2D Prompt Learning(Ziyang Luo, Nian Liu, Wangbo Zhao, Xuguang Yang, Dingwen Zhang, Deng-Ping Fan, Fahad Khan, Junwei Han, 2023, 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR))
- Spectral Camouflage Characteristics and Recognition Ability of Targets Based on Visible/Near-Infrared Hyperspectral Images(Jiale Zhao, Bin Zhou, Guanglong Wang, Jiaju Ying, Jie Liu, Qi Chen, 2022, Photonics)
本次综述涵盖的文献主要分为三类:一是以优化RGB图像特征提取和建模为核心的单模态伪装检测方法;二是通过引入红外、深度等模态进行多源信息融合,以提升复杂场景下伪装目标辨识能力的模型;三是侧重于通用视觉任务协同学习的框架以及基于物理光谱特性的伪装机理研究。
总计11篇相关文献
Salient object detection (SOD) and camouflaged object detection (COD) are related yet distinct binary mapping tasks. These tasks involve multiple modalities, sharing commonalities and unique cues. Existing research often employs intricate task-specific specialist models, potentially leading to redundancy and suboptimal results. We introduce VS-Code, a generalist model with novel 2D prompt learning, to jointly address four SOD tasks and three COD tasks. We utilize VST as the foundation model and introduce 2D prompts within the encoder-decoder architecture to learn domain and task-specific knowledge on two separate dimensions. A prompt discrimination loss helps disentangle peculiarities to benefit model optimization. VSCode outperforms state-of-the-art methods across six tasks on 26 datasets and exhibits zero-shot generalization to unseen tasks by combining 2D prompts, such as RGB-D COD. Source code has been available at https://github.com/Sssssuperior/VSCode.
… , accurately detecting … ) for camouflage object detection, comprehensively exploiting complementary useful information from depth and RGB images to boost the accuracy of detection. …
Although great progress has been made in Camouflaged Object Detection (COD), it still faces challenges in complex real-world scenes. Existing methods are primarily designed for visible images but face limitations when detecting highly camouflaged or partially occluded objects. Integrating multiple complementary information sources, such as visible images and infrared images, is an effective way to improve the performance of COD. However, research in this field is limited by the lack of comprehensive and high-quality benchmark datasets. To solve this problem, a Visible-Infrared Artificial Camouflage (VIAC) dataset is constructed. Building on this dataset, we propose a novel Visible-Infrared Camouflaged Object Detection (VICOD) framework, termed the Confidence-Guided Fusion and Inpainting Network (CGFINet). The network utilizes a cross-modal collaborative fusion module (CMCF) to achieve adaptive integration of visible and infrared information. Simultaneously, low-confidence regions segmentation boundaries are refined by leveraging high-confidence pixel information within the confidence-driven inpainting module (CDIM). To focus on low-confidence areas, pixel-level uncertainty is incorporated into the loss function as a dynamic weight factor, which prompts the model to focus on high-uncertainty areas. Extensive experiments on VIAC demonstrate that our method achieves state-of-the-art performance, surpassing existing COD and visible-infrared SOD approaches.
Hyperspectral imaging can simultaneously obtain the spatial morphological information of the ground objects and the fine spectral information of each pixel. Through the quantitative analysis of the spectral characteristics of objects, it can complete the task of classification and recognition of ground objects. The appearance of imaging spectrum technology provides great advantages for military target detection and promotes the continuous improvement of military reconnaissance levels. At the same time, spectral camouflage materials and methods that are relatively resistant to hyperspectral reconnaissance technology are also developing rapidly. In order to study the reconnaissance effect of visible/near-infrared hyperspectral images on camouflage targets, this paper analyzes the spectral characteristics of different camouflage targets using the hyperspectral images obtained in the visible and near-infrared bands under natural conditions. Two groups of experiments were carried out. The first group of experiments verified the spectral camouflage characteristics and camouflage effects of different types of camouflage clothing with grassland as the background; the second group of experiments verified the spectral camouflage characteristics and camouflage effects of different types of camouflage paint sprayed on boards and steel plates. The experiment shows that the hyperspectral image based on the near-infrared band has a good reconnaissance effect for different camouflage targets, and the near-infrared band is an effective “window” band for detecting and distinguishing true and false targets. However, the stability of the visible/near-infrared band detection for the target identification under camouflage paint is poor, and it is difficult to effectively distinguish the object materials under the same camouflage paint. This research confirms the application ability of detection based on the visible/near-infrared band, and points out the direction for the development of imaging detectors and camouflage materials in the future.
Substantial advancements have been made in the field of salient object detection in image processing in recent years. This research introduces the three-decoder cross-modal interaction network (TCINet) for salient object detection in unregistered red-green-blue (RGB)–thermal image pairs, modeling information from different modal perspectives. TCINet employs a three-decoder framework to process RGB, thermal, and fused feature maps concurrently. To ensure robust integration between the modalities, mitigating the impact of unregistered images and addressing modality imbalances, we introduce the fusion complementary registration (FCR) module. This module guides attention to connect the two modalities and uses atrous spatial pyramid pooling (ASPP) to adapt to image scale changes. To fully utilize the differences between modalities, we designed two distinct decoders: fusion feature decoder (FFD) for decoding the fused features and single-modal decoder (SMD) for decoding single-modal features. Additionally, we incorporated feature enhancement (FE) units into the modal decoding to mitigate the blurring effect caused by high-speed autonomous aerial vehicle (AAV) flight. We use a weighted fusion module (WFM) to dynamically integrate the features decoded by the three decoders to increase the network’s generalization ability. Extensive experiments show that TCINet outperforms existing methods, achieving excellent results on a variety of challenging scenarios containing complex details. The code will be published at https://github.com/zqiuqiu235/TCINet.git.
RGB-T salient object detection enhances the performance of detection in complex scenes by integrating RGB and thermal data, but effective fusion remains challenging. To this end, we propose a dual-stream cross-domain integration network (DSCDNet), which explores the effective integration of spatial and frequency domain features, demonstrating remarkable accuracy and stability. Specifically, in the bimodal integration stage, to deeply extract high-level fusion information from multiple modalities, we introduce the spatial domain bimodal integration module and the frequency domain bimodal integration module. This parallel integration strategy facilitates the deep integration of RGB and thermal image from multiple dimensions. After that, our proposed feature decomposition strategy decomposes the spatial fused features into two streams: region perception stream and detail-aware stream, which interpret spatial features from the perspectives of regional and detail understanding, making the model more flexible and efficient in complex scenes. Further, to maximize the complementary advantages between spatial and frequency domain, we design a cross domain feature alignment module to facilitate the interactive learning between the two domains, thereby providing the model with a more comprehensive and enriched representation perspective. Experimental results demonstrate that the proposed DSCDNet outperforms 11 state-of-the-art (SOTA) methods.
… of camouflaged objects, we propose a dual-task collaborative network based on edge information maps and coarse segmentation … of camouflaged objects; third, a Coarse Segmentation …
Camouflaged object detection (COD) aims to segment objects that closely resemble their surroundings. Accurately recognizing camouflaged objects in these complex environments is challenging due to factors such as low illumination, object occlusion, small size, and similar background. To this end, we propose a novel network for camouflaged object detection, the Multi-Level Feature Cross-Fusion Network (MFCF-Net). This framework aims to learn and utilize background features at different scales through cross-fusion, thereby improving detection accuracy. The core of our approach is to use a modified version of the Pyramid Vision Transformer (PVTv2) as a backbone network to effectively capture contextual information at different scales. Then, we design the Multi-scale Feature Enhancement (MFE) module to optimize features at each scale. In addition, to enhance the model’s ability to recognize camouflaged objects in complex contexts, we cross-fused these enhanced features. Finally, we designed the Balanced Multilevel Feature Cross-Fusion (BMFCF) module. This module improves the accuracy of camouflaged object detection by deeply learning and effectively utilizing contextual feature information and cross-fusing these multi-scale features. Extensive research results show that our MFCF-Net significantly outperforms 18 leading methods on four widely used standard datasets.
Camouflaged objects, exhibiting high similarity with their surroundings, pose a substantial challenge for both humans and machines to detect when concealed within the environment. Existing methods for camouflage object detection (COD) struggle in accurately segmenting the overall structure of camouflaged objects. To address this issue, we propose a novel boundary-guided fusion of multi-level features network (BGFM-Net) for COD. In contrast to existing boundary-guided methods, we pay more attention to addressing the significant imbalance in the pixel quantities between boundary and background features, allowing for a more comprehensive representation of boundary features. BGFM-Net primarily consists of a multi-scale aggregation module (MSAM), a boundary-guided feature module (BFM), and a cross-Level fusion module (CLFM). MSAM effectively integrates contextual semantics at different scales, achieving a powerful and efficient feature representation. BFM adeptly combines edge features while constraining interference from background features, guiding the learning of camouflaged object boundary representation. CLFM integrates multi-level features for predicting camouflaged objects while adaptively adjusting channel weights to emphasize important channels and diminish the impact of less relevant channels for the task. Extensive experiments on three benchmark camouflage datasets demonstrate that our BGFM-Net outperforms other state-of-the-art COD models.
In the field of camouflaged object detection (COD), effectively distinguishing the intrinsic similarity between objects and their backgrounds is a critical factor for improving detection performance. Existing approaches typically leverage boundary constraints to provide additional auxiliary information during the training phase. To capture more discriminative detailed cues, we introduce texture labels as supervisory signals and propose a context- and texture-aware hierarchical interaction network (CTHINet) for COD. In the coding phase, the network is divided into two separate branches, a context and a texture encoder. Specifically, a context encoder is employed to generate contextual information. Subsequently, the features at different scales are refined by implementing a Multi-head Feature Aggregation Module (MFAM). The diversity of features is subsequently enhanced by leveraging the interactions among their distinct feature receptive fields, facilitating the matching of candidate areas for camouflaged objects with varying sizes and shapes. Following this, the enhanced features are combined with texture features generated by the texture encoder, fully exploiting imperceptible cues within candidate objects through utilizing the Hierarchical mixed-scale Interaction Modules (HMIM). This module continuously integrates texture cues with contextual information within a single feature scale, aiming for more accurate detection. Extensive experiments conducted on three challenging benchmark datasets, e.g., CAMO, COD10K, and NC4K, illustrate that our model has superior performance compared to state-of-the-art methods. Furthermore, the evaluation results on the polyp segmentation dataset underscore the promising potential of CTHINet for downstream applications.
Camouflaged Object Detection (COD) aims to identify objects that are intentionally concealed within their surroundings through appearance, texture, or pattern adaptations. Despite recent advances, extreme object–background similarity causes existing methods struggle with accurately capturing discriminative features and effectively modeling multiscale patterns while preserving fine details. To address these challenges, we propose Iterative Refinement Fusion Network (IRFNet), a novel framework that mimics human visual cognition through progressive feature enhancement and iterative optimization. Our approach incorporates the following: (1) a Hierarchical Feature Enhancement Module (HFEM) coupled with a dynamic channel-spatial attention mechanism, which enriches multiscale feature representations through bilateral and trilateral fusion pathways; and (2) a Context-guided Iterative Optimization Framework (CIOF) that combines transformer-based global context modeling with iterative refinement through dual-branch supervision. Extensive experiments on three challenging benchmark datasets (CAMO, COD10K, and NC4K) demonstrate that IRFNet consistently outperforms fourteen state-of-the-art methods, achieving improvements of 0.9–13.7% across key metrics. Comprehensive ablation studies validate the effectiveness of each proposed component and demonstrate how our iterative refinement strategy enables progressive improvement in detection accuracy.
本次综述涵盖的文献主要分为三类:一是以优化RGB图像特征提取和建模为核心的单模态伪装检测方法;二是通过引入红外、深度等模态进行多源信息融合,以提升复杂场景下伪装目标辨识能力的模型;三是侧重于通用视觉任务协同学习的框架以及基于物理光谱特性的伪装机理研究。