Abstract Visual Reasoning Deep Learning
基于Raven's Progressive Matrices (RPM)的矩阵推理模型
聚焦于RPM及其衍生数据集的深度学习架构设计,通过视觉感知与逻辑规则推导解决结构化抽象类比任务。
- RAVEN: A Dataset for Relational and Analogical Visual REasoNing(Chi Zhang, Feng Gao, Baoxiong Jia, Yixin Zhu, Song-Chun Zhu, 2019, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR))
- Deep Learning Methods for Abstract Visual Reasoning: A Survey on Raven's Progressive Matrices(Mikolaj Malki'nski, Jacek Ma'ndziuk, 2022, ACM Computing Surveys)
- Hierarchical Perceptual and Predictive Analogy-Inference Network for Abstract Visual Reasoning(Wentao He, Jianfeng Ren, Ruibin Bai, Xudong Jiang, 2024, Proceedings of the 32nd ACM International Conference on Multimedia)
- Learning abstract visual reasoning via task decomposition: A case study in Raven progressive matrices(Jakub Kwiatkowski, Krzysztof Krawiec, 2024, International Journal of Applied Mathematics and Computer Science)
- Learning Visual Abstract Reasoning through Dual-Stream Networks(Kai Zhao, Chang Xu, Bailu Si, 2024, Proceedings of the AAAI Conference on Artificial Intelligence)
- MLRQA: A Dataset with Multimodal Logical Reasoning Challenges(Jing Xiao, Guijin Lin, Ping Li, 2024, Lecture Notes in Computer Science)
- Hierarchical Attention Network with Slot-weighted Modeling for Visual Reasoning(Haoning Yu, Jinlin Guo, Wenfeng Hu, Zhiping Shi, Jun Li, 2026, SSRN Electronic Journal)
- Stratified Rule-Aware Network for Abstract Visual Reasoning(Sheng Hu, Yuqing Ma, Xianglong Liu, Yanlu Wei, Shihao Bai, 2020, Proceedings of the AAAI Conference on Artificial Intelligence)
- Abstract Visual Reasoning: An Algebraic Approach for Solving Raven's Progressive Matrices(Jingyi Xu, Tushar Vaidya, Y. Blankenship, Saket Chandra, Zhangsheng Lai, Kai Fong Ernest Chong, 2023, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR))
- Abstract visual reasoning with hybrid relation modeling(Jinlin Guo, Wenhua Shen, Yancheng Zhao, L. Kang, 2025, Displays)
- Dextroamphetamine Enhances “Neural Network-Specific” Physiological Signals: A Positron-Emission Tomography rCBF Study(V. Mattay, K. Berman, J. Ostrem, G. Esposito, J. V. Van Horn, L. Bigelow, D. Weinberger, 1996, The Journal of Neuroscience)
神经符号推理系统与逻辑架构
研究如何融合神经网络的感知性能与显式逻辑符号的推理能力,重点在于提升模型的解释性、泛化能力及约束处理机制。
- Learning Where and When to Reason in Neurosymbolic Inference(Cristina Cornelio, 2025, Metacognitive Artificial Intelligence)
- Symmetric Graph-Based Visual Question Answering Using Neuro-Symbolic Approach(Jiyoun Moon, 2023, Symmetry)
- Abductive learning: towards bridging machine learning and logical reasoning(Zhi-Hua Zhou, 2019, Science China Information Sciences)
- AI Reasoning in Deep Learning Era: From Symbolic AI to Neural–Symbolic AI(Baoyu Liang, Yucheng Wang, Chao Tong, 2025, Mathematics)
- Neuro-Symbolic AI: Bringing a new era of Machine Learning(Rabinandan Kishor, 2022, International Journal of Research Publication and Reviews)
- A bidirectional neuro-symbolic framework for clinical decision support via dynamic integration of deep learning and symbolic reasoning(R. Chavda, K. Suresh, Sushil Kumar, K. L. R. Reddy, B. Jayaprakash, P. Sahu, B. Bharathi, Devendra Singh, Saurabh Namdev, 2026, Network Modeling Analysis in Health Informatics and Bioinformatics)
- Ontology Reasoning with Deep Neural Networks(Patrick Hohenecker, Thomas Lukasiewicz, 2018, Journal of Artificial Intelligence Research)
- Neuro-Symbolic Homogeneous Concept Reasoning for Scene Interpretation(Sangwon Kim, ByoungChul Ko, In-su Jang, Kwangju Kim, 2024, 2024 International Conference on Platform Technology and Service (PlatCon))
- Neuro-Symbolic Reasoning: Performance, Challenges, and Benchmarks: A Systematic Literature Review(P. Obike, P. U. Usip, E. Udo, Aniekan J. Ananga, 2025, ABUAD Journal of Engineering and Applied Sciences)
- Neural Logic Reasoning(Shaoyun Shi, H. Chen, Weizhi Ma, Jiaxin Mao, Min Zhang, Yongfeng Zhang, 2020, Proceedings of the 29th ACM International Conference on Information & Knowledge Management)
认知机制与通用抽象推理
探讨ARC数据集、人类类比能力、工作记忆与抽象思维模型,旨在揭示机器认知与人类抽象思维之间的内在机制。
- Learn to abstract via concept graph for weakly-supervised few-shot learning(Baoquan Zhang, Ka-Cheong Leung, Xutao Li, Yunming Ye, 2021, Pattern Recognition)
- Same/different in visual reasoning(Kenneth D. Forbus, A. Lovett, 2021, Current Opinion in Behavioral Sciences)
- Few-Shot Semantic Segmentation in Remote Sensing: A Review on Definitions, Methods, Datasets, Advances and Future Trends(M. Petrov, Ema Pandilova, I. Dimitrovski, D. Trajanov, Vlatko Spasev, Ivan Kitanovski, 2026, Remote Sensing)
- Development of Analogical Reasoning: A Novel Perspective From Cross‐Cultural Studies(S. Christie, Yang Gao, Qiuchen Ma, 2020, Child Development Perspectives)
- Abstract Visual Reasoning Enabled by Language(Giacomo Camposampiero, Loic Houmard, Benjamin Estermann, Joël Mathys, R. Wattenhofer, 2023, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW))
- The relationships among working memory, inhibitory control, and mathematical skills in primary school children: Analogical reasoning matters(Yue Qi, Yinghe Chen, Xiao Yu, Xiujie Yang, Xinyi He, Xiaoyu Ma, 2024, Cognitive Development)
- Neural networks for abstraction and reasoning(Mikel Bober-Irizar, Soumya Banerjee, 2024, Scientific Reports)
- Neural network (overview)(A. Murphy, J. Seah, 2017, Radiopaedia.org)
- Working memory as a moderator of training and transfer of analogical reasoning in children(C. Stevenson, W. Heiser, W. Resing, 2013, Contemporary Educational Psychology)
- The association between relational reasoning in nonverbal and verbal representations and mathematics achievement.(Eason Sai-Kit Yip, T. Wong, 2024, Journal of Educational Psychology)
- Abstraction and analogy in AI(Melanie Mitchell, 2023, Annals of the New York Academy of Sciences)
- Patterns of analogical reasoning among beginning readers(Lee Farrington-Flint, C. Wood, K. H. Canobi, D. Faulkner, 2004, Journal of Research in Reading)
- Towards a Theory of Commonsense Visual Reasoning(B. Chandrasekaran, N. Narayanan, 1990, Lecture Notes in Computer Science)
- Neural oscillatory dynamics serving abstract reasoning reveal robust sex differences in typically-developing children and adolescents(Brittany K. Taylor, C. Embury, E. Heinrichs-Graham, Michaela R. Frenzel, Jacob A. Eastman, Alex I. Wiesman, Yu-ping Wang, V. Calhoun, J. Stephen, T. Wilson, 2020, Developmental Cognitive Neuroscience)
表征几何与关系推理的理论分析
从表征几何学、关系建模及信号处理的角度研究深度学习模型内部推理能力的演进规律。
- Unraveling the geometry of visual relational reasoning.(Jiaqi Shang, G. Kreiman, H. Sompolinsky, 2026, Scientific Reports)
- Neural representational geometry underlies few-shot concept learning(Ben Sorscher, S. Ganguli, H. Sompolinsky, 2022, Proceedings of the National Academy of Sciences)
- Object-centric Learning with Capsule Networks: A Survey(Fabio De Sousa Ribeiro, Kevin Duarte, Miles Everett, G. Leontidis, M. Shah, 2024, ACM Computing Surveys)
视觉模拟、场景理解与多模态推理
专注于复杂视觉场景中的关系推理、类比任务及视频/图像的多模态特征融合,强调模型在实际应用中的场景解析能力。
- A Neural Model of Rule Generation in Inductive Reasoning(Daniel Rasmussen, C. Eliasmith, 2011, Topics in Cognitive Science)
- Enhancing Advanced Visual Reasoning Ability of Large Language Models(Zhiyuan Li, Dongnan Liu, Chaoyi Zhang, Heng Wang, Tengfei Xue, Weidong Cai, 2024, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing)
- Deep neural reasoning(Herbert Jaeger, 2016, Nature)
- RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning(Deyi Ji, Yuekui Yang, Haiyang Wu, Shaoping Ma, Tianrun Chen, Lanyun Zhu, 2025, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track))
- UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning(Huy Le, Nhat Chung, T. Kieu, Jingkang Yang, Ngan Le, 2025, 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV))
- Multimodal feature fusion by relational reasoning and attention for visual question answering(Weifeng Zhang, J. Yu, Hua Hu, Haiyang Hu, Zengchang Qin, 2020, Information Fusion)
- Neural algorithmic reasoning(Petar Velickovic, C. Blundell, 2021, Patterns)
- Understanding the What and When of Analogical Reasoning Across Analogy Formats: An Eye‐Tracking and Machine Learning Approach(J. Thibaut, Yannick Glady, R. French, 2022, Cognitive Science)
- See Beyond: Benchmarking MLLMs' Visual Relational Reasoning Ability(Yifan Wang, Haizhou Wang, 2025, Lecture Notes in Computer Science)
- Relational Reasoning Image Captioning Via Multi-Agent Retrieval-Augmented Generation(Aiwen Jiang, Duan Wang, Chao Peng, Mingwen Wang, 2025, Knowledge-Based Systems)
- VASR: Visual Analogies of Situation Recognition(Yonatan Bitton, Ron Yosef, Eliyahu Strugo, Dafna Shahaf, Roy Schwartz, Gabriel Stanovsky, 2023, Proceedings of the AAAI Conference on Artificial Intelligence)
- Broadcasting Convolutional Network for Visual Relational Reasoning(Simyung Chang, John Yang, Seonguk Park, Nojun Kwak, 2017, Lecture Notes in Computer Science)
- Relational Reasoning Over Spatial-Temporal Graphs for Video Summarization(Wencheng Zhu, Yucheng Han, Jiwen Lu, Jie Zhou, 2022, IEEE Transactions on Image Processing)
- NERO: NEural algorithmic reasoning for zeRO-day attack detection in the IoT: A hybrid approach(Jesús Fernando Cevallos Moreno, A. Rizzardi, S. Sicari, A. Coen-Porisini, 2024, Computers & Security)
- A unified graph neural network-based approach for few-shot learning with task nodes and DiffPool abstraction(Poupak Azad, Arash Heidari, C. Akcora, Ahmad Khonsari, Seyed Hamed Rastegar, 2026, Neurocomputing)
本报告对抽象视觉推理研究进行了系统整合,分为五大核心维度:RPM矩阵推理、神经符号推理、认知科学视角、表征几何理论以及复杂场景的多模态推理。研究重心已从早期的特定基准评测,扩展至融合人类认知机制的架构设计,以及通过表征几何等数学工具揭示深层逻辑泛化能力的理论探索,展示了AI从简单逻辑匹配迈向通用推理的趋势。
总计53篇相关文献
While artificial intelligence (AI) models have achieved human or even superhuman performance in many well-defined applications, they still struggle to show signs of broad and flexible intelligence. The Abstraction and Reasoning Corpus (ARC), a visual intelligence benchmark introduced by François Chollet, aims to assess how close AI systems are to human-like cognitive abilities. Most current approaches rely on carefully handcrafted domain-specific program searches to brute-force solutions for the tasks present in ARC. In this work, we propose a general learning-based framework for solving ARC. It is centered on transforming tasks from the vision to the language domain. This composition of language and vision allows for pre-trained models to be leveraged at each stage, enabling a shift from handcrafted priors towards the learned priors of the models. While not yet beating state-of-the-art models on ARC, we demonstrate the potential of our approach, for instance, by solving some ARC tasks that have not been solved previously.
Abstract reasoning refers to the ability to analyze information, discover rules at an intangible level, and solve problems in innovative ways. Raven's Progressive Matrices (RPM) test is typically used to examine the capability of abstract reasoning. The subject is asked to identify the correct choice from the answer set to fill the missing panel at the bottom right of RPM (e.g., a 3×3 matrix), following the underlying rules inside the matrix. Recent studies, taking advantage of Convolutional Neural Networks (CNNs), have achieved encouraging progress to accomplish the RPM test. However, they partly ignore necessary inductive biases of RPM solver, such as order sensitivity within each row/column and incremental rule induction. To address this problem, in this paper we propose a Stratified Rule-Aware Network (SRAN) to generate the rule embeddings for two input sequences. Our SRAN learns multiple granularity rule embeddings at different levels, and incrementally integrates the stratified embedding flows through a gated fusion module. With the help of embeddings, a rule similarity metric is applied to guarantee that SRAN can not only be trained using a tuplet loss but also infer the best answer efficiently. We further point out the severe defects existing in the popular RAVEN dataset for RPM test, which prevent from the fair evaluation of the abstract reasoning ability. To fix the defects, we propose an answer set generation algorithm called Attribute Bisection Tree (ABT), forming an improved dataset named Impartial-RAVEN (I-RAVEN for short). Extensive experiments are conducted on both PGM and I-RAVEN datasets, showing that our SRAN outperforms the state-of-the-art models by a considerable margin.
Dramatic progress has been witnessed in basic vision tasks involving low-level perception, such as object recognition, detection, and tracking. Unfortunately, there is still enormous performance gap between artificial vision systems and human intelligence in terms of higher-level vision problems, especially ones involving reasoning. Earlier attempts in equipping machines with high-level reasoning have hovered around Visual Question Answering (VQA), one typical task associating vision and language understanding. In this work, we propose a new dataset, built in the context of Raven's Progressive Matrices (RPM) and aimed at lifting machine intelligence by associating vision with structural, relational, and analogical reasoning in a hierarchical representation. Unlike previous works in measuring abstract reasoning using RPM, we establish a semantic link between vision and reasoning by providing structure representation. This addition enables a new type of abstract reasoning by jointly operating on the structure representation. Machine reasoning ability using modern computer vision is evaluated in this newly proposed dataset. Additionally, we also provide human performance as a reference. Finally, we show consistent improvement across all models by incorporating a simple neural module that combines visual understanding and structure reasoning.
… Context-Aware Image Description vs General Image Caption In this section, we investigate CaID’s impact at an abstract level and design a novel method to quantitatively demonstrate …
… facts for improving the machine learning models. … machine learning model and the logical reasoning model jointly. We demonstrate that by using abductive learning, machines can learn …
We introduce algebraic machine reasoning, a new reasoning framework that is well-suited for abstract reasoning. Effectively, algebraic machine reasoning reduces the difficult process of novel problem-solving to routine algebraic computation. The fundamental algebraic objects of interest are the ideals of some suitably initialized polynomial ring. We shall explain how solving Raven's Progressive Matrices (RPMs) can be realized as computational problems in algebra, which combine various well-known algebraic subroutines that include: Computing the Gröbner basis of an ideal, checking for ideal containment, etc. Crucially, the additional algebraic structure satisfied by ideals allows for more operations on ideals beyond set-theoretic operations. Our algebraic machine reasoning framework is not only able to select the correct answer from a given answer set, but also able to generate the correct answer with only the question matrix given. Experiments on the I-RAVEN dataset yield an overall 93.2% accuracy, which significantly out-performs the current state-of-the-art accuracy of77.0% and exceeds human performance at 84.4% accuracy.
For half a century, artificial intelligence research has attempted to reproduce the human qualities of abstraction and reasoning - creating computer systems that can learn new concepts from a minimal set of examples, in settings where humans find this easy. While specific neural networks are able to solve an impressive range of problems, broad generalisation to situations outside their training data has proved elusive. In this work, we look at several novel approaches for solving the Abstraction & Reasoning Corpus (ARC). This is a dataset of abstract visual reasoning tasks introduced to test algorithms on broad generalization. Despite three international competitions with $100,000 in prizes, the best algorithms still fail to solve a majority of ARC tasks. The best solvers today rely on complex hand-crafted rules, without using machine learning at all. We revisit whether recent advances in neural networks allow progress on this task, or whether an entirely different class of models are required. First, we adapt the DreamCoder neurosymbolic reasoning solver to ARC. DreamCoder automatically writes programs in a bespoke domain-specific language to perform reasoning, using a neural network to mimic human intuition. We present the Perceptual Abstraction and Reasoning Language (PeARL) language, which allows DreamCoder to solve ARC tasks, and propose a new recognition model that allows us to significantly improve on the previous best implementation. We also propose a new encoding and augmentation scheme that allows large language models (LLMs) to solve ARC tasks, and find that the largest models can solve some ARC tasks. LLMs are able to solve a different group of problems to state-of-the-art solvers, and provide an interesting way to complement other approaches. We perform an ensemble analysis, combining systems to achieve better results than any system alone and analysing individual strengths. However, it is sobering to see that approaches based on neural networks still lag behind existing hand-crafted solvers, and we suggest avenues for future improvements. Our findings with the ensemble model may indicate that a diversity of methods might be necessary to solve problems in ARC. Humans likely employ diverse strategies to solve ARC. Studies involving human participants to identify the strategies they employ to solve ARC could provide valuable insights for future AI approaches. Finally, we publish the arckit Python library to make future research on ARC easier.
The pursuit of Artificial General Intelligence (AGI) demands AI systems that not only perceive but also reason in a human-like manner. While symbolic systems pioneered early breakthroughs in logic-based reasoning, such as MYCIN and DENDRAL, they suffered from brittleness and poor scalability. Conversely, modern deep learning architectures have achieved remarkable success in perception tasks, yet continue to fall short in interpretable and structured reasoning. This dichotomy has motivated growing interest in Neural–Symbolic AI, a paradigm that integrates symbolic logic with neural computation to unify reasoning and learning. This survey provides a comprehensive and technically grounded overview of AI reasoning in the deep learning era, with a particular focus on Neural–Symbolic AI. Beyond a historical narrative, we introduce a formal definition of AI reasoning and propose a novel three-dimensional taxonomy that organizes reasoning paradigms by representation form, task structure, and application context. We then systematically review recent advances—including Differentiable Logic Programming, abductive learning, program induction, logic-aware Transformers, and LLM-based symbolic planning—highlighting their technical mechanisms, capabilities, and limitations. In contrast to prior surveys, this work bridges symbolic logic, neural computation, and emergent generative reasoning, offering a unified framework to understand and compare diverse approaches. We conclude by identifying key open challenges such as symbolic–continuous alignment, dynamic rule learning, and unified architectures, and we aim to provide a conceptual foundation for future developments in general-purpose reasoning systems.
… capacity to reason about abstract concepts has … abstract reasoning in learning machines, and reveals some important insights about the nature of generalization itself. Artificial neural …
Recent years have witnessed the success of deep neural networks in many research areas. The fundamental idea behind the design of most neural networks is to learn similarity patterns from data for prediction and inference, which lacks the ability of cognitive reasoning. However, the concrete ability of reasoning is critical to many theoretical and practical problems. On the other hand, traditional symbolic reasoning methods do well in making logical inference, but they are mostly hard rule-based reasoning, which limits their generalization ability to different tasks since difference tasks may require different rules. Both reasoning and generalization ability are important for prediction tasks such as recommender systems, where reasoning provides strong connection between user history and target items for accurate prediction, and generalization helps the model to draw a robust user portrait over noisy inputs. In this paper, we propose Logic-Integrated Neural Network (LINN) to integrate the power of deep learning and logic reasoning. LINN is a dynamic neural architecture that builds the computational graph according to input logical expressions. It learns basic logical operations such as AND, OR, NOT as neural modules, and conducts propositional logical reasoning through the network for inference. Experiments on theoretical task show that LINN achieves significant performance on solving logical equations and variables. Furthermore, we test our approach on the practical task of recommendation by formulating the task into a logical inference problem. Experiments show that LINN significantly outperforms state-of-the-art recommendation models in Top-K recommendation, which verifies the potential of LINN in practice.
The ability to conduct logical reasoning is a fundamental aspect of intelligent human behavior, and thus an important problem along the way to human-level artificial intelligence. Traditionally, logic-based symbolic methods from the field of knowledge representation and reasoning have been used to equip agents with capabilities that resemble human logical reasoning qualities. More recently, however, there has been an increasing interest in using machine learning rather than logic-based symbolic formalisms to tackle these tasks. In this paper, we employ state-of-the-art methods for training deep neural networks to devise a novel model that is able to learn how to effectively perform logical reasoning in the form of basic ontology reasoning. This is an important and at the same time very natural logical reasoning task, which is why the presented approach is applicable to a plethora of important real-world problems. We present the outcomes of several experiments, which show that our model is able to learn to perform highly accurate ontology reasoning on very large, diverse, and challenging benchmarks. Furthermore, it turned out that the suggested approach suffers much less from different obstacles that prohibit logic-based symbolic reasoning, and, at the same time, is surprisingly plausible from a biological point of view.
… the neuron's firing, as it is what truly differentiates how a neuron will respond to a given input. In summary, the activity of neuron a i … a nonlinear neuron model in order to generate spikes. …
We present neural algorithmic reasoning—the art of building neural networks that are able to execute algorithmic computation—and provide our opinion on its transformative potential for running classical algorithms on inputs previously considered inaccessible to them.
Humans readily generalize abstract relations, such as recognizing "constant" in shape or color, whereas neural networks struggle to do so, limiting their flexible reasoning. We ask what properties of internal representations enable abstract relational generalization and propose a geometric framework for understanding relational reasoning. We introduce SimplifiedRPM, a controlled benchmark that isolates abstract relational structure and enables rigorous evaluation of generalization to unseen rules. We collect human behavioral data on the task to quantify relational difficulty and enable direct comparison between models and human reasoning. Testing four models-ResNet-50, Vision Transformer, Wild Relation Network, and Scattering Compositional Learner (SCL)-we find that SCL generalizes best and most closely aligns with human difficulty ordering. Using a geometric approach, we model relational rules as manifolds in representation space and show that interpretable geometric quantities accurately predict generalization performance. Layer-wise analysis reveals distinct geometric strategies adopted by different architectures. We further uncover a trade-off between representation signal and dimensionality, with learned and unseen relations aligning with a common low-dimensional subspace. Finally, we show that directly optimizing the representation geometry improves relational generalization in this controlled benchmark, demonstrating that it can be leveraged to guide training. Together, our results demonstrate that representation geometry provides a principled framework for relational generalization in SimplifiedRPM and opens promising avenues for extending geometric analysis to broader cognitive tasks.
In this paper, we propose the Broadcasting Convolutional Network (BCN) that extracts key object features from the global field of an entire input image and recognizes their relationship with local features. BCN is a simple network module that collects effective spatial features, embeds location information and broadcasts them to the entire feature maps. We further introduce the Multi-Relational Network (multiRN) that improves the existing Relation Network (RN) by utilizing the BCN module. In pixel-based relation reasoning problems, with the help of BCN, multiRN extends the concept of ‘pairwise relations’ in conventional RNs to ‘multiwise relations’ by relating each object with multiple objects at once. This yields in \(\mathcal {O}(n)\) complexity for n objects, which is a vast computational gain from RNs that take \(\mathcal {O}(n^2)\). Through experiments, multiRN has achieved a state-of-the-art performance on CLEVR dataset, which proves the usability of BCN on relation reasoning problems.
… Recent studies suggest that visual relational reasoning is … Visual Relational Reasoning abilities of MLLMs. The dataset is divided into Non-Relational and Visual Relational reasoning …
A core process in human cognition is analogical mapping: the ability to identify a similar relational structure between different situations. We introduce a novel task, Visual Analogies of Situation Recognition, adapting the classical word-analogy task into the visual domain. Given a triplet of images, the task is to select an image candidate B' that completes the analogy (A to A' is like B to what?). Unlike previous work on visual analogy that focused on simple image transformations, we tackle complex analogies requiring understanding of scenes. We leverage situation recognition annotations and the CLIP model to generate a large set of 500k candidate analogies. Crowdsourced annotations for a sample of the data indicate that humans agree with the dataset label ~80% of the time (chance level 25%). Furthermore, we use human annotations to create a gold-standard dataset of 3,820 validated analogies. Our experiments demonstrate that state-of-the-art models do well when distractors are chosen randomly (~86%), but struggle with carefully chosen distractors (~53%, compared to 90% human accuracy). We hope our dataset will encourage the development of new analogy-making models. Website: https://vasr-dataset.github.io/
… Others have investigated the use of analogical representations for reasoning about the … many similarities with our approach to visual reasoning. Direct representations have the same …
Capsule networks emerged as a promising alternative to convolutional neural networks for learning object-centric representations. The idea is to explicitly model part-whole hierarchies by using groups of neurons called capsules to encode visual entities, then learn the relationships between these entities dynamically from data. However, a major hurdle for capsule network research has been the lack of a reliable point of reference for understanding their foundational ideas and motivations. This survey provides a comprehensive and critical overview of capsule networks, which aims to serve as a main point of reference going forward. To that end, we introduce the fundamental concepts and motivations behind capsule networks, such as equivariant inference. We then cover various technical advances in capsule routing algorithms as well as alternative geometric and generative formulations. We provide a detailed explanation of how capsule networks relate to the attention mechanism in Transformers and uncover non-trivial conceptual similarities between them in the context of object-centric representation learning. We also review the extensive applications of capsule networks in computer vision, video and motion, graph representation learning, natural language processing, medical imaging, and many others. To conclude, we provide an in-depth discussion highlighting promising directions for future work.
By integrating hard constraints into neural network outputs, we not only improve the reliability of AI systems but also pave the way for meta-cognitive capabilities that ensure the alignment of predictions with domain-specific knowledge. This topic has received a lot of attention, however, existing methods either impose the constraints in a “weak” form at training time, with no guarantees at inference, or fail to provide a general framework that supports different tasks and constraint types. We tackle this open problem from a neuro-symbolic perspective, developing a pipeline that enhances a conventional neural predictor with a symbolic reasoning module capable of correcting structured prediction errors and a neural attention module that learns to direct the reasoning effort to focus on potential prediction errors, while keeping other outputs unchanged. This framework provides an appealing trade-off between the efficiency of constraint-free neural inference and the prohibitive cost of exhaustive reasoning at inference time that satisfies the rigorous demands of meta-cognitive assurance.
Processing Natural Language using machines is not a new concept. Back in 1940 researchers estimated the importance of a machine that could translate one language to another. Further, during 1957-1970 researchers split into two divisions concerning NLP: symbolic and stochastic. This paper presents an extensive review of recent breakthroughs in Neuro Symbolic Artificial Intelligence (NSAI), an area of AI research which seeks to combine traditional rules-based AI approaches with modern deep learning techniques. Neuro Symbolic models have already demonstrated the capability to outperform state-of-the-art deep learning models in domains such as image and video reasoning. Such models not only performed better when trained on a fraction of dataset compared with traditional machine learning models, but also solved an underlined issue called generalization of deep neural network systems. We also find that symbolic models are good in visual question answering (VQA). In this paper, we also review research results related to Neuro Symbolic AI with the objective of exploring the importance of such AI systems and how it would shape the future of AI as a whole. We discuss different types of dataset of Visual Question Answering (VQA) tasks based on NSAI and extensive comparison of performance of different NSAI models. Later, the article focuses on the contemporary real time application of NSAI systems and how NSAI is shaping the world’s different sectors including finance, healthcare, and cyber security.
Traditional AI struggles with interpretability and generalisation in complex reasoning tasks, limiting its effectiveness in domains like healthcare and robotics. This systematic review aims to evaluate Neuro-Symbolic Reasoning (NeSy) frameworks, which integrate symbolic reasoning with neural networks to address these challenges. 28 empirical studies (2017–2024) were analysed from arXiv, IEEE Xplore, PubMed, and conferences, using a PRISMA-guided methodology with inclusion criteria focusing on NeSy frameworks, performance, and scalability. Results show NeSy systems achieve a mean accuracy of 93.00% (SD 5.35%) across visual reasoning, NLP, robotics, and healthcare, outperforming neural baselines by 26.00% on average (SD 18.29%). Methodologies like pLogicNet, DiffLogic, and NSFR enhance generalisation, e.g., in spatial reasoning tasks. However, computational inefficiencies and explainability gaps persist (mean quality score 7.53/9, SD 1.04). NeSyBench, using datasets like MIMIC-III and CLEVR, and NeSyEval for standardised metrics (accuracy, F1-score, interpretability), was proposed to refine NeSy systems. This review provides a roadmap for developing interpretable, scalable AI, advancing applications in diagnostics and autonomous systems.
Interpreting visual scenes is a complex challenge in artificial intelligence (AI), requiring the integration of pattern recognition and logical reasoning. This paper introduces neuro-symbolic based scene interpretation (NeuroScene), a novel approach combining neural networks and symbolic AI for enhanced scene understanding. NeuroScene employs a pre-trained ResNet for feature extraction and Mask R-CNN for identifying regions of interest. These features are converted into homogeneous concept representations using vector quantization. The semantic parsing module translates natural language questions into executable programs using operations from the CLEVR dataset, which are then used for neuro-symbolic reasoning over the constructed scene graph. Experiments on the CLEVR dataset demonstrate NeuroScene's effectiveness in interpreting complex visual scenes and providing interpretable answers. By integrating neural and symbolic components, NeuroScene ensures that high-level reasoning is informed by precise visual features, enhancing the clarity and usefulness of the model's outputs. This approach shows promise for applications requiring detailed scene understanding and logical reasoning.
As the applications of robots expand across a wide variety of areas, high-level task planning considering human–robot interactions is emerging as a critical issue. Various elements that facilitate flexible responses to humans in an ever-changing environment, such as scene understanding, natural language processing, and task planning, are thus being researched extensively. In this study, a visual question answering (VQA) task was examined in detail from among an array of technologies. By further developing conventional neuro-symbolic approaches, environmental information is stored and utilized in a symmetric graph format, which enables more flexible and complex high-level task planning. We construct a symmetric graph composed of information such as color, size, and position for the objects constituting the environmental scene. VQA, using graphs, largely consists of a part expressing a scene as a graph, a part converting a question into SPARQL, and a part reasoning the answer. The proposed method was verified using a public dataset, CLEVR, with which it successfully performed VQA. We were able to directly confirm the process of inferring answers using SPARQL queries converted from the original queries and environmental symmetric graph information, which is distinct from existing methods that make it difficult to trace the path to finding answers.
Abstract visual reasoning (AVR) domain encompasses problems solving which requires the ability to reason about relations among entities present in a given scene. While humans, generally, solve AVR tasks in a “natural” way, even without prior experience, this type of problem has proven difficult for current machine learning systems. The paper summarises recent progress in applying deep learning methods to solving AVR problems, as a proxy for studying machine intelligence. We focus on the most common type of AVR tasks—the Raven’s Progressive Matrices (RPMs)—and provide a comprehensive review of the learning methods and deep neural models applied to solve RPMs, as well as present the RPM benchmark sets. Performance analysis of the state-of-the-art approaches to solving RPMs leads to formulation of certain insights and remarks on the current and future trends in this area. We also attempt to put RPM studies in a more general perspective and demonstrate how real-world problems from the outside of AVR area can benefit from the presented research.
Visual abstract reasoning tasks present challenges for deep neural networks, exposing limitations in their capabilities. In this work, we present a neural network model that addresses the challenges posed by Raven’s Progressive Matrices (RPM). Inspired by the two-stream hypothesis of visual processing, we introduce the Dual-stream Reasoning Network (DRNet), which utilizes two parallel branches to capture image features. On top of the two streams, a reasoning module first learns to merge the high-level features of the same image. Then, it employs a rule extractor to handle combinations involving the eight context images and each candidate image, extracting discrete abstract rules and utilizing an multilayer perceptron (MLP) to make predictions. Empirical results demonstrate that the proposed DRNet achieves state-of-the-art average performance across multiple RPM benchmarks. Furthermore, DRNet demonstrates robust generalization capabilities, even extending to various out-of-distribution scenarios. The dual streams within DRNet serve distinct functions by addressing local or spatial information. They are then integrated into the reasoning module, leveraging abstract rules to facilitate the execution of visual reasoning tasks. These findings indicate that the dual-stream architecture could play a crucial role in visual abstract reasoning.
… reasoning ability in abstract visual reasoning tasks. The extensive experiments on the major datasets (RAVEN, I-RAVEN … the state-of-the-art reasoning performance. Especially, we can …
<abstract xmlns="http://www.w3.org/1999/xhtml"> Learning to perform abstract reasoning often requires decomposing the task in question into intermediate subgoals that are not specified upfront, but need to be autonomously devised by the learner. In Raven progressive matrices (RPMs), the task is to choose one of the available answers given a context, where both the context and answers are composite images featuring multiple objects in various spatial arrangements. As this high-level goal is the only guidance available, learning to solve RPMs is challenging. In this study, we propose a deep learning architecture based on the transformer blueprint which, rather than directly making the above choice, addresses the subgoal of predicting the visual properties of individual objects and their arrangements. The multidimensional predictions obtained in this way are then directly juxtaposed to choose the answer. We consider a few ways in which the model parses the visual input into tokens and several regimes of masking parts of the input in self-supervised training. In experimental assessment, the models not only outperform state-of-the-art methods but also provide interesting insights and partial explanations about the inference. The design of the method also makes it immune to biases that are known to be present in some RPM benchmarks. </abstract>
… Datasets To promote research on abstract visual reasoning in computers, several datasets are designed to simulate Raven’… in the RAVEN dataset, the IRAVEN dataset is proposed [25]. …
Advances in computer vision research enable human-like high-dimensional perceptual induction over analogical visual reasoning problems, such as Raven's Progressive Matrices (RPMs). In this paper, we propose a Hierarchical Perception and Predictive Analogy-Inference network (HP^2AI), consisting of three major components that tackle key challenges of RPM problems. Firstly, in view of the limited receptive fields of shallow networks in most existing RPM solvers, a perceptual encoder is proposed, consisting of a series of hierarchically coupled Patch Attention and Local Context (PALC) blocks, which could capture local attributes at early stages and capture the global panel layout at deep stages. Secondly, most methods seek for object-level similarities to map the context images directly to the answer image, while failing to extract the underlying analogies. The proposed reasoning module, Predictive Analogy-Inference (PredAI), consists of a set of Analogy-Inference Blocks (AIBs) to model and exploit the inherent analogical reasoning rules instead of object similarity. Lastly, the Squeeze-and-Excitation Channel-wise Attention (SECA) in the proposed PredAI discriminates essential attributes and analogies from irrelevant ones. Extensive experiments over four benchmark RPM datasets show that the proposed HP^2AI achieves significant performance gains over all the state-of-the-art methods consistently on all four datasets.
… Existing multimodal reasoning datasets primarily depend on automated construction … For instance, the RAVEN [23] dataset assesses a model’s logical reasoning ability using 2D shapes …
Visual reasoning tasks involving comparison provide interesting insights into how people make similarity and difference judgments. This review summarizes work that provides evidence that the same structure-mapping comparison processes that appear to be used elsewhere in cognition can also be used to model comparison in human visual reasoning tasks. These models rely on qualitative representations, which provide symbolic descriptions of continuous properties, an important kind of relational representation. Cognitive simulations of multiple human visual reasoning tasks, using the same model of high-level vision to compute relational representations, achieve human-like performance, both in terms of accuracy and estimating the relative difficulty of problems.
Significance Humans have an extraordinary ability to learn new concepts from just a few examples, but the cognitive mechanism behind this ability remains mysterious. A long line of work in systems neuroscience has revealed that familiar concepts can be discriminated by the patterns of activity they elicit across neurons in higher-order cortical layers. How might downstream brain areas use this same neural code to learn new concepts? Here, we show that this ability is governed by key geometric properties of the neural code. We find that these geometric quantities undergo orchestrated transformations along the primate visual pathway and along artificial neural networks, so that in higher-order layers, a simple plasticity rule may allow novel concepts to be learned from few examples.
… for the weakly-supervised few-shot learning, and propose a … with multi-level conceptual abstraction to train a universal meta-… on two weakly-supervised few-shot learning benchmarks, …
The abilities to form concepts and abstractions, and to make analogies, are key to human intelligence, but AI systems have a long way to go before they can match the abilities of humans in these areas. To develop machines that can abstract and analogize, researchers typically focus on idealized problem domains that are meant to capture the essence of human abstraction abilities without having to deal with the complexity of real‐world situations. This commentary describes why solving problems in these domains remains difficult for AI systems, and discusses how AI researches can make progress on imbuing machines with these essential abilities.
Highlights • A cohort of 10–16 year-olds completed an abstract reasoning task during MEG.• Performance on the abstract reasoning task correlated with fluid intelligence.• The task was associated with increased cortical dynamics in frontoparietal areas.• Youth showed sexually divergent patterns of distributed cortical activity with age.• Specific frontoparietal activity differentially predicted aspects of task behavior.
… The human brain can solve highly abstract reasoning problems using a neural network that is entirely physical. The underlying mechanisms are only partially understood, but an …
Previous studies in animals and humans suggest that monoamines enhance behavior-evoked neural activity relative to nonspecific background activity (i.e., increase signal-to-noise ratio). We studied the effects of dextroamphetamine, an indirect monoaminergic agonist, on cognitively evoked neural activity in eight healthy subjects using positron-emission tomography and the O15 water intravenous bolus method to measure regional cerebral blood flow (rCBF). Dextroamphetamine (0.25 mg/kg) or placebo was administered in a double-blind, counterbalanced design 2 hr before the rCBF study in sessions separated by 1–2 weeks. rCBF was measured while subjects performed four different tasks: two abstract reasoning tasks—the Wisconsin Card Sorting Task (WCST), a neuropsychological test linked to a cortical network involving dorsolateral prefrontal cortex and other association cortices, and Ravens Progressive Matrices (RPM), a nonverbal intelligence test linked to posterior cortical systems—and two corresponding sensorimotor control tasks. There were no significant drug or task effects on pCO2 or on global blood flow. However, the effect of dextroamphetamine (i.e., dextroamphetamine vs placebo) on task-dependent rCBF activation (i.e., task − control task) showed double dissociations with respect to task and region in the very brain areas that most distinctly differentiate the tasks. In the superior portion of the left inferior frontal gyrus, dextroamphetamine increased rCBF during WCST but decreased it during RPM (ANOVA F(1,7) = 16.72, p < 0.0046). In right hippocampus, blood flow decreased during WCST but increased during RPM (ANOVAF(1,7) = 18.7, p < 0.0035). These findings illustrate that dextroamphetamine tends to “focus” neural activity, to highlight the neural network that is specific for a particular cognitive task. This capacity of dextroamphetamine to induce cognitively specific signal augmentation may provide a neurobiological explanation for improved cognitive efficiency with dextroamphetamine.
Anomaly detection approaches for network intrusion detection learn to identify deviations from normal behavior on a data-driven basis. However, current approaches strive to infer the degree of abnormality of out-of-distribution samples when these appertain to different zero-day attacks. Inspired by the successes of the neural algorithmic reasoning paradigm to leverage the generalization of rule-based behavior, this paper presents a deep learning strategy for solving zero-day network attack detection and categorization. Moreover, focusing on the particular scenario of the Internet of Things (IoT), the privacy preservation requirement may imply a low training data regime for any learning algorithm. To this respect, the presented framework uses metric-based meta-learning to achieve few-shot learning capabilities. The presented pipeline is called NERO , as it imports the encode-process-decode architecture from the NE ural algorithmic reasoning blueprint to converge ze RO -day attack detection policies within constrained training data.
Abstract The recently emerged research of Visual Question Answering (VQA) has become a hot topic in computer vision. A key solution to VQA exists in how to fuse multimodal features extracted from image and question. In this paper, we show that combining visual relationship and attention together achieves more fine-grained feature fusion. Specifically, we design an effective and efficient module to reason complex relationship between visual objects. In addition, a bilinear attention module is learned for question guided attention on visual objects, which allows us to obtain more discriminative visual features. Given an image and a question in natural language, our VQA model learns visual relational reasoning network and attention network in parallel to fuse fine-grained textual and visual features, so that answers can be predicted accurately. Experimental results show that our approach achieves new state-of-the-art performance of single model on both VQA 1.0 and VQA 2.0 datasets.
… multimodal large language models to extract visual details as contextual prompts for caption … In this paper, we have proposed a novel relational reasoning image captioning framework …
… Emerging evidence has demonstrated the close association between relational reasoning (… In addition to the existing measures of nonverbal relational reasoning (RR), the current study …
In this paper, we propose a dynamic graph modeling approach to learn spatial-temporal representations for video summarization. Most existing video summarization methods extract image-level features with ImageNet pre-trained deep models. Differently, our method exploits object-level and relation-level information to capture spatial-temporal dependencies. Specifically, our method builds spatial graphs on the detected object proposals. Then, we construct a temporal graph by using the aggregated representations of spatial graphs. Afterward, we perform relational reasoning over spatial and temporal graphs with graph convolutional networks and extract spatial-temporal representations for importance score prediction and key shot selection. To eliminate relation clutters caused by densely connected nodes, we further design a self-attention edge pooling module, which disregards meaningless relations of graphs. We conduct extensive experiments on two popular benchmarks, including the SumMe and TVSum datasets. Experimental results demonstrate that the proposed method achieves superior performance against state-of-the-art video summarization methods.
Abstract Starting with the hypothesis that analogical reasoning consists of a search of semantic space, we used eye‐tracking to study the time course of information integration in adults in various formats of analogies. The two main questions we asked were whether adults would follow the same search strategies for different types of analogical problems and levels of complexity and how they would adapt their search to the difficulty of the task. We compared these results to predictions from the literature. Machine learning techniques, in particular support vector machines (SVMs), processed the data to find out which sets of transitions best predicted the output of a trial (error or correct) or the type of analogy (simple or complex). Results revealed common search patterns, but with local adaptations to the specifics of each type of problem, both in terms of looking‐time durations and the number and types of saccades. In general, participants organized their search around source‐domain relations that they generalized to the target domain. However, somewhat surprisingly, over the course of the entire trial, their search included, not only semantically related distractors, but also unrelated distractors, depending on the difficulty of the trial. An SVM analysis revealed which types of transitions are able to discriminate between analogy tasks. We discuss these results in light of existing models of analogical reasoning.
… In the present analyses, we found that the children's performance on the visual analogical reasoning task could offer some concurrent prediction of children's orthographic …
… pretest to posttest in figural analogy solving, as … reasoning skills to an analogy construction task was related to initial ability, but not working memory; transfer to two inductive reasoning …
… Two hundred fifty-one students from first to third grades were tested on visual-spatial … roles of analogical reasoning. These results highlight the important role of analogical reasoning in …
Analogical reasoning allows children to generate new abstractions from experience, which drives early learning. But our current understanding of analogical learning is based primarily on evidence from the West, and new data from Asia seem to call into question that information. In this article, we describe our finding that East Asian children do not share their Western peers’ strong bias for object similarity—often cited as the major reason for difficulties in relational reasoning. We analyze how this difference affects the ways analogy shapes learning in different cultures, such as what children use as base analogs and their likelihood of using comparison that results in relational abstraction. We also address crosscultural differences, which are evident in classroom contexts, since teachers from the United States and East Asia use analogy differently. Overall, cross-cultural data are necessary to answer critical questions in theories of analogical learning; in this article, we chart pressing research questions and look ahead at directions for the field. KEYWORDS—analogy; cross-cultural development; learning; cognitive development Learning and development are about making sense of the new: acquiring new words and concepts, and discovering methods to solve novel problems. Analogical reasoning—the ability to perceive similarity of relations across events—is a powerful mechanism that facilitates learning. In reasoning analogically, learners map structures or relations from a familiar situation onto a novel one (Gentner, 1983). For example, if Jen knows that putting larger blocks atop small ones makes her block towers topple, she can put the smaller stool on top of the larger one when reaching for a high-placed cookie jar. The essence of analogical reasoning as a learning tool is that Jen makes a prediction about a novel situation given her knowledge of earlier setups, even if the earlier situation differs from the current task. The prominence of analogical reasoning in learning theories is paramount. It plays a fundamental role in language learning, problem solving, creativity, scientific discovery, and social cognition (Christie, 2017; see Gentner & Hoyos, 2017, for a review). Some have even suggested that higher-order relational reasoning is what separates the cognition of humans and nonhuman animals (Gentner, 2003, 2010; Penn, Holyoak, & Povinelli, 2008). Analogical reasoning also plays an important role in educational theories and testing (Spearman, 1927). In fact, one of the most widely used intelligence tests—Raven’s Progressive Matrices (Raven, 1941)—is essentially an analogical reasoning test. However, despite the centrality and claimed universality of analogy, almost everything we know about the development of analogical reasoning comes from studies conducted in the West (i.e., the United States, Europe, and Australia). The lack of evidence from other cultures is a critical gap in how we understand analogy and in how much confidence we place in our understanding of the concept. This gap would not be so concerning if preliminary data from other cultural milieus mirrored the results of studies done in the West; however, for the most part, this seems not to be the case. A well-known early example is a study of scene analogy tasks in which preschoolers from Hong Kong outperformed their U.S. peers when complex relations were involved (Richland, Chan, Morrison, & Au, 2010). Stella Christie, Tsinghua Laboratory of Brain and Intelligence and Department of Psychology, Tsinghua University; Yang Gao, Department of Psychology, Tsinghua University; Qiuchen Ma, Department of Psychology, Tsinghua University. This work is supported by Tsinghua University Initiative Scientific Research Program (grant number 20197010006). This research has no known conflict of interest to disclose. Stella Christie thanks Bartlomiej Czech for discussions on ideas in this article Correspondence concerning this article should be addressed to Stella Christie, Department of Psychology and Tsinghua Laboratory for Brain and Intelligence, Tsinghua University, Beijing, 100084, China. e-mail: christie@tsinghua.edu.cn. © 2020 Society for Research in Child Development DOI: 10.1111/cdep.12380 Volume 0, Number 0, 2020, Pages 1–7 CHILD DEVELOPMENT PERSPECTIVES
Video Scene Graph Generation (VidSGG) aims to represent dynamic visual content by detecting objects and modeling their temporal interactions as structured graphs. Prior studies typically target either coarse-grained box-level or fine-grained panoptic pixel-level VidSGG, often requiring task-specific architectures and multi-stage training pipelines. In this paper, we present UNO (UNified Object-centric VidSGG), a single-stage, unified framework that jointly addresses both tasks within an end-to-end architecture. UNO is designed to minimize task-specific modifications and maximize parameter sharing, enabling generalization across different levels of visual granularity. The core of UNO is an extended slot attention mechanism that decomposes visual features into object and relation slots. To ensure robust temporal modeling, we introduce object temporal consistency learning, which enforces consistent object representations across frames without relying on explicit tracking modules. Additionally, a dynamic triplet prediction module links relation slots to corresponding object pairs, capturing evolving interactions over time. We evaluate UNO on standard box-level and pixel-level VidSGG benchmarks. Results demonstrate that UNO not only achieves competitive performance across both tasks but also offers improved efficiency through a unified, object-centric design.
… perception and reasoning, the neuro-symbolic framework … , in a way that perception can inform reasoning and vice-versa. In … neuro-symbolic framework that enables symbolic reasoning …
Advertisement (Ad) video violation detection is critical for ensuring platform compliance, but existing methods struggle with precise temporal grounding, noisy annotations, and limited generalization. We propose RAVEN, a novel framework that integrates curriculum reinforcement learning with multimodal large language models (MLLMs) to enhance reasoning and cognitive capabilities for violation detection. RAVEN employs a progressive training strategy, combining precisely and coarsely annotated data, and leverages Group Relative Policy Optimization (GRPO) to develop emergent reasoning abilities without explicit reasoning annotations. Multiple hierarchical sophisticated reward mechanism ensures precise temporal grounding and consistent category prediction. Experiments on industrial datasets and public benchmarks show that RAVEN achieves superior performances in violation category accuracy and temporal interval localization. We also design a pipeline to deploy the RAVEN on the online Ad services, and online A/B testing further validates its practical applicability, with significant improvements in precision and recall. RAVEN also demonstrates strong generalization, mitigating the catastrophic forgetting issue associated with supervised fine-tuning.
Semantic segmentation in remote sensing images, which is the task of classifying each pixel of the image in a specific category, is widely used in areas such as disaster management, environmental monitoring, precision agriculture, and many others. However, traditional semantic segmentation methods face a major challenge: they require large amounts of annotated data to train effectively. To tackle this challenge, few-shot semantic segmentation has been introduced, where the models can learn and adapt quickly to new classes from just a few annotated samples. This paper presents a comprehensive review of recent advances in few-shot semantic segmentation (FSSS) for remote sensing, covering datasets, methods, and emerging research directions. We first outline the fundamental principles of few-shot learning and summarize commonly used remote-sensing benchmarks, emphasizing their scale, geographic diversity, and relevance to episodic evaluation. Next, we categorize FSSS methods into major families (meta-learning, conditioning-based, and foundation-assisted approaches) and analyze how architectural choices, pretraining strategies, and inference protocols influence performance. The discussion highlights empirical trends across datasets, the behavior of different conditioning mechanisms, the impact of self-supervised and multimodal pretraining, and the role of reproducibility and evaluation design. Finally, we identify key challenges and future trends, including benchmark standardization, integration with foundation and multimodal models, efficiency at scale, and uncertainty-aware adaptation. Collectively, they signal a shift toward unified, adaptive models capable of segmenting novel classes across sensors, regions, and temporal domains with minimal supervision.
… We constructed a few-shot node classification job during training, treating the training labels as known. We chose 400 few-shot tasks from the validation and test labels to measure …
本报告对抽象视觉推理研究进行了系统整合,分为五大核心维度:RPM矩阵推理、神经符号推理、认知科学视角、表征几何理论以及复杂场景的多模态推理。研究重心已从早期的特定基准评测,扩展至融合人类认知机制的架构设计,以及通过表征几何等数学工具揭示深层逻辑泛化能力的理论探索,展示了AI从简单逻辑匹配迈向通用推理的趋势。