基于图数据库的财务、税务领域的报表数据逻辑验证与自动化测试方法
财务领域知识图谱构建与数据增强
这些文献主要关注如何从非结构化财务报告或外部数据中自动提取信息,构建高质量的财务知识图谱,并探讨如何通过大语言模型增强数据处理能力。
- Data Set and Evaluation of Automated Construction of Financial Knowledge Graph(Wenguang Wang, Yonglin Xu, Chunhui Du, Yunwen Chen, Yijie Wang, Hui Wen, 2021, Data Intelligence)
- An automated information extraction system from the knowledge graph based annual financial reports(Syed Farhan Mohsin, S. I. Jami, Shaukat Wasi, Muhammad Shoaib Siddiqui, 2024, PeerJ Computer Science)
- Financial Knowledge Graph Based Financial Report Query System(Samreen Zehra, Syed Farhan Mohsin, Shaukat Wasi, S. I. Jami, Muhammad Shoaib Siddiqui, Muhammad Khaliq-Ur-Rahman Raazi Syed, 2021, IEEE Access)
- FinKario: Event-Enhanced Automated Construction of Financial Knowledge Graph(Xiang Li, Penglei Sun, Wanyun Zhou, Zikai Wei, Yongqi Zhang, Xiaowen Chu, 2026, Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers))
- HybridRAG-Finance: Agent-Guided Information Summarization with Vector and Graph Retrieval for Financial Documents(Mazen Wael Baioumy, Konstantinos N. Plataniotis, Yuri Lawryshyn, 2026, SSRN Electronic Journal)
智能审计与财务欺诈检测
该组文献集中研究如何利用图技术在审计领域识别风险、关联交易分析及欺诈检测,重点在于提升审计的透明度和识别准确性。
- From data to insights: the application and challenges of knowledge graphs in intelligent audit(Hao Zhong, Dong Yang, Shengdong Shi, Lai Wei, Yanyan Wang, 2024, Journal of Cloud Computing)
- Using graph databases to detect financial fraud(Richard Henderson, 2020, Computer Fraud & Security)
- Financial fraud risk analysis based on audit information knowledge graph(Huidong Wu, Ya-Yun Chang, Jianping Li, Xiaoqian Zhu, 2021, Procedia Computer Science)
- A Graph Mining Approach to Identify Financial Reporting Patterns: An Empirical Examination of Industry Classifications(Steve Y. Yang, Fang-Chun Liu, Xiaodi Zhu, David C. Yen, 2018, Decision Sciences)
- Intelligent audit decision support for enterprise related-party transactions based on knowledge graph and graph neural network(Ruining Hao, Yao Sun, 2026, Scientific Reports)
- A Study of Intelligent Financial Fraud Detection Incorporating Knowledge Graph and Graph Visualisation(Huxiang Xu, 2025, 2025 10th International Conference on Cyber Security and Information Engineering (ICCSIE))
- A Knowledge Graph-Based Data Provenance Framework for Risk Indicators in Intelligent Auditing(Zhichao Jiang, Linfan Xia, Zexin Qu, Bo Han, Yifan Wu, 2026, 2026 6th International Conference on Neural Networks, Information and Communication Engineering (NNICE))
多源数据一致性验证与质量控制
这些文献探讨如何利用图数据库和逻辑推理方法,在复杂的业务系统、合规性要求或多源数据背景下实现数据一致性检查与验证。
- Consistency Meets Inconsistency: A Unified Graph Learning Framework for Multi-view Clustering(Youwei Liang, Dong Huang, Changdong Wang, 2019, 2019 IEEE International Conference on Data Mining (ICDM))
- Consistency and Visual Checking for Exploration and Development Data Standards Based on Knowledge Graph(Min Cui, Weichong Li, Zhiyong Deng, Jinbo Zhang, 2025, 2025 10th International Conference on Intelligent Computing and Signal Processing (ICSP))
- A Knowledge Graph-Based Consistency Detection Method for Network Security Policies(Yaang Chen, Teng Hu, Fang Lou, Mingyong Yin, Tao Zeng, Guo Wu, Hao Wang, 2024, Applied Sciences)
- A graph-based algorithm for consistency maintenance in incremental and interactive integration tools(S. Becker, Sebastian Herold, Sebastian Lohmann, B. Westfechtel, 2007, Software & Systems Modeling)
- Integrity constraints in graph databases(J. Pokorný, M. Valenta, Jirí Kovacic, 2017, Procedia Computer Science)
- Graph-based inter-domain consistency maintenance for BIM models(Zijian Wang, Boyuan Ouyang, Rafael Sacks, 2023, Automation in Construction)
- Efficient and Scalable Integrity Verification of Data and Query Results for Graph Databases(M. Arshad, A. Kundu, E. Bertino, A. Ghafoor, Chinmay Kundu, 2018, 2018 IEEE 34th International Conference on Data Engineering (ICDE))
- Network evaluation from the consistency of the graph structure with the measured data(Shigeru Saito, S. Aburatani, K. Horimoto, 2008, BMC Systems Biology)
- A Semi-automatic Semantic Consistency-Checking Method for Learning Ontology from Relational Database(Chuangtao Ma, B. Molnár, A. Benczúr, 2021, Information)
合规监管与法律契约审计
该组论文聚焦于金融和安全领域中复杂的监管框架(如GDPR)及组织内规制的落地,通过图技术实现契约、政策的自动化合规性验证。
- An Organization-Scoped LLM Agent Runtime Architecture for Regulated Cybersecurity Operations(George Fatouros, Georgios Makridis, George Kousiouris, John Soldatos, D. Kyriazis, 2026, arXiv.org)
- Automated GDPR Contract Compliance Verification Using Knowledge Graphs(Amar Tauqeer, Anelia Kurteva, T. Chhetri, Albin Ahmeti, A. Fensel, 2022, Information)
- Connecting the dots: Graph neural networks for auditing accounting journal entries(Q Huang, M Schreyer, NR Michiles Jr, 2026, Auditing: A Journal …)
图模型设计与企业网络分析
这些文献涵盖了图数据库的建模方法、数据集成技术,以及如何利用图分析处理大规模企业关系网络,以实现风险关联挖掘与复杂数据处理。
- 一种基于概率图模型的关联规则更新方法(蔡鹏飞, 岳昆, 李雪, 刘惟一)
- Commonsense Reasoning and Large Network Analysis: A Computational Study of ConceptNet 4(Dimitrios I. Diochnos, 2013, arXiv.org)
- Comprehensive management method of financial data based on knowledge graph(N.A. Yang, 2026, International Journal of Computer Applications in Technology)
- Business Registers Data Processing for Entity Resolution and Relationship Mapping Using Graph Database(Richard Marko, Lukáš Grejták, 2025, 2025 International Conference on Innovation and Intelligence for Informatics, Computing, and Technologies (3ICT))
- Designing Graph Databases With GRAPHED(G. V. Erven, Rommel N. Carvalho, Waldeyr M. C. Silva, Sérgio Lifschitz, H. V. Olivera, M. Holanda, 2019, Journal of Database Management)
- 基于杜邦分析与大模型驱动的交互式财务数据决策支持可视分析(张潇文, 万怡, 王芬, 陈玄, 赵宇恒, 高琳, 陈思明, 2024, 计算机辅助设计与图形学学报)
- Establishing the representational faithfulness of financial accounting information using multiparty security, network analysis and a blockchain(John McCallig, Alastair Robb, Fiona H. Rohde, 2019, International Journal of Accounting Information Systems)
- Does automation improve financial reporting? Evidence from internal controls(Musaib Ashraf, 2024, Review of Accounting Studies)
- Hybrid Neuro-Symbolic Models for Ethical AI in Risk-Sensitive Domains(Chaitanya Kumar Kolli, 2025, arXiv.org)
- 基于知识图谱网络特征的中国外汇市场系统性风险测度研究(沈嘉贤, 陈浩智, 张卫国, 2025, 中国管理科学)
- A Graph-Based Data Model and its Ramifications(M. Levene, G. Loizou, 1995, IEEE Transactions on Knowledge and Data Engineering)
- Leveraging Graph-RAG and Prompt Engineering to Enhance LLM-Based Automated Requirement Traceability and Compliance Checks(Arsalan Masoudifard, Mohammad Mowlavi Sorond, Moein Madadi, M. Sabokrou, Elahe Habibi, 2024, arXiv.org)
该研究集合围绕图数据库在财务与税务领域中的应用,探讨了从知识图谱自动化构建、智能审计风险识别、跨系统数据一致性验证到合规性自动监测等核心技术路径,展示了图技术在复杂金融业务数据处理与决策支持中的深度价值。
总计36篇相关文献
针对事务库发生变化后关联规则的更新问题,讨论了一种只对具有实用价值的关联规则更新其前件的方法.首先分析了关联规则各组件间的依赖关系及其不确定性,进而建立描述其中所蕴含不确定性知识的贝叶斯网模型(称为规则贝叶斯网),并提出了基于Gibbs采样的规则贝叶斯网近似推理算法,从而实现关联规则的更新.实验结果表明,作者提出的基于概率图模型的关联规则前件更新方法具有高效性和可行性.
… as part of the financial statement audit. This responsibility is … Conformance checking: Journal entries are compliant if they … in labeled multi-graph databases. ACM Transactions on …
… such as (10-k forms, and annual reports), extracting key … Neo4j is a graph database that represents data as nodes with … This enables what we call "double retrieval" verification, where …
Annual Financial Reports are the core in the Banking Sector to publish its financial statistics. Extracting useful information from these complex and lengthy reports involves manual process to resolve the financial queries, resulting in delays and ambiguity in investment decisions. One of the major reasons is the lack of any standardization in the format and vocabulary used in the reports. An automated system for resolution of intelligent financial queries is therefore difficult to design. Several works have been proposed to overcome these problems using Information Extraction; however, they do not address the semantic interoperability of the reports across different institutions. This work proposed an automated querying engine to answer the financial queries using Ontology based Information Extraction. For Semantic modeling of financial reports, a Financial Knowledge Graph, assisted by Financial Ontology, has been proposed. The nodes are populated with entities, while links are populated with relationships using Information Extraction applied on annual reports. Two benefits have been provided by this system to stakeholders through automation: decision making through queries and generation of custom financial stories. The work can further be extended to other domains including healthcare and academia where physical reports are used for communication.
… is constructed to verify signature data … financial news text features with financial report data while extracting time-series information. The method analyses both structured financial report …
We evaluate the current state of international business registers by analyzing their accessibility, completeness, level of digitalization, and associated costs across jurisdictions. We introduce a unified data pipeline designed to collect, clean, tokenize, normalize, and integrate heterogeneous records, with a focus on resolving entities and mapping cross-border relationships between individuals and legal entities. The system automates the acquisition of structured data from Slovak and Czech business registers and currently maintains information on over 500,000 individuals. Our methodology combines string normalization, fuzzy and phonetic matching, and graph database techniques to improve the accuracy of entity resolution. In a case study involving 84 companies, the system achieved a deduplication rate of 69.2 %, reducing 1,952 person records to 601 unique individuals. The deduplication logic is applied not only to persons, but also to legal entities, addresses, and positions, enabling consistent and reliable record consolidation across the entire data model. The system constructs large-scale corporate graphs with hundreds of thousands of nodes and relationships, uncovering previously unlinked ownership and management connections. The tool is accessible via a web interface and REST API, and provides a foundation for scalable, automated corporate network analysis for financial analysts, regulators, and investigative researchers. These results confirm that the approach yields high-accuracy consolidation of fragmented registry records and reveals connections that are not directly observable in raw sources. A public demo of the web interface is available at https://pep.adversea.com.
… results for graph redaction and verification times for shortest-path and neighborhood queries as executed on the graph … generation and present the average results for 10 such queries. …
In recent years, graph database systems have become very popular and been deployed mainly in situations where the relationship between data is significant, such as in social networks. Although they do not require a particular schema design, a data model contributes to their consistency. Designing diagrams is an approach to satisfying this demand for a conceptual data model. While researchers and companies have been developing concepts and notations for graph database modeling, their notations focus on their specific implementations. In this article, the authors propose a diagram to address this lack of a generic and comprehensive notation for graph databases modeling, named GRAPHED (Graph Description Diagram for Graph Databases). The authors verified the effectiveness and compatibility of GRAPHED in two case studies: fraud identification, and a biological network model.
Network security policy is regarded as a guideline for the use and management of the network environment, which usually formulates various requirements in the form of natural language. It can help network managers conduct standardized network attack detection and situation awareness analysis in the overall time and space environment of network security. However, in most cases, due to configuration updates or policy conflicts, there are often differences between the real network environment and network security policies. In this case, the consistency detection of network security policies is necessary. The previous consistency detection methods of security policies have some problems. Firstly, the detection direction is single, only focusing on formal reasoning methods to achieve logical consistency detection and solve problems. Secondly, the detection policy field is not comprehensive, focusing only on a certain type of problem in a certain field. Thirdly, there are numerous forms of data structures used for consistency detection, and it is difficult to unify the structured processing and analysis of rule library carriers and target information carriers. With the development of intelligent graph and data mining technology, the above problems have the possibility of optimization. This article proposes a new consistency detection approach for network security policy, which uses an intelligent graph database as a visual information carrier, which can widely connect detection information and achieve comprehensive detection across knowledge domains, physical devices, and detection methods. At the same time, it can also help users grasp the security associations with the real network environment based on the graph algorithm of the knowledge graph and intelligent reasoning. Furthermore, these actual network situations and knowledge bases can help managers improve policies more tailored to local conditions. This article also introduces the consistency detection process of typical cases of network security policies, demonstrating the practical details and effectiveness of this method.
Data standards are significant achievements of data governance, and they are also the foundation for data applications, especially the construction of artificial intelligence scenarios. The consistency of data items within data standards is a crucial part of standard construction. The data standards for offshore exploration and development have given rise to approximately 10000 standard tables and 190000 data items. Due to reasons such as domain-based governance, complex sources, and manual verification, phenomena like diverse and and inconsistent encodings for the same data item, as well as spelling mistakes, have seriously affected data modeling and data services. Meanwhile, data standards presented in tabular form have poor readability, and it is almost impossible to achieve consistency checking between tables. Because of the huge workload, it is difficult to achieve full coverage, and the quality highly depends on experts. There is an urgent need for data visualization to aid in standard writing and checking. The solution based on knowledge graphs, which consists of nodes and edges, can represent entities and the relationships between them. By sorting out and establishing the relationships between data items, constructing a graph database, and presenting the relationships between all data tables, data items, and their encodings in the form of knowledge graphs, it can effectively help with the comparison and checking of data standardization and consistency, and realize the visualization of data content and relationships.
A Semi-automatic Semantic Consistency-Checking Method for Learning Ontology from Relational Database
To tackle the issues of semantic collision and inconsistencies between ontologies and the original data model while learning ontology from relational database (RDB), a semi-automatic semantic consistency checking method based on graph intermediate representation and model checking is presented. Initially, the W-Graph, as an intermediate model between databases and ontologies, was utilized to formalize the semantic correspondences between databases and ontologies, which were then transformed into the Kripke structure and eventually encoded with the SMV program. Meanwhile, description logics (DLs) were employed to formalize the semantic specifications of the learned ontologies, since the OWL DL showed good semantic compatibility and the DLs presented an excellent expressivity. Thereafter, the specifications were converted into a computer tree logic (CTL) formula to improve machine readability. Furthermore, the task of checking semantic consistency could be converted into a global model checking problem that could be solved automatically by the symbolic model checker. Moreover, an example is given to demonstrate the specific process of formalizing and checking the semantic consistency between learned ontologies and RDB, and a verification experiment was conducted to verify the feasibility of the presented method. The results showed that the presented method could correctly check and identify the different kinds of inconsistencies between learned ontologies and its original data model.
Abstract: One thing that is still being developed for graph databases is integrity constraint (IC) support. One possibility to IC proposal is to consider a graph conceptual schema and a graph database schema. At least inherent ICs coming from a graph conceptual schema should be considered as explicit ICs on the graph databases level, i.e., using a DDL. In the paper, we focus on graph database Neo4j and its possibilities to express a database schema and ICs. We extend these possibilities through new constructs in Neo4j DDL including their prototype implementation and experiments.
In response to prevalent issues in traditional auditing—specifically obscure indicator origins, fragmented data lineage, and the underutilization of unstructured evidence—this study presents a novel data provenance framework for risk indicators in intelligent auditing. We construct a comprehensive indicator-data mapping model that systematically delineates data lineage, including sources, schema structures, and processing pipelines. By integrating metadata management, field-level semantic alignment, and knowledge graph modeling, our approach ensures verifiable traceability of structured data throughout its lifecycle. Furthermore, we leverage NLP techniques, such as entity extraction and vectorization, to convert unstructured assets like audit reports and rectification records into an accessible semantic knowledge base, offering robust supplementary evidence. Empirical evaluations demonstrate that this framework significantly improves source transparency, data consistency, and query efficiency. Ultimately, it establishes a solid data foundation for risk identification, model explainability, and audit decision-making.
Ensuring that Software Requirements Specifications (SRS) align with higher-level organizational or national requirements is vital, particularly in regulated environments such as finance and aerospace. In these domains, maintaining consistency, adhering to regulatory frameworks, minimizing errors, and meeting critical expectations are essential for the reliable functioning of systems. The widespread adoption of large language models (LLMs) highlights their immense potential, yet there remains considerable scope for improvement in retrieving relevant information and enhancing reasoning capabilities. This study demonstrates that integrating a robust Graph-RAG framework with advanced prompt engineering techniques, such as Chain of Thought and Tree of Thought, can significantly enhance performance. Compared to baseline RAG methods and simple prompting strategies, this approach delivers more accurate and context-aware results. While this method demonstrates significant improvements in performance, it comes with challenges. It is both costly and more complex to implement across diverse contexts, requiring careful adaptation to specific scenarios. Additionally, its effectiveness heavily relies on having complete and accurate input data, which may not always be readily available, posing further limitations to its scalability and practicality.
Regulated cybersecurity workflows lack a runtime substrate that enforces organization-level scope across retrieval, tool calls, memory, findings, reports, and audit while remaining model-agnostic and locally deployable. Recent large language model (LLM) agent systems report strong results on isolated cybersecurity tasks, yet they do not by themselves define an auditable platform architecture for regulated security operations centre (SOC) and compliance workflows, where a single analyst may trigger actions that bind the organization, and where the runtime must integrate with existing SIEM/XDR stacks as a primary source of context and alert-driven triggers rather than operate as a standalone analytical layer. This paper proposes an organization-scoped LLM agent runtime architecture for financial cybersecurity. The contribution is a typed Security Context that is created at every entry point, including SIEM/XDR notifications ingested as first-class triggers, and enforced at every component boundary, combined with a shared Runtime Core, logical specialist subagents, a governed Tool Adapter Layer exposing SIEM/XDR query, enrichment, and response primitives under uniform policy and audit, structured findings with evidence references, tiered human-in-the-loop (HITL) gates, and append-only audit. Model Context Protocol (MCP), extended telemetry, digital twins for pentesting, graph retrieval, and federated knowledge sharing are treated as optional extension paths rather than mandatory runtime assumptions. We describe an implementable slice as the architecture's testability surface, and we propose a falsifiable evaluation plan with metric-level pass criteria for architecture readiness, security-policy enforcement, evidence traceability, output quality, and operational observability.
Artificial intelligence deployed in risk-sensitive domains such as healthcare, finance, and security must not only achieve predictive accuracy but also ensure transparency, ethical alignment, and compliance with regulatory expectations. Hybrid neuro symbolic models combine the pattern-recognition strengths of neural networks with the interpretability and logical rigor of symbolic reasoning, making them well-suited for these contexts. This paper surveys hybrid architectures, ethical design considerations, and deployment patterns that balance accuracy with accountability. We highlight techniques for integrating knowledge graphs with deep inference, embedding fairness-aware rules, and generating human-readable explanations. Through case studies in healthcare decision support, financial risk management, and autonomous infrastructure, we show how hybrid systems can deliver reliable and auditable AI. Finally, we outline evaluation protocols and future directions for scaling neuro symbolic frameworks in complex, high stakes environments.
Related-party transactions (RPTs) occupy a central position in contemporary corporate governance, serving as essential channels for resource allocation among affiliated business entities. The growing complexity of ownership arrangements and globalized operations has made RPT auditing increasingly challenging for financial professionals. Traditional audit methodologies struggle with intricate enterprise relationship networks that often span multiple jurisdictions and involve layered ownership through shell companies, rendering manual inspection approaches both time-consuming and incomplete. Despite advances in computational approaches, significant gaps persist in current research and practice. Existing audit systems inadequately capture the multi-dimensional relationships connecting enterprises, shareholders, executives, and transactions. Most machine learning methods treat transactions as independent observations, thereby ignoring the relational context that distinguishes legitimate RPTs from irregular ones. Furthermore, the dynamic nature of enterprise relationships demands adaptive models capable of tracking temporal evolution, yet static snapshots of relationship networks may produce outdated or misleading risk assessments. Four critical challenges must be addressed for effective RPT audit automation: modeling heterogeneous entity types linked through diverse relationship categories with different audit implications; capturing temporal dynamics as ownership structures change through acquisitions and personnel movements; providing interpretable outputs that support professional judgment and regulatory inspection; and integrating domain knowledge with data-driven learning in a coherent framework. This study develops an intelligent decision support system that bridges knowledge graph technology with graph neural network architectures to address these challenges. The proposed framework constructs a domain-specific knowledge graph capturing multi-dimensional enterprise relationships through entity resolution and temporal modeling. An enhanced graph neural network employs heterogeneous attention mechanisms alongside temporal fusion components to learn relationship-aware representations. Interpretable risk assessment emerges through attention weight visualization and risk propagation path analysis. Experimental evaluation on real-world enterprise data demonstrates substantial improvements over baseline methods, with practical deployment confirming utility for detecting undisclosed relationships and pricing anomalies in accounting firm engagements.
The application of intelligent algorithms in the field of financial fraud detection faces the dual challenges of interpretability and integrity verification. In order to overcome the financial fraud detection challenges, the study deeply integrates knowledge graph and graph visualisation techniques to construct an intelligent detection model. The model mines the potential rules of financial data through a two-layer knowledge graph, and uses graph visualisation to achieve linkage verification of business and financial data. Experiments show that the model achieves 94.87%, 94.38%, 93.21% and 94.71% accuracy in detecting four types of frauds, namely, inflated revenues, concealment of related transactions, capitalisation of expenses, and asset impairment manipulation, respectively, and the F1 score is maintained at over 92%. The overall accuracy, precision, recall and F1 score of the model are 94.53%, 93.54%, 93.87% and 93.70%, respectively, which are 4.23%, 5.02%, 5.09% and 5.05% higher than other models. The overall accuracy reached 94.53%, an average improvement of 4.23% compared to other models, and the F1 score was 93.70%, an average improvement of 5.05%. Meanwhile, the Area under the Curve (AUC) value reached 0.951, and specificity improved by 3.00%. The results show that the structured reasoning of knowledge graph and the dynamic interaction of graph visualisation synergistically improve the accuracy and transparency of fraud identification, providing a new paradigm for intelligent auditing.
Online fraud will cost businesses more than $200bn between 2020 and 2024, according to Juniper Research. 1 This stunning amount is driven by the increased sophistication of fraud attempts and the rising number of attack vectors. And, while banks are fighting back harder than ever, fraudsters have adjusted their techniques to remain below the radar. Online fraud will cost businesses more than $200bn between 2020 and 2024. This stunning amount is driven by the increased sophistication of fraud attempts and the rising number of attack vectors. Banks, however, have a new weapon in the war against fraud – graph analytics. These techniques can be used for fighting financial fraud by analysing the links between people, phones and bank accounts to reveal indicators of fraudulent behaviour, helping banks pinpoint suspicious activity in a sea of data, as Richard Henderson of TigerGraph explains.
Individual investors are disadvantaged in financial markets, overwhelmed by abundant information and lacking professional analysis.Equity research reports are crucial resources, offering valuable insights.By leveraging these reports, large language models (LLMs) can enhance investors' decision-making and strengthen financial analysis.However, two key challenges limit their effectiveness: (1) the rapid evolution of market events outpaces the slow update cycles of existing knowledge bases, and (2) the long-form, unstructured nature of financial reports hinders timely, context-aware integration by LLMs.To address these challenges, we tackle both data and methodology aspects.We introduce the Event-Enhanced Automated Construction of Financial Knowledge Graph (FinKario), a dataset with over 305,360 entities, 210,328 relational triples, and 19 relation types.FinKario integrates real-time company fundamentals and events through promptdriven extraction guided by institutional templates, providing structured, accessible financial insights for LLMs.We further propose a Two-stage, Graph-based retrieval strategy (FinKario-RAG) to optimize retrieval over evolving, large-scale financial knowledge.Experiments show that FinKario with FinKario-RAG achieves superior trend prediction accuracy, outperforming financial LLMs by 18.81% and institutional strategies by 17.85% on average in backtesting.
… a financial statement structure industry classification (FSSIC) method based on firms' financial statement similarity measured by the graph … In this section, we provide validation results …
This paper aims to develop a design for an accounting information system that will enhance the representational faithfulness of financial reporting information. One of the functions of financial reporting is to aggregate and report the entity's private data. This paper shows that recognizing that some of the firm's private data is already shared with others allows the methods of multiparty security to be applied to the reporting and audit processes. We contend that using both public key cryptography and network analysis allows the identity of an entity to be modelled as a place on a network. We also develop accounting recordkeeping techniques to balance public access with privacy using a blockchain. Taken together, these three design ideas can enhance the representational faithfulness of financial reporting systems because they use shared data from independent entities, a transparent system, and open-access immutable storage. Faithful representation is enhanced because information from this system can be used by auditors to support their audit opinion or by stakeholders who need credible information about the entity.
In the past few years, the main research efforts regarding General Data Protection Regulation (GDPR)-compliant data sharing have been focused primarily on informed consent (one of the six GDPR lawful bases for data processing). In cases such as Business-to-Business (B2B) and Business-to-Consumer (B2C) data sharing, when consent might not be enough, many small and medium enterprises (SMEs) still depend on contracts—a GDPR basis that is often overlooked due to its complexity. The contract’s lifecycle comprises many stages (e.g., drafting, negotiation, and signing) that must be executed in compliance with GDPR. Despite the active research efforts on digital contracts, contract-based GDPR compliance and challenges such as contract interoperability have not been sufficiently elaborated on yet. Since knowledge graphs and ontologies provide interoperability and support knowledge discovery, we propose and develop a knowledge graph-based tool for GDPR contract compliance verification (CCV). It binds GDPR’s legal basis to data sharing contracts. In addition, we conducted a performance evaluation in terms of execution time and test cases to validate CCV’s correctness in determining the overhead and applicability of the proposed tool in smart city and insurance application scenarios. The evaluation results and the correctness of the CCV tool demonstrate the tool’s practicability for deployment in the real world with minimum overhead.
This article presents a semantic web-based solution for extracting the relevant information automatically from the annual financial reports of the banks/financial institutions and presenting this information in a queryable form through a knowledge graph. The information in these reports is significantly desired by various stakeholders for making key investment decisions. However, this information is available in an unstructured format making it much more complex and challenging to understand and query manually or even through digital systems. Another challenge that makes the understanding of information more complex is the variation of terminologies among financial reports of different banks or financial institutions. The solution presented in this article signifies an ontological approach to solving the standardization problems of the terminologies in this domain. It further addresses the issue of semantic differences to extract relevant data sharing common semantics. Such semantics are then incorporated by implementing their representation as a Knowledge Graph to make the information understandable and queryable. Our results highlight the usage of Knowledge Graph in search engines, recommender systems and question-answering (Q-A) systems. This financial knowledge graph can also be used to serve the task of financial storytelling. The proposed solution is implemented and tested on the datasets of various banks and the results are presented through answers to competency questions evaluated on precision and recall measures.
With the technological development of entity extraction, relationship extraction, knowledge reasoning, and entity linking, the research on knowledge graph has been carried out in full swing in recent years. To better promote the development of knowledge graph, especially in the Chinese language and in the financial industry, we built a high-quality data set, named financial research report knowledge graph (FR2KG), and organized the automated construction of financial knowledge graph evaluation at the 2020 China Knowledge Graph and Semantic Computing Conference (CCKS2020). FR2KG consists of 17,799 entities, 26,798 relationship triples, and 1,328 attribute triples covering 10 entity types, 19 relationship types, and 6 attributes. Participants are required to develop a constructor that will automatically construct a financial knowledge graph based on the FR2KG. In addition, we summarized the technologies for automatically constructing knowledge graphs, and introduced the methods used by the winners and the results of this evaluation.
… automation improves financial reporting quality. Specifically, firms’ use of automation in the financial reporting … fuzzy logic), the fact that data inputs are not always well structured (which …
… This work presents a theoretical basis for systematic consistency maintenance by 1) defining constraint classes to relate objects across disciplines and 2) devising a mechanism to …
BackgroundA knowledge-based network, which is constructed by extracting as many relationships identified by experimental studies as possible and then superimposing them, is one of the promising approaches to investigate the associations between biological molecules. However, the molecular relationships change dynamically, depending on the conditions in a living cell, which suggests implicitly that all of the relationships in the knowledge-based network do not always exist. Here, we propose a novel method to estimate the consistency of a given network with the measured data: i) the network is quantified into a log-likelihood from the measured data, based on the Gaussian network, and ii) the probability of the likelihood corresponding to the measured data, named the graph consistency probability (GCP), is estimated based on the generalized extreme value distribution.ResultsThe plausibility and the performance of the present procedure are illustrated by various graphs with simulated data, and with two types of actual gene regulatory networks in Escherichia coli: the SOS DNA repair system with the corresponding data measured by fluorescence, and a set of 29 networks with data measured under anaerobic conditions by microarray. In the simulation study, the procedure for estimating GCP is illustrated by a simple network, and the robustness of the method is scrutinized in terms of various aspects: dimensions of sampling data, parameters in the simulation study, magnitudes of data noise, and variations of network structures.In the actual networks, the former example revealed that our method operates well for an actual network with a size similar to those of the simulated networks, and the latter example illustrated that our method can select the activated network candidates consistent with the actual data measured under specific conditions, among the many network candidates.ConclusionThe present method shows the possibility of bridging between the static network from the literature and the corresponding measurements, and thus will shed light on the network structure variations in terms of the changes in molecular interaction mechanisms that occur in response to the environment in a living cell.
A graph-based algorithm for consistency maintenance in incremental and interactive integration tools
… In the case of reactive integration, the integration tool performs incremental transformations and consistency checks on each user command. In contrast, integration on demand means …
… We then compare it with other graph-based data models and with setbased data … graph-based data modelling, we show how to bridge the gap between graph-based and set-based data …
Graph Learning has emerged as a promising technique for multi-view clustering, and has recently attracted lots of attention due to its capability of adaptively learning a unified and probably better graph from multiple views. However, the existing multi-view graph learning methods mostly focus on the multi-view consistency, but neglect the potential multi-view inconsistency (which may be incurred by noise, corruptions, or view-specific characteristics). To address this, this paper presents a new graph learning-based multi-view clustering approach, which for the first time, to our knowledge, simultaneously and explicitly formulates the multi-view consistency and the multi-view inconsistency in a unified optimization model. To solve this model, a new alternating optimization scheme is designed, where the consistent and inconsistent parts of each single-view graph as well as the unified graph that fuses the consistent parts of all views can be iteratively learned. It is noteworthy that our multi-view graph learning model is applicable to both similarity graphs and dissimilarity graphs, leading to two graph fusion-based variants, namely, distance (dissimilarity) graph fusion and similarity graph fusion. Experiments on various multi-view datasets demonstrate the superiority of our approach. The MATLAB source code is available at https://github.com/youweiliang/ConsistentGraphLearning.
… paper constructs an audit information knowledge graph based on the audit information of … corporations in the knowledge graph. This method using knowledge graph reasoning not only …
In recent years, knowledge graph technology has been widely applied in various fields such as intelligent auditing, urban transportation planning, legal research, and financial analysis. In traditional auditing methods, there are inefficiencies in data integration and analysis, making it difficult to achieve deep correlation analysis and risk identification among data. Additionally, decision support systems in the auditing process may face issues of insufficient information interpretability and limited predictive capability, thus affecting the quality of auditing and the scientificity of decision-making. However, knowledge graphs, by constructing rich networks of entity relationships, provide deep knowledge support for areas such as intelligent search, recommendation systems, and semantic understanding, significantly improving the accuracy and efficiency of information processing. This presents new opportunities to address the challenges of traditional auditing techniques. In this paper, we investigate the integration of intelligent auditing and knowledge graphs, focusing on the application of knowledge graph technology in auditing work for power engineering projects. We particularly emphasize mainstream key technologies of knowledge graphs, such as data extraction, knowledge fusion, and knowledge graph reasoning. We also introduce the application of knowledge graph technology in intelligent auditing, such as improving auditing efficiency and identifying auditing risks. Furthermore, considering the environment of cloud-edge collaboration to reduce computing latency, knowledge graphs can also play an important role in intelligent auditing. By integrating knowledge graph technology with cloud-edge collaboration, distributed computing and data processing can be achieved, reducing computing latency and improving the response speed and efficiency of intelligent auditing systems. Finally, we summarize the current research status, outlining the challenges faced by knowledge graph technology in the field of intelligent auditing, such as scalability and security. At the same time, we elaborate on the future development trends and opportunities of knowledge graphs in intelligent auditing.
In this report a computational study of ConceptNet 4 is performed using tools from the field of network analysis. Part I describes the process of extracting the data from the SQL database that is available online, as well as how the closure of the input among the assertions in the English language is computed. This part also performs a validation of the input as well as checks for the consistency of the entire database. Part II investigates the structural properties of ConceptNet 4. Different graphs are induced from the knowledge base by fixing different parameters. The degrees and the degree distributions are examined, the number and sizes of connected components, the transitivity and clustering coefficient, the cores, information related to shortest paths in the graphs, and cliques. Part III investigates non-overlapping, as well as overlapping communities that are found in ConceptNet 4. Finally, Part IV describes an investigation on rules.
高效的财务分析对于财务报表使用者获得洞察并支持其金融决策至关重要. 针对财务报表中多维复杂数据难以快速提取并理解关键维度的问题, 提出一种融合杜邦分析与大语言模型的智能可视分析方法. 该方法以解析报表为起点, 提取关键财务指标; 通过杜邦分析对指标进行拆解, 构建财务分析路径; 结合大语言模型, 生成趋势洞察与可视化视图, 辅助识别数据背后的关系与变化. 在此基础上设计并实现了一个原型系统FinDecipher进行方法验证. 通过该方法对上市公司财务数据进行案例分析, 基于交互式分析与大模型辅助洞察支持定位财务异常、识别关键驱动因素. 结果表明, FinDecipher能够提升用户对财务数据的理解和决策的准确性, 展现了所提出方法在实际应用中的可行性, 及其在金融科技中的应用潜力.
外汇市场系统性风险不仅影响外汇市场健康发展,而且影响整个金融系统稳定。鉴于外汇做市商在外汇市场中扮演着关键角色,本文分析中国外汇市场系统性风险形成与聚集的机理,厘清负面冲击下的外汇市场系统性风险形成渠道,结合知识图谱网络特征与现有三类系统性风险测度指标ΔCoVaR、MES、SRISK,构建中国外汇市场系统性风险测度指标。通过国内外重大事件验证测度指标的有效性,识别外汇市场的系统重要性外汇做市商银行,探究不同冲击渠道下的影响特征。研究结果表明:1)MES能够体现短事件冲击下外汇市场系统性风险水平的变化,即其短期测度效果最好,而SRISK能够体现长事件冲击下外汇市场系统性风险水平的整体变化,即其长期测度效果最好;2)考虑知识图谱网络特征下的指标测度效果要优于不使用网络特征,并且使用CFETS指数作为中国外汇市场指数时的测度效果要优于使用SDR指数和BIS指数;3)长期来看,四大国有银行和交通银行等系统重要性更高,短期来看,地方性银行系统重要性更高;4)在长事件冲击下,对系统重要性高的银行需要加强监管,考虑监管成本时可以适当下调系统重要性较低银行的监管水平。
该研究集合围绕图数据库在财务与税务领域中的应用,探讨了从知识图谱自动化构建、智能审计风险识别、跨系统数据一致性验证到合规性自动监测等核心技术路径,展示了图技术在复杂金融业务数据处理与决策支持中的深度价值。