流式 多模态 DLM
流式生成与流匹配模型框架
这些文献主要集中于基于流匹配(Flow Matching/Rectified Flow)框架的生成模型研究,重点在于处理多模态之间的条件生成与转换,强调生成过程的高效性与一致性。
- Semantic-guided flow matching for multimodal audio generation(S Shan, H Wang, J Chen, J Li, 2026, … Conference on AI …)
- VAFlow: Video-to-Audio Generation with Cross-Modality Flow Matching(Xihua Wang, Xin Cheng, Yuyue Wang, Ruihua Song, Yunfeng Wang, 2025, 2025 IEEE/CVF International Conference on Computer Vision (ICCV))
- CTFlow: Video-Inspired Latent Flow Matching for 3D CT Synthesis(Jiayi Wang, Hadrien Reynaud, F. Erick, Bernhard Kainz, 2025, 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW))
- OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows(Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Zichun Liao, Yusuke Kato, Kazuki Kozuka, Aditya Grover, 2024, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR))
- Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching(Yongqi Wang, Wenxiang Guo, Rongjie Huang, Jia-Bin Huang, Zehan Wang, Fuming You, Ruiqi Li, Zhou Zhao, 2024, Advances in Neural Information Processing Systems 37)
- Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching(A. Vosoughi, Yongyi Zang, Qihui Yang, Nathan Paek, Randal Leistikow, Chenliang Xu, 2026, ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP))
- Variational Flow Matching for Graph Generation(Grigory Bartosh, Floor Eijkelboom, Christian A. Naesseth, Jan-Willem van de Meent, Max Welling, 2024, Advances in Neural Information Processing Systems 37)
- GoMatch: Goal-Guided Truncated Flow Matching for Multimodal Trajectory Prediction(Yamei Xu, Yucheng Shi, Zhenghan Gao, Hong Zhang, Chengming Liu, Lei Shi, 2026, IEEE Internet of Things Journal)
- AvaeFlow: Unified Latent Space Learning for NoiseFree Multimodal Synthesis(Zhenhai Li, Jinnan Zhang, Zheyu Liu, Zhimeng Jiao, Xiaotian Yang, Yutao Shi, Xia Zhang, 2025, 2025 17th International Conference on Advanced Infocomm Technology (ICAIT))
- MusFlow: Multimodal Music Generation via Conditional Flow Matching(Jiahao Song, Yu-Zhao Wang, 2025, Proceedings of the 33rd ACM International Conference on Multimedia)
- GoalFlow: Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous Driving(Zebin Xing, Xingyu Zhang, Yang Hu, Bo Jiang, Tong He, Qian Zhang, Xiaoxiao Long, Wei Yin, 2025, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR))
交互式流式多模态大模型与推理
该组文献关注大规模多模态模型(LMMs)在连续流式数据处理、实时对话及交互式理解场景下的应用,重点讨论如何提升推理效率、长上下文理解及 proactive 推理能力。
- ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large Language Models(Danae S'anchez Villegas, Ingo Ziegler, Desmond Elliott, 2025, 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV))
- MMLSCU: A Dataset for Multi-modal Multi-domain Live Streaming Comment Understanding(Zixiang Meng, Qiang Gao, Di Guo, Yunlong Li, Bobo Li, Hao Fei, Shengqiong Wu, Fei Li, Chong Teng, Donghong Ji, 2024, Proceedings of the ACM Web Conference 2024)
- OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts(Yuxuan Wang, Yueqian Wang, Borun Chen, Tong Wu, Dongyan Zhao, Zilong Zheng, 2025, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR))
- AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling(Jun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou, Dong Zhang, Zhigeng Liu, Xin Zhang, Ruibin Yuan, Ge Zhang, Linyang Li, Hang Yan, Jie Fu, Tao Gui, Tianxiang Sun, Yu–Gang Jiang, Xipeng Qiu, 2024, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers))
- VideoLLM-online: Online Video Large Language Model for Streaming Video(Joya Chen, Zhaoyang Lv, Shiwei Wu, Kevin Qinghong Lin, Chenan Song, Difei Gao, Jia-Wei Liu, Ziteng Gao, Dongxing Mao, M. Shou, 2024, 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR))
- Streaming Long Video Understanding with Large Language Models(Shuangrui Ding, Xiaoyi Dong, Dahua Lin, Rui Qian, Jiaqi Wang, Yuhang Zang, Pan Zhang, 2024, Advances in Neural Information Processing Systems 37)
- EventGPT: Event Stream Understanding with Multimodal Large Language Models(Shaoyu Liu, Jianing Li, Guanghui Zhao, Yunjian Zhang, Xin Meng, F. Yu, Xiangyang Ji, Ming Li, 2024, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR))
- StreamSense: Streaming Social Task Detection with Selective Vision–Language Model Routing(Han Wang, Deyi Ji, Lanyun Zhu, Jiebo Luo, Roy Ka-Wei Lee, 2026, Proceedings of the ACM Web …)
实时多模态感知与行为反馈系统
这些文献聚焦于工业、教育、医疗等垂直领域,构建实时多模态感知系统,旨在通过融合多传感器(如EEG、姿态、生理信号等)进行实时状态评估、行为识别与协同交互。
- Multimodal Mean Adaptive Backgrounding for Embedded Real-Time Video Surveillance(S. Apewokin, B. Valentine, L. Wills, S. Wills, A. Gentile, 2007, 2007 IEEE Conference on Computer Vision and Pattern Recognition)
- The Coordinating Role of Language in Real-Time Multimodal Learning of Cooperative Tasks(Maxime Petit, S. Lallée, Jean-David Boucher, G. Pointeau, Pierrick Cheminade, D. Ognibene, E. Chinellato, U. Pattacini, I. Gori, Uriel Martinez-Hernandez, Hector Barron-Gonzalez, Martin Inderbitzin, Andre L. Luvizotto, V. Vouloutsi, Y. Demiris, G. Metta, Peter Ford Dominey, 2013, IEEE Transactions on Autonomous Mental Development)
- Multimodal sequence dynamics and convergence optimization in dual-stream LSTM networks for complex physiological state estimation(Xiaoxiao Cao, 2026, Frontiers in Neurorobotics)
- Multimodal Streaming Speech Synthesis and Zero-Sample Clone Framework for Smart Education(Renfei He, 2026, Highlights in Science, Engineering and Technology)
- Real-Time early Stress and Anxiety Detection Using Deep Learning with Attention-Based Sequence Modeling(K. Lokesh, Balaram Nadiya, K. Sathvika, V. Bhavitha, C. C. Kumar, R. Bharathi, 2026, 2026 4th International Conference on Knowledge Engineering and Communication Systems (ICKECS))
- Real-Time Driver Cognitive Workload Recognition: Attention-Enabled Learning With Multimodal Information Fusion(Haohan Yang, Jingda Wu, Zhongxu Hu, Chen Lv, 2024, IEEE Transactions on Industrial Electronics)
- Smart sensor integration: A framework for multimodal emotion recognition in real-time(J. Wagner, E. André, Frank Jung, 2009, 2009 3rd International Conference on Affective Computing and Intelligent Interaction and Workshops)
- Receipt Recognition Technology Driven by Multimodal Alignment and Lightweight Sequence Modeling(Jin-ming Yu, Huijun Ma, Jianlei Kong, 2025, Electronics)
- Real-Time Multimodal Feedback with the CPR Tutor(Daniele Di Mitri, Ján Schneider, Kevin Trebing, Saša Sopka, Marcus Specht, Hendrik Drachsler, 2020, Lecture Notes in Computer Science)
- Multimodal AI for Real‐Time Food Safety and Quality: From Sensors to Foundation Models, Edge Deployment, and Regulation(Zhaojie Chen, Guangyu Zhang, Fan Zhang, 2026, Food Science & Nutrition)
- Real-Time Multi-modal Human-Robot Collaboration Using Gestures and Speech(Haodong Chen, M. C. Leu, Zhaozheng Yin, 2022, Journal of Manufacturing Science and Engineering)
- Sequence-based multimodal behavior modeling for social agents(S. Dermouche, C. Pelachaud, 2016, Proceedings of the 18th ACM International Conference on Multimodal Interaction)
- Real-Time Phishing Detection for Brand Protection Using Temporal Convolutional Network-Driven URL Sequence Modeling(Marie-Laure E. Alorvor, Sajjad Dadkhah, 2025, Electronics)
调研文献涵盖了流式多模态深度学习模型的三大核心方向:一是基于流匹配的高效生成技术,解决了模态转换与对齐;二是基于大语言模型的实时交互式推理,推动了端到端流式多模态理解;三是面向特定垂直领域的嵌入式实时感知与反馈系统,通过多模态数据融合优化了复杂人机交互与实时决策流程。
总计32篇相关文献
Recent Large Language Models (LLMs) have been en-hanced with vision capabilities, enabling them to compre-hend images, videos, and interleaved vision-language con-tent. However, the learning methods of these large multi-modal models (LMMs) typically treat videos as predeter-mined clips, rendering them less effective and efficient at handling streaming video inputs. In this paper, we pro-pose a novel Learning-In- Video-Stream (LIVE) framework, which enables temporally aligned, long-context, and real-time dialogue within a continuous video stream. Our LIVE framework comprises comprehensive approaches to achieve video streaming dialogue, encompassing: (1) a training ob-jective designed to perform language modeling for contin-uous streaming inputs, (2) a data generation scheme that converts offline temporal annotations into a streaming di-alogue format, and (3) an optimized inference pipeline to speed up interactive chat in real-world video streams. With our LIVE framework, we develop a simplified model called VideoLLM-online and demonstrate its significant advan-tages in processing streaming videos. For instance, our VideoLLM-online-7B model can operate at over 10 FPS on an A100 GPU for a 5-minute video clip from Ego4D narration. Moreover, VideoLLM-online also showcases state-of-the-art performance on public offline video bench-marks, such as recognition, captioning, and forecasting. The code, model, data, and demo have been made available at showlab.github. iolvideollm-online.
The rapid advancement of multi-modal language models (MLLMs) like GPT-4o has propelled the development of Omni language models, designed to process and proactively respond to continuous streams of multi-modal data. Despite their potential, evaluating their real-world interactive capabilities in streaming video contexts remains a formidable challenge. In this work, we introduce OmniMMI, a comprehensive multi-modal interaction benchmark tailored for OmniLLMs in streaming video contexts. OmniMMI encompasses over 1,121 videos and 2,290 questions, addressing two critical yet underexplored challenges in existing video benchmarks: streaming video understanding and proactive reasoning, across six distinct subtasks. Moreover, we propose a novel framework, Multi-modal Multiplexing Modeling (M4), designed to enable an inference-efficient streaming model that can see, listen while generating. Extensive experimental results reveal that the existing MLLMs fall short in interactive streaming understanding, particularly struggling with proactive tasks and multi-turn queries. Our proposed M4, though lightweight, demonstrates a significant improvement in handling proactive tasks and real-time interactions.
Live streaming platforms require real-time monitoring and reaction to social signals, utilizing partial and asynchronous evidence from video, text, and audio. We propose StreamSense, a streaming detector that couples a lightweight streaming encoder with selective routing to a Vision–Language Model (VLM) expert. StreamSense handles most timestamps with the lightweight streaming encoder, escalates hard/ambiguous cases to the VLM, and defers decisions when context is insufficient. The encoder is trained using (i) a cross-modal contrastive term to align visual/audio cues with textual signals, and (ii) an IoU-weighted loss that down-weights poorly overlapping target segments, mitigating label interference across segment boundaries. We evaluate StreamSense on multiple social streaming detection tasks (e.g., sentiment classification and hate content moderation), and the results show that StreamSense achieves higher accuracy than VLM-only streaming while only occasionally invoking the VLM, thereby reducing average latency and compute. Our results indicate that selective escalation and deferral are effective primitives for understanding streaming social tasks. Code is publicly available on GitHub.
… Timechat: A time-sensitive multimodal large language model for long video understanding. arXiv preprint arXiv:2312.02051, 2023. [63] Dustin Schwenk, Apoorv Khandelwal, …
Event cameras capture visual information as asynchronous pixel change streams, excelling in challenging lighting and high-dynamic scenarios. Existing multimodal large language models (MLLMs) concentrate on natural RGB images, failing in scenarios where event data fits better. In this paper, we introduce EventGPT, the first MLLM for event stream understanding, pioneering the integration of large language models (LLMs) with event-based vision. To bridge the huge domain gap, we propose a three-stage optimization paradigm to progressively equip a pre-trained LLM with event understanding. Our EventGPT consists of an event encoder, a spatio-temporal aggregator, a linear projector, an event-language adapter, and an LLM. Firstly, GPT-generated RGB image-text pairs warm up the linear projector, following LLaVA, as the gap between natural images and language is smaller. Secondly, we construct N-ImageNet-Chat, a large synthetic dataset of event data and corresponding texts to enable the use of the spatio-temporal aggregator and to train the event-language adapter, thereby aligning event features more closely with the language space. Finally, we gather an instruction dataset, EventChat, which contains extensive real-world data to fine-tune the entire model, further enhancing its generalization ability. We construct a comprehensive benchmark, and experiments show that EventGPT surpasses previous state-of-the-art MLLMs in generation quality, descriptive accuracy, and reasoning capability. Code: EventGPT
We propose GoalFlow, an end-to-end autonomous driving method for generating high-quality multimodal trajectories. In autonomous driving scenarios, there is rarely a single suitable trajectory. Recent methods have increasingly focused on modeling multimodal trajectory distributions. However, they suffer from trajectory selection complexity and reduced trajectory quality due to high trajectory divergence and inconsistencies between guidance and scene information. To address these issues, we introduce GoalFlow, a novel method that effectively constrains the generative process to produce high-quality, multimodal trajectories. To resolve the trajectory divergence problem inherent in diffusion-based methods, GoalFlow constrains the generated trajectories by introducing a goal point. GoalFlow establishes a novel scoring mechanism that selects the most appropriate goal point from the candidate points based on scene information. Furthermore, GoalFlow employs an efficient generative method, Flow Matching, to generate multimodal trajectories, and incorporates a refined scoring mechanism to select the optimal trajectory from the candidates. Our experimental results, validated on the Navsim[7], demonstrate that GoalFlow achieves state-of-the-art performance, delivering robust multimodal trajectories for autonomous driving. GoalFlow achieved PDMS of 90.3, significantly surpassing other methods. Compared with other diffusion-policy-based methods, our approach requires only a single denoising step to obtain excellent performance. The code is available at https://github.com/YvanYin/GoalFlow.
Music generation aims to create music segments that align with human aesthetics based on diverse conditions. Despite advancements in generating music from specific textual descriptions (e.g., style, genre, instruments), the practical application is still hindered by ordinary users' limited expertise to write accurate prompts. To bridge this application gap, this paper introduces MusFlow, a novel multimodal music generation model using Conditional Flow Matching (CFM). We employ multiple Multi-Layer Perceptrons to align multimodal conditions into the audio's CLAP embedding space. CFM is trained to reconstruct the compressed Mel-spectrogram in the VAE latent space guided by aligned feature embedding. MusFlow can generate music from images, story texts, and music captions. To collect data for model training, inspired by multi-agent collaboration, we construct an intelligent annotation workflow centered around a fine-tuned Qwen2-VL model. Using this workflow, we build a new multimodal music dataset, MMusSet, with each sample containing a quadruple of image, story text, music caption, and music piece. We conduct four sets of experiments: image-to-music, story-to-music, caption-to-music, and multimodal music generation. Experimental results demonstrate that MusFlow can generate high-quality music pieces from multimoal conditions. We hope this work can advance the application of music generation in multimedia field, making music creation more accessible. Our generated samples are available at https://anonymous22356.github.io/musflow.github.io/
… iteratively denoising Gaussian-sampled latents, govenmented by various designed schedulers achieving state-ofthe-art performance in multimodal generation tasks. Unlike diffusions, …
… Rectified Flow Matching We train a DiT with conditional rectified flow matching [16] to learn the transformation from noise to RIR latents. The model learns a velocity field vθ(xt, t, c) that …
… flow matching, which jointly optimizes these objectives. Specifically, we introduce a truncated flow matching … , enabling efficient and stable generation. In this paper, we employ a …
… We adopt the Rectified Flow Matching framework to learn a deterministic probability path from a standard Gaussian distribution to the data distribution. Given the data x1 (audio latent) …
Video-to-audio (V2A) generation aims to synthesize content-matching audio from silent video, and it remains challenging to build V2A models with high generation quality, efficiency, and visual-audio temporal synchrony. We propose Frieren, a V2A model based on rectified flow matching. Frieren regresses the conditional transport vector field from noise to spectrogram latent with straight paths and conducts sampling by solving ODE, outperforming autoregressive and score-based models in terms of audio quality. By employing a non-autoregressive vector field estimator based on a feed-forward transformer and channel-level cross-modal feature fusion with strong temporal alignment, our model generates audio that is highly synchronized with the input video. Furthermore, through reflow and one-step distillation with guided vector field, our model can generate decent audio in a few, or even only one sampling step. Experiments indicate that Frieren achieves state-of-the-art performance in both generation quality and temporal alignment on VGGSound, with alignment accuracy reaching 97.22%, and 6.2% improvement in inception score over the strong diffusion-based baseline. Audio samples are available at http://frieren-v2a.github.io.
Jun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou, Dong Zhang, Zhigeng Liu, Xin Zhang, Ruibin Yuan, Ge Zhang, Linyang Li, Hang Yan, Jie Fu, Tao Gui, Tianxiang Sun, Yu-Gang Jiang, Xipeng Qiu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
With the growing adoption of cross-modal generation tasks - such as text-to-image and image-to-text generation-in natural language processing and computer vision, achieving efficient and controllable modality alignment while preserving semantic consistency has become a critical challenge. Existing methods often rely on task-specific architectural designs or complex noise modeling, which limits their generalizability and scalability.To address this, we propose AvaeFlow, a unified framework for cross-modal generation. The design is centered around two core innovations: (1) the incorporation of an Autoencoding Variational Autoencoder (AVAE) to regularize the encoding of the source modality, ensuring alignment with the target modality in the latent space; and (2) the introduction of a binary conditioning indicator during training, which enables effective integration of Classifier-Free Guidance (CFG) into the flow-matching process to enhance conditional controllability. Unlike conventional approaches, AvaeFlow performs direct mapping between heterogeneous modalities without requiring auxiliary conditioning modules or noise-based modeling, thereby offering superior generality, scalability, and controllability. Experiments were conducted on a large-scale text-to-image dataset containing approximately 2 million image-text pairs. The results demonstrate that AvaeFlow consistently outperforms existing mainstream methods in terms of generation quality, latent space manipulation, and scalability.
We introduce OmniFlow, a novel generative model designed for any-to-any generation tasks such as text-to-image, text-to-audio, and audio-to-image synthesis. OmniFlow advances the rectified flow (RF) framework used in text-to-image models to handle the joint distribution of multiple modalities. It outperforms previous any-to-any models on a wide range of tasks, such as text-to-image and text-to-audio synthesis. Our work offers three key contributions: First, we extend RF to a multi-modal setting and introduce a novel guidance mechanism, enabling users to flexibly control the alignment between different modalities in the generated outputs. Second, we propose a novel architecture that extends the text-to-image MMDiT architecture of Stable Diffusion 3 and enables audio and text generation. The extended modules can be efficiently pretrained individually and merged with the vanilla text-to-image MMDiT for fine-tuning. Lastly, we conduct a comprehensive study of the design choices of rectified flow transformers for large-scale audio and text generation, providing valuable insights into optimizing performance across various modalities. Code is available at https://github.com/jacklishufan/OmniFlows.
Generative modelling of entire CT volumes conditioned on clinical reports has the potential to accelerate research through data augmentation, privacy-preserving synthesis and reducing regulator-constraints on patient data while preserving diagnostic signals. With the recent release of CT-RATE, a large-scale collection of 3D CT volumes paired with their respective clinical reports, training large text-conditioned CT volume generation models has become achievable. In this work, we introduce CTFlow, a 0.5B latent flow matching transformer model, conditioned on clinical reports. We leverage the A-VAE from FLUX to define our latent space, and rely on the CT-Clip text encoder to encode the clinical reports. To generate consistent whole CT volumes while keeping the memory constraints tractable, we rely on a custom autoregressive approach, where the model predicts the first sequence of slices of the volume from text-only, and then relies on the previously generated sequence of slices and the text, to predict the following sequence. We evaluate our results against state-of-the-art generative CT model, and demonstrate the superiority of our approach in terms of temporal coherence, image diversity and text-image alignment, with FID, FVD, IS scores and CLIP score.
With the deepening of the digital transformation of education, smart education has put forward higher requirements for natural and expressive real-time voice interaction technologies. However, traditional speech synthesis (TTS) systems face core challenges such as insufficient understanding of professional domain terms, monotonous emotional expression, and the lack of cross-modal collaboration in educational scenarios, making it difficult to meet the needs of immersive and interactive teaching. To overcome these limitations, this paper proposes a multimodal end-to-end streaming speech synthesis and intelligent processing framework for educational scenarios. Firstly, this framework builds a multimodal fusion network based on cross-modal attention mechanisms, which dynamically aligns text semantics, acoustic features, and speaker identity at multiple scales of phonemes, syllables, and sentences, significantly improving the naturalness and semantic consistency of the synthesized speech. Secondly, in terms of personalized speech modeling, the system introduces zero-sample speech cloning technology that integrates semantic understanding of large language models (LLMs) and progressive fine-tuning strategies, enabling high-fidelity replication of teacher-specific voice and cross-language synthesis with only a few seconds of audio samples. To meet the low latency requirements of real-time classroom interaction, the architecture integrates a streaming generation engine based on Chunk-Aware Causal Flow Matching, effectively supporting generation and transmission simultaneously, strictly controlling the system's end-to-end latency within 150 milliseconds. Experimental verification and system analysis show that this multi-task joint optimization framework can precisely handle speechization of complex subject content, adaptively adjust teaching emotional expression, and provide a solid multimodal speech technology foundation for building a highly inclusive and personalized intelligent education ecosystem.
With the increasing popularity of live streaming, the interactions from viewers during a live streaming can provide more specific and constructive feedback for both the streamer and platform. In such scenario, the primary and most direct feedback method from the audience is through comments. Thus, mining these live streaming comments to unearth the intentions behind them and, in turn, aiding streamers to enhance their live streaming quality is significant for the well development of live streaming ecosystem. To this end, we introduce the MMLSCU dataset, containing 50,129 intention-annotated comments across multiple modalities (text, images, vi-deos, audio) from eight streaming domains. Using multimodal pretrained large model and drawing inspiration from the Chain of Thoughts (CoT) concept, we implement an end-to-end model to sequentially perform the following tasks: viewer comment intent detection ➛ intent cause mining ➛ viewer comment explanation ➛ streamer policy suggestion. We employ distinct branches for video and audio to process their respective modalities. After obtaining the video and audio representations, we conduct a multimodal fusion with the comment. This integrated data is then fed into the large language model to perform inference across the four tasks following the CoT framework. Experimental results indicate that our model outperforms three multimodal classification baselines on comment intent detection and streamer policy suggestion, and one multimodal generation baselines on intent cause mining and viewer comment explanation. Compared to the models using only text, our multimodal setting yields superior outcomes. Moreover, incorporating CoT allows our model to enhance comment interpretation and more precise suggestions for the streamers. Our proposed dataset and model will bring new research attention on multimodal live streaming comment understanding.
… Generative flows on discrete state-spaces: Enabling multimodal flows with applications to protein co-design. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria …
With the rapid advancement of global digital transformation, enterprises and financial institutions face increasing challenges in managing and processing receipt-like financial documents. Traditional manual document processing methods can no longer meet the demands of modern office operations and business expansion. To address these issues, automated document recognition systems based on computer vision and deep learning technologies have emerged. This paper proposes a receipt recognition technology based on multimodal alignment and lightweight sequence modeling, integrating the CLIP (Contrastive Language-Image Pretraining) and Bidirectional Gated Recurrent Unit (BiGRU) framework. The framework aims to achieve synergistic optimization of image and text information through semantic correction. By leveraging dynamic threshold classification, geometric regression loss, and multimodal feature alignment, the framework significantly improves text detection and recognition accuracy in complex layouts and low-quality images. Experimental results show that the model achieves a detection F1 score of 93.1% and a Character Error Rate (CER) of 5.1% on the CORD dataset. Through a three-stage compression strategy of quantization, pruning, and distillation, the model size is reduced to 18 MB, achieving real-time inference speeds of 25 FPS on the Jetson AGX Orin edge device, with power consumption stabilized below 12 W. This framework provides an efficient, accurate, and edge-computing-friendly solution for automated receipt processing. Practical implications include its potential to enhance the efficiency of financial audits, improve tax compliance, and streamline the operational management of financial institutions, making it a valuable tool for real-world applications in receipt automation.
Driver workload inference is significant for the design of intelligent human–machine cooperative driving schemes since it allows the systems to alert drivers before potentially dangerous maneuvers and achieve a safer control transition. However, pattern variations among individual drivers and sensor artifacts pose great challenges to the existing cognitive workload recognition approaches. In this article, we develop an attention-enabled recognition network with a decision-level fusion architecture to further improve the workload estimation performance. Specifically, the cross-attention mechanism can enhance useful feature representations learned by hyper long-short-term-memory-based modules from time-series multimodal information, i.e., electroencephalogram signals, eye movements, and vehicle states. A novel dataset containing multiple driving scenarios is constructed to evaluate the model performance across different historical horizons and decision thresholds, and test results demonstrate the superior performance of the proposed model to other existing methods. Furthermore, robustness tests and driver-in-the-loop experiments are conducted to verify the effectiveness of the developed model in real-time workload levels inference. The code and supplementary materials are available at https://yanghh.io/Driver-Workload-Recognition.
Phishing, especially brand impersonation attacks, is a critical cybersecurity threat that harms user trust and organization security. This paper establishes a lightweight model for real-time detection that relies on URL-only sequences, addressing limitations for multimodal methods that leverage HTML, images, or metadata. This approach is based on a Temporal Convolutional Network with Attention (TCNWithAttention) that utilizes character-level URLs to capture both local and long-range dependencies, while providing interpretability with attention visualization and Shapley additive explanations (SHAP). The model was trained and tested on the balanced GramBeddings dataset (800,000 URLs) and validated on the PhiUSIIL dataset of real-world phishing URLs. The model achieved 97.54% accuracy on the GramBeddings dataset, and 81% recall on the PhiUSIIL dataset. The model demonstrated strong generalization, fast inference, and CPU-only deployability. It outperformed CNN, BiLSTM and BERT baselines. Explanations highlighted phishing indicators, such as deceptive subdomains, brand impersonation, and suspicious tokens. It also affirmed real patterns in the legitimate domains. To our knowledge, a Streamlit application to facilitate single and batch URL analysis and log feedback to maintain usability is the first phishing detection framework to integrate TCN, attention, and SHAP, bridging academic innovation with practical cybersecurity techniques.
Stress and anxiety are critical intellectual health issues that significantly affect cognitive average overall performance, emotional balance, and standard properly-being, particularly amongst college students and working specialists. Traditional assessment techniques on the aspect of self-opinions and clinical reviews are subjective, infrequent, and unable to capture rapid changes in emotional states. To overcome those limitations, this work proposes a actual-time pressure and anxiety detection framework that integrates multi-modal physiological signals—along with electroencephalography (EEG), heart charge variability (HRV), and galvanic skin reaction (GSR)—with superior deep reading fashions. The accrued indicators are preprocessed, segmented into temporal home windows, and converted into informative characteristic representations, which might be then analyzed the use of a hybrid CNN–BiLSTM structure with an attention mechanism for greater interpretability and series-level studying. The system demonstrates excessive accuracy in identifying every day, mild, and high-strain states, outperforming conventional system studying tactics and single-sensor strategies. The proposed framework offers a scalable, non-invasive, and real-time answer for early stress and anxiety detection, enabling timely intervention and improved intellectual health support in educational and occupational settings
As artificial intelligence and industrial automation are developing, human-robot collaboration (HRC) with advanced interaction capabilities has become an increasingly significant area of research. In this paper, we design and develop a real-time, multi-model HRC system using speech and gestures. A set of sixteen dynamic gestures is designed for communication from a human to an industrial robot. A data set of dynamic gestures is designed and constructed, and it will be shared with the community. A convolutional neural network (CNN) is developed to recognize the dynamic gestures in real time using the Motion History Image (MHI) and deep learning methods. An improved open-source speech recognizer is used for real-time speech recognition of the human worker. An integration strategy is proposed to integrate the gesture and speech recognition results, and a software interface is designed for system visualization. A multi-threading architecture is constructed for simultaneously operating multiple tasks, including gesture and speech data collection and recognition, data integration, robot control, and software interface operation. The various methods and algorithms are integrated to develop the HRC system, with a platform constructed to demonstrate the system performance. The experimental results validate the feasibility and effectiveness of the proposed algorithms and the HRC system.
… methods, including the multimodal Mixture of Gaussians (MoG) [… Multimodal Mean (MM), for real-time background modeling. Our technique achieves accuracy comparable to multimodal …
ABSTRACT Real‐time assurance of food safety and quality requires decisions at line speed, from farm to retail, using signals that span vision, spectroscopy, volatiles, biosensing, and process telemetry. This review investigates and summarizes evidence on multimodal artificial intelligence that fuses such heterogeneous data to detect hazards, verify authenticity, and predict freshness within seconds. We outline sensing coverage along the chain, typical response times, and reported limits of detection, then detail data engineering practices that make disparate streams analysis‐ready, including time synchronization, co‐registration to ground truth, and robust sampling for multisite and multiseason generalization. We appraise fusion strategies, from early and late schemes to attention‐based hybrids that learn joint embeddings across images, spectra, and gas sensor time series, and we summarize head‐to‐head studies where multimodality improves accuracy or reduces error against unimodal baselines. We discuss the maturation of foundation scale encoders and vision language systems for food tasks, together with efficient adaptation, knowledge infusion from HACCP, and bias control. Finally, we examine edge deployment and validation in industrial settings, including hardware constraints, latency budgets, repeatability and reproducibility, documentation for audits, and perspectives on regulatory alignment in EU and US contexts, extended to China's standards‐driven framework where the National Health Commission (NHC) and the State Administration for Market Regulation (SAMR) jointly issue and update National Food Safety Standards (GB) that govern key compliance requirements for labelling and contaminant limits. Evidence gaps persist, notably few multisite deployments over long durations, limited public benchmarks for hyperspectral and e‐nose fusion, and sparse cost–benefit analyses in the scholarly record. Addressing these gaps will enable trustworthy, auditable multimodal AI that complements existing controls and reduces waste while protecting consumers.
Reasoning over sequences of images remains a challenge for multimodal large language models (MLLMs). While recent models incorporate multi-image data during pretraining, they still struggle to recognize sequential structures, often treating images independently. This work introduces ImageChain, a framework that enhances MLLMs with sequential reasoning capabilities over image data by modeling visual sequences as a multi-turn conversation. In ImageChain, images are interleaved with corresponding textual descriptions to form a controlled dialogue that explicitly captures temporal dependencies and narrative progression. Our method optimizes for the task of next-scene description, where the model generates a context-aware description of an upcoming scene based on preceding visual and textual cues. We demonstrate that our approach improves performance on the next-scene description task – achieving an average improvement from 3.7% to 19% in SimRate, a metric that quantifies semantic similarity to human-annotated ground truths. Moreover, ImageChain achieves robust zero-shot out-of-domain performance in applications ranging from comics to robotics. Extensive experiments validate that instruction-tuning in a multimodal, multi-turn conversation design is key to bridging the gap between static image understanding and temporally-aware reasoning.1
Introduction The integration of virtual simulation with intelligent modeling is crucial for advancing the scientization and personalization of volleyball physical training. This study aims to overcome the convergence instability and feature misalignment in modeling multimodal kinematic and physiological sequences. Methods A dynamical framework based on a Dual-Stream Long Short-Term Memory network integrated with a temporal attention mechanism is proposed. The framework decouples heterogeneous feature learning and optimizes temporal weight distribution. Results Experimental validation on complex motion state estimation demonstrates that the proposed model reduces load modeling error to 3.8% and achieves a motion classification accuracy of 93.1%. The velocity trajectory fitting coefficient of determination is 0.91 with a peak deviation of 0.05 m/s. Discussion These results confirm the effectiveness of the attention-based DS-LSTM in optimizing multimodal sequence modeling for training state estimation and feedback.
… for the offline model training of the CPR mistakes as well as for the real-time multimodal data … The System Architecture proposed is the first complete implementation of the Multimodal …
One of the defining characteristics of human cognition is our outstanding capacity to cooperate. A central requirement for cooperation is the ability to establish a “shared plan”—which defines the interlaced actions of the two cooperating agents—in real time, and even to negotiate this shared plan during its execution. In the current research we identify the requirements for cooperation, extending our earlier work in this area. These requirements include the ability to negotiate a shared plan using spoken language, to learn new component actions within that plan, based on visual observation and kinesthetic demonstration, and finally to coordinate all of these functions in real time. We present a cognitive system that implements these requirements, and demonstrate the system's ability to allow a Nao humanoid robot to learn a nontrivial cooperative task in real-time. We further provide a concrete demonstration of how the real-time learning capability can be easily deployed on a different platform, in this case the iCub humanoid. The results are considered in the context of how the development of language in the human infant provides a powerful lever in the development of cooperative plans from lower-level sensorimotor capabilities.
… Investigating multimodal real-time patterns of joint attention in an HRI … Real-time adaptive behaviors in multimodal human-avatar interactions. In International Conference on Multimodal …
… The modeling and simulation of affect-aware … models of emotion, such as EMA [10] and ALMA [3], hardly any support is provided for multimodal emotion recognition in real-time which is …
调研文献涵盖了流式多模态深度学习模型的三大核心方向:一是基于流匹配的高效生成技术,解决了模态转换与对齐;二是基于大语言模型的实时交互式推理,推动了端到端流式多模态理解;三是面向特定垂直领域的嵌入式实时感知与反馈系统,通过多模态数据融合优化了复杂人机交互与实时决策流程。