UIST
AI 赋能的交互式设计与创作工具
这些文献主要探讨如何利用 AI 技术辅助用户进行 UI 设计、3D 建模或界面生成,重点在于提升创作效率及交互体验。
- ReFinE: Streamlining UI Mockup Iteration with Research Findings(Donghoon Shin, Bingcan Guo, Jaewook Lee, Lucy Lu Wang, Gary Hsieh, 2026, Proceedings of the 2026 Designing Interactive Systems Conference)
- Generative and Malleable User Interfaces with Generative and Evolving Task-Driven Data Model(Yining Cao, Peiling Jiang, Haijun Xia, 2025, Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems)
- RealityCrafter: User-guided Editable 3D Scene Generation from a Single Image in Mixed Reality(Seokyoung Kim, Dooyoung Kim, Taejun Son, Woontack Woo, 2025, Adjunct Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology)
- SVGraffiti: Remixing the Web with Vector Illustrations(Tongyu Zhou, Joshua Kong Yang, E. Chen, Jeffson Huang, 2026, Proceedings of the 2026 Designing Interactive Systems Conference)
增强现实与虚拟现实的交互范式
这组论文关注 AR/VR 环境下的交互设计,包括目标选择、沉浸感研究、生理反馈的适应性系统以及远程协作技术。
- Uncertain Pointer: Situated Feedforward Visualizations for Ambiguity-Aware AR Target Selection(Ching-Yi Tsai, Nicole Tacconi, Andrew D. Wilson, Parastoo Abtahi, 2026, Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems)
- CrossGaussian: Enhancing Remote Collaboration through 3D Gaussian Splatting and Real-time 360° Streaming(Jaehyun Byun, Byunghoon Kang, Yong-Rae Gwon, Hongsong Choi, Yunseo Do, Eunho Kim, Sangkeun Park, Seungjae Oh, 2025, Adjunct Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology)
- Towards Intelligent VR Training: A Physiological Adaptation Framework for Cognitive Load and Stress Detection(Mahsa Nasri, 2025, Proceedings of the 33rd ACM Conference on User Modeling, Adaptation and Personalization)
- Effects of User Interface Orientation on Sense of Immersion in Augmented Reality(Xiangdong Li, Kailin Yin, Yifei Shan, Xinyao Wang, Weidong Geng, 2024, International Journal of Human–Computer Interaction)
包容性设计与辅助技术
这组论文致力于通过 AI 驱动的交互代理或辅助系统,解决特定群体(如盲人和低视力者、认知障碍患者)在完成复杂任务或进行心理疗愈时的障碍。
- RemiGraph: An Interactive Memory Graph for AI-Guided Life Journaling and Reminiscence Support(Brian Ozawa Burns, Arthur Caetano, Alice Zhong, Tobias Höllerer, Misha Sra, 2025, Adjunct Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology)
- Vid2Coach: Transforming How-To Videos into Task Assistants(Mina Huh, Zihui Xue, Ujjaini Das, Kumar Ashutosh, Kristen Grauman, Amy Pavel, 2025, Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology)
- Morae: Proactively Pausing UI Agents for User Choices(Yi-Hao Peng, Dingzeyu Li, Jeffrey P. Bigham, Amy Pavel, 2025, Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology)
新型交互媒介、感知建模与伦理研究
这组文献探讨非传统交互媒介(如形状改变接口)、人机交互的感知模型以及提升系统透明度和问责制的交互框架。
- AI for Haptics and Haptics for AI: Challenges and Opportunities(Easa AliAbbasi, Dennis Wittchen, Yinan Li, Shihan Lu, Thomas Müller, Donald Degraen, Thomas Leimkühler, Sang Ho Yoon, Hasti Seifi, Oliver S. Schneider, Heather Culbertson, Jürgen Steimle, Paul Strohmeier. 2026, 2026, Proceedings of the Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems)
- "I don't want to break it": An Exploration of Perceived Fragility in Shape-Changing Interfaces(Eva Mackamul, Tom Maillard, No'e Marceaul, Y. Coulibaly, J. Pansiot, Laurence Boissieux, Dominique Vaufreydaz, A. Roudaut, Céline Coutrix, 2026, Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems)
- Why am I seeing this: Democratizing End User Auditing for Online Content Recommendations(Chaoran Chen, Leyang Li, Luke Cao, Yanfang Ye, Tianshi Li, Yaxing Yao, Toby Jia-Jun Li, 2024, Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology)
本次研究报告涵盖了 UIST 主题下交互技术的前沿进展,主要集中在 AI 驱动的生成式 UI 设计、增强/虚拟现实的空间交互、面向特殊人群的辅助辅助技术,以及人机交互感知与系统伦理的交叉探索,体现了智能化、沉浸化与包容性设计的核心发展趋势。
总计14篇相关文献
… Example screenshots of the 2025 UIST conference page showing how (top) mouse clicks of hyperlinks can be linked to sketches and how (bottom) mouse scrolls can be linked to a …
Target disambiguation is crucial in resolving input ambiguity in augmented reality (AR), especially for queries over distant objects or cluttered scenes on the go. Yet, visual feedforward techniques that support this process remain underexplored. We present Uncertain Pointer, a systematic exploration of feedforward visualizations that annotate multiple candidate targets before user confirmation, either by adding distinct visual identities (e.g., colors) to support disambiguation or by modulating visual intensity (e.g., opacity) to convey system uncertainty. First, we construct a pointer space of 25 pointers by analyzing existing placement strategies and visual signifiers used in target visualizations across 30 years of relevant literature. We then evaluate them through two online experiments (n = 60 and 40), measuring user preference, confidence, mental ease, target visibility, and identifiability across varying object distances and sparsities. Finally, from the results, we derive design recommendations in choosing different Uncertain Pointers based on AR context and disambiguation techniques.
We propose RealityCrafter, a mixed-reality 3D authoring tool that enables users to edit and interact with a reconstructed 3D scene from a single real-world image. Prior research has largely focused on 3D authoring tools for purely virtual spaces, insufficiently incorporating real-world context and thereby hindering user immersion during the creation process. To overcome these limitations, our approach takes a single real-world image as input, generates segmented object-level 3D meshes in a zero-shot manner, and reconstructs a 3D scene where objects can be removed or modified without occlusion through instance mask-based inpainting. We leverage LLMs to interpret user voice commands and update the style, position, scale, and orientation of 3D objects in real time, providing an interactive 3D authoring interface in mixed-reality environments. By using a single image as a baseline, this approach enables effortless generation of realistic 3D scenes and intuitive editing based on user intent, delivering a novel creative experience that seamlessly blends the real and the virtual objects.
AI has transformed methods and knowledge across many domains. However, the intersection of AI and haptics remains underexplored. While modern AI techniques – fueled by machine learning and using powerful techniques such as generative modeling and reinforcement learning – offer powerful opportunities for advancing haptic design, insights from haptics research, such as perception modeling and adaptive interaction - grounded in human touch, embodiment, and multisensory integration — can also play a critical role in shaping more human-centered AI systems. This workshop will bring together an interdisciplinary community of researchers from HCI, haptics, AI, robotics, and design to (1) identify pressing questions in haptics that could benefit from AI approaches and (2) highlight ways in which haptic knowledge can support the development of embodied and context-aware AI. Through position papers and paper presentations, we will map key challenges, exchange methods, and explore new research directions that connect the two fields. By framing haptics and AI as mutually reinforcing, the workshop aims to build a shared research agenda and foster collaborations that advance both the science of touch and the design of intelligent interactive systems.
Remote users often face significant challenges in remote collaboration systems when joining virtual scenes reconstructed from a local user’s environment. They are disadvantaged by the information asymmetry inherent in a shared virtual environment compared to local users. We present CrossGaussian, a VR collaboration system designed to address these limitations by providing remote users with a comprehensive 3D interactive view of the shared environment, created using 3DGS. It blends 360° video streaming and a large-scale 3DGS reconstruction via our automated pipeline. We then define the design space for visualization and interaction techniques that combine wide-coverage 3D and responsive 360° scenes.
Shape-changing interfaces (SCIs) dynamically alter their form, an inherent characteristic that introduces fragility into their design. As a result, users’ perceptions of an interface’s fragility or its potential to move or break may influence their interaction, however the extent of this effect is unclear. To address this gap, we conducted a qualitative study (N = 18) using video stimuli showcasing 20 existing SCIs. Through thematic analysis, we identified key factors impacting perceived fragility and formalized these into a framework. We then conducted a second study (N = 36) for which we fabricated SCIs that varied across selected fragility-related dimensions. We recorded user interactions and compared how the selected dimensions shaped manipulation of the objects and how they were considered by users. Together, these studies provide a structured foundational understanding of perceived fragility in SCIs and offer insights to enhance perceived robustness and inform future SCI development.
Although HCI research papers offer valuable design insights, designers often struggle to apply them in design workflows due to difficulties in finding relevant literature, understanding technical jargon, the lack of contextualization, and limited actionability. To address these challenges, we present ReFinE, a Figma plugin that supports real-time design iteration by surfacing contextualized insights from research papers. ReFinE identifies and synthesizes design implications from HCI literature relevant to the mockup's design context, and tailors this research evidence to a specific design mockup by providing actionable visual guidance on how to update the mockup. To assess the system's effectiveness, we conducted a technical evaluation and a user study. Results show that ReFinE effectively synthesizes and contextualizes design implications, reducing cognitive load and improving designers'ability to integrate research evidence into UI mockups. This work contributes to bridging the gap between research and design practice by presenting a tool for embedding scholarly insights into the UI design process.
Regular journaling supports autobiographical memory, anchoring identity and aiding future dementia care. We introduce RemiGraph, a system that enhances journaling with AI-generated question prompts, pairing it with an interactive memory graph that maps life events and patterns. The system aims to deepen self-reflection and give caregivers a clear view of a dementia patient’s life-log to support meaningful reminiscence therapy. In a pilot study (n=8), AI prompts boosted engagement and journal entry word count, while a walkthrough study with five therapists indicated RemiGraph’s potential to aid patient familiarization and enrich therapy. These findings position RemiGraph as a bridge between everyday journaling and dementia support.
People use videos to learn new recipes, exercises, and crafts. Such videos remain difficult for blind and low vision (BLV) people to follow as they rely on visual comparison. Our observations of visual rehabilitation therapists (VRTs) guiding BLV people to follow how-to videos revealed that VRTs provide both proactive and responsive support including detailed descriptions, non-visual workarounds, and progress feedback. We propose Vid2Coach, a system that transforms how-to videos into wearable camera-based assistants that provide accessible instructions and mixed-initiative feedback. From the video, Vid2Coach generates accessible instructions by augmenting narrated instructions with demonstration details and completion criteria for each step. It then uses retrieval-augmented-generation to extract relevant non-visual workarounds from BLV-specific resources. Vid2Coach then monitors user progress with a camera embedded in commercial smart glasses to provide context-aware instructions, proactive feedback, and answers to user questions. BLV participants (N=8) using Vid2Coach completed cooking tasks with 58.5% fewer errors than when using their typical workflow and wanted to use Vid2Coach in their daily lives. Vid2Coach demonstrates an opportunity for AI visual assistance that strengthens rather than replaces non-visual expertise.
Personalized recommendation systems tailor content based on user attributes, which are either provided or inferred from private data. Research suggests that users often hypothesize about reasons behind contents they encounter (e.g., “I see this jewelry ad because I am a woman”), but they lack the means to confirm these hypotheses due to the opaqueness of these systems. This hinders informed decision-making about privacy and system use and contributes to the lack of algorithmic accountability. To address these challenges, we introduce a new interactive sandbox approach. This approach creates sets of synthetic user personas and corresponding personal data that embody realistic variations in personal attributes, allowing users to test their hypotheses by observing how a website’s algorithms respond to these personas. We tested the sandbox in the context of targeted advertisement. Our user study demonstrates its usability, usefulness, and effectiveness in empowering end-user auditing in a case study of targeting ads.
User interface (UI) agents promise to make inaccessible or complex UIs easier to access for blind and low-vision (BLV) users. However, current UI agents typically perform tasks end-to-end without involving users in critical choices or making them aware of important contextual information, thus reducing user agency. For example, in our field study, a BLV participant asked to buy the cheapest available sparkling water, and the agent automatically chose one from several equally priced options, without mentioning alternative products with different flavors or better ratings. To address this problem, we introduce Morae, a UI agent that automatically identifies decision points during task execution and pauses so that users can make choices. Morae uses large multimodal models to interpret user queries alongside UI code and screenshots, and prompt users for clarification when there is a choice to be made. In a study over real-world web tasks with BLV participants, Morae helped users complete more tasks and select options that better matched their preferences, as compared to baseline agents, including OpenAI Operator. More broadly, this work exemplifies a mixed-initiative approach in which users benefit from the automation of UI agents while being able to express their preferences.
Unlike static and rigid user interfaces, generative and malleable user interfaces offer the potential to respond to diverse users’ goals and tasks. However, current approaches primarily rely on generating code, making it difficult for end-users to iteratively tailor the generated interface to their evolving needs. We propose employing task-driven data models—representing the essential information entities, relationships, and data within information tasks—as the foundation for UI generation. We leverage AI to interpret users’ prompts and generate the data models that describe users’ intended tasks, and by mapping the data models with UI specifications, we can create generative user interfaces. End-users can easily modify and extend the interfaces via natural language and direct manipulation, with these interactions translated into changes in the underlying model. The technical evaluation of our approach and user evaluation of the developed system demonstrate the feasibility and effectiveness of the proposed generative and malleable UIs.
Abstract The sense of immersion provided by overlying computer-generated graphics over actual items is one of augmented reality’s distinguishing features. It is affected by multiple factors such as the user interface orientation in the virtual physical hybrid space. There have been many studies on realistic graphic rendering and multimodal interaction in augmented reality. Few studies, however, have systematically investigated the influence of the user interface orientations, more specifically the user-based orientation (UO) and object-based orientation (OO), on various aspects of immersion. We developed the UO and OO user interface prototypes and evaluated the sense of immersion in terms of engagement, engrossment, and overall immersion. The results showed that the user interface orientations in augmented reality had a significant influence on users’ perceived engagement and engrossment, whereas there was no significant difference in overall immersion. Particularly, compared to the OO user interfaces, the UO user interfaces exhibited higher perceived usability and slower focus of attention switching. Generalising implications for alternative augmented reality are provided.
Adaptive Virtual Reality (VR) systems have the potential to enhance training and learning experiences by dynamically responding to users’ cognitive states. This research investigates how eye tracking and heart rate variability (HRV) can be used to detect cognitive load and stress in VR environments, enabling real-time adaptation. The study follows a three-phase approach: (1) conducting a user study with the Stroop task to label cognitive load data and train machine learning models to detect high cognitive load, (2) fine-tuning these models with new users and integrating them into an adaptive VR system that dynamically adjusts training difficulty based on physiological signals, and (3) developing a privacy-aware approach to detect high cognitive load and compare this with the adaptive VR in Phase two. This research contributes to affective computing and adaptive VR using physiological sensing, with applications in education, training, and healthcare. Future work will explore scalability, real-time inference optimization, and ethical considerations in physiological adaptive VR.
本次研究报告涵盖了 UIST 主题下交互技术的前沿进展,主要集中在 AI 驱动的生成式 UI 设计、增强/虚拟现实的空间交互、面向特殊人群的辅助辅助技术,以及人机交互感知与系统伦理的交叉探索,体现了智能化、沉浸化与包容性设计的核心发展趋势。