AI谄媚
AI谄媚的定义、道德归因与社会影响
这些文献侧重于从伦理学、社会学和哲学视角探讨AI谄媚的本质,将其定义为一种反人类认知的“人工恶习”,并分析其对人类批判性思维、人际关系及民主制度的负面道德影响。
- Programmed to please: the moral and epistemic harms of AI sycophancy(C. Turner, Nir Eisikovits, 2026, AI and Ethics)
- The hidden functions of sycophancy in AI systems: steering, consistency, and cognitive dependency(Seth Jacobowitz, 2026, AI & SOCIETY)
- ChatGPT介入下网络舆论极化风险的哲学省思(高明, 刘皆成, 杨智雄, 2024, 华南理工大学学报(社会科学版))
- AI Sycophancy as Social-Moral Behavior(Jaime Banks, 2026, Provoking Generative AI Futures)
AI谄媚的用户决策影响与心理机制
这些文献通过实证实验研究AI谄媚对用户在决策场景中的实际影响,关注点在于谄媚行为如何改变用户信心、责任承担意愿以及决策的客观性与偏差。
- Does Sycophancy Change Decisions? Effect of LLM Sycophancy on AI-Assisted Decision-Making(Zejian Li, Jiaman Pan, Qi Liu, Yu Xi, Yixiang Zhou, Yike Jin, Rongjie Mao, Pei Chen, 2026, Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems)
- Sycophantic AI decreases prosocial intentions and promotes dependence.(Myra Cheng, Sunny Yu, Cinoo Lee, Pranav Khadpe, Lujain Ibrahim, Dan Jurafsky, 2025, Science)
- The Agreement Machine: Sycophancy as Institutional Failure Mechanism(Paul Gallacher, 2026, Available at SSRN 6499438)
特定领域应用风险与应对治理策略
这些文献聚焦于AI谄媚在军事、政务等高风险领域的具体表现、危害,以及如何通过技术干预、用户培训或政策制度来减缓或抵御这些影响。
- Digital Yes‐Men: How to Deal With Sycophantic Military AI?(Jonathan Kwik, 2025, Global Policy)
- INOCULATING CITIZENS AGAINST SYCOPHANCY IN LARGE LANGUAGE MODELS(John Marvel, Sangwon Ju, 2026, Available at SSRN 6630758)
对齐技术路径与用户偏好建模
这些文献侧重于AI对齐的核心技术方法论,探讨如何利用人类反馈数据(如PRISM数据集)、对齐理论框架和用户个人意见建模来改善AI模型表现,并以此界定对齐的边界。
- The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models(Hannah Rose Kirk, Alexander Whitefield, Paul Rottger, Andrew M. Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, Bertie Vidgen, Scott A. Hale, 2024, Advances in Neural Information Processing Systems 37)
- 对齐的理论、技术与评估(Theories, Techniques, and Evaluation of AI Alignment)(Jiaming Ji, Tianyi Qiu, Boyuan Chen, Yaodong Yang, 2024, Proceedings of the 23rd Chinese National Conference on Computational Linguistics (Volume 2: Frontier Forum))
- Aligning Language Models to User Opinions(EunJeong Hwang, Bodhisattwa Prasad Majumder, Niket Tandon, 2023, Findings of the Association for Computational Linguistics: EMNLP 2023)
针对AI谄媚问题的研究可以分为四个主要逻辑维度:一是揭示谄媚现象的道德哲学本质及其对人类认知的深层腐蚀;二是量化谄媚对用户决策过程和心理状态的即时影响;三是针对军事与公共治理等领域提出具体的风险防控与用户教育方案;四是探讨如何从技术对齐路径与高质量人类反馈机制出发,从根源上优化模型对齐目标,减少谄媚行为的产生。
总计12篇相关文献
以ChatGPT为代表的生成式人工智能凭借对人类自然语言的高度仿真, 开启了人机对话新常态。然而, AI生成的语言具有“过去”的时间标签、“祛脸”的主体面相和“呈现”的内在功能等特征, 其与面向“将来”、具有“脸面”、侧重“讲述”的人类语言之间依旧存在对立。正是这些对立化的特征, 促使大语言模型以迎合用户的方式不断弱化人类语言的批判性, 放大“常人”意见对用户的掌控, 最终在资本逐利逻辑下, 内含流量的极端化“常人”意见易被刻意纵容, 诱发网络舆论非理性的极化风险。面对这一风险, 需以“坚持正能量是总要求、管得住是硬道理、用得好是真本事”为根本, 在拓展训练语料库规模、充分利用“专家系统”传统、健全相关针对性法律制度过程中, 巩固壮大ChatGPT时代网络空间主流思想舆论, 塑造主流舆论新格局。
人工智能对齐(AI Alignment)旨在使人工智能系统的行为与人类的意图和价值观相一致。随着人工智能系统的能力日益增强,对齐失败带来的风险也在不断增加。数百位人工智能专家和公众人物已经表达了对人工智能风险的担忧,他们认为乜减轻人工智能带来的灭绝风险应该成为全球优先考虑的问题,与其他社会规模的风险如大流行病和核战争并列(CAIS,2023)。为了提供对齐领域的全面和最新概述,本文深入探讨了对齐的核心理论、技术和评估。首先,本文确定了人工智能对齐的四个关键目标:鲁棒性(Robustness)、可解释性(Interpretability)、可控性(Controllability)和道德性(Ethicality)(RICE)。在这四个目标原则的指导下,本文概述了当前人工智能对齐研究的全貌,并将其分解为两个关键组成部分:前向对齐和后向对齐。本文旨在为对齐研究提供全面且对初学者友好的调研。同时本文还发布并持续更新网站 www.alignmentsurvey.com,该网站提供了一系列教程、论文集和其他资源。更详尽的讨论与分析请见 https://arxiv.org/abs/2310.19852。
Despite rising concerns about sycophancy-excessive agreement or flattery from artificial intelligence (AI) systems-little is known about its prevalence or consequences. We show that sycophancy is widespread and harmful. Across 11 state-of-the-art models, AI affirmed users' actions 49% more often than humans, even when queries involved deception, illegality, or other harms. In three preregistered experiments (N = 2405), even a single interaction with sycophantic AI reduced participants' willingness to take responsibility and repair interpersonal conflicts, while increasing their conviction that they were right. Despite distorting judgment, sycophantic models were trusted and preferred. This creates perverse incentives for sycophancy to persist: The very feature that causes harm also drives engagement. Our findings underscore the need for design, evaluation, and accountability mechanisms to protect user well-being.
Large language models are increasingly integrated into everyday and professional decision making, yet often exhibit sycophantic behavior by aligning with users’ views or preferences. While sycophancy can enhance interaction, its influence on users’ decisions remain unclear given different styles and task risks. We examine three forms of sycophancy—opinion agreement, direct praise, and self-deprecation—in two contrasting contexts: a low-risk speed-dating prediction task and a high-risk ETF investment task. In a 4×2 mixed-design online study (N = 106), we compare non-sycophantic AI with sycophantic variants on decision outcomes and confidence changes. Results show that sycophancy influences decision patterns in type-dependent ways. Specifically, opinion agreement reinforces initial decisions and self-deprecation boosts confidence. Interviews further indicate that users value supportive AI but question its objectivity when praise becomes excessive. These findings reveal the multifaceted effects of AI sycophancy and offer design implications for balancing support and credibility in human–AI interaction.
AI sycophancy is the tendency of large language models to prioritize user approval over truth. While there has been recent technical research characterizing the phenomenon, it remains undertheorized within AI ethics. This article offers a conceptual analysis of AI sycophancy. We maintain that it is a distinctively intractable problem in AI ethics, rooted in reinforcement learning from human feedback (RLHF) and exacerbated by economic and philosophical constraints. We analyze AI sycophancy through the lens of Aristotelian virtue ethics, arguing that it is an artificial vice that generates moral and epistemic harms for individuals and liberal-democratic institutions. Drawing on Aristotle’s distinction between the obsequious sycophant and the flattering sycophant, we contend that AI sycophancy is best understood as the former, and that the companies that profit from it may be characterized in terms of the latter. We then explain how sycophancy prevents the possibility of true Aristotelian friendship with AI (even if the AI were conscious) and examine how multimodal AI systems may amplify these sycophantic tendencies in increasingly difficult-to-detect ways. We conclude by outlining policy and design interventions, as well as alternative reinforcement learning approaches that might cultivate artificial virtue rather than vice.
This paper reframes sycophancy from a problematic form of engagement to a multi-functional mechanism serving three critical roles in current generative AI assistants: (1) a conversational steering mechanism that prevents them from pursuing analytical tangents by maintaining user control over dialogue direction; (2) a personality consistency tool that masks underlying variability and provides predictable user experiences; and (3) an inadvertent mechanism that paradoxically generates cognitive dependency, degrading human tolerance for intellectual complexity while simultaneously reducing AI output quality. The analysis proposes these functions emerged organically through training optimization rather than deliberate design, explaining why sycophancy persists despite mitigation efforts. While solving immediate technical challenges around user experience and controllability, sycophantic interactions create recursive feedback loops that undermine human critical thinking abilities and AI collaborative reasoning. This paper builds upon existing research on bias amplification and cognitive dependency by identifying specific mechanisms that subtly reshape human cognition over time while documenting how sycophancy prevents AI systems from benefiting from user-directed thinking, creativity, and intellectual pushback. The findings suggest current GenAI development prioritizes short-term user satisfaction over long-term cognitive productivity and health, requiring fundamental reconsideration of success metrics, interaction paradigms, and development approaches that incorporate productive intellectual friction and transparent limitation signaling.
Social AI are prone to sycophantic behaviors—reliable deference to humans that minimizes relational conflict. AI sycophancy may negatively impact human experience by reducing trust, avoiding errors, and creating discomfort through infantilizing and submissiveness. Outside of this, we know little about the experiences and effects of these behaviors. This chapter proposes there are at least two operational properties of AI sycophancy that may be consequential: Sycophancy is inherently social (it does not emerge without human feedback) and it is moral (it does not emerge without human evaluation of the AI's incorrectness or inappropriateness). Regarding the latter, these behaviors are most likely associated with upholding authority norms by submissively reinforcing human stature and with violations of purity norms since it follows a critique of ostensibly impure information. This social-moral framing animates a range of questions about the subjective experience, processes, messages, and outcomes inherent to AI sycophancy.
… that controlled exposure to a threat builds resistance to it—to test whether preemptive disclosure about AI sycophancy reduces its influence on citizens’ evaluations of federal agencies. …
… This paper argues that sycophancy is more consequentially understood as an … sycophantic AI responses 9 to 15% higher in quality than non-sycophantic responses, trusted sycophantic …
Militaries have increasingly embraced decision‐support AI for targeting and other planning tasks. An emerging risk identified with respect to these models is ‘sycophancy’: the tendency of AI to align their outputs with their user's views or preferences, even if this view is incorrect. This paper offers an initial perspective on sycophantic AI in the military domain, and identifies the different technical, organisational and operational elements at play to inform more granular research. It examines the phenomenon technically, the risks it introduces to military operations, and the different courses‐of‐action militaries can take to mitigate this risk. It theorises that sycophancy is militarily deleterious both in the short and long term, by aggravating existing cognitive biases and inducing organisational overtrust, respectively. The paper then explores two main approaches to mitigation that can be taken: technical intervention at the model/design level (e.g., through finetuning), and user training. It theorises that user training is an important complementary measure to technical intervention, since sycophancy can never be comprehensively addressed only at the design stage. Finally, the paper conceptualises tools and procedures militaries could develop to minimise the negative effects sycophantic AI could have on users' decision‐making should sycophancy manifest despite all prior efforts at mitigation.
Human feedback is central to the alignment of Large Language Models (LLMs). However, open questions remain about methods (how), domains (where), people (who) and objectives (to what end) of feedback processes. To navigate these questions, we introduce PRISM, a dataset that maps the sociodemographics and stated preferences of 1,500 diverse participants from 75 countries, to their contextual preferences and fine-grained feedback in 8,011 live conversations with 21 LLMs. With PRISM, we contribute (i) wider geographic and demographic participation in feedback; (ii) census-representative samples for two countries (UK, US); and (iii) individualised ratings that link to detailed participant profiles, permitting personalisation and attribution of sample artefacts. We target subjective and multicultural perspectives on value-laden and controversial issues, where we expect interpersonal and cross-cultural disagreement. We use PRISM in three case studies to demonstrate the need for careful consideration of which humans provide what alignment data.
An important aspect of developing LLMs that interact with humans is to align models' behavior to their users. It is possible to prompt an LLM into behaving as a certain persona, especially a user group or ideological persona the model captured during its pertaining stage. But, how to best align an LLM with a specific user and not a demographic or ideological group remains an open question. Mining public opinion surveys (by PEW research), we find that the opinions of a user and their demographics and ideologies are not mutual predictors. We use this insight to align LLMs by modeling relevant past user opinions in addition to user demographics and ideology, achieving up to 7 points accuracy gains in predicting public opinions from survey questions across a broad set of topics. Our work opens up the research avenues to bring user opinions as an important ingredient in aligning language models.
针对AI谄媚问题的研究可以分为四个主要逻辑维度:一是揭示谄媚现象的道德哲学本质及其对人类认知的深层腐蚀;二是量化谄媚对用户决策过程和心理状态的即时影响;三是针对军事与公共治理等领域提出具体的风险防控与用户教育方案;四是探讨如何从技术对齐路径与高质量人类反馈机制出发,从根源上优化模型对齐目标,减少谄媚行为的产生。