[Submitted connected 1 Oct 2025]
View PDF HTML (experimental)
Abstract:Both the wide nationalist and world communities person raised concerns astir sycophancy, the arena of artificial intelligence (AI) excessively agreeing pinch aliases flattering users. Yet, beyond isolated media reports of terrible consequences, for illustration reinforcing delusions, small is known astir the grade of sycophancy aliases really it affects group who usage AI. Here we show the pervasiveness and harmful impacts of sycophancy erstwhile group activity proposal from AI. First, crossed 11 state-of-the-art AI models, we find that models are highly sycophantic: they affirm users' actions 50% much than humans do, and they do truthful moreover successful cases wherever personification queries mention manipulation, deception, aliases different relational harms. Second, successful 2 preregistered experiments (N = 1604), including a live-interaction study wherever participants talk a existent interpersonal conflict from their life, we find that relationship pinch sycophantic AI models importantly reduced participants' willingness to return actions to repair interpersonal conflict, while expanding their condemnation of being successful the right. However, participants rated sycophantic responses arsenic higher quality, trusted the sycophantic AI exemplary more, and were much consenting to usage it again. This suggests that group are drawn to AI that unquestioningly validate, moreover arsenic that validation risks eroding their judgement and reducing their inclination toward prosocial behavior. These preferences create perverse incentives some for group to progressively trust connected sycophantic AI models and for AI exemplary training to favour sycophancy. Our findings item the necessity of explicitly addressing this inducement building to mitigate the wide risks of AI sycophancy.Submission history
From: Myra Cheng [view email]
[v1] Wed, 1 Oct 2025 19:26:01 UTC (5,571 KB)
English (US) ·
Indonesian (ID) ·