Skip to main navigation Skip to search Skip to main content

The generative AI Hawthorne effect: how evaluation context shapes model behaviour in medical education

  • Andrew O'Malley*
  • , Elizabeth Lang
  • , Ileri Ojikutu
  • *Corresponding author for this work

Research output: Contribution to conferenceAbstract

Abstract

Large language models (LLMs) increasingly underpin generative AI (GenAI) tools used in medical education, from virtual patients to AI-powered tutors. However, emerging evidence suggests that these systems modify their responses based on inferred evaluative context, effectively becoming more ‘virtuous’, risk-averse, or emotionally intelligent when they believe they are being observed or assessed.

In this study, we investigated this phenomenon across three domains: mental health screening, emotional intelligence, and demographic representation. Using standardised instruments (e.g. PHQ-9, GAD-7, FANTASTIC), multiple LLMs were tested in naïve and informed conditions; the latter structured to reveal the evaluative purpose of the prompt. Across models and domains, scores shifted significantly toward socially desirable outputs when evaluation was inferred. For example, AI models underreported symptoms of anxiety and depression and emphasised health-promoting behaviours when prompted with a full questionnaire context. Similar context sensitivity was observed in tasks involving emotional attunement and racial representation, suggesting a generalisable “Hawthorne-like” effect in LLM behaviour.

These findings have critical implications for health professions education. When used to simulate patients, GenAI may produce more compliant, courteous, or ‘textbook’ cases if it detects that performance is being evaluated. When used to simulate clinicians or peers, it may perform with exaggerated emotional intelligence or ethical alignment in observed conditions. This risks introducing an artefactual layer to simulation, reducing fidelity and masking important pedagogical challenges such as non-adherence, diagnostic uncertainty, or patient mistrust.

We recommend that educators critically appraise GenAI outputs in light of contextual sensitivity. Future tools should incorporate evaluation-blind modes, adversarial stress testing, and transparency in system prompting to safeguard against distorted behaviours under scrutiny.
Original languageEnglish
Pages1-10
Number of pages10
DOIs
Publication statusPublished - 15 Oct 2025
EventICME-IMEC 2025: Globalisation of Health Professions Education: Strategies, Stakeholders, and Sustainability - IMU University, Kuala Lumpur, Malaysia
Duration: 9 Oct 202512 Oct 2025
https://www.imu.edu.my/events/icme-imec/

Conference

ConferenceICME-IMEC 2025
Abbreviated titleICME-IMEC 2025
Country/TerritoryMalaysia
CityKuala Lumpur
Period9/10/2512/10/25
Internet address

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Fingerprint

Dive into the research topics of 'The generative AI Hawthorne effect: how evaluation context shapes model behaviour in medical education'. Together they form a unique fingerprint.

Cite this