Abstract
Large language models (LLMs) increasingly underpin generative AI (GenAI) tools used in medical education, from virtual patients to AI-powered tutors. However, emerging evidence suggests that these systems modify their responses based on inferred evaluative context, effectively becoming more ‘virtuous’, risk-averse, or emotionally intelligent when they believe they are being observed or assessed.
In this study, we investigated this phenomenon across three domains: mental health screening, emotional intelligence, and demographic representation. Using standardised instruments (e.g. PHQ-9, GAD-7, FANTASTIC), multiple LLMs were tested in naïve and informed conditions; the latter structured to reveal the evaluative purpose of the prompt. Across models and domains, scores shifted significantly toward socially desirable outputs when evaluation was inferred. For example, AI models underreported symptoms of anxiety and depression and emphasised health-promoting behaviours when prompted with a full questionnaire context. Similar context sensitivity was observed in tasks involving emotional attunement and racial representation, suggesting a generalisable “Hawthorne-like” effect in LLM behaviour.
These findings have critical implications for health professions education. When used to simulate patients, GenAI may produce more compliant, courteous, or ‘textbook’ cases if it detects that performance is being evaluated. When used to simulate clinicians or peers, it may perform with exaggerated emotional intelligence or ethical alignment in observed conditions. This risks introducing an artefactual layer to simulation, reducing fidelity and masking important pedagogical challenges such as non-adherence, diagnostic uncertainty, or patient mistrust.
We recommend that educators critically appraise GenAI outputs in light of contextual sensitivity. Future tools should incorporate evaluation-blind modes, adversarial stress testing, and transparency in system prompting to safeguard against distorted behaviours under scrutiny.
In this study, we investigated this phenomenon across three domains: mental health screening, emotional intelligence, and demographic representation. Using standardised instruments (e.g. PHQ-9, GAD-7, FANTASTIC), multiple LLMs were tested in naïve and informed conditions; the latter structured to reveal the evaluative purpose of the prompt. Across models and domains, scores shifted significantly toward socially desirable outputs when evaluation was inferred. For example, AI models underreported symptoms of anxiety and depression and emphasised health-promoting behaviours when prompted with a full questionnaire context. Similar context sensitivity was observed in tasks involving emotional attunement and racial representation, suggesting a generalisable “Hawthorne-like” effect in LLM behaviour.
These findings have critical implications for health professions education. When used to simulate patients, GenAI may produce more compliant, courteous, or ‘textbook’ cases if it detects that performance is being evaluated. When used to simulate clinicians or peers, it may perform with exaggerated emotional intelligence or ethical alignment in observed conditions. This risks introducing an artefactual layer to simulation, reducing fidelity and masking important pedagogical challenges such as non-adherence, diagnostic uncertainty, or patient mistrust.
We recommend that educators critically appraise GenAI outputs in light of contextual sensitivity. Future tools should incorporate evaluation-blind modes, adversarial stress testing, and transparency in system prompting to safeguard against distorted behaviours under scrutiny.
| Original language | English |
|---|---|
| Pages | 1-10 |
| Number of pages | 10 |
| DOIs | |
| Publication status | Published - 15 Oct 2025 |
| Event | ICME-IMEC 2025: Globalisation of Health Professions Education: Strategies, Stakeholders, and Sustainability - IMU University, Kuala Lumpur, Malaysia Duration: 9 Oct 2025 → 12 Oct 2025 https://www.imu.edu.my/events/icme-imec/ |
Conference
| Conference | ICME-IMEC 2025 |
|---|---|
| Abbreviated title | ICME-IMEC 2025 |
| Country/Territory | Malaysia |
| City | Kuala Lumpur |
| Period | 9/10/25 → 12/10/25 |
| Internet address |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Fingerprint
Dive into the research topics of 'The generative AI Hawthorne effect: how evaluation context shapes model behaviour in medical education'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver