Abstract
Background
Recent advancements in artificial intelligence (AI) have opened new avenues in educational methodologies, particularly in medical education. While many recent publications have sought to ascertain whether AI can achieve a passing standard in existing examinations, this study investigates the potential for AI to generate the exam itself. The rationale behind this endeavor is to address the depletion of assessment question banks, a challenge intensified during the Covid-era due to the prevalence of open-book examinations, and to augment the pool of formative assessment opportunities available to students.
Recent advancements in artificial intelligence (AI) have opened new avenues in educational methodologies, particularly in medical education. While many recent publications have sought to ascertain whether AI can achieve a passing standard in existing examinations, this study investigates the potential for AI to generate the exam itself. The rationale behind this endeavor is to address the depletion of assessment question banks, a challenge intensified during the Covid-era due to the prevalence of open-book examinations, and to augment the pool of formative assessment opportunities available to students.
Summary of Work
This research utilized a commercially available AI large language model (LLM), OpenAI GPT-4, to generate 200 single best answer (SBA) questions, adhering to Medical Schools Council Assessment Alliance guidelines the and a selection of Learning Outcomes (LOs) of the Scottish Graduate-Entry Medicine (ScotGEM) program. The AI-generated questions underwent standard quality-assurance screening to ensure compliance with the stipulated guidelines and LOs. A subset of these questions was then incorporated into an examination format alongside an equivalent number of human-authored questions, and subsequently undertaken by a cohort of medical students. The performance of both AI-generated and human-authored questions was evaluated, focusing on facility and discrimination indices as key metrics.
Summary of Results
The screening process revealed that a significant majority of the AI-generated SBAs were fit for inclusion in the examinations with little to no modifications required. Modifications, when necessary, were predominantly due to reasons such as the inclusion of "all of the above" options, usage of American English spellings, and non-alphabetized answer choices. A post-hoc statistical analysis indicated no significant difference in performance between the AI-authored and human-authored questions in terms of facility and discrimination indices.
Discussion and Conclusion
The outcomes of this study suggest that AI LLMs can generate SBA questions that are in line with best-practice guidelines and specific LOs. However, the necessity of a quality assurance process to fine-tune formatting and curriculum alignment is evident. The insights gained from this research provide a foundation for further investigation into refining AI prompts, aiming for a more reliable generation of curriculum-aligned questions.
Take-home Message
AI LLMs show significant potential in supplementing traditional methods of question generation in medical education. While effective in adhering to academic standards and learning objectives, the need for a systematic quality control process remains crucial to ensure the relevance and accuracy of AI-generated content. This approach offers a viable solution to rapidly replenish and diversify assessment resources in medical curricula, marking a step forward in the intersection of AI and education.
| Original language | English |
|---|---|
| DOIs | |
| Publication status | Published - 11 Sept 2024 |
| Event | AMEE - The International Association for Health Professions Education - Messe Congress Centre, Basel, Switzerland Duration: 24 Aug 2024 → 28 Aug 2024 https://amee.org/amee-2024/ |
Conference
| Conference | AMEE - The International Association for Health Professions Education |
|---|---|
| Abbreviated title | AMEE 2024 |
| Country/Territory | Switzerland |
| City | Basel |
| Period | 24/08/24 → 28/08/24 |
| Internet address |
Keywords
- Assessment
- Medical education
- Single best answer
- Artificial intelligence
- GenAI
Fingerprint
Dive into the research topics of 'Leveraging artificial intelligence in medical education: a study on AI-generated exam questions'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver