TPTh 1.10 Can Evidence-Based Prompting Enable AI to Generate UKMLA-Quality MCQs? A Blinded Evaluation
Clicks: 1
ID: 323906
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
0.0
/100
1 views
0 readers
AI Quality Assessment
Not analyzed
Readership in this journal
Ranked #340 of 432 articles by views in the british journal of surgery
Most read
Least read
Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 432 in total.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Abstract Aims Developing high-quality multiple-choice questions (MCQs) is a time-consuming process that places additional pressure on already stretched clinician-educators. Large language models (LLMs), such as ChatGPT, have demonstrated promise in streamlining MCQ development and optimising the question creation workflow. While emerging evidence suggests that LLM-generated MCQs are comparable to those written by human experts, their suitability for UK Medical Licensing Assessment (UKMLA)–aligned learning has not been established. Furthermore, concerns remain regarding factual accuracy and cognitive depth. This study evaluates whether embedding an evidence-based framework into question generation enables LLMs to produce high-quality MCQs suitable for UKMLA-aligned learning and revision. Methods This is an ongoing, blinded, cross-sectional comparative evaluation study. Twenty UKMLA-mapped MCQs generated using ChatGPT and guided by the five A’s of evidence-based medicine will be compared with twenty educator-authored MCQs. Independent consultant experts, blinded to question origin, will evaluate all items. Outcomes will include perceived origin, content validity, educational value and item performance metrics, alongside time taken for question generation. Results We predict that ChatGPT-generated MCQs will be comparable to human-authored questions in quality, educational value and item performance, while being produced significantly faster. Experts are expected to be unable to distinguish question origin. Conclusion This study will provide evidence on the feasibility of using evidence-informed LLMs to support UKMLA-aligned MCQ development. Findings will help inform the role of AI in medical education, with potential to reduce educator workload while maintaining assessment quality.
| Reference Key |
openalex_W7172540361
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Arthur Bowers-Barnard, Evripidis Tokidis |
| Journal | the british journal of surgery |
| Year | 2026 |
| DOI |
10.1093/bjs/znag087.423
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.