Development and validation of a biology item bank using item response theory – four parameter logistic (IRT–4PL) model

Clicks: 1
ID: 287502
2025
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal

Ranked #3,749 of 3,757 articles by views in Malay Journal

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 3,757 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
This study developed and psychometrically validated a Biology Item Bank using the Item Response Theory Four-Parameter Logistic (IRT–4PL) model, aimed at providing a standardized pool of calibrated items for Senior High School STEM students preparing for biology-intensive and health-allied college programs. The development process followed a multi-phase validation protocol integrating expert evaluation, empirical testing, and advanced psychometric modeling. An initial pool of 120 multiple-choice items was constructed and reviewed by five biology educators through online focus group discussions. Items were evaluated for content accuracy, linguistic clarity, and curricular relevance, and were classified based on Bloom’s revised taxonomy across six cognitive levels. A pilot validation confirmed semantic and content appropriateness, after which the test was administered to 1,017 STEM students from a private university in Metro Manila. Both dichotomous and polytomous scoring were employed, enabling robust distractor analysis. Reliability analysis yielded a strong Cronbach’s alpha (α = 0.920), which improved slightly (α = 0.923) after the removal of underperforming items. Additional distractor diagnostics resulted in revisions and refinements, producing an 88-item calibrated pool. Structural validity was established through exploratory factor analysis (KMO = 0.879; Bartlett’s test, p < .001) and confirmatory factor analysis, which demonstrated acceptable fit indices (RMSEA = 0.013, SRMR = 0.025, TLI = 0.916, CFI = 0.932). Item-level calibration under the IRT–4PL model provided parameter estimates for discrimination (a), difficulty (b), guessing (c), and slipping (d). Results indicated a small number of items with misfit or overfitting, while the majority performed within psychometric expectations. The Item Characteristic Curves (ICCs) displayed the psychometric soundness of retained items across cognitive domains. Based on integrated statistical and expert criteria, the final classification consisted of 18 retained items, 20 revised, 16 reassigned to alternative domains, and 66 rejected due to psychometric flaws. This study affirms the utility of the IRT–4PL model in developing item banks for high-stakes assessments. The finalized test, rigorously validated, provides a dependable source of calibrated items for biology assessments and diagnostic purposes. Moreover, the study recommends extending the IRT–4PL framework to the development of item banks in other science domains, ensuring validity, fairness, and pedagogical alignment in assessment design. Keywords: Biology test, IRT–4PL, item bank, discrimination, difficulty, guessing, slipping, factor analysis, latent trait/construct
Reference Key
persistent_1760661788_68f1911c6f794 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Mirabete, Glen S.
Journal Malay Journal
Year 2025
DOI
DOI not found
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.