rater severity drift in peer assessment

Clicks: 158
ID: 200454
2017
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal
Emerging

Ranked #12 of 30 articles by views in bmc international health and human rights

Most read Least read

Bar heights use a square-root scale.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
There are not enough psychometrically sound studies about the validity and reliability of the scores obtained from peer assessment. This study examined degree of rater severity drift, a rater effect, in peer assessment. The college students’ presentations were scored by 29 peers in the class using a rating scale. Nine presentations were held on four separate days, two presentations on each of the first three days and three presentations on the fourth day. Drift was investigated with two many-facet Rasch measurement models (separate models and dummy time MFRM). Standardized differences were calculated from the estimates obtained with separate models and interaction terms were calculated with the dummy time MFRM. In drift analysis, shifts in estimations were examined from Day-1 which is a baseline to other three days. Results showed that peer raters varied according to their level of severity and they tend to be lenient. Statistics showed that the quality of the scale was acceptable and its items behaved as expected. In drift analysis, standardized differences and interaction term provided very similar results. Between Day-1 and 2, there was no statistically significant difference in the estimates of the rater severity. Between Day-1 and 3, the percentage of scorers with significant drift in the estimates was 38.10%. The raters’ severity shifts on the average of about 0.14 logit and they displayed more severe scoring behavior. Between Day-1 and 4, the number of raters who had significant shifts in their estimates was three according to the standardized difference method, while one according to interaction method. On the average, the raters became more severe. Among three comparison, Day-4 had the largest rater severity drift on the average; Day 3, however, has the highest number of raters with rater severity drift.
Reference Key
brkan2017journalrater Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors ;Bengü BÖRKAN
Journal bmc international health and human rights
Year 2017
DOI
10.21031/epod.328119
URL
Keywords

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.