How Swedish teachers used RM Compare to support assessment calibration at scale

The challenge

School systems and education ministries increasingly focus assessment on complex, open-ended outcomes and artefacts.

Consistently evaluating skills, knowledge and understanding in this context can be difficult. Other drawbacks include:

  • variation between assessors can be the result of traditional rubric-based marking
  • an emphasis on surface-level features rather than the quality of learning
  • substantial moderation workload

Dr Eva Hartell of Stockholm’s KTH Royal Institute of Technology led a repeated-measures study in Sweden to explore whether a distributed assessors group could develop a shared understanding of quality through Adaptive Comparative Judgement (ACJ). The study also examined whether that shared standard could help create reliable assessment scales with fewer judgements over time.

Summary of findings

A short comparative judgement experience, even for teachers new to the system, delivered highly impactful results.

  • Assessors were making more consistent decisions after taking part in the comparative judgement exercise.
  • Assessors had internalised a clearer standard of quality. With less cognitive friction, they could focus on making informed professional judgements rather than repeatedly interpreting criteria from first principles.
  • Assessors also used more precise, shared language connected to curriculum and subject expectations. Their comments became shorter and more decisive, indicating greater confidence in how they defined and recognised quality.

The findings suggest that, once a panel has developed a shared standard, RM Compare can support the construction of reliable scales with reduced judging volume.

The results support previous research about the power and impact of Learning by Evaluating.

The study shows that there is now a clear way forward for sharing and securing learners’ performance standards across schools. This can be done at scale across regions, countries or even continents.

The approach

The study used RM Compare across two linked sessions with the same assessor panel and assessment context.

Establishing a baseline:

In the first session, assessors judged complex student work using their existing assessment habits and internal standards. This provided a baseline rank order of work, while revealing the natural variation in evaluator alignment, decision times and approaches to judging quality.

the approach image
Testing calibration:

The same panel then completed a shorter second session in the same assessment context. The platform configuration, assessment domain and item pool remained unchanged, allowing the study to explore the effect of assessors’ experience.

The central question was simple: after working comparatively through the first session, would assessors have developed a clearer, more shared view of quality - and could RM Compare produce a stable scale more efficiently as a result?

The impacts

1. More consistent evaluator judgement

RM Compare uses a judge-misfit measure to show how closely an assessor’s decisions align with the emerging consensus of the group.

During the first session, assessor alignment varied widely. Some judges’ decisions differed noticeably from the group’s emerging view of quality.

Evaluator judgement

In the second session, those patterns moved closer to the cohort consensus. Misfit values clustered into a tighter range, suggesting that the same assessors were making more consistent decisions after taking part in the initial comparative-judgement exercise.

Rather than relying on lengthy moderation meetings to achieve alignment, the comparative process helped assessors calibrate through the act of judging.

2. Greater fluency in decision making

In the baseline session, decision times varied considerably, including longer periods of hesitation and occasional rapid decisions. In the second session, decision times became more consistent, with assessors making judgements at a steadier professional pace.

This suggests that assessors had begun to internalise a clearer standard of quality. With less cognitive friction, they could focus on making informed professional judgements rather than repeatedly interpreting criteria from first principles.

Greater fluency
3. A meaningful change in how assessors described their decisions - from checklists to holistic judgement

In the first session, comments often reflected a checklist mindset, with attention given to layout, formatting and other surface-level features. By the second session, comments focused more often on clarity of reasoning, conceptual depth and the coherence of each response.

Assessors also used more precise, shared language connected to curriculum and subject expectations. Their comments became shorter and more decisive, indicating greater confidence in how they defined and recognised quality.

4.  Stable scales with fewer judgements

The second session was intentionally shorter but still produced a stable scale.

High-quality anchor pieces and lower performing work occupied similar positions across both sessions, while the overall ordering of work showed strong agreement. Difficult boundary cases - work that sits between performance levels - also settled more consistently into intermediate positions.

The findings suggest that, once a panel has developed a shared standard, RM Compare can support the construction of reliable scales with reduced judging volume.

 Stable scales diagram

Conclusion

A framework for scalable assessment

Dr Hartell’s study points towards a practical approach for systems seeking to build more consistent, efficient and transparent assessment.

  • Use a high-density comparative-judgement session with a trusted panel to establish a shared rank order, align evaluator expectations and identify meaningful examples of quality.

  • Capture representative work at key points on the scale and retain it as a shared reference. These exemplars become visual and conceptual anchors for the agreed standard.

  • Use those anchor items within ongoing assessment sessions across schools, trusts or regions. New work can then be positioned against an established system-wide scale with relatively few judgements.

What next?

The next phase of the project is underway with a larger sample of teachers from across Sweden. The intention is to also test the
approach across more curriculum areas.