Every teacher has met the student who hands in an essay, gets it back covered in comments, glances at the grade and files it away. The feedback was careful. It changed nothing. The problem was not the feedback; it was that the student had no way to use it, because they had never been asked to judge the work themselves.
Self-evaluation, the habit of a learner checking their own work against a clear standard and acting on what they find, is one of the better-supported ideas in education research. It is also one of the most misunderstood, and the gap between “students should reflect” and a classroom where reflection actually moves learning is where most of the value is lost. This post walks through what the evidence says, where it is honest about its limits, and what it looks like on a Tuesday in an Islamic school.
What the research actually shows
The foundational argument comes from Royce Sadler’s 1989 paper on formative assessment. He set out three conditions a learner must meet to improve: hold a concept of the standard being aimed for, compare their current performance with that standard, and take action to close the gap. His conclusion is blunt: “It is insufficient for students to rely upon evaluative judgments made by the teacher.” The student has to come to hold, in his words, “a concept of quality roughly similar to that held by the teacher”.
Black and Wiliam’s 1998 review of classroom assessment built on this and found the gains from formative assessment to be “amongst the largest ever reported for educational interventions”. They were direct about the student’s role: “self-assessment by the student is not an interesting option or luxury; it has to be seen as essential.”
More recent meta-analyses put numbers on it. Panadero, Jonsson and Botella (2017) pooled 19 studies with 2,305 students and found self-assessment interventions produced effect sizes of 0.23, 0.65 and 0.43 on three measures of self-regulated learning, and 0.73 on self-efficacy. Interventions that included self-monitoring as a component did better, and girls benefited more than boys. Yan, Wang, Boud and Lao (2021), working with 26 studies in higher education, found an overall effect on academic performance of g = 0.455, and a striking split: where self-assessment was paired with explicit feedback from others, the effect was 0.664; without it, 0.213.
Barry Zimmerman’s overview of self-regulated learning gives the mechanism. Learning runs in a cycle of forethought, performance and self-reflection, and self-reflection is where a student compares their result to a standard and decides what caused it. His line is worth pinning above a desk: “Self-regulation is not a mental ability or an academic performance skill; rather it is the self-directive process by which learners transform their mental abilities into academic skills.” He also noted, with some frustration, that “students are rarely asked to self-evaluate their work”.
Where the evidence is mixed, and where the hype is
You may have seen “self-reported grades” near the top of John Hattie’s rankings, listed at d = 1.33. Treat that number carefully. Critics have pointed out that the underlying studies measure how well students’ self-reports correlate with their actual grades. They do not show that the act of predicting a grade causes anyone to learn more. A correlation between knowing yourself and doing well is interesting; it is not an intervention.
Students are not especially accurate judges of their own work. Brown, Andrade and Chen (2015) reviewed the accuracy literature and found that correlations between self-ratings and teacher ratings “tended to be only weakly positive”. The pattern has a shape. Younger children overestimate more: in one study they cite, the proportion of poor performers who overrated themselves fell from 85% at ages five to nine to 43% at ages ten and eleven. Higher achievers judge themselves more in line with their teachers; lower achievers tend to overrate. This is the classroom version of what Kruger and Dunning (1999) found in adults, where participants scoring in the 12th percentile estimated themselves in the 62nd. The hopeful part of that study is often forgotten: when participants’ skills improved, so did their ability to judge them.
Practice alone does not fix this. Brown and colleagues are explicit that self-assessment “needs to be taught”, that “practice alone is an insufficient condition”, and that accuracy improves when the task is concrete and “the more specific and concrete the reference criteria”. Self-assessment without criteria is a mood survey. And once self-assessment counts toward a grade, students are tempted to inflate it; the authors’ advice is to keep it formative.
There is a subtler point in Heidi Andrade’s 2019 review. Accuracy may not even be the goal. She argues the benefit “may come from active engagement in the learning process, rather than by being ‘veridical’”. A student who compares their paragraph to a rubric, notices a gap and revises has learned something, whether or not their final self-score matches yours.
One more caution. Dunlosky and colleagues (2013) reviewed ten study techniques and rated rereading and highlighting as low utility, despite students reporting they rely on them, while practice testing and distributed practice rated highest. Students’ sense of what is working is often wrong, which is exactly why self-evaluation has to be trained against something more solid than a feeling.
What it looks like in a classroom

The conditions above collapse into three moves: a clear standard, a structured comparison, and a required next action. Here is how they land in three ordinary lessons.
A Grade 6 writing lesson. Before anyone drafts, put two short model paragraphs on the board, one strong and one weak, and have the class say what makes the difference. Turn what they say into three or four criteria in their own words. Students draft, then, before you see it, mark their own paragraph against those criteria with a highlighter: green where the criterion is met, yellow where it is not, and one sentence on what they will change. Only then do they revise and hand in. Andrade, Du and Wang (2008) tested exactly this sequence with Grade 3 and 4 students, a model, criteria generation and rubric-referenced self-assessment, and the treatment group wrote measurably better. The order matters: the model comes before the criteria, and the self-check comes before the teacher’s mark.
A Grade 8 maths practice set. Give the answers with the questions. Students work, check each one, and sort their errors into two piles: “I know what I did wrong” and “I do not know what I did wrong”. The second pile is the lesson. Ask each student to write one line beside each item in that pile about what they think went wrong, then compare with a partner before you intervene. This turns marking into diagnosis and gives you a map of the room’s misconceptions in five minutes. Brown and colleagues note that Grade 5 and 6 students taught a self-checking strategy in long division became more accurate self-assessors, and Zimmerman’s distinction is the one to teach explicitly: experts attribute errors to strategy, novices to ability.
A three-minute exit reflection. Skip “How do you feel about today?” Ask instead: What was the hardest part? What would you do differently on a similar question? What is one thing you still cannot do? The third question is the one that matters, because it is the one a student will avoid unless it is asked. Collect the slips; read them in a minute; open tomorrow’s lesson with the most common answer.
Three things make these work:
- Criteria first, in plain language, ideally co-written with the class. Students cannot compare their work with a standard they have never seen.
- Model the judgement out loud. Think through a piece of anonymous work in front of them, showing the hesitation and the second look. This is Sadler’s “guild knowledge” being handed over.
- Follow every judgement with an action. A self-assessment that ends in a score and goes nowhere is the version the research finds ineffective. The Yan meta-analysis is the reminder: pair self-assessment with explicit feedback and the effect roughly triples.
Expect the first attempts to be off. Younger and weaker students will overrate themselves; that is the starting point, not a reason to stop. Their calibration improves as their skill does, and it improves faster when a teacher is showing them what quality looks like.
The habit has a name

Islamic tradition has a word for this discipline. Muhasabah, self-accounting, is the practice of taking stock of your own actions before someone else does. Umar ibn al-Khattab put it plainly: “Hold yourselves accountable before you are held accountable and evaluate yourselves before you are evaluated” (Ibn Abi al-Dunya, Muhasabat al-Nafs, 2). The Qur’an frames the same instruction: “let every soul look to what it has sent forth for tomorrow” (59:18).
A student who learns to look honestly at their own paragraph, their own working, their own understanding, and then to act on what they see, is practising a habit that reaches well past the classroom.
We would not push the parallel too hard. A rubric is not a spiritual discipline, and a reflection slip is not repentance. But teachers in Islamic schools are in the unusual position of teaching a skill the research supports with a vocabulary their students already own. The student who marks their own work with honesty and then does something about it is doing, in miniature, what they are asked to do with their whole life. For what it is worth, that is why every RISE chapter deck closes on a reflection prompt rather than a summary slide.
The research says the same thing the tradition does: the judgement that changes you is the one you make yourself, against a standard you understand, followed by a decision to do better.
Sources
- Andrade, H. L. (2019). A critical review of research on student self-assessment. Frontiers in Education, 4, 87. doi:10.3389/feduc.2019.00087
- Andrade, H. L., Du, Y., & Wang, X. (2008). Putting rubrics to the test: The effect of a model, criteria generation, and rubric-referenced self-assessment on elementary school students’ writing. Educational Measurement: Issues and Practice, 27(2), 3–13. doi:10.1111/j.1745-3992.2008.00118.x
- Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 5(1), 7–74. doi:10.1080/0969595980050102
- Brown, G. T. L., Andrade, H. L., & Chen, F. (2015). Accuracy in student self-assessment: Directions and cautions for research. Assessment in Education: Principles, Policy & Practice, 22(4), 444–457. doi:10.1080/0969594X.2014.996523
- Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. (2013). Improving students’ learning with effective learning techniques: Promising directions from cognitive and educational psychology. Psychological Science in the Public Interest, 14(1), 4–58. doi:10.1177/1529100612453266
- Kruger, J., & Dunning, D. (1999). Unskilled and unaware of it: How difficulties in recognizing one’s own incompetence lead to inflated self-assessments. Journal of Personality and Social Psychology, 77(6), 1121–1134. doi:10.1037/0022-3514.77.6.1121
- Panadero, E., Jonsson, A., & Botella, J. (2017). Effects of self-assessment on self-regulated learning and self-efficacy: Four meta-analyses. Educational Research Review, 22, 74–98. doi:10.1016/j.edurev.2017.08.004
- Sadler, D. R. (1989). Formative assessment and the design of instructional systems. Instructional Science, 18(2), 119–144. doi:10.1007/BF00117714
- Yan, Z., Wang, X., Boud, D., & Lao, H. (2021). The effect of self-assessment on academic performance and the role of explicitness: A meta-analysis. Assessment & Evaluation in Higher Education, 48(1), 1–15. doi:10.1080/02602938.2021.2012644
- Zimmerman, B. J. (2002). Becoming a self-regulated learner: An overview. Theory Into Practice, 41(2), 64–70. doi:10.1207/s15430421tip4102_2
