Article
Designing Pre and Post Assessment That Actually Shows Growth
A pre and post assessment shows growth only when the two measurements use the same standard. In student self-assessment, that is often not true. After an internship, practicum, or capstone, students frequently understand the skill better than they did at the start, and that new understanding can make their post score drop even when their actual performance improved. That is response-shift bias, and if you ignore it, you can mistake better judgment for weaker learning.
Why simple pre and post scores can mislead
A basic pre test asks students to rate their own ability before an experience. A post test asks the same question afterward. On paper, the comparison is straightforward. In practice, the two ratings may not be on the same scale.
A student who has never given professional feedback may rate themselves a 4 out of 5 on communication because they can write clearly in class. After several employer check-ins, they may realize that workplace communication includes tone, timing, audience awareness, and follow-through. On the post assessment, that same student may rate themselves a 3 because their standard changed. The score fell, but the skill likely improved.
This is not a flaw unique to students. It appears in any setting where the experience changes the learner’s frame of reference. Higher education sees it often in internships, clinical placements, service learning, and leadership roles because those settings make the target behavior more concrete.
What response-shift bias looks like
Response-shift bias happens when the experience changes one or more of these elements:
The meaning of the competency
Before the experience, a student may think “teamwork” means showing up and being polite. After the experience, they may define it as anticipating conflict, sharing workload, and adapting to group norms.
The standard for good performance
A student may rate a reflection paper highly before learning what “strong reflection” looks like in a professional context. After supervisor feedback, they may judge their own work more critically because they now recognize gaps they could not see before.
The evidence they use to judge themselves
At the start, students often rate themselves by effort or intent. At the end, they may rate themselves by observed behavior, examples, or feedback from others. That change can produce lower post scores even when the underlying competence improved.
The practical problem is not just measurement noise. If a program uses only raw pre and post self-ratings, it may conclude that a placement was ineffective when the real issue is that students learned enough to judge themselves more accurately.
Why the retrospective pre-test helps
A retrospective pre-test asks students, after the experience, to rate their skill at the start and at the end using the same current frame of reference. Instead of asking, “How capable were you before?” and “How capable are you now?” it asks, in effect, “Looking back, how capable were you then, compared with now?”
This design reduces response-shift bias because the student uses the same understanding of the competency for both ratings. It does not eliminate every problem, but it often gives a more believable picture of change in self-perception.
For example, a student may initially rate their professionalism as a 4. After a semester-long internship, they may rate their past professionalism as a 2 and their current professionalism as a 4. The retrospective pre-test reveals growth that a traditional pre test might miss if the original self-rating was inflated.
That said, retrospective pre-tests have a real weakness: they rely on memory. Students may misremember their earlier ability, especially if they want to show improvement or please a supervisor. The method is useful, but it is not neutral.
When to use which design
Use a traditional pre and post when the construct is stable
If you are measuring something students already understand well, and the experience is unlikely to change how they interpret the skill, a standard pre and post comparison is reasonable. Examples might include a narrowly defined software task or a procedural requirement with clear criteria.
Use retrospective pre when the experience is developmental
If the experience is designed to teach students what the competency actually looks like in practice, the retrospective pre-test is often the better self-assessment tool. This is especially true for broad skills such as communication, teamwork, professionalism, and critical thinking, where students commonly recalibrate their standard.
Use both when the stakes justify it
For high-value programs, pairing a traditional pre test with a retrospective pre can be informative. If the two versions tell different stories, that difference itself is useful. It may indicate that students entered with inflated confidence or that they learned a more sophisticated definition of the skill.
Build assessments that can handle disagreement
Self-assessment should not be the only evidence of growth. If you want a pre and post design that actually shows change, collect at least one external perspective.
Faculty supervisors, employer mentors, and site preceptors can each rate the same competencies from observation. Their ratings will not match student self-ratings exactly, and that is the point. A student may believe they communicate well because they are confident in meetings, while a supervisor notices missed deadlines or incomplete follow-up. That mismatch is not a failure of the assessment. It is a signal worth examining.
A useful design often includes three parts:
- Student self-assessment before the experience.
- Student self-assessment after the experience, ideally with a retrospective pre item.
- Supervisor or faculty ratings at the end of the experience.
When all three are collected on the same competency definitions, you can see whether growth appears in self-perception, observed behavior, or both. That matters for accreditation, because a program can show stronger evidence when it demonstrates alignment and disagreement rather than pretending all raters should match perfectly.
Keep the scale simple and the anchors specific
A pre and post assessment works better when students know what each point on the scale means. Vague anchors like “poor” and “excellent” invite inconsistent use. Behavior-based anchors are stronger.
For example:
- 1 = needs frequent direction, misses expectations often
- 3 = meets expectations with occasional guidance
- 5 = meets expectations independently and consistently
That kind of scale gives students and supervisors a shared reference. It also makes change easier to interpret. A shift from 2 to 3 means more than “feels better,” it suggests movement from frequent support to occasional support.
Avoid too many scale points unless your respondents can reliably distinguish them. A five-point scale is usually enough for student self-assessment. A nine-point scale may look more precise, but the extra detail is often not real.
Treat negative change as data, not failure
A post score that drops is not automatically bad news. It may mean the student learned the competency is harder than expected. It may also mean the experience was weak, the supervisor did not provide enough feedback, or the student never had a chance to practice the skill.
The question is not whether every score rises. The question is whether the assessment captures the kind of growth the experience was supposed to produce. If a student’s confidence falls while their supervisor rating rises, that may indicate better calibration, not regression.
That distinction is important when reporting results to faculty, employers, or accreditors. Growth is not always a straight line upward on a self-report survey. Sometimes the most meaningful development is that students become more accurate judges of their own performance.
Use the results to improve the experience
A strong pre and post design should tell you more than whether students changed. It should help you see where the experience worked and where it did not.
If student self-ratings rise but supervisor ratings stay flat, the experience may have improved confidence more than skill. If supervisor ratings rise but self-ratings fall, the program may be teaching students the complexity of the competency without enough time for them to internalize it. If both rise, you have evidence that the placement supported both learning and calibration.
That is why many programs pair self-assessment with 360-degree feedback, then benchmark the results across sections, sites, or terms. The value is not only in the numbers. It is in the pattern they reveal when multiple raters use the same competencies over time. A platform such as the Career Readiness Report can help organize that kind of evidence across the eight NACE competencies, but the core design principle is the same whether you use software or spreadsheets: measure the same construct from more than one angle, and interpret self-ratings with caution.
The Career Readiness Report is free for every college and university. Open now, in beta.
Create your institution