Article

What to Do About Supervisors Who Rate Everyone Excellent

July 29, 2026 ยท 6 min read

What to Do About Supervisors Who Rate Everyone Excellent
Photograph by Werner Pfennig on Pexels.

Stop trying to make supervisors rate more harshly. It does not work, and the harshness is not what you need anyway. What you need is variance, and you can engineer variance into a rating form even when every score lands in the top two points of the scale: use behavioral anchors instead of adjectives, force differentiation across competencies within each student, and report distributions rather than means. The supervisor who rates everyone excellent can still tell you that this student's communication is the strongest thing about her and her critical thinking is the weakest, and that comparison is more useful than any absolute score she would have given you.

Leniency and range restriction are not the same problem

Leniency is a shifted mean. Every rating is a point or two higher than the behavior warrants. If the shift is roughly constant across students, leniency costs you almost nothing analytically. You lose the ability to say "our students average 4.2 out of 5 on teamwork" and have that number mean anything absolute, but you can still rank, compare, and track change over time.

Range restriction is a compressed variance, and it is the expensive one. When 78 percent of a competency's ratings sit at 4 or 5 on a five-point scale and the standard deviation is 0.41, nearly every downstream use fails. You cannot detect a difference between two majors. You cannot see growth from midterm to final. Correlations with other measures get attenuated toward zero, which makes your data look like it has no relationship to anything, when the real problem is that the ruler has three marks on it.

Diagnose which one you have before you change anything. Pull the percentage of ratings at the top scale point, per competency, and the standard deviation within each supervisor who rated more than one student. A supervisor whose ratings average 4.6 with an SD of 0.7 is lenient and still discriminating. A supervisor whose ratings average 4.6 with an SD of 0.1 is giving you a signature, not an assessment.

Why supervisors compress

Some of it is real. Internship supervisors are evaluating a population that has already been screened twice, once by the employer's hiring process and once by your program's eligibility requirements. Genuinely poor performers are rare in that pool, and a distribution that leans high is partly accurate. Do not design your instrument on the assumption that the ceiling is entirely artifact.

The rest is structural. The supervisor has an ongoing relationship with the student and often wants to keep them. The rating has ambiguous consequences: nobody has told her whether a 3 costs the student a grade, a reference, or a job. She is rating eight competencies she has never seen defined, using labels like "proficient" and "exceeds expectations" that mean whatever she decides they mean. And halo is doing heavy work: when a supervisor likes a student, correlations among the eight competency ratings frequently run above 0.70, which means she is effectively answering one question eight times.

Behavioral anchors, written for the rater's actual vocabulary

Replace scale labels with observable behavior at each level. Not "3 = meets expectations" but "3 = writes clearly for a familiar audience; needs a second draft or supervisor edits before external circulation." The anchor gives the supervisor permission to score honestly, because she is not calling the student mediocre, she is describing what happened.

Two honest limits. Anchors take real work to write, roughly a full afternoon per competency with two or three employer partners in the room, and they must be rewritten for meaningfully different work contexts, because "presents to external stakeholders" describes a nonprofit development intern and not a lab intern. And the research on behaviorally anchored scales shows modest gains, not transformation. Expect the ceiling rate to drop by ten or fifteen points, not to disappear.

Force differentiation within the student, not across students

Do not impose a curve on supervisors. Telling an employer partner that only 20 percent of her interns may be rated excellent invites her to stop responding to your survey, and she is right to be annoyed, because she may have had three genuinely excellent interns.

Force differentiation the other direction. Ask each supervisor, after the competency ratings, to identify the two competencies where this specific student is strongest and the one where developing further would most improve her effectiveness. This is a within-person, within-rater comparison, so the rater's general generosity cancels out. It survives a complete ceiling: even if all eight ratings are 5, the rank ordering carries information about the student's profile, and it aggregates cleanly. If 41 percent of your business interns are flagged on technology and 9 percent on teamwork, you have a curricular finding that no mean score would have surfaced.

The cost is that ranking annoys some raters and produces artificial spread when the student really is even across competencies. Keep an explicit "no meaningful difference" option and count how often it gets used. If it is chosen 60 percent of the time, believe it.

Report distributions, not means

A mean of 4.3 tells your dean nothing and tells your accreditor less. Report the percentage of ratings at each anchor level, per competency, with the n visible. "Sixty-one percent of supervisors rated our students at the top level on teamwork; 34 percent at the second level; 5 percent below" is defensible. It also makes the ceiling visible instead of hiding it inside an average, which matters when a reviewer asks whether your instrument discriminates.

For comparisons across cohorts or programs, use within-rater centered scores: subtract each supervisor's own mean from her ratings before aggregating. This is the single highest-value analytic step available for 360-degree data, and it costs one column in a spreadsheet. It removes the leniency shift entirely. The tradeoff is that you lose absolute level, so you can no longer say students improved in an absolute sense, only that their relative profile changed. Report both, and label which is which.

Self-assessment paired with supervisor ratings gives you a second lever here, because the gap between the two is informative even when the supervisor rating is pinned at the ceiling. Students who rate themselves above a lenient supervisor are unusual and worth a conversation. The Career Readiness Report structures both sides on the same eight NACE competencies for exactly that comparison, and reports the distribution rather than the average.

Three things to stop doing

Do not drop lenient raters from your dataset. You will systematically remove the most engaged employer partners and the sites where students had the best experiences.

Do not statistically adjust individual student scores upward or downward based on their supervisor's tendency, then report the adjusted number to the student. The adjustment is defensible in aggregate and indefensible at the individual level, where your n per supervisor is usually one.

Do not benchmark programs against each other on means when the pooled standard deviation is below 0.5. The difference you are seeing between a 4.31 and a 4.19 is who happened to supervise whom.

The Career Readiness Report is free for every college and university. Open now, in beta.

Create your institution