Article
Who Should Rate a Student, and How Many Raters Is Enough
One rater, chosen correctly, beats three chosen by convenience. The person who should rate a student is the one who directly observed and directed the student's work, week to week, for the duration of the placement: usually the line supervisor, not the site coordinator, not the HR contact who signed the affiliation agreement, and not the executive whose name appears on the placement letter. Two raters is the point where reliability gains justify the administrative cost. Three is defensible for high-stakes uses. Past three, you are trading measurable reliability gains of four or five points for double-digit losses in completion.
Observation opportunity is the only qualification that matters
The research on multi-source feedback is consistent on one point: rating quality tracks observation, not seniority. Conway and Huffcutt's meta-analysis of 360 data put mean inter-rater reliability for a single supervisor rating at roughly .50, for peers around .37, and for subordinates around .30. Those differences are largely explained by how much relevant behavior each source actually sees.
This is why the site coordinator is a poor default. At a hospital, an accounting firm, or a school district, the coordinator manages the pipeline: they place four students, sign the timesheets, and field your calls. They may have spoken with the student twice. When they complete a competency rating, they do one of three things: rate everything at the midpoint, rate everything high to protect the relationship, or reconstruct impressions from the line supervisor's offhand comments. All three degrade your data, and the third is the most dangerous because it looks like an independent rating and is not.
A workable eligibility test
Before you accept a rater, require an affirmative answer to two questions on the form itself:
- Did you directly observe this student's work at least weekly?
- Did you observe at least 100 hours of the student's work, or at least six weeks of it?
Those thresholds are conventions, not findings, but they do useful work. They give the coordinator a graceful way to redirect the request, and they let you flag ratings that came from insufficient exposure rather than silently averaging them in. In practice, roughly one in ten nominated raters will fail this test, and most of those are exactly the administrative contacts you did not want.
What the second rater buys
Apply the Spearman-Brown formula to a single-rater reliability of .40 and the sequence is instructive. One rater: .40. Two: .57. Three: .67. Four: .73. Five: .77. The second rater adds 17 points. The third adds 10. The fourth adds 6. The fifth adds 4.
But the reliability arithmetic understates what a second rater does, because the useful second rater usually comes from a different context. Employer supervisors see professionalism, teamwork, and equity and inclusion in practice, in situations faculty never witness. Faculty see written communication, critical thinking, and technology through artifacts the employer never reads. When you collect both, disagreement is not noise to be averaged away. A student rated 4.5 on communication by a site supervisor and 2.8 by the faculty sponsor who read their reflection memos is telling you something specific: they present well in person and write poorly. Two raters from the same office would never surface that.
A second rater also protects the record. Supervisors change jobs, go on leave, and ignore email. If a single required rater does not respond, that student contributes nothing. With two, partial data survives.
The response-rate math
This is the part that gets skipped. Supervisor evaluation response rates in internship programs typically land between 65 and 85 percent with a named contact, a deadline, and two reminders. Call it 78 percent as a working figure.
If you require complete data from two raters and each responds independently at 78 percent, your rate of fully complete student records is .78 x .78, or 61 percent. Three raters: 47 percent. Four: 37 percent. A program with 200 placements that requires four raters per student will finish the term with roughly 74 complete records, and those 74 will not be a random sample. They will overrepresent well-organized host sites, longer placements, and students who nagged effectively. That selection bias will do more damage to your program-level conclusions than the reliability you gained.
The fix is to distinguish required from optional. Require the direct supervisor. Invite the faculty sponsor or a second workplace observer, and treat their response as a bonus that strengthens the subset of records that has it. You can then report the full sample on supervisor ratings and a smaller high-confidence subset where two independent sources agree. Platforms that collect this, including the Career Readiness Report, will show you both cuts, and the comparison between them is itself worth reading: if the two-rater subset looks systematically stronger than the full sample, your optional raters are being recruited selectively.
Do not count self-assessment as a rater
Self-ratings correlate weakly with supervisor ratings. Harris and Schaubroeck found a mean corrected correlation of about .35 between self and supervisor, against .62 between peer and supervisor. Averaging a self-rating into an aggregate score does not improve reliability; it introduces a different construct. The self-rating is valuable as a gap measure, particularly because the students who overestimate themselves most are often the ones with the least workplace exposure, and that gap is coachable. Keep it separate in your reporting.
A default configuration
For most programs: one required rater (the direct supervisor, named in the learning agreement at the start of the placement, not identified in week fourteen), one invited second rater from a different context, and a self-assessment reported separately. Send the request within five days of the placement's end date, while memory is fresh and the supervisor still cares.
The honest limit: a two-rater design gives you per-student reliability around .55 to .60. That is adequate for program-level reporting, where you have hundreds of students and standard errors shrink accordingly. It is not adequate to deny a student a credential or assign a large portion of a grade. If a competency score carries individual consequences, you need three raters and you need to accept the completion penalty, or you need to use the rating as one input among several rather than as the decision itself.
The Career Readiness Report is free for every college and university. Open now, in beta.
Create your institution