Article

Writing Behavioral Anchors So Two Raters Agree

August 29, 2026 ยท 5 min read

Writing Behavioral Anchors So Two Raters Agree
Photograph by Armin Rimoldi on Pexels.

Two raters will only agree on a student's competency score when the rating scale describes observable actions rather than internal attitudes or vague adjectives. Writing reliable behavioral anchors requires defining the exact actions a student takes, specifying the magnitude or frequency of those actions, and assigning a distinct, concrete scenario to every point on the assessment scale.

The limitation of adjective-based scales

Many assessment rubrics rely on adjectives to measure performance. Scales ranging from "Poor" to "Excellent," or "Rarely" to "Always," force supervisors to rely on their personal baselines. If a rubric asks an internship supervisor to rate a student's professionalism on a scale of one to five where five is "Excellent," the resulting data is highly subjective. To a supervisor in a rigid corporate environment, excellent professionalism might mean adhering to a strict dress code and hierarchical communication. To a supervisor at a startup, excellent professionalism might mean anticipating client needs and taking unprompted initiative.

When accreditation bodies or academic departments look at these scores, they cannot know what behavior earned the rating. A four out of five from a lenient rater often represents the exact same student behavior as a two out of five from a strict rater. Behaviorally Anchored Rating Scales eliminate this variance by replacing adjectives with descriptions of specific actions.

Translating abstract concepts into observable actions

To write a functional behavioral anchor, focus exclusively on actions a video camera could record. You cannot observe an internal state like "understands" or "appreciates." You can observe a student submitting a report, asking a clarifying question, or correcting a mistake after receiving feedback.

Consider the National Association of Colleges and Employers (NACE) competency of Teamwork. A traditional rubric might describe a competent student as one who "works well with others." This phrase is too vague for consistent rating. A behavioral anchor translates that abstract idea into concrete mechanics.

Weak anchor: Works well with others, communicates effectively, and is a good team player.

Strong anchor: Solicits ideas from team members, completes assigned share of collaborative work by the agreed deadline, and shares relevant project information with the group without prompting.

The strong anchor provides three distinct behavioral markers. If a student misses team deadlines or hoards information, the supervisor knows immediately that the student does not meet this level of the competency, regardless of the supervisor's personal definition of a "good team player."

Anchoring the entire rating scale

A single anchor defines one level of performance. To build a complete scale, you must write distinct anchors for every point, typically ranging from a beginner level to advanced mastery. The progression between points must represent a clear increase in complexity, autonomy, or scope of impact.

For the Teamwork competency, a five-point scale requires distinct behavioral profiles at each step.

Level 1 (Beginning): Fails to complete assigned collaborative tasks, ignores communication from peers, or dismisses others' ideas during group discussions.

Level 2 (Emerging): Completes assigned tasks only after repeated reminders, participates in group discussions when directly questioned, but rarely initiates communication with team members.

Level 3 (Competent): Solicits ideas from team members, completes assigned share of collaborative work by the agreed deadline, and shares relevant project information with the group without prompting.

Level 4 (Advanced): Identifies gaps in the team's workflow, volunteers to take on unassigned tasks, and actively encourages quieter members to share their perspectives.

Level 5 (Mastery): Mediates disagreements among team members to reach consensus, helps peers complete their tasks when they encounter roadblocks, and synthesizes multiple conflicting perspectives into a cohesive project plan.

The cost of precision and calibration

The primary tradeoff of using a Behaviorally Anchored Rating Scale is the intense labor required to create it. Drafting these scales involves gathering subject matter experts, defining each competency, writing the anchors, and testing them with real raters to ensure the language is neither too broad nor too narrow. If an anchor is too specific, such as requiring a student to "use Slack to update the team," raters in organizations that use Microsoft Teams will not know how to score the student. The language must be platform-agnostic but behaviorally specific.

This calibration process takes months of iteration. That barrier to entry is why the Career Readiness Report provides pre-calibrated behavioral anchors for the eight NACE competencies to collect 360-degree feedback from employers and faculty. Because the platform uses established scales, institutions can generate reliable benchmark data without building and testing their own rubrics. However, if your internship program evaluates specialized technical skills, such as statistical software proficiency for data science majors or clinical triage for nursing students, you must invest the time to draft and test custom behavioral anchors for those specific domains.

Managing edge cases in the field

Even with rigorously tested scales, supervisors will encounter edge cases. Human behavior is rarely perfectly consistent. A student might display Level 5 mastery by mediating a major team conflict but also display Level 2 behavior by needing a reminder to submit a weekly progress report. Raters often struggle when a student's actions split two different scale points.

To maintain high inter-rater reliability, provide supervisors with clear instructions on how to handle these discrepancies before the assessment begins. Instruct raters to score based on the student's consistent, majority behavior over the entire assessment period rather than isolated incidents or their most recent interaction. If a student consistently performs at a Level 3 but had one exceptional day at Level 5, the accurate score is Level 3.

Finally, ensure supervisors know that higher scores require the student to consistently demonstrate the behaviors in the lower scores. A student cannot achieve Level 5 mastery if they are failing to meet the basic deadline requirements outlined in Level 3. By providing concrete behavioral anchors and clear rules for applying them, you remove the guesswork from supervisor evaluations and generate assessment data your institution can actually trust.

The Career Readiness Report is free for every college and university. Open now, in beta.

Create your institution