MAKE QUALITY OBSERVABLE
Make quality
observable.
Use AI to draft criteria, examples, and feedback, then test whether those distinctions fit the actual student task.

Move from polished descriptors to defensible judgment
A rubric can look professional while leaving its expectations unclear. In the campus activity, students need to compare proposals, interpret evidence, explain a tradeoff, and acknowledge a limitation. The rubric should help learners see those requirements and help instructors explain judgments using observable features of the response.
Try the rubric on a real distinction.
These are fictional responses from the campus teaching scenario. The explanation is authored calibration guidance, not an automated grade.
One criterion: the recommendation uses relevant evidence accurately and recognizes what it cannot establish. Judge this criterion, not the policy choice.
| Emerging | Developing | Convincing |
|---|---|---|
| Relies on unsupported claims, or extends the evidence beyond what it can show. | Uses relevant evidence but leaves a key interpretation or limitation unclear. | Connects relevant evidence to the decision and explains important limits on the conclusion. |
I recommend cycling support because cycling is always the fastest and everyone can do it. The cycle trips are cheaper than car trips, so the college will save all of its transport costs. The repair station is clearly the best option, and we do not need any more information.
I would pilot the evening shuttle because the proposal explicitly targets evening learners, although the supplied dataset does not tell us how many students attend at that time. The four bus records have a mean journey time of 39.25 minutes, compared with 24.5 minutes for cars, but their distances and routes differ; these figures do not show that a shuttle will be faster. Cycling support also deserves consideration, but we lack evidence about seasonal uptake and access. My choice prioritizes a defined group rather than a proven improvement. Before implementation, I would ask intended users about schedules and routes; during the pilot I would record uptake and collect feedback about whether the service fits their needs.
Start from the assignment and outcome
Supply the task, approved material, audience, and intended evidence before requesting criteria. Ask for distinct dimensions and descriptions of performance that matter for this assignment. “Uses evidence effectively” needs an explanation: is the evidence relevant, accurately interpreted, and connected to the recommendation? Keep limitations visible within the criteria rather than treating caution as a weakness. Check for overlap that would count the same problem twice. An agent can suggest a structure and identify ambiguities; the instructor decides standards and weighting. A useful descriptor should give a student a reasoned next step, not merely a label such as excellent or weak.
Rehearse the rubric on contrasting work
Use the fictional responses in the calibration exercise. Rate a sample yourself before viewing an agent's proposed rating, then compare the cited evidence criterion by criterion. Check numerical claims against the supplied dataset. A disagreement may reveal an ambiguous boundary, an unstated expectation, or a mistaken interpretation. Revise that specific issue and apply the descriptor again. Two agents agreeing does not establish grading validity, and averaging ratings can hide the point that needs discussion. Keep the original rubric and the proposed revision so you can explain how the standard became clearer. Final assessment decisions remain with the instructor.
Connect the judgment to improvement
Feedback becomes usable when a learner can act on it. Identify an evidenced strength, a priority linked to a criterion, and a concrete revision the learner could attempt. For example, ask the student to explain why a transport mean supports—or cannot support—the recommendation instead of simply requesting more critical thinking. The same discipline helps with question banks: keep answers separate from student questions, explain the key, acknowledge defensible alternatives, and solve items independently before release. Treat distractors as possible indicators of reasoning rather than proof of a student's belief. The quality of the evidence matters more than the number of generated items.
PUT THE IDEA TO WORK