Every evaluation scores 7 dimensions, combined with public baseline weights: Role Assignment 15%, Context 20%, Specific Objective 20%, Constraints 15%, Output Format 10%, Audience 10%, Token Efficiency 10%. Weights can vary by scenario (a contract weighs context more); when they do, the scenario declares it.
Evaluation is deterministic (same prompt, same score), feedback is dimension-specific, and the whole method is formalized in the versioned Grace Syllabus.
In problem-solving cases you write your solution rather than a prompt, so the 7 dimensions don't apply and the judgment rests on 4 criteria: it addresses the core problem rather than the symptoms, it is feasible within the case constraints, it weighs risks and consequences, and it avoids the typical mistakes of that case. The verdict is solved or not solved with a 0-100 score, and even an approach different from ours passes if the reasoning holds up.
From our network
AGORÀ Intelligence: Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.