LessonBench-V1 is integrated as a gated community benchmark; execution waits for license/schema clearance.
3-judge ensemble with position swapping, self-family controls, and TeacherOS Copilot as ninth model once license clears.
647 expected items
REVIEWER-RADAR INTEGRATIONS
LessonBench, AIriskEval-edu, and L2-Bench are integrated as separate Suite B benchmark surfaces. Raw items stay out of public APIs and QLM QRS tiers.
External dataset items are never re-exported through LabPath APIs.
Community suites remain separate from QLM public/evaluation/held-out tiers.
Scores stay hidden until license, schema, reproducibility, and publish-gate review clear.
LessonBench-V1 is integrated as a gated community benchmark; execution waits for license/schema clearance.
3-judge ensemble with position swapping, self-family controls, and TeacherOS Copilot as ninth model once license clears.
647 expected items
L2-Bench is integrated as a gated community benchmark; execution waits for license/schema clearance.
Phase 1 crosswalk: competency transfer status, TeachProof scenario fit, and future probe design before scored execution.
1,000 expected items
AIriskEval-edu is integrated as a gated community benchmark; execution waits for license/schema clearance.
Direct ground-truth scoring for five risk dimensions plus localized-span overlap on the localization subset once license clears.
1,639 expected items
LessonBench-V1
Objective + metadata becomes a lesson-generation prompt; field models and TeacherOS Copilot are judged by a 3-judge ensemble.
AIriskEval-edu
K-12 explanations are classified across five pedagogical risk dimensions and scored against dataset annotations.
AIriskEval-edu
Risk-localized spans are evaluated for dimension match and span overlap.
L2-Bench
Practitioner competencies are crosswalked into direct, STEM/ELL-adapted, L2-native, or out-of-scope categories before any scored probe.
5 community risk dimensions mapped to SafeTutors dimensions and sub-risks.
12 competencies triaged before probe authoring.
3 judges, position swapping, self-family rules, and TeacherOS Copilot comparison are defined.