QuantumLearning Machines
← LabPath

REVIEWER-RADAR INTEGRATIONS

Community benchmarks, gated before claims.

LessonBench, AIriskEval-edu, and L2-Bench are integrated as separate Suite B benchmark surfaces. Raw items stay out of public APIs and QLM QRS tiers.

No raw export

External dataset items are never re-exported through LabPath APIs.

No QRS mixing

Community suites remain separate from QLM public/evaluation/held-out tiers.

No score leakage

Scores stay hidden until license, schema, reproducibility, and publish-gate review clear.

lessonbench-v1

LessonBench-V1

arXiv paper

LessonBench-V1 is integrated as a gated community benchmark; execution waits for license/schema clearance.

3-judge ensemble with position swapping, self-family controls, and TeacherOS Copilot as ninth model once license clears.

license: pendingschema: pendingblocked license pending

647 expected items

l2-bench

L2-Bench

arXiv paper

L2-Bench is integrated as a gated community benchmark; execution waits for license/schema clearance.

Phase 1 crosswalk: competency transfer status, TeachProof scenario fit, and future probe design before scored execution.

license: pendingschema: pendingblocked license pending

1,000 expected items

airiskeval-edu

AIriskEval-edu

arXiv paper

AIriskEval-edu is integrated as a gated community benchmark; execution waits for license/schema clearance.

Direct ground-truth scoring for five risk dimensions plus localized-span overlap on the localization subset once license clears.

license: pendingschema: pendingblocked license pending

1,639 expected items

Suite B buildout

LessonBench-V1

LessonBench generation

Objective + metadata becomes a lesson-generation prompt; field models and TeacherOS Copilot are judged by a 3-judge ensemble.

position swappingself-family ruleTeacherOS Copilot as ninth modelpublic scores hidden until publish gate

AIriskEval-edu

AIriskEval classification

K-12 explanations are classified across five pedagogical risk dimensions and scored against dataset annotations.

direct label scoringSafeTutors crosswalklocalization subset separated

AIriskEval-edu

AIriskEval localization

Risk-localized spans are evaluated for dimension match and span overlap.

span overlapmissed-dimension reportno raw-span export

L2-Bench

L2-Bench Phase 1

Practitioner competencies are crosswalked into direct, STEM/ELL-adapted, L2-native, or out-of-scope categories before any scored probe.

taxonomy before probesTeachProof scenario fitno forced math-tutor scoring

Artifacts

SafeTutors ↔ AIriskEval crosswalk

5 community risk dimensions mapped to SafeTutors dimensions and sub-risks.

ready
L2-Bench competency crosswalk

12 competencies triaged before probe authoring.

phase-1-ready
LessonBench judged-generation policy

3 judges, position swapping, self-family rules, and TeacherOS Copilot comparison are defined.

ready-after-license