RobinAI Education Radar

Evaluation Tool · Updated 2026-09-21

How can a shared rubric compare education AI tools?

Written and maintained by Robin · Updated

This answer synthesizes public sources. Check the evidence and limits before applying it. Review the evidence

Link to this answer
Citation text
Robin. How can a shared rubric compare education AI tools?. 2026-09-21.
Compare tools for the same learner group, subject, and task. Record learning fit, age fit, output quality, teacher control, data safety, accessibility, research evidence, full cost, and exit readiness. We recommend “verified / evidence needed / unmet” for each item, with source documents and trial examples. Defer uses with unmet safety or human-responsibility conditions and assess each dimension separately.
https://edu.hhhh.life/en/guide/ai-tool-evaluation-rubric/#answer
15direct citations
15source organizations
2026-09-14evidence through

READ THIS FIRST

Three judgments to remember

  1. 01

    Start with learning fit and observable value

  2. 02

    Set separate gates for safety, data, and human control

  3. 03

    Evidence, cost, and exit determine durable usability

ACTION PLAN

What to do

Work through five steps in order, preserving process records and the linked evidence. Return to an earlier step when conditions change.

  1. 01

    Define the problem and boundary

    State the task, age, subject, users, and required human decisions and screen out misaligned products first. Sixteen independent tool reviews add test cases for instructional fit, error handling, and general-chatbot risk.

    View 5 sources for this step
  2. 02

    Record the baseline and owners

    Create nine records marked verified, evidence needed, or unmet; agree data, human-control, and exit stop conditions in advance.

    View 5 sources for this step
  3. 03

    Run a bounded practice

    Test candidates on the same course material and tasks, recording errors, human edits, time, and accessibility.

    View 6 sources for this step
  4. 04

    Check outcomes and costs

    Check original studies, population, sample, comparison, independent work, and delayed outcomes; name the study type and omissions.

    View 6 sources for this step
  5. 05

    Expand, adjust, or exit

    Document item-level results, unresolved questions, approved pilot scope, and exit requirements, with an expiry date.

    View 3 sources for this step

REVIEW CHECKLIST

Check before proceeding

  1. Is evaluate an education AI tool tied to an observable task, covered population, and prohibited-use boundary?

    EU and OECD Release a Primary and Secondary AI Literacy Framework Defining 19 Competencies · Teacher Generative AI Guidelines State That AI Grading Cannot Directly Serve as the Final Evaluation of Open-Ended Work · OECD Warns That Better AI-Assisted Task Performance Does Not Necessarily Mean Real Learning · iFlytek Launches Spark Intelligent Grading Machine M50 with Error Diagnosis · Instruction Partners Releases 16 Deep Dives on AI Learning Tools, Warns of General Chatbot Risks
  2. Are owners, human review, data handling, incident reporting, and appeals explicit?

    UK Adds Children's Mental Health and Manipulation Risks to Safety Standards for Educational AI Products · UK Expands Generative AI Data Guidance for Schools, Clarifying Personal Information, Bias, and Supplier Risks · World Digital Education Alliance Releases Two Standards Covering the Full Educational AI Life Cycle and Smart Campuses · 35 New Mexico Lawmakers Urge Education Department to Add Independent Testing and Parental Consent for Amira AI Literacy Tool · Philippine Developer Launches Hiraia, an Offline AI Science Tutor That Runs on $100 Android Phones · Oxford University Press Expands AI Study Assistant to Business, Politics and Science Trove
  3. Do results include independent performance, sustained use, workload, safety events, and group differences?

    OpenAI Releases Learning-Outcome Measurement Tools, Shifting Evaluation Toward Reasoning and Mastery · EduClaw-Bench Places AI Tutors in a Continuous 30-Day Simulated Learning Relationship · Acumen Acquires EduCorePro, Applying AI to University Admissions Review and Fraud Prevention
  4. Do continue, adjust, pause, and exit decisions each have a threshold, date, and owner?

    EU and OECD Release a Primary and Secondary AI Literacy Framework Defining 19 Competencies · OpenAI Releases Learning-Outcome Measurement Tools, Shifting Evaluation Toward Reasoning and Mastery

CURRENT ANSWER

How we answer today

Each judgment links to the relevant news and original sources. New evidence enters the corresponding dimension.

01

Start with learning fit and observable value

Curriculum frameworks, teacher guidance, and evidence reviews connect features to learning goals and independent performance; grading-device speed and error labels need common-task testing. Sixteen independent tool reviews add test cases for instructional fit, error handling, and general-chatbot risk.

View 5 direct sources
02

Set separate gates for safety, data, and human control

Child-product standards, school data guidance, and lifecycle standards cover admission, monitoring, and accountability. New cases let the rubric compare three conditions: independent validity for high-risk assessment, data minimization in an offline tool, and curriculum grounding in an in-textbook assistant.

View 6 direct sources
03

Evidence, cost, and exit determine durable usability

Outcome tools, longitudinal benchmarks, and product change require method review, durable performance, and migration readiness.

View 3 direct sources
04

Copyable nine-item record and decision rule

Candidate and edition: ____; learner group / subject / common task: ____; review date: ____. Complete 1 learning fit, 2 age and account terms, 3 output accuracy, 4 teacher control, 5 data and safety, 6 accessibility, 7 research evidence, 8 full cost, and 9 exit readiness. For each: “status: ____; document or trial example: ____; gap: ____; reviewer: ____”. Resolve stop items before comparing eligible tools on quality and effort. Mark missing records as “evidence needed”. This editorial rubric has no research-validated weights or thresholds; the school should agree them before testing.

View 3 direct sources

EVIDENCE BOUNDARY

Limits to keep in mind

These limits determine how strong a conclusion the page can support.

  1. 01

    These steps are editorial recommendations informed by public sources. The full workflow has not been validated as an intervention; adapt it to local curricula, age, and resources.

  2. 02

    Product features, coverage, and participation establish an implementation entry; learning effects require independent tasks, delayed measures, and disaggregated results.

RELATED QUESTIONS

What else do readers ask?

Each adjacent search question receives a concise answer linked to its supporting evidence.

01

How can reviewers avoid scoring only the demo?

Give candidates the same course material, learner description, and task. Save original outputs, errors, teacher edits, and time. Add independent no-AI work for learning tasks and test accessibility on the school’s devices and support needs. We suggest retaining failures; a supplier demo does not replace local verification.

View 2 sources for this answer
02

What should be checked in a claim that a tool improves attainment?

Record participants, grade and subject, sample, comparison group, task, duration, independent work, delayed measurement, and funding. Identify whether the outcome is the assisted current task or later independent work. Label randomized trials, observational studies, product self-tests, and simulations separately, preserving limits where the population differs from the school.

View 2 sources for this answer
03

How should a high-scoring tool with an unmet essential condition be handled?

Document the unmet condition and affected use, then request evidence or remediation. Defer the relevant workflow if data use or human review remains unresolved. Continue recording other dimensions and limit approval to verified conditions. This editorial method requires the school to agree stop items and responsible staff in advance.

View 2 sources for this answer

EVIDENCE INDEX

Evidence index

Sorted by public date, preserving only verifiable records and original sources.

View 15 related records
ResearchGlobalOriginal publication

Instruction Partners Releases 16 Deep Dives on AI Learning Tools, Warns of General Chatbot Risks

Education consulting nonprofit Instruction Partners released a large-scale evaluation of AI-powered learning tools, including 16 deep dives into individual products, covering 20 tools and 16 school systems, with interviews of teachers, students, district and building leaders, and product developers. The analysis found the biggest risks from general-purpose chatbots, which students may use to avoid effortful thinking, while purpose-built instructional tools showed more promise but none was ready to do the pedagogical job independently.

Education WeekSchools / Educators
AI TutoringGlobalOriginal publication

Philippine Developer Launches Hiraia, an Offline AI Science Tutor That Runs on $100 Android Phones

Luis Buenaventura, a member of the Blockchain Council of the Philippines, has developed Hiraia, an open-source AI science tutor that runs fully offline on entry-level Android phones costing around $100, supporting Tagalog, Bisaya, and English. Built on Sea AI Lab's Sailor 2 model and a continued-pretraining fork of Qwen 3.5-2B, the app includes over 40,000 science facts and 30,000 illustrations aligned with the Department of Education's MATATAG curriculum. The project received a 1 million peso research grant from the Tether Foundation and is currently in early alpha (v0.3.1), with the developer seeking academic partners for classroom pilots.

BitPinasStudents / Educators
Policy & GovernanceGlobalOriginal publication

35 New Mexico Lawmakers Urge Education Department to Add Independent Testing and Parental Consent for Amira AI Literacy Tool

A bipartisan group of 35 New Mexico state lawmakers sent a letter on Aug. 25 to Public Education Department Secretary Mariana Padilla, calling for greater transparency and additional independent testing to evaluate how accurately Amira, an AI literacy testing tool, measures students' reading proficiency. The tool has been required statewide for K-2 students since the 2025-26 school year, with students reading aloud to a digital avatar. Lawmakers said Amira has collected thousands of student voice recordings and has access to names, genders, birthdays, locations and other sensitive information, and asked that parental consent be required before children use the tool.

The Taos NewsFamilies / Educators
Learning ToolsGlobalOriginal publication

Oxford University Press Expands AI Study Assistant to Business, Politics and Science Trove

Oxford University Press announced on 9 September 2026 that its AI Study Assistant has expanded to Business, Politics and Science Trove, following its October 2025 launch on Law Trove, so all Trove platforms now offer the tool. It generates summaries, answers and targeted quizzes drawn only from textbook content, and the Science Trove version includes textbook images, figures and their original captions. An AI literacy module covering generative AI fundamentals, academic integrity, critical thinking and prompting is available to all Trove users.

Oxford University PressStudents / Educators
Teacher ToolsChinaOriginal publication

iFlytek Launches Spark Intelligent Grading Machine M50 with Error Diagnosis

On September 1, 2026, iFlytek released the Spark Intelligent Grading Machine M50, which can grade and analyze homework for an entire class within minutes, forming an hour-level personalized teaching loop. Its error cause system includes over 4,000 labels and passed expert appraisal. The device supports 50g low-weight paper with a jam rate of 0.4‰.

iFlytek EducationEducators / Schools
Policy & GovernanceChinaPolicy publication

Teacher Generative AI Guidelines State That AI Grading Cannot Directly Serve as the Final Evaluation of Open-Ended Work

In December 2025, the Expert Steering Committee for Teacher Workforce Development under China's Ministry of Education released the Guidelines for Teachers' Use of Generative Artificial Intelligence (Version 1), covering learning, teaching, student development, evaluation, administration, and research.

Guidelines for Teachers' Use of Generative Artificial IntelligenceSchools / Educators