Robin. How should readers judge the progress, value, and lines of responsibility around AI assessment and feedback?. 2026-09-21.
AI assessment and feedback should track tool-assisted performance, process disclosure, and tool-free transfer together. An iFlytek grading device reports class-level processing within minutes and more than 4,000 error labels, while a Zhuji speaking-and-listening project reports regional reach and stage gains. Both require clear vendor or organizer attribution and independent validation. The return to oral examinations adds a process-authentication option whose validity, workload, and accessibility need review.
https://edu.hhhh.life/en/guide/ai-assessment-feedback/#answer
20direct citations
19source organizations
2026-09-18evidence through
READ THIS FIRST
Three judgments to remember
01
Classroom products now connect personalized feedback and writing-process support to assignments.
02
Regulation and teacher guidance require validation and human responsibility in high-stakes evaluation.
03
Reasoning, mastery, sustained engagement, and independent performance require different measures.
CURRENT ANSWER
How we answer today
Each judgment links to the relevant news and original sources. New evidence enters the corresponding dimension.
01
Formative feedback works best as a draft for teacher confirmation
Google Classroom and Writing Coach support teacher feedback workflows, while an iFlytek device connects grading, error labels, and learning analysis to assignments. Workflows should preserve teacher edits, student revisions, and final rationale, and independently test vendor claims about speed and accuracy. The return to oral examinations adds a process-authentication option whose validity, workload, and accessibility need review.
Exams and high-stakes decisions need a higher evidence threshold
UK assessment regulation restricts independent AI scoring, Ireland's advisory group prioritizes exams and evaluation, and outcome tools are beginning to measure reasoning and mastery. High-stakes use needs validation samples, human review, bias checks, and appeals. A Brazilian writing platform, literacy-assessment concerns, curriculum-text risks, a skills-validation platform, and New South Wales rules add different assessment contexts. High-stakes decisions still require independent validation, process evidence, and human judgment.
Feedback speed and genuine learning need separate measures
Tutor CoPilot offers evidence on teacher support and outcomes, Khanmigo's learning-history experiment observes next-question performance, and StudyFetch evaluates responsible prompting. These results cover teachers, immediate tasks, and process behavior rather than one durable effect.
These limits determine how strong a conclusion the page can support.
01
Product features and company cases show workflow design, while independent evidence on accuracy, bias, and long-term outcomes is often sparse.
02
Assessment criteria vary by subject, age, and task, so one model-accuracy figure cannot summarize educational quality. Needed evidence includes blind-rating agreement, teacher edit rates, revision trails, appeal outcomes, and delayed tests.
RELATED QUESTIONS
What else do readers ask?
Each adjacent search question receives a concise answer linked to its supporting evidence.
01
What can currently be confirmed about AI assessment and feedback?
+
Current public evidence can establish policy, curriculum, program, or product progress. Reach and launch figures should retain their own definitions and remain separate from sustained use and learning outcomes.
Does the available material establish learning outcomes?
+
The available material mainly supports policy, implementation, product, or participation progress. Learning effects still require independent tasks, delayed measures, subgroup results, and reproducible methods.
eSchool News reported on September 18, 2026 that Catherine Hartmann, a professor at the University of Wyoming, redesigned her upper-level humanities course around oral assessment after a 2023 assignment in a history of meditation course came back with a copy-pasted AI prompt left at the top. Students practiced discussion and oral explanation throughout the semester, and the final oral exam became the culmination of that work. At the University of Pennsylvania, mathematics professor Robin Pemantle has students work calculus problems at the board while explaining their reasoning aloud.
On 9 September 2026, UNESCO awarded its 2026 ICT in Education Prize during Digital Learning Week in Paris to Finland's Generation AI and Brazil's Redação Paraná. Generation AI, developed by the University of Eastern Finland, the University of Helsinki, the University of Oulu and other partners, provides classroom activities and tools to help students understand and critically assess AI systems; its materials have been accessed more than 200,000 times across over 50 countries, and workshops and school interventions have involved 1,500 students and teachers. Redação Paraná, developed by the State Secretariat of Education of Paraná in Brazil, combines AI-assisted feedback with teacher involvement to support writing and revision, reaching approximately 850,000 students across 2,000 schools, with teachers trained to integrate its materials. The two projects were selected from more than 100 nominations, and the 2026 prize focused on approaches that use AI while encouraging learners to retain critical thinking, creativity and independent judgement.
A bipartisan group of 35 New Mexico state lawmakers sent a letter on Aug. 25 to Public Education Department Secretary Mariana Padilla, calling for greater transparency and additional independent testing to evaluate how accurately Amira, an AI literacy testing tool, measures students' reading proficiency. The tool has been required statewide for K-2 students since the 2025-26 school year, with students reading aloud to a digital avatar. Lawmakers said Amira has collected thousands of student voice recordings and has access to names, genders, birthdays, locations and other sensitive information, and asked that parental consent be required before children use the tool.
Education Week reports that AI-generated text has entered elementary school classrooms alongside classroom library books and early reading curricula. Tools for educators can rewrite articles to different reading levels or generate decodable text. Jean Gunderson, a Title I reading interventionist in South Dakota, said AI can write stories and comprehension questions in minutes, a task that used to take hours, but she must prompt precisely and sort through all output. Researchers flagged three risks: AI's middle-ground tone, tools struggling to hit requested grade levels, and weak alignment with academic standards and curricula.
Coursera previewed Project Helix, an AI-native skills platform, at its annual FWD customer event on September 8, 2026. The platform aims to help organizations address talent gaps and verify workforce capabilities, drawing on over 30,000 global content partners and instructors from Coursera and Udemy. It will generate adaptive learning paths based on business goals expressed in natural language, with broad availability to enterprise customers expected in the first half of 2027.
Recently, the practice of AI-empowered English listening and speaking teaching in Zhuji, submitted by the Zhuji Education Research Center, was selected for the public list of '2026 Smart Education Excellent Cases'. The case uses AI listening and speaking classrooms, now in regular use in 16 schools, covering 92 classes, 58 English teachers, and over 4,000 students. The project adopts a mechanism of pilot verification, scale-up, and dynamic optimization, and has established a tiered training system.
The University of Glasgow's School of Education has released a free 92-page toolkit to help teachers make deliberate decisions about digital and AI technology in learning. Developed by Mark Peart and colleagues, it guides teachers through examining assumptions, lesson design, and ethical checks, without recommending specific tools.
A study tracking about 27,000 students aged 12-18 in China for 30 months found that after adopting generative AI, homework scores rose by 18% while time per assignment fell from 64 to 45 minutes. However, in monthly closed-book exams without AI, scores dropped by 20% within six months, and high-stakes entrance exam performance also declined. The research was conducted by scholars from Stockholm University and the University of Hong Kong.
The Computing Research Association's Education Committee has issued a white paper calling on universities to rethink how computer science students are taught and assessed in the age of generative AI. It proposes four principles, including treating learning as a process, and suggests alternatives like oral exams and code walkthroughs. The paper cites a University of Illinois facility that proctors over 90,000 exams annually.
On September 1, 2026, iFlytek released the Spark Intelligent Grading Machine M50, which can grade and analyze homework for an entire class within minutes, forming an hour-level personalized teaching loop. Its error cause system includes over 4,000 labels and passed expert appraisal. The device supports 50g low-weight paper with a jam rate of 0.4‰.
The New South Wales government released new rules on 1 September 2026 limiting schools to a maximum of one take-home assessment task worth no more than 15 per cent of the school-based assessment mark, or 7.5 per cent of the total HSC mark. The advice from the NSW Education Standards Authority applies to the Class of 2027 beginning HSC studies in Term 4 this year and to students starting Year 11 in Term 1, 2027. HSC major works and some courses including creative arts, technologies and English Extension 2 are exempt, but schools must still authenticate students' work.
On July 16, Ofqual updated its approach to regulating AI in qualifications, continuing to prohibit AI as the sole scorer while allowing validated supporting and quality-assurance uses.
Khan Academy summarized about 20 Khanmigo product experiments conducted between October 2025 and April 2026. The organization says adding recent practice history and prerequisite skills not yet mastered to responses produced a combined 6.1% increase in the rate of independently answering the next question correctly.
Khan Academy official blogEducators / Product Teams
On May 14, StudyFetch launched Active AI Literacy, scoring each student prompt in real time on quality and responsibility and offering improvement suggestions before it is sent.
On April 7, the Irish government announced an external advisory working group on AI in schools as an ongoing, multi-stakeholder governance mechanism to study AI's effects on teaching, learning, and assessment.
On March 4, OpenAI announced a set of tools for measuring learning outcomes, disclosed early research on Study Mode, and said it planned to continue validation through randomized trials.
On February 19, Google launched AI-suggested feedback in Classroom. Gemini can draft personalized written guidance using a student's assignment, grade level, and focus areas specified by the teacher.
On January 21, Google and Khan Academy announced a partnership to enhance Khan Academy's Writing Coach with Gemini models. The product is focused on guidance and feedback during the writing process.
In December 2025, the Expert Steering Committee for Teacher Workforce Development under China's Ministry of Education released the Guidelines for Teachers' Use of Generative Artificial Intelligence (Version 1), covering learning, teaching, student development, evaluation, administration, and research.
Guidelines for Teachers' Use of Generative Artificial IntelligenceSchools / Educators
In live K–12 mathematics tutoring, Tutor CoPilot suggests guiding questions, hints, and conceptual scaffolds to human tutors. A Stanford research summary reports that the randomized trial involved more than 700 tutors and more than 1,000 students.
Stanford SCALEEducators / Schools
KEEP READING
Continue reading
Enter through an adjacent search question or return to a long-term topic for its full evidence base.