ARIA Research InitiativeBack to research status

Research note 01 · Updated July 2026

How ARIA plans to earn its claims.

ARIA separates software checks, educator judgment, real-language validation, student feasibility, and learning outcomes. A result advances only the claim it actually tests.

Current statusResearch infrastructure readyIndependent human evidence and student outcomes remain pending

The research question

Can ARIA provide grounded help without taking over the work?

The program tests several different questions: whether tasks are correct, whether responses use the student’s actual reasoning, whether language labels match independent human judgments, and eventually whether students learn and transfer strategies.

Evidence sequence

Five gates, from a working system to a learning claim.

  1. 01

    Validate the task bank.

    Two subject-qualified educators independently review answer models, solution paths, misconceptions, hints, and scoring criteria. Disagreements stay visible.

  2. 02

    Blindly compare response quality.

    Educators rate five paired conditions for problem grounding, student grounding, usefulness, actionability, ownership, answer leakage, and invented student actions.

  3. 03

    Validate language on real student messages.

    Two trained annotators label observable reasoning moves with exact evidence spans. Evaluation is split by complete student or session, never random messages.

  4. 04

    Run a reviewed feasibility pilot.

    With institutional review, parent permission, and student assent, test whether the tool is understandable, usable, and safe before asking whether it is effective.

  5. 05

    Measure learning against an active control.

    Use independent outcome tasks, concealed assignment, blinded scoring, intention-to- treat analysis, and a delayed no-ARIA transfer measure.

Current evidence ledger

100

Structured task drafts

Schema-valid and ready for independent educator review.

13

Observable reasoning moves

Each automatic label keeps the exact supporting words visible.

0

Completed outcome studies

No classroom learning or ADHD-specific effectiveness claim is being made.

Evidence layerCurrent statusWhat remainsClaim allowed today
Task models100 schema-checked draftsEducator correctness review is pending.The bank is structured and ready for review.
Observable language13 reasoning-move codesIndependent annotation on real student language is pending.ARIA can show the exact phrase behind a tentative label.
Response quality5 blinded conditions designedTwo qualified educators must rate the locked responses.A fair comparison can be run; no winner is claimed yet.
Synthetic stress test84.6% same-style accuracyThe balanced score is 0.837; the cross-generator gap is 9.05 points.Useful for debugging only.
Student outcomes0 completed controlled studiesFeasibility, learning, retention, and transfer remain untested with students.No effectiveness claim.

Reading the language

Four terms that prevent inflated claims.

Development benchmark
A test used to find software weaknesses. Synthetic benchmark scores are not evidence that students learn more.
Human ground truth
Labels created independently by trained people using a locked codebook, without seeing ARIA’s prediction.
Active control
A comparison tool with the same tasks, interface, time, and base model, but without ARIA’s learner-conditioned pipeline.
Transfer
A student uses planning, checking, or self-correction on a new task while ARIA is absent.

Sources behind the design

Evidence for the method, not proof of the product.

These sources support how ARIA should be studied. They do not establish that ARIA is effective.

Observable self-regulated learning
Human coding of planning, monitoring, evaluation, and related think-aloud activity.
Human evaluation of AI tutors
Mistake identification, guidance, and actionability evaluated with human labels.
Grounded tutoring dialogue
Teacher-authored scaffolding and documented risks of incorrect feedback or answer revelation.
Metacognition guidance
Planning, monitoring, and evaluation taught explicitly inside subject learning.
What Works Clearinghouse standards
Randomization, attrition, baseline equivalence, eligible outcomes, and study confounds.
Research with children
Institutional review, parental permission, affirmative assent, and risk requirements.

Limitations

What ARIA does not know yet.

  • The 100 task models have not yet been approved by independent educators.
  • The observable-move baseline has not yet been tested against real human labels.
  • The five response conditions have not yet received locked, blinded ratings.
  • No student study has established usability, learning, retention, or transfer.
  • Synthetic benchmark performance does not establish real-student understanding.
  • Personal experience with ADHD motivates the question; it is not clinical evidence.

Evidence base

Built from learning science and rigorous evaluation standards.

The design draws on observable self-regulated-learning coding, human evaluation of AI tutoring, grounded tutoring dialogue, metacognitive transfer research, and What Works Clearinghouse standards. Those sources justify the study design, not ARIA’s effectiveness.